Cpf1-related methods and compositions for gene editing

JP2024023294A5Pending Publication Date: 2025-06-24EDITAS MEDICINE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023193282
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-10-16
Filing Date
2023-11-13
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Current CRISPR/Cas9 systems face limitations in efficiently targeting and editing specific nucleic acid sequences, particularly in eukaryotic cells, and there is a need for improved methods to modulate gene expression and evaluate editing efficiency in therapeutically relevant cell populations.

Method used

The use of CRISPR/Cpf1 systems, including modified Cpf1 proteins with nuclear localization signals and reduced cysteine residues, to target and edit specific nucleic acid sequences in cells such as CD8+ T cells, CD4+ T cells, and hematopoietic stem cells, with strategies for evaluating editing efficiency and modulating gene expression, particularly in the HBG and BCL11a genes to increase fetal hemoglobin expression.

Benefits of technology

The CRISPR/Cpf1 systems demonstrate enhanced editing efficiency and increased fetal hemoglobin expression in cells, potentially alleviating conditions like sickle cell disease and beta-thalassemia, and improve T cell function and persistence for cancer therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000675_0000
    Figure 00000675_0000
  • Figure 00000675_0001
    Figure 00000675_0001
  • Figure 00000675_0002
    Figure 00000675_0002
Patent Text Reader

Abstract

To provide CRISPR / Cpf1-related methods and components for editing a target nucleic acid sequence and / or modulating expression of a target nucleic acid sequence, and to provide methods and compositions for evaluating such editing and / or modulation of expression.SOLUTION: In one aspect, the present disclosure relates to the use of CRISPR / Cpf1-mediated editing of therapeutically-relevant target sites in therapeutically-relevant cell populations. For example, the present disclosure provides isolated cells that include a modification of a therapeutically-relevant target site. In certain embodiments, the cell is, e.g., a CD8+ T cell or the like. In certain embodiments, the present disclosure provides an isolated cell or a population of cells that include a modification, e.g., disruption, in an HBG locus, e.g., generated by the delivery of an RNP complex comprising a Cpf1 RNA-guided nuclease and a gRNA molecule that targets the HBG locus including, for example, the regulatory region of an HBG gene.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 597,118, filed December 11, 2017, U.S. Provisional Patent Application No. 62 / 623,501, filed January 29, 2018, U.S. Provisional Patent Application No. 62 / 664,905, filed April 30, 2018, and U.S. Provisional Patent Application No. 62 / 746,494, filed October 16, 2018, the contents of each of which are claimed and incorporated herein by reference in their entirety.

[0002] Array List

[0001] This specification further incorporates by reference the Sequence Listing submitted via EFS on December 11, 2018. In accordance with 37 C.FR § 1.52(e)(5), the Sequence Listing text file identified as 0841770210SL.txt is 444,032 bytes and was created on December 11, 2018. The entire contents of the Sequence Listing are incorporated herein by reference. The Sequence Listing does not extend beyond the scope of the present specification and therefore does not contain any new material.

[0003] The present disclosure relates to CRISPR / Cpf1-related methods and components for editing target nucleic acid sequences and / or modulating expression of target nucleic acid sequences, as well as methods and compositions for assessing such editing and / or modulation of expression. [Background technology]

[0004] CRISPR (clustered regularly interspaced short palindromic repeats) evolved in bacteria and archaea as an adaptive immune system to defend against viral attack. Upon exposure to a virus, a short segment of viral DNA is integrated into the CRISPR locus. RNA is transcribed from the portion of the CRISPR locus that contains the viral sequence. The RNA, which contains a sequence complementary to the viral genome, mediates targeting of the Cpf1 protein to the target sequence in the viral genome. The Cpf1 protein, also known as Cas12a ("CRISPR1 from Prevotella and Francisella"), then cleaves, thereby silencing the viral target.

[0005] Recently, the CRISPR / Cpf1 system has been applied for genome editing in eukaryotic cells. Introduction of site-specific double-strand breaks (DSBs) allows targeted sequence modification through endogenous DNA repair mechanisms, such as non-homologous end joining (NHEJ) or homology-directed repair (HDR). Summary of the Invention

[0006] The present disclosure provides improvements to CRISPR / Cpfl-related methods and components for editing target nucleic acid sequences and / or modulating expression of target nucleic acid sequences, e.g., in therapeutically relevant cell lines and with respect to therapeutically relevant target sequences, as well as strategies for assessing the efficiency of such target editing and / or modulation of expression.

[0007] In one aspect, the present disclosure relates to the use of CRISPR / Cpf1-mediated editing of a therapeutically relevant target site in a therapeutically relevant cell population. For example, but not limited to, the present disclosure provides isolated cells comprising a therapeutically relevant target site modification. In certain embodiments, the cells are, for example, CD8 + T cells, CD8 + Naive T cells, CD4 + central memory T cells, CD8 + central memory T cells, CD4 + Effector memory T cells, CD4+ Effector memory T cells, CD4 + T cells, CD4 + Stem cell memory T cell, CD8 + Stem cell memory T cell, CD4 + Helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, CD4+ naive T cells, TH17CD4 + T cells, TH1CD4 + T cells, TH2CD4 + T cells, TH9CD4 + T cells, CD4 + Foxp3 + T cells, CD4 + CD25 + CD127 - T cells or CD4 + CD25 + CD127 - Foxp3 + The cell is a T cell, such as a T cell. In certain embodiments, the cell is a lymphoid progenitor cell, a hematopoietic stem cell (HSC), a human umbilical cord blood-derived erythroid progenitor (HUDEP) cell, a natural killer cell, or a dendritic cell. In certain embodiments, the cell is an HSC or a HUDEP cell.

[0008] In certain embodiments, the present disclosure provides isolated cells or cell populations comprising a modification, e.g., a disruption of the HBG locus, generated, for example, by delivery of an RNP complex comprising a Cpf1 RNA-guided nuclease and a gRNA molecule targeting the HBG locus, e.g., including a regulatory region of the HBG gene. In certain embodiments, the RNP complex comprises a complex between a Cpf1 RNA-guided nuclease and a gRNA molecule. In certain embodiments, any region of the HBG locus can be targeted. In certain embodiments, a cis-regulatory region of the HBG gene is targeted. In certain embodiments, the present disclosure relates to the use of CRISPR / Cpf1-mediated editing, e.g., disruption of the promoter region of the HBG locus. In certain embodiments, the present disclosure relates to the use of CRISPR / Cpf1-mediated editing of the −800 to −60 nt promoter region of the HBG locus, e.g., the −110 nt promoter region. In certain embodiments, a cis-regulatory region of the HBG locus can be edited, e.g., disrupted. For example, but not by way of limitation, CRISPR / Cpf1-mediated editing can be used to disrupt the CAAT box present in the cis-regulatory region of the HBG locus. Generally, disruption of the HBG promoter region and disruption of the CAAT box can be achieved through delivery of a CRISPR / Cpf1 editing system that targets these sequences. Non-limiting examples of gRNA molecules for use with such CRISPR / Cpf1 editing systems that target these sequences of the HBG locus are identified in Figures 6, 9, and 11 and Table 19. In certain embodiments, the gRNA molecule that targets the HBG gene sequence comprises the sequence of a gRNA molecule designated HBG1-1.

[0009] In certain embodiments, the present disclosure relates to isolated CRISPR / Cpf1-edited cells in which the -110 nt promoter region of the HBG locus has been disrupted using a complex comprising a CRISPR / Cpf1 RNA-guided nuclease and a guide RNA that targets the -110 nt promoter region of the HBG locus. In certain embodiments, such CRISPR / Cpf1-edited cells may comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, such CRISPR / Cpf1-edited cells do not comprise one or more components of a CRISPR / Cpf1 editing system, as determined using a suitable method used to detect such components. In certain embodiments, the present disclosure relates to a CRISPR / Cpf1-edited cell population in which the -110 nt promoter region of the HBG locus has been disrupted using a complex comprising a CRISPR / Cpf1 RNA-guided nuclease and a guide RNA that targets the -110 nt promoter region of the HBG locus. In certain embodiments, such CRISPR / Cpf1 edited cell populations may include cells comprising one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, such CRISPR / Cpf1 edited cell populations do not comprise one or more components of a CRISPR / Cpf1 editing system, as determined using a suitable method used to detect such components. In certain embodiments, the present disclosure relates to CRISPR / Cpf1 edited cells in which a CAAT box present in the HBG promoter region has been disrupted using a complex comprising a CRISPR / Cpf1 RNA-guided nuclease and a guide RNA that targets the CAAT box present in the promoter region of the HBG locus. In certain embodiments, such cells comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, such CRISPR / Cpf1 edited cells do not comprise one or more components of a CRISPR / Cpf1 editing system, as determined using a suitable method used to detect such components.In certain embodiments, the present disclosure relates to a CRISPR / Cpf1 edited cell population in which a CAAT box present in the HBG promoter region has been disrupted using a complex comprising a CRISPR / Cpf1 RNA-guided nuclease and a guide RNA that targets the CAAT box present in the promoter region of the HBG locus. In certain embodiments, such a CRISPR / Cpf1 edited cell population may include cells comprising one or more components of a CRISPR / Cpf1 editing system.

[0010] In certain embodiments, the present disclosure provides CRISPR / Cpf1-edited cells or cell populations that have been edited using CRISPR / Cpf1, including modifications such as disrupting erythroid cell-specific expression of the transcriptional repressor BCL11a, e.g., generated by delivery of a complex comprising a Cpf1 RNA-guided nuclease and a gRNA molecule targeted to the BCL11a gene sequence. In certain embodiments, any region of the BCL11a gene sequence can be targeted. For example, but not limited to, the erythroid enhancer region of the BCL11a gene can be targeted, e.g., +55 kb to +62 kb from the transcription start site (TSS). In certain embodiments, CRISPR / Cpf1-mediated editing can be used to disrupt the GATA1-binding motif of BCL11a, located at +58 DHS in intron 2 of the BCL11a gene. Disruption of the GATA1 binding motif in BCL11a can be achieved through delivery of a CRISPR / Cpf1 editing system that targets that motif. Non-limiting examples of gRNA molecules for use in such CRISPR / Cpf1 editing systems that target the GATA1 motif in BCL11a are identified in Figures 7, 10, and 12.

[0011] In certain embodiments, the disclosure relates to CRISPR / Cpf1-edited cells in which the +58 DHS region of intron 2 of the BCL11a gene has been disrupted. In certain embodiments, such CRISPR / Cpf1-edited cells may comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, the disclosure relates to CRISPR / Cpf1-edited cell populations in which the +58 DHS region of intron 2 of the BCL11a gene has been disrupted. In certain embodiments, such CRISPR / Cpf1-edited cell populations may comprise cells comprising one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, the disclosure relates to CRISPR / Cpf1-edited cells in which the GATA1 motif of the BCL11a gene has been disrupted. In certain embodiments, such CRISPR / Cpf1-edited cells may comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, the present disclosure relates to a CRISPR / Cpf1-edited cell population in which the GATA1 motif in the BCL11a gene has been disrupted. In certain embodiments, such a CRISPR / Cpf1-edited cell population may include cells that contain one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, one or more components of the CRISPR / Cpf1 system used to modify or disrupt the BCL11a gene in a cell or cell population are undetectable using appropriate means used to detect such components.

[0012] In certain embodiments, the present disclosure provides isolated CRISPR / Cpf1-edited T cells or CRISPR / Cpf1-edited T cell populations comprising a modification, e.g., a disruption, of one or more endogenous genes of a T cell. In certain embodiments, the disclosure relates to the use of CRISPR / Cpf1-mediated editing of an endogenous gene of a T cell selected from the group consisting of FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, TRBC, and any combination thereof. For example, but not limited to, the modification is generated by delivery of one or more complexes comprising a Cpf1 RNA-guided nuclease, such as an RNP complex, and a gRNA molecule, targeting, e.g., a portion of the FAS gene sequence, a portion of the BID gene sequence, a portion of the CTLA4 gene sequence, a portion of the PDCD1 gene sequence, a portion of the CBLB gene sequence, a portion of the PTPN6 gene sequence, a portion of the B2M gene sequence, a portion of the TRAC gene sequence, a portion of the CIITA gene sequence, a portion of the TRBC gene sequence, or a combination thereof. For example, but not limited to, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten complexes, e.g., RNP complexes, can be delivered, each complex targeting a different gene. In certain embodiments, the gRNA can be complementary to either strand of the targeted gene. In certain embodiments, the gRNA molecule can target a regulatory region, intron, or exon of the targeted gene.

[0013] In certain embodiments, a CRISPR / Cpf1 system encompassed by the disclosure herein targets the TRAC gene, resulting in, e.g., isolated CRISPR / Cpf1-edited T cells or CRISPR / Cpf1-edited T cell populations that contain a modification, e.g., a disruption, of the TRAC gene. In certain embodiments, the CRISPR system includes a gRNA that is complementary to a portion of the TRAC gene sequence. In certain embodiments, the gRNA may be complementary to either strand of the TRAC gene. In certain embodiments, the targeting portion of the TRAC gene sequence is within the coding sequence of the TRAC gene. In certain embodiments, the targeting portion of the TRAC gene sequence is within an exon. In certain embodiments, the targeting portion of the TRAC gene sequence is within an intron. In certain embodiments, the targeting portion of the TRAC gene sequence is within a regulatory region of the gene. In certain embodiments, more than one sequence is targeted, and the targeting portion of the TRAC gene sequence is within one or more exons, one or more introns, one or more regulatory regions, or one or more exons, one or more introns, and one or more regulatory regions. In certain embodiments, gRNA molecule targeting domains for use in such CRISPR / Cpf1 systems targeting TRAC comprise the targeting domain sequences listed in Tables 2 and 3.

[0014] In certain embodiments, the CRISPR / Cpf1 system encompassed by the present disclosure targets the TRBC gene, generating, for example, isolated CRISPR / Cpf1-edited T cells or CRISPR / Cpf1-edited T cell populations containing modifications, such as, for example, disruptions, of the TRBC gene. In certain embodiments, the CRISPR system includes a gRNA complementary to a portion of the TRBC gene sequence. In certain embodiments, the gRNA may be complementary to either strand of the TRBC gene. In certain embodiments, the targeting portion of the TRBC gene sequence is within the coding sequence of the TRBC gene. In certain embodiments, the targeting portion of the TRBC gene sequence is within an exon. In certain embodiments, the targeting portion of the TRBC gene sequence is within an intron. In certain embodiments, the targeting portion of the TRBC gene sequence is within a regulatory region of the gene. In certain embodiments, more than one sequence is targeted, and the targeting portion of the TRBC gene sequence is within one or more exons, one or more introns, one or more regulatory regions, or one or more exons, one or more introns, and one or more regulatory regions. In certain embodiments, gRNA molecule targeting domains for use in such CRISPR / Cpf1 systems targeting TRBCs comprise the targeting domain sequences listed in Tables 4 and 5.

[0015] In certain embodiments, a CRISPR / Cpf1 system encompassed by the disclosure herein targets the B2M gene, resulting in, e.g., isolated CRISPR / Cpf1-edited T cells or CRISPR / Cpf1-edited T cell populations that contain a modification, e.g., a disruption, of the B2M gene. In certain embodiments, the CRISPR system includes a gRNA that is complementary to a portion of the B2M gene sequence. In certain embodiments, the gRNA can be complementary to either strand of the B2M gene. In certain embodiments, the targeting portion of the B2M gene sequence is within the coding sequence of the B2M gene. In certain embodiments, the targeting portion of the B2M gene sequence is within an exon. In certain embodiments, the targeting portion of the B2M gene sequence is within an intron. In certain embodiments, the targeting portion of the B2M gene sequence is within a regulatory region of the gene. In certain embodiments, more than one sequence is targeted, and the targeted portion of the B2M gene sequence is within one or more exons, one or more introns, one or more regulatory regions, or one or more exons, one or more introns and one or more regulatory regions. In certain embodiments, the targeting domain of a gRNA molecule for use in such a CRISPR / Cpf1 system that targets B2M comprises a targeting domain sequence listed in Tables 6, 7, and 8. In certain embodiments, the targeting domain of a gRNA molecule for use in such a CRISPR / Cpf1 system that targets B2M comprises the nucleic acid sequence AGUGGGGGUGAAUUCAGUGU.

[0016] In certain embodiments, a CRISPR / Cpf1 system encompassed by the present disclosure targets the CIITA gene, resulting in, e.g., isolated CRISPR / Cpf1-edited T cells or CRISPR / Cpf1-edited T cell populations that contain a modification, e.g., a disruption, of the CIITA gene. In certain embodiments, the CRISPR system includes a gRNA that is complementary to a portion of the CIITA gene sequence. In certain embodiments, the gRNA may be complementary to either strand of the CIITA gene. In certain embodiments, the targeting portion of the CIITA gene sequence is within the coding sequence of the CIITA gene. In certain embodiments, the targeting portion of the CIITA gene sequence is within an exon. In certain embodiments, the targeting portion of the CIITA gene sequence is within an intron. In certain embodiments, the targeting portion of the CIITA gene sequence is within a regulatory region of the gene. In certain embodiments, two or more sequences are targeted, and the targeted portions of the CIITA gene sequence are within one or more exons, one or more introns, one or more regulatory regions, or one or more exons, one or more introns and one or more regulatory regions. In certain embodiments, the gRNA molecule targeting domain for use in such a CRISPR / Cpf1 system targeting CIITA comprises a targeting domain sequence listed in Table 9.

[0017] In certain embodiments, the CRISPR / Cpf1 systems encompassed by the disclosure herein target a combination of two or more of the TRAC, CIITA, TRBC, and B2M genes using gRNAs that target one or more exons, one or more introns, or one or more regulatory regions of two or more of these genes, resulting in isolated CRISPR / Cpf1-edited T cells or CRISPR / Cpf1-edited T cell populations that contain modifications, e.g., disruptions, of two or more of the TRAC, CIITA, TRBC, and B2M genes. In certain embodiments, the CRISPR / Cpf1 systems of the present disclosure may include one or more complexes comprising a Cpf1 RNA-guided nuclease and a gRNA molecule that targets, for example, one or more of the genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC. For example, but not limited to, the CRISPR / Cpf1 system of the present disclosure may include: (a) a first RNP complex comprising a first gRNA molecule having a first targeting domain complementary to a target sequence of a first gene and a first Cpf1 RNA-guided nuclease; and (b) a second RNP complex comprising a second gRNA molecule having a second targeting domain complementary to a target sequence of a second gene and a second Cpf1 RNA-guided nuclease. In certain embodiments, the first gene and the second gene are selected from the group consisting of B2M, TRAC, CIITA, and TRBC. The CRISPR / Cpf1 system may further include additional RNP complexes targeting one or more additional genes. For example, but not limited to, in the case of multiplexing, each RNP complex may contain the same Cpf1 protein, or each RNP complex may contain a different Cpf1 protein, such as, for example, a Cpf1 protein variant.

[0018] In certain embodiments, an isolated cell, such as, for example, an isolated CRISPR / Cpf1-edited HSC or CRISPR / Cpf1-edited T cell, or a population of such CRISPR / Cpf1-edited cells, does not comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, less than about 10%, less than about 5%, or less than about 1% of the CRISPR / Cpf1-edited cells in the cell population comprise one or more components of a CRISPR / Cpf1 editing system, as determined using a suitable means for detecting such components. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in a cell population are edited and / or modified, such as having a disruption of the BCL11a gene, a disruption of the HBG locus, and / or a disruption of one or more genes selected from FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC. In certain embodiments, the cell population has more than about 15% editing, more than about 20% editing, more than about 25% editing, more than about 30% editing, more than about 35% editing, more than about 40% editing, more than about 45% editing, more than about 50% editing, more than about 55% editing, or more than about 60% editing. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in the cell population have productive indels.

[0019] In another aspect, the present disclosure relates to modified Cpf1 proteins and their use in CRISPR / Cpf1-related methods for editing target nucleic acid sequences and / or modulating expression of target nucleic acid sequences. The present disclosure further provides nucleic acids encoding the modified Cpf1 proteins.

[0020] In certain embodiments, the modified Cpf1 protein is selected from the group consisting of Acidaminococcus sp. strain BV3L6 Cpf1 protein (AsCpf1), Francisella novicida U112 (FnCpf1), Moraxella bovoculi 237 (MbCpf1), Candidatus Methanomethylphilus alvus Mx1201 (CMaCpf1), Sneathia amnii (SaCpfq), Moraxella lacunata (MlCpf1), Moraxella bovoculi 237 (MbCpf1), and the like. bovoculi AAX08_00205 (Mb2Cpf1), Moraxella bovoculi AAX11_00205 (Mb3Cpf1), Lachnospiraceae bacterium ND2006 Cpf1 protein (LbCpf1), Lachnospiraceae bacterium MA2020 (Lb5Cpf1), Lachnospiraceae bacterium MC2017 (Lb4Cpf1), Flavobacterium branchiophilum branchiophilum FL-15 (FbCpf1), Thiomicrospira species XS5 (TsCpf1), Parcubacteria group GW2011 (PgCpf1), Candidatus Roizmanbacteria GW2011 (CRbCpf1), Candidatus Peregrinbacteria GW2011 (CPbCpf1), Butyrivibrio species NC3005 (BsCpf1), Butyrivibrio fibrisolvens (BfCpf1), Prevotella bryantiibryantii B14 (Pb2Cpf1) and Bacteroidetes oral taxon 274 (BoCpf1) (see, e.g., Zetsche et al., bioRxiv 134015; doi: https: / / doi.org / 10.1101 / 134015, the entire contents of which are incorporated herein by reference).

[0021] In certain embodiments, the modified Cpf1 protein comprises a nuclear localization signal (NLS). For example, but not limited to, such an NLS sequence is selected from the group consisting of a nucleoplasmin NLS (nNLS) (SEQ ID NO: 1) and a simian virus 40 "SV40" NLS (sNLS) (SEQ ID NO: 2).

[0022] In certain embodiments, the NLS sequence of the modified Cpf1 protein is located at or near the C-terminus of the Cpf1 protein sequence. For example, without limitation, the modified Cpf1 protein can be selected from His-AsCpf1-nNLS (SEQ ID NO: 3); His-AsCpf1-sNstaneyLS (SEQ ID NO: 4); and His-AsCpf1-sNLS-sNLS (SEQ ID NO: 5). In certain embodiments, the NLS sequence of the modified Cpf1 protein is located at or near the N-terminus of the Cpf1 protein sequence. For example, without limitation, the modified Cpf1 protein can be selected from His-sNLS-AsCpf1 (SEQ ID NO: 6), His-sNLS-sNLS-AsCpf1 (SEQ ID NO: 7), and sNLS-sNLS-AsCpf1 (SEQ ID NO: 8). In certain embodiments, the modified Cpf1 protein comprises NLS sequences located at or near both the N-terminus and C-terminus of the Cpf1 protein sequence. For example, but not by way of limitation, the modified Cpf1 protein can be selected from His-sNLS-AsCpf1-sNLS (SEQ ID NO: 9) and His-sNLS-sNLS-AsCpf1-sNLS-sNLS (SEQ ID NO: 10). Additional permutations of NLS sequence identity and N-terminal / C-terminal position are within the scope of the presently disclosed subject matter, such as combinations of two or more nNLS sequences or nNLS and sNLS sequences (or other NLS sequences), as well as adding sequences with or without purification sequences, such as a 6-histidine sequence.

[0023] In certain embodiments, the modified Cpfl protein comprises an alteration (e.g., a deletion or substitution) at one or more cysteine ​​residues in the Cpfl protein sequence. For example, without limitation, the modified Cpfl protein comprises an alteration at a position selected from the group consisting of C65, C205, C334, C379, C608, C674, C1025, and C1248. In certain embodiments, the modified Cpfl protein comprises a substitution of one or more cysteine ​​residues with serine or alanine. In certain embodiments, the modified Cpfl protein comprises an alteration selected from the group consisting of C65S, C205S, C334S, C379S, C608S, C674S, C1025S, and C1248S. In certain embodiments, the modified Cpfl protein comprises modifications selected from the group consisting of C65A, C205A, C334A, C379A, C608A, C674A, C1025A, and C1248A. In certain embodiments, the modified Cpfl protein comprises modifications at positions C334 and C674 or C334, C379, and C674. In certain embodiments, the modified Cpfl protein comprises modifications at C334S and C674S or C334S, C379S, and C674S. In certain embodiments, the modified Cpfl protein comprises modifications at C334A and C674A or C334A, C379A, and C674A. In certain embodiments, modified Cpf1 proteins include both one or more cysteine ​​residue modifications and the introduction of one or more NLS sequences, such as, for example, His-AsCpf1-nNLS Cys-less (SEQ ID NO: 11) or His-AsCpf1-nNLS Cys-low (SEQ ID NO: 12). In certain embodiments, Cpf1 proteins that include deletion or substitution of one or more cysteine ​​residues exhibit reduced aggregation.

[0024] In a further aspect, the present disclosure provides methods for modifying one or more target sequences in a cell. In certain embodiments, such methods comprise contacting a cell or cell population with (a) a gRNA molecule complementary to the target sequence of interest; and (b) a Cpf1 RNA-guided nuclease. In certain embodiments, the Cpf1 RNA-guided nuclease modifies the target sequence of interest in the cell or cell population. In certain embodiments, the cells are T cells, hematopoietic stem cells (HSCs), or human umbilical cord blood-derived erythroid progenitor (HUDEP) cells. In certain embodiments, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in the cell population are modified. In certain embodiments, the target sequence of interest is an HBG1 gene sequence, e.g., a promoter region, and the gRNA molecule comprises the sequence of the gRNA molecule HBG1-1. In certain embodiments, the target sequence of interest is the BCL11a gene sequence. Alternatively, the target nucleic acid sequence is selected from the group consisting of a portion of the FAS gene sequence, a portion of the BID gene sequence, a portion of the CTLA4 gene sequence, a portion of the PDCD1 gene sequence, a portion of the CBLB gene sequence, a portion of the PTPN6 gene sequence, a portion of the B2M gene sequence, a portion of the TRAC gene sequence, a portion of the CIITA gene sequence, a portion of the TRBC gene sequence, and combinations thereof.

[0025] The present disclosure further provides a method for modifying one or more genes, such as two or more, three or more, or four or more, in a cell, comprising, for example, contacting the cell with (a) a first RNP complex comprising a first gRNA that includes a first targeting domain that is complementary to a target sequence of a first gene and a first Cpf1 RNA-guided nuclease; and (b) a second RNP complex comprising a second gRNA molecule that includes a second targeting domain that is complementary to a target sequence of a second gene and a second Cpf1 RNA-guided nuclease. In certain embodiments, the method may further include (c) a third RNP complex comprising a third gRNA molecule comprising a third targeting domain complementary to a target sequence of a third gene and a third Cpf1 RNA-guided nuclease, and / or (d) a fourth RNP complex comprising a fourth gRNA molecule comprising a fourth targeting domain complementary to a target sequence of a fourth gene and a fourth Cpf1 RNA-guided nuclease. In certain embodiments, each RNP complex may comprise the same Cpf1 protein, or each RNP complex may comprise a different Cpf1 protein, e.g., a Cpf1 protein variant. In certain embodiments, a method for modifying one or more genes, e.g., two or more, three or more, or four or more, in a cell may include contacting the cell with (a) a first gRNA molecule comprising a first targeting domain complementary to a target sequence of a first gene; (b) a second gRNA molecule comprising a second targeting domain complementary to a target sequence of a second gene; and (c) a Cpf1 RNA-guided nuclease disclosed herein or a Cpf1 RNA-guided nuclease encoded by a nucleic acid encoding a disclosed Cpf1 RNA-guided nuclease. In certain embodiments, the method may further include (d) a third gRNA molecule comprising a third targeting domain complementary to a target sequence of a third gene, and / or (e) a fourth gRNA molecule comprising a fourth targeting domain complementary to a target sequence of a fourth gene. The Cpf1 RNA-guided nuclease modifies the first gene, the second gene, the third gene, and / or the fourth gene. In certain embodiments, the first gene, the second gene, the third gene, and the fourth gene are selected from the group consisting of B2M, TRAC, CIITA, and TRBC genes. In certain embodiments, the cell is a T cell.

[0026] In another aspect, the present disclosure relates to a method of treating a subject by administering to the subject one or more cells modified using a CRISPR / Cpf1 system encompassed by the present disclosure. In certain embodiments, the one or more cells are modified ex vivo or in vitro and then administered to the subject. In certain embodiments, the method for treating the subject comprises contacting cells obtained from the subject with a CRISPR / Cpf1 system comprising: (a) a gRNA molecule complementary to a target sequence of a target nucleic acid; and (b) a Cpf1 RNA-guided nuclease disclosed herein. In certain embodiments, the present disclosure relates to a method of treating a subject in need thereof by administering to the subject one or more cells obtained from a donor and genetically modified ex vivo or in vitro using the CRISPR / Cpf1 system of the present disclosure prior to administration to the subject. In certain embodiments, the subject is suffering from a hemoglobinopathy, such as sickle cell disease or beta-thalassemia. In certain embodiments, the subject is suffering from cancer or an autoimmune disorder.

[0027] In certain embodiments, the present disclosure further provides a method of administering a cell population to a subject suffering from a hemoglobinopathy, wherein the cell population comprises a modification in the HBG gene sequence or the BCL11a gene sequence generated by delivery of a complex comprising a Cpf1 RNA-guided nuclease and a gRNA molecule targeting the HBG gene sequence or the BCL11a gene sequence. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in the cell population are modified. In certain embodiments, the cells are hematopoietic stem cells (HSCs) or human umbilical cord blood-derived erythroid progenitor (HUDEP) cells.

[0028] In a further aspect, the present disclosure provides gRNA molecules for targeting a nucleic acid sequence of interest and generating modified cells, e.g., CRISPR / Cpf1-edited cells. In certain embodiments, the gRNA molecule comprises a first targeting domain complementary to a target sequence, where the target sequence is an HBG gene sequence or a BCL11a gene sequence. Non-limiting examples of such gRNAs are provided in Figures 6-12 and 46 and Table 19. In certain embodiments, the present disclosure provides CRISPR / Cpf1 systems comprising gRNA molecules that, when introduced into a cell, form indels at or near the target sequence complementary to the first targeting domain of the gRNA molecule and / or create deletions in the sequence complementary to the first targeting domain of the gRNA within the HBG1 or HBG2 promoter region. In certain embodiments, a CRISPR / Cpf1 system comprising a gRNA molecule of the present disclosure, when introduced into a cell, results in increased expression of fetal hemoglobin. In certain embodiments, a CRISPR / Cpfl system comprising a gRNA molecule of the present disclosure results in increased expression of fetal hemoglobin in an amount suitable to partially or completely alleviate symptoms of a hemoglobinopathy, such as, for example, sickle cell disease or beta-thalassemia. For example, without limitation, fetal hemoglobin expression can be increased by at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95% compared to the level of fetal hemoglobin expression in a cell or cell population without a disruption in the BCL11a gene or HBG locus and / or genes. In certain embodiments, the increase in fetal hemoglobin expression may be greater than about 1 picogram (pg), greater than about 2 pg, greater than about 3 pg, greater than about 4 pg, greater than about 5 pg, greater than about 6 pg, greater than about 7 pg, greater than about 8 pg, greater than about 9 pg, or greater than about 10 pg.

[0029] The present disclosure further provides gRNA molecules comprising a first targeting domain complementary to a target sequence selected from the group consisting of a portion of the B2M gene sequence, a portion of the TRAC gene sequence, a portion of the CIITA gene sequence, a portion of the TRBC gene sequence, and combinations thereof. Non-limiting examples of such gRNAs are provided in Tables 2-9.

[0030] The present disclosure provides compositions comprising the gRNA molecules disclosed herein. In certain embodiments, the gRNA molecules comprise gRNAs disclosed in Tables 2-9 and 19 and Figures 6-12. In certain embodiments, the gRNA targets a chromosomal location (e.g., genomic coordinates) provided in Table 18. In certain embodiments, the composition may further comprise a Cpf1 protein, for example, to generate an RNP complex. In certain embodiments, the present disclosure provides compositions comprising one or more RNP complexes, e.g., a population of RNP complexes, each RNP complex targeting a different gene or gene region. In certain embodiments, the compositions can be used to treat a subject in need thereof, e.g., a subject suffering from cancer, an autoimmune disorder, or a hemoglobinopathy.

[0031] In another aspect, the present disclosure relates to a genome editing system for modifying a target nucleic acid sequence. In certain embodiments, the genome editing system may include a gRNA molecule; and a Cpf1 RNA-guided nuclease as disclosed herein. The present disclosure further provides a multiplex genome editing system for editing two or more genes selected from the group consisting of, for example, B2M, TRAC, CIITA, and TRBC.

[0032] In a further aspect, the present disclosure relates to methods for assessing CRISPR / Cpf1-mediated editing of a target nucleic acid sequence and / or modulation of expression of a target nucleic acid sequence and components for accomplishing the same.

[0033] In certain embodiments, methods for assessing CRISPR / Cpf1-mediated editing of a target nucleic acid sequence and / or modulation of expression of a target nucleic acid sequence include comparing the activity of a test Cpf1 protein with a control Cpf1 protein with respect to the target nucleic acid sequence. In certain embodiments, the test Cpf1 protein comprises one or more modifications compared to a control, such as a wild-type Cpf1 protein. Examples of such modifications include, but are not limited to, incorporation of one or more NLS sequences, incorporation of a 6-histidine purification sequence, and modification of cysteine ​​amino acids in the Cpf1 protein, and combinations thereof.

[0034] In certain embodiments, a method for assessing CRISPR / Cpf1-mediated editing of a target nucleic acid sequence and / or modulation of expression of a target nucleic acid sequence involves comparing the activity of a test Cpf1 protein with a control Cas9 protein with respect to a "match site" target nucleic acid sequence. As used herein, a match site target nucleic acid sequence incorporates both the requirements for editing by Cpf1 and Cas9, such as, for example, the TTTV AsCpf1 wild-type protospacer adjacent motif ("PAM") and the NGG SpCas9 wild-type PAM. As described above, the test Cpf1 protein may contain one or more modifications compared to the wild-type Cpf1 protein. Examples of such modifications include, but are not limited to, the modifications described above to incorporate one or more NLS sequences, the modifications described above to incorporate a 6-histidine purification sequence, and alterations of cysteine ​​amino acids in the Cpf1 protein, as well as combinations thereof.

[0035] In certain embodiments, the present disclosure relates to assays for comparing CRISPR / Cpf1-mediated editing of a target nucleic acid sequence and / or modulation of expression of a target nucleic acid sequence by a test CRISPR / Cpf1 genome editing system with a control RNA-guided nuclease genome editing system. For example, without limitation, the test and control genome editing systems can differ in any one or more of the following aspects: the sequence of the RNA-guided nuclease; the source, e.g., the method of manufacture of a component of the genome editing system; the formulation of one or more components of the genome editing system; and the identity of the cells into which the genome editing system is introduced, e.g., the cell type or method of preparation of the cells. In certain embodiments, the assays described herein enable quality control analysis of test genome editing systems. In certain embodiments, the assays of the present disclosure evaluate CRISPR / Cpf1-mediated editing of a target nucleic acid sequence and / or modulation of expression of a target nucleic acid sequence, wherein the target comprises a match site sequence.

[0036] In certain embodiments, the use of a matched site target nucleic acid allows for assaying and / or evaluating CRISPR / Cpf1-mediated editing versus CRISPR / Cas9-mediated editing of the target nucleic acid sequence (or editing by another CRISPR-based system) and / or modulation of expression of the target nucleic acid sequence.

[0037] In certain embodiments, the use of coincident site target nucleic acids allows for the assay and / or evaluation of CRISPR / Cas9-mediated editing of the target nucleic acid sequence (or editing by another CRISPR-based system) versus CRISPR / Cpfl-mediated editing versus modulation of expression of the target nucleic acid sequence in a particular cell type. For example, but not limited to, such methods can be used to assay and / or evaluate the expression of T cells, hematopoietic stem cells (CD34), and hematopoietic stem cells (CD34), among several other cell types. + CRISPR / Cas9-mediated editing of target nucleic acid sequences and / or modulation of expression of target nucleic acid sequences can be assessed in human umbilical cord blood-derived erythroid progenitor (HUDEP) cells, including but not limited to HSCs.

[0038] In certain embodiments, the use of matched site target nucleic acids allows for assays and / or evaluation of CRISPR / Cpf1-mediated editing versus CRISPR / Cas9-mediated editing of a target nucleic acid sequence (or editing by another CRISPR-based system) and / or modulation of expression of the target nucleic acid sequence with respect to specific attributes of the CRISPR / Cpf1-mediated editing system used. For example, but not by way of limitation, such methods can be used to evaluate CRISPR / Cas9-mediated editing of a target nucleic acid sequence and / or modulation of expression of the target nucleic acid sequence to identify differences in the activity of Cpf1 RNA-guided nucleases and / or gRNAs prepared by different manufacturing processes. Such methods can also identify differences in the activity of Cpf1 RNA-guided nucleases and / or gRNAs present in different formulations and using different delivery strategies.

[0039] In certain embodiments, the match site target nucleic acid sequence is selected from the group consisting of match site 1 ("MS1"; SEQ ID NO: 13), match site 5 ("MS5"; SEQ ID NO: 14), match site 11 ("MS11"; SEQ ID NO: 15), and match site 18 ("MS18"; SEQ ID NO: 16). In certain embodiments, the match site target nucleic acid sequence is MS5.

[0040] The CRISPR / Cpf1 editing system of the present disclosure can be delivered to cells using various strategies. For example, but not limited to, vectors, such as AAV or other viral vectors, encoding the components of the CRISPR / Cpf1 editing system can be used to induce expression of the components of the CRISPR / Cpf1 editing system in cells. Alternatively, RNP complexes containing various components of the CRISPR / Cpf1 editing system can be delivered to cells by, for example, electroporation or any other suitable method that can be used to deliver RNP complexes to cells. In certain embodiments, lipid nanoparticles can be used to deliver RNP complexes to cells.

[0041] The accompanying drawings are intended to provide illustrative, schematic examples of certain aspects and embodiments of the present disclosure, rather than being comprehensive. The drawings are not intended to be limited to or bound by any particular theory or model, and are not necessarily drawn to scale. Without limiting the above, nucleic acids and polypeptides may be depicted as linear sequences or as schematic two- or three-dimensional structures; these depictions are intended to be illustrative, rather than limiting to or bound by any particular model or theory regarding their structure. [Brief explanation of the drawings]

[0042] [Figure 1] We provide an overview of how engineered Cpf1 variants expand the PAM targeting space. [Figure 2] We provide an overview of the four match site sequences (MS1, MS5, MS11, and MS18) from Kleinstiver et al., Nature Biotechnology, 34(8):869-74 Aug. 2016, and the cell types used to evaluate the performance of Cpf1 and Cas9 in relation to these match site target sequences. [Figure 3A] Depicted are the results of a dose-response experiment comparing increasing concentrations of Cpf1 / gRNA RNP with Cas9 / gRNA RNP at two matched site loci (MS1 and MS5) (Figure 3A) and an assay comparing the activity of AsCpf1 and SpCas9 against matched site targets MS1, MS5, MS11, and MS18, with Cpf1 editing specific target sites more efficiently than Cas9 (Figure 3B). [Figure 3B] Depicted are the results of a dose-response experiment comparing increasing concentrations of Cpf1 / gRNA RNP with Cas9 / gRNA RNP at two matched site loci (MS1 and MS5) (Figure 3A) and an assay comparing the activity of AsCpf1 and SpCas9 against matched site targets MS1, MS5, MS11, and MS18, with Cpf1 editing specific target sites more efficiently than Cas9 (Figure 3B). [Figure 4]Figure 1 shows a comparison of various AsCpf1 NLS mutants across multiple cell types at a fixed 4.4 µM RNP dose using the matched site 5 guide. Data is normalized to the mutant exhibiting maximum editing for each cell type. [Figure 5A] Figure 5A shows a comparison of two optimal AsCpf1 NLS variants at a 4.4 μM RNP dose using guide RNA GWED545 targeting the TRAC locus in primary T cells, and Figure 5B shows a comparison of a His-AsCpf1-sNLS-sNLS variant at a 4.4 μM RNP dose using guide RNA B2M-12 targeting the TRAC locus in primary T cells. In both cases, data are normalized to the variant exhibiting the greatest editing. [Figure 5B] Figure 5A shows a comparison of two optimal AsCpf1 NLS variants at a 4.4 μM RNP dose using guide RNA GWED545 targeting the TRAC locus in primary T cells, and Figure 5B shows a comparison of a His-AsCpf1-sNLS-sNLS variant at a 4.4 μM RNP dose using guide RNA B2M-12 targeting the TRAC locus in primary T cells. In both cases, data are normalized to the variant exhibiting the greatest editing. [Figure 6] The gRNA sequences used in the HBG1 assay in HSCs and HUDEPs are shown. [Figure 7] The gRNA sequences used in the BCL11a assay in HSCs and HUDEPs are shown. [Figure 8] The specific sequences of HBG1 or BCL11a in either HSCs or HUDEPs and their corresponding % editing are shown. Proposed gRNAs targeting HBG1 are also provided. [Figure 9] The HBG1 promoter region with gRNA AsCpf1 WT HBG1-1 bound to the CAAT box motif is shown. [Figure 10] A portion of the BCL11a enhancer region where gRNA BCL11a AsCpf1 RR-8 binds to the GATA1 motif is shown. [Figure 11] The region of the HBG1 promoter that was screened using the gRNAs identified in Figure 6 is shown. This region spans approximately 150 bp. HBG1-1 is shown overlapping the CAAT box motif. [Figure 12] Shown is the region of the BCL11a erythroid enhancer that was screened using the gRNAs identified in Figure 7. This region spans approximately 600 base pairs, and BCL11a RR-8 is shown overlapping with the GATA1 motif. [Figure 13] The AsCpf1 low cysteine ​​construct shows the identified cysteine ​​mutants. [Figure 14] 1 shows the results of an AlexaFluor maleimide assay demonstrating significantly reduced accessibility of cysteine ​​residues in AsCpf1 C334S C379S C674S. [Figure 15] 1 depicts the demonstration of equivalent endonuclease activity of WT AsCpf1, AsCpf1 cysteine-less and two low-cysteine ​​mutants on MS5 substrate DNA. [Figure 16] Targeting of the HBG1 promoter region by AsCpf1 WT and RR PAM mutants in HUDEP and HSCs is shown. HUDEP experiments were performed using the optimal CA-137 pulse program and Lonza solution SE. HSC screening was performed using pulse code EO-100 and Lonza solution P3 as recommended by the manufacturer. The dose was 4.4 μM RNP at a 2:1 guide:protein ratio for all guides. 50,000 HSCs were treated per condition. AsCpf1 WT and RR proteins had endotoxin levels of <5 EU / mL. [Figure 17]Screening of the BCL11a enhancer region with AsCpf1 WT and RR and RVR PAM mutants, along with one WT FnCpf1 target, in HUDEP and HSCs is shown. HUDEP screening runs were performed using an optimized CA-137 pulse program with Lonza solution SE. HSC screening was performed using pulse code EO-100 and Lonza solution P3, as recommended by the manufacturer. A control guide for BCL11a (designated KOBEH) is also shown. The dose was 4.4 μM RNP at a 2:1 guide:protein ratio for all guides. 50,000 HSCs were treated per condition. AsCpf1 WT, RR, and RVR proteins had endotoxin levels of <5 EU / mL. [Figure 18] Nucleofection screening of AsCpf1 in HUDEP is shown. The dose was 2.2 μM AsCpf1 RNP using consensus site 5 (MS5) guide RNA at a 2:1 guide:protein ratio. AsCpf1 WT protein had endotoxin levels of <5 EU / mL. Lonza solutions SE, SF, and SG were tested with 50,000 HUDEPs per condition using different pulse programs. Pulse codes CA-137 and CA-138 in solution SE demonstrated optimal editing. [Figure 19] Nucleofection screening of AsCpf1 in HSCs is shown. The dose was 2.2 μM AsCpf1 RNP using consensus site 5 (MS5) guide RNA at a 2:1 guide:protein ratio. AsCpf1 WT protein had endotoxin levels of <5 EU / mL. Lonza solutions P1, P2, P3, P4, and P5 were tested on 50,000 HSCs per condition using different pulse programs. Pulse codes CA-137 and CA-138 with solution P2 demonstrated optimal editing, as did FF-100 and FF-104. [Figure 20]We demonstrate that the use of specific pulse codes in Lonza Amaxa increases editing of HSCs across targets and PAM variants. The dose was 4.4 μM RNP at a 2:1 guide:protein ratio for all guides. 50,000 HSCs were treated per condition. AsCpf1 WT, RR, and RVR proteins had endotoxin levels of <5 EU / mL. [Figure 21] Screening of T cell therapeutic targets with AsCpf1 and its RR and RVR PAM variants at the TRBC, TRAC, and B2M loci is shown. In preliminary screening, approximately 30% of gRNAs showed greater than 50% editing, which is comparable to the commonly observed SpCas9 hit rate, demonstrating that Cpf1 can potentially be used for gene editing of patient T cells at therapeutic loci, including, but not limited to, TRAC, TRBC, and / or B2M. [Figure 22] We show that altering the electroporation pulse code significantly improves maximal editing of T cells at multiple therapeutic target loci. [Figure 23A] Figure 23 shows efficient knockout editing of primary T cells at disease-associated loci by Cpf1 RNP. Figure 23A shows the RNP workflow for ex vivo cell therapy. Figure 23B shows efficient single knockout at multiple therapeutically relevant T cell loci using AsCpf1 or engineered PAM mutants. [Figure 23B] Figure 23 shows efficient knockout editing of primary T cells at disease-associated loci by Cpf1 RNP. Figure 23A shows the RNP workflow for ex vivo cell therapy. Figure 23B shows efficient single knockout at multiple therapeutically relevant T cell loci using AsCpf1 or engineered PAM mutants. [Figure 24] Figure 1 shows highly efficient double knockout of two therapeutic targets in T cells treated with Cpf1 RNP as measured by flow cytometry. [Figure 25]1 shows screening of T cell therapy targets with AsCpf1 and its RR and RVR PAM variants at the TRBC, TRAC, and B2M loci. [Figure 26] 1 summarizes the high editing efficiency of AsCpf1 WT, RR, and RVR in T cells on three allogeneic T cell targets. [Figure 27] 1 illustrates the dual knockout of two T cell targets by Cpf1 or Cas9 in human primary T cells. [Figure 28] 1 shows screening of T cell therapy targets by Cpf1 at the CIITA locus. [Figure 29] We summarize the high editing efficiency of Cpf1 in T cells on three allogeneic T cell targets, TRAC, CIITA, and B2M, compared to SpCas9. [Figure 30] Illustrates the efficiency of triple knockout of three T cell targets by Cpf1 RNP in T cells. [Figure 31A] Figure 31A summarizes the specificity of the top Cpfl candidate guides for three T cell targets, CIITA, TRAC, and B2M, and shows the number of off-targets detected. Figure 31B shows that targeted amplicon sequencing found no detectable off-targets. [Figure 31B] Figure 31A summarizes the specificity of the top Cpfl candidate guides for three T cell targets, CIITA, TRAC, and B2M, and shows the number of off-targets detected. Figure 31B shows that targeted amplicon sequencing found no detectable off-targets. [Figure 32]

[0033] Figure 1 shows the identification of electroporation conditions that improve maximal editing in T cells. Condition 1 was DS-130 and condition 2 was CA-137. [Figure 33]

[0023] Figure 1 shows the identification of NLS configurations that improve the efficacy of gene editing in T cells. NLS v1 represents the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 1) and NLS v2 represents the sequence 2xPKKKRKV (SEQ ID NO: 2). [Figure 34]Shows editing efficiency at the HBG-1 locus in HSCs using AsCpf1 and the HBG1-1 guide. [Figure 35] 1 shows the editing efficiency of NLS variants in T cells at match site 5 using MS5 guide RNA. [Figure 36] 1 shows the reduction of MHC II in T cells edited at the CIITA locus as measured by flow cytometry. [Figure 37A] 1 shows the editing efficiency in T cells edited at the CIITA locus. [Figure 37B] The genomic locations targeted by CIITA gRNAs CIITA-34, CIITA-41, CIITA-45, and CIITA-10 are shown. [Figure 38] Summarizes the percentage reduction of MHC II in T cells edited at the CIITA locus. [Figure 39] The editing efficiency of Cpf1 CIITA gRNA is shown, along with the number of off-targets detected by the gRNA. [Figure 40] The editing efficiencies of AspCpf1 RR and WT TRAC, CIITA, and B2M gRNAs are shown. [Figure 41] The editing efficiencies of AspCpf1 RR and WT B2M gRNAs of different lengths are shown. [Figure 42] The editing efficiencies of AspCpf1 RR and WT TRAC gRNAs of different lengths are shown. [Figure 43] The editing efficiencies of AspCpf1 RR and WT CIITA gRNA of different lengths are shown. [Figure 44A]Schematic diagram of an unedited genomic DNA targeting site, an exemplary DNA donor template for targeted integration, potential insertion outcomes (i.e., non-targeted integration at the cleavage site or targeted integration at the cleavage site), and three potential PCR amplicons resulting from using a primer pair targeting the P1 and P2 primer sites (amplicon X), a primer pair targeting the P1 and P2' priming sites (amplicon Y), or a primer pair targeting the P1' and P2 primer sites (amplicon Z). The depicted exemplary DNA donor template contains integrated primer sites (P1' and P2') and stuffer sequences (S1 and S2). A1 / A2: donor homology arms; S1 / S2: donor stuffer sequences; P1 / P2: genomic primer site; P1' / P2': integrated primer site; H1 / H2: genomic homology arms; N: cargo; X: cleavage site. [Figure 44B] Schematic diagram of an unedited genomic DNA targeting site, an exemplary DNA donor template for targeted integration, potential insertion outcomes (i.e., non-targeted integration at the cleavage site or targeted integration at the cleavage site), and two potential PCR amplicons resulting from using a primer pair targeting the P1 and P2 primer sites (amplicon X) or a primer pair targeting the P1' and P2 primer sites (amplicon Y). The exemplary DNA donor template contains an integrated primer site (P1') and a stuffer sequence (S2). A1 / A2: donor homology arms; S1 / S2: donor stuffer sequences; P1 / P2: genomic primer site; P1': integrated primer site; H1 / H2: genomic homology arms; N: cargo; X: cleavage site. [Figure 44C]Schematic diagram of an unedited genomic DNA targeting site, an exemplary DNA donor template for targeted integration, potential insertion outcomes (i.e., non-targeted integration at the cleavage site or targeted integration at the cleavage site), and two potential PCR amplicons resulting from using a primer pair targeting the P1 and P2 primer sites (amplicon X) or a primer pair targeting the P1 and P2' primer sites (amplicon Y). The exemplary DNA donor template contains an integrated primer site (P2') and a stuffer sequence (S1). A1 / A2: donor homology arms; S1 / S2: donor stuffer sequences; P1 / P2: genomic primer site; P2': integrated primer site; H1 / H2: genomic homology arms; N: cargo; X: cleavage site. [Figure 45] An exemplary DNA donor template designed for gRNA targeting of the T-cell receptor alpha constant (TRAC) locus is shown. [Figure 46] The gRNAs identified from screening of the promoter regions of HBG1 and HBG2 are shown. DETAILED DESCRIPTION OF THE INVENTION

[0043] Definitions and Abbreviations Unless otherwise specified, each of the following terms has the meaning associated with it in this section.

[0044] The indefinite articles "a" and "an" refer to at least one of the associated noun and are used interchangeably with the terms "at least one" and "one or more." For example, "a module" means at least one module or one or more modules.

[0045] The conjunctions "or" and "and / or" are used interchangeably as non-exclusive disjunctions.

[0046] The terms "about" or "approximately," as used herein, can mean within an acceptable error range for a particular value, as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" can mean within one standard deviation or more than one standard deviation according to convention for a given value. When a particular value is described in this application and in the claims, unless otherwise specified, the term "about" can mean an acceptable error range for the particular value, such as, for example, ±10% of the value modified by the term "about."

[0047] The phrase "consisting essentially of" means that the recited chemical species are the predominant species, but that other chemical species may be present in trace amounts or quantities that do not affect the structure, function, or behavior of the subject composition. For example, a composition consisting essentially of a particular chemical species generally contains 90%, 95%, 96%, or more of that chemical species.

[0048] "Domain" is used to describe a segment of a protein or nucleic acid. Unless otherwise specified, a domain need not have any particular functional properties.

[0049] An "indel" is an insertion and / or deletion in a nucleic acid sequence. An indel can be the product of repair of a DNA double-strand break, such as a double-strand break formed by the genome editing system of the present disclosure. Indels are most commonly formed when the break is repaired by an "error-prone" repair pathway, such as the NHEJ pathway described below.

[0050] A "productive indel" with respect to HSCs refers to an indel (deletion and / or insertion) that results in HbF expression. In certain embodiments, a productive indel in HSCs can induce HbF expression. In certain embodiments, a productive indel in HSCs can result in increased levels of HbF expression. A "productive indel" with respect to T cells refers to an indel (deletion and / or insertion) that reduces expression of a target gene in T cells, such as, for example, an endogenous T cell gene. In certain embodiments, a "productive indel" in a T cell results in reduced or eliminated expression of a cell surface protein or marker on the T cell.

[0051] "Gene conversion" refers to the modification of a DNA sequence by the incorporation of an endogenous homologous sequence (e.g., a homologous sequence in a gene array). "Gene correction" refers to the modification of a DNA sequence by the incorporation of an exogenous homologous sequence, such as an exogenous single-stranded or double-stranded donor template DNA. Gene conversion and gene correction are the products of repair of DNA double-strand breaks by HDR pathways, such as those described below.

[0052] Indels, gene conversions, gene corrections, and other genome editing results are typically assessed by sequencing (most commonly by "next-generation" or "sequencing by synthesis" methods, although Sanger sequencing can also be used) and quantified by the relative frequency of numerical changes (e.g., ±1, ±2, or more bases) in all sequencing reads. DNA samples for sequencing can be prepared by a variety of methods known in the art, including amplification of target sites by polymerase chain reaction (PCR), capture of DNA ends generated by double-strand breaks, such as in the GUIDEseq process described by Tsai et al. (Nat. Biotechnol. 34(5):483(2016) (incorporated herein by reference), or by other means known in the art. Genome editing results can also be assessed by in situ hybridization methods, such as the FiberComb™ system commercialized by Genomic Vision (Bagneux, France), and any other suitable method known in the art.

[0053] As used herein, the phrase "modification in a target sequence" and its equivalents include, but are not limited to, deletion, insertion, gene conversion, gene correction, and / or introduction of indels into a target sequence. Modifications in a target sequence can result in altered expression of the target sequence; for example, a modification in a coding sequence can prevent expression of the protein encoded by that sequence, while a modification in a regulatory sequence can result in increased or decreased expression of a protein under the control of that regulatory sequence, depending on whether the regulatory sequence activates or inhibits expression of the protein.

[0054] "Alt-HDR," "alternative homology-directed repair," or "alternative HDR" are used interchangeably to refer to the process of repairing DNA damage using homologous nucleic acids (e.g., endogenous homologous sequences, such as sister chromatids, or exogenous nucleic acids, such as template nucleic acids). Alt-HDR differs from canonical HDR in that the process utilizes a different pathway than canonical HDR and can be inhibited by canonical HDR mediators RAD51 and BRCA2. Alt-HDR is also distinguished by the involvement of single-stranded or nicked homologous nucleic acid templates, whereas canonical HDR generally involves double-stranded homologous templates.

[0055] "Standard HDR," "standard homology-directed repair," or "cHDR" refers to the process of repairing DNA damage using a homologous nucleic acid (e.g., an endogenous homologous sequence such as a sister chromatid, or an exogenous nucleic acid such as a template nucleic acid). Standard HDR typically functions when there is significant excision at the double-stranded break, resulting in the formation of at least one single-stranded portion of DNA. In normal cells, cHDR typically involves a series of steps, including break recognition, break stabilization, excision, stabilization of single-stranded DNA, formation of a DNA crossover intermediate, resolution of the crossover intermediate, and ligation. The process requires RAD51 and BRCA2, and the homologous nucleic acid is typically double-stranded.

[0056] Unless otherwise specified, the term "HDR" as used herein encompasses both standard HDR and alt-HDR.

[0057] "Non-homologous end joining" or "NHEJ" refers to ligation-mediated and / or non-template-mediated repair such as canonical NHEJ (cNHEJ) and alternative NHEJ (altNHEJ), which in turn includes microhomology-mediated end joining (MMEJ), single-strand annealing (SSA), and synthesis-dependent microhomology-mediated end joining (SD-MMEJ).

[0058] "Substitution" or "substituted," when used in reference to the modification of a molecule (e.g., a nucleic acid or protein), does not require a process restriction, but simply indicates that a replacement entity is present.

[0059] "Subject" means a human or non-human animal. A human subject can be of any age (e.g., infant, child, young adult, or adult) and can be suffering from a disease or in need of genetic modification. Alternatively, a subject can be an animal, which term includes, but is not limited to, mammals, birds, fish, reptiles, amphibians, and more particularly non-human primates, rodents (e.g., mice, rats, hamsters), rabbits, guinea pigs, dogs, cats, and the like. In certain embodiments of the present disclosure, the subject is a livestock animal, such as, for example, a cow, horse, sheep, or goat. In certain embodiments, the subject is poultry. In certain embodiments, the subject is a plant.

[0060] "Treat," "treating," and "treatment" refer to the treatment of a disease in a subject (e.g., a human subject), including one or more of: inhibiting the disease, i.e., arresting or preventing its onset or progression; palliating the disease, i.e., causing regression of the disease state; alleviating one or more symptoms of the disease; or curing the disease.

[0061] "Prevent," "preventing," and "prevention" refer to the prevention of disease in a mammal, such as a human, including (a) avoiding or eliminating disease; (b) affecting predisposition to disease; or (c) preventing or delaying the onset of at least one symptom of the disease.

[0062] A "kit" refers to any collection of two or more components that together constitute a functional unit that can be used for a particular purpose. By way of example (and not limitation), a kit according to the present disclosure may include a guide RNA complexed or complexable with an RNA-guided nuclease, along with a pharmaceutically acceptable carrier (e.g., suspended or suspendable therein). For example, the kit may be used to introduce the complex into a cell or subject for the purpose of causing a desired genomic modification in such a cell or subject. The components of the kit may be packaged together, or they may be packaged separately. A kit according to the present disclosure also optionally includes a instructions for use (DFU) that describes, for example, the use of the kit according to the methods of the present disclosure. The DFU may be physically packaged with the kit, or it may be provided to the kit user by, for example, electronic means.

[0063] The terms "polynucleotide," "nucleotide sequence," "nucleic acid," "nucleic acid molecule," "nucleic acid sequence," and "oligonucleotide" refer to a series of nucleotide bases (also referred to as "nucleotides") in DNA and RNA, and refer to any chain of two or more nucleotides. Polynucleotides, nucleotide sequences, nucleic acids, etc., can be single-stranded or double-stranded chimeric mixtures, or derivatives or modified versions thereof. They can be modified, for example, at the base moiety, sugar moiety, or phosphate backbone to improve the molecule's stability, its hybridization parameters, etc. Nucleotide sequences typically carry genetic information, including, but not limited to, information used by cellular machinery to produce proteins and enzymes. These terms include double- or single-stranded genomic DNA, RNA, any synthetic and genetically engineered polynucleotides, and both sense and antisense polynucleotides. These terms also include nucleic acids containing modified bases.

[0064] Conventional IUPAC notation is used in the nucleotide sequences presented herein, as shown in Table 1 below (see also Cornish-Bowden A, Nucleic Acids Res. 1985 May 10;13(9):3021-30, incorporated herein by reference). However, it should be noted that where a sequence can be encoded by either DNA or RNA, for example, in a gRNA targeting domain, "T" indicates "thymine or uracil."

[0065] TIFF2024023294000001.tif113170

[0066] The terms "protein," "peptide," and "polypeptide" are used interchangeably and refer to a continuous chain of amino acids linked together through peptide bonds. The term includes individual proteins, groups or complexes of proteins linked together, as well as fragments or portions, variants, derivatives, and analogs of such proteins. Peptide sequences are presented herein using conventional notation, starting from the amino or N-terminus on the left and proceeding to the carboxyl or C-terminus on the right. Standard one-letter or three-letter abbreviations may be used.

[0067] The term "variant" refers to an entity, such as a polypeptide, polynucleotide, or small molecule, that exhibits significant structural identity with a reference entity (e.g., a wild-type or naturally occurring entity), but that differs structurally from the reference entity in the presence or level of one or more chemical moieties, e.g., amino acids in the context of a polypeptide or nucleotides in the context of a polynucleotide, compared to the reference entity. As used herein, the term variant also encompasses entities, such as polypeptides, polynucleotides, or small molecules, that are functionally better or superior to the reference entity in one or more properties associated with such entity. In many embodiments, variants also differ functionally from their reference entities. For example, without limitation, "variant Cpf1 polypeptide" encompasses an AsCpf1 variant containing S542R / K607R substitutions that recognizes the TYCV PAM, as well as an AsCpf1 variant containing S542R / K548V / N552R substitutions that recognizes the TATV PAM.

[0068] As used herein, the term "cleavage event" refers to the cleavage of a nucleic acid molecule. A cleavage event can be a single-strand cleavage event or a double-strand cleavage event. A single-strand cleavage event can result in a 5' overhang or a 3' overhang. A double-strand cleavage event can result in a blunt end, two 5' overhangs, or two 3' overhangs.

[0069] The term "cleavage site," as used herein with reference to a site on a target nucleic acid sequence, refers to a target position between two nucleotide residues of a target nucleic acid at which a double-stranded break occurs, or alternatively, a target position within a span of several nucleotide residues of a target nucleic acid at which two single-stranded breaks mediated by an RNA-guided nuclease-dependent process occur. The cleavage site can be, for example, a target position for a blunt-type double-stranded break. Alternatively, the cleavage site can be, for example, a target position within a span of several nucleotide residues of a target nucleic acid for two single-stranded breaks or nicks, separated by, for example, about 10 base pairs, forming a double-stranded break. The closer of the double-stranded break or pair of single-stranded nicks is ideally within 0-500 bp of the target position (e.g., 450, 400, 350, 300, 250, 200, 150, 100, 50, or 25 bp or less from the target position). When dual nickases are used, the two nicks in a pair are within 25-55 bp of each other (e.g., 25-50, 25-45, 25-40, 25-35, 25-30, 50-55, 45-55, 40-55, 35-55, 30-55, 30-50, 35-50, 40-50, 45-50, 35-45, or 40-45 bp) and are not more than 100 bp apart from each other (e.g., 90, 80, 70, 60, 50, 40, 30, 20, or 10 bp or less).

[0070] The present disclosure provides CRISPR / Cpf1-related methods and components for editing a target nucleic acid sequence and / or modulating the expression of a target nucleic acid sequence. For example, the present disclosure provides CRISPR / Cpf1-related methods for targeting nucleic acid sequences that affect the proliferation, survival, persistence, and / or function of hematopoietic stem cells (HSCs). In certain non-limiting embodiments, the present disclosure provides methods for targeting CD34 by Cpf1 RNA-guided nuclease. +This disclosure provides the first evidence of efficient editing of a target nucleic acid sequence in a cell. Furthermore, this disclosure provides the first evidence of efficient editing of BCL11a and HBG1, genes associated with the hereditary persistence of hemoglobin-like fetal heart failure (referred to herein as "HPFH"), by Cpf1 RNA-guided nuclease. This disclosure also provides CRISPR / Cpf1-related methods for targeting nucleic acid sequences that affect T cell proliferation, survival, persistence, and / or function. This disclosure further provides modified Cpf1 proteins that exhibit significant editing efficiency and exhibit improved properties, and strategies for evaluating the efficiency of such modified Cpf1 proteins.

[0071] Modified Cpf1 protein In one aspect, the present disclosure relates to modified Cpf1 proteins and their use in CRISPR / Cpf1-related methods for editing target nucleic acid sequences and / or modulating expression of target nucleic acid sequences.

[0072] In certain embodiments, the modified Cpf1 protein is selected from the group consisting of Acidaminococcus sp. strain BV3L6 Cpf1 protein (AsCpf1), Francisella novicida U112 (FnCpf1), Moraxella bovoculi 237 (MbCpf1), Candidatus Methanomethylphilus alvus Mx1201 (CMaCpf1), Sneathia amnii (SaCpfq), Moraxella lacunata (MlCpf1), Moraxella bovoculi 237 (MbCpf1), and the like. bovoculi AAX08_00205 (Mb2Cpf1), Moraxella bovoculi AAX11_00205 (Mb3Cpf1), Lachnospiraceae bacterium ND2006 Cpf1 protein (LbCpf1), Lachnospiraceae bacterium MA2020 (Lb5Cpf1), Lachnospiraceae bacterium MC2017 (Lb4Cpf1), Flavobacterium branchiophilum branchiophilum FL-15 (FbCpf1), Thiomicrospira species XS5 (TsCpf1), Parcubacteria group GW2011 (PgCpf1), Candidatus Roizmanbacteria GW2011 (CRbCpf1), Candidatus Peregrinbacteria GW2011 (CPbCpf1), Butyrivibrio species NC3005 (BsCpf1), Butyrivibrio fibrisolvens (BfCpf1), Prevotella bryantiibryantii B14 (Pb2Cpf1) and Bacteroidetes oral taxon 274 (BoCpf1) (see, e.g., Zetsche et al., bioRxiv 134015; doi: https: / / doi.org / 10.1101 / 134015, the entire contents of which are incorporated herein by reference). In certain embodiments, the Cpf1 protein comprises a sequence selected from the group consisting of SEQ ID NOs: 17-19, which have the codon-optimized nucleic acid sequences of SEQ ID NOs: 20-22, respectively.

[0073] Cpf1 nuclear localization signal (NLS) mutant In certain embodiments, the modified Cpf1 protein comprises a nuclear localization signal (NLS) (also referred to herein as a "Cpf1 NLS variant"). For example, without limitation, an NLS sequence useful in connection with the methods and compositions disclosed herein would comprise an amino acid sequence capable of promoting protein import into a cell nucleus. NLS sequences useful in connection with the methods and compositions disclosed herein are known in the art. Non-limiting examples of such NLS sequences include the nucleoplasmin NLS, having the amino acid sequence: KRPAATKKAGQAKKKK (SEQ ID NO: 1), and the simian virus 40 "SV40" NLS, having the amino acid sequence PKKKRKV (SEQ ID NO: 2).

[0074] In certain embodiments, the modified Cpf1 protein may have one or more NLS sequences, such as, for example, two or more, three or more, or four or more. For example, without limitation, the modified Cpf1 protein may have two NLS sequences, three NLS sequences, or four NLS sequences. In certain embodiments, the modified Cpf1 protein may have two NLS sequences. In certain embodiments, the NLS sequence of the modified Cpf1 protein is located at or near the C-terminus of the Cpf1 protein sequence. In certain embodiments, the NLS sequence of the modified Cpf1 protein is located at or near the N-terminus of the Cpf1 protein sequence. In certain embodiments, the modified Cpf1 protein of the present disclosure may have one or more NLS sequences located at or near the N-terminus of the Cpf1 protein sequence and one or more NLS sequences located at or near the C-terminus of the Cpf1 protein sequence; for example, the modified Cpf1 protein includes NLS sequences located at or near both the N-terminus and the C-terminus of the Cpf1 protein sequence.

[0075] In certain embodiments, modified Cpf1 proteins having an NLS sequence located at or near the C-terminus of the Cpf1 protein sequence may be selected from His-AsCpf1-nNLS (also referred to herein as "Asp Cpf1 NLS v1") (SEQ ID NO: 3); His-AsCpf1-sNLS (SEQ ID NO: 4); and His-AsCpf1-sNLS-sNLS (also referred to herein as "Asp Cpf1 NLS v2") (SEQ ID NO: 5) (where "His" refers to the 6-histidine purification sequence, "AsCpf1" refers to the Acidaminococcus sp. Cpf1 protein sequence, "nNLS" refers to the nucleoplasmin NLS, and "sNLS" refers to the SV40 NLS). Additional permutations of the identity and C-terminal position of the NLS sequence are within the scope of the presently disclosed subject matter, such as the addition of two or more nNLS sequences or the addition of a combination of nNLS and sNLS sequences (or other NLS sequences), as well as the addition of sequences with or without refining sequences, such as a 6-histidine sequence.

[0076] In certain embodiments, modified Cpf1 proteins having an NLS sequence located at or near the N-terminus of the Cpf1 protein sequence can be selected from His-sNLS-AsCpf1 (SEQ ID NO: 6), His-sNLS-sNLS-AsCpf1 (SEQ ID NO: 7), and sNLS-sNLS-AsCpf1 (SEQ ID NO: 8). Additional permutations of NLS sequence identity and N-terminal position are within the scope of the presently disclosed subject matter, such as the addition of two or more nNLS sequences or the addition of a combination of nNLS and sNLS sequences (or other NLS sequences), as well as the addition of sequences with or without a refining sequence, such as a 6-histidine sequence.

[0077] In certain embodiments, modified Cpf1 proteins having NLS sequences located at or near both the N- and C-termini of the Cpf1 protein sequence can be selected from His-sNLS-AsCpf1-sNLS (SEQ ID NO: 9) and His-sNLS-sNLS-AsCpf1-sNLS-sNLS (SEQ ID NO: 10). Additional permutations of NLS sequence identity and N- / C-terminal position are within the scope of the presently disclosed subject matter, such as the addition of two or more nNLS sequences or combinations of nNLS and sNLS sequences (or other NLS sequences) to either the N- or C-terminal positions, as well as the addition of sequences with or without purification sequences, such as a 6-histidine sequence.

[0078] CD34 + To determine which Cpf1 protein modifications, such as NLS modifications, favor editing in CD34 and T cells, AsCpf1 proteins containing different positions and types of NLS sequences were synthesized. The protein variants were complexed with gRNA targeting consensus site 5 and expressed CD34. + Cells, T cells, and HUDEP (4.4 μM RNP) were electroporated into the CD34 nuclease-containing medium. In Figure 4, the results are depicted as % editing normalized to the variant that displayed the greatest editing in each cell type. The data show that different species of nucleases interact with CD34 nuclease-containing medium. + It has diverse activities at the same target site in CD34 and T cells (among other cells), and efficient editing by AsCpf1 is +We show that this can be achieved in cells and T cells (among other cells).

[0079] Cysteine-modified Cpf1 protein and RNP Disulfide bond formation is known to promote protein aggregation, therefore, the Cpf1 crystal structure and known Cpf1 primary amino acid sequence were analyzed in an effort to identify cysteines that could be modified to reduce the likelihood of such disulfide bond formation (Figure 13).

[0080] The modified Cpfl proteins of the present disclosure can include modifications (e.g., deletions or substitutions) at one or more cysteine ​​residues in the Cpfl protein sequence. Such modified Cpfl proteins exhibit reduced aggregation, which is particularly useful when scaling up production of the protein. For example, but not limited to, the modified Cpfl protein includes modifications at one or more positions, such as two or more, three or more, four or more, five or more, six or more, seven or more, or eight positions selected from the group consisting of C65, C205, C334, C379, C608, C674, C1025, and C1248. In certain embodiments, the modified Cpfl protein includes substitutions of one or more cysteine ​​residues with serine or alanine. In certain embodiments, the modified Cpfl protein comprises one or more modifications, such as, for example, substitutions selected from the group consisting of C65S, C205S, C334S, C379S, C608S, C674S, C1025S, and C1248S. In certain embodiments, the modified Cpfl protein comprises one or more modifications selected from the group consisting of C65A, C205A, C334A, C379A, C608A, C674A, C1025A, and C1248A. In certain embodiments, the modified Cpfl protein comprises modifications at positions C334 and C674 or C334, C379, and C674. In certain embodiments, the modified Cpfl protein comprises modifications at C334S and C674S or C334S, C379S, and C674S. In certain embodiments, the modified Cpf1 protein includes modifications at C334A and C674A or C334A, C379A, and C674A. In certain embodiments, the modified Cpf1 protein includes both one or more cysteine ​​residue modifications and the introduction of one or more NLS sequences, such as, for example, His-AsCpf1-nNLS Cys-less (SEQ ID NO: 11) or His-AsCpf1-nNLS Cys-low (SEQ ID NO: 12) as described herein.

[0081] CD34 at target sites associated with hemoglobinopathies + Cpf1 editing in HSCs The present disclosure further provides CRISPR / Cpf1-related methods for editing target nucleic acid sequences to treat hemoglobinopathies, such as beta-thalassemia and sickle cell disease. For example, but not limited to, CRISPR / Cpf1-related methods can be used to modify CD34, which regulates the expression of fetal hemoglobin (HbF). + This results in the disruption of one or more genes in the cell.

[0082] One therapeutic strategy for treating hemoglobinopathy involves increasing the expression of HbF. HbF expression can be induced through targeted disruption of the erythroid cell-specific expression of transcriptional repressor BCL11a (Canvers et al., Nature, 527(12):192-197). One strategy for increasing HbF expression is to disrupt BCL11a expression using gene editing. For example, but not limited to, RNA-guided nucleases, such as Cpf1 RNA-guided nuclease, can target specific target sequences that affect the expression of the BCL11a gene. In certain embodiments, any region of the BCL11a gene can be targeted.

[0083] The present disclosure provides cells or cell populations comprising modifications in the BCL11a gene, e.g., to disrupt, knockdown, or knockout BCL11a expression. For example, but not limited to, the cells or cell populations can be generated by delivery of a complex comprising a Cpf1 RNA-guided nuclease and a gRNA molecule, e.g., an RNP complex targeting the BCL11a gene sequence. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in the cell population are modified. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in the cell population have productive indels.

[0084] In certain embodiments, the Cpf1 RNA-guided nuclease may target intron 2 of the BCL11a gene. In certain embodiments, the Cpf1 RNA-guided nuclease will be targeted to disrupt the GATA1 binding motif of the erythroid-specific enhancer of BCL11a, located in the +58 DHS region of intron 2 of the BCL11a gene. Exemplary gRNA molecules for use in such a CRISPR / Cpf1 editing system targeting BCL11a are identified in Figures 7, 10, and 12.

[0085] In certain embodiments, the present disclosure relates to cells in which the BCL11a gene has been disrupted. In certain embodiments, the erythroid enhancer region of the BCL11a gene may be targeted, for example, the erythroid enhancer region between +55 kb and +62 kb from the transcription start site (TSS). For example, but not by way of limitation, the present disclosure relates to cells in which the +58 DHS region of intron 2 of the BCL11a gene has been disrupted. In certain embodiments, such cells may comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, the present disclosure relates to cell populations in which the BCL11a gene has been disrupted, for example, the +58 DHS region of intron 2 of the BCL11a gene has been disrupted. In certain embodiments, such cell populations include cells comprising one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, the present disclosure relates to cells in which the GATA1 motif of the BCL11a gene has been disrupted. In certain embodiments, such cells may comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, the present disclosure relates to a cell population in which the GATA1 motif of the BCL11a gene has been disrupted. In certain embodiments, such a cell population may comprise cells comprising one or more components of a CRISPR / Cpf1 editing system.

[0086] As outlined in Example 3 below, AsCpf1 successfully mediated editing of a target site within the +58 DHS region of intron 2 of the BCL11a gene. First, several AsCpf1 mutant guide RNAs with different PAMs (Figure 1) were screened in HUDEP2 cells, and the most efficient guide RNA and nuclease variants were then cloned into mPB CD34 cells. + The effects of AsCpf1 on the BCL11a enhancer region were examined in cells (Figure 17). In particular, Figure 17 shows the screening of the BCL11a enhancer region with AsCpf1 WT and RR and RVR PAM mutants, along with one WT FnCpf1 target in HUDEPs and HSCs.

[0087] For example, in the context of treating hemoglobinopathies such as β-thalassemia and sickle cell disease, another strategy to induce fetal hemoglobin expression is to disrupt expression of the HBG locus, particularly HGB1 and / or HGB2.

[0088] In certain embodiments, the present disclosure relates to the use of CRISPR / Cpf1-mediated editing of the HBG locus. In certain embodiments, any region of the HBG locus may be targeted. In certain embodiments, CRISPR / Cpf1-mediated editing as described herein may be used to disrupt a non-coding region of the HBG locus (see, e.g., Table 18). In certain embodiments, CRISPR / Cpf1-mediated editing as described herein may be used to disrupt an intron of the HBG locus. In certain embodiments, CRISPR / Cpf1-mediated editing as described herein may be used to disrupt a targeted cis-regulatory region of the HBG gene. For example, but not limited to, a cis-regulatory region may include a promoter and / or an enhancer. In certain embodiments, the present disclosure relates to the use of CRISPR / Cpf1-mediated editing of the promoter region of the HBG locus. In certain embodiments, CRISPR / Cpf1-mediated editing as described herein may be used to disrupt a region from -800 to -60 nt of the promoter region of the HBG locus. For example, but not by way of limitation, CRISPR / Cpf1-mediated editing can be used to disrupt the -110 nt promoter region of the HBG promoter region and / or the CAAT box present within the HBG promoter region. Generally, disruption of the HBG promoter region and disruption of the CAAT box can be achieved through delivery of a CRISPR / Cpf1 editing system that targets these sequences. Exemplary gRNA molecules for use with such CRISPR / Cpf1 editing systems that target these sequences of the HBG locus are identified in Figures 6, 9, and 11 and Table 19. Chromosomal regions (e.g., genomic coordinates) that can be targeted to disrupt the HBG locus are provided in Table 18. In certain embodiments, the gRNA molecule used to disrupt the HBG1 locus is HBG1-1.

[0089] The present disclosure provides cells or cell populations comprising modifications at the HBG locus, e.g., to disrupt, knock down, or knock out HBG expression. For example, but not limited to, the cells or cell populations can be generated by delivery of a complex comprising a Cpf1 RNA-guided nuclease and a gRNA molecule, e.g., an RNP complex that targets the HBG locus. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in the cell population are modified. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in the cell population have productive indels.

[0090] In certain embodiments, the present disclosure relates to cells, such as CD34+ hematopoietic stem and progenitor cells, in which the HBG locus has been disrupted. For example, but not by way of limitation, the present disclosure relates to cells in which the promoter region of the HBG locus has been disrupted. In certain embodiments, the −110 nt promoter region of the HBG locus is disrupted. In certain embodiments, such cells may comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, the present disclosure relates to cell populations in which the −110 nt promoter region of the HBG locus has been disrupted. In certain embodiments, such cell populations may include cells that comprise one or more components of a CRISPR / Cpf1 editing system, as determined using a suitable method for detecting such components. In certain embodiments, the present disclosure relates to cells in which the CAAT box present in the HBG promoter region has been disrupted. In certain embodiments, such cells comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, the present disclosure relates to cell populations in which the CAAT box present in the HBG promoter region has been disrupted. In certain embodiments, such cell populations may include cells that contain one or more components of a CRISPR / Cpf1 editing system, as determined using a suitable method for detecting such components. In certain embodiments, the present disclosure provides cell populations in which the HBG1 locus has been disrupted by use of a CRISPR / Cpf1 editing system that includes gRNA HBG1-1.

[0091] In certain embodiments, a CRISPR / Cpf1 edited cell or population of CRISPR / Cpf1 edited cells comprising a modification in the HBG locus or BCL11a gene does not contain one or more components of such a CRISPR / Cpf1 editing system, as determined using a suitable method for detecting such components. In certain embodiments, less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% of the cells in the CRISPR / Cpf1 edited cell population contain one or more components of the CRISPR / Cpf1 editing system, as determined using a suitable method for detecting such components. In certain embodiments, the disclosure provides a CRISPR / Cpf1 edited cell population administered to a subject in need thereof, wherein less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% of the cells in the CRISPR / Cpf1 edited cell population comprise one or more components of the CRISPR / Cpf1 editing system.

[0092] In certain embodiments, disruption of the BCL11a gene or HBG gene in a cell using the CRISPR / Cpf1 editing system of the present disclosure can result in increased expression of fetal hemoglobin in the cell compared to a cell lacking the disruption of the BCL11a gene or HBG gene. For example, but not limited to, compared to the expression level of fetal hemoglobin in a cell lacking the disruption of the BCL11a gene or HBG locus and / or gene, expression of fetal hemoglobin can be increased by at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, or at least about 95%.

[0093] In certain embodiments, disruption of the BCL11a gene or HBG gene in a cell by the CRISPR / Cpf1 editing system of the present disclosure can result in increased expression of fetal hemoglobin in an amount suitable to partially or completely alleviate symptoms of a hemoglobinopathy, such as, for example, sickle cell disease or beta-thalassemia. For example, but not limited to, the increase in fetal hemoglobin expression can be greater than about 1 picogram (pg), greater than about 2 pg, greater than about 3 pg, greater than about 4 pg, greater than about 5 pg, greater than about 6 pg, greater than about 7 pg, greater than about 8 pg, greater than about 9 pg, greater than about 10 pg, greater than about 11 pg, greater than about 12 pg, greater than about 13 pg, greater than about 14 pg, or greater than about 15 pg.

[0094] In certain embodiments, disruption of the BCL11a gene or the HBG gene in a cell with the CRISPR / Cpf1 editing system of the disclosure can result in production of at least about 1 picogram, at least about 2 picograms, at least about 3 picograms, at least about 4 picograms, at least about 5 picograms, at least about 6 picograms, at least about 7 picograms, at least about 8 picograms, at least about 9 picograms, at least about 10 picograms, or about 8 to about 9 picograms or about 9 to about 10 picograms of fetal hemoglobin per cell.

[0095] The present disclosure also relates to a cell population modified by the above-described genome editing system, in which a higher percentage of the cell population can differentiate into a cell population of erythroid lineage that expresses HbF compared to a cell population not modified by the genome editing system. In certain embodiments, the higher percentage can be at least about 15%, at least about 20%, at least about 25%, at least about 30%, or at least about 40% higher. In certain embodiments, the cells can be hematopoietic stem cells. In certain embodiments, the cells can differentiate into erythroblasts, erythrocytes, or erythroid precursors or erythroblasts.

[0096] In certain embodiments, expression levels, such as the relative expression level of HbF (eg, relative to total β-like globin chains), may be measured by ultra-performance liquid chromatography (UPLC).

[0097] The CRISPR / Cpf1 editing system of the present disclosure can be delivered to cells using various strategies. For example, but not limited to, vectors, such as AAV or other viral vectors, encoding the components of the CRISPR / Cpf1 editing system can be used to induce expression of the components of the CRISPR / Cpf1 editing system in cells. Alternatively, RNP complexes containing the components of the CRISPR / Cpf1 editing system can be introduced into cells, for example, through electroporation. In certain embodiments, the RNP complexes can be delivered to cells by lipid nanoparticles.

[0098] As outlined in Example 3 below, FIG. 16 shows successful targeting of the HBG1 promoter region by AsCpf1 WT and RR PAM mutants in HUDEPs and HSCs.

[0099] Together, these data on disruption of the BCL11a gene and HBG locus support the finding that disruption of the CD34 gene at a clinically relevant locus (i.e., a known HPFH target site) + 1 shows efficient editing by AsCpf1 mutants in cells.

[0100] Cpf1 editing in T cells at target sites associated with T cell proliferation, survival, and / or function One therapeutic strategy proposed for treating cancer involves adoptive T cell transfer. Factors limiting the effectiveness of genetically modified T cells as cancer therapeutics include (1) T cell proliferation, e.g., limited proliferation of T cells following adoptive transfer; (2) T cell survival, e.g., induction of T cell apoptosis by factors in the tumor environment; and (3) T cell function, e.g., inhibition of cytotoxic T cell function by inhibitory factors secreted by host immune cells and cancer cells. One strategy to enhance efficacy is to use gene editing to modify or disrupt T cell genes associated with T cell proliferation, survival, and / or function. For example, but not limited to, RNA-guided nucleases, such as Cpf1 RNA-guided nuclease, can target specific sequences that affect the expression of T cell genes.

[0101] Methods and compositions encompassed by the present disclosure may be used to affect T cell proliferation, survival, persistence, and / or function by modifying one or more T cell-expressed genes, such as, for example, one or more of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes. In certain embodiments, the methods and compositions disclosed herein may be used to affect T cell proliferation by modifying one or more T cell-expressed genes, such as, for example, the CBLB and / or PTPN6 genes. In certain embodiments, the methods and compositions disclosed herein may be used to affect T cell survival by modifying one or more T cell-expressed genes, such as, for example, the FAS and / or BID genes. In certain embodiments, the methods and compositions disclosed herein may be used to affect T cell function by modifying one or more T cell-expressed genes, such as, for example, the CTLA4, PDCD1, TRAC, CIITA, and / or TRBC genes. In certain embodiments, the methods and compositions disclosed herein may be used to improve T cell persistence by modifying the B2M gene.

[0102] In certain embodiments, one or more T cell expressed genes, including but not limited to, FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes, are independently targeted for targeted knockout, e.g., to affect T cell proliferation, survival, persistence, and / or function. In certain embodiments, the presently disclosed methods involve knocking out one T cell expressed gene (e.g., one selected from the group consisting of FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes). In certain embodiments, the presently disclosed methods involve independently knocking out two T cell expressed genes (e.g., two selected from the group consisting of FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes). In certain embodiments, the presently disclosed methods involve independently knocking out three T cell expressed genes, such as, for example, three selected from the group consisting of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes. In certain embodiments, the presently disclosed methods involve independently knocking out four T cell expressed genes, such as, for example, four selected from the group consisting of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes. In certain embodiments, the presently disclosed methods involve independently knocking out five T cell expressed genes, such as, for example, five selected from the group consisting of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes. In certain embodiments, the presently disclosed methods involve independently knocking out six T cell expressed genes, such as six selected from the group consisting of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes.In certain embodiments, the presently disclosed methods involve independently knocking out seven T cell expressed genes, such as seven selected from the group consisting of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes. In certain embodiments, the presently disclosed methods involve independently knocking out eight T cell expressed genes, such as seven selected from the group consisting of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes. In certain embodiments, the presently disclosed methods involve independently knocking out nine T cell expressed genes, such as nine selected from the group consisting of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes. In certain embodiments, the presently disclosed methods involve independently knocking out nine T cell expressed genes, such as, for example, the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes.

[0103] In addition to the genes described above, several other T cell-expressed genes can be targeted to affect the efficacy of genetically engineered T cells. These genes include, but are not limited to, TGFBRI, TGFBRII, and TGFBRIII (Kershaw et al. 2013 Nat. Rev. Cancer 13, 525-541). In certain embodiments, using the methods disclosed herein, one or more of the TGFBRI, TGFBRII, and TGFBRIII genes can be modified either individually or in combination. In certain embodiments, using the methods disclosed herein, one or more of the TGFBRI, TGFBRII, and TGFBRIII genes can be modified either individually or in combination with any one or more of the eight genes listed above (i.e., FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes).

[0104] In certain embodiments, the methods and compositions disclosed herein modify the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and / or TRBC genes by targeting a location of the gene (e.g., a knockout location), such as a location within a non-coding region (e.g., a promoter region or regulatory region) or a location within the coding region, or by targeting a transcribed sequence of the gene, such as an intronic sequence or an exon sequence. In certain embodiments, a coding sequence, e.g., a coding region, e.g., an early coding region of a gene (e.g., the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and / or TRBC genes), is targeted for expression modification and knockout. In certain embodiments, locations within non-coding regions (e.g., promoter or regulatory regions) of T cell expressed genes (e.g., FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and / or TRBC genes) are targeted for modification and knockout of expression of T cell expressed genes.

[0105] In certain embodiments, the methods and compositions disclosed herein modify the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and / or TRBC genes by targeting the coding sequence of the gene. In certain embodiments, the coding sequence is an early coding sequence. In certain embodiments, the coding sequence of the gene is targeted to knock out expression of a T cell expressible gene.

[0106] In certain embodiments, the methods and compositions disclosed herein modify the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and / or TRBC genes by targeting a non-coding sequence of the gene. In certain embodiments, the non-coding sequence comprises a sequence within a promoter region, an enhancer sequence, an intron sequence, a sequence within a 3'UTR, a polyadenylation signal sequence, or a combination thereof. In certain embodiments, the non-coding sequence of the gene is targeted to knock out expression of the gene.

[0107] In certain embodiments, the presently disclosed methods include knocking out one or two alleles of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and / or TRBC genes, for example, by inducing genetic modifications, which in certain embodiments include insertions, deletions, mutations, or combinations thereof.

[0108] In certain embodiments, the targeted knockout approach is mediated by non-homologous end joining (NHEJ) using the CRISPR / Cpf1 system, which includes the Cpf1 enzyme.

[0109] In certain embodiments, the CRISPR / Cpfl system disclosed herein targets the TRAC gene. In certain embodiments, the CRISPR system includes a gRNA that is complementary to a portion of the TRAC gene sequence. In certain embodiments, the gRNA can be complementary to either strand of the TRAC gene. In certain embodiments, the targeting portion of the TRAC gene sequence is within the coding sequence of the TRAC gene. In certain embodiments, the targeting portion of the TRAC gene sequence is within an exon. In certain embodiments, the targeting portion of the TRAC gene sequence is within an intron. In certain embodiments, the targeting portion of the TRAC gene sequence is within a regulatory region of the gene. In certain embodiments, more than one sequence is targeted, and the targeting portion of the TRAC gene sequence is within one or more exons, one or more introns, one or more regulatory regions, or one or more exons, one or more introns and one or more regulatory regions. In certain embodiments, the portion of the TRAC gene sequence is within the first 500 bp of the coding sequence of the TRAC gene. In certain embodiments, gRNA molecule targeting domains for use in such CRISPR / Cpf1 systems targeting TRAC comprise targeting domain sequences listed in Tables 2 and 3. The present disclosure provides compositions comprising one or more of the gRNAs provided in Tables 2 and 3. The present disclosure further provides compositions comprising one or more RNP complexes comprising one or more of the gRNAs provided in Tables 2 and 3.

[0110] In certain embodiments, the CRISPR / Cpf1 system disclosed herein targets the TRBC gene. In certain embodiments, the CRISPR system includes a gRNA complementary to a portion of the TRBC gene sequence. In certain embodiments, the gRNA can be complementary to either strand of the TRBC gene. In certain embodiments, the targeting portion of the TRBC gene sequence is within the coding sequence of the TRBC gene. In certain embodiments, the targeting portion of the TRBC gene sequence is within an exon. In certain embodiments, the targeting portion of the TRBC gene sequence is within an intron. In certain embodiments, the targeting portion of the TRBC gene sequence is within a regulatory region of the gene. In certain embodiments, more than one sequence is targeted, and the targeting portion of the TRBC gene sequence is within one or more exons, one or more introns, one or more regulatory regions, or one or more exons, one or more introns and one or more regulatory regions. In certain embodiments, the portion of the TRBC gene sequence is within the first 500 bp of the coding sequence of the TRBC gene. In certain embodiments, gRNA molecule targeting domains for use in such CRISPR / Cpf1 systems targeting TRBCs comprise targeting domain sequences listed in Tables 4 and 5. The present disclosure provides compositions comprising one or more of the gRNAs provided in Tables 4 and 5. The present disclosure further provides compositions comprising one or more RNP complexes comprising one or more of the gRNAs provided in Tables 4 and 5.

[0111] In certain embodiments, the CRISPR / Cpfl system disclosed herein targets the B2M gene. In certain embodiments, the CRISPR system includes a gRNA that is complementary to a portion of the B2M gene sequence. In certain embodiments, the gRNA can be complementary to either strand of the B2M gene. In certain embodiments, the targeting portion of the B2M gene sequence is within the coding sequence of the B2M gene. In certain embodiments, the targeting portion of the B2M gene sequence is within an exon. In certain embodiments, the targeting portion of the B2M gene sequence is within an intron. In certain embodiments, the targeting portion of the B2M gene sequence is within a regulatory region of the gene. In certain embodiments, more than one sequence is targeted, and the targeting portion of the B2M gene sequence is within one or more exons, one or more introns, one or more regulatory regions, or one or more exons, one or more introns and one or more regulatory regions. In certain embodiments, the portion of the B2M gene sequence is within the first 500 bp of the coding sequence of the B2M gene. In certain embodiments, the portion of the B2M gene sequence is between the 501st and last nucleotides of the coding sequence of the B2M gene. In certain embodiments, the targeting domain of a gRNA molecule for use in such a CRISPR / Cpf1 system that targets B2M comprises a targeting domain sequence listed in Tables 6, 7, and 8. In certain embodiments, the targeting domain of a gRNA molecule for use in such a CRISPR / Cpf1 system that targets B2M comprises AGUGGGGGUGAAUUCAGUGU. The present disclosure provides compositions comprising one or more gRNAs provided in Tables 6, 7, and 8. The present disclosure further provides compositions comprising one or more RNP complexes comprising one or more gRNAs provided in Tables 6, 7, and 8.

[0112] In certain embodiments, the CRISPR / Cpf1 system disclosed herein targets the CIITA gene. In certain embodiments, the CRISPR system includes a gRNA complementary to a portion of the CIITA gene sequence. In certain embodiments, the CRISPR system includes a gRNA complementary to a portion of the CIITA gene sequence. In certain embodiments, the gRNA may be complementary to either strand of the CIITA gene. In certain embodiments, the targeting portion of the CIITA gene sequence is within the coding sequence of the CIITA gene. In certain embodiments, the targeting portion of the CIITA gene sequence is within an exon. In certain embodiments, the targeting portion of the CIITA gene sequence is within an intron. In certain embodiments, the targeting portion of the CIITA gene sequence is within a regulatory region of the gene. In certain embodiments, two or more sequences are targeted, and the targeting portion of the CIITA gene sequence is within one or more exons, one or more introns, one or more regulatory regions, or one or more exons, one or more introns and one or more regulatory regions. In certain embodiments, the portion of the CIITA gene sequence is within the first 500 bp of the coding sequence of the CIITA gene. In certain embodiments, a gRNA molecule targeting domain for use in such a CRISPR / Cpf1 system targeting CIITA comprises a targeting domain sequence listed in Table 9. The present disclosure provides compositions comprising one or more gRNAs provided in Table 9. The present disclosure further provides compositions comprising one or more RNP complexes comprising one or more gRNAs provided in Table 9.

[0113] TIFF2024023294000002.tif231170

[0114] TIFF2024023294000003.tif82170

[0115] TIFF2024023294000004.tif153170

[0116] TIFF2024023294000005.tif231170

[0117] TIFF2024023294000006.tif230170

[0118] TIFF2024023294000007.tif231170

[0119] TIFF2024023294000008.tif233170

[0120] TIFF2024023294000009.tif22170

[0121] TIFF2024023294000010.tif216170

[0122] TIFF2024023294000011.tif47170

[0123] TIFF2024023294000012.tif189170

[0124] TIFF2024023294000013.tif234170

[0125] TIFF2024023294000014.tif234170

[0126] TIFF2024023294000015.tif230170

[0127] TIFF2024023294000016.tif58170

[0128] TIFF2024023294000017.tif178170

[0129] TIFF2024023294000018.tif234170

[0130] TIFF2024023294000019.tif234170

[0131] TIFF2024023294000020.tif54170

[0132] TIFF2024023294000021.tif185170

[0133] TIFF2024023294000022.tif180170

[0134] TIFF2024023294000023.tif60170

[0135] TIFF2024023294000024.tif232170

[0136] TIFF2024023294000025.tif234170

[0137] TIFF2024023294000026.tif126170

[0138] Knockout and / or knockdown of FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC may be useful in a variety of contexts, including, but not limited to, those related to adoptive immunotherapy for treating cancer and non-cancer diseases such as autoimmune disorders. According to certain embodiments of the present disclosure, FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC are knocked out in immune cells, such as T cells, used in therapy. As a non-limiting example, T cells may express an engineered receptor, such as a chimeric antigen receptor (CAR) or a heterologous T cell receptor (TCR), which may be configured to recognize an antigen on cells or tissues implicated in pathology, such as tumor cells. Regardless of whether they express engineered receptors, TCR, MHC I, and / or MHC II knockout T cells according to the present disclosure may be used to target tissues or organs in which GvH or HvG responses may present safety or efficacy concerns.

[0139] TCR, MHC I and / or MHC II knockout and / or knockdown cells may be used in "allogeneic" cell therapy, in which cells are harvested from a subject, modified to knock out or knock down, for example, by disrupting FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC expression, and then returned to a different subject. In either approach, between harvest and administration, the TCR, MHC I and / or MHC II knockout and / or knockdown cells of the present disclosure may be manipulated in various ways, such as expanded, stimulated, purified or sorted, transduced with a transgene, frozen and / or thawed, etc.

[0140] Knocking out or knocking down the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and / or TRBC genes as described herein can (1) prevent GvH responses; (2) prevent HvG responses; and / or (3) improve the safety and efficacy of T cells. Knocking down expression of the FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and / or TRBC proteins as described herein can similarly (1) prevent GvH responses; (2) prevent HvG responses; and / or (3) improve the safety and efficacy of T cells.

[0141] In certain embodiments, the presently disclosed methods comprise independently knocking out and / or knocking down one or more genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC in T cells. In certain embodiments, the presently disclosed methods comprise independently knocking out and / or knocking down two genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC in T cells. In certain embodiments, the presently disclosed methods comprise independently knocking out and / or knocking down three genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC in T cells. In certain embodiments, the presently disclosed methods comprise independently knocking out and / or knocking down all four genes, B2M, TRAC, CIITA, and TRBC, in T cells.

[0142] In certain embodiments, the presently disclosed methods comprise knocking out and / or knocking down the B2M gene in T cells. In certain embodiments, the presently disclosed methods comprise knocking out and / or knocking down the TRAC gene in T cells. In certain embodiments, the presently disclosed methods comprise knocking out and / or knocking down the CIITA gene in T cells. In certain embodiments, the presently disclosed methods comprise knocking out and / or knocking down the TRBC gene in T cells. In certain embodiments, the presently disclosed methods comprise knocking out and / or knocking down the B2M and TRAC genes in T cells. In certain embodiments, the presently disclosed methods comprise knocking out and / or knocking down the B2M and CIITA genes in T cells. In certain embodiments, the presently disclosed methods comprise knocking out and / or knocking down the B2M and TRBC genes in T cells. In certain embodiments, the presently disclosed methods comprise knocking out and / or knocking down the TRAC and CIITA genes in T cells. In certain embodiments, the presently disclosed methods involve knocking out and / or knocking down the TRAC and TRBC genes in T cells. In certain embodiments, the presently disclosed methods involve knocking out and / or knocking down the CIITA and TRBC genes in T cells. In certain embodiments, the presently disclosed methods involve knocking out and / or knocking down the B2M, TRAC, and CIITA genes in T cells. In certain embodiments, the presently disclosed methods involve knocking out and / or knocking down the B2M, TRAC, and TRBC genes in T cells. In certain embodiments, the presently disclosed methods involve knocking out and / or knocking down the B2M, CIITA, and TRBC genes in T cells. In certain embodiments, the presently disclosed methods involve knocking out and / or knocking down the TRAC, CIITA, and TRBC genes in T cells.In certain embodiments, the presently disclosed methods involve knocking out and / or knocking down the B2M, TRAC, CIITA, and TRBC genes in T cells.

[0143] In certain embodiments, knockout and / or knockdown of one or more genes, two or more genes, three or more genes, or four or more genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC in T cells may (1) prevent a GvH response; (2) prevent an HvG response; and / or (3) improve the safety and efficacy of the T cells. For example, without limitation, knockout and / or knockdown of one or more genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC in T cells may be used to generate "allogeneic" cells, such as allogeneic T cells. In certain embodiments, knockout and / or knockdown of one or more genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC may be used in "allogeneic" cell therapy, in which cells are removed from a subject, modified to be knocked out or knocked down, e.g., to disrupt B2M, TRAC, CIITA, and / or TRBC expression, and then returned to a different subject.

[0144] In certain embodiments, knocking out and / or knocking down one or more genes, two or more genes, three or more genes, or four or more genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC in T cells results in a decrease in MHC II receptor expression in the T cells compared to unmodified T cells. In certain embodiments, a cell population that has been modified to knock out and / or knock down one or more genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC exhibits at least about a 10%, at least about a 20%, at least about a 30%, at least about a 40%, at least about a 50%, at least about a 60%, at least about a 70%, at least about a 80%, or at least about a 90% decrease in MHC II receptor, TCR, or B2M expression compared to the amount of MHC II receptor, TCR, or B2M expression in an unmodified cell population.

[0145] In certain embodiments, knocking out and / or knocking down two or more genes may involve using different nucleases for editing each target gene. For example, but not limited to, a CRISPR / Cpfl editing system may be used to knock out and / or knock down one target gene, and a CRISPR / Cas9 editing system may be used to knock out and / or knock down a second target gene.

[0146] The present disclosure provides isolated CRISPR / Cpf1-edited T cells or CRISPR / Cpf1-edited T cell populations comprising one or more modifications in one or more endogenous genes of the T cells disclosed herein. In certain embodiments, the CRISPR / Cpf1-edited T cells or CRISPR / Cpf1-edited T cell populations comprise one or more components of a CRISPR / Cpf1 editing system. Alternatively, the CRISPR / Cpf1-edited T cells or CRISPR / Cpf1-edited T cell populations do not comprise one or more components of a CRISPR / Cpf1 editing system. In certain embodiments, less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% of the cells in the CRISPR / Cpf1-edited cell population comprise one or more components of the CRISPR / Cpf1 editing system.

[0147] In certain embodiments, the T cells are CD8 + T cells, CD8 + Naive T cells, CD4 + central memory T cells, CD8 + central memory T cells, CD4 + Effector memory T cells, CD4 + Effector memory T cells, CD4 + T cells, CD4 + Stem cell memory T cell, CD8 + Stem cell memory T cell, CD4 + Helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, CD4+ naive T cells, TH17CD4 + T cells, TH1CD4 + T cells, TH2CD4 + T cells, TH9CD4 + T cells, CD4 + Foxp3 + T cells, CD4 + CD25 + CD127 - T cells or CD4 + CD25 + CD127 - Foxp3 + T cells.

[0148] In certain embodiments, the present disclosure relates to the use of CRISPR / Cpf1-mediated editing of endogenous genes in T cells selected from the group consisting of FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, TRBC, and any combination thereof. For example, but not limited to, the modification is generated by delivery of one or more complexes comprising a Cpf1 RNA-guided nuclease, e.g., an RNP complex, and a gRNA molecule, targeting, e.g., a portion of the FAS gene sequence, a portion of the BID gene sequence, a portion of the CTLA4 gene sequence, a portion of the PDCD1 gene sequence, a portion of the CBLB gene sequence, a portion of the PTPN6 gene sequence, a portion of the B2M gene sequence, a portion of the TRAC gene sequence, a portion of the CIITA gene sequence, a portion of the TRBC gene sequence, or a combination thereof. In certain embodiments, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten complexes, e.g., RNP complexes, may be delivered, each targeting a different gene. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in a T cell population are edited and / or modified. In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells of the T cell population have a productive indel in at least one endogenous T cell gene selected from the group consisting of, for example, FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC.

[0149] Benchmarking assays for Cpf1 variants, different cell types and formulations CRISPR / Cpf1-mediated editing of a target nucleic acid sequence and / or modulation of expression of a target nucleic acid sequence can be assessed by comparing the activity of a control CRISPR / RNA-guided nuclease editing system and a test CRISPR / Cpf1 editing system with respect to a target nucleic acid sequence, such as, for example, a "match site" target nucleic acid sequence.

[0150] The matched site target nucleic acid sequence incorporates both the requirement that it be edited by Cpf1 and a second RNA-guided nuclease, such as Cas9. For example, the TTTV AsCpf1 wild-type protospacer adjacent motif ("PAM") and the NGG SpCas9 wild-type PAM can be used in this example. As described above, the test Cpf1 protein can contain one or more modifications compared to the wild-type Cpf1 protein. Examples of such modifications include, but are not limited to, the aforementioned modifications to incorporate one or more NLS sequences, the aforementioned modifications to incorporate a 6-histidine purification sequence, and alterations of the cysteine ​​amino acids in the Cpf1 protein, as well as combinations thereof.

[0151] Exemplary match site target nucleic acid sequences that can be used in this example include match site 1 ("MS1"; SEQ ID NO: 13), match site 5 ("MS5"; SEQ ID NO: 14), match site 11 ("MS11"; SEQ ID NO: 15), and match site 18 ("MS18"; SEQ ID NO: 16).

[0152] For example, CD34 +To evaluate CRISPR / Cpf1-mediated target nucleic acid sequence editing and / or modulation of target nucleic acid sequence expression compared to CRISPR / Cas9-mediated editing in a particular cell type, such as HSCs, a CRISPR / Cpf1 genome editing system, i.e., a system comprising a Cpf1 RNA-guided nuclease and a gRNA complementary to at least a portion of a target nucleic acid containing a matched site target, is introduced into cells of the cell type of interest, for example, as an RNP or via the use of a vector encoding the system components. The target nucleic acid sequence editing and / or modulation of target nucleic acid sequence expression can be detected as disclosed herein. The detected target nucleic acid sequence editing and / or modulation of target nucleic acid sequence expression can then be compared to the target nucleic acid sequence editing and / or modulation of target nucleic acid sequence expression detected when using the CRISPR / Cas9 genome editing system on the same matched site target and the same cell type.

[0153] The above-described method for comparing CRISPR / Cpf1-mediated editing of target nucleic acid sequences (or editing by another CRISPR-based system) and / or the modulation of expression of target nucleic acid sequences allows for the evaluation of specific attributes of the CRISPR / Cpf1-mediated editing system used.For example, but not limited to, such a method can be used to evaluate CRISPR / Cpf1-mediated editing of target nucleic acid sequences and / or the modulation of expression of target nucleic acid sequences, and identify differences in the activity of Cpf1 RNA-guided nucleases and / or gRNAs prepared by different manufacturing processes.Such a method can also identify differences in the activity of Cpf1 RNA-guided nucleases and / or gRNAs that exist in different formulations and use different delivery strategies.

[0154] In certain embodiments, the present disclosure relates to assays for comparing CRISPR / Cpf1-mediated editing of a target nucleic acid sequence and / or modulation of expression of a target nucleic acid sequence by a test CRISPR / Cpf1 genome editing system with a control RNA-guided nuclease genome editing system. More specifically, the present disclosure provides assays using a match site (e.g., a cell containing match site 5) to which a gene editing system (e.g., CRISPR / Cas9 or CRISPR / Cpf1 or a variant thereof and a gRNA complementary to the match site) is targeted, such that the level or efficiency of editing at the match site is indicative of how efficient the gene editing system is at editing at any other sites. In other words, editing efficiency can be assessed by varying various components of the gene editing system and measuring the level or efficiency of editing achieved at the match site (e.g., match site 5).

[0155] For example, without limitation, the test and control gene or genome editing systems can differ in any one or more of the following aspects: the sequence of the RNA-guided nuclease; the source, e.g., the method of manufacture, of the components of the genome editing system; the formulation of one or more components of the genome editing system; and the identity of the cells into which the genome editing system is introduced, e.g., the cell type, or the method of preparation of the cells. In certain embodiments, the assays described herein enable quality control analysis of test genome editing systems. In certain embodiments, the assays disclosed assess CRISPR / Cpf1-mediated editing of a target nucleic acid sequence and / or modulation of expression of a target nucleic acid sequence, wherein the target comprises a match site sequence.

[0156] Electroporation Pulse Code Screening The present disclosure further provides electroporation pulse codes that result in higher editing at target sites. As shown in the Examples, screening of electroporation pulse codes allows for the identification of codes that result in higher efficiency editing by the Cpf1 RNA-guided nuclease of the present disclosure. For example, but not by way of limitation, Figure 18 shows a nucleofection screen of AsCpf1 in HUDEPs using a series of specific pulse codes and solutions. Similarly, Figure 19 shows an exemplary nucleofection screen of AsCpf1 in HSCs. In certain embodiments, pulse codes CA-137 and CA-138 can be used to promote higher efficiency editing by the Cpf1 RNA-guided nuclease. For example, but not by way of limitation, Figures 20 and 23C demonstrate the improved efficiency of the CA-137 pulse code.

[0157] Treatment method The present disclosure further provides methods for treating diseases and / or disorders by administering cells that have been edited using the disclosed genome editing methods. In certain embodiments, the present disclosure relates to methods for treating a subject by modifying one or more cells of the subject. In certain embodiments, one or more cells are modified ex vivo and then administered to the subject. For example, without limitation, a method for treating a subject may include contacting cells from the subject, e.g., ex vivo, with (a) a gRNA molecule complementary to a target sequence of a target nucleic acid; and (b) a Cpf1 RNA-guided nuclease as disclosed herein. In certain embodiments, the present disclosure provides methods for treating a subject, including administering to the subject one or more cells modified by the CRISPR / Cpf1 system of the present disclosure. In certain embodiments, one or more cells are obtained from a donor, genetically modified using the CRISPR / Cpf1 system of the present disclosure, and then administered to the subject.

[0158] In certain embodiments, the disclosed methods may include administering to a subject in need thereof T cells that have been edited using the disclosed genome editing methods, e.g., to produce allogeneic T cells. For example, but not limited to, the disclosed methods may include administering one or more T cells that have been edited to knock out or knock down expression of FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and / or TRBC. In certain embodiments, the T cells are edited to knock out or knock down B2M, TRAC, CIITA, and / or TRBC expression. In certain embodiments, one or more T cells are edited ex vivo and then administered to a subject. In certain embodiments, the one or more cells are obtained from a donor. In certain embodiments, such T cells may be used to treat a subject with cancer or an autoimmune disorder. In certain embodiments, in the CRISPR / Cpf1 edited T cell population administered to a subject, less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% of the cells in the CRISPR / Cpf1 edited cell population comprise one or more components of the CRISPR / Cpf1 editing system.

[0159] In certain embodiments, the methods of the present disclosure may include administering CD34+ hematopoietic stem and progenitor cells (HSPCs) that have been edited using the disclosed genome editing methods to a subject in need thereof. In certain embodiments, the CD34+ cells may be edited to knock out or knock down BCL11a or HBG expression. For example, but not limited to, CD34+ hematopoietic stem and progenitor cells (HSPCs) that have been edited using the genome editing methods disclosed herein may be used to treat a hemoglobinopathy in a subject in need thereof. In certain embodiments, the hemoglobinopathy may be severe sickle cell disease (SCD) or a thalassemia, such as β-thalassemia, δ-thalassemia, or β / δ-thalassemia. In certain embodiments, an exemplary protocol for treating a hemoglobinopathy may include harvesting CD34+ HSPCs from a subject in need thereof, editing the autologous CD34+ HSPCs ex vivo using the genome editing methods disclosed herein, and subsequently reinfusing the edited autologous CD34+ HSPCs into the subject. In certain embodiments, treatment with the edited autologous CD34+ HSPCs may result in increased HbF induction.

[0160] In certain embodiments, prior to collection of CD34+ HSPCs, the subject may discontinue hydroxyurea treatment, if applicable, and receive a blood transfusion to maintain adequate hemoglobin (Hb) levels. In certain embodiments, the subject may receive plerixafor (e.g., 0.24 mg / kg) intravenously to mobilize CD34+ HSPCs from the bone marrow into the peripheral blood. In certain embodiments, the subject may undergo one or more leukapheresis cycles (e.g., one cycle defined as two plerixafor-mobilized leukapheresis collections performed on consecutive days, with approximately one month between cycles). In certain embodiments, the number of leukapheresis cycles a subject undergoes depends on the dose of unedited autologous CD34+ HSPCs / kg for backup storage, along with a dose of unedited autologous CD34+ HSPCs / kg for reinfusion of the subject (e.g., ≥ 1.5 x 10 6 cells / kg), the dose of edited autologous CD34+ HSPCs (e.g., ≥ 2 × 10 6cells / kg, ≧3×10 6 cells / kg, ≧4×10 6 cells / kg, ≧5×10 6 cells / kg, 2×10 6 cells / kg~3×10 6 cells / kg, 3×10 6 cells / kg~4×10 6 cells / kg, 4×10 6 cells / kg~5×10 6

[0013] The number of times may be as many as necessary to achieve a 100% genomic DNA fragment size (1000 x 1000 cells / kg). In certain embodiments, CD34+ HSPCs harvested from a subject may be edited using any of the genome editing methods discussed herein. In certain embodiments, any one or more of the gRNAs and one or more of the RNA-guided nucleases disclosed herein may be used in the genome editing method.

[0161] In certain embodiments, treatment may include autologous stem cell transplantation. In certain embodiments, subjects may undergo myeloablative conditioning with busulfan conditioning (e.g., a test dose of 1 mg / kg, titrated based on initial dose pharmacokinetic analysis). In certain embodiments, conditioning may be performed for four consecutive days. In certain embodiments, after a three-day busulfan washout period, edited autologous CD34+ HSPCs (e.g., ≥ 2 x 10 6 cells / kg, ≧3×10 6 cells / kg, ≧4×10 6 cells / kg, ≧5×10 6 cells / kg, 2×10 6 cells / kg~3×10 6 cells / kg, 3×10 6 cells / kg~4×10 6 cells / kg, 4×10 6 cells / kg~5×10 6 cells / kg) can be reinfused (e.g., into peripheral blood) into the subject. In certain embodiments, edited autologous CD34+ HSPCs can be produced for a particular subject and cryopreserved. In certain embodiments, a subject can achieve neutrophil engraftment following a sequential myeloablative conditioning regimen and infusion of edited autologous CD34+ cells. Neutrophil engraftment occurs when ≥ 0.5 x 10 9 / L on three consecutive measurements of ANC. In certain embodiments, in the CRISPR / Cpf1-edited CD34+ HSPC population administered to a subject, less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% of the cells in the CRISPR / Cpf1-edited CD34+ HSPC population comprise one or more components of the CRISPR / Cpf1 editing system.

[0162] In certain embodiments, the CRISPR / Cpf1-mediated editing system of the present disclosure may result in a clinically or therapeutically relevant editing efficiency of about 10% or greater. For example, without limitation, the CRISPR / Cpf1-mediated editing system of the present disclosure may result in a clinically or therapeutically relevant editing efficiency of about 5% or greater, about 10% or greater, 15% or greater, about 20% or greater, about 25% or greater, about 30% or greater, about 35% or greater, about 40% or greater, about 45% or greater, about 50% or greater, about 55% or greater, about 60% or greater, about 65% or greater, about 70% or greater, about 75% or greater, about 80% or greater, about 85% or greater, about 90% or greater, about 95% or greater, about 96% or greater, about 97% or greater, about 98% or greater, or about 99% or greater.

[0163] In certain embodiments, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90% of the cells in a cell population administered in a therapeutic method disclosed herein are modified.

[0164] In certain embodiments, less than about 10%, less than about 5%, less than about 1%, less than about 0.5%, less than about 0.25%, or less than about 0.1% of the cells in the CRISPR / Cpf1 edited cell population comprise one or more components of the CRISPR / Cpf1 editing system.

[0165] Genome editing system The term "genome editing system" or "gene editing system" refers to any system with RNA-guided DNA editing activity. The genome editing system of the present disclosure includes at least two components, a guide RNA (gRNA) and an RNA-guided nuclease, adapted from the naturally occurring CRISPR system. These two components bind to a specific nucleic acid sequence to form a complex that can edit DNA in or around the nucleic acid sequence, for example, by generating one or more single-strand breaks (SSBs or nicks), double-strand breaks (DSCs), and / or point mutations.

[0166] Naturally occurring CRISPR systems are evolutionarily organized into two classes and five types (Makarova et al. Nat Rev Microbiol. 2011 Jun;9(6):467-477 (Makarova), incorporated herein by reference), and while the genome editing system of the present disclosure can adapt components of any type or class of naturally occurring CRISPR system, the embodiments presented herein are generally adapted from class 2 and type II or type V CRISPR systems. Class 2 systems, including types II and V, are characterized by a relatively large, multidomain RNA-guided nuclease protein (e.g., Cas9 or Cpfl) and one or more guide RNAs (e.g., crRNA, optionally tracrRNA), which form a ribonucleoprotein (RNP) complex that associates with (targets) and cleaves specific genetic loci complementary to the target (or spacer) sequence of the crRNA. The genome editing system according to the present disclosure similarly targets and edits cellular DNA sequences, but is significantly different from naturally occurring CRISPR systems. For example, the unimolecular guide RNAs described herein do not occur in nature, and guide RNAs and RNA-guided nucleases according to the present disclosure may incorporate any number of modifications that do not occur in nature.

[0167] Genome editing systems can be implemented in a variety of ways (e.g., administered or delivered to a cell or subject), and different implementations may be suitable for different uses. For example, in certain embodiments, genome editing systems are implemented as protein / RNA complexes (ribonucleoproteins or RNPs), which may be included in pharmaceutical compositions that optionally include a pharmaceutically acceptable carrier and / or encapsulating agent, such as lipid or polymer microparticles or nanoparticles, micelles, or liposomes. In certain embodiments, genome editing systems are implemented as one or more nucleic acids (optionally with one or more additional components) encoding the RNA-guided nuclease and guide RNA components described above; in certain embodiments, genome editing systems are implemented as one or more vectors, e.g., viral vectors such as adeno-associated viruses, containing such nucleic acids; and in certain embodiments, genome editing systems are implemented as any combination of the foregoing. Additional or modified implementations that operate according to the principles described herein will be apparent to those skilled in the art and are within the scope of this disclosure.

[0168] It should be noted that the genome editing system of the present disclosure can target or target a single specific nucleotide sequence, and the use of two or more guide RNAs can edit two or more specific nucleotide sequences in parallel. The use of multiple gRNAs, referred to throughout this disclosure as "multiplexing," can be used to target multiple unrelated target sequences of interest or to create multiple SSBs or DSBs within a single target domain, and optionally to make specific edits within such a target domain. For example, International Publication No. 2015 / 138510 by Maeder et al. (Maeder), incorporated herein by reference, describes a genome editing system for correcting a point mutation in the human CEP290 gene (C.2991+1655A to G), which results in the creation of a cryptic splice site and consequently reduces or eliminates gene function. Maeder's genome editing system utilizes two guide RNAs that target sequences on either side of (i.e., flanking) the point mutation to create a DSB flanking the mutation. This then facilitates the deletion of the intervening sequence containing the mutation, thereby eliminating the cryptic splice site and restoring normal gene function.

[0169] As another example, International Publication No. 2016 / 073990 by Cotta-Ramusino et al. ("Cotta-Ramusino et al."), incorporated herein by reference in its entirety, describes a genome editing system that utilizes two gRNAs in combination with a Cas9 nickase (a Cas9 that generates a single-stranded nick, such as S. pyogenes D10A), an arrangement referred to as a "dual nickase system." The dual nickase system of Cotta-Ramusino et al. is configured to generate two nicks in opposite strands of the sequence of interest, offset by one or more nucleotides, which combine to create a double-stranded break with an overhang (5' in the case of Cotta-Ramusino et al., although a 3' overhang is also possible). The overhang can then facilitate homology-directed repair events in some circumstances. Also, as another example, WO 2015 / 070083 by Palestrant et al. ("Palestrant," incorporated herein by reference in its entirety) describes a gRNA (referred to as a "controller RNA") that targets a nucleotide sequence encoding Cas9, which can be included in a genome editing system that includes one or more additional gRNAs, for example, to allow transient expression of Cas9 that may otherwise be constitutively expressed in some virally transduced cells. These multiplexing applications are intended to be exemplary rather than limiting, and one of skill in the art will understand that other multiplexing applications are generally compatible with the genome editing systems described herein.

[0170] Genome editing systems can optionally create double-strand breaks that are repaired by cellular DNA double-strand break mechanisms, such as NHEJ or HDR. These mechanisms are described throughout the literature, for example, in Davis & Maizels, PNAS, 111(10):E924-932, March 11, 2014 (Davis) (describing Alt-HDR); Frit et al. DNA Repair 17(2014)81-97 (Frit) (describing Alt-NHEJ); and Iyama and Wilson III, DNA Repair (Amst.) 2013-Aug;12(8):620-636 (Iyama) (describing canonical HDR and NHEJ pathways generally).

[0171] When a genome editing system functions by forming DSBs, such a system optionally includes one or more components that promote or facilitate a specific type of double-strand break repair or a specific repair outcome. For example, Cotta-Ramusino et al. also describe a genome editing system in which a single-stranded oligonucleotide "donor template" is added; the donor template can be incorporated into the target region of cellular DNA that is cut by the genome editing system, resulting in a change in the target sequence.

[0172] In certain embodiments, the genome editing system modifies a target sequence or modifies the expression of a gene in or near the target sequence without causing single- or double-strand breaks. For example, the genome editing system can include an RNA-guided nuclease fused to a functional domain that acts on DNA, thereby modifying the target sequence or its expression. As an example, the RNA-guided nuclease can be bound (e.g., fused) to a cytidine deaminase functional domain and function by generating a targeted C to A substitution. Exemplary nuclease / deaminase fusions are described in Komor et al. Nature 533, 420-424 (19 May 2016) ("Komor"), which is incorporated by reference. Instead, genome editing systems may utilize cleaving-inactivating (i.e., "dead") nucleases, such as dead Cas9 (dCas9), which may function by forming stable complexes on one or more targeted regions of cellular DNA, thereby disrupting functions involving the targeted regions, including, but not limited to, mRNA transcription, chromatin remodeling, etc.

[0173] In certain embodiments, the genome editing systems encompassed by the present disclosure will exhibit a certain minimum percentage of editing in a standard assay. For example, but not limited to, certain genome editing systems encompassed by the present disclosure will exhibit at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 95% editing in a certain standard assay. One or more assays known in the art or described herein, such as those described in Example 1 below, can be used to evaluate CRISPR / Cpf1-mediated editing of a target nucleic acid sequence. For example, in Example 1 below, for example, CD34 +Evaluation of CRISPR / Cpf1-mediated target nucleic acid sequence editing and / or modulation of target nucleic acid sequence expression compared to CRISPR / Cas9-mediated editing in a specific cell type, such as HSC, is described. A CRISPR / Cpf1 genome editing system, i.e., a system comprising a Cpf1 RNA-guided nuclease and a gRNA complementary to at least a portion of a target nucleic acid including a matched site target, is introduced into cells of the cell type of interest, for example, as an RNP or via the use of a vector encoding the system components. Editing of the target nucleic acid sequence and / or modulation of target nucleic acid sequence expression is detected as disclosed herein. The detected editing of the target nucleic acid sequence and / or modulation of target nucleic acid sequence expression can then be compared to the editing of the target nucleic acid sequence and / or modulation of target nucleic acid sequence expression detected when using the CRISPR / Cas9 genome editing system on the same matched site target and the same cell type.

[0174] In certain embodiments, the genome editing system of the present disclosure can knock out or knock down one or more, two or more, three or more, or four or more genes selected from the group consisting of B2M, TRAC, CIITA, and TRBC simultaneously in a cell population. In certain embodiments, the genome editing system of the present disclosure can include one or more, two or more, three or more, or four or more gRNA molecules, each gRNA molecule comprising a targeting domain for a different gene, such as a gene selected from the group consisting of B2M, TRAC, CIITA, and TRBC genes. For example, but not limited to, the multiplexed genome editing system of the present disclosure may include: (i) a first RNP complex comprising a first guide RNA (gRNA) comprising a first targeting domain complementary to a target sequence of a first gene and a first Cpf1 RNA-guided nuclease; (ii) a second RNP complex comprising a second gRNA molecule comprising a second targeting domain complementary to a target sequence of a second gene and a second Cpf1 RNA-guided nuclease; (iii) a third RNP complex comprising a third gRNA molecule comprising a third targeting domain complementary to a target sequence of a third gene and a fourth Cpf1 RNA-guided nuclease; and / or (iv) a fourth RNP complex comprising a fourth gRNA molecule comprising a fourth targeting domain complementary to a target sequence of a fourth gene and a fourth Cpf1 RNA-guided nuclease. In certain embodiments, the first gene, the second gene, the third gene, and the fourth gene are selected from the group consisting of B2M, TRAC, CIITA, and TRBC. In certain embodiments, the targeting domain of a gRNA molecule for targeting B2M comprises a targeting domain sequence listed in Tables 6, 7, and 8. In certain embodiments, the targeting domain of a gRNA molecule for targeting TRAC comprises a targeting domain sequence listed in Tables 2 and 3. In certain embodiments, the targeting domain of a gRNA molecule for targeting CIITA comprises a targeting domain sequence listed in Table 9. In certain embodiments, the targeting domain of a gRNA molecule for targeting TRBC comprises a targeting domain sequence listed in Tables 4 and 5. In certain embodiments, the editing efficiency may be >80%, >85%, >90%, >95%, >98%, or >99% for all target genes.In certain embodiments, the cell population may be a T cell population.

[0175] Guide RNA (gRNA) molecules The terms "guide RNA" and "gRNA" refer to any nucleic acid that facilitates the specific binding (or "targeting") of an RNA-guided nuclease, such as Cpf1, to a target sequence, such as a genomic or episomal sequence within a cell. gRNAs can be unimolecular (comprising a single RNA molecule or also referred to as chimeric) or modular (comprising two or more, typically two separate RNA molecules, such as crRNA and tracrRNA, that are usually linked to each other by duplexing). gRNAs and their component parts are described throughout the literature, for example, in Briner et al. (Molecular Cell 56(2), 333-339, October 23, 2014 (Briner), incorporated by reference) and Cotta-Ramusino.

[0176] In bacteria and archaea, type II CRISPR systems generally include an RNA-guided nuclease protein, such as Cas9; a CRISPR RNA (crRNA) that includes a 5' region complementary to an exogenous sequence; and a trans-activating crRNA (tracrRNA) that includes a 5' region that is complementary to and duplexes with the 3' region of the crRNA. Without intending to be bound by any theory, it is believed that this duplex promotes the formation of a Cas9 / gRNA complex and is required for its activity. While adapting type II CRISPR systems for use in gene editing, it was discovered that, in one non-limiting example, the crRNA and tracrRNA can be linked into a single, unimolecular or chimeric guide RNA by a four-nucleotide (e.g., GAAA) "tetraloop" or "linker" sequence that bridges the complementary regions of the crRNA (at its 3' end) and the tracrRNA (at its 5' end). (Mali et al. Science. 2013 Feb 15;339(6121):823-826 (“Mali”); Jiang et al. Nat Biotechnol. 2013 Mar;31(3):233-239 (“Jiang”); and Jinek et al., 2012 Science Aug. 17;337(6096):816-821 (“Jinek”), all of which are incorporated herein by reference.

[0177] Guide RNAs, whether monopartite or modular, contain a "targeting domain" that is fully or partially complementary to a target domain in a target sequence, such as a DNA sequence in the genome of a cell where editing is desired. Targeting domains have been referred to by various names in the literature, including, but not limited to, "guide sequence" (Hsu et al., Nat Biotechnol. 2013 Sep;31(9):827-832, ("Hsu"), incorporated herein by reference), "complementary region" (Cotta-Ramusino et al.), "spacer" (Briner), and collectively "crRNA" (Jiang). Regardless of the name given to them, targeting domains are typically 10-30 nucleotides in length, and in certain embodiments 16-24 nucleotides in length (e.g., 16, 17, 18, 19, 20, 21, 22, 23, or 24 nucleotides in length), and are located at or near the 5' end in the case of Cas9 gRNAs and at or near the 3' end in the case of Cpf1 gRNAs.

[0178] In addition to the targeting domain, gRNAs typically (though not necessarily, as discussed below) contain multiple domains that can affect the formation or activity of gRNA / Cas9 and gRNA / Cpf1 complexes. For example, as described above, the duplex structure formed by the first and second complementary domains of the gRNA (also referred to as the repeat:anti-repeat duplex) can interact with the recognition (REC) lobe of Cas9 to mediate the formation of the Cas9 / gRNA complex. (Nishimasu et al., Cell 156, 935-949, February 27, 2014 (Nishimasu 2014) and Nishimasu et al., Cell 162, 1113-1126, August 27, 2015 (Nishimasu 2015), both of which are incorporated herein by reference.) It should be noted that the first and / or second complementary domains may contain one or more polyA stretches that may be recognized as termination signals by RNA polymerase. As such, the sequences of the first and second complementary domains are optionally modified, for example, through the use of AG swaps or AU swaps as described in Briner, to remove these regions and facilitate complete in vitro transcription of the gRNA. These and other similar modifications to the first and second complementary domains are within the scope of the present disclosure.

[0179] Along with the first and second complementary domains, Cas9g RNAs typically contain two or more additional double-stranded regions that are involved in nuclease activity in vivo, but not necessarily in vitro (Nishimasu 2015). The first stem-loop 1, located near the 3' portion of the second complementary domain, is variously referred to as the "proximal domain" (Cotta-Ramusino), "stem-loop 1" (Nishimasu 2014 and 2015), and "nexus" (Briner). One or more additional stem-loop structures are generally present near the 3' end of the gRNA, with the number varying by species: S. pyogenes gRNAs typically contain two 3' stem-loops (for a total of four stem-loop structures including the repeat:anti-repeat duplex), while S. aureus and other species have only one (for a total of three stem-loop structures). A description of conserved stem-loop structures (and gRNA structures more generally) organized by species is provided in Briner.

[0180] While the foregoing description has focused on gRNAs for use with Cas9, it should be understood that other RNA-guided nucleases have been discovered or invented (or may be discovered in the future) that utilize gRNAs that differ in some respects from those described thus far. For example, Cpf1 (CRISPR from "Prevotella and Francicella 1") is a recently discovered RNA-guided nuclease that does not require a tracrRNA to function. (Zetsche et al., 2015, Cell 163, 759-771 October 22, 2015 (Zetsche I), incorporated herein by reference). gRNAs for use in the Cpf1 genome editing system generally include a targeting domain and a complementary domain (alternately referred to as a "handle"). It should also be noted that in gRNAs for use with Cpf1, the targeting domain is typically located at or near the 3' end rather than the 5' end, as described above for Cas9 gRNAs (the handle is at or near the 5' end of Cpf1 gRNA).

[0181] Those skilled in the art will understand that while structural differences may exist between gRNAs from different prokaryotic species or between the gRNAs of Cpfl and Cas9, the principles by which gRNAs operate are generally consistent. Because of this consistency of operation, gRNAs may be broadly defined by their targeting domain sequence, and those skilled in the art will understand that a given targeting domain sequence may be incorporated into any suitable gRNA, including monomolecular or chimeric gRNAs or gRNAs containing one or more chemical and / or sequence modifications (substitutions, additional nucleotides, truncations, etc.). Therefore, for economy of presentation in this disclosure, gRNAs may be described solely in terms of their targeting domain sequence.

[0182] More generally, those skilled in the art will understand that some aspects of the present disclosure relate to systems, methods, and compositions that can be implemented using multiple RNA-guided nucleases. For this reason, unless otherwise specified, the term gRNA should be understood to encompass not only gRNAs compatible with particular species of Cas9 or Cpfl, but any suitable gRNA that can be used with any RNA-guided nuclease. By way of example, the term gRNA may, in certain embodiments, include gRNAs for use with, or derived from, or adapted from, any RNA-guided nuclease present in a class 2 CRISPR system, such as type II or type V, or a CRISPR system.

[0183] The present disclosure provides gRNA molecules and compositions thereof comprising the sequence of any one of the gRNAs provided in Tables 2-9 and 19. The present disclosure further provides compositions and compositions thereof comprising one or more gRNAs comprising the sequence of a gRNA set forth in Tables 2-9 and 19. The present disclosure provides gRNAs and compositions thereof targeting the chromosomal regions (e.g., genomic coordinates) provided in Table 18.

[0184] The present disclosure provides gRNAs that result in greater than about 10% editing at a target site in, for example, a cell population. For example, without limitation, gRNAs of the present disclosure result in greater than about 15% editing, greater than about 20% editing, greater than about 25% editing, greater than about 30% editing, greater than about 35% editing, greater than about 40% editing, greater than about 45% editing, greater than about 50% editing, greater than about 55% editing, greater than about 60% editing, greater than about 65% editing, greater than about 70% editing, greater than about 75% editing, greater than about 80% editing, greater than about 85% editing, greater than about 90% editing, greater than about 95% editing, greater than about 96% editing, greater than about 97% editing, greater than about 98% editing, or greater than about 99% editing at a target site in, for example, a cell population.

[0185] gRNA design Methods for target sequence selection and validation and off-target analysis have been previously described, for example, in Mali; Hsu; Fu et al., 2014 Nat biotechnol 32(3):279-84, Heigwer et al., 2014 Nat methods 11(2):122-3; Bae et al. (2014) Bioinformatics 30(10):1473-5; and Xiao A et al. (2014) Bioinformatics 30(8):1180-1182. Each of these references is incorporated herein by reference. As a non-limiting example, gRNA design can involve the use of software tools to optimize the selection of potential target sequences corresponding to the user's target sequence, e.g., to minimize total off-target activity across the genome. While off-target activity is not limited to cleavage, the cleavage efficiency at each off-target sequence can be predicted, for example, using an empirically derived weighting scheme. These and other guide selection methods are described in detail in Maeder and Cotta-Ramusino et al.

[0186] gRNA modification The activity, stability, or other characteristics of gRNAs can be altered by incorporating certain modifications. As one example, transiently expressed or delivered nucleic acids may be susceptible to degradation by, for example, cellular nucleases. Accordingly, the gRNAs described herein may contain one or more modified nucleosides or nucleotides that confer stability against nucleases. Without wishing to be bound by theory, it is also believed that certain modified gRNAs described herein may exhibit a reduced innate immune response when introduced into cells. Those skilled in the art are aware of certain cellular responses commonly observed in cells, such as mammalian cells, in response to exogenous nucleic acids, particularly those of viral or bacterial origin. Such responses, which may include the induction of cytokine expression and release and cell death, can be reduced or completely eliminated by the modifications provided herein.

[0187] Certain exemplary modifications discussed in this section can be included anywhere in the gRNA sequence, including, but not limited to, at or near the 5' end (e.g., within 1-10, 1-5, or 1-2 nucleotides of the 5' end) and / or at or near the 3' end (e.g., within 1-10, 1-5, or 1-2 nucleotides of the 3' end). In some cases, modifications are located within functional motifs, such as the repeat:anti-repeat duplex of a Cas9 gRNA, the stem-loop structure of a Cas9 or Cpf1 gRNA, and / or the targeting domain of a gRNA.

[0188] As an example, the 5' end of the gRNA may include a eukaryotic mRNA cap structure or a G-cap analog (e.g., G(5')ppp(5')G-cap analog, m7G(5')ppp(5')G-cap analog, or 3'-O-Me-m7G(5')ppp(5')G-anti-cap analog (ARCA)), as shown below. TIFF2024023294000027.tif48161

[0189] The cap or cap analog can be included during chemical synthesis or in vitro transcription of the gRNA.

[0190] Along similar lines, the 5' end of the gRNA can lack a 5' triphosphate group, for example, the in vitro transcribed gRNA can be phosphatase treated (e.g., using calf intestinal alkaline phosphatase) to remove the 5' triphosphate group.

[0191] Another common modification involves adding multiple (e.g., 1-10, 10-20, or 25-200) adenine (A) residues, called a polyA tract, to the 3' end of the gRNA. PolyA sequences can be added to gRNAs during chemical synthesis following in vitro transcription using a polyadenosine polymerase (e.g., E. coli poly(A) polymerase) or in vivo by polyadenylation sequences as described by Maeder.

[0192] It should be noted that the modifications described herein may be combined in any suitable manner, for example, a gRNA transcribed from a DNA vector in vivo or transcribed in vitro may contain either or both a 5' cap structure or cap analog, and a 3' polyA sequence.

[0193] The guide RNA can be modified with a 3'-terminal U ribose. For example, the two terminal hydroxyl groups of U ribose can be oxidized to aldehyde groups, with concomitant opening of the ribose ring resulting in the modified nucleoside shown below: TIFF2024023294000028.tif29161 where "U" can be unmodified or modified uridine.

[0194] The 3' terminal U ribose may be modified with a 2'3' cyclic phosphate as shown below: TIFF2024023294000029.tif36161 where "U" can be unmodified or modified uridine.

[0195] Guide RNAs may contain 3' nucleotides that may be stabilized against degradation, for example, by incorporating one or more of the modified nucleotides described herein. In certain embodiments, uridine may be substituted with modified uridines, such as, for example, 5-(2-amino)propyluridine and 5-bromouridine, or any modified uridines described herein; adenosine and guanosine may be substituted with modified adenosines and guanosines modified, for example, at position 8, such as, for example, 8-bromoguanosine, or any modified adenosines or guanosines described herein.

[0196] In certain embodiments, sugar-modified ribonucleotides may be incorporated into the gRNA, e.g., where the 2'OH group is replaced with a group selected from H, -OR, -R (where R can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or a sugar), halo, -SH, -SR (where R can be, e.g., alkyl, cycloalkyl, aryl, aralkyl, heteroaryl, or a sugar), amino (where amino can be, e.g., NH; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, diheteroarylamino, or an amino acid); or cyano (-CN). In certain embodiments, the phosphate backbone may be modified as described herein, e.g., with a phosphorothioate (PhTx) group. In certain embodiments, one or more of the nucleotides of the gRNA may each independently be a modified or unmodified nucleotide, including, but not limited to, a 2'-sugar modification such as 2'-O-methyl, 2'-O-methoxyethyl, or a 2'-fluoro modification including, for example, 2'-F or 2'-O-methyl, adenosine (A), 2'-F or 2'-O-methyl, cytidine (C), 2'-F or 2'-O-methyl, uridine (U), 2'-F or 2'-O-methyl, thymidine (T) 1,2'-F or 2'-O-methyl, guanosine (G), 2'-O-methoxyethyl-5-methyluridine (Teo), 2'-O-methoxyethyl adenosine (Aeo), 2'-O-methoxyethyl-5-methylcytidine (m5CeO), and any combination thereof.

[0197] The guide gRNA may also comprise a "locked" nucleic acid (LNA), in which the 2'OH group can be linked, for example, by a C1-6 alkylene or C1-6 heteroalkylene bridge to the 4' carbon of the same ribose sugar. Examples of bridges that can be used include, but are not limited to, methylene, propylene, ether, or amino bridges; O-amino (wherein amino can be, for example, NH2; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino, or diheteroarylamino, ethylenediamine, or polyamino); and aminoalkoxy or O(CH2). n Any suitable moiety can be used, including -amino (wherein amino can be, for example, NH; alkylamino, dialkylamino, heterocyclyl, arylamino, diarylamino, heteroarylamino or diheteroarylamino, ethylenediamine or polyamino).

[0198] In certain embodiments, gRNAs may include modified nucleotides that are polycyclic (e.g., tricyclic; and "unlocked" forms such as glycol nucleic acids (GNAs) (e.g., R-GNAs or S-GNAs, in which the ribose is replaced with a glycol unit attached to a phosphodiester bond) or threose nucleic acids (TNAs, in which the ribose is replaced with α-L-threofuranosyl-(3'→2')).

[0199] Generally, gRNAs contain ribose, a five-membered sugar ring containing oxygen. Exemplary modified gRNAs include, but are not limited to, substitution of oxygen in ribose (e.g., with sulfur (S), selenium (Se), or alkylenes such as methylene or ethylene); addition of double bonds (e.g., replacing ribose with cyclopentenyl or cyclohexenyl); ribose ring contraction (e.g., forming a four-membered cyclobutane or oxetane ring); and ribose ring expansion (e.g., forming a six- or seven-membered ring with additional carbon or heteroatoms, such as anhydrohexitol, altritol, mannitol, cyclohexanyl, cyclohexenyl, and morpholino, which also has a phosphoramidate backbone). While the majority of sugar analog modifications are located at the 2' position, other sites, including the 4' position, are also amenable to modification. In certain embodiments, gRNAs contain 4'-S, 4'-Se, or 4'-C-aminomethyl-2'-O-Me modifications.

[0200] In certain embodiments, deazanucleotides, such as 7-deaza-adenosine, may be incorporated into the gRNA. In certain embodiments, O- and N-alkylated nucleotides, such as N6-methyladenosine, may be incorporated into the gRNA. In certain embodiments, one or more or all of the nucleotides in the gRNA are deoxyribonucleotides.

[0201] In certain embodiments, the gRNA will comprise one or more linkers and / or processes for gRNA synthesis selected from those described in International Patent Application having Serial No. PCT / US17 / 69019, the entire contents of which are incorporated herein by reference.

[0202] RNA-guided nuclease RNA-guided nucleases according to the present disclosure include, but are not limited to, naturally occurring class 2 CRISPR nucleases such as Cpf1, as well as other nucleases derived from or obtained therefrom, such as mutants. RNA-guided nucleases can also be defined functionally. For example, an RNA-guided nuclease can be defined as a nuclease that (a) interacts with (e.g., complexes with) a gRNA; and (b) together with the gRNA, binds to and optionally cleaves or modifies a target region of DNA that contains (i) a sequence complementary to the targeting domain of the gRNA, and optionally (ii) an additional sequence referred to as a "protospacer adjacent motif" or "PAM," as described in more detail below. As the following examples illustrate, RNA-guided nucleases can be broadly defined by their PAM specificity and cleavage activity, even though variations may exist between individual RNA-guided nucleases that share the same PAM specificity or cleavage activity. Those skilled in the art will appreciate that some aspects of the present disclosure relate to systems, methods, and compositions that can be implemented using any suitable RNA-guided nuclease with a particular PAM specificity and / or cleavage activity. For this reason, unless otherwise specified, the term RNA-guided nuclease should be understood as a generic term and is not limited to any particular type (e.g., Cpfl versus Cas9), species (e.g., S. aureus versus S. pyogenes), or variation (e.g., truncated or split versus full-length; engineered versus native PAM specificity, etc.) of RNA-guided nuclease.

[0203] The PAM sequence takes its name from its sequence relationship to a "protospacer" sequence that is complementary to the gRNA targeting domain (or "spacer"). Together with the protospacer sequence, the PAM sequence defines the target region or sequence for a particular RNA-guided nuclease / gRNA combination.

[0204] Various RNA-guided nucleases may require different sequence relationships between the PAM and the protospacer. Generally, Cas9 recognizes the PAM sequence 3' of the protospacer, while Cpfl generally recognizes the PAM sequence 5' of the protospacer.

[0205] In addition to recognizing specific sequence orientations of the PAM and protospacer, RNA-guided nucleases can also recognize specific PAM sequences. For example, S. aureus Cas9 recognizes the NNGRRT or NNGRRV PAM sequence, in which the N residue is adjacent 3' to the region recognized by the gRNA targeting domain. S. pyogenes Cas9 recognizes the NGG PAM sequence. F. novicida Cpfl recognizes the TTN PAM sequence. PAM sequences have been identified for various RNA-guided nucleases, and a strategy for identifying novel PAM sequences has been described by Shmakov et al., 2015, Molecular Cell 60, 385-397, November 5, 2015. It should also be noted that an engineered RNA-guided nuclease may have a PAM specificity that differs from that of the reference molecule (e.g., in the case of an engineered RNA-guided nuclease, the reference molecule may be a naturally occurring variant from which the RNA-guided nuclease is derived, or may be a naturally occurring variant with the greatest amino acid sequence homology with the engineered RNA-guided nuclease).

[0206] In addition to their PAM specificity, RNA-guided nucleases can be characterized by their DNA cleavage activity; while natural RNA-guided nucleases typically form DSBs in target nucleic acids, engineered mutants have been generated that produce only SSBs or do not cleave at all (discussed above; Ran & Hsu, et al., Cell 154(6), 1380-1389, September 12, 2013 (Ran), incorporated herein by reference).

[0207] Cpf1 The crystal structure of Acidaminococcus sp. Cpf1 complexed with crRNA and a double-stranded (ds) DNA target, such as the TTTN PAM sequence, has been analyzed by Yamano et al. (Cell. 2016 May 5;165(4):949-962 (Yamano), incorporated herein by reference). Similar to Cas9, Cpf1 has two lobes: the REC (recognition) lobe and the NUC (nuclease) lobe. The REC lobe contains the REC1 and REC2 domains, which bear no similarity to any known protein structure. Meanwhile, the NUC lobe contains three RuvC domains (RuvC-I, -II, and -III) and one BH domain. However, in contrast to Cas9, the Cpf1 REC lobe lacks an HNH domain and contains a structurally unique PI domain, three wedge (WED) domains (WED-I, -II, and -III), and a nuclease (Nuc) domain, among other domains that lack similarity to known protein structures.

[0208] While Cas9 and Cpf1 share similarities in structure and function, it should be understood that certain Cpf1 activities are mediated by structural domains that are not similar to either Cas9 domain. For example, cleavage of the complementary strand of target DNA appears to be mediated by the Nuc domain, which is sequence- and spatially distinct from the HNH domain of Cas9. Furthermore, the non-targeting portion (handle) of the Cpf1 gRNA adopts a pseudoknot structure rather than the stem-loop structure formed by the repeat:antirepeat duplex in the Cas9 gRNA.

[0209] Modification of RNA-guided nucleases While the RNA-guided nucleases described above have activities and properties that may be useful in a variety of applications, one of skill in the art will recognize that RNA-guided nucleases may also be optionally modified to alter cleavage activity, PAM specificity, or other structural or functional characteristics.

[0210] Turning first to modifications that alter cleavage activity, mutations that reduce or eliminate the activity of domains within the NUC lobe are described above. Exemplary mutations that can be made in the RuvC domain, the Cas9 HNH domain, or the Cpf1 Nuc domain are described in Ran and Yamano and Cotta-Ramusino. Generally, mutations that reduce or eliminate the activity of one of the two nuclease domains result in an RNA-guided nuclease with nickase activity, although it should be noted that the type of nickase activity varies depending on which domain is inactivated. For example, inactivation of the RuvC domain of Cas9 results in a nickase that cleaves the complementary or upper strand. On the other hand, inactivation of the Cas9 HNH domain results in a nickase that cleaves the lower or non-complementary strand.

[0211] Modifications of PAM specificity relative to a naturally occurring Cas9 reference molecule have been described by Kleinstiver et al. for both S. pyogenes (Kleinstiver et al., Nature. 2015 Jul 23;523(7561):481-5 (Kleinstiver I) and S. aureus (Kleinstiver et al., Nat Biotechnol. 2015 Dec;33(12):1293-1298 (Kleinstiver II)). Kleinstiver et al. also described modifications that improve the targeting fidelity of Cas9 (Nature, 2016 January 28;529,490-495 (Kleinstiver III)). Modifications of PAM specificity relative to a naturally occurring Cas9 reference molecule have been described by Kleinstiver et al. and by Kleinstiver et al. for S. pyogenes (Kleinstiver et al., Nature. 2015 Jul 23;523(7561):481-5 (Kleinstiver I)). Each of these references is incorporated herein by reference.

[0212] Modifications of PAM specificity compared to a naturally occurring Cpf1 reference molecule have been described by Gao et al. (Gao et al., Nat Biotechnol. 2017 Aug;35(8):789-792, incorporated herein by reference). In certain embodiments, the RNA-guided nuclease can be a Cpf1 variant, such as an AsCpf1 variant. In certain embodiments, the Cpf1 variant is an AsCpf1 variant comprising an S542R / K607R mutation, which recognizes the TYCV PAM. In certain embodiments, the Cpf1 variant is an AsCpf1 variant comprising an S542R / K548V / N552R mutation, which recognizes the TATV PAM.

[0213] RNA-guided nucleases are split into two or more parts, as described by Zetsche et al. (Nat Biotechnol. 2015 Feb;33(2):139-42 (Zetsche II), incorporated by reference) and Fine et al. (Sci Rep. 2015 Jul. 1;5:10777 (Fine), incorporated by reference).

[0214] In certain embodiments, the RNA-guided nuclease may be size-optimized or truncated, for example, through one or more deletions that reduce the size of the nuclease while still maintaining gRNA binding, target and PAM recognition, and cleavage activity. In certain embodiments, the RNA-guided nuclease is covalently or non-covalently linked to another polypeptide, nucleotide, or other structure, optionally via a linker. Exemplary linked nucleases and linkers are described by Guilinger et al., Nature Biotechnology 32, 577-582 (2014), which is incorporated by reference for all purposes herein.

[0215] The RNA-guided nuclease optionally also includes a label, including but not limited to, a nuclear localization signal, to facilitate translocation of the RNA-guided nuclease protein into the nucleus. In certain embodiments, the RNA-guided nuclease may incorporate a C-terminal and / or N-terminal nuclear localization signal. Nuclear localization sequences are known in the art and are described in Maeder and elsewhere.

[0216] The foregoing list of modifications is intended to be exemplary in nature, and those skilled in the art will understand in light of the present disclosure that other modifications may be possible or desirable in particular applications. Thus, for the sake of brevity, the exemplary systems, methods, and compositions of the present disclosure are presented with reference to specific RNA-guided nucleases, but it should be understood that the RNA-guided nucleases used can be modified in ways that do not alter their principles of operation. Such modifications are within the scope of the present disclosure.

[0217] Nucleic acid encoding an RNA-guided nuclease For example, provided herein are nucleic acids encoding RNA-guided nucleases, such as Cpf1 or functional fragments thereof. Exemplary nucleic acids encoding RNA-guided nucleases have been previously described (see, e.g., Cong 2013; Wang 2013; Mali 2013; Jinek 2012).

[0218] In some cases, the nucleic acid encoding the RNA-guided nuclease may be a synthetic nucleic acid sequence. For example, the synthetic nucleic acid molecule may be chemically modified. In certain embodiments, the mRNA encoding the RNA-guided nuclease may have one or more (e.g., all) of the following properties: it may be capped; it may be polyadenylated; or it may be substituted with 5-methylcytidine and / or pseudouridine.

[0219] Synthetic nucleic acid sequences can also be codon-optimized, e.g., at least one uncommon or less common codon is replaced with a common codon. For example, the synthetic nucleic acid can direct the synthesis of an optimized messenger mRNA, e.g., optimized for expression in a mammalian expression system as described herein. Examples of codon-optimized Cas9 coding sequences are provided in Cotta-Ramusino.

[0220] Additionally or alternatively, the nucleic acid encoding the RNA-guided nuclease may comprise a nuclear localization sequence (NLS). Nuclear localization sequences are known in the art.

[0221] Functional analysis of candidate molecules Candidate RNA-guided nucleases, gRNAs, and their complexes can be evaluated by standard methods known in the art. See, e.g., Cotta-Ramusino et al. The stability of RNP complexes can be evaluated by differential scanning fluorimetry, as described below.

[0222] Differential Scanning Fluorometry (DSF) The thermal stability of ribonucleoprotein (RNP) complexes containing gRNA and RNA-guided nucleases can be measured by DSF. The DSF technique measures the thermal stability of proteins, which can increase under favorable conditions, such as the addition of a binding RNA molecule, such as gRNA.

[0223] The DSF assay can be performed according to any suitable protocol and can be used in any suitable setting, including, but not limited to, (a) testing different conditions (e.g., different stoichiometric ratios of gRNA:RNA-guided nuclease protein, different buffers, etc.) to identify optimal conditions for RNP formation; and (b) testing modifications (e.g., chemical modifications, sequence alterations, etc.) of the RNA-guided nuclease and / or gRNA to identify modifications that improve RNP formation or stability. One readout of the DSF assay is the change in melting temperature of the RNP complex; a relatively high change suggests that the RNP complex is more stable (and therefore may have greater activity or more favorable formation kinetics, disassembly kinetics, or another functional property) compared to a standard RNP complex characterized by a lower change. When the DSF assay is deployed as a screening tool, a threshold melting temperature change can be identified, whereby the result is one or more RNPs with a melting temperature change above the threshold. For example, the threshold can be 5-10°C (e.g., 5°C, 6°C, 7°C, 8°C, 9°C, 10°C) or more, and the result can be one or more RNPs characterized by a melting temperature shift at or above the threshold.

[0224] Two non-limiting examples of DSF assay conditions are as follows (conditions refer to the use of Cas9, while similar conditions can be used for Cpf1):

[0225] To determine the best solution for RNP complex formation, a fixed concentration (e.g., 2 μM) of Cas9 + 10× SYPRO Orange® (Life Technologies, Catalog No. S-6650) in water is dispensed into a 384-well plate. Equimolar amounts of gRNA diluted in solutions with various pH and salt concentrations are then added. After 10 minutes of incubation at room temperature and brief centrifugation to remove any air bubbles, a gradient is run from 20°C to 90°C in 1°C increments every 10 seconds using a Bio-Rad CFX384™ Real-Time System C1000 Touch™ thermal cycler with Bio-Rad CFX Manager software.

[0226] The second assay consisted of mixing various concentrations of gRNA with a fixed concentration (e.g., 2 μM) of Cas9i in the optimal buffer from assay 1 above and incubating in a 384-well plate (e.g., at room temperature for 10 minutes). An equal volume of optimal buffer + 10x SYPRO Orange® (Life Technologies, catalog number S-6650) was added, and the plate was sealed with Microseal® B adhesive (MSB-1001). After a brief centrifugation to remove any air bubbles, a gradient was run from 20°C to 90°C in 1°C increments every 10 seconds using a Bio-Rad CFX384™ Real-Time System C1000 Touch™ thermal cycler with Bio-Rad CFX Manager software.

[0227] Genome editing strategies The genome editing systems described above are used in various embodiments of the present disclosure to edit (i.e., modify) target regions of DNA within a cell or obtained from a cell. Various strategies for making specific edits are described herein, and these strategies are generally described by the desired repair outcome, the number and location of individual edits (e.g., single-stranded subunits (SSBs) or double-stranded subunits (DSBs)), and the target sites of such edits.

[0228] Genome editing strategies involving the formation of SSBs or DSBs are characterized by repair outcomes, including (a) deletion of all or part of the target region; (b) insertion into or replacement of all or part of the target region; or (c) disruption of all or part of the target region. This grouping is not intended to limit or be bound by any particular theory or model, and is provided solely for economy of presentation. Those skilled in the art will understand that the listed outcomes are not mutually exclusive, and that some repairs may lead to other outcomes. Description of a particular editing strategy or method should not be construed as requiring a particular repair outcome unless otherwise specified.

[0229] Replacement of the target region generally involves the replacement of all or part of the sequence present within the target region with a homologous sequence, for example, through gene correction or gene conversion, two repair outcomes mediated by the HDR pathway. HDR is facilitated by the use of a donor template, which can be single-stranded or double-stranded, as described in more detail below. The single-stranded or double-stranded template can be exogenous, in which case they facilitate gene correction, or they can be endogenous (e.g., homologous sequences in the cellular genome) and facilitate gene conversion. The exogenous template can have an asymmetric overhang (i.e., the portion of the template complementary to the site of the DSB can be offset in the 3' or 5' direction, rather than being centrally located within the template), as described, for example, by Richardson et al. (Nature Biotechnology 34, 339-344 (2016), (Richardson), incorporated by reference). If the template is single-stranded, it can correspond to either the complementary (top) or non-complementary (bottom) strand of the target region.

[0230] As described in Ran and Cotta-Ramusino et al., gene conversion and gene correction are optionally facilitated by creating one or more nicks within or surrounding the target region. Optionally, a double nickase strategy is used to create two offset single stranded nucleotide subunits (SSBs), followed by a single double stranded nucleotide subunit (DSB) with an overhang (e.g., a 5' overhang).

[0231] The interruption and / or deletion of all or part of the target sequence can be achieved by various repair outcomes. As one example, as described by Maeder for LCA10 mutations, the sequence can be deleted by simultaneously generating two or more DSBs flanking the target region, which are then excised when the DSBs are repaired. As another example, the sequence can be interrupted by the formation of a double-stranded break with a single-stranded overhang, followed by a deletion generated by nucleotide terminal hydrolysis of the overhang before repair.

[0232] One particular subset of target sequence interruptions is mediated by the formation of indels in the target sequence, in which repair typically occurs via the NHEJ pathway (including Alt-NHEJ). NHEJ is referred to as an "error-prone" repair pathway due to its association with indel mutations. However, in some cases, DSBs are repaired by NHEJ without altering the surrounding sequence (so-called "perfect" or "scarless" repair); this generally requires complete ligation of both ends of the DSB. On the other hand, indels are thought to result from enzymatic processing of free DNA ends before they are ligated, which adds and / or removes nucleotides from one or both strands of one or both free ends.

[0233] Because enzymatic processing of free DSB ends can be stochastic in nature, indel mutations tend to be variable, occur along a distribution, and can be influenced by a variety of factors, including the specific target site, the cell type used, and the genome editing strategy employed. Nevertheless, limited generalizations about indel formation can be made: deletions formed by repair of a single DSB are most commonly in the 1-50 bp range, but can exceed 100-200 bp. Insertions formed by repair of a single DSB tend to be shorter, often containing short duplications of the sequence immediately surrounding the break site. However, larger insertions can be obtained, and in these cases, the inserted sequence is often traced to other regions of the genome or to plasmid DNA present in the cell.

[0234] Indel mutations and genome editing systems configured to generate indels are useful for interrupting target sequences, for example, when the generation of a specific final sequence is not required and / or frameshift mutations are tolerated. They can also be useful in situations where a specific sequence is preferred, as long as the desired specific sequence tends to preferentially result from repair of an SSB or DSB at a given site. Indel mutations are also useful tools for evaluating or screening the activity of specific genome editing systems and their components. In these and other settings, indels can be characterized by (a) their relative and absolute frequency in the genome of cells contacted with the genome editing system and (b) the distribution of numerical differences, e.g., ±1, ±2, ±3, relative to the unedited sequence. As an example, in a lead discovery setting, multiple gRNAs can be screened to identify the gRNA that most efficiently drives cleavage at the target site based on indel reads under controlled conditions. Guides that generate indels at or above a threshold frequency or a specific distribution of indels can be selected for further research and development. The frequency and distribution of indels can also be useful as a readout to evaluate different genome editing system implementations or formulations and delivery methods, for example, by keeping the gRNA constant and varying other specific reaction conditions or delivery methods.

[0235] Multiple Strategies While the above exemplary strategies focus on the repair results mediated by a single DSB, two or more DSBs can be generated at either the same or different loci using a genome editing system according to the present disclosure. Editing strategies involving the formation of multiple DSBs or SSBs are described, for example, in Cotta-Ramusino et al. As described herein, the methods and compositions encompassed by the present disclosure can be used to affect T cell proliferation, survival, persistence, and / or function by modifying two or more T cell-expressed genes, such as, for example, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more FAS, BID, CTLA4, PDCD1, CBLB, PTPN6, B2M, TRAC, CIITA, and TRBC genes.

[0236] Donor Mold Design Donor template design is described in detail in, for example, Cotta-Ramusino et al. DNA oligomer donor templates (oligodeoxynucleotides or ODNs), which can be single-stranded (ssODN) or double-stranded (dsODN), can be used to facilitate HDR-based repair of DSBs, which are particularly useful for introducing modifications into target DNA sequences, inserting new sequences into target sequences, or completely replacing target sequences.

[0237] Whether single-stranded or double-stranded, the donor template generally contains regions that are homologous to regions of DNA in or near (e.g., flanking or adjacent to) the target sequence to be cleaved. These regions of homology are referred to herein as "homology arms" and are shown schematically below. [5' homology arm]--[substitution sequence]--[3' homology arm]

[0238] The homology arms can have any suitable length (including no nucleotides when only one homology arm is used), and the 3' and 5' homology arms can have the same length or can be different lengths. The selection of an appropriate homology arm length can be influenced by various factors, such as the desire to avoid homology or microhomology with specific sequences, such as Alu repeat sequences or other highly common elements. For example, the 5' homology arm can be shortened to avoid sequence repeat elements. In other embodiments, the 3' homology arm can be shortened to avoid sequence repeat elements. In certain embodiments, both the 5' and 3' homology arms can be shortened to avoid the inclusion of specific sequence repeat elements. Furthermore, some homology arm designs can improve editing efficiency or increase the frequency of the desired repair outcome. For example, Richardson et al. Nature Biotechnology 34, 339-344 (2016) (Richardson), incorporated by reference, found that the relative asymmetry of the 3' and 5' homology arms of a single-stranded donor template affects the rate and / or outcome of repair.

[0239] Replacement sequences in donor templates are described elsewhere, including Cotta-Ramusino et al. Replacement sequences can be of any suitable length (including zero nucleotides when the desired repair outcome is a deletion) and typically contain one, two, three, or more sequence modifications relative to the native sequence in the cell for which editing is desired. One common sequence modification involves altering the native sequence to repair a mutation associated with a disease or condition for which treatment is desired. Another common sequence modification involves altering one or more sequences complementary to or encoding the PAM sequence of an RNA-guided nuclease or the targeting domain of a gRNA used to generate an SSB or DSB, to reduce or eliminate recurrent cleavage at the target site after the replacement sequence is integrated into the target site.

[0240] When a linear ssODN is used, it can be configured to (i) anneal to a nicked strand of a target nucleic acid, (ii) anneal to an intact strand of a target nucleic acid, (iii) anneal to a plus strand of a target nucleic acid, and / or (iv) anneal to a minus strand of a target nucleic acid. The ssODN can have any suitable length, such as, for example, about or at least 150-200 nucleotides or less (e.g., 150, 160, 170, 180, 190, or 200 nucleotides).

[0241] It should be noted that the template nucleic acid can also be a nucleic acid vector such as a viral genome or a circular double-stranded DNA such as a plasmid. The nucleic acid vector containing the donor template can contain other coding or non-coding elements. For example, the template nucleic acid can be delivered as part of a viral genome (e.g., in an AAV or lentiviral genome) that contains specific genomic backbone elements (e.g., inverted terminal repeats in the case of an AAV genome) and, optionally, additional sequences encoding gRNAs and / or RNA-guided nucleases. In certain embodiments, the donor template is flanked by or flanked by target sites recognized by one or more gRNAs, which can facilitate the formation of free DSBs at one or both ends of the donor template, which can participate in the repair of corresponding SSBs or DSBs formed in cellular DNA using the same gRNA. Exemplary nucleic acid vectors suitable for use as donor templates are described in Cotta-Ramusino et al.

[0242] Whatever format is used, the template nucleic acid can be designed to avoid undesired sequences, in certain embodiments, one or both homology arms can be shortened to avoid overlap with certain sequence repeat elements, such as Alu repeats, LINE elements, etc.

[0243] Targeted integration The present disclosure further provides a genome editing system including a donor template specifically designed to enable quantitative evaluation of gene editing events occurring upon cleavage separation events at the cleavage site of a target nucleic acid in a cell. The donor template of the genome editing system described herein is a DNA oligodeoxynucleotide (ODN), which may be single-stranded (ssODN) or double-stranded (dsODN), and can be used to promote HDR-based repair of double-strand breaks. The donor template is particularly useful for introducing modifications into a target DNA sequence, inserting a new sequence into a target sequence, or replacing the entire target sequence. The present disclosure provides a donor template comprising a cargo, one or two homology arms, and one or more priming sites. The priming sites of the donor template are spatially arranged in such a manner that the frequency of incorporation of a portion of the donor template into the target nucleic acid can be easily assessed and quantified.

[0244] 44A, 44B, and 44C are schematic diagrams illustrating representative donor templates and potential targeted integration outcomes resulting from the use of these donor templates. Use of the exemplary donor templates described herein results in targeted integration of at least one priming site in a targeting nucleic acid, which can be used to generate amplicons that can be sequenced to determine the frequency of targeted integration of a cargo (e.g., a transgene) into the targeting nucleic acid in a target cell.

[0245] For example, Figure 44A illustrates an exemplary donor template comprising, from 5' to 3', a first homology arm (A1), a first stuffer sequence (S1), a second priming site (P2'), cargo, a first priming site, a second stuffer sequence, and a second homology arm. The first homology arm (A1) of the donor template is substantially identical to the first homology arm of the target nucleic acid, while the second homology arm (A2) of the donor template is substantially identical to the second homology arm of the target nucleic acid. The donor template is designed so that the second priming site (P2') is substantially identical to the first priming site (P1) of the target nucleic acid, and so that the first priming site (P1') is substantially identical to the second priming site (P2) of the target nucleic acid. During a cleavage separation event of a target nucleic acid using a nuclease described herein, a single primer pair set can be used to amplify the nucleic acid sequence surrounding the cleavage site of the target nucleic acid (i.e., the nucleic acid present between P1 and P2, between P1 and P2', and between P1' and P2). Advantageously, the sizes of the amplicons (denoted as amplicons X, Y, and Z) obtained from a cleavage separation event without or with targeted integration are approximately the same. The amplicons can then be evaluated, for example, by sequencing or hybridization to a probe sequence, to determine the frequency of targeted integration.

[0246] Instead, Figures 44B and 44C depict exemplary donor templates containing a single priming site located either 3' (Figure 44B) or 5' (Figure 44C) from the cargo nucleic acid sequence. Again, upon cleavage separation of the target nucleic acid using a nuclease described herein, these exemplary donor templates are designed such that a single primer pair can be used to amplify the nucleic acid sequence surrounding the cleavage site of the target nucleic acid, resulting in two amplicons of approximately the same size. When the priming site of the donor template is located 3' from the cargo nucleic acid, an amplicon corresponding to a non-targeted integration event or an amplicon corresponding to the 5' junction of the targeted integration site can be amplified. When the priming site of the donor template is located 5' from the cargo nucleic acid, an amplicon corresponding to a non-targeted integration event or an amplicon corresponding to the 3' junction of the targeted integration site can be amplified. These amplicons can be sequenced to determine the frequency of targeted integration.

[0247] A donor template according to the present disclosure can be implemented in any suitable manner, including, without limitation, linear or circular, naked single-stranded or double-stranded DNA, or can be contained within a vector, and / or can be covalently or non-covalently linked to a guide RNA (e.g., by direct hybridization or splint hybridization). In certain embodiments, the donor template is a ssODN. When a linear ssODN is used, it can be configured to (i) anneal to a nicked strand of a target nucleic acid, (ii) anneal to an intact strand of a target nucleic acid, (iii) anneal to a plus strand of a target nucleic acid, and / or (iv) anneal to a minus strand of a target nucleic acid. The ssODN can have any suitable length, such as, for example, about 150-200 nucleotides or less (e.g., 150, 160, 170, 180, 190, or 200 nucleotides). In other embodiments, the donor template is a dsODN. In certain embodiments, the donor template comprises a first strand. In other embodiments, the donor template comprises a first strand and a second strand. In certain embodiments, the donor template is an exogenous oligonucleotide, e.g., an oligonucleotide that does not naturally occur in a cell.

[0248] It should be noted that the donor template can also be contained within a nucleic acid vector, such as a viral genome or a circular double-stranded DNA, e.g., a plasmid. In certain embodiments, the donor template can be a dog-bone-shaped DNA (see, e.g., U.S. Pat. No. 9,499,847). The nucleic acid vector containing the donor template can contain other coding or non-coding elements. For example, the donor template nucleic acid can be delivered as part of a viral genome (e.g., in an AAV or lentiviral genome) that contains specific genomic backbone elements (e.g., inverted terminal repeats in the case of an AAV genome) and, optionally, additional sequences encoding gRNAs and / or RNA-guided nucleases. In certain embodiments, the donor template can be flanked by or flanked by target sites recognized by one or more gRNAs, facilitating the formation of free DSBs at one or both ends of the donor template, which can participate in the repair of corresponding SSBs or DSBs formed in cellular DNA using the same gRNA. Exemplary nucleic acid vectors suitable for use as donor templates are described in Cotta-Ramusino et al.

[0249] homology arm Whether single-stranded or double-stranded, the donor template generally contains one or more regions homologous to a region of DNA, e.g., a target nucleic acid, within or near (e.g., flanking or adjacent to) the target sequence to be cleaved, e.g., the cleavage site. These regions of homology are referred to herein as "homology arms" and are shown schematically below. [5' homology arm]-[substitution sequence]-[3' homology arm]

[0250] The homology arm of the donor template described herein can be any suitable length, as long as it is long enough to allow the DNA repair process that requires donor template to efficiently separate the cleavage site of target nucleic acid.For example, in certain embodiments where PCR is desired to amplify homology arm, the homology arm is long enough to allow amplification to be carried out.In certain embodiments where homology arm sequencing is desired, the homology arm is long enough to allow sequencing to be carried out.In certain embodiments where quantitative evaluation of amplicons is desired, the homology arm is long enough to achieve similar amplification times for each amplicon, for example, by having similar G / C content, amplification temperature, etc.In certain embodiments, the homology arm is double-stranded.In certain embodiments, the double-stranded homology arm is single-stranded.

[0251] In certain embodiments, the 5' homology arm is 50 to 250 nucleotides in length. In certain embodiments, the 5' homology arm is 50 to 2000 nucleotides in length. In certain embodiments, the 5' homology arm is 50 to 1500 nucleotides in length. In certain embodiments, the 5' homology arm is 50 to 1000 nucleotides in length. In certain embodiments, the 5' homology arm is 50 to 500 nucleotides in length. In certain embodiments, the 5' homology arm is 150 to 250 nucleotides in length. In certain embodiments, the 5' homology arm is 2000 nucleotides or less in length. In certain embodiments, the 5' homology arm is 1500 nucleotides or less in length. In certain embodiments, the 5' homology arm is 1000 nucleotides or less in length. In certain embodiments, the 5' homology arm is 700 nucleotides or less in length. In certain embodiments, the 5' homology arm is 650 nucleotides or less in length. In certain embodiments, the 5' homology arm is 600 nucleotides or less in length. In certain embodiments, the 5' homology arm is 550 nucleotides or less in length. In certain embodiments, the 5' homology arm is 500 nucleotides or less in length. In certain embodiments, the 5' homology arm is 400 nucleotides or less in length. In certain embodiments, the 5' homology arm is 300 nucleotides or less in length. In certain embodiments, the 5' homology arm is 250 nucleotides or less in length. In certain embodiments, the 5' homology arm is 200 nucleotides or less in length. In certain embodiments, the 5' homology arm is 150 nucleotides or less in length. In certain embodiments, the 5' homology arm is less than 100 nucleotides in length. In certain embodiments, the 5' homology arm is 50 nucleotides or less in length. In certain embodiments, the 5' homology arm is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21 or 20 nucleotides in length.In certain embodiments, the 5' homology arm is at least 20 nucleotides in length. In certain embodiments, the 5' homology arm is at least 40 nucleotides in length. In certain embodiments, the 5' homology arm is at least 50 nucleotides in length. In certain embodiments, the 5' homology arm is at least 70 nucleotides in length. In certain embodiments, the 5' homology arm is at least 100 nucleotides in length. In certain embodiments, the 5' homology arm is at least 200 nucleotides in length. In certain embodiments, the 5' homology arm is at least 300 nucleotides in length. In certain embodiments, the 5' homology arm is at least 400 nucleotides in length. In certain embodiments, the 5' homology arm is at least 500 nucleotides in length. In certain embodiments, the 5' homology arm is at least 600 nucleotides in length. In certain embodiments, the 5' homology arm is at least 700 nucleotides in length. In certain embodiments, the 5' homology arm is at least 1000 nucleotides in length. In certain embodiments, the 5' homology arm is at least 1500 nucleotides in length. In certain embodiments, the 5' homology arm is at least 2000 nucleotides in length. In certain embodiments, the 5' homology arm is about 20 nucleotides in length. In certain embodiments, the 5' homology arm is about 40 nucleotides in length. In certain embodiments, the 5' homology arm is 250 nucleotides or less in length. In certain embodiments, the 5' homology arm is about 100 nucleotides in length. In certain embodiments, the 5' homology arm is about 200 nucleotides in length.

[0252] In certain embodiments, the 3' homology arm is 50 to 250 nucleotides in length. In certain embodiments, the 3' homology arm is 50 to 2000 nucleotides in length. In certain embodiments, the 3' homology arm is 50 to 1500 nucleotides in length. In certain embodiments, the 3' homology arm is 50 to 1000 nucleotides in length. In certain embodiments, the 3' homology arm is 50 to 500 nucleotides in length. In certain embodiments, the 3' homology arm is 150 to 250 nucleotides in length. In certain embodiments, the 3' homology arm is 2000 nucleotides or less in length. In certain embodiments, the 3' homology arm is 1500 nucleotides or less in length. In certain embodiments, the 3' homology arm is 1000 nucleotides or less in length. In certain embodiments, the 3' homology arm is 700 nucleotides or less in length. In certain embodiments, the 3' homology arm is 650 nucleotides or less in length. In certain embodiments, the 3' homology arm is 600 nucleotides or less in length. In certain embodiments, the 3' homology arm is 550 nucleotides or less in length. In certain embodiments, the 3' homology arm is 500 nucleotides or less in length. In certain embodiments, the 3' homology arm is 400 nucleotides or less in length. In certain embodiments, the 3' homology arm is 300 nucleotides or less in length. In certain embodiments, the 3' homology arm is 200 nucleotides or less in length. In certain embodiments, the 3' homology arm is 150 nucleotides or less in length. In certain embodiments, the 3' homology arm is 100 nucleotides or less in length. In certain embodiments, the 3' homology arm is 50 nucleotides or less in length. In certain embodiments, the 3' homology arm is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 nucleotides in length. In certain embodiments, the 3' homology arm is at least 20 nucleotides in length.In certain embodiments, the 3' homology arm is at least 40 nucleotides in length. In certain embodiments, the 3' homology arm is at least 50 nucleotides in length. In certain embodiments, the 3' homology arm is at least 70 nucleotides in length. In certain embodiments, the 3' homology arm is at least 100 nucleotides in length. In certain embodiments, the 3' homology arm is at least 200 nucleotides in length. In certain embodiments, the 3' homology arm is at least 300 nucleotides in length. In certain embodiments, the 3' homology arm is at least 400 nucleotides in length. In certain embodiments, the 3' homology arm is at least 500 nucleotides in length. In certain embodiments, the 3' homology arm is at least 600 nucleotides in length. In certain embodiments, the 3' homology arm is at least 700 nucleotides in length. In certain embodiments, the 3' homology arm is at least 1000 nucleotides in length. In certain embodiments, the 3' homology arm is at least 1500 nucleotides in length. In certain embodiments, the 3' homology arm is at least 2000 nucleotides in length. In certain embodiments, the 3' homology arm is about 20 nucleotides in length. In certain embodiments, the 3' homology arm is about 40 nucleotides in length. In certain embodiments, the 3' homology arm is 250 nucleotides or less in length. In certain embodiments, the 3' homology arm is about 100 nucleotides in length. In certain embodiments, the 3' homology arm is about 200 nucleotides in length.

[0253] In certain embodiments, the 5' homology arm is 50 to 250 base pairs in length. In certain embodiments, the 5' homology arm is 50 to 2000 base pairs in length. In certain embodiments, the 5' homology arm is 50 to 1500 base pairs in length. In certain embodiments, the 5' homology arm is 50 to 1000 base pairs in length. In certain embodiments, the 5' homology arm is 50 to 500 base pairs in length. In certain embodiments, the 5' homology arm is 150 to 250 base pairs in length. In certain embodiments, the 5' homology arm is 2000 base pairs or less in length. In certain embodiments, the 5' homology arm is 1500 base pairs or less in length. In certain embodiments, the 5' homology arm is 1000 base pairs or less in length. In certain embodiments, the 5' homology arm is 700 base pairs or less in length. In certain embodiments, the 5' homology arm is 650 base pairs or less in length. In certain embodiments, the 5' homology arm is 600 base pairs or less in length. In certain embodiments, the 5' homology arm is 550 base pairs or less in length. In certain embodiments, the 5' homology arm is 500 base pairs or less in length. In certain embodiments, the 5' homology arm is 400 base pairs or less in length. In certain embodiments, the 5' homology arm is 300 base pairs or less in length. In certain embodiments, the 5' homology arm is 250 base pairs or less in length. In certain embodiments, the 5' homology arm is 200 base pairs or less in length. In certain embodiments, the 5' homology arm is 150 base pairs or less in length. In certain embodiments, the 5' homology arm is less than 100 base pairs in length. In certain embodiments, the 5' homology arm is 50 base pairs or less in length. In certain embodiments, the 5' homology arm is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 base pairs in length. In certain embodiments, the 5' homology arm is at least 20 base pairs in length.In certain embodiments, the 5' homology arm is at least 40 base pairs in length. In certain embodiments, the 5' homology arm is at least 50 base pairs in length. In certain embodiments, the 5' homology arm is at least 70 base pairs in length. In certain embodiments, the 5' homology arm is at least 100 base pairs in length. In certain embodiments, the 5' homology arm is at least 200 base pairs in length. In certain embodiments, the 5' homology arm is at least 300 base pairs in length. In certain embodiments, the 5' homology arm is at least 400 base pairs in length. In certain embodiments, the 5' homology arm is at least 500 base pairs in length. In certain embodiments, the 5' homology arm is at least 600 base pairs in length. In certain embodiments, the 5' homology arm is at least 700 base pairs in length. In certain embodiments, the 5' homology arm is at least 1000 base pairs in length. In certain embodiments, the 5' homology arm is at least 1500 base pairs in length. In certain embodiments, the 5' homology arm is at least 2000 base pairs in length. In certain embodiments, the 5' homology arm is about 20 base pairs in length. In certain embodiments, the 5' homology arm is about 40 base pairs in length. In certain embodiments, the 5' homology arm is 250 base pairs or less in length. In certain embodiments, the 5' homology arm is about 100 base pairs in length. In certain embodiments, the 5' homology arm is about 200 base pairs in length.

[0254] In certain embodiments, the 3' homology arm is 50 to 250 base pairs in length. In certain embodiments, the 3' homology arm is 50 to 2000 base pairs in length. In certain embodiments, the 3' homology arm is 50 to 1500 base pairs in length. In certain embodiments, the 3' homology arm is 50 to 1000 base pairs in length. In certain embodiments, the 3' homology arm is 50 to 500 base pairs in length. In certain embodiments, the 3' homology arm is 150 to 250 base pairs in length. In certain embodiments, the 3' homology arm is 2000 base pairs or less in length. In certain embodiments, the 3' homology arm is 1500 base pairs or less in length. In certain embodiments, the 3' homology arm is 1000 base pairs or less in length. In certain embodiments, the 3' homology arm is 700 base pairs or less in length. In certain embodiments, the 3' homology arm is 650 base pairs or less in length. In certain embodiments, the 3' homology arm is 600 base pairs or less in length. In certain embodiments, the 3' homology arm is 550 base pairs or less in length. In certain embodiments, the 3' homology arm is 500 base pairs or less in length. In certain embodiments, the 3' homology arm is 400 base pairs or less in length. In certain embodiments, the 3' homology arm is 300 base pairs or less in length. In certain embodiments, the 3' homology arm is 250 base pairs or less in length. In certain embodiments, the 3' homology arm is 200 base pairs or less in length. In certain embodiments, the 3' homology arm is 150 base pairs or less in length. In certain embodiments, the 3' homology arm is less than 100 base pairs in length. In certain embodiments, the 3' homology arm is 50 base pairs or less in length. In certain embodiments, the 3' homology arm is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 base pairs in length. In certain embodiments, the 3' homology arm is at least 20 base pairs in length.In certain embodiments, the 3' homology arm is at least 40 base pairs in length. In certain embodiments, the 3' homology arm is at least 50 base pairs in length. In certain embodiments, the 3' homology arm is at least 70 base pairs in length. In certain embodiments, the 3' homology arm is at least 100 base pairs in length. In certain embodiments, the 3' homology arm is at least 200 base pairs in length. In certain embodiments, the 3' homology arm is at least 300 base pairs in length. In certain embodiments, the 3' homology arm is at least 400 base pairs in length. In certain embodiments, the 3' homology arm is at least 500 base pairs in length. In certain embodiments, the 3' homology arm is at least 600 base pairs in length. In certain embodiments, the 3' homology arm is at least 700 base pairs in length. In certain embodiments, the 3' homology arm is at least 1000 base pairs in length. In certain embodiments, the 3' homology arm is at least 1500 base pairs in length. In certain embodiments, the 3' homology arm is at least 2000 base pairs in length. In certain embodiments, the 3' homology arm is about 20 base pairs in length. In certain embodiments, the 3' homology arm is about 40 base pairs in length. In certain embodiments, the 3' homology arm is 250 base pairs or less in length. In certain embodiments, the 3' homology arm is about 100 base pairs in length. In certain embodiments, the 3' homology arm is about 200 base pairs in length. In certain embodiments, the 3' homology arm is 250 base pairs or less in length. In certain embodiments, the 3' homology arm is 200 base pairs or less in length. In certain embodiments, the 3' homology arm is 150 base pairs or less in length. In certain embodiments, the 3' homology arm is 100 base pairs or less in length. In certain embodiments, the 3' homology arm is 50 base pairs or less in length.In certain embodiments, the 3' homology arm is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 base pairs in length. In certain embodiments, the 3' homology arm is 40 base pairs in length.

[0255] The 5' and 3' homology arms can be the same length or can be different lengths. In certain embodiments, the 5' and 3' homology arms are amplified to enable quantitative assessment of gene editing events, such as targeted integration, in the target nucleic acid. In certain embodiments, quantitative assessment of gene editing events can rely on amplifying both the 5' junction and the 3' junction at the targeted integration site by amplifying all or part of the homology arms using a pair of PCR primers in a single amplification reaction. Thus, the lengths of the 5' and 3' homology arms can be different, but the length of each homology arm must be amplifiable (e.g., using PCR) as needed. Furthermore, if both amplification of the 5' homology arm and the difference in length between the 5' and 3' homology arms are desired in a single PCR reaction, the difference in length between the 5' and 3' homology arms must be PCR amplifiable using a pair of PCR primers.

[0256] In certain embodiments, the lengths of the 5' and 3' homology arms do not differ by more than 75 nucleotides. Thus, in certain embodiments, if the 5' and 3' homology arms are different in length, the difference in length between the homology arms is less than 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide or base pair. In certain embodiments, the 5' and 3' homology arms differ in length by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74 or 75 nucleotides. In certain embodiments, the difference in length between the 5' and 3' homology arms is less than 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base pairs. In certain embodiments, the 5' and 3' homology arms differ in length by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74 or 75 base pairs.

[0257] The donor template of the present disclosure is designed to promote homologous recombination with a target nucleic acid having a cleavage site, the target nucleic acid comprising, in a 5' to 3' direction, P1--H1--X--H2--P2; wherein P1 is a first priming site; H1 is a first homology arm; X is a cleavage site; H2 is a second homology arm; P2 is a second priming site; and the donor template comprises, from 5' to 3', A1--P2'--N--A2 or A1--N--P1'--A2; wherein A1 is a homology arm substantially identical to H1; P2' is a priming site substantially identical to P2; N is cargo; P1' is a priming site substantially identical to P1; and A2 is a homology arm substantially identical to H2. In certain embodiments, the target nucleic acid is double-stranded. In certain embodiments, the target nucleic acid comprises a first strand and a second strand. In other embodiments, the target nucleic acid is single-stranded. In certain embodiments, the target nucleic acid comprises a first strand.

[0258] In certain embodiments, the donor template comprises, in the 5' to 3' direction, A1--P2'--N--A2.

[0259] In certain embodiments, the donor template comprises, from 5' to 3', A1--P2'--N--P1'--A2.

[0260] In certain embodiments, the target nucleic acid comprises, in the 5' to 3' direction, P1--H1--X--H2--P2, wherein P1 is a first priming site; H1 is a first homology arm; X is a cleavage site; H2 is a second homology arm; P2 is a second priming site; the first strand of the donor template comprises, from 5' to 3', A1--P2'--N--A2 or A1--N--P1'--A2; wherein A1 is a homology arm substantially identical to H1; P2' is a priming site substantially identical to P2; N is cargo; P1' is a priming site substantially identical to P1; and A2 is a homology arm substantially identical to H2.

[0261] In certain embodiments, the first strand of the donor template comprises, from 5' to 3', A1--P2'--N--P1'--A2.

[0262] In certain embodiments, the first strand of the donor template comprises, from 5' to 3', A1--N--P1'--A2.

[0263] In certain embodiments, A1 is 700 base pairs or less in length. In certain embodiments, A1 is 650 base pairs or less in length. In certain embodiments, A1 is 600 base pairs or less in length. In certain embodiments, A1 is 550 base pairs or less in length. In certain embodiments, A1 is 500 base pairs or less in length. In certain embodiments, A1 is 400 base pairs or less in length. In certain embodiments, A1 is 300 base pairs or less in length. In certain embodiments, A1 is less than 250 base pairs in length. In certain embodiments, A1 is less than 200 base pairs in length. In certain embodiments, A1 is less than 150 base pairs in length. In certain embodiments, A1 is less than 100 base pairs in length. In certain embodiments, A1 is less than 50 base pairs in length. In certain embodiments, A1 is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 base pairs in length. In certain embodiments, A1 is 40 base pairs in length. In certain embodiments, A1 is 30 base pairs in length. In certain embodiments, A1 is 20 base pairs in length.

[0264] In certain embodiments, A2 is 700 base pairs or less in length. In certain embodiments, A2 is 650 base pairs or less in length. In certain embodiments, A2 is 600 base pairs or less in length. In certain embodiments, A2 is 550 base pairs or less in length. In certain embodiments, A2 is 500 base pairs or less in length. In certain embodiments, A2 is 400 base pairs or less in length. In certain embodiments, A2 is 300 base pairs or less in length. In certain embodiments, A2 is less than 250 base pairs in length. In certain embodiments, A2 is less than 200 base pairs in length. In certain embodiments, A2 is less than 150 base pairs in length. In certain embodiments, A2 is less than 100 base pairs in length. In certain embodiments, A2 is less than 50 base pairs in length. In certain embodiments, A2 is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 base pairs in length. In certain embodiments, A2 is 40 base pairs in length. In certain embodiments, A2 is 30 base pairs in length. In certain embodiments, A2 is 20 base pairs in length.

[0265] In certain embodiments, A1 is 700 nucleotides or less in length. In certain embodiments, A1 is 650 nucleotides or less in length. In certain embodiments, A1 is 600 nucleotides or less in length. In certain embodiments, A1 is 550 nucleotides or less in length. In certain embodiments, A1 is 500 nucleotides or less in length. In certain embodiments, A1 is 400 nucleotides or less in length. In certain embodiments, A1 is 300 nucleotides or less in length. In certain embodiments, A1 is less than 250 nucleotides in length. In certain embodiments, A1 is less than 200 nucleotides in length. In certain embodiments, A1 is less than 150 nucleotides in length. In certain embodiments, A1 is less than 100 nucleotides in length. In certain embodiments, A1 is less than 50 nucleotides in length. In certain embodiments, A1 is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 nucleotides in length. In certain embodiments, A1 is at least 40 nucleotides in length. In certain embodiments, A1 is at least 30 nucleotides in length. In certain embodiments, A1 is at least 20 nucleotides in length.

[0266] In certain embodiments, A2 is 700 nucleotides or less in length. In certain embodiments, A2 is 650 base pairs or less in length. In certain embodiments, A2 is 600 nucleotides or less in length. In certain embodiments, A2 is 550 nucleotides or less in length. In certain embodiments, A2 is 500 nucleotides or less in length. In certain embodiments, A2 is 400 nucleotides or less in length. In certain embodiments, A2 is 300 nucleotides or less in length. In certain embodiments, A2 is less than 250 nucleotides in length. In certain embodiments, A2 is less than 200 nucleotides in length. In certain embodiments, A2 is less than 150 nucleotides in length. In certain embodiments, A2 is less than 100 nucleotides in length. In certain embodiments, A2 is less than 50 nucleotides in length. In certain embodiments, A2 is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 nucleotides in length. In certain embodiments, A2 is at least 40 nucleotides in length. In certain embodiments, A2 is at least 30 nucleotides in length. In certain embodiments, A2 is at least 20 nucleotides in length.

[0267] In certain embodiments, the nucleic acid sequence of A1 is substantially identical to the nucleic acid sequence of H1. In certain embodiments, A1 has a sequence that is identical to or differs from H1 by no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides. In certain embodiments, A1 has a sequence that is identical to or differs from H1 by no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 base pairs.

[0268] In certain embodiments, the nucleic acid sequence of A2 is substantially identical to the nucleic acid sequence of H2. In certain embodiments, A2 has a sequence that is identical to or differs from H2 by no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides. In certain embodiments, A2 has a sequence that is identical to or differs from H2 by no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 base pairs.

[0269] Whatever the format, the donor template can be designed to avoid undesired sequences, in certain embodiments, one or both homology arms can be shortened to avoid overlap with certain sequence repeat elements, such as Alu repeats, LINE elements, etc.

[0270] Priming Site The donor templates described herein contain at least one priming site having a sequence substantially similar to or identical to the sequence of a priming site in a target nucleic acid, but in a different spatial order or orientation relative to the homologous sequences / arms in the donor template. When the donor template homologously recombines with the target nucleic acid, the priming site is advantageously incorporated into the target nucleic acid, thereby allowing amplification of the portion of the modified nucleic acid sequence resulting from the recombination event. In certain embodiments, the donor template contains at least one priming site. In certain embodiments, the donor template contains a first and a second priming site. In certain embodiments, the donor template contains three or more priming sites.

[0271] In certain embodiments, the donor template comprises a priming site P1' that is substantially similar to or identical to priming site P in the target nucleic acid, such that P1' is incorporated downstream of P1 when the donor template is incorporated into the target nucleic acid. In certain embodiments, the donor template comprises a first priming site P1' and a second priming site P2'; P1' is substantially similar to or identical to the first priming site P1 in the target nucleic acid; P2' is substantially similar to or identical to the second priming site P2 in the target nucleic acid; and P1 and P2 are not substantially similar or identical. In certain embodiments, the donor template comprises a first priming site P1' and a second priming site P2'; P1' is substantially similar to or identical to the first priming site P1 within the target nucleic acid; P2' is substantially similar to or identical to the second priming site P2 within the target nucleic acid; P2 is located downstream of P1 on the target nucleic acid; P1 and P2 are not substantially similar or identical; when the donor template is incorporated into the target nucleic acid, P1' is incorporated downstream of P1. P2' is incorporated upstream of P2, and P2' is incorporated upstream of P1.

[0272] In certain embodiments, the target nucleic acid comprises a first priming site (P1) and a second priming site (P2). The first priming site of the target nucleic acid can be within the first homology arm. Alternatively, the first priming site of the target nucleic acid can be 5' and adjacent to the first homology arm. The second priming site of the target nucleic acid can be within the second homology arm. Alternatively, the second priming site of the target nucleic acid can be 3' and adjacent to the second homology arm.

[0273] The donor template may comprise a cargo sequence, a first priming site (P1'), and a second priming site (P2'), where P2' is located 5' of the cargo sequence and P1' is located 3' of the cargo sequence (i.e., A1--P2'--N--P1'--A2), where P1' is substantially identical to P1, and P2' is substantially identical to P2. In this scenario, a primer pair comprising oligonucleotides targeting P1' and P1 and an oligonucleotide comprising P2' and P2 may be used to amplify the targeted locus, thereby generating three amplicons of similar size, which may be sequenced to determine whether targeted integration occurred. A first amplicon, amplicon X, results from amplification of the nucleic acid sequence between P1 and P2 as a result of non-targeted integration in the target nucleic acid. The second amplicon, amplicon Y, is generated by amplification of the nucleic acid sequence between P1 and P2' after the targeted incorporation event of the target nucleic acid, thereby amplifying the 5' junction. The third amplicon, amplicon Z, is generated by amplification of the nucleic acid sequence between P1' and P2 after the targeted incorporation event of the target nucleic acid, thereby amplifying the 3' junction. In other embodiments, P1' can be identical to P1. Additionally, P2' can be identical to P2.

[0274] In certain embodiments, the donor template comprises a cargo and a priming site (P1'), where P1' is located 3' of the cargo nucleic acid sequence (rnpA1--N--P1'--A2), and P1' is substantially identical to P1. In this scenario, a primer pair comprising an oligonucleotide targeting P1' and P1 and an oligonucleotide targeting P2 can be used to amplify the targeted locus, thereby generating two amplicons of similar size, which can be sequenced to determine whether targeted integration occurred. A first amplicon, amplicon X, is generated by amplification of the nucleic acid sequence between P1 and P2 as a result of non-targeted integration in the target nucleic acid. A second amplicon, amplicon Z, is generated by amplification of the nucleic acid sequence between P1' and P2 after the targeted integration event of the target nucleic acid, thereby amplifying the 3' junction. In other embodiments, P1' can be identical to P1. Furthermore, P2' can be identical to P2.

[0275] In certain embodiments, the target nucleic acid comprises a first priming site (P1) and a second priming site (P2), and the donor template comprises a priming site P2', where P2' is located 5' of the cargo nucleic acid sequence (i.e., A1--P2'--N--A2), and P2' is substantially identical to P2. In this scenario, a primer pair comprising an oligonucleotide targeting P2' and P2 and an oligonucleotide targeting P1 can be used to amplify the targeted locus, thereby generating two amplicons of similar size, which can be sequenced to determine whether targeted integration occurred. A first amplicon, amplicon X, arises as a result of non-targeted integration in the target nucleic acid by amplification of the nucleic acid sequence between P1 and P2'. A second amplicon, amplicon Y, arises by amplification of the nucleic acid sequence between P1 and P2' after the targeted integration event of the target nucleic acid, thereby amplifying the 5' junction. In other embodiments, P1' can be identical to P1. Additionally, P2' can be identical to P2.

[0276] The priming site of the donor template can be of any length that allows for quantitative assessment of gene editing events in the target nucleic acid by amplification and / or sequencing of a portion of the target nucleic acid. For example, in certain embodiments, the target nucleic acid comprises a first priming site (P1) and the donor template comprises a priming site (P1'). In these embodiments, the lengths of the P1' priming site and the P1 primer site are such that a single primer can specifically anneal to both priming sites (e.g., in certain embodiments, the lengths of the P1' priming site and the P1 priming site are such that they both have the same or very similar GC content).

[0277] In certain embodiments, the priming site of the donor template is 60 nucleotides in length. In certain embodiments, the priming site of the donor template is less than 60 nucleotides in length. In certain embodiments, the priming site of the donor template is less than 50 nucleotides in length. In certain embodiments, the priming site of the donor template is less than 40 nucleotides in length. In certain embodiments, the priming site of the donor template is less than 30 nucleotides in length. In certain embodiments, the priming site of the donor template is 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides in length. In certain embodiments, the priming site of the donor template is 60 base pairs in length. In certain embodiments, the priming site of the donor template is less than 60 base pairs in length. In certain embodiments, the priming site of the donor template is less than 50 base pairs in length. In certain embodiments, the priming site of the donor template is less than 40 base pairs in length. In certain embodiments, the priming site of the donor template is less than 30 base pairs in length. In certain embodiments, the priming site of the donor template is 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 base pairs in length.

[0278] In certain embodiments, upon a cleavage separation event at the cleavage site of the target nucleic acid and homologous recombination of the donor template with the target nucleic acid, the distance between the first priming site (P1) of the target nucleic acid and the now incorporated P2' priming site is 600 base pairs or less. In certain embodiments, upon a cleavage separation event at the cleavage site of the target nucleic acid and homologous recombination of the donor template with the target nucleic acid, the distance between the first priming site (P1) of the target nucleic acid and the now incorporated P2' priming site is 550, 500, 450, 400, 350, 300, 250, 200, 150 base pairs or less. In certain embodiments, upon a cleavage separation event at the target nucleic acid and homologous recombination of the donor template with the target nucleic acid, the distance between the first priming site (P1) of the target nucleic acid and the now incorporated P2' priming site is 600 nucleotides or less. In certain embodiments, upon cleavage separation events in the target nucleic acid and homologous recombination of the donor template with the target nucleic acid, the distance between the first priming site (P1) of the target nucleic acid and the now incorporated P2' priming site is 550, 500, 450, 400, 350, 300, 250, 200, 150 nucleotides or less.

[0279] In certain embodiments, the target nucleic acid comprises a second priming site (P2), and the donor template comprises a priming site (P2') substantially identical to P2. In certain embodiments, upon a cleavage separation event at the target nucleic acid and homologous recombination of the donor template with the target nucleic acid, the distance between the second priming site (P2) of the target nucleic acid and the now incorporated P1' priming site is 600 base pairs or less. In certain embodiments, upon a cleavage separation event at the target nucleic acid and homologous recombination of the donor template with the target nucleic acid, the distance between the second priming site (P2) of the target nucleic acid and the now incorporated P1' priming site is 550, 500, 450, 400, 350, 300, 250, 200, 150 base pairs or less. In certain embodiments, upon a cleavage separation event at the target nucleic acid and homologous recombination of the donor template with the target nucleic acid, the distance between the second priming site (P2) of the target nucleic acid and the now incorporated P1' priming site is 600 nucleotides or less. In certain embodiments, upon cleavage separation events in the target nucleic acid and homologous recombination of the donor template with the target nucleic acid, the distance between the second priming site (P2) of the target nucleic acid and the now incorporated P1' priming site is 550, 500, 450, 400, 350, 300, 250, 200, 150 nucleotides or less.

[0280] In certain embodiments, the nucleic acid sequence of P2' is contained within the nucleic acid sequence of A1. In certain embodiments, the nucleic acid sequence of P2' is immediately adjacent to the nucleic acid sequence of A1. In certain embodiments, the nucleic acid sequence of P2' is immediately adjacent to the nucleic acid sequence of N. In certain embodiments, the nucleic acid sequence of P2' is contained within the nucleic acid sequence of N.

[0281] In certain embodiments, the nucleic acid sequence of P1' is contained within the nucleic acid sequence of A2. In certain embodiments, the nucleic acid sequence of P1' is immediately adjacent to the nucleic acid sequence of A2. In certain embodiments, the nucleic acid sequence of P1' is immediately adjacent to the nucleic acid sequence of N. In certain embodiments, the nucleic acid sequence of P1' is contained within the nucleic acid sequence of N.

[0282] In certain embodiments, the nucleic acid sequence of P2' is contained within the nucleic acid sequence of S1. In certain embodiments, the nucleic acid sequence of P2' is immediately adjacent to the nucleic acid sequence of S1. In certain embodiments, the nucleic acid sequence of P1' is contained within the nucleic acid sequence of S2. In certain embodiments, the nucleic acid sequence of P1' is immediately adjacent to the nucleic acid sequence of S2.

[0283] cargo The donor template of the gene editing system described herein comprises a cargo (N). The cargo can be of any length necessary to achieve the desired result. For example, the cargo sequence can be less than 2500 base pairs or less than 2500 nucleotides in length. In other embodiments, the cargo sequence can be 12 kb or less. In other embodiments, the cargo sequence can be 10 kb or less. In other embodiments, the cargo sequence can be 7 kb or less. In other embodiments, the cargo sequence can be 5 kb or less. In other embodiments, the cargo sequence can be 4 kb or less. In other embodiments, the cargo sequence can be 3 kb or less. In other embodiments, the cargo sequence can be 2 kb or less. In other embodiments, the cargo sequence can be 1 kb or less. In certain embodiments, the cargo can be about 5-10 kb in length. In other embodiments, the cargo can be about 1-5 kb in length. In other embodiments, the cargo can be about 0-1 kb in length. For example, in exemplary embodiments, the cargo can be about 1000, 900, 800, 700, 600, 500, 400, 300, 200, or 100 base pairs or nucleotides in length. In other exemplary embodiments, the cargo can be about 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1, or 0 base pairs or nucleotides in length. One of skill in the art will readily ascertain that when delivering a donor template using a size-constrained delivery vehicle (e.g., a viral delivery vehicle such as an adeno-associated virus (AAV), adenovirus, lentivirus, integration-deficient lentivirus (IDLV), or herpes simplex virus (HSV) delivery vehicle), the size of the donor template, including the cargo, must not exceed the size limit of the delivery system.

[0284] In certain embodiments, the cargo comprises a replacement sequence. In certain embodiments, the cargo comprises an exon of a gene sequence. In certain embodiments, the cargo comprises an intron of a gene sequence. In certain embodiments, the cargo comprises a cDNA sequence. In certain embodiments, the cargo comprises a transcriptional regulatory element. In certain embodiments, the cargo comprises a reverse complement of a replacement sequence, an exon of a gene sequence, an intron of a gene sequence, a cDNA sequence, or a transcriptional regulatory element. In certain embodiments, the cargo comprises a portion of a replacement sequence, an exon of a gene sequence, an intron of a gene sequence, a cDNA sequence, or a transcriptional regulatory element. In certain embodiments, the cargo is a transgene sequence. In certain embodiments, the cargo introduces a deletion in a target nucleic acid. In certain embodiments, the cargo comprises an exogenous sequence. In other embodiments, the cargo comprises an endogenous sequence.

[0285] Replacement sequences in donor templates are described elsewhere, including Cotta-Ramusino et al. Replacement sequences can be of any suitable length (including zero nucleotides when the desired repair outcome is a deletion) and typically contain one, two, three, or more sequence modifications relative to the native sequence in the cell for which editing is desired. One common sequence modification involves altering the native sequence to repair a mutation associated with a disease or condition for which treatment is desired. Another common sequence modification involves altering one or more sequences complementary to or encoding the PAM sequence of an RNA-guided nuclease or the targeting domain of a gRNA used to generate an SSB or DSB, to reduce or eliminate recurrent cleavage at the target site after the replacement sequence is integrated into the target site.

[0286] A particular cargo can be selected for a given application based on the cell type to be edited, the target nucleic acid, and the effect to be achieved.

[0287] For example, in certain embodiments, it may be desirable to "knock in" a desired gene sequence at a selected chromosomal locus in a target cell. In such cases, the cargo may comprise the desired gene sequence. In certain embodiments, the gene sequence encodes a desired protein, such as, for example, an exogenous protein, an orthologous protein, or an endogenous protein, or a combination thereof.

[0288] In certain embodiments, the cargo may contain a wild-type sequence or a sequence that includes one or more modifications relative to the wild-type sequence. For example, in some embodiments where it is desirable to correct a mutation in a target gene in a cell, the cargo may be designed to restore the wild-type sequence to the target protein.

[0289] In other embodiments, it may also be desirable to "knock out" a gene sequence at a selected chromosomal locus in a target cell. In such cases, the cargo may be designed to integrate into a site that disrupts expression of the target gene sequence, such as, for example, the coding region of the target gene sequence or an expression control region of the target gene sequence, such as, for example, a promoter or enhancer of the target gene sequence. In other embodiments, the cargo may be designed to disrupt the target gene sequence. For example, in certain embodiments, the cargo may introduce a deletion, insertion, stop codon, or frameshift mutation into the target nucleic acid.

[0290] In certain embodiments, the donor is designed to delete all or a portion of the target nucleic acid sequence. In certain embodiments, the homology arms of the donor can be designed to flank the desired deletion site. In certain embodiments, the donor does not contain a cargo sequence between the homology arms, and targeted integration of the donor results in deletion of the portion of the target nucleic acid located between the homology arms. In other embodiments, the donor contains a cargo sequence homologous to the target nucleic acid, and one or more nucleotides of the target nucleic acid sequence are absent from the cargo. Following targeted integration of the donor, the target nucleic acid will contain a deletion at residues absent from the cargo sequence. The size of the deletion can be selected based on the size of the target nucleic acid and the desired effect. In certain embodiments, the donor is designed to introduce a deletion of 1 to 2,000 nucleotides into the target nucleic acid following targeted integration. In other embodiments, the donor is designed to introduce a deletion of 1 to 1,000 nucleotides into the target nucleic acid following targeted integration. In other embodiments, the donor is designed to introduce a deletion of 1 to 500 nucleotides into the target nucleic acid following targeted integration. In other embodiments, the donor is designed to introduce a deletion of 1 to 100 nucleotides into the target nucleic acid following targeted integration. In exemplary embodiments, the donor is designed to introduce a deletion of about 2000, 1500, 1000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide into the target nucleic acid following targeted integration. In other embodiments, the donor is designed to introduce a deletion of more than 2000 nucleotides from the target nucleic acid following targeted integration, for example, a deletion of about 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000 or more nucleotides.

[0291] In certain embodiments, the cargo may include a promoter sequence, hi other embodiments, the cargo is designed to be integrated into a site under the control of a promoter endogenous to the target cell.

[0292] In certain embodiments, cargo encoding an exogenous or orthologous protein or polypeptide may be integrated into a protein-encoding chromosomal sequence such that the chromosomal sequence is inactivated but the exogenous sequence is expressed. In other embodiments, the cargo sequence may be integrated into a chromosomal sequence without altering expression of the chromosomal sequence. This may be accomplished by integrating the cargo into a "safe harbor" locus, such as the Rosa26 locus, the HPRT locus, or the AAV locus.

[0293] In certain embodiments, the cargo encodes a protein associated with a disease or disorder. In certain embodiments, the cargo may encode the wild-type form of the protein or is designed to restore expression of the wild-type form of the protein when the protein is deficient in a subject suffering from a disease or disorder. In other embodiments, the cargo encodes a protein associated with a disease or disorder, and the protein encoded by the cargo contains at least one modification such that a modified version of the protein protects against the onset of the disease or disorder. In other embodiments, the cargo encodes a protein containing at least one modification such that a modified version of the protein causes or enhances the disease or disorder.

[0294] In certain embodiments, cargo can be used to insert genes from one species into the genome of a different species. For example, "humanized" animal models and / or "humanized" animal cells can be generated through targeted integration of human genes into the genome of a non-human animal species, such as, for example, a mouse, a rat, or a non-human primate species. In certain embodiments, such humanized animal models and animal cells contain integrated sequences encoding one or more human proteins.

[0295] In another embodiment, the cargo encodes a protein that provides a benefit to a plant species, including crops such as grains, fruits, or vegetables. For example, the cargo may encode a protein that allows the plant to be grown at higher temperatures, may confer extended shelf life following harvest, or may confer disease resistance. In certain embodiments, the cargo may encode a protein that confers resistance to diseases and pests (see, e.g., Jones et al. (1994) Science 266:789 (cloning of the tomato Cf-9 gene for resistance to Cladosporium fulvum); Martin et al. (1993) Science 262:1432; Mindrinos et al. (1994) Cell 78:1089 (the RSP2 gene for resistance to Pseudomonas syringae); WO 96 / 30517 (resistance to soybean cyst nematode)). In other embodiments, the cargo may encode a protein that encodes resistance to an herbicide, as described in U.S. Patent Application Publication No. 2013 / 0326645A1, the entire contents of which are incorporated herein by reference.In another embodiment, the cargo encodes a protein that confers value-added traits to plant cells, such as, but not limited to, modified fatty acid metabolism, reduced phytate content, and altered carbohydrate composition, resulting from, for example, transforming plants with genes encoding enzymes that alter the branching pattern of starch (see, e.g., Shiroza et al. (1988) J. Bacteol. 170:810 (nucleotide sequence of a Streptococcus variant fructosyltransferase gene); Steinmetz et al. (1985) Mol. Gen. Genet. 20:220 (levansucrase gene); Pen et al. (1992) Bio / Technology 10:292 (α-amylase); Elliot et al. (1993) Plant Molec. Biol. 21:515 (nucleotide sequence of a tomato invertase gene); Sogaard et al. (1993) J. Biol. Chem. 268:22480 (barley α-amylase gene); and Fisher et al. (1993) Plant Physiol. 102:1045 (corn endosperm starch branching enzyme II). Other exemplary cargoes useful for targeted integration in plant cells are described in U.S. Patent Application Publication No. 2013 / 0326645 A1, the entire contents of which are incorporated herein by reference.

[0296] The additional cargo can be selected by one of skill in the art for a given application based on the cell type to be edited, the target nucleic acid, and the effect to be achieved.

[0297] Staffer In certain embodiments, the donor template may optionally include one or more stuffer sequences. Generally, the stuffer sequence is a heterologous or random nucleic acid sequence selected to (a) promote (or not inhibit) targeted integration of the donor template into the target site and subsequent amplification of the amplicon containing the stuffer sequence by certain methods of the present disclosure, but (b) avoid promoting integration of the donor template into another site. The stuffer sequence may be placed, for example, between the homology arm A1 and the primer site P2' to adjust the size of the amplicon generated when the donor template sequence is integrated into the target site. Such size adjustment can be used, for example, to balance the size of the amplicons generated by integrated and unintegrated target sites, thereby balancing the efficiency with which each amplicon is generated in a single PCR reaction; this can then facilitate quantitative assessment of targeted integration based on the relative abundance of the two amplicons in the reaction mixture.

[0298] To facilitate targeted integration and amplification, stuffer sequences can be selected to minimize the formation of secondary structures that may prevent DNA repair mechanisms (e.g., via homologous recombination) from separating the break site or that may prevent amplification. In certain embodiments, the donor template comprises, in the 5' to 3' direction, A1--S1--P2'--N--A2 or A1--N--P1'--S2--A2; where S1 is the first stuffer sequence and S2 is the second stuffer sequence.

[0299] In certain embodiments, the donor template comprises, in the 5' to 3' direction, A1--S1--P2'--N--P1'--S2--A2, where S1 is the first stuffer sequence and S2 is the second stuffer sequence.

[0300] In certain embodiments, the stuffer sequence contains approximately the same guanine-cytosine content ("GC content") as the genome of the entire cell. In certain embodiments, the stuffer sequence contains approximately the same GC content as the targeted locus. For example, if the target cell is a human cell, the stuffer sequence contains about 40% GC content. In certain embodiments, the stuffer sequence can be designed by creating random nucleic acid sequences containing the desired GC content. For example, to generate a stuffer sequence containing 40% GC content, a nucleic acid sequence can be designed with the following nucleotide distribution: A=30%, T=30%, G=20%, C=20%. Methods for determining the GC content of a genome or the GC content of a target locus are known to those of skill in the art. Thus, in certain embodiments, the stuffer sequence contains 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, or 75% GC content. Exemplary 2.0 kilobase stuffer sequences having a GC content of 40±5% are provided herein as SEQ ID NOs: 23-123.

[0301] In certain embodiments, the first stuffer comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 160, at least 170, at least 180, at least 190, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, at least 310, at least 320, at least 330, at least 340, at least 350, at least 360, at least 370, at least 380, at least 390, at least 410, at least 420, at least 430, at least 440, at least 450, at least 460, at least 470, at least 480, at least 490, at least 500, at least 510, at least 520, at least 530, at least 540, at least 550, at least 550, at least 550, at least 600, at least 650, at least 700, at least 750, at least 8 In some embodiments, the sequence comprises 0, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, at least 200, at least 205, at least 210, at least 215, at least 220, at least 225, at least 230, at least 235, at least 240, at least 245, at least 250, at least 275, at least 300, at least 325, at least 350, at least 375, at least 400, at least 425, at least 450, at least 475 or at least 500 nucleotides.In another embodiment, the second stuffer is at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least 145, at least 150, at least 160, at least 170, at least 180, at least 190, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, at least 310, at least 320, at least 330, at least 340, at least 350, at least 360, at least 370, at least 380, at least 390, at least 410, at least 420, at least 430, at least 440, at least 450, at least 460, at least 470, at least 480, at least 490, at least 500, at least 510, at least 520, at least 530, at least 540, at least 550, at least 550, at least 560, at least 570, at least 580, at least 590, at least 610, at least 6 In some embodiments, the sequence comprises 0, at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, at least 200, at least 205, at least 210, at least 215, at least 220, at least 225, at least 230, at least 235, at least 240, at least 245, at least 250, at least 275, at least 300, at least 325, at least 350, at least 375, at least 400, at least 425, at least 450, at least 475 or at least 500 nucleotides.

[0302] It is preferred that the stuffer sequence does not interfere with the separation of the cleavage site in the target nucleic acid. Therefore, the stuffer sequence should have minimal sequence identity with the nucleic acid sequence at the cleavage site of the target nucleic acid. In certain embodiments, the stuffer sequence is less than 80%, 70%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, or 10% identical to any nucleic acid sequence within 500, 450, 400, 350, 300, 250, 200, 150, 100, or 50 nucleotides of the cleavage site of the target nucleic acid. In certain embodiments, the stuffer sequence is less than 80%, 70%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, or 10% identical to any nucleic acid sequence within 500, 450, 400, 350, 300, 250, 200, 150, 100, 50 base pairs of the cleavage site of the target nucleic acid.

[0303] To avoid off-target molecular recombination events, it is preferred that the stuffer sequence have minimal homology to nucleic acid sequences in the genome of the target cell. In certain embodiments, the stuffer sequence has minimal sequence identity to nucleic acids in the genome of the target cell. In certain embodiments, the stuffer sequence is less than 80%, 70%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, or 10% identical to any nucleic acid sequence of the same length (measured in base pairs or nucleotides) in the genome of the target cell. In certain embodiments, a 20 base pair stretch of the stuffer sequence is less than 80%, 70%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, or 10% identical to any stretch of at least 20 base pairs of nucleic acid in the genome of the target cell. In certain embodiments, a 20 nucleotide stretch of the stuffer sequence is less than 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20% or 10% identical to any stretch of at least 20 nucleotides of nucleic acid in the target cell genome.

[0304] In certain embodiments, the stuffer sequence has minimal sequence identity with a nucleic acid sequence in the donor template (e.g., a nucleic acid sequence of the cargo or a nucleic acid sequence of a priming site present in the donor template). In certain embodiments, the stuffer sequence is less than 80%, 70%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, or 10% identical to any nucleic acid sequence of the same length (measured in base pairs or nucleotides) within the donor template. In certain embodiments, a 20 base pair stretch of the stuffer sequence is less than 80%, 70%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, or 10% identical to any 20 base pair stretch of nucleic acid in the donor template. In certain embodiments, a 20 nucleotide stretch of the stuffer sequence is less than 80%, 70%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, or 10% identical to any 20 nucleotide stretch of the donor template nucleic acid.

[0305] In certain embodiments, the length of the first homology arm and its adjacent stuffer sequence (i.e., A1+S1) is approximately equal to the length of the second homology arm and its adjacent stuffer sequence (i.e., A2+S2). For example, in certain embodiments, the length of A1+S1 is the same as the length of A2+S2 (as determined in base pairs or nucleotides). In certain embodiments, the length of A1+S1 differs from the length of A2+S2 by no more than 25 nucleotides. In certain embodiments, the length of A1+S1 differs from the length of A2+S2 by no more than 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 nucleotides. In certain embodiments, the length of A1+S1 differs from the length of A2+S2 by no more than 25 base pairs. In certain embodiments, the length of A1+S1 differs from the length of A2+S2 by no more than 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or 2 base pairs.

[0306] In certain embodiments, the length of A1+H1 is 250 base pairs or less. In certain embodiments, the length of A1+H1 is 200 base pairs or less. In certain embodiments, the length of A1+H1 is 150 base pairs or less. In certain embodiments, the length of A1+H1 is 100 base pairs or less. In certain embodiments, the length of A1+H1 is 50 base pairs or less. In certain embodiments, the length of A1+H1 is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 base pairs. In certain embodiments, the length of A1+H1 is 40 base pairs. In certain embodiments, the length of A2+H2 is 250 base pairs or less. In certain embodiments, the length of A2+H2 is 200 base pairs or less. In certain embodiments, the length of A2+H2 is 150 base pairs or less. In certain embodiments, the length of A2+H2 is 100 base pairs or less. In certain embodiments, the length of A2+H2 is 50 base pairs or less. In certain embodiments, the length of A2+H2 is 250, 200, 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, or 20 base pairs. In certain embodiments, the length of A2+H2 is 40 base pairs.

[0307] In certain embodiments, the length of A1+S1 is the same as the length of H1+X+H2 (as determined by nucleotides or base pairs). In certain embodiments, the length of A1+S1 differs from the length of H1+X+H2 by fewer than 25 nucleotides. In certain embodiments, the length of A1+S1 differs from the length of H1+X+H2 by 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 nucleotides. In certain embodiments, the length of A1+S1 differs from the length of H1+X+H2 by fewer than 25 base pairs. In certain embodiments, the length of A1+S1 differs from the length of H1+X+H2 by 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or 2 base pairs.

[0308] In certain embodiments, the length of A2+S2 is the same as the length of H1+X+H2 (as determined by nucleotides or base pairs). In certain embodiments, the length of A2+S2 differs from the length of H1+X+H2 by fewer than 25 nucleotides. In certain embodiments, the length of A2+S2 differs from the length of H1+X+H2 by 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 nucleotides. In certain embodiments, the length of A2+S2 differs from the length of H1+X+H2 by fewer than 25 base pairs. In certain embodiments, the length of A2+S2 differs from the length of H1+X+H2 by 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or 2 base pairs.

[0309] target cell Using a genome editing system according to the present disclosure, cells can be manipulated or modified, for example, to edit or modify a target nucleic acid. Manipulation can, in various embodiments, be performed in vivo or ex vivo.

[0310] A variety of cell types can be manipulated or modified in accordance with embodiments of the present disclosure, and in some cases, such as in vivo applications, multiple cell types can be modified or manipulated, for example, by delivering a genome editing system according to the present disclosure to multiple cell types. However, in other cases, it may be desirable to limit the manipulation or modification to specific cell types. For example, in some instances, it may be desirable to edit cells with limited differentiation potential or terminally differentiated cells, such as photoreceptor cells in the Maeder example, where modifying the genotype is expected to result in a change in cell phenotype. However, in other cases, it may be desirable to edit less differentiated, multipotent, or pluripotent stem or progenitor cells. By way of example, the cells may be embryonic stem cells, induced pluripotent stem cells (iPSCs), hematopoietic stem / progenitor cells (HSPCs), or other stem or progenitor cell types that differentiate into cell types relevant to a given use or indication.

[0311] In certain embodiments, the engineered cell is a eukaryotic cell. For example, but not limited to, the cell is a vertebrate, mammalian, rodent, goat, pig, bird, chicken, turkey, bovine, equine, ovine, fish, primate, or human cell. In certain embodiments, the engineered cell is a somatic cell, an embryonic cell, or a prenatal cell. In certain embodiments, the engineered cell is a zygote cell, a blastocyst cell, an embryonic cell, a stem cell, a mitosis-competent cell, or a meiosis-competent cell. In certain embodiments, the engineered cell is not part of a human embryo. In certain embodiments, the engineered cell is a T cell, a CD8 + T cells, CD8 + Naive T cells, CD4 + central memory T cells, CD8 + central memory T cells, CD4 + Effector memory T cells, CD4 + Effector memory T cells, CD4 + T cells, CD4 + Stem cell memory T cell, CD8 + Stem cell memory T cell, CD4 + Helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, CD4+ naive T cells, TH17CD4 +T cells, TH1CD4 + T cells, TH2CD4 + T cells, TH9CD4 + T cells, CD4 + Foxp3 + T cells, CD4 + CD25 + CD127 - T cells, CD4 + CD25 + CD127 - Foxp3 + In certain embodiments, the engineered cells are long-term hematopoietic stem cells, short-term hematopoietic stem cells, multipotent progenitor cells, lineage-restricted progenitor cells, lymphoid progenitor cells, myeloid progenitor cells, common myeloid progenitor cells, erythroid progenitor cells, megakaryocytic erythroid progenitor cells, retinal cells, photoreceptor cells, rod cells, cone cells, retinal pigment epithelial cells, trabecular meshwork cells, cochlear hair cells, outer hair cells, inner hair cells, pulmonary epithelial cells, bronchial epithelial cells, alveolar epithelial cells, pulmonary epithelial progenitor cells, striated muscle cells, cardiac muscle cells, muscle satellite cells, neurons, neural stem cells, mesenchymal stem cells, induced pluripotent stem (iPS) cells, embryonic stem cells, monocytes, megakaryocytes, neutrophils, eosinophils, basophils, mast cells, reticulocytes, B cells (e.g., precursor B cells, pre-B cells, pro-B cells, memory B cells, plasma B cells, etc.), gastrointestinal epithelial cells, biliary epithelial cells, pancreatic duct epithelial cells, intestinal stem cells, hepatic parenchymal cells, hepatic stellate cells, Kupffer cells, osteoblasts, osteoclasts, adipocytes, pre-adipocytes, pancreatic islet cells (e.g., beta cells, alpha cells, delta cells), pancreatic exocrine cells, Schwann cells, or oligodendrocytes. In certain embodiments, the engineered cells are plant cells, such as, for example, monocotyledonous or dicotyledonous plant cells.

[0312] In certain embodiments, the target cell is a circulating blood cell, such as, for example, a reticulocyte, a megakaryocyte erythroid progenitor (MEP) cell, a myeloid progenitor (CMP / GMP), a lymphoid progenitor (LP) cell, a hematopoietic stem / progenitor cell (HSC), or an endothelial cell (EC). In certain embodiments, the target cell is a bone marrow cell (e.g., a reticulocyte, an erythroid cell (e.g., an erythroblast), an MEP cell, a myeloid progenitor (CMP / GMP), a LP cell, an erythroid progenitor (EP) cell, an HSC, a multipotent progenitor (MPP) cell, an endothelial cell (EC), a hemogenic endothelial (HE) cell, or a mesenchymal stem cell). In certain embodiments, the target cell is a myeloid progenitor cell (e.g., a common myeloid progenitor (CMP) cell or a granulocyte-macrophage progenitor (GMP) cell). In certain embodiments, the target cell is a lymphoid progenitor cell, such as, for example, a common lymphoid progenitor (CLP) cell. In certain embodiments, the target cell is an erythroid progenitor cell (e.g., an MEP cell). In certain embodiments, the target cell is a hematopoietic stem / progenitor cell (e.g., a long-term HSC (LT-HSC), a short-term HSC (ST-HSC), an MPP cell, or a lineage-restricted progenitor (LRP) cell). In certain embodiments, the target cell is a CD34 + cells, CD34 + CD90 + cells, CD34 + CD38 - cells, CD34 + CD90 + CD49f + CD38 - CD45RA - cells, CD105 + cells, CD31 + or CD133 + cells or CD34 + CD90 + CD133 + In certain embodiments, the target cells are cord blood CD34 + HSPC, umbilical vein endothelial cells, umbilical artery endothelial cells, amniotic fluid CD34 + cells, amniotic fluid endothelial cells, placental endothelial cells or placental hematopoietic CD34 + In certain embodiments, the target cells are mobilized peripheral blood hematopoietic CD34 cells. +cells (after the patient has been treated with a mobilizing agent such as G-CSF or plerixafor). In certain embodiments, the target cells are peripheral blood endothelial cells.

[0313] Of course, the cells being modified or engineered may be variously dividing or non-dividing cells, depending on the cell type being targeted and / or the desired editing outcome.

[0314] If cells are manipulated or modified ex vivo, the cells can be used immediately (e.g., administered to a subject) or they can be maintained or stored for later use. Those skilled in the art will understand that cells can be maintained or stored in culture (e.g., frozen in liquid nitrogen) using any suitable method known in the art.

[0315] Implementation of genome editing systems: delivery, formulation, and administration routes As discussed above, the genome editing system of the present disclosure can be implemented in any suitable manner, meaning that the components of the system, including, but not limited to, RNA-guided nucleases, gRNAs, and any donor template nucleic acids, can be delivered, formulated, or administered in any suitable form or combination of forms that result in the transduction, expression, or introduction of the genome editing system and / or cause the desired repair result in a cell, tissue, or subject. Some non-limiting examples of genome editing system implementations are described in Tables 10 and 11. However, those skilled in the art will understand that these lists are not comprehensive and that other implementations are possible. With particular reference to Table 10, this table lists several exemplary implementations of genome editing systems that include a single gRNA and any donor template. However, genome editing systems according to the present disclosure can incorporate other components, such as multiple gRNAs, multiple RNA-guided nucleases, and proteins, and various implementations will be apparent to those skilled in the art based on the principles set forth in the table. In the table, [N / A] indicates that the genome editing system does not include the indicated component.

[0316] TIFF2024023294000030.tif237170

[0317] Table 11 summarizes various delivery methods for the components of the genome editing system as described herein. Again, the list is intended to be exemplary rather than limiting.

[0318] TIFF2024023294000031.tif242170

[0319] Nucleic acid-based delivery of genome editing systems Nucleic acids encoding various elements of the genome editing system according to the present disclosure can be administered to a subject or delivered intracellularly by methods known in the art or as described herein. For example, the DNA encoding the RNA-guided nuclease and / or the DNA encoding the gRNA and the donor template nucleic acid can be delivered by, for example, a vector (e.g., a viral vector or a non-viral vector), a non-vector-based method (e.g., using naked DNA or a DNA complex), or a combination thereof.

[0320] Nucleic acids encoding genome editing systems or components thereof can be delivered directly to cells as naked DNA or RNA, for example, by transfection or electroporation, or can be conjugated to molecules (e.g., N-acetylgalactosamine) that facilitate uptake by target cells (e.g., erythrocytes, HSCs). Nucleic acid vectors, such as those summarized in Table 11, can also be used.

[0321] The nucleic acid vector may include one or more sequences encoding genome editing system components, such as an RNA-guided nuclease, a gRNA, and / or a donor template. The vector may also include a sequence encoding a signal peptide (e.g., for nuclear localization, nucleolar localization, or mitochondrial localization) that is linked to (e.g., inserted into or fused to) the protein-encoding sequence. As an example, the nucleic acid vector may include a Cpf1-encoding sequence that includes one or more nuclear localization sequences (e.g., a nuclear localization sequence from SV40).

[0322] The nucleic acid vector can also include any suitable number of regulatory / control elements, such as, for example, promoters, enhancers, introns, polyadenylation signals, Kozak consensus sequences, or internal ribosome entry sites (IRES), etc. These elements are well known in the art and are described in Cotta-Ramusino et al.

[0323] Nucleic acid vectors according to the present disclosure include recombinant viral vectors. Exemplary viral vectors are listed in Table 11, and additional suitable viral vectors and their use and manufacture are described in Cotta-Ramusino et al. Other viral vectors known in the art can also be used. Furthermore, viral particles can be used to deliver components of genome editing systems in the form of nucleic acids and / or peptides. For example, "empty" viral particles can be assembled to contain any suitable cargo. Viral vectors and viral particles can also be engineered to incorporate targeting ligands to alter target tissue specificity.

[0324] In addition to viral vectors, non-viral vectors can be used to deliver nucleic acids encoding genome editing systems according to the present disclosure. One important category of non-viral nucleic acid vectors is nanoparticles, which can be organic or inorganic. Nanoparticles are well known in the art and are summarized in Cotta-Ramusino et al. Any suitable nanoparticle design can be used to deliver genome editing system components or nucleic acids encoding such components. For example, organic (e.g., lipid and / or polymer) nanoparticles may be suitable for use as delivery vehicles in certain embodiments of the present disclosure. Exemplary lipids for use in nanoparticle formulations and / or gene transfer are shown in Table 12, and Table 13 lists exemplary polymers for use in gene transfer and / or nanoparticle formulations.

[0325] TIFF2024023294000032.tif138170

[0326] TIFF2024023294000033.tif214170

[0327] TIFF2024023294000034.tif250170

[0328] Non-viral vectors optionally contain targeting modifications to improve uptake and / or selectively target specific cell types. These targeting modifications include, for example, cell-specific antigens, monoclonal antibodies, single-chain antibodies, aptamers, polymers, sugars (e.g., N-acetylgalactosamine (GalNAc)), and cell-penetrating peptides. Such vectors also optionally use fusogenic and endosome-destabilizing peptides / polymers, undergo acid-induced conformational changes (e.g., to promote endosomal leakage of cargo), and / or incorporate stimulus-cleavable polymers, e.g., for release into cellular compartments. For example, disulfide-based cationic polymers that are cleaved in the reducing cellular environment can be used.

[0329] In certain embodiments, one or more nucleic acid molecules (e.g., DNA molecules) other than components of a genome editing system are delivered, such as, for example, the RNA-guided nuclease component and / or gRNA component described herein. In certain embodiments, the nucleic acid molecule is delivered simultaneously with one or more components of the genome editing system. In certain embodiments, the nucleic acid molecule is delivered before or after one or more components of the genome editing system are delivered (e.g., less than about 30 minutes, 1 hour, 2 hours, 3 hours, 6 hours, 9 hours, 12 hours, 1 day, 2 days, 3 days, 1 week, 2 weeks, or 4 weeks). In certain embodiments, the nucleic acid molecule is delivered by a different means than one or more components of the genome editing system, such as, for example, the RNA-guided nuclease component and / or gRNA component. The nucleic acid molecule may be delivered by any of the delivery methods described herein. For example, the nucleic acid molecule can be delivered by a viral vector, such as, for example, an integration-deficient lentivirus, so that toxicity caused by the nucleic acid (e.g., DNA) can be reduced, and the RNA-guided nuclease molecule component and / or the gRNA component can be delivered by electroporation. In certain embodiments, the nucleic acid molecule encodes a therapeutic protein, such as, for example, a protein described herein. In certain embodiments, the nucleic acid molecule encodes an RNA molecule, such as, for example, an RNA molecule described herein.

[0330] Delivery of RNP and / or RNA-encoding genome editing system components RNA encoding RNP (a complex of gRNA and RNA-guided nuclease) and / or RNA-guided nuclease and / or gRNA can be delivered to cells or administered to a subject by methods known in the art, some of which are described in Cotta-Ramusino et al. In vitro, RNA encoding RNA-guided nuclease and / or gRNA can be delivered by, for example, microinjection, electroporation, transient cell compaction or squeezing (see, e.g., Lee 2012). Lipid-mediated transfection, peptide-mediated delivery, GalNAc or other conjugate-mediated delivery, and combinations thereof can also be used for in vitro and in vivo delivery.

[0331] In vitro, delivery via electroporation involves mixing cells with RNA-guided nucleases and / or gRNA-encoding RNA in a cartridge, chamber, or cuvette, in the presence or absence of donor template nucleic acid molecules, and applying one or more electrical impulses of defined length and amplitude. Electroporation systems and protocols are known in the art, and any suitable electroporation tool and / or protocol may be used in connection with various embodiments of the present disclosure. Exemplary systems include, but are not limited to, Nucleofector™ technologies (Lonza), Gene Pulser Xcell™ (BioRad), Flow Electroporation™ transfection system (MaxCyte), and Neon™ transfection system (ThermoFisher).

[0332] Administration route Genome editing systems or cells modified or engineered using such systems can be administered to a subject by any suitable mode or route, whether local or systemic. Systemic administration methods include oral and parenteral routes. Parenteral routes include, for example, intravenous, intramedullary, intraarterial, intramuscular, intradermal, subcutaneous, intranasal, and intraperitoneal routes. Systemically administered components can be modified or formulated to target, for example, HSCs, hematopoietic stem / progenitor cells, or erythroid precursor or progenitor cells.

[0333] Examples of local administration modes include intramedullary injection into the trabecular bone or intrafemoral injection into the bone marrow cavity, and injection into the portal vein. In certain embodiments, significantly smaller amounts of components (compared to systemic approaches) may be effective when administered locally (e.g., directly into the bone marrow) compared to systemic administration (e.g., intravenously). Local administration methods may reduce or eliminate the occurrence of potentially toxic side effects that may occur when therapeutically effective amounts of components are administered systemically.

[0334] Doses can be provided as periodic boluses (e.g., intravenously) or as continuous infusion from an internal or external reservoir (e.g., from an IV bag or implantable pump). The components can be administered locally, for example, by continuous release from a sustained release drug delivery device.

[0335] Additionally, components can be formulated to be released over an extended period of time. The release system can include a matrix of biodegradable materials or materials that release incorporated components by diffusion. Components can be uniformly or non-uniformly distributed within the release system. A variety of release systems can be useful, and the selection of an appropriate system depends on the release rate required by a particular application. Both non-degradable and degradable release systems can be used. Suitable release systems include polymers and polymeric matrices, non-polymeric matrices, or inorganic and organic excipients and diluents, such as, but not limited to, calcium carbonate and sugars (e.g., trehalose). Release systems can be natural or synthetic. However, synthetic release systems are preferred because they typically result in more reliable, more reproducible, and more defined release profiles. Release system materials can be selected so that components with different molecular weights are released by diffusion through the material or by degradation of the material.

[0336] Exemplary synthetic biodegradable polymers include, for example, polyamides, such as poly(amino acids) and poly(peptides); polyesters, such as poly(lactic acid), poly(glycolic acid), poly(lactic-co-glycolic acid), and poly(caprolactone); poly(anhydrides); polyorthoesters; polycarbonates; and chemical derivatives thereof (e.g., substitution and addition of chemical groups such as alkyl and alkylene, hydroxylation, oxidation, and other modifications routinely performed by one of ordinary skill in the art), copolymers, and mixtures thereof. Typical synthetic non-degradable polymers include, for example, polyethers such as poly(ethylene oxide), poly(ethylene glycol), and poly(tetramethylene oxide); vinyl polymers such as methyl, ethyl, other alkyl, hydroxyethyl methacrylate, acrylic and methacrylic acid, and poly(vinyl alcohol), poly(vinylpyrrolidone), and poly(vinyl acetate)-polyacrylates and polymethacrylates; poly(urethanes); cellulose and its derivatives such as alkyl, hydroxyalkyl, ether, ester, nitrocellulose, and various cellulose acetates; polysiloxanes; any of its chemical derivatives (e.g., substitution of alkyl, alkylene, and other chemical groups, addition, hydroxylation, oxidation, and other modifications routinely made by one skilled in the art), copolymers, and mixtures thereof.

[0337] Poly(lactide-co-glycolide) microspheres may also be used. Typically, the microspheres are composed of lactic acid and glycolic acid polymers structured to form hollow spheres. The spheres may be approximately 15-30 microns in diameter and may be loaded with the components described herein.

[0338] Multimodal or differential delivery of components Those skilled in the art will understand, in light of the present disclosure, that different components of the genome editing systems disclosed herein can be delivered together or separately, and simultaneously or non-simultaneously. Separate and / or asynchronous delivery of genome editing system components can be particularly desirable to provide temporal or spatial control over the function of the genome editing system and limit certain effects caused by their activity.

[0339] As used herein, different or differential modes refer to delivery modes that impart different pharmacodynamic or pharmacokinetic properties to a component molecule of interest, such as an RNA-guided nuclease molecule, gRNA, template nucleic acid, or payload. For example, the delivery mode may result in different tissue distribution, different half-life, or different temporal distribution, e.g., in selected compartments, tissues, or organs.

[0340] Some delivery modes result in more sustained expression and presence of the component, such as delivery by nucleic acid vectors that remain within the cell or its progeny by autonomous replication or insertion into cellular nucleic acid. Examples include viral delivery, such as AAV or lentiviral.

[0341] For example, components of a genome editing system, such as an RNA-guided nuclease and a gRNA, can be delivered by modalities that result in different half-lives or persistence of the delivered components within the body, in particular compartments, tissues, or organs. In certain embodiments, the gRNA can be delivered by such modalities. The RNA-guided nuclease molecule component can be delivered by a modality that results in lower persistence or less exposure to the body or to particular compartments, tissues, or organs.

[0342] More generally, in certain embodiments, a first delivery mode is used to deliver a first component and a second delivery mode is used to deliver a second component. The first delivery mode confers a first pharmacodynamic or pharmacokinetic property. The first pharmacodynamic property can be, for example, the distribution, persistence, or exposure of the component or the nucleic acid encoding the component within the body, compartment, tissue, or organ. The second delivery mode confers a second pharmacodynamic or pharmacokinetic property. The second pharmacodynamic property can be, for example, the distribution, persistence, or exposure of the component or the nucleic acid encoding the component within the body, compartment, tissue, or organ.

[0343] In certain embodiments, a first pharmacodynamic or pharmacokinetic property, eg, distribution, duration, or exposure, is more specific than a second pharmacodynamic or pharmacokinetic property.

[0344] In certain embodiments, the first delivery modality is selected to optimize, eg, minimize, a pharmacodynamic or pharmacokinetic property, such as, for example, distribution, duration, or exposure.

[0345] In certain embodiments, the second delivery modality is selected to optimize, eg, maximize, a pharmacodynamic or pharmacokinetic property, such as, for example, distribution, duration, or exposure.

[0346] In certain embodiments, the first delivery modality involves the use of relatively persistent elements, such as nucleic acids, e.g., plasmids or viral vectors, e.g., AAV or lentivirus, etc. Because such vectors are relatively persistent, the products transcribed from them will be relatively persistent.

[0347] In certain embodiments, the second delivery modality comprises a relatively transient element, such as, for example, RNA or a protein.

[0348] In certain embodiments, the first component comprises a gRNA, and the delivery mode is relatively persistent, e.g., the gRNA is transcribed from a plasmid or viral vector, e.g., AAV or lentivirus. Transcription of these genes is physiologically inconsequential because these genes do not encode protein products and the gRNA cannot act alone. The second component, the RNA-guided nuclease molecule, is delivered in a transient manner, e.g., as mRNA or protein, ensuring that only the complete RNA-guided nuclease molecule / gRNA complex is present and active for a short period of time.

[0349] Furthermore, the components may be delivered in different molecular forms or different delivery vectors that complement each other to enhance safety and tissue specificity.

[0350] The use of differential delivery modes may enhance performance, safety, and / or efficacy, and may reduce the likelihood of eventual off-target modifications, for example. Because peptides from bacterial Cas enzymes are presented on the cell surface by MHC molecules, delivery of immunogenic components, such as Cas9 molecules, in a less persistent manner may reduce immunogenicity. A two-part delivery system may alleviate these drawbacks.

[0351] Differential delivery modes can be used to deliver components to different but overlapping target regions. Formation of active complexes outside the overlapping target regions is minimized. Thus, in certain embodiments, a first component, such as a gRNA, is delivered by a first delivery mode, resulting in a first spatial distribution, such as tissue distribution. A second component, such as an RNA-guided nuclease molecule, is delivered by a second delivery mode, resulting in a second spatial distribution, such as tissue distribution. In certain embodiments, the first mode includes a first component selected from liposomes, nanoparticles, such as polymeric nanoparticles, and nucleic acids, such as viral vectors. The second mode includes a second component selected from the group. In certain embodiments, the first delivery mode includes a first targeting component, such as a cell-specific receptor or antibody, while the second delivery mode does not include that component. In certain embodiments, the second delivery mode includes a second targeting component, such as a second cell-specific receptor or a second antibody.

[0352] When RNA-guided nuclease molecules are delivered in viral delivery vectors, liposomes, or polymeric nanoparticles, there is the potential for delivery to and therapeutic activity in multiple tissues, where it may be desirable to target only a single tissue. A two-part delivery system may solve this challenge and increase tissue specificity. If the gRNA and RNA-guided nuclease molecules are packaged in separate delivery vehicles with distinct but overlapping tissue tropisms, a fully functional complex will form only in the tissues targeted by both vectors.

[0353] Exemplary Non-Limiting Embodiments A. In certain non-limiting embodiments, the presently disclosed subject matter provides isolated CRISPRs from Prevotella and Francisella 1 (Cpf1) RNA-guided nucleases that contain a nuclear localization signal (NLS).

[0354] A1. The Cpf1 RNA-guided nuclease is the Cpf1 RNA-guided nuclease of A above, which contains an NLS at or near the N-terminus of the nuclease.

[0355] A2. The Cpf1 RNA-guided nuclease is the Cpf1 RNA-guided nuclease of A above, which contains an NLS at or near the C-terminus of the nuclease.

[0356] A3. The Cpf1 RNA-guided nuclease is the Cpf1 RNA-guided nuclease of A1 described above, which contains two NLS sequences at or near the N-terminus of the nuclease.

[0357] A4. The Cpf1 RNA-guided nuclease is the Cpf1 RNA-guided nuclease described in A2 above, which contains two NLS sequences at or near the C-terminus of the nuclease.

[0358] A5. The Cpf1 RNA-guided nuclease of A above, which contains an NLS at or near both the N-terminus and C-terminus of the nuclease.

[0359] A6. The Cpf1 RNA-guided nuclease of A above, wherein when the Cpf1 RNA-guided nuclease comprises two or more NLS sequences, the NLS sequences are the same or different.

[0360] A7. The Cpf1 RNA-guided nuclease of A above, wherein the NLS sequence or sequences are selected from the group consisting of nucleoplasmin NLS (nNLS) (SEQ ID NO: 1) and simian virus 40 "SV40" NLS (sNLS) (SEQ ID NO: 2).

[0361] A8. The sequence of the Cpf1 RNA-guided nuclease is selected from the group consisting of NLS (SEQ ID NO: 3); His-AsCpf1-sNLS (SEQ ID NO: 4; His-AsCpf1-sNLS-sNLS (SEQ ID NO: 5); His-sNLS-AsCpf1 (SEQ ID NO: 6); His-sNLS-sNLS-AsCpf1 (SEQ ID NO: 7); sNLS-sNLS-AsCpf1 (SEQ ID NO: 8); His-sNLS-AsCpf1-sNLS (SEQ ID NO: 9); and His-sNLS-sNLS-AsCpf1-sNLS-sNLS (SEQ ID NO: 10).

[0362] B. In certain non-limiting embodiments, the presently disclosed subject matter provides isolated Cpf1 RNA-guided nucleases that include a deletion or substitution of a cysteine ​​amino acid.

[0363] B1. The Cpf1 RNA-guided nuclease of B above, wherein the Cpf1 RNA-guided nuclease comprises a deletion or substitution at C65, C205, C334, C379, C608, C674, C1025, or C1248 of the wild-type AsCpf1 amino acid sequence.

[0364] B2. The Cpf1 RNA-guided nuclease of B1, which comprises, relative to the wild-type AsCpf1 amino acid sequence, a substitution selected from the group consisting of C65S / A, C205S / A, C334S / A, C379S / A, C608S / A, C674S / A and C1025S / A.

[0365] B3. The Cpf1 RNA-guided nuclease of B1 above, which comprises a deletion or substitution at either C334 and C674 or C334, C379 and C674 of the wild-type AsCpf1 amino acid sequence.

[0366] B4. The Cpf1 RNA-guided nuclease of B3 above, which contains, relative to the wild-type AsCpf1 amino acid sequence, substitutions selected from the group consisting of: (1) C334S / A and C674S / A; and (2) C334S / A, C379S / A, and C674S / A.

[0367] B5. The Cpf1 RNA-guided nuclease of B, described above, further contains an NLS.

[0368] B6. The Cpf1 RNA-guided nuclease of B5 as defined above, wherein the sequence of the Cpf1 RNA-guided nuclease is selected from His-AsCpf1-nNLSCys-less (SEQ ID NO: 11) and His-AsCpf1-nNLSCys-low (SEQ ID NO: 12).

[0369] C. In certain embodiments, the presently disclosed subject matter provides an isolated nucleic acid encoding any of the aforementioned Cpf1 RNA-guided nucleases A through A8 and B through B8.

[0370] D. In certain embodiments, the presently disclosed subject matter comprises: guide RNA (gRNA); and Any of the Cpf1 RNA-guided nucleases A to A8 and B to B8 or the Cpf1 RNA-guided nuclease encoded by the nucleic acid C The present invention provides a genome editing system comprising:

[0371] E. In certain embodiments, the presently disclosed subject matter provides a method for modifying a target sequence of interest in a cell, the method comprising: a gRNA complementary to the target sequence of interest; and Any of the Cpf1 RNA-guided nucleases A to A8 and B to B8 or the Cpf1 RNA-guided nuclease encoded by the nucleic acid C wherein the Cpf1 RNA-guided nuclease modifies a target sequence of interest.

[0372] E1. The method of E above, wherein the cells are T cells, hematopoietic stem cells (HSCs), or human umbilical cord blood-derived erythroid progenitor cells (HUDEP cells).

[0373] E2. The method of E1 above, wherein the HSCs are CD34+ cells, CD34+CD90+ cells, CD34+CD38- cells, CD34+CD90+CD49f+CD38-CD45RA- cells, CD105+ cells, CD31+ or CD133+ cells, or CD34+CD90+CD133+ cells.

[0374] E3.T cells are CD8 + T cells, CD8 + Naive T cells, CD4 + central memory T cells, CD8 + central memory T cells, CD4 + Effector memory T cells, CD4 + Effector memory T cells, CD4 + T cells, CD4 + Stem cell memory T cell, CD8 + Stem cell memory T cell, CD4 + Helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, CD4+ naive T cells, TH17CD4 + T cells, TH1CD4 + T cells, TH2CD4 + T cells, TH9CD4 + T cells, CD4 + Foxp3 + T cells, CD4 + CD25 + CD127 - T cells or CD4 + CD25 + CD127 - Foxp3 + T cells, the aforementioned E1 method.

[0375] E4. The method of E above, wherein the Cpf1 RNA-guided nuclease modifies the target sequence of interest to achieve at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% editing.

[0376] E5. The method of E above, further including a second gRNA complementary to a second target sequence of interest.

[0377] E6. The method of E above, further comprising a second RNA-guided nuclease.

[0378] E7. The method of E above, wherein the target sequence of interest is selected from the group consisting of: a portion of the HBG1 gene sequence; and a portion of the BCL11a gene sequence.

[0379] E8. The method of E7 described above, wherein the part of the HBG1 gene sequence is the −110 nt promoter region of the HBG gene.

[0380] E9. The method of E8 described above, in which part of the HBG1 gene sequence is the CAAT box in the −110 nt promoter region of the HBG gene.

[0381] E10. The method of E7 described above, in which the part of the Bcl11a gene sequence is the +58 DHS region of intron 2 of the BCL11a gene.

[0382] E11. The method of E10 described above, in which part of the Bcl11a gene sequence is the GATA1 motif in the +58 DHS region of intron 2 of the BCL11a gene.

[0383] E12. The method of E above, wherein the target sequence of interest is selected from the group consisting of a portion of the FAS gene sequence; a portion of the BID gene sequence; a portion of the CTLA4 gene sequence; a portion of the PDCD1 gene sequence; a portion of the CBLB gene sequence; a portion of the PTPN6 gene sequence; a portion of the B2M gene sequence; a portion of the TRAC gene sequence; and a portion of the TRBC gene sequence.

[0384] E13. The method of E12 above, wherein the target sequence of interest is selected from the group consisting of: a portion of the B2M gene sequence; a portion of the TRAC gene sequence; and a portion of the TRBC gene sequence.

[0385] E14. The method of E13 described above, wherein a portion of the B2M gene sequence is within the first 500 bp of the coding sequence of the B2M gene.

[0386] E15. The method of E13 above, wherein the portion of the B2M gene sequence is between the 501st and last nucleotides of the coding sequence of the B2M gene.

[0387] E16. A portion of the TRAC gene sequence is within the first 500 bp of the coding sequence of the TRAC gene in the aforementioned E12 cells.

[0388] E17. A portion of the TRBC gene sequence is within the first 500 bp of the coding sequence of the TRBC gene in the aforementioned E12 cells.

[0389] F. In certain embodiments, the presently disclosed subject matter provides a method of treating a subject, comprising: a gRNA complementary to a target sequence of the target nucleic acid; and Any of the aforementioned Cpf1 RNA-guided nucleases A to A8 and B to B8 The method includes contacting the

[0390] F1. The method of F above, wherein the Cpf1 molecule forms a double-stranded break in the target nucleic acid.

[0391] F2. The method of any one of F and F1 above, wherein the Cpf1 molecule is selected from the group consisting of Acidaminococcus sp. strain BV3L6 Cpf1 molecule (AsCpf1), Lachnospiraceae bacterium ND2006 Cpf1 molecule (LbCpf1), and Lachnospiraceae bacterium MA2020 (Lb2Cpf1).

[0392] F3. Any one of the methods F-F2 above, wherein the subject is suffering from a hemoglobinopathy.

[0393] F4. The method of F3, wherein the hemoglobinopathy is sickle cell disease or beta thalassemia.

[0394] F5. The method of any one of F-F4 above, wherein the cells are T cells, hematopoietic stem cells (HSCs), or human umbilical cord blood-derived erythroid progenitor cells (HUDEP cells).

[0395] F6. T cells are CD8 + T cells, CD8 + Naive T cells, CD4 + central memory T cells, CD8 + central memory T cells, CD4 + Effector memory T cells, CD4 + Effector memory T cells, CD4 + T cells, CD4 + Stem cell memory T cell, CD8 + Stem cell memory T cell, CD4 + Helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, CD4+ naive T cells, TH17CD4 + T cells, TH1CD4 + T cells, TH2CD4 + T cells, TH9CD4 + T cells, CD4 + Foxp3 + T cells, CD4 + CD25 + CD127 - T cells or CD4 + CD25 + CD127 - Foxp3 + The method of F5 mentioned above, which is a T cell.

[0396] F7.HSC cells are CD34 + cells, CD34 + CD90 + cells, CD34 + CD38 - cells, CD34 + CD90 + CD49f + CD38 - CD45RA - cells, CD105 +cells, CD31 + or CD133 + cells or CD34 + CD90 + CD133 + The aforementioned F5 method is a cell.

[0397] F8. The method of any one of F-F7 above, wherein the contacting is performed ex vivo.

[0398] F9. The method of any one of F-F8 above, wherein the contacted cells are returned to the subject's body.

[0399] G. In certain embodiments, the presently disclosed subject matter comprises: (a) any one of the Cpf1 RNA-guided nucleases A to A8 and B to B8; (b) a gRNA complementary to a target sequence of the target nucleic acid; and (c) cells from the subject that would benefit from one or more modifications of the target nucleic acid A reaction mixture comprising:

[0400] H. In certain embodiments, the presently disclosed subject matter comprises: (a) any one of A to A8 and B to B8, or a nucleic acid composition encoding a Cpf1 RNA-guided nuclease; and (b) a gRNA complementary to the target sequence of the target nucleic acid or the nucleic acid composition of the gRNA A kit comprising:

[0401] I. In certain embodiments, the presently disclosed subject matter provides cells comprising a modification in a target nucleic acid sequence introduced via the genome editing system of D described above.

[0402] I1. The cell of I above, wherein the modification is to the HBG1 gene sequence or the Bcl11a gene sequence.

[0403] I2. The cells of I1 above, wherein the modified HBG1 gene sequence is the −110 nt promoter region of the HBG gene.

[0404] I3. The cell of claim I1, wherein the modified HBG1 gene sequence is the CAAT box in the -110 nt promoter region of the HBG gene.

[0405] I4. The cell of claim I1, wherein the modified Bcl11a gene sequence is the +58 DHS region of intron 2 of the BCL11a gene.

[0406] I5. The cell of claim I1, wherein the modified Bcl11a gene sequence is a GATA1 motif in the +58 DHS region of intron 2 of the BCL11a gene.

[0407] J. In certain embodiments, the presently disclosed subject matter provides a method for assessing CRISPR / Cpf1-mediated editing of a target nucleic acid sequence and / or modulation of expression of a target nucleic acid sequence by a test Cpf1 RNA-guided nuclease, comprising: (a) determining the activity of the test Cpf1 RNA-guided nuclease with respect to editing and / or modulating expression of a target nucleic acid sequence that includes a matched site target nucleic acid sequence; (b) comparing the activity of the test Cpf1 RNA-guided nuclease with the activity of a control RNA-guided nuclease with respect to editing and / or modulating expression of a target nucleic acid sequence that includes a matched site target nucleic acid sequence. The present invention provides a method comprising:

[0408] J1. The method of J above, wherein the matching target nucleic acid sequence is selected from the group consisting of matching site 1 (sequence number 13), matching site 5 (sequence number 14), matching site 11 (sequence number 15) and matching site 18 (sequence number 16).

[0409] J2. Test Cpf1 RNA-guided nuclease and control RNA-guided nuclease (a) have the ...

Claims

1. A genome editing system comprising: (a) a gRNA molecule comprising a targeting domain complementary to a target nucleic acid sequence, wherein the target nucleic acid sequence comprises a part of the TRAC gene sequence; (b) a Cpf1 RNA-guided nuclease or a nucleic acid encoding a Cpf1 RNA-guided nuclease; and (c) a DNA donor template comprising a 5' homology arm, a 3' homology arm, a first stuffer sequence, a second stuffer sequence, and a cargo, wherein the 5' homology arm is homologous to a sequence upstream of the target nucleic acid sequence, the 3' homology arm is homologous to a sequence downstream of the target nucleic acid sequence, and the DNA donor template comprises, in the 5' to 3' direction, a 5' homology arm, a first stuffer sequence, a cargo, a second stuffer sequence, and a 3' homology arm; A genome editing system comprising the above components.

2. The genome editing system according to claim 1, wherein the 5' homology arm and / or the 3' homology arm has a length between 50 and 1000 nucleotides.

3. The genome editing system according to claim 1 or 2, wherein the cargo comprises a substitution sequence.

4. The cargo comprises: (a) an exon of a gene sequence; (b) an intron of a gene sequence; (c) a cDNA sequence; (d) a transcriptional regulatory element; and / or (e) a promoter sequence The genome editing system according to any one of claims 1 to 3.

5. The genome editing system according to any one of claims 1 to 4, wherein the cargo comprises a sequence encoding an exogenous protein.

6. The genome editing system according to any one of claims 1 to 5, wherein the Cpf1 RNA-guided nuclease comprises one or more nuclear localization signal (NLS) sequences at the N-terminus of the nuclease, one or more NLS sequences at the C-terminus of the nuclease, or one or more NLS sequences at the N-terminus and C-terminus of the nuclease.

7. The genome editing system according to claim 6, wherein the NLS sequence is selected from the group consisting of a nucleoplasmin NLS (nNLS) (SEQ ID NO: 1) and a simian virus 40 (SV40) NLS (sNLS) (SEQ ID NO: 2).

8. The genome editing system according to any one of claims 1 to 7, wherein the Cpf1 RNA-guided nuclease comprises substitutions of C334S, C379S, and / or C674S in the wild-type AsCpf1 amino acid sequence of SEQ ID NO:

17.

9. The genome editing system according to any one of claims 1 to 8, wherein the targeting domain comprises the sequences of SEQ ID NOs: 216 to 475.

10. (i) (a) The first stuffer sequence comprises a sequence comprising at least 25 nucleotide sequences of the sequences of SEQ ID NOs: 23 to 123; (b) The second stuffer sequence comprises a sequence comprising at least 25 nucleotide sequences of the sequences of SEQ ID NOs: 23 to 123; or (c) The first stuffer sequence and the second stuffer sequence each comprise a sequence comprising at least 25 nucleotide sequences of the sequences of SEQ ID NOs: 23 to 123; (ii) The lengths of the 5' homology arm and the first stuffer sequence differ from the lengths of the 3' homology arm and the second stuffer sequence by 25 nucleotides or less; and / or (iii) The first stuffer sequence and the second stuffer sequence have less than 10% identity with any nucleic acid sequence of the same length in the genome of the target cell; The genome editing system according to any one of claims 1 to 9.

11. The genome editing system according to any one of claims 1 to 10, for use in the treatment of a subject suffering from cancer.

12. The genome editing system according to any one of claims 1 to 10, for use in modifying a cell.

13. A cell comprising the genome editing system according to any one of claims 1 to 10.

14. The cell according to claim 13, which is a T cell, a lymphoid progenitor cell, a natural killer cell, a dendritic cell, a hematopoietic stem cell (HSC) or a human umbilical cord blood-derived erythroid progenitor cell (HUDEP cell).

15. T cells, CD8 + T cells, CD8 + Naive T cells, CD4 + Central memory T cells, CD8 + Central memory T cells, CD4 + Effector memory T cells, CD4 + T cells, CD4 + Stem cell memory T cells, CD8 + Stem cell memory T cells, CD4 + Helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, CD4 + Naive T cells, TH17 CD4 + T cells, TH1 CD4 + T cells, TH2 CD4 + T cells, TH9 CD4 + T cells, CD4 + Foxp3 + T cells, CD4 + CD25 + CD127 - T cells and CD4 + CD25 + CD127 - Foxp3 + The cell according to claim 14, selected from the group consisting of T cells.

16. An isolated T cell comprising a modification in a target nucleic acid sequence, wherein the isolated T cell (a) a CRISPR 1 (Cpf1) RNA-guided nuclease from the genera Prevotella and Francisella, (b) a gRNA molecule comprising a targeting domain complementary to a target nucleic acid sequence that is a part of the TRAC gene sequence, A DNA donor template comprising a 5' homology arm, a 3' homology arm, a first stuffer sequence, a second stuffer sequence, and a cargo, wherein the 5' homology arm is homologous to a sequence upstream of the target nucleic acid sequence, the 3' homology arm is homologous to a sequence downstream of the target nucleic acid sequence, the DNA donor template comprises, in the 5' to 3' direction, a 5' homology arm, a first stuffer sequence, a cargo, a second stuffer sequence, and a 3' homology arm, and the modification is a deletion, insertion, stop codon, or frameshift mutation in the target nucleic acid sequence, and an isolated T cell comprising the DNA donor template.

17. The isolated T cell according to claim 16, wherein the 5' homology arm and / or the 3' homology arm is between 50 and 1000 nucleotides in length.

18. The isolated T cell according to claim 16 or 17, wherein the cargo comprises a substitution sequence.

19. A population of T cells comprising a modification in the TRAC gene, wherein the modification is (i) (a)a CRISPR / Cpf1 RNA-guided nuclease and (b)a gRNA comprising a targeting domain complementary to a target nucleic acid sequence that is a part of the TRAC gene sequence, one or more RNP complexes comprising, and (ii)a DNA donor template comprising a 5' homology arm, a 3' homology arm, a first stuffer sequence, a second stuffer sequence, and a cargo, wherein the 5' homology arm is homologous to a sequence upstream of the target nucleic acid sequence, the 3' homology arm is homologous to a sequence downstream of the target nucleic acid sequence, and the DNA donor template comprises, in the 5' to 3' direction, a 5' homology arm, a first stuffer sequence, a cargo, a second stuffer sequence, and a 3' homology arm, and the modification is an interference with the expression of the TRAC gene, and the DNA donor template, generated using, at least 60% of the T cells in the population of T cells do not contain a detectable level of TCR on the surface of the T cells. A population of T cells.

20. The isolated T cell according to any one of claims 16 to 18 or the population of T cells according to claim 19, further comprising a chimeric antigen receptor (CAR) inserted into the modified TRAC locus or an engineered T cell receptor (eTCR).

21. T cells, CD8 + T cells, CD8 + Naive T cells, CD4 + Central memory T cells, CD8 + Central memory T cells, CD4 + Effector memory T cells, CD4 + T cells, CD4 + Stem cell memory T cells, CD8 + Stem cell memory T cells, CD4 + Helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, CD4 + Naive T cells, TH17 CD4 + T cells, TH1 CD4 + T cells, TH2 CD4 + T cells, TH9 CD4 + T cells, CD4 + Foxp3 + T cells, CD4 + CD25 + CD127 - T cells or CD4 + CD25 + CD127 - Foxp3 + The isolated T cells according to any one of claims 16 to 18 and 20, which are T cells, or the T cell population according to claim 19 or claim 20.

22. A method of modifying a target sequence of interest in a population of T cells, comprising the population of cells (i) (a) A gRNA molecule comprising a targeting domain complementary to a target nucleic acid sequence, wherein the target nucleic acid sequence comprises a portion of the TRAC gene sequence; and (b) A Cpf1 RNA-guided nuclease; One or more RNP complexes comprising, and (ii) A DNA donor template comprising a 5' homology arm, a 3' homology arm, a first stuffer sequence, a second stuffer sequence, and a cargo, wherein the 5' homology arm is homologous to a sequence upstream of the target nucleic acid sequence, the 3' homology arm is homologous to a sequence downstream of the target nucleic acid sequence, and the DNA donor template comprises, in the 5' to 3' direction, a 5' homology arm, a first stuffer sequence, a cargo, a second stuffer sequence, and a 3' homology arm, a DNA donor template Comprising contacting ex vivo or in vitro with, The one or more RNP complexes modify a target sequence of interest within the population of cells, and The modification is a deletion, insertion, stop codon, or frameshift mutation in the target nucleic acid sequence, Method.

23. A medicament for treating or alleviating a condition in a subject, comprising a population of cells, wherein the population of cells (i) (a) A CRISPR1 (Cpf1) RNA-guided nuclease from the genus Prevotella and the genus Francisella; and (b) A gRNA molecule comprising a targeting domain complementary to a target nucleic acid sequence, wherein the target nucleic acid sequence comprises a portion of the TRAC gene sequence; A complex comprising, and (ii) A DNA donor template comprising a 5' homology arm, a 3' homology arm, a first stuffer sequence, a second stuffer sequence, and a cargo, wherein the 5' homology arm is homologous to a sequence upstream of the target nucleic acid sequence, the 3' homology arm is homologous to a sequence downstream of the target nucleic acid sequence, and the DNA donor template comprises, in the 5' to 3' direction, a 5' homology arm, a first stuffer sequence, a cargo, a second stuffer sequence, and a 3' homology arm, a DNA donor template Generated by delivery of, a medicament comprising a modification in the TRAC gene.

24. The medicament according to claim 23, wherein the subject is suffering from cancer or an autoimmune disorder.

25. The medicament according to claim 23 or 24, wherein at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80% or at least 90% of the cells in the cell population contain productive indels.

26. The medicament according to any one of claims 23 to 25, wherein the 5' homology arm and / or the 3' homology arm is between 50 and 1000 nucleotides in length and / or the cargo contains a substitution sequence.

27. (a)A gRNA molecule comprising a first targeting domain complementary to a nucleic acid encoding a target sequence or a gRNA molecule, wherein the target sequence comprises a part of the TRAC gene sequence; and (b)A DNA donor template comprising a 5' homology arm, a 3' homology arm, a first stuffer sequence, a second stuffer sequence and a cargo, wherein the 5' homology arm is homologous to a sequence upstream of the target nucleic acid sequence, the 3' homology arm is homologous to a sequence downstream of the target nucleic acid sequence, and the DNA donor template comprises, in the 5' to 3' direction, a 5' homology arm, a first stuffer sequence, a cargo, a second stuffer sequence and a 3' homology arm A composition comprising.

28. The composition according to claim 27, wherein the gRNA molecule comprises the sequences of SEQ ID NOs: 216 to 475.

29. The composition according to claim 27 or 28, further comprising a nucleic acid encoding a CRISPR1 (Cpf1) RNA-guided nuclease or a Cpf1 RNA-guided nuclease from the genus Prevotella and the genus Francisella.

30. The composition according to claim 29, for use in the treatment of a subject suffering from cancer.

31. A medicament for treating or alleviating symptoms in a subject, the medicament comprising the composition according to claim 29.