Compositions and methods for manipulating cells
By employing engineered dual-guide CRISPR-Cas systems to modify CAR T cells, the challenges of immunogenicity and specificity in current CAR T-cell therapies are addressed, resulting in reduced graft-versus-host disease and improved therapeutic efficacy.
Patent Information
- Application Number
- JP2024568584
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-16
- Filing Date
- 2023-05-16
- Publication Date
- 2025-06-17
AI Technical Summary
Current CAR T-cell therapies face challenges in distinguishing between host and foreign cells, leading to potential graft-versus-host disease and adverse immune responses.
The development of engineered non-natural dual-guide CRISPR-Cas systems that incorporate protecting groups, donor template recruitment sequences, and editing enhancers to specifically modify genomic DNA in cells, thereby reducing immunogenicity and improving CAR T-cell specificity.
This approach enables the generation of CAR T cells with reduced immunogenicity, minimizing graft-versus-host disease and enhancing the therapeutic efficacy of CAR T-cell therapies by improving their ability to target specific antigens while avoiding host cells.
Smart Images

Figure 2025518552000001_ABST
Abstract
Description
Technical Field
[0001] Description of Research and Development Funded by the Federal Government None
[0002] Cross - Reference to Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 342,472, filed May 16, 2022, the disclosure of which is hereby incorporated by reference in its entirety for all purposes.
Background Art
[0003] The immune system recognizes specific antigen patterns on the cell surface, for example, in human, human leukocyte antigen (HLA) proteins. These protein patterns help the immune system distinguish "host" cells from "foreign" cells, that is, distinguish "self" from "non - self". Chimeric antigen receptor (CAR) T cells are T cells that have been genetically engineered to produce an artificial T - cell receptor for use in one or more therapies. The receptor is "chimeric" because it combines both antigen - binding and T - cell activation functions in a single receptor. Since recognizing a "host" cell as "foreign" can trigger an immune response, that is, "graft - versus - host disease" (GvH), and the therapy can show adverse and / or harmful effects on the recipient, T cells containing CARs engineered for therapeutic purposes also need to distinguish "self" from "non - self". Therefore, considering many disease targets, there is a continuing need to develop CARs that show suitable binding to the target antigen of interest with little or no binding to host antigens.
[0004] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0005] The novel features of the present invention are particularly set forth in the appended claims. A further understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description which describes exemplary embodiments in which the principles of the invention are utilized, and to the appended drawings.
Brief Description of the Drawings
[0006]
Figure 1A
Figure 1B
Figures 2A-C
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figures 8A-B
Figure 9
[0007] Overview I. Cells Containing Engineered Chimeric Antigen Receptors (CARs) A. CAR B. Polynucleotide Encoding a Polypeptide and / or a Polypeptide Containing a CAR or a Portion Thereof C. Cells Containing a Polynucleotide Encoding a Polypeptide and / or a Polypeptide Containing a CAR or a Portion Thereof D. Populations of Cells Containing a CAR E. Compositions and / or Kits for Engineering Cells to Contain a CAR F. Methods for Engineering Cells to Contain a CAR II. Engineered Non-Natural Dual-Guide CRISPR-Cas Systems A. Cas Protein B. Guide Nucleic Acid C. gNA Modification III. Compositions and Methods for Targeting, Editing, and / or Modifying Genomic DNA A. Ribonucleoprotein (RNP) Delivery and “cas RNA” Delivery B. CRISPR Expression Systems C. Donor Template D. Efficiency and Specificity E. Multiplexing F. Genomic Safe Harbors V. Therapeutic Use A. Gene Therapy VI. Kits VII. Embodiments VIII. Examples IX. Equivalents
[0008] I. Cells Comprising an Engineered Chimeric Antigen Receptor (CAR) Compositions, methods, and / or kits are provided herein that comprise a polynucleotide encoding a polypeptide comprising a CAR or a portion thereof and / or a polypeptide comprising a CAR or a portion thereof. Further provided herein are compositions, methods, and / or kits that comprise a polynucleotide encoding a polypeptide comprising a CAR or a portion thereof and / or a polypeptide comprising a CAR or a portion thereof, and / or a cell comprising such a polypeptide, e.g., a progeny of a cell comprising a polypeptide comprising a CAR or a portion thereof. Further provided herein are compositions, methods, and / or kits for generating and / or using a cell, cell population, and / or plurality of cell populations that comprise a polynucleotide encoding a polypeptide comprising a CAR or a portion thereof and / or a polypeptide comprising a CAR or a portion thereof. Any suitable CAR can be used. In certain embodiments, the CAR comprises a CAR as described in the CAR section below. Any suitable cell can be used. In certain embodiments, the cell comprises a human cell, e.g., a human immune cell, e.g., a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, lymphocyte, or a combination thereof, preferably a T cell, and / or a human stem cell, e.g., a human pluripotent, multipotent stem cell, embryonic stem cell, induced pluripotent stem cell, hematopoietic stem cell, CD34+ cell, or a combination thereof, preferably a hematopoietic stem cell, more preferably a CD34+ stem cell, even more preferably an induced pluripotent stem cell (iPSC).
[0009] In certain embodiments, the cells include allogeneic cells. Any suitable allogeneic cells can be used. As used herein, the term "allogeneic" includes cells derived from the same species that are genetically different and thus not immunologically compatible with the host. In certain embodiments, the allogeneic cells include one or more genetic recombinations that reduce the immunogenicity of the cells in the host, e.g., the recipient. In certain embodiments, the cells can be engineered to include one or more genomic modifications. In certain embodiments, the cells can be engineered to include one or more genomic modifications that reduce the immunogenicity of the cells, e.g., the modified cells elicit little or no immune response in vitro and / or in vivo. In certain embodiments, the allogeneic cells for a host (recipient, patient, or suitable alternative) can be engineered to include one or more genomic modifications that reduce the immunogenicity of the one or more allogeneic cells in the host. In certain embodiments, the cells can be engineered to cause 90, 80, 70, 60, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1% or less of the immune response compared to an unengineered equivalent. In certain embodiments, the cells can be engineered to not cause an immune response in the host. The immune response can be measured using any suitable technique, e.g., flow cytometry or ELISA.
[0010] In certain embodiments, the cells comprise one or more genomic modifications that reduce the immunogenicity of the cells. Any suitable genomic modification and / or combination of genomic modifications that reduce the immunogenicity of the cells can be used. In certain embodiments, the cells comprise (1) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of the HLA-1 protein, (2) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of the HLA-2 protein or transcription factors that regulate the expression of one or more subunits of the HLA-2 protein, and / or (3) one or more genomic modifications that partially or completely inactivate one or more genes encoding subunits of the TCR protein. In preferred embodiments, the cells comprise all three genomic modifications. In certain embodiments, one or more of the genomic modifications completely inactivate one or more genes. In certain embodiments, one or more of the genomic modifications at least partially or completely eliminate the surface expression of the active (immunogenic) protein. In certain embodiments, one or more of the genomic modifications completely eliminate the surface expression of the active (immunogenic) protein. In certain embodiments, the cells comprising one or more genomic modifications can further comprise one or more additional modifications including, but not limited to, the introduction of one or more heterologous genes, such as transgenes. The one or more transgenes can be introduced at any suitable location within the genome. In certain embodiments, the one or more transgenes are introduced into a safe harbor site (SHS), such as a safe harbor, as described in the following genomic safe harbor sections. In certain embodiments, the one or more transgenes are introduced into one or more of the sites comprising genomic modifications (1)-(3), for example, a CAR transgene can be introduced into one or more genes encoding subunits of the TCR protein, such as the TRAC gene, and / or a B2M-HLA-E and / or B2M HLA-G fusion protein can be introduced into one or more genes encoding subunits of the HLA-1 protein, such as the B2M gene.
[0011] Cells can be engineered using any suitable compositions and methods. In certain embodiments, cells can be engineered by delivering to the cell a composition comprising a site-specific nuclease and / or one or more polynucleotides encoding a site-specific nuclease. The site-specific nuclease can be any suitable nuclease, such as a homing endonuclease, TALEN, meganuclease, Argonaute, and / or a CRISPR / Cas nuclease, i.e., a nucleic acid-guided nuclease. In preferred embodiments, the site-specific nuclease comprises a nucleic acid-guided nuclease. The site-specific nuclease can hydrolyze the backbone, i.e., generate one or more cleavages or strand breaks in a polynucleotide, such as within the genome, at or near the recognition site of the nuclease, i.e., the target site. One or more strand breaks in at least one strand of the polynucleotide can be repaired via any suitable natural cellular repair mechanism, such as non-homologous end joining (NHEJ) and / or homologous recombination repair (HDR). In certain embodiments, repair of one or more strand breaks in at least one strand of the polynucleotide by NHEJ results in one or more genomic modifications, such as insertions and / or deletions (INDELS). In addition or alternatively, one or more portions of heterologous DNA, such as a donor template, can be introduced into the cell, and at least a portion of the heterologous DNA can be inserted by the cell at or near one or more strand breaks in the DNA by HDR.
[0012] In certain embodiments, the site-specific nuclease includes a nucleic acid-guided nuclease, such as a CRISPR / Cas nuclease. Any suitable nucleic acid-guided nuclease can be used to generate one or more strand breaks in a target polynucleotide. In certain embodiments, the nucleic acid-guided nuclease includes one or more engineered, non-native components. In certain embodiments, the nucleic acid-guided nuclease includes a class 1 or class 2 Cas nuclease, such as type V-A, V-B, V-C, V-D, or V-E. In certain embodiments, the nucleic acid-guided nuclease includes a MAD, ART, or ABW nuclease, such as MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, MAD20, ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11 * , ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART29, ART30, ART31, ART32, ART33, ART34, and / or ART35 nuclease, MAD2, MAD7, ART11, ART11 * , or an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of an ART2 nuclease. In a more preferred embodiment, the nucleic acid-guided nuclease is MAD2, MAD7, ART2, ART11, or ART11 *An amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence, and more preferably, an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of SEQ ID NO: 37. In certain embodiments, the nucleic acid-guided nuclease comprises one or more nuclear localization signals (NLSs), e.g., 1, 4, or 5 nuclear localization signals, e.g., 1-5 NLSs at the carboxy terminus, 1-5 NLSs at the amino terminus, or a combination thereof, preferably, 1 N-terminal NLS and 3 C-terminal NLSs, more preferably, 5 N-terminal NLSs. Any suitable NLS sequence, e.g., SEQ ID NOs: 40-56, preferably any one of SEQ ID NOs: 40, 51, and 56 can be used. Any suitable combination of NLS sequences can be used. Additional nucleases and their modifications are found in the Cas protein section below.
[0013] In certain embodiments, the nucleic acid-guided nuclease further comprises a guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid, e.g., a dual guide nucleic acid. In certain embodiments, the targeter nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence. In certain embodiments, the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence, and optionally, a 5' sequence. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In certain embodiments, the guide nucleic acid comprises an engineered non-natural guide nucleic acid. In certain embodiments, the guide nucleic acid comprises a dual guide nucleic acid (as described in the guide nucleic acid section below), where the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments where the guide nucleic acid is a dual guide nucleic acid, the stem of the targeter nucleic acid and the stem of the modulator nucleic acid hydrolyze. In certain embodiments, the dual guide nucleic acid can bind to and activate the nucleic acid-guided nuclease, which is activated by a single crRNA in the absence of tracrRNA in the native system.
[0014] In certain embodiments, the guide nucleic acid has one or more chemical modifications to one or more nucleotides and / or internucleotide linkages, such as 2'-O-alkyl, 2'-O-methyl, phosphorothioate, phosphonoacetate, thiophosphonoacetate, 2'-O-methyl-3'-phosphorothioate, 2'-O-methyl-3'-phosphonoacetate, 2'-O-methyl-3'-thiophosphonoacetate, 2'-deoxy-3'-phosphonoacetate, 2'-deoxy-3'-thiophosphonoacetate, or combinations thereof, at and / or near the 5' end, 3' end, and / or both ends as described in the gNA modification section below.
[0015] In certain embodiments, one or more guide nucleic acids can form a complex with one or more nucleases, such as a nucleic acid-guided nuclease complex. In certain embodiments, one or more guide nucleic acids, one or more nucleic acid-guided nucleases, and / or one or more nucleic acid-guided nucleases can further comprise one or more additives that stabilize the nucleic acid-guided nuclease complex.
[0016] Such cells and / or populations of cells comprising a CAR can be used for various purposes, and one such purpose can be CAR T cells.
[0017] A.CAR In certain embodiments, compositions, methods, and / or kits are provided herein that include a CAR or a portion thereof (Figures 3(302) and (303)). In certain embodiments, compositions, methods, and / or kits are provided herein that include a dual CAR, e.g., a CAR fusion protein or two separate CARs (Figure 3(304) and Figure 4). As used herein, the term "dual CAR" includes a polypeptide that includes a first CAR or a portion thereof and a second CAR or a portion thereof that are either separate or joined via one or more polypeptide linkers. In certain embodiments, the second CAR or a portion thereof targets the same antigen as the first CAR or a portion thereof. In certain embodiments, the second CAR or a portion thereof targets a different antigen than the first CAR or a portion thereof. Further, polypeptides are disclosed herein that include several CARs or portions thereof that are either separate or joined via one or more polypeptide linkers. In certain embodiments, a cell can include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 and / or 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 or fewer CARs or portions thereof, e.g., 1 - 15, preferably 1 - 10, more preferably 2 - 10, even more preferably 2 - 7, even more preferably 2 - 5 CARs or portions thereof, that are either separate or joined via one or more polypeptide linkers. The polypeptide linker can include any suitable linker that includes natural or non-natural amino acids. In certain embodiments, the CAR or a portion thereof is expressed on the cell surface. In addition or alternatively, the CAR or a portion thereof is secreted.
[0018] In certain embodiments, the CAR or a portion thereof comprises a polypeptide that binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, CD3ζ, a portion thereof, or a combination thereof, preferably B7H3, BCMxA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3ζ. In certain embodiments, the CAR or a portion thereof is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124 or 2044-2070 in Table 1, preferably at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104, 116-124, or 2044-2070 in Table 1, even more preferably at least 95% identical, even more preferably at least 99% identical polypeptide.
[0019] In certain embodiments, provided herein are compositions, methods, and / or kits comprising a polynucleotide encoding a polypeptide comprising the CAR or a portion thereof. The polynucleotide can be inserted at any suitable location within the genome, such as a safe harbor site (as described in the section on genomic safe harbors below) and / or within a suitable gene such as the TRAC gene. In certain embodiments, the polynucleotide can comprise a first portion comprising a first CAR or a portion thereof and a second portion comprising a second CAR or a portion thereof. The first portion and the second portion of the polynucleotide can be expressed separately, for example, a first mRNA transcript encoding the first CAR or a portion thereof can be produced and a second mRNA transcript encoding the second CAR or a portion thereof can be produced. Alternatively, or additionally, the first portion and the second portion can be expressed as a single mRNA transcript, for example, comprising an internal ribosome entry site. Any suitable genomic construct for expression can be used in combination with any suitable number of CARs and / or dual CARs.
[0020] The CAR or dual CAR or a portion thereof can include any suitable form, such as an antibody, nanobody (VHH), Fab, scFv, diabody, triabody, minibody, single-domain antibody, first-generation CAR, second-generation CAR, third-generation CAR, and / or fourth-generation CAR. Any suitable antigen-binding domain can be combined with any suitable combination of additional CAR domains as needed for the application.
[0021] In certain embodiments, the CAR or a portion thereof can be an antibody, an antigen-binding fragment of an antibody, or a fusion protein derived from such an antibody, such as a single-chain variable fragment (scFv). According to these embodiments herein, the scFv can include a single-chain Fv antibody in which the variable domains of the heavy and light chains of a conventional two-chain antibody are joined so as to form a single polypeptide chain. In some embodiments, the single-chain antibodies contemplated herein can be derived from any species including human or an animal (e.g., mouse, rabbit, pig, dog, cow, horse, goat, camel, or other animal). In some embodiments, the intracellular signaling domain can include a signaling domain and a co-stimulatory domain.
[0022] In other embodiments, the CAR or a portion thereof further comprises a spacer domain that links the antigen-binding domain to the transmembrane domain. In some embodiments, a spacer domain of appropriate length can improve the mobility and flexibility of the antigen-binding domain to allow for optimal binding to the target antigen. In certain embodiments, the spacer domain comprises at least a portion or segment of the hinge region of IgGl, IgG2, IgG3, or IgG4. In some embodiments, the spacer domain can be derived from the CH2 region and / or CH3 region of IgGl, IgG2, IgG3, or IgG4. In other embodiments, the spacer domain comprises upper hinge amino acids found between the variable heavy chain and the core, and core hinge amino acids including a polyproline region. In other embodiments, the spacer region comprises at least a portion of the hinge region of a human IgG4 hinge spacer. In some embodiments, the spacer region comprises a human IgG4 hinge-CH3 spacer.
[0023] In some embodiments, the CAR or a portion thereof further comprises a transmembrane domain. The transmembrane domain can provide for the anchoring of the CAR in the cell membrane. In some embodiments, the transmembrane domain comprises a membrane-bound or transmembrane protein. In other embodiments, the transmembrane domain comprises the transmembrane region of the α, β, or ζ chain of a T cell receptor, such as CD28, CD3, CD45, CD4, CD8, CD8a CD9, CD16, CD22, CD33, CD37, CD64, CD80, CD86, CD134, CD137, or CD154. In some embodiments, the transmembrane domain comprises the CD28 transmembrane domain (CD28tm).
[0024] In other embodiments, the CAR or a portion thereof further comprises an intracellular signaling domain linked to the transmembrane domain. According to these embodiments, the intracellular signaling domain can activate the function of the cell when the antigen-binding domain binds to the target antigen. In some embodiments, the intracellular signaling domain can activate the function of a cell expressing the CAR, such as a T cell expressing the CAR. In other embodiments, the intracellular signaling domain comprises one or more intracellular signaling domains. In some embodiments, the intracellular signaling domain comprises a functional domain of a primary cytoplasmic signaling protein. In some embodiments, the intracellular signaling domain comprises a functional domain of a primary cytoplasmic signaling protein and at least one functional domain of one or more secondary cytoplasmic signaling proteins. In certain embodiments, the stimulatory primary cytoplasmic signaling protein can contain a signaling motif known as an immunoreceptor tyrosine-based activation motif (ITAM). According to these embodiments, examples of ITAMs containing a primary cytoplasmic signaling domain for use herein include, but are not limited to, those derived from CD3ζ, FcRγ, CD3γ, CD3δ, CD3ε, CD5, CD22, CD79a, CD79b, or CD66d. In some embodiments, the intracellular signaling domain and / or co-stimulatory domain herein can include all or a biologically active fragment of a ligand that specifically binds to CD27, CD28, 4-1BB, OX40, CD30, CD40, ICOS, lymphocyte function-associated antigen-1 (LFA-1), CD2, CD7, LIGHT, NKG2C, or B7H3, and / or CD83. In some embodiments, the intracellular signaling domain herein can include all or a biologically relevant segment of the signaling domain of CD3-ζ or a variant thereof and all or a part of the signaling domain of 4-1BB or a variant thereof.
[0025]
Table 1
[0026]
Table 2
[0027]
Table 3
[0028]
Table 4
[0029]
Table 5
[0030]
Table 6
[0031]
Table 7
[0032] B. A polynucleotide encoding a polypeptide and / or a polypeptide containing a CAR or a portion thereof In certain embodiments, compositions comprising a polypeptide are provided herein. In certain embodiments, the polypeptide comprises a CAR or a portion thereof. In certain embodiments, the polypeptide is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of SEQ ID NOs: 86-124 or 2044-2070, preferably at least 90% identical, more preferably at least 95% identical, even more preferably at least 99% identical, even more preferably 99.5% identical, even more preferably 100% identical, and comprises a sequence. In a more preferred embodiment, the polypeptide is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of SEQ ID NOs: 86-104, 116-124, or 2044-2070, preferably at least 90% identical, more preferably at least 95% identical, even more preferably at least 99% identical, even more preferably 99.5% identical, even more preferably 100% identical, and comprises a sequence.
[0033] In certain embodiments, compositions are provided herein that include polynucleotides encoding polypeptides. In certain embodiments, the polynucleotide encodes a polypeptide that includes a CAR or a portion thereof. In certain embodiments, the polynucleotide is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of SEQ ID NOs: 86-124 or 2044-2070, preferably at least 90% identical, more preferably at least 95% identical, even more preferably at least 99% identical, even more preferably 99.5% identical, even more preferably 100% identical, and encodes a polypeptide comprising a sequence. In a more preferred embodiment, the polynucleotide is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of SEQ ID NOs: 86-104, 116-124, or 2044-2070, preferably at least 90% identical, more preferably at least 95% identical, even more preferably at least 99% identical, even more preferably 99.5% identical, even more preferably 100% identical, and encodes a polypeptide comprising a sequence.
[0034] In certain embodiments, the polynucleotide encodes two or more polypeptides. In certain embodiments, the polynucleotide comprises a first portion and a second portion, where the first portion encodes a first polypeptide comprising a first CAR or a portion thereof, and the second portion encodes a second polypeptide comprising a second CAR or a portion thereof. The polynucleotide can encode any suitable number of polypeptides, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 and / or 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 or fewer polypeptides, such as 1-15, preferably 1-10, more preferably 2-10, even more preferably 2-7, even more preferably 2-5. In certain embodiments, the polypeptides are distinct. In certain embodiments, the polypeptides are linked. Any suitable linker can be used, for example, the linker can comprise one or more amino acids or polypeptides, such as a self-cleaving polypeptide. The polypeptide linker can comprise any suitable linker comprising natural or non-natural amino acids.
[0035] In certain embodiments, the first CAR or a portion thereof binds to a first site of a first binding partner, such as an antigen or epitope, and the second CAR or a portion thereof binds to a second site of a second binding partner. In certain embodiments, the first and second sites are the same site on the same binding partner. In certain embodiments, the first and second sites are different sites on the same binding partner. In certain embodiments, the first site binds to a site on the first binding partner and the second site binds to a site on the second binding partner.
[0036] C. Cells comprising a polynucleotide encoding a polypeptide and / or a polypeptide comprising a CAR or a portion thereof In certain embodiments, provided herein are compositions comprising cells comprising a first polynucleotide encoding a first polypeptide. The first polynucleotide can be inserted at any suitable location in the genome, such as a safe harbor site or a gene, such as the TRAC gene. In certain embodiments, the first polynucleotide encoding the first polypeptide comprises a first CAR or a portion thereof. The first CAR or a portion thereof can comprise any suitable CAR or a portion thereof as described in the CAR section above. The first polypeptide comprising the CAR or a portion thereof can be expressed on the surface of the cell or secreted. In a preferred embodiment, the first polypeptide comprising the CAR or a portion thereof is expressed on the surface of the cell. Exemplary examples are shown in FIG. 3. Specifically, FIG. 3 shows a first cell comprising a first polypeptide comprising a first CAR or a portion thereof (302) and a second cell comprising a second polypeptide comprising a second CAR or a portion thereof (303), wherein the second CAR or a portion thereof is different from the first CAR or a portion thereof. The cell can comprise any suitable number of polypeptides comprising the CAR or a portion thereof, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, or 9 and / or 10, 9, 8, 7, 6, 5, 4, 3, or 2 or fewer CARs, for example, 1 to 10 polypeptides or portions thereof, preferably 1 to 5 polypeptides or portions thereof, more preferably 1 to 5 polypeptides or portions thereof, even more preferably 1 to 3 polypeptides or portions thereof. It should be understood that the cells shown comprise polynucleotides encoding the first and / or second polypeptide or a portion thereof.
[0037] In certain embodiments, the first polynucleotide further encodes a second polypeptide comprising a second CAR or a portion thereof, such as a dual CAR. The first CAR or a portion thereof can comprise any suitable CAR or a portion thereof, as described in the CAR sections above. The first and / or second polypeptide can be expressed on the surface of the cell or secreted. In certain embodiments, the first and second polypeptides comprise the same CAR or a portion thereof. In a preferred embodiment, the second CAR or a portion thereof is different from the first CAR or a portion thereof.
[0038] In certain embodiments, a cell comprising the first polynucleotide further comprises a second polynucleotide encoding a second polypeptide comprising a second CAR or a portion thereof, such as a dual CAR. The second CAR or a portion thereof can comprise any suitable CAR or a portion thereof, as described in the CAR sections above. The second CAR or a portion thereof can be expressed on the surface of the cell or secreted. In certain embodiments, the first and second polypeptides comprise the same CAR or a portion thereof. In a preferred embodiment, the second CAR or a portion thereof is different from the first CAR or a portion thereof.
[0039] Exemplary examples of cells comprising the first and second CARs or portions thereof (304) are shown in FIG. 3.
[0040] In certain embodiments, a polypeptide comprising a CAR or a portion thereof comprises a dual CAR comprising a first CAR or a portion thereof linked via a polypeptide linker to a second CAR or a portion thereof. The first and second CARs or portions thereof can comprise any suitable CAR or portion thereof, as described in the CAR sections above. The dual CAR can be expressed on the surface of a cell or secreted. In certain embodiments, the first and second CARs or portions thereof comprise the same CAR or portion thereof. In preferred embodiments, the second CAR or portion thereof is different from the first CAR or portion thereof. An exemplary example of a cell comprising a dual CAR linked via a polypeptide linker is shown in FIG. 4. Specifically, FIG. 4 shows a cell surface, e.g., a cell membrane, separating an intracellular (401) and extracellular (402) space, comprising a first polypeptide comprising a first dual CAR (403 and 404) expressed on the surface of the cell and a second polypeptide comprising a second dual CAR (405 and 406), where the first dual CAR comprises a second CAR or portion thereof (404) linked to a first CAR or portion thereof (403), and the second dual CAR comprises a first CAR or portion thereof (403) linked to a second CAR or portion thereof (404). The cell can comprise any suitable number of polypeptides comprising the dual CAR, e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, or 9 and / or 10, 9, 8, 7, 6, 5, 4, 3, or 2 or fewer polypeptides, e.g., 1 to 10 polypeptides, preferably 1 to 5 polypeptides, more preferably 1 to 5 polypeptides, even more preferably 1 to 3 polypeptides. The dual CAR can comprise any suitable combination of CARs or portions thereof. Further dual CARs, e.g., second, third, etc. dual CARs, can comprise any suitable combination of CARs or portions thereof different from the first and / or further dual CARs.
[0041] In certain embodiments, the first and / or second CAR or portions or dual CAR thereof, and the first and / or second polypeptides containing the same, can be secreted. Exemplary examples are shown in FIG. 5. Specifically, FIG. 5 shows a cell surface, for example, a cell membrane-separated intracellular space (501) and extracellular space (502), including a first polypeptide containing a first CAR or a portion thereof (503) expressed on the cell surface and a second polypeptide containing a second CAR or a portion thereof (504) expressed in the intracellular space (501) of the cell and secreted into the extracellular space (502) of the cell. The cell can contain any suitable number of secreted polypeptides containing the CAR or a portion thereof, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, or 9 and / or 10, 9, 8, 7, 6, 5, 4, 3, or 2 or fewer secreted polypeptides, for example, 1 to 10 secreted polypeptides, preferably 1 to 5 secreted polypeptides, more preferably 1 to 5 secreted polypeptides, and even more preferably 1 to 3 secreted polypeptides.
[0042] In certain embodiments, the cell includes a human cell, for example, a human immune cell, for example, a neutrophil, eosinophil, basophil, mast cell, monocyte, macrophage, dendritic cell, natural killer cell, lymphocyte, or a combination thereof, preferably a T cell, and / or a human stem cell, for example, a human totipotent, pluripotent stem cell, embryonic stem cell, induced pluripotent stem cell, hematopoietic stem cell, CD34+ cell, or a combination thereof, preferably a hematopoietic stem cell, more preferably a CD34+ stem cell, and even more preferably an induced pluripotent stem cell (iPSC). In certain embodiments, the cell includes allogeneic cells.
[0043] A cell population containing D.CAR In certain embodiments where the cell comprises two or more polynucleotides, provided herein is a composition comprising one or more cell populations comprising a first polynucleotide encoding a first polypeptide comprising a first CAR or portion thereof and / or a dual CAR and / or a second polynucleotide encoding a second polypeptide comprising a second CAR or portion thereof and / or a dual CAR. In certain embodiments, the composition comprises a single cell population, where each of the cells comprises both the first and second polynucleotides. In certain embodiments, provided herein is a composition comprising a plurality of cell populations, where each cell population comprises a different set of polynucleotides. Generally, at least one cell population comprises both the first and second polynucleotides in addition to one or more additional cell populations that do not comprise both the first and second polynucleotides. The cell population can comprise any one of the cells as described in the section of cells comprising a polypeptide comprising the CAR or portion thereof above. Exemplary examples are shown in FIG. 3. Specifically, FIG. 3 shows a target cell (301) that does not comprise either the first or second polynucleotide, whereby introduction of the first and second polynucleotides at suitable positions in the genome using suitable genome engineering techniques results in the following four possible cell populations: (301) a cell in which neither the first nor the second polynucleotide was introduced at a suitable position in the genome; (302) a cell in which the first polynucleotide was introduced at a suitable position in the genome, but the second polynucleotide was not introduced at a suitable position in the genome; (303) a cell in which the second polynucleotide was introduced at a suitable position in the genome, but the first polynucleotide was not introduced at a suitable position in the genome; and (304) a cell in which both the first and second polynucleotides were introduced at suitable positions in the genome. Any suitable number of polynucleotides can be introduced into the genome and thus any suitable number of possible cell populations can result.In certain embodiments, the plurality of cell populations comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or 45 and / or 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, or 50 or fewer populations, e.g., a cell population of 1-50, preferably a cell population of 1-20.
[0044] In certain embodiments, each cell population in a plurality of cell populations can be present at any percentage relative to other cell populations, where the relative percentage of each population is influenced by many factors including, but not limited to, the delivery efficiency of the editing component, the quality of the editing component, the concentration of the editing component, the relative efficiency and specificity of the editing event, the viability of the cells, and / or the viability of the cells before or after one or more genomic modifications. In certain embodiments, the first cell population comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 70% and / or 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 75% or less of all of the cells in the plurality of cell populations, for example, 1-75% of all of the cells in the plurality of cell populations, preferably 5-75%, more preferably 10-75%, even more preferably 15-75%, and even still more preferably 20-75%. In certain embodiments, the second cell population comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 70% and / or 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 75% or less of all of the cells in the plurality of cell populations, for example, 1-75% of all of the cells in the plurality of cell populations, preferably 50% or less, more preferably 30% or less, even more preferably 20% or less, and even still more preferably 10% or less. In certain embodiments, the third cell population comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 70% and / or 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 75% or less of all of the cells in the plurality of cell populations, for example, 1-75% of all of the cells in the plurality of cell populations, preferably 50% or less, more preferably 30% or less, even more preferably 20% or less, and even still more preferably 10% or less.In certain embodiments, the fourth cell population comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 70% and / or 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 75% or less, for example, 1-75%, preferably 50% or less, more preferably 30% or less, even more preferably 20% or less, and even still more preferably 10% or less of all the cells in the plurality of cell populations. It is understood that the sum of the percentages of each cell population in the plurality of cell populations is 100%.
[0045] The number, relative abundance, and / or identity of cell populations in the plurality of cell populations can be measured by any suitable method. In certain embodiments, the number, relative abundance, and / or identity of cell populations in the plurality of cell populations can be measured by analyzing one or more nucleic acids in a sample using one or more methods, such as PCR, multiplex PCR, FISH, and / or sequencing. In certain embodiments, the number and / or identity of cell populations in the plurality of cell populations can be measured by analyzing one or more cell surface proteins and / or their absence in a sample using one or more methods, such as immunostaining and microscopy, ELISA, pull-down, and / or flow cytometry.
[0046] Compositions and / or kits for engineering cells to comprise E.CAR In certain embodiments, compositions are provided herein that include a guide nucleic acid, a nucleic acid-guided nuclease, a nucleic acid-guided nuclease complex, and / or one or more polynucleotides encoding the same. In certain embodiments, the composition further includes a donor template. In certain embodiments, the composition further includes an additive that stabilizes the nucleic acid-guided nuclease complex, such as poly-L-glutamic acid. In certain embodiments, one or more components of the composition are combined in the presence of an aqueous buffer. In certain embodiments, the composition further includes an excipient. In certain embodiments, the composition is dehydrated, for example, lyophilized or freeze-dried.
[0047] In certain embodiments, compositions are provided herein that include a first nucleic acid-guided nuclease, a first guide nucleic acid, and / or a first donor template, comprising a first polynucleotide encoding a first CAR or a portion thereof. In certain embodiments, the composition further includes a second nucleic acid-guided nuclease, a second guide nucleic acid, and / or a second donor template, comprising a second polynucleotide encoding a second CAR or a portion thereof. In certain embodiments, the composition further includes a third nucleic acid-guided nuclease, a third guide nucleic acid, and / or a third donor template, comprising a third polynucleotide encoding a third CAR or a portion thereof. Any suitable number of additional nucleic acid-guided nucleases, guide nucleic acids, and / or donor templates can be used depending on the number of genomic modifications to be introduced. In other words, the composition can include a nucleic acid-guided nuclease (N) x , a guide nucleic acid (gNA) x , and / or a donor template (D) xcan include, where the integer (x) represents a genomic modification to be introduced at a suitable position in the genome. For example, for a single genomic modification (x = 1), the composition includes a first nucleic acid-guided nuclease (N)1, a first guide nucleic acid (gNA)1, and / or a first donor template (D)1. For multiple genomic modifications, x can be any suitable integer, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, or 40 and / or 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, or 2 or less, where the composition includes a suitable number of each component. In certain embodiments, multiple genomic modifications, two or more, can be generated sequentially. In certain embodiments, multiple genomic modifications can be generated simultaneously. The method is further described in the section of the method for engineering a cell to include the following CARs.
[0048] The nucleic acid-guided nuclease can be any suitable nuclease, for example, a homing endonuclease, a TALEN, a meganuclease, an Argonaute, and / or a CRISPR / Cas nuclease, preferably a CRISPR / Cas nuclease, more preferably a class 1 or class 2 Cas nuclease, even more preferably a class 2 nuclease, even more preferably a type V-A, V-B, V-C, V-D, or V-E nuclease, even more preferably a type V-A nuclease. In certain embodiments, the nucleic acid-guided nuclease is a MAD, ART, or ABW nuclease, for example, MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, MAD20, ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11, ART11 *, ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART29, ART30, ART31, ART32, ART33, ART34, and / or ART35 nuclease, MAD2, MAD7, ART11, ART11 * , or an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of ART2 nuclease, preferably MAD2, MAD7, ART2, ART11, or ART11 * and an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence thereof, more preferably an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of SEQ ID NO: 37. In certain embodiments, the nucleic acid-inducible nuclease comprises one or more nuclear localization signals (NLSs), e.g., 1, 4, or 5 nuclear localization signals, e.g., 1-5 NLSs at the carboxy terminus, 1-5 NLSs at the amino terminus, or combinations thereof, preferably 1 N-terminal NLS and 3 C-terminal NLSs. In certain embodiments, the NLS comprises any one of SEQ ID NOs: 40-56, preferably SEQ ID NOs: 40, 51, and 56. Additional nucleases and their modifications are found in the Cas protein section below.
[0049] The guide nucleic acid can be any gNA that is compatible with the nucleic acid-guided nuclease. In certain embodiments, the guide nucleic acid comprises a targeter nucleic acid and a modulator nucleic acid, such as a dual guide nucleic acid. In certain embodiments, the targeter nucleic acid comprises a targeter nucleic acid comprising a targeter stem sequence and a spacer sequence. In certain embodiments, the modulator nucleic acid comprises a modulator stem sequence complementary to the targeter stem sequence, and optionally, a 5' sequence. In certain embodiments, the guide nucleic acid comprises a single polynucleotide. In certain embodiments, the guide nucleic acid comprises an engineered non-natural guide nucleic acid. In a preferred embodiment, the guide nucleic acid comprises a dual guide nucleic acid (as described in the section on guide nucleic acids below), where the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. In certain embodiments where the guide nucleic acid is a dual guide nucleic acid, the stems of the targeter nucleic acid and the modulator nucleic acid hydrolyze. In certain embodiments, the dual guide nucleic acid is capable of binding to and activating the nucleic acid-guided nuclease, which in the native system is activated by a single crRNA in the absence of tracrRNA. In certain embodiments, the guide nucleic acid comprises a donor recruit sequence. The gNA comprises any suitable spacer sequence. In certain embodiments, the gNA comprises a spacer sequence as shown in Table 2.
[0050]
Table 8
[0051]
Table 9
[0052]
Table 10
[0053]
Table 11
[0054]
Table 12
[0055]
Table 13
[0056]
Table 14
[0057]
Table 15
[0058]
Table 16
[0059]
Table 17
[0060]
Table 18
[0061]
Table 19
[0062]
Table 20
[0063]
Table 21
[0064] [Table 22]
[0065] [Table 23]
[0066] [Table 24]
[0067] [Table 25]
[0068] [Table 26]
[0069] [Table 27]
[0070] [Table 28]
[0071] [Table 29]
[0072] [Table 30]
[0073] [Table 31]
[0074]
Table 32
[0075]
Table 33
[0076]
Table 34
[0077]
Table 35
[0078]
Table 36
[0079]
Table 37
[0080]
Table 38
[0081]
Table 39
[0082]
Table 40
[0083]
Table 41
[0084]
Table 42
[0085]
Table 43
[0086]
Table 44
[0087]
Table 45
[0088]
Table 46
[0089]
Table 47
[0090]
Table 48
[0091]
Table 49
[0092]
Table 50
[0093]
Table 51
[0094]
Table 52
[0095]
Table 53
[0096]
Table 54
[0097]
Table 55
[0098]
Table 56
[0099]
Table 57
[0100]
Table 58
[0101]
Table 59
[0102]
Table 60
[0103]
Table 61
[0104]
Table 62
[0105]
Table 63
[0106] [Table 64]
[0107] [Table 65]
[0108] [Table 66]
[0109] In certain embodiments, the guide nucleic acid has one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at and / or near the 5′ end, 3′ end, and / or both ends, as described in the gNA modification section below, such as 2′-O-alkyl, 2′-O-methyl, phosphorothioate, phosphonoacetate, thiophosphonoacetate, 2′-O-methyl-3′-phosphorothioate, 2′-O-methyl-3′-phosphonoacetate, 2′-O-methyl-3′-thiophosphonoacetate, 2′-deoxy-3′-phosphonoacetate, 2′-deoxy-3′-thiophosphonoacetate, or combinations thereof.
[0110] In certain embodiments, one or more guide nucleic acids can form a complex with one or more nucleases, and can be, for example, a nucleic acid-guided nuclease complex. In certain embodiments, one or more guide nucleic acids, one or more nucleic acid-guided nucleases, and / or one or more nucleic acid-guided nucleases can further include one or more additives that stabilize the nucleic acid-guided nuclease complex. Exemplary examples of nucleic acid-guided nuclease complexes are shown in FIG. 6. Specifically, FIG. 6 shows a type V-A nucleic acid-guided nuclease (601) complexed with a dual gNA comprising a modulator nucleic acid (606) and a targeter nucleic acid (607), where the modulator nucleic acid and the targeter nucleic acid are hybridized through a stem. The targeter nucleic acid further includes a spacer sequence (605), i.e., a protospacer, that is at least partially complementary to a target nucleotide sequence (604) in a target polynucleotide (602) adjacent to a suitable PAM (603). When bound to the target nucleotide sequence, the nucleic acid-guided nuclease complex can generate one or more strand breaks (608) in the target polynucleotide at or near the target nucleotide sequence.
[0111] Any suitable donor template can be used as described in the Donor Template section below. In certain embodiments, the donor template includes a polynucleotide encoding a polypeptide, where the polypeptide includes a chimeric antigen receptor (CAR) or a portion thereof. The polynucleotide can encode any suitable number of CARs or dual CARs as described above.
[0112] In certain embodiments, the donor template has one or more chemical modifications to one or more nucleotides and / or internucleotide linkages at and / or near the 5′ end, 3′ end, and / or both ends, as described in the section on gNA modifications below, such as 2′-O-alkyl, 2′-O-methyl, phosphorothioate, phosphonoacetate, thiophosphonoacetate, 2′-O-methyl-3′-phosphorothioate, 2′-O-methyl-3′-phosphonoacetate, 2′-O-methyl-3′-thiophosphonoacetate, 2′-deoxy-3′-phosphonoacetate, 2′-deoxy-3′-thiophosphonoacetate, or combinations thereof.
[0113] Method for engineering cells to contain F.CAR In certain embodiments, methods are provided herein. In certain embodiments, methods for engineering cells, such as human cells, are provided herein. In certain embodiments, the methods are for engineering cells to contain one or more CARs or dual CARs. In certain preferred embodiments, a method is provided herein for generating any of the compositions described in the section on compositions and / or kits for engineering cells to contain the above-described CARs, using any of the compositions described in the section on cells or cell populations containing a polypeptide comprising the above-described CAR or a portion thereof. In certain embodiments, a nucleic acid-guided nuclease, guide nucleic acid, nucleic acid-guided nuclease complex, one or more polynucleotides encoding the same, donor template, and / or suitable combinations thereof are delivered into a cell, where one or more components are transferred to the nuclease, thereby generating one or more genome modifications.
[0114] An exemplary method is shown in FIG. 7. Specifically, FIG. 7A shows a nucleic acid-guided nuclease complex (701) and a donor template (702) delivered to a cell (703), where one or more components are transferred to a nuclease (704). In certain embodiments, the nucleic acid-guided nuclease comprises one or more nuclease localization signals (such as those described above and in the section on ribonucleoprotein (RNP) below). In such cases, as shown in FIG. 7B, a nucleic acid-guided nuclease complex comprising one or more nuclear localization signals (705) and a donor template (706) is delivered to a cell (708), where the donor template (706) and the nucleic acid-guided nuclease complex (705) bind together (707) and are transferred into a nuclease (709) assisted by one or more nuclear localization signals.
[0115] In certain embodiments, methods are provided herein for generating one or more modifications in the genome of a target cell. In certain embodiments, the method can generate at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or 100 or fewer genomic modifications, e.g., 1 to 100 genomic modifications, preferably 1 to 20 genomic modifications, simultaneously or sequentially (see the section on multiplex methods below). In certain embodiments, a first genomic modification is introduced into one or more target cells, where the target cells include wild-type cells or cells containing one or more genomic modifications (see the section on cells containing a polypeptide comprising a CAR or a portion thereof or a cell population comprising a CAR). In certain embodiments, the target cells include one or more of the modified cells as described in the section on cells containing a polypeptide comprising a CAR or a portion thereof or a cell population comprising a CAR. In certain embodiments, the method includes generating one or more genomic modifications in one or more target cells, where the one or more genomic modifications are generated simultaneously. In certain embodiments, the method includes generating one or more genomic modifications in one or more target cells, where one or more of the genomic modifications are generated sequentially. In certain embodiments where one or more genomic modifications are introduced sequentially, the one or more genomic modifications may be introduced in any suitable amounts, orders, and / or combinations.For example, when introducing three genomic modifications (A, B, and C) into one or more cells, the three genomic modifications can be introduced in any one of the following orders: (1) A, then B, then C; (2) A, then C, then B; (3) A and B, then C; (4) A, then B and C; (5) A and C, then B; (6) A, then C and B; (7) B, then A, then C; (8) B, then C, then A; (9) B and A, then C; (10) B, then A and C; (11) B and C, then A; (12) B, then C and A; (13) C, then A, then B; (14) C, then B, then A; (15) C and A, then B; (16) C, then A and B; (17) C, then B and A; (18) C and B, then A; or (19) A and B and C. An exemplary method is shown in FIG. 8A. Specifically, FIG. 8A shows a first set of components for generating one or more genomic modifications (801) to be delivered to starting cells (802), whereby a first engineered cell or population of cells (803) is generated. The first engineered cell or population of cells (803) can be the final resulting cell or population of cells (804), and / or a further set of components for generating one or more genomic modifications (805) can be delivered to (803), resulting in a second engineered cell or population of cells (806). This process can be repeated as many times as necessary. In other words, a method for introducing a plurality of exogenous nucleic acids into the genome of a target cell, the method comprising contacting the target cell with a composition comprising a nucleic acid-guided nuclease (N). x , a guide nucleic acid (gNA) x , and / or a donor template (D) x is provided herein, wherein one or more components of the composition are capable of initiating at least partial HDR or homologous recombination of (D) at a target site (TS) selected from a plurality of target sites of the genome of the host cell, e.g., x in (D) x , and wherein for each target site, (N) x cleaves at (TS) x to enable (D) xis capable of effecting at least partial homologous recombination, where the integer (x) represents a genomic modification to be introduced at a suitable location in the genome. For example, for a single genomic modification (x = 1), the composition comprises a first nucleic acid-induced nuclease (N)1, a first guide nucleic acid (gNA)1, and / or a first donor template (D)1. For multiple genomic modifications, x can be any suitable integer, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, or 40 and / or 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, 3, or 2 or less, where the composition comprises a suitable number of each component. In certain embodiments, multiple genomic modifications, two or more, can be generated sequentially. In certain embodiments, multiple genomic modifications can be generated simultaneously.
[0116] In certain embodiments, methods for engineering one or more suitable human cells are provided herein. In certain embodiments, the cells comprise one or more suitable human stem cells or human immune cells. In certain embodiments, the cells comprise one or more human cells comprising immune cells including neutrophils, eosinophils, basophils, mast cells, monocytes, macrophages, dendritic cells, natural killer cells, lymphocytes, or combinations thereof. In certain embodiments, the cells comprise one or more T cells. In certain embodiments, the cells comprise one or more chimeric antigen receptor (CAR)-T cells. In certain embodiments, the CAR T cells comprise a CAR or a portion thereof. In certain embodiments, the CAR T cells comprise two or more CAR polypeptides or portions thereof. In certain embodiments, the CAR T cells comprise two CAR polypeptides or portions thereof, wherein the second CAR polypeptide is different from the first CAR polypeptide. In certain embodiments, the CAR T cells comprise a dual CAR or a portion thereof. In certain embodiments, the cells comprise one or more human stem cells including human totipotent, pluripotent stem cells, embryonic stem cells, induced pluripotent stem cells, hematopoietic stem cells, CD34+ cells, combinations thereof. In preferred embodiments, the cells comprise one or more hematopoietic stem cells. In preferred embodiments, the cells comprise one or more CD34+ stem cells. In preferred embodiments, the cells comprise one or more induced pluripotent stem cells (iPSCs). In certain embodiments, the cells comprise allogeneic cells.
[0117] In certain embodiments, one or more cells comprising one or more introduced genomic modifications are grown, e.g., proliferated, or differentiated, e.g., iPSCs are differentiated into T cells. In certain embodiments where two or more genomic modifications are introduced sequentially, one or more target cells are grown after the first set of genomic modifications, where the second set of genomic modifications is introduced into the progeny of the first set of cells. In certain embodiments, the stem cells are differentiated before or after the introduction of one or more genomic modifications. In certain embodiments, the stem cells are differentiated after the introduction of one or more genomic modifications. An exemplary method is shown in FIG. 8B. Specifically, the cells can be grown or proliferated (807) after the introduction of the first set of genomic modifications, where the second set of genomic modifications is introduced into the progeny of the cell or cell population. In certain embodiments, the cells can be differentiated (808) at any stage of the manipulation process, e.g., iPSCs can be differentiated into immune cells after receiving multiple genomic modifications.
[0118] In certain embodiments, one or more genomic modifications are introduced into a population of cells, where the resulting cell population comprises multiple cell populations each having received a different set of genomic modifications (see section on cell populations comprising CARs above). For example, when three genomic modifications (A, B, C) are introduced into a population of cells, either sequentially and / or simultaneously, the resulting multiple cell populations can comprise any number and / or combination of the following cell populations: (1) A, (2) AB, (3) AC, (4) ABC, (5) B, (6) BC, (7) C, and / or (8) no genomic modification. In certain embodiments, each cell population in the multiple cell populations can be present at any percentage relative to the other cell populations, where the relative percentage of each population is influenced by many factors including, but not limited to, delivery efficiency of the editing component, quality of the editing component, concentration of the editing component, relative efficiency and specificity of the editing event, viability of the cells, and / or viability of the cells before or after one or more genomic modifications.
[0119] In certain embodiments, provided herein are methods for manipulating cells, the methods comprising delivering one or more site-specific nucleases to one or more target cells. In certain embodiments, the one or more site-specific nucleases are delivered to the target cells as polypeptides. In certain embodiments, the one or more site-specific nucleases comprise a nucleic acid-guided nuclease system, such as a CRISPR / Cas system. In certain embodiments, one or more polynucleotides encoding one or more components of the nuclease system are delivered to the target cells. In preferred embodiments, the nucleic acid-guided nuclease system is a type V nuclease, more preferably a type V-A nuclease, even more preferably MAD2, MAD7, ART2, ART11, ART11 * nuclease, even more preferably comprising a MAD7 nuclease.
[0120] In certain embodiments, one or more guide nucleic acids comprising a spacer sequence that is at least partially complementary to a target nucleotide sequence within a site where one or more genomic modifications are to be introduced are delivered to the target cells. In certain embodiments, one or more nucleic acid-guided nucleases are delivered to the target cells. In certain embodiments, a combination of one or more guide nucleic acids and nucleic acid-guided nucleases is delivered to the target cells, wherein the one or more nucleic acid-guided nucleases are optionally complexed with the guide nucleic acid (see the section on ribonucleoprotein (RNP) below). In certain embodiments, one or more fully formed nucleic acid-guided nuclease complexes, such as RNPs, are delivered. In some cases, any one of the embodiments described in the section on compositions and / or kits for manipulating cells to comprise a CAR can be delivered to the target cells.
[0121] II. Engineered Non-Natural Dual-Guide CRISPR-Cas System The CRISPR-Cas system generally includes a Cas protein and one or more guide nucleic acids (gNAs). The Cas protein can be directed to a specific position in a double-stranded DNA target by recognizing a protospacer adjacent motif (PAM) in the non-target strand of the DNA, and one or more guide nucleic acids can be directed to a specific position by hybridizing to a target nucleotide sequence, also referred to herein as a target sequence, in the target strand of the target polynucleotide. Typically, both PAM recognition and target nucleotide sequence hybridization are required for stable binding of the CRISPR-Cas complex to the DNA target and, if the Cas protein has an effector function (e.g., nuclease activity), activation of the effector function. As a result, when generating a CRISPR-Cas system, the guide nucleic acid can be designed to include a nucleotide sequence called a spacer sequence that is at least partially complementary to and can hybridize with the target nucleotide sequence, where the target nucleotide sequence is positioned adjacent to the PAM in an orientation operable with the Cas protein. It has been observed that not all CRISPR-Cas systems designed according to these criteria are equally effective. The larger polynucleotide in which the target nucleotide sequence is located can be referred to as the target polynucleotide; for example, a chromosome or other genomic DNA, or a portion thereof, or any other suitable polynucleotide in which the target nucleotide sequence is located. The target polynucleotide in double-stranded DNA includes two strands. The strand of the DNA duplex to which the spacer sequence is complementary is referred to herein as the "target strand," while the strand that shares sequence identity with the spacer sequence is referred to herein as the "non-target strand."
[0122] Two different classes of CRISPR-Cas systems have been identified. Class 1 CRISPR-Cas systems use multi-protein effector complexes, while Class 2 CRISPR-Cas systems use single-protein effectors (see Makarova et al. (2017) CELL, 168:328). Among the types of Class 2 CRISPR-Cas systems, type II and type V systems typically target DNA, while type VI systems typically target RNA (id.). The native type II effector complex contains Cas9, CRISPR RNA (crRNA), and trans-activating CRISPR RNA (tracrRNA), although crRNA and tracrRNA can be fused as single-guide RNA in engineered systems for simplicity (see Wang et al. (2016) ANNU. REV. BIOCHEM., 85:227). Some native type V systems, such as type V-A, type V-C, and type V-D systems, do not require tracrRNA and use single crRNA as a guide for target DNA cleavage (see Zetsche et al. (2015) CELL, 163:759; Makarova et al. (2017) CELL, 168:328).
[0123] Natural type II CRISPR-Cas systems (e.g., CRISPR-Cas9 systems) generally include two guide nucleic acids called crRNA and tracrRNA, which form a complex by nucleotide hybridization. Single guide nucleic acids capable of activating type II Cas nucleases have been developed, for example, by ligating crRNA and tracrRNA (see, e.g., U.S. Patent Nos. 10,266,850 and 8,906,616). Natural type II Cas proteins contain an RuvC-like nuclease domain and an HNH endonuclease domain and recognize a 3' G-rich PAM located immediately downstream of the target nucleotide sequence, which is an orientation determined using the non-target strand (i.e., the strand that does not hybridize to the spacer sequence) as a cofactor. The CRISPR-Cas system cleaves double-stranded DNA to generate blunt ends. The cleavage site is generally 3 to 4 nucleotides upstream from the PAM in the non-target strand.
[0124] Natural type V-A, V-C, and V-D CRISPR-Cas systems lack tracrRNA and rely on a single crRNA to direct the CRISPR-Cas complex to a target polynucleotide. Dual guide nucleic acids capable of activating a type V-A, V-C, or V-D Cas nuclease have been developed, for example, by splitting a single crRNA into a targeter nucleic acid and a modulator nucleic acid (see, e.g., International (PCT) Application Publication No. WO 2021 / 067788). Natural type V-A Cas proteins contain an RuvC-like nuclease domain but lack an HNH endonuclease domain and recognize a 5′ T-rich PAM located immediately upstream of the target nucleotide sequence (i.e., the strand that does not hybridize to the spacer sequence), which is the orientation determined using the non-target strand. These CRISPR-Cas systems cleave double-stranded DNA to generate staggered double-strand breaks rather than blunt ends. The cleavage site is distant from the PAM site (e.g., at least 10, 11, 12, 13, 14, or 15 nucleotides downstream from the PAM on the non-target strand and / or at least 15, 16, 17, 18, or 19 nucleotides upstream from the sequence complementary to the PAM on the target strand).
[0125] Elements in an exemplary single-guide CRISPR Cas system, e.g., a type V-A CRISPR-Cas system, are shown in FIG. 1A. A single gNA, when it exists in the form of RNA, may also be referred to as a "crRNA" or a "single gRNA". It includes, from 5' to 3', an optional 5' sequence, e.g., a tail, a modulator stem sequence, a loop, a targeter stem sequence complementary to the modulator stem sequence, and a spacer sequence that is at least partially complementary to and hybridizable with a target sequence in the target strand of the target polynucleotide. When a 5' tail is present, the sequence including the 5' tail and the modulator stem sequence is also referred to herein as a "modulator sequence". A fragment of the single-guide nucleic acid from any 5' tail, also referred to herein as a "scaffold sequence", to the targeter stem sequence binds to the Cas protein. Further, the PAM in the non-target strand of the target DNA binds to the Cas protein.
[0126] Elements in an exemplary dual-guide type CRISPR Cas system, such as a dual-guide type V-A CRISPR-Cas system, are shown in FIG. 1B. A first guide nucleic acid, which may be referred to herein as a "modulator nucleic acid", includes, from 5' to 3', an optional 5' tail and a modulator stem sequence. When a 5' tail is present, the sequence including the 5' tail and the modulator stem sequence may also be referred to herein as a "modulator sequence". A second guide nucleic acid, which may be referred to herein as a "targeter nucleic acid", includes, from 5' to 3', a targeter stem sequence complementary to the modulator stem sequence and a spacer sequence at least partially complementary to and hybridizable with a target sequence in a target strand of a target polynucleotide. The duplex between the modulator stem sequence and the targeter stem sequence, and any 5' tail, constitute a structure that binds to the Cas protein. Further, a PAM in the non-target strand of the target DNA binds to the Cas protein. In a dual gNA, such as a dual gRNA, the targeter nucleic acid and the modulator nucleic acid are not in the same nucleic acid, i.e., they are not linked end-to-end via a conventional polynucleotide bond, while they may be conjugated to each other covalently via one or more chemical modifications introduced into these nucleic acids, thereby enhancing the stability of the double-stranded complex and / or improving other characteristics of the system.
[0127] As used herein, the terms "targeter stem sequence" and "modulator stem sequence" can refer to a pair of nucleotide sequences in one or more guide nucleic acids that hybridize to each other. When the targeter stem sequence and the modulator stem sequence are included in a single guide nucleic acid, the targeter stem sequence is proximal to the spacer sequence designed to hybridize to the target nucleotide sequence, and the modulator stem sequence is proximal to the targeter stem sequence. When the targeter stem sequence and the modulator stem sequence are in separate nucleic acids, the targeter stem sequence is in the same nucleic acid as the spacer sequence designed to hybridize to the target nucleotide sequence. In a CRISPR-Cas system that naturally contains separate crRNAs and tracrRNAs (e.g., type II systems), the duplex formed between the targeter stem sequence and the modulator stem sequence corresponds to the duplex formed between the crRNA and the tracrRNA. In a CRISPR-Cas system that naturally contains a single crRNA but no tracrRNA (e.g., type V-A systems), the duplex formed between the targeter stem sequence and the modulator stem sequence corresponds to the stem portion of the stem-loop structure in the scaffold sequence of the crRNA. It is understood that 100% complementarity is not required between the targeter stem sequence and the modulator stem sequence. However, in type V-A CRISPR-Cas systems, the targeter stem sequence is typically 100% complementary to the modulator stem sequence.
[0128] A.Cas protein A guide nucleic acid, as a single single-guide nucleic acid (where the targeter and modulator nucleic acids are part of a single polynucleotide) or as a dual gNA comprising a separate targeter nucleic acid used in combination with a homologous modulator nucleic acid, is capable of binding to a CRISPR-associated (Cas) protein, such as a Cas nuclease. In certain embodiments, a guide nucleic acid, as a single single-guide nucleic acid (where the targeter and modulator nucleic acids are part of a single polynucleotide) or as a dual gNA comprising a separate targeter nucleic acid used in combination with a homologous modulator nucleic acid, is capable of activating a Cas nuclease. A gNA capable of activating a particular Cas nuclease is said to be "compatible" with the Cas nuclease; a Cas nuclease capable of being activated by a particular gNA is said to be "compatible" with the gNA.
[0129] The terms "CRISPR-associated protein", "Cas protein", and "Cas", used synonymously herein, can refer to a native Cas protein or an engineered Cas protein. Non-limiting examples of Cas protein engineering include, but are not limited to, mutating and modifying the Cas protein to modify its activity, modify its PAM specificity, expand the range of recognized PAMs, and / or reduce the ability to modify one or more off-target loci compared to the corresponding unmodified Cas. In certain embodiments, the modified activity of the engineered Cas includes a modified ability to bind to a native gNA, such as a gRNA, or an engineered gNA, such as a gRNA (e.g., specificity or reaction rate), a modified ability to bind to a target nucleotide sequence (e.g., specificity or reaction rate), a modified processivity of nucleic acid scanning, and / or a modified effector (e.g., nuclease) activity. A Cas protein having nuclease activity can be referred to as a "CRISPR-associated nuclease" or "Cas nuclease", or simply "nuclease", as used synonymously herein.
[0130] In certain embodiments, the Cas protein is a type V-A, V-C, or V-D Cas protein. In certain embodiments, the Cas protein is a type V-A Cas protein. In other embodiments, the Cas protein is a type II Cas protein, such as, for example, Cas9 protein.
[0131] In certain embodiments, the V-A type Cas nuclease comprises Cpf1. The Cpf1 protein is known in the art and is described, for example, in U.S. Patent Nos. 9,790,490 and 10,113,179. Cpf1 orthologs are found in various bacterial and archaeal genomes. For example, in certain embodiments, the Cpf1 protein is Francisella novicida U112 (Fn), Acidaminococcus sp. BV3L6 (As), Lachnospiraceae bacterium ND2006 (Lb), Lachnospiraceae bacterium MA2020 (Lb2), Candidatus Methanoplasma termitum (CMt), Moraxella bovoculi 237 (Mb), Porphyromonas crevioricanis (Pc), Prevotella disiens (Pd), Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC2017 1, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacterium GW2011_GWC2_44_17, Smithella sp.)It is derived from the genes of SCADC, Eubacterium eligens, Leptospira inadai, Porphyromonas macacae, Prevotella bryantii, Proteocatella sphenisci, Anaerovibrio sp. RM50, Moraxella caprae, Lachnospiraceae bacterium COE1, or Eubacterium coprostanoli.
[0132] In certain embodiments, the type V-A Cas nuclease comprises AsCpf1 or a variant thereof. In certain embodiments, the type V-A Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 3 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the type V-A Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 3 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0133] In certain embodiments, the V-A type Cas nuclease comprises LbCpf1 or a variant thereof. In certain embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 4 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the V-A type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 4 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0134] In certain embodiments, the V-A type Cas nuclease comprises FnCpf1 or a variant thereof. In certain embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 5 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the V-A type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 5 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0135] In certain embodiments, the type V-A Cas nuclease comprises Prevotella bryantii Cpf1 (PbCpf1) or a variant thereof. In certain embodiments, the type V-A Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 6 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the type V-A Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 6 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0136] In certain embodiments, the type V-A Cas nuclease comprises Proteocatella sphenisci Cpf1 (PsCpf1) or a variant thereof. In certain embodiments, the type V-A Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 7 of the pamphlet of International (PCT) Application Publication No. WO 2021158918. In certain embodiments, the type V-A Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 7 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0137] In certain embodiments, the V-A type Cas nuclease comprises Anaerovibrio sp. RM50 Cpf1 (As2Cpf1) or a variant thereof. In certain embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 8 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the V-A type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 8 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0138] In certain embodiments, the V-A type Cas nuclease comprises Moraxella caprae Cpf1 (McCpf1) or a variant thereof. In certain embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 9 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the V-A type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 9 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0139] In certain embodiments, the type V-A Cas nuclease comprises Lachnospiraceae bacterium COE1 Cpf1 (Lb3Cpf1) or a variant thereof. In certain embodiments, the type V-A Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 10 of the International (PCT) Application Publication No. WO 2021 / 158918 pamphlet. In certain embodiments, the type V-A Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 10 of the International (PCT) Application Publication No. WO 2021 / 158918 pamphlet.
[0140] In certain embodiments, the type V-A Cas nuclease comprises Eubacterium coprostanoli gene Cpf1 (EcCpf1) or a variant thereof. In certain embodiments, the type V-A Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 11 of the International (PCT) Application Publication No. WO 2021 / 158918 pamphlet. In certain embodiments, the type V-A Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 11 of the International (PCT) Application Publication No. WO 2021 / 158918 pamphlet.
[0141] In certain embodiments, the type V-A Cas nuclease is not Cpf1. In certain embodiments, the type V-A Cas nuclease is not AsCpf1.
[0142] In certain embodiments, the V-A type Cas nuclease comprises MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, or MAD20, or variants thereof. MAD1-MAD20 are known in the art and are described in U.S. Patent No. 9,982,279.
[0143] In certain embodiments, the V-A type Cas nuclease comprises MAD7 or a variant thereof. In certain embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 37. In certain embodiments, the V-A type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 37.
[0144] MAD7 (SEQ ID NO: 37)
Chemical formula
[0145] In certain embodiments, the V-A type Cas nuclease comprises MAD2 or a variant thereof. In certain embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 38. In certain embodiments, the V-A type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 38.
[0146] MAD2 (SEQ ID NO: 38)
Chemical formula
[0147] In certain embodiments, the V-A type Cas nuclease comprises Csm1. The Csm1 protein is known in the art and is described in U.S. Patent No. 9,896,696. Csm1 orthologs are found in various bacterial and archaeal genomes. For example, in certain embodiments, the Csm1 protein is derived from Smithella sp. SCADC (Sm), Sulfuricurvum sp. (Ss), or Microgenomates (Roizmanbacteria) bacterium (Mb).
[0148] In certain embodiments, the V-A type Cas nuclease comprises SmCsm1 or a variant thereof. In certain embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 12 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the V-A type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 12 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0149] In certain embodiments, the V-A type Cas nuclease comprises SsCsm1 or a variant thereof. In certain embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 13 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the V-A type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 13 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0150] In certain embodiments, the V-A type Cas nuclease comprises MbCsm1 or a variant thereof. In certain embodiments, the V-A type Cas protein comprises an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in SEQ ID NO: 14 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918. In certain embodiments, the V-A type Cas protein comprises the amino acid sequence set forth in SEQ ID NO: 14 of the pamphlet of International (PCT) Application Publication No. WO 2021 / 158918.
[0151] In certain embodiments, the V-A type Cas nuclease comprises an ART nuclease or a variant thereof. Generally, such nuclease sequences have <60% AA sequence similarity to Cas12a, <60% AA sequence similarity to a positive control nuclease, and >80% query coverage. In certain embodiments, the V-A type nuclease is ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11, ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART28, ART30, ART31, ART32, ART33, ART34, ART35, or ART11 *(That is, ART11_L679F, that is, it includes ART11, where leucine (L) at amino acid position 679 is replaced by phenylalanine (F) nuclease. In certain embodiments, the V-A type Cas protein has at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the amino acid sequence specified for each individual ART nuclease as shown in Table 3. In certain embodiments, provided is a nucleic acid-derived nuclease polypeptide having at least 85% identity to the amino acid sequence represented by SEQ ID NOs: 1-36 or a nucleic acid encoding a nucleic acid-derived nuclease polypeptide comprising at least 85% identity to the polynucleotide represented by SEQ ID NOs: 1-36. In certain embodiments, provided is a nucleic acid-derived nuclease comprising a polypeptide having at least 90% identity to the amino acid sequence represented by SEQ ID NOs: 1-36, wherein the polypeptide does not contain the peptide motif of YLFQIYNKDF (SEQ ID NO: 39). In certain embodiments, provided is a nucleic acid-derived nuclease comprising a nucleic acid encoding a polypeptide having at least 90% identity to the nucleic acid represented by SEQ ID NOs: 808-845, wherein the encoded polypeptide does not contain the peptide motif of YLFQIYNKDF (SEQ ID NO: 39). In certain embodiments, provided is a nucleic acid-derived nuclease wherein the polypeptide comprises at least 90% identity to the amino acid sequence represented by SEQ ID NOs: 1-9. In certain embodiments, provided is a nucleic acid-derived nuclease comprising a polypeptide having at least 90% identity to the amino acid sequence represented by SEQ ID NO: 2, 11, or 36.)
[0152] [Table 67]
[0153] [Table 68]
[0154]
Table 69
[0155]
Table 70
[0156]
Table 71
[0157]
Table 72
[0158]
Table 73
[0159]
Table 74
[0160]
Table 75
[0161]
Table 76
[0162]
Table 77
[0163]
Table 78
[0164]
Table 79
[0165]
Table 80
[0166]
Table 81
[0167]
Table 82
[0168]
Table 83
[0169]
Table 84
[0170]
Table 85
[0171]
Table 86
[0172]
Table 87
[0173]
Table 88
[0174]
Table 89
[0175]
Table 90
[0176]
Table 91
[0177]
Table 92
[0178]
Table 93
[0179]
Table 94
[0180]
Table 95
[0181]
Table 96
[0182] In certain embodiments, the Cas nuclease is ABW1 (SEQ ID NO: 3), ABW2 (SEQ ID NO: 16), ABW3 (SEQ ID NO: 29), ABW4 (SEQ ID NO: 42), ABW5 (SEQ ID NO: 55), ABW6 (SEQ ID NO: 68), ABW7 (SEQ ID NO: 81), ABW8 (SEQ ID NO: 94), or ABW9 (SEQ ID NO: 107) (all SEQ ID NOs of ABW1-9 and variants thereof from International (PCT) Application Publication No. WO 2021 / 108324 Pamphlet), or variants thereof, for example, any one of variants 1-10 of ABW1 (SEQ ID NOs: 4-13 respectively), any one of variants 1-10 of ABW2 (SEQ ID NOs: 17-26 respectively), any one of variants 1-10 of ABW3 (SEQ ID NOs: 30-39 respectively), any one of variants 1-10 of ABW4 (SEQ ID NOs: 43-52 respectively), any one of variants 1-10 of ABW5 (SEQ ID NOs: 56-65 respectively), any one of variants 1-10 of ABW6 (SEQ ID NOs: 69-78 respectively), any one of variants 1-10 of ABW7 (SEQ ID NOs: 82-91 respectively), any one of variants 1-10 of ABW8 (SEQ ID NOs: 95-104 respectively), any one of variants 1-10 of ABW9 (SEQ ID NOs: 108-117 respectively). ABW1-ABW9, and variants thereof, are known in the art and are described in International (PCT) Application Publication No. WO 2021 / 108324 Pamphlet.
[0183] Additional V-A type Cas nucleases and their corresponding native CRISPR-Cas systems can be identified by computational and experimental methods known in the art, such as those described in U.S. Patent No. 9,790,490 and Shmakov et al. (2015) Mol. Cell, 60:385. Exemplary computational methods include analysis of putative Cas proteins by homology modeling, structural BLAST, PSI-BLAST, or HHPred, and analysis of putative CRISPR loci by identification of CRISPR arrays. Exemplary experimental methods include in vitro cleavage assays and intracellular nuclease assays (e.g., Surveyor assay) as described in Zetsche et al. (2015) Cell, 163:759.
[0184] In certain embodiments, the Cas protein is a Cas nuclease that directs cleavage of one or both strands at a target locus, e.g., the target strand (i.e., the strand having a target nucleotide sequence that is at least partially complementary to and hybridizable with a single guide nucleic acid or dual guide nucleic acid) and / or the non-target strand. In certain embodiments, the Cas nuclease directs cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more nucleotides from the first or last nucleotide of the target nucleotide sequence or its complementary sequence. In certain embodiments, the cleavages are staggered, i.e., generate sticky ends. In certain embodiments, the cleavages generate staggered cleavages with 5' overhangs. In certain embodiments, the cleavages generate staggered cleavages with 5' overhangs of 1 to 5 nucleotides, e.g., 4 or 5 nucleotides. In certain embodiments, the cleavage site is distal from the PAM, e.g., the cleavage occurs after the 18th nucleotide on the non-target strand and after the 23rd nucleotide on the target strand.
[0185] In certain embodiments, the compositions provided herein include a compatibility guide nucleic acid (gNA), e.g., a Cas nuclease that can be activated by a gRNA. In certain embodiments, the compositions provided herein further include a Cas protein associated with a compatibility guide nucleic acid (gNA), e.g., a Cas nuclease that can be activated by a gRNA. For example, in certain embodiments, the Cas protein includes an amino acid sequence that is at least 80% (e.g., at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) identical to the Cas nuclease amino acid sequence. In certain embodiments, the Cas protein includes a nuclease-inactive mutant of the Cas nuclease. In certain embodiments, the Cas protein further includes an effector domain.
[0186] In certain embodiments, the Cas protein lacks substantially all DNA cleavage activity. Such Cas proteins can be generated, for example, by introducing one or more mutations into an active Cas nuclease (e.g., a native Cas nuclease). A mutated Cas protein is considered to lack substantially all DNA cleavage activity if the protein's DNA cleavage activity is about 25%, 10%, 5%, 1%, 0.1%, 0.01% or less, or less than that, of the DNA cleavage activity of the corresponding non-mutated form, e.g., nil or very little compared to the non-mutated form. Thus, the Cas protein may contain one or more mutations (e.g., mutations in the RuvC domain of a V-A type Cas protein) and can be used as a general DNA binding protein with or without fusion to an effector domain. Exemplary mutations include D908A, E993A, and D1263A with respect to the amino acid positions in AsCpf1; D832A, E925A, and D1180A with respect to the amino acid positions in LbCpf1; and D917A, E1006A, and D1255A with respect to the amino acid position numbering of FnCpf1. Additional mutations can be designed and generated according to the crystal structures described in Yamano et al. (2016) CELL, 165:949.
[0187] It is understood that the Cas protein may lose the ability to cleave only the target strand or only the non-target strand of double-stranded DNA, rather than losing nuclease activity to cleave all DNA, thereby functioning as a nickase (see Gao et al. (2016) CELL RES., 26:901). Thus, in certain embodiments, the Cas nuclease is a Cas nickase. In certain embodiments, the Cas nuclease has activity to cleave the non-target strand but substantially lacks activity to cleave the target strand, e.g., due to a mutation in the Nuc domain. In certain embodiments, the Cas nuclease has cleavage activity to cleave the target strand but substantially lacks activity to cleave the non-target strand.
[0188] In certain embodiments, the Cas nuclease has the activity of cleaving double-stranded DNA, resulting in double-strand breaks.
[0189] Cas proteins that substantially lack all DNA cleavage activity or have the ability to cleave only one strand can also be identified from natural systems. For example, certain natural CRISPR-Cas systems may retain the ability to bind to target nucleotide sequences but lose overall or partial DNA cleavage activity in eukaryotic (e.g., mammalian or human) cells. Such type V-A proteins are disclosed, for example, in Kim et al. (2017) ACS SYNTH. BIOL. 6(7):1273-82 and Zhang et al. (2017) CELL DISCOV. 3:17018.
[0190] The activity of a Cas protein (e.g., a Cas nuclease) can be modified, for example, by generating a engineered Cas protein. In certain embodiments, the modified activity of the engineered Cas protein includes increased targeting efficiency and / or decreased off-target binding. Without wishing to be bound by theory, it is hypothesized that off-target binding can be recognized by the Cas protein, for example, due to the presence of one or more mismatches between the spacer sequence and the target nucleotide sequence, which can affect the stability and / or conformation of the CRISPR-Cas complex. In certain embodiments, the modified activity includes modified binding, e.g., increased binding to a target locus (e.g., the target strand or non-target strand) and / or decreased binding to an off-target locus. In certain embodiments, the modified activity includes a modified charge in a region of the protein associated with a single guide nucleic acid or dual guide nucleic acid. In certain embodiments, the modified activity of the engineered Cas protein includes a modified charge in a region of the protein associated with the target strand and / or non-target strand. In certain embodiments, the modified activity of the engineered Cas protein includes a modified charge in a region of the protein associated with an off-target locus. The modified charge can include a decreased positive charge, a decreased negative charge, an increased positive charge, or an increased negative charge. For example, a decreased negative charge and an increased positive charge can generally enhance binding to nucleic acids, while a decreased positive charge and an increased negative charge can weaken binding to nucleic acids. In certain embodiments, the modified activity includes an increased or decreased steric hindrance between the protein and the single guide nucleic acid or dual guide nucleic acid. In certain embodiments, the modified activity includes an increased or decreased steric hindrance between the protein and the target strand and / or non-target strand. In certain embodiments, the modified activity includes an increased or decreased steric hindrance between the protein and an off-target locus. In certain embodiments, the modification or mutation includes one or more substitutions of Lys, His, Arg, Glu, Asp, Ser, Gly, and / or Thr.In certain embodiments, the modification or mutation comprises one or more substitutions with Gly, Ala, Ile, Glu, and / or Asp. In certain embodiments, the modification or mutation comprises one or more amino acid substitutions in the groove between the WED and RuvC domains of the Cas protein (e.g., a type V-A Cas protein).
[0191] In certain embodiments, the modified activity of the engineered Cas protein comprises increased nuclease activity that cleaves the target locus. In certain embodiments, the modified activity of the engineered Cas protein comprises decreased nuclease activity that cleaves off-target loci. In certain embodiments, the modified activity of the engineered Cas protein comprises a modified helicase reaction rate. In certain embodiments, the engineered Cas protein comprises a modification that modifies the formation of the CRISPR complex.
[0192] In certain embodiments, the protospacer adjacent motif (PAM) or PAM-like motif directs the binding of the Cas protein complex to the target locus. Many Cas proteins have PAM specificity. The exact sequence and length requirements of the PAM vary depending on the Cas protein used. The PAM sequence is typically 2-5 base pairs in length and is adjacent to the target nucleotide sequence (although it is located on a strand different from the target nucleotide sequence in the target DNA). The PAM sequence can be identified by any suitable method, such as testing for cleavage, targeting, or modifying oligonucleotides having the target nucleotide sequence and different PAM sequences.
[0193] Exemplary PAM sequences are shown in Tables 2 and 3. In certain embodiments, the Cas protein comprises MAD7 and the PAM is TTTN, where N is A, C, G, or T. In certain embodiments, the Cas protein comprises MAD7 and the PAM is CTTN, where N is A, C, G, or T. In certain embodiments, the Cas protein comprises AsCpf1 and the PAM is TTTN, where N is A, C, G, or T. In certain embodiments, the Cas protein comprises FnCpf1 and the PAM is 5’TTN, where N is A, C, G, or T. PAM sequences for certain other type V-A Cas proteins are disclosed in Zetsche et al. (2015) CELL, 163:759 and U.S. Patent No. 9,982,279. Furthermore, engineering of the PAM-interacting (PI) domain of the Cas protein can enable programming of PAM specificity, improve target site recognition fidelity, and / or increase the versatility of engineered non-native systems. Exemplary methods for modifying the PAM specificity of Cpf1 are described in Gao et al. (2017) NAT. BIOTECHNOL., 35:789.
[0194] In certain embodiments, engineered Cas proteins include modifications that modify Cas protein specificity in conjunction with modifications to the targeting scope. Cas mutants can be designed to have increased target specificity and corresponding modifications in PAM recognition, for example, by selecting mutations that modify PAM specificity (e.g., in the PI domain) and combining those mutations with groove mutations that increase (or decrease as needed) specificity for the on-target locus compared to off-target loci. The Cas modifications described herein can be used to suppress a decrease in specificity resulting from modification of PAM recognition, enhance an increase in specificity resulting from modification of PAM recognition, suppress an increase in specificity resulting from modification of PAM recognition, or enhance a decrease in specificity resulting from modification of PAM recognition.
[0195] In certain embodiments, the engineered Cas protein comprises one or more nuclear localization signal (NLS) motifs. In certain embodiments, the engineered Cas protein comprises at least two (e.g., at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs. Non-limiting examples of NLS motifs include: the NLS of SV40 large T antigen having the amino acid sequence of PKKKRKV (SEQ ID NO: 40); the NLS from the nucleoplasm, e.g., the bipartite nucleoplasmic NLS, having the amino acid sequence of KRPAATKKAGQAKKKK (SEQ ID NO: 41); the C-myc NLS having the amino acid sequence of PAAKRVKLD (SEQ ID NO: 42) or RQRRNELKRSP (SEQ ID NO: 43); the hRNPA1 M9 NLS having the amino acid sequence of NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 44); the importin-α IBB domain NLS having the amino acid sequence of RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 45); the myogenic T protein NLS having the amino acid sequence of VSRKRPRP (SEQ ID NO: 46) or PPKKARED (SEQ ID NO: 47); the human p53 NLS having the amino acid sequence of PQPKKKPL (SEQ ID NO: 48); the mouse C-abl IV NLS having the amino acid sequence of SALIKKKKKMAP (SEQ ID NO: 49); the influenza virus NS1 NLS having the amino acid sequence of DRLRR (SEQ ID NO: 50) or PKQKKRK (SEQ ID NO: 51); the hepatitis delta antigen NLS having the amino acid sequence of RKLKKKIKKL (SEQ ID NO: 52); the mouse Mx1 protein NLS having the amino acid sequence of REKKKFLKRR (SEQ ID NO: 53); the human poly(ADP-ribose) polymerase NLS having the amino acid sequence of KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 54); the human glucocorticoid receptor NLS having the amino acid sequence of RKCLQAGMNLEARKTKK (SEQ ID NO: 55); and synthetic NLS motifs such as PAAKKKKLD (SEQ ID NO: 56).
[0196] Generally, one or more NLS motifs are of sufficient strength to promote the accumulation of detectable amounts of Cas protein within the nucleus of eukaryotic cells. The strength of the nuclear localization activity is obtained by the NLS motif in the Cas protein, the number of specific NLS motifs used, the position of the NLS motif, or a combination of these and / or other factors. In certain embodiments, the engineered Cas protein comprises at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs at or near the N-terminus (e.g., within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the N-terminus). In certain embodiments, the engineered Cas protein comprises at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs at or near the C-terminus (e.g., within about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more amino acids along the polypeptide chain from the C-terminus). In certain embodiments, the engineered Cas protein comprises at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs at or near the C-terminus and at least one (e.g., at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten) NLS motifs at or near the N-terminus. In certain embodiments, the engineered Cas protein comprises one, two, or three NLS motifs at or near the C-terminus. In certain embodiments, the engineered Cas protein comprises one NLS motif at or near the N-terminus and one, two, or three NLS motifs at or near the C-terminus.In certain embodiments, the engineered Cas protein comprises a nuclear localization signal (NLS) at or near its C-terminus.
[0197] Detection of accumulation in the nucleus can be performed by any suitable technique. For example, a detectable marker can be fused to the nucleic acid-targeting protein such that the intracellular location can be visualized. The cell nucleus can also be isolated from the cell and then its contents analyzed by any suitable process for detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assay. Accumulation in the nucleus can also be determined indirectly by assays that detect the effect of transport of the Cas protein complex into the nucleus as compared to a control that is not exposed to the Cas protein or is exposed to a Cas protein lacking one or more of the NLS motifs (e.g., an assay for DNA cleavage or mutation at the target locus, or an assay for altered gene expression activity).
[0198] The Cas protein can include chimeric Cas proteins, e.g., Cas proteins having enhanced function by virtue of being chimeric. A chimeric Cas protein can be a new Cas protein containing fragments from two or more native Cas proteins or variants thereof. For example, fragments of multiple type V-A Cas homologs (e.g., orthologs) can be fused to form a chimeric Cas protein. In certain embodiments, the chimeric Cas protein comprises fragments of Cpf1 orthologs from multiple species and / or strains.
[0199] In certain embodiments, the Cas protein comprises one or more effector domains. The one or more effector domains can be located at or near the N-terminus of the Cas protein and / or at or near the C-terminus of the Cas protein. In certain embodiments, the effector domain included in the Cas protein is a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., KRAB domain or SID domain), an exogenous nuclease domain (e.g., FokI), a deaminase domain (e.g., cytidine deaminase or adenine deaminase), or a reverse transcriptase domain (e.g., high-fidelity reverse transcriptase domain). Other activities of the effector domain include, but are not limited to, methylase activity, demethylase activity, transcriptional release factor activity, translation initiation activity, translational activation activity, translational repression activity, histone modification (e.g., acetylation or demethylation) activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity.
[0200] In certain embodiments, the Cas protein comprises one or more protein domains that enhance homologous recombination repair (HDR) and / or inhibit non-homologous end joining (NHEJ). Exemplary protein domains having such functions are described in Jayavaradhan et al. (2019) Nat. Commun. 10(1):2866 and Janssen et al. (2019) Mol. Ther. Nucleic Acids 16:141-54. In certain embodiments, the Cas protein comprises a dominant-negative form of p53-binding protein 1 (53BP1), e.g., a fragment of 53BP1 comprising the minimal focus-forming region (e.g., amino acids 1231-1644 of human 53BP1). In certain embodiments, the Cas protein comprises a motif targeted by APC-Cdh1, e.g., amino acids 1-110 of human geminin, thereby resulting in degradation of the fusion protein during the HDR-incompetent G1 phase of the cell cycle.
[0201] In certain embodiments, the Cas protein includes an inducible or regulatable domain. Non-limiting examples of inducers or regulators include light, hormones, and small molecule drugs. In certain embodiments, the Cas protein includes a light-inducible or regulatable domain. In certain embodiments, the Cas protein includes a chemically inducible or regulatable domain.
[0202] In certain embodiments, the Cas protein includes a tag protein or peptide to facilitate tracking and / or purification. Non-limiting examples of tag proteins and peptides include fluorescent proteins (e.g., green fluorescent protein (GFP), YFP, RFP, CFP, mCherry, tdTomato), HIS tags (e.g., 6×His tag, or gly-6xHis; 8xHis, or gly-8xHis), hemagglutinin (HA) tags, FLAG tags, 3xFLAG tags, and Myc tags.
[0203] In certain embodiments, the Cas protein is conjugated to a non-protein moiety, such as a fluorophore useful for genomic imaging. In certain embodiments, the Cas protein is covalently conjugated to a non-protein moiety. The terms "CRISPR-associated protein", "Cas protein", "Cas", "CRISPR-associated nuclease", and "Cas nuclease" are used herein to include such conjugates regardless of the presence of one or more non-protein moieties.
[0204] B. Guide Nucleic Acid The guide nucleic acid can be a single guide nucleic acid (sgNA, e.g., sgRNA) where the gNA is a single polynucleotide, or a dual guide nucleic acid (e.g., dual gRNA) where the gNA includes two separate polynucleotides (which in some cases can be covalently linked but not via conventional inter-nucleotide linkages). In certain embodiments, the single guide nucleic acid is capable of activating the Cas nuclease alone (e.g., in the absence of tracrRNA).
[0205] Generally, a gNA includes a modulator nucleic acid and a targeter nucleic acid. In an sgNA, the modulator and targeter nucleic acids are part of a single polynucleotide. In a dual gNA, the modulator and targeter nucleic acids are separate and, for example, are not linked by conventional nucleotide binding, for example, are not linked at all. The targeter nucleic acid includes a spacer sequence and a targeter stem sequence. The modulator nucleic acid includes a modulator stem sequence and generally further includes nucleotides, for example, nucleotides including a 5’ tail. The modulator stem sequence and the targeter stem sequence can each include any suitable number of nucleotides and have sufficient complementarity to hybridize. In a single gNA, additional NTs may be present between the targeter stem sequence and the modulator stem sequence; in some cases, these can form secondary structures such as loops.
[0206] In certain embodiments, the guide nucleic acid includes a targeter nucleic acid that can bind to a Cas protein in combination with the modulator nucleic acid. In certain embodiments, the guide nucleic acid includes a targeter nucleic acid that can activate a Cas nuclease in combination with the modulator nucleic acid. In certain embodiments, the system further includes a Cas protein to which the targeter nucleic acid and the modulator nucleic acid can bind or a Cas nuclease that can be activated by the targeter nucleic acid and the modulator nucleic acid.
[0207] A single or dual guide nucleic acid is thought to need to be compatible with a Cas protein (e.g., a Cas nuclease) to provide an operative CRISPR system. For example, the targeter stem sequence and the modulator stem sequence can be derived from a native crRNA capable of activating a Cas nuclease in the absence of tracrRNA. Alternatively, the targeter stem sequence and the modulator stem sequence can be derived from the respective native pair of a crRNA and a tracrRNA capable of activating a Cas nuclease. In certain embodiments, the nucleotide sequences of the targeter stem sequence and the modulator stem sequence are identical to the corresponding stem sequences of the stem-loop structure in such native crRNAs.
[0208] Guide nucleic acid sequences that act with type II or type V Cas proteins are known in the art and are disclosed, for example, in U.S. Patent Nos. 9,790,490, 9,896,696, 10,113,179, and 10,266,850, and U.S. Patent Application Publication No. 2014 / 0242664. It is understood that these sequences are merely exemplary and that other guide nucleic acid sequences can also be used with these Cas proteins.
[0209] [Table 97]
[0210] [Table 98]
[0211] [Table 99]
[0212] [Table 100]
[0213]
Table 101
[0214]
Table 102
[0215] In certain embodiments, the guide nucleic acid comprises a targeter stem sequence listed in Table 5 with respect to a type V-A CRISPR-Cas system. The targeter stem sequences that are the same as a part of the scaffold sequence are underlined in bold in Table 4.
[0216] In certain embodiments, the guide nucleic acid is a single guide nucleic acid that includes, from 5' to 3', a modulator stem sequence, a loop sequence, a targeter stem sequence, and a spacer sequence. In certain embodiments, the targeter stem sequence in the single guide nucleic acid is listed in Table 4 as the underlined and bolded portion of the scaffold sequence, and the modulator stem sequence is complementary to the targeter stem sequence (e.g., 100% complementary). In certain embodiments, the single guide nucleic acid includes, from 5' to 3', a modulator sequence listed in Table 4 as the underlined portion of the scaffold sequence, a loop sequence, a targeter stem sequence that is the underlined and bolded portion of the same scaffold sequence, and a spacer sequence. In certain embodiments, the engineered non-natural system includes a single guide nucleic acid that includes the scaffold sequence listed in Table 4. In certain embodiments, the system further includes a Cas protein (e.g., a Cas nuclease) that includes an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in the SEQ ID NO listed in the same row of Table 4. In certain embodiments, the system further includes a Cas protein (e.g., a Cas nuclease) that includes the amino acid sequence set forth in the SEQ ID NO listed in the same row of Table 4. In certain embodiments, the system is useful for targeting, editing, or modifying a nucleic acid that includes a target nucleotide sequence near or adjacent to (e.g., immediately downstream of) the PAM listed in the same row of Table 4 when using a non-target strand as a cofactor (i.e., a strand that does not hybridize to the spacer sequence).
[0217] In certain embodiments, a guide nucleic acid, e.g., a dual gNA, comprises a target guide nucleic acid that includes a target stem sequence and a spacer sequence, 5' to 3'. In certain embodiments, the target stem sequence in the target nucleic acid is listed in Table 5. In certain embodiments, the engineered non-natural system comprises a target nucleic acid that is complementary (e.g., 100% complementary) to the target stem sequence and a modulator stem sequence. In certain embodiments, the modulator nucleic acid comprises a modulator sequence listed in the same row of Table 5. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) that has an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence set forth in the SEQ ID NO listed in the same row of Table 5. In certain embodiments, the system further comprises a Cas protein (e.g., a Cas nuclease) that has the amino acid sequence set forth in the SEQ ID NO listed in the same row of Table 5. In certain embodiments, the system is useful for targeting, editing, or modifying a nucleic acid that comprises a target nucleotide sequence near or adjacent to (e.g., immediately downstream of) a PAM listed in the same row of Table 5 when using a non-target strand as a cofactor.
[0218] The single-guide nucleic acid, targeter nucleic acid, and / or modulator nucleic acid can be chemically synthesized or generated in a biological process (e.g., catalyzed by RNA polymerase in an in vitro reaction). Such a reaction or process can limit the length of the single-guide nucleic acid, targeter nucleic acid, and / or modulator nucleic acid. In certain embodiments, the single-guide nucleic acid is 100, 90, 80, 70, 60, 50, 40, 30, or 25 nucleotides or less in length. In certain embodiments, the single-guide nucleic acid is at least 20, 25, 30, 40, 50, 60, 70, 80, or 90 nucleotides in length. In certain embodiments, the single-guide nucleic acid is 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, 20-30, 20-25, 25-100, 25-90, 25-80, 25-70, 25-60, 25-50, 25-40, 25-30, 30-100, 30-90, 30-80, 30-70, 30-60, 30-50, 30-40, 40-100, 40-90, 40-80, 40-70, 40-60, 40-50, 50-100, 50-90, 50-80, 50-70, 50-60, 60-100, 60-90, 60-80, 60-70, 70-100, 70-90, 70-80, 80-100, 80-90, or 90-100 nucleotides in length. In certain embodiments, the targeter nucleic acid is 100, 90, 80, 70, 60, 50, 40, 30, or 25 nucleotides or less in length. In certain embodiments, the targeter nucleic acid is at least 20, 25, 30, 40, 50, 60, 70, 80, or 90 nucleotides in length.In certain embodiments, the targeter nucleic acid is 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 20 to 25, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to 60, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 90, 30 to 80, 30 to 70, 30 to 60, 30 to 50, 30 to 40, 40 to 100, 40 to 90, 40 to 80, 40 to 70, 40 to 60, 40 to 50, 50 to 100, 50 to 90, 50 to 80, 50 to 70, 50 to 60, 60 to 100, 60 to 90, 60 to 80, 60 to 70, 70 to 100, 70 to 90, 70 to 80, 80 to 100, 80 to 90, or 90 to 100 nucleotides in length. In certain embodiments, the modulator nucleic acid is 100, 90, 80, 70, 60, 50, 40, 30, or 20 nucleotides or less in length. In certain embodiments, the modulator nucleic acid is at least 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, or 90 nucleotides in length. In certain embodiments, the modulator nucleic acid is 10 to 100, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 15 to 100, 15 to 90, 15 to 80, 15 to 70, 15 to 60, 15 to 50, 15 to 40, 15 to 30, 15 to 20, 20 to 100, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 25 to 100, 25 to 90, 25 to 80, 25 to 70, 25 to 60, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 90, 30 to 80, 30 to 70, 30 to 60, 30 to 50, 30 to 40, 40 to 100, 40 to 90, 40 to 80, 40 to 70, 40 to 60, 40 to 50, 50 to 100, 50 to 90, 50 to 80, 50 to 70, 50 to 60, 60 to 100, 60 to 90, 60 to 80, 60 to 70, 70 to 100, 70 to 90, 70 to 80, 80 to 100, 80 to 90, or 90 to 100 nucleotides in length.
[0219] The length of the duplex formed within a single-guide nucleic acid or between a targeter nucleic acid and a modulator nucleic acid, e.g., in a dual gNA, can be a factor in providing an operative CRISPR system. In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4 to 10 nucleotides that base pair with each other. In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 10, 5 to 9, 5 to 8, 5 to 7, or 5 to 6 nucleotides that base pair with each other. In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of 4, 5, 6, 7, 8, 9, or 10 nucleotides. It is understood that the composition of the nucleotides in each sequence affects the stability of the duplex, and that C-G base pairs confer higher stability than A-U base pairs. In certain embodiments, 20% to 80%, 20% to 70%, 20% to 60%, 20% to 50%, 20% to 40%, 20% to 30%, 30% to 80%, 30% to 70%, 30% to 60%, 30% to 50%, 30% to 40%, 40% to 80%, 40% to 70%, 40% to 60%, 40% to 50%, 50% to 80%, 50% to 70%, 50% to 60%, 60% to 80%, 60% to 70%, or 70% to 80% of the base pairs are C-G base pairs.
[0220] In certain embodiments, the targeter stem sequence and the modulator stem sequence each consist of five nucleotides. Accordingly, the targeter stem sequence and the modulator stem sequence form a double-strand of five base pairs. In certain embodiments, 0 to 4, 0 to 3, 0 to 2, 0 to 1, 1 to 5, 1 to 4, 1 to 3, 1 to 2, 2 to 5, 2 to 4, 2 to 3, 3 to 5, 3 to 4, or 4 to 5 of the five base pairs are C-G base pairs. In certain embodiments, 0, 1, 2, 3, 4, or 5 of the five base pairs are C-G base pairs. In certain embodiments, the targeter stem sequence consists of 5'-GUAGA-3', and the modulator stem sequence consists of 5'-UCUAC-3'. In certain embodiments, the targeter stem sequence consists of 5'-GUGGG-3', and the modulator stem sequence consists of 5'-CCCAC-3'.
[0221] In certain embodiments, in a V-A type system, the 3' end of the targeter stem sequence is linked to the 5' end of the spacer sequence by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or fewer nucleotides. In certain embodiments, the targeter stem sequence and the spacer sequence are adjacent to each other and are directly linked by a polynucleotide bond. In certain embodiments, the targeter stem sequence and the spacer sequence are linked by one nucleotide, for example, uridine. In certain embodiments, the targeter stem sequence and the spacer sequence are linked by two or more nucleotides. In certain embodiments, the targeter stem sequence and the spacer sequence are linked by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides.
[0222] In certain embodiments, the targeter nucleic acid further comprises an additional nucleotide sequence 5' to the targeter stem sequence. In certain embodiments, the additional nucleotide sequence comprises at least one (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50) nucleotides. In certain embodiments, the additional nucleotide sequence consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides. In certain embodiments, the additional nucleotide sequence consists of two nucleotides. In certain embodiments, the additional nucleotide sequence is similar to a loop or a fragment thereof (e.g., 1, 2, 3, or 4 nucleotides at the 3' end of the loop) in the crRNA of the corresponding single-guide CRISPR-Cas system. It can be understood that the additional nucleotide sequence 5' to the targeter stem sequence may not be essential. Thus, in certain embodiments, the targeter nucleic acid does not comprise any additional nucleotides 5' to the targeter stem sequence.
[0223] In certain embodiments, the targeter nucleic acid or single guide nucleic acid further comprises an additional nucleotide sequence containing one or more nucleotides at the 3' end that do not hybridize to the target nucleotide sequence. The additional nucleotide sequence may protect the targeter nucleic acid from degradation by 3'-5' exonucleases. In certain embodiments, the additional nucleotide sequence is 100 nucleotides in length or less. In certain embodiments, the additional nucleotide sequence is 90, 80, 70, 60, 50, 40, 30, 20, or 10 nucleotides in length or less. In certain embodiments, the additional nucleotide sequence is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides in length. In certain embodiments, the additional nucleotide sequence is 5 to 100, 5 to 50, 5 to 40, 5 to 30, 5 to 25, 5 to 20, 5 to 15, 5 to 10, 10 to 100, 10 to 50, 10 to 40, 10 to 30, 10 to 25, 10 to 20, 10 to 15, 15 to 100, 15 to 50, 15 to 40, 15 to 30, 15 to 25, 15 to 20, 20 to 100, 20 to 50, 20 to 40, 20 to 30, 20 to 25, 25 to 100, 25 to 50, 25 to 40, 25 to 30, 30 to 100, 30 to 50, 30 to 40, 40 to 100, 40 to 50, or 50 to 100 nucleotides in length.
[0224] In certain embodiments, additional nucleotide sequences form a hairpin with the spacer sequence. Such secondary structures can increase the specificity of guide nucleic acids or engineered non-natural systems (see Kocak et al. (2019) Nat. Biotech. 37:657-66). In certain embodiments, the free energy change during hairpin formation is -20 kcal / mol, -15 kcal / mol, -14 kcal / mol, -13 kcal / mol, -12 kcal / mol, -11 kcal / mol, or -10 kcal / mol or greater. In certain embodiments, the free energy change during hairpin formation is -5 kcal / mol, -6 kcal / mol, -7 kcal / mol, -8 kcal / mol, -9 kcal / mol, -10 kcal / mol, -11 kcal / mol, -12 kcal / mol, -13 kcal / mol, -14 kcal / mol, or -15 kcal / mol or greater. In certain embodiments, the free energy change during hairpin formation is in the range of -20 to -10 kcal / mol, -20 to -11 kcal / mol, -20 to -12 kcal / mol, -20 to -13 kcal / mol, -20 to -14 kcal / mol, -20 to -15 kcal / mol, -15 to -10 kcal / mol, -15 to -11 kcal / mol, -15 to -12 kcal / mol, -15 to -13 kcal / mol, -15 to -14 kcal / mol, -14 to -10 kcal / mol, -14 to -11 kcal / mol, -14 to -12 kcal / mol, -14 to -13 kcal / mol, -13 to -10 kcal / mol, -13 to -11 kcal / mol, -13 to -12 kcal / mol, -12 to -10 kcal / mol, -12 to -11 kcal / mol, or -11 to -10 kcal / mol. In other embodiments, the targeter nucleic acid or single guide nucleic acid does not contain any nucleotides 3' to the spacer sequence.
[0225] In certain embodiments, the modulator nucleic acid further comprises an additional nucleotide sequence 3' to the modulator stem sequence. In certain embodiments, the additional nucleotide sequence comprises at least one (e.g., at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50) nucleotides. In certain embodiments, the additional nucleotide sequence consists of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides. The additional nucleotide sequence consists of one nucleotide (e.g., uridine). In certain embodiments, the additional nucleotide sequence consists of two nucleotides. In certain embodiments, the additional nucleotide sequence is similar to a loop or a fragment thereof (e.g., 1, 2, 3, or 4 nucleotides at the 5' end of the loop) in the crRNA of the corresponding single-guide CRISPR-Cas system. It is understood that the additional nucleotide sequence 3' to the modulator stem sequence may not be essential. Thus, in certain embodiments, the modulator nucleic acid does not comprise any additional nucleotides 3' to the modulator stem sequence.
[0226] It is understood that if there are additional nucleotide sequences on the 5' side relative to the targeter stem sequence and additional nucleotide sequences on the 3' side relative to the modulator stem sequence, they may interact with each other. For example, the nucleotide immediately on the 5' side relative to the targeter stem sequence and the nucleotide immediately on the 3' side relative to the modulator stem sequence do not form Watson-Crick type base pairs (originally, they could each constitute part of the targeter stem sequence and part of the modulator stem sequence), but other nucleotides in the additional nucleotide sequence on the 5' side relative to the targeter stem sequence and the additional nucleotide sequence on the 3' side relative to the modulator stem sequence may form one, two, three, or more base pairs (e.g., Watson-Crick type base pairs). Such interactions can affect the stability of the complex containing the targeter nucleic acid and the modulator nucleic acid.
[0227] The stability of a complex comprising a targeter nucleic acid and a modulator nucleic acid can be evaluated by the Gibbs free energy change (ΔG) during complex formation, either calculated or actually measured. If all predicted base pairings of the complex occur between the bases in the targeter nucleic acid and the bases in the modulator nucleic acid, i.e., when there is no intra-strand secondary structure, the ΔG during complex formation generally correlates with the ΔG during the formation of the secondary structure within the corresponding single-guide nucleic acid. Methods for calculating or measuring ΔG are known in the art. An exemplary method is RNAfold (rna.tbi.univie.ac.at / cgi-bin / RNAWebSuite / RNAfold.cgi) as disclosed in Gruber et al. (2008) Nucleic Acids Res., 36(Web Server issue):W70-W74. Unless otherwise indicated, the ΔG values in this disclosure are calculated by RNAfold for the formation of the secondary structure within the corresponding single-guide nucleic acid. In certain embodiments, ΔG is -1 kcal / mol or less, for example, -2 kcal / mol or less, -3 kcal / mol or less, -4 kcal / mol or less, -5 kcal / mol or less, -6 kcal / mol or less, -7 kcal / mol or less, -7.5 kcal / mol or less, or -8 kcal / mol or less. In certain embodiments, ΔG is -10 kcal / mol or more, for example, -9 kcal / mol or more, -8.5 kcal / mol or more, or -8 kcal / mol or more. In certain embodiments, ΔG is in the range of -10 to -4 kcal / mol. In certain embodiments, ΔG is in the range of -8 to -4 kcal / mol, -7 to -4 kcal / mol, -6 to -4 kcal / mol, -5 to -4 kcal / mol, -8 to -4.5 kcal / mol, -7 to -4.5 kcal / mol, -6 to -4.5 kcal / mol, or -5 to -4.5 kcal / mol.In certain embodiments, ΔG is about -8 kcal / mol, -7 kcal / mol, -6 kcal / mol, -5 kcal / mol, -4.9 kcal / mol, -4.8 kcal / mol, -4.7 kcal / mol, -4.6 kcal / mol, -4.5 kcal / mol, -4.4 kcal / mol, -4.3 kcal / mol, -4.2 kcal / mol, -4.1 kcal / mol, or -4 kcal / mol.
[0228] It is understood that ΔG can be affected by sequences in the targeter nucleic acid that are not within the targeter stem sequence and / or by sequences in the modulator nucleic acid that are not within the modulator stem sequence. For example, one or more base pairs (e.g., Watson-Crick type base pairs) between a further sequence 5' to the targeter stem sequence and a further sequence 3' to the modulator stem sequence can lower ΔG, i.e., can stabilize the nucleic acid complex. In certain embodiments, the nucleotide immediately 5' to the targeter stem sequence contains uracil or is uridine, and the nucleotide immediately 3' to the modulator stem sequence contains uracil or is uridine, thereby forming a non-conventional U-U base pair.
[0229] In certain embodiments, the modulator nucleic acid or single-guide nucleic acid comprises a nucleotide sequence herein referred to as the "5' tail" that is located 5' to the modulator stem sequence. In the native V-A type CRISPR-Cas system, the 5' tail is the nucleotide sequence located 5' to the stem-loop structure of the crRNA. The 5' tail in an engineered V-A type CRISPR-Cas system can be similar to the 5' tail in the corresponding native V-A type CRISPR-Cas system, whether the single-guide or dual-guide.
[0230] Although not restricted by theory, the 5' tail is thought to be able to participate in the formation of the CRISPR-Cas complex. For example, in certain embodiments, the 5' tail forms a pseudoknot structure with the modulator stem sequence, which is recognized by the Cas protein (see Yamano et al. (2016) Cell, 165:949). In certain embodiments, the 5' tail is at least 3 (e.g., at least 4 or at least 5) nucleotides in length. In certain embodiments, the 5' tail is 3, 4, or 5 nucleotides in length. In certain embodiments, the nucleotide at the 3' end of the 5' tail contains uracil or is uridine. In certain embodiments, the nucleotide at the second position counting from the 3' end in the 5' tail contains uracil or is uridine. In certain embodiments, the nucleotide at the third position counting from the 3' end in the 5' tail contains adenine or is adenosine. This third nucleotide can form a base pair (e.g., a Watson-Crick type base pair) with the nucleotide 5' to the modulator stem sequence. Thus, in certain embodiments, the modulator nucleic acid contains a uridine or uracil-containing nucleotide 5' to the modulator stem sequence. In certain embodiments, the 5' tail contains the nucleotide sequence 5'-AUU-3'. In certain embodiments, the 5' tail contains the nucleotide sequence 5'-AAUU-3'. In certain embodiments, the 5' tail contains the nucleotide sequence 5'-UAAUU-3'. In certain embodiments, the 5' tail is located immediately 5' to the modulator stem sequence.
[0231] In certain embodiments, the single guide nucleic acid, the targeter nucleic acid, and / or the modulator nucleic acid are designed to reduce the degree of secondary structure other than hybridization between the targeter stem sequence and the modulator stem sequence. In certain embodiments, about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1% or less, or less than that of the nucleotides of the single guide nucleic acid other than the targeter stem sequence and the modulator stem sequence are involved in self-complementary base pairing when optimally folded. In certain embodiments, about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1% or less, or less than that of the nucleotides of the targeter nucleic acid and / or the modulator nucleic acid are involved in self-complementary base pairing when optimally folded. Optimal folding can be determined by any suitable polynucleotide folding algorithm. Some programs are based on the calculation of the minimum Gibbs free energy. An example of such an algorithm is mFold as described in Zuker and Stiegler (Nucleic Acids Res. 9(1981), 133-148). Another example of a folding algorithm is the online web server RNAfold developed by the Institute for Theoretical Chemistry at the University of Vienna and using the centroid structure prediction algorithm (see, for example, A.R. Gruber et al., 2008, Cell 106(1):23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12):1151-62).
[0232] The targeter nucleic acid is directed to a specific target nucleotide sequence, and the donor template can be designed to modify the target nucleotide sequence or a sequence in its vicinity. Thus, it is understood that the binding of the single-guide nucleic acid, targeter nucleic acid, or modulator nucleic acid to the donor template can enhance the editing efficiency and reduce off-target effects. Thus, in certain embodiments, the single-guide nucleic acid or modulator nucleic acid further comprises a donor template recruitment sequence capable of hybridizing to the donor template (see FIG. 2B). The donor template is described in the "Donor Template" subsection of Section II below. The donor template and the donor template recruitment sequence can be designed such that they have sequence complementarity. In certain embodiments, the donor template recruitment sequence is at least 90% (e.g., at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) complementary to at least a portion of the donor template. In certain embodiments, the donor template recruitment sequence is 100% complementary to at least a portion of the donor template. In certain embodiments, where the donor template comprises a engineered sequence that is not homologous to the sequence to be repaired, the donor template recruitment sequence can hybridize to the engineered sequence in the donor template. In certain embodiments, the donor template recruitment sequence is at least 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides in length. In certain embodiments, the donor template recruitment sequence is located at or near the 5' end of the single-guide nucleic acid or at or near the 5' end of the modulator nucleic acid. In certain embodiments, the donor template recruitment sequence is linked to the 5' tail of the single-guide nucleic acid or modulator nucleic acid, if present, or to the modulator stem sequence via a polynucleotide linkage or nucleotide linker.
[0233] In certain embodiments, the single guide nucleic acid or modulator nucleic acid further comprises an editing enhancer sequence, which enhances the efficiency of gene editing and / or homologous recombination repair (HDR) (see Figure 2C). Exemplary editing enhancer sequences are described in Park et al. (2018) Nat. Commun. 9:3313. In certain embodiments, the editing enhancer sequence is located 5' to the 5' tail, if present, or 5' to the single guide nucleic acid or modulator stem sequence. In certain embodiments, the editing enhancer sequence is 1-50, 4-50, 9-50, 15-50, 25-50, 1-25, 4-25, 9-25, 15-25, 1-15, 4-15, 9-15, 1-9, 4-9, or 1-4 nucleotides in length. In certain embodiments, the editing enhancer sequence is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, or 55 nucleotides in length. The editing enhancer sequence is designed to minimize homology with the target nucleotide sequence or any other sequence that the engineered non-native system may contact, e.g., the genomic sequence of the cell into which the engineered non-native system is delivered. In certain embodiments, the editing enhancer is designed to minimize the presence of hairpin structures. The editing enhancer may comprise one or more of the chemical modifications disclosed herein.
[0234] The single-guide nucleic acid, modulator nucleic acid, and / or targeter nucleic acid may further comprise a protective nucleotide sequence that prevents or reduces nucleic acid degradation. In certain embodiments, the protective nucleotide sequence is at least 5 (e.g., at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50) nucleotides in length. The length of the protective nucleotide sequence increases the time for the exonuclease to reach the 5' tail, modulator stem sequence, targeter stem sequence, and / or spacer sequence, thereby protecting these portions of the single-guide nucleic acid, modulator nucleic acid, and / or targeter nucleic acid from degradation by the exonuclease. In certain embodiments, the protective nucleotide sequence forms a secondary structure, such as a hairpin or tRNA structure, to reduce the rate of degradation by the exonuclease (see, e.g., Wu et al. (2018) Cell. Mol. Life Sci., 75(19):3593-3607). The secondary structure can be predicted by an online web server RNAfold, which is developed by the University of Vienna and uses the centroid structure prediction algorithm, a method known in the art (see Gruber et al. (2008) Nucleic Acids Res., 36:W70). Specific chemical modifications that may be present in the protective nucleotide sequence may also prevent or reduce nucleic acid degradation, as disclosed in the following "RNA Modifications" subsection.
[0235] The protecting nucleotide sequence is typically located at the 5' or 3' end of the single-guide nucleic acid, the modulator nucleic acid, and / or the targeter nucleic acid. In certain embodiments, the single-guide nucleic acid comprises the protecting nucleotide sequence at the 5' end, 3' end, or both ends, optionally via a nucleotide linker. In certain embodiments, the modulator nucleic acid comprises the protecting nucleotide sequence at the 5' end, 3' end, or both ends, optionally via a nucleotide linker. In certain embodiments, the modulator nucleic acid comprises a protecting nucleotide sequence at the 5' end (see Figure 2A). In certain embodiments, the targeter nucleic acid comprises the protecting nucleotide sequence at the 5' end, 3' end, or both ends, optionally via a nucleotide linker.
[0236] As described above, without limitation, various nucleotide sequences including, but not limited to, donor template recruit arrays, editing enhancer arrays, protective nucleotide sequences, and linkers that couple such sequences, when present, to the 5' tail or to the modulator stem array, can be present in the 5' portion of a single nucleic acid or a modulator nucleic acid. It is understood that the functions of donor template recruitment, editing enhancement, protection against degradation, and binding are not mutually exclusive, and that one nucleotide sequence can have one or more of such functions. For example, in certain embodiments, a single guide nucleic acid or a modulator nucleic acid comprises a nucleotide sequence that is both a donor template recruit array and an editing enhancer array. In certain embodiments, a single guide nucleic acid or a modulator nucleic acid comprises a nucleotide sequence that is both a donor template recruit array and a protective sequence. In certain embodiments, a single guide nucleic acid or a modulator nucleic acid comprises a nucleotide sequence that is both an editing enhancer array and a protective sequence. In certain embodiments, a single guide nucleic acid or a modulator nucleic acid comprises a nucleotide sequence that is a donor template recruit array, an editing enhancer array, and a protective sequence. In certain embodiments, the nucleotide sequence 5' to the 5' tail, when present, or 5' to the modulator stem array is 1 to 90, 1 to 80, 1 to 70, 1 to 60, 1 to 50, 1 to 40, 1 to 30, 1 to 20, 1 to 10, 10 to 90, 10 to 80, 10 to 70, 10 to 60, 10 to 50, 10 to 40, 10 to 30, 10 to 20, 20 to 90, 20 to 80, 20 to 70, 20 to 60, 20 to 50, 20 to 40, 20 to 30, 30 to 90, 30 to 80, 30 to 70, 30 to 60, 30 to 50, 30 to 40, 40 to 90, 40 to 80, 40 to 70, 40 to 60, 40 to 50, 50 to 90, 50 to 80, 50 to 70, 50 to 60, 60 to 90, 60 to 80, 60 to 70, 70 to 90, 70 to 80, or 80 to 90 nucleotides in length.
[0237] In certain embodiments, the engineered non-natural system further comprises one or more compounds (e.g., small molecule compounds) that enhance HDR and / or inhibit NHEJ. Exemplary compounds having such functions are described in Maruyama et al. (2015) Nat Biotechnol. 33(5):538-42; Chu et al. (2015) Nat Biotechnol. 33(5):543-48; Yu et al. (2015) Cell Stem Cell 16(2):142-47; Pinder et al. (2015) Nucleic Acids Res. 43(19):9379-92; and Yagiz et al. (2019) Commun. Biol. 2:198. In certain embodiments, the engineered non-natural system further comprises one or more compounds selected from the group consisting of DNA ligase IV antagonists (e.g., SCR7 compound, Ad4 E1B55K protein, and Ad4 E4orf6 protein), RAD51 agonists (e.g., RS-1), DNA-dependent protein kinase (DNA-PK) antagonists (e.g., NU7441 and KU0060648), β3 adrenergic receptor agonists (e.g., L755507), inhibitors of intracellular protein transport from the ER to the Golgi apparatus (e.g., brefeldin A), and any combination thereof.
[0238] In certain embodiments, the engineered non-natural system comprising a targeter nucleic acid and a modulator nucleic acid is regulatable or inducible. For example, in certain embodiments, the targeter nucleic acid, the modulator nucleic acid, and / or the Cas protein can be introduced into the target nucleotide sequence at different times, and the system becomes active only when all components are present. In certain embodiments, the amounts of the targeter nucleic acid, the modulator nucleic acid, and / or the Cas protein can be adjusted to achieve the desired efficiency and specificity. In certain embodiments, an excess amount of nucleic acid comprising the targeter stem sequence or the modulator stem sequence can be added to the system, thereby dissociating the complex of the targeter nucleic acid and the modulator nucleic acid and turning the system off.
[0239] C.gNA modification Guide nucleic acids comprising a single guide nucleic acid, a targeter nucleic acid, and / or a modulator nucleic acid can include DNA (e.g., modified DNA), RNA (e.g., modified RNA), or combinations thereof. In certain embodiments, the single guide nucleic acid includes DNA (e.g., modified DNA), RNA (e.g., modified RNA), or combinations thereof. In certain embodiments, the targeter nucleic acid includes DNA (e.g., modified DNA), RNA (e.g., modified RNA), or combinations thereof. In certain embodiments, the modulator nucleic acid includes DNA (e.g., modified DNA), RNA (e.g., modified RNA), or combinations thereof. The spacer sequence can be represented as a DNA sequence by including thymidine (T) instead of uridine (U). It is understood that corresponding RNA sequences and DNA / RNA chimeric sequences are also contemplated. For example, if the spacer sequence is RNA, the sequence can be derived from the DNA sequences disclosed herein by replacing each T with U. As a result, T and U are used synonymously herein to describe nucleotide sequences.
[0240] In certain embodiments, an engineered non - natural system comprising a targeter nucleic acid comprising: a target nucleotide sequence and a spacer sequence designed to hybridize to a targeter stem sequence; and a modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence, and optionally, a 5' sequence, e.g., a tail sequence, wherein in a single - guide nucleic acid, the targeter nucleic acid and the modulator nucleic acid are part of a single polynucleotide, and in a dual - guide nucleic acid, the targeter nucleic acid and the modulator nucleic acid are separate nucleic acids; the modification can comprise one or more chemical modifications to one or more nucleotides or internucleotide linkages at or near the 3' end of the targeter nucleic acid (dual and single gNA), at or near the 5' end of the targeter nucleic acid (dual gNA), at or near the 3' end of the modulator nucleic acid (dual gNA), at or near the 5' end of the modulator nucleic acid (single and dual gNA), or, as needed, combinations thereof for single or dual gNA. In certain embodiments, the Cas nuclease is a type V - A Cas nuclease. The modulator and / or targeter nucleic acid sequences can comprise additional sequences as detailed in the guide nucleic acid section, and the modification can be within these additional sequences as needed and as will be apparent to one of ordinary skill in the art. In the embodiments described in this section below, in certain embodiments, the guide nucleic acid is oriented from the 5' side in the modulator nucleic acid to the 3' side in the modulator stem sequence, and from the 5' side in the targeter stem sequence to the 3' side in the targeter sequence (e.g., see FIGS. 1A and 1B); in certain embodiments, optionally, the guide nucleic acid is oriented from the 3' side in the modulator nucleic acid to the 5' side in the modulator stem sequence, and from the 3' side in the targeter stem sequence to the 5' side in the targeter sequence.
[0241] The target nucleic acid can include DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. The modulator nucleic acid can include DNA (e.g., modified DNA), RNA (e.g., modified RNA), or a combination thereof. In certain embodiments, the target nucleic acid is RNA and the modulator nucleic acid is RNA. The target nucleic acid in the form of RNA is also referred to as target RNA, and the modulator nucleic acid in the form of RNA is also referred to as modulator RNA. The nucleotide sequences disclosed herein are presented as DNA sequences by including RNA sequences that contain thymidine (T) and / or uridine (U). It is understood that corresponding DNA sequences, RNA sequences, and DNA / RNA chimeric sequences are also contemplated. For example, if a spacer sequence is presented as a DNA sequence, a nucleic acid containing this spacer sequence as RNA can be derived from the DNA sequences disclosed herein by replacing each T with U. As a result, for the purpose of describing nucleotide sequences, T and U are used synonymously herein.
[0242] In certain embodiments, some or all of the gNAs are RNAs, such as gRNAs. In certain embodiments, 5-100%, 10-100%, 20-100%, 30-100%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, 95-100%, 99-100%, 99.5-100% of the gNAs are gRNAs. In certain embodiments, 20%-80%, 20%-70%, 20%-60%, 20%-50%, 20%-40%, 20%-30%, 30%-80%, 30%-70%, 30%-60%, 30%-50%, 30%-40%, 40%-80%, 40%-70%, 40%-60%, 40%-50%, 50%-80%, 50%-70%, 50%-60%, 60%-80%, 60%-70%, or 70%-80% of the gNAs are RNAs. In certain embodiments, 50% of the gNAs are RNAs. In certain embodiments, 70% of the gNAs are RNAs. In certain embodiments, 90% of the gNAs are RNAs. In certain embodiments, 100% of the gNAs are RNAs, such as gRNAs. In further embodiments, the remaining portion of the gNAs that are not RNAs are modified ribonucleotides, deoxyribonucleotides, modified deoxyribonucleotides, or synthetic, e.g., unnatural nucleotides, such as, but not limited to, threose nucleic acids, locked nucleic acids, peptide nucleic acids, arabino nucleic acids, hexose nucleic acids.
[0243] In certain embodiments, the targeter nucleic acid and / or the modulator nucleic acid is an RNA having one or more modifications in the ribose group, one or more modifications in the phosphate group, one or more modifications in the nucleobase, one or more terminal modifications, or combinations thereof. Exemplary modifications are disclosed in U.S. Pat. Nos. 10,900,034 and 10,767,175, U.S. Patent Application Publication No. 2018 / 0119140, Watts et al. (2008) Drug Discov. Today 13:842-55, and Hendel et al. (2015) NAT. BIOTECHNOL. 33:985.
[0244] In certain embodiments, the targeter nucleic acid, e.g., RNA, comprises at least one nucleotide at or near the 3'-end that includes modifications to ribose, phosphate groups, nucleobases, or terminal modifications. In certain embodiments, the 3'-end of the targeter nucleic acid comprises a spacer sequence. In certain embodiments, the 3'-end of the targeter nucleic acid comprises a targeter stem sequence. Exemplary modifications are disclosed in Dang et al. (2015) Genome Biol. 16:280, Kocaz et al. (2019) Nature Biotech. 37:657-66, Liu et al. (2019) Nucleic Acids Res. 47(8):4169-4180, Schubert et al. (2018) J. Cytokine Biol. 3(1):121, Teng et al. (2019) Genome Biol. 20(1):15, Watts et al. (2008) Drug Discov. Today 13(19-20):842-55, and Wu et al. (2018) Cell Mol. Life. Sci. 75(19):3593-607.
[0245] Modifications in the ribose group include, but are not limited to, modifications at the 2'-position or modifications at the 4'-position. For example, in certain embodiments, the ribose includes 2'-O-C1-4 alkyl, such as 2'-O-methyl (2'-OMe, or M). In certain embodiments, the ribose includes 2'-O-C1-3 alkyl-O-C1-3 alkyl, such as 2'-methoxyethoxy (2'-O-CH2CH2OCH3), also known as 2'-O-(2-methoxyethyl) or 2'-MOE. In certain embodiments, the ribose includes 2'-O-allyl. In certain embodiments, the ribose includes 2'-O-2,4-dinitrophenol (DNP). In certain embodiments, the ribose includes 2'-halo, such as 2'-F, 2'-Br, 2'-Cl, or 2'-I. In certain embodiments, the ribose includes 2'-NH2. In certain embodiments, the ribose includes 2'-H (e.g., deoxynucleotide). In certain embodiments, the ribose includes 2'-arabino or 2'-F-arabino. In certain embodiments, the ribose includes 2'-LNA or 2'-ULNA. In certain embodiments, the ribose includes 4'-thioglucosyl.
[0246] Modifications may also include deoxy groups, such as 2'-deoxy-3'-phosphonoacetate (DP), 2'-deoxy-3'-thiophosphonoacetate (DSP).
[0247] Nucleotide internucleotide bond modifications in the phosphate group include, but are not limited to, phosphorothioate (S), chiral phosphorothioate, phosphorodithioate, boranophosphonate, C 1~4Examples of modified phosphates include alkyl phosphonates such as methyl phosphonate, borano phosphonate, phosphonocarboxylates such as phosphonoacetate (P), phosphonocarboxylic acid esters such as phosphonoacetate ester, amides, thiophosphonocarboxylates such as thiophosphonoacetate (SP), thiophosphonocarboxylic acid esters such as thiophosphonoacetate ester, and phosphodiesters or any of the above modified phosphates having a 2′,5′-linkage. Also included are various salts, mixed salts, and free acid forms.
[0248] Examples of modifications to nucleobases include, but are not limited to, 2-thiouracil, 2-thiocytosine, 4-thiouracil, 6-thioguanine, 2-aminoadenine, 2-aminopurine, pseudouracil, hypoxanthine, 7-deazaguanine, 7-deaza-8-azaguanine, 7-deazaadenine, 7-deaza-8-azaadenine, 5-methylcytosine, 5-methyluracil, 5-hydroxymethylcytosine, 5-hydroxymethyluracil, 5,6-dihydrouracil, 5-propynylcytosine, 5-propynyluracil, 5-ethynylcytosine, 5-ethynyluracil, 5-allyluracil, 5-allylcytosine, 5-aminoallyluracil, 5-aminoallyl-cytosine, 5-bromouracil, 5-iodouracil, diaminopurine, difluorotoluene, dihydrouracil, abasic nucleotides, Z base, P base, unstructured nucleic acids, isoguanine, isocytosine (see Piccirilli et al. (1990) NATURE, 343:33), 5-methyl-2-pyrimidine (see Rappaport (1993) BIOCHEMISTRY, 32:3047), x(A, G, C, T), and y(A, G, C, T).
[0249] Examples of terminal modifications include, but are not limited to, polyethylene glycol (PEG), hydrocarbon linkers (heteroatom (O, S, N)-substituted hydrocarbon spacers; halo-substituted hydrocarbon spacers; keto-, carboxyl-, amide-, thionyl-, carbamoyl-, thiocarbamoyl-containing hydrocarbon spacers, propanediol), spermine linkers, dyes such as fluorescent dyes (e.g., fluorescein, rhodamine, cyanine), quenchers (e.g., dabsyl, BHQ), and other labels (e.g., biotin, digoxigenin, acridine, streptavidin, avidin, peptides and / or proteins, etc.). In certain embodiments, the terminal modification comprises conjugation (or ligation) of the RNA with another molecule comprising an oligonucleotide (such as deoxyribonucleotides and / or ribonucleotides), peptide, protein, sugar, oligosaccharide, steroid, lipid, folic acid, vitamin, and / or other molecule. In certain embodiments, the terminal modification incorporated into the RNA is incorporated as a phosphodiester bond within the interior of the RNA sequence and can be located anywhere between two nucleotides in the RNA via a linker such as a 2-(4-butylamidofluorescein)propane-1,3-diol bis(phosphodiester) linker.
[0250] The modifications disclosed above can be combined in the target nucleic acid and / or modulator nucleic acid in the form of RNA. In certain embodiments, the modification in the RNA is selected from the group consisting of incorporation of 2'-O-methyl-3'-phosphorothioate (MS), 2'-O-methyl-3'-phosphonoacetate (MP), 2'-O-methyl-3'-thiophosphonoacetate (MSP), 2'-halo-3'-phosphorothioate (e.g., 2'-fluoro-3'-phosphorothioate), 2'-halo-3'-phosphonoacetate (e.g., 2'-fluoro-3'-phosphonoacetate), and 2'-halo-3'-thiophosphonoacetate (e.g., 2'-fluoro-3'-thiophosphonoacetate).
[0251] In certain embodiments, modifications may include 2'-O-methyl (M), phosphorothioate (S), phosphonoacetate (P), thiophosphonoacetate (SP), 2'-O-methyl-3'-phosphorothioate (MS), 2'-O-methyl-3'-phosphonoacetate (MP), 2'-O-methyl-3'-thiophosphonoacetate (MSP), 2'-deoxy-3'-phosphonoacetate (DP), 2'-deoxy-3'-thiophosphonoacetate (DSP), or combinations thereof, at or near the 3' or 5' end of either the targeter or modulator nucleic acid, as appropriate for single or dual gNAs. In certain embodiments, modifications may include either 5' or 3' propanediol or C3 linker modifications.
[0252] In certain embodiments, the modifications alter the stability of the RNA. In certain embodiments, the modifications improve the stability of the RNA, for example, by increasing the nuclease resistance of the RNA compared to the corresponding RNA without the modification. Stability-improving modifications include, but are not limited to, 2'-O-methyl, 2'-OC 1~4 Alkyl, 2'-halo (e.g., 2'-F, 2'-Br, 2'-Cl, or 2'-I), 2'MOE, 2'-OC 1~3 These include the incorporation of alkyl-O-C1-3 alkyl, 2'-NH2, 2'-H (or 2'-deoxy), 2'-arabino, 2'-F-arabino, 4'-thioribosyl sugar moieties, 3'-phosphorothioates, 3'-phosphonoacetates, 3'-thiophosphonoacetates, 3'-methylphosphonates, 3'-boranophosphates, 3'-phosphorodithioates, locked nucleic acid ("LNA") nucleotides containing a methylene bridge between the 2' and 4' carbons of the ribose ring, and unlocked nucleic acid ("ULNA") nucleotides. Such modifications are suitable for use as protecting groups to prevent or reduce degradation of 5' sequences, e.g., tail sequences, modulator stem sequences (dual guide nucleic acids), targeter stem sequences (dual guide nucleic acids), and / or spacer sequences (see the "Targeter and Modulator Nucleic Acids" subsection).
[0253] In certain embodiments, the modification modifies the specificity of the engineered non-natural system. In certain embodiments, the modification improves the specificity of the engineered non-natural system, for example, by improving off-target binding and / or cleavage, or reducing off-target binding and / or cleavage, or a combination thereof. Specificity-improving modifications include, but are not limited to, 2-thiouracil, 2-thiocytosine, 4-thiouracil, 6-thioguanine, 2-aminoadenine, and pseudouracil. The modification is within 10, 5, 4, 3, 2, or 1 nucleotide of the 3' end, for example, the 3' terminal nucleotide is modified.
[0254] In certain embodiments, the modification modifies the immunostimulatory effect of the RNA as compared to the corresponding unmodified RNA. For example, in certain embodiments, the modification reduces the ability of the RNA to activate TLR7, TLR8, TLR9, TLR3, RIG-I, and / or MDA5.
[0255] In certain embodiments, the targeter nucleic acid and / or the modulator nucleic acid comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 modified nucleotides or internucleotide linkages. The modifications can be made at one or more positions in the targeter nucleic acid and / or the modulator nucleic acid such that these nucleic acids retain functionality. For example, the modified nucleic acid can still direct the Cas protein to the target nucleotide sequence and allow the Cas protein to exert its effector function. It is understood that a particular modification at a position can be selected based on the functionality of the nucleotide or internucleotide linkage at that position. For example, specificity-improving modifications can be suitable for nucleotides or internucleotide linkages in the spacer sequence, the targeter stem sequence, or the modulator stem sequence. Stability-improving modifications can be suitable for one or more terminal nucleotides or internucleotide linkages in the targeter nucleic acid and / or the modulator nucleic acid. In certain embodiments, at least 1 (e.g., at least 2, at least 3, at least 4, or at least 5) terminal nucleotides or internucleotide linkages at or near the 5' end and / or at least 1 (e.g., at least 2, at least 3, at least 4, or at least 5) terminal nucleotides or internucleotide linkages at or near the 3' end of the targeter nucleic acid are modified. In certain embodiments, 5 or fewer (e.g., 1 or fewer, 2 or fewer, 3 or fewer, or 4 or fewer) terminal nucleotides or internucleotide linkages at or near the 5' end and / or 5 or fewer (e.g., 1 or fewer, 2 or fewer, 3 or fewer, or 4 or fewer) terminal nucleotides or internucleotide linkages at or near the 3' end of the targeter nucleic acid are modified.In certain embodiments, at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide linkages at or near the 5′ end and / or at least one (e.g., at least two, at least three, at least four, or at least five) terminal nucleotides or internucleotide linkages at or near the 3′ end of the modulator nucleic acid are modified. In certain embodiments, five or fewer (e.g., one or fewer, two or fewer, three or fewer, or four or fewer) terminal nucleotides or internucleotide linkages at or near the 5′ end and / or five or fewer (e.g., one or fewer, two or fewer, three or fewer, or four or fewer) terminal nucleotides or internucleotide linkages at or near the 3′ end of the modulator nucleic acid are modified. The selection of positions for modification is described in U.S. Patent Nos. 10,900,034 and 10,767,175. As used in this paragraph, when the targeter or modulator nucleic acid is a combination of DNA and RNA, the nucleic acid as a whole is considered to be RNA, and the DNA nucleotides are considered to be modifications of RNA, including 2′-H modification of ribose and optionally modification of the nucleobase.
[0256] In a dual guide nucleic acid system, the targeter nucleic acid and the modulator nucleic acid are not in the same nucleic acid, i.e., they are not linked end-to-end via a conventional polynucleotide linkage, while they can be conjugated to each other covalently via one or more chemical modifications introduced into these nucleic acids, thereby enhancing the stability of the double-stranded complex and / or improving other characteristics of the system.
[0257] III. Compositions and Methods for Targeting, Editing, and / or Modifying Genomic DNA Engineered non-natural systems such as those disclosed herein can be useful for targeting, editing, and / or modifying target nucleic acids, e.g., DNA (e.g., genomic DNA) within a cell or organism.
[0258] The present invention provides a method for cleaving a target nucleic acid (e.g., DNA) containing a preselected target sequence or a sequence of a portion thereof, comprising contacting the target DNA with an engineered non-natural system disclosed herein, thereby resulting in cleavage of the target DNA.
[0259] Furthermore, the present invention provides a method for binding a target nucleic acid (e.g., DNA) containing a preselected target sequence or a sequence of a portion thereof, comprising contacting the target DNA with an engineered non-natural system disclosed herein, thereby resulting in binding of the system to the target DNA. This method is useful, for example, for detecting the presence and / or location of a preselected target gene when a component of the system (e.g., a Cas protein) contains a detectable marker.
[0260] Furthermore, the present invention provides a method for modifying a target nucleic acid (e.g., DNA) containing a preselected target sequence or a sequence of a portion thereof, or a structure (e.g., a protein) associated with the target DNA (e.g., a histone protein in a chromosome), comprising contacting the target DNA with an engineered non-natural system disclosed herein, thereby resulting in modification of the target DNA or a structure associated with the target DNA, wherein the Cas protein contains an effector domain or is associated with an effector protein. The modification corresponds to the function of the effector domain or effector protein. The exemplary functions described in the "Cas protein" subsection in Section I above are applicable herein.
[0261] The engineered non-natural system can be contacted with the target nucleic acid as a complex. Thus, in certain embodiments, the method comprises contacting the target nucleic acid with a CRISPR-Cas complex comprising a targeter nucleic acid, a modulator nucleic acid, and a Cas protein disclosed herein. In certain embodiments, the Cas protein is a type V-A, V-C, or V-D Cas protein (e.g., a Cas nuclease). In certain embodiments, the Cas protein is a type V-A Cas protein (e.g., a Cas nuclease).
[0262] In certain embodiments, provided is a method of editing a human genomic sequence at one of a preselected group of target loci, the method comprising delivering into a human cell an engineered non-natural system disclosed herein, thereby effecting editing of the genomic sequence at the target locus in the human cell. In certain embodiments, provided herein is a method of detecting a human genomic sequence at one of a preselected group of target loci, the method comprising delivering into a human cell an engineered non-natural system disclosed herein, thereby detecting the target locus in the human cell, wherein a component of the system (e.g., a Cas protein) comprises a detectable marker. In certain embodiments, provided herein is a method of modifying a human chromosome at one of a preselected group of target loci, the method comprising delivering into a human cell an engineered non-natural system disclosed herein, thereby effecting modification of the chromosome at the target locus in the human cell, wherein the Cas protein comprises an effector domain or is associated with an effector protein.
[0263] The CRISPR-Cas complex can be delivered to a cell by introducing a pre-formed ribonucleoprotein (RNP) complex into the cell. Alternatively, one or more components of the CRISPR-Cas complex can be expressed intracellularly. Exemplary methods of delivery are known in the art and are described, for example, in U.S. Patent Nos. 8,697,359; 10,113,167; 10,570,418; 10,829,787; 11,118,194; and 11,125,739 and U.S. Patent Application Publication Nos. 2015 / 0344912; 2018 / 0119140; and 2018 / 0282763.
[0264] It is understood that contacting DNA within a cell (e.g., genomic DNA) with a CRISPR-Cas complex does not require delivery of all components of the complex into the cell. For example, one or more of the components may already be present within the cell. In certain embodiments, the cell (or its parental / ancestral cell) has been engineered to express a Cas protein, and a single-guide nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the single-guide nucleic acid), a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the targeter nucleic acid), and / or a modulator nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the modulator nucleic acid) is delivered into the cell. In certain embodiments, the cell (or its parental / ancestral cell) has been engineered to express a modulator nucleic acid, and a Cas protein (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the Cas protein) and a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the targeter nucleic acid) are delivered into the cell. In certain embodiments, the cell (or its parental / ancestral cell) and the modulator nucleic acid have been engineered to express a Cas protein, and a targeter nucleic acid (or a nucleic acid comprising a regulatory element operably linked to a nucleotide sequence encoding the targeter nucleic acid) is delivered into the cell.
[0265] In certain embodiments, the target DNA is within the genome of the target cell. Accordingly, the invention also provides a cell comprising a non-natural system or a CRISPR expression system described herein. Further, the invention provides a cell whose genome has been modified by a CRISPR-Cas system or complex disclosed herein.
[0266] Target cells can be mitotic cells or post-mitotic cells derived from any organism, such as bacterial cells (e.g., Escherichia coli), archaeal cells, cells of unicellular eukaryotes, plant cells, algal cells, such as Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens C. Agardh, etc., fungal cells (e.g., yeast cells, such as S. cerevisiae), animal cells, cells derived from invertebrates (e.g., Drosophila, cnidarians, echinoderms, nematodes, etc.), cells derived from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells derived from mammals, cells derived from rodents, or cells derived from humans. Examples of target cell types include, but are not limited to, stem cells (e.g., embryonic stem (ES) cells, induced pluripotent stem (iPS) cells, germ cells), somatic cells (e.g., fibroblasts, hematopoietic cells, T lymphocytes (e.g., CD8+ T lymphocytes), NK cells, neurons, muscle cells, bone cells, hepatocytes, pancreatic cells), in vitro or in vivo embryonic cells at any stage of the embryo (e.g., 1-cell stage, 2-cell stage, 4-cell stage, 8-cell stage zebrafish embryos). The cells may be derived from an established cell line or may be primary cells (i.e., cells and cell cultures that are derived from a subject and have been grown in vitro for a limited number of passages). For example, a primary culture may be one that has been passaged 0, 1, 2, 4, 5, 10, or 15 times or less, but not enough times to reach the transformation stage. Typically, a primary cell line is maintained in vitro for fewer than 10 passages. When the cells are primary cells, they can be collected from an individual by any suitable method. For example, white blood cells can be collected by apheresis, leukapheresis, or density gradient separation, while cells from tissues such as skin, muscle, bone marrow, spleen, liver, pancreas, lung, intestine, or stomach can be collected by biopsy.The cells taken may be used immediately or stored with a cryoprotectant under cryogenic conditions and thawed at a later time by methods generally known in the art.
[0267] A. Ribonucleoprotein (RNP) delivery and “casRNA” delivery The engineered non-native systems disclosed herein can be delivered into cells by suitable methods known in the art, including but not limited to ribonucleoprotein (RNP) delivery and “CasRNA” delivery as described hereinafter.
[0268] In certain embodiments, a CRISPR-Cas system comprising a single guide nucleic acid and a Cas protein, or a CRISPR-Cas system comprising a targeter nucleic acid, a modulator nucleic acid, and a Cas protein, can be combined into an RNP complex and then delivered into cells as a pre-formed complex. This method is suitable for the active modification of a cell's genetic or epigenetic information for a limited period. For example, if the Cas protein has nuclease activity for modifying the genomic DNA of a cell, the nuclease activity only needs to be maintained for a period to enable DNA cleavage, and long-term nuclease activity can increase off-target effects. Similarly, some specific epigenetic modifications can be maintained within the cell once established and inherited by daughter cells.
[0269] "Ribonucleoprotein" or "RNP", as used herein, may refer to a complex comprising a nuclear protein and ribonucleic acid. As used herein, "nuclear protein" may refer to a protein capable of binding to a nucleic acid (e.g., RNA, DNA). When a nuclear protein binds to ribonucleic acid, it may be referred to as a "ribonucleoprotein". The interaction between a ribonucleoprotein and ribonucleic acid may be direct, for example, by a covalent bond, or indirect, for example, by non-covalent bonds (e.g., electrostatic interactions (e.g., ionic bonds, hydrogen bonds, halogen bonds), van der Waals interactions (e.g., dipole-dipole, dipole-induced dipole, London dispersion forces), ring stacking (π effects), hydrophobic interactions, etc.). In certain embodiments, the ribonucleoprotein comprises an RNA-binding motif non-covalently bound to ribonucleic acid. For example, positively charged aromatic amino acid residues (e.g., lysine residues) in the RNA-binding motif may form electrostatic interactions with the negatively charged nucleic acid backbone of RNA.
[0270] To ensure efficient loading of the Cas protein, a single-guide nucleic acid, or a combination of a targeter nucleic acid and a modulator nucleic acid, may be provided in a molar excess (e.g., at least 2-fold, at least 3-fold, at least 4-fold, or at least 5-fold) compared to the Cas protein. In certain embodiments, the targeter nucleic acid and the modulator nucleic acid are annealed under suitable conditions prior to forming a complex with the Cas protein. In other embodiments, the targeter nucleic acid, the modulator nucleic acid, and the Cas protein are directly mixed together to form an RNP.
[0271] Using various delivery methods, the RNPs disclosed herein can be introduced into cells. Exemplary delivery methods or vehicles include, but are not limited to, microinjection, liposomes (see, e.g., U.S. Patent No. 10,829,787), e.g., molecular trojan horse liposomes that deliver molecules across the blood-brain barrier (see Pardridge et al. (2010) Cold Spring Harb. Protoc., doi:10.1101 / pdb.prot5407), immunoliposomes, virosomes, microvesicles (e.g., exosomes and ARMM), polycations, lipid:nucleic acid conjugates, electroporation, cell-penetrating peptides (see U.S. Patent No. 11,118,194), nanoparticles, nanowires (see Shalek et al. (2012) Nano Letters, 12:6498), exosomes, and perturbation of the cell membrane (e.g., by passing cells through constrictions in a microfluidic system, see U.S. Patent No. 11,125,739). When the target cells are proliferative cells, the efficiency of RNP delivery can be improved by cell cycle synchronization (see U.S. Patent No. 10,570,418). In certain embodiments, the RNPs are delivered into cells by electroporation.
[0272] In certain embodiments, the CRISPR-Cas system is delivered into a cell by a method, i.e., by delivering (a) a single guide nucleic acid, or a combination of a targeter nucleic acid and a modulator nucleic acid, and (b) an RNA encoding a Cas protein (e.g., messenger RNA (mRNA)). The RNA encoding the Cas protein is translated intracellularly and can form a complex intracellularly with the single guide nucleic acid or the combination of the targeter nucleic acid and the modulator nucleic acid. Similar to the RNP method, the RNA has a limited half-life intracellularly even if a stability-improving modification is performed on one or more of the RNAs. Thus, the "Cas RNA" method is suitable for the active modification of the genetic or epigenetic information of cells for a limited period, e.g., DNA cleavage, and has the advantage of reducing off-target effects.
[0273] The mRNA can be produced by transcription of DNA containing regulatory elements operably linked to the Cas coding sequence. Considering that multiple copies of the Cas protein can be generated from one mRNA, the single guide nucleic acid, or the targeter nucleic acid and the modulator nucleic acid are generally provided in a molar excess (e.g., at least 5-fold, at least 10-fold, at least 20-fold, at least 30-fold, at least 50-fold, or at least 100-fold) compared to the mRNA. In certain embodiments, the targeter nucleic acid and the modulator nucleic acid are annealed under suitable conditions prior to delivery into the cell. In other embodiments, the targeter nucleic acid and the modulator nucleic acid are delivered into the cell without annealing in vitro.
[0274] Using various delivery systems, the "Cas RNA" system can be introduced into cells. Non-limiting examples of delivery methods or vehicles include microinjection, biolistic particles, liposomes (see, e.g., U.S. Patent No. 10,829,787), e.g., Trojan horse liposomes of molecular trojans that deliver molecules across the blood-brain barrier (Pardridge et al. (2010) Cold Spring Harb. Protoc., doi:10.1101 / pdb.prot5407), immunoliposomes, virosomes, polycations, lipid:nucleic acid conjugates, electroporation, nanoparticles, nanowires (see Shalek et al. (2012) Nano Letters, 12:6498), exosomes, and perturbation of cell membranes (e.g., by passing cells through constrictions in a microfluidic system, see U.S. Patent No. 11,125,739). A specific example of the "nucleic acid only" approach by electroporation is described in International (PCT) Publication No. WO 2016 / 164356 pamphlet.
[0275] In certain embodiments, the CRISPR-Cas system is delivered into cells in the form of DNA comprising (a) a combination of a single guide nucleic acid or targeter nucleic acid and a modulator nucleic acid, and (b) regulatory elements operably linked to the Cas coding sequence. The DNA can be provided in the form of a plasmid, viral vector, or any other form described in the "CRISPR Expression Systems" subsection. Such delivery methods can result in constitutive expression of the Cas protein in the target cells (e.g., when the DNA is maintained intracellularly by an episomal vector or integrated into the genome), and when the Cas protein has nuclease activity, it can increase the risk of unwanted off-target effects. Nevertheless, this approach is useful when the Cas protein comprises a non-nuclease effector (e.g., a transcriptional activator or repressor). It is also useful for research purposes and genome editing of plants.
[0276] B. CRISPR Expression System Also provided herein are nucleic acids comprising regulatory elements operably linked to a nucleotide sequence encoding a guide nucleic acid disclosed herein. In certain embodiments, the nucleic acid comprises a regulatory element operably linked to a nucleotide sequence encoding a single guide nucleic acid; this nucleic acid alone may constitute a CRISPR expression system. In certain embodiments, the nucleic acid comprises a regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid. In certain embodiments, the nucleic acid further comprises a nucleotide sequence encoding a modulator nucleic acid, wherein the nucleotide sequence encoding the modulator nucleic acid is operably linked to the same or a different regulatory element as the nucleotide sequence encoding the targeter nucleic acid; this nucleic acid alone may constitute a CRISPR expression system.
[0277] Further provided by the present invention is a CRISPR expression system comprising (a) a nucleic acid comprising a first regulatory element operably linked to a nucleotide sequence encoding a targeter nucleic acid and (b) a nucleic acid comprising a second regulatory element operably linked to a nucleotide sequence encoding a modulator nucleic acid.
[0278] In certain embodiments, the CRISPR expression system further comprises a nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein, such as a Cas protein disclosed herein. In certain embodiments, the Cas protein is a type V-A, V-C, or V-D Cas protein (e.g., a Cas nuclease). In certain embodiments, the Cas protein is a type V-A Cas protein (e.g., a Cas nuclease).
[0279] As used in this context, the term "operably linked" means that the nucleotide sequence of interest is linked to the regulatory element in a manner that enables expression of the nucleotide sequence (e.g., when the vector is introduced into a host cell, within an in vitro transcription / translation system or within a host cell).
[0280] The nucleic acids of the above-described CRISPR expression system can independently be selected from various nucleic acids, such as DNA (e.g., modified DNA) and RNA (e.g., modified RNA). In certain embodiments, the nucleic acid comprising a regulatory element operably linked to one or more nucleotide sequences encoding a guide nucleic acid is in the form of DNA. In certain embodiments, the nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein is in the form of DNA. The third regulatory element can be a constitutive or inducible promoter that promotes the expression of the Cas protein. In other embodiments, the nucleic acid comprising a third regulatory element operably linked to a nucleotide sequence encoding a Cas protein is in the form of RNA (e.g., mRNA).
[0281] The nucleic acids of the CRISPR expression system can be provided in one or more vectors. As used herein, the term "vector" can refer to a nucleic acid molecule capable of transporting another nucleic acid to which it is linked. Using conventional virus-based and non-virus-based gene delivery methods, nucleic acids can be introduced into cells, such as prokaryotic cells, eukaryotic cells, mammalian cells, or target tissues. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., the transcription products of the vectors described herein), naked nucleic acids, and nucleic acids complexed with delivery vehicles, such as liposomes. Virus vector delivery systems include DNA and RNA viruses that have either an episomal or integrated genome after delivery to the cell. Gene therapy procedures are known in the art and are disclosed in Van Brunt (1988) BIOTECHNOLOGY, 6:1149; Anderson (1992) SCIENCE, 256:808; Nabel & Feigner (1993) TIBTECH, 11:211; Mitani & Caskey (1993) TIBTECH, 11:162; Dillon (1993) TIBTECH, 11:167; Miller (1992) NATURE, 357:455; Vigne, (1995) RESTORATIVE NEUROLOGY AND NEUROSCIENCE, 8:35; Kremer & Perricaudet (1995) BRITISH MEDICAL BULLETIN, 51:31; Haddada et al. (1995) CURRENT TOPICS IN MICROBIOLOGY AND IMMUNOLOGY, 199:297; Yu et al. (1994) GENE THERAPY, 1:13; and Doerfler and Bohm (Eds.) (2012) The Molecular Repertoire of Adenoviruses II: Molecular Biology of Virus-Cell Interactions. In certain embodiments, at least one of the vectors is a DNA plasmid.In certain embodiments, at least one of the vectors is a viral vector (e.g., a retrovirus, an adenovirus, or an adeno-associated virus).
[0282] Certain vectors are capable of self-replicating within the host cells into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Other vectors (e.g., non-episomal mammalian vectors and replication-defective viral vectors) do not self-replicate within host cells. However, certain vectors may integrate into the genome of the host cell and thereby be replicated along with the host genome. Those skilled in the art will understand that various vectors may be suitable for various delivery methods, have various host tropisms, and that one or more vectors suitable for use can be selected.
[0283] As used herein, the term "regulatory element" can refer to transcriptional control sequences and / or translational control sequences, such as promoters, enhancers, transcriptional termination signals (e.g., polyadenylation signals), internal ribosome entry sites (IRES), proteolytic signals, etc., that effect and / or regulate the transcription of non-coding sequences (e.g., targeter nucleic acids or modulator nucleic acids) or coding sequences (e.g., Cas proteins) and / or regulate the translation of the encoded polypeptide. Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY, 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that direct constitutive expression of nucleotide sequences in many types of host cells and those that direct expression of nucleotide sequences only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired tissue of interest, such as muscle, neuron, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocyte). Regulatory elements can also direct expression in a time-dependent manner, which may or may not be tissue- or cell-type specific, e.g., in a cell cycle-dependent or developmental stage-dependent manner. In certain embodiments, the vector includes one or more pol III promoters (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5, or more pol I promoters), or combinations thereof. Examples of pol III promoters include, but are not limited to, the U6 and H1 promoters.Examples of pol II promoters include, but are not limited to, the Rous sarcoma virus (RSV) LTR promoter of retrovirus (optionally with RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. The term "regulatory element" includes enhancer elements such as the WPRE; the CMV enhancer; the R-U5' segment in the LTR of HTLV-I (see Takebe et al. (1988) MOL. CELL. BIOL., 8:466); the SV40 enhancer; and the intron sequence between exons 2 and 3 of rabbit β-globin (see O'Hare et al. (1981) PROC. NATL. ACAD. SCI. USA., 78:1527). It will be understood by those skilled in the art that the design of the expression vector may depend on factors such as the choice of host cell to be transformed and the desired level of expression. The vector is introduced into the host cell to produce a transcript, protein, or peptide (including fusion protein or peptide) encoded by the nucleic acid described herein, (e.g., a CRISPR transcript, protein, enzyme, its mutant form, or its fusion protein).
[0284] In certain embodiments, the nucleotide sequence encoding the Cas protein is codon-optimized for expression in a prokaryotic cell, such as E. coli, a eukaryotic host cell, such as a yeast cell (e.g., S. cerevisiae), a mammalian cell (e.g., a mouse cell, a rat cell, or a human cell), or a plant cell. Various species exhibit a particular bias for some codons of a particular amino acid. Codon bias (the difference in codon usage frequency among organisms) often correlates with the translation efficiency of messenger RNA (mRNA), which in turn is thought to depend particularly on the properties of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of the selected tRNA within a cell is generally a reflection of the codons most frequently used in peptide synthesis. Thus, genes can be adjusted for optimal gene expression in a given organism based on codon optimization. Tables of codon usage frequencies are readily available, for example, in the "Codon Usage Database" available at kazusa.or.jp / codon / , and these tables can be adapted in several ways (see Nakamura et al. (2000) NUCL. ACIDS RES., 28:292). Computer algorithms for codon-optimizing a particular sequence for expression in a particular host cell, such as Gene Forge (Aptagen; Jacobus, Pa.), are also available. In certain embodiments, codon optimization facilitates or improves the expression of the Cas protein within the host cell.
[0285] C. Donor template Cleavage of a target nucleotide sequence in the genome of a cell by a CRISPR-Cas system or complex can activate the DNA damage pathway, whereby the cleaved DNA fragment can be religated by NHEJ or HDR. HDR requires an endogenous or exogenous repair template for transferring sequence information from the repair template to the target.
[0286] In certain embodiments, the engineered non-natural system or CRISPR expression system further comprises a donor template. As used herein, the term "donor template" can refer to a nucleic acid designed to act as a repair template at or near a target nucleotide sequence upon introduction into a cell or organism. In certain embodiments, the donor template is complementary to a polynucleotide that includes the target nucleotide sequence or a portion thereof. When optimally aligned, the donor template can overlap with one or more nucleotides of the target nucleotide sequence (e.g., about 1, 5, 10, 15, 20, 25, 30, 35, 40 or more nucleotides, or more than about 1, 5, 10, 15, 20, 25, 30, 35, 40 nucleotides). The nucleotide sequence of the donor template is typically not identical to the genomic sequence it replaces. Instead, the donor template can include one or more substitutions, insertions, deletions, inversions or rearrangements relative to the genomic sequence, so long as there is sufficient homology to assist homologous recombination repair. In certain embodiments, the donor template includes a non-homologous sequence flanked by two homologous regions (i.e., homology arms), such that homologous recombination repair between the target DNA region and the two flanking sequences results in the insertion of the non-homologous sequence into the target region. In certain embodiments, the donor template includes a non-homologous sequence that is 10 to 100 nucleotides, 50 to 500 nucleotides, 100 to 1,000 nucleotides, 200 to 2,000 nucleotides, or 500 to 5,000 nucleotides in length located between the two homology arms.
[0287] Generally, the homologous region of the donor template has at least 50% sequence identity with the genomic sequence where recombination is desired. The homology arms are designed or selected such that they are capable of undergoing recombination with the nucleotide sequence adjacent to the target nucleotide sequence under intracellular conditions. In certain embodiments, when HDR of the non-target strand is desired, the donor template comprises a first homology arm homologous to the sequence 5' of the target nucleotide sequence and a second homology arm homologous to the sequence 3' of the target nucleotide sequence. In certain embodiments, the first homology arm is at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identical to the sequence 5' of the target nucleotide sequence. In certain embodiments, the second homology arm is at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identical to the sequence 3' of the target nucleotide sequence. In certain embodiments, when the polynucleotide comprising the donor template sequence and the target nucleotide sequence is optimally aligned, the closest nucleotide of the donor template is within about 1, 5, 10, 15, 20, 25, 50, 75, 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, or more nucleotides from the target nucleotide sequence.
[0288] In certain embodiments, the donor template further comprises an engineered sequence that is not homologous to the sequence to be repaired. Such an engineered sequence may have a barcode and / or a sequence capable of hybridizing to the donor template recruitment sequences disclosed herein.
[0289] In certain embodiments, the donor template further comprises one or more mutations to the genomic sequence, where the one or more mutations reduce or prevent cleavage of the donor template by the same CRISPR-Cas system or cleavage of a modified genomic sequence into which at least a portion of the donor template sequence has been incorporated. In certain embodiments, in the donor template, the PAM that is adjacent to the target nucleotide sequence and recognized by the Cas nuclease is mutated to a sequence not recognized by the same Cas nuclease. In certain embodiments, in the donor template, the target nucleotide sequence (e.g., the seed region) is mutated. In certain embodiments, the one or more mutations are silent with respect to the reading frame of the protein-coding sequence containing the mutation site.
[0290] The donor template can be provided to the cell as single-stranded DNA, single-stranded RNA, double-stranded DNA, or double-stranded RNA. It is understood that a CRISPR-Cas system, such as the systems disclosed herein, can have nuclease activity that cleaves the target strand, the non-target strand, or both. When HDR of the target strand is desired, a donor template having a nucleic acid sequence complementary to the target strand is also contemplated.
[0291] Donor templates can be introduced into cells in linear or circular form. When introduced in linear form, the ends of the donor template can be protected by methods known to those of skill in the art (e.g., from exonuclease digestion). For example, one or more dideoxynucleotide residues can be added to the 3’ of the linear molecule, and / or self-complementary oligonucleotides can be ligated to one or both ends (see, e.g., Chang et al. (1987) Proc. Natl. Acad. Sci. USA, 84:4959; Nehls et al. (1996) Science, 272:886; see also chemical modifications to increase the stability and / or specificity of RNA disclosed above). Additional methods for protecting exogenous polynucleotides from degradation include, but are not limited to, the addition of terminal amino groups and the use of modified internucleotide linkages, such as phosphorothioates, phosphoramidates, and O-methyl ribose or deoxyribose residues. Instead of protecting the ends of the linear donor template, additional lengths of sequence that can be degraded without affecting recombination can be included outside of the homologous regions.
[0292] Donor templates can be components of the vectors described herein, can be included in separate vectors, or can be provided as separate polynucleotides, such as oligonucleotides, linear polynucleotides, or synthetic polynucleotides. In certain embodiments, the donor template is DNA. In certain embodiments, the donor template, where applicable, is in the same nucleic acid as the sequence encoding the single guide nucleic acid, the sequence encoding the targeter nucleic acid, the sequence encoding the modulator nucleic acid, and / or the sequence encoding the Cas protein. In certain embodiments, the donor template is provided as a separate nucleic acid. The donor template polynucleotide can be of any suitable length, e.g., about or at least about 50, 75, 100, 150, 200, 500, 1000, 2000, 3000, 4000, or more nucleotides in length.
[0293] The donor template can be introduced into cells as an isolated nucleic acid. Alternatively, the donor template can be introduced into cells as part of a vector (e.g., a plasmid) that has additional sequences not intended to be inserted into the DNA region of interest, such as an origin of replication, a promoter, and genes encoding antibiotic resistance. Alternatively, the donor template can be delivered by a virus (e.g., an adenovirus, an adeno-associated virus (AAV)). In certain embodiments, the donor template is introduced as an AAV, e.g., a pseudotyped AAV. The capsid protein of the AAV can be selected by one of ordinary skill in the art based on the tropism of the AAV and the target cell type. For example, in certain embodiments, the donor template is introduced into hepatocytes as AAV8 or AAV9. In certain embodiments, the donor template is introduced into hematopoietic stem cells, hematopoietic progenitor cells, or T lymphocytes (e.g., CD8+ T lymphocytes) as AAV6 or AAVHSC (see U.S. Patent No. 9,890,396). It is understood that the sequence of the capsid protein (VP1, VP2, or VP3) can be modified from a wild-type AAV capsid protein to have at least 50% (e.g., at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) sequence identity to the wild-type AAV capsid sequence.
[0294] Donor templates can be delivered to cells (e.g., primary cells) by a variety of delivery methods, such as viral or non-viral methods disclosed herein. In certain embodiments, non-viral donor templates are introduced into target cells as naked nucleic acids or as complexes with liposomes or poloxamers. In certain embodiments, non-viral donor templates are introduced into target cells by electroporation. In other embodiments, viral donor templates are introduced into target cells by infection. The engineered non-natural system can be delivered before, after, or simultaneously with the donor template (see International (PCT) Application Publication No. WO 2017 / 053729 pamphlet). One of ordinary skill in the art will be able to select an appropriate timing based on the form of delivery (e.g., considering the time required for transcription and translation of RNA and protein components) and the half-life of the molecule within the cell. In certain embodiments, when a CRISPR-Cas system comprising a Cas protein is delivered by electroporation (e.g., as an RNP), the donor template (e.g., as an AAV) is introduced into the cell within 4 hours (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 90, 120, 150, 180, 210, or 240 minutes) after introduction of the engineered non-natural system.
[0295] In certain embodiments, the donor template is covalently conjugated to the modulator nucleic acid. Covalent bonds suitable for this conjugation are known in the art and are described, for example, in U.S. Patent No. 9,982,278 and Savic et al. (2018) ELIFE 7:e33761. In certain embodiments, the donor template is covalently linked to the modulator nucleic acid (e.g., the 5' end of the modulator nucleic acid) via an internucleotide bond. In certain embodiments, the donor template is covalently linked to the modulator nucleic acid (e.g., the 5' end of the modulator nucleic acid) via a linker.
[0296] In certain embodiments, the donor template can comprise any nucleic acid chemistry. In certain embodiments, the donor template can comprise DNA and / or RNA nucleotides. In certain embodiments, the donor template can comprise single-stranded DNA, linear single-stranded RNA, linear double-stranded DNA, linear double-stranded RNA, circular single-stranded DNA, circular single-stranded RNA, circular double-stranded DNA, or circular double-stranded RNA. In certain embodiments, the donor template comprises a mutation in the PAM sequence to partially or completely abolish the binding of the RNP to DNA. In certain embodiments, the donor template is present at a concentration of at least 0.05, 0.01, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.25, 1.5, 1.75, 2, 3, or 4, and / or 0.01, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.25, 1.5, 1.75, 2, 3, 4, or 5 μg μL -1 Hereinafter, for example, 0.01 to 5 μg μL -1 and is present at a concentration of. In certain embodiments, the donor template comprises one or more promoters. In certain embodiments, the donor template comprises a promoter that is at least 70, 75, 80, 85, 90, 95, 99.5, or 100% identical to any one of SEQ ID NOs: 78-85 in Table 6.
[0297]
Table 103
[0298]
Table 104
[0299]
Table 105
[0300]
Table 106
[0301]
Table 107
[0302] D. Efficiency and Specificity The engineered non-natural system can be evaluated for efficiency and / or specificity in nucleic acid targeting, cleavage, or modification.
[0303] In certain embodiments, the engineered non-natural system has high efficiency. For example, in certain embodiments, at least 1, 1.5, 2, 2.5, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% of a population of nucleic acids having a target nucleotide sequence and a cognate PAM is targeted, cleaved, or modified when contacted with the engineered non-natural system. In certain embodiments, at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 100% of the genome of a population of cells is targeted, cleaved, or modified when the engineered non-natural system is delivered into the cells.
[0304] For a given spacer array, it is observed that the occurrence of on-target events and the occurrence of off-target events generally correlate. For certain therapeutic purposes, a lower on-target efficiency may be acceptable, and a lower off-target frequency is more desirable. For example, when editing or modifying proliferating cells that are delivered to a subject and grow in vivo, the tolerance to off-target events is low. Prior to delivery, it is possible to evaluate on-target and off-target events and thereby select one or more colonies that have the desired edit or modification and no unwanted edits or modifications. Nevertheless, the on-target efficiency may need to meet certain criteria to be suitable for therapeutic use. The high editing efficiency in a standard CRISPR-Cas system allows for adjustment of the system, for example, by reducing the binding of guide nucleic acid to the Cas protein, without losing therapeutic applicability.
[0305] In certain embodiments, when a population of nucleic acids having a target nucleotide sequence and a cognate PAM is contacted with the engineered non-natural systems disclosed herein, the frequency of off-target events (e.g., targeting, cleavage, or modification depending on the function of the CRISPR-Cas system) is reduced. Methods for assessing off-target events are summarized in Lazzarotto et al. (2018) Nat Protoc. 13(11):2615-42 and include discovery and sequencing of in situ Cas off-targets (DISCOVER-seq) as disclosed in Wienert et al. (2019) Science 364(6437):286-89; genome-wide unbiased identification of double-strand breaks (DSBs) enabled by sequencing (GUIDE-seq) as disclosed in Kleinstiver et al. (2016) Nat. Biotech. 34:869-74; and cyclization for in vitro reporting of cleavage effects by sequencing (CIRCLE-seq) as described in Kocak et al. (2019) Nat. Biotech. 37:657-66. In certain embodiments, an off-target event includes targeting, cleavage, or modification at a given off-target locus (e.g., the locus at which the most occurrences of off-target events are detected). In certain embodiments, off-target events collectively include targeting, cleavage, or modification at all loci having detectable off-target events.
[0306] In certain embodiments, the genomic mutations are detected at 0.0001%, 0.0002%, 0.0003%, 0.0004%, 0.0005%, 0.0006%, 0.0007%, 0.0008%, 0.0009%, 0.001%, 0.002%, 0.003%, 0.004%, 0.005%, 0.006%, 0.007%, 0.008%, 0.009%, 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, or 5% or less (in total) in cells at any off-target locus. In certain embodiments, the ratio of the percentage of cells having an on-target event to the percentage of cells having any off-target event (e.g., the ratio of the percentage of cells having an on-target editing event to the percentage of cells having a mutation at any off-target locus) is at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000. It is understood that genetic mutations may be present within a population of cells, e.g., by spontaneous mutations, and such mutations are not included as off-target events.
[0307] E. Multiplex method The methods for targeting, editing, and / or modifying genomic DNA disclosed herein can be carried out in a multiplexed manner. For example, a library of targeter nucleic acids can be used to target multiple genomic loci; a library of donor templates can also be used to generate multiple insertions, deletions, and / or substitutions. Multiplex assays can be performed in screening methods, where each individual cell culture (e.g., within the wells of a 96-well plate or a 384-well plate) is exposed to different guide nucleic acids having different targeter stem sequences and / or different donor templates. Multiplex assays can also be performed in selection methods, where a cell culture is exposed to a mixed population of different guide nucleic acids and / or donor templates, and cells having a desired characteristic (e.g., functionality) are enriched or selected by advantageous survival or proliferation, resistance to a particular agent, expression of a detectable protein (e.g., a fluorescent protein detectable by flow cytometry), etc.
[0308] In certain embodiments, a plurality of guide nucleic acids and / or a plurality of donor templates are designed for saturation editing. For example, in certain embodiments, each nucleotide position in a sequence of interest is systematically modified with each of the conventional four bases, A, T, G, and C. In other embodiments, at least one sequence in each gene from a pool of genes of interest is modified, e.g., according to a CRISPR design algorithm. In certain embodiments, each sequence from a pool of exogenous elements of interest (e.g., protein-coding sequences, non-protein-coding genes, regulatory elements) is inserted into one or more given loci of the genome.
[0309] It is understood that multiplex methods suitable for performing screening or selection methods typically performed for research purposes may differ from those suitable for therapeutic purposes. For example, constitutive expression of certain elements (e.g., Cas nuclease and / or guide nucleic acid) may be undesirable for therapeutic purposes due to the potential for increased off-target effects. Conversely, for research purposes, constitutive expression of Cas nuclease and / or guide nucleic acid may be desirable. For example, constitutive expression provides a broad time frame in which other elements can be introduced. When stable cell lines are established for constitutive expression, the number of exogenous elements that need to be co-delivered into a single cell is also reduced. Thus, constitutive expression of certain elements can enhance the efficiency of the screening or selection process and reduce complexity. Inducible expression of certain elements of the systems disclosed herein can also be used for research purposes, considering similar advantages. Expression can be induced by exogenous factors (e.g., small molecules) or by endogenous molecules or complexes present in specific cell types (e.g., of a specific differentiation stage). Methods known in the art, such as those described herein, can be used to constitutively or inducibly express one or more elements. For example, the specificity of a CRISPR nuclease is at least partially determined by the uniqueness of the spacer (in combination with the proximity of the spacer sequence to the required PAM), and its off-target score can be calculated using an algorithm, e.g., crispr.mit.edu (Hsu et al. (2013) Nat. Biotech. 31:827-832). The highest possible score is 100, which indicates high specificity and low off-target potential. Since our SHS library targets intergenic regions, the algorithm for gRNA prediction should be able to perform alignments using repetitive regions and low-complexity sequences.
[0310] Regardless of the need to introduce multiple elements - a single guide nucleic acid and a Cas protein; or a targeter nucleic acid, a modulator nucleic acid, and a Cas protein - it is further understood that these elements can be delivered intracellularly as a single complex of pre-formed RNPs. Thus, the efficiency of the screening or selection process can also be achieved by pre-assembling multiple RNP complexes in a multiplexed manner.
[0311] In certain embodiments, the methods disclosed herein further include identifying, by a screening or selection process, a guide nucleic acid, a Cas protein, a donor template, or a combination of two or more of these elements. A set of barcodes can be used, for example, in the donor template between two homology arms to facilitate identification. In certain embodiments, the methods include obtaining a population of cells; selectively amplifying genomic DNA or RNA samples and / or barcodes that contain a target nucleotide sequence; and / or sequencing the selectively amplified genomic DNA or RNA samples and / or barcodes.
[0312] Furthermore, the invention provides libraries that include multiple guide nucleic acids, for example, libraries that include the multiple guide nucleic acids disclosed herein. In another aspect, the invention provides libraries that include multiple nucleic acids each comprising a regulatory element operably linked to a different guide nucleic acid, for example, the different guide nucleic acids disclosed herein. These libraries can be used in combination with one or more Cas proteins or Cas-encoding nucleic acids, for example, those disclosed herein, and / or one or more donor templates, for example, those disclosed herein, for screening or selection methods.
[0313] F. Genomic Safe Harbor Genome engineering is a research field that aims to modify the genes of organisms in order to improve the understanding of gene function, and in particular, to develop methods for genome engineering to treat hereditary or acquired diseases. To modify the genome of target cells, those skilled in the art use one or more available means to introduce changes into the genome at the targeted location to modify the sequence of a target polynucleotide, such as a target gene, in a desired manner, for example, to regulate gene expression, to regulate gene sequence, to remove gene sequence, to introduce a gene, such as exogenous DNA, such as a transgene, etc. Efficient transgene insertion can be achieved by non-precise methods including, but not limited to, viral vectors, such as retroviral vectors, such as adeno-associated virus (AAV), etc., or by precise methods including, but not limited to, inducible nucleases, such as zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), homing endonucleases, such as restriction endonucleases, or nucleic acid-induced nucleases, such as CRISPR-cas, such as Cas9 and Cas12a and their engineered forms.
[0314] Exogenous genes, such as transgenes, inserted into the genome of target human cells can interact with other genomic elements in an unpredictable manner, randomly, for example, via a retroviral vector, or in a tandem manner, for example, by the action of a nucleic acid-induced nuclease, such as Cas. Attempts to modify the default genome structure by the integration of exogenous DNA, such as a transgene, or a synthetic sequence, for the complex transcriptional regulation of genes in mammalian cells by cis and trans regulatory elements, such as proximal and distal enhancers, and networks of multiple transcription factors, can affect the expression of the transgene itself, which can, for example, compromise safety checkpoints present in healthy cells and dramatically change cell behavior, i.e., promote clonal expansion or malignant transformation of the host, including complete attenuation or complete silencing of both proximal and distal endogenous genes, and / or expression resulting in downregulation of the expression of key genes, such as oncogenes and tumor suppressor genes.
[0315] Integration of genes adjacent to regulatory elements of oncogenes has been shown to cause oncogenic transformation, which is particularly important when manipulating cells for therapeutic use. Accordingly, there is a desire to identify suitable target polynucleotides comprising target nucleotide sequences in the human genome that result in suitable expression of the transgene without disruption of adjacent genes upon insertion of the transgene. In particular, for gene and cell therapy applications, there is a desire for suitable target polynucleotides comprising target nucleotide sequences in the human genome that result in sufficient expression of the transgene in therapeutic cells, such as T cells, such as CAR T cells; or progenitor cells, such as stem cells, such as hematopoietic stem cells, without any other disruption that would be harmful to the malignant transformation or the individual after transplantation.
[0316] Expression of an exogenous gene, such as a transgene, in a desired cell type and / or stage of development / differentiation depends on integration into a suitable target polynucleotide comprising a target nucleotide sequence that provides sufficient expression to a sufficient extent for the intended purpose from a candidate locus. Expression from a particular genomic locus can be affected by many factors, including cell type and stage of differentiation, as well as changes in chromatin structure, although not limited to, when one or more components of the target polynucleotide are activated during differentiation while other components are silenced. Accordingly, there is a desire to identify suitable target polynucleotides comprising target nucleotide sequences in the human genome such that insertion of exogenous DNA, such as a transgene, results in sufficient expression in target human cells and, in the case of stem cells, the expression is maintained at sufficient levels (1) by differentiation and (2) by clonal expansion. The present disclosure provides a significant advance in the ability to manipulate the human genome by providing compositions and methods for targeting and neutralizing an exogenous gene, such as a transgene, to a suitable target polynucleotide comprising a target nucleotide sequence.
[0317] Compositions and methods for genome manipulation are provided herein. Certain embodiments include compositions. Certain embodiments include compositions for editing the genome. Embodiments disclosed herein relate to novel guide nucleic acids (gNAs), such as gRNAs, that are complementary to a target nucleotide sequence in a target polynucleotide. As used herein, a "target polynucleotide" includes a polynucleotide in which a target nucleotide sequence is located. As used herein, a "target nucleotide sequence" includes a sequence to which a guide sequence can bind, e.g., having complementarity, wherein binding between the target nucleotide sequence and the guide sequence can enable the activity of a nucleic acid-guided nuclease complex. Further embodiments disclosed herein relate to novel gNAs, such as gRNAs, that are complementary to a target nucleotide sequence in a target polynucleotide, wherein the insertion of exogenous DNA, such as a transgene, does not adversely affect the cell, e.g., does not significantly affect the expression of one or more endogenous genes or result in malignant transformation of the cell. In further embodiments disclosed herein, gene expression shown in human target cells is maintained at a level sufficient for the ultimate use of the cells, by differentiation of the human target cells and / or by proliferation of one or more progeny cells. Certain embodiments disclosed herein relate to novel nucleic acid-guided nuclease complexes, such as RNPs, such as Cas bound to a gNA, which is complementary to a target nucleotide sequence within a target polynucleotide and hybridizes (also referred to as cleaves or cuts) the phosphodiester backbone at at least one position in at least one strand of the target polynucleotide. Certain embodiments disclosed herein relate to methods for selecting and using a gNA, such as a gRNA, for genome manipulation. Certain embodiments relate to methods for using a gNA that is complementary to a target nucleotide sequence within a target polynucleotide, synthesizing the gNA and a nucleic acid-guided nuclease, and / or combining the nucleic acid-guided nuclease with the gNA to form a nucleic acid-guided nuclease complex, such as an RNP. Certain embodiments disclosed herein relate to methods. Certain embodiments disclosed herein relate to methods for manipulating the genome.Certain embodiments disclosed herein relate to a method in which a nucleic acid-guided nuclease complex, e.g., an RNP, is introduced, e.g., transfected, into a human target cell together with a donor template, e.g., exogenous DNA, e.g., a transgene, where the nucleic acid-guided nuclease cleaves the backbone at at least one position in at least one strand of a target polynucleotide and uses the donor template to repair the cleaved target polynucleotide and introduce at least a portion of the donor template into the target polynucleotide. As used herein, "exogenous DNA" or "transgene" includes any natural or synthetic gene that is introduced into the genome of an organism or cell that is not endogenous thereto. The transgene may or may not retain the ability to be expressed in the human target cell and / or produce RNA or protein. The transgene may or may not modify the resulting phenotype of the human target cell. Certain embodiments are human target cells, e.g., eukaryotic cells, e.g., mammalian cells, e.g., human cells, e.g., stem cells or immune cells, produced by a method in which a nucleic acid-guided nuclease complex, e.g., an RNP, is introduced, e.g., transfected, into a human target cell together with a donor template, e.g., exogenous DNA or a transgene, e.g., a chimeric antigen receptor (CAR), where the nucleic acid-guided nuclease is cleaved at or near a target sequence in the target polynucleotide and uses the donor template to introduce at least a portion of the donor template into the target polynucleotide to repair the cleaved target polynucleotide. Certain embodiments disclosed herein include a promoter sequence adjacent to an exogenous gene, e.g., a transgene; in some cases, a construct containing the promoter, when introduced into the target polynucleotide of a human target cell, e.g., an immune cell or a stem cell, maintains sufficient gene expression in the edited human target cell for the intended purpose of the cell or its progeny. In certain embodiments, the human target cell is viable after introduction of the exogenous DNA.
[0318] As used herein, "human target cell" includes cells into which an exogenous product, such as a protein, nucleic acid, or combination thereof, has been introduced. In certain cases, the human target cell can be used to produce a gene product from exogenous DNA, such as a transgene, such as an exogenous protein, such as a CAR. In certain cases, the human target cell may contain a target nucleotide sequence within a target polynucleotide, where a nucleic acid-guided nuclease hybridizes and cleaves at a cleavage site at one or more positions in one or more strands of the target polynucleotide at or near the target nucleotide sequence.
[0319] As used herein, "cleavage site" includes one or more positions at which a nucleic acid-guided nuclease complex hydrolyzes the phosphodiester backbone of a single-stranded or double-stranded target polynucleotide after binding to a target nucleotide sequence in the target polynucleotide. In certain cases where the target polynucleotide of the nucleic acid-guided nuclease complex is double-stranded, binding of the nucleic acid-guided nuclease complex to the target nucleotide sequence in the target polynucleotide can result in hydrolysis of one strand of the target polynucleotide at or near the target nucleotide sequence, resulting in strand breakage. In such cases, the nucleic acid-guided nuclease complex can cleave either strand of the target polynucleotide. In certain cases, binding of the nucleic acid-guided nuclease complex to the target nucleotide sequence in the target polynucleotide can result in hydrolysis of both strands of the target polynucleotide at or near the target nucleotide sequence, resulting in cleavage of both strands. The cleavage sites are the same for both strands and can result in blunt ends, or the cleavage sites for each strand can offset to result in a single-stranded overhang, such as a sticky end. In certain cases, mismatches at or near the cleavage site may or may not affect the cleavage efficiency of the nucleic acid-guided nuclease complex.
[0320] In some cases, uncontrolled gene integration adjacent to regulatory elements of proto-oncogenes has been shown to cause oncogenic transformation, which is particularly important when manipulating cells for therapeutic purposes. Accordingly, it is desirable to identify suitable target polynucleotides comprising target nucleotide sequences that provide safe and stable integration of exogenous DNA with sufficient expression in human target cells and their resulting progeny.
[0321] Exemplary characteristics of target nucleotide sequences that can exhibit predictable functions without potentially harmful changes in human target cell genome activity are: (1) >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb; (2) >150 kb away from any miRNA / other functional small RNA, e.g., >200, e.g., >250, and in some cases, >300 kb; (3) >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb; (4) >10 kb away from any origin of replication, e.g., >20, e.g., >30, and in some cases, >50 kb; (5) >10 kb away from any ultra-conserved element, e.g., >20, e.g., >30, and in some cases, >50 kb; (6) exhibiting low transcriptional activity; (7) outside of copy number variable regions; (8) located in open chromatin; and (9) being unique, i.e., comprising one or more out of one copy per genome.
[0322] In certain embodiments, compositions are provided herein. In certain embodiments, compositions for manipulating human target cells at suitable target nucleotide sequences within the target polynucleotides of human target cells are provided herein.
[0323] In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least one of the exemplary features. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least two of the exemplary features. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least three of the exemplary features. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least four of the exemplary features. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least five of the exemplary features. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least six of the exemplary features. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least seven of the exemplary features. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has at least eight of the exemplary features. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence has all of the exemplary features.
[0324] In certain embodiments, suitable target polynucleotides are >10 kb, e.g., >20, e.g., >30, and in some cases, >50 kb away from any 5' gene end. In certain embodiments, suitable target polynucleotides are >10 kb, e.g., >20, e.g., >30, and in some cases, >50 kb away from any 5' gene end and further include at least one additional exemplary feature. In certain embodiments, suitable target polynucleotides are >10 kb, e.g., >20, e.g., >30, and in some cases, >50 kb away from any 5' gene end and further include at least two additional exemplary features. In certain embodiments, suitable target polynucleotides are >10 kb, e.g., >20, e.g., >30, and in some cases, >50 kb away from any 5' gene end and further include at least three additional exemplary features. In certain embodiments, suitable target polynucleotides are >10 kb, e.g., >20, e.g., >30, and in some cases, >50 kb away from any 5' gene end and further include at least four additional exemplary features. In certain embodiments, suitable target polynucleotides are >10 kb, e.g., >20, e.g., >30, and in some cases, >50 kb away from any 5' gene end and further include at least five additional exemplary features. In certain embodiments, suitable target polynucleotides are >10 kb, e.g., >20, e.g., >30, and in some cases, >50 kb away from any 5' gene end and further include at least six additional exemplary features. In certain embodiments, suitable target polynucleotides are >10 kb, e.g., >20, e.g., >30, and in some cases, >50 kb away from any 5' gene end and further include at least seven additional exemplary features. In certain embodiments, suitable target polynucleotides are >10 kb, e.g., >20, e.g., >30, and in some cases, >50 kb away from any 5' gene end and further include all eight additional exemplary features.
[0325] In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and further comprise at least one additional exemplary feature. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and further comprise at least two additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and further comprise at least three additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and further comprise at least four additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and further comprise at least five additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and further comprise at least six additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and further comprise at least seven additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and further comprise all eight additional exemplary features.
[0326] In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb away. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb away, and further comprise at least one additional exemplary feature. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb away, and further comprise at least two additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb away, and further comprise at least three additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb away, and further comprise at least four additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, and >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb away, and further comprise at least five additional exemplary features.In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb away, and further include at least six additional exemplary features. In certain embodiments, suitable target polynucleotides are >150 kb away from known cancer-related genes, e.g., >200, e.g., >250, and in some cases, >300 kb away, >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb away, and further include all seven additional exemplary features.
[0327] In preferred embodiments, suitable target polynucleotides are >10 kb away from any 5' gene end, e.g., >20, e.g., >30, and in some cases, >50 kb away, and >150, e.g., >200, e.g., >250, and in some cases, >300 kb away from known cancer-related genes.
[0328] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise any one of SEQ ID NOs: 2020 to 2043 in Table 7. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or identical to any one of SEQ ID NOs: 2020 to 2043. In a preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 98% identical to any one of SEQ ID NOs: 2020 to 2043. In a more preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 99% identical to any one of SEQ ID NOs: 2020 to 2043.
[0329] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise any one of SEQ ID NOs: 2020 to 2042 in Table 7. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or identical to any one of SEQ ID NOs: 2020 to 2042. In a preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 98% identical to any one of SEQ ID NOs: 2020 to 2042. In a more preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 99% identical to any one of SEQ ID NOs: 2020 to 2042.
[0330] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise any one of SEQ ID NOs: 2020-2041 and 2043 in Table 7. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or identical to any one of SEQ ID NOs: 2020-2041 and 2043. In a preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 98% identical to any one of SEQ ID NOs: 2020-2041 and 2043. In a more preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 99% identical to any one of SEQ ID NOs: 2020-2041 and 2043.
[0331] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise any one of SEQ ID NOs: 2020 to 2041 in Table 7. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or identical to any one of SEQ ID NOs: 2020 to 2041. In a preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 98% identical to any one of SEQ ID NOs: 2020 to 2041. In a more preferred embodiment, a suitable target polynucleotide comprising a target nucleotide sequence is at least 99% identical to any one of SEQ ID NOs: 2020 to 2041.
[0332] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise at least a portion of any one of SEQ ID NOs: 2020 to 2030 in Table 7, such as nucleotides 1 to 495, 1 to 490, 1 to 485, 1 to 480, 1 to 475, 1 to 470, 1 to 465, 1 to 460, 1 to 455, 1 to 450, 1 to 445, 1 to 440, 1 to 435, 1 to 430, 1 to 425, 1 to 420, 1 to 415, 1 to 410, 1 to 405, or 1 to 400. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or identical to a portion of any one of SEQ ID NOs: 2020 to 2030.
[0333] In certain embodiments, for example, for transgene insertion, a suitable target polynucleotide comprising a target nucleotide sequence may comprise at least a portion of any one of SEQ ID NOs: 2031 - 2041 in Table 7, such as nucleotides 5 - 500, 10 - 500, 15 - 500, 20 - 500, 25 - 500, 30 - 500, 35 - 500, 40 - 500, 45 - 500, 50 - 500, 55 - 500, 60 - 500, 65 - 500, 70 - 500, 75 - 500, 80 - 500, 85 - 500, 90 - 500, 95 - 500, or 100 - 500. In certain embodiments, a suitable target polynucleotide comprising a target nucleotide sequence is at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or identical to a portion of any one of SEQ ID NOs: 2031 - 2041.
[0334]
Table 108
[0335]
Table 109
[0336]
Table 110
[0337]
Table 111
[0338]
Table 112
[0339]
Table 113
[0340]
Table 114
[0341]
Table 115
[0342]
Table 116
[0343]
Table 117
[0344]
Table 118
[0345]
Table 119
[0346]
Table 120
[0347]
Table 121
[0348]
Table 122
[0349] In some cases, the expression of exogenous DNA inserted in or near the target nucleotide sequence or in the target polynucleotide, e.g., a transgene, may or may not correlate with the rearrangement of chromatin structure rearrangement during differentiation, while one or more components of the target polynucleotide are activated during differentiation and other components are silenced, which may depend on the cell type and differentiation stage. To overcome this, in certain embodiments, in addition to the exemplary features described above, a suitable target polynucleotide comprising the target nucleotide sequence demonstrates suitable expression of the inserted exogenous DNA, e.g., a transgene, through differentiation and clonal expansion.
[0350] IV. Pharmaceutical Compositions Compositions (e.g., pharmaceutical compositions) are provided herein that include a guide nucleic acid, an engineered non-natural system, or a eukaryotic cell, e.g., a guide nucleic acid, an engineered non-natural system, or a eukaryotic cell as disclosed herein. In certain embodiments, the composition includes an RNP that includes a guide nucleic acid, e.g., a guide nucleic acid as disclosed herein, and a Cas protein (e.g., a Cas nuclease). In certain embodiments, the composition includes a single guide nucleic acid, e.g., a single guide nucleic acid as disclosed herein. In certain embodiments, the composition includes an RNP that includes a single guide nucleic acid and a Cas protein (e.g., a Cas nuclease). In certain embodiments, the composition includes an RNP that includes a targeter nucleic acid, a modulator nucleic acid, and a Cas protein (e.g., a Cas nuclease). In certain embodiments, the composition includes a complex of a targeter nucleic acid and a modulator nucleic acid, e.g., a complex of a targeter nucleic acid and a modulator nucleic acid as disclosed herein. In certain embodiments, the composition includes an RNP that includes a targeter nucleic acid, a modulator nucleic acid, and a Cas protein (e.g., a Cas nuclease).
[0351] In certain embodiments, provided herein is a method of generating a composition, comprising incubating a single guide nucleic acid, such as a single guide nucleic acid disclosed herein, with a Cas protein, thereby generating a complex of the single guide nucleic acid and the Cas protein (e.g., an RNP). In certain embodiments, the method further comprises purifying the complex (e.g., the RNP).
[0352] In certain embodiments, provided herein is a method of generating a composition, comprising incubating a targeter nucleic acid and a modulator nucleic acid, such as a targeter nucleic acid and a modulator nucleic acid disclosed herein, under suitable conditions, thereby generating a composition (e.g., a pharmaceutical composition) comprising a complex of the targeter nucleic acid and the modulator nucleic acid. In certain embodiments, the method further comprises a method comprising incubating the targeter nucleic acid and the modulator nucleic acid with a Cas protein (e.g., a Cas nuclease or a related Cas protein capable of being activated by the targeter nucleic acid and the modulator nucleic acid), thereby generating a complex of the targeter nucleic acid, the modulator nucleic acid, and the Cas protein (e.g., an RNP). In certain embodiments, the method further comprises purifying the complex (e.g., the RNP).
[0353] For therapeutic use, a guide nucleic acid, an engineered non-natural system, a CRISPR expression system, or a cell comprising or modified by such a system as disclosed herein can be combined with a pharmaceutically acceptable carrier. As used herein, the term “pharmaceutically acceptable” can refer to compounds, materials, compositions, and / or dosage forms that are suitable for use in contact with human and animal tissues within the scope of sound medical judgment, commensurate with a reasonable benefit / risk ratio, and without undue toxicity, irritation, allergic response, or other problems or complications.
[0354] As used herein, the term "pharmaceutically acceptable carrier" includes buffers, carriers, and excipients suitable for use in contact with human and animal tissues without undue toxicity, irritation, allergic response, or other problems or complications commensurate with a reasonable benefit / risk ratio. Pharmaceutically acceptable carriers include any of the standard pharmaceutical carriers, such as phosphate buffered saline, water, emulsions (such as oil-in-water or water-in-oil emulsions, etc.), and various types of wetting agents. The compositions may also include stabilizers and preservatives. For examples of carriers, stabilizers, and adjuvants, see, e.g., Martin, Remington’s Pharmaceutical Sciences, 15th Ed., Mack Publ. Co., Easton, PA (1975). Pharmaceutically acceptable carriers include buffers, solvents, dispersion media, coatings, isotonic agents, and absorption delaying agents, etc., which are compatible with the administration of pharmaceuticals. The use of such media and agents for pharmaceutically active substances is known in the art.
[0355] In certain embodiments, the pharmaceutical compositions disclosed herein include salts such as NaCl, MgCl2, KCl, MgSO4, etc.; buffers such as Tris buffer, N-(2-hydroxyethyl)piperazine-N'-(2-ethanesulfonic acid) (HEPES), 2-(N-morpholino)ethanesulfonic acid (MES), MES sodium salt, 3-(N-morpholino)propanesulfonic acid (MOPS), N-tris[hydroxymethyl]methyl-3-aminopropanesulfonic acid (TAPS), etc.; solubilizing agents; detergents such as nonionic detergents such as Tween-20, etc.; nuclease inhibitors, etc. For example, in certain embodiments, the compositions of the invention include a buffer for stabilizing the DNA-targeting RNA of the invention, such as gRNA, and nucleic acids.
[0356] In certain embodiments, the pharmaceutical composition can contain formulation materials for adjusting, maintaining, or preserving, for example, the pH, osmolality, viscosity, clarity, color, isotonicity, odor, sterility, stability, rate of dissolution or release, adsorption, or permeability of the composition. In such embodiments, suitable formulation materials include, but are not limited to, amino acids (such as glycine, glutamine, asparagine, arginine, or lysine); antibacterial agents; antioxidants (such as ascorbic acid, sodium sulfite, or sodium bisulfite); buffers (such as borate, bicarbonate, Tris-HCl, citrate, phosphate, or other organic acids); bulking agents (such as mannitol or glycine); chelating agents (such as ethylenediaminetetraacetic acid (EDTA)); complexing agents (such as caffeine, polyvinylpyrrolidone, β-cyclodextrin, or hydroxypropyl-β-cyclodextrin); fillers; monosaccharides; disaccharides; and other carbohydrates (such as glucose, mannose, or dextrin); proteins (such as serum albumin, gelatin, or immunoglobulins); coloring agents, flavoring agents, and diluents; emulsifying agents; hydrophilic polymers (such as polyvinylpyrrolidone); low molecular weight polypeptides; salt-forming counterions (such as sodium); preservatives (such as benzalkonium chloride, benzoic acid, salicylic acid, thimerosal, phenethyl alcohol, methylparaben, propylparaben, chlorhexidine, sorbic acid, or hydrogen peroxide); solvents (such as glycerin, propylene glycol, or polyethylene glycol); sugar alcohols (such as mannitol or sorbitol); suspending agents; surfactants or wetting agents (such as pluronics, PEG, sorbitan esters, polysorbates, for example, polysorbate 20, polysorbate, triton, tromethamine, lecithin, cholesterol, tyloxapal, etc.); stability enhancers (such as sucrose or sorbitol); tonicity enhancers (such as alkali metal halides, preferably sodium or potassium chloride, mannitol, sorbitol, etc.); delivery vehicles; diluents; excipients; and / or pharmaceutical adjuvants (see Remington’s Pharmaceutical Sciences, 18th ed. (Mack Publishing Company, 1990)).
[0357] In certain embodiments, the pharmaceutical composition can contain nanoparticles, such as polymeric nanoparticles, liposomes, or micelles (see Anselmo et al. (2016) Bioeng. Transl. Med. 1:10 - 29). In certain embodiments, the pharmaceutical composition includes inorganic nanoparticles. Exemplary inorganic nanoparticles include, for example, magnetic nanoparticles (e.g., Fe3MnO2) or silica. The outer surface of the nanoparticles can be conjugated with a positively charged polymer (e.g., polyethyleneimine, polylysine, polyarginine) that enables binding (e.g., conjugation or encapsulation) of the payload. In certain embodiments, the pharmaceutical composition includes organic nanoparticles (e.g., encapsulation of the payload inside the nanoparticles). Exemplary organic nanoparticles include, for example, SNALP liposomes coated with polyethylene glycol (PEG) and containing a cationic lipid together with a neutral helper lipid, as well as protamine and nucleic acid complexes coated with a lipid coating. In certain embodiments, the pharmaceutical composition includes liposomes, such as those disclosed in International (PCT) Application Publication No. WO 2015 / 148863 pamphlet.
[0358] In certain embodiments, the pharmaceutical composition includes a targeting moiety for enhancing binding to target cells or turnover of the nanoparticles and liposomes. Exemplary targeting moieties include cell - specific antigens, monoclonal antibodies, single - chain antibodies, aptamers, polymers, saccharides, and cell - membrane - permeable peptides. In certain embodiments, the pharmaceutical composition includes a fusogenic or endosome - destabilizing peptide or polymer.
[0359] In certain embodiments, the pharmaceutical composition may contain a sustained release or controlled release formulation. Techniques for formulating sustained release or controlled release means, such as liposomal carriers, biodegradable microparticles or porous beads, and depot injections, are also known to those skilled in the art. Sustained release formulations may include, for example, porous polymer microparticles or semipermeable polymer matrices in the form of shaped articles, such as films, or microcapsules. Sustained release matrices may include polyesters, hydrogels, polylactides, copolymers of L-glutamic acid and γ-ethyl-L-glutamate, poly(2-hydroxyethyl-in methacrylate), ethylene vinyl acetate, or poly-D(-)-3-hydroxybutyric acid. Sustained release compositions may also include liposomes that can be prepared by any of several methods known in the art.
[0360] The pharmaceutical compositions of the present invention can be administered by various methods known in the art. The route and / or method of administration varies depending on the desired result. Administration can be intravenous, intramuscular, intraperitoneal, or subcutaneous, or can be administered proximal to the target site. The pharmaceutically acceptable carrier should be suitable for intravenous, intramuscular, subcutaneous, parenteral, spinal, or epidermal administration (e.g., by injection or infusion). Depending on the route of administration, the active compound (e.g., the guide nucleic acid, engineered non-natural system, or CRISPR expression system disclosed herein) can be coated with materials to protect the compound from the action of acids and other natural conditions that can inactivate the compound.
[0361] Formulation ingredients suitable for parenteral administration include sterile diluents such as water for injection, physiological saline solution, fixed oils, polyethylene glycol, glycerin, propylene glycol, or other synthetic solvents; antibacterial agents such as benzyl alcohol or methylparaben; antioxidants such as ascorbic acid or sodium bisulfite; chelating agents such as EDTA; buffers such as acetate, citrate, or phosphate; and agents for adjusting tonicity such as sodium chloride or dextrose.
[0362] For intravenous administration, suitable carriers include physiological saline, bacteriostatic water, Cremophor ELTM (BASF, Parsippany, NJ), or phosphate buffered saline (PBS). The carrier should be stable under the conditions of manufacture and storage and should be preserved against microorganisms. The carrier can be, for example, a solvent or dispersion medium containing water, ethanol, polyols (such as glycerol, propylene glycol, and liquid polyethylene glycol), and suitable mixtures thereof.
[0363] The pharmaceutical preparation is preferably sterile. Sterilization can be carried out by any suitable method, such as filtration through a sterile filtration membrane. If the composition is lyophilized, filtration sterilization can be carried out before or after lyophilization and reconstitution. In certain embodiments, the pharmaceutical composition is lyophilized and then reconstituted with buffered saline at the time of administration.
[0364] The pharmaceutical compositions of the present invention are well-known in the art and can be prepared according to methods routinely practiced. See, for example, Remington: The Science and Practice of Pharmacy, Mack Publishing Co., 20th ed., 2000; and Sustained and Controlled Release Drug Delivery Systems, J.R. Robinson, ed., Marcel Dekker, Inc., New York, 1978. The pharmaceutical compositions are preferably manufactured under GMP conditions. Typically, a therapeutically effective dose or an effective dose of the guide nucleic acids, engineered non-natural systems, or CRISPR expression systems disclosed herein is used in the pharmaceutical compositions of the present invention. The compositions disclosed herein are formulated into pharmaceutically acceptable dosage forms by conventional methods known to those skilled in the art. The dosage regimen is adjusted to provide the optimum desired response (e.g., a therapeutic response). For example, a single bolus may be administered, several divided doses may be administered over time, or the dose may be proportionally decreased or increased as indicated by the exigencies of the therapeutic situation. It is particularly advantageous to formulate parenteral compositions in unit dosage form for ease of administration and uniformity of dosage. As used herein, a unit dosage form refers to physically discrete units suitable as unitary dosages for the subjects to be treated; each unit contains a predetermined quantity of the active compound calculated to produce the desired therapeutic effect, together with the necessary pharmaceutical carrier.
[0365] The actual dosage level of the active ingredient in the pharmaceutical composition of the present invention can be varied to obtain an amount of the active ingredient that is effective in obtaining the desired therapeutic response in a specific patient, composition, and method of administration and that is not toxic to the patient. The selected dosage level depends on various pharmacokinetic factors including the activity of the specific composition, or an ester, salt or amide thereof, disclosed herein, the route of administration, the time of administration, the rate of excretion of the specific compound used, the treatment period, other drugs, compounds and / or materials used in combination with the specific composition used, and factors such as the age, sex, weight, condition, general health and medical history of the patient being treated.
[0366] V. Therapeutic Use For example, the guide nucleic acids, engineered non-natural systems, and CRISPR expression systems disclosed herein are useful for targeting, editing, and / or modifying genomic DNA within a cell or organism. These guide nucleic acids and systems, and cells whose genomes have been modified by one of the systems, can be used to treat diseases or disorders in which modification of genetic or epigenetic information is desirable. Accordingly, provided herein is a method of treating a disease or disorder, the method comprising administering to a subject in need thereof a guide nucleic acid, non-natural system, CRISPR expression system, or cell disclosed herein.
[0367] The term "subject" includes humans and non-human animals. Non-human animals include all vertebrates, e.g., mammals and non-mammals, e.g., non-human primates, sheep, dogs, cows, chickens, amphibians, and insects. Unless noted otherwise, the terms "patient" and "subject" are used interchangeably herein.
[0368] As used herein, the terms "treatment", "treating", "treat", "being treated", etc. can refer to obtaining a desired pharmacological and / or physiological effect. The effect can be therapeutic in terms of partial or complete cure of a disease and / or adverse effect caused by the disease, or delay of the progression of the disease. As used herein, "treatment" encompasses any treatment of a disease in a mammal, e.g., a human, and includes (a) suppressing the disease, i.e., arresting its progression; and (b) alleviating the disease, i.e., causing regression of the disease. It is understood that a disease or disorder can be identified by genetic methods and treated before the manifestation of any medical symptoms.
[0369] For minimizing toxicity and off-target effects, it can be important to control the concentration of the delivered CRISPR-Cas system. The optimal concentration can be determined by testing various concentrations in a cell model, tissue model, or non-human eukaryotic animal model and analyzing the degree of modification at potential off-target genomic loci using deep sequencing. The concentration that yields the highest on-target modification level while minimizing the level of off-target modification is generally selected for ex vivo or in vivo delivery.
[0370] It is understood that the guide nucleic acids, engineered non-natural systems, and CRISPR expression systems disclosed herein can be used to treat any suitable disease or disorder that can be ameliorated by the systems in cells.
[0371] For therapeutic purposes, the specific methods disclosed herein are particularly suitable for editing or modifying proliferating cells, such as stem cells (e.g., hematopoietic stem cells), progenitor cells (e.g., hematopoietic progenitor cells or lymphoid progenitor cells), or memory cells (e.g., memory T cells). Considering that such cells are delivered to a subject and proliferate in vivo, the tolerance to off-target events is low. However, it is possible to evaluate on-target and off-target events prior to delivery and thereby select one or more colonies that have the desired editing or modification and do not have the undesired editing or modification. Thus, a lower editing or modification efficiency may be tolerated for such cells. The engineered non-natural systems of the invention have the advantage of increasing or decreasing the efficiency of nucleic acid cleavage, for example, by modulating the hybridization of dual guide nucleic acids. As a result, it can be used to minimize off-target events when generating genetically modified proliferating cells.
[0372] In certain embodiments, the guide nucleic acids, engineered non-natural systems, and / or CRISPR expression systems disclosed herein can be used to engineer immune cells. Immune cells include, but are not limited to, lymphocytes (e.g., B lymphocytes or B cells, T lymphocytes or T cells, and natural killer cells), myeloid cells (e.g., monocytes, macrophages, eosinophils, mast cells, basophils, and granulocytes), and stem and progenitor cells that can differentiate into these cell types (e.g., hematopoietic stem cells, hematopoietic progenitor cells, and lymphoid progenitor cells). The cells can include autologous cells derived from the subject being treated or allogeneic cells derived from a donor.
[0373] In certain embodiments, the immune cells are T cells, which can be, for example, cultured T cells, primary T cells, T cells derived from cultured T cell lines (e.g., Jurkat, SupTi), or T cells obtained from a mammal, such as a subject to be treated. When obtained from a mammal, T cells can be obtained from many sources including, but not limited to, blood, bone marrow, lymph nodes, thymus, or other tissues or body fluids. T cells can also be enriched or purified. The T cells can be of any type of T cell, including, but not limited to, CD4 + / CD8 + double positive T cells, CD4 + helper T cells (e.g., Th1 and Th2 cells), CD8 + T cells (e.g., cytotoxic T cells), tumor infiltrating lymphocytes (TILs), memory T cells (e.g., central memory T cells and effector memory T cells), regulatory T cells, naive T cells, etc., and can be of any developmental stage.
[0374] In certain embodiments, immune cells, such as T cells, are engineered to express an exogenous gene. For example, in certain embodiments, the engineered CRISPR system disclosed herein can catalyze DNA cleavage at a locus and enable site-specific integration of an exogenous gene at the locus by HDR.
[0375] In certain embodiments, immune cells, such as T cells, are engineered to express a chimeric antigen receptor (CAR), i.e., the T cells contain an exogenous nucleotide sequence encoding the CAR. As used herein, the term “chimeric antigen receptor” or “CAR” includes any artificial receptor that contains an antigen-specific binding portion and one or more signaling chains derived from an immunoreceptor. The CAR may include a single-chain variable fragment (scFv) of an antibody specific for an antigen linked, via a hinge and transmembrane region, to the cytoplasmic domain of a T cell signaling molecule, such as a T cell trigger domain (e.g., derived from CD3ζ) and a T cell costimulatory domain (e.g., derived from CD28, CD137, OX40, ICOS, or CD27) in tandem. T cells that express a chimeric antigen receptor are referred to as CAR T cells. Exemplary CAR T cells include CD19-targeted CTL019 cells (see Grupp et al. (2015) BLOOD, 126:4983), 19-28z cells (see Park et al. (2015) J. CLIN. ONCOL., 33:7010), and KTE-C19 cells (see Locke et al. (2015) BLOOD, 126:3991). Further exemplary CAR T cells are described in U.S. Patent Nos. 7,446,190, 8,399,645, 8,906,682, 9,181,527, 9,272,002, 9,266,960, 10,253,086, 10640569, and 10,808,035, and in International (PCT) Publication Nos. WO 2013 / 142034, WO 2015 / 120180, WO 2015 / 188141, WO 2016 / 120220, and WO 2017 / 040945.Exemplary methods of expressing a CAR using a CRISPR system are described in Hale et al. (2017) Mol Ther Methods Clin Dev., 4:192, MacLeod et al. (2017) Mol Ther, 25:949, and Eyquem et al. (2017) Nature, 543:113.
[0376] In certain embodiments, an immune cell, e.g., a T cell, binds to an antigen, e.g., a cancer antigen, via an endogenous T cell receptor (TCR). In certain embodiments, an immune cell, e.g., a T cell, is engineered to express an exogenous TCR, e.g., an exogenous native TCR or an exogenous engineered TCR. The T cell receptor comprises two chains called the α-chain and the β-chain, which combine at the surface of the T cell to form a heterodimeric receptor capable of recognizing MHC-restricted antigens. Each of the α-chain and the β-chain comprises a constant region and a variable region. Each variable region of the α-chain and the β-chain defines three loops called complementarity-determining regions (CDRs), known as CDR1, CDR2, and CDR3, which confer antigen-binding activity and binding specificity to the T cell receptor.
[0377] In certain embodiments, the CAR or TCR binds to a cancer antigen selected from B cell maturation antigen (BCMA), mesothelin, prostate-specific membrane antigen (PSMA), prostate stem cell antigen (PSCA), carbonic anhydrase IX (CAIX), carcinoembryonic antigen (CEA), CD5, CD7, CD10, CD19, CD20, CD22, CD30, CD33, CD34, CD38, CD41, CD44, CD49f, CD56, CD70, CD74, CD123, CD133, CD138, epithelial glycoprotein 2 (EGP2), epithelial glycoprotein-40 (EGP-40), epithelial cell adhesion molecule (EpCAM), receptor tyrosine-protein kinase (FLT3), folate-binding protein (FBP), fetal acetylcholine receptor (AChR), folate receptors -a and β (FRa and β), ganglioside G2 (GD2), ganglioside G3 (GD3), epidermal growth factor receptor 2 (HER-2 / ERB2), epidermal growth factor receptor vIII (EGFRvIII), ERB3, ERB4, human telomerase reverse transcriptase (hTERT), interleukin-13 receptor subunit alpha-2 (IL-13Ra2), K-light chain, kinase insert domain receptor (KDR), Lewis A (CA19.9), Lewis Y (LeY), LI cell adhesion molecule (LICAM), melanoma-associated antigen 1 (melanoma antigen family Al, MAGE-A1), mucin 16 (MUC-16), mucin 1 (MUC-1; e.g., cleaved MUC-1), KG2D ligand, cancer-testis antigen NY-ESO-1, tumor fetal antigen (h5T4), tumor-associated glycoprotein 72 (TAG-72), vascular endothelial growth factor R2 (VEGF-R2), Wilms tumor protein (WT-1), receptor tyrosine-protein kinase transmembrane receptor type 1 (ROR1), B7-H3 (CD276), B7-H6 (Nkp30), chondroitin sulfate proteoglycan-4 (CSPG4), DNAX accessory molecule (DNAM-1), Ephrin type-A receptor 2 (EpHA2), fibroblast activation protein (FAP), Gpl00 / HLA-A2, glypican 3 (GPC3), HA-IH, HERK-V, IL-1 IRa, latent membrane protein 1 (LMP1), neural cell adhesion molecule (N-CAM / CD56), and TRAIL receptor (TRAIL-R).
[0378] Suitable loci for insertion of a CAR coding sequence or an exogenous TCR coding sequence include, but are not limited to, safe harbor loci (e.g., the AAVS1 locus), TCR subunit loci (e.g., the TCRα constant region (TRAC) locus, the TCRβ constant region 1 (TRBC1) locus, and the TCRβ constant region 2 (TRBC2) locus). Insertion at the TRAC locus is understood to reduce tonic CAR signaling and improve T cell function (see Eyquem et al. (2017) NATURE, 543:113). Furthermore, inactivation of the endogenous TRAC, TRBC1, or TRBC2 gene can reduce the graft-versus-host disease (GVHD) response, thereby enabling the use of allogeneic T cells as a starting material for the preparation of CAR T cells. Thus, in certain embodiments, immune cells, such as T cells, are engineered to have reduced expression of an endogenous TCR or TCR subunit, such as TRAC, TRBC1, and / or TRBC2. The cells can be engineered to have partially reduced expression or no expression of the endogenous TCR or TCR subunit. For example, in certain embodiments, immune cells, such as T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) expression of the endogenous TCR or TCR subunit compared to the corresponding unmodified or parental cells. In certain embodiments, immune cells, such as T cells, are engineered to have no detectable expression of the endogenous TCR or TCR subunit. Exemplary methods for reducing TCR expression using the CRISPR system are described in U.S. Patent No. 9,181,527, Liu et al. (2017) CELL RES, 27:154, Ren et al. (2017) CLIN CANCER RES, 23:2255, Cooper et al. (2018) LEUKEMIA, 32:1970, and Ren et al. (2017) ONCOTARGET, 8:17002.
[0379] Certain immune cells, such as T cells, also express the major histocompatibility complex (MHC) or human leukocyte antigen (HLA) genes, and inactivation of these endogenous genes can reduce the immune response, thereby enabling the use of allogeneic T cells as starting materials for the preparation of CAR T cells. Thus, in certain embodiments, immune cells, such as T cells, are engineered to have reduced expression of one or more endogenous class I or class II MHC or HLA (e.g., β2-microglobulin (B2M), class II major histocompatibility complex transactivator (CIITA)). The cells can be engineered to have partially reduced or no expression of the endogenous MHC or HLA. For example, in certain embodiments, immune cells, such as T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) of the expression of the endogenous MHC (e.g., B2M, CIITA) compared to the corresponding unmodified or parental cells. In certain embodiments, immune cells, such as T cells, are engineered to have no detectable expression of the endogenous MHC (e.g., B2M, CIITA). In some cases, the cells can be engineered to have expression of, for example, HLA-E and / or HLA-G to avoid attack by natural killer (NK) cells. Exemplary methods for reducing MHC expression using the CRISPR system are described in Liu et al. (2017) CELL RES, 27:154, Ren et al. (2017) CLIN CANCER RES, 23:2255, and Ren et al. (2017) ONCOTARGET, 8:17002.
[0380] Other genes that can be inactivated include, but are not limited to, CD3, CD52, and deoxycytidine kinase (DCK). For example, inactivation of DCK can render immune cells (e.g., T cells) resistant to purine nucleotide analog (PNA) compounds that are often used to suppress the host immune system to reduce the GVHD response during immune cell therapy. In certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) of the endogenous CD52 or DCK expression compared to the corresponding unmodified or parental cells.
[0381] It is understood that the activity of immune cells (e.g., T cells) can be enhanced by inactivating or reducing the expression of immunosuppressive factors, such as immune checkpoint proteins. Thus, in certain embodiments, immune cells, e.g., T cells, are engineered to have a reduced expression of immune checkpoint proteins. Exemplary immune checkpoint proteins expressed by wild-type T cells include, but are not limited to, PDCD1 (PD-1), CTLA4, ADORA2A (A2AR), B7-H3, B7-H4, BTLA, KIR, LAG3, HAVCR2 (TIM3), TIGIT, VISTA, PTPN6 (SHP-1), and FAS. Cells can be modified to have a partially reduced or no expression of immune checkpoint proteins. For example, in certain embodiments, immune cells, e.g., T cells, are engineered to have less than 80% (e.g., less than 70%, less than 60%, less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5%) of the expression of immune checkpoint proteins compared to the corresponding unmodified or parental cells. In certain embodiments, immune cells, e.g., T cells, are engineered to have no detectable expression of immune checkpoint proteins. Exemplary methods for reducing the expression of immune checkpoint proteins using the CRISPR system are described in International (PCT) Publication No. WO 2017 / 017184 Pamphlet, Cooper et al. (2018) LEUKEMIA, 32:1970, Su et al. (2016) ONCOIMMUNOLOGY, 6:e1249558, and Zhang et al. (2017) FRONT MED, 11:554.
[0382] Immune cells can be engineered by gene editing or modification to have reduced expression of an endogenous gene, such as the endogenous genes described above. For example, in certain embodiments, the engineered CRISPR systems disclosed herein can result in DNA cleavage at a locus, thereby inactivating the targeted gene. In other embodiments, the engineered CRISPR systems disclosed herein can be fused to an effector domain (e.g., a transcriptional repressor or a histone methyltransferase) to reduce the expression of a target gene.
[0383] Immune cells can also be engineered to express an exogenous protein (in addition to the antigen-binding proteins described above) at the locus of a human ADORA2A, B2M, CD52, CIITA, CTLA4, DCK, FAS, HAVCR2, LAG3, PDCD1, PTPN6, TIGIT, TRAC, TRBC1, TRBC2, CARD11, CD247, IL7R, LCK, or PLCG1 gene.
[0384] In certain embodiments, immune cells, such as T cells, are modified to express a dominant-negative form of an immune checkpoint protein. In certain embodiments, the dominant-negative form of the checkpoint inhibitor can act as a decoy receptor that binds to the natural ligand that would otherwise bind to and activate the wild-type immune checkpoint protein or sequesters it. Examples of engineered immune cells, such as T cells, containing a dominant-negative form of an immunosuppressive factor are described, for example, in International (PCT) Publication No. WO 2017 / 040945 pamphlet.
[0385] In certain embodiments, immune cells, such as T cells, are modified to express a gene (e.g., a transcription factor, cytokine, or enzyme) that regulates the survival, proliferation, activity, or differentiation (e.g., into memory cells) of the immune cells. In certain embodiments, the immune cells are modified to express TET2, FOXO1, IL-12, IL-15, IL-18, IL-21, IL-7, GLUT1, GLUT3, HK1, HK2, GAPDH, LDHA, PDK1, PKM2, PFKFB3, PGK1, ENO1, GYS1, and / or ALDOA. In certain embodiments, the modification is the insertion of a nucleotide sequence encoding a protein operably linked to a regulatory element. In certain embodiments, the modification is the substitution of a single nucleotide polymorphism (SNP) site in an endogenous gene. In certain embodiments, immune cells, such as T cells, are modified to express a variant of a gene, e.g., a variant having higher activity than its respective wild-type gene. In certain embodiments, the immune cells are modified to express a variant of CARD11, CD247, IL7R, LCK, or PLCG1. For example, some gain-of-function variants of IL7R are disclosed in Zenatti et al., (2011) NAT. GENET. 43(10):932-39. The variants can be expressed from the native locus of their respective wild-type genes by delivering the engineered system described herein in combination with a donor template having the variant or a portion thereof to target the native locus.
[0386] In certain embodiments, immune cells, such as T cells, are modified to express a protein (e.g., a cytokine or enzyme) that regulates a microenvironment (e.g., the tumor microenvironment) in which the immune cells are designed to migrate. In certain embodiments, the immune cells are modified to express CA9, CA12, a V-ATPase subunit, NHE1, and / or MCT-1.
[0387] Gene Therapy For example, it is understood that the engineered non-natural systems and CRISPR expression systems disclosed herein can be used to treat a genetic disease or disorder, i.e., a disease or disorder associated with or mediated by an undesirable mutation in the genome of a subject.
[0388] Exemplary genetic diseases or disorders include age-related macular degeneration, adrenoleukodystrophy (ALD), Alagille syndrome, α1-antitrypsin deficiency, argininemia, argininosuccinic aciduria, ataxia (e.g., Friedreich ataxia, spinocerebellar ataxia, ataxia telangiectasia, essential tremor, spastic paraplegia), autism, biliary atresia, biotinidase deficiency, carbamoyl phosphate synthetase I deficiency, congenital disorder of glycosylation (CDGS), central nervous system (CNS)-related diseases (e.g., Alzheimer's disease, amyotrophic lateral sclerosis (ALS), Canavan disease (CD), ischemia, multiple sclerosis (MS), neuropathic pain, Parkinson's disease), Bloom syndrome, cancer, Charcot-Marie-Tooth disease (e.g., peroneal muscular atrophy, hereditary motor and sensory neuropathy), congenital hepatic porphyria, citrullinemia, Crigler-Najjar syndrome, cystic fibrosis (CF), dentatorubral-pallidoluysian atrophy (DRPLA), diabetes insipidus, Fabry disease, familial hypercholesterolemia (LDL receptor deficiency), Fanconi anemia, fragile X syndrome, fatty acid metabolism disorder, galactosemia, glucose-6-phosphate dehydrogenase (G6PD), glycogenosis (e.g., type I (glucose-6-phosphatase deficiency, von Gierke II (α-glucosidase deficiency, Pompe disease)), III (debranching enzyme deficiency, Cori disease), IV (branching enzyme deficiency, Andersen disease), V (muscle glycogen phosphorylase deficiency, McArdle disease), VII (muscle phosphofructokinase deficiency, Tarui disease (Tauri)), VI (liver phosphorylase deficiency, Hers disease), IX (liver glycogen phosphorylase kinase deficiency)), hemophilia A (associated with factor VIII deficiency), hemophilia B (associated with factor IX deficiency), Huntington disease, glutaric aciduria, hypophosphatemia, Krabbe disease, lactic acidosis, Lafora disease, Leber congenital amaurosis, Lesch-Nyhan syndrome, lysosomal storage disease, metachromatic leukodystrophy disease (MLD), mucopolysaccharidosis (MPS) (e.g., Hunter syndrome, Hurler syndrome, Maroteaux-Lamy syndrome, Sanfilippo syndrome, Scheie syndrome, Morquio syndrome, others, MPSI, MPSII, MPSIII, MSIV, MPS7) Musculoskeletal disorders (e.g., muscular dystrophy, Duchenne muscular dystrophy), myotonic dystrophy (DM), neovascularization, N-acetylglutamate synthase deficiency, ornithine transcarbamylase deficiency, phenylketonuria, primary open-angle glaucoma, retinitis pigmentosa, schizophrenia, severe combined immunodeficiency (SCID), spinal and bulbar muscular atrophy (SBMA), sickle cell anemia, Asherman syndrome, Tay-Sachs disease, thalassemia (e.g., β-thalassemia), trinucleotide repeat diseases, tyrosinemia, Wilson's disease, Wiskott-Aldrich syndrome, X-linked chronic granulomatous disease (CGD), X-linked severe combined immunodeficiency, and xeroderma pigmentosum.
[0389] Additional exemplary genetic diseases or disorders and related information are available on the World Wide Web at kumc.edu / gec / support, genome.gov / 10001200, and ncbi.nlm.nih.gov / books / NBK22183 / . Additional exemplary genetic diseases or disorders, related gene mutations, and gene therapy techniques for treating genetic diseases or disorders are described in International (PCT) Publication Nos. WO 2013 / 126794, WO 2013 / 163628, WO 2015 / 048577, WO 2015 / 070083, WO 2015 / 089354, WO 2015 / 134812, WO 2015 / 138510, WO 2015 / 148670, WO 2015 / 148860, WO 2015 / 148863, WO 2015 / 153780, WO 2015 / 153789, and WO 2015 / 153791, U.S. Pat. Nos. 8,383,604, 8,859,597, 8,956,828, 9,255,130, and 9,273,296, and U.S. Patent Application Publication Nos. 2009 / 0222937, 2009 / 0271881, 2010 / 0229252, 2010 / 0311124, 2011 / 0016540, 2011 / 0023139, 2011 / 0023144, 2011 / 0023145, 2011 / 0023146, 2011 / 0023153, 2011 / 0091441, 2012 / 0159653, and 2013 / 0145487.
[0390] VI. Kit It is understood that the guide nucleic acids, engineered non-natural systems, CRISPR expression systems, and / or libraries disclosed herein can be packaged in kits suitable for use by a healthcare provider. Accordingly, in another aspect, the invention provides a kit comprising any one or more of the elements disclosed in the above systems, libraries, methods, and compositions. In certain embodiments, the kit comprises an engineered non-natural system disclosed herein and instructions for use for using the kit. The instructions for use can be specific to the uses and methods described herein. In certain embodiments, one or more of the elements of the system are provided in solution. In certain embodiments, one or more of the elements of the system are provided in lyophilized form and the kit further comprises a diluent. The elements can be provided individually or in combination and can be provided in any suitable container, such as a vial, bottle, tube, or immobilized on the surface of a solid substrate (e.g., a chip or microarray). In certain embodiments, the kit comprises one or more of the nucleic acids and / or proteins described herein. In certain embodiments, the kit comprises all of the elements of the system of the invention.
[0391] In certain embodiments of a kit comprising an engineered non-natural dual guide system, the targeter nucleic acid and the modulator nucleic acid are provided in separate containers. In other embodiments, the targeter nucleic acid and the modulator nucleic acid are pre-complexed and the complex is provided in a single container.
[0392] In certain embodiments, the kit comprises a nucleic acid comprising a regulatory element operably linked to a Cas protein or a nucleic acid encoding a Cas protein provided in a separate container. In other embodiments, the kit comprises a Cas protein pre-complexed with a single guide nucleic acid or a combination of a targeter nucleic acid and a modulator nucleic acid, and the complex is provided in a single container.
[0393] In certain embodiments, the kit further comprises one or more donor templates provided in one or more separate containers. In certain embodiments, the kit comprises a plurality of donor templates disclosed herein (e.g., immobilized in separate tubes or on the surface of a solid substrate such as a chip or microarray), one or more guide nucleic acids disclosed herein, and optionally, a Cas protein or a regulatory element operably linked to a nucleic acid encoding a Cas protein disclosed herein. Such kits are useful for identifying donor templates that introduce optimal gene modifications in multiplex assays. The CRISPR expression systems disclosed herein are also suitable for use in kits.
[0394] In certain embodiments, the kit further comprises one or more reagents and / or buffers for use in a process using one or more of the elements described herein. The reagents may be provided in any suitable container and may be provided in a form that is usable in a particular assay or in a form that requires the addition of one or more other components prior to use (e.g., a concentrate or lyophilized form). The buffer may be a reaction or storage buffer including, but not limited to, sodium carbonate buffer, sodium bicarbonate buffer, boric acid buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In certain embodiments, the buffer has a pH of from about 7 to about 10. In certain embodiments, the kit further comprises a pharmaceutically acceptable carrier. In certain embodiments, the kit further comprises one or more devices or other materials for administration to a subject.
[0395] VII. Embodiments In Embodiment 1, a polypeptide that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of SEQ ID NOs: 86 to 124 or 2044 to 2070 is provided herein. In Embodiment 2, a polypeptide according to Embodiment 1, wherein the polypeptide is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of SEQ ID NOs: 86 to 104, 116 to 124, or 2044 to 2070, is provided herein. In Embodiment 3, a polypeptide according to any one of Embodiments 1 or 2, wherein the sequence identity is at least 90%, is provided herein. In Embodiment 4, a polypeptide according to any one of Embodiments 1 or 2, wherein the sequence identity is at least 95%, is provided herein. In Embodiment 5, a polypeptide according to any one of Embodiments 1 or 2, wherein the sequence identity is at least 99%, is provided herein. In Embodiment 6, a polypeptide according to any one of Embodiments 1 or 2, wherein the sequence identity is at least 99.5%, is provided herein. In Embodiment 7, a polypeptide according to any one of Embodiments 1 or 2, wherein the sequence identity is 100%, is provided herein.
[0396] In Embodiment 8, a polynucleotide encoding the polypeptide according to any one of Embodiments 1 to 7 is provided herein.
[0397] In Embodiment 9, (a) a polynucleotide encoding the polypeptide according to any one of Embodiments 1 to 7; and / or (b) a cell containing the polypeptide according to any one of Embodiments 1 to 7 are provided herein.
[0398] A polynucleotide comprising a first portion comprising a first sequence encoding a first polypeptide, wherein the first polypeptide comprises a first chimeric antigen receptor (CAR) or a portion thereof, is provided herein. In Embodiment 11, a polynucleotide as described in Embodiment 10 further comprising a second portion comprising a second sequence encoding a second polypeptide different from the first CAR or a portion thereof is provided herein. In Embodiment 12, a polynucleotide as described in Embodiment 11, wherein the first and second polypeptides are separate polypeptides, is provided herein. In Embodiment 13, a polynucleotide as described in Embodiment 11, wherein the first and second polypeptides are linked, is provided herein. In Embodiment 14, a polynucleotide as described in Embodiment 13, wherein the first and second polypeptides are linked by one or more amino acids, is provided herein. In Embodiment 15, a polynucleotide as described in any one of Embodiments 10-14, wherein the first CAR or a portion thereof binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof, is provided herein. In Embodiment 16, a polynucleotide as described in Embodiment 14, wherein the second CAR or a portion thereof binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof, different from the binding partner of the first CAR or a portion thereof, is provided herein. In Embodiment 17, a polynucleotide as described in Embodiment 14, wherein the second CAR or a portion thereof binds to a different side of the same binding partner, is provided herein.In Embodiment 18, the polynucleotide according to any one of Embodiments 10 to 17, which comprises a polypeptide that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86 to 124 or 2044 to 2070, is provided herein. In Embodiment 19, the polynucleotide according to Embodiment 18, which comprises a polypeptide that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86 to 104, 116 to 124, or 2044 to 2070, is provided herein.
[0399] In Embodiment 20, cells comprising the first polynucleotide described in any one of Embodiments 10 to 19 are provided herein. In Embodiment 21, cells are provided herein, wherein the cells further comprise the second polynucleotide described in any one of Embodiments 10 to 19, and the second polynucleotide is different from the first polynucleotide. In Embodiment 22, cells are provided herein according to any one of Embodiments 20 or 21, wherein the first and / or second polynucleotide further comprises homology arms adjacent to the first and / or second sequences encoding the first and / or second polypeptides. In Embodiment 23, cells are provided herein according to any one of Embodiments 20 or 22, further comprising a nucleic acid-inducible nuclease. In Embodiment 24, cells are provided herein according to Embodiment 23, wherein the nucleic acid-inducible nuclease comprises an engineered non-natural nuclease. In Embodiment 25, cells are provided herein according to any one of Embodiments 23 or 24, wherein the nucleic acid-inducible nuclease comprises a Class 1 or Class 2 nuclease. In Embodiment 26, cells are provided herein according to Embodiment 25, wherein the nucleic acid-inducible nuclease comprises a Type II or Type V nuclease. In Embodiment 27, cells are provided herein according to Embodiment 26, wherein the nucleic acid-inducible nuclease comprises a Type V-A, V-B, V-C, V-D, or V-E nuclease. In Embodiment 28, cells are provided herein according to Embodiment 27, wherein the nucleic acid-inducible nuclease comprises a Type V-A nuclease. In Embodiment 29, cells are provided herein according to Embodiment 28, wherein the nucleic acid-inducible nuclease comprises a MAD nuclease, an ART nuclease, or an ABW nuclease. In Embodiment 30, cells are provided herein according to Embodiment 29, wherein the nucleic acid-inducible nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of a MAD, ART, or ABW nuclease.In Embodiment 31, a cell according to Embodiment 29, wherein the nucleic acid-inducible nuclease comprises a MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, or MAD20 nuclease, is provided herein. In Embodiment 32, a cell according to Embodiment 29, wherein the nucleic acid-inducible nuclease comprises an ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11, ART11. * , ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART29, ART30, ART31, ART32, ART33, ART34, or ART35 nuclease, is provided herein. In Embodiment 33, the nucleic acid-inducible nuclease is MAD2, MAD7, ART2, ART11, or ART11 *Cells as described in Embodiment 29 are provided herein, which comprise an amino acid sequence that is at least 80, 85, 90, 95, 99% identical, or 100% identical, to the amino acid sequence. In Embodiment 34, cells as described in Embodiment 29 are provided herein, wherein the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99% identical, or 100% identical, to the amino acid sequence of SEQ ID NO: 37. In Embodiment 35, cells as described in any one of Embodiments 24-34 are provided herein, wherein the nucleic acid-guided nuclease further comprises a nuclear localization signal (NLS), a purification tag, and / or a cleavage site. In Embodiment 36, cells as described in Embodiment 35 are provided herein, wherein the nucleic acid-guided nuclease comprises at least 4 NLSs. In Embodiment 37, cells as described in Embodiment 36 are provided herein, wherein the nucleic acid-guided nuclease comprises 1 N-terminal and 3 C-terminal NLSs. In Embodiment 38, cells as described in Embodiment 35 are provided herein, wherein the nucleic acid-guided nuclease comprises at least 5 NLSs. In Embodiment 39, cells as described in Embodiment 38 are provided herein, wherein the nucleic acid-guided nuclease comprises 5 N-terminal NLSs. In Embodiment 40, cells as described in any one of Embodiments 35-39 are provided herein, wherein the NLS comprises any one of SEQ ID NOs: 40-56. In Embodiment 41, cells as described in Embodiment 40 are provided herein, wherein the NLS comprises any one of SEQ ID NOs: 40, 51, and 56. In Embodiment 42, cells as described in any one of Embodiments 20-41 are provided herein, which further comprise a guide nucleic acid (gNA). In Embodiment 43, cells as described in Embodiment 42 are provided herein, wherein the gNA comprises: (i) a target nucleic acid comprising a targeter stem sequence and a spacer sequence; and (ii) a modulator nucleic acid comprising a modulator stem sequence complementary to the targeter stem sequence and, optionally, a 5' sequence. In Embodiment 44, cells as described in Embodiment 43 are provided herein, wherein the gNA is an engineered non-natural guide nucleic acid.In Embodiment 45, a cell according to any one of Embodiments 43 or 44, wherein the gNA comprises a single polynucleotide, is provided herein. In Embodiment 46, a cell according to any one of Embodiments 43 or 44, wherein the gNA comprises a dual guide nucleic acid, wherein the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides, is provided herein. In Embodiment 47, a cell according to Embodiment 46, wherein the dual gNA is capable of binding to and activating a nucleic acid-guided nuclease that is activated by a single crRNA in the absence of tracrRNA in a native system, is provided herein. In Embodiment 48, a cell according to any one of Embodiments 42-47, wherein the gNA and the nucleic acid-guided nuclease form a nucleic acid-guided nuclease complex, is provided herein. In Embodiment 49, a cell according to one of Embodiments 42-48, wherein the gNA further comprises a donor template recruitment sequence, is provided herein. In Embodiment 50, a cell according to any one of Embodiments 20-49, wherein the first and / or second polynucleotide is integrated at the first and / or second position in the genome of the cell, is provided herein. In Embodiment 51, a cell according to Embodiment 50, wherein the first and / or second position in the genome of the cell comprises a safe harbor site or a gene encoding a subunit of HLA-1, HLA-2, or TRC protein, is provided herein. In Embodiment 52, a cell according to Embodiment 50, wherein the position in the genome of the cell comprises the TRAC gene, is provided herein. In Embodiment 53, a cell according to any one of Embodiments 50-52, wherein the spacer sequence is at least partially complementary to the target sequence within or near the first and / or second position in the genome of the cell, is provided herein. In Embodiment 54, a cell according to Embodiment 53, wherein the spacer sequence is at least 50, 60, 70, 80, 90, 95, 99, 99.5% identical or 100% identical to the target sequence within or near the first and / or second position in the genome of the cell, is provided herein.In Embodiment 55, the cells described in any one of Embodiments 20 to 54, in which the first and / or second polypeptide is expressed, are provided herein. In Embodiment 56, the cells described in Embodiment 55, in which the first and second polypeptides are expressed on the cell surface, are provided herein. In Embodiment 57, the cells described in Embodiment 55, in which the expression of the second polypeptide is initiated by a signal indicating a change in the cell state, are provided herein. In Embodiment 58, the cells described in Embodiment 57, in which the binding of a binding partner to the first polypeptide initiates a signal indicating a change in the cell state, are provided herein. In Embodiment 59, the cells described in Embodiment 55, in which the first polypeptide is expressed on the cell surface and the second polypeptide is secreted, are provided herein. In Embodiment 60, the cells described in Embodiment 59, in which the expression of the second polypeptide is initiated by a signal indicating a change in the cell state, are provided herein. In Embodiment 61, the cells described in Embodiment 60, in which the binding of a binding partner to the first polypeptide initiates a signal indicating a change in the cell state, are provided herein. In Embodiment 62, the cells described in any one of Embodiments 20 to 61, in which the cell is a human cell, are provided herein. In Embodiment 63, the cells described in Embodiment 62, in which the human cell is an immune cell or a stem cell, are provided herein. In Embodiment 64, the cells described in Embodiment 62, in which the human cell is an immune cell including neutrophils, eosinophils, basophils, mast cells, monocytes, macrophages, dendritic cells, natural killer cells, or lymphocytes, are provided herein. In Embodiment 65, the cells described in Embodiment 62, in which the human cell is a T cell, are provided herein. In Embodiment 66, the cells described in Embodiment 62, in which the human cell is a human totipotent, pluripotent stem cell, embryonic stem cell, induced pluripotent stem cell, hematopoietic stem cell, CD34+ cell stem cell, are provided herein.
[0400] In Embodiment 67, provided herein are cells comprising a first polynucleotide comprising a first portion comprising a first sequence encoding a first polypeptide comprising a first CAR or a portion thereof. In Embodiment 68, provided herein are cells according to Embodiment 67, further comprising a nucleic acid-inducible nuclease system and / or one or more polynucleotides encoding one or more portions of the system, the system comprising (i) a nucleic acid-inducible nuclease; and (ii) a guide nucleic acid (gNA) compatible with the nucleic acid-inducible nuclease. In Embodiment 69, provided herein are cells according to Embodiment 68, wherein the nucleic acid-inducible nuclease comprises a type V nucleic acid-inducible nuclease. In Embodiment 70, provided herein are cells according to any one of Embodiments 67-69, wherein the first polynucleotide comprises a donor template. In Embodiment 71, provided herein are cells according to any one of Embodiments 67-70, wherein the first polynucleotide further comprises a second portion comprising a second sequence encoding a second polypeptide comprising a second CAR or a portion thereof. In Embodiment 72, provided herein are cells according to any one of Embodiments 67-70, wherein the cells further comprise a second polynucleotide comprising a second sequence encoding a second polypeptide comprising a second CAR or a portion thereof. In Embodiment 73, provided herein are cells according to any one of Embodiments 71 or 72, wherein the second CAR or a portion thereof is different from the first CAR or a portion thereof. In Embodiment 74, provided herein are cells according to any one of Embodiments 67-73, wherein at least a portion of the first and / or second polynucleotide is inserted within a locus in the genome of the cell. In Embodiment 75, provided herein are cells according to Embodiment 74, wherein the first polynucleotide is inserted within a first locus in the genome of the cell, the second polynucleotide is inserted within a second locus in the genome of the cell, and the first locus is different from the second locus.In Embodiment 76, the cells according to any one of Embodiments 67 to 75 are provided herein, wherein the first CAR or a portion thereof comprises a polypeptide that binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof. In Embodiment 77, the cells according to Embodiment 76 are provided herein, wherein the second CAR or a portion thereof comprises a polypeptide that binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof, which is different from the binding partner of the first CAR or a portion thereof. In Embodiment 78, the cells according to any one of Embodiments 76 or 77 are provided herein, wherein the first CAR or a portion thereof comprises a polypeptide that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86 to 124 or 2044 to 2070. In Embodiment 79, the cells according to any one of Embodiments 67 to 78 are provided herein, wherein the first and / or second polypeptide is expressed. In Embodiment 80, the cells according to any one of Embodiments 67 to 79 are provided herein, wherein the nucleic acid-inducible nuclease comprises an engineered non-natural nuclease. In Embodiment 81, the cells according to any one of Embodiments 67 to 80 are provided herein, wherein the nucleic acid-inducible nuclease comprises a V-A, V-B, V-C, V-D, or V-E type nuclease. In Embodiment 82, the cells according to Embodiment 81 are provided herein, wherein the nucleic acid-inducible nuclease comprises a V-A type nuclease. In Embodiment 83, the cells according to Embodiment 82 are provided herein, wherein the nucleic acid-inducible nuclease comprises a MAD nuclease, an ART nuclease, or an ABW nuclease.In Embodiment 84, a cell according to Embodiment 83 is provided herein, wherein the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, or 99% identical, or 100% identical, to the amino acid sequence of a MAD, ART, or ABW nuclease. In Embodiment 85, a cell according to Embodiment 83 is provided herein, wherein the nucleic acid-guided nuclease comprises a MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, or MAD20 nuclease. In Embodiment 86, a cell according to Embodiment 83 is provided herein, wherein the nucleic acid-guided nuclease comprises an ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11, * , ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART29, ART30, ART31, ART32, ART33, ART34, or ART35 nuclease. In Embodiment 87, a cell according to Embodiment 83 is provided herein, wherein the nucleic acid-guided nuclease comprises a MAD2, MAD7, ART2, ART11, or ART11 *Cells as described in Embodiment 83 are provided herein, which comprise an amino acid sequence that is at least 80, 85, 90, 95, 99% identical, or 100% identical, to the amino acid sequence. In Embodiment 88, cells as described in Embodiment 83 are provided herein, wherein the nucleic acid-guided nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99% identical, or 100% identical, to the amino acid sequence of SEQ ID NO: 37. In Embodiment 89, cells as described in any one of Embodiments 67-88 are provided herein, wherein the nucleic acid-guided nuclease further comprises a nuclear localization signal (NLS), a purification tag, and / or a cleavage site. In Embodiment 90, cells as described in Embodiment 89 are provided herein, wherein the nucleic acid-guided nuclease comprises at least 4 NLSs. In Embodiment 91, cells as described in Embodiment 90 are provided herein, wherein the nucleic acid-guided nuclease comprises 1 N-terminal and 3 C-terminal NLSs. In Embodiment 92, cells as described in Embodiment 90 are provided herein, wherein the nucleic acid-guided nuclease comprises at least 5 NLSs. In Embodiment 93, cells as described in Embodiment 92 are provided herein, wherein the nucleic acid-guided nuclease comprises 5 N-terminal NLSs. In Embodiment 94, cells as described in any one of Embodiments 89-93 are provided herein, wherein the NLS comprises any one of SEQ ID NOs: 40-56. In Embodiment 95, cells as described in Embodiment 94 are provided herein, wherein the NLS comprises SEQ ID NOs: 40, 51, and 56. In Embodiment 96, gNA comprises: (1) a target nucleic acid comprising a target stem sequence and a spacer sequence is provided herein; (2) a modulator nucleic acid comprising a modulator stem sequence complementary to the target stem sequence and, optionally, a 5' sequence is provided herein. Cells as described in any one of Embodiments 67-95 are provided herein. In Embodiment 97, cells as described in any one of Embodiments 67-96 are provided herein, wherein the gNA is an engineered non-natural guide nucleic acid.In Embodiment 98, cells according to any one of Embodiments 67-97, wherein the gNA comprises a single polynucleotide, are provided herein. In Embodiment 99, cells according to any one of Embodiments 67-97, wherein the gNA comprises a dual guide nucleic acid, where the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides, are provided herein. In Embodiment 100, cells according to Embodiment 99, wherein the dual gNA is capable of binding to and activating a nucleic acid-guided nuclease that is activated by a single crRNA in the absence of tracrRNA in a native system, are provided herein. In Embodiment 101, cells according to any one of Embodiments 67-100, wherein the gNA and the nucleic acid-guided nuclease form a nucleic acid-guided nuclease complex, are provided herein. In Embodiment 102, cells according to any one of Embodiments 67-101, wherein the gNA further comprises a donor template recruit sequence, are provided herein. In Embodiment 103, cells according to any one of Embodiments 67-102, wherein the cells are human cells, are provided herein. In Embodiment 104, cells according to Embodiment 103, wherein the human cells are immune cells or stem cells, are provided herein. In Embodiment 105, cells according to Embodiment 103, wherein the human cells are immune cells including neutrophils, eosinophils, basophils, mast cells, monocytes, macrophages, dendritic cells, natural killer cells, or lymphocytes, are provided herein. In Embodiment 106, cells according to Embodiment 103, wherein the human cells are T cells, are provided herein. In Embodiment 107, cells according to Embodiment 103, wherein the human cells are stem cells that are human totipotent, pluripotent stem cells, embryonic stem cells, induced pluripotent stem cells, hematopoietic stem cells, CD34+ cells, are provided herein. In Embodiment 108, cells according to Embodiment 103, wherein the human cells are induced pluripotent stem cells, are provided herein. In Embodiment 109, cells according to any one of Embodiments 103-108, wherein the cells exhibit reduced immunogenicity when placed within an allogeneic host, are provided herein.In Embodiment 110, the cells provided herein are the cells described in Embodiment 109, which are non-immunogenic when placed in an allogeneic host.
[0401] In Embodiment 111, a composition comprising a plurality of cell populations including first and second cell populations, wherein: (a) the first cell population comprises: (i) a first genomic modification comprising insertion of a first polynucleotide encoding a f...
Claims
1. A polypeptide that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of SEQ ID NOs: 86 to 124 or 2044 to 2070.
2. The polypeptide according to claim 1, wherein the polypeptide is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5, or 100% identical to any one of SEQ ID NOs: 86 to 104, 116 to 124, or 2044 to 2070.
3. The polypeptide according to claim 1 or 2, wherein the sequence identity is at least 90%.
4. The polypeptide according to claim 1 or 2, wherein the sequence identity is at least 95%.
5. The polypeptide according to claim 1 or 2, wherein the sequence identity is at least 99%.
6. The polypeptide according to claim 1 or 2, wherein the sequence identity is at least 99.5%.
7. The polypeptide according to claim 1 or 2, wherein the sequence identity is 100%.
8. A polynucleotide encoding the polypeptide according to any one of claims 1 to 7.
9. (a) A polynucleotide encoding the polypeptide according to any one of claims 1 to 7; and / or (b) The polypeptide according to any one of claims 1 to 7 A cell comprising the same.
10. A polynucleotide comprising a first portion comprising a first sequence encoding a first polypeptide, wherein the first polypeptide comprises a first chimeric antigen receptor (CAR) or a portion thereof.
11. The polynucleotide according to claim 10, further comprising a second portion comprising a second sequence encoding a second polypeptide comprising a second CAR or a portion thereof that is different from the first CAR or a portion thereof. **Claim 12** The polynucleotide according to claim 11, wherein the first and second polypeptides are separate polypeptides. **Claim 13** The polynucleotide according to claim 11, wherein the first and second polypeptides are linked. **Claim 14** The polynucleotide according to claim 13, wherein the first and second polypeptides are linked by one or more amino acids. **Claim 15** The polynucleotide according to any one of claims 10 to 14, wherein the first CAR or a portion thereof binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof. **Claim 16** The polynucleotide according to claim 14, wherein the second CAR or a portion thereof binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof, which is different from the binding partner of the first CAR or a portion thereof. **Claim 17** The polynucleotide according to claim 14, wherein the second CAR or a portion thereof binds to a different side of the same binding partner. **Claim 18** The polynucleotide according to any one of claims 10 to 17, wherein the first or second polypeptide comprises a polypeptide that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, or 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86 to 124 or 2044 to 2070. **Claim 19** The polynucleotide according to claim 18, wherein the first or second polypeptide is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86 to 104, 116 to 124, or 2044 to 2070.
20. A cell comprising the first polynucleotide according to any one of claims 10 to 19.
21. The cell according to claim 20, wherein the cell further comprises a second polynucleotide according to any one of claims 10 to 19, and the second polynucleotide is different from the first polynucleotide.
22. The cell according to claim 20 or 21, wherein the first and / or second polynucleotide further comprises a homology arm adjacent to the first and / or second sequence encoding the first and / or second polypeptide.
23. The cell according to claim 20 or 22, further comprising a nucleic acid-induced nuclease.
24. The cell according to claim 23, wherein the nucleic acid-induced nuclease comprises an engineered non-natural nuclease.
25. The cell according to claim 23 or 24, wherein the nucleic acid-induced nuclease comprises a class 1 or class 2 nuclease.
26. The cell according to claim 25, wherein the nucleic acid-induced nuclease comprises a type II or type V nuclease.
27. The cell according to claim 26, wherein the nucleic acid-induced nuclease comprises a type V-A, V-B, V-C, V-D, or V-E nuclease.
28. The cell according to claim 27, wherein the nucleic acid-induced nuclease comprises a type V-A nuclease.
29. The cell according to claim 28, wherein the nucleic acid-inducible nuclease comprises a MAD nuclease, an ART nuclease, or an ABW nuclease. **Claim 30** The cell according to claim 29, wherein the nucleic acid-inducible nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99% identical, or 100% identical, to the amino acid sequence of a MAD, ART, or ABW nuclease. **Claim 31** The cell according to claim 29, wherein the nucleic acid-inducible nuclease comprises a MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, or MAD20 nuclease. **Claim 32** The nucleic acid-inducible nuclease comprises an ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11 * , ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART29, ART30, ART31, ART32, ART33, ART34, or ART35 nuclease, and the cell according to claim 29. **Claim 33** The cell according to claim 29, wherein the nucleic acid-inducible nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99, or 100% identical to the amino acid sequence of MAD2, MAD7, ART2, ART11, or ART11 * . **Claim 34** The cell according to claim 29, wherein the nucleic acid-inducible nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99% identical, or 100% identical, to the amino acid sequence of SEQ ID NO:
37. **Claim 35** The cell according to any one of claims 24 to 34, wherein the nucleic acid-induced nuclease further comprises a nuclear localization signal (NLS), a purification tag, and / or a cleavage site.
36. The cell according to claim 35, wherein the nucleic acid-induced nuclease comprises at least four NLSs.
37. The cell according to claim 36, wherein the nucleic acid-induced nuclease comprises one N-terminal and three C-terminal NLSs.
38. The cell according to claim 35, wherein the nucleic acid-induced nuclease comprises at least five NLSs.
39. The cell according to claim 38, wherein the nucleic acid-induced nuclease comprises five N-terminal NLSs.
40. The cell according to any one of claims 35 to 39, wherein the NLS comprises any one of SEQ ID NOs: 40 to 56.
41. The cell according to claim 40, wherein the NLS comprises any one of SEQ ID NOs: 40, 51, and 56.
42. The cell according to any one of claims 20 to 41, further comprising a guide nucleic acid (gNA).
43. The gNA is (i) a target nucleic acid comprising a target stem sequence and a spacer sequence; and (ii) a modulator nucleic acid comprising a modulator stem sequence complementary to the target stem sequence and, optionally, a 5' sequence The cell according to claim 42, comprising.
44. The cell according to claim 43, wherein the gNA is an engineered non-natural guide nucleic acid.
45. The cell according to claim 43 or 44, wherein the gNA comprises a single polynucleotide.
46. The cell according to claim 43 or 44, wherein the gNA comprises a dual guide nucleic acid, wherein the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. **Claim 47** The cell according to claim 46, wherein the dual gNA is capable of binding to and activating a nucleic acid-induced nuclease that is activated by a single crRNA in the absence of tracrRNA in a native system. **Claim 48** The cell according to any one of claims 42 to 47, wherein the gNA and the nucleic acid-induced nuclease form a nucleic acid-induced nuclease complex. **Claim 49** The cell according to any one of claims 42 to 48, wherein the gNA further comprises a donor template recruit sequence. **Claim 50** The cell according to any one of claims 20 to 49, wherein the first and / or second polynucleotide is integrated into a first and / or second position in the genome of the cell. **Claim 51** The cell according to claim 50, wherein the first and / or second position in the genome of the cell comprises a safe harbor site or a gene encoding a subunit of HLA-1, HLA-2, or TRC protein. **Claim 52** The cell according to claim 50, wherein the position in the genome of the cell comprises the TRAC gene. **Claim 53** The cell according to any one of claims 50 to 52, wherein the spacer sequence is at least partially complementary to the target sequence within or near the first and / or second position in the genome of the cell. **Claim 54** The cell according to claim 53, wherein the spacer sequence is at least 50, 60, 70, 80, 90, 95, 99, 99.5% identical to or 100% identical to the target sequence within or near the first and / or second position in the genome of the cell.
55. The cell according to any one of claims 20 to 54, wherein the first and / or second polypeptide is expressed.
56. The cell according to claim 55, wherein the first and second polypeptides are expressed on the surface of the cell.
57. The cell according to claim 55, wherein the expression of the second polypeptide is initiated by a signal indicating a change in the state of the cell.
58. The cell according to claim 57, wherein the binding of a binding partner to the first polypeptide initiates the signal indicating a change in the state of the cell.
59. The cell according to claim 55, wherein the first polypeptide is expressed on the surface of the cell and the second polypeptide is secreted.
60. The cell according to claim 59, wherein the expression of the second polypeptide is initiated by a signal indicating a change in the state of the cell.
61. The cell according to claim 60, wherein the binding of a binding partner to the first polypeptide initiates the signal indicating a change in the state of the cell.
62. The cell according to any one of claims 20 to 61, wherein the cell is a human cell.
63. The cell according to claim 62, wherein the human cell is an immune cell or a stem cell.
64. The cell according to claim 62, wherein the human cell is an immune cell including neutrophils, eosinophils, basophils, mast cells, monocytes, macrophages, dendritic cells, natural killer cells, or lymphocytes.
65. The cell according to claim 62, wherein the human cell is a T cell.
66. The cell according to claim 62, wherein the human cell is a stem cell that is a human totipotent, pluripotent stem cell, embryonic stem cell, induced pluripotent stem cell, hematopoietic stem cell, CD34+ cell.
67. (a) A cell comprising a first polynucleotide comprising a first portion comprising a first sequence encoding a first polypeptide comprising a first CAR or a portion thereof.
68. (b) Further comprising a nucleic acid-inducible nuclease system and / or one or more polynucleotides encoding one or more portions of said system, said system comprising (i) a nucleic acid-inducible nuclease; and (ii) a guide nucleic acid (gRNA) compatible with said nucleic acid-inducible nuclease The cell according to claim 67.
69. The cell according to claim 68, wherein the nucleic acid-inducible nuclease comprises a type V nucleic acid-inducible nuclease.
70. The cell according to any one of claims 67 to 69, wherein the first polynucleotide comprises a donor template.
71. The cell according to any one of claims 67 to 70, wherein the first polynucleotide further comprises a second portion comprising a second sequence encoding a second polypeptide comprising a second CAR or a portion thereof.
72. The cell according to any one of claims 67 to 70, wherein the cell further comprises a second polynucleotide comprising a second sequence encoding a second polypeptide comprising a second CAR or a portion thereof.
73. The cell according to claim 71 or 72, wherein the second CAR or a portion thereof is different from the first CAR or a portion thereof.
74. The cell according to any one of claims 67 to 73, wherein at least a portion of the first and / or second polynucleotide is inserted within a locus in the genome of the cell.
75. The cell according to claim 74, wherein the first polynucleotide is inserted into a first position in the genome of the cell, the second polynucleotide is inserted into a second position in the genome of the cell, and the first position is different from the second position.
76. The cell according to any one of claims 67 to 75, wherein the first CAR or a portion thereof comprises a polypeptide that binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof.
77. The cell according to claim 76, wherein the second CAR or a portion thereof comprises a polypeptide that binds to a binding partner different from the binding partner of the first CAR or a portion thereof and comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof.
78. The cell according to claim 76 or 77, wherein the first CAR or a portion thereof comprises a polypeptide that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86 to 124 or 2044 to 2070.
79. The cell according to any one of claims 67 to 78, wherein the first and / or second polypeptide is expressed.
80. The cell according to any one of claims 67 to 79, wherein the nucleic acid-inducible nuclease comprises an engineered non-natural nuclease.
81. The cell according to any one of claims 67 to 80, wherein the nucleic acid-inducible nuclease comprises a V-A, V-B, V-C, V-D, or V-E type nuclease.
82. The cell according to claim 81, wherein the nucleic acid-induced nuclease comprises a V-A type nuclease. **Claim 83** The cell according to claim 82, wherein the nucleic acid-induced nuclease comprises a MAD nuclease, an ART nuclease, or an ABW nuclease. **Claim 84** The cell according to claim 83, wherein the nucleic acid-induced nuclease comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 99% identical, or 100% identical to the amino acid sequence of a MAD, ART, or ABW nuclease. **Claim 85** The cell according to claim 83, wherein the nucleic acid-induced nuclease comprises a MAD1, MAD2, MAD3, MAD4, MAD5, MAD6, MAD7, MAD8, MAD9, MAD10, MAD11, MAD12, MAD13, MAD14, MAD15, MAD16, MAD17, MAD18, MAD19, or MAD20 nuclease. **Claim 86** The cell according to claim 83, wherein the nucleic acid-induced nuclease comprises an ART1, ART2, ART3, ART4, ART5, ART6, ART7, ART8, ART9, ART10, ART11 * , ART12, ART13, ART14, ART15, ART16, ART17, ART18, ART19, ART20, ART21, ART22, ART23, ART24, ART25, ART26, ART27, ART28, ART29, ART30, ART31, ART32, ART33, ART34, or ART35 nuclease. **Claim 87** The cell according to claim 83, wherein the nucleic acid-induced nuclease comprises an amino acid sequence that is at least 80%, 85%, 90%, 95%, 99% identical, or 100% identical to the amino acid sequence of MAD2, MAD7, ART2, ART11, or ART11 * . **Claim 88** The cell according to claim 83, wherein the nucleic acid-induced nuclease comprises an amino acid sequence that is at least 80, 85, 90, 95, 99% identical to the amino acid sequence of SEQ ID NO: 37 or 100% identical.
89. The cell according to any one of claims 67 to 88, wherein the nucleic acid-induced nuclease further comprises a nuclear localization signal (NLS), a purification tag, and / or a cleavage site.
90. The cell according to claim 89, wherein the nucleic acid-induced nuclease comprises at least 4 NLSs.
91. The cell according to claim 90, wherein the nucleic acid-induced nuclease comprises 1 N-terminal and 3 C-terminal NLSs.
92. The cell according to claim 90, wherein the nucleic acid-induced nuclease comprises at least 5 NLSs.
93. The cell according to claim 92, wherein the nucleic acid-induced nuclease comprises 5 N-terminal NLSs.
94. The cell according to any one of claims 89 to 93, wherein the NLS comprises any one of SEQ ID NOs: 40 to 56.
95. The cell according to claim 94, wherein the NLS comprises SEQ ID NOs: 40, 51, and 56.
96. The gNA is (1) a target nucleic acid comprising a target stem sequence and a spacer sequence; and (2) a modulator nucleic acid comprising a modulator stem sequence complementary to the target stem sequence and, optionally, a 5' sequence The cell according to any one of claims 67 to 95, comprising.
97. The cell according to any one of claims 67 to 96, wherein the gNA is an engineered non-natural guide nucleic acid.
98. The cell according to any one of claims 67 to 97, wherein the gNA comprises a single polynucleotide. **Claim 99** The cell according to any one of claims 67 to 97, wherein the gNA comprises a dual guide nucleic acid, wherein the targeter nucleic acid and the modulator nucleic acid are separate polynucleotides. **Claim 100** The cell according to claim 99, wherein the dual gNA is capable of binding to and activating a nucleic acid-inducible nuclease that is activated by a single crRNA in the absence of tracrRNA in a native system. **Claim 101** The cell according to any one of claims 67 to 100, wherein the gNA and the nucleic acid-inducible nuclease form a nucleic acid-inducible nuclease complex. **Claim 102** The cell according to any one of claims 67 to 101, wherein the gNA further comprises a donor template recruit sequence. **Claim 103** The cell according to any one of claims 67 to 102, wherein the cell is a human cell. **Claim 104** The cell according to claim 103, wherein the human cell is an immune cell or a stem cell. **Claim 105** The cell according to claim 103, wherein the human cell is an immune cell comprising neutrophils, eosinophils, basophils, mast cells, monocytes, macrophages, dendritic cells, natural killer cells, or lymphocytes. **Claim 106** The cell according to claim 103, wherein the human cell is a T cell. **Claim 107** The cell according to claim 103, wherein the human cell is a stem cell that is a human totipotent, pluripotent stem cell, embryonic stem cell, induced pluripotent stem cell, hematopoietic stem cell, or CD34+ cell. **Claim 108** The cell according to claim 103, wherein the human cell is an induced pluripotent stem cell.
109. The cell according to any one of claims 103 to 108, wherein the cell exhibits reduced immunogenicity when placed in an allogeneic host.
110. The cell according to claim 109, wherein the cell is non-immunogenic when placed in an allogeneic host.
111. A composition comprising a plurality of cell populations including first and second cell populations, (a) the first cell population comprises (i) a first genomic modification comprising insertion of a first polynucleotide encoding a first CAR or a portion thereof at a first position in the genome; and (ii) a second genomic modification comprising insertion of a second polynucleotide encoding a second CAR or a portion thereof at a second position in the genome comprising; (b) the second cell population comprises the first genomic modification or the second genomic modification, but not both.
112. The composition according to claim 111, further comprising a third cell population different from the first and second cell populations, wherein the third cell population comprises the first genomic modification or the second genomic modification, but not both.
113. The composition according to claim 112, further comprising a fourth cell population different from the first, second, and third cell populations.
114. A composition comprising a pharmaceutically acceptable excipient and a composition according to any one of claims 20 to 113.
115. A composition, (a) a first type V nucleic acid-guided nuclease; (b) a first guide nucleic acid (gNA); and (c) a first donor template comprising a first polynucleotide encoding a first polypeptide comprising a first CAR or a portion thereof A composition comprising
116. (d) A second type V nucleic acid-guided nuclease; (e) A second gRNA; and (f) A second donor template comprising a second polynucleotide encoding a second polypeptide comprising a CAR or a portion thereof The composition according to claim 115, further comprising
117. (g) A third type V nucleic acid-guided nuclease; (h) A third gRNA; and (i) A third donor template comprising a third polynucleotide different from the first and / or second polynucleotides The composition according to claim 116, further comprising
118. The composition according to any one of claims 115 to 117, wherein the nucleic acid-guided nuclease comprises MAD7.
119. The composition according to any one of claims 115 to 118, wherein the gRNA comprises a dual gRNA.
120. The composition according to any one of claims 115 to 119, wherein the first and second polypeptides are a CAR or a portion thereof that binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof.
121. The cell according to claim 120, wherein the first and second polypeptides are at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124 or 2044-2070.
122. The cell according to any one of claims 115 to 119, wherein the first and second polypeptides bind to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof, and comprises a CAR or a portion thereof.
123. The cell according to claim 122, wherein the first and second polypeptides are at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86 to 104, 116 to 124, or 2044 to 2070.
124. The cell according to any one of claims 115 to 123, wherein the first CAR or a portion thereof binds to a binding partner different from the second CAR or a portion thereof.
125. A composition comprising a pharmaceutically acceptable excipient and a cell according to any one of claims 20 to 124.
126. A composition for editing a plurality of sites in the genome of a target cell, wherein for each integer x, the composition comprises: (a) a donor template (D) comprising a polynucleotide encoding a polypeptide comprising a CAR or a portion thereof (CAR) x ; x ; (b) a type V nucleic acid-guided nuclease (N) x ; (c) a guide nucleic acid (gNA) x ; wherein (CAR) x is different for each integer.
127. (N) x comprises MAD7, the composition according to claim 126.
128. (gNA) x comprises a dual gNA, the composition according to claim 126 or 127.
129. (D) x encodes a polypeptide comprising a CAR or a portion thereof that binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD19, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof, The composition according to any one of claims 126 to 128.
130. (D) x encodes a polypeptide comprising a CAR or a portion thereof that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-124 or 2044-2070, The composition according to claim 129.
131. (D) x encodes a polypeptide comprising a CAR or a portion thereof that binds to a binding partner comprising B7H3, BCMA, GPRC5D, CD8, CD8a, CD20, CD22, CD28, 4-1BB, or CD3ζ or a portion thereof, The composition according to any one of claims 126 to 128.
132. (D) x encodes a polypeptide comprising a CAR or a portion thereof that is at least 60, 70, 75, 80, 85, 90, 95, 96, 97, 98, 99, 99.5% identical or 100% identical to any one of the amino acid sequences of SEQ ID NOs: 86-104, 116-124, or 2044-2070, The composition according to claim 132.
133. The number of different integers x is at least 2, 3, 4, 5, 6, 7, 8, or 9 and 10, 9, 8, 6, 5, 4, or 3 or less, The composition according to any one of claims 126 to 132.
134. The number of different integers x is 2 to 10, The composition according to claim 133.
135. The composition according to claim 134, wherein the number of different integers x is from 2 to 5.
136. A method for editing the genome of a target cell, comprising introducing the first composition according to any one of claims 115 to 125 into the target cell.
137. The method according to claim 136, further comprising adding an additive for stabilizing the nucleic acid-induced nuclease complex to the first composition before introducing the composition into the cell.
138. The method according to claim 136 or 137, wherein the introduction comprises electroporation.
139. The method according to any one of claims 136 to 138, further comprising adding an additive for reducing non-homologous end joining (NHEJ) to the cell growth medium before or after introducing the composition into the cell.
140. The method according to any one of claims 136 to 139, further comprising growing the cell after introducing the composition into the cell.
141. The method according to any one of claims 136 to 139, further comprising differentiating the cell after introducing and / or growing the composition in the cell.
142. The method according to any one of claims 136 to 141, further comprising introducing a second composition comprising the composition according to any one of claims 115 to 125, wherein the second composition is different from the first composition in the cell and / or the progeny of the cell.
143. The method according to claim 142, further comprising introducing a third composition comprising the composition according to any one of claims 115 to 125, wherein the third composition is different from the first and / or second composition in the cell and / or the progeny of the cell.
144. One or more progeny of one or more of the cells according to any one of claim 9 or claims 20 to 110.