Modified immune cells with adenosine deaminase base editor for modifying nucleobases in target sequence

By using the novel adenosine base editor (ABE8) for genetic modification, the problem of insufficient genomic rearrangement and immunosuppressive resistance in the prior art was solved, and the efficient anti-tumor activity and immunosuppressive resistance of CAR-T cells were achieved.

CN120174005APending Publication Date: 2025-06-20BEAM THERAPEUTICS INC
View PDF 38 Cites 0 Cited by

Patent Information

Application Number
CN202510195465.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-01-27
Filing Date
2020-02-13
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art can easily lead to genomic rearrangement when gene editing CAR-T cells, affecting cell function and immunosuppressive resistance, and has a higher risk of graft-versus-host response.

Method used

Genetic modification is performed using a novel adenosine base editor (such as ABE8), by expressing or introducing nucleobase editor polypeptides and contacting cells with guide RNA, affecting the changes in nucleic acid molecules to encode polypeptides that enhance CAR-T cell function.

Benefits of technology

It improves the anti-tumor activity of CAR-T cells, enhances resistance to immunosuppression, and reduces the risk of causing graft-versus-host responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120174005A_ABST
    Figure CN120174005A_ABST
Patent Text Reader

Abstract

The present invention relates to a modified immune cell having an adenosine deaminase base editor for modifying nucleobases in a target sequence. A genetically modified immune cell characterized by a novel adenosine base editor (e.g., ABE8) having enhanced anti-tumor activity, resistance to immunosuppression, and reduced risk of causing a graft-versus-host response or host-versus-graft response, or a combination thereof. The invention also features methods of producing and using these modified immune effector cells.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This invention is a divisional application of a patent application with Chinese Patent Application No. 202080028181.2, invention title "Modified immune cells with adenosine deaminase base editors for modifying nucleobases in target sequences", and application date February 13, 2020.

[0002] Related applications

[0003] This application is an international PCT application claiming priority and benefit of U.S. Provisional Application No. 62 / 805,271 filed on February 13, 2019; No. 62 / 852,228 filed on May 23, 2019; No. 62 / 852,224 filed on May 23, 2019; No. 62 / 931,722 filed on November 6, 2019; No. 62 / 941,523 filed on November 27, 2019; No. 62 / 941,569 filed on November 27, 2019; and No. 62 / 966,526 filed on January 27, 2020, the entire contents of which are hereby incorporated herein by reference.

[0004] Incorporation by reference

[0005] All publications, patents, and patent applications mentioned in this specification are hereby incorporated herein by reference to the extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Unless otherwise specified, the publications, patents, and patent applications mentioned in this specification are incorporated herein by reference in their entirety. Technical field

[0006] This application relates to genetically modified immune cells comprising novel adenosine base editors, which have enhanced anti-tumor activity, resistance to immunosuppression, and reduced risk of causing graft-versus-host reaction or host-versus-graft reaction or a combination thereof. Background art

[0007] Autologous and allogeneic immunotherapies are cancer treatment methods in which immune cells expressing chimeric antigen receptors are administered to a subject. To generate immune cells expressing chimeric antigen receptors (CARs), immune cells are first collected from a subject (autologous) or a donor isolated from a subject being treated (allogeneic) and genetically modified to express a chimeric antigen receptor. The resulting cells express the chimeric antigen receptor on their cell surface (e.g., CAR T cells), and after administration to a subject, the chimeric antigen receptor binds to a marker expressed on tumor cells. This interaction with the tumor marker activates the CAR-T cells, which then kill the tumor cells. However, for autologous or allogeneic cell therapies to be effective and efficient, significant conditions and cellular responses, such as inhibition of T cell signaling, must be overcome or avoided. For allogeneic cell therapies, graft-versus-host disease (GVHD) and host rejection of CAR-T cells can pose additional challenges. Editing genes involved in these processes can enhance CAR-T cell function and resistance to immunosuppression or inhibition, but current methods for performing such editing have the potential to induce substantial genomic rearrangements in CAR-T cells, thereby negatively impacting their efficacy. Thus, there is an urgent need for techniques to more precisely modify immune cells, particularly CAR-T cells. This application addresses this need and other important needs. SUMMARY OF THE INVENTION

[0008] The present invention features genetically modified immune cells comprising a novel adenosine base editor (e.g., ABE8), which have enhanced anti-tumor activity, resistance to immunosuppression, and a reduced risk of causing graft-versus-host or host-versus-graft responses, or a combination thereof. The present invention also features methods for producing and using these modified immune effector cells.

[0009] On the one hand, the present invention provides a method for generating modified immune cells, the method comprising expressing or introducing a nucleobase editor polypeptide in immune cells and contacting the cells with two or more guide RNAs targeting the nucleobase editor polypeptide to effect alteration of a nucleic acid molecule encoding at least one polypeptide selected from the group consisting of T cell receptor alpha constant (TRAC), beta-2 microglobulin (B2M), programmed cell death 1 (PD1), cluster of differentiation 7 (CD7), cluster of differentiation 5 (CD5), cluster of differentiation 33 (CD33), cluster of differentiation 123 (CD123), Cbl proto-oncogene B (CBLB), and class II major histocompatibility complex transactivator (CIITA) polypeptide, wherein the nucleobase editor polypeptide comprises a nucleic acid programmable DNA binding protein (napDNAbp) and a base editor domain, the base editor domain comprising an adenosine deaminase variant domain comprising an alteration at amino acid position 82 and / or 166 of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD. In one embodiment, the immune cell is a T cell. In one embodiment, the immune cell is taken from a healthy subject.

[0010] In one embodiment, the adenosine deaminase variant domain comprises alterations at amino acid positions 82 and 166. In one embodiment, the adenosine deaminase variant domain comprises the alteration V82S. In one embodiment, the adenosine deaminase variant domain comprises the alteration T166R. In one embodiment, the adenosine deaminase variant domain comprises the alterations V82S and T166R. In one embodiment, the adenosine deaminase variant domain further comprises one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, and / or Q154R. In one embodiment, the adenosine deaminase variant domain comprises an alteration combination selected from the group consisting of: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R. In one embodiment, the adenosine deaminase variant domain comprises the alteration combination: V82S+Q154R. In one embodiment, the adenosine deaminase variant domain comprises the alteration combination: Y147R+Q154R+Y123H. In one embodiment, the adenosine deaminase variant domain comprises the alteration combination: Y147R+Q154R+Y123H+I76Y. In one embodiment, the adenosine deaminase variant domain comprises the alteration combination: I76Y+V82S+Y123H+Y147R+Q154R. In one embodiment, the adenosine deaminase variant is TadA*8. In one embodiment, the TadA*8 is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, TadA*8.24.

[0011] In one embodiment, the adenosine deaminase variant domain comprises a C-terminal deletion starting from a residue selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157. In one embodiment, the base editor domain is an adenosine deaminase variant monomer. In one embodiment, the base editor domain is ABE8.1-m, ABE8.2-m, ABE8.3-m, ABE8.4-m, ABE8.5-m, ABE8.6-m, ABE8.7-m, ABE8.8-m, ABE8.9-m, ABE8.10-m, ABE8.11-m, ABE8.12-m, ABE8.13-m, ABE8.14-m, ABE8.15-m, ABE8.16-m, ABE8.17-m, ABE8.18-m, ABE8.19-m, ABE8.20-m, ABE8.21-m, ABE8.22-m, ABE8.23-m, ABE8.24-m.

[0012] In one embodiment, the base editor domain is an adenosine deaminase variant heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant domain. In one embodiment, the base editor domain is ABE8.1-d, ABE8.2-d, ABE8.3-d, ABE8.4-d, ABE8.5-d, ABE8.6-d, ABE8.7-d, ABE8.8-d, ABE8.9-d, ABE8.10-d, ABE8.11-d, ABE8.12-d, ABE8.13-d, ABE8.14-d, ABE8.15-d, ABE8.16-d, ABE8.17-d, ABE8.18-d, ABE8.19-d, ABE8.20-d, ABE8.21-d, ABE8.22-d, ABE8.23-d, or ABE8.24-d.

[0013] In one embodiment, the base editor domain is an adenosine deaminase variant heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain. In one embodiment, the adenosine deaminase variant domain lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length adenosine deaminase. In one embodiment, the adenosine deaminase variant domain comprises or consists essentially of the following sequences having adenosine deaminase activity or fragments thereof:

[0014] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD。

[0015] In one embodiment, the napDNAbp comprises the following sequence:

[0016]

[0017] wherein the bold sequence represents the sequence derived from Cas9, the italic sequence represents the linker sequence, and the underlined sequence represents the dual nuclear localization sequence.

[0018] In various embodiments of any aspect described herein, the napDNAbp is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In one embodiment, the napDNAbp comprises a variant of SpCas9 that has an altered protospacer adjacent motif (PAM) specificity or is specific for a non-G PAM. In one embodiment, the altered PAM is specific for the nucleic acid sequence 5'-NGC-3'. In one embodiment, the modified SpCas9 comprises amino acid substitutions of D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R, or corresponding amino acid substitutions. In various embodiments of any aspect described herein, the napDNAbp comprises nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9. In one embodiment, the nickase variant comprises the amino acid substitution D10A or a corresponding amino acid substitution. In various embodiments of any aspect described herein, the base editor polypeptide further comprises a zinc finger domain. In various embodiments of any aspect described herein, the base editor polypeptide further comprises one or more uracil glycosylase inhibitors. In various embodiments of any aspect described herein, the adenosine deaminase variant domain is capable of deaminating adenine in deoxyribonucleic acid (DNA). In various embodiments of any aspect described herein, the adenosine deaminase variant domain is a modified adenosine deaminase not found in nature. In various embodiments of any aspect described herein, the adenosine deaminase variant is TadA*8. In some embodiments, the TadA*8 is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24.

[0019] In various embodiments of any aspect described herein, the nucleobase editor polypeptide further comprises a linker between the napDNAbp and the adenosine deaminase variant domain. In one embodiment, the linker comprises the amino acid sequence:

[0020] SGGSSGGSSGSETPGTSESATPES。

[0021] In various embodiments of any aspect described herein, the base editor polypeptide further comprises one or more nuclear localization signals (NLSs). In one embodiment, the NLS is a bidirectional NLS. In one embodiment, the nuclear base editor polypeptide comprises an N-terminal NLS and a C-terminal NLS. In various embodiments of any aspect described herein, the napDNAbp is a modified Staphylococcus aureus Cas9 (SaCas9). In one embodiment, the modified SaCas9 comprises amino acid substitutions of E782K, N968K, and R1015H, or corresponding amino acid substitutions thereof. In one embodiment, the modified SaCas9 comprises the amino acid sequence:

[0022]

[0023] In various embodiments of any aspect described herein, two or more guide RNAs are expressed in or contacted with a cell, each targeting a separate polynucleotide. In various embodiments, multiplex base editing involves simultaneously modifying 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more target genomic loci. In various embodiments of any aspect described herein, two or more guide RNAs are expressed in or contacted with a cell, each individually targeting the B2M or TRAC polynucleotide. In various embodiments of any aspect described herein, three guide RNAs are expressed in or contacted with a cell. In various embodiments of any aspect described herein, three guide RNAs are expressed in or contacted with a cell, each individually targeting the B2M, CD7, TRAC, CIITA, PDCD1, and / or CBLC polynucleotides. In various embodiments of any aspect described herein, three guide RNAs are expressed in or contacted with a cell, each individually targeting the B2M, TRAC, and PDCD1 polynucleotides. In various embodiments of any aspect described herein, three guide RNAs are expressed in or contacted with a cell, each individually targeting the B2M, TRAC, and CIITA polynucleotides. In various embodiments of any aspect described herein, four guide RNAs are expressed in or contacted with a cell, each individually targeting the B2M, CD7, TRAC, CIITA, PDCD1, and / or CBLC polynucleotides. In various embodiments of any aspect described herein, the two or more guide RNAs target the TRAC exon 4 splice acceptor site, the B2M exon 1 splice donor site, and / or the PDCD1 exon 1 splice donor site. In various embodiments of any aspect described herein, the two or more guide RNAs target a splice acceptor site or a splice donor site in the target polynucleotide. In various embodiments of any aspect described herein, the nucleobase editor polypeptide generates a stop codon in the target polynucleotide.. In various embodiments of any aspect described herein, the nucleobase editor polypeptide generates a stop codon in PDCD1 exon 2. In various embodiments, by introducing a base editor and a guide RNA targeting one or more genes encoding a polypeptide, the expression of one or more of the above polypeptides is reduced by 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99% or more, or even 100% relative to a reference.

[0024] On the other hand, the present invention provides for the expression of a chimeric antigen receptor (CAR) in modified immune cells of any of the aspects described herein. In various embodiments of any of the aspects described herein, the immune cells are modified ex vivo. In various embodiments of any of the aspects described herein, the immune cells are cytotoxic T cells, regulatory T cells, or T helper cells. In various embodiments of any of the aspects described herein, the modified immune cells do not contain detectable translocations.

[0025] On the other hand, the present invention provides modified immune cells produced by the methods of any of the aspects described herein. In various embodiments of any of the aspects described herein, the cells have reduced immunogenicity and enhanced anti-tumor activity. In various embodiments of any of the aspects described herein, the immune cells express a chimeric antigen receptor.

[0026] In various embodiments of any aspect described herein, the immune cell is a T cell. In various embodiments of any aspect described herein, the cell comprises one or more mutations in a polynucleotide encoding B2M, CD7, CIITA, PD1, CBLB, and / or TRAC. In one embodiment, the cell comprises one or more mutations in a polynucleotide encoding B2M, TRAC, and CIITA. In various embodiments of any aspect described herein, the cell comprises mutations in one or more polynucleotides encoding TIGIT, TGFBR2, ZAP70, NFATc1, or TET2. In various embodiments of any aspect described herein, the cell comprises mutations in one or more polynucleotides encoding V-Set immunoregulatory receptor (VISTA), T cell immunoglobulin mucin 3 (Tim-3), T cell immunoreceptor with Ig and ITIM domains (TIGIT), transforming growth factor beta receptor II (TGFbRII), regulatory factor X-associated ankyrin-containing protein (RFXANK), PVR-related immunoglobulin domain-containing protein (PVRIG), lymphocyte activation gene 3 (Lag3), cytotoxic T-lymphocyte-associated protein 4 (CTLA-4), chitinase 3-like 1 (Chi3l1), cluster of differentiation 96 (CD96), B and T lymphocyte associated (BTLA), Tet methylcytosine dioxygenase 2 (TET2), Sprouty RTK signaling antagonist 1 (Spry1), Sprouty RTK signaling antagonist 2 (Spry2), class II major histocompatibility complex transactivator (CIITA), cluster of differentiation 7 (CD7), cluster of differentiation 33 (CD33), cluster of differentiation 52 (CD52), cluster of differentiation 123 (CD123), T cell receptor beta constant 1 (TRBC1), T cell receptor beta constant 2 (TRBC2), cytokine-inducible SH2-containing protein (CISH), acetyl-CoA acetyltransferase 1 (ACAT1), cytochrome P450 family 11 subfamily A member 1 (Cyp11a1), GATA-binding protein 3 (GATA3), nuclear receptor subfamily 4A group member 1 (NR4A1), nuclear receptor subfamily 4A group member 2 (NR4A2), nuclear receptor subfamily 4A group member 3 (NR4A3), methylation-controlled J protein (MCJ), Fas cell surface death receptor (FAS), or selectin P ligand / P-selectin glycoprotein ligand-1 (SELPG / PSGL1).

[0027] In various embodiments of any aspect described herein, the chimeric antigen receptor comprises an extracellular domain having affinity for a tumor-associated marker. In one embodiment, the tumor is multiple myeloma. In various embodiments of any aspect described herein, the marker is B cell maturation antigen (BCMA).

[0028] On the other hand, the present invention provides a method for modulating an immune response in a subject, the method comprising administering an effective amount of modified immune cells according to any aspect described herein. In various embodiments of any aspect described herein, the method increases or decreases the immune response.

[0029] On the other hand, the present invention provides a method for treating a tumor in a subject, the method comprising administering an effective amount of modified immune cells according to any aspect described herein.

[0030] On the other hand, the present invention provides a pharmaceutical composition for treating a tumor, the pharmaceutical composition comprising an effective amount of modified immune cells according to any aspect described herein.

[0031] On the other hand, the present invention provides a pharmaceutical composition, the pharmaceutical composition comprising an effective amount of modified immune cells according to any aspect described herein in a pharmaceutically acceptable excipient.

[0032] On the other hand, the present invention provides a kit for treating a tumor, the kit comprising modified immune cells according to any aspect described herein. In various embodiments of any aspect described herein, the kit comprises written instructions for treating a tumor using the modified immune effector cells.

[0033] In various embodiments of any aspect described herein, the modified immune cells further comprise a chimeric antigen receptor having an affinity for a tumor-associated marker. In certain embodiments, the chimeric antigen receptor is introduced into the cell by a viral vector such as a lentiviral vector. In certain embodiments, the chimeric antigen receptor is introduced into the cell by a double-stranded DNA template to insert at a locus cleaved by a nuclease. In various embodiments of any aspect described herein, the chimeric antigen receptor comprises an extracellular domain having an affinity for a tumor-associated marker.

[0034] In various embodiments of any aspect described herein, the tumor is a B-cell carcinoma. In various embodiments of any aspect described herein, the B-cell carcinoma is lymphoma or leukemia. In various embodiments of any aspect described herein, the B-cell carcinoma is multiple myeloma.

[0035] On the other hand, the present invention provides a method of treating a subject having or at risk of developing graft-versus-host disease (GVHD) with an effective amount of a modified immune cell according to any aspect described herein. On the other hand, the present invention provides a pharmaceutical composition for treating GVHD, the pharmaceutical composition comprising an effective amount of a modified immune cell according to any aspect described herein. On the other hand, the present invention provides a kit for treating GVHD, the kit comprising a modified immune cell according to any aspect described herein. In various embodiments of any aspect described herein, the modified immune cell lacks or has reduced levels of functional TRAC.

[0036] On the other hand, the present invention provides a method of treating a subject having or at risk of developing host-versus-graft disease (HVGD) with an effective amount of a modified immune cell according to any aspect described herein. On the other hand, the present invention provides a pharmaceutical composition for treating HVGD, the pharmaceutical composition comprising an effective amount of a modified immune cell according to any aspect described herein. On the other hand, the present invention provides a kit for treating HVGD, the kit comprising a modified immune cell according to any aspect described herein. In various embodiments of any aspect described herein, the modified immune cell lacks or has reduced levels of functional B2M.

[0037] On the other hand, the present invention provides a method of generating a modified immune cell, the method comprising expressing or introducing a nucleobase editor polypeptide in an immune cell and contacting the cell with two or more guide RNAs capable of targeting a nucleic acid molecule encoding at least one polypeptide selected from the group consisting of T cell receptor alpha constant (TRAC), beta-2 microglobulin (B2M), programmed cell death 1 (PD1), cluster of differentiation 7 (CD7), cluster of differentiation 5 (CD5), cluster of differentiation 33 (CD33), cluster of differentiation 123 (CD123), Cbl proto-oncogene B (CBLB), and class II major histocompatibility complex transactivator (CIITA) polypeptide, wherein the nucleobase editor polypeptide comprises at least one base adenosine deaminase variant domain inserted within a nucleic acid programmable DNA binding protein (napDNAbp).

[0038] In one embodiment, the adenosine deaminase variant domain comprises the amino acid sequence:

[0039] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD, wherein the amino acid sequence comprises at least one alteration. In one embodiment, the adenosine deaminase variant domain comprises an alteration at amino acid position 82 and / or 166. In one embodiment, the at least one alteration comprises: V82S, T166R, Y147T, Y147R, Q154S, Y123H, and / or Q154R. In one embodiment, the adenosine deaminase variant comprises one of the following combinations of alterations: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R. In one embodiment, the adenosine deaminase variant is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, TadA*8.24. In one embodiment, the adenosine deaminase variant comprises a C-terminal deletion starting from a residue selected from the group consisting of 149, 150, 151, 152, 153, 154, 155, 156, and 157. In one embodiment, the adenosine deaminase variant domain is an adenosine deaminase monomer. In one embodiment, the adenosine deaminase variant is an adenosine deaminase heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant domain.In one embodiment, the adenosine deaminase variant is an adenosine deaminase heterodimer comprising a TadA domain and an adenosine deaminase variant domain.

[0040] In one embodiment, the napDNAbp is a Cas9 or Cas12 polypeptide. In one embodiment, the adenosine deaminase variant is inserted within a flexible loop, α-helical region, unstructured portion, or solvent-accessible portion of the napDNAbp. In one embodiment, the adenosine deaminase variant is flanked by an N-terminal fragment and a C-terminal fragment of the napDNAbp. In one embodiment, the nucleobase editor polypeptide has the structure NH2-[N-terminal fragment of napDNAbp]-[adenosine deaminase variant]-[C-terminal fragment of napDNAbp]-COOH, where each instance of “]-[” is an optional linker. In one embodiment, the C-terminus of the N-terminal fragment or the N-terminus of the C-terminal fragment comprises a portion of the flexible loop of the napDNAbp. In one embodiment, the flexible loop comprises amino acids proximal to the target nucleobase. In one embodiment, the target nucleobase is 1 to 20 nucleobases from the PAM sequence in the target polynucleotide sequence. In some embodiments, the target nucleobase is 2-12 nucleobases upstream of the PAM sequence. In one embodiment, the N-terminal fragment or the C-terminal fragment of the napDNAbp binds the target polynucleotide sequence.

[0041] In some embodiments, the N-terminal fragment or the C-terminal fragment comprises a RuvC domain; the N-terminal fragment or the C-terminal fragment comprises an NHN domain; neither the N-terminal fragment nor the C-terminal fragment comprises an HNH domain; or neither the N-terminal fragment nor the C-terminal fragment comprises a RuvC domain. In some embodiments, the napDNAbp comprises a partial or complete deletion in one or more domains, and wherein the deaminase is inserted at the partial or complete deletion of the napDNAbp. In some embodiments, the deletion is within the RuvC domain; the deletion is within the HNH domain; or the deletion bridges the RuvC domain and the C-terminal domain, the L-I domain and the HNH domain, or the RuvC domain and the L-I domain.

[0042] In another embodiment, the napDNAbp is a Cas9 or Cas12 polypeptide. In one embodiment, the napDNAbp comprises a Cas9 polypeptide. In one embodiment, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a variant thereof. In one embodiment, the Cas9 polypeptide has the following amino acid sequence (Cas9 reference sequence):

[0043]

[0044] (Single underline: HNH domain; double underline: RuvC domain; (Cas9 reference sequence), or its corresponding region.

[0045] In some embodiments, the Cas9 polypeptide comprises a deletion of amino acids numbered 1017 to 1069 or their corresponding amino acids in the Cas9 polypeptide reference sequence; the Cas9 polypeptide comprises a deletion of amino acids numbered 792 to 872 or their corresponding amino acids in the Cas9 polypeptide reference sequence; or the Cas9 polypeptide comprises a deletion of amino acids numbered 792 to 906 or their corresponding amino acids in the Cas9 polypeptide reference sequence. In one embodiment, the adenosine deaminase variant is inserted into a flexible loop of the Cas9 polypeptide. In one embodiment, the flexible loop comprises a region selected from the group consisting of amino acid residues numbered 530 to 537, 569 to 579, 686 to 691, 768 to 793, 943 to 947, 1002 to 1040, 1052 to 1077, 1232 to 1248, and 1298 to 1300 in the Cas9 reference sequence, or their corresponding amino acid positions. In one embodiment, the deaminase is inserted at an amino acid position between amino acids numbered 768 to 769, 791 to 792, 792 to 793, 1015 to 1016, 1022 to 1023, 1026 to 1027, 1029 to 1030, 1040 to 1041, 1052 to 1053, 1054 to 1055, 1067 to 1068, 1068 to 1069, 1247 to 1248, or 1248 to 1249 in the Cas9 reference sequence, or their corresponding amino acid positions. In one embodiment, the deaminase is inserted at an amino acid position between amino acids numbered 768 to 769, 792 to 793, 1022 to 1023, 1026 to 1027, 1040 to 1041, 1068 to 1069, or 1247 to 1248 in the Cas9 reference sequence, or their corresponding amino acid positions. In one embodiment, the deaminase is inserted at an amino acid position between amino acids numbered 1016 to 1017, 1023 to 1024, 1029 to 1030, 1040 to 1041, 1069 to 1070, or 1247 to 1248 in the Cas9 reference sequence, or their corresponding amino acid positions. In one embodiment, the adenosine deaminase variant is inserted into the Cas9 polypeptide at the locus identified in Table 13A. In one embodiment, the N-terminal fragment comprises amino acid residues 1 to 529, 538 to 568, 580 to 685, 692 to 942, 948 to 1001, 1026 to 1051, 1078 to 1231, and / or 1248 to 1297 of the Cas9 reference sequence, or their corresponding residues. In one embodiment, the C-terminal fragment comprises amino acid residues 1301 to 1368, 1248 to 1297, 1078 to 1231, 1026 to 1051, 948 to 1001, 692 to 942, 580 to 685, and / or 538 to 568 of the Cas9 reference sequence, or their corresponding residues.

[0046] In another embodiment, the Cas9 polypeptide is a modified Cas9 and is specific for an altered PAM. In one embodiment, the Cas9 polypeptide is a nickase or wherein the Cas9 polypeptide is nuclease-inactive. In one embodiment, the Cas9 polypeptide is a modified SpCas9 polypeptide. In one embodiment, the modified SpCas9 polypeptide comprises the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and is specific for the altered PAM 5'-NGC-3'.

[0047] In some embodiments, the adenosine deaminase variant is inserted within the Cas12 polypeptide. In one embodiment, the Cas12 polypeptide is Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, or Cas12i. In one embodiment, the adenosine deaminase variant is inserted between amino acid positions: a) 153 to 154, 255 to 256, 306 to 307, 980 to 981, 1019 to 1020, 534 to 535, 604 to 605, or 344 to 345 of BhCas12b or the corresponding amino acid residues of Cas12a, Cas12c, Cas2d, Cas12e, Cas12g, Cas12h, or Cas12i; b) 147 and 148, 248 and 249, 299 and 300, 991 and 992, or 1031 and 1032 of BvCas12b or the corresponding amino acid residues of Cas12a, Cas12c, Cas2d, Cas12e, Cas12g, Cas12h, or Cas12i; or c) 157 and 158, 258 and 259, 310 and 311, 1008 and 1009, or 1044 and 1045 of AaCas12b or the corresponding amino acid residues of Cas12a, Cas12c, Cas2d, Cas12e, Cas12g, Cas12h, or Cas12i. In one embodiment, the deaminase variant is inserted within the Cas12 polypeptide at the locus identified in Table 13B. In one embodiment, the Cas12 polypeptide is Cas12b. In one embodiment, the Cas12 polypeptide comprises a BhCas12b domain, a BvCas12b domain, or an AACas12b domain.

[0048] On the one hand, the present invention provides modified immune cells according to any aspect described herein. In one embodiment, the immune cells are T cells. In one embodiment, the immune cells express a chimeric antigen receptor. In one embodiment, the method comprises administering an effective amount of the modified immune cells according to any aspect described herein. On the one hand, the present invention provides a pharmaceutical composition comprising an effective amount of the modified immune cells according to any aspect described herein in a pharmaceutically acceptable excipient. On the other hand, the present invention provides a kit comprising the modified immune cells according to any aspect described herein.

[0049] On the one hand, the present invention provides a base editor system comprising a polynucleotide programmable DNA binding domain and at least one base editor domain, the base editor domain comprising

[0050] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEI

[0051] An adenosine deaminase variant with an alteration at amino acid position 82 or 166 of MALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD, and two or more guide RNAs targeting the nucleobase editor polypeptide to effect an alteration of a nucleic acid molecule encoding at least one polypeptide selected from the group consisting of T cell receptor alpha constant (TRAC), beta-2 microglobulin (B2M), programmed cell death 1 (PD1), cluster of differentiation 7 (CD7), cluster of differentiation 5 (CD5), cluster of differentiation 33 (CD33), cluster of differentiation 123 (CD123), Cbl proto-oncogene B (CBLB), and class II major histocompatibility complex transactivator (CIITA) polypeptides. In some embodiments, the adenosine deaminase variant comprises the alteration V82S and / or the alteration T166R. In some embodiments, the adenosine deaminase variant further comprises one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, and Q154R. In some embodiments, the base editor domain is an adenosine deaminase heterodimer comprising a wild-type adenosine deaminase domain and an adenosine deaminase variant. In some embodiments, the adenosine deaminase variant is a truncated TadA*8 that lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues relative to full-length TadA*8. In some embodiments, the adenosine deaminase variant lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C-terminal amino acid residues relative to full-length TadA8. In some embodiments, the polynucleotide programmable DNA binding domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In some embodiments, the polynucleotide programmable DNA binding domain is a variant of SpCas9 that has an altered protospacer adjacent motif (PAM) specificity or is specific for a non-G PAM. In some embodiments, the polynucleotide programmable DNA binding domain is a nuclease-inactive Cas9. In some embodiments, the polynucleotide programmable DNA binding domain is a Cas9 nickase.

[0052] On the one hand, the present invention provides a base editor system comprising two or more guide RNAs and a fusion protein, the fusion protein comprising a polynucleotide programmable DNA binding domain comprising the following sequences:

[0053]

[0054] wherein the bold sequences represent sequences derived from Cas9, the italic sequences represent linker sequences, the underlined sequences represent dual nuclear localization sequences, and at least one base editor domain, the base editor domain comprising an adenosine deaminase variant having an alteration at amino acid position 82 or 166 of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD and two or more guide RNAs targeting the nuclear base editor polypeptide to effect alteration of a nucleic acid molecule encoding at least one polypeptide selected from the group consisting of T cell receptor alpha constant (TRAC), beta-2 microglobulin (B2M), programmed cell death 1 (PD1), cluster of differentiation 7 (CD7), cluster of differentiation 5 (CD5), cluster of differentiation 33 (CD33), cluster of differentiation 123 (CD123), Cbl proto-oncogene B (CBLB), and class II major histocompatibility complex transactivator (CIITA) polypeptides.

[0055] On the one hand, a cell comprising any of the base editor systems described above is provided. Any cell is a human cell or a mammalian cell. In some embodiments, the cell is ex vivo, in vivo, or in vitro.

[0056] The descriptions and examples herein detail embodiments of the present disclosure. It should be understood that the present disclosure is not limited to the specific embodiments described herein and may thus vary. Those skilled in the art will recognize that there are various changes and modifications to the present disclosure, and such changes and modifications are included within its scope.

[0057] Unless otherwise indicated, the practice of some embodiments disclosed herein employs conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which are within the skill of the art. See, e.g., Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F.M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B.D. Hames and G.R. Taylor eds (1995)), Harlow and Lane, eds (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)).

[0058] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0059] Although various features of the present disclosure may be described in the context of a single embodiment, these features may also be provided singly or in any suitable combination. Conversely, although the present disclosure may be described herein in the context of separate embodiments for clarity, the present disclosure may also be implemented in a single embodiment. The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0060] The features of the present disclosure are specifically set forth in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the present disclosure are utilized, and in view of the drawings described below.

[0061] Definitions

[0062] The following definitions supplement those in the art and are specific to the current application and are not attributed to any related or unrelated cases, e.g., any co-owned patents or applications. Although any methods and materials similar or equivalent to those described herein may be used in testing the practice of the present disclosure, the preferred materials and methods are described herein. Accordingly, the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide one of ordinary skill in the art with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991).

[0064] In this application, unless otherwise specifically stated, the use of the singular includes the plural. It must be noted that the singular forms “a,” “an,” and “the” used in the specification include plural references unless the context clearly dictates otherwise. In this application, unless otherwise indicated, the use of “or” means “and / or” and is understood to be inclusive. Further, the use of the term “including” and other forms such as “include,” “includes,” and “included” is not limiting.

[0065] As used in this specification and the claims, the terms "comprising" (and any form of comprising, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of including, such as "includes" and "include") or "containing" (and any form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed in this specification may be implemented with respect to any method or combination of the present disclosure, and vice versa. Additionally, the compositions of the present disclosure may be used to implement the methods of the present disclosure.

[0066] The terms "about" or "approximately" mean within an acceptable error range of a particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the measurement system. For example, in accordance with the practice in the art, "about" can mean within one standard deviation or more than one standard deviation. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly for biological systems or processes, the term can mean within an order of magnitude, such as within 5-fold or 2-fold of the value. Where a particular value is described in the application and claims, unless otherwise stated, the meaning of the term "about" should be assumed to be within the acceptable error range of the particular value.

[0067] The ranges provided herein should be understood as shorthand for all values within the range. For example, a range of 1 to 50 is understood to include from 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0068] References in the specification to "some embodiments", "an embodiment", "one embodiment", or "other embodiments" refer to particular features, structures, or characteristics described in connection with the embodiments being included in at least some embodiments, but not necessarily all embodiments of the present disclosure.

[0069] "Adenosine deaminase" refers to a polypeptide or a fragment thereof that can catalyze the hydrolysis and deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolysis and deamination of adenosine to inosine or the hydrolysis and deamination of deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolysis and deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism, such as bacteria.

[0070] In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*8. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase. For example, deaminase domains are described in International PCT Application Nos. PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), which are each incorporated herein by reference in their entirety.In addition, see Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017)), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788. doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.

[0071] The wild-type TadA (wt) adenosine deaminase has the following sequence (also referred to as the TadA reference sequence):

[0072] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD

[0073] In some embodiments, the adenosine deaminase comprises an alteration of the following sequence:

[0074]

[0075] (Also known as TadA*7.10).

[0076] In some embodiments, TadA*7.10 contains at least one alteration. In some embodiments, TadA*7.10 contains an alteration at amino acid 82 and / or 166. In certain embodiments, variants of the above sequence contain one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. The alteration Y123H is also referred to herein as H123H (the alteration H123Y in TadA*7.10 reverts back to Y123H (wt)). In other embodiments, variants of the TadA*7.10 sequence contain a combination of alterations selected from the group consisting of: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R.

[0077] In other embodiments, the present invention provides adenosine deaminase variants that lack, for example, TadA*8, which contain a C-terminal deletion starting from residue 149, 150, 151, 152, 153, 154, 155, 156, or 157, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a TadA (e.g., TadA*8) monomer that contains one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA 7. In other embodiments, the adenosine deaminase variant is a monomer that contains a combination of alterations selected from the group consisting of: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA 7.

[0078] In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*8), each domain having one or more of the following alterations Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*8), each domain having a combination of alterations selected from the group consisting of: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA 7.

[0079] In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type TadA adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8), which comprises one or more of the following alterations Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type TadA adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8), which comprises a combination of alterations selected from the group consisting of: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA 7.

[0080] In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8), which comprises one or more of the following alterations Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain (e.g., TadA*8), which comprises a combination of alterations selected from: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R or I76Y + V82S + Y123H + Y147R + Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA 7.

[0081] In one embodiment, the adenosine deaminase is TadA*8, which comprises or consists essentially of the following sequence or a fragment thereof having adenosine deaminase activity:

[0082] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD.

[0083] In some embodiments, the TadA*8 is truncated. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length TadA*8. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length TadA*8. In some embodiments, the adenosine deaminase variant is the full-length TadA*8.

[0084] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from one of the following:

[0085] Staphylococcus aureus (S. aureus) TadA:

[0086] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN

[0087] Bacillus subtilis (B. subtilis) TadA:

[0088] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE

[0089] Salmonella typhimurium (S. typhimurium) TadA:

[0090] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV

[0091] Shewanella putrefaciens TadA:

[0092] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE

[0093] Haemophilus influenzae F3031 TadA:

[0094] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTAHAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK

[0095] Caulobacter crescentus TadA:

[0096] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI

[0097] Geobacter sulfurreducens TadA:

[0098] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP

[0099] TadA*7.10

[0100] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD

[0101] "Administering" as used herein refers to providing to a patient or subject one or more of the compositions described herein. For example but not limited to, administering a composition, such as by injection, can be performed by intravenous (iv) injection, subcutaneous (sc) injection, intradermal (id) injection, intraperitoneal (ip) injection, or intramuscular (im) injection. One or more such routes can be employed. Parenteral administration can be, for example, by bolus injection or by infusion over time. Alternatively or concurrently, administration can be by the oral route.

[0102] "Agent" refers to any small molecule compound, antibody, nucleic acid molecule, or polypeptide, or fragment thereof.

[0103] As used herein, "Allogeneic" means that cells of the same species are genetically distinct from the cells being compared.

[0104] "Alter" refers to a change (e.g., an increase or decrease) in the structure, expression level, or activity of a gene or polypeptide, as detected by standard methods known in the art such as those described herein. As used herein, an alter includes a change in the polynucleotide or polypeptide sequence or a change in the expression level, such as a 25% change, 40% change, 50% change, or greater.

[0105] "Ameliorate" means to reduce, inhibit, attenuate, mitigate, prevent, or stabilize the development or progression of a disease.

[0106] "Analog" refers to a molecule that is not identical but has similar functions or structural features. For example, a polynucleotide or polypeptide analog retains the biological activity of the corresponding naturally occurring polynucleotide or polypeptide, while having certain modifications that enhance the function of the analog relative to the naturally occurring polynucleotide or polypeptide. Such modifications can increase the affinity, efficiency, specificity, protease or nuclease resistance, membrane permeability, and / or half-life of the analog, without altering, for example, ligand binding. An analog can include unnatural nucleotides or amino acids.

[0107] "Anti-neoplasia activity" refers to preventing or inhibiting the maturation and / or proliferation of a tumor.

[0108] As used herein, "autologous" refers to cells from the same subject.

[0109] "Base editor (BE)" or "nucleobase editor (NBE)" refers to a reagent that binds to a polynucleotide and has nucleobase modification activity. In various embodiments, the base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) and a nucleic acid programmable nucleotide-binding domain that binds to a guide polynucleotide (e.g., a guide RNA). In various embodiments, the reagent is a biomolecular complex comprising a protein domain having base editing activity, i.e., capable of modifying a base (e.g., in DNA) within a nucleic acid molecule (e.g., A, T, C, G, or U). In some embodiments, the nucleic acid programmable DNA-binding domain is fused or linked to a deaminase domain.. In one embodiment, the reagent is a fusion protein comprising a domain having base editing activity. In another embodiment, the protein domain having base editing activity is linked to a guide RNA (e.g., via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to the deaminase). In some embodiments, the domain having base editing activity is capable of deaminating a base within a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases within a DNA molecule. In some embodiments, the base editor is capable of deaminating one or more bases within a DNA molecule. In some embodiments, the base editor is an adenosine base editor (ABE).

[0110] In some embodiments, base editors are generated by cloning an adenosine deaminase variant (e.g., TadA*8) into a scaffold comprising a circularly permuted Cas9 (e.g., spCAS9 or saCAS9) and a bipartite nuclear localization sequence (e.g., ABE8). Circularly permuted Cas9s are known in the art and described, for example, in Oakes et al., Cell 176, 254–267, 2019. Exemplary circular arrangements are as follows, where bold sequences represent sequences derived from Cas9, italic sequences represent linker sequences, and underlined sequences represent bipartite nuclear localization sequences.

[0111] CP5 (with MSP “NGC = Pam variant conventional Cas9-like NGG with mutations” PID = Protein Interaction Domain and “D10A” nickase):

[0112]

[0113]

[0114] In some embodiments, the ABE8 is selected from the base editors of Tables 8, 9, 10, or 11 below. In some embodiments, ABE8 contains an adenosine deaminase variant evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is the TadA*8 variant as described in Table 9 below. In some embodiments, the adenosine deaminase variant is a TadA*7.10 variant (e.g., TadA*8) comprising one or more alterations selected from Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In various embodiments, ABE8 comprises a TadA*7.10 variant (e.g., TadA*8) having a combination of alterations selected from the following groups: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R. In some embodiments, ABE8 is a monomeric construct. In some embodiments, ABE8 is a heterodimeric construct. In some embodiments, the ABE8 comprises the sequence:

[0115] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD。

[0116] In some embodiments, the polynucleotide programmable DNA binding domain is a CRISPR-associated (e.g., Cas or Cpf1) enzyme. In some embodiments, the base editor is catalytically inactive Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is Cas9 nickase (nCas9) fused to a deaminase domain. Base editors are described in International PCT Application Nos. PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), each of which is incorporated herein by reference in its entirety. See also Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3: eaao4774 (2017), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788.doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.

[0117] For example, an adenine base editor (ABE) for use in the base editing compositions, systems, and methods described herein has a nucleic acid sequence (8,877 base pairs) (Addgene, Watertown, MA.; Komor NM et al., 2017, SciAdv., 30; 3(8):2017 Nov 23; 551(7681):464-471. doi:10.1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct; 36(9):843-846. doi:10.1038 / nbt.4172.) provided as follows. Also included are polynucleotide sequences having at least 95% or greater identity to the ABE nucleic acid sequence.

[0118]

[0119] "Base editing activity" refers to the chemical alteration of bases within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, such as converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity, such as converting a target A·T to C·G. In another embodiment, the base editing activity is cytidine deaminase activity, such as converting a target C·G to T·A, and adenosine or adenine deaminase activity, such as converting A·T to G·C. In some embodiments, base editing activity is evaluated by editing efficiency. Base editing efficiency can be measured by any suitable means, e.g., by Sanger sequencing or next-generation sequencing. In some embodiments, base editing efficiency is measured as the percentage of total sequencing reads with nucleobase conversions affected by the base editor, e.g., the percentage of total sequencing reads of target A·T base pairs that are converted to G·C base pairs. In some embodiments, when base editing is performed in a cell population, base editing efficiency is measured as the percentage of total cells with nucleobase conversions affected by the base editor.

[0120] The term "base editor system" refers to a system for editing the nucleobases of a target nucleotide sequence. In various embodiments, the base editor system comprises (1) a polynucleotide programmable nucleotide binding domain (e.g., Cas9); (2) a deaminase domain for deaminating the nucleobase (e.g., adenosine deaminase and / or cytidine deaminase); and (3) one or more guide polynucleotides (e.g., guide RNA). In some embodiments, the polynucleotide programmable acid binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor system is ABE8.

[0121] In some embodiments, a base editor system can include more than one base editing component. For example, a base editor system can include more than one deaminase. In some embodiments, a base editor system can include one or more adenosine deaminases. In some embodiments, a single guide polynucleotide can be used to target different deaminases to a target nucleic acid sequence. In some embodiments, a pair of guide polynucleotides can be used to target different deaminases to a target nucleic acid sequence.

[0122] The deaminase domain and the polynucleotide programmable nucleotide-binding component of the base editor system can be associated with each other covalently or non-covalently, or any combination of their association and interaction. For example, in some embodiments, the deaminase domain can be targeted to a target nucleotide sequence by a polynucleotide programmable nucleotide-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain can be fused or linked to the deaminase domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain can target the deaminase domain to the target nucleotide sequence by non-covalent interaction or association with the deaminase domain. For example, in some embodiments, the deaminase domain can comprise an additional heterologous moiety or domain that is capable of interacting, associating, or forming a complex with an additional heterologous moiety or domain that is part of the polynucleotide programmable nucleotide-binding domain. In some embodiments, the additional heterologous moiety may be capable of binding, interacting, associating, or forming a complex with a polypeptide. In some embodiments, the additional heterologous moiety may be capable of binding, interacting, associating, or forming a complex with a polynucleotide. In some embodiments, the additional heterologous moiety may be capable of binding a guide polynucleotide. In some embodiments, the additional heterologous moiety may be capable of binding a polypeptide linker. In some embodiments, the additional heterologous moiety may be capable of binding a polynucleotide linker. The additional heterologous moiety can be a protein domain. In some embodiments, the additional heterologous moiety can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.

[0123] The base editor system can further comprise a guide polynucleotide component. It should be understood that the components of the base editor system can be associated with each other by covalent bonds, non-covalent interactions, or any combination of their associations and interactions. In some embodiments, the deaminase domain can be targeted to a target nucleotide sequence by a guide polynucleotide. For example, in some embodiments, the deaminase domain can comprise additional heterologous portions or domains (e.g., polynucleotide binding domains, such as RNA or DNA binding proteins) that are capable of interacting with, associating with, or forming a complex with a portion or segment of the guide polynucleotide (e.g., a polynucleotide motif). In some embodiments, the additional heterologous portion or domain (e.g., a polynucleotide binding domain, such as an RNA or DNA binding protein) can be fused or linked to the deaminase domain. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, associating with, or forming a complex with a polypeptide. In some embodiments, the additional heterologous portion may be capable of binding to, interacting with, associating with, or forming a complex with a polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a guide polynucleotide. In some embodiments, the additional heterologous portion may be capable of binding to a polypeptide linker. In some embodiments, the additional heterologous portion may be capable of binding to a polynucleotide linker. The additional heterologous portion can be a protein domain. In some embodiments, the additional heterologous portion can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku binding motif and Ku protein, a telomerase Sm7 binding motif and Sm7 protein, or an RNA recognition motif.

[0124] In some embodiments, the base editor system may further comprise an inhibitor of a base excision repair (BER) component. It should be understood that the components of the base editor system may be associated with each other by covalent bonds, non-covalent interactions, or any combination of their associations and interactions. The inhibitor of the BER component may include a BER inhibitor. In some embodiments, the inhibitor of BER may be a uracil DNA glycosylase inhibitor (UGI). In some embodiments, the inhibitor of BER may be an inosine BER glycosylase inhibitor. In some embodiments, the inhibitor of BER may target the target nucleotide sequence through the polynucleotide programmable nucleotide binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain may be fused or linked to the inhibitor of BER. In some embodiments, the polynucleotide programmable nucleotide binding domain may be fused or linked to a deaminase domain and an inhibitor of BER. In some embodiments, the polynucleotide programmable nucleotide binding domain may target the inhibitor of BER to the target nucleotide sequence by non-covalent interaction or association with the inhibitor of BER. For example, in some embodiments, the inhibitor of BER may comprise an additional heterologous moiety or domain that is capable of interacting, associating, or forming a complex with an additional heterologous moiety or domain that is part of the polynucleotide programmable nucleotide binding domain.

[0125] In some embodiments, the inhibitor of BER may target the target nucleotide sequence through the guide polynucleotide. For example, in some embodiments, the inhibitor of BER may comprise an additional heterologous moiety or domain (e.g., a polynucleotide binding domain, such as an RNA or DNA binding protein) that is capable of interacting, associating, or forming a complex with a part or segment (e.g., a polynucleotide motif) of the guide polynucleotide. In some embodiments, the additional heterologous moiety or domain (e.g., a polynucleotide binding domain, such as an RNA or DNA binding protein) of the guide polynucleotide may be fused or linked to the inhibitor of BER. In some embodiments, the additional heterologous moiety may be capable of binding to a polynucleotide, interacting, associating, or forming a complex. In some embodiments, the additional heterologous moiety may be capable of binding to the guide polynucleotide. In some embodiments, the additional heterologous moiety may be capable of binding to a polypeptide linker. In some embodiments, the additional heterologous moiety may be capable of binding to a polynucleotide linker. The additional heterologous moiety may be a protein domain. In some embodiments, the additional heterologous moiety may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.

[0126] "B-cell maturation antigen, or tumor necrosis factor receptor superfamily member 17 polypeptide, (BCMA)" refers to a protein identical to NCBI accession number NP_001183 or a fragment thereof expressed on mature B lymphocytes having at least about 85% amino acid sequence. Exemplary BCMA polypeptide sequences are provided below.

[0127] >NP_001183.2 tumor necrosis factor receptor superfamily member 17 [Homo sapiens]

[0128] MLQMAGQCSQNEYFDSLLHACIPCQLRCSSNTPPLTCQRYCNASVTNSVKGTNAILWTCLGLSLIISLAVFVLMFLLRKINSEPLKDEFKNTGSGLLGMANIDLEKSRTGDEIILPRGLEYTVEECTCEDCIKSKPKVDSDHCFPLPAMEEGATILVTTKTNDYCKSLPAALSATEIEKSISAR

[0129] This antigen can be targeted for the treatment of relapsed or refractory multiple myeloma and other hematological malignancies.

[0130] "B-cell maturation antigen, or tumor necrosis factor receptor superfamily member 17, (BCMA) polynucleotide" refers to a nucleic acid molecule encoding a BCMA polypeptide. The BCMA gene encodes a cell surface receptor that recognizes the B-cell activating factor. Exemplary B2M polypeptide sequences are provided below.

[0131] > Homo sapiens TNF receptor superfamily member 17 (TNFRSF17),mRNA AAGACTCAAACTTAGAAACTTGAATTAGATGTGGTATTCAAATCCTTAGCTGCCGCGAAGACACAGACAGCCCCCGTAAGAACCCACGAAGCAGGCGAAGTTCATTGTTCTCAACATTCTAGCTGCTCTTGCTGCATTTGCTCTGGAATTCTTGTAGAGATATTACTTGTCCTTCCAGGCTGTTCTTTCTGTAGCTCCCTTGTTTTCTTTTTGTGATCATGTTGCAGATGGCTGGGCAGTGCTCCCAAAATGAATATTTTGACAGTTTGTTGCATGCTTGCATACCTTGTCAACTTCGATGTTCTTCTAATACTCCTCCTCTAACATGTCAGCGTTATTGTAATGCAAGTGTGACCAATTCAGTGAAAGGAACGAATGCGATTCTCTGGACCTGTTTGGGACTGAGCTTAATAATTTCTTTGGCAGTTTTCGTGCTAATGTTTTTGCTAAGGAAGATAAACTCTGAACCATTAAAGGACGAGTTTAAAAACACAGGATCAGGTCTCCTGGGCATGGCTAACATTGACCTGGAAAAGAGCAGGACTGGTGATGAAATTATTCTTCCGAGAGGCCTCGAGTACACGGTGGAAGAATGCACCTGTGAAGACTGCATCAAGAGCAAACCGAAGGTCGACTCTGACCATTGCTTTCCACTCCCAGCTATGGAGGAAGGCGCAACCATTCTTGTCACCACGAAAACGAATGACTATTGCAAGAGCCTGCCAGCTGCTTTGAGTGCTACGGAGATAGAGAAATCAATTTCTGCTAGGTAATTAACCATTTCGACTCGAGCAGTGCCACTTTAAAAATCTTTTGTCAGAATAGATGATGTGTCAGATCTCTTTAGGATGACTGTATTTTTCAGTTGCCGATACAGCTTTTTGTCCTCTAACTGTGGAAACTCTTTATGTTAGATATATTTCTCTAGGTTACTGTTGGGAGCTTAATGGTAGAAACTTCCTTGGTTTCATGATTAAACTCTTTTTTTTCCTGA,

[0132] The “β-2 microglobulin (B2M) polypeptide” refers to a protein having at least about 85% amino acid sequence identity with UniProt accession number P61769 or a fragment thereof and having immunomodulatory activity. Exemplary B2M polypeptide sequences are provided below.

[0133] >sp|P61769|B2MG_HUMAN Beta-2-microglobulin OS=Homo sapiens OX=9606 GN=B2M PE=1 SV=1

[0134] MSRSVALAVLALLSLSGLEAIQRTPKIQVYSRHPAENGKSNFLNCYVSGFHPSDIEVDLL

[0135] KNGERIEKVEHSDLSFSKDWSFYLLYYTEFTPTEKDEYACRVNHVTLSQPKIVKWDRDM

[0136] The “β-2-microglobulin (B2M) polynucleotide” refers to a nucleic acid molecule encoding a B2M polypeptide. The β-2-microglobulin gene encodes a serum protein associated with the major histocompatibility complex. B2M is involved in the non-self recognition of host CD8+ T cells. Exemplary B2M polynucleotide sequences are provided below.

[0137] >DQ217933.1 Homo sapiens beta-2-microglobulin (B2M) gene, complete cds

[0138]

[0139] The term “Cas9” or “Cas9 domain” refers to an RNA-guided nuclease that comprises a Cas9 protein or a fragment thereof (e.g., a protein that contains the DNA cleavage domain of Cas9 that is active, inactive, or partially active, and / or the binding domain for gRNA Cas9). The Cas9 nuclease is sometimes also referred to as the Casnl nuclease or a CRISPR (clustered regularly interspaced short palindromic repeats)-associated nuclease. CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). The CRISPR cluster contains spacer sequences, sequences complementary to antecedent mobile elements, and target invading nucleic acids. The CRISPR cluster is transcribed and processed into CRISPR RNA (crRNA). In type II CRISPR systems, proper processing of pre-crRNA requires a trans-encoded small RNA (tracrRNA), the endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3 to assist in processing the pre-crRNA. Subsequently, the Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer sequence. The target strand that is not complementary to the crRNA is first cleaved endonucleolytically and then trimmed 3′-5′ exonucleolytically. In nature, DNA binding and cleavage generally require a protein and two RNAs. However, a single-guide RNA (“sgRNA,” or simply “gRNA”) can be engineered to incorporate aspects of the crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif within the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self.The Cas9 nuclease sequence and structure are well-known to those skilled in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012). The entire contents of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus. Based on the present disclosure, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire contents of which are incorporated herein by reference.

[0140] An exemplary Cas9 is Streptococcus pyogenes Cas9 (spCas9), the amino acid sequence of which is provided below:

[0141]

[0142]

[0143] (Single underline: HNH domain; double underline: RuvC domain)

[0144] The nuclease-inactivated Cas9 protein can be interchangeably referred to as the "dCas9" protein (for nuclease - "dead" Cas9) or catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with an inactive DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337:816 - 821 (2012); Qi et al., "Repurposing CRISPR as an RNA - Guided Platform for Sequence - Specific Control of Gene Expression" (2013) Cell. 28; 152(5):1173 - 83, the entire contents of which are incorporated herein by reference). For example, the DNA cleavage domain of Cas9 is known to include two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non - complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337:816 - 821 (2012); Qi et al., Cell. 28:152(5):1173 - 83 (2013)). In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase, referred to as the "nCas9" protein (for "nickase" Cas9). In some embodiments, proteins comprising fragments of Cas9 are provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA - binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, proteins comprising Cas9 or fragments thereof are referred to as "Cas9 variants". Cas9 variants are homologous to Cas9 or fragments thereof. For example, Cas9 variants are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 96% identical, at least about 97% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical or at least about 99.9% identical to wild - type Cas9. In some embodiments, compared to wild - type Cas9, Cas9 variants can have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes.In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., the gRNA binding domain or the DNA cleavage domain) such that the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% the amino acid length of the corresponding wild-type Cas9.

[0145] In some embodiments, the fragment is at least 100 amino acids in length. In some embodiments, the fragment is at least 100, 150, 200, 250, 300, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, or 1300 amino acids in length.

[0146] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows).

[0147]

[0148]

[0149] (Single underline: HNH domain; double underline: RuvC domain)

[0150] In some embodiments, wild-type Cas9 corresponds to or comprises the following nucleotide and / or amino acid sequences:

[0151]

[0152]

[0153] (Single underline: HNH domain; double underline: RuvC domain)

[0154] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_002737.2 (nucleotide sequence as follows); and Uniprot reference sequence: Q99ZW2 (amino acid sequence as follows).

[0155]

[0156]

[0157] (SEQ ID NO:1. Single underline: HNH domain; double underline: RuvC domain)

[0158] In some embodiments, Cas9 refers to Cas9 from the following: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Refs: NC_021284.1); Prevotella intermedia (NCBI Refs: NC_017861.1); Spiroplasma taiwanense, China (NCBI Refs: NC_021846.1); Streptococcus iniae (NCBI Refs: NC_021314.1); Belliella baltica (NCBI Refs: NC_018010.1); Psychroflexus torquisI (NCBI Refs: NC_018721.1); Streptococcus thermophilus (NCBI Refs: YP_820832.1); Listeria innocua (NCBI Refs: NP_472073.1); Campylobacter jejuni (NCBI Refs: YP_002344900.1); Neisseria meningitidis (NCBI Refs: YP_002342100.1) or Cas9 from any other organism.

[0159] In some embodiments, Cas9 is from Neisseria meningitidis (Nme). In some embodiments, the Cas9 is Nme1, Nme2, or Nme3. In some examples, the PAM interaction domains of Nme1, Nme2, or Nme3 are N4GAT, N4CC, and N4CAAA, respectively (see, e.g., Edraki, A. et al., A Compact, High-Accuracy Cas9 with a Dinucleotide PAM for In Vivo Genome Editing, Molecular Cell (2018)). The exemplary Neisseria meningitidis Cas9 protein Nme1Cas9, (NCBI Reference: WP_002235162.1; type II CRISPR RNA-guided endonuclease Cas9) has the following amino acid sequence:

[0160]

[0161] Another exemplary Neisseria meningitidis Cas9 protein Nme2Cas9, (NCBI Reference: WP_002230835; type II CRISPR RNA-guided endonuclease Cas9) has the following amino acid sequence:

[0162]

[0163] In some embodiments, dCas9 corresponds to or partially or fully comprises a Cas9 amino acid sequence having one or more mutations that inactivate the Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain comprises the D10A and H840A mutations or corresponding mutations in another Cas9. In some examples, dCas9 comprises the amino acid sequence of dCas9 (D10A and H840A):

[0164]

[0165]

[0166] (Single underline: HNH domain; double underline: RuvC domain).

[0167] In some embodiments, the Cas9 domain comprises the D10A mutation, and the residue at position 840 remains histidine at the corresponding position in the amino acid sequence provided above or in any amino acid sequence provided herein

[0168] In other embodiments, dCas9 variants with mutations other than D10A and H840A are provided, which for example result in nuclease-inactivated Cas9 (dCas9). For example, such mutations include other amino acid substitutions at D10 and H840, or other substitutions within the Cas9 nuclease domain (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, there are provided that are shorter or longer by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more.

[0169] In some embodiments, the Cas9 fusion proteins provided herein comprise the full-length amino acid sequence of a Cas9 protein, such as one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not comprise the full-length Cas9 sequence, but only one or more of its fragments. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and other suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art.

[0170] It should be understood that additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including their variants and homologs, are within the scope of the present disclosure. Exemplary Cas9 proteins include, but are not limited to, those provided below. In some embodiments, the Cas9 protein is nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein is a Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is a Cas9 with nuclease activity.

[0171] Exemplary catalytically inactive Cas9 (dCas9):

[0172]

[0173] Exemplary catalytically active Cas9 nickase (nCas9):

[0174]

[0175] Cas9 with exemplary catalytic activity:

[0176]

[0177] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., Nanoarchaeum), which constitutes the domain and kingdom of single-celled prokaryotic microorganisms. In some embodiments, Cas9 refers to CasX or CasY, which have been described in, for example, Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the entire content of which is incorporated herein by reference. Using genome-resolved metagenomics, many CRISPR-Cas systems have been identified, including Cas9 first reported in the archaea domain. This divergent Cas9 protein was found in the little-studied Nanoarchaeum as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were found, which are among the most compact systems discovered to date. In some embodiments, Cas9 refers to CasX, or a variant of CasX. In some embodiments, Cas9 refers to CasY, or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins can be used as nucleic acid programmable DNA-binding proteins (napDNAbps) and are within the scope of the present disclosure.

[0178] In certain embodiments, the napDNAbps useful in the methods of the present invention include circular permutants known in the art and described, for example, by Oakes et al., Cell 176, 254–267, 2019. Exemplary circular permutations are as follows, where the bold sequences represent sequences derived from Cas9, the italic sequences represent linker sequences, and the underlined sequences represent dual nuclear localization,

[0179] CP5 (with MSP “NGC = conventional Cas9-like NGG with mutated Pam variant” PID = protein interaction domain and “D10A” nickase):

[0180]

[0181]

[0182] Non-limiting examples of polynucleotide programmable nucleotide binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction nucleases, meganucleases, TAL nucleases (TALENs), and zinc finger nucleases (ZFNs).

[0183] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) or any fusion protein provided herein can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp is a naturally occurring CasX or CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to any CasX or CasY protein described herein. It should be understood that Cas12b / C2c1, CasX, and CasY from other bacterial species can also be used according to the present disclosure.

[0184] Cas12b / C2c1(uniprot.org / uniprot / T0D7A2#2)

[0185]

[0186] CasX (uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53)

[0187] >tr|F0NN87|F0NN87_SULIH CRISPR-associated Casx protein OS=Sulfolobus islandicus (strain HVE10 / 4) GN=SiH_0402 PE=4 SV=1

[0188] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG

[0189] >tr|F0NH53|F0NH53_SULIR CRISPRassociated protein, Casx OS=Sulfolobus islandicus (strain REY15A) GN=SiRe_0771 PE=4 SV=1

[0190] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG

[0191] Delta Proteobacterium CasX

[0192] MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDAYNEVIARVRMWVNLNLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA

[0193] CasY (ncbi.nlm.nih.gov / protein / APG80656.1)

[0194] >APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group]

[0195]

[0196]

[0197] An amino acid sequence having at least 85% or higher identity with the BhCas12b amino acid sequence can also be used in the methods of the present disclosure.

[0198] The "Cbl proto-oncogene B (CBLB) polypeptide" refers to a protein having at least about 85% amino acid sequence identity with GenBank accession number ABC86700.1 or a fragment thereof involved in the regulation of immune response. Exemplary CBLB polypeptide sequences are provided below.

[0199] >ABC86700.1 CBL-B [Homo sapiens]

[0200] MANSMNGRNPGGRGGNPRKGRILGIIDAIQDAVGPPKQAAADRRTVEKTWKLMDKVVRLCQNPKLQLKNSPPYILDILPDTYQHLRLILSKYDDNQKLAQLSENEYFKIYIDSLMKKSKRAIRLFKEGKERMYEEQSQDRRNLTKLSLIFSHMLAEIKAIFPNGQFQGDNFRITKADAAEFWRKFFGDKTIVPWKVFRQCLHEVHQISSGLEAMALKSTIDLTCNDYISVFEFDIFTRLFQPWGSILRNWNFLAVTHPGYMAFLTYDEVKARLQKYSTKPGSYIFRLSCTRLGQWAIGYVTGDGNILQTIPHNKPLFQALIDGSREGFYLYPDGRSYNPDLTGLCEPTPHDHIKVTQEQYELYCEMGSTFQLCKICAENDKDVKIEPCGHLMCTSCLTAWQESDGQGCPFCRCEIKGTEPIIVDPFDPRDEGSRCCSIIDPFGMPMLDLDDDDDREESLMMNRLANVRKCTDRQNSPVTSPGSSPLAQRRKPQPDPLQIPHLSLPPVPPRLDLIQKGIVRSPCGSPTGSPKSSPCMVRKQDKPLPAPPPPLRDPPPPPPERPPPIPPDNRLSRHIHHVESVPSRDPPMPLEAWCPRDVFGTNQLVGCRLLGEGSPKPGITASSNVNGRHSRVGSDPVLMRKHRRHDLPLEGAKVFSNGHLGSEEYDVPPRLSPPPPVTTLLPSIKCTGPLANSLSEKTRDPVEEDDDEYKIPSSHPVSLNSQPSHCHNVKPPVRSCDNGHCMLNGTHGPSSEKKSNIPDLSIYLKGDVFDSASDPVPLPPARPPTRDNPKHGSSLNRTPSDYDLLIPPLGEDAFDALPPSLPPPPPPARHSLIEHSKPPGSSSRPSSGQDLFLLPSDPFVDLASGQVPLPPARRLPGENVKTNRTSQDYDQLPSCSDGSQAPARPPKPRPRRTAPEIHHRKPHGPEAALENVDAKIAKLMGEGYAFEEVKRALEIAQNNVEVARSILREFAFPPPVSPRLNL

[0201] "Cbl proto-oncogene B (CBLB) polynucleotide" refers to a nucleic acid molecule encoding a CBLB polypeptide. The CBLB gene encodes an E3 ubiquitin ligase. Exemplary CBLB nucleic acid sequences are provided below.

[0202] >DQ349203.1 Homo sapiens CBL-B mRNA, complete cds

[0203]

[0204] "Chimeric antigen receptor" or "CAR" refers to a synthetic receptor that includes an extracellular antigen-binding domain, a transmembrane domain, and an intracellular signaling domain that confers antigen specificity to immune cells.

[0205] "Class II major histocompatibility complex transactivator (CIITA) polypeptide" refers to a protein that has at least about 85% amino acid sequence identity with NCBI Reference Sequence: NP_000237.2 or a fragment thereof, or a protein that functions as a transcriptional coactivator. Exemplary CIITA polypeptide sequences are provided below.

[0206]

[0207] "Class II major histocompatibility complex transactivator (CIITA) polynucleotide" refers to a nucleic acid molecule that encodes a CIITA polypeptide. Exemplary CIITA nucleic acid sequences are provided below.

[0208]

[0209]

[0210]

[0211] "Cluster of differentiation 7 (CD7) polypeptide" refers to a protein that has at least about 85% amino acid sequence identity with NCBI Reference Sequence: NP_006128.1 or a fragment thereof, and is involved in T cell and T cell / B cell interactions. Exemplary CD7 polypeptide sequences are provided below.

[0212]

[0213] "Cluster of differentiation 7 (CD7) polynucleotide" refers to a nucleic acid molecule that encodes a CD7 polypeptide. The CD7 gene encodes a transmembrane protein. Exemplary CD7 nucleic acid sequences are provided below.

[0214]

[0215] "Cluster of differentiation 5 (CD5) polypeptide" refers to a protein that has at least about 85% amino acid sequence identity with NCBI Reference Sequence: NP_001333385.1 or a fragment thereof, and is expressed on the surface of T cells. Exemplary CD5 polypeptide sequences are provided below.

[0216]

[0217] "Cluster of differentiation 5 (CD5) polynucleotide" refers to a nucleic acid molecule that encodes a CD5 polypeptide. The CD5 gene encodes a transmembrane protein. Exemplary CD5 nucleic acid sequences are provided below.

[0218]

[0219]

[0220]

[0221] The term "conservative amino acid substitution" or "conservative mutation" refers to the substitution of one amino acid by another amino acid having common properties. One functional way to define the common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between the corresponding proteins of homologous organisms (Schulz, G.E. and Schirmer, R.H., Principles of Protein Structure, Springer-Verlag, New York (1979)). Based on such an analysis, groups of amino acids can be defined in which the amino acids within a group preferentially exchange with each other and are thus most similar to each other in terms of their effect on the overall protein structure (Schulz, G.E. and Schirmer, R.H., supra). Non-limiting examples of conservative mutations include amino acid substitutions of amino acids such as lysine for arginine and vice versa, so that a positive charge can be maintained; glutamic acid for aspartic acid and vice versa to maintain a negative charge; serine for threonine so that a free -OH can be maintained; and glutamine for asparagine so that a free -NH2 can be maintained.

[0222] The term "coding sequence" or "protein coding sequence", as used interchangeably herein, refers to a polynucleotide fragment that encodes a protein. This region or sequence has a start codon near the 5'-end and a stop codon near the 3'-end. The coding sequence may also be referred to as an open reading frame.

[0223] "Cytotoxic T lymphocyte-associated protein 4 (CTLA-4) polypeptide" refers to a protein having at least about 85% sequence identity with NCBI accession number EAW70354.1 or a fragment thereof. Exemplary amino acid sequences are provided below:

[0224] >EAW70354.1 Cytotoxic T-lymphocyte-associated protein 4 [Homo sapiens] MACLGFQRHKAQLNLATRTWPCTLLFFLLFIPVFCKAMHVAQPAVVLASSRGIASFVCEYASPGKATEVRVTVLRQADSQVTEVCAATYMMGNELTFLDDSICTGTSSGNQVNLTIQGLRAMDTGLYICKVELMYPPPYYLGIGNGTQIYVIDPEPCPDSDFLLWILAAVSSGLFFYSFLLTAVSLSKMLKKRSPLTTGVYVKMPPTEPECEKQFQPYFIPIN

[0225] "Cytotoxic T-lymphocyte-associated protein 4 (CTLA-4) polynucleotide" refers to a nucleic acid molecule encoding a CTLA-4 polypeptide. The CTLA-4 gene encodes an immunoglobulin superfamily and encodes a protein that transmits an inhibitory signal to T cells. Exemplary CTLA-4 nucleic acid sequences are provided below.

[0226] >BC074842.2 Homo sapiens cytotoxic T-lymphocyte-associated protein 4, mRNA (cDNA clone MGC:104099 IMAGE:30915552), complete cds GACCTGAACACCGCTCCCATAAAGCCATGGCTTGCCTTGGATTTCAGCGGCACAAGGCTCAGCTGAACCTGGCTACCAGGACCTGGCCCTGCACTCTCCTGTTTTTTCTTCTCTTCATCCCTGTCTTCTGCAAAGCAATGCACGTGGCCCAGCCTGCTGTGGTACTGGCCAGCAGCCGAGGCATCGCCAGCTTTGTGTGTGAGTATGCATCTCCAGGCAAAGCCACTGAGGTCCGGGTGACAGTGCTTCGGCAGGCTGACAGCCAGGTGACTGAAGTCTGTGCGGCAACCTACATGATGGGGAATGAGTTGACCTTCCTAGATGATTCCATCTGCACGGGCACCTCCAGTGGAAATCAAGTGAACCTCACTATCCAAGGACTGAGGGCCATGGACACGGGACTCTACATCTGCAAGGTGGAGCTCATGTACCCACCGCCATACTACCTGGGCATAGGCAACGGAACCCAGATTTATGTAATTGATCCAGAACCGTGCCCAGATTCTGACTTCCTCCTCTGGATCCTTGCAGCAGTTAGTTCGGGGTTGTTTTTTTATAGCTTTCTCCTCACAGCTGTTTCTTTGAGCAAAATGCTAAAGAAAAGAAGCCCTCTTACAACAGGGGTCTATGTGAAAATGCCCCCAACAGAGCCAGAATGTGAAAAGCAATTTCAGCCTTATTTTATTCCCATCAATTGAGAAACCATTATGAAGAAGAGAGTCCATATTTCAATTTCCAAGAGCTGAGG

[0227] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is adenosine deaminase, which catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism, such as bacteria. In some embodiments, the adenosine deaminase is from bacteria, such as Escherichia coli, Staphylococcus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter.

[0228] In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant. In some embodiments, the TadA variant is TadA*8. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase. For example, deaminase domains are described in International PCT Application Nos. PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), which are each incorporated herein by reference in their entirety.In addition, see Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017)), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788.doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference...

[0229] "Detecting" refers to identifying the presence, absence, or amount of an analyte to be detected. In one embodiment, detecting a sequence alteration in a polynucleotide or polypeptide. In another embodiment, detecting the presence of an indel.

[0230] "Detectable label" refers to a composition that renders the latter detectable by spectroscopic, photochemical, biochemical, immunochemical, or chemical means when linked to a molecule of interest. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., commonly used in ELISA), biotin, digoxin, or haptens.

[0231] "Disease" refers to any disorder or condition that impairs or interferes with the normal function of a cell, tissue, or organ. In one embodiment, the disease is a tumor or cancer.

[0232] As used herein, the term "effective amount" refers to the amount of a bioactive agent sufficient to elicit a desired biological response. The effective amount of an active agent for practicing the present disclosure for treating a disease varies depending on the mode of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage regimen. Such amount is referred to as an "effective" amount. In one embodiment, the effective amount is the amount of a base editor of the present invention (e.g., a fusion protein comprising a programmable DNA-binding protein, a nucleobase editor, and a gRNA) sufficient to introduce a change in a gene of interest in a cell (e.g., in vitro or in vivo cells). In one embodiment, the effective amount is the amount of a base editor required to achieve a therapeutic effect (e.g., alleviating or controlling a disease or its symptoms or conditions). Such a therapeutic effect does not require sufficient to alter the gene of interest in all cells of the subject, tissue, or organ, but only about 1%, 5%, 10%, 25%, 50%, 75% or more of the gene of interest in the cells present in the subject, tissue, or organ.

[0233] As used herein, "Epitope" refers to an antigenic determinant. An epitope is a part of an antigen molecule that determines the specific antibody molecule that recognizes and binds to it through its structure.

[0234] "Fragment" refers to a portion of a polypeptide or nucleic acid molecule. This portion contains at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the full length of the reference nucleic acid molecule or polypeptide. Fragments can contain 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1000 nucleotides or amino acids.

[0235] "Graft-versus-host disease" (GVHD) refers to a pathological condition in which transplanted cells from a donor mount an immune response against the host cells.

[0236] "Guide RNA" or "gRNA" refers to a polynucleotide that can be specific to a target sequence and can form a complex with a polynucleotide programmable nucleotide-binding domain protein (such as Cas9 or Cpf1). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. The gRNA that exists as a single RNA molecule can be referred to as a single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNAs that exist as a single molecule or as a complex of two or more molecules. Generally, the gRNA that exists as a single RNA species contains two domains: (1) a domain that is homologous to the target nucleic acid (e.g., directs the binding of the Cas9 complex to the target); (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA provided in Jinek et al., Science 337:816-821 (2012), the entire content of which is incorporated herein by reference. Other examples of gRNAs (e.g., those that include domain 2) can be found in U.S. Provisional Patent Application, U.S.S.N. 61 / 874,682, filed September 6, 2013, entitled "Switchable Cas9 Nucleases and Uses Thereof", and U.S. Provisional Patent Application, U.S.S.N. 61 / 874,746, filed September 6, 2013, named "Delivery System For Functional Nucleases", the entire content of each of which is incorporated herein by reference in its entirety. In some embodiments, the gRNA contains two or more of domains (1) and (2) and can be referred to as an "extended gRNA". As described herein, the extended gRNA will bind two or more Cas9 proteins and bind to the target nucleic acid at two or more different regions. The gRNA contains a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site, providing sequence specificity for the nuclease:RNA complex. As will be understood by those skilled in the art, RNA polynucleotide sequences, such as gRNA sequences, include the nucleobase uracil (U), a pyrimidine derivative, rather than the nucleobase thymine (T) contained in DNA polynucleotide sequences. In RNA, uracil pairs with the adenine base and replaces thymine during DNA transcription.

[0237] "Heterodimer" refers to a fusion protein containing two domains, such as a wild-type TadA domain and a variant of the TadA domain (e.g., TadA*8) or two variant TadA domains (e.g., TadA*7.10 and TadA*8 or two TadA*8 domains).

[0238] "Host-versus-graft disease" (HVGD) refers to a pathological condition in which the host's immune system mounts an immune response against transplanted cells from a donor.

[0239] "Hybridization" refers to hydrogen bonding between complementary nucleobases, which can be Watson-Crick, Hoogsteen, or reverse Hoogsteen hydrogen bonds. For example, adenine and thymine are complementary nucleobases that pair by forming hydrogen bonds.

[0240] "Immune cell" refers to a cell of the immune system capable of mounting an immune response.

[0241] "Immune effector cell" refers to a lymphocyte that, once activated, is capable of mounting an immune response against a target cell. T cells are exemplary immune effector cells.

[0242] The term "base repair inhibitor" or "IBR" refers to a protein capable of inhibiting the activity of a nucleic acid repair enzyme such as a base excision repair (BER) enzyme. In some embodiments, the IBR is an inhibitor of inosine base excision repair. Exemplary inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGG1, hNEIL1, T7Endol, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is a catalytically inactive EndoV or a catalytically inactive hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, the base repair inhibitor is a catalytically inactive EndoV or a catalytically inactive hAAG.

[0243] In some embodiments, the base excision repair inhibitor is a uracil glycosylase inhibitor (UGI). UGI refers to a protein that can inhibit the uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain contains wild-type UGI or a fragment of wild-type UGI. In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments. In some embodiments, the base excision repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base excision repair inhibitor is a catalytically inactive inosine-specific nuclease or a "dead inosine-specific nuclease". Without wishing to be bound by any particular theory, a catalytically inactive inosine glycosylase (e.g., alkyladenine glycosylase (AAG)) can bind inosine but cannot generate an abasic site or remove inosine, thereby spatially blocking the newly formed inosine moiety from DNA damage / repair mechanisms. In some embodiments, the catalytically inactive inosine-specific nuclease is capable of binding inosine in a nucleic acid but not cleaving the nucleic acid. Non-limiting exemplary catalytically inactive inosine-specific nucleases include catalytically inactive alkyladenosine glycosylase (AAG nuclease), e.g., from humans, and catalytically inactive endonuclease V (EndoV nuclease), e.g., from Escherichia coli. In some embodiments, the catalytically inactive AAG nuclease contains an E125Q mutation or a corresponding mutation in another AAG nuclease.

[0244] "Increase" means a positive change of at least 10%, 25%, 50%, 75% or 100%.

[0245] "Intein" is a protein fragment that can self-excise and join the remaining fragments (exteins) with peptide bonds in a process called protein splicing. Inteins are also called "protein introns". The process by which an intein excises itself and joins the remaining part of the protein is referred to herein as "protein splicing" or "intein-mediated protein splicing". In some embodiments, the intein of a precursor protein (the protein containing the intein prior to intein-mediated protein splicing) comes from two genes. Such an intein is referred to herein as a split intein (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE of the catalytic subunit a of DNA polymerase III is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene can be referred to herein as "intein-N". The intein encoded by the dnaE-c gene can be referred to herein as "intein-C".

[0246] Other intein systems can also be used. For example, synthetic inteins based on dnaE intein, Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C) intein pairs have been described (e.g., in Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5, incorporated herein by reference). Non-limiting examples of intein pairs that can be used according to the present disclosure include: Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Patent No. 8,394,604, incorporated herein by reference).

[0247] Exemplary nucleotide and amino acid sequences of the intein are provided.

[0248] DnaE Intein-N DNA:

[0249] TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT

[0250] DnaE Intein-N protein:

[0251] CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN

[0252] DnaE Intein-C DNA:

[0253] ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGATATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT

[0254] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN

[0255] Cfa-N DNA:

[0256] TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA

[0257] Cfa-N protein:

[0258] CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP

[0259] Cfa-C DNA:

[0260] ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC

[0261] Cfa-C protein:

[0262] MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0263] Intein-N and intein-C can be fused to the N-terminal portion and the C-terminal portion of split Cas9, respectively, for linking the N-terminal portion and the C-terminal portion of split Cas9. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, i.e., an N--[N-terminal portion of split Cas9]-[Intein-N]--C structure is formed. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., an N-[Intein-C]--[C-terminal portion of split Cas9]-C structure is formed. The intein-mediated protein splicing mechanism for linking the proteins to which the intein is fused (e.g., split Cas9) is known in the art, for example, as described by Shah et al., Chem Sci. 2014; 5(1):446-461, which is incorporated herein by reference. Methods for designing and using inteins are known in the art and are described, for example, in WO2014004336, WO2017132580, US20150344549, and US20180127780, each of which is incorporated herein by reference in its entirety.

[0264] The terms “isolated,” “purified,” or “biologically pure” refer to a material that is, to varying degrees, free from components that normally accompany it in its native state. “Isolated” refers to the degree of separation from the original source or the surrounding environment. “Purified” refers to a degree of separation that is higher than isolation. A “purified” or “biologically pure” protein is sufficiently free of other materials such that any impurities do not substantially affect the biological properties of the protein or cause other adverse consequences. That is, a nucleic acid or peptide of the present invention is purified if it is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA techniques, or is substantially free of chemical precursors or other chemicals when chemically synthesized. Purity and homogeneity are typically determined using analytical chemistry techniques such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term “purified” can denote that a nucleic acid or protein gives rise to substantially one band in an electrophoretic gel. For proteins that can be modified (e.g., phosphorylated or glycosylated), different modifications may result in different isolated proteins, which can be purified separately.

[0265] "Isolated polynucleotide" refers to a nucleic acid (e.g., DNA) that does not contain a gene that is flanked by that gene in the naturally occurring genome of the organism from which the nucleic acid molecule of the present invention is derived. Thus, the term includes, for example, recombinant DNA incorporated into a vector; into a plasmid or virus that replicates autonomously; or into the genomic DNA of a prokaryote or eukaryote; or exists as an independent molecule separate from other sequences (e.g., cDNA or genomic or cDNA fragments produced by PCR or restriction endonuclease digestion). In addition, the term includes RNA molecules transcribed from DNA molecules, and recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence.

[0266] "Isolated polypeptide" refers to a polypeptide of the present invention that has been separated from its naturally associated components. Generally, a polypeptide is isolated when it is at least 60% (by weight) free of proteins and naturally occurring organic molecules. Preferably, the polypeptides of the present invention are at least 75% by weight, more preferably at least 90% by weight, and most preferably at least 99% by weight. The isolated polypeptides of the present invention can be obtained, for example, by extraction from natural sources, by expression of recombinant nucleic acids encoding such polypeptides; or by chemical synthesis of the protein. Purity can be measured by any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or by HPLC analysis.

[0267] As used herein, the term "linker" can refer to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that connects two molecules or moieties (e.g., two components of a protein complex or ribonucleoprotein complex), or two domains of a fusion protein, such as a polynucleotide programmable DNA-binding domain (e.g., dCas9) and a deaminase domain (e.g., an adenosine deaminase, a cytidine deaminase, or an adenosine deaminase and a cytidine deaminase). The linker can connect different components or different parts of a base editor system. For example, in some embodiments, the linker can connect the guide polynucleotide-binding domain of a polynucleotide programmable nucleotide-binding domain and the catalytic domain of a deaminase. In some embodiments, the linker can connect a CRISPR polypeptide and a deaminase. In some embodiments, the linker can connect Cas9 and a deaminase. In some embodiments, the linker can connect dCas9 and a deaminase. In some embodiments, the linker can connect nCas9 and a deaminase. In some embodiments, the linker can connect a guide polynucleotide and a deaminase. In some embodiments, the linker can connect a deaminating component and a polynucleotide programmable nucleotide-binding component of a base editor system. In some embodiments, the linker can connect the RNA-binding portion of a deaminating component and a polynucleotide programmable nucleotide-binding component of a base editor system. In some embodiments, the linker can connect the RNA-binding portion of a deaminating component and the RNA-binding portion of a polynucleotide programmable nucleotide-binding component of a base editor system. The linker can be located between or on both sides of two groups, molecules, or other moieties and is connected to each by a covalent bond or non-covalent interaction, thereby connecting the two. In some embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker can be an RNA linker. In some embodiments, the linker can comprise an aptamer capable of binding a ligand. In some embodiments, the ligand can be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker can comprise an aptamer derivable from a riboswitch. The riboswitch from which the aptamer is derived can be selected from theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosylcobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or pre-queosine1 (PreQ1) riboswitch. In some embodiments, the linker can comprise an aptamer that binds to a polypeptide or protein domain, such as a polypeptide ligand.In some embodiments, the polypeptide ligand can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand can be part of a base editor system component. For example, a nucleobase editing component can comprise a deaminase domain and an RNA recognition motif.

[0268] In some embodiments, the linker can be one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker can be about 5 - 100 amino acids in length, such as about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 - 30, 30 - 40, 40 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, or 90 - 100 amino acids in length. In some embodiments, the linker can be about 100 - 150, 150 - 200, 200 - 250, 250 - 300, 300 - 350, 350 - 400, 400 - 450, or 450 - 500 amino acids. Longer or shorter linkers can also be contemplated.

[0269] In some embodiments, the linker links the gRNA-binding domain of an RNA-programmable nuclease, including the Cas9 nuclease domain and the catalytic domain of a nucleic acid editing protein (e.g., a cytidine or adenosine deaminase). In some embodiments, the linker links dCas9 and a nucleic acid editing protein. For example, the linker is located between or on both sides of two groups, molecules, or other moieties and is covalently linked to each, thereby linking the two. In some embodiments, the linker can be one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker can be about 5 - 200 amino acids in length, such as 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or more amino acids in length. Longer or shorter linkers can also be contemplated.

[0270] In some embodiments, the domains of the base editor are fused via a linker comprising the amino acid sequence: SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the domains of the base editor are fused via a linker comprising the amino acid sequence SGSETPGTSESATPES, which linker may also be referred to as an XTEN linker. In some embodiments, the linker comprises the amino acid sequence SGGS. In some embodiments, the linker comprises an (SGGS)n, (GGGS)n, (GGGGS)n, (G)n, (EAAAK)n, (GGS)n, SGSETPGTSESATPES, or (XP)n motif, or any one of these, where n is independently an integer between 1 and 30, and where X is any amino acid. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15.

[0271] In some embodiments, the length of the linker is 24 amino acids. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the length of the linker is 40 amino acids. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the length of the linker is 64 amino acids. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGS SGGS. In some embodiments, the length of the linker is 92 amino acids. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.

[0272] "Marker" refers to any protein or polynucleotide that has an altered expression level or activity in relation to a disease or disorder.

[0273] As used herein, the term "mutation" refers to the replacement of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) with another residue, or the deletion or insertion of one or more residues within the sequence. Mutations are typically described herein by identifying the original residue followed by the position of that residue within the sequence and by the identity of the new replacement residue. The various methods for making amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)). In some embodiments, the presently disclosed base editors can effectively generate "desired mutations" in nucleic acids (e.g., nucleic acids within a subject's genome), such as point mutations, without generating a large number of undesired mutations, such as unexpected point mutations. In some embodiments, the desired mutation is a mutation generated by a specific base editor (e.g., a cytidine base editor or an adenosine base editor) that binds to a guide polynucleotide (e.g., a gRNA), which is specifically designed to generate the desired mutation.

[0274] Generally, mutations generated or identified in a sequence (e.g., an amino acid sequence as described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. Those skilled in the art will readily understand how to determine the position of a mutation in an amino acid and nucleic acid sequence relative to the reference sequence.

[0275] "Neoplasia" refers to cells or tissues that exhibit abnormal growth or proliferation. The term neoplasia includes cancers and solid tumors.

[0276] The term "non-conservative mutation" involves amino acid substitutions between different groups, e.g., tryptophan for lysine, or serine for phenylalanine, etc. In such cases, the non-conservative amino acid substitution preferably does not interfere with, or inhibits the biological activity of the functional variant. The non-conservative amino acid substitution can enhance the biological activity of the functional variant, such that the biological activity of the functional variant is increased compared to the wild-type protein.

[0277] "Nuclear factor of activated T cells 1 (NFATc1) polypeptide" refers to a protein having at least about 85% amino acid sequence identity to NCBI accession number NM_172390.2 or a fragment thereof, and is a component of the activated T cell DNA-binding transcriptional complex. Exemplary amino acid sequences are provided below.

[0278] >NP_765978.1 Nuclear factor of activated T cells, cytoplasmic 1 isoform A [Homo sapiens]

[0279] MPSTSFPVPSKFPLGPAAAVFGRGETLGPAPRAGGTMKSAEEEHYGYASSNVSPALPLPTAHSTLPAPCHNLQTSTPGIIPPADHPSGYGAALDGGPAGYFLSSGHTRPDGAPALESPRIEITSCLGLYHNNNQFFHDVEVEDVLPSSKRSPSTATLSLPSLEAYRDPSCLSPASSLSSRSCNSEASSYESNYSYPYASPQTSPWQSPCVSPKTTDPEEGFPRGLGACTLLGSPRHSPSTSPRASVTEESWLGARSSRPASPCNKRKYSLNGRQPPYSPHHSPTPSPHGSPRVSVTDDSWLGNTTQYTSSAIVAAINALTTDSSLDLGDGVPVKSRKTTLEQPPSVALKVEPVGEDLGSPPPPADFAPEDYSSFQHIRKGGFCDQYLAVPQHPYQWAKPKPLSPTSYMSPTLPALDWQLPSHSGPYELRIEVQPKSHHRAHYETEGSRGAVKASAGGHPIVQLHGYLENEPLMLQLFIGTADDRLLRPHAFYQVHRITGKTVSTTSHEAILSNTKVLEIPLLPENSMRAVIDCAGILKLRNSDIELRKGETDIGRKNTRVRLVFRVHVPQPSGRTLSLQVASNPIECSQRSAQELPLVEKQSTDSYPVVGGKKMVLSGHNFLQDSKVIFVEKAPDGHHVWEMEAKTDRDLCKPNSLVVEIPPFRNQRITSPVHVSFYVCNGKRKRSQYQRFTYLPANGNAIFLTVSREHERVGCFF

[0280] "Nuclear factor of activated T cells 1 (NFATc1) polynucleotide" refers to a nucleic acid molecule encoding an NFATc1 polypeptide. The NFATc1 gene encodes a protein that is involved in the inducible expression of cytokine genes, particularly IL-2 and IL-4, in T cells. Exemplary nucleic acids that have been sequenced are provided below.

[0281] >NM_172390.2 Homo sapiens nuclear factor of activated T cells 1 (NFATC1), transcript variant 1, mRNA

[0282]

[0283] The terms "nuclear localization sequence", "nuclear localization signal", or "NLS" refer to an amino acid sequence that facilitates the import of a protein into the nucleus. Nuclear localization sequences are known in the art and are described, for example, in the international PCT application of Plank et al., PCT / EP2000 / 011690, filed Nov. 23, 2000, and published as WO / 2001 / 038547 on May 31, 2001, the content of which is incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, such as that described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETY.

[0284] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to compounds that contain nucleobases and an acidic moiety, such as nucleosides, nucleotides, or polymers of nucleotides. Generally, polymeric nucleic acids, such as nucleic acid molecules that contain three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other by phosphodiester bonds. In some embodiments, "nucleic acid" refers to a single nucleic acid residue (e.g., nucleotide and / or nucleoside). In some embodiments, "nucleic acid" refers to an oligonucleotide chain that contains three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" are used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" includes RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids can be naturally occurring, such as in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, a nucleic acid molecule can be a non-naturally occurring molecule, such as recombinant DNA or RNA, artificial chromosome, engineered genome or fragment thereof, or synthetic DNA, RNA, DNA / RNA hybrid, or that includes non-naturally occurring nucleotides or nucleosides. In addition, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In suitable cases, such as in the case of chemically synthesized molecules, nucleic acids can contain nucleoside analogs, such as analogs having chemically modified bases or sugar and backbone modifications. Unless otherwise indicated, nucleic acid sequences are presented in the 5' to 3' direction. In some embodiments, the nucleic acid is or contains natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanosine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioates and 5'-N-phosphoramidite linkages).

[0285] The term "nucleic acid programmable DNA binding protein" or "napDNAbp" can be used interchangeably with "polynucleotide programmable nucleotide binding domain" to refer to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA), that directs the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable acid binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable acid binding domain is a polynucleotide programmable RNA binding domain. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 protein. The Cas9 protein can be associated with a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactivated Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, and Cas12i.Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI proteins, CARF, DinG, their homologs, or their modified or engineered versions. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically listed in the present disclosure. See, e.g., Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 Oct; 1: 325-336. doi:10.1089 / crispr.2018.0033; Yan et al., “Functionally diverse type V CRISPR-Cas systems” Science. 2019 Jan 4; 363(6422): 88-91. doi:10.1126 / science.aav7271, the entire contents of which are incorporated herein by reference.

[0286] The terms "nucleobase", "nitrogenous base", or "base" are used interchangeably herein and refer to nitrogen-containing biological compounds that form nucleosides, which are components of nucleotides. The ability of nucleobases to form base pairs and stack on top of each other directly leads to long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases - adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U) - are referred to as primary or canonical. Adenine and guanine are derived from purine, and cytosine, uracil, and thymine are derived from pyrimidine. DNA and RNA can also contain other (non-primary) modified bases. Non-limiting exemplary modified nucleobases can include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Hypoxanthine and xanthine can be produced by the presence of mutagens, and they are both produced by deamination (replacing an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be produced by deamination of cytosine. A "nucleoside" consists of a nucleobase and a five-carbon sugar (ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five-carbon sugar (ribose or deoxyribose), and at least one phosphate group.

[0287] As used herein, the term "nucleoside base editing domain" or "nucleoside base editing protein" refers to a protein or enzyme that can catalyze the modification of nucleoside bases in RNA or DNA, such as the deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine) and adenine (or adenosine) to hypoxanthine (or inosine), as well as non-templated nucleotide addition and insertion. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase; or cytidine deaminase or cytosine deaminase). In some embodiments, the nucleobase editing domain is multiple deaminase domains (e.g., adenine deaminase or adenosine deaminase and cytidine or cytosine deaminase). In some embodiments, the nucleobase editing domain can be a naturally occurring nucleobase editing domain. In some embodiments, the nucleobase editing domain can be a modified or evolved nucleobase editing domain derived from a naturally occurring nucleobase editing domain. The nucleobase editing domain can be from any organism, such as bacteria, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse.

[0288] As used herein, "obtaining," as in "obtaining an agent," includes synthesizing, purchasing, or otherwise obtaining the agent.

[0289] As used herein, "patient" or "subject" refers to a mammalian subject or individual diagnosed with, at risk of developing, or suspected of having or developing a disease or disorder. In some embodiments, the term "patient" refers to a mammalian subject having a higher than average likelihood of developing a disease or disorder. Exemplary patients can be humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female.

[0290] "A patient in need" or "a subject in need" as used herein refers to a patient diagnosed with, at risk of, or having, or suspected of having a disease or disorder.

[0291] The terms "pathogenic mutation," "pathogenic variant," "disease shell mutation," "pathogenic variant," "harmful mutation," or "susceptibility mutation" refer to a genetic alteration or mutation that increases an individual's susceptibility or predisposition to a particular disease or disorder. In some embodiments, a pathogenic mutation comprises the substitution of at least one wild-type amino acid with at least one pathogenic amino acid in a protein encoded by a gene.

[0292] The terms “protein,” “peptide,” “polypeptide,” and their grammatical equivalents are used interchangeably herein and refer to polymers of amino acid residues joined together by peptide (amide) bonds. These terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Generally, a protein, peptide, or polypeptide is at least three amino acids in length. A protein, peptide, or polypeptide can refer to a single protein or a collection of proteins. One or more of the amino acids in a protein, peptide, or polypeptide can be modified, e.g., by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl, geranylgeranyl, fatty acid groups, linkers for conjugation, functionalization, or other modification, etc. A protein, peptide, or polypeptide can also be a single molecule or can be a multimolecular complex. A protein, peptide, or polypeptide can be just a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof. As used herein, the term “fusion protein” refers to a hybrid polypeptide that contains protein domains from at least two different proteins. One protein can be located in the amino-terminal (N-terminal) portion or carboxy-terminal (C-terminal) portion of the fusion protein, thereby forming an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively. A protein can contain different domains, e.g., a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the protein to bind to a target site) and a nucleic acid cleavage domain, or the catalytic domain of a nucleic acid editing protein. In some embodiments, a protein contains a protein portion, e.g., an amino acid sequence that constitutes a nucleic acid binding domain, and an organic compound, e.g., a compound that can act as a nucleic acid cleavage agent. In some embodiments, a protein forms a complex or associates with a nucleic acid such as RNA or DNA. Any protein provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced by recombinant protein expression and purification, which is particularly applicable to fusion proteins that contain peptide linkers. Methods for recombinant protein expression and purification are well known and include those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire content of which is incorporated by reference.

[0293] The polypeptides and proteins disclosed herein (including functional portions and functional variants thereof) may contain synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-aminodecanoic acid, homoserine, S-acetamidomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxy phenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 3-quinoline carboxylic acid, 3-hydroxy phenylalanine, 3,4-dihydroxy phenylalanine, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, homophenylalanine, and α-tert-butylglycine. The polypeptides and proteins may be associated with post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include phosphorylation, acylation (including acetylation and formylation), glycosylation (including N-linked and O-linked), amidation, hydroxylation, alkylation (including methylation and ethylenation), ubiquitination, addition of pyrrolidone carboxylic acid, formation of disulfide bonds, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylgeranylation, glycosylation, lipoylation, and iodination.

[0294] The "programmed cell death 1 (PDCD1 or PD-1) polypeptide" refers to a protein having at least about 85% amino acid sequence identity with NCBI accession number AJS10360.1 or a fragment thereof. The PD-1 protein is thought to be involved in the regulation of T cell function in immune responses and tolerance conditions. Exemplary B2M polypeptide sequences are provided below.

[0295] >AJS10360.1 Programmed cell death 1 protein [Homo sapiens]

[0296] MQIPQAPWPVVWAVLQLGWRPGWFLDSPDRPWNPPTFSPALLVVTEGDNATFTCSFSNTSESFVLNWYRMSPSNQTDKLAAFPEDRSQPGQDCRFRVTQLPNGRDFHMSVVRARRNDSGTYLCGAISLAPKAQIKESLRAELRVTERRAEVPTAHPSPSPRPAGQFQTLVVGVVGGLLGSLVLLVWVLAVICSRAARGTIGARRTGQPLKEDPSAVPVFSVDYGELDFQWREKTPEPPVPCVPEQTEYATIVFPSGMGTSSPARRGSADGPRSAQPLRPEDGHCSWPL

[0297] "Programmed cell death 1 (PDCD1 or PD-1) polynucleotide" refers to a nucleic acid molecule encoding a PD-1 polypeptide. The PDCD1 gene encodes an inhibitory cell surface receptor that inhibits T cell effector function in an antigen-specific manner. Exemplary PDCD1 nucleic acid sequences are provided below.

[0298] >AY238517.1 Homo sapiens programmed cell death 1 (PDCD1) mRNA, complete cds

[0299] ATGCAGATCCCACAGGCGCCCTGGCCAGTCGTCTGGGCGGTGCTACAACTGGGCTGGCGGCCAGGATGGTTCTTAGACTCCCCAGACAGGCCCTGGAACCCCCCCACCTTCTCCCCAGCCCTGCTCGTGGTGACCGAAGGGGACAACGCCACCTTCACCTGCAGCTTCTCCAACACATCGGAGAGCTTCGTGCTAAACTGGTACCGCATGAGCCCCAGCAACCAGACGGACAAGCTGGCCGCCTTCCCCGAGGACCGCAGCCAGCCCGGCCAGGACTGCCGCTTCCGTGTCACACAACTGCCCAACGGGCGTGACTTCCACATGAGCGTGGTCAGGGCCCGGCGCAATGACAGCGGCACCTACCTCTGTGGGGCCATCTCCCTGGCCCCCAAGGCGCAGATCAAAGAGAGCCTGCGGGCAGAGCTCAGGGTGACAGAGAGAAGGGCAGAAGTGCCCACAGCCCACCCCAGCCCCTCACCCAGGCCAGCCGGCCAGTTCCAAACCCTGGTGGTTGGTGTCGTGGGCGGCCTGCTGGGCAGCCTGGTGCTGCTAGTCTGGGTCCTGGCCGTCATCTGCTCCCGGGCCGCACGAGGGACAATAGGAGCCAGGCGCACCGGCCAGCCCCTGAAGGAGGACCCCTCAGCCGTGCCTGTGTTCTCTGTGGACTATGGGGAGCTGGATTTCCAGTGGCGAGAGAAGACCCCGGAGCCCCCCGTGCCCTGTGTCCCTGAGCAGACGGAGTATGCCACCATTGTCTTTCCTAGCGGAATGGGCACCTCATCCCCCGCCCGCAGGGGCTCAGCTGACGGCCCTCGGAGTGCCCAGCCACTGAGGCCTGAGGATGGACACTGCTCTTGGCCCCTCTGA

[0300] As used herein, the term "recombinant" in the context of a protein or nucleic acid refers to a protein or nucleic acid that does not exist in nature but is a product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule comprises an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.

[0301] "Reduce" means a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0302] "Reference" means a standard or control condition. In one embodiment, the reference is a wild-type or healthy cell. In other embodiments and without limitation, the reference is an untreated cell that has not been subjected to the test condition, or has been subjected to a placebo or saline, medium, buffer, and / or a control vector that does not contain the target polynucleotide.

[0303] "Reference sequence" is a defined sequence that serves as a basis for sequence comparison. The reference sequence can be a subset or all of a particular sequence; for example, a fragment of a full-length cDNA or gene sequence, or a complete cDNA or gene sequence. For a polypeptide, the length of the reference polypeptide sequence is typically at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For a nucleic acid, the length of the reference nucleic acid sequence is typically at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides or any integer near or between them. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is the polynucleotide sequence encoding the wild-type protein.

[0304] The terms “RNA programmable nuclease” and “RNA-guided nuclease” are used (e.g., in combination or association) with one or more RNAs that are not the target for cleavage. In some embodiments, an RNA programmable nuclease may be referred to as a nuclease:RNA complex when complexed with RNA. Generally, the bound RNA is referred to as guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. A gRNA that exists as a single RNA molecule may be referred to as a single guide RNA (sgRNA), although “gRNA” may be used interchangeably to refer to guide RNA that exists as a single molecule or as a complex of two or more molecules. Generally, a gRNA that exists as a single RNA species contains two domains: (1) a domain that is homologous to the target nucleic acid (e.g., directs binding of the Cas9 complex to the target); (2) a domain that binds the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence referred to as tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA provided in Jinek et al., Science 337:816-821 (2012), the entire content of which is incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in U.S. Provisional Patent Application, U.S.S.N. 61 / 874,682, filed Sep. 6, 2013, entitled “Switchable Cas9 Nucleases and Uses Thereof,” and U.S. Provisional Patent Application, U.S.S.N. 61 / 874,746, filed Sep. 6, 2013, named “Delivery System For Functional Nucleases,” the entire content of each of which is incorporated herein by reference in its entirety. In some embodiments, the gRNA contains two or more of domains (1) and (2) and may be referred to as an “extended gRNA.” For example, as described herein, an extended gRNA will bind two or more Cas9 proteins and bind the target nucleic acid at two or more distinct regions. The gRNA contains a nucleotide sequence that is complementary to the target site, which mediates binding of the nuclease / RNA complex to the target site, providing sequence specificity to the nuclease:RNA complex.

[0305] In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, such as Cas9 (Casnl) from Streptococcus pyogenes (see, e.g., "Complete genome sequence of an Mlstrain of Streptococcus pyogenes." Ferretti J.J., McShan W.M., Ajdic D.J., SavicD.J., Savic G., Lyon K., Primeaux C, Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factorRNase III." Deltcheva E., Chylinski K., Sharma CM., Gonzales K., Chao Y., PirzadaZ.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607 (2011).

[0306] Because RNA-programmable nucleases (such as Cas9) use RNA:DNA hybridization to target DNA cleavage sites, these proteins are in principle capable of targeting any sequence specified by the guide RNA. Methods for site-specific cleavage (e.g., modifying the genome) using RNA-programmable nucleases (such as Cas9) are known in the art (see, e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, W.Y. et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); DiCarlo, J.E. et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013); the entire contents of which are incorporated herein by reference).

[0307] The term "single nucleotide polymorphism (SNP)" refers to a variation of a single nucleotide at a specific position in the genome, where each variation exists at a certain frequency (e.g., >1%) in the population. For example, at a specific base position in the human genome, the C nucleotide may be present in most individuals, but in a few individuals, this position is occupied by an A. This means that there is an SNP at this specific position, and the two possible nucleotide variations, C or A, are called the alleles at this position. SNPs are the basis for differences in disease susceptibility. The severity of a disease and the way our body responds to treatment are also manifestations of genetic variation. SNPs can fall within the coding region of a gene, the non-coding region of a gene, or the intergenic region (the region between genes). In some embodiments, due to the degeneracy of the genetic code, SNPs within the coding sequence do not necessarily change the amino acid sequence of the resulting protein. There are two types of SNPs in the coding region: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, and non-synonymous SNPs change the amino acid sequence of the protein. The non-synonymous SNPs have two types: missense and nonsense. SNPs that are not in the protein-coding region can still affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNAs. Gene expression affected by such SNPs is called eSNP (expression SNP) and can be located upstream or downstream of the gene. A single nucleotide variant (SNV) is a variation of a single nucleotide without any frequency limitation and can occur in somatic cells. Somatic single nucleotide variants can also be referred to as single nucleotide alterations.

[0308] "Specifically binds" means a nucleic acid molecule, polypeptide, or a complex thereof (e.g., a nucleic acid programmable DNA binding domain and a guide nucleic acid), a compound, or a molecule that recognizes and binds to the polypeptide and / or nucleic acid molecule of the present invention, but substantially does not recognize and bind to other molecules in the sample, such as a biological sample.

[0309] Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecule encoding a polypeptide of the present disclosure or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will generally exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence is generally capable of hybridizing to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecule encoding a polypeptide of the present disclosure or a fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will generally exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence is generally capable of hybridizing to at least one strand of a double-stranded nucleic acid molecule. "Hybridization" refers to the pairing between complementary polynucleotide sequences (e.g., the genes described herein) or portions thereof to form a double-stranded molecule under various stringent conditions. (See, e.g., Wahl, G.M. and S.L. Berger (1987) Methods Enzymol. 152:399; Kimmel, A.R. (1987) Methods Enzymol. 152:507).

[0310] For example, stringent salt concentrations are typically less than about 750 mM NaCl and 75 mM sodium citrate, preferably less than about 500 mM NaCl and 50 mM sodium citrate, and more preferably less than about 250 mM NaCl and 25 mM sodium citrate. Low stringency hybridization can be achieved in the absence of organic solvents such as formamide, while high stringency hybridization can be achieved in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions generally include temperatures of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. For example, sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those of skill in the art. Different degrees of stringency are achieved by combining these different conditions as needed. In one embodiment, hybridization will occur at 30°C in 750 mM NaCl, 75 mM sodium citrate, and 1% SDS. In another embodiment, hybridization will occur at 37°C in 500 mM NaCl, 50 mM sodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In one embodiment, hybridization will occur at 42°C in 250 mM NaCl, 25 mM sodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions will be apparent to those of skill in the art.

[0311] For most applications, the washing steps after hybridization also vary in terms of stringency. Washing stringency conditions can be defined by salt concentration and temperature. As described above, washing stringency can be increased by decreasing the salt concentration or increasing the temperature. For example, the stringent salt concentration for the washing step is preferably less than about 30 mM NaCl and 3 mM trisodium citrate, and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. The stringent temperature conditions for the washing step generally include a temperature of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, the washing step will occur at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step will be carried out at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step will be carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Other variations of these conditions will be apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0312] "Split" means to divide into two or more segments.

[0313] "Split Cas9 protein" or "split Cas9" refers to the Cas9 protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal portion and the C-terminal portion of the Cas9 protein can be spliced to form a "reconstituted" Cas9 protein. In certain embodiments, the Cas9 protein is split into two fragments within a disordered region of the protein, e.g., as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as in Jiang et al. (2016) Science 351:867-871. PDB file: 5F9R, each of which is incorporated herein by reference. In some embodiments, the protein is split into two fragments at any C, T, A, or S between approximately amino acids A292-G364, F445-K483, or E565-T637 within the SpCas9 region, or at any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "splitting" the protein.

[0314] In other embodiments, the N-terminal portion of the Cas9 protein comprises amino acids 1-573 or 1-637 of Streptococcus pyogenes Cas9 wild-type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot Reference Sequence: Q99ZW2) and the C-terminal portion of the Cas9 protein comprises a portion of amino acids 574-1368 or 638-1368 of SpCas9 wild-type or their corresponding positions.

[0315] The C-terminal portion of split Cas9 can be joined to the N-terminal portion of split Cas9 to form a complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein begins where the N-terminal portion of the Cas9 protein ends. Thus, in some embodiments, the C-terminal portion of split Cas9 comprises a portion of amino acids (551-651)-1368 of spCas9. "(551-651)-1368" means starting from the amino acids between amino acids 551-651 (inclusive) and ending at amino acid 1368.For example, the C-terminal portion of split Cas9 can include a portion of any amino acids of SpCas9: 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556-1368, 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368, 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368, 573-1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368, 581-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368, 589-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368, 597-1368, 598-1368, 599-1368, 600-1368, 601-1368, 602-1368, 603-1368, 604-1368, 605-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368, 613-1368, 614-1368, 615-1368, 616-1368, 617-1368, 618-1368, 619-1368, 620-1368, 621-1368, 622-1368, 623-1368, 624-1368, 625-1368, 626-1368, 627-1368, 628-1368, 629-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368, 637-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368, 645-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368 or 651-1368. In some embodiments, the C-terminal portion of split Cas9 includes a portion of 574-1368 or 638-1368 of the SpCas9 protein.

[0316] "Subject" means a mammal, including but not limited to a human or non - human mammal such as a cow, horse, dog, sheep or cat. Subjects include domestic animals, domesticated animals raised for the production of labor and the provision of goods such as food, including but not limited to cows, goats, chickens, horses, pigs, rabbits and sheep.

[0317] "Substantially identical" means a polypeptide or nucleic acid molecule to a reference amino acid sequence (e.g., any of the amino acid sequences described herein) or nucleic acid sequence (e.g., any of the nucleic acid sequences described herein). In one embodiment, such a sequence is at least 60%, 80% or 85%, 90%, 95% or even 99% identical to the sequence being compared at the amino acid level or nucleic acid.

[0318] Sequence identity is generally determined using sequence analysis software (e.g., the Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning a degree of homology for various substitutions, deletions and / or other modifications. Conservative substitutions generally include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary method for determining the degree of identity, the BLAST program can be used, where a probability score between e - 3 and e - 100 indicates closely related sequences.

[0319] For example, COBALT is used with the following parameters:

[0320] a) alignment parameters: Gap penalties - 11, - 1 and End - Gap penalties - 5, - 1,

[0321] b) CDD Parameters: Use RPS BLAST on; Blast E - value 0.003; Find Conserved columns and Recompute on, and

[0322] c) Query Clustering Parameters: Use query clusters on; Word Size 4; Maxcluster distance 0.8; Alphabet Regular.

[0323] For example, EMBOSS is used with the following parameters:

[0324] a) Matrix: BLOSUM62;

[0325] b) GAP OPEN: 10;

[0326] c) GAP EXTEND: 0.5;

[0327] d) OUTPUT FORMAT: pair;

[0328] e) END GAP PENALTY: false;

[0329] f) END GAP OPEN: 10; and

[0330] g) END GAP EXTEND: 0.5.

[0331] The term "target site" refers to a sequence within a nucleic acid molecule that is modified by a base editor. In one embodiment, the target site is deaminated by a deaminase or a fusion protein comprising a deaminase (e.g., a cytidine or adenine deaminase).

[0332] "Tet methylcytosine dioxygenase 2 (TET2) polypeptide" refers to a protein having at least about 85% amino acid sequence identity with NCBI accession number FM992369.1 or a fragment thereof and having catalytic activity for converting methylcytosine to 5-hydroxymethylcytosine. Defects in this gene are associated with myeloproliferative disorders, and the ability of this enzyme to methylate cytosine contributes to transcriptional regulation. Exemplary TET2 amino acid sequences are provided below.

[0333] >CAX30492.1 tet oncogene family member 2 [Homo sapiens]

[0334]

[0335] "tet methylcytosine dioxygenase 2 (TET2) polynucleotide" refers to a nucleic acid molecule encoding a TET2 polypeptide. The TETs polypeptides encode methylcytosine dioxygenases and have transcriptional regulatory activity. Exemplary TET2 nucleic acids are presented below.

[0336]

[0337] The "transforming growth factor receptor 2 (TGFBRII) polypeptide" refers to a protein having at least about 85% sequence identity with NCBI accession number ABG65632.1 or a fragment thereof and having immunosuppressive activity. Exemplary amino acid sequences are provided below:

[0338] >ABG65632.1 transforming growth factor beta receptor II [Homo sapiens]

[0339] MGRGLLRGLWPLHIVLWTRIASTIPPHVQKSVNNDMIVTDNNGAVKFPQLCKFCDVRFSTCDNQKSCMSNCSITSICEKPQEVCVAVWRKNDENITLETVCHDPKLPYHDFILEDAASPKCIMKEKKKPGETFFMCSCSSDECNDNIIFSEEYNTSNPDLLLVIFQVTGISLLPPLGVAISVIIIFYCYRVNRQQKLSSTWETGKTRKLMEFSEHCAIILEDDRSDISSTCANNINHNTELLPIELDTLVGKGRFAEVYKAKLKQNTSEQFETVAVKIFPYEEYASWKTEKDIFSDINLKHENILQFLTAEERKTELGKQYWLITAFHAKGNLQEYLTRHVISWEDLRKLGSSLARGIAHLHSDHTPCGRPKMPIVHRDLKSSNILVKNDLTCCLCDFGLSLRLDPTLSVDDLANSGQVGTARYMAPEVLESRMNLENVESFKQTDVYSMALVLWEMTSRCNAVGEVKDYEPPFGSKVREHPCVESMKDNVLRDRGRPEIPSFWLNHQGIQMVCETLTECWDHDPEARLTAQCVAERFSELEHLDRLSGRSCSEEKIPEDGSLNTTK

[0340] The "transforming growth factor receptor 2 (TGFBRII) polynucleotide" refers to a nucleic acid encoding the TGFBRII polypeptide. The TGFBRII gene encodes a transmembrane protein having serine / threonine kinase activity. Exemplary TGFBRII nucleic acids are provided below:

[0341]

[0342] "T cell immunoreceptor with Ig and ITIM domains (TIGIT) polypeptide" refers to a protein having at least about 85% sequence identity with NCBI accession number ACD74757.1 or a fragment thereof and having immunomodulatory activity. Exemplary TIGIT amino acid sequences are provided below.

[0343] >ACD74757.1 T cell immunoreceptor with Ig and ITIM domains [Homo sapiens] MRWCLLLIWAQGLRQAPLASGMMTGTIETTGNISAEKGGSIILQCHLSSTTAQVTQVNWEQQDQLLAICNADLGWHISPSFKDRVAPGPGLGLTLQSLTVNDTGEYFCIYHTYPDGTYTGRIFLEVLESSVAEHGARFQIPLLGAMAATLVVICTAVIVVVALTRKKKALRIHSVEGDLRRKSAGQEEWSPSAPSPPGSCVQAEAAPAGLCGEQRGEDCAELHDYFNVLSYRSLGNCSFFTETG

[0344] "T cell immunoreceptor with Ig and ITIM domains (TIGIT) polynucleotide" refers to a nucleic acid encoding a TIGIT polypeptide. The TIGIT gene encodes an inhibitory immunoreceptor associated with tumor formation and T cell exhaustion. Exemplary nucleic acid sequences are provided below:

[0345] >EU675310.1 Homo sapiens T cell immunoreceptor with Ig and ITIM domains (TIGIT) mRNA, complete cds

[0346] CGTCCTATCTGCAGTCGGCTACTTTCAGTGGCAGAAGAGGCCACATCTGCTTCCTGTAGGCCCTCTGGGCAGAAGCATGCGCTGGTGTCTCCTCCTGATCTGGGCCCAGGGGCTGAGGCAGGCTCCCCTCGCCTCAGGAATGATGACAGGCACAATAGAAACAACGGGGAACATTTCTGCAGAGAAAGGTGGCTCTATCATCTTACAATGTCACCTCTCCTCCACCACGGCACAAGTGACCCAGGTCAACTGGGAGCAGCAGGACCAGCTTCTGGCCATTTGTAATGCTGACTTGGGGTGGCACATCTCCCCATCCTTCAAGGATCGAGTGGCCCCAGGTCCCGGCCTGGGCCTCACCCTCCAGTCGCTGACCGTGAACGATACAGGGGAGTACTTCTGCATCTATCACACCTACCCTGATGGGACGTACACTGGGAGAATCTTCCTGGAGGTCCTAGAAAGCTCAGTGGCTGAGCACGGTGCCAGGTTCCAGATTCCATTGCTTGGAGCCATGGCCGCGACGCTGGTGGTCATCTGCACAGCAGTCATCGTGGTGGTCGCGTTGACTAGAAAGAAGAAAGCCCTCAGAATCCATTCTGTGGAAGGTGACCTCAGGAGAAAATCAGCTGGACAGGAGGAATGGAGCCCCAGTGCTCCCTCACCCCCAGGAAGCTGTGTCCAGGCAGAAGCTGCACCTGCTGGGCTCTGTGGAGAGCAGCGGGGAGAGGACTGTGCCGAGCTGCATGACTACTTCAATGTCCTGAGTTACAGAAGCCTGGGTAACTGCAGCTTCTTCACAGAGACTGGTTAGCAACCAGAGGCATCTTCTGG

[0347] The "T cell receptor alpha constant (TRAC) polypeptide" refers to a protein having at least about 85% amino acid sequence identity with NCBI accession number P01848.2 or a fragment thereof and having immunomodulatory activity. Exemplary amino acid sequences are provided below.

[0348] >sp|P01848.2|TRAC_HUMAN RecName:Full=T cell receptor alpha constant IQNPDPAVYQLRDSKSSDKSVCLFTDFDSQTNVSQSKDSDVYITDKTVLDMRSMDFKSNSAVAWSNKSDFACANAFNNSIIPEDTFFPSPESSCDVKLVEKSFETDTNLNFQNLSVIGFRILLLKVAGFNLLMTLRLWSS

[0349] "T cell receptor alpha constant (TRAC) polynucleotide" refers to a nucleic acid encoding a TRAC polypeptide. Exemplary TRAC nucleic acid sequences are provided below.

[0350] >X02592.1 Human mRNA for T cell receptor alpha chain (TCR-alpha)

[0351]

[0352] As used herein, "transduction" refers to the transfer of a gene or genetic material into a cell by a viral vector.

[0353] As used herein, "transformation" refers to the process of introducing a genetic change in a cell produced by the introduction of foreign nucleic acid.

[0354] "Transfection" refers to the transfer of a gene or genetic material into a cell by chemical or physical means.

[0355] "Translocation" refers to the rearrangement of a nucleic acid fragment between non-homologous chromosomes.

[0356] As used herein, the terms "treat", "treating", "treatment", etc. refer to reducing or ameliorating a disorder and / or its associated symptoms or obtaining a desired pharmacological and / or physiological effect. It should be understood that, although not excluded, treating a disorder or condition does not require complete elimination of the disorder, condition or its associated symptoms. In some embodiments, the effect is therapeutic, i.e., but not limited to, the effect partially or completely reduces, attenuates, eliminates, alleviates, mitigates, decreases the intensity of the disease and / or adverse symptoms attributable to the disease or cures the disease and / or adverse symptoms. In some embodiments, the effect is prophylactic, i.e., the effect protects or prevents the occurrence or recurrence of a disease or disorder. To this end, the presently disclosed methods include administering a therapeutically effective amount of a composition as described herein.

[0357] "Uracil glycosylase inhibitor" or "UGI" refers to an agent that inhibits the uracil excision repair system. In one embodiment, the agent is a protein or fragment thereof that binds to the host uracil-DNA glycosylase and prevents the removal of uracil residues from DNA. In one embodiment, UGI is a protein, fragment or domain thereof that is capable of inhibiting the uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or a modified version thereof. In some embodiments, the UGI domain comprises a fragment of the exemplary amino acid sequence set forth below. In some embodiments, the amino acid sequence comprised by the UGI fragment is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% or 100% identical to the exemplary UGI sequence provided below. In some embodiments, UGI comprises an amino acid sequence that is homologous to the exemplary UGI amino acid sequence or a fragment thereof, as described below. In some embodiments, the UGI or a portion thereof is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or at least 99% or 100% identical to wild-type UGI or the UGI sequence or a portion thereof, as described below. Exemplary UGI comprises the following amino acid sequence:

[0358] >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSD APEYKPWALVIQDSNGENKIKML.

[0359] The term "vector" refers to a means for introducing a nucleic acid sequence into a cell to produce a transformed cell. Vectors include plasmids, transposons, bacteriophages, viruses, liposomes and episomes. An "expression vector" is a nucleic acid sequence that contains a nucleotide sequence to be expressed in a recipient cell. An expression vector may include additional nucleic acid sequences to facilitate and / or promote the expression of the introduced sequence, such as initiation, termination, enhancer, promoter and secretion sequences.

[0360] "Zeta chain of the T cell receptor-associated protein kinase 70 (ZAP70) polypeptide" refers to a protein having at least about 85% amino acid sequence identity with NCBI accession number AAH53878.1 and having kinase activity. Exemplary amino acid sequences are provided below.

[0361] >AAH53878.1 Zeta-chain (TCR)-associated protein kinase 70 kDa [Homo sapiens] MPDPAAHLPFFYGSISRAEAEEHLKLAGMADGLFLLRQCLRSLGGYVLSLVHDVRFHHFPIERQLNGTYAIAGGKAHCGPAELCEFYSRDPDGLPCNLRKPCNRPSGLEPQPGVFDCLRDAMVRDYVRQTWKLEGEALEQAIISQAPQVEKLIATTAHERMPWYHSSLTREEAERKLYSGAQTDGKFLLRPRKEQGTYALSLIYGKTVYHYLISQDKAGKYCIPEGTKFDTLWQLVEYLKLKADGLIYCLKEACPNSSASNASGAAAPTLPAHPSTLTHPQRRIDTLNSDGYTPEPARITSPDKPRPMPMDTSVYESPYSDPEELKDKKLFLKRDNLLIADIELGCGNFGSVRQGVYRMRKKQIDVAIKVLKQGTEKADTEEMMREAQIMHQLDNPYIVRLIGVCQAEALMLVMEMAGGGPLHKFLVGKREEIPVSNVAELLHQVSMGMKYLEEKNFVHRDLAARNVLLVNRHYAKISDFGLSKALGADDSYYTARSAGKWPLKWYAPECINFRKFSSRSDVWSYGVTMWEALSYGQKPYKKMKGPEVMAFIEQGKRMECPPECPPELYALMSDCWIYKWEDRPDFLTVEQRMRACYYSLASKVEGPPGSTQKAEAACA

[0362] "Zeta-chain of T cell receptor-associated protein kinase 70 (ZAP70) polynucleotide" refers to a nucleic acid encoding a ZAP70 polypeptide. The ZAP70 gene encodes a tyrosine kinase involved in T cell development and lymphocyte activation. Lack of functional ZAP10 can lead to severe combined immunodeficiency, characterized by the absence of CD8+ T cells. Exemplary ZAP70 nucleic acid sequences are provided below.

[0363]

[0364] Any composition or method provided herein can be combined with one or more of any other compositions and methods provided herein.

[0365] DNA editing has emerged as a viable means of altering disease states by correcting pathogenic mutations at the gene level. Until recently, all DNA editing platforms functioned by inducing DNA double-strand breaks (DSBs) at specific genomic loci and relying on endogenous DNA repair pathways to determine product outcomes in a semi-random manner, resulting in complex populations of genetic products. While precise, user-defined repair outcomes can be achieved via the homology-directed repair (HDR) pathway, numerous challenges have hindered the efficient use of HDR for repair in therapeutically relevant cell types. In practice, this pathway is inefficient relative to the competing, error-prone non-homologous end joining pathway. Additionally, HDR is strictly restricted to the G1 and S phases of the cell cycle, precluding the precise repair of DSBs in post-mitotic cells. As a result, it has proven difficult or impossible to efficiently alter genomic sequences in these populations in a user-defined, programmable manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0366] Figure 1A and 1B are diagrams of three proteins that affect T cell function. Figure 1A is a diagram of the TRAC protein, which is a key component of graft-versus-host disease. Figure 1B is a diagram of the B2M protein, which is a component of the MHC class I antigen presentation complex present on nucleated cells and can be recognized by the host's CD8+ T cells. Figure 1C is a diagram of T cell signaling that leads to the expression of the PDCD1 gene, and the resulting PD-1 protein functions to inhibit T cell signaling.

[0367] Figures 2A to 2D depicts the A·T to G·C conversion and phenotypic outcomes in primary cells. Figure 2A is a violin plot depicting the reduction in protein expression measured by flow cytometry after electroporating primary human T cells with the indicated mRNA and 41 individual sgRNAs targeting 6 genes. The individual values shown represent the average percentage of cells with reduced protein expression from two replicate cells edited with the indicated mRNA and one of the 41 sgRNAs tested. Figure 2B is a heat map depicting the NGS analysis of A·T to G·C conversion by 8 ABE8 mRNAs and ABE7.10-m / d at 6 target sites. The values shown reflect the average of three independent biological replicates. The positions of the edited nucleotides for each target site are shown above the heat map. Figure 2CThe NGS analysis diagram depicting A·T to G·C conversion in multiplex-edited T cells at locus 21 (B2M), locus 25 (TRAC), and locus 24 (CIITA) after primary human T cells were electroporated is shown. The mRNA and three sgRNAs are presented in a multiplex-edited format. Figure 2D (Upper panel) Protein expression diagrams of B2M, CIITA, and TRAC proteins, as measured by flow cytometry on the cell population in Figure 1. Five days after electroporation at 2C. The values shown are from a representative donor. Figure 2D (Lower panel) A table describing the percentage of cell expression measured by flow cytometry after ABE editing using the specified one.

[0368] Figure 3 A heatmap depicting protein knockdown by ABE editors in primary T cells measured by flow cytometry. Eight mRNAs encoding the ABE8 editor and two mRNAs encoding ABE7.10-m / d were transfected into T cells with 41 sgRNAs targeting six genes respectively, and their effects on protein expression were measured using flow cytometry. The values shown are the mean of n = 2 independent replicates.

[0369] Figure 4 A chart depicting that ABE-edited CAR-T cells have effective cytotoxic activity against antigen-positive tumor cells. Fluorescently labeled RPMI-8226 cells were inoculated at time = 0 h, and their growth was monitored using the IncuCyte live cell imaging system within 28 h before the introduction of CAR-T cells. T cells multiplex-edited using the specified ABE ( Figure 1C ) were transduced with lentivirus encoding an anti-BCMACAR molecule and introduced into RPMI-8226 cells at time = 28 h, and the growth of RPMI-8226 cells was monitored for an additional 68 h. The values shown are the mean of n = 3 independent biological replicates.

[0370] Figure 5A and 5B Depicts RNA amplicon sequencing to detect cellular A-to-I editing in RNA associated with ABE treatment. Individual data points are shown, and error bars represent the standard deviation. For n = 3 independent biological replicates, performed on different days. Figure 5A A graph depicting the A-to-I editing frequency in the targeted RNA amplicon of the core ABE 8 construct compared to the ABE7 and Cas9 (D10A) nickase controls. Figure 5B A graph depicting the A-to-I editing frequency in the targeted RNA amplicon of ABE8 with mutations reported to improve RNA off-target editing.

[0371] Figure 6A and6B It is a diagram depicting examples of gates for evaluating protein knockdown in T cells. A representative gating strategy for population analysis of live, single lymphocytes to determine the reduction of surface proteins by flow cytometry.

[0372] Figure 7 It is a graph depicting alleles generated by ABE across 8 different genomic loci in HEK293T cells.

[0373] Figure 8A and 8B Depicts whole transcriptome and whole genome sequencing data from cells treated with base editor mRNA. Figure 8A It is a strip chart depicting whole transcriptome sequencing in HEK293T cells treated with the specified mRNA. Variant allele frequencies of A->G mutations within the RNA transcriptome were observed in repeated HEK293T cell experiments. Total A->G mutations are shown above each sample. Figure 8B It is a strip chart depicting whole transcriptome sequencing in T cells treated with the specified mRNA. Variant allele frequencies of A-to-G mutations within the RNA transcriptome were observed in three different T cell donors. Total A to G mutations are shown above each sample.

[0374] Figure 9A and 9B Depicts representative examples of gates for flow sorting of B2M-positive and B2M-negative cells prior to whole genome sequencing. Figure 9A Depicts a representative diagram and gates of live B2M-positive HEK293T cells sorted into single cell clones under untreated conditions. Figure 9B Depicts a representative diagram and gates of live, B2M-negative HEK293T cells sorted for all treatment conditions (ABE, CBE, or Cas9-treated cells).

[0375] Figure 10 It is a table describing Cas9 variants for accessing all possible PAMs within the NRNN PAM space. Only Cas9 variants that require three or fewer defined nucleotides in their PAMs are listed. Non-G PAM variants include SpCas9-NRRH, SpCas9-NRTH, and SpCas9-NRCH. Detailed implementation

[0376] The present invention features genetically modified immune cells comprising a novel adenosine base editor (e.g., ABE8), which have enhanced anti-tumor activity, resistance to immunosuppression, and a reduced risk of causing graft-versus-host reaction or host-versus-graft reaction, or a combination thereof. The present invention also features methods of producing and using such modified immune effector cells (e.g., immune effector cells such as T cells). The present invention also features methods of treating a subject having or at risk of developing a tumor, graft-versus-host disease (GVHD), or host-versus-graft disease (HVGD) with an effective amount of the modified immune effector cells (e.g., CAR-T cells).

[0377] Use a base editor system comprising an adenosine deaminase as described herein to modify immune effector cells to express chimeric antigen receptors (CARs) and to knockout or knockdown specific genes to reduce the negative impact that their expression may have on immune cell function.

[0378] Autologous, patient-derived chimeric antigen receptor-T cell (CAR-T) therapy has shown significant efficacy in treating some blood cancers. While these products have provided significant clinical benefits to patients, the need to generate personalized therapies has presented significant manufacturing challenges and financial burdens. Allogeneic CAR-T therapy has been developed as a potential solution to these challenges, with similar clinical efficacy characteristics to autologous products while treating many patients with cells from a single healthy donor, thus significantly reducing the cost of goods and batch-to-batch variability.

[0379] Most first-generation allogeneic CAR-Ts use nucleases to introduce two or more targeted genomic DNA double-strand breaks (DSBs) in a population of target T cells, relying on error-prone DNA repair to generate mutations that knockout target genes in a semi-random manner. This nuclease-based gene knockout strategy is designed to reduce the risk of graft-versus-host disease and host rejection of CAR-Ts. However, the simultaneous induction of multiple DSBs results in the final cell product containing large-scale genomic rearrangements, such as balanced and unbalanced translocations, as well as a relatively high abundance of local rearrangements, including inversions and large deletions. In addition, as more and more simultaneous genetic modifications are made by induced DSBs, significant genotoxicity is observed in the treated cell population. This has the potential to significantly reduce the cell expansion potential per production run, thereby reducing the number of patients that can be treated per healthy donor.

[0380] Base editors (BEs) are a class of emerging gene editing reagents that enable efficient, user-defined modification of target genomic DNA without creating DSBs. Here, an alternative method for producing allogeneic CAR-T cells is presented, by using base editing technology to reduce or eliminate detectable genomic rearrangements while enhancing cell expansion. As shown herein, compared to nuclease-only editing strategies, multiplex base editing of three loci simultaneously by base editing generates efficient gene knockouts with no detectable translocation events. In one embodiment, the base editor (e.g., ABE8) is used for multiplex base editing of at least one cell surface target (e.g., including but not limited to TRAC, B2M, CD7, PDCD1, CBLB, and / or CITA) in T cells. In one embodiment, ABE8 is used for multiplex base editing of TRAC, B2M, and CIITA in T cells. Multiplex editing of genes may contribute to creating CAR-T cell therapies with improved therapeutic properties. This approach addresses known limitations of multiplex-edited T cell products and is a promising development towards next-generation cell-based precision therapies.

[0381] Chimeric antigen receptors and CAR-T cells

[0382] The present invention provides immune cells modified with a nucleobase editor expressing a chimeric antigen receptor (CARs) as described herein. Modifying immune cells to express a chimeric antigen receptor can enhance the immune activity of the immune cells, wherein the chimeric antigen receptor has an affinity for an epitope on an antigen, wherein the antigen is related to an adaptive change in an organism. For example, the chimeric antigen receptor can have an affinity for an epitope on a protein expressed in tumor cells. Since CAR-T cells can function independently of the major histocompatibility complex (MHC), activated CAR-T cells can kill tumor cells expressing the antigen. The direct action of CAR-T cells evades tumor cell defense mechanisms that have evolved in response to MHC presentation of antigens to immune cells.

[0383] In some embodiments, the present invention provides immune effector cells expressing a chimeric antigen receptor that targets B cells involved in an autoimmune response (e.g., B cells of an individual expressing antibodies produced against the individual's own tissues).

[0384] Some embodiments include autologous immune cell immunotherapy, in which immune cells are obtained from an individual with a disease or an adaptive change, characterized by cancerous or other altered cells that express surface markers. The obtained immune cells are genetically engineered to express a chimeric antigen receptor and are effectively redirected against a specific antigen. Thus, in some embodiments, the immune cells are obtained from an individual in need of CAR-T immunotherapy. In some embodiments, these autologous immune cells are cultured and modified shortly after being obtained from the individual. In other embodiments, autologous cells are obtained and then stored for future use. This may be desirable for individuals who may be undergoing parallel treatments that will reduce immune cell counts in the future. In allogeneic immune cell immunotherapy, the immune cells can be obtained from a donor other than the individual to be treated. After being modified to express a chimeric antigen receptor, the immune cells are administered to the individual to treat a tumor. In some embodiments, the immune cells to be modified to express a chimeric antigen receptor can be obtained from a pre-existing immune cell stock culture.

[0385] Standard techniques known in the art can be used to isolate or purify immune cells and / or immune effector cells from samples collected from an individual or a donor. For example, immune effector cells can be isolated or purified from a whole blood sample by lysing red blood cells and removing peripheral mononuclear blood cells by centrifugation. Immune effector cells can be further isolated or purified using selective purification methods that separate immune effector cells based on cell-specific markers such as CD25, CD3, CD4, CD8, CD28, CD45RA, or CD45RO. In one embodiment, CD25+ is used as a marker to select regulatory T cells. In another embodiment, the present invention provides T cells having a targeted gene knockout at the TCR constant region (TRAC) responsible for TCRαβ surface expression. TCRαβ-deficient CAR T cells are compatible with allogeneic immunotherapy (Qasim et al., Sci. Transl. Med. 9, eaaj2013 (2017); Valton et al., Mol Ther. 2015 Sep;23(9):1507–1518). If desired, residual TCRαβ T cells can be removed using the CliniMACS bead depletion method to minimize the risk of GVHD. In another embodiment, the present invention provides in vitro selected donor T cells that recognize minor histocompatibility antigens expressed on recipient hematopoietic cells, thereby minimizing the risk of graft-versus-host disease (GVHD), which is a major cause of post-transplant morbidity and mortality (Warren et al., Blood 2010;115(19):3869-3878). Another technique for isolating or purifying immune effector cells is flow cytometry. In fluorescence-activated cell sorting, fluorescently labeled antibodies that have an affinity for immune effector cell markers are used to label immune effector cells in a sample. Gating strategies applicable to cells expressing the marker are used to isolate the cells. For example, T lymphocytes can be separated from other cells in a sample by using fluorescently labeled antibodies specific for immune effector cell markers (such as CD4, CD8, CD28, CD45) and corresponding gating strategies. In one embodiment, a CD45 gating strategy is employed. In some embodiments, gating strategies using other markers specific for immune effector cells are used instead of or in combination with the CD45 gating strategy.

[0386] The immune effector cells contemplated in the present invention are effector T cells. In some embodiments, the effector T cells are naive CD8 + T cells, cytotoxic T cells, or regulatory T (Treg) cells. In some embodiments, the effector T cells are thymocytes, immature T lymphocytes, mature T lymphocytes, resting T lymphocytes, or activated T lymphocytes. In some embodiments, the immune effector cells are CD4 + CD8+ T cells or CD4 - CD8 - T cells. In some embodiments, the immune effector cell is a helper T cell. In some embodiments, the helper T cell is a helper T cell 1 (Th1), a helper T cell 2 (Th2), or a CD4-expressing helper T cell (CD4+ T cell). In some embodiments, the immune effector cell is any other subset of T cells. In addition to the chimeric antigen receptor, the modified immune effector cell may also express an exogenous cytokine, a different chimeric receptor, or any other reagent that enhances the signaling or function of the immune effector cell. For example, co-expression of the chimeric antigen receptor and a cytokine can enhance the ability of CAR-T cells to lyse target cells.

[0387] The chimeric antigen receptor contemplated in the present invention comprises an extracellular binding domain, a transmembrane domain, and an intracellular domain. Binding of an antigen to the extracellular binding domain can activate CAR-T cells and generate an effector response, including CAR-T cell proliferation, cytokine production, and other processes leading to the death of antigen-expressing cells. In some embodiments of the present invention, the chimeric antigen receptor further comprises a linker.

[0388] The extracellular binding domain of the chimeric antigen receptor contemplated herein comprises the amino acid sequence of an antibody or an antigen-binding fragment thereof that has an affinity for a specific antigen. In various embodiments, the CAR specifically binds 5T4. Exemplary anti-5T4 CARs include, but are not limited to, CART-5T4 (Oxford BioMedica plc) and UCART-5T4 (Cellectis SA).

[0389] In various embodiments, the CAR specifically binds alpha-fetoprotein. Exemplary anti-alpha-fetoprotein CARs include, but are not limited to, ET-1402 (Eureka Therapeutics Inc). In various embodiments, the CAR specifically binds Axl. Exemplary anti-Axl CARs include, but are not limited to, CCT-301-38 (F1Oncology Inc). In various embodiments, the CAR specifically binds B7H6. Exemplary anti-B7H6 CARs include, but are not limited to, CYAD-04 (Celyad SA).

[0390] In various embodiments, the CAR specifically binds BCMA. Exemplary anti-BCMA CARs include, but are not limited to, ACTR-087+SEA-BCMA (Seattle Genetics Inc), ALLO-715 (Cellectis SA), ARI-0002 (Institut d'Investigacions Biomediques August Pi I Sunyer), bb-2121 (bluebird bio Inc), bb-21217 (bluebird bio Inc), CART-BCMA (University of Pennsylvania), CT-053 (Carsgen Therapeutics Ltd), Descartes-08 (Cartesian Therapeutics), FCARH-143 (Juno Therapeutics Inc), ICTCAR-032 (Innovative Cellular Therapeutics Co Ltd), IM21CART (Beijing Immunochina Medical Science & Technology Co Ltd), JCARH-125 (Memorial Sloan-Kettering Cancer Center), KITE-585 (Kite Pharma Inc), LCAR-B38M (Nanjing Legend Biotech Co Ltd), LCAR-B4822M (Nanjing Legend Biotech Co Ltd), MCARH-171 (Memorial Sloan-Kettering Cancer Center), P-BCMA-101 (Poseida Therapeutics Inc), P-BCMA-ALLO1 (Poseida Therapeutics Inc), spCART-269 (Shanghai Unicar-Therapy Bio-medicine Technology Co Ltd), and BCMA02 / bb2121 (bluebird bio Inc). The polypeptide sequence of the BCMA02 / bb2121 CAR is as follows:

[0391] MALPVTALLLPLALLLHAARPDIVLTQSPPSLAMSLGKRATISCRASESVTILGSHLIHWYQQKPGQPPTLLIQLASNVQTGVPARFSGSGSRTDFTLTIDPVEEDDVAVYYCLQSRTIPRTFGGGTKLEIKGSTSGSGKPGSGEGSTKGQIQLVQSGPELKKPGETVKISCKASGYTFTDYSINWVKRAPGKGLKWMGWINTETREPAYAYDFRGRFAFSLETSASTAYLQINNLKYEDTATYFCALDYSYAMDYWGQGTSVTVSSAAATTTPAPRPPTPAPTIASQPLSLRPEACRPAAGGAVHTRGLDFACDIYIWAPLAGTCGVLLLSLVITLYCKRGRKKLLYIFKQPFMRPVQTTQEEDGCSCRFPEEEEGGCELRVKFSRSADAPAYQQGQNQLYNELNLGRREEYDVLDKRRGRDPEMGGKPRRKNPQEGLYNELQKDKMAEAYSEIGMKGERRRGKGHDGLYQGLSTATKDTYDALHMQALPPR

[0392] In various embodiments, the CAR specifically binds to CCK2R. Exemplary anti-CCK2R CAR

[0393] includes, but is not limited to, anti-CCK2R CAR-T adaptor molecule (CAM) + anti-FITC CAR T cell therapy (cancer), Endocyte / Purdue (Purdue University).

[0394] In various embodiments, the CAR specifically binds to a CD antigen. Exemplary anti-CD antigen CARs include, but are not limited to, VM-802 (ViroMed Co Ltd). In various embodiments, the CAR specifically binds to CD123. Exemplary anti-CD123 CARs include, but are not limited to, MB-102 (Fortress Biotech Inc), RNACART123 (University of Pennsylvania), SFG-iMC-CD123.zeta (Bellicum Pharmaceuticals Inc), and UCART-123 (Cellectis SA). In various embodiments, the CAR specifically binds to CD133. Exemplary anti-CD133 CARs include, but are not limited to, KD-030 (Nanjing Kaedi Biotech Inc). In various embodiments, the CAR specifically binds to CD138. Exemplary anti-CD138 CARs include, but are not limited to, ATLCAR.CD138 (UNC Lineberger Comprehensive Cancer Center) and CART-138 (Chinese PLA General Hospital). In various embodiments, the CAR specifically binds to CD171. Exemplary anti-CD171 CARs include, but are not limited to, JCAR-023 (Juno Therapeutics Inc). In various embodiments, the CAR specifically binds to CD19. Exemplary anti-CD19 CARs include, but are not limited to, 1928z-41BBL (Memorial Sloan-Kettering Cancer Center), 1928z-E27 (Memorial Sloan-Kettering Cancer Center), 19-28z-T2 (Guangzhou Institutes of Biomedicine and Health), 4G7-CARD (University College London), 4SCAR19 (Shenzhen Institute of Gene Immunology), ALLO-501 (Pfizer Inc), ATA-190 (QIMR Berghofer Medical Research Institute), AUTO-1 (University College London), AVA-008 (Avacta Ltd), axicabtagene ciloleucel (Kite Pharma Inc), BG-T19 (Guangzhou Bio-gene Technology Co Ltd.), BinD-19 (Shenzhen BinDeBio Ltd)), BPX-401 (Bellicum Pharmaceuticals Inc), CAR19h28TM41BBz (Westmead Institute for Medical Research), C-CAR-011 (Chinese PLA General Hospital), CD19CART (Innovative Cellular Therapeutics Co Ltd), CIK-CAR.CD19 (Formula Pharmaceuticals Inc), CLIC-1901 (Ottawa Hospital Research Institute), CSG-CD19 (Carsgen Therapeutics Ltd), CTL-119 (University of Pennsylvania), CTX-101 (CRISPR Therapeutics AG), DSCAR-01 (Shanghai Hengrui Biotechnology Co Ltd), ET-190 (Eureka Therapeutics Inc), FT-819 (Memorial Sloan-Kettering Cancer Center), ICAR-19 (Immune Cell Therapy Inc), IM19CAR-T (Beijing Immunochina Medical Science & Technology Co Ltd), JCAR-014 (Juno Therapeutics Inc), JWCAR-029 (MingJu Therapeutics (Shanghai) Co., Ltd), KD-C-19 (Nanjing Kaedi Biotech Inc), LinCART19 (iCellGene Therapeutics), lisocabtagene maraleucel (Juno Therapeutics Inc), MatchCART (Shanghai Hrain Biotechnology), MB-CART19.1 (Shanghai Children's Medical Center), PBCAR-0191 (PrecisionBioSciences Inc), PCAR-019 (PersonGen Biomedicine (Suzhou) Co Ltd), pCAR-19B (Chongqing Precision Biotech Co Ltd), PZ-01 (Pinze Life Sciences Co Ltd), RB-1916 (Refuge Biotechnologies Inc), SKLB-083019 (Chengdu Galaxy Biopharmaceuticals Co Ltd), spCART-19 (Shanghai Unika-Therapeutic Biomedicine Technology Co Ltd), TBI-1501 (Takara Bio Inc), TC-110 (TCR2 Therapeutics Inc), TI-1007 (Timmune Biotech Inc), tisagenlecleucel (Abramson Cancer Center of the University of Pennsylvania), U-CART (Shanghai Bioray Laboratory Inc), UCART-19 (Wugen Inc), UCART-19 (Cellectis SA), vadacabtagene leraleucel (Memorial Sloan-Kettering Cancer Center), XLCART-001 (Nanjing Medical University), and yinnuokati-19 (Shenzhen Innovation Immunotechnology Co Ltd). In various embodiments, the CAR specifically binds to CD2. Exemplary anti-CD2 CARs include, but are not limited to, UCART-2 (Wugen Inc). In various embodiments, the CAR specifically binds to CD20. Exemplary anti-CD20 CARs include, but are not limited to, ACTR-087 (National University of Singapore), ACTR-707 (Unum Therapeutics Inc), CBM-C20.1 (Chinese PLA General Hospital), MB-106 (Fred Hutchinson Cancer Research Center), and MB-CART20.1 (Miltenyi Biotec GmbH).

[0395] In various embodiments, the CAR specifically binds CD22. Exemplary anti-CD22 CARs include, but are not limited to, anti-CD22 CAR T cell therapy (B-cell acute lymphoblastic leukemia), University of Pennsylvania, CD22-CART (Shanghai Unika Therapeutic Biomedicine Technology Co., Ltd.), JCAR-018 (Opus Bio Inc.), MendCART (Shanghai Hengrun Biotechnology Co., Ltd.), and UCART-22 (Cellectis SA). In various embodiments, the CAR specifically binds CD30. Exemplary anti-CD30 CARs include, but are not limited to, ATLCAR.CD30 (UNC Lineberger Comprehensive Cancer Center), CBM-C30.1 (Chinese PLA General Hospital), and Hu30-CD28zeta (National Cancer Institute). In various embodiments, the CAR specifically binds CD33. Exemplary anti-CD33 CARs include, but are not limited to, anti-CD33 CAR γδ T cell therapy (acute myeloid leukemia), TC BioPharm / University College London, CAR33VH (Opus Bio Inc.), CART-33 (Chinese PLA General Hospital), CIK-CAR.CD33 (Formula Pharmaceuticals Inc.), UCART-33 (Cellectis SA), and VOR-33 (Columbia University).

[0396] In various embodiments, the CAR specifically binds to CD38. Exemplary anti-CD38 CARs include, but are not limited to, UCART-38 (Cellectis SA). In various embodiments, the CAR specifically binds to CD38A2. Exemplary anti-CD38A2 CARs include, but are not limited to, T-007 (TNK Therapeutics Inc). In various embodiments, the CAR specifically binds to CD4. Exemplary anti-CD4 CARs include, but are not limited to, CD4CAR (iCell Gene Therapeutics). In various embodiments, the CAR specifically binds to CD44. Exemplary anti-CD44 CARs include, but are not limited to, CAR-CD44v6 (Istituto Scientifico H SanRaffaele). In various embodiments, the CAR specifically binds to CD5. Exemplary anti-CD5 CARs include, but are not limited to, CD5CAR (iCell Gene Therapeutics). In various embodiments, the CAR specifically binds to CD7. Exemplary anti-CD7 CARs include, but are not limited to, CAR-pNK (PersonGen Biomedicine (Suzhou) Co Ltd) and CD7.CAR / 28zeta CAR T cells (Baylor College of Medicine), UCART7 (Washington University in St Louis).

[0397] In various embodiments, the CAR specifically binds to CDH17. Exemplary anti-CDH17 CARs include, but are not limited to, ARB-001.T (Arbele Ltd). In various embodiments, the CAR specifically binds to CEA. Exemplary anti-CEA CARs include, but are not limited to, HORC-020 (HumOrigin Inc). In various embodiments, the CAR specifically binds to the chimeric TGF-β receptor (CTBR). Exemplary anti-chimeric TGF-β receptor (CTBR) CARs include, but are not limited to, CAR-CTBR T cells (bluebirdbio Inc). In various embodiments, the CAR specifically binds to Claudin18.2. Exemplary anti-Claudin18.2 CARs include, but are not limited to, CAR-CLD18 T cells (Carsgen Therapeutics Ltd) and KD-022 (Nanjing Kaedi Biotech Inc).

[0398] In various embodiments, the CAR specifically binds to CLL1. Exemplary anti-CLL1 CARs include, but are not limited to, KITE-796 (Kite Pharma Inc). In various embodiments, the CAR specifically binds to DLL3. Exemplary anti-DLL3 CARs include, but are not limited to, AMG-119 (Amgen Inc). In various embodiments, the CAR specifically binds to dual BCMA / TACI (APRIL). Exemplary anti-dual BCMA / TACI (APRIL) CARs include, but are not limited to, AUTO-2 (Autolus Therapeutics Limited). In various embodiments, the CAR specifically binds to dual CD19 / CD22. Exemplary anti-dual CD19 / CD22 CARs include, but are not limited to, AUTO-3 (Autolus Therapeutics Limited) and LCAR-L10D (Nanjing Legend Biotech Co Ltd). In various embodiments, the CAR specifically binds to CD19. In various embodiments, the CAR specifically binds to dual CLL1 / CD33. Exemplary anti-dual CLL1 / CD33 CARs include, but are not limited to, ICG-136 (iCell Gene Therapeutics). In various embodiments, the CAR specifically binds to dual EpCAM / CD3. Exemplary anti-dual EpCAM / CD3 CARs include, but are not limited to, IKT-701 (IcellKealex Therapeutics). In various embodiments, the CAR specifically binds to Dual ErbB / 4ab. Exemplary anti-dual ErbB / 4ab CARs include, but are not limited to, LEU-001 (King's College London). In various embodiments, the CAR specifically binds to dual FAP / CD3. Exemplary anti-dual FAP / CD3 CARs include, but are not limited to, IKT-702 (Icell KealexTherapeutics). In various embodiments, the CAR specifically binds to EBV. Exemplary anti-EBV CARs include, but are not limited to, TT-18 (Tessa Therapeutics Pte Ltd).

[0399] In various embodiments, the CAR specifically binds to EGFR. Exemplary anti-EGFR CARs include, but are not limited to, anti-EGFR CAR T cell therapy (CBLB MegaTAL, cancer), bluebird bio (bluebird bio Inc), anti-EGFR CAR T cell therapy expressing CTLA-4 checkpoint inhibitor + PD-1 checkpoint inhibitor monoclonal antibody (EGFR-positive advanced solid tumors), Shanghai Institute of Cellular Therapy (Shanghai Institute of Cellular Therapy), CSG-EGFR (Carsgen Therapeutics Ltd), and EGFR-IL12-CART (Pregene (Shenzhen) Biotechnology Co., Ltd).

[0400] In various embodiments, the CAR specifically binds to EGFRvIII. Exemplary anti-EGFRvIII CARs include, but are not limited to, KD-035 (Nanjing Kaedi Biotech Inc) and UCART-EgfrVIII (Cellectis SA). In various embodiments, the CAR specifically binds to Flt3. Exemplary anti-Flt3 CARs include, but are not limited to, ALLO-819 (Pfizer Inc) and AMG-553 (Amgen Inc). In various embodiments, the CAR specifically binds to the folate receptor. Exemplary anti-folate receptor CARs include, but are not limited to, EC17 / CAR T (Endocyte Inc). In various embodiments, the CAR specifically binds to G250. Exemplary anti-G250 CARs include, but are not limited to, autologous T lymphocyte therapy (G250-scFV transduced renal cell carcinoma), Erasmus Medical Center (Daniel den Hoed Cancer Center).

[0401] In various embodiments, the CAR specifically binds to GD2. Exemplary anti-GD2 CARs include, but are not limited to, 1RG-CART (University College London), 4SCAR-GD2 (Shenzhen Institute of Gene Immunology), C7R-GD2.CART cells (Baylor College of Medicine), CMD-501 (Baylor College of Medicine), CSG-GD2 (Carsgen Therapeutics Ltd), GD2-CART01 (Bambino Gesu Hospital and Research Institute), GINAKIT cells (Baylor College of Medicine), iC9-GD2-CAR-IL-15 T cells (UNC Lineberger Comprehensive Cancer Center), and IKT-703 (Icell Kealex Therapeutics). In various embodiments, the CAR specifically binds to GD2 and MUC1. Exemplary anti-GD2 / MUC1 CARs include, but are not limited to, PSMACAR-T (University of Pennsylvania).

[0402] In various embodiments, the CAR specifically binds to GPC3. Exemplary anti-GPC3 CARs include, but are not limited to, ARB-002.T (Arbele Ltd), CSG-GPC3 (Carsgen Therapeutics Ltd), GLYCAR (Baylor College of Medicine), and TT-14 (Tessa Therapeutics Pte Ltd). In various embodiments, the CAR specifically binds to Her2. Exemplary anti-Her2 CARs include, but are not limited to, ACTR-087 + trastuzumab (Unum Therapeutics Inc), ACTR-707 + trastuzumab (Unum Therapeutics Inc), CIDeCAR (Bellicum Pharmaceuticals Inc), MB-103 (Mustang Bio Inc), RB-H21 (Refuge Biotechnologies Inc), and TT-16 (Baylor College of Medicine). In various embodiments, the CAR specifically binds to IL13R. Exemplary anti-IL13R CARs include, but are not limited to, MB-101 (City of Hope) and YYB-103 (YooYoung Pharmaceuticals Co Ltd). In various embodiments, the CAR specifically binds to integrin β-7. Exemplary anti-integrin β-7 CARs include, but are not limited to, the MMG49 CAR T cell therapy (Osaka University). In various embodiments, the CAR specifically binds to the LC antigen. Exemplary anti-LC antigen CARs include, but are not limited to, VM-803 (ViroMed Co Ltd) and VM-804 (ViroMed Co Ltd).

[0403] In various embodiments, the CAR specifically binds to mesothelin. Exemplary anti-mesothelin CARs include, but are not limited to, CARMA-hMeso (Johns Hopkins University), CSG-MESO (Carsgen Therapeutics Ltd), iCasp9M28z (Memorial Sloan-Kettering Cancer Center), KD-021 (Nanjing Kaedi Biotech Inc), m-28z-T2 (Guangzhou Institute of Biomedicine and Health), MesoCART (University of Pennsylvania), meso-CAR-T+PD-78 (MirImmune LLC), RB-M1 (Refuge Biotechnologies Inc), and TC-210 (TCR2 Therapeutics Inc).

[0404] In various embodiments, the CAR specifically binds to MUC1. Exemplary anti-MUC1 CARs include, but are not limited to, anti-MUC1 CAR T cell therapy + PD-1 knockout T cell therapy (esophageal cancer / NSCLC), Guangzhou Anjie Biomedical Technology / Sydney University of Technology (Guangzhou Anjie Biomedical Technology Co., Ltd), ICTCAR-043 (Innovative Cellular Therapeutics Co Ltd), ICTCAR-046 (Innovative Cellular Therapeutics Co Ltd), P-MUC1C-101 (Poseida Therapeutics Inc), and TAB-28z (OncoTab Inc). In various embodiments, the CAR specifically binds to MUC16. Exemplary anti-MUC16 CARs include, but are not limited to, 4H1128Z-E27 (Eureka Therapeutics Inc) and JCAR-020 (Memorial Sloan-Kettering Cancer Center).

[0405] In various embodiments, the CAR specifically binds to nfP2X7. Exemplary anti-nfP2X7 CARs include, but are not limited to, BIL-022c (Biosceptre International Ltd). In various embodiments, the CAR specifically binds to PSCA. Exemplary anti-PSCA CARs include, but are not limited to, BPX-601 (Bellicum Pharmaceuticals Inc). In various embodiments, the CAR specifically binds to PSMA. CIK-CAR.PSMA (Formula Pharmaceuticals Inc) and P-PSMA-101 (Poseida Therapeutics Inc). In various embodiments, the CAR specifically binds to ROR1. Exemplary anti-ROR1 CARs include, but are not limited to, JCAR-024 (Fred Hutchinson Cancer Research Center). In various embodiments, the CAR specifically binds to ROR2. Exemplary anti-ROR2 CARs include, but are not limited to, CCT-301-59 (F1OncologyInc). In various embodiments, the CAR specifically binds to SLAMF7. Exemplary anti-SLAMF7 CARs include, but are not limited to, UCART-CS1 (Cellectis SA). In various embodiments, the CAR specifically binds to TRBC1. Exemplary anti-TRBC1 CARs include, but are not limited to, AUTO-4 (Autolus Therapeutics Limited). In various embodiments, the CAR specifically binds to TRBC2. Exemplary anti-TRBC2 CARs include, but are not limited to, AUTO-5 (Autolus Therapeutics Limited). In various embodiments, the CAR specifically binds to TSHR. Exemplary anti-TSHR CARs include, but are not limited to, ICTCAT-023 (Innovative Cellular Therapeutics Co Ltd). In various embodiments, the CAR specifically binds to VEGFR-1. Exemplary anti-VEGFR-1 CARs include, but are not limited to, SKLB-083017 (Sichuan University).

[0406] In various embodiments, the CAR is AT-101 (AbClon Inc); AU-101, AU-105, and AU-180 (Aurora Biopharma Inc); CARMA-0508 (Carisma Therapeutics); CAR-T (Fate Therapeutics Inc); CAR-T (Cell Design Labs Inc); CM-CX1 (Celdara Medical LLC); CMD-502, CMD-503, and CMD-504 (Baylor College of Medicine); CSG-002 and CSG-005 (Carsgen Therapeutics Ltd); ET-1501, ET-1502, and ET-1504 (Eureka Therapeutics Inc); FT-61314 (Fate Therapeutics Inc); GB-7001 (Shanghai GeneChem Co., Ltd); IMA-201 (Immatics Biotechnologies GmbH); IMM-005 and IMM-039 (Immunome Inc); ImmuniCAR (TC BioPharm Ltd); NT-0004 and NT-0009 (BioNTech Cell and Gene Therapies GmbH), OGD-203 (OGD2Pharma SAS), PMC-005B (PharmAbcine), and TI-7007 (Timmune Biotech Inc).

[0407] In some embodiments, the chimeric antigen receptor comprises the amino acid sequence of an antibody. In some embodiments, the chimeric antigen receptor comprises the amino acid sequence of an antigen-binding fragment of an antibody. The antibody (or fragment thereof) portion of the extracellular binding domain recognizes and binds to an epitope of an antigen. In some embodiments, the antibody fragment portion of the chimeric antigen receptor is a single-chain variable fragment (scFv). The scFV comprises the light and variable fragments of a monoclonal antibody. In other embodiments, the antibody fragment portion of the chimeric antigen receptor is a multi-chain variable fragment, which may comprise more than one extracellular binding domain and thus bind more than one antigen simultaneously. In multi-chain variable fragment embodiments, a hinge region may separate the different variable fragments, providing the necessary spatial arrangement and flexibility.

[0408] In other embodiments, the antibody portion of the chimeric antigen receptor comprises at least one heavy chain and at least one light chain. In some embodiments, the antibody portion of the chimeric antigen receptor comprises two heavy chains and two light chains linked by disulfide bonds, wherein each light chain is linked to one of the heavy chains by a disulfide bridge. In some embodiments, the light chain comprises a constant region and a variable region. The complementarity-determining regions located in the variable region of the antibody are responsible for the affinity of the antibody for a specific antigen. Thus, antibodies that recognize different antigens contain different complementarity-determining regions. The complementarity-determining regions are located in the variable domain of the extracellular binding domain, and the variable domains (i.e., variable heavy chain and variable light chain) can be linked to a linker or, in some embodiments, by a disulfide bond.

[0409] In some embodiments, the antigen recognized and bound by the extracellular domain is a protein or peptide, nucleic acid, lipid, or polysaccharide. The antigen can be heterologous, such as an antigen expressed in a pathogenic bacterium or virus. The antigen can also be synthetic; for example, some people are extremely allergic to synthetic latex, and exposure to this antigen can result in an extreme immune response. In some embodiments, the antigen is autologous and is expressed on diseased or otherwise altered cells. For example, in some embodiments, the antigen is expressed in tumor cells. In some embodiments, the tumor cells are solid tumor cells. In other embodiments, the tumor cells are blood cancers, such as B cell cancers. In some embodiments, the B cell cancer is lymphoma (such as Hodgkin lymphoma or non-Hodgkin lymphoma) or leukemia (such as B cell acute lymphoblastic leukemia). Exemplary B cell lymphomas include diffuse large B cell lymphoma (DLBCL), primary mediastinal B cell lymphoma, follicular lymphoma, chronic lymphocytic leukemia (CLL), small lymphocytic lymphoma (SLL), mantle cell lymphoma, marginal zone lymphoma, Burkitt lymphoma, Burkitt-like lymphoma, lymphoplasmacytic lymphoma (Waldenstrom macroglobulinemia), and hairy cell leukemia. In some embodiments, the B cell cancer is multiple myeloma.

[0410] The antibody-antigen interaction is a non-covalent interaction caused by hydrogen bonds, electrostatic or hydrophobic interactions, or van der Waals forces. The affinity of the extracellular binding domain of the chimeric antigen receptor for the antigen can be calculated by the following formula:

[0411] KA = [antibody-antigen] / [antibody][antigen], where

[0412] [Ab] = the molar concentration of unoccupied binding sites on the antibody;

[0413] [Ag] = the molar concentration of unoccupied binding sites on the antigen; and

[0414] [Ab-Ag] = the molar concentration of the antibody-antigen complex.

[0415] Antibody-antigen interactions can also be characterized based on the dissociation of the antigen from the antibody. The dissociation constant (KD) is the ratio of the association rate to the dissociation rate and is inversely proportional to the affinity constant. Thus, KD = 1 / KA. Those skilled in the art will be familiar with these concepts and will know that traditional methods, such as ELISA assays, can be used to calculate these constants.

[0416] The transmembrane domain of the chimeric antigen receptor described herein spans the lipid bilayer cell membrane of the CAR-T cell and separates the extracellular binding domain and the intracellular signaling domain. In some embodiments, the domain is derived from other receptors having a transmembrane domain, while in other embodiments, the domain is synthetic. In some embodiments, the transmembrane domain can be derived from a non-human transmembrane domain and, in some embodiments, humanized. "Humanized" means optimizing the nucleic acid sequence encoding the transmembrane domain so that it is expressed more reliably or efficiently in a human individual. In some embodiments, the transmembrane domain is derived from another transmembrane protein expressed in human immune effector cells. Examples of such proteins include, but are not limited to, subunits of the T cell receptor (TCR) complex, PD1, or any cluster of differentiation proteins or other proteins that are expressed in immune effector cells and have a transmembrane domain. In some embodiments, the transmembrane domain will be synthetic, and such sequences will contain a number of hydrophobic residues.

[0417] In some embodiments, the chimeric antigen receptor is designed to include a spacer sequence between the transmembrane domain and the extracellular domain, the intracellular domain, or both. The length of such spacer sequences can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids. In some embodiments, the length of the linker can be 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acids. In other embodiments, the length of the spacer sequence can be between 100 and 500 amino acids. The spacer sequence can be any polypeptide that links one domain to another and is used to position such linked domains to enhance or optimize the function of the chimeric antigen receptor.

[0418] The intracellular signaling domain of the chimeric antigen receptor considered in this article includes a primary signaling domain. In some embodiments, the chimeric antigen receptor includes a primary signaling domain and a secondary or co-stimulatory signaling domain. In some embodiments, the domain includes one or more tyrosine-based immunoreceptor activation motifs or ITAMs. In some embodiments, the primary signaling domain includes more than one ITAM. The ITAMs incorporated into the chimeric antigen receptor can be derived from the ITAMs of other cell receptors. In some embodiments, the primary signaling domain containing an ITAM can be derived from subunits of the TCR complex, such as CD3γ, CD3ε, CD3ζ, or CD3δ (see Figure 1A ). In some embodiments, the primary signaling domain containing an ITAM can be derived from FcRγ, FcRβ, CD5, CD22, CD79a, CD79b, or CD66d. In some embodiments, the secondary signaling domain is derived from CD28. In other embodiments, the secondary signaling domain is derived from CD2, CD4, CDS, CD8α, CD83, CD134, CD137, ICOS, or CD154.

[0419] This article also provides nucleic acids encoding the chimeric antigen receptors described herein. In some embodiments, the nucleic acids are isolated or purified. Ex vivo delivery of the nucleic acids can be accomplished using methods known in the art. For example, immune cells obtained from an individual can be transformed with a nucleic acid vector encoding the chimeric antigen receptor. The recipient immune cells can then be transformed using the vector, such that these cells express the chimeric antigen receptor. Effective methods for transforming immune cells include transfection and transduction. Such methods are well known in the art. For example, suitable methods for delivering nucleic acid molecules encoding chimeric antigen receptors (and nucleic acids encoding base editors) can be found in International Patent Application No. PCT / US2009 / 040040 and U.S. Patent Nos. 8,450,112; 9,132,153; and 9,669,058, each of which is incorporated herein by reference in its entirety. In addition, those methods and vectors described herein for delivering nucleic acids encoding base editors (e.g., ABE8) are applicable to delivering nucleic acids encoding chimeric antigen receptors.

[0420] Some aspects of the present invention provide immune cells comprising a chimeric antigen and an altered endogenous gene, which enhance immune cell function, resistance to immune suppression or inhibition, or a combination thereof. Allogeneic immune cells expressing an endogenous immune cell receptor as well as a chimeric antigen receptor can recognize and attack host cells, a condition known as graft-versus-host disease (GVHD). The α component of the immune cell receptor complex is encoded by the TRAC gene, and in some embodiments, the gene is edited such that the α subunit of the TCR complex is non-functional or absent. Since the subunit is required for endogenous immune cell signaling, editing the gene can reduce the risk of graft-versus-host disease caused by allogeneic immune cells.

[0421] Host immune cells can potentially recognize allogeneic CAR-T cells as non-self cells and initiate an immune response to eliminate the non-self cells. B2M is expressed in almost all nucleated cells and is associated with the MHC class I complex ( Figure 1B ). Circulating host CD8 + T cells can recognize this B2M protein as non-self and kill allogeneic cells. To overcome this transplant rejection, in some embodiments, the B2M gene is edited to knockout or knockdown expression.

[0422] In some embodiments of the present invention, the PDCD1 gene is edited in CAR-T cells to knockout or knockdown expression. The PDCD1 gene encodes the cell surface receptor PD-1, an immune system checkpoint expressed in immune cells, which is involved in reducing autoimmunity by promoting apoptosis of antigen-specific immune cells. By knocking out or knocking down the expression of the PDCD1 gene, the modified CAR-T cells are less likely to undergo apoptosis, more likely to proliferate, and can escape the programmed cell death immune checkpoint.

[0423] The CBLB gene encodes an E3 ubiquitin ligase that plays an important role in suppressing the activation of immune effector cells. As shown in Figure 1C , the CBLB protein favors the signaling pathway leading to immune effector cell tolerance and actively inhibits the signal transduction leading to immune effector cell activation. Since immune effector cell activation is required for the in vivo proliferation of CAR-T cells after transplantation, in some embodiments of the present invention, CBLB is edited to knockout or knockdown expression.

[0424] In some embodiments, gene editing can be performed in immune cells to enhance the function of immune cells or reduce immunosuppression or inhibition before the cells are transformed to express a chimeric antigen receptor. In other aspects, gene editing can be performed in CAR-T cells to enhance the function of immune cells or reduce immunosuppression or inhibition, i.e., after the immune cells are transformed to express a chimeric antigen receptor. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TRAC, B2M, PDCD1, CD7, CIITA, CBLB gene or a combination thereof, wherein the expression of the edited gene is knocked out or knocked down.

[0425] In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TRAC gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TRAC gene and one or more of the B2M, PDCD1, CD7, CIITA, and / or CBLB genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TRAC gene and the B2M gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TRAC gene and the PDCD1 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TRAC gene and the CBLB gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TRAC gene and the CD7 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TRAC gene and the CIITA gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, and PDCD1 genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, PDCD1, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, and CIITA genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, and CD7 genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, PDCD1, and CD7 genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, PDCD1, and CIITA genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, PDCD1, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, CD7, and CIITA genes, wherein the expression of the edited gene is knocked out or knocked down.In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, CD7, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, CIITA, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down.

[0426] In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, PDCD1, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, PDCD1, and CD7 genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, CD7, and CIITA genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, CD7, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, PDCD1, and CIITA genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, CBLB, and CIITA genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, PDCD1, CD7, and CIITA genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, PDCD1, CD7, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, PDCD1, CIITA, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, CIITA, CD7, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down.

[0427] In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, PDCD1, CD7, and CIITA genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, PDCD1, CD7, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, CD7, CIITA, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, PDCD1, CIITA, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, PDCD1, CD7, CIITA, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited TRAC, B2M, PDCD1, CD7, and CBLB genes, wherein the expression of the edited genes is knocked out or knocked down.

[0428] In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited B2M gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited B2M gene and one or more of the CBLB, PDCD1, CD7, CIITA, and / or TRAC genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited B2M gene and a PDCD1 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited B2M gene and a CBLB gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited B2M gene and a CIITA gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited B2M gene and a CD7 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, CIITA, and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, PDCD1, and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, PDCD1, and CIITA genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, CD7, and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, CD7, and PDCD1 genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, CD7, and CIITA genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, PDCD1, CIITA, and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, PDCD1, CIITA, and CD7 genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, PDCD1, CD7, and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited B2M, PDCD1, CD7, CIITA, and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down.

[0429] In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited PDCD1 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited PDCD1 gene and one or more of the B2M, CBLB, CD7, CIITA, and / or TRAC genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited PDCD1 and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited PDCD1 and CD7 genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited PDCD1 and CIITA genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited PDCD1, CIITA, and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down.

[0430] In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CD7, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CBLB, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CD7 and CIITA genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CD7 and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CD7, PDCD1, and CIITA genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CD7, PDCD1, CIITA, and CBLB genes, wherein the expression of the edited gene is knocked out or knocked down.

[0431] In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CBLB, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CBLB gene and one or more of the B2M, PDCD1, CD7, CIITA, and / or TRAC genes, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CBLB and CIITA genes, wherein the expression of the edited gene is knocked out or knocked down.

[0432] In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CIITA, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited CBLB gene and one or more of the B2M, PDCD1, CD7, CBLB, and / or TRAC genes, wherein the expression of the edited gene is knocked out or knocked down.

[0433] In some embodiments, immune cells, including but not limited to any immune cells comprising any of the above gene edits, can be edited to generate mutations in other genes, thereby enhancing the function of CAR-T or reducing immunosuppression or inhibition of the cells. For example, in some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2, ZAP70, NFATc1, TET2 gene, or a combination thereof, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2 and ZAP70 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2 and ZAP70 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2 and NFATC1 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2 and TET2 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2, ZAP70, and NFATC1 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2, ZAP70, and TET2 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2, NFATC1, and TET2 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited TGFBR2, ZAP70, NFATC1, and TET2 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited ZAP70 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited ZAP70 and NFATC1 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited ZAP70 and TET2 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited ZAP70, PDCD1, and TET2 gene, wherein the expression of the edited gene is knocked out or knocked down.In some embodiments, the immune cells comprise a chimeric antigen receptor and an edited PDCD1 gene, wherein the expression of the edited gene is knocked out or knocked down. In some embodiments, the immune cells comprise a chimeric antigen receptor and edited PDCD1 as well as TET2 genes, wherein the expression of the edited genes is knocked out or knocked down. And in some embodiments, the immune cells comprise a chimeric antigen receptor and edited TET2, wherein the expression of the edited gene is knocked out or knocked down.

[0434] In some embodiments, the chimeric antigen receptor is inserted into the TRAC gene. This has advantages. First, since TRAC is highly expressed in immune cells, when the construct is designed to insert the chimeric antigen receptor into the TRAC gene, the chimeric antigen receptor will be similarly expressed, such that the expression of the receptor is driven by the TRAC promoter. Second, inserting the chimeric antigen receptor into the TRAC gene will knock out TRAC expression. In some embodiments, the gene editing systems described herein can be used to insert a chimeric antigen receptor into the TRAC locus. A gRNA specific to the TRAC locus can direct the gene editing system to the locus and initiate double-stranded DNA cleavage. In certain embodiments, the gRNA is used in combination with Cas12b. In various embodiments, the gene editing system is used in combination with a nucleic acid having a sequence encoding a CAR receptor. Exemplary guide RNAs are provided in Table 1A below.

[0435] Table 1A: TRAC guide RNAs

[0436]

[0437] A DNA construct encoding a chimeric antigen receptor and a nucleic acid, which comprises an extended TRAC DNA fragment flanking the gRNA targeting sequence. Without being bound by theory, the construct binds to the complementary TRAC sequence and then inserts the chimeric antigen receptor DNA near the TRAC sequence on the construct into the lesion site, effectively knocking out the TRAC gene and knocking out in the chimeric antigen receptor nucleic acid. Table 1B provides guide RNAs for the TRAC gene that can direct a base editing mechanism to the TRAC locus, enabling the insertion of a chimeric antigen receptor nucleic acid. The first 11 gRNAs are for the BhCas12b nuclease. The second set of 11 are for the BvCas12b nuclease. These are all for inserting a CAR into TRAC by creating a double-strand break, rather than for base editing.

[0438] Table 1B: TRAC guide RNAs

[0439]

[0440] In some embodiments, ABE8 can be used to target the nucleic acid encoding the chimeric antigen receptor of the present invention to the TRAC locus. In some embodiments, the chimeric antigen receptor targets the TRAC locus using the CRISPR / Cas9 base editing system. To generate the above gene editing, immune cells are collected from a subject and contacted with two or more guide RNAs and a nucleobase editor polypeptide comprising a nucleic acid programmable DNA binding protein (napDNAbp) and an adenosine deaminase (such as TadA*8). In some embodiments, the collected immune cells are contacted with at least one nucleic acid, wherein the at least one nucleic acid encodes two or more guide RNAs and a nucleobase editor polypeptide comprising a nucleic acid programmable DNA binding protein (napDNAbp) and an adenosine deaminase. In some embodiments, the gRNA comprises nucleotide analogs. These nucleotide analogs can inhibit the degradation of the gRNA during cellular processes. Table 2 provides target sequences for the gRNA.

[0441] Table 2: Exemplary target sequences

[0442]

[0443]

[0444]

[0445] The adenosine deaminase nucleobase editors (such as ABE8) used in the present invention can act on DNA, including single-stranded DNA. Methods for generating modifications in the target nucleobase sequences in immune cells using them are described. In certain embodiments, the fusion proteins provided herein comprise one or more features that improve the base editing activity of the fusion proteins. For example, any fusion protein provided herein can comprise a Cas9 domain with reduced nuclease activity. In some embodiments, any fusion protein provided herein can have a Cas9 domain that does not have nuclease activity (dCas9), or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule, called Cas9 nickase (nCas9). Without being bound by any particular theory, the presence of catalytic residues (e.g., H840) maintains the activity of Cas9 to cleave the non-edited (e.g., non-methylated) strand opposite the targeted nucleobase. Mutations in catalytic residues (e.g., D10 to A10) can prevent the cleavage of the edited strand containing the targeted A residue. Such Cas9 variants can generate single-stranded DNA breaks (nicks) at specific positions according to the target sequence defined by the gRNA, thereby repairing the non-edited strand and ultimately resulting in a change in the nucleobase on the non-edited strand.

[0446] Nucleobase editor

[0447] The present disclosure relates to base editors or nucleobase editors for editing, modifying or altering the target nucleotide sequence of a polynucleotide. Nucleobase editors or base editors are described herein that comprise a polynucleotide programmable nucleotide binding domain and a nucleobase editing domain (e.g., an adenosine deaminase). The polynucleotide programmable nucleotide binding domain, when bound to a bound guide polynucleotide (e.g., gRNA), can specifically bind to a target polynucleotide sequence (i.e., through complementary base pairing sequences between the bases of the bound guide nucleic acid and the bases of the target polynucleotide), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.

[0448] Polynucleotide programmable nucleotide binding domain

[0449] It should be understood that the polynucleotide programmable nucleotide binding domain may also include a nucleic acid programmable protein that binds RNA. For example, the polynucleotide programmable nucleotide binding domain may be associated with a nucleic acid that guides the polynucleotide programmable nucleotide binding domain to RNA. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they are not specifically listed in the present disclosure.

[0450] The polynucleotide programmable nucleotide binding domain of the base editor itself may comprise one or more domains. For example, the polynucleotide programmable nucleotide binding domain may comprise one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain may comprise an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to a protein or polypeptide capable of digesting nucleic acids (e.g., RNA or DNA) from free ends, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) within a nucleic acid (e.g., DNA or RNA). In some embodiments, the endonuclease may cleave a single strand of a double-stranded nucleic acid. In some embodiments, the endonuclease may cleave both strands of a double-stranded nucleic acid. In some embodiments the polynucleotide programmable nucleotide binding domain may be a deoxyribonuclease. In some embodiments the polynucleotide programmable nucleotide binding domain may be a ribonuclease.

[0451] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide-binding domain can cleave zero, one, or both strands of a target polynucleotide. In some embodiments, the polynucleotide programmable nucleotide-binding domain can comprise a nickase domain. As used herein, the term "nickase" refers to a polynucleotide programmable nucleotide-binding domain comprising a nuclease domain that is capable of cleaving only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide programmable nucleotide-binding domain by introducing one or more mutations into the active polynucleotide programmable nucleotide-binding domain. For example, in the case where the polynucleotide programmable nucleotide-binding domain comprises a nickase domain derived from Cas9, the nickase domain derived from Cas9 can comprise a D10A mutation and histidine at position 840. In such embodiments, residue H840 retains catalytic activity and can thus cleave the single strand of the nucleic acid duplex. In another example, the nickase domain derived from Cas9 can comprise an H840A mutation, while the amino acid residue at position 10 remains D. In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide programmable nucleotide-binding domain by removing all or part of the nuclease domain that is not required for nickase activity. For example, in the case where the polynucleotide programmable nucleotide-binding domain comprises a nickase domain derived from Cas9, the nickase domain derived from Cas9 can comprise a complete or partial deletion of the RuvC domain or the HNH domain.

[0452] The amino acid sequence of an exemplary catalytically active Cas9 is as follows:

[0453]

[0454] Base editors comprising a polynucleotide programmable nucleotide-binding domain that includes a nickase domain are thus capable of creating a single-stranded DNA break (nick) at a specific polynucleotide target sequence (e.g., determined by the complementary sequence of the bound guide nucleic acid). In some embodiments, the strand of the nucleic acid duplex target polynucleotide sequence that is nicked by the base editor comprising the nickase domain (e.g., a Cas9-derived nickase domain) is the strand that is not edited by the base editor (i.e., the strand that is nicked by the base editor is opposite the strand that contains the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain) can nick the strand of the DNA molecule that is targeted for editing. In such embodiments, the non-targeted strand is not nicked.

[0455] Also provided herein are base editors that comprise a catalytically dead polynucleotide programmable nucleotide-binding domain (i.e., that is unable to cleave the target polynucleotide sequence). As used herein, the terms “catalytically dead” and “nuclease dead” are used interchangeably and refer to a polynucleotide programmable nucleotide-binding domain that has one or more mutations and / or deletions that result in its inability to cleave a nucleic acid strand. In some embodiments, a catalytically dead polynucleotide programmable nucleotide-binding domain base editor may lack nuclease activity due to specific point mutations in one or more nuclease domains. For example, where the base editor comprises a Cas9 domain, the Cas9 can comprise D10A and H840A mutations. Such mutations inactivate both nuclease domains, resulting in a loss of nuclease activity. In other embodiments, a catalytically dead polynucleotide programmable nucleotide-binding domain can comprise one or more deletions of all or part of a catalytic domain (e.g., the RuvC1 and / or HNH domains). In further embodiments, a catalytically dead polynucleotide programmable nucleotide-binding domain comprises a point mutation (e.g., D10A or H840A) as well as a deletion of all or part of a nuclease domain.

[0456] The present disclosure also contemplates mutations that can generate catalytically dead polynucleotide programmable nucleotide binding domains from previous functional versions of polynucleotide programmable nucleotide binding domains. For example, in the case of catalytically dead Cas9 (“dCas9”), variants are provided that have mutations other than D10A and H840A, which result in nuclease-inactivated Cas9. Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the Cas9 nuclease domain (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). Based on the present disclosure and knowledge in the art, other suitable nuclease-inactive dCas9 domains will be apparent to those of skill in the art and are within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013;31(9):833-8338, the entire contents of which are incorporated by reference).

[0457] Non-limiting examples of polynucleotide programmable nucleotide binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction nucleases, meganucleases, transcription activator-like effector nucleases (TALENs), and zinc finger nucleases (ZFNs). In some embodiments, a base editor comprises a polynucleotide programmable nucleotide binding domain that comprises a native or modified protein or portion thereof that, through an associated guide nucleic acid, is capable of binding to a nucleic acid sequence during CRISPR (i.e., clustered regularly interspaced short palindromic repeats)-mediated nucleic acid modification. Such a protein is referred to herein as a "CRISPR protein." Accordingly, base editors comprising a polynucleotide programmable nucleotide binding domain that comprises all or a portion of a CRISPR protein are disclosed herein (i.e., base editors that comprise all or a portion of a CRISPR protein as a domain, also referred to as a "derived domain" of a "CRISPR protein" base editor). The CRISPR protein-derived domain incorporated into the base editor can be modified compared to the wild-type or native version of the CRISPR protein. For example, as described below, the CRISPR protein-derived domain can comprise one or more mutations, insertions, deletions, rearrangements, and / or recombinations relative to the wild-type or native form of the CRISPR protein.

[0458] CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacer sequences, sequences complementary to antecedent mobile elements, and target invading nucleic acids. The CRISPR clusters are transcribed and processed into CRISPR RNAs (crRNAs). In type II CRISPR systems, proper processing of pre-crRNA requires the trans-encoded small RNA (tracrRNA), the endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3 to assist in processing pre-crRNA. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer sequence. The target strand that is not complementary to the crRNA is first cleaved endonucleolytically and then trimmed 3′-5' exonucleolytically. In nature, DNA binding and cleavage generally require a protein and two RNAs. However, a single guide RNA (“sgRNA,” or simply “gRNA” for short) can be engineered to incorporate aspects of both the crRNA and the tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif within the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self.

[0459] In some embodiments, the methods described herein can utilize engineered Cas proteins. A guide RNA (gRNA) is a short synthetic RNA that consists of a scaffold sequence required for Cas binding and a user-defined ~20 nucleotide spacer sequence that defines the genomic target to be modified. Thus, the genomic target specificity of the Cas protein can be altered by a person skilled in the art depending on the specificity of the gRNA targeting sequence for the genomic target compared to the rest of the genome.

[0460] In some embodiments, the gRNA scaffold sequence is as follows:

[0461] In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is an endonuclease (e.g., a deoxyribonuclease or ribonuclease) that is capable of binding to a target polynucleotide when bound to a bound guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a nickase that is capable of binding to a target polynucleotide when bound to a bound guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a catalytically dead domain that is capable of binding to a target polynucleotide when bound to a bound guide nucleic acid. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is DNA. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is RNA.

[0462] Cas proteins useful herein include Class 1 and Class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, CARF, DinG, homologs thereof or modified versions thereof. Unmodified CRISPR enzymes can have DNA cleavage activity, such as Cas9, which has two functional endonuclease domains: RuvC and HNH. The CRISPR enzyme can direct cleavage of one or both strands at the target sequence, e.g., within the target sequence and / or within the complementary sequence of the target sequence. For example, the CRISPR enzyme can direct cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 or more base pairs from the first or last nucleotide of the target sequence.

[0463] Vectors encoding CRISPR enzymes can be used, where the CRISPR enzyme is mutated relative to the corresponding wild-type enzyme such that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing a target sequence. Cas9 can refer to a polypeptide having at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity and / or sequence homology with a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from Streptococcus pyogenes). Cas9 can refer to a polypeptide having at most or at most about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity and / or sequence homology with a wild-type exemplary Cas9 polypeptide (e.g., from Streptococcus pyogenes). Cas9 can refer to the wild-type or a modified form of the Cas9 protein, which can contain amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.

[0464] In some embodiments, the CRISPR protein-derived domain of the base editor can include those from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Refs: NC_021284.1); Prevotella intermedia (NCBI Refs: NC_017861.1); Spiroplasma taiwanense, China (NCBI Refs: NC_021846.1); Streptococcus iniae (NCBI Refs: NC_021314.1); Belliella baltica (NCBI Refs: NC_018010.1); Psychroflexus torquis I (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Refs: YP_820832.1); Listeria innocua (NCBI Refs: NP_472073.1); Campylobacter jejuni (NCBI Refs: YP_002344900.1); Neisseria meningitidis (NCBI Refs: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus aureus.

[0465] The Cas9 domain of the nucleobase editor

[0466] The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus. Based on the present disclosure, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5,726-737; the entire contents of which are incorporated herein by reference.

[0467] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain (dCas9), or a Cas9 nickase (nCas9). In some embodiments, the Cas9 domain is a domain having nuclease activity. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any of the amino acid sequences described herein. In some embodiments, the amino acid sequence comprised by the Cas9 domain is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or 99.5% identical to any of the amino acid sequences described herein. In some embodiments, the amino acid sequence comprised by the Cas9 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences described herein. In some embodiments, compared to any of the amino acid sequences described herein, the Cas9 domain comprises at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100 or at least 1200 identical consecutive amino acid residues.

[0468] In some embodiments, proteins comprising Cas9 fragments are provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, the protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". The Cas9 variant is homologous to Cas9 or a fragment thereof. For example, the Cas9 variant is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5% or at least about 99.9% identical to wild-type Cas9. In some embodiments, compared to wild-type Cas9, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., the gRNA-binding domain or the DNA cleavage domain) such that the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5% or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% the amino acid length of the corresponding wild-type Cas9. In some embodiments, the length of the fragment is at least 100 amino acids. In some embodiments, the length of the fragment is at least 100, 150, 200, 250, 300, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250 or 1300 amino acids.

[0469] In some embodiments, the Cas9 fusion proteins provided herein comprise the full-length amino acid sequence of a Cas9 protein, such as one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not comprise the full-length Cas9 sequence, but only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and other suitable sequences of Cas9 domains and fragments will be apparent to those of skill in the art.

[0470] The Cas9 protein can be associated with a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide binding domain is a Cas9 domain, such as nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactivated Cas9 (dCas9). Examples of nucleic acid programmable DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), CasX, CasY, Cpf1, Cas12b / C2C1, and Cas12c / C2C3.

[0471] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows).

[0472]

[0473]

[0474] (Single underline: HNH domain; double underline: RuvC domain)

[0475] In some embodiments, wild-type Cas9 corresponds to or comprises the following nucleotide and / or amino acid sequences:

[0476]

[0477]

[0478]

[0479] (Single underline: HNH domain; double underline: RuvC domain).

[0480] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes

[0481] (NCBI reference sequence: NC_002737.2 (nucleotide sequence as follows); and Uniprot reference sequence: Q99ZW2 (amino acid sequence as follows):

[0482]

[0483]

[0484] (Single underline: HNH domain; double underline: RuvC domain)

[0485] In some embodiments, Cas9 refers to Cas9 from the following: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Refs: NC_021284.1); Prevotella intermedia (NCBI Refs: NC_017861.1); Spiroplasma taiwanense, China (NCBI Refs: NC_021846.1); Streptococcus iniae (NCBI Refs: NC_021314.1); Belliella baltica (NCBI Refs: NC_018010.1); Psychroflexus torquisI (NCBI Refs: NC_018721.1); Streptococcus thermophilus (NCBI Refs: YP_820832.1); Listeria innocua (NCBI Refs: NP_472073.1); Campylobacter jejuni (NCBI Refs: YP_002344900.1); Neisseria meningitidis (NCBI Refs: YP_002342100.1) or Cas9 from any other organism.

[0486] It should be understood that additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including variants and homologs thereof, are within the scope of the present disclosure. Exemplary Cas9 proteins include, but are not limited to, those provided below. In some embodiments, the Cas9 protein is nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein is a Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is a Cas9 with nuclease activity.

[0487] In some embodiments, the Cas9 domain is a nuclease-inactivated domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-inactivated dCas9 domain contains the D10X mutation and the H840X mutation of the amino acid sequence described herein, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactivated dCas9 domain contains the D10A mutation and the H840A mutation of the amino acid sequence described herein, or corresponding mutations in any amino acid sequence provided herein. As an example, the nuclease-inactive Cas9 domain contains the amino acid sequence listed in the cloning vector pPlatTET-gRNA2 (accession number BAV54124).

[0488]

[0489] Based on the present disclosure and knowledge in the art, other suitable nuclease-inactive dCas9 domains will be apparent to those skilled in the art and are within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013;31(9):833 - 8338, the entire content of which is incorporated herein by reference).

[0490] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase, referred to as the “nCas9” protein (for “nickase” Cas9). The nuclease-inactivated Cas9 protein may be interchangeably referred to as the “dCas9” protein (for nuclease-“dead” Cas9) or catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with an inactive DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337:816 - 821 (2012); Qi et al. “Repurposing CRISPR as an RNA-Guided Platform for Sequence-Specific Control of Gene Expression” (2013) Cell. 28;152(5):1173 - 83, the entire content of which is incorporated herein by reference). For example, it is known that the DNA cleavage domain of Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337:816 - 821 (2012); Qi et al., Cell. 28:152(5):1173 - 83 (2013)).

[0491] In some embodiments, the amino acid sequence contained in the dCas9 domain is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% or 99.5% identical to any of the Cas9 domains described herein. In some embodiments, the amino acid sequence contained in the Cas9 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences described herein. In some embodiments, compared to any of the amino acid sequences described herein, the Cas9 domain contains at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100 or at least 1200 identical consecutive amino acid residues.

[0492] In some embodiments, dCas9 corresponds to or partially or fully contains a Cas9 amino acid sequence having one or more mutations that inactivate the Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain contains the D10A and H840A mutations or corresponding mutations in another Cas9.

[0493] In some embodiments, dCas9 contains the amino acid sequence of dCas9 (D10A and H840A):

[0494] (Single underline: HNH domain; double underline: RuvC domain).

[0495] In some embodiments, the Cas9 domain contains the D10A mutation, and the residue at position 840 remains histidine at the corresponding position in the amino acid sequence provided above or in any of the amino acid sequences provided herein.

[0496] In other embodiments, dCas9 variants having mutations other than D10A and H840A are provided, which for example result in nuclease-inactivated Cas9 (dCas9). For example, such mutations include other amino acid substitutions at D10 and H840, or other substitutions within the Cas9 nuclease domain (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical or at least about 99.9% identical. In some embodiments, there are provided about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids or more shorter or longer.

[0497] In some embodiments, the Cas9 domain is a Cas9 nickase. A Cas9 nickase can be a Cas9 protein that is only capable of cleaving one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that base pairs (is complementary) with the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase contains a D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-target, non-base editing strand of the double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that does not base pair with the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase contains an H840A mutation and has an aspartic acid residue or a corresponding mutation at position 10. In some embodiments, the Cas9 nickase contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% or 99.5% identical to any of the Cas9 nickases described herein. Based on the present disclosure and the knowledge in the art, other suitable Cas9 nickases will be apparent to those skilled in the art and are within the scope of the present disclosure.

[0498]

[0499] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., Nanoarchaeum), which constitutes the domain and kingdom of single-celled prokaryotic microorganisms. In some embodiments, the programmable nucleotide-binding protein can be a CasX or CasY protein, which has been described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the entire content of which is incorporated herein by reference. Using genome-resolved metagenomics, many CRISPR-Cas systems have been identified, including Cas9 first reported in the archaea domain. This divergent Cas9 protein was found in the little-studied Nanoarchaeum as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were found, which are among the most compact systems discovered to date. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasX or a variant of CasX. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasY or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins can be used as nucleic acid programmable DNA-binding proteins (napDNAbp) and are within the scope of the present disclosure.

[0500] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) or any fusion protein provided herein can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide-binding protein is a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide-binding protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to any CasX or CasY protein described herein. It should be understood that CasX and CasY from other bacterial species can also be used according to the present disclosure.

[0501] The amino acid sequence of exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN87|F0NN87_SULIHC RISPR-associated Casx protein OS = Sulfolobus islandicus (strain HVE10 / 4) GN = SiH_0402 PE = 4 SV = 1) is as follows:

[0502] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.

[0503] The amino acid sequence of exemplary CasX (>tr|F0NH53|F0NH53_SULIR CRISPR associated protein, Casx OS = Sulfolobus islandicus (strain REY15A) GN = SiRe_0771 PE = 4 SV = 1) is as follows:

[0504] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.

[0505] Delta Proteobacterium CasX

[0506] MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA

[0507] The exemplary CasY ((ncbi.nlm.nih.gov / protein / APG80656.1)>APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group]) amino acid sequence is as follows:

[0508]

[0509] The Cas9 nuclease has two functional endonuclease domains: RuvC and HNH. Upon target binding, Cas9 undergoes a conformational change that positions the nuclease domains to cleave opposite strands of the target DNA. The end result of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (about 3-4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired by one of two general repair pathways: (1) the highly efficient but error-prone non-homologous end joining (NHEJ) pathway; or (2) the less efficient but high-fidelity homologous directed repair (HDR) pathway.

[0510] The "efficiency" of non-homologous end joining (NHEJ) and / or homologous directed repair (HDR) can be calculated by any convenient method. For example, in some embodiments, efficiency can be expressed as the percentage of successful HDR. For example, the Surveyor nuclease assay can be used to generate cleavage products, and the ratio of products to substrate can be used to calculate the percentage. For example, the Surveyor nuclease can be used to directly cleave DNA containing a newly integrated restriction sequence as a result of successful HDR. More cleaved substrate indicates a higher percentage of HDR (higher HDR efficiency). As an illustrative example, the fraction (percentage) of HDR can be calculated using the following equation: [(cleavage products) / (substrate plus cleavage products)] (e.g., (b + c) / (a + b + c)), where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products.

[0511] In some embodiments, efficiency can be expressed as the percentage of successful NHEJ. For example, the T7 endonuclease I assay can be used to generate cleavage products, and the ratio of products to substrate can be used to calculate the NHEJ percentage. T7 endonuclease I cleaves mismatched heteroduplex DNA generated by hybridization of wild-type and mutant DNA strands (NHEJ generates small random insertions or deletions (indels) at the original break site). More cleavage indicates a higher percentage of NHEJ (higher NHEJ efficiency). As an illustrative example, the fraction (percentage) of NHEJ can be calculated using the following equation: (1 - (1 - (b + c) / (a + b + c)) 1 / 2 ) × 100, where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products (Ran et al., Cell. 2013 Sep. 12; 154(6):1380 - 9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11):2281–2308).

[0512] The NHEJ repair pathway is the most active repair mechanism and often results in small nucleotide insertions or deletions (indels) at the DSB site. The randomness of NHEJ-mediated DSB repair has important practical implications because a population of cells expressing Cas9 and gRNA or a guide polynucleotide will result in a variety of mutations. In most embodiments, NHEJ generates small indels in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations that lead to premature stop codons within the open reading frame (ORF) of the target gene. The ideal end result is a loss-of-function mutation within the target gene.

[0513] Although NHEJ-mediated DSB repair often disrupts the open reading frame of a gene, homology-directed repair (HDR) can be used to generate specific nucleotide changes, ranging from single nucleotide changes to large insertions such as the addition of a fluorophore or a tag.

[0514] To utilize HDR for gene editing, a DNA repair template containing the desired sequence can be delivered into the cell type of interest using a gRNA and Cas9 or Cas9 nickase. The repair template can contain the desired edit as well as additional homologous sequences (referred to as left and right homology arms) immediately upstream and downstream of the target. The length of each homology arm depends on the size of the introduced change, with larger insertions requiring longer homology arms. The repair template can be a single-stranded oligonucleotide, a double-stranded oligonucleotide, or a double-stranded DNA plasmid. Even in cells expressing Cas9, gRNA, and an exogenous repair template, the efficiency of HDR is typically low (<10% modified alleles). The efficiency of HDR can be increased by synchronizing the cells because HDR occurs during the S and G2 phases of the cell cycle. Chemical or genetic inhibitors of genes involved in NHEJ can also increase the HDR frequency.

[0515] In some embodiments, Cas9 is a modified Cas9. A given gRNA target sequence can have additional sites throughout the genome where there is partial homology. These sites are called off-target sites and need to be considered when designing the gRNA. In addition to optimizing gRNA design, the specificity of CRISPR can also be improved by modifications to Cas9. Cas9 generates double-stranded breaks (DSBs) through the combined activity of two nuclease domains, RuvC and HNH. Cas9 nickase is a D10A mutant of SpCas9 that retains one nuclease domain and generates a DNA nick instead of a DSB. The nickase system can also be combined with HDR-mediated gene editing for specific gene editing.

[0516] In some embodiments, Cas9 is a variant Cas9 protein. The variant Cas9 polypeptide has an amino acid sequence that differs by one amino acid from the amino acid sequence of the wild-type Cas9 protein (e.g., has a deletion, insertion, substitution, fusion). In some cases, the variant Cas9 polypeptide has an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nuclease activity of the Cas9 polypeptide. For example, in some cases, the variant Cas9 polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas9 protein. In some embodiments, the variant Cas9 protein has no substantial nuclease activity. When the subject Cas9 protein is a variant Cas9 protein with no substantial nuclease activity, it can be referred to as "dCas9".

[0517] In some embodiments, the variant Cas9 protein has reduced nuclease activity. For example, the variant Cas9 protein exhibits less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the endonuclease activity of the wild-type Cas9 protein, such as the wild-type Cas9 protein.

[0518] In some embodiments, the variant Cas9 protein can cleave the complementary strand of the guide target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10) and can thus cleave the complementary strand of the double-stranded guide target sequence but not the complementary strand of the non-double-stranded guide target sequence (thus resulting in a single-strand break (SSB) rather than a double-strand break (DSB) when the variant Cas9 protein cleaves the double-stranded target nucleic acid) (see, e.g., Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21).

[0519] In some embodiments, the variant Cas9 protein can cleave the non-complementary strand of the double-stranded guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has an H840A (histidine to alanine at amino acid position 840) mutation, and thus can cleave the non-complementary strand of the guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence (resulting in the use of SSB instead of DSB when the variant Cas9 protein cleaves the double-stranded guide target sequence). Such Cas9 proteins have a reduced ability to cleave the guide target sequence (e.g., single-stranded guide target sequence), but retain the ability to bind to the guide target sequence (e.g., single-stranded guide target sequence).

[0520] In some embodiments, the variant Cas9 protein has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. As a non-limiting example, in some embodiments, the variant Cas9 protein contains both D10A and H840A mutations, such that the polypeptide has a reduced ability to cleave both the complementary and non-complementary strands of double-stranded target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0521] As another non-limiting example, in some embodiments, the variant Cas9 protein contains W476A and W1126A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0522] As another non-limiting example, in some embodiments, the variant Cas9 protein contains P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the polypeptide has a reduced ability to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0523] As another non-limiting example, in some embodiments, the variant Cas9 protein contains H840A, W476A, and W1126A mutations such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, the variant Cas9 protein contains H840A, D10A, W476A, and W1126A mutations such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 restores the catalytic His residue at position 840 in the Cas9 HNH domain (A840H).

[0524] As another non-limiting example, in some embodiments, the variant Cas9 protein contains the H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some embodiments, the variant Cas9 protein contains the D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind target DNA (e.g., single-stranded target DNA). In some embodiments, when the variant Cas9 protein contains the W476A and W1126A mutations or when the variant Cas9 protein contains the P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, the variant Cas9 protein does not bind effectively to the PAM sequence. Thus, in some such embodiments, when such a variant Cas9 protein is used in a binding method, the method does not require a PAM sequence. In other words, in some embodiments, when such a variant Cas9 protein is used in a binding method, the method can include a guide RNA, but the method can proceed in the absence of a PAM sequence (and the specificity of the binding is thus provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effect (i.e., inactivate one or the other nuclease moiety). As non-limiting examples, the residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., substituted). Additionally, mutations other than alanine substitutions are suitable.

[0525] In some embodiments, a variant Cas9 protein with reduced catalytic activity (e.g., when the Cas9 protein has the D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutations, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A), the variant Cas9 protein can still bind to target DNA in a site-specific manner (because it is still guided by the guide RNA to the target DNA sequence), as long as it retains the ability to interact with the guide RNA.

[0526] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.

[0527] In some embodiments, a modified SpCas9 is used that includes the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and is specific for the altered PAM 5'-NGC.

[0528] Alternatives to Streptococcus pyogenes Cas9 can include RNA-guided endonucleases from the Cpf1 family, which show cleavage activity in mammalian cells. CRISPR from Prevotella and Francisella 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the type II CRISPR / Cas system. This acquired immune mechanism exists in Prevotella and Francisella. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cut viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some of the limitations of the CRISPR / Cas9 system. Different from Cas9 nuclease, the result of Cpf1-mediated DNA cleavage is a double-strand break with short 3' overhangs. The staggered cleavage pattern of Cpf1 can open up the possibility of directional gene transfer, similar to traditional restriction enzyme cloning, which can improve the efficiency of gene editing. Like the above-mentioned Cas9 variants and orthologs, Cpf1 can also expand the number of CRISPR-targetable sites to AT-rich regions or AT-rich genomes that lack the NGG PAM sites favored by SpCas9. The Cpf1 locus contains a mixed alpha / beta domain, a RuvC-I followed by a helical region, a RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. In addition, Cpf1 does not have an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain architecture indicates that Cpf1 is functionally unique and is classified as a type 2 class V CRISPR system. The Cas1, Cas2, and Cas4 proteins encoded by the Cpf1 locus are more similar to type I and type III than those from type II systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA), and thus, only CRISPR (crRNA) is needed. This is beneficial for genome editing because Cpf1 is not only smaller than Cas9, but its sgRNA molecule is also smaller (about half the nucleotides of Cas9). Compared with the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex cuts the target DNA or RNA by recognizing the protospacer adjacent motif 5'-YTN-3'. After identifying the PAM, Cpf1 introduces a sticky-end-like DNA double-strand break with 4 or 5 nucleotide overhangs.

[0529] In some embodiments, Cas9 is a Cas9 variant that is specific for an altered PAM sequence. In some embodiments, additional Cas9 variants and PAM sequences are described in Miller, S.M. et al., Continuous evolution of SpCas9 variants compatible with non-G PAMs, Nat. Biotechnol. (2020), the entire content of which is incorporated herein by reference. In some embodiments, the Cas9 variant has no specific PAM requirement. In some embodiments, the Cas9 variant, such as the SpCas9 variant, is specific for an NRNH PAM, where R is A or G and H is A, C, or T. In some embodiments, the SpCas9 variant is specific for the PAM sequences AAA, TAA, CAA, GAA, TAT, GAT, or CAC. In some embodiments, the SpCas9 variant contains amino acid substitutions at position 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1218, 1219, 1221, 1249, 134, 137, 127, 137 1320, 1321, 1323, 1332, 1333, 1335, 1337, or 1339 of SEQ ID NO:1 or their corresponding positions. In some embodiments, the SpCas9 variant contains amino acid substitutions at position 1114, 1135, 1218, 1219, 1221, 1249, 1320, 1321, 1323, 1332, 1333, 1335, or 1337 of SEQ ID NO:1 or their corresponding positions. In some embodiments, the SpCas9 variant contains amino acid substitutions at position 1114, 1134, 1135, 1137, 1139, 1151, 1180, 1188, 1211, 1219, 1221, 1256, 1264, 1290, 1318, 1317, 1320, 1323, 1333, or their corresponding positions of SEQ ID NO:1. In some embodiments, the SpCas9 variant contains amino acid substitutions at position 1114, 1131, 1135, 1150, 1156, 1180, 1191, 1218, 1219, 1221, 1227, 1249, 1253, 1286, 1293, 1320, 1321, 1332, 1335, 1339, or their corresponding positions of SEQ ID NO:1.In some embodiments, the SpCas9 variant contains amino acid substitutions at position 1114, 1127, 1135, 1180, 1207, 1219, 1234, 1286, 1301, 1332, 1335, 1337, 1338, 1349 of SEQ ID NO:1 or their corresponding positions. Exemplary amino acid substitutions and PAM specificities of the SpCas9 variant are shown in Tables 3A - 3D.

[0530] Table 3A

[0531]

[0532]

[0533] Table 3B

[0534]

[0535] Table 3C

[0536]

[0537] Table 3D

[0538]

[0539] 5 In some embodiments, Cas9 is Neisseria meningitidis Cas9 (NmeCas9) or a variant thereof. In some embodiments, NmeCas9 is specific for the NNNNGAYW PAM, where

[0540] Y is C or T and W is A or T. In some embodiments, NmeCas9 is specific for the NNNNGYTT PAM, where Y is C or T. In some embodiments, NmeCas9 is specific for

[0541] The NNNNGTCT PAM is specific. In some embodiments, the NmeCas9 is Nme110Cas9. In some embodiments, NmeCas9 is specific for NNNNGATT PAM, NNNNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM, NNNNCCTG PAM, NNNNCCGT PAM, NNNNCCGGPAM, NNNNCCCA PAM, NNNNCCCT PAM, NNNNCCCC PAM, NNNNCCAT PAM, NNNNCCAG PAM, NNNNCCAT PAM or NNNGATT PAM. In some embodiments, Nme1Cas9 is specific for NNNNGATT PAM, NNNNCCTA PAM, NNNNCCTC PAM, NNNNCCTT PAM or NNNNCCTGPAM. In some embodiments, NmeCas9 is specific for CAA PAM, CAAA PAM or CCA PAM. In some embodiments, the NmeCas9 is Nme2Cas9. In some embodiments, NmeCas9 is specific for NNNNCC(N4CC)PAM, where N is any one of A, G, C or T. In some embodiments, NmeCas9 is specific for NNNNCCGTPAM, NNNNCCGGPAM, NNNNCCCA PAM, NNNNCCCT PAM, NNNNCCCC PAM, NNNNCCAT PAM, NNNNCCAGPAM, NNNNCCAT PAM or NNNGATT PAM. In some embodiments, the NmeCas9 is Nme3Cas9. In some embodiments, NmeCas9 is specific for NNNNCAAA PAM, NNNNCC PAM or NNNNCNNN PAM. Additional NmeCas9 features and PAM sequences, as described by Edraki et al. Mol. Cell. (2019) 73(4):714-726, are incorporated herein by reference in their entirety.

[0542] The following provides an exemplary amino acid sequence of Nme1Cas9:

[0543] Type II CRISPR RNA-guided endonuclease Cas9 [Neisseria meningitidis] WP_002235162.1

[0544]

[0545] The following provides an exemplary amino acid sequence of Nme2Cas9:

[0546] Type II CRISPR RNA-guided endonuclease Cas9 [Neisseria meningitidis] WP_002230835.1

[0547]

[0548] Cas12 domain of base editors

[0549] Generally, microbial CRISPR-Cas systems are divided into Class 1 and Class 2 systems. Class 1 systems have multi-subunit effector complexes, while Class 2 systems have single protein effectors. For example, Cas9 and Cpf1 are Class 2 effectors, although they are of different types (Type II and Type V, respectively). In addition to Cpf1, Class 2 Type V CRISPR-Cas systems also include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i). See, for example, Shmakov et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR Cas Systems,” Mol. Cell, 2015 Nov. 5; 60(3):385-397; Makarova et al., “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR Journal, 2018, 1(5):325-336; and Yan et al., “Functionally diverse type V CRISPR-Cas systems,” Science, 2019 Jan. 4; 363:88-91, which are hereby incorporated by reference in their entirety. Type V Cas proteins contain a RuvC (or RuvC-like) endonuclease domain. Although the production of mature CRISPR RNA (crRNA) generally does not depend on tracrRNA, for example, Cas12b / C2c1 requires tracrRNA to produce crRNA. Cas12b / C2c1 relies on crRNA and tracrRNA for DNA cleavage.

[0550] The nucleic acid programmable DNA-binding proteins contemplated in the present disclosure include Cas proteins classified as Class 2, Type V (Cas12 proteins). Non-limiting examples of Class 2, Type V Cas proteins include Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i, homologs thereof, or modified forms thereof. As used herein, a Cas12 protein may also be referred to as a Cas12 nuclease, a Cas12 domain, or a Cas12 protein domain. In some embodiments, the Cas12 protein of the present invention comprises an amino acid sequence interrupted by an internal fusion protein domain such as a deaminase domain.

[0551] In some embodiments, the Cas12 domain is a Cas12 domain or a Cas12 nickase having no nuclease activity. In some embodiments, the Cas12 domain is a domain having nuclease activity. For example, the Cas12 domain may be a Cas12 domain that makes a nick in one strand of a double-stranded nucleic acid (e.g., a double-stranded DNA molecule). In some embodiments, the Cas12 domain comprises any of the amino acid sequences described herein. In some embodiments, the amino acid sequence comprised by the Cas12 domain is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or 99.5% identical to any of the amino acid sequences described herein. In some embodiments, compared to any of the amino acid sequences described herein, the amino acid sequence comprised by the Cas12 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations. In some embodiments, compared to any of the amino acid sequences described herein, the Cas12 domain comprises at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100 or at least 1200 identical consecutive amino acid residues.

[0552] In some embodiments, proteins comprising Cas12 fragments are provided. For example, in some embodiments, the protein comprises one of two Cas12 domains: (1) the gRNA-binding domain of Cas12; or (2) the DNA cleavage domain of Cas12. In some embodiments, a protein comprising Cas12 or a fragment thereof is referred to as a "Cas12 variant". The Cas12 variant is homologous to Cas12 or a fragment thereof. For example, the Cas12 variant is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, ...

Claims

1. A method of generating modified immune cells, the method comprising expressing or introducing a nucleobase editor polypeptide in immune cells and contacting the cells with two or more guide RNAs targeting the nucleobase editor polypeptide to effect alteration of a nucleic acid molecule encoding at least one polypeptide selected from the group consisting of T cell receptor alpha constant (TRAC), beta-2 microglobulin (B2M), programmed cell death 1 (PD1), cluster of differentiation 7 (CD7), cluster of differentiation 5 (CD5), cluster of differentiation 33 (CD33), cluster of differentiation 123 (CD123), Cbl proto-oncogene B (CBLB), and class II major histocompatibility complex transactivator (CIITA) polypeptide, wherein the nucleobase editor polypeptide comprises a nucleic acid programmable DNA binding protein (napDNAbp) and at least one base editor domain, the base editor domain comprising an adenosine deaminase variant domain comprising an alteration at amino acid position 82 and / or 166 of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD.

2. A modified immune cell produced by the method according to claim 1.

3. Use of an effective amount of the modified immune cell according to claim 2 in the manufacture of a medicament for modulating the immune response of a subject.

4. Use of an effective amount of the modified immune cell according to claim 2 in the manufacture of a medicament for treating a tumor in a subject.

5. Use of an effective amount of the modified immune cell according to claim 2 in the manufacture of a medicament for treating a subject suffering from or having a tendency to develop graft-versus-host disease (GVHD).

6. Use of an effective amount of the modified immune cell according to claim 2 in the manufacture of a medicament for treating a subject suffering from or having a tendency to develop host-versus-graft disease (HVGD).

7. A pharmaceutical composition comprising an effective amount of the modified immune cell according to claim 2 in a pharmaceutically acceptable excipient.

8. A base editor comprising a polynucleotide programmable DNA binding domain and at least one base editor domain, wherein the base editor domain comprises an adenosine deaminase variant having an alteration at amino acid position 82 or 166 of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD and two or more guide RNAs targeting the nucleobase editor polypeptide to effect an alteration of a nucleic acid molecule encoding at least one polypeptide selected from the group consisting of T cell receptor alpha constant (TRAC), beta-2 microglobulin (B2M), programmed cell death 1 (PD1), cluster of differentiation 7 (CD7), cluster of differentiation 5 (CD5), cluster of differentiation 33 (CD33), cluster of differentiation 123 (CD123), Cbl proto-oncogene B (CBLB), and class II major histocompatibility complex transactivator (CIITA) polypeptides.

9. A base editor system comprising the base editor of claim 8, wherein the adenosine deaminase variant comprises a V82S alteration and / or a T166R alteration.

10. A base editor system comprising two or more guide RNAs and a fusion protein comprising a polynucleotide programmable DNA binding domain comprising the following sequence: Wherein the bold sequence represents a sequence derived from Cas9, the italic sequence represents a linker sequence, the underlined sequence represents a dual nuclear localization sequence, and at least one base editor domain, the base editor domain comprising an adenosine deaminase variant having an alteration at amino acid position 82 or 166 of MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD and two or more guide RNAs targeting the nuclear base editor polypeptide to effect an alteration of a nucleic acid molecule encoding at least one polypeptide selected from the group consisting of T cell receptor alpha constant (TRAC), beta-2 microglobulin (B2M), programmed cell death 1 (PD1), cluster of differentiation 7 (CD7), cluster of differentiation 5 (CD5), cluster of differentiation 33 (CD33), cluster of differentiation 123 (CD123), Cbl proto-oncogene B (CBLB), and class II major histocompatibility complex transactivator (CIITA) polypeptide.

Citation Information

Patent Citations

  • Sedan protective covering

    CN2558543Y

  • Towel rings

    CN3315821D

  • cell phone

    CN3329834D

  • Split inteins, conjugates and uses thereof

    US20150344549A1

  • AAV delivery of nucleobase editors

    US20180127780A1