Compositions and methods for non-toxic conditioning

By expressing nucleobase editing peptides in hematopoietic stem cells and combining guide RNA to target cell surface proteins, the side effects of existing conditioning methods are solved, non-toxic or low-toxic conditioning is achieved, and the implantation efficiency of hematopoietic stem cell grafts is improved.

CN120249248APending Publication Date: 2025-07-04BEAM THERAPEUTICS INC
View PDF 25 Cites 0 Cited by

Patent Information

Application Number
CN202510339660.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-08-29
Filing Date
2020-08-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing hematopoietic stem cell transplantation conditioning methods such as the use of Bupleuri An has a significant risk of side effects, which is difficult to effectively promote the implantation of hematopoietic stem cell grafts in the host, and existing methods may lead to an increase in chimeric level.

Method used

Gene editing using nucleobase editing polypeptides in hematopoietic stem cells or their progenitor cells, binding to guide RNA to target specific cell surface proteins, introducing mutations to improve the characteristics of cell surface proteins, thereby facilitating transplantation, using nucleic acid programmable DNA binding proteins (napDNAbp) and deaminases.

Benefits of technology

Non-toxic or low-toxic conditioning is achieved, improving the implantation efficiency of hematopoietic stem cell grafts in the host, reducing the chimeric level and reducing the risk of side effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120249248A_ABST
    Figure CN120249248A_ABST
Patent Text Reader

Abstract

The present invention features compositions and methods for conditioning a patient (e.g., to promote transplantation and / or implantation). The invention provides a base editing strategy for targeting a cell surface protein. The strategy can be used for conditioning.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of a Chinese patent application with the application number 202080076279.5, the filing date of August 28, 2020, and the invention title of "Compositions and Methods for Aseptic Conditioning".

[0002] Related Applications

[0003] This application claims the benefit of U.S. Provisional Application No. 62 / 893,677, filed on January 28, 2019, under 35 U.S.C. 119(e), the entire content of which is incorporated herein by reference. Technical Field

[0004] This application relates to compositions and methods for aseptic conditioning and a base editing strategy that targets cell surface proteins and can be used for such aseptic conditioning. Background Art

[0005] Gene therapy using autologous hematopoietic stem cells (HSCs) is desirable because it is generally independent of the requirement to suppress the immune system and reduces the risk of host-versus-graft or graft-versus-host disease. However, ensuring engraftment of hematopoietic stem cell grafts in the host remains challenging. Conditioning (HSC ablation) prior to hematopoietic stem cell transplantation (HSCT) is used to facilitate transplantation. In fact, conditioning efficacy is associated with improved engraftment. One drawback of myeloablative (toxic) conditioning is that it leads to increased levels of chimerism.

[0006] Current conditioning methods for autologous gene therapy for treating, for example, thalassemia, sickle cell disease (SCD), or adenosine deaminase deficiency generally involve using 2 - 5 g / kg of intravenous busulfan for 2 - 4 days. Busulfan is a DNA alkylating agent originally designed for treating blood diseases such as acute myeloid leukemia (AML) and others. However, busulfan carries a risk of significant side effects, including infertility, primary or secondary malignancies, and other acute and chronic toxicities.

[0007] Conditioning prior to HSCT is an unmet medical need. Therefore, there is an urgent need for compositions and methods for conditioning to facilitate engraftment of hematopoietic stem cell grafts. Summary of the Invention

[0008] As described below, the present invention features compositions and methods for aseptic conditioning.

[0009] On the one hand, the present invention provides a method for manufacturing hematopoietic stem cells or their progenitor cells for treating hemoglobinopathies, blood cancers, or myeloproliferative diseases. In some embodiments, the method comprises (a) expressing a nucleobase editing polypeptide in hematopoietic stem cells or their progenitor cells, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA binding protein (napDNAbp) and a deaminase; (b) contacting the hematopoietic stem cells or their progenitor cells with a guide RNA that targets a nucleic acid molecule encoding a cell surface protein selected from the group consisting of CD117, CXCR4, CD135, CD90, CD45, and CD34, and introducing a mutation in the cell surface protein.

[0010] In some embodiments, the method comprises (a) expressing a nucleobase editing polypeptide in hematopoietic stem cells or their progenitor cells comprising CD117 protein, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA binding protein (napDNAbp) and a deaminase; (b) contacting the hematopoietic stem cells or progenitor cells with a guide RNA capable of targeting a polynucleotide encoding a cell surface protein, thereby manufacturing hematopoietic stem cells or their progenitor cells for treating hemoglobinopathies, blood cancers, or myeloproliferative diseases.

[0011] In some embodiments, the hemoglobinopathies are selected from the group consisting of sickle cell anemia, thalassemia, Fanconi anemia, aplastic anemia, and Wiskott - Aldrich syndrome. In some embodiments, the blood cancers are selected from the group consisting of acute myeloid leukemia, acute lymphoblastic leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, multiple myeloma, diffuse large B - cell lymphoma, and non - Hodgkin's lymphoma. In some embodiments, the myeloproliferative disease is myelodysplastic syndrome. On the one hand, the present invention provides a method for manufacturing hematopoietic stem cells or their progenitor cells for treating immunodeficiency. In some embodiments, the immunodeficiency is severe combined immunodeficiency (SCID).

[0012] On the other hand, the present invention provides a method for identifying mutations that alter the binding of an antibody to a cell surface protein. In some embodiments, the method comprises (a) expressing a nucleobase editing polypeptide in a cell comprising a cell surface protein, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA binding protein (napDNAbp) and a deaminase; (b) contacting the cell with a guide RNA capable of targeting a nucleic acid molecule encoding the cell surface protein and introducing a mutation in the cell surface protein; and (c) contacting the cell with an antibody that specifically binds to the wild - type cell surface protein but exhibits reduced binding to the cell surface protein comprising the mutation, thereby identifying a mutation that alters the binding of the antibody to the cell surface protein.

[0013] In some embodiments, the method further comprises determining the biological activity of the cells. In some embodiments, the cell surface protein is selected from the group consisting of CD117, CXCR4, CD135, CD90, CD45, and CD34. In some embodiments, the cell surface protein is CD117. In some embodiments, the method further comprises contacting the cells with one or more additional guide RNAs that target cell surface proteins selected from the group consisting of CXCR4, CD135, CD90, CD45, and CD34.

[0014] In some embodiments, the cells are hematopoietic stem cells or their progenitor cells. In some embodiments, the mutation is a missense mutation. In some embodiments, the missense mutation does not alter the biological activity of the cell surface protein. In some embodiments, one or more amino acid substitutions are introduced by a nucleobase editing polypeptide and a guide RNA.

[0015] In some embodiments, the method comprises: (a) expressing a nucleobase editing polypeptide in a hematopoietic stem cell or its progenitor cell, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA binding protein (napDNAbp) and a deaminase; (b) contacting the cell with a guide RNA capable of targeting a nucleic acid encoding a cell surface protein selected from the group consisting of CD117, CXCR4, CD135, CD90, CD45, and CD34 and introducing a mutation in the cell surface protein; and (c) contacting the cell with an antibody that specifically binds to the wild-type cell surface protein but exhibits reduced binding to the cell surface protein comprising the mutation, thereby identifying a mutation that alters antibody binding to the cell surface protein.

[0016] In another aspect, the present invention provides a method of base editing a gene encoding a cell surface protein expressed by a hematopoietic stem cell or its progenitor cell. In some embodiments, the method comprises: (a) expressing a nucleobase editing polypeptide in a hematopoietic stem cell or its progenitor cell comprising a CD117 protein, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA binding protein (napDNAbp) and a deaminase; (b) contacting the cell with a guide RNA capable of targeting a nucleic acid encoding CD117, thereby base editing the gene encoding the cell surface protein.

[0017] In one aspect, the present invention provides a method of conditioning a subject either concurrently with or subsequent to hematopoietic stem cell transplantation (HSCT). In some embodiments, the method comprises (a) expressing in isolated hematopoietic stem cells of the subject or donor a nucleobase editing polypeptide comprising a nucleic acid programmable DNA binding protein (napDNAbp) and a deaminase; (b) contacting the hematopoietic stem cells with a guide RNA capable of targeting a nucleic acid molecule encoding a cell surface protein selected from the group consisting of CD117, CXCR4, CD135, CD90, CD45, and CD34, thereby introducing a mutation in the cell surface protein and generating edited hematopoietic stem cells; (c) administering the edited hematopoietic stem cells to the subject; (d) administering to the subject an antibody, antibody-drug conjugate, or chimeric antigen receptor-expressing T cell (CAR-T) that selectively binds the wild-type cell surface protein, wherein the administration of step (d) is concurrent with or subsequent to step (c).

[0018] In some embodiments, the method comprises (a) expressing in hematopoietic stem cells of the subject a nucleobase editing polypeptide comprising a nucleic acid programmable DNA binding protein (napDNAbp) and a deaminase; (b) contacting the hematopoietic stem cells with a guide RNA capable of targeting a nucleic acid molecule encoding the CD117 protein, thereby introducing a mutation in the CD117 protein and generating edited hematopoietic stem cells; (c) administering the edited hematopoietic stem cells to the subject; (d) administering to the subject an antibody, antibody-drug conjugate, or chimeric antigen receptor-expressing T cell (CAR-T) that selectively binds the wild-type CD117 protein, wherein the administration of step (d) is concurrent with or subsequent to step (c).

[0019] In some embodiments, the deaminase domain is adenosine deaminase or cytidine deaminase. In some embodiments, the adenosine deaminase domain is a TadA deaminase domain. In some embodiments, the adenosine deaminase contains a TadA*8 variant. In some embodiments, the adenosine deaminase is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In some embodiments, the deaminase is a monomer or a heterodimer. In some embodiments, the base editing polypeptide is an ABE8 base editor. In some embodiments, the ABE8 base editor is ABE8.1-m, ABE8.2-m, ABE8.3-m, ABE8.4-m, ABE8.5-m, ABE8.6-m, ABE8.7-m, ABE8.8-m, ABE8.9-m, ABE8.10-m, ABE8.11-m, ABE8.12-m, ABE8.13-m, ABE8.14-m, ABE8.15-m, ABE8.16-m, ABE8.17-m, ABE8.18-m, ABE8.19-m, ABE8.20-m, ABE8.21-m, ABE8.22-m, ABE8.23-m, ABE8.24-m, ABE8.1-d, ABE8.2-d, ABE8.3-d, ABE8.4-d, ABE8.5-d, ABE8.6-d, ABE8.7-d, ABE8.8-d, ABE8.9-d, ABE8.10-d, ABE8.11-d, ABE8.12-d, ABE8.13-d, ABE8.14-d, ABE8.15-d, ABE8.16-d, ABE8.17-d, ABE8.18-d, ABE8.19-d, ABE8.20-d, ABE8.21-d, ABE8.22-d, ABE8.23-d, or ABE8.24-d.

[0020] In some embodiments, the nucleobase editor polypeptide is an internal base editor (IBE) that includes a deaminase domain inserted at an internal position within the napDNAbp. In some embodiments, the nucleobase editor polypeptide further includes one or more uracil glycosylase inhibitors (UGIs). In some embodiments, the nucleobase editor polypeptide further includes one or more nuclear localization sequences (NLSs).

[0021] In some embodiments, the mutation is a missense mutation. In some embodiments, the missense mutation does not alter the biological activity of the cell surface protein. In some embodiments, the CD117 cell surface protein comprising the mutation is capable of binding stem cell factor (SCF). In some embodiments, the CD117 cell surface protein comprising the mutation is capable of signaling stem cell factor (SCF). In some embodiments, the mutation is at least one amino acid substitution caused by modification of one or more single target nucleobases. In some embodiments, the single target nucleobase is cytosine (C), and wherein the modification comprises converting the C to thymine (T). In some embodiments, the single target nucleobase is adenosine (A), and wherein the modification comprises converting the A to guanine (G). In some embodiments, the at least one amino acid substitution is a naturally occurring mutation. In some embodiments, the at least one amino acid substitution is in domain 1, 2, 3, 4, or 5 of CD117. In some embodiments, the at least one amino acid substitution in CD117 is selected from the group consisting of: T13A, S35P, I39V, H40R, K43G, K43R, S44P, D45G, I47V, D52G, E53G, I54V, R55G, L56P, L57P, T59A, F63P, V64A, K65E, K65R, W66R, T67A, D72G, E73G, T74A, N75D, N75G, E76G, N77G, N77S, N77Y, K78E, K78R, Q79R, N80G, N80S, E81G, E81D, W82R, I83T, I83V, T84A, E85G, K86E, E88G, T90A, N99G, K100G, H101R, K116R, V120A, S123P, L124P, Y125H, K127G, K127R, E128G, D129G, D129E, N130G, N130D, D131G, D131N, T132A, T144A, N145D, N145G, N145Y, N145S, Y146C, K149E, K149G, Q152R, K154E, K154G, R161G, F162P, F162L, I163T, I163V, D165G, M171T, I172T, I172V, K173G, S174G, K176G, Q190R, E191G, K193G, V195A, L196P, S197P, E198G, K199G, F200P, I201T, I201V, L202P, V213A, V214A, S215P, V216A, K218G, K218R, S220G, Y221C, E225G, E227G,E228G, T230A, S240G, Y243C, K247G, R248G, Q256R, E257G, E257D, K258E, K258G, Y259C, Y259H, N260G, D266G, N268D, N268G, Y269C, T276A, I277T, I277V, S279P, R281G, V282A, S285P, N293G, N294G, T295A, F296P, S298P, N300G, N300S, T302A, T303A, T304A, M318V, T322A, V323A, F324L, N326G, D327G, D332G, I334V, K342E, K342G, K342R, Q347R, Y350H, M351T, R353G, T354A, T354I, K358G, E360G, D361G, K364G, E366G, N367G, H378R, T380A, R381G, K383G, T385A, T389A, D398G, V399A, N400G, V407A, Y408H, E414G, I415V, T417A, Y418C, Y418H, D419G, R420G, L421P, V422A, N423S, M425V, E435G, I438M, D439G, V454A, L455P, V457A, V459A, Q460R, T461A, N463G, S464P, S465P, F469P, K471E, K471G, L472P, V473A, Q475R, S476G, I478M, I478V, D479G, S481G, F483P, K484G, N486G, N486S, T488A, Y494C, N495D, D496G, K499E, Y503C, Y503H, F504P, and N505G. In some embodiments, the amino acid substitution in the at least one CD117 is selected from the group consisting of: T13A, I39V, H40R, D45G, D52G, E53G, I54V, R55G, T59A, K65E, K65R, T67A, E76G, N77G, N77S, N77Y, K78E, Q79R, N80G, E81G, E81D, N99G, K100G, H101R, L124P, Y125H, D129E, D129G, N130D, N130G, D131G, D131N, T132A, T144A, N145D, N145G, N145Y, N145S, Y146C, F162P, F162L, I163T, I163V, M171T, I172T, V195A, L196P,K199G, I201V, S220G, Y221C, Q256R, E257D, E257G, K258E, Y259C, T303A, T304A, M318V, T322A, V323A, F324L, K342G, K342R, Y350H, M351T, R353G, T354A, T354I, R381G, K383G, T385A, Y418C, D419G, R420G, I438M, D439G, Q460R, T461A, K484G, N486G, and T488A.,

[0022] In some embodiments, at least one amino acid substitution in CD117 is a naturally occurring mutation selected from the group consisting of: T13A, N77S, D129E, N130D, D131N, T144A, Y221C, E257D, T322A, T354I, D419G.,

[0023] In some embodiments, the cell is a hematopoietic stem cell, a common myeloid progenitor cell, a proerythroblast, or an erythrocyte. In some embodiments, the cell is a CD34+ cell. In some embodiments, the cell is from a subject having a hemoglobinopathy, a blood cancer, or a myeloproliferative disorder. In some embodiments, the cell is from a subject having sickle cell disease (SCD). In some embodiments, the cell is from a subject having hereditary persistence of fetal hemoglobin (HPFH). In some embodiments, the cell is from a subject having severe combined immunodeficiency (SCID). In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell.,

[0024] In some embodiments, the adenosine deaminase domain is fused to napDNAbp. In some embodiments, the deaminase domain is inserted into an internal position of napDNAbp. In some embodiments, the napDNAbp is an inactivated nuclease or nickase variant. In some embodiments, the napDNAbp comprises a Cas9, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i or Cas12j / CasΦ polynucleotide or a portion thereof. In some embodiments, the napDNAbp comprises a Cas9 polynucleotide or a portion thereof. In some embodiments, the napDNAbp comprises inactivated Cas9 (dCas9) or Cas9 nickase (nCas9). In some embodiments, the napDNAbp is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9) or a variant thereof. In some embodiments, the napDNAbp comprises a modified protospacer adjacent motif (PAM) specificity. In some embodiments, the modified PAM is specific for the nucleic acid sequence 5'-NGC-3'.

[0025] In some embodiments, the deaminase domain is capable of deaminating cytidine or adenine within DNA. In some embodiments, the deaminase domain is a cytidine deaminase domain. In some embodiments, the cytidine deaminase is an APOBEC deaminase domain or a derivative thereof. In some embodiments, the deaminase domain is an adenosine deaminase domain. In some embodiments, the adenosine deaminase domain is a TadA deaminase domain. In some embodiments, the adenosine deaminase is a TadA*8 variant. In some embodiments, the adenosine deaminase is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In some embodiments, the deaminase is a monomer or a heterodimer. In some embodiments, the guide polynucleotide comprises a nucleic acid sequence that comprises at least 10 consecutive nucleotides that are complementary to a CD117 nucleic acid sequence. In some embodiments, the one or more guide polynucleotides comprise a nucleic acid sequence selected from Table 23. In some embodiments, the one or more guide polynucleotides comprise a nucleic acid sequence that hybridizes to a complementary sequence of a CD117 target sequence, the complementary sequence of the CD117 target sequence being selected from the group consisting of: CAAGCTATCTTCTTAGGGA; GCAAGCTATCTTCTTAGGGAA; ACTTACGACAGGCTCGTGAA; TGACCAATTATTCCCTCAAG; AGTGACCAATTATTCCCTCA; ACTACAGTATTTGTAAACGA; AAGACAACGACACGCTGGTC; and GGCTGTTATGCACTGATCCG. In some embodiments, the guide RNA targets the following scaffold nucleic acid sequence:

[0026] GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC.

[0027] In some embodiments, the antibody is an anti-CD117 antibody. In some embodiments, the antibody binds to domains 1, 2, 3, 4, or 5 of CD117. In some embodiments, the antibody is a monoclonal antibody. In some embodiments, the hemoglobinopathy is sickle cell disease (SCD). In some embodiments, the hemoglobinopathy is hereditary persistence of fetal hemoglobin (HPFH). In some embodiments, the sickle cell disease is associated with a mutation in the β-globin (HBB) polynucleotide. In some embodiments, the HPFH is associated with a mutation in the hemoglobin subunit gamma 1 (HBG1) and / or hemoglobin subunit gamma 2 (HBG2) polynucleotide.

[0028] In another aspect, the present invention provides a base editing system comprising a fusion protein or a polynucleotide encoding the fusion protein, wherein the fusion protein comprises a nucleic acid programmable DNA binding protein (napDNAbp) and a deaminase domain, and a guide polynucleotide comprising a nucleic acid sequence selected from Table 23.

[0029] Furthermore, in another aspect, the present invention provides a base editing system comprising a fusion protein or a polynucleotide encoding the fusion protein, wherein the fusion protein comprises a nucleic acid programmable DNA binding protein (napDNAbp) and a deaminase domain, and a guide polynucleotide that hybridizes to a complementary sequence of a target nucleic acid sequence, the complementary sequence of the target nucleic acid sequence being selected from the group consisting of: CAAGCTATCTTCTTAGGGA; GCAAGCTATCTTCTTAGGGAA; ACTTACGACAGGCTCGTGAA; TGACCAATTATTCCCTCAAG; AGTGACCAATTATTCCCTCA; ACTACAGTATTTGTAAACGA; AAGACAACGACACGCTGGTC; and GGCTGTTATGCACTGATCCG. In some embodiments, the fusion protein further comprises one or more uracil glycosylase inhibitors (UGIs). In some embodiments, the fusion protein further comprises one or more nuclear localization sequences (NLSs).

[0030] In some embodiments, the base editing system is capable of modifying one or more single target nuclear bases to effect at least one amino acid substitution in the CD117 polypeptide. In some embodiments, the single target nuclear base is cytosine (C), and wherein the modification comprises converting the C to thymine (T). In some embodiments, the single target nuclear base is adenosine (A), and wherein the modification comprises converting the A to guanine (G). In some embodiments, the at least one amino acid substitution is a naturally occurring mutation. In some embodiments, the at least one amino acid substitution is located in domain 1, 2, 3, 4, or 5 of CD117. In some embodiments, the at least one amino acid substitution is selected from the group consisting of: T13A, S35P, I39V, H40R, K43G, K43R, S44P, D45G, I47V, D52G, E53G, I54V, R55G, L56P, L57P, T59A, F63P, V64A, K65E, K65R, W66R, T67A, D72G, E73G, T74A, N75D, N75G, E76G, N77G, N77S, N77Y, K78E, K78R, Q79R, N80G, N80S, E81G, E81D, W82R, I83T, I83V, T84A, E85G, K86E, E88G, T90A, N99G, K100G, H101R, K116R, V120A, S123P, L124P, Y125H, K127G, K127R, E128G, D129G, D129E, N130G, N130D, D131G, D131N, T132A, T144A, N145D, N145G, N145Y, N145S, Y146C, K149E, K149G, Q152R, K154E, K154G, R161G, F162P, F162L, I163T, I163V, D165G, M171T, I172T, I172V, K173G, S174G, K176G, Q190R, E191G, K193G, V195A, L196P, S197P, E198G, K199G, F200P, I201T, I201V, L202P, V213A, V214A, S215P, V216A, K218G, K218R, S220G, Y221C, E225G, E227G, E228G, T230A, S240G, Y243C, K247G, R248G, Q256R, E257G, E257D, K258E, K258G, Y259C, Y259H, N260G, D266G, N268D, N268G, Y269C, T276A, I277T, I277V, S279P, R281G,V282A, S285P, N293G, N294G, T295A, F296P, S298P, N300G, N300S, T302A, T303A, T304A, M318V, T322A, V323A, F324L, N326G, D327G, D332G, I334V, K342E, K342G, K342R, Q347R, Y350H, M351T, R353G, T354A, T354I, K358G, E360G, D361G, K364G, E366G, N367G, H378R, T380A, R381G, K383G, T385A, T389A, D398G, V399A, N400G, V407A, Y408H, E414G, I415V, T417A, Y418C, Y418H, D419G, R420G, L421P, V422A, N423S, M425V, E435G, I438M, D439G, V454A, L455P, V457A, V459A, Q460R, T461A, N463G, S464P, S465P, F469P, K471E, K471G, L472P, V473A, Q475R, S476G, I478M, I478V, D479G, S481G, F483P, K484G, N486G, N486S, T488A, Y494C, N495D, D496G, K499E, Y503C, Y503H, F504P, and N505G. In some embodiments, the amino acid substitutions in the at least one CD117 are selected from the group consisting of: T13A, I39V, H40R, D45G, D52G, E53G, I54V, R55G, T59A, K65E, K65R, T67A, E76G, N77G, N77S, N77Y, K78E, Q79R, N80G, E81G, E81D, N99G, K100G, H101R, L124P, Y125H, D129E, D129G, N130D, N130G, D131G, D131N, T132A, T144A, N145D, N145G, N145Y, N145S, Y146C, F162P, F162L, I163T, I163V, M171T, I172T, V195A, L196P, K199G, I201V, S220G, Y221C, Q256R, E257D, E257G, K258E, Y259C, T303A, T304A, M318V, T322A, V323A, F324L, K342G, K342R, Y350H, M351T, R353G, T354A, T354I, R381G,K383G, T385A, Y418C, D419G, R420G, I438M, D439G, Q460R, T461A, K484G, N486G, and T488A. In some embodiments, the at least one amino acid substitution is a naturally occurring mutation selected from the group consisting of: T13A, N77S, D129E, N130D, D131N, T144A, Y221C, E257D, T322A, T354I, D419G.,

[0031] In some embodiments, the napDNAbp is an inactivated nuclease or nickase variant. In some embodiments, the napDNAbp comprises a Cas9, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ polynucleotide or a portion thereof. In some embodiments, the napDNAbp comprises a Cas9 polynucleotide or a portion thereof. In some embodiments, the napDNAbp comprises inactivated Cas9 (dCas9) or Cas9 nickase (nCas9). In some embodiments, the napDNAbp is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In some embodiments, the napDNAbp comprises a modified protospacer adjacent motif (PAM) specificity. In some embodiments, the modified PAM is specific for the nucleic acid sequence 5'-NGC-3'. In some embodiments, the deaminase domain is capable of deaminating cytidine or adenine within DNA. In some embodiments, the deaminase domain is a cytidine deaminase domain. In some embodiments, the cytidine deaminase is an APOBEC deaminase or a derivative thereof. In some embodiments, the deaminase domain is an adenosine deaminase domain. In some embodiments, the adenosine deaminase domain is a TadA deaminase domain. In some embodiments, the adenosine deaminase is a TadA*8 variant. In some embodiments, the adenosine deaminase is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.21, TadA*8.22, TadA*8.23, or TadA*8.24. In some embodiments, the deaminase is a monomer or a heterodimer.

[0032] In one aspect, the present invention provides a polynucleotide encoding any base editor system as provided herein. In another aspect, the present invention provides a cell produced by any method provided herein. Also, in one aspect, the present invention provides a cell produced by introducing any base editor system or any polynucleotide provided herein into a cell or its progenitor cell. In some embodiments, the cell is produced ex vivo or in vitro. In some embodiments, the cell is a hematopoietic stem cell, a common myeloid progenitor cell, a proerythroblast, or an erythrocyte. In some embodiments, the cell is a CD34+ cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell.

[0033] In one aspect, the present invention provides an isolated cell or cell population for any cell proliferation or expansion as provided herein.

[0034] Also, in one aspect, the present invention provides a pharmaceutical composition comprising an effective amount of any cell provided herein.

[0035] In another aspect, the present invention provides a method of treating a subject having a hemoglobinopathy, a blood cancer, or a myeloproliferative disorder. In some embodiments, the method comprises administering to the subject any pharmaceutical composition provided herein. In some embodiments, the method comprises administering to the subject a conditioning regimen comprising an antibody, an antibody-drug conjugate, or a chimeric antigen receptor-expressing T cell that selectively binds to a cell surface protein and any cell provided herein, thereby treating a hemoglobinopathy, a blood cancer, or a myeloproliferative disorder. In some embodiments, the cell surface protein is a CD117 antibody. In some embodiments, the antibody, antibody-drug conjugate, or chimeric antigen receptor-expressing T cell selectively binds to domain 1, 2, 3, 4, or 5 of CD117. In some embodiments, the antibody is a monoclonal antibody. In some embodiments, the administration of the antibody is sequential or simultaneous. In some embodiments, the hemoglobinopathy is sickle cell disease or hereditary persistence of fetal hemoglobin (HPFH). In another aspect, the present invention provides a method of treating a subject having an immunodeficiency. In some embodiments, the immunodeficiency is severe combined immunodeficiency (SCID).

[0036] In some embodiments, the cell is autologous to the subject. In some embodiments, the cell is allogeneic to the subject. In some embodiments, the subject is a mammalian. In some embodiments, the mammalian is a human.

[0037] In one aspect, the present invention provides a kit comprising any cell provided herein, any base editor system provided herein, any polynucleotide provided herein, or any pharmaceutical composition provided herein. In some embodiments, the kit includes written instructions for treating hemoglobinopathy, blood cancer, myeloproliferative disease, or immunodeficiency using the kit.

[0038] The descriptions and examples herein detail embodiments of the present disclosure. Any composition or method provided herein can be combined with one or more of any other compositions and methods provided herein. It should be understood that the present disclosure is not limited to the specific embodiments described herein and can thus vary. Those skilled in the art will recognize that there are various changes and modifications to the present disclosure, and these are included within its scope.

[0039] Unless otherwise indicated, the practice of some embodiments disclosed herein employs conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which are within the skill of the art. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (F.M. Ausubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR 2: A Practical Approach (M.J. MacPherson, B.D. Hames and G.R. Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (R.I. Freshney, ed. (2010)).

[0040] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0041] Although various features of the present disclosure may be described in the context of a single embodiment, these features may also be provided singly or in any suitable combination. Conversely, although the present disclosure may be described herein in the context of separate embodiments for clarity, the present disclosure may also be implemented in a single embodiment. The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.

[0042] The features of the present disclosure are specifically set forth in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, which utilize the principles of the present disclosure, and in view of the drawings described below.

[0043] Definitions

[0044] The following definitions supplement those in the art and are for the current application and are not attributed to any related or unrelated cases, e.g., any co-owned patents or applications. Although any methods and materials similar or equivalent to those described herein may be used in the practice of testing the present disclosure, the preferred materials and methods are described herein. Thus, the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting.

[0045] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. The following references provide one of ordinary skill in the art with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings given to them below, unless otherwise noted.

[0046] In this application, unless otherwise specifically stated, the use of the singular includes the plural. It must be noted that the singular forms "a", "an", and "the" used in the specification include plural referents unless the context clearly dictates otherwise. In this application, unless otherwise specified, the use of "or" means "and / or". Additionally, the use of the term "including" and other forms such as "include", "includes", and "included" is not restrictive.

[0047] As used in this specification and the claims, the terms "comprising" (and any form of comprising, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of including, such as "includes" and "include"), or "containing" (and any form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method or combination of the present disclosure, and vice versa. Additionally, the compositions of the present disclosure can be used to implement the methods of the present disclosure.

[0048] The term "about" or "approximately" means within an acceptable error range of a particular value as determined by a person of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the measurement system. For example, in accordance with the practice in the art, "about" can mean within one standard deviation or more than one standard deviation. Alternatively, "about" can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly for biological systems or processes, the term can mean within an order of magnitude, such as within 5 times, within 2 times of the value. Where a particular value is described in the application and the claims, unless otherwise stated, the meaning of the term "about" should be assumed to be within the acceptable error range of the particular value.

[0049] References in the specification to "some embodiments", "an embodiment", "one embodiment", or "other embodiments" refer to specific features, structures, or characteristics described in connection with the examples being included in at least some embodiments, but not necessarily all embodiments of the present disclosure.

[0050] "Adenosine deaminase" refers to a polypeptide or a fragment thereof that can catalyze the hydrolysis and deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolysis and deamination of adenosine to inosine or the hydrolysis and deamination of deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolysis and deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism, such as bacteria.

[0051] In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to a naturally occurring deaminase. In some embodiments, the adenosine deaminase is from bacteria, such as Escherichia coli, Staphylococcus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae, or Clostridium crescentus. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is an Escherichia coli TadA (ecTadA) deaminase or a fragment thereof.

[0052] For example, the deaminase domains are described in International PCT Application Nos. PCT / 2017 / 045381 (WO 2018 / 027078) and PCT / US2016 / 058344 (WO 2017 / 070632), which are hereby incorporated by reference in their entireties. See also Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3:eaao4774 (2017)), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec;19(12):770-788.doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.

[0053] The wild-type TadA (wt) adenosine deaminase has the following sequence (also referred to as the TadA reference sequence):

[0054] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD.

[0055] In some embodiments, the adenosine deaminase comprises an alteration of the following sequence:

[0056] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD

[0057] (also referred to as TadA*7.10).

[0058] In some embodiments, TadA*7.10 comprises at least one alteration. In some embodiments, TadA*7.10 comprises an alteration at amino acid 82 and / or 166. In certain embodiments, variants of the above sequence comprise one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. The alteration Y123H refers to the alteration H123Y in TadA*7.10 reverting back to Y123H TadA (wt). In other embodiments, variants of the TadA*7.10 sequence comprise a combination of alterations selected from the group consisting of: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R.

[0059] In other embodiments, the present invention provides adenosine deaminase variants that lack, for example, TadA*8 and that contain a C-terminal deletion starting from residue 149, 150, 151, 152, 153, 154, 155, 156, or 157, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a TadA (e.g., TadA*8) monomer that contains one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA 7. In other embodiments, the adenosine deaminase variant is a TadA (e.g., TadA*8) monomer that contains the following alterations: Y147T+Q154R; Y147T+Q154S; Y147R+Q154S; V82S+Q154S; V82S+Y147R; V82S+Q154R; V82S+Y123H; I76Y+V82S; V82S+Y123H+Y147T; V82S+Y123H+Y147R; V82S+Y123H+Q154R; Y147R+Q154R+Y123H; Y147R+Q154R+I76Y; Y147R+Q154R+T166R; Y123H+Y147R+Q154R+I76Y; V82S+Y123H+Y147R+Q154R; and I76Y+V82S+Y123H+Y147R+Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA 7.

[0060] In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains, each domain having one or more of the following alterations Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a homodimer comprising two adenosine deaminase domains (e.g., TadA*8), each domain having a combination of alterations selected from the group consisting of: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA 7.

[0061] In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type TadA adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8), which comprises one or more of the following alterations Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a wild-type TadA adenosine deaminase domain and an adenosine deaminase variant domain (e.g., TadA*8), which comprises a combination of alterations selected from the group consisting of: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA 7.

[0062] In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain (such as TadA*8), which comprises one or more of the following alterations: Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA. In other embodiments, the adenosine deaminase variant is a heterodimer comprising a TadA*7.10 domain and an adenosine deaminase variant domain (such as TadA*8), which comprises a combination of alterations selected from: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R or I76Y + V82S + Y123H + Y147R + Q154R, relative to the corresponding mutations in TadA*7.10, the TadA reference sequence, or another TadA. In one embodiment, the adenosine deaminase is TadA*8, which comprises or consists essentially of the following sequence or a fragment thereof having adenosine deaminase activity:

[0063] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD。

[0064] In some embodiments, the TadA*8 is truncated. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 N-terminal amino acid residues relative to the full-length TadA*8. In some embodiments, the truncated TadA*8 lacks 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 C-terminal amino acid residues relative to the full-length TadA*8. In some embodiments, the adenosine deaminase variant is the full-length TadA*8.

[0065] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from one of the following:

[0066] Staphylococcus aureus (S. aureus) TadA:

[0067] MGSHMTNDIYFMTLAIEEAKKAAQLGEVPIGAIITKDDEVIARAHNLRETLQQPTAHAEHIAIERAAKVLGSWRLEGCTLYVTLEPCVMCAGTIVMSRIPRVVYGADDPKGGCSGSLMNLLQQSNFNHRAIVDKGVLKEACSTLLTTFFKNLRANKKSTN

[0068] Bacillus subtilis (B. subtilis) TadA:

[0069] MTQDELYMKEAIKEAKKAEEKGEVPIGAVLVINGEIIARAHNLRETEQRSIAHAEMLVIDEACKALGTWRLEGATLYVTLEPCPMCAGAVVLSRVEKVVFGAFDPKGGCSGTLMNLLQEERFNHQAEVVSGVLEEECGGMLSAFFRELRKKKKAARKNLSE

[0070] Salmonella typhimurium (S. typhimurium) TadA:

[0071] MPPAFITGVTSLSDVELDHEYWMRHALTLAKRAWDEREVPVGAVLVHNHRVIGEGWNRPIGRHDPTAHAEIMALRQGGLVLQNYRLLDTTLYVTLEPCVMCAGAMVHSRIGRVVFGARDAKTGAAGSLIDVLHHPGMNHRVEIIEGVLRDECATLLSDFFRMRRQEIKALKKADRAEGAGPAV

[0072] Shewanella putrefaciens TadA:

[0073] MDEYWMQVAMQMAEKAEAAGEVPVGAVLVKDGQQIATGYNLSISQHDPTAHAEILCLRSAGKKLENYRLLDATLYITLEPCAMCAGAMVHSRIARVVYGARDEKTGAAGTVVNLLQHPAFNHQVEVTSGVLAEACSAQLSRFFKRRRDEKKALKLAQRAQQGIE

[0074] Haemophilus influenzae F3031 TadA:

[0075] MDAAKVRSEFDEKMMRYALELADKAEALGEIPVGAVLVDDARNIIGEGWNLSIVQSDPTΑΗAEIIALRNGAKNIQNYRLLNSTLYVTLEPCTMCAGAILHSRIKRLVFGASDYKTGAIGSRFHFFDDYKMNHTLEITSGVLAEECSQKLSTFFQKRREEKKIEKALLKSLSDK Clostridium crescentus TadA:

[0076] MRTDESEDQDHRMMRLALDAARAAAEAGETPVGAVILDPSTGEVIATAGNGPIAAHDPTAHAEIAAMRAAAAKLGNYRLTDLTLVVTLEPCAMCAGAISHARIGRVVFGADDPKGGAVVHGPKFFAQPTCHWRPEVTGGVLADESADLLRGFFRARRKAKI

[0077] Geobacter sulfurreducens TadA:

[0078] MSSLKKTPIRDDAYWMGKAIREAAKAAARDEVPIGAVIVRDGAVIGRGHNLREGSNDPSAHAEMIAIRQAARRSANWRLTGATLYVTLEPCLMCMGAIILARLERVVFGCYDPKGGAAGSLYDLSADPRLNHQVRLSPGVCQEECGTMLSDFFRDLRRRKKAKATPALFIDERKVPPEP

[0079] TadA*7.10

[0080] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAE IMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD。

[0081] "Adenosine deaminase base editor 8 (ABE8) polypeptide" or "ABE8" refers to a base editor as defined herein that comprises an adenosine deaminase variant that comprises a change at amino acid position 82 and / or 166 of the following reference sequence:

[0082] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD

[0083] In some embodiments, ABE8 comprises further changes relative to the reference sequence, as described herein.

[0084] "Adenosine deaminase base editor 8 (ABE8) polypeptide" refers to a polynucleotide encoding ABE8.

[0085] "Administering" as used herein refers to providing a patient or subject with one or more of the compositions described herein. For example but not limited to, administering a composition, such as by injection, can be performed by intravenous (iv) injection, subcutaneous (sc) injection, intradermal (id) injection, intraperitoneal (ip) injection or intramuscular (im) injection. One or more such routes can be employed. Parenteral administration can be, for example, by bolus injection or by infusion over time. Alternatively, or concurrently, administration can be by the oral route.

[0086] "Agent" refers to any small molecule compound, antibody, nucleic acid molecule or polypeptide, or fragment thereof.

[0087] "Alter" refers to a change (increase or decrease) in the structure, expression level or activity of a gene or polypeptide, as detected by standard methods known in the art such as those described herein. As used herein, an alter includes a 10% change, 25% change, 40% change and a 50% or greater change in the expression level.

[0088] "Improve" refers to reducing, inhibiting, attenuating, weakening, preventing or stabilizing the development or progression of a disease.

[0089] "Analogue" refers to a molecule that is not identical but has similar functional or structural characteristics. For example, a polypeptide analogue retains the biological activity of the corresponding naturally occurring polypeptide while having certain biochemical modifications that enhance analogue function relative to the naturally occurring polypeptide. Such biochemical modifications can increase the protease resistance, membrane permeability or half-life of the analogue without altering, for example, ligand binding. Analogues can include non-natural amino acids.

[0090] As used herein, the term "antibody" refers to an immunoglobulin molecule that specifically binds to a particular antigen or reacts immunologically with a particular antigen, including polyclonal, monoclonal, genetically engineered and other modified forms of antibodies, including but not limited to chimeric antibodies, humanized antibodies, heteroconjugate antibodies (e.g., bispecific, trispecific and tetra-specific antibodies, diabodies, tribodies and tetrabodies) and antigen-binding fragments of antibodies, including for example Fab', F(ab')2, Fab, Fv, rIgG and scFv fragments. Unless otherwise specified, the term "monoclonal antibody" (mAb) includes both the intact molecule and antibody fragments capable of specifically binding to the target protein (including for example Fab and F(ab’)2 fragments). As used herein, Fab and F(ab')2 fragments refer to antibody fragments that lack the Fc fragment of the intact antibody. Examples of these antibody fragments are described herein.

[0091] "Base editor (BE)" or "nucleobase editor polypeptide (NBE)" refers to a reagent that binds to a polynucleotide and has nucleobase modification activity. In various embodiments, the base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) and a polynucleotide-programmable nucleotide-binding domain that binds to a guide polynucleotide (e.g., guide RNA). In various embodiments, the reagent is a biomolecular complex comprising a protein domain having base editing activity, i.e., capable of modifying a base (e.g., in DNA) within a nucleic acid molecule (e.g., A, T, C, G, or U). In some embodiments, the polynucleotide-programmable DNA-binding domain is fused or linked to a deaminase domain. In one embodiment, the reagent is a fusion protein comprising one or more domains having base editing activity. In another embodiment, the protein domain having base editing activity is linked to a guide RNA (e.g., via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to the deaminase). In some embodiments, the domain having base editing activity is capable of deaminating a base within a nucleic acid molecule. In some embodiments, the base editor is capable of deaminating one or more bases within a DNA molecule.

[0092] In some embodiments, the base editor is capable of deaminating cytosine (C) or adenosine (A) in DNA. In some embodiments, the base editor is capable of deaminating both cytosine (C) and adenosine (A) in DNA. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenosine base editor (ABE). In some embodiments, the base editor is an adenosine base editor (ABE) and a cytidine base editor (CBE). In some embodiments, the base editor is a nuclease-free Cas9 (dCas9) fused to an adenosine deaminase. In some embodiments, the Cas9 is a circularly permuted Cas9 (e.g., spCas9 or saCas9). Circularly permuted Cas9s are known in the art and described, for example, in Oakes et al., Cell 176, 254–267, 2019. In some embodiments, the base editor is fused to a base excision repair inhibitor, such as the UGI domain or the dISN domain. In some embodiments, the fusion protein comprises a Cas9 nickase fused to a deaminase and a base excision repair inhibitor, such as the UGI or dISN domain. In other embodiments, the base editor is a base-free base editor.

[0093] In some embodiments, the adenosine deaminase is evolved from TadA. In some embodiments, the polynucleotide programmable DNA binding domain is a CRISPR-associated (e.g., Cas or Cpf1) enzyme. In some embodiments, the base editor is a catalytically inactive Cas9 (dCas9) fused to a deaminase domain. In some embodiments, the base editor is a Cas9 nickase (nCas9) fused to a deaminase domain. In some embodiments, the base editor is fused to a base excision repair inhibitor (BER). In some embodiments, the base excision repair inhibitor is a uracil DNA glycosylase inhibitor (UGI). In some embodiments, the base excision repair inhibitor is an inosine base excision repair inhibitor. Details of the base editors are described in International PCT Application Nos. PCT / 2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), which are hereby incorporated by reference in their entirety. Additionally, see Komor, A.C., et al., “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016); Gaudelli, N.M., et al., “Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage” Nature 551, 464-471 (2017); Komor, A.C., et al., “Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity” Science Advances 3: eaao4774 (2017), and Rees, H.A., et al., “Base editing: precision chemistry on the genome and transcriptome of living cells.” Nat Rev Genet. 2018 Dec; 19(12):770-788. doi:10.1038 / s41576-018-0059-1, the entire contents of which are incorporated herein by reference.

[0094] In some embodiments, the base editor is generated by cloning an adenosine deaminase variant (e.g., TadA*8) into a scaffold comprising circularly permuted Cas9 (e.g., saCAS9) and a bipartite nuclear localization sequence (e.g., ABE8). Circularly permuted Cas9s are known in the art and described, for example, in Oakes et al., Cell 176, 254–267, 2019. Exemplary circularly permuted sequences are as follows, where the bold sequences represent sequences derived from Cas9, the italic sequences represent linker sequences, and the underlined sequences represent the bipartite nuclear localization sequence.

[0095] CP5 (with MSP “NGC = Pam variant conventional Cas9-like NGG with mutations” PID = protein interaction domain and “D10A” nickase):

[0096]

[0097]

[0098] In some embodiments, the ABE8 is selected from the base editors of Tables 10, 11, or 13 below. In some embodiments, ABE8 contains an adenosine deaminase variant evolved from TadA. In some embodiments, the adenosine deaminase variant of ABE8 is the TadA*8 variant as described in Tables 8, 10, 11, or 13 below. In some embodiments, the adenosine deaminase variant is a TadA*7.10 variant (e.g., TadA*8) comprising one or more alterations selected from the group consisting of Y147T, Y147R, Q154S, Y123H, V82S, T166R, and / or Q154R. In various embodiments, ABE8 comprises a TadA*7.10 variant (e.g., TadA*8) having a combination of alterations selected from the group: Y147T + Q154R; Y147T + Q154S; Y147R + Q154S; V82S + Q154S; V82S + Y147R; V82S + Q154R; V82S + Y123H; I76Y + V82S; V82S + Y123H + Y147T; V82S + Y123H + Y147R; V82S + Y123H + Q154R; Y147R + Q154R + Y123H; Y147R + Q154R + I76Y; Y147R + Q154R + T166R; Y123H + Y147R + Q154R + I76Y; V82S + Y123H + Y147R + Q154R; and I76Y + V82S + Y123H + Y147R + Q154R.

[0099] In some embodiments, ABE8 is a monomeric construct. In some embodiments, ABE8 is a heterodimeric construct. In some embodiments, the ABE8 base editor comprises the sequence:

[0100] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCTFFRMPRQVFNAQKKAQSSTD

[0101] For example, the adenine base editor ABE for the base editing compositions, systems, and methods described herein has a nucleic acid sequence (8,877 base pairs) (Addgene, Watertown, MA.; Komor NM et al., 2017, Sci Adv., 30; 3(8):2017 Nov 23; 551(7681):464 - 471. doi:10.1038 / nature24644; Koblan LW, et al., Nat Biotechnol. 2018 Oct; 36(9):843 - 846. doi:10.1038 / nbt.4172.) as provided below. Also included are polynucleotide sequences having at least 95% or higher identity to the ABE nucleic acid sequence.

[0102]

[0103] For example, the cytidine base editor (CBE) used in the base editing compositions, systems, and methods described herein has the following nucleic acid sequence (8,877 base pairs) (Addgene, Watertown, MA.; Komor AC et al., 2017, Sci Adv., 30; 3(8): eaao4774. doi: 10.1126 / sciadv.aao4774) provided below. Also included are polynucleotide sequences having at least 95% or higher identity to the BE4 nucleic acid sequence.

[0104]

[0105]

[0106]

[0107]

[0108] In some embodiments, the cytidine base editor is BE4 having a nucleic acid sequence selected from:

[0109] Original BE4 nucleic acid sequence:

[0110]

[0111] BE4 codon-optimized 1 nucleic acid sequence:

[0112]

[0113] BE4 codon-optimized 2 nucleic acid sequence:

[0114]

[0115] "Base editing activity" refers to the action of chemically altering a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is cytidine deaminase activity, for example converting a target C·G to T·A. In another embodiment, the base editing activity is adenosine or adenine deaminase activity, for example converting a target A·T to C·G. In another embodiment, the base editing activity is cytidine deaminase activity, for example converting a target C·G to T·A, and adenosine or adenine deaminase activity, for example converting A·T to G·C.

[0116] The term "base editor system" refers to a system for editing a nucleobase of a target nucleotide sequence. In various embodiments, a base editor (BE) system comprises (1) a polynucleotide programmable nucleotide binding domain, a deaminase domain for deaminating a nucleobase in a target nucleotide sequence (e.g., a cytidine deaminase or an adenosine deaminase); (2) one or more guide polynucleotides (e.g., guide RNA) that bind to the polynucleotide programmable nucleotide binding domain. In various embodiments, a base editor (BE) system comprises a nucleobase editor domain selected from adenosine deaminase or cytidine deaminase, and a domain having nucleic acid sequence specific binding activity. In some embodiments, a base editor system comprises (1) a base editor (BE) comprising a polynucleotide programmable DNA binding domain and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence; and (2) one or more guide RNAs that bind to the polynucleotide programmable DNA binding domain. In some embodiments, the polynucleotide programmable acid binding domain is a polynucleotide programmable DNA binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is an adenine or adenosine base editor (ABE) or a cytidine base editor (CBE).

[0117] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease that comprises a Cas9 protein or a fragment thereof (e.g., a protein comprising an active, inactive or partially active DNA cleavage domain of Cas9, and / or a binding domain for gRNA Cas9). The Cas9 nuclease is sometimes also referred to as the casnl nuclease or a CRISPR (clustered regularly interspaced short palindromic repeats)-associated nuclease. An exemplary Cas9 is Streptococcus pyogenes Cas9 (spCas9), the amino acid sequence of which is provided below:

[0118] (Single underline: HNH domain; double underline: RuvC domain)

[0119] The term "conservative amino acid substitution" or "conservative mutation" refers to the substitution of one amino acid with another amino acid having common properties. One functional way to define common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of homologous organisms (Schulz, G.E. and Schirmer, R.H., Principles of Protein Structure, Springer-Verlag, New York (1979)). Based on such an analysis, groups of amino acids can be defined where amino acids within a group preferentially exchange with each other and are thus most similar to each other in terms of their effect on the overall protein structure (Schulz, G.E. and Schirmer, R.H., supra). Non-limiting examples of conservative mutations include amino acid substitutions of amino acids such as lysine for arginine and vice versa, so that a positive charge can be maintained; glutamic acid for aspartic acid and vice versa to maintain a negative charge; serine for threonine so that a free -OH can be maintained; and glutamine for asparagine so that a free -NH2 can be maintained.

[0120] The term "coding sequence" or "protein-coding sequence", as used interchangeably herein, refers to a polynucleotide fragment that encodes a protein. The coding sequence may also be referred to as an open reading frame. The region or sequence has a start codon near the 5'-end and a stop codon near the 3'-end. Stop codons useful for the base editors described herein include the following:

[0121]

[0122] As used herein, the terms "condition" and "conditioning" refer to the process by which a patient is prepared to receive a graft containing hematopoietic stem cells. Such procedures facilitate the engraftment of the hematopoietic stem cell graft (e.g., as inferred by the sustained increase in the number of viable hematopoietic stem cells in a blood sample isolated from a patient who has undergone a conditioning procedure and subsequent hematopoietic stem cell transplantation). According to the methods described herein, a patient can be conditioned for hematopoietic stem cell transplantation therapy by administering to the patient an antibody or antigen-binding fragment thereof that is capable of binding to an antigen expressed by hematopoietic stem cells, such as CD117, CXCR4, CD135, CD90, CD45, and / or CD34. Such antibodies are expected to act via complement-mediated cytotoxicity and antibody-dependent cell-mediated cytotoxicity. As described herein, the transplanted cells have been edited such that the antibody no longer binds to the antigen (e.g., CD117, CXCR4, CD135, CD90, CD45, and / or CD34). Administration of an antibody, an antigen-binding fragment thereof, a drug-antibody conjugate, or a chimeric antigen receptor-expressing T cell (CAR-T) that is capable of binding to one or more antigens (e.g., CD117, CXCR4, CD135, CD90, CD45, CD34) to a patient in need of hematopoietic stem cell transplantation therapy can facilitate the engraftment of the hematopoietic stem cell graft, e.g., by selectively depleting endogenous hematopoietic stem cells, thereby creating a niche to be filled by the exogenous hematopoietic stem cell graft.

[0123] "Cytidine deaminase" refers to a polypeptide or fragment thereof that is capable of catalyzing a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, the cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. PmCDA1 derived from Petromyzon marinus (cytidine deaminase 1), AID (activation-induced cytidine deaminase; AICDA) derived from mammals (e.g., human, pig, cow, horse, monkey, etc.), and APOBEC are exemplary cytidine deaminases.

[0124] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes the hydrolytic deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine to uracil. In some embodiments, the deaminase is an adenosine deaminase that catalyzes the hydrolytic deamination of adenine to hypoxanthine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine or adenine (A) to inosine (I). In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases provided herein (e.g., engineered adenosine deaminases, evolved adenosine deaminases) can be from any organism, such as a bacterium. In some embodiments, the adenosine deaminase is from a bacterium, such as Escherichia coli, Staphylococcus aureus, Salmonella typhimurium, Shewanella putrefaciens, Haemophilus influenzae, or Clostridium crescentus. In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the deaminase or deaminase domain is a variant of a naturally occurring deaminase from an organism such as a human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse. In some embodiments, the deaminase or deaminase domain does not exist in nature. For example, in some embodiments, the deaminase or deaminase domain is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to a naturally occurring deaminase.

[0125] "Detecting" refers to identifying the presence, absence, or amount of an analyte to be detected. In one embodiment, detecting a sequence alteration in a polynucleotide or polypeptide. In another embodiment, detecting the presence of an indel.

[0126] "Detectable label" refers to a composition that renders the latter detectable by spectroscopic, photochemical, biochemical, immunochemical, or chemical means when linked to a molecule of interest. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., commonly used in enzyme-linked immunosorbent assay (ELISA)), biotin, digoxin, or haptens.

[0127] "Disease" refers to any disorder or condition that impairs or interferes with the normal function of cells, tissues, or organs. Exemplary diseases include diseases suitable for treatment by hematopoietic stem cell transplantation, such as thalassemia, sickle cell disease (SCD), or adenosine deaminase deficiency.

[0128] "Effective amount" refers to the amount of a reagent or active compound (such as a base editor described herein) required to improve the symptoms of a disease relative to an untreated patient or an individual without the disease (i.e., a healthy person), or the amount of a reagent or active compound sufficient to elicit a desired biological response. The effective amount of an active compound for practicing the present disclosure to treat a disease varies depending on the mode of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosage regimen. Such an amount is referred to as an "effective" amount. In one embodiment, the effective amount is an amount of the base editor of the present invention sufficient to introduce a change in a gene of interest in a cell (e.g., an in vitro or in vivo cell). In one embodiment, the effective amount is the amount of base editor required to achieve a therapeutic effect. Such a therapeutic effect does not require sufficient to alter the disease-causing gene in all cells of the subject, tissue, or organ, but only to alter the disease-causing gene present in about 1%, 5%, 10%, 25%, 50%, 75% or more of the cells in the subject, tissue, or organ. In one embodiment, the effective amount is sufficient to improve one or more symptoms of the disease.

[0129] In some embodiments, the effective amount of a fusion protein provided herein, such as the effective amount of a nuclear base editor comprising an nCas9 domain and a deaminase domain (such as an adenosine deaminase, a cytidine deaminase), refers to the amount sufficient to induce editing of a target site specifically bound and edited by the nuclear base editor described herein. As will be understood by those skilled in the art, the effective amount of a reagent (e.g., a fusion protein) can vary depending on various factors, e.g., depending on the desired biological response, e.g., depending on the specific allele, genome, or target site being edited, depending on the cell or tissue being targeted, and / or depending on the reagent being used.

[0130] In some embodiments, an effective amount of a fusion protein provided herein, such as an effective amount of a fusion protein comprising an nCas9 domain and a deaminase domain, can refer to an amount of the fusion protein sufficient to induce editing of a target site that is specifically bound and edited by the fusion protein described herein. As will be understood by those skilled in the art, the effective amount of a reagent (e.g., a fusion protein, nuclease, hybrid protein, protein dimer, protein complex (or protein dimer) and polynucleotide complex, or polynucleotide) can vary depending on various factors, e.g., depending on the desired biological response, e.g., depending on the particular allele, genome or target site being edited, the cell or tissue being targeted, and / or the reagent being used.

[0131] "Fragment" refers to a portion of a polypeptide or nucleic acid molecule. The portion comprises at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% of the full length of the reference nucleic acid molecule or polypeptide. Fragments can comprise 10, 20, 30, 40, 50, 60, 70, 80, 90 or 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1000 nucleotides or amino acids.

[0132] "Guide RNA" or "gRNA" refers to a polynucleotide that is specific for a target sequence and can form a complex with a polynucleotide programmable nuclease domain protein (such as Cas9 or Cpf1). In one embodiment, the guide polynucleotide is guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule. The gRNA that exists as a single RNA molecule can be referred to as single guide RNA (sgRNA), but "gRNA" is used interchangeably to refer to guide RNA that exists as a single molecule or a complex of two or more molecules. Generally, the gRNA that exists as a single RNA species contains two domains: (1) a domain that is homologous to the target nucleic acid (e.g., directs the binding of the Cas9 complex to the target); (2) a domain that binds to the Cas9 protein. In some embodiments, domain (2) corresponds to a sequence called tracrRNA and contains a stem-loop structure. For example, in some embodiments, domain (2) is identical or homologous to the tracrRNA provided in Jinek et al., Science 337:816-821 (2012), the entire content of which is incorporated herein by reference. Other examples of gRNAs (e.g., those including domain 2) can be found in the patent applications titled "Switchable Cas9 Nucleases and Uses Thereof" (US20160208288) and "Delivery System For Functional Nucleases" (US 9,737,604), the entire contents of which are incorporated herein by reference in their entirety. In some embodiments, the gRNA contains two or more of domains (1) and (2) and can be referred to as an "extended gRNA". As described herein, the extended gRNA will bind two or more Cas9 proteins and bind to the target nucleic acid at two or more different regions. The gRNA contains a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to the target site, providing sequence specificity to the nuclease:RNA complex.

[0133] As used herein, the term "hematopoietic stem cell" ("HSC") refers to immature blood cells that have the ability to self-renew and differentiate into mature blood cells, including but not limited to granulocytes (e.g., promyelocytes, neutrophils), eosinophils, basophils), erythrocytes (e.g., reticulocytes, red blood cells), platelets (e.g., megakaryocytes, platelet-producing megakaryocytes, platelets), monocytes (e.g., monocytes, macrophages), dendritic cells, microglia, osteoclasts, and lymphocytes (e.g., NK cells, B cells, and T cells). Such cells can include CD34+ cells. CD34+ cells are immature cells that express the CD34 cell surface marker. In humans, CD34+ cells are thought to include a subset of cells with the above stem cell characteristics, while in mice, HSCs are CD34-. In addition, HSCs also refer to long-term repopulating HSCs (LT-HSCs) and short-term repopulating HSCs (ST-HSCs). LT-HSCs and ST-HSCs are distinguished based on functional potential and cell surface marker expression. For example, human HSCs are CD34+, CD38-, CD45RA-, CD90+, CD49F+, and lin-

[0134] (negative for mature lineage markers, including CD2, CD3, CD4, CD7, CD8, CD10, CD11B, CD19, CD20, CD56, CD235A). In mice, bone marrow LT-HSCs are CD34-, SCA-1+, C-kit+, CD135-, Slamfl / CD150+, CD48-, and lin-(negative for mature lineage markers, including Ter119, CD11b, Gr1, CD3, CD4, CD8, B220, IL7ra), while ST-HSCs are CD34+, SCA-1+, C-kit+, CD135-, Slamfl / CD150+, and lin-

[0135] (negative for mature lineage markers, including Ter119, CD1 1b, Gr1, CD3, CD4, CD8, B220, IL7ra). In addition, under steady-state conditions, ST-HSCs have lower quiescence and higher proliferative capacity compared to LT-HSCs. However, LT-HSCs have greater self-renewal potential (i.e., they can survive throughout adulthood and can be serially transplanted into successive recipients), while ST-HSCs have limited self-renewal (i.e., they can only survive for a limited time and do not have serial transplantation potential). Any of these HSCs can be used in the methods described herein. ST-HSCs are particularly useful because they are highly proliferative and can therefore generate differentiated progeny more quickly.

[0136] As used herein, the term "hematopoietic stem cell functional potential" refers to the functional characteristics of hematopoietic stem cells, including 1) pluripotency (the ability to differentiate into multiple different blood lineages, including but not limited to granulocytes (e.g., promyelocytes, neutrophils, eosinophils, basophils), erythrocytes (e.g., reticulocytes, red blood cells), platelets (e.g., megakaryocytes, platelet-producing megakaryocytes, platelets), monocytes (e.g., monocytes, macrophages), dendritic cells, microglia, osteoclasts, and lymphocytes (such as NK cells, B cells, and T cells)), 2) self-renewal (the ability of hematopoietic stem cells to produce daughter cells with the same potential as the parent cell, and the ability to recur repeatedly throughout the lifetime of an individual without exhaustion), and 3) the ability of hematopoietic stem cells or their progeny to be re-introduced into a transplant recipient, whereupon they return to the hematopoietic stem cell niche and re-establish productive and sustained hematopoietic function.

[0137] "Heterodimer" refers to a fusion protein containing two domains, such as the wild-type TadA domain and a variant of the TadA domain (e.g., TadA*8) or two variant TadA domains (e.g., TadA*7.10 and TadA*8 or two TadA*8 domains).

[0138] "Hybridization" refers to hydrogen bonding between complementary nucleobases, which can be Watson-Crick, Hoogsteen, or reverse Hoogsteen hydrogen bonds. For example, adenine and thymine are complementary nucleobases that pair by forming hydrogen bonds.

[0139] "Increase" refers to a positive change of at least 10%, 25%, 50%, 75%, or 100%.

[0140] The term "inhibitor of base repair", "base repair inhibitor", "IBR", or other grammatical equivalents refers to a protein that can inhibit the activity of a nucleic acid repair enzyme, such as a base excision repair enzyme. In some embodiments, the IBR is an inhibitor of inosine base excision repair. Exemplary inhibitors of base repair include inhibitors of APE1, Endo III, Endo IV, Endo V, Endo VIII, Fpg, hOGG1, hNEIL1, T7 Endol, T4PDG, UDG, hSMUG1, and hAAG. In some embodiments, the base repair inhibitor is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is an inhibitor of Endo V or hAAG. In some embodiments, the IBR is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is catalytically inactive EndoV or catalytically inactive hAAG. In some embodiments, the base repair inhibitor is uracil DNA glycosylase inhibitor (UGI). UGI refers to a protein that can inhibit the base excision repair enzyme uracil-DNA glycosylase. In some embodiments, the UGI domain comprises wild-type UGI or a fragment of wild-type UGI. In some embodiments, the UGI proteins provided herein include fragments of UGI and proteins homologous to UGI or UGI fragments. In some embodiments, the base repair inhibitor is an inhibitor of inosine base excision repair. In some embodiments, the base repair inhibitor is a catalytically inactive inosine-specific nuclease or a "dead inosine-specific nuclease". Without wishing to be bound by any particular theory, a catalytically inactive inosine glycosylase (e.g., alkyladenine glycosylase (AAG)) can bind inosine but cannot generate an abasic site or remove inosine, thereby spatially blocking the newly formed inosine moiety from the DNA damage / repair machinery. In some embodiments, the catalytically inactive inosine-specific nuclease is capable of binding inosine in a nucleic acid but not cleaving the nucleic acid. Non-limiting exemplary catalytically inactive inosine-specific nucleases include catalytically inactive alkyladenosine glycosylase (AAG nuclease), e.g., from human, and catalytically inactive endonuclease V (EndoV nuclease), e.g., from Escherichia coli. In some embodiments, the catalytically inactive AAG nuclease comprises the E125Q mutation or a corresponding mutation in another AAG nuclease.

[0141] An "intein" is a protein segment that can excise itself and join the remaining segments (exteins) with peptide bonds in a process called protein splicing. Inteins are also called "protein introns". The process by which an intein excises itself and joins the remaining part of the protein is referred to herein as "protein splicing" or "intein-mediated protein splicing". In some embodiments, the intein of a precursor protein (the protein containing the intein prior to intein-mediated protein splicing) is from two genes. Such an intein is referred to herein as a split intein (e.g., split intein-N and split intein-C). For example, in cyanobacteria, DnaE of the catalytic subunit a of DNA polymerase III is encoded by two separate genes, dnaE-n and dnaE-c. The intein encoded by the dnaE-n gene can be referred to herein as "intein-N". The intein encoded by the dnaE-c gene can be referred to herein as "intein-C".

[0142] Other intein systems can also be used. For example, synthetic inteins based on dnaE intein, Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C) intein pairs have been described (e.g., in Stevens et al., J Am Chem Soc. 2016 Feb. 24; 138(7):2162-5, incorporated herein by reference). Non-limiting examples of intein pairs that can be used according to the present disclosure include: Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Patent No. 8,394,604, incorporated herein by reference.

[0143] Exemplary nucleotide and amino acid sequences of inteins are provided.

[0144] DnaE intein-N DNA:

[0145] TGCCTGTCATACGAAACCGAGATACTGACAGTAGAATATGGCCTTCTGCCAATCGGGAAGATTGTGGAGAAACGGATAGAATGCACAGTTTACTCTGTCGATAACAATGGTAACATTTATACTCAGCCAGTTGCCCAGTGGCACGACCGGGGAGAGCAGGAAGTATTCGAATACTGTCTGGAGGATGGAAGTCTCATTAGGGCCACTAAGGACCACAAATTTATGACAGTCGATGGCCAGATGCTGCCTATAGACGAAATCTTTGAGCGAGAGTTGGACCTCATGCGAGTTGACAACCTTCCTAAT

[0146] DnaE intein-N protein: CLSYETEILTVEYGLLPIGKIVEKRIECTVYSVDNNGNIYTQPVAQWHDRGEQEVFEYCLEDGSLIRATKDHKFMTVDGQMLPIDEIFERELDLMRVDNLPN

[0147] DnaE intein-C DNA: ATGATCAAGATAGCTACAAGGAAGTATCTTGGCAAACAAAACGTTTATGATATTGGAGTCGAAAGAGATCACAACTTTGCTCTGAAGAACGGATTCATAGCTTCTAAT

[0148] Intein-C: MIKIATRKYLGKQNVYDIGVERDHNFALKNGFIASN

[0149] Cfa-N DNA:

[0150] TGCCTGTCTTATGATACCGAGATACTTACCGTTGAATATGGCTTCTTGCCTATTGGAAAGATTGTCGAAGAGAGAATTGAATGCACAGTATATACTGTAGACAAGAATGGTTTCGTTTACACACAGCCCATTGCTCAATGGCACAATCGCGGCGAACAAGAAGTATTTGAGTACTGTCTCGAGGATGGAAGCATCATACGAGCAACTAAAGATCATAAATTCATGACCACTGACGGGCAGATGTTGCCAATAGATGAGATATTCGAGCGGGGCTTGGATCTCAAACAAGTGGATGGATTGCCA

[0151] Cfa-N protein:

[0152] CLSYDTEILTVEYGFLPIGKIVEERIECTVYTVDKNGFVYTQPIAQWHNRGEQEVFEYCLEDGSIIRATKDHKFMTTDGQMLPIDEIFERGLDLKQVDGLP

[0153] Cfa-C DNA:

[0154] ATGAAGAGGACTGCCGATGGATCAGAGTTTGAATCTCCCAAGAAGAAGAGGAAAGTAAAGATAATATCTCGAAAAAGTCTTGGTACCCAAAATGTCTATGATATTGGAGTGGAGAAAGATCACAACTTCCTTCTCAAGAACGGTCTCGTAGCCAGCAAC

[0155] Cfa-C protein:

[0156] MKRTADGSEFESPKKKRKVKIISRKSLGTQNVYDIGVEKDHNFLLKNGLVASN

[0157] The intein-N and intein-C can be fused to the N-terminal portion and the C-terminal portion of split Cas9, respectively, for connecting the N-terminal portion and the C-terminal portion of split Cas9. For example, in some embodiments, the intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, i.e., forming a structure of N--[N-terminal portion of split Cas9]-[intein-N]--C. In some embodiments, the intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., forming a structure of N-[intein-C]--[C-terminal portion of split Cas9]-C. The intein-mediated protein splicing mechanism for connecting the protein to which the intein is fused (e.g., split Cas9) is known in the art, for example, as described by Shah et al., Chem Sci. 2014; 5(1):446-461, which is incorporated herein by reference. Methods for designing and using inteins are known in the art and are described, for example, in WO2014004336, WO2017132580, US20150344549, and US20180127780, each of which is incorporated herein by reference in its entirety.

[0158] The terms "isolated", "purified", or "biologically pure" refer to a material that is, to varying degrees, free of the components that are normally associated with it in its natural state. "Isolated" indicates the degree of separation from the original source or the surrounding environment. "Purified" indicates a degree of separation higher than isolation. A "purified" or "biologically pure" protein is sufficiently free of other materials such that any impurities do not substantially affect the biological properties of the protein or cause other adverse consequences. That is, if the nucleic acid or peptide of the present invention is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA technology, or is substantially free of chemical precursors or other chemicals when chemically synthesized, then the nucleic acid or peptide is purified. Purity and homogeneity are generally determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high performance liquid chromatography. The term "purified" may indicate that the nucleic acid or protein gives rise to substantially a single band in an electrophoretic gel. For proteins that can be modified (e.g., phosphorylated or glycosylated), different modifications may result in different separated proteins, which can be purified separately.

[0159] "Isolated polynucleotide" refers to a nucleic acid (e.g., DNA) that does not contain a gene that is flanked in the naturally occurring genome of the organism from which the nucleic acid molecule of the present invention is derived. Thus, the term includes, for example, recombinant DNA incorporated into a vector; into a plasmid or virus that replicates autonomously; or into the genomic DNA of a prokaryote or eukaryote; or exists as an independent molecule separate from other sequences (e.g., cDNA or genomic or cDNA fragments produced by PCR or restriction endonuclease digestion). In addition, the term includes RNA molecules transcribed from a DNA molecule, and recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence.

[0160] "Isolated polypeptide" refers to a polypeptide of the present invention that has been separated from its naturally associated components. Generally, a polypeptide is isolated when it is at least 60% (by weight) free of proteins and naturally occurring organic molecules. Preferably, the polypeptides of the present invention are at least 75% by weight, more preferably at least 90% by weight, and most preferably at least 99% by weight. The isolated polypeptides of the present invention can be obtained, for example, by extraction from natural sources, by expression of recombinant nucleic acids encoding such polypeptides; or by chemical synthesis of the protein. Purity can be measured by any suitable method, such as column chromatography, polyacrylamide gel electrophoresis, or by HPLC analysis.

[0161] "CD117 (C-kit; SCFR) polypeptide" refers to a polypeptide or a fragment thereof having at least about 95% amino acid sequence identity to the amino acid sequence of the polypeptide that binds anti-CD117 antibody provided by GenBank accession number NP_000213. In some embodiments, the CD117 polypeptide or its fragment has SCF signaling activity. The sequences of exemplary CD117 polypeptides are provided below:

[0162] >NP_000213.1 Mast / stem cell growth factor receptor Kit isoform 1 precursor

[0163] [Homo sapiens]

[0164] MRGARGAWDFLCVLLLLLRVQTGSSQPSVSPGEPSPPSIHPGKSDLIVRVGDEIRLLCTDPGFVKWTFEILDETNENKQNEWITEKAEATNTGKYTCTNKHGLSNSIYVFVRDPAKLFLVDRSLYGKEDNDTLVRCPLTDPEVTNYSLKGCQGKPLPKDLRFIPDPKAGIMIKSVKRAYHRLCLHCSVDQEGKSVLSEKFILKVRPAFKAVPVVSVSKASYLLREGEEFTVTCTIKDVSSSVYSTWKRENSQTKLQEKYNSWHHGDFNYERQATLTISSARVNDSGVFMCYANNTFGSANVTTTLEVVDKGFINIFPMINTTVFVNDGENVDLIVEYEAFPKPEHQQWIYMNRTFTDKWEDYPKSENESNIRYVSELHLTRLKGTEGGTYTFLVSNSDVNAAIAFNVYVNTKPEILTYDRLVNGMLQCVAAGFPEPTIDWYFCPGTEQRCSASVLPVDVQTLNSSGPPFGKLVVQSSIDSSAFKHNGTVECKAYNDVGKTSAYFNFAFKGNNKEQIHPHTLFTPLLIGFVIVAGMMCIIVMILTYKYLQKPMYEVQWKVVEEINGNNYVYIDPTQLPYDHKWEFPRNRLSFGKTLGAGAFGKVVEATAYGLIKSDAAMTVAVKMLKPSAHLTEREALMSELKVLSYLGNHMNIVNLLGACTIGGPTLVITEYCCYGDLLNFLRRKRDSFICSKQEDHAEAALYKNLLHSKESSCSDSTNEYMDMKPGVSYVVPTKADKRRSVRIGSYIERDVTPAIMEDDELALDLEDLLSFSYQVAKGMAFLASKNCIHRDLAARNILLTHGRITKICDFGLARDIKNDSNYVVKGNARLPVKWMAPESIFNCVYTFESDVWSYGIFLWELFSLGSSPYPGMPVDSKFYKMIKEGFRMLSPEHAPAEMYDIMKTCWDADPLKRPTFKQIVQLIEKQISESTNHIYSNLANCSPNRQKPVVDHSVRINSVGSTASSSQPLLVHDDV

[0165] "CD117 polynucleotide" refers to a nucleic acid molecule encoding a CD117 polypeptide. The sequences of exemplary CD117 polynucleotides are provided below:

[0166] >NM_000222.2 Homo sapiens KIT proto-oncogene, receptor tyrosine kinase (KIT), transcript

[0167] Variant 1, mRNA

[0168]

[0169] "C-X-C chemokine receptor type 4 (CXCR4) polypeptide" refers to a polypeptide or a fragment thereof that has at least about 95% amino acid sequence identity with the amino acid sequence that binds to an anti-CXCR4 antibody provided by GenBank accession number NP_001008540. The following provides the sequences of exemplary CXCR4 polypeptides:

[0170] >NP_001008540.1 C-X-C chemokine receptor type 4 isoform a [Homo sapiens]

[0171] MSIPLPLLQIYTSDNYTEEMGSGDYDSMKEPCFREENANFNKIFLPTIYSIIFLTGIVGNGLVILVMGYQKKLRSMTDKYRLHLSVADLLFVITLPFWAVDAVANWYFGNFLCKAVHVIYTVNLYSSVLILAFISLDRYLAIVHATNSQRPRKLLAEKVVYVGVWIPALLLTIPDFIFANVSEADDRYICDRFYPNDLWVVVFQFQHIMVGLILPGIVILSCYCIIISKLSHSKGHQKRKALKTTVILILAFFACWLPYYIGISIDSFILLEIIKQGCEFENTVHKWISITEALAFFHCCLNPILYAFLGAKFKTSAQHALTSVSRGSSLKILSKGKRGGHSSVSTESESSSFHSS

[0172] "CXCR4 polynucleotide" refers to a nucleic acid molecule encoding a CXCR4 polypeptide. The following provides the sequences of exemplary CXCR4 polynucleotides:

[0173] >NM_003467.2 Homo sapiens C-X-C motif chemokine receptor 4 (CXCR4), transcript variant 2, mRNA

[0174]

[0175] "CD135 polypeptide" refers to a polypeptide or a fragment thereof having at least about 95% amino acid sequence identity with the amino acid sequence of the anti-CD135 antibody-binding provided by GenBank accession number NP_004110. The sequences of exemplary CD135 polypeptides are provided below:

[0176] >NP_004110.2 receptor-type tyrosine-protein kinase FLT3 precursor [Homo sapiens]

[0177] MPALARDGGQLPLLVVFSAMIFGTITNQDLPVIKCVLINHKNNDSSVGKSSSYPMVSESPEDLGCALRPQSSGTVYEAAAVEVDVSASITLQVLVDAPGNISCLWVFKHSSLNCQPHFDLQNRGVVSMVILKMTETQAGEYLLFIQSEATNYTILFTVSIRNTLLYTLRRPYFRKMENQDALVCISESVPEPIVEWVLCDSQGESCKEESPAVVKKEEKVLHELFGTDIRCCARNELGRECTRLFTIDLNQTPQTTLPQLFLKVGEPLWIRCKAVHVNHGFGLTWELENKALEEGNYFEMSTYSTNRTMIRILFAFVSSVARNDTGYYTCSSSKHPSQSALVTIVEKGFINATNSSEDYEIDQYEEFCFSVRFKAYPQIRCTWTFSRKSFPCEQKGLDNGYSISKFCNHKHQPGEYIFHAENDDAQFTKMFTLNIRRKPQVLAEASASQASCFSDGYPLPSWTWKKCSDKSPNCTEEITEGVWNRKANRKVFGQWVSSSTLNMSEAIKGFLVKCCAYNSLGTSCETILLNSPGPFPFIQDNISFYATIGVCLLFIVVLTLLICHKYKKQFRYESQLQMVQVTGSSDNEYFYVDFREYEYDLKWEFPRENLEFGKVLGSGAFGKVMNATAYGISKTGVSIQVAVKMLKEKADSSEREALMSELKMMTQLGSHENIVNLLGACTLSGPIYLIFEYCCYGDLLNYLRSKREKFHRTWTEIFKEHNFSFYPTFQSHPNSSMPGSREVQIHPDSDQISGLHGNSFHSEDEIEYENQKRLEEEEDLNVLTFEDLLCFAYQVAKGMEFLEFKSCVHRDLAARNVLVTHGKVVKICDFGLARDIMSDSNYVVRGNARLPVKWMAPESLFEGIYTIKSDVWSYGILLWEIFSLGVNPYPGIPVDANFYKLIQNGFKMDQPFYATEEIYIIMQSCWAFDSRKRPSFPNLTSFLGCQLADAEEAMYQNVDGRVSECPHTYQNRRPFSREMDLGLLSPQAQVEDS

[0178] "CD135 polynucleotide" refers to a nucleic acid molecule encoding a CD135 polypeptide. The sequences of exemplary CD135 polynucleotides are provided below:

[0179]

[0180] "CD90 polypeptide" refers to a polypeptide or a fragment thereof having at least about 95% amino acid sequence identity with the amino acid sequence that binds to an anti-CD90 antibody provided by GenBank accession number NP_001298089. The sequences of exemplary CD90 polypeptides are provided below:

[0181] >NP_001298089.1 thy-1 membrane glycoprotein isoform 1 preproprotein [Homo sapiens]

[0182] MNLAISIALLLTVLQVSRGQKVTSLTACLVDQSLRLDCRHENTSSSPIQYEFSLTRETKKHVLFGTVGVPEHTYRSRTNFTSKYNMKVLYLSAFTSKDEGTYTCALHHSGHSPPISSQNVTVLRDKLVKCEGISLLAQNTSWLLLLLLSLSLLQATDFMSL

[0183] "CD90 polynucleotide" refers to a nucleic acid molecule encoding a CD90 polypeptide. The sequences of exemplary CD90 polynucleotides are provided below:

[0184] >NM_006288.5 Homo sapiens Thy-1 cell surface antigen (THY1), transcript variant 1, mRNA

[0185]

[0186] "CD45 polypeptide" refers to a polypeptide or a fragment thereof having at least about 95% amino acid sequence identity to the amino acid sequence that binds an anti-CD45 antibody provided by GenBank accession number NP_001254727. The sequences of exemplary CD45 polypeptides are provided below:

[0187] >NP_001254727.1 Receptor-type tyrosine-protein phosphatase C isoform 5 precursor [Homo sapiens]

[0188] MTMYLWLKLLAFGFAFLDTEVFVTGQSPTPSPTGHLQAEEQGSQSKSPNLKSREADSSAFSWWPKAREPL TNHWSKSKSPKAEELGV

[0189] "CD45 polynucleotide" refers to a nucleic acid molecule encoding a CD45 polypeptide. The sequences of exemplary CD45 polynucleotides are provided below:

[0190] >NM_001267798.2 Homo sapiens protein tyrosine phosphatase receptor type C (PTPRC), transcript variant 5, mRNA

[0191]

[0192] "CD34 polypeptide" refers to a polypeptide or a fragment thereof having at least about 95% amino acid sequence identity with the amino acid sequence that binds to an anti-CD34 antibody provided by GenBank accession number NP_001020280. The following provides the sequences of exemplary CD34 polypeptides:

[0193] >NP_001020280.1 Hematopoietic progenitor cell antigen CD34 isoform precursor [Homo sapiens]

[0194] MLVRRGARAGPRMPRGWTALCLLSLLPSGFMSLDNNGTATPELPTQGTFSNVSTNVSYQETTTPSTLGSTSLHPVSQHGNEATTNITETTVKFTSTSVITSVYGNTNSSVQSQTSVISTVFTTPANVSTPETTLKPSLSPGNVSDLSTTSTSLATSPTKPYTSSSPILSDIKAEIKCSGIREVKLTQGICLEQNKTSSCAEFKKDRGEGLARVLCGEEQADADAGAQVCSLLLAQSEVRPQCLLLVLANRTEISSKLQLMKKHQSDLKKLGILDFTEQDVASHQSYSQKTLIALVTSGALLAVLGITGYFLMNRRSWSPTGERLGEDPYYTENGGGQGYSSGPGTSPEAQGKASVNRGAQENGTGQATSRNGHSARQHVVADTEL

[0195] "CD34 polynucleotide" refers to a nucleic acid molecule encoding a CD34 polypeptide. The following provides the sequences of exemplary CD34 polynucleotides:

[0196] >NM_001025109.2 Homo sapiens CD34 molecule (CD34), transcript variant 1, mRNA

[0197]

[0198] "Stem cell factor (SCF) polypeptide" refers to a polypeptide or a fragment thereof having at least about 95% amino acid sequence identity with the hematopoietic amino acid sequence provided by GenBank accession number NP_000890. In some embodiments, the SCF polypeptide or its fragment binds to CD117. The sequences of exemplary SCF polypeptides are provided below:

[0199] >NP_000890.1 Kit ligand isoform b precursor [Homo sapiens]

[0200] MKKTQTWILTCIYLQLLLFNPLVKTEGICRNRVTNNVKDVTKLVANLPKDYMITLKYVPGMDVLPSHCWISEMVVQLSDSLTDLLDKFSNISEGLSNYSIIDKLVNIVDDLVECVKENSSKDLKKSFKSPEPRLFTPEEFFRIFNRSIDAFKDFVVASETSDCVVSSTLSPEKDSRVSVTKPFMLPPVAASSLRNDSSSSNRKAKNPPGDSSLHWAAMALPALFSLIIGFAFGALYWKKRQPSLTRAVENIQINEEDNEISMLQEKEREFQEV

[0201] "SCF polynucleotide" refers to a nucleic acid molecule encoding an SCF polypeptide. The sequences of exemplary SCF polynucleotides are provided below:

[0202] >NM_003994.5 Homo sapiens KIT ligand (KITLG), transcript variant a, mRNA

[0203]

[0204] As used herein, the term "linker" can refer to a covalent linker (e.g., a covalent bond), a non-covalent linker, a chemical group, or a molecule that connects two molecules or moieties (e.g., two components of a protein complex or ribonucleoprotein complex), or two domains of a fusion protein, such as a polynucleotide-programmable DNA-binding domain (e.g., dCas9) and a deaminase domain (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase). The linker can connect different components or different parts of a base editor system. For example, in some embodiments, the linker can connect the guide polynucleotide-binding domain of a polynucleotide-programmable nucleotide-binding domain and the catalytic domain of a deaminase. In some embodiments, the linker can connect a CRISPR polypeptide and a deaminase. In some embodiments, the linker can connect Cas9 and a deaminase. In some embodiments, the linker can connect dCas9 and a deaminase. In some embodiments, the linker can connect nCas9 and a deaminase. In some embodiments, the linker can connect a guide polynucleotide and a deaminase. In some embodiments, the linker can connect a deaminating component and the polynucleotide-programmable nucleotide-binding component of a base editor system. In some embodiments, the linker can connect the RNA-binding portion of a deaminating component and the polynucleotide-programmable nucleotide-binding component of a base editor system. In some embodiments, the linker can connect the RNA-binding portion of a deaminating component and the RNA-binding portion of the polynucleotide-programmable nucleotide-binding component of a base editor system. The linker can be located between or on both sides of two groups, molecules, or other moieties and is connected to each by a covalent bond or non-covalent interaction, thereby connecting the two. In some embodiments, the linker can be an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker can be a polynucleotide. In some embodiments, the linker can be a DNA linker. In some embodiments, the linker can be an RNA linker. In some embodiments, the linker can comprise an aptamer capable of binding a ligand. In some embodiments, the ligand can be a carbohydrate, peptide, protein, or nucleic acid. In some embodiments, the linker can comprise an aptamer derivable from a riboswitch. The riboswitch from which the aptamer is derived can be selected from the theophylline riboswitch, thiamine pyrophosphate (TPP) riboswitch, adenosylcobalamin (AdoCbl) riboswitch, S-adenosylmethionine (SAM) riboswitch, SAH riboswitch, flavin mononucleotide (FMN) riboswitch, tetrahydrofolate riboswitch, lysine riboswitch, glycine riboswitch, purine riboswitch, GlmS riboswitch, or pre-queosine1 (PreQ1) riboswitch. In some embodiments, the linker can comprise an aptamer that binds to a polypeptide or protein domain, such as a polypeptide ligand.In some embodiments, the polypeptide ligand can be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif. In some embodiments, the polypeptide ligand can be part of a base editor system component. For example, a nucleobase editing component can include a deaminase domain and an RNA recognition motif.

[0205] In some embodiments, the linker can be one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the length of the linker can be about 5 - 100 amino acids, such as about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20 - 30, 30 - 40, 40 - 50, 50 - 60, 60 - 70, 70 - 80, 80 - 90, or 90 - 100 amino acids in length. In some embodiments, the length of the linker can be about 100 - 150, 150 - 200, 200 - 250, 250 - 300, 300 - 350, 350 - 400, 400 - 450, or 450 - 500 amino acids. Longer or shorter linkers can also be considered.

[0206] In some embodiments, the linker connects the gRNA-binding domain of an RNA-programmable nuclease, including the Cas9 nuclease domain and the catalytic domain of a nucleic acid editing protein (e.g., a cytidine or adenosine deaminase). In some embodiments, the linker connects dCas9 and a nucleic acid editing protein. For example, the linker is located between or on both sides of two groups, molecules, or other moieties and is covalently linked to each, thereby connecting the two. In some embodiments, the linker can be one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker is one amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the length of the linker can be about 5 - 200 amino acids, such as 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140, 150, 160, 175, 180, 190, or more amino acids in length.

[0207] In some embodiments, the domains of the base editor are fused via a linker comprising the amino acid sequence: SGGSSGSETPGTSESATPESSGGS, SGGSSGGSSGSETPGTSESATPESSGGSSGGS, or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS. In some embodiments, the domains of the base editor are fused via a linker comprising the amino acid sequence SGSETPGTSESATPES, which may also be referred to as the XTEN linker. In some embodiments, the length of the linker is 24 amino acids. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES. In some embodiments, the length of the linker is 40 amino acids. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS. In some embodiments, the length of the linker is 64 amino acids. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS. In some embodiments, the length of the linker is 92 amino acids. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS.

[0208] "Makassar" or "Hb G-Makassar" refers to human β-hemoglobin variant, G-Makassar variant or mutant (HB Makassar variant) of human hemoglobin (Hb), which is an asymptomatic, naturally occurring variant (E6A) hemoglobin. Hb G-Makassar was first discovered in Indonesia. (Mohamad, A.S. et al., 2018, Hematol. Rep., 10(3):7210 (doi:10.4081 / hr.2018.7210). When electrophoresis is performed, Hb G-Makassar has a slower migration rate. The Makassar β-hemoglobin variant has an anatomical abnormality at the β-6 or A3 position, where the glutamyl residue is usually replaced by an alanyl residue. Substituting a single amino acid in the gene encoding the β-6 glutamyl of the β-globin subunit with valine will result in sickle cell disease. Conventional procedures, such as isoelectric focusing, hemoglobin electrophoresis separation by cation exchange high performance liquid chromatography (HPLC), and cellulose acetate electrophoresis, cannot separate Hb G-Makassar and HbS globin forms because they are found to have the same properties when analyzed by these methods. Therefore, Hb G-Makassar and HbS have been misidentified and misinterpreted by those skilled in the art, leading to misdiagnosis of sickle cell disease (SCD). In one embodiment, the valine at amino acid position 6 that causes sickle cell disease is replaced with alanine, resulting in an Hb variant (HbMakassar) that does not produce a sickle cell phenotype. In some embodiments, the Val Ala (GTG GCG) substitution (i.e., Hb Makassar variant) can occur using an A·T to G·C base editor (ABE).

[0209] Accordingly, the present invention includes compositions and methods for base editing the thymidine (T) base in the codon of the sixth amino acid of the sickle cell variant (Sickle HbS; E6V) of β-globin to cytidine (C), thereby replacing valine with alanine (V6A) at said amino acid position. Substituting valine with alanine at position 6 of HbS produces a β-globin variant that does not have a sickle cell phenotype (e.g., does not have the potential to polymerize like the pathogenic variant HbS). Accordingly, the compositions and methods of the present invention can be used to treat sickle cell disease (SCD).

[0210] "Marker" refers to any protein or polynucleotide that has an altered expression level or activity in relation to a disease or disorder.

[0211] As used herein, the term "mutation" refers to the replacement of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) with another residue, or the deletion or insertion of one or more residues within the sequence. Mutations are typically described herein by identifying the position of the residue within the sequence that follows the original residue and by the identity of the newly substituted residue. The various methods for making amino acid substitutions (mutations) provided herein are well known in the art and are provided, for example, by Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)).

[0212] In some embodiments, the presently disclosed base editors can effectively generate "desired mutations", such as point mutations, in a nucleic acid (e.g., a nucleic acid within a subject's genome) without generating a large number of undesired mutations, such as unexpected point mutations. In some embodiments, the desired mutations are mutations generated by a particular base editor (e.g., a cytidine base editor or an adenosine base editor) that binds to a guide polynucleotide (e.g., a gRNA), which is specifically designed to generate the desired mutations.

[0213] Typically, mutations generated or identified in a sequence (e.g., an amino acid sequence as described herein) are numbered relative to a reference (or wild-type) sequence, i.e., a sequence that does not contain the mutation. Those of skill in the art will readily understand how to determine the position of mutations in amino acid and nucleic acid sequences relative to the reference sequence.

[0214] The term "non-conservative mutation" refers to an amino acid substitution between different groups, e.g., tryptophan for lysine, or serine for phenylalanine, etc. In such cases, the non-conservative amino acid substitution preferably does not interfere with, or inhibit, the biological activity of the functional variant. The non-conservative amino acid substitution may enhance the biological activity of the functional variant, such that the biological activity of the functional variant is increased compared to the wild-type protein.

[0215] The terms "nuclear localization sequence", "nuclear localization signal", or "NLS" refer to an amino acid sequence that facilitates the import of a protein into the cell nucleus. Nuclear localization sequences are known in the art and are described, for example, in the international PCT application by Plank et al., PCT / EP2000 / 011690, filed Nov. 23, 2000 and published as WO / 2001 / 038547 on May 31, 2001, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In other embodiments, the NLS is an optimized NLS, such as that described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRK, PKKKRKV, or MDSLLMNRRKFLYQFKNVRWAKGRRETY.

[0216] The terms "nucleobase", "nitrogenous base", or "base" are used interchangeably herein and refer to a nitrogen-containing biological compound that forms part of a nucleoside, which in turn is a component of a nucleotide. The ability of nucleobases to form base pairs and stack on top of each other directly results in long helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases - adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U) - are referred to as primary or canonical. Adenine and guanine are derived from purine, while cytosine, uracil, and thymine are derived from pyrimidine. DNA and RNA can also contain other (non-primary) modified bases. Non-limiting exemplary modified nucleobases can include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Hypoxanthine and xanthine can be produced in the presence of a mutagen, and they are both produced by deamination (replacing an amine group with a carbonyl group). Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can be produced by deamination of cytosine. A "nucleoside" consists of a nucleobase and a five-carbon sugar (ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides with modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five-carbon sugar (ribose or deoxyribose), and at least one phosphate group.

[0217] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to compounds that contain nucleobases and an acidic moiety, such as nucleosides, nucleotides, or polymers of nucleotides. Generally, polymeric nucleic acids, such as nucleic acid molecules that contain three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other by phosphodiester bonds. In some embodiments, "nucleic acid" refers to a single nucleic acid residue (e.g., nucleotide and / or nucleoside). In some embodiments, "nucleic acid" refers to an oligonucleotide chain that contains three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" are used interchangeably to refer to a polymer of nucleotides (e.g., a string of at least three nucleotides). In some embodiments, "nucleic acid" includes RNA as well as single-stranded and / or double-stranded DNA. Nucleic acids can be naturally occurring, such as in the context of a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, nucleic acid molecules can be non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or those that include non-naturally occurring nucleotides or nucleosides. In addition, the terms "nucleic acid", "DNA", "RNA", and / or similar terms include nucleic acid analogs, such as analogs having a backbone other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems and optionally purified, chemically synthesized, etc. In appropriate cases, such as in the case of chemically synthesized molecules, nucleic acids can contain nucleoside analogs, such as analogs having chemically modified bases or sugar and backbone modifications. Unless otherwise indicated, nucleic acid sequences are presented in the 5' to 3' direction. In some embodiments, the nucleic acid is or contains natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine); nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine); chemically modified bases; biologically modified bases (e.g., methylated bases); inserted bases; modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose); and / or modified phosphate groups (e.g., phosphorothioates and 5'-N-phosphoramidite linkages).

[0218] The term "nucleic acid programmable DNA-binding protein" or "napDNAbp" can be used interchangeably with "polynucleotide programmable nucleotide-binding domain" to refer to a protein associated with a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA), that directs the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide programmable acid-binding domain is a polynucleotide programmable DNA-binding domain. In some embodiments, the polynucleotide programmable acid-binding domain is a polynucleotide programmable RNA-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a Cas9 protein. The Cas9 protein can be associated with a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, such as nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ.Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, their homologs or their modified or engineered forms. Other nucleic acid programmable DNA-binding proteins are also within the scope of the present disclosure, although they may not be specifically listed in the present disclosure. See, e.g., Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Wherefrom Here?" CRISPR J. 2018 Oct;1:325-336. doi:10.1089 / crispr.2018.0033; Yan et al., "Functionally diverse type V CRISPR-Cas systems" Science. 2019 Jan 4;363(6422):88-91. doi:10.1126 / science.aav7271, the entire contents of which are incorporated herein by reference.

[0219] As used herein, the term "nucleobase editing domain" or "nucleobase editing protein" refers to a protein or enzyme that can catalyze the modification of nucleobases in RNA or DNA, such as cytosine

[0220] (Or cytidine) is the deamination of uracil (or uridine) or thymine (or thymidine) and adenine (or adenosine) to hypoxanthine (or inosine), as well as non-template nucleotide addition and insertion. In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., an adenine deaminase or adenosine deaminase; or a cytidine deaminase or cytosine deaminase). In some embodiments, the nucleobase editing domain is multiple deaminase domains (e.g., an adenine deaminase or adenosine deaminase and a cytidine or cytosine deaminase). In some embodiments, the nucleobase editing domain can be a naturally occurring nucleobase editing domain. In some embodiments, the nucleobase editing domain can be a modified or evolved nucleobase editing domain from a naturally occurring nucleobase editing domain. The nucleobase editing domain can be from any organism, such as bacteria, human, chimpanzee, gorilla, monkey, cow, dog, rat, or mouse.

[0221] As used herein, "obtaining" as in "obtaining an agent" includes synthesizing, purchasing, or otherwise obtaining the agent.

[0222] As used herein, "patient" or "subject" refers to a mammalian subject or individual diagnosed with, at risk of developing, or suspected of having or developing a disease or disorder. In some embodiments, the term "patient" refers to a mammalian subject with a higher than average likelihood of developing a disease or disorder. Exemplary patients can be human, non-human primate, cat, dog, pig, cow, cat, horse, camel, llama, goat, sheep, rodent (e.g., mouse, rabbit, rat, or guinea pig), and other mammals that can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female.

[0223] "A patient in need" or "a subject in need" as used herein refers to a patient diagnosed with, at risk of, or having, or suspected of having a disease or disorder.

[0224] The terms "pathogenic mutation", "pathogenic variant", "disease-causing mutation", "pathogenic variant", "harmful mutation", or "susceptibility mutation" refer to a genetic alteration or mutation that increases an individual's susceptibility or predisposition to a particular disease or disorder. In some embodiments, a pathogenic mutation comprises the replacement of at least one wild-type amino acid with at least one pathogenic amino acid in a protein encoded by a gene.

[0225] The terms "protein", "peptide", "polypeptide" and their grammatical equivalents are used interchangeably herein and refer to polymers of amino acid residues joined together by peptide (amide) bonds. These terms refer to proteins, peptides or polypeptides of any size, structure or function. Generally, a protein, peptide or polypeptide is at least three amino acids in length. A protein, peptide or polypeptide can refer to a single protein or a collection of proteins. One or more of the amino acids in a protein, peptide or polypeptide can be modified, e.g., by the addition of chemical entities such as carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl, geranylgeranyl, fatty acid groups, linkers for conjugation, functionalization or other modification, etc. A protein, peptide or polypeptide can also be a single molecule or can be a multimolecular complex. A protein, peptide or polypeptide can be just a fragment of a naturally occurring protein or peptide. A protein, peptide or polypeptide can be naturally occurring, recombinant or synthetic, or any combination thereof. As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains from at least two different proteins. One protein can be located in the amino-terminal (N-terminal) portion or the carboxy-terminal (C-terminal) portion of the fusion protein, thereby forming an amino-terminal fusion protein or a carboxy-terminal fusion protein, respectively. A protein can comprise different domains, e.g., a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the protein to bind to a target site) and a nucleic acid cleavage domain, or the catalytic domain of a nucleic acid editing protein. In some embodiments, a protein comprises a protein portion, e.g., an amino acid sequence that constitutes a nucleic acid binding domain, and an organic compound, e.g., a compound that can act as a nucleic acid cleavage agent. In some embodiments, a protein forms a complex or associates with a nucleic acid such as RNA or DNA. Any protein provided herein can be produced by any method known in the art. For example, the proteins provided herein can be produced by recombinant protein expression and purification, which is particularly applicable to fusion proteins comprising peptide linkers. Methods for recombinant protein expression and purification are well known and include those described in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2012)), the entire contents of which are incorporated by reference.

[0226] The polypeptides and proteins disclosed herein (including their functional portions and functional variants) may contain synthetic amino acids in place of one or more naturally occurring amino acids. Such synthetic amino acids are known in the art and include, for example, aminocyclohexanecarboxylic acid, norleucine, α-aminodecanoic acid, homoserine, S-acetamidomethyl-cysteine, trans-3- and trans-4-hydroxyproline, 4-aminophenylalanine, 4-nitrophenylalanine, 4-chlorophenylalanine, 4-carboxyphenylalanine, β-phenylserine β-hydroxy phenylalanine, phenylglycine, α-naphthylalanine, cyclohexylalanine, cyclohexylglycine, indoline-2-carboxylic acid, 3-quinoline carboxylic acid, 3-hydroxy phenylalanine, 3,4-dihydroxy phenylalanine, aminomalonic acid, aminomalonic acid monoamide, N'-benzyl-N'-methyl-lysine, N',N'-dibenzyl-lysine, 6-hydroxylysine, ornithine, α-aminocyclopentanecarboxylic acid, α-aminocyclohexanecarboxylic acid, α-aminocycloheptanecarboxylic acid, α-(2-amino-2-norbornane)-carboxylic acid, α,γ-diaminobutyric acid, α,β-diaminopropionic acid, norphenylalanine, and α-tert-butylglycine. The polypeptides and proteins may be associated with post-translational modifications of one or more amino acids of the polypeptide construct. Non-limiting examples of post-translational modifications include phosphorylation, acylation (including acetylation and formylation), glycosylation (including N-linked and O-linked), amidation, hydroxylation, alkylation (including methylation and ethylization), ubiquitination, addition of pyrrolidone carboxylic acid, formation of disulfide bonds, sulfation, myristoylation, palmitoylation, isoprenylation, farnesylation, geranylgeranylation, glycosylation, lipoylation, and iodination.

[0227] As used herein, the term "recombinant" in the context of a protein or nucleic acid refers to a protein or nucleic acid that does not exist in nature but is a product of human engineering. For example, in some embodiments, a recombinant protein or nucleic acid molecule contains an amino acid or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.

[0228] "Reduce" means a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0229] "Reference" means a standard or control condition. In one embodiment, the reference is a wild-type or healthy cell. In other embodiments and without limitation, the reference is an untreated cell that has not been subjected to the test condition, or has been subjected to a placebo or saline, medium, buffer, and / or a control vector that does not contain the target polynucleotide.

[0230] "Reference sequence" is a defined sequence used as a basis for sequence comparison. The reference sequence can be a subset or the entirety of a particular sequence; for example, a fragment of a full-length cDNA or gene sequence, or a full-length cDNA or gene sequence. For polypeptides, the length of the reference polypeptide sequence is typically at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, about 35 amino acids, about 50 amino acids, or about 100 amino acids. For nucleic acids, the length of the reference nucleic acid sequence is typically at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, about 100 nucleotides, or about 300 nucleotides, or any integer near or between them. In some embodiments, the reference sequence is the wild-type sequence of the protein of interest. In other embodiments, the reference sequence is the polynucleotide sequence encoding the wild-type protein.

[0231] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" are used (e.g., in combination or association) with one or more RNAs that are not the target for cleavage. In some embodiments, an RNA-programmable nuclease may be referred to as a nuclease:RNA complex when complexed with RNA. Generally, the bound RNA is referred to as guide RNA (gRNA). In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, such as Cas9 from Streptococcus pyogenes (Csnl) (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C, Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607 (2011).

[0232] As used herein, the term "scFv" refers to a single-chain Fv antibody in which the variable domains from the heavy and light chains of an antibody have been joined to form a single chain. The scFv fragment comprises a polypeptide chain that includes the variable region of the antibody light chain (VL) (e.g., CDR-L1, CDR-L2, and / or CDR-L3) and the variable region of the antibody heavy chain (VH) (e.g., CDR-H1, CDR-H2, and / or CDR-H3) separated by a linker. The linker that joins the VL and VH regions of the scFv fragment can be a peptide linker composed of protein amino acids. Alternative linkers can be used to increase the resistance of the scFv fragment to proteolytic degradation (e.g., linkers containing D-amino acids), to enhance the solubility of the scFv fragment (e.g., hydrophilic linkers such as polyethylene glycol-containing linkers or polypeptides containing repeating glycine and serine residues), to improve the biophysical stability of the molecule (e.g., linkers containing cysteine residues that form intramolecular or intermolecular disulfide bonds), or to attenuate the immunogenicity of the scFv fragment (e.g., linkers containing glycosylation sites). One of ordinary skill in the art will also understand that the variable regions of the scFv molecules described herein can be modified such that their amino acid sequences are different from the antibody molecules from which they are derived. For example, nucleotide or amino acid substitutions can be made that result in conservative substitutions or changes in amino acid residues (e.g., in CDR and / or framework residues) to maintain or enhance the ability of the scFv to bind to the antigen recognized by the corresponding antibody.

[0233] "Selectively binds" means specifically binds to the wild-type form of a cell surface protein but exhibits reduced binding or fails to bind to a cell surface protein that contains a mutation.

[0234] The term "single nucleotide polymorphism (SNP)" refers to a variation of a single nucleotide at a specific position in the genome, where each variation exists at a certain frequency (e.g., > 1%) in the population. For example, at a specific base position in the human genome, the C nucleotide may be present in most individuals, but in a few individuals, that position is occupied by an A. This means that there is an SNP at that specific position, and the two possible nucleotide variations, C or A, are referred to as the alleles at that position. SNPs are the basis for differences in disease susceptibility. The severity of diseases and the way our bodies respond to treatment are also manifestations of genetic variations. SNPs can fall within the coding region of a gene, the non-coding region of a gene, or the intergenic region (the region between genes). In some embodiments, due to the degeneracy of the genetic code, SNPs within the coding sequence do not necessarily change the amino acid sequence of the resulting protein. There are two types of SNPs in the coding region: synonymous SNPs and non-synonymous SNPs. Synonymous SNPs do not affect the protein sequence, and non-synonymous SNPs change the amino acid sequence of the protein. The non-synonymous SNPs have two types: missense and nonsense. SNPs that are not in the protein-coding region can still affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNAs. Gene expression affected by such SNPs is called eSNP (expressed SNP), which can be located upstream or downstream of the gene. A single nucleotide variant (SNV) is a variation of a single nucleotide without any frequency limitation and can occur in somatic cells. Somatic single nucleotide variants can also be referred to as single nucleotide alterations.

[0235] "Specifically binds" means a nucleic acid molecule, polypeptide, or complex thereof (e.g., a nucleic acid programmable DNA binding domain and a guide nucleic acid), compound, or molecule that recognizes and binds to the polypeptide and / or nucleic acid molecule of the present invention, but substantially does not recognize and bind to other molecules in the sample, such as a biological sample.

[0236] Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but will generally exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence is generally capable of hybridizing to at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present invention include any nucleic acid molecule encoding a polypeptide of the present invention or a fragment thereof. Such nucleic acid molecules need not be 100% identical to the endogenous nucleic acid sequence, but will generally exhibit substantial identity. A polynucleotide having "substantial identity" to an endogenous sequence is generally capable of hybridizing to at least one strand of a double-stranded nucleic acid molecule.

[0237] "Hybridization" refers to the pairing between complementary polynucleotide sequences (e.g., genes as described herein) or portions thereof to form double-stranded molecules under various stringent conditions. (See, e.g., Wahl, G.M. and S.L.Berger (1987) Methods Enzymol. 152:399; Kimmel, A.R. (1987) Methods Enzymol. 152:507).

[0238] For example, stringent salt concentrations are typically less than about 750 mM NaCl and 75 mM trisodium citrate, preferably less than about 500 mM NaCl and 50 mM trisodium citrate, and more preferably less than about 250 mM NaCl and 25 mM trisodium citrate. Low stringency hybridization can be achieved in the absence of organic solvents such as formamide, while high stringency hybridization can be achieved in the presence of at least about 35% formamide, more preferably at least about 50% formamide. Stringent temperature conditions generally include a temperature of at least about 30°C, more preferably at least about 37°C, and most preferably at least about 42°C. For example, sodium dodecyl sulfate (SDS), and the inclusion or exclusion of carrier DNA, are well known to those skilled in the art. Different degrees of stringency are achieved by combining these different conditions as needed. In a preferred embodiment, hybridization will occur at 30°C in 750 mM NaCl, 75 mM trisodium citrate, and 1% SDS. In a more preferred embodiment, hybridization will occur at 37°C in 500 mM NaCl, 50 mM trisodium citrate, 1% SDS, 35% formamide, and 100 μg / ml denatured salmon sperm DNA (ssDNA). In a most preferred embodiment, hybridization will occur at 42°C in 250 mM NaCl, 25 mM trisodium citrate, 1% SDS, 50% formamide, and 200 μg / ml ssDNA. Useful variations of these conditions will be apparent to those skilled in the art.

[0239] For most applications, the washing steps after hybridization also vary in terms of stringency. Washing stringency conditions can be defined by salt concentration and temperature. As described above, washing stringency can be increased by decreasing the salt concentration or increasing the temperature. For example, the stringent salt concentration for the washing step is preferably less than about 30 mM NaCl and 3 mM trisodium citrate, and most preferably less than about 15 mM NaCl and 1.5 mM trisodium citrate. The stringent temperature conditions for the washing step generally include a temperature of at least about 25°C, more preferably at least about 42°C, and even more preferably at least about 68°C. In one embodiment, the washing step will occur at 25°C in 30 mM NaCl, 3 mM trisodium citrate, and 0.1% SDS. In another embodiment, the washing step will be carried out at 42°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. In a more preferred embodiment, the washing step will be carried out at 68°C in 15 mM NaCl, 1.5 mM trisodium citrate, and 0.1% SDS. Other variations of these conditions will be apparent to those skilled in the art. Hybridization techniques are well known to those skilled in the art and are described, for example, in Benton and Davis (Science 196:180, 1977); Grunstein and Hogness (Proc. Natl. Acad. Sci., USA 72:3961, 1975); Ausubel et al. (Current Protocols in Molecular Biology, Wiley Interscience, New York, 2001); Berger and Kimmel (Guide to Molecular Cloning Techniques, 1987, Academic Press, New York); and Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, New York.

[0240] "Fragmentation" means breaking into two or more pieces.

[0241] "Split Cas9 protein" or "split Cas9" refers to a Cas9 protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal portion and the C-terminal portion of the Cas9 protein can be spliced to form a "reconstituted" Cas9 protein. In certain embodiments, the Cas9 protein is split into two fragments within a disordered region of the protein, e.g., as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as in Jiang et al. (2016) Science 351:867-871. PDB file: 5F9R, each of which is incorporated herein by reference. In some embodiments, the protein is split into two fragments at any C, T, A, or S between approximately amino acids A292-G364, F445-K483, or E565-T637 within the SpCas9 region, or in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp. In some embodiments, the protein is split into two fragments at SpCas9 T310, T313, A456, S469, or C574. In some embodiments, the process of splitting the protein into two fragments is referred to as "splitting" the protein.

[0242] In other embodiments, the N-terminal portion of the Cas9 protein comprises amino acids 1-573 or 1-637 of Streptococcus pyogenes Cas9 wild type (SpCas9) (NCBI reference sequence: NC_002737.2, Uniprot Reference Sequence: Q99ZW2) and the C-terminal portion of the Cas9 protein comprises the portion of amino acids 574-1368 or 638-1368 of SpCas9 wild type.

[0243] The C-terminal portion of split Cas9 can be joined to the N-terminal portion of split Cas9 to form a complete Cas9 protein. In some embodiments, the C-terminal portion of the Cas9 protein begins where the N-terminal portion of the Cas9 protein ends. Thus, in some embodiments, the C-terminal portion of split Cas9 comprises a portion of amino acids (551-651)-1368 of spCas9. "(551-651)-1368" refers to starting from the amino acids between amino acids 551-651 (inclusive) to the end at amino acid 1368.For example, the C-terminal portion of split Cas9 can include a portion of any amino acid of SpCas9: 551-1368, 552-1368, 553-1368, 554-1368, 555-1368, 556-1368, 557-1368, 558-1368, 559-1368, 560-1368, 561-1368, 562-1368, 563-1368, 564-1368, 565-1368, 566-1368, 567-1368, 568-1368, 569-1368, 570-1368, 571-1368, 572-1368, 573-1368, 574-1368, 575-1368, 576-1368, 577-1368, 578-1368, 579-1368, 580-1368, 581-1368, 582-1368, 583-1368, 584-1368, 585-1368, 586-1368, 587-1368, 588-1368, 589-1368, 590-1368, 591-1368, 592-1368, 593-1368, 594-1368, 595-1368, 596-1368, 597-1368, 598-1368, 599-1368, 600-1368, 601-1368, 602-1368, 603-1368, 604-1368, 605-1368, 606-1368, 607-1368, 608-1368, 609-1368, 610-1368, 611-1368, 612-1368, 613-1368, 614-1368, 615-1368, 616-1368, 617-1368, 618-1368, 619-1368, 620-1368, 621-1368, 622-1368, 623-1368, 624-1368, 625-1368, 626-1368, 627-1368, 628-1368, 629-1368, 630-1368, 631-1368, 632-1368, 633-1368, 634-1368, 635-1368, 636-1368, 637-1368, 638-1368, 639-1368, 640-1368, 641-1368, 642-1368, 643-1368, 644-1368, 645-1368, 646-1368, 647-1368, 648-1368, 649-1368, 650-1368, or 651-1368. In some embodiments, the C-terminal portion of split Cas9 includes a portion of 574-1368 or 638-1368 of the SpCas9 protein.

[0244] "Substantially identical" refers to a polypeptide or nucleic acid molecule to a reference amino acid sequence (e.g., any of the amino acid sequences described herein) or nucleic acid sequence (e.g., any of the nucleic acid sequences described herein). In one embodiment, such a sequence is at least 60%, 80%, 85%, 90%, 95% or even 99% identical to the sequence being compared at the amino acid level or nucleic acid level.

[0245] Sequence identity is typically determined using sequence analysis software (e.g., the Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine. In an exemplary method for determining the degree of identity, the BLAST program can be used, where a probability score between e-3 and e-100 indicates closely related sequences.

[0246] For example, COBALT is used with the following parameters:

[0247] a) alignment parameters: Gap penalties -11, -1 and End-Gap penalties -5, -1,

[0248] b) CDD Parameters: Use RPS BLAST on; Blast E-value 0.003; Find Conserved columns and Recompute on, and

[0249] c) Query Clustering Parameters: Use query clusters on; Word Size 4; Max cluster distance 0.8; Alphabet Regular.

[0250] For example, EMBOSS is used with the following parameters:

[0251] a) Matrix: BLOSUM62;

[0252] b) GAP OPEN: 10;

[0253] c) GAP EXTEND: 0.5;

[0254] d) OUTPUT FORMAT: pair;

[0255] e) END GAP PENALTY: false;

[0256] f) END GAP OPEN: 10; and

[0257] g) END GAP EXTEND: 0.5.

[0258] The term "target site" refers to a sequence within a nucleic acid molecule that is deaminated by a deaminase (such as cytidine or adenine deaminase) or a fusion protein containing a deaminase (such as the dCas9-adenosine deaminase fusion protein or base editor disclosed herein).

[0259] Because RNA-programmable nucleases (such as Cas9) use RNA:DNA hybridization to target DNA cleavage sites, these proteins are in principle capable of targeting any sequence specified by the guide RNA. Methods for site-specific cleavage (e.g., modifying the genome) using RNA-programmable nucleases (such as Cas9) are known in the art (see, e.g., Cong, L. et al., Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al., RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, W.Y. et al., Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al., RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); DiCarlo, J.E. et al., Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013); Jiang, W. et al., RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013); the entire contents of which are incorporated herein by reference).

[0260] As used herein, the terms "treat", "treating", "treatment", etc. refer to reducing or ameliorating a disorder and / or symptoms associated therewith or obtaining a desired pharmacological and / or physiological effect. It should be understood that, although not excluded, treating a disorder or condition does not require complete elimination of the disorder, condition or symptoms associated therewith. In some embodiments, the effect is therapeutic, i.e., but not limited to, the effect partially or completely reduces, attenuates, eliminates, alleviates, mitigates, decreases the intensity of the disease and / or adverse symptoms attributable to the disease or cures the disease and / or adverse symptoms. In some embodiments, the effect is prophylactic, i.e., the effect protects or prevents the occurrence or recurrence of a disease or condition. To this end, the presently disclosed methods include administering a therapeutically effective amount of a composition as described herein.

[0261] "Uracil glycosylase inhibitor" or "UGI" refers to an agent that inhibits the uracil excision repair system. In one embodiment, the agent is a protein or fragment thereof that binds to the host uracil-DNA glycosylase and prevents the removal of uracil residues from DNA. In one embodiment, UGI is a protein, fragment or domain thereof that is capable of inhibiting the uracil-DNA glycosylase base excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or a modified version thereof. In some embodiments, the UGI domain comprises a fragment of the exemplary amino acid sequence set forth below. In some embodiments, the amino acid sequence comprised by the UGI fragment comprises at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% or 100% identical to the exemplary UGI sequence provided below. In some embodiments, UGI comprises an amino acid sequence homologous to the exemplary UGI amino acid sequence or a fragment thereof, as described below. In some embodiments, the UGI or a portion thereof is at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.9% or at least 100% identical to wild-type UGI or a UGI sequence or a portion thereof, as described below. Exemplary UGI comprises the following amino acid sequence:

[0262] >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor

[0263] MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENV MLLTSDAPEYKPWALVIQDSNGENKIKML.

[0264] The ranges provided herein are to be understood as shorthand for all values within the range. For example, a range of 1 to 50 is to be understood as including 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0265] The recitation of a list of chemical groups in any definition of a variable herein includes defining the variable as any single group or combination of the listed groups. The recitation of embodiments of a variation or aspect herein includes the embodiment as any single embodiment or in combination with any other embodiment or portion thereof.

[0266] All terms are intended to be understood in the manner understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0267] DNA editing has emerged as a viable means of altering disease states by correcting disease-causing mutations at the gene level. Until recently, all DNA editing platforms functioned by inducing DNA double-strand breaks (DSBs) at specific genomic loci and relying on endogenous DNA repair pathways to determine product outcomes in a semi-random manner, resulting in a complex population of genetic products. While precise, user-defined repair outcomes can be achieved via the homology-directed repair (HDR) pathway, numerous challenges have hindered the efficient use of HDR for repair in therapeutically relevant cell types. In practice, the pathway is inefficient relative to the competing, error-prone non-homologous end-joining pathway. Additionally, HDR is strictly restricted to the G1 and S phases of the cell cycle, precluding precise repair of DSBs in post-mitotic cells. Thus, it has proven difficult or impossible to efficiently alter genomic sequences in these populations in a user-defined, programmable manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0268] Figure 1 is a schematic diagram depicting that the anti-CD117 antibody AMG 191 (humanized SR1) can be used for patient conditioning prior to hematopoietic stem cell transplantation.

[0269] Figure 2 is a schematic diagram depicting that the anti-CD117 antibody SR1 (AMG191) blocks the binding of stem cell factor (SCF) to CD117 in vivo and depletes human hematopoietic stem and progenitor cells (HSPCs) and myelodysplastic syndrome (MDS) cells.

[0270] Figure 3 It is a schematic diagram depicting a non-toxic conditioning strategy that enables hematopoietic stem cell transplantation (HSCT) with autogene-edited cells for patients with blood diseases. Stem cell factor (SCF) drives HSC self-renewal and differentiation into progenitor cells. Administration of anti-CD117 antibody blocks the binding of SCF to CD117, thereby depleting HSCs and progenitor cells in the patient (conditioning). Autogene-edited HSCs are transplanted into the patient. Base editing is used to generate amino acid substitutions in CD117, which can prevent the anti-CD117 antibody from binding to the edited cells but does not interfere with normal SCF signaling. The gene-edited cells compete with residual host HSCs to repopulate the bone marrow (BM). The anti-CD117 antibody blocks the binding of SCF to wild-type (WT) CD117 but cannot bind to the gene-edited CD117 HSCs. Therefore, natural wild-type HSCs are targeted by the anti-CD117 antibody, but gene-edited HSCs are not.

[0271] Figure 4 It is a schematic diagram depicting the binding of SCF to wild-type and gene-edited HSCs. Both types of cells express CD117, which is activated by SCF binding.

[0272] Figure 5 It is a schematic diagram depicting the effect of anti-CD117 antibody on wild-type and gene-edited HSCs. The binding of anti-CD117 antibody to wild-type CD117 disrupts SCF binding and leads to inhibition of SCF signal transduction in wild-type cells. In contrast, gene-edited HSCs are unresponsive to the anti-CD117 antibody because the amino acid substitutions introduced in CD117 prevent the binding of the anti-CD117 antibody but do not interfere with normal SCF binding and signaling.

[0273] Figure 6 It is a table depicting the sequence reads of base-edited cells from Makassar-WT guides (top) and hereditary persistent fetal hemoglobin elevation (HPFH) guides (bottom).

[0274] Figure 7A and 7B It is a graph depicting on-target editing ( Figure 7A ) and the insertion rate of five (5) guides (ABE8.8) ( Figure 7B ). 7B) of five (5) guides (ABE8.8). HPFH guides were used as controls.

[0275] Figure 8A and 8B It is a graph depicting on-target editing ( Figure 8A) and the insertion rate of eleven (11) guides (ABE8.8) Figure 8B )。8B) of eleven (11) guides (ABE8.8). The HPFH guide was used as a control.

[0276] Figure 9A and 9B is a graph depicting the editing on the target Figure 9A ) and the insertion rate of nine (9) guides (ABE8.8) Figure 9B )。9B) of nine (9) guides (ABE8.8). The HPFH guide was used as a control.

[0277] Figure 10 is a graph depicting the on - target editing of four (4) guides for ABE8.8 and thirteen (13) guides for IBE - NGC (Makassar).

[0278] Figures 11A to 11T is a schematic diagram depicting the positions of guide - targeted mutations for anti - CD117 antibody activity.

[0279] Figures 12A to 12N Depicts the guide that introduces a naturally occurring mutation in the CD117 target sequence. Figure 12A is a schematic diagram describing the CD117 target nucleotide sequence and the corresponding amino acid sequence for guiding the hybridization of cc - 102 and cc - 103. Figure 12B and 12C is a table describing the editing efficiency of guide cc - 102 Figure 12B ) or guide cc - 103 (Figure 12) at different base - pair positions along the CD117 target site. Figure 12C is a schematic diagram describing the CD117 target nucleotide sequence and the corresponding amino acid sequence for guiding the hybridization of cc - 78. Figure 12E is a table describing the editing efficiency of guide cc - 78 at different base - pair positions along the CD117 target site. Figure 12F is a schematic diagram describing the CD117 target sequence and the corresponding amino acid sequence for guiding the hybridization of cc - 87 and cc - 89. Figure 12G and Figure 12H is a table describing the editing efficiency of guide cc - 87 Figure 12G ) or guide cc - 89 Figure 12H ) at different base - pair positions along the CD117 target site. Figure 12I is a schematic diagram describing the CD117 target sequence and the corresponding amino acid sequence for guiding the hybridization of cc - 110. Figure 12J is a table describing the editing efficiency of guide cc - 110 at different base - pair positions along the CD117 target site. Figure 12KA schematic diagram depicting the CD117 target sequence and the corresponding amino acid sequence for guiding cc-146 hybridization. Figure 12L A table depicting the editing efficiency for guiding cc-146 at different base pair positions along the CD117 target site. Figure 12M A schematic diagram depicting the CD117 target sequence and the corresponding amino acid sequence for guiding cc-182 hybridization. Figure 12N A table depicting the editing efficiency for guiding cc-182 at different base pair positions along the CD117 target site.

[0280] Figure 13A and Figure 13B A chart depicting multiplex editing. Figure 13A A chart depicting c-KIT editing using c-KIT and HPFH dual-guide (day 3), c-KIT and HPFH dual-guide (day 5), or only c-KIT guide without HPFH guide. Figure 13B A chart depicting HPFH editing using a dual-guide of c-KIT and HPFH at 72 or 120 hours. HPFH is only used as a control. Detailed Description

[0281] The present invention provides a base editing strategy targeting cell surface proteins (e.g., one or more of CD117, CXCR4, CD135, CD90, CD45, CD34), which can be used for non-toxic conditioning.

[0282] The present invention is at least partially based on the discovery that base editing can be used to generate single base substitutions in the CD117 gene, resulting in amino acid substitutions, which in turn alter the epitope-binding domain in the encoded protein, thereby preventing anti-CD117 antibodies, antibody-drug conjugates (ADCs), or chimeric antigen receptor (CAR)-T cells from binding to CD117 on the edited cells. Advantageously, this alteration of CD117 does not change stem cell factor (SCF) binding or CD117 biological activity. Thus, the present invention provides for the use after HSCT of anti-CD117 therapies, as well as therapies targeting other antigens expressed on the surface of hematopoietic stem cells. It does not bind or deplete CD117-edited cells. Thus, gene-edited cells can expand in vivo. Advantageously, this allows for cancer treatment without hematotoxicity.

[0283] Conditioning methods in the prior art are limited to administering a conditioning regimen prior to HSC transplantation (i.e., a patient cannot be re-dosed to expand edited cells in vivo or to continue treating a malignant disease). To avoid donor influence / edited cells, the half-life of the antibody needs to be very short so that it can be cleared from the body prior to transplantation. This ensures that the antibody does not target the transplanted cells. Currently, there is no method / therapy that allows for selective targeting of a patient's HSCs. This applies to all current conditioning steps (e.g., Abs, ADCs, CAR-T cells). Thus, clinical trial design is extremely challenging for conditioning prior to HSCT and must consider the dose amount / concentration versus HSC depletion and hematopoietic failure. However, studies have shown that effective antibody-based conditioning can only be achieved with multiple doses in immunocompetent patients (as opposed to SCID trials).

[0284] Accordingly, the conditioning method of the present invention has been developed to overcome the limitations of prior methods. The conditioning method of the present invention has several advantages. The methods described herein provide selective targeting of endogenous HSCs while preserving edited HSCs. Thus, antibodies or ADC therapies can be continued after HSCT to expand gene-edited cells in vivo or to treat a malignant disease by repeated dosing. This minimizes the risk of killing edited cells. Edited cells will allow for the administration of antibodies or ADCs without Fc modification to reduce their half-life. This has the potential to enable the use of antibodies with longer half-lives, such as AMG191, and to simplify the development of ADCs. Clinical trial design is also simplified - HSCs can be infused before or at the same time as conditioning with little or no risk of depletion (e.g., in MDS or AML patients). These methods can benefit all patients regardless of their immune status.

[0285] In one embodiment, CD117 is altered (e.g., using base editing) in cells for transplantation to prevent binding of anti-CD117 antibodies, but without interfering with normal SCF signaling. Using base editing, changes in the bases can be made to effect amino acid substitutions in CD117. Stem cell factor (SCF) drives HSC self-renewal and differentiation into progenitor cells. Administration of anti-CD117 antibodies blocks the binding of SCF to CD117, thereby depleting HSCs and progenitor cells (conditioning) in the patient. Autologous gene-edited HSCs are transplanted into the patient. The gene-edited cells compete with residual host HSCs to repopulate the bone marrow (BM). Anti-CD117 antibodies block the binding of SCF to wild-type (WT) CD117, but not to HSCs with edited CD117. Thus, native wild-type HSCs are targeted by anti-CD117 antibodies, but gene-edited HSCs are not. Thus, both cell types express the CD117 polypeptide activated by SCF binding. However, binding of anti-CD117 antibodies to wild-type CD117 disrupts SCF binding and results in inhibition of SCF signaling in wild-type cells. In contrast, gene-edited HSCs are unresponsive to anti-CD117 antibodies because the amino acid substitutions introduced in CD117 prevent binding of anti-CD117 antibodies, but do not interfere with normal SCF binding and signaling.

[0286] In various embodiments, the antibodies of the invention are currently used clinically (e.g., if the edited epitope is required for binding of AMG 191). In one embodiment, CD117 is altered to be similar to murine CD117, which does not cross-react with AMG191 or SR1. Since murine SCF activates murine and human cells, interspecies variation in the epitope provides potential variants with SCF activity while having reduced or no anti-CD117 binding. To develop new antibodies, ADCs, CAR-T cells, etc. capable of targeting a new epitope in one of the extracellular domains of CD117 (D1-D5). Thus, methods for identifying candidate agents for selectively depleting or ablating endogenous stem cell populations are also within the scope of the invention. Such methods can include the steps of: (a) contacting a sample containing a stem cell population with a test agent (e.g., an antibody); and (b) detecting whether one or more cells of the stem cell population are depleted or ablated from the sample; wherein depletion or ablation of one or more cells of the stem cell population after the contacting step will identify the test agent as a candidate agent. In a further step, the edited cells are identified as not being similarly depleted or ablated by the reagent. In some embodiments, the cells are contacted with the test agent for at least about 2-24 hours.

[0287] Exemplary but non-limiting instances of compositions and methods for treating hemoglobinopathies are described in International Publication No. WO2020168133, which is incorporated herein by reference in its entirety.

[0288] Treatment methods

[0289] The methods and compositions disclosed herein can be used to condition a subject's tissue (e.g., bone marrow) for implantation or transplantation, and after such conditioning, a population of stem cells is administered to the subject. The transplanted cells (e.g., HSCs) can be autologous or allogeneic cells. In certain aspects, the population of stem cells includes an exogenous population of stem cells. In some embodiments, the population of stem cells comprises endogenous stem cells of the subject (e.g., endogenous stem cells that have been genetically modified to correct a disease or genetic defect). Preferably, such methods and compositions can be used to treat such diseases without causing the toxicities observed in response to conventional conditioning therapies (e.g., radiation).

[0290] A hematopoietic stem cell transplantation therapy can be administered to a subject in need of treatment to repopulate or reconstitute one or more blood cell types. Hematopoietic stem cells generally exhibit pluripotency and can thus differentiate into multiple different blood lineages, including but not limited to granulocytes (e.g., promyelocytes, neutrophils, eosinophils, basophils), erythrocytes (e.g., reticulocytes, red blood cells), platelets (e.g., megakaryocytes, platelet-producing megakaryocytes, platelets), monocytes (e.g., monocytes, macrophages), dendritic cells, microglia, osteoclasts, and lymphocytes (e.g., NK cells, B cells, and T cells). Hematopoietic stem cells also have the ability to self-renew and can thus produce daughter cells with the same potential as the parent cell, and also have the ability to be re-introduced into a transplant recipient, so they home to hematopoietic stem cell niches and reconstitute productive and sustained hematopoietic function.

[0291] Thus, hematopoietic stem cells can be administered to a patient having a defect or deficiency in one or more cell types of the hematopoietic lineage to reconstitute in vivo the defective or absent cell population, thereby treating pathologies associated with a defect or depletion in the endogenous blood cell population. Accordingly, the compositions and methods described herein can be used to treat non-malignant hemoglobinopathies (e.g., hemoglobinopathies selected from the group consisting of sickle cell anemia, thalassemia, Fanconi anemia, aplastic anemia, and Wiskott-Aldrich syndrome). Additionally or alternatively, the compositions and methods described herein can be used to treat malignancies or proliferative diseases, such as blood cancers, myeloproliferative diseases. In the case of cancer treatment, the compositions and methods described herein can be administered to a patient to deplete the endogenous hematopoietic stem cell population prior to hematopoietic stem cell transplantation therapy, in which case the transplanted cells can home to the resulting niche through the step of endogenous cell depletion and establish productive hematopoiesis. In turn, this can reconstitute the cell population depleted during eradication of cancer cells, such as during systemic chemotherapy. Exemplary blood cancers that can be treated using the compositions and methods described herein include, but are not limited to, acute myeloid leukemia, acute lymphoblastic leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia, multiple myeloma, diffuse large B-cell lymphoma, and non-Hodgkin lymphoma, as well as other cancerous conditions, including neuroblastoma.

[0292] The antibodies, antigen-binding fragments, and ligands described herein can be administered to a patient (e.g., a human patient having cancer, an autoimmune disease, or in need of hematopoietic stem cell transplantation therapy) in a variety of dosage forms. For example, the antibodies, antigen-binding fragments, and ligands described herein can be administered in an aqueous solution form, e.g., an aqueous solution containing one or more pharmaceutically acceptable excipients, to a patient having cancer, an autoimmune disease, or in need of hematopoietic stem cell transplantation therapy. Pharmaceutically acceptable excipients used in conjunction with the compositions and methods described herein include viscosity modifiers. The aqueous solution can be sterilized using techniques known in the art.

[0293] The antibodies, antigen-binding fragments, and ligands described herein can be administered by a variety of routes, such as orally, transdermally, subcutaneously, intranasally, intravenously, intramuscularly, intraocularly, or parenterally. The most suitable route of administration in any given case will depend on the particular antibody, antigen-binding fragment, or ligand being administered, the patient, the method of drug formulation, the method of administration (e.g., the time and route of administration), the age, body weight, sex of the patient, the severity of the disease being treated, the diet of the patient, and the excretion rate of the patient.

[0294] Nucleobase editors

[0295] The present disclosure relates to base editors or nucleobase editors for editing, modifying or altering a target nucleotide sequence of a polynucleotide. Nucleobase editors or base editors are described herein that comprise a polynucleotide programmable nucleotide binding domain and a nucleobase editing domain (e.g., adenosine deaminase, cytidine deaminase). The polynucleotide programmable nucleotide binding domain, when bound to a bound guide polynucleotide (e.g., gRNA), can specifically bind to a target polynucleotide sequence (i.e., through complementary base pairing sequences between the bases of the bound guide nucleic acid and the bases of the target polynucleotide), thereby localizing the base editor to the target nucleic acid sequence to be edited. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.

[0296] Polynucleotide programmable nucleotide binding domain

[0297] It should be understood that the polynucleotide programmable nucleotide binding domain may also include a nucleic acid programmable protein that binds RNA. For example, the polynucleotide programmable nucleotide binding domain may be associated with a nucleic acid that directs the polynucleotide programmable nucleotide binding domain to RNA. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they are not specifically listed in the present disclosure.

[0298] The polynucleotide programmable nucleotide binding domain of the base editor itself may comprise one or more domains. For example, the polynucleotide programmable nucleotide binding domain may comprise one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain may comprise an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to a protein or polypeptide capable of digesting nucleic acids (e.g., RNA or DNA) from free ends, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) within a nucleic acid (e.g., DNA or RNA). In some embodiments, the endonuclease may cleave a single strand of a double-stranded nucleic acid. In some embodiments, the endonuclease may cleave both strands of a double-stranded nucleic acid. In some embodiments the polynucleotide programmable nucleotide binding domain may be a deoxyribonuclease. In some embodiments the polynucleotide programmable nucleotide binding domain may be a ribonuclease.

[0299] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide-binding domain can cleave zero, one, or both strands of a target polynucleotide. In some cases, the polynucleotide programmable nucleotide-binding domain can comprise a nickase domain. As used herein, the term "nickase" refers to a polynucleotide programmable nucleotide-binding domain that comprises a nuclease domain that is capable of cleaving only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide-binding domain by introducing one or more mutations into the active polynucleotide programmable nucleotide-binding domain. For example, in the case where the polynucleotide programmable nucleotide-binding domain comprises a nickase domain derived from Cas9, the nickase domain derived from Cas9 can comprise a D10A mutation and histidine at position 840. In such cases, residue H840 retains catalytic activity and can thus cleave a single strand of the nucleic acid duplex. In another example, the nickase domain derived from Cas9 can comprise an H840A mutation, while the amino acid residue at position 10 remains D. In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide-binding domain by removing all or part of the nuclease domain that is not required for nickase activity. For example, in the case where the polynucleotide programmable nucleotide-binding domain comprises a nickase domain derived from Cas9, the nickase domain derived from Cas9 can comprise a complete or partial deletion of the RuvC domain or the HNH domain.

[0300] The amino acid sequence of an exemplary catalytically active Cas9 is as follows:

[0301]

[0302] Base editors comprising a polynucleotide programmable nucleotide binding domain that includes a nickase domain are thus capable of creating single-stranded DNA breaks (nicks) at a specific polynucleotide target sequence (e.g., determined by the complementary sequence of the bound guide nucleic acid). In some embodiments, the strand of the nucleic acid duplex target polynucleotide sequence that is nicked by the base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain) is the strand that is not edited by the base editor (i.e., the strand that is nicked by the base editor is opposite the strand that contains the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain) can nick the strand of the DNA molecule that is targeted for editing. In such cases, the non-targeted strand is not nicked.

[0303] Also provided herein are base editors that comprise a catalytically dead polynucleotide programmable nucleotide binding domain (i.e., cannot cleave a target polynucleotide sequence). As used herein, the terms “catalytically dead” and “nuclease dead” are used interchangeably and refer to a polynucleotide programmable nucleotide binding domain having one or more mutations and / or deletions that result in its inability to cleave a nucleic acid strand. In some embodiments, a catalytically dead polynucleotide programmable nucleotide binding domain base editor may lack nuclease activity due to specific point mutations in one or more nuclease domains. For example, in the case where the base editor comprises a Cas9 domain, the Cas9 can comprise D10A and H840A mutations. Such mutations inactivate both nuclease domains, resulting in loss of nuclease activity. In other embodiments, a catalytically dead polynucleotide programmable nucleotide binding domain can comprise one or more deletions of all or part of a catalytic domain (e.g., the RuvC1 and / or HNH domains). In further embodiments, a catalytically dead polynucleotide programmable nucleotide binding domain comprises a point mutation (e.g., D10A or H840A) and a deletion of all or part of a nuclease domain.

[0304] The present disclosure also contemplates mutations that can generate catalytically dead polynucleotide programmable nucleotide binding domains from previous functional versions of polynucleotide programmable nucleotide binding domains. For example, in the case of catalytically dead Cas9 (“dCas9”), variants are provided that have mutations other than D10A and H840A, which result in nuclease-inactivated Cas9. Such mutations include, for example, other amino acid substitutions at D10 and H840, or other substitutions within the Cas9 nuclease domain (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). Based on the present disclosure and knowledge in the art, other suitable nuclease-inactive dCas9 domains will be apparent to those of ordinary skill in the art and are within the scope of the present disclosure. Such additional exemplary suitable nuclease-inactive Cas9 domains include, but are not limited to, the D10A / H840A, D10A / D839A / H840A, and D10A / D839A / H840A / N863A mutant domains (see, e.g., Prashant et al., CAS9 transcriptional activators for target specificity screening and paired nickases for cooperative genome engineering. Nature Biotechnology. 2013;31(9):833 - 8338, the entire content of which is incorporated by reference).

[0305] Non-limiting examples of polynucleotide programmable nucleotide binding domains that can be incorporated into base editors include CRISPR protein-derived domains, restriction nucleases, meganucleases, transcription activator-like effector nucleases (TALENs), and zinc finger nucleases (ZFNs). In some cases, a base editor comprises a polynucleotide programmable nucleotide binding domain that comprises a native or modified protein or portion thereof that, through an associated guide nucleic acid, is capable of binding to a nucleic acid sequence during CRISPR (i.e., clustered regularly interspaced short palindromic repeats)-mediated nucleic acid modification. Such a protein is referred to herein as a "CRISPR protein". Accordingly, base editors comprising a polynucleotide programmable nucleotide binding domain that comprises all or a portion of a CRISPR protein are disclosed herein (i.e., base editors that comprise all or a portion of a CRISPR protein as a domain, also referred to as a "derived domain" of a "CRISPR protein" base editor). The CRISPR protein-derived domain incorporated into the base editor can be modified as compared to the wild-type or native version of the CRISPR protein. For example, as described below, the CRISPR protein-derived domain can comprise one or more mutations, insertions, deletions, rearrangements, and / or recombinations relative to the wild-type or native form of the CRISPR protein.

[0306] CRISPR is an adaptive immune system that provides protection against mobile genetic elements (viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to antecedent mobile elements, and target invading nucleic acids. The CRISPR clusters are transcribed and processed into CRISPR RNAs (crRNAs). In type II CRISPR systems, proper processing of pre-crRNA requires the trans-encoded small RNA (tracrRNA), the endogenous ribonuclease 3 (rnc), and the Cas9 protein. The tracrRNA serves as a guide for ribonuclease 3 to assist in processing pre-crRNA. Subsequently, the Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand that is not complementary to the crRNA is first cleaved endonucleolytically and then trimmed 3′-5' exonucleolytically. In nature, DNA binding and cleavage generally require a protein and two RNAs. However, a single guide RNA (“sgRNA”, or simply “gRNA”) can be engineered to incorporate aspects of both the crRNA and tracrRNA into a single RNA species. See, e.g., Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire content of which is incorporated herein by reference. Cas9 recognizes a short motif within the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self.

[0307] In some embodiments, the methods described herein can utilize engineered Cas proteins. A guide RNA (gRNA) is a short synthetic RNA that consists of a scaffold sequence required for Cas binding and a user-defined ~20 nucleotide spacer that defines the genomic (or polynucleotide, e.g., DNA or RNA) target to be modified. Thus, one of ordinary skill in the art can change the genomic or polynucleotide target of the Cas protein by changing the target sequence present in the gRNA. The specificity of the Cas protein depends in part on the specificity of the gRNA targeting sequence for the genomic nucleotide targeting sequence as compared to the rest of the genome.

[0308] In some embodiments, the gRNA scaffold sequence is as follows: GUUUUAGAGC UAGAAAUAGC AAGUUAAAAUAAGGCUAGUC CGUUAUCAAC UUGAAAAAGU GGCACCGAGU CGGUGCUUUU.

[0309] In one embodiment, the RNA scaffold comprises a stem-loop. In one embodiment, the RNA scaffold comprises the nucleic acid sequence:

[0310] GUUUUUGUACUCUCAAGAUUUAAGUAACUGUACAACGAAACUUACACAGUUACUUAAAUCUUGCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUG.

[0311] In one embodiment, the RNA scaffold comprises the nucleic acid sequence:

[0312] GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU.

[0313] In some embodiments, the Streptococcus pyogenes sgRNA scaffold polynucleotide sequence is as follows:

[0314] GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC.

[0315] In one embodiment, the Staphylococcus aureus scaffold polynucleotide sequence is as follows:

[0316] GUUUUAGUACUCUGUAAUGAAAAUUACAGAAUCUACUAAAACAAGGCAAAAUGCCGUGUUUAUCUCGUCAACUUGUUGGCGAGA.

[0317] In one embodiment, the BhCas12b sgRNA scaffold has the following polynucleotide sequence:

[0318] GUUCUGTCUUUUGGUCAGGACAACCGUCUAGCUAUAAGUGCUGCAGGGUGUGAGAAACUCCUAUUGCUGGACGAUGUCUCUUACGAGGCAUUAGCAC.

[0319] In one embodiment, the BvCas12b sgRNA scaffold has the following polynucleotide sequence:

[0320] GACCUAUAGGGUCAAUGAAUCUGUGCGUGUGCCAUAAGUAAUUAAAAAUUACCCACCACAGGAGCACCUGAAAACAGGUGCUUGGCAC.

[0321] In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is an endonuclease (e.g., a deoxyribonuclease or ribonuclease) capable of binding to a target polynucleotide when bound to a bound guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a nickase capable of binding to a target polynucleotide when bound to a bound guide nucleic acid. In some embodiments, the CRISPR protein-derived domain incorporated into the base editor is a catalytically dead domain capable of binding to a target polynucleotide when bound to a bound guide nucleic acid. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is DNA. In some embodiments, the target polynucleotide bound by the CRISPR protein-derived domain of the base editor is RNA.

[0322] Cas proteins that can be used herein include Class 1 and Class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i and Cas12j / CasΦ, CARF, DinG, their homologs or their modified versions. Unmodified CRISPR enzymes can have DNA cleavage activity, such as Cas9, which has two functional endonuclease domains: RuvC and HNH. CRISPR enzymes can direct cleavage of one or both strands at the target sequence, for example within the target sequence and / or within the complementary sequence of the target sequence. For example, CRISPR enzymes can direct cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 or more base pairs from the first or last nucleotide of the target sequence.

[0323] Vectors encoding CRISPR enzymes can be used, where the CRISPR enzyme is mutated relative to the corresponding wild-type enzyme such that the mutated CRISPR enzyme lacks the ability to cleave one or both strands of a target polynucleotide containing a target sequence. Cas9 can refer to a polypeptide having at least or at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity and / or sequence homology with a wild-type exemplary Cas9 polypeptide (e.g., Cas9 from Streptococcus pyogenes). Cas9 can refer to a polypeptide having at most or at most about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity and / or sequence homology with a wild-type exemplary Cas9 polypeptide (e.g., from Streptococcus pyogenes). Cas9 can refer to the wild-type or a modified form of the Cas9 protein, which can contain amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras or any combination thereof.

[0324] In some embodiments, the CRISPR protein-derived domain of the base editor can include those from Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheriae (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Refs: NC_021284.1); Prevotella intermedia (NCBI Refs: NC_017861.1); Spiroplasma taiwanense, China (NCBI Refs: NC_021846.1); Streptococcus iniae (NCBI Refs: NC_021314.1); Belliella baltica (NCBI Refs: NC_018010.1); Psychroflexus torquis I (NCBI Ref: NC_018721.1); Streptococcus thermophilus (NCBI Refs: YP_820832.1); Listeria innocua (NCBI Refs: NP_472073.1); Campylobacter jejuni (NCBI Refs: YP_002344900.1); Neisseria meningitidis (NCBI Refs: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus aureus.

[0325] The Cas9 domain of the nucleobase editor

[0326] The Cas9 nuclease sequence and structure are well known to those skilled in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti et al., J.J., McShan W.M., Ajdic D.J., Savic D.J., Savic G., Lyon K., Primeaux C., Sezate S., Suvorov A.N., Kenton S., Lai H.S., Lin S.P., Qian Y., Jia H.G., Najar F.Z., Ren Q., Zhu H., Song L., White J., Yuan X., Clifton S.W., Roe B.A., McLaughlin R.E., Proc. Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E., Chylinski K., Sharma C.M., Gonzales K., Chao Y., Pirzada Z.A., Eckert M.R., Vogel J., Charpentier E., Nature 471:602-607 (2011); and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna J.A., Charpentier E. Science 337:816-821 (2012), the entire contents of each of which are hereby incorporated by reference. Cas9 orthologs have been described in a variety of species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus.Based on the present disclosure, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5, 726-737; the entire content of which is incorporated herein by reference.

[0327] In some aspects, the nucleic acid programmable DNA binding protein (napDNAbp) is a Cas9 domain. Non-limiting exemplary Cas9 domains are provided herein. The Cas9 domain can be a nuclease-active Cas9 domain, a nuclease-inactive Cas9 domain, or a Cas9 nickase. In some embodiments, the Cas9 domain is a domain having nuclease activity. For example, the Cas9 domain can be a Cas9 domain that cleaves both strands of a double-stranded nucleic acid (e.g., both strands of a double-stranded DNA molecule). In some embodiments, the Cas9 domain comprises any of the amino acid sequences described herein. In some embodiments, the amino acid sequence comprised by the Cas9 domain is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or 99.5% identical to any of the amino acid sequences described herein. In some embodiments, the amino acid sequence comprised by the Cas9 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences described herein. In some embodiments, compared to any of the amino acid sequences described herein, the Cas9 domain comprises at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100 or at least 1200 identical consecutive amino acid residues.

[0328] In some embodiments, proteins comprising Cas9 fragments are provided. For example, in some embodiments, the protein comprises one of two Cas9 domains: (1) the gRNA-binding domain of Cas9; or (2) the DNA cleavage domain of Cas9. In some embodiments, a protein comprising Cas9 or a fragment thereof is referred to as a "Cas9 variant". The Cas9 variant is homologous to Cas9 or a fragment thereof. For example, the Cas9 variant is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5% or at least about 99.9% identical to wild-type Cas9. In some embodiments, compared to wild-type Cas9, the Cas9 variant may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acid changes. In some embodiments, the Cas9 variant comprises a fragment of Cas9 (e.g., the gRNA-binding domain or the DNA cleavage domain) such that the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5% or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, the fragment is at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% the amino acid length of the corresponding wild-type Cas9. In some embodiments, the length of the fragment is at least 100 amino acids. In some embodiments, the length of the fragment is at least 100, 150, 200, 250, 300, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250 or 1300 amino acids.

[0329] In some embodiments, the Cas9 fusion proteins provided herein comprise the full-length amino acid sequence of a Cas9 protein, such as one of the Cas9 sequences provided herein. However, in other embodiments, the fusion proteins provided herein do not comprise the full-length Cas9 sequence, but only one or more fragments thereof. Exemplary amino acid sequences of suitable Cas9 domains and Cas9 fragments are provided herein, and other suitable sequences of Cas9 domains and fragments will be apparent to those skilled in the art.

[0330] The Cas9 protein can be associated with a guide RNA that directs the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the polynucleotide programmable nucleotide-binding domain is a Cas9 domain, such as nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactivated Cas9 (dCas9). Examples of nucleic acid programmable DNA-binding proteins include, but are not limited to, Cas9 (such as dCas9 and nCas9), CasX, CasY, Cpf1, Cas12b / C2C1, Cas12c / C2C3, and Cas12j / CasΦ.

[0331] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, nucleotide and amino acid sequences are as follows).

[0332]

[0333]

[0334]

[0335] (Single underline: HNH domain; double underline: RuvC domain)

[0336] In some embodiments, wild-type Cas9 corresponds to or comprises the following nucleotide and / or amino acid sequences:

[0337]

[0338]

[0339] (Single underline: HNH domain; double underline: RuvC domain).

[0340] In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes

[0341] (NCBI reference sequence: NC_002737.2 (nucleotide sequence is as follows); and Uniprot reference sequence: Q99ZW2 (amino acid sequence is as follows):

[0342]

[0343] (Single underline: HNH domain; double underline: RuvC domain)

[0344] In some embodiments, Cas9 refers to Cas9 from the following: Corynebacterium ulcerans (NCBI Refs: NC_015683.1, NC_017317.1); Corynebacterium diphtheria (NCBI Refs: NC_016782.1, NC_016786.1); Spiroplasma syrphidicola (NCBI Refs: NC_021284.1); Prevotella intermedia (NCBI Refs: NC_017861.1); Spiroplasma taiwanense, China (NCBI Refs: NC_021846.1); Streptococcus iniae (NCBI Refs: NC_021314.1); Belliella baltica (NCBI Refs: NC_018010.1); Psychroflexus torquis I (NCBI Refs: NC_018721.1); Streptococcus thermophilus (NCBI Refs: YP_820832.1); Listeria innocua (NCBI Refs: NP_472073.1); Campylobacter jejuni (NCBI Refs: YP_002344900.1); Neisseria meningitidis (NCBI Refs: YP_002342100.1); or Cas9 from any other organism.

[0345] It should be understood that additional Cas9 proteins (e.g., nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9), including variants and homologs thereof, are within the scope of the present disclosure. Exemplary Cas9 proteins include, but are not limited to, those provided below. In some embodiments, the Cas9 protein is nuclease-inactive Cas9 (dCas9). In some embodiments, the Cas9 protein is a Cas9 nickase (nCas9). In some embodiments, the Cas9 protein is a Cas9 with nuclease activity.

[0346] In some embodiments, the Cas9 domain is a nuclease-inactivated domain (dCas9). For example, the dCas9 domain can bind to a double-stranded nucleic acid molecule (e.g., via a gRNA molecule) without cleaving either strand of the double-stranded nucleic acid molecule. In some embodiments, the nuclease-inactivated dCas9 domain contains the D10X mutation and the H840X mutation of the amino acid sequence described herein, or the corresponding mutations in any amino acid sequence provided herein, where X is any amino acid change. In some embodiments, the nuclease-inactivated dCas9 domain contains the D10A mutation and the H840A mutation of the amino acid sequence described herein, or the corresponding mutations in any amino acid sequence provided herein. As an example, the nuclease-inactive Cas9 domain contains the amino acid sequence listed in the cloning vector pPlatTET-gRNA2 (accession number BAV54124).

[0347] An exemplary amino acid sequence of catalytically inactive Cas9 (dCas9) is as follows:

[0348]

[0349] (See, e.g., Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression.” Cell. 2013;152(5):1173-83, the entire content of which is incorporated herein by reference).

[0350] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase, referred to as the “nCas9” protein (for “nickase” Cas9). The nuclease-inactivated Cas9 protein may be interchangeably referred to as the “dCas9” protein (for nuclease-“dead” Cas9) or catalytically inactivated Cas9. Methods for generating Cas9 proteins (or fragments thereof) with an inactive DNA cleavage domain are known (see, e.g., Jinek et al., Science. 337:816-821 (2012); Qi et al., “Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression” (2013) Cell. 28;152(5):1173-83, the entire content of which is incorporated herein by reference). For example, it is known that the DNA cleavage domain of Cas9 includes two subdomains, the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A and H840A completely inactivate the nuclease activity of Streptococcus pyogenes Cas9 (Jinek et al., Science. 337:816-821 (2012); Qi et al., Cell. 28:152(5):1173-83 (2013)).

[0351] In some embodiments, the amino acid sequence contained in the dCas9 domain is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% or 99.5% identical to any of the Cas9 domains described herein. In some embodiments, the amino acid sequence contained in the Cas9 domain has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences described herein. In some embodiments, compared to any of the amino acid sequences described herein, the Cas9 domain contains at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100 or at least 1200 identical consecutive amino acid residues.

[0352] In some embodiments, dCas9 corresponds to or partially or fully contains a Cas9 amino acid sequence having one or more mutations that inactivate the Cas9 nuclease activity. For example, in some embodiments, the dCas9 domain contains the D10A and H840A mutations or corresponding mutations in another Cas9.

[0353] In some embodiments, dCas9 contains the amino acid sequence of dCas9 (D10A and H840A):

[0354] (Single underline: HNH domain; double underline: RuvC domain).

[0355] In some embodiments, the Cas9 domain contains the D10A mutation, while the residue at position 840 remains histidine at the corresponding position in the amino acid sequence provided above or in any of the amino acid sequences provided herein.

[0356] In other embodiments, dCas9 variants having mutations other than D10A and H840A are provided, which result in, for example, nuclease-inactivated Cas9 (dCas9). For example, such mutations include other amino acid substitutions at D10 and H840, or other substitutions within the Cas9 nuclease domain (e.g., substitutions in the HNH nuclease subdomain and / or the RuvC1 subdomain). In some embodiments, variants or homologs of dCas9 are provided that are at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical. In some embodiments, variants are provided that are shorter or longer by about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 40 amino acids, about 50 amino acids, about 75 amino acids, about 100 amino acids, or more.

[0357] In some embodiments, the Cas9 domain is a Cas9 nickase. A Cas9 nickase can be a Cas9 protein that is only capable of cleaving one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase cleaves the target strand of a double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that base pairs (is complementary) with the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase contains a D10A mutation and has a histidine at position 840. In some embodiments, the Cas9 nickase cleaves the non-target, non-base-editing strand of a double-stranded nucleic acid molecule, meaning that the Cas9 nickase cleaves the strand that does not base pair with the gRNA (e.g., sgRNA) bound to the Cas9. In some embodiments, the Cas9 nickase contains an H840A mutation and has an aspartic acid residue or a corresponding mutation at position 10. In some embodiments, the Cas9 nickase contains an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% or 99.5% identical to any of the Cas9 nickases described herein. Based on the present disclosure and knowledge in the art, other suitable Cas9 nickases will be apparent to those skilled in the art and are within the scope of the present disclosure.

[0358] The amino acid sequences of exemplary catalytically active Cas9 nickases (nCas9) are as follows:

[0359]

[0360] In some embodiments, Cas9 refers to Cas9 from archaea (e.g., Nanoarchaeum), which constitutes the domain and kingdom of single-celled prokaryotic microorganisms. In some embodiments, the programmable nucleotide-binding protein can be a CasX or CasY protein, which has been described, for example, in Burstein et al., "New CRISPR-Cas systems from uncultivated microbes." Cell Res. 2017 Feb 21. doi:10.1038 / cr.2017.21, the entire content of which is incorporated herein by reference. Using genome-resolved metagenomics, many CRISPR-Cas systems have been identified, including Cas9 first reported in the archaea domain. This divergent Cas9 protein was found in the little-studied Nanoarchaeum as part of an active CRISPR-Cas system. In bacteria, two previously unknown systems, CRISPR-CasX and CRISPR-CasY, were found, which are one of the most compact systems discovered to date. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasX or a variant of CasX. In some embodiments, in the base editor systems described herein, Cas9 is replaced by CasY or a variant of CasY. It should be understood that other RNA-guided DNA-binding proteins can be used as nucleic acid programmable DNA-binding proteins (napDNAbp) and are within the scope of the present disclosure.

[0361] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) or any fusion protein provided herein can be a CasX or CasY protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp is a CasY protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide-binding protein is a naturally occurring CasX or CasY protein. In some embodiments, the programmable nucleotide-binding protein comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to any CasX or CasY protein described herein. It should be understood that CasX and CasY from other bacterial species can also be used according to the present disclosure.

[0362] The amino acid sequence of exemplary CasX ((uniprot.org / uniprot / F0NN87; uniprot.org / uniprot / F0NH53) tr|F0NN87|F0NN87_SULI HCRISPR-associated Casx protein OS = Sulfolobus islandicus (strain HVE10 / 4) GN = SiH_0402 PE = 4 SV = 1) is as follows:

[0363] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYEFGRSPGMVERTRRVKLEVEPHYLIIAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVRIYTISDAVGQNPTTINGGFSIDLTKLLEKRYLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTG SKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.

[0364] The amino acid sequence of exemplary CasX (>tr|F0NH53|F0NH53_SULI R CRISPR associated protein, Casx OS = Sulfolobus islandicus (strain REY15A) GN = SiRe_0771 PE = 4 SV = 1) is as follows:

[0365] MEVPLYNIFGDNYIIQVATEAENSTIYNNKVEIDDEELRNVLNLAYKIAKNNEDAAAERRGKAKKKKGEEGETTTSNIILPLSGNDKNPWTETLKCYNFPTTVALSEVFKNFSQVKECEEVSAPSFVKPEFYKFGRSPGMVERTRRVKLEVEPHYLIMAAAGWVLTRLGKAKVSEGDYVGVNVFTPTRGILYSLIQNVNGIVPGIKPETAFGLWIARKVVSSVTNPNVSVVSIYTISDAVGQNPTTINGGFSIDLTKLLEKRDLLSERLEAIARNALSISSNMRERYIVLANYIYEYLTGSKRLEDLLYFANRDLIMNLNSDDGKVRDLKLISAYVNGELIRGEG.

[0366] Delta Proteobacterium CasX

[0367] MEKRINKIRKKLSADNATKPVSRSGPMKTLLVRVMTDDLKKRLEKRRKKPEVMPQVISNNAANNLRMLLDDYTKMKEAILQVYWQEFKDDHVGLMCKFAQPASKKIDQNKLKPEMDEKGNLTTAGFACSQCGQPLFVYKLEQVSEKGKAYTNYFGRCNVAEHEKLILLAQLKPVKDSDEAVTYSLGKFGQRALDFYSIHVTKESTHPVKPLAQIAGNRYASGPVGKALSDACMGTIASFLSKYQDIIIEHQKVVKGNQKRLESLRELAGKENLEYPSVTLPPQPHTKEGVDfAYNEVIARVRMWVNLNLWQKLKLSRDDAKPLLRLKGFPSFPVVERRENEVDWWNTINEVKKLIDAKRDMGRVFWSGVTAEKRNTILEGYNYLPNENDHKKREGSLENPKKPAKRQFGDLLLYLEKKYAGDWGKVFDEAWERIDKKIAGLTSHIEREEARNAEDAQSKAVLTDWLRAKASFVLERLKEMDEKEFYACEIQLQKWYGDLRGNPFAVEAENRVVDISGFSIGSDGHSIQYRNLLAWKYLENGKREFYLLMNYGKKGRIRFTDGTDIKKSGKWQGLLYGGGKAKVIDLTFDPDDEQLIILPLAFGTRQGREFIWNDLLSLETGLIKLANGRVIEKTIYNKKIGRDEPALFVALTFERREVVDPSNIKPVNLIGVARGENIPAVIALTDPEGCPLPEFKDSSGGPTDILRIGEGYKEKQRAIQAAKEVEQRRAGGYSRKFASKSRNLADDMVRNSARDLFYHAVTHDAVLVFANLSRGFGRQGKRTFMTERQYTKMEDWLTAKLAYEGLTSKTYLSKTLAQYTSKTCSNCGFTITYADMDVMLVRLKKTSDGWATTLNNKELKAEYQITYYNRYKRQTVEKELSAELDRLSEESGNNDISKWTKGRRDEALFLLKKRFSHRPVQEQFVCLDCGHEVHAAEQAALNIARSWLFLNSNSTEFKSYKSGKQPFVGAWQAFYKRRLKEVWKPNA

[0368] The amino acid sequence of exemplary CasY ((ncbi.nlm.nih.gov / protein / APG80656.1)>APG80656.1 CRISPR-associated protein CasY [uncultured Parcubacteria group]) is as follows:

[0369]

[0370] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) is a single effector of the microbial CRISPR-Cas system. The single effectors of the microbial CRISPR-Cas system include, but are not limited to, Cas9, Cpf1, Cas12b / C2c1, and Cas12c / C2c3. Generally, the microbial CRISPR-Cas system is divided into class 1 and class 2 systems. Class 1 systems have multi-subunit effector complexes, while class 2 systems have single protein effectors. For example, Cas9 and Cpf1 are class 2 effectors. In addition to Cas9 and Cpf1, three different class 2 CRISPR-Cas systems (Cas12b / C2c1 and Cas12c / C2c3) were described by Shmakov et al. in "Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems", Mol. Cell, 2015 Nov. 5; 60(3): 385-397, the entire content of which is incorporated herein by reference. The effectors of two of the systems, Cas12b / C2c1 and Cas12c / C2c3, contain a RuvC-like endonuclease domain related to Cpf1. The third system contains an effector with two predicted HEPN RNase domains. The production of mature CRISPR RNA is independent of tracrRNA, which is different from the CRISPR RNA produced by Cas12b / C2c1. Cas12b / C2c1 relies on CRISPR RNA and tracrRNA for DNA cleavage.

[0371] It has been reported that the crystal structure of Alicyclobaccillus acidoterrastris Cas12b / C2c1 (AacC2c1) is complexed with a chimeric single-guide RNA (sgRNA). See, e.g., Liu et al., “C2c1-sgRNA Complex Structure Reveals RNA-guided DNA Cleavage Mechanism”, Mol. Cell, Jan. 19, 2017; 65(2):310-322, the entire content of which is incorporated herein by reference. Crystal structures have also been reported for AacC2c1 in complex with target DNA in a ternary complex form. See, e.g., Yang et al., “PAM-dependent Target DNA Recognition and Cleavage by C2C1 CRISPR-Cas Endonuclease”, Cell, Dec. 15, 2016; 167(7):1814-1828, the entire content of which is incorporated herein by reference. The catalytically competent conformation of AacC2c1, including the target DNA strand and the non-target DNA strand, has been independently captured within a single RuvC catalytic pocket, and Cas12b / C2c1-mediated cleavage results in a seven-nucleotide staggered break in the target DNA. Structural comparisons between the Cas12b / C2c1 ternary complex and previously characterized Cas9 and Cpf1 counterparts demonstrate the diversity of the mechanisms used by the CRISPR-Cas9 system.

[0372] In some embodiments, the nucleic acid programmable DNA-binding protein (napDNAbp) or any fusion protein provided herein can be a Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a Cas12b / C2c1 protein. In some embodiments, the napDNAbp is a Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp is a naturally occurring Cas12b / C2c1 or Cas12c / C2c3 protein. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to any of the napDNAbp sequences provided herein. It should be understood that Cas12b / C2c1 or Cas12c / C2c3 from other bacterial species can also be used according to the present disclosure.

[0373] The amino acid sequence of a Cas12b / C2c1 ((uniprot.org / uniprot / T0D7A2#2) sp|T0D7A2|C2C1_ALIAGCRISPR-associated endonuclease C2c1 OS= Alicyclobacillus acidoterrestris (strain ATCC 49025 / DSM3922 / CIP 106132 / NCIMB 13137 / GD3B) GN=c2c1 PE=1 SV=1) is as follows:

[0374]

[0375] BhCas12b (Bacillus hiramatsui) NCBI Reference Sequence: WP_095142515

[0376]

[0377] In some embodiments, Cas12b is BvCas12B, which is a variant of BhCas12b and contains the following changes relative to BhCas12B: S893R, K846R, and E837G.

[0378] BvCas12b (Bacillus sp. V3-13) NCBI Reference Sequence: WP_101661451.1

[0379]

[0380] In some embodiments, the nucleic acid programmable DNA binding protein (napDNAbp) or any fusion protein provided herein can be a Cas12j / CasΦ protein. Cas12j / CasΦ was described by Pausch et al., “CRISPR-CasΦ from huge phages is a hypercompact genome editor”, Science, July 17, 2020, Vol. 369, No. 6501, pp. 333-337, which is hereby incorporated by reference in its entirety. In some embodiments, the napDNAbp comprises an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or at least 99.5% identical to a naturally occurring Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a naturally occurring Cas12j / CasΦ protein. In some embodiments, the napDNAbp is a nuclease-inactivated (dead) Cas12j / CasΦ protein. It should be understood that Cas12j / CasΦ from other species can also be used according to the present disclosure.

[0381] Exemplary Cas12j / CasΦ amino acid sequences are as follows:

[0382] >CasΦ-1

[0383] MADTPTLFTQFLRHHLPGQRFRKDILKQAGRILANKGEDATIAFLRGKSEESPPDFQPPVKCPIIACSRPLTEWPIYQASVAIQGYVYGQSLAEFEASDPGCSKDGLLGWFDKTGVCTDYFSVQGLNLIFQNARKRYIGVQTKVTNRNEKRHKKLKRINAKRIAEGLPELTSDEPESALDETGHLIDPPGLNTNIYCYQQVSPKPLALSEVNQLPTAYAGYSTSGDDPIQPMVTKDRLSISKGQPGYIPEHQRALLSQKKHRRMRGYGLKARALLVIVRIQDDWAVIDLRSLLRNAYWRRIVQTKEPSTITKLLKLVTGDPVLDATRMVATFTYKPGIVQVRSAKCLKNKQGSKLFSERYLNETVSVTSIDLGSNNLVAVATYRLVNGNTPELLQRFTLPSHLVKDFERYKQAHDTLEDSIQKTAVASLPQGQQTEIRMWSMYGFREAQERVCQELGLADGSIPWNVMTATSTILTDLFLARGGDPKKCMFTSEPKKKKNSKQVLYKIRDRAWAKMYRTLLSKETREAWNKALWGLKRGSPDYARLSKRKEELARRCVNYTISTAEKRAQCGRTIVALEDLNIGFFHGRGKQEPGWVGLFTRKKENRWLMQALHKAFLELAHHRGYHVIEVNPAYTSQTCPVCRHCDPDNRDQHNREAFHCIGCGFRGNADLDVATHNIAMVAITGESLKRARGSVASKTPQPLAAE*

[0384] >CasΦ-2

[0385] MPKPAVESEFSKVLKKHFPGERFRSSYMKRGGKILAAQGEEAVVAYLQGKSEEEPPNFQPPAKCHVVTKSRDFAEWPIMKASEAIQRYIYALSTTERAACKPGKSSESHAAWFAATGVSNHGYSHVQGLNLIFDHTLGRYDGVLKKVQLRNEKARARLESINASRADEGLPEIKAEEEEVATNETGHLLQPPGINPSFYVYQTISPQAYRPRDEIVLPPEYAGYVRDPNAPIPLGVVRNRCDIQKGCPGYIPEWQREAGTAISPKTGKAVTVPGLSPKKNKRMRRYWRSEKEKAQDALLVTVRIGTDWVVIDVRGLLRNARWRTIAPKDISLNALLDLFTGDPVIDVRRNIVTFTYTLDACGTYARKWTLKGKQTKATLDKLTATQTVALVAIDLGQTNPISAGISRVTQENGALQCEPLDRFTLPDDLLKDISAYRIAWDRNEEELRARSVEALPEAQQAEVRALDGVSKETARTQLCADFGLDPKRLPWDKMSSNTTFISEALLSNSVSRDQVFFTPAPKKGAKKKAPVEVMRKDRTWARAYKPRLSVEAQKLKNEALWALKRTSPEYLKLSRRKEELCRRSINYVIEKTRRRTQCQIVIPVIEDLNVRFFHGSGKRLPGWDNFFTAKKENRWFIQGLHKAFSDLRTHRSFYVFEVRPERTSITCPKCGHCEVGNRDGEAFQCLSCGKTCNADLDVATHNLTQVALTGKTMPKREEPRDAQGTAPARKTKKASKSKAPPAEREDQTPAQEPSQTS

[0386] >CasΦ-3

[0387] MEKEITELTKIRREFPNKKFSSTDMKKAGKLLKAEGPDAVRDFLNSCQEIIGDFKPPVKTNIVSISRPFEEWPVSMVGRAIQEYYFSLTKEELESVHPGTSSEDHKSFFNITGLSNYNYTSVQGLNLIFKNAKAIYDGTLVKANNKNKKLEKKFNEINHKRSLEGLPIITPDFEEPFDENGHLNNPPGINRNIYGYQGCAAKVFVPSKHKMVSLPKEYEGYNRDPNLSLAGFRNRLEIPEGEPGHVPWFQRMDIPEGQIGHVNKIQRFNFVHGKNSGKVKFSDKTGRVKRYHHSKYKDATKPYKFLEESKKVSALDSILAIITIGDDWVVFDIRGLYRNVFYRELAQKGLTAVQLLDLFTGDPVIDPKKGVVTFSYKEGVVPVFSQKIVPRFKSRDTLEKLTSQGPVALLSVDLGQNEPVAARVCSLKNINDKITLDNSCRISFLDDYKKQIKDYRDSLDELEIKIRLEAINSLETNQQVEIRDLDVFSADRAKANTVDMFDIDPNLISWDSMSDARVSTQISDLYLKNGGDESRVYFEINNKRIKRSDYNISQLVRPKLSDSTRKNLNDSIWKLKRTSEEYLKLSKRKLELSRAVVNYTIRQSKLLSGINDIVIILEDLDVKKKFNGRGIRDIGWDNFFSSRKENRWFIPAFHKAFSELSSNRGLCVIEVNPAWTSATCPDCGFCSKENRDGINFTCRKCGVSYHADIDVATLNIARVAVLGKPMSGPADRERLGDTKKPRVARSRKTMKRKDISNSTVEAMVTA*

[0388] >CasΦ-4

[0389] MYSLEMADLKSEPSLLAKLLRDRFPGKYWLPKYWKLAEKKRLTGGEEAACEYMADKQLDSPPPNFRPPARCVILAKSRPFEDWPVHRVASKAQSFVIGLSEQGFAALRAAPPSTADARRDWLRSHGASEDDLMALEAQLLETIMGNAISLHGGVLKKIDNANVKAAKRLSGRNEARLNKGLQELPPEQEGSAYGADGLLVNPPGLNLNIYCRKSCCPKPVKNTARFVGHYPGYLRDSDSILISGTMDRLTIIEGMPGHIPAWQREQGLVKPGGRRRRLSGSESNMRQKVDPSTGPRRSTRSGTVNRSNQRTGRNGDPLLVEIRMKEDWVLLDARGLLRNLRWRESKRGLSCDHEDLSLSGLLALFSGDPVIDPVRNEVVFLYGEGIIPVRSTKPVGTRQSKKLLERQASMGPLTLISCDLGQTNLIAGRASAISLTHGSLGVRSSVRIELDPEIIKSFERLRKDADRLETEILTAAKETLSDEQRGEVNSHEKDSPQTAKASLCRELGLHPPSLPWGQMGPSTTFIADMLISHGRDDDAFLSHGEFPTLEKRKKFDKRFCLESRPLLSSETRKALNESLWEVKRTSSEYARLSQRKKEMARRAVNFVVEISRRKTGLSNVIVNIEDLNVRIFHGGGKQAPGWDGFFRPKSENRWFIQAIHKAFSDLAAHHGIPVIESDPQRTSMTCPECGHCDSKNRNGVRFLCKGCGASMDADFDAACRNLERVALTGKPMPKPSTSCERLLSATTGKVCSDHSLSHDAIEKAS*

[0390] >CasΦ-5

[0391] MSSLPTPLELLKQKHADLFKGLQFSSKDNKMAGKVLKKDGEEAALAFLSERGVSRGELPNFRPPAKTLVVAQSRPFEEFPIYRVSEAIQLYVYSLSVKELETVPSGSSTKKEHQRFFQDSSVPDFGYTSVQGLNKIFGLARGIYLGVITRGENQLQKAKSKHEALNKKRRASGEAETEFDPTPYEYMTPERKLAKPPGVNHSIMCYVDISVDEFDFRNPDGIVLPSEYAGYCREINTAIEKGTVDRLGHLKGGPGYIPGHQRKESTTEGPKINFRKGRIRRSYTALYAKRDSRRVRQGKLALPSYRHHMMRLNSNAESAILAVIFFGKDWVVFDLRGLLRNVRWRNLFVDGSTPSTLLGMFGDPVIDPKRGVVAFCYKEQIVPVVSKSITKMVKAPELLNKLYLKSEDPLVLVAIDLGQTNPVGVGVYRVMNASLDYEVVTRFALESELLREIESYRQRTNAFEAQIRAETFDAMTSEEQEEITRVRAFSASKAKENVCHRFGMPVDAVDWATMGSNTIHIAKWVMRHGDPSLVEVLEYRKDNEIKLDKNGVPKKVKLTDKRIANLTSIRLRFSQETSKHYNDTMWELRRKHPVYQKLSKSKADFSRRVVNSIIRRVNHLVPRARIVFIIEDLKNLGKVFHGSGKRELGWDSYFEPKSENRWFIQVLHKAFSETGKHKGYYIIECWPNWTSCTCPKCSCCDSENRHGEVFRCLACGYTCNTDFGTAPDNLVKIATTGKGLPGPKKRCKGSSKGKNPKIARSSETGVSVTESGAPKVKKSSPTQTSQSSSQSAP*

[0392] >CasΦ-6

[0393] MNKIEKEKTPLAKLMNENFAGLRFPFAIIKQAGKKLLKEGELKTIEYMTGKGSIEPLPNFKPPVKCLIVAKRRDLKYFPICKASCEIQSYVYSLNYKDFMDYFSTPMTSQKQHEEFFKKSGLNIEYQNVAGLNLIFNNVKNTYNGVILKVKNRNEKLKKKAIKNNYEFEEIKTFNDDGCLINKPGINNVIYCFQSISPKILKNITHLPKEYNDYDCSVDRNIIQKYVSRLDIPESQPGHVPEWQRKLPEFNNTNNPRRRRKWYSNGRNISKGYSVDQVNQAKIEDSLLAQIKIGEDWIILDIRGLLRDLNRRELISYKNKLTIKDVLGFFSDYPIIDIKKNLVTFCYKEGVIQVVSQKSIGNKKSKQLLEKLIENKPIALVSIDLGQTNPVSVKISKLNKINNKISIESFTYRFLNEEILKEIEKYRKDYDKLELKLINEA

[0394] >CasΦ-7

[0395] MSNTAVSTREHMSNKTTPPSPLSLLLRAHFPGLKFESQDYKIAGKKLRDGGPEAVISYLTGKGQAKLKDVKPPAKAFVIAQSRPFIEWDLVRVSRQIQEKIFGIPATKGRPKQDGLSETAFNEAVASLEVDGKSKLNEETRAAFYEVLGLDAPSLHAQAQNALIKSAISIREGVLKKVENRNEKNLSKTKRRKEAGEEATFVEEKAHDERGYLIHPPGVNQTIPGYQAVVIKSCPSDFIGLPSGCLAKESAEALTDYLPHDRMTIPKGQPGYVPEWQHPLLNRRKNRRRRDWYSASLNKPKATCSKRSGTPNRKNSRTDQIQSGRFKGAIPVLMRFQDEWVIIDIRGLLRNARYRKLLKEKSTIPDLLSLFTGDPSIDMRQGVCTFIYKAGQACSAKMVKTKNAPEILSELTKSGPVVLVSIDLGQTNPIAAKVSRVTQLSDGQLSHETLLRELLSNDSSDGKEIARYRVASDRLRDKLANLAVERLSPEHKSEILRAKNDTPALCKARVCAALGLNPEMIAWDKMTPYTEFLATAYLEKGGDRKVATLKPKNRPEMLRRDIKFKGTEGVRIEVSPEAAEAYREAQWDLQRTSPEYLRLSTWKQELTKRILNQLRHKAAKSSQCEVVVMAFEDLNIKMMHGNGKWADGGWDAFFIKKRENRWFMQAFHKSLTELGAHKGVPTIEVTPHRTSITCTKCGHCDKANRDGERFACQKCGFVAHADLEIATDNIERVALTGKPMPKPESERSGDAKKSVGARKAAFKPEEDAEAAE*

[0396] >CasΦ-8

[0397] MIKPTVSQFLTPGFKLIRNHSRTAGLKLKNEGEEACKKFVRENEIPKDECPNFQGGPAIANIIAKSREFTEWEIYQSSLAIQEVIFTLPKDKLPEPILKEEWRAQWLSEHGLDTVPYKEAAGLNLIIKNAVNTYKGVQVKVDNKNKNNLAKINRKNEIAKLNGEQEISFEEIKAFDDKGYLLQKPSPNKSIYCYQSVSPKPFITSKYHNVNLPEEYIGYYRKSNEPIVSPYQFDRLRIPIGEPGYVPKWQYTFLSKKENKRRKLSKRIKNVSPILGIICIKKDWCVFDMRGLLRTNHWKKYHKPTDSINDLFDYFTGDPVIDTKANVVRFRYKMENGIVNYKPVREKKGKELLENICDQNGSCKLATVDVGQNNPVAIGLFELKKVNGELTKTLISRHPTPIDFCNKITAYRERYDKLESSIKLDAIKQLTSEQKIEVDNYNNNFTPQNTKQIVCSKLNINPNDLPWDKMISGTHFISEKAQVSNKSEIYFTSTDKGKTKDVMKSDYKWFQDYKPKLSKEVRDALSDIEWRLRRESLEFNKLSKSREQDARQLANWISSMCDVIGIENLVKKNNFFGGSGKREPGWDNFYKPKKENRWWINAIHKALTELSQNKGKRVILLPAMRTSITCPKCKYCDSKNRNGEKFNCLKCGIELNADIDVATENLATVAITAQSMPKPTCERSGDAKKPVRARKAKAPEFHDKLAPSYTVVLREAV*

[0398] >CasΦ-9

[0399] MRSSREIGDKILMRQPAEKTAFQVFRQEVIGTQKLSGGDAKTAGRLYKQGKMEAAREWLLKGARDDVPPNFQPPAKCLVVAVSHPFEEWDISKTNHDVQAYIYAQPLQAEGHLNGLSEKWEDTSADQHKLWFEKTGVPDRGLPVQAINKIAKAAVNRAFGVVRKVENRNEKRRSRDNRIAEHNRENGLTEVVREAPEVATNADGFLLHPPGIDPSILSYASVSPVPYNSSKHSFVRLPEEYQAYNVEPDAPIPQFVVEDRFAIPPGQPGYVPEWQRLKCSTNKHRRMRQWSNQDYKPKAGRRAKPLEFQAHLTRERAKGALLVVMRIKEDWVVFDVRGLLRNVEWRKVLSEEAREKLTLKGLLDLFTGDPVIDTKRGIVTFLYKAEITKILSKRTVKTKNARDLLLRLTEPGEDGLRREVGLVAVDLGQTHPIAAAIYRIGRTSAGALESTVLHRQGLREDQKEKLKEYRKRHTALDSRLRKEAFETLSVEQQKEIVTVSGSGAQITKDKVCNYLGVDPSTLPWEKMGSYTHFISDDFLRRGGDPNIVHFDRQPKKGKVSKKSQRIKRSDSQWVGRMRPRLSQETAKARMEADWAAQNENEEYKRLARSKQELARWCVNTLLQNTRCITQCDEIVVVIEDLNVKSLHGKGAREPGWDNFFTPKTENRWFIQILHKTFSELPKHRGEHVIEGCPLRTSITCPACSYCDKNSRNGEKFVCVACGATFHADFEVATYNLVRLATTGMPMPKSLERQGGGEKAGGARKARKKAKQVEKIVVQANANVTMNGASLHSP*

[0400] >CasΦ-10

[0401] MDMLDTETNYATETPAQQQDYSPKPPKKAQRAPKGFSKKARPEKKPPKPITLFTQKHFSGVRFLKRVIRDASKILKLSESRTITFLEQAIERDGSAPPDVTPPVHNTIMAVTRPFEEWPEVILSKALQKHCYALTKKIKIKTWPKKGPGKKCLAAWSARTKIPLIPGQVQATNGLFDRIGSIYDGVEKKVTNRNANKKLEYDEAIKEGRNPAVPEYETAYNIDGTLINKPGYNPNLYITQSRTPRLITEADRPLVEKILWQMVEKKTQSRNQARRARLEKAAHLQGLPVPKFVPEKVDRSQKIEIRIIDPLDKIEPYMPQDRMAIKASQDGHVPYWQRPFLSKRRNRRVRAGWGKQVSSIQAWLTGALLVIVRLGNEAFLADIRGALRNAQWRKLLKPDATYQSLFNLFTGDPVVNTRTNHLTMAYREGVVNIVKSRSFKGRQTREHLLTLLGQGKTVAGVSFDLGQKHAAGLLAAHFGLGEDGNPVFTPIQACFLPQRYLDSLTNYRNRYDALTLDMRRQSLLALTPAQQQEFADAQRDPGGQAKRACCLKLNLNPDEIRWDLVSGISTMISDLYIERGGDPRDVHQQVETKPKGKRKSEIRILKIRDGKWAYDFRPKIADETRKAQREQLWKLQKASSEFERLSRYKINIARAIANWALQWGRELSGCDIVIPVLEDLNVGSKFFDGKGKWLLGWDNRFTPKKENRWFIKVLHKAVAELAPHRGVPVYEVMPHRTSMTCPACHYCHPTNREGDRFECQSCHVVKNTDRDVAPYNILRVAVEGKTLDRWQAEKKPQAEPDRPMILIDNQES*

[0402] The asterisk (*) in the above sequence represents the STOP codon. Alternatively, CasΦ-1 is also known as Cas12j ortholog 1. Thus, CasΦ-1 - CasΦ-10 can also be referred to as Cas12j orthologs 1 - 10, respectively.

[0403] The Cas9 nuclease has two functional endonuclease domains: RuvC and HNH. Upon target binding, Cas9 undergoes a conformational change that positions the nuclease domains to cleave opposite strands of the target DNA. The end result of Cas9-mediated DNA cleavage is a double-strand break (DSB) within the target DNA (approximately 3-4 nucleotides upstream of the PAM sequence). The resulting DSB is then repaired by one of two general repair pathways: (1) the highly efficient but error-prone non-homologous end joining (NHEJ) pathway; or (2) the less efficient but high-fidelity homologous directed repair (HDR) pathway.

[0404] The "efficiency" of non-homologous end joining (NHEJ) and / or homologous directed repair (HDR) can be calculated by any convenient method. For example, in some cases, efficiency can be expressed as the percentage of successful HDR. For example, the Surveyor nuclease assay can be used to generate cleavage products, and the ratio of products to substrate can be used to calculate the percentage. For example, the Surveyor nuclease can be used to directly cleave DNA containing a newly integrated restriction sequence as a result of successful HDR. More cleaved substrate indicates a higher percentage of HDR (higher HDR efficiency). As an illustrative example, the fraction (percentage) of HDR can be calculated using the following equation: [(cleavage products) / (substrate plus cleavage products)] (e.g., (b + c) / (a + b + c)), where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products.

[0405] In some cases, efficiency can be expressed as the percentage of successful NHEJ. For example, the T7 endonuclease I assay can be used to generate cleavage products, and the ratio of products to substrate can be used to calculate the NHEJ percentage. T7 endonuclease I cleaves mismatched heteroduplex DNA generated by hybridization of wild-type and mutant DNA strands (NHEJ generates small random insertions or deletions (indels) at the original break site). More cleavage indicates a higher percentage of NHEJ (higher NHEJ efficiency). As an illustrative example, the fraction (percentage) of NHEJ can be calculated using the following equation: (1 - (1 - (b + c) / (a + b + c)) 1 / 2 ) × 100, where "a" is the band intensity of the DNA substrate, and "b" and "c" are the cleavage products (Ran et al., Cell. 2013 Sep. 12; 154(6):1380-9; and Ran et al., Nat Protoc. 2013 Nov.; 8(11):2281–2308).

[0406] The NHEJ repair pathway is the most active repair mechanism and often results in small nucleotide insertions or deletions (indels) at the DSB site. The randomness of NHEJ-mediated DSB repair has important practical implications because a population of cells expressing Cas9 and gRNA or a guide polynucleotide will result in a variety of mutations. In most cases, NHEJ generates small indels in the target DNA, resulting in amino acid deletions, insertions, or frameshift mutations that lead to premature stop codons within the open reading frame (ORF) of the target gene. The ideal end result is a loss-of-function mutation within the target gene.

[0407] Although NHEJ-mediated DSB repair often disrupts the open reading frame of a gene, homology-directed repair (HDR) can be used to generate specific nucleotide changes, ranging from single nucleotide changes to large insertions such as the addition of fluorophores or tags.

[0408] To utilize HDR for gene editing, a DNA repair template containing the desired sequence can be delivered into the cell type of interest using a gRNA and Cas9 or Cas9 nickase. The repair template can contain the desired edit as well as additional homologous sequences immediately upstream and downstream of the target (referred to as left and right homology arms). The length of each homology arm depends on the size of the introduced change, with larger insertions requiring longer homology arms. The repair template can be a single-stranded oligonucleotide, double-stranded oligonucleotide, or double-stranded DNA plasmid. Even in cells expressing Cas9, gRNA, and an exogenous repair template, the efficiency of HDR is typically low (<10% modified alleles). The efficiency of HDR can be increased by synchronizing the cells because HDR occurs during the S and G2 phases of the cell cycle. Chemical or genetic inhibitors of genes involved in NHEJ can also increase the HDR frequency.

[0409] In some embodiments, Cas9 is a modified Cas9. A given gRNA target sequence can have additional sites throughout the genome where there is partial homology. These sites are called off-target sites and need to be considered when designing the gRNA. In addition to optimizing gRNA design, the specificity of CRISPR can also be improved by modifications to Cas9. Cas9 generates double-stranded breaks (DSBs) through the combined activity of two nuclease domains, RuvC and HNH. Cas9 nickase is a D10A mutant of SpCas9 that retains one nuclease domain and generates DNA nicks instead of DSBs. The nickase system can also be combined with HDR-mediated gene editing for specific gene editing.

[0410] In some cases, Cas9 is a variant Cas9 protein. The variant Cas9 polypeptide has an amino acid sequence that differs by one amino acid from the amino acid sequence of the wild-type Cas9 protein (e.g., has a deletion, insertion, substitution, fusion). In some cases, the variant Cas9 polypeptide has an amino acid change that reduces the nuclease activity of the Cas9 polypeptide (e.g., a deletion, insertion, or substitution). For example, in some cases, the variant Cas9 polypeptide has less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, less than 5%, or less than 1% of the nuclease activity of the corresponding wild-type Cas9 protein. In some cases, the variant Cas9 protein has no substantial nuclease activity. When the subject Cas9 protein is a variant Cas9 protein with no substantial nuclease activity, it can be referred to as "dCas9".

[0411] In some cases, the variant Cas9 protein has reduced nuclease activity. For example, the variant Cas9 protein exhibits less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% of the endonuclease activity of the wild-type Cas9 protein, e.g., the wild-type Cas9 protein.

[0412] In some cases, the variant Cas9 protein can cleave the complementary strand of the guide target sequence, but has a reduced ability to cleave the non-complementary strand of the double-stranded guide target sequence. For example, the variant Cas9 protein can have a mutation (amino acid substitution) that reduces the function of the RuvC domain. As a non-limiting example, in some embodiments, the variant Cas9 protein has D10A (aspartic acid to alanine at amino acid position 10) and can thus cleave the complementary strand of the double-stranded guide target sequence but not the complementary strand of the non-double-stranded guide target sequence (thus resulting in a single-strand break (SSB) rather than a double-strand break (DSB) when the variant Cas9 protein cleaves the double-stranded target nucleic acid) (see, e.g., Jinek et al., Science. 2012 Aug. 17; 337(6096):816-21).

[0413] In some cases, variant Cas9 proteins can cleave the non-complementary strand of a double-stranded guide target sequence, but have a reduced ability to cleave the complementary strand of the guide target sequence. For example, variant Cas9 proteins can have mutations (amino acid substitutions) that reduce the function of the HNH domain (RuvC / HNH / RuvC domain motif). As a non-limiting example, in some embodiments, the variant Cas9 protein has an H840A (histidine to alanine at amino acid position 840) mutation, and thus can cleave the non-complementary strand of the guide target sequence, but has a reduced ability to cleave the complementary strand of the guide target sequence (resulting in the use of SSB instead of DSB when the variant Cas9 protein cleaves a double-stranded guide target sequence). Such Cas9 proteins have a reduced ability to cleave a guide target sequence (e.g., a single-stranded guide target sequence), but retain the ability to bind to a guide target sequence (e.g., a single-stranded guide target sequence).

[0414] In some cases, the ability of variant Cas9 proteins to cleave the complementary and non-complementary strands of double-stranded target DNA is reduced. As a non-limiting example, in some cases, the variant Cas9 protein contains both D10A and H840A mutations, such that the ability of the polypeptide to cleave the complementary and non-complementary strands of double-stranded target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0415] As another non-limiting example, in some cases, the variant Cas9 protein contains W476A and W1126A mutations, such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0416] As another non-limiting example, in some cases, the variant Cas9 protein contains P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA).

[0417] As another non-limiting example, in some cases, the variant Cas9 protein contains H840A, W476A, and W1126A mutations, such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, the variant Cas9 protein contains H840A, D10A, W476A, and W1126A mutations, such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some embodiments, the variant Cas9 restores the catalytic His residue at position 840 in the Cas9 HNH domain (A840H).

[0418] As another non-limiting example, in some cases, the variant Cas9 protein contains the H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, the variant Cas9 protein contains the D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, such that the ability of the polypeptide to cleave target DNA is reduced. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA), but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some cases, when the variant Cas9 protein contains the W476A and W1126A mutations or when the variant Cas9 protein contains the P475A, W476A, N477A, D1125A, W1126A, and D1127A mutations, the variant Cas9 protein does not bind effectively to the PAM sequence. Thus, in some such cases, when such variant Cas9 proteins are used in a binding method, the method does not require a PAM sequence. In other words, in some cases, when such variant Cas9 proteins are used in a binding method, the method can include a guide RNA, but the method can proceed in the absence of a PAM sequence (and the specificity of binding is thus provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effect (i.e., inactivate one or the other nuclease moiety). As non-limiting examples, the residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., substituted). In addition, mutations other than alanine substitutions are suitable.

[0419] In some embodiments, a variant Cas9 protein with reduced catalytic activity (e.g., when the Cas9 protein has the D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 mutations, e.g., D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A), the variant Cas9 protein can still bind to target DNA in a site-specific manner (because it is still guided to the target DNA sequence by the guide RNA), as long as it retains the ability to interact with the guide RNA.

[0420] In some embodiments, the variant Cas protein can be spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK or spCas9-LRVSQL.

[0421] In some embodiments, a modified SpCas9 is used that includes the amino acid substitutions D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (SpCas9-MQKFRAER) and is specific for the altered PAM 5'-NGC.

[0422] Alternatives to Streptococcus pyogenes Cas9 can include RNA-guided endonucleases from the Cpf1 family, which show cleavage activity in mammalian cells. CRISPR from Prevotella and Francisella novicida 1 (CRISPR / Cpf1) is a DNA editing technology similar to the CRISPR / Cas9 system. Cpf1 is an RNA-guided endonuclease of the type II CRISPR / Cas system. This acquired immune mechanism exists in Prevotella and Francisella novicida. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses guide RNA to find and cut viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9, overcoming some limitations of the CRISPR / Cas9 system. Different from Cas9 nuclease, the result of Cpf1-mediated DNA cleavage is a double-strand break with short 3' overhangs. The staggered cleavage pattern of Cpf1 can open up the possibility of directed gene transfer, similar to traditional restriction enzyme cloning, which can improve the efficiency of gene editing. Like the above-mentioned Cas9 variants and orthologs, Cpf1 can also expand the number of CRISPR-targetable sites to AT-rich regions or AT-rich genomes that lack the NGG PAM sites favored by SpCas9. The Cpf1 locus contains a mixed alpha / beta domain, a RuvC-I followed by a helical region, a RuvC-II, and a zinc finger-like domain. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. In addition, Cpf1 does not have an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain architecture indicates that Cpf1 is functionally unique and is classified as a type 2 class V CRISPR system. The Cas1, Cas2, and Cas4 proteins encoded by the Cpf1 locus are more similar to type I and type III than those from type II systems. Functional Cpf1 does not require trans-activating CRISPR RNA (tracrRNA), and thus, only requires CRISPR (crRNA). This is beneficial for genome editing because Cpf1 is not only smaller than Cas9, but its sgRNA molecule is also smaller

[0423] (Approximately half the nucleotides of Cas9). Compared with the G-rich PAM targeted by Cas9, the Cpf1-crRNA complex cuts the target DNA or RNA by recognizing the protospacer adjacent motif 5'-YTN-3'. After identifying the PAM, Cpf1 introduces a sticky-end-like DNA double-strand break with 4 or 5 nucleotide overhangs.

[0424] Some aspects of the present disclosure provide fusion proteins that include domains that act as nucleic acid programmable DNA binding proteins that can be used to direct proteins, such as base editors, to specific nucleic acid (e.g., DNA or RNA) sequences. In certain embodiments, the fusion protein includes a nucleic acid programmable DNA binding protein domain and a deaminase domain. DNA binding proteins include, but are not limited to, Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ. An example of a programmable polynucleotide binding protein with a PAM specificity different from Cas9 is the clustered regularly interspaced short palindromic repeats from Prevotella and Francisella 1 (Cpf1). Similar to Cas9, Cpf1 is also a type II CRISPR effector. It has been shown that Cpf1 mediates potent DNA interference, which is characterized differently from Cas9. Cpf1 is a single RNA-guided endonuclease that lacks tracrRNA and it utilizes a T-rich protospacer adjacent motif (TTN, TTTN, or YTN). In addition, Cpf1 cleaves DNA by staggered DNA double-strand breaks. Among 16 Cpf1 family proteins, two enzymes from Acidaminococcus and Lachnospiraceae have been shown to have efficient genome editing activity in human cells. Cpf1 proteins are known in the art and have been previously described, e.g., Yamano et al., “Crystal structure of Cpf1 in complex with guide RNA and target DNA.” Cell (165) 2016, p. 949-962; the entire content of which is incorporated herein by reference.

[0425] Useful in the present compositions and methods are nuclease-inactivated Cpf1 (dCpf1) variants, which can be used as programmable polynucleotide-binding protein domains for guide nucleotide sequences. The Cpf1 protein has a RuvC-like nuclease domain that is similar to the RuvC domain of Cas9 but does not have an HNH endonuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition lobe of Cas9. Zetsche et al., Cell, 163, 759-771, 2015 (incorporated herein by reference) showed that the RuvC-like domain of Cpf1 is responsible for cleaving both DNA strands and inactivating the RuvC-like domain inactivates Cpf1 nuclease activity. For example, mutations corresponding to D917A, E1006A, or D1255A in Francisella novicida Cpf1 inactivate Cpf1 nuclease activity. In some embodiments, the dCpf1 of the present disclosure contains mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D122. It should be understood that any mutation that inactivates the RuvC domain of Cpf1 can be used according to the present disclosure, such as substitution mutations, deletions, or insertions.

[0426] In some embodiments, the nucleic acid programmable nucleotide-binding protein or any fusion protein provided herein can be a Cpf1 protein. In some embodiments, the Cpf1 protein is a Cpf1 nickase (nCpf1). In some embodiments, the Cpf1 protein is nuclease-inactivated Cpf1 (dCpf1). In some embodiments, the Cpf1, the nCpf1, or the dCpf1 contains an amino acid sequence that is at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any Cpf1 protein described herein. In some embodiments, the dCpfl contains at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cpf1 sequence disclosed herein and contains mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A. It should be understood that Cpf1 from other bacterial species can also be used according to the present disclosure.

[0427] The amino acid sequence of wild-type Francisella novicida Cpf1 is as follows. D917, E1006, and D1255 are shown in bold and underlined.

[0428]

[0429]

[0430] The amino acid sequence of Francisella novicida Cpf1 D917A is as follows. (A917, E1006, and D1255 are shown in bold and underlined).

[0431]

[0432] The amino acid sequence of Francisella novicida Cpf1 E1006A is as follows. (D917, A1006, and D1255 are shown in bold and underlined).

[0433]

[0434] The amino acid sequence of Francisella novicida Cpf1 D1255A is as follows. (The mutation positions of D917, E1006, and A1255 are shown in bold and underlined).

[0435]

[0436]

[0437] The amino acid sequence of Francisella novicida Cpf1 D917A / E1006A is as follows. (A917, A1006, and D1255 are shown in bold and underlined).

[0438]

[0439]

[0440] The amino acid sequence of Francisella novicida Cpf1 D917A / D1255A is as follows. (A917, E1006, and A1255 are shown in bold and underlined).

[0441]

[0442] The amino acid sequence of Francisella novicida Cpf1 E1006A / D1255A is as follows. (D917, A1006, and A1255 are shown in bold and underlined).

[0443]

[0444]

[0445] The amino acid sequence of Francisella novicida Cpf1 D917A / E1006A / D1255A is as follows. (A917, A1006, and A1255 are shown in bold and underlined).

[0446]

[0447]

[0448] In some embodiments, one of the Cas9 domains present in the fusion protein can be replaced with a DNA-binding protein domain programmable by a guide nucleotide sequence that has no requirement for a PAM sequence.

[0449] In some embodiments, the Cas domain is the Cas9 domain from Staphylococcus aureus (SaCas9). In some embodiments, the SaCas9 domain is nuclease-active SaCas9, nuclease-inactivated SaCas9 (SaCas9d), or SaCas9 nickase (SaCas9n). In some embodiments, the SaCas9 domain contains the N579A mutation, or a corresponding mutation in any of the amino acid sequences provided herein.

[0450] In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having a non-canonical PAM. In some embodiments, the SaCas9 domain, the SaCas9d domain, or the SaCas9n domain can bind to a nucleic acid sequence having an NNGRRT or NNGRRT PAM sequence. In some embodiments, the SaCas9 domain comprises one or more of the E781X, N967X, and R1014X mutations, or the corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SaCas9 domain comprises one or more of the E781K, N967K, and R1014H mutations, or one or more corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SaCas9 domain comprises the E781K, N967K, or R1014H mutation, or the corresponding mutation in any of the amino acid sequences provided herein.

[0451] The amino acid sequence of exemplary SaCas9 is as follows:

[0452] In this sequence, residue N579 (underlined and bolded) can be mutated (e.g., to A579) to generate SaCas9 nickase.

[0453] The amino acid sequence of exemplary SaCas9n is as follows:

[0454]

[0455] In this sequence, residue A579 can be mutated from N579 to generate SaCas9 nickase, shown underlined and bolded.

[0456] The amino acid sequence of exemplary SaKKH Cas9 is as follows:

[0457]

[0458]

[0459] The above residue A579 can be mutated from N579 to generate SaCas9 nickase, shown underlined and bolded. The above residues K781, K967, and H1014 can be mutated from E781, N967, and R1014 to generate SaKKH Cas9, shown underlined and italicized.

[0460] High-fidelity Cas9 domain

[0461] Some aspects of the present disclosure provide high-fidelity Cas9 domains. In some embodiments, the high-fidelity Cas9 domain is an engineered Cas9 domain that includes one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA, relative to the corresponding wild-type Cas9 domain. A high-fidelity Cas9 domain with reduced electrostatic interaction with the sugar-phosphate backbone of DNA can have fewer off-target effects. In some embodiments, the Cas9 domain (e.g., wild-type Cas9 domain) includes one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA. In some embodiments, the Cas9 domain includes one or more mutations that reduce the association between the Cas9 domain and the sugar-phosphate backbone of DNA by at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, or at least 70%.

[0462] In some embodiments, any Cas9 fusion protein provided herein includes one or more N497X, R661X, Q695X, and / or Q926X mutations, or corresponding mutations in any amino acid sequence provided herein, where X is any amino acid. In some embodiments, any Cas9 fusion protein provided herein includes one or more N497A, R661A, Q695A, and / or Q926A mutations, or corresponding mutations in any amino acid sequence provided herein. In some embodiments, the Cas9 domain includes the D10A mutation, or a corresponding mutation in any amino acid sequence provided herein. High-fidelity Cas9 domains are known in the art and will be apparent to those skilled in the art. For example, high-fidelity Cas9 domains have been described in Kleinstiver, B.P. et al., “High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects.” Nature 529, 490-495 (2016); and Slaymaker, I.M., et al., “Rationally engineered Cas9 nucleases with improved specificity.” Science 351, 84-88 (2015); the entire contents of which are incorporated by reference.

[0463] In some embodiments, the modified Cas9 is a high-fidelity Cas9 enzyme. In some embodiments, the high-fidelity Cas9 enzyme is SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, or the hyper-accurate Cas9 variant (HypaCas9). The modified Cas9 eSpCas9(1.1) contains alanine substitutions that weaken the interaction between the HNH / RuvC groove and the non-target DNA strand, preventing strand separation and cleavage at off-target sites. Similarly, SpCas9-HF1 reduces off-target editing through alanine substitutions that disrupt the interaction of Cas9 with the DNA phosphate backbone. HypaCas9 contains mutations in the REC3 domain (SpCas9N692A / M694A / Q695A / H698A) that increase Cas9 proofreading and target discrimination. All three high-fidelity enzymes produce less off-target editing compared to wild-type Cas9.

[0464] Exemplary high-fidelity Cas9s are provided below.

[0465] High-fidelity Cas9 domain mutations relative to Cas9 are shown in bold and underlined.

[0466]

[0467]

[0468] Guide polynucleotide

[0469] In one embodiment, the guide polynucleotide is a guide RNA. The RNA / Cas complex can help “guide” the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets complementary to the spacer. The target strand that is not complementary to the crRNA is first endonucleolytically cleaved and then trimmed 3'-5' exonucleolytically. In nature, DNA binding and cleavage generally require a protein and two RNAs. However, single guide RNAs (“sgRNAs,” or simply “gRNAs”) can be engineered to incorporate aspects of crRNAs and tracrRNAs into a single RNA species. See, e.g., Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes a short motif within the CRISPR repeat sequence (PAM or protospacer adjacent motif) to help distinguish self from non-self. Cas9 nuclease sequences and structures are well known to those of skill in the art (see, e.g., “Complete genome sequence of an M1 strain of Streptococcus pyogenes.” Ferretti, J.J. et al., Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III.” Deltcheva E. et al., Nature 471:602-607 (2011); and

[0470] "Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M. et al., Science 337:816-821 (2012), the entire content of which is incorporated herein by reference). Cas9 orthologs have been described in various species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus. Based on the present disclosure, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun, and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737; the entire content of which is incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase.

[0471] In some embodiments, the guide polynucleotide is at least one single guide RNA

[0472] ("sgRNA" or "gNRA"). In some embodiments, the guide polynucleotide is at least one tracrRNA. In some embodiments, the guide polynucleotide does not require a PAM sequence to direct the polynucleotide-programmable DNA binding domain (e.g., Cas9 or Cpf1) to the target nucleotide sequence.

[0473] The polynucleotide programmable nucleotide binding domain (e.g., CRISPR-derived domain) of the base editors disclosed herein can recognize a target polynucleotide sequence by association with a guide polynucleotide. The guide polynucleotide (e.g., gRNA) is typically single-stranded and can be programmed to site-specifically bind (i.e., by complementary base pairing) to the target sequence of a polynucleotide, thereby directing the base editor that binds to the guide nucleic acid to the target sequence. The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. In some cases, the guide polynucleotide comprises natural nucleotides (e.g., adenosine). In some cases, the guide polynucleotide comprises non-natural (or unnatural) nucleotides (e.g., peptide nucleic acids or nucleotide analogs). In some cases, the length of the targeting region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The length of the targeting region of the guide nucleic acid can be between 10 - 30 nucleotides, or between 15 - 25 nucleotides, or between 15 - 20 nucleotides.

[0474] In some embodiments, the guide polynucleotide comprises two or more separate polynucleotides that can interact with each other, e.g., by complementary base pairing (e.g., dual guide polynucleotide). For example, the guide polynucleotide can comprise a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA). For example, the guide polynucleotide can comprise one or more trans-activating CRISPR RNAs (tracrRNAs).

[0475] In type II CRISPR systems, targeting of nucleic acids by a CRISPR protein (e.g., Cas9) typically requires complementary base pairing repeats between a first RNA molecule (crRNA) containing a sequence that recognizes the target sequence and a second RNA molecule (trRNA) containing a sequence that recognizes the target sequence to form a scaffold region that stabilizes the guide RNA-CRISPR protein complex. Such dual guide RNA systems can be used as guide polynucleotides to direct the base editors disclosed herein to a target polynucleotide sequence.

[0476] In some embodiments, the base editors provided herein utilize a single guide polynucleotide (e.g., gRNA). In some embodiments, the base editors provided herein utilize a dual guide polynucleotide (e.g., dual gRNA). In some embodiments, the base editors provided herein utilize one or more guide polynucleotides (e.g., multiplex gRNA). In some embodiments, a single guide polynucleotide is used for different base editors described herein. For example, a single guide polynucleotide can be used for a cytidine base editor and an adenosine base editor.

[0477] In other embodiments, the guide polynucleotide can comprise a polynucleotide targeting portion of a nucleic acid and a scaffold portion of a nucleic acid in a single molecule (i.e., a single-molecule guide nucleic acid). For example, the single-molecule guide polynucleotide can be a single guide RNA (sgRNA or gRNA). As used herein, the term guide polynucleotide sequence encompasses any single-, double-, or multi-molecule nucleic acid capable of interacting with a base editor and guiding the base editor to a target polynucleotide sequence.

[0478] Typically, a guide polynucleotide (e.g., a crRNA / trRNA complex or gRNA) comprises a "polynucleotide targeting segment" that includes a sequence capable of recognizing and binding to a target polynucleotide sequence, and a "protein-binding segment" that anchors the guide polynucleotide within the polynucleotide programmable nucleotide-binding domain assembly of the base editor. In some embodiments, the polynucleotide targeting segment of the guide polynucleotide recognizes and binds to a DNA polynucleotide, thereby facilitating editing of a base in the DNA. In other instances, the polynucleotide targeting segment of the guide polynucleotide recognizes and binds to an RNA polynucleotide, thereby facilitating editing of a base in the RNA. As used herein, a "segment" refers to a portion or region of a molecule, e.g., a contiguous stretch of nucleotides in a guide polynucleotide. A segment can also refer to a region / segment of a complex such that the segment can comprise regions of more than one molecule. For example, when a guide polynucleotide comprises multiple nucleic acid molecules, the protein-binding segment can comprise, e.g., all or a portion of multiple individual molecules hybridized along complementary regions. In some embodiments, the protein-binding segment of an RNA targeting DNA that comprises two separate molecules can comprise (i) 40-75 base pairs of a first RNA molecule that is 100 base pairs in length; (ii) 10-25 base pairs of a second RNA molecule that is 50 base pairs in length. Unless otherwise specifically defined in a particular context, the definition of "segment" is not limited to a particular total number of base pairs, not limited to any particular number of base pairs from a given RNA molecule, not limited to the number of separate molecules in a particular complex, and can include regions of RNA molecules of any total length and can include regions complementary to other molecules.

[0479] The guide RNA or guide polynucleotide can comprise two or more RNAs, such as CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA). The guide RNA or guide polynucleotide can sometimes comprise a single-stranded RNA or a single guide RNA (sgRNA) formed by fusion of a portion (e.g., a functional portion) of the crRNA and tracrRNA. The guide RNA or guide polynucleotide can also be a dual RNA comprising the crRNA and tracrRNA. Additionally, the crRNA can hybridize to the target DNA.

[0480] As described above, the guide RNA or guide polynucleotide can be an expression product. For example, the DNA encoding the guide RNA can be a vector containing the sequence encoding the guide RNA. By transfecting cells with an isolated guide RNA or plasmid DNA containing the sequence encoding the guide RNA and a promoter, the guide RNA or guide polynucleotide can be transferred into the cells. The guide RNA or guide polynucleotide can also be transferred into the cells in other ways, such as using virus-mediated gene delivery.

[0481] The guide RNA or guide polynucleotide can be isolated. For example, the guide RNA can be transfected into cells or organisms in the form of isolated RNA. The guide RNA can be prepared by in vitro transcription using any in vitro transcription system known in the art. The guide RNA can be transferred into cells in the form of isolated RNA rather than in the form of a plasmid containing the coding sequence of the guide RNA.

[0482] The guide RNA or guide polynucleotide can comprise three regions: a first region at the 5' end can be complementary to a target site in the chromosomal sequence, a second internal region can form a stem-loop structure, and a third 3' region can be single-stranded. The first region of each guide RNA can also be different such that each guide RNA directs the fusion protein to a specific target site. In addition, the second and third regions of each guide RNA can be the same in all guide RNAs.

[0483] The first region of the guide RNA or guide polynucleotide can be complementary to the sequence of the target site in the chromosomal sequence such that the first region of the guide RNA can base pair with the target site. In some cases, the first region of the guide RNA can comprise or be from about 10 nucleotides to 25 nucleotides (i.e., from 10 nucleotides to nucleotides; or from about 10 nucleotides to about 25 nucleotides; or from 10 nucleotides to about 25 nucleotides; or from about 10 nucleotides to 25 nucleotides) or more. For example, the base pairing region between the first region of the guide RNA and the target site in the chromosomal sequence can be or can be about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25 or more nucleotides in length. Sometimes, the length of the first region of the guide RNA can be or can be about 19, 20 or 21 nucleotides.

[0484] The guide RNA or guide polynucleotide may also comprise a second region that forms a secondary structure. For example, the secondary structure formed by the guide RNA may comprise a stem (or hairpin) and a loop. The lengths of the loop and the stem can vary. For example, the length of the loop can range from about 3 to 10 nucleotides, while the length of the stem can range from about 6 to 20 base pairs. The stem may comprise one or more bulges of 1 to 10 or about 10 nucleotides. The total length of the second region can range from about 16 to 60 nucleotides. For example, the length of the loop can be or can be about 4 nucleotides, and the stem can be or can be about 12 base pairs.

[0485] The guide RNA or guide polynucleotide may also comprise a third region at the 3' end that is substantially single-stranded. For example, the third region is sometimes not complementary to any chromosomal sequence in the target cell and sometimes not complementary to the rest of the guide RNA. In addition, the length of the third region can vary. The length of the third region can be more than or more than about 4 nucleotides. For example, the total length of the third region can range from about 5 to 60 nucleotides.

[0486] The guide RNA or guide polynucleotide can target any exon or intron of a gene target. In some cases, the guide can target exon 1 or 2 of a gene, and in other cases; the guide can target exon 3 or 4 of a gene. The composition can comprise multiple guide RNAs that all target the same exon, or in some cases, can comprise multiple guide RNAs that target different exons. Exons and introns of a gene can be targeted.

[0487] The guide RNA or guide polynucleotide can target a nucleic acid sequence of about 20 nucleotides or about 20 nucleotides. The target nucleic acid can be less than or less than about 20 nucleotides. The length of the target nucleic acid can be at least or at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or anywhere between 1 - 100 nucleotides. The length of the target nucleic acid can be at most or at most about 5, 10, 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, or anywhere between 1 - 100 nucleotides. The target nucleic acid sequence can be or can be about 20 bases at the 5' of the first nucleotide adjacent to the PAM. The guide RNA can target a nucleic acid sequence. The length of the target nucleic acid can be at least or at least about 1 - 10, 1 - 20, 1 - 30, 1 - 40, 1 - 50, 1 - 60, 1 - 70, 1 - 80, 1 - 90, or 1 - 100 nucleotides.

[0488] Guide polynucleotides, such as guide RNAs, can refer to nucleic acids that can hybridize to another nucleic acid, such as a target nucleic acid or protospacer in a cellular genome. The guide polynucleotide can be RNA. The guide polynucleotide can be DNA. The guide polynucleotide can be programmed or designed to bind site-specifically to a nucleic acid sequence. The guide polynucleotide can comprise a polynucleotide chain and can be referred to as a single guide polynucleotide. The guide polynucleotide can comprise two polynucleotide chains and can be referred to as a dual guide polynucleotide. The guide RNA can be introduced into a cell as an RNA molecule. For example, the RNA molecule can be transcribed in vitro and / or can be chemically synthesized. The RNA can be transcribed from a synthetic DNA molecule, such as a gene fragment. The guide RNA can then be introduced into the cell as an RNA molecule. The guide RNA can also be introduced into the cell in the form of a non-RNA nucleic acid molecule, such as a DNA molecule. For example, the DNA encoding the guide RNA can be operably linked to a promoter control sequence to express the guide RNA in a cell of interest. The RNA coding sequence can be operably linked to a promoter sequence recognized by RNA polymerase III (Pol III). Plasmid vectors that can be used to express the guide RNA include, but are not limited to, the px330 vector and the px333 vector. In some cases, a plasmid vector (e.g., the px333 vector) can comprise at least two DNA sequences encoding the guide RNA.

[0489] Methods for selecting, designing, and validating guide polynucleotides, such as guide RNAs and targeting sequences, are described herein and are known to those of skill in the art. For example, to minimize the effects of potential substrate promiscuity of deaminase domains in a base editor system (e.g., an AID domain), the number of residues that might inadvertently become deamination targets (e.g., off-target C residues that might be present on ssDNA within a target nucleic acid locus) can be minimized. In addition, software tools can be used to optimize the gRNA corresponding to a target nucleic acid sequence, e.g., to minimize the total off-target activity across the genome. For example, for each possible targeting domain selection using Streptococcus pyogenes Cas9, all off-target sequences (preceded by the selected PAM, e.g., NAG or NGG) can be identified in the genome that contain up to a specific number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) of mismatched base pairs. The first region of the gRNA that is complementary to the target site can be identified, and all first regions (e.g., crRNAs) can be ranked according to their total predicted off-target score; the top-ranked targeting domains represent those that are likely to have the greatest on-target and least off-target activity. Candidate targeting gRNAs can be functionally evaluated using methods known in the art and / or as described herein.

[0490] As a non-limiting example, a DNA sequence search algorithm can be used to identify the target DNA hybridization sequences in the crRNA of the guide RNA to be used with Cas9. gRNA design can be performed using custom gRNA design software based on the public tool cas-offinder, such as Bae S., Park J., and Kim J.-S. Cas-OFFinder: A fast and versatile algorithm that searches for potential off-target sites of Cas9 RNA-guided dendonucleases. Bioinformatics 30, 1473-1475 (2014). The software scores the guides after calculating the genome-wide off-target propensity. Typically, for guides ranging in length from 17 to 24, matches from perfect match to 7 mismatches are considered. Once the off-target sites are determined computationally, a total score is calculated for each guide and summarized in a tabular output using a web interface. In addition to identifying potential target sites adjacent to the PAM sequence, the software also identifies all PAM-adjacent sequences that differ from the selected target site by 1, 2, 3, or more than 3 nucleotides. The target nucleic acid sequence, such as the genomic DNA sequence of the target gene, can be obtained and repetitive elements can be screened using publicly available tools, such as the RepeatMasker program. RepeatMasker searches for repetitive elements and low complexity regions in the input DNA sequence. The output is a detailed annotation of the repeats present in the given query sequence.

[0491] After identification, the first regions of the guide RNA, such as the crRNA, can be ranked according to their distance from the target site, their orthogonality, and the presence of 5' nucleotides, in order to pair with the relevant PAM sequence (e.g., 5'G based on the identification of close matches containing the relevant PAM in the human genome, such as the NGG PAM of Streptococcus pyogenes, the NNGRRT or NNGRRV PAM of Staphylococcus aureus). As used herein, orthogonality refers to the number of sequences in the human genome that contain the minimum number of mismatches with the target sequence. For example, "high-level orthogonality" or "good orthogonality" can refer to a 20-mer targeting domain that has no identical sequences in the human genome other than the expected target, nor any sequence order that contains one or two mismatches in the target. Targeting domains with good orthogonality can be selected to minimize off-target DNA cleavage.

[0492] In some embodiments, a reporting system can be used to detect base editing activity and test candidate guide polynucleotides. In some embodiments, the reporting system can include a reporter gene-based assay, wherein base editing activity results in the expression of the reporter gene. For example, the reporting system can include a reporter gene containing an inactivated start codon, such as a mutation from 3'-TAC-5' to 3'-CAC-5' on the template strand. After successful deamination of the target C, the corresponding mRNA will be transcribed as 5'-AUG-3' instead of 5'-GUG-3', enabling translation of the reporter gene. Suitable reporter genes will be apparent to those skilled in the art. Non-limiting examples of reporter genes include genes encoding green fluorescent protein (GFP), red fluorescent protein (RFP), luciferase, secreted alkaline phosphatase (SEAP), or any other gene whose expression is detectable and apparent to those skilled in the art. The reporting system can be used to test many different gRNAs, e.g., to determine which residues of a target DNA sequence will be targeted by the corresponding deaminase. sgRNAs targeting the non-template strand can also be tested to evaluate off-target effects of a particular base editing protein (such as a Cas9 deaminase fusion protein). In some embodiments, such gRNAs can be designed such that the mutated start codon does not base pair with the gRNA. The guide polynucleotide can comprise standard ribonucleotides, modified ribonucleotides (such as pseudouridine), ribonucleotide homologs, and / or ribonucleotide analogs. In some embodiments, the guide polynucleotide can comprise at least one detectable label. The detectable label can be a fluorophore (e.g., FAM, TMR, Cy3, Cy5, Texas Red, Oregon Green, Alexa Fluors, HaloTag, or a suitable fluorescent dye), a detection tag (e.g., biotin, digoxin, etc.), a quantum dot, or a gold particle.

[0493] The guide polynucleotide can be chemically synthesized, enzymatically synthesized, or a combination thereof. For example, guide RNA can be synthesized using standard phosphoramidite-based solid-phase synthesis methods. Alternatively, guide RNA can be synthesized in vitro by operably linking a DNA encoding the guide RNA to a promoter control sequence recognized by a phage RNA polymerase. Examples of suitable phage promoter sequences include T7, T3, SP6 promoter sequences, or variants thereof. In embodiments where the guide RNA comprises two separate molecules (e.g., crRNA and tracrRNA), the crRNA can be chemically synthesized and the tracrRNA can be enzymatically synthesized.

[0494] In some embodiments, the base editor system can comprise multiple guide polynucleotides, such as gRNAs. For example, the gRNAs can target one or more target loci (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs) included in the base editor system. The multiple gRNA sequences can be arranged in tandem and preferably separated by direct repeats.

[0495] The DNA sequence encoding the guide RNA or guide polynucleotide can also be part of a vector. Additionally, the vector can contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences, etc.), selection marker sequences (e.g., GFP or antibiotic resistance genes, such as puromycin), origins of replication, and the like. The DNA molecule encoding the guide RNA can also be linear. The DNA molecule encoding the guide RNA or guide polynucleotide can also be circular.

[0496] In some embodiments, one or more components of the base editor system can be encoded by DNA sequences. Such DNA sequences can be introduced into an expression system, such as a cell, together or individually. For example, the DNA sequences encoding the polynucleotide programmable nucleotide-binding domain and the guide RNA can be introduced into a cell, each DNA sequence can be part of a separate molecule (e.g., one vector containing the polynucleotide programmable nucleotide-binding domain encoding sequence and a second vector containing the guide RNA encoding sequence) or both can be part of the same molecule (e.g., a vector containing the encoding (and regulatory) sequences for the polynucleotide programmable nucleotide-binding domain and the guide RNA).

[0497] The guide polynucleotide can comprise one or more modifications to provide a nucleic acid with new or enhanced characteristics. The guide polynucleotide can comprise a nucleic acid affinity tag. The guide polynucleotide can include synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.

[0498] In some cases, the gRNA or guide polynucleotide can comprise modifications. Modifications can be made at any position of the gRNA or guide polynucleotide. More than one modification can be made to a single gRNA or guide polynucleotide. The gRNA or guide polynucleotide can be subject to quality control after modification. In some cases, quality control can include PAGE, HPLC, MS, or any combination thereof.

[0499] The modifications to the gRNA or guide polynucleotide can be substitutions, insertions, deletions, chemical modifications, physical modifications, stabilization, purification, or any combination thereof.

[0500] The gRNA or guide polynucleotide can also be 5'-adenylate, 5'-guanosine-triphosphate cap, 5'-N7-methylguanosine-triphosphate cap, 5'-triphosphate cap, 3'-phosphate, 3'-thiophosphate, 5'-phosphate, 5'-modified thiophosphate ester, cis-Syn thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, dSpacer, PC spacer, rSpacer, spacer 18, spacer 9, 3'-3' modification, 5'-5' modification, abasic, acridine, azobenzene, biotin, biotin-BB, biotin-TEG, cholesterol-TEG, desthiobiotin-TEG, DNP-TEG, DNP-X, DOTA, dT-biotin, dibiotin, PC-biotin, psoralen C2, psoralen C6, TINA, 3'-DABCYL, black hole quencher 1, black hole quencher 2, DABCYL-SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analog purine, 2'-deoxyribonucleoside analog pyrimidine, ribonucleoside analog, 2'-O-methyl ribonucleoside analog, sugar-modified analog, wobble / universal base, fluorescent dye label, 2'-fluoro RNA, 2'-O-methyl RNA, methylphosphonate, phosphodiester DNA, phosphorothioate DNA, phosphorothioate RNA, UNA, pseudouridine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate or any combination thereof.

[0501] In some cases, the modification is permanent. In other cases, the modification is temporary. In some cases, multiple modifications are made to the gRNA or guide polynucleotide. The gRNA or guide polynucleotide modification can alter the physicochemical properties of the nucleotides, such as their conformation, polarity, hydrophobicity, chemical reactivity, base pairing interactions or any combination thereof.

[0502] The PAM sequence can be any PAM sequence known in the art. Suitable PAM sequences include but are not limited to NGG, NGA, NGC, NGN, NGT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TATV, NNNNGATT, NNAGAAW or NAAAAC. Y is a pyrimidine; N is any nucleobase; W is A or T.

[0503] The modification can also be a phosphorothioate substitute. In some cases, natural phosphodiester bonds can be readily degraded rapidly by cellular nucleases; modification of the internucleotide bond with a phosphorothioate (PS) bond substitute can be more stable to hydrolysis by cellular degradation. The modification can increase the stability of the gRNA or guiding polynucleotide. The modification can also enhance biological activity. In some cases, phosphorothioate-enhanced RNA gRNA can inhibit RNase A, RNase T1, calf serum nuclease, or any combination thereof. These properties can enable PS-RNA gRNA to be used in applications with a high likelihood of exposure to nucleases in vivo or in vitro. For example, phosphorothioate (PS) bonds can be introduced between the last 3-5 nucleotides at the 5'- or 3'-end of the gRNA, which can inhibit exonuclease degradation. In some cases, phosphorothioate bonds can be added throughout the gRNA to reduce endonuclease attack.

[0504] Protospacer adjacent motif

[0505] The term "protospacer adjacent motif (PAM)" or PAM-like motif refers to a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. In some embodiments, the PAM can be a 5' PAM (i.e., upstream of the 5' end of the protospacer). In some embodiments, the PAM can be a 3' PAM

[0506] (i.e., downstream of the 5' end of the protospacer).

[0507] The PAM sequence is crucial for target binding, but the exact sequence depends on the type of Cas protein.

[0508] The base editors provided herein may include CRISPR protein-derived domains that are capable of binding to nucleotide sequences containing canonical or non-canonical protospacer adjacent motif (PAM) sequences. The PAM site is a nucleotide sequence near the target polynucleotide sequence. Some aspects of the present disclosure provide base editors that include all or part of a CRISPR protein with different PAM specificities. For example, a typical Cas9 protein, such as Cas9 from Streptococcus pyogenes (spCas9), requires a canonical NGG PAM sequence to bind to a specific nucleic acid region, where "N" in "NGG" is adenine (A), thymine (T), guanine (G), or cytosine (C), and G is guanine. The PAM can be CRISPR protein-specific and can vary between different base editors that include different CRISPR protein-derived domains. The PAM can be 5' or 3' of the target sequence. The PAM can be upstream or downstream of the target sequence. The length of the PAM can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides. Generally, the length of the PAM is between 2-6 nucleotides. Several PAM variants are described in Table 1 below.

[0509] Table 1. Cas9 proteins and corresponding PAM sequences

[0510] Variant PAM spCas9 NGG spCas9-VRQR NGA spCas9-VRER NGCG spCas9-MQKFRAER NGC xCas9(sp) NGN saCas9 NNGRRT saCas9-KKH NNNRRT spCas9-MQKSER NGCG spCas9-MQKSER NGCN spCas9-LRKIQK NGTN spCas9-LRVSQK NGTN spCas9-LRVSQL NGTN SpyMacCas9 NAA Cpf1 5’(TTTV)

[0511] In some embodiments, the PAM is NGC. In some embodiments, the NGC PAM is recognized by a Cas9 variant. In some embodiments, the NGC PAM variant includes one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (collectively referred to as "MQKFRAER").

[0512] In some embodiments, the PAM is NGT. In some embodiments, the NGT PAM is recognized by a Cas9 variant. In some embodiments, the NGT PAM variant is generated by targeted mutagenesis at one or more of residues 1335, 1337, 1135, 1136, 1218, and / or 1219. In some embodiments, the NGT PAM variant is generated by targeted mutagenesis at one or more of residues 1219,

[0513] 1335, 1337, 1218. In some embodiments, the NGT PAM variant is generated by targeted mutagenesis at one or more of residues 1135, 1136, 1218, 1219, and 1335. In some embodiments, the NGT PAM variant is selected from a group of targeted mutations provided in Tables 2 and 3 below.

[0514] Table 2: NGT PAM variant mutations at residues 1219, 1335, 1337, 1218

[0515] Variant E1219V R1335Q T1337 G1218 1 F V T 2 F V R 3 F V Q 4 F V L 5 F V T R 6 F V R R 7 F V Q R 8 F V L R 9 L L T 10 L L R 11 L L Q 12 L L L 13 F I T 14 F I R 15 F I Q 16 F I L 17 F G C 18 H L N 19 F G C A 20 H L N V 21 L A W 22 L A F 23 L A Y 24 I A W 25 I A F 26 I A Y

[0516] Table 3: NGT PAM variant mutations at residues 1135, 1136, 1218, 1219 and 1335

[0517] Variant D1135L S1136R G1218S E1219V R1335Q 27 G 28 V 29 I 30 A 31 W 32 H 33 K 34 K 35 R 36 Q 37 T 38 N 39 I 40 A 41 N 42 Q 43 G 44 L 45 S 46 T 47 L 48 I 49 V 50 N 51 S 52 T 53 F 54 Y 55 N1286Q I1331F

[0518] In some embodiments, the NGT PAM variant is selected from variant 5, 7, 28, 31 or 36 in Tables 2 and 3. In some embodiments, the variant has improved NGT PAM recognition.

[0519] In some embodiments, the NGT PAM variant has a mutation at residue 1219, 1335, 1337 and / or 1218. In some embodiments, the NGT PAM variant having a mutation for improved recognition is selected from the variants provided in Table 4 below.

[0520] Table 4: NGT PAM variant mutations at residues 1219, 1335, 1337 and 1218

[0521] Variant E1219V R1335Q T1337 G1218 1 F V T 2 F V R 3 F V Q 4 F V L 5 F V T R 6 F V R R 7 F V Q R 8 F V L R

[0522] In some embodiments, the NGT PAM is selected from the variants provided in Table 5 below.

[0523] Table 5. NGT PAM variants

[0524]

[0525] In some embodiments, the NGTN variant is variant 1. In some embodiments, the NGTN variant is variant 2. In some embodiments, the NGTN variant is variant 3. In some embodiments, the NGTN variant is variant 4. In some embodiments, the NGTN variant is variant 5. In some embodiments, the NGTN variant is variant 6.

[0526] In some embodiments, the Cas9 domain is a Cas9 domain from Streptococcus pyogenes (SpCas9). In some embodiments, the SpCas9 domain is nuclease-active SpCas9, nuclease-inactivated SpCas9 (SpCas9d), or SpCas9 nickase (SpCas9n). In some embodiments, SpCas9 contains a D9X mutation, or a corresponding mutation in any of the amino acid sequences provided herein, where X is any amino acid other than D. In some embodiments, SpCas9 contains a D9A mutation, or the corresponding mutation is in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain can bind to a nucleic acid sequence having a non-canonical PAM. In some embodiments, the SpCas9 domain, the SpCas9d domain, or the SpCas9n domain can bind to a nucleic acid sequence having an NGG, NGA, or NGCG PAM sequence.

[0527] In some embodiments, the SpCas9 domain comprises one or more of the D1135X, R1335X, and T1337X mutations, or the corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1135E, R1335Q, and T1337R mutations, or one or more of the corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1135E, R1335Q, and T1337R mutations, or the corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises one or more of the D1135X, R1335X, and T1337T1337X mutations, or the corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1135V, R1335Q, and T1337R mutations, or one or more of the corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1135V, R1335Q, and T1337R mutations, or the corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises one or more of the D1135X, G1218X, R1335X, and T1337X mutations, or the corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid. In some embodiments, the SpCas9 domain comprises one or more of the D1135V, G1218R, R1335Q, and T1337R mutations, or one or more of the corresponding mutations in any of the amino acid sequences provided herein. In some embodiments, the SpCas9 domain comprises the D1135V, G1218R, R1335Q, and T1337R mutations, or the corresponding mutations in any of the amino acid sequences provided herein.

[0528] In some embodiments, the amino acid sequence comprised by the Cas9 domain of any of the fusion proteins provided herein is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to the Cas9 polypeptide described herein. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein comprises the amino acid sequence of any of the Cas9 polypeptides described herein. In some embodiments, the Cas9 domain of any of the fusion proteins provided herein consists of the amino acid sequence of any of the Cas9 polypeptides described herein.

[0529] In some instances, the PAM recognized by the CRISPR protein-derived domain of the base editors disclosed herein can be provided to cells on an oligonucleotide that is different from the insert encoding the base editor (e.g., an AAV insert). In such an embodiment, providing the PAM on a separate oligonucleotide can allow cleavage of a target sequence that otherwise would not be cleaved because there is no adjacent PAM on the same polynucleotide as the target sequence.

[0530] In one embodiment, Streptococcus pyogenes Cas9 (SpCas9) can be used as the CRISPR endonuclease for genome engineering. However, others can also be used. In some embodiments, different endonucleases can be used to target certain genomic targets. In some embodiments, synthetic SpCas9-derived variants with non-NGG PAM sequences can be used. Additionally, other Cas9 orthologs from different species have been identified, and these "non-SpCas9" can bind to a variety of PAM sequences that can also be used in the present disclosure. For example, the relatively large SpCas9 (approximately 4 kb coding sequence) can result in inefficient expression of plasmids carrying the SpCas9 cDNA in cells. In contrast, the coding sequence of Staphylococcus aureus Cas9 (SaCas9) is approximately 1 kilobase shorter than SpCas9, which may enable its efficient expression in cells. Similar to SpCas9, the SaCas9 endonuclease is capable of modifying target genes in mammalian cells both in vitro and in mice. In some embodiments, the Cas protein can target different PAM sequences. In some embodiments, the target gene can be adjacent to, for example, the Cas9 PAM, 5'-NGG. In other embodiments, other Cas9 orthologs may have different PAM requirements. For example, other PAMs, such as those of Streptococcus thermophilus (5'-NNAGAA for CRISPR1 and 5'-NGGNG for CRISPR3) and Neisseria meningitidis (5'-NNNNGATT), can also be adjacent to the target gene.

[0531] In some embodiments, for the Streptococcus pyogenes system, the target gene sequence can be before (i.e., 5' to) the 5'-NGG PAM, and the 20-nt guide RNA sequence can base pair with the opposite strand to mediate Cas9 cleavage adjacent to the PAM. In some embodiments, the adjacent cleavage can be or can be about 3 base pairs upstream of the PAM. In some embodiments, the adjacent cleavage can be or can be about 10 base pairs upstream of the PAM. In some embodiments, the adjacent cleavage can be or can be about 0 - 20 base pairs upstream of the PAM. For example, the adjacent cleavage can be adjacent to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 base pairs upstream of the PAM. The adjacent cleavage can also be 1 to 30 base pairs downstream of the PAM. Sequences of exemplary SpCas9 proteins capable of binding the PAM sequence are as follows:

[0532] The amino acid sequences of exemplary PAM-binding SpCas9 are as follows:

[0533]

[0534] The amino acid sequence of exemplary PAM-bound SpCas9n is as follows:

[0535]

[0536] The amino acid sequence of exemplary PAM-binding SpEQR Cas9 is as follows:

[0537] In this sequence, residues E1135, Q1335, and R1337 can be mutated from D1135, R1335, and T1337 to generate SpEQR Cas9, which are underlined and in bold.

[0538] The amino acid sequence of exemplary PAM-binding SpVQR Cas9 is as follows:

[0539] In this sequence, residues V1135, Q1335, and R1337 can be mutated from D1135, R1335, and T1337 to generate SpVQR Cas9, which are underlined and in bold.

[0540] The amino acid sequence of exemplary PAM-binding SpVRER Cas9 is as follows:

[0541]

[0542] In the above sequence, residues V1135, R1218, Q1335, and R1337 can be mutated from D1134, G1217, R1335, and T1337 to generate SpVRER Cas9, which are underlined and in bold.

[0543] In some embodiments, the Cas9 domain is a recombinant Cas9 domain. In some embodiments, the Cas9 domain is a SpyMacCas9 domain. In some embodiments, the SpyMacCas9 domain is nuclease-active SpyMacCas9, nuclease-inactive SpyMacCas9 (SpyMacCas9d), or SpyMacCas9 nickase (SpyMacCas9n). In some embodiments, the SaCas9 domain, SaCas9d domain, or SaCas9n domain can bind to nucleic acid sequences with non-canonical PAMs. In some embodiments, the SpyMacCas9 domain, SpCas9d domain, or SpCas9n domain can bind to nucleic acid sequences with an NAA PAM sequence.

[0544] Exemplary SpyMacCas9

[0545]

[0546] In some cases, the variant Cas9 protein contains H840A, P475A, W476A, N477A, D1125A, W1126A, and D1218A mutations, which results in a reduced ability of the polypeptide to cleave target DNA or RNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). As another non-limiting example, in some cases, the variant Cas9 protein contains D10A, H840A, P475A, W476A, N477A, D1125A, W1126A, and D1218A mutations, which results in a reduced ability of the polypeptide to cleave target DNA. Such Cas9 proteins have a reduced ability to cleave target DNA (e.g., single-stranded target DNA) but retain the ability to bind to target DNA (e.g., single-stranded target DNA). In some cases, when the variant Cas9 protein contains W476A and W1126A mutations or when the variant Cas9 protein contains P475A, W476A, N477A, D1125A, W1126A, and D1218A mutations, the variant Cas9 protein does not bind effectively to the PAM sequence. Thus, in some such cases, when such variant Cas9 proteins are used in a binding method, the method does not require a PAM sequence. In other words, in some cases, when such variant Cas9 proteins are used in a binding method, the method can include guide RNA, but the method can be carried out in the absence of a PAM sequence (and the specificity of binding is thus provided by the targeting segment of the guide RNA). Other residues can be mutated to achieve the above effect (i.e., inactivate one or the other nuclease moiety). As non-limiting examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 can be altered (i.e., substituted). In addition, mutations other than alanine substitutions are also suitable.

[0547] In some embodiments, the CRISPR protein-derived domain of a base editor can include all or part of a Cas9 protein with a canonical PAM sequence (NGG). In other embodiments, the Cas9-derived domain of a base editor can employ a non-canonical PAM sequence. Such sequences have been described in the art and will be apparent to those skilled in the art. For example, Cas9 domains that bind non-canonical PAM sequences have been described in Kleinstiver, B.P. et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015); Kleinstiver, B.P. et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition” Nature Biotechnology 33, 1293-1298 (2015); R.T. Walton et al., “Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants” Science 10.1126 / science.aba8853 (2020); Hu et al., “Evolved Cas9 variants with broad PAM compatibility and high DNA specificity,” Nature, 2018 Apr. 5, 556(7699), 57-63; Miller et al., “Continuous evolution of SpCas9 variants compatible with non-G PAMs” Nat. Biotechnol., 2020 Apr; 38(4):471-481, the entire contents of each of which are incorporated herein by reference.Walton et al., “Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants,” Science 10.1126 / science.aba8853 (2020); Hu et al., “Evolved Cas9 variants with broad PAM compatibility and high DNA specificity,” Nature, Apr. 5, 2018, 556(7699), 57-63; Miller et al., “Continuous evolution of SpCas9 variants compatible with non-G PAMs,” Nat. Biotechnol., Apr. 2020; 38(4):471-481; the entire contents of each are hereby incorporated by reference.

[0548] A fusion protein comprising a Cas9 domain and a cytidine deaminase and / or an adenosine deaminase

[0549] Some aspects of the present disclosure provide a fusion protein comprising a Cas9 domain or other nucleic acid programmable DNA-binding protein and one or more adenosine deaminase domains, cytidine deaminase domains, and / or DNA glycosylase domains. It should be understood that the Cas9 domain can be any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9). In some embodiments, any Cas9 domain or Cas9 protein provided herein (e.g., dCas9 or nCas9) can be fused with any cytidine deaminase and adenosine deaminase provided herein. The domains of the base editors disclosed herein can be arranged in any order.

[0550] For example but not limited to, in some embodiments, the fusion protein comprises the following structures:

[0551] NH2 - [cytidine deaminase] - [Cas9 domain] - [adenosine deaminase] - COOH;

[0552] NH2 - [adenosine deaminase] - [Cas9 domain] - [cytidine deaminase] - COOH;

[0553] NH2 - [adenosine deaminase] - [cytidine deaminase] - [Cas9 domain] - COOH;

[0554] NH2 - [Cytidine deaminase] - [Adenosine deaminase] - [Cas9 domain] - COOH;

[0555] NH2 - [Cas9 domain] - [Adenosine deaminase] - [Cytidine deaminase] - COOH; or

[0556] NH2 - [Cas9 domain] - [Cytidine deaminase] - [Adenosine deaminase] - COOH.

[0557] In some embodiments, the adenosine deaminase of the fusion protein comprises TadA*8 and cytidine deaminase. In some examples, TadA*8 is TadA*8.1, TadA*8.2, TadA*8.3, TadA*8.4, TadA*8.5, TadA*8.6, TadA*8.7, TadA*8.8, TadA*8.9, TadA*8.10, TadA*8.11, TadA*8.12, TadA*8.13, TadA*8.14, TadA*8.15, TadA*8.16, TadA*8.17, TadA*8.18, TadA*8.19, TadA*8.20, TadA*8.28.2A., or TadA*8.24.

[0558] Exemplary fusion protein structures include the following:

[0559] NH2 - [Adenosine deaminase] - [Cas9] - [[Cytidine deaminase] - COOH;

[0560] NH2 - [Cytidine deaminase] - [Cas9] - [Adenosine deaminase] - COOH;

[0561] NH2 - [TadA*8] - [Cas9] - [Cytidine deaminase] - COOH; or

[0562] NH2 - [Cytidine deaminase] - [Cas9] - [TadA*8] - COOH.

[0563] In some embodiments, the fusion protein comprising cytidine deaminase, base editor, adenosine deaminase, and napDNAbp (such as Cas9 domain) does not include a linker sequence. In some embodiments, there is a linker between the cytidine deaminase and adenosine deaminase domains and napDNAbp. In some examples, the "-" used in the above general architecture indicates the presence of an optional linker. In some embodiments, the cytidine deaminase, adenosine deaminase, and napDNAbp are fused via any linker provided herein. For example, in some embodiments, the cytidine deaminase, adenosine deaminase, and napDNAbp are fused via any linker provided in the following paragraph titled "Linker".

[0564] For example, in some embodiments, the cytidine deaminase, adenosine deaminase, and napDNAbp are fused via any linker provided in the section entitled "Linker" below.

[0565] NH2-NLS-[Cytidine deaminase]-[Cas9 domain]-[Adenosine deaminase]-COOH;

[0566] NH2-NLS-[Adenosine deaminase]-[Cas9 domain]-[Cytidine deaminase]-COOH;

[0567] NH2-NLS-[Adenosine deaminase][Cytidine deaminase]-[Cas9 domain]-COOH;

[0568] NH2-NLS-[Cytidine deaminase]-[Adenosine deaminase]-[Cas9 domain]-COOH;

[0569] NH2-NLS-[Cas9 domain]-[Adenosine deaminase]-[Cytidine deaminase]-COOH;

[0570] NH2-NLS-[Cas9 domain]-[Cytidine deaminase]-[Adenosine deaminase]-COOH;

[0571] NH2-[Cytidine deaminase]-[Cas9 domain]-[Adenosine deaminase]-NLS-COOH;

[0572] NH2-[Adenosine deaminase]-[Cas9 domain]-[Cytidine deaminase]-NLS-COOH;

[0573] NH2-[Adenosine deaminase][Cytidine deaminase]-[Cas9 domain]-NLS-COOH;

[0574] NH2-[Cytidine deaminase]-[Adenosine deaminase]-[Cas9 domain]-NLS-COOH;

[0575] NH2-[Cas9 domain]-[Adenosine deaminase]-[Cytidine deaminase]-NLS-COOH; or

[0576] NH2-[Cas9 domain]-[Cytidine deaminase]-[Adenosine deaminase]-NLS-COOH.

[0577] In some embodiments, the NLS is present in or flanked by a linker, as described herein. In some embodiments, the N-terminal or C-terminal NLS is a bipartite NLS. A bipartite NLS contains two clusters of basic amino acids that are separated by a relatively short spacer sequence (thus bipartite - two parts, while a monopartite NLS is not). The NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK, is the prototype of a ubiquitous bipartite signal: two clusters of basic amino acids separated by a spacer of about 10 amino acids. Sequences of exemplary bipartite NLSs are as follows: PKKKRKVEGADKRTADGSEFESPKKKRKV.

[0578] In some embodiments, a fusion protein comprising a cytidine deaminase, an adenosine deaminase, a Cas9 domain, and an NLS does not contain a linker sequence. In some embodiments, there is a linker sequence between one or more domains or proteins (e.g., a cytidine deaminase, an adenosine deaminase, a Cas9 domain, or an NLS).

[0579] It should be understood that the fusion proteins of the present disclosure can include one or more additional features. For example, in some embodiments, the fusion protein can include an inhibitor, a cytoplasmic localization sequence, an export sequence, such as a nuclear export sequence or other localization sequences, and sequence tags that can be used for solubilizing, purifying, or detecting the fusion protein. Suitable protein tags provided herein include, but are not limited to, biotin carboxyl carrier protein (BCCP) tag, myc tag, calmodulin tag, FLAG tag, hemagglutinin (HA) tag, polyhistidine tag, also known as a histidine tag or His-tag, maltose binding protein (MBP)-tag, nus-tag, glutathione-S-transferase (GST)-tag, green fluorescent protein (GFP)-tag, thioredoxin-tag, S-tag, Softags (e.g., Softag 1, Softag 3), streptag, biotin ligase tag, Flash tag, V5 tag, and SBP tag. Other suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein contains one or more His tags.

[0580] Exemplary but non-limiting fusion proteins are described in International PCT Application Nos. PCT / 2017 / 044935 and PCT / US2020 / 016288, which are hereby incorporated by reference in their entireties.

[0581] Fusion proteins comprising a nuclear localization sequence (NLS)

[0582] In some embodiments, the fusion proteins provided herein further comprise one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, such as nuclear localization sequences (NLSs). In one embodiment, a bipartite NLS is used. In some embodiments, an NLS comprises an amino acid sequence that facilitates the import of a protein comprising the NLS into the cell nucleus (e.g., by nuclear transport). In some embodiments, any of the fusion proteins provided herein further comprises a nuclear localization sequence (NLS). In some embodiments, the NLS is fused to the N-terminus of the fusion protein. In some embodiments, the NLS is fused to the C-terminus of the fusion protein. In some embodiments, the NLS is fused to the N-terminus of the Cas9 domain. In some embodiments, the NLS is fused to the C-terminus of the nCas9 domain or the dCas9 domain. In some embodiments, the NLS is fused to the N-terminus of the deaminase. In some embodiments, the NLS is fused to the C-terminus of the deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. In some embodiments, the NLS comprises the amino acid sequence of any of the NLS sequences provided or mentioned herein. Additional nuclear localization sequences are known in the art and will be apparent to the skilled person. For example, NLS sequences are described in Plank et al., PCT / EP2000 / 011690, the contents of which are incorporated herein by reference as they disclose exemplary nuclear localization sequences. In some embodiments, the NLS comprises the amino acid sequence PKKKRKVEGADKRTADGSEFESPKKKRKV, KRTADGSEFESPKKKRKV, KRPAATKKAGQAKKKK, KKTELQTTNAENKTKKL, KRGINDRNFWRGENGRKTR, RKSGKIAAIVVKRPRKPKKKRKV or MDSLLMNRRKFLYQFKNVRWAKGRRETYLC. In some embodiments, the NLS is present within a linker or flanked by linkers such as those described herein. In some embodiments, the N-terminal or C-terminal NLS is a bipartite NLS. A bipartite NLS comprises two clusters of basic amino acids that are separated by a relatively short spacer sequence (hence bipartite - 2 parts, whereas monopartite NLSs are not). The NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK, is the prototype of the ubiquitous bipartite signal: two clusters of basic amino acids separated by a spacer of approximately 10 amino acids. Sequences of exemplary bipartite NLSs are as follows: PKKKRKVEGADKRTADGSEFESPKKKRKV.

[0583] In some embodiments, the fusion proteins of the invention do not contain a linker sequence. In some embodiments, there is a linker sequence between one or more domains or proteins. In some embodiments, the general structure of an exemplary Cas9 fusion protein having an adenosine deaminase or a cytidine deaminase and a Cas9 domain comprises any of the following structures, where NLS is a nuclear localization sequence (e.g., any NLS provided herein), NH2 is the N-terminus of the fusion protein, and COOH is the C-terminus of the fusion protein:

[0584] NH2-NLS-[adenosine deaminase]-[Cas9 domain]-COOH;

[0585] NH2-NLS[Cas9 domain]-[adenosine deaminase]-COOH;

[0586] NH2-[adenosine deaminase]-[Cas9 domain]-NLS-COOH;

[0587] NH2-[Cas9 domain]-[adenosine deaminase]-NLS-COOH;

[0588] NH2-NLS-[cytidine deaminase]-[Cas9 domain]-COOH;

[0589] NH2-NLS[Cas9 domain]-[cytidine deaminase]-COOH;

[0590] NH2-[cytidine deaminase]-[Cas9 domain]-NLS-COOH; or

[0591] NH2-[Cas9 domain]-[cytidine deaminase]-NLS-COOH.

[0592] It should be understood that the fusion proteins of the present disclosure can include one or more additional features. For example, in some embodiments, the fusion protein can include inhibitors, cytoplasmic localization sequences, export sequences such as nuclear export sequences or other localization sequences, and sequence tags that can be used for solubilization, purification, or detection of the fusion protein. Suitable protein tags provided herein include, but are not limited to, biotin carboxyl carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, polyhistidine tags, also known as histidine tags or His-tags, maltose binding protein (MBP)-tags, nus-tags, glutathione-S-transferase (GST)-tags, green fluorescent protein (GFP)-tags, thioredoxin-tags, S-tags, Softags (e.g., Softag 1, Softag 3), streptavidin tags, biotin ligase tags, Flash tags, V5 tags, and SBP tags. Other suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein includes one or more His tags.

[0593] Vectors encoding CRISPR enzymes containing one or more nuclear localization sequences (NLSs) can be used. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 NLSs can be used or approximately used. The CRISPR enzyme can include an NLS located at or near the amino terminus, about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 NLSs located at or near the carboxyl terminus, or any combination of these (e.g., one or more NLSs at the amino terminus and one or more NLSs at the carboxyl terminus). When more than one NLS is present, each NLS can be selected independently of the others such that a single NLS can be present in more than one copy and / or be present in more than one copy with one or more other NLSs.

[0594] The CRISPR enzyme used in the method can include about 6 NLSs. When the amino acid closest to the NLS is within about 50 amino acids of the polypeptide chain from the N or C terminus, e.g., within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, or 50 amino acids.

[0595] Fusion proteins with internal insertions

[0596] The present disclosure provides fusion proteins that comprise a heterologous polypeptide fused to a nucleic acid programmable nucleic acid binding protein (e.g., napDNAbp). The heterologous polypeptide can be a polypeptide not found in the native or wild-type napDNAbp polypeptide sequence. The heterologous polypeptide can be fused to the napDNAbp at the C-terminus of the napDNAbp, at the N-terminus of the napDNAbp, or inserted at an internal position within the napDNAbp. In some embodiments, the heterologous polypeptide is inserted at an internal position within the napDNAbp.

[0597] In some embodiments, the heterologous polypeptide is a deaminase or a functional fragment thereof. For example, the fusion protein can comprise a deaminase flanked by an N-terminal fragment and a C-terminal fragment of a Cas9 or Cas12 (e.g., Cas12b / C2c1) polypeptide. The deaminase in the fusion protein can be an adenosine deaminase. In some embodiments, the adenosine deaminase is TadA (e.g., TadA7.10 or TadA*8). In some embodiments, the TadA is TadA*8. The TadA sequences (e.g., TadA7.10 or TadA*8) as described herein are deaminases suitable for the above-described fusion proteins.

[0598] The deaminase can be a circularly permuted deaminase. For example, the deaminase can be a circularly permuted adenosine deaminase. In some embodiments, the deaminase is circularly permuted TadA, circularly permuted at amino acid residue 116 numbered in the TadA reference sequence. In some embodiments, the deaminase is circularly permuted TadA, circularly permuted at amino acid residue 136 numbered in the TadA reference sequence. In some embodiments, the deaminase is circularly permuted TadA, circularly permuted at amino acid residue 65 numbered in the TadA reference sequence.

[0599] The fusion protein can comprise more than one deaminase. The fusion protein can comprise, for example, 1, 2, 3, 4, 5 or more deaminases. In some embodiments, the fusion protein comprises one deaminase. In some embodiments, the fusion protein comprises two deaminases. Two or more deaminases in the fusion protein can be adenosine deaminases, cytidine deaminases, or a combination thereof. Two or more deaminases can be homodimers. Two or more deaminases can be heterodimers. Two or more deaminases can be inserted in tandem into the napDNAbp. In some embodiments, two or more deaminases may not be in tandem in the napDNAbp.

[0600] In some embodiments, the napDNAbp in the fusion protein is a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide can be a variant Cas9 polypeptide. In some embodiments, the Cas9 polypeptide is a Cas9 nickase (nCas9) polypeptide or a fragment thereof. In some embodiments, the Cas9 polypeptide is a nuclease-dead Cas9 (dCas9) polypeptide or a fragment thereof. The Cas9 polypeptide in the fusion protein can be a full-length Cas9 polypeptide. In some cases, the Cas9 polypeptide in the fusion protein may not be a full-length Cas9 polypeptide. The Cas9 polypeptide can be truncated, for example, at the N-terminus or C-terminus relative to the naturally occurring Cas9 protein. The Cas9 polypeptide can be a circularly permuted Cas9 protein. The Cas9 polypeptide can be a fragment, portion, or domain of a Cas9 polypeptide that is still capable of binding to a target polynucleotide and a guide nucleic acid sequence.

[0601] In some embodiments, the Cas9 polypeptide is Streptococcus pyogenes Cas9 (SpCas9), Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), or a fragment or variant thereof.

[0602] The Cas9 polypeptide of the fusion protein can comprise at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to a naturally occurring Cas9 polypeptide.

[0603] The Cas9 polypeptide of the fusion protein can comprise at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity to the Cas9 amino acid sequence set forth below (hereinafter referred to as the "Cas9 reference sequence"):

[0604] (Single underline: HNH domain; double underline: RuvC domain).

[0605] Fusion proteins comprising heterologous catalytic domains with N- and C-terminal fragments of Cas9 polypeptides flanking them can also be used for base editing in the methods described herein. Fusion proteins comprising Cas9 and one or more deaminase domains, such as adenosine deaminase, or comprising an adenosine deaminase domain flanked by Cas9 sequences can also be used for highly specific and efficient base editing of target sequences. In one embodiment, the chimeric Cas9 fusion protein comprises a heterologous catalytic domain (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) inserted within the Cas9 polypeptide. In some embodiments, the fusion protein comprises an adenosine deaminase domain and a cytidine deaminase domain inserted within Cas9. In some embodiments, the adenosine deaminase is fused within Cas9 and the cytidine deaminase is fused to the C-terminus. In some embodiments, the adenosine deaminase is fused within Cas9 and the cytidine deaminase is fused to the N-terminus. In some embodiments, the cytidine deaminase is fused within Cas9 and the adenosine deaminase is fused to the C-terminus. In some embodiments, the cytidine deaminase is fused within Cas9 and the adenosine deaminase is fused to the N-terminus.

[0606] Exemplary structures of fusion proteins having adenosine deaminase and cytidine deaminase as well as Cas9 are provided below:

[0607] NH2-[Cas9 (adenosine deaminase)]-[cytidine deaminase]-COOH;

[0608] NH2-[cytidine deaminase]-[Cas9 (adenosine deaminase)]-COOH;

[0609] NH2-[Cas9 (cytidine deaminase)]-[adenosine deaminase]-COOH; or

[0610] NH2-[adenosine deaminase]-[Cas9 (cytidine deaminase)]-COOH.

[0611] In some embodiments, the "-" used in the above general architecture indicates the presence of an optional linker.

[0612] In various embodiments, the catalytic domain has DNA modification activity (e.g., deaminase activity), such as adenosine deaminase activity. In some embodiments, the adenosine deaminase is TadA (e.g., TadA7.10). In some embodiments, the TadA is TadA*8. In some embodiments, TadA*8 is fused within Cas9 and a cytidine deaminase is fused to the C-terminus. In some embodiments, TadA*8 is fused within Cas9 and a cytidine deaminase is fused to the N-terminus. In some embodiments, a cytidine deaminase is fused within Cas9 and TadA*8 is fused to the C-terminus. In some embodiments, a cytidine deaminase is fused within Cas9 and TadA*8 is fused to the N-terminus. Exemplary structures of fusion proteins having TadA*8 and a cytidine deaminase and Cas9 are provided below:

[0613] NH2-[Cas9(TadA*8)]-[cytidine deaminase]-COOH;

[0614] NH2-[cytidine deaminase]-[Cas9]-[Cas9(TadA*8)]-COOH;

[0615] NH2-[Cas9(cytidine deaminase)]-[TadA*8]-COOH; or

[0616] NH2-[TadA*8]-[Cas9(cytidine deaminase)]-COOH.

[0617] In some embodiments, the "-" used in the above general architecture indicates the presence of an optional linker.

[0618] A heterologous polypeptide (e.g., a deaminase) can be inserted into a suitable position of a napDNAbp (e.g., Cas9 or Cas12 (e.g., Cas12b / C2c1)) such that the napDNAbp retains its ability to bind a target polynucleotide and a guide nucleic acid. A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into the napDNAbp without impairing the function of the deaminase (e.g., base editing activity) or the napDNAbp (e.g., the ability to bind a target nucleic acid and a guide nucleic acid). A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into the napDNAbp, such as into a disordered region shown by crystallographic studies or a region containing a temperature factor or B-factor. Less ordered, disordered, or unstructured protein regions, such as solvent-exposed regions and loops, can be used for insertion without impairing structure or function. A deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) can be inserted into a flexible loop region or a solvent-exposed region of the napDNAbp. In some embodiments, a deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted into a flexible loop of a Cas9 or Cas12b / C2c1 polypeptide.

[0619] In some embodiments, the insertion position of a deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is determined by B-factor ...

Claims

1. A method for in vitro manufacturing hematopoietic stem cells or their progenitor cells for treating hemoglobinopathy, blood cancer, or myeloproliferative disease, comprising: (a) expressing a nucleobase editing polypeptide in the hematopoietic stem cells or their progenitor cells, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA-binding protein and a deaminase; and (b) contacting the hematopoietic stem cells or their progenitor cells with a guide RNA that targets a nucleic acid molecule encoding a cell surface protein selected from the group consisting of CD117, CXCR4, CD135, CD90, CD45, and CD34 and introducing a mutation in the cell surface protein.

2. A method for in vitro manufacturing hematopoietic stem cells or their progenitor cells for treating hemoglobinopathy, blood cancer, or myeloproliferative disease, comprising: (a) expressing a nucleobase editing polypeptide in hematopoietic stem cells or their progenitor cells comprising CD117 protein, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA-binding protein and a deaminase; (b) contacting the hematopoietic stem cells or their progenitor cells with a guide RNA that can target a cell surface protein encoded by a polynucleotide, thereby manufacturing the hematopoietic stem cells or their progenitor cells for treating hemoglobinopathy, blood cancer, or myeloproliferative disease.

3. A method for in vitro identifying a mutation that alters the binding of an antibody to a cell surface protein, comprising: (a) expressing a nucleobase editing polypeptide in a cell comprising a cell surface protein, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA-binding protein and a deaminase; (b) contacting the cell with a guide RNA that can target the nucleic acid molecule encoding the cell surface protein and introducing a mutation in the cell surface protein; and (c) contacting the cell with an antibody that specifically binds to the wild-type cell surface protein but exhibits reduced binding to the cell surface protein comprising the mutation, thereby identifying a mutation that alters the binding of the antibody to the cell surface protein.

4. A method for in vitro identifying a mutation that alters the binding of an antibody to a cell surface protein, comprising: (a) expressing a nucleobase editing polypeptide in hematopoietic stem cells or their progenitor cells, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA-binding protein and a deaminase; (b) contacting the cell with a guide RNA that can target a nucleic acid molecule encoding a cell surface protein selected from the group consisting of CD117, CXCR4, CD135, CD90, CD45, and CD34 and introducing a mutation in the cell surface protein; (c) contacting the cell with an antibody that specifically binds to the wild-type cell surface protein but exhibits reduced binding to the cell surface protein comprising the mutation, thereby identifying a mutation that alters the binding of the antibody to the cell surface protein.

5. A method for in vitro base editing the gene encoding a cell surface protein expressed by hematopoietic stem cells or their progenitor cells, comprising (a) expressing a nucleobase editing polypeptide in hematopoietic stem cells or their progenitor cells comprising CD117 protein, wherein the nucleobase editing polypeptide comprises a nucleic acid programmable DNA-binding protein and a deaminase; and (b) contacting the cells with a guide RNA capable of targeting a nucleic acid molecule encoding the CD117 protein, thereby performing base editing on the gene encoding the cell surface protein.

6. Use of the edited hematopoietic stem cells generated in step (b) in the preparation of a medicament for conditioning a subject simultaneously with or subsequent to hematopoietic stem cell transplantation, comprising: (a) expressing, in isolated hematopoietic stem cells of a subject or a donor, a nuclear base editing polypeptide comprising a nucleic acid programmable DNA-binding protein and a deaminase; (b) contacting the hematopoietic stem cells with a guide RNA capable of targeting a nucleic acid molecule encoding a cell surface protein selected from the group consisting of CD117, CXCR4, CD135, CD90, CD45, and CD34, thereby introducing a mutation in the cell surface protein and generating edited hematopoietic stem cells; (c) administering the edited hematopoietic stem cells to the subject; and (d) administering to the subject an antibody, antibody-drug conjugate, or chimeric antigen receptor-expressing T cell that selectively binds to the wild-type cell surface protein, wherein the administration in step (d) is simultaneous with or subsequent to step (c).

7. Use of the edited hematopoietic stem cells generated in step (b) in the preparation of a medicament for conditioning a subject simultaneously with or subsequent to hematopoietic stem cell transplantation, comprising (a) expressing a nuclear base editing polypeptide in the hematopoietic stem cells of a subject, the nuclear base editing polypeptide comprising a nucleic acid programmable DNA-binding protein and a deaminase; (b) contacting the hematopoietic stem cells with a guide RNA capable of targeting a nucleic acid molecule encoding the CD117 cell surface protein, thereby introducing a mutation in the CD117 protein and generating edited hematopoietic stem cells; (c) administering the edited hematopoietic stem cells to the subject; and (d) administering to the subject an antibody, antibody-drug conjugate, or chimeric antigen receptor-expressing T cell that selectively binds to wild-type CD117, wherein the administration in step (d) is simultaneous with or subsequent to step (c).

8. A base editing system comprising a fusion protein or a polynucleotide encoding the fusion protein, wherein, The fusion protein comprises a nucleic acid programmable DNA-binding protein, a deaminase domain, and a guide polynucleotide, and the guide polynucleotide comprises a nucleic acid sequence selected from Table 23.

9. A base editing system comprising a fusion protein or a polynucleotide encoding the fusion protein, wherein, The fusion protein comprises a nucleic acid programmable DNA-binding protein, a deaminase domain, and a guide polynucleotide, and the guide polynucleotide hybridizes to a complementary sequence of a target nucleic acid sequence, and the complementary sequence of the target nucleic acid sequence is selected from the group consisting of: CAAGCTATCTTCTTAGGGA; GCAAGCTATCTTCTTAGGGAA; ACTTACGACAGGCTCGTGAA; TGACCAATTATTCCCTCAAG; AGTGACCAATTATTCCCTCA; ACTACAGTATTTGTAAACGA; AAGACAACGACACGCTGGTC; and GGCTGTTATGCACTGATCCG.

10. Use of a pharmaceutical composition for the preparation of a medicament for treating a subject suffering from a hemoglobinopathy, a blood cancer or a myeloproliferative disorder, wherein the pharmaceutical composition comprises: an antibody, an antibody-drug conjugate or a chimeric antigen receptor-expressing T cell that selectively binds to a cell surface protein; and a cell generated by introducing the base editing system according to claim 8 or 9 into a cell or its progenitor cell, thereby treating a hemoglobinopathy, a blood cancer or a myeloproliferative disorder.

Citation Information

Patent Citations

  • Towel rings

    CN3315821D

  • cell phone

    CN3329834D

  • Split inteins, conjugates and uses thereof

    US20150344549A1

  • Switchable cas9 nucleases and uses thereof

    US20160208288A1

  • AAV delivery of nucleobase editors

    US20180127780A1