Synthetic polypeptides and their uses

Synthetic polypeptides with trans-splicing inteins and their fragments enable efficient delivery of large polynucleotide payloads using AAVs, addressing the packaging limitations of AAVs and facilitating genome editing.

JP2025533556APending Publication Date: 2025-10-07BEAM THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517549
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-04-21
Filing Date
2023-09-26
Publication Date
2025-10-07

AI Technical Summary

Technical Problem

The limited packaging capacity of adeno-associated viruses (AAVs) poses a challenge in delivering all necessary elements for genome editing, particularly large polynucleotide payloads encoding polypeptides.

Method used

The use of synthetic polypeptides comprising trans-splicing inteins and their functional fragments, along with encoding polynucleotides, to facilitate delivery of polynucleotides encoding split-polypeptides using vectors with limited packaging capacity, such as adeno-associated viral vectors.

Benefits of technology

Enables efficient delivery of large polynucleotide payloads, overcoming the packaging limitations of AAVs and facilitating genome editing processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533556000053
    Figure 2025533556000053
  • Figure 2025533556000054
    Figure 2025533556000054
  • Figure 2025533556000055
    Figure 2025533556000055
Patent Text Reader

Abstract

Synthetic polypeptides comprising a trans-splicing intein, functional fragments thereof, and polynucleotides encoding them are provided for use in systems, compositions, kits, and methods for delivering one or more polynucleotides (e.g., polynucleotides encoding split-polypeptides) to cells using vectors with limited packaging capacity (e.g., viral vectors such as adeno-associated viral vectors).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 497,632, filed April 21, 2023, and U.S. Provisional Application No. 63 / 377,131, filed September 26, 2022, the entire contents of which are incorporated herein by reference in their entirety.

[0002] Sequence Listing

[0001] This application contains a Sequence Listing that has been filed electronically in XML format, the entire contents of which are incorporated herein by reference. The XML file of the Sequence Listing, created on September 25, 2023, is named 180802-046603PCT_SL.xml and is 974,566 bytes in size. [Background technology]

[0003] The discovery of forms of genome editing such as clustered regularly interspaced short palindromic repeats (CRISPR) has revolutionized the field of molecular biology. Much of the enthusiasm for genome editing has focused on its clinical potential to treat human diseases. One of the challenges to achieving this goal is delivering the elements necessary for genome editing. Due to the limited packaging capacity of adeno-associated viruses (AAVs), it is difficult to deliver all the elements necessary to achieve the desired gene editing goal.

[0004] Thus, there is a need for improved methods for delivering large polynucleotide payloads (e.g., polynucleotides encoding polypeptides) to cells using AAV or other vectors with limited packaging capacity. Summary of the Invention

[0005] As described below, the disclosure features synthetic polypeptides comprising trans-splicing inteins, functional fragments thereof, and polynucleotides encoding them, for use in systems, compositions, kits, and methods for delivering one or more polynucleotides (e.g., polynucleotides encoding split-polypeptides) to cells using vectors with limited packaging capacity (e.g., viral vectors such as adeno-associated viral vectors).

[0006] As described below, the disclosure features synthetic polypeptides comprising trans-splicing inteins, functional fragments thereof, and polynucleotides encoding them, for use in systems, compositions, kits, and methods for delivering one or more polynucleotides (e.g., polynucleotides encoding split-polypeptides) to cells using vectors with limited packaging capacity (e.g., viral vectors such as adeno-associated viral vectors).

[0007] In one aspect, the disclosure features a synthetic polypeptide including an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427), Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432).

[0008] In another aspect, the disclosure features a polynucleotide encoding a synthetic polypeptide of any of the aspects or embodiments thereof provided herein.

[0009] In another aspect, the disclosure features a cell including a polynucleotide of any of the aspects or embodiments thereof provided herein.

[0010] In another aspect, the disclosure features a pair of vectors, one member of which includes a polynucleotide sequence encoding a synthetic polypeptide N (Syn-N) having at least about 85% amino acid sequence identity to a sequence selected from one or more of the following: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO:427). The other member of the pair of vectors comprises a polynucleotide sequence encoding a synthetic polypeptide C (Syn-C) having at least about 85% amino acid sequence identity to a sequence selected from one or more of the following: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432).

[0011] In another aspect, the disclosure features a pair of adeno-associated virus (AAV) vectors, wherein one member of the AAV vector pair includes a polynucleotide sequence encoding a synthetic polypeptide N (Syn-N) selected from one or more of the following: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427). The other member of the AAV vector pair is a synthetic polypeptide C (Syn-C) selected from one or more of the following: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432).

[0012] In another aspect, the disclosure features a cell including a pair of vectors of any of the aspects or embodiments thereof provided herein.

[0013] In another aspect, the disclosure features a fusion protein comprising a heterologous polypeptide fragment fused at its C-terminus to a synthetic polypeptide comprising an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427).

[0014] In another aspect, the disclosure features a fusion protein that includes a heterologous polypeptide fused at its N-terminus to a synthetic polypeptide that includes an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432).

[0015] In another aspect, the disclosure features a polynucleotide encoding a fusion protein of any of the aspects or embodiments thereof provided herein.

[0016] In another aspect, the disclosure features a vector including a polynucleotide of any of the aspects or embodiments thereof provided herein.

[0017] In another aspect, the disclosure features a cell including the fusion protein, polynucleotide, or vector of any of the aspects or embodiments thereof provided herein.

[0018] In another aspect, the disclosure features a composition including a fusion protein, polynucleotide, vector, or cell of any of the aspects or embodiments thereof provided herein.

[0019] In another aspect, the disclosure features a pharmaceutical composition, the pharmaceutical composition including a fusion protein, polynucleotide, vector, or cell of any of the aspects or embodiments thereof provided herein and a pharmaceutically acceptable excipient.

[0020] In another aspect, the disclosure features a polynucleotide delivery system that includes: (a) a first polynucleotide that encodes a fusion protein that includes a heterologous polypeptide fused at its C-terminus to a first synthetic polypeptide, the first synthetic polypeptide comprising an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427). The system further includes (b) a second polynucleotide encoding a fusion protein comprising another heterologous polypeptide fused at its N-terminus to the second synthetic polypeptide. The second synthetic polypeptide comprises an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432).

[0021] In another aspect, the disclosure features a polynucleotide delivery system. The system includes: (a) a first polynucleotide, the first polynucleotide encoding a fusion protein comprising an N-terminal fragment of a base editor. The base editor comprises a deaminase domain, a nucleic acid-programmable DNA-binding protein (napDNAbp) domain, and a first synthetic polypeptide fused to the C-terminus of the N-terminal fragment of the base editor. The first synthetic polypeptide comprises an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427). The system further comprises (b) a second polynucleotide, wherein the second polynucleotide encodes a fusion protein comprising a second synthetic polypeptide fused to the N-terminus of the C-terminal fragment of the base editor. The second synthetic polypeptide comprises an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432).

[0022] In another aspect, the disclosure features a composition including a polynucleotide delivery system of any of the aspects or embodiments thereof provided herein.

[0023] In another aspect, the disclosure features a pharmaceutical composition including a composition of any of the aspects or embodiments thereof provided herein and a pharmaceutical excipient.

[0024] In another aspect, the disclosure features a method for delivering a polynucleotide encoding a heterologous polypeptide to a cell. The method includes (a) contacting the cell with a first polynucleotide, the first polynucleotide encoding a fusion protein comprising a heterologous polypeptide, or a fragment thereof, fused at its C-terminus to a first synthetic polypeptide. The first synthetic polypeptide comprises an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO:427). The method further comprises (a) contacting the cell with a second polynucleotide, wherein the second polynucleotide encodes a fusion protein comprising another heterologous polypeptide fused at its N-terminus to a second synthetic polypeptide. The second synthetic polypeptide comprises an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432).

[0025] In another aspect, the disclosure features a method for delivering a polynucleotide encoding a base-edited fragment to a cell. The method includes (a) contacting the cell with a first polynucleotide, wherein the first polynucleotide encodes a fusion protein comprising an N-terminal fragment of a base editor. The base editor comprises a deaminase domain and a nucleic acid-programmable DNA-binding protein (napDNAbp) domain. The first polynucleotide also comprises a first synthetic polypeptide fused to the C-terminus of the N-terminal fragment of the base editor. The first synthetic polypeptide comprises an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427). The method further comprises contacting the cell with (b) a second polynucleotide, wherein the second polynucleotide encodes a fusion protein comprising a second synthetic polypeptide fused to the N-terminus of the C-terminal fragment of the base editor. The second synthetic polypeptide comprises an amino acid sequence having at least 85% sequence identity to one of the following sequences, or a functional fragment thereof: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432).

[0026] In another aspect, the disclosure features a method of editing a target polynucleotide in a cell, the method including delivering to the cell a polynucleotide encoding a base editor fragment according to the methods of any of the aspects or embodiments thereof provided herein.

[0027] In another aspect, the disclosure features a kit suitable for use in the method of any of the aspects or embodiments thereof provided herein, the kit including a polynucleotide, polypeptide, vector, or composition of any of the aspects or embodiments thereof provided herein.

[0028] In any of the aspects or embodiments thereof provided herein, the polypeptide has at least 95% sequence identity to SEQ ID NO: 425, 426, 427, 428, 429, 430, 431, or 432. In any of the aspects or embodiments thereof provided herein, the polypeptide comprises SEQ ID NO: 425, 426, 427, 428, 429, 430, 431, or 432. In any of the aspects or embodiments thereof provided herein, the polypeptide comprises only SEQ ID NO: 425, 426, 427, 428, 429, 430, 431, or 432.

[0029] In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide is fused to a heterologous polypeptide.

[0030] In any of the aspects or embodiments thereof provided herein, the vector(s) are selected from one or more of a retroviral vector, an adenoviral vector, a lentiviral vector, a herpesvirus vector, and an adeno-associated virus vector. In any of the aspects or embodiments thereof provided herein, the vector comprises a lipid nanoparticle. In any of the aspects or embodiments thereof provided herein, the vector is an adeno-associated virus (AAV) vector.

[0031] In any of the aspects or embodiments thereof provided herein, Syn-N has at least 95% sequence identity to SEQ ID NO: 425, 426, or 427, and Syn-C has at least 95% sequence identity to SEQ ID NO: 428, 429, 430, 431, or 432. In any of the aspects or embodiments thereof provided herein, Syn-N comprises SEQ ID NO: 425, 426, or 427, and Syn-C comprises SEQ ID NO: 428, 429, 430, 431, or 432. In any of the aspects or embodiments thereof provided herein, Syn-N comprises only SEQ ID NO: 425, 426, or 427, and Syn-C comprises only SEQ ID NO: 428, 429, 430, 431, or 432.

[0032] In any of the aspects or embodiments thereof provided herein, Syn-N and Syn-C are each fused to a heterologous polypeptide.

[0033] In any of the aspects provided herein or embodiments thereof, Syn-N and Syn-C can mediate binding between heterologous polypeptides to which they are fused. In any of the aspects provided herein, the binding is a non-covalent bond. In any of the aspects provided herein, the binding is a covalent bond. In any of the aspects provided herein, base editing activity is restored by the binding. In any of the aspects provided herein, base editing activity is restored by binding of a base editor fragment. In any of the aspects provided herein, Syn-N and Syn-C mediate peptide bond formation between heterologous polypeptides to which they are fused. In any of the aspects provided herein or embodiments thereof, Syn-N and Syn-C are cleaved during peptide bond formation. In any of the aspects provided herein or embodiments thereof, the first and second synthetic polypeptides mediate peptide bond formation between polypeptides to which they are fused. In any of the aspects provided herein or embodiments thereof, the first and second synthetic polypeptides are cleaved during peptide bond formation.

[0034] In any of the aspects or embodiments thereof provided herein, the heterologous polypeptides are each fragments of a base editor, wherein the base editor comprises a deaminase domain and a nucleic acid-programmable DNA-binding protein domain.

[0035] In any of the aspects or embodiments thereof provided herein, Syn-N is an N intein and Syn-C is a C intein, which together can function in protein splicing. In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide is an N intein that can function in protein splicing. In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide is a C intein that can function in protein splicing.

[0036] In any of the aspects or embodiments thereof provided herein, Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn1-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn4-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn5-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn9-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn10-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn1-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn4-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn5-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn9-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn10-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn1-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn4-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn5-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn9-C; or Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn10-C.

[0037] In any of the aspects or embodiments thereof provided herein, Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn1-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn4-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn5-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn9-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn2-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn10-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn1-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn4-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn5-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn9-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn3-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn10-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn1-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn4-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn5-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn9-C; or Syn-N comprises an amino acid sequence having at least about 95% sequence identity with Syn5-N, and Syn-C comprises an amino acid sequence having at least about 95% sequence identity with Syn10-C.

[0038] In any of the aspects or embodiments thereof provided herein, Syn-N contains Syn2-N and Syn-C contains Syn1-C, or Syn-N contains Syn2-N and Syn-C contains Syn4-C, or Syn-N contains Syn2-N and Syn-C contains Syn5-C, Syn-N includes Syn2-N and Syn-C includes Syn9-C, or Syn-N includes Syn2-N and Syn-C includes Syn10-C, or Syn-N includes Syn3-N and Syn-C includes Syn1-C, or Syn-N contains Syn3-N and Syn-C contains Syn4-C, or Syn-N includes Syn3-N and Syn-C includes Syn5-C, or Syn-N includes Syn3-N and Syn-C includes Syn9-C, or Syn-N includes Syn3-N and Syn-C includes Syn10-C, or Syn-N includes Syn5-N and Syn-C includes Syn1-C, or Syn-N contains Syn5-N and Syn-C contains Syn4-C, or Syn-N contains Syn5-N and Syn-C contains Syn5-C, or Syn-N includes Syn5-N and Syn-C includes Syn9-C, or Syn-N includes Syn5-N, and Syn-C includes Syn10-C.

[0039] In any of the aspects or embodiments thereof provided herein, the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn9-C; or The first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn10-C.

[0040] In any of the aspects or embodiments thereof provided herein, the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity to Syn9-C; or The first synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 90% sequence identity with Syn10-C.

[0041] In any of the aspects or embodiments thereof provided herein, the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity to Syn9-C; or The first synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity with Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 95% sequence identity with Syn10-C.

[0042] In any of the aspects or embodiments thereof provided herein, the first synthetic polypeptide comprises Syn2-N and the second synthetic polypeptide comprises Syn1-C; or the first synthetic polypeptide comprises Syn2-N and the second synthetic polypeptide comprises Syn4-C; or the first synthetic polypeptide comprises Syn2-N and the second synthetic polypeptide comprises Syn5-C; or the first synthetic polypeptide comprises Syn2-N and the second synthetic polypeptide comprises Syn9-C; or the first synthetic polypeptide comprises Syn2-N and the second synthetic polypeptide comprises Syn10-C; or the first synthetic polypeptide comprises Syn3-N and the second synthetic polypeptide comprises Syn1-C; or the first synthetic polypeptide comprises Syn3-N and the second synthetic polypeptide comprises Syn4-C; or the first synthetic polypeptide comprises Syn3-N and the second synthetic polypeptide comprises Syn5-C; or the first synthetic polypeptide comprises Syn3-N and the second synthetic polypeptide comprises Syn9-C; or the first synthetic polypeptide comprises Syn3-N and the second synthetic polypeptide comprises Syn10-C; or the first synthetic polypeptide comprises Syn5-N and the second synthetic polypeptide comprises Syn1-C; or the first synthetic polypeptide comprises Syn5-N and the second synthetic polypeptide comprises Syn4-C; or the first synthetic polypeptide comprises Syn5-N and the second synthetic polypeptide comprises Syn5-C; or the first synthetic polypeptide comprises Syn5-N and the second synthetic polypeptide comprises Syn9-C, or The first synthetic polypeptide comprises Syn5-N and the second synthetic polypeptide comprises Syn10-C.

[0043] In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 425, 426, or 427. In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 425, 426, or 427. In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide comprises an amino acid sequence corresponding to SEQ ID NO: 425, 426, or 427. In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide comprises only an amino acid sequence corresponding to SEQ ID NO: 425, 426, or 427. In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 428, 429, 430, 431, or 432. In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 428, 429, 430, 431, or 432. In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide comprises an amino acid sequence corresponding to SEQ ID NO: 428, 429, 430, 431, or 432. In any of the aspects or embodiments thereof provided herein, the synthetic polypeptide comprises only an amino acid sequence corresponding to SEQ ID NO: 428, 429, 430, 431, or 432.

[0044] In any of the aspects or embodiments thereof provided herein, the heterologous polypeptide comprises at least a fragment of a deaminase domain. In any of the aspects or embodiments thereof provided herein, the heterologous polypeptide comprises at least a fragment of a nucleic acid-programmable DNA-binding protein (napDNAbp) domain. In any of the aspects or embodiments thereof provided herein, the heterologous polypeptide is a fragment of a base editor, wherein the base editor comprises a deaminase domain and a nucleic acid-programmable DNA-binding protein domain.

[0045] In any of the aspects or embodiments thereof provided herein, the napDNAbp comprises a Cas9, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, or Cas12j / CasΦ polypeptide, or a functional fragment thereof. In any of the aspects or embodiments thereof provided herein, the napDNAbp is a Cas9 polypeptide, or a functional fragment thereof. In any of the aspects or embodiments thereof provided herein, the napDNAbp is an inactive Cas9 (dCas9) or Cas9 nickase (nCas9). In any of the aspects or embodiments thereof provided herein, the napDNAbp is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus 1 Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), Neisseria meningitidis (NmeCas9), Nme2 Cas9, or a variant thereof. In any of the aspects or embodiments thereof provided herein, the napDNAbp is SpCas9 or a variant thereof. In any of the aspects or embodiments thereof provided herein, the napDNAbp is a variant of SpCas9 with modified protospacer adjacent motif (PAM) specificity. In any of the aspects or embodiments thereof provided herein, the SpCas9 variant recognizes a PAM sequence selected from one or more of NGA, NGCG, NNNRRT, NGCG, NGCN, NGTN, and NGC.

[0046] In any of the aspects or embodiments thereof provided herein, the C-terminus of the heterologous polypeptide is the C-terminal amino acid of a napDNAbp fragment. In any of the aspects or embodiments thereof provided herein, the C-terminal amino acid of the napDNAbp fragment corresponds to an amino acid selected from one or more of amino acids A292 to G364, F445 to K438, and E565 to T637 in the following sequence: spCas9 1 mdkkysigld igtnsvgwav itdeykvpsk kfkvlgntdr hsikknliga llfdsgetae 61 atrlkrtarr rytrrknric ylqeifsnem akvddsffhr leesflveed kkherhpifg 121 nivdevayhe kyptiyhlrk klvdstdkad lrliylalah mikfrghfli egdlnpdnsd 181 vdklfiqlvq tynqlfeenp inasgvdaka ilsarlsksr rlenliaqlp gekknglfgn 241 lialslgltp nfksnfdlae daklqlskdt ydddldnlla qigdqyadlf laaknlsdai 301 llSdilrvnT eiTkaplsas mikrydehhq dltllkalvr qqlpekykei ffdqSkngya 361 gyidggasqe efykfikpil ekmdgteell vklnredllr kqrtfdngsi phqihlgelh 421 ailrrqedfy pflkdnreki ekiltfripy yvgplArgnS rfAwmTrkSe eTiTpwnfee 481 vvdkgasaqs fiermtnfdk nlpnekvlpk hsllyeyftv yneltkvkyv tegmrkpafl <h2 style=";text-align:left;direction:ltr">541 sgeqkkaivd llfktnrkvt vkqlkedyfk kieCfdSvei sgvedrfnAS lgtyhdllki<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 601<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 661 rlsrklingi rdkqsgktil dflksdgfan rnfmqlihdd sltfkediqk aqvsgqgdsl<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 721 hehianlags paikkgilqt vkvvdelvkv mgrhkpeniv iemarenqtt qkgqknsrer<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 781 mkrieegike lgsqilkehp ventqlqnek lylyylqngr dmyvdqeldi nrlsdydvdh<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 841 ivpqsflkdd sidnkvltrs dknrgksdnv pseevvkkmk nywrqllnak litqrkfdnl<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 901 tkaergglse ldkagfikrq lvetrqitkh vaqildsrmn tkydendkli revkvitlks<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 961 klvsdfrkdf qfykvreinn yhhahdayln avvgtalikk ypklesefvy gdykvydvrk<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1021 miakseqeig katakyffys nimnffktei tlangeirkr plietngetg eivwdkgrdf<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1081 atvrkvlsmp qvnivkktev qtggfskesi lpkrnsdkli arkkdwdpkk yggfdsptva<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1141 ysvlvvakve kgkskklksv kellgitime rssfeknpid fleakgykev kkdliiklpk<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1201 yslfelengr krmlasagel qkgnelalps kyvnflylas hyeklkgspe dneqkqlfve<h2 style=";text-align:left;direction:ltr"> 1261 qhkhyldeii eqisefskrv iladanldkv lsaynkhrdk pireqaenii hlftltnlga 1321 paafkyfdtt idrkrytstk evldatlihq sitglyetri dlsqlggd (SEQ ID NO:197) In any of the aspects or embodiments thereof provided herein, the C-terminal amino acid of the napDNAbp fragment corresponds to amino acid position 302, 309, 312, 354, 455, 459, 462, 465, 471, 473, 573, 576, 588, or 589 of SEQ ID NO:197.

[0047] In any of the aspects or embodiments thereof provided herein, the deaminase domain is fused to the N-terminus of the napDNAbp domain.

[0048] In any of the aspects or embodiments thereof provided herein, the C-terminus of the heterologous polypeptide is the C-terminal amino acid of a fragment of the deaminase domain.

[0049] In any of the aspects or embodiments thereof provided herein, the deaminase domain is an adenosine deaminase domain, a cytidine deaminase domain, or a cytidine adenosine deaminase domain. In any of the aspects or embodiments thereof provided herein, the adenosine deaminase domain converts a target A·T to G·C in a polynucleotide. In any of the aspects or embodiments thereof provided herein, the cytidine deaminase domain converts a target C·G to T·A in a polynucleotide. In any of the aspects or embodiments thereof provided herein, the cytidine deaminase domain comprises an APOBEC deaminase domain or a derivative thereof. In any of the aspects or embodiments thereof provided herein, the adenosine deaminase domain is a TadA deaminase domain. In any of the aspects or embodiments thereof provided herein, the adenosine deaminase domain is a TadA*8 mutant or a TadA*9 mutant. In any of the aspects or embodiments thereof provided herein, the adenosine deaminase domain is TadA*8.5. In any of the aspects or embodiments thereof provided herein, the deaminase domain is a cytidine adenosine deaminase domain.

[0050] In any of the aspects or embodiments thereof provided herein, the deaminase domain and the napDNAbp fragment are connected by a linker. In some embodiments, the linker is a peptide linker.

[0051] In any of the aspects or embodiments thereof provided herein, the fusion protein further comprises one or more uracil glycosylase inhibitors (UGIs). In any of the aspects or embodiments thereof provided herein, the fusion protein further comprises two uracil glycosylase inhibitors (UGIs).

[0052] In any of the aspects or embodiments thereof provided herein, the fusion protein further comprises a nuclear localization signal (NLS). In any of the aspects or embodiments thereof provided herein, the NLS is a bipartite NLS (binate nuclear localization signal).

[0053] In any of the aspects or embodiments thereof provided herein, the N-terminus of the heterologous polypeptide is the N-terminal amino acid of a napDNAbp fragment. In any of the aspects or embodiments thereof provided herein, the N-terminal amino acid of the napDNAbp fragment corresponds to an amino acid selected from one or more of amino acids A292 to G364, F445 to K438, and E565 to T637 in the following sequence: spCas9 1 mdkkysigld igtnsvgwav itdeykvpsk kfkvlgntdr hsikknliga llfdsgetae 61 atrlkrtarr rytrrknric ylqeifsnem akvddsffhr leesflveed kkherhpifg 121 nivdevayhe kyptiyhlrk klvdstdkad lrliylalah mikfrghfli egdlnpdnsd 181 vdklfiqlvq tynqlfeenp inasgvdaka ilsarlsksr rlenliaqlp gekknglfgn 241 lialslgltp nfksnfdlae daklqlskdt ydddldnlla qigdqyadlf laaknlsdai 301 llSdilrvnT eiTkaplsas mikrydehhq dltllkalvr qqlpekykei ffdqSkngya 361 gyidggasqe efykfikpil ekmdgteell vklnredllr kqrtfdngsi phqihlgelh <h2 style=";text-align:left;direction:ltr">421 ailrrqedfy pflkdnreki ekiltfripy yvgplArgnS rfAwmTrkSe eTiTpwnfee<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 481 vvdkgasaqs fiermtnfdk nlpnekvlpk hsllyeyftv yneltkvkyv tegmrkpafl<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 541 sgeqkkaivd llfktnrkvt vkqlkedyfk kieCfdSvei sgvedrfnAS lgtyhdllki<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 601<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 661 rlsrklingi rdkqsgktil dflksdgfan rnfmqlihdd sltfkediqk aqvsgqgdsl<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 721 hehianlags paikkgilqt vkvvdelvkv mgrhkpeniv iemarenqtt qkgqknsrer<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 781 mkrieegike lgsqilkehp ventqlqnek lylyylqngr dmyvdqeldi nrlsdydvdh<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 841 ivpqsflkdd sidnkvltrs dknrgksdnv pseevvkkmk nywrqllnak litqrkfdnl<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 901 tkaergglse ldkagfikrq lvetrqitkh vaqildsrmn tkydendkli revkvitlks<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 961 klvsdfrkdf qfykvreinn yhhahdayln avvgtalikk ypklesefvy gdykvydvrk<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1021 miakseqeig katakyffys nimnffktei tlangeirkr plietngetg eivwdkgrdf<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1081 atvrkvlsmp qvnivkktev qtggfskesi lpkrnsdkli arkkdwdpkk yggfdsptva<h2 style=";text-align:left;direction:ltr"> 1141 ysvlvvakve kgkskklksv kellgitime rssfeknpid freakgykev kkdliiklpk 1201 yslfelengr krmlasagel qkgnelalps kyvnflylas hyeklkgspe dneqkqlfve 1261 qhkhyldeii eqisefskrv iladanldkv lsaynkhrdk pireqaenii hlftltnlga 1321 paafkyfdtt idrkrytstk evldatlihq sitglyetri dlsqlggd (SEQ ID NO: 197)

[0054] In any of the aspects or embodiments thereof provided herein, the N-terminal amino acid of the napDNAbp fragment corresponds to amino acid position 303, 310, 313, 355, 456, 460, 463, 466, 472, 474, 574, 577, 589, or 590 of SEQ ID NO:197.

[0055] In any of the aspects or embodiments thereof provided herein, the deaminase domain is fused to the N-terminus of the napDNAbp domain. In any of the aspects or embodiments thereof provided herein, the N-terminus of the heterologous polypeptide is the N-terminal amino acid of a fragment of the deaminase domain.

[0056] In any of the aspects or embodiments thereof provided herein, the N-terminal amino acid of the napDNAbp fragment is Cys substituted for Ala, Ser, or Thr.

[0057] In any of the aspects or embodiments provided herein, the polynucleotide comprises DNA. In any of the aspects or embodiments provided herein, the polynucleotide comprises RNA. In any of the aspects or embodiments provided herein, the polynucleotide further comprises a promoter. In any of the aspects or embodiments provided herein, the promoter is a constitutive promoter. In an embodiment, the constitutive promoter is a CMV promoter or a CAG promoter.

[0058] In any of the aspects or embodiments thereof provided herein, the cell is a mammalian cell. In any of the aspects or embodiments thereof provided herein, the cell is present in a subject. In some embodiments, the subject is a mammal. In some embodiments, the mammal is a human or a non-human primate. In some embodiments, the mammal is a human.

[0059] In any of the aspects or embodiments thereof provided herein, the synthetic polypeptides are capable of mediating binding between polypeptides to which they are fused.

[0060] In any of the aspects or embodiments thereof provided herein, the polynucleotide delivery system further comprises a guide polynucleotide. In any of the aspects or embodiments thereof provided herein, the guide polynucleotide is a single guide RNA (sgRNA) or a polynucleotide encoding the sgRNA. In any of the aspects or embodiments thereof provided herein, the sgRNA is complementary to a target polynucleotide associated with a disease or disorder.

[0061] In any of the aspects or embodiments thereof provided herein, the first polynucleotide and the second polynucleotide are not covalently linked.

[0062] In any of the aspects or embodiments thereof provided herein, the C-terminal amino acid of the N-terminal fragment of the base editor and / or the N-terminal amino acid of the C-terminal fragment of the base editor corresponds to an amino acid in the napDNAbp domain selected from one or more of amino acids A292 to G364, F445 to K438, and E565 to T637 in the following sequences: spCas9 1 mdkkysigld igtnsvgwav itdeykvpsk kfkvlgntdr hsikknliga llfdsgetae 61 atrlkrtarr rytrrknric ylqeifsnem akvddsffhr leesflveed kkherhpifg 121 nivdevayhe kyptiyhlrk klvdstdkad lrliylalah mikfrghfli egdlnpdnsd 181 vdklfiqlvq tynqlfeenp inasgvdaka ilsarlsksr rlenliaqlp gekknglfgn 241 lialslgltp nfksnfdlae daklqlskdt ydddldnlla qigdqyadlf laaknlsdai 301 llSdilrvnT eiTkaplsas mikrydehhq dltllkalvr qqlpekykei ffdqSkngya 361 gyidggasqe efykfikpil ekmdgteell vklnredllr kqrtfdngsi phqihlgelh 421 ailrrqedfy pflkdnreki ekiltfripy yvgplArgnS rfAwmTrkSe eTiTpwnfee 481 vvdkgasaqs fiermtnfdk nlpnekvlpk hsllyeyftv yneltkvkyv tegmrkpafl <h2 style=";text-align:left;direction:ltr">541 sgeqkkaivd llfktnrkvt vkqlkedyfk kieCfdSvei sgvedrfnAS lgtyhdllki<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 601<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 661 rlsrklingi rdkqsgktil dflksdgfan rnfmqlihdd sltfkediqk aqvsgqgdsl<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 721 hehianlags paikkgilqt vkvvdelvkv mgrhkpeniv iemarenqtt qkgqknsrer<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 781 mkrieegike lgsqilkehp ventqlqnek lylyylqngr dmyvdqeldi nrlsdydvdh<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 841 ivpqsflkdd sidnkvltrs dknrgksdnv pseevvkkmk nywrqllnak litqrkfdnl<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 901 tkaergglse ldkagfikrq lvetrqitkh vaqildsrmn tkydendkli revkvitlks<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 961 klvsdfrkdf qfykvreinn yhhahdayln avvgtalikk ypklesefvy gdykvydvrk<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1021 miakseqeig katakyffys nimnffktei tlangeirkr plietngetg eivwdkgrdf<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1081 atvrkvlsmp qvnivkktev qtggfskesi lpkrnsdkli arkkdwdpkk yggfdsptva<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1141 ysvlvvakve kgkskklksv kellgitime rssfeknpid fleakgykev kkdliiklpk<h2 style=";text-align:left;direction:ltr"> <h2 style=";text-align:left;direction:ltr"> 1201 yslfelengr krmlasagel qkgnelalps kyvnflylas hyeklkgspe dneqkqlfve<h2 style=";text-align:left;direction:ltr"> 1261 qhkhyldeii eqisefskrv iladanldkv lsaynkhrdk pireqaenii hlftltnlga 1321 paafkyfdtt idrkrytstk evldatlihq sitglyetri dlsqlggd (SEQ ID NO: 197) In any of the aspects or embodiments thereof provided herein, the C-terminal amino acid of the N-terminal fragment of the base editor corresponds to an amino acid within a napDNAbp domain selected from one or more of 302, 309, 312, 354, 455, 459, 462, 465, 471, 473, 573, 576, 588, or 589 of SEQ ID NO: 197. In any of the aspects or embodiments thereof provided herein, the N-terminal amino acid of the C-terminal fragment of the base editor corresponds to an amino acid within a napDNAbp domain selected from one or more of 303, 310, 313, 355, 456, 460, 463, 466, 472, 474, 574, 577, 589, or 590 of SEQ ID NO: 197. In any of the aspects or embodiments thereof provided herein, the C-terminal amino acid of the N-terminal fragment of the base editor and / or the N-terminal amino acid of the C-terminal fragment of the base editor corresponds to an amino acid within the deaminase domain.

[0063] In any of the aspects or embodiments provided herein, a base editor comprising a fusion of an N-terminal fragment of a base editor and a C-terminal fragment of a base editor has base editing activity. In any of the aspects or embodiments provided herein, fusion of the N-terminal fragment and the C-terminal fragment of a base editor generates a reconstituted full-length base editor.

[0064] In any of the aspects or embodiments thereof provided herein, the first polynucleotide and / or the second polynucleotide encodes one or more uracil glycosylase inhibitors (UGIs) fused to the N-terminal and / or C-terminal fragments of the napDNAbp. In any of the aspects or embodiments thereof provided herein, the first polynucleotide and / or the second polynucleotide encodes two uracil glycosylase inhibitors (UGIs) fused to the N-terminal and / or C-terminal fragments of the napDNAbp.

[0065] In any of the aspects or embodiments thereof provided herein, the first polynucleotide and / or the second polynucleotide encodes a nuclear localization signal (NLS) fused to the N-terminal fragment and / or the C-terminal fragment of the napDNAbp. In any of the aspects or embodiments thereof provided herein, the NLS is a bipartite NLS (binatomic nuclear localization signal).

[0066] In any of the aspects or embodiments thereof provided herein, the N-terminal amino acid of the C-terminal napDNAbp fragment is Cys substituted for Ala, Ser, or Thr. In any of the aspects or embodiments thereof provided herein, the N-terminal amino acid of the C-terminal fragment of the base editor is Cys substituted for Ala, Ser, or Thr.

[0067] In any of the aspects or embodiments thereof provided herein, the synthetic polypeptides are capable of mediating binding between polypeptides to which they are fused.

[0068] In any of the aspects or embodiments thereof provided herein, the method further comprises contacting the cell with a guide polynucleotide. In some embodiments, the guide polynucleotide is a single guide RNA (sgRNA). In some embodiments, the sgRNA is complementary to the target polynucleotide.

[0069] In any of the aspects or embodiments provided herein, the target polynucleotide is associated with a disease or disorder. In any of the aspects or embodiments provided herein, the disease or disorder is a congenital disease or disorder. In any of the aspects or embodiments provided herein, the target polynucleotide is present in the genome of an organism. In some embodiments, the organism is an animal, a plant, or a prokaryotic organism.

[0070] In any of the aspects or embodiments thereof provided herein, the method further comprises contacting the cell with a vector comprising the first polynucleotide and / or a vector comprising the second polynucleotide.

[0071] In any of the aspects or embodiments thereof provided herein, the first polynucleotide and the second polynucleotide are not covalently linked.

[0072] In any of the aspects or embodiments thereof provided herein, the C-terminal amino acid of the N-terminal fragment of the base editor and / or the N-terminal amino acid of the C-terminal fragment of the base editor corresponds to an amino acid in the napDNAbp domain selected from one or more of amino acids A292 to G364, F445 to K438, and E565 to T637 in the following sequences: spCas9 1 mdkkysigld igtnsvgwav itdeykvpsk kfkvlgntdr hsikknliga llfdsgetae 61 atrlkrtarr rytrrknric ylqeifsnem akvddsffhr leesflveed kkherhpifg 121 nivdevayhe kyptiyhlrk klvdstdkad lrliylalah mikfrghfli egdlnpdnsd 181 241 301 llSdilrvnT eiTkaplsas mikridehhq dltllkalvr qqlpekykei ffdqSkngya 361 421 ailrrqedfy pflkdnreki ekiltfripy yvgplArgnS rfAwmTrkSe eTiTpwnfee 481 541 sgeqkkaivd llfktnrkvt vkqlkedyfk kieCfdSvei sgvedrfnAS lgtyhdllki 601 ikdkdfldne enedilediv ltltlfedre mieerlktya hlfddkvmkq lkrrrytgwg 661 rlsrklingi rdkqsgktil dflksdgfan rnfmqlihdd sltfkediqk aqvsgqgdsl 721 781 841 ivpqsflkdd sidnkvltrs dknrgksdnv pseevvkkmk niwrqllnak litqrkfdnl 901 tkaergglse ldkagfikrq lvetrqitkh vaqildsrmn tkydendkli revkvitlks 961 klvsdfrkdf qfykvreinn yhhahdayln avvgtalikk ypklesefvy gdykvydvrk 1021 miakseqeig katakyffys nimnffktei tlangeirkr plietngetg eivwdkgrdf 1081 atvrkvlsmp qvnivkktev qtggfskesi lpkrnsdkli arkkdwdpkk yggfdsptva 1141 ysvlvvakve kgkskklksv kellgitime rssfeknpid freakgykev kkdliiklpk 1201 yslfelengr krmlasagel qkgnelalps kyvnflylas hyeklkgspe dneqkqlfve 1261 qhkhyldeii eqisefskrv iladanldkv lsaynkhrdk pireqaenii hlftltnlga 1321 paafkyfdtt idrkrytstk evldatlihq sitglyetri dlsqlggd (SEQ ID NO: 197)

[0073] In any of the aspects provided herein or embodiments thereof, the C-terminal amino acid of the N-terminal fragment of the base editor corresponds to an amino acid within the napDNAbp domain selected from one or more of 302, 309, 312, 354, 455, 459, 462, 465, 471, 473, 573, 576, 588, or 589 of SEQ ID NO: 197.

[0074] In any of the aspects provided herein or embodiments thereof, the deaminase domain of the base editor is fused to the N-terminus of the napDNAbp domain.

[0075] In any of the aspects or embodiments thereof provided herein, the cell is a hepatocyte, a neuron, a hematopoietic stem cell, an immune cell, or a progenitor cell thereof. In any of the aspects or embodiments thereof provided herein, the cell is a mammalian cell. In any of the aspects or embodiments thereof provided herein, the cell is in vitro or in vivo.

[0076] In any of the aspects or embodiments provided herein, the method achieves a base editing efficiency of at least about 10%. In any of the aspects or embodiments provided herein, the method achieves a base editing efficiency of at least about 30%. In any of the aspects or embodiments provided herein, the method achieves a base editing efficiency of at least about 50%. In any of the aspects or embodiments provided herein, the method achieves a base editing efficiency equal to or greater than the base editing efficiency achieved when the first polynucleotide and the second polynucleotide are replaced with a single polynucleotide encoding a full-length base editor comprising the deaminase domain and napDNAbp. In any of the aspects or embodiments provided herein, the method achieves a base editing efficiency equal to or greater than the base editing efficiency achieved when the first polynucleotide and the second polynucleotide do not encode an N intein or a C intein. In any of the aspects or embodiments provided herein, the synthetic polypeptide comprises an intein.

[0077] In any of the aspects or embodiments thereof provided herein, the vector(s) can cross the blood-brain barrier. In any of the aspects or embodiments thereof provided herein, the AAV vector(s) are AAV9, PHP.EB, PHP.B, AAV.CAP-B10, AAV, CAP-B22, AAV-rh10, or PAL family AAV vectors. In any of the aspects or embodiments thereof provided herein, the AAV vector(s) are generated using a RepCap plasmid containing a Rep2Cap5 V2 nucleotide sequence or a Rep2Cap5 V3 nucleotide sequence. In any of the aspects or embodiments provided herein, the PAL family AAV vector(s) comprise a VP1 capsid polypeptide, wherein the VP1 capsid polypeptide has an amino acid sequence that is at least 95% identical to the following AAV9 VP1 capsid polypeptide amino acid sequence, in which a 7-amino acid-long peptide is inserted between amino acid positions Q588 and A589 relative to the following AAV9 VP1 capsid polypeptide amino acid sequence: TIFF2025533556000001.tif721657 amino acid long peptides are selected from those shown in Table 7B. In any of the aspects provided herein or embodiments thereof, the AAV vector(s) comprise the amino acid mutations A587D and Q588G relative to the AAV9 VP1 capsid polypeptide sequence.

[0078] In any of the aspects or embodiments thereof provided herein, the polynucleotide delivery system further comprises a guide RNA or further comprises a polynucleotide encoding the guide RNA. In any of the aspects or embodiments thereof provided herein, the guide RNA targets a base editor to generate a pathogenic nucleotide modification in an ABCA4 polypeptide in a cell, wherein the pathogenic nucleotide is associated with Stargardt disease.

[0079] In any of the aspects or embodiments thereof provided herein, the cell comprises a single nucleotide polymorphism (SNP) associated with Stargardt disease.

[0080] In any of the aspects or embodiments thereof provided herein, the method results in a pathogenic nucleotide modification in an ABCA4 polypeptide in a cell, wherein the pathogenic nucleotide is associated with Stargardt disease.

[0081] In any of the aspects or embodiments thereof provided herein, the guide RNA comprises a spacer sequence, and the spacer sequence is GUGUCG A AGUUCGCCCUGGAG (SEQ ID NO: 444), GUGUCG G In any of the aspects or embodiments thereof provided herein, the guide RNA comprises a nucleotide sequence selected from one or more of: AGUUCGCCCUGGAG (SEQ ID NO:445), CACCUCUCCAGGGCGAACUUCGACACAGC (SEQ ID NO:446), CACCUCUCCAGGGCGAACUCCGACACAGC (SEQ ID NO:447), and CUCCAGGGCGAACUUCGACACAGC (SEQ ID NO:448), or 1 nt (nucleotide), 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt fragments and / or extensions thereof. AGG TG (SEQ ID NO: 449) and GCTGTGTGCGGAGTTCGCCCTGGAG AGG The base editor can be targeted to edit a nucleobase in a target sequence selected from TG (SEQ ID NO: 450).

[0082] In any aspect or embodiment thereof provided herein, the method is not a process for altering the genetic identity of a human germline.

[0083] definition Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this disclosure belongs. The following references provide those of ordinary skill in the art with general definitions for many of the terms used in this disclosure: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994), The Cambridge Dictionary of Science and Technology (Walker ed., 1988), The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991), and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, unless otherwise specified, the following terms have the meanings indicated below.

[0084] By "ATP-binding cassette, subfamily A, member 4 (ABCA4) polypeptide" is meant a protein having an amino acid sequence with at least about 85% amino acid sequence identity to the following UniProtKB / Swiss-Prot Accession No. P78363.3, or a fragment thereof that has ATPase activity. In one embodiment, the ABCA4 polypeptide contains an amino acid mutation of A1038V, L541P, or G1961E (where positions A1038, L541, and G1961 are indicated in bold underlined text in the sequence below), or a combination thereof. The amino acid sequence of an exemplary ABCA4 protein is shown below: TIFF2025533556000002.tif227165

[0085] "ATP-binding cassette subfamily A member 4 (ABCA4) polynucleotide" refers to a nucleic acid molecule encoding an ABCA4 polypeptide, as well as introns, exons, 3' untranslated regions, 5' untranslated regions, and regulatory sequences associated with its expression, or fragments thereof. In certain embodiments, the ABCA4 polynucleotide is a genomic sequence, cDNA, mRNA, or gene associated with and / or required for the expression of ABCA4. An exemplary ABCA4 polynucleotide sequence from Homo sapiens is shown below (NCBI Reference Sequence Accession No. NM_000350.3). In various embodiments, the ABCA4 polynucleotide comprises a single nucleotide polymorphism (SNP) associated with Stargardt disease. An exemplary ABCA4 gene sequence is provided in Ensembl Accession No. ENSG00000198691 (SEQ ID NO: 442). An exemplary ABCA4 polynucleotide sequence is shown below.

[0086] "Adenine" or "9H-purin-6-amine" has the molecular formula C5H5N5 and the structure [ka] and refers to the purine nucleobase corresponding to CAS number 73-24-5.

[0087] "Adenosine" or "4-amino-1-[(2R,3R,4S,5R)-3,4-dihydroxy-5-(hydroxymethyl)oxolan-2-yl]pyrimidin-2(1H)-one" is a compound attached to a ribose sugar via a glycosidic bond and has the structure [ka] It refers to the adenine molecule having the formula C and corresponding to the CAS number 65-46-3. 10 H 13 It is N5O4.

[0088] "Adenosine deaminase" or "adenine deaminase" refers to a polypeptide or fragment thereof capable of catalyzing the hydrolytic deamination of adenine or adenosine. In some embodiments, the deaminase or deaminase domain is an adenosine deaminase that catalyzes the hydrolytic deamination of adenosine to inosine or deoxyadenosine to deoxyinosine. In some embodiments, the adenosine deaminase catalyzes the hydrolytic deamination of adenine or adenosine in deoxyribonucleic acid (DNA). The adenosine deaminases (e.g., engineered adenosine deaminases, evolved adenosine deaminases) provided herein can be from any organism (e.g., eukaryotes, prokaryotes), including, but not limited to, algae, bacteria, fungi, plants, invertebrates (e.g., insects), and vertebrates (e.g., amphibians, mammals). In some embodiments, the adenosine deaminase is an adenosine deaminase mutant having one or more mutations and capable of deaminating both adenine and cytosine in a target polynucleotide (e.g., DNA, RNA), and can be referred to as a "dual deaminase." Non-limiting examples of dual deaminases include those described in PCT / US22 / 22050. In some embodiments, the target polynucleotide is single-stranded or double-stranded. In some embodiments, the adenosine deaminase mutant is capable of deaminating both adenine and cytosine in DNA. In some embodiments, the adenosine deaminase mutant is capable of deaminating both adenine and cytosine in single-stranded DNA. In some embodiments, the adenosine deaminase mutant is capable of deaminating both adenine and cytosine in RNA. In some embodiments, the adenosine deaminase variant is selected from those described in PCT / US2020 / 018192, PCT / US2020 / 049975, PCT / US2017 / 045381, and PCT / US2020 / 028568, the entire contents of which are incorporated herein by reference in their entirety for all purposes.

[0089] "Adenosine deaminase activity" means catalyzing the deamination of adenine or adenosine to guanine in a polynucleotide. In some embodiments, the adenosine deaminase variants provided herein maintain adenosine deaminase activity (e.g., maintain at least about 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more of the activity of a reference adenosine deaminase (e.g., TadA*8.20 or TadA*8.19)).

[0090] "Adenosine base editor (ABE)" means a base editor that includes an adenosine deaminase.

[0091] By "adenosine base editor (ABE) polynucleotide" is meant a polynucleotide that encodes an ABE. By "adenosine base editor 8 (ABE8) polypeptide" or "ABE8" is meant a base editor as defined herein that comprises an adenosine deaminase, or a base editor as defined herein that comprises an adenosine deaminase mutant having one or more of the mutations listed in Table 5B, one of the combinations of mutations listed in Table 5B, or a mutation at one or more of the amino acid positions listed in Table 5B (where such mutations are relative to the following reference sequence MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID NO: 1) or relative to a corresponding position in another adenosine deaminase). In certain embodiments, ABE8 comprises a mutation at amino acid 82 and / or 166 of SEQ ID NO: 1. In some embodiments, ABE8 comprises additional mutations relative to the reference sequence, as described herein.

[0092] By "adenosine base editor 8 (ABE8) polynucleotide" is meant a polynucleotide that encodes an ABE8 polypeptide.

[0093] As used herein, "administering" refers to providing one or more compositions described herein to a patient or subject. By way of example and not limitation, administration (e.g., injection) of a composition can be performed by intravenous (iv), subcutaneous (sc), intradermal (id), intraperitoneal (ip), or intramuscular (im) injection. One or more such routes can be used. Parenteral administration can be performed, for example, by bolus injection or gradual perfusion over time. In some embodiments, parenteral administration includes intravascular, intravenous, intramuscular, intraarterial, intrathecal, intratumoral, intradermal, intraperitoneal, transtracheal, subcutaneous, subcuticular, intraarticular, subcapsular, intrathecal, and intrasternal infusion or injection. Alternatively, or additionally, administration can be performed by the oral route. In certain embodiments, one or more compositions described herein are administered by subretinal or subfoveal injection. In some cases, subretinal injection results in the formation of a bleb in the fovea.

[0094] By "agent" is meant any small molecule chemical compound, antibody, nucleic acid molecule, or polypeptide, or fragment thereof.

[0095] "Modification" refers to a change in the level, structure, or activity of an analyte, gene, or polypeptide, as detected by standard methods known in the art, e.g., methods described herein. As used herein, modification includes a change (e.g., an increase or decrease) in expression levels. In certain embodiments, the increase or decrease in expression levels is 10%, 25%, 40%, 50%, or more. In some embodiments, the modification includes an insertion, deletion, or substitution (e.g., by genetic engineering) of a nucleic acid base or amino acid.

[0096] By "ameliorating" is meant slowing, inhibiting, attenuating, reducing, arresting, or stabilizing the onset or progression of a disease.

[0097] "Analog" refers to a molecule that has similar, but not identical, functional or structural characteristics. For example, a polypeptide analog retains the biological activity of the corresponding naturally occurring polypeptide while possessing certain biochemical modifications that enhance the analog's function relative to the naturally occurring polypeptide. Such biochemical modifications may, for example, increase the analog's protease resistance, membrane permeability, or half-life without altering ligand binding. Analogs may contain unnatural amino acids.

[0098] "Base editor (BE)" or "nucleobase editor polypeptide (NBE)" refers to an agent that binds to a polynucleotide and has nucleobase-modifying activity. In various embodiments, a base editor comprises a nucleobase-modifying polypeptide (e.g., a deaminase) and a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9 or Cpf1). Exemplary nucleic acid and protein sequences of base editors include those with about 85% sequence identity or at least about 85% sequence identity to any of the base editor sequences set forth in the Sequence Listing, such as those corresponding to SEQ ID NOS: 2-11.

[0099] "BE4 cytidine deaminase (BE4) polypeptide" refers to a base editor that comprises a nucleic acid-programmable DNA-binding protein (napDNAbp) domain, a cytidine deaminase domain, and two uracil glycosylase inhibitor (UGI) domains. In one embodiment, the napDNAbp is a Cas9n (D10A) polypeptide. Non-limiting examples of cytidine deaminase domains include rAPOBEC, ppAPOBEC, RrA3F, AmAPOBEC1, and SsAPOBEC3B.

[0100] By "BE4 cytidine deaminase (BE4) polynucleotide" is meant a polynucleotide that encodes a BE4 polypeptide.

[0101] "Base editing activity" refers to acting to chemically modify a base within a polynucleotide. In one embodiment, a first base is converted to a second base. In one embodiment, the base editing activity is a cytidine deaminase activity, e.g., converting a target C·G to T·A. In another embodiment, the base editing activity is an adenosine deaminase activity or an adenine deaminase activity, e.g., converting A·T to G·C.

[0102] The term "base editor system" refers to an intermolecular complex for editing nucleobases of a target nucleotide sequence. In various embodiments, a base editor (BE) system comprises: (1) a polynucleotide-programmable nucleotide-binding domain, a deaminase domain (e.g., cytidine deaminase or adenosine deaminase) for deaminating nucleobases in a target nucleotide sequence, and (2) one or more guide polynucleotides (e.g., guide RNAs) in combination with the polynucleotide-programmable nucleotide-binding domain. In various embodiments, the base editor (BE) system comprises a nucleobase editor domain selected from adenosine deaminase or cytidine deaminase, and a domain with nucleic acid sequence-specific binding activity. In some embodiments, the base editor system comprises: (1) a base editor (BE) comprising a polynucleotide-programmable DNA-binding domain and a deaminase domain for deaminating one or more nucleobases in a target nucleotide sequence, and (2) one or more guide RNAs in combination with the polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the base editor is a cytidine base editor (CBE). In some embodiments, the base editor is an adenine base editor or an adenosine base editor (ABE). In some embodiments, the base editor is an adenine base editor or an adenosine base editor (ABE), or a cytidine base editor or a cytosine base editor (CBE). In some embodiments, the base editor system (e.g., a base editor system comprising cytidine deaminase) comprises a uracil glycosylase inhibitor or other agent or peptide that inhibits the inosine base excision repair system (e.g., a uracil stabilizing protein as set forth in WO2022015969, the disclosure of which is incorporated herein by reference in its entirety for all purposes).

[0103] The term "Cas9" or "Cas9 domain" refers to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., a protein comprising an active, inactive, or partially active DNA cleavage domain of Cas9 and / or a gRNA-binding domain of Cas9). Cas9 nucleases are also sometimes referred to as casnl nucleases or CRISPR (clustered regularly interspaced short palindromic repeats)-associated nucleases.

[0104] The term "conservative amino acid substitution" or "conservative mutation" refers to the replacement of one amino acid with another amino acid that shares common properties. A functional method for defining common properties between individual amino acids is to analyze the normalized frequency of amino acid changes between corresponding proteins of the same organism (Schulz, GE and Schirmer, RH, Principles of Protein Structure, Springer-Verlag, New York (1979)). Such analysis allows the definition of groups of amino acids in which amino acids within a group preferentially substitute for each other and are therefore most similar to each other in their effect on overall protein structure (Schulz, GE and Schirmer, RH, supra). Non-limiting examples of conservative mutations include amino acid substitutions, such as arginine to lysine and vice versa to maintain a positive charge, aspartic acid to glutamic acid and vice versa to maintain a negative charge, threonine to serine to maintain a free -OH, and asparagine to glutamine to maintain a free -NH.

[0105] Amino acids can generally be classified according to common side chain properties as follows: (1) Hydrophobic: Norleucine, Met, Ala, Val, Leu, He, (2) Neutral hydrophilicity: Cys, Ser, Thr, Asn, Gin, (3) Acidic: Asp, Glu, (4) Basic: His, Lys, Arg, (5) Residues that affect chain orientation: Gly, Pro, (6) Aromatic: Trp, Tyr, Phe.

[0106] In some embodiments, conservative substitutions may involve exchanging a member of any of these classes for another member of the same class, while in some embodiments, non-conservative amino acid substitutions may involve exchanging a member of any of these classes for one of another class.

[0107] The terms "coding sequence" or "protein-coding sequence," used interchangeably herein, refer to a segment of a polynucleotide that encodes a protein. A coding sequence can also be referred to as an open reading frame. The region or sequence is bounded proximal to the 5' end by a start codon and proximal to the 3' end by a stop codon. Stop codons useful in the base editors described herein include TAG, TAA, and TGA.

[0108] A "complex" refers to a combination of two or more molecules whose interaction is due to intermolecular forces. Non-limiting examples of intermolecular forces include covalent and non-covalent interactions. Non-limiting examples of non-covalent interactions include hydrogen bonds, ionic bonds, halogen bonds, hydrophobic bonds, van der Waals interactions (e.g., dipole-dipole interactions, dipole-induced dipole interactions, and London dispersion forces), and the π effect. In one embodiment, the complex comprises a polypeptide, a polynucleotide, or a combination of one or more polypeptides and one or more polynucleotides. In one embodiment, the complex comprises one or more polypeptides that bind to form a base editor (e.g., a nucleic acid-programmable DNA-binding protein such as Cas9, and a base editor comprising a deaminase) and a polynucleotide (e.g., a guide RNA). In one embodiment, the complex is held together by hydrogen bonds. As will be apparent, one or more components of a base editor (e.g., a deaminase or a nucleic acid-programmable DNA-binding protein) can be bound by covalent or non-covalent bonds. As an example, a base editor can comprise a deaminase covalently linked (e.g., by a peptide bond) to a nucleic acid-programmable DNA-binding protein. Alternatively, a base editor can comprise a deaminase and a nucleic acid-programmable DNA-binding protein that are non-covalently linked (e.g., where one or more components of the base editor are provided in trans, bound directly or via another molecule, such as a protein or nucleic acid). In one embodiment, one or more components of the complex are held together by hydrogen bonds.

[0109] "Cytosine" or "4-aminopyrimidin-2(1H)-one" has the molecular formula C4H5N3O and the structure [ka] and refers to the purine nucleobase corresponding to CAS number 71-30-7.

[0110] "Cytidine" is attached to the ribose sugar via a glycosidic bond and has the structure [ka] and corresponds to CAS number 65-46-3. Its molecular formula is CH 13 It is N3O5.

[0111] "Cytidine base editor (CBE)" means a base editor that includes a cytidine deaminase.

[0112] By "cytidine base editor (CBE) polynucleotide" is meant a polynucleotide that encodes a CBE.

[0113] "Cytidine deaminase" or "cytosine deaminase" refers to a polypeptide or fragment thereof capable of deaminating cytidine or cytosine. In certain embodiments, the cytidine or cytosine is present in a polynucleotide. In one embodiment, the cytidine deaminase converts cytosine to uracil or 5-methylcytosine to thymine. The terms "cytidine deaminase" and "cytosine deaminase" are used interchangeably throughout this application. Petromyzon marinus cytosine deaminase 1 (PmCDA1) (SEQ ID NOS: 13-14), activation-induced cytidine deaminase (AICDA) (SEQ ID NOS: 15-21), and APOBEC (SEQ ID NOS: 12-61) are exemplary cytidine deaminases. Further exemplary cytidine deaminase (CDA) sequences are set forth in the Sequence Listing as SEQ ID NOS: 62-66 and 67-189. Non-limiting examples of cytidine deaminases include those described in PCT / US20 / 16288, PCT / US2018 / 021878, 180802-021804 / PCT, PCT / US2018 / 048969, and PCT / US2016 / 058344.

[0114] "Cytosine deaminase activity" means catalyzing the deamination of cytosine or cytidine. In one embodiment, a polypeptide having cytosine deaminase activity converts an amino group to a carbonyl group. In one embodiment, a cytosine deaminase converts cytosine to uracil (i.e., C to U) or 5-methylcytosine to thymine (i.e., 5mC to T). In some embodiments, the cytosine deaminase provided herein has higher cytosine deaminase activity (e.g., at least 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, or more) than a reference cytosine deaminase.

[0115] As used herein, the term "deaminase" or "deaminase domain" refers to a protein or fragment thereof that catalyzes a deamination reaction.

[0116] "Detection" refers to determining the presence, absence, or amount of an analyte to be detected. In one embodiment, a sequence alteration in a polynucleotide or polypeptide is detected. In another embodiment, the presence of an indel is detected.

[0117] "Detectable label" refers to a composition that, when attached to a molecule of interest, renders the molecule of interest detectable by spectroscopic, photochemical, biochemical, immunochemical, or chemical means. For example, useful labels include radioisotopes, magnetic beads, metal beads, colloidal particles, fluorescent dyes, electron-dense reagents, enzymes (e.g., enzymes commonly used in enzyme-linked immunosorbent assays (ELISAs)), biotin, digoxigenin, or haptens.

[0118] "Disease" refers to any illness or disorder that damages or interferes with the normal function of a cell, tissue, or organ. Exemplary diseases include diseases treatable with a base editor, base editor system, or nuclease (e.g., a Cas9 mutant or a Cas12b mutant). In certain embodiments, diseases treatable with the compositions of the present disclosure are associated with point mutations, splicing events, premature stop codons, or misfolding events. Non-limiting examples of diseases include retinitis pigmentosa (RP), Leber's congenital amaurosis (LCA), Stargardt's disease (STGD), Usher's disease (USH), Alström syndrome, congenital non-progressive night blindness (CSNB), macular dystrophy, latent macular dystrophy, diseases caused by mutations in the ABCA4 gene, Duchenne muscular dystrophy, cystic fibrosis, hemophilia A, Wilson's disease, phenylketonuria, dysporinopathy, Rett's syndrome, polycystic kidney disease, Niemann-Pick disease type C, and Huntington's disease. In some cases, the disease is associated with a mutation in a gene selected from one or more of ABCA4, MY07A, CEP290, CDH23, EYS, PCDH15, CACNA1, SNRNP200, RP1, PRPF8, RP1L1, ALMS1, USH2A, GPR98, HMCN1, DMD, CFTR, F8, ATP7B, PAH, DYSF, MECP2, PKD, NPC1, and HTT.

[0119] "Dual editing activity" or "dual deaminase activity" means having adenosine deaminase activity and cytidine deaminase activity. In one embodiment, a base editor with dual editing activity has both A→G and C→T activity, where the two activities are approximately equal or within about 10% or 20% of each other. In another embodiment, in the dual editor, the A→G activity is at most about 10% or 20% higher than the C→T activity. In another embodiment, in the dual editor, the A→G activity is at most about 10% or 20% lower than the C→T activity. In some embodiments, an adenosine deaminase mutant has primarily cytosine deaminase activity, with little, if any, adenosine deaminase activity. In some embodiments, an adenosine deaminase mutant has cytosine deaminase activity and no significant or detectable adenosine deaminase activity.

[0120] By "effective amount" is meant the amount of an agent (e.g., a base editor, a cell) described herein necessary to ameliorate disease symptoms compared to an untreated patient or a disease-free individual, i.e., a healthy individual, or the amount of an agent sufficient to produce a desired biological response. The effective amount of the active compound(s) used to practice embodiments of the present disclosure for the therapeutic treatment of a disease will vary depending on the mode of administration, the age, weight, and general health of the subject. Ultimately, the attending physician or veterinarian will determine the appropriate amount and dosing regimen. Such an amount is referred to as an "effective" amount. In one embodiment, an effective amount is the amount of a base editor of the present disclosure sufficient to introduce a modification into a gene of interest in a cell (e.g., a cell in vitro or in vivo). In one embodiment, an effective amount is the amount of a base editor necessary to achieve a therapeutic effect. Such a therapeutic effect need not be sufficient to modify pathogenic genes in all cells of a subject, tissue, or organ, but need only be sufficient to modify pathogenic genes in about 1%, 5%, 10%, 25%, 50%, 75%, or more of the cells present in the subject, tissue, or organ. In one embodiment, the effective amount is sufficient to ameliorate one or more symptoms of the disease.

[0121] The term "exonuclease" refers to a protein or polypeptide that can remove consecutive nucleotides from either the 5' or 3' end of a polynucleotide.

[0122] The term "endonuclease" refers to a protein or polypeptide capable of catalyzing the cleavage of an internal region of a polynucleotide.

[0123] By "fragment" is meant a portion of a polypeptide or nucleic acid molecule, the portion comprising at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% of the entire length of the reference nucleic acid molecule or polypeptide. A fragment can comprise 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or amino acids. In some embodiments, a fragment is a functional fragment.

[0124] "Guide polynucleotide" refers to a polynucleotide or polynucleotide complex that is specific to a target sequence and can form a complex with a polynucleotide-programmable nucleotide-binding domain protein (e.g., Cas9 or Cpf1). In one embodiment, the guide polynucleotide is a guide RNA (gRNA). The gRNA can exist as a complex of two or more RNAs or as a single RNA molecule.

[0125] "Heterologous" or "exogenous" means that a polynucleotide or polypeptide is 1) experimentally incorporated into a polynucleotide or polypeptide sequence in which it is not normally found in nature, or 2) experimentally placed into a cell that does not normally contain that polynucleotide or polypeptide. In some embodiments, "heterologous" means that a polynucleotide or polypeptide is experimentally placed into a non-native context. In some embodiments, a heterologous polynucleotide or polypeptide is obtained from a first species or host organism and incorporated into a polynucleotide or polypeptide from a second species or host organism. In some embodiments, the first species or host organism is different from the second species or host organism. In some embodiments, the heterologous polynucleotide is DNA. In some embodiments, the heterologous polynucleotide is RNA.

[0126] "Hybridization" refers to hydrogen bonding, which may be Watson-Crick, Hoogsteen, or reversed Hoogsteen hydrogen bonding, between complementary nucleobases. For example, adenine and thymine are complementary nucleobases that pair through the formation of hydrogen bonds.

[0127] By "increase" is meant a positive change of at least 10%, 25%, 50%, 75%, or 100%, or about 1.5-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 7-fold, about 8-fold, about 9-fold, about 10-fold, about 15-fold, about 20-fold, about 25-fold, about 30-fold, about 35-fold, about 40-fold, about 45-fold, about 50-fold, or about 100-fold.

[0128] The terms "inhibitor of base repair," "base repair inhibitor," "IBR," or grammatical equivalents thereof, refer to a protein that is capable of inhibiting the activity of a nucleic acid repair enzyme, e.g., a base excision repair enzyme.

[0129] By "intein" is meant a protein segment that can simultaneously excise itself and join adjacent exteins in a process known as protein splicing. The process by which an intein excises itself and joins the remainder of a protein is referred to herein as "protein splicing" or "intein-mediated protein splicing." In some embodiments, the intein is a trans-splicing intein (also referred to as a "split intein"). In the case of a trans-splicing intein, a full-length polypeptide is split into two separate fragments, with the C-terminus of the N-terminal fragment fused to the N-terminal fragment of the split intein (N intein) and the N-terminus of the remaining C-terminal fragment fused to the C-terminal fragment of the split intein (C intein). Without being bound by theory or mechanism of action, contacting these two polypeptide sequences with each other results in excision of the intein and joining of the two polypeptide sequences to form the full-length polypeptide sequence. In some embodiments, contacting two polypeptide fragments, each fused to an intein fragment or a peptide derived from an intein fragment, results in catalytic activity (e.g., deamination of a nucleobase in a polynucleotide sequence) in a cell that is greater than the catalytic activity observed when the two polypeptide fragments are in contact with each other in a cell without the intein fragment. Non-limiting examples of N and C intein sequences include sequences that share at least 85% sequence identity with an amino acid sequence listed in Table 1A or Table 1B, or functional fragments thereof.

[0130] The terms "isolated," "purified," or "biologically pure" refer to a material that is free, to varying degrees, from components that normally accompany it as found in its native state. "Isolated" indicates some degree of separation from the original source or environment. "Purified" indicates a greater degree of separation than "isolation." A "purified" or "biologically pure" protein is substantially free from other substances, such that impurities do not substantially affect the biological properties of the protein or cause other adverse effects. That is, a nucleic acid or peptide of the present disclosure is purified if, when produced by recombinant DNA technology, it is substantially free from cellular material, viral material, and culture medium, or, when chemically synthesized, it is substantially free from chemical precursors and other chemicals. Purity and homogeneity are typically determined using analytical chemistry techniques, such as polyacrylamide gel electrophoresis or high-performance liquid chromatography. The term "purified" can mean that the nucleic acid or protein gives rise to substantially one band in an electrophoretic gel. For example, in the case of proteins that can be subject to modifications such as phosphorylation or glycosylation, different modifications can result in different isolated proteins that can be separately purified.

[0131] "Isolated polynucleotide" refers to a nucleic acid molecule that is free of the genes adjacent to the nucleic acid molecule of the present disclosure in the naturally occurring genome of the organism from which it is derived. Thus, the term encompasses, for example, recombinant DNA incorporated into a vector, an autonomously replicating plasmid or virus, or integrated into the genomic DNA of a prokaryote or eukaryote, or recombinant DNA that exists as a separate molecule independent of other sequences (e.g., cDNA, or genomic or cDNA fragments produced by PCR or restriction endonuclease digestion). In addition, the term encompasses RNA molecules transcribed from DNA molecules, as well as recombinant DNA that is part of a hybrid gene encoding an additional polypeptide sequence.

[0132] "Isolated polypeptide" refers to a polypeptide of the present disclosure that has been separated from components that naturally accompany it. Typically, a polypeptide is isolated when it is at least 60%, by weight, free from the proteins and naturally-occurring organic molecules with which it is naturally associated. In some embodiments, the preparation is at least 75%, at least 90%, or at least 99%, by weight, a polypeptide of the present disclosure. Isolated polypeptides of the present disclosure can be obtained, for example, by extraction from a natural source, by expression of a recombinant nucleic acid encoding such a polypeptide, or by chemically synthesizing the protein. Purity can be measured by any appropriate method, for example, column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.

[0133] As used herein, the term "linker" refers to a molecule that connects two moieties. In one embodiment, the term "linker" refers to a covalent linker (e.g., a covalent bond) or a non-covalent linker.

[0134] "Marker" refers to any protein or polynucleotide whose expression, level, structure, or activity is altered in association with a disease or disorder. In one embodiment, the disease or disorder is retinitis pigmentosa (RP), Leber's congenital amaurosis (LCA), Stargardt's disease (STGD), Usher's disease (USH), Alström syndrome, congenital non-progressive night blindness (CSNB), macular dystrophy, latent macular dystrophy, diseases caused by mutations in the ABCA4 gene, Duchenne muscular dystrophy, cystic fibrosis, hemophilia A, Wilson's disease, phenylketonuria, dysporinopathy, Rett syndrome, polycystic kidney disease, Niemann-Pick disease type C, or Huntington's disease. In some cases, the marker is selected from one or more of the following polypeptides: ABCA4, MY07A, CEP290, CDH23, EYS, PCDH15, CACNA1, SNRNP200, RP1, PRPF8, RP1L1, ALMS1, USH2A, GPR98, HMCN1, DMD, CFTR, F8, ATP7B, PAH, DYSF, MECP2, PKD, NPC1, and HTT, or polynucleotides encoding same.

[0135] As used herein, the term "mutation" refers to the substitution of a residue in a sequence (e.g., a nucleic acid sequence or an amino acid sequence) with another residue, or the deletion or insertion of one or more residues in a sequence. Mutations are typically indicated herein by identifying the original residue, followed by the position of the residue in the sequence, and then the newly substituted residue. Various methods for making the amino acid substitutions (mutations) indicated herein are well known in the art and can be found, for example, in Green and Sambrook, Molecular Cloning: A Laboratory Manual (4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (2012)).

[0136] As used herein, the terms "nucleic acid" and "nucleic acid molecule" refer to a compound comprising a nucleobase and an acid moiety, e.g., a nucleoside, a nucleotide, or a polymer of nucleotides. Typically, polymeric nucleic acids, e.g., nucleic acid molecules comprising three or more nucleotides, are linear molecules in which adjacent nucleotides are linked to each other via phosphodiester bonds. In some embodiments, "nucleic acid" refers to an individual nucleic acid residue (e.g., a nucleotide and / or a nucleoside). In some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising three or more individual nucleotide residues. As used herein, the terms "oligonucleotide" and "polynucleotide" can be used interchangeably to refer to a polymer of nucleotides (e.g., a chain of at least three nucleotides). In some embodiments, "nucleic acid" encompasses RNA and single- and / or double-stranded DNA. Nucleic acids can occur in nature, for example, when associated with a genome, transcript, mRNA, tRNA, rRNA, siRNA, snRNA, plasmid, cosmid, chromosome, chromatid, or other naturally occurring nucleic acid molecule. On the other hand, nucleic acid molecules include non-naturally occurring molecules, such as recombinant DNA or RNA, artificial chromosomes, genetically engineered genomes, or fragments thereof, as well as synthetic DNA, RNA, DNA / RNA hybrids, or non-naturally occurring nucleotides or nucleosides. Furthermore, the terms "nucleic acid," "DNA," "RNA," and / or similar terms encompass nucleic acid analogs, e.g., analogs having other than a phosphodiester backbone. Nucleic acids can be purified from natural sources, produced using recombinant expression systems, and optionally purified or chemically synthesized. Where appropriate, for example, in the case of chemically synthesized molecules, nucleic acids include nucleoside analogs, such as analogs having chemically modified bases or sugar and backbone modifications. Nucleic acid sequences are presented in a 5' to 3' direction unless otherwise indicated.In some embodiments, nucleic acids are selected from natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleotide analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyl-uridine, C5-propynyl-cytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalated bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose), and / or modified phosphate groups (e.g., phosphorothioate and 5'-N-phosphoramidite linkages).

[0137] The term "nuclear localization sequence," "nuclear localization signal," or "NLS" refers to an amino acid sequence that promotes the import of proteins into the cell nucleus.Nuclear localization sequences are known in the art and are described, for example, in International PCT Application PCT / EP2000 / 011690 filed on November 23, 2000 by Plank et al. (published on May 31, 2001 as WO / 2001 / 038547), the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences.In other embodiments, the NLS is, for example, the optimized NLS described by Koblan et al., Nature Biotech. 2018 doi:10.1038 / nbt.4172. In some embodiments, the NLS comprises the amino acid sequence KRTADGSEFESPKKKRKV (SEQ ID NO: 190), KRPAATKKAGQAKKKK (SEQ ID NO: 191), KKTELQTTNAENKTKKL (SEQ ID NO: 192), KRGINDRNFWRGENGRKTR (SEQ ID NO: 193), RKSGKIAAIVVKRPRK (SEQ ID NO: 194), PKKKRKV (SEQ ID NO: 195), MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 196), PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 328), or RKSGKIAAIVVKRPRKPKKKRKV (SEQ ID NO: 329).

[0138] The terms "nucleobase," "nitrogenous base," or "base" are used interchangeably herein to refer to nitrogen-containing biological compounds that form nucleosides, the building blocks of nucleotides. The ability of nucleobases to base pair and stack with one another directly leads to long-chain helical structures such as ribonucleic acid (RNA) and deoxyribonucleic acid (DNA). The five nucleobases (adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U)) are referred to as primary or standard nucleobases. Adenine and guanine are derived from purines, while cytosine, uracil, and thymine are derived from pyrimidines. DNA and RNA can also contain other (non-primary) modified bases. Non-limiting exemplary modified nucleobases include hypoxanthine, xanthine, 7-methylguanine, 5,6-dihydrouracil, 5-methylcytosine (m5C), and 5-hydroxymethylcytosine. Both hypoxanthine and xanthine can be produced through deamination (replacing an amine group with a carbonyl group) in the presence of mutagens. Hypoxanthine can be modified from adenine. Xanthine can be modified from guanine. Uracil can result from the deamination of cytosine. A "nucleoside" consists of a nucleobase and a five-carbon sugar (either ribose or deoxyribose). Examples of nucleosides include adenosine, guanosine, uridine, cytidine, 5-methyluridine (m5U), deoxyadenosine, deoxyguanosine, thymidine, deoxyuridine, and deoxycytidine. Examples of nucleosides having modified nucleobases include inosine (I), xanthosine (X), 7-methylguanosine (m7G), dihydrouridine (D), 5-methylcytidine (m5C), and pseudouridine (Ψ). A "nucleotide" consists of a nucleobase, a five-carbon sugar (either ribose or deoxyribose), and at least one phosphate group. Non-limiting examples of modified nucleobases and / or chemical modifications that the modified nucleobases may contain include the following:Pseudouridine, 5-methyl-cytosine, 2'-O-methyl-3'-phosphonoacetate, 2'-O-methylthioPACE (MSP), 2'-O-methyl-PACE (MP), 2'-fluoroRNA (2'-F-RNA), constrained ethyl (S-cEt), 2'-O-methyl ("M"), 2'-O-methyl-3'-phosphorothioate ("MS"), 2'-O-methyl-3'-thiophosphonoacetate ("MSP"), 5-methoxyuridine, phosphorothioate, and N1-methylpseudouridine.

[0139] The term "nucleic acid-programmable DNA-binding protein" or "napDNAbp" can be used interchangeably with "polynucleotide-programmable nucleotide-binding domain" and can refer to a protein that binds to a nucleic acid (e.g., DNA or RNA), such as a guide nucleic acid or guide polynucleotide (e.g., gRNA), that guides the napDNAbp to a specific nucleic acid sequence. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable RNA-binding domain. In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a Cas9 protein. The Cas9 protein can bind to a guide RNA that guides the Cas9 protein to a specific DNA sequence complementary to the guide RNA. In some embodiments, the napDNAbp is a Cas9 domain, e.g., nuclease-active Cas9, Cas9 nickase (nCas9), or nuclease-inactive Cas9 (dCas9). Non-limiting examples of nucleic acid programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ (Cas12j / Casphi).Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, and Cas9. (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Cas12j / CasΦ, Cpf1, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4 , Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, homologs thereof, or modified or engineered versions thereof. Other nucleic acid-programmable DNA binding proteins are also within the scope of this disclosure, even if not specifically enumerated herein. See, for example, Makarova et al. "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?" CRISPR J. 2018 Oct;1:325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., "Functionally diverse type V CRISPR-Cas systems" Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science.aav7271 (the entire contents of each are incorporated herein by reference).Exemplary nucleic acid-programmable DNA-binding proteins and nucleic acid sequences encoding nucleic acid-programmable DNA-binding proteins are set forth in the Sequence Listing as SEQ ID NOs: 197-245, 254-260, and 378.

[0140] As used herein, the term "nucleobase editing domain" or "nucleobase editing protein" refers to a protein or enzyme that can catalyze nucleobase modifications in RNA or DNA (e.g., deamination of cytosine (or cytidine) to uracil (or uridine) or thymine (or thymidine), and deamination of adenine (or adenosine) to hypoxanthine (or inosine), as well as non-templated nucleotide addition and insertion). In some embodiments, the nucleobase editing domain is a deaminase domain (e.g., adenine deaminase or adenosine deaminase, or cytidine deaminase or cytosine deaminase).

[0141] As used herein, "obtaining" in "obtaining an agent" includes synthesizing, purchasing, or otherwise obtaining the agent.

[0142] The term "single nucleotide polymorphism (SNP)" refers to a single nucleotide variation occurring at a specific position in the genome, where each variation is present to a measurable extent (e.g., >1%) in a population. For example, at a particular base position in the human genome, a C nucleotide occurs in most individuals, while an A occupies that position in a minority of individuals. This means that there is an SNP at this specific position, and the two possible nucleotide variations, C or A, are said to be alleles at this position. SNPs underlie differences in susceptibility to disease. Disease severity and the body's response to treatment are also manifestations of genetic variation. SNPs can occur within the coding region of a gene, within the noncoding region of a gene, or in intergenic regions (regions between genes). In some embodiments, SNPs within a coding sequence do not necessarily alter the amino acid sequence of the resulting protein due to the degeneracy of the genetic code. SNPs within coding regions are of two types: synonymous and nonsynonymous SNPs. Synonymous SNPs do not affect the protein sequence, while nonsynonymous SNPs alter the amino acid sequence of the protein. There are two types of nonsynonymous SNPs: missense and nonsense. SNPs not located in protein-coding regions can still affect gene splicing, transcription factor binding, messenger RNA degradation, or the sequence of non-coding RNA. Gene expression affected by this type of SNP is called an eSNP (expressed SNP) and can be upstream or downstream of the gene. Single nucleotide variants (SNVs) are single-nucleotide variations with unlimited frequency that can occur somatically. Somatic single-nucleotide variants are also called single-nucleotide mutations.

[0143] "Subject" or "patient" means a mammal. Non-limiting examples of mammals include, but are not limited to, humans or non-human mammals. In certain embodiments, the mammal is a cow, horse, dog, sheep, rabbit, rodent, non-human primate, or cat. In one embodiment, "patient" refers to a mammalian subject who has a higher than average likelihood of developing a disease or disorder. Exemplary patients can be humans, non-human primates, cats, dogs, pigs, cows, cats, horses, camels, llamas, goats, sheep, rodents (e.g., mice, rabbits, rats, or guinea pigs), and other mammals that can benefit from the therapies disclosed herein. Exemplary human patients can be male and / or female.

[0144] As used herein, a "patient in need thereof" or a "subject in need thereof" refers to a patient who has been diagnosed with, is at risk of, has, is predetermined to have, or is suspected of having a disease or disorder.

[0145] The terms "pathogenic mutation," "pathogenic variant," "disease-causing mutation," "pathogenic variant," "deleterious mutation," or "predisposing mutation" refer to a genetic change or mutation that is associated with a disease or disorder, or that increases an individual's susceptibility or predisposition to a particular disease or disorder. In some embodiments, a pathogenic mutation comprises a substitution of at least one wild-type amino acid with at least one pathogenic amino acid in a protein encoded by a gene. In some embodiments, a pathogenic mutation is within a termination region (e.g., a stop codon). In some embodiments, a pathogenic mutation is within a non-coding region (e.g., an intron, promoter, etc.).

[0146] The terms "protein," "peptide," "polypeptide," and their grammatical equivalents are used interchangeably herein to refer to a polymer of amino acid residues joined by peptide (amide) bonds. A protein, peptide, or polypeptide can be natural, recombinant, or synthetic, or any combination thereof.

[0147] As used herein, the term "fusion protein" refers to a hybrid polypeptide comprising protein domains derived from at least two different proteins.

[0148] The term "recombinant" as used herein with respect to a protein or nucleic acid refers to a protein or nucleic acid that is an artificial product that does not occur in nature. For example, in some embodiments, a recombinant protein or recombinant nucleic acid molecule comprises an amino acid sequence or nucleotide sequence that contains at least one, at least two, at least three, at least four, at least five, at least six, or at least seven mutations compared to any naturally occurring sequence.

[0149] By "reduction" is meant a negative change of at least 10%, 25%, 50%, 75%, or 100%.

[0150] "Reference" refers to a standard or control condition. In one embodiment, the reference is a wild-type cell or a normal cell. In other embodiments, without limitation, the reference is an untreated cell that has not been subjected to the test condition or that has been subjected to a placebo or normal saline, medium, buffer, and / or a control vector that does not carry the polynucleotide of interest. Optionally, the reference is a full-size polypeptide (e.g., a base editor). Optionally, the reference is a cell that expresses a full-length polypeptide (e.g., a base editor) from a polynucleotide sequence that does not encode a trans-splicing intein peptide (e.g., an N intein or a C intein). Optionally, the reference is a full-length polypeptide that is expressed as two separate fragments in a cell, where one fragment is an N-terminal fragment of the full-length polypeptide and the other fragment is a C-terminal fragment that includes the remaining C-terminal portion of the full-length polypeptide. Thus, the amino acid sequences of the N-terminal fragment and the C-terminal fragment collectively comprise the complete sequence of the full-length polypeptide, wherein in some embodiments, the N-terminal amino acid of the C-terminal fragment of the polypeptide is replaced with a methionine (M) amino acid, and one or both fragments do not contain a trans-spliced ​​intein peptide (e.g., the N intein or the C intein). In one embodiment, the reference is a cell that does not express the synthetic intein sequences provided herein (see, e.g., the sequences and fragments thereof listed in Tables A-C).

[0151] A "reference sequence" is a predetermined sequence used as a basis for sequence comparison. A reference sequence can be a subset or the entirety of a particular sequence, for example, a segment of a full-length cDNA or gene sequence, or the complete cDNA or gene sequence. For polypeptides, the length of a reference polypeptide sequence is generally at least about 16 amino acids, at least about 20 amino acids, at least about 25 amino acids, at least about 35 amino acids, at least about 50 amino acids, or at least about 100 amino acids. For nucleic acids, the length of a reference nucleic acid sequence is generally at least about 50 nucleotides, at least about 60 nucleotides, at least about 75 nucleotides, at least about 100 nucleotides, or at least about 300 nucleotides, or any integer number of nucleotides thereabout or therebetween. In some embodiments, the reference sequence is the wild-type sequence of a protein of interest. In other embodiments, the reference sequence is a polynucleotide sequence encoding a wild-type protein.

[0152] The terms "RNA-programmable nuclease" and "RNA-guided nuclease" refer to a nuclease that forms a complex with (e.g., binds to or associates with) one or more RNA(s) that are not targets for cleavage. In some embodiments, when an RNA-programmable nuclease is complexed with an RNA, it can be referred to as a nuclease-RNA complex. Typically, the bound RNA(s) are referred to as guide RNAs (gRNAs). In some embodiments, the RNA-programmable nuclease is a (CRISPR-associated system) Cas9 endonuclease, such as Cas9 from Streptococcus pyogenes (Csnl) (e.g., SEQ ID NO: 197), Cas9 from Neisseria meningitidis (NmeCas9, SEQ ID NO: 208), Nme2Cas9 (SEQ ID NO: 209), Streptococcus constellatus (ScoCas9), or a derivative thereof (e.g., a sequence having at least about 85% sequence identity to Cas9, such as Nme2Cas9 or spCas9).

[0153] "Substantially identical" means that a polypeptide or nucleic acid molecule exhibits at least 50% identity to a reference sequence. In one embodiment, the reference sequence is a wild-type amino acid or nucleic acid sequence. In another embodiment, the reference sequence is any one of the amino acid or nucleic acid sequences described herein. In one embodiment, such a sequence is at least about 60%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.9%, or even 99.99% identical at the amino acid or nucleic acid level to the sequence used for comparison.

[0154] Sequence identity is typically measured using sequence analysis software (e.g., the BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs in the sequence analysis software package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. Conservative substitutions typically include substitutions within the following groups: glycine, alanine; valine, isoleucine, leucine; aspartic acid, glutamic acid, asparagine, glutamine; serine, threonine; lysine, arginine; and phenylalanine, tyrosine.

[0155] Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecule encoding a polypeptide of the present disclosure or a functional fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit sufficient identity. A polynucleotide having "sufficient identity" to an endogenous sequence will typically be able to hybridize with at least one strand of a double-stranded nucleic acid molecule. Nucleic acid molecules useful in the methods of the present disclosure include any nucleic acid molecule encoding a polypeptide of the present disclosure or a functional fragment thereof. Such nucleic acid molecules need not be 100% identical to an endogenous nucleic acid sequence, but will typically exhibit sufficient identity. A polynucleotide having "sufficient identity" to an endogenous sequence will typically be able to hybridize with at least one strand of a double-stranded nucleic acid molecule. "Hybridizing" means that the pair forms a double-stranded molecule between complementary polynucleotide sequences (e.g., genes described herein) or portions thereof under various stringency conditions. (See, e.g., Wahl, GM and SL Berger (1987) Methods Enzymol. 152:399; Kimmel, AR (1987) Methods Enzymol. 152:507).

[0156] "Split" means divided into two or more fragments.

[0157] A "split-polypeptide" or "split-protein" refers to a protein that results from an N-terminal fragment and a C-terminal fragment that are translated from a nucleotide sequence(s) as two separate polypeptides. The polypeptides corresponding to the N-terminal and C-terminal portions of the split-protein can, in some embodiments, be spliced ​​to form a "reassembled" protein. In certain embodiments, the split-polypeptide is a nucleic acid-programmable DNA-binding protein (e.g., Cas9) or a base editor.

[0158] The term "target site" refers to a nucleotide sequence or nucleobase of interest within a nucleic acid molecule to be modified. In some embodiments, the modification is base deamination. The deaminase can be a cytidine deaminase or an adenine deaminase. The fusion protein or base editing complex comprising the deaminase can include a dCas9-adenosine deaminase fusion protein, a Cas12b-adenosine deaminase fusion, or a base editor disclosed herein.

[0159] As used herein, terms such as "treat," "treating," and "treatment" refer to reducing or ameliorating a disorder and / or its associated symptoms, or achieving a desired pharmacological and / or physiological effect. As should be apparent, treating a disorder or disease does not necessarily require, although does not preclude, that the disorder, disease, or its associated symptoms be completely eliminated. In some embodiments, the effect is therapeutic, i.e., without limitation, the effect partially or completely reduces, diminishes, suppresses, alleviates, relieves, reduces the intensity of, or cures, the disease and / or adverse symptoms resulting from the disease. In some embodiments, the effect is prophylactic, i.e., the effect protects against or prevents the onset or recurrence of the disease or condition. To this end, the methods disclosed herein comprise administering a therapeutically effective amount of a composition as described herein.

[0160] "Uracil glycosylase inhibitor" or "UGI" refers to an agent that inhibits the uracil excision repair system. Base editors, including cytidine deaminase, convert cytosine to uracil, which is then converted to thymine during DNA replication or repair. In various embodiments, a uracil DNA glycosylase inhibitor (UGI) blocks base excision repair that results in a U back to C. In some examples, contacting a cell and / or polynucleotide with a UGI and a base editor blocks base excision repair that results in a U back to C. Exemplary UGIs include the following amino acid sequences: >splP14739IUNGI_BPPB2 Uracil-DNA glycosylase inhibitor MTNLSDIIEKETGKQLVIQESILMLPEEVEEVIGNKPESDILVHTAYDESTDENVMLLTSDAPEYKPWALVIQDSNGENKIKML (SEQ ID NO: 231)

[0161] In some embodiments, the agent that inhibits the uracil excision repair system is a uracil stabilizing protein (USP). See, e.g., WO2022015969A1, which is incorporated herein by reference.

[0162] As used herein, the term "vector" refers to a means for introducing a nucleic acid sequence into a cell, resulting in a transformed cell. Vectors include plasmids, transposons, phages, viruses, liposomes, lipid nanoparticles, and episomes.

[0163] Ranges provided herein are understood to be shorthand for all values ​​within that range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or subrange from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50.

[0164] Any recitation of a list of chemical groups in a definition of a symbol herein includes a definition of that symbol as any single group or any combination of the listed groups. Any recitation of an embodiment of a symbol, variable, or aspect herein includes any single embodiment or any combination of the embodiment with any other embodiment or portion thereof.

[0165] All terms are intended to be understood as commonly understood by one of ordinary skill in the art. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0166] As used herein, the use of the singular includes the plural unless specifically stated otherwise. It should be noted that as used herein, the singular forms "a," "an," and "the" include plural references unless the context clearly indicates otherwise. As used herein, the use of "or" means "and / or" unless specifically stated otherwise. Furthermore, the use of the term "including," as well as other forms such as "include," "includes," and "included," is meant to be open-ended.

[0167] As used in this specification and claim(s), the words "comprising" (and any of its forms, such as "comprise" and "comprises"), "having" (and any of its forms, such as "have" and "has"), "including" (and any of its forms, such as "includes" and "include"), or "containing" (and any of its forms, such as "contains" and "contain") are inclusive or open-ended. This language indicates the presence of specified elements, features, components, and / or method steps, but does not exclude the presence of other elements, features, components, and / or method steps. Any embodiment that specifies "comprising" particular component(s) or element(s) also contemplates that in some embodiments, "consisting of" or "consisting essentially of" the particular component(s) or element(s). It is contemplated that any embodiment described herein can be implemented with respect to any method or composition of the disclosure, and vice versa. Furthermore, the compositions of the disclosure can be used to achieve the methods of the disclosure.

[0168] The term "about" or "approximately" means that a particular value is within a range of acceptable error as determined by one of ordinary skill in the art, which will depend, in part, on how the value is measured or determined (i.e., the limitations of the measurement system).

[0169] References herein to "some embodiments," "embodiments," "one embodiment," or "other embodiments" mean that a particular feature, structure, or characteristic described in an embodiment is included in at least some embodiments of the present disclosure, but not necessarily in all embodiments. [Brief explanation of the drawings]

[0170] [Figure 1]1 is a bar graph showing the rate of A>G conversion achieved in an experiment (Experiment 1) in which cells were contacted with the base editor (e.g., split base editor) indicated below each bar. In Figure 1, "ABE8.5 full length" indicates an ABE8.5 base editor that does not contain an intein, "CFA" indicates an ABE8.5 base editor split using the CFA split intein, "no intein (T310)" indicates a base editor that was split at position T310 of the CAS9 domain without the use of an intein (i.e., amino acid T310 was the N-terminal amino acid of the C-terminal fragment of CAS9), "GP41.1-SC" indicates an ABE8.5 base editor that was split using the GP41.1-SC split intein, "SSPDNAX-S1" indicates an ABE8.5 base editor that was split using the SSPDNAX-S1 split intein, and "SSPGYRB-S11" indicates an ABE8.5 base editor that was split using the SSPGYRB-S11 split intein. The notation "SYNY-N+SYNX-C" in Figure 1 indicates a base editor split using the "SYNY-N" N intein and the "SNYX-C" C intein, where Y (equal to 2, 3, or 5) and X (equal to 1, 4, 5, 9, or 10) are the identifying numbers of the N intein and the C intein (see Tables 1A and 1B). In Figure 1, the dashed line indicates the percent A>G conversion observed for the ABE8.5 full-length base editor.

[0171] [Figure 2]2 is a bar graph showing the A>G conversion rate achieved in an experiment (Experiment 2). In the experiment, cells were contacted with the base editor (e.g., split base editor) indicated below each bar. In Figure 2, "ABE8.5 full-length" indicates the ABE8.5 base editor without an intein, "CFA" indicates the ABE8.5 base editor split using the CFA split intein, "GP41.1-SC" indicates the ABE8.5 base editor split using the GP41.1-SC split intein, "SSPDNAX-S1" indicates the ABE8.5 base editor split using the SSPDNAX-S1 split intein, and "SSPGYRB-S11" indicates the ABE8.5 base editor split using the SSPGYRB-S11 split intein. The notation "SYNY-N+SYNX-C" in Figure 2 indicates a base editor split using the "SYNY-N" N intein and the "SNYX-C" C intein, where Y (equal to 2, 3, or 5) and X (equal to 1, 4, 5, 9, or 10) are the identifying numbers of the N intein and the C intein (see Tables 1A and 1B). In Figure 2, the dashed line indicates the percent A>G conversion observed for the ABE8.5 full-length base editor.

[0172] [Figure 3A] These are bar graphs juxtaposing the data from Figures 1 and 2. For each pair of bar graphs shown in Figures 3A and 3B, the bar corresponding to Figure 1 (i.e., "Experiment 1") is on the left and the bar corresponding to Figure 2 (i.e., "Experiment 2") is on the right. [Figure 3B] These are bar graphs juxtaposing the data from Figures 1 and 2. For each pair of bar graphs shown in Figures 3A and 3B, the bar corresponding to Figure 1 (i.e., "Experiment 1") is on the left and the bar corresponding to Figure 2 (i.e., "Experiment 2") is on the right.

[0173] [Figure 4]Figure 4 is a bar graph showing the A>G conversion rate achieved in an experiment (Experiment 3). In the experiment, cells were contacted with the base editor (e.g., split base editor) indicated below each bar. In Figure 4, "Cfa" indicates the ABE8.5 base editor split using the Cfa split intein, and "No Intein (T310)" indicates the base editor split at position T310 of the Cas9 domain without the use of an intein (i.e., amino acid T310 was the N-terminal amino acid of the C-terminal fragment of Cas9). The notation "SynY-N + SynX-C" in Figure 4 indicates the base editor split using the "SynY-N" N intein and the "SynX-C" C intein, where Y (equal to 2 or 3) and X (equal to 5, 9, or 10) are the identifying numbers of the N intein and C intein (see Tables 1A and 1B). For each pair of bars in Figure 4, the left bar corresponds to the A7G modification and the right bar corresponds to the A8G modification.

[0174] [Figure 5] 4 is a bar graph showing the A>G conversion rate achieved in an experiment. In the experiment, cells were transfected with an adeno-associated virus (AAV) vector encoding the base editor (e.g., split base editor) shown below each bar. In Figure 4, "Cfa" indicates an ABE8.5 base editor split using the Cfa split intein, "No intein (T310)" indicates a base editor split at position T310 of the Cas9 domain without the use of an intein (i.e., amino acid T310 was the N-terminal amino acid of the C-terminal fragment of Cas9), and "Cfa RbGlob" indicates an ABE8.5 base editor split using the split intein Cfa RbGlob. The notation "SynY-N+SynX-C" in Figure 5 indicates a base editor split using a "SynY-N" N intein and a "SynX-C" C intein, where Y (equal to 2 or 3) and X (equal to 5, 9, or 10) are the identifying numbers of the N intein and the C intein (see Tables 1A and 1B).

[0175] [Figure 6A]1 is a bar graph showing the % A>G base editing rate at the targeted nucleobase (i.e., adenosine at position 8 of the spacer) as measured by next-generation sequencing (NGS) in HEK293T cells transfected with a base editor system comprising the ABE8.5m base editor split using the indicated intein. [Figure 6B] 1 is a bar graph showing the % A>G base editing rate at the targeted nucleobase (i.e., adenosine at position 8 of the spacer) as measured by next-generation sequencing (NGS) in HEK293T cells transfected with a base editor system comprising the ABE8.5m base editor split using the indicated intein. [Figure 6C]6A and 6B are bar graphs showing the percentage of A>G base editing at the targeted nucleobase (i.e., adenosine at position 8 of the spacer) as measured by next-generation sequencing (NGS) in HEK293T cells transfected with a base editor system containing the ABE8.5m base editor split using the indicated intein. The base editor was split at position 309 within the Cas9 domain, resulting in 309 as the C-terminal amino acid of the N-extein and 310 as the N-terminal amino acid of the C-extein. Data corresponding to each of Figures 6A and 6B were collected in two separate experiments, each performed by a different individual, serving as biological replicates for the other. Data from the first experiment corresponds to Figure 6A, and data from the second experiment corresponds to Figure 6B. Similarly, data corresponding to Figure 6C were collected in two separate experiments, each performed by a different individual, serving as biological replicates for the other. The data corresponding to the first five columns from the left of the bar graph correspond to the first experiment, and the data corresponding to the last five columns from the left of the bar graph correspond to the second experiment. In Figure 6C, the term "UNT" indicates "untreated," "CFA N" indicates the Cfa N intein from Table 2A, "CFA C" indicates the Cfa C intein from Table 2C, "3N" indicates Syn3-N from Table 1A, "5C" indicates Syn5-C from Table 1B, "CFA N only" indicates a split base editor whose N-terminal fragment has a CFA N peptide fused to its C-terminus and no peptide fused to its C-terminal fragment, "CFA C only" indicates a split base editor whose C-terminal fragment has a CFA C peptide fused to its N-terminus and no peptide fused to its N-terminal fragment, and "CFA N+C" indicates a split base editor whose N-terminal fragment has a CFA N peptide fused to its C-terminus and no peptide fused to its C-terminal fragment."3N only" refers to a split base editor having a 3N peptide fused to the C-terminus of its N-terminal fragment and no peptide fused to the C-terminal fragment; "3N+5C" refers to a split base editor having a 3N peptide fused to the C-terminus of its N-terminal fragment and a 5C peptide fused to the N-terminus of its C-terminal fragment; "3N+9C" refers to a split base editor having a 3N peptide fused to the C-terminus of its N-terminal fragment and a 9C peptide fused to the N-terminus of its C-terminal fragment; and "5C only" refers to a split base editor having a 3N peptide fused to the C-terminus of its C-terminal fragment and no peptide fused to the N-terminus of its C-terminal fragment. "9C only" indicates a split base editor with a 5C peptide fused to its N-terminal fragment and no peptide fused to its N-terminal fragment; "No Intein N" indicates the N-terminal fragment of a split base editor not fused to any peptide; "No Intein" indicates the C-terminal fragment of a split base editor not fused to any peptide (e.g., an intein); "No Intein N+C" indicates a split base editor with no peptide fused to its N-terminal or C-terminal fragment; and "Full Length" indicates a full-length (i.e., unsplit) base editor. All of the datasets in Figures 6A-6C showed similar base editing rate patterns and similar editing rates between synthetic inteins (i.e., 3N, 5C, and 9C) and Cfa inteins (i.e., CFA N and CFA C).

[0176] [Figure 7A] 1 is a Western blot image showing the level and molecular weight of HEK293T cells expressing the indicated polypeptide(s). [Figure 7B]Figure 7A shows Western blot images showing the levels and molecular weights of HEK293T cells expressing the indicated polypeptide(s). In Figure 7A, the Western blot was stained using an antibody specific for the C-terminal portion of SpCas9 (Abcam189380 [EPR18991]) and an anti-rabbit monoclonal antibody (1:1000 dilution) as the secondary antibody. Samples transfected with all synthetic N+C fragments (i.e., 3N+5C and 3N+9C, corresponding to rows 5 and 6, respectively) show the full-length Cas9 band (higher molecular weight), while samples transfected with all C fragments show a shorter band size (lower molecular weight). In Figure 7B, the Western blot was stained using an antibody specific for the N-terminal portion of SpCas9 (CST14697 [7A9-3A3]) and an anti-mouse monoclonal antibody (1:1000 dilution) as the secondary antibody. Samples transfected with all synthetic N+C (i.e., 3N+5C and 3N+9C, corresponding to rows 5 and 6, respectively) show a full-length Cas9 band (higher molecular weight), while samples transfected with all N-split products show shorter band sizes (lower molecular weight). In Figures 7A and 7B, the numbers above each lane indicate the following samples: 1) CFA N only, 2) CFA C only, 3) CFA N+C, 4) 3N only, 5) 3N+5C, 6) 3N+9C, 7) 5C, 8) 9C, 9) N without intein, 10) C without intein, 11) N without intein, 12) full-length (i.e., unsplit), and 13) untreated. Here, "CFA N" refers to the Cfa N intein in Table 2A, "CFA C" refers to the Cfa C intein in Table 2C, "3N" refers to Syn3-N in Table 1A, "5C" refers to Syn5-C in Table 1B, "CFA N only" refers to a split base editor in which the N-terminal fragment has a CFA N peptide fused to its C-terminus and no peptide fused to its C-terminal fragment, "CFA C only" refers to a split base editor in which the C-terminal fragment has a CFA C peptide fused to its N-terminus and no peptide fused to its N-terminal fragment, and "CFA"N+C" indicates a split base editor having a CFA N peptide fused to the C-terminus of its N-terminal fragment and a CFA C peptide fused to the N-terminus of its C-terminal fragment; "9C" indicates Syn9-C in Table 1B; "3N only" indicates a split base editor having a 3N peptide fused to the C-terminus of its N-terminal fragment and no peptide fused to the C-terminal fragment; "3N+5C" indicates a split base editor having a 3N peptide fused to the C-terminus of its N-terminal fragment and a 5C peptide fused to the N-terminus of its C-terminal fragment; "3N+9C" indicates a split base editor having a 3N peptide fused to the C-terminus of its N-terminal fragment and a 9C peptide fused to the N-terminus of its C-terminal fragment; and "5C only" indicates a split base editor having a C-terminal fragment fused to the C-terminus of its C-terminal fragment. 7A and 7B, glyceraldehyde-3-phosphate dehydrogenase (GAPDH) was stained as a control for normalizing protein concentration.

[0177] [Figure 8]This figure shows an image of a denaturing gel used to evaluate the packaging efficiency of AAV particles. The chart on the left side of the image indicates the RepCap plasmids used to prepare the AAV particles corresponding to each lane of the denaturing gel. AAV8 ("reference AAV") and PHP.eBAAV particles were run on the denaturing gel as controls. The "reference AAV" contained a polynucleotide encoding green fluorescent protein (GFP) rather than a polynucleotide encoding the base editing system or its components. The terms "N-split" and "C-split" refer to the respective halves of the split base editor, which were fused together via protein splicing in cells using split inteins to form the full-length base editor.

[0178] [Figure 9] 9 is a bar graph showing the A8G base editing rate measured in HEK293T cells transduced with a base editor system comprising a split base editor and a guide polynucleotide targeting the ABCA4 G1961 codon for base editing. In Figure 9, the left bar of each pair of bar graphs corresponds to a multiplicity of infection (MOI) of 1e6, and the right bar corresponds to an MOI of 5e6. In Figure 9, the terms "CMF" and "CBA" refer to the promoters used to express the base editing system in HEK293T cells.

[0179] [Figure 10] 10 is a bar graph showing the A8G base editing rate measured in G1961E lentiviral-integrated HEK293T cells transduced with a base editor system comprising a split base editor and a guide polynucleotide targeting the ABCA4 G1961 codon for base editing. In Figure 10, the left bar of each pair of bar graphs corresponds to a multiplicity of infection (MOI) of 1e6, and the right bar corresponds to an MOI of 5e6. In Figure 10, the terms "CMF" and "CBA" refer to the promoters used to express the base editing system in HEK293T cells.

[0180] [Figure 11A]FIG. 1 shows A·T to G·C conversions in the ABCA4 G1961E allele by ABE7.10 and ABE8 mutants in model cell lines. [Figure 11B] Figure 11A shows A·T to G·C conversion by ABE7.10 and ABE8 mutants in the ABCA4 G1961E allele in model cell lines. Figure 11A: A·T to G·C conversion in HEK293T cells with the integrated disease allele and wobble base of the ABCA4 G1961E codon after plasmid lipofection of a 21-nucleotide spacer sgRNA and base editor mutants. Cells were incubated for 5 days after lipofection and then assessed for editing. Figure 11B: DNA sequence at the site of interest, including the ABCA4G 1961E disease allele, the wobble base of the codon, and the -NGG PAM used by the 21-nucleotide spacer sgRNA. Error bars represent the standard deviation of three replicates. In each dataset, the disease allele is on the left and the wobble base is on the right. FIG. 11B shows, in order of appearance, the nucleotide sequence GCTGTGTGTCGAAGTTCGCCCTGGAGAGGTG (SEQ ID NO: 449) and the amino acid sequence LCVEVRPGEV (SEQ ID NO: 453), respectively.

[0181] [Figure 12] Figure 1 shows A·T to G·C conversion by sgRNA spacer length variants in the ABCA4 G1961E allele in a model cell line. A·T to G·C conversion of the ABCA4 G1961E codon with the integrated disease allele and wobble base in HEK293T cells after lipofection of sgRNAs with various spacer lengths and the ABE7.10 plasmid. Cells were incubated for 5 days after lipofection and then assessed for editing. hRz = integration of a self-cleaving hammerhead ribozyme into the 5' end of the sgRNA. Error bars represent the standard deviation of three replicates. In each dataset, the disease allele is on the left and the wobble base is on the right.

[0182] [Figure 13]This diagram shows dual AAV delivery of a split base editor using split intein reconstitution. Two AAV particles are separately packaged with the components required for base editing. One virus encodes the C-terminal region of the base editor with an N-terminal split intein fusion, while the complementary virus encodes the N-terminal region of the base editor with a C-terminal split intein fusion and an sgRNA. Co-transduction of the complementary viruses transcribes the sgRNA, expressing each half of the base editor, which then recombine by split intein-mediated protein trans-splicing.

[0183] [Figure 14A] Figure 1 shows A·T to G·C conversion by dual AAV delivery of the split ABE mutant at ABCA4 G1961 in wild-type cells. [Figure 14B] Figure 14A shows the A·T to G·C conversion of the split ABE mutant at ABCA4 G1961 in wild-type cells by dual AAV delivery. Figure 14A: A·T to G·C conversion and C·G to T·A conversion at the wild-type ABCA4 G1961 target site in wild-type ARPE-19 cells. Here, editing position 8A serves as an alternative target for editing in these cells. Cells were infected with the viral genome at a multiplicity of infection (MOI) of 5E+4 per virus. Cells were incubated for 2 weeks post-infection and then assessed for editing. Error bars represent the standard deviation (sd) of six replicates. For each data point, the sample treated with the alternative site at position 8 (A>G) is shown on the left, and the sample treated with the alternative site at position 5 (C>T) is shown on the right. Figure 14B: DNA sequence of the wild-type target site, including the -NGG PAM and ABCA4 G1961 allele utilized by the 21-nucleotide spacer sgRNA targeting the wild-type sequence. Figure 14B shows, in order of appearance, the nucleotide sequence GCTGTGTGTCGGAGTTCGCCCTGGAGAGGTG (SEQ ID NO:450) and the amino acid sequence LCVGVRPGEV (SEQ ID NO:454), respectively. In each pair of bar graphs in Figure 14A, the left bar corresponds to position 8 and the right bar corresponds to position 5.

[0184] [Figure 15A] FIG. 1 shows off-target base editing in wild-type ARPE-19 cells dually infected with AAV2 expressing split ABE7.10 and an sgRNA targeting the ABCA4 G1961E disease allele. [Figure 15B] Figure 15 shows off-target base editing in wild-type ARPE-19 cells dually infected with AAV2 expressing split ABE7.10 and sgRNA targeting the ABCA4 G1961E disease allele. Figure 15A: Maximum A·T to G·C conversion across the target or off-target protospacer compared to untreated controls two weeks after coinfection with the dual AAV. Figure 15B: Maximum A·T to G·C non-conversion across the target or off-target protospacer compared to untreated controls two weeks after coinfection with the dual AAV. For each data point, treated wild-type (wt) ARPE-19 cells are shown on the left, and untreated wtARPE-19 cells are shown on the right.

[0185] [Figure 16]

[0039] Figure 1 shows the generation of indels due to base editing in wild-type ARPE-19 cells dually infected with AAV2 expressing split ABE7.10 and an sgRNA targeting the ABCA4 G1961E disease allele. Percentage of indels generated within or proximal to the on-target or off-target protospacer compared to untreated controls two weeks after dual AAV co-infection. For each data point, treated wild-type (wt) ARPE-19 cells are shown on the left, and untreated wtARPE-19 cells are shown on the right.

[0186] [Figure 17]This figure shows the integrity and GFP expression of primate retina at 22 days in culture. Sections were immunolabeled overnight at 4°C with anti-rhodopsin, anti-GFP, and biotinylated peanut agglutinin antibodies. Anc80L65.hGRK.eGFP demonstrated that GFP was only found in the outer nuclear layer (ONL), which contains photoreceptors, confirming photoreceptor-specific activity of the GRK promoter. Row 1: Day 0, untransduced. Row 2: Day 22, untransduced. Row 3: Day 22, GRK. Row 4: Day 22, CMB. Columns are unstained (column 1), DAPI (column 2), GFP (column 3), PNA (column 4), and rhodopsin (column 5).

[0187] [Figure 18] Figure 1. Expression of Cas9 in NHPs. Cas9 expression is detected in the primate retina as early as day 6 after culture. Results are shown for ABE7.10 (rows 1 and 2), ABE8.5 (rows 2 and 3), and ABE8.9 (rows 3 and 4). Top row: day 6 after culture. Bottom row: day 17 after culture. Results demonstrate that the AAV system delivers split intein expressing Cas9. Scale bar: 100 μm. DETAILED DESCRIPTION OF THE INVENTION

[0188] The present disclosure features synthetic polypeptides comprising trans-splicing inteins, functional fragments thereof, and polynucleotides encoding them, for use in systems, compositions, kits, and methods for delivering one or more polynucleotides (e.g., polynucleotides encoding split-polypeptides) to cells using vectors with limited packaging capacity (e.g., viral vectors such as adeno-associated viral vectors).

[0189] The present disclosure is based, at least in part, on the discovery that synthetic trans-splicing intein sequences can be used to "split" a base editor into fragments that can be reconstituted into full-length polypeptides within a cell. Polynucleotides encoding the N- and C-terminal fragments of a base editor can be fused to the N and C inteins of a trans-splicing intein (or "split intein") pair, respectively, and delivered to a cell in separate vectors (e.g., two separate adeno-associated virus (AAV) vectors). While not intending to be bound by the theory or mechanism of action proposed herein, in certain embodiments, the encoded base editor fragments, when co-expressed within a cell, are spliced ​​together to reconstitute a full-length, functional base editor polypeptide useful, among other things, for modifying nucleobases within a polynucleotide sequence.

[0190] Intein Inteins (intervening proteins) are auto-processing domains present in a wide variety of organisms that carry out a process known as protein splicing. Protein splicing is a multi-step biochemical reaction consisting of both the cleavage and formation of peptide bonds. The endogenous substrates for protein splicing are proteins present in the intein-containing organism, but inteins can also be used to chemically engineer virtually any polypeptide backbone.

[0191] In protein splicing, an intein excises itself from a precursor polypeptide by cleaving two peptide bonds, thereby joining adjacent extein (external protein) sequences through the formation of new peptide bonds. Intein-mediated protein splicing requires only folding of the intein domain and occurs spontaneously.

[0192] Approximately 5% of inteins are split inteins, which are transcribed and translated as an N intein component and a C intein component, where each component of the full-length intein is fused to a different extein. The N intein is fused to the C-terminus of one extein, and the C intein is fused to the N-terminus of another extein. Without being bound by theory or proposed mechanism of action, in some embodiments, during translation, the N and C inteins spontaneously non-covalently assemble into the canonical intein structure and carry out protein splicing in trans. For this reason, split inteins are often referred to as "trans-splicing inteins."

[0193] In some embodiments, the N intein and C intein can be fused to the N-terminal fragment of a base editor (e.g., within a nucleic acid-programmable DNA-binding domain, such as a Cas9 domain) and the remaining C-terminal fragment of the base editor, respectively, thereby allowing for the conjugation or functional coupling of the N-terminal portion of the base editor (extein) with the C-terminal portion of the base editor (another extein) (e.g., coupling deamination, nuclease, nickase, and / or nucleic acid-programmable DNA-binding activities to each other to modify nucleobases within a polynucleotide sequence). For example, in some embodiments, the N intein is fused to the C-terminus of the N-terminal fragment (extein) of the split-polypeptide, i.e., forming the structure N-[N-terminal fragment of the split-polypeptide (e.g., base editor)]-[N intein]-C. In some embodiments, the C intein is fused to the N-terminus of the C-terminal fragment (another extein) of the split-polypeptide, i.e., forming the structure N-[C intein]-[C-terminal fragment of the split-polypeptide]-C. Without intending to be bound by theory, i.e., the mechanism of intein-mediated protein splicing that binds exteins, inteins are described in Shah et al., Chem Sci. 2014;5(1):446-461 (incorporated herein by reference). Methods for designing and using inteins are described, for example, in WO2014004336, WO2017132580, WO2013045632, US20150344549, and US20180127780, the disclosures of each of which are incorporated herein by reference in their entirety for all purposes.

[0194] In certain embodiments, the base editor is split into two fragments (i.e., an N-terminal fragment and a C-terminal fragment) within its Cas9 domain. In some embodiments, the Cas9 polypeptide is split into two fragments (i.e., an N-terminal fragment and a C-terminal fragment). In certain embodiments, the two fragments associate within a disordered region of Cas9, as described, for example, in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014 or Jiang et al. (2016) Science 351: 867-871. PDB file: 5F9R (each of which is incorporated by reference herein in its entirety for all purposes). The disordered region can be determined by one or more protein structure determination methods known in the art. Such methods include, but are not limited to, X-ray crystallography, NMR spectroscopy, electron microscopy (e.g., cryo-EM), and / or in silico protein modeling. In some embodiments, the base editor or Cas9 polypeptide is split into two fragments at any C, T, A, or S within the region approximately between amino acids A292 and G364, between F445 and K483, or between E565 and T637, referenced to the SpCas9 amino acid sequence, or at a corresponding position in any other Cas9, Cas9 variant (e.g., nCas9, dCas9), or other napDNAbp, where the C, T, A, or S corresponds to the N-terminal amino acid in the C-terminal fragment of the split-polypeptide. In some embodiments, the base editor or Cas9 is split into two fragments at positions T310, T313, A456, S469, or C574, referenced to the SpCas9 amino acid sequence, or at a corresponding position in other Cas9, Cas9 variant, or other napDNAbp, where the amino acid position corresponds to the N-terminal amino acid in the C-terminal fragment of the split-polypeptide. The process of dividing a protein into two fragments is called "splitting" the protein.In various examples, the N-terminal fragment is fused at its C-terminus to an N-intein, and the C-terminal fragment is fused at its N-terminus to a C-intein, where the N-terminal amino acid of the C-terminal fragment corresponds to a position selected from S303, T310, T313, S355, A456, S460, A463, T466, S469, T472, T474, C574, S577, A589, and S590, shown in uppercase and bold in the spCas9 amino acid sequence below, or corresponds to a corresponding position in another of the nucleic acid-programmable DNA-binding protein amino acid sequences. In one embodiment, the C-terminal amino acid of the N-terminal fragment corresponds to a position selected from 302, 309, 312, 354, 455, 459, 462, 465, 471, 473, 573, 576, 588, or 589 with reference to the following sequences, or to a corresponding position in another of the nucleic acid-programmable DNA-binding protein amino acid sequences. TIFF2025533556000007.tif149165

[0195] In some embodiments, a portion or fragment of a polypeptide (e.g., a Cas9 mutant, a nuclease, or a base editor) is fused to an N intein or a C intein, or a functional variant thereof. The polypeptide can be fused to the N-terminus or C-terminus of the N intein or C intein, or a functional variant thereof. In some embodiments, a fragment of a polypeptide is fused to an N intein or a C intein, which is further fused to an AAV capsid protein. The N intein or C intein, polypeptide, and capsid protein can be fused to each other in any configuration (e.g., nuclease-C / N intein-capsid, C / N intein-nuclease-capsid, capsid-C / N intein-nuclease, etc.). In some embodiments, an N-terminal fragment of a polypeptide is fused to an N intein or a functional variant thereof, and a C-terminal fragment of a polypeptide is fused to a C intein or a functional variant thereof. In some embodiments, an N-terminal fragment of a base editor (e.g., ABE, CBE, CABE) is fused to an N-intein or a functional variant thereof, and a C-terminal fragment is fused to a C-intein or a functional variant thereof. In some embodiments, an N-terminal fragment of a nucleic acid-programmable DNA-binding protein (napDNAbp) (e.g., Cas9) is fused to an N-intein or a functional variant thereof, and a C-terminal fragment is fused to a C-intein or a functional variant thereof. In some embodiments, an N-terminal fragment of a deaminase domain (e.g., adenosine deaminase, cytidine deaminase, or dual deaminase) is fused to an N-intein or a variant thereof, and a C-terminal fragment is fused to a C-intein or a functional variant thereof.

[0196] In various embodiments, the methods of the disclosure comprise modifying a nucleobase of a polynucleotide in a cell by contacting the cell with a base editor split into two fragments (i.e., an N-terminal fragment and a C-terminal fragment), where each fragment is fused to a C-intein or an N-intein, a functional variant thereof, or a polynucleotide encoding same, and a guide polynucleotide or a polynucleotide encoding same. The N-intein or functional variant thereof is fused to the C-terminus of the N-terminal fragment of the base editor (e.g., within a nucleic acid-programmable DNA-binding domain such as a Cas9 domain), and the C-intein or functional variant thereof is fused to the N-terminus of the remaining C-terminal fragment of the base editor. When the cell is contacted with the two fragments, the base editing rate is about 10% or more, about 15% or more, about 20% or more, about 25% or more, about 30% or more, about 35% or more, about 40% or more, about 45% or more, about 50% or more, about 55% or more, about 60% or more, about 65% or more, about 70% or more, about 75% or more, about 80% or more, about 85% or more, about 90% or more, or about 95% or more. In some embodiments, contacting a cell with the two fragments results in a rate of base editing that is about 1% or more, about 5% or more, about 10% or more, about 15%, about 20% or more, about 25% or more, about 30% or more, about 35% or more, about 40% or more, about 45% or more, about 50% or more, about 55% or more, about 60% or more, about 65% or more, about 70% or more, about 75% or more, about 80% or more, about 85% or more, about 90% or more, or about 95% or more than that for a cell contacted with two fragments of a base editor that do not comprise the N or C intein or functional variants thereof, or a polynucleotide encoding same. In some embodiments, contacting a cell with the two fragments results in a base editing rate that is within about 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, or 50% of the base editing rate obtained with a cell contacted with a full-length, non-split base editor or a polynucleotide encoding same.

[0197] In some embodiments, the N-intein and C-intein, or functional variants thereof, fused to the C- and N-terminal fragments of the two split polynucleotides, respectively, mediate binding between the N- and C-terminal fragments, sufficient to measure activity (e.g., base editing activity) associated with the full-length polynucleotide in cells expressing the two protein fragments fused to the N- and C-inteins. Optionally, this binding is non-covalent.

[0198] In some embodiments, the C intein does not contain the amino acid sequence GEP. In some embodiments, the C intein does not contain a mutation from the EKD amino acid sequence to the GEP amino acid sequence (e.g., an "EKD" to "GEP" loop mutation at residues 122-124 of the Cfa intein).

[0199] In some embodiments, the polypeptide fragments fused to the C intein or N intein are packaged into two or more AAV vectors. In some embodiments, the N-terminus of the C intein or N intein is fused to the C-terminus of the fusion protein, and the C-terminus of the C intein or N intein is fused to the N-terminus of the AAV capsid protein.

[0200] In one embodiment, a trans-splicing intein is utilized to link fragments or portions of a cytidine editor protein, an adenosine editor protein, or a dual base editor protein that are grafted onto an AAV capsid protein. The use of specific inteins to link heterologous protein fragments is described, for example, in Wood et al., J. Biol. Chem. 289(21);14512-9 (2014). For example, when fused to separate protein fragments, the inteins IntN and IntC recognize each other, splice out, and simultaneously join the adjacent N- and C-terminal exteins of the protein fragments to which they are fused, thereby reconstituting a full-length protein from the two protein fragments.

[0201] In some embodiments, the base editor or nucleic acid-programmable DNA-binding protein (napDNAbp) is split into N- and C-terminal fragments at Ala, Ser, Thr, or Cys residues within selected regions of SpCas9, where these regions correspond to loop regions identified by Cas9 crystallography, and the Ala, Ser, Thr, or Cys residue corresponds to the N-terminal amino acid of the base editor or C-terminal fragment of the napDNAbp.

[0202] In some embodiments, the N intein or C intein is a synthetic N intein or synthetic C intein. Non-limiting examples of synthetic N inteins and C inteins include peptides having about 60% or more, about 65% or more, about 70% or more, about 75% or more, about 80% or more, about 85% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or 100% sequence identity to a sequence set forth in any of Tables 1A and 1B, or functional fragments thereof. Optionally, the N intein or C intein is truncated and / or extended by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 amino acids at the N-terminus and / or C-terminus. In some embodiments, the C intein is truncated at the N-terminus by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids, with the new N-terminal amino acid replaced with a methionine (M). In some embodiments, the C intein is extended at the N-terminus by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids, with the new N-terminal amino acid being a methionine (M). Optionally, the C intein is 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 amino acids in length. In some embodiments, the N intein is truncated at the C-terminus by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the N intein is extended at the C-terminus by 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some cases, the N intein is 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 amino acids in length. In various embodiments, any of the N inteins provided herein can be used in combination with any of the C inteins provided herein to trans-splice two extein sequences to form a full-length polypeptide sequence.

[0203] In various embodiments, the N-terminal portion of the split-polypeptide is fused at its C-terminus to an N intein, and the remaining C-terminal portion of the split-polypeptide is fused at its N-terminus to a C intein. In certain embodiments, the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn2-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn1-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn2-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn4-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn2-N. The C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn5-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn2-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn9-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn2-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn10-C. the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn3-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn1-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn3-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn4-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn3-N. and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn5-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn3-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn9-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn3-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn10-C;The N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn5-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn1-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn5-N, and the C intein comprises an amino acid sequence having at least about 85% sequence identity with Syn4-C; the N intein comprises an amino acid sequence having at least about 85% sequence identity with Syn5-N, and the C intein , comprising an amino acid sequence having at least about 85% sequence identity with Syn5-C; the N intein comprising an amino acid sequence having at least about 85% sequence identity with Syn5-N, and the C intein comprising an amino acid sequence having at least about 85% sequence identity with Syn9-C; or the N intein comprising an amino acid sequence having at least about 85% sequence identity with Syn5-N, and the C intein comprising an amino acid sequence having at least about 85% sequence identity with Syn10-C (see Tables 1A and 1B).

[0204] [Table 1A]

[0205] [Table 1B]

[0206] Further non-limiting examples of inteins suitable for use in embodiments of the present disclosure include any of the inteins provided herein, C inteins, N inteins, or functional variants, functional fragments, and various combinations thereof. Non-limiting examples of inteins include the dnaE intein, the Cfa-N (e.g., N intein) and Cfa-C (e.g., C intein) intein pair (described in Stevens et al., J Am Chem Soc. 2016 Feb. 24;138(7):2162-5, incorporated herein by reference), and synthetic inteins based on DnaE. Non-limiting examples of intein pairs that can be used in accordance with the present disclosure include the Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein, and Cne Prp8 intein (e.g., as described in U.S. Patent No. 8,394,604, incorporated herein by reference). Exemplary nucleotide and amino acid sequences of inteins are set forth in SEQ ID NOs: 370-377 in the Sequence Listing. Inteins suitable for use in embodiments of the present disclosure and methods for their use are described in U.S. Patent Nos. 10,526,401 and 11,142,550, International Patent Application Publication No. WO2013 / 045632, and U.S. Patent Application Publication No. US2020 / 0055900, and further described in Stevens, et al. "Design of a Split Intein with Exceptional Protein Splicing Activity," Journal of the American Chemical Society, 138:2162-2165 (2016), the entire disclosures of which are incorporated herein by reference in their entireties for all purposes.

[0207] In some embodiments, the split intein is selected from Gp41.1, IMPDH.1, NrdJ.1, and Gp41.8 (Carvajal-Vallejos, Patricia et al. "Unprecedented rates and efficiencies revealed for new natural split inteins from metagenomic sources." J. Biol. Chem., vol. 287,34 (2012)). Further non-limiting examples of amino acid and nucleotide sequences of N and C inteins suitable for use as intein pairs include those having at least 85% sequence identity to the amino acid or nucleotide sequences listed in Tables 2A-2C below, or functional fragments thereof.

[0208] [Table 2A-1] [Table 2A-2] [Table 2A-3]

[0209] [Table 2B-1] [Table 2B-2]

[0210] [Table 2C-1] [Table 2C-2]

[0211] Stargardt disease Stargardt disease (also known as Stargardt macular dystrophy, juvenile macular degeneration, or fundus flava) is a genetic disorder of the retina (i.e., the light-sensing tissue at the back of the eye). Stargardt disease is one of several genetic disorders that cause macular degeneration. The disease typically causes vision loss during childhood or adolescence, but in some cases, vision loss may not be noticed until adulthood. The disease rarely progresses to complete blindness. Generally, vision loss progresses slowly over time, and as progressive damage (degeneration) of the macula occurs, vision deteriorates to 20 / 200 or worse. In one example, Stargardt disease treated by the methods described herein includes juvenile Stargardt disease. In one example, Stargardt disease treated by the methods described herein includes late-onset Stargardt disease. In another example, Stargardt disease treated by the methods described herein includes Stargardt dominant macular dystrophy. In another example, Stargardt's disease treated by the methods described herein includes dominant Stargardt-like macular dystrophy.

[0212] The progression of Stargardt disease symptoms can vary from patient to patient. Patients with early onset of the disease generally tend to experience more rapid vision loss. Vision loss progresses slowly at first, then rapidly worsens until it plateaus. Most patients with Stargardt disease eventually have vision of 20 / 200 or worse. Patients with Stargardt disease may also begin to lose some peripheral (side) vision as they age.

[0213] In some embodiments, the pathogenic SNP is associated with Stargardt disease, optionally wherein the pathogenic SNP is located in the ABCA4 gene, and optionally wherein the pathogenic mutation comprises A1038V, L541P, G1961E, or a combination thereof. In some embodiments, the pathogenic SNP is associated with pseudoxanthoma elasticum, optionally wherein the pathogenic SNP is located in the ABCC6 gene, and optionally wherein the pathogenic mutation comprises R1141* (a nonsense mutation). In some embodiments, the pathogenic SNP is associated with medium-chain acyl-CoA dehydrogenase deficiency, optionally wherein the pathogenic SNP is located in the ACADM gene, and optionally wherein the pathogenic mutation comprises K329E. In some embodiments, the pathogenic SNP is associated with severe combined immunodeficiency, optionally wherein the pathogenic SNP is located in the ADA gene, and optionally wherein the pathogenic mutation comprises G216R, Q3*, or a combination thereof.

[0214] One or more symptoms of Stargardt disease include, but are not limited to, fluctuating and gradual loss of central vision in the form of gray, black, or blurred spots in the center of the visual field of both eyes; the eyes taking longer than normal to adapt when moving from a bright to a dark environment; the eyes becoming more sensitive to bright light; color vision defects occurring in the later stages of the disease; accumulation of toxic lipofuscin pigments such as A2E in cells of the retinal pigment epithelium (RPE); photoreceptor death; increased synthesis of 11-cis-retinaldehyde (11cRAL or retinal); increased rhodopsin regeneration; accumulation of lipofuscin; formation of lipofuscin pigments; retinal degeneration; waste product production; formation of A2E (and A2E-related molecules); accumulation of A2E (and A2E-related molecules); choroidal neovascularization; chorioretinal atrophy; or a combination thereof. The subject may show improvement in one or more of the symptoms of Stargardt disease. In one embodiment, the improvement in one or more of the symptoms is at least 5%. In another embodiment, one or more of the symptoms are improved by at least 10%. In another embodiment, one or more of the symptoms are improved by at least 15%. In another embodiment, one or more of the symptoms are improved by at least 20%. In another embodiment, one or more of the symptoms are improved by at least 25%. In another embodiment, one or more of the symptoms are improved by at least 30%. In another embodiment, one or more of the symptoms are improved by at least 35%. In another embodiment, one or more of the symptoms are improved by at least 40%. In another embodiment, one or more of the symptoms are improved by at least 50%. In another embodiment, one or more of the symptoms are improved by at least 60%. In another embodiment, one or more of the symptoms are improved by at least 70%. In another embodiment, one or more of the symptoms are improved by at least 75%. In another embodiment, one or more of the symptoms are improved by at least 80%. In another embodiment, one or more of the symptoms are improved by at least 85%. In another embodiment, one or more of the symptoms are improved by at least 90%. In another embodiment, one or more of the symptoms are improved by at least 95%.

[0215] Targeted gene editing To modify nucleic acid bases in a polynucleotide sequence, a subject is administered one or more guide polynucleotides and a base editor polypeptide or a polynucleotide encoding the same, or a cell is contacted with one or more guide polynucleotides and a base editor polypeptide or a polynucleotide encoding the same. The base editor polypeptide comprises a nucleic acid-programmable DNA-binding protein (napDNAbp) and a cytidine deaminase or an adenosine deaminase, or comprises one or more deaminases with cytidine deaminase activity and / or adenosine deaminase activity (e.g., a "dual deaminase" having cytidine deaminase activity and adenosine deaminase activity). In some embodiments, the base editor and / or endonuclease is introduced into a cell or administered to a subject using a polynucleotide sequence (e.g., mRNA) encoding the base editor and / or endonuclease. In some embodiments, the base editor and / or guide RNA is administered to a subject or contacted with a cell using an appropriate vector (e.g., an AAV vector or lipid nanoparticle). In some cases, the vector targets rod cells and / or cone cells. Non-limiting examples of vectors suitable for targeting rod cells and / or cone cells include the AAV5 vector or the PHB.EB AAV vector. In some embodiments, at least one nucleic acid is administered to a subject and / or a cell is contacted with at least one nucleic acid, wherein the at least one nucleic acid encodes one or more guide RNAs and encodes a nucleic acid-programmable DNA-binding protein (napDNAbp) and a nucleobase editor polypeptide comprising a cytidine deaminase. In some embodiments, the gRNA comprises nucleotide analogs. In some examples, the gRNA is added directly to the cell. These nucleotide analogs can inhibit degradation of the gRNA by cellular processes.Guide RNAs containing spacer sequences can be used to target base editors (e.g., adenosine base editors (ABEs), cytidine base editors (CBEs), and / or cytidine adenosine base editors (CABEs)) to edit genes.

[0216] In various instances, it is advantageous for a spacer sequence to include a 5' "G" nucleotide and / or a 3' "G" nucleotide. In some cases, for example, any spacer sequence or guide polynucleotide provided herein includes or further includes a 5' "G", where in some embodiments, the 5' "G" is complementary or not complementary to the target sequence. In some embodiments, a 5' "G" is added to a spacer sequence that does not already include a 5' "G". For example, when a guide RNA is expressed under the control of a U6 promoter or the like, it may be advantageous to include a 5'-terminal "G" in the guide RNA because the U6 promoter prefers a "G" at the transcription start site (see Cong, L. et al. "Multiplex genome engineering using CRISPR / Cas systems." Science 339:819-823 (2013) doi: 10.1126 / science.1231143). In some cases, a 5'-terminal "G" is added to a guide polynucleotide that is expressed under the control of a promoter. However, when a guide polynucleotide is not expressed under the control of a promoter, a 5'-terminal "G" is not added to the guide polynucleotide as needed.

[0217] In an embodiment, a guide suitable for use in targeting a base editor to effect a modification to the nucleobase at codon 1961 (e.g., E1961) of an ABCA4 polynucleotide, thereby altering a pathogenic G1961E amino acid mutation in the ABCA4 polynucleotide, comprises the following spacer sequence: GUGUCG AAGUUCGCCCUGGAG (SEQ ID NO:444), or a 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, or 20 nt fragment and / or extension thereof, wherein the nucleobase of the spacer corresponding to the target nucleobase is underlined in the sequence. In one embodiment, a guide suitable for use in targeting a base editor to effect a modification to a nucleobase of an ABCA4 polynucleotide comprises the following spacer sequence: GUGUCG G AGUUCGCCCUGGAG (SEQ ID NO:445), or a 1-nt, 2-nt, 3-nt, 4-nt, 5-nt, 6-nt, 7-nt, 8-nt, 9-nt, 10-nt, 11-nt, 12-nt, 13-nt, 14-nt, 15-nt, 16-nt, 17-nt, 18-nt, 19-nt, or 20-nt fragment and / or extension thereof, wherein the nucleobase of the spacer corresponding to the targeted nucleobase is underlined in the sequence. In certain embodiments, a guide suitable for use in targeting a base editor to effect a modification to a nucleobase of an ABCA4 polynucleotide comprises one of the following spacer sequences: CACCUCUCCAGGGCGAACUUCGACACAGC (SEQ ID NO:446), CACCUCUCCAGGGCGAACUCCGACACAGC (SEQ ID NO:447), or CUCCAGGGCGAACUUCGACACACAGC (SEQ ID NO:448).

[0218] In various embodiments, a guide suitable for use in targeting a base editor to effect a modification to a nucleobase of an ABCA4 polynucleotide as part of a treatment for Stargardt disease is GCTGTGTGCGAAGTTCGCCCTGGAG AGG TG (SEQ ID NO: 449) or GCTGTGTGCGGAGTTCGCCCTGGAG AGG TG (SEQ ID NO: 450), where representative PAM sequences in each target sequence are underlined.

[0219] Nucleic acid base editor Nucleobase editors that edit, modify, or alter a target nucleotide sequence of a polynucleotide are useful in the methods and compositions described herein. The nucleobase editors described herein typically comprise a polynucleotide-programmable nucleotide-binding domain and a nucleobase-editing domain (e.g., an adenosine deaminase, a cytidine deaminase, or a dual deaminase). The polynucleotide-programmable nucleotide-binding domain, when coupled with a bound guide polynucleotide (e.g., a gRNA), can specifically bind to a target polynucleotide sequence, thereby localizing the base editor to the target nucleic acid sequence desired to be edited.

[0220] Polynucleotide-programmable nucleotide-binding domains The polynucleotide-programmable nucleotide binding domain binds to a polynucleotide (e.g., RNA, DNA). The polynucleotide-programmable nucleotide binding domain of a base editor may itself comprise one or more domains (e.g., one or more nuclease domains). In some embodiments, the nuclease domain of the polynucleotide-programmable nucleotide binding domain comprises an endonuclease or an exonuclease.

[0221] Disclosed herein are base editors that comprise a nucleotide-binding domain that can be programmed by a polynucleotide comprising all or a portion (e.g., a functional portion) of a CRISPR protein (i.e., the base editor comprises all or a portion (e.g., a functional portion) of a CRISPR protein (e.g., a Cas protein) as a domain, also referred to as the "CRISPR protein-derived domain" of the base editor). The CRISPR protein-derived domain incorporated into the base editor may be modified relative to a wild-type or naturally occurring CRISPR protein. The CRISPR protein-derived domain may comprise one or more mutations, insertions, deletions, rearrangements, and / or recombinations relative to a wild-type or naturally occurring CRISPR protein.

[0222] Cas proteins that can be used in the present disclosure include class 1 and class 2. Non-limiting examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 or Csx12), Cas10, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csx1, Csx2, Csx3, Csx4, Csx5, Csx6, Csx7, Csx8, Csx9, Csx10, Csx11, Csx12, Csx13, Csx14, Csx15, Csx16, Csx17, Csx18, Csx19, Csx20, Csx21, Csx22, Csx23, Csx24, Csx25, Csx26, Csx27, Csx28, Csx29, Csx30, Csx31, Csx32, Csx33, Csx34, Csx35, Csx36, Csx37, Csx38, Csx39, Csx40, Csx41, Csx42, Csx43, Csx44, Csx45, Csx46, Csx47, Csx48, Csx49, Csx51, Csx52, Csx53, Examples of such proteins include sb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas12a / Cpf1, Cas12b / C2c1 (e.g., SEQ ID NO: 232), Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, and Cas12j / CasΦ, CARF, DinG, homologs thereof, or modifications thereof. CRISPR enzymes can induce cleavage of one or both strands in a target sequence, for example, within the target sequence and / or within a sequence complementary to the target sequence. For example, CRISPR enzymes can induce cleavage of one or both strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence.

[0223] A vector encoding a CRISPR enzyme mutated relative to the corresponding wild-type enzyme can be used. Such mutant CRISPR enzymes lack the ability to cleave one or both strands of a target polynucleotide containing a target sequence. A Cas protein (e.g., Cas9, Cas12) or Cas domain (e.g., Cas9, Cas12) can refer to a polypeptide or domain that has about 50% or more, about 60% or more, about 70% or more, about 80% or more, about 90% or more, about 91% or more, about 92% or more, about 93% or more, about 94% or more, about 95% or more, about 96% or more, about 97% or more, about 98% or more, about 99% or more, or 100% sequence identity and / or sequence homology to a wild-type exemplary Cas polypeptide or Cas domain. Cas (e.g., Cas9, Cas12) can refer to wild-type or modified forms of Cas proteins that can include amino acid changes such as deletions, insertions, substitutions, mutants, mutations, fusions, chimeras, or any combination thereof.

[0224] In some embodiments, the base editor CRISPR protein-derived domains are selected from the group consisting of Corynebacterium ulcerans (NCBI Reference Nos. NC_015683.1, NC_017317.1), Corynebacterium diphtheria (NCBI Reference Nos. NC_016782.1, NC_016786.1), Spiroplasma syrphidicola (NCBI Reference No. NC_021284.1), Prevotella intermedia (NCBI Reference No. NC_017861.1), Spiroplasma taiwanense (NCBI Reference No. NC_021846.1), Streptococcus iniae (NCBI Reference No. NC_021314.1), Belliella baltica (NCBI Reference No. NC_018010.1), Psychroflexus torquis (NCBI Reference Number: NC_018721.1), Streptococcus thermophilus (NCBI Reference Number: YP_820832.1), Listeria innocua (NCBI Reference Number: NP_472073.1), Campylobacter jejuni (NCBI Reference Number: YP_002344900.1), Neisseria meningitidis (NCBI Reference Number: YP_002342100.1), Streptococcus pyogenes, or Staphylococcus aureus.

[0225] Some embodiments of the present disclosure provide a high-fidelity Cas9 domain. High-fidelity Cas9 domains are known in the art and are described, for example, in Kleinstiver, BP, et al. "High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects." Nature 529, 490-495 (2016) and Slaymaker, IM, et al. "Rationally engineered Cas9 nucleases with improved specificity." Science 351, 84-88 (2015), the entire contents of each of which are incorporated herein by reference. An exemplary high-fidelity Cas9 domain is set forth in the Sequence Listing as SEQ ID NO: 233.

[0226] In some embodiments, any of the Cas9 fusion proteins or complexes provided herein comprises one or more of the following mutations: D10A, N497X, R661X, Q695X, and / or Q926X, or the corresponding mutations in any of the amino acid sequences provided herein, where X is any amino acid.

[0227] Typically, Cas9 proteins, such as Cas9 from S. pyogenes (spCas9), require a "protospacer adjacent motif (PAM)" or PAM-like motif, which is a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by the Cas9 nuclease in the CRISPR bacterial adaptive immune system. The presence of the NGG PAM sequence is required to bind to a specific nucleic acid region. Here, the "N" in "NGG" is adenosine (A), thymidine (T), or cytosine (C), and G is guanosine. In some embodiments, any of the fusion proteins or complexes provided herein may contain a Cas9 domain capable of binding to a nucleotide sequence that does not contain a standard (e.g., NGG) PAM sequence. Cas9 domains that bind to non-standard PAM sequences have been described in the art and would be apparent to one of skill in the art. For example, Cas9 domains that bind to non-canonical PAM sequences are described in Kleinstiver, BP, et al., "Engineered CRISPR-Cas9 nucleases with altered PAM specificities" Nature 523, 481-485 (2015), and Kleinstiver, BP, et al., "Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition" Nature Biotechnology 33, 1293-1298 (2015), the entire contents of each of which are incorporated herein by reference.

[0228] In some embodiments, the napDNAbp is a circular permutant (eg, SEQ ID NO: 238).

[0229] In some embodiments, the polynucleotide-programmable nucleotide-binding domain comprises a nickase domain. As used herein, the term "nickase" refers to a polynucleotide-programmable nucleotide-binding domain that includes a nuclease domain that can cleave only one of the two strands of a double-stranded nucleic acid molecule (e.g., DNA). For example, when the polynucleotide-programmable nucleotide-binding domain comprises a nickase domain derived from Cas9, the Cas9-derived nickase domain can comprise a D10A mutation and a histidine at position 840. In another example, the Cas9-derived nickase domain comprises a H840A mutation, but the amino acid residue at position 10 remains D.

[0230] In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA-cleaving domain, i.e., (in the case of "nickase" Cas9, SEQ ID NO: 201), and the Cas9 is a nickase and is referred to as an "nCas9" protein. A Cas9 nickase can be a Cas9 protein that can cleave only one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). In some embodiments, the Cas9 nickase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the Cas9 nickases provided herein. Additional suitable Cas9 nickases will be apparent to those of skill in the art based on this disclosure and knowledge in the art, and are within the scope of this disclosure.

[0231] Also provided herein are base editors that are catalytically inactive (i.e., unable to cleave a target polynucleotide sequence) and comprise a polynucleotide-programmable nucleotide-binding domain. For example, in the case of a base editor comprising a Cas9 domain, the Cas9 may comprise both a D10A mutation and an H840A mutation. In further embodiments, the catalytically inactive polynucleotide-programmable nucleotide-binding domain comprises a point mutation (e.g., D10A or H840A) and a deletion of all or a portion (e.g., a functional portion) of the nuclease domain. dCas9 domains are known in the art and are described, for example, in Qi et al., "Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression." Cell. 2013;152(5):1173-83, the entire contents of which are incorporated herein by reference.

[0232] The term "protospacer adjacent motif (PAM)" or PAM-like motif refers to a 2-6 base pair DNA sequence immediately following the DNA sequence targeted by a nucleic acid-programmable DNA-binding protein. In some embodiments, the PAM can be a 5' PAM (i.e., a PAM located upstream of the 5' end of the protospacer). In other embodiments, the PAM can be a 3' PAM (i.e., a PAM located downstream of the 5' end of the protospacer). The PAM sequence can be any PAM sequence known in the art. Suitable PAM sequences include, but are not limited to, NGG, NGA, NGC, NGN, NGT, NGTT, NGCG, NGAG, NGAN, NGNG, NGCN, NGCG, NGTN, NNGRRT, NNNRRT, NNGRR(N), TTTV, TYCV, TYCV, TYCV, TATV, NNNNGATT, NNAGAAW, or NAAAAC. Y is a pyrimidine, N is any nucleotide base, and W is A or T.

[0233] The base editors provided herein can comprise a domain from a CRISPR protein, which can bind to a nucleotide sequence containing a canonical or non-canonical protospacer adjacent motif (PAM) sequence.

[0234] In some embodiments, the PAM is an "NRN" PAM (wherein the "N" in "NRN" is adenine (A), thymine (T), guanine (G), or cytosine (C), and R is adenine (A) or guanine (G)), or the PAM is an "NYN" PAM (wherein the "N" in "NYN" is adenine (A), thymine (T), guanine (G), or cytosine (C), and Y is cytidine (C) or thymine (T)), e.g., as described in R.T. Walton et al., 2020, Science, 10.1126 / science.aba8853(2020), the entire contents of which are incorporated herein by reference.

[0235] Table 3 below lists some PAM variants. [Table 3]

[0236] In some embodiments, the PAM is NGC. In some embodiments, the NGC PAM is recognized by a Cas9 mutant. In some embodiments, the NGC PAM Cas9 mutant comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219F, A1322R, D1332A, R1335E, and T1337R (collectively "MQKFRAER") of spCas9 (SEQ ID NO: 197), or the corresponding mutations of another Cas9. In some embodiments, the Cas9 mutant comprises one or more amino acid substitutions selected from D1135V, G1218R, R1335Q, and T1337R (collectively "VRQR") of spCas9 (SEQ ID NO: 197), or the corresponding mutations of another Cas9. In some embodiments, the Cas9 mutant comprises one or more amino acid substitutions selected from D1135V, G1218R, R1335E, and T1337R (collectively VRER) of spCas9 (SEQ ID NO: 197), or the corresponding mutations of another Cas9. In some embodiments, the Cas9 mutant comprises one or more amino acid substitutions selected from E782K, N968K, and R1015H (collectively KHH) of saCas9 (SEQ ID NO: 218). In some embodiments, the Cas9 mutant comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219S, R1335E, and T1337R (collectively "MQKSER") of spCas9 (SEQ ID NO: 197), or the corresponding mutations of another Cas9. In some embodiments, the Cas9 mutant comprises one or more amino acid substitutions selected from D1135M, S1136Q, G1218K, E1219S, R1335E, and T1337R (collectively "MQKSER") of spCas9 (SEQ ID NO: 197), or the corresponding mutations of another Cas9.

[0237] In some embodiments, the base editor CRISPR protein-derived domain comprises all or a portion (e.g., a functional portion) of a Cas9 protein that uses the standard PAM sequence (NGG). In other embodiments, the base editor Cas9-derived domain can use a non-standard PAM sequence. Such sequences are described in the art and will be apparent to one of skill in the art. For example, Cas9 domains that bind to non-canonical PAM sequences have been reported in Kleinstiver, B.P., et al., “Engineered CRISPR-Cas9 nucleases with altered PAM specificities” Nature 523, 481-485 (2015), and Kleinstiver, B.P., et al., “Broadening the targeting range of Staphylococcus aureus CRISPR-Cas9 by modifying PAM recognition” Nature Biotechnology 33, 1293-1298 (2015), RT Walton et al. “Unconstrained genome targeting with near-PAMless engineered CRISPR-Cas9 variants” Science 10.1126 / science.aba8853 (2020), Hu et al. “Evolved Cas9 variants with broad PAM compatibility and high DNA specificity,” Nature, 2018 Apr. 5, 556(7699), 57-63, and Miller et al., “Continuous evolution of SpCas9 variants compatible with non-G PAMs,” Nat. Biotechnol., 2020 Apr;38(4):471-481, the entire contents of each of which are incorporated herein by reference.

[0238] Fusion protein or complex containing NapDNAbp and cytidine deaminase and / or adenosine deaminase Some aspects of the present disclosure provide fusion proteins or complexes comprising a Cas9 domain or other nucleic acid-programmable DNA-binding protein (e.g., Cas12) and one or more cytidine deaminase domains, adenosine deaminase domains, or cytidine adenosine deaminase domains. As is apparent, the Cas9 domain can be any of the Cas9 domains or Cas9 proteins provided herein (e.g., dCas9 or nCas9). In some embodiments, any of the Cas9 domains or Cas9 proteins provided herein (e.g., dCas9 or nCas9) can be fused to any of the cytidine deaminases and / or adenosine deaminases provided herein. The domains of the base editors disclosed herein can be arranged in any order.

[0239] In some embodiments, a fusion protein or complex comprising a cytidine deaminase or adenosine deaminase and a napDNAbp (e.g., a Cas9 or Cas12 domain) does not include a linker sequence. In some embodiments, a linker is present between the cytidine deaminase or adenosine deaminase and the napDNAbp. In some embodiments, the cytidine deaminase or adenosine deaminase and the napDNAbp are fused via any of the linkers provided herein. For example, in some embodiments, the cytidine deaminase or adenosine deaminase and the napDNAbp are fused via any of the linkers provided herein.

[0240] It should be understood that the fusion proteins or complexes of the present disclosure may have one or more additional features. For example, in some embodiments, the fusion proteins or complexes may include an inhibitor, a cytoplasmic localization sequence, a transport sequence (e.g., a nuclear export sequence), or other localization sequence, as well as a sequence tag useful for solubilizing, purifying, or detecting the fusion protein or complex. Suitable protein tags provided herein include, but are not limited to, biotin carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, polyhistidine tags (also known as histidine tags or His tags), maltose-binding protein (MBP) tags, nus tags, glutathione-S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, soft tags (e.g., soft tag 1, soft tag 3), streptags, biotin ligase tags, FIAsH tags, V5 tags, and SBP tags. Additional suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein or complex comprises one or more His tags. Exemplary, but non-limiting, fusion proteins are described in International PCT Applications PCT / 2017 / 045381, PCT / US2019 / 044935, and PCT / US2020 / 016288 (each of which is incorporated by reference in its entirety).

[0241] Fusion proteins or complexes with internal insertions Provided herein are nucleic acid-programmable nucleic acid-binding proteins, such as fusion proteins or complexes, comprising a heterologous polypeptide fused to a napDNAbp. The heterologous polypeptide can be fused to the napDNAbp at the C-terminus, the N-terminus, or inserted at an internal position of the napDNAbp. In some embodiments, the heterologous polypeptide is a deaminase (e.g., cytidine deaminase or adenosine deaminase) or a functional fragment thereof. For example, the fusion protein can comprise a deaminase adjacent to the N-terminal and C-terminal fragments of a Cas9 or Cas12 (e.g., Cas12b / C2c1) polypeptide.

[0242] The deaminase can be a circularly permuted deaminase. In some embodiments, the deaminase is a circularly permuted TadA that is circularly permuted at amino acid residues 116, 136, or 65, as numbered in the TadA reference sequence.

[0243] The fusion protein or complex may comprise multiple deaminases, for example, 1, 2, 3, 4, 5, or more deaminases. The deaminases of the fusion protein or complex may be adenosine deaminase, cytidine deaminase, or a combination thereof.

[0244] In some embodiments, the napDNAbp in the fusion protein or complex comprises a Cas9 polypeptide or a fragment thereof. The Cas9 polypeptide can be a mutant Cas9 polypeptide. The Cas9 polypeptide can be a circularly permuted Cas9 protein.

[0245] A heterologous polypeptide (e.g., a deaminase) can be inserted into the napDNAbp (e.g., Cas9 or Cas12 (e.g., Cas12b / C2c1)) at a suitable position, for example, so that the napDNAbp retains the ability to bind to a target polynucleotide and a guide nucleic acid. A deaminase (e.g., an adenosine deaminase, a cytidine deaminase, or an adenosine deaminase and a cytidine deaminase (dual deaminase)) can be inserted into the napDNAbp without impairing the function of the deaminase (e.g., base editing activity) or the function of the napDNAbp (e.g., the ability to bind to a target nucleic acid and a guide nucleic acid).

[0246] In some embodiments, the deaminase (e.g., adenosine deaminase, cytidine deaminase, or adenosine deaminase and cytidine deaminase) is inserted into a region of a Cas9 polypeptide having a higher than average B factor (e.g., a higher B factor compared to the entire protein or a protein domain containing a disordered region). Positions in the Cas9 polypeptide having a higher than average B factor can include, for example, residues 768, 792, 1052, 1015, 1022, 1026, 1029, 1067, 1040, 1054, 1068, 1246, 1247, and 1248, as numbered in the above Cas9 reference sequence. Regions of the Cas9 polypeptide having a higher than average B factor can include, for example, residues 792-872, 792-906, and 2-791, as numbered in the above Cas9 reference sequence.

[0247] In some embodiments, the heterologous polypeptide (e.g., a deaminase) is inserted into a flexible loop of a Cas9 polypeptide. The flexible loop portion can be selected from the group consisting of amino acid residues 530-537, 569-570, 686-691, 943-947, 1002-1025, 1052-1077, 1232-1247, or 1298-1300, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide. The flexible loop portion can be selected from the group consisting of amino acid residues 1-529, 538-568, 580-685, 692-942, 948-1001, 1026-1051, 1078-1231, or 1248-1297, as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0248] A heterologous polypeptide (e.g., adenine deaminase) can be inserted into a region of a Cas9 polypeptide corresponding to amino acid residues 1017-1069, 1242-1247, 1052-1056, 1060-1077, 1002-1003, 943-947, 530-537, 568-579, 686-691, 1242-1247, 1298-1300, 1066-1077, 1052-1056, or 1060-1077 as numbered in the above Cas9 reference sequence, or the corresponding amino acid residues in another Cas9 polypeptide.

[0249] A heterologous polypeptide (e.g., adenine deaminase) can be inserted in place of the deleted region of the Cas9 polypeptide. The deleted region can correspond to the N-terminal or C-terminal portion of the Cas9 polypeptide. Exemplary internal fusion base editors are shown in Table 4A below. [Table 4A]

[0250] A heterologous polypeptide (e.g., a deaminase) can be inserted within a structural or functional domain of a Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) can be inserted between two structural or functional domains of a Cas9 polypeptide. A heterologous polypeptide (e.g., a deaminase) can be inserted in place of a structural or functional domain, for example, after removing the structural or functional domain from the Cas9 polypeptide. A structural or functional domain of a Cas9 polypeptide can include, for example, RuvCI, RuvC II, RuvC III, Rec1, Rec2, PI, or HNH.

[0251] The fusion protein can include a linker between the deaminase and the napDNAbp polypeptide. The linker can be a peptide linker or a non-peptide linker. For example, the linker can be XTEN, (GGGS) n (SEQ ID NO: 246), SGGSSGGS (SEQ ID NO: 330), (GGGGS) n (SEQ ID NO: 247), (G) n , (EAAAK) n (SEQ ID NO: 248), (GGS) n , SGSETPGTSESATPES (SEQ ID NO: 249). In some embodiments, the fusion protein comprises a linker between the N-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein comprises a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the N-terminal and C-terminal fragments of napDNAbp are linked to the deaminase with a linker. In some embodiments, the N-terminal and C-terminal fragments are linked to the deaminase domain without a linker. In some embodiments, the fusion protein comprises a linker between the N-terminal Cas9 fragment and the deaminase, but does not comprise a linker between the C-terminal Cas9 fragment and the deaminase. In some embodiments, the fusion protein comprises a linker between the C-terminal Cas9 fragment and the deaminase, but does not comprise a linker between the N-terminal Cas9 fragment and the deaminase.

[0252] In some embodiments, the napDNAbp in the fusion protein or complex is a Cas12 polypeptide (e.g., Cas12b / C2c1) or a functional fragment thereof capable of binding to a nucleic acid (e.g., gRNA) that directs Cas12 to a specific nucleic acid sequence. The Cas12 polypeptide can be a mutant Cas12 polypeptide. In other embodiments, an N-terminal or C-terminal fragment of a Cas12 polypeptide comprises a nucleic acid-programmable DNA-binding domain or a RuvC domain. In other embodiments, the fusion protein comprises a linker between the Cas12 polypeptide and the catalytic domain. In other embodiments, the amino acid sequence of the linker is GGSGGS (SEQ ID NO:250) or GSSGSETPGTSESATPESSG (SEQ ID NO:251). In other embodiments, the linker is a rigid linker. In other embodiments of the above aspects, the linker is encoded by GGAGGCTCTGGAGGAAGC (SEQ ID NO:252) or GGCTCTTCTGGATCTGAAACACCTGGCACAAGCGAGAGCGCCACCCCTGAGAGCTCTGGC (SEQ ID NO:253).

[0253] In other embodiments, the fusion protein or complex comprises a nuclear localization signal (e.g., a bipartite nuclear localization signal). In other embodiments, the amino acid sequence of the nuclear localization signal is MAPKKKRKVGIHGVPAA (SEQ ID NO: 261). In other embodiments of the above aspects, the nuclear localization signal is encoded by the following sequence: ATGGCCCCAAAGAAGAAGCGGAAGGTCGGTATCCACGGAGTCCCAGCAGCC (SEQ ID NO: 262). In other embodiments, the Cas12b polypeptide comprises a mutation that abolishes the catalytic activity of the RuvC domain. In other embodiments, the Cas12b polypeptide comprises a D574A, D829A, and / or D952A mutation.

[0254] In some embodiments, the fusion protein or complex comprises a napDNAbp domain (e.g., a Cas12-derived domain) with an internally fused nucleobase editing domain (e.g., all or a portion (e.g., a functional portion) of a deaminase domain (e.g., an adenosine deaminase domain)). In some embodiments, the napDNAbp is Cas12b. In some embodiments, the base editor comprises a BhCas12b domain with an internally fused TadA*8 domain inserted at the position shown in Table 4B below.

[0255] [Table 4B]

[0256] In some embodiments, the base editor system described herein is an ABE in which TadA has been inserted into Cas9. The polypeptide sequences of suitable ABEs with TadA inserted into Cas9 are set forth in the accompanying sequence listing as SEQ ID NOs: 263-308.

[0257] Exemplary, non-limiting fusion proteins are described in International PCT Application No. PCT / US2020 / 016285 and U.S. Provisional Application Nos. 62 / 852,228 and 62 / 852,224, the contents of which are incorporated by reference in their entireties.

[0258] Editing A to G In some embodiments, a base editor described herein comprises an adenosine deaminase domain. Such an adenosine deaminase domain of a base editor can deaminate A to form inosine (I), which exhibits the base pairing properties of G, thereby facilitating the editing of an adenine (A) nucleobase to a guanine (G) nucleobase. In some embodiments, an A to G base editor further comprises an inhibitor of inosine base excision repair, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease. Without wishing to be bound by any particular theory, the UGI domain or catalytically inactive inosine-specific nuclease can inhibit or prevent base excision repair of deaminated adenosine residues (e.g., inosine), which can improve the activity or efficiency of the base editor.

[0259] Base editors comprising adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In one embodiment, the adenosine deaminase domain of a base editor comprises all or a portion (e.g., a functional portion) of ADAT containing one or more mutations that enable ADAT to deaminate a target A in DNA. For example, a base editor can comprise all or a portion (e.g., a functional portion) of ADAT from Escherichia coli (EcTadA) containing one or more of D108N, A106V, D147Y, E155V, L84F, H123Y, I156F, or corresponding mutations in another adenosine deaminase. Exemplary ADAT homolog polypeptide sequences are set forth in the Sequence Listing as SEQ ID NOS: 1 and 309-315.

[0260] The adenosine deaminase can be derived from any suitable organism (e.g., E. coli). In some embodiments, the adenosine deaminase is derived from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenine deaminase is a naturally occurring adenosine deaminase that includes one or more mutations corresponding to any of the mutations set forth herein (e.g., mutations in ecTadA). Corresponding residues in any homologous protein can be identified, for example, by sequence alignment and determination of homologous residues. Mutations in any naturally occurring adenosine deaminase (e.g., having homology to ecTadA) that correspond to any of the mutations described herein (e.g., any of the mutations identified in ecTadA) can be generated accordingly.

[0261] In some embodiments, the adenosine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the amino acid sequences set forth in any of the adenosine deaminases provided herein. As will be apparent, the adenosine deaminases provided herein can comprise one or more mutations (e.g., any of the mutations set forth herein). The present disclosure provides any deaminase domain having a particular percent identity and, in addition, any of the mutations or combinations of mutations described herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to a reference sequence or any of the adenosine deaminases provided herein.

[0262] As will be apparent, any of the mutations provided herein (e.g., based on a TadA reference sequence such as TadA*7.10 (SEQ ID NO: 1)) can be introduced into other adenosine deaminases, such as E. coli TadA (ecTadA), S. aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). In some embodiments, the TadA reference sequence is TadA*7.10 (SEQ ID NO: 1). As will be apparent to one of skill in the art, additional deaminases can be similarly sequenced to identify homologous amino acid residues, which can then be mutated as provided herein. Thus, any of the mutations identified in the TadA reference sequence can be made in other adenosine deaminases (e.g., ecTadA) that have homologous amino acid residues. As will be apparent, any of the mutations provided herein can be made individually or in any combination in the TadA reference sequence or another adenosine deaminase.

[0263] In some embodiments, the adenosine deaminase comprises a mutation or set of mutations selected from those listed in Tables 5A-5E below.

[0264] [Table 5A-1] [Table 5A-2]

[0265] [Table 5B]

[0266] [Table 5C-1] [Table 5C-2]

[0267] In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant further comprising an F149Y amino acid mutation. In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant (TadA*8.10+) further comprising amino acid mutations R147D, F149Y, T166I, and D167N. In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant (TadA*9v1) further comprising amino acid mutations S82T and F149Y. In some embodiments, the adenosine deaminase comprises a TadA*8.20 adenosine deaminase variant (TadA*9v2) further comprising amino acid mutations Y147D, F149Y, T166I, D167N, and S82T.

[0268] In some embodiments, the adenosine deaminase is selected from the group consisting of M1I, S2A, S2E, V4D, V4E, V4M, F6S, H8E, H8Y, E9Y, M12S, R13H, R13I, R13Y, T17L, T17S, L18A, L18E, A19N, R21N, K20K, K20R, R21A, G22P, W23D, R23H, W23G, W23Q, W23L, W23R, D24E, D24G, E25F, E25M, E25D, E25A, E25G, E25R, E25V, E25S, E25Y, R26D, R26E, R26G, R26N, R26Q, R26C, R26L, R26K, R26W, E2 7V, E27D, P29V, V30G, L34S, L34V, L36H, H36L, H36N, N37N, N37T, N37S, N38G, N38R, W45A, W45L, W45N, N46N, R46W, R46F, R46Q, R46M, R47A, R47Q, R47F, R4 7K, R47P, R47W, R47M, P48T, P48L, P48A, P48I, P48S, I49G, I49H, I49V, I49F, I49H, G50L, R51H, R51L, R51N, L51W, R51Y, H52D, H52Y, D53P, P54C, P54T, A5 5H, T55A, A56E, A56S, E59A, E59G, E59I, E59Q, E59W, M61A, M61I, M61L, M61V, L63S, L63V, Q65V, G66C, G67D, G67L, G67V, L68Q, M70H, M70Q, L84F, M70V, M7 0L, E70A, M70V, Q71M, Q71N, Q71L, Q71R, N72A, N72K, N72S, N72D, N72Y, Y73G, Y73I, Y73K, Y73R, Y73S, R74A, R74Q, R74G, R74K, R74L, R74N, I76D, I76F, I7 6I, I76N, I76T, I76Y, D77G, A78I, T79M, L80M, L80Y, V82A, V82S, V82G, V82T, L84E, L84F, L84Y, E85K, E85G, E85P, E85S, S87C, S87L, S87V, V88A, V88M, C9 0S, A91A, A91G, A91S, A91V, A91T, G92T, A93I, M94A, M94V, M94L, M94I, M94H,I95S、I95G、I95L、I95H、I95V、H96A、H96L、H96R、H96S、S97C、S97G、S97I、S97M、S97R、S97S、R98K、R98I、R98N、R98Q、G100R、G100V、R101V、R101R、V102A、V102F、V102I、V102V、D103A、F104G、D104N、F104V、F104I、F104L、A106T、V106Q、V106F、V106W、V106M、A106A、A106Q、A106F、A106G、A106W、A106M、A106V、A106R、R107C、R107G、R107P、R107K、R107A、R107N、R107W、R107H、R107S、D108N、D108F、D108G、D108V、D108A、D108Y、D108H、D108I、D108K、D108L、D108M、D108Q、N108Q、N108F、N108W、N108M、N108K、D108K、D108F、D108M、D108Q、D108R、D108W、D108S、A109H、A109K、A109R、A109S、A109T、A109V、K110G、K110H、K110I、K110R、K110T、T111A、T111G、T111H、T111R、G112A、A114G、A114H、A114V、G115S、L117M、L117N、L117V、M118D、M118G、M118K、M118N、M118V、D119L、D119N、D119S、D119V、V120H、V120L、H122H、H122N、H122P、H122R、H122S、H122Y、H123C、H123G、H123P、H123V、H123Y、Y123H、P124G、P124I、P124L、P124W、G125H、G125I、G125A、G125M、G125K、M126D、M126H、M126K、M126I、M126N、M126O、M126S、M126Y、N127H、N127S、N127D、N127K、N127R、H128R、R129H、R129Q、R129V、R129I、R129E、R129V、I132I、I132F、T133V、T133E、T133G、T133K、E134A、E134E、E134G、E134I、G135G、G135V、I136G、I136L、I136T、L137A、L137D, L137E, L137M, L137S, A138D, A138E, A138G, S138A, A138N, A138S, A138T, A138V, A138Y, D139E, D139I, D139C, D139L, D13 9M, E140A, E140C, E140L, E140R, A142N, A142D, A142G, A142A, A142L, A142S, A142T, A142N, A142S, A142V, A143D, A143E, A143G, , A143D, A143G, A143E, A143L, A143W, A143M, A143S, A143Q, A143R, C146R, S146A, S146C, S146D, S146F, S146R, S146T, D 147D, D147L, D147F, D147G, D147Y, Y147T, Y147R, Y147D, D147R, F148L, F148F, F148R, F148Y, F149C, F149M, F149R, F149 Y, M151F, M151P, M151R, M151V, R152C, R152F, R152H, R152P, R152R, R153C, R153Q, R153R, R153V, Q154E, Q154H, Q154M, Q154R, Q154L, Q154S, Q154V, E155F, E155G, E155I, E155K, E155P, E155V, E155D, I156A, I156F, I156D, I156K, I156N, I15 6R, I156Y, E157A, E157F, E157I, E157P, E157T, E157V, N157K, K157N, K157R, A158Q, A158K, A158V, Q159F, Q159K, Q159L , Q159N, K160A, K160S, K160E, K160K, K160N, K161I, K161A, K161N, K161Q, K161S, K161T, A162D, A162Q, R162H, R162P, A1 and / or D167N mutations, and any alternative mutations at the corresponding positions, or R26, W23, E27, H36, R46, W56, W62S, Q163G, Q163H, Q163N, Q163R, S164I, S164R, S164Y, S165A, S165D, S165I, S165T, S165Y, T166D, T166K, T166I, T166N, T166P, T166R, D167S, and / or D167N mutations, with respect to the TadA reference sequence.including any substitution from R47, P48, R51, H52, R74, I76, V82, V88, M94, I95, H96, A106, D108, A109, K110, T111, A114, D119, H122, H123, M126, N127, A142, S146, D147, F149, R152, Q154, E155, I156, E157, K161, T166, and / or D167, or W23R, E27D, H36L, R47K, P48A, R51H, R51L, I The TadA sequence may comprise 2 to 50 amino acid substitutions in the TadA reference sequence, which may be selected from 76F, I76Y, V82S, A106V, D108G, A109S, K110R, T111H, A114V, D119N, H122R, H122N, H123Y, M126I, N127K, S146C, D147R, R152P, Q154R, E155V, 1156F, K157N, K161N, T166I, and D167N, or one or more corresponding mutations in another adenosine deaminase. Further mutations are described in U.S. Patent Application Publication No. 2022 / 0307003A1, International Patent Application Publication No. WO2023 / 288304A2, and WO2023 / 034959A2, the disclosures of which are incorporated herein by reference in their entireties for all purposes.

[0269] In certain embodiments, the mutant of TadA*7.10 comprises one or more mutations selected from any of the mutations provided herein.

[0270] In certain embodiments, the adenosine deaminase heterodimer comprises a TadA*8 domain and an adenosine deaminase domain selected from Staphylococcus aureus (S. aureus) TadA, Bacillus subtilis (B. subtilis) TadA, Salmonella typhimurium (S. typhimurium) TadA, Shewanella putrefaciens (S. putrefaciens) TadA, Haemophilus influenzae F3031 (H. influenzae) TadA, Caulobacter crescentus (C. crescentus) TadA, Geobacter sulfurreducens (G. sulfurreducens) TadA, or TadA*7.10.

[0271] In some embodiments, TadA*8 is a variant shown in Table 5D. Table 5D lists the numbers of specific amino acid positions in the TadA amino acid sequence and the amino acids present at those positions in TadA-7.10 adenosine deaminase. Table 5D also lists the amino acid mutations of TadA variants relative to TadA-7.10 after phage-assisted non-continuous evolution (PANCE) and phage-assisted continuous evolution (PACE) as described in M. Richter et al., 2020, Nature Biotechnology, doi.org / 10.1038 / s41587-020-0453-z, the entire contents of which are incorporated herein by reference. In some embodiments, TadA*8 is TadA*8a, TadA*8b, TadA*8c, TadA*8d, or TadA*8e. In some embodiments, TadA*8 is TadA*8e. In one embodiment, the adenosine deaminase is TadA*8 comprising SEQ ID NO: 316 or a fragment thereof having adenosine deaminase activity, or TadA*8 consisting essentially of SEQ ID NO: 316 or a fragment thereof having adenosine deaminase activity.

[0272] [Table 5D]

[0273] In some embodiments, the TadA mutant is a mutant shown in Table 5E. Table 5E shows the number of specific amino acid positions in the TadA amino acid sequence and the amino acid present at those positions in TadA*7.10 adenosine deaminase. In some embodiments, the TadA mutant is MSP605, MSP680, MSP823, MSP824, MSP825, MSP827, MSP828, or MSP829. In some embodiments, the TadA mutant is MSP828. In some embodiments, the TadA mutant is MSP829.

[0274] [Table 5E]

[0275] In certain embodiments, a fusion protein or complex comprises a single (e.g., provided as a monomer) TadA* (e.g., TadA*8 or TadA*9). Throughout this disclosure, adenosine deaminase base editors comprising a single TadA* domain are referred to using the term ABEm or ABE#m, where "#" is an identifier number (e.g., ABE8.20m) and "m" indicates "monomer." In some embodiments, TadA* is linked to a Cas9 nickase. In some embodiments, a fusion protein or complex of the disclosure comprises wild-type TadA (TadA(wt)) linked to TadA* as a heterodimer. Throughout this disclosure, adenosine deaminase base editors comprising a single TadA* domain and a TadA(wt) domain are referred to using the term ABEd or ABE#d, where "#" is an identifier number (e.g., ABE8.20d) and "d" indicates "dimer." In other embodiments, a fusion protein or complex of the disclosure comprises TadA*7.10 linked to TadA* as a heterodimer. In some embodiments, the base editor is ABE8, which comprises a TadA* mutant monomer. In some embodiments, the base editor is an ABE, which comprises a heterodimer of TadA* and TadA(wt). In some embodiments, the base editor is an ABE, which comprises a heterodimer of TadA* and TadA*7.10. In some embodiments, the base editor is an ABE, which comprises a heterodimer of TadA*. In some embodiments, TadA* is selected from Tables 5A-5E.

[0276] In some embodiments, adenosine deaminase is expressed as a monomer. In other embodiments, adenosine deaminase is expressed as a heterodimer. In some embodiments, the deaminase or other polypeptide sequence lacks methionine, for example, when included as a component of a fusion protein. This may result in a change in the numbering of positions. However, it is clear to those skilled in the art that such corresponding mutations refer to the same mutation.

[0277] Any of the mutations presented herein, and any additional mutations (e.g., based on the ecTadA amino acid sequence), can be introduced into any other adenosine deaminase. Any of the mutations presented herein can be made individually or in any combination in the TadA reference sequence or another adenosine deaminase (e.g., ecTadA).

[0278] Details of the A to G nucleobase editing protein are described in International PCT Application No. PCT / US2017 / 045381 (WO2018 / 027078) and Gaudelli, NM, et al., "Programmable base editing of A ·T to G ·C in genomic DNA without DNA cleavage" Nature, 551, 464-471 (2017), the entire contents of which are incorporated herein by reference.

[0279] C to T Edit In some embodiments, a base editor disclosed herein comprises a fusion protein or complex comprising a cytidine deaminase, which can deaminate a targeted cytidine (C) base of a polynucleotide to produce a uridine (U), which has the base pairing properties of thymine. In some embodiments, for example, when the polynucleotide is double-stranded (e.g., DNA), the uridine base can then be replaced with a thymidine base (e.g., by cellular repair machinery), resulting in a C:G to T:A transition. In other embodiments, deamination of a nucleic acid by a base editor from C to U cannot be accompanied by a U to T substitution.

[0280] Deamination of a targeted C in a polynucleotide to generate a U is a non-limiting example of the type of base editing that can be performed by the base editors described herein. In another example, a base editor comprising a cytidine deaminase domain can mediate the conversion of a cytosine (C) base to a guanine (G) base. For example, a U in a polynucleotide produced by deamination of a cytidine by the cytidine deaminase domain of a base editor can be removed from the polynucleotide by a base excision repair mechanism (e.g., by a uracil DNA glycosylase (UDG) domain) to generate an abasic site. The nucleobase opposite the abasic site can then be replaced with another base (e.g., C) by, for example, a translesion polymerase (e.g., by a base repair mechanism). Typically, the nucleobase opposite the abasic site is replaced with C, although other substitutions (e.g., A, G, or T) can also occur.

[0281] Thus, in some embodiments, the base editors described herein comprise a deamination domain (e.g., a cytidine deaminase domain) that can deaminate a targeted C in a polynucleotide to U. Additionally, as described below, the base editor, in some embodiments, can comprise an additional domain that promotes the conversion of the U resulting from deamination to T or G. For example, a base editor comprising a cytidine deaminase domain can further comprise a uracil glycosylase inhibitor (UGI) domain that mediates the substitution of U with T, completing the base editing of C to T. In another example, the base editor can comprise a uracil stabilizing protein, as described herein. In another example, the base editor can incorporate a translesion polymerase to improve the efficiency of C to G base editing, as the translesion polymerase can promote the incorporation of a C opposite the abasic site (i.e., resulting in the incorporation of a G at the abasic site, completing the base editing of C to G).

[0282] Base editors that contain cytidine deaminase as a domain can deaminate target Cs in any polynucleotide, including DNA, RNA, and DNA-RNA hybrids.

[0283] In some embodiments, the base editor cytidine deaminase comprises all or a portion (e.g., a functional portion) of an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. APOBEC is an evolutionarily conserved family of cytidine deaminases. Members of this family are C-to-U editing enzymes. The N-terminal domain of APOBEC-like proteins is the catalytic domain, and the C-terminal domain is the pseudocatalytic domain. More specifically, the catalytic domain is a zinc-dependent cytidine deaminase domain and is important for cytidine deamination. APOBEC family members include APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now referred to as "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine) deaminase.

[0284] Other exemplary deaminases that can be fused to Cas9 according to aspects of the present disclosure are shown below. In some embodiments, the deaminase is activation-induced deaminase (AID). Of course, in some embodiments, the active domain of each sequence can be used, for example, a domain without a localization signal (a domain without a nuclear localization sequence, a nuclear export signal, or a cytoplasmic localization signal).

[0285] Some aspects of the present disclosure are based on the recognition that modulating the catalytic activity of the deaminase domain of any of the fusion proteins or complexes described herein, for example, by introducing point mutations into the deaminase domain, alters the processivity of the fusion protein (e.g., base editor) or complex. For example, mutations that reduce, but do not eliminate, the catalytic activity of the deaminase domain within a base editing fusion protein or complex can reduce the likelihood that the deaminase domain will catalyze the deamination of residues adjacent to a target residue, thereby narrowing the scope of deamination. The ability to narrow the scope of deamination can prevent undesired deamination of residues adjacent to a particular target residue, reducing or preventing off-target effects.

[0286] In some embodiments, the APOBEC deaminase incorporated into the base editor can comprise one or more mutations selected from the group consisting of H121R, H122R, R126A, R126E, R118A, W90A, W90Y, and R132E of rAPOBEC1, D316R, D317R, R320A, R320E, R313A, W285A, W285Y, and R326E of hAPOBEC3G, and any alternative mutations at the corresponding positions, or one or more corresponding mutations in another APOBEC deaminase.

[0287] Several engineered cytidine deaminases are commercially available, including, but not limited to, SaBE3, SaKKH-BE3, VQR-BE3, EQR-BE3, VRER-BE3, YE1-BE3, EE-BE3, YE2-BE3, and YEE-BE3, which are available from Addgene (plasmids 85169, 85170, 85171, 85172, 85173, 85174, 85175, 85176, 85177). In some embodiments, the deaminase incorporated into the base editor comprises all or a portion (e.g., a functional portion) of APOBEC1 deaminase.

[0288] In some embodiments, the fusion protein or complex of the present disclosure comprises one or more cytidine deaminase domains. In some embodiments, the cytidine deaminase provided herein is capable of deaminating cytosine or 5-methylcytosine to uracil or thymine. In some embodiments, the cytidine deaminase provided herein is capable of deaminating cytosine in DNA. The cytidine deaminase can be derived from any suitable organism. In some embodiments, the cytidine deaminase is a naturally occurring cytidine deaminase containing one or more mutations corresponding to any of the mutations set forth herein. One of skill in the art can identify corresponding residues in any homologous protein, for example, by sequence alignment and determination of homologous residues. Thus, one of skill in the art can generate mutations in any naturally occurring cytidine deaminase corresponding to any of the mutations described herein. In some embodiments, the cytidine deaminase is derived from a prokaryote. In some embodiments, the cytidine deaminase is derived from a bacterium. In some embodiments, the cytidine deaminase is from a mammal (e.g., a human).

[0289] In some embodiments, the cytidine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any one of the cytidine deaminase amino acid sequences set forth herein. As will be apparent, the cytidine deaminases provided herein can include one or more mutations (e.g., any of the mutations set forth herein). Some embodiments provide polynucleotide molecules, which encode the cytidine deaminase nucleobase editor polypeptide of any of the above aspects or described herein. In some embodiments, the polynucleotide is codon-optimized.

[0290] In some embodiments, a fusion protein of the present disclosure comprises two or more nucleic acid editing domains.

[0291] Details of C to T nucleobase editing proteins are described in International PCT Application No. PCT / US2016 / 058344 (WO2017 / 070632) and Komor, AC, et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016), the entire contents of which are incorporated herein by reference.

[0292] Cytidine Adenosine Base Editor (CABE) In some embodiments, base editors described herein comprise adenosine deaminase mutants with increased cytidine deaminase activity. Such base editors may be referred to as "cytidine adenosine base editors (CABEs)" or "TadA*-derived cytosine base editors (CBE-T)," and their corresponding deaminase domains are "DNA cytosine (T)" deaminase mutants. ADC) domain." In some examples, the adenosine deaminase mutant has both adenine deaminase activity and cytosine deaminase activity (i.e., is a dual deaminase). In some embodiments, the adenosine deaminase mutant deaminates adenine and cytosine in DNA. In some embodiments, the adenosine deaminase mutant deaminates adenine and cytosine in single-stranded DNA. In some embodiments, the adenosine deaminase mutant deaminates adenine and cytosine in RNA. In some embodiments, the adenosine deaminase variant primarily deaminates cytosines in DNA and / or RNA (e.g., greater than 30%, greater than 40%, greater than 50%, greater than 60%, greater than 70%, greater than 80%, greater than 90%, greater than 95%, or greater than 99% of all deamidations catalyzed by the adenosine deaminase variant, or the number of cytosine deamidations catalyzed by the variant is about 2-fold or more, about 3-fold or more, about 4-fold or more, about 5-fold or more, about 6-fold or more, about 7-fold or more, about 8-fold or more, about 9-fold or more, about 10-fold or more, about 25-fold or more, about 50-fold or more, about 75-fold or more, about 100-fold or more, about 500-fold or more, or about 1,000-fold or more relative to the number of adenine deamidations catalyzed by the variant). In some embodiments, the adenosine deaminase mutant has approximately equal cytosine deaminase activity and adenosine deaminase activity (e.g., the two activities are within about 10% or 20% of each other). In some embodiments, the adenosine deaminase mutant has primarily cytosine deaminase activity and little, if any, adenosine deaminase activity. In some embodiments, the adenosine deaminase mutant has cytosine deaminase activity and no significant or detectable adenosine deaminase activity. In some embodiments, the target polynucleotide is present in a cell in vitro or in vivo. In some embodiments, the cell is a bacterial, yeast, fungal, insect, plant, or mammalian cell.

[0293] In some embodiments, the CABE comprises a bacterial TadA deaminase mutant (e.g., ecTadA). In some embodiments, the CABE comprises a truncated TadA deaminase mutant. In some embodiments, the CABE comprises a fragment of a TadA deaminase mutant. In some embodiments, the CABE comprises a TadA*8.20 mutant.

[0294] In some embodiments, an adenosine deaminase variant of the disclosure is a TadA adenosine deaminase that comprises one or more mutations that maintain adenosine deaminase activity (e.g., at least about 30%, 40%, 50%, or more of the activity of a reference adenosine deaminase (e.g., TadA*8.20 or TadA*8.19)) while increasing cytosine deaminase activity (e.g., at least about a 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, or more increase). In some examples, the adenosine deaminase variant comprises one or more mutations that increase cytosine deaminase activity compared to the activity of a reference adenosine deaminase (e.g., at least about a 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, or more increase), and result in undetectable adenosine deaminase activity or an adenosine deaminase activity that is less than 30%, 20%, 10%, or 5% of the reference adenosine deaminase activity. In some embodiments, the reference adenosine deaminase is TadA*8.20 or TadA*8.19.

[0295] In some embodiments, the adenosine deaminase variant is an adenosine deaminase that includes two or more mutations at amino acid positions selected from the group consisting of 2, 4, 6, 8, 13, 17, 23, 27, 29, 30, 47, 48, 49, 67, 76, 77, 82, 84, 96, 100, 107, 112, 114, 115, 118, 119, 122, 127, 142, 143, 147, 149, 158, 159, 162 165, 166, and 167 of an amino acid sequence having at least about 70%, 75%, 80%, 85%, 90%, 95%, 98%, 99% or more identity to SEQ ID NO:1, or corresponding mutations in another deaminase.

[0296] In some embodiments, the adenosine deaminase variants are selected from the group consisting of S2H, V4K, V4S, V4T, V4Y, F6G, F6H, F6Y, H8Q, R13G, T17A, T17W, R23Q, E27C, E27G, E27H, E27K, E27Q, E27S, E27G, P29A, P29G, P29K, V30F, V30I, R47G, R47S, A48G, I49K, I49M, I49N ... and S165P, or a corresponding mutation in another deaminase.

[0297] In some embodiments, the adenosine deaminase mutant is an adenosine deaminase that includes an amino acid mutation or combination of amino acid mutations selected from those shown in any of Tables 6A-6F.

[0298] The residue identities of exemplary adenosine deaminase mutants capable of deaminating adenine and / or cytidine in a target polynucleotide (e.g., DNA) are shown in Tables 6A-6F below. Further examples of adenosine deaminase mutants include the following mutant 1.17 (see Table 6A): 1.17+E27H, 1.17+E27K, 1.17+E27S, 1.17+E27S+I49K, 1.17+E27G, 1.17+I49N, 1.17+E27G+I49N, and 1.17+E27Q. In some embodiments, any of the amino acid mutations shown herein are substituted with a conservative amino acid. Additional mutations known in the art can be added to any of the adenosine deaminase mutants provided herein. In some embodiments, a base editor system comprising a CABE provided herein has at least about 30%, 40%, 50%, 60%, 70%, or more C to T editing activity in a target polynucleotide (e.g., DNA). In some embodiments, a base editor system comprising a CABE provided herein has increased C to T base editing activity (e.g., at least about 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, or more increase) compared to a reference base editor system comprising a reference adenosine deaminase (e.g., TadA*8.20 or TadA*8.19).

[0299] [Table 6A-1] [Table 6A-2]

[0300] [Table 6B-1] [Table 6B-2] [Table 6B-3] [Table 6B-4] [Table 6B-5] [Table 6B-6]

[0301] [Table 6C-1] [Table 6C-2]

[0302] [Table 6D-1] [Table 6D-2]

[0303] [Table 6E]

[0304] [Table 6F]

[0305] Guide polynucleotide The polynucleotide-programmable nucleotide-binding domain, when in conjunction with a bound guide polynucleotide (e.g., gRNA), can specifically bind to a target polynucleotide sequence (i.e., via complementary base pairing between bases of the bound guide nucleic acid and bases of the target polynucleotide sequence), thereby localizing the base editor to the target nucleic acid sequence requiring editing. In some embodiments, the target polynucleotide sequence comprises single-stranded DNA or double-stranded DNA. In some embodiments, the target polynucleotide sequence comprises RNA. In some embodiments, the target polynucleotide sequence comprises a DNA-RNA hybrid.

[0306] In one embodiment, the guide polynucleotide described herein can be RNA or DNA. In one embodiment, the guide polynucleotide is a gRNA.

[0307] In some embodiments, the guide polynucleotide is at least one single guide RNA ("sgRNA" or "gRNA"). In some embodiments, the guide polynucleotide comprises two or more individual polynucleotides, which can interact with each other, e.g., through complementary base pairing (e.g., dual guide polynucleotides, dual gRNAs). For example, the guide polynucleotide may comprise a CRISPR RNA (crRNA) and a trans-activating CRISPR RNA (tracrRNA), or may comprise one or more trans-activating CRISPR RNAs (tracrRNAs).

[0308] A guide polynucleotide can include natural or unnatural (or non-natural) nucleotides (e.g., peptide nucleic acids or nucleotide analogs). Optionally, the target region (e.g., spacer) of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length.

[0309] In some embodiments, the methods described herein can utilize engineered Cas proteins. Guide RNAs (gRNAs) are short, synthetic RNAs composed of a scaffold sequence required for Cas binding and a user-specified, approximately 20-nucleotide spacer that identifies the genomic target to be modified. Exemplary gRNA scaffold sequences are shown in the Sequence Listing as SEQ ID NOS: 317-327 and 441. Thus, one skilled in the art can vary the genomic target of a Cas protein. Specificity is determined in part by how specific the gRNA target sequence is for the genomic target relative to the rest of the genome. In certain embodiments, the spacer is approximately 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 22, 23, 24, 25, or more nucleotides in length. The gRNA spacer can be 19, 20, or 21 nucleotides in length, or approximately 19, 20, or 21 nucleotides in length.

[0310] The gRNA or guide polynucleotide can target any exon or intron of the gene target. In some embodiments, the composition includes multiple gRNAs that all target the same exon, or multiple gRNAs that target different exons. The exons and / or introns of the gene can be targeted. The gRNA or guide polynucleotide can target a nucleic acid sequence of about 20 nucleotides or less than about 20 nucleotides (e.g., at least about 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 nucleotides), or any number between about 1 and 100 nucleotides (e.g., 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 60, 70, 80, 90, 100). The target nucleic acid sequence can be 20 (or about 20) bases immediately 5' to the first nucleotide of the PAM. The gRNA can target the nucleic acid sequence. The target nucleic acid can be at least (or at least about) 1-10, 1-20, 1-30, 1-40, 1-50, 1-60, 1-70, 1-80, 1-90, or 1-100 nucleotides.

[0311] The guide polynucleotide can comprise standard ribonucleotides, modified ribonucleotides (e.g., pseudouridine), ribonucleotide isomers, and / or ribonucleotide analogs.

[0312] In some embodiments, a base editor system can include multiple guide polynucleotides (e.g., gRNAs). For example, gRNAs can target one or more target sites included in the base editor system (e.g., at least 1 gRNA, at least 2 gRNAs, at least 5 gRNAs, at least 10 gRNAs, at least 20 gRNAs, at least 30 gRNAs, at least 50 gRNAs). Multiple gRNA sequences can be arranged in tandem and separated by direct repeats.

[0313] Modified Polynucleotides To enhance expression, stability, and / or genome / base editing efficiency, and / or reduce potential toxicity, the base editor coding sequence (e.g., mRNA) and / or guide polynucleotide (e.g., gRNA) can be modified to include one or more modified nucleotides and / or chemical modifications, for example, using pseudouridine, 5-methylcytosine, 2'-O-methyl-3'-phosphonoacetic acid, 2'-O-methylthioPACE (MSP), 2'-O-methyl-PACE (MP), 2'-fluoroRNA (2'-F-RNA), =constrained ethyl (S-cEt), 2'-O-methyl ("M"), 2'-O-methyl-3'-phosphorothioate ("MS"), 2'-O-methyl-3'-thiophosphonoacetic acid ("MSP"), 5-methoxyuridine, phosphorothioate, and N1-methylpseudouridine. Chemically protected gRNAs can improve stability and editing efficiency in vivo and ex vivo. Methods for using chemically modified mRNAs and guide RNAs are known in the art and are described, for example, in Jiang et al., "Chemical modifications of adenine base editor mRNA and guide RNA expand its application scope. Nat Commun 11, 1979 (2020). doi.org / 10.1038 / s41467-020-15892-8," Callum et al., "N1-Methylpseudouridine substitution enhances the performance of synthetic mRNA switches in cells," Nucleic Acids Research, Volume 48, Issue 6, 06 April 2020, Page e35, and Andries et al., "Journal of Controlled Release," Volume 217, 10 November 2015, Pages 337-344, each of which is incorporated herein by reference in its entirety.

[0314] In some embodiments, the guide polynucleotide comprises one or more modified nucleotides at the 5' and / or 3' end of the guide. In some embodiments, the guide polynucleotide comprises two, three, four, or more modified nucleosides at the 5' and / or 3' end of the guide. In some embodiments, the guide polynucleotide comprises two, three, four, or more modified nucleosides at the 5' and / or 3' end of the guide.

[0315] In some embodiments, the guide comprises at least about 50%-75% modified nucleotides. In some embodiments, the guide comprises at least about 85% or more modified nucleotides. In some embodiments, at least about 1-5 nucleotides are modified at the 5' end of the gRNA and at least about 1-5 nucleotides are modified at the 3' end of the gRNA. In some embodiments, at least about 3-5 consecutive nucleotides are modified at each of the 5' and 3' ends of the gRNA. In some embodiments, at least about 20% of the nucleotides present in the direct repeat or inverted repeat are modified. In some embodiments, at least about 50% of the nucleotides present in the direct repeat or inverted repeat are modified. In some embodiments, at least about 50-75% of the nucleotides present in the direct repeat or inverted repeat are modified. In some embodiments, at least about 100 nucleotides present in the direct repeat or inverted repeat are modified. In some embodiments, at least about 20% or more of the nucleotides in the hairpin present in the gRNA scaffold are modified. In some embodiments, at least about 50% or more of the nucleotides in the hairpin present in the gRNA scaffold are modified. In some embodiments, the guide comprises a spacer of variable length. In some embodiments, the guide comprises a spacer of 20-40 nucleotides. In some embodiments, the guide comprises a spacer comprising at least about 20-25 nucleotides or at least about 30-35 nucleotides. In some embodiments, the spacer comprises modified nucleotides. In some embodiments, the guide has two or more of the following characteristics: at least about 1-5 nucleotides at the 5' end of the gRNA are modified and at least about 1-5 nucleotides at the 3' end of the gRNA are modified; at least about 20% of the nucleotides present in the direct repeat or inverted repeat are modified; At least about 50-75% of the nucleotides present in the direct or inverted repeat sequences are modified; At least approximately 20% or more of the nucleotides in the hairpin present in the gRNA scaffold are modified; a spacer of variable length; and A spacer comprising a modified nucleotide.

[0316] In some embodiments, the gRNA contains multiple modified nucleotides and / or chemical modifications ("heavy modifications"). Such heavy modifications can increase base editing in vivo or in vitro by about two-fold. In some embodiments, the gRNA contains 2'-O-methyl or phosphorothioate modifications. In one embodiment, the gRNA contains 2'-O-methyl and phosphorothioate modifications. In one embodiment, the modifications increase base editing by at least about two-fold.

[0317] The guide polynucleotide can include one or more modifications to result in a nucleic acid with new or enhanced characteristics. The guide polynucleotide can include a nucleic acid affinity tag. The guide polynucleotide can include synthetic nucleotides, synthetic nucleotide analogs, nucleotide derivatives, and / or modified nucleotides.

[0318] Additionally, gRNA or guide polynucleotides can be modified with the following: 5' adenylate, 5' guanosine triphosphate cap, 5' N7-methylguanosine triphosphate cap, 5' triphosphate cap, 3' phosphate, 3' thiophosphate, 5' phosphate, 5' thiophosphate, Cis-Syn thymidine dimer, trimer, C12 spacer, C3 spacer, C6 spacer, d spacer, PC spacer, r spacer, spacer 18, spacer 9, 3'-3' modification, 2'-O-methylthioPACE (MSP), 2'-O-methyl-PACE (MP), and constrained ethyl (S-cEt), 5'-5' modification, abasic, acridine, azobenzene, biotin, biotin BB, biotin TEG, cholesteryl TEG, desthiobiotin TEG, DNP. TEG, DNP-X, DOTA, dT-biotin, dual biotin, PC-biotin, psoralen C2, psoralen C6, TINA, 3'DABCYL, black hole quencher 1, black hole quencher 2, DABCYL SE, dT-DABCYL, IRDye QC-1, QSY-21, QSY-35, QSY-7, QSY-9, carboxyl linker, thiol linker, 2'-deoxyribonucleoside analogs purine, 2'-deoxyribonucleoside analogs pyrimidine, ribonucleoside analog, 2'-O-methylribonucleoside analog, sugar modified analog, wobble / universal base, fluorescent dye label, 2'-fluoro RNA, 2'-O-methyl RNA, methyl phosphonate, phosphodiester DNA, phosphodiester RNA, phosphothioate DNA, phosphorothioate RNA, UNA, pseudouridine-5'-triphosphate, 5'-methylcytidine-5'-triphosphate, or any combination thereof.

[0319] In some cases, phosphorothioate-enhanced gRNAs can inhibit RNase A, RNase T1, bovine serum nuclease, or any combination thereof. These properties allow the use of PS-RNA gRNAs in applications where exposure to nucleases is likely to occur in vivo or in vitro. For example, phosphorothioate (PS) bonds can be introduced between the last three to five nucleotides at the 5' or 3' end of the gRNA, which can inhibit exonuclease degradation. In some cases, phosphorothioate bonds can be added throughout the gRNA to reduce attack by endonucleases.

[0320] Fusion proteins or complexes containing nuclear localization sequences (NLS) In some embodiments, the fusion proteins or complexes provided herein further comprise one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, e.g., nuclear localization sequences (NLSs). In one embodiment, a bipartite NLS is used. In some embodiments, the NLS comprises an amino acid sequence that promotes import of the NLS-containing protein into the cell nucleus (e.g., by nuclear transport). In some embodiments, the NLS is fused to the N- or C-terminus of the fusion protein. In some embodiments, the NLS is fused to the C- or N-terminus of the nCas9 domain or dCas9 domain. In some embodiments, the NLS is fused to the N- or C-terminus of the Cas12 domain. In some embodiments, the NLS is fused to the N- or C-terminus of the cytidine deaminase or adenosine deaminase. In some embodiments, the NLS is fused to the fusion protein via one or more linkers. In some embodiments, the NLS is fused to the fusion protein without a linker. In some embodiments, NLS comprises any one of the amino acid sequences of the NLS sequences shown or referenced herein.Other nuclear localization sequences are known in the art and will be clear to those skilled in the art.For example, NLS sequences are described in Plank et al., PCT / EP2000 / 011690 (the contents of which are incorporated herein by reference for disclosure of exemplary nuclear localization sequences).

[0321] In some embodiments, the NLS is present within a linker, or the NLS is adjacent to a linker, for example, as described herein. A bi-knot NLS contains two clusters of basic amino acids separated by a relatively short spacer sequence (thus, a bi-knot is two parts, while a mono-knot NLS is not). The nucleoplasmin NLS, KR[PAATKKAGQA]KKKK (SEQ ID NO: 191), is a prototype of a ubiquitous bi-knot signal, and is two clusters of basic amino acids separated by a spacer of about 10 amino acids. The sequence of an exemplary bi-knot NLS is as follows: PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 328).

[0322] In some embodiments, any of the fusion proteins or complexes provided herein comprises an NLS comprising the amino acid sequence EGADKRTADGSEFESPKKKRKV (amino acids 8-29 of SEQ ID NO: 328). In some embodiments, any of the adenosine base editors provided herein comprises an NLS comprising the amino acid sequence EGADKRTADGSEFESPKKKRKV (amino acids 8-29 of SEQ ID NO: 328). In some embodiments, the NLS is at the C-terminal portion of the adenosine base editor. In some embodiments, the NLS is at the C-terminus of the adenosine base editor.

[0323] Additional Domains The base editors described herein can include any domain that assists in facilitating nucleobase editing, modification, or alteration of a polynucleotide. In some embodiments, the base editor comprises a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9), a nucleobase editing domain (e.g., a deaminase domain), and one or more additional domains. In some embodiments, the additional domains can facilitate the enzymatic or catalytic function of the base editor, the binding function of the base editor, or can be inhibitors of cellular machinery (e.g., enzymes) that can interfere with the desired base editing result. In some embodiments, the base editor comprises a nuclease, nickase, recombinase, deaminase, methyltransferase, methylase, acetylase, acetyltransferase, transcriptional activator, or transcriptional repressor domain.

[0324] In some embodiments, the base editor comprises a uracil glycosylase inhibitor (UGI) domain. Optionally, the base editor is expressed in trans in a cell along with a UGI polypeptide. In some embodiments, a cell's DNA repair response to the presence of U:G heteroduplex DNA can cause reduced nucleic acid base editing efficiency in the cell. In such embodiments, uracil DNA glycosylase (UDG) can catalyze the removal of U from DNA in the cell and initiate base excision repair (BER), primarily resulting in the reversion of U:G pairings to C:G pairings. In such embodiments, BER can be inhibited in base editors that include one or more domains that bind to a single strand, block the edited base, inhibit UGI, inhibit BER, protect the edited base, and / or promote repair of the unedited strand. Accordingly, the present disclosure contemplates base editor fusion proteins or complexes that include a UGI domain and / or a uracil stabilizing protein (USP) domain.

[0325] Base editors Provided herein are systems, compositions, and methods for editing nucleobases using a base editor system. In some embodiments, the base editor system includes: (1) a base editor (BE) comprising a polynucleotide-programmable nucleotide-binding domain and a nucleobase editing domain (e.g., a deaminase domain) for editing nucleobases; and (2) a guide polynucleotide (e.g., a guide RNA) used in combination with the polynucleotide-programmable nucleotide-binding domain. In some embodiments, the base editor system is a cytidine base editor (CBE) or an adenosine base editor (ABE). In some embodiments, the polynucleotide-programmable nucleotide-binding domain is a polynucleotide-programmable DNA-binding domain or RNA-binding domain. In some embodiments, the nucleobase editing domain is a deaminase domain. In some embodiments, the deaminase domain can be a cytidine deaminase or cytosine deaminase. In some embodiments, the deaminase domain can be an adenine deaminase or an adenosine deaminase. In some embodiments, the adenosine base editor can deaminate adenines in DNA. In some embodiments, the base editor is capable of deaminating cytidines in DNA.

[0326] Use of the base editor system provided herein includes: (a) contacting a target nucleotide sequence of a polynucleotide of interest (e.g., double-stranded or single-stranded DNA or RNA) with a base editor system comprising a nucleobase editor (e.g., an adenosine base editor or cytidine base editor) and a guide polynucleotide (e.g., gRNA), wherein the target nucleotide sequence comprises a target nucleobase pair; (b) inducing strand separation of the target region; (c) converting a first nucleobase of the target nucleobase pair to a second nucleobase in one strand of the target region; and (d) cleaving no more than one strand of the target region, wherein a third nucleobase complementary to the first nucleobase is replaced with a fourth nucleobase complementary to the second nucleobase. It should be understood that in some embodiments, step (b) is omitted. In some embodiments, the target nucleobase pair is a plurality of nucleobase pairs in one or more genes. In some embodiments, the base editor systems provided herein are capable of multiplex editing of multiple nucleobase pairs in one or more genes. In some embodiments, the multiple nucleobase pairs are located in the same gene. In some embodiments, the multiple nucleobase pairs are located in one or more genes, and at least one gene is located at a different locus.

[0327] The components of the base editor system (e.g., the deaminase domain, the guide RNA, and / or the polynucleotide-programmable nucleotide binding domain) may be linked to each other by covalent or non-covalent bonds. For example, in some embodiments, the deaminase domain can be targeted to a target nucleotide sequence by a polynucleotide-programmable nucleotide binding domain, and optionally, the polynucleotide-programmable nucleotide binding domain is complexed with a polynucleotide (e.g., a guide RNA). In some embodiments, the polynucleotide-programmable nucleotide binding domain can be fused or linked to the deaminase domain. In some embodiments, the polynucleotide-programmable nucleotide binding domain can non-covalently interact with or bind to the deaminase domain, thereby enabling the deaminase domain to target the target nucleotide sequence. For example, in some embodiments, the nucleobase editing component (e.g., a deaminase component) comprises an additional heterologous moiety or heterologous domain that can interact with, bind to, or form a complex with a corresponding heterologous moiety, heterologous antigen, or heterologous domain that is part of the polynucleotide-programmable nucleotide-binding domain and / or the guide polynucleotide (e.g., guide RNA) complexed therewith. In some embodiments, the polynucleotide-programmable nucleotide-binding domain and / or the guide polynucleotide (e.g., guide RNA) complexed therewith comprises an additional heterologous moiety or heterologous domain that can interact with, bind to, or form a complex with a corresponding heterologous moiety, heterologous antigen, or heterologous domain that is part of the nucleobase editing domain (e.g., a deaminase component). In some embodiments, the additional heterologous moiety can bind to, interact with, associate with, or form a complex with a polypeptide. In some embodiments,The additional heterologous moiety may bind to, interact with, associate with, or form a complex with a polynucleotide. In some embodiments, the additional heterologous moiety may be attached to a guide polynucleotide. In some embodiments, the additional heterologous moiety may be attached to a polypeptide linker. In some embodiments, the additional heterologous moiety may be attached to a polynucleotide linker. The additional heterologous moiety may be a protein domain. In some embodiments, the additional heterologous moiety is a polypeptide, e.g., the 22 amino acid RNA binding domain of the anti-terminator protein N of bacteriophage lambda (N22p), a 2G12 IgG homodimer domain, an ABI, an antibody (e.g., an antibody that binds to a component of a base editor system or a heterologous portion thereof) or a fragment thereof (e.g., heavy chain domain 2 (CH2) of IgM (MHD2) or IgE (EHD2), an immunoglobulin Fc region, heavy chain domain 3 (CH3) of IgG or IgA, heavy chain domain 4 (CH4) of IgM or IgE, Fab, Fab2, miniantibody, and / or ZIP antibody), a barnase-barstar dimer domain, a Bcl-xL domain, a calcineurin A (CAN) domain, a cardiac phospholamban transmembrane pentamer domain, a collagen domain, a ComRNA binding protein domain (e.g., an SfMu Com coat protein domain and an SfMU Com binding protein domain), a cyclophilin-Fas fusion protein (CyP-Fas) domain, Fab domain, Fe domain, fibritin foldon domain, FK506-binding protein (FKBP) domain, FKBP-binding domain of mTOR (FRB) domain, foldon domain, fragment X domain, GAI domain, GID1 domain, glycophorin A transmembrane domain, GyrB domain, Halo tag, HIV Gp41 trimerization domain, HPV45 oncoprotein E7 C-terminal dimerization domain, hydrophobic polypeptide, K homology (KH) domain, Ku protein domain (e.g., Ku heterodimer), leucine zipper, LOV domain,These include a mitochondrial antiviral signaling protein CARD filament domain, an MS2 coat protein domain (MCP), a non-natural RNA aptamer ligand that binds to a corresponding RNA motif / aptamer, a parathyroid hormone dimerization domain, a PP7 coat protein (PCP) domain, a PSD95-Dlgl-zo-1 (PDZ) domain, a PYL domain, a SNAP tag, a SpyCatcher moiety, a SpyTag moiety, a streptavidin domain, a streptavidin-binding protein domain, a streptavidin-binding protein (SBP) domain, a telomerase Sm7 protein domain (e.g., an Sm7 homoheptamer or a monomeric Sm-like protein), and / or fragments thereof. In certain embodiments, the additional heterologous moiety comprises a polynucleotide (e.g., an RNA motif), such as an MS2 phage operator stem-loop (e.g., MS2, MS2 C-5 mutant, or MS2 F-5 mutant), a non-naturally occurring RNA motif, a PP7 operator stem-loop, an SfMu phage Com stem-loop, a sterile alpha motif (SAM), a telomerase Ku binding motif, a telomerase Sm7 binding motif, and / or a fragment thereof. Non-limiting examples of additional heterologous moieties include polypeptides having at least about 85% sequence identity to any one or more of SEQ ID NOs: 380, 382, ​​384, 386-388, or fragments thereof. Non-limiting examples of additional heterologous moieties include polynucleotides having at least about 85% sequence identity to any one or more of SEQ ID NOs: 379, 381, 383, 385, or fragments thereof.

[0328] In some examples, components of the base editing system bind to each other through the interaction of leucine zipper domains (e.g., SEQ ID NOs: 387 and 388). Optionally, components of the base editing system bind to each other through polypeptide domains (e.g., FokI domains). The polypeptide domains bind to form a protein complex comprising about 1, 2 (i.e., dimerized), 3, 4, 5, 6, 7, 8, 9, 10 polypeptide domain units, or at least about 1, 2 (i.e., dimerized), 3, 4, 5, 6, 7, 8, 9, 10 polypeptide domain units, or no more than about 1, 2 (i.e., dimerized), 3, 4, 5, 6, 7, 8, 9, 10 polypeptide domain units, and optionally, the polypeptide domains may contain mutations that reduce or eliminate their activity.

[0329] In some examples, the components of the base editing system bind to each other through interactions of a multimeric antibody or fragment thereof (e.g., heavy chain domain 2 (CH2) of IgG, IgD, IgA, IgM, IgE, IgM (MHD2) or IgE (EHD2), immunoglobulin Fc region, heavy chain domain 3 (CH3) of IgG or IgA, heavy chain domain 4 (CH4) of IgM or IgE, Fab, and Fab2). In some examples, the antibody is a dimer, trimer, or tetramer. In certain embodiments, the dimeric antibody binds to a polypeptide component or a polynucleotide component of the base editing system.

[0330] In some cases, the components of the base editing system bind to each other through interactions between polynucleotide binding protein domain(s) and polynucleotide(s). In some examples, the components of the base editing system bind to each other through interactions between one or more polynucleotide binding protein domains and self-complementary and / or mutually complementary polynucleotides, such that their respective bound polynucleotide binding protein domain(s) are bound by complementary binding between the polynucleotides.

[0331] In some examples, components of a base editing system bind to each other through interactions between a polypeptide domain(s) and a small molecule(s) (e.g., a dimerization-inducing compound (CID), also known as a "dimerizer"). Non-limiting examples of CIDs include Amara, et al., "A versatile synthetic dimerizer for the regulation of protein-protein interactions," PNAS, 94:10618-10623 (1997), and Voss, et al. "Chemically induced dimerization: reversible and spatiotemporal control of protein function in cells," Current Opinion in Chemical Biology, 28:194-201 (2015), the disclosures of each of which are incorporated by reference in their entirety for all purposes. In some embodiments, the base editor inhibits base excision repair (BER) of the edited strand. In some embodiments, the base editor protects or binds to the unedited strand. In some embodiments, the base editor has UGI activity or USP activity. In some embodiments, the base editor comprises a catalytically incompetent inosine-specific nuclease.

[0332] The base editors of the present disclosure can have any domain, feature, or amino acid sequence that facilitates editing of a target polynucleotide sequence. For example, in some embodiments, the base editor comprises a nuclear localization sequence (NLS). In some embodiments, the NLS of the base editor is located between the deaminase domain and the polynucleotide-programmable nucleotide-binding domain. In some embodiments, the NLS of the base editor is located C-terminal to the polynucleotide-programmable nucleotide-binding domain.

[0333] The protein domain included in the fusion protein may be a heterologous functional domain. Non-limiting examples of protein domains that can be included in the fusion protein include a deaminase domain (e.g., cytidine deaminase and / or adenosine deaminase), a uracil glycosylase inhibitor (UGI) domain, an epitope tag, and a reporter gene sequence.

[0334] In some embodiments, an adenosine base editor (ABE) can deaminate adenines in DNA. In some embodiments, an ABE is generated by replacing the APOBEC1 component of BE3 with native or engineered E. coli TadA, human ADAR2, mouse ADA, or human ADAT2. In some embodiments, the ABE comprises an evolved TadA variant. In some embodiments, the base editor is ABE8.1, which comprises or consists essentially of SEQ ID NO:331 or a fragment thereof having adenosine deaminase activity. Other ABE8 sequences are set forth in the accompanying sequence listing (SEQ ID NOs:332-354).

[0335] In some embodiments, the base editor comprises an adenosine deaminase mutant, wherein the adenosine deaminase mutant has an amino acid sequence that includes a mutation relative to the ABE 7*10 reference sequence, as described herein. As used in Table 7A, the term "monomer" refers to a monomeric form of TadA*7.10 that includes the described mutation. As used in Table 7, the term "heterodimer" refers to a particular wild-type E. coli TadA adenosine deaminase fused to TadA*7.10 that includes the described mutation.

[0336] [Table 7A]

[0337] In some embodiments, a base editor comprises a domain that includes all or a portion (e.g., a functional portion) of a uracil glycosylase inhibitor (UGI) domain or a uracil stabilizing protein (USP) domain.

[0338] Linker In certain embodiments, a linker may be used to link any of the peptides or peptide domains of the present disclosure. The linker may be as simple as a covalent bond, or may be a polymeric linker of multiple atoms in length. In certain embodiments, the linker is a polypeptide or based on amino acids. In other embodiments, the linker is not peptide-like. In certain embodiments, the linker is a covalent bond (e.g., a carbon-carbon bond, a disulfide bond, a carbon-heteroatom bond, etc.).

[0339] In some embodiments, any of the fusion proteins provided herein comprise a cytidine deaminase or adenosine deaminase and a Cas9 domain fused to each other via a linker. Linkers of various lengths and flexibility (e.g., (GGGS) n (SEQ ID NO: 246), (GGGGS) n (SEQ ID NO: 247), and (G)n from a highly flexible linker of the form (EAAAK) n (SEQ ID NO: 248), (SGGS) n(SEQ ID NO: 355), SGSETPGTSESATPES (SEQ ID NO: 249) (see, e.g., Guilinger JP, et al. Fusion of catalytically inactive Cas9 to FokI nuclease improves the specificity of genome modification. Nat. Biotechnol. 2014;32(6): 577-82, the entire contents of which are incorporated herein by reference), and more rigid linkers of the form (XP)n) can be used to achieve an optimal length for cytidine deaminase or adenosine deaminase nucleobase editor activity. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises a (GGS)n motif, where n is 1, 3, or 7. In some embodiments, the cytidine deaminase or adenosine deaminase and Cas9 domain of any of the fusion proteins provided herein are fused via a linker (which may also be referred to as an XTEN linker) comprising the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 249).

[0340] In some embodiments, the base editor domain has the amino acid sequence SGGSSGSETPGTSESATPESSGGS (SEQ ID NO: 356), The fusion occurs via a linker comprising SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 357), or GGSGGSPGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATSGGSGGS (SEQ ID NO: 358).

[0341] In some embodiments, the base editor domains are fused via a linker (which may also be referred to as an XTEN linker) comprising the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 249). In some embodiments, the linker comprises the amino acid sequence SGGS (SEQ ID NO: 355). In some embodiments, the linker is 24 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPES (SEQ ID NO: 359). In some embodiments, the linker is 40 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGS (SEQ ID NO: 360). In some embodiments, the linker is 64 amino acids in length. In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSSGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 361). In some embodiments, the linker is 92 amino acids in length. In some embodiments, the linker comprises the amino acid sequence PGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSEGSAPGTSESATPESGPGSEPATS (SEQ ID NO: 362).

[0342] In some embodiments, the linker comprises multiple proline residues and is 5 to 21, 5 to 14, 5 to 9, or 5 to 7 amino acids in length, e.g., PAPAP (SEQ ID NO: 363), PAPAPA (SEQ ID NO: 364), PAPAPAP (SEQ ID NO: 365), PAPAPAPA (SEQ ID NO: 366), P(AP) (SEQ ID NO: 367), P(AP) (SEQ ID NO: 368), or P(AP) (SEQ ID NO: 369) (see, e.g., Tan J, Zhang F, Karcher D, Bock R. Engineering of high-precision base editors for site-specific single nucleotide replacement. Nat Commun. 2019 Jan 25;10(1):439, the entire contents of which are incorporated herein by reference). Such proline-rich linkers are also referred to as "rigid" linkers.

[0343] Nucleic acid-programmable DNA-binding proteins used in conjunction with guide RNAs Compositions and methods for base editing in cells are provided herein. Further provided herein are compositions comprising a guide polynucleotide sequence (e.g., a guide RNA sequence) or a combination of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more guide RNAs provided herein. In some embodiments, the compositions for base editing provided herein further comprise a polynucleotide encoding a base editor (e.g., a C base editor or an A base editor). For example, a composition for base editing may comprise an mRNA sequence encoding BE, BE4, ABE, and a combination of one or more of the provided guide RNAs. A composition for base editing may comprise a base editor polypeptide and a combination of one or more of any of the guide RNAs provided herein. Such compositions may be used to perform base editing in cells using different delivery approaches (e.g., electroporation, nucleofection, viral transduction, or transfection). In some embodiments, the composition for base editing comprises an mRNA sequence encoding a base editor and a combination of one or more guide RNA sequences provided herein for electroporation.

[0344] Some aspects of the present disclosure provide systems including any of the fusion proteins or complexes provided herein and a guide RNA that binds to a nucleic acid-programmable DNA-binding protein (napDNAbp) domain (e.g., Cas9 (e.g., dCas9, nuclease-active Cas9, or Cas9 nickase) or Cas12), wherein the napDNAbp domain is of the fusion protein or complex. Such complexes are also referred to as ribonucleoproteins (RNPs). In some embodiments, the guide nucleic acid (e.g., guide RNA) is 15-100 nucleotides in length and comprises a sequence of at least 10 contiguous nucleotides complementary to a target sequence. In some embodiments, the target sequence is a DNA sequence. In some embodiments, the target sequence is an RNA sequence. In some embodiments, the target sequence is a sequence within a bacterial, yeast, fungal, insect, plant, or animal genome. In some embodiments, the target sequence is a sequence within a human genome. In some embodiments, the 3' end of the target sequence is immediately adjacent to a canonical PAM sequence (NGG). In some embodiments, the 3' end of the target sequence is immediately adjacent to a non-canonical PAM sequence (e.g., a sequence shown in Table 3 or 5'-NAA-3'). In some embodiments, the guide nucleic acid (e.g., guide RNA) is complementary to a sequence of a gene of interest (e.g., a gene associated with a disease or disorder).

[0345] Some embodiments of the present disclosure provide methods of using the fusion proteins or complexes provided herein. For example, some embodiments of the present disclosure provide methods that include contacting a DNA molecule with any of the fusion proteins or complexes provided herein and at least one guide RNA, wherein the guide RNA is about 15 to 100 nucleotides in length and comprises a sequence of at least 10 consecutive nucleotides complementary to a target sequence.

[0346] The domains of the base editors disclosed herein can be arranged in any order.

[0347] The predetermined target region can be a deamination window (deamidatable region). The deamination window can be a predetermined region in which a base editor acts on and deaminates a target nucleotide. In some embodiments, the deamination window is within a region of 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases. In some embodiments, the deamination window is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 bases upstream of the PAM.

[0348] Base editors of the present disclosure can have any domain, feature, or amino acid sequence that facilitates editing of a target polynucleotide sequence.

[0349] Methods of using fusion proteins or complexes comprising cytidine deaminase or adenosine deaminase and a Cas9 domain

[0350] Some embodiments of the present disclosure provide methods for using the fusion proteins or complexes provided herein.For example, some embodiments of the present disclosure provide methods comprising contacting a DNA molecule with any of the fusion proteins or complexes provided herein and at least one guide RNA described herein.

[0351] In some embodiments, the fusion proteins or complexes of the present disclosure are used to edit a target gene of interest. In particular, the cytidine deaminase or adenosine deaminase nucleobase editors described herein can introduce multiple mutations within the target sequence. These mutations can affect the function of the target. For example, targeting a regulatory region with a cytidine deaminase or adenosine deaminase nucleobase editor alters the function of the regulatory region, reducing or eliminating expression of a downstream protein.

[0352] Base editor efficiency In some embodiments, the purpose of the methods provided herein is to modify genes and / or gene products through gene editing. The nucleobase editing proteins provided herein can be used in vitro or in vivo for gene editing-based human therapy. As will be apparent to one of skill in the art, the nucleobase editing proteins provided herein (e.g., fusion proteins or complexes comprising a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9) and a nucleobase editing domain (e.g., an adenosine deaminase domain or a cytidine deaminase domain)) can be used to edit nucleotides from A to G or C to T.

[0353] Advantageously, the base editing systems provided herein provide genome editing without generating double-stranded DNA breaks, without requiring a donor DNA template, and without inducing excessive accidental insertions and deletions like CRISPR. In some embodiments, the base editors provided herein efficiently generate intended mutations, such as stop codons, in nucleic acids (e.g., nucleic acids in a subject's genome) without generating a significant number of unintended mutations, e.g., unintended point mutations.

[0354] The disclosed base editors advantageously modify specific protein-encoding nucleotide bases without generating a significant proportion of indels (i.e., insertions or deletions), which can result in frameshift mutations within the coding region of a gene.

[0355] In some embodiments, the base editors provided herein are capable of producing a ratio of intended mutations to indels (i.e., intended point mutations:unintended point mutations) of greater than 1:1. In some embodiments, base editors provided herein can result in a ratio of intended mutations to indels that is at least 1.5:1, at least 2:1, at least 2.5:1, at least 3:1, at least 3.5:1, at least 4:1, at least 4.5:1, at least 5:1, at least 5.5:1, at least 6:1, at least 6.5:1, at least 7:1, at least 7.5:1, at least 8:1, at least 10:1, at least 12:1, at least 15:1, at least 20:1, at least 25:1, at least 30:1, at least 40:1, at least 50:1, at least 100:1, at least 200:1, at least 300:1, at least 400:1, at least 500:1, at least 600:1, at least 700:1, at least 800:1, at least 900:1, or at least 1000:1, or more. The number of intended mutations and indels can be determined using any suitable method.

[0356] In some embodiments, the base editors provided herein can limit the generation of indels in a region of a nucleic acid. In some embodiments, this region is at a nucleotide targeted by the base editor or within 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of a nucleotide targeted by the base editor. In some embodiments, any of the base editors provided herein can limit the generation of indels in a region of a nucleic acid to less than 1%, less than 1.5%, less than 2%, less than 2.5%, less than 3%, less than 3.5%, less than 4%, less than 4.5%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, less than 10%, less than 12%, less than 15%, or less than 20%.

[0357] Base editing is often referred to as a "modification," such as a genetic modification, gene modification, and modification of a nucleic acid sequence, and can be clearly understood based on the context in which the modification is a base editing modification. Thus, a base editing modification is a modification at the nucleotide base level, for example, as a result of deaminase activity as discussed throughout this disclosure, which can then result in a change in the gene sequence and affect the gene product.

[0358] In some embodiments, the modification, e.g., a single base edit, results in a reduction of gene target expression by about 10% or more, about 15% or more, about 20% or more, about 25% or more, about 30% or more, about 35% or more, about 40% or more, about 45% or more, about 50% or more, about 55% or more, about 60% or more, about 65% or more, about 70% or more, about 75% or more, about 80% or more, about 85% or more, about 90% or more, about 95% or more, about 99% or more, or 100%, or to an undetectable level.

[0359] The present disclosure provides adenosine deaminase mutants (e.g., ABE8 mutants) with improved efficiency and specificity. In particular, the adenosine deaminase mutants described herein are more likely to edit desired bases in a polynucleotide and less likely to edit bases that are not intended to be modified (e.g., "bystanders").

[0360] In some embodiments, any of the base editor systems comprising one of the ABE8 base editor variants described herein exhibits at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% reduction in bystander editing or mutations compared to a base editor system comprising an ABE7 base editor (e.g., ABE7.10). In some embodiments, any of the ABE8 base editor variants described herein have higher base editing efficiency compared to an ABE7 base editor. In some embodiments, any of the ABE8 base editor variants described herein have at least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 100%, 105%, 110%, 115%, 120%, 125%, 130%, 135%, 140%, 145%, 150%, 155%, 160%, 165%, 170%, 175%, 180%, 185%, 185%, 190%, 195%, 200%, 205%, 210%, 215%, 220%, 225%, 230%, 235%, 240%, 245%, 250%, 255%, 260%, 265%, 270%, 275%, 280%, 285%, 290%, 300%, 310%, 315%, 320%, 325%, 330%, 335%, 340%, 345%, 350%, 355%, 360%, 365%, 370%, 375%, 380%, 385%, 390%, 400%, 410%, 410%, 420%, 425%, 430%, 430%, 440%, 445%, 450%, 450%, 460%, 470%, 475%, 480%, 485%, 30%, 135%, 140%, 145%, 150%, 155%, 160%, 165%, 170%, 175%, 180%, 185%, 190%, 195%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 310%, 320%, 330%, 340%, 350%, 360%, 370%, 380%, 390%, 400%, 450%, or 500% higher base editing efficiency.

[0361] The ABE8 base editor mutants described herein can be delivered to a host cell via a plasmid, vector, LNP complex, or mRNA, hi some embodiments, any of the ABE8 base editor mutants described herein are delivered to a host cell as mRNA.

[0362] In some embodiments, the methods described herein, e.g., base editing methods, have minimal or no off-target effects, hi some embodiments, the methods described herein, e.g., base editing methods, result in minimal or no chromosomal translocations.

[0363] In some embodiments, the base editing methods described herein successfully edit about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of a population of cells.

[0364] In some embodiments, the percentage of viable cells in the cell population after base editing intervention is at least 60%, 70%, 80%, or more than 90% of the starting cell population at the time of base editing. In some embodiments, the percentage of viable cells in the edited cell population is about 70%. In some embodiments, the percentage of viable cells in the edited cell population is about 75%. In some embodiments, the percentage of viable cells in the edited cell population is about 80%. In some embodiments, the percentage of viable cells in the cell population is about 85%. In some embodiments, the percentage of viable cells in the cell population is about 90%, or about 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the cell population at the time of base editing.

[0365] In an embodiment, the cell population is a population of cells that has been contacted with a base editor, complex, or base editor system of the disclosure.

[0366] The number of intended mutations and indels can be determined, for example, as described in International PCT Applications PCT / US2017 / 045381 (WO2018 / 027078) and PCT / US2016 / 058344 (WO2017 / 070632), Komor, AC, et al., "Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage" Nature 533, 420-424 (2016); Gaudelli, NM, et al., "Programmable base editing of A·T to G·C in genomic DNA without DNA cleavage" Nature 551, 464-471 (2017), and Komor, AC, et al., "Improved base excision repair inhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editors with higher efficiency and product purity" Science Advances 3:eaao4774 (2017), the entire contents of which are incorporated herein by reference.

[0367] In some embodiments, to calculate indel frequency, sequencing reads are scanned for exact matches with two 10-bp sequences flanking the region where an indel may occur. If an exact match is not found, the read is excluded from analysis. If the length of this indel region perfectly matches the reference sequence, the read is classified as not containing an indel. If the indel region is two or more bases longer or shorter than the reference sequence, the sequencing read is classified as an insertion or deletion, respectively. In some embodiments, the base editors provided herein can limit the generation of indels in a region of a nucleic acid. In some embodiments, this region is located at the nucleotide targeted by the base editor or within 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides of the nucleotide targeted by the base editor.

[0368] Multiple Editing In some embodiments, the base editor systems provided herein are capable of multiplex editing of multiple nucleic acid base pairs in one or more genes or polynucleotide sequences. In some embodiments, the multiple nucleic acid base pairs are located within the same gene or within one or more genes, wherein at least one gene is located within a different locus. In some embodiments, the multiplex editing comprises one or more guide polynucleotides. In some embodiments, the multiplex editing comprises one or more base editor systems. In some embodiments, the multiplex editing comprises one or more base editor systems with a single guide polynucleotide or multiple guide polynucleotides. In some embodiments, the multiplex editing comprises one or more guide polynucleotides associated with a single base editor system. As will be apparent, the features of multiplex editing using any of the base editors described herein can be applied to any combination of methods using any of the base editors provided herein. As will be apparent, the multiplex editing using any of the base editors described herein can also comprise sequential editing of multiple nucleic acid base pairs.

[0369] In some embodiments, the base editor system capable of multiplex editing of multiple nucleic acid base pairs in one or more genes comprises one of an ABE7 base editor, an ABE8 base editor, and / or an ABE9 base editor.

[0370] Expression of the fusion protein or complex in a host cell The fusion protein or complex of the present disclosure containing a deaminase can be expressed in virtually any desired host cell using conventional methods known to those skilled in the art. Such host cells include, but are not limited to, bacteria, yeast, fungi, insect, plant, and animal cells. For example, DNA encoding the adenosine deaminase of the present disclosure can be cloned by designing appropriate primers upstream and downstream of the CDS based on the cDNA sequence. The cloned DNA may be ligated to DNA encoding one or more additional components of the base editing system directly, or after digestion with a restriction enzyme, if desired, or after adding an appropriate linker and / or nuclear localization signal. The base editing system is translated in the host cell to form a complex.

[0371] Polynucleotides encoding the polypeptides described herein can be obtained by chemically synthesizing the polynucleotide, or by constructing a polynucleotide (e.g., DNA) encoding the full length of the polypeptide by connecting synthesized overlapping short oligo-DNA strands using PCR and Gibson assembly. The advantage of constructing a full-length polynucleotide by a combination of chemical synthesis, PCR, and Gibson assembly is that the codons used can be selected depending on the host into which the polynucleotide is to be introduced. When expressing from a heterologous DNA molecule, converting the DNA sequence to codons that are more frequently used in the host organism is expected to increase protein expression levels. Host cell codon usage data (e.g., codon usage data available at kazusa.or.jp / codon / index.html) can be used to guide codon optimization of a polynucleotide sequence encoding a polypeptide. Codons that are less frequently used in the host can be converted to more frequently used codons that encode the same amino acid.

[0372] An expression vector containing a polynucleotide encoding a nucleic acid sequence recognition module and / or a nucleic acid base conversion enzyme can be produced, for example, by ligating the DNA downstream of a promoter in an appropriate expression vector.

[0373] The following expression vectors can be used: Escherichia coli-derived plasmids (e.g., pBR322, pBR325, pUC12, pUC13), Bacillus subtilis-derived plasmids (e.g., pUB110, pTP5, pC194), yeast-derived plasmids (e.g., pSH19, pSH15), insect cell expression plasmids (e.g., pFast-Bac), animal cell expression plasmids (e.g., pA1-11, pXT1, pRc / CMV, pRc / RSV, pcDNAI / Neo), bacteriophages such as lambda phage, insect virus vectors such as baculovirus (e.g., BmNPV, AcNPV), and animal virus vectors such as retrovirus, vaccinia virus, and adenovirus.

[0374] Any promoter suitable for the host used for gene expression can be used. Conventional methods using double-strand breaks can significantly reduce the viability of host cells due to toxicity, so it is desirable to use an induction promoter to increase the cell number before induction begins. However, since sufficient cell growth can be achieved by expressing the nucleic acid-modifying enzyme complex of the present disclosure, constitutive promoters can be used without restriction.

[0375] For example, when the host is an animal cell, the SRα promoter, SV40 promoter, LTR promoter, cytomegalovirus (CMV) promoter, Rous sarcoma virus (RSV) promoter, Moloney murine leukemia virus (MoMuLV), LTR, herpes simplex virus thymidine kinase (HSV-TK) promoter, etc. may be used. Of these, the CMV promoter, SRα promoter, etc. may also be used.

[0376] When the host is Escherichia coli, the trp promoter, lac promoter, recA promoter, lambda P.sub.L promoter, lpp promoter, T7 promoter, etc. can be used.

[0377] When the host is a Bacillus species, the SPO1 promoter, SPO2 promoter, penP promoter, etc. can be used.

[0378] When the host is yeast, the Gal1 / 10 promoter, PHO5 promoter, PGK promoter, GAP promoter, ADH promoter, etc. can be used.

[0379] When the host is an insect cell, the polyhedrin promoter, P10 promoter, etc. can be used.

[0380] When the host is a plant cell, the CaMV35S promoter, CaMV19S promoter, NOS promoter, etc. can be used.

[0381] In addition to the above, the expression vector used in the present disclosure may include an enhancer, a splicing signal, a terminator, a polyA addition signal, a selection marker such as a drug resistance gene or an auxotrophy-complementing gene, a replication origin, and the like.

[0382] RNA encoding the protein domains described herein can be prepared, for example, by in vitro transcription of a nucleic acid sequence encoding any of the fusion proteins or complexes disclosed herein.

[0383] The fusion proteins or conjugates of the present disclosure can be expressed in a cell by introducing into the cell an expression vector containing a nucleic acid sequence encoding the fusion protein or conjugate.

[0384] Host cells of interest include, but are not limited to, bacteria, yeast, fungi, insects, plants, and animal cells. For example, the host cell may be a bacterium of the genus Escherichia, such as Escherichia coli K12.cndot.DH1 (Proc. Natl. Acad. Sci. USA, 60, 160 (1968)), Escherichia coli JM103 (Nucleic Acids Research, 9, 309 (1981)), Escherichia coli JA221 (Journal of Molecular Biology, 120, 517 (1978)), Escherichia coli HB101 (Journal of Molecular Biology, 41, 459 (1969)), or Escherichia coli C600 (Genetics, 39, 440 (1954)).

[0385] The host cell may be a bacterium of the genus Bacillus, for example, Bacillus subtilis M1114 (Gene, 24, 255 (1983)) or Bacillus subtilis 207-21 (Journal of Biochemistry, 95, 87 (1984)).

[0386] The host cell may be a yeast cell. Examples of yeast cells include Saccharomyces cerevisiae AH22, AH22R.sup.-, NA87-11A, DKD-5D, and 20B-12, Schizosaccharomyces pombe NCYC1913 and NCYC2036, and Pichia pastoris KM71. When the viral delivery method uses the virus AcNPV, cells of an established line derived from armyworm larvae (Spodoptera frugiperda cells, Sf cells), MG1 cells derived from the midgut of Trichoplusia ni, High Five™ cells derived from Trichoplusia ni eggs, cells derived from Mamestra brassicae, cells derived from Estigmena acrea, etc. can be used. When the virus is BmNPV, cells of an established line derived from Bombyx mori (Bombyx mori N cells, BmN cells) can be used. Examples of Sf cells that can be used include Sf9 cells (ATCC CRL1711) and Sf21 cells (all of the above, In Vivo, 13, 213-217 (1977)).

[0387] The insect may be any insect, such as a silkworm moth, a fruit fly, or a cricket larva (Nature, 315, 592 (1985)).

[0388] Animal cells contemplated by the present disclosure include, but are not limited to, cell lines such as monkey COS-7 cells, monkey Vero cells, Chinese hamster ovary (CHO) cells, dhfr gene-deficient CHO cells, mouse L cells, mouse AtT-20 cells, mouse myeloma cells, rat GH3 cells, and human FL cells; pluripotent stem cells such as iPS cells and ES cells derived from humans and other mammals; and primary culture cells prepared from various tissues. In addition, zebrafish embryos, Xenopus oocytes, and the like can also be used.

[0389] Plant cells are also contemplated in the present disclosure, including, but not limited to, suspension culture cells, callus, protoplasts, leaf segments, root segments, and the like prepared from various plants (e.g., grains such as rice, wheat, and corn; crops such as tomato, cucumber, and eggplant; horticultural plants such as carnation and lisianthus; and other plants such as tobacco and Arabidopsis).

[0390] All of the above host cells can be monoploid (haploid) or polyploid (e.g., diploid, triploid, tetraploid, etc.). Using conventional methods, in principle, mutations are introduced into only one homologous chromosome to generate heterozygous cells. Therefore, if the mutation is not dominant, the desired phenotype will not be expressed. In the case of recessive mutations, obtaining homozygous cells can be inconvenient because it requires labor and time. In contrast, according to the present disclosure, mutations can be introduced into any allele on a homologous chromosome within the genome, so that even in the case of recessive mutations, the desired phenotype can be expressed in a single generation, which overcomes the problems of conventional mutagenesis methods.

[0391] The expression vector can be introduced by known methods (e.g., lysozyme method, competent method, PEG method, CaCl2 co-precipitation method, electroporation method, microinjection method, particle gun method, lipofection method, delivery by Agrobacterium, etc.) depending on the type of host.

[0392] Escherichia coli can be transformed according to the method described in, for example, Proc. Natl. Acad. Sci. USA, 69, 2110 (1972), Gene, 17, 107 (1982).

[0393] Vectors can be introduced into Bacillus bacteria according to the method described in, for example, Molecular & General Genetics, 168, 111 (1979).

[0394] Vectors can be introduced into yeast according to the methods described in, for example, Methods in Enzymology, 194, 182-187 (1991) and Proc. Natl. Acad. Sci. USA, 75, 1929 (1978).

[0395] Insect cells and insects can be introduced with vectors according to the method described in, for example, Bio / Technology, 6, 47-55 (1988).

[0396] Animal cells can be introduced with vectors according to the methods described in, for example, Cell Engineering additional volume 8, New Cell Engineering Experiment Protocol, 263-267 (1995) (published by Shujunsha), and Virology, 52, 456 (1973).

[0397] Cells into which a vector has been introduced can be cultured by known methods appropriate for the type of host. For example, when culturing Escherichia coli or Bacillus species, a liquid medium can be used for culture. The medium can contain carbon sources, nitrogen sources, inorganic substances, and the like necessary for the growth of the transformant. Examples of carbon sources include glucose, dextrin, soluble starch, and sucrose. Examples of nitrogen sources include inorganic or organic substances such as ammonium salts, nitrate salts, corn steep liquor, peptone, casein, meat extract, soybean cake, and potato extract. Examples of inorganic substances include calcium chloride, sodium dihydrogen phosphate, and magnesium chloride. The medium may also contain yeast extract, vitamins, growth promoting factors, and the like. In one embodiment, the pH of the medium is about 5 to about 8.

[0398] As a medium for culturing Escherichia coli, for example, M9 medium containing glucose and casamino acids (Journal of Experiments in Molecular Genetics, 431-433, Cold Spring Harbor Laboratory, New York 1972) can be used. If necessary, agents such as 3β-indolylacrylic acid may be added to the medium to ensure efficient promoter function. Escherichia coli is generally cultured at about 15 to about 43°C. Aeration and agitation may be performed as necessary.

[0399] The genus Bacillus is generally cultured at about 30° C. to about 40° C. Aeration and stirring may be carried out as necessary.

[0400] Examples of media for culturing yeast include Burkholder's minimal medium (Proc. Natl. Acad. Sci. USA, 77, 4505 (1980)) and SD medium containing 0.5% casamino acids (Proc. Natl. Acad. Sci. USA, 81, 5330 (1984)). The pH of the medium can be about 5 to about 8. Cultivation is generally carried out at about 20°C to about 35°C. Aeration and stirring may be performed as necessary.

[0401] As a medium for culturing insect cells or insects, for example, Grace's insect medium (Nature, 195, 788 (1962)) containing additives such as inactivated 10% bovine serum is used. The pH of the medium can be about 6.2 to about 6.4. Culture is generally carried out at about 27°C. Aeration and stirring may be performed as necessary.

[0402] Examples of media for culturing animal cells include minimum essential medium (MEM) containing about 5 to about 20% fetal bovine serum (Science, 122, 501 (1952)), Dulbecco's modified Eagle's medium (DMEM) (Virology, 8, 396 (1959)), RPMI 1640 medium (The Journal of the American Medical Association, 199, 519 (1967)), and 199 medium (Proceedings of the Society for the Biological Medicine, 73, 1 (1950)). The pH of the medium can be about 6 to about 8. Culture is generally carried out at about 30°C to about 40°C. Aeration and agitation may be performed as necessary.

[0403] Examples of media that can be used to culture plant cells include MS medium, LS medium, and B5 medium. The pH of the medium can be about 5 to about 8. Culture is generally carried out at about 20°C to about 30°C. Aeration and stirring may be performed as necessary.

[0404] When higher eukaryotic cells such as animal cells, insect cells, or plant cells are used as host cells, transient expression of the base editing system can be achieved by introducing a polynucleotide encoding the base editing system of the present disclosure (e.g., comprising an adenosine deaminase mutant) into the host cell under the control of an inducible promoter (e.g., a metallothionein promoter (induced by heavy metal ions), a heat shock protein promoter (induced by heat shock), a Tet-ON / Tet-OFF system promoter (induced by the addition or removal of tetracycline or its derivatives), a steroid-responsive promoter (induced by a steroid hormone or its derivatives), etc.), adding an inducer to the culture medium (or removing it from the culture medium) at an appropriate stage to induce expression of the nucleic acid-modifying enzyme complex, culturing for a certain period of time to perform base editing, and introducing a mutation into the target gene.

[0405] Prokaryotic cells such as Escherichia coli can utilize inducible promoters, including, but not limited to, the lac promoter (induced by IPTG), the cspA promoter (induced by cold shock), and the araBAD promoter (induced by arabinose).

[0406] Alternatively, the inducible promoters described above can also be used as a vector removal mechanism when higher eukaryotic cells such as animal cells, insect cells, and plant cells are used as host cells. That is, the vector contains a replication origin that functions in the host cell and nucleic acids encoding proteins necessary for replication (e.g., in the case of animal cells, SV40 large T antigen, oriP, EBNA-1, etc.), and expression of the nucleic acids encoding the proteins is regulated by the inducible promoter. As a result, the vector is capable of autonomous replication in the presence of an inducer, but is unable to self-replicate when the inducer is removed, and the vector is naturally lost with cell division (Tet-OFF vectors are not capable of autonomous replication upon the addition of tetracycline or doxycycline).

[0407] delivery system

[0408] Nucleic acid-based delivery of base editor systems Nucleic acid molecules encoding base editor systems according to the present disclosure can be administered to a subject or delivered to cells in vitro or in vivo by methods known in the art or as described herein. For example, a base editor system comprising a deaminase (e.g., cytidine deaminase or adenine deaminase) can be delivered by a vector (e.g., a viral vector or a non-viral vector), or by naked DNA, a DNA complex, a lipid nanoparticle, or a combination of the aforementioned compositions. The base editor system can be delivered to cells using any method available in the art. Such methods include, but are not limited to, physical methods (e.g., electroporation, particle gun, calcium phosphate transfection), viral methods, non-viral methods (e.g., liposomes, cationic methods, lipid nanoparticles, polymer nanoparticles), or biological non-viral methods (e.g., attenuated bacteria, engineered bacteriophages, mammalian virus-like particles, biological liposomes, erythrocyte ghosts, exosomes).

[0409] Nanoparticles, which can be organic or inorganic, are useful for delivering base editor systems or components thereof. Nanoparticles are well known in the art, and any suitable nanoparticles can be used to deliver base editor systems or components thereof, or nucleic acid molecules encoding such components. In one example, organic (e.g., lipid and / or polymer) nanoparticles are suitable for use as delivery vehicles in certain embodiments of the present disclosure. Non-limiting examples of lipid nanoparticles suitable for use in the methods of the present disclosure include those described in International Patent Application Publications WO2022140239, WO2022140252, WO2022140238, WO2022159421, WO2022159472, WO2022159475, WO2022159463, WO2021113365, and WO2021141969, the disclosures of which are each incorporated herein by reference in their entirety for all purposes.

[0410] viral vectors The base editors described herein can be delivered with a viral vector. In some embodiments, the base editors disclosed herein can be encoded in a nucleic acid contained in a viral vector. In some embodiments, one or more components of the base editor system can be encoded in one or more viral vectors.

[0411] Viral vectors can include lentiviral vectors (e.g., HIV- and FIV-based vectors), adenoviral vectors (e.g., AD100), retroviral vectors (e.g., Moloney murine leukemia virus, MML-V), herpesvirus vectors (e.g., HSV-2), and adeno-associated virus (AAV) vectors, or other plasmid or viral vector types, particularly those described in, for example, U.S. Pat. No. 8,454,972 (formulations, dosages for adenovirus), U.S. Pat. No. 8,404,658 (formulations, dosages for AAV), and U.S. Pat. No. 5,846,946 (formulations, dosages for DNA plasmids), as well as formulations and dosages from clinical trials and publications involving lentivirus, AAV, and adenovirus. For example, in the case of AAV, the route of administration, formulation, and dosage can be similar to those described in U.S. Pat. No. 8,454,972 and clinical trials involving AAV. In the case of adenovirus, the route of administration, formulation, and dosage can be similar to those described in U.S. Patent No. 8,404,658 and clinical trials involving adenovirus. In the case of plasmid delivery, the route of administration, formulation, and dosage can be similar to those described in U.S. Patent No. 5,846,946 and clinical trials involving plasmids. Dosages can be based on or estimated for an average 70 kg individual (e.g., an adult human male) and can be adjusted for patients, subjects, or mammals of different weights and species. Dosage frequency is a matter within the discretion of a medical or veterinary practitioner (e.g., a physician or veterinarian) depending on routine factors, including age, sex, general health, other conditions of the patient or subject, and the specific disease or condition being addressed. The viral vector can be injected into the tissue of interest. In the case of cell-type-specific base editors, expression of the base editor and any guide nucleic acid can be driven by a cell-type-specific promoter.

[0412] Viral vectors can be selected based on the application. For example, in the case of in vivo gene delivery, AAV may be more advantageous than other viral vectors. In some embodiments, AAV allows for low toxicity, which may be due to the purification method not requiring ultracentrifugation of cell particles that may activate immune responses. In some embodiments, AAV does not integrate into the host genome, making it less likely to cause insertional mutagen...

Claims

1. Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427), Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432) A synthetic polypeptide comprising an amino acid sequence having at least 85% sequence identity with any one of the sequences listed above, or a functional fragment thereof, or consisting of an amino acid sequence having at least 85% sequence identity with any one of the sequences listed above, or a functional fragment thereof.

2. 2. The synthetic polypeptide of claim 1, wherein the polypeptide has at least 95% sequence identity to, comprises, or consists of SEQ ID NO: 425, 426, 427, 428, 429, 430, 431, or 432.

3. 3. The synthetic polypeptide of claim 1 or claim 2, wherein the synthetic polypeptide comprises an intein.

4. A polynucleotide encoding the synthetic polypeptide of claim 1 or claim 2.

5. A cell comprising the polynucleotide of claim 4.

6. A pair of vectors, a) one member of the pair of vectors is Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427) a polynucleotide sequence encoding a synthetic polypeptide N (Syn-N) having at least about 85% amino acid sequence identity to a sequence selected from the group consisting of: b) the other member of said pair of vectors is Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432) a pair of vectors comprising a polynucleotide sequence encoding a synthetic polypeptide C (Syn-C) having at least about 85% amino acid sequence identity with a sequence selected from the group consisting of:

7. 7. The pair of vectors of claim 6, wherein the vectors are selected from the group consisting of retroviral vectors, adenoviral vectors, lentiviral vectors, herpesvirus vectors, and adeno-associated virus (AAV) vectors.

8. 8. The pair of vectors of claim 7, wherein the vectors are capable of crossing the blood-brain barrier.

9. 8. The pair of vectors of claim 7, wherein the AAV vectors are AAV9, PHP.EB, PHP.B, AAV.CAP-B10, AAV, CAP-B22, AAV-rh10, or PAL family AAV vectors.

10. The pair of vectors of claim 9, wherein the AAV vector is generated using a RepCap plasmid containing a Rep2Cap5 V2 nucleotide sequence or a Rep2Cap5 V3 nucleotide sequence.

11. The PAL family AAV vector comprises a VP1 capsid polypeptide, the VP1 capsid polypeptide comprising:

10. The pair of vectors of claim 9, comprising an amino acid sequence that has at least 95% amino acid sequence identity to the AAV9 VP1 capsid polypeptide amino acid sequence of claim 9, wherein a 7-amino acid peptide is inserted between amino acid positions Q588 and A589 of the AAV9 VP1 capsid polypeptide amino acid sequence of claim 9, wherein the 7-amino acid peptide is selected from those shown in Table 7B.

12. 12. The pair of vectors of claim 11, wherein the AAV vector comprises amino acid mutations A587D and Q588G relative to the AAV9 VP1 capsid polypeptide sequence.

13. 8. A pair of vectors according to claim 6 or claim 7, wherein the Syn-N and the Syn-C are each fused to a heterologous polypeptide.

14. The pair of vectors described in claim 13, wherein the Syn-N and the Syn-C are capable of mediating binding between heterologous polypeptides fused thereto.

15. 15. The pair of vectors of claim 14, wherein the heterologous polypeptides are each fragments of a base editor, the base editor comprising a deaminase domain and a nucleic acid-programmable DNA-binding protein domain, and wherein binding of the base editor fragment restores base editing activity.

16. 8. The pair of vectors of claim 6 or claim 7, wherein the Syn-N is an N intein and the Syn-C is a C intein, and the Syn-N and Syn-C together can function in protein splicing.

17. the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn1-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn4-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn5-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn9-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn2-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn10-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn1-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn4-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn5-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn9-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn3-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn10-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn1-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn4-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn5-C; or the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn9-C; or 8. A pair of vectors according to claim 6 or claim 7, wherein the Syn-N comprises an amino acid sequence having at least about 90% sequence identity with Syn5-N, and the Syn-C comprises an amino acid sequence having at least about 90% sequence identity with Syn10-C.

18. A cell comprising a pair of vectors according to any one of claims 6 to 17.

19. 1. A fusion protein comprising or consisting of a heterologous polypeptide fragment, said heterologous polypeptide fragment comprising: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427) or a synthetic polypeptide comprising an amino acid sequence having at least 85% sequence identity with one of them, or a functional fragment thereof, and a C-terminal end of the fusion protein.

20. 20. The fusion protein of claim 19, wherein the synthetic polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 425, 426, or 427, or an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 425, 426, or 427, or comprises or consists of SEQ ID NO: 425, 426, or 427.

21. 21. The fusion protein of claim 19 or claim 20, wherein the synthetic polypeptide is an N intein capable of functioning in protein splicing.

22. 21. The fusion protein of claim 19 or claim 20, wherein the heterologous polypeptide is a fragment of a base editor, the base editor comprising a deaminase domain and a nucleic acid-programmable DNA-binding protein domain.

23. the C-terminus of the heterologous polypeptide is the C-terminal amino acid of the napDNAbp fragment; and spCas9 1 mdkkysigld igtnsvgwav itdeykvpsk kfkvlgntdr hsikknliga llfdsgetae 61 atrlkrtarr rytrrknric ylqeifsnem akvddsffhr leesflveed kkherhpifg 121 nivdevayhe kyptiyhlrk klvdstdkad lrliylalah mikfrghfli egdlnpdnsd 181 vdklfiqlvq tynqlfeenp inasgvdaka ilsarlsksr rlenliaqlp gekknglfgn 241 lialslgltp nfksnfdlae daklqlskdt ydddldnlla qigdqyadlf laaknlsdai 301 llSdilrvnT eiTkaplsas mikrydehhq dltllkalvr qqlpekykei ffdqSkngya 361 gyidggasqe efykfikpil ekmdgteell vklnredllr kqrtfdngsi phqihlgelh 421 ailrrqedfy pflkdnreki ekiltfripy yvgplArgnS rfAwmTrkSe eTiTpwnfee 481 vvdkgasaqs fiermtnfdk nlpnekvlpk hsllyeyftv yneltkvkyv tegmrkpafl 541 sgeqkkaivd llfktnrkvt vkqlkedyfk kieCfdSvei sgvedrfnAS lgtyhdllki 601 661 rlsrklingi rdkqsgktil dflksdgfan rnfmqlihdd sltfkediqk aqvsgqgdsl 721 hehianlags paikkgilqt vkvvdelvkv mgrhkpeniv iemarenqtt qkgqknsrer 781 mkrieegike lgsqilkehp ventqlqnek lylyylqngr dmyvdqeldi nrlsdydvdh 841 ivpqsflkdd sidnkvltrs dknrgksdnv pseevvkkmk nywrqllnak litqrkfdnl 901 tkaergglse ldkagfikrq lvetrqitkh vaqildsrmn tkydendkli revkvitlks 961 klvsdfrkdf qfykvreinn yhhahdayln avvgtalikk ypklesefvy gdykvydvrk 1021 miakseqeig katakyffys nimnffktei tlangeirkr plietngetg eivwdkgrdf 1081 atvrkvlsmp qvnivkktev qtggfskesi lpkrnsdkli arkkdwdpkk yggfdsptva 1141 ysvlvvakve kgkskklksv kellgitime rssfeknpid freakgykev kkdliiklpk 1201 yslfelengr krmlasagel qkgnelalps kyvnflylas hyeklkgspe dneqkqlfve 1261 qhkhyldeii eqisefskrv iladanldkv lsaynkhrdk pireqaenii hlftltnlga 1321 paafkyfdtt idrkrytstk evldatlihq sitglyetri dlsqlggd (SEQ ID NO: 197) 23. The fusion protein of claim 22, wherein the amino acids correspond to amino acids selected from the group consisting of amino acids A292 to G364, F445 to K438, and E565 to T637 in the sequence:

24. 23. The fusion protein of claim 22, wherein the C-terminus of the heterologous polypeptide is the C-terminal amino acid of a fragment of the deaminase domain.

25. 24. The fusion protein of claim 23, wherein the deaminase domain is an adenosine deaminase, a cytidine deaminase domain, or a cytidine adenosine deaminase domain.

26. 26. The fusion protein of claim 25, wherein the adenosine deaminase domain is a TadA*8 mutant or a TadA*9 mutant.

27. A fusion protein comprising a heterologous polypeptide, said heterologous polypeptide comprising: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432) or a synthetic polypeptide comprising an amino acid sequence having at least 85% sequence identity with one of them, or a functional fragment thereof, wherein the fusion protein is fused at its N-terminus to a synthetic polypeptide comprising an amino acid sequence having at least 85% sequence identity with one of them, or a functional fragment thereof.

28. 28. The fusion protein of claim 27, wherein the synthetic polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 428, 429, 430, 431, or 432, or comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO: 428, 429, 430, 431, or 432, or comprises or consists of SEQ ID NO: 428, 429, 430, 431, or 432.

29. 29. The fusion protein of claim 27 or claim 28, wherein the synthetic polypeptide is a C intein capable of functioning in protein splicing.

30. 29. The fusion protein of Claim 27 or Claim 28, wherein the heterologous polypeptide is a fragment of a base editor, wherein the base editor comprises a deaminase domain and a nucleic acid-programmable DNA-binding protein domain.

31. the N-terminus of the heterologous polypeptide is the N-terminal amino acid of the napDNAbp fragment; and spCas9 1 mdkkysigld igtnsvgwav itdeykvpsk kfkvlgntdr hsikknliga llfdsgetae 61 atrlkrtarr rytrrknric ylqeifsnem akvddsffhr leesflveed kkherhpifg 121 nivdevayhe kyptiyhlrk klvdstdkad lrliylalah mikfrghfli egdlnpdnsd 181 vdklfiqlvq tynqlfeenp inasgvdaka ilsarlsksr rlenliaqlp gekknglfgn 241 lialslgltp nfksnfdlae daklqlskdt ydddldnlla qigdqyadlf laaknlsdai 301 llSdilrvnT eiTkaplsas mikrydehhq dltllkalvr qqlpekykei ffdqSkngya 361 gyidggasqe efykfikpil ekmdgteell vklnredllr kqrtfdngsi phqihlgelh 421 ailrrqedfy pflkdnreki ekiltfripy yvgplArgnS rfAwmTrkSe eTiTpwnfee 481 vvdkgasaqs fiermtnfdk nlpnekvlpk hsllyeyftv yneltkvkyv tegmrkpafl 541 sgeqkkaivd llfktnrkvt vkqlkedyfk kieCfdSvei sgvedrfnAS lgtyhdllki 601 661 rlsrklingi rdkqsgktil dflksdgfan rnfmqlihdd sltfkediqk aqvsgqgdsl 721 hehianlags paikkgilqt vkvvdelvkv mgrhkpeniv iemarenqtt qkgqknsrer 781 mkrieegike lgsqilkehp ventqlqnek lylyylqngr dmyvdqeldi nrlsdydvdh 841 ivpqsflkdd sidnkvltrs dknrgksdnv pseevvkkmk nywrqllnak litqrkfdnl 901 tkaergglse ldkagfikrq lvetrqitkh vaqildsrmn tkydendkli revkvitlks 961 klvsdfrkdf qfykvreinn yhhahdayln avvgtalikk ypklesefvy gdykvydvrk 1021 miakseqeig katakyffys nimnffktei tlangeirkr plietngetg eivwdkgrdf 1081 atvrkvlsmp qvnivkktev qtggfskesi lpkrnsdkli arkkdwdpkk yggfdsptva 1141 ysvlvvakve kgkskklksv kellgitime rssfeknpid freakgykev kkdliiklpk 1201 yslfelengr krmlasagel qkgnelalps kyvnflylas hyeklkgspe dneqkqlfve 1261 qhkhyldeii eqisefskrv iladanldkv lsaynkhrdk pireqaenii hlftltnlga 1321 paafkyfdtt idrkrytstk evldatlihq sitglyetri dlsqlggd (SEQ ID NO: 197) 31. The fusion protein of claim 30, wherein the amino acids correspond to amino acids selected from the group consisting of amino acids A292 to G364, F445 to K438, and E565 to T637 in the sequence:

32. 31. The fusion protein of Claim 30, wherein the N-terminus of the heterologous polypeptide is the N-terminal amino acid of a fragment of the deaminase domain.

33. 32. The fusion protein of claim 31 , wherein the deaminase domain is an adenosine deaminase, a cytidine deaminase domain, or a cytidine adenosine deaminase domain.

34. 34. The fusion protein of claim 33, wherein the adenosine deaminase domain is a TadA*8 mutant or a TadA*9 mutant.

35. 32. The fusion protein of claim 31, wherein the N-terminal amino acid of the napDNAbp fragment is Cys substituted for Ala, Ser, or Thr.

36. A polynucleotide encoding the fusion protein of any one of claims 19 to 35.

37. A vector comprising the polynucleotide of claim 36.

38. 7. The vector of claim 6, wherein the vector is selected from the group consisting of a retroviral vector, an adenoviral vector, a lentiviral vector, a herpesvirus vector, and an adeno-associated virus (AAV) vector.

39. 39. The vector of claim 38, wherein the vector is capable of crossing the blood-brain barrier.

40. 40. The vector of claim 39, wherein the AAV vector is an AAV9, PHP.EB, PHP.B, AAV.CAP-B10, AAV, CAP-B22, AAV-rh10, or PAL family AAV vector.

41. The vector of claim 40, wherein the AAV vector is generated using a RepCap plasmid comprising a Rep2Cap5 V2 nucleotide sequence or a Rep2Cap5 V3 nucleotide sequence.

42. The PAL family AAV vector comprises a VP1 capsid polypeptide, the VP1 capsid polypeptide comprising:

41. The vector of claim 40, comprising an amino acid sequence that has at least 95% amino acid sequence identity to an AAV9 VP1 capsid polypeptide amino acid sequence of the present invention, wherein a 7-amino acid peptide is inserted between amino acid positions Q588 and A589 of the AAV9 VP1 capsid polypeptide amino acid sequence of the present invention, wherein the 7-amino acid peptide is selected from those shown in Table 7B.

43. The vector of claim 42, wherein the AAV vector comprises the amino acid mutations A587D and Q588G relative to the AAV9V P1 capsid polypeptide sequence.

44. 38. The vector of Claim 37, further comprising a polynucleotide sequence encoding a gRNA capable of targeting the base editor to effect modification of the polynucleotide.

45. A cell comprising the fusion protein, polynucleotide, or vector of any one of claims 19 to 44.

46. A pharmaceutical composition comprising the fusion protein, polynucleotide, vector, or cell of any one of claims 19 to 44 and a pharmaceutically acceptable excipient.

47. (a) a first polynucleotide encoding a fusion protein comprising a heterologous polypeptide fused at its C-terminus to a first synthetic polypeptide, said first synthetic polypeptide comprising: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427) the first polynucleotide comprising an amino acid sequence having at least 85% sequence identity with one of the sequences of (b) a second polynucleotide encoding a fusion protein comprising another heterologous polypeptide fused at its N-terminus to a second synthetic polypeptide, said second synthetic polypeptide comprising: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432) and the second polynucleotide comprising an amino acid sequence having at least 85% sequence identity with one of the sequences of or a functional fragment thereof.

48. (a) a first polynucleotide encoding a fusion protein comprising an N-terminal fragment of a base editor, the base editor comprising a deaminase domain, a nucleic acid-programmable DNA-binding protein (napDNAbp) domain, and a first synthetic polypeptide fused to the C-terminus of the N-terminal fragment of the base editor, wherein the first synthetic polypeptide Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427) the first polynucleotide comprising an amino acid sequence having at least 85% sequence identity with one of the sequences of (b) a second polynucleotide encoding a fusion protein comprising a second synthetic polypeptide fused to the N-terminus of the C-terminal fragment of the base editor, wherein the second synthetic polypeptide comprises: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432) and the second polynucleotide comprising an amino acid sequence having at least 85% sequence identity with one of the sequences of or a functional fragment thereof.

49. 49. The polynucleotide delivery system of claim 47 or claim 48, wherein the first synthetic polypeptide and the second synthetic polypeptide mediate the formation of a peptide bond between the polypeptides to which they are fused.

50. 49. The polynucleotide delivery system of Claim 47 or Claim 48, further comprising a guide RNA or a polynucleotide encoding said guide RNA.

51. 51. The polynucleotide delivery system of Claim 50, wherein said guide RNA targets said base editor to effect modification of a polynucleotide in a cell.

52. 49. The polynucleotide delivery system of claim 47 or claim 48, wherein the first polynucleotide and the second polynucleotide are not covalently linked.

53. the C-terminal amino acid of the N-terminal fragment of the base editor and / or the N-terminal amino acid of the C-terminal fragment of the base editor spCas9 1 mdkkysigld igtnsvgwav itdeykvpsk kfkvlgntdr hsikknliga llfdsgetae 61 atrlkrtarr rytrrknric ylqeifsnem akvddsffhr leesflveed kkherhpifg 121 nivdevayhe kyptiyhlrk klvdstdkad lrliylalah mikfrghfli egdlnpdnsd 181 vdklfiqlvq tynqlfeenp inasgvdaka ilsarlsksr rlenliaqlp gekknglfgn 241 lialslgltp nfksnfdlae daklqlskdt ydddldnlla qigdqyadlf laaknlsdai 301 llSdilrvnT eiTkaplsas mikrydehhq dltllkalvr qqlpekykei ffdqSkngya 361 gyidggasqe efykfikpil ekmdgteell vklnredllr kqrtfdngsi phqihlgelh 421 ailrrqedfy pflkdnreki ekiltfripy yvgplArgnS rfAwmTrkSe eTiTpwnfee 481 vvdkgasaqs fiermtnfdk nlpnekvlpk hsllyeyftv yneltkvkyv tegmrkpafl 541 sgeqkkaivd llfktnrkvt vkqlkedyfk kieCfdSvei sgvedrfnAS lgtyhdllki 601 ikdkdfldne enedilediv ltltlfedre mieerlktya hlfddkvmkq lkrrrytgwg 661 rlsrklingi rdkqsgktil dflksdgfan rnfmqlihdd sltfkediqk aqvsgqgdsl 721 hehianlags paikkgilqt vkvvdelvkv mgrhkpeniv iemarenqtt qkgqknsrer 781 mkrieegike lgsqilkehp ventqlqnek lylyylqngr dmyvdqeldi nrlsdydvdh 841 ivpqsflkdd sidnkvltrs dknrgksdnv pseevvkkmk nywrqllnak litqrkfdnl 901 tkaergglse ldkagfikrq lvetrqitkh vaqildsrmn tkydendkli revkvitlks 961 klvsdfrkdf qfykvreinn yhhahdayln avvgtalikk ypklesefvy gdykvydvrk 1021 miakseqeig katakyffys nimnffktei tlangeirkr plietngetg eivwdkgrdf 1081 atvrkvlsmp qvnivkktev qtggfskesi lpkrnsdkli arkkdwdpkk yggfdsptva 1141 ysvlvvakve kgkskklksv kellgitime rssfeknpid freakgykev kkdliiklpk 1201 yslfelengr krmlasagel qkgnelalps kyvnflylas hyeklkgspe dneqkqlfve 1261 qhkhyldeii eqisefskrv iladanldkv lsaynkhrdk pireqaenii hlftltnlga 1321 paafkyfdtt idrkrytstk evldatlihq sitglyetri dlsqlggd (SEQ ID NO: 197) 49. The polynucleotide delivery system of claim 48, wherein the amino acids in the napDNAbp domain correspond to amino acids A292 to G364, F445 to K438, and E565 to T637 in the sequence:

54. 49. The polynucleotide delivery system of Claim 48, wherein the C-terminal amino acid of the N-terminal fragment of the base editor and / or the N-terminal amino acid of the C-terminal fragment of the base editor corresponds to an amino acid within the deaminase domain.

55. 49. The polynucleotide delivery system of Claim 48, wherein the base editor comprising a fusion of an N-terminal fragment of the base editor and a C-terminal fragment of the base editor has base editing activity.

56. 49. The polynucleotide delivery system of Claim 48, wherein a fusion of an N-terminal fragment and a C-terminal fragment of the base editor generates a reconstituted full-length base editor.

57. 49. The polynucleotide delivery system of claim 48, wherein the deaminase domain is an adenosine deaminase, a cytidine deaminase domain, or a cytidine adenosine deaminase domain.

58. 58. The polynucleotide delivery system of claim 57, wherein the adenosine deaminase domain is a TadA*8 mutant or a TadA*9 mutant.

59. 59. The polynucleotide delivery system of Claim 58, wherein the adenosine deaminase domain is TadA*8.

5.

60. 49. The polynucleotide delivery system of any one of Claims 48, wherein the N-terminal amino acid of the C-terminal fragment of the base editor is Cys substituted for Ala, Ser, or Thr.

61. the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn9-C; or 49. The polynucleotide delivery system of claim 47 or claim 48, wherein the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn10-C.

62. A pharmaceutical composition comprising the polynucleotide delivery system of any one of claims 47 to 61 and a pharmaceutical excipient.

63. 1. A method for delivering a polynucleotide encoding a heterologous polypeptide to a cell, comprising: (a) a first polynucleotide encoding a fusion protein comprising a heterologous polypeptide or fragment thereof fused at its C-terminus to a first synthetic polypeptide, wherein the first synthetic polypeptide comprises: Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427) the first polynucleotide comprising an amino acid sequence having at least 85% sequence identity with one of the sequences of (b) a second polynucleotide encoding a fusion protein comprising another heterologous polypeptide fused at its N-terminus to a second synthetic polypeptide, said second synthetic polypeptide comprising: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432) and a second polynucleotide comprising an amino acid sequence having at least 85% sequence identity to one of the sequences of (a) to (c), or a functional fragment thereof.

64. 1. A method for delivering a polynucleotide encoding a base editor fragment to a cell, comprising: (a) a first polynucleotide encoding a fusion protein comprising an N-terminal fragment of a base editor, the base editor comprising a deaminase domain, a nucleic acid-programmable DNA-binding protein (napDNAbp) domain, and a first synthetic polypeptide fused to the C-terminus of the N-terminal fragment of the base editor, wherein the first synthetic polypeptide Syn2-N CLSYDTEILTVEYGLIPIGEIVEKKIECTVYTIDNNGLIYTQSIEQWHHRGYQELFEYILEDGSTIRATKDHKFMTSERQMLPIEEIFERGWELKQVL (SEQ ID NO: 425), Syn3-N CLSSDTEVITEEYGPIAIGKIVDEGIRCSVYSVDNNGNLYTQPISQWHDRGRQEIYEYYLENGSVIRATKDHKFMTKDGEMLPIDEIFEKGLELKQVLP (SEQ ID NO: 426), and Syn5-N CLSYETEVLTVEYGFMPIGKIVEERIRCSVYTVDKNGFIYSQPIAQWHQRGLQEVYEYDLENGSIIRATKEHQFMTNDGQMLAIHEIFTRKLDLLQSQE (SEQ ID NO: 427) the first polynucleotide comprising an amino acid sequence having at least 85% sequence identity with one of the sequences of (b) a second polynucleotide encoding a fusion protein comprising a second synthetic polypeptide fused to the N-terminus of the C-terminal fragment of the base editor, wherein the second synthetic polypeptide comprises: Syn1-C MKVISRKSLGTQPVYDICVTHDHNFLMKNGLIASN (SEQ ID NO: 428), Syn4-C MDVKIVSYKFLGSENVYDILERDHNFLIKNGLVASN (SEQ ID NO: 429), Syn5-C MVKIITYKSLGRQKVYDLGLEQDHNFVLANGLVASN (SEQ ID NO: 430), Syn9-C MVKIISRKYLDTQPVYDVGVQKDHNFLISNGSIASN (SEQ ID NO: 431), and Syn10-C MVKIATRRSLGTEPVYDIGLQQEHNFLLANGLVASN (SEQ ID NO: 432) and a second polynucleotide comprising an amino acid sequence having at least 85% sequence identity to one of the sequences of (a) to (c), or a functional fragment thereof.

65. 65. The method of claim 63 or claim 64, wherein the synthetic polypeptides are capable of mediating binding between the polypeptides to which they are fused.

66. 66. The method of claim 65, wherein the bond is a covalent bond.

67. 66. The method of Claim 65, wherein said binding restores base editing activity.

68. 65. The method of claim 63 or claim 64, wherein the first synthetic polypeptide and the second synthetic polypeptide mediate the formation of a peptide bond between the polypeptides to which they are fused.

69. 65. The method of Claim 64, further comprising contacting the cell with a guide RNA or a polynucleotide encoding said guide RNA.

70. 70. The polynucleotide delivery system of Claim 69, wherein said guide RNA targets said base editor to effect modification of a polynucleotide in a cell.

71. 65. The method of claim 63 or claim 64, further comprising contacting the cell with a vector comprising the first polynucleotide and / or a vector comprising the second polynucleotide.

72. 72. The method of claim 71, wherein the vector is an adeno-associated virus (AAV) vector.

73. 73. The method of claim 72, wherein the vector is selected from the group consisting of a retroviral vector, an adenoviral vector, a lentiviral vector, a herpesvirus vector, and an adeno-associated virus (AAV) vector.

74. 74. The vector of claim 73, wherein the vector is capable of crossing the blood-brain barrier.

75. 74. The method of claim 73, wherein the AAV vector is an AAV9, PHP.EB, PHP.B, AAV.CAP-B10, AAV, CAP-B22, AAV-rh10, or PAL family AAV vector.

76. 76. The method of claim 75, wherein the AAV vector is generated using a RepCap plasmid comprising a Rep2Cap5 V2 nucleotide sequence or a Rep2Cap5 V3 nucleotide sequence.

77. The PAL family AAV vector comprises a VP1 capsid polypeptide, the VP1 capsid polypeptide comprising:

76. The method of claim 75, wherein the AAV9 VP1 capsid polypeptide comprises an amino acid sequence that has at least 95% amino acid sequence identity to the AAV9 VP1 capsid polypeptide amino acid sequence of

78. 78. The method of claim 77, wherein the AAV vector comprises the amino acid mutations A587D and Q588G relative to the AAV9V P1 capsid polypeptide sequence.

79. 65. The method of claim 63 or claim 64, wherein the first polynucleotide and the second polynucleotide are not covalently linked.

80. the C-terminal amino acid of the N-terminal fragment of the base editor and / or the N-terminal amino acid of the C-terminal fragment of the base editor spCas9 1 mdkkysigld igtnsvgwav itdeykvpsk kfkvlgntdr hsikknliga llfdsgetae 61 atrlkrtarr rytrrknric ylqeifsnem akvddsffhr leesflveed kkherhpifg 121 levels 181 241 301 llSdilrvnT eiTkaplsas mikridehhq dltllkalvr qqlpekykei ffdqSkngya 361 421 ailrrqedfy pflkdnreki ekiltfripy yvgplArgnS rfAwmTrkSe eTiTpwnfee 481 541 sgeqkkaivd llfktnrkvt vkqlkedyfk kieCfdSvei sgvedrfnAS lgtyhdllki 601 ikdkdfldne enedilediv ltltlfedre mieerlktya hlfddkvmkq lkrrrytgwg 661 rlsrklingi rdkqsgktil dflksdgfan rnfmqlihdd sltfkediqk aqvsgqgdsl 721 781 841 ivpqsflkdd sidnkvltrs dknrgksdnv pseevvkkmk nywrqllnak litqrkfdnl 901 tkaergglse ldkagfikrq lvetrqitkh vaqildsrmn tkydendkli revkvitlks 961 klvsdfrkdf qfykvreinn yhhahdayln avvgtalikk ypklesefvy gdykvydvrk 1021 miakseqeig katakyffys nimnffktei tlangeirkr plietngetg eivwdkgrdf 1081 atvrkvlsmp qvnivkktev qtggfskesi lpkrnsdkli arkkdwdpkk yggfdsptva 1141 ysvlvvakve kgkskklksv kellgitime rssfeknpid freakgykev kkdliiklpk 1201 yslfelengr krmlasagel qkgnelalps kyvnflylas hyeklkgspe dneqkqlfve 1261 qhkhyldeii eqisefskrv iladanldkv lsaynkhrdk pireqaenii hlftltnlga 1321 paafkyfdtt idrkrytstk evldatlihq sitglyetri dlsqlggd (SEQ ID NO: 197) 65. The method of claim 64, wherein the amino acids in the napDNAbp domain correspond to amino acids selected from the group consisting of amino acids A292 to G364, F445 to K438, and E565 to T637 in the sequence

81. 65. The method of Claim 64, wherein the C-terminal amino acid of the N-terminal fragment of the base editor and / or the N-terminal amino acid of the C-terminal fragment of the base editor corresponds to an amino acid within the deaminase domain.

82. 65. The method of Claim 64, wherein the base editor comprising a fusion of an N-terminal fragment of the base editor and a C-terminal fragment of the base editor has base editing activity.

83. 65. The method of Claim 64, wherein a reconstituted full-length base editor is generated by fusing an N-terminal fragment and a C-terminal fragment of the base editor.

84. 65. The method of claim 64, wherein the deaminase domain is an adenosine deaminase, a cytidine deaminase domain, or a cytidine adenosine deaminase domain.

85. 65. The method of claim 64, wherein the adenosine deaminase domain is a TadA*8 mutant or a TadA*9 mutant.

86. 66. The method of any one of Claims 51-65, wherein the N-terminal amino acid of the C-terminal fragment of the base editor is Cys substituted for Ala, Ser, or Thr.

87. the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn2-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn9-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn3-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn10-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn1-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn4-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-C; or the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity to Syn9-C; or 65. The method of claim 63 or claim 64, wherein the first synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn5-N, and the second synthetic polypeptide comprises an amino acid sequence having at least about 85% sequence identity with Syn10-C.

88. 88. A method for editing a target polynucleotide in a cell, comprising delivering to the cell a polynucleotide encoding a base-edited fragment according to the method of any one of claims 64-87.

89. 89. The method of Claim 88, wherein the method achieves a base editing efficiency of at least about 10%, 30%, or 50%.

90. 90. The method of Claim 88 or Claim 89, wherein the method achieves a base editing efficiency that is equal to or greater than a base editing efficiency achieved when the first polynucleotide and the second polynucleotide are replaced with a single polynucleotide encoding a full-length base editor comprising the deaminase domain and the napDNAbp.

91. 91. The method of any one of Claims 88-90, wherein the method achieves a base editing efficiency that is equal to or greater than the base editing efficiency achieved when the first polynucleotide and the second polynucleotide encode neither an N intein nor a C intein.

92. 10. A kit suitable for use in a method according to any one of the preceding claims, said kit comprising a polynucleotide, polypeptide, vector or composition according to any one of the preceding claims.