Programmable DNA pyrimidine base editing via engineered uracil-DNA glycosylase-based excision

Engineered UNG variants with targeted amino acid substitutions enable efficient and precise editing of thymine and cytosine bases, addressing the limitations of current DNA editing tools and enhancing therapeutic applications.

WO2025232923A1PCT designated stage Publication Date: 2025-11-13PEKING UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/094182
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-10
Filing Date
2025-05-12
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Current DNA editing tools lack the capability to directly edit G and T bases, limiting the diversity of available nucleic acid editing tools and their therapeutic applications.

Method used

Engineered thymine-modifying polypeptides comprising uracil-DNA glycosylase (UNG) variants with specific amino acid substitutions, such as Y85A and additional substitutions like L80V, K113E, K197E, S204A, and E237Q, are used to excise thymine or cytosine bases, combined with a DNA recognition domain for targeted base editing.

Benefits of technology

The engineered UNG variants achieve efficient and precise base editing, with editing efficiencies up to sixfold higher than wild-type UNG, reducing off-target effects and enabling correction of up to 70% of human disease-related point mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2025094182-FTAPPB-I100001
    Figure PCTCN2025094182-FTAPPB-I100001
  • Figure PCTCN2025094182-FTAPPB-I100002
    Figure PCTCN2025094182-FTAPPB-I100002
  • Figure PCTCN2025094182-FTAPPB-I100003
    Figure PCTCN2025094182-FTAPPB-I100003
Patent Text Reader

Abstract

Provided are compositions, methods, and systems for DNA pyrimidine-base editing. In some embodiments, provided are engineered thymine-modifying polypeptides or engineered cytosine-modifying polypeptides comprising a variant of a uracil-DNA glycosylate (UNG). In some embodiments, provided are such engineered polypeptides and a DNA recognition domain, which are configured to target a nucleic acid for excision. Also provided herein are nucleic acid encoding the polypeptides described herein, additional components useful for editing, such as sgRNA, kits, medicines, composition, and method of use thereof.
Need to check novelty before this filing date? Find Prior Art

Description

PROGRAMMABLE DNA PYRIMIDINE BASE EDITING VIA ENGINEERED URACIL-DNA GLYCOSYLASE-BASED EXCISIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority benefit of International Application No. PCT / CN2024 / 092138, filed on May 10, 2024, the content of which is hereby incorporated herein by reference in its entirety. REFERENCE TO AN ELECTRONIC SEQUENCE LISTING

[0002] The content of the electronic sequence listing (165392002041seqlist. xml; Size: 172, 464 bytes; and Date of Creation: May 8, 2025) is herein incorporated by reference in its entirety.TECHNICAL FIELD

[0003] The present application is directed to compositions, methods, and systems for DNA pyrimidine-base editing. In some embodiments, the present application is directed to engineered thymine-modifying polypeptides or engineered cytosine-modifying polypeptides comprising a variant of a uracil-DNA glycosylate (UNG) . In some embodiments, the present application is directed to such engineered polypeptides and a DNA recognition domain, which are configured to target a nucleic acid for excision. Also provided herein are nucleic acid encoding the polypeptides described herein, additional components useful for editing, such as sgRNA, kits, medicines, composition, and method of use thereof.BACKGROUND

[0004] Much effort has been spent to develop tools for editing DNA nucleic acids, which holds significant promise for advancing the treatment of many diseases including those caused by a DNA mutation or that can be treated by a DNA mutation. Currently, there are developed cytosine base editors and adenine base editors comprising deaminases. Generally, the deaminase in these tools transforms cytosine (C) or adenine (A) into uracil (U) or inosine (I) , which are then subsequently recognized as thymine (T) and guanine (G) in DNA repair mechanisms to complete C-to-T and A-to-G base editing. Building on this, the generation of apurinic / apyrimidinic sites (AP sites) occurs by cutting the intermediate products uracil or inosine with DNA glycosylase. Subsequent translesion DNA synthesis (TLS) incorporates other bases opposite the AP site, leading to the development of C-to-G base editors (CGBE) and A-to-Y base editors (AYBE) . These base editing tools all rely on deaminase enzymes, resulting in a current absence of tools capable of directly editing G and T and those utilizing other enzymes. Thus, there remains a need to diversify available nucleic acid editing tools including tools capable of programmable C or T base editing. BRIEF SUMMARY

[0005] In some aspects, provided herein is an engineered thymine-modifying polypeptide comprising a uracil-DNA glycosylase (UNG) variant comprising an amino acid substitution selected from a group consisting of Y85A, Y85S, Y85N, Y85G, and Y85C; and one or more amino acid substitutions of L80V, K113E, K197E, S204A, S204G, or E237Q, wherein the positions of amino acid substitutions are in reference to the amino acid sequence set forth in SEQ ID NO: 7.

[0006] In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitution of Y85A. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of L80V, K113E, K197E, S204A, and E237Q. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of K113E and S204G. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of K113E, S204A, and E237Q. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitution of L80V. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of K197E, S204A, and E237Q. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of K113E, K197E, and E237Q.

[0007] In some embodiments, the UNG is derived from an organism selected from a group consisting of Deinococcus radiodurans, Homo sapiens, Escherichia coli, Human Herpesvirus 1, and Vaccinia virus. In some embodiments, the UNG is derived from Deinococcus radiodurans.

[0008] In some aspects, provided herein is an engineered thymine-modifying polypeptide comprising a uracil-DNA glycosylase (UNG) variant comprising an amino acid substitution comprising Y85A, wherein the UNG variant is derived from Deinococcus radiodurans. In some embodiments, the engineered thymine-modifying polypeptide further comprises one or more amino acid substitutions of L80V, K113E, K197E, S204A, S204G, or E237Q, wherein the positions of amino acid substitutions are in reference to the amino acid sequence set forth in SEQ ID NO: 7.

[0009] In some embodiments, the UNG variant has at least about 70%sequence identity to SEQ ID NO: 7. In some embodiments, the UNG variant comprises an N-terminal deletion relative to the corresponding UNG from which the UNG variant was derived. In some embodiments, the N-terminal deletion comprises a deletion of up to the first 23 amino acids. In some embodiments, the N-terminal deletion comprises a deletion of the first two amino acids.

[0010] In some embodiments, the engineered thymine-modifying polypeptide further comprises a DNA recognition domain. In some embodiments, the DNA recognition domain is fused to the UNG variant. In some embodiments, the DNA recognition domain is fused to the C-terminus of the UNG variant. In some embodiments, the DNA recognition domain is operably connected to the UNG variant by a linker domain. In some embodiments, the linker domain is a 32-amino acid linker or a 64-amino acid linker, and is selected from a list composed of GS, SGGS, PAPAP, XTEN, or repeats thereof.

[0011] In some embodiments, the DNA recognition domain comprises a Cas nuclease. In some embodiments, the Cas nuclease is an nCas9. In some embodiments, the nCas9 comprises a D10A mutation. In some embodiments, the nCas9 is derived from Streptococcus pyogenes. In some embodiments, the Cas nuclease is a dCas9.

[0012] In some embodiments, the engineered thymine-modifying polypeptide does not comprise a deaminase domain having deamination activity.

[0013] In some aspects, provided herein is an engineered thymine-modifying polypeptide comprising SEQ ID NO: 9. In some aspects, provided herein is an engineered thymine-modifying polypeptide comprising SEQ ID NO: 10. In some aspects, provided herein is an engineered thymine-modifying polypeptide comprising SEQ ID NO: 12. In some aspects, provided herein is an engineered thymine-modifying polypeptide comprising SEQ ID NO: 13. In some aspects, provided herein is an engineered thymine-modifying polypeptide comprising SEQ ID NO: 14. In some aspects, provided herein is an engineered thymine-modifying polypeptide comprising SEQ ID NO: 15. In some embodiments, the engineered thymine-modifying polypeptide further comprises a DNA recognition domain.

[0014] In some aspects, provided herein is a nucleic acid encoding an engineered thymine-modifying polypeptide described herein.

[0015] In some aspects, provided herein is a thymine editing system comprising an engineered thymine-modifying polypeptide described herein, or a nucleic acid encoding the same. In some embodiments, the thymine editing system further comprises an sgRNA. In some embodiments, the sgRNA targets a protospacer comprising a target thymine.

[0016] In some aspects, provided herein is a method of modifying a target thymine in a nucleic acid sequence present in one or more nucleic acid molecules, the method comprising contacting at least one nucleic acid molecule of the one or more nucleic acid molecules with an engineered thymine-modifying polypeptide described herein, wherein the DNA recognition domain is configured to associate with the at least one nucleic acid molecule such that the UNG variant is positioned to modify the target thymine.

[0017] In some embodiments, the one or more nucleic acid molecules are in a cell. In some embodiments, the one or more nucleic acid molecules are in a mitochondrion. In some embodiments, the target thymine undergoes a T-to-C, T-to-G, or T-to-Amodification. In some embodiments, about 15%to about 75%of the modified target thymine are modified to a cytosine. In some embodiments, about 1%to about 65%of the modified target thymine are modified to a guanine. In some embodiments, about 5%to about 50%of the modified target thymine are modified to an alanine.

[0018] In some embodiments, an NGG PAM site is 14 to 19 base pairs away from the target thymine. In some embodiments, an NGG PAM site is 14 to 19 base pairs away from the target thymine, wherein counting starts at the first base outside of the NGG PAM in a protospacer, and wherein the NGG PAM site is associated with DNA recognized by the DNA recognition domain, such as a Cas nuclease. In some embodiments, an NG PAM site is 14 to 19 base pairs away from the target thymine. In some embodiments, an NG PAM site is 14 to 19 base pairs away from the target thymine, wherein counting starts at the first base outside of the NG PAM in a protospacer, and wherein the NG PAM site is associated with DNA recognized by the DNA recognition domain, such as a Cas nuclease

[0019] In some embodiments, the method exhibits an editing efficiency at least 2-fold greater than that of a method comprising contacting a nucleic acid molecule comprising a target thymine with a wild-type UNG. In some embodiments, the method exhibits an editing efficiency of about 6%to about 40%.

[0020] In some aspects, provided herein is a method of treating a disease associated with a DNA mutation in an individual, the method comprising contacting a nucleic acid comprising a target thymine associated with a disease with a thymine editing system described herein. In some embodiments, the thymine editing system targets a splicing site on a gene. In some embodiments, the disease is Duchene Muscular Dystrophy.

[0021] In some embodiments, the contacting reduces the amount of a premature stop codon. In some embodiments, the contacting results in the formation of a TAG-to-XAG mutation in exon 9 of a IDUA gene. In some embodiments, the administration of the thymine editing system results in restoration of IDUA catalytic activity. In some embodiments, the disease is Hurler syndrome.

[0022] In some aspects, provided herein is a method of treating Hurler syndrome in an individual in need thereof, comprising contacting a nucleic acid comprising a target thymine in a mutation associated with Hurler syndrome with a thymine editing system described herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] FIG. 1A shows a schematic of screen hUNG variants that can specifically excise thymine and to produce AP sites using reporter system (SEQ ID NO: 22 and SEQ ID NO: 23) . FIG. 1B shows a schematic diagram of nCas9 (D10A) -hUNG variant targeting reporter system (Top sequence (SEQ ID NO: 123) ; middle sequence (SEQ ID NO: 124) ; bottom sequence (SEQ ID NO: 125) ) . FIG. 1C shows the eGFP+ ratio of hUNG key amino acid saturation mutations for thymine excision (SEQ ID NO: 126) . FIG. 1D shows the editing rate of the hUNG (Y147A) -nCas9 (D10A) targeting reporter system.

[0024] FIG. 2A shows a schematic of screen hUNG variant that can specifically excise cytosine to produce AP sites using reporter system (SEQ ID NO: 25 and SEQ ID NO: 26) . FIG. 2B shows the eGFP+ ratio of hUNG key amino acid saturation mutations for cytosine excision (SEQ ID NO: 126) . FIG. 2C shows the editing rate of the hUNG (N204D) -nCas9 (D10A) targeting reporter system.

[0025] FIG. 3A shows the evaluation and results of cytosine excision by native and engineered UNG variants from various species (From top to bottom: SEQ ID NO: 127, 128, 129, 30, 131, 132) . FIG. 3B shows the evaluation and results of thymine excision by native and engineered UNG variants from various species (From top to bottom: SEQ ID NO: 133, 134, 135, 136, 137, 138) .

[0026] FIG. 4A shows a flowchart for screening DrUNG variants. FIG. 4B shows the enrichment percentage of mutations at particular amino acids obtained from the screen. FIG. 4C shows the eGFP+ ratio of different DrUNG variants.

[0027] FIG. 5A shows separate plots the thymine base editing rate of DrUNG variants 2 and 6 and indel percentage at 16 endogenous sites. FIG. 5B shows the comparison of thymine base editing efficiency of DrUNG variant 2 and 6 at 16 endogenous sites. FIG. 5C shows the distribution of thymine editing results at 16 endogenous sites of DrUNG mutant 6 (also referred to as “thymine base editor” [TBE] ) . FIG. 5D shows a plot of the editing rate for various amino acid positions away from a NGG PAM. FIG. 5E shows the editing efficiency of DrUNG mutant 6 comprising various linkers. FIG. 5F shows the editing rate of DrUNG mutant 6 comprising various linkers at various endogenous sites. FIG. 5G shows the editing efficiency of DrUNG mutant 6 in various cell lines. FIG. 5H shows indel levels of DrUNG mutant 6 in various cell lines.

[0028] FIG. 6A shows the genome wide off-target effects of TBE at site 31. Sample transfected with eGFP-expressing plasmid as control. FIG. 6B shows the transcriptome wide off-target effects of TBE at site 31. Sample transfected with eGFP-expressing plasmid as control. FIG. 6C shows the editing efficiency of top 10 off-target sites predicted by Cas-OFFinder at site 15 (from top to bottom: SEQ ID NO: 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149) , 16 (from top to bottom: SEQ ID NO: 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160) , 17 (from top to bottom: SEQ ID NO: 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171) and 18 (from top to bottom: SEQ ID NO: 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182) .

[0029] FIG. 7A shows a diagram illustrating the human Hurler syndrome reporter system and sgRNA design for TBE. (From top to bottom: SEQ ID NO: 183, 184, 185, 186, 187, 188) . FIG. 7B shows Editing efficiency of SpCas9, SpCas9-NG and SpG-Cas9 respectively fused to DrUNG mutant 6 and co-transfected with the corresponding sgRNA in IDUA reporter system. FIG. 7C shows a schematic diagram of co-transfection of mRNA of the fusion protein of SpCas9 and DrUNG mutant 6 and sgRNA into GM06214 (IDUAW402X) cells derived from patient with Hurler syndrome, which contain the IDUAW402X mutation. FIG. 7D shows the editing efficiency of on-target in GM06214 cells using TBE. FIG. 7E shows the relative IDUA catalytic activity of GM06214 cells after transfection with TBE. Data are presented as mean ±s.d. of n = 3 independent biological replicates.DETAILED DESCRIPTION

[0030] Provided herein, in some aspects, are engineered polypeptides for DNA pyrimidine base modification. In certain aspects, the DNA pyrimidine base modifying polypeptides taught herein are derived from a uracil-DNA glycosylase (UNG) , which natively excises uracil in DNA to prevent mutagenesis, and engineered to excise a DNA pyrimidine base such as a thymine or a cytosine. In some embodiments, the engineered DNA pyrimidine base modifying polypeptide is configured to excise a target thymine. In other embodiments, the engineered DNA pyrimidine base modifying polypeptide is configured to excise a target cytosine. In some embodiments, the engineered DNA pyrimidine base modifying polypeptide comprises a DNA recognition domain, such as to provide targeting specificity. Further described herein are methods of using the taught engineered pyrimidine base modifying polypeptides (e.g., in a method of treatment) , kits, and systems thereof.

[0031] The disclosure provided herein is based, at least in part, on the inventors’ unique perspectives and findings associated with UNG variants that transform the native function of a UNG and allow for highly efficient excision of DNA pyrimidine bases, such as thymine or cytosine bases (as noted above, UNGs natively excise uracil bases in DNA) . As demonstrated herein, the inventors engineered novel UNGs from species not previously explored, such as engineered Deinococcus radiodurans UNG and surprisingly found mutant variants with increase DNA pyrimidine excision activity. Specifically, it was found that EcUNG (N123D) , HHV1_UNG (N147D) , and DrUNG (Y85A) each demonstrated increasingly efficient excision of cytosine and thymine, respectively. The activity of DrUNG (Y85A) surpasses that of hUNG (Y147A) by a factor of five, and through further mutations, the inventors found that the activity of such engineered polypeptides could further be enhanced by about sixfold. The engineered UNG variants provided herein represent a novel approach to editing DNA pyrimidine bases that do not involve a deaminase. Moreover, the inventors developed an effective thymine base editing tool by fusing a UNG variant to a DNA recognition domain, with the editing window situated at positions 14-18 of the spacer between the two. More specifically, the inventors also found that by using an engineered UNG variant provided herein, one can selectively generate an apurinic / apyrimidinic site (AP site) , induce nicking on the complementary strand with nCas9, and then achieve base editing of cytosine or thymine during translesion DNA synthesis via incorporation of alternative bases opposite the AP site. As demonstrated in the Examples, we found a preference for incorporating G opposite the AP site after thymine removal, resulting in T-to-C base editing. A smaller fraction of incorporations involves C or T, allowing for T-to-G and T-to-Abase conversions. In the context of human disease-related point mutations, the base editing tool provided herein holds the potential to correct up to 70%of these mutations. Notably, existing base editing tools relying on deaminases often exhibit significant off-target effects on RNA, posing a clear drawback for disease treatment. The base editing tool, circumventing the need for deaminases, provides a more precise approach to base editing for therapeutic applications.

[0032] Thus, in some aspects, provided herein is an engineered uracil-DNA glycosylase (UNG) variant configured to excise a DNA pyrimidine such as a target thymine or a target cytosine. In some embodiments, the engineered polypeptide is a thymine-modifying polypeptide comprising a variant of a UNG. In some embodiments, the engineered polypeptide is a cytosine-modifying polypeptide comprising a variant of a UNG.

[0033] In some embodiments, provided herein is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) variant comprising an amino acid substitution of Y85 (such as to an amino acid that is smaller than tyrosine) ; and one or more amino acid substitutions at L80, K113, K197, S204, S204, or E237, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7.

[0034] In some embodiments, provided herein is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution selected from a group consisting of Y85A, Y85S, Y85N, Y85G, and Y85C; and (b) one or more amino acid substitutions of L80V, K113E, K197E, S204A, S204G, or E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitution of Y85A. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of L80V, K113E, K197E, S204A, and E237Q. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of K113E and S204G. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of K113E, S204A, and E237Q. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitution of L80V. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of K197E, S204A, and E237Q. In some embodiments, the engineered thymine-modifying polypeptide comprises the amino acid substitutions of K113E, K197E, and E237Q. In some embodiments, the engineered thymine-modifying polypeptide is derived from Deinococcus radiodurans. In some embodiments, the engineered thymine-modifying polypeptide comprises an N-terminal deletion (such as a deletion of one or two amino acids) .

[0035] In some embodiments, the engineered thymine-modifying polypeptide described herein comprises a DNA recognition domain, such as a Cas nuclease (e.g., nCas9 or dCas9) or TAL effector.

[0036] In other aspects, provided herein is a composition comprising a nucleic acid encoding an engineered polypeptide for DNA pyrimidine base modification described herein such as an engineered thymine-modifying polypeptide described herein.

[0037] In other aspects, provided herein is a method of excising a DNA pyrimidine, such as thymine, from a nucleic acid molecule, the method comprising delivering an engineered polypeptide for DNA pyrimidine base modification described herein, such as an engineered thymine-modifying polypeptide, or a nucleic acid composition encoding such a polypeptide.

[0038] In other aspects, provided is a nucleic acid editing system comprising an engineered polypeptide for DNA pyrimidine base modification described herein, such as an engineered thymine-modifying polypeptide comprising a DNA recognition domain, and an sgRNA.

[0039] In other aspects, provided herein is a method of treating a disease associated with a DNA mutation in an individual, the method comprising administering an engineered polypeptide for DNA pyrimidine base modification described herein, such as an engineered thymine-modifying polypeptide, or a nucleic acid composition encoding such a polypeptide, such that the engineered thymine-modifying polypeptide, or the nucleic acid composition, contacts a target nucleic acid and excises a target DNA pyrimidine base.

[0040] In other aspects, provided herein is a kit for modifying nucleic acids, the kit comprising an engineered polypeptide for DNA pyrimidine base modification described herein, such as an engineered thymine-modifying polypeptide, or a nucleic acid composition encoding such a polypeptide. I. Definitions

[0041] For purposes of interpreting this specification, the following definitions will apply and whenever appropriate, terms used in the singular will also include the plural and vice versa. In the event that any definition set forth below conflicts with any document incorporated herein by reference, the definition set forth shall control.

[0042] The terms “polypeptide” and “protein, ” as used herein, may be used interchangeably to refer to a polymer comprising amino acid residues, and are not limited to a minimum length. Such polymers may be translational fusions of two or more proteins. Such polymers may contain natural or non-natural amino acid residues, or combinations thereof, and include, but are not limited to, peptides, polypeptides, oligopeptides, dimers, trimers, and multimers of amino acid residues. Full-length polypeptides or proteins, and fragments thereof, are encompassed by this definition. The terms also include modified species thereof, e.g., post-translational modifications of one or more residues, for example, methylation, phosphorylation glycosylation, sialylation, or acetylation.

[0043] The term “polynucleotide, ” as used herein, refers to a polymeric form of nucleotides of any length, and may be either ribonucleotides (RNA) or deoxyribonucleotides (DNA) . Thus, this term includes, but is not limited to unless specifically stated to be so limited, single-, double-or multi-stranded DNA or RNA, genomic DNA, mitochondrial DNA (mtDNA) , cDNA, DNA-RNA hybrids, or a polymer comprising purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The backbone of the polynucleotide can comprise sugars and phosphate groups (as may typically be found in RNA or DNA) , or modified or substituted sugar or phosphate groups. Alternatively, the backbone of the polynucleotide can comprise a polymer of synthetic subunits such as phosphoramidates and phosphorothioates, and thus can be an oligodeoxynucleoside phosphoramidate (P-NH2) or a mixed phosphoramidate-phosphodiester oligomer. In addition, a double-stranded polynucleotide can be obtained from the single stranded polynucleotide product of chemical synthesis either by synthesizing the complementary strand and annealing the strands under appropriate conditions, or by synthesizing the complementary strand de novo using a DNA polymerase with an appropriate primer.

[0044] As used herein, “treatment” or “treating” is an approach for obtaining beneficial or desired results, including clinical results. For purposes of this invention, beneficial or desired clinical results include, but are not limited to, alleviating one or more symptoms of a disease associated with a DNA mutation, e.g., a mitochondrial disease, reducing one or more symptoms of a disease, preventing one or more symptoms of a disease, treating one or more symptoms of a disease, ameliorating one or more symptoms of a disease, delaying onset of one or more symptoms associated with having a disease, diminishing the extent of one or more symptoms of a disease, stabilizing the disease (e.g., preventing or delaying the worsening of the disease) , delaying or slowing the progression of the disease, ameliorating one or more symptoms of a disease, decreasing the dose of one or more other medications and / or treatments required to treat the disease, increasing the quality of life of the individual, and / or prolonging survival of the individual. Also encompassed by “treatment” is a reduction of a pathological consequence of a disease associated with a DNA mutation, e.g., a mitochondrial disease. The methods of the invention contemplate any one or more of these aspects of treatment.

[0045] The term “individual” refers to a mammal and includes, but is not limited to, human, bovine, horse, feline, canine, rodent, or primate. In some embodiments, the individual is human.

[0046] The terms “comprising, ” “having, ” “containing, ” and “including, ” and other similar forms, and grammatical equivalents thereof, as used herein, are intended to be equivalent in meaning and to be open ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items, or meant to be limited to only the listed item or items. For example, an article “comprising” components A, B, and C can consist of (i.e., contain only) components A, B, and C, or can contain not only components A, B, and C but also one or more other components. As such, it is intended and understood that “comprises” and similar forms thereof, and grammatical equivalents thereof, include disclosure of embodiments of “consisting essentially of” or “consisting of. ”

[0047] Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, unless the context clearly dictate otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.

[0048] Reference to “about” a value or parameter herein includes (and describes) variations that are directed to that value or parameter per se. For example, description referring to “about X” includes description of “X. ”

[0049] As used herein, including in the appended claims, the singular forms “a, ” “or, ” and “the” include plural referents unless the context clearly dictates otherwise. II. Engineered DNA pyrimidine base-modifying polypeptides

[0050] In certain aspects, provided herein are engineered DNA pyrimidine base-modifying polypeptides for editing DNA comprising an engineered variant of a uracil-DNA glycosylase (UNG) . In some embodiments, the UNG is mutated to enable modification (via excision) of a DNA pyrimidine base such as thymine or cytosine. In some embodiments, the engineered DNA pyrimidine base-modifying polypeptides taught herein further comprise a DNA recognition domain, such as a double-stranded (ds) DNA binding polypeptide. In some embodiments, the engineered DNA pyrimidine base-modifying polypeptide comprises a dsDNA binding polypeptide fused to a UNG variant via translational fusion. In certain embodiments, the components of the polypeptide are configured such that the engineered UNG variant is brought into proximity of a target DNA pyrimidine to catalyze, at least in part, a desired nucleotide base edit. In some embodiments, the engineered DNA pyrimidine base-modifying polypeptide is configured as a thymine-modifying system, e.g., generating a T-to-C, T-to-G, or T-to-Aconversion. In some embodiments, the engineered DNA pyrimidine base-modifying polypeptide is configured as a cytosine-modifying system, e.g., generating a C-to-G, C-to-T, or C-to-Aconversion.

[0051] Certain aspects of the engineered DNA pyrimidine base-modifying polypeptides taught herein are discussed in more detail in a modular fashion below. One of ordinary skill in the art will readily understand how the aspects of the present description can be combined to obtain any engineered DNA pyrimidine base-modifying polypeptide encompassed by the teachings provided herein. The discussion of engineered DNA pyrimidine base-modifying polypeptides, including components and configurations thereof, in a modular fashion does not limit the scope of the description encompassed herein. A. Uracil-DNA glycosylases and variants thereof

[0052] Provided herein are engineered polypeptides comprising a variant of a uracil-DNA glycosylase (UNG) configured DNA pyrimidine base editing. UNGs are enzymes that natively cleave N-glycosidic bonds in DNA for the purpose of performing uracil base excision repair when a uracil (which is a nucleobase typically found in RNA) is improperly incorporated into DNA thereby helping to prevent mutagenesis. UNGs work by positioning the uracil base out of the DNA double helix and cleaving the N-glycosidic bond thereby excising the uracil while keeping the sugar-phosphate backbone intact. This generates an apurinic / apyrimidinic site, herein referred to as an “AP site, ” at which a different base can be inserted. The engineered DNA pyrimidine base-modifying polypeptides described herein comprise UNG variants modified to excise DNA pyrimidines, such as thymine or cytosine. In some embodiments, the UNG variant comprises at least the catalytic domain of the UNG from which it is derived. Following formation of the AP site, canonical base repair aids in completion of the desired edit.

[0053] In some embodiments, provided herein is a UNG variant comprising one or more amino acid substitutions which enlarge the active site pocket such that the UNG variant is capable of excising a target thymine or a target cytosine. The position of the one or more amino acid substitutions provided herein may be based on UNGs of different species. One of ordinary skill in the art will readily understand how to convert an amino acid numbering from one UNG to another, such as across species, e.g., by comparing sequences and / or looking for sequence and / or structural homology. For example, in some embodiments, the position of amino acid substitution is based on reference to the amino acid sequence set forth in SEQ ID NO: 1, a human UNG. In some embodiments, the UNG variant comprises a mutation of Y147 (based on SEQ ID NO: 1) to an amino acid having a smaller residue (such as determined by occupied volume or surface area) the tyrosine, wherein the UNG variant is capable of excising thymine. In some embodiments, the UNG variant comprises an amino acid substitution selected from a group consisting of Y147A, Y147S, Y147N, Y147G, and Y147C (based on SEQ ID NO: 1) , wherein the UNG variant is capable of excising thymine. In some embodiments, the UNG variant comprises an amino acid substitution of Y147A (based on SEQ ID NO: 1) . In some embodiments, the position of amino acid substitution is based on reference to the amino acid sequence set forth in SEQ ID NO: 7, a Deinococcus radiodurans UNG. In some embodiments, the UNG variant comprises an amino acid substitution selected from a group consisting of Y85A, Y85S, Y85N, Y85G, and Y85C (based on SEQ ID NO: 7) , wherein the UNG variant is capable of excising thymine. In some embodiments, the amino acid substitution is Y85A (based on SEQ ID NO: 7) .

[0054] In some embodiments, the UNG variant comprises an amino acid substitution of N204D (based on SEQ ID NO: 1) , wherein the UNG variant is capable of excising cytosine.

[0055] In certain embodiments, the UNG variant further comprises one or more additional amino acid substitutions. In some embodiments, the one or more additional amino acid substitutions comprise substitutions in conserved motifs of UNG. Examples of conserved UNG motifs are the catalytic water-activating loop (143-GQDPYH-148 in hUNG) , the Pro-rich loop, which compresses the DNA backbone 5’ to the lesion (165-PPPPS-169 in hUNG) , the Ura-binding motif (201-LLLN-204 in hUNG) , the Gly-Ser loop that compresses the DNA backbone 3’ to the lesion (246-GS-247 in hUNG) 5) , and the Leu-intercalation loop, which penetrates the minor groove (268-HPSPLS-273 in hUNG) . In some embodiments, the one or more amino acid substitutions comprise substitutions of one or more of L80V, K113E, K197E, S204A, S204G, or E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain embodiments, the UNG variant comprises amino acid substitutions of L80V, K113E, K197E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, UNG variant comprises amino acid substitutions of K113E and S204G, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain embodiments, the UNG variant comprises amino acid substitutions of K113E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the UNG variant comprises amino acid substitutions of L80V, wherein the positions of amino acid substitution is based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain embodiments, the UNG variant comprises amino acid substitutions of K197E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the UNG variant comprises amino acid substitutions of K113E, K197E, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7.

[0056] In some embodiments, the UNG variant comprises amino acid substitutions of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) and one or more of L80V, K113E, K197E, S204A, S204G, or E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the UNG variant comprises amino acid substitutions of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) , L80V, K113E, K197E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, UNG variant comprises amino acid substitutions of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) , K113E and S204G, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain embodiments, the UNG variant comprises amino acid substitutions of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) , K113E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the UNG variant comprises amino acid substitutions of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) , L80V, wherein the positions of amino acid substitution is based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain embodiments, the UNG variant comprises amino acid substitutions of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) , K197E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the UNG variant comprises amino acid substitutions of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) , K113E, K197E, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7.

[0057] The UNG variants described herein can be engineered to further comprise an N-terminal deletion relative to the UNG (such as a full-length UNG) from which it is derived. In some embodiments, the N-terminal deletion is a deletion of up to the first 30 amino acids based on reference to the amino acid sequence set forth in SEQ ID NO: 7, such as deletion of any of 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 amino acids. In certain embodiments, the N-terminal deletion comprises a deletion of the first 2 amino acids. For example, a deletion of the first two amino acids based on reference to the amino acid sequence set forth in SEQ ID NO: 7 would comprise deleting a methionine and threonine from the N-terminus. In some embodiments, the UNG variant is at least about 70%identical to the wildtype sequence of a corresponding truncated version of the UNG.

[0058] In certain embodiments, the variant of a UNG has at least about 65%, such as at least about any of 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%, sequence identity to a corresponding portion of a naturally occurring UNG, such as a UNG from Deinococcus radiodurans (e.g., SEQ ID NO: 7) . In some embodiments, the UNG variant has at least about 65%, such as at least about any of 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%, sequence identity to SEQ ID NO: 1, including a corresponding portion of SEQ ID NO: 1. In some embodiments, the UNG variant has at least about 65%, such as at least about any of 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%, sequence identity to SEQ ID NO: 16, including a corresponding portion of SEQ ID NO: 16. In some embodiments, the UNG variant has at least about 65%, such as at least about any of 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%, sequence identity to SEQ ID NO: 4, including a corresponding portion of SEQ ID NO: 4. In some embodiments, the UNG variant has at least about 65%, such as at least about any of 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99%, sequence identity to SEQ ID NO: 19, including a corresponding portion of SEQ ID NO: 19.

[0059] Many organisms possess UNG enzymes for the purpose of uracil base excision. Thus, engineered UNG variants described herein may be derived from a specific species, including but not limited to mammals (human, mouse, swine, bovine, monkey, horse, etc. ) , bacteria (Escherichia coli, Deinococcus radiodurans, etc. ) , viruses (Poxviridae, Herpesviridae, etc. ) , and yeast. In some embodiments of the invention, an engineered UNG variant described herein is derived from Homo sapiens (UniprotKB No. P13051-2) , Escherichia coli (UniprotKB No. P12295) , Deinococcus radiodurans (UniprotKB No. Q9RWH9) , Human Herpesvirus 1 (UniprotKB No. P10186) , or Vaccinia virus (UniprotKB No. P04303, P20536, or P04303) . In some embodiments, the engineered UNG variant described herein is derived from Deinococcus radiodurans. B. DNA recognition domains

[0060] In certain aspects, provided herein is an engineered polypeptide for DNA pyrimidine base modification, wherein the engineered polypeptide comprises a variant of a UNG and a DNA recognition domain. In some embodiments, the DNA recognition domain is a double-stranded (ds) DNA recognition domain.

[0061] As described herein, the DNA recognition domain comprises a polypeptide (and in some embodiments comprises a nucleic acid component, such as a guide) configured to bind to a dsDNA, such as at a specific location to enable action of the engineered UNG variant at a specific target site. As described in more detail below, the DNA recognition domains can be configured (e.g., is programmable) to bind to a specific location, including strand, of dsDNA relative to the target DNA pyrimidine to be edited. In some embodiments, the DNA recognition domain does not comprise a nucleic acid component, e.g., the DNA recognition domains is a TALE or zinc finger, or a portion thereof capable of binding to the dsDNA. In other embodiments, the DNA recognition domain comprises a domain of a nucleic acid programmable endonuclease such as Cas9, CasX, CasY, Cpf1, C2c1, C2c2, or C2c3, that binds DNA. In some embodiments, the DNA recognition domains is a catalytically inactive form of a nickase, such as an endonuclease. In some embodiments, the DNA recognition domain does not comprise nickase activity.

[0062] In some embodiments, the DNA recognition domain is a Cas9 protein. Cas9 recognizes a protospacer adjacent motif (PAM) of 5’-NGG-3’. Cas9 comprises a RuvC domain for cleavage of the non-target strand and a HNH domain for cleavage of the target strand. In certain embodiments, the DNA recognition domain is nicking Cas9, i.e., nCas9. nCas9 is a mutated form of Cas9 wherein an amino acid substitution disables cleavage of one strand of dsDNA. In some embodiments, this mutation is D10A in the RuvC domain, enabling cleavage of only the target strand. In some embodiments, this mutation is H840A in the HNH domain, enabling cleavage of only the non-target strand. In certain embodiments, the DNA recognition domain is dead Cas9, i.e., dCas9. dCas9 is a mutated form of Cas9 with no endonuclease activity, comprising mutations of D10A in the RuvC domain and H840A in the HNH domain. As used herein, D10A is based on the Streptococcus pyogenes sequence. In some embodiments, the DNA recognition domain is derived from Streptococcus pyogenes.

[0063] In some embodiments, the DNA recognition domain comprises a guide (such as a guide RNA) for targeting a specific DNA site. In some embodiments, the guide is a Crispr RNA (targeting) and tracr RNA (scaffolding to Cas protein) as a single RNA molecule known as single guide RNA (sgRNA) . sgRNA is known in the field, e.g., US Pat. No. 11, 015, 193, which is hereby incorporated by reference herein in its entirety. C. Systems, configurations, and targeting of polypeptides comprising a UNG variant and a  DNA recognition domain

[0064] Provided herein, in certain aspects, are DNA pyrimidine base-editing systems comprising a UNG variant and a DNA recognition domain. In some embodiments, the DNA pyrimidine base-editing system further comprises a nucleic acid for targeted editing and such will be apparent based on the type of DNA recognition domain used. The DNA pyrimidine base-editing systems encompassed herein can be formed from any combination or arrangement of a described UNG variant and a DNA recognition domain, and, optionally, in some embodiments a targeting nucleic acid such as a sgRNA.

[0065] In certain aspects, the DNA pyrimidine base-editing systems are configured such that a UNG variant is brought into proximity of a target DNA pyrimidine to catalyze, at least in part, a target thymine or cytosine nucleotide base edit. In some embodiments, the DNA pyrimidine base-editing system targets a thymine or cytosine on a target DNA strand. In some embodiments, the DNA pyrimidine base-editing system targets a thymine or cytosine in a non-strand specific manner, i.e., can excise a thymine or cytosine on either strand of DNA. In some embodiments, the provided systems are configured to target a thymine that is not in close proximity to another thymine on the opposite DNA strand, such as to avoid formation of a double strand break. In such embodiments, the provided systems are configured to target a thymine that is not within 2 base pairs of another thymine on the opposite DNA strand. Thus, in some aspects provided herein, is a selection step of picking a suitable target thymine for editing. In some embodiments, the provided systems are configured to target a cytosine that is not in close proximity to another cytosine on the opposite DNA strand, such as to avoid formation of a double strand break. In such embodiments, the provided systems are configured to target a cytosine that is not within 2 base pairs of another cytosine on the opposite DNA strand. Thus, in some aspects provided herein, is a selection step of picking a suitable target cytosine for editing.

[0066] In some embodiments, the DNA pyrimidine base-editing system comprises a UNG variant and a DNA recognition domain, wherein the UNG variant and DNA recognition domain form a complex configured such that the DNA recognition domain, when associated with a dsDNA, positions the UNG variant such that it can act on the target DNA pyrimidine. If the DNA recognition domain uses an sgRNA, then it can be selected / designed to be complementary to a sequence in proximity to the target site. Proper sgRNA design, including construction and / or selection for a desired specificity to limit off-target editing, is well known in the art (for example, Naeem M, et al, Cells, 9, 2020) . Such known design aspects are encompassed by the teachings provided herein and include optimization based on any one or more of biased or unbiased off-target detection, GC content, length, mismatches, or chemical modifications. Polypeptide only DNA recognition domains are conceptual similar as can be selected / designed to bind DNA in proximity to the target site. In some embodiments, the UNG variant and DNA recognition domain are guided to the editing region by a nucleic acid, e.g., an sgRNA. In some embodiments, the target DNA pyrimidine is at a specified position on a protospacer targeted by the sgRNA. In certain embodiments, the target DNA pyrimidine is at positions 1 to 30 base pairs, 5 to 25 base pairs, 10 to 20 base pairs, or 14 to 18 base pairs on the protospacer. In some embodiments, the target DNA pyrimidine on the dsDNA is 0-24 base pairs in length, including any of 1-20 base pairs in length, 1-10 base pairs in length, 10-20 base pairs in length, 10-16 base pairs in length, 14-20 base pairs in length, or 14-16 base pairs in length. In some embodiments, directly preceding the target protospacer is a protospacer adjacent motif (PAM) , which is required for Cas endonuclease activity. In certain embodiments, the PAM sequence is NGG. In some embodiments, the PAM sequence is a relaxed NG.

[0067] The complex of a UNG variant and a DNA recognition domain can be configured in various ways, such as via direct fusion (e.g., as a single expressed polypeptide) , non-covalent interaction, or covalent linkage (e.g., via a polypeptide or non-polypeptide linker) . In some embodiments, the UNG variant is fused to the DNA recognition domain. In some embodiments, the DNA recognition domain is fused to the C-terminus of a UNG variant. In some embodiments, the DNA recognition domain is fused to the N-terminus of a UNG variant. In some embodiments, the DNA pyrimidine base-editing system comprises a linker associating a DNA recognition domain and a UNG variant. In some embodiments, the linker is a polypeptide linker. In some embodiments, additional adjustments may be included for configuring the position of the UNG variant, such as by adjusting a linker length connecting a DNA recognition domain and a UNG variant. The present disclosure encompasses such variations of the DNA pyrimidine base-editing systems described herein, e.g., many variations of a pyrimidine base-editing system may be designed based on the teachings provided herein that perform the same DNA edit. In some embodiments, the DNA recognition domain is configured to bind 0-25 base pairs away, such as any of 10-20, 12-20, or 14-20 base pairs away, from a target base of a UNG variant described herein. In some embodiments, the linker comprises a polypeptide linker, such as a GGGGS linker. In some embodiments, the polypeptide linker is from 1-100 amino acids in length, such as any of 1-90 amino acids in length, 1-80 amino acids in length, 1-70 amino acids in length, or 1-65 amino acids in length. In some embodiments, the polypeptide linker is any of the following amino acids in lengths: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 2, 4 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 amino acids. In some embodiments, the linker is selected from a list composed of GS, SGGS, PAPAP, XTEN, or repeats thereof. In some embodiments, the linker is at least about 30 amino acids in length, such as a 32-amino acid linker or a 64-amino acid linker. In some embodiments, the linker is at least about 30 amino acids in length, such as a 32-amino acid linker or a 64-amino acid linker, and is selected from a list composed of GS, SGGS, PAPAP, XTEN, or repeats thereof. In some embodiments, the linker is a 64-amino acid linker. In some embodiments, the linker is a 32-amino acid linker. In some embodiments, the linker is GGGGS (SEQ ID NO: 27) . In some embodiments, the linker is GGGGSGGGGSGGGGS (SEQ ID NO: 28) .

[0068] In some embodiments, the DNA pyrimidine base-modifying system can generate a specific thymine conversion at a target site. In certain embodiments, the conversion is a T-to-C, T-to-G, or T-to-Aconversion. In certain embodiments, when an edit occurs, the DNA pyrimidine base-modifying system generates a T-to-C conversion at a rate of about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, or about 75%. In certain embodiments, when an edit occurs, the DNA pyrimidine base-modifying system generates a T-to-G conversion at a rate of about 1%, about 2%, about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, or about 65%. In certain embodiments, when an edit occurs, the DNA pyrimidine base-modifying system generates a T-to-Aconversion at a rate of about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, or about 50%. In some embodiments, the DNA pyrimidine base-editing system can generate a specific cytosine conversion at a target site. In some embodiments, conversion is a C-to-G, C-to-T, or C-to-Aconversion. In some embodiments, the conversion is C-to-G.

[0069] As described herein, the DNA recognition domains of a pyrimidine base-editing system may be configured to position an associated UNG variant relative to one or more target DNA pyrimidines. In some embodiments, the DNA recognition domain recognizes DNA that occurs at a single target site. In some embodiments, the DNA recognition domain recognizes DNA that occurs at different target sites (e.g., the target sequence recognized by the DNA recognition domain occurs at two or more different locations of the genome) .

[0070] In some embodiments, the dsDNA targeted for editing by the DNA pyrimidine base-editing systems described herein can be any type of double-stranded DNA. In some embodiments, the dsDNA is a circularized dsDNA. In some embodiments, the dsDNA is a B-DNA conformation. In some embodiments, the dsDNA is an A-DNA conformation. In some embodiments, the dsDNA is a Z-DNA conformation.

[0071] Nucleotide editing can occur outside of the editing region, such as off-target editing. Notably, however, the invention described herein exhibits significantly lower off-target editing compared to existing deaminase-based editing tools. In some embodiments, the systems provided herein have an off-target editing rate of 30%or less, such as any of about 15%or less, 10%or less, 9%or less, 8%or less, 7%or less, 6%or less, 5%or less, 4%or less, 3%or less, 2%or less, or 1%or less. In some embodiments, the systems provided herein have an off-target editing rate of about 1-30%, such as any of about 1-15%, about 1-10%, about 1-9%, about 1-8%, about 1-7%, about 1-6%, about 1-5%, about 1-4%, about 1-3%, about 1-2%, or about 1%. In some embodiments, the off-target editing is based on off-target DNA editing. In some embodiments, the off-target editing is based on off-target DNA editing. In some embodiments, the off-target editing is based on off-target RNA editing. In some embodiments, the off-target editing is based on off-target DNA and RNA editing. In some embodiments, off-target editing rates are determined by high-throughput sequencing. D. Additional components

[0072] In certain aspects, the DNA pyrimidine base-editing systems provided herein may comprise one or more additional features, such as to aid in delivery and / or function of the engineered DNA pyrimidine base-modifying polypeptides.

[0073] Mitochondria are unique sub-cellular organelles that possess their own DNA and RNA and mechanisms for their translation, yet they express only 10%of the proteins that they contain. Instead, mitochondria rely in part on the translation products of nuclear genes. These products traverse the cytoplasm and are ‘imported’ into the mitochondria via a system of outer-and inner-membrane-bound protein complexes, where they are delivered to the appropriate mitochondrial compartment and rendered active. This mitochondrial import process is regulated by an N-terminal pre-sequence in the nuclear gene of the protein that tags the protein with a sequence that tells the import machinery where the protein should be delivered-these are known as mitochondrial location signal (MLS) , which can also be referred to as a mitochondrial targeting signal (MTS) , peptides. Once the protein has been transported to the desired compartment, the MTS portion of the protein may be removed by a mitochondrial peptidase, allowing the protein to fold into its functional state and become active.

[0074] In some embodiments, the DNA pyrimidine base-editing system, or one or more components thereof, comprise a mitochondrial location signal (MLS) , which can also be referred to as a mitochondrial targeting signal (MTS) . In some embodiments, the MLS is fused to a UNG variant. In some embodiments, the MLS is fused to a DNA recognition domain. In some embodiments, wherein the DNA pyrimidine base-editing system comprises more than one unit, any one or more, including all, units of the engineered DNA pyrimidine base-modifying polypeptide may comprise a MLS.

[0075] In some embodiments, the MLS is about 10 to about 80 amino acids in length, such as about any of 10-15, 15-20, 20-25, 25-30, 30-35, 35-40, 40-45, 45-50, 50-55, 55-60, 60-65, 65-70, 70-75, or 75-80 amino acids in length.

[0076] In some embodiments, the MLS comprises an amphipathic helix structural motif. In order to adopt the amphipathic helix structural motif, the MLS can be enriched in basic (e.g., Arg, Lys) , hydroxylated (e.g., Ser, Thr) and / or hydrophobic (e.g., Ala, Leu, Ile) residues. In some embodiments, the MLS comprising an amphipathic helix structural motif exhibit alternating hydrophobic and hydrophilic segments. In some embodiments, at least about 20% (such as at least about any of 30%, 40%, 50%, or 60%) of the amino acid residues in the MLS are basic amino acid residues. In some embodiments, at least about 20% (such as at least about any of 30%, 40%, 50%, or 60%) of the amino acid residues in the MLS are hydrophobic amino acid residues. In some embodiments, the MLS is amphipathic, for example forms an amphipathic helix. In some embodiments, the MLS comprises an alternating pattern of hydrophobic and basic residues. In some embodiments, the MLS is derived from a protein selected from the group consisting of ATP synthase, cytochrome C oxidase peptide VIII, Su9, and HSP60. In some embodiments, the MLS can selectively direct a compound to an outer membrane, an inner membrane, and inter-membrane space, or a mitochondrial matrix.

[0077] In some embodiments, the DNA pyrimidine base editing system does not comprise a deaminase domain having deamination activity. In some embodiments, the engineered DNA pyrimidine base-modifying polypeptides, or systems thereof, do not comprise a deaminase, or a functional unit therefrom. E. Example engineered DNA pyrimidine base-modifying polypeptides

[0078] In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) L80V, K113E, K197E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) L80V, K113E, K197E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7, and a DNA recognition domain. In some embodiments, the Y85 substitution is Y85A. In some embodiments, the DNA recognition domain is a Cas nuclease (such as nCas9 or dCas9) or a TAL effector. In some embodiments, the engineered thymine-modifying polypeptide comprises a linker linking the UNG variant and DNA recognition domain (such as a polypeptide linker, e.g., a 64-amino acid linker) . In some embodiments, the UNG variant comprises an N-terminal deletion relative to the full-length UNG from which it is derived, such as deletion of 1 or 2 amino acids. In some embodiments, the engineered thymine-modifying polypeptide comprises, including is or consists of, SEQ ID NO: 9. In some embodiments, the UNG is derived from Deinococcus radiodurans.

[0079] In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) K113E and S204G, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) K113E and S204G, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7, and a DNA recognition domain. In some embodiments, the Y85 substitution is Y85A. In some embodiments, the DNA recognition domain is a Cas nuclease (such as nCas9 or dCas9) or a TAL effector. In some embodiments, the engineered thymine-modifying polypeptide comprises a linker linking the UNG variant and DNA recognition domain (such as a polypeptide linker, e.g., a 64-amino acid linker) . In some embodiments, the UNG variant comprises an N-terminal deletion relative to the full-length UNG from which it is derived, such as deletion of 1 or 2 amino acids. In some embodiments, the engineered thymine-modifying polypeptide comprises, including is or consists of, SEQ ID NO: 10. In some embodiments, the UNG is derived from Deinococcus radiodurans.

[0080] In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) K113E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) K113E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7, and a DNA recognition domain. In some embodiments, the Y85 substitution is Y85A. In some embodiments, the DNA recognition domain is a Cas nuclease (such as nCas9 or dCas9) or a TAL effector. In some embodiments, the engineered thymine-modifying polypeptide comprises a linker linking the UNG variant and DNA recognition domain (such as a polypeptide linker, e.g., a 64-amino acid linker) . In some embodiments, the UNG variant comprises an N-terminal deletion relative to the full-length UNG from which it is derived, such as deletion of 1 or 2 amino acids. In some embodiments, the engineered thymine-modifying polypeptide comprises, including is or consists of, SEQ ID NO: 12. In some embodiments, the UNG is derived from Deinococcus radiodurans.

[0081] In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) L80V, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) L80V, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7, and a DNA recognition domain. In some embodiments, the Y85 substitution is Y85A. In some embodiments, the DNA recognition domain is a Cas nuclease (such as nCas9 or dCas9) or a TAL effector. In some embodiments, the engineered thymine-modifying polypeptide comprises a linker linking the UNG variant and DNA recognition domain (such as a polypeptide linker, e.g., a 64-amino acid linker) . In some embodiments, the UNG variant comprises an N-terminal deletion relative to the full-length UNG from which it is derived, such as deletion of 1 or 2 amino acids. In some embodiments, the engineered thymine-modifying polypeptide comprises, including is or consists of, SEQ ID NO: 13. In some embodiments, the UNG is derived from Deinococcus radiodurans.

[0082] In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) K197E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) K113E, S204A, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7, and a DNA recognition domain. In some embodiments, the Y85 substitution is Y85A. In some embodiments, the DNA recognition domain is a Cas nuclease (such as nCas9 or dCas9) or a TAL effector. In some embodiments, the engineered thymine-modifying polypeptide comprises a linker linking the UNG variant and DNA recognition domain (such as a polypeptide linker, e.g., a 64-amino acid linker) . In some embodiments, the UNG variant comprises an N-terminal deletion relative to the full-length UNG from which it is derived, such as deletion of 1 or 2 amino acids. In some embodiments, the engineered thymine-modifying polypeptide comprises, including is or consists of, SEQ ID NO: 14. In some embodiments, the UNG is derived from Deinococcus radiodurans.

[0083] In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) K113E, K197E, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7. In certain aspects, provided is an engineered thymine-modifying polypeptide comprising a variant of a uracil-DNA glycosylase (UNG) comprising: (a) an amino acid substitution of Y85 (such as any of Y85A, Y85S, Y85N, Y85G, and Y85C) ; and (b) K113E, K197E, and E237Q, wherein the positions of amino acid substitutions are based on reference to the amino acid sequence set forth in SEQ ID NO: 7, and a DNA recognition domain. In some embodiments, the Y85 substitution is Y85A. In some embodiments, the DNA recognition domain is a Cas nuclease (such as nCas9 or dCas9) or a TAL effector. In some embodiments, the engineered thymine-modifying polypeptide comprises a linker linking the UNG variant and DNA recognition domain (such as a polypeptide linker, e.g., a 64-amino acid linker) . In some embodiments, the UNG variant comprises an N-terminal deletion relative to the full-length UNG from which it is derived, such as deletion of 1 or 2 amino acids. In some embodiments, the engineered thymine-modifying polypeptide comprises, including is or consists of, SEQ ID NO: 15. In some embodiments, the UNG is derived from Deinococcus radiodurans. F. Polynucleotide forms of the polypeptides described herein

[0084] Provided herein, in certain aspects, are polynucleotide forms of engineered DNA pyrimidine base-modifying polypeptides described herein. The disclosure provided herein covers a multitude of formats capable of introducing functional engineered DNA pyrimidine base-modifying polypeptides described herein in a cell, including different types of polynucleotides (e.g., DNA or RNA, such as circular RNA) and different designs of polynucleotides.

[0085] “Introducing” or “introduction” used herein in reference to delivering engineered DNA pyrimidine base-modifying polypeptides means delivering one or more components of a DNA pyrimidine base-modifying polypeptide, or a precursor thereof (e.g., one or more polynucleotides encoding an engineered DNA pyrimidine base-modifying polypeptide or a component thereof) , to a cell. The compositions and methods of the present application can employ many delivery systems, including but not limited to, viral, liposome, electroporation, nanoparticle, microinjection and conjugation, to achieve the introduction of the construct as described herein into a cell. Conventional viral and non-viral based gene transfer methods can be used to introduce nucleic acids into mammalian cells or target tissues. Such methods can be used to administer nucleic acids of the present application to cells in culture, or in a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., a transcript of a construct described herein) , naked nucleic acid, and nucleic acid complexed with a delivery vehicle, such as a liposome. Viral vector delivery systems include DNA and RNA viruses, which have either episomal or integrated genomes for delivery to the cell. As described herein, the polynucleotides taught herein encoding an engineered DNA pyrimidine base-modifying polypeptide may be one or more polynucleotides. In some embodiments, introducing the engineered DNA pyrimidine base-modifying polypeptide may comprise introducing two or more different polynucleotides, wherein said two or more different polynucleotides may be introduced simultaneously, sequentially, or concurrently, including introduced simultaneously, sequentially, or concurrently into the cell.

[0086] Methods of non-viral delivery of one or more components of an engineered DNA pyrimidine base-modifying polypeptide, including nucleic acids, include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid: nucleic acid conjugates, electroporation, nanoparticles, exosomes, microvesicles, or gene-gun, naked DNA and artificial virions.

[0087] In some embodiments, the polynucleotides described herein comprise additional features useful for expression of a nucleobase editor system in a cell, such as a promoter sequence. A promoter can be constitutively active, meaning that the promoter is always active in a given cellular context, or conditionally active, meaning that the promoter is only active in the presence of a specific condition. For example, a conditional promoter may only be active in the presence of a specific protein that connects a protein associated with a regulatory element in the promoter to the basic transcriptional machinery, or only in the absence of an inhibitory molecule. A subclass of conditionally active promoters are inducible promoters that require the presence of a small molecule “inducer” for activity. Examples of inducible promoters include, but are not limited to, arabinose-inducible promoters, Tet-on promoters, and tamoxifen-inducible promoters. A variety of constitutive, conditional, and inducible promoters are well known to the skilled artisan, and the skilled artisan will be able to ascertain a variety of such promoters useful in carrying out the instant invention, which is not limited in this respect.

[0088] The use of RNA or DNA viral based systems for the delivery of nucleic acids has high efficiency in targeting a virus to specific cells and trafficking the viral payload to the cellular nuclei. In some embodiments, delivery comprises introducing a viral vector (such as lentiviral vector) encoding the nucleic acid (s) to the cell. In some embodiments, the viral vector is an AAV, e.g., AAV8. In some embodiments, delivery comprises introducing a plasmid encoding one or more engineered DNA pyrimidine base-modifying polypeptide components to the cell. In some embodiments, delivery comprises introducing (e.g., by electroporation) one or more engineered DNA pyrimidine base-modifying polypeptide components into the cell. In some embodiments, delivery comprises transfection of one or more engineered DNA pyrimidine base-modifying polypeptide components into the cell.

[0089] In some embodiments, the polynucleotide, such as the polynucleotide introduced to a cell, is DNA or RNA. In some embodiments, the RNA is linear RNA. In some embodiments, the RNA is circular RNA. In some embodiments, the linear RNA is capable of forming a circular RNA. The circulation can be performed, for example, by using the Tornado expression system ( “Twister-optimized RNA for durable overexpression” ) as described in Litke, J.L. &Jaffrey, S.R. Highly efficient expression of circular RNA aptamers in cells using autocatalytic transcripts. Nat Biotechnol 37, 667-675 (2019) , which is hereby incorporated herein by reference in its entirety. Briefly, Tornado-expressed transcripts contain an RNA of interest flanked by Twister ribozymes. A twister ribozyme is any catalytic RNA sequences that are capable of self-cleavage. The ribozymes rapidly undergo autocatalytic cleavage, leaving termini that are ligated by an RNA ligase. Non-limiting examples of RNA ligase include: RtcB, T4 RNA Ligase 1, T4 RNA Ligase 2, Rnl3 and Trl1. In some embodiments, the RNA ligase is expressly endogenously in the cell. In some embodiments, the RNA ligase is RNA ligase RtcB. In some embodiments, the method further comprises introducing an RNA ligase (e.g., RtcB) into the cell. In some embodiments, the RNA is circularized before being introduced to the cell. In some embodiments, the RNA is chemically synthesized. In some embodiments, the RNA is circularized through in vitro enzymatic ligation (e.g., using RNA or DNA ligase) or chemical ligation (e.g., using cyanogen bromide or a similar condensing agent) .

[0090] In some embodiments, the polynucleotides described herein comprise additional features useful for expression of an engineered DNA pyrimidine base-modifying polypeptide in a cell, such as a promoter sequence. III. Methods of making and use

[0091] In other aspects, provided herein are methods of using the engineered DNA pyrimidine base-modifying polypeptides and systems described herein. In some embodiments, provided is a method of excising a target thymine from a nucleic acid molecule. In other aspects, provided herein are methods of excising a target cytosine from a nucleic acid molecule.

[0092] In some embodiments, provided is method of modifying a target DNA pyrimidine (such as a target thymine or a target cytosine) in a nucleic acid sequence present in one or more nucleic acid molecules, the method comprising contacting at least one nucleic acid molecule of the one or more nucleic acid molecules with a pyrimidine base-editing system described herein. In some embodiments, the DNA pyrimidine base-editing system, or at least a component thereof, is in a polypeptide form. In some embodiments, the DNA pyrimidine base-editing system, or at least a component thereof, is in a polynucleotide form, such as one or more polynucleotides encoding the engineered DNA pyrimidine base-modifying polypeptide, or a component thereof. In such methods, the one or more polynucleotides are configured such that the associated polypeptide of the engineered DNA pyrimidine base-modifying polypeptides is expressed in the cell. In some embodiments, the method comprises delivering one or more nucleic acids encoding a DNA pyrimidine base-editing system described herein.

[0093] As relevant to certain methods provided herein, specificity of a Cas endonuclease is enabled by guide RNA. Crispr RNA (targeting) and tracr RNA (scaffolding to Cas protein) can be synthetically produced as a single RNA molecule known as single guide RNA (sgRNA) . In some embodiments, the methods comprise, including further comprise, delivering a sgRNA to a cell.

[0094] In some embodiments, provided is a method of modifying a target DNA pyrimidine (such as a target thymine or a target cytosine) in a nucleic acid sequence present in one or more nucleic acid molecules, the method comprising contacting at least one nucleic acid molecule of the one or more nucleic acid molecules with an engineered DNA pyrimidine-modifying polypeptide described herein comprising a DNA recognition domain, wherein the DNA recognition domain of said engineered DNA pyrimidine-modifying polypeptide is configured to associate with the at least one nucleic acid molecule such that the UNG variant of the engineered DNA pyrimidine-modifying polypeptide is positioned to modify the target DNA pyrimidine. For example, in some embodiments, the method comprises a method of modifying a target thymine in a nucleic acid sequence present in one or more nucleic acid molecules, wherein the target thymines in the one or more nucleic acid molecules undergo a T-to-C, T-to-G, or T-to-Amodification. In some embodiments, an NGG PAM site is 14 to 19 base pairs away from the target thymine. In some embodiments, an NGG PAM site is 14 to 19 base pairs away from the target thymine, wherein counting starts at the first base outside of the NGG PAM in a protospacer, and wherein the NGG PAM site is associated with DNA recognized by the DNA recognition domain, such as a Cas nuclease. In some embodiments, a relaxed NG PAM site is 14 to 19 base pairs away from the target thymine. In some embodiments, a relaxed NG PAM site is 14 to 19 base pairs away from the target thymine, wherein counting starts at the first base outside of the NG PAM in a protospacer, and wherein the relaxed NG PAM site is associated with DNA recognized by the DNA recognition domain, such as a Cas nuclease.

[0095] In some embodiments, the efficiency of editing of a target DNA nucleotide base is at least about 2%, such as at least about any of 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, or 65%. In some embodiments, the method exhibits an editing efficiency at least 2-fold greater than that of a method comprising delivering an engineered thymine-modifying polypeptide which does not comprise one or more amino acid substitutions of L80V, K113E, K197E, S204A, S204G, or E237Q, to a cell. In some embodiments, the efficiency of editing is determined by Sanger sequencing. In some embodiments, the efficiency of editing is determined by next-generation sequencing.

[0096] In some embodiments, the method has a low off-target editing rate. In some embodiments, the method has lower than about 1% (e.g., no more than about any one of 0.5%, 0.1%, 0.05%, 0.01%, 0.001%or lower) editing efficiency on a non-target DNA nucleotide base as compared to the target DNA nucleotide base. In some embodiments, the method does not edit non-target DNA nucleotide bases.

[0097] In some embodiments, the methods comprise delivering a pyrimidine base-editing system to a cellular location. In some embodiments, the cellular location is the cytoplasm. In some embodiments, the cellular location is a mitochondrion and the engineered DNA pyrimidine base-modifying polypeptide comprises a MLS.

[0098] In some embodiments, the method can be practiced in a variety of cell lines. In certain embodiments, the cell line is selected from a group consisting of human cell lines HEK293T and HCT116, mouse cell line N2A, and monkey cell lines NIH3T3 and Cos-7. In some embodiments, the cell line is HEK293T.

[0099] Use of base editing systems can risk the formation of double-stranded breaks in DNA; thus, methods of base editing may include measures to mitigate double-stranded break formation. In some embodiments described herein, the method of modifying a target DNA pyrimidine in a nucleic acid sequence comprises selecting a target DNA pyrimidine that is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 base pairs away from an editable base on the opposite strand.

[0100] In some embodiments, the method results in low cytotoxicity. In some embodiments, cytotoxicity can be determined by percent of viable cells 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 days (e.g., 3 days) after transfection with an editor system. In some embodiments, the methods of modifying a target DNA pyrimidine described herein result in less than about a 1.5-fold, less than about a 1.4-fold, less than about a 1.3-fold, less than about a 1.2-fold, or less than about a 1.1-fold decrease in viable cell count compared to untreated cells. In some embodiments, the methods of modifying a target DNA pyrimidine described herein result in about no decrease in viable cell count compared to untreated cells. In some embodiments, cytotoxicity can be determined by cell growth / proliferation, as measured by optical density, 72 hours after transfection with an editor system. In some embodiments, the methods of modifying a target DNA pyrimidine described herein result in less than about a 1.2-fold or less than about a 1.1-fold decrease in optical density at 72 hours compared to untreated cells. In some embodiments, the methods of modifying a target DNA pyrimidine described herein result in about no decrease in optical density at 72 hours compared to untreated cells.

[0101] The purposes of editing a target nucleotide in a cell, are diverse, all of which are encompassed by the description provided herein. For example, in some embodiments, the method of editing a target nucleotide in a cell using a pyrimidine base-editing system provided herein is performed to create a cell model. In some embodiments, the target nucleotide is the site of a known SNP, wherein the nucleobase editor system is configured to edit the target nucleotide to revert the SNP to the wild-type nucleotide, create the SNP, or adjust the SNP associated with a disease to another nucleotide base.

[0102] In some embodiments, provided is a method of treating a disease (such as a disease associated with a DNA mutation) in an individual, the method comprising administering to the individual a pyrimidine base-editing system described herein, or a precursor thereof. In some embodiments, the DNA pyrimidine base-editing system is configured to edit a base associate with the disease. In some embodiments, the DNA pyrimidine base-editing system is configured to edit a base associated with treatment, in whole or in part, of the disease. In some embodiments, the DNA pyrimidine base-editing system described herein, or a precursor thereof, administered to an individual is formulated for intracellular delivery.

[0103] An exemplary disease-relevant mutation that can be corrected by the provided fusion proteins in vitro or in vivo is a mutation resulting in Hurler syndrome. Hurler syndrome is caused by the presence of a premature stop codon on exon 9 of IDUA gene, which prevents the production of functional IDUA protein. In some embodiments, the method of treating a disease comprises contacting a cell carrying a mutation to be corrected, e.g., a GM06214 (IDUAW402X) cell derived from patient with Hurler syndrome, with a pyrimidine base-editing system described herein. In some embodiments, the method of treating a disease results in at least about 50%, at least about 45%, at least about 40%, at least about 35%, at least about 30%, at least about 25%, or at least about 20%editing efficiency of the targeted site. In some embodiments, the method of treating a disease results in restoration of IDUA catalytic activity in a GM06214 (IDUAW402X) cell derived from patient with Hurler syndrome. In some embodiments, the restoration of IDUA catalytic activity is at least about 3-fold, at least about 2.5-fold, at least about 2-fold, or at least about 1.5-fold.

[0104] Another exemplary disease-relevant mutation that can be corrected by the provided fusion proteins in vitro or in vivo is the H1047R (A3140G) polymorphism in the PI3KCA protein. The phosphoinositide-3-kinase, catalytic alpha subunit (PI3KCA) protein acts to phosphorylate the 3-OH group of the inositol ring of phosphatidylinositol. The PI3KCA gene has been found to be mutated in many different carcinomas, and thus it is considered to be potent oncogene. In fact, the A3140G mutation is present in several NCI-60 cancer cell lines, such as, for example, the HCT116, SKOV3, and T47D cell lines, which are readily 51 available from the American Type Culture Collection (ATCC) .

[0105] In some embodiments, a cell carrying a mutation to be corrected, e.g., a cell carrying a point mutation, e.g., an A3140G point mutation in exon 20 of the PI3KCA gene, resulting in a H1047R substitution in the PI3KCA protein, is contacted with an expression construct encoding a Cas9 DNA-editing fusion protein and an appropriately designed sgRNA targeting the fusion protein to the respective mutation site in the encoding PI3KCA gene. Control experiments can be performed where the sgRNAs are designed to target the fusion enzymes to non-C residues that are within the PI3KCA gene. Genomic DNA of the treated cells can be extracted, and the relevant sequence of the PI3KCA genes PCR amplified and sequenced to assess the activities of the fusion proteins in human cell culture.

[0106] It will be understood that the example of correcting point mutations in PI3KCA is provided for illustration purposes and is not meant to limit the instant disclosure. The skilled artisan will understand that the instantly disclosed DNA-editing fusion proteins can be used to correct other point mutations and mutations associated with other cancers and with diseases other than cancer including other proliferative diseases.

[0107] The target nucleotide sequence may comprise a target base (e.g., a point mutation) associated with a disease, disorder, or condition, such as sickle cell anemia, Fanconi anemia, ectodermal dysplasia skin fragility syndrome, lattice corneal dystrophy Type III, or Noonan syndrome. The target sequence may comprise a T to A point mutation associated with a disease, disorder, or condition, and wherein the methylation of the mutant A base results in mismatch repair-mediated correction to a sequence that is not associated with a disease, or disorder, or condition. The target sequence may instead comprise an A to T point mutation associated with a disease, disorder, or condition, and wherein the methylation of the A base paired with the mutant T results in mismatch repair-mediated correction to a sequence that is not associated with a disease, or disorder, or condition. The target sequence may encode a protein, and where the point mutation is in a codon and results in a change in the amino acid encoded by the mutant codon as compared to a wild-type codon. The target sequence may also be at a splice site, and the point mutation results in a change in the splicing of an mRNA transcript as compared to a wild-type transcript. In addition, the target may be at a non-coding sequence of a gene, such as a promoter, and the point mutation results in increased or decreased expression of the gene.

[0108] Exemplary target genes include HBB, in which an A to T point mutation at residue 334 results in a sickle cell anemia phenotype; and FANCC, in which an A to T point mutation at residue 456 results in a Fanconi anemia phenotype. Additional target genes include TGFBI (associated with lattice corneal dystrophy type III) , PKP1 (associated with ectodermal dysplasiaskin fragility syndrome) , KRAS and SOS1 (both associated with Noonan syndrome) , for which the disease phenotype is frequently caused by T: Ato A: T point mutations.

[0109] In some embodiments, the edit results in alleviation of a premature stop codon. In some embodiments, the edit is a TAG-to-XAG mutation in exon 9 of a IDUA gene. In some embodiments, the method results in restoration of IDUA catalytic activity. In some embodiments, the disease is Hurler syndrome. In some embodiments, the engineered DNA pyrimidine modifying polypeptide targets a splicing site on a gene. In some embodiments, the disease is Duchene Muscular Dystrophy.

[0110] In some embodiments, the DNA pyrimidine base editing system has a first editing efficiency (such as measured by an editing percentage) in a first cell type and a second editing efficiency in a second cell type, wherein the first editing efficiency is different than the second editing efficiency. For example, in some embodiments, the nucleobase editor system may be configured to have cell type or tissue specificity, wherein the editing efficiency of the nucleobase editor system is higher in a targeted cell type or tissue and lower or not substantially occurring (such as an editing percentage of about 5%or less) in a different cell type or tissue. IV. Kits, medicines, and compositions

[0111] Provided herein, in certain aspects, are kits, medicines, and compositions of the DNA pyrimidine base-editing system taught herein.

[0112] In some embodiments, provided herein is a kit for a pyrimidine base-editing system, the kit comprising: a UNG variant described herein and, optionally, a DNA recognition domain, such as a UNG variant for thymine or cytosine editing. In some embodiments, provided herein is a kit for a pyrimidine base-editing system, the kit comprising one or more polynucleotides encoding a UNG variant described herein and, optionally, a DNA recognition domain, such as a UNG variant for thymine or cytosine editing. Kits provided herein may include one or more containers, and instruction for use thereof according to the methods provided herein. Instructions supplied in the kits of the invention are typically written instructions on a label or package insert (e.g., a paper sheet included in the kit) , but machine-readable instructions (e.g., instructions carried on a magnetic or optical storage disk) are also acceptable.

[0113] The kits provided herein are in suitable packaging. Suitable packaging include, but is not limited to, vials, bottles, jars, flexible packaging (e.g., sealed Mylar or plastic bags) , and the like. Kits may optionally provide additional components such as buffers and interpretative information. The present application thus also provides articles of manufacture, which include vials (such as sealed vials) , bottles, jars, flexible packaging, and the like.

[0114] Also provided are medicines, compositions, and unit dosage forms useful for the methods described herein. In some embodiments, the medicine, composition, or unit dosage form comprising a pyrimidine base-editing system described herein. For example, in some embodiments, the DNA pyrimidine base-editing system comprises a UNG variant described herein and, optionally, a DNA recognition domain, such as a UNG variant for thymine or cytosine editing. V. Sequences

[0115] Provided in this section are sequences referenced in the instant description. UNG (Homo sapiens) SEQ ID NO: 1 UNG Y147A (Homo sapiens; *engineered amino acid) SEQ ID NO: 2 UNG N204D (Homo sapiens; *engineered amino acid) SEQ ID NO: 3 UNG (Escherichia coli) SEQ ID NO: 4 UNG Y66A (Escherichia coli; *engineered amino acid) SEQ ID NO: 5 UNG N123D (Escherichia coli; *engineered amino acid) ; protein sequence) SEQ ID NO: 6 UNG (Deinococcus radiodurans) SEQ ID NO: 7 UNG Y85A (Deinococcus radiodurans; *engineered amino acid) SEQ ID NO: 8 UNG L80V, Y85A, K113E, K197E, S204A, E237Q, N* (Deinococcus radiodurans; *engineered amino acid) SEQ ID NO: 9 UNG Y85A, K113E, S204G (Deinococcus radiodurans; *engineered amino acid) SEQ ID  NO: 10 UNG Y85A, K113E, S204G (Deinococcus radiodurans; *engineered amino acid) SEQ ID  NO: 11 UNG Y85A, K113E, S204A, E237Q (Deinococcus radiodurans; *engineered amino acid) SEQ  ID NO: 12 UNG L80V, Y85A (Deinococcus radiodurans; *engineered amino acid) SEQ ID NO: 13 UNG Y85A, K113E, S204A, E237Q (Deinococcus radiodurans; *engineered amino acid) SEQ  ID NO: 14 UNG Y85A, K113E, K197E, E237Q (Deinococcus radiodurans; *engineered amino acid) SEQ  ID NO: 15 UNG (Human Herpesvirus 1) SEQ ID NO: 16 UNG Y90A (Human Herpesvirus 1; *engineered amino acid) SEQ ID NO: 17 UNG N147D (Human Herpesvirus 1; *engineered amino acid) SEQ ID NO: 18 UNG (Vaccinia virus) SEQ ID NO: 19 UNG Y70A (Vaccinia virus; *engineered amino acid) SEQ ID NO: 20 UNG N120D (Vaccinia virus; *engineered amino acid) SEQ ID NO: 21 Pre-edit of FIG. 1A; nucleotide sequence (3’ to 5’) ; *target thymine; SEQ ID NO: 22 Post-edit of FIG. 1A; nucleotide sequence (3’ to 5’) ; _excised base prior to repair; SEQ ID  NO: 23 GFP system protospacer of FIG. 1B; nucleotide sequence; SEQ ID NO: 24 Pre-edit of FIG. 2A; nucleotide sequence (3’ to 5’) ; *target cytosine; SEQ ID NO: 25 Post-edit of FIG. 2A; nucleotide sequence (3’ to 5’) ; _excised base prior to repair; SEQ ID  NO: 26 GGGGS repeat linker motif; SEQ ID NO: 27 3x GGGGS linker; SEQ ID NO: 28 UNG (Coxiella burnetiid) SEQ ID NO: 29

[0116] Those skilled in the art will recognize that several embodiments are possible within the scope and spirit of the disclosure of this application. The disclosure is illustrated further by the examples below, which are not to be construed as limiting the disclosure in scope or spirit to the specific procedures described therein. Additional sequences are disclosed within the Examples. EXAMPLES Materials and methods Plasmid construction

[0117] PCR was performed using PrimeSTAR GXL DNA polymerase (TaKaRa) or Q5 Hot Start High-Fidelity DNA Polymerase (NEB) . Wild-type hUNG (SEQ ID NO: 1) , EcUNG (SEQ ID NO: 4) , DrUNG (SEQ ID NO: 7) , HHV1_UNG (SEQ ID NO: 16) , VACV_UNG (SEQ ID NO: 19) and its variants, Cas9 and other genes were synthesized as gene blocks and codon optimized for mammalian expression (Tsingke Biotechnology Co., Ltd. ) . The mutation UNG fragment was inserted into the pCMV vector by Gibson assembly using Gibson Assembly Master Mix (NEB) . Individual sgRNA oligonucleotides (Table 1) were synthesized and cloned into pCG-2.0 sgRNA-expressing vector through Golden-Gate assembly. Ligated plasmids were transformed into Trans1-T1 chemically competent cells (TransGene Biotech) and subjected to Sanger sequencing to analyse the identity of the constructs (Rui Biotech) . Final plasmids were prepared (TianGen) for cell transfection. Table 1 Production and purification of mRNA

[0118] The production of mRNAs followed manufacturer’s instructions. Briefly, the mRNAs were produced using the commercial HiScribeTM T7 High Yield RNA Synthesis Kit (New England Biolabs) according to the manufacturer’s instructions with the linearized plasmids containing the T7 promotor, UNG mutation, nCas9 (D10A) and -225-nt polyA elements. Final IVT products were column purified and concentrated with the RNA Clean &Concentrator Kit (ZYMO Research) . Cell line construction

[0119] To create the dual-fluorescence reporter, mCherry and eGFP coding sequences (the ATG start codon of eGFP was deleted) were PCR amplified and digested using BsmBI (Thermo Fisher Scientific, no. ER0452) before being subjected to T4 DNA ligase (NEB, no. M0202L) -mediated ligation with 3 × GGGGS linkers (GGGGS SEQ ID NO: 27; and GGGGSGGGGSGGGGS SEQ ID NO: 28) . The ligation product was subsequently inserted into the pLenti-CMV-MCS-PURO backbone. To construct stable reporter cell lines, reporter constructs (pLenti-CMV-MCS-PURO backbone) were cotransfected into HEK293T cells together with two viral packaging plasmids, pR8.74 and pVSVG. After 72 hours, viral supernatant was collected and stored at -80 ℃. HEK293T cells were infected with lentivirus, and then mCherry+ cells were sorted via fluorescence-activated cell sorting (FACS) and cultured to select a single clone cell line stably expressing a dual-fluorescence reporter system with no detectable eGFP background. Cell culture and transfection

[0120] HEK293T, HCT116, NIH3T3, N2A and Cos-7 cell lines were maintained at Peking University. All the cells above and the dual-fluorescence reporter cells were cultured in Dulbecco’s modified Eagle’s medium (DMEM, Gibco) with 10%fetal bovine serum (Biological Industries) and penicillin / streptomycin (Sigma) at 37 ℃ with 5%CO2. GM06214 were from MEISEN CELL and cultured in Dulbecco’s modified Eagle’s medium (DMEM, Gibco) with 15%fetal bovine serum (Biological Industries) , 1%Non-Essential Amino Acid (NEAA) and penicillin / streptomycin (Sigma) at 37 ℃ with 5%CO2. For lipofection, cells were plated in 12-well cell culture plates at a density that approximately reached 70%after 20 hours. Cells in each well were transfected with 2, 250 ng of UNG mutation-nCas9 (D10A) and 2, 250 ng of sgRNA using 9 μL of PEI (ProteinTech) or transfected with 2, 500 ng of each UNG monomer mRNA using 5 μL of Lipofectamine MessengerMAX Reagent (Invitrogen) . Cells were collected after 72 hours of transfection. Genomic DNA was extracted using the DNeasy Blood &Tissue Kit (Qiagen) and stored at -20 ℃. UNG variants screening by fluorescence-activated cell sorting (FACS) analysis

[0121] To assess UNG variants editing efficiency with the dual-fluorescence reporter system, HEK293T reporter cells were plated in 12-well cell culture plates. After 72 hours of transfection, mCherry, BFP and EGFP fluorescence were analyzed by flow cytometer. The mCherry signal served as a fluorescent selection marker for reporter expressing cells, and BFP signal served as a fluorescent selection marker for sg RNA expressing cells. Percentages of eGFP+ / mCherry+ cells were calculated as the readout for editing efficiency. FACS data were analyzed with FlowJo X (v. 10.0.7) . For the mutant library, the NKK-scanned mutant library was utilized as a background and Error Prone PCR was performed (Agilent, GeneMorph II) . Subsequently, the PCR products were constructed into the pLenti vector using GoldenGate. After packaging into lentivirus, the virus was used to infect the reporter system cells, and FACS sorting was conducted two weeks post-infection. Targeted deep sequencing

[0122] Genomic sites of interest were amplified into fragments of approximately 200 bp from genomic DNA samples using PrimeSTAR GXL DNA polymerase (TaKaRa) . See Table 2 for the list of primers used. PCR products were purified using DNA Clean &Concentrator-25 (Zymo Research) for Sanger sequencing and targeted deep sequencing. Targeted deep sequencing libraries were prepared using the VAHTS Universal DNA Library Prep Kit for Illumina V3 (Vazyme) . Briefly, the PCR fragments were sequentially subjected to end repair, adapter ligation, and then PCR amplification. DNA purification in library preparation was performed using Agencourt Ampure XP beads (Beckman Coulter) , and library amplification was performed using Q5U Hot Start High-Fidelity DNA Polymerase (NEB) and VAHTS Multiplex Oligos Set 4 / 5 for Illumina (Vazyme) . The final library was subjected to quantification using the Qubit dsDNA HS assay kit (Invitrogen) and sequenced using Illumina HiSeq X Ten. Table 2 Analysis of high-throughput sequencing data for targeted amplicon sequencing

[0123] For high-throughput sequencing data analysis, an index was generated using the targeted site sequences (upstream and downstream ~100 nt) of the protospacer. All data were analyzed by CRISPResso2 follow parameter --quantification_window_size 10 --quantification_window_center -10 --base_editor_output. The mutations that appeared in the control and experimental groups simultaneously were due to single nucleotide polymorphisms. Cytotoxicity assay

[0124] HEK293T cells were seeded in 96-well plates (Corning) at 2×104 cells per well in 200 μl of complete growth medium. 24 hours after seeding, cells were transfected with 1 μl PEI (ProteinTech) and 250 ng sgRNA and 250 ng editors. 0 h, 24 h, 48 h, and 72 h after transfection, 20 μl CCK8 solution was added to each well and the absorbance at 450 nm was measured after 1.5 hours using a microplate reader. Electroporation in primary cells

[0125] For mRNA electroporation in GM06214 cells, 1.5 μg sgRNA and 4.5 μg DrUNG-nCas9 mRNA were electroporated with Nucleofector 2b Device (Lonza) and Basic Nucleofector Kit (Lonza) , and the electroporation program was U-012. Then the cells were transferred to warm culture medium for use in one or more assays. IDUA catalytic activity assay The gathered cell pellet was resuspended and lysed with 28 μl 0.5%Triton X-100 in 1 × PBS  buffer on ice for 30 min. Then 25 μl of the cell lysis was added to 25 μl 190 μM 4-methylumbelliferyl-α-liduronidase substrate (Cayman) , which was dissolved in 0.4 M sodium formate buffer containing 0.2%Triton X-100, pH 3.5 and incubated for 90 min at 37 ℃ in the dark. The catalytic reaction was quenched by adding 200 μl 0.5 M NaOH / Glycine buffer, pH 10.3 and then centrifuged for 2 min at 4 ℃. The supernatant was transferred to a 96-well plate, and fluorescence was measured at 365 nm excitation wavelength and 450 nm emission wavelength with Infinite M200 reader (TECAN) . sgRNA-dependent DNA off-target sequencing

[0126] Cas-OFFinder (CRISPR RGEN Tools (rgenome. net) ) used for prediction of potential off-target sites of Cas9 RNA-guided endonucleases, the top 10 off-target sites were selected for validation. Ten off-target sites of each target site were amplified from genomic DNA prepared and sequenced on Hi-TOM NGS platform. Genome-wide DNA off-target sequencing

[0127] We input 500 to 1,000 ng of genomic DNA for library preparation using the VAHTS Universal Plus DNA Library Prep Kit for Illumina (Vazyme) . The process of library preparation was as follows: fragmentation, end preparation and dA-tailing, adapter ligation and library amplification. Among them, 500 to 1000 ng of genomic DNA was fragmented with FEA enzyme mix at 37 ℃ for 10 min, and end repair and dA-tailing were simultaneously completed in the process. The final library was subjected to quantification using the Qubit dsDNA HS assay kit (Invitrogen) and fragment analyzer. All libraries were finally sequenced using Illumina HiSeq X Ten (Illumina) . Transcriptome-wide RNA off-target sequencing

[0128] The control eGFP-expressing or TBE-expressing plasmid with sgRNA 584 18 were transfected into HEK293T cells. 72h after transfection, RNAs were purified with Direct-zol RNA Miniprep Kits (Zymo Research) . The mRNA was then purified using NEBNext Poly (A) mRNA Magnetic Isolation Module (New England Biolabs) , processed with the NEBNext Ultra II RNA Library Prep Kit for Illumina (New England Biolabs) , followed by deep sequencing analysis using Illumina HiSeq X Ten platform. Analysis of nuclear genome off-target editing

[0129] Whole-genome sequencing reads underwent quality control using FastQC (v0.12.1) , and adapters were removed using fastp (0.23.2) . Following trimming, reads were aligned to the GRCh38-hg38 reference genome using bwa-mem2 (2.2.1) with default parameters. Subsequently, the picard AddOrReplaceReadGroups (v2.25.5) , MarkDuplicatesSpark, BaseRecalibratorSpark, and ApplyBQSRSpark were applied to add read group information, remove duplicates, and correct base quality. After preprocessing, GATK Mutect2 was utilized to identify somatic short variants. Variant calls were filtered based on the FilterMutectCalls criteria, excluding positions annotated as position, slippage, weak evidence, or low mapping quality. Additionally, mutations with a frequency exceeding 1%in control experiments were excluded. Only mutations at positions where the reference genome contained a T and the mutated allele was A / C / G were retained. To identify potential off-target genome editing events, stringent criteria were applied to mitigate high noise levels. Additional requirements for base quality and mapping quality were imposed based on quality control 611 criteria. Only mutations with a high median base quality (MBQ ≥ 32) and high mapping quality (MMQ ≥ 50) were considered potential off-target editing sites. The Mann-Whitney U test was employed to assess the significance of differences in mutation frequency between each experimental group and the control group (p-value < 0.1) . Example 1: Engineered uracil-DNA glycosylase (UNG) variants for editing of thymine

[0130] This example demonstrates the development and use of engineered uracil-DNA glycosylases (UNG) variants and UNG variant / nCas9 fusions to achieve base editing of thymine during translesion DNA synthesis.

[0131] Specifically, mutations were introduced to amino acids of UNG which influence the native binding pocket size for uracil, enabling the entry of thymine into the active site pocket and facilitating programmable thymine base editing (FIG. 1A) .

[0132] Saturation mutagenesis was conducted on several amino acids, including G143-D145, Y147, and F158, which could affect the entry of thymine into the active site pocket. Y147 was identified as hindering the entry of the methyl group at the 5th position of thymine. By fusing the mutated form of human UNG (SEQ ID NO: 2) at the N-terminal of nCas9 (D10A) and targeting a previously developed reporter system (FIG. 1B) , editing efficiency was assessed based on the eGFP ratio.

[0133] Mutating Tyr147 of human UNG (SEQ ID NO: 2) to amino acids with smaller side chains, such as Ala, Cys, Gly, or Ser, led to a detectable but low eGFP-positive ratio (FIG. 1C) . This suggested that amino acids with smaller side chains may reduce obstacles for thymine to enter the active site pocket.

[0134] Through targeted sequencing of Y147A mutants, the Y147A mutant exhibited 0.4%editing efficiency, mainly resulting in T-to-C conversion (FIG. 1D) . These results indicated that direct editing of thymine can be achieved through engineered of UNG variants described herein. Example 2: Engineered UNG variant for editing of cytosine

[0135] This example demonstrates the development and use of an engineered uracil-DNA glycosylases (UNG) variant and UNG variant / nCas9 fusions to achieve base editing of cytosine during translesion DNA synthesis.

[0136] In this example, mutations were introduced to amino acids which influence the binding pocket size of uracil, enabling the entry of cytosine into the active site pocket and facilitating programmable cytosine base editing (FIG. 2A) .

[0137] Amino acid residues in the hUNG binding pocket were mutagenized as described in Example 1. N204 affected the entry of cytosine into the active site pocket. The proportion of eGFP-positive cells reached 30%only when Asn204 was mutated to Asp (FIG. 2B) . Through targeted sequencing of N204D mutants, N204D mutant (SEQ ID NO: 3) showed the highest editing efficiency at 30%, which predominantly resulted in C-to-G or C-to-T conversion (FIG. 2C) . These results indicated that direct editing of cytosine can be achieved through engineering of UNG variants described herein. Example 3: Engineered UNG derived from various species for editing of thymine and  cytosine

[0138] This example demonstrates the development and use of non-human engineered uracil-DNA glycosylases (UNG) variants and UNG variant / nCas9 fusions to achieve base editing of thymine or cytosine during translesion DNA synthesis.

[0139] Employing protein sequence alignment, UNGs were specifically selected from Escherichia coli (Ec) (SEQ ID NO: 4) , Deinococcus radiodurans (Dr) (SEQ ID NO: 7) , Human Herpesvirus 1 (HHV1) (SEQ ID NO: 16) , and Vaccinia virus (VACV) (SEQ ID NO: 19) . Utilizing homologous protein sequence alignment, the corresponding amino acids in these UNG variants analogous to N204 and Y147 of human UNG were identified. Mutations were introduced to these amino acids, changing them to D and A, respectively, to assess their efficacy in excising cytidine and thymine.

[0140] EcUNG (N123D) (SEQ ID NO: 6) and HHV1_UNG (N147D) (SEQ ID NO: 18) achieved editing efficiency on par with hUNG (N204D) (SEQ ID NO: 3) (see Example 2) for cytidine (FIG. 3A) . For thymine editing, DrUNG (Y85A) (SEQ ID NO: 8) exhibited the highest efficiency, nearly five times greater than that of hUNG (Y147A) (SEQ ID NO: 2) (FIG. 3B) . Example 4: Enhancement of DrUNG thymine editing efficiency

[0141] This example demonstrates the development and use of optimized UNG variants with high thymine-editing efficiency.

[0142] To enhance the activity of DrUNG variants, a library of DrUNG-Y85A (SEQ ID NO: 1) (SEQ ID NO: 8) variants was generated. Subsequently, the targeting reporter system was introduced using nCas9 and sgRNA, the top 1%of eGFP+ cells were sorted through FACS, and then the sequences of DrUNG (Y85A) mutants were amplified (FIG. 4A) . Of the resulting mutations, substitutions at S204, Y85, E237, and K113 were most common (FIG. 4B) . In total, 25 mutants were identified. The 6th mutant variant (SEQ ID NO: 9) exhibited the highest editing efficiency, followed by the 2nd mutant variant (SEQ ID NO: 10) (FIG. 4C) . These two variants were designated DrUNG mutant 6 (SEQ ID NO: 9, also referred to as “thymine base editor” [TBE] ) and DrUNG mutant 2 (SEQ ID NO: 10) , respectively. DrUNG mutant 6 (SEQ ID NO: 9) contains mutations L80V, Y85A, K113E, K197E, S204A, E237Q, and an N-terminal deletion of two amino acids. DrUNG mutant 2 contains mutations Y85A, K113E, and S204G (SEQ ID NO: 10) . Compared to DrUNG-Y85A (SEQ ID NO: 8) , both DrUNG mutant 6 (SEQ ID NO: 9) and DrUNG mutant 2 (SEQ ID NO: 10) demonstrated a 5.4-fold and 2.2-fold increase in editing efficiency, respectively. Additionally, when compared to hUNG (Y147A) (SEQ ID NO: 2) , these two mutants exhibited a 27-fold and 11-fold increase in editing efficiency, respectively (FIG. 4B) . Example 5: Use of DrUNG to edit host cells

[0143] This example demonstrates the use of engineered UNG variants to edit thymine at endogenous sites in a host cell.

[0144] DrUNG mutant 6 (SEQ ID NO: 9) and DrUNG (SEQ ID NO: 10) mutant 2 were utilized to target 14 endogenous sites in HEK293T cells. Notably, DrUNG mutant 6 (SEQ ID NO: 9) exhibited effective editing at all 16 targeted sites, whereas DrUNG mutant 2 (SEQ ID NO: 10) exhibited effective editing at only 9 sites. The peak editing efficiency for DrUNG mutant 6 (SEQ ID NO: 9) reached 40%, and for DrUNG mutant 2 (SEQ ID NO: 10) it reached 33% (FIG. 5A) . Concurrent with targeted editing, the thymine base editor induced indels at the targeted sites, with the highest indel rate being 10% (FIG. 5A) . Of the indels caused by the thymine base editor based on DrUNG, 76%were deletions and 24%were insertions.

[0145] The average editing efficiency of the base editor mediated by DrUNG mutant 6 (SEQ ID NO: 9) was 20%, marking a substantial 3-fold increase compared to the editing efficiency of DrUNG mutant 2 (SEQ ID NO: 10) (FIG. 5B) . In the editing window of the thymine base editor, approximately 52%of thymine converted to cytosine, around 30%transformed into guanine, and approximately 18%changed to adenosine (FIG. 5C) . Notably, there are variations in the distribution of thymine conversion among different sites. Upon analyzing the editing positions, it became apparent that the majority of edits occurred at positions 14 to 18 away from the NGG PAM, exhibiting a similar distribution pattern to other types of CRISPR-based base editors (FIG. 5D) . Different linkers between DrUNG mutant 6 and nCas9 (D10A) were also tested. It was found that using a 64-amino acid linker on the reporter system was the most efficient, while using GS, SGGS, PAPAP, XTEN, 3×EAAAK, and 32-amino acid linkers showed similar editing efficiency (FIG. 5E) . Using XTEN, 3×EAAAK, 32-amino acid, and 64-amino acid linkers show higher editing efficiency compared to GS, SGGS and PAPAP linkers on multiple endogenous sites (FIG. 5F) . Further, the DrUNG mutant 6 base editor showed strong editing activity and low indel levels in a variety of cell lines on reporter system, including human cell line HCT116, mouse cell line N2A, monkey cell lines NIH3T3 and Cos-7, among which N2A and NIH3T3 cell lines achieved approximately 70%reporter system lighting efficiency (FIG. 5G and 5H) . To conclude, the thymine base editors we developed exhibit effective editing in multiple cell lines. Example 6: High editing specificity of DrUNG mutant 6 at genome and transcriptome level

[0146] The editing specificity of DrUNG mutant 6 was evaluated on a genome-wide and transcriptome scale. Through genome-wide high-throughput sequencing detection, it was determined that DrUNG mutant 6 transfection group showed some off-targets compared to the control group (FIG. 6A) . The 50 bp upstream and downstream of these off-target sites were removed, and it was determined there were no potential sgRNA binding sites. Furthermore, Cas-OFFinder was used to predict the potential off-target sites of four sgRNAs on the genome, and the 10 highest-scoring off-target sites were selected for sequencing. None of these off-target sites had been edited (FIG. 6C) . This shows that the off-target sites may be random off-target caused by DrUNG mutant 6 (FIG. 6A) . In addition, significant off-target sites were not detected at the transcriptome level (FIG. 3B) . In summary, DrUNG mutant 6 will cause some random off-target sites on the genome but will not cause off-target on the transcriptome, which shows that DrUNG mutant 6 is a relatively safe thymine base editor. Example 7: Significant restoration of Hurler syndrome disease cell phenotypes using  DrUNG mutant 6

[0147] DrUNG mutant 6 is a tool that can be used in the treatment of many mutation-related diseases, such as Hurler syndrome. Hurler syndrome is caused by the presence of a premature stop codon (resulting from a TGG-to-TAG mutation) on exon 9 of IDUA gene, which prevents the production of functional IDUA protein. Thus, performing a TAG-to-XAG mutation can potentially restore IDUA activity. Three different Cas9 proteins were selected and corresponding sgRNA were designed. By transfecting cells of the previously reported Hurler syndrome disease premature stop codon reporter system (FIG. 7A) , it was observed that all three Cas9 proteins fused to DrUNG mutant 6 can achieve thymine base editing, among which spCas9 protein has the highest editing efficiency (FIG. 7B) . Subsequently, the mRNA of the spCas9 and DrUNG mutant 6 fusion protein and sgRNA was co-transfected into GM06214 (IDUAW402X) cells derived from patient with Hurler syndrome, which contained the IDUAW402X mutation (FIG. 7C) , achieving an editing efficiency of approximately 25%at the targeted site (FIG. 7D) and significantly restoring the IDUA catalytic activity of the cells (FIG.7E) .

Claims

1.An engineered thymine-modifying polypeptide comprising a uracil-DNA glycosylase (UNG) variant comprising:(a) an amino acid substitution selected from a group consisting of Y85A, Y85S, Y85N, Y85G, and Y85C; and(b) one or more amino acid substitutions of L80V, K113E, K197E, S204A, S204G, or E237Q,wherein the positions of amino acid substitutions are in reference to the amino acid sequence set forth in SEQ ID NO: 7.2.The engineered thymine-modifying polypeptide of claim 1, comprising the amino acid substitution of Y85A.3.The engineered thymine-modifying polypeptide of claim 1 or 2, comprising the amino acid substitutions of L80V, K113E, K197E, S204A, and E237Q.4.The engineered thymine-modifying polypeptide of claim 1 or 2, comprising the amino acid substitutions of K113E and S204G.5.The engineered thymine-modifying polypeptide of claim 1 or 2, comprising the amino acid substitutions of K113E, S204A, and E237Q.6.The engineered thymine-modifying polypeptide of claim 1 or 2, comprising the amino acid substitution of L80V.7.The engineered thymine-modifying polypeptide of claim 1 or 2, comprising the amino acid substitutions of K197E, S204A, and E237Q.8.The engineered thymine-modifying polypeptide of claim 1 or 2, comprising the amino acid substitutions of K113E, K197E, and E237Q.9.The engineered thymine-modifying polypeptide of any one of claims 1-8, wherein the UNG is derived from an organism selected from a group consisting of Deinococcus radiodurans, Homo sapiens, Escherichia coli, Human Herpesvirus 1, and Vaccinia virus.10.The engineered thymine-modifying polypeptide of claim 9, wherein the UNG is derived from Deinococcus radiodurans.11.An engineered thymine-modifying polypeptide comprising a uracil-DNA glycosylase (UNG) variant comprising an amino acid substitution comprising Y85A, wherein the UNG variant is derived from Deinococcus radiodurans.12.The engineered thymine-modifying polypeptide of claim 11, further comprising one or more amino acid substitutions of L80V, K113E, K197E, S204A, S204G, or E237Q, wherein the positions of amino acid substitutions are in reference to the amino acid sequence set forth in SEQ ID NO: 7.13.The engineered thymine-modifying polypeptide of any one of claims 1-12, wherein the UNG variant has at least about 70%sequence identity to SEQ ID NO: 7.14.The engineered thymine-modifying polypeptide of any one of claims 1-13, wherein the UNG variant comprises an N-terminal deletion relative to the corresponding UNG from which the UNG variant was derived.15.The engineered thymine-modifying polypeptide of claim 14, wherein the N-terminal deletion comprises a deletion of up to the first 23 amino acids.16.The engineered thymine-modifying polypeptide of claim 14 or 15, wherein the N-terminal deletion comprises a deletion of the first two amino acids.17.The engineered thymine-modifying polypeptide of any one of claims 1-16, further comprising a DNA recognition domain.18.The engineered thymine-modifying polypeptide of claim 17, wherein the DNA recognition domain is fused to the UNG variant.19.The engineered thymine-modifying polypeptide of claim 17 or 18, wherein the DNA recognition domain is fused to the C-terminus of the UNG variant.20.The engineered thymine-modifying polypeptide of claim 17, wherein the DNA recognition domain is operably connected to the UNG variant by a linker domain.21.The engineered thymine-modifying polypeptide of claim 20, wherein the linker domain is a 32-amino acid linker or a 64-amino acid linker, and is selected from a list composed of GS, SGGS, PAPAP, XTEN, or repeats thereof.22.The engineered thymine-modifying polypeptide of any one of claims 17-21, wherein the DNA recognition domain comprises a Cas nuclease.23.The engineered thymine-modifying polypeptide of claim 20, wherein the Cas nuclease is an nCas9.24.The engineered thymine-modifying polypeptide of claim 23, wherein the nCas9 comprises a D10A mutation.25.The engineered thymine-modifying polypeptide of claim 23 or 24, wherein the nCas9 is derived from Streptococcus pyogenes.26.The engineered thymine-modifying polypeptide of claim 22, wherein the Cas nuclease is a dCas9.27.The engineered thymine-modifying polypeptide of any one of claims 1-26, wherein the engineered thymine-modifying polypeptide does not comprise a deaminase domain having deamination activity.28.An engineered thymine-modifying polypeptide comprising SEQ ID NO: 9.29.An engineered thymine-modifying polypeptide comprising SEQ ID NO: 10.30.An engineered thymine-modifying polypeptide comprising SEQ ID NO: 12.31.An engineered thymine-modifying polypeptide comprising SEQ ID NO: 13.32.An engineered thymine-modifying polypeptide comprising SEQ ID NO: 14.33.An engineered thymine-modifying polypeptide comprising SEQ ID NO: 15.34.The engineered thymine-modifying polypeptide of any one of claims 28-33, further comprising a DNA recognition domain.35.A nucleic acid encoding an engineered thymine-modifying polypeptide of any one of claims 1-34.36.A thymine editing system comprising an engineered thymine-modifying polypeptide of any one of claims 1-34, or a nucleic acid of claim 35.37.The thymine editing system of claim 36, further comprising an sgRNA.38.The thymine editing system of claim 37, wherein the sgRNA targets a protospacer comprising a target thymine.39.A method of modifying a target thymine in a nucleic acid sequence present in one or more nucleic acid molecules, the method comprising contacting at least one nucleic acid molecule of the one or more nucleic acid molecules with an engineered thymine-modifying polypeptide of any one of claims 17-27 or 34, wherein the DNA recognition domain is configured to associate with the at least one nucleic acid molecule such that the UNG variant is positioned to modify the target thymine.40.The method of claim 39, wherein the one or more nucleic acid molecules are in a cell.41.The method of claim 39 or 40, wherein the one or more nucleic acid molecules are in a mitochondrion.42.The method of any one of claims 39-41, wherein the target thymine undergoes a T-to-C, T-to-G, or T-to-A modification.43.The method of claim 42, wherein about 15%to about 75%of the modified target thymine are modified to a cytosine.44.The method of claim 42, wherein about 1%to about 65%of the modified target thymine are modified to a guanine.45.The method of claim 42, wherein about 5%to about 50%of the modified target thymine are modified to an alanine.46.The method of any one of claims 42-45, wherein an NGG PAM site is 14 to 19 base pairs away from the target thymine.47.The method of any one of claims 42-45, wherein an NG PAM site is 14 to 19 base pairs away from the target thymine.48.The method of any one of claims 37-45, wherein the method exhibits an editing efficiency at least 2-fold greater than that of a method comprising contacting a nucleic acid molecule comprising a target thymine with a wild-type UNG.49.The method of any one of claims 39-48, wherein the method exhibits an editing efficiency of about 6%to about 40%.50.A method of treating a disease associated with a DNA mutation in an individual, the method comprising contacting a nucleic acid comprising a target thymine associated with a disease with the thymine editing system of any one of claims 36-38.51.The method of treatment of claim 50, wherein the contacting reduces the amount of a premature stop codon.52.The method of treatment of claim 51, wherein the contacting results in the formation of a TAG-to-XAG mutation in exon 9 of a IDUA gene.53.The method of treatment of any one of claims 50-52, wherein the administration of the thymine editing system results in restoration of IDUA catalytic activity.54.The method of treatment of any one of claims 50-53, wherein the disease is Hurler syndrome.55.A method of treating Hurler syndrome in an individual in need thereof, comprising contacting a nucleic acid comprising a target thymine in a mutation associated with Hurler syndrome with the thymine editing system of any one of claims 36-38.56.The method of treatment of claim 50, wherein the thymine editing system targets a splicing site on a gene.57.The method of treatment of claim 56, wherein the disease is Duchene Muscular Dystrophy.

Citation Information

Patent Citations

  • Novel DNA glycosylases and their use

    US20060134631A1

  • Method for modifying target site in double-stranded DNA in cell

    US20200270631A1

  • T:a to a:t base editing through adenine excision

    WO2020181195A1

  • Novel base editors and uses thereof

    WO2024222812A1