Novel base editor and use thereof

By developing a fusion protein that combines a nucleic acid programmable DNA binding domain and a base excision domain, the problem that existing base editors cannot directly edit dG or dT has been solved, enabling direct base editing of dG and dT and expanding the application scope of base editing.

CN120936712APending Publication Date: 2025-11-11HUIGENE THERAPEUTICS CO LTD +1

Patent Information

Application Number
CN202480020620.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2024-04-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing base editors cannot directly edit deoxyribonucleotides such as dG or dT, and cannot achieve direct base editing of G or T.

Method used

A fusion protein has been developed, comprising a nucleic acid programmable DNA binding domain and a base excision domain, which can directly edit the dG and dT in the target dsDNA without deamination by binding to the prototype spacer sequence on the non-target strand of the target dsDNA and excising the corresponding bases.

Benefits of technology

It enables direct base editing of dG and dT, surpassing the limitations of existing base editors and providing a wider range of editing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005605321870000741
    Figure BDA0005605321870000741
  • Figure BDA0005605321870000751
    Figure BDA0005605321870000751
  • Figure BDA0005605321870000752
    Figure BDA0005605321870000752
Patent Text Reader

Abstract

The present disclosure provides a novel base editor and uses thereof.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority and benefit from the following applications filed on the following dates: PCT / CN2023 / 090660, filed April 25, 2023; PCT / CN2023 / 091734, filed April 28, 2023; PCT / CN2023 / 094565, filed May 16, 2023; PCT / CN2024 / 070217, filed January 2, 2024; and PCT / CN2024 / 084498, filed March 28, 2024, the entire contents of which (including any figures and sequence listings) are incorporated herein by reference.

[0003] Reference to electronic sequence listing

[0004] This disclosure contains a sequence list XML file that has been electronically submitted in XML format and is hereby incorporated by reference in its entirety. A copy of the XML created on April 25, 2024, using the software “WIPOSequence” in accordance with WIPO Standard ST.26 is named HGP032PCT.xml and has a size of 2,772,603 ​​bytes.

[0005] According to WIPO standard ST.26, the symbol “t” is used to represent both T in DNA and U in RNA. Therefore, in this sequence listing prepared according to ST.26, in any case where the sequence is RNA, T in that sequence should be considered as U. Background Technology

[0006] Base editing is a powerful technology for basic research and therapeutic applications [1, 2]. Current base editors mainly contain nucleic acid-programmable DNA-binding proteins, such as catalytically damaged CRISPR-associated (Cas) nucleases, which fuse with single-stranded DNA deaminases, and sometimes with other proteins that can regulate DNA repair mechanisms [3, 4]. In the past few years, two main classes of DNA base editors have been developed, namely adenine base editors (ABE) [5] and cytosine base editors (CBE) [4], and have been widely used for A-to-G and C-to-T conversions, respectively. Figure 1 a). Recently, C-to-G base editors (CGBE) [6-10] and adenine transversion base editors (AYBE)

[11] have been constructed by fusing existing CBEs or ABEs with DNA glycosylation enzyme variants to produce new tools for achieving more diverse base editing outcomes, including C-to-G, A-to-C, and A-to-T editing. Figure 9A CRISPR-free CBE (DdCBE) has been reported for C-to-T base editing in mitochondrial DNA by fusing two halves of a double-stranded DNA cytidine deaminase (DddA) variant with two separate TALE (transcription activator-like effector) proteins [12-14]. To date, these base editing methods have all begun with the deamination of C or A as a necessary step to generate uridine (U) or inosine (I) intermediates, which are then converted to another base via endogenous DNA repair or replication mechanisms [4-11]. Although G and T in the non-edited strand can be modified in conjunction with the editing of C and A in the edited strand, respectively, no existing base editor can directly edit G or T.

[0007] The goal is to develop novel base editors and base editing methods that surpass current base editors for editing deoxyribonucleotides (such as dG or dT).

[0008] References to or identification of any document in this disclosure are not an admission that such document is available as prior art. Each reference mentioned or cited in this disclosure is incorporated herein by reference in its entirety. Summary of the Invention

[0009] At least a portion of the information provided in this disclosure includes base editors and base editing methods capable of directly editing target deoxyribonucleotides (e.g., dG, dT) in target dsDNA. At least a portion of the information provided in this disclosure includes base editors and base editing methods capable of editing target deoxyribonucleotides (e.g., dC) in target dsDNA without the presence of deamination.

[0010] In one aspect, this disclosure provides a fusion protein comprising:

[0011] (1) A nucleic acid programmable DNA-binding domain (napDNAbd) capable of binding to target dsDNA, wherein the target dsDNA comprises:

[0012] (a) The first deoxyribonucleotide (e.g., dG (deoxyguanosine), dT (thymidine), dC (deoxycytidine)) in the prototype spacer sequence on the non-target strand (editing strand) of the target dsDNA, and

[0013] (b) a second deoxyribonucleotide (e.g., dC (deoxycytidine), dA (deoxyadenosine), dG (deoxyguanosine)) in a target sequence paired with the first deoxyribonucleotide (e.g., dG, dT, dC) and located on the target strand (non-edited strand) of the target dsDNA, wherein the prototype spacer sequence is completely inversely complementary to the target sequence; and

[0014] (2) A base-removal domain capable of removing a base of the first deoxyribonucleotide (e.g., guanine, thymine, cytosine).

[0015] In some embodiments, the fusion protein does not contain a deaminase domain, such as an adenine or cytosine deaminase domain, such as TadA and its variants.

[0016] In some embodiments, the first deoxyribonucleotide is deoxyguanosine (dG), thymidine (dT), or deoxycytidine (dC).

[0017] In some embodiments, the conversion of the first deoxyribonucleotide to the fourth deoxyribonucleotide is dG to dA, dG to dT, dG to dC, dT to dA, dT to dC, dT to dG, dC to dA, dC to dT, or dC to dG.

[0018] In some embodiments, the base-removing domain comprises a glycosylation enzyme.

[0019] In some embodiments, the glycosylation enzyme is selected from N-methylpurine DNA glycosylation enzyme (MPG), 8-oxoguanine DNA glycosylation enzyme (OGG1), methyl-CpG-binding domain 4 DNA glycosylation enzyme (MBD4), thymine DNA glycosylation enzyme (TDG), uracil DNA glycosylation enzyme (UNG), single-stranded selective monofunctional uracil DNA glycosylation enzyme 1 (SMUG1), mutY DNA glycosylation enzyme (MUTYH), nth-like DNA glycosylation enzyme 1 (NTHL1), nei-like DNA glycosylation enzyme 1 (NEIL1), nei-like DNA glycosylation enzyme 2 (NEIL2), nei-like DNA glycosylation enzyme 3 (NEIL3), and mutants thereof capable of recognizing and excising bases from nucleotides of nucleic acids.

[0020] In some embodiments, the base-excision domain comprises N-methylpurine DNA glycosylation enzyme (MPG).

[0021] In some embodiments, the MPG contains amino acid substitutions at positions corresponding to or selected from the following positions: N169, D175, C178 and / or Q294 of the reference MPG, wherein the positions are numbered according to SEQ ID NO:1.

[0022] In some embodiments, the amino acid substitution is performed using R, A, N, or G.

[0023] In some embodiments, the MPG contains an amino acid substitution relative to the reference MPG of SEQ ID NO:7, the amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following substitutions: N169G, D175R, C178N, Q294R, wherein the position is numbered according to SEQ ID NO:1.

[0024] In some embodiments, the MPG contains a combination substitution relative to the reference MPG of SEQ ID NO:7, the combination substitution corresponding to the combination substitution of N169G, D175R, C178N and Q294R, wherein the positions are numbered according to SEQ ID NO:1.

[0025] In some embodiments, the MPG comprises, is substantially composed of, or consists of an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO: 8, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38) sequence identity.

[0026] In some implementations, the MPG is (substantially) capable of cleaving guanine from dG.

[0027] In some embodiments, the base-excision domain comprises uracil DNA glycosylase (UNG).

[0028] In some embodiments, the UNG contains an amino acid substitution at a position corresponding to a position selected from or selected from the following positions: K184, A214, Q259 and / or Y284 of the reference UNG, wherein the position is numbered according to SEQ ID NO:133.

[0029] In some embodiments, the amino acid substitution is performed using A, D, V, or T.

[0030] In some embodiments, the UNG relative to the reference UNG of SEQ ID NO:135 or 137 comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following substitutions: K184A, A214V, A214T, Q259A, Y284D, and the positions thereof, wherein the positions are numbered according to SEQ ID NO:133.

[0031] In some embodiments, the UNG comprises the deletion of amino acids at positions corresponding to or corresponding to positions 1-65, 1-66, 1-67, 1-68, 1-69, 1-70, 1-71, 1-72, 1-73, 1-74, 1-75, 1-76, 1-77, 1-78, 1-79, 1-80, 1-81, 1-82, 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, or 1-100 of the reference UNG of SEQ ID NO: 133.

[0032] In some embodiments, the UNG, relative to the reference UNG of SEQ ID NO:135, comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following substitutions: A214T, Q259A, Y284D, and the like, and the UNG comprises the deletion of an amino acid at a position corresponding to or a position of positions 1-88 of the reference UNG, wherein the positions are numbered according to SEQ ID NO:133.

[0033] In some embodiments, the UNG, relative to the reference UNG of SEQ ID NO:137, comprises an amino acid substitution corresponding to a substitution selected from or a combination of K184A, A214V, and the two substitutions, and the UNG comprises a deletion of an amino acid at a position corresponding to or a position of positions 1-88 of the reference UNG, wherein the positions are numbered according to SEQ ID NO:133.

[0034] In some embodiments, the UNG containing the amino acid mutation comprises, substantially consists of, or consists of: (with SEQ ID) The amino acid sequence of NO:56, 58, 60, 62, 135, 137, 139, 141, 144, 146, 148, 150, 152, 155, 157, and 159, or the N-terminal truncated amino acid sequence lacking the most N-terminal methionine (M) (encoded by the start codon ATG), has at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity.

[0035] In some implementations, the UNG is (substantially) capable of cleaving thymine from dT.

[0036] In some implementations, the UNG is (substantially) capable of cleaving dC cytosine.

[0037] In some embodiments, the napDNAbd is an RNA-programmable DNA-binding protein.

[0038] In some embodiments, the napDNAbd is selected from CRISPR-associated (Cas) proteins, IscB, IsrB, Argonaute, and TnpB.

[0039] In some embodiments, the napDNAbd is a nicking enzyme, such as Cas9 nicking enzyme or IscB nicking enzyme.

[0040] In some embodiments, the napDNAbd is nuclease-free, for example, inactivated Cas9 or inactivated Cas12i.

[0041] In some embodiments, the napDNAbd comprises an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO:2, 48, 50, 52, or 163.

[0042] In some embodiments, the fusion protein comprises (1) the napDNAbp and the base excision domain from the N-terminus to the C-terminus; or (2) the base excision domain and the napDNAbp.

[0043] In some embodiments, the napDNAbd (e.g., Cas9) is a two-part napDNAbd comprising an N-terminal portion and a C-terminal portion, such as a two-part splitting Cas9, and wherein the fusion protein comprises from the N-terminus to the C-terminus (1) the N-terminal portion of the napDNAbd, the base-removing domain, and the C-terminal portion of the napDNAbd; (2) the C-terminal portion of the napDNAbd, the base-removing domain, and the N-terminal portion of the napDNAbd; or (3) the base-removing domain, the C-terminal portion of the napDNAbd (e.g., amino acids at positions 1249-1368), and the N-terminal portion of the napDNAbd (e.g., amino acids at positions 1-1248).

[0044] In some embodiments, the napDNAbd is SpCas9 (e.g., SpCas9 nickase) or a mutant thereof (e.g., SpG Cas9 nickase).

[0045] In some embodiments, the N-terminal portion of the napDNAbd is the amino acid at position 1 or 2 to 1012, 1028, 1041, 1046, 1047, 1248, 1249, or 1300 of the napDNAbp.

[0046] In some embodiments, the C-terminal portion of the napDNAbd is the amino acid at positions 1013, 1029, 1042, 1047, 1048, 1249, 1063, 1064, 1230, 1249 or 1301 to 1368 of the napDNAbp.

[0047] In some embodiments, the fusion protein includes the base-removing domain, which is embedded between positions 2-1248 and 1249-1368 of nCas9 (SEQ ID NO:2), wherein the first amino acid residue D of nCas9 (SEQ ID NO:2) is designated as position 2; or embedded between positions 2-1047 and 1064-1368 of nCas9 (SEQ ID NO:2), wherein the first amino acid residue D of nCas9 (SEQ ID NO:2) is designated as position 2.

[0048] In some embodiments, the fusion protein comprises having at least about 60 ppm of any one of SEQ ID NO: 12, 14, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 55, 57, 59, 61, 63, 136, 138, 140, 142, 143, 145, 147, 149, 151, 153, 154, 156, 158, 160, 161, 162, and 164. The amino acid sequence is of % (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity.

[0049] On the other hand, this disclosure provides a system comprising:

[0050] (i) the fusion protein of this disclosure or the polynucleotide encoding said fusion protein; and

[0051] (ii) a guiding nucleic acid or a polynucleotide encoding the guiding nucleic acid, the guiding nucleic acid comprising:

[0052] (1) A scaffold sequence capable of forming a complex with the napDNAbd; and

[0053] (2) It is capable of hybridizing with the target sequence on the target strand of the target dsDNA, thereby directing the complex to the guide sequence of the target dsDNA.

[0054] In some implementations, the guiding nucleic acid is a guiding RNA (gRNA).

[0055] In some embodiments, the stent sequence has a secondary structure that is substantially the same as that of the sequence of SEQ ID NO:40, 73 or 74.

[0056] In some embodiments, the scaffold sequence comprises (1) a 5' or 3' truncated form of SEQ ID NO:40, 73, or 74, or a 5' or 3' truncated form thereof by 1, 2, 3, 4, 5, or 6 nucleotides; or (2) a sequence having at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO:40, 73, or 74, or a 5' or 3' truncated form thereof by 1, 2, 3, 4, 5, or 6 nucleotides; or (3) a sequence differing from SEQ ID NO:40, 73, or 74 by at most 1, 2, 3, 4, 5, or 6 nucleotides, regardless of whether the difference is continuous.

[0057] In some embodiments, the fusion protein or system of this disclosure further comprises a trans-damage synthesis (TLS) polymerase or a recruitment domain or component capable of recruiting a TLS polymerase.

[0058] In some embodiments, the TLS polymerase is selected from the group consisting of Pol α (α), Pol β (β), Pol δ (δ) (PCNA), Pol γ (γ), Pol η (n), Pol ι (i), Pol κ (κ), Pol λ (λ), Pol μ (μ), Pol ν (ν), Pol θ (θ), and REV1.

[0059] In another aspect, this disclosure provides a polynucleotide encoding the fusion protein of this disclosure and optionally the guiding nucleic acid of this disclosure.

[0060] In another aspect, this disclosure provides a delivery system comprising (1) a fusion protein of this disclosure, a polynucleotide of this disclosure, or a system of this disclosure; and (2) a delivery vehicle.

[0061] In another aspect, this disclosure provides a carrier containing the polynucleotides of this disclosure.

[0062] In another aspect, this disclosure provides a complex comprising a fusion protein of this disclosure or a polynucleotide (e.g., mRNA) encoding said fusion protein and a guiding nucleic acid (e.g., gRNA) of this disclosure.

[0063] In another aspect, this disclosure provides a pharmaceutical composition comprising (1) the system of this disclosure, the carrier of this disclosure, the ribonucleoprotein of this disclosure, the lipid nanoparticles of this disclosure, or the cell of this disclosure; and (2) a pharmaceutically acceptable excipient. In another aspect, this disclosure provides a cell or its progeny comprising the system of this disclosure. In yet another aspect, this disclosure provides a cell or its progeny modified by the system of this disclosure or the method of this disclosure.

[0064] In another aspect, this disclosure provides a method for modifying target dsDNA, the method comprising contacting the target dsDNA with a system of this disclosure.

[0065] The target dsDNA contains:

[0066] (a) The first deoxyribonucleotide (e.g., dG (deoxyguanosine), dT (thymidine), dC (deoxycytidine)) in the prototype spacer sequence on the non-target strand (editing strand) of the target dsDNA, and

[0067] (b) A second deoxyribonucleotide (e.g., dC (deoxycytidine), dA (deoxyadenosine), dG (deoxyguanosine)) in a target sequence that is paired with the first deoxyribonucleotide (e.g., dG, dT, dC) and located on the target strand (non-edited strand) of the target dsDNA, wherein the prototype spacer sequence is completely inversely complementary to the target sequence.

[0068] In some embodiments, the method does not include deamination of the bases of the first deoxyribonucleotide before removing the bases of the first deoxyribonucleotide.

[0069] In some embodiments, the method does not include deamination of the bases of the first deoxyribonucleotide.

[0070] In another aspect, this disclosure provides an MPG as described herein or disclosed herein.

[0071] In another aspect, this disclosure provides a UNG described herein or disclosed herein.

[0072] Details of one or more embodiments of this disclosure are set forth in the following description. Other features or advantages of this disclosure will become apparent from the following drawings and detailed description, as well as from the appended claims. It should be understood that, unless otherwise indicated, any aspect or embodiment of this disclosure may be combined with any other aspect or embodiment of this disclosure (including aspects or embodiments described only in a subsection, only in the embodiments, or only in the claims) to constitute another embodiment explicitly or implicitly disclosed herein.

[0073] definition

[0074] This disclosure describes specific embodiments, but is not limited thereto. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Unless otherwise indicated, the terms set forth below are generally understood in their simple and common sense or as commonly understood.

[0075] Overview

[0076] Programmable nucleic acid binding proteins (napBPs) (e.g., programmable nucleic acid DNA binding proteins (napDNAbp) (such as Cas9, Cas12, IscB), programmable nucleic acid RNA binding proteins (napRNAbp) (such as Cas13)) can bind to target nucleic acids (e.g., dsDNA, mRNA), as directed by a guide nucleic acid (e.g., guide RNA) containing a guide sequence targeting the target nucleic acid. In some embodiments, the target nucleic acid is eukaryotic.

[0077] To avoid being bound by theory, in some implementations, the guiding nucleic acid includes a scaffold sequence responsible for forming a complex with napBP, and a guiding sequence intentionally designed to hybridize with a target sequence of the target nucleic acid, thereby guiding the complex containing napBP and the guiding nucleic acid to the target nucleic acid.

[0078] refer to Figure 6 An exemplary target dsDNA is depicted as comprising 5' to 3' single DNA strands and 3' to 5' single DNA strands.

[0079] An exemplary guide nucleic acid (e.g., guide RNA) is depicted as comprising a guide sequence and a scaffold sequence. The guide sequence is designed to hybridize with a portion of a 3' to 5' single DNA strand, and thus the guide sequence "targets" that portion. Therefore, the 3' to 5' single DNA strand is referred to as the "target strand (TS)" of the target dsDNA, while the opposite 5' to 3' single DNA strand is referred to as the "non-target strand (NTS)" of the target dsDNA. The portion of the target strand that serves as the basis for designing the guide sequence and can hybridize with the guide sequence is referred to as the "target sequence," and the opposite portion of the non-target strand corresponding to that portion is referred to as the "prototype spacer sequence," which is 100% (perfectly) anticomplementary to the target sequence and is referred to herein as "corresponding to" the target sequence.

[0080] refer to Figure 6 The exemplary target dsDNA is depicted as comprising a 5' to 3' single DNA strand and a 3' to 5' single DNA strand. Following a conventional transcription process, the exemplary target RNA (transcription, e.g., pre-mRNA) can be transcribed using the 3' to 5' single DNA strand as a template for synthesis; therefore, the 3' to 5' single DNA strand is referred to as the "template strand" or "antisense strand." The transcript thus transcribed has the same primary sequence as the 5' to 3' single DNA strand (except for the substitution of T with U); therefore, the 5' to 3' single DNA strand is referred to as the "coding strand" or "sense strand."

[0081] An exemplary guide nucleic acid (e.g., guide RNA) is depicted as comprising a guide sequence and a scaffold sequence. The guide sequence is designed to hybridize with a portion of the transcript (target RNA), thus the guide sequence "targets" that portion. Therefore, the portion of the target RNA that serves as the basis for the designed guide sequence and can hybridize with the guide sequence is referred to as the "target sequence." In some embodiments, the guide sequence is 100% (perfectly) anticomplementary to the target sequence. In some other embodiments, the guide sequence is anticomplementary to the target sequence and contains mismatches with the target sequence.

[0082] Typically, as is customary in the art, unless otherwise explicitly instructed, nucleic acid sequences (e.g., DNA sequences, RNA sequences) are written in a 5' to 3' orientation.

[0083] For example, the DNA sequence for ATGC is generally understood as 5'-ATGC-3' unless otherwise specified. Its reverse sequence is 5'-CGTA-3'. Its complete complement is 5'-TACG-3'. Its complete reverse complement is 5'-GCAT-3'. Note that complete complement sequences typically do not have the ability to pair / hybridize with the original sequence.

[0084] Typically, unless otherwise indicated, the double-stranded sequence of dsDNA can be represented by the sequence of its 5' to 3' single DNA strand, written conventionally in the 5' to 3' direction / orientation.

[0085] For example, for a dsDNA with a 5' to 3' single DNA strand of 5'-ATGC-3' and a 3' to 5' single DNA strand of 3'-TACG-5', the dsDNA can be simply represented as 5'-ATGC-3'.

[0086] 5'-----ATGC-----3'

[0087] 3'-----TACG-----5'

[0088] It should be noted that the 5' to 3' single DNA strand or the 3' to 5' single DNA strand of dsDNA can be a non-target strand from which the prototype spacer sequence is selected.

[0089] Typically, for genes that are dsDNA, the 5' to 3' single DNA strand is the sense strand of the gene, and the 3' to 5' single DNA strand is the antisense strand. It should be noted that the sense or antisense strand of a gene can be a non-target strand from which the prototypical spacer sequence is selected.

[0090] Normally, transcripts transcribed from dsDNA (target RNA) have a 5'-AUGC-3' (target) sequence.

[0091] In order to hybridize with the target dsDNA, in one embodiment, the guide sequence of the guiding nucleic acid is designed to have a 5'-AUGC-3' sequence that is completely anticomplementary to the 3' to 5' strands of the target dsRNA, and this sequence will be shown in the electronic sequence listing as ATGC but labeled as an RNA sequence; and in another embodiment, the guide sequence of the guiding nucleic acid is designed to have a 5'-GCAU-3' sequence that is completely anticomplementary to the 5' to 3' strands of the target dsRNA, and this sequence will be shown in the electronic sequence listing as GCAT but labeled as an RNA sequence.

[0092] In cases where the guide sequence of the guiding nucleic acid is completely anticomplementary to the target sequence and the target sequence is completely anticomplementary to the prototype spacer sequence, the guide sequence is identical to the prototype spacer sequence except for the U in the guide sequence (due to its RNA nature) and the T in the corresponding prototype spacer sequence (due to its DNA nature). According to WIPO standard ST.26, the symbol “t” is used to represent both T in DNA and U in RNA (see “Table 1: List of Nucleotide Symbols”, where the symbol “t” is defined as “thymine in DNA / uracil in RNA (t / u)”). Therefore, in the electronic sequence listing of this disclosure prepared according to ST.26, such a guide sequence can be shown with the same sequence as the corresponding prototype spacer sequence. For convenience, a single SEQ ID NO in the electronic sequence listing can be used to represent both such a guide sequence and the prototype spacer sequence, regardless of whether such a single SEQ ID NO is labeled as DNA or RNA in the electronic sequence listing. When a reference is given as such a SEQ ID NO as a prototype spacer sequence / guide sequence, it refers to either a prototype spacer sequence as a DNA sequence or a guide sequence as an RNA sequence, as the case may be, regardless of whether it is labeled as DNA or RNA in the electronic sequence listing.

[0093] In order to hybridize with the target RNA, in one embodiment, the guide sequence of the guiding nucleic acid is designed to have a 5'-GCAU-3' sequence that is completely reverse complementary to the (target) sequence of the target RNA. This sequence will be shown in the electronic sequence listing as GCAT, but is labeled as an RNA sequence.

[0094] the term

[0095] As used herein, if a DNA sequence (e.g., 5'-ATGC-3') is transcribed into an RNA sequence in which each dT (deoxythymidine, or simply "T") in the primary sequence is replaced with U (uridine), and the other dA (deoxyadenosine, or simply "A"), dG (deoxyguanosine, or simply "G"), and dC (deoxycytidine, or simply "C") are replaced with A (adenosine), G (guanosine), and C (cytidine), respectively, for example 5'-AUGC-3', then in this disclosure the DNA sequence is referred to as "encoding" the RNA sequence.

[0096] As used herein, the term "activity" refers to biological activity. In some embodiments, activity includes enzymatic activity, such as the catalytic ability of an effector. For example, activity may include nuclease activity, such as dsDNA endonuclease activity or RNA endonuclease activity.

[0097] As used herein, the term "nucleic acid programmable binding protein (napBP)" is used interchangeably with "nucleic acid programmable binding domain (napBD)" and refers to a protein that can be associated (e.g., bound) to a programmable nucleic acid (e.g., DNA or RNA) (such as a guide nucleic acid (e.g., gRNA)), which can be programmed to direct the protein to a specific sequence of a target nucleic acid via an interaction (e.g., hybridization) between the programmable nucleic acid (e.g., a guide sequence of the programmable nucleic acid) and a target nucleic acid (e.g., a target sequence of the target nucleic acid). napBP can be indirectly associated (e.g., bound) to a target nucleic acid via an interaction (e.g., binding) between the napBP and a programmable nucleic acid (e.g., a scaffold sequence of the programmable nucleic acid) and an interaction (e.g., hybridization) between the programmable nucleic acid (e.g., a guide sequence of the programmable nucleic acid) and a target nucleic acid (e.g., a target sequence of the target nucleic acid). In some embodiments, napBP is a nucleic acid programmable DNA binding protein (napDNAbp). In some embodiments, napBP is a nucleic acid programmable RNA binding protein (napRNAbp).

[0098] As used herein, the term "complex" refers to the grouping of two or more molecules. In some embodiments, the complex comprises polypeptides and nucleic acids that interact with each other (e.g., bind, contact, adhere). As used herein, the term "complex" can refer to the grouping of a nucleic acid and a polypeptide (e.g., napBP). As used herein, the term "complex" can refer to the grouping of a nucleic acid, a polypeptide (e.g., napBP), and a target nucleic acid.

[0099] As used herein, the term "prototype spacer adjacent motif" or "PAM" refers to a short DNA sequence (or DNA motif) on the non-target strand of dsDNA that is adjacent to the prototype spacer sequence. As used herein, the term "adjacent" includes the case where there are no nucleotides between the prototype spacer sequence and the PAM, and the case where there are a small number of nucleotides (e.g., 1, 2, 3, 4, or 5) between the prototype spacer sequence and the PAM. As used herein, A "adjacent to" B, A "adjacent to the 5' of B," and A "adjacent to the 3' of B" mean that there are no nucleotides between A and B. In some embodiments, the PAM is adjacent to the 5' of the prototype spacer sequence. In some embodiments, the PAM is adjacent to the 3' of the prototype spacer sequence.

[0100] As used herein, the term "guide nucleic acid" refers to any nucleic acid that facilitates the targeting of napBP to a target nucleic acid. For this purpose, a guide nucleic acid can be designed to include a guide sequence capable of hybridizing to a specific sequence of the target nucleic acid, and the guide nucleic acid may also contain a scaffold sequence that facilitates the guidance of napBP to the target nucleic acid. In some embodiments, the guide nucleic acid is guide RNA. In some embodiments, the guide nucleic acid is a nucleic acid encoding guide RNA.

[0101] As used herein, the terms “nucleic acid,” “polynucleotide,” and “nucleotide sequence” are used interchangeably to refer to a polymeric form of nucleotides of any length, including deoxyribonucleotides, ribonucleotides, combinations thereof, and analogs or modifications thereof.

[0102] As used in the case of CRISPR-Cas technology (e.g., CRISPR-Cas12 technology), the term “guide RNA” is used interchangeably with the terms “CRISPR RNA (crRNA),” “single guide RNA (sgRNA),” or “RNA guide,” the term “guide sequence” is used interchangeably with the term “spacer sequence,” and the term “scaffold sequence” is used interchangeably with the term “directed repeat sequence.”

[0103] As described herein, the guide sequence is designed to hybridize with the target sequence. As used herein, the term "hybridize (hybridize, hybridizing, or hybridization)" refers to a reaction in which one or more polynucleotide sequences react to form a complex that is stabilized via hydrogen bonds between the bases of the polynucleotide sequence. Hydrogen bonds can occur through Watson-Crick base pairing, Hoogstein binding, or any other sequence-specific manner. A polynucleotide sequence capable of hybridizing with a given polynucleotide sequence is referred to as its "complement." As used herein, the hybridization of the guide sequence with the target sequence is stabilized to allow functional domains of an effector peptide (e.g., napBP) or associated with (e.g., fusion) a nucleic acid containing the guide sequence to act on the target sequence or its complement or a nearby sequence (e.g., cleavage, deamination).

[0104] For hybridization purposes, in some embodiments, the guide sequence and the target sequence are reverse complementary. As used herein, the term "reverse complementary" refers to the ability of the nucleobases of a first polynucleotide sequence (such as the guide sequence) to pair with the nucleobases of a second polynucleotide sequence (such as the target sequence) via conventional Watson-Crick base pairing. Two reverse complementary polynucleotide sequences can bind nonvalently under appropriate temperature and solution ionic strength conditions. In some embodiments, the first polynucleotide sequence (e.g., the guide sequence) contains 100% (complete) reverse complementarity with the second nucleic acid (e.g., the target sequence). In some embodiments, if the first polynucleotide sequence contains at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementarity with the second nucleic acid (i.e., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the nucleotides in the first polynucleotide sequence can pair with the nucleotide bases of the second polynucleotide sequence), then the first polynucleotide sequence (e.g., the guide sequence) is reverse complementary to the second polynucleotide sequence (e.g., the target sequence). As used herein, the term “substantially complementary” means that the first polynucleotide sequence (e.g., the guide sequence) and the second polynucleotide sequence (e.g., the target sequence) have a certain level of complementarity (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the nucleotides in the first polynucleotide sequence can pair with the nucleotide bases of the second polynucleotide sequence, or at most 1, 2, 3, 4, or 5 adjacent or non-adjacent nucleotides in the first polynucleotide sequence are mismatched with the nucleotides of the second polynucleotide sequence). In some embodiments, the complementarity level enables the first polynucleotide sequence (e.g., a guide sequence) to hybridize with a second polynucleotide sequence (e.g., a target sequence) with sufficient affinity to allow functional domains of an effector peptide (e.g., napBP) complexed with a nucleic acid containing the first polynucleotide sequence, or associated with (e.g., fusion) the effector peptide, to act on the target sequence or its complement or a nearby sequence (e.g., cleavage, deamination). In some embodiments, the guide sequence, which is substantially complementary to the target sequence, has less than 100% complementarity with the target sequence.In some implementations, the guide sequence, which is substantially complementary to the target sequence, has at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% complementarity to the target sequence, and / or has at most 1, 2, 3, 4, or 5 adjacent or non-adjacent nucleotide mismatches with the target sequence.

[0105] As used herein, the term “sequence identity” is related to sequence homology. Homology comparisons can be performed visually or, more commonly, with the aid of readily available sequence comparison programs. These commercially available computer programs can calculate the percentage (%) of sequence identity between two or more sequences (peptide or polynucleotide sequences). Sequence homology can be generated using any of a variety of computer programs known in the art, such as BLAST and FASTA. A suitable computer program for performing such comparisons is the GCG Wisconsin Bestfit package (University of Wisconsin, USA; Devereux et al., 1984, Nucleic Acids Research, 12:387). Examples of other software capable of performing sequence comparisons include, but are not limited to, the BLAST package (see Ausubel et al., 1999, ibid. – Chapter 18), FASTA (Atschul et al., 1990, J. Mol. Biol., 403-410), and the GNEWORKS comparison tool suite. Both BLAST and FASTA can be used for offline and online searches (see Ausubel et al., 1999, ibid., pp. 7-58-7-60). Commonly used online tools for calculating the percentage of sequence identity between two or more sequences (peptide or polynucleotide sequences) are available on the website of the European Institute of Bioinformatics at EMBL (www.ebi.ac.uk.slash.jdispatcherslash), allowing for rapid online calculation of the percentage of sequence identity through global or local alignment.

[0106] As used herein, the terms “polypeptide” and “peptide” are used interchangeably to refer to amino acid polymers of any length. Proteins may have one or more polypeptides. Amino acid polymers may also be modified, for example, by disulfide bond formation, glycosylation, esterification, acetylation, phosphorylation, or any other operation, such as conjugation with a labeled component.

[0107] As used herein, “variant” is interpreted as meaning a polynucleotide or polypeptide that differs from a reference polynucleotide or polypeptide but retains essential properties (e.g., the binding properties of napBP). Typical variants of a polynucleotide differ in their nucleic acid sequence from another reference polynucleotide. Changes in the nucleic acid sequence of a polynucleotide variant may or may not alter the amino acid sequence of the polypeptide encoded by the reference polynucleotide. Changes in the nucleic acid sequence of a polynucleotide variant may result in amino acid substitutions, additions, and / or deletions in the polypeptide encoded by the reference polynucleotide. Typical variants of a polypeptide differ in their amino acid sequence from another reference polypeptide. Typically, the differences are limited, making the sequences of the reference polypeptide and the polypeptide variant very similar overall and identical in many regions. Polypeptide variants and reference polypeptides may differ in their amino acid sequences by one or more substitutions, additions, and / or deletions in any combination. Polynucleotide or polypeptide variants may be naturally occurring, such as allelic variants, or they may be variants known not to be naturally occurring. Non-naturally occurring variants of polynucleotides and polypeptides can be prepared by mutagenesis, by direct synthesis, and by other recombinant methods known to those skilled in the art.

[0108] As used herein, the terms "upstream" and "downstream" refer to the relative positions of two or more elements within a nucleic acid in the 5' to 3' direction. When the 3' end of the first sequence is to the left of the 5' end of the second sequence, the first sequence is upstream of the second sequence. When the 5' end of the first sequence is to the right of the 3' end of the second sequence, the first sequence is downstream of the second sequence. In some embodiments, PAM is upstream of a napBP-induced indel, and the napBP-induced indel is downstream of PAM. In some embodiments, PAM is downstream of a napBP-induced indel, and the napBP-induced indel is upstream of PAM.

[0109] As used herein, the term "wildtype" has the meaning commonly understood by those skilled in the art as referring to the typical form of an organism, strain, gene, or trait that distinguishes it from mutants or variants found in nature. It can be isolated from its natural source without intentional modification.

[0110] As used herein, the terms “non-naturally occurring” and “engineered” are used interchangeably and refer to artificial involvement. When these terms are used to describe a nucleic acid or polypeptide, it means that the nucleic acid or polypeptide is at least substantially free of at least one other component that is found in nature or is associated with it as such.

[0111] As used herein, the term "regulatory element" is intended to include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Such regulatory elements are described, for example, in the following literature: Goeddel, *Gene Expression Technology: Methods of Enzymology*, 185, Academic Press, San Diego, Calif. (1990). Regulatory elements include those that guide constitutive expression of nucleotide sequences in many types of host cells, as well as those that guide expression of nucleotide sequences only in certain host cells (e.g., tissue-specific regulatory sequences). Regulatory elements can also guide expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which may or may not be tissue- or cell type-specific.

[0112] As used herein, the term "cell" is understood not only to refer to a specific, individual cell, but also to the offspring or potential offspring of that cell. Because certain modifications may occur in subsequent generations due to mutations or environmental influences, such offspring may indeed differ from the parent cell, but are still included within the scope of this term.

[0113] As used herein, the term “in vivo” refers to the interior of an organism’s body, and the terms “out vivo” or “ex vivo” refer to the exterior of an organism’s body.

[0114] As used herein, the term "treat" (or "treatment") is a method for achieving a beneficial or desired outcome (including clinical outcomes). For the purposes of this disclosure, a beneficial or desired clinical outcome includes, but is not limited to, one or more of the following: alleviating one or more symptoms caused by a disease; reducing the severity of the disease; stabilizing the disease (e.g., delaying disease progression); delaying the spread of the disease (e.g., metastasis); delaying the recurrence of the disease; reducing the recurrence rate of the disease; delaying or slowing the progression of the disease; improving the disease state; providing remission of the disease (partial or complete); reducing the dosage of one or more other medications required to treat the disease; delaying disease progression; improving quality of life and / or prolonging survival. "Treatment" also encompasses reducing the pathological consequences of a disease (e.g., cancer). The methods of this disclosure contemplate any one or more of these aspects of treatment.

[0115] As used in this article, the term “disease” includes, but is not limited to, the terms “disorder” and “symptom”.

[0116] As used herein, the term “transcription” includes any transcript obtained by transcription from DNA, including subgenomic RNA, mRNA, noncoding RNA and any variants, derivatives or ancestors thereof, such as premRNA, and any transcript or variant produced from DNA or premRNA by, for example, the use of a selective promoter, selective splicing, selective initiation, and any naturally occurring variants thereof or products processed therefrom.

[0117] As used in this article, when referring to something as "not" a value or parameter, it usually means and describes something as "different" from that value or parameter. For example, "not using a method to treat type X cancer" means that the method can be used to treat cancers different from type X.

[0118] Unless otherwise explicitly stated in the context, as used herein, the singular forms “a / an” and “the” include plural indicators.

[0119] As used herein, the term "and / or" in phrases such as "A and / or B" is intended to include both A and B; A or B; A (alone); and B (alone). Similarly, the term "and / or" in phrases such as "A, B, and / or C" is intended to cover each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).

[0120] As used herein, when the term “about” precedes a series of numbers (e.g., about 1, 2, 3), it should be understood that each of those numbers is modified by the term “about” (i.e., about 1, about 2, about 3). The terms “about XY” or “about X to Y” as used herein have the same meaning as “about X to about Y”.

[0121] It should be understood that the embodiments described herein include “consisting of embodiments” and / or “substantially consisting of embodiments”.

[0122] It should also be noted that claims can be drafted without any optional elements. Therefore, this statement is intended to serve as a pre-basis for combining the elements of the claims using exclusive terms such as “solely” or “only” or using “negative” restrictions. Attached Figure Description

[0123] The features and advantages of this disclosure will be understood by referring to the following detailed description and accompanying drawings, which illustrate illustrative embodiments that can utilize the principles of this disclosure, in which:

[0124] Figure 1The design and mechanism of base editing in conventional base editors and the glycosylation-based base editor disclosed herein. Figure 1 a: Schematic diagram of ABE (left) and CBE (middle) and the deaminase-free glycosylase-based guanine base editor (gGBE, right) of this disclosure. The nCas9-sgRNA complex generates an R loop at the target site in DNA. In ABE and CBE, evolved adenine deaminases (tRNA adenosine deaminase, TadA) and AID / APOBEC-like cytidine deaminase convert exposed adenine (A) to deoxyinosine (I) and cytosine (C) to deoxyuridine (U), respectively. In CBE, an additional linker protein, uracil glycosylase inhibitor (UGI), protects U from uracil DNA N-glycosylase (UNG). After deamination, during DNA repair or replication, the resulting I is recognized as G by DNA polymerase, and U is recognized as T. In the gGBE disclosed herein, the prototype form of the glycosylation-based guanine base editor (gGBE) is designed to remove G, and the resulting AP site is repaired via TLS and / or DNA replication, leading to G-to-C or G-to-T conversion. PAM, prototype spacer adjacent motif. AP, apurinol / pyrimidine-free site. Figure 1 b: A screening and reporting system for detecting the G-to-T transformation of gGBE. P2A, a 2A peptide from porcine swine clonal virus-1. Figure 1 c: EGFP used for evaluating the G-editing efficiency of gGBE with various MPGs + Percentage of cells (mean ± sem, n = 3). WT: Wild type.

[0125] Figure 2 Mutagenesis of the MPG portion in .gGBE. Figure 2 a: Schematic diagram of the mutagenesis and screening strategy for engineered gGBE. The EGFP reporter plasmid and gGBE expression plasmid were transiently co-transfected into cultured cells. Figure 2 b: Genotypes of engineered gGBE subsets, where EGFP is present for each gGBE subset. + Cell percentage is in the rightmost column (more engineered gGBE is listed in...). Figure 13 (In the middle). Different mutagenesis steps are marked with different shade colors. Figure 2 cd: Site 3 in transfected HEK293T cells was determined by target depth sequencing. Figure 2 c) and site 10 ( Figure 2 In d), the edited G7 position was obtained using G base editing results with different gGBEs (mean ± sem, n = 3).

[0126] Figure 3 Characterization of gGBE editing profile via target depth sequencing. Figure 3 a: A bar chart showing the mid-target DNA base editing (mean ± sem, n = 3) at each genomic locus in HEK293T cells at the location with the highest G conversion frequency. G number: G location with the highest mid-target base editing frequency in prototype spacer positions 1-20. Locus number: Genomic locus number. Figure 3 b: In Figure 3 At the site shown in a, the ratio of G to C / T to G to A / C / T conversion frequencies edited by gGBEv6.3. Figure 3 c: In Figure 3 The edit sites in 'a' are located in the prototype spacer subpositions 1-20, and the G transition frequencies of gGBEv6.3 (where PAM is located at positions 21-23). ​​A single point represents a separate data point from 3 independent replicates at each site. The box spans the interquartile range (25th to 75th percentiles); the horizontal line within the box indicates the median (50th percentile); and small horizontal bars mark the minimum and maximum values. Figure 3 d: Guided sequence-dependent off-target analysis of gGBEv6.3 editing efficiency at site 7 (mean ± sem, n = 3). OT: Off-target. Figure 3 e: Guide sequence-independent off-target editing efficiency (mean ± sem, n = 3) detected by orthogonal R-loop determination at each R-loop site.

[0127] Figure 4 Gene editing applications of .gGBE. Figure 4 a: gGBEv6.3 is used for editing splice sites, introducing early stop codons (PTCs), and editing applications that bypass PTCs. Figure 4 b: A schematic diagram showing gGBE-induced skipping of DMD exon 45. Figure 4 c: Bar graph showing the frequency of mid-target DNA base editing by G editing of gGBEv6.3 and the ratio of G to C / T to G to A / C / T editing frequencies (mean ± sem, n = 3) at two DMD sites in HEK293T cells. SpG (PAM flexible Cas9 variant (SEQ ID NO: 163)) was used to target DMD site 2. Figure 4 d: A schematic diagram showing the introduction of PTC into the mouse Tyr gene via gGBE. Figure 4 e: A bar graph showing the mid-target G editing frequency and G-to-Y ratio (mean ± sem, n = 3) of gGBEv6.3 at three Tyr sites in N2a cultured cells. Figure 4f: Percentage of PTCs induced by G conversion via gGBEv6.3 at the two Tyr sites in mouse pups (mean ± sem, n = 25 and 21 pups for sites 2 and 3, respectively). Figure 4 g: Phenotype of F0 mice generated by gGBE editing in mouse zygotes. Images of the presence of edited P6 mice are shown. Red arrows indicate albinism; blue arrows indicate chimeric pigmentation. Figure 4 h: This shows a bar graph of the mid-target G-editing frequency in individual mouse pups when gGBEv6.3 targets Tyr site 3. Figure 4 i: from ( Figure 4 Genotyping of representative F0 offspring (h). Allele frequencies of mutants were determined using high-throughput sequencing. Red arrows indicate albino offspring.

[0128] Figure 5 An exemplary nucleotide conversion achieved through base excision and trans-damage synthesis (TLS) is shown.

[0129] Figure 6 Exemplary target dsDNA, exemplary guide nucleic acid, and exemplary napDNAbp containing a first exemplary deoxyribonucleotide dG prior to base editing are shown.

[0130] Figure 7 Exemplary target dsDNA, exemplary guide nucleic acid, and exemplary napDNAbp are shown, which contain a fourth exemplary deoxyribonucleotide dC after base editing.

[0131] Figure 8 Exemplary target dsDNA, exemplary guide nucleic acid, and exemplary napDNAbp containing a fourth exemplary deoxyribonucleotide dT after base editing are shown.

[0132] Figure 9 An overview of AYBE and CGBE. Figure 9 a: Schematic diagram of AYBE. N-methylpurine DNA glycosylase (MPG) removes inosine (I) produced by the deamination of adenine (A) by adenine deaminase (TadA), triggering the base excision repair (BER) pathway in the cell, resulting in more diverse base editing outcomes, including A-to-C and A-to-T editing. Figure 9 b: Schematic diagram of CGBE. Uracil DNA N-glycosylation (UNG) excision: Urosine (U) produced by the deamination of cytosine (C) by AID / APOBEC-like cytidine deaminase triggers the base excision repair (BER) pathway in the cell, resulting in dominant C-to-G editing. PAM, prototypical spacer adjacent motif. AP, apurinol / pyrimidine-free site.

[0133] Figure 10 Characterization of A-to-T and G-to-T editing using an intron-splitting EGFP reporter system. Figure 10 a: Design of reporter for A-to-T or G-to-T edit detection. P2A, a 2A peptide from porcine swine cyclovir-1. Figure 10 b: EGFP used to evaluate the A-editing efficiency of gABE with various MPGs. + Percentage of cells. WT: wild type; v0.2: MPG-N169S; v1: MPG with mutations in N169S, S198A, K202A, G203A, S206A, and K210A; v2: MPG with mutations in N169S and G163R; v3: MPG with mutations in G163R, N169S, S198A, K202A, G203A, S206A, and K210A (mean ± sem, n = 3). Figure 10 c: EGFP representing the G-to-T conversion efficiency of gGBE containing various MPG and sgRNAs. + Percentage of cells. dMPG, inactive MPG; T: targeting sgRNA; NT: non-targeting sgRNA (mean ± sem, n = 3). Figure 10 d: This shows the gating strategy of gGBEv3 and EGFP. + A flow cytometry scatter plot representing the percentage of cells.

[0134] Figure 11 A view of the MPG structure and the first round of mutations in MPG. Figure 11 a: The structure of the human MPG protein in regions aa 78-298 (left) and 163-179 (right) as predicted by AlphaFold (alphafold.com / entry / P29372), compared with the crystal structure of MPG (PDB entry 1ewn, not shown), where εA is mutated to G in DNA. Figure 11 b: EGFP used to evaluate the G-editing efficiency of different engineered gGBEs with various MPG mutants. + Percentage of cells (mean ± sem, n = 3). Figure 11 c: Various engineered gGBEs passed the first round of screening using EGFP. + Performance of cell percentage measurement. Dashed line, mean of the MPGv3 group. Calculate the fold change relative to gGBEv3 (mean ± sem, n = 3). Figure 11 d: EGFP for each gGBE + Percentage of cells (mean ± sem, n = 3).

[0135] Figure 12 Performance of engineered gGBE in the second round of screening. Figure 12 ad: EGFP with gGBE containing various MPG mutants + The percentage of cells from which these mutants originate (glutamate) Figure 12 a) Valine ( Figure 12 b) Glycine ( Figure 12 c) and tyrosine ( Figure 12 d) The sequential replacement of (X to E, V, G or Y). n = 3. All values ​​are presented as mean ± sem.

[0136] Figure 13 The progressive engineering and G-editing efficiency of .gGBE. Figure 13 a: Progressive mutation of gGBE. Mutations from different rounds are marked with different shaded colors. Figure 13 b: EGFP for each gGBE + Percentage of cells. n=3. All values ​​are presented as mean ± sem.

[0137] Figure 14 Further characterization of the editing profile of .gGBEv6.3. Figure 14 ac: from Figure 3 In the prototype spacer positions 1-20 of the edit site in a (where PAM is at positions 21-23), the C( of gGBEv6.3) Figure 14 a) T( Figure 14 b) and A( Figure 14 c) Frequency of conversion. Figure 14 d: The frequency of G to T and G to C editing in gGBEv6.3. Figure 14 In the ad, a single dot represents a single replicate (n = 3 independent replicates / locus), and the box spans the interquartile range (25th to 75th percentile); the horizontal line inside the box indicates the median (50%); and the whisker line extends to the minimum and maximum values. Figure 14 e.g.: in Figure 3 At each edit site shown in a, G to C of gGBEv6.3 ( Figure 14 e), G to T ( Figure 14 f) or G to A( Figure 14 g) Percentage of edits (mean ± sem, n = 3). Figure 14 h: Insertion / deletion frequency (mean ± sem, n = 3) at 24 intermediate target sites using gGBEv6.3. Figure 14 i: Bar graph showing the mid-target DNA base editing (mean ± sem, n = 3) at each genomic locus in HEK293T cells at locations with a G conversion frequency >10%. Figure 14 j: from ( Figure 14 Statistical analysis of target DNA base editing at each NG motif at the edit site in i). Each point represents the mean of three biological replicates for each edit position at each edit site.

[0138] Figure 15 Off-target analysis of guided sequence dependency and guided sequence independence. Figure 15 a: Off-target analysis of guide sequence dependence of gGBEv6.3 editing at different sites (n=3). OT: Off-target. Figure 15 b: Guide sequence-independent off-target editing of gGBE, AYBE, and ABE8e detected at each R-loop site by orthogonal R-loop determination (n=3). Data for AYBE and ABE8e from Tong et al. [1] were used. All values ​​are presented as mean ± sem.

[0139] Figure 16 Percentage of G to C and G to T in all G to C / T / A conversion events at the DMD or Tyr site targeted by gGBEv6.3. Figure 16 a: In HEK293T cells, at two DMD sites (corresponding to...) Figure 4 At point C), the percentage of G-to-C and G-to-T edit events in gGBEv6.3. n = 3. SA, cleavage acceptor site. PAM, prototypical spacer neighbor motif. Figure 16 b: In N2a cells, at three Tyr sites (corresponding to...) Figure 4 e) Percentage of G to C and G to T edit events in gGBEv6.3. n = 3. All values ​​are presented as mean ± sem.

[0140] Figure 17 G editing of gGBEv6.3 in mouse embryos. Figure 17 a: Mid-target base editing efficiency of gGBEv6.3 targeting three Tyr sites in mouse embryos (mean ± sem, n = 20). Figure 17 bc: Percentage of PTCs induced by G conversion via gGBEv6.3 at three Tyr sites in mouse embryos (mean ± sem, n = 20 embryos). Figure 17 d: through Figure 17 Figure a shows the gGBEv6.3-induced insertion / deletion frequencies (n=20) of three Tyr sites in mouse embryos targeted by gGBEv6.3. Figure 17 e: This shows a bar graph illustrating the mid-target G-editing frequencies in individual mouse embryos when gGBEv6.3 targets Tyr sites 1, 2, and 3.

[0141] Figure 18 Phenotypic and genotypic classification of F0 mouse pups. Figure 18 ac: via microinjection of mRNA encoding gGBEv6.3 and for targeting Tyr site 1 ( Figure 18 a) Site 2 ( Figure 18 b) and site 3 ( Figure 18 c) Phenotype of F0 mice generated by sgRNA. Images of P6 mice were obtained. Arrows: Red, albino mice; Blue, mice with chimeric pigmentation. Figure 18 d: A bar graph showing the mid-target G-editing frequency of individual mouse pups when gGBEv6.3 targets Tyr site 1 and Tyr site 2.

[0142] Figure 19 Design and mechanism of two orthogonal glycosylation enzyme-based base editors. Figure 19 a, Prototypes of the deaminase-free glycosylation-based thymine base editor (gTBE) and the deaminase-free glycosylation-based cytosine base editor (gCBE). PAM, Prototype spacer neighbor motif. AP, Purine-free / Pyrimidine-free site. Magenta asterisk (*) indicates a nick produced by nCas9. Figure 19 b. Schematic diagram of potential pathways and outcomes of T (or C) editing. Glycosylation enzyme mutants are engineered to remove normal T or C, the nCas9-sgRNA complex generates an R loop at the target site and creates a nick in the non-edited strand, leading to T or C editing via AP sites generated through trans-damage synthesis (TLS) and / or DNA replication repair. DSB, double-strand breaks. Insertion / deletion, insertion and deletion. Figure 19 c. Schematic diagram of various gTBE and gCBE candidate architectures. Note that Y156A and N213D of UNG2 are equivalent to Y147A and N204D of UNG1, respectively. Figure 19 d, EGFP used to evaluate the T-editing efficiency of different gTBEs using T to G reporter substances. + Percentage of cells (n=3). NT: Non-target sgRNA. T: Target sgRNA. Figure 19 e, EGFP used to evaluate the C-editing efficiency of different gCBEs using C to G reporter substances. + Percentage of cells (n=3). NT: Non-target sgRNA. T: Target sgRNA. Figure 19 f, Orthogonality of gTBE and gCBE for base editing evaluated using two different reporter materials (n=3). All values ​​are presented as mean ± sem.

[0143] Figure 20 Protein engineering and evolution of .gTBE. Figure 20a. Schematic diagram of the mutagenesis and screening strategy for engineered gTBE. The EGFP reporter plasmid and gTBE plasmid were transiently co-transfected into cultured cells, and the fluorescence intensity of EGFP was detected by flow cytometry. Figure 20 b, left side, used in human UNG-DNA complex (PDB entry 1EMH) 24 Selected residues (shown as the surface) near the catalytic site pocket of ) are mutagenized, where dΨU is mutated to T(dT) in DNA. On the right, the positions of effective residues in gTBEv3 are shown as red spheres in the three-dimensional structure. Figure 20 c. Gradual improvement in EGFP activation for each gTBE (n=3). WT, wild-type UNG2Δ88. Inactivated, catalytically inactive UNG2Δ88 (carrying D154N and H277N mutations, equivalent to D145N and H268N of UNG1). 60 . Figure 20 d. Frequency of T-base editing outcomes (left) and insertions / deletions (right) using different gTBEs at the edited T5 position in site 9 (CLYBL gene) in transfected HEK293T cells, by target depth sequencing (n=3). All values ​​are presented as mean ± sem.

[0144] Figure 21 Characterization of gTBE editing profile via target depth sequencing. Figure 21 a. Bar graph showing the mid-target DNA base editing (mean ± sem, n = 3) at each genomic locus in HEK293T cells at the location with the highest T conversion frequency. T number: T position with the highest mid-target base editing frequency among prototype spacer positions 1-20. Locus number: Genomic locus number. Figure 21 b, in Figure 21 At the site shown in a, the ratio of T to C / G to T to A / C / G conversion frequencies edited by gTBEv3. Figure 21 c. Use the insertion / deletion frequency of gTBEv3 at 20 intermediate target sites (n=3). Figure 21 off-target analysis of guided sequence dependence of gTBEv3 editing efficiency at sites 1 and 15 (n=3). OT: off-target. Figure 21 f represents the guide sequence-independent off-target editing efficiency detected at each R-loop site by orthogonal R-loop determination (n=3). All values ​​are presented as mean ± sem.

[0145] Figure 22 Enhanced gCBE editing efficiency through protein engineering. Figure 22 a. Schematic diagram of the mutagenesis and screening strategies for engineered gCBE. Figure 22b, Gradual improvement in EGFP activation for each gCBE (n=3). WT, wild-type UNG2Δ88. Inactivated, catalytically inactive UNG2Δ88 (carrying D154N and H277N mutations, equivalent to D145N and H268N of UNG1). 60 . Figure 22 c shows a bar chart of mid-target DNA base editing (n=3) at each genomic locus in HEK293T cells, located at the position with the highest C conversion frequency. C number: The C position with the highest mid-target base editing frequency among the prototype spacer positions 1-20. Locus number: Genomic locus number. Figure 22 d shows a bar graph illustrating the mid-target DNA base editing (n=3) of gCBEv2 or CGBE1 at different locations in three loci. Figure 22 e, for orthogonal R-loop assays, the mid-target base editing frequency (n=3) of gCBEv2 at C6 of site 22 in HEK293T cells. Figure 22 f, the gRNA-independent cumulative off-target editing frequency detected at each R-loop site by orthogonal R-loop assay. Each R-loop was performed by co-transfecting each base editor with SpCas9 sgRNA, which targets dSaCas9 and

[0146] The sites corresponding to SaCas9 sgRNA (n=3). All values ​​are presented as mean ± sem.

[0147] Figure 23 Gene editing applications of gTBE and gCBE. Figure 23 a. The principle of exon skipping using a base editor. Figure 23 b shows a bar graph illustrating the number of sgRNA candidates targeting splicing sites in 16 genes for different base editors. gCBE, gCBEv2; gGBE, gGBEv6.3; gTBE, gTBEv3. The 16 genes are AGT, ANGPTL3, APOC3, B2M, CD33, DMD, DNMT3A, HPD, KLKB1, PCSK9, PDCD1, PRDM1, TGFBR2, TRAC, TTR, and VEGFA. Figure 23 c, shows Figure 23 Venn diagram of the distribution of sgRNAs of the four base editors in b. Figure 23 d shows a schematic diagram of sgRNA candidates that specifically target the SD or SA sites in human DMD using gTBEv3 (red line) or gCBEv2 (black line) (instead of ABE or CBE). Figure 23 e shows a schematic diagram of skipping human DMD exon 45 induced by disruption of splice donor sites via gTBE. Figure 23 f, the mid-target base editing efficiency of gTBEv3 targeting the splicing donor site of humanized DMD exon 45 in mouse embryos (mean ± sem, n = 20). Figure 23 g, DNA sequencing chromatograms from wild-type (WT) and representative embryos co-injected with gTBEv3 mRNA and sgRNA targeting the SD site of human DMD exon 45.

[0148] Figure 24 Comparison of different gTBEs. Figure 24 a. Strategies used for protein engineering and screening in the three studies. Figure 24 b. Basic architectural diagram of various base editors. UNG2*, UNG2 mutant from the corresponding base editor. ΔNTD, deletion of the N-terminal domain. Figure 24 c. Frequency of T-transformation at 17 endogenous loci. Thymines with an editing frequency >25% by any base editor are shown. The highest frequency at the corresponding position is highlighted in the heatmap (n=3). Figure 24 d, from Figure 24 The T-transformation frequencies of various base editors are located at the prototype spacer positions 1-20 (where PAM is at positions 21-23) of the edit sites in c. A single point represents an individual repeat (n = 3 independent repeats / site), and the box spans the interquartile range (25th to 75th percentile); the horizontal line within the box indicates the median (50%); and the whisker lines extend to the minimum and maximum values.

[0149] Figure 25 Characteristic sequences and motifs of human UNG1 and UNG2. Figure 25 a. UNG1-specific N-terminal residues (amino acids 1-35) are marked in gray. UNG2-specific N-terminal residues (amino acids 1-44) are light blue. The common RPA binding site (yellow) and globular catalytic domain (light green) are indicated. RPA, replication protein A. Figure 25 b. UNG contains five conserved motifs, numbered according to UNG2 as follows: a catalytic water-activating loop (152-GQDPYH-157); a proline-rich (Pro) loop that compresses the DNA backbone 5' into a damaged loop (174-PPPPS-178); a uracil-binding motif (210-LLLN-213); a glycine-serine (Gly-Ser) loop that compresses the DNA backbone 3' into a damaged loop (255-GS-256); and a leucine (Leu) intercalation loop that penetrates the minor groove (277-HPSPLS-282).

[0150] Figure 26 Characterization of the .T to G and C to G reporting systems. Figure 26a, a schematic construct design for reports used in T-to-G or C-to-G editing detection. PAM, prototype spacer neighbor motif. Figure 26 b shows a representative flow cytometry scatter plot of the percentage of gated cells for the gating strategy, negative control (top inset), and gCBEv0.3 (bottom inset).

[0151] Figure 27 It has editing efficiency for gTBE and gCBE candidates with various UNG-NTD truncations. Figure 27 a) EGFP used to evaluate the T-editing efficiency of gTBE candidates with various UNG mutants. + Percentage of cells (n=3). Figure 27 b, EGFP used to evaluate the C-editing efficiency of gCBE candidates with various UNG mutants. + Percentage of cells (n=3). WT, wild-type UNG2Δ88. Inactivated, catalytically inactive UNG2Δ88 (carrying D154N and H277N mutations, equivalent to D145N and H268N of UNG1). All values ​​are presented as mean ± sem.

[0152] Figure 28 Performance of the .UNG mutant in the context of gTBEv0.3. Figure 28 a, EGFP from gTBE scan-mutated with alanine from the region covering the catalytic water-activated ring, the Pro-rich ring, and the uracil-binding motif. + Percentage of cells (n=3). Replacing alanine with valine (A to V) aims to cover all residues in the region of interest. Figure 28 b, EGFP from the gTBE variant resulting from site-saturated mutagenesis at residue 214 + Percentage of cells (n=3). Figure 28 c, EGFP of gTBE with a selected spatially neighboring residue mutation of residue T214. + Percentage of cells (n=3). Inactivated, catalytically inactive UNG2Δ88 (carrying D154N and H277N mutations, equivalent to D145N and H268N of UNG1). All values ​​are presented as mean ± sem.

[0153] Figure 29 Performance of the .UNG mutant in the context of gTBEv2. Figure 29 ac, EGFP with different UNG mutants of gTBE + The percentage of cells from which these mutants originate (arginine) Figure 29 a) Aspartic acid ( Figure 29 b) and valine ( Figure 29c) Sequential substitutions (X to R, D, or V) (n = 3). Inactivated, catalytically inactive UNG2Δ88 (carrying D154N and H277N mutations, equivalent to D145N and H268N of UNG1). All values ​​are presented as mean ± sem.

[0154] Figure 30 Further characterization of the editing profile of .gTBEv3. Figure 30 a) shows a bar graph of mid-target DNA base editing (mean ± sem, n = 3) at each genomic locus in HEK293T cells at locations with a T conversion frequency >10%. Figure 30 be, in the context of Figure 21 In the prototype spacer positions 1-20 of the edit site in a (where PAM is at positions 21-23), the T( of gTBEv3) Figure 30 b), G( Figure 30 c), C( Figure 30 d) and A( Figure 30 e) Frequency of conversion. Figure 30 f, the frequency of T to G and T to C edits in gTBEv3. (See figure)

[0155] In 30b-f, a single point represents a single replicate (n = 3 independent replicates / locus), and the box spans the interquartile range (25th to 75th percentile); the horizontal line within the box indicates the median (50%); and the whisker extends to the minimum and maximum values. Figure 30 gi, in Figure 3 At each edit site shown in a, T to G of gTBEv3 ( Figure 30 g), T to C ( Figure 30 h) or T to A( Figure 30 i) Percentage of edits (mean ± sem, n = 3). T number: T position with the highest mid-target base editing frequency in prototype spacer positions 1-20. Site number: Genome site number. Figure 30 j, in Figure 21 At the sites shown in a, the ratio of T to S edits by gTBEv3 to total T edits (base conversions and insertions / deletions). Figure 30 k. from ( Figure 30 Statistical analysis of target DNA base editing at each NT motif at the editing sites in (a). Each point represents the mean of three biological replicates for each editing position at each editing site. For motifs AT, CT, GT, TT, n = 8, 14, 6, 6.

[0156] Figure 31Off-target analysis of guided sequence dependence of gTBEv3 at more sites. Off-target analysis of guided sequence dependence of gTBEv3 editing at sites 9(a) and 15(b) (n=3). OT: Off-target. All values ​​are presented as mean ± sem.

[0157] Figure 32 Performance of the .UNG mutant in the context of gCBEv0.3. Figure 32 a. Percentage of EGFP+ cells with gCBE by introducing mutant A214V (n=3). Figure 32 b, Percentage of EGFP+ cells from gCBE with scan-induced mutagenesis of alanine from regions covering the catalytic water-activated loop and the Pro-rich loop (n=3). Valine was used to replace alanine (A to V) to cover all residues in the region of interest. Inactivated, catalytically inactive UNG2Δ88 (carrying D154N and H277N mutations, equivalent to D145N and H268N of UNG1). All values ​​are presented as mean ± sem.

[0158] Figure 33 Further characterization of the editing profile of .gCBEv2. Figure 33 a. By target depth sequencing, the frequency (mean ± sem, n = 3) of C-base editing outcomes using different gCBEs at the edited C2 position in site 28 in transfected HEK293T cells. Figure 33 b, shows a bar graph of mid-target DNA base editing (mean ± sem, n = 3) at two or more locations with a C conversion frequency >10% at each genomic locus in HEK293T cells. Figure 33 cd, from Figure 4 In the prototype spacer positions 1-20 of the edit site in c (where PAM is at positions 21-23), the C of gCBEv2 ( Figure 33 c) and T( Figure 33 d) Frequency of transformation. A single point represents a single replicate (n = 3 independent replicates / locus), and the box spans the interquartile range (25th to 75th percentile); the horizontal line within the box indicates the median (50%); and the whisker extends to the minimum and maximum values. Figure 33 e, in Figure 22 At the site shown in c, the ratio of C to G / T to C to A / G / T conversion frequencies edited by gCBEv2. Figure 33 fh, in Figure 3 At each edit site shown in a, C to G of gCBEv2 ( Figure 33 f), C to T ( Figure 33 g) or C to A ( Figure 33h) Percentage of edits (mean ± sem, n = 3). Figure 33 i. Insertion / deletion frequencies (mean ± sem, n = 3) of gCBEv2 were used at 16 target sites. Figure 33 In ei, C number: the C position with the highest mid-target base editing frequency among prototype spacer positions 1-20. Site number: the genomic site number. Figure 33 j, Statistical analysis of mid-target DNA base editing from each NC motif at 16 editing sites. Each point represents the mean of three biological replicates for each editing position at each editing site. For motifs AC, CC, GC, TC, n = 8, 9, 10, 13. Figure 33 k, from Figure 4 Insertion and deletion frequencies (mean ± sem, n = 3) of gCBEv2 and CGBE1 were used at the three intermediate target sites of d. Figure 33 l, For orthogonal R-loop assays, the mid-target base editing frequency of CGBE1 at C6 of site 22 in HEK293T cells (mean ± sem, n = 3).

[0159] Figure 34 .gTBEv3 performs base editing at splice sites. Figure 34 a. The best editing window for various base editors. Figure 34 b, shows Figure 23 Venn diagram of the sgRNA distribution of CBE and gCBEv2 in b. Figure 34 c. A bar graph showing the T or C conversion frequencies (mean ± sem, n = 3) at several splice acceptor (SA) or splice donor (SD) sites of interest targeted by gTBEv3 or gCBEv2. T or C number: the location of the T or C targeted in the prototype spacer positions 1-20. Figure 34 d. DNA sequencing chromatograms targeting the SD sites of exon 37 and exon 12 in human DMD with gTBEv3. The Sanger sequencing results were quantified using EditR.

[0160] Figure 35 PTC editing and importing of various base editors. Figure 35 a. The principle of bypassing the early stop codon (PTC) using various base editors. Figure 35 b. Possible codon endings from different base editors' stop codons (TAA, TAG, or TGA). Figure 35 c. The principle of introducing PTC using various base editors. Figure 35 d, the available codons used to edit into stop codons (TAA, TAG, or TGA) using different base editors. Figure 35e, a 10×10 dot plot, showing the percentage of potential sgRNAs (the number of available sgRNAs) for gene and cell therapy studies using gGBEv6.3 and CBE across 15 fully investigated genes (AGT, ANGPTL3, APOC3, B2M, CD33, DNMT3A, HPD, KLKB1, PCSK9, PDCD1, PRDM1, TGFBR2, TRAC, TTR, VEGFA) to introduce early stop codons (PTCs) by targeting different codons (the number of available sgRNAs is shown on the right). Figure 35 b and Figure 35 In d, AYBE, AYBEv3; gCBE, gCBEv2; gGBE, gGBEv6.3; gTBE, gTBEv3.

[0161] Figure 36 Further comparisons of different gTBEs. Figure 36 ab, T to G( Figure 36 a) or T to G( Figure 36 b) Frequency of conversion. The highest frequency (>20%) of edited thymine at the corresponding location is highlighted in the heatmap (n=3). Figure 36 cd shows the T-to-S and total base conversions for various base editors ( Figure 36 c) or T to S with total T editing (base conversion and insertion / deletion); Figure 36 Heatmap of the percentage of d) (n=3). Figure 36 e shows a heatmap of insertion and deletion frequencies for various base editors (n=3). Figure 36 fh, T base editing ( Figure 36 f, n=35), percentage from T to S ( Figure 36 g, n=35) and insertion / deletion ( Figure 36 Statistical analysis was performed on h, n=17. All values ​​are presented as mean ± sem. Each point represents the mean of three biological replicates for each edit location at each edit site. Figure 36 Ah, these images are from Figure 24 Data for various base editors are shown in c. Figure 36 In fh, the Dunnett multiple comparison test following one-way ANOVA is used to compare gTBEv3 or gTBEv5 with other base editors.

[0162] Figure 37 T editing in dsDNA upstream of the target site. Figure 37 a) shows a bar graph (n=3) of the T-transformation frequencies at -5T, -3T, -2T, and -1T in the dsDNA upstream of target site 38 (VEGFA). Figure 37b shows a bar graph (n=3) of the T-transformation frequencies at -4T, -3T, -2T, and -1T in the dsDNA upstream of target site 44 (EMX1-site H1F). All values ​​are presented as mean ± sem. These graphs are from [source missing]. Figure 24 Data for various base editors are shown in c.

[0163] Figure 38 A comparison of various glycosylation-based base editors used for cytosine editing. Figure 38 a. A schematic diagram of the basic architecture of various base editors. UNG2*, a UNG2 mutant from the corresponding base editor. ΔNTD, the deletion of the N-terminal domain. Figure 38 bc, total C-editing (base transformation and insertion / deletion) of various base editors at 19 endogenous loci. Figure 38 b) or C base conversion ( Figure 38 c) Frequency. This shows cytosines with an edit frequency >25% for any base editor. The highest frequency at the corresponding position is highlighted in the heatmap (n=3). Figure 38 d shows a heatmap of insertion and deletion frequencies for various base editors (n=3). Figure 38 e.g., Chief Editor ( Figure 38 e, n=56), C base conversion ( Figure 38 f, n=49) and insertion / deletion ( Figure 38 Statistical analysis was performed on gCBEv2 (n=19). All values ​​are presented as mean ± sem. Each point represents the mean of three biological replicates for each edit site. Dunnett's multiple comparison test following one-way ANOVA was used to compare gCBEv2 or gCBEv3 with other base editors.

[0164] Figure 39 Further comparison of various glycosylation-based base editors used for cytosine editing. Figure 39 ab shows the C to G frequencies of various base editors ( Figure 39 a) or C to G purity ( Figure 39 b)(n=3) heatmap. Cytosine is highlighted for any base editor C to G editing frequency >25%. Figure 39 c. C-base transition frequencies of various base editors in prototype spacer positions 1–20 (where PAM is at positions 21–23). A single point represents an individual repeat (n = 3 independent repeats / site), and the box spans the interquartile range (25th to 75th percentile); the horizontal line within the box indicates the median (50%); and the whisker lines extend to the minimum and maximum values. These figures are from [source missing]. Figure 38 Data for various base editors are shown in c.

[0165] Figure 40 Off-target analysis of various glycosylation enzyme-based base editors. Figure 40 ab, cumulative T-editing of various base editors at the corresponding site ( Figure 40 a) or C edit ( Figure 40 b) Off-target analysis of frequency-guided sequence dependence (n=3). OT: Off-target. Figure 40 c, guide sequence-independent cumulative off-target T-editing detected at each R-loop site by orthogonal R-loop determination (n=3). Figure 40 d. Statistical analysis of sequence-independent off-target T-editing (n=5). Each point represents the mean of three biological replicates for each edit site. gTBEv3 or gTBEv5 was compared with other base editors using Dunnett's multiple comparison test after one-way ANOVA. Figure 40 e. RNA off-target analysis of various base editors (n=3). mCherry was used as a control. D=A, G, or U; V=A, C, or G. For multiple comparisons, Dunnett's multiple comparison test following one-way ANOVA was used to compare gCBEv3 or gTBEv5 with other groups. For comparisons of RNA U to V SNVs of gTBEv5 and mCherry, a two-tailed unpaired two-sample t-test was used. All values ​​are presented as mean ± sem.

[0166] Figure 41 Comparison between .gTBE or gCBE and PE. Figure 41 a. Bar graph showing the mid-target T to G editing frequencies (n=3) of various editors at each genomic locus in HEK293T cells. OT: Off-target. Figure 41 b, a bar graph showing the mid-target C to G editing frequencies (n=3) of various editors at each genomic locus in HEK293T cells. PE6d was used in conjunction with epigRNA and nicking sgRNA. For PE6d max, PE6d was co-expressed with codon-optimized hMLH1dn (a dominant-negative MMR protein). All values ​​are presented as mean ± sem.

[0167] Figure 42 Characterization of gTBE or gCBE editing profiles in HEK293T, HuH-7, and U2OS cells. Figure 42 a) Bar graphs showing the mid-target T base editing frequency, T-to-G editing purity, or T-to-C editing purity (n=3) of gTBEv3 and gTBEv5 in different cell lines. Figure 42b. Bar graphs showing the mid-target C base editing frequency, C-to-G editing purity, or C-to-T editing purity (n=3) of gCBEv2 and gCBEv3 in different cell lines. HuH-7, a cell line established from human hepatocellular carcinoma; U2OS, a cell line established from human osteosarcoma. All values ​​are presented as mean ± sem.

[0168] The accompanying figures in this article are for illustrative purposes only and are not necessarily drawn to scale. Detailed Implementation

[0169] Overview

[0170] Traditional DNA base editors (such as ABE and CBE) can directly edit A and C with the aid of deamination. However, G and T can only be edited indirectly by directly editing C and T on the strand opposite to G and T. In contrast, at least part of the base editors and base editing methods provided in this disclosure include those capable of direct base editing of target deoxyribonucleotides (e.g., dG, dT, dC) in target dsDNA.

[0171] Traditional DNA base editors (such as ABE and CBE) require deaminases for editing A and C bases. In contrast, the base editors and base editing methods provided in this disclosure include, in at least part, base editing of target deoxyribonucleotides (e.g., dG, dT, dC) in target dsDNA in the absence of deaminases.

[0172] The base editor and base editing method disclosed herein rely on the base excision domain of the base editor. The base excision domain can directly excise the base of the target deoxyribonucleotide in the target dsDNA, creating a base-free site in situ, thereby triggering the base excision repair (BER) pathway. As a result of base excision and subsequent repair, the target deoxyribonucleotide can be converted into another deoxyribonucleotide, leading to base editing of the target deoxyribonucleotide.

[0173] The base editor and base editing method disclosed herein also rely on a nucleic acid programmable DNA-binding domain (napDNAbd) to specifically guide the base editor to the target dsDNA via a guide nucleic acid capable of interacting with both napDNAbd and the target dsDNA. napDNAbd can be associated (e.g., complexed) with a guide nucleic acid (e.g., guide RNA). The guide nucleic acid is programmed to localize or target napDNAbd to the target dsDNA via hybridization between a target sequence of the target dsDNA and a corresponding guide sequence of the guide nucleic acid. Specifically, the guide nucleic acid contains a guide sequence that is capable of hybridizing with the target sequence of the target dsDNA due to substantial complementarity between the guide sequence and the target sequence. The guide nucleic acid also contains a scaffold sequence capable of forming a complex with napDNAbd. In this way, the guide nucleic acid “programs” napDNAbd so that napDNAbd can be specifically localized via the guide nucleic acid and (indirectly) bound to the target sequence of the target dsDNA and to regions surrounding the target sequence. The binding of napDNAbd to the target dsDNA enables the base excision domain associated with napDNAbd to specifically approach and act on the bases of the target deoxyribonucleotide in the target sequence of the target dsDNA in a guided sequence-specific / dependent manner.

[0174] refer to Figure 1 a, 6-8, and 19a, the base-removal domain of the base editor of this disclosure directly removes the target deoxyribonucleotide ( Figure 6 The first deoxyribonucleotide in DNA is removed to create a base-free site, which is either a purine-free site (e.g., guanine) or a pyrimidine-free site (e.g., thymine, cytosine). Base-free sites can be repaired via trans-damage synthesis (TLS) (by, for example, TLS polymerase) and / or DNA replication, resulting in base editing in which case a nick may not be necessary in the strand opposite the base-free site.

[0175] Alternatively, if a nucleic acid programmable DNA cleavage enzyme (such as Cas9 cleavage enzyme) is used as napDNAbd, it creates a cleavage in the strand opposite the baseless site (the non-edited strand), and the pyrimidine-free site can be removed by AP lyase to create another cleavage in the edited strand. These two cleavages trigger double-strand break (DSB) repair and the introduction of insertion / deletion mutations, which also result in a high potential change in the target deoxyribonucleotide.

[0176] As a specific example Figure 6This diagram illustrates the first deoxyribonucleotide of dG on the edit strand (non-target strand) as the target deoxyribonucleotide to be edited, and the second deoxyribonucleotide of dC on the opposing strand (non-editing strand / target strand) that pairs with the dG base, prior to base editing. dG is located in the prototype spacer sequence on the non-target strand of the target dsDNA, and dC is located in the target sequence on the target strand of the target dsDNA. The guiding nucleic acid is designed to contain a guiding sequence capable of hybridizing with the target sequence and a scaffold sequence capable of forming a complex with napDNAbd. napDNAbd enables nicks to be created on the target strand and fuses with a base-removing domain capable of excising the guanine of dG. As a specific example, Figure 7 The diagram shows the fourth deoxyribonucleotide of dC, which is the final deoxyribonucleotide on the edited strand (non-target strand), and the third deoxyribonucleotide of dG, which pairs with the base of dC on the opposing strand (non-edited strand / target strand), after base editing. Figure 6 and Figure 7 Together, they illustrate direct dG to dC base editing. As a concrete example, Figure 8 The diagram shows the fourth deoxyribonucleotide of dT, which is the final deoxyribonucleotide on the edited strand (non-target strand), and the third deoxyribonucleotide of dA, which pairs with the dT base on the opposing strand (non-edited strand / target strand), after base editing. Figure 6 and Figure 8 Together they demonstrate direct dG to dT base editing.

[0177] The base editing method disclosed herein allows for direct base editing of target deoxyribonucleotides (e.g., dG, dT, dC) in target dsDNA, thereby expanding the scope of target design and screening for direct base editing. For example, if the goal is to edit target dG to dA, a conventional base editor that cannot directly edit dG must be applied to edit the opposite strand dC to dT, thus indirectly editing dG to dA. However, the editing capabilities of conventional CBEs may not be able to edit dC to the desired outcome dT; the PAM limitations of CBEs may not allow for the design of target / guide sequences targeting dC to specifically guide the CBE to dC; and even if such guide sequences could be designed, the base editing efficiency of CBEs may be insufficient. Using the base editing method disclosed herein, direct base editing of target dG is possible, and therefore, developers will have significantly more opportunities to design, screen, and obtain suitable target / guide sequences targeting dG to specifically guide the base editor of this disclosure to dG.

[0178] The base editing method disclosed herein can work in the absence of deamination of the target deoxyribonucleotide base before the base of the target deoxyribonucleotide is removed, or in the absence of deamination at all. In contrast, conventional ABE requires deamination of the target dA base (adenine) to hypoxanthine, thereby converting the target dA to inosine (I), which is read as dG in DNA repair replication; and conventional CBE requires deamination of the target dC base (cytosine) to uracil, thereby converting the target dC to uridine (U), which is read as dT in DNA repair or replication. Deamination is unlikely for G (due to spontaneous repair)

[18] and impossible for T (due to the absence of amine), making the development of deaminase-based G and T base editors a challenging task. The omission of deaminase domains for G and T in the base editor disclosed herein paves the way for base editing, and also avoids the undesirable effects caused by deaminase domains and deamination in traditional deaminase-based base editors, while reducing the size of the base editor.

[0179] As a non-limiting example of this disclosure, a G-editing-capable, deaminase-free, glycosylase-based guanine base editor (gGBE) was developed by fusing a nucleic acid-programmable DNA cleavage enzyme (such as Cas9 cleavage enzyme) with a human N-methylpurine DNA glycosyltransferase (MPG) mutant capable of cleaving dG guanine. This mutant was developed through several rounds of MPG mutagenesis via unbiased and rational screening. gGBE demonstrated high G-editing efficiency. Furthermore, gGBE exhibited high base editing efficiency (up to 81.2%) and high G-to-T or G-to-C (i.e., G-to-Y) conversion ratios (up to 0.95) in both cultured human cells and mouse embryos. The editing profile of gGBE was characterized by targeting dozens of endogenous genomic loci in cultured mammalian cells and mouse embryos, demonstrating its high G-to-Y (Y = C or T) base editing efficiency.

[0180] As another non-limiting example of this disclosure, two deaminase-free, glycosylation-based base editors for direct T-editing (gTBE) and direct C-editing (gCBE) were developed to achieve orthogonal base editing by fusing a nucleic acid-programmable DNA cleavage enzyme (such as Cas9 cleavage enzyme) with a human uracil DNA glycosylation enzyme (UNG) mutant developed separately via UNG mutagenesis capable of cleaving dT thymine or dC cytosine. These are gTBE for direct T-editing and gCBE for direct C-editing, respectively. Through several rounds of structure-informed, rational mutagenesis of UNG in cultured human cells, highly active gTBEs and gCBEs exhibiting T-to-S (i.e., T-to-C or T-to-G) and C-to-G conversions, respectively, were obtained. Furthermore, by embedding the UNG mutant into a nucleic acid-programmable DNA cleavage enzyme (such as Cas9 cleavage enzyme), more gTBEs and gCBEs were generated, exhibiting enhanced average editing efficiency and selective editing windows. The editing profiles of gTBE and gCBE were characterized by targeting dozens of endogenous genomic loci in cultured mammalian cells and mouse embryos, demonstrating their high base editing efficiency.

[0181] Base Editor

[0182] The base editor disclosed herein can be provided in the form of a fusion protein. In one aspect, this disclosure provides a fusion protein comprising:

[0183] (1) A nucleic acid programmable DNA-binding domain (napDNAbd) capable of binding to target dsDNA, wherein the target dsDNA comprises:

[0184] (a) The first deoxyribonucleotide (e.g., dG (deoxyguanosine), dT (thymidine), dC (deoxycytidine)) in the prototype spacer sequence on the non-target strand (editing strand) of the target dsDNA, and

[0185] (b) a second deoxyribonucleotide (e.g., dC (deoxycytidine), dA (deoxyadenosine), dG (deoxyguanosine)) in a target sequence paired with the first deoxyribonucleotide (e.g., dG, dT, dC) and located on the target strand (non-edited strand) of the target dsDNA, wherein the prototype spacer sequence is completely inversely complementary to the target sequence; and

[0186] (2) A base-removal domain capable of removing a base of the first deoxyribonucleotide (e.g., guanine, thymine, cytosine).

[0187] In some embodiments, the fusion protein does not contain a deaminase domain, such as an adenine or cytosine deaminase domain, such as TadA and its variants.

[0188] The components of the fusion protein are described in more detail in other subsections of this paper.

[0189] Base editing system

[0190] On the other hand, this disclosure provides a system comprising:

[0191] (i) a fusion protein or a polynucleotide encoding the fusion protein, the fusion protein comprising:

[0192] (1) A nucleic acid programmable DNA-binding domain (napDNAbd) capable of binding to target dsDNA, wherein the target dsDNA comprises:

[0193] (a) The first deoxyribonucleotide (e.g., dG (deoxyguanosine), dT (thymidine), dC (deoxycytidine)) in the prototype spacer sequence on the non-target strand (editing strand) of the target dsDNA, and

[0194] (b) a second deoxyribonucleotide (e.g., dC (deoxycytidine), dA (deoxyadenosine), dG (deoxyguanosine)) in a target sequence paired with the first deoxyribonucleotide (e.g., dG, dT, dC) and located on the target strand (non-edited strand) of the target dsDNA, wherein the prototype spacer sequence is completely inversely complementary to the target sequence; and

[0195] (2) A base-removing domain capable of cleaving a base (e.g., guanine, thymine, cytosine) of the first deoxyribonucleotide; and

[0196] (ii) a guiding nucleic acid or a polynucleotide encoding the guiding nucleic acid, the guiding nucleic acid comprising:

[0197] (1) A scaffold sequence capable of forming a complex with the napDNAbd; and

[0198] (2) It is capable of hybridizing with the target sequence on the target strand of the target dsDNA, thereby directing the complex to the guide sequence of the target dsDNA.

[0199] In some embodiments, the system is a complex comprising a fusion protein that complexes with a guiding nucleic acid. In some embodiments, the complex further comprises target dsDNA that hybridizes with a guiding sequence. In some embodiments, the system is a composition comprising component (i) and component (ii).

[0200] In some implementations, the guide nucleic acid, as described herein, is guide RNA (gRNA). As used herein, the term "gRNA" is used interchangeably with single guide RNA (sgRNA).

[0201] The system’s components are described in more detail in other subsections of this paper.

[0202] Base editing methods

[0203] In another aspect, this disclosure provides a method for modifying target dsDNA, the method comprising contacting the target dsDNA with a system.

[0204] The target dsDNA contains:

[0205] (a) The first deoxyribonucleotide (e.g., dG (deoxyguanosine), dT (thymidine), dC (deoxycytidine)) in the prototype spacer sequence on the non-target strand (editing strand) of the target dsDNA, and

[0206] (b) a second deoxyribonucleotide (e.g., dC (deoxycytidine), dA (deoxyadenosine), dG (deoxyguanosine)) in a target sequence paired with the first deoxyribonucleotide (e.g., dG, dT, dC) and located on the target strand (non-edited strand) of the target dsDNA, wherein the prototype spacer sequence is completely inversely complementary to the target sequence; and

[0207] The system includes:

[0208] (i) a fusion protein or a polynucleotide encoding the fusion protein, the fusion protein comprising:

[0209] (1) Capable of binding the target dsDNA to a programmable DNA-binding domain (napDNAbd); and

[0210] (2) A base-removing domain capable of cleaving a base (e.g., guanine, thymine, cytosine) of the first deoxyribonucleotide; and

[0211] (ii) a guiding nucleic acid or a polynucleotide encoding the guiding nucleic acid, the guiding nucleic acid comprising:

[0212] (1) A scaffold sequence capable of forming a complex with the napDNAbd; and

[0213] (2) It is capable of hybridizing with the target sequence on the target strand of the target dsDNA, thereby directing the complex to the guide sequence of the target dsDNA.

[0214] In some embodiments, the method does not include deamination of the bases of the first deoxyribonucleotide before removing the bases of the first deoxyribonucleotide.

[0215] In some embodiments, the method does not include deamination of the bases of the first deoxyribonucleotide.

[0216] In some embodiments, the method includes inducing strand separation of the target dsDNA.

[0217] The components of the method are described in more detail in other subsections of this document.

[0218] Target deoxyribonucleotides and target dsDNA

[0219] refer to Figure 6-8 The target deoxyribonucleotide to be edited in the target dsDNA can be called the "first deoxyribonucleotide", and the resulting deoxyribonucleotide transformed from the target deoxyribonucleotide through base editing can be called the "fourth deoxyribonucleotide".

[0220] In some embodiments, the first deoxyribonucleotide is deoxyguanosine (dG), thymidine (dT), deoxyadenosine (dA), or deoxycytidine (dC). In some embodiments, the first deoxyribonucleotide is dG. In some embodiments, the first deoxyribonucleotide is dT. In some embodiments, the fourth deoxyribonucleotide is dA, dT, dC, or dG. In some embodiments, the second deoxyribonucleotide is dC, dA, dT, or dG. In some embodiments, the third deoxyribonucleotide is dA, dT, dC, or dG.

[0221] As desired, in some embodiments, the first deoxyribonucleotide is converted into a fourth deoxyribonucleotide, which is different from the first deoxyribonucleotide. In some embodiments, the conversion from the first deoxyribonucleotide to the fourth deoxyribonucleotide is dG to dA, dG to dT, dG to dC, dT to dA, dT to dC, dT to dG, dC to dA, dC to dT, or dC to dG.

[0222] In some embodiments, the target dsDNA is wild-type or naturally occurring. In some embodiments, the target dsDNA is neither wild-type nor naturally occurring. In some embodiments, the target dsDNA is eukaryotic or prokaryotic. In some embodiments, the target dsDNA is derived from an animal (e.g., human, monkey, mouse) or plant. In some embodiments, the target dsDNA is a target gene. In some embodiments, the gene is an animal (e.g., human, monkey, mouse) gene or a plant gene. In some embodiments, the dsDNA is in a target cell.

[0223] In some embodiments, the first deoxyribonucleotide is natural or non-natural with respect to the target dsDNA. In some embodiments, the first deoxyribonucleotide is a mutation in the target dsDNA. In some embodiments, the first deoxyribonucleotide is a pathogenic mutation in the target dsDNA. In some embodiments, the first deoxyribonucleotide is a mutation in the target dsDNA that produces a stop codon.

[0224] In some embodiments, the conversion from the first deoxyribonucleotide to the fourth deoxyribonucleotide directly or indirectly converts a stop codon on the target or non-target strand into a non-stop codon, or directly or indirectly converts a non-stop codon on the target or non-target strand into a stop codon. In some embodiments, the stop codon is located on the sense strand of the dsDNA.

[0225] In some embodiments, the conversion from the first deoxyribonucleotide to the fourth deoxyribonucleotide occurs on the sense or senseless strand of the dsDNA. In some embodiments, the conversion from the first deoxyribonucleotide to the fourth deoxyribonucleotide (e.g., dG to dT) occurs on the sense strand of the dsDNA, converting a stop codon on the sense strand to a non-stop codon, or converting a non-stop codon (e.g., GAA) on the sense strand to a stop codon (e.g., TAA). In some embodiments, the conversion from the first deoxyribonucleotide to the fourth deoxyribonucleotide (e.g., dG to dC) occurs on the senseless strand of the dsDNA, converting a stop codon on the sense strand to a non-stop codon, or converting a non-stop codon (e.g., TCA) on the sense strand to a stop codon (e.g., TGA).

[0226] In some embodiments, the conversion from a first deoxyribonucleotide to a fourth deoxyribonucleotide (e.g., dG to dC) occurs at the splicing site (e.g., splice donor, splice acceptor) of the target dsDNA. In some embodiments, the conversion from a first deoxyribonucleotide to a fourth deoxyribonucleotide (e.g., dG to dC) occurring at the splicing site (e.g., splice donor, splice acceptor) increases or decreases the translation of transcripts transcribed from the target dsDNA.

[0227] In some embodiments, the first deoxyribonucleotide is located in the prototype spacer sequence at a position selected from the following: position 1, position 2, position 3, position 4, position 5, position 6, position 7, position 8, position 9, position 10, position 11, position 12, position 13, position 14, position 15, position 16, position 17, position 18, position 19, position 20, and combinations thereof; or wherein the first deoxyribonucleotide is located in the prototype spacer sequence between position 1 and position 20 (including both end values); or wherein the first deoxyribonucleotide is located in the prototype spacer sequence between position 1 and position 14 (including both end values). In some embodiments, the first deoxyribonucleotide is located in the prototype spacer sequence at a position selected from the following: position 6, position 7, position 8, position 9, position 10, position 11, and combinations thereof; or wherein the first deoxyribonucleotide is located in the prototype spacer sequence between position 6 and position 11 (including both end values). In some implementations, the first deoxyribonucleotide is located at position 7 of the prototype spacer sequence.

[0228] In some embodiments, the first deoxyribonucleotide is either N1 or N2 nucleotide in the motif of N1N2, wherein N1 or N2 is A, T, G, or C. In some embodiments, the first deoxyribonucleotide is N2 nucleotide in the motif of N1N2, wherein N1 is A or T, and N2 is C. In some embodiments, the first deoxyribonucleotide is either N1, N2, or N3 nucleotide in the motif of N1N2N3, wherein N1, N2, or N3 is A, T, G, or C.

[0229] Base excision domain

[0230] As used herein, the term “base-removing domain (BED)” is used interchangeably with “base-removing protein (BEP)” or “base-removing enzyme (BEE)” and refers to a protein capable of recognizing and removing bases (e.g., A, T, C, G, or U) of nucleotides in nucleic acids (e.g., DNA (ssDNA or dsDNA) or RNA). As used herein, the term “base” is used interchangeably with “nucleobase” or “nitrogenous base”. Bases include, for example, adenine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), and they may be referred to as primary bases, normal bases, or canonical bases. As is well known in the art, deoxyribonucleotides consist of a base, a deoxyribose, and a phosphate ester, and deoxyribonucleotides consist of a base and a deoxyribose. Removing a base from a deoxyribonucleotide releases the base from the deoxyribonucleotide. In some embodiments, removing the base of a deoxyribonucleoside includes cleaving or hydrolyzing the glycosidic bond that links the base to the deoxyribose of the first deoxyribonucleotide, thereby releasing the base from the first deoxyribonucleotide.

[0231] In some embodiments, the base-removing domain is (substantially) capable of cleaving a base (e.g., guanine, thymine, cytosine) of a first deoxyribonucleotide (e.g., dG, dT, dC). In some embodiments, the first deoxyribonucleotide is dG, dT, dA, or dC. In some embodiments, the base-removing domain is (substantially) capable of cleaving guanine in dG. In some embodiments, the base-removing domain is (substantially) capable of cleaving thymine in dT. In some embodiments, the base-removing domain is (substantially) capable of cleaving cytosine in dC. In some embodiments, the base-removing domain is (substantially) capable of cleaving adenine in dA.

[0232] In some cases, highly specific base editing is desired, for example, in disease treatment. Therefore, it is desirable for the base-removing domain to remove only one type of base without removing other types. In some embodiments, the base-removing domain (substantially) cannot remove guanine from dG. In some embodiments, the base-removing domain (substantially) cannot remove thymine from dT. In some embodiments, the base-removing domain (substantially) cannot remove cytosine from dC. In some embodiments, the base-removing domain (substantially) cannot remove adenine from dA.

[0233] In some cases, multiplexed base editing is desired. Therefore, it is desirable for the base-removing domain to be able to remove more than one type of base. In some embodiments, the base-removing domain is (essentially) capable of removing any two, three, or four of the following: guanine in dG, thymine in dT, cytosine in dC, and adenine in dA. For example, in some embodiments, the base-removing domain is (essentially) capable of removing both guanine in dG and thymine in dT.

[0234] In some embodiments, the base-removing domain is (substantially) capable of cleaving uracil. In some embodiments, the base-removing domain is (substantially) not capable of cleaving uracil. In some embodiments, the base-removing domain is (substantially) capable of cleaving hypoxanthine. In some embodiments, the base-removing domain is (substantially) not capable of cleaving hypoxanthine.

[0235] In some embodiments, the fusion protein of this disclosure does not contain (substantially) base-removing domains capable of cleaving guanine in dG, thymine in dT, cytosine in dC, adenine, uracil, and / or hypoxanthine in dA.

[0236] In some embodiments, the base-removing domain (substantially) cannot remove bases from either strand of the target dsDNA. In some embodiments, the base-removing domain (substantially) cannot remove two bases from a pair of base-paired deoxyribonucleotides on either strand of the dsDNA.

[0237] In some embodiments, the base-removing domain comprises an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with a naturally occurring base-removing domain (such as the naturally occurring base-removing domains provided herein).

[0238] In some embodiments, the fusion protein of this disclosure comprises one, two, three or more base-removing domains. In some embodiments, the fusion protein comprises two, three or more identical or different base-removing domains.

[0239] The base-removal domain can be a glycosylation enzyme with the desired base-removal capability. In some embodiments, the base-removal domain comprises a glycosylation enzyme. In some embodiments, the glycosylation enzyme is selected from N-methylpurine DNA glycosylation enzyme (MPG), 8-oxoguanine DNA glycosylation enzyme (OGG1), methyl-CpG-binding domain 4 DNA glycosylation enzyme (MBD4), thymine DNA glycosylation enzyme (TDG), uracil DNA glycosylation enzyme (UNG), single-stranded selective monofunctional uracil DNA glycosylation enzyme 1 (SMUG1), mutY DNA glycosylation enzyme (MUTYH), nth-like DNA glycosylation enzyme 1 (NTHL1), nei-like DNA glycosylation enzyme 1 (NEIL1), nei-like DNA glycosylation enzyme 2 (NEIL2), nei-like DNA glycosylation enzyme 3 (NEIL3), and mutants thereof capable of recognizing and removing bases from nucleotides of nucleic acids.

[0240] Exemplary glycosylases capable of base removal include, but are not limited to, UDG-N204D and UDG-Y147A, as described in the following reference: Kavli, B et al., Excision of cytosine and thymine from DNA by mutants of human uracil-DNA glycosylase. EMBO J [Journal of the European Society for Molecular Biology] 15, 3442-3447 (1996); the entire contents of which are hereby incorporated by reference.

[0241] In some implementations, the base-excision domain is not wild-type or naturally occurring.

[0242] MPG

[0243] In some embodiments, the base-excision domain comprises an N-methylpurine DNA glycosylation enzyme (MPG). In some embodiments, the MPG comprises the motif GxxYxxxxYGxxxxxN, where x represents any amino acid. In some embodiments, the MPG is obtained from a species selected from Table A. In some embodiments, the MPG contains an amino acid mutation relative to (compared to; with respect to) the wild-type or reference MPG. In some embodiments, the wild-type or reference MPG is a human MPG (SEQ ID NO:9), or an MPG obtained from a species selected from Table A, or any MPG as shown in Table D, or a homolog or mutant thereof (e.g., containing the amino acid sequence of SEQ ID NO:5, 6, or 7), or an N-terminal truncated form lacking the N-terminal methionine (Met) (encoded by the start codon ATG) (e.g., SEQ ID NO:1). In some embodiments, the wild-type or reference MPG contains the amino acid sequence of SEQ ID NO:7.

[0244] In some embodiments, the amino acid mutation comprises amino acid substitutions at positions selected from the following: wild type or reference MPG 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65. 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 1 68, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 19 9, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297 and / or 298, where the position is numbered according to SEQ ID NO:1.

[0245] In some embodiments, the amino acid mutation confers the ability to remove a base from MPG. In some embodiments, this base is guanine.

[0246] In some embodiments, compared to a otherwise identical control MPG without the stated amino acid mutation, the mutation results in an increased base-excision capacity, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, or 160%. 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0247] In some embodiments, compared to a otherwise identical control fusion protein without the said amino acid mutation, the amino acid mutation results in an increase in the guide sequence-specific base editing efficiency of the fusion protein, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, or 150%. %, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0248] In some embodiments, the amino acid mutation includes amino acid substitutions at positions corresponding to or selected from the following positions: wild type or reference MPG: G163, N169, D175, C178, S198, K202, G203, S206, K210 and / or Q294, wherein the position is numbered according to SEQ ID NO:1.

[0249] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position corresponding to a position selected from or equivalent to N169, D175, C178, and / or Q294 of the wild-type or reference MPG, wherein the position is numbered according to SEQ ID NO:1. In some embodiments, the wild-type or reference MPG comprises the amino acid sequence of SEQ ID NO:7.

[0250] In some embodiments, the amino acid substitution is either a conserved or non-conserved amino acid substitution. In some embodiments, the amino acid substitution is performed using an amino acid residue that is different from the amino acid residue at that position in the wild type or the reference MPG. In some embodiments, amino acid substitution is performed using the following amino acid substitutions: (1) nonpolar amino acid residues (such as glycine (Gly / G), alanine (Ala / A), valine (Val / V), cysteine ​​(Cys / C), proline (Pro / P), leucine (Leu / L), isoleucine (Ile / I), methionine (Met / M), tryptophan (Trp / W), phenylalanine (Phe / F); (2) polar amino acid residues (such as serine (Ser / S), threonine (Thr / T), tyrosine (Tyr / Y), asparagine (Asn / N), glutamine (Gln / Q)); (3) positively charged amino acid residues (such as lysine (Lys / K), arginine (Arg / R), histidine (His / H)); or (4) negatively charged amino acid residues (such as aspartic acid (Asp / D), glutamic acid (Glu / E)). In some embodiments, amino acid substitution is performed using R, A, N, or G.

[0251] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: G163R, N169G, D175R, C178N, S198A, K202A, G203A, S206A, K210A, Q294R, and the position numbered according to SEQ ID NO:1.

[0252] In some embodiments, the MPG containing the amino acid mutation comprises, substantially comprises, or consists of the following: an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100% sequence identity with the amino acid sequence of the wild-type or reference MPG.

[0253] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: N169G, D175R, C178N, Q294R, wherein the position is numbered according to SEQ ID NO:1. In some embodiments, the wild type or reference MPG comprises the amino acid sequence of SEQ ID NO:7.

[0254] In some embodiments, the amino acid mutation comprises a combination of substitutions corresponding to the combination of N169G, D175R, C178N, and Q294R, wherein the position is numbered according to SEQ ID NO:1. In some embodiments, the wild-type or reference MPG comprises the amino acid sequence of SEQ ID NO:7.

[0255] In some embodiments, the amino acid mutation comprises a combination of substitutions corresponding to the combination of G163R, N169G, D175R, C178N, S198A, K202A, G203A, S206A, K210A, and Q294R, wherein the position is numbered according to SEQ ID NO:1. In some embodiments, the wild-type or reference MPG comprises the amino acid sequence of SEQ ID NO:1 or 9.

[0256] In some embodiments, the MPG containing the amino acid mutation comprises, substantially comprises, or consists of an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO:8, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38) sequence identity.

[0257] In some embodiments, MPG is (substantially) capable of cleaving guanine in dG. In some embodiments, MPG is (substantially) incapacitated by thymine in dT. In some embodiments, MPG is (substantially) incapacitated by cytosine in dC. In some embodiments, MPG is (substantially) incapacitated by adenine in dA.

[0258] In some implementations, MPG is neither wild-type nor naturally occurring.

[0259] UNG

[0260] In some embodiments, the base-excision domain comprises uracil DNA glycosyltransferase (UNG). In some embodiments, the UNG comprises the motif GQDPYH. In some embodiments, the UNG is obtained from species selected from Table C. In some embodiments, the UNG comprises an amino acid mutation relative to (compared to; with respect to) the wild-type or reference UNG. In some embodiments, the wild-type or reference UNG is human UNG1 (SEQ ID NO: 54) or human UNG2 (SEQ ID NO: 133), or an UNG obtained from species selected from Table C, or any UNG as shown in Table D, or its homologs or mutants, or its N-terminal truncated form lacking the N-terminal methionine (Met) (encoded by the start codon ATG). In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO: 56, 58, 135, or 137.

[0261] In some embodiments, the amino acid mutation comprises amino acid substitutions at positions selected from the following: wild type or reference UNG 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65. 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 1 68, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 19 9, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312 and / or 313, wherein the position is numbered according to SEQ ID NO: 133.

[0262] Alternative splicing and transcription from two different initiation sites result in two distinct human UNG isoforms: mitochondrial UNG1 (304 amino acids (aa)) (SEQ ID NO:54) and nuclear UNG2 (313 aa) (SEQ ID NO:133), each possessing a unique N-terminus that mediates translocation to the mitochondria or nucleus. 16 ( Figure 25 Sequence alignment of human UNG1 and human UNG2 shows that amino acid residues at positions 1-35 of UNG1 are different from amino acid residues at positions 1-44 of UNG2, and the rest of UNG1 and UNG2 are identical. The correspondence between the positions of UNG1 (SEQ ID NO: 54) and UNG2 (SEQ ID NO: 133) can be determined by sequence alignment. For example, residues Y156 and N213 of UNG2 correspond to residues Y147 and N204 of UNG1, respectively.

[0263] In some embodiments, the amino acid mutation confers the ability to remove the base at UNG. In some embodiments, the base is thymine. In some embodiments, the base is cytosine.

[0264] In some embodiments, compared to a otherwise identical control UNG without the stated amino acid mutation, the amino acid mutation results in an increased base-excision capacity, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, or 160%. 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0265] In some embodiments, compared to a otherwise identical control fusion protein without the said amino acid mutation, the amino acid mutation results in an increase in the guide sequence-specific base editing efficiency of the fusion protein, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, or 150%. %, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0266] In some embodiments, the amino acid mutation includes an amino acid substitution at a position corresponding to a position selected from or selected from the following positions: wild type or reference UNG Y156, K184, N213, A214, Q259 and / or Y284, wherein the position is numbered according to SEQ ID NO:133.

[0267] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position corresponding to a position selected from or selected from the wild-type or reference UNG, specifically K184, A214, Q259, and / or Y284, wherein the position is numbered according to SEQ ID NO: 133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO: 135 or 137.

[0268] In some embodiments, the amino acid substitution is either a conserved or non-conserved amino acid substitution. In some embodiments, the amino acid substitution is performed using an amino acid residue that is different from the amino acid residue at that position in the wild type or the reference UNG. In some embodiments, the amino acid substitution is performed using the following amino acid substitutions: (1) nonpolar amino acid residues (such as glycine (Gly / G), alanine (Ala / A), valine (Val / V), cysteine ​​(Cys / C), proline (Pro / P), leucine (Leu / L), isoleucine (Ile / I), methionine (Met / M), tryptophan (Trp / W), phenylalanine (Phe / F); (2) polar amino acid residues (such as serine (Ser / S), threonine (Thr / T), tyrosine (Tyr / Y), asparagine (Asn / N), glutamine (Gln / Q)); (3) positively charged amino acid residues (such as lysine (Lys / K), arginine (Arg / R), histidine (His / H)); or (4) negatively charged amino acid residues (such as aspartic acid (Asp / D), glutamic acid (Glu / E)). In some embodiments, the amino acid substitution is performed using A, D, V, or T.

[0269] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: Y156A, K184A, N213D, A214V, A214T, Q259A, Y284D, and the position numbered according to SEQ ID NO:133.

[0270] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: A214T, Q259A, Y284D, and the position specified in SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:135.

[0271] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of K184A, A214V, and two such substitutions, wherein the position is designated according to SEQ ID NO: 133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO: 137.

[0272] In some embodiments, the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or corresponding to one of the following positions: wild type or reference UNG 1-65, 1-66, 1-67, 1-68, 1-69, 1-70, 1-71, 1-72, 1-73, 1-74, 1-75, 1-76, 1-77, 1-78, 1-79, 1-80, 1-81, 1-82, 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, or 1-100, wherein the position is numbered according to SEQ ID NO:133.

[0273] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: A214T, Q259A, Y284D, and any two or more thereof, and the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or at positions 1-88 of the wild-type or reference UNG, wherein the position is numbered according to SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:135.

[0274] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of K184A, A214V, and two such substitutions, and the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or a position 1-88 of the wild-type or reference UNG, wherein the position is numbered according to SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:137.

[0275] In some embodiments, the UNG containing the amino acid mutation comprises, is substantially composed of, or consists of, an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100% sequence identity with the amino acid sequence of the wild-type or reference UNG.

[0276] In some embodiments, the amino acid mutation comprises a combination of substitutions corresponding to the combination of A214T, Q259A, and Y284D, and the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or corresponding to positions 1-88 of the wild-type or reference UNG, wherein the position is numbered according to SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:135.

[0277] In some embodiments, the amino acid mutation comprises a combination substitution corresponding to a combination substitution of K184A and A214V, and the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or corresponding to positions 1-88 of the wild-type or reference UNG, wherein the position is numbered according to SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:137.

[0278] In some embodiments, the UNG containing the amino acid mutation comprises, substantially consists of, or consists of: with SEQ ID The amino acid sequence of NO:56, 58, 60, 62, 135, 137, 139, 141, 144, 146, 148, 150, 152, 155, 157, and 159, or the N-terminal truncated amino acid sequence lacking the most N-terminal methionine (M) (encoded by the start codon ATG), has at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity.

[0279] In some embodiments, the UNG (substantially) can cleave thymine from dT. In some embodiments, the UNG (substantially) can cleave cytosine from dC. In some embodiments, the UNG (substantially) cannot cleave thymine from dT. In some embodiments, the UNG (substantially) cannot cleave cytosine from dC. In some embodiments, the UNG (substantially) cannot cleave adenine from dA. In some embodiments, the UNG (substantially) cannot cleave guanine from dG.

[0280] In some implementations, UNG is neither wild-type nor naturally occurring.

[0281] Other BED

[0282] In some embodiments, the base-removing domain comprises TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 contains an amino acid mutation relative to (compared to; with respect to) wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1. In some embodiments, wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3 or NTHL1 is human TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3 or NTHL1 (SEQ ID NO: 64, 65, 66, 67, 68, 69, 70, 71 or 72 respectively) or its homologs or mutants or its N-terminal truncated form lacking the N-terminal methionine (Met) (encoded by the start codon ATG). In some embodiments, the wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 each comprises an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with any of SEQ ID NO:64-72.

[0283] In some embodiments, the amino acid mutation includes amino acid substitutions at positions selected from the following: wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 at positions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50. 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 1 88, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 21 9, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250251、252、253、254、255、256、257、258、259、260、261、262、263、264、265、266、267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, and / or 604, wherein the position is based on any one of the numbers in SEQ ID NO: 64-72.

[0284] In some embodiments, the amino acid mutation confers the ability to remove bases from TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1. In some embodiments, the bases are guanine, thymine, cytosine, adenine, uracil, or hypoxanthine.

[0285] In some embodiments, compared to otherwise identical controls TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 without the said amino acid mutation, the amino acid mutation results in an increased base-cutting capacity, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, or 1... 20%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0286] In some embodiments, compared to a otherwise identical control fusion protein without the said amino acid mutation, the amino acid mutation results in an increase in the guide sequence-specific base editing efficiency of the fusion protein, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, or 150%. %, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0287] In some embodiments, the amino acid substitution is either a conserved amino acid substitution or a non-conserved amino acid substitution. In some embodiments, the amino acid substitution is an amino acid substitution made with an amino acid residue that is different from the amino acid residue at that position of the wild type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3 or NTHL1. In some embodiments, amino acid substitution is performed using the following amino acid substitutions: (1) nonpolar amino acid residues (such as glycine (Gly / G), alanine (Ala / A), valine (Val / V), cysteine ​​(Cys / C), proline (Pro / P), leucine (Leu / L), isoleucine (Ile / I), methionine (Met / M), tryptophan (Trp / W), phenylalanine (Phe / F); (2) polar amino acid residues (such as serine (Ser / S), threonine (Thr / T), tyrosine (Tyr / Y), asparagine (Asn / N), glutamine (Gln / Q)); (3) positively charged amino acid residues (such as lysine (Lys / K), arginine (Arg / R), histidine (His / H)); or (4) negatively charged amino acid residues (such as aspartic acid (Asp / D), glutamic acid (Glu / E)).

[0288] In some embodiments, the TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 containing the said amino acid mutation comprises, substantially comprises, or comprises the following: an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100% sequence identity with the wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1.

[0289] In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cleaves guanine from dG. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cleaves thymine from dT. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cleaves cytosine from dC. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cleaves adenine from dA.

[0290] In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cannot cleave thymine in dT. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cannot cleave cytosine in dC. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cannot cleave adenine in dA. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cannot cleave guanine in dG.

[0291] In some implementations, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 are not wild-type or naturally occurring.

[0292] napDNAbd

[0293] As used herein, the term “napDNAbd” is used interchangeably with “nucleic acid programmable DNA-binding protein (napDNAbp)”. In some embodiments, the napDNAbd is an RNA programmable DNA-binding protein. Various napDNAbds are known in the art, including, for example, those listed in WO 2020 / 181195, which is incorporated herein by reference in its entirety. Representative napDNAbds include, for example, CRISPR-associated (Cas) proteins, IscB, IsrB, Argonaute, and TnpB.

[0294] In some embodiments, napDNAbd is substantially lacking dsDNA cleavage activity (endonuclease activity). In some embodiments, napDNAbd is substantially lacking both dsDNA cleavage activity (endonuclease) and cleavage enzyme activity. In some embodiments, napDNAbd is non-nuclease-active, for example, Cas-inactivated. In some embodiments, napDNAbd is non-nuclease-active, for example, Cas-inactivated.

[0295] In some embodiments, napDNAbd is a nicking enzyme. In some embodiments, napDNAbd has nicking enzyme activity. In some embodiments, napDNAbd has nicking enzyme activity to create a nick in the target strand. In some embodiments, napDNAbd creates a nick in the target strand. In some embodiments, the method includes creating a nick in the target strand. In some embodiments, the nick on the target strand or the creation of a nick in the target strand incorporates an insertion / deletion (insertion and / or deletion) into the target strand.

[0296] In some implementations, napDNAbd can induce strand separation of the target dsDNA.

[0297] In some embodiments, napDNAbd contains a Cas domain. In some embodiments, napDNAbd contains either the Cas cleavage enzyme (nCas) of the Cas protein or inactivated (nuclease-free) Cas (dCas).

[0298] In some implementations, the Cas protein is selected from Cas9 proteins (e.g., SpCas9, SaCas9, GeoCas9, CjCas9, Cas9-KKH, circularly permuted Cas9, Argonaute (Ago), SmacCas9, Spy-macCas9, xCas9, SpCas9-NG, SpG). Cas9); Cas12 proteins (e.g., Cas12a (Cpf1), AsCas12a, LbCas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f (Cas14), Cas12g, Cas12h, Cas12i, xCas12i, Cas12Max, hfCas12Max, Cas12j, Cas12k, Cas12l, Cas12m, Cas12n, Cas12o, Cas12p, Cas12q, Cas12r, Cas12s, Cas12t, Cas12u, Cas12v, Cas12w, Cas12x, Cas12y, Cas12z); Cas13 proteins (e.g., Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, Cas13x, Cas13y); Csn2; and their mutants. xCas12i, Cas12Max and hfCas12Max are listed in PCT / CN2022 / 129376, which is incorporated herein by reference in its entirety.

[0299] In some implementations, the Cas nicking enzyme is the Cas9 nicking enzyme (nCas9), such as the SpCas9 nicking enzyme (SpCas9-D10A).

[0300] In some implementations, the deactivated Cas is deactivated Cas9 (dCas9), such as deactivated SpCas9 (SpCas9-D10A+H840A).

[0301] In some embodiments, the Cas cleavage enzyme is a Cas12i cleavage enzyme (nCas12i) or an inactivated Cas12i (dCas12i), such as an inactivated Cas12i of the xCas12i peptide.

[0302] In some embodiments, napDNAbd comprises an IscB protein (e.g., OgeuIscB) or an IscB cleavage enzyme (nIscB) or an inactivated IscB (dIscB) of the IscB protein described in PCT / CN2023 / 129167, PCT / CN2023 / 142506, PCT / CN2024 / 071744 and PCT / CN2023 / 125069, which are incorporated herein by reference in their entirety.

[0303] In some embodiments, napDNAbd comprises an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO:2, 48, 50, 52, or 163.

[0304] In some implementations, napDNAbd contains a TnpB cleavage enzyme or an inactivated TnpB protein.

[0305] BE construction

[0306] In some embodiments, the fusion protein comprises (1) a napDNAbp and a base excision domain from the N-terminus to the C-terminus; or (2) a base excision domain and a napDNAbp.

[0307] In some embodiments, napDNAbd (e.g., Cas9) is a two-part napDNAbd comprising an N-terminal portion and a C-terminal portion, such as a two-part splitting Cas9, and wherein the fusion protein comprises from the N-terminus to the C-terminus (1) the N-terminal portion of napDNAbd, a base-removing domain, and the C-terminal portion of napDNAbd; (2) the C-terminal portion of napDNAbd, a base-removing domain, and the N-terminal portion of napDNAbd; or (3) a base-removing domain, the C-terminal portion of napDNAbd (e.g., amino acids at positions 1249-1368), and the N-terminal portion of napDNAbd (e.g., amino acids at positions 1-1248).

[0308] In some embodiments, napDNAbd is SpCas9 (e.g., SpCas9 nickase) or a mutant thereof (e.g., SpG Cas9 nickase). In some embodiments, the N-terminal portion of napDNAbd is amino acids at positions 1 or 2 to 1012, 1028, 1041, 1046, 1047, 1248, 1249, or 1300 of napDNAbp. In some embodiments, the C-terminal portion of napDNAbd is amino acids at positions 1013, 1029, 1042, 1047, 1048, 1249, 1063, 1064, 1230, 1249, or 1301 to 1368 of napDNAbp.

[0309] For example, in some embodiments, the fusion protein includes a base-cutting domain embedded between positions 2-1248 and 1249-1368 of nCas9 (SEQ ID NO:2), wherein the first amino acid residue D of nCas9 (SEQ ID NO:2) is designated as position 2; or embedded between positions 2-1047 and 1064-1368 of nCas9 (SEQ ID NO:2), wherein the first amino acid residue D of nCas9 (SEQ ID NO:2) is designated as position 2.

[0310] Typical proteins usually have an N-terminal Met at their N-terminus (position 1) because it requires translation from a polynucleotide containing a start codon ATG (encoding Met) at its 5' end. However, if such a protein is fused to the C-terminus of a second protein (e.g., NLS, napDNAbd), the start codon ATG may not be necessary for that protein, because typically a start codon is present upstream of the second protein for translation of the fused protein as a whole, and thus the N-terminal Met of the protein can be removed. Any protein described in this disclosure refers to both the protein itself and its N-terminal truncated form with its N-terminal Met removed (if present).

[0311] In some embodiments, the fusion protein comprises an NLS at the N-terminus and / or C-terminus of the napDNAbp. In some embodiments, the fusion protein comprises an NLS at the N-terminus and / or C-terminus of a base-excision domain. In some embodiments, the NLS is or comprises an SV40 NLS, a bpSV40 NLS (e.g., SEQ ID NO: 10 or 11), or an NP NLS (Xenopus laevis nucleoplasmic protein NLS, nucleoplasmic protein NLS). Other NLS suitable for this disclosure, or any manner of linking an NLS to a component of the fusion protein of this disclosure, includes, for example, a SGGS adapter or those listed in WO 2020 / 181195, which is incorporated herein by reference in its entirety.

[0312] In some embodiments, components of the fusion protein (e.g., napDNAbp and a base-removing domain, NLS and napDNAbp, or NLS and a base-removing domain) are fused together with or without a linker. Suitable linkers include, for example, SGGS, the linker of SEQ ID NO:134, and those listed in WO 2020 / 181195, which is incorporated herein by reference in its entirety.

[0313] In some embodiments, the fusion protein comprises having at least about 60 ppm of any one of SEQ ID NO: 12, 14, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 55, 57, 59, 61, 63, 136, 138, 140, 142, 143, 145, 147, 149, 151, 153, 154, 156, 158, 160, 161, 162, and 164. The amino acid sequence is of % (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity.

[0314] Deaminase domain

[0315] While omitting the deaminase domain in the base editor of this disclosure has advantages, deamination can be permitted after the bases of the target deoxyribonucleotide are removed. In some respects, the fusion proteins of this disclosure can be used in combination with the deaminase domain for various purposes (e.g., to improve the purity of the outcome). "Purity" refers to a percentage / proportion of a particular outcome among all possible outcomes. For example, the purity of dT refers to the percentage / proportion of dT as an outcome among all possible outcomes, including, for example, dA, dT, dG, and dC. As an example, the introduction of the deaminase domain can facilitate the further conversion of an undesirable deoxyribonucleotide (e.g., dC) as a byproduct (e.g., dT) into a desired deoxyribonucleotide (e.g., dT) via A-to-T base editing. Thus, in summary, if dG to dT is desired, a two-stage conversion can be achieved, first, by converting the target dG portion into dC via base editing without deamination as described herein, and second, by converting dC into dT via C-to-T base editing with deamination, thereby achieving high-purity dG to dT. Therefore, in some embodiments, the fusion protein further comprises a deaminase domain. The deaminase domain can be fused to components of the fusion protein, with or without a linker as described herein.

[0316] Various deaminases are known in the art, including, for example, those listed in WO 2020 / 181195, which is incorporated herein by reference in its entirety. Representative adenine deaminases include, for example, TadA and its homologues and variants, and APOBEC and its homologues and variants.

[0317] In some embodiments, the deaminase domain is (substantially) a deaminase domain capable of deaminating adenine, guanine, hypoxanthine, cytidine, thymine, and / or uracil. In some embodiments, the deaminase domain is an adenine deaminase domain or a cytosine deaminase domain.

[0318] In some embodiments, the deaminase domain comprises tRNA adenosine deaminase (TadA) or a functional variant or fragment thereof, such as TadA8e (SEQ ID NO:3), TadA8.17, TadA8.20, TadA9, or TadA8E. V106W 、TadA8E V106W+D108Q TadA-CDa, TadA-CDb, TadA-CDc, TadA-CDd, TadA-CDe, TadA-dual, T AD AC-1.2, T AD AC-1.14, T AD AC-1.17, T AD AC-1.19, T AD AC-2.5, T AD AC-2.6, TAD AC-2.9, T AD AC-2.19, T AD AC-2.23, TadA8e-N46L, TadA8e-N46P.

[0319] In some embodiments, the deaminase domain comprises apolipoprotein B mRNA editing complex (APOBEC) family deaminases, activation-inducible deaminases (AID), cytidine deaminase 1 (pmCDA1) from petrified marinus, or functional variants or fragments thereof, such as APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, APOBEC3G, and APOBEC3H.

[0320] Prototype spacer subsequence

[0321] In some embodiments, the prototype spacer sequence comprises about, at least about, or at most about 14 adjacent nucleotides of the target dsDNA, such as about, at least about, or at most about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, on the non-target strand of the target dsDNA. 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more adjacent nucleotides, or within a range of any two of the foregoing values, for example, about 16 to about 50 or about 17 to about 22 adjacent nucleotides of the target dsDNA. In some embodiments, the prototype spacer sequence contains about 20, 30 or 50 adjacent nucleotides on the non-target strand of the target dsDNA.

[0322] In some embodiments, the prototype spacer sequence is an extension of about, at least about or at most about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more adjacent nucleotides on the non-target strand of the target dsDNA, or an extension of about 16 to about 50 adjacent nucleotides in the range of any two of the aforementioned values ​​on the non-target strand of the target dsDNA. In some implementations, the prototype spacer sequence is an extension of about 20, 30, or 50 adjacent nucleotides on the non-target strand of the target dsDNA.

[0323] In some embodiments, the prototype spacer subsequence is immediately adjacent to the 5' or 3' of a prototype spacer neighboring motif (PAM) containing the sequence 5'-NN-3', 5'-NNN-3', 5'-NNNN-3', 5'-NNNNN-3', or 5'-NNNNNN-3', where N is A, T, G, or C. For example, in some embodiments, the prototype spacer subsequence is immediately adjacent to the 5' of a prototype spacer neighboring motif (PAM) containing the sequence 5'-NGG-3' or 5'-NTN-3', where N is A, T, G, or C. In some embodiments, the prototype spacer subsequence is immediately adjacent to the 3' of a prototype spacer neighboring motif (PAM) containing the sequence 5'-TTN-3', where N is A, T, G, or C.

[0324] Guide sequence

[0325] In some embodiments, the length of the guide sequence is about, at least about, or at most about 14 nucleotides, for example about, at least about, or at most about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more nucleotides, or the length of the guide sequence is within a numerical range between any two of the foregoing values, for example, about 16 to about 50 nucleotides. In some implementations, the guide sequence is about 20, 30, or 50 nucleotides in length.

[0326] In some embodiments, (1) the guide sequence is at least about 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% (completely) anticomplementary to the target sequence, optionally about 100% (completely) anticomplementary; (2) the guide sequence contains no more than 5, 4, 3, 2, or 1 mismatch with the target sequence or contains no mismatch; or (3) the guide sequence is first 1, 2, 3, 4, 5, ... Nucleotides 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, or 70 do not contain mismatches with the target sequence. In some embodiments, the guide sequence is approximately 100% (completely) anticomplementary to the target sequence.

[0327] In some embodiments, the guide sequence contains one mismatch with the target sequence. In some embodiments, the guide sequence and the target sequence are approximately 98% anticomplementary. In some embodiments, the mismatch in the guide sequence is located at a position corresponding to the nucleotide in the target sequence that is intended to be substituted.

[0328] In some embodiments, the guide sequence comprises (1) a sequence of SEQ ID NO:75-103 and 165-3030 (excluding PAM, if PAM is present) or a 5' or 3' truncated form thereof with 1, 2, 3, 4, 5 or 6 nucleotides shortened at the 5' or 3' end; or (2) a sequence having at least about 70%, 75%, 80%, 85%, 90%, 95% or 100% sequence identity with SEQ ID NO:75-103 and 165-3030 (excluding PAM, if PAM is present) or a 5' or 3' truncated form thereof with 1, 2, 3, 4, 5 or 6 nucleotides shortened at the 5' or 3' end; or (3) a sequence having at most 1, 2, 3, 4, 5 or 6 nucleotide differences compared with SEQ ID NO:75-103 and 165-3030 (excluding PAM, if PAM is present), whether or not such differences are consecutive.

[0329] In some implementations, the guide sequence comprises the sequence of either SEQ ID NO:75-103 or 165-3030 (excluding PAM, if PAM is present).

[0330] In some aspects, this disclosure provides a guide nucleic acid comprising a guide sequence as described herein and a scaffold sequence capable of forming a complex with napDNAbd. The scaffold sequence and napDNAbd may be as described herein.

[0331] stent sequence

[0332] For the purposes of this disclosure, the scaffold sequence is compatible with and capable of complexing with napDNAbd. The scaffold sequence may be a naturally occurring scaffold sequence identified together with napDNAbd, or a variant thereof that maintains its ability to complex with napDNAbd. Generally, the ability to complex with napDNAbd is maintained as long as the secondary structure of the variant is substantially identical to that of the naturally occurring scaffold sequence. Nucleotide deletions, insertions, or substitutions in the primary sequence of the scaffold sequence may not necessarily alter the secondary structure of the scaffold sequence (e.g., the relative positions and / or sizes of the stems, protrusions, and loops of the scaffold sequence do not significantly deviate from the relative positions and / or sizes of the original stems, protrusions, and loops). For example, nucleotide deletions, insertions, or substitutions may be located in the protrusion or loop regions of the scaffold sequence such that the overall symmetry of the protrusions and therefore the secondary structure remains substantially the same. Nucleotide deletions, insertions, or substitutions may also be located in the stems of the scaffold sequence such that the length of the stem does not significantly deviate from the length of the original stem (e.g., adding or deleting one base pair in each of the two stems corresponds to a total of 4 base changes). In some embodiments, the scaffold sequence is located at the 5' or 3' of the guide sequence.

[0333] In some embodiments, the scaffold sequence has a secondary structure substantially identical to that of the sequence of SEQ ID NO:40, 73, or 74. In some embodiments, the scaffold sequence comprises (1) a 5' or 3' truncated form of SEQ ID NO:40, 73, or 74, or a 5' or 3' truncated form thereof by 1, 2, 3, 4, 5, or 6 nucleotides; or (2) a sequence having at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% sequence identity with SEQ ID NO:40, 73, or 74, or a 5' or 3' truncated form thereof by 1, 2, 3, 4, 5, or 6 nucleotides; or (3) a sequence differing from SEQ ID NO:40, 73, or 74 by at most 1, 2, 3, 4, 5, or 6 nucleotides, regardless of whether such differences are consecutive. In some embodiments, the scaffold sequence comprises the sequence of SEQ ID NO:40, 73, or 74.

[0334] Trans-damage synthesis (TLS) polymerase

[0335] In some respects, the fusion proteins of this disclosure can be used in combination with transdamage synthesis (TLS) polymerases to improve outcome purity. "Purity" refers to the percentage / proportion of a particular outcome among all possible outcomes. For example, the purity of dT refers to the percentage / proportion of dT as an outcome among all possible outcomes, including, for example, dA, dT, dG, and dC. As listed in Table 5, TLS polymerases may have a tendency to incorporate various deoxyribonucleotides opposite to the no-base site during polymerization. By utilizing this tendency, the base-editing outcome can be intentionally controlled to improve outcome purity. For example, human Polη (SEQ ID NO: 118) is a TLS polymerase that preferentially incorporates dA opposite to the no-base site. By using human Polη in combination, the base-editing outcome of dT can be adjusted, thereby increasing the purity of the dT product.

[0336] Table 5. DNA polymerases used to incorporate perfect bases that are opposite to base-free sites.

[0337] DNA polymerase Perfect bases opposite to baseless sites Polα(α) dA Polδ(δ) / PCNA dA Polγ(γ) dA Polη(η) dT and dA Polι(ι) dT, dG and dA Polκ(κ) dC and dA Polθ(θ) dA REV1 dC

[0338] In some embodiments, the fusion protein or system of this disclosure further comprises a transdamage synthesis (TLS) polymerase or a recruitment domain or component capable of recruiting a TLS polymerase. In some embodiments, the TLS polymerase or recruitment domain or component is fused to a component of the fusion protein, either without or with a linker as described herein.

[0339] Non-limiting examples of TLS polymerases are selected from the group consisting of Pol α (α), Pol β (β), Pol δ (δ) (PCNA), Pol γ (γ), Pol η (η), Pol ι (ι), Pol κ (κ), Pol λ (λ), Pol μ (μ), Pol ν (ν), Pol θ (θ), and REV1.

[0340] In some embodiments, the TLS polymerase comprises an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with the amino acid sequence of SEQ ID NO: 118. In some embodiments, the TLS polymerase comprises the amino acid sequence (Poln) of SEQ ID NO: 118.

[0341] In some embodiments, a fusion protein or system comprising a trans-damage synthesis (TLS) polymerase or capable of recruiting the recruitment domain of a TLS polymerase results in the conversion of the first deoxyribonucleotide to dG, dC, dT, or dA.

[0342] MPG

[0343] In another aspect, this disclosure provides an MPG as described herein or disclosed herein.

[0344] In some embodiments, the MPG comprises the motif GxxYxxxxYGxxxxxN, where x represents any amino acid. In some embodiments, the MPG is obtained from species selected from Table A. In some embodiments, the MPG contains an amino acid mutation relative to (compared to; with respect to) the wild-type or reference MPG. In some embodiments, the wild-type or reference MPG is a human MPG (SEQ ID NO:9), or an MPG obtained from a species selected from Table A, or any MPG as shown in Table D, or a homolog or mutant thereof (e.g., containing the amino acid sequence of SEQ ID NO:5, 6, or 7), or an N-terminal truncated form of the MPG lacking the N-terminal methionine (Met) (encoded by the start codon ATG) (e.g., SEQ ID NO:1). In some embodiments, the wild-type or reference MPG contains the amino acid sequence of SEQ ID NO:7.

[0345] In some embodiments, the amino acid mutation comprises amino acid substitutions at positions selected from the following: wild type or reference MPG 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65. 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 1 68, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 19 9, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297 and / or 298, where the position is numbered according to SEQ ID NO:1.

[0346] In some embodiments, the amino acid mutation confers the ability to remove a base from MPG. In some embodiments, the base is guanine.

[0347] In some embodiments, compared to a otherwise identical control MPG without the stated amino acid mutation, the mutation results in an increased base-excision capacity, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, or 160%. 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0348] In some embodiments, compared to a otherwise identical control fusion protein without the said amino acid mutation, the amino acid mutation results in an increase in the guide sequence-specific base editing efficiency of the fusion protein, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, or 150%. %, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0349] In some embodiments, the amino acid mutation includes amino acid substitutions at positions corresponding to or selected from the following positions: wild type or reference MPG: G163, N169, D175, C178, S198, K202, G203, S206, K210 and / or Q294, wherein the position is numbered according to SEQ ID NO:1.

[0350] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position corresponding to a position selected from or equivalent to N169, D175, C178, and / or Q294 of the wild-type or reference MPG, wherein the position is numbered according to SEQ ID NO:1. In some embodiments, the wild-type or reference MPG comprises the amino acid sequence of SEQ ID NO:7.

[0351] In some embodiments, the amino acid substitution is either a conserved or non-conserved amino acid substitution. In some embodiments, the amino acid substitution is performed using an amino acid residue that is different from the amino acid residue at that position in the wild type or the reference MPG. In some embodiments, the amino acid substitution is performed using the following amino acid substitutions: (1) nonpolar amino acid residues (such as glycine (Gly / G), alanine (Ala / A), valine (Val / V), cysteine ​​(Cys / C), proline (Pro / P), leucine (Leu / L), isoleucine (Ile / I), methionine (Met / M), tryptophan (Trp / W), phenylalanine (Phe / F); (2) polar amino acid residues (such as serine (Ser / S), threonine (Thr / T), tyrosine (Tyr / Y), asparagine (Asn / N), glutamine (Gln / Q)); (3) positively charged amino acid residues (such as lysine (Lys / K), arginine (Arg / R), histidine (His / H)); or (4) negatively charged amino acid residues (such as aspartic acid (Asp / D), glutamic acid (Glu / E)). In some embodiments, the amino acid substitution is performed using R, A, N, or G.

[0352] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: G163R, N169G, D175R, C178N, S198A, K202A, G203A, S206A, K210A, Q294R, and the position numbered according to SEQ ID NO:1.

[0353] In some embodiments, the MPG containing the amino acid mutation comprises, substantially comprises, or consists of the following: an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100% sequence identity with the amino acid sequence of the wild-type or reference MPG.

[0354] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: N169G, D175R, C178N, Q294R, wherein the position is numbered according to SEQ ID NO:1. In some embodiments, the wild type or reference MPG comprises the amino acid sequence of SEQ ID NO:7.

[0355] In some embodiments, the amino acid mutation comprises a combination of substitutions corresponding to the combination of N169G, D175R, C178N, and Q294R, wherein the position is numbered according to SEQ ID NO:1. In some embodiments, the wild-type or reference MPG comprises the amino acid sequence of SEQ ID NO:7.

[0356] In some embodiments, the amino acid mutation comprises a combination of substitutions corresponding to the combination of G163R, N169G, D175R, C178N, S198A, K202A, G203A, S206A, K210A, and Q294R, wherein the position is numbered according to SEQ ID NO:1. In some embodiments, the wild-type or reference MPG comprises the amino acid sequence of SEQ ID NO:1 or 9.

[0357] In some embodiments, the MPG containing the amino acid mutation comprises, substantially comprises, or consists of an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO:8, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38) sequence identity.

[0358] In some embodiments, MPG is (substantially) capable of cleaving guanine in dG. In some embodiments, MPG is (substantially) incapacitated by thymine in dT. In some embodiments, MPG is (substantially) incapacitated by cytosine in dC. In some embodiments, MPG is (substantially) incapacitated by adenine in dA.

[0359] In some respects, this disclosure provides fusion proteins that contain the MPG and functional domains described herein or disclosed herein, such as napDNAbd.

[0360] In some respects, this disclosure provides for the use of MPG, as described herein or disclosed herein, for base editing as described herein.

[0361] In some implementations, MPG is neither wild-type nor naturally occurring.

[0362] UNG

[0363] In another aspect, this disclosure provides a UNG described herein or disclosed herein.

[0364] In some embodiments, the UNG contains the motif GQDPYH. In some embodiments, the UNG is obtained from species selected from Table C. In some embodiments, the UNG contains an amino acid mutation relative to (compared to; with respect to) the wild-type or reference UNG. In some embodiments, the wild-type or reference UNG is human UNG1 (SEQ ID NO: 54) or human UNG2 (SEQ ID NO: 133), or an UNG obtained from species selected from Table C, or any UNG as shown in Table D, or its homologs or mutants, or its N-terminal truncated form lacking the N-terminal methionine (Met) (encoded by the start codon ATG). In some embodiments, the wild-type or reference UNG contains the amino acid sequence of SEQ ID NO: 56, 58, 135, or 137.

[0365] In some embodiments, the amino acid mutation comprises amino acid substitutions at positions selected from the following: wild type or reference UNG 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65. 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 1 68, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 19 9, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312 and / or 313, wherein the position is numbered according to SEQ ID NO: 133.

[0366] Alternative splicing and transcription from two different initiation sites result in two distinct human UNG isoforms: mitochondrial UNG1 (304 amino acids (aa)) (SEQ ID NO:54) and nuclear UNG2 (313 aa) (SEQ ID NO:133), each possessing a unique N-terminus that mediates translocation to the mitochondria or nucleus. 16 ( Figure 25 Sequence alignment of human UNG1 and human UNG2 shows that amino acid residues at positions 1-35 of UNG1 are different from amino acid residues at positions 1-44 of UNG2, and the rest of UNG1 and UNG2 are identical. The correspondence between the positions of UNG1 (SEQ ID NO: 54) and UNG2 (SEQ ID NO: 133) can be determined by sequence alignment. For example, residues Y156 and N213 of UNG2 correspond to residues Y147 and N204 of UNG1, respectively.

[0367] In some embodiments, the amino acid mutation confers the ability to remove the base at UNG. In some embodiments, the base is thymine. In some embodiments, the base is cytosine.

[0368] In some embodiments, compared to a otherwise identical control UNG without the stated amino acid mutation, the amino acid mutation results in an increased base-excision capacity, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, 150%, or 160%. 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0369] In some embodiments, compared to a otherwise identical control fusion protein without the said amino acid mutation, the amino acid mutation results in an increase in the guide sequence-specific base editing efficiency of the fusion protein, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, or 150%. %, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0370] In some embodiments, the amino acid mutation includes an amino acid substitution at a position corresponding to a position selected from or selected from the following positions: wild type or reference UNG Y156, K184, N213, A214, Q259 and / or Y284, wherein the position is numbered according to SEQ ID NO:133.

[0371] In some embodiments, the amino acid mutation comprises an amino acid substitution at a position corresponding to a position selected from or selected from the wild-type or reference UNG, specifically K184, A214, Q259, and / or Y284, wherein the position is numbered according to SEQ ID NO: 133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO: 135 or 137.

[0372] In some embodiments, the amino acid substitution is either a conserved or non-conserved amino acid substitution. In some embodiments, the amino acid substitution is performed using an amino acid residue that is different from the amino acid residue at that position in the wild type or the reference UNG. In some embodiments, the amino acid substitution is performed using the following amino acid substitutions: (1) nonpolar amino acid residues (such as glycine (Gly / G), alanine (Ala / A), valine (Val / V), cysteine ​​(Cys / C), proline (Pro / P), leucine (Leu / L), isoleucine (Ile / I), methionine (Met / M), tryptophan (Trp / W), phenylalanine (Phe / F); (2) polar amino acid residues (such as serine (Ser / S), threonine (Thr / T), tyrosine (Tyr / Y), asparagine (Asn / N), glutamine (Gln / Q)); (3) positively charged amino acid residues (such as lysine (Lys / K), arginine (Arg / R), histidine (His / H)); or (4) negatively charged amino acid residues (such as aspartic acid (Asp / D), glutamic acid (Glu / E)). In some embodiments, the amino acid substitution is performed using A, D, V, or T.

[0373] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: Y156A, K184A, N213D, A214V, A214T, Q259A, Y284D, and the position numbered according to SEQ ID NO:133.

[0374] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: A214T, Q259A, Y284D, and the position specified in SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:135.

[0375] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of K184A, A214V, and two such substitutions, wherein the position is designated according to SEQ ID NO: 133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO: 137.

[0376] In some embodiments, the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or corresponding to one of the following positions: wild type or reference UNG 1-65, 1-66, 1-67, 1-68, 1-69, 1-70, 1-71, 1-72, 1-73, 1-74, 1-75, 1-76, 1-77, 1-78, 1-79, 1-80, 1-81, 1-82, 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, or 1-100, wherein the position is numbered according to SEQ ID NO:133.

[0377] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: A214T, Q259A, Y284D, and any two or more thereof, and the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or at positions 1-88 of the wild-type or reference UNG, wherein the position is numbered according to SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:135.

[0378] In some embodiments, the amino acid mutation comprises an amino acid substitution corresponding to a substitution selected from or a combination of K184A, A214V, and two such substitutions, and the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or a position 1-88 of the wild-type or reference UNG, wherein the position is numbered according to SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:137.

[0379] In some embodiments, the UNG containing the amino acid mutation comprises, is substantially composed of, or consists of, an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100% sequence identity with the amino acid sequence of the wild-type or reference UNG.

[0380] In some embodiments, the amino acid mutation comprises a combination of substitutions corresponding to the combination of A214T, Q259A, and Y284D, and the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or corresponding to positions 1-88 of the wild-type or reference UNG, wherein the position is numbered according to SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:135.

[0381] In some embodiments, the amino acid mutation comprises a combination substitution corresponding to a combination substitution of K184A and A214V, and the amino acid mutation comprises the deletion of an amino acid at a position corresponding to or corresponding to positions 1-88 of the wild-type or reference UNG, wherein the position is numbered according to SEQ ID NO:133. In some embodiments, the wild-type or reference UNG comprises the amino acid sequence of SEQ ID NO:137.

[0382] In some embodiments, the UNG containing the amino acid mutation comprises, substantially consists of, or consists of: with SEQ ID The amino acid sequence of NO:56, 58, 60, 62, 135, 137, 139, 141, 144, 146, 148, 150, 152, 155, 157, and 159, or the N-terminal truncated amino acid sequence lacking the most N-terminal methionine (M) (encoded by the start codon ATG), has at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity.

[0383] In some embodiments, the UNG (substantially) can cleave thymine from dT. In some embodiments, the UNG (substantially) can cleave cytosine from dC. In some embodiments, the UNG (substantially) cannot cleave thymine from dT. In some embodiments, the UNG (substantially) cannot cleave cytosine from dC. In some embodiments, the UNG (substantially) cannot cleave adenine from dA. In some embodiments, the UNG (substantially) cannot cleave guanine from dG.

[0384] In some respects, this disclosure provides fusion proteins that contain the UNG and functional domains described herein or disclosed herein, such as napDNAbd.

[0385] In some respects, this disclosure provides for the use of UNG as described herein or in this disclosure for base editing as described herein.

[0386] In some implementations, UNG is neither wild-type nor naturally occurring.

[0387] Other BEP

[0388] In another respect, this disclosure provides TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3 or NTHL1 as described herein or disclosed herein.

[0389] In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 contains an amino acid mutation relative to (compared to; with respect to) wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1. In some embodiments, wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 is human TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (SEQ ID NO: 64, 65, 66, 67, 68, 69, 70, 71, or 72, respectively) or its homologs or mutants, or an N-terminal truncated form lacking the N-terminal methionine (Met) (encoded by the start codon ATG). In some embodiments, the wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 each comprises an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with any of SEQ ID NO: 64-72.

[0390] In some embodiments, the amino acid mutation includes amino acid substitutions at positions selected from the following: wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 at positions 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50. 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 1 88, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 21 9, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250251、252、253、254、255、256、257、258、259、260、261、262、263、264、265、266、267、268、269、270、271、272、273、274、275、276、277、278、279、280、281、282、283、284、285、286、287、288、289、290、291、292、293、294、295、296、297、298、299、300、301、302、303、304、305、306、307、308、309、310、311、312、313、314、315、316、317、318、319、320、321、322、323、324、325、326、327、328、329、330、331、332、333、334、335、336、337、338、339、340、341、342、343、344、345、346、347、348、349、350、351、352、353、354、355、356、357、358、359、360、361、362、363、364、365、366、367、368、369、370、371、372、373、374、375、376、377、378、379、380、381、382、383、384、385、386、387、388、389、390、391、392、393、394、395、396、397、398、399、400、401、402、403、404、405、406、407、408、409、410、411、412、413、414、415、416、417、418、419、420、421、422、423、424、425、426、427、428、429、430、431、432、433、434、435、436、437、438、439、440、441、442、443、444、445、446、447、448、449、450、451、452、453、454、455、456、457、458、459、460、461、462、463、464、465、466、467、468、469、470、471、472、473、474、475、476、477、478、479、480、481、482、483、484、485、486、487、488、489、490、491、492、493、494、495、496、497、498、499、500、501, 502, 503, 504, 505, 506, 507, 508, 509, 510, 511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525, 526, 527, 528, 529, 530, 531, 532, 533, 534, 535, 536, 537, 538, 539, 540, 541, 542, 543, 544, 545, 546, 547, 548, 549, 550, 551, 552, 553, 554 555, 556, 557, 558, 559, 560, 561, 562, 563, 564, 565, 566, 567, 568, 569, 570, 571, 572, 573, 574, 575, 576, 577, 578, 579, 580, 581, 582, 583, 584, 585, 586, 587, 588, 589, 590, 591, 592, 593, 594, 595, 596, 597, 598, 599, 600, 601, 602, 603, and / or 604, wherein the position is based on any one of the numbers in SEQ ID NO: 64-72.

[0391] In some embodiments, the amino acid mutation confers the ability to remove bases from TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1. In some embodiments, the bases are guanine, thymine, cytosine, adenine, uracil, or hypoxanthine.

[0392] In some embodiments, compared to otherwise identical controls TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 without the said amino acid mutation, the amino acid mutation results in an increased base-cutting capacity, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, or 1... 20%, 130%, 140%, 150%, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0393] In some embodiments, compared to a otherwise identical control fusion protein without the said amino acid mutation, the amino acid mutation results in an increase in the guide sequence-specific base editing efficiency of the fusion protein, for example, an increase of at least about 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 100%, 110%, 120%, 130%, 140%, or 150%. %, 160%, 170%, 180%, 190%, 200%, 210%, 220%, 230%, 240%, 250%, 260%, 270%, 280%, 290%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, 1100%, 1200%, 1300%, 1400%, 1500%, 1600%, 1700%, 1800%, 1900%, 2000% or more.

[0394] In some embodiments, the amino acid substitution is either a conserved amino acid substitution or a non-conserved amino acid substitution. In some embodiments, the amino acid substitution is an amino acid substitution made with an amino acid residue that is different from the amino acid residue at that position of the wild type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3 or NTHL1. In some embodiments, amino acid substitution is performed using the following amino acid substitutions: (1) nonpolar amino acid residues (such as glycine (Gly / G), alanine (Ala / A), valine (Val / V), cysteine ​​(Cys / C), proline (Pro / P), leucine (Leu / L), isoleucine (Ile / I), methionine (Met / M), tryptophan (Trp / W), phenylalanine (Phe / F); (2) polar amino acid residues (such as serine (Ser / S), threonine (Thr / T), tyrosine (Tyr / Y), asparagine (Asn / N), glutamine (Gln / Q)); (3) positively charged amino acid residues (such as lysine (Lys / K), arginine (Arg / R), histidine (His / H)); or (4) negatively charged amino acid residues (such as aspartic acid (Asp / D), glutamic acid (Glu / E)).

[0395] In some embodiments, the TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 containing the said amino acid mutation comprises, substantially comprises, or comprises the following: an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%) and less than 100% sequence identity with the wild-type or reference TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1.

[0396] In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cleaves guanine from dG. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cleaves thymine from dT. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cleaves cytosine from dC. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cleaves adenine from dA.

[0397] In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cannot cleave thymine in dT. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cannot cleave cytosine in dC. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cannot cleave adenine in dA. In some embodiments, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 (substantially) cannot cleave guanine in dG.

[0398] In some respects, this disclosure provides fusion proteins comprising TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3 or NTHL1 as described herein or disclosed herein, and functional domains such as napDNAbd.

[0399] In some respects, this disclosure provides for the use of TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3 or NTHL1 described herein or disclosed herein for base editing as described herein.

[0400] In some implementations, TDG, SMUG1, OGG1, MBD4, MUTYH, NEIL1, NEIL2, NEIL3, or NTHL1 are not wild-type or naturally occurring.

[0401] Guiding the regulation of nucleic acids

[0402] This disclosure also provides polynucleotides that contain or encode guiding nucleic acids.

[0403] In some embodiments, the polynucleotide comprising or encoding the guide nucleic acid is DNA, RNA, or a DNA / RNA mixture. A “DNA / RNA mixture” refers to a nucleic acid comprising one or more modified or unmodified ribonucleotides and one or more modified or unmodified deoxyribonucleotides, regardless of whether these ribonucleotides and deoxyribonucleotides are sequential. However, “DNA” or “RNA” may also refer to DNA containing one or more modified or unmodified ribonucleotides (whether sequential) or RNA containing one or more modified or unmodified deoxyribonucleotides (whether sequential).

[0404] In some implementations, nucleic acids are operatively linked to promoters or regulated by promoters.

[0405] In some implementations, the promoter is a ubiquitous promoter, a tissue-specific promoter, a cell type-specific promoter, a constitutive promoter, or an inducible promoter.

[0406] Suitable promoters are those known in the art and include, for example, the Cbh promoter, Cba promoter, pol I promoter, pol II promoter, pol III promoter, T7 promoter, U6 promoter, H1 promoter, retroviral Rous sarcoma virus LTR promoter, cytomegalovirus (CMV) promoter, SV40 promoter, dihydrofolate reductase promoter, β-actin promoter, elongation factor 1α short (EFS) promoter, β-glucuronidase (GUSB) promoter, cytomegalovirus (CMV) immediate early (Ie) enhancer and / or promoter, chicken β-actin (CBA) promoter or its derivatives such as CAG promoter, CB promoter, (human) elongation factor 1α-subunit (EF1α) promoter, ubiquitin C (UBC) promoter, prion promoter, neuron-specific enolase (NSE), neurofilament light chain (NFL) promoter, neurofilament heavy chain (NFH) promoter, etc. The promoters for platelet-derived growth factor (PDGF), platelet-derived growth factor B chain (PDGF-β) promoter, synaptoprotein (Syn) promoter, synaptoprotein 1 (Syn1) promoter, methyl-CpG-binding protein 2 (MeCP2) promoter, Ca2+ / calmodulin-dependent protein kinase II (CaMKII) promoter, metabolite glutamate receptor 2 (mGluR2) promoter, neurofilament light chain (NFL) promoter, neurofilament heavy chain (NFH) promoter, β-globin small gene nβ2 promoter, proenkephalinogen (PPE) promoter, enkephalin (Enk) promoter, excitatory amino acid transporter 2 (EAAT2) promoter, glial fibrillary acidic protein (GFAP) promoter, and myelin basic protein (MBP) promoter.

[0407] Regulation of fusion proteins

[0408] In another aspect, this disclosure provides a polynucleotide encoding the fusion protein of this disclosure and optionally the guiding nucleic acid of this disclosure.

[0409] In some embodiments, the polynucleotide encoding the fusion protein is DNA, RNA, or a DNA / RNA mixture. "DNA / RNA mixture" refers to a nucleic acid comprising one or more modified or unmodified ribonucleotides and one or more modified or unmodified deoxyribonucleotides, regardless of whether these ribonucleotides and deoxyribonucleotides are sequential. However, "DNA" or "RNA" may also refer to DNA containing one or more modified or unmodified ribonucleotides (whether sequential or not), or RNA containing one or more modified or unmodified deoxyribonucleotides (whether sequential or not).

[0410] In some implementations, the polynucleotide encoding napDNAbd is mRNA.

[0411] In some implementations, the polynucleotide encoding napDNAbd comprises a sequence encoding napDNAbd and a promoter operatively linked to the sequence encoding napDNAbd.

[0412] In some implementations, the polynucleotide encoding napDNAbd is operatively linked to a promoter or regulated by a promoter.

[0413] In some implementations, the promoter is a broad-spectrum promoter, a tissue-specific promoter, a cell type-specific promoter, a constitutive promoter, or an inducible promoter.

[0414] Suitable promoters are known in the art and include, for example, the Cbh promoter, the Cba promoter, the pol I promoter, the pol II promoter, and the pol... III promoter, T7 promoter, U6 promoter, H1 promoter, retroviral Rouss sarcoma virus LTR promoter, cytomegalovirus (CMV) promoter, SV40 promoter, dihydrofolate reductase promoter, β-actin promoter, elongation factor 1α short (EFS) promoter, β-glucuronidase (GUSB) promoter, cytomegalovirus (CMV) early (Ie) enhancer and / or promoter, chicken β-actin (CBA) promoter or its derivatives such as CAG promoter, CB promoter, (human) elongation factor 1α-subunit (EF1α) promoter, ubiquitin C (UBC) promoter, prion promoter, neuron-specific enolase (NSE), neurofilament light chain (NFL) promoter, neurofilament heavy chain (NFH) promoter, platelet-derived growth factor (PDGF) promoter, platelet-derived growth factor B chain (PDGF-β) promoter, synaptic protein (Syn) promoter, human synaptic protein (hSyn) promoter, synaptic protein 1 (Syn1) promoter, methyl-CpG binding protein 2 (MeCP2) promoter, Ca2+ / calmodulin-dependent protein kinase II (CaMKII) promoter, metabolite glutamate receptor 2 (mGluR2) promoter, neurofilament light chain (NFL) promoter, neurofilament heavy chain (NFH) promoter, β-globin small gene nβ2 promoter, proenkephalinogen (PPE) promoter, enkephalin (Enk) promoter, excitatory amino acid transporter 2 (EAAT2) promoter, glial fibrillary acidic protein (GFAP) promoter, myelin basic protein (MBP) promoter, OTOF promoter, GRK1 promoter, CRX promoter, NRL promoter, MECP2 promoter, mMECP2 promoter, hMECP2 promoter, APP promoter, and RCRRN promoter.

[0415] deliver

[0416] Depending on practical needs, various delivery methods can be applied to the fusion protein or system disclosed herein.

[0417] In another aspect, this disclosure provides a delivery system comprising (1) a fusion protein of this disclosure, a polynucleotide of this disclosure, or a system of this disclosure; and (2) a delivery medium.

[0418] In another aspect, this disclosure provides a vector comprising the polynucleotides of this disclosure. In some embodiments, the vector encodes the guide nucleic acid of this disclosure. In some embodiments, the vector is a plasmid vector, a recombinant AAV (rAAV) vector (vector genome), or a recombinant lentiviral vector.

[0419] In another aspect, this disclosure provides a recombinant AAV (rAAV) particle containing the rAAV vector genome of this disclosure. A brief introduction to the AAV used for delivery can be found in the "Adeno-Associated Virus (AAV) Guide".

[0420] (addgene.org / guides / aav / ).

[0421] Adeno-associated viruses (AAVs) that are engineered to deliver, for example, protein-coding sequences of interest, can be referred to as (r)AAV vectors, (r)AAV vector particles, or (r)AAV particles, where "r" stands for "recombinant." The genome packaged in an AAV vector for delivery can be referred to as (r)AAV vector genome, vector genome, or simply vg, while the viral genome can refer to the original viral genome of the native AAV.

[0422] The serotype of the rAAV particle capsid can be matched with the type of target cells. For example, Table 2 of WO 2018002719 A1 lists exemplary cell types that can be transduced by the indicated AAV serotype (incorporated herein by reference).

[0423] In some embodiments, the rAAV particles comprise a capsid having a serotype suitable for delivery to desired target cells. In some embodiments, the rAAV particles comprise a capsid having a serotype of AAV1, AAV2, AAV3A, AAV3B, AAV4, AAV5, AAV6, AAV7, AAVrh74, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV-DJ, or AAV.PHP.eB, a member of any of the clades belonging to AAV1-AAV13, or a functional variant thereof (e.g., a functional truncated form), thereby encapsulating the rAAV vector genome. In some embodiments, the capsid serotype is a wild-type serotype or a functional variant thereof.

[0424] The general principles of rAAV particle generation are known in the art. In some embodiments, rAAV particles can be generated using a triple transfection method (described in detail in U.S. Patent No. 6,001,650).

[0425] Vector titer is typically expressed as vector genome / ml (vg / ml). In some implementations, the vector titer is higher than 1×10⁻⁶. 9 Higher than 5×10 10 Higher than 1×10 11 Higher than 5×10 11 Higher than 1×10 12 Higher than 5×10 12 or higher than 1×10 13 vg / ml.

[0426] A system and method for packaging RNA sequences as vector genomes into rAAV particles have recently been developed and are applicable hereto, replacing single-stranded (ss) DNA sequences as vector genomes for rAAV particles. See PCT / CN2022 / 075366, which is incorporated herein by reference in its entirety.

[0427] When the vector genome is RNA, as in, for example, PCT / CN2022 / 075366, for the sake of simplicity in description and declaration, the sequence elements described herein for DNA vector genomes should generally be considered applicable to RNA vector genomes when present in RNA vector genomes, except that the deoxyribonucleotides in the DNA sequence are the corresponding ribonucleotides in the RNA sequence (e.g., dT is equivalent to U and dA is equivalent to A), and / or elements in the DNA sequence are replaced in the RNA sequence with corresponding elements having the corresponding functions, or elements are omitted because their function is not required in the RNA sequence, and / or other elements required for the introduction of the RNA vector genome.

[0428] As used herein, coding sequences (e.g., as sequence elements of the rAAV vector genome in this document) are interpreted, understood, and considered to cover both DNA-coding sequences and RNA-coding sequences, and cover both. When it is a DNA-coding sequence, an RNA sequence can be transcribed from the DNA-coding sequence, and optionally, proteins can be further translated from the transcribed RNA sequence as needed. When it is an RNA-coding sequence, the RNA-coding sequence itself can be a functional RNA sequence for use, or an RNA sequence can be generated from the RNA-coding sequence (e.g., through RNA processing), or proteins can be translated from the RNA-coding sequence.

[0429] For example, the fusion protein coding sequence that encodes the fusion protein covers the fusion protein DNA coding sequence that expresses (indirectly expressed via transcription and translation) the fusion protein or the fusion protein RNA coding sequence that is translated (directly translated) the fusion protein.

[0430] For example, the gRNA coding sequence that encodes gRNA covers the gRNA DNA coding sequence that transcribes gRNA or the following gRNA RNA coding sequences: (1) are themselves functional gRNAs for use, or (2) are gRNAs produced from them (e.g., through RNA processing).

[0431] In some implementations of the rAAV RNA vector genome, the 5'-ITR and / or 3'-ITR, which serve as DNA packaging signals, may be unnecessary and can be at least partially omitted, while RNA packaging signals can be introduced.

[0432] In some implementations of the rAAV RNA vector genome, the promoter that drives the transcription of the DNA sequence may be unnecessary and can be omitted at least partially.

[0433] In some implementations of the rAAV RNA vector genome, the sequence encoding the polyA signal may be unnecessary and can be at least partially omitted, while a polyA tail may be introduced.

[0434] Similarly, other DNA elements of the rAAV DNA vector genome can be omitted or replaced with corresponding RNA elements, and / or additional RNA elements can be introduced to adapt to strategies for delivering RNA vector genomes via rAAV particles.

[0435] In another aspect, this disclosure provides a complex comprising a fusion protein of this disclosure or a polynucleotide (e.g., mRNA) encoding the fusion protein and a guiding nucleic acid (e.g., gRNA) of this disclosure.

[0436] In another aspect, this disclosure provides a ribonucleoprotein (RNP) comprising the fusion protein of this disclosure and the guide nucleic acid of this disclosure.

[0437] In another aspect, this disclosure provides a lipid nanoparticle (LNP) comprising RNA (e.g., mRNA) encoding a fusion protein of this disclosure and a guiding nucleic acid of this disclosure.

[0438] Pharmaceutical Composition

[0439] In another aspect, this disclosure provides a pharmaceutical composition comprising (1) the system of this disclosure, the carrier of this disclosure, the ribonucleoprotein of this disclosure, the lipid nanoparticle of this disclosure, or the cell of this disclosure; and (2) a pharmaceutically acceptable excipient.

[0440] In some embodiments, the pharmaceutical composition comprises rAAV particles at a concentration selected from about 1 × 10⁻⁶. 10 vg / mL, 2×10 10 vg / mL, 3×10 10 vg / mL, 4×10 10 vg / mL, 5×10 10 vg / mL, 6×10 10 vg / mL, 7×10 10 vg / mL, 8×10 10 vg / mL, 9×10 10 vg / mL, 1×10 11 vg / mL, 2×10 11 vg / mL, 3×10 11 vg / mL, 4×10 11 vg / mL, 5×10 11 vg / mL, 6×10 11 vg / mL, 7×10 11 vg / mL, 8×10 11 vg / mL, 9×10 11 vg / mL, 1×10 12 vg / mL, 2×10 12 vg / mL, 3×10 12 vg / mL, 4×10 12 vg / mL, 5×10 12 vg / mL, 6×10 12 vg / mL, 7×10 12 vg / mL, 8×10 12 vg / mL, 9×10 12 vg / mL, 1×10 13 vg / mL, or a concentration within a numerical range between any two of the aforementioned values, for example, a concentration of approximately 9 × 10 10 vg / mL to approximately 8×10 11 vg / mL. In some embodiments, the pharmaceutical composition is an injectable formulation.

[0441] In some embodiments, the volume of the injection solution is selected from about 1 μL, 10 μL, 50 μL, 100 μL, 150 μL, 200 μL, 250 μL, 300 μL, 350 μL, 400 μL, 450 μL, 500 μL, 550 μL, 600 μL, 650 μL, 700 μL, 750 μL, 800 μL, 850 μL, 900 μL, 950 μL, 1000 μL, and volumes within a numerical range between any two of the aforementioned values, for example, a concentration of about 10 μL to about 750 μL.

[0442] cell

[0443] The methods disclosed herein can be used to introduce the systems of this disclosure into cells and induce changes in the production of one or more cell products (such as antibodies, starch, ethanol, or any other desired product). Such cells and their progeny are within the scope of this disclosure.

[0444] In another aspect, this disclosure provides a cell or its descendants comprising a system of this disclosure. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a human cell.

[0445] In another aspect, this disclosure provides a cell or its progeny modified by the systems or methods of this disclosure. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is modified in vitro, in vivo, or ex vivo.

[0446] In some embodiments, the cell is a stem cell. In some embodiments, the cell is not a human embryonic stem cell. In some embodiments, the cell is not a human germ cell.

[0447] In some implementations, the cells are prokaryotic cells.

[0448] In some embodiments, the cell is a eukaryotic cell (e.g., animal cell, vertebrate cell, mammalian cell, non-human mammalian cell, non-human primate cell, rodent (e.g., mouse or rat) cell, human cell, plant cell, or yeast cell) or a prokaryotic cell (e.g., bacterial cell).

[0449] In some implementations, the cells are derived from plants or animals.

[0450] In some embodiments, the plant is a dicotyledonous plant. In some embodiments, the dicotyledonous plant is selected from soybean, cabbage (e.g., Chinese cabbage), rapeseed, brassica, watermelon, cantaloupe, potato, tomato, tobacco, eggplant, pepper, cucumber, cotton, alfalfa, and grape.

[0451] In some embodiments, the plant is a monocotyledonous plant. In some embodiments, the monocotyledonous plant is selected from rice, corn, wheat, barley, oats, sorghum, millet, grasses, Poaceae, Zizania, Avena, Coix, Hordeum, Oryza, Panicum (e.g., millet), Secale, Setaria (e.g., foxtail millet), Sorghum, Triticum, Zea, Cymbopogon, Saccharum (e.g., sugarcane), Phyllostachys, Dendrocalamus, Bambusa, and Yushania.

[0452] In some implementations, the animal is selected from pigs, cattle, sheep, goats, mice, rats, alpacas, monkeys, rabbits, chickens, ducks, geese, and fish (e.g., zebrafish).

[0453] In some embodiments, the cells are eukaryotic cells, such as mammalian cells, including human cells (primary human cells or established human cell lines). In some embodiments, the cells are non-human mammalian cells, such as cells derived from: non-human primates (e.g., monkeys), cows / bulls / cattle, sheep, goats, pigs, horses, dogs, cats, rodents (e.g., rabbits, mice, rats, hamsters, etc.). In some embodiments, the cells are derived from fish (e.g., salmon), birds (e.g., poultry, including chickens, ducks, geese), reptiles, shellfish (e.g., oysters, clams, lobsters, shrimp), insects, worms, yeast, etc. In some embodiments, the cells are derived from plants, such as monocots or dicots. In some embodiments, the plants are food crops, such as barley, cassava, cotton, peanuts or peanuts, corn, millet, oil palm fruit, potatoes, dried beans, rapeseed or canola, rice, rye, sorghum, soybeans, sugarcane, sugar beets, sunflowers, and wheat. In some embodiments, the plant is a cereal (barley, corn, millet, rice, rye, sorghum, and wheat). In some embodiments, the plant is a tuber (cassava and potato). In some embodiments, the plant is a sugar crop (beet and sugarcane). In some embodiments, the plant is an oil crop (soybean, peanut, canola, sunflower, and oil palm). In some embodiments, the plant is a fiber crop (cotton). In some embodiments, the plant is a tree (such as a peach or nectarine tree, apple or pear tree, nut tree (such as an apricot, walnut, or pistachio tree), or citrus tree (e.g., orange, grapefruit, or lemon tree), grass, vegetable, fruit, or algae. In some embodiments, the plant is a nightshade; a brassica plant; a lettuce plant; a spinach plant; or a capsicum plant.

[0454] Cotton, tobacco, asparagus, carrots, cabbage, broccoli, cauliflower, tomatoes, eggplant, peppers, lettuce, spinach, strawberries, blueberries, raspberries, blackberries, grapes, coffee, cocoa, etc.

[0455] In some embodiments, the cell is not located within an organism (such as a human or animal). In some embodiments, the cell is not a human embryonic stem cell. In some embodiments, the cell is not a human germ cell.

[0456] method

[0457] In another aspect, this disclosure provides a method for modifying target DNA, the method comprising contacting the target dsDNA with a system of the present disclosure, wherein the guide sequence is capable of hybridizing with a target sequence of the target dsDNA, wherein the target dsDNA is modified by the complex.

[0458] In another aspect, this disclosure provides a method for diagnosing, preventing, and / or treating a disease in a subject in need, the method comprising administering to the subject (e.g., a therapeutically effective amount / therapeutic dose) a system of this disclosure, a carrier of this disclosure, a ribonucleoprotein of this disclosure, a lipid nanoparticle of this disclosure, a cell of this disclosure, or a pharmaceutical composition of this disclosure, wherein the disease is associated with a target dsDNA, wherein the guide sequence is capable of hybridizing with a target sequence of the target dsDNA, wherein the target dsDNA is modified by the complex, and wherein the modification of the target dsDNA diagnoses, prevents, and / or treats the disease.

[0459] In some embodiments, the disease is selected from Angelman syndrome (AS), Alzheimer's disease (AD), thyroxine transporter amyloidosis (ATTR), thyroxine transporter amyloid cardiomyopathy (ATTR-CM), cystic fibrosis (CF), hereditary angioedema, diabetes, progressive pseudohypertrophic muscular dystrophy, Duchenne muscular dystrophy (DMD), Becker muscular dystrophy (BMD), spinal muscular atrophy (SMA), alpha-1-antitrypsin deficiency, Pompey's disease, myotonic dystrophy, Huntington's disease (HTT), fragile X syndrome, and Friedreich's ataxia. Amyotrophic lateral sclerosis (ALS), frontotemporal dementia, hereditary chronic kidney disease, hyperlipidemia, Leber congenital amaurosis (LCA), sickle cell disease, thalassemia (e.g., β-thalassemia), Parkinson's disease (PD), myelodysplastic syndrome (MDS), retinitis pigmentosa (RP), age-related macular degeneration (AMD), hepatitis B, non-alcoholic fatty liver disease (NAFLD), acquired immunodeficiency syndrome, corneal dystrophy (CD), hypercholesterolemia, familial hypercholesterolemia (FH), heart disease (e.g., hypertrophic cardiomyopathy (HCM)), and cancer.

[0460] In some implementations, the target dsDNA encodes mRNA, tRNA, ribosomal RNA (rRNA), microRNA (miRNA), noncoding RNA, long noncoding (lnc) RNA, nuclear RNA, interfering RNA (iRNA), small interfering RNA (siRNA), ribozyme, riboswitch, satellite RNA, microswitch, microzyme, or viral RNA.

[0461] In some implementations, the target dsDNA is eukaryotic DNA.

[0462] In some embodiments, the eukaryotic DNA is mammalian DNA (such as non-human mammalian DNA), non-human primate DNA, human DNA, plant DNA, insect DNA, bird DNA, reptile DNA, rodent (e.g., mouse, rat) DNA, fish DNA, nematode DNA, or yeast DNA.

[0463] In some implementations, the target dsDNA is in eukaryotic cells (e.g., human cells, non-human primate cells, or mouse cells).

[0464] In some implementations, application includes local or systemic application.

[0465] In some embodiments, administration includes intrathecal administration, intramuscular administration, intravenous administration, transdermal administration, intranasal administration, oral administration, mucosal administration, intraperitoneal administration, intracranial administration, intraventricular administration, or stereotactic administration.

[0466] In some implementations, administration is by injection or infusion.

[0467] In some implementations, the subjects are humans, non-human primates, or mice.

[0468] In some implementations, the level of the target dsDNA transcript (e.g., mRNA) is reduced in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the target dsDNA transcript (e.g., mRNA) in the subject prior to administration.

[0469] In some implementations, the level of the target dsDNA transcript (e.g., mRNA) is increased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the target dsDNA transcript (e.g., mRNA) in the subject prior to administration.

[0470] In some implementations, the level of the target dsDNA expression product (e.g., protein) is reduced in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the target dsDNA expression product (e.g., protein) in the subject prior to administration.

[0471] In some embodiments, the level of the target dsDNA expression product (e.g., protein) is increased in the subject by at least about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or more compared to the level of the target dsDNA expression product (e.g., protein) prior to administration. In some embodiments, the expression product is a functional mutant of the target dsDNA expression product.

[0472] In some implementations, the median survival of subjects who have the disease but receive the treatment is 5 days, 10 days, 20 days, 30 days, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 1.5 years, 2 years, 2.5 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, or longer than the median survival of subjects who have the disease but do not receive the treatment or the subject population.

[0473] The effective therapeutic dose can be administered via a single dose or multiple doses. Those skilled in the art will understand that the actual dose can vary significantly depending on a variety of factors, such as the choice of carrier, target cells, organism, tissue, general condition of the subject to be treated, the degree of transformation / modification sought, route of administration, mode of administration, and type of transformation / modification sought.

[0474] For example, the therapeutically effective dose of rAAV particles can be approximately 1.0E+8, 2.0E+8, 3.0E+8, 4.0E+8, 6.0E+8, 8.0E+8, 1.0E+9, 2.0E+9, 3.0E+9, 4.0E+9, 6.0E+9, 8.0E+9, 1.0E+10, 2.0E+10, 3.0E+10, 4.0E+10, 6.0E+10, 8.0E+10, 1.0E+11, 2.0E+11, 3.0E+11, 4.0E+11, 6.0E+11, 8.0E+11, 1.0E+12, 2.0E+12, 3.0E+12, 4.0E+12, 6 .0E+12, 8.0E+12, 1.0E+13, 2.0E+13, 3.0E+13, 4.0E+13, 6.0E+13, 8.0E+13, 1.0E+14, 2.0E+14, 3.0E+14, 4.0E+14, 6.0E+14, 8.0E+14, 1.0E+15, 2.0E+15, 3.0E+15, 4.0E+15, 6.0E+15, 8.0E+15, 1.0E+16, 2.0E+16, 3.0E+16, 4.0E+16, 6.0E+16, 8.0E+16, or 1.0E+17vg, or within a range of any two of those point values. vg represents the vector genome used for the applied rAAV particles.

[0475] In another aspect, this disclosure provides a method for detecting target dsDNA, the method comprising contacting the target dsDNA with a system of the present disclosure, wherein the target dsDNA is modified by the complex, and wherein the modification detects the target dsDNA. In some embodiments, the modification generates a detectable signal, such as a fluorescence signal.

[0476] Reagent test kit

[0477] In another aspect, this disclosure provides a kit comprising the fusion protein of this disclosure, the system of this disclosure, the polynucleotide of this disclosure, the vector of this disclosure, the RNP of this disclosure, the LNP of this disclosure, the delivery system of this disclosure, the cell of this disclosure, or the pharmaceutical composition of this disclosure, or any one, two or all of the components thereof.

[0478] In some embodiments, the kit further includes instructions for using the components contained therein, and / or instructions for combining with other components that may be available or required elsewhere.

[0479] In some embodiments, the kit further comprises one or more buffer solutions that can be used to dissolve any of the components contained therein, and / or to provide suitable reaction conditions for one or more of the components. Such buffer solutions may include one or more of the following: PBS, HEPES, Tris, MOPS, Na₂CO₃, NaHCO₃, NaB, or combinations thereof. In some embodiments, reaction conditions include an appropriate pH, such as an alkaline pH. In some embodiments, the pH is between 7 and 10.

[0480] In some implementations, any one or more of the kit components may be stored in a suitable container or at a suitable temperature (e.g., 4 degrees Celsius).

[0481] Further embodiments are described in the following examples, which are for illustrative purposes only and are not intended to limit the scope of this disclosure.

[0482] Example

[0483] The following embodiments are provided to further illustrate some implementations of this disclosure, but are not intended to limit the scope of the invention; by the exemplary nature of these embodiments, it will be understood that other procedures, methods or techniques known to those skilled in the art may be used alternatively.

[0484] Example 1. Development of a deaminase-free, glycosylation-based base editor

[0485] The aim is to develop a novel base editing system by utilizing a BER pathway triggered by depurination of G or A. A prototype of a deaminase-free, glycosylation-based base editor (gBE) is generated by fusing human MPG to the C-terminus of the Cas9-D10A nickase (nCas9) (SEQ ID NO:2), where MPG is designed to remove normal G or A (two purines with similar structures to hypoxanthine (Hx)) to create an AP site, and nCas9 causes a nick on the opposite non-edited strand. This allows damaged DNA to preferentially use the edited strand as a template for DNA repair and / or DNA replication. Therefore, incorporation of a nucleotide opposite the AP site via TLS will lead to different base editing outcomes ( Figure 1 a).

[0486] To facilitate the evaluation of gBE editing efficiency, a simple intron-splitting EGFP reporter system was developed, as previously reported

[11] . A disruptive point mutation (AG to CG or TG) is introduced at the intron boundary to generate an inactive splice acceptor (SA) signal, and a G to T or A to T conversion (relative to C to A or T to A conversion in the strand) is required to correct the mutation for proper splicing of the EGFP coding sequence, thereby activating EGFP expression and generating an EGFP signal ( Figure 1b and Figure 10 a). Flow cytometry can be used to detect EGFP fluorescence intensity ( Figure 10 d). Reporters G to T will be used to evaluate the guanine base editing efficiency of gBE (for this purpose, it will be referred to as glycosylation-based guanine base editor (gGBE)), and reporters A to T will be used to evaluate the adenine base editing efficiency of gBE (for this purpose, it will be referred to as glycosylation-based adenine base editor (gABE)).

[0487] gBE expression vectors were constructed that had wild-type human MPG (SEQ ID NO:1) (without the first N-terminal methionine (M) compared to the full-length wild-type human MPG (SEQ ID NO:9)) or unique forms of human MPG mutants reported in previous studies

[11] (i.e., MPGv0.2 (SEQ ID NO:4), MPGv1 (SEQ ID NO:5), MPGv2 (SEQ ID NO:6), and MPGv3 (SEQ ID NO:7)). After co-transfection of G-to-T or A-to-T reporter vectors with gBE expression vectors encoding a single target guide RNA (targeting sgRNA) that targets an intron missplicing mutation, all gBEs showed G-to-T conversion activity but not A-to-T conversion activity. Figure 1 c and Figure 10 b). The difference in substrate recognition capabilities between the normal G and A substrates of the tested MPG can explain the A editing failure in this system.

[0488] As used herein, the term "conversion activity" refers to the activity of the gBE of this disclosure in converting a target deoxyribonucleotide into an outcome deoxyribonucleotide, and the outcome deoxyribonucleotide may or may not be specified as a specific type of deoxyribonucleotide, such as G to T. As used herein, the term "(base) editing efficiency" refers to the activity of the gBE of this disclosure in converting a target deoxyribonucleotide into an outcome deoxyribonucleotide, and the outcome deoxyribonucleotide may or may not be specified as a specific type of deoxyribonucleotide, such as G to T. Where neither the outcome deoxyribonucleotide with respect to conversion activity nor the outcome deoxyribonucleotide with respect to (base) editing efficiency is specified, or where both the outcome deoxyribonucleotide with respect to conversion activity and (base) editing efficiency are specified as the same one or more specific types of deoxyribonucleotide, they may refer to the same performance of the gBE of this disclosure and may be used interchangeably.

[0489] For G base editing, in cultured HEK293T cells, gBE (hereinafter referred to as gGBEv3) (SEQ ID NO:15) containing MPGv3 (SEQ ID NO:7) exhibited the highest G-to-T base editing efficiency (4.33%) compared to gGBE (SEQ ID NO:14) with WT hMPG (SEQ ID NO:1) (0.03%). Figure 1 c) showed a significant 144-fold increase in G base editing efficiency. As two negative controls, no transforming activity was observed with gBE (SEQ ID NO:13) containing inactivated MPG (SEQ ID NO:3) (dMPG, carrying E125A, Y127A, and H136A mutations) or gGBEv3 (“NT”) along with non-targeting sgRNA. Figure 10 c). Thus, it was demonstrated that the G-to-T conversion of gGBEv3 depends on the catalytic G-removal domain of MPG guided by specific sgRNA.

[0490] Example 2. Further enhancement of G-editing efficiency in gGBE

[0491] To further enhance the G-to-T activity of gGBEv3, the protein was engineered by subjecting several rounds of appropriate mutagenesis to MPGv3 contained in gGBEv3, and the G-editing efficiency was evaluated using the G-to-T reporter from Example 1. Figure 2 a). Based on MPG structural analysis [32, 33], a 17 aa-long region (R163-V179) of MPGv3 from R163 to V179 was selected. This region forms a pocket around the targeted G base in the MPG-DNA complex model. Figure 11 a).

[0492] On the one hand, mutating gGBEv3 to have G174R or D175R produces new base editors gGBEv3.1 (SEQ ID NO:17) (containing MPGv3 with an additional G174R substitution; referred to as MPGv3.1 (SEQ ID NO:16)) and gGBEv3.2 (SEQ ID NO:19) (containing MPGv3 with an additional D175R substitution; referred to as MPGv3.2 (SEQ ID NO:18)). Figure 2 a). The G-to-T conversion activity of gGBEv3.2 (10.22%) was found to be approximately 1.78 times that of gGBEv3 (5.73%). Figure 2 b and Figure 11 b).

[0493] On the other hand, the R163-V179 region was scanned to show sequential substitutions (X to N) to asparagine (Asn / N). Interestingly, gGBEv3.3 (SEQ ID NO:21) (containing MPGv3 carrying an additional substituted C178N; referred to as MPGv3.3 (SEQ ID NO:20)) was obtained, which showed a significantly increased G to T conversion activity (31.37%), approximately 5.5 times that of gGBEv3 (5.73%). Figure 2 a and Figure 11 c).

[0494] In summary, gGBEv4 (SEQ ID NO:23) (containing MPGv3 carrying additional substitutes for both D175R and C178N; referred to as MPGv4 (SEQ ID NO:22)) achieves a synergistic enhancement in G-to-T editing efficiency (39.57%), approximately 6.9 times that of gGBEv3 (5.73%). Figure 2 b and Figure 11 b).

[0495] Furthermore, the orientations of nCas9 and MPG were varied to observe whether editing efficiency was related to their positional relationship. It was found that gGBEv4 fused with MPG at the C-terminus of nCas9 had slightly higher editing efficiency than gGBEv4 fused with MPG at the N-terminus of nCas9 (34.6% vs. 25.9%). Figure 11 d).

[0496] In addition, another round of mutagenesis and screening was conducted based on gGBEv4 to improve G editing efficiency. The R163-V179 region of MPGv4 was mutated by sequentially replacing amino acids with different properties, including glutamic acid (with a negatively charged side chain), valine (with a small hydrophobic side chain), glycine (without a side chain), or tyrosine (with a large hydrophobic side chain) (X to E, V, G, or Y). Compared to gGBEv4, the three base editors, namely gGBEv4.1 (SEQ ID NO:25) (containing MPGv4 with an additional I170V substitution; referred to as MPGv4.1 (SEQ ID NO:24)), gGBEv4.2 (SEQ ID NO:27) (containing MPGv4 with an additional S169G (or N169G, if compared to WT MPG); referred to as MPGv4.2 (SEQ ID NO:26)), and gGBEv4.3 (SEQ ID NO:29) (containing MPGv4 with an additional R163Y (or G163Y, if compared to WT MPG); referred to as MPGv4.3 (SEQ ID NO:28)), showed editing efficiencies increased by 1.06 times, 1.28 times, and 1.09 times, respectively. Figure 12 Combining the above three effective mutations yielded two additional base editors (gGBEv5.1 with S169G and I170V and gGBEv5.2 with R163Y, S169G, and I170Y), but no further enhancement was found compared to gGBEv4.1 and gGBEv4.2. Figure 13 ).

[0497] Encouraged by the above findings regarding gGBEv3.2, T199R, S230R, Q294R, or D295R were further individually introduced into gGBEv4.2. It was found that compared to gGBEv4.2 (SEQ ID NO:27) (49.93%) and gGBE with WT MPG (SEQ ID NO:14) (0.03%), gGBEv6.3 (SEQ ID NO:12) (containing MPGv4.2 plus an additional substituted Q294R; referred to as MPGv6.3 (SEQ ID NO:8)) (50.83%) further enhanced editing efficiency by approximately 1.02 times and 1.694 times, respectively. Figure 2 b and Figure 13 The amino acid sequences of gGBEv5.1, v5.2, v6.1, v6.2 and v6.4 are shown in SEQ ID NO:31, 33, 35, 37 and 39, respectively, and the corresponding MPGv5.1, v5.2, v6.1, v6.2 and v6.4 are shown in SEQ ID NO:30, 32, 34, 36 and 38, respectively.

[0498] Next, the enhanced G-editing efficiency of the obtained gGBEs (v3, v4, v4.2, v6.3, and gGBE with WT MPG) was verified at two endogenous genomic loci in cultured HEK293T cells. Cells were transfected with constructs encoding each gGBE, as well as mCherry and sgRNA targeting site 3 or site 10, and mCherry was targeted. + Cells were sorted by FACS for target depth sequencing analysis. A gradual increase in overall G-editing efficiency was observed at G7, specifically at site 3 from 6.4% to 78.5%, and at site 10 from 7.5% to 80.3%. Figure 2 c and Figure 2d) This confirms that gGBEv6.3 is indeed the optimal form of gGBE. G-to-C and G-to-T conversions were the main events at these two sites, and a small percentage of G-to-A conversions (4.6% at site 3 and 5.0% at site 10) and a small number of insertions and deletions (8.5% for site 3 and 4.3% for site 10) were detected in gGBEv6.3. These results indicate that the four rounds of mutagenesis described above have effectively optimized the activity of gGBE for G-to-C and G-to-T base editing. In summary, engineered gGBEv6.3 (SEQ ID NO:12) (carrying the G163R, N169G, D175R, C178N, S198A, K202A, G203A, S206A, K210A, and Q294R mutations on MPGv6.3 (SEQ ID NO:8) exhibits the highest G editing efficiency and will be used for further research.

[0499] Example 3. Characterization of gGBEv6.3 at human genomic DNA sites

[0500] The editing profile of gGBEv6.3 was characterized by targeting 24 endogenous genomic loci, most of which had been used in previous base editing studies [9, 35, 36]. gGBEv6.3 was found to achieve efficient G base editing (ranging from 27.6% to 76.5%), primarily through G-to-C and G-to-T conversions, with virtually no A, C, or T editing at all 24 sites examined. Figure 3 a and 3c and Figure 14 (ac). In all cells examined, the ratio of G to C / T (G to Y, Y = C and T) to G to A / T / C conversion was high (range 0.72 to 0.94). Figure 3 b). A small percentage of G to A conversions were detected. Figure 3 a and Figure 14 e.g., this is consistent with previous results from AYBE

[11] and CGBE[6-10]. gGBEv6.3 was also found to induce insertions and deletions at 24 edit sites, with frequencies ranging from 4.7% to 30.0%. Figure 14 h). Furthermore, it was found that the editable range of gGBEv6.3 is positions 1 to 14, and the optimal editing window with high G-conversion efficiency covers prototype interval sub-positions 6 to 11 (h). Figure 3 c) has the highest editing efficiency at position 7. Figure 3 c. Figure 14 d and S6i). Analysis of mid-target editing and sequences at all sites showed that gGBEv6.3 did not exhibit a significant NG motif preference for G transformation (d and S6i). Figure 14 j).

[0501] The efficiency of gGBEv6.3 in directive sequence-dependent off-target base editing was analyzed at several previously reported [11, 35] and computer simulation-predicted

[37] directive sequence-dependent off-target sites, and the ability of gGBEv6.3 to mediate directive sequence-independent off-target DNA editing in five dSaCas9 R loops was characterized using orthogonal R-loop assays, as previously reported in [11, 35]. Similar or lower percentages of editing were found at directive sequence-dependent off-target loci compared to previously identified adenine base editors [11, 35]. Figure 3 d and Figure 15 a). Furthermore, a very low frequency (average 1.1%) was detected at four of the five guide sequence-independent off-target sites. Figure 3 e and Figure 15 (b) and a slightly higher frequency was detected at one site with 12 Gs spanning the entire prototype spacer. In summary, gGBEv6.3 represents a highly efficient G-to-Y base editor with low off-target effects in mammalian cells.

[0502] On the other hand, gGBEv6.3 was also tested with A-to-T, C-to-G, and T-to-G reporters, and the editing efficiencies were 0.68%, 0.58%, and 1.81%, respectively, demonstrating its specificity for G editing, which is desirable for targeted base editing applications and reducing unwanted off-target editing.

[0503] Example 4. Potential gene editing applications of gGBE

[0504] gGBE's G-to-Y conversion capability allows for a variety of gene editing applications, including editing splice sites, introducing early stop codons (PTCs), and editing that bypasses PTCs. Figure 4 a). The inactive splice acceptor (SA) signaling with destructive point mutations, as illustrated by the intron-splitting EGFP reporter system used above, can be repaired using gGBE. Figure 1 b). Conversely, gGBE can be used to disrupt splicing signals by converting G within the splice donor site (“GT”) or splice acceptor site (“AG”) to other bases, resulting in exon skipping. To illustrate this application, the splice acceptor site of exon 45 in DMD (Duchenne muscular dystrophy) was edited with gGBEv6.3, and high G editing efficiency (up to 30.3%) was achieved with a high G-to-Y ratio (up to 0.88) when targeting DMD site 1. Figure 4 b、 Figure 4 c and Figure 16 a).

[0505] Another application of gGBE is to introduce PTCs to disrupt gene expression by converting TCA, TAC, or GAA codons to stop codons TGA, TAG, or TAA. Note that PTCs induced by GAA-to-TAA conversion can only be introduced using gGBEv6.3; currently, no other base editor can induce this type of PTC. PTCs were generated in cultured N2a cells by targeting three sites in the mouse Tyr (tyrosinase, associated with coat color) gene with gGBEv6.3. Figure 4 d) Achieve high G-editing efficiency (up to 46.3%) with a high G-to-Y ratio (up to 0.95). Figure 4 e and Figure 16 b). To further illustrate potential in vivo applications, mRNA encoding gGBEv6.3 and Tyr-targeting sgRNAs were co-injected into C57BL6 background (black fur) mouse zygotes, using 20 mouse embryos for each of the three Tyr-targeting sgRNAs. Two of the three sgRNAs were found to have highly efficient G editing ( Figure 17 a) When targeting Tyr site 3, the average PTC introduction efficiency was 50.9% ( Figure 17 bc). Similar to the data obtained above in cultured cell lines ( Figure 4 e), gGBE induces very few insertions and deletions in mouse embryos ( Figure 17 When gGBEv6.3 was used to target Tyr site 3, all 21 F0 pups showed G conversion, with an average efficiency of 59.4%. Figure 4 fh). This gGBEv6.3-induced G-to-Y editing results in albino or chimeric coat phenotypes in F0 mice. Figure 18 This indicates effective disruption of tyrosinase activity. Figure 4 i). Thus, gGBEv6.3 is demonstrated to be an effective G base editor not only in cell lines but also in mouse embryos.

[0506] discuss

[0507] Two major classes of deaminase-based base editors (dBEs), ABEs and CBEs, and their derivatives (such as AYBEs and CGBEs), perform base editing with the deamination of A or C as the first key step [3-11]. In this disclosure, a deaminase-free base editor was designed based on engineered MPGs, and a gGBE editor capable of achieving efficient G-to-C and G-to-T conversions in both cultured human cells and mouse embryos was generated. Engineered MPGs demonstrate that DNA glycosylases can be engineered into proteins that selectively remove specific nucleotide bases (such as G). The high editing efficiency of the gGBEs of this disclosure can be attributed to mutations in the MPG moiety that can promote its specific substrate selection or DNA binding activity, or both.

[0508] Of all pathogenic SNPs (60,372 in total), approximately 10% are C to G and 5% are T to G SNPs

[11] . Although C to G SNPs can be corrected by CGBE [6-10], efficient sgRNAs are not frequently found due to PAM limitations and narrow editing windows. Compared to CGBE, the gGBE of this disclosure increases the chance of finding efficient sgRNAs by targeting the relative strand. For T to G SNPs, no base editor currently can efficiently induce G to T (or C to A in the relative strand). Therefore, the gGBE of this disclosure greatly broadens the targeting range of base editors.

[0509] method

[0510] Molecular cloning

[0511] Using standard molecular cloning techniques, the base editor construct used in this study was cloned into a mammalian expression plasmid backbone under the control of the EF1α promoter. Intron splitting EGFP reporters were engineered as previously described

[11] . Briefly, corresponding mutations were performed at the splice acceptor site via PCR to construct A-to-T or G-to-T reporters, respectively. Mutations at the splice acceptor site resulted in inactive EGFP from the unspliced ​​EGFP transcript. Inspired by previous discoveries of base editors [7, 10], the 68th base of the intron sequence was modified (G to C) to introduce artificial PAM onto the template strand, so that the corresponding mutation at the splice acceptor site was at position 6 of the prototype spacer. Proper splicing of the EGFP coding sequence requires transversion correction in either the A-to-T or G-to-T reporter. MPG mutagenesis libraries were designed and generated as previously described

[46] . Seventeen aa-length regions (R163-V179) of MPGv3, from R163 to V179, were selected for protein engineering. MPG mutants with BpiI, MPG-G174R / D175R / T199R / S230R / Q294R / D295R mutants, or corresponding combinations thereof, were constructed via site-directed mutagenesis using PCR. Sequential asparagine / glutamate / valine / glycine / tyrosine substitutions (X to N, E, V, G, or Y) were designed, and the oligonucleotides encoding these mutants were annealed and ligated into the corresponding BpiI-digested backbone vectors. gRNA oligonucleotides were annealed and ligated into the BpiI site. Unless otherwise indicated, each mutation in MPG was numbered based on a full-length wild-type human MPG (SEQ ID NO:9) with a first N-terminal Met.

[0512] Analysis of target sequencing data

[0513] The target sequencing data analysis was previously described

[11] . In short, target amplicon sequencing reads were processed using fastp with default parameters

[47] . Cleaned pairs were then merged using FLASH v1.2.11. Sequences amplified from individual targets were demultiplexed using fastx barcode splitter.pl from fastx_toolkit (0.0.14).

[0514] Further amplicon sequencing analysis was performed using CRISPResso2

[48] . Modifications were quantified using a 10bp window centered around the center of the 20bp gRNA. Otherwise, the default parameters were used for analysis. G-to-C purity was calculated as G-to-C editing efficiency / (G-to-C editing efficiency + G-to-T editing efficiency + G-to-A editing efficiency). G-to-Y conversion ratio was calculated as (G-to-C editing efficiency + G-to-T editing efficiency) / (G-to-C editing efficiency + G-to-T editing efficiency + G-to-A editing efficiency).

[0515] Cell culture, transfection and flow cytometry analysis

[0516] HEK293T cells were cultured in DMEM (catalog number 11995065, Gibco) supplemented with 10% fetal bovine serum (catalog number 04-001-1ACS, BI) and 0.1 mM non-essential amino acids (catalog number 11140-050, Gibco). Cells were grown in an incubator at 37°C and 5% CO2. MPG mutant selection was performed in 48-well plates. One day before transfection, 3 × 10⁶ cells were cultured in each well. 4 HEK293T cells were seeded in 250 μL of complete growth medium in 48-well plates. After 12 h, 500 ng of GBE plasmid and 500 ng of A-to-T or G-to-T reporter plasmid were co-transfected into the cells with 2 μg polyethyleneimine (PEI) (DNA / PEI ratio 1:2) / well. For HEK293T cell transfection for FACS, 5 × 10⁶ cells per well were seeded the day before transfection. 5 Cells were seeded in 12-well plates containing 1 ml of complete growth medium. After 14–16 h, 2 μg of gGBE-sgRNA plasmid was transfected into the cells using PEI (DNA / PEI ratio 1:2). To target DMD site 2 with NG PAM, the PAM flexible Cas9 variant SpG was used. Orthogonal R-loop assays were performed as previously described [1, 2]. Briefly, 1 μg of gGBE plasmid containing sgRNA targeting site 3 and 1 μg of dSaCas9 plasmid containing the corresponding sgRNA targeting five off-target sites to generate R-loops were co-transfected into 12-well plates using PEI (DNA / PEI ratio 1:2).

[0517] In HEK293T cells, the expression of mCherry, BFP, and EGFP fluorescence was analyzed by BD FACS Aria III or Beckman CytoFLEX S 48 h post-transfection. Flow cytometry results were analyzed using FlowJo V10.5.3. mCherry was used for mid-target editing efficiency evaluation. + BFP + and EGFP + Gating strategies in cell identification provide Figure 10 d.

[0518] animal

[0519] The experiments involving mice were approved by the Biomedical Research Ethics Committee of HuidaGene Therapeutics Co., Ltd. Superovulated C57BL / 6 females (4 weeks old) were mated with C57BL / 6 males (8 weeks old), and females from the ICR strain were used as foster mothers. Mice were maintained in a specific pathogen-free facility under 12-hour dark-light cycles and constant temperature (20℃–26℃) and humidity (40%–60%) conditions.

[0520] In vitro transcription of gGBE mRNA and Tyr-sgRNA

[0521] The gGBE plasmid was structured as follows: standard PCR amplification was performed using Phanta Max ultra-fidelity DNA polymerase (Vazyme Biotech Co., Ltd.), followed by assembly using Gibson assembly premix (NEB E2611L) and transformation into chemicompetent DH5α cells. The gGBE plasmid was linearized using FastDigest KpnI restriction enzyme (Thermo Fisher Scientific), purified using a gel extraction kit (Omega Scientific), and used as a template for in vitro transcription (IVT) using the mMESSAGE mMACHINE T7 Ultra kit (Life Technologies). For Tyr-sgRNA preparation, the T7 promoter sequence was added to the sgRNA template via PCR amplification using px330 (Addgene, serial number 42230). The T7-Tyr-sgRNA PCR product was purified using a gel extraction kit (Omega) and used as a template for sgRNA IVT using the MEGAshortscript T7 kit (Lifetech Corporation). gGBE mRNA and Tyr-sgRNA were purified using the MEGAclear kit (Lifetech Corporation) and eluted in RNase-free water. The in vitro transcribed RNA was aliquoted and stored at -80°C until use. Prior to microinjection, a mixture of gGBE mRNA and Tyr-sgRNA was prepared by centrifugation at 14,000 rpm for 10 min at 4°C, and the supernatant was transferred to a new 0.2 mL PCR tube for injection.

[0522] Microinjection of gGBE mRNA and Tyr-sgRNA into mouse fertilized eggs

[0523] Superovulated C57BL / 6 females (4 weeks old) were mated with C57BL / 6 males, and fertilized embryos were collected from the fallopian tubes 21 hours after hCG injection. For fertilized egg injection, a mixture of gGBE mRNA (100 ng / μL) and Tyr-sgRNA (100 ng / μL) was injected into the cytoplasm of single-cell embryos in M2 medium droplets using a FemtoJet microinjector (Eppendorf) at a constant flow rate setting. The injected embryos were cultured at 37°C in air at 5% CO2 in M16 medium containing amino acids for 2 hours, and then transferred to the fallopian tubes of pseudopregnant ICR surrogate mothers 0.5 days (0.5-dpc) after mating.

[0524] Target sequencing of endogenous sites

[0525] 72 h post-transfection, 10,000 mCherry-positive cells were isolated using FACS. Genomic DNA was extracted by adding 40 μl of lysis buffer and 1 μL of proteinase K (catalog number PD101-01, Novizan) directly to each sorted cell tube. The genomic DNA / lysis buffer mixture was incubated at 55 °C for 45 min, followed by enzyme inactivation at 95 °C for 10 min. Site-specific primers were used to amplify regions of interest at the target sites via PCR. PCR reactions were performed using Max Ultra Fidelity DNA Polymerase (catalog number P505-d3, Novizan) as follows: 5 min at 95°C, 30 cycles of 95°C for 15 s, 60°C for 15 s, 72°C for 30 s, and a final extension at 72°C for 5 min. PCR products were purified using a universal DNA purification kit (TIANGEN) according to the manufacturer's instructions and analyzed by Sanger sequencing (Genewiz). Amplicones were ligated to adaptors and sequenced on an Illumina MiSeq platform. Prototype spacer sequences for each genomic locus are listed in Table 1.

[0526] Statistical analysis

[0527] Statistical tests performed using Graphpad Prism 8 include a two-tailed unpaired two-sample t-test following one-way ANOVA or Dunnett's multiple comparison test.

[0528] References

[0529] 1.Porto EM,Komor AC,Slaymaker IM,et al.;Base editing:advances andtherapeutic opportunities.Nat Rev Drug Discov 2020;19(12):839-859.doi:10.1038 / s41573-020-0084-6.

[0530] 2. Rees HA, Liu DR; Base editing: precision chemistry on the genome and transcriptome of living cells. Nat Rev Genet 2018; 19(12):770-788.doi:10.1038 / s41576-018-0059-1.

[0531] 3.Komor AC,Zhao KT,Packer MS,et al.;Improved base excision repairinhibition and bacteriophage Mu Gam protein yields C:G-to-T:A base editorswith higher efficiency and product purity.Sci Adv 2017;3(8):eaao4774.doi:10.1126 / sciadv.aao4774.

[0532] 4.Komor AC,Kim YB,Packer MS,et al.;Programmable editing of a targetbase in genomic DNA without double-stranded DNA cleavage.Nature 2016;533(7603):420-4.doi:10.1038 / nature17946.

[0533] 5.Gaudelli NM,Komor AC,Rees HA,et al.;Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage.Nature 2017;551(7681):464-471.doi:10.1038 / nature24644.

[0534] 6.Zhao D,Li J,Li S,et al.;Glycosylase base editors enable C-to-A andC-to-G base changes.Nat Biotechnol 2021;39(1):35-40.doi:10.1038 / s41587-020-0592-2.

[0535] 7.Kurt IC,Zhou R,Iyer S,et al.;CRISPR C-to-G base editors forinducing targeted DNA transversions in human cells.Nat Biotechnol 2021;39(1):41-46.doi:10.1038 / s41587-020-0609-x.

[0536] 8.Koblan LW,Arbab M,Shen MW,et al.;Efficient C*G-to-G*C base editorsdeveloped using CRISPRi screens,target-library analysis,and machinelearning.Nat Biotechnol 2021;39(11):1414-1425.doi:10.1038 / s41587-021-00938-z.

[0537] 9.Chen L,Park JE,Paa P,et al.;Programmable C:G to G:C genome editingwith CRISPR-Cas9-directed base excision repair proteins.Nat Commun 2021;12(1):1384.doi:10.1038 / s41467-021-21559-9.

[0538] 10.Yuan T,Yan N,Fei T,et al.;Optimization of C-to-G base editors withsequence context preference predictable by machine learning methods.NatCommun 2021;12(1):4902.doi:10.1038 / s41467-021-25217-y.

[0539] 11.Tong H,Wang X,Liu Y,et al.;Programmable A-to-Y base editing byfusing an adenine base editor with an N-methylpurine DNA glycosylase.NatBiotechnol 2023.doi:10.1038 / s41587-022-01595-6.

[0540] 12.Mok BY,de Moraes MH,Zeng J,et al.;A bacterial cytidine deaminasetoxin enables CRISPR-free mitochondrial base editing.Nature 2020;583(7817):631-637.doi:10.1038 / s41586-020-2477-4.

[0541] 13.Lei Z,Meng H,Liu L,et al.;Mitochondrial base editor inducessubstantial nuclear off-target mutations.Nature 2022;606(7915):804-811.doi:10.1038 / s41586-022-04836-5.

[0542] 14.Cho SI,Lee S,Mok YG,et al.;Targeted A-to-G base editing in humanmitochondrial DNA with programmable deaminases.Cell 2022;185(10):1764-1776e12.doi:10.1016 / j.cell.2022.03.039.

[0543] 15.Ladstatter S,Tachibana-Konwalski K;A Surveillance MechanismEnsures Repair of DNA Lesions during Zygotic Reprogramming.Cell 2016;167(7):1774-1787 e13.doi:10.1016 / j.cell.2016.11.009.

[0544] 16.Marrone A,Ballantyne J;Hydrolysis of DNA and its molecularcomponents in the dry state.Forensic Sci Int Genet 2010;4(3):168-77.doi:10.1016 / j.fsigen.2009.08.007.

[0545] 17.Lindahl T;Instability and decay of the primary structure ofDNA.Nature1993;362(6422):709-15.doi:10.1038 / 362709a0.

[0546] 18.Lee SJ,Choi MY,Miller RE;Vibrational spectroscopy of xanthine insuperfluid helium nanodroplets.Chemical Physics Letters 2009;475(1):24-29.doi:10.1016 / j.cplett.2009.05.016.

[0547] 19.Hindi NN,Elsakrmy N,Ramotar D;The base excision repair process:comparison between higher and lower eukaryotes.Cell Mol Life Sci 2021;78(24):7943-7965.doi:10.1007 / s00018-021-03990-9.

[0548] 20.Thompson PS,Cortez D;New insights into abasic site repair andtolerance.DNA Repair(Amst)2020;90:102866.doi:10.1016 / j.dnarep.2020.102866.

[0549] 21.Bauer NC,Corbett AH,Doetsch PW;The current state of eukaryotic DNAbase damage and repair.Nucleic Acids Res 2015;43(21):10083-101.doi:10.1093 / nar / gkv1136.

[0550] 22.Jacobs AL,Schar P;DNA glycosylases:in DNA repair andbeyond.Chromosoma 2012;121(1):1-20.doi:10.1007 / s00412-011-0347-4.

[0551] 23.Robertson AB,Klungland A,Rognes T,et al.;DNA repair in mammaliancells:Base excision repair:the long and short of it.Cell Mol Life Sci 2009;66(6):981-93.doi:10.1007 / s00018-009-8736-z.

[0552] 24.Gallagher PE,Brent TP;Partial purification and characterizationof3-methyladenine-DNA glycosylase from human placenta.Biochemistry 1982;21(25):6404-9.doi:10.1021 / bi00268a013.

[0553] 25.Asaeda A,Ide H,Asagoshi K,et al.;Substrate specificity of humanmethylpurine DNA N-glycosylase.Biochemistry 2000;39(8):1959-65.doi:10.1021 / bi9917075.

[0554] 26.Lau AY,Wyatt MD,Glassner BJ,et al.;Molecular basis fordiscriminating between normal and damaged bases by the human alkyladenineglycosylase,AAG.Proc Natl Acad Sci U S A 2000;97(25):13573-8.doi:10.1073 / pnas.97.25.13573.

[0555] 27.Wyatt MD,Samson LD;Influence of DNA structure on hypoxanthine and1,N(6)-ethenoadenine removal by murine 3-methyladenine DNAglycosylase.Carcinogenesis 2000;21(5):901-8.doi:10.1093 / carcin / 21.5.901.

[0556] 28.Saparbaev M,Laval J;Excision of hypoxanthine from DNA containingdIMP residues by the Escherichia coli,yeast,rat,and human alkylpurine DNAglycosylases.Proc Natl Acad Sci U S A 1994;91(13):5873-7.doi:10.1073 / pnas.91.13.5873.

[0557] 29.O'Brien PJ,Ellenberger T;Dissecting the broad substratespecificity of human3-methyladenine-DNA glycosylase.J Biol Chem 2004;279(11):9750-7.doi:10.1074 / jbc.M312232200.

[0558] 30.Berdal KG,Johansen RF,Seeberg E;Release of normal bases fromintact DNA by a native DNA repair enzyme.EMBO J 1998;17(2):363-7.doi:10.1093 / emboj / 17.2.363.

[0559] 31.Kay JE,Corrigan JJ,Armijo AL,et al.;Excision of mutagenicreplication-blocking lesions suppresses cancer but promotes cytotoxicity andlethality in nitrosamine-exposed mice.Cell Rep 2021;34(11):108864.doi:10.1016 / j.celrep.2021.108864.

[0560] 32.Vallur AC,Maher RL,Bloom LB;The efficiency of hypoxanthineexcision by alkyladenine DNA glycosylase is altered by changes in nearestneighbor bases.DNA Repair(Amst)2005;4(10):1088-98.doi:10.1016 / j.dnarep.2005.05.008.

[0561] 33.Lau AY,Scharer OD,Samson L,et al.;Crystal structure of a humanalkylbase-DNA repair enzyme complexed to DNA:mechanisms for nucleotideflipping and base excision.Cell 1998;95(2):249-58.doi:10.1016 / s0092-8674(00)81755-9.

[0562] 34.Connor EE,Wyatt MD;Active-site clashes prevent the human 3-methyladenine DNA glycosylase from improperly removing bases.Chem Biol 2002;9(9):1033-41.doi:10.1016 / s1074-5521(02)00215-6.

[0563] 35.Richter MF,Zhao KT,Eton E,et al.;Phage-assisted evolution of anadenine base editor with improved Cas domain compatibility and activity.NatBiotechnol2020;38(7):883-891.doi:10.1038 / s41587-020-0453-z.

[0564] 36.Doman JL,Raguram A,Newby GA,et al.;Evaluation and minimization ofCas9-independent off-target DNA editing by cytosine base editors.NatBiotechnol2020;38(5):620-628.doi:10.1038 / s41587-020-0414-6.

[0565] 37.Bae S,Park J,Kim JS;Cas-OFFinder:a fast and versatile algorithmthat searches for potential off-target sites of Cas9 RNA-guided endonucleases.Bioinformatics2014;30(10):1473-5.doi:10.1093 / bioinformatics / btu048.

[0566] 38.Zuo E,Sun Y,Wei W,et al.;GOTI,a method to identify genome-wideoff-target effects of genome editing in mouse embryos.Nat Protoc 2020;15(9):3009-3029.doi:10.1038 / s41596-020-0361-1.

[0567] 39.Zuo E,Sun Y,Wei W,et al.;Cytosine base editor generatessubstantial off-target single-nucleotide variants in mouse embryos.Science2019;364(6437):289-292.doi:10.1126 / science.aav9973.

[0568] 40.Yan N,Feng H,Sun Y,et al.;Cytosine base editors induce off-targetmutations and adverse phenotypic effects in transgenic mice.Nat Commun 2023;14(1):1784.doi:10.1038 / s41467-023-37508-7.

[0569] 41.Lei Z,Meng H,Lv Z,et al.;Detect-seq reveals out-of-protospacerediting and target-strand editing by cytosine base editors.Nat Methods 2021;18(6):643-651.doi:10.1038 / s41592-021-01172-w.

[0570] 42.Chen L,Zhang S,Xue N,et al.;Engineering a precise adenine baseeditor with minimal bystander editing.Nat Chem Biol 2023;19(1):101-110.doi:10.1038 / s41589-022-01163-8.

[0571] 43.Kim YB,Komor AC,Levy JM,et al.;Increasing the genome-targetingscope and precision of base editing with engineered Cas9-cytidine deaminasefusions.Nat Biotechnol2017;35(4):371-376.doi:10.1038 / nbt.3803.

[0572] 44.Wang Y,Zhao D,Sun L,et al.;Engineering of the Translesion DNASynthesis Pathway Enables Controllable C-to-G and C-to-A Base Editing inCorynebacterium glutamicum.ACS Synth Biol 2022;11(10):3368-3378.doi:10.1021 / acssynbio.2c00265.

[0573] 45.Sun N,Zhao D,Li S,et al.;Reconstructed glycosylase base editorsGBE2.0 with enhanced C-to-G base editing efficiency and purity.Mol Ther 2022;30(7):2452-2463.doi:10.1016 / j.ymthe.2022.03.023.

[0574] 46.Tong H, Huang J, Xiao Q, et al.; High-fidelity Cas13 variants fortargeted RNA degradation with minimal collateral effects. Nat Biotechnol 2023; 41(1):108-119.doi:10.1038 / s41587-022-01419-7.

[0575] 47. Chen S, Zhou Y, Chen Y, et al.; fastp: an ultra-fast all-in-one FASTQpreprocessor. Bioinformatics 2018; 34(17):i884-i890.doi:10.1093 / bioinformatics / bty560.

[0576] 48. Clement K, Rees H, Canver MC, et al.; CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat Biotechnol 2019; 37(3):224-226.doi:10.1038 / s41587-019-0032-3.

[0577] Example 5. Development of an orthogonal base editor based on engineered glycosylation enzymes

[0578] Inspired by the development of gGBE in the above examples, a glycosylation-based strategy without deaminases was developed to create thymine and cytosine base editors. Since the three pyrimidine bases (i.e., T, C, and U) are structurally similar, it is hypothesized that canonical T or C excision could be achieved by engineering certain uracil DNA glycosylation enzymes (UNG). T or C excision would generate apurinol / pyrimidine-free (AP) sites, which would then trigger the base excision repair (BER) pathway and facilitate direct T or C editing. Figure 19 ab).

[0579] Alternative splicing and transcription from two different initiation sites result in two distinct human UNG isoforms: mitochondrial UNG1 (304 amino acids (aa)) (SEQ ID NO:54) and nuclear UNG2 (313 aa) (SEQ ID NO:133), each possessing a unique N-terminus that mediates translocation to the mitochondria or nucleus. 16 ( Figure 25Sequence alignment of human UNG1 and human UNG2 shows that amino acid residues at positions 1-35 of UNG1 are different from amino acid residues at positions 1-44 of UNG2, and the rest of UNG1 and UNG2 are the same.

[0580] Two human UNG1 variants (UNG1-Y147A and UNG1-N204D) have been engineered to remove T and C from the DNA, respectively. 17 Residues Y156 and N213 of UNG2 correspond to residues Y147 and N204 of UNG1, as determined by sequence alignment of UNG1 and UNG2. To edit nuclear DNA, two prototype gBEs were generated by fusing UNG2-Y156A (SEQ ID NO:135) and UNG2-N213D (SEQ ID NO:137) (where Y156A and N213D of human UNG2 are equivalent to Y147A and N204D of human UNG1, respectively) (with their N-terminal Met retained) to the C-terminus of Cas9-D10A nickase (nCas9) (SEQ ID NO:2) via adapter (SEQ ID NO:134). These prototypes are: a deaminase-free glycosylase-based thymine base editor (gTBE) (gTBEv0.1; SEQ ID NO:136) and a deaminase-free glycosylase-based cytosine base editor (gCBE) (gCBEv0.1; SEQ ID NO:138). Figure 19 ac).

[0581] T to G reporter items and C to G reporter items were established (as previously reported). 9 Two similar intron-splitting EGFP reporter systems were used to evaluate the editing efficiency of gTBE and gCBE, respectively. Figure 26 a). In these reporter compounds, the inactive splice acceptor AG to AT or AG to AC (SA) can only be repaired by T to G or C to G conversion, which leads to proper splicing of the EGFP coding sequence and EGFP expression and emission of the EGFP signal. Figure 26 b). Co-transfect the gTBE or gCBE encoding vector with a T-to-G or C-to-G reporter vector containing a target single guide RNA (target sgRNA) with a missplicing mutation.

[0582] It was observed that, compared with the negative control, gTBEv0.1 showed T-to-G conversion activity, and gCBEv0.1 showed C-to-G conversion activity. Figure 19 ce).

[0583] Given that the disordered N-terminal domain (NTD) of UNG contains protein-binding motifs and sites for post-translational modifications. 18This may constrain the targeted cleavage activity of the glycosylation enzyme domain in ssDNA. 19、20 Therefore, gTBEv0.2 (SEQ ID NO:140) (with UNG2Δ88-Y156A (SEQ ID NO:139) fused to the C-terminus of nCas9) and gCBEv0.2 (SEQ ID NO:142) (with UNG2Δ88-N213D (SEQ ID NO:141) fused to the C-terminus of nCas9) were constructed by deleting 88 amino acid residues at the N-terminus of UNG2 (denoted as "Δ88"). Figure 19 c) to eliminate potentially undesirable protein-protein interactions. 20-22 As described, in the case of an 88-amino acid residue deletion at the N-terminus, the remainder of human UNG2 will be identical to the remainder of human UNG1 with a corresponding 79-amino acid residue deletion at the N-terminus, and in this case, both UNG2 and UNG1 can be referred to as UNG, unless the context requires a distinction between UNG2 and UNG1. Unless otherwise indicated, amino acid mutations on UNG are numbered based on the full-length human UNG2.

[0584] gTBEv0.2 exhibited comparable T-to-G conversion activity to gTBEv0.1 (1.0% vs. 1.1%). Figure 19 d), while gCBEv0.2 exhibited a significantly increased C to G conversion activity compared to gCBEv0.1 (13.3% compared to 1.0%). Figure 19 e). Furthermore, both gTBEv0.3 (SEQ ID NO:143) (with UNG2Δ88-Y156A fused to the N-terminus of nCas9) and gCBEv0.3 (SEQ ID NO:154) (with UNG2Δ88-N213D fused to the N-terminus of nCas9) showed significantly higher editing efficiencies than gTBEv0.2 and gCBEv0.2 (with UNG2 mutant fused to the C-terminus of nCas9) (10.2% vs. 1.0% for gTBE and 51.4% vs. 13.3% for gCBE). Figure 19 The editing efficiency is improved by approximately 10 times and 3.9 times, respectively. This indicates that having the N'-UNG-nCas9-C' construction would be desirable for both gTBE and gCBE. Unless otherwise indicated, this construction is used in all subsequent embodiments.

[0585] For the negative control, no editing efficiency was found for any of the gTBE and gCBE mentioned above, as well as for non-targeted sgRNAs. Figure 19 de).

[0586] Furthermore, various forms of N-terminal deletions were applied to UNG2 in gTBE and gCBE. For gTBE, it was demonstrated that the deletion of 88 amino acid residues in gTBEv0.3 achieved a significantly increased editing efficiency compared to deletions of 44, 72, and 93 amino acid residues; and for gCBE, it was demonstrated that the deletion of 88 amino acid residues in gCBEv0.3 achieved comparable editing efficiency compared to deletions of 93 amino acid residues, and a significantly increased editing efficiency compared to deletions of 44 and 72 amino acid residues. Figure 27 ).

[0587] Furthermore, the orthogonality of gTBE and gCBE for base editing was examined. Despite being engineered from the same original DNA glycosylase UNG, no C-editing efficiency was observed for gTBEv0.3, and no T-editing efficiency was observed for gCBEv0.3. Figure 19 f). Therefore, two orthogonal base editors were developed: gTBE for direct T editing and gCBE for C editing.

[0588] Example 6. Evolution of gTBE with Enhanced Editing Efficiency

[0589] To further enhance the T-to-G conversion activity of gTBEv0.3, rational mutagenesis was performed to engineer the UNG moiety. The editing efficiency in cultured mammalian cells (HEK293T cells) was evaluated using a T-to-G reporter. Figure 20 a) Based on structural and functional analysis, WT UNG contains five conserved motifs required for effective glycosylation enzyme activity: a catalytic water activation ring, a proline-rich ring, a uracil-binding motif, a glycine-serine motif, and a leucine ring. 23-25 ( Figure 25 b). Since Y156 in the catalytic water-activated ring and N213 in the uracil-binding motif are crucial for the active transition from U-cleavage to T or C-cleavage, the sequential and spatially adjacent residues of these two residues were selected to examine their roles in the regulation of base cleavage activity. Figure 20 ab). Alanine scanning mutagenesis was performed by replacing all non-alanines (X to A) with alanine and alanine with valine (A to V) to cover all residues in the I150-L179 and L210-T217 regions.

[0590] Interestingly, gTBEv1.1 (SEQ ID NO:145) (v0.3 plus A214V) with UNG2Δ88-Y156A+A214V (SEQ ID NO:144) was obtained, which showed a significantly increased T-to-G conversion activity, approximately 2.68 times that of gTBEv0.3. Figure 28a). To examine whether any amino acid at position 214 performs better than valine, further site-saturation mutagenesis was performed focusing on the residues at position 214. gTBEv1.2 (SEQ ID NO:147) (v0.3 plus A214T) with UNG2Δ88-Y156A+A214T (SEQ ID NO:146) was obtained, exhibiting increased editing efficiency, approximately 1.06 times that of gTBEv1.1 (a). Figure 28 b).

[0591] On the other hand, the spatially adjacent residues of A214 that compress the DNA backbone 3' to the damaged Gly-Ser loop were examined. Figure 20 b), and obtained gTBEv1.3 (SEQ ID NO:149) (v0.3 plus Q259A) with UNG2Δ88-Y156A+Q259A (SEQ ID NO:148), which has increased editing efficiency, approximately 1.46 times that of gTBEv0.3. Figure 28 c).

[0592] Furthermore, it was found that compared to the T editing efficiency of gTBEv0.3, gTBEv2 (SEQ ID NO:151) (v0.3 plus the combination of A214T and Q259A) with UNG2Δ88-Y156A+A214T+Q259A (SEQ ID NO:150) showed a synergistic enhancement of approximately 2.7 times in T-to-G editing efficiency. Figure 20 c).

[0593] The study also scanned residues in the Q274-Y284 region, within or near the Leu insertion ring, by sequentially substituting amino acids with different properties (including arginine (with a positively charged side chain), aspartic acid (with a negatively charged side chain), or valine (with a small hydrophobic side chain)) (X to R, D, or V). gTBEv3 (SEQ ID NO:153) (v2 plus Y284D) with UNG2Δ88-Y156A+A214T+Q259A+Y284D (SEQ ID NO:152) showed an approximately 1.22-fold increase compared to gTBEv2. Figure 29 b) and increased by approximately 3.09 times compared to gTBEv0.3 ( Figure 20 c) Editing efficiency.

[0594] Improved T-editing efficiency with different gTBEs was validated at an endogenous genomic site in cultured mammalian cells (HEK293T cells). mCherry-positive cells were FACS-sorted after transfection with an integrated construct encoding each gTBE, a targeting sgRNA targeting site 9 in the CLYBL gene, and mCherry for fluorescence-activated cell sorting (FACS). Target depth sequencing analysis revealed a gradual increase in overall T-editing efficiency at T5, from 26.9% for gTBEv1.1 to 67.4% for gTBEv3, confirming gTBEv3 as the optimal form for base editing at the endogenous target, where T-to-S (i.e., T-to-C or T-to-G; S=C or G bases) transitions were the dominant event at this site. Figure 20 d).

[0595] These results demonstrate that the aforementioned mutagenesis effectively optimized the performance of gTBE for T-to-C and T-to-G base editing. In subsequent examples, gTBEv3, which exhibits the highest T-editing efficiency, was used.

[0596] Example 7. Characterization of gTBEv3 at human genomic DNA sites

[0597] The editing profile of gTBEv3 was characterized by targeting 20 endogenous genomic loci, most of which were used in previous base editing studies. 11、12、26、27 gTBEv3 was found to achieve efficient T-base editing (ranging from 24.3% to 81.5%). Figure 21 a and Figure 30 ab), and there were essentially no A, C, or G edits at any of the examination sites ( Figure 30 ce). The T-to-C or T-to-G transformation is the primary event ( Figure 30 fh), only a low percentage of T to A conversions were detected ( Figure 21 a and Figure 30 i) with previous information regarding gGBE 3 AYBE 9 and CGBE 11-15 The findings are consistent. The ratio of T to S and T conversions ranged from 0.68 to 0.97 (without insertions or deletions). Figure 21 b) and 0.41 to 0.92 (with insertions or deletions). Figure 30 j). gTBEv3 was found to induce insertions and deletions at 20 edit sites, with frequencies ranging from 5.2% to 45.2%. Figure 21 c). Furthermore, it was found that the editable range of gTBEv3 is positions 2 to 11, and the optimal editing window with high T-conversion efficiency covers prototype interval positions 3 to 7, with the highest editing efficiency at position 5. Figure 30b). Analysis of the mid-target edits and sequences at all test sites revealed no significant motif preference for T-transformation in gTBEv3. Figure 30 k).

[0598] In several computer simulations, predictions were made. 28 The off-target activity of gTBEv3 was analyzed at sequence-dependent off-target sites, and its activity was determined using orthogonal R-loop assays at five previously reported dSaCas9 R-loop sites. 9 , 29 The study characterized the ability of gTBEv3 to mediate guide sequence-independent off-target DNA editing. Very low editing percentages were found at all guide sequence-dependent off-target loci. Figure 21 de and Figure 31 Furthermore, it was detected at a very low frequency (average 1.1%) at all five guide sequence-independent off-target sites. Figure 21 f). In summary, gTBEv3 represents a highly efficient T-to-S base editor with low off-target effects in mammalian cells.

[0599] Example 8. Enhanced C editing efficiency of gCBE

[0600] To examine whether mutations arising from the engineering of gTBE would benefit the enhancement of gCBE activity, gCBEv1.1 was generated by introducing A214V into gCBEv0.3. Figure 22 a). When evaluated using C-to-G reporter compounds, gCBEv1.1 (SEQ ID NO:156) with UNG2Δ88-N213D+A214V (SEQ ID NO:155) was found to have significantly increased C-to-G conversion activity, approximately 1.34 times that of gCBEv0.3. Figure 32 a).

[0601] On the other hand, alanine scanning mutagenesis was performed on the D154-D189 region of UNG2 to examine its role in regulating base excision activity, and gCBEv1.2 (SEQ ID NO:158) (v0.3 plus K184A) with UNG2Δ88-N213D+K184A (SEQ ID NO:157) was obtained, which had a significantly increased C to G conversion activity of about 1.55 times compared with gCBEv0.3. Figure 32 b).

[0602] The combination of A214V and K184A was further investigated by combining these two mutations to generate gCBEv2 (SEQ ID NO: 160) with UNG2Δ88-N213D+K184A+A214V (SEQ ID NO: 159), thereby achieving approximately 1.3 times the C to G editing efficiency compared to gCBEv0.3. Figure 22 b). Further validation of improved C-editing efficiency in different gCBEs was achieved by targeting endogenous genomic sites, and the overall C-editing efficiency at C2 of site 28 gradually increased from 18.2% to 37.2%. Figure 33 a).

[0603] The editing profile of gCBEv2 was characterized by targeting 16 endogenous genomic loci, demonstrating an effective C-base editing efficiency ranging from 31.8% to 77.7%. Figure 22 c and Figure 33 bd). gCBEv2 was found to primarily induce C-to-G and C-to-T conversions, with the C-to-G / T to C-to-A / G / T conversion ratio reaching as high as 0.97, and very little C-to-A conversion was detected. Figure 22 c. Figure 33 eh). gCBEv2 induced insertions and deletions at the examined sites, with frequencies ranging from 3.1% to 48.3%. Figure 33 i). After analyzing the sequences of all test sites, it was found that the editable range of gCBEv2 is positions 2 to 9 (i). Figure 33 c), and gCBEv2 shows a preference for editing at AC or TC motifs as being more efficient than at other motifs. Figure 33 j).

[0604] When with CGBE1 12 Compared to (C to G base editors), gCBEv2 was found to exhibit higher editing efficiency at certain positions toward the distal end of the target sequence. Figure 22 d and Figure 33 c) indicates its positional preference within different optimal editing windows (gCBEv2 at positions 2 to 6 compared to CGBE1). 12 At positions 5 to 7). Compared to CGBE1, gCBEv2 induces fewer insertions / deletions at position 36 and more insertions / deletions at positions 28 and 29. Figure 33 k). It should be noted that the orthogonal R-loop determination mentioned above should be used. 9 , 29 It was found that gCBEv2 showed a frequency comparable to CGBE1 at two guide sequence-independent off-target sites. Figure 22 ef and Figure 33 l).

[0605] Furthermore, gCBEv2 was found to promote only C editing and virtually no T editing at all examined sites. Figure 33 cd). The editing specificity of gCBEv2 along with the editing specificity of gTBEv3 ( Figure 30 Together, be reinforces the orthogonality of these two base editors for direct base editing.

[0606] Example 9. Application of gTBE and gCBE

[0607] The potential applications of gTBE and gCBE were further evaluated. gTBE can not only repair inactive splicing signals in the intron-splitting EGFP reporter system used above, but also... Figures 19-20 and Figure 26 Furthermore, it can be used for exon skipping by disrupting the splicing signal at the splice donor (SD) or splice acceptor (SA) sites. Figure 23 a). In analyzing research used in gene and cell therapy 30-32 After examining splicing sites in 16 well-studied genes, gTBE and gCBE, along with other existing base editors, provided 1904 sgRNA candidates (prototype spacer sequences / guide sequences shown in Table 3), with SD or SA sites located in each optimal editing window. Figure 23 b and Figure 34 a). Of the 771 sgRNA candidates targeting ABE and CBE, 156 and 103 candidates overlapped with candidates targeting gGBE and gTBE, respectively. Figure 23 c). Furthermore, 232 and 223 sgRNA candidates could only be screened by targeting with gGBE or gTBE, respectively. Figure 23 c). For gCBE, in addition to the 205 sgRNA candidates that overlap with those for CBE, there are 148 unique sgRNA candidates ( Figure 34 b). The availability of these base editors can greatly expand the range of guided RNA screenings for effective editing at splice sites. Figure 34 Furthermore, these newly developed base editors can be used to bypass early termination codons (PTCs) and introduce PTCs. Figure 35 gTBE and gCBE can provide more diverse codon endings from PTC editing. Figure 35 b), and by editing more codons encoding various amino acids to introduce PTC ( Figure 35d). To investigate the potential disruption of gene function through the introduction of PTCs, 851 sgRNA candidates targeting various codons for PTC introduction in 15 genes were analyzed using gGBE and CBE (prototype spacer sequences / guide sequences shown in Table 4), of which 191 TACs and 124 TCAs were used for gGBE targeting ( Figure 35 e).

[0608] To illustrate these applications, splicing sites in the human DMD gene (Duchenne muscular dystrophy, encoding dystrophin) that cannot be targeted by ABE or CBE were developed. A series of sgRNAs specifically targeting SD or SA sites were designed and screened using gTBEv3 or gCBEv2. Figure 23 d and Figure 34 c) This includes targeting DMD exon 45, which is uniquely targeted by gTBEv3. Figure 23 e), 12 and 37 ( Figure 34 d) Three sgRNAs at the SD site. Disruption of the SD site at exon 45, leading to exon skipping, will be applicable to restoring dystrophin expression in 9% of DMD patients. 33 Therefore, the mRNA encoding gTBEv3 and the sgRNA targeting the SD site of exon 45 of DMD were co-injected into humanized mouse zygotes to explore the potential applications of gTBE. It was found that 100% (20 / 20) of mouse embryos exhibited efficient base conversion at the desired location T3 (range 35.0% to 97.0%). Figure 23 The results (fg) demonstrate the significant potential of gTBE for human disease modeling and gene therapy. Overall, gBEs (including gTBE, gCBE, and gGBE) offer more options for targeting sites that deaminase-based base editors cannot, greatly expanding the targeting scope of base editors.

[0609] Example 10. Comparison of different base editing systems

[0610] gTBEv4 (SEQ ID NO:161) and gTBEv5 (SEQ ID NO:162) were generated by inserting the UNG2 mutant (SEQ ID NO:152) contained in gTBEv3 (SEQ ID NO:153) into different positions in the fission-type nCas9 domain. Figure 24b). For gTBEv4 (SEQ ID NO:161), the UNG2 mutant (SEQ ID NO:152) is embedded between positions 2-1248 of nCas9 (SEQ ID NO:2) and positions 1249-1368 of nCas9 (SEQ ID NO:2). For gTBEv5 (SEQ ID NO:162), the UNG2 mutant (SEQ ID NO:152) is embedded between positions 2-1047 of nCas9 (SEQ ID NO:2) and positions 1064-1368 of nCas9 (SEQ ID NO:2). Note that the first amino acid residue D of nCas9 (SEQ ID NO:2) is numbered at position 2 instead of position 1.

[0611] Similarly, gCBEv3 (SEQ ID NO:164) was generated by replacing the UNG mutant (SEQ ID NO:152) in gTBEv5 (SEQ ID NO:162) with the UNG mutant (SEQ ID NO:159) in gCBEv2 (SEQ ID NO:160). Figure 38 a).

[0612] To better characterize the performance of various deaminase-free base editors, the base editor in this study was compared side-by-side with base editors from the following two other studies: He et al. developed TSBE3 for T-to-G / C substitution using a protein language model (PLM)-assisted strategy. 34 Ye et al. used error-prone PCR to perform multiple rounds of random mutagenesis in Escherichia coli for directed evolution and obtained several deaminase-free base editors (DAF-TBE and DAF-CBE). 35 ( Figure 24 a). The basic architectures of the base editors mentioned above are different; for example, TSBE3 is built using an embedding strategy, while DAF-TBE2 is built using a circular rearrangement strategy. Figure 24 b).

[0613] The T-editing efficiency of various thymine base editors was compared at 17 endogenous sites, including those from He studies. 34 Five sites and studies from Ye 35 The five sites ( Figure 24 c and Figure 36 For base editors with UNG mutants fused to the N-terminus of nCas9, gTBEv3 showed higher editing efficiency than DAF-TBE at the vast majority of T sites (29 out of 35) at the test sites. Figure 24 c. Figure 36f) indicates that the UNG mutants generated through rational mutagenesis are superior to those generated through random mutagenesis.

[0614] gTBEv3 is also compared to two base editors, gTBEv4 and gTBEv5, built using an embedding strategy. gTBEv4 shows the editing window moving from position 3-7 to position 7-13. Figure 24 d) Its average editing efficiency was not significantly different from gTBEv3 (23.2% compared to 23.1%). Figure 36 f). For gTBEv5, editing efficiency was 23.1% compared to gTBEv3 (average 39.3%). Figure 36 f) and gTBEv4 show significantly improved editing efficiency compared to other base editors, with the same major T-to-S conversion (f) Figure 36 ad and g), and the optimal editing window covers prototype interval subpositions 5 to 9 (ad ...). Figure 24 d).

[0615] TSBE3 (carrying L83Q and G116E mutations, equivalent to L74Q and G107E in UNG1) is an nCas9-intercalated base editor with nearly identical insertion sites to gTBEv5. Figure 24 c). gTBEv5 showed higher editing efficiency than TSBE3 at the vast majority of T sites (29 out of 35) (39.3% vs. 22.5%). Figure 36 f)( Figure 24 c) indicates that the UNG mutants generated through rational mutagenesis are superior to those generated through PLM-assisted mutagenesis. The optimal editing window of TSBE3 covers prototype spacer positions 4 to 9. Figure 24 d).

[0616] The circularly rearranged DAF-TBE2 shows low average editing efficiency and editing windows at positions 9-13, unlike the editing windows of DAF-TBE (positions 2-6). Figure 24 d).

[0617] Despite showing the highest average editing efficiency, gTBEv5 induced insertion / deletion rates compared to DAF-TBE (14.4% vs. 14.4%), DAF-TBE2 (14.4% vs. 10.3%), and TSBE3 (14.4% vs. 13.5%). Figure 36 The insertion and deletion rates (e.g.) are comparable. It should be noted that gTBE induces significantly less unintended T editing in the proximal DNA sequence upstream of the two sites with unintended editing (sites 38 and 44) ​​than TSBE3 and DAF-TBE (Figure B13), which is related to the fact that UNG's NTD can promote enzyme targeting to the ssDNA-dsDNA junction.19 The findings are consistent.

[0618] Similarly, the C-editing efficiency of various base editors was compared at 19 endogenous sites. Figure 38 a) These endogenous sites include those from He studies 34 Five sites and studies from Ye 35 The five sites ( Figure 38 bd). Both gCBEv2 and gCBEv3 were found to exhibit higher overall average editing efficiency than all other base editors. Figure 38 (ef), especially gCBEv3. In terms of average base conversion efficiency, gCBEv2 outperforms DAF-CBE (30.1% vs. 21.3%) and CGBE-CDG (30.1% vs. 19.3%). Figure 38 The results show that the UNG mutants induced by rational mutagenesis are superior to those induced by random mutagenesis. The average insertion / deletion rate induced by gCBEv2 is comparable to that of other deaminase-free base editors, including DAF-CBE (16.8% vs. 16.9%), DAF-CBE2 (16.8% vs. 12.1%), and CGBE-CDG (16.8% vs. 13.6%). Figure 38 dg). The C to G editing frequencies and purities of different base editors demonstrate the respective advantages of CGBE1 and various deaminase-free base editors at different pyrimidine positions in the prototype spacer. Figure 39 ab). Each base editor allows editing of its target bases within a single editable window, specifically positions 2 to 9 for gCBEv2, positions 2 to 11 for gCBEv3, positions 4 to 10 for CGBE1, positions 2 to 9 for CGBE-CDG, positions 2 to 9 for DAF-CBE, and positions 9 to 12 for DAF-CBE2. Figure 39 c).

[0619] By analyzing the off-target effects at some guide sequence-dependent and guide sequence-independent off-target sites, it was found that gTBE and gCBE induce low-level off-target editing comparable to other base editors at most sites. Figure 40 Furthermore, whole-transcriptome RNA analysis revealed that gTBEv5 and gCBEv3 did not exhibit significant off-target RNA editing or affect the cell's intrinsic DNA repair process. Figure 40 d), which is related to DAF-TBE, DAF-CBE, CGBE-CDG and TSBE3 34、35 Those consistent with each other.

[0620] The lead editing (PE) system can theoretically mediate all types of base substitutions, including T-to-G and C-to-G conversions. 39 Comparing gTBEv3 and gTBEv5 with the recently evolved PE6d system 4 0 in six previously reported endogenous sites in HEK293T cells 35 The comparison was performed at four test sites. gTBEv3 and gTBEv5 outperformed PE6d or PE6d max in T-to-G conversion. Figure 41 a). At all five test sites, gCBEv2 and gCBEv3 outperformed PE6d or PE6d max in C-to-G conversion. Figure 41 b). These findings suggest that base editing and leader editing offer complementary advantages, and that base editors generally demonstrate more effective editing if the target base is optimally positioned. Furthermore, gTBE and gCBE exhibited efficient T and C editing in three different human cell lines (HEK293T, U2OS, and HuH-7 cells), with slight variations in product purity for gTBE, and comparable substitution frequencies of certain bases in gCBE across different cell lines. Figure 42 ).

[0621] In summary, gTBE and gCBE were found to outperform other base editors in this study, including DAF-TBE, DAF-CBE, TSBE3, and CGBE-CDG from two other studies. Furthermore, the selective editing windows of different base editors will provide more options for appropriate base conversions.

[0622] discuss

[0623] Deaminase-based base editors (dBEs) and their derivatives enable the direct editing of adenine (A) and cytosine (C), but not thymine (T). In humans, approximately 19% of pathogenic single nucleotide polymorphisms (SNPs) can be corrected via T-to-G conversion. 9 In this study, two orthogonal base editors, gTBE and gCBE, were developed that enable efficient T and C editing in both cultured human cells and mouse embryos. gTBE and gCBE significantly broaden the targeting scope of base editors by overcoming the limitations of PAM and narrow editing windows, thereby increasing the opportunity to obtain efficient strategies for further research. The T-to-S conversion capability of gTBE allows for a variety of gene editing applications, including editing splice sites and bypassing PTCs.

[0624] It has been shown that the same primitive DNA glycosylation enzyme can be engineered into an enzyme that selectively removes different specific nucleotide bases and can be used to develop novel base editors using the deaminase-free glycosylation-based strategy of this disclosure. The enhanced editing efficiency can be attributed to mutations in the UNG moiety that promote its specific substrate preference or ssDNA binding activity, or both. The high editing efficiency of gTBEv5 suggests that inserting UNG mutants into split-type nCas9 can enhance target DNA accessibility by modulating the interaction between the UNG moiety and the target DNA.

[0625] This study systematically compared glycosylation enzyme-mediated base editors developed in different studies. Using structure-informed rational design, gTBE and gCBE were successfully engineered for efficient T and C editing, respectively. He et al. used PLM to assist in the engineering of TSBE3, while Ye et al. obtained DAF-TBE and DAF-CBE through random mutagenesis. Figure 24 a). It was found that gTBE / gCBE in this disclosure outperforms DAF-TBE, DAF-CBE, TSBE3, and CGBE-CDG, exhibiting higher average editing efficiency and a wider selective editing window. Figure 24 cd and Figures 38-39 ).

[0626] Wild-type UNG proteins exhibit high specificity for uracil in both ssDNA and dsDNA, showing a preference for ssDNA. 43 NTDs containing UNG motifs and sites for unintended protein-protein interactions and post-translational modifications can facilitate enzyme targeting to the ssDNA-dsDNA junction. 19、20 TSBE3 and DAF-TBE with full-length UNG2 induce more unintended editing than gTBE in the proximal DNA sequence upstream of the two sites with unintended editing. Figure 37 ).

[0627] In summary, two orthogonal base editors based on the same primitive DNA glycosylation enzyme have been engineered for direct T-editing and direct C-editing, and structure-informed rational design represents an efficient and effective protein engineering strategy, providing reference and solutions for the subsequent evolution of other proteins.

[0628] method

[0629] Molecular cloning

[0630] The base editor construct used in this study was cloned into the mammalian expression plasmid backbone using standard molecular cloning techniques. It was under the control of the EF1α promoter and, like the previously described ones... 9Similarly, two intron-split EGFP reporters were constructed, but an engineered sequence containing the last 86 base pairs (bp) of human RPS5 introns was inserted between the BFP and EGFP coding sequences. Corresponding mutations were then performed at the splice acceptor site to construct T-G or C-G reporters via site-directed mutagenesis using PCR, respectively. Mutations at the splice acceptor site resulted in inactive EGFP. (The last sentence appears to be incomplete and possibly refers to a previous base editor.) 12、15 Encouraged by the discovery, the corresponding mutation at the splice acceptor site was placed at position 6 of the prototype spacer.

[0631] The wild-type UNG2 sequence (313 amino acids long) (SEQ ID NO: 133) was amplified by PCR from the cDNA of truncated mutants of HEK293T, UNG2-Y156A, UNG2-N213D, and UNG-NTD, and the corresponding combinations were constructed by site-directed mutagenesis via PCR. UNG mutants were fused with nCas9 in different orientations using the Gibson assembly method. The PE6d construct contains an evolved and engineered M-MLV variant of the human codon-optimized RNase H truncated with the R221K / N394K / H840A mutation in SpCas9. A cleavage sgRNA and an epigRNA with the tevoPreQ1 motif were cloned into the PE6d construct using Golden Gate assembly to generate an integrated plasmid. For PE6d max, codon-optimized hMLH1dn was co-expressed with PE6d.

[0632] As previously stated 52 A UNG mutagenesis library was designed and generated, with some modifications. In short, the region of UNG2 containing 98-313 amino acids was divided into 8-amino acid segments. Mutants with BpiI containing Y156A or N213D were introduced via site-directed mutagenesis using PCR. For the evolution of gTBE, regions I150-L179, A158-K261, L210-T217, and Q274-Y284 were selected for multiple rounds of sequential alanine / arginine / aspartic acid / valine substitution (X to A, R, D, or V). Site-saturation mutagenesis was performed on residue 214 to check for any amino acid at that position that outperformed valine. For the evolution of gCBE, regions D154-D189 were selected for sequential alanine substitution (X to A). To cover all residues in the corresponding segments used for sequential alanine substitution, alanine was mutated to valine (A to V). The oligonucleotides encoding the mutant were annealed and ligated into the corresponding BpiI-digested backbone vector.

[0633] Using Cas-OFFinder 28Search for potential guide sequence-dependent off-target sites of Cas9 RNA-guided endonucleases, with a maximum of 3 mismatches (no protrusions). For sgRNAs with NGN PAM targeting DMD splicing sites, use the PAM flexible Cas9 variant SpG (SEQ ID NO:163) instead of nCas9 (SEQ ID NO:2). Anneal the sgRNA oligonucleotides and ligate them to the BpiI site.

[0634] Cell culture, transfection and flow cytometry analysis

[0635] HEK293T, HuH-7, and U2OS cells were cultured in DMEM (catalog number 11995065, Gibco) supplemented with 10% fetal bovine serum (catalog number 04-001-1ACS, BI) and 0.1 mM non-essential amino acids (catalog number 11140-050, Gibco) in an incubator at 37°C with 5% CO2.

[0636] Mutant screening was performed in 48-well plates, with 3 × 10⁻⁶ cells per well the day before transfection. 4 HEK293T cells were seeded in 250 μL of complete growth medium. Between 16 and 24 h post-seeding, cells were co-transfected with 250 ng gTBE (or gCBE) plasmid, 250 ng T-G (or C-G) reporter plasmid, and 1 μg polyethyleneimine (PEI) per well (DNA / PEI ratio 1:2). For HEK293T, HuH-7, and U2OS cells used in FACS, 5 × 10⁶ cells per well were seeded one day prior to transfection. 5 Cells were seeded in 12-well plates containing 1 ml of complete growth medium. After 14–16 h, 2 μg of a monoclonal plasmid expressing gTBE or gCBE and its corresponding sgRNA was transfected into the cells using PEI (1:2 DNA / PEI ratio) or FuGENE HD transfection reagent (1:3 DNA:FuGENE ratio; E2311, Promega). As previously described. 9 , 29Orthogonal R-loop assays were performed. In short, HEK293T cells were co-transfected with 1 μg of a gTBE or gCBE expression plasmid (with mCherry as a reporter) containing sgRNA targeting the corresponding site and 1 μg of a dSaCas9 expression plasmid (with EGFP as a reporter) containing corresponding sgRNA targeting five off-target sites to generate R-loops using PEI (1:2 DNA / PEI ratio). For leader editing, cells were co-transfected with 2 μg of a monoclonal plasmid containing PE6d, nick sgRNA, and epigRNA, or 1 μg of a monoclonal plasmid and 1 μg of hMLH1dn plasmid using PEI (1:2 DNA / PEI ratio).

[0637] Forty-eight hours post-transfection, the expression of mCherry, BFP, and EGFP fluorescence was analyzed using BD FACS Aria III or Beckman CytoFLEX S. Flow cytometry results were analyzed using FlowJo V10.5.3. mCherry was used for mid-target editing efficiency evaluation. + BFP + and EGFP + Gating strategies in cell identification provide Figure 26 b in.

[0638] Target sequencing and data analysis of endogenous sites

[0639] As previously stated 9 The endogenous target sites of interest were amplified from genomic DNA. In short, 72 hours after transfection, 10,000 mCherry-positive cells were isolated using FACS, and genomic DNA was extracted. Site-specific primers were then used to amplify regions of interest at the target sites via PCR. The purified PCR products were analyzed using Sanger sequencing (Kingwiz).

[0640] The target sequencing data analysis is described in previous papers. 3 In short, the amplicon is ligated to the adapter, and sequencing is performed on the Illumina MiSeq platform, followed by fastp with default parameters. 53 Process targeted amplicon sequencing reads and via CRISPResso2 54 Further amplicon sequencing analysis was performed. T-G purity was calculated as T-G editing efficiency / (T-C editing efficiency + T-G editing efficiency + T-A editing efficiency). T-S conversion ratio was calculated as (T-C editing efficiency + T-G editing efficiency) / (T-C editing efficiency + T-G editing efficiency + T-A editing efficiency). The prototype spacer sequence guide is shown in Table 2.

[0641] In vitro transcription of gTBEv3 mRNA and DMD sgRNA

[0642] As previously stated 9 mRNA and sgRNA preparation were performed. The gTBEv3 expression plasmid was linearized using FastDigest KpnI restriction enzyme (catalog number FD0524, Thermo Fisher Scientific), purified using a gel extraction kit (catalog number D2500-03, Omega), and used as a template for in vitro transcription (IVT) using the mMESSAGE mMACHINE T7 Ultra kit (catalog number AM1345, Thermo Ambion). For DMD-sgRNA preparation, the T7 promoter sequence was added to the sgRNA template by PCR amplification. The T7-DMD-sgRNA PCR product was purified using a gel extraction kit (catalog number D2500-03, Omega), and used as a template for sgRNA IVT using the MEGAshortscript T7 kit (catalog number AM1354, Invitrogen). The mRNA and DMD-sgRNA encoding gTBEv3 were purified using the MEGAclear kit (catalog number AM1908, Ingenium Technologies), eluted in RNase-free water, and stored at -80°C until use.

[0643] Microinjection of fertilized eggs in animals and mice

[0644] Animal manipulation and those previously reported 3 Consistent. The experiments involving mice were approved by the Biomedical Research Ethics Center Committee of Huida (Shanghai) Biotechnology Co., Ltd. Mice were kept in a specific pathogen-free facility under 12-hour dark-light cycles and constant temperature (20℃-26℃) and humidity (40%-60%) conditions.

[0645] Humanized DMD females (4 weeks old) with superovulation and exon 45 of human DMD in a C57BL / 6 background were mated with C57BL / 6 males (8 weeks old), and females from the ICR strain were used as surrogate mothers. Fertilized embryos were collected from the fallopian tubes 21 hours after hCG injection. For fertilized egg injection, mRNA encoding gTBEv3 (250 ng / μL) was injected with hCG at a constant flow rate using a FemtoJet microinjector (Eppendorf).

[0646] A mixture of DMD-sgRNA (100 ng / μL) was injected into the cytoplasm of single-cell embryos in M2 medium droplets. The injected embryos were cultured in M16 medium containing amino acids until blastocysts formed for three days (37°C and 5% CO2), after which genomic DNA was extracted and target amplification was performed.

[0647] RNA sequencing experiment

[0648] HEK293T cells were seeded in 12-well plates as described above and transfected with 2 μg of gTBEv5, gCBEv3, CGBE1, or mCherry plasmid using PEI (1:2 DNA / PEI ratio). Approximately 5 × 10⁶ cells were collected 48 hours post-transfection. 6 Total RNA was extracted from cells using a TRIzol-based method, fragmented, and reverse transcribed into cDNA using HiScript Q RTSuperMix according to the manufacturer's instructions. Total RNA integrity was quantified using an Agilent 2100 bioanalyzer. RNA-seq library qualification was confirmed using the Illumina NovaSeq 6000 platform (by Genewiz). Trimmomatic (v.0.39-2) was used. 55 Filter the raw RNA-seq data using Hisat2 (v.2.2.1). 56 The clean reads were aligned with the hg38 reference genome. REDItools2 with default parameters was used. 57 Calculate RNA editing sites. Use the dbSNP (v.146) database downloaded from NCBI to filter sites that overlap with common single nucleotide variants (SNVs). Further filter sites with fewer than five mutated or non-mutated reads.

[0649] Using StringTie 58 Calculate expression values. Use DESeq2. 59 Calculate differentially expressed genes with FDR < 0.05 and fold change > 1.

[0650] Statistical analysis

[0651] Statistical tests performed using Graphpad Prism 8 include a two-tailed unpaired two-sample t-test following one-way ANOVA or Dunnett's multiple comparison test.

[0652] References

[0653] 1.Porto,E.M.,Komor,A.C.,Slaymaker,I.M.&Yeo,G.W.Base editing:advancesand therapeutic opportunities.Nat Rev Drug Discov 19,839-859(2020).

[0654] 2.Rees,H.A.&Liu,D.R.Base editing:precision chemistry on the genomeand transcriptome of living cells.Nat Rev Genet 19,770-788(2018).

[0655] 3.Tong,H.et al.Programmabledeaminase-free base editors for G-to-Yconversion by engineered glycosylase.Natl Sci Rev 10,nwad143(2023).

[0656] 4.Gaudelli,N.M.et al.Programmable base editing of A*T to G*C ingenomic DNA without DNA cleavage.Nature 551,464-471(2017).

[0657] 5.Komor,A.C.,Kim,Y.B.,Packer,M.S.,Zuris,J.A.&Liu,D.R.Programmableediting of a target base in genomic DNA without double-stranded DNAcleavage.Nature533,420-424(2016).

[0658] 6.Mok,B.Y.et al.A bacterial cytidine deaminase toxin enables CRISPR-free mitochondrial base editing.Nature 583,631-637(2020).

[0659] 7.Lei,Z.et al.Mitochondrial base editor induces substantial nuclearoff-target mutations.Nature 606,804-811(2022).

[0660] 8.Zhang,X.et al.Dual base editor catalyzes both cytosine and adeninebase conversions in human cells.Nat Biotechnol 38,856-860(2020).

[0661] 9.Tong,H.et al.Programmable A-to-Y base editing by fusing an adeninebase editor with an N-methylpurine DNA glycosylase.Nat Biotechnol 41,1080-1084(2023).

[0662] 10.Chen,L.et al.Adenine transversion editors enable precise,efficientA*T-to-C*G base editing in mammalian cells and embryos.Nat Biotechnol(2023).

[0663] 11.Zhao,D.et al.Glycosylase base editors enable C-to-A and C-to-Gbase changes.Nat Biotechnol 39,35-40(2021).

[0664] 12.Kurt,I.C.et al.CRISPR C-to-G base editors for inducing targetedDNA transversions in human cells.Nat Biotechnol 39,41-46(2021).

[0665] 13.Koblan,L.W.et al.Efficient C*G-to-G*C base editors developed usingCRISPRi screens,target-library analysis,and machine learning.Nat Biotechnol39,1414-1425(2021).

[0666] 14.Chen,L.et al.Programmable C:G to G:C genome editing with CRISPR-Cas9-directed base excision repair proteins.Nat Commun 12,1384(2021).

[0667] 15.Yuan,T.et al.Optimization of C-to-G base editors with sequencecontext preference predictable by machine learning methods.Nat Commun 12,4902(2021).

[0668] 16.Nilsen,H.et al.Nuclear and mitochondrial uracil-DNA glycosylasesare generated by alternative splicing and transcription from differentpositions in the UNG gene.Nucleic Acids Res 25,750-755(1997).

[0669] 17.Kavli,B.et al.Excision of cytosine and thymine from DNA by mutantsof human uracil-DNA glycosylase.EMBO J 15,3442-3447(1996).

[0670] 18.Rodriguez,G.et al.Disordered N-Terminal Domain of Human Uracil DNAGlycosylase(hUNG2)Enhances DNA Translocation.ACS Chem Biol 12,2260-2263(2017).

[0671] 19.Weiser,B.P.,Rodriguez,G.,Cole,P.A.&Stivers,J.T.N-terminal domainof human uracil DNA glycosylase(hUNG2)promotes targeting to uracil sitesadjacent to ssDNA-dsDNA junctions.Nucleic Acids Res 46,7169-7178(2018).

[0672] 20.Perkins,J.L.&Zhao,L.The N-terminal domain of uracil-DNAglycosylase:Roles for disordered regions.DNA Repair(Amst)101,103077(2021).

[0673] 21.Nagelhus,T.A.et al.A sequence in the N-terminal region of humanuracil-DNA glycosylase with homology to XPA interacts with the C-terminalpart of the 34-kDa subunit of replication protein A.J Biol Chem 272,6561-6566(1997).

[0674] 22.Torseth,K.et al.The UNG2 Arg88Cys variant abrogates RPA-mediatedrecruitment of UNG2 to single-stranded DNA.DNA Repair(Amst)11,559-569(2012).

[0675] 23.Schormann,N.,Ricciardi,R.&Chattopadhyay,D.Uracil-DNA glycosylases-structural and functional perspectives on an essential family of DNA repairenzymes.Protein Sci 23,1667-1685(2014).

[0676] 24.Parikh,S.S.et al.Uracil-DNA glycosylase-DNA substrate and productstructures:conformational strain promotes catalytic efficiency by coupledstereoelectronic effects.Proc Natl Acad Sci U S A 97,5083-5088(2000).

[0677] 25.Parikh,S.S.et al.Base excision repair initiation revealed bycrystal structures and binding kinetics of human uracil-DNA glycosylase withDNA.EMBO J 17,5214-5226(1998).

[0678] 26.Chen,L.et al.Re-engineering the adenine deaminase TadA-8e forefficient and specific CRISPR-based cytosine base editing.Nat Biotechnol 41,663-672(2023).

[0679] 27.Jeong,Y.K.et al.Adenine base editor engineering reduces editing ofbystander cytosines.Nat Biotechnol 39,1426-1433(2021).

[0680] 28.Bae,S.,Park,J.&Kim,J.S.Cas-OFFinder:a fast and versatile algorithmthat searches for potential off-target sites of Cas9 RNA-guidedendonucleases.Bioinformatics30,1473-1475(2014).

[0681] 29.Richter,M.F.et al.Phage-assisted evolution of an adenine baseeditor with improved Cas domain compatibility and activity.Nat Biotechnol 38,883-891(2020).

[0682] 30.Uddin,F.,Rudin,C.M.&Sen,T.CRISPR Gene Therapy:Applications,Limitations,and Implications for the Future.Front Oncol 10,1387(2020).

[0683] 31.Nordestgaard,B.G.,Nicholls,S.J.,Langsted,A.,Ray,K.K.&Tybjaerg-Hansen,A.Advances in lipid-lowering therapy through gene-silencingtechnologies.Nat Rev Cardiol 15,261-272(2018).

[0684] 32.Zhang,X.et al.Gene knockout in cellular immunotherapy:Applicationand limitations.Cancer Lett 540,215736(2022).

[0685] 33.Bladen,C.L.et al.The TREAT-NMD DMD Global Database:analysis ofmore than 7,000 Duchenne muscular dystrophy mutations.Hum Mutat 36,395-402(2015).

[0686] 34.He,Y.et al.Protein language models-assisted optimization of auracil-N-glycosylase variant enables programmable T-to-G and T-to-C baseediting.Mol Cell,Online ahead of print(2024).

[0687] 35.Ye,L.et al.Glycosylase-based base editors for efficient T-to-G andC-to-G editing in mammalian cells.Nat Biotechnol,Online ahead of print(2024).

[0688] 36.Li,S.et al.Docking sites inside Cas9 for adenine base editingdiversification and RNA off-target elimination.Nat Commun 11,5827(2020).

[0689] 37.Liu,Y.et al.A Cas-embedding strategy for minimizing off-targeteffects of DNA base editors.Nat Commun 11,6073(2020).

[0690] 38.Nguyen Tran,M.T.et al.Engineering domain-inlaid SaCas9 adeninebase editors with reduced RNA off-targets and increased on-target DNAediting.Nat Commun 11,4871(2020).

[0691] 39.Anzalone,A.V.et al.Search-and-replace genome editing withoutdouble-strand breaks or donor DNA.Nature 576,149-157(2019).

[0692] 40.Doman,J.L.et al.Phage-assisted evolution and protein engineeringyield compact,efficient prime editors.Cell 186,3983-4002 e3926(2023).

[0693] 41.Zuo,E.et al.Cytosine base editor generates substantial off-targetsingle-nucleotide variants in mouse embryos.Science 364,289-292(2019).

[0694] 42.Yan,N.et al.Cytosine base editors induce off-target mutations andadverse phenotypic effects in transgenic mice.Nat Commun 14,1784(2023).

[0695] 43.Slupphaug,G.et al.Properties of a recombinant human uracil-DNAglycosylase from the UNG gene and evidence that UNG encodes the major uracil-DNA glycosylase.Biochemistry 34,128-138(1995).

[0696] 44.Chen,L.et al.Engineering a precise adenine base editor withminimal bystander editing.Nat Chem Biol 19,101-110(2023).

[0697] 45.Kim,Y.B.et al.Increasing the genome-targeting scope and precisionof base editing with engineered Cas9-cytidine deaminase fusions.NatBiotechnol 35,371-376(2017).

[0698] 46.Huang,M.E.et al.C-to-G editing generates double-strand breakscausing deletion,transversion and translocation.Nat Cell Biol 26,294-304(2024).

[0699] 47.Hindi,N.N.,Elsakrmy,N.&Ramotar,D.The base excision repair process:comparison between higher and lower eukaryotes.Cell Mol Life Sci 78,7943-7965(2021).

[0700] 48.Thompson,P.S.&Cortez,D.New insights into abasic site repair andtolerance.DNA Repair(Amst)90,102866(2020).

[0701] 49.Wang,Y.et al.Engineering of the Translesion DNA Synthesis PathwayEnables Controllable C-to-G and C-to-A Base Editing in Corynebacteriumglutamicum.ACS Synth Biol 11,3368-3378(2022).

[0702] 50.Sun,N.et al.Reconstructed glycosylase base editors GBE2.0 withenhanced C-to-G base editing efficiency and purity.Mol Ther 30,2452-2463(2022).

[0703] 51.Komor,A.C.et al.Improved base excision repair inhibition andbacteriophage Mu Gam protein yields C:G-to-T:A base editors with higherefficiency and product purity.SCI ADV 3,eaao4774(2017).

[0704] 52.Tong,H.et al.High-fidelity Cas13 variants for targeted RNAdegradation with minimal collateral effects.Nat Biotechnol 41,108-119(2023).

[0705] 53.Chen,S.,Zhou,Y.,Chen,Y.&Gu,J.fastp:an ultra-fast all-in-one FASTQpreprocessor.Bioinformatics 34,i884-i890(2018).

[0706] 54.Clement,K.et al.CRISPResso2 provides accurate and rapid genomeediting sequence analysis.Nat Biotechnol 37,224-226(2019).

[0707] 55.Bolger,A.M.,Lohse,M.&Usadel,B.Trimmomatic:a flexible trimmer forIllumina sequence data.Bioinformatics 30,2114-2120(2014).

[0708] 56.Kim,D.,Paggi,J.M.,Park,C.,Bennett,C.&Salzberg,S.L.Graph-basedgenome alignment and genotyping with HISAT2 and HISAT-genotype.Nat Biotechnol37,907-915(2019).

[0709] 57.Flati,T.et al.HPC-REDItools:a novel HPC-aware tool for improvedlarge scale RNA-editing analysis.BMC Bioinformatics 21,353(2020).

[0710] 58.Pertea,M.et al.StringTie enables improved reconstruction of atranscriptome from RNA-seq reads.Nat Biotechnol 33,290-295(2015).

[0711] 59. Anders, S. & Huber, W. Differential expression analysis for sequence count data. Genome Biol 11, R106 (2010).

[0712] 60.Krusong,K.,Carpenter,EP,Bellamy,SR,Savva,R.&Baldwin,GSAcomparative study of uracil-DNA glycosylases from human and herpes simplexvirus type 1.J Biol Chem 281,4983-4992(2006).

[0713] Exemplary sequence

[0714] SEQ ID NO:1, wild-type human MPG without a first N-terminal M, 297aa

[0715] VTPALQMKKPKQFCRRMGQKKQRPARAGQPHSSSDAAQAPAEQPHSSSDAAQAPCPRERCLGPPTTPGPYRSIYFSSPKGHLTRLGLEFFDQPAVPLARAFLGQVLVRRLPNGTELRGRIVETEAYLGPEDEAAHSRGGRQTPRNRGM FMKPGTLYVYIIYGMYFCMNISSQGDGACVLLRALEPLEGLETMRQLRSTLRKGTASRVLKDRELCSGPSKLCQALAINKSFDQRDLAQDEAVWLERGPLEPSEPAVVAAARVGVGHAGEWARKPLRFYVRGSPWVSVVDRVAEQDTQA

[0716] SEQ ID NO:2, SpCas9-D10A, Cas9 nickase, nCas9, without a first N-terminal Met, 1367aa

[0717]

[0718] SEQ ID NO:3, dMPG (inactivated MPG, inactivated MPG-E125A+Y127A+H136A mutant)

[0719]

[0720] SEQ ID NO:4,MPGv0.2(MPG-N169S)

[0721]

[0722] SEQ ID NO:5,MPGv1(MPG-N169S+S198A+K202A+G203A+S206A+K210A)

[0723]

[0724] SEQ ID NO:6,MPGv2(MPG-G163R+N169S)

[0725]

[0726] SEQ ID NO:7, MPGv3

[0727] (MPG-G163R+N169S+S198A+K202A+G203A+S206A+K210A)

[0728]

[0729] SEQ ID NO:8, MPGv6.3

[0730] (MPG-G163R+ N169G + D175R+C178N +S198A+K202A+G203A+S206A+K210A+ Q294R )

[0731]

[0732] SEQ ID NO:9, wild-type human MPG with a first N-terminal M, 298aa

[0733] MVTPALQMKKPKQFCRRMGQKKQRPARAGQPHSSSDAAQAPAEQPHSSSDAAQAPCPRERCLGPPTTPGPYRSIYFSSPKGHLTRLGLEFFDQPAVPLARAFLGQVLVRRLPNGTELRGRIVETEAYLGPEDEAAHSRGGRQTPRNRGMFMKPGTLYVYIIYGMYFCMNISSQGDGACVLLRALEPLEGLETMRQLRSTLRKGTASRVLKDRELCSGPSKLCQALAINKSFDQRDLAQDEAVWLERGPLEPSEPAVVAAARVGVGHAGEWARKPLRFYVRGSPWVSVVDRVAEQDTQA

[0734] SEQ ID NO:10, bpSV40 NLS1

[0735] KRTADGSEFESPKKKRKV

[0736] SEQ ID NO:11, bpSV40 NLS2

[0737] KRTADGSEFEPKKKRKV

[0738] SEQ ID NO:12, gGBEv6.3

[0739] MKRTADGSEFESPKKKRKVSGGSDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKN LIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNI VDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEE NPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDD LDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEI FFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQ EDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNL PNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSV EISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRR RYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGS PAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEGIGELGSQILKEHPVENTQL QNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKNY WRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITL KSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATA KYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESI LPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKG YKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELA...

Claims

1. A fusion protein comprising: (1) A nucleic acid programmable DNA-binding domain (napDNAbd) capable of binding to target dsDNA, wherein the target dsDNA comprises: (a) The first deoxyribonucleotide (e.g., dG (deoxyguanosine), dT (thymidine), dC (deoxycytidine)) in the prototype spacer sequence on the non-target strand (editing strand) of the target dsDNA, and (b) a second deoxyribonucleotide (e.g., dC (deoxycytidine), dA (deoxyadenosine), dG (deoxyguanosine)) in a target sequence paired with the first deoxyribonucleotide (e.g., dG, dT, dC) and located on the target strand (non-edited strand) of the target dsDNA, wherein the prototype spacer sequence is completely inversely complementary to the target sequence; and (2) A base-removing domain capable of removing a base (e.g., guanine, thymine, cytosine) of the first deoxyribonucleotide; The fusion protein described herein does not contain a deaminase domain, such as adenine or cytosine deaminase domains, such as TadA and its variants; and The first deoxyribonucleotide is deoxyguanosine (dG), thymidine (dT), or deoxycytidine (dC).

2. The fusion protein according to any of the preceding claims, wherein the conversion of the first deoxyribonucleotide to the fourth deoxyribonucleotide is dG to dA, dG to dT, dG to dC, dT to dA, dT to dC, dT to dG, dC to dA, dC to dT, or dC to dG.

3. The fusion protein of any of the preceding claims, wherein the base-excision domain comprises a glycosylation enzyme.

4. The fusion protein according to any of the preceding claims, wherein the glycosylation enzyme is selected from N-methylpurine DNA glycosylation enzyme (MPG), 8-oxoguanine DNA glycosylation enzyme (OGG1), methyl-CpG binding domain 4 DNA glycosylation enzyme (MBD4), thymine DNA glycosylation enzyme (TDG), uracil DNA glycosylation enzyme (UNG), single-stranded selective monofunctional uracil DNA glycosylation enzyme 1 (SMUG1), mutY DNA glycosylation enzyme (MUTYH), nth-like DNA glycosylation enzyme 1 (NTHL1), nei-like DNA glycosylation enzyme 1 (NEIL1), nei-like DNA glycosylation enzyme 2 (NEIL2), nei-like DNA glycosylation enzyme 3 (NEIL3), and mutants thereof capable of recognizing and excising bases from nucleotides of nucleic acids.

5. The fusion protein of any of the preceding claims, wherein the base excision domain comprises N-methylpurine DNA glycosylation enzyme (MPG).

6. The fusion protein according to any of the preceding claims, wherein the MPG contains amino acid substitutions at positions corresponding to or selected from the following positions: N169, D175, C178 and / or Q294 of the reference MPG, wherein the positions are numbered according to SEQ ID NO:

1.

7. The fusion protein of any of the preceding claims, wherein the amino acid substitution is made by substitution with R, A, N or G.

8. The fusion protein according to any of the preceding claims, wherein the MPG comprises an amino acid substitution relative to the reference MPG of SEQ ID NO:7, the amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following substitutions: N169G, D175R, C178N, Q294R, wherein the position is numbered according to SEQ ID NO:

1.

9. The fusion protein according to any of the preceding claims, wherein the MPG comprises a combination substitution relative to the reference MPG of SEQ ID NO:7, the combination substitution corresponding to a combination substitution of N169G, D175R, C178N and Q294R, wherein the position is numbered according to SEQ ID NO:

1.

10. The fusion protein of any of the preceding claims, wherein the MPG comprises, is substantially composed of, or is composed of: An amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO:8, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, or 38.

11. The fusion protein of any of the preceding claims, wherein the MPG is (substantially) capable of cleaving guanine from dG.

12. The fusion protein of any of the preceding claims, wherein the base excision domain comprises uracil DNA glycosylase (UNG).

13. The fusion protein of any of the preceding claims, wherein the UNG comprises an amino acid substitution at a position corresponding to a position selected from or selected from the following positions: K184, A214, Q259 and / or Y284 of the reference UNG, wherein the position is numbered according to SEQ ID NO:

133.

14. The fusion protein of any of the preceding claims, wherein the amino acid substitution is made by substitution with A, D, V or T.

15. The fusion protein of any of the preceding claims, wherein the UNG comprises an amino acid substitution relative to the reference UNG of SEQ ID NO:135 or 137, the amino acid substitution corresponding to a substitution selected from or a combination of two or more of the following: K184A, A214V, A214T, Q259A, Y284D, and wherein the position is numbered according to SEQ ID NO:

133.

16. The fusion protein according to any of the preceding claims, wherein the UNG comprises the deletion of amino acids at positions corresponding to or corresponding to positions 1-65, 1-66, 1-67, 1-68, 1-69, 1-70, 1-71, 1-72, 1-73, 1-74, 1-75, 1-76, 1-77, 1-78, 1-79, 1-80, 1-81, 1-82, 1-83, 1-84, 1-85, 1-86, 1-87, 1-88, 1-89, 1-90, 1-91, 1-92, 1-93, 1-94, 1-95, 1-96, 1-97, 1-98, 1-99, or 1-100 of the reference UNG of SEQ ID NO:

133.

17. The fusion protein of any of the preceding claims, wherein the UNG comprises an amino acid substitution relative to the reference UNG of SEQ ID NO:135, the amino acid substitution corresponding to a substitution selected from or a combination of two or more substitutions selected from A214T, Q259A, Y284D, and thereof, and the UNG comprises a deletion of an amino acid at a position corresponding to or a position of positions 1-88 of the reference UNG, wherein the positions are numbered according to SEQ ID NO:

133.

18. The fusion protein of any of the preceding claims, wherein the UNG comprises an amino acid substitution relative to the reference UNG of SEQ ID NO:137, the amino acid substitution corresponding to a substitution selected from or a combination of K184A, A214V, and the two substitutions, and the UNG comprises a deletion of an amino acid at a position corresponding to or a position of positions 1-88 of the reference UNG, wherein the positions are numbered according to SEQ ID NO:

133.

19. The fusion protein of any of the preceding claims, wherein the UNG containing the amino acid mutation comprises, substantially comprises, or comprises: An amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with any of SEQ ID NO: 56, 58, 60, 62, 137, 139, 141, 144, 146, 148, 150, 152, 155, 157, and 159, or an N-terminal truncated form lacking the N-terminal methionine (M) (encoded by the start codon ATG).

20. The fusion protein as claimed in any of the preceding claims, wherein the UNG is (substantially) capable of cleaving thymine from dT.

21. The fusion protein as claimed in any of the preceding claims, wherein the UNG is (substantially) capable of cleaving dC cytosine.

22. The fusion protein of any of the preceding claims, wherein the napDNAbd is an RNA-programmable DNA-binding protein.

23. The fusion protein of any of the preceding claims, wherein the napDNAbd is selected from CRISPR-associated (Cas) proteins, IscB, IsrB, Argonaute, and TnpB.

24. The fusion protein of any of the preceding claims, wherein the napDNAbd is a nicking enzyme, such as Cas9 nicking enzyme or IscB nicking enzyme.

25. The fusion protein of any of the preceding claims, wherein the napDNAbd is nuclease-free, for example, inactivated Cas9 or inactivated Cas12i.

26. The fusion protein of any of the preceding claims, wherein the napDNAbd comprises an amino acid sequence having at least about 60% (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity with SEQ ID NO:2, 48, 50, 52, or 163.

27. The fusion protein of any of the preceding claims, wherein the fusion protein comprises (1) the napDNAbp and the base excision domain from the N-terminus to the C-terminus; or (2) the base excision domain and the napDNAbp.

28. The fusion protein of any of the preceding claims, wherein the napDNAbd (e.g., Cas9) is a two-part napDNAbd comprising an N-terminal portion and a C-terminal portion, such as a two-part splitting Cas9, and wherein the fusion protein comprises from the N-terminus to the C-terminus (1) the N-terminal portion of the napDNAbd, the base-removing domain, and the C-terminal portion of the napDNAbd; (2) the C-terminal portion of the napDNAbd, the base-removing domain, and the N-terminal portion of the napDNAbd; or (3) the base-removing domain, the C-terminal portion of the napDNAbd (e.g., amino acids at positions 1249-1368), and the N-terminal portion of the napDNAbd (e.g., amino acids at positions 1-1248).

29. The fusion protein of any of the preceding claims, wherein the napDNAbd is SpCas9 (e.g., SpCas9 nickase) or a mutant thereof (e.g., SpG Cas9 nickase).

30. The fusion protein of any of the preceding claims, wherein the N-terminal portion of the napDNAbd is an amino acid at position 1 or 2 to 1012, 1028, 1041, 1046, 1047, 1248, 1249 or 1300 of the napDNAbp.

31. The fusion protein of any of the preceding claims, wherein the C-terminal portion of the napDNAbd is an amino acid at positions 1013, 1029, 1042, 1047, 1048, 1249, 1063, 1064, 1230, 1249 or 1301 to 1368 of the napDNAbp.

32. The fusion protein according to any of the preceding claims, wherein the fusion protein comprises the base-removing domain, the base-removing domain being embedded between positions 2-1248 and 1249-1368 of nCas9 (SEQ ID NO:2), wherein the first amino acid residue D of nCas9 (SEQ ID NO:2) is designated as position 2; or embedded between positions 2-1047 and 1064-1368 of nCas9 (SEQ ID NO:2), wherein the first amino acid residue D of nCas9 (SEQ ID NO:2) is designated as position 2.

33. The fusion protein according to any of the preceding claims, wherein the fusion protein comprises having at least about 60 The amino acid sequence is of % (e.g., at least about 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100%) sequence identity.

34. A system comprising: (i) the fusion protein of any of the preceding claims or the polynucleotide encoding said fusion protein; and (ii) a guiding nucleic acid or a polynucleotide encoding the guiding nucleic acid, the guiding nucleic acid comprising: (1) A scaffold sequence capable of forming a complex with the napDNAbd; and (2) It is capable of hybridizing with the target sequence on the target strand of the target dsDNA, thereby directing the complex to the guide sequence of the target dsDNA.

35. The fusion protein of any of the preceding claims, wherein the guiding nucleic acid is a guiding RNA (gRNA).

36. The fusion protein of any of the preceding claims, wherein the scaffold sequence has a secondary structure substantially identical to that of the sequence of SEQ ID NO:40, 73 or 74, or wherein the scaffold sequence comprises (1) a 5' or 3' truncated form of SEQ ID NO:40, 73 or 74 or thereof truncated by 1, 2, 3, 4, 5 or 6 nucleotides at the 5' or 3' end; or (2) a sequence having at least about 70%, 75%, 80%, 85%, 90%, 95% or 100% sequence identity with SEQ ID NO:40, 73 or 74 or thereof truncated by 1, 2, 3, 4, 5 or 6 nucleotides at the 5' or 3' end; or (3) a sequence differing from SEQ ID NO:40, 73 or 74 by at most 1, 2, 3, 4, 5 or 6 nucleotides, whether or not the difference is continuous.

37. The fusion protein or system according to any of the preceding claims, wherein the fusion protein or system further comprises a trans-damage synthesis (TLS) polymerase or a recruitment domain or component capable of recruiting a TLS polymerase.

38. The fusion protein or system as claimed in any of the preceding claims, wherein the TLS polymerase is selected from Polα(α), Polβ(β), Polδ(δ)(PCNA), Polγ(γ), Polη(η), Polι(ι), Polκ(κ), Polλ(λ), Polμ(μ), Polν(ν), Polθ(θ) and REV1.

39. A method for modifying target dsDNA, the method comprising contacting the target dsDNA with the system of any of the preceding claims, The target dsDNA contains: (a) The first deoxyribonucleotide (e.g., dG (deoxyguanosine), dT (thymidine), dC (deoxycytidine)) in the prototype spacer sequence on the non-target strand (editing strand) of the target dsDNA, and (b) A second deoxyribonucleotide (e.g., dC (deoxycytidine), dA (deoxyadenosine), dG (deoxyguanosine)) in a target sequence that is paired with the first deoxyribonucleotide (e.g., dG, dT, dC) and located on the target strand (non-edited strand) of the target dsDNA, wherein the prototype spacer sequence is completely inversely complementary to the target sequence; The method described herein does not include deamination of the bases of the first deoxyribonucleotide before removing the bases of the first deoxyribonucleotide.

40. The method of any of the preceding claims, wherein the method does not include deamination of the bases of the first deoxyribonucleotide.

41. An MPG as defined in any of the preceding claims.

42. A UNG as defined in any of the preceding claims.

Citation Information

Patent Citations

  • High-efficiency wild-type-free AAV helper functions

    US6001650A

  • Compositions and methods for gene editing

    WO2018002719A1

  • T:a to a:t base editing through adenine excision

    WO2020181195A1

Cited By

  • Novel uracil DNA glycosylase Ehi and application thereof in base editing

    CN122278807A

  • Novel uracil DNA glycosylase ehi and its application in base editing

    CN122278807B