Deaminases for base editing

By designing cytidine base editing agents containing polynucleotide programmable DNA binding domains and cytidine deaminases, the specificity and efficiency of existing base editing agents are solved, and more efficient nucleic acid sequence editing is achieved.

CN118475691BActive Publication Date: 2025-07-22ACCUREDIT THERAPEUTICS (SUZHOU) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480001057.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-06-09
Filing Date
2024-02-27
Publication Date
2025-07-22
Estimated Expiration
2044-02-27

AI Technical Summary

Technical Problem

When existing base editers target the editing of nucleic acid sequences, there are problems of specificity and efficiency, especially when converting C·G base pairs to T·A base pairs and A·T base pairs to G·C base pairs.

Method used

A cytidine base editor is provided that contains a polynucleotide programmable DNA binding domain and a cytidine deaminase, which enhances the cis/trans activity ratio and improves editing efficiency and specificity through the introduction of uracil glycosylate enzyme inhibitors (UGI) and nuclear localization signals (NLS).

Benefits of technology

The editing efficiency and specificity of cytidine base editing agents are improved, especially the editing efficiency of GC, AC, and TC sites is increased, which enhances cis activity and reduces trans activity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118475691B_ABST
    Figure CN118475691B_ABST
Patent Text Reader

Abstract

The present invention provides a nucleobase editor comprising a cytosine deaminase, a composition comprising the editor, and a method of using the editor to generate a modification in a target nucleobase sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] [Cross-reference to related applications]

[0002] This application claims priority to PCT / CN2023 / 078869, filed on February 28, 2023, entitled “Deaminases for use in base editing,” and PCT / CN2023 / 099407, filed on June 9, 2023, entitled “Deaminases for use in base editing,” the contents of which are incorporated herein by reference in their entirety.

Technical field

[0003] The present disclosure relates to deaminases, such as nucleic acid-guided nucleases, for use in modifying target nucleic acid sequences. [Background Technology]

[0004] Targeted editing of nucleic acid sequences, such as targeted cutting or targeted modification of genomic DNA, is an effective method for studying gene function and has the potential to provide new therapies for human genetic diseases. Currently available base editors include cytidine base editors (such as BE4) that convert target C·G base pairs to T·A base pairs and adenine base editors (such as ABE7.10) that convert A·T base pairs to G·C base pairs. There is a need in the art for improved base editors that can induce modifications within the target sequence with greater specificity and efficiency. [Summary of the Invention]

[0005] In a first aspect, the present disclosure provides a cytidine base editor comprising: (i) a polynucleotide programmable DNA binding domain; (ii) a cytidine deaminase, wherein the cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 1-24 or 70-124. In some embodiments, the cytidine base editor further comprises one or more uracil glycosylase inhibitor (UGI) domains. In some embodiments, the cytidine base editor comprises two glycosylase inhibitor (UGI) domains. In some embodiments, the cytidine base editor further comprises one or more nuclear localization signals (NLS). In some embodiments, the cytidine base editor comprises an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.

[0006] In some embodiments, the cytidine base editor has an increased cis / trans activity ratio compared to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the increased cis / trans activity ratio is at least 2, 2.5, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60-fold or more relative to the BE4max cytidine base editor. In some embodiments, the cytidine base editor has a 5%, 10%, 20%, 30%, 40%, or 50% increased editing efficiency at a GC site relative to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the cytidine base editor has a 5%, 10%, 20%, 30%, 40%, or 50% increased editing efficiency at an AC site relative to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the cytidine base editor has a 5%, 10%, 20%, 30%, 40%, or 50% increased editing efficiency at the TC site relative to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the cytidine base editor has a 5%, 10%, 20%, 30%, 40%, or 50% increased editing efficiency at the CC site relative to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the cytidine base editor has at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120% or more cis activity compared to the BE4max cytidine base editor. In some embodiments, the cytidine base editor has at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120% or less trans activity compared to the BE4max cytidine base editor.

[0007] In some embodiments, the programmable DNA binding domain is a Cas9 selected from the group consisting of Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9. In some embodiments, the cytidine deaminase is a cytidine deaminase from Burkholderia cenocepacia, or a cytidine deaminase having an amino acid sequence that is at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 1. In some embodiments, the cytidine deaminase is an SCP1.201-like deaminase from Streptacidiphilus jeojiense, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 2. In some embodiments, the cytidine deaminase is an Imml family immunoprotease from Brevilactibacter coleopterorum, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 3. In some embodiments, the cytidine deaminase is a DUF6531 domain-containing protease from Pseudomonas koroiensis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 4. In some embodiments, the cytidine deaminase is a PAAR domain-containing protease from Yersinia similis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 5. In some embodiments, the cytidine deaminase is a RHS repeat protease from Salmonella enterica, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6. In some embodiments, the cytidine deaminase is a PAAR domain-containing protease from Pseudomonas, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO:7.In some embodiments, the cytidine deaminase is an APOBEC-1 enzyme from Tupaia chinensis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 8. In some embodiments, the cytidine deaminase is an APOBEC-1-like enzyme from Suncus etruscus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 9. In some embodiments, the cytidine deaminase is an APOBEC-1 enzyme from Fukomys damarensis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 10. In some embodiments, the cytidine deaminase is an APOBEC-3D-like enzyme from Hyaena hyaena, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 11. In some embodiments, the cytidine deaminase is an APOBEC-1-like enzyme from Falco cherrug, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 12. In some embodiments, the cytidine deaminase is an APOBEC-1-like enzyme from Chelonia mydas, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 13. In some embodiments, the cytidine deaminase is APOBEC-1 from a snapping turtle (Chelydra serpentina), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 14. In some embodiments, the cytidine deaminase is APOBEC-1 from a starling (Sturnus vulgaris), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 15. In some embodiments, the cytidine deaminase is an APOBEC-3C-like enzyme from a Peruvian night monkey (Aotus nancymaae), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 16.In some embodiments, the cytidine deaminase is APOBEC-1 from Cricetulus griseus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 17. In some embodiments, the cytidine deaminase is a single-chain cytosine deaminase-like enzyme from Syngnathus scovelli, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 18. In some embodiments, the cytidine deaminase is an ABEC1 enzyme from Edolisoma coerulescens, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 19. In some embodiments, the cytidine deaminase is an AICDA deaminase from Trogon melanurus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 20. In some embodiments, the cytidine deaminase is an APOBEC-3A from Ochotona princeps, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 21. In some embodiments, the cytidine deaminase is an APOBEC-3A-like enzyme from Trichechus manatuslatirostris, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 22. In some embodiments, the cytidine deaminase is an APOBEC-3-like enzyme from Acipenser ruthenus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 23. In some embodiments, the cytidine deaminase is an APOBEC-3-like enzyme from Chelonoidis abingdonii, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 24.

[0008] In some embodiments, the cytidine deaminase comprises one or more alterations at positions R57X, W114X, and H146X relative to SEQ ID NO: 8. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of R57A, W114Y, and H146A relative to SEQ ID NO: 8. In some embodiments, the cytidine deaminase comprises one or more alterations at positions K34X, W89X, and H121X relative to SEQ ID NO: 9. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of K34A, W89Y, and H121A relative to SEQ ID NO: 9. In some embodiments, the cytidine deaminase comprises one or more alterations at positions R218X, W275X, and H307X relative to SEQ ID NO: 11. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of R218A, W275Y, and H307A relative to SEQ ID NO: 11. In some embodiments, the cytidine deaminase comprises one or more alterations at positions R53X, W122X, and Y152X relative to SEQ ID NO: 16, or one or more corresponding alterations thereof, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of R53A, W122Y, and Y152F relative to SEQ ID NO: 16.

[0009] In another aspect, the present disclosure provides a cell comprising a cytidine base editor as described above. In some embodiments, the cell is a bacterial cell, a plant cell, an insect cell, or a mammalian cell.

[0010] In another aspect, the present invention discloses a molecular complex comprising a cytidine base editor as described above and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences.

[0011] On the other hand, the present disclosure provides a method for editing the nucleobase of a nucleic acid sequence, the method comprising contacting the nucleic acid sequence with a cytidine base editor as described above, and converting the first nucleobase of the nucleic acid sequence into a second nucleobase. In some embodiments, the method further comprises contacting the nucleic acid sequence with a guide polynucleotide to achieve the conversion. In some embodiments of the method, the first nucleobase is cytosine and the second nucleobase is thymidine.

[0012] In another aspect, the present disclosure provides a nucleic acid vector encoding a cytidine base editor comprising: (i) a polynucleotide programmable DNA binding domain; (ii) an APOBEC cytidine deaminase, wherein the APOBEC cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 1-24 or 70-124. In some embodiments, the vector further comprises a polynucleotide sequence encoding one or more uracil glycosylase inhibitor (UGI) domains. In some embodiments, the vector further comprises a polynucleotide sequence encoding two glycosylase inhibitor (UGI) domains. In some embodiments, the vector further comprises a polynucleotide sequence encoding one or more nuclear localization signals (NLS). In some embodiments, the vector further comprises a polynucleotide sequence encoding an N-terminal NLS and / or a C-terminal NLS. In some embodiments of the vector, the NLS is a bipartite NLS. In some embodiments, the programmable DNA-binding domain is a Cas9 selected from the group consisting of Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9.

[0013] On the other hand, the present disclosure provides a nucleic acid vector encoding a cytidine base editor, comprising: (i) a polynucleotide programmable DNA binding domain; (ii) a cytidine deaminase, wherein the nucleic acid encoding the cytidine deaminase has a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 25-48 or 125-179. In some embodiments, the vector further comprises a polynucleotide sequence encoding one or more uracil glycosylase inhibitor (UGI) domains. In some embodiments, the vector further comprises a polynucleotide sequence encoding two glycosylase inhibitor (UGI) domains. In some embodiments, the vector further comprises a polynucleotide sequence encoding one or more nuclear localization signals (NLS). In some embodiments, the vector further comprises a polynucleotide sequence encoding an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS. In some embodiments, the programmable DNA-binding domain is a Cas9 selected from the group consisting of Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9.

[0014] In another aspect, the present disclosure provides a composition comprising: (i) a nucleic acid, or a vector comprising a nucleic acid, encoding a guide RNA; (ii) a cytidine base editor comprising: (i) a polynucleotide programmable DNA binding domain; (ii) a cytidine deaminase, wherein the cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 1-24 or 70-124. In some embodiments, the composition further comprises one or more uracil glycosylase inhibitor (UGI) domains. In some embodiments, the cytidine base editor comprises two glycosylase inhibitor (UGI) domains. In some embodiments, the composition further comprises one or more nuclear localization signals (NLS). In some embodiments, the cytidine base editor comprises an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.

[0015] In some embodiments, the cytidine base editor has an increased cis / trans activity ratio compared to the BE4max cytidine base editor. In some embodiments, the increased cis / trans activity ratio is at least 2, 2.5, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60-fold or more relative to the BE4max cytidine base editor. In some embodiments, the cytidine base editor has at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120% or more cis activity compared to the BE4max cytidine base editor. In some embodiments, the cytidine base editor has at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120% or less trans activity compared to the BE4max cytidine base editor. In some embodiments, the programmable DNA binding domain is a Cas9 selected from the group consisting of Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), or nuclease-active Cas9.

[0016] On the other hand, the present disclosure provides a fusion protein comprising: a polynucleotide programmable DNA binding domain and at least one nucleobase editing subdomain comprising a cytidine deaminase, wherein the cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 1 to 24 or 70 to 124. In some embodiments, the fusion protein further comprises one or more uracil glycosylase inhibitor (UGI) domains. In some embodiments, the fusion protein further comprises two glycosylase inhibitor (UGI) domains. In some embodiments, the fusion protein further comprises one or more nuclear localization signals (NLS). In some embodiments, the fusion protein further comprises an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.

[0017] In some embodiments, the cytidine deaminase is a cytidine deaminase from Burkholderia cenocepacia, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 1. In some embodiments, the cytidine deaminase is an SCP1.201-like deaminase from Streptacidiphilus jeojiense, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 2. In some embodiments, the cytidine deaminase is an Imml family immunoprotease from Brevilactibacter coleopterorum, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 3. In some embodiments, the cytidine deaminase is a DUF6531 domain-containing protease from Pseudomonas koroiensis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 4. In some embodiments, the cytidine deaminase is a PAAR domain-containing protease from Yersinia similis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 5. In some embodiments, the cytidine deaminase is a RHS repeat protease from Salmonella enterica, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6. In some embodiments, the cytidine deaminase is a PAAR domain-containing protease from Pseudomonas, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 7. In some embodiments, the cytidine deaminase is an APOBEC-1 enzyme from Tupaia chinensis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:8.In some embodiments, the cytidine deaminase is an APOBEC-1-like enzyme from Suncus etruscus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 9. In some embodiments, the cytidine deaminase is an APOBEC-1 enzyme from Fukomys damarensis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 10. In some embodiments, the cytidine deaminase is an APOBEC-3D-like enzyme from Hyaena hyaena, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 11. In some embodiments, the cytidine deaminase is an APOBEC-1-like enzyme from a hunting falcon (Falco cherrug), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 12. In some embodiments, the cytidine deaminase is an APOBEC-1-like enzyme from a green sea turtle (Chelonia mydas), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 13. In some embodiments, the cytidine deaminase is an APOBEC-1 from a snapping turtle (Chelydra serpentina), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 14. In some embodiments, the cytidine deaminase is APOBEC-1 from Sturnus vulgaris, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 15. In some embodiments, the cytidine deaminase is an APOBEC-3C-like enzyme from Aotus nancymaae, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 16. In some embodiments, the cytidine deaminase is APOBEC-1 from Cricetulus griseus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 17.In some embodiments, the cytidine deaminase is a single-chain cytosine deaminase-like enzyme from Syngnathus scovelli, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 18. In some embodiments, the cytidine deaminase is an ABEC1 enzyme from Edolisoma coerulescens, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 19. In some embodiments, the cytidine deaminase is an AICDA deaminase from Trogon melanurus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 20. In some embodiments, the cytidine deaminase is an APOBEC-3A enzyme from Ochotona princeps, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 21. In some embodiments, the cytidine deaminase is an APOBEC-3A-like enzyme from Trichechus manatus latirostris, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 22. In some embodiments, the cytidine deaminase is an APOBEC-3-like enzyme from Acipenser ruthenus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 23. In some embodiments, the cytidine deaminase is an APOBEC-3-like enzyme from the Pinta Island giant tortoise (Chelonoidis abingdonii), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:24.

[0018] In some embodiments, the cytidine deaminase comprises one or more alterations relative to SEQ ID NO: 8 at positions R57X, W114X, and H146X, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations relative to SEQ ID NO: 8 at positions selected from the group consisting of R57A, W114Y, and H146A. In some embodiments, the cytidine deaminase comprises one or more alterations relative to SEQ ID NO: 9 at positions K34X, W89X, and H121X, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations relative to SEQ ID NO: 9 at positions selected from the group consisting of K34A, W89Y, and H121A. In some embodiments, the cytidine deaminase comprises one or more alterations relative to SEQ ID NO: 11 at positions R218X, W275X, and H307X. In some embodiments, the cytidine deaminase comprises one or more alterations relative to SEQ ID NO: 11 at a position selected from the group consisting of R218A, W275Y, and H307A. In some embodiments, the cytidine deaminase comprises one or more alterations relative to SEQ ID NO: 16 at positions R53X, W122X, and Y152X. In some embodiments, the cytidine deaminase comprises one or more alterations relative to SEQ ID NO: 16 at a position selected from the group consisting of R53A, W122Y, and Y152F.

[0019] In another aspect, the present disclosure provides an engineered system comprising a catalytically inactive Cas13 effector protein (dCas13) or a nucleotide sequence encoding a catalytically inactive Cas13 effector protein, wherein the Cas13 effector protein is Cas13a, wherein the system further comprises a functional component or the nucleotide sequence further encodes the functional component, wherein the functional component is a base editing component, wherein the base editing component comprises a cytidine deaminase or a catalytic domain thereof, and wherein the adenosine deaminase or a catalytic domain thereof is inserted into an internal loop of dCas13, wherein the cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 1-24 or 70-124.

[0020] In some embodiments, the cytidine deaminase of the engineered system is a cytidine deaminase from Burkholderia cenocepacia, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 1. In some embodiments, the cytidine deaminase of the engineered system is an SCP1.201-like deaminase from Streptacidiphilus jeojiense, or a cytidine deaminase having an amino acid deaminase at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 2. In some embodiments, the cytidine deaminase in the engineered system is an Imml family immunoprotease from Brevilactibacter coleopterorum, or a cytidine deaminase having an amino acid sequence that is at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 3. In some embodiments, the cytidine deaminase in the engineered system is a protease containing a DUF6531 domain from Pseudomonas koroiensis, or a cytidine deaminase having an amino acid sequence that is at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 4. In some embodiments, the cytidine deaminase of the engineered system is a PAAR domain-containing protease from Yersinia similis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 5. In some embodiments, the cytidine deaminase of the engineered system is a RHS repeat protease from Salmonella enterica, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 6. In some embodiments, the cytidine deaminase of the engineered system is a PAAR domain-containing protease from Pseudomonas, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 7. In some embodiments, the cytidine deaminase of the engineered system is an APOBEC-1 enzyme from Tupaia chinensis, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:8.In some embodiments, the cytidine deaminase in the engineered system is an APOBEC-1-like enzyme from Suncus etruscus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 9. In some embodiments, the cytidine deaminase in the engineered system is an APOBEC-1 enzyme from Damaraland mole-rat (Fukomys damarensis), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 10. In some embodiments, the cytidine deaminase of the engineered system is an APOBEC-3D-like enzyme from Hyaena hyaena, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 11. In some embodiments, the cytidine deaminase of the engineered system is an APOBEC-1-like enzyme from Falco cherrug, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 12. In some embodiments, the cytidine deaminase of the engineered system is an APOBEC-1-like enzyme from Chelonia mydas, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 13. In some embodiments, the cytidine deaminase of the engineered system is APOBEC-1 from the snapping turtle (Chelydra serpentina), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 14. In some embodiments, the cytidine deaminase of the engineered system is APOBEC-1 from the starling (Sturnus vulgaris), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 15. In some embodiments, the cytidine deaminase of the engineered system is an APOBEC-3C-like enzyme from the Peruvian night monkey (Aotus nancymaae), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 16.In some embodiments, the cytidine deaminase of the engineered system is APOBEC-1 from Chinese hamster (Cricetulus griseus), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 17. In some embodiments, the cytidine deaminase of the engineered system is a single-chain cytosine deaminase-like enzyme from gulf sea dragon (Syngnathus scovelli), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 18. In some embodiments, the cytidine deaminase in the engineered system is an ABEC1 enzyme from Edolisoma coerulescens, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 19. In some embodiments, the cytidine deaminase in the engineered system is an AICDA deaminase from Trogon melanurus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 20. In some embodiments, the cytidine deaminase in the engineered system is APOBEC-3A from Ochotona princeps, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 21. In some embodiments, the cytidine deaminase of the engineered system is an APOBEC-3A-like enzyme from Florida manatee (Trichechus manatus latirostris), or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 22. In some embodiments, the cytidine deaminase of the engineered system is an APOBEC-3-like enzyme from Acipenser ruthenus, or a cytidine deaminase having an amino acid sequence at least 80%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 23. In some embodiments, the cytidine deaminase of the engineered system is an APOBEC-3-like enzyme from the Pinta Island giant tortoise (Chelonoidis abingdonii), or a cytidine deaminase having an amino acid sequence that is at least 80%, 90%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 24.

[0021] In another aspect, the present disclosure provides a cytidine base editor comprising: (i) a polynucleotide programmable DNA binding domain; (ii) a cytidine deaminase, wherein the cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124 and 201-289.

[0022] In some embodiments, the cytidine deaminase comprises a cytidine deaminase numbered relative to SEQ ID NO: 18, or one or more corresponding alterations thereof, or one or more corresponding alterations thereof, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of K20A, R21A, Y22A, Y22F, R26A, T28A, L30A, N46A, A52D, C73F, C73W, L75A, W77A, W77F, S78A, P79A, L87A, I101A, Y107A, Y107F, Y108A, and Y108F as numbered in SEQ ID NO: 18, or one or more corresponding alterations thereof. In some embodiments, the cytidine deaminase further comprises one or more uracil glycosylase inhibitor (UGI) domains.

[0023] In some embodiments, the cytidine base editor comprises two glycosylase inhibitor (UGI) domains. In some embodiments, the cytidine base editor further comprises one or more nuclear localization signals (NLS). In some embodiments, the cytidine base editor comprises an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.

[0024] In some embodiments, the cytidine base editor has an increased cis / trans activity ratio compared to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the increased cis / trans activity ratio is at least 2, 2.5, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60-fold or more relative to the BE4max cytidine base editor. In some embodiments, the cytidine base editor has a 5%, 10%, 20%, 30%, 40%, or 50% increased editing efficiency at a GC site relative to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the cytidine base editor has a 5%, 10%, 20%, 30%, 40%, or 50% increased editing efficiency at an AC site relative to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the cytidine base editor has a 5%, 10%, 20%, 30%, 40%, or 50% increased editing efficiency at the TC site relative to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the cytidine base editor has a 5%, 10%, 20%, 30%, 40%, or 50% increased editing efficiency at the CC site relative to a BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by a polynucleotide that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 49. In some embodiments, the cytidine base editor has at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120% or more cis activity compared to the BE4max cytidine base editor. In some embodiments, the cytidine base editor has at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120% or less trans activity compared to the BE4max cytidine base editor.

[0025] In another aspect, the present disclosure provides a cell comprising a cytidine base editor. In some embodiments, the cell is a bacterial cell, a plant cell, an insect cell, or a mammalian cell.

[0026] In another aspect, the present disclosure provides a molecular complex comprising a cytidine base editor and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences.

[0027] On the other hand, the present disclosure provides a method for editing a nucleic acid sequence, comprising contacting the nucleic acid sequence with a cytidine base editor and converting a first nucleic acid base of the nucleic acid sequence to a second nucleic acid base. In some embodiments, the method further comprises contacting the nucleic acid sequence with a guide polynucleotide to achieve the conversion. In some embodiments, the first nucleic acid base is cytosine and the second nucleic acid base is thymidine.

[0028] In another aspect, the present disclosure provides a nucleic acid vector encoding a cytidine base editor comprising: (i) a polynucleotide programmable DNA binding domain; (ii) a cytidine deaminase, wherein the cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 18, 1-24, 70-124 and 201-289. In some embodiments, the vector further comprises a polynucleotide sequence encoding one or more uracil glycosylase inhibitor (UGI) domains. In some embodiments, the vector further comprises a polynucleotide sequence encoding two glycosylase inhibitor (UGI) domains. In some embodiments, the vector further comprises a polynucleotide sequence encoding one or more nuclear localization signals (NLS). In some embodiments, the vector further comprises a polynucleotide sequence encoding an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.

[0029] In some embodiments, the programmable DNA binding domain is a Cas9 selected from the group consisting of Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), nuclease-dead Cas12, nuclease-dead TnpB, nuclease-dead IscB, nickase IscB, nuclease-active Cas12, nuclease-active TnpB, nuclease-active IscB, or nuclease-active Cas9.

[0030] On the other hand, the present disclosure provides a nucleic acid vector encoding a cytidine base editor, comprising: (i) a polynucleotide programmable DNA binding domain; (ii) a cytidine deaminase, wherein the nucleic acid encoding the cytidine deaminase has a polynucleotide sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 42, 25-41, 43-48, 125-179 or 297-385. In some embodiments, the vector further comprises a polynucleotide sequence encoding one or more uracil glycosylase inhibitor (UGI) domains. In some embodiments, the vector further comprises a polynucleotide sequence encoding two glycosylase inhibitor (UGI) domains. In some embodiments, the vector further comprises a polynucleotide sequence encoding one or more nuclear localization signals (NLS). In some embodiments, the vector further comprises a polynucleotide sequence encoding an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.

[0031] In some embodiments, the programmable DNA binding domain is a Cas9 selected from the group consisting of Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), nuclease-dead Cas12, nuclease-dead TnpB, nuclease-dead IscB, IscB nickase, nuclease-active Cas12, nuclease-active TnpB, nuclease-active IscB, or nuclease-active Cas9.

[0032] In another aspect, the present disclosure provides a composition comprising: (i) a nucleic acid, or a vector comprising a nucleic acid, encoding a guide RNA; (ii) a cytidine base editor comprising: (i) a polynucleotide programmable DNA binding domain; (ii) a cytidine deaminase, wherein the cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124 and 201-289. In some embodiments, the composition further comprises one or more uracil glycosylase inhibitor (UGI) domains. In some embodiments, the cytidine base editor comprises two glycosylase inhibitor (UGI) domains. In some embodiments, the cytidine base editor further comprises one or more nuclear localization signals (NLS). In some embodiments, the cytidine base editor comprises an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.

[0033] In some embodiments, the cytidine base editor has an increased cis / trans activity ratio compared to the BE4max cytidine base editor. In some embodiments, the increased cis / trans activity ratio is at least 2, 2.5, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60-fold or more relative to the BE4max cytidine base editor. In some embodiments, the cytidine base editor has at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120% or more cis activity compared to the BE4max cytidine base editor. In some embodiments, the cytidine base editor has at least 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 100%, 105%, 110%, 115%, 120% or less trans activity compared to the BE4max cytidine base editor.

[0034] In some embodiments, the programmable DNA binding domain is a Cas9 selected from the group consisting of Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), nuclease-dead Cas12, nuclease-dead TnpB, nuclease-dead IscB, IscB nickase, nuclease-active Cas12, nuclease-active TnpB, nuclease-active IscB, or nuclease-active Cas9.

[0035] In another aspect, the present disclosure provides a fusion protein comprising: a polynucleotide programmable DNA binding domain and at least one nucleobase editor domain, including a cytidine deaminase, wherein the cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124 and 201-289.

[0036] In some embodiments, the fusion protein further comprises one or more uracil glycosylase inhibitor (UGI) domains. In some embodiments, the fusion protein further comprises two glycosylase inhibitor (UGI) domains. In some embodiments, the fusion protein further comprises one or more nuclear localization signals (NLS). In some embodiments, the fusion protein further comprises an N-terminal NLS and / or a C-terminal NLS. In some embodiments, the NLS is a bipartite NLS.

[0037] In another aspect, the present disclosure provides an engineered system comprising a catalytically inactive Cas13 effector protein (dCas13) or a nucleotide sequence encoding a catalytically inactive Cas13 effector protein, wherein the Cas13 effector protein is Cas13a, wherein the system further comprises a functional component or the nucleotide sequence further encodes a functional component, wherein the functional component is a base editing component, wherein the base editing component comprises a cytidine deaminase or a catalytic domain thereof, and wherein the cytidine deaminase or the catalytic domain thereof is inserted into an internal loop of dCas13, wherein the cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124 and 201-289.

[0038] In another aspect, the present disclosure provides a biomolecule deaminase complementary base editing system, the base editing system comprising at least one of the following:

[0039] (1) a base editor fusion protein A comprising an amino acid sequence at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 201, a base editor fusion protein B comprising an amino acid sequence at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 202, and a guide RNA;

[0040] (2) an expression construct comprising a nucleotide sequence encoding: a base editor fusion protein A comprising an amino acid sequence that is at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 201, a base editor fusion protein B comprising an amino acid sequence that is at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 202, and a guide RNA;

[0041] (3) a base editor fusion protein A comprising an amino acid sequence that is at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 201, a base editor fusion protein B comprising an amino acid sequence that is at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 202, and an expression construct comprising a nucleotide sequence encoding a guide RNA;

[0042] (4) an expression construct comprising a nucleotide sequence encoding: a base editor fusion protein A comprising an amino acid sequence that is at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 201, a base editor fusion protein B comprising an amino acid sequence that is at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 202, and an expression construct comprising a nucleotide sequence encoding a guide RNA;

[0043] (5) An expression construct comprising a nucleotide sequence encoding: a base editor fusion protein A comprising an amino acid sequence that is at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 201, a base editor fusion protein B comprising an amino acid sequence that is at least 80%, 85%, 90% or 95% identical to SEQ ID NO: 202, and a nucleotide sequence encoding a guide RNA.

[0044] In some embodiments, the base editing fusion protein A comprises a first nCas9 polypeptide fragment, a flexible linker peptide, and a first nucleobase deaminase polypeptide fragment arranged in sequence from N-terminus to C-terminus; the base editing fusion protein B comprises a second nucleobase deaminase polypeptide fragment, a flexible linker peptide, and a second nCas9 polypeptide fragment from N-terminus to C-terminus; the first nucleobase deaminase polypeptide fragment and the second nucleobase deaminase polypeptide fragment are selected from the same nucleobase deaminase.

[0045] On the other hand, the present disclosure provides a fusion protein comprising a nucleic acid-guided nuclease protein and a chimeric nucleotide base editor (NBE) domain. In some embodiments, the embedded NBE domain is a cytidine base editor (CBE) domain. In some embodiments, the chimeric CBE domain is a chimeric cytosine deaminase protein domain. In some embodiments, the nucleic acid-guided nuclease protein comprises a protospacer-adjacent motif interaction domain (PID) having a peptide that hybridizes to an N4CC nucleotide sequence or an N4C nucleotide sequence. In some embodiments, the fusion protein further comprises a nuclear localization signal (NLS) protein selected from the group consisting of: nucleoplasmin NLS, SV40 NLS, and C-myc NLS. In some embodiments, the fusion protein further comprises a uracil glycosylase inhibitor. In some embodiments, the nucleic acid-guided nuclease protein comprises a mutation. In some embodiments, the nucleic acid-guided nuclease protein comprises a linker flanking at least one chimeric domain protein.

[0046] In another aspect, the present disclosure provides a biomolecule deaminase complementary base editing system, the base editing system comprising at least one of the following:

[0047] (1) a base editing fusion protein A comprising all or part of any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124, and 201-289; a base editing fusion protein B comprising all or part of any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124, and 201-289; and a guide RNA;

[0048] (2) an expression construct comprising a nucleotide sequence encoding: a base editing fusion protein A comprising all or part of any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124, and 201-289; a base editing fusion protein B comprising all or part of any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124, and 201-289; and a guide RNA;

[0049] (3) a base editing fusion protein A comprising all or part of any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124, and 201-289; a base editing fusion protein B comprising all or part of any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124, and 201-289; and an expression construct comprising a nucleotide sequence encoding a guide RNA;

[0050] (4) an expression construct comprising a nucleotide sequence encoding: a base editing fusion protein A comprising all or part of any one of SEQ ID NOs: 18, 1 to 17, 19 to 24, 70 to 124, and 201 to 289; a base editing fusion protein B comprising all or part of any one of SEQ ID NOs: 18, 1 to 17, 19 to 24, 70 to 124, and 201 to 289; and an expression construct comprising a nucleotide sequence encoding a guide RNA;

[0051] (5) An expression construct comprising a nucleotide sequence encoding: a base editing fusion protein A comprising all or part of any one of SEQ ID NOs: 18, 1 to 17, 19 to 24, 70 to 124, and 201 to 289; a base editing fusion protein B comprising all or part of any one of SEQ ID NOs: 18, 1 to 17, 19 to 24, 70 to 124, and 201 to 289; and a nucleotide sequence encoding a guide RNA.

[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials for use in the present invention are described herein; other suitable methods and materials known in the art may also be used. The materials, methods, and examples are illustrative only and are not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In the event of conflict, this specification (including definitions) will control.

[0053] Other features and advantages of the invention will be apparent from the following detailed description and drawings, and from the claims.

Brief Description of the Drawings

[0054] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication and color drawings will be provided by the Office upon request and payment of the necessary fee.

[0055] Picture 1 Figure 2 is a graph of C to T editing efficiency using the eight unique guide RNAs disclosed herein and each cytosine base editor.

[0056] Picture 2A Figure 3 is a plot of the average C to T editing efficiency of the BE4max cytidine base editor for 14 different sgRNAs at each position within the protospacer.

[0057] Picture 2B Figure 3 is a plot of the average C to T editing efficiency of 14 different sgRNAs for the CE_1 cytidine base editor at each position within the protospacer.

[0058] Picture 2C Figure 3 is a plot of the average C to T editing efficiency of 14 different sgRNAs for the CE_2 cytidine base editor at each position within the protospacer.

[0059] Picture 2D Figure 3 is a plot of the average C to T editing efficiency of 14 different sgRNAs for the CE_4 cytidine base editor at each position within the protospacer.

[0060] Picture 2E Figure 3 is a plot of the average C to T editing efficiency of 14 different sgRNAs for the CE_9 cytidine base editor at each position within the protospacer.

[0061] Picture 2F Figure 3 is a plot of the average C to T editing efficiency of 14 different sgRNAs for the CE_12 cytidine base editor at each position within the protospacer.

[0062] Picture 2GFigure 3 is a plot of the average C to T editing efficiency of 14 different sgRNAs for the CE_16 cytidine base editor at each position within the protospacer.

[0063] Picture 3A Figure 2 is a plot of the percentage of next-generation sequencing (NGS) reads that include C to T edits by the BE4max cytidine base editor across positions within the protospacer.

[0064] Picture 3B is a plot of the percentage of next-generation sequencing (NGS) reads that include C to T edits across positions within the protospacer of the CE_1 cytidine base editor.

[0065] Picture 3C is a plot of the percentage of next-generation sequencing (NGS) reads that include C to T edits spanning positions within the protospacer of the CE_2 cytidine base editor.

[0066] Photo 3D is a plot of the percentage of next-generation sequencing (NGS) reads that include C to T edits spanning positions within the protospacer of the CE_4 cytidine base editor.

[0067] Picture 3E Figure 3 is a plot of the percentage of next-generation sequencing (NGS) reads that include C to T edits at various positions within the protospacer of the CE_9 cytidine base editor.

[0068] Picture 3F Figure 3 is a plot of the percentage of next-generation sequencing (NGS) reads that include C to T edits across positions within the protospacer of the CE_12 cytidine base editor.

[0069] Picture 4A Figure 2 is a graph of the percentage of NGS reads containing insertions and / or deletions (indels) following editing with each of several cytidine base editors compared to BE4max.

[0070] Picture 4B Figure 2 is a graph showing the ratio of editing efficiency to indel rate after editing with several cytidine base editors compared to BE4max.

[0071] Picture 5 Figure 2 is a scatter plot of the orthogonal R-loop assay results, showing on-target and off-target editing of cytidine base editors.

[0072] Picture 6 The present invention provides the results of an orthogonal R-loop assay for measuring the outcomes of on-target editing and off-guide single-stranded DNA deamination mediated by engineered variants of the base editors disclosed herein.

[0073] Picture 7Figure 2 is a graph of C to T editing efficiency using the five unique guide RNAs disclosed herein and each cytosine base editor.

[0074] Picture 8 is a graph showing the C to T editing efficiency of the cytosine base editor using sgRNA15 disclosed herein.

[0075] Picture 9A Graph showing the C to T editing efficiency of cytosine base editors at different positions of sgRNA5 disclosed herein.

[0076] Picture 9B Graph showing the C to T editing efficiency of cytosine base editors at different positions of sgRNA14 disclosed herein.

[0077] Picture 10 The figure shows the calculated accuracy ratio of cytosine base editors at two target sites relative to BE4max.

[0078] Picture 11 Figure 3 is a plot of off-target / on-target activity of gRNA relative to BE4max, measuring cytosine base editors at two R-loop sites using the R-loop assay.

[0079] Picture 12 Figure 2 is a plot of guide-dependent off-target / on-target activity relative to cytosine base editor activity of BE4max at three target sites.

[0080] Picture 13 Figure 3 is a graph of the total number of C-to-U edited RNA SNVs by cytosine base editors.

[0081] Picture 14 This is a plot of the calculated accuracy ratio of the engineered variant BE4max of the cytosine base editor.

[0082] Picture 15 Figure 3 is a plot of off-target / on-target activity of gRNA relative to BE4max, an engineered variant of the cytosine base editor measured by the R-loop assay.

[0083] Picture 16 Figure 2 is a graph of C to T editing efficiency using the three unique guide RNAs disclosed herein in combination with each cytosine base editor.

[0084] Picture 17 is a graph of the efficiency of C to T on-target or guide-dependent off-target editing using sgRNA4 in combination with each of the cytosine base editors disclosed herein.

[0085]

Detailed description

[0086] Cytidine base editor

[0087] The nucleobase editors disclosed herein, such as cytidine base editors, are used to edit, modify or change the target nucleotide sequence of a polynucleotide. The nucleobase editors described herein comprise a programmable nucleotide binding domain, such as a polynucleotide programmable nucleotide binding domain (e.g., Cas9) or a zinc finger protein DNA binding domain or a TALE DNA binding domain and at least one nucleobase editing domain, such as a cytidine deaminase. The polynucleotide programmable nucleotide binding domain (e.g., Cas9 or Cas12) when present in a cell and bound to a bound guide polynucleotide (e.g., gRNA) can specifically bind to a target polynucleotide sequence (i.e., by complementary base pairing between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence), thereby positioning the base editor to the target nucleic acid sequence that needs to be edited.

[0088] Programmable nucleotide-binding domain

[0089] The programmable nucleotide binding domain of the base editor may itself comprise one or more domains. For example, a polynucleotide programmable nucleotide binding domain may comprise one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain may comprise an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to a protein or polypeptide capable of digesting a nucleic acid (e.g., RNA or DNA) from a free end, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cutting) an internal region of a nucleic acid (e.g., DNA or RNA). In some embodiments, an endonuclease can cut a single strand of a double-stranded nucleic acid.

[0090] Any combination of any DNA unstable molecules can be used in the compositions described herein, including but not limited to Cas9 or Cas12 nickases, IscB, TnpB, Cas9 or Cas12 proteins (e.g., dCas) operably linked to a single guide RNA (sgRNA), any RNA programmable system, zinc finger nuclease nickases (ZFN nickases), TALEN nickases and / or one or more nucleotides (e.g., one or more peptide nucleic acids (PNAs), locked nucleic acids (LNAs) and / or bridging nucleic acids (BNAs)). In certain embodiments, the base editing composition comprises a plurality of DNA unstable molecules, such as one or more proteins (e.g., nickases, etc.) and / or one or more nucleotides. In certain embodiments, the composition comprises a ZFN nickase and one or more additional proteins and / or nucleotide DNA unstable molecules (e.g., one or more nucleotides described herein). In certain aspects, the base editing composition does not comprise Cas9 protein, but may comprise other Cas proteins (e.g., non-Cas9 RNA programmable systems). In certain embodiments, the DNA unstable molecule comprises a zinc finger nuclease (ZFN) nickase.

[0091] In some embodiments, the nuclease is a zinc finger nuclease (ZFN) or a TALE DNA binding domain-nuclease fusion (TALEN). ZFNs and TALENs contain a DNA binding domain (zinc finger protein or TALE DNA binding domain) that has been designed to bind to a target site in a selection gene and a cleavage domain or cleavage half-domain.

[0092] At least one zinc finger protein (ZFP) DNA binding domain of the base editing composition can be operably connected to one or more other components of the base editing composition, for example, to one or more DNA unstable molecules (for example, to Cas9 nickase, dCas9, etc.) and / or to at least one adenine or cytosine deaminase. In certain embodiments, at least one ZFP DNA binding domain is operably connected to adenine or cytosine deaminase. In other embodiments, the base editing composition comprises a first and a second ZFP DNA binding domain, wherein the first ZFP DNA binding domain is operably connected to the Cas9 nickase. The ZFP DNA binding domain may include 3, 4, 5, 6 or more fingers and may be bound to a target site on either side (5' or 3') of the target base to be edited. In certain embodiments, the ZFP binds to a target site of 1 to 100 (or any number therebetween) nucleotides on both sides of the target base. In other embodiments, the ZFP binds to a target site of 1 to 50 (or any number therebetween) nucleotides on both sides of the target base.

[0093] In some embodiments, the endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide programmable nucleotide binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide programmable nucleotide binding domain can be a ribonuclease.

[0094] In some embodiments, the programmable nucleotide binding domain of the nuclease domain can cut zero strands, one strand or two strands of the target polynucleotide. In some embodiments, the polynucleotide programmable nucleotide binding domain may include a nickase domain. The term "nickase" herein refers to a polynucleotide programmable nucleotide binding domain comprising a nuclease domain that is only capable of cutting one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, a nickase can be derived from a fully catalytically active (e.g., natural) form of a polynucleotide programmable nucleotide binding domain by introducing one or more mutations into an active polynucleotide programmable nucleotide binding domain. For example, when a polynucleotide programmable nucleotide binding domain comprises a nickase domain derived from Cas9, the nickase domain derived from Cas9 may include a D10A mutation and a histidine at position 840. In such an embodiment, residue H840 maintains catalytic activity, thereby enabling the cutting of a single strand of a nucleic acid duplex. In another example, a nickase domain derived from Cas9 may include an H840A mutation, while the amino acid residue at position 10 remains D. In some embodiments, the nickase can be derived from a fully catalytically active (e.g., native) form of a polynucleotide programmable nucleotide binding domain by removing all or part of the nuclease domain that is not required for nickase activity. For example, when the polynucleotide programmable nucleotide binding domain comprises a nickase domain derived from Cas9, the Cas9-derived nickase domain can include a deletion of all or part of the RuvC domain or the HNH domain.

[0095] The base editor comprises a polynucleotide programmable nucleotide binding domain comprising a nickase domain, and is therefore capable of producing single-stranded DNA breaks (nicks) at a specific polynucleotide target sequence (e.g., determined by binding to a complementary sequence of a guide nucleic acid). In some embodiments, the chain of the double-stranded target polynucleotide sequence of nucleic acid cut by a base editor comprising a nickase domain (e.g., a nickase domain derived from Cas9) is a chain that is not edited by the base editor (i.e., the chain cut by the base editor is opposite to the chain comprising the base to be edited). In other embodiments, the base editor comprising a nickase domain (e.g., a nickase domain derived from Cas9) can cut the chain of the DNA molecule that is edited by the target. In such an embodiment, the non-targeting chain is not cut.

[0096] Also provided herein are base editors comprising a polynucleotide programmable nucleotide binding domain that is catalytically dead (i.e., unable to cut the target polynucleotide sequence). Here, the terms "catalytically dead" and "nuclease dead" are used interchangeably to refer to a polynucleotide programmable nucleotide binding domain with one or more mutations and / or deletions, resulting in its inability to cut the nucleic acid chain while maintaining its ability and specificity to bind to the target polynucleotide. In some embodiments, the catalytically dead polynucleotide programmable nucleotide binding domain base editor lacks nuclease activity due to specific point mutations in one or more nuclease domains. For example, in the case of a base editor comprising a Cas9 domain, Cas9 can simultaneously comprise a D10A mutation and an H840A mutation. This mutation inactivates two nuclease domains, resulting in loss of nuclease activity. In other embodiments, the catalytically dead polynucleotide programmable nucleotide binding domain can comprise one or more deletions of all or part of a catalytic domain (e.g., RuvC and / or HNH domain). In further embodiments, the catalytically dead polynucleotide programmable nucleotide binding domain comprises a point mutation (eg, D10A or H840A) and a deletion of all or part of the nuclease domain.

[0097] Some aspects of the disclosure provide fusion proteins comprising domains acting as polynucleotide programmable DNA binding proteins that can be used to guide proteins, such as base editors, to specific nucleic acid (e.g., DNA or RNA) sequences. In a particular embodiment, the fusion protein comprises a nucleic acid programmable DNA binding protein domain and one or more deaminase domains. Non-limiting examples of polynucleotide programmable DNA binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpf1, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpf1, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, C asl2h, Casl2i, Csyl, Csy2, Csy3, Csy4, Csel, Cse2, Cse3, Cse4, Cse5e, Cscl, Csc2, Csa5, Csnl, Csn2, Csml, Csm2, Csm3, Cs m4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csxl0, Csxl6, CsaX, Csx3, Csxl, CsxlS, Csxl 1, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Cshl, Csh2, Csal, Csa2, Csa3, Csa4, Csa5, type II Cas effector protein, type V Cas effector protein, type VI Cas effector protein, CARF, DinG, homologs thereof, or modified or engineered versions thereof. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically listed in the present disclosure.See, e.g., Makarova et alia, “Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?” CRISPR J. 2018 October; 1: 325-336. doi: 10.1089 / crispr.2018.0033; Yan et alia, “Functionally diverse type VCRISPR-Cas systems” Science. 2019 January 4; 363(6422): 88-91. doi: 10.1 126 / science.aav7271, the entire contents of which are incorporated herein by reference.

[0098] In some embodiments, the present disclosure provides a fusion protein comprising a V-type CRISPR / Cas effector protein. The V-type CRISPR / Cas effector protein is a subtype of a 2-type CRISPR / Cas effector protein. For examples of V-type CRISPR / Cas systems and effector proteins thereof (e.g., Cas12 family proteins such as Cas12a), see, for example, Shmakov et alia, Nat Rev Microbial. 2017 March; 15 (3): 169-182: "Diversity and evolution of class 2 CRISPR-Cas systems". Examples include, but are not limited to: Cas12 family (Cas12a, Cas12b, Cas12c), C2c4, C2c8, C2c5, C2c10, and C2c9; and CasX (Cas12e) and CasY (Cas12d). See also, e.g., Koonin et alia, Curr Opin Microbial. 2017 June; 37: 67-78: “Diversity, classification and evolution of CRISPR-Cas systems.” In some embodiments, a CBE disclosed herein comprises a type V CRISPR / Cas effector protein.

[0099] In some embodiments, the present disclosure provides a fusion protein comprising a programmable DNA binding protein. For non-limiting examples of programmable DNA binding proteins, for example, see Makarova et alia, "Classification and Nomenclature of CRISPR-Cas Systems: Where from Here?" CRISPR J. 2018 October; 1: 325-336. doi: 10.1089 / crispr.2018.0033; Yan et alia, "Functionally diverse type VCRISPR-Cas systems" Science. 2019 January 4; 363(6422): 88-91. doi: 10.1 126 / science.aav7271, the entire contents of each of which are hereby cited. In some embodiments, examples of programmable DNA binding proteins may include, but are not limited to, Cas12a / Cpf1 (UniProtKB: A0A7C9H0Z9), Cas12b / C2cl (UniProtKB: T0D7A2), Cas12c / C2c3 (UniProtKB: A0A9E2NRP4), Cas12d / CasY (UniProtKB: A0A9E2QZL7), Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Non-limiting examples of Cas enzymes include Cas12a / Cpf1, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12l, Cas12m, Cas12n, V-type Cas effector proteins, IscB proteins, TnpB proteins, Fanzor proteins, or modified or engineered versions thereof. In some embodiments, the programmable DNA binding protein may include Cas12a proteins, wherein the programmable DNA binding protein has DNA binding ability and nuclease activity. In some embodiments, the programmable DNA binding protein may include Cas12m proteins, wherein the programmable DNA binding protein has DNA binding ability and RNA endonuclease activity. Other nucleic acid programmable DNA binding proteins are also within the scope of the present disclosure, although they may not be specifically listed in the present disclosure.

[0100] In some embodiments, TALE base editors are disclosed, which may include a TALE domain, a deaminase domain and / or a cofactor protein (e.g., FokI endonuclease) domain, comprising a base editor having NH2-[TALE]-[deaminase domain]-COOH, NH2-[deaminase domain]-[TALE]-COOH, NH2-[TALE]-[deaminase domain]-[cofactor protein]-COOH, NH2-[cofactor protein]-[deaminase domain]-[TALE]-COOH, NH2-[cofactor protein]-[deaminase domain]-[TALE]-COOH, NH2-[cofactor protein]-[TALE]-[deaminase domain]-[TALE]-COOH, or NH2-[deaminase domain]-[TALE]-[cofactor protein]-COOH; wherein each instance of "]-[" includes an optional linker, such as a peptide linker.

[0101] In some embodiments, the disclosed methods involve transducing (e.g., by transfection) cells with multiple complexes, each complex comprising a fusion protein comprising a TAL effector domain and a deaminase domain and a cofactor protein, wherein each cofactor protein localizes the fusion protein to a different target sequence. See Yang L. et al., Engineering and optimizing deaminase fusions for genome editing, Nature Comms., 2016. In specific embodiments, the methods disclosed herein involve TAL effector domains that bind to target sites not by Watson-Crick hybridization, but by binding to the major groove of the DNA double helix. In certain embodiments, these methods involve transfecting nucleic acid constructs (e.g., plasmids), each of which (or together) encodes components of multiple complexes of TALE base editors, including a TALE domain and a deaminase domain, and a cofactor protein. In certain embodiments, the disclosed fusion protein comprises a cofactor protein domain—that is, the domain is incorporated into the fusion protein construct. In other embodiments, the TALE base editor comprises a TALE domain and a deaminase domain, and the cofactor protein is introduced into the cell separately from the base editor.

[0102] In certain embodiments of the disclosed methods, constructs encoding TALE base editors are transfected into cells separately from constructs encoding cofactor proteins. In certain embodiments, these components are encoded on a single construct and transfected together. In specific embodiments, these single constructs encoding TALE base editors and cofactor proteins can be iteratively transfected into cells, with each iteration being associated with a subset of target sequences. In specific embodiments, these single constructs can be transfected into cells within a few days. In other embodiments, they can be transfected into cells within a few weeks.

[0103] Guide polynucleotides

[0104] In some embodiments, the guide polynucleotide is a guide RNA. The RNA / Cas complex can help "guide" the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA endonucleolytically cleaves linear or circular dsDNA targets that are complementary to the spacer region. The target chain complementary to the crRNA is first subjected to endonucleolytic cleavage, followed by 3'-5' exonucleolytic cleavage. In nature, DNA binding and cutting typically require proteins and two RNAs. However, single guide RNA ("sgRNA" or simply "gRNA") can be engineered to integrate various aspects of crRNA and tracrRNA into a single RNA species. For example, see Jinek M. et alia, Science 337: 816-821 (2012), the entire contents of which are hereby cited. Cas9 recognizes short motifs (PAM or protospacer adjacent motifs) in CRISPR repeat sequences to help distinguish between self and non-self. The Cas9 nuclease sequence and structure are well known to those skilled in the art (e.g., see "Complete genome sequence of an M1 strain of Streptococcus pyogenes", Ferretti, J. et alia, Natl. Acad. Sci. USA 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III", Deltcheva E. et alia, Nature 471:602-607 (2011); and "Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity", Jinek M. et alia, Science 337:816-821 (2012), the entire contents of each of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus.Based on this disclosure, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems” (2013) RNA Biology 10:5,726-737; the entire contents of which are incorporated herein by reference. In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase. In some embodiments, the guide polynucleotide is at least a single guide RNA (“sgRNA” or “gRNA”). In some embodiments, the guide polynucleotide is at least one tracrRNA. In some embodiments, the guide polynucleotide does not require a PAM sequence to guide the polynucleotide programmable DNA binding domain (e.g., Cas9 or Cpfl) to the target nucleotide sequence. The polynucleotide programmable nucleotide binding domains (e.g., CRISPR-derived domains) of the base editors disclosed herein can identify target polynucleotide sequences by binding to guide polynucleotides. Guide polynucleotides (e.g., gRNA) are typically single-stranded and can be programmed to site-specifically bind (i.e., by complementary base pairing) to the target sequence of the polynucleotide, thereby guiding the base editor bound to the guide nucleic acid to the target sequence. The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. In some embodiments, the guide polynucleotide comprises natural nucleotides (e.g., adenosine). In some embodiments, the guide polynucleotide comprises non-natural (or non-natural) nucleotides (e.g., nucleic acid peptides or nucleotide analogs). In some embodiments, the targeting region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides in length. The targeting region of the guide nucleic acid can be between 10-30 nucleotides in length, or between 15-25 nucleotides in length, or between 15-20 nucleotides in length.

[0105]

C to T Edit

[0106] In some embodiments, the fusion protein provided herein comprises one or more nucleic acid editing domains. In some embodiments, the base editor disclosed herein comprises a fusion protein comprising a target cytidine (C) base capable of deaminating a polynucleotide to produce uridine (U), which has the base pairing properties of thymine. In some embodiments, for example, the polynucleotide is double-stranded (e.g., DNA), and then the uridine base can be replaced by a thymidine base (e.g., by a cell repair mechanism) to produce a C: G to T: A transition. In other embodiments, the deamination of C in a nucleic acid to U by a base editor may not be accompanied by substitution of U for T.

[0107] Deamination of target C in a polynucleotide to produce U is a non-limiting example of a base editing type, which can be performed by a base editor as described herein. In another example, a base editor comprising a cytidine deaminase domain can mediate the conversion of cytosine (C) base to guanine (G) base. For example, the U of a polynucleotide produced by the cytidine deaminase domain of a base editor can be removed from the polynucleotide by a base excision repair mechanism (e.g., by uracil DNA glycosylase (UDG) domain), thereby producing a baseless site. The non-basic site can then be replaced (e.g., by a base repair mechanism) with another base (e.g., C), for example, by a translesion polymerase. Although the core base relative to the non-basic site is typically replaced with C, other replacements (e.g., A, G, or T) may also occur.

[0108] Thus, in some embodiments, the base editors described herein comprise a deaminase or a deaminase domain (e.g., a cytidine deaminase domain) capable of deaminating a target C in a polynucleotide to U. In addition, as described below, base editors may include additional domains that facilitate conversion of the U produced by the deamination reaction to T or G. For example, a base editor comprising a cytidine deaminase domain may further comprise a uracil glycosylase inhibitor (UGI) domain to mediate the substitution of T for U, completing a C to T base editing event. In another example, a base editor may be incorporated into a translesion polymerase to improve the efficiency of C to G base editing, because the translesion polymerase may promote the incorporation of C opposite the base site (i.e., resulting in the incorporation of G at the base site, completing a C to G base editing event).

[0109] Base editors comprising cytidine deaminase as a domain can deaminate target C in any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. Typically, cytidine deaminase catalyzes the C nucleobase located in the single-stranded portion of the polynucleotide. In some embodiments, the entire polynucleotide comprising target C can be single-stranded. For example, a cytidine deaminase incorporated into a base editor can deaminate target C in a single-stranded RNA polynucleotide. In other embodiments, a base editor comprising a cytidine deaminase domain can act on a double-stranded polynucleotide, but the target C can be positioned in a portion of the polynucleotide that is in a single-stranded state during the deamination reaction. For example, in an embodiment comprising a Cas9 domain, several nucleotides can remain unpaired during the formation of the Cas9-gRNA-target DNA complex, thereby forming a Cas9 "R-loop complex". These unpaired nucleotides can form bubbles of single-stranded DNA, which can serve as substrates for single-stranded specific nucleotide deaminases (e.g., cytidine deaminase).

[0110] Details of C to T nucleobase editing proteins are described in International PCT Application No. PCT / US2016 / 058344 (WO2017 / 070632) and Komor, AC., et alia, “Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage” Nature 533, 420-424 (2016), the entire contents of which are hereby incorporated by reference into this article.

[0111] Cytidine deaminase

[0112] Fusion proteins disclosed herein can include cytidine deaminase. In some embodiments, the cytidine deaminase provided herein can deaminate cytosine or 5-methylcytosine to uracil or thymine. In some embodiments, the cytidine deaminase provided herein can deaminate cytosine in DNA. The cytidine deaminase can be derived from any suitable organism. In some embodiments, the cytidine deaminase is from prokaryotes. In some embodiments, the cytidine deaminase is from bacteria. In some embodiments, the cytidine deaminase is from mammals (e.g., humans).

[0113] In some embodiments, the cytidine deaminase is a naturally occurring (e.g., wild-type) cytidine deaminase, and in some embodiments, the cytidine deaminase is a naturally occurring cytidine deaminase that includes one or more mutations corresponding to any of the mutations provided herein. One of skill in the art will be able to identify corresponding residues in any homologous protein, e.g., by sequence alignment and determination of homologous residues. Thus, one of skill in the art will be able to generate a mutation in any naturally occurring cytidine deaminase that corresponds to any of the mutations described herein.

[0114] In some embodiments, the cytidine deaminase comprises an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 70%, at least 80%, at least 85%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identical to any of the cytidine deaminase amino acid sequences set forth in SEQ ID NOs: 18, 1-17, 19-24, 70-124, and 201-289. It should be understood that the cytidine deaminases provided herein can include one or more mutations (e.g., any mutation provided herein). The present disclosure provides any deaminase domain having a certain percentage of identity and any mutation described herein, or any combination thereof. In some embodiments, the cytidine deaminase comprises a sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to a reference sequence or any of the cytidine deaminases provided herein. In some embodiments, the cytidine deaminase comprises at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, or at least 170 identical consecutive amino acid residues compared to any amino acid sequence known in the art or described herein.

[0115] In some embodiments, the cytidine deaminase of the base editor may comprise all or part of the deaminase of the apolipoprotein B mRNA editing complex (APOBEC) family. APOBEC is an evolutionarily conserved family of cytidine deaminases. Members of this family are C to U editing enzymes. The N-terminal domain of the APOBEC-like protein is the catalytic domain, while the C-terminal domain is a pseudocatalytic domain. More specifically, the catalytic domain is a zinc-dependent cytidine deaminase domain, which is important for cytidine deaminase. Members of the APOBEC family include APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D ("APOBEC3E" now refers to this), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, and activation-induced (cytidine) deaminase (AID). In some embodiments, the deaminase incorporated into the base editor comprises all or part of the APOBEC1 deaminase. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC2. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC3. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC3A. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC3B. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC3C. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC3D. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC3E. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC3F. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC3G. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC3H. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the deaminase APOBEC4. In some embodiments, the deaminase incorporated into the base editor comprises all or part of the activation-induced deaminase (AID). In some embodiments, the deaminase incorporated into the base editor comprises all or part of the cytidine deaminase 1 (CDA1).

[0116] In some embodiments, the cytidine deaminase is (a) a cytidine deaminase from Burkholderia cenocepacia, (b) an SCP1.201-like deaminase from Streptacidiphilus jeojiense, (c) an Imm1 family immunoprotease from Brevilactibacter coleopterorum, (d) a DUF6531 domain-containing protease from Pseudomonas koroiiensis, (e) a PAAR domain protease from Yersinia similis, (f) a protease from Salmonella enterica. enterica), (g) a PAAR domain-containing protease from Pseudomonas, (h) an APOBEC-1 enzyme from the Burmese tree shrew (Tupaiachinensis), (i) an APOBEC-1-like enzyme from the small shrew (Suncus etruscus), (j) an APOBEC-1 enzyme from the Damaraland mole (Fukomys damarensis), (k) an APOBEC-3D-like enzyme from the striped hyena (Hyaena hyaena), (l) an APOBEC-1-like enzyme from the hunting falcon (Falco cherrug), (m) an APOBEC-1-like enzyme from the green sea turtle (Cheloniamydas), (n) an APOBEC-1 from the alligator snapping turtle (Chelydra serpentina), (o) an APOBEC-1 from the purple-winged starling (Sturnus vulgaris), (p) APOBEC-3C-like enzyme from Peruvian night monkey (Aotus nancymaae), (q) APOBEC-1 from Chinese hamster (Cricetulus griseus), (r) single-chain cytosine deaminase-like enzyme from Gulf sea dragon (Syngnathus scovel li), (s) ABEC1 enzyme from Philippine black cuckoo (Edolisoma coerulescens), (t) AICDA deaminase from black-tailed trogon (Trogon melanurus), (u) APOBEC-3A from North American pika (Ochotona princeps), (v) APOBEC-3A-like enzyme from Florida manatee (Trichechusmanatus latirostris), (w) APOBEC-3-like enzyme from sterlet (Acipenser ruthenus),(x) an APOBEC-3-like enzyme from the Pinta Island giant tortoise (Chelonoidis abingdonii), or a cytidine deaminase having an amino acid sequence at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to any one of (a) to (x). In some embodiments, the cytidine deaminase is an enzyme having an amino acid sequence at least identical to any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124, and 201-289 (see Table 1).

[0117] Certain aspects of the present disclosure are based on the recognition that regulating the catalytic activity of the deaminase domain of any fusion protein described herein, for example by making point mutations in the deaminase domain, can affect the synthetic ability of the fusion protein (e.g., base editor). For example, a mutation that reduces but does not eliminate the catalytic activity of the deaminase domain in a base editing fusion protein can reduce the likelihood that the deaminase domain catalyzes the deamination of residues adjacent to the target residue, thereby narrowing the deamination window. The ability to narrow the deamination window can prevent unnecessary deamination of residues adjacent to a specific target residue, thereby reducing or preventing off-target effects.

[0118] In some embodiments, the cytidine deaminase comprises one or more mutations at positions R57X, W114X, and H146X, numbered in SEQ ID NO: 8, or one or more corresponding mutations thereof, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of R57A, W114Y, and H146A, numbered in SEQ ID NO: 8, or one or more corresponding alterations thereof.

[0119] In some embodiments, the cytidine deaminase comprises one or more alterations at positions K34X, W89X, and H121X as numbered in SEQ ID NO: 9, or one or more corresponding alterations thereof, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations in the group consisting of K34A, W89Y, and H121A as numbered in SEQ ID NO: 9, or one or more corresponding alterations thereof.

[0120] In some embodiments, the cytidine deaminase comprises one or more alterations at positions R218X, W275X, and H307X, as numbered in SEQ ID NO: 11, or one or more corresponding alterations thereof, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of R218A, W275Y, and H307A, as numbered in SEQ ID NO: 11, or one or more corresponding alterations thereof.

[0121] In some embodiments, the cytidine deaminase comprises one or more alterations at positions R53X, W122X, and Y152X as numbered in SEQ ID NO: 16, or one or more corresponding alterations thereof, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of R53A, W122Y, and Y152F as numbered in SEQ ID NO: 16, or one or more corresponding alterations thereof.

[0122] In some embodiments, the cytidine deaminase comprises one or more mutations at positions K20X, R21X, Y22X, R26X, T28X, L30X, N46X, A52X, C73X, L75X, W77X, S78X, P79X, L87X, I101X, Y107X, and Y108X as numbered in SEQ ID NO: 18, or one or more corresponding alterations thereof, wherein X is any amino acid. In some embodiments, the cytidine deaminase comprises one or more alterations selected from the group consisting of K20A, R21A, Y22A, Y22F, R26A, T28A, L30A, N46A, A52D, C73F, C73W, L75A, W77A, W77F, S78A, P79A, L87A, I101A, Y107A, Y107F, Y108A, and Y108F, or one or more corresponding alterations thereof.

[0123] [Table 1: Amino acid sequence and base editor of cytidine deaminase]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153] The DNA sequences of exemplary mammalian expression plasmids for CBE are provided below. The DNA sequence of the BE4-rAPOBEC1 expression vector is provided in SEQ ID NO:49, with the deaminase enzyme underlined. For other constructs, only the deaminase sequence is shown, as the backbone BE4 vector sequence is identical. In some embodiments, a polynucleotide sequence encoding a deaminase enzyme provided by any of SEQ ID NOs:42, 25-41, 43-48, 125-179, or 297-385 is introduced into the vector backbone exemplified by SEQ ID NO:49, with the polynucleotide sequence provided by any of SEQ ID NOs:42, 25-41, 43-48, 125-179, or 297-385 replacing the underlined portion of SEQ ID NO:49 as follows:

[0154] DNA sequence encoding BE4-rAPOBEC1

[0155]

[0156]

[0157]

[0158]

[0159] [Table 2: Polynucleotide sequences encoding cytidine deaminase and base editors]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172]

[0173]

[0174]

[0175]

[0176]

[0177]

[0178]

[0179]

[0180]

[0181]

[0182]

[0183]

[0184]

[0185]

[0186]

[0187]

[0188]

[0189]

[0190]

[0191]

[0192]

[0193]

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201]

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225]

[0226]

[0227]

[0228]

[0229]

[0230]

[0231]

[0232]

[0233]

[0234]

[0235]

[0236]

[0237]

[0238]

[0239]

[0240]

[0241]

[0242] RNA editing

[0243] In some embodiments, the compositions disclosed herein comprise an RNA binding protein or a functional domain thereof. The RNA binding protein or its functional domain may comprise an RNA recognition motif. In some embodiments, the RNA binding protein is a Cas effector protein, such as a Cas13 effector protein. In some embodiments, the RNA binding protein is a catalytically inactive Cas effector protein, such as dCas13.

[0244] In some embodiments, the RNA binding protein comprises a zinc finger motif. In some embodiments, the RNA binding protein or its functional domain comprises a Cyst-Hist, Gag-knuckle, Treble-clet, zinc ribbon, or Zn2 / Cys6 motif.

[0245] In some embodiments, a method for the targeted deamination of cytidine in RNA (more specifically in an RNA sequence of interest) is provided. Cytidine deaminase (CD) protein is specifically recruited to the relevant cytidine in the target RNA sequence by a CRISPR-Cas complex that can specifically bind to the target sequence. To achieve this, the cytidine deaminase protein can be covalently linked to the CRISPR-Cas enzyme or provided as a separate protein, but adjusted to ensure that it is recruited into the CRISPR-Cas complex. In some embodiments, the recruitment of cytidine deaminase to the target locus is ensured by fusing cytidine deaminase or its catalytic domain to a CRISPR-Cas protein (which can be a Cas13 protein).

[0246] Methods for generating fusion proteins from two independent proteins are known in the art and may involve the use of spacers or linkers. The Cas13 protein may be fused to a cytidine deaminase protein or its catalytic domain at its N-terminus or C-terminus. In some embodiments, the CRISPR-Cas protein is an inactivated or dead Cas13 protein and is connected to the N-terminus of the deaminase protein or its catalytic domain. In some embodiments, the CD functionalized CRISPR system comprises (a) a cytidine deaminase fused or connected to a CRISPR-Cas protein, wherein the CRISPR-Cas protein is a catalytically inactive Cas13, and (b) a guide molecule comprising a guide sequence, optionally designed to introduce a C-AU mismatch in the RNA duplex formed between the guide sequence and the target sequence (A) upstream or downstream of the target cytosine or (B) between the guide sequence and the target sequence. In some embodiments, the CRISPR-Cas protein and / or the cytidine deaminase is NES-tagged at the N-terminus or C-terminus or both. In some embodiments, a CD-functionalized CRISPR system comprises (a) a CRISPR-Cas protein that is a catalytically inactive Cas13, (b) a guide molecule comprising a guide sequence, optionally designed to introduce a CA / U mismatch (A) upstream or downstream of the target cytosine or (B) in the RNA duplex formed between the guide sequence and the target sequence, and an aptamer sequence (e.g., an MS2 RNA motif or a PP7 RNA motif) capable of binding to an adapter protein (e.g., an MS2 coat protein or a PP7 coat protein), and (c) a cytidine deaminase fused or linked to the adapter protein, wherein binding of the aptamer and adapter protein recruits the cytidine deaminase to the RNA duplex formed between the guide sequence and the target sequence for targeted deamination, either at a C outside the target sequence or at a C with an optional CA / U mismatch. In some embodiments, the adapter protein and / or the cytidine deaminase are NES-tagged, either at the N-terminus or the C-terminus or both. The CRISPR-Cas protein can also be NES-tagged.

[0247] Split System

[0248] In some embodiments, the CBE disclosed herein comprises a "split" system in which the Cas protein, cytidine deaminase, or both are split into two components and can be reconstituted to form a functional protein to perform base editing. "Split Cas9 protein," "split Cas9," or "split cytidine deaminase" refers to a Cas9 or cytidine deaminase protein provided as an N-terminal fragment and a C-terminal fragment encoded by two independent nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein or cytidine deaminase can be spliced ​​to form a "recombinant" Cas9 protein or cytidine deaminase. In some embodiments, the Cas9 protein or cytidine deaminase is split into two fragments within a disordered region of the protein, for example, as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al., (2016) Science 351: 867-871. PDB file: 5F9R, each of which is incorporated herein by reference.

[0249]

Chimeric system

[0250] In some embodiments, the CBE system disclosed herein comprises a chimeric system in which a cytidine deaminase is inserted into Cas9, as shown in Example 16. In some embodiments, the present disclosure provides a fusion protein comprising a nucleic acid-guided nuclease protein (e.g., Cas9 or any nuclease disclosed herein) and a chimeric nucleotide deaminase protein domain. In some embodiments, the chimeric nucleotide deaminase protein domain is a chimeric cytidine deaminase protein domain. In some embodiments, the chimeric cytidine deaminase has an amino acid sequence that is at least 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to any one of SEQ ID NOs: 18, 1-17, 19-24, 70-124, and 201-289. [Example]

[0251] The methods and compositions are further described in the following examples, which do not limit the scope of the claims.

[0252] Example 1: Evaluation of on-target editing by cytidine base editors (CBEs)

[0253] To evaluate the efficiency of the novel cytidine base editors (CBEs) disclosed herein, CBEs were co-expressed with 8 unique single guide RNAs (sgRNAs) targeting 8 different loci (sgRNA1-sgRNA8 are provided in Table 11 below) in HEK293T cells. After editing, on-target editing efficiency was assessed by next-generation sequencing (NGS) of the target loci. BE4max CBE was included as a positive control. HEK293T cells (1.5×10 cells per experiment) were transfected 20-24 hours before transfection. 4 Cells were seeded in 96-well plates. According to the manufacturer's experimental protocol, cells were transfected with 80 ng of expression vector encoding CBE, 40 ng of spCas9-gRNA expression plasmid and 0.36 μL of FuGENE HD transfection reagent (E2312). After 48 hours, cells were collected, genomic DNA was extracted, and DNA was analyzed by NGS to assess editing efficiency. Table 3 below describes the C to T editing efficiency of each CBE using each of the 8 different sgRNAs. A subset of these results is plotted in Picture 1 In. Picture 1 As shown, the editing efficiencies of CE_2, CE_13, and CE_16 are comparable to those of BE4max. The average editing efficiency of CE_9 is higher than that of BE4max.

[0254] Table 3: C to T editing efficiency of CBE using 8 sgRNAs

[0255]

[0256]

[0257] To evaluate the editing window of the CBEs disclosed herein, the average C to T editing efficiency of 14 different sgRNAs at each position within the protospacer for each CBE was measured by calculating the percentage of NGS reads that included C to T edits at each position within the protospacer (Table 11). The data from these experiments are shown in Table 4 below and plotted in Picture 2A-2G In. With BE4max( Picture 2A ) compared to CE_1( Picture 2B ) and CE_12( Picture 2F ) has a narrower editing window, while CE_9( Picture 2E ) and CE_16( Picture 2G )'s editing window is wider than that of BE4max.

[0258] Table 4: Average C to T editing efficiency of 14 different sgRNAs at each position within the protospacer

[0259]

[0260] To assess the sequence bias of the CBEs disclosed herein, the effect of sequence context was measured by the percentage of reads (including C to T edits) at each position within the protospacer based on adjacent bases (i.e., whether the edited C was 3' to a T, C, A, or G). Table 5 lists the average editing efficiency data for "TC" sites. Table 6 lists the average editing efficiency data for "AC" sites. Table 7 lists the average editing efficiency data for "GC" sites. Table 8 lists the average editing efficiency data for "CC" sites. "NA" = No data available. A subset of these data is plotted in Picture 3A-3F CBE CE_1( Picture 3B ) and CE_12( Picture 3F ) is relatively inefficient at GC sites, similar to the pattern of BE4max ( Picture 3A ). CBE CE_2( Picture 3C )、CE_4( Photo 3D ) and CE_9( Picture 3E ) showed no significant sequence-context dependence.

[0261] Table 5: Average editing efficiency of “TC” sites at different positions within the protospacer

[0262]

[0263]

[0264] Table 6: Average editing efficiency of “AC” sites at different positions within the protospacer

[0265] 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 BE4max 10.88 NA 33.50 12.65 49.98 61.34 47.82 NA 12.80 NA 1.22 6.50 NA NA 0.15 0.13 0.03 NA 0.00 NA CE_1 5.37 NA 8.21 3.97 21.07 31.21 13.37 NA 5.43 NA 0.86 2.29 NA NA 0.14 0.07 0.01 NA 0.02 NA CE_2 8.15 NA 16.09 11.00 43.93 40.06 15.62 NA 6.42 NA 0.30 1.23 NA NA 0.14 0.04 0.03 NA 0.03 NA CE_3 4.00 NA 2.54 1.18 1.63 1.92 1.45 NA 0.90 NA 0.18 2.32 NA NA 0.02 0.03 0.02 NA 0.01 NA CE_4 28.99 NA 51.70 20.67 44.29 33.66 26.44 NA 19.42 NA 7.55 33.63 NA NA 0.49 0.82 0.06 NA 0.01 NA CE_5 4.52 NA 2.20 0.80 1.44 3.04 0.89 NA 0.58 NA 0.25 0.84 NA NA 0.02 0.04 0.02 NA 0.01 NA CE_6 4.42 NA 4.80 2.96 4.30 12.95 3.79 NA 2.10 NA 0.70 1.87 NA NA 0.05 0.02 0.02 NA 0.00 NA CE_7 4.72 NA 0.59 0.66 0.53 0.91 0.99 NA 0.34 NA 0.13 0.25 NA NA 0.05 0.05 0.01 NA 0.02 NA CE_8 6.19 NA 2.97 1.20 3.77 2.95 1.64 NA 1.08 NA 0.29 0.58 NA NA 0.02 0.02 0.01 NA 0.00 NA CE_9 49.27 NA 59.25 50.30 51.43 59.10 50.14 NA 43.77 NA 7.18 13.51 NA NA 0.47 0.11 0.03 NA 0.01 NA CE_12 4.57 NA 1.41 0.91 2.42 24.71 1.76 NA 1.01 NA 0.14 0.53 NA NA 0.04 0.02 0.02 NA 0.04 NA CE_13 48.27 NA 64.88 26.27 52.32 61.33 48.02 NA 40.25 NA 8.36 56.76 NA NA 1.13 0.97 0.30 NA 0.02 NA CE_16 32.69 NA 59.21 NA 42.34 46.34 40.41 NA 38.21 NA NA 83.14 NA NA NA NA 2.44 NA 0.01 NA Dd_3 5.74 NA 5.52 1.41 6.93 6.62 5.10 NA 2.47 NA 0.31 3.45 NA NA 0.05 0.04 0.02 NA 0.01 NA Dd_6 10.90 NA 1.66 NA 0.63 0.96 0.65 NA 0.53 NA NA 0.80 NA NA NA NA 0.01 NA 0.03 NA Ss_1 9.31 NA 1.56 NA 0.57 0.95 0.44 NA 0.30 NA NA 0.87 NA NA NA NA 0.02 NA 0.00 NA

[0266] Table 7: Average editing efficiency of "GC" sites at different positions within the protospacer

[0267]

[0268] Table 8: Average editing efficiency of “CC” sites at different positions within the original spacer

[0269]

[0270]

[0271] To assess the accuracy of the CBEs disclosed herein, the frequency of unexpected insertions and / or deletions (indels) and C to A or C to G substitutions was quantified in the NGS data after editing. The data for indel frequencies after editing are shown in Table 9. Picture 4AA subset of these data is shown, showing the percentage of NGS reads containing insertions and / or deletions (indels) after editing with each of several cytidine base editors compared to BE4max. Table 10 lists data on the ratio of editing efficiency to indel rate after editing with several cytidine base editors. A subset of these data is shown, showing the percentage of NGS reads containing insertions and / or deletions (indels) after editing with each of several cytidine base editors compared to BE4max. Picture 4B shown.

[0272]

Table 9: Frequency of insertions / deletions after editing

[0273]

[0274]

Table 10: Ratio of editing efficiency to insertion / deletion rate

[0275]

[0276]

[0277] Example 2: Evaluation of off-target editing of CBE

[0278] To evaluate the CBE-mediated, non-guided single-stranded DNA deamination disclosed herein, an R-loop assay was performed. The R-loop assay is a high-throughput method for evaluating sgRNA-independent off-target effects of CBEs that utilizes orthogonal R-loops generated by the SaCas9 nickase to mimic actively transcribed genomic loci that are more susceptible to cytidine deaminases. The SaCas9 nickase was fused to two copies of a uracil glycosylase inhibitor (UGI) to amplify the potential for off-target effects, which were then detected by the assay. As shown in this example, sgRNA 10, provided in Table 11, was used to guide on-target editing, and six different saCas9 gRNAs, provided in Table 12, were used to mimic off-target editing.

[0279] [Table 11: sgRNAs with C nucleotide positions in target genes]

[0280] sgRNA name hierarchy SEQ ID NO fundamental cause sgRNA1 TGCCCCTCCC TCCCTGGCCC 50 EMX1 sgRNA2 AGAGCCCCCC CTCAAAGAGA 51 DNMT3B sgRNA3 CACACACACT TAGAATCTGT 52 basic cause sgRNA4 GGCACTGCGG CTGGAGGTGG 53 basic cause sgRNA5 GGAATCCCTT CTGCAGCACC 54 FANCF sgRNA6 GAGTCCGAGC AGAAGAAGAA 55 EMX1 sgRNA7 GACCCCCTCC ACCCCGCCTC 56 VEGFA sgRNA8 CCCCACCGTT GAAGAACCAG 57 CEACAM16 sgRNA9 GCAGAGAGTC GCCGTCTCCA 58 basic cause sgRNA10 CGGAATCCTG GCTGGGAGCT 59 PCSK9 sgRNA11 TAACGGAACC CCCGGACTGG 60 PCSK9 sgRNA12 TCATCCGCCC GGTACCGTGG 61 PCSK9 sgRNA13 TACCCCTCCA CGGTACCGGG 62 PCSK9 sgRNA14 CCCGCACCTT GGCGCAGCGG 63 PCSK9

[0281] For R-loop experiments, 48 ​​ng of expression vector encoding each CBE, 32 ng of SpCas9 guide RNA plasmid (sgRNA10 provided in Table 11), 48 ng of nSaCas9-2xUGI plasmid, and 32 ng of SaCas9 guide RNA plasmid were co-transfected into HEK293T cells in 96-well plates. Two days later, off-target deamination at the six dSaCas9 loci was detected using NGS analysis. The results of the CBEs disclosed herein were compared with those of spCas9 gRNA BE4max with YE1, SECURE (R33A+K34A), SsAPOBEC3B, and PpAPOBEC1-H122A, which are well-studied rAPOBEC1 variants or APOBEC orthologs with reduced off-target effects relative to BE4max.

[0282] The results of orthogonal R-ring determination are as follows Picture 5 As shown. The x-axis represents the average off-target editing mediated by saCas9 gRNA. The y-axis represents the on-target editing mediated by sgRNA10. The results of the CBEs disclosed herein are plotted as triangles. BE4max is plotted as circles. The previously reported CBE (BE4max with YE1, SECURE(R33A+K34A), SsAPOBEC3B, and PpAPOBEC1-H122A) is plotted as squares. Several CBEs disclosed herein exhibit reduced off-target editing relative to BE4max while retaining substantial on-target activity as measured by the R-loop assay.

[0283] Table 12: saCas9 gRNAs used in off-target experiments

[0284]

[0285]

[0286] Example 3: Evaluating Off-Target Editing of Engineered CBEs

[0287] To evaluate guideless single-stranded DNA deamination mediated by engineered variants of the base editors disclosed herein, an R-loop assay was performed in the same manner as described in Example 2 above. For CBE CE_1, variants R57A, W114Y, and H146A were tested using the R-loop assay along with unmodified CE_1. For CBE CE_2, variants K34A, W89Y, H121A, and unmodified CE_2 were tested. For CBE CE_4, variants R218A, W275Y, H307A, and unmodified CE_4 were tested. For CBE CE_9, variants R53A, W122Y, Y152F, R53A+Y152F, and unmodified CE_9 were tested. The on-target editing efficiency of the three "C" positions in the sgRNA10 guide RNA (C7, C8, and C12) was measured. Off-target editing was measured for three sgRNAs (saRNA1, saRNA2, and saRNA3) as shown in Table 12. Off-target and on-target editing were measured for CBE and its variants relative to BE4max and BE4max_YE1.

[0288] Picture 6 Results of the R-loop assay for engineered CBE variants are provided in . CE_1R57A, CE_1W114Y, CE_1H146A, CE_2K34A, CE_2W89Y, CE_9W122Y, CE_9Y152F, and CE_9R53A+Y152F all exhibited off-target editing below BE4max for one or more of the sgRNAs tested.

[0289] Example 4: Splitting or Co-expressing Cytidine Base Editors

[0290] To utilize the cytidine deaminase disclosed herein in the context of a split CRISPR / Cas system, the cytidine deaminase was introduced into a split intein-based vector that cleaves the required CRISPR / Cas9 components in the dual-vector system. The CRISPR / Cas and deaminase were reconstituted into active forms upon co-infection to achieve targeted genome editing in cells. In one series of experiments, an RNA aptamer was included on the polynucleotide sequence encoding the sgRNA to recruit the aptamer-binding protein-fused deaminase. In a second series of experiments, Cas9 was fused to a GCN4 tag to recruit the deaminase fused to a scFv. The scFv recognized and was recruited to GCN4. In a third series of experiments, Cas9 was fused to a C-intein to recruit the deaminase fused to an N-intein. In each type of experiment, an intermediate mechanism was utilized to recruit the deaminase to the Cas domain to reconstitute the CBE.

[0291] Example 5: Cytidine Base Editors for RNA Editing

[0292] To utilize the cytidine deaminases disclosed herein in the context of RNA editing, each cytidine deaminase was introduced into a CRISPR / Cas13-based RNA editing system. Following RNA editing, the on-target and off-target RNA editing efficiency of each CBE disclosed herein was quantified by NGS.

[0293] Example 6: Evaluation of Highly Efficient Cytidine Base Editors (CBEs)

[0294] To evaluate the efficiency of the cytidine base editors (CBEs) disclosed herein, CBEs were co-expressed in HEK293T cells with 6 unique single guide RNAs (sgRNAs) targeting 6 different loci provided in Table 13 below. After editing, on-target editing efficiency was assessed by next generation sequencing (NGS) of the target loci. BE4max CBE was included as a positive control. HEK293T cells (1.5×10 per experiment) were transfected 20-24 hours before transfection. 4 cells) were seeded in 96-well plates. According to the manufacturer's experimental protocol, the cells were transfected with 80 ng of expression vector encoding CBE, 40 ng of spCas9-gRNA expression plasmid and 0.36 μL of FuGENE HD transfection reagent (E2312). After 48 hours, the cells were collected, genomic DNA was extracted, and the DNA was analyzed by NGS to evaluate the editing efficiency. Table 14 below describes the C to T editing efficiency of each CBE using each of the 6 different sgRNAs. All of these CBEs disclosed in this article showed an editing efficiency of more than 30% at at least 1 target site. Some of them showed editing efficiencies comparable to or higher than BE4max, and the results of the best performance are shown in Table 14. Picture 7 shown.

[0295] [Table 13: spCas9 sgRNA sequence]

[0296]

[0297]

[0298] Table 14: C to T editing efficiency of CBE using different sgRNAs

[0299]

[0300]

[0301] BE4max and sgRNA15(GAAG GCTTTA CTGTATTACA GA) has poor activity, which has GC dinucleotides in the classical editing window 4-8 (PAM definition 21-23). ​​Some CBEs disclosed herein do not have this sequence context restriction (see Table 14). The best performance is as follows Picture 8 The editing efficiency of GC dinucleotides in sgRNA4 also supports the same conclusion (see Table 15).

[0302] [Table 15: C to T editing efficiency of CBE using sgRNA4]

[0303]

[0304]

[0305]

[0306] Compared to commonly used base editing tools, some CBEs disclosed herein exhibit wider editing windows (see, e.g., Tables 16-17 and Picture 9A-9B ), which significantly increases the target range and expands the application of base editing in genome editing.

[0307] Table 16: C to T editing efficiency of sgRNA5 at different positions

[0308]

[0309]

[0310]

[0311] Table 17: C to T editing efficiency of sgRNA14 at different positions

[0312]

[0313]

[0314] Example 7: Evaluation of CBE with high specificity

[0315] To determine the precision of the editing window of the CBEs disclosed herein, a precision ratio was estimated for each CBE. The precision ratio was calculated by dividing the sum of edits per editable nucleotide within the canonical 4-8 base pair window by the sum of edits per editable nucleotide within the entire protospacer region. Editors with lower precision are expected to induce more editing events outside the canonical window and are therefore estimated to have lower precision ratios. The precision ratio results for the subset of CBEs disclosed herein are shown in Picture 10and Table 18. They show a higher precision ratio than BE4max, indicating that they can induce less bystander editing.

[0316] Table 18: Editing accuracy of CBE relative to BE4max

[0317]

[0318]

[0319]

[0320] Cytosine base editors like BE4max have been found to exhibit off-target effects, including guide-dependent DNA off-target, guide-independent DNA off-target, and RNA off-target. These off-target effects may limit their application, especially in therapeutic settings. To evaluate the non-guided single-stranded DNA deamination mediated by the CBEs disclosed herein, an R-loop assay was performed. The R-loop assay is a high-throughput method for evaluating the sgRNA-independent off-target effects of CBEs that utilizes orthogonal R-loops generated by the SaCas9 nickase to mimic actively transcribed genomic loci that are more susceptible to cytidine deaminases. The SaCas9 nickase was fused to two copies of a uracil glycosylase inhibitor (UGI) to amplify the possibility of off-target effects, which were then detected by the assay. As shown in this example, sgRNA10 provided in Table 13 was used to guide on-target editing, and two different saCas9 gRNAs provided in Table 19 were used to simulate off-target editing.

[0321] [Table 19: saCas9 gRNA used in off-target experiments]

[0322] R placement point sgRNA name hierarchy SEQ ID NO fundamental cause R rank 1 saRNA5 GCCACAGACT TTTCCATTTG C 181 EMX1 R-loop site 2 saRNA6 GCCCAGCTCC AGCCTCTGAT G 182 EMX1

[0323] For the R-loop experiment, 48 ng of expression vector encoding each CBE, 32 ng of SpCas9 guide RNA plasmid, 48 ng of nSaCas9-2xUGI plasmid, and 32 ng of SaCas9 guide RNA plasmid were co-transfected into HEK293T cells in 96-well plates. Two days later, NGS analysis was used to detect off-target deamination at the three nSaCas9 loci. Compared with BE4max, some base editors showed lower off-target editing. The results are as follows Figure 11 and as shown in Table 20.

[0324] [Table 20: R-loop determination of CBE variants relative to BE4max]

[0325]

[0326]

[0327]

[0328] The predicted guide-dependent off-target sites (OT) of sgRNAs 4-6 listed in Table 21 were measured by NGS. The off-target editing efficiency was normalized to the on-target editing efficiency. Compared with BE4max, a portion of these base editors showed lower off-target properties, as shown in Tables 22 and Figure 12 shown.

[0329] [Table 21: sgRNA guide-dependent OTs 4-6]

[0330] OT name OT sequence SEQ ID NO sgRNA4-OT1 TGCACTGCGG CCGGAGGAGG TGG 183 sgRNA4-OT2 GGCTCTGCGG CTGGAGGGGG TGG 184 sgRNA4-OT4 GGCATCACGG CTGGAGGTGG AGG 185 sgRNA6-OT GAGTCTAAGC AGAAGAAGAA GAG 186 sgRNA5-OT GGAACCCCGT CTGCAGCACC AGG 187

[0331] Table 22: CBE-dependent OT C to T editing efficiency relative to BE4max

[0332]

[0333]

[0334] RNA-seq was used to test the RNA off-target editing of CBE. HEK293T cells (1.2 × 10 cells per experiment) were cultured 20–24 h before transfection. 5 Cells were seeded in 12-well plates. According to the manufacturer's protocol, cells were transfected with 600 ng of expression vector encoding CBE, 200 ng of spCas9-gRNA expression plasmid, and 2.4 μL of FuGENE HD transfection reagent (E2312). After 48 hours, cells were harvested and RNA was extracted for poly(A) mRNA isolation and library preparation. The results of RNA off-target C to U editing are shown in Tables 23 and Figure 13 The base editors shown here result in low off-target RNA editing close to background levels.

[0335] Table 23: Number of C to U edited RNA SNVs in CBE variants relative to BE4max

[0336]

[0337]

[0338] Example 8: Further engineering of cytosine base editors leads to higher specificity

[0339] Compared to BE4max, CE_1 and CE_2 are CBEs with higher precision and lower off-target editing. To further improve the specificity of CBEs, mutations were introduced into CBEs by rational design. On-target and R-loop assays were performed in the same manner as in Examples 6 and 7 above. For the targeting assay, sgRNA 1, sgRNA4, and sgRNA5 were used. CE_1R57A, CE_1W114Y, and CE_1H146A showed higher precision ratios than unmodified CE_1. CE_2W89Y showed higher precision ratios than unmodified CE_2 (see Tables 24 and 25). Figure 14 ).

[0340] Table 24: Editing accuracy of CBE variants relative to BE4max

[0341] sgRNA4 sgRNA5 sgRNA1 CE_1 1.32 1.25 1.1 CE_1vR57A 1.64 1.33 1.36 CE_1_W114Y 1.5 1.33 1.34 CE_1_H146A 1.51 1.32 1.36 CE_2 1.42 1.16 1.17 CE_2_W89Y 1.7 1.3 1.43

[0342] The results of the R-loop assay are shown in Tables 25 and Figure 15 As shown. CE_1 R57A, CE_1 W114Y, and CE_1 H146A showed lower off-target editing than unmodified CE_1. Similarly, CE_2 K34A showed lower off-target editing compared to unmodified CE_2. CE_9 is a highly efficient cytosine base editor with higher off-target activity than BE4max. However, off-target effects were significantly reduced in the CE_9 W122Y and CE_9 Y152F variants.

[0343] Table 25: Editing accuracy of CBE variants relative to BE4max

[0344]

[0345] [Example 9: Designing CBE CE_12 to further improve effectiveness]

[0346] The CBE variant CE_12, with its specificity and small size, shows significant advantages in AAV-mediated delivery. To further enhance its on-target editing ability, a series of mutations were introduced into CE_12 by rational design. The on-target editing experiment was performed in the same manner as described in Example 6. sgRNA4 (SEQ ID NO: 53), sgRNA5 (SEQ ID NO: 54) and sgRNA6 (SEQ ID NO: 55) were used. Figure 16As shown, variants with single mutations such as S42N, L129Y, R20H, R7H, and K101R, as well as variants including CE_12_C41 (R20H+S42N+T73W+F103Y) and CE_12_C42 (R20H+S42N+T73F+F103Y), improved the editing efficiency of the tested sgRNA sites relative to the CE_12 variant. Among all the CBE variants tested, CE_12_C41 and CE_12_C42 had the most significant effects on editing efficiency.

[0347] [Example 10: Evaluation of CBE activity in the form of RNA in Huh7 cells]

[0348] To further evaluate the base editing activity of the CBE variants, mRNA for selected CBE editors was produced and combined with selected sgRNAs to characterize the editing activity. Table 26 provides the spacer sequences of the sgRNAs. In the experiment, Huh7 cells were cultured in DMEM medium supplemented with 10% fetal bovine serum. 24 hours before transfection, cells were seeded in 96-well plates at a density of 8,000 cells / well. For each sample, 100 ng of base editor mRNA and 100 ng of end-modified synthetic sgRNA were transfected using Lipofectamine MessengerMAX (ThermoFisher, Cat. LMRNA003) according to the manufacturer's protocol. After 72 hours, genomic DNA was harvested from the cells and the editing efficiency was analyzed by amplicon sequencing using NGS. The editing efficiency of each combination was calculated and is shown in Table 27. Among all the variants tested, CE_11, CE_43, CE_115, and CE_108 showed higher average on-target activity at the selected loci than BE4max.

[0349] Table 26: sgRNAs (RNA components) used for characterization of on-target activity in Huh7 cells

[0350] gRNA Spacer sequence SEQ ID NO gRNA_h1 TCCTCAGAGG AGAAAATCTG 188 gRNA_h2 CTCCACACCG CGCTGTAATT 189 gRNA_h3 TCTAAATCAG GTGAGACTGC 190 gRNA_h4 CACTCACCCA AAAATGTCCT 191 gRNA_h5 ATACTTACCA ATATGGGATG 192 gRNA_h6 ATTCAATTTG AAGCAGTGGT 193 gRNA_PCSK9 CCCGCACCTT GGCGCAGCGG 194

[0351] [Table 27: C to T editing efficiency of CBE in Huh7 cells (RNA component)]

[0352] Editor gRNA_h1 gRNA_h2 gRNA_h3 gRNA_h4 gRNA_h5 gRNA_h6 average value CE_1 53.7 59.17 37.2 60.88 27.58 20.69 43.20 CE_2 46.95 52.33 38.37 62.77 32.27 48.07 46.79 CE_4 35.23 45.89 34.1 39.05 54.63 25.41 39.05 CE_11 79.84 88.52 78.83 87.05 87.7 85.27 84.54 CE_108 52.92 77.47 3.88 68.16 74.67 58.53 55.94 CE_115 59.29 74.22 8.65 62.54 69.87 61.06 55.94 CE_125 25.09 53.55 11.22 37.82 45.76 8.32 30.29 CE_43 44.4 84.8 76.14 70.25 81.06 78.38 72.51 CE_82 39.6 16.54 31.33 71.34 47.09 48.72 42.44 BE4max 50.58 25.03 23.57 74.07 55.78 48.77 46.30

[0353] Example 11: Evaluation of CBE activity delivered as RNA in primary human hepatocytes (PHH)

[0354] The editing activity of the selected base editors and sgRNAs was further evaluated in primary human hepatocytes (PHH). PHHs were cultured according to the manufacturer's protocol. The cells were thawed and resuspended in hepatocyte thawing medium containing supplements and then centrifuged at 100g for 10 minutes. The supernatant was discarded and the precipitated cells were resuspended in hepatocyte seeding medium and supplements. The cells were counted and plated on 48-well plates coated with Bio-coat Collagen I (ThermoFisher, Cat. No. 877272) at a density of 132,000 cells / well. The seeded cells were allowed to settle and adhere for 4-6 hours in a tissue culture incubator at 37°C and 5% CO2 atmosphere, and then each well was transfected with 250ng base editor mRNA and 250ng end-modified sgRNA using RNAiMax according to the manufacturer's protocol. After 24 hours, the culture medium was replaced with fresh culture medium. 72 h after transfection, genomic DNA and RNA were extracted from the cells and analyzed for editing efficiency by amplicon sequencing using NGS. Table 28 shows the editing efficiency achieved by each combination in two PHH donors. In this experiment, the CBE variant CE_11 showed higher on-target activity relative to BE4max.

[0355] Table 28: C to T editing efficiency of CBEs delivered as RNA in PHH

[0356] Base editors PHH donors gRNA_h1 gRNA_h2 gRNA_h3 gRNA_h4 gRNA_h5 gRNA_h6 CE_11 Donor 1 58.86 65.54 N / A 73.53 66.83 N / A BE4max Donor 1 39.95 24.51 13.05 41.56 17.62 5.24 CE_11 Donor 2 29.92 33.6 19.72 38.34 33.11 19.36 BE4max Donor 2 18.18 15.93 6.53 17.13 6.49 1.83

[0357] [Example 12: Evaluation of CBE activity in a humanized mouse model]

[0358] The editing effect of the selected base editors was further evaluated using a humanized PCSK9 mouse model. In the experiment, a guide RNA (gRNA_PCSK9 is shown in Table 26) was specially designed to disrupt the human PCSK9 gene. According to the method outlined in patent application WO / 2023 / 185697A2, the mRNA and modified gRNA were combined and encapsulated into lipid nanoparticles (LNPs). These formulated LNPs were then administered to humanized mice by tail vein injection at different doses (0.3, 1, or 3 mg / kg).

[0359] After three weeks, liver and blood samples were collected from the mice for analysis. DNA editing in the liver and PCSK9 protein levels in the blood were measured. The results showed that across all doses, mice given CE_11 showed significantly improved editing efficiency and protein knockdown compared to mice given BE4max (see Table 29 for details).

[0360] Table 29: Liver editing and protein reduction levels achieved by CBE in a humanized PCSK9 mouse model

[0361] Editor Dosage (mg / kg) edit% Protein reduction% BE4max 0.3 N / A N / A CE_11 0.3 19.87 43.40 BE4max 1 8.20 49.93 CE_11 1 61.61 96.09 BE4max 3 37.45 79.31 CE_11 3 64.94 97.91

[0362] [Example 13: CBE variant CE_11 shows strong editing efficiency in combination with other programmable DNA binding proteins]

[0363] In order to evaluate the potential of CE_11 deaminase with other programmable DNA binding domains, dead LbCpf1, ISAam1-D363A (TnpB) and enIscB-D61A (IscB) were fused to CE_11 deaminase (SEQ ID NO: 210-212). After transfection with CE_11 fusion base editor (80 ng) and gRNA (40 ng) expression plasmid (gRNA sequence in Table 30), genomic DNA was collected from these cells after 72 hours of incubation period and used for NGS analysis to determine DNA editing efficiency as described above. The results are shown in Table 31.

[0364] [Table 30: sgRNAs used for on-target activity characterization]

[0365] sgRNA sequence SEQ ID NO target gene sgRNA21 Ggctggcagg ctccttgttc 195 ISAam1-PGK1 sgRNA22 Gcaaatcagc atgttcctca 196 ISAam1-AGBL1-1 sgRNA24 GCGTAGGTCT CCCTCC 197 EMX1-S1 sgRNA30 TGACTCACAC TCACCCAAAA 198 HSD17B13 sgRNA31 CCCAGAGCAT CCCGTGGAAC 199 PCSK9 sgRNA32 CCTCACTCCT GCTCGGTGAA 200 DNMT1

[0366] Table 31: C to T editing efficiency of CE_11 in combination with different programmable DNA binding proteins

[0367]

[0368]

[0369] [Example 14: Further design of CBE variant CE_11 to improve specificity]

[0370] In order to alleviate the off-target effects of base editors (including CE_11 deaminase), targeted mutations were introduced into CE_11 deaminase by rational design. These new variants were subjected to on-target and R-loop assays in the same manner as described in Examples 6 and 7. As shown in Table 32, A52D, W77A, and Y107A mutations reduced the editing efficiency of CE_11. Compared with the original CE_11 enzyme, other variants showed comparable or even higher on-target activity. The predicted guide-dependent off-target sites (OT) of sgRNA 4 and 6 listed in Table 21 were measured for these variants by NGS. The off-target editing efficiency of each sample was normalized to the on-target editing efficiency and listed in Table 33. Compared with unmodified CE_11, some variants in this subset showed significant advantages by significantly reducing off-target editing.

[0371] [Table 32: On-target editing of engineered CE_11 variants relative to CE_11]

[0372] CE_11 changes sgRNA2 sgRNA4 sgRNA6 CE_11_A52D 0.02 0.02 0.02 CE_11_C73F 1.24 1.27 1.17 CE_11_C73W 1.26 1.28 1.18 CE_11_I101A 0.88 0.98 0.92 CE_11_K20A 0.99 1.11 1.04 CE_11_L30A 1.02 1.06 1.00 CE_11_L75A 1.17 1.15 1.09 CE_11_L87A 1.33 1.43 1.25 CE_11_N46A 1.12 1.03 1.07 CE_11_P79A 0.94 1.37 1.08 CE_11_R21A 1.01 1.05 1.01 CE_11_R26A 1.32 1.07 1.18 CE_11_S78A 0.90 0.99 0.94 CE_11_T28A 1.12 1.17 1.09 CE_11_W77A 0.10 0.01 0.08 CE_11_W77F 1.24 0.61 1.12 CE_11_Y107A 0.36 0.02 1.14 CE_11_Y107F 0.89 1.08 1.05 CE_11_Y108A 1.34 0.77 1.21 CE_11_Y108F 1.11 1.09 1.07 CE_11_Y22A 1.29 1.15 1.21 CE_11_Y22F 1.16 1.16 1.11

[0373] [Table 33: C to T editing efficiency of engineered CE_11 variants relative to guide-dependent off-target editing of CE_11]

[0374]

[0375]

[0376] In the R-loop assay, several variants also showed significant reductions in off-target editing compared to the original CE_11, as shown in Table 34. Specifically, variants including C73W, R26A, N46A, W77F, P79A, I101A, and Y108A showed lower off-target editing at all three R-loop sites tested compared to CE_11.

[0377] [Table 34: R-loop assay activity of engineered CE_11 variants relative to CE_11]

[0378]

[0379]

[0380] To further reduce off-target editing of CE_11, combinatorial mutations were introduced into the deaminase, as shown in Table 35. On-target and R-loop assays were performed in the same manner as described in Examples 6 and 7. The on-target editing efficiencies of these CE_11 combinatorial variants when used with sgRNA4, sgRNA5, and sgRNA6 are shown in Table 36.

[0381] [Table 35: Overview of engineered CE_11 variants containing combined mutations]

[0382]

[0383]

[0384] [Table 36: On-target editing of engineered CE_11 variants containing combined mutations relative to CE_11]

[0385]

[0386]

[0387] The guide-dependent off-target editing results of these CE_11 combination variants are shown in Table 37. The R-loop editing results of these CE_11 combination variants are shown in Table 38. Compared to unmodified CE_11, several combination variants showed a reduction in off-target editing. Compared to CE_11, variants 9_WY, 10_PY, 31_KT, 35_TL, 38_LI, and 39_LC showed the most significant effect in minimizing off-target editing.

[0388] [Table 37: C to T editing efficiency of engineered CE_11 variants relative to guide-dependent off-target editing of CE_11]

[0389]

[0390]

[0391] [Table 38: Ratio of R-loop editing to on-target editing of engineered CE_11 variants relative to CE_11]

[0392]

[0393]

[0394] Example 15: Evaluation of engineered CE_11 variants delivered as RNA

[0395] To verify the efficacy and specificity of the CE_11 variant, the variant was transcribed into mRNA in vitro and co-transfected into Huh7 cells with synthetic sgRNA of equivalent mass. Huh7 cells (8×10 cells per well) were plated 20–24 h before transfection. 3 Cells) were seeded in 96-well plates. According to the manufacturer's protocol, lipofectamine MessengerMax reagent (LMRNA015) was used for transfection with different doses of mRNA and sgRNA4, as shown in Table 11. After 48 hours, cells were collected, genomic DNA was extracted, and the editing efficiency was analyzed by NGS. The on-target editing results of these variants are shown in Table 39. The off-target editing results of these variants are shown in Table 40. [Table 39: C to T editing efficiency of sgRNA4 with engineered CE_11 variants]

[0396]

[0397]

[0398] [Table 40: Ratio of C to T on-target editing and guide-dependent off-target editing of engineered CE_11 variants]

[0399]

[0400] [Example 16: Evaluation of Editing Activity and Specificity of CE_11 Chimeric Variants]

[0401] To test whether deaminase chimerism can improve the specificity of CE_11 deaminase in the context of base editing, where the deaminase is chimeric at a certain position within the nuclease protein, CE_11 and its variants are inserted into nCas9 at different sites (I1029, R535 and P1249) via 20 amino acid linkers. The on-target and R-loop determinations of these chimeric variants were performed in the same manner as described in Examples 6 and 7. The results of these experiments are shown in Tables 41 and 42. In particular, P1249_CE_11_R26A showed comparable editing efficiency compared to the original CE_11, while off-target editing was lower.

[0402] [Table 41: C to T editing efficiency of CE_11 chimeric variants relative to CE_11]

[0403] CE_11 changes sgRNA4 sgRNA5 sgRNA6 CE_11 1.00 1.00 1.00 I1029_CE_11 0.77 0.42 0.71 R535_CE_11_R26A 0.31 0.34 0.33 I1029_CE_11_R26A 0.83 0.67 0.80 P1249_CE_11_R26A 0.98 1.22 0.90 R535_CE_11_35_TL 0.52 0.27 0.35 I1029_CE_11_35_TL 0.84 0.52 0.69 P1249_CE_11_35_TL 0.84 0.88 0.69

[0404] [Table 42: Ratio of R-loop editing to on-target editing of engineered CE_11 variants relative to CE_11]

[0405]

[0406] [Example 17: Evaluation of the activity and specificity of CE_11 split variants]

[0407] To generate the split version of CE_11, the initial step involved inserting a deaminase between G1247 and S1248 of nCas9. Subsequently, based on the predicted structure, CE_11 was cleaved at Q148 and S149 to generate the split variant Split-CE_11. N and Split-CE_11 C The mRNAs of the two split components were then transcribed. The mRNAs of split CE_11 and unmodified CE_11 were transfected into Huh7 cells at different doses in a similar manner as described in Example 15. The on-target and off-target editing efficiencies of these split variants are shown in Table 11. Figure 17 As shown in Figure 3, split-CE_11 showed comparable or higher on-target editing efficiency compared to unmodified CE_11 at all doses tested. Split-CE_11 also reduced guide-dependent off-target (OT) effects.

[0408] [Other implementation methods]

[0409] It should be understood that although the invention has been described in conjunction with its detailed description, the foregoing description is intended to illustrate and not limit the scope of the invention, which is defined by the scope of the appended claims. Other aspects, advantages and modifications are within the scope of the following claims.

Claims

1. A cytidine base editor, comprising: (i) a polynucleotide programmable DNA binding domain; and (ii) a cytidine deaminase, wherein the cytidine deaminase consists of an amino acid sequence selected from any one of SEQ ID NO: 18 and 228-240, 242, 244-289.

2. The cytidine base editor according to claim 1, further comprising one or more uracil glycosylase inhibitor (UGI) domains.

3. The cytidine base editor according to claim 2, wherein the cytidine base editor comprises 2 glycosylase inhibitor (UGI) domains.

4. The cytidine base editor according to claim 1, further comprising one or more nuclear localization signals (NLS).

5. The cytidine base editor according to claim 4, wherein the cytidine base editor comprises an N-terminal NLS and / or a C-terminal NLS.

6. The cytidine base editor according to claim 4, wherein the NLS is a bipartite NLS.

7. The cytidine base editor according to claim 1, wherein the cytidine base editor has an increased cis / trans activity ratio compared to the BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by the polynucleotide shown in SEQ ID NO:

49.

8. The cytidine base editor according to claim 7, wherein the increased cis / trans activity ratio is at least 2-fold higher than that of the BE4max cytidine base editor.

9. The cytidine base editor according to claim 1, wherein the cytidine base editor has a 5%, 10%, 20%, 30%, 40% or 50% increased editing efficiency at GC sites compared to the BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by the polynucleotide shown in SEQ ID NO:

49.

10. The cytidine base editor according to claim 1, wherein the cytidine base editor has a 5%, 10%, 20%, 30%, 40% or 50% increased editing efficiency at AC sites compared to the BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by the polynucleotide shown in SEQ ID NO:

49.

11. The cytidine base editor according to claim 1, wherein the cytidine base editor has a 5%, 10%, 20%, 30%, 40% or 50% increased editing efficiency at TC sites compared to the BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by the polynucleotide shown in SEQ ID NO:

49.

12. The cytidine base editor according to claim 1, wherein the cytidine base editor has a 5%, 10%, 20%, 30%, 40% or 50% increased editing efficiency at CC sites compared to the BE4max cytidine base editor, wherein the BE4max cytidine base editor is encoded by the polynucleotide shown in SEQ ID NO:

49.

13. The cytidine base editor according to any one of claims 1 to 12, wherein the cytidine base editor has at least 5% cis activity compared to the BE4max cytidine base editor.

14. The cytidine base editor according to any one of claims 1 to 12, wherein the cytidine base editor has at least 5% trans activity compared to the BE4max cytidine base editor.

15. A cell comprising the cytidine base editor according to any one of claims 1 to 14.

16. The cell according to claim 15, wherein the cell is a bacterial cell, a plant cell, an insect cell or a mammalian cell.

17. A molecular complex comprising: the cytidine base editor according to any one of claims 1 to 14, and one or more of the following: a guide RNA sequence, a tracrRNA sequence or a target DNA sequence.

18. A method of editing a nucleobase of a nucleic acid sequence, the method comprising: contacting the nucleic acid sequence with the cytidine base editor according to any one of claims 1 to 14, and converting a first nucleobase of the nucleic acid sequence to a second nucleobase.

19. The method according to claim 18, further comprising contacting the nucleic acid sequence with a guiding polynucleotide to effect the conversion.

20. The method according to any one of claims 18 to 19, wherein the first nucleobase is cytosine, the second nucleobase is thymidine.

21. A nucleic acid vector encoding a cytidine base editor, the cytidine base editor comprising: (i) a polynucleotide programmable DNA binding domain; and (ii) a cytidine deaminase, wherein the cytidine deaminase consists of an amino acid sequence selected from any one of SEQ ID NO: 18 and 228 to 240, 242, 244 to 289.

22. The nucleic acid vector according to claim 21, further comprising a polynucleotide sequence encoding one or more uracil glycosylase inhibitor (UGI) domains.

23. The nucleic acid vector according to claim 21, further comprising a polynucleotide sequence encoding 2 glycosylase inhibitor (UGI) domains.

24. The nucleic acid vector according to claim 21, further comprising a polynucleotide sequence encoding one or more nuclear localization signals (NLSs).

25. The nucleic acid vector according to claim 24, further comprising a polynucleotide sequence encoding an N-terminal NLS and / or a C-terminal NLS.

26. The nucleic acid vector according to claim 24, wherein the NLS is a bipartite NLS.

27. The nucleic acid vector according to any one of claims 21 to 26, wherein the programmable DNA binding domain is Cas9 selected from the following: Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), nuclease-dead Cas12, nuclease-dead TnpB, nuclease-dead IscB, nickase IscB, nuclease-active Cas12, nuclease-active TnpB, nuclease-active IscB or nuclease-active Cas9.

28. A nucleic acid vector encoding a cytidine base editor, the cytidine base editor comprising: (i) a polynucleotide programmable DNA binding domain; and (ii) a cytidine deaminase, wherein the nucleic acid encoding the cytidine deaminase consists of a polynucleotide sequence selected from any one of SEQ ID NO: 42 and 324 to 336, 338, 340 to 385.

29. The nucleic acid vector according to claim 28, further comprising a polynucleotide sequence encoding one or more uracil glycosylase inhibitor (UGI) domains.

30. The nucleic acid vector according to claim 28, further comprising a polynucleotide sequence encoding two glycosylase inhibitor (UGI) domains.

31. The nucleic acid vector according to claim 28, further comprising a polynucleotide sequence encoding one or more nuclear localization signals (NLS).

32. The nucleic acid vector according to claim 31, further comprising a polynucleotide sequence encoding an N-terminal NLS and / or a C-terminal NLS.

33. The nucleic acid vector according to claim 31, wherein the NLS is a bipartite NLS.

34. The nucleic acid vector according to any one of claims 28 to 33, wherein the programmable DNA binding domain is Cas9 selected from the following: Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), nuclease-dead Cas12, nuclease-dead TnpB, nuclease-dead IscB, nickase IscB, nuclease-active Cas12, nuclease-active TnpB, nuclease-active IscB or nuclease-active Cas9.

35. A composition comprising: (a) a nucleic acid or a vector comprising a nucleic acid, the nucleic acid encoding a guide RNA; and (b) a cytidine base editor comprising: (i) a polynucleotide programmable DNA binding domain; and (ii) a cytidine deaminase, wherein the cytidine deaminase consists of an amino acid sequence selected from any one of SEQ ID NO: 18 and 228 to 240, 242, 244 to 289.

36. The composition according to claim 35, further comprising one or more uracil glycosylase inhibitor (UGI) domains.

37. The composition according to claim 35, wherein the cytidine base editor comprises two glycosylase inhibitor (UGI) domains.

38. The composition according to claim 35, further comprising one or more nuclear localization signals (NLS).

39. The composition according to claim 38, wherein the cytidine base editor comprises an N-terminal NLS and / or a C-terminal NLS.

40. The composition according to claim 38, wherein the NLS is a bipartite NLS.

41. The composition according to claim 35, wherein the cytidine base editor has an increased cis / trans activity ratio compared to the BE4max cytidine base editor.

42. The composition according to claim 41, wherein the increased cis / trans activity ratio is at least 2-fold higher than that of the BE4max cytidine base editor.

43. The composition according to claim 35, wherein the cytidine base editor has at least 5% cis activity compared to the BE4max cytidine base editor.

44. The composition according to claim 35, wherein the cytidine base editor has at least 5% trans activity compared to the BE4max cytidine base editor.

45. The composition according to any one of claims 35 to 44, wherein the programmable DNA-binding domain is a Cas9 selected from the following: Staphylococcus aureus Cas9 (SaCas9), Streptococcus pyogenes Cas9 (Sp-Cas9), nuclease-dead Cas9 (dCas9), Cas9 nickase (nCas9), nuclease-dead Cas12, nuclease-dead TnpB, nuclease-dead IscB, nickase IscB, nuclease-active Cas12, nuclease-active TnpB, nuclease-active IscB or nuclease-active Cas9.

46. A fusion protein comprising: a polynucleotide programmable DNA-binding domain, and at least one nucleobase editor domain comprising a cytidine deaminase, wherein the cytidine deaminase consists of an amino acid sequence selected from any one of SEQ ID NO: 18 and 228 to 240, 242, 244 to 289.

47. The fusion protein according to claim 46, further comprising one or more uracil glycosylase inhibitor (UGI) domains.

48. The fusion protein according to claim 46, further comprising two glycosylase inhibitor (UGI) domains.

49. The fusion protein according to claim 46, further comprising one or more nuclear localization signals (NLS).

50. The fusion protein according to claim 49, further comprising an N-terminal NLS and / or a C-terminal NLS.

51. The fusion protein according to any one of claims 49 to 50, wherein the NLS is a bipartite NLS.

52. An engineered system comprising: a catalytically inactivated Cas13 effector protein (dCas13), or a nucleotide sequence encoding a catalytically inactivated Cas13 effector protein, wherein the Cas13 effector protein is Cas13a, wherein the system further comprises a functional component, or wherein the nucleotide sequence further encodes a functional component, wherein the functional component is a base editing component, wherein the base editing component comprises a cytidine deaminase or its catalytic domain, wherein the cytidine deaminase or its catalytic domain is inserted into the inner loop of dCas13, and wherein the cytidine deaminase consists of an amino acid sequence selected from any one of SEQ ID NOs: 18 and 228 to 240, 242, 244 to 289.

Citation Information

Patent Citations

  • Nucleobase editors and uses thereof

    WO2017070632A2