Deaminases for base editing and their variants

By developing adenosine deaminase and cytidine deaminase and binding with polynucleotide programmable DNA binding domains, a nucleobase editor is formed, which solves the problem of insufficient specificity and efficiency of base editing in the prior art, and achieves efficient and specific base editing effects.

CN119162157BActive Publication Date: 2025-06-03ACCUREDIT THERAPEUTICS (SUZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411624246.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-06-03
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The specificity and efficiency of existing base editors induce modifications within the targeting sequence are insufficient, making it difficult to meet the needs of efficient targeted editing.

Method used

A deaminase and cytidine deaminase were developed to improve the editing efficiency of adenosine and cytosine in DNA by binding to specific amino acid sequences or variants thereof. These deaminases can bind to polynucleotide programmable DNA binding domains to form a nucleobase editor for efficient base editing.

Benefits of technology

By using these efficient deaminases, efficient and specific base editing within the targeted sequence is achieved, improving the efficiency and accuracy of base editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119162157B_ABST
    Figure CN119162157B_ABST
Patent Text Reader

Abstract

The present application provides deaminases for base editing and variants thereof, nucleobase editors or fusion proteins comprising the deaminases, their coding polynucleotides or expression vectors, and methods of using such editors to generate modifications in a target nucleobase sequence. The present application also provides mutated Cas9.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to deaminases for use with, for example, nucleic acid-guided nucleases to modify target nucleic acid sequences. More specifically, the present application relates to adenosine deaminases and cytidine deaminases with high base editing efficiency. Background Art

[0002] Targeted editing of nucleic acid sequences, such as targeted cleavage or targeted modification of genomic DNA, is an effective method for studying gene function and also has the potential to provide new treatment methods for human genetic diseases. Currently available base editors include cytidine base editors (such as CBE4) that convert target C•G base pairs to T•A base pairs, and adenosine base editors (such as ABE7.10) that convert A•T base pairs to G•C base pairs. There is a need in the art for improved base editors that can induce modifications within target sequences with higher specificity and efficiency. Summary of the Invention

[0003] In a first aspect, the present application provides an adenosine deaminase comprising:

[0004] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with any one of SEQ ID NOs: 1-47 (e.g., SEQ ID NOs: 2-47); or

[0005] (ii) one or more alterations selected from the group consisting of D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, K169X relative to position in SEQ ID NO: 1, wherein X is any amino acid, preferably one or more substitutions selected from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, K169N relative to position in SEQ ID NO: 1, wherein the amino acid positions are numbered relative to SEQ ID NO: 1.

[0006] In some embodiments, the adenosine deaminase comprises:

[0007] (i) an amino acid having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5% or 95% to 99.5%) sequence identity with SEQ ID NO: 1, and

[0008] (ii) one or more alterations selected from the group consisting of positions D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, K169X relative to SEQ ID NO: 1, wherein X is any amino acid, and preferably, one or more substitutions selected from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, K169N relative to SEQ ID NO: 1, wherein the amino acid positions are numbered relative to SEQ ID NO: 1.

[0009] In some embodiments, the adenosine deaminase is capable of deaminating adenosine in DNA. In some embodiments, the adenosine deaminase is capable of deaminating cytosine in DNA, but with lower efficiency.

[0010] In some embodiments, the adenosine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 1-47 (e.g., SEQ ID NOs: 2-47), or consists of the amino acid sequence of any one of SEQ ID NOs: 1-47 (e.g., SEQ ID NOs: 2-47).

[0011] In some embodiments, the adenosine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 1-19 (e.g., SEQ ID NOs: 2-19), or consists of the amino acid sequence of any one of SEQ ID NOs: 1-19 (e.g., SEQ ID NOs: 2-19).

[0012] In some embodiments, the adenosine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 20-47, or consists of the amino acid sequence of any one of SEQ ID NOs: 20-47.

[0013] In a second aspect, the present application provides a cytidine deaminase, comprising:

[0014] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with any one of SEQ ID NOs: 48-63 or 64-99 (e.g., SEQ ID NOs: 48-63 or 65-99); or

[0015] (ii) one or more alterations selected from the group consisting of A17X, G64X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, K153X relative to SEQ ID NO: 64, wherein X is any amino acid, preferably one or more substitutions selected from the group consisting of A17N, G64C, Y70G or Y70H or Y70K, L73F, V79T, F81L, M91L, H93N, G102S, D112N, A119G, N124V or N124P, R130V, S139K, A140R, C143A, R151M, K153G relative to SEQ ID NO: 64, wherein the amino acid positions are numbered relative to SEQ ID NO: 64.

[0016] In some embodiments, the cytidine deaminase comprises:

[0017] (i) an amino acid having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5% or 95% to 99.5%) sequence identity with SEQ ID NO: 64, and

[0018] (ii) one or more alterations selected from the group consisting of A17X, G64X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, K153X relative to SEQ ID NO: 64, wherein X is any amino acid, preferably one or more substitutions selected from the group consisting of A17N, G64C, Y70G or Y70H or Y70K, L73F, V79T, F81L, M91L, H93N, G102S, D112N, A119G, N124V or N124P, R130V, S139K, A140R, C143A, R151M, K153G relative to SEQ ID NO: 64, wherein the amino acid positions are numbered relative to SEQ ID NO: 64.

[0019] In some embodiments, the cytidine deaminase is capable of deaminating cytidine in DNA. In some embodiments, the cytidine deaminase is capable of deaminating adenosine in DNA, but with lower efficiency.

[0020] In some embodiments, the cytidine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 48-63 or 64-99 (e.g., SEQ ID NOs: 48-63 or 65-99), or consists of the amino acid sequence of any one of SEQ ID NOs: 48-63 or 64-99 (e.g., SEQ ID NOs: 48-63 or 65-99).

[0021] In a third aspect, the present application provides a base editor, comprising:

[0022] (i) a polynucleotide programmable DNA binding domain; and

[0023] (ii) a deaminase, wherein the deaminase is the adenosine deaminase described in the first aspect or the cytidine deaminase described in the second aspect.

[0024] In some embodiments, the base editor further comprises one or more uracil glycosylase inhibitor (UGI) domains.

[0025] In some embodiments, the base editor is an adenosine base editor (ABE), wherein the deaminase is an adenosine deaminase. In a preferred embodiment, the adenosine deaminase comprises:

[0026] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with any one of SEQ ID NOs: 1-47 (e.g., SEQ ID NOs: 2-47); or

[0027] (ii) one or more alterations selected from the group consisting of positions D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, K169X relative to SEQ ID NO: 1, wherein X is any amino acid, preferably one or more substitutions selected from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, K169N relative to SEQ ID NO: 1, wherein the amino acid positions are numbered relative to SEQ ID NO: 1.

[0028] In some embodiments, the base editor is a cytidine base editor (CBE). In a preferred embodiment, the cytidine deaminase comprises: The cytidine deaminase comprises:

[0029] (i)An amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 48 - 63 or 64 - 99 (e.g., SEQ ID NOs: 48 - 63 or 65 - 99); or

[0030] (ii)One or more alterations selected from the group consisting of A17X, G64X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, K153X relative to position in SEQ ID NO: 64, where X is any amino acid, preferably one or more substitutions selected from the group consisting of A17N, G64C, Y70G or Y70H or Y70K, L73F, V79T, F81L, M91L, H93N, G102S, D112N, A119G, N124V or N124P, R130V, S139K, A140R, C143A, R151M, K153G relative to SEQ ID NO: 64, wherein the amino acid positions are numbered relative to SEQ ID NO: 64.

[0031] In some embodiments, the polynucleotide programmable DNA - binding domain comprises a domain selected from the group consisting of a Cas9 domain, a Cas12 domain, a TnpB domain, an IscB domain, a homing endonuclease (meganuclease) domain, a zinc - finger DNA - binding domain, and a transcription activator - like effector (TALE) DNA - binding domain. In some embodiments, the polynucleotide programmable DNA - binding domain comprises a Cas9 domain. In some embodiments, the Cas9 domain is Staphylococcus aureus Cas9 (SaCas9), Streptococcus thermophilus l Cas9 (St1Cas9), Streptococcus pyogenes Cas9 (SpCas9), or a variant thereof. In some embodiments, the polynucleotide programmable DNA - binding domain is an inactive nuclease or a nickase variant.

[0032] In some embodiments, the polynucleotide programmable DNA - binding domain is a Cas9 domain, preferably a mutant Cas9 comprising one or two amino acid substitutions selected from D700R or E1243R, more preferably Cas9 comprising D700R or Cas9 comprising E1243R. In a preferred embodiment, the mutant Cas9 comprises the amino acid sequence of SEQ ID NO: 115 or 116, or consists of the amino acid sequence of SEQ ID NO: 115 or 116.

[0033] In some embodiments, the base editor comprises a nucleic acid-guided nuclease protein and an embedded nucleobase editor (NBE) domain. In some embodiments, the embedded NBE domain is an adenosine base editor (ABE) domain. In some embodiments, the embedded ABE domain is an embedded adenosine deaminase protein domain. In some embodiments, the polynucleotide programmable DNA-binding domain comprises a linker flanking at least one embedded domain protein. In some embodiments, the embedded NBE domain is a cytidine base editor (CBE) domain. In some embodiments, the embedded CBE domain is an embedded cytidine deaminase protein domain.

[0034] In some embodiments, the polynucleotide programmable DNA-binding domain is associated with a deaminase, e.g., directly linked together or linked together via a linker. Generally, those skilled in the art are able to select a suitable linker. In one embodiment, the linker comprises the amino acid sequence of any one of SEQ ID NO: 104 - 107, or consists of the amino acid sequence of any one of SEQ ID NO: 104 - 107. Preferably, the linker is SEQ ID NO: 106.

[0035] In some embodiments, the nucleobase editor can be in the form of a kit.

[0036] In a fourth aspect, the present application provides a fusion protein comprising a polynucleotide programmable DNA-binding domain and at least one nucleobase editor domain comprising a deaminase, which are linked together. The polynucleotide programmable DNA-binding domain and at least one nucleobase editor domain comprising a deaminase can be directly linked or linked via a linker.

[0037] Generally, those skilled in the art are able to select a suitable linker. In one embodiment, the linker comprises the amino acid sequence of any one of SEQ ID NO: 104 - 107, or consists of the amino acid sequence of any one of SEQ ID NO: 104 - 107. Preferably, the linker is SEQ ID NO: 106.

[0038] In some embodiments, the polynucleotide programmable DNA binding domain is a Cas9 domain, preferably a mutant Cas9 comprising one or more amino acid mutations (such as amino acid substitutions), for example, a mutant Cas9 comprising one or two amino acid substitutions selected from D700R or E1243R, more preferably Cas9 comprising D700R or Cas9 comprising E1243R. In a preferred embodiment, the mutant Cas9 comprises the amino acid sequence of SEQ ID NO: 115 or 116, or consists of the amino acid sequence of SEQ ID NO: 115 or 116.

[0039] In some embodiments, the deaminase is the adenosine deaminase or cytidine deaminase described in the present invention.

[0040] In a fifth aspect, the present application provides a polynucleotide encoding an adenosine deaminase, the adenosine deaminase comprising:

[0041] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 1-47 (such as SEQ ID NOs: 2-47); or

[0042] (ii) one or more alterations selected from the group consisting of positions D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, K169X relative to SEQ ID NO: 1, wherein X is any amino acid, preferably one or more substitutions selected from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, K169N relative to SEQ ID NO: 1, wherein the amino acid positions are numbered relative to SEQ ID NO: 1.

[0043] In some embodiments, the adenosine deaminase comprises:

[0044] (i) an amino acid having 80% to 99.5% (such as 85% to 99.5%, 90% to 99.5% or 95% to 99.5%) sequence identity to SEQ ID NO: 1, and

[0045] (ii) one or more alterations selected from the group consisting of positions D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, K169X relative to SEQ ID NO: 1, wherein X is any amino acid, preferably, one or more substitutions selected from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, K169N relative to SEQ ID NO: 1, wherein the amino acid positions are numbered relative to SEQ ID NO: 1.

[0046] In some embodiments, the polynucleotide may be codon-optimized for better expression in a host cell.

[0047] In a sixth aspect, the present application provides a polynucleotide encoding a cytidine deaminase, the cytidine deaminase comprising:

[0048] (i) an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 48-63 or 64-99 (e.g., SEQ ID NOs: 48-63 or 65-99); or

[0049] (ii) one or more alterations selected from the group consisting of positions A17X, G64X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, K153X relative to SEQ ID NO: 64, wherein X is any amino acid, preferably, one or more substitutions selected from the group consisting of A17N, G64C, Y70G or Y70H or Y70K, L73F, V79T, F81L, M91L, H93N, G102S, D112N, A119G, N124V or N124P, R130V, S139K, A140R, C143A, R151M, K153G relative to SEQ ID NO: 64, wherein the amino acid positions are numbered relative to SEQ ID NO: 64.

[0050] In some embodiments, the cytidine deaminase comprises:

[0051] (i)amino acids having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5%, or 95% to 99.5%) sequence identity with SEQ ID NO: 64, and

[0052] (ii)one or more alterations selected from the group consisting of A17X, G64X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, K153X relative to SEQ ID NO: 64, wherein X is any amino acid, preferably one or more substitutions selected from the group consisting of A17N, G64C, Y70G or Y70H or Y70K, L73F, V79T, F81L, M91L, H93N, G102S, D112N, A119G, N124V or N124P, R130V, S139K, A140R, C143A, R151M, K153G relative to SEQ ID NO: 64, and the amino acid positions are numbered relative to SEQ ID NO: 64.

[0053] In some embodiments, the polynucleotide may be codon-optimized for better expression in a host cell.

[0054] In a seventh aspect, the present application provides an expression vector comprising a polynucleotide encoding the adenosine deaminase of the present invention. In some embodiments, the expression vector is a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vector, adenoviral vector, lentiviral vector, Sendai virus vector, and herpesvirus vector.

[0055] In an eighth aspect, the present application provides an expression vector comprising a polynucleotide encoding the cytidine deaminase of the present invention. In some embodiments, the expression vector is a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vector, adenoviral vector, lentiviral vector, Sendai virus vector, and herpesvirus vector.

[0056] In a ninth aspect, the present application provides a polynucleotide encoding a base editor, the base editor comprising: (i) a polynucleotide programmable DNA binding domain; and (ii) a deaminase, wherein the deaminase is the adenosine deaminase described in the first aspect or the cytidine deaminase described in the second aspect.

[0057] In some embodiments, the deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 1-47, SEQ ID NOs: 48-63 or 64-99 (e.g., SEQ ID NOs: 2-47, SEQ ID NOs: 48-63 or 65-99).

[0058] In some embodiments, the deaminase is an adenosine deaminase, and the adenosine deaminase comprises:

[0059] (i) amino acids having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5% or 95% to 99.5%) sequence identity to SEQ ID NO: 1, and

[0060] (ii) one or more alterations selected from the group consisting of D12X, H21X, A61X, V95X, S155X, S156X, C159X, E164X, E167X, K169X relative to the position in SEQ ID NO: 1, wherein X is any amino acid, preferably one or more substitutions selected from the group consisting of D12H, H21D or H21Q, A61C, V95T, S155R, S156N or S156K, C159R, E164V or E164W, E167R or E167W, K169N relative to the position in SEQ ID NO: 1, wherein the amino acid positions are numbered relative to SEQ ID NO: 1.

[0061] In some embodiments, the deaminase is a cytidine deaminase, and the cytidine deaminase comprises:

[0062] (i) amino acids having 80% to 99.5% (e.g., 85% to 99.5%, 90% to 99.5% or 95% to 99.5%) sequence identity to SEQ ID NO: 64, and

[0063] (ii) selecting one or more alterations from the group consisting of A17X, G64X, Y70X, L73X, V79X, F81X, M91X, H93X, G102X, D112X, A119X, N124X, R130X, S139X, A140X, C143X, R151X, K153X relative to SEQ ID NO: 64, wherein X is any amino acid, preferably, selecting one or more substitutions from the group consisting of A17N, G64C, Y70G or Y70H or Y70K, L73F, V79T, F81L, M91L, H93N, G102S, D112N, A119G, N124V or N124P, R130V, S139K, A140R, C143A, R151M, K153G relative to SEQ ID NO: 64, wherein the amino acid positions are numbered relative to SEQ ID NO: 64.

[0064] In some embodiments, the polynucleotide may be codon-optimized for better expression in a host cell. In a tenth aspect, the present application provides an expression vector comprising a polynucleotide encoding the nucleobase editor of the present invention. In some embodiments, the vector is a viral vector selected from the group consisting of adeno-associated virus (AAV), retroviral vector, adenoviral vector, lentiviral vector, Sendai virus vector, and herpesvirus vector.

[0065] In an eleventh aspect, the present application provides a polynucleotide encoding the fusion protein of the present invention and an expression vector comprising the polynucleotide.

[0066] In some embodiments, the polynucleotide may be codon-optimized for better expression in a host cell.

[0067] It will be readily understood by those skilled in the art that the expression vector of the present invention may comprise transcriptional regulatory elements such as promoters, enhancers, etc.

[0068] In a twelfth aspect, the present application provides a cell comprising the expression vector of the present invention.

[0069] In some embodiments, the cell may be any suitable type of cell, e.g., a prokaryotic cell or a eukaryotic cell, including, but not limited to: bacterial cells, fungal cells, plant cells, insect cells, or mammalian cells.

[0070] In a thirteenth aspect, the present application provides a base editing method, comprising contacting a polynucleotide sequence with a base editor, wherein the base editor comprises: (i) a polynucleotide programmable DNA binding domain; and (ii) a deaminase, wherein the deaminase is the adenosine deaminase described in the first aspect or the cytidine deaminase described in the second aspect, and wherein the deaminase deaminates a nucleobase in the polynucleotide, thereby editing the polynucleotide sequence.

[0071] In some embodiments, the deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 1-47, SEQ ID NOs: 48-63 or 64-99 (e.g., SEQ ID NOs: 2-47, SEQ ID NOs: 48-63 or 65-99).

[0072] In a fourteenth aspect, the present application provides a method for correcting a genetic defect in a subject, the method comprising administering to the subject a base editor or a polynucleotide encoding a base editor, wherein the base editor comprises: (i) a polynucleotide programmable DNA binding domain; and (ii) a deaminase, wherein the deaminase is the adenosine deaminase described in the first aspect or the cytidine deaminase described in the second aspect, to deaminate a target nucleobase in the target nucleotide sequence of the subject, thereby correcting the genetic defect.

[0073] In some embodiments, the deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 1-47, SEQ ID NOs: 48-63 or 64-99 (e.g., SEQ ID NOs: 2-47, SEQ ID NOs: 48-63 or 65-99).

[0074] In some embodiments, the method comprises delivering the base editor or the polynucleotide encoding the base editor, and one or more guide polynucleotides, to the cells of the subject. In some embodiments, the subject is a mammal, including a non-human mammal or a human. In some embodiments, deamination of the target nucleobase causes the target nucleobase to be replaced by a wild-type nucleobase.

[0075] In a fifteenth aspect, the present invention provides the pharmaceutical use of the adenosine deaminase or cytidine deaminase, base editor or fusion protein of the present invention, or their encoding polynucleotides.

[0076] In some embodiments, the present application provides the use of the adenosine deaminase or cytidine deaminase of the present invention or their encoding polynucleotides in the preparation of a drug or a kit for base editing or correcting genetic defects in a subject.

[0077] In some embodiments, the present application provides the use of the fusion protein of the present invention or its encoding polynucleotide in the preparation of a drug or a kit for base editing or correcting genetic defects in a subject.

[0078] In some embodiments, the present application provides the use of the nucleobase editor of the present invention or the polynucleotide encoding the nucleobase editor in the preparation of a drug or a kit for base editing or correcting genetic defects in a subject.

[0079] In a sixteenth aspect, the present application provides a molecular complex comprising the base editor of the present invention and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences.

[0080] In a seventeenth aspect, the present application provides a molecular complex comprising the fusion protein of the present invention and one or more guide RNA sequences, tracrRNA sequences, or target DNA sequences.

[0081] In an eighteenth aspect, the present application provides a pharmaceutical composition comprising the base editor or fusion protein of the present invention, or an expression vector comprising a polynucleotide encoding the base editor or fusion protein, or a molecular complex, and a pharmaceutically acceptable excipient.

[0082] In a nineteenth aspect, the present application provides a kit comprising: the base editor or fusion protein of the present invention, or an expression vector comprising a polynucleotide encoding the base editor or fusion protein, or a molecular complex, and instructions for use.

[0083] The pharmaceutical composition or kit of the present invention can be used for base editing or correcting genetic defects in a subject.

[0084] In a twentieth aspect, the present application provides a mutant Cas9 comprising one or two amino acid substitutions selected from D700R or E1243R.

[0085] In some embodiments, the mutant Cas comprises D700R, and in other embodiments, the mutant Cas comprises E1243R. In a preferred embodiment, the mutant Cas9 comprises the amino acid sequence of SEQ ID NO: 115 or 116, or consists of the amino acid sequence of SEQ ID NO: 115 or 116.

[0086] In the present invention, the terms "mutated Cas" and "Cas variant" are used interchangeably.

[0087] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Methods and materials for the present invention are described herein; other suitable methods and materials known in the art may also be used. The materials, methods, and examples are illustrative only and are not intended to be limiting. All publications, patent applications, patents, sequences, database entries, and other references mentioned herein are incorporated by reference in their entirety. In case of conflict, the present specification (including definitions) shall prevail.

[0088] Other features and advantages of the present invention will be apparent from the following detailed description, the accompanying drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 Showing the A-to-G editing efficiency of different TadA variants verified in Example 1 in the sgRNA library cell line.

[0090] Figure 2 Showing the fold change in the percentage of mutations in the surviving colonies before and after selection (16 or 64 μg / ml chloramphenicol) of Escherichia coli cells transformed with the cytidine deaminase variant library in Example 7.

[0091] Figure 3 Showing the percentage of mutations identified by next-generation sequencing (NGS) in the surviving colonies on the selective plates for the engineered CBE variants of the present invention in Example 7.

[0092] Figure 4 Showing the C-to-T editing efficiency of variants T88.74-V10, T88.74-V14, T88.74-V16, and T88.74-V17 relative to T88.74-V0 detected in Example 8.

[0093] Figure 5 Showing the fold change in the percentage of mutations in the surviving colonies before and after selection (64 or 128 μg / ml chloramphenicol) of Escherichia coli cells transformed with the cytidine deaminase variant library in Example 8.

[0094] Figure 6 Showing the C-to-T editing efficiency of variants T88.74-V21 and T88.74-V22 relative to variant T88.74-V14 detected in Example 8.

[0095] Figure 7Show the C-to-T editing efficiency of the variants T88.74-V27, T88.74-V29, T88.74-V32, and T88.74-V33 detected in Example 8 relative to the variant T88.74-V17. Detailed Description

[0096] Base Editor

[0097] Disclosed herein are nucleobase editors for editing, modifying, or altering a target nucleotide sequence of a polynucleotide, such as an adenosine base editor (ABE) or a cytidine base editor (CBE). Nucleobase editors are described herein that comprise a programmable nucleotide-binding domain (e.g., a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9), or a zinc finger DNA-binding domain, or a TALE DNA-binding domain) and at least one nucleobase editing domain (e.g., an adenosine deaminase or a cytidine deaminase). When a polynucleotide-programmable nucleotide-binding domain (e.g., Cas9 or Cas12) is present in a cell and binds to a bound guide polynucleotide (e.g., gRNA), it can specifically bind to a target polynucleotide sequence (i.e., by complementary base pairing between the bases of the bound guide nucleic acid and the bases of the target polynucleotide sequence) and thereby localize the base editor to the target nucleic acid sequence to be edited.

[0098] In some embodiments, base editing activity is evaluated by editing efficiency. Base editing can be determined by any suitable method, such as by Sanger sequencing or next-generation sequencing. In some embodiments, base editing efficiency is measured as the total sequencing read percentage of nucleobase conversions achieved by the base editor. For example, for adenosine deaminase, the total sequencing read percentage of target A-T base pair conversions to G-C base pairs is measured; for cytidine deaminase, the total sequencing read percentage of target C-G base pair conversions to T-A base pairs is measured. In some embodiments, when base editing is performed in a cell population, base editing efficiency is measured as the total cell percentage having nucleobase conversions achieved by the base editor.

[0099] The term "base editor system" refers to the components required to edit the nucleobases of a target nucleotide sequence. In various embodiments, a base editor system comprises: (1) a polynucleotide programmable nucleotide-binding domain (e.g., Cas9); (2) a deaminase domain for deaminating the nucleobase (e.g., adenosine deaminase and / or cytidine deaminase; see PCT / US2019 / 044935, PCT / US2020 / 016288, the entire contents of each of which are incorporated herein by reference); and (3) one or more guide polynucleotides (e.g., guide RNA). In some embodiments, the polynucleotide programmable nucleotide-binding domain is a polynucleotide programmable DNA-binding domain. In some embodiments, the base editor is an adenosine base editor (ABE). In other embodiments, the base editor is a cytidine base editor (CBE).

[0100] In some embodiments, a base editor system can comprise more than one base editing component. For example, a base editor system can comprise more than one deaminase. In some embodiments, a base editor system can comprise one or more adenosine deaminases or one or more cytidine deaminases. In some embodiments, different deaminases can be targeted to a target nucleic acid sequence using a single guide polynucleotide. In some embodiments, different deaminases can be targeted to a target nucleic acid sequence using a pair of guide polynucleotides.

[0101] The deaminase domain and the polynucleotide programmable nucleotide-binding component of a base editor system can be associated with each other covalently or non-covalently, or in any combination of their association and interaction. For example, in some embodiments, the deaminase domain can be targeted to a target nucleotide sequence by the polynucleotide programmable nucleotide-binding domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain can be fused or linked to the deaminase domain. In some embodiments, the polynucleotide programmable nucleotide-binding domain can target the deaminase domain to a target nucleotide sequence by non-covalent interaction or association with the deaminase domain. For example, in some embodiments, the deaminase domain can comprise additional heterologous portions or domains that are capable of interacting with, associating with, or forming a complex with additional heterologous portions or domains that are part of the polynucleotide programmable nucleotide-binding domain.

[0102] In some embodiments, an additional heterologous moiety may be capable of binding to, interacting with, associating with, or forming a complex with a polypeptide. In some embodiments, an additional heterologous moiety may be capable of binding to, interacting with, associating with, or forming a complex with a polynucleotide. In some embodiments, an additional heterologous moiety may be capable of binding to a guide polynucleotide. In some embodiments, an additional heterologous moiety may be capable of binding to a polypeptide linker. In some embodiments, an additional heterologous moiety may be capable of binding to a polynucleotide linker. The additional heterologous moiety may be a protein domain. In some embodiments, the additional heterologous moiety may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.

[0103] The base editor system may further comprise a guide polynucleotide component. It should be understood that the components of the base editor system may associate with each other by covalent bonds, non-covalent interactions, or any combination of their association and interaction. In some embodiments, the deaminase domain may be targeted to a target nucleotide sequence by a guide polynucleotide. For example, in some embodiments, the deaminase domain may comprise an additional heterologous moiety or domain (such as a polynucleotide-binding domain, such as an RNA or DNA-binding protein) capable of interacting with, associating with, or capable of forming a complex with a portion or segment (such as a polynucleotide motif) of the guide polynucleotide. In some embodiments, the additional heterologous moiety or domain (such as a polynucleotide-binding domain, such as an RNA or DNA-binding protein) may be fused or linked to the deaminase domain. In some embodiments, the additional heterologous moiety may be capable of binding to, interacting with, associating with, or forming a complex with a polypeptide. In some embodiments, the additional heterologous moiety may be capable of binding to, interacting with, associating with, or forming a complex with a polynucleotide. In some embodiments, the additional heterologous moiety may be capable of binding to a guide polynucleotide. In some embodiments, the additional heterologous moiety may be capable of binding to a polypeptide linker. In some embodiments, the additional heterologous moiety may be capable of binding to a polynucleotide linker. The additional heterologous moiety may be a protein domain. In some embodiments, the additional heterologous moiety may be a K homology (KH) domain, an MS2 coat protein domain, a PP7 coat protein domain, an SfMu Com coat protein domain, a sterile alpha motif, a telomerase Ku-binding motif and Ku protein, a telomerase Sm7-binding motif and Sm7 protein, or an RNA recognition motif.

[0104] Guide polynucleotide

[0105] In some embodiments, the guide polynucleotide is a guide RNA. The RNA / Cas complex can help "guide" the Cas protein to the target DNA. Cas9 / crRNA / tracrRNA cleaves linear or circular dsDNA targets complementary to the spacer in an endonuclease manner. The target strand not complementary to the crRNA is first cleaved in an endonuclease manner and then trimmed 3'-5' in an exonuclease manner. In nature, DNA binding and cleavage generally require a protein and two RNAs. However, a single guide RNA ("sgRNA" or simply "gRNA") can be engineered to incorporate aspects of both the crRNA and the tracrRNA into a single RNA species. See, e.g., Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference. Cas9 recognizes short motifs (PAM or protospacer adjacent motif) in the CRISPR repeat to help distinguish self from non-self.

[0106] The Cas9 nuclease sequences and structures are well known to those skilled in the art (see, e.g., "Complete genome sequence of an Ml strain of Streptococcus pyogenes." Ferretti, J.J. et al., Natl. Acad. Sci. U.S.A. 98:4658-4663 (2001); "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E. et al., Nature 471:602-607 (2011); and "Programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity." Jinek M. et al., Science 337:816-821 (2012), the entire contents of which are incorporated herein by reference). Cas9 orthologs have been described in various species, including but not limited to Streptococcus pyogenes and Streptococcus thermophilus. Based on the present application, other suitable Cas9 nucleases and sequences will be apparent to those skilled in the art, and such Cas9 nucleases and sequences include Cas9 sequences from the organisms and loci disclosed in Chylinski, Rhun and Charpentier, "The tracrRNA and Cas9 families of type II CRISPR-Cas immunity systems" (2013) RNA Biology 10:5, 726-737 (the entire contents of this reference are incorporated herein by reference). In some embodiments, the Cas9 nuclease has an inactive (e.g., inactivated) DNA cleavage domain, i.e., Cas9 is a nickase. In some embodiments, the guide polynucleotide is at least one single guide RNA ("sgRNA" or "gNRA"). In some embodiments, the guide polynucleotide is at least one tracrRNA. In some embodiments, the guide polynucleotide does not require a protospacer adjacent motif (PAM) sequence to direct the polynucleotide programmable DNA binding domain (e.g., Cas9 or Cpfl) to the target nucleotide sequence. The polynucleotide programmable nucleotide binding domain (e.g., a CRISPR-derived domain) of the base editors disclosed herein can recognize a target polynucleotide sequence by associating with a guide polynucleotide.Guide polynucleotides (e.g., gRNAs) are typically single-stranded and can be programmed to bind site-specifically (i.e., by complementary base pairing) to a target sequence of a polynucleotide, thereby guiding a base editor that binds to the guide nucleic acid to the target sequence. The guide polynucleotide can be DNA. The guide polynucleotide can be RNA. In some embodiments, the guide polynucleotide comprises natural nucleotides (e.g., adenosine). In some embodiments, the guide polynucleotide comprises non-natural (or not natural) nucleotides (e.g., peptide nucleic acids or nucleotide analogs). In some embodiments, the length of the targeting region of the guide nucleic acid sequence can be at least 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The length of the targeting region of the guide nucleic acid can be 10-30 nucleotides, or 15-25 nucleotides, or 15-20 nucleotides.

[0107] Programmable nucleotide binding domain

[0108] The programmable nucleotide binding domain of a base editor can itself comprise one or more domains. For example, a polynucleotide programmable nucleotide binding domain can comprise one or more nuclease domains. In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide binding domain can comprise an endonuclease or an exonuclease. As used herein, the term "exonuclease" refers to a protein or polypeptide capable of digesting nucleic acids (e.g., RNA or DNA) from a free end, and the term "endonuclease" refers to a protein or polypeptide capable of catalyzing (e.g., cleaving) an internal region in a nucleic acid (e.g., DNA or RNA). In some embodiments, the endonuclease can cleave a single strand in a double-stranded nucleic acid.

[0109] Any DNA destabilizing molecule can be used in the compositions described herein in any combination, including but not limited to Cas9 or Cas12 nickases, Cas9 or Cas12 proteins (e.g., dCas) operably linked to a single guide RNA (sgRNA), any RNA programmable system, zinc finger nuclease nickases (ZFN nickases), TALEN nickases, and / or one or more nucleotides (e.g., one or more peptide nucleic acids (PNAs), locked nucleic acids (LNAs), and / or bridged nucleic acids (BNAs)). In certain embodiments, the base editing composition comprises more than one DNA destabilizing molecule, such as one or more proteins (e.g., nickases, etc.) and / or one or more nucleotides. In certain embodiments, the composition comprises a ZFN nickase and one or more additional protein and / or nucleotide DNA destabilizing molecules (e.g., one or more nucleotides as described herein). In certain aspects, the base editing composition does not comprise a Cas9 protein, but may comprise other Cas proteins (e.g., non-Cas9 RNA programmable systems). In certain embodiments, the DNA destabilizing molecule comprises a zinc finger nuclease (ZFN) nickase.

[0110] In some embodiments, the nuclease is a zinc finger nuclease (ZFN) or a TALE DNA binding domain-nuclease fusion (TALEN). ZFNs and TALENs comprise a DNA binding domain (zinc finger protein or TALE DNA binding domain) and a cleavage domain or cleavage half-domain, the DNA binding domain having been engineered to bind to a target site in a selected gene.

[0111] At least one zinc finger protein (ZFP) DNA binding domain of the base editing composition can be operably linked to one or more other components of the base editing composition, such as to one or more DNA destabilizing molecules (e.g., to a Cas9 nickase, dCas9, etc.) and / or to at least one adenine or cytosine deaminase. In certain embodiments, at least one ZFP DNA binding domain is operably linked to an adenine or cytosine deaminase. In other embodiments, the base editing composition comprises first and second ZFP DNA binding domains, wherein the first ZFP DNA binding domain is operably linked to a Cas9 nickase. The ZFP DNA binding domain can comprise 3, 4, 5, 6, or more fingers and can bind to a target site on either side (5' or 3') of the targeted base to be edited. In certain embodiments, the target site bound by the ZFP is located 1 to 100 (or any number in between) nucleotides from either side of the targeted base. In other embodiments, the target site bound by the ZFP is located 1 to 50 (or any number in between) nucleotides from either side of the targeted base.

[0112] In some embodiments, the endonuclease can cleave both strands of a double-stranded nucleic acid molecule. In some embodiments, the polynucleotide programmable nucleotide-binding domain can be a deoxyribonuclease. In some embodiments, the polynucleotide programmable nucleotide-binding domain can be a ribonuclease.

[0113] In some embodiments, the nuclease domain of the polynucleotide programmable nucleotide-binding domain can cleave zero, one, or both strands of the target polynucleotide. In some embodiments, the polynucleotide programmable nucleotide-binding domain can include a nickase domain. As used herein, the term "nickase" refers to a polynucleotide programmable nucleotide-binding domain that includes a nuclease domain capable of cleaving only one of the two strands in a double-stranded nucleic acid molecule (e.g., DNA). In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide-binding domain by introducing one or more mutations into the active polynucleotide programmable nucleotide-binding domain. For example, when the polynucleotide programmable nucleotide-binding domain includes a nickase domain derived from Cas9, the Cas9-derived nickase domain can contain a D10A mutation and a histidine at position 840. In such an embodiment, residue H840 retains catalytic activity such that it can cleave a single strand of the nucleic acid duplex. In another example, the Cas9-derived nickase domain can contain an H840A mutation while the amino acid residue at position 10 remains D. In some embodiments, a nickase can be derived from a fully catalytically active (e.g., native) form of the polynucleotide programmable nucleotide-binding domain by removing all or part of the nuclease domain not required for nickase activity. For example, when the polynucleotide programmable nucleotide-binding domain contains a nickase domain derived from Cas9, the Cas9-derived nickase domain can contain a deletion of all or part of the RuvC domain or the HNH domain.

[0114] Thus, a base editor comprising a polynucleotide programmable nucleotide-binding domain (comprising a nickase domain) is capable of generating a single-stranded DNA break (nick) at a specific polynucleotide target sequence (e.g., determined by the complementary sequence of the bound guide nucleic acid). In some embodiments, the strand of the nucleic acid duplex target polynucleotide sequence cleaved by a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain) is the strand that is not edited by the base editor (i.e., the strand cleaved by the base editor is opposite the strand containing the base to be edited). In other embodiments, a base editor comprising a nickase domain (e.g., a Cas9-derived nickase domain) can cleave the strand of the DNA molecule that is targeted for editing. In such embodiments, the non-targeted strand is not cleaved.

[0115] The present disclosure also provides base editors that include a catalytically inactivated (i.e., unable to cleave a target polynucleotide sequence) polynucleotide programmable nucleotide binding domain. As used herein, the terms "catalytically inactivated" and "nuclease inactivated" are used interchangeably and refer to a polynucleotide programmable nucleotide binding domain having one or more mutations and / or deletions that result in its inability to cleave a nucleic acid strand while retaining its ability and specificity to bind to a target polynucleotide. In some embodiments, a catalytically inactivated polynucleotide programmable nucleotide binding domain base editor may lack nuclease activity due to specific point mutations in one or more nuclease domains. For example, in the case of a base editor that includes a Cas9 domain, the Cas9 may include both a D10A mutation and an H840A mutation. Such mutations inactivate both nuclease domains, resulting in the loss of nuclease activity. In other embodiments, a catalytically inactivated polynucleotide programmable nucleotide binding domain may include one or more deletions of all or part of a catalytic domain (e.g., the RuvC1 and / or HNH domains). In further embodiments, a catalytically inactivated polynucleotide programmable nucleotide binding domain includes a point mutation (e.g., D10A or H840A) and a deletion of all or part of a nuclease domain.

[0116] Some aspects of the present application provide fusion proteins that include a domain that acts as a polynucleotide-programmable DNA-binding protein, which can be used to direct a protein (such as a base editor) to a specific nucleic acid (e.g., DNA or RNA) sequence. In certain embodiments, the fusion protein includes a nucleic acid-programmable DNA-binding protein domain and one or more deaminase domains. Non-limiting examples of polynucleotide-programmable DNA-binding proteins include Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, and Cas12i. Non-limiting examples of Cas enzymes include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (also known as Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csx11, Csf1, Csf2, Csf3, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V Cas effector proteins, type VI Cas effector proteins, CARF, DinG, their homologs, or their modified or engineered forms. Other polynucleotide-programmable DNA-binding proteins are also within the scope of the present application, although they may not be specifically listed in the present application.For example, see Makarova et al., "Classification and Nomenclature of CRISPR-Cas Systems: Wherefrom Here", CRISPR J. 2018 Oct; 1:325-336. doi: 10.1089 / crispr.2018.0033; Yan et al., "Functionally diverse type V CRISPR-Cas systems", Science. 2019 Jan 4;363(6422):88-91. doi: 10.1126 / science.aav7271, the entire contents of each reference are incorporated herein by reference.

[0117] In some embodiments, the polynucleotide programmable DNA binding domain is a Cas9 domain, preferably a mutant Cas9 comprising one or two amino acid substitutions selected from D700R or E1243R, more preferably Cas9 comprising D700R or Cas9 comprising E1243R. In a preferred embodiment, the mutant Cas9 comprises the amino acid sequence of SEQ ID NO: 115 or 116, or consists of the amino acid sequence of SEQ ID NO: 115 or 116.

[0118] In some embodiments, the present application provides a fusion protein comprising a type V CRISPR / Cas effector protein. Type V CRISPR / Cas effector proteins are subtypes of class 2 CRISPR / Cas effector proteins. For examples of type V CRISPR / Cas systems and their effector proteins (e.g., Cas12 family proteins such as Cas12a), see, e.g., Shmakov et al., Nat Rev Microbial. 2017 March; 15(3):169-182: "Diversity and evolution of class 2 CRISPR-Cas systems." Examples include but are not limited to: the Cas12 family (Cas12a, Cas12b, Cas12c), C2c4, C2c8, C2c5, C2c10, and C2c9; and CasX (Cas12e) and CasY (Cas12d). See also, e.g., Koonin et al., Curr Opin Microbial. 2017 June; 37:67-78: "Diversity, classification and evolution of CRISPR-Cas systems." In some embodiments, the CBEs disclosed herein comprise a type V CRISPR / Cas effector protein.

[0119] In some embodiments, the present application provides a TALE base editor that can comprise a TALE domain, a deaminase domain, and / or a cofactor protein (e.g., FokI endonuclease) domain, the TALE base editor comprising a fusion protein having a general structure of NH2-[TALE]-[deaminase domain]-COOH, NH2-[deaminase domain]-[TALE]-COOH, NH2-[TALE]-[deaminase domain]-[cofactor protein]-COOH, NH2-[cofactor protein]-[deaminase domain]-[TALE]-COOH, NH2-[cofactor protein]-[TALE]-[deaminase]-COOH, or NH2-[deaminase domain]-[TALE]-[cofactor protein]-COOH; wherein each instance of "]-[ " includes an optional linker, such as a peptide linker.

[0120] In some embodiments, the disclosed methods involve transducing (e.g., by transfection) cells with multiple complexes, each complex comprising a fusion protein and a cofactor protein, the fusion protein comprising a TAL effector domain and a deaminase domain, wherein each cofactor protein localizes the fusion protein to a different target sequence. See Yang L. et al., Engineering and optimizing deaminase fusions for genome editing, Nature Comms., 2016. In certain embodiments, the disclosed methods involve a TAL effector domain that binds to a target site by binding to the major groove of the DNA double helix rather than by Watson-Crick hybridization. In some embodiments, the method involves transfecting nucleic acid constructs (e.g., plasmids), each nucleic acid construct (or together) encoding components of multiple complexes of a TALE base editor and a cofactor protein, the TALE base editor comprising a TALE domain and a deaminase domain. In certain embodiments, the disclosed fusion protein comprises a cofactor protein domain, i.e., the domain is integrated into the fusion protein construct. In other embodiments, the TALE base editor comprises a TALE domain and a deaminase domain, and the cofactor protein is introduced into the cell separately from the base editor.

[0121] In some embodiments of the disclosed methods, the construct encoding the TALE base editor is transfected into the cell separately from the construct encoding the cofactor protein. In some embodiments, these components are encoded on a single construct and transfected together. In certain embodiments, these single constructs encoding the TALE base editor and the cofactor protein can be iteratively transfected into the cell, each iteration being associated with a subset of the target sequences. In certain embodiments, these single constructs can be transfected into the cell over several days. In other embodiments, they can be transfected into the cell over several weeks.

[0122] A to G editing

[0123] In some embodiments, the base editor described herein may comprise a deaminase domain that includes an adenosine deaminase. This adenosine deaminase domain of the base editor can facilitate the editing of an adenine (A) nucleobase to a guanine (G) nucleobase by deaminating adenosine or adenine (A) to form inosine (I), which exhibits base pairing properties similar to G, i.e., inosine can pair with cytosine, thereby introducing guanine at the site of the adenine that has been deaminated by the adenosine deaminase during subsequent transcription. The adenosine deaminase is capable of deaminating (i.e., removing the amine group) the adenine of deoxyadenosine residues in deoxyribonucleic acid (DNA).

[0124] In some embodiments, the base editors provided herein can be prepared by fusing together one or more protein domains, resulting in a fusion protein. In certain embodiments, the fusion proteins provided herein contain one or more features that improve the base editing activity (e.g., efficiency, selectivity, and specificity) of the fusion protein. For example, the fusion proteins provided herein can contain a Cas9 domain with reduced nuclease activity. In some embodiments, the fusion proteins provided herein can have a Cas9 domain that lacks nuclease activity (dCas9), or a Cas9 domain that cleaves one strand of a double-stranded DNA molecule (referred to as Cas9 nickase (nCas9)). Without being bound by any particular theory, the presence of catalytic residues (e.g., H840) maintains the activity of Cas9 to cleave the unedited (e.g., non-deaminated) strand containing a T opposite the targeted A. Mutation of the catalytic residues of Cas9 (e.g., D10 to A10) prevents cleavage of the edited strand containing the targeted A residue. Such Cas9 variants are capable of generating a single-stranded DNA break (nick) at a specific position based on the gRNA-defined target sequence, resulting in the repair of the unedited strand and ultimately a T-to-C change on the unedited strand. In some embodiments, the A-to-G base editors also contain an inhibitor of inosine base excision repair, such as a uracil glycosylase inhibitor (UGI) domain or a catalytically inactive inosine-specific nuclease. Without being bound by any particular theory, the UGI domain or the catalytically inactive inosine-specific nuclease can inhibit or prevent base excision repair of the deaminated adenosine residue (e.g., inosine), which can improve the activity or efficiency of the base editor.

[0125] Base editors containing an adenosine deaminase can act on any polynucleotide, including DNA, RNA, and DNA-RNA hybrids. In certain embodiments, base editors containing an adenosine deaminase can deaminate the target A of a polynucleotide (including RNA). For example, the base editor can contain an adenosine deaminase domain capable of deaminating the target A of an RNA polynucleotide and / or a DNA-RNA hybrid polynucleotide. In one embodiment, the adenosine deaminase incorporated into the base editor contains all or part of an adenosine deaminase acting on RNA (ADAR, e.g., ADAR1 or ADAR2). In another embodiment, the adenosine deaminase incorporated into the base editor contains all or part of an adenosine deaminase acting on tRNA (ADAT). Base editors containing an adenosine deaminase domain can also be capable of deaminating the A nucleobase of a DNA polynucleotide. In one embodiment, the adenosine deaminase domain of the base editor contains all or part of an ADAT that contains one or more mutations that allow the ADAT to deaminate the target A in DNA.

[0126] The adenosine deaminase can be derived from any suitable organism (e.g., Escherichia coli). In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase that includes one or more mutations corresponding to the amino acid site mutations provided herein (e.g., mutations occurring in any of SEQ ID NOs: 1-47 (e.g., SEQ ID NOs: 2-47)). The corresponding residues in any homologous protein can be identified, for example, by sequence alignment and determination of homologous residues. Mutations can be generated accordingly in any naturally occurring adenosine deaminase corresponding to any of the mutations described herein.

[0127] Adenosine deaminase

[0128] Some aspects of the present application provide an adenosine deaminase. In some embodiments, the adenosine deaminase provided herein is capable of deaminating adenine. In some embodiments, the adenosine deaminase provided herein is capable of deaminating adenine in the deoxyadenosine residues of DNA. As mentioned herein, the term "adenosine deaminase" refers to a deaminase capable of deaminating adenine in the deoxyadenosine residues of DNA, and in certain cases, the adenosine deaminase can also deaminate cytosine in the deoxycytidine residues of DNA. In some embodiments, the adenosine deaminase provided herein is capable of deaminating cytosine. In some embodiments, the adenosine deaminase provided herein is capable of deaminating cytosine in the deoxycytidine residues of DNA, but with lower efficiency. The adenosine deaminase can be derived from any suitable organism (e.g., Escherichia coli). In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase that contains one or more mutations corresponding to the amino acid site mutations provided herein. Those skilled in the art will be able to identify the corresponding residues in any homologous protein and the respective encoding nucleic acids by methods well known in the art (e.g., by sequence alignment and determination of homologous residues). Thus, those skilled in the art will be able to generate mutations corresponding to any of the mutations described herein in any naturally occurring adenosine deaminase. In some embodiments, the adenosine deaminase is from a prokaryote. In some embodiments, the adenosine deaminase is from a bacterium. In some embodiments, the adenosine deaminase is from Escherichia coli, Staphylococcus aureus, Salmonella typhi, Shewanella putrefaciens, Haemophilus influenzae, Caulobacter crescentus, or Bacillus subtilis. In some embodiments, the adenosine deaminase is from Escherichia coli.

[0129] In some embodiments, the adenosine deaminase is a TadA deaminase. In some embodiments, the TadA deaminase is a TadA variant.

[0130] In some embodiments, the adenosine deaminase comprises any amino acid sequence shown in any of SEQ ID NOs: 1-47 (e.g., SEQ ID NOs: 2-47) or an amino acid sequence that is at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% identical to any adenosine deaminase provided herein. It should be understood that the adenosine deaminases provided herein may comprise one or more mutations (e.g., any of the mutations provided herein). This application provides any deaminase domain having a certain percentage identity plus any of the mutations described herein or combinations thereof. In some embodiments, the adenosine deaminase comprises an amino acid sequence having 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 21, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more mutations compared to any of the amino acid sequences shown in SEQ ID NOs: 1-47 or any adenosine deaminase provided herein. In some embodiments, the adenosine deaminase comprises an amino acid sequence having at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160 or at least 170 identical consecutive amino acid residues compared to any of the amino acid sequences shown in SEQ ID NOs: 1-47 or any adenosine deaminase provided herein.

[0131] In some embodiments, the deaminases provided herein are capable of deaminating adenine. In some embodiments, the adenosine deaminases provided herein are capable of deaminating adenine in the deoxyadenosine residues of DNA.

[0132] In some embodiments, the deaminases provided herein are capable of deaminating cytosine or 5-methylcytosine to uracil or thymine. In some embodiments, the deaminases provided herein are capable of deaminating cytosine in DNA.

[0133] In some embodiments, the adenosine deaminase comprises one or more amino acid substitutions selected from the group consisting of D12H, H21D or H21Q, A61C, V95T, S156K, C159R, E164V or E164W, E167W numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having an amino acid sequence of any one of SEQ ID NOs: 1-47 (e.g., SEQ ID NOs: 2-47). In one embodiment, the adenosine deaminase consists of an amino acid sequence of any one of SEQ ID NOs: 1-47 (e.g., SEQ ID NOs: 2-47).

[0134] In some embodiments, the adenosine deaminase comprises the V95T amino acid substitution numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 2.

[0135] In some embodiments, the adenosine deaminase comprises the V95T and A61C amino acid substitutions numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 3.

[0136] In some embodiments, the adenosine deaminase comprises the V95T and C159R amino acid substitutions numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 4.

[0137] In some embodiments, the adenosine deaminase comprises the V95T and D12H amino acid substitutions numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 5.

[0138] In some embodiments, the adenosine deaminase comprises the V95T and E164V amino acid substitutions numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 6.

[0139] In some embodiments, the adenosine deaminase comprises the V95T and E164W amino acid substitutions numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 7.

[0140] In some embodiments, the adenosine deaminase comprises V95T and E167R amino acid substitutions relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 8.

[0141] In some embodiments, the adenosine deaminase comprises V95T and E167W amino acid substitutions relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 9.

[0142] In some embodiments, the adenosine deaminase comprises V95T, E164V and E167W amino acid substitutions relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 10.

[0143] In some embodiments, the adenosine deaminase comprises V95T and H21D amino acid substitutions relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 11.

[0144] In some embodiments, the adenosine deaminase comprises V95T, C159R, E164V and E167W amino acid substitutions selected from relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 12.

[0145] In some embodiments, the adenosine deaminase comprises A61C, E164V and E167W amino acid substitutions relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 13.

[0146] In some embodiments, the adenosine deaminase comprises C159R, E164V and E167W amino acid substitutions relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 14.

[0147] In some embodiments, the adenosine deaminase comprises C159R and E167W amino acid substitutions relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 15.

[0148] In some embodiments, the adenosine deaminase comprises the H21Q and A61C amino acid substitutions numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 16.

[0149] In some embodiments, the adenosine deaminase comprises the S155R and K169N amino acid substitutions numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 17.

[0150] In some embodiments, the adenosine deaminase comprises one or more amino acid substitutions from the group consisting of S156N and E164V numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 18.

[0151] In some embodiments, the adenosine deaminase comprises one or more amino acid substitutions from the group consisting of V95T, S156K, C159R, E164V, and E167W numbered relative to SEQ ID NO: 1. In one embodiment, the adenosine deaminase is an adenosine deaminase having the amino acid sequence of SEQ ID NO: 19.

[0152] It should be appreciated that any of the mutations provided herein can be introduced into other adenosine deaminases, such as Staphylococcus aureus TadA (saTadA), or other adenosine deaminases (e.g., bacterial adenosine deaminases). Those skilled in the art will appreciate how to identify amino acid residues homologous to the mutated residues in SEQ ID NOs: 1-47 in other adenosine deaminases. Thus, any of the mutations in SEQ ID NOs: 1-47 can be made in other adenosine deaminases. It should also be understood that any of the mutations provided herein can be made in another adenosine deaminase singly or in any combination.

[0153] Table A. Amino acid sequences of adenosine deaminases

[0154] Name SEQ ID NO Amino acid sequence T8.4-V0 1 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYREPREQKNALKRASQDPRSSNA T8.4-V1 2 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYREPREQKNALKRASQDPRSSNA T8.4-V2 3 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRCIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYREPREQKNALKRASQDPRSSNA T8.4-V3 4 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLRDFYREPREQKNALKRASQDPRSSNA T8.4-V4 5 MYNAPRFSTGVHALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYREPREQKNALKRASQDPRSSNA T8.4-V5 6 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYRVPREQKNALKRASQDPRSSNA T8.4-V6 7 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYRWPREQKNALKRASQDPRSSNA T8.4-V7 8 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYREPRRQKNALKRASQDPRSSNA T8.4-V8 9 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYREPRWQKNALKRASQDPRSSNA T8.4-V9 10 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYRVPRWQKNALKRASQDPRSSNA T8.4-V10 11 MYNAPRFSTGVDALSETELNDEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYREPREQKNALKRASQDPRSSNA T8.4-V11 12 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLRDFYRVPRWQKNALKRASQDPRSSNA T8.4-V12 13 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRCIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYRVPRWQKNALKRASQDPRSSNA T8.4-V13 14 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLRDFYRVPRWQKNALKRASQDPRSSNA T8.4-V14 15 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLRDFYREPRWQKNALKRASQDPRSSNA T8.4-V15 16 MYNAPRFSTGVDALSETELNQEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRCIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSSLLCDFYREPREQKNALKRASQDPRSSNA T8.4-V16 17 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCRSLLCDFYREPREQNNALKRASQDPRSSNA T8.4-V17 18 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSNLLCDFYRVPREQKNALKRASQDPRSSNA T8.4-V18 19 MYNAPRFSTGVDALSETELNHEYWMRHALNLAQRAREEGEVPVGAVLVYQDKVIGEGWNRAIGLHDPTAHAEIMALRQGGMVLQNYRLIDATLYTTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLINVLNYPGMNHRVEVTEGVLAESCSKLLRDFYRVPRWQKNALKRASQDPRSSNA A-DN-1 20 SGTDEDWMARVLELAEEARAAREVPVGAILVLNNELVGEGWNRAIILNDPTAHAEIIALNEGKKKLKSYRLIGATLYVNFEPCVMCANAMIHARIGKVVYGVRNSKRGAAGSTQNILNYPGMNHKVKVVSGVLAEEVGKLLCDFYRMPRQVFNAQKKAQSSIN A-DN-2 21 SGTREDWMERVIELARKAREAREVPVGAILVLNNELVGEGWNRAIVLNDPTAHAEIIALEEGKKKLKSYRLIGATLYVNFEPCVMCALAMIHARIGEVVYGVRNSKRGAAGSVRNILNYPGMNHKVKVVSGVLAEEVGRLLCDFYRMPRQVFNAQKKAQSSIN A-DN-3 22 PGTLEDWMRRALELARKAREAREVPVGAILVLDGKLIGEGWNRAIVLNDPTAHAEIIALEKGKEVTKSYRLIGATLYVNFEPCVMCALAMIHARIGRVVYGVRNSKRGAAGSLMNVLNYPGMNHRVEIVEGVLADEVGKLLCDFYRMPRQVFNAQKKAQSSIN A-DN-4 23 AGTDEDWMARALALAEQARAAREVPVGAILVLNDELVGEGWNRAIILNDPTAHAEIIALEEGKKKTNSYRLIGATLYVNFEPCVMCALAMIHARIGKVVYGVRNSKRGAAGSLMNVLNYPGMNHRVEVVSGVLAEEVGKLLCDFYRMPRQVFNAQKKAQSSIN A-DN-5 24 AGTEEDWMRRALALAEQARAAREVPVGAILVLDGELVGEGWNRAIVLHDPTAHAEIIALREAGKKLQNYRLIGATLYVNFEPCVMCAGAMIHSRIGKVVYGVRNSKRGAAGSLMNVLNYPGMNHRVEVVSGVLAEEVGRLLCDFYRMPRQVFNAQKKAQSSIN A-DN-6 25 AGTEEDWMRRALALAEQARQAREVPVGAILVLDGELVGEGWNRAIVLHDPTAHAEIIALREAGRRTQNYRLIGATLYVNFEPCVMCAGAMIHSRIGKVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEVVEGVLADEVGKLLCDFYRMPRQVFNAQKKAQSSIN A-DN-7 26 SGTDEDWMRRALELAEQAREAREVPVGAVLVLDGELVGEGWNRAIVLHDPTAHAEIIALRQGGERLQNYRLIGATLYVTFEPCVMCAGAMIHSRIGRVVYGVRNSKRGAAGSLMNVLNYPGMNHRVEVVSGVLAKEVGKLLCDFYRMPRQVFNAQKKAQSSIN A-DN-8 27 MGTDEDWMERALELAKKAREAREVPVGAVLVLNDELIGEGWNRAIVLHDPTAHAEIIALREGGRKTQNYRLIGATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEVVSGVLAEEVGKLLCDFYRMPRQVFNAQKKAQSSIN A-DN-9 28 AGTDEDWMREALALAEKAREAREVPVGAVLVLDNEVIGEGWNRAIVLHDPTAHAEIIALRKGGERMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEIVEGILAEECGKLLCDFYRMPRQVFNAQKKAQSSIN A-DN-10 29 PGTDEYWMRRALELARRAREAREVPVGAVLVLNNKVIGEGWNRAIILHDPTAHAEIIALRKGGERMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEIVSGILAEECGRLLCDFYRMPRQVFNAQKKAQSSIN A-DN-11 30 SGTDEDWMRRALRLAEKAREAREVPVGAVLVLDNEVIGEGWNRAIVLHDPTAHAEIIALRKGGEKMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEIVSGILAEECGKLLCDFYRMPRQVFNAQKKAQSSIN A-DN-12 31 SGSDEYWMRLALELAEKAREAREVPVGAVLVLNNEVIGEGWNRAIVLHDPTAHAEIIALRQGGEKMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHKVEIVEGILAEECGKLLCDFYRMPRQVFNAQKKAQSSIN A-DN-13 32 MTDADWMARALELARKALEAGEVPVGAVLVLDGEVVGEGFNNAIRLNDPTAHAEIIALRKGGKKKKDYRLIDATLYVTFEPCVMCAGAMIHARIGKVVFGVRNAKTGAAGSLMDVLKHPGMNHQVEVVSGVLADECAELLCEFFRMPRRVFNAQKKAQSSTD A-DN-14 33 MSDADWMALALEEARKALAAGEVPVGAVLVLDGEVVGRGYNNAIRLNDPTAHAEIIAIRKGGKKKKDYRLIDATLYVTFEPCVMCAGAMIHARIKRVVFGVRNAKTGAAGSLMDVLHHPGMNHRVEVVSGVLADECAALLCEFFRMPRRVFNAQKKAQSSTD A-DN-15 34 MTHEDWMRMALEEARKAKEAREVPVGAVLVLNNEVIGRGYNSAIRLNDPTAHAEIIALRKGGKKMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLNHPGMNHKVEIVSGILADECAELLCDFFRMPRRVFNAQKKAQSSTD A-DN-16 35 MSHEDWMRLALEEARKALEAREVPVGAVLVLNNEVIGRGYNNAIKLNDPTAHAEIIALRKGGKTMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLRHPGMNHKVEIVSGILAEECAELLCKFFRMPRRVFNAQKKAQSSTD A-DN-17 36 MTDADWMRRALAEARKALAAGEVPVGAVLVLDGEVVGEGYNAAIRLHDPTAHAEIIALRKGGLTKQNYRLIDATLYVTFEPCVMCAGAMIHSRIARVVFGVRNAKTGAAGSLMDVLNHPGMNHKVEVVSGVLADEAAALLCDFFRMPRRVFNAQKKAQSSTD A-DN-18 37 MSDEDWMRLALEEARKALEAGEVPVGAVLVLDGEVVGRGYNSAIRLHDPTAHAEIIALRKGGLRKQNYRLIDAVLYVTFEPCVMCAGAMIHSRIKKVVFGVRNAKTGAAGSLMDVLHHPGMNHRVEVVSGVLAEECAELLCEFFRMPRRVFNAQKKAQSSTD A-DN-19 38 MTHEDWMRLALAEAEKALADREVPVGAVLVLNNEVIGSGYNNAIILHDPTAHAEIIALRKGGLKMQNYRLIDATLYVTFEPCVMCAGAMIHSRIKRVVFGVRNAKTGAAGSLMDVLNHPGMNHKVEIVSGILADECAELLCRFFRMPRRVFNAQKKAQSSTD A-DN-20 39 MSHEDWMRLALEEARKALEAREVPVGAVLVLNNEVIGRGYNNAITLHDPTAHAEIIALRKGGLKMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAAGSLMDVLNHPGMNHKVEIVSGILADECAELLCEFFRMPRRVFNAQKKAQSSTD A-DN-21 40 MTDEDWMAEALKLAEKAASLREVPVGAVLVLNDEIIGKGYNAAITLSDTTAHAEINALREGGKKTKNYRLYDATLYSTFEPCVMCAGAMLHARIARYVYGVRNAKTGAAGSLMDVLNHPGMNHRVEVVSGVRAEECAELLCRFFRMPRSVFNAQKKAQSSTD A-DN-22 41 MTDEDWMAIALEEAKKAAELREVPVGAVLVLNDEIIGTGYNAAIVLNDTTAHAEINALRDGGKKLKDYRLYDATLYSTFEPCVMCAGAMLHARIKRVVYGVRNAKTGAAGSLMDVLNHPGMNHRVEVVRGVRADEAAELLCRFFRMPRSVFNAQKKAQSSTD A-DN-23 42 MTHEDWMALALELAKKAAANREVPVGAVLVLDNRVIGKGYNAAITLNDPTAHAEINALRDGGKVMQNYRLYDATLYSTFEPCVMCAGAMLHSRIKRVVFGVRNAKTGAAGSLMDVLNHPGMNHRVEIVEGILADECAELLCRFFRMPRTVFNAQKKAQSSTD A-DN-24 43 MTHEDWMRIALELAKIAEENREVPVGAVLVLNNEVIGKGYNAAITLNDPTAHAEINALRDGGKKMQNYRLYDATLYSTFEPCVMCAGAMLHSRIKRVVFGVRNAKTGAAGSLMDVLNHPGMNHRVEIVEGILAEECAELLCRFFRMPRSVFNAQKKAQSSTD A-DN-25 44 MTDEDWMALALELAEKAAALREVPVGAVLVLDDEVIGRGYNAAIVLHDPTAHAEINALRDGGLRLQNYRLYDATLYSTFEPCVMCAGAMIHSRIKRYVYGVRNAKTGAAGSLMDVLNHPGMNHRVEVVRGVRAEECAELLCRFFRMPRSVFNAQKKAQSSTD A-DN-26 45 MTDEDWMARALELAEKAAAAREVPVGAVLVLDDEIIGEGYNAAITLHDPTAHAEILALRQGGLKLQNYRLYDATLYSTFEPCVMCAGAMIHSRIKRYVYGVRNAKTGAAGSLMDVLNHPGMNHRVEVVRGVRAEECAELLCRFFRMPRSVFNAQKKAQSSTD A-DN-27 46 MTHEDWMRIALELAEKALAAREVPVGAVLVLNNEVIGKGWNAAITLHDPTAHAEINALRHGGLVMQNYRLYDATLYSTFEPCVMCAGAMIHSRIKRVVFGVRNAKTGAAGSLMDVLNHPGMNHRVEIVEGILAEECAELLCRFFRMPRSVFNAQKKAQSSTD A-DN-28 47 MTHEDWMRIALELAEKAAAAREVPVGAVLVLNNEVIGKGWNRAITLHDPTAHAEINALRDGGLKMQNYRLYDATLYSTFEPCVMCAGAMIHSRIKRVVFGVRNAKTGAAGSLMDVLNHPGMNHRVEIVEGILAEECAELLCRFFRMPRSVFNAQKKAQSSTD

[0155] Cytidine deaminase and C to T (U) editing

[0156] "Cytidine deaminase" means a polypeptide or fragment thereof capable of catalyzing a deamination reaction that converts an amino group to a carbonyl group. In one embodiment, the cytidine deaminase converts cytosine to uracil, or 5-methylcytosine to thymine. The cytidine deaminases provided herein (e.g., engineered cytidine deaminases) can be from any organism, such as bacteria.

[0157] In some embodiments, the cytidine deaminase of the base editor can include all or part of the apolipoprotein B mRNA editing complex (APOBEC) family deaminases or variants thereof. APOBEC is a family of evolutionarily conserved cytidine deaminases. Members of this family are C-to-U editing enzymes. In some embodiments, the cytidine deaminase includes, but is not limited to: APOBEC family members, including but not limited to: APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D (now called "APOBEC3E"), APOBEC3F, APOBEC3G, APOBEC3H, APOBEC4, activation-induced (cytidine) deaminase (AID), hAPOBEC1 (derived from Homo sapiens), rAPOBEC1 (derived from Rattus norvegicus), ppAPOBEC1 (derived from Pongo pygmaeus), AmAPOBEC1 (BEM3.31) (derived from Alligator mississippiensis), ocAPOBEC1 (derived from Oryctolagus cuniculus), SsAPOBEC2 (BEM3.39) (derived from Sus scrofa), hAPOBEC3A (derived from Homo sapiens), maAPOBEC1 (derived from Mesocricetus auratus), mdAPOBEC1 (derived from Monodelphis domestica); cytidine deaminase 1 (CDA1), hA3A (i.e., APOBEC3A derived from Homo sapiens), RrA3F (BEM3.14) (which is APOBEC3F derived from Rhinopithecus roxellana); PmCDA1 (derived from Petromyzon marinus); AID (activation-induced cytidine deaminase; AICDA), derived from mammals (e.g., human, pig, cow, horse, monkey, etc.); hAID (derived from Homo sapiens); and FENRY.

[0158] The term "deaminase" or "deaminase domain", as used herein, refers to a protein or enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase or deaminase domain is a cytidine deaminase that catalyzes the hydrolytic deamination of cytidine to uridine or deoxycytidine to deoxyuridine. In some embodiments, the deaminase or deaminase domain is a cytosine deaminase that catalyzes the hydrolytic deamination of cytosine to uracil, thereby enabling editing from C:G to T:A (i.e., C to T editing).

[0159] In some embodiments, the cytidine deaminase is a de novo designed TadA variant having cytidine deaminase activity. In some embodiments, the cytidine deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 48 - 63

[0160] In some embodiments, the cytidine deaminase is generated by directed evolution based on SEQ ID NO: 64. In some embodiments, the cytidine deaminase comprises one or more substitutions selected from the group consisting of A17N, G64C, Y70G or Y70H or Y70K, L73F, V79T, F81L, M91L, H93N, G102S, D112N, A119G, N124V or N124P, R130V, S139K, A140R, C143A, R151M, K153G relative to SEQ ID NO: 64, wherein the amino acid positions are numbered relative to SEQ ID NO: 64. In some embodiments, the cytidine deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to any one of SEQ ID NOs: 64 - 99.

[0161] In some embodiments, the cytidine deaminase comprises the Y70H and M91L amino acid substitutions numbered relative to SEQ ID NO: 64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO: 74.

[0162] In some embodiments, the cytidine deaminase comprises the Y70H, M91L and F81L amino acid substitutions numbered relative to SEQ ID NO: 64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO: 78.

[0163] In some embodiments, the cytidine deaminase comprises amino acid substitutions Y70H, M91L, F81L, C143A, G64C, and L73F numbered relative to SEQ ID NO: 64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO: 80.

[0164] In some embodiments, the cytidine deaminase comprises amino acid substitutions Y70H, M91L, F81L, G64C, and L73F numbered relative to SEQ ID NO: 64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO: 81.

[0165] In some embodiments, the cytidine deaminase comprises amino acid substitutions Y70H, M91L, F81L, and N124V numbered relative to SEQ ID NO: 64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO: 85.

[0166] In some embodiments, the cytidine deaminase comprises amino acid substitutions Y70H, M91L, F81L, G64C, L73F, and N124V numbered relative to SEQ ID NO: 64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO: 91.

[0167] In some embodiments, the cytidine deaminase comprises amino acid substitutions Y70H, M91L, F81L, G64C, L73F, R130V, A140R, D112N, and N124V numbered relative to SEQ ID NO: 64. In one embodiment, the cytidine deaminase is a cytidine deaminase having the amino acid sequence of SEQ ID NO: 97.

[0168] In some embodiments, the cytidine deaminase comprises the amino acid sequence of any one of SEQ ID NOs: 48 - 63 and SEQ ID NOs: 64 - 99. Preferably, the cytidine deaminase consists of the amino acid sequence of any one of SEQ ID NOs: 48 - 63 and SEQ ID NOs: 64 - 99.

[0169] Table B - 1. Amino acid sequences of de novo designed TadA variants with cytidine deaminase activity

[0170] Name SEQ ID NO Amino acid sequence C-DN-1 48 GMHSDEEWMAIALEEARKARDAGKAPVGAVLVLDDEIIGKGYNAAIDLNDPTAHAEIIAIREGGKKLKDYRLIDATLYVTFEPCVMCAGALLNARIARVVYGVRNSKRGAAGSLMNVLNYPGLNHRVEVRSGVLADEALALLCDFYRMPRQVFNAQKKAQSSIN C-DN-2 49 GMHSDEYWMDLALEEARKAREAGKAPVGAVLVLDDEVIGRGYNSAIDLNDPTAHAEIIALREGGKRLKDYRLIDATLYVTFEPCVMCAGALLNARIKRVVYGVRNSKRGAAGSLMNVLNYPGLNHRVEVVSGVKAEEALQLLCDFYRMPRQVFNAQKKAQSSIN C-DN-3 50 GMHSHEEWMKIALAEAEKARKERKAPVGAVLVLNNEVIGKGYNSAIELNDPTAHAEIIALREGGKKMQDYRLIDATLYVTFEPCVMCAGAMLNSRIKRVVFGVRNSKRGAAGSLMNVLNYPGLNHKVEIVEGILADECLELLCDFYRMPRQVFNAQKKAQSSIN C-DN-4 51 SSHSHEYWMEKALALARKAREERKAPVGAVLVLDNEVIGEGYNSAIDLNDPTAHAEIIALREGGKKMQDYRLIDATLYVTFEPCVMCAGAMLNSRIKRVVFGVRNSKRGAAGSLMNVLNYPGLNHRVEIVSGILADECLGLLCDFYRMPRQVFNAQKKAQSSIN C-DN-5 52 MTHSDEEWMRVALAEAEKAREAGKAPVGAVLVLNDEIIGRGYNSAIDLHDPTAHAEIIALRQGGLYLQNYRLIDATLYVTFEPCVMCAGALINSRIRRVVYGVRNSKRGAAGDLMNVLNYPGMNHRVEVRSGVLAEEALALLCEFYRMPRQVFNAQKKAQSSIN C-DN-6 53 HMHSDEHWMRLALEEARKAREEGKAPVGAVLVLDDEVIGRGYNNAISLHDPTAHAEIIALRQGGLRRQNYRLIDATLYVTFEPCVMCAGALINSRIKRVVYGVRNSKRGAAGGLMNVLNYPGMNHKVEVVSGVLAEEALALLCEFYRMPRQVFNAQKKAQSSIN C-DN-7 54 GSHSHEHWMRLALAEARKARDARKAPVGAVLVLNNEVIGRGYNSAIALHDPTAHAEIIALRQGGLKMQNYRLIDATLYVTFEPCVMCAGAMINSRIKRVVFGVRNSKRGAAGSLMNVLNYPGMNHKVEIVEGILADECLELLCEFYRMPRQVFNAQKKAQSSIN C-DN-8 55 DMHSHEYWMRLALEEARKAREERKAPVGAVLVLDNEVIGKGYNSAIKLHDPTAHAEIIALRQGGLKMQNYRLIDATLYVTFEPCVMCAGAMINSRIKRVVFGVRNSKRGAAGSLMNVLNYPGMNHKVEIVSGILADECLDLLCEFYRMPRQVFNAQKKAQSSIN C-DN-9 56 SSHSDEHWMAKALELAEKARAAGHVPVGAVLVLNDEIVGTGFNAAKTLNDPTAHAEIIAIRKGGKTLKDYRLWGATLYTTFEPCVMCAGAMIHARIKRVVFGVCNAKTHACMGMMDVLGLPGLNHRVEVVSGVLADEALDLLCRFFRMPRRVFNAQKKAQSSTD C-DN-10 57 ASHSDAHWMARALELAEKAREDGHVPVGAVLVLDDEVIGEGYNAAKTLNDPTAHAEILALRAGGKRLRDYRLWGATLYTTFEPCVMCAGAMIHARIKRVVFGVCNAKTHACMERMDVLGTPGLNHRVEVVRGVLADECLELLCRFFRMPRRVFNAQKKAQSSTD C-DN-11 58 GLHSDEEWMAIALREAEKARENRHVPVGAVLVLDNEVIGTGYNAAKTLNDPTAHAEILALRAGGKKMQDYRLWGATLYTTFEPCVMCAGAMIHSRIERVVFGVCNAKTHACMSRMDVLGLPGLNHKVEIVSGILADECLDLLCRFFRMPRRVFNAQKKAQSSTD C-DN-12 59 GAHSDEEWMAKALELAEKAREARHVPVGAVLVLDNEVIGTGFNAAKTLNDPTAHAEIIALRRGGQTMQDYRLWGATLYTTFEPCVMCAGAMIHSRIKRVVFGVCNAKTHACMSMMDVLGLPGLNHRVEIVAGILADECLDLLCRFFRMPRRVFNAQKKAQSSTD C-DN-13 60 SSHSDAHWMARALELAEKAREAGHVPVGAVLVLDDEVIGEGYNAAKTLHDPTAHAEIIALRQGGLRLQNYRLWGATLYTTFEPCVMCAGAMIHSRIKRVVFGVCNAKTHACMGLMDVLGHPGMNHKVEVVGGVLADECLDLLCRFFRMPRRVFNAQKKAQSSTD C-DN-14 61 ASHSDEHWMRRALELAERAREAGHVPVGAVLVLDDEVVGEGYNAAKTLHDPTAHAEIIALRQGGLRRQNYRLWGATLYTTFEPCVMCAGAMIHSRIARVVYGVCNAKTHACMGLMDVLGHPGMNHRVEVVGGVLADEALELLCRFFRMPRRVFNAQKKAQSSTD C-DN-15 62 DMHSDEYWMRRALELAEKARERRHVPVGAVLVLNNEVIGEGFNAAKSLHDPTAHAEIIALRKGGLKMQNYRLWGATLYTTFEPCVMCAGAMIHSRIKRVVFGVCNAKTHACMSLMDVLGHPGMNHRVEIVRGILADECLDLLCRFFRMPRRVFNAQKKAQSSTD C-DN-16 63 GSHSDEYWMRRALELAEKAREDRHVPVGAVLVLNNEVIGEGYNAAKTLHDPTAHAEIIALRKGGLKMQNYRLWGATLYTTFEPCVMCAGAMIHSRIKRVVFGVCNAKTHACMSLMDVLGHPGMNHKVEIVRGILREECEELLCRFFRMPRRVFNAQKKAQSSTD

[0171] Table B - 2. Amino acid sequences of cytidine deaminases directionally evolved based on SEQ ID NO: 64

[0172] Name SEQ ID NO Amino acid sequence T88.74-V0 64 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNYRLLECTLYVTFEPCMMCAGAMVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V1 65 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNGRLLECTLYVTFEPCMMCAGAMVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V2 66 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNYRLLECTLYVTFEPCMMCAGAMVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACKAQLCRFFRRPRRVKNAARKQAAGG T88.74-V3 67 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNYRLLECTLYVTFEPCMMCAGAMVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLARFFRRPRRVKNAARKQAAGG T88.74-V4 68 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNYRLLECTLYVTFEPCMMCAGAMVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V5 69 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNYRLLECTLYVTLEPCMMCAGAMVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V6 70 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNYRLFECTLYVTFEPCMMCAGAMVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V7 71 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNYRLLECTLYTTFEPCMMCAGAMVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V8 72 MNEQDEYWMRRAMALANRAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNYRLLECTLYVTFEPCMMCAGAMVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V9 73 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNKRLLECTLYVTFEPCMMCAGAMVHSRIQRLVYSASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V10 74 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTFEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V11 75 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNKRLLECTLYVTFEPCMMCAGAMVNSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V12 76 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTFEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLARFFRRPRRVKNAARKQAAGG T88.74-V13 77 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLLECTLYVTFEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V14 78 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V15 79 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLFECTLYVTFEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V16 80 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLARFFRRPRRVKNAARKQAAGG T88.74-V17 81 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V18 82 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V19 83 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLLECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V20 84 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLARFFRRPRRVKNAARKQAAGG T88.74-V21 85 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLVHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V22 86 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLPHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V23 87 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRMVGNAARKQAAGG T88.74-V24 88 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVNSVLQLLGHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V25 89 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVNSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V26 90 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQAGAHLGNHRLLECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWVAGVLADACSRQLCRFFRRPRRVKNAARKQAAGG T88.74-V27 91 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLVHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V28 92 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLPHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V29 93 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWVAGVLADACSRQLCRFFRRPRRVKNAARKQAAGG T88.74-V30 94 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRMVGNAARKQAAGG T88.74-V31 95 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVNSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V32 96 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLVHRVDWVAGVLADACSRQLCRFFRRPRRVKNAARKQAAGG T88.74-V33 97 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVNSVLQLLAHPGLVHRVDWVAGVLADACSRQLCRFFRRPRRVKNAARKQAAGG T88.74-V34 98 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYSASNAKTGAVDSVLQLLAHPGLNHRVDWRAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG T88.74-V35 99 MNEQDEYWMRRAMALAARAEQEGEVPVGALVVYHGDCVGEGWNRSIGHHDATAHAEIMALRQACAHLGNHRLFECTLYVTLEPCMMCAGALVHSRIQRLVYGASNAKTGAVDSVLQLLAHPGLNHRVDWVAGVLADACSAQLCRFFRRPRRVKNAARKQAAGG

[0173] Table C. Coding nucleotide sequences of adenosine deaminase and cytidine deaminase

[0174]

[0175] Other sequences used in this application, such as linker sequences or sgRNA sequences, are listed in Tables D and E.

[0176] Table D. Linker sequences

[0177] SEQ ID NO Amino acid sequence 104 SGGSSGGSSGSETPGTSESATPESSGGSSGGS 105 SGGS 106 PAPAPAP 107 SGSETPGTSESATPES

[0178] Table E. sgRNA sequences

[0179] Name SEQ ID NO Nucleotide sequence gRNA_1 108 CCCATACCTTGGAGCAACGG gRNA_2 109 CCCGCACCTTGGCGCAGCGG sgRNA_3 110 GGAATCCCTTCTGCAGCACC CAM-gRNA1 111 TGATCCGAACGTGGCCAATA CAM-gRNA2 112 TACGGCGCGGTGCACCTGGA

[0180] Table F. Other sequences

[0181]

[0182] Splitting system

[0183] In some embodiments, the base editors disclosed herein comprise "split" systems, where the Cas protein, deaminase, or both are split into two components and can be reconstituted to form a functional protein to perform base editing. "Split Cas9 protein", "split Cas9", or "split deaminase" refers to a Cas9 or deaminase protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal portion and the C-terminal portion of the Cas9 protein or deaminase can be spliced to form a "reconstituted" Cas9 protein or deaminase. In some embodiments, the Cas9 protein or deaminase is split into two fragments within a disordered region of the protein, e.g., as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351: 867-871. PDB file: 5F9R, each of which is incorporated herein by reference.

[0184] Embedded system

[0185] In some embodiments, the base editor systems disclosed herein comprise an embedded system, where a deaminase (e.g., adenosine deaminase or cytidine deaminase) is inserted into Cas9. In some embodiments, the present application provides a fusion protein comprising a nuclease protein guided by a nucleic acid (e.g., Cas9 or any nuclease disclosed herein) and an embedded nucleotide deaminase protein domain. In some embodiments, the embedded nucleotide deaminase protein domain is an embedded adenosine deaminase protein domain. In some embodiments, the embedded adenosine deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-47 (e.g., SEQ ID NOs: 2-47). In some embodiments, the embedded nucleotide deaminase protein domain is an embedded cytidine deaminase protein domain. In some embodiments, the embedded cytidine deaminase comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NOs: 48-63 or 64-99 (e.g., SEQ ID NOs: 48-63 or 65-99).

[0186] Examples

[0187] The compositions and methods disclosed herein are further described in the following examples, which are illustrative only and are not meant to limit the scope of the invention thereto. The scope of the invention is defined by the claims.

[0188] Example 1. Evaluation of Engineered TadA Variants in sgRNA Library Cells

[0189] The TadA deaminase is from nature. Mutations were introduced into it by rational design to obtain the DNA adenosine deaminase activity mutant T8.4-V0 (SEQ ID NO: 1) with high activity. Subsequently, directed evolution was used to further improve its activity, and a series of TadA mutants were obtained by introducing single or multiple mutations (Table 1).

[0190] Table 1. Mutations Contained in Different Adenosine Deaminase Variants Relative to T8.4-V0 (SEQ ID NO: 1)

[0191] Variant name Mutation introduced compared to T8.4-V0 T8.4-V1 V95T T8.4-V2 V95T-A61C T8.4-V3 V95T-C159R T8.4-V4 V95T-D12H T8.4-V5 V95T-E164V T8.4-V6 V95T-E164W T8.4-V7 V95T-E167R T8.4-V8 V95T-E167W T8.4-V9 V95T-E164V-E167W T8.4-V10 V95T-H21D T8.4-V11 V95T-C159R-E164V-E167W T8.4-V12 A61C-E164V-E167W T8.4-V13 C159R-E164V-E167W T8.4-V14 C159R-E167W T8.4-V16 S155R-K169N T8.4-V17 S156N-E164V T8.4-V18 V95T-S156K-C159R-E164V-E167W

[0192] To systematically evaluate the ability of each mutant to perform base editing in mammalian cells, we used the sgRNA library cell line (Max W. Shen et al., Nature, 2018 Nov;563(7733):646-651. doi:10.1038 / s41586-018-0686-x, cloned and constructed by GenScript Biotech Corporation). This library cell is derived from HEK293T and integrates more than 10,000 target sites and their paired homologous sgRNAs.

[0193] The nucleotide sequence encoding ecTadA8e in the ABE8e expression vector (Michelle F. Richter et al., Nat Biotechnol, 2020 Jul;38(7):883-891. doi: 10.1038 / s41587-020-0453-z, gene cloning completed by GenScript Biotech Corporation) was replaced with the nucleotide sequence encoding the TadA variant designed by the inventors. A series of mRNAs encoding ABE were generated by in vitro transcription in the experiment and tested for activity in the expressed library cell line. HEK293T library cells (1x10 5Individuals) were seeded in 24-well plates (Corning). According to the manufacturer's protocol, 400 ng of mRNA encoding the ABE variant was transfected into the cells using Lipofectamine MessengerMAX (ThermoFisher, catalog number LMRNA003). After 72 hours, genomic DNA of the cells was extracted and the editing efficiency was analyzed by next-generation sequencing (NGS). The sgRNA detected in all samples was defined as the effective sgRNA, and the A-to-G editing efficiency of these paired target sites was analyzed. To quantify the potency of each variant, its average editing efficiency value was calculated and normalized against the average editing efficiency value of T8.4-V0. As shown in Table 2 and Figure 1 as shown, the editing efficiency of T8.4-V0 was slightly higher than that of ABE8.8 (a known ABE variant, used as a control; Gaudelli NM et al., Nat Biotechnol. 2020 Jul;38(7):892-900. doi: 10.1038 / s41587-020-0491-6.), and the average editing efficiency of many engineered variants of the present invention was higher than that of T8.4-V0, demonstrating that the engineered variants of the present invention are suitable as adenine base editors with improved activity.

[0194] Table 2. A-to-G editing efficiency of different TadA variants in library cells

[0195] Editor Average A-to-G editing efficiency ratios of A to G normalized by the activity of T8.4-V0 in the same experiment T8.4-V0 1.00 T8.4-V1 1.13 T8.4-V2 1.20 T8.4-V3 1.19 T8.4-V4 1.16 T8.4-V5 1.24 T8.4-V6 1.18 T8.4-V7 1.17 T8.4-V8 1.28 T8.4-V9 1.24 T8.4-V10 1.15 T8.4-V11 1.21 T8.4-V12 1.25 T8.4-V13 1.30 T8.4-V14 1.26 T8.4-V16 1.10 T8.4-V17 1.17 T8.4-V18 1.23 ABE8.8 (control) 0.97

[0196] Figure 1 Shows the overall A-to-G editing efficiency (i.e., the conversion rate of A:T to G:C) of some adenosine deaminase variants of the present invention (T8.4-V1, T8.4-V2, T8.4-V4, T8.4-V7, T8.4-V10, T8.4-V12) and the known ABE variant ABE8.8 (SEQ ID NO: 113) compared to T8.4-V0. Compared with T8.4-V0, some variants of the present invention (T8.4-V1, T8.4-V2, T8.4-V4, T8.4-V10) showed similar editing windows but higher editing efficiencies, and some other variants (T8.4-V7, T8.4-V12) showed wider editing windows and higher editing efficiencies. Here, the "editing window" refers to the position of A within the protospacer that can be edited. A "wider editing window" means more positions of A within the protospacer that can be edited.

[0197] Example 2. Evaluation of the activity of adenine base editor (ABE) variants in mice

[0198] Wild-type mice (mouse strain , purchased from Beijing Biocytogen, with an age of about 7 - 9 weeks and a body weight of 20 - 25 g when used in experiments) to further evaluate the in vivo editing effect of the base editor constructed based on the adenosine deaminase variant selected in Example 1. The present inventors designed a specific guide RNA (gRNA_1: CCCATACCTTGGAGCAACGG, SEQ ID NO: 108) for disrupting the mouse Pcsk9 gene. The mRNA encoding the ABE variant obtained by in vitro transcription and the chemically synthesized gRNA_1 were combined and encapsulated into lipid nanoparticles (LNP) according to the method described in patent application WO / 2023 / 185697A2. Then these formulated LNPs were injected into mice via the tail vein at a dose of 0.15 mg per kilogram of body weight (0.15 mg / kg body weight (mpk)).

[0199] One week after injection, mouse liver samples were collected. After DNA extraction, PCR amplification and NGS sequencing analysis were performed on the region near the editing site. As shown in Table 3, compared with T8.4-V0, the mice treated with the engineered mutants of the present invention showed higher A-to-G editing efficiency.

[0200] Table 3. DNA editing efficiency (%) achieved by the adenosine deaminase variant of the present invention in wild-type mice

[0201] Experiment batch Editor Number of animals Average A-to-G editing efficiency (%) B10 T8.4-V0 3 24.37 B10 T8.4-V12 3 45.61 B10 T8.4-V1 3 47.75 B10 T8.4-V8 3 43.59 B10 T8.4-V14 3 52.13 B10 T8.4-V13 3 45.12 B12 T8.4-V0 3 26.17 B12 T8.4-V3 3 35.04 B13 T8.4-V0 3 21.44 B13 T8.4-V2 3 25.18 B13 T8.4-V10 3 31.25

[0202] Example 3. Effect of different linker sequences on the activity of adenine base editor (ABE) variants

[0203] The adenosine deaminase variants tested in Examples 1 and 2 were linked to nCas9 (D10A) through a 32 - amino - acid linker (SGGSSGGSSGSETPGTSESATPESSGGSSGGS, SEQ ID NO: 104). To test the effect of different linker sequences on the editing activity of ABE variants, we tested the effect of changing the length and type of the linker or using no linker on the activity of T8.4-V0. Detection was performed using the same experimental system as in Example 1. The results are shown in Table 4. Taking the editor (T8.4-L0) obtained by linking T8.4-V0 and nCas9 (D10A) through the linker shown in SEQ ID NO: 104 as a control, the replacement of the linker did not reduce the base editing activity of T8.4-V0. Among them, T8.4-L3 obtained by linking T8.4-V0 and nCas9 with the linker sequence PAPAPAP (SEQ ID NO: 106) had significantly improved editing activity.

[0204] Table 4. A-to-G editing efficiency of different linker sequences in library cells

[0205] Editor Adapter sequence Average A-to-G editing efficiency ratios of A to G normalized by the activity of T8.4-L0 T8.4-L0 SGGSSGGSSGSETPGTSESATPESSGGSSGGS 1.00 T8.4-L1 No adapter 0.99 T8.4-L2 SGGS 0.99 T8.4-L3 PAPAPAP 1.04 T8.4-L4 SGSETPGTSESATPES 1.01

[0206] Example 4. Modifying Cas9 to Enhance Base Editing Activity

[0207] The editing activity of adenine base editors (ABEs) variants was enhanced by modifying nCas9 (nickase D10A, SEQ ID NO: 114). Rational design based on structure was used to introduce mutations into Cas9, and the activity was tested using the same experimental system as in Example 1. Using the editor T8.4-C0 composed of T8.4-V0 and non-mutated nCas9 as a control, the results are shown in Table 5. Introducing D700R or E1243R into nCas9 can improve the editing activity and efficiency of adenine base editors.

[0208] Table 5. A-to-G Editing Efficiency of T8.4-V0 and Mutated Cas9 in Library Cells

[0209] Editor Mutations introduced in nCas9 Average A-to-G editing efficiency ratios of A to G normalized by the activity of T8.4-C0 T8.4-C0 None 1.00 T8.4-C1 nCas9-D700R (SEQ ID NO: 115) 1.26 T8.4-C2 nCas9-E1243R (SEQ ID NO: 116) 1.25 T8.4-C3 nCas9-A987R (SEQ ID NO: 117) 0.45

[0210] Example 5. De Novo Design of Adenosine Deaminases with High Activity

[0211] Based on structural and sequence information, a new TadA with adenosine deaminase activity was designed. The polynucleotide sequence encoding ecTadA8e was replaced with the polynucleotide sequence encoding the newly designed TadA variant in the ABE8e expression vector to generate a series of ABE variants. The editing activities of these variants were measured in HEK293T cells as described below.

[0212] Twenty to twenty-four hours before transfection, HEK293T cells (1.5x10 4 cells) were seeded into 96-well plates (Corning). For targeted editing experiments, cells were transfected according to the manufacturer's protocol with 80 ng of the expression vector encoding the ABE variant, 40 ng of the spCas9 gRNA_2 (CCCGCACCTTGGCGCAGCGG, SEQ ID NO: 109) expression plasmid, and 3.6 μL of FuGENE HD transfection reagent (E2312, purchased from Promega). Seventy-two hours later, genomic DNA was extracted from the cells and the editing efficiency of the ABE variants was analyzed by NGS. The results are shown in Table 6. Many of the newly designed ABE variant sequences have A-to-G editing activity, and A-DN-8, A-DN-9, A-DN-10, A-DN-11, and A-DN-12 exhibit an A-to-G editing efficiency of over 60%.

[0213] Table 6. A-to-G Editing Efficiency of Different ABEs

[0214] Editor A-to-G editing efficiency (%) C-to-T editing efficiency (%) A-DN-1 4.64 0.09 A-DN-2 6.45 0.08 A-DN-3 5.67 0.09 A-DN-4 3.94 0.09 A-DN-5 4.59 0.10 A-DN-6 3.43 0.07 A-DN-7 36.12 0.20 A-DN-8 70.07 0.34 A-DN-9 69.55 0.45 A-DN-10 65.40 0.82 A-DN-11 66.30 0.81 A-DN-12 62.01 0.28 A-DN-13 8.27 0.12 A-DN-14 22.62 0.12 A-DN-15 21.27 0.11 A-DN-16 25.79 0.08 A-DN-17 25.98 0.09 A-DN-18 16.84 0.12 A-DN-19 21.58 0.12 A-DN-20 19.45 0.14 A-DN-21 9.18 0.06 A-DN-22 12.91 0.02 A-DN-23 40.40 0.29 A-DN-24 49.87 0.24 A-DN-25 2.17 0.06 A-DN-26 33.58 0.27 A-DN-27 40.30 0.16 A-DN-28 40.49 0.33

[0215] Example 6. De novo design of TadA variants with cytidine deaminase activity

[0216] TadA was engineered to mediate C to T editing and has advantages such as smaller size and fewer off-targets compared to natural cytidine deaminases such as rAPOBEC used in BE4max. For this purpose, based on structural and sequence information, new TadA with cytidine deaminase activity was designed, and the polynucleotide sequence encoding rAOPBEC (SEQ ID NO: 103) was replaced with the polynucleotide sequence encoding the newly designed TadA variant in the BE4max (encoding nucleotide sequence is SEQ ID NO: 102) expression vector to generate a series of CBE variants. Using the experimental method described in Example 5, its editing activity was measured in the HEK293T cell line, and the sgRNA sequence used was sgRNA_3 (GGAATCCCTTCTGCAGCACC, SEQ ID NO: 110), and its A to G and C to T editing activities were detected by NGS sequencing. As shown in Table 7, many newly designed CBE variants have C to T editing activity, and the editing activities of variants C-DN-7 and C-DN-8 are comparable to that of the control BE4max. Moreover, these CBE variants do not have A to G editing activity, demonstrating that they are specific cytosine deaminases.

[0217] Table 7. A to G and C to T editing efficiencies of different CBE variants

[0218] Editor Editing efficiency from A to G (%) Editing efficiency from C to T (%) C-DN-1 0.02 14.07 C-DN-2 0.03 17.10 C-DN-3 0.20 36.64 C-DN-4 0.08 20.73 C-DN-5 0.02 2.98 C-DN-6 0.04 22.96 C-DN-7 0.33 41.63 C-DN-8 0.31 42.63 C-DN-9 0.03 5.20 C-DN-10 0.05 6.37 C-DN-11 0.03 5.75 C-DN-12 0.02 2.03 C-DN-13 0.13 23.38 C-DN-14 0.06 16.00 C-DN-15 0.06 30.29 C-DN-16 0.37 32.73 BE4max 0.03 42.36

[0219] Example 7. Improving its C to T editing activity by directed evolution of T88.74

[0220] To enhance the activity of the cytidine deaminase T88.74-V0 (SEQ ID NO: 64) constructed by the inventors, a site-saturation mutagenesis library of T88.74 was constructed and subjected to directed evolution. The experimental system used three plasmids: a selection plasmid containing a chloramphenicol resistance gene (CamR) with two inactivating mutations (L158P and H193R); an sgRNA expression plasmid that produces two sgRNAs targeting the mutated sites in the CamR gene; and an editor (i.e., CBE variant) plasmid library that expresses T88.74-dCas9-2xUGI variants from the library. Library members capable of introducing C:G to T:A editing at the two mutated sites will restore chloramphenicol resistance, allowing bacteria to survive in the presence of chloramphenicol.

[0221] We co-transformed all three plasmids into Escherichia coli DH10B strain. After overnight induction of the editor library expression, the bacteria were inoculated onto culture plates containing chloramphenicol (selective; 16 or 64 µg / ml) or without chloramphenicol (non-selective). After overnight incubation at 37°C, plasmids were extracted from colonies grown on selective and non-selective plates and analyzed using next-generation sequencing (NGS). We calculated the percentage of mutations and the enrichment fold on the selective plates. The top 50 variants are shown in Figure 3 and Figure 2 respectively.

[0222] Example 8. Validation of Beneficial Mutations in Bacterial and Mammalian sgRNA Library Cells

[0223] We first used the system described in Example 7 to evaluate the CBE activity of the enriched T88.74 variants in Escherichia coli, and measured and compared the editing efficiency of each variant at two target sites on the selection plasmid by NGS (CAM-gRNA1: TGATCCGAACGTGGCCAATA, SEQ ID NO: 111; CAM-gRNA2: TACGGCGCGGTGCACCTGGA, SEQ ID NO: 112). The 11 highly active variants and their corresponding editing efficiencies are listed in Table 8.

[0224] Table 8. C to T Editing Efficiency of Different Variant Versions of T88.74 in Bacteria

[0225] Editor Mutation sites based on T88.74 Editing efficiency of CAM-gRNA1 Editing efficiency of CAM-gRNA2 T88.74-V0 Original sequence 0.91 7.11 T88.74-V8 A17N 1.02 4.18 T88.74-V3 C143A 2.18 15.78 T88.74-V5 F81L 1.59 7.96 T88.74-V4 G64C 0.87 8.25 T88.74-V6 L73F 2.23 16.31 T88.74-V2 S139K 1.37 6.51 T88.74-V7 V79T 0.94 3.51 T88.74-V1 Y70G 1.42 12.18 T88.74-V10 Y70H; M91L 16.49 39.80 T88.74-V9 Y70K; G102S 5.29 23.95 T88.74-V11 Y70K; H93N 2.36 2.83

[0226] Then, we selected mutations in the top 7 T88.74 variants validated in Escherichia coli for combination and further tested their CBE activity in HEK293T sgRNA library cells. In this HEK293T sgRNA library, more than 10,000 different sgRNAs and their corresponding target sequences were inserted into hotspots of the HEK293T genome, allowing analysis of the activity of CBE on multiple sgRNAs in a single simple experiment.

[0227] To evaluate the CBE activity of the selected T88.74 variants, 400 ng of the T88.74 variant-based CBE mRNA was transfected into the sgRNA library cells using Lipofectamine MessengerMAX (ThermoFisher, catalog number LMRNA003) according to the manufacturer's instructions. After 72 hours, genomic DNA of the cells was extracted, and the editing efficiency was analyzed by NGS. The sgRNAs detected in all samples were defined as valid sgRNAs. The C-to-T editing efficiency of these paired target sites was analyzed. The average editing efficiency value of each variant in the sgRNA library was calculated. As shown in Table 9 and Figure 4 as shown, the average editing efficiency of many engineered CBE variants was significantly higher than that of the original T88.74-V0.

[0228] Table 9. C-to-T editing efficiency of different variant versions of T88.74 in library cells

[0229] Editor Mutation sites based on T88.74 Average C-to-T editing efficiency ratio normalized with T88.74-V0 T88.74-V0 Original sequence 1.00 T88.74-V10 Y70H; M91L 2.03 T88.74-V12 Y70H; M91L; C143A 2.24 T88.74-V13 Y70H; M91L; G64C 2.24 T88.74-V14 Y70H; M91L; F81L 2.67 T88.74-V15 Y70H; M91L; L73F 2.28 T88.74-V16 Y70H; M91L; F81L; C143A; G64C; L73F 3.28 T88.74-V17 Y70H; M91L; F81L; G64C; L73F 3.27 T88.74-V18 Y70H; M91L; F81L; L73F 3.05 T88.74-V19 Y70H; M91L; F81L; G64C 3.02 T88.74-V20 Y70H; M91L; F81L; C143A 2.80

[0230] Then, we introduced 3 enhanced mutations (T88.74-V14) verified in bacteria and mammalian cells into the original directed evolution library, and selection was carried out again using a higher concentration of chloramphenicol (64 or 128 μg / ml) according to the method described in Example 7. The top 50 variants in the second round of directed evolution are shown in Figure 5 as follows.

[0231] Next, we selected some enhanced mutations and evaluated their activities in HEK293T sgRNA library cells according to the above method. Variants T88.74-V21 and T88.74-V22 showed significantly further enhanced C-to-T editing efficiency compared with the T88.74-V14 variant, as shown in Table 10 and Figure 6 as shown.

[0232] Table 10. C-to-T editing efficiency of different variant versions of T88.74 in library cells

[0233] Editor Mutation sites based on T88.74 Average C-to-T editing efficiency ratio normalized with T88.74-V14 T88.74-V14 Y70H; M91L; F81L 1.00 T88.74-V21 Y70H; M91L; F81L; N124V 1.25 T88.74-V22 Y70H; M91L; F81L; N124P 1.20 T88.74-V23 Y70H; M91L; F81L; R151M; K153G 1.05 T88.74-V24 Y70H; M91L; F81L; D112N; A119G 1.07 T88.74-V25 Y70H; M91L; F81L; D112N 1.06 T88.74-V26 Y70H; M91L; F81L; R130V; A140R 1.08

[0234] Then, we added a part of the newly confirmed enhanced mutations (Table 10) based on T88.74-V14 to T88.74-V17, one of the strongest variants screened previously, and evaluated its activity in HEK293T sgRNA library cells according to the above method. Variants T88.74-V27 and T88.74-V33 showed significantly further enhanced C-to-T editing efficiency compared with the T88.74-V17 variant, as shown in Table 11 and Figure 7 as shown.

[0235] Table 11. C to T editing efficiency of different variant versions of T88.74 in library cells

[0236] Editor Mutation sites based on T88.74 Average C-to-T editing efficiency ratio with activity standardized using T88.74-V17 T88.74-V17 Y70H; M91L; F81L; G64C; L73F 1.00 T88.74-V27 Y70H; M91L; F81L; G64C; L73F; N124V 1.17 T88.74-V28 Y70H; M91L; F81L; G64C; L73F; N124P 1.04 T88.74-V29 Y70H; M91L; F81L; G64C; L73F; R130V; A140R 1.09 T88.74-V30 Y70H; M91L; F81L; G64C; L73F; R151M; K153G 1.00 T88.74-V31 Y70H; M91L; F81L; G64C; L73F; D112N 1.05 T88.74-V32 Y70H; M91L; F81L; G64C; L73F; N124V; R130V; A140R 1.06 T88.74-V33 Y70H; M91L; F81L; G64C; L73F; N124V; R130V; A140R; D112N 1.10 T88.74-V35 Y70H; M91L; F81L; G64C; L73F; R130V 1.02

[0237] Other embodiments

[0238] It should be understood that although the present invention has been described in conjunction with its detailed description, the foregoing description is intended to be illustrative and not restrictive of the scope of the present invention, which is defined by the scope of the appended claims. Other aspects, advantages, and modifications are within the scope of the following claims.

Claims

1. An adenosine deaminase, wherein the adenosine deaminase consists of the amino acid sequence shown in any one of SEQ ID NOs: 2-19.

2. A nucleobase editor comprising: (i) a polynucleotide programmable DNA binding domain; and (ii) a deaminase, wherein the deaminase is the adenosine deaminase according to claim 1.

3. The nucleobase editor of claim 2, wherein the nucleobase editor further comprises one or more uracil glycosylase inhibitor domains.

4. The nucleobase editor of any one of claims 2-3, wherein the polynucleotide programmable DNA binding domain comprises any one domain selected from the group consisting of a Cas9 domain, a Cas12 domain, a TnpB domain, an IscB domain, a homing endonuclease domain, a zinc finger DNA binding domain, and a transcription activator-like effector DNA binding domain.

5. The nucleobase editor of claim 4, wherein the polynucleotide programmable DNA binding domain comprises a Cas9 domain.

6. The nucleobase editor of claim 5, wherein the Cas9 domain is Staphylococcus aureus Cas9, Streptococcus thermophilus Cas9, Streptococcus pyogenes Cas9 or a variant thereof.

7. A nucleobase editor according to claim 5, wherein the Cas9 domain consists of an amino acid sequence of SEQ ID NO: 115 or 116.

8. The nucleobase editor of any one of claims 2-3, wherein the polynucleotide programmable DNA binding domain is an inactive nuclease or nickase variant.

9. The nucleobase editor of any one of claims 2-3, wherein the nucleobase editor comprises a nucleic acid-guided nuclease protein and an embedded nucleotide base editor domain.

10. The nuclear base editor of claim 9, wherein the embedded nucleotide base editor domain is an embedded adenosine base editor domain.

11. The nucleobase editor of claim 10, wherein the embedded adenosine base editor domain is an embedded adenosine deaminase protein domain.

12. The nucleobase editor of claim 11, wherein the polynucleotide programmable DNA binding domain comprises a linker flanked by at least one embedded domain protein.

13. A fusion protein comprising a polynucleotide programmable DNA binding domain and at least one nucleobase editor domain comprising a deaminase linked together, wherein the deaminase is the adenosine deaminase of claim 1. A polynucleotide encoding the adenosine deaminase according to claim 1 .

15. A polynucleotide encoding a nucleobase editor according to any one of claims 2-12. A polynucleotide encoding the fusion protein according to claim 13 .

17. An expression vector comprising the polynucleotide according to any one of claims 14-16.

18. The expression vector according to claim 17, wherein the vector is any one viral vector selected from the group consisting of adeno-associated virus, retroviral vector, adenoviral vector, lentiviral vector, Sendai virus vector and herpes virus vector.

19. A cell comprising the expression vector according to claim 17 or 18, wherein the cell is not a plant cell.

20. Use of a nuclear base editor according to any one of claims 2-12 in the preparation of a drug or kit for base editing, wherein the base editing comprises contacting a polynucleotide sequence with a nuclear base editor according to any one of claims 2-12, wherein a deaminase deaminates a nuclear base in the polynucleotide, thereby editing the polynucleotide sequence.

21. Use of a nucleobase editor according to any one of claims 2-12 in the preparation of a drug or kit for correcting a genetic defect in a subject, wherein the correction of the genetic defect in the subject comprises administering to the subject: a nucleobase editor according to any one of claims 2-12 or a polynucleotide encoding a nucleobase editor according to any one of claims 2-12, and one or more guide polynucleotides that guide the nucleobase editor to deaminize a target nucleobase in the target nucleotide sequence of the subject, thereby correcting the genetic defect.

22. The use according to claim 21, comprising delivering the nucleobase editor or a polynucleotide encoding the nucleobase editor, and one or more guide polynucleotides to cells of the subject.

23. The use according to claim 21, wherein the subject is a non-human mammal or a human.

24. The use according to any one of claims 21-23, wherein deamination of the target nucleobase results in replacement of the target nucleobase with a wild-type nucleobase.

25. A molecular complex comprising: a nucleobase editor according to any one of claims 2-12 or a fusion protein according to claim 13, and one or more guide RNA sequences, tracrRNA sequences or target DNA sequences.

26. A kit comprising: a nucleobase editor according to any one of claims 2-12, or a fusion protein according to claim 13, or an expression vector according to claim 17, or a molecular complex according to claim 25, and instructions for use.

27. A pharmaceutical composition comprising: a nucleobase editor according to any one of claims 2-12, or a fusion protein according to claim 13, or an expression vector according to claim 17, or a molecular complex according to claim 25, and a pharmaceutically acceptable excipient.

28. A base editing method for non-therapeutic purposes, the method comprising contacting a polynucleotide sequence with a nuclear base editor according to any one of claims 2-12, wherein a deaminase deaminates a nuclear base in the polynucleotide, thereby editing the polynucleotide sequence.

Citation Information

Patent Citations

  • Cas12 protein, gene editing system containing Cas12 protein and application thereof

    CN113373130A

  • Adenosine deaminase variants and uses thereof

    CN117561074A