Deaminase, single base editor, and application thereof
By providing a deaminase to form a complex with nucleic acid-binding peptides, precise deamination of multinucleotide sequences is achieved, solving the problem of inefficient single-base editing in existing technologies and improving the accuracy and efficiency of gene editing.
Patent Information
- Application Number
- PCT/CN2025/091312
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-28
- Filing Date
- 2025-04-25
- Publication Date
- 2025-11-06
AI Technical Summary
Existing gene editing technologies are unable to achieve efficient and precise single-base editing, which limits the scope and effectiveness of gene editing applications.
A deaminase is provided, comprising a polypeptide with high homology to a specific amino acid sequence, which can form a complex with a nucleic acid-binding polypeptide to precisely deaminate polynucleotide sequences and perform targeted editing by binding to guide RNA.
It enables precise point mutations in multinucleotide sequences, expands the gene editing toolbox, and improves the accuracy and efficiency of gene editing.
Smart Images

Figure PCTCN2025091312-FTAPPB-I100001 
Figure PCTCN2025091312-FTAPPB-I100002 
Figure PCTCN2025091312-FTAPPB-I100003
Abstract
Description
Deaminases, single base editors, and uses thereof TECHNICAL FIELD
[0001] The present application relates to the field of gene editing, in particular to deaminases, single base editors, and uses thereof. BACKGROUND
[0002] CRISPR-Cas system is an adaptive immune defense formed by bacteria and archaea in the long-term evolution process, which can be used to resist invading viruses and exogenous DNA. SpCas9 derived from CRISPR / Cas9 system of Streptococcus pyogenes is widely used in genetic engineering due to its simple operation and high efficiency.
[0003] On April 20, 2016, David Liu et al. published a paper in Nature, describing single base editors that can achieve targeted modification of a single base.
[0004] Exploring more efficient and precise adenine deaminases for the construction of base editors, expanding the toolbox of gene editing, and achieving precise point mutation of target sites have become an important research direction of gene editing. SUMMARY
[0005] One aspect of the present application provides a deaminase comprising an amino acid sequence having at least 50% sequence identity to any one of SEQ ID NOs: 1-29, 34, 35, 40-50.
[0006] In some embodiments of the present disclosure, the deaminase is capable of deaminating at least one nucleotide in a polynucleotide sequence.
[0007] In some embodiments of the present disclosure, the deaminase is capable of deaminating at least one nucleotide in a target polynucleotide.
[0008] In some embodiments of the present disclosure, the deaminase deaminates at least one nucleotide in a polynucleotide sequence.
[0009] In some embodiments of the present disclosure, the deaminase deaminates at least one nucleotide in a target polynucleotide.
[0010] In some embodiments of the present disclosure, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to a sequence set forth in any one of SEQ ID NOs: 1-29, 34, 35, 40-50.
[0011] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 16, 24, 26, 34, and 35.
[0012] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 16, 24, 26, 34, and 35.
[0013] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 16, 24, 26, 34, and 35.
[0014] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 16, 24, 26, 34, and 35.
[0015] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 16, 24, 26, 34, and 35.
[0016] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 35.
[0017] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 40.
[0018] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 41.
[0019] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 43.
[0020] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 49.
[0021] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 16, and comprises a mutation at a position corresponding to residue Q123 of the sequence set forth in SEQ ID NO: 16. Without limitation, the corresponding position described herein can be determined by sequence alignment; for example, for a statement like “comprises a mutation at a position corresponding to residue Q123 of the sequence set forth in SEQ ID NO: 16,” it is meant that the deaminase polypeptide is one-dimensionally aligned (2-sequence alignment or multiple sequence alignment) with the sequence set forth in SEQ ID NO: 16, and the residue of the deaminase polypeptide that corresponds to residue Q123 of SEQ ID NO: 16 is identified, typically the residue that is in the same position in the alignment result figure.
[0022] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 16, and comprises a mutation at a position corresponding to residue Q123 of the sequence set forth in SEQ ID NO: 16.
[0023] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 16, and comprises a mutation at a position corresponding to residue Q123 of the sequence set forth in SEQ ID NO: 16.
[0024] In some embodiments of the disclosure, the deaminase polypeptide comprises a mutation at a position corresponding to at least one, at least two, at least three, at least four, at least five, or at least six of residues Q123, T105, E141, A109, T144, D120, L135, Y71 of the sequence set forth in SEQ ID NO: 16. In some embodiments of the disclosure, the deaminase polypeptide comprises a mutation corresponding to at least one, at least two, at least three, at least four, at least five, or at least six of Q123R, T105K, E141K, A109S, T144K, D120N, L135M, Y71H mutations of the sequence set forth in SEQ ID NO: 16.
[0025] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 16, and comprises a mutation at a position corresponding to residue Q123 of the sequence set forth in SEQ ID NO: 16, and the mutated amino acid residue is R.
[0026] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 26, and comprises a mutation at a position corresponding to residue L108, S141, S153, L108, S145, H155, Q154, A157, N126, and / or A18 of the sequence set forth in SEQ ID NO: 26.
[0027] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 26, and comprises a mutation at a position corresponding to residue L108, S141, S153, L108, S145, H155, Q154, A157, N126, and / or A18 of the sequence set forth in SEQ ID NO: 26, and the mutated amino acid residue corresponds to the same as L108F, S141R, S153R, L108H, S145N, H155P, Q154P, Q154K, A157T, N126H, A18S, Q154R of the sequence set forth in SEQ ID NO: 26.
[0028] In some embodiments of the disclosure, the deaminase polypeptide comprises a mutation at a position corresponding to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 of residues L108, S141, S153, L108, S145, H155, Q154, A157, N126, A18 of the sequence set forth in SEQ ID NO: 26.
[0029] In some embodiments of the disclosure, the deaminase polypeptide comprises a mutation at a position corresponding to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 of residues L108F, S141R, S153R, L108H, S145N, H155P, Q154P, Q154K, A157T, N126H, A18S, Q154R of the sequence set forth in SEQ ID NO: 26.
[0030] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 26, and comprises a mutation at a position corresponding to residue S141 of the sequence set forth in SEQ ID NO: 26.
[0031] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 26, and comprises a mutation at a position corresponding to the S141 residue of the sequence set forth in SEQ ID NO: 26, and the mutated amino acid residue is R.
[0032] In some embodiments of the disclosure, the deaminase comprises an amino acid sequence having 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 1-29, 34, 35, 40-50.
[0033] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 16, 24, 26, 34, and 35.
[0034] In some embodiments of the disclosure, the deaminase comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 1-29, 34, 35, 40-50.
[0035] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 16, 24, 26, 34, and 35.
[0036] In some embodiments of the disclosure, the deaminase comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 1-29, 34, 35, 40-50.
[0037] In some embodiments of the disclosure, the deaminase comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 16, 24, 26, 34, and 35.
[0038] In some embodiments of the disclosure, the deaminase can form a complex with a nucleic acid binding polypeptide, which in turn deaminates a base in a polynucleotide sequence.
[0039] In some embodiments of the disclosure, the deaminase is capable of forming a complex with a nucleic acid binding polypeptide and a guide RNA (gRNA), which directs the complex to bind to a target polynucleotide and deaminates a base in a sequence of the target polynucleotide.
[0040] In some embodiments of the disclosure, the nucleic acid binding polypeptide forms a complex with a guide RNA (gRNA) that directs the complex to bind to a target polynucleotide and deaminate a base in the sequence of the target polynucleotide.
[0041] The deaminase and the nucleic acid binding polypeptide in the complex can be two separate molecules or a single molecule (e.g., forming a fusion protein) in which the deaminase is linked to the nucleic acid binding polypeptide.
[0042] Non-limiting examples: the deaminase can be used to replace the deaminase domain in the tBE base editing system as described in patent documents WO2022206986A1 or WO2023155901A1, thereby obtaining a new base editing system.
[0043] In some embodiments of the disclosure, the deaminase is capable of forming a complex with a nucleic acid binding polypeptide and a guide RNA (gRNA), the nucleic acid binding polypeptide being a CRISPR-Cas nuclease, the guide RNA directing the complex to bind to a target polynucleotide and deaminate a base in the sequence of the target polynucleotide.
[0044] In some embodiments of the disclosure, the deaminase is capable of forming a complex with a nucleic acid binding polypeptide and a guide RNA (gRNA), the nucleic acid binding polypeptide being optionally selected from Cas9, Cas12, Cas13, TnpB, IscB, IsrB and Fancor nuclease, the guide RNA directing the complex to bind to a target polynucleotide and deaminate a base in the sequence of the target polynucleotide.
[0045] An aspect of the present application provides a fusion protein comprising: an amino acid sequence of a deaminase as described herein.
[0046] The amino acid sequence of the deaminase refers to the deaminase.
[0047] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 16.
[0048] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 24.
[0049] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 26.
[0050] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 34.
[0051] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 35.
[0052] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 40.
[0053] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 41.
[0054] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 43.
[0055] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 49.
[0056] In some embodiments of the disclosure, the deaminase comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in any one of SEQ ID NOs: 1-29, 34, 35, 40-50.
[0057] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 16, and comprises a mutation at a position corresponding to residue Q123, T105, E141, A109, T144, D120, L135, and / or Y71 of the sequence set forth in SEQ ID NO: 16, and the mutated amino acid residue corresponds to Q123R, T105K, E141K, A109S, T144K, D120N, L135M, Y71H of the sequence set forth in SEQ ID NO: 16.
[0058] In some embodiments of the disclosure, the deaminase polypeptide comprises a mutation at a position corresponding to at least 1, at least 2, at least 3, at least 4, at least 5, or at least 6 of residues Q123, T105, E141, A109, T144, D120, L135, Y71 of the sequence set forth in SEQ ID NO: 16. In some embodiments of the disclosure, the deaminase polypeptide comprises a mutation corresponding to at least 1, at least 2, at least 3, at least 4, at least 5, or at least 6 of mutations Q123R, T105K, E141K, A109S, T144K, D120N, L135M, Y71H of the sequence set forth in SEQ ID NO: 16.
[0059] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 16, and comprises a mutation at a position corresponding to residue Q123 of the sequence set forth in SEQ ID NO: 16, and the mutated amino acid residue is R.
[0060] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 26, and comprises a mutation at a position corresponding to the L108, S141, S153, L108, S145, H155, Q154, A157, N126, and / or A18 residues of the sequence set forth in SEQ ID NO: 26.
[0061] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 26, and comprises a mutation at a position corresponding to the L108, S141, S153, L108, S145, H155, Q154, A157, N126, and / or A18 residues of the sequence set forth in SEQ ID NO: 26, and the mutated amino acid residue corresponds to the same mutation as L108F, S141R, S153R, L108H, S145N, H155P, Q154P, Q154K, A157T, N126H, A18S, Q154R of the sequence set forth in SEQ ID NO: 26.
[0062] In some embodiments of the disclosure, the deaminase polypeptide comprises a mutation at a position corresponding to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 of the L108, S141, S153, L108, S145, H155, Q154, A157, N126, A18 residues of the sequence set forth in SEQ ID NO: 26.
[0063] In some embodiments of the disclosure, the deaminase polypeptide comprises a mutation at a position corresponding to at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 of the L108F, S141R, S153R, L108H, S145N, H155P, Q154P, Q154K, A157T, N126H, A18S, Q154R mutations of the sequence set forth in SEQ ID NO: 26.
[0064] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 26, and comprises a mutation at a position corresponding to the S141 residue of the sequence set forth in SEQ ID NO: 26.
[0065] In some embodiments of the disclosure, the deaminase polypeptide comprises an amino acid sequence having at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to the sequence set forth in SEQ ID NO: 26, and comprises a mutation at a position corresponding to the S141 residue of the sequence set forth in SEQ ID NO: 26, and the mutated amino acid residue is R.
[0066] In some embodiments of the disclosure, the fusion protein comprises:
[0067] a) a nucleic acid binding polypeptide that binds to a polynucleotide of interest; and
[0068] b) an amino acid sequence of a deaminase as described herein.
[0069] The fusion protein can be used for targeted editing of a nucleic acid (e.g., DNA nucleic acid, RNA nucleic acid) in vitro, ex vivo, or in vivo, such that an adenine nucleoside is converted to a guanine nucleoside, e.g., for the production of mutant cells. These mutant cells can be in a plant or an animal. Such a fusion protein can also be used to introduce a targeted mutation, e.g., for correction of a genetic deficiency in vitro in a mammalian cell (e.g., a cell obtained from a subject, which is then reintroduced into the same or another subject); and for introducing a targeted mutation, e.g., to correct a genetic deficiency or introduce a mutation in a disease-associated gene in a mammalian subject. Such a fusion protein can also be used to introduce a targeted mutation in a plant cell, e.g., for introducing a beneficial or agriculturally valuable trait or allele. The targeted editing, e.g., deaminates at least one nucleotide in the nucleic acid.
[0070] Optionally, the nucleic acid binding polypeptide is a DNA binding polypeptide.
[0071] Optionally, the nucleic acid binding polypeptide is a RNA binding polypeptide.
[0072] Optionally, the deaminase deaminates at least one nucleotide in the target polynucleotide.
[0073] In some embodiments of the disclosure, the nucleic acid binding polypeptide is a TALEN nuclease, a zinc finger nuclease, a CRISPR-Cas nuclease, or a meganuclease.
[0074] In some embodiments of the disclosure, the nucleic acid binding polypeptide is an RNA guided nuclease. Optionally, the DNA binding polypeptide is a CRISPR associated nuclease, i.e., a Cas enzyme (also known as Cas protein, CRISPR-Cas nuclease).
[0075] In some embodiments of the disclosure, the nucleic acid binding polypeptide is selected from Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f / CasZ, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, TnpB, IscB, IsrB, and Fancor; or fragments thereof. Non-limiting examples of the fragments are, for example, nucleic acid binding domain fragments.
[0076] In some embodiments of the disclosure, the nucleic acid binding polypeptide is selected from Cas9, Cas12, Cas13, TnpB, IscB, IsrB, Fancor nuclease; or fragments thereof, including but not limited to nucleic acid binding domain fragments.
[0077] In some embodiments of the disclosure, the Cas9 is selected from the group consisting of SpCas9, SaCas9, Nme2Cas9, Nme3Cas9, CjCas9, NmCas9, FnCas9, PpnCas9, FrCas9, SauCas9, SauriCas9, ScaCas9, St1Cas9, BlatCas9, CdiCas9, and GeoCas9, fragments thereof, and mutants thereof or fragments of mutants.
[0078] In some embodiments of the disclosure, the Cas12 is selected from C12-279, C12-334, C13-2, AsCpf1, enAsCas12a (addgene plasmid #196724), dFnCas12a (addgene plasmid #136379), ErCas12a, LbCas12a D832A, LbCas12a H759A, LbCas12a E795L, FnCas12a3, FnCas12a D917A, AsCas12a R1226A, AsCas12a D908A, AsCas12a E174R / S542R, AsCas12a (S542R / K548V / N552R), PrCas12a, PxCas12a, PcCas12a, PdCas12a, Mb2Cas12a, Mb3Cas12a, MlCas12a, CMaCas12a, CMtCas12a, HkCas12a, Lb5Cas12a, ErCas12a, TsCas12a, FnCpf1, LbCas12a, ttHsCas12a, AaCas12b, AaCas12b D570A, AaCas12b Q119F / E475R / E758R, BhCas12b, BvCas12b, BrCas12b, AkCas12b, AmCas12b, BsCas12b, OspCas12c, Cas12c2 (addgene plasmid #183072), Cas12c_4 (addgene plasmid #183071), Cas12c1 (addgene plasmid #120872), CasY.1 (from Katanobacteria), CasY.2 (from Vogelbacteria), CasY.3 (from Vogelbacteria), CasY.4 (from Parcubacteria), CasY.5 (from Komeilibacteria), CasY.6 (from Kerfeldbacteria), PlmCasX, DpbCasX, UnlCas12f, CnCas12fl, enRhCas12fl, AsCas12fl, SpaCas12fl, Cas12gl (addgene plasmid #120879), Cas12h of WO2021113522A1 (SEQ ID NO: 1 in that patent), Cas12il (addgene plasmid #171670), Cas12i2 (addgene plasmid #171671), Cas12j (addgene plasmid #120878), Cas12k (addgene plasmid #120877), Cas12l (addgene plasmid #120876), Cas12m (addgene plasmid #120875), Cas12n (addgene plasmid #120874), Cas12o (addgene plasmid #120873), Cas12p (addgene plasmid #120881), Cas12q (addgene plasmid #120882), Cas12r (addgene plasmid #120883), Cas12s (addgene plasmid #120884), Cas12t (addgene plasmid #120885), Cas12u (addgene plasmid #120886), Cas12v (addgene plasmid #120887), Cas12w (addgene plasmid #120888), Cas12x (addgene plasmid #120889), Cas12y (addgene plasmid #120890), Cas12z (addgene plasmid #120891), or a variant thereof.plasmid #188275), Cas12i1 (addgene plasmid #120882), Cas12i2 (addgene plasmid #120883), Cas12i protein named Cas12f.4 / Cas12f.5 / Cas12f.6 in CN111757889B, dSiCas12i (D1049A), SiCas12i, Si2Cas12i, WiCas12i, Wi2Cas12i, Wi3Cas12i, SaCas12i, Sa2Cas12i, Sa3Cas12i, WaCas12i, Wa2Cas12i, xCas12i, hfCas12Max, Cas12i-Max (addgene plasmid #188276), Cas12i1 D647A (addgene plasmid #171671), Cas12i-HiFi (addgene plasmid #188269), Cas12i1 D647A, Cas12j3 (addgene plasmid #188497), Cas12j2 (addgene plasmid #188498), AsCas12j-2 (addgene plasmid #191655), Cas12j-8 (addgene plasmid #194966), ShCas12k, N7Cas12k, AcCas12k, Cas12k-TniQ (addgene plasmid #181787), Cas12k-TnsC (addgene plasmid #181789), Cas12l, MmCas12m, MmCas12mΔZF (H549A, C552A), dCas12m-ΔZF (D485A, H549A, C552A), AcCas12n, dAcCas12n (D240), TnpB Actinomadura cellulosilytica strain DSM 45823, TnpB Actinomadura namibiensis strain DSM 44197, TnpB Actinomadura umbrina strain DSM 43927 $, TnpB Actinoplanes lobatus strain DSM 43150 (TnpB-1 and TnpB-2), TnpB Alicyclobacillus macrosporagiidus strain DSM 17980, TnpBHaloactinospora alba Strain DSM 45015, TnpB Lipingzhangella halophila strain DSM 102030, TnpB Meiothermus silvanus DSM 9946, TnpB QNFX01000004, ISDra2 TnpB (PDB: 8H1J), KraIscB-1, AwaIscB, OgeuIscB, GtFz1 (from Guillardia theta), SpuFz1 (from Spizellomyces punctatus), NlovFz2 (from Percolozoa Naegleria lovaniensis), and MmeFz2 (from Mercenaria mercenaria), fragments thereof, and mutants or fragments of mutants thereof.
[0079] In some embodiments of the disclosure, the DNA-binding polypeptide is capable of cleaving one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). Alternatively, the DNA-binding polypeptide is an RNA-guided nuclease with nickase activity.
[0080] In some embodiments of the disclosure, the Cas enzyme cleaves a target strand of a double-stranded nucleic acid molecule, meaning that the Cas enzyme cleaves the strand that base pairs (is complementary to) the gRNA (e.g., sgRNA) bound to the Cas enzyme.
[0081] In some embodiments of the disclosure, the Cas enzyme cleaves a non-target strand of a double-stranded nucleic acid molecule.
[0082] In some embodiments of the disclosure, the nucleic acid-binding polypeptide is an RNA-guided nuclease with nuclease activity inactivated.
[0083] In some embodiments of the disclosure, the nucleic acid-binding polypeptide is a variant of an RNA-guided nuclease with nuclease activity completely inactivated compared to the wild type, e.g., a dCas enzyme, including but not limited to, Cas9, Cas12, Cas13, TnpB, IscB, IsrB, and Fancor with nuclease activity completely inactivated (may be referred to as dead Cas9, dead Cas12, dead Cas13, dead TnpB, dead IscB, dead IsrB, and dead Fancor); or fragments or mutants thereof, e.g., SpRY Cas9 variants.
[0084] In some embodiments of the disclosure, the nucleic acid binding polypeptide is a variant of an RNA-guided nuclease that is partially inactivated for nuclease activity compared to the wild type, e.g., a nCas enzyme, e.g., a polypeptide fragment that retains nuclease activity for cleaving a single strand of double stranded DNA, including but not limited to nickase Cas9 and nickase Casl2.
[0085] In some embodiments of the disclosure, the fusion protein comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more cellular localization signals. The cellular localization signals include but are not limited to nuclear localization signal NLS, nuclear export signal NES, chloroplast localization signal, mitochondrial localization signal.
[0086] In some embodiments of the disclosure, the fusion protein comprises 1, 2, 3, 4, or more NES (nuclear export signal).
[0087] In some embodiments of the disclosure, the fusion protein comprises 1, 2, 3, 4, or more NLS (nuclear localization signal).
[0088] In some embodiments of the disclosure, the fusion protein comprises a nuclear export signal and a nuclear localization signal.
[0089] Nuclear localization signals are known in the art and generally comprise a stretch of basic amino acids (see, e.g., Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In particular embodiments, the fusion protein comprises 2, 3, 4, 5, 6, or more nuclear localization signals. One or more of the nuclear localization signals can be a heterologous NLS. Non-limiting examples of nuclear localization signals that can be used with the presently disclosed RGNs are the nuclear localization signals of the SV40 large T antigen, nucleoplasmin, and c-Myc (see, e.g., Ray et al., (2015) Bioconjug Chem 26(6): 1004-7). The fusion protein can comprise one or more NLS sequences at its N-terminus, C-terminus, or both N-terminus and C-terminus. For example, the fusion protein can comprise two NLS sequences in the N-terminal region and four NLS sequences in the C-terminal region.
[0090] Optionally, the NLS comprises an amino acid sequence as set forth in SPKKKRKVEAS, GPKKKRKVAAA, PKKKRKV, KRPAATKKAGQAKKKK, PAAKRVKLD, RQRRNELKRSP, NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY, RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV, VSRKRPRP, PPKKARED, POPKKKPL, SALIKKKKKMAP, DRLRR, PKQKKRK, RKLKKKIKKL, REKKKFLKRR, KRKGDEVDGVDEVAKKKSKK, RKCLQAGMNLEARKTKK, and PAAKKKKLD.
[0091] In some embodiments of the disclosure, the nucleic acid binding polypeptide, the deaminase polypeptide, and the cellular localization signal are directly connected or connected via linkers.
[0092] In some embodiments of the disclosure, when the nucleic acid binding polypeptide, the deaminase polypeptide, and the cellular localization signal are connected via linkers, the linkers can be the same or different.
[0093] In some embodiments of the disclosure, the linker comprises an amino acid sequence as set forth in SGGSSGGSSGSETPGTSESATPESSGGSSGGS, SGGSGGSGGS, and / or SGGSGGSGGSETRTADGSRLEFE. In some embodiments of the disclosure, the linker is an XTEN linker.
[0094] In some embodiments of the disclosure, the fusion protein comprises, in order from N-terminus to C-terminus, the following structures: deaminase, nucleic acid binding polypeptide. In some embodiments of the disclosure, the fusion protein comprises, in order from N-terminus to C-terminus, the following structures: nucleic acid binding polypeptide, deaminase.
[0095] In some embodiments of the disclosure, the fusion protein comprises, in order from N-terminus to C-terminus, the following structures: deaminase, linker 1, nucleic acid binding polypeptide. In some embodiments of the disclosure, the fusion protein comprises, in order from N-terminus to C-terminus, the following structures: nucleic acid binding polypeptide, linker 1, deaminase.
[0096] In some embodiments of the disclosure, the fusion protein comprises, in order from N-terminus to C-terminus, the following structures: deaminase, linker 1, nucleic acid binding polypeptide, linker 2, NLS. In some embodiments of the disclosure, the fusion protein comprises, in order from N-terminus to C-terminus, the following structures: NLS, linker 1, nucleic acid binding polypeptide, linker 2, deaminase.
[0097] In some embodiments of the disclosure, the fusion protein comprises the following structures in order from N-terminus to C-terminus: NLS, deaminase, linker 1, nucleic acid binding polypeptide, linker 2, NLS.
[0098] In some embodiments of the disclosure, the fusion protein comprises the following structures in order from N-terminus to C-terminus: SV 40 NLS, deaminase, linker 1, nucleic acid binding polypeptide, linker 2, SV 40 NLS.
[0099] The numbers in the linker 1 and the linker 2 are only used to distinguish the same type of nouns, and do not represent the actual meaning.
[0100] In some embodiments of the disclosure, the fusion protein is obtained by replacing the deaminase in the prior art ABE1, ABE2, ABE3, ABE4, ABE5, ABE6, ABE7, ABE8, ABE7.10-m, ABE7.10-d, ABE8e, ABE8.8-m, ABE8.8-d, ABE8.13-m, ABE8.13-d, ABE8.17-m, ABE8.17-d, ABE8.20-m or ABE8.20-d with the deaminase of the present application.
[0101] In some embodiments of the disclosure, the fusion protein forms a complex with a guide RNA (gRNA), which guides the complex to bind to a target polynucleotide and deaminate a base in the sequence of the target polynucleotide.
[0102] Another aspect of the present application provides an isolated nucleic acid encoding a deaminase as described herein or a fusion protein as described herein.
[0103] In some embodiments of the disclosure, the nucleic acid is codon-optimized, for example, prokaryotic codon-optimized for a prokaryotic expression system, or eukaryotic codon-optimized for a eukaryotic expression system.
[0104] Another aspect of the present application provides a vector comprising a nucleic acid as described herein.
[0105] A person skilled in the art can select a suitable vector, for example, the vector can be, including but not limited to, a viral vector, a bacterial vector, a fungal vector, an animal vector (for example, insects, fruit flies, nematodes, fish, zebrafish, mammals, mice, rats, rabbits, pigs, monkeys, humans, etc.), a plant vector, etc.
[0106] In some embodiments of the disclosure, the vector comprises a regulatory sequence. In some embodiments of the disclosure, the vector comprises a regulatory sequence operably linked to the nucleic acid. The regulatory sequence is used to regulate expression of the deaminase or fusion protein.
[0107] In some embodiments of the disclosure, the regulatory sequence is optionally selected from one or more of a promoter, an enhancer, an internal ribosome entry site, and a transcription termination signal; the promoter is, for example, a constitutive promoter, an inducible promoter, a broad-spectrum promoter, or a tissue-specific promoter, and / or the transcription termination signal is, for example, a polyadenylation signal or a poly-U sequence.
[0108] In some embodiments, the vector comprises a pol III promoter (e.g., U6 and H1 promoters), a pol II promoter (e.g., a Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), a cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), an SV40 promoter, a dihydrofolate reductase promoter, a beta-actin promoter, a phosphoglycerol kinase (PGK) promoter, or an EF1a promoter.
[0109] In some embodiments, the promoter is a constitutive promoter, which is continuously active and not regulated by external signals or molecules. Suitable constitutive promoters include, but are not limited to, CMV, RSV, SV40, EF1a, CAG, and beta-actin promoters. In some embodiments, the promoter is an inducible promoter, which is regulated by external signals or molecules (e.g., transcription factors).
[0110] In some embodiments, the promoter is a tissue-specific promoter, which can be used to drive tissue-specific expression of the deaminase or fusion protein. Suitable muscle-specific promoters include, but are not limited to, CK8, MHCK7, Myoglobin promoter (Mb), Desmin promoter, muscle creatine kinase promoter (MCK) and variants thereof, and SPc5-12 synthetic promoter. Suitable immune cell-specific promoters include, but are not limited to, B29 promoter (B cells), CD14 promoter (monocytes), CD43 promoter (leukocytes and platelets), CD68 (macrophages), and SV40 / CD43 promoter (leukocytes and platelets). Suitable blood cell-specific promoters include, but are not limited to, CD43 promoter (leukocytes and platelets), CD45 promoter (hematopoietic cells), INF-beta (hematopoietic cells), WASP promoter (hematopoietic cells), SV40 / CD43 promoter (leukocytes and platelets), and SV40 / CD45 promoter (hematopoietic cells). Suitable pancreas-specific promoters include, but are not limited to, Elastase-1 promoter. Suitable endothelial cell-specific promoters include, but are not limited to, Fit-1 promoter and ICAM-2 promoter. Suitable neuronal tissue / cell-specific promoters include, but are not limited to, GFAP promoter (astrocytes), SYN1 promoter (neurons), and NSE / RU5' (mature neurons). Suitable kidney-specific promoters include, but are not limited to, Nphsl promoter (podocytes). Suitable bone-specific promoters include, but are not limited to, OG-2 promoter (osteoblasts, odontoblasts). Suitable lung-specific promoters include, but are not limited to, SP-B promoter (lung). Suitable liver-specific promoters include, but are not limited to, SV40 / Alb promoter. Suitable heart-specific promoters include, but are not limited to, alpha-MHC.
[0111] Another aspect of the present application provides a composition comprising a deaminase as described herein, a fusion protein as described herein, a nucleic acid as described herein, or a vector as described herein.
[0112] Another aspect of the present application provides a composition comprising:
[0113] a deaminase as described herein, a fusion protein as described herein, a nucleic acid encoding a deaminase or fusion protein as described herein, or a vector as described herein.
[0114] In some embodiments of the disclosure, the composition further comprises one or more guide RNAs, or nucleic acids encoding the guide RNAs. Alternatively, the deaminase is capable of forming a complex with a nucleic acid binding polypeptide and a guide RNA (gRNA) that directs the complex to bind to a target polynucleotide and deaminate a base in the sequence of the target polynucleotide.
[0115] In some embodiments of the disclosure, the composition comprises a deaminase and a nucleic acid binding polypeptide; alternatively, the deaminase and the nucleic acid binding polypeptide can be two separate molecules, or a single molecule formed by linking the deaminase and the nucleic acid binding polypeptide (e.g., forming a fusion protein or a conjugate).
[0116] Non-limiting examples: the deaminase can be used to replace the deaminase domain in the tBE base editing system as described in patent documents WO2022206986A1 or WO2023155901A1, thereby obtaining a new base editing system.
[0117] In some embodiments of the disclosure, the deaminase is capable of forming a complex with a nucleic acid binding polypeptide and a guide RNA (gRNA), the nucleic acid binding polypeptide being a CRISPR-Cas nuclease, the guide RNA directing the complex to bind to a target polynucleotide and deaminate a base in the sequence of the target polynucleotide.
[0118] In some embodiments of the disclosure, a composition is provided, characterized in that the composition comprises:
[0119] (a) a fusion protein comprising an RNA-guided nuclease and a deaminase as described herein, or a nucleic acid encoding the fusion protein; and
[0120] (b) one or more guide RNAs, or nucleic acids encoding the guide RNAs;
[0121] the fusion protein and the guide RNA form a complex, the guide RNA directing the complex to bind to a target polynucleotide and deaminate a base in the sequence of the target polynucleotide.
[0122] In some embodiments of the disclosure, the composition further comprises a pharmaceutically acceptable delivery system.
[0123] In some embodiments of the disclosure, the delivery system is optionally selected from the group consisting of: a viral vector, a VLP (virus like particle), a nanoparticle such as a Lipid Nanoparticle (LNP), a liposome, an extracellular vesicle (EV) such as an Exosome and a microvesicle, and a gene gun.
[0124] In some embodiments of the disclosure, the viral vector is an adeno-associated virus, an adenovirus, a lentivirus, or a herpes virus.
[0125] In some embodiments of the disclosure, the adeno-associated virus is a recombinant adeno-associated viral vector of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShH10, AAVrh10, or AAVrh74.
[0126] In some embodiments of the disclosure, the deaminase, fusion protein, nucleic acid, or vector described herein is packaged in an AAV, for example, in an AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShH10, AAVrh10, or AAVrh74 capsid.
[0127] In some embodiments of the disclosure, the lentivirus is pseudotyped with an envelope protein. Preferably, the nucleic acid is linked to an aptamer sequence.
[0128] In some embodiments of the disclosure, in the virus like particle, the nucleic acid is linked to a gene encoding a gag protein.
[0129] Another aspect of the disclosure provides an expression system comprising a vector as described herein, a composition as described herein, or a genome having integrated therein an exogenous nucleic acid as described herein. The expression system can be a host cell, which can express a fusion protein as described herein, which is capable of complexing with a gRNA, thereby allowing the fusion protein to be localized to a target polynucleotide for base editing of the target polynucleotide.
[0130] In some embodiments of the disclosure, the host cell can be a eukaryotic cell and / or a prokaryotic cell. The eukaryotic cell includes a fungus, a plant cell, an animal cell, such as a mammalian cell, such as a mouse cell and a human cell, etc., such as a mouse brain neuroblastoma cell, a human embryonic kidney cell, a human cervical cancer cell, a human colon cancer cell, or a human osteosarcoma cell, etc., and more specifically, an N2a cell, a HEK293FT cell, a Hela cell, a HCT116 cell, or a U2OS cell, etc.
[0131] Another aspect of the present application provides a system for modifying a target polynucleotide, the system comprising:
[0132] a) one or more guide RNAs (gRNAs) or one or more nucleic acids encoding the one or more guide RNAs, the one or more guide RNAs being capable of hybridizing to the target polynucleotide;
[0133] The system further comprises any one or more of:
[0134] b) a deaminase as described herein, or a nucleic acid encoding the deaminase; and a nucleic acid binding polypeptide, or a nucleic acid encoding the nucleic acid binding polypeptide;
[0135] c) a fusion protein as described herein, or a nucleic acid encoding the fusion protein;
[0136] d) a vector as described herein;
[0137] e) a composition as described herein;
[0138] f) an expression system as described herein.
[0139] The target polynucleotide can also be referred to as a target nucleic acid.
[0140] In some embodiments of the disclosure, the system for modifying a target polynucleotide deaminates at least one nucleotide in the target polynucleotide. Further, the system for modifying a target polynucleotide deaminates at least one adenine nucleotide in the target polynucleotide.
[0141] In some embodiments of the disclosure, the deaminase in b) and the nucleic acid binding polypeptide form a complex, which in turn deaminates a base in a nucleic acid sequence.
[0142] In some embodiments of the disclosure, b) and a) together form a complex, which in turn deaminates a base in a target polynucleotide sequence.
[0143] In some embodiments of the disclosure, b), c), d), e) and / or f) are directed to a target polynucleotide by a), which in turn deaminates a base in a target polynucleotide sequence.
[0144] In some embodiments of the disclosure, the guide RNA is a guide RNA of a CRISPR-Cas system, or the system for modifying a target polynucleotide comprises an RNA-guided nuclease. In some embodiments of the disclosure, the system for modifying a target polynucleotide is a single base editing system comprising a guide RNA, an RNA-guided nuclease, and a deaminase as described herein.
[0145] In some embodiments of the disclosure, the guide RNA is not a guide RNA of a CRISPR-Cas system.
[0146] In some embodiments of the disclosure, the system for modifying a target polynucleotide does not comprise a nucleic acid binding polypeptide.
[0147] In some embodiments of the disclosure, the system for modifying a target polynucleotide does not comprise an RNA-guided nuclease.
[0148] In some embodiments of the disclosure, the guide RNA comprises a targeting domain that is reverse complementary to a target polynucleotide, and the guide RNA is capable of recruiting a deaminase or a fusion protein as described herein in a host cell comprising the system for modifying a target polynucleotide, which in turn deaminates an adenine nucleoside in the target polynucleotide.
[0149] In some embodiments of the disclosure, the guide RNA comprises an adenine deaminase recruiting domain. In some embodiments of the disclosure, the guide RNA does not comprise an adenine deaminase recruiting domain.
[0150] In some embodiments of the disclosure, the guide RNA is capable of forming a complex with a deaminase or a fusion protein as described herein in a host cell comprising the system for modifying a target polynucleotide through a covalent bond, a hydrogen bond, an acid-base based ionic bond, an aromatic ring π-π interaction, and / or a hydrophobic-hydrophobic interaction, or the guide RNA does not form a complex with a deaminase or a fusion protein as described herein in a host cell comprising the system for modifying a target polynucleotide.
[0151] In some embodiments of the disclosure, the guide RNA comprises a RNA sequence similar to dRNA, arRNA (ADAR-recruiting RNA) or circular arRNA in LEAPER or LEAPER2.0 technology (as described in patents / patent applications with publication number CN114269919B, CN113631708B, CN115651927B, TW202339774A, CN116515814A), which can recruit the deaminase as claimed in claim 1 or 2 or the fusion protein as claimed in claim 3 or 4.
[0152] In some embodiments of the disclosure, the guide RNA comprises a RNA sequence similar to CLUSTER guide RNA in CLUSTER guide RNA technology (as described in Reautschnig P, Wahn N, Wettengel J, Schulz AE, Latifi N, Vogel P, Kang TW, Pfeiffer LS, Zarges C, Naumann U, Zender L, Li JB, Stafforst T. CLUSTER guide RNAs enable precise and efficient RNA editing with endogenous ADAR enzymes in vivo. Nat Biotechnol. 2022 May; 40(5): 759-768. doi: 10.1038 / s41587-021-01105-0. Epub 2022 Jan 3. PMID: 34980913.), which contains a recruiting domain and a RE cluster and can recruit the deaminase or fusion protein as described in the present application.
[0153] In some embodiments of the disclosure, the guide RNA comprises an antisense oligonucleotide (AON) sequence, a small interfering RNA (siRNA) sequence, a short hairpin RNA (shRNA) sequence, and / or a microRNA (miRNA) sequence that targets to a target polynucleotide. Non-limiting examples include, for example, the guide RNA comprises an antisense oligonucleotide sequence similar to AONs in RESTORE technology (as described in Merkle T, Merz S, Reautschnig P, Blaha A, Li Q, Vogel P, Wettengel J, Li JB, Stafforst T. Precise RNA editing by recruiting endogenous ADARs with antisense oligonucleotides. Nat Biotechnol. 2019 Feb;37(2):133-138. doi: 10.1038 / s41587-019-0013-6. Epub 2019 Jan 28. PMID: 30692694.) that is capable of recruiting a deaminase or a fusion protein as described herein.
[0154] Another aspect of the present application provides a method of modifying a target polynucleotide, the method comprising contacting the target polynucleotide with a deaminase as described herein, a fusion protein as described herein, a nucleic acid as described herein, a vector as described herein, a composition as described herein, an expression system as described herein, or a system for modifying a target polynucleotide as described herein.
[0155] In some embodiments of the disclosure, the vector or expression system comprises a nucleotide sequence encoding the deaminase or fusion protein.
[0156] In some embodiments of the disclosure, the modification is deamination of at least one nucleotide in the target polynucleotide.
[0157] Optionally, the method is for non-therapeutic or diagnostic purposes.
[0158] The above conditions can be combined in any manner, consistent with common general knowledge in the art. BRIEF DESCRIPTION OF DRAWINGS
[0159] Figure 1 illustrates a plasmid map of ABE-16-01 single base editor.
[0160] Figure 2 shows the editing results of ABE-16-01 single base editor targeting the endogenous NG_053265.1 gene in 293T cells.
[0161] Figure 3 shows the results of editing by ABE-24-01 single base editor targeting the endogenous NG_053265.1 gene in 293T cells.
[0162] Figure 4 shows the results of editing by ABE-28-01 single base editor targeting the endogenous NG_053265.1 gene in 293T cells.
[0163] Figure 5 shows the results of editing by ABE8e single base editor targeting the endogenous NG_053265.1 gene in 293T cells.
[0164] Figure 6 shows the single base editing efficiency of nSpRY fused ABE-16-01 deaminase single base editor in targeting the genome in 293T cells.
[0165] Figure 7 shows the single base editing efficiency of nSpRY fused ABE-16-01-P1 deaminase mutant single base editor in targeting the genome in 293T cells.
[0166] Figure 8 shows the single base editing efficiency of nSpRY fused ABE-28-01-P2 deaminase mutant single base editor in targeting the genome in 293T cells.
[0167] All sequences in Figures 2-8 are gaacacaaagcatagactgc (SEQ ID NO: 30). DETAILED DESCRIPTION
[0168] TERMINOLOGY
[0169] DEAMINASE
[0170] The term "deaminase" refers to an enzyme that catalyzes a deamination reaction (i.e., removal of an amino group from a base of a nucleotide). In some embodiments, the deaminase is an adenine deaminase. The deaminase can act on DNA and / or RNA. In some embodiments, the adenine deaminase catalyzes deamination of an adenine nucleoside of an RNA molecule. In some embodiments, the adenine deaminase catalyzes deamination of an adenine nucleoside of a DNA molecule.
[0171] The term deaminase is used interchangeably with deaminase polypeptide.
[0172] DEAMINATION
[0173] In some embodiments, the deamination refers to catalyzing the deamination of a base in a DNA and / or RNA molecule. In some embodiments, the deamination refers to catalyzing the deamination of a base in a DNA molecule. In some embodiments, the deamination refers to catalyzing the deamination of a base in an RNA molecule. In some embodiments, the deamination refers to catalyzing the deamination of an adenine to generate a hypoxanthine.
[0174] In some embodiments, the base expected to be edited is upstream of the PAM site. In some embodiments, the base expected to be edited is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides upstream of the PAM site. In some embodiments, the base expected to be edited is downstream of the PAM site. In some embodiments, the base expected to be edited is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides downstream of the PAM site. The base that is edited is the base that undergoes deamination. The position at which the base that is edited is located is the editing window. In some embodiments, the target sequence comprises the editing window. In some embodiments, the editing window comprises 1-10 nucleotides. In some embodiments, the editing window is 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2, or 1 nucleotides in length. In some embodiments, the editing window is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides in length.
[0175] Proteins, peptides, and polypeptides
[0176] The terms "protein," "peptide," and "polypeptide" are used interchangeably herein to refer to polymers of amino acids linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide is at least three amino acids in length. A protein, peptide, or polypeptide can refer to a single protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified, e.g., by the addition of a chemical entity such as a carbohydrate group, a hydroxyl group, a phosphate group, a famesyl group, an isofamesyl group, a fatty acid group, a linker for conjugation, functionalization, or other modification, etc. A protein, peptide, or polypeptide can also be a single molecule or can be a multimeric complex. A protein, peptide, or polypeptide can be a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be naturally occurring, recombinant, or synthetic, or any combination thereof.
[0177] Nucleic acid binding polypeptide
[0178] Optionally, the nucleic acid binding polypeptide is a DNA binding polypeptide.
[0179] Optionally, the nucleic acid binding polypeptide is a RNA binding polypeptide.
[0180] Optionally, the nucleic acid binding polypeptide is a TALEN nuclease, a zinc finger nuclease, a CRISPR-Cas nuclease, or a meganuclease.
[0181] In some embodiments of the disclosure, the nucleic acid binding polypeptide is an RNA guided nuclease.
[0182] In some embodiments of the disclosure, the DNA binding polypeptide is an RNA guided nuclease.
[0183] In some embodiments of the disclosure, the RNA binding polypeptide is an RNA guided nuclease.
[0184] RNA guided nuclease
[0185] The term RNA guided nuclease (RGN) refers to a polypeptide that binds to a specific target polynucleotide in a sequence specific manner, and the polypeptide is guided to the target polynucleotide by a guide RNA molecule that is complexed with the polypeptide and hybridized to the target polynucleotide. The term RNA guided nuclease includes nucleases that cleave the target polynucleotide when bound to the target polynucleotide, and also includes nucleases that are capable of binding to, but do not cleave, the target polynucleotide, and are inactive for nuclease activity. Cleavage of a target sequence by an RNA guided nuclease can result in a single-stranded or double-stranded break. An RNA guided nuclease that is only capable of cleaving one of the single strands of a double-stranded nucleic acid molecule is referred to herein as a nickase.
[0186] Optionally, the DNA-binding polypeptide is a CRISPR-associated nuclease, i.e., a Cas enzyme (also referred to as a Cas protein).
[0187] In some embodiments of the disclosure, the nucleic acid-binding polypeptide is selected from Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12f / CasZ, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, Cas13a, Cas13b, Cas13c, Cas13d, Cas13e, Cas13f, TnpB, IscB, IsrB, and Fancor; or fragments thereof. Non-limiting examples of the fragments are, for example, nucleic acid-binding domain fragments.
[0188] In some embodiments of the disclosure, the nucleic acid-binding polypeptide is selected from Cas9, Cas12, Cas13, TnpB, IscB, IsrB, Fancor nuclease; or fragments thereof, including but not limited to nucleic acid-binding domain fragments.
[0189] In some embodiments of the disclosure, the Cas9 is selected from SpCas9, SaCas9, Nme2Cas9, Nme3Cas9, CjCas9, NmCas9, FnCas9, PpnCas9, FrCas9, SauCas9, SauriCas9, ScaCas9, St1Cas9, BlatCas9, CdiCas9, and GeoCas9, fragments thereof, and mutants thereof or fragments of mutants.
[0190] In some embodiments disclosed herein, the Cas12 is selected from AsCpf1, enAsCas12a (addgene plasmid#196724), dFnCas12a (addgene plasmid#136379), ErCas12a, LbCas12a D832A, LbCas12aH759A, LbCas12a E795L, FnCas12a3, FnCas12a D917A, AsCas12a R1226A, AsCas12a D908A, and AsCas12a E174R / S542R, AsCas12a(S542R / K548V / N552R), PrCas12a, PxCas12a, PcCas12a, PdCas12a, Mb2Cas12a, Mb3Cas12a, MlCas12a, CMaCa s12a, CMtCas12a, HkCas12a, Lb5Cas12a, ErCas12a, TsCas12a, FnCpf1, LbCas12a, ttHsCas12a, AaCas12b, AaCas12bD570A, AaCas12b Q119F / E475R / E758R, BhCas12b, BvCas12b, BrCas12b, AkCas12b, AmCas12b, BsCas12b, OspCas12c, Cas12c2(addgene plasmid#183072), Cas12c_4(addgene plasmid#183071), Cas12c1(addgene plasmid#120872), CasY.1 (from Katanobacteria), CasY.2 (from Vogelbacteria), CasY.3 (from Vogelbacteria), CasY.4 (from Parcubacteria), CasY.5 (from Komeilibacteria), CasY.6 (from Kerfeldbacteria), PlmCasX, DpbCasX, Un1Cas12f, CnCas12f1, enRhCas12f1, AsCas12f1, SpaCas12f1, Cas12g1 (addgene plasmid#120879), Cas12h of WO2021113522A1 (SEQ ID NO:1 in this patent), Cas12i1 (addgene plasmid#171670), Cas12i2 (addgene plasmid#188275), Cas12i1 (addgeneplasmid #120882), Cas12i2 (addgene plasmid #120883), Cas12i protein named Cas12f.4 / Cas12f.5 / Cas12f.6 in CN111757889B, dSiCas12i (D1049A), SiCas12i, Si2Cas12i, WiCas12i, Wi2Cas12i, Wi3Cas12i, SaCas12i, Sa2Cas12i, Sa3Cas12i, WaCas12i, Wa2Cas12i, xCas12i, hfCas12Max, Cas12i-Max (addgene plasmid #188276), Cas12i1 D647A (addgene plasmid #171671), Cas12i-HiFi (addgene plasmid #188269), Cas12i1 D647A, Cas12j3 (addgene plasmid #188497), Cas12j2 (addgene plasmid #188498), AsCas12j-2 (addgene plasmid #191655), Cas12j-8 (addgene plasmid #194966), ShCas12k, N7Cas12k, AcCas12k, Cas12k-TniQ (addgene plasmid #181787), Cas12k-TnsC (addgene plasmid #181789), Cas12l, MmCas12m, MmCas12mΔZF (H549A, C552A), dCas12m-ΔZF (D485A, H549A, C552A), AcCas12n, dAcCas12n (D240), TnpB Actinomadura cellulosilytica strain DSM 45823, TnpB Actinomadura namibiensis strain DSM 44197, TnpB Actinomadura umbrina strain DSM 43927 $, TnpB Actinoplanes lobatus strain DSM 43150 (TnpB-1 and TnpB-2), TnpB Alicyclobacillus macrosporagiidus strain DSM 17980, TnpB Haloactinospora alba Strain DSM 45015, TnpBLipingzhangella halophila strain DSM 102030, TnpB Meiothermus silvanus DSM 9946, TnpB QNFX01000004, ISDra2 TnpB (PDB: 8H1J), KraIscB-1, AwaIscB, OgeuIscB, GtFz1 (from Guillardia theta), SpuFz1 (from Spizellomyces punctatus), NlovFz2 (from Percolozoa Naegleria lovaniensis), and MmeFz2 (from Mercenaria mercenaria), fragments thereof, and mutants or fragments of mutants thereof.
[0191] In some embodiments of the disclosure, the DNA-binding polypeptide is capable of cleaving one strand of a double-stranded nucleic acid molecule (e.g., a double-stranded DNA molecule). Alternatively, the DNA-binding polypeptide is an RNA-guided nuclease with nickase activity.
[0192] In some embodiments of the disclosure, the Cas enzyme cleaves the target strand of a double-stranded nucleic acid molecule, meaning that the Cas enzyme cleaves the strand that base pairs (is complementary to) the gRNA (e.g., sgRNA) bound to the Cas enzyme.
[0193] In some embodiments of the disclosure, the nucleic acid-binding polypeptide is an RNA-guided nuclease with inactivated nuclease activity.
[0194] In some embodiments of the disclosure, the nucleic acid-binding polypeptide is a variant of an RNA-guided nuclease with completely inactivated nuclease activity compared to the wild-type, e.g., a dCas enzyme, including but not limited to, Cas9, Cas12, Cas13, TnpB, IscB, IsrB, and Fancor (may be referred to as dead Cas9, dead Cas12, dead Cas13, dead TnpB, dead IscB, dead IsrB, and dead Fancor) with completely inactivated nuclease activity.
[0195] In some embodiments of the disclosure, the nucleic acid-binding polypeptide is a variant of an RNA-guided nuclease with partially inactivated nuclease activity compared to the wild-type, e.g., a nCas enzyme, e.g., a polypeptide fragment that retains nuclease activity to cleave a single strand of a double-stranded DNA, including but not limited to, nickase Cas9 and nickase Cas12.
[0196] The target polynucleotide is bound by the nucleic acid binding polypeptide provided herein and hybridizes to the guide RNA associated with the nucleic acid binding polypeptide. Then, if the nucleic acid binding polypeptide has nuclease activity, the target sequence can be subsequently cleaved by the nucleic acid binding polypeptide. For example, the nucleic acid binding polypeptide can optionally cleave either single strand of a double stranded DNA or cleave both single strands, or can not cleave either single strand and only bind the target polynucleotide; the deaminase deaminates at least one base of the target polynucleotide at the same time or after the nucleic acid binding polypeptide binds or cleaves the target polynucleotide.
[0197] In some embodiments, the binding is dependent on recognition of a PAM sequence by the RNA-guided nuclease. In some embodiments, the binding is not dependent on recognition of a PAM sequence by the RNA-guided nuclease. The PAM sequences recognized by many of the nucleases described herein are known in the art.
[0198] The term "cleave" or "cleavage" refers to the breaking of at least one phosphodiester bond within the backbone of a target polynucleotide, which can result in a single- or double-stranded break within the target sequence. The presently disclosed nucleic acid binding polypeptides can cleave nucleotides within a polynucleotide as endonucleases or can be exonucleases (remove consecutive nucleotides from the end (5' and / or 3' end) of a polynucleotide). In other embodiments, the disclosed nucleic acid binding polypeptides can endonucleolytically cleave nucleotides of a target sequence within any position of a polynucleotide, thus acting as both endonucleases and exonucleases. Cleavage of a target polynucleotide by the presently disclosed nucleic acid binding polypeptides can result in staggered breaks or blunt ends.
[0199] Nuclear localization signal, nuclear localization sequence, NLS
[0200] The terms nuclear localization signal, nuclear localization sequence, and NLS are used interchangeably. The presently disclosed fusion proteins can comprise at least one nuclear localization signal (NLS) to enhance delivery of the nucleic acid binding polypeptide to the nucleus. Nuclear localization signals are known in the art and generally comprise a stretch of basic amino acids (see, e.g., Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In particular embodiments, the fusion protein comprises 2, 3, 4, 5, 6, or more nuclear localization signals. The nuclear localization signal(s) can be a heterologous NLS. Non-limiting examples of nuclear localization signals that can be used with the presently disclosed RGNs are the nuclear localization signals of the SV40 large T antigen, nucleoplasmin, and c-Myc (see, e.g., Ray et al., (2015) Bioconjug Chem 26(6): 1004-7).
[0201] Linker
[0202] The term "linker" as used herein refers to a chemical group or molecule that connects two molecules or moieties (e.g., the binding domain and cleavage domain of a nuclease). Typically, a linker is positioned between or flanking two groups, molecules, or other moieties and is connected to each group, molecule, or other moiety via a covalent bond, thereby connecting the two. In some embodiments, a linker is an amino acid or a plurality of amino acids (e.g., a peptide or protein). In some embodiments, a linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, a linker is 5-100 amino acids in length, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. Longer or shorter linkers can also be contemplated.
[0203] In some embodiments, a linker comprises a (GGGGS)n, (G)n, (EAAAK)n, (XP)n, or (SGG)n motif or a combination of any of these, wherein n is independently an integer between 1 and 30. In some embodiments, n is independently 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30, or any combination thereof if more than one linker or more than one linker motif is present. Other suitable linker motifs and linker configurations will be apparent to one of skill in the art. In some embodiments, suitable linker motifs and configurations include those described in Chen et al., Fusion protein linkers: property, design and functionality (Adv Drug Deliv Rev. 2013; 65(10): 1357-69, which is incorporated by reference herein in its entirety). Other suitable linker sequences will be apparent to one of skill in the art based on the present disclosure.
[0204] guide RNA, gRNA, guide polynucleotide
[0205] The terms guide RNA, gRNA, guide polynucleotide of the present application can be used interchangeably in some instances depending on the context. In some embodiments of the present disclosure, the guide RNA comprises a guide sequence and a scaffold sequence. The scaffold sequence interacts with the RNA-guided nuclease. The scaffold sequence is the sequence that is usually kept unchanged in the design of a guide RNA molecule. For example, the scaffold sequence can refer to the portion of the guide RNA molecule other than the guide sequence. For example, when the nucleic acid binding polypeptide is Cas9, Cas12 or Cas13, the guide RNA can be a crRNA or a complex of crRNA and tracrRNA, for example, a single molecule gRNA formed by the chimerization of crRNA and tracrRNA.
[0206] In some embodiments of the present disclosure, the guide sequence hybridizes to the target polynucleotide. The guide sequence refers to a contiguous nucleotide sequence in the gRNA that has partial or complete complementarity to the target polynucleotide, in particular, to the target sequence in the target polynucleotide [target nucleic acid], and can hybridize to the target sequence in the target polynucleotide through base pairing facilitated by the RNA-guided nuclease. The complete complementarity of the guide sequence to the target sequence is not required by the present disclosure, as long as there is sufficient complementarity to cause hybridization and to facilitate the formation of a gene editing complex. As used herein, the term target sequence refers to a small stretch of sequence in the target polynucleotide molecule that can be complementary (completely complementary or partially complementary) to the guide sequence of the gRNA molecule. The gene editing complex is sequence-specifically localized to the target sequence by the guide sequence and performs the corresponding function at or near this location. The length of the target sequence is often tens of nt (nucleotides), for example, can be 10-60 nt, 15-50 nt, 15-40 nt, 15-35 nt, 20-35 nt, 20-30 nt, about 10 nt, about 20 nt, about 30 nt, about 40 nt, about 50 nt, about 60 nt.
[0207] In some embodiments of the present disclosure, the guide sequence has at least 50% sequence identity to the target polynucleotide. In some embodiments of the present disclosure, the guide sequence has at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to the target polynucleotide.
[0208] The length of the guide sequence is often tens of nt (nucleotides), for example, can be 10-60 nt, 15-50 nt, 15-40 nt, 15-35 nt, 20-35 nt, 20-30 nt, about 10 nt, about 20 nt, about 30 nt, about 40 nt, about 50 nt, about 60 nt.
[0209] In some embodiments of the disclosure, the guide sequence is located at the 3' end or the 5' end of the backbone sequence. In some embodiments of the disclosure, the guide sequence is located at the 3' end of the backbone sequence. In some embodiments of the disclosure, the guide sequence is located at the 5' end of the backbone sequence.
[0210] In some embodiments of the disclosure, the guide RNA comprises an aptamer sequence.
[0211] In some embodiments of the disclosure, the aptamer sequence is inserted into a loop of a stem-loop structure of the secondary structure of the guide RNA.
[0212] In some embodiments of the disclosure, the guide RNA comprises modified nucleotides. The modifications include, but are not limited to, 2'-O-methyl, 2'-O-methyl-3'-phosphorothioate or 2'-O-methyl-3'-thioPACE modifications. In some embodiments of the disclosure, the guide RNA comprises modified nucleotides selected from deoxyribonucleotides, locked nucleic acids (LNAs). In some embodiments, the guide RNA comprises at least one chemically modified nucleotide.
[0213] In some embodiments, the guide RNA is a hybrid RNA-DNA guide, i.e., some RNA nucleotides in the guide RNA are replaced by DNA nucleotides. In some embodiments, the guide RNA is a hybrid RNA-LNA (locked nucleic acid) guide, i.e., some RNA nucleotides in the guide RNA are replaced by LNA nucleotides.
[0214] In some embodiments of the disclosure, the target RNA is located in the nucleus and / or cytoplasm of a eukaryotic cell.
[0215] Target polynucleotide
[0216] The term target polynucleotide is used interchangeably with target nucleic acid, target gene, depending on the context.
[0217] The deaminase or fusion protein of the disclosure can be used to deaminate adenosine in a target polynucleotide. It can be used to effect base conversion on a DNA molecule, or to effect base conversion on an RNA molecule. For example, it can be used to edit an exon region, an intron region, and / or a regulatory region (e.g., a promoter region, an enhancer region) of a gene; it can be used to edit a pre-mRNA molecule, to edit a mature mRNA molecule, or to edit a non-coding RNA molecule.
[0218] Alternatively, the target polynucleotide is a DNA sequence or an RNA sequence. In some embodiments of the disclosure, the target polynucleotide is a DNA sequence. In some embodiments of the disclosure, the target polynucleotide is an RNA sequence.
[0219] Patent application publication CN108513575A discloses a number of mutations that cause disease, which is incorporated herein by reference. These mutations can be corrected by the deaminases or fusion proteins of the present disclosure to treat the corresponding disease.
[0220] Table 27 of patent application publication WO2025061113A1 discloses a list of target nucleic acids / target genes and corresponding diseases or conditions, which is incorporated herein by reference. These target nucleic acids / target genes can be single base edited by the deaminases or fusion proteins of the present disclosure to treat the corresponding disease or condition.
[0221] Without limitation, the corresponding positions of the 2 amino acid sequences (or 2 proteins, or 1 protein and 1 amino acid sequence) described herein can be determined by sequence alignment; for example, for a statement like “comprises a mutation at a position corresponding to residue Q123 of the sequence set forth in SEQ ID NO: 16”, it refers to one-dimensional sequence alignment (2-sequence alignment or multiple sequence alignment) of the deaminase polypeptide with the sequence set forth in SEQ ID NO: 16, identifying the residue of the deaminase polypeptide corresponding to residue Q123 of SEQ ID NO: 16, which is usually the residue at the same position in the alignment result figure.
[0222] In some embodiments of the present disclosure, the deaminases provide higher editing efficiency compared to prior art deaminases. In some embodiments of the present disclosure, the deaminases provide narrower editing window compared to prior art deaminases. In some embodiments of the present disclosure, the deaminases provide higher editing specificity compared to prior art deaminases. In some embodiments of the present disclosure, the deaminases provide better safety compared to prior art deaminases. In some embodiments of the present disclosure, the deaminases provide different editing sites compared to prior art deaminases.
[0223] When referring to RNA sequences, “t” in the sequence can be used interchangeably with “u”. When referring to “guide sequence”, “t” in the sequence can be used interchangeably with “u”.
[0224] Depending on the context, “multiple” can refer to 2, 3, 4, or more, for example, “multiple” can refer to 2, 3, 4, or more. When “multiple” is used in parallel with “2”, “multiple” can also refer to 3, 4, or more.
[0225] The application will be further described in the following examples, but the application is not limited to the examples described. The experimental methods in the following examples, if not specified, are selected according to conventional methods and conditions, or according to the instructions of the commercial product.
[0226] Examples
[0227] Experimental Example 1. Design and synthesis of plasmids for expression of different adenine deaminases
[0228] The inventors, through screening, de novo design and / or mutagenesis, combined with experimental verification, have discovered a plurality of adenine deaminases (SEQ ID NOs: 1-29, 34, 35, 40-50), including adenine deaminases ABE-16-01 (SEQ ID NO: 16), ABE-24-01 (SEQ ID NO: 24), and ABE-28-01 (SEQ ID NO: 26).
[0229] >ABE-16-01
[0230] >ABE-24-01
[0231] >ABE-28-01
[0232] An expression framework for adenine base editors containing different adenine deaminases (SEQ ID NOs: 1-29, 34, 35, 40-50) and nSpRYCas9 sequences was designed and synthesized. A gRNA (guide sequence: gaacacaaagcatagactgc, SEQ ID NO: 30) targeting the NG_053265.1 gene sequence (NCBI, Homo sapiens VISTA enhancer hs267 (LOC110120638) on chromosome 5) was designed for detecting the function of adenine deaminases, and an expression framework for the gRNA was simultaneously synthesized. Then, the two frameworks were inserted into a pCDNA3.1 expression vector to obtain different adenine base editor plasmids, including an ABE-16-01 editor plasmid (SEQ ID NO: 31), an ABE-24-01 editor plasmid (SEQ ID NO: 32), and an ABE-28-01 editor plasmid (SEQ ID NO: 33). Other adenine base editor plasmids were constructed in the same way as ABE-16-01, except that the deaminase coding sequence was changed to the specific sequence of each. At the same time, ABE8e, which is currently reported and used more, was synthesized as a control, also known as ABE8e-UPC, and was simultaneously tested.
[0233] Specific sequences are shown below:
[0234] >ABE-16-01
[0235] The lower italic part is the expression region of the adenine base editor fusion protein, the lower underlined part is the deaminase coding sequence, and the upper bold underlined part is the coding sequence of the gRNA guide sequence.
[0236] Experimental Example 2. Detection of editing efficiency of base editors with different adenine deaminases
[0237] a. Cell plating: the 293T cell line was plated after digestion when the confluence reached 70-80%, and the number of cells seeded in a 24-well plate was 5*10^5 cells / well.
[0238] b. Transfection: 12-14 hours after plating, 1.5 ul PEI (Yincheng Bio, MW 25000) and 500 ng adenine base editor plasmid were added to 100 ul Opti-MEM, mixed, and incubated at room temperature for 20 minutes before being added to 293T cells for cell transfection. After overnight transfection, fresh medium was replaced for continued culture. Meanwhile, one well of cells was not transfected as a negative control, labeled as NC.
[0239] c. Cell lysis, PCR amplification of the edited region, and Sanger sequencing: after 72 hours of culture, the cells were washed with PBS, and then 100 ul of cell lysis solution (Viagen, 302-C, Lysis Reagent (Cell)) was added for lysis to obtain a lysis solution containing genomic DNA. The region near the target of the genomic DNA was amplified, and the PCR product was subjected to Sanger sequencing. Lysis Reagent(Cell)) for lysis to obtain a lysis solution containing genomic DNA. The region near the target of the genomic DNA was amplified, and the PCR product was subjected to Sanger sequencing.
[0240] d. Analysis of sequencing data: the sequencing peak chart and sgRNA target sequence were submitted to the editing efficiency analysis website (https: / / moriaritylab.shinyapps.io / editr_v10 / ) for single-base editing efficiency analysis to obtain the adenine editing efficiency of different deaminases at different positions of the target, as shown in Table 1.
[0241] Table 1. Editing efficiency of different adenine base editors
[0242] Although the ABE8e editor has high editing efficiency, the editing window is wide and the specificity is not good. ABE-16-01, ABE-24-01, and ABE-28-01 only edit the 5th adenine to A→G, and the specificity is very good.
[0243] In this experiment, the editing efficiency of other editors (ABE-01 to ABE-15, ABE-17-01 to ABE-23-01, ABE-25-01, ABE-29-01 to ABE-31-01) on all bases was 0.
[0244] Example 3. Obtaining ABE-16-01 and ABE-28-01 mutants and verification
[0245] Screening ABE-16-01 and ABE-28-01 high-efficiency / high-specificity mutants
[0246] The wild-type ABE-16-01 (152 aa) amino acid sequence is as follows:
[0247] The wild-type ABE-28-01 (173 aa) amino acid sequence is as follows:
[0248] According to the method in the literature (Neugebauer M E, Hsu A, Arbab M, et al. Evolution of an adenine base editor into a small, efficient cytosine base editor with low off-target activity [J]. Nature biotechnology, 2023, 41(5): 673-685.), phages expressing ABE-16-01 and ABE-28-01 were prepared, respectively, and using the screening system and strains in the literature, ABE-16-01 and ABE-28-01 were evolved by PANCE method, and finally a series of mutants as shown in Tables 2 and 3 were screened, including mutants ABE-16-01-P1 (Q123R) and ABE-28-01-P2 (S141R). Among them, multiple mutants have high editing efficiency and / or strict editing window.
[0249] Table 2. ABE-16-01 mutants evolved
[0250] The ABE-16-01-P1 (Q123R) mutant (152 aa) amino acid sequence is as follows:
[0251] Table 3. ABE-28-01 mutants evolved
[0252] The ABE-28-01-P2 (S141R) mutant (173 aa) amino acid sequence is as follows:
[0253] Synthesis of validation vectors ABE-16-01, ABE-28-01 targeting the target site on the genome using nSPRY, and the corresponding mutant vectors, wherein the spacer targeted by nSPRY on the genome is gaacacaaagcatagactgc (SEQ ID NO: 30). The rest of the sequence of each mutant vector is exactly the same except for the coding nucleic acid of the mutation site. Exemplarily, the sequence of ABE-16-01-P1 vector is SEQ ID NO: 36, and the sequence of ABE-28-01-P2 vector is SEQ ID NO: 37. The ABE base editor architecture expressed by the vector is NLS-deaminase-Linker-nSpRY-Linker-NLS.
[0254] Transfect 293T cells with different validation vectors and control vectors.
[0255] Transfect 24-well plates according to the instructions of Lipofectamine 2000 (Thermo), and detect at 108h.
[0256] Genomic DNA extraction kit (Tiangen) was used to extract the cell genome from the cells at 108h after transfection, and the extraction product was subjected to PCR amplification, and the primers are as follows:
[0257] ChkABE-PF1 ACTGCCATTCTACCAACAATAGAGGC (SEQ ID NO: 38)
[0258] ChkABE-PR1 CCAAGTGAGAAGCCAGTGGAATACAA (SEQ ID NO: 39)
[0259] Amplification was performed using 2X Pro Taq Premix Ver.2 (Aikangrui Biological), and the product length was 866bp. The product was sent to Jinweizhi for sequencing.
[0260] The sequencing results were analyzed by EditR 1.0.10, mainly analyzing the A5 site, and the results of 3 batches are shown in Table 4:
[0261] Table 4. Editing results of A5 site
[0262] In the experiments of 3 batches, the editing effects of P1, P2, P3, P4, P5, P6 mutants of ABE-16-01 were better than those of ABE-16-01, and the editing efficiency of P1 was the highest. The editing effects of P2, P5, P6, P8, P12 mutants of ABE-28-01 were better than those of ABE-28-01, and the editing efficiency of P2 was the highest.
[0263] Experimental Example 4. More adenine deaminases were tested
[0264] The same as the foregoing Example 1, Example 2, more different adenine deaminases were designed, synthesized, and tested, as shown in Table 5. The adenine deaminases were combined with nSPRYCas9 sequences to form an adenine base editor expression framework, to obtain each base editing plasmid, and then the editing efficiency was detected in 293T cells, as shown in Table 6. Each deaminase had a significantly narrower editing window and better specificity compared to the positive control group ABE8e-UPC. Among them, ABE-65-01, ABE-66-01, ABE-77-01, and ABE-90-01 had relatively high editing efficiency.
[0265] Table 5. Adenine deaminases
[0266] Table 6. Editing efficiency of different deaminases for A to G base mutation at different positions of the target Note: All blank spaces in the table represent 0% A to G editing efficiency.
Claims
1. A deaminase, characterized in that, the deaminase comprises an amino acid sequence having at least 50% sequence identity to any one of SEQ ID NOs: 1-29, 34, 35, 40-50; optionally, the deaminase comprises an amino acid sequence having at least 50% sequence identity to any one of SEQ ID NOs: 16, 24, 26, 34, and 35; optionally, the deaminase polypeptide deaminates at least one nucleotide in a polynucleotide.
2. The deaminase of claim 1, wherein, the deaminase comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 1-29, 34, 35, 40-50; optionally, the deaminase comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 16, 24, 26, 34, and 35.
3. A fusion protein, characterized in that, the fusion protein comprises: an amino acid sequence of the deaminase as set forth in claim 1 or 2.
4. The fusion protein of claim 3, wherein, the fusion protein comprises: a) a nucleic acid binding polypeptide that binds to a target polynucleotide; and b) an amino acid sequence of the deaminase as set forth in claim 1 or 2; optionally, the deaminase deaminates at least one nucleotide in a target polynucleotide; optionally, the target polynucleotide is a DNA sequence or an RNA sequence; optionally, the nucleic acid binding polypeptide is a TALEN nuclease, a zinc finger nuclease, a CRISPR-Cas nuclease, or a meganuclease; optionally, the nucleic acid binding polypeptide is an RNA-guided nuclease; optionally, the nucleic acid binding polypeptide is a Cas enzyme (CRISPR-Cas nuclease); optionally, the nucleic acid binding polypeptide is selected from Cas9, Cas12, Cas13, TnpB, IscB, IsrB, Fancor nuclease; or a fragment thereof, including but not limited to a nucleic acid binding domain fragment; further optionally, the fusion protein comprises one or more cell localization signals, including but not limited to a nuclear localization signal NLS, a nuclear export signal NES, a chloroplast localization signal, a mitochondrial localization signal; further optionally, the nucleic acid binding polypeptide, the deaminase polypeptide, and the one or more cell localization signals are directly linked or connected via a linker.
5. An isolated nucleic acid, comprising, the nucleic acid encodes the deaminase as set forth in claim 1 or 2, or the fusion protein as set forth in claim 3 or 4.
6. A vector, characterized in that, the vector comprises the nucleic acid of claim 5; optionally, the vector comprises a regulatory sequence operably linked to the nucleic acid.
7. A composition characterized in that, the composition comprises: the deaminase as set forth in claim 1 or 2, the fusion protein as set forth in claim 3 or 4, the nucleic acid as set forth in claim 5, or the vector as set forth in claim 6; optionally, the composition further comprises one or more guide RNAs, or nucleic acids encoding the guide RNAs; further optionally, the deaminase is capable of forming a complex with a nucleic acid binding polypeptide and a guide RNA (gRNA) that directs the complex to bind to a target polynucleotide and deaminate a base in a target polynucleotide sequence.
8. An expression system comprising, The expression system comprises the vector of claim 6, the composition of claim 7, or the nucleic acid of claim 5 integrated in the genome.
9. A system for modifying a target polynucleotide, comprising, The system comprises: a) one or more guide RNAs or one or more nucleic acids encoding the one or more guide RNAs, the one or more guide RNAs being capable of hybridizing to the target polynucleotide; The system further comprises any one or more of: b) the deaminase of claim 1 or 2, or a nucleic acid encoding the deaminase; and a nucleic acid binding polypeptide, or a nucleic acid encoding the nucleic acid binding polypeptide; c) the fusion protein of claim 3 or 4, or a nucleic acid encoding the fusion protein; d) the vector of claim 6; e) the composition of claim 7; f) the expression system of claim 8; Optionally, the system for modifying a target polynucleotide deaminates at least one nucleotide in the target polynucleotide. Optionally, the target polynucleotide is a DNA sequence or an RNA sequence.
10. The system for modifying a target polynucleotide of claim 9, wherein (a) the guide RNA is a guide RNA of a CRISPR-Cas system, or the system for modifying a target polynucleotide comprises an RNA-guided nuclease; optionally, the system for modifying a target polynucleotide is a single base editing system comprising a guide RNA, an RNA-guided nuclease, and the deaminase of claim 1 or 2; or, (b) the guide RNA is not a guide RNA of a CRISPR-Cas system, the system for modifying a target polynucleotide does not contain an RNA-guided nuclease, or, does not contain a nucleic acid-binding polypeptide; optionally, the guide RNA comprises a targeting domain that is reverse-complementary to a target polynucleotide, the guide RNA is capable of recruiting the deaminase of claim 1 or 2 or the fusion protein of claim 3 or 4 in a host cell containing the system for modifying a target polynucleotide, followed by deamination of an adenine nucleoside in the target polynucleotide; further optionally, the guide RNA comprises an adenine deaminase recruiting domain, or, the guide RNA does not comprise an adenine deaminase recruiting domain; further optionally, the guide RNA is capable of forming a complex with the deaminase of claim 1 or 2 or the fusion protein of claim 3 or 4 in a host cell containing the system for modifying a target polynucleotide through covalent bond, hydrogen bond, acid-base based ionic bond, aromatic ring π-π interaction and / or hydrophobic-hydrophobic interaction, or, the guide RNA does not form a complex with the deaminase of claim 1 or 2 or the fusion protein of claim 3 or 4 in a host cell containing the system for modifying a target polynucleotide; further optionally, (a) the guide RNA comprises an RNA sequence similar to a dRNA, an arRNA (ADAR-recruiting RNA) or a circular arRNA in the LEAPER or LEAPER2.0 technology (as described in the patents / patent applications with publication numbers CN114269919B, CN113631708B, CN115651927B, TW202339774A, CN116515814A), capable of recruiting the deaminase of claim 1 or 2 or the fusion protein of claim 3 or 4, (b) the guide RNA comprises a CLUSTER guide RNA technology (as described in Reautschnig P, Wahn N, Wettengel J, Schulz AE, Latifi N, Vogel P, Kang TW, Pfeiffer LS, Zarges C, Naumann U, Zender L, Li JB, Stafforst T. CLUSTER guide RNAs enable precise and efficient RNA editing with endogenous ADAR enzymes in vivo. Nat Biotechnol. 2022 May; 40(5): 759-768. doi: 10.1038 / s41587-021-01105-0. Epub 2022 Jan 3. PMID: 34980913.(b) the guide RNA comprises a sequence of an RNA similar to the CLUSTER guide RNA described in (c) the guide RNA comprises an antisense oligonucleotide (AON) sequence, a small interfering RNA (siRNA) sequence, a short hairpin RNA (shRNA) sequence, and / or a microRNA (miRNA) sequence, non-limiting examples of which are, for example, antisense oligonucleotide sequences similar to the AONs described in RESTORE technology (as described in Merkle T, Merz S, Reautschnig P, Blaha A, Li Q, Vogel P, Wettengel J, Li JB, Stafforst T. Precise RNA editing by recruiting endogenous ADARs with antisense oligonucleotides. Nat Biotechnol. 2019 Feb;37(2):133-138. doi: 10.1038 / s41587-019-0013-6. Epub 2019 Jan 28. PMID: 30692694.), which are capable of recruiting the deaminase of claim 1 or 2 or the fusion protein of claim 3 or 4; and / or (d) the guide RNA comprises a sequence of an RNA similar to the CLUSTER guide RNA described in (e) the guide RNA comprises an antisense oligonucleotide (AON) sequence, a small interfering RNA (siRNA) sequence, a short hairpin RNA (shRNA) sequence, and / or a microRNA (miRNA) sequence, non-limiting examples of which are, for example, antisense oligonucleotide sequences similar to the AONs described in RESTORE technology (as described in Merkle T, Merz S, Reautschnig P, Blaha A, Li Q, Vogel P, Wettengel J, Li JB, Stafforst T. Precise RNA editing by recruiting endogenous ADARs with antisense oligonucleotides. Nat Biotechnol. 2019 Feb;37(2):133-138. doi: 10.1038 / s41587-019-0013-6. Epub 2019 Jan 28. PMID: 30692694.), which are capable of recruiting the deaminase of claim 1 or 2 or the fusion protein of claim 3 or 4. Optionally, the system for modifying a target polynucleotide deaminates at least one nucleotide in the target polynucleotide. Optionally, the target polynucleotide is a DNA sequence or an RNA sequence.
11. A method of modifying a target polynucleotide, comprising: The method comprises contacting the target polynucleotide with the deaminase of claim 1 or 2, the fusion protein of claim 3 or 4, the nucleic acid of claim 5, the vector of claim 6, the composition of claim 7, the expression system of claim 8, or the system for modifying a target polynucleotide of claim 9 or 10.
Citation Information
Patent Citations
Uses of adenosine base editors
CN111757937A
Adenosine deaminase base editors and methods of using same to modify a nucleobase in a target sequence
CN114072496A
Adenosine deaminase, base editor fusion protein, base editor system and application
CN114634923A
Adenosine deaminase, base editor fusion protein, base editor system and application
CN117925585A
Adenine deaminase, adenine base editor containing same, and applications thereof
WO2023036189A1