Fusion protein comprising Cas protein and bacterial toxin and uses thereof

By designing the fusion protein of Cas protein and bacterial toxin SsdA, the vector problem caused by the large size of APOBEC or AID proteins is solved, and the efficient base editing and cell therapy application of the CRISPR-Cas system is achieved.

CN120092083APending Publication Date: 2025-06-03THE ASAN FOUND +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073912.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-10-19
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Due to its large size, APOBEC or AID proteins are difficult to make as carriers for cell therapeutic agents, which limits the application of the CRISPR-Cas system.

Method used

A fusion protein, which contains the Cas protein and the bacterial toxin SsdA, is designed to connect through a linker to form a CRISPR-Cas system that is easy to produce and express.

Benefits of technology

Highly efficient base editing of the CRISPR-Cas system is achieved, and due to the small size of the peptide, it is easy to deliver through vectors, has potential cellular therapeutic applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120092083A_ABST
    Figure CN120092083A_ABST
Patent Text Reader

Abstract

The present invention relates to a fusion protein comprising a Cas protein and a bacterial toxin and a use thereof, a fusion protein according to one aspect, a polypeptide thereof, and a CRISPR-Cas system comprising the fusion protein, which can perform efficient base editing. In addition, the size of the polypeptide is small, so that the polypeptide is easy to deliver through a carrier. In addition, when insertion deletion is induced, the insertion deletion efficiency is remarkably improved, when insertion deletion is induced, the insertion deletion efficiency is remarkably improved, and the size of nucleotide forming insertion deletion is also increased, so that the gene knockout can be effectively applied to gene knockout.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a fusion protein comprising a Cas protein and a bacterial toxin and uses thereof. Background Art

[0002] Genome editing is a technology for freely editing the genetic information of organisms. Through the progress in the field of life sciences and the development of genome sequencing technology, we can widely understand various genetic information. For example, information has been obtained regarding the reproduction, diseases and growth of animals and plants, gene mutations that cause various human genetic diseases, and genes for biofuel production, etc. However, to reach the level of directly applying it to improve organisms and treat human diseases, further technological progress is required.

[0003] The whole editing technology can greatly expand its application scope by changing the genetic information of animals, plants, and microorganisms including humans. Gene scissors are molecular tools designed and manufactured for precisely cutting the required genetic information and play a key role in genome editing technology. Similar to next-generation sequencing technology that has pushed the field of gene sequencing to a new height, gene scissors are becoming the core technology for expanding the application speed range of gene information and creating new industrial fields.

[0004] However, APOBEC or AID proteins, which are commonly used for cytosine base editing using Cas9, are relatively large in size and have the problem of being difficult to be made into a carrier for cell therapy agents. Therefore, there is a need to produce a CRISPR-Cas system that is easy to produce as a carrier and has excellent base editing efficiency. Summary of the Invention

[0005] Technical Problem

[0006] In one aspect, there is provided a fusion protein comprising a Cas protein (CRISPR-associated protein) and a bacterial toxin.

[0007] In another aspect, there is provided a polynucleotide encoding the fusion protein.

[0008] In still another aspect, there is provided a vector comprising the polynucleotide.

[0009] In still another aspect, there is provided a CRISPR-Cas system comprising the fusion protein or a polynucleotide encoding the fusion protein, and a guide polynucleotide.

[0010] In yet another aspect, there is provided a method of editing a nucleic acid, which comprises the step of contacting a nucleic acid molecule with the CRISPR-Cas system.

[0011] Technical solution

[0012] In one aspect, there is provided a fusion protein, which comprises a Cas protein (CRISPR-associated protein) and a bacterial toxin.

[0013] In the present specification, the term "Cas protein (CRISPR-associated protein)" may be a CRISPR-associated endonuclease. The Cas protein may cleave all or part of a specific target polynucleotide sequence.

[0014] The Cas protein may be a Class 2 Cas protein. The Class 2 Cas proteins may be included in Type II, Type V, or Type VI systems.

[0015] The Type II system may have cas1, cas2, and cas9 genes. The Type II system can be further divided into three subtypes, namely, II-A, II-B, and II-C subtypes. The II-A subtype may contain an additional gene csn2. Organisms having a Type II-A subtype system may include Streptococcus thermophilus. The II-B subtype lacks csn2 but may have cas4. Organisms having a Type II-B subtype system may include Legionella pneumophila. The II-C subtype is the most common Type II system in bacteria and may have only three proteins, Cas1, Cas2, and Cas9. Organisms having a Type II-C subtype system may include Neisseria lactamica.

[0016] The Type V system may have a cas12 gene and cas1 and cas2 genes. The cas12 gene may encode a Cas12 protein, which has a RuvC-like nuclease domain homologous to various regions of Cas9 but lacks the HNH nuclease domain present in the Cas9 protein.

[0017] The Type VI system may have a cas13 gene and cas1 and cas2 genes.

[0018] In the type II system, the RuvC-like nuclease (RNase H fold) domain and the HNH (McrA-like) nuclease domain of Cas9 can each cleave one strand of the target nucleic acid. The Cas9 cleavage activity of the type II system may also require hybridization of the crRNA with the tracrRNA in order to form a duplex that facilitates crRNA and target binding induced by Cas9.

[0019] In the type V system, the RuvC-like nuclease domain of Cas12 can cleave both strands of the target nucleic acid in a staggered structure, generating 5' overhangs. These 5' overhangs may facilitate DNA insertion by the non-homologous end joining method. The Cas12 cleavage activity of the type V system also does not require hybridization of the crRNA with the tracrRNA to form a duplex, and the crRNA of the type V system can use a single crRNA with a stem-loop structure that forms an internal duplex. The type V system can induce single-stranded or double-stranded breaks at the site of the target sequence. The strand breaks can be staggered cleavage using the 5' overhangs.

[0020] The Cas protein may comprise Cas9 or Cas12.

[0021] The Cas12 protein may refer to proteins derived from a variety of bacterial species. The Cas protein may be derived from Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helcococcus, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium, or Acidaminococcus.More specifically, the Cas12 protein can be derived from a bacterial species selected from the group consisting of Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella oralis, Lachnospiraceae bacterium MC20171, Proteiniphilum acetatigenes of the genus Butyrivibrio, bacterium GW2011_GWA2_33_10 of the genus Peregrinibacteria, bacterium GW2011_GWC2_44_17 of the genus Parvarchaeum, Treponema succinifaciens strain SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas canoris 3, Prevotella dentalis, and Porphyromonas macacae.

[0022] In addition, the Cas12 protein can be one selected from Cas12a, mgCas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, and Cas12j. The Cas12 protein can comprise a modification of the Cas12 protein. When the Cas12 protein has nuclease activity, the Cas12 protein can be modified to have reduced nuclease activity, for example, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or 100% nuclease inactivation compared to the wild-type enzyme.

[0023] The Cas9 can be S. pneumoniae, S. pyogenes, S. thermophilus, or C. jejuni Cas9, and can comprise mutant Cas9 derived from these organisms. The enzyme can be a Cas9 homolog or ortholog. In a specific example, the CRISPR enzyme can be codon-optimized for expression in eukaryotic cells. In a specific example, the CRISPR enzyme can induce single-stranded or double-stranded breaks at the site of the target sequence. The Cas9 protein can comprise a modification of the Cas9 protein. When the Cas9 protein has nuclease activity, the Cas9 protein can be modified to have reduced nuclease activity, for example, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or 100% nuclease inactivation compared to the wild-type enzyme. According to a specific example, the Cas9 can be Cas9 D10A.

[0024] In a specific example, the Cas protein can be a Cas9 protein or a Cas12 protein.

[0025] In one embodiment of the present invention, at least one nuclear localization signal (NLS) can be attached to the nucleic acid sequence encoding the Cas protein. In a specific example, at least one or more C-terminal or N-terminal nuclear localization signals can be attached. It can encode a Cas protein, its ortholog or homolog, which contains more than one nuclear localization sequence (NLS), for example, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nuclear localization sequences. In a preferred example of the Cas protein complex described herein, the codon-optimized Cas protein can contain a nuclear localization sequence attached to the C-terminus of the protein. In a specific example, other localization tags can be fused with the Cas protein, for example, to localize the Cas to a specific location within the cell, such as an organelle, such as: mitochondria, plastids, chloroplasts, vesicles, Golgi apparatus, (nuclear or cell) membrane, ribosome, nucleolus, endoplasmic reticulum (ER), cytoskeleton, vacuole, centrosome, nucleosome, granule, centriole, etc. (not limited thereto).

[0026] The Cas protein and the bacterial toxin can be fused through a linker. The linker can be located at the C-terminus, N-terminus, or both the C-terminus and N-terminus of the Cas protein, and through the linker, the bacterial toxin can bind to the Cas protein. The suitable linker motifs and linker formulations include those described in the literature [Chen et al., Fusion protein linkers: property, design and functionality. Adv Drug Deliv Rev. 2013; 65(10): 1357-69], the entire content of which can be incorporated into this specification by reference.

[0027] In a specific example, the bacterial toxin can be single-stranded DNA deaminase toxin A (SsdA).

[0028] The SsdA can be derived from a Pseudomonas SP. strain. The Pseudomonas strain can be Pseudomonas syringae, Pseudomonas congelans, Pseudomonas savastanoi, Pseudomonas viridiflava, Pseudomonas coronafaciens, Pseudomonas fluorescens, Pseudomonas sp. MPC6, Pseudomonas sp. GL-R-26, or Pseudomonas sp. GL-RE-26.

[0029] The N-terminus of the SsdA can have a PAAR domain, and the C-terminus of the SsdA can have a DYW deaminase domain. The SsdA is the same as the deaminases used in existing base editing in that it has the common amino acid motifs HxE and CxxC motifs, but is different from the deaminases used in existing base editing in that it also has an SGW motif. The SsdA is a deaminase that is structurally and evolutionarily different from other deaminases used in existing base editing technologies and is classified as a DYW-like deaminase. The phylogenetic tree and the differences in the domain composition between SsdA and other deaminases used in existing base editing technologies are as Figure 1 shown.

[0030] According to a specific example, the SsdA can comprise the amino acid sequence of SEQ ID NO: 1.

[0031] The SsdA can comprise a toxin domain. The toxin domain of the SsdA is the part with deaminase activity and can have a length of 100 to 200 amino acids, for example, a length of 120 to 180 amino acids, a length of 120 to 170 amino acids, a length of 130 to 160 amino acids, or a length of 140 to 160 amino acids. Specifically, the toxin domain of the SsdA can comprise the amino acid sequence of SEQ ID NO: 2. The base sequence of the toxin domain is shown in Table 1.

[0032] Table 1

[0033]

[0034] The amino acid sequence of SEQ ID NO: 2 of SsdA (toxin domain) may contain a catalytic active site. The catalytic active site may contain an HxE motif, a CxxC motif, or an SGW motif. The SGW motif is an additional motif unique to the SsdA enzyme and is different from existing deaminases such as APOBEC and AID. Specifically, the HxE motif may contain the amino acids at positions 301 to 303 of SEQ ID NO: 1 (PAAR domain-containing protein), the HxE motif may contain the amino acids at positions 347 to 349, and the SGW motif may contain the amino acids at positions 301 to 303.

[0035] In a specific example, the SsdA may have a sequence identity of at least 80% or more, 85% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more with the amino acid sequence of SEQ ID NO: 1 (PAAR domain-containing protein).

[0036] In a specific example, the fusion protein may be formed by binding a Cas protein and a DYW deaminase protein.

[0037] In a specific example, the SsdA may be inactivated SsdA.

[0038] The inactivated SsdA may have an amino acid mutation at the catalytic active site of the activated SsdA. The inactivated SsdA may have lower cytotoxicity than SsdA.

[0039] According to a specific example, the inactivated SsdA may have amino acid mutations at positions G302 and E349 in the amino acid sequence of SEQ ID NO: 1 (a protein containing a PAAR domain). The amino acid mutation means that the wild-type protein is replaced by an amino acid other than the amino acids at positions G302 and E349. The other amino acid may be any one selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine (C), selenocysteine (U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the amino acids, excluding the amino acids that the wild-type protein has at the mutation sites. Specifically, the inactivated SsdA may have G302D, E349A or their corresponding amino acid mutations in the amino acid sequence of SEQ ID NO: 1. More specifically, the inactivated SsdA having a G302D mutation in the amino acid sequence of SEQ ID NO: 1 (a protein containing a PAAR domain) may be SEQ ID NO: 17.

[0040] In a specific example, the amino acid sequence of the inactivated SsdA and SEQ ID NO: 17 may have a sequence identity of at least 80% or more, 85% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more.

[0041] The SsdA can induce deamination of single-stranded DNA, and the SsdA can be a cytidine deaminase.

[0042] In this specification, the term "cytidine deaminase" refers to an enzyme having the activity of removing the amino (-NH2) group of cytosine, cytidine, or deoxycytidine. In this specification, the cytidine deaminase is used as a concept including cytosine deaminase. In this specification, the cytidine deaminase can be used in combination with the cytosine deaminase.

[0043] The cytidine deaminase refers to all enzymes having the activity of converting the base cytosine present in a nucleotide (e.g., cytosine present in DNA or RNA) into uracil (C-to-U conversion or C-to-U editing), and converting cytosine on the strand where the PAM sequence of the target site sequence (target nucleic acid sequence) is located into uracil.

[0044] The bacterial toxin may be bound to the end of the Cas protein. For example, the bacterial toxin may be bound to the C-terminus, N-terminus, or both the C-terminus and N-terminus of the Cas protein.

[0045] In a specific example, the fusion protein may further comprise a DNA glycosylase inhibitor.

[0046] The DNA glycosylase inhibitor may be a thymine glycosylase inhibitor, uracil glycosylase inhibitor, oxoguanine glycosylase inhibitor, or alkylguanine DNA glycosylase inhibitor.

[0047] The uracil DNA glycosylase inhibitor may be a uracil DNA glycosylase inhibitor derived from the Bacillus subtilis phage, PBS1, a uracil DNA glycosylase inhibitor derived from the Bacillus subtilis phage or PBS2, but is not limited thereto.

[0048] In another aspect, a polynucleotide encoding the fusion protein is provided.

[0049] In yet another aspect, a vector comprising the polynucleotide is provided.

[0050] In this specification, the term "vector" can refer to a nucleic acid molecule that can transport another nucleic acid to which it is linked. A vector can be a single-stranded, double-stranded, or partially double-stranded nucleic acid molecule; a nucleic acid molecule containing more than one free end, a nucleic acid molecule without a free end (e.g., circular); a nucleic acid molecule containing DNA, RNA, or both; and various other polynucleotides known in the art. One type of vector is a "plasmid", which can refer to a circular double-stranded DNA loop into which additional DNA fragments can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which virus-derived DNA or RNA sequences can be present in the vector for packaging into a virus (e.g., retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses). A recombinant expression vector can contain a nucleic acid of the present invention in a form suitable for expression in a host cell, which can mean that the recombinant expression vector contains more than one regulatory element, and the more than one regulatory element can be selected based on the host cell to be used for expression and can be operably linked to the nucleic acid sequence to be expressed. In a recombinant expression vector, "operably linked" can mean that a nucleotide sequence of interest is linked to a regulatory element in a manner that permits expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell if the vector is introduced into the host cell).

[0051] In this specification, the term "regulatory element" may include a promoter, an enhancer, an internal ribosomal entry site (IRES), and other expression control elements (e.g., a transcription termination signal, e.g., a polyadenylation signal and a poly-U sequence). Regulatory elements are described, for example, in the literature [Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press San Diego, Calif. (1990)]. Regulatory elements may include elements that direct constitutive expression of a nucleotide sequence in many types of host cells and elements that direct expression of a nucleotide sequence only in specific host cells (e.g., tissue-specific regulatory sequences). In one specific example, a vector may contain more than one Pol III promoter (e.g., 1, 2, 3, 4, 5, or more Pol III promoters), more than one Pol II promoter (e.g., 1, 2, 2, 3, 4, 5, or more POL II promoters), more than one Pol I promoter (e.g., 1, 2, 3,, 4, 5, or more pol I promoters), or a combination thereof. Examples of Pol III promoters may include, but are not limited to, the U6 and H1 promoters. Examples of Pol II promoters may include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. For example, the vector may include lentivirus and adeno-associated virus (AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, or AAV9), and the type of vector may also be selected to target specific types of cells.

[0052] In addition, multiple nucleic acid molecules within the vector system may be located on the same or different vectors.

[0053] In one specific example, a vector, such as a plasmid or viral vector, is delivered to a tissue of interest, for example, by intramuscular injection, while in other cases, the delivery can be carried out by intravenous, percutaneous, intranasal, oral, mucosal, or other methods. Such delivery can be carried out by either a single dose or multiple doses. Those skilled in the art will understand that the actual dose delivered in this specification can vary greatly depending on a variety of factors, such as vector selection, target cells, organisms, or tissues, the general condition of the subject to be treated, the degree of transformation / modification sought, the route of administration, the method of administration, the type of transformation / modification sought, etc. The dosage administered can also include, for example, a vector (water, saline, ethanol, glycerol, lactose, sucrose, calcium phosphate, gelatin, dextran, agar, pectin, peanut oil, sesame oil, etc.), a diluent, a pharmaceutically acceptable carrier (such as phosphate-buffered saline), a pharmaceutically acceptable excipient, and / or other compounds known in the art. The dosage administered can also include more than one pharmaceutically acceptable salt, for example, inorganic acid salts such as hydrochloride, hydrobromide, phosphate, sulfate, etc.; and organic acid salts, such as acetate, propionate, malonate, benzoate, etc. Further, auxiliary components can also be given in this specification, such as wetting agents or emulsifiers, pH buffering components, gels or gelling substances, flavoring agents, coloring agents, microspheres, polymers, suspending agents, etc. Further, there can also be more than one other conventional pharmaceutical ingredient, such as preservatives, water-retaining agents, suspending agents, surfactants, antioxidants, fillers, chelating agents, coating agents, chemical stabilizers, etc. Suitable exemplary ingredients include microcrystalline cellulose, sodium carboxymethyl cellulose, polysorbate 80, phenethyl alcohol, chlorobutanol, potassium sorbate, sorbic acid, sulfur dioxide, propyl gallate, parabens, ethyl vanillin, glycerol, phenol, parachlorophenol, gelatin, albumin, and combinations thereof. For example, delivery for disease treatment can be achieved through AAV. A therapeutically effective dose for in vivo delivery of AAV to humans can be about 20 ml to about 50 ml of a saline solution, with each ml of the solution containing about 1×10¹⁰ to about 1×10¹⁰⁰ AAV. The dosage administered can be adjusted to balance the therapeutic benefit and any side effects.

[0054] In yet another aspect, there is provided a CRISPR-Cas system comprising: a fusion protein comprising a Cas protein and a bacterial toxin or a polynucleotide encoding said fusion protein; and a guide polynucleotide.

[0055] The fusion protein and the polynucleotide encoding the fusion protein are as described above.

[0056] The guide polynucleotide can comprise a targeting sequence and / or an activation sequence.

[0057] In this specification, the term "targeting sequence" may refer to a polynucleotide comprising DNA that is complementary to a sequence in a target nucleic acid, or a mixture of DNA and RNA. In certain embodiments, the targeting sequence may further comprise other nucleic acids, or nucleic acid analogs, or a combination thereof. In certain embodiments, the targeting sequence may consist only of DNA, as such constructs are less likely to be degraded in host cells. In some embodiments, such configurations may increase targeting sequence recognition specificity and / or reduce the occurrence of off-target binding / hybridization. The targeting sequence may comprise a guide sequence or a spacer sequence. The length of the domain of the targeting sequence may be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides.

[0058] In this specification, the term "activation sequence" may refer to a portion of a polynucleotide comprising RNA or DNA, or a mixture of DNA and RNA, that is capable of interacting with, or associating with, or binding to a Cas protein. In a specific example, the activation region may further comprise other nucleic acids, or nucleic acid analogs, or a combination thereof. In a specific example, the activation sequence may be adjacent to or linked to the target sequence. In a specific example, the activation region may be located downstream of the targeting region. In a specific example, the activation region may be located upstream of the targeting region. The activation sequence may comprise direct repeats, CRISPR RNA (CRISPR RNA: crRNA), and / or trans-activating RNA (trans-activating RNA: tracrRNA).

[0059] In a specific example, a guide polynucleotide, guide RNA, mature crRNA, and immature crRNA may comprise, or consist of, direct repeats and a guide sequence or spacer sequence. In a specific example, a guide RNA or mature crRNA may comprise, or consist of, a direct repeat linked to a guide sequence or spacer sequence. In a specific example, the direct repeat may be located upstream (i.e., 5') of the guide sequence or spacer sequence.

[0060] In a specific example, the guide polynucleotide may comprise crRNA and tracrRNA.

[0061] In a specific example, the guide polynucleotide may be a dual guide RNA or a single-chain guide RNA (sgRNA).

[0062] The system can form a deletion, insertion, substitution, or insertion and deletion (indel) of at least one nucleotide in the nucleotide sequence of a target nucleic acid molecule.

[0063] The nucleic acid can be RNA or DNA.

[0064] According to a specific example, the system can form a deletion of nucleotides from 1 bp to 60 bp in the nucleotide sequence of the target nucleic acid molecule, for example, 1 bp to 55 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 60 bp, 5 bp to 55 bp, 5 bp to 50 bp, 5 bp to 45 bp, 5 bp to 40 bp, 5 bp to 35 bp, 5 bp to 30 bp, 5 bp to 25 bp, 5 bp to 20 bp, 5 bp to 15 bp, 5 bp to 10 bp, 10 bp to 60 bp, 10 bp to 55 bp, 10 bp to 50 bp, 10 bp to 45 bp, 10 bp to 40 bp, 10 bp to 35 bp, 10 bp to 30 bp, 10 bp to 25 bp, 10 bp to 20 bp, 10 bp to 15 bp, 15 bp to 60 bp, 15 bp to 55 bp, 15 bp to 50 bp, 15 bp to 45 bp, 15 bp to 40 bp, 15 bp to 35 bp, 15 bp to 30 bp, 15 bp to 25 bp, 15 bp to 20 bp, 20 bp to 60 bp, 20 bp to 55 bp, 20 bp to 50 bp, 20 bp to 45 bp, 20 bp to 40 bp, 20 bp to 35 bp, 20 bp to 30 bp, 20 bp to 25 bp, 25 bp to 60 bp, 25 bp to 55 bp, 25 bp to 50 bp, 25 bp to 45 bp, 25 bp to 40 bp, 25 bp to 35 bp, 25 bp to 30 bp, 30 bp to 60 bp, 30 bp to 55 bp, 30 bp to 50 bp, 30 bp to 45 bp, 30 bp to 40 bp, 30 bp to 35 bp, 35 bp to 60 bp, 35 bp to 55 bp, 35 bp to 50 bp, 35 bp to 45 bp, 35 bp to 40 bp, 40 bp to 60 bp, 40 bp to 55 bp, 40 bp to 50 bp, 40 bp to 45 bp, 45 bp to 60 bp, 45 bp to 55 bp, 45 bp to 50 bp, 50 bp to 60 bp, 50 bp to 55 bp, or 55 bp to 60 bp.

[0065] According to one specific example, the system can form an insertion of nucleotides from 1 bp to 60 bp in the nucleotide sequence of the target nucleic acid molecule. For example, 1 bp to 55 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 60 bp, 5 bp to 55 bp, 5 bp to 50 bp, 5 bp to 45 bp, 5 bp to 40 bp, 5 bp to 35 bp, 5 bp to 30 bp, 5 bp to 25 bp, 5 bp to 20 bp, 5 bp to 15 bp, 5 bp to 10 bp, 10 bp to 60 bp, 10 bp to 55 bp, 10 bp to 50 bp, 10 bp to 45 bp, 10 bp to 40 bp, 10 bp to 35 bp, 10 bp to 30 bp, 10 bp to 25 bp, 10 bp to 20 bp, 10 bp to 15 bp, 15 bp to 60 bp, 15 bp to 55 bp, 15 bp to 50 bp, 15 bp to 45 bp, 15 bp to 40 bp, 15 bp to 35 bp, 15 bp to 30 bp, 15 bp to 25 bp, 15 bp to 20 bp, 20 bp to 60 bp, 20 bp to 55 bp, 20 bp to 50 bp, 20 bp to 45 bp, 20 bp to 40 bp, 20 bp to 35 bp, 20 bp to 30 bp, 20 bp to 25 bp, 25 bp to 60 bp, 25 bp to 55 bp, 25 bp to 50 bp, 25 bp to 45 bp, 25 bp to 40 bp, 25 bp to 35 bp, 25 bp to 30 bp, 30 bp to 60 bp, 30 bp to 55 bp, 30 bp to 50 bp, 30 bp to 45 bp, 30 bp to 40 bp, 30 bp to 35 bp, 35 bp to 60 bp, 35 bp to 55 bp, 35 bp to 50 bp, 35 bp to 45 bp, 35 bp to 40 bp, 40 bp to 60 bp, 40 bp to 55 bp, 40 bp to 50 bp, 40 bp to 45 bp, 45 bp to 60 bp, 45 bp to 55 bp, 45 bp to 50 bp, 50 bp to 60 bp, 50 bp to 55 bp, or 55 bp to 60 bp.

[0066] According to a specific example, the system can form insertions and deletions of nucleotides of 1 bp to 60 bp in the nucleotide sequence of the target nucleic acid molecule. For example, 1 bp to 55 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 60 bp, 5 bp to 55 bp, 5 bp to 50 bp, 5 bp to 45 bp, 5 bp to 40 bp, 5 bp to 35 bp, 5 bp to 30 bp, 5 bp to 25 bp, 5 bp to 20 bp, 5 bp to 15 bp, 5 bp to 10 bp, 10 bp to 60 bp, 10 bp to 55 bp, 10 bp to 50 bp, 10 bp to 45 bp, 10 bp to 40 bp, 10 bp to 35 bp, 10 bp to 30 bp, 10 bp to 25 bp, 10 bp to 20 bp, 10 bp to 15 bp, 15 bp to 60 bp, 15 bp to 55 bp, 15 bp to 50 bp, 15 bp to 45 bp, 15 bp to 40 bp, 15 bp to 35 bp, 15 bp to 30 bp, 15 bp to 25 bp, 15 bp to 20 bp, 20 bp to 60 bp, 20 bp to 55 bp, 20 bp to 50 bp, 20 bp to 45 bp, 20 bp to 40 bp, 20 bp to 35 bp, 20 bp to 30 bp, 20 bp to 25 bp, 25 bp to 60 bp, 25 bp to 55 bp, 25 bp to 50 bp, 25 bp to 45 bp, 25 bp to 40 bp, 25 bp to 35 bp, 25 bp to 30 bp, 30 bp to 60 bp, 30 bp to 55 bp, 30 bp to 50 bp, 30 bp to 45 bp, 30 bp to 40 bp, 30 bp to 35 bp, 35 bp to 60 bp, 35 bp to 55 bp, 35 bp to 50 bp, 35 bp to 45 bp, 35 bp to 40 bp, 40 bp to 60 bp, 40 bp to 55 bp, 40 bp to 50 bp, 40 bp to 45 bp, 45 bp to 60 bp, 45 bp to 55 bp, 45 bp to 50 bp, 50 bp to 60 bp, 50 bp to 55 bp, or 55 bp to 60 bp.

[0067] According to a specific example, the efficiency of the system in forming nucleotide insertions and deletions can be 5% to 50%, for example, 5% to 45%, 5% to 40%, 5% to 35%, 5% to 30%, 5% to 25%, 5% to 20%, 5% to 15%, 5% to 10%, 10% to 10%, 50%, 10% to 45%, 10% to 40%, 10% to 35%, 10% to 30%, 10% to 25%, 10% to 20%, 10% to 15%, 15% to 50%, 15% to 45%, 15% to 40%, 15% to 35%, 15% to 30%, 15% to 25%, 15% to 20%, 20% to 50%, 20% to 45%, 20% to 40%, 20% to 35%, 20% to 30%, 20% to 25%, 25% to 50%, 25% to 45%, 25% to 40%, 25% to 35%, 25% to 30%, 30% to 50%, 30% to 45%, 30% to 40%, 30% to 35%, 35% to 50%, 35% to 45%, 35% to 40%, 40% to 50%, 40% to 45%, or 45% to 50%.

[0068] According to a specific example, the efficiency of the system in forming nucleotide substitutions can be 1% to 20%, for example, 1% to 18%, 1% to 16%, 1% to 14%, 1% to 12%, 1% to 10%, 1% to 8%, 1% to 6%, 1% to 4%, 1% to 2%, 2% to 20%, 2% to 18%, 2% to 16%, 2% to 14%, 2% to 12%, 2% to 10%, 2% to 8%, 2% to 6%, 2% to 4%, 4% to 20%, 4% to 18%, 4% to 16%, 4% to 14%, 4% to 12%, 4% to 10%, 4% to 8%, 4% to 6%, 6% to 20%, 6% to 18%, 6% to 16%, 6% to 14%, 6% to 12%, 6% to 10%, 6% to 8%, 8% to 20%, 8% to 18%, 8% to 16%, 8% to 14%, 8% to 12%, 8% to 10%, 10% to 20%, 10% to 18%, 10% to 16%, 10% to 14%, 10% to 12%, 12% to 20%, 12% to 18%, 12% to 16%, 12% to 14%, 14% to 20%, 14% to 18%, 14% to 16%, 16% to 20%, 16% to 18%, or 18% to 20%.

[0069] The system can form an editing window of at least 4 nucleotides in the nucleotide sequence of a target nucleic acid molecule. According to a specific example, the system can have an editing window of at least 50 nucleotides, such as at least 49 nucleotides, at least 48 nucleotides, at least 47 nucleotides, at least 46 nucleotides, at least 45 nucleotides, at least 44 nucleotides, at least 43 nucleotides, at least 42 nucleotides, at least 41 nucleotides, at least 40 nucleotides, at least 39 nucleotides, at least 38 nucleotides, at least 37 nucleotides, at least 36 nucleotides, at least 35 nucleotides, at least 34 nucleotides, at least 33 nucleotides, at least 32 nucleotides, at least 31 nucleotides, at least 30 nucleotides, at least 29 nucleotides, at least 28 nucleotides, at least 27 nucleotides, at least 26 nucleotides, at least 25 nucleotides, at least 24 nucleotides, at least 23 nucleotides, at least 22 nucleotides, at least 21 nucleotides, at least 20 nucleotides, at least 19 nucleotides, at least 18 nucleotides, at least 17 nucleotides, at least 16 nucleotides, at least 15 nucleotides, at least 14 nucleotides, at least 13 nucleotides, at least 12 nucleotides, at least 11 nucleotides, at least 10 nucleotides, at least 9 nucleotides, at least 8 nucleotides, at least 7 nucleotides, at least 6 nucleotides, at least 5 nucleotides, or at least 4 nucleotides.

[0070] According to a specific example, the system may have an editing window that is 1 bp to 20 bp, 1 bp to 19 bp, 1 bp to 18 bp, 1 bp to 17 bp, 1 bp to 16 bp, 1 bp to 15 bp, 1 bp to 14 bp, 1 bp to 13 bp, 1 bp to 12 bp, 1 bp to 11 bp, 1 bp to 10 bp, 1 bp to 9 bp, 1 bp to 8 bp, 2 bp to 20 bp, 2 bp to 19 bp, 2 bp to 18 bp, 2 bp to 17 bp, 2 bp to 16 bp, 2 bp to 15 bp, 2 bp to 14 bp, 2 bp to 13 bp, 2 bp to 12 bp, 2 bp to 11 bp, 2 bp to 10 bp, 2 bp to 9 bp, 2 bp to 8 bp, 3 bp to 20 bp, 3 bp to 19 bp, 3 bp to 18 bp, 3 bp to 17 bp, 3 bp to 16 bp, 3 bp to 15 bp, 3 bp to 14 bp, 3 bp to 13 bp, 3 bp to 12 bp, 3 bp to 11 bp, 3 bp to 10 bp, 3 bp to 9 bp, 3 bp to 8 bp, 4 bp to 20 bp, 4 bp to 19 bp, 4 bp to 18 bp, 4 bp to 17 bp, 4 bp to 16 bp, 4 bp to 15 bp, 4 bp to 14 bp, 4 bp to 13 bp, 4 bp to 12 bp, 4 bp to 11 bp, 4 bp to 10 bp, 4 bp to 9 bp, or 4 bp to 8 bp from the 5` end of the gRNA target base sequence.

[0071] According to a specific example, the nucleotide editing window may show a cytosine (C 1 ) at the 1st position to a cytosine (C 20 ) at the 20th position from the 5` end of the gRNA target base sequence, C 2 to C 20 , C 3 to C 20 , C 4 to C 20 , C 1 to C 19 , C 2 to C 19 , C 3 to C 19 , C 4 to C 19 , C 1 to C 18 , C 2 to C 18 , C 3 to C 18 , C 4 to C 18 , C 1 to C 17 , C2 to C 17 、C 3 to C 17 、C 4 to C 17 、C 1 to C 16 、C 2 to C 16 、C 3 to C 16 、C 4 to C 16 、C 1 to C 15 、C 2 to C 15 、C 3 to C 15 、C 4 to C 15 、C 1 to C 14 、C 2 to C 14 、C 3 to C 14 、C 4 to C 14 、C 1 to C 13 、C 2 to C 13 、C 3 to C 13 、C 4 to C 13 、C 1 to C 12 、C 2 to C 12 、C 3 to C 12 、C 4 to C 12 、C 1 to C 11 、C 2 to C 11 、C 3 to C 11 、C 4 to C 11 、C 1 to C 10 、C 2 to C 10 、C 3 to C 10 、C 4 to C 10 、C 1 to C 9 、C 2 to C 9 、C 3to C 9 , C 4 to C 9 , C 1 to C 8 , C 2 to C 8 , C 3 to C 8 , or C 4 to C 8 base editing in the range of...

[0072] In yet another aspect, there is provided a method of editing a nucleic acid, comprising the step of contacting a nucleic acid molecule with a CRISPR-Cas system.

[0073] The nucleic acid, and the CRISPR-Cas system are as described above.

[0074] In a specific example, the editing can result in a deletion, insertion, substitution, or indel of at least one nucleotide sequence in the nucleotide sequence of the nucleic acid molecule.

[0075] According to a specific example, in the method of modifying the nucleic acid, the modification may be a deletion of nucleotides forming 1 bp to 60 bp in the nucleotide sequence of the nucleic acid molecule. For example, 1 bp to 55 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 60 bp, 5 bp to 55 bp, 5 bp to 50 bp, 5 bp to 45 bp, 5 bp to 40 bp, 5 bp to 35 bp, 5 bp to 30 bp, 5 bp to 25 bp, 5 bp to 20 bp, 5 bp to 15 bp, 5 bp to 10 bp, 10 bp to 60 bp, 10 bp to 55 bp, 10 bp to 50 bp, 10 bp to 45 bp, 10 bp to 40 bp, 10 bp to 35 bp, 10 bp to 30 bp, 10 bp to 25 bp, 10 bp to 20 bp, 10 bp to 15 bp, 15 bp to 60 bp, 15 bp to 55 bp, 15 bp to 50 bp, 15 bp to 45 bp, 15 bp to 40 bp, 15 bp to 35 bp, 15 bp to 30 bp, 15 bp to 25 bp, 15 bp to 20 bp, 20 bp to 60 bp, 20 bp to 55 bp, 20 bp to 50 bp, 20 bp to 45 bp, 20 bp to 40 bp, 20 bp to 35 bp, 20 bp to 30 bp, 20 bp to 25 bp, 25 bp to 60 bp, 25 bp to 55 bp, 25 bp to 50 bp, 25 bp to 45 bp, 25 bp to 40 bp, 25 bp to 35 bp, 25 bp to 30 bp, 30 bp to 60 bp, 30 bp to 55 bp, 30 bp to 50 bp, 30 bp to 45 bp, 30 bp to 40 bp, 30 bp to 35 bp, 35 bp to 60 bp, 35 bp to 55 bp, 35 bp to 50 bp, 35 bp to 45 bp, 35 bp to 40 bp, 40 bp to 60 bp, 40 bp to 55 bp, 40 bp to 50 bp, 40 bp to 45 bp, 45 bp to 60 bp, 45 bp to 55 bp, 45 bp to 50 bp, 50 bp to 60 bp, 50 bp to 55 bp, or 55 bp to 60 bp.

[0076] According to a specific example, in the method of modifying the nucleic acid, the modification can be the insertion of nucleotides forming 1 bp to 60 bp in the nucleotide sequence of the nucleic acid molecule. For example, 1 bp to 55 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 60 bp, 5 bp to 55 bp, 5 bp to 50 bp, 5 bp to 45 bp, 5 bp to 40 bp, 5 bp to 35 bp, 5 bp to 30 bp, 5 bp to 25 bp, 5 bp to 20 bp, 5 bp to 15 bp, 5 bp to 10 bp, 10 bp to 60 bp, 10 bp to 55 bp, 10 bp to 50 bp, 10 bp to 45 bp, 10 bp to 40 bp, 10 bp to 35 bp, 10 bp to 30 bp, 10 bp to 25 bp, 10 bp to 20 bp, 10 bp to 15 bp, 15 bp to 60 bp, 15 bp to 55 bp, 15 bp to 50 bp, 15 bp to 45 bp, 15 bp to 40 bp, 15 bp to 35 bp, 15 bp to 30 bp, 15 bp to 25 bp, 15 bp to 20 bp, 20 bp to 60 bp, 20 bp to 55 bp, 20 bp to 50 bp, 20 bp to 45 bp, 20 bp to 40 bp, 20 bp to 35 bp, 20 bp to 30 bp, 20 bp to 25 bp, 25 bp to 60 bp, 25 bp to 55 bp, 25 bp to 50 bp, 25 bp to 45 bp, 25 bp to 40 bp, 25 bp to 35 bp, 25 bp to 30 bp, 30 bp to 60 bp, 30 bp to 55 bp, 30 bp to 50 bp, 30 bp to 45 bp, 30 bp to 40 bp, 30 bp to 35 bp, 35 bp to 60 bp, 35 bp to 55 bp, 35 bp to 50 bp, 35 bp to 45 bp, 35 bp to 40 bp, 40 bp to 60 bp, 40 bp to 55 bp, 40 bp to 50 bp, 40 bp to 45 bp, 45 bp to 60 bp, 45 bp to 55 bp, 45 bp to 50 bp, 50 bp to 60 bp, 50 bp to 55 bp, or 55 bp to 60 bp.

[0077] According to a specific example, in the method of modifying the nucleic acid, the modification can be an insertion or deletion of nucleotides forming 1 bp to 60 bp of nucleotides in the nucleotide sequence of the nucleic acid molecule, such as 1 bp to 55 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 60 bp, 5 bp to 55 bp, 5 bp to 50 bp, 5 bp to 45 bp, 5 bp to 40 bp, 5 bp to 35 bp, 5 bp to 30 bp, 5 bp to 25 bp, 5 bp to 20 bp, 5 bp to 15 bp, 5 bp to 10 bp, 10 bp to 60 bp, 10 bp to 55 bp, 10 bp to 50 bp, 10 bp to 45 bp, 10 bp to 40 bp, 10 bp to 35 bp, 10 bp to 30 bp, 10 bp to 25 bp, 10 bp to 20 bp, 10 bp to 15 bp, 15 bp to 60 bp, 15 bp to 55 bp, 15 bp to 50 bp, 15 bp to 45 bp, 15 bp to 40 bp, 15 bp to 35 bp, 15 bp to 30 bp, 15 bp to 25 bp, 15 bp to 20 bp, 20 bp to 60 bp, 20 bp to 55 bp, 20 bp to 50 bp, 20 bp to 45 bp, 20 bp to 40 bp, 20 bp to 35 bp, 20 bp to 30 bp, 20 bp to 25 bp, 25 bp to 60 bp, 25 bp to 55 bp, 25 bp to 50 bp, 25 bp to 45 bp, 25 bp to 40 bp, 25 bp to 35 bp, 25 bp to 30 bp, 30 bp to 60 bp, 30 bp to 55 bp, 30 bp to 50 bp, 30 bp to 45 bp, 30 bp to 40 bp, 30 bp to 35 bp, 35 bp to 60 bp, 35 bp to 55 bp, 35 bp to 50 bp, 35 bp to 45 bp, 35 bp to 40 bp, 40 bp to 60 bp, 40 bp to 55 bp, 40 bp to 50 bp, 40 bp to 45 bp, 45 bp to 60 bp, 45 bp to 55 bp, 45 bp to 50 bp, 50 bp to 60 bp, 50 bp to 55 bp, or 55 bp to 60 bp.

[0078] According to a specific example, the efficiency of the method for modifying the nucleic acid to form an indel of a nucleotide sequence can be 5% to 50%, for example, 5% to 45%, 5% to 40%, 5% to 35%, 5% to 30%, 5% to 25%, 5% to 20%, 5% to 15%, 5% to 10%, 10% to 50%, 10% to 45%, 10% to 40%, 10% to 35%, 10% to 30%, 10% to 25%, 10% to 20%, 10% to 15%, 15% to 50%, 15% to 45%, 15% to 40%, 15% to 35%, 15% to 30%, 15% to 25%, 15% to 20%, 20% to 50%, 20% to 45%, 20% to 40%, 20% to 35%, 20% to 30%, 20% to 25%, 25% to 50%, 25% to 45%, 25% to 40%, 25% to 35%, 25% to 30%, 30% to 50%, 30% to 45%, 30% to 40%, 30% to 35%, 35% to 50%, 35% to 45%, 35% to 40%, 40% to 50%, 40% to 45%, or 45% to 50%.

[0079] According to a specific example, the efficiency of the method for modifying the nucleic acid to form a substitution of a nucleotide sequence can be 1% to 20%, for example, 1% to 18%, 1% to 16%, 1% to 14%, 1% to 12%, 1% to 10%, 1% to 8%, 1% to 6%, 1% to 4%, 1% to 2%, 2% to 20%, 2% to 18%, 2% to 16%, 2% to 14%, 2% to 12%, 2% to 10%, 2% to 8%, 2% to 6%, 2% to 4%, 4% to 20%, 4% to 18%, 4% to 16%, 4% to 14%, 4% to 12%, 4% to 10%, 4% to 8%, 4% to 6%, 6% to 20%, 6% to 18%, 6% to 16%, 6% to 14%, 6% to 12%, 6% to 10%, 6% to 8%, 8% to 20%, 8% to 18%, 8% to 16%, 8% to 14%, 8% to 12%, 8% to 10%, 10% to 20%, 10% to 18%, 10% to 16%, 10% to 14%, 10% to 12%, 12% to 20%, 12% to 18%, 12% to 16%, 12% to 14%, 14% to 20%, 14% to 18%, 14% to 16%, 16% to 20%, 16% to 18%, or 18% to 20%.

[0080] The method of modifying the nucleic acid can form an editing window of at least 4 nucleotides in the nucleotide sequence of the target nucleic acid molecule. According to a specific example, the method of modifying the nucleic acid has an editing window of at least 50 nucleotides, for example, at least 49 nucleotides, at least 48 nucleotides, at least 47 nucleotides, at least 46 nucleotides, at least 45 nucleotides, at least 44 nucleotides, at least 43 nucleotides, at least 42 nucleotides, at least 41 nucleotides, at least 40 nucleotides, at least 39 nucleotides, at least 38 nucleotides, at least 37 nucleotides, at least 36 nucleotides, at least 35 nucleotides, at least 34 nucleotides, at least 33 nucleotides, at least 32 nucleotides, at least 31 nucleotides, at least 30 nucleotides, at least 29 nucleotides, at least 28 nucleotides, at least 27 nucleotides, at least 26 nucleotides, at least 25 nucleotides, at least 24 nucleotides, at least 23 nucleotides, at least 22 nucleotides, at least 21 nucleotides, at least 20 nucleotides, at least 19 nucleotides, at least 18 nucleotides, at least 17 nucleotides, at least 16 nucleotides, at least 15 nucleotides, at least 14 nucleotides, at least 13 nucleotides, at least 12 nucleotides, at least 11 nucleotides, at least 10 nucleotides, at least 9 nucleotides, at least 8 nucleotides, at least 7 nucleotides, at least 6 nucleotides, or at least 5 nucleotides.

[0081] According to a specific example, the method of modifying the nucleic acid may have an editing window that is 1 bp to 20 bp, 1 bp to 19 bp, 1 bp to 18 bp, 1 bp to 17 bp, 1 bp to 16 bp, 1 bp to 15 bp, 1 bp to 14 bp, 1 bp to 13 bp, 1 bp to 12 bp, 1 bp to 11 bp, 1 bp to 10 bp, 1 bp to 9 bp, 1 bp to 8 bp, 2 bp to 20 bp, 2 bp to 19 bp, 2 bp to 18 bp, 2 bp to 17 bp, 2 bp to 16 bp, 2 bp to 15 bp, 2 bp to 14 bp, 2 bp to 13 bp, 2 bp to 12 bp, 2 bp to 11 bp, 2 bp to 10 bp, 2 bp to 9 bp, 2 bp to 8 bp, 3 bp to 20 bp, 3 bp to 19 bp, 3 bp to 18 bp, 3 bp to 17 bp, 3 bp to 16 bp, 3 bp to 15 bp, 3 bp to 14 bp, 3 bp to 13 bp, 3 bp to 12 bp, 3 bp to 11 bp, 3 bp to 10 bp, 3 bp to 9 bp, 3 bp to 8 bp, 4 bp to 20 bp, 4 bp to 19 bp, 4 bp to 18 bp, 4 bp to 17 bp, 4 bp to 16 bp, 4 bp to 15 bp, 4 bp to 14 bp, 4 bp to 13 bp, 4 bp to 12 bp, 4 bp to 11 bp, 4 bp to 10 bp, 4 bp to 9 bp, or 4 bp to 8 bp from the 5` end of the gRNA target base sequence.

[0082] According to a specific example, the editing window of the method of modifying the nucleic acid may show a cytosine (C 1 ) at the first position to a cytosine (C 20 ) at the 20th position from the 5` end of the gRNA target base sequence. 2 to C 20 , C 3 to C 20 , C 4 to C 20 , C 1 to C 19 , C 2 to C 19 , C 3 to C 19 , C 4 to C 19 , C 1 to C 18 , C 2 to C 18 , C 3 to C 18 , C 4 to C 18 , C 1 to C17 , C 2 to C 17 , C 3 to C 17 , C 4 to C 17 , C 1 to C 16 , C 2 to C 16 , C 3 to C 16 , C 4 to C 16 , C 1 to C 15 , C 2 to C 15 , C 3 to C 15 , C 4 to C 15 , C 1 to C 14 , C 2 to C 14 , C 3 to C 14 , C 4 to C 14 , C 1 to C 13 , C 2 to C 13 , C 3 to C 13 , C 4 to C 13 , C 1 to C 12 , C 2 to C 12 , C 3 to C 12 , C 4 to C 12 , C 1 to C 11 , C 2 to C 11 , C 3 to C 11 , C 4 to C 11 , C 1 to C 10 , C 2 to C 10 , C 3 to C 10 , C 4 to C 10 , C 1 to C 9 , C 2 to C 9, C 3 to C 9 , C 4 to C 9 , C 1 to C 8 , C 2 to C 8 , C 3 to C 8 , or C 4 to C 8 Base editing within the range.

[0083] Advantages of the Invention

[0084] According to one aspect, the fusion protein, its polypeptide, and the CRISPR-Cas system containing the same can perform effective base editing. In addition, due to the small size of the polypeptide, it is easy to bind to the Cas protein, and when used as a cell therapeutic agent, it has the effect of promoting delivery by a vector.

[0085] In addition, when inducing indels through this method, the indel efficiency and the size of the nucleotides forming the indels also increase, so it can be effectively used for gene knock-out. Brief Description of the Drawings

[0086] Figure 1 is a schematic diagram showing the structural / evolutionary differences between SsdA and other deaminases used in existing base editing technologies.

[0087] Figure 2 is a diagram showing the process of deaminating single-stranded DNA with the base sequence of SEQ ID NO: 4 using SsdA and cleaving the single-stranded DNA with UDG and NaOH, as well as a photograph showing the results of a protein immunoblot for confirming whether SsdA deaminates or not.

[0088] Figure 3 is a diagram showing the base editing efficiency of the CRISPR-Cas system containing a Cas protein, SsdA, and a guide polynucleotide at the target site of target DNA: Figure 3a is a diagram showing the base editing efficiency of the CRISPR-Cas system containing a Cas protein, SsdA, and a guide polynucleotide at the target site of RNF2 DNA, Figure 3b is a diagram showing the base editing efficiency of the CRISPR-Cas system containing a Cas protein, SsdA, and a guide polynucleotide at the target site of HEK2 DNA.

[0089] Figure 4It is a figure showing the cytotoxicity of Cas9(D10A)-SsdA fusion protein, Cas9(D10A)-SsdA-UGI fusion protein, dCas9-SsdA fusion protein, and dCas9-SsdA-UGI fusion protein respectively.

[0090] Figure 5 is a figure showing the base editing efficiency and indel formation efficiency of the CRISPR-Cas system containing the Cas9 and SsdA fusion protein:

[0091] Figure 5a It is a figure showing the base editing efficiency and indel formation efficiency of each CRISPR-Cas system containing Cas9(D10A)-SsdA fusion protein, Cas9(D10A)-SsdA-UGI fusion protein, dCas9-SsdA fusion protein, or dCas9-SsdA-UGI fusion protein in HEK2, HEK3, HEK4, and RNF2. Figure 5b It is a figure showing the base editing efficiency of each CRISPR-Cas system containing cjCas9(D8A)-SsdA fusion protein, cjCas9(D8A)-SsdA-UGI fusion protein, or cjCas9(L58Y / D900K)(D8A)-SsdA-UGI fusion protein in EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, and TFPi. Figure 5c It is a figure showing the base editing efficiency of each CRISPR-Cas system containing cjCas9(D8A)-SsdA fusion protein, cjCas9(D8A)-SsdA-UGI fusion protein, or cjCas9(L58Y / D900K)(D8A)-SsdA-UGI fusion protein in EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, and TFPi in a uracil-DNA glycosylase (UDG) expression-inhibited cell line.

[0092] Figure 6 It is a schematic diagram of a plasmid for preparing a fusion protein in which SsdA and uracil glycosylase inhibitor (UGI) are bound to the C-terminus, N-terminus, and both the N-terminus and C-terminus of the Cas protein.

[0093] Figure 7 is a figure showing the base editing efficiency confirmed according to the binding site of SsdA and the Cas protein:

[0094] Figure 7aGraph showing the cytosine base editing efficiency of each CRISPR-Cas system containing a Cas9(D10A)-SsdA fusion protein (SsdA-C), an SsdA-Cas9(D10A) fusion protein (SsdA-N, SsCBE), or an SsdA-Cas9(D10A)-SsdA (SsdA-NC) fusion protein in HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1. Figure 7b Graph showing the cytosine base editing efficiency of each CRISPR-Cas system containing a C-terminal, N-terminal, or both C-terminal and N-terminal fusion protein of SsdA-Cas9(D10A) (SsdA-N, SsCBE) bound to one or two UGIs in HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1. Figure 7c Graph showing the cytosine base editing efficiency of a UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) with two UGIs with significantly higher base editing efficiency bound to the N-terminal or a UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2) with two UGIs bound to each of the N-terminal and C-terminal for accurate comparison of base editing efficiency at 20 target sites (HEK2-1, HEK2-2, HEK2-3, HEK2-4, HEK3-1, HEK3-2, HEK3-3, HEK3-6, HEK3-7, HEK3-8, HEK4-1, HEK4-2, HEK4-3, HEK4-4, HEK4-5, HEK4-6, HEK4-7, HEK4-8, RNF2-3, RNF2-4). Figure 7d Is for Figure 7c Graph analyzing the characteristics of base editing technology, i.e., the editing range (editing window), for base editing at the 20 target sites shown.

[0095] Figure 8 Graph showing the substitution frequency (%) of a UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2), BE3, and BE4max at target sites (HEK2, HEK3, HEK2-2, HEK3-8, HEK4-2, or HEK4-7 sites).

[0096] Figure 9A graph showing the base editing efficiency and indel formation efficiency of each CRISPR-Cas system containing the Cas9(D10A)-SsdA(G54D)-UGI fusion protein, Cas9(D10A)-SsdA(G54D) fusion protein, dCas9-SsdA(G54D)-UGI fusion protein, or dCas9-SsdA(G54D) fusion protein in HEK2, HEK3, and HEK4. Detailed Description of the Invention

[0097] It will be described in more detail by the following examples. However, these examples are for illustrative purposes only, and the scope of the present invention is not limited to these examples.

[0098] Example 1 Plasmid Cloning

[0099] The gBlock double-stranded DNA fragment (Integrated DNA Technologies) encoding His6-SsdA with the amino acid sequence of SEQ ID NO: 1 and SsdAI with the amino acid sequence of SEQ ID NO: 3 and the bacterial expression vector pET-28b(+) DNA (Novagen) were respectively treated with XbaI and XhoI restriction enzymes (New England Biolabs) at 37 °C for 3 hours. Then, the linearized pET-28b vector was purified using agarose gel extraction (Geneall) and ligated to the gBlock double-stranded DNA fragment using a quick ligase (New England Biolabs).

[0100] Then, the coding sequences of Cas9 and UGI were obtained by PCR amplification using pCMV plasmid DNA, the SsdA sequence was amplified in the gBlock using Gibson Assembly Master Mix (New England Biolabs), and subcloned into the pCMV plasmid.

[0101] The amino acid sequence of the PAAR domain-containing protein of SEQ ID NO: 1 and the SsdAI amino acid sequence of SEQ ID NO: 3 are shown in Table 2.

[0102] Table 2

[0103]

[0104]

[0105] Example 2 Purification of SsdA

[0106] Purify the SsdA protein cloned in Example 1 using Escherichia coli BL21.

[0107] More specifically, pET-28b-His6-SsdA-SsdAI was applied to Escherichia coli BL21 using 0.5 mM IPTG, and the His6-SsdA and SsdAI complex protein was purified using Ni-NTA agarose beads (Qiagen). Then, in order to isolate His6-SsdA from SsdAI, the His6-SsdA and SsdAI complex protein was denatured with a denaturing buffer (8 M urea, 50 mM Tris-HCl pH 7.5, 500 mM NaCl, and 1 mM DTT) and incubated at 4 °C for 16 hours. The suspension buffer containing the denatured protein complex was mixed with Ni-NTA agarose beads (Qiagen) and loaded onto a gravity-flow column to remove the unbound SsdAI. Subsequently, urea with decreasing concentrations (6 M, 4 M, 2 M, 1 M, and 0 M) was used to treat the denaturing buffer in advance for the refolding of SsdA. After eluting the refolded protein bound to the Ni-NTA agarose beads with an elution buffer containing 300 mM imidazole, the eluted protein was dialyzed against 20 mM Tris-HCl pH 7.5, 200 mM NaCl, 1 mM DTT, and 40% (w / v) glycerol, and concentrated using an Amicon Ultra-15 centrifugal filtration device (Millipore). Then, the concentration of the His6-SsdA protein was analyzed by SDS-PAGE.

[0108] Example 3 Cell Culture and Transfection

[0109] HEK293T cells (ATCC CRL-11268) were maintained in Dulbecco's modified Eagle's medium (DMEM) containing 10% fetal bovine serum (FBS) and 1% penicillin / streptomycin (Welgene), and then the HEK293T cells were seeded at a density of 6×10 4Cells were seeded into a TC-treated 48-well plate (Corning Life Sciences). 24 hours after seeding, at a cell confluence of approximately 60%, transfection was performed using 500 ng of plasmid (250 ng of Cas9-SsdA expression plasmid and 250 ng of gRNA expression plasmid) and 1.5 μL of Lipofectamine 2000 (Thermo Fisher Scientific). Then, the transfected cells were incubated at 37 °C for 3 days, and the cells were directly lysed using lysis buffer (10 mM Tris-HCl, 0.05% SDS, 100 mg / mL proteinase K at pH 7.5; QIAGEN), and genomic DNA was prepared. The cell lysate was incubated at 56 °C for 30 minutes and further incubated at 99 °C for 15 minutes to inactivate proteinase K.

[0110] Example 4 Targeted deep sequencing and data analysis

[0111] The target sites were amplified by PCR (a total of 2 or 3 times) and sequenced using an Illumina MiniSeq or iSeq100 sequencing system.

[0112] More specifically, after applying 3 μL of cell lysate or 1 μL of isolated genomic DNA to the first PCR, 1 μL of the first PCR product was used for the second PCR. Using 1 μL of the second PCR product, the Illumina TruSeq HT dual-tag adapter sequences were attached to the tagged PCR primer pair. The size of the PCR amplicons was confirmed on a 2% agarose gel, and the amplicons were sequenced using an Illumina MiniSeq or iSeq100 sequencing system. Targeted deep sequencing analysis was performed using MAUND (https: / / github.com / ibscge / maund), and all results were confirmed by Cas-Analyzer (http: / / www.rgenome.net / cas-analyzer / ).

[0113] Example 5 Confirmation of cytosine deamination in single-stranded DNA of SsdA

[0114] To confirm whether SsdA converts the cytosine base of single-stranded DNA to uracil, single-stranded DNA containing FAM was treated with SsdA.

[0115] More specifically, after treating single-stranded DNA having a base sequence of SEQ ID NO: 4 (5'-Aaaaaaaaaaaaaaaagcgaaaaaaaaaaaaaaaaaaa-3') with 1 to 200 nM of SsdA at 37°C for 1 hour, it was treated with uracil DNA glycosylase (UDG) at 37°C for 30 minutes to remove DNA bases that had become uracil, thereby generating abasic sites. Then, it was treated with 100 mM NaOH and incubated at 95°C for 2 minutes to cleave the abasic sites. And the cleavage was confirmed by Western blot, and the results are as Figure 2 shown.

[0116] Figure 2 FIG. is a diagram showing the process of deaminating single-stranded DNA having a base sequence of SEQ ID NO: 4 with SsdA, treating and cleaving the single-stranded DNA with UDG and NaOH, and a photograph showing the Western blot results for confirming whether SsdA deamination occurred or not.

[0117] As Figure 2 shown, it was confirmed that DNA cleavage occurred only when the single-stranded DNA was treated with SsdA, UDG, and NaOH. This means that SsdA can effectively convert cytosine in single-stranded DNA to uracil by deamination.

[0118] Example 6 confirmed the base editing efficiency of the CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide

[0119] To confirm the base editing efficiency of the CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide, the target DNA (RNF2 and HEK2) was treated with the CRISPR-Cas system comprising SsdA, dCas9, and gRNA to determine whether deamination of cytosine occurred. The base sequences of the target DNA (RNF2 and HEK2) are shown in Table 3.

[0120] Table 3

[0121] Target DNA Target site SEQ ID NO RNF2 GTC3ATC6TTAGTC12ATTACCTGAGG SEQ ID NO: 5 HEK2 GAAC4AC6AAAGC11ATAGACTGCGGG SEQ ID NO: 6

[0122] More specifically, after treating RNF2 DNA and HEK2 DNA with 100 nM Cas9, 300 nM sgRNA, and 40 nM SsdA, respectively, they were incubated at 37 °C for 8 hours to induce the conversion of cytosine at the target site of the target DNA to uracil. Then, 8 hours later, sgRNA, Cas9, and SsdA were removed by treatment with RNase and Protease K. And the DNA was purified using a Qiagen DNA extraction kit. Subsequently, PCR was performed using primers containing the target site, and the base editing efficiency was measured by deep sequencing, and the results are shown in Figure 3.

[0123] Figure 3 is a graph showing the base editing efficiency of the CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide at the target site of a target DNA:

[0124] Figure 3a is a graph showing the base editing efficiency of the CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide at the target site of RNF2 DNA, Figure 3b is a graph showing the base editing efficiency of the CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide at the target site of HEK2 DNA.

[0125] As Figure 3a shown, when using the gRNA targeting RNF2, the efficiency of cytosine conversion to uracil was confirmed to be approximately 7% at C 3 and approximately 11% at C 12 . In addition, as Figure 3b shown, when using the gRNA targeting HEK2, the efficiency of cytosine conversion to uracil was confirmed to be approximately 18% to 19% at C 4 , C 6 , and C 11 .

[0126] This means that the CRISPR-Cas system can only perform gene editing on one strand of the DNA double strand. Therefore, considering that the maximum base editing efficiency is 50%, this result shows a rather high level of base editing efficiency.

[0127] In addition, further, to confirm whether the base conversion caused by the CRISPR-Cas system containing the Cas protein, SsdA, and the guide polynucleotide is the conversion of cytosine to uracil, the deaminated DNA was treated with a uracil-specific excision reagent (USER), and the target site was amplified by PCR. As a result, the reads of the conversion of cytosine to uracil disappeared.

[0128] These results indicate that the CRISPR-Cas system can effectively convert cytosine to uracil.

[0129] Example 7 confirmed the cytotoxicity of the fusion protein containing the Cas protein and SsdA

[0130] Confirmed the cytotoxicity of the fusion protein containing the Cas protein, SsdA, and UGI, and the results are shown in Figure 4 .

[0131] To confirm whether SsdA of the CRISPR-Cas system is toxic in eukaryotic cells, a plasmid expressing a protein that binds to Cas9 (D10A) was constructed. Then, HEK293 cells were dispensed into a 48-well plate at a concentration of 6×10 4 cells / well, and the plasmid was transfected into HEK293 cells. 48 hours after transfection, the living cells were digested with trypsin and the number of HEK293 cells was counted using a hemacytometer.

[0132] Figure 4 is a graph showing the cytotoxicity of each of the Cas9 (D10A)-SsdA fusion protein, Cas9 (D10A)-SsdA-UGI fusion protein, dCas9-SsdA fusion protein, and dCas9-SsdA-UGI fusion protein.

[0133] As Figure 4 shown, it was confirmed that the fusion protein containing Cas9, SsdA, and UGI has almost no cytotoxicity.

[0134] Example 8 confirmed the base editing efficiency and indel formation efficiency of the CRISPR-Cas system of the fusion protein containing the Cas protein and SsdA

[0135] To confirm whether SsdA of the CRISPR-Cas system can effectively cause cytosine deamination in eukaryotic cells, plasmids expressing proteins that bind to Cas9 were constructed and transfected into HEK293 cells. Then, the base sequences of the target sites (HEK2, HEK3, HEK4, RNF2, EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, or TFPi sites) of the transfected cells were amplified by PCR, and next-generation sequencing (NGS) was used to analyze the introduction of mutations and the formation of indels. The results are shown in Figure 5.

[0136] The Cas9 used was spCas9 (Cas9(D10A)) or cjCas9(D8A). The uracil-DNA glycosylase (UDG) knockdown cell line (UNG KD) was generated by transfecting HEK293 cells with a plasmid expressing shRNA (5`-GTCTACAGACATAGAGGATTT-3: SEQ ID NO: 7) that knockdowns UDG and an HIV-based packaging plasmid (containing genes: Gag / Pol, Rev, VSV-G) mixed with Lipofectamine 3000 (Invitrogen) reagent in Opti MEM (Invitrogen).

[0137] Figure 5 is a graph showing the base editing efficiency and indel formation efficiency of the CRISPR-Cas system containing the Cas9 and SsdA fusion protein:

[0138] Figure 5a is a graph showing the base editing efficiency and indel formation efficiency of each CRISPR-Cas system containing the Cas9(D10A)-SsdA fusion protein, Cas9(D10A)-SsdA-UGI fusion protein, dCas9-SsdA fusion protein, or dCas9-SsdA-UGI fusion protein in HEK2, HEK3, HEK4, and RNF2, Figure 5b is a graph showing the base editing efficiency of each CRISPR-Cas system containing the cjCas9(D8A)-SsdA fusion protein, cjCas9(D8A)-SsdA-UGI fusion protein, or cjCas9(L58Y / D900K)(D8A)-SsdA-UGI fusion protein in EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, and TFPi, Figure 5cGraph showing the base editing efficiency of each CRISPR-Cas system containing the cjCas9(D8A)-SsdA fusion protein, cjCas9(D8A)-SsdA-UGI fusion protein, or cjCas9(L58Y / D900K)(D8A)-SsdA-UGI fusion protein in EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, and TFPi in uracil-DNA glycosylase (UDG)-knockdown cell lines.

[0139] As Figure 5a shown, in the case of Cas9(D10A)-SsdA-UGI, cytosine base editing of 4% to 11% was confirmed at the HEK2, HEK3, HEK4, and RNF2 sites, and indels in the base sequence were formed at the target sites. In addition, in the case of dCas9-SsdA-UGI, cytosine base editing was confirmed at the 1% to 3.5% level, and no indels were formed.

[0140] As Figure 5b shown, in the cases of cjCas9(D8A)-SsdA-UGI and cjCas9(L58Y / D900K)(D8A)-SsdA-UGI, both caused base editing of approximately 3% and showed improved base editing efficiency compared to the fusion protein not bound to UGI.

[0141] As Figure 5c shown, in the UDG-knockdown cell lines, in the cases of cjCas9(D8A)-SsdA-UGI and cjCas9(L58Y / D900K)(D8A)-SsdA-UGI, both showed a base editing efficiency of approximately 15%, showing a significantly improved base editing efficiency compared to the cjCas9(D8A)-SsdA fusion protein not bound to UGI.

[0142] Table 4 shows the analysis results of the sequences of base editing generated by Cas9(D10A)-Ssda-UGI and the HEK2 target gRNA.

[0143] Table 4

[0144]

[0145] Example 9 confirmed the base editing efficiency according to the binding sites of SsdA and Cas proteins

[0146] Uracil-DNA glycosylase (UDG) is a well-known protein that repairs cytosine deamination generated in intracellular DNA. To accurately confirm that SsdA indeed causes cytosine deamination in intracellular DNA, a HEK293 cell line with knocked-out UDG (HEK293 UDG-KO) was first generated.

[0147] To confirm the base editing efficiency and indel formation efficiency of the fusion proteins of SsdA with the C-terminus, N-terminus, and both the N-terminus and C-terminus of the Cas protein, fusion proteins were prepared by changing the position of SsdA, and the base editing efficiency and indel formation efficiency were thereby confirmed.

[0148] More specifically, a pCMV plasmid as shown in Figure 6 was constructed and introduced into HEK293 UDG-KO cells together with the gRNA for each target site according to Example 3. Then, the base sequences of the target sites (HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1 sites) of the transfected cells were amplified by PCR, and the introduction of mutations and the formation of indels were analyzed using next-generation sequencing (NGS). The results are shown in Fig. 7.

[0149] Figure 6 is a schematic diagram showing the plasmids for preparing the fusion proteins of SsdA and uracil glycosylase inhibitor (UGI) bound to the C-terminus, N-terminus, and both the N-terminus and C-terminus of the Cas protein.

[0150] Fig. 7 is a graph showing the base editing efficiency confirmed according to the binding site of SsdA and the Cas protein:

[0151] Figure 7a is a graph showing the cytosine base editing efficiency of each CRISPR-Cas system containing the Cas9(D10A)-SsdA fusion protein (SsdA-C), SsdA-Cas9(D10A) fusion protein (SsdA-N, SsCBE), or SsdA-Cas9(D10A)-SsdA (SsdA-NC) fusion protein in HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1, Figure 7b is a graph showing the cytosine base editing efficiency of each CRISPR-Cas system containing the fusion proteins of the C-terminus, N-terminus, and both the C-terminus and N-terminus of the SsdA-Cas9(D10A) fusion protein (SsdA-N, SsCBE) bound with 1 or 2 UGIs in HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1, Figure 7cThe figure shows the cytosine base editing efficiency of the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) with two UGIs with significantly higher base editing efficiency bound to the N-terminus or the UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2) with two UGIs bound to each of the N-terminus and C-terminus, respectively, at 20 target sites (HEK2-1, HEK2-2, HEK2-3, HEK2-4, HEK3-1, HEK3-2, HEK3-3, HEK3-6, HEK3-7, HEK3-8, HEK4-1, HEK4-2, HEK4-3, HEK4-4, HEK4-5, HEK4-6, HEK4-7, HEK4-8, RNF2-3, RNF2-4) for accurate comparison of base editing efficiency. Figure 7d It is for Figure 7c The figure analyzes the characteristics of the base editing technology, i.e., the editing range (editing window), for the base editing of the 20 target sites shown.

[0152] The base sequences of the target sites of the target genes are shown in Table 5.

[0153] Table 5

[0154]

[0155]

[0156] As Figure 7a shown, in the case of SsdA-Cas9(D10A) (SsdA-C) where SsdA is bound to the C-terminus of Cas, it was confirmed that the cytosine base editing efficiency was slightly reduced compared to the cases where SsdA was bound to the N-terminus or both the N-terminus and C-terminus. On the other hand, in the case where SsdA was bound to the N-terminus of Cas (SsdA-N, SsCBE), an increase in base editing efficiency was confirmed.

[0157] As Figure 7b shown, when further binding UGI to SsdA-Cas9(D10A) (SsdA-N, SsCBE) with high base editing efficiency, the base editing efficiency was increased compared to the case without UGI. In the case of the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) with two UGIs bound to the N-terminus or the UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2) with two UGIs bound to each of the N-terminus and C-terminus, a significant increase in base editing efficiency was confirmed.

[0158] As Figure 7cAs shown, the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) that binds two UGIs with the highest base editing efficiency at the N-terminus; or the UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2) that binds two UGIs at both the N-terminus and the C-terminus; the results of comparison at 20 target sites show that the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) that binds two UGIs at the N-terminus; is confirmed to show slightly higher base editing efficiency.

[0159] As Figure 7d shown, the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) that binds two UGIs with the highest base editing efficiency at the N-terminus; or the UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2) that binds two UGIs at both the N-terminus and the C-terminus; the results of comparison of the editing window at 20 target sites show that at the sites between 4 bp and 8 bp from the 5` end of the gRNA target base sequence, it is confirmed to show high base editing efficiency.

[0160] Example 10 confirmed the base editing ability of the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) in wild-type HEK293 cells

[0161] It was confirmed whether the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) with the highest base editing ability also shows base editing efficiency in wild-type HEK293 cells.

[0162] More specifically, according to Example 3, the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2); and the existing BE3 or BE4max were introduced into HEK293 cells together with the gRNA of each target site. Then, the base sequences of the target sites (HEK2, HEK3, HEK2-2, HEK3-8, HEK4-2, or HEK4-7 sites) in the transfected cells were amplified by PCR, and the introduction of mutations and the formation of indels were analyzed by next-generation sequencing (NGS), and the results are shown in Figure 8 .

[0163] Figure 8It is a graph showing the substitution frequencies (%) of the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2), BE3, and BE4max at the target sites (HEK2, HEK3, HEK2-2, HEK3-8, HEK4-2, or HEK4-7 sites).

[0164] As Figure 8 shown, the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) successfully demonstrated base editing ability at six target sites (HEK2, HEK3, HEK2-2, HEK3-8, HEK4-2, or HEK4-7 sites).

[0165] Example 11 confirmed the cytotoxicity and base editing ability of the CRISPR-Cas system using the SsdA mutant

[0166] To confirm the cytotoxicity and base editing ability of the SsdA mutant, inactivated SsdA was produced and its effect was confirmed.

[0167] More specifically, after introducing the G54D mutation into the catalytic active site of SsdA, the mutated SsdA protein was bound to Cas9, and its intracellular function was confirmed. The results are shown in Figure 9 . The G54D mutation in the catalytic active site of SsdA is the G302D mutation in the amino acid sequence of SEQ ID NO: 1 (protein containing the PAAR domain). Specifically, the G54D mutation in the catalytic active site of SsdA may have the sequence of SEQ ID NO: 17. The amino acid sequence of the G54D mutation in the catalytic active site of SsdA is shown in Table 6.

[0168] Table 6

[0169]

[0170] Figure 9 It is a chart showing the base editing efficiency and indel formation efficiency of each CRISPR-Cas system containing the Cas9(D10A)-SsdA(G54D)-UGI fusion protein, Cas9(D10A)-SsdA(G54D) fusion protein, dCas9-SsdA(G54D)-UGI fusion protein, or dCas9-SsdA(G54D) fusion protein in HEK2, HEK3, and HEK4.

[0171] As Figure 9As shown, when the G54D mutation of SsdA binds to Cas9(D10A), indels occur with a very high efficiency of about 18% to 34% regardless of the presence or absence of UGI. This means that regardless of the UGI function, SsdA(G54D) can cause indels at the target site with higher efficiency.

[0172] In addition, when using SsdA(G54D), it was confirmed that not only was the intracellular toxicity significantly reduced, but also the size of the deleted base sequence was significantly increased.

[0173] Table 7 shows the analysis results of the base-edited sequences generated by SsdA(G54D)-Cas9(D10A).

[0174] As shown in Table 7, the results of analyzing the indel sequences generated by SsdA(G54D)-Cas9(D10A) confirmed that, unlike Cas9 which usually forms a 1bp deletion, in the case of SsdA(G54D)-Cas9(D10A), the size of the deleted base sequence was significantly increased. These results mean that the CRISPR-Cas system containing SsdA(G54D) can improve the gene knockout efficiency compared to the system that only uses Cas9.

[0175] Table 7

[0176]

Claims

1. A fusion protein comprising a Cas protein and a bacterial toxin, wherein, the bacterial toxin is single-stranded DNA deaminase toxin A.

2. The fusion protein according to claim 1, wherein, the Cas protein is Cas9 protein or Cas12 protein.

3. The fusion protein according to claim 1, wherein, the single-stranded DNA deaminase toxin A is cytidine deaminase.

4. The fusion protein according to claim 1, wherein, the single-stranded DNA deaminase toxin A is inactivated single-stranded DNA deaminase toxin A.

5. The fusion protein according to claim 1, wherein, the single-stranded DNA deaminase toxin A comprises the amino acid sequence of SEQ ID NO:

1.

6. The fusion protein according to claim 1, wherein, the single-stranded DNA deaminase toxin A comprises a sequence having at least 85% sequence identity with the amino acid sequence of SEQ ID NO:

1.

7. The fusion protein according to claim 4, wherein, the inactivated single-stranded DNA deaminase toxin A has an amino acid mutation at the catalytic active site of single-stranded DNA deaminase toxin A.

8. The fusion protein according to claim 4, wherein, the inactivated single-stranded DNA deaminase toxin A has G302D, E349A, or their corresponding amino acid mutations in the amino acid sequence of SEQ ID NO:

1.

9. The fusion protein according to claim 4, wherein, the inactivated single-stranded DNA deaminase toxin A has the amino acid sequence of SEQ ID NO:

17.

10. The fusion protein according to claim 4, wherein, the inactivated single-stranded DNA deaminase toxin A comprises a sequence having at least 85% sequence identity with the amino acid sequence of SEQ ID NO:

17.

11. The fusion protein according to claim 1, wherein, the bacterial toxin binds to the C-terminus, N-terminus, or both the C-terminus and N-terminus of the Cas protein.

12. The fusion protein according to claim 1, wherein, the fusion protein further comprises a DNA glycosylase inhibitor.

13. The fusion protein according to claim 12, wherein, the DNA glycosylase inhibitor is thymine glycosylase inhibitor, uracil glycosylase inhibitor, oxoguanine glycosylase inhibitor, or alkylguanine DNA glycosylase inhibitor.

14. The fusion protein according to claim 1, wherein, the fusion protein has an editing window ranging from 1 bp to 20 bp at the 5'-end of the target sequence.

15. A polynucleotide encoding the fusion protein according to any one of claims 1 to 14.

16. A vector comprising the polynucleotide according to claim 15.

17. A CRISPR-Cas system comprising a fusion protein or a polynucleotide encoding the fusion protein, and a guide polynucleotide, wherein, the fusion protein comprises a Cas protein and a bacterial toxin, and the bacterial toxin is single-stranded DNA deaminase toxin A.

18. The CRISPR-Cas system according to claim 17, wherein, The guide polynucleotide comprises a CRISPR RNA and a trans-activating RNA, and the guide polynucleotide is a double-stranded guide RNA or a single-stranded guide RNA.

19. The CRISPR-Cas system according to claim 17, wherein, the system forms a deletion, insertion, substitution or indel of at least one nucleotide in the nucleotide sequence of the target nucleic acid molecule.

20. The CRISPR-Cas system according to claim 19, wherein, the system forms a deletion of 1 bp to 60 bp of nucleotides, an insertion of 1 bp to 60 bp of nucleotides, or an indel of 1 bp to 60 bp of nucleotides in the nucleotide sequence of the target nucleic acid molecule.

21. The CRISPR-Cas system according to claim 17, wherein, the system has an editing window ranging from 1 bp to 20 bp at the 5'-end of the target sequence.

22. A method for editing a nucleic acid, comprising the step of contacting the nucleic acid molecule with a CRISPR-Cas system, wherein, the editing forms a deletion, insertion, substitution or indel of at least one nucleotide sequence in the nucleotide sequence of the nucleic acid molecule, the CRISPR-Cas system comprises a fusion protein or a polynucleotide encoding the fusion protein, and a guide polynucleotide, the fusion protein comprises a Cas protein and a bacterial toxin, and the bacterial toxin is a single-stranded DNA deaminase toxin A.

23. The method for editing a nucleic acid according to claim 22, wherein, the editing is a deletion of 1 bp to 60 bp of nucleotides, an insertion of 1 bp to 60 bp of nucleotides, or an indel of 1 bp to 60 bp of nucleotides in the nucleotide sequence of the nucleic acid molecule.