Fusion proteins comprising Cas proteins and bacterial toxins and uses thereof

A Cas protein-fusion with a modified bacterial toxin, like SsdA, addresses the production challenges of large CRISPR-Cas systems, enabling efficient and targeted nucleic acid editing in organisms.

JP2025535373APending Publication Date: 2025-10-24THE ASAN FOUND +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025522573
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-19
Filing Date
2023-10-19
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing CRISPR-Cas systems, such as those using APOBEC or AID proteins for cytosine base editing, are large and difficult to produce in vectors for cell therapy applications, necessitating a more efficient and producible system.

Method used

A fusion protein comprising a Cas protein, such as Cas9 or Cas12, and a bacterial toxin like SsdA, which is codon-optimized for eukaryotic expression and modified for reduced nuclease activity, is developed, along with a vector system for efficient delivery and editing.

Benefits of technology

The fusion protein enables efficient and targeted nucleic acid editing with reduced cytotoxicity, facilitating improved genome editing capabilities in various organisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535373000001_ABST
    Figure 2025535373000001_ABST
Patent Text Reader

Abstract

Regarding a fusion protein comprising a Cas protein and a bacterial toxin and its use, the fusion protein, its polypeptide, and a CRISPR-Cas system comprising them according to one embodiment enable effective base proofreading. Furthermore, the polypeptide is small in size, making it easy to deliver via a vector. This significantly increases indel efficiency when indels are induced, and the size of the nucleotide at which indels are formed also increases, making it effective for gene knockout.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a fusion protein comprising a Cas protein and a bacterial toxin and uses thereof. [Background technology]

[0002] Genome editing is a technology that freely edits the genetic information of living organisms. Advances in the field of life sciences and the development of genome sequence analysis technology have enabled us to gain a broad understanding of various genetic information. For example, we already have a solid understanding of the genes that govern plant and animal reproduction, disease and growth, genetic mutations that cause various human genetic diseases, and biofuel production. However, further technological advances are required to directly utilize this knowledge to improve living organisms and treat human diseases.

[0003] Genome editing technology can dramatically expand the scope of applications by modifying the genetic information of animals, plants, and microorganisms, including humans. Genetic scissors are molecular tools designed to precisely cut desired genetic information and play a vital role in genome editing technology. Like next-generation sequencing, which has one-dimensionally advanced the field of gene sequence information analysis, genetic scissors are expanding the speed and scope of genetic information utilization, becoming a core technology that will create new industrial fields.

[0004] However, the APOBEC or AID proteins commonly used for cytosine base editing using Cas9 are relatively large and difficult to produce in vectors for use as cell therapy. Therefore, there is a need for a CRISPR-Cas system that is easy to produce in vectors and has excellent base editing efficiency. Summary of the Invention [Problem to be solved by the invention]

[0005] One aspect is to provide a fusion protein comprising a Cas protein (CRISPR-associated protein) and a bacterial toxin.

[0006] Another aspect is to provide a polynucleotide encoding said fusion protein.

[0007] Another aspect is to provide a vector comprising said polynucleotide.

[0008] Another aspect is to provide a CRISPR-Cas system comprising the fusion protein or a polynucleotide encoding the same, and a guide polynucleotide.

[0009] Another aspect is to provide a method of editing a nucleic acid comprising contacting a nucleic acid molecule with said CRISPR-Cas system. [Means for solving the problem]

[0010] One aspect provides a fusion protein comprising a Cas protein (CRISPR-associated protein) and a bacterial toxin.

[0011] As used herein, the term "CRISPR-associated protein" refers to a CRISPR-associated endonuclease that can cleave all or part of a specific target polynucleotide sequence.

[0012] The Cas protein can be a Class 2 Cas (Class 2 CRISPR associated system) protein. The Class 2 Cas protein can be included in a Type II, Type V, or Type VI system.

[0013] The type II system can have the cas1, cas2, and cas9 genes. Type II systems can be further divided into three subtypes: subtype II-A, II-B, and II-C. Subtype II-A can contain an additional gene, csn2. Organisms with subtype II-A systems can include Streptococcus thermophilus. Subtype II-B systems lack csn2 but can have cas4. Organisms with subtype II-B systems can include Legionella pneumophila. Subtype II-C, the most common type II system found in bacteria, can have only three proteins: Cas1, Cas2, and Cas9. Organisms with subtype II-C systems can include Neisseria lactamica.

[0014] The Type V system can have the cas12 gene and the cas1 and cas2 genes. The cas12 gene can encode a protein, Cas12, that has RuvC-like nuclease domains homologous to regions of Cas9 but lacks the HNH nuclease domain present in the Cas9 protein.

[0015] The type VI system can have the cas13 gene and the cas1 and cas2 genes.

[0016] In the Type II system, the RuvC-like nuclease (RNase H-fold domain and HNH (McrA-like) nuclease) domain of Cas9 can each cleave one strand of the target nucleic acid. The Cas9 cleavage activity in Type II systems may also require hybridization of the crRNA with the tracrRNA to form a duplex that facilitates crRNA and target binding by Cas9.

[0017] In the Type V system, the RuvC-like nuclease domain of Cas12 can cleave both strands of a target nucleic acid in a staggered manner to generate 5' overhangs. Such 5' overhangs can facilitate DNA insertion by non-homologous end joining. The Cas12 cleavage activity of the Type V system also does not require hybridization of the crRNA to the tracrRNA to form a duplex; instead, the crRNA in the Type V system can use a single crRNA with a stem-loop structure that forms an internal duplex. The Type V system can induce single- or double-strand breaks at the target sequence. The strand breaks can be staggered cleavages using the 5' overhangs.

[0018] The Cas protein may include Cas9 or Cas12.

[0019] The Cas12 protein may refer to proteins derived from various bacterial species, including Streptococcus, Campylobacter, Nitratifractor, Staphylococcus, Parvibaculum, Roseburia, Neisseria, Gluconacetobacter, Azospirillum, Sphaerocerca, and the like. Sphaerochaeta, Lactobacillus, Eubacterium, Corynebacter, Carnobacterium, Rhodobacter, Listeria, Paludibacter, Clostridium, Lachnospiraceae, Clostridia Clostridiaridium, Leptotrichia, Francisella, Legionella, Alicyclobacillus, Methanomethyophilus, Porphyromonas, Prevotella, Bacteroidetes, Helicococcus ccus, Letospira, Desulfovibrio, Desulfonatronum, Opitutaceae, Tuberibacillus, Bacillus, Brevibacilus, Methylobacterium, or Acidaminococcus. More specifically, the Cas12 protein is derived from Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella arvensis, Lachnospiraceae bacteria. The bacterial species may be selected from the group consisting of Butyrivibrio proteoclasticus, Peregrinibacterium sp. GW2011_GWA2_33_10, Parcubacterium bacterium GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella boehmii 237, Leptospira inadae, Lachnospiraceae bacterium ND2006, Porphyromonas creviolicanis 3, Prevotella dysciens, and Porphyromonas macacae.

[0020] The Cas12 protein may be any one selected from the group consisting of Cas12a, mgCas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, and Cas12j. The Cas12 protein may include a modified Cas12 protein. If the Cas12 protein has nuclease activity, the Cas12 protein may be modified to have reduced nuclease activity, for example, at least 70%, at least 80%, at least 90%, at least 95%, at least 97%, or 100% nuclease inactivation compared to the wild-type enzyme.

[0021] The Cas9 may be Streptococcus pneumoniae (S. pneumoniae), Streptococcus pyogenes (S. pyogenes), Streptococcus thermophilus (S. thermophilus), or Campylobacter jejuni (C. jejuni) Cas9, including mutated Cas9s derived from these organisms. The enzyme may be a Cas9 homolog or ortholog. In one embodiment, the CRISPR enzyme can be codon-optimized for expression in eukaryotic cells. In one embodiment, the CRISPR enzyme can induce single- or double-stranded breaks at the target sequence. The Cas9 protein may include modifications of the Cas9 protein. If the Cas9 protein has nuclease activity, the Cas9 protein may be modified to have reduced nuclease activity, for example, at least 70%, at least 80%, at least 90%, at least 95% or more, at least 97%, or 100% nuclease inactivation compared to the wild-type enzyme. According to one embodiment, the Cas9 may be Cas9 D10A.

[0022] In one embodiment, the Cas protein may be a Cas9 protein or a Cas12 protein.

[0023] In one embodiment of the present invention, at least one nuclear localization signal (NLS) can be added to the nucleic acid sequence encoding the Cas protein. In one embodiment, at least one or more NLSs can be added to the C-terminus or N-terminus. Cas proteins, orthologues, or homologs thereof, can be encoded that contain one or more nuclear localization sequences (NLSs), for example, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs. In a preferred embodiment of the Cas protein complex described herein, the codon-optimized Cas protein can contain an NLS added to the C-terminus of the protein. In certain embodiments, other localization tags can be fused to the Cas protein to localize the Cas to specific locations within the cell, such as, but not limited to, organelles, such as mitochondria, plastids, chloroplasts, vesicles, Golgi, (nuclear or cytoplasmic) membranes, ribosomes, nucleosomes, ER, cytoskeleton, vacuoles, centrosomes, nucleosomes, granules, centrioles, etc.

[0024] The Cas protein and bacterial toxin can be fused via a linker. The linker can be located at the C-terminus, N-terminus, or both the C-terminus and N-terminus of the Cas protein, and the bacterial toxin can be bound to the Cas protein via the linker. Suitable linker motifs and linker configurations include those described in Chen et al., Fusion protein linkers: property, design, and functionality. Adv Drug Deliv Rev. 2013;65(10):1357-69, the entire contents of which are incorporated herein by reference.

[0025] In one embodiment, the bacterial toxin may be SsdA (single-stranded DNA deaminase toxin A).

[0026] The SsdA may be derived from a Pseudomonas sp. strain. The Pseudomonas sp. strain may be Pseudomonas syringae, Pseudomonas congelans, Pseudomonas savastanoi, Pseudomonas coronafaciens, Pseudomonas fluorescens, Pseudomonas sp. MPC6, Pseudomonas sp. GL-R-26, or Pseudomonas sp. GL-RE-26.

[0027] SsdA may have a PAAR domain at its N-terminus and a DYW deaminase domain at its C-terminus. SsdA is similar to deaminases used in existing base proofreading in that it contains the common amino acid motifs HxE and CxxC, but differs from existing deaminases in that it also contains an SGW motif. SsdA is classified as a DYW-like deaminase, a deaminases that is structurally and evolutionarily different from other deaminases used in existing base proofreading technologies. The phylogenetic tree and differences in constituent domains between SsdA and other deaminases used in existing base proofreading technologies are shown in Figure 1.

[0028] According to one embodiment, the SsdA may comprise the amino acid sequence of SEQ ID NO:1.

[0029] The SsdA may contain a toxin domain. The toxin domain of SsdA is a portion having deaminase activity and may have a length of 100 to 200 amino acids, for example, 120 to 180 amino acids, 120 to 170 amino acids, 130 to 160 amino acids, or 140 to 160 amino acids. Specifically, the toxin domain of SsdA may contain the amino acid sequence of SEQ ID NO: 2. The nucleotide sequence of the toxin domain is shown in Table 1 below.

[0030] [Table 1]

[0031] The amino acid sequence of SsdA of SEQ ID NO: 2 (toxin domain) may contain a catalytically active site. The catalytically active site may contain an HxE motif, a CxxC motif, or an SGW motif. The SGW motif is a motif that is unique to the SsdA enzyme, unlike existing deaminase enzymes such as APOBEC and AID. Specifically, the HxE motif may include amino acids 301 to 303 of SEQ ID NO: 1 (PAAR domain-containing protein), the HxE motif may include amino acids 347 to 349, and the SGW motif may include amino acids 301 to 303.

[0032] In one embodiment, the SsdA may have at least 80% or more, 85% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more sequence identity with the amino acid sequence of SEQ ID NO: 1 (PAAR domain-containing protein).

[0033] In one embodiment, the fusion protein may be a Cas protein bound to a DYW deaminase protein.

[0034] In one embodiment, the SsdA can be an inactivated SsdA.

[0035] The inactivated SsdA may have an amino acid mutation in the catalytically active site of activated SsdA, and may have lower cytotoxicity than activated SsdA.

[0036] According to one embodiment, the inactivated SsdA may have amino acid mutations at positions G302 and E349 in the amino acid sequence of SEQ ID NO: 1 (PAAR domain-containing protein). The amino acid mutations refer to substitutions of amino acids other than those present at positions G302 and E349 in the wild-type protein. The other amino acid may be any amino acid selected from the group consisting of arginine (R), histidine (H), lysine (K), aspartic acid (D), glutamic acid (E), serine (S), threonine (T), asparagine (N), glutamine (Q), cysteine ​​(C), selenocysteine ​​(U), glycine (G), proline (P), alanine (A), valine (V), isoleucine (I), leucine (L), methionine (M), phenylalanine (F), tyrosine (Y), tryptophan (W), and all variants of the above amino acids, excluding the amino acid present at the mutated position in the wild-type protein. Specifically, the inactivated SsdA may have G302D, E349A, or a corresponding amino acid mutation in the amino acid sequence of SEQ ID NO: 1. More specifically, the inactivated SsdA having a G302D mutation in the amino acid sequence of SEQ ID NO: 1 (PAAR domain-containing protein) can be SEQ ID NO: 17.

[0037] In one embodiment, the inactivated SsdA may have at least 80% or more, 85% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more sequence identity with the amino acid sequence of SEQ ID NO: 17.

[0038] The SsdA can induce deamination of single-stranded DNA, and the SsdA can be a cytidine deaminase.

[0039] As used herein, the term "cytidine deaminase" refers to an enzyme that has the activity of removing the amino (-NH2) group of cytosine, cytidine, or deoxycytidine. As used herein, the term "cytidine deaminase" encompasses cytosine deaminase. As used herein, the term "cytidine deaminase" may be used interchangeably with the term "cytosine deaminase."

[0040] The cytidine deaminase refers to any enzyme that has the activity of converting cytosine, a base present in a nucleotide (e.g., cytosine present in DNA or RNA), to uracil (C-to-U conversion or C-to-U editing), and converts cytosine located in the strand containing the PAM sequence of the target site sequence (target nucleic acid sequence) to uracil.

[0041] The bacterial toxin can be attached to the terminus of the Cas protein, for example, the bacterial toxin can be attached to the C-terminus, the N-terminus, or both the C-terminus and the N-terminus of the Cas protein.

[0042] In one embodiment, the fusion protein may further comprise a DNA glycosylase inhibitor.

[0043] The DNA glycosylase inhibitor can be a thymine glycosylase inhibitor, a uracil glycosylase inhibitor, an oxoguanine glycosylase inhibitor, or an alkylguanine DNA glycosylase inhibitor.

[0044] The uracil DNA glycosylase inhibitor can be, but is not limited to, a uracil DNA glycosylase inhibitor derived from Bacillus subtilis bacteriophage, PBS1, a uracil DNA glycosylase inhibitor derived from Bacillus subtilis bacteriophage, or PBS2.

[0045] Another aspect provides a polynucleotide encoding the fusion protein.

[0046] Another aspect provides a vector comprising the polynucleotide.

[0047] As used herein, the term "vector" can refer to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors can include nucleic acid molecules that are single-stranded, double-stranded, or partially double-stranded; nucleic acid molecules containing one or more free ends; nucleic acid molecules without free ends (e.g., circular); nucleic acid molecules containing DNA, RNA, or both; and various other polynucleotides known in the art. One type of vector is a "plasmid," which can refer to a circular double-stranded DNA loop into which additional DNA segments can be inserted, for example, by standard molecular cloning techniques. Another type of vector is a viral vector, in which viral-derived DNA or RNA sequences can be present in a vector for packaging into a virus (e.g., retrovirus, replication-defective retrovirus, adenovirus, replication-defective adenovirus, and adeno-associated virus). A recombinant expression vector can contain a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, meaning that the recombinant expression vector also contains one or more regulatory elements, which can be selected based on the host cell used for expression and operably linked to the nucleic acid sequence to be expressed. Within a recombinant expression vector, "operably linked" may mean that the nucleotide sequence of interest is linked to regulatory elements that allow for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system or in a host cell when the vector is introduced into the host cell).

[0048] As used herein, the term "regulatory element" may include promoters, enhancers, internal ribosomal entry sites (IRES), and other expression control elements (e.g., transcription termination signals, e.g., polyadenylation signals and poly-U sequences). Regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif. (1990). Regulatory elements may include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of a nucleotide sequence only in specific host cells (e.g., tissue-specific regulatory sequences). In some embodiments, a vector can include one or more Pol III promoters (e.g., 1, 2, 3, 4, 5, or more Pol III promoters), one or more Pol II promoters (e.g., 1, 2, 3, 4, 5, or more Pol II promoters), one or more Pol I promoters (e.g., 1, 2, 3, 4, 5, or more Pol I promoters), or a combination thereof. Examples of Pol III promoters can include, but are not limited to, the U6 and H1 promoters. Examples of Pol II promoters can include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the β-actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1α promoter. For example, vectors can include lentiviruses and adeno-associated viruses (AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, or AAV9), and such vector types can also be selected for targeting specific cell types.

[0049] Furthermore, the multiple nucleic acid molecules within the vector system can be located on the same or different vectors.

[0050] In one embodiment, a vector, e.g., a plasmid or viral vector, is delivered to a tissue of interest, e.g., by intramuscular injection; in other cases, delivery can be intravenous, transdermal, intranasal, oral, mucosal, or other delivery methods. Such delivery can be via either a single dose or multiple doses. Those skilled in the art will appreciate that the actual dosage delivered herein can vary widely depending on a variety of factors, such as the vector selection, the target cell, organism, or tissue, the general condition of the subject being treated, the degree of transformation / modification obtained, the route and method of administration, and the form of transformation / modification obtained. Dosages can further contain, for example, carriers (e.g., water, saline, ethanol, glycerol, lactose, sucrose, calcium phosphate, gelatin, dextran, agar, pectin, peanut oil, sesame oil, etc.), diluents, pharmaceutically acceptable carriers (e.g., phosphate-buffered saline), pharmaceutically acceptable excipients, and / or other compounds known in the art. The dosage may further contain one or more pharmaceutically acceptable salts, such as inorganic acid salts, e.g., hydrochlorides, hydrobromides, phosphates, sulfates, etc., and organic acid salts, e.g., acetates, propanoates, malonates, benzoates, etc. Additionally, auxiliary ingredients, such as wetting or emulsifying agents, pH buffering components, gels or gelling materials, flavoring agents, coloring agents, microspheres, polymers, suspending agents, etc., may also be present herein. Additionally, one or more other conventional pharmaceutical ingredients, such as preservatives, humectants, suspending agents, surfactants, antioxidants, fillers, chelating agents, coating agents, chemical stabilizers, etc., may also be present. Suitable exemplary ingredients include microcrystalline cellulose, sodium carboxymethylcellulose, polysorbate 80, phenylethyl alcohol, chlorobutanol, potassium sorbate, sorbic acid, sulfur dioxide, propyl gallate, parabens, ethyl vanillin, glycerin, phenol, parachlorophenol, gelatin, albumin, and combinations thereof. For example, delivery for disease treatment can be via AAV. A therapeutically effective dose for in vivo delivery of AAV to humans can be in the range of about 20 ml to about 50 ml of saline solution containing about 1 Y10 10 to about 1 Y10 100 AAV per ml of solution.Dosage can be adjusted to balance the therapeutic benefit against any side effects.

[0051] Another aspect provides a CRISPR-Cas system comprising a fusion protein comprising a Cas protein and a bacterial toxin, or a polynucleotide encoding the fusion protein, and a guide polynucleotide.

[0052] The fusion protein and the polynucleotide encoding it are as described above.

[0053] The guide polynucleotide may comprise a targeting sequence and / or an activating sequence.

[0054] As used herein, the term "targeting sequence" can refer to a polynucleotide comprising DNA or a mixture of DNA and RNA that is complementary to a sequence within a target nucleic acid. In certain embodiments, the targeting sequence can also comprise other nucleic acids, nucleic acid analogs, or combinations thereof. In certain embodiments, the targeting sequence can consist solely of DNA, as such constructs are less likely to be degraded inside host cells. In some embodiments, such constructs can increase / enhance the specificity of target sequence recognition or reduce the occurrence of off-target binding / hybridization. The targeting sequence can also include a guide sequence or spacer sequence. The length of the targeting sequence domain can be at least 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides.

[0055] As used herein, the term "activation sequence" can refer to a portion of a polynucleotide comprising RNA or DNA, or a mixture of DNA and RNA, that can interact with, associate with, or bind to a Cas protein. In one embodiment, the activation sequence can also comprise other nucleic acids, nucleic acid analogs, or combinations thereof. In one embodiment, the activation sequence can be adjacent to or linked to a target sequence. In one embodiment, the activation region can be downstream of a targeting region. In one embodiment, the activation region can be upstream of a targeting region. The activation sequence can include direct repeats, CRISPR RNA (crRNA) and / or trans-activating RNA (tracrRNA).

[0056] In one embodiment, the guide polynucleotide, guide RNA, mature crRNA, and premature crRNA can comprise or consist of direct repeat sequences and guide or spacer sequences. In one embodiment, the guide RNA or mature crRNA can comprise or consist of direct repeat sequences linked to a guide or spacer sequence. In one embodiment, the direct repeat sequences can be located upstream (i.e., 5') from the guide or spacer sequence.

[0057] In one embodiment, the guide polynucleotides may include crRNA and tracrRNA.

[0058] In one embodiment, the guide polynucleotide may be a dual guide RNA or a single-chain guide RNA (sgRNA).

[0059] The system can form deletions, insertions, substitutions, or indels of at least one nucleotide in the nucleotide sequence of a target nucleic acid molecule.

[0060] The nucleic acid can be RNA or DNA.

[0061] According to one embodiment, the system may be configured to detect a sequence of nucleotides in the target nucleic acid molecule that is between 1 bp and 60 bp, e.g., between 1 bp and 55 bp, between 1 bp and 50 bp, between 1 bp and 45 bp, between 1 bp and 40 bp, between 1 bp and 35 bp, between 1 bp and 30 bp, between 1 bp and 25 bp, between 1 bp and 20 bp, between 1 bp and 15 bp, between 1 bp and 10 bp, between 1 bp and 5 bp, between 5 bp and 60 bp, between 5 bp and 55 bp, between 5 bp and 50 bp, between 5 bp and 45 bp, between 5 bp and 40 bp, between 5 bp and 35 bp, between 5 bp and 30 bp, 0bp, 5bp~25bp, 5bp~20bp, 5bp~15bp, 5bp~10bp, 10bp~60bp, 10bp~55bp, 10bp~50bp, 10bp~45bp, 10bp~40bp, 10bp~35bp, 10 bp~30bp, 10bp~25bp, 10bp~20bp, 10bp~15bp, 15bp~60bp, 15bp~55bp, 15bp~50bp, 15bp~45bp, 15bp~40bp, 15bp~35bp, 15bp~ 30bp, 15bp~25bp, 15bp~20bp, 20bp~60bp, 20bp~55bp, 20bp~50bp, 20bp~45bp, 20bp~40bp, 20bp~35bp, 20bp~30bp, 20bp~25 bp, 25bp~60bp, 25bp~55bp, 25bp~50bp, 25bp~45bp, 25bp~40bp, 25bp~35bp, 25bp~30bp, 30bp~60bp, 30bp~55bp, 30bp~50bp, The deletion may be of nucleotides of 30bp to 45bp, 30bp to 40bp, 30bp to 35bp, 35bp to 60bp, 35bp to 55bp, 35bp to 50bp, 35bp to 45bp, 35bp to 40bp, 40bp to 60bp, 40bp to 55bp, 40bp to 50bp, 40bp to 45bp, 45bp to 60bp, 45bp to 55bp, 45bp to 50bp, 50bp to 60bp, 50bp to 55bp, or 55bp to 60bp.

[0062] According to one embodiment, the system detects a sequence of nucleotides in the target nucleic acid molecule that is between 1 bp and 60 bp, for example, between 1 bp and 55 bp, between 1 bp and 50 bp, between 1 bp and 45 bp, between 1 bp and 40 bp, between 1 bp and 35 bp, between 1 bp and 30 bp, between 1 bp and 25 bp, between 1 bp and 20 bp, between 1 bp and 15 bp, between 1 bp and 10 bp, between 1 bp and 5 bp, between 5 bp and 60 bp, between 5 bp and 55 bp, between 5 bp and 50 bp, between 5 bp and 45 bp, between 5 bp and 40 bp, between 5 bp and 35 bp, between 5 bp and 5 bp. p~30bp, 5bp~25bp, 5bp~20bp, 5bp~15bp, 5bp~10bp, 10bp~60bp, 10bp~55bp, 10bp~50bp, 10bp~45bp, 10bp~40bp, 10bp~35bp, 10bp~30bp, 10bp~25bp, 10bp~20bp, 10bp~15bp, 15bp~60bp, 15bp~55bp, 15bp~50bp, 15bp~45bp, 15bp~40bp, 15bp~35bp, 15b p~30bp, 15bp~25bp, 15bp~20bp, 20bp~60bp, 20bp~55bp, 20bp~50bp, 20bp~45bp, 20bp~40bp, 20bp~35bp, 20bp~30bp, 20bp~2 5bp, 25bp~60bp, 25bp~55bp, 25bp~50bp, 25bp~45bp, 25bp~40bp, 25bp~35bp, 25bp~30bp, 30bp~60bp, 30bp~55bp, 30bp~50bp , 30bp to 45bp, 30bp to 40bp, 30bp to 35bp, 35bp to 60bp, 35bp to 55bp, 35bp to 50bp, 35bp to 45bp, 35bp to 40bp, 40bp to 60bp, 40bp to 55bp, 40bp to 50bp, 40bp to 45bp, 45bp to 60bp, 45bp to 55bp, 45bp to 50bp, 50bp to 60bp, 50bp to 55bp, or 55bp to 60bp.

[0063] According to one embodiment, the system detects a sequence of nucleotides in the target nucleic acid molecule that is between 1 bp and 60 bp, for example, between 1 bp and 55 bp, between 1 bp and 50 bp, between 1 bp and 45 bp, between 1 bp and 40 bp, between 1 bp and 35 bp, between 1 bp and 30 bp, between 1 bp and 25 bp, between 1 bp and 20 bp, between 1 bp and 15 bp, between 1 bp and 10 bp, between 1 bp and 5 bp, between 5 bp and 60 bp, between 5 bp and 55 bp, between 5 bp and 50 bp, between 5 bp and 45 bp, between 5 bp and 40 bp, between 5 bp and 35 bp, between 5 bp and 5 bp. ~30bp, 5bp~25bp, 5bp~20bp, 5bp~15bp, 5bp~10bp, 10bp~60bp, 10bp~55bp, 10bp~50bp, 10bp~45bp, 10bp~40bp, 10bp~35bp, 1 0bp~30bp, 10bp~25bp, 10bp~20bp, 10bp~15bp, 15bp~60bp, 15bp~55bp, 15bp~50bp, 15bp~45bp, 15bp~40bp, 15bp~35bp, 15bp ~30bp, 15bp~25bp, 15bp~20bp, 20bp~60bp, 20bp~55bp, 20bp~50bp, 20bp~45bp, 20bp~40bp, 20bp~35bp, 20bp~30bp, 20bp~25 bp, 25bp~60bp, 25bp~55bp, 25bp~50bp, 25bp~45bp, 25bp~40bp, 25bp~35bp, 25bp~30bp, 30bp~60bp, 30bp~55bp, 30bp~50bp, The indel may be formed from nucleotides of 30bp to 45bp, 30bp to 40bp, 30bp to 35bp, 35bp to 60bp, 35bp to 55bp, 35bp to 50bp, 35bp to 45bp, 35bp to 40bp, 40bp to 60bp, 40bp to 55bp, 40bp to 50bp, 40bp to 45bp, 45bp to 60bp, 45bp to 55bp, 45bp to 50bp, 50bp to 60bp, 50bp to 55bp, or 55bp to 60bp.

[0064] According to one embodiment, the system has an efficiency of forming nucleotide indels of between 5% and 50%, for example, between 5% and 45%, between 5% and 40%, between 5% and 35%, between 5% and 30%, between 5% and 25%, between 5% and 20%, between 5% and 15%, between 5% and 10%, between 10% and 50%, between 10% and 45%, between 10% and 40%, between 10% and 35%, between 10% and 30%, between 10% and 25%, between 10% and 20%, between 10% and 15%, between 15% and 50%, between 15% and 45%, between 15% and 40%, between 15% and 35%, It can be 15%-30%, 15%-25%, 15%-20%, 20%-50%, 20%-45%, 20%-40%, 20%-35%, 20%-30%, 20%-25%, 25%-50%, 25%-45%, 25%-40%, 25%-35%, 25%-30%, 30%-50%, 30%-45%, 30%-40%, 30%-35%, 35%-50%, 35%-45%, 35%-40%, 40%-50%, 40%-45%, or 45%-50%.

[0065] According to one embodiment, the system is configured to generate substitutions in a sequence of nucleotides with an efficiency of 1% to 20%, for example, 1% to 18%, 1% to 16%, 1% to 14%, 1% to 12%, 1% to 10%, 1% to 8%, 1% to 6%, 1% to 4%, 1% to 2%, 2% to 20%, 2% to 18%, 2% to 16%, 2% to 14%, 2% to 12%, 2% to 10%, 2% to 8%, 2% to 6%, 2% to 4%, 4% to 20%, 4% to 18%, 4% to 16%, 4% to 14%, 4% to 12%, 4% to 10%, 4% to 8%, 4% to It can be 6%, 6%-20%, 6%-18%, 6%-16%, 6%-14%, 6%-12%, 6%-10%, 6%-8%, 8%-20%, 8%-18%, 8%-16%, 8%-14%, 8%-12%, 8%-10%, 10%-20%, 10%-18%, 10%-16%, 10%-14%, 10%-12%, 12%-20%, 12%-18%, 12%-16%, 12%-14%, 14%-20%, 14%-18%, 14%-16%, 16%-20%, 16%-18%, or 18%-20%.

[0066] The system can generate an editing window of at least four nucleotides in the nucleotide sequence of a target nucleic acid molecule. According to one embodiment, the system provides a proofreading window of at least 50 nucleotides, e.g., at least 49 nucleotides, at least 48 nucleotides, at least 47 nucleotides, at least 46 nucleotides, at least 45 nucleotides, at least 44 nucleotides, at least 43 nucleotides, at least 42 nucleotides, at least 41 nucleotides, at least 40 nucleotides, at least 39 nucleotides, at least 38 nucleotides, at least 37 nucleotides, at least 36 nucleotides, at least 35 nucleotides, at least 34 nucleotides, at least 33 nucleotides, at least 32 nucleotides, at least 31 nucleotides, at least 30 nucleotides, at least 29 nucleotides, at least 28 nucleotides, at least 27 nucleotides, at least 26 nucleotides, at least 25 nucleotides, at least 24 nucleotides, at least 23 nucleotides, at least 22 nucleotides, at least 21 nucleotides, at least 20 nucleotides, at least 19 nucleotides, at least 18 nucleotides, at least 17 nucleotides, at least 16 nucleotides, at least 15 nucleotides, at least 14 nucleotides, at least 13 nucleotides, at least 12 nucleotides, at least 11 nucleotides, at least 10 nucleotides, at least 9 nucleotides, at least 8 nucleotides, at least 7 nucleotides, at least 6 nucleotides, at least 5 nucleotides, or at least 4 nucleotides. The optical fiber may have a window.

[0067] According to one embodiment, the system is configured to amplify 1 bp to 20 bp, 1 bp to 19 bp, 1 bp to 18 bp, 1 bp to 17 bp, 1 bp to 16 bp, 1 bp to 15 bp, 1 bp to 14 bp, 1 bp to 13 bp, 1 bp to 12 bp, 1 bp to 11 bp, 1 bp to 10 bp, 1 bp to 9 bp, 1 bp to 8 bp, 2 bp to 20 bp, 2 bp to 19 bp, 2 bp to 18 bp, 2 bp to 17 bp, 2 bp to 16 bp, 2 bp to 15 bp, 2 bp to 14 bp, 2 bp to 13 bp, 2 bp to 12 bp, 2 bp to 11 bp, 2 bp to 10 bp, or 2 bp to 9 bp from the 5' end of the gRNA target sequence. , 2bp to 8bp, 3bp to 20bp, 3bp to 19bp, 3bp to 18bp, 3bp to 17bp, 3bp to 16bp, 3bp to 15bp, 3bp to 14bp, 3bp to 13bp, 3bp to 12bp, 3bp to 11bp, 3bp to 10bp, 3bp to 9bp, 3bp to 8bp, 4bp to 20bp, 4bp to 19bp, 4bp to 18bp, 4bp to 17bp, 4bp to 16bp, 4bp to 15bp, 4bp to 14bp, 4bp to 13bp, 4bp to 12bp, 4bp to 11bp, 4bp to 10bp, 4bp to 9bp, or 4bp to 8bp.

[0068] According to one embodiment, the nucleotide editing window is from the first cytosine (C1) to the 20th cytosine (C2) from the 5' end of the gRNA target base sequence. 20 ), C2~C 20 , C3~C 20 , C4~C 20 , C1~C 19 , C2~C 19 , C3~C 19 , C4~C 19 , C1~C 18 , C2~C 18 , C3~C 18 , C4~C 18 , C1~C 17 , C2~C 17 , C3~C 17 , C4~C 17 , C1~C 16 , C2~C 16 , C3~C 16, C4~C 16 , C1~C 15 , C2~C 15 , C3~C 15 , C4~C 15 , C1~C 14 , C2~C 14 , C3~C 14 , C4~C 14 , C1~C 13 , C2~C 13 , C3~C 13 , C4~C 13 , C1~C 12 , C2~C 12 , C3~C 12 , C4~C 12 , C1~C 11 , C2~C 11 , C3~C 11 , C4~C 11 , C1~C 10 , C2~C 10 , C3~C 10 , C4~C 10 , C1 to C9, C2 to C9, C3 to C9, C4 to C9, C1 to C8, C2 to C8, C3 to C8 or C4 to C8 ranges.

[0069] Another aspect provides a method of editing a nucleic acid comprising contacting a nucleic acid molecule with said CRISPR-Cas system.

[0070] The nucleic acid and CRISPR-Cas system are as described above.

[0071] In one embodiment, the editing may result in the formation of a deletion, insertion, substitution, or indel in at least one or more nucleotide sequences within the nucleotide sequence of the nucleic acid molecule.

[0072] According to one embodiment, the modification in the method for modifying a nucleic acid is a modification of 1 bp to 60 bp, for example, 1 bp to 55 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 60 bp, 5 bp to 55 bp, 5 bp to 50 bp, 5 bp to 45 bp, 5 bp to 40 bp, 5 bp to 3 5bp, 5bp~30bp, 5bp~25bp, 5bp~20bp, 5bp~15bp, 5bp~10bp, 10bp~60bp, 10bp~55bp, 10bp~50bp, 10bp~45bp, 10bp~40bp, 10bp~ 35bp, 10bp~30bp, 10bp~25bp, 10bp~20bp, 10bp~15bp, 15bp~60bp, 15bp~55bp, 15bp~50bp, 15bp~45bp, 15bp~40bp, 15bp~35bp , 15bp~30bp, 15bp~25bp, 15bp~20bp, 20bp~60bp, 20bp~55bp, 20bp~50bp, 20bp~45bp, 20bp~40bp, 20bp~35bp, 20bp~30bp, 20b p~25bp, 25bp~60bp, 25bp~55bp, 25bp~50bp, 25bp~45bp, 25bp~40bp, 25bp~35bp, 25bp~30bp, 30bp~60bp, 30bp~55bp, 30bp~50 bp, 30bp to 45bp, 30bp to 40bp, 30bp to 35bp, 35bp to 60bp, 35bp to 55bp, 35bp to 50bp, 35bp to 45bp, 35bp to 40bp, 40bp to 60bp, 40bp to 55bp, 40bp to 50bp, 40bp to 45bp, 45bp to 60bp, 45bp to 55bp, 45bp to 50bp, 50bp to 60bp, 50bp to 55bp, or 55bp to 60bp.

[0073] According to one embodiment, the modification in the method for modifying a nucleic acid is a modification of 1 bp to 60 bp, for example, 1 bp to 55 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 60 bp, 5 bp to 55 bp, 5 bp to 50 bp, 5 bp to 45 bp, 5 bp to 40 bp, 5 bp to 3 5bp, 5bp~30bp, 5bp~25bp, 5bp~20bp, 5bp~15bp, 5bp~10bp, 10bp~60bp, 10bp~55bp, 10bp~50bp, 10bp~45bp, 10bp~40bp, 10bp~ 35bp, 10bp~30bp, 10bp~25bp, 10bp~20bp, 10bp~15bp, 15bp~60bp, 15bp~55bp, 15bp~50bp, 15bp~45bp, 15bp~40bp, 15bp~35bp , 15bp~30bp, 15bp~25bp, 15bp~20bp, 20bp~60bp, 20bp~55bp, 20bp~50bp, 20bp~45bp, 20bp~40bp, 20bp~35bp, 20bp~30bp, 20b p~25bp, 25bp~60bp, 25bp~55bp, 25bp~50bp, 25bp~45bp, 25bp~40bp, 25bp~35bp, 25bp~30bp, 30bp~60bp, 30bp~55bp, 30bp~50 bp, 30bp to 45bp, 30bp to 40bp, 30bp to 35bp, 35bp to 60bp, 35bp to 55bp, 35bp to 50bp, 35bp to 45bp, 35bp to 40bp, 40bp to 60bp, 40bp to 55bp, 40bp to 50bp, 40bp to 45bp, 45bp to 60bp, 45bp to 55bp, 45bp to 50bp, 50bp to 60bp, 50bp to 55bp, or 55bp to 60bp.

[0074] According to one embodiment, the modification in the method for modifying a nucleic acid is performed by modifying a portion of the nucleotide sequence of a nucleic acid molecule that is 1 bp to 60 bp, for example, 1 bp to 55 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 60 bp, 5 bp to 55 bp, 5 bp to 50 bp, 5 bp to 45 bp, 5 bp to 40 bp, 5 bp to 35 bp, bp, 5bp~30bp, 5bp~25bp, 5bp~20bp, 5bp~15bp, 5bp~10bp, 10bp~60bp, 10bp~55bp, 10bp~50bp, 10bp~45bp, 10bp~40bp, 10bp~3 5bp, 10bp~30bp, 10bp~25bp, 10bp~20bp, 10bp~15bp, 15bp~60bp, 15bp~55bp, 15bp~50bp, 15bp~45bp, 15bp~40bp, 15bp~35bp, 15bp~30bp, 15bp~25bp, 15bp~20bp, 20bp~60bp, 20bp~55bp, 20bp~50bp, 20bp~45bp, 20bp~40bp, 20bp~35bp, 20bp~30bp, 20bp ~25bp, 25bp~60bp, 25bp~55bp, 25bp~50bp, 25bp~45bp, 25bp~40bp, 25bp~35bp, 25bp~30bp, 30bp~60bp, 30bp~55bp, 30bp~50b p, 30bp to 45bp, 30bp to 40bp, 30bp to 35bp, 35bp to 60bp, 35bp to 55bp, 35bp to 50bp, 35bp to 45bp, 35bp to 40bp, 40bp to 60bp, 40bp to 55bp, 40bp to 50bp, 40bp to 45bp, 45bp to 60bp, 45bp to 55bp, 45bp to 50bp, 50bp to 60bp, 50bp to 55bp, or 55bp to 60bp of nucleotides.

[0075] According to one embodiment, the method for modifying a nucleic acid has an efficiency of forming indels in a nucleotide sequence of 5% to 50%, for example, 5% to 45%, 5% to 40%, 5% to 35%, 5% to 30%, 5% to 25%, 5% to 20%, 5% to 15%, 5% to 10%, 10% to 50%, 10% to 45%, 10% to 40%, 10% to 35%, 10% to 30%, 10% to 25%, 10% to 20%, 10% to 15%, 15% to 50%, 15% to 45%, 15% to 40%, 15% to It can be 35%, 15%-30%, 15%-25%, 15%-20%, 20%-50%, 20%-45%, 20%-40%, 20%-35%, 20%-30%, 20%-25%, 25%-50%, 25%-45%, 25%-40%, 25%-35%, 25%-30%, 30%-50%, 30%-45%, 30%-40%, 30%-35%, 35%-50%, 35%-45%, 35%-40%, 40%-50%, 40%-45%, or 45%-50%.

[0076] According to one embodiment, the method for modifying a nucleic acid has an efficiency of forming substitutions in a nucleotide sequence of 1% to 20%, for example, 1% to 18%, 1% to 16%, 1% to 14%, 1% to 12%, 1% to 10%, 1% to 8%, 1% to 6%, 1% to 4%, 1% to 2%, 2% to 20%, 2% to 18%, 2% to 16%, 2% to 14%, 2% to 12%, 2% to 10%, 2% to 8%, 2% to 6%, 2% to 4%, 4% to 20%, 4% to 18%, 4% to 16%, 4% to 14%, 4% to 12%, 4% to 10%, 4% to 8 ... It can be % to 6%, 6% to 20%, 6% to 18%, 6% to 16%, 6% to 14%, 6% to 12%, 6% to 10%, 6% to 8%, 8% to 20%, 8% to 18%, 8% to 16%, 8% to 14%, 8% to 12%, 8% to 10%, 10% to 20%, 10% to 18%, 10% to 16%, 10% to 14%, 10% to 12%, 12% to 20%, 12% to 18%, 12% to 16%, 12% to 14%, 14% to 20%, 14% to 18%, 14% to 16%, 16% to 20%, 16% to 18%, or 18% to 20%.

[0077] The method for modifying nucleic acid can form an editing window of at least four nucleotides in the nucleotide sequence of a target nucleic acid molecule. According to one embodiment, the method of modifying a nucleic acid comprises a proofreading window of at least 50 nucleotides, such as at least 49 nucleotides, at least 48 nucleotides, at least 47 nucleotides, at least 46 nucleotides, at least 45 nucleotides, at least 44 nucleotides, at least 43 nucleotides, at least 42 nucleotides, at least 41 nucleotides, at least 40 nucleotides, at least 39 nucleotides, at least 38 nucleotides, at least 37 nucleotides, at least 36 nucleotides, at least 35 nucleotides, at least 34 nucleotides, at least 33 nucleotides, at least 32 nucleotides, at least 31 nucleotides, at least 30 nucleotides, at least 29 nucleotides, at least 28 nucleotides, at least 27 nucleotides, at least 26 nucleotides, at least 25 nucleotides, at least 24 nucleotides, at least 23 nucleotides, at least 22 nucleotides, at least 21 nucleotides, at least 20 nucleotides, at least 19 nucleotides, at least 18 nucleotides, at least 17 nucleotides, at least 16 nucleotides, at least 15 nucleotides, at least 14 nucleotides, at least 13 nucleotides, at least 12 nucleotides, at least 11 nucleotides, at least 10 nucleotides, at least 9 nucleotides, at least 8 nucleotides, at least 7 nucleotides, at least 6 nucleotides, or at least 5 nucleotides. The optical fiber may have a window.

[0078] According to one embodiment, the method for modifying a nucleic acid includes modifying a gRNA target nucleotide sequence by adding 1 bp to 20 bp, 1 bp to 19 bp, 1 bp to 18 bp, 1 bp to 17 bp, 1 bp to 16 bp, 1 bp to 15 bp, 1 bp to 14 bp, 1 bp to 13 bp, 1 bp to 12 bp, 1 bp to 11 bp, 1 bp to 10 bp, 1 bp to 9 bp, 1 bp to 8 bp, 2 bp to 20 bp, 2 bp to 19 bp, 2 bp to 18 bp, 2 bp to 17 bp, 2 bp to 16 bp, 2 bp to 15 bp, 2 bp to 14 bp, 2 bp to 13 bp, 2 bp to 12 bp, 2 bp to 11 bp, 2 bp to 10 bp, 2 bp to 8 bp, The editing window may be 9 bp, 2 bp to 8 bp, 3 bp to 20 bp, 3 bp to 19 bp, 3 bp to 18 bp, 3 bp to 17 bp, 3 bp to 16 bp, 3 bp to 15 bp, 3 bp to 14 bp, 3 bp to 13 bp, 3 bp to 12 bp, 3 bp to 11 bp, 3 bp to 10 bp, 3 bp to 9 bp, 3 bp to 8 bp, 4 bp to 20 bp, 4 bp to 19 bp, 4 bp to 18 bp, 4 bp to 17 bp, 4 bp to 16 bp, 4 bp to 15 bp, 4 bp to 14 bp, 4 bp to 13 bp, 4 bp to 12 bp, 4 bp to 11 bp, 4 bp to 10 bp, 4 bp to 9 bp, or 4 bp to 8 bp.

[0079] According to one embodiment, the editing window of the nucleic acid modification method is a window that modifies the first cytosine (C1) to the 20th cytosine (C2) from the 5' end of the gRNA target base sequence. 20 ), C2~C 20 , C3~C 20 , C4~C 20 , C1~C 19 , C2~C 19 , C3~C 19 , C4~C 19 , C1~C 18 , C2~C 18 , C3~C 18 , C4~C 18 , C1~C 17 , C2~C 17 , C3~C 17 , C4~C 17 , C1~C 16 , C2~C 16 , C3~C16 , C4~C 16 , C1~C 15 , C2~C 15 , C3~C 15 , C4~C 15 , C1~C 14 , C2~C 14 , C3~C 14 , C4~C 14 , C1~C 13 , C2~C 13 , C3~C 13 , C4~C 13 , C1~C 12 , C2~C 12 , C3~C 12 , C4~C 12 , C1~C 11 , C2~C 11 , C3~C 11 , C4~C 11 , C1~C 10 , C2~C 10 , C3~C 10 , C4~C 10 , C1 to C9, C2 to C9, C3 to C9, C4 to C9, C1 to C8, C2 to C8, C3 to C8 or C4 to C8 ranges. [Effects of the Invention]

[0080] The fusion protein, its polypeptide, and the CRISPR-Cas system containing the same according to one embodiment enable effective base proofreading. Furthermore, the polypeptide has a small size, allowing it to easily bind to a Cas protein. When used as a cell therapy agent, this facilitates vector-mediated delivery.

[0081] Furthermore, when indels are induced, the efficiency of indels and the size of the nucleotides at which indels are formed also increase, making it possible to effectively use the method for gene knockout. [Brief explanation of the drawings]

[0082] [Figure 1]A schematic diagram showing the structural and evolutionary differences between SsdA and other deaminases used in existing base proofreading technologies. [Figure 2] FIG. 1 shows a diagram illustrating the process of cleavage of single-stranded DNA by treating single-stranded DNA having the base sequence of SEQ ID NO: 4 with SsdA to deaminate it, and then treating it with UDG and NaOH; and FIG. 2 shows a photograph illustrating the results of Western blot analysis to confirm whether or not SsdA was deaminated. [Figure 3] Figure 3A shows the base proofreading efficiency of a CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide at a target site in target DNA. Figure 3B shows the base proofreading efficiency of a CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide at a target site in RNF2 DNA. Figure 3B shows the base proofreading efficiency of a CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide at a target site in HEK2 DNA. [Figure 4] 1 is a graph showing the cytotoxicity of the Cas9(D10A)-SsdA fusion protein, the Cas9(D10A)-SsdA-UGI fusion protein, the dCas9-SsdA fusion protein, and the dCas9-SsdA-UGI fusion protein. [Figure 5]Figure 5A shows the base proofreading efficiency and indel formation efficiency of CRISPR-Cas systems containing Cas9 and SsdA fusion proteins. Figure 5A shows the base proofreading efficiency and indel formation efficiency of CRISPR-Cas systems containing Cas9(D10A)-SsdA fusion protein, Cas9(D10A)-SsdA-UGI fusion protein, dCas9-SsdA fusion protein, or dCas9-SsdA-UGI fusion protein in HEK2, HEK3, HEK4, and RNF2, respectively. Figure 5B is a graph showing the base proofreading efficiency in CRISPR-Cas systems containing the cjCas9(D8A)-SsdA fusion protein, the cjCas9(D8A)-SsdA-UGI fusion protein, or the cjCas9(L58Y / D900K)(D8A)-SsdA-UGI fusion protein, as well as EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, and TFPi, respectively. Figure 5C is a graph showing the base proofreading efficiency of the CRISPR-Cas systems containing the cjCas9(D8A)-SsdA fusion protein, the cjCas9(D8A)-SsdA-UGI fusion protein, or the cjCas9(L58Y / D900K)(D8A)-SsdA-UGI fusion protein, as well as EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, and TFPi, in a cell line with suppressed uracil-DNA glycosylase (UDG) expression. [Figure 6] FIG. 1 is a schematic diagram showing plasmids for producing fusion proteins in which SsdA and a uracil glycosylase inhibitor (UGI) are linked to the C-terminus, N-terminus, and both the N- and C-terminus of a Cas protein. [Figure 7]Figure 7A shows the cytosine base proofreading efficiency depending on the binding site of SsdA and the Cas protein. Figure 7A shows the cytosine base proofreading efficiency in CRISPR-Cas systems containing the Cas9(D10A)-SsdA fusion protein (SsdA-C), the SsdA-Cas9(D10A) fusion protein (SsdA-N, SsCBE), or the SsdA-Cas9(D10A)-SsdA (SsdA-NC) fusion protein, in HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1. Figure 7B shows the cytosine base proofreading efficiency in CRISPR-Cas systems containing fusion proteins with one or two UGIs attached to the C-terminus, N-terminus, or both the C- and N-terminus of the SsdA-Cas9(D10A) fusion protein (SsdA-N, SsCBE), HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1. Figure 7C shows the accurate base proofreading efficiency for a UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) with two UGIs attached to the N-terminus, or a UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2) with two UGIs attached to both the N- and C-terminus. Figure 7D is a graph showing the cytosine base proofreading efficiency at 20 target positions (HEK2-1, HEK2-2, HEK2-3, HEK2-4, HEK3-1, HEK3-2, HEK3-3, HEK3-6, HEK3-7, HEK3-8, HEK4-1, HEK4-2, HEK4-3, HEK4-4, HEK4-5, HEK4-6, HEK4-7, HEK4-8, RNF2-3, and RNF2-4). Figure 7C is a graph showing the base proofreading range (Editing window), which is a characteristic of the base proofreading technology, for base proofreading at the 20 target positions shown in Figure 7C. [Figure 8]Graph showing the substitution frequency (%) of the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2), BE3, and BE4max at the target position (HEK2, HEK3, HEK2-2, HEK3-8, HEK4-2, or HEK4-7 site). [Figure 9] 10A-10C are graphs showing the base proofreading efficiency and indel formation efficiency of CRISPR-Cas systems containing the Cas9(D10A)-SsdA(G54D)-UGI fusion protein, the Cas9(D10A)-SsdA(G54D) fusion protein, the dCas9-SsdA(G54D)-UGI fusion protein, or the dCas9-SsdA(G54D) fusion protein in HEK2, HEK3, and HEK4, respectively. DETAILED DESCRIPTION OF THE INVENTION

[0083] The present invention will be described in more detail through the following examples, but these examples are for illustrative purposes only and the scope of the present invention is not limited to these examples.

[0084] Example 1. Plasmid cloning

[0085] The gBlock double-stranded DNA fragment (Integrated DNA Technologies) encoding His6-SsdA having the amino acid sequence of SEQ ID NO: 1 and SsdAI having the amino acid sequence of SEQ ID NO: 3 and the bacterial expression vector pET-28b(+) DNA (Novagen) were treated with XbaI and XhoI restriction enzymes (New England Biolabs) for 3 hours at 37°C. The linearized pET-28b vector was then purified using agarose gel extraction (GeneAll) and ligated to the gBlock double-stranded DNA fragment using Quick Ligase (New England Biolabs).

[0086] The coding sequences of Cas9 and UGI were then obtained by PCR amplification using pCMV plasmid DNA, and the SsdA sequence was amplified in gBlock using Gibson Assembly Master Mix (New England Biolabs) and subcloned into the pCMV plasmid.

[0087] The PAAR-domin-containing protein amino acid sequence of SEQ ID NO: 1 and the SsdAI amino acid sequence of SEQ ID NO: 3 are shown in Table 2.

[0088] [Table 2]

[0089] Example 2. Purification of SsdA

[0090] The SsdA protein cloned in Example 1 was purified using Escherichia coli BL21.

[0091] Specifically, pET-28b-His6-SsdA-SsdAI was introduced into E. coli BL21 using 0.5 mM IPTG, and the His6-SsdA and SsdAI complex was purified using Ni-NTA agarose beads (Qiagen). To separate His6-SsdA from SsdAI, the His6-SsdA and SsdAI complex was denatured in denaturing buffer (8 M urea, 50 mM Tris-HCl pH 7.5, 500 mM NaCl, and 1 mM DTT) and then cultured at 4°C for 16 hours. The suspension containing the denatured protein complex was mixed with Ni-NTA agarose beads (Qiagen) and loaded onto a gravity-flow column to remove unbound SsdAI. SsdA was then refolded using denaturing buffers containing decreasing concentrations of urea (6 M, 4 M, 2 M, 1 M, and 0 M). The refolded protein bound to the Ni-NTA agarose beads was eluted with elution buffer containing 300 mM imidazole. The eluted protein was then dialyzed using 20 mM Tris-HCl pH 7.5, 200 mM NaCl, 1 mM DTT, and 40% (w / v) glycerol and concentrated using an Amicon Ultra-15 Centrifugal Filter Unit (Millipore). The concentration of His6-SsdA protein was then analyzed by SDS-PAGE.

[0092] Example 3. Cell culture and transfection

[0093] HEK293T cells (ATCC CRL-11268) were preserved in Dulbecco's modified Eagle's medium (DMEM) containing 10% fetal bovine serum (FBS) and 1% penicillin / streptomycin (Welgene), and then plated at a density of 6 x 10 cells per well. 4Cells were seeded into TC-treated 48-well plates (Corning Life Sciences). 24 hours after seeding, at approximately 60% cell confluence, transfection was performed using 500 ng of plasmids (250 ng of Cas9-SsdA expression plasmid and 250 ng of gRNA expression plasmid) and 1.5 μL of Lipofectamine 2000 (Thermo Fisher Scientific). The transfected cells were then incubated at 37°C for 3 days. Genomic DNA was then prepared by direct cell lysis using lysis buffer (10 mM Tris-HCl, pH 7.5, 0.05% SDS, 100 mg / mL proteinase K; QIAGEN). The cell lysate was incubated at 56°C for 30 minutes and then further incubated at 99°C for 15 minutes to inactivate the proteinase K.

[0094] Example 4. Targeted deep sequencing and data analysis

[0095] The target sites were amplified by PCR (two or three times in total) and sequenced via an Illumina MiniSeq or iSeq 100 sequencing system.

[0096] More specifically, 3 mL of cell lysate or 1 mL of isolated genomic DNA was applied to the primary PCR, and 1 mL of the primary PCR product was used for the secondary PCR. The Illumina TruSeq HT Dual Index Adapter Sequence was ligated with an index PCR primer pair using 1 mL of the secondary PCR product. The size of the PCR amplicon was confirmed on a 2% agarose gel, and the amplicons were analyzed using an Illumina MiniSeq or iSeq 100 sequencing system. Targeted deep sequencing analysis was performed using MAUND (https: / / github.com / ibscge / maund), and all results were analyzed using Cas-Analyzer ( http: / / www.rgenome.net / cas-analyzer / ) was confirmed.

[0097] Example 5. Confirmation of deamination of cytosine in single-stranded DNA by SsdA

[0098] To confirm that SsdA converts cytosine bases in single-stranded DNA to uracil, we treated single-stranded DNA containing FAM with SsdA.

[0099] More specifically, single-stranded DNA containing the base sequence of SEQ ID NO: 4 (5'-Aaaaaaaaaaaaaagcgaaaaaaaaaaaaaaaaaa-3') was treated with 1-200 nM SsdA at 37°C for 1 hour, followed by treatment with uracil DNA glycosylase (UDG) at 37°C for 30 minutes to remove the uracil-converted DNA bases and create abasic sites. The DNA was then treated with 100 mM NaOH and incubated at 95°C for 2 minutes to cleave the abasic sites. The presence or absence of cleavage was confirmed by Western blot analysis, and the results are shown in Figure 2.

[0100] Figure 2 shows a diagram illustrating the process of cleaving single-stranded DNA by treating a single-stranded DNA having the base sequence of SEQ ID NO: 4 with SsdA to deaminate it, and then treating it with UDG and NaOH. It also shows a photograph showing the results of Western blot analysis to confirm whether or not SsdA was deaminated.

[0101] As shown in Figure 2, we confirmed that DNA cleavage occurred only when single-stranded DNA was treated with SsdA, UDG, and NaOH, indicating that SsdA can effectively deaminate cytosine in single-stranded DNA and convert it to uracil.

[0102] Example 6. Confirmation of base proofreading efficiency of CRISPR-Cas system containing Cas protein, SsdA, and guide polynucleotide

[0103] To confirm the base proofreading efficiency of the CRISPR-Cas system containing the Cas protein, SsdA, and guide polynucleotide, we treated target DNA (RNF2 and HEK2) with the CRISPR-Cas system containing SsdA, dCas9, and gRNA to determine whether cytosine deamination was achieved. The base sequences of the target DNA (RNF2 and HEK2) are shown in Table 3.

[0104] [Table 3]

[0105] More specifically, RNF2 DNA and HEK2 DNA were treated with 100 nM Cas9, 300 nM sgRNA, and 40 nM SsdA, respectively, and then incubated at 37°C for 8 hours to induce cytosine-to-uracil conversion at the target position in the target DNA. After 8 hours, the sgRNA, Cas9, and SsdA were removed by treatment with RNase and Protease K. The DNA was then purified using a Qiagen DNA extraction kit. PCR was then performed using primers containing the target site, and base proofreading efficiency was measured using deep sequencing. The results are shown in Figure 3.

[0106] FIG. 3 is a graph showing the base proofreading efficiency of a CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide at a target site in a target DNA.

[0107] Figure 3A is a graph showing the base proofreading efficiency of a CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide at a target site in RNF2 DNA. Figure 3B is a graph showing the base proofreading efficiency of a CRISPR-Cas system comprising a Cas protein, SsdA, and a guide polynucleotide at a target site in HEK2 DNA.

[0108] As shown in Figure 3A, when using gRNA targeting RNF2, the cytosine to uracil conversion efficiency was approximately 7% for C3 and 4% for C4. 12 In addition, as shown in Figure 3B, when gRNA targeting HEK2 was used, the C4, C6, and C 11 The conversion efficiency of cytosine to uracil was confirmed to be approximately 18-19%.

[0109] This means that the CRISPR-Cas system can only cause gene editing in one strand of a DNA duplex, and therefore, considering that the maximum base proofreading efficiency is 50%, such results indicate a fairly high level of base proofreading efficiency.

[0110] In addition, to confirm whether the base conversion generated by the CRISPR-Cas system containing the Cas protein, SsdA, and guide polynucleotide was a cytosine-to-uracil conversion, deaminated DNA was treated with Uracil-Specific Excision Reagent (USER), a uracil-specific excision enzyme, and the target region was amplified by PCR. As a result, it was confirmed that the number of reads in which cytosine was converted to uracil was eliminated.

[0111] These results indicate that the CRISPR-Cas system effectively converts cytosine to uracil.

[0112] Example 7. Confirmation of cytotoxicity of fusion proteins containing Cas protein and SsdA

[0113] The cytotoxicity of the fusion protein containing the Cas protein, SsdA, and UGI was confirmed, and the results are shown in Figure 4.

[0114] To confirm whether the SsdA gene in the CRISPR-Cas system is toxic in eukaryotic cells, we constructed a plasmid expressing the protein fused to Cas9(D10A) and inoculated 6 × 10 HEK293 cells into a 48-well plate. 4 After dispensing at a cell / well concentration, the plasmid was transfected into HEK293 cells. 48 hours after transduction, viable cells were trypsinized and the number of HEK293 cells was counted using a hemacytometer.

[0115] Figure 4 is a graph showing the cytotoxicity of the Cas9(D10A)-SsdA fusion protein, the Cas9(D10A)-SsdA-UGI fusion protein, the dCas9-SsdA fusion protein, and the dCas9-SsdA-UGI fusion protein.

[0116] As shown in Figure 4, the fusion protein containing Cas9, SsdA, and UGI was confirmed to have almost no cytotoxicity.

[0117] Example 8. Confirmation of base proofreading efficiency and indel formation efficiency of CRISPR-Cas system containing fusion protein containing Cas protein and SsdA

[0118] To confirm that the SsdA CRISPR-Cas system effectively induces cytosine deamination in eukaryotic cells, we constructed a plasmid expressing the Cas9-bound protein and transfected it into HEK293 cells. The target sequences (HEK2, HEK3, HEK4, RNF2, EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, or TFPi sites) were amplified by PCR from the transfected cells and analyzed for mutations and indel formation using next-generation sequencing (NGS). The results are shown in Figure 5.

[0119] Cas9 was used as spCas9 (Cas9(D10A)) or cjCas9(D8A). Uracil-DNA glycosylase (UDG) expression-depleted cell lines (UNG KD) were prepared by transfecting HEK293 cells with a UDG knockdown shRNA (5′-GTCTACAGACATAGAGGATTT-3: SEQ ID NO: 7) expression plasmid and an HIV-based packaging plasmid (containing the Gag / Pol, Rev, and VSV-G genes) in OptiMEM (Invitrogen) mixed with Lipofectamine 3000 reagent.

[0120] FIG. 5 is a graph showing the base proofreading efficiency and indel formation efficiency of a CRISPR-Cas system containing Cas9 and SsdA fusion proteins.

[0121] Figure 5A is a graph showing the base proofreading efficiency and indel formation efficiency of CRISPR-Cas systems containing the Cas9(D10A)-SsdA fusion protein, the Cas9(D10A)-SsdA-UGI fusion protein, the dCas9-SsdA fusion protein, or the dCas9-SsdA-UGI fusion protein, in HEK2, HEK3, HEK4, and RNF2, respectively. Figure 5b is a graph showing the base proofreading efficiency of CRISPR-Cas systems containing the cjCas9(D8A)-SsdA fusion protein, the cjCas9(D8A)-SsdA-UGI fusion protein, or the cjCas9(L58Y / D900K)(D8A)-SsdA-UGI fusion protein, in EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, and TFPi, respectively. Figure 5C is a graph showing the base proofreading efficiency of Uracil-DNA This is a graph showing the base proofreading efficiency of CRISPR-Cas systems containing the cjCas9(D8A)-SsdA fusion protein, the cjCas9(D8A)-SsdA-UGI fusion protein, or the cjCas9(L58Y / D900K)(D8A)-SsdA-UGI fusion protein, as well as EPAS1_e2, EPAS1_e5, HIF_e8, HIF_e9, and TFPi, respectively, in a glycosylase (UDG) expression-suppressed cell line.

[0122] As shown in Figure 5A, Cas9(D10A)-SsdA-UGI showed 4-11% cytosine base proofreading at all HEK2, HEK3, HEK4, and RNF2 sites, resulting in the formation of indels at the target sites. dCas9-SsdA-UGI also showed 1-3.5% cytosine base proofreading, resulting in the formation of no indels.

[0123] As shown in Figure 5B, both cjCas9(D8A)-SsdA-UGI and cjCas9(L58Y / D900K)(D8A)-SsdA-UGI showed approximately 3% base proofreading, demonstrating improved base proofreading efficiency compared to the fusion proteins without UGI.

[0124] As shown in Figure 5C, in the UDG expression-suppressed cell line, both cjCas9(D8A)-SsdA-UGI and cjCas9(L58Y / D900K)(D8A)-SsdA-UGI showed a base proofreading efficiency of approximately 15%, which was significantly improved compared to the cjCas9(D8A)-SsdA fusion protein without UGI.

[0125] Table 4 shows the results of analyzing the base-edited sequences generated by Cas9(D10A)-Ssda-UGI and HEK2-targeting gRNA.

[0126] [Table 4]

[0127] Example 9. Confirmation of base proofreading efficiency depending on the binding position of SsdA and Cas protein

[0128] Uracil-DNA glycosylase (UDG) is a well-known protein that repairs cytosine deamination in cellular DNA. To accurately confirm that SsdA actually causes cytosine deamination in cellular DNA, we first generated a UDG-knockout HEK293 cell line (HEK293 UDG-KO).

[0129] To confirm the base proofreading and indel formation efficiencies of fusion proteins in which SsdA is bound to the C-terminus, N-terminus, and both the N- and C-terminus of the Cas protein, fusion proteins were prepared with different SsdA positions, and the base proofreading and indel formation efficiencies were confirmed.

[0130] More specifically, a pCMV plasmid like the one shown in Figure 6 was prepared and transfected into HEK293 UDG-KO cells together with gRNAs for each target site according to Example 3. The nucleotide sequence of the target site (HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1 site) was then amplified by PCR from the transfected cells, and next-generation sequencing (NGS) was used to analyze whether mutations were introduced and whether indels were formed. The results are shown in Figure 7.

[0131] Figure 6 is a schematic diagram showing plasmids for producing fusion proteins in which SsdA and a uracil glycosylase inhibitor (UGI) are bound to the C-terminus, N-terminus, and both the N- and C-terminus of a Cas protein.

[0132] Figure 7 is a graph confirming the base proofreading efficiency depending on the binding position of SsdA and Cas protein.

[0133] Figure 7A shows the cytosine base proofreading efficiency in CRISPR-Cas systems containing the Cas9(D10A)-SsdA fusion protein (SsdA-C), the SsdA-Cas9(D10A) fusion protein (SsdA-N, SsCBE), or the SsdA-Cas9(D10A)-SsdA (SsdA-NC) fusion protein in HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1, respectively. 7b is a graph showing the cytosine base proofreading efficiency in CRISPR-Cas systems containing fusion proteins with one or two UGIs attached to the C-terminus, N-terminus, and both the C-terminus and N-terminus of the SsdA-Cas9(D10A) fusion protein (SsdA-N, SsCBE), HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1, respectively. FIG. 7C is a graph showing the cytosine base proofreading efficiency in CRISPR-Cas systems containing fusion proteins with one or two UGIs attached to the C-terminus, N-terminus, and both the C-terminus and N-terminus of the SsdA-Cas9(D10A) fusion protein (SsdA-N, SsCBE), HEK2, HEK3, HEK4, RNF2, FANCF, TYRO3, CCR5, or EMX1, respectively. For comparison of accurate base proofreading efficiency, we used a UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) with two UGIs attached to the N-terminus, or a UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2) with two UGIs attached to both the N-terminus and C-terminus. We then tested 20 target positions (HEK2-1, HEK2-2, HEK2-3, HEK2-4). , HEK3-1, HEK3-2, HEK3-3, HEK3-6, HEK3-7, HEK3-8, HEK4-1, HEK4-2, HEK4-3, HEK4-4, HEK4-5, HEK4-6, HEK4-7, HEK4-8, RNF2-3, RNF2-4), and Figure 7D is a graph showing an analysis of the base proofreading range (Editing window), which is a characteristic of the base proofreading technology, for base proofreading at the 20 target positions shown in Figure 7C.

[0134] The base sequences of the target sites of the target genes are shown in Table 5.

[0135] [Table 5] TIFF2025535373000007.tif12170

[0136] As shown in Figure 7A, when SsdA was attached to the C-terminus of Cas (SsdA-Cas9(D10A) (SsdA-C), the cytosine base proofreading efficiency was slightly reduced compared to when SsdA was attached to the N-terminus or both the N- and C-termini. On the other hand, when SsdA was attached to the N-terminus of Cas (SsdA-N, SsCBE), the base proofreading efficiency was increased.

[0137] As shown in Figure 7B, when UGI was further conjugated to SsdA-Cas9(D10A) (SsdA-N, SsCBE), which has high base proofreading efficiency, the base proofreading efficiency was higher than that without UGI.We confirmed that the base proofreading efficiency was significantly higher in the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) in which two UGIs were conjugated to the N-terminus, and in the UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2) in which two UGIs were conjugated to both the N- and C-termini.

[0138] As shown in Figure 7C, the base proofreading efficiency was highest when compared with the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2), which has two UGIs attached to the N-terminus, or the UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2), which has two UGIs attached to both the N-terminus and C-terminus. The results showed that the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) with two UGIs attached to the N-terminus exhibited slightly increased base proofreading efficiency.

[0139] As shown in Figure 7D, we compared the editing window at 20 target positions for the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2), which has two UGIs attached to the N-terminus, and the UGI-UGI-SsdA-Cas9(D10A)-UGI-UGI fusion protein (SsCBE-UGI-N2C2), which has two UGIs attached to both the N-terminus and C-terminus, which have the highest base editing efficiency. As a result, we confirmed that high base editing efficiency was observed between 4 bp and 8 bp from the 5' end of the gRNA target sequence.

[0140] Example 10: Confirmation of base proofreading ability using UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) in wild-type HEK293 cells

[0141] We confirmed that base proofreading efficiency was also observed in wild-type HEK293 cells using the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2), which was confirmed to have the highest base proofreading ability.

[0142] More specifically, the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) and the existing BE3 or BE4max were transfected into HEK293 cells together with gRNAs for each target site as described in Example 3. The nucleotide sequences of the target sites (HEK2, HEK3, HEK2-2, HEK3-8, HEK4-2, or HEK4-7 sites) were then amplified by PCR from the transfected cells, and next-generation sequencing (NGS) was used to analyze whether mutations were introduced and whether indels were formed. The results are shown in Figure 8.

[0143] Figure 8 is a graph showing the substitution frequency (%) of the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2), BE3, and BE4max at the target site (HEK2, HEK3, HEK2-2, HEK3-8, HEK4-2, or HEK4-7 site).

[0144] As shown in Figure 8, the UGI-UGI-SsdA-Cas9(D10A) fusion protein (SsCBE-UGI-N2) successfully demonstrated base proofreading ability at six target positions (HEK2, HEK3, HEK2-2, HEK3-8, HEK4-2, or HEK4-7 sites).

[0145] Example 11. Confirmation of cytotoxicity and base proofreading ability of the CRISPR-Cas system using SsdA mutants

[0146] To confirm the cytotoxicity and base proofreading ability of SsdA mutants, we created inactivated SsdA and confirmed its effects.

[0147] More specifically, a G54D mutation was introduced into the catalytically active site of SsdA, and the mutated SsdA protein was then bound to Cas9 to confirm its intracellular function. The results are shown in Figure 9. The G54D mutation in the catalytically active site of SsdA is a G302D mutation in the amino acid sequence of SEQ ID NO: 1 (PAAR domain-containing protein). Specifically, the G54D mutation in the catalytically active site of SsdA may have the sequence of SEQ ID NO: 17. The amino acid sequence of the G54D mutation in the catalytically active site of SsdA is shown in Table 6.

[0148] [Table 6]

[0149] Figure 9 shows the base proofreading efficiency and indel formation efficiency of CRISPR-Cas systems containing the Cas9(D10A)-SsdA(G54D)-UGI fusion protein, the Cas9(D10A)-SsdA(G54D) fusion protein, the dCas9-SsdA(G54D)-UGI fusion protein, or the dCas9-SsdA(G54D) fusion protein in HEK2, HEK3, and HEK4, respectively.

[0150] As shown in Figure 9, the G54D mutation in SsdA, when combined with Cas9(D10A), exhibited very high indel generation efficiency (approximately 18–34%), regardless of the presence or absence of UGI. This suggests that SsdA(G54D) may generate indels at the target site with higher efficiency, regardless of UGI function.

[0151] Furthermore, when SsdA(G54D) was used, not only was intracellular toxicity significantly reduced, but the size of the deleted base sequence was also significantly increased.

[0152] Table 7 shows the results of analyzing the base-edited sequences generated by SsdA(G54D)-Cas9(D10A).

[0153] As shown in Table 7, analysis of the indel sequence generated by SsdA(G54D)-Cas9(D10A) confirmed that, unlike the usual Cas9 deletion of 1 bp, the size of the deleted sequence was significantly increased with SsdA(G54D)-Cas9(D10A). These results suggest that the CRISPR-Cas system including SsdA(G54D) may increase gene knockout efficiency compared to a system using only Cas9.

[0154] [Table 7]

Claims

1. A fusion protein comprising a Cas protein (CRISPR-associated protein) and a bacterial toxin, wherein the bacterial toxin is SsdA (single-stranded DNA deaminase toxin A).

2. The fusion protein of claim 1, wherein the Cas protein is a Cas9 protein or a Cas12 protein.

3. The fusion protein of claim 1, wherein the SsdA is a cytidine deaminase.

4. The fusion protein of claim 1, wherein the SsdA is an inactivated SsdA.

5. The fusion protein of claim 1 , wherein the SsdA comprises the amino acid sequence of SEQ ID NO:

1.

6. The fusion protein of claim 1 , wherein the SsdA comprises a sequence having at least 85% or more sequence identity with the amino acid sequence of SEQ ID NO:

1.

7. The fusion protein according to claim 4, wherein the inactivated SsdA has an amino acid mutation in the catalytic active site of SsdA.

8. The fusion protein of claim 4, wherein the inactivated SsdA has G302D, E349A, or their corresponding amino acid mutations in the amino acid sequence of SEQ ID NO:

1.

9. The fusion protein of claim 4, wherein the inactivated SsdA has the amino acid sequence of SEQ ID NO:

17.

10. The fusion protein of claim 4 , wherein the inactivated SsdA comprises a sequence having at least 85% or more sequence identity with the amino acid sequence of SEQ ID NO:

17.

11. 2. The fusion protein of claim 1, wherein the bacterial toxin is attached to the C-terminus, the N-terminus, or both the C-terminus and the N-terminus of the Cas protein.

12. The fusion protein of claim 1 , further comprising a DNA glycosylase inhibitor.

13. 13. The fusion protein of claim 12, wherein the DNA glycosylase inhibitor is a thymine glycosylase inhibitor, a uracil glycosylase inhibitor, an oxoguanine glycosylase inhibitor, or an alkylguanine glycosylase inhibitor.

14. The fusion protein of claim 1, wherein the fusion protein has an editing window ranging from 1 bp to 20 bp from the 5' end of the target sequence.

15. A polynucleotide encoding the fusion protein of any one of claims 1 to 14.

16. A vector comprising the polynucleotide of claim 15.

17. A CRISPR-Cas system comprising a fusion protein comprising a Cas protein (CRISPR-associated protein) and a bacterial toxin, or a polynucleotide encoding the fusion protein, and a guide polynucleotide, wherein the bacterial toxin is SsdA (single-stranded DNA deaminase toxin A).

18. The CRISPR-Cas system of claim 17, wherein the guide polynucleotide comprises a CRISPR-RNA (crRNA) and a trans-activating RNA (tracrRNA), and the guide polynucleotide is a dual guide RNA (dual guide RNA) or a single-chain guide RNA (sgRNA).

19. The CRISPR-Cas system of claim 17, wherein the system forms a deletion, insertion, substitution, or indel (insertion and deletion; indel) of at least one nucleotide in a nucleotide sequence of a target nucleic acid molecule.

20. 20. The CRISPR-Cas system of claim 19, wherein the system creates a deletion of 1 bp to 60 bp of nucleotides, an insertion of 1 bp to 60 bp of nucleotides, or an indel of 1 bp to 60 bp of nucleotides in a nucleotide sequence of a target nucleic acid molecule.

21. 18. The CRISPR-Cas system of claim 17, wherein the system has an editing window ranging from 1 bp to 20 bp from the 5' end of the target sequence.

22. 1. A method of editing a nucleic acid comprising contacting a nucleic acid molecule with a CRISPR-Cas system, The editing is to form a deletion, insertion, substitution, or indel (insertion and deletion; indel) of at least one nucleotide sequence among the nucleotide sequences of the nucleic acid molecule, The method for editing nucleic acid, wherein the CRISPR-Cas system comprises a fusion protein comprising a Cas protein (CRISPR-associated protein) and a bacterial toxin, or a polynucleotide encoding the fusion protein, and a guide polynucleotide, and the bacterial toxin is SsdA (single-stranded DNA deaminase toxin A).

23. The method for editing nucleic acids according to claim 22, wherein the edit is a deletion of 1 bp to 60 bp of nucleotides, an insertion of 1 bp to 60 bp of nucleotides, or an indel of 1 bp to 60 bp of nucleotides in the nucleotide sequence of the nucleic acid molecule.

Citation Information

Patent Citations

  • Nucleobase editors and uses thereof

    JP2019509012A

  • Method for screening target-specific nucleases using on-target and off-target multi-targeting systems and uses thereof

    JP2019517802A

  • Methods of increasing biotic stress resistance in plants

    WO2021048272A1

  • Bacterial DNA cytosine deaminases for mapping DNA methylation sites

    WO2022212584A1