Gene editing system and base editing system comprising engineered deinococcus radiodurans-derived tnpb protein, and use thereof

WO2026182507A1PCT designated stage Publication Date: 2026-09-03NSAGE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/003053
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-09
Filing Date
2026-02-24
Publication Date
2026-09-03

Smart Images

  • Figure KR2026003053_03092026_PF_FP_ABST
    Figure KR2026003053_03092026_PF_FP_ABST
Patent Text Reader

Abstract

The present specification provides: 1) an engineered TnpB protein; 2) a reRNA comprising an engineered scaffold; and 3) an RNA-induced nucleic acid cleavage system comprising 1) and 2). Furthermore, the present specification provides a method for cleaving a target nucleic acid and editing a cell genome using the RNA-induced nucleic acid cleavage system. In addition, the present specification provides: 1) an engineered dTnpB protein; 2) a base editing protein comprising the engineered dTnpB protein; and 3) a base editing system comprising 2) and a reRNA. Furthermore, the present specification provides a method for changing at least one base of a target nucleic acid into a different base by using the base editing system.
Need to check novelty before this filing date? Find Prior Art

Description

Gene editing system and base editing system containing engineered Deinococcus radiodurans-derived TnpB protein, and their uses

[0001] The invention disclosed in this specification relates to a technique for using TnpB proteins and their associated genes found in the IS200 / IS605 family of transposon systems for target-specific nucleic acid cleavage.

[0002] The TnpB gene is a gene found in the IS200 / IS605 family, a transposon system. Previous researchers studied the function of TnpB in ISDra2, a member of the IS200 / IS605 family of Deinococcus radiodurans, and revealed that it possesses target-specific nucleic acid cleavage activity similar to the CRISPR / Cas system. The TnpB protein forms a complex by binding to RNA containing a specific sequence, referred to as reRNA. This complex: recognizes a transposon-associated motif (TAM) contained in a double-stranded nucleic acid (target nucleic acid); is guided by the reRNA to a region containing the specific sequence near the TAM; binds to the corresponding region of the target nucleic acid; and causes double-stranded nucleic acid cleavage in the target nucleic acid. If the TnpB protein is correlated with the Cas protein and the reRNA with the guide RNA, the complex formed by the binding of the TnpB protein and reRNA can be understood as similar to a CRISPR / Cas complex. Therefore, the technologies introduced in the process of utilizing CRISPR / Cas complexes can be applied to the above TnpB complex almost as they are.

[0003] This specification sets forth the technical objective of enhancing the function of an RNA-induced nucleic acid cleavage complex comprising Deinococcus radiodurans-derived TnpB (ISDra2 TnpB) protein and right-end element RNA (reRNA). Furthermore, the technical objective is to develop an RNA-induced nucleic acid cleavage complex that can be utilized for actual cell genome editing.

[0004] This specification sets forth the technical objective of developing a base editor complex comprising a Deinococcus radiodurans-derived TnpB (ISDra2 TnpB) protein or a variant thereof. Specifically, the technical objective is to develop a base editor complex having an actual cellular genome base editing effect.

[0005] The inventors of this application identified a variant capable of enhancing the protein-nucleic acid interaction of the wild-type ISDra2 TnpB protein and demonstrated that the function of the RNA-induced nucleic acid cleavage complex is improved by the actual enhancement of the protein-nucleic acid interaction. Furthermore, the inventors identified a variant that stabilizes the scaffold of reRNA functioning in conjunction with the wild-type ISDra2 TnpB protein and demonstrated that the corresponding modification enhances the function of the RNA-induced nucleic acid cleavage complex.

[0006] In summary, this specification provides, as a solution to the above technical problem, 1) an engineered TnpB protein, 2) reRNA comprising an engineered scaffold, and 3) an RNA-induced nucleic acid cleavage system comprising 1) and 2). Furthermore, this specification provides a method for cleaving a target nucleic acid and further editing a cell genome using the RNA-induced nucleic acid cleavage system.

[0007] The inventors of this application prepared an engineered dTnpB protein by combining a mutation that removes nucleic acid cleavage activity and a mutation that enhances protein-nucleic acid interactions, as disclosed in the prior art, and demonstrated that it functions as intended. The inventors prepared a base editor protein by fusing a base editing domain to the TnpB protein from which nucleic acid cleavage activity was removed. Furthermore, they prepared a base editing complex comprising the base editor protein and reRNA. Moreover, they demonstrated that the base editing complex can edit a base at a specific position of a target nucleic acid into a different base, and specifically identified which base at which position is edited into which base depending on the composition of the base editing complex.

[0008] In summary, this specification provides, as a solution to the above technical problem, 1) an engineered dTnpB protein, 2) a base editor protein comprising the engineered dTnpB protein, and 3) a base editing system comprising 2) and reRNA. Furthermore, this specification provides a method for changing one or more bases of a target nucleic acid to another base using the base editing system.

[0009] By using the engineered TnpB protein disclosed in this specification, the base editor protein containing the engineered TnpB protein, and reRNA, intended gene editing can be induced with high efficiency by targeting the genome of a cell. In addition, since it exhibits gene editing efficiency comparable to the Streptococcus pyogenes-derived Cas9 system currently widely used for gene editing, it can be utilized in various fields of gene editing technology.

[0010] Figure 1 shows the results of structural analysis and visualization of a protein-RNA complex containing wild-type TnpB protein bound to a target nucleic acid by observing it with a cryo-electron microscope (cryo-EM).

[0011] Figure 2 is a graph showing the frequency of indel introduction into target nucleic acids by TnpB complexes containing each variant, after introducing a single amino acid variant into wild-type TnpB protein according to Experimental Example 2.1. Each label below the graph corresponds to a label in Table 1, and the height of the graph represents the frequency (%) of indel introduction. Below each label, it is indicated which domain of the TnpB protein the corresponding variant is located in.

[0012] Figure 3 is a graph showing the frequency of indel introduction into specific target nucleic acids by TnpB complexes containing each variant, in which the top five single amino acid variants were introduced into wild-type TnpB protein in all possible combinations according to Experimental Example 2.2. The graph title is a label indicating the target nucleic acid targeted by the reRNA used in each experiment. Each label on the vertical axis of the graph corresponds to a label in Table 3, and the horizontal axis represents the frequency of indel introduction (%). The vertical line shown in the graph indicates the frequency (%) of indel introduction by the complex containing wild-type TnpB.

[0013] Figure 4 is a graph showing the frequency of indel introduction into specific target nucleic acids by TnpB complexes containing each variant, in which the top five single amino acid variants were introduced into wild-type TnpB protein in all possible combinations according to Experimental Example 2.2. The graph title is a label indicating the target nucleic acid targeted by the reRNA used in each experiment. Each label on the vertical axis of the graph corresponds to a label in Table 3, and the horizontal axis represents the frequency (%) of indel introduction. The vertical line shown in the graph indicates the frequency (%) of indel introduction by the complex containing wild-type TnpB.

[0014] Figure 5 is a box-and-whisker plot summarizing frequency data of indel introduction into target nucleic acids by TnpB complexes containing each variant, in which the top 5 single amino acid variants were introduced into wild-type TnpB protein in all possible combinations according to Experimental Example 2.2. The horizontal axis of the graph represents each variant, and the vertical axis represents the relative fold change rate of indel introduction frequency compared to the complex containing wild-type TnpB protein. Each row below the horizontal axis represents the top 5 amino acid variants (Q64R, S72R, Q128R, N255R, K281R), uncolored circles indicate that the corresponding variant was not introduced, and colored circles indicate that the corresponding variant was introduced. For example, the first graph represents data for wild-type TnpB with no mutations introduced, and the graph on the far right represents data for TnpB variants with all of the top 5 amino acid mutations introduced.

[0015] Figure 6 is a graph showing the frequency (%) of indel introduction into the target nucleic acid by the TnpB complex containing each variant, created by introducing a combination of 5 additional variants (S57R, A237R, T350R, V254R, K310R) to the TnpB variant (enTnpB) in which all of the top 5 single amino acid variants were introduced according to Experimental Example 2.3. The graph title is the label of the target nucleic acid targeted by the reRNA used in the experiment (Site 9). The horizontal axis of the graph represents each variant, and the vertical axis represents the frequency of indel introduction. Each row below the horizontal axis represents the wild-type TnpB (TnpB), the TnpB variant with the top 5 amino acid variants introduced (enTnpB), and the additionally introduced variants (S57R, A237R, T350R, V254R, K310R). Uncolored circles indicate that the variant in the corresponding row was not introduced, and colored circles indicate that the variant in the corresponding row was introduced. For example, the rightmost label represents data for the TnpB variant in which all 5 additional variants (S57R, A237R, T350R, V254R, K310R) were introduced into enTnpB.

[0016] Figure 7 is a graph showing the frequency (%) of indel introduction into the target nucleic acid by the TnpB complex containing each variant, created by introducing a combination of five additional variants (S57R, A237R, T350R, V254R, K310R) to the TnpB variant (enTnpB) in which all of the top five single amino acid variants were introduced according to Experimental Example 2.3. The graph title is the label of the target nucleic acid targeted by the reRNA used in the experiment (EMX1-1). The horizontal axis of the graph represents each variant, and the vertical axis represents the frequency of indel introduction. Each row below the horizontal axis represents the wild-type TnpB (TnpB), the TnpB variant (enTnpB) in which the top five amino acid variants were introduced, and the additionally introduced variants (S57R, A237R, T350R, V254R, K310R). Uncolored circles indicate that the variant in the corresponding row was not introduced, and colored circles indicate that the variant in the corresponding row was introduced. For example, the rightmost label represents data for the TnpB variant in which all 5 additional variants (S57R, A237R, T350R, V254R, K310R) were introduced into enTnpB.

[0017] Figure 8 is a graph showing the frequency (%) of indel introduction into the target nucleic acid by the TnpB complex containing each variant, created by introducing a combination of 5 additional variants (S57R, A237R, T350R, V254R, K310R) to the TnpB variant (enTnpB) in which all of the top 5 single amino acid variants were introduced according to Experimental Example 2.3. The graph title is the label of the target nucleic acid targeted by the reRNA used in the experiment (AGBL1). The horizontal axis of the graph represents each variant, and the vertical axis represents the frequency of indel introduction. Each row below the horizontal axis represents the wild-type TnpB (TnpB), the TnpB variant (enTnpB) in which the top 5 amino acid variants were introduced, and the additionally introduced variants (S57R, A237R, T350R, V254R, K310R). Uncolored circles indicate that the variant in the corresponding row was not introduced, and colored circles indicate that the variant in the corresponding row was introduced. For example, the rightmost label represents data for the TnpB variant in which all 5 additional variants (S57R, A237R, T350R, V254R, K310R) were introduced into enTnpB.

[0018] Figure 9 is a graph showing the frequency (%) of indel introduction into the target nucleic acid by the TnpB complex containing each variant, created by introducing a combination of five additional variants (S57R, A237R, T350R, V254R, K310R) into the TnpB variant (enTnpB) in which all of the top five single amino acid variants were introduced according to Experimental Example 2.3. The graph title is the label of the target nucleic acid targeted by the reRNA used in the experiment (HEK1-1). The horizontal axis of the graph represents each variant, and the vertical axis represents the frequency of indel introduction. Each row below the horizontal axis represents the wild-type TnpB (TnpB), the TnpB variant with the top five amino acid variants introduced (enTnpB), and the additionally introduced variants (S57R, A237R, T350R, V254R, K310R). Uncolored circles indicate that the variant in the corresponding row was not introduced, and colored circles indicate that the variant in the corresponding row was introduced. For example, the rightmost label represents data for the TnpB variant in which all 5 additional variants (S57R, A237R, T350R, V254R, K310R) were introduced into enTnpB.

[0019] Figure 10 is a graph showing the frequency (%) of indel introduction into the target nucleic acid by the TnpB complex containing each variant, created by introducing a combination of five additional variants (S57R, A237R, T350R, V254R, K310R) to the TnpB variant (enTnpB) in which all of the top five single amino acid variants were introduced according to Experimental Example 2.3. The graph title is the label of the target nucleic acid targeted by the reRNA used in the experiment (EMX1-2). The horizontal axis of the graph represents each variant, and the vertical axis represents the frequency of indel introduction. Each row below the horizontal axis represents the wild-type TnpB (TnpB), the TnpB variant (enTnpB) in which the top five amino acid variants were introduced, and the additionally introduced variants (S57R, A237R, T350R, V254R, K310R). Uncolored circles indicate that the variant in the corresponding row was not introduced, and colored circles indicate that the variant in the corresponding row was introduced. For example, the rightmost label represents data for the TnpB variant in which all 5 additional variants (S57R, A237R, T350R, V254R, K310R) were introduced into enTnpB.

[0020] Figure 11 is a box-and-whisker plot summarizing frequency data of indel introduction into target nucleic acids by TnpB complexes containing each variant, created by introducing a combination of 5 additional variants (S57R, A237R, T350R, V254R, K310R) into a TnpB variant (enTnpB) in which all of the top 5 single amino acid variants were introduced according to Experimental Example 2.3. The horizontal axis of the graph represents each variant, and the vertical axis represents the frequency (%) of indel introduction. Each row below the horizontal axis represents the wild-type TnpB (TnpB), the TnpB variant (enTnpB) in which the top 5 amino acid variants were introduced, and the additionally introduced variants (S57R, A237R, T350R, V254R, K310R). Uncolored circles indicate that the variant in the corresponding row was not introduced, and colored circles indicate that the variant in the corresponding row was introduced. For example, the rightmost label represents data for the TnpB variant in which all 5 additional variants (S57R, A237R, T350R, V254R, K310R) were introduced into enTnpB.

[0021] Figure 12 is a graph comparing the gene editing efficiency of a TnpB complex containing a TnpB variant (enTnpB) in which all top 5 single amino acid variants were introduced according to Experimental Example 2.4 and a TnpB complex containing a wild-type TnpB variant. Each label on the horizontal axis represents the target gene targeted by reRNA, and the vertical axis represents the frequency (%) of indel introduction. For each label, the leftmost bar represents the negative control (UNT), the middle bar represents the data for a TnpB complex (TnpB) containing wild-type TnpB, and the rightmost bar represents the data for a TnpB complex (enTnpB) containing an engineered TnpB variant, respectively.

[0022] Figure 13 is a graph comparing the gene editing efficiency of TnpB complexes containing TnpB variants (engineered TnpB) or wild-type TnpB, respectively, into which all top five single amino acid variants were introduced according to Experimental Example 2.5, for four target genes (HEK1-1, EMX1-2, Site9, RUNX1-1) in K562 cells. The labels at the bottom of each graph represent the target genes targeted by reRNA, and the vertical axis represents the frequency (%) of indel introduction. For each graph, the leftmost bar represents data for the negative control (UNT), the middle bar represents data for the TnpB complex containing wild-type TnpB (TnpB), and the rightmost bar represents data for the TnpB complex containing the engineered TnpB variant (enTnpB). The graph on the far right is a violin chart summarizing the frequency data of indel introduction by complex type. The labels on the horizontal axis represent wild-type TnpB (TnpB) and engineered TnpB (enTnpB), respectively, and the vertical axis represents the relative multiple change when the average indel introduction frequency of wild-type TnpB is set to 1.

[0023] Figure 14 is a graph comparing the gene editing efficiency of four target genes (HEK1-1, EMX1-2, Site9, RUNX1-1) in iPSCs containing TnpB complexes containing TnpB variants (engineered TnpB) or wild-type TnpB, respectively, in which all top five single amino acid variants were introduced according to Experimental Example 2.5. The labels at the bottom of each graph represent the target genes targeted by reRNA, and the vertical axis represents the frequency (%) of indel introduction. For each graph, the leftmost bar represents data for the negative control (UNT), the middle bar represents data for the TnpB complex containing wild-type TnpB (TnpB), and the rightmost bar represents data for the TnpB complex containing the engineered TnpB variant (enTnpB). The graph on the far right is a violin chart summarizing the frequency data of indel introduction by complex type. The labels on the horizontal axis represent wild-type TnpB (TnpB) and engineered TnpB (enTnpB), respectively, and the vertical axis represents the relative multiple change when the average indel introduction frequency of wild-type TnpB is set to 1.

[0024] Figure 15 shows a graph comparing the gene editing efficiency of the TnpB complex (enTnpB) and the SpCas9 complex (SpCas9) engineered according to Experimental Example 2.6. The labels on the horizontal axis represent the target genes, and the vertical axis represents the frequency (%) of indel introduction. For each target gene, starting from the left, the solid rectangular bar represents the negative control (UNT), the dotted rectangular bar represents the SpCas9 complex (SpCas9), and the bar without a border represents the experimental results for the engineered TnpB complex (enTnpB), respectively.

[0025] Figure 16 shows a violin chart comparing the gene editing efficiency of the engineered TnpB complex (enTnpB) and the SpCas9 complex (SpCas9) according to Experimental Example 2.7. The horizontal axis represents each complex, and the vertical axis represents the frequency of indel introduction (%). The average frequency of indel introduction (values ​​indicated by the dotted line in the violin chart) was statistically significantly higher for the engineered TnpB.

[0026] Figure 17 shows the target nucleic acid sequences of the engineered TnpB complex (enTnpB) and SpCas9 complex (SpCas9) used in Experimental Example 2.7. The dotted box represents the spacer portion of enTnpB, the dashed box represents the spacer portion of SpCas9, and the solid box represents the spacer portion shared between the two complexes.

[0027] Figure 18 is a graph showing the frequency of indel introduction in a complex containing a TnpB variant with a portion of CTD removed according to Experimental Example 3. Each label on the horizontal axis represents a TnpB variant of a different length with a portion of the C-terminal sequence removed, and the vertical axis represents the frequency of indel introduction (%). The region marking immediately below the horizontal axis indicates which region (RuvC or CTD) the removed sequence belongs to. Specifically, UNT represents the negative control, TnpB-WT-408aa represents the wild-type TnpB protein, and the label in the form of TnpB-ΔCTD-(number) represents a TnpB protein variant with the C-terminal removed, where (number) represents the total length of the corresponding variant. For example, TnpB-ΔCTD-375 represents a TnpB variant with a continuous sequence removed from the C-terminal amino acid to achieve a total length of 375aa. A statistically significant difference in gene editing efficiency is observed between TnpB-WT-408aa and TnpB-ΔCTD-372 (or variants shorter than this) (****), but no statistically significant difference in editing efficiency is observed between TnpB-WT-408aa and TnpB-ΔCTD-375 (or variants longer than this) (ns).

[0028] Figure 19 is a graph showing the frequency of indel introduction in a complex containing a TnpB variant with a portion of CTD removed according to Experimental Example 3. Each label on the horizontal axis represents a TnpB variant of a different length with a portion of the C-terminal sequence removed, and the vertical axis represents the frequency of indel introduction (%). Specifically, UNT represents the negative control, TnpB-WT represents the wild-type TnpB protein, and labels in the form of TnpB-ΔCTD-(number)aa represent TnpB protein variants with the C-terminal portion removed, where (number) represents the total length of the corresponding variant. For example, TnpB-ΔCTD-373 represents a TnpB variant with a continuous sequence removed from the C-terminal amino acid to achieve a total length of 373aa. A statistically significant difference in gene editing efficiency is observed between TnpB-WT and TnpB-ΔCTD-372 (****), but no statistically significant difference in editing efficiency is observed between TnpB-WT and TnpB-ΔCTD-373aa (or variants longer than this) (ns).

[0029] Figure 20 shows a visualization figure obtained by analyzing, through cryo-EM structural analysis, where the mutation site of the engineered TnpB variant is located during the TnpB complex-target nucleic acid interaction according to Experimental Example 4.1.

[0030] Figure 21 is a schematic diagram of the structures of target nucleic acids and reRNA. Here, the amino acid positions (Q128, S72, Q64, S57, K281, N255) shown on the top of some nucleic acids represent the positions of the TnpB protein expected to interact with the corresponding nucleic acids, which were analyzed by cryo-EM structure according to Experimental Example 4.1.

[0031] Figure 22 shows the experimental results confirming the presence of additional TAMs recognized by the TnpB protein engineered according to Experimental Example 4. The figure at the top is a schematic representation of which bases of the TAM were expanded and identified through structure-based TAM scanning. The figure at the bottom is a graph showing the target-specific gene editing activity for each TAM in the form of a heatmap. Each row of the heatmap represents the TnpB variant used, where UNT is the negative control, TnpB is the wild-type TnpB, enTnpB is the engineered TnpB of SEQ ID NO. 2, and enTnpB+S57R is the engineered TnpB of SEQ ID NO. 3. Each column of the heatmap represents the TAM used as a target, and the number following TAM indicates that a different type of target nucleic acid was used. The intensity of the color in each cell of the heatmap indicates the frequency (%) of indel introduction, and the intensity scale is shown on the right bar.

[0032] Figure 23 shows the experimental results confirming the existence of additional TAMs recognized by the TnpB protein engineered according to Experimental Example 4. The figure at the top is a schematic diagram illustrating which bases of the TAM were expanded and identified through combinations of functionally acceptable TAM features. The figure at the bottom is a graph showing the target-specific gene editing activity for each TAM in the form of a heatmap. Each row of the heatmap represents the TnpB variant used, where UNT represents the negative control, TnpB represents the wild-type TnpB, and enTnpB represents the engineered TnpB of SEQ ID NO. 2. Each column of the heatmap represents the TAM used as a target, and the number following TAM indicates that a different type of target nucleic acid was used. The intensity of each cell in the heatmap indicates the frequency (%) of indel introduction, and the intensity scale is shown on the right bar.

[0033] Figure 24 shows the TAM sequence logo formed by synthesizing the data obtained in Experimental Example 4.

[0034] Figure 25 is a graph in the form of a heatmap showing whether the TnpB protein engineered according to Experimental Example 4 recognizes the TAM sequence of 5'-TNGAT-3' or 5'-TTGVT-3'. Each row of the heatmap represents the TnpB variant used, where UNT represents the negative control, TnpB represents the wild-type TnpB, and enTnpB represents the engineered TnpB of SEQ ID NO. 2. Each column of the heatmap represents the target sequence used as a target. The numbers represent the type of target, and the letters grouped from 1 to 9 displayed at the bottom represent the TAM used. Each column of the heatmap represents the TAM used as a target, and the number following the TAM indicates that different types of target nucleic acids were used. The intensity of the color in each cell of the heatmap indicates the frequency (%) of indel introduction, and the intensity scale is shown on the right bar.

[0035] Figure 26 schematically shows a base editing system constructed according to Experimental Example 5. The upper figure schematically shows the cytosine base editing system (TnpB-CBE) and the adenine base editing system (TnpB-ABE) binding to and functioning on a target nucleic acid, respectively. The lower figure schematically shows the structure of a vector for expressing each base editing system.

[0036] Figure 27 is a schematic diagram showing the relationship between the target nucleic acid, the correction window, the guide domain of reRNA, and the TAM. Here, m and l represent the number of nucleotides included in the corresponding region.

[0037] Figure 28 shows a heatmap graph confirming the gene editing efficiency of the cytosine base editing system according to Experimental Example 5.1. The target genes targeted by reRNA are plotted at the top of each heatmap (site2, site4, site12, EMX1-2). Each row of the heatmap represents the cytosine base editing system, where UNT represents the negative control, TnpB-CBE represents the base editing system containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-CBE represents the base editing system containing the dead-enTnpB protein of SEQ ID NO. 9. Each column of the heatmap represents the TAM sequence (indicated by a shaded box) and protospacer sequence of the target nucleic acid. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0038] Figure 29 shows a graph confirming the gene editing efficiency of the cytosine base editing system according to Experimental Example 5.1. The target genes targeted by reRNA are plotted at the top of each graph (site2, site4, site12, EMX1-2). The horizontal axis of each graph indicates the position of cytidine that undergoes nucleobase correction at the corresponding target. For example, in the case of C5, it refers to the cytidine at position 5 of the protospacer. The vertical axis of each graph indicates the frequency (%) of the corresponding cytidine being changed to thymine. For each cytidine position, the leftmost bar represents data for the negative control (UNT), the middle bar represents data for the base editing system containing the dead-TnpB protein of SEQ ID NO. 8 (TnpB-CBE), and the rightmost bar represents data for the base editing system containing the dead-enTnpB of SEQ ID NO. 9 (enTnpB-CBE).

[0039] Figure 30 shows a heatmap graph confirming the gene editing efficiency of the cytosine base editing system according to Experimental Example 5.1. The target genes targeted by reRNA are plotted at the top of each heatmap (site5, HEK1-1, RUNX1-1, AGBL1, RUNX1-2, TTR). Each row of the heatmap represents the cytosine base editing system, UNT represents the negative control, TnpB-CBE represents the base editing system containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-CBE represents the base editing system containing the dead-enTnpB of SEQ ID NO. 9. Each column of the heatmap represents the TAM sequence (indicated by a shaded box) and protospacer sequence of the target nucleic acid. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0040] Figure 31 shows a graph confirming the gene editing efficiency of the cytosine base editing system according to Experimental Example 5.1. The target genes targeted by reRNA are plotted at the top of each graph (site5, HEK1-1, RUNX1-1, AGBL1, RUNX1-2, TTR). The horizontal axis of each graph indicates the position of cytidine that is nuclear base-corrected at the corresponding target. For example, in the case of C6, it refers to the cytidine at position 6 of the protospacer. The vertical axis of each graph indicates the frequency (%) of the corresponding cytidine being changed to thymine. For each cytidine position, the leftmost bar represents data for the negative control (UNT), the middle bar represents data for the base editing system containing the dead-TnpB protein of SEQ ID NO. 8 (TnpB-CBE), and the rightmost bar represents data for the base editing system containing the dead-enTnpB of SEQ ID NO. 9 (enTnpB-CBE).

[0041] Figure 32 shows a heatmap graph confirming the gene editing efficiency of the adenine base editing system according to Experimental Example 5.1. The target genes targeted by reRNA are plotted at the top of each heatmap (site9, TTR, AGBL1, EMX1-2). Each row of the heatmap represents the adenine base editing system, UNT represents the negative control, TnpB-ABE represents the base editing system containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-ABE represents the base editing system containing the dead-enTnpB protein of SEQ ID NO. 9. Each column of the heatmap represents the TAM sequence (indicated by the shaded box) and protospacer sequence of the target nucleic acid. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0042] Figure 33 shows a graph confirming the gene editing efficiency of the adenine base editing system according to Experimental Example 5.1. The target genes targeted by reRNA are plotted at the top of each graph (site9, TTR, AGBL1, EMX1-2). The horizontal axis of each graph indicates the position of the adenosine that is nuclear base-corrected at the corresponding target. For example, in the case of A4, it refers to the adenosine at position 4 of the protospacer. The vertical axis of each graph indicates the frequency (%) of the corresponding adenosine being changed to guanosine. For each cytidine position, the leftmost bar represents data for the negative control (UNT), the middle bar represents data for the base editing system containing the dead-TnpB protein of SEQ ID NO. 8 (TnpB-ABE), and the rightmost bar represents data for the base editing system containing the dead-enTnpB of SEQ ID NO. 9 (enTnpB-ABE).

[0043] Figure 34 shows a heatmap graph confirming the gene editing efficiency of the adenine base editing system according to Experimental Example 5.1. The target genes targeted by reRNA are plotted at the top of each heatmap (RUNX1-3, site4, site12, site3, site5, site7). Each row of the heatmap represents the adenine base editing system, UNT represents the negative control, TnpB-ABE represents the base editing system containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-ABE represents the base editing system containing the dead-enTnpB protein of SEQ ID NO. 9. Each column of the heatmap represents the TAM sequence (indicated by the shaded box) and protospacer sequence of the target nucleic acid. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0044] Figure 35 shows a graph confirming the gene editing efficiency of the adenine base editing system according to Experimental Example 5.1. The target genes targeted by reRNA are plotted at the top of each graph (RUNX1-3, site4, site12, site3, site5, site7). The horizontal axis of each graph indicates the position of the adenosine that is nuclear base-corrected at the corresponding target. For example, in the case of A6, it refers to the adenosine at position 6 of the protospacer. The vertical axis of each graph indicates the frequency (%) of the corresponding adenosine being changed to guanosine. For each cytidine position, the leftmost bar represents data for the negative control (UNT), the middle bar represents data for the base editing system containing the dead-TnpB protein of SEQ ID NO. 8 (TnpB-ABE), and the rightmost bar represents data for the base editing system containing the dead-enTnpB of SEQ ID NO. 9 (enTnpB-ABE).

[0045] FIG. 36 is a graph showing the nucleobase correction efficiency of a base editing system containing dead-enTnpB of SEQ ID NO. 9 according to Experimental Example 5.1, organized by position. The upper graph represents the cytosine base editing system, and the lower graph represents the adenine base editing system. In each graph, the horizontal axis represents cytidine or adenosine at each position included in the protospacer of the target nucleic acid, and the vertical axis represents the respective nucleobase correction efficiency.

[0046] Figure 37 shows the results of gene editing targeting an extended TAM sequence using a cytosine base editing system according to Experimental Example 5.2 as a heatmap graph. Each row of the heatmap represents the cytosine base editing system, UNT represents the negative control, TnpB-CBE represents the base editing system containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-CBE represents the base editing system containing the dead-enTnpB protein of SEQ ID NO. 9. Each column of the heatmap indicates the TAM sequence of the target nucleic acid, the target gene, and the cytidine location where nucleobase editing occurs. For example, TAGAT-site1-C3 is a label for the frequency of editing of the C at position 3 of the protospacer by targeting the site1 target containing the TAM sequence of 5'-TAGAT-3'. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0047] Figure 38 shows the results of gene editing targeting an extended TAM sequence using an adenine base editing system according to Experimental Example 5.2 as a heatmap graph. Each row of the heatmap represents an adenine base editing system, where UNT represents a negative control, TnpB-ABE represents a base editing system containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-ABE represents a base editing system containing the dead-enTnpB protein of SEQ ID NO. 9. Each column of the heatmap indicates the TAM sequence of the target nucleic acid, the target gene, and the adenosine location where nucleobase editing occurs. For example, TAGAT-site1-A4 is a label for the frequency of editing of A at position 4 of the protospacer by targeting the site1 target containing the TAM sequence of 5'-TAGAT-3'. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0048] Figure 39 shows the results of editing each gene in K562 cells using a base editing system according to Experimental Example 5.3 as a heatmap graph. The top three heatmaps represent the results using an adenine base editing system, and the bottom three heatmaps represent the results using a cytosine base editing system. The targeted genes are listed at the top of each heatmap (HEK1-1, EMX1-2, site9). Each row of each heatmap represents a base editing system, where UNT represents the negative control, TnpB-ABE and TnpB-CBE represent base editing systems containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-ABE and enTnpB-CBE represent base editing systems containing the dead-enTnpB protein of SEQ ID NO. 9. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0049] Figure 40 shows the results of editing each gene of an iPSC using a base editing system according to Experimental Example 5.3 as a heatmap graph. The top three heatmaps represent the results using an adenine base editing system, and the bottom three heatmaps represent the results using a cytosine base editing system. The targeted genes are listed at the top of each heatmap (HEK1-1, EMX1-2, site9). Each row of each heatmap represents a base editing system, where UNT represents a negative control, TnpB-ABE and TnpB-CBE represent base editing systems containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-ABE and enTnpB-CBE represent base editing systems containing the dead-enTnpB protein of SEQ ID NO. 9. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0050] Figure 41 is a heatmap graph showing the results of gene editing using various base editing systems according to Experimental Example 5.4. The targeted gene is listed at the top of each heatmap (EMX1, EMX2). Each row of each heatmap represents a base editing system, UNT represents a negative control, and each label represents a base editing system with the configuration listed in Table 14. Each column of the heatmap represents the TAM sequence (indicated by a shaded box) and protospacer sequence of the target nucleic acid. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0051] Figure 42 shows the results of in-silico selection of targets expected to show therapeutic effects using a dead-enTnpB-based base editing system utilizing pathological SNV data listed in the ClinVar database.

[0052] Figure 43 is a schematic diagram showing the method of adding a stabilization motif to reRNA according to Experimental Example 6.

[0053] Figure 44 is a graph showing the frequency (%) of indel introduction into the TnpB complex according to the addition of a stabilization motif to the reRNA in accordance with Experimental Example 6.1. The horizontal axis of the graph represents the reRNA of each composition, and the composition by label is as disclosed in Table 14.

[0054] Figure 45 is a heatmap graph showing the gene editing frequency (%) of a base editing system according to the addition of a stabilization motif to reRNA in accordance with Experimental Example 6.2. The experiment was performed using a cytosine base editing system targeting the HEK1-1 gene. Each row of the heatmap represents the reRNA of each composition, and the composition by label is as disclosed in Table 14. Each column of the heatmap represents the TAM sequence (indicated by a shaded box) and protospacer sequence of the target nucleic acid. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0055] Figure 46 is a graph showing the gene editing frequency (%) of a base editing system according to the addition of a stabilization motif to reRNA in accordance with Experimental Example 6.2. The experiment was performed using a cytosine base editing system targeting the HEK1-1 gene. The labels on the horizontal axis of the graph indicate the location of the cytidine targeted for correction. For example, C2 refers to the cytidine at position 2 relative to the protospacer. For each label, from the left, the 1st bar represents ω*RNA, the 2nd bar represents 5' evopreQ1-ω*RNA, the 3rd bar represents 5' evopreQ1-linker-ω*RNA, the 4th bar represents 5' mpknot-ω*RNA, the 5th bar represents 5' mpknot-linker-ω*RNA, the 6th bar represents 3' evopreQ1-ω*RNA, the 7th bar represents 3' evopreQ1-linker-ω*RNA, the 8th bar represents 3' mpknot-ω*RNA, the 9th bar represents 3' mpknot-linker-ω*RNA, and the 10th bar represents the negative control (UNT).

[0056] Figure 47 is a heatmap graph showing the gene editing frequency (%) of a base editing system according to the addition of a stabilization motif to reRNA in accordance with Experimental Example 6.2. The experiment was performed using an adenine base editing system targeting the HEK1-1 gene. Each row of the heatmap represents the reRNA of each composition, and the composition by label is as disclosed in Table 14. Each column of the heatmap represents the TAM sequence (indicated by a shaded box) and protospacer sequence of the target nucleic acid. The nucleobase editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0057] Figure 48 is a graph showing the gene editing frequency (%) of a base editing system according to the addition of a stabilization motif to reRNA in accordance with Experimental Example 6.2. The experiment was performed using an adenine base editing system targeting the HEK1-1 gene. The labels on the horizontal axis of the graph indicate the location of the adenosine targeted for correction. For example, A3 refers to the adenosine at position 3 relative to the protospacer. For each label, from the left, the 1st bar represents ω*RNA, the 2nd bar represents 5' evopreQ1-ω*RNA, the 3rd bar represents 5' evopreQ1-linker-ω*RNA, the 4th bar represents 5' mpknot-ω*RNA, the 5th bar represents 5' mpknot-linker-ω*RNA, the 6th bar represents 3' evopreQ1-ω*RNA, the 7th bar represents 3' evopreQ1-linker-ω*RNA, the 8th bar represents 3' mpknot-ω*RNA, the 9th bar represents 3' mpknot-linker-ω*RNA, and the 10th bar represents the negative control (UNT).

[0058] Figure 49 is a schematic diagram showing the structure of the multiplexing vector according to Experimental Example 7, and the process in which reRNA is self-cleaved and self-matured to edit each target.

[0059] Figure 50 is a schematic diagram showing the experimental process of Experimental Example 7.1.

[0060] Figure 51 is a graph showing the frequency (%) of indel introduction according to Experimental Example 7.1. The horizontal axis of the graph represents the target genes (HEK1-1, Site2, RUNX1-2). The results of the three groups on the left represent the experimental results using wild-type reRNA (reRNA), and the results of the three groups on the right represent the results using modified reRNA (reRNA*). For each label, the first bar from the left represents the negative control (UNT), the second bar represents the wild-type TnpB complex, and the third bar represents the results for the engineered TnpB complex containing the engineered TnpB of Sequence No. 2. The vertical axis of the graph represents the frequency (%) of indel introduction.

[0061] Figure 52 is a heatmap graph showing the results of base editing multiple targets according to Experimental Example 7.2. Target genes are listed at the top of the heatmap (HEK1-1, Site 2, RUNX1-2). The types of reRNA scaffolds used are listed on the left side of the heatmap. Here, reRNA refers to reRNA for wild-type TnpB, and reRNA* refers to reRNA with a portion of the scaffold removed. Each row of the heatmap represents an adenine base editing system, UNT refers to a negative control, TnpB-ABE refers to a base editing system containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-ABE refers to a base editing system containing the dead-enTnpB of SEQ ID NO. 9. Each column of the heatmap represents the TAM sequence (indicated by a shaded box) and protospacer sequence of the target nucleic acid. The nucleus-base editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0062] Figure 53 is a heatmap graph showing the results of base editing multiple targets according to Experimental Example 7.2. Target genes are listed at the top of the heatmap (HEK1-1, Site 2, RUNX1-2). The types of reRNA scaffolds used are listed on the left side of the heatmap. Here, reRNA refers to reRNA for wild-type TnpB, and reRNA* refers to reRNA with a portion of the scaffold removed. Each row of the heatmap represents a cytosine base editing system, UNT refers to a negative control, TnpB-CBE refers to a base editing system containing the dead-TnpB protein of SEQ ID NO. 8, and enTnpB-CBE refers to a base editing system containing the dead-enTnpB of SEQ ID NO. 9. Each column of the heatmap represents the TAM sequence (indicated by a shaded box) and protospacer sequence of the target nucleic acid. The nucleus-base editing efficiency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the right bar.

[0063] Figure 54 is a heatmap graph showing the results of confirming the frequency of indel introduction according to the expression order (order of proximity to the promoter) of reRNA within the multiplexing vector according to Experimental Example 7.3. The left side is a schematic diagram of the vector structure, where the arrows represent the promoter and the boxes represent the reRNAs targeting each respective target gene. Specifically, the boxes without boundaries represent the reRNA targeting HEK1-1, the boxes with dotted boundaries represent the reRNA targeting Site2, and the boxes with solid boundaries represent the reRNA targeting RUNX1-2. Each row of the heatmap box represents the frequency of indel introduction according to the vector structure disclosed in the corresponding row on the left. Each column of the heatmap box represents the target gene, and from the left, they represent the HEK1-1 target, the Site 2 target (dotted box), and the RUNX1-2 target (solid box). The frequency of indel introduction is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the bar at the bottom of each column.

[0064] Figure 55 is a heatmap graph showing the results of confirming the adenine base editing efficiency according to the expression order of reRNA within the multiplexing vector (order closest to the promoter) according to Experimental Example 7.3. The left side is a schematic diagram showing the vector structure, where the arrows represent the promoter and the boxes represent the reRNAs targeting each target gene. Specifically, the boxes without boundaries represent the reRNA targeting HEK1-1, the boxes with dotted boundaries represent the reRNA targeting Site 2, and the boxes with solid boundaries represent the reRNA targeting RUNX1-2. Each row of the heatmap box represents the frequency of nucleotide editing according to the vector structure disclosed in the corresponding row on the left. Each column of the heatmap box represents the nucleotide to be edited for each target gene. Specifically, the first three columns represent the HEK1-1 target, the next three columns represent the Site 2 target (dotted boxes), and the last three columns represent the RUNX1-2 target (solid boxes). A3 in the first column on the left refers to the adenosine at the 3rd position relative to the protospacer of the HEK1-1 target. The other columns are specified in the same way. The nucleobase edit frequency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the bottom bar of each column group.

[0065] Figure 56 is a heatmap graph showing the results of confirming the cytosine base editing efficiency according to the expression order of reRNA within the multiplexing vector (order closest to the promoter) according to Experimental Example 7.3. The left side is a schematic diagram showing the vector structure, where the arrows represent the promoter and the boxes represent the reRNAs targeting each target gene. Specifically, the boxes without boundaries represent the reRNA targeting HEK1-1, the boxes with dotted boundaries represent the reRNA targeting Site2, and the boxes with solid boundaries represent the reRNA targeting RUNX1-2. Each row of the heatmap box represents the frequency of nuclear base editing according to the vector structure disclosed in the corresponding row on the left. Each column of the heatmap box represents the nuclear bases targeted for editing for each target gene. Specifically, the first three columns represent the HEK1-1 target, the next three columns represent the Site 2 target (dotted boxes), and the last three columns represent the RUNX1-2 target (solid boxes). C2 in the first column on the left refers to the cytidine at the 2nd position relative to the protospacer of the HEK1-1 target. The other columns are specified in the same way. The nucleobase edit frequency is indicated by the brightness of each cell in the heatmap, and the brightness scale is shown on the bottom bar of each column group.

[0066] FIG. 57 schematically shows an AAV vector loaded with a TnpB system engineered according to Experimental Example 8.1. The upper vector configuration represents an AAV vector configured to target the Pcsk9 gene alone, and the lower vector configuration represents an AAV configured to target both the Pcsk9 gene and the Angptl3 gene.

[0067] Figure 58 schematically illustrates the experimental process of injecting an AAV vector into a mouse according to Experimental Example 8.1.

[0068] Figure 59 shows the target nucleic acid of a reRNA designed to target mouse Pcsk9. Specifically, the target nucleic acid of the corresponding reRNA is a region of Exon 8 of the mouse Pcsk9 gene.

[0069] Figure 60 shows the target nucleic acid of a reRNA designed to target mouse Angptl3. Specifically, the target nucleic acid of the corresponding reRNA is a region of Exon 3 of the mouse Angptl3 gene.

[0070] Figure 61 is a diagram showing the experimental results according to Experimental Example 8.2. The graph on the left shows the frequency (%) of indel introduction for the negative control group (UNT) and the experimental group (PCSK9-KO) administered a vector configured to target Pcsk9. The graph on the right shows the mRNA expression levels of the Pcsk9 gene and the expression levels of the Pcsk9 protein in the blood of the experimental group as relative multiples of the negative control group.

[0071] Figure 62 is a diagram showing the experimental results according to Experimental Example 8.2. The labels on the horizontal axis represent the negative control group (UNT) and the experimental group administered a vector configured to target Pcsk9 (PCSK9-KO), respectively. Starting from the left graph, the blood LDL concentration, triglyceride concentration, ALT concentration, and AST concentration are shown in order.

[0072] Figure 63 is a diagram showing the experimental results according to Experimental Example 8.2. The labels on the horizontal axis represent the liver, adrenal gland, heart, kidney, diaphragm, lung, and spleen, respectively, and the vertical axis represents the frequency of indel introduction within the corresponding tissue.

[0073] Figure 64 is a diagram showing the experimental results according to Experimental Example 8.3. The graph on the left shows the frequency (%) of indel introduction for the negative control group (UNT) and the experimental group (Multi-KO) administered a vector configured to target Pcsk9 and Angptl3. The graph on the right shows the mRNA expression levels of the Pcsk9 gene, Pcsk9 protein expression levels, Angptl3 mRNA expression levels, and Angptl3 protein expression levels in the blood of the experimental group as relative multiples of the negative control group.

[0074] Figure 65 is a diagram showing the experimental results according to Experimental Example 8.3. The labels on the horizontal axis represent the negative control group (UNT) and the experimental group (Multi-KO) administered with a vector configured to target Pcsk9 and Angptl3, respectively. Starting from the left graph, the LDL concentration, triglyceride concentration, ALT concentration, and AST concentration in the blood are shown in order.

[0075] Figure 66 is a diagram showing the experimental results according to Experimental Example 8.2. The graph on the left shows the frequency of indel introduction for Pcsk9, and the graph on the right shows the frequency of indel introduction for Angptl3. The labels on the horizontal axis of each graph represent the liver, adrenal gland, heart, kidney, diaphragm, lung, and spleen, respectively, and the vertical axis represents the frequency of indel introduction within the corresponding tissue.

[0076] The best modes for carrying out the invention are disclosed below by way of example. These include some embodiments of the invention disclosed herein, but not all embodiments. The embodiments described in this paragraph are merely illustrative and should not be understood as the only "best modes of the invention." A person skilled in the art would be able to conceive of many variations and more preferred embodiments of the examples described in this paragraph, and such should also be considered to be included in the best modes for carrying out the invention.

[0077] This specification provides an engineered Deinococcus radiodurans ISDra2-derived TnpB protein (ISDra2 TnpB), and

[0078] Herein, the engineered ISDra2 TnpB protein comprises an amino acid sequence in which one to five variations selected from the following are introduced into the amino acid sequence of SEQ ID NO. 1:

[0079] N4R; V8R; G49R; S57R; Q64R; S72R; K84R; Q128R; G140R; G146R; Q148R; V204R; T228R; K234R; A237R; Y239R; K243R; A247R; V254R; N255R; K256R; K263R; T266R; K281R; K310R; P339R; And T350R.

[0080] In one embodiment,

[0081] The above-mentioned engineered ISDra2 TnpB protein may include the amino acid sequence of SEQ ID NO. 2 or SEQ ID NO. 3.

[0082] In one embodiment,

[0083] The above-mentioned engineered ISDra2 TnpB protein may further include one or more Nuclear Localization Signals (NLS) at the N-terminus, C-terminus, or both the N-terminus and C-terminus.

[0084] This specification provides a composition for gene editing comprising the following:

[0085] The aforementioned engineered ISDra2 TnpB protein, or a nucleic acid encoding the engineered ISDra2 TnpB protein; and

[0086] right-end element RNA (reRNA), or nucleic acid encoding said reRNA;

[0087] Here, the reRNA comprises a scaffold and a spacer, and

[0088] The above scaffold can form a complex with ISDra2-derived TnpB protein, and

[0089] The above spacer can bind complementarily to a predetermined target nucleic acid.

[0090] In one embodiment,

[0091] The above reRNA may include the following:

[0092] 5'-[Scaffold]-[Spacer]-[Linker]-[Stabilization Domain]-3'

[0093] Here, the scaffold comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 5 to 7, and

[0094] The above linker includes or is absent the nucleic acid sequence of SEQ ID NO. 23, and

[0095] The above stabilization domain includes a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22, or is absent.

[0096] In one embodiment,

[0097] The gene editing composition may include the engineered ISDra2 TnpB protein and an RNA-guided nucleic acid cleavage complex to which the reRNA is bound.

[0098] In one embodiment,

[0099] The gene editing composition may include a vector comprising a nucleic acid encoding the engineered ISDra2 TnpB protein and a nucleic acid encoding the reRNA.

[0100] This specification provides an intracellular target gene editing method comprising the following:

[0101] A process of delivering any one of the aforementioned gene editing compositions to the cell;

[0102] Here, the spacer of the reRNA of the gene editing composition is configured to bind complementarily to a predetermined target nucleic acid included in the target gene, and

[0103] By the above delivery, an RNA-guided nucleic acid cleavage complex bound to the engineered ISDra2 TnpB protein of the composition and reRNA is delivered into the cell, or an RNA-guided nucleic acid cleavage complex is formed within the cell, and

[0104] Contact between the RNA-guided nucleic acid cleavage complex and the target gene is induced, and

[0105] By the above contact, the target gene is edited.

[0106] This specification provides an engineered ISDra2-derived TnpB protein (ISDra2 TnpB), and

[0107] Herein, the ISDra2 TnpB comprises the amino acid sequence of SEQ ID NO. 8; or

[0108] The amino acid sequence comprises an amino acid sequence specified by applying one to six variations selected from the following to the amino acid sequence of SEQ ID NO. 8:

[0109] N4R; V8R; G49R; S57R; Q64R; S72R; K84R; Q128R; G140R; G146R; Q148R; V204R; T228R; K234R; A237R; Y239R; K243R; A247R; V254R; N255R; K256R; K263R; T266R; K281R; K310R; P339R; And the T350R,

[0110] The above-mentioned engineered ISDra2-derived TnpB protein has its nucleic acid cleavage activity removed.

[0111] In one embodiment,

[0112] The above-mentioned engineered ISDra2 TnpB protein may include an amino acid sequence selected from the group consisting of SEQ ID NOs 9 to 10.

[0113] This specification provides a base editor protein comprising the following:

[0114] Any one of the aforementioned engineered ISDra2 TnpB proteins;

[0115] One or more nuclear localization signals (NLS); and

[0116] One or more base editing domains;

[0117] Here, the base editing domain is adenosine deaminase or cytidine deaminase.

[0118] In one embodiment,

[0119] The above base editor protein may include the following:

[0120] Any one of the aforementioned engineered ISDra2-derived TnpB proteins;

[0121] One or more nuclear localization signals (NLS); and

[0122] One or more adenosine deaminases.

[0123] In one embodiment,

[0124] The above base editor protein may include the following structure:

[0125] [NLS1] - [AD] - [Dead TnpB] - [NLS2]

[0126] Here, the NLS1 is a first nuclear localization signal, and

[0127] Here, the NLS2 is a second nuclear localization signal, and

[0128] The above AD is an adenosine deaminase, and

[0129] The above Dead TnpB is an engineered ISDra2-derived TnpB protein selected from any one of claims 9 to 10.

[0130] In one embodiment,

[0131] The above base editor protein may include the following:

[0132] Any one of the aforementioned engineered ISDra2-derived TnpB proteins;

[0133] One or more nuclear localization signals (NLS);

[0134] One or more cytidine deaminases; and

[0135] One or more uracil glycosylase inhibitors (UGI).

[0136] In one embodiment,

[0137] The above base editor protein may be a protein of the following structure:

[0138] [NLS1] - [CD] - [Dead TnpB] - [UGI] - [NLS2]

[0139] Here, the NLS1 is a first nuclear localization signal, and

[0140] Here, the NLS2 is a second nuclear localization signal, and

[0141] The above CD is cytidine deaminase, and

[0142] The above Dead TnpB is an engineered ISDra2-derived TnpB protein selected from any one of claims 9 to 10.

[0143] This specification provides a base editor composition comprising the following:

[0144] The aforementioned base editor protein, or a nucleic acid encoding the base editor protein; and

[0145] reRNA, or a nucleic acid encoding the reRNA;

[0146] Here, the reRNA comprises a scaffold and a spacer, and

[0147] The scaffold can bind to the base editor protein to form a base editor complex, and

[0148] The above spacer can bind complementarily to a predetermined target nucleic acid.

[0149] This specification provides a base editor composition comprising the following:

[0150] Any one of the aforementioned base editor proteins, or a nucleic acid encoding said base editor protein; and

[0151] reRNA, or a nucleic acid encoding the reRNA;

[0152] Here, the reRNA comprises a scaffold and a spacer, and

[0153] The scaffold can bind to the base editor protein to form a base editor complex, and

[0154] The above spacer can bind complementarily to a predetermined target nucleic acid.

[0155] In one embodiment,

[0156] The above reRNA may include the following:

[0157] 5'-[Scaffold]-[Spacer]-[Linker]-[Stabilization Domain]-3'

[0158] Here, the scaffold comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 5 to 7, and

[0159] The above linker includes or is absent the nucleic acid sequence of SEQ ID NO. 23, and

[0160] The above stabilization domain includes a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22, or is absent.

[0161] This specification provides a base editor composition comprising the following:

[0162] A base editor protein selected from any one of claims 14 to 15, or a nucleic acid encoding said base editor protein; and

[0163] reRNA, or a nucleic acid encoding the reRNA;

[0164] Here, the reRNA comprises a scaffold and a spacer, and

[0165] The scaffold can bind to the base editor protein to form a base editor complex, and

[0166] The above spacer can bind complementarily to a predetermined target nucleic acid.

[0167] In one embodiment,

[0168] The above reRNA may include the following:

[0169] 5'-[Scaffold]-[Spacer]-[Linker]-[Stabilization Domain]-3'

[0170] Here, the scaffold comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 5 to 7, and

[0171] The above linker includes or is absent the nucleic acid sequence of SEQ ID NO. 23, and

[0172] The above stabilization domain includes a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22, or is absent.

[0173] This specification provides a method for editing target DNA within a cell, and

[0174] Here:

[0175] The above target DNA is double-stranded DNA comprising a first strand and a second strand inversely complementary thereto;

[0176] The above first strand includes a first correction window;

[0177] The second strand includes a second correction window at a position corresponding to the first correction window;

[0178] The first correction window and the second correction window are inversely complementary to each other; and

[0179] The first calibration window and the second calibration window are collectively referred to as calibration windows.

[0180] The above method includes the following:

[0181] The process of delivering any one of the aforementioned base editing compositions to the cell,

[0182] Here:

[0183] In the above cell, the base editor protein of the composition and reRNA bind to form a base editor complex; and

[0184] One or more C:G base pairs contained in the correction window are edited into T:A base pairs by the base editor complex.

[0185] In one embodiment,

[0186] The length of the above correction window is l-bp;

[0187] The base editor complex recognizes a sequence consisting of the m-th nucleotide to the (m + 4)-th nucleotide in the upstream direction relative to the 5'-terminus of the first correction window as a transposon-associated motif (TAM);

[0188] The above reRNA includes a spacer of length n-nt;

[0189] The above m is an integer between 2 and 10;

[0190] The above l is an integer between 1 and 9; and

[0191] The above n may be an integer between 18 and 22.

[0192] In one embodiment,

[0193] The above m can be 2, the above l can be 6, and the above n can be 20.

[0194] This specification provides a method for editing target DNA within a cell, and

[0195] Here:

[0196] The above target DNA is double-stranded DNA comprising a first strand and a second strand inversely complementary thereto;

[0197] The above first strand includes a first correction window;

[0198] The second strand includes a second correction window at a position corresponding to the first correction window;

[0199] The first correction window and the second correction window are inversely complementary to each other; and

[0200] The first calibration window and the second calibration window are collectively referred to as calibration windows.

[0201] The above method includes the following:

[0202] The process of delivering any one of the aforementioned base editing compositions to the cell,

[0203] Here:

[0204] In the above cell, the base editor protein of the composition and reRNA bind to form a base editor complex; and

[0205] One or more A:T base pairs included in the correction window are edited into G:C base pairs by the base editor complex.

[0206] In one embodiment,

[0207] The length of the above correction window is l-bp;

[0208] The base editor complex recognizes a sequence consisting of the m-th nucleotide to the (m + 4)-th nucleotide in the upstream direction relative to the 5'-terminus of the first correction window as a transposon-associated motif (TAM);

[0209] The above reRNA includes a spacer of length n-nt;

[0210] The above m is an integer between 2 and 11;

[0211] The above l is an integer between 1 and 10; and

[0212] The above n may be an integer between 18 and 22.

[0213] In one embodiment,

[0214] The above m can be 2, the above l can be 6, and the above n can be 20.

[0215] Hereinafter, with reference to the attached drawings, the content of the invention will be described in more detail through specific embodiments and examples. It should be noted that the attached drawings include some embodiments of the invention, but not all embodiments. The content of the invention disclosed by this specification may be implemented in various ways and is not limited to the specific embodiments described herein. Such embodiments should be considered as provided to satisfy the legal requirements applicable to this specification. A person skilled in the art to which the invention disclosed in this specification belongs would be able to conceive of many variations and other embodiments of the content of the invention disclosed in this specification. Therefore, the content of the invention disclosed in this specification is not limited to the specific embodiments described herein, and variations and other embodiments thereof should be understood to be included within the scope of the claims.

[0216] Definition of Terms

[0217] nearby, approximately, about

[0218] As used in this specification, the terms “near,” “approximately,” or “about” mean a quantity, level, value, number, frequency, percentage, dimension, size, amount, weight, or length that varies by about 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1%, or 0% with respect to a reference quantity, level, value, number, frequency, percentage, dimension, size, amount, weight, or length. Additionally, the terms are used to include a reasonable range that takes into account measurement errors, experimental conditions, manufacturing deviations, etc.

[0219] Singular and plural expressions

[0220] In this specification, unless otherwise specified in the context, terms expressed in the singular form are used to include the plural form, and terms expressed in the plural form are also used to include the singular form. For example, the term "bilirubin-amino sugar complex" may include not only a single bilirubin-amino sugar complex but also multiple bilirubin-amino sugar complexes. As another example, the expression "multiple bilirubin-amino sugar complexes" may be used to include individual bilirubin-amino sugar complexes.

[0221] Includes

[0222] Expressions such as "includes" as used in this specification are used to mean that they do not exclude other components not enumerated, unless otherwise specified.

[0223] and / or

[0224] As used in this specification, "and / or" is used to mean including all combinations of one or more of the listed items.

[0225] It could be

[0226] As used in this specification, expressions such as "... may be" or ""... may be" indicate optional configurations or optional features. They do not imply mandatory components.

[0227] Amino acid sequence notation

[0228] Unless otherwise stated, amino acid sequences described in this specification shall be written from the N-terminal to the C-terminal using either the single-letter or triple-letter amino acid notation. For example, if denoted as PAAK, it signifies a peptide in which proline, alanine, alanine, and lysine are linked in sequence from the N-terminal to the C-terminal. As another example, if denoted as Thr-Leu-Lys, it signifies a peptide in which threonine, leucine, and lysine are linked in sequence from the N-terminal to the C-terminal. For amino acids that cannot be represented by the above single-letter notation, other letters shall be used, and further supplementary explanations shall be provided. The notation methods for each amino acid are as follows: Alanine (Ala, A); Arginine (Arg, R); Asparagine (Asn, N); Aspartic acid (Asp, D); cysteine ​​(Cys, C); glutamic acid (Glu, E); glutamine (Gln, Q); glycine (Gly, G); histidine (His, H); isoleucine (Ile, I); leucine (Leu, L); lysine (Lys, K); methionine (Met, M); phenylalanine (Phe, F); proline (Pro, P); serine (Ser, S); threonine (Thr, T); tryptophan (Trp, W); tyrosine (Tyr, Y); and valine (Val, V).

[0229] Nucleic acid sequence notation

[0230] The symbols A, T, C, G, and U used in this specification are interpreted as having the meaning understood by a person skilled in the art. Depending on the context and description, they may be appropriately interpreted as bases, nucleosides, or nucleotides on DNA or RNA. For example, when referring to a base, it may be interpreted as adenine (A), thymine (T), cytosine (C), guanine (G), or uracil (U) itself, respectively; when referring to a nucleoside, it may be interpreted as adenosine (A), thymidine (T), cytidine (C), guanosine (G), or uridine (U), respectively; and when referring to a nucleotide in a sequence, it should be interpreted as a nucleotide containing each of the above nucleosides.

[0231] Target gene or target nucleic acid

[0232] As used in this specification, "target gene" or "target nucleic acid" essentially refers to an intracellular gene or nucleic acid that is the subject of gene editing. The terms "target gene" or "target nucleic acid" may be used interchangeably and may refer to the same subject. Unless otherwise stated, the target gene or target nucleic acid may refer to both the genes or nucleic acids inherent to the target cell and genes or nucleic acids of external origin, and is not particularly limited as long as they can be the subject of gene editing. The target gene or target nucleic acid may be single-stranded DNA, double-stranded DNA, and / or RNA. Furthermore, the terms encompass all meanings recognizable by a person skilled in the art and may be interpreted appropriately according to the context.

[0233] vector

[0234] As used herein, the term "vector" refers collectively to any material capable of transporting genetic material into a cell, unless otherwise specified. For example, a vector may be, but is not limited to, a DNA molecule containing the target genetic material, such as a nucleic acid encoding the TnpB protein or a variant thereof, and a nucleic acid encoding reRNA or a variant thereof. The above term encompasses all meanings recognizable by a person skilled in the art and may be interpreted appropriately according to the context.

[0235] Method for determining the position of elements included in a specific sequence

[0236] Unless otherwise stated, the position of each element within a specific sequence in this specification is defined as follows. Here, for nucleic acid sequences, each element is an individual nucleotide, and for amino acid sequences, each element is an individual amino acid. For other sequences, a range appropriate to be treated as a single unit in context is defined as an element.

[0237] 1) It is assumed that a specific sequence is specified enough to determine the position of at least each element.

[0238] 2) The upstream, upstream direction, or upstream region of the sequence, and the downstream, downstream direction, or downstream region are determined as follows:

[0239] In the case of a nucleic acid sequence, the 5'-terminal direction of a specific nucleotide is defined as the upstream direction, and the region located in the upstream direction is defined as the upstream region. Conversely, the 3'-terminal direction of a specific nucleotide is defined as the downstream direction, and the region located in the downstream direction is defined as the downstream region.

[0240] In the case of an amino acid sequence, the N-terminal direction of a specific amino acid is defined as the upstream direction, and the region located in the upstream direction is defined as the upstream region. Conversely, the C-terminal direction of a specific amino acid is defined as the downstream direction, and the region located in the downstream direction is defined as the downstream region.

[0241] For other sequences, when the above sequence is expressed in the conventional manner of the relevant technical field, the direction of the element expressed before a specific element is defined as the upstream direction, and the region located in the upstream direction is defined as the upstream region. Conversely, when the above sequence is expressed in the conventional manner of the relevant technical field, the direction of the element expressed after a specific element is defined as the downstream direction, and the region located in the downstream direction is defined as the downstream region.

[0242] 3) Unless otherwise noted, the first element is the element at the top of the sequence.

[0243] In the case of a nucleic acid sequence, the nucleotide located at the 5'-terminus is the 1st nucleotide.

[0244] In the case of an amino acid sequence, the amino acid located at the N-terminus is the first amino acid.

[0245] For other sequences, the first element is the element that is listed first when the above sequence is written in the conventional manner of the relevant technical field.

[0246] 4) Unless otherwise noted, the last element is the downstream element of the sequence.

[0247] In the case of a nucleic acid sequence, the nucleotide located at the 3'-terminus is the last nucleotide.

[0248] In the case of an amino acid sequence, the amino acid located at the C-terminus is the last amino acid.

[0249] For other sequences, the last element is the element listed last when the above sequence is written in the conventional manner of the relevant technical field.

[0250] 5) Unless otherwise noted, the position of the first element is position 1, and position numbers are assigned in ascending order downstream.

[0251] In the case of nucleic acid sequences, the position where the 1st nucleotide at the 5' end is located is position 1, and they are listed in ascending order in the downstream direction.

[0252] In the case of amino acid sequences, the position where the 1st amino acid of the N-terminus is located is position 1, and they are listed in ascending order in the downstream direction.

[0253] For other sequences, when the above sequence is expressed in the conventional manner of the relevant technical field, the position of the first element expressed is position 1, and they are expressed in ascending order in the downstream direction.

[0254] 6) When a reference element is set, the position of the reference element is position 0. In other words, the reference element is the 0th element.

[0255] 7) Where a reference element is defined and no other direction is defined, "n consecutive sequences from the reference element" means a sequence consisting of a total of n consecutive elements in the downstream direction, including the reference element. Here, n is an integer.

[0256] 8) Where a reference element and a direction are defined, "n consecutive sequences in the direction defined from the reference element" means a sequence consisting of a total of n consecutive elements in the direction defined, including the reference element. Here, n is an integer. However, regardless of the direction defined above, the sequence is always indicated from upstream to downstream.

[0257] 9) When two different elements are specified, "sequence consisting of one element or another element" means a sequence consisting of the two elements and all consecutive elements in between.

[0258] 10) The above numbers may also be applied to elements outside of a specific sequence. That is, when the first sequence includes the second sequence, the numbers of elements that are not included in the second sequence but are included in the first sequence can be defined based on the elements of the second sequence.

[0259] 11) The terms “m-th element,” “m-th element,” and “m-th position” within the same sequence and based on the same criteria may be used interchangeably and should be interpreted appropriately according to the context.

[0260] Example of element order within a sequence

[0261] For example, if the specific amino acid sequence above is ARNDCEQGHI (Sequence No. 80):

[0262] The 1st amino acid, the 1st amino acid, the amino acid at the 1st position, or the amino acid at the 1st position is A;

[0263] The last amino acid, or the amino acid at the last position, is I;

[0264] The 4th amino acid, the 4th amino acid, the amino acid at the 4th position, or the amino acid at the 4th position is D;

[0265] Based on the above 4th amino acid (D), the 1st amino acid in the upstream direction is N, and the 1st amino acid in the downstream direction is itself, C;

[0266] Based on the 4th amino acid (D) above, the 2nd amino acid in the downstream direction is E;

[0267] Based on the 4th amino acid (D) above, the 1st amino acid in the upstream direction is N;

[0268] With respect to the 1st amino acid (N) in the upstream direction of the 4th amino acid (D) above, three consecutive sequences are NDCs;

[0269] The 5th amino acid, the 5th amino acid, the amino acid at the 5th position, or the amino acid at the 5th position is C;

[0270] Based on the 5th amino acid (C) above, three consecutive sequences upstream are NDCs;

[0271] The sequence consisting of the 6th to 8th amino acids is EQG.

[0272] Prior art literature

[0273] The prior art literature cited in this specification for the description of the technical content is as follows:

[0274] Prior Literature 1: Karvelis, T., Druteika, G., Bigelyte, G. et al. Transposon-associated TnpB is a programmable RNA-guided DNA endonuclease. Nature 599, 692-696 (2021). https: / / doi.org / 10.1038 / s41586-021-04058-1.

[0275] Prior Literature 2: Li, Z., Guo, R., Sun, X. et al. Engineering a transposon-associated TnpB-ωRNA system for efficient gene editing and phenotypic correction of a tyrosinaemia mouse model. Nat Commun 15, 831 (2024). https: / / doi.org / 10.1038 / s41467-024-45197-z.

[0276] Prior Literature 3: Han Altae-Tran et al. ,The widespread IS200 / IS605 transposon family encodes diverse programmable RNA-guided endonucleases.Science374,57-65(2021). DOI:10.1126 / science.abj6856.

[0277] Deinococcus radiodurans-derived TnpB (ISDra2 TnpB) protein and reRNA

[0278] TnpB protein and reRNA derived from Deinococcus radiodurans

[0279] The TnpB gene is found in the IS200 / IS605 family, a transposon system. While the function of the TnpA gene from the same family has been elucidated in detail, the function of the TnpB gene has not been clearly determined. However, bioinformatics analysis has merely raised the possibility that the TnpB gene may be an ancestor of CRISPR / Cas12.

[0280] Under these circumstances, the authors of Prior Art 1 investigated the function of TnpB in ISDra2, a member of the IS200 / IS605 family of Deinococcus radiodurans. The study revealed that the TnpB protein expressed by the gene forms a complex with RNA containing a specific sequence, recognizes specific motifs in double-stranded nucleic acids, and, induced by said RNA, can cause double-stranded nucleic acid cleavage. Here, the RNA containing said specific sequence was named right-end element RNA (reRNA) (Prior Art 1 et al.) or omega RNA (Obligate Mobile Element-Guided Activity RNA; OMEGA RNA) (Prior Art 2 and Prior Art 3 et al.). This specification refers to the RNA containing said specific sequence as reRNA, but the term omega RNA may also be used interchangeably. The authors of Prior Art 1 revealed that a nucleic acid cleavage system similar to the CRISPR / Cas system can be constructed using the said TnpB protein and reRNA. According to this, if the above TnpB protein and reRNA are well designed, the intended target nucleic acid can be cleaved.

[0281] Transposon-Associated Motif (TAM)

[0282] The above TnpB protein forms a complex with reRNA and recognizes a specific motif of double-stranded nucleic acid. Specifically, the ISDra2 TnpB protein derived from Deinococcus radiodurans is known to recognize the short nucleic acid sequence 5'-TTGAT-3'. This specific motif has been named the Transposon-Associated Motif (TAM). In some literature, it is also referred to as the Target Adjacent Motif (TAM). The above TAM corresponds to the Protospacer Adjacent Motif (PAM) recognized by Cas proteins in the CRISPR / Cas system.

[0283] Nucleic acid cleavage activity of TnpB protein-reRNA complex

[0284] According to prior art document 1, a TnpB protein and reRNA bind to form a complex, and the complex can recognize and cleave a target nucleic acid. It has been revealed that the nucleic acid cleavage mechanism of the complex is very similar to that of a CRISPR / Cas complex. Specifically, when the complex of the TnpB protein and reRNA comes into contact with a target nucleic acid, 1) the TnpB protein recognizes TAM, 2) a region of the reRNA that recognizes the target sequence (referred to as a spacer or guide domain) binds complementarily to the target sequence near the TAM sequence, and 3) the nucleic acid cleavage domain (RuvC-Like domain) of the TnpB protein acts to cleave the target nucleic acid. If the TnpB protein is correlated with a Cas protein and the reRNA is correlated with a guide RNA, the RNA-induced nucleic acid cleavage complex containing the TnpB protein and reRNA (hereinafter referred to as the TnpB complex) can be understood as similar to a CRISPR / Cas complex. Therefore, the technologies introduced in the process of utilizing CRISPR / Cas complexes can be applied to the above TnpB complex almost as they are.

[0285]

[0286] Chapter 1. Gene Editing Method Using Engineered ISDra2 TnpB Protein

[0287] Engineered ISDra2 TnpB Protein #1 - Variant with Enhanced Nucleic Acid Cleavage Activity

[0288] Overview of variants with enhanced nucleic acid cleavage activity

[0289] This specification discloses engineered ISDra2 TnpB proteins, specifically TnpB protein variants with enhanced nucleic acid cleavage activity. The inventors of this application identified sites that interact critically with the target nucleic acid or reRNA during the process in which TnpB protein-reRNA cleaves the target nucleic acid, and by substituting these sites with arginine, obtained TnpB protein variants with enhanced nucleic acid cleavage efficiency. In this specification, "engineered ISDra2 TnpB protein" is also referred to simply as "engineered TnpB protein" or "TnpB protein variant." Additionally, the mutation that must be introduced from the wild-type TnpB protein to obtain the engineered TnpB protein is referred to as the "protein-nucleic acid interaction enhancing mutation." The engineered TnpB protein forms a complex by binding to reRNA that acts in conjunction with the wild-type TnpB protein, and is capable of target-specific cleaving of nucleic acids. The above-described engineered TnpB protein is a protein comprising a variant in which at least one amino acid residue is substituted with arginine, based on the amino acid sequence of the wild-type TnpB protein. A TnpB protein-reRNA complex containing the above-described engineered TnpB protein exhibits enhanced nucleic acid cleavage activity compared to a TnpB protein-reRNA complex containing the wild-type TnpB protein.

[0290] The specific process of discovering TnpB protein variants is explained in more detail in the following paragraphs.

[0291] Engineered TnpB Protein #1 - Substitution Target Site

[0292] The above-mentioned engineered TnpB protein is a protein in which, based on the amino acid sequence of the wild-type TnpB protein, i.e., SEQ ID NO. 1, the corresponding amino acid is substituted with another amino acid at 1 to 5 selected positions among the positions described below:

[0293] 4th asparagine; 8th valine; 49th glycine; 57th serine; 64th glutamine; 72nd serine; 84th lysine; 128th glutamine; 140th glycine; 146th glycine; 148th glutamine; 204th valine; 228th threonine; 234th lysine; 237th alanine; 239th tyrosine; 243rd lysine; 247th alanine; 254th valine; 255th asparagine; 256th lysine; 263rd lysine; 266th threonine; 281st lysine; 310th lysine; 339th proline; and 350th threonine.

[0294] The amino acids at the above positions are amino acid residues located close to the reRNA or target DNA during the process in which the TnpB protein-reRNA complex cleaves target DNA, and thus are expected to play an important role in the interaction between protein-target DNA or protein-reRNA.

[0295] Engineered TnpB Protein #2 - Target Substitution Amino Acid

[0296] The engineered TnpB protein above may be a protein in which the amino acid at the selected position above is substituted with arginine. Arginine is known to carry a strong positive charge, while nucleic acids are generally known to carry a negative charge. Therefore, it can be expected that substituting the amino acid at the above position with arginine will enhance the interaction between the protein and the nucleic acid (DNA or RNA).

[0297] Engineered TnpB Protein #3 - May contain Nuclear Localization Signal (NLS)

[0298] The engineered TnpB protein may include one or more nuclear localization signals at the N-terminus, C-terminus, or both ends. If the engineered TnpB protein includes nuclear localization signals, it can be delivered into the nucleus of a eukaryotic cell. Therefore, the engineered TnpB protein containing the nuclear localization signals can edit the eukaryotic cell genome. For example, delivering the engineered TnpB protein and appropriately designed reRNA to a eukaryotic cell can induce indels in the genome of the eukaryotic cell.

[0299] Engineered TnpB Protein #4 - Transposon-Associated Motif (TAM) Recognition

[0300] The engineered TnpB protein can recognize a TAM contained in a target nucleic acid. For example, the TAM may be the nucleic acid sequence 5'-TTGAT-3'. As another example, the TAM may be the nucleic acid sequence 5'-TNGAT-3' or 5'-TTGVT-3'. Here, N is A, T, C, or G, and V is A, C, or G.

[0301] Engineered TnpB Protein Feature #1 - Operates with wild-type TnpB protein reRNA

[0302] The above-mentioned engineered TnpB protein can function with reRNA that functions with the wild-type TnpB protein. In other words, if the scaffold of the reRNA can interact with the wild-type TnpB protein to form a complex, the reRNA can also interact with the above-mentioned engineered TnpB protein to form a complex. Therefore, the above-mentioned engineered TnpB protein can be used not only with the reRNA of the previously disclosed ISDra2 TnpB protein, but also with reRNA in which the scaffold has been mutated to the extent that it can function with the wild-type TnpB.

[0303] Variant Features with Enhanced Nucleic Acid Cleavage Activity #2 - Cell Genome Editing Capable

[0304] The engineered TnpB protein described above can edit the cell genome. In particular, if the engineered TnpB protein contains a nuclear localization signal, it can be delivered to the cell nucleus and used to edit the eukaryotic cell genome. The inventors of this application have demonstrated through experiments that the engineered TnpB protein-reRNA complex actually edits the eukaryotic cell genome.

[0305] Variant with enhanced nucleic acid cleavage activity Feature #3 - Possesses extended TAM

[0306] The engineered TnpB protein recognizes not only the 5'-TTGAT-3' sequence, which the wild-type TnpB protein recognizes as a TAM, but also additional TAMs. As a result, the engineered TnpB protein can be designed to target a wider variety of gene sequences compared to the wild-type TnpB protein.

[0307] Characteristics of the variant with enhanced nucleic acid cleavage activity #4 - Enhanced nucleic acid cleavage activity

[0308] The above-described engineered TnpB protein is characterized by enhanced nucleic acid cleavage activity compared to the wild-type TnpB protein. This is because, during the process in which the TnpB protein-reRNA complex cleaves target nucleic acids, an amino acid that plays a crucial role in protein-nucleic acid interactions is substituted with arginine, which strengthens these interactions. The inventors of this application 1) formulated the hypothesis that "strengthening protein-nucleic acid interactions improves nucleic acid cleavage efficiency," 2) identified candidate sites capable of strengthening protein-nucleic acid interactions and designed candidate TnpB variants in which the amino acid at those sites is substituted with arginine, and 3) synthesized the aforementioned candidate TnpB variants and verified their nucleic acid cleavage activity to discover the above-described engineered TnpB protein. The process of discovering the engineered TnpB is described below.

[0309] Variant Discovery Process #1 - Substitution Site Selection

[0310] The inventors of this application anticipated that during the process in which TnpB protein-reRNA binds to and cleaves a target nucleic acid, the strength of the interaction between the TnpB protein and the target nucleic acid or reRNA would have a significant influence on the cleavage of the target nucleic acid. Specifically, they hypothesized that if the protein-nucleic acid interaction becomes stronger, the nucleic acid cleavage activity would be enhanced. Accordingly, the inventors analyzed the three-dimensional structure when TnpB protein-reRNA binds to the target nucleic acid to identify amino acid positions where protein-nucleic acid interactions are expected to occur. Specifically, these amino acids are located within 4 angstroms of the target nucleic acid or reRNA in the three-dimensional structure. A known Cryo-EM structure was utilized for the analysis of the three-dimensional structure. The inventors determined that the amino acids at these positions were highly likely to contribute to protein-nucleic acid interactions and intended to create variants by substituting these positions with other amino acids.

[0311] Variant Discovery Process #2 - Selection of Substitution Targets

[0312] Nucleic acids generally carry a negative charge in the body. This is because the phosphate groups contained in nucleic acids carry a negative charge. Therefore, when a positively charged amino acid comes into contact with a nucleic acid during protein-nucleic acid interaction, it can be expected that they will interact strongly due to electrostatic attraction. Amino acids that carry a positive charge in the body include arginine, lysine, and histidine. Among these, arginine has properties suitable for application in functional sites of proteins, such as carrying a strong positive charge, being structurally flexible, and being able to form hydrogen bonds. Accordingly, the inventors of this application selected arginine as the target amino acid for substitution.

[0313] Variant Discovery Process #3 - Verification of the Effect of Single Position Substitution

[0314] The inventors of this application prepared TnpB protein variants in which only one of the specific substitution sites above was substituted with arginine and evaluated their nucleic acid cleavage activity. The specific evaluation process is disclosed in the experimental examples. As a result of the evaluation, TnpB protein variants incorporating the following mutations exhibited higher nucleic acid cleavage efficiency compared to the wild-type TnpB protein:

[0315] 4th asparagine; 8th valine; 49th glycine; 57th serine; 64th glutamine; 72nd serine; 84th lysine; 128th glutamine; 140th glycine; 146th glycine; 148th glutamine; 204th valine; 228th threonine; 234th lysine; 237th alanine; 239th tyrosine; 243rd lysine; 247th alanine; 254th valine; 255th asparagine; 256th lysine; 263rd lysine; 266th threonine; 281st lysine; 310th lysine; 339th proline; and 350th threonine.

[0316] Not all of the variants specified above improve nucleic acid cleavage efficiency, and only the variants introduced at specific locations improve nucleic acid cleavage efficiency. TnpB protein variants containing each of the above variants correspond to the "engineered TnpB protein with improved nucleic acid cleavage efficiency" of this specification.

[0317] Variant Discovery Process #4 - Verification of the Effects of Combining Variants

[0318] The inventors of this application prepared TnpB protein variants in which two or more amino acids were substituted with arginine by combining single-position substitutions that improve nucleic acid cleavage efficiency as described above, and evaluated the nucleic acid cleavage activity. The specific evaluation process is disclosed in the experimental examples. As a result of the evaluation, a consistent tendency for nucleic acid cleavage efficiency to increase was observed when single-position substitutions were combined. However, the gene editing effect saturates at a combination of five variants, and when six or more combinations of variants are introduced, nucleic acid cleavage efficiency tends to decrease, except for some combinations. In conclusion, TnpB protein variants composed of two to six single-position substitutions found as described above also correspond to the "engineered TnpB protein with improved nucleic acid cleavage efficiency" of this specification.

[0319] Important variant location

[0320] For example, the engineered TnpB protein may include one or more of the following selected variants compared to the wild-type TnpB protein:

[0321] Based on the amino acid sequence of wild-type TnpB protein, glutamine at position 128 is replaced with arginine,

[0322] Here, the location is expected to interact with target nucleic acids, particularly nucleic acids near the TAM;

[0323] Based on the amino acid sequence of wild-type TnpB protein, the 72nd serine is replaced with arginine,

[0324] Here, the corresponding location is expected to interact with target nucleic acids, specifically the TAM portion;

[0325] Based on the amino acid sequence of wild-type TnpB protein, the 64th glutamine is replaced with arginine.

[0326] Here, the corresponding location is expected to interact with target nucleic acids, specifically the TAM portion;

[0327] Based on the amino acid sequence of wild-type TnpB protein, the 57th serine is replaced with arginine.

[0328] Here, the corresponding location is expected to interact with target nucleic acids, specifically the TAM portion;

[0329] Based on the amino acid sequence of wild-type TnpB protein, the 281st lysine is substituted with arginine,

[0330] Here, the corresponding site is expected to interact with the region where the guide domain of the target nucleic acid, in particular, the reRNA binds complementarily;

[0331] Based on the amino acid sequence of wild-type TnpB protein, the 255th asparagine is replaced with arginine,

[0332] Here, the site is expected to interact with reRNA, particularly the pseudonot portion;

[0333] Figure 21 schematically shows exactly which position the amino acid at each position interacts with the nucleic acid at which position.

[0334] Engineered ISDra2 TnpB Protein #2 - Shortened Amino Acid Sequence Variant

[0335] Overview of variants with shortened amino acid sequences

[0336] This specification discloses engineered TnpB proteins, specifically TnpB protein variants with shortened amino acid sequence lengths. The wild-type TnpB protein is a protein composed of 408 amino acids. The inventors of this application discovered that a portion of the 3'-terminal region of the gene encoding the TnpB protein overlaps with the region encoding reRNA. The inventors hypothesized that although this overlapping gene region is translated into an amino acid sequence and incorporated into the C-terminal region of the wild-type TnpB protein, it is an unnecessary region with no particular function. Based on this hypothesis, the region overlapping with the reRNA-coding region was identified, and after sequentially removing the corresponding region from the TnpB protein and expressing it, gene editing activity was verified. As a result of the verification,

[0337] Location to be removed

[0338] The above-mentioned engineered TnpB protein may have an amino acid sequence in which one or more selected from amino acids 374 to 408 are removed compared to the amino acid sequence of the wild-type TnpB protein (Sequence No. 1). For example, the engineered TnpB protein may have a sequence in which 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, or 34 amino acids are removed from the C-terminal amino acid of SEQ ID NO. 1 in the N-terminal direction.

[0339] Characteristic of the variant with shortened amino acid sequence length #1 - No effect on nucleic acid cleavage activity

[0340] It is confirmed that the variant with the shortened amino acid sequence length does not show a significant change in gene editing activity compared to the wild-type TnpB protein. Therefore, since the variant has a shorter coding sequence length compared to the wild-type TnpB protein when loaded into a vector, it has the advantage of being easy to utilize in genetic engineering technology.

[0341] Characteristics of variant with shortened amino acid sequence length #2 - Can be combined with the aforementioned variant

[0342] Based on the amino acid sequence of the wild-type TnpB protein (SEQ No. 1), the position where the amino acid substitution variant is introduced is not included in the position that is removed by the amino acid length reduction variant. Therefore, an engineered TnpB protein can be formed by introducing both the amino acid substitution variant and the amino acid length reduction variant into the wild-type TnpB protein. Accordingly, in this specification, "engineered TnpB protein" refers to a TnpB protein variant in which the variant described in the [Engineered ISDra2 TnpB protein #1 - variant with enhanced nucleic acid cleavage activity] paragraph and / or the variant described in the [Engineered ISDra2 TnpB protein #2 - variant with shortened amino acid sequence length] paragraph is introduced into the wild-type TnpB protein.

[0343] reRNA

[0344] ReRNA Overview

[0345] This specification discloses a reRNA that binds to the engineered TnpB protein to form an RNA-induced nucleic acid cleavage complex. The reRNA is similar in structure and function to the guide RNA of the CRISPR / Cas system. The reRNA comprises a guide domain and a scaffold. The guide domain, also referred to as a spacer, directs the RNA-induced nucleic acid cleavage complex to a target nucleic acid near the TAM. The scaffold interacts with the engineered TnpB protein, causing the reRNA to form a complex with the engineered TnpB protein.

[0346] This specification discloses reRNA containing a wild-type scaffold as well as reRNA containing an engineered scaffold.

[0347] Wild-type scaffold

[0348] As previously stated, prior art document 1 discloses a scaffold nucleic acid sequence of reRNA that binds to a wild-type ISDra2 TnpB protein to form an RNA-induced nucleic acid cleavage complex. This is referred to as the wild-type scaffold. The wild-type scaffold comprises the nucleic acid sequence GATTCAAGAATCCCGAAGTGAAGAATCTTGCCGTCCGTACATGGACTTGCCCGAACTGTGGGGAAACCCATGACCGAGACGAGAACGCTGCGCTGAACATTCGGCGTGAAGCGTTGGTGGCTGCGGGAATCTCAGACACCTTAAACGCTCATGGAGGCTATGTCAGACCTGCTTCGGCGGGCAATGGTCTGCGAAGTGAGAATCACGCGACTTTAGTCGTGTGAGGTTCAA (Sequence No. 5). The engineered TnpB protein disclosed in this specification may also interact with the wild-type scaffold to form a complex. In other words, the above reRNA may include a wild-type scaffold.

[0349] Engineered scaffold

[0350] Through Prior Art 2, it was revealed that even if the nucleic acid sequence of the above-mentioned wild-type scaffold is partially modified, its function is maintained, and furthermore, the RNA-induced nucleic acid cleavage complex can function more efficiently. The modified scaffold disclosed in the above prior study is referred to as a "known scaffold variant." The above-mentioned known scaffold variant includes the nucleic acid sequences of TGGTGGCTGCGGGAATCTCAGACACCTTAAACGCTCATGGAGGCTATGTCAGACCTGCTTCGGCGGGCAATGGTCTGCGAAGTGAGAATCACGCGACTTTAGTCGTGTGAGGTTCAA (Sequence No. 6) or tGGTGGCTGCGGGAATCTCAGACACCTTAAACGCTCATGGAGGCTATgaaaATGGTCTGCGAAGTGAGAATCACGCGACTTTAGTCGTGTGAGGTTCAA (Sequence No. 7). The engineered TnpB protein disclosed in this specification may also interact with the known scaffold variant to form a complex. Alternatively, the reRNA may include the known scaffold variant.

[0351] Guide Domain

[0352] The above reRNA contains a guide domain. The guide domain is also referred to as a spacer. The guide domain functions identically to the guide domain of the guide RNA in the CRISPR / Cas system. The guide domain targets a target nucleic acid to guide the engineered TnpB protein-reRNA complex toward the target nucleic acid. The guide domain can hybridize with the target nucleic acid near a transposon-associated motif (TAM). The guide domain is designed to target a predetermined target nucleic acid. The nucleic acid sequence of the guide domain is a nucleic acid sequence complementary to all or part of the target strand of the target nucleic acid, or a nucleic acid sequence equivalent to all or part of the non-target strand of the target nucleic acid.

[0353] Connection structure between guide domain and scaffold

[0354] The above reRNA is an RNA in which the scaffold and the guide domain are connected. Specifically, it has a structure in which the 3' end of the scaffold and the 5' end of the guide domain are connected. Here, the scaffold and the guide domain may be directly connected or may be connected through an RNA linker.

[0355] It may include additional stabilization domains

[0356] The above reRNA may further include one or more stabilization domains. If the above reRNA includes a stabilization domain, the stabilization domain is located at the 3' end, the 5' end, or both ends of the above reRNA. For example, the above reRNA includes a stabilization domain connected to the 5' end of the (engineered) scaffold. As another example, the above reRNA includes a stabilization domain connected to the 3' end of the guide domain. As yet another example, the above reRNA includes a first stabilization domain connected to the 5' end of the (engineered) scaffold and a second stabilization domain connected to the 3' end of the guide domain. The above reRNA further including a stabilization domain is characterized by high in vivo stability due to the stabilization domain. The stabilization domain may be evopreQ1 or mpknot. Specifically, the stabilization domain may include any one of the selected nucleic acid sequences among TTGACGCGGTTCTATCTAGTTACGCGTTAAACCAACTAGAAA (Sequence No. 21) and GGGTCAGGAGCCCCCCCCCTGAACCCAGGATAACCCTCAAAGTCGGGGGGCAACCC (Sequence No. 22).

[0357] Engineered TnpB Systems and Their Various Implementations

[0358] Overview of Engineered TnpB Systems and Their Various Implementations

[0359] This specification discloses a nucleic acid cleavage system comprising the aforementioned engineered TnpB protein and reRNA (hereinafter referred to as the engineered TnpB system). The engineered TnpB system recognizes and cleaves target nucleic acids and, when applied to cells, can induce cellular genome editing. In practice, the object editing the cellular genome is the engineered TnpB complex, which is a combination of the engineered TnpB protein and reRNA described above. However, with the advancement of biotechnology, it is possible to deliver nucleic acids encoding proteins and RNA to cells and express each protein and RNA within the cell to form an RNA-protein complex. Therefore, the term "engineered TnpB system" as used in this specification encompasses 1) the engineered TnpB complex, 2) a vector capable of expressing each component of the engineered TnpB complex (engineered TnpB system expression vector), and 3) the engineered TnpB composition. The engineered TnpB system is implemented and utilized in various ways depending on the mode of use. For example, the engineered TnpB system may be implemented as an "engineered TnpB complex" in which an engineered TnpB protein and reRNA are combined. As another example, the engineered TnpB system may be implemented as an "engineered TnpB system expression vector" comprising a nucleic acid encoding an engineered TnpB protein and a nucleic acid encoding reRNA. As yet another example, the engineered TnpB system may be implemented as a composition comprising mRNA encoding an engineered TnpB protein and reRNA.

[0360] Engineered TnpB complex

[0361] This specification discloses an RNA-induced nucleic acid cleavage complex comprising an engineered TnpB protein and reRNA. The RNA-induced nucleic acid cleavage complex is also referred to as an "engineered TnpB protein-reRNA complex" or an "engineered TnpB complex." When the engineered TnpB complex recognizes a TAM sequence contained in a target nucleic acid and recognizes and binds to a target sequence adjacent to the TAM sequence, it cleaves the target nucleic acid. Here, the engineered TnpB protein and reRNA are as described in the paragraph above.

[0362] Engineered TnpB System Expression Vector

[0363] This specification discloses an engineered TnpB system expression vector. The vector comprises a nucleic acid encoding an engineered TnpB protein, a nucleic acid encoding reRNA, and components that enable said engineered TnpB protein and reRNA to be expressed within a cell. The expression vector is not limited in its composition as long as it can express said components and may further include known components.

[0364] Engineered TnpB composition

[0365] This specification discloses a composition comprising an engineered TnpB protein, or a nucleic acid encoding said engineered TnpB protein; and reRNA or a nucleic acid encoding said reRNA. For example, said composition may comprise mRNA encoding said engineered TnpB protein and reRNA. said composition is referred to as an engineered TnpB composition. More specific compositions are disclosed in the [Numbered Examples] section.

[0366] Gene Editing Method Using Engineered TnpB Protein #1 - Introducing Indels into the Cell Genome

[0367] This specification discloses a gene editing method using an engineered TnpB protein. Specifically, the method comprises the process of delivering an engineered TnpB system to a cell. The engineered TnpB system may be implemented in various ways as an engineered TnpB complex, an expression vector, a composition, or any combination thereof. The process of delivering the engineered TnpB system to a cell may be performed in an appropriate manner depending on the specific implementation of the engineered TnpB system. Techniques used when delivering CRISPR / Cas systems to cells may be applied when delivering the engineered TnpB system to a cell. A person skilled in the art may select an appropriate method according to the purpose, environment, and circumstances. Specific implementations of the method are exemplified in the [Numbered Examples] section.

[0368] When the above-described engineered TnpB system is delivered to a cell, the engineered TnpB complex comes into contact with the cell's genome. Through this contact, the engineered TnpB complex binds to a predetermined target nucleic acid, cleaves it, and consequently functions to introduce an indel.

[0369] Gene Editing Method Using Engineered TnpB Protein #2 - Multiplexing Capable

[0370] This specification discloses a multiplex gene editing method using an engineered TnpB protein and multiple reRNAs. Specifically, the method comprises the process of delivering 1) an engineered TnpB protein or a nucleic acid encoding it, and 2) a nucleic acid encoding two or more types of reRNAs to a cell. Here, the nucleic acid encoding the reRNAs is a nucleic acid in which the nucleic acids encoding two or more types of reRNAs are linked without being separated by spacers, terminators, etc. Only one promoter is operatively linked to the nucleic acid encoding two or more types of reRNAs. The nucleic acids encoding two or more types of reRNAs 1) are independently cleaved during the transcription process without a special separator to mature into respective reRNAs, 2) each reRNA binds to the engineered TnpB protein to form a complex, and 3) each complex recognizes and cleaves the nucleic acid targeted by each reRNA. For example, the reRNAs may recognize the 5'-GAAC-3' motif in the scaffold to self-cleave and self-mature. By using the above method, only the engineered TnpB protein and one molecule of "nucleic acid encoding reRNA" can be delivered into the cell to induce editing of multiple target nucleic acids. More specific details are disclosed in the [Numbered Examples] section.

[0371]

[0372] Chapter 2. Base Editing Method Using Engineered ISDra2 TnpB Protein

[0373] Engineered dTnpB protein

[0374] Overview of Engineered dTnpB Protein

[0375] This specification discloses an engineered ISDra2 TnpB protein from which nucleic acid cleavage activity has been removed. The wild-type TnpB protein contains a protein domain similar to the RuvC domain of the Cas protein (RuvC-Like Domain) to cleave target nucleic acids. Prior Art 1 revealed that mutating specific amino acids of the nucleic acid cleavage domain of the wild-type TnpB protein inhibits nucleic acid cleavage activity. This specification developed an engineered TnpB protein by combining the mutation that inhibits the nucleic acid cleavage activity with the mutation that enhances the protein-nucleic acid interaction described in Chapter 1 and / or the mutation that shortens the length. The engineered TnpB protein is characterized by inhibiting nucleic acid cleavage activity but interacting more strongly with target nucleic acids or reRNA. The engineered TnpB protein recognizes the target nucleic acid in a sequence-specific manner and binds strongly to it, but does not cleave the target nucleic acid. The engineered TnpB protein can be utilized to construct a base editing system. To distinguish it from the engineered TnpB protein disclosed in Chapter 1, the above TnpB protein variant is hereinafter referred to as the engineered dTnpB (dead-TnpB) protein.

[0376] Introduce a mutation that removes nucleic acid cleavage activity

[0377] As previously stated, according to Prior Art 1, it is known that mutating the 191st amino acid of a wild-type TnpB protein to alanine inhibits the nucleic acid cleavage activity of the said TnpB protein. The said position corresponds to the RuvC-like domain, which is the nucleic acid cleavage domain of the TnpB protein. The engineered TnpB protein with removed nucleic acid cleavage activity disclosed herein also includes the said mutation. In other words, the engineered TnpB protein with removed nucleic acid cleavage activity includes the D191A mutation compared to the wild-type TnpB protein and may include additional mutations.

[0378] Can be combined with additional variations

[0379] The mutation that removes the nucleic acid cleavage activity inactivates the active amino acid of the nucleic acid cleavage domain of the TnpB protein. The mutation does not significantly affect the binding of the TnpB protein to the target nucleic acid, or the binding of the TnpB protein to reRNA. Additionally, the D191A mutation is not included in the amino acid region removed in the amino acid length reduction mutation. Therefore, the mutation that removes the nucleic acid cleavage activity can be combined with the protein-nucleic acid interaction enhancing mutation and / or the amino acid length reduction mutation disclosed in the paragraph [Chapter 1. Gene editing method using engineered ISDra2 TnpB protein]. When combining the mutations, a TnpB protein variant can be obtained in which 1) nucleic acid cleavage activity is suppressed, 2) protein-nucleic acid interactions are enhanced to bind more strongly to the target nucleic acid, reRNA, or both, and / or 3) the length is short, making it easy to load into a vector. For example, the variant of the above TnpB protein can be prepared by introducing both the D191A variant and the variant disclosed in the paragraph [Chapter 1. Gene editing method using engineered ISDra2 TnpB protein] into the wild-type TnpB protein. Hereinafter, in this specification, "engineered dTnpB protein" refers collectively to all TnpB protein variants in which the D191A variant is introduced into the wild-type TnpB protein and, optionally, one or more variants disclosed in the paragraph [Chapter 1. Gene editing method using engineered ISDra2 TnpB protein] are additionally introduced.

[0380] Can be used as a base editor by fusing the base editing domain.

[0381] The engineered dTnpB protein specifically recognizes and binds to target nucleic acids. Furthermore, the engineered dTnpB protein, into which a protein-nucleic acid interaction enhancing variant has been introduced, exhibits a stronger binding affinity to target nucleic acids compared to the wild-type TnpB protein. Therefore, by fusing a base editing domain to the engineered dTnpB protein, it can be utilized as a base editor. For example, the engineered dTnpB protein can be fused with adenosine deaminase to be used as an adenine base editor. As another example, the engineered dTnpB protein can be fused with cytidine deaminase to be used as a cytosine base editor. The inventors of this application have constructed a base editor system comprising the engineered dTnpB protein and have further demonstrated that the base editor system can correct the bases of target nucleic acids (see Experimental Examples).

[0382] Base editor protein containing engineered dTnpB protein

[0383] Base Editor Protein Overview

[0384] This specification discloses a base editor protein comprising an engineered dTnpB protein. The base editor protein comprises an engineered dTnpB protein; one or more base editing domains; and one or more Nuclear Localization Signals (NLS), and may optionally include additional domains. The base editor protein is very similar in composition and operating principle to a Dead-Form Cas protein-based base editor protein. Therefore, the base editor protein comprising the engineered dTnpB protein can be treated similarly to a base editor protein comprising a Dead-Form Cas protein, and the technologies associated with the base editor protein comprising the Dead-Form Cas protein can be applied as is.

[0385] Representative base editing domains include deaminases, which function to change nucleotides into other bases. The manner in which nucleotides are edited varies depending on which deaminase the base editor protein contains. For example, adenosine deaminase acts on adenine among nucleotides to change adenine to inosine, and consequently can induce a change to guanine. As another example, cytidine deaminase acts on cytosine among nucleotides to change it to uracil, and consequently can induce a change to thymine.

[0386] The above base editor protein includes one or more nuclear localization signals. This configuration is intended to deliver the base editor protein into the nucleus of a eukaryotic cell. If base editing is performed on a prokaryotic genome, the base editor protein does not necessarily need to include nuclear localization signals.

[0387] In some cases, the base editor protein may include an additional domain. The additional domain is not otherwise limited as long as it does not interfere with the function of the dTnpB protein and the deaminase. For example, the additional domain may be various amino acid linkers. As another example, when the deaminase is a cytosine deaminase, the additional domain may include a uracil glycosylase inhibitor. The composition of the base editor protein is very well known to researchers in the past and may include known compositions.

[0388] Engineered dTnpB protein

[0389] The above base editor protein includes an engineered dTnpB protein. This was described in the "Engineered dTnpB protein" paragraph.

[0390] Base Editing Domain #1 - Adenosine Deaminase

[0391] The above deaminase may be an adenosine deaminase. If the above base editor protein contains an adenosine deaminase, it may be referred to as an adenine base editor. The above adenosine deaminase can act on adenine among nucleobases to deaminate it and convert it into inosine. The above inosine can be converted into guanine during the nucleic acid base repair process, and consequently, A can be changed to G. The above adenosine deaminase is not limited to a single type that performs the above function, and known compositions may be utilized. For example, the above adenosine deaminase may be tRNA adenosine deaminase (TadA) or a variant of said TadA. Specifically, said TadA may be an enzyme derived from E. coli. More diverse examples are described in the 'Possible Examples of the Invention' section.

[0392] Base Editing Domain #2 - Cytidine Deaminase

[0393] The above deaminase may be cytidine deaminase. If the above base editor protein contains cytidine deaminase, it may be referred to as a cytosine base editor. The above cytidine deaminase can act on cytosine among nucleobases to deaminate it and convert it into uracil. The above uracil can be converted into thymine during the nucleic acid base repair process, and consequently, C can be changed to T. The above cytidine deaminase is not limited to a single type that performs the above function, and known compositions may be utilized. For example, the above cytidine deaminase may be APOBEC1, recombinant APOBEC1 (rAPOBEC1), or BE4max. More diverse examples are described in the 'Possible Examples of the Invention' section.

[0394] Nuclear Localization Signal (NLS)

[0395] The base editor protein may include one or more Nuclear Localization Signals (NLS). These NLS play an important role in delivering the base editor protein into the nucleus of a eukaryotic cell. If base editing is performed on a prokaryotic genome, the base editor protein does not necessarily need to include a Nuclear Localization Signal. The Nuclear Localization Signal is not limited to a single type that performs the above function, and known configurations may be utilized. For example, the Nuclear Localization Signal may be SV40. More diverse examples are described in the 'Possible Embodiments of the Invention' section.

[0396] Selective Configuration #2 - Uracyl Glycosylase Inhibitor

[0397] The above base editor protein may include one or more uracil glycosylase inhibitors (UGIs). The uracil glycosylase inhibitor is a domain that inhibits uracil glycosylase. Uracil glycosylase is an enzyme involved in a repair mechanism that recognizes uracil bases contained in DNA and removes them from the DNA. The uracil glycosylase inhibitor prevents uracil from being removed by uracil glycosylase when cytosine deaminase acts to convert cytosine among nucleobases into uracil. In other words, the uracil glycosylase inhibitor is an additional domain that helps the cytosine base editor protein operate more efficiently. The type of the uracil glycosylase inhibitor is not otherwise limited as long as it has the aforementioned function. For example, the uracil glycosylase inhibitor may be a protein comprising an amino acid sequence selected from SEQ ID NOs 19 to 20. As another example, the cytosine base editor protein may comprise one or more units of the uracil glycosylase inhibitor.

[0398] Base editing system including base editor protein and reRNA and various implementations thereof

[0399] Overview of Base Editing Systems and Their Various Implementations

[0400] This specification discloses a base editing system comprising the aforementioned base editor protein and reRNA. In practice, the object editing the nucleobase is a base editor complex formed by the combination of the base editor protein and reRNA described above. However, with the advancement of biotechnology, it is possible to deliver nucleic acids encoding proteins and RNA to cells and express each protein and RNA within the cell to form an RNA-protein complex. Therefore, the term "base editing system" as used in this specification encompasses 1) a base editing complex, 2) a vector capable of expressing each component of the base editing complex (base editing system expression vector), and 3) a base editing composition. The base editor protein was described in the paragraph "Base editor protein comprising engineered dTnpB protein." In addition, the above reRNA was described in the detailed section on "reRNA" in the paragraph "Chapter 1. Gene editing method using engineered ISDra2 TnpB protein." The base editing system 1) recognizes and binds to a target nucleic acid, and 2) changes a specific base of the target nucleic acid to another base. At this time, the manner in which a specific base is changed to another base varies depending on the type of base editing domain included in the base editor protein. The base editing system is implemented and utilized in various ways depending on the mode of use. For example, the base editing system can be implemented as a "base editing complex" in which a base editor protein and reRNA are bound. As another example, the base editing system can be implemented as a "base editing system expression vector" comprising a nucleic acid encoding a base editor protein and a nucleic acid encoding reRNA. As another example, the base editing system may be implemented as a composition comprising mRNA encoding a base editor protein and reRNA.

[0401] Base Editing Complex

[0402] This specification discloses a base editor complex comprising the base editor protein and reRNA. When the base editor complex recognizes and binds to a target nucleic acid, a base editing domain included in the base editor protein changes one or more specific bases of the target nucleic acid to other bases. Here, the principle by which the base editor complex recognizes and binds to the target nucleic acid is the same as the principle by which an "engineered TnpB complex" recognizes and binds to the target nucleic acid.

[0403] Base editing system expression vector

[0404] This specification discloses a base editing system expression vector. The vector comprises a nucleic acid encoding a base editor protein, a nucleic acid encoding reRNA, and components that enable the base editor protein and reRNA to be expressed within a cell. The expression vector is not limited in its composition as long as it can express the components and may further include known components.

[0405] Base editing composition

[0406] This specification discloses a composition comprising a base editor protein, or a nucleic acid encoding said base editor protein; and reRNA or a nucleic acid encoding said reRNA. The composition is referred to as a base editing composition. More specific configurations are disclosed in the [Numbered Examples] section.

[0407] Base Editing Method Using a Base Editing System #1 - Cell Genome Base Pair Editing

[0408] Overview of Cell Genome Base Pair Editing Methods

[0409] This specification discloses a gene editing method using a base editing system. Specifically, the method comprises the step of delivering the base editing system to a cell. The base editing system may be implemented in various ways as a base editing complex, an expression vector, a composition, or any combination thereof. The step of delivering the base editing system to a cell may be performed in an appropriate manner depending on the specific implementation of the base editing system. Specific implementations of the method are exemplified in the [Numbered Examples] section.

[0410] When the base editing system described above is delivered to a cell, the base editing complex comes into contact with the genome of the cell. Through this contact, the base editing complex binds to a predetermined target nucleic acid, and the base editing domain of the base editor protein functions to change (edit) one or more specific bases to other bases. For example, if the base editing domain is an adenosine deaminase, one or more AT base pairs of the target nucleic acid are edited to GC base pairs. As another example, if the base editing domain is a cytidine deaminase, one or more CG base pairs of the target nucleic acid are edited to TA base pairs. Here, the locations within the target nucleic acid of the one or more base pairs to be edited are structurally predictable. Therefore, the base editing system can be designed to change base pairs at intended locations to other base pairs. The specific base pair editing mode is described in more detail below.

[0411] The relationship between the correction window, target nucleic acid, and reRNA

[0412] The relationship between the target nucleic acid, the correction window, and the reRNA is schematically shown in Figure 27.

[0413] 1) The base editing system is designed so that when delivered to a cell, the base editor complex can come into contact with the target nucleic acid.

[0414] 2) The target nucleic acid is a double-stranded nucleic acid and includes a target strand and a non-target strand.

[0415] 3) The correction window is a double-stranded nucleic acid with a length of l-bp at a 'specific location' within the target nucleic acid.

[0416] 4) The portion of the correction window included in the non-target strand is referred to as the first correction window, and the portion included in the target strand is referred to as the second correction window.

[0417] 5) As a result of performing the base editing method, the first base of the first correction window is changed to the third base, and the second base of the corresponding second correction window is changed to the fourth base. For example, if the base editor protein of the base editing system contains adenosine deaminase, A of the first correction window is changed to G, and T of the corresponding second correction window is changed to C. As another example, if the base editor protein contains cytidine deaminase, C of the first correction window is changed to T, and G of the corresponding second correction window is changed to A.

[0418] 6) The above base editor complex is designed to satisfy the following conditions:

[0419] 7) The base editor protein recognizes the nucleic acid sequence containing the m-th base upstream from the 5'-terminal base of the first correction window to the (m+4)-th base as TAM. That is, there are a total of m nucleotides located between the 3'-terminal base of TAM and the 5'-terminal base of the first correction window.

[0420] 8) The guide domain of the reRNA can hybridize with the region containing the 1st to mth bases downstream from the 3'-terminal base of the second correction window, and with the second correction window, and optionally can also hybridize with a portion of the region located upstream of said region.

[0421] The specific location and specific length of the correction window, the range for recognizing TAM, and the range in which the guide domain is hybridized may vary slightly depending on the implementation of the base editing system. For example, if the base editor protein of the base editing system includes adenine deaminase, l may be an integer between 1 and 9, and m may be an integer between 2 and 10. As another example, if the base editor protein includes cytidine deaminase, l may be an integer between 1 and 10, and m may be an integer between 2 and 11. More specific examples are disclosed in the [Numbered Examples] section.

[0422] Bass Editing Method Using a Bass Editing System #2 - Multiplexing Possible

[0423] This specification discloses a multiplex nucleic acid base pair editing method using an engineered TnpB protein and multiple reRNAs. Specifically, the method comprises the steps of: 1) delivering a base editor protein or a nucleic acid encoding the same, and 2) delivering a nucleic acid encoding two or more types of reRNAs to a cell. Here, the nucleic acid encoding the reRNAs is a nucleic acid in which the nucleic acids encoding two or more types of reRNAs are linked without being separated by spacers, terminators, etc. Only one promoter is operatively linked to the nucleic acid encoding two or more types of reRNAs. The nucleic acids encoding two or more types of reRNAs 1) are independently cleaved during the transcription process without a special separator to mature into respective reRNAs, 2) each reRNA binds to the base editor protein to form a complex, and 3) each complex recognizes and cleaves the nucleic acid targeted by each reRNA. For example, the reRNAs may recognize the 5'-GAAC-3' motif in the scaffold to self-cleave and self-mature. By using the above method, only the base editor protein and one molecule of "nucleic acid encoding reRNA" can be delivered into the cell to induce editing of base pairs of multiple target nucleic acids. More specific details are disclosed in the [Numbered Examples] section.

[0424]

[0425] [Numbered Examples (Enumerated Embodiments)]

[0426] Hereinafter, various numbered embodiments (or modes) are described sequentially to aid in understanding the invention. However, the present invention should not be interpreted as being limited to the embodiments listed below, but should be understood to include all variations, equivalents, and additional combinations that a person skilled in the art can easily derive from this specification, the accompanying drawings, and the claims.

[0427] ISDra2 TnpB protein

[0428] Example 1, ISDra2 TnpB protein

[0429] ISDra2 TnpB protein of the IS200 / IS605 family of Deinococcus radiodurans,

[0430] Here, the above protein is also referred to as "Deinococcus radiodurans-derived ISDra2 TnpB protein", "ISDra2 TnpB protein", or "wild-type TnpB protein".

[0431] Example 2, TnpB protein sequence limitation

[0432] In the wild-type TnpB protein of Example 1,

[0433] The above wild-type TnpB protein contains the amino acid sequence of SEQ ID NO. 1.

[0434] Example 3, TnpB complex formation function

[0435] In any one of Examples 1 to 2, a wild-type TnpB protein selected,

[0436] The above wild-type TnpB protein can bind to reRNA to form a protein-RNA complex, and

[0437] The above protein-RNA complex can recognize a target nucleic acid, bind to the target nucleic acid, and cleave the target nucleic acid.

[0438] Reduced length ISDra2 TnpB protein

[0439] Example 4, length reduction

[0440] A shortened TnpB protein containing the following sequence:

[0441] Based on the amino acid sequence of SEQ ID NO. 1,

[0442] With an amino acid sequence determined by removing the K-th amino acid and all amino acids located downstream therefrom,

[0443] The above K is 374, 375, 376, 377, 378, 379, 380, 381, 382, ​​383, 384, 385, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, or 408.

[0444] Example 5, length reduction

[0445] A shortened TnpB protein containing the following sequence:

[0446] Based on the amino acid sequence of SEQ ID NO. 1,

[0447] A sequence in which 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, or 34 amino acids are removed consecutively from the C-terminal amino acid toward the N-terminal.

[0448] Example 6, sequence specific

[0449] A shortened TnpB protein comprising the amino acid sequences of SEQ ID NOs 445 to 479 and SEQ ID NOs 483 to 517, or composed of the amino acid sequences of SEQ ID NOs 445 to 479 and SEQ ID NOs 483 to 517.

[0450] A variant of ISDra2 TnpB protein with inhibited nucleic acid cleavage activity

[0451] Example 7, TnpB protein with inhibited nucleic acid cleavage activity

[0452] TnpB protein with inhibited nucleic acid cleavage activity,

[0453] Subsequently, the above protein is also referred to as "Dead-TnpB protein" or "dTnpB protein", and

[0454] The above dTnpB protein is a protein artificially modified from any one of the TnpB proteins selected from Examples 1 to 3 or Examples 4 to 6, and

[0455] Nucleic acid cleavage activity is inactivated, deactivated, removed, or suppressed.

[0456] Example 8, TnpB protein with inhibited nucleic acid cleavage activity

[0457] In the dTnpB protein of Example 7,

[0458] The above dTnpB protein comprises a sequence selected from the following:

[0459] A sequence specified by introducing a D191A variant into the amino acid sequence of SEQ ID NO. 1; and

[0460] A sequence specified by introducing a D191A variant into the amino acid sequences of SEQ ID NOs 445 to 479 and SEQ ID NOs 483 to 517.

[0461] Example 9, dTnpB sequence limitation

[0462] In any one of the dTnpB proteins selected from Examples 7 to 8,

[0463] The above dTnpB protein includes the amino acid sequence of SEQ ID NO. 8, or is composed of the amino acid sequence of SEQ ID NO. 8.

[0464] Engineered ISDra2 TnpB Protein #1 - Enhanced Nucleic Acid Interactions

[0465] Example 10, engineered TnpB protein

[0466] Engineered TnpB protein,

[0467] Hereinafter, the above-mentioned engineered TnpB protein is also referred to as a "modified TnpB protein" or "artificially engineered TnpB protein," and

[0468] The above-mentioned engineered TnpB protein is a protein engineered from any one of the TnpB proteins selected from Examples 1 to 3 or Examples 4 to 6.

[0469] Example 11, Specific substitution purpose

[0470] In any one of the engineered TnpB proteins selected in Example 10,

[0471] The above-mentioned engineered TnpB protein can bind to reRNA to form a protein-RNA complex, and

[0472] The above protein-RNA complex can bind to a target nucleic acid, and

[0473] The above-mentioned engineered TnpB protein has enhanced interaction with reRNA, target nucleic acid, or both compared to the above-mentioned wild-type TnpB protein.

[0474] Example 12, Meaning of Interaction

[0475] In Example 11, the interaction between the protein and the nucleic acid refers to the strength of the mutual electrical attraction, and

[0476] The above-mentioned engineered TnpB protein binds to reRNA, target nucleic acids, or both with a stronger electrical attraction compared to the above-mentioned wild-type TnpB protein.

[0477] Example 13, Substitution Site Specification #1

[0478] In any one of the engineered TnpB proteins selected from Examples 10 to 12,

[0479] The above-mentioned engineered TnpB protein comprises an amino acid sequence in which one or more amino acids selected from the amino acid sequence of SEQ ID NO. 1 are substituted with other amino acids:

[0480] 4th Asparagine; 8th Valine; 49th Glycine; 57th Serine; 64th Glutamine; 72nd Serine; 84th Lysine; 128th Glutamine; 140th Glycine; 146th Glycine; 148th Glutamine; 204th Valine; 228th Threonine; 234th Lysine; 237th Alanine; 239th Tyrosine; 243rd Lysine; 247th Alanine; 254th Valine; 255th Asparagine; 256th Lysine; 263rd Lysine; 266th Threonine; 281st Lysine; 310th Lysine; 339th Proline; and 350th Threonine,

[0481] The aforementioned target amino acid for substitution ("other amino acid" in the above description) is an amino acid that strongly interacts with nucleic acids.

[0482] Example 14, specification of the number of substitutions

[0483] In any one of Examples 10 to 13, the engineered TnpB protein

[0484] The above-mentioned engineered TnpB protein is a protein in which one or more and six or fewer amino acids are substituted in the amino acid sequence of SEQ ID NO. 1.

[0485] Example 15, Substituted Amino Acid Specific

[0486] In any one of the engineered TnpB proteins selected from Examples 13 to 14,

[0487] The above-mentioned target amino acid for substitution is arginine, lysine, or histidine, and

[0488] In the case where two or more amino acids in the amino acid sequences of SEQ ID NO. 1 or SEQ ID NOs 445 to 479 and SEQ ID NOs 483 to 517 are substituted with other amino acids,

[0489] The amino acid to be substituted is independently selected from arginine, lysine, or histidine, respectively.

[0490] Example 16, Substituted Amino Acid Specific

[0491] In the engineered TnpB protein of Example 15,

[0492] All of the above-mentioned substituted amino acids are arginine.

[0493] Example 17, variant specification

[0494] In any one of the engineered TnpB proteins selected from Examples 10 to 16,

[0495] The above-mentioned engineered TnpB protein comprises one or more variations selected from the following in the amino acid sequences of SEQ ID NO. 1 or SEQ ID NOs 445 to 479, and SEQ ID NOs 483 to 517:

[0496] N4X; V8X; G49X; S57X; Q64X; S72X; K84X; Q128X; G140X; G146X; Q148X; V204X; T228X; K234X; A247X; Y239X; K243X; V254X; N255X; K256X; K263X; T266X; K281X; K310X; P339X; And T350X;

[0497] Here, X is independently arginine, lysine, or histidine, respectively.

[0498] Example 18, specification of the number of variations

[0499] In the engineered TnpB protein of Example 17,

[0500] The above-mentioned engineered TnpB protein contains 1 to 6 variants.

[0501] Example 19, Arginine Specific

[0502] In any one of the engineered TnpB proteins selected from Examples 17 to 18,

[0503] All of the above X are arginine.

[0504] Example 20, sequence specific, X used

[0505] In any one of the engineered TnpB proteins selected from Examples 10 to 19,

[0506] The above-mentioned engineered TnpB protein comprises the amino acid sequence of SEQ ID NO. 4, or is composed of the amino acid sequence of SEQ ID NO. 4, and

[0507] The amino acid denoted by X in the amino acid sequence of SEQ ID NO. 4 means the following:

[0508] X at position 4 is arginine or asparagine; X at position 8 is arginine or valine; X at position 49 is arginine or glycine; X at position 57 is arginine or serine; X at position 64 is arginine or glutamine; X at position 72 is arginine or serine; X at position 84 is arginine or lysine; X at position 128 is arginine or glutamine; X at position 140 is arginine or glycine; X at position 146 is arginine or glycine; X at position 148 is arginine or glutamine; X at position 204 is arginine or valine; X at position 228 is arginine or threonine; X at position 234 is arginine or lysine; X at position 237 is arginine or alanine; X at position 239 is arginine or tyrosine; X at position 243 is arginine or lysine; X at position 247 is arginine or alanine; X at position 254 is arginine or valine; X at position 255 is arginine or asparagine; X at position 256 is arginine or lysine; X at position 263 is arginine or lysine; X at position 266 is arginine or threonine; X at position 281 is arginine or lysine; X at position 310 is arginine or lysine; X at position 339 is arginine or proline; and X at position 350 is arginine or threonine;

[0509] Here, the amino acid sequence of SEQ ID NO. 4 is different from the amino acid sequence of SEQ ID NO. 1.

[0510] Example 21, arginine number specification

[0511] In the engineered TnpB protein of Example 20,

[0512] Among the amino acids labeled X in the amino acid sequence of SEQ ID NO. 4 above, arginine is 1 or more and 6 or less.

[0513] Example 22, TAM recognition

[0514] In any one of the engineered TnpB proteins selected from Examples 10 to 21,

[0515] The above-mentioned engineered TnpB protein can recognize transposon-associated motifs (TAMs) contained in target nucleic acids.

[0516] Example 23, TAM specific

[0517] In the engineered TnpB protein of Example 22,

[0518] The above TAM consists of the following nucleic acid sequence:

[0519] 5'-TTGAT-3'; 5'-TNGAT-3'; and / or 5'-TTGVT-3'.

[0520] Example 24, specific sequence limitation

[0521] In any one of the engineered TnpB proteins selected from Examples 10 to 23,

[0522] The above-mentioned engineered TnpB protein comprises the amino acid sequence of SEQ ID NO. 2 or SEQ ID NO. 3, or is composed of the amino acid sequence of SEQ ID NO. 2 or SEQ ID NO. 3.

[0523] Engineered ISDra2 TnpB Protein #2 - Inhibition of Nucleic Acid Cleavage Activity

[0524] Example 25, engineered dTnpB protein

[0525] Engineered dTnpB protein,

[0526] Hereinafter, the above-mentioned engineered dTnpB protein is also referred to as a "modified dTnpB protein" or "artificially engineered dTnpB protein," and

[0527] The above-mentioned engineered dTnpB protein is a protein engineered from any one of the dTnpB proteins selected from Examples 7 to 9.

[0528] Example 26, Specific purpose of substitution

[0529] In the engineered dTnpB protein of Example 25,

[0530] The above-mentioned engineered dTnpB protein can bind to reRNA to form a protein-RNA complex, and

[0531] The above protein-RNA complex can bind to a target nucleic acid, and

[0532] The above-mentioned engineered dTnpB protein has enhanced interaction with reRNA, target nucleic acid, or both compared to the above-mentioned dTnpB protein.

[0533] Example 27, Meaning of Interaction

[0534] In Example 26, the interaction between the protein and the nucleic acid refers to the strength of the mutual electrical attraction, and

[0535] The engineered dTnpB protein binds to reRNA, target nucleic acid, or both with a stronger electrical attraction compared to the dTnpB protein.

[0536] Example 28, Substitution Position Specification #1

[0537] In any one of the engineered dTnpB proteins selected from Examples 25 to 27,

[0538] The above-mentioned engineered dTnpB protein is a protein in which one or more amino acids selected from the following in the amino acid sequence of SEQ ID NO. 8 are substituted with other amino acids:

[0539] 4th Asparagine; 8th Valine; 49th Glycine; 57th Serine; 64th Glutamine; 72nd Serine; 84th Lysine; 128th Glutamine; 140th Glycine; 146th Glycine; 148th Glutamine; 204th Valine; 228th Threonine; 234th Lysine; 237th Alanine; 239th Tyrosine; 243rd Lysine; 247th Alanine; 254th Valine; 255th Asparagine; 256th Lysine; 263rd Lysine; 266th Threonine; 281st Lysine; 310th Lysine; 339th Proline; and 350th Threonine,

[0540] The aforementioned target amino acid for substitution ("other amino acid" in the above description) is an amino acid that strongly interacts with nucleic acids.

[0541] Example 29, specifying the number of substitutions

[0542] In any one of Examples 25 to 28, the engineered dTnpB protein

[0543] The above-mentioned engineered dTnpB protein is a protein in which one or more and six or fewer amino acids are substituted in the amino acid sequence of SEQ ID NO. 8.

[0544] Example 30, Substituted Amino Acid Specific

[0545] In any one of the engineered dTnpB proteins selected from Examples 28 to 29,

[0546] The above-mentioned target amino acid for substitution is arginine, lysine, or histidine, and

[0547] In the case where two or more amino acids in the amino acid sequence of SEQ ID NO. 8 are substituted with other amino acids,

[0548] The amino acid to be substituted is independently selected from arginine, lysine, or histidine, respectively.

[0549] Example 31, Substituted Amino Acid Specific

[0550] In the engineered dTnpB protein of Example 30,

[0551] All of the above-mentioned substituted amino acids are arginine.

[0552] Example 32, variant specific

[0553] In any one of the engineered dTnpB proteins selected from Examples 25 to 31,

[0554] The above-mentioned engineered dTnpB protein comprises one or more variations selected from the following in the amino acid sequence of SEQ ID NO. 8:

[0555] N4X; V8X; G49X; S57X; Q64X; S72X; K84X; Q128X; G140X; G146X; Q148X; V204X; T228X; K234X; A237X; Y239X; K243X; A247X; V254X; N255X; K256X; K263X; T266X; K281X; K310X; P339X; And T350X;

[0556] Here, X is independently arginine, lysine, or histidine, respectively.

[0557] Example 33, specification of the number of variations

[0558] In the engineered dTnpB protein of Example 32,

[0559] The above-mentioned engineered dTnpB protein contains 1 to 6 variants.

[0560] Example 34, Arginine Specific

[0561] In any one of the engineered TnpB proteins selected from Examples 32 to 33,

[0562] All of the above X are arginine.

[0563] Example 35, sequence specific, use of X

[0564] In any one of the engineered dTnpB proteins selected from Examples 25 to 34,

[0565] The above-mentioned engineered dTnpB protein comprises the amino acid sequence of SEQ ID NO. 11, or is composed of the amino acid sequence of SEQ ID NO. 11, and

[0566] The amino acid denoted by X in the amino acid sequence of SEQ ID NO. 11 means the following:

[0567] X at position 4 is arginine or asparagine; X at position 8 is arginine or valine; X at position 49 is arginine or glycine; X at position 57 is arginine or serine; X at position 64 is arginine or glutamine; X at position 72 is arginine or serine; X at position 84 is arginine or lysine; X at position 128 is arginine or glutamine; X at position 140 is arginine or glycine; X at position 146 is arginine or glycine; X at position 148 is arginine or glutamine; X at position 204 is arginine or valine; X at position 228 is arginine or threonine; X at position 234 is arginine or lysine; X at position 237 is arginine or alanine; X at position 239 is arginine or tyrosine; X at position 243 is arginine or lysine; X at position 247 is arginine or alanine; X at position 254 is arginine or valine; X at position 255 is arginine or asparagine; X at position 256 is arginine or lysine; X at position 263 is arginine or lysine; X at position 266 is arginine or threonine; X at position 281 is arginine or lysine; X at position 310 is arginine or lysine; X at position 339 is arginine or proline; and X at position 350 is arginine or threonine;

[0568] Here, the amino acid sequence of SEQ ID NO. 11 is different from the amino acid sequence of SEQ ID NO. 1.

[0569] Example 36, arginine number specification

[0570] In the engineered dTnpB protein of Example 35,

[0571] Among the amino acids labeled X in the amino acid sequence of SEQ ID NO. 11 above, arginine is 1 or more and 6 or less.

[0572] Example 37, specific sequence limitation

[0573] In any one of the engineered dTnpB proteins selected from Examples 25 to 36,

[0574] The above-mentioned engineered dTnpB protein comprises the amino acid sequence of SEQ ID NO. 9 or SEQ ID NO. 10, or is composed of the amino acid sequence of SEQ ID NO. 9 or SEQ ID NO. 10.

[0575] Nuclear location signal

[0576] Example 38, nuclear location signal

[0577] Nuclear Localization Signal (NLS).

[0578] Example 39, nuclear location signal, sequence

[0579] In the nuclear location signal of Example 38,

[0580] The above nuclear location signal comprises one or more amino acid sequences selected from the group consisting of SEQ ID NOs 24 to 52, or

[0581] Any combination of one or more amino acid sequences selected from the group consisting of SEQ ID NOs 24 to 52.

[0582] Example 40, nuclear location signal, meaning of combination

[0583] In the nuclear location signal of Example 39,

[0584] "Any combination of one or more amino acid sequences" means that one or more amino acid sequences are directly linked or linked through an appropriate amino acid linker.

[0585] Example 41, amino acid linker limited

[0586] In the nuclear location signal of Example 40,

[0587] The above amino acid linkers each independently comprise an amino acid sequence selected from the group consisting of SEQ ID NOs 53 to 73.

[0588] TnpB protein for use in eukaryotic cells

[0589] Example 42, including nuclear location signal

[0590] TnpB proteins, including the following:

[0591] Any one wild-type TnpB protein selected from Examples 1 to 3, any one engineered TnpB protein selected from Examples 10 to 24, any one dTnpB protein selected from Examples 7 to 9, or any one engineered dTnpB protein selected from Examples 25 to 37; and

[0592] One or more nuclear localization signals,

[0593] Here, the nuclear location signal is any one of the nuclear location signals selected from Examples 38 to 41.

[0594] Example 43, nuclear location signal location specification

[0595] In the TnpB protein of Example 42,

[0596] Each nuclear localization signal contained in the above TnpB protein is located at the N-terminus or C-terminus.

[0597] Example 44, TnpB protein structure containing NLS

[0598] TnpB proteins, including the following:

[0599] [NLS1]-[TnpB protein]-[NLS2]

[0600] Here, the NLS1 is a first nuclear location signal or is non-existent,

[0601] The above TnpB protein is any one wild-type TnpB protein selected from Examples 1 to 3, any one engineered TnpB protein selected from Examples 10 to 24, any one dTnpB protein selected from Examples 7 to 9, or any one engineered dTnpB protein selected from Examples 25 to 37, and

[0602] The above NLS2 is a second nuclear location signal, or is non-existent,

[0603] The first nuclear position signal and the second nuclear position signal are each independently any one of the nuclear position signals selected from Examples 38 to 41, and

[0604] At least one of the first nuclear location signal or the second nuclear location signal exists.

[0605] Example 45, including TnpB

[0606] In any one of the TnpB proteins selected from Examples 42 to 44,

[0607] The above TnpB protein comprises any one wild-type TnpB protein selected from Examples 1 to 3, or any one engineered TnpB protein selected from Examples 10 to 24.

[0608] Example 46, including dTnpB

[0609] In any one of the TnpB proteins selected from Examples 42 to 44,

[0610] The above TnpB protein comprises any one of the dTnpB proteins selected from Examples 7 to 9, or any one of the engineered dTnpB proteins selected from Examples 25 to 37, and is hereinafter referred to as the dTnpB protein.

[0611] reRNA and engineered reRNA

[0612] Example 47, reRNA functioning with TnpB protein

[0613] reRNA,

[0614] Here, the above reRNA may form a protein-RNA complex by binding to any one wild-type TnpB protein selected from Examples 1 to 3, any one engineered TnpB protein selected from Examples 10 to 24, any one dTnpB protein selected from Examples 7 to 9, any one engineered dTnpB protein selected from Examples 25 to 37, or any one TnpB protein selected from Examples 42 to 44.

[0615] Example 48, reRNA domain separation

[0616] In the reRNA of Example 47,

[0617] The above reRNA includes a guide domain and a scaffold, and

[0618] Here, the above guide domain is also referred to as a spacer.

[0619] Example 49, reRNA structure specification

[0620] In the reRNA of Example 48,

[0621] The 3' end of the scaffold is connected to the 5' end of the guide domain.

[0622] Example 50, including additional stabilization domain

[0623] In any one of the reRNAs selected from Examples 47 to 49,

[0624] The above reRNA further includes one or more stabilization domains.

[0625] Example 51, stabilization domain structure specification

[0626] In the reRNA of Example 50,

[0627] Each of the above stabilization domains is connected to the 5' end of the scaffold or the 3' end of the guide domain.

[0628] Example 52, Stabilization Domain Linkage Linker

[0629] In the reRNA of Example 51,

[0630] One or more of the above stabilization domains are connected to the 5' end of the scaffold via a nucleic acid linker, or to the 3' end of the guide domain via a nucleic acid linker.

[0631] Example 53, linker limited

[0632] In the reRNA of Example 52,

[0633] The above nucleic acid linker includes any one of the nucleic acid linkers selected from SEQ ID NO. 23, or is composed of any one of the nucleic acid linkers selected from SEQ ID NO. 23.

[0634] Example 54, stabilization domain specific

[0635] In either Example 51 or Example 52, the reRNA selected

[0636] The above stabilization domains are each independently evopreQ1 or mpknot.

[0637] Example 55, stabilization domain sequence identification

[0638] In the reRNA of Example 54,

[0639] Each of the above stabilization domains independently includes the nucleic acid sequence of SEQ ID NO. 21 or SEQ ID NO. 22, or is composed of the nucleic acid sequence of SEQ ID NO. 21 or SEQ ID NO. 22.

[0640] Example 56, specification of scaffold function

[0641] In any one of the reRNAs selected from Examples 47 to 55,

[0642] The scaffold of the reRNA above causes the reRNA to bind to any one wild-type TnpB protein selected from Examples 1 to 3, any one engineered TnpB protein selected from Examples 10 to 24, any one dTnpB protein selected from Examples 7 to 9, or any one engineered dTnpB protein selected from Examples 25 to 37 to form a protein-RNA complex.

[0643] Example 57, scaffold sequence limitation

[0644] In any one of the reRNAs selected from Examples 47 to 56,

[0645] The scaffold of the above reRNA comprises the nucleic acid sequence of SEQ ID NO. 5, SEQ ID NO. 6, or SEQ ID NO. 7, or

[0646] Composed of the nucleic acid sequence of SEQ ID NO. 5, SEQ ID NO. 6, or SEQ ID NO. 7.

[0647] Example 58, guide domain function

[0648] In any one of the reRNAs selected from Examples 47 to 57,

[0649] The above guide domain is configured to target a target nucleic acid.

[0650] Example 59, Meaning of Targeting Target Nucleic Acid

[0651] In the reRNA of Example 58,

[0652] The above target nucleic acid is a double-stranded nucleic acid comprising a target strand and a non-target strand, and

[0653] The above-mentioned non-target strand includes a protospacer and a transposon-associated motif (TAM), and

[0654] The above target strand includes a target site that binds complementarily to the above protospacer, and

[0655] The fact that the above guide domain targets the target nucleic acid means one of the following:

[0656] The nucleic acid sequence of the guide domain and the sequence of the protospacer are identical, matched, or equivalent;

[0657] The nucleic acid sequence of the guide domain and the sequence of the protospacer are identical, matched, or equivalent except for one, two, three, four, or five bases;

[0658] The nucleic acid sequence of the above guide domain and the sequence of the above target site are complementary;

[0659] The nucleic acid sequence of the guide domain and the sequence of the target site are complementary except for 1, 2, 3, 4, or 5 bases; or

[0660] The above guide domain is configured to complementarily bind or hybridize with the target site of the above target strand.

[0661] Example 60, guide domain length

[0662] In any one of the reRNAs selected from Examples 47 to 59,

[0663] The length of the above guide domain is 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, or 30 nt, or satisfies a range consisting of the two values ​​mentioned above, for example, a length of 18 nt or more and 24 nt or less.

[0664] Example 61, wRNA structure

[0665] As reRNA, it includes the following:

[0666] 5'-[1st Stabilization Domain]-[1st Linker]-[Scaffold]-[Guide Domain]-[2nd Linker]-[2nd Stabilization Domain]-3'

[0667] Here, the first stabilization domain is absent or comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22, and

[0668] The first linker above is absent or comprises a nucleic acid sequence selected from the group consisting of SEQ ID NO. 23, and

[0669] The above scaffold comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 5 to 7, and

[0670] The above guide domain is configured to target a target nucleic acid, and

[0671] The above second linker is absent or comprises a nucleic acid sequence selected from the group consisting of SEQ ID NO. 23, and

[0672] The second stabilization domain is absent or comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22.

[0673] TnpB-RNA complex

[0674] Example 62, RNA-protein complex

[0675] TnpB complex containing the following:

[0676] Any one wild-type TnpB protein selected from Examples 1 to 3, any one engineered TnpB protein selected from Examples 10 to 24, any one dTnpB protein selected from Examples 7 to 9, any one engineered dTnpB protein selected from Examples 25 to 37, or any one TnpB protein selected from Examples 42 to 44 (hereinafter collectively referred to as TnpB protein); and

[0677] Any one reRNA selected from Examples 47 to 61;

[0678] Here, the TnpB protein and the scaffold of the reRNA interact with each other to form a complex, and

[0679] The guide domain of the above reRNA targets the target nucleic acid, and

[0680] The above TnpB complex recognizes the TAM sequence contained in the target nucleic acid and can bind to the target nucleic acid through the guide domain of the reRNA, and

[0681] If the TnpB protein is any one of the wild-type TnpB proteins selected from Examples 1 to 3, any one of the engineered TnpB proteins selected from Examples 10 to 24, or any one of the TnpB proteins selected from Examples 42 to 44, the TnpB complex can cleave the target nucleic acid.

[0682] Example 63, preferred example

[0683] In the TnpB complex of Example 62,

[0684] The above TnpB protein comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 2 to 3, and

[0685] The above reRNA contains the following structure:

[0686] 5'-[1st Stabilization Domain]-[1st Linker]-[Scaffold]-[Guide Domain]-[2nd Linker]-[2nd Stabilization Domain]-3'

[0687] Here, the first stabilization domain is absent or comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22, and

[0688] The first linker above is absent or comprises a nucleic acid sequence selected from the group consisting of SEQ ID NO. 23, and

[0689] The above scaffold comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 5 to 7, and

[0690] The above guide domain targets the target nucleic acid, and

[0691] The above second linker is absent or comprises a nucleic acid sequence selected from the group consisting of SEQ ID NO. 23, and

[0692] The second stabilization domain is absent or comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22.

[0693] A vector configured to express TnpB protein and / or reRNA

[0694] Example 64, component expression vector

[0695] TnpB system component expression vectors including the following:

[0696] nucleic acid encoding the TnpB protein,

[0697] Here, the TnpB protein is any one wild-type TnpB protein selected from Examples 1 to 3, any one engineered TnpB protein selected from Examples 10 to 24, any one dTnpB protein selected from Examples 7 to 9, any one engineered dTnpB protein selected from Examples 25 to 37, or any one TnpB protein selected from Examples 42 to 44; and

[0698] A nucleic acid encoding any one of the reRNAs selected from Examples 47 to 61;

[0699] Here, the TnpB protein and the scaffold of the reRNA interact with each other to form a complex, and

[0700] The guide domain of the above reRNA targets the target nucleic acid, and

[0701] The above TnpB complex recognizes the TAM sequence contained in the target nucleic acid and can bind to the target nucleic acid through the guide domain of the reRNA, and

[0702] If the TnpB protein is any one of the wild-type TnpB proteins selected from Examples 1 to 3, any one of the engineered TnpB proteins selected from Examples 10 to 24, or any one of the TnpB proteins selected from Examples 42 to 44, the TnpB complex can cleave the target nucleic acid.

[0703] Example 65, multiplexing vector

[0704] TnpB system component expression vector comprising the following (hereinafter also referred to as TnpB system multiplexing expression vector):

[0705] nucleic acid encoding the TnpB protein,

[0706] Here, the TnpB protein is any one wild-type TnpB protein selected from Examples 1 to 3, any one engineered TnpB protein selected from Examples 10 to 24, any one dTnpB protein selected from Examples 7 to 9, any one engineered dTnpB protein selected from Examples 25 to 37, or any one TnpB protein selected from Examples 42 to 44; and

[0707] nucleic acid encoding the first reRNA,

[0708] Here, the first reRNA is any one selected from Examples 47 to 61; and

[0709] nucleic acid encoding 2nd reRNA,

[0710] Here, the second reRNA is any one selected from Examples 47 to 61;

[0711] Here, the 3' end of the nucleic acid encoding the first reRNA and the 5' end of the nucleic acid encoding the second reRNA are directly connected without a separate spacer, and

[0712] The guide domain of the first reRNA targets a first target nucleic acid, and

[0713] The guide domain of the second reRNA above targets a second target nucleic acid, and

[0714] The above TnpB complex recognizes the TAM sequence contained in the first target nucleic acid and the second target nucleic acid, and

[0715] The above TnpB complex can bind to a first target nucleic acid through the guide domain of the first reRNA, and

[0716] The above TnpB complex can bind to a second target nucleic acid through the guide domain of the second reRNA, and

[0717] If the TnpB protein is any one of the wild-type TnpB proteins selected from Examples 1 to 3, any one of the engineered TnpB proteins selected from Examples 10 to 24, or any one of the TnpB proteins selected from Examples 42 to 44, the TnpB complex can cleave the target nucleic acid.

[0718] Example 66, self-cutting motif

[0719] In the TnpB system multiplexing expression vector of Example 65,

[0720] The first reRNA is operably linked to the first promoter, and

[0721] The nucleic acid encoding a scaffold among the first reRNA and the nucleic acid encoding a scaffold among the second reRNA include a self-cleavage motif, and

[0722] The above self-cleavage motif is a portion that is self-cleaved after the nucleic acid encoding the first reRNA and the nucleic acid encoding the second reRNA are transcribed when the first promoter is activated.

[0723] Self-maturation of the first reRNA and the second reRNA occurs through the above self-cleavage.

[0724] Example 67, identification of self-cleavage motif sequence

[0725] In the TnpB system multiplexing expression vector of Example 66,

[0726] The above self-cleavage motif includes a nucleic acid sequence of 5'-GAAC-3' or consists of a nucleic acid sequence of 5'-GAAC-3'.

[0727] Example 68, including additional components

[0728] In any one of Examples 64 to 67, a TnpB system component expression vector selected,

[0729] The above TnpB system component expression vector further comprises one or more selected from the following:

[0730] Promoter; enhancer; intron; polyadenylation signal; Kozak consensus sequence; Internal Ribosome Entry Site (IRES); splice acceptor; 2A sequence; and replication origin.

[0731] Example 69, promoter specific

[0732] In the TnpB system component expression vector of Example 68,

[0733] The above promoter is one or more of the following selected:

[0734] SV40 early promoter; mouse mammary tumor virus long terminal repeat (LTR) promoter; adenovirus major late promoter (Ad MLP); herpes simplex virus (HSV) promoter; cytomegalovirus (CMV) promoters such as the CMV immediate early promoter region (CMVIE); rous sarcoma virus (RSV) promoter; human U6 small nuclear promoter (U6) (Miyagishi et al.; Nature Biotechnology 20; 497-500 (2002)); enhanced U6 promoter (e.g., Xia et al.; Nucleic Acids Res. 2003 Sep 1;31(17)); human H1 promoter (H1); and 7SK.

[0735] Example 70, identification of replication origin

[0736] In the TnpB system component expression vector of Example 68 or Example 69,

[0737] The above replication origin is one or more of the following selected:

[0738] f1 replication origin; SV40 replication origin; pMB1 replication origin; adeno replication origin; AAV replication origin; and BBV replication origin.

[0739] Example 71, Vector Type

[0740] In any one of the selected TnpB system component expression vectors among Examples 64 to 70,

[0741] The above vector is a viral vector, a non-viral vector, or any combination thereof.

[0742] Example 72, virus vector

[0743] In the TnpB system component expression vector of Example 71,

[0744] The above vector is a selected virus vector from the following:

[0745] Retrovirus; lentivirus; adenovirus; adeno-associated virus; vacciniavirus; poxvirus; herpes simplex virus; and any combination thereof.

[0746] Example 73, non-viral vector

[0747] In the TnpB system component expression vector of Example 71,

[0748] The above vector is a non-viral vector selected from the following:

[0749] Plasmid; phage; naked DNA; DNA complex; PCR amplicon; mRNA; and any combination thereof.

[0750] Example 74, Plasmid Specific

[0751] In the TnpB system component expression vector of Example 73,

[0752] The above plasmid is any one of the following selected:

[0753] pcDNA series; pS456; p326; pACYC177; ColE1; pKT230; pME290; pBR322; pUC8 / 9; pUC6; pBD9; pHC79; pIJ61; pLAFR1; pHV14; pGEX series; pET series; pUC19 and any combination thereof.

[0754] Example 75, phage specific

[0755] In the TnpB system component expression vector of Example 73,

[0756] The above waste is any one of the following selected:

[0757] λgt4λB; λ-Charon; λΔz1; M13; and any combination thereof.

[0758] Example 76, single vector

[0759] In any one of Examples 64 to 75, a TnpB system component expression vector selected,

[0760] The above vector is a single vector composed of one molecule.

[0761] Example 77, DNA and RNA specific

[0762] In any one of Examples 64 to 76, a TnpB system component expression vector,

[0763] The nucleic acid encoding the TnpB protein and the nucleic acid encoding the reRNA are each independently DNA or RNA, or

[0764] The nucleic acid encoding the TnpB protein, the nucleic acid encoding the first reRNA, and the nucleic acid encoding the second reRNA are each independently DNA or RNA.

[0765] TnpB composition

[0766] Example 78, TnpB composition

[0767] TnpB composition comprising the following:

[0768] TnpB protein, or nucleic acid encoding the said TnpB protein,

[0769] Here, the TnpB protein is any one wild-type TnpB protein selected from Examples 1 to 3, any one engineered TnpB protein selected from Examples 10 to 24, any one dTnpB protein selected from Examples 7 to 9, any one engineered dTnpB protein selected from Examples 25 to 37, or any one TnpB protein selected from Examples 42 to 44; and

[0770] reRNA, or nucleic acid encoding the reRNA,

[0771] Here, the reRNA is any one of the reRNAs selected from Examples 47 to 61.

[0772] Here, the TnpB protein and the scaffold of the reRNA interact with each other to form a complex, and

[0773] The guide domain of the above reRNA targets the target nucleic acid, and

[0774] The above TnpB complex recognizes the TAM sequence contained in the target nucleic acid and can bind to the target nucleic acid through the guide domain of the reRNA, and

[0775] If the TnpB protein is any one of the wild-type TnpB proteins selected from Examples 1 to 3, any one of the engineered TnpB proteins selected from Examples 10 to 24, or any one of the TnpB proteins selected from Examples 42 to 44, the TnpB complex can cleave the target nucleic acid.

[0776] Example 79, mRNA delivery form

[0777] In the TnpB composition of Example 78,

[0778] The above TnpB composition comprises mRNA encoding the TnpB protein and the reRNA.

[0779] Gene editing method

[0780] Example 80, Target cell genome editing method

[0781] A method for editing the genome of a target cell, comprising the following:

[0782] A process of delivering the TnpB system to the target cells;

[0783] Here, the above TnpB system is selected from the following:

[0784] Any one of the TnpB complexes selected from Examples 62 to 63;

[0785] Any one of the TnpB system component expression vectors selected from Examples 64 to 77; or

[0786] Any one of the selected TnpB compositions from Example 78;

[0787] Here, as a result of the TnpB system being delivered to the target cell, a TnpB complex is delivered within the target cell or the formation of a TnpB complex is induced, and

[0788] The guide domain of the reRNA of the above TnpB complex targets the target nucleic acid contained in the target cell genome, and

[0789] The above TnpB complex recognizes TAM contained in the target nucleic acid and binds to the target nucleic acid through the guide domain of the reRNA, and

[0790] The target nucleic acid is edited by the above TnpB complex.

[0791] Example 81, eukaryotic cell genome editing method

[0792] In the method for editing the genome of a target cell of Example 80,

[0793] The above target cell is a eukaryotic cell, and

[0794] The TnpB complex delivered to the target cell or induced to form within the target cell comprises the TnpB protein of Example 45, and

[0795] The above TnpB complex is delivered to the nucleus of the target cell and edits the genome of the target cell.

[0796] Example 82, multiplex editing method

[0797] A method for editing the genome of a target cell, comprising the following:

[0798] A process of delivering the TnpB system multiplexing expression vector of Example 65, or the TnpB system multiplexing expression vector of Examples 68 to 77 to the TnpB system multiplexing expression vector of Example 65, to the target cells;

[0799] Here, as a result of the TnpB system being delivered to a target cell, the TnpB protein, the first reRNA, and the second reRNA of the TnpB system multiplexing expression vector are expressed within the target cell, and

[0800] The formation of the first TnpB complex and the second TnpB complex is induced within the above target cell, and

[0801] The first TnpB complex comprises a TnpB protein and a first reRNA, and

[0802] The above-mentioned second TnpB complex comprises a TnpB protein and a second reRNA, and

[0803] The guide domain of the first reRNA of the first TnpB complex targets the first target nucleic acid included in the target cell genome, and

[0804] The guide domain of the second reRNA of the second TnpB complex targets the second target nucleic acid included in the target cell genome, and

[0805] The first TnpB complex recognizes TAM contained in the target nucleic acid and binds to the first target nucleic acid through the guide domain of the first reRNA, and

[0806] The second TnpB complex recognizes TAM contained in the target nucleic acid and binds to the second target nucleic acid through the guide domain of the second reRNA, and

[0807] The first target nucleic acid is edited by the first TnpB complex, and

[0808] The second target nucleic acid is edited by the second TnpB complex.

[0809] Example 83, multiplex editing method, eukaryotic cell

[0810] In the method for editing the genome of a target cell of Example 82,

[0811] The above target cell is a eukaryotic cell, and

[0812] The first TnpB complex and the second TnpB complex, whose formation is induced within the target cells, comprise the TnpB protein of Example 45, and

[0813] The first TnpB complex and the second TnpB complex are delivered to the nucleus of the target cell to edit the genome of the target cell.

[0814] Example 84, alternative expression of transfer

[0815] A method for editing the genome of any one of the selected target cells in Examples 80 to 83,

[0816] The above expression "Deliver to target cells" is replaced with the following expression:

[0817] Introduce to target cells; administer to target cells; inject to target cells; transfect to target cells; transduct to target cells; or any combination of the above expressions.

[0818] Base Editing Domain #1 - Adenosine Deaminase

[0819] Example 85, Adenosine Deaminase

[0820] Adenosine Deaminase

[0821] Example 86, TadA and its variants

[0822] In the adenosine deaminase of Example 85,

[0823] The above adenosine deaminase is tRNA adenosine deaminase (TadA) derived from E. coli or a variant of TadA.

[0824] Example 87, specific sequence

[0825] In any one of the selected adenosine deaminases among Examples 85 to 86,

[0826] The above adenosine deaminase comprises an amino acid sequence selected from the group consisting of SEQ ID NOs 13 to 15.

[0827] Base Editing Domain #2 - Cytosine Deaminase

[0828] Example 88, cytidine deaminase

[0829] Cytidine deaminase.

[0830] Example 89, Cytidine Deaminase Type

[0831] In the cytidine deaminase of Example 88,

[0832] The above cytidine deaminase is activation-induced cytidine deaminase (AID), APOBEC3G, APOBEC1, rAPOBEC1, APOBEC3A, APOBEC3B, CDA, Anc689 APOBEC, or PmCDA1.

[0833] Example 90, specific sequence

[0834] In any one of Examples 88 to 89, the cytidine deaminase

[0835] The above cytidine deaminase comprises an amino acid sequence selected from the group consisting of SEQ ID NOs 16 to 18.

[0836] Uracil glycosylase inhibitors

[0837] Example 91, Uracyl glycosylase inhibitor

[0838] Uracil Glycosylase Inhibitor (UGI).

[0839] Example 92, UGI sequence

[0840] In the uracil glycosylase inhibitor of Example 91,

[0841] The above uracil glycosylase inhibitor comprises an amino acid sequence selected from the group consisting of SEQ ID NOs 19 to 20.

[0842] amino acid linker

[0843] Example 93, amino acid linker

[0844] Amino acid linker.

[0845] Example 94, sequence specific

[0846] In the amino acid linker of Example 93,

[0847] The above amino acid linker comprises an amino acid sequence selected from SEQ ID NOs 53 to 73.

[0848] Base editor protein #1 containing ISDra2 TnpB protein or its variant - Adenine base editor protein

[0849] Example 95, Adenine-based editor protein

[0850] Adenine-based editor proteins, including the following:

[0851] Any one dTnpB protein selected from Examples 7 to 9, or any one engineered dTnpB protein selected from Examples 25 to 37;

[0852] One or more adenosine deaminases,

[0853] Here, the adenosine deaminase is each independently selected from Examples 85 to 87;

[0854] Optionally, one or more nuclear location signals,

[0855] Here, the nuclear location signal is each independently selected from Examples 38 to 41.

[0856] Example 96, fusion protein

[0857] In the adenine-based editor protein of Example 95,

[0858] Each component included in the above adenine-based editor protein is directly linked or linked through any one of the amino acid linkers selected from Examples 93 to 94.

[0859] Example 97, fusion protein structure specification

[0860] In any one of the selected adenine-based editor proteins among Examples 95 to 96,

[0861] The above adenine base editor protein includes the following structure:

[0862] [1st nuclear localization signal]-[1st amino acid linker]-[adenosine deaminase]-[2nd amino acid linker]-[dTnpB protein]-[3rd amino acid linker]-[2nd nuclear localization signal]; or

[0863] [1st nuclear localization signal]-[1st amino acid linker]-[dTnpB protein]-[2nd amino acid linker]-[adenosine deaminase]-[3rd amino acid linker]-[2nd nuclear localization signal];

[0864] Here, the first nuclear location signal and the second nuclear location signal are each independently any one of the selected nuclear location signals among Examples 38 to 41, or are non-existent, and

[0865] The first amino acid linker, the second amino acid linker, and the third amino acid linker are each independently any one of the amino acid linkers selected from Examples 93 to 94, or are absent, and

[0866] The above adenosine deaminase is any one of the adenosine deaminases selected from Examples 85 to 87, and

[0867] The above dTnpB protein is any one of the dTnpB proteins selected from Examples 7 to 9, or any one of the engineered dTnpB proteins selected from Examples 25 to 37.

[0868] Example 98, identification of adenine base editor protein sequence

[0869] In any one of the selected adenine-based editor proteins from Examples 95 to 97,

[0870] The above adenine base editor protein comprises an amino acid sequence selected from the group consisting of SEQ ID NOs 74 to 76.

[0871] Base editor protein #2 containing ISDra2 TnpB protein or its variant - Cytosine base editor protein

[0872] Example 99, cytosine-based editor protein

[0873] Cytosine-based editor proteins, including:

[0874] Any one dTnpB protein selected from Examples 7 to 9, or any one engineered dTnpB protein selected from Examples 25 to 37;

[0875] One or more cytidine deaminases,

[0876] Here, the cytidine deaminase is each independently selected from Examples 88 to 90;

[0877] Optionally, one or more nuclear location signals,

[0878] Here, the nuclear location signal is each independently selected from Examples 38 to 41.

[0879] Example 100, fusion protein

[0880] In the cytosine-based editor protein of Example 99,

[0881] Each component included in the above cytosine-based editor protein is directly connected or connected through any one of the amino acid linkers selected from Examples 93 to 94.

[0882] Example 101, fusion protein structure specification

[0883] In any one of the cytosine-based editor proteins selected from Examples 99 to 100,

[0884] The above cytosine-based editor protein comprises the following structure:

[0885] [1st nuclear localization signal] - [1st linker] - [cytidine deaminase] - [2nd linker] - [dTnpB protein] - [3rd linker] - [UGI] - [4th linker] - [2nd nuclear localization signal]; or

[0886] [1st nuclear localization signal] - [1st linker] - [UGI] - [2nd linker] - [cytidine deaminase] - [3rd linker] - [dTnpB protein] - [4th linker] - [2nd nuclear localization signal];

[0887] Here, the first nuclear location signal and the second nuclear location signal are each independently any one of the selected nuclear location signals among Examples 38 to 41, or are non-existent, and

[0888] The first amino acid linker, the second amino acid linker, the third amino acid linker, and the fourth amino acid linker are each independently any one of the amino acid linkers selected from Examples 93 to 94, or are absent, and

[0889] The above cytidine deaminase is any one of the cytidine deaminases selected from Examples 88 to 90, and

[0890] The above dTnpB protein is any one of the dTnpB proteins selected from Examples 7 to 9, or any one of the engineered dTnpB proteins selected from Examples 25 to 37.

[0891] Example 102, identification of cytosine base editor protein sequence

[0892] In any one of the cytosine-based editor proteins selected from Examples 99 to 101,

[0893] The above cytosine base editor protein comprises an amino acid sequence selected from the group consisting of SEQ ID NOs 77 to 79 and SEQ ID NOs 443 to 444.

[0894] Base Editing Complex

[0895] Example 103, Adenine-based editing complex

[0896] Adenine base editing complexes including the following:

[0897] Any one of the adenine base editor proteins selected from Examples 95 to 98; and

[0898] Any one reRNA selected from Examples 47 to 61;

[0899] Here, the adenine base editor protein and the reRNA bind to form a complex, and

[0900] The guide domain of the above reRNA targets the target nucleic acid, and

[0901] The above adenine base editing complex changes the A:T base pairs of the target nucleic acid or an adjacent nucleic acid into C:G base pairs.

[0902] Example 104, Correction window of an adenine-based editing complex

[0903] In the adenine-based editing complex of Example 103,

[0904] The target nucleic acid targeted by the guide domain of the reRNA of the adenine base editing complex is double-stranded DNA, and

[0905] The above target nucleic acid includes a target strand and a non-target strand, and

[0906] The above-mentioned non-target strand includes a transposon-associated motif (TAM) and a protospacer, and

[0907] The above target strand includes a target site that binds complementarily to the above protospacer, and

[0908] The adenine-based editor protein recognizes the TAM, and the guide domain of the reRNA hybridizes with the target site, and

[0909] The above adenine base editor complex can convert adenine contained in the 7th to 15th nuclei, and thymine of the target strand that binds complementarily to it, into cytosine and guanine, respectively, when the nucleus at the 5' end of the TAM is referred to as the 1st.

[0910] Example 105, Cytosine-based Editing Complex

[0911] Cytosine-based editing complexes including the following:

[0912] Any one cytosine-based editor protein selected from Examples 99 to 102; and

[0913] Any one reRNA selected from Examples 47 to 61;

[0914] Here, the cytosine base editor protein and the reRNA bind to form a complex, and

[0915] The guide domain of the above reRNA targets the target nucleic acid, and

[0916] The above cytosine base editing complex changes the C:G base pairs of the target nucleic acid or an adjacent nucleic acid into T:A base pairs.

[0917] Example 106, Correction window of a cytosine-based editing complex

[0918] In the cytosine-based editing complex of Example 105,

[0919] The target nucleic acid targeted by the guide domain of the reRNA of the above cytosine base editing complex is double-stranded DNA, and

[0920] The above target nucleic acid includes a target strand and a non-target strand, and

[0921] The above-mentioned non-target strand includes a transposon-associated motif (TAM) and a protospacer, and

[0922] The above target strand includes a target site that binds complementarily to the above protospacer, and

[0923] The cytosine-based editor protein recognizes the TAM, and the guide domain of the reRNA hybridizes with the target site, and

[0924] The above cytosine base editor complex can convert cytidine contained in the 7th to 15th nucleotides, and guanine of the target strand that binds complementarily to it, into thymine and adenine, respectively, when the nucleotide at the 5' end of the TAM is referred to as the 1st.

[0925] A vector configured to express each component of base editing

[0926] Example 107, expression vector of adenine-based editing system components

[0927] An adenine base editing system component expression vector comprising the following (hereinafter also referred to as the base editing system component expression vector):

[0928] A nucleic acid encoding any one of the adenine base editor proteins selected from Examples 95 to 98 (hereinafter also referred to as base editor protein); and

[0929] A nucleic acid encoding any one of the reRNAs selected from Examples 47 to 61;

[0930] Here, the base editing system component expression vector is configured to express the adenine base editor protein and the reRNA within target cells, and

[0931] The above adenine base editor protein and the above reRNA combine to form any one of the selected adenine base editor complexes in Example 104.

[0932] Example 108, Adenine-based editing system component expression vector, multiplexing

[0933] An adenine base editing system component expression vector comprising the following (hereinafter also referred to as the base editing system component expression vector):

[0934] A nucleic acid encoding any one of the adenine base editor proteins selected from Examples 95 to 98 (hereinafter also referred to as base editor protein);

[0935] nucleic acid encoding the first reRNA,

[0936] Here, the first reRNA is any one selected from Examples 47 to 61; and

[0937] nucleic acid encoding 2nd reRNA,

[0938] Here, the second reRNA is any one selected from Examples 47 to 61;

[0939] Here, the base editing system component expression vector is configured to express the adenine base editor protein, the first reRNA, and the second reRNA within target cells, and

[0940] The above adenine base editor protein and the above first reRNA combine to form a first adenine base editor protein, and

[0941] The first adenine base editor is any one of the adenine base editor complexes selected in Example 104, and targets a first target nucleic acid, and

[0942] The above adenine base editor protein and the above second reRNA combine to form a second adenine base editor protein, and

[0943] The second adenine base editor is any one of the adenine base editor complexes selected in Example 104, and targets the second target nucleic acid.

[0944] Example 109, expression vector of cytosine-based editing system components

[0945] A cytosine base editing system component expression vector comprising the following (hereinafter also referred to as the base editing system component expression vector):

[0946] A nucleic acid encoding any one of the cytosine base editor proteins selected from Examples 99 to 102 (hereinafter also referred to as base editor protein); and

[0947] A nucleic acid encoding any one of the reRNAs selected from Examples 47 to 61;

[0948] Here, the base editing system component expression vector is configured to express the cytosine base editor protein and the reRNA within target cells, and

[0949] The above cytosine base editor protein and the above reRNA combine to form a cytosine base editor complex selected from any one of Examples 105 to 106.

[0950] Example 110, cytosine-based editing system component expression vector, multiplexing

[0951] A cytosine base editing system component expression vector comprising the following (hereinafter also referred to as the base editing system component expression vector):

[0952] A nucleic acid encoding any one of the cytosine base editor proteins selected from Examples 99 to 102 (hereinafter also referred to as base editor protein);

[0953] nucleic acid encoding the first reRNA,

[0954] Here, the first reRNA is any one selected from Examples 47 to 61; and

[0955] nucleic acid encoding 2nd reRNA,

[0956] Here, the second reRNA is any one selected from Examples 47 to 61;

[0957] Here, the base editing system component expression vector is configured to express the cytosine base editor protein, the first reRNA, and the second reRNA within a target cell, and

[0958] The above cytosine base editor protein and the above first reRNA combine to form a first cytosine base editor protein, and

[0959] The first cytosine base editor is a cytosine base editor complex selected from any one of Examples 105 to 106, which targets a first target nucleic acid, and

[0960] The above cytosine base editor protein and the above second reRNA combine to form a second cytosine base editor protein, and

[0961] The second cytosine base editor is a cytosine base editor complex selected from any one of Examples 105 to 106, and targets a second target nucleic acid.

[0962] Example 111, including additional components

[0963] In any one of the base editing system component expression vectors selected from Examples 107 to 110,

[0964] The above base editing system component expression vector further includes one or more selected from the following:

[0965] Promoter; enhancer; intron; polyadenylation signal; Kozak consensus sequence; Internal Ribosome Entry Site (IRES); splice acceptor; 2A sequence; and replication origin.

[0966] Example 112, promoter specific

[0967] In the base editing system component expression vector of Example 111,

[0968] The above promoter is one or more of the following selected:

[0969] SV40 early promoter; mouse mammary tumor virus long terminal repeat (LTR) promoter; adenovirus major late promoter (Ad MLP); herpes simplex virus (HSV) promoter; cytomegalovirus (CMV) promoters such as the CMV immediate early promoter region (CMVIE); rous sarcoma virus (RSV) promoter; human U6 small nuclear promoter (U6) (Miyagishi et al.; Nature Biotechnology 20; 497-500 (2002)); enhanced U6 promoter (e.g., Xia et al.; Nucleic Acids Res. 2003 Sep 1;31(17)); human H1 promoter (H1); and 7SK.

[0970] Example 113, identification of replication origin

[0971] In the base editing system component expression vector of Example 111 or Example 112,

[0972] The above replication origin is one or more of the following selected:

[0973] f1 replication origin; SV40 replication origin; pMB1 replication origin; adeno replication origin; AAV replication origin; and BBV replication origin.

[0974] Example 114, Vector Type

[0975] In any one of the base editing system component expression vectors selected from Examples 107 to 113,

[0976] The above vector is a viral vector, a non-viral vector, or any combination thereof.

[0977] Example 115, virus vector

[0978] In the base editing system component expression vector of Example 114,

[0979] The above vector is a selected virus vector from the following:

[0980] Retrovirus; lentivirus; adenovirus; adeno-associated virus; vacciniavirus; poxvirus; herpes simplex virus; and any combination thereof.

[0981] Example 116, non-viral vector

[0982] In the base editing system component expression vector of Example 114,

[0983] The above vector is a non-viral vector selected from the following:

[0984] Plasmid; phage; naked DNA; DNA complex; PCR amplicon; mRNA; and any combination thereof.

[0985] Example 117, Plasmid Specific

[0986] In the base editing system component expression vector of Example 116,

[0987] The above plasmid is any one of the following selected:

[0988] pcDNA series; pS456; p326; pACYC177; ColE1; pKT230; pME290; pBR322; pUC8 / 9; pUC6; pBD9; pHC79; pIJ61; pLAFR1; pHV14; pGEX series; pET series; pUC19 and any combination thereof.

[0989] Example 118, phage specific

[0990] In the base editing system component expression vector of Example 116,

[0991] The above waste is any one of the following selected:

[0992] λgt4λB; λ-Charon; λΔz1; M13; and any combination thereof.

[0993] Example 119, single vector

[0994] In any one of the base editing system component expression vectors selected from Examples 107 to 118,

[0995] The above vector is a single vector composed of one molecule.

[0996] Example 120, DNA and RNA specific

[0997] In any one of the base editing system component expression vectors selected from Examples 107 to 119,

[0998] The nucleic acid encoding the base editor protein and the nucleic acid encoding the reRNA are each independently DNA or RNA, or

[0999] The nucleic acid encoding the base editor protein, the nucleic acid encoding the first reRNA, and the nucleic acid encoding the second reRNA are each independently DNA or RNA.

[1000] Base editing composition

[1001] Example 121, base editing composition

[1002] A base editing composition comprising the following:

[1003] A base editor protein, or a nucleic acid encoding the base editor protein,

[1004] Here, the base editor protein is any one of the adenine base editor proteins selected from Examples 95 to 98, or any one of the cytosine base editor proteins selected from Examples 99 to 102; and

[1005] reRNA, or nucleic acid encoding the reRNA,

[1006] Here, the reRNA is any one of the reRNAs selected from Examples 47 to 61.

[1007] Here, the base editor protein and the reRNA combine to form a base editing complex, and

[1008] The guide domain of the above reRNA targets the target nucleic acid, and

[1009] The base editing complex changes a base pair of the target nucleic acid or an adjacent nucleic acid to another base pair.

[1010] Example 122, mRNA delivery form

[1011] In the base editing composition of Example 121,

[1012] The above base editing composition comprises mRNA encoding the base editor protein and the reRNA.

[1013] Example 123, Adenine-based editing composition

[1014] In any one of the base editing compositions selected from Examples 121 to 122,

[1015] The above base editor protein is any one of the adenine base editor proteins selected from Examples 95 to 98, and

[1016] The above base editing complex is any one of the adenine base editing complexes selected in Example 104, and

[1017] The above base editing complex changes the A:T base pairs of the target nucleic acid or an adjacent nucleic acid into C:G base pairs, and

[1018] The above base editing composition is also referred to as an adenine base editing composition.

[1019] Example 124, Cytosine-based editing composition

[1020] In any one of the base editing compositions selected from Examples 121 to 122,

[1021] The above base editor protein is any one of the cytosine base editor proteins selected from Examples 99 to 102, and

[1022] The above base editing complex is any one of the cytosine base editing complexes selected from Examples 105 to 106, and

[1023] The above cytosine base editing complex changes the C:G base pairs of the target nucleic acid or an adjacent nucleic acid into T:A base pairs, and

[1024] The above base editing composition is also referred to as a cytosine base editing composition.

[1025] Base Editing Method #2 - Adenine Base Editing

[1026] Example 125, delivery of an adenine-based editing system

[1027] As a method for editing base pairs of target nucleic acids within target cells,

[1028] The above method includes the following:

[1029] Delivering an adenine-based editing system to the target cells;

[1030] Here, the adenine-based editing system is any one of the adenine-based editing complexes selected in Example 104, any one of the adenine-based editing system component expression vectors selected in Example 107, or the adenine-based editing composition of Example 123;

[1031] Here, by the above process, any one of the selected adenine base editing complexes of Example 104 is delivered into the target cell, or the expression of any one of the selected adenine base editing complexes of Example 104 is induced, and

[1032] The guide domain of the reRNA of the adenine base editing complex targets the target nucleic acid within the target cell, and

[1033] The adenine base editing complex recognizes the TAM contained in the target nucleic acid and binds to the target nucleic acid through the guide domain of the reRNA, and

[1034] The above target nucleic acid is a double-stranded nucleic acid and includes a correction window, and

[1035] A:T base pairs within the correction window are changed to G:C base pairs by the adenine base editing complex.

[1036] Example 126, specification of target nucleic acid structure

[1037] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 125,

[1038] The above target nucleic acid includes a target strand and a non-target strand, and

[1039] The guide domain of the reRNA of the adenine-based editing complex is configured to hybridize with the target strand, and

[1040] The target strand of the above target nucleic acid includes a target site that binds complementarily to the guide domain, and

[1041] The non-target strand of the above-mentioned target nucleic acid comprises a transposon-associated motif (TAM) and a protospacer complementary to the above-mentioned target site, and

[1042] The above correction window includes a second correction window that is part of the target strand, and a first correction window that is part of the non-target strand and is complementary (or corresponding) to it.

[1043] When the above method is performed, A included in the first calibration window is changed to G, and the corresponding T in the second calibration window is changed to C.

[1044] Example 127, Refinement Window Structure Specification #1 - Guide Domain Standard

[1045] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 126,

[1046] The first correction window above is a sequence composed of the i-th nucleotide to the (i + j)-th nucleotide of the protospacer, and

[1047] The second calibration window is a portion of the target strand that can be coupled complementarily to the first calibration window, or a portion corresponding to the first calibration window, and

[1048] The above i and j are integers.

[1049] Example 128, i and j limitations

[1050] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 127,

[1051] The above i is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13, or is an integer within the range consisting of the two aforementioned numbers (e.g., an integer between 2 and 7);

[1052] The above j is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11, or is an integer within the range consisting of the two aforementioned numbers (e.g., an integer between 5 and 11).

[1053] Example 129, limiting the relationship between i and j

[1054] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 128,

[1055] The range of j above is an integer between 0 and (13 - i) inclusive.

[1056] Example 130, preferred i and j

[1057] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 128,

[1058] The above i is 2 and the above j is 10; or

[1059] The above i is 2, and the above j is 5.

[1060] Example 131, Specification of Calibration Window Structure #2 - Calibration Window Standard

[1061] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 126,

[1062] The length of the above correction window is l-bp, and

[1063] The adenine base editor complex recognizes a sequence consisting of the m-th nucleotide to the (m + 4)-th nucleotide in the upstream direction relative to the 1st nucleotide of the first correction window as a transposon-associated motif (TAM), and

[1064] The length of the guide domain of the above reRNA is n-nt, and

[1065] The guide domain of the above reRNA is,

[1066] Based on the nucleotide of the target strand corresponding to the first nucleotide downstream from the last nucleotide of the above TAM, it binds complementarily to n consecutive nucleotide sequences in the upstream direction, and

[1067] The above l, m, and n are all integers.

[1068] Example 132, ranges of l, m, and n

[1069] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 131,

[1070] The above m is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13, or is an integer within the range consisting of the two aforementioned numbers (e.g., an integer between 2 and 7),

[1071] The above l is an integer between 1 and 13 inclusive, and

[1072] The above n is 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30, or is an integer within the range consisting of the two aforementioned numbers (e.g., an integer between 18 and 23).

[1073] Example 133, relationship between l and m

[1074] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 132,

[1075] The above l is an integer greater than or equal to (14 - m).

[1076] Example 134, preferred l, m, and n limitations

[1077] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 132,

[1078] The above m is 2, the above l is 13, or

[1079] The above m is 2, and the above l is 6.

[1080] Bass Editing Method #3 - Cytosine Bass Editing

[1081] Example 135, Cytosine-based editing system delivery

[1082] As a method for editing base pairs of target nucleic acids within target cells,

[1083] The above method includes the following:

[1084] Delivering a cytosine-based editing system to the target cell;

[1085] Here, the cytosine-based editing system is any one of the cytosine-based editing complexes selected in Example 104, any one of the cytosine-based editing system component expression vectors selected in Example 107, or the cytosine-based editing composition of Example 123;

[1086] Here, by the above process, any one of the cytosine-based editing complexes selected in Example 104 is delivered into the target cell, or the expression of any one of the cytosine-based editing complexes selected in Example 104 is induced, and

[1087] The guide domain of the reRNA of the above cytosine-based editing complex targets the target nucleic acid within the target cell, and

[1088] The above cytosine-based editing complex recognizes the TAM contained in the target nucleic acid and binds to the target nucleic acid through the guide domain of the reRNA, and

[1089] The above target nucleic acid is a double-stranded nucleic acid and includes a correction window, and

[1090] C:G base pairs within the correction window are changed to T:A base pairs by the above cytosine base editing complex.

[1091] Example 136, specification of target nucleic acid structure

[1092] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 135,

[1093] The above target nucleic acid includes a target strand and a non-target strand, and

[1094] The guide domain of the reRNA of the above cytosine-based editing complex is configured to hybridize with the target strand, and

[1095] The target strand of the above target nucleic acid includes a target site that binds complementarily to the guide domain, and

[1096] The non-target strand of the above-mentioned target nucleic acid comprises a transposon-associated motif (TAM) and a protospacer complementary to the above-mentioned target site, and

[1097] The above correction window includes a second correction window that is part of the target strand, and a first correction window that is part of the non-target strand and is complementary (or corresponding) to it.

[1098] When the above method is performed, C included in the first calibration window is changed to T, and the corresponding G of the second calibration window is changed to A.

[1099] Example 137, Refinement Window Structure Specification #1 - Guide Domain Standard

[1100] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 136,

[1101] The first correction window above is a sequence composed of the i-th nucleotide to the (i + j)-th nucleotide of the protospacer, and

[1102] The second calibration window is a portion of the target strand that can be coupled complementarily to the first calibration window, or a portion corresponding to the first calibration window, and

[1103] The above i and j are integers.

[1104] Example 138, i and j limitations

[1105] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 137,

[1106] The above i is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13, or is an integer within the range consisting of the two aforementioned numbers (e.g., 3 to 7),

[1107] The above j is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11, or is an integer within the range consisting of the two aforementioned numbers (e.g., 1 to 5).

[1108] Example 139, limiting the relationship between i and j

[1109] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 138,

[1110] The range of j above is an integer between 0 and (13 - i) inclusive.

[1111] Example 140, preferred i and j

[1112] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 138,

[1113] The above i is 2 and the above j is 9; or

[1114] The above i is 2, and the above j is 5.

[1115] Example 141, Specification of Calibration Window Structure #2 - Calibration Window Standard

[1116] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 136,

[1117] The length of the above correction window is l-bp, and

[1118] The cytosine base editor complex recognizes a sequence consisting of the m-th nucleotide to the (m + 4)-th nucleotide in the upstream direction relative to the 1st nucleotide of the first correction window as a transposon-associated motif (TAM), and

[1119] The length of the guide domain of the above reRNA is n-nt, and

[1120] The guide domain of the above reRNA is,

[1121] Based on the nucleotide of the target strand corresponding to the first nucleotide downstream from the last nucleotide of the above TAM, it binds complementarily to n consecutive nucleotide sequences in the upstream direction, and

[1122] The above l, m, and n are all integers.

[1123] Example 142, ranges of l, m, and n

[1124] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 141,

[1125] The above m is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or 13, or is an integer within the range consisting of the two aforementioned numbers (e.g., an integer between 2 and 7),

[1126] The above l is an integer between 1 and 13 inclusive, and

[1127] The above n is 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30, or is an integer within the range consisting of the two aforementioned numbers (e.g., an integer between 18 and 23).

[1128] Example 143, relationship between l and m

[1129] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 142,

[1130] The above l is an integer greater than or equal to (14 - m).

[1131] Example 144, preferred l, m, and n limitations

[1132] In the method for editing base pairs of a target nucleic acid within a target cell according to Example 142,

[1133] The above m is 2, the above l is 13, or

[1134] The above m is 2, and the above l is 5.

[1135]

[1136] [Experimental Examples]

[1137] The invention provided by this specification will be described in more detail below through experimental examples and embodiments. These embodiments are intended solely to illustrate the contents disclosed by this specification, and it will be obvious to those skilled in the art that the scope of the contents disclosed by this specification is not to be interpreted as being limited by these embodiments.

[1138] Experimental Example 1. Experimental Method and Materials

[1139] Experimental Example 1.1. Construction of a Plasmid for Mammalian Cell Experiments

[1140] The engineered TnpB protein was optimized for human codons using the codon optimization tool from IDT (Integrated DNA Technologies) and ordered from IDT in gBlock form. The engineered TnpB protein was constructed using site-directed mutagenesis (Q5-site directed mutagenesis kit, NEB #E0554S). The plasmid encoding the engineered reRNA was constructed by assembling PCR-amplified oligonucleotides into pLb-Cpf1-pGL3-U6-sgRNA (Addgene, #107682) using Gibson assembly (NEBuilder HiFi DNA Assembly Cloning Kit, NEB #E5520). The scaffolds for the reRNA were constructed using the wild-type sequence (SEQ ID NO: 5) and the Truncated Scaffold sequences (SEQ ID NO: 6, 7), respectively. To prepare cytidine base editors and adenosine base editors containing TnpB proteins, pCMV-dCpf1-BE (Addgene, #107685) and ABE8e (Addgene, #138489) were modified using Gibson Assembly (NEBuilder HiFi DNA Assembly Cloning Kit, NEB #E5520) to express the intended base editor proteins.

[1141] Experimental Example 1.2. Mammalian Cell Culture and Transfection

[1142] HEK293T (ATCC CRL-11268) cells were cultured in Dulbecco's Modified Eagle Medium supplemented with 10% fetal bovine serum and 1% penicillin / streptomycin (Welgene). Cells were not tested for Mycoplasma contamination. HEK293T cells (7.5X10 4(dogs) were seeded in a 48-well plate, and at a cell density of approximately 80%, plasmids (750 ng) expressing the engineered TnpB protein according to Experimental Example 1.1, the cytidine-based editor protein (TnpB-CBE), and the adenosine-based editor protein (TnpB-ABE), and plasmids (250 ng) expressing the appropriate reRNA were transduced using Lipofectamine 2000 (Invitrogen) according to the manufacturer's protocol. 72 hours after transduction, genomic DNA was isolated using the DNeasy Blood & Tissue Kit (Qiagen).

[1143] K562 cells were prepared under known culture conditions. K562 cells (8X10 5 (dogs) were seeded in well plates, and plasmids (3 µg) expressing the engineered TnpB protein according to Experimental Example 1.1, the cytidine-based editor protein (TnpB-CBE), and the adenosine-based editor protein (TnpB-ABE), and a plasmid (1 µg) expressing the appropriate reRNA were transfected by electroporation using a Neon Transfection System (Thermo Fisher) under conditions of 1300V, 10ms, and 3 pulses.

[1144] iPSC cells were prepared under known culture conditions. iPSC cells (4X10 5 (dogs) were seeded into well plates, and plasmids (1500 ng) expressing the engineered TnpB protein according to Experimental Example 1.1, the cytidine-based editor protein (TnpB-CBE), and the adenosine-based editor protein (TnpB-ABE), and plasmids (500 ng) expressing the appropriate reRNA were transfected using the Lonza 4D nucleofector system (Code: CB150).

[1145] Experimental Example 1.3. Targeted Deep Sequencing

[1146] To analyze edit frequencies, a deep sequencing library was generated by amplifying target sites using nested primary PCR, secondary PCR, and tertiary PCR with PrimeSTAR® GXL DNA Polymerase (TAKARA) and primers containing the Tru-Seq HT Dual Index. The library was sequenced using a paired sequencing system with Illumina iSeq. Base edit and unwanted indel frequencies were expressed as the percentage of sequencing reads containing correct edits or indels out of the total sequencing reads. The computer program used to analyze edit frequencies is available at the following address: https: / / github.com / ibs-cge2 / prime_editor_analysis.

[1147] Experimental Example 1.4. Confirmation of Gene Editing Efficiency

[1148] After transfecting cells prepared according to Experimental Example 1.2 with the plasmid constructed according to Experimental Example 1.1, the genomic DNA of the cells was isolated and the results were verified by targeted deep sequencing according to Experimental Example 1.3. The specific composition of the TnpB protein, base editor protein, and reRNA included in the plasmid was disclosed for each experimental example. Through the targeted deep sequencing, the frequency of indel introduction (%) was calculated for experiments in which the TnpB complex was introduced, and the frequency of nucleobase correction (%) was calculated for experiments in which the base editing system was introduced.

[1149] Experimental Example 1.5. Code Availability

[1150] Base edit frequencies and unwanted indel frequencies in targeted deep sequencing data were calculated using source code written in Python (version 3.9) (https: / github.com / ibs-cge2 / prime_editor_analysis, written by BotBot Inc. DOI: 10.5281 / zenodo.7726909). The 'maund_default.py' script was used to analyze base edit frequencies and unwanted indel frequencies. In this experiment, the 'prime_editor.py' script was used with default parameters to calculate base edit frequencies from insertions and deletions.

[1151] Experimental Example 1.6. Cryo-EM structural analysis using PyMOL

[1152] Cryo-EM structural analysis was performed using the PyMOL Molecular Graphics System (version 1.20, Schrφdinger, LLC). The structure of the target protein was obtained from the Protein Data Bank (PDB ID: 8EXA). To analyze the interactions between the protein and nucleic acids, the structure was visualized, and residues within 4 angstroms of the target DNA / reRNA were identified. Interatomic distance measurements were performed using PyMOL's built-in distance measurement tool.

[1153] Experimental Example 2. Discovery of variants with enhanced nucleic acid cleavage activity

[1154] Experimental Example 2.1. Introduction of a single-position variation

[1155] For the wild-type ISDra2 TnpB protein, Cryo-EM structural analysis was performed according to Experimental Example 1.5, and all amino acid positions within 4 angstrom distance from the target nucleic acid or reRNA were identified.

[1156] The results of the cryo-EM structural analysis are shown in Figure 1.

[1157] We aimed to identify TnpB variants with enhanced nucleic acid cleavage efficiency by substituting the amino acid at the corresponding position with arginine. First, we prepared TnpB variant proteins in which only one of the amino acids at a specific position above was substituted with arginine and experimented to see if indel introduction efficiency was improved.

[1158] The synthesized TnpB variant proteins are shown in the following table:

[1159]

[1160] Plasmids encoding the TnpB variant proteins and reRNA, respectively, of the table above were prepared according to Experimental Example 1.1 and transfected into HEK293T cells cultured according to Experimental Example 1.2. Subsequently, targeted deep sequencing was performed according to Experimental Example 1.3 to measure the frequency of indel introduction.

[1161] The composition of reRNA in the above experiment is shown in the following table:

[1162] ScaffoldGuide Domain (RNA)tGGTGGCTGCGGGAATCTCAGACACCTTAAACGCTCATGGAGGCTATgaaaATGGTCTGCGAAGTGAGAATCACGCGACTTTAGTCGTGTGAGGTTCAA (SEQ ID NO: 7)TCAAGTACCACCAGTTTTAT (SEQ ID NO: 86)

[1163] The experimental results were measured for the frequency of introducing indels into the target nucleic acids of cells and are shown in Figure 2.

[1164] Experimental results showed that the indel induction efficiency of all TnpB protein variants with introduced arginine substitutions was not improved; in fact, many variants exhibited lower indel induction efficiency than the wild-type TnpB protein, or showed almost no indel induction efficiency. The inventors of this application identified the mutations introduced into the TnpB variants that showed higher indel induction efficiency than the wild-type TnpB protein, as well as their amino acid sequences.

[1165] Experimental Example 2.2. Combinations of Single-Location Variants #1 - Top 5 Variant Combinations

[1166] Based on the results of Experimental Example 2.1, five variants with the highest indel introduction efficiency were selected. TnpB protein variants were constructed by introducing the above variants in all possible combinations.

[1167] The prepared TnpB variant proteins are shown in the following table:

[1168] LabelVariantLabelVariantTnpB-TnpB-(Q128R + N255R + K281R)Q128R, N255R, K281RTnpB-(Q128R)Q128RTnpB-(Q128R + N255R + S72R)Q128R, N255R, S72RTnpB-(N255R)N255RTnpB-(Q128R + N255R + Q64R)Q128R, N255R, Q64RTnpB-(K281R)K281RTnpB-(Q128R + K281R + S72R)Q128R, K281R, S72RTnpB-(S72R)S72RTnpB-(Q128R + K281R + Q64R)Q128R, K281R, Q64RTnpB-(Q64R)Q64RTnpB-(Q128R + S72R + Q64R)Q128R, S72R, Q64RTnpB-(Q128R + N255R)Q128R, N255RTnpB-(N255R + K281R + S72R)N255R, K281R, S72RTnpB-(Q128R + K281R)Q128R, K281RTnpB-(N255R + K281R + Q64R)N255R, K281R, Q64RTnpB-(Q128R + S72R)Q128R, S72RTnpB-(N255R + S72R + Q64R)N255R, S72R, Q64RTnpB-(Q128R + Q64R)Q128R, Q64RTnpB-(K281R + S72R + Q64R)K281R, S72R, Q64RTnpB-(N255R + K281R)N255R, K281RTnpB-(Q128R + N255R + K281R + S72R)Q128R, N255R, K281R, S72RTnpB-(N255R + S72R)N255R, S72RTnpB-(Q128R + N255R + K281R + Q64R)Q128R, N255R, K281R, Q64RTnpB-(N255R + Q64R)N255R, Q64RTnpB-(Q128R + N255R + S72R + Q64R)Q128R, N255R, S72R, Q64RTnpB-(K281R + S72R)K281R,S72RTnpB-(Q128R + K281R + S72R + Q64R)Q128R, K281R, S72R, Q64RTnpB-(K281R + Q64R)K281R, Q64RTnpB-(N255R + K281R + S72R + Q64R)N255R, K281R, S72R, Q64RTnpB-(S72R + Q64R)S72R, Q64RTnpB-(Q128R + N255R + K281R + S72R + Q64R)Q128R, N255R, K281R, S72R, Q64R,

[1169] In the above, each variant was introduced based on the amino acid sequence of SEQ ID NO. 1, and if multiple variants are listed, it means that all variants were introduced. The above TnpB variant protein and reRNA were introduced into cells according to the method of Experimental Example 2.1, and the frequency of introducing indels into target nucleic acids was measured.

[1170] The composition of the reRNA used in the above experiment is shown in the following table:

[1171] LabelScaffoldGuide Domain (RNA)EMX1-2tGGTGGCTGCGGGAATCTCAGACACCTTAAACGCTCATGGAGGCTATgaaaATGGTCTGCGAAGTGAGAATCACGCGACTTTAGTCGTGTGAGGTTCAA (SEQ ID NO: 7)GCATTTCTGTTTTTAATTTAT (SEQ ID NO: 84)EMX1-1SEQ ID NO: 7GTGATGGGAGCCCTTCTTCT (SEQ ID NO: 83)HEK1-1SEQ ID NO: 7TCAAGTACCACCAGTTTTAT (SEQ ID NO: 86)AGBL1SEQ ID NO: 7TGTTGGCTCAAACACCAGAT (SEQ ID NO: 81)

[1172] The experimental results are shown in Figures 3 to 5.

[1173] Experimental results confirmed that the more modifications included to improve indel introduction efficiency, the greater the improvement in indel introduction efficiency.

[1174] Based on the above experimental results, an engineered TnpB protein containing the amino acid sequence of SEQ ID NO. 2, which includes five variations that increase the efficiency of indel introduction, was identified. Hereinafter, this is referred to as enTnpB.

[1175] Experimental Example 2.3. Combination between single-position variants #1 - Variant addition combination

[1176] Five additional mutations were added to the engineered TnpB protein with increased indel introduction efficiency derived in Experimental Example 2.2, and the indel introduction efficiency was measured.

[1177] The prepared TnpB variant proteins are shown in the following table:

[1178] BaseAdditional VariantAll Variants Compared to Wild-type TnpBenTnpB (SEQ ID NO: 2)S57RQ128R, N255R, K281R, S72R, Q64R, S57RSame as aboveA237RQ128R, N255R, K281R, S72R, Q64R, A237RSame as aboveT350RQ128R, N255R, K281R, S72R, Q64R, T350RSame as aboveV254RQ128R, N255R, K281R, S72R, Q64R, V254RSame as aboveK310RQ128R, N255R, K281R, S72R, Q64R, K310RSame as aboveS57R, A237RQ128R, N255R, K281R, S72R, Q64R, S57R, A237RSame as aboveS57R, T350RQ128R, N255R, K281R, S72R, Q64R, S57R, T350RSame as aboveS57R, V254RQ128R, N255R, K281R, S72R, Q64R, S57R, V254RSame as aboveS57R, K310RQ128R, N255R, K281R, S72R, Q64R, S57R, K310RSame as aboveA237R, T350RQ128R, N255R, K281R, S72R, Q64R, A237R, T350RSame as aboveA237R, V254RQ128R, N255R, K281R, S72R, Q64R, A237R, V254RSame as aboveA237R, K310RQ128R, N255R, K281R, S72R, Q64R, A237R, K310RSame as aboveT350R, V254RQ128R, N255R, K281R, S72R, Q64R, T350R, V254RSame as aboveT350R, K310RQ128R, N255R, K281R, S72R, Q64R, T350R, K310RSame as aboveV254R, K310RQ128R,N255R, K281R, S72R, Q64R, V254R, K310RSame as aboveS57R, A237R, T350RQ128R, N255R, K281R, S72R, Q64R, S57R, A237R, T350RSame as aboveS57R, A237R, V254RQ128R, N255R, K281R, S72R, Q64R, S57R, A237R, V254RSame as aboveS57R, A237R, K310RQ128R, N255R, K281R, S72R, Q64R, S57R, A237R, K310RSame as aboveS57R, T350R, V254RQ128R, N255R, K281R, S72R, Q64R, S57R, T350R, V254RSame as aboveS57R, T350R, K310RQ128R, N255R, K281R, S72R, Q64R, S57R, T350R, K310RSame as aboveS57R, V254R, K310RQ128R, N255R, K281R, S72R, Q64R, S57R, V254R, K310RSame as aboveA237R, T350R, V254RQ128R, N255R, K281R, S72R, Q64R, A237R, T350R, V254RSame as aboveA237R, T350R, K310RQ128R, N255R, K281R, S72R, Q64R, A237R, T350R, K310RSame as aboveA237R, V254R, K310RQ128R, N255R, K281R, S72R, Q64R, A237R, V254R, K310RSame as aboveT350R, V254R, K310RQ128R, N255R, K281R, S72R, Q64R, T350R, V254R, K310RSame as aboveS57R, A237R, T350R, V254RQ128R, N255R, K281R, S72R, Q64R, S57R, A237R, T350R, V254RSame as aboveS57R, A237R, T350R,K310RQ128R, N255R, K281R, S72R, Q64R, S57R, A237R, T350R, K310RSame as aboveS57R, A237R, V254R, K310RQ128R, N255R, K281R, S72R, Q64R, S57R, A237R, V254R, K310RSame as aboveS57R, T350R, V254R, K310RQ128R, N255R, K281R, S72R, Q64R, S57R, T350R, V254R, K310RSame as aboveA237R, T350R, V254R, K310RQ128R, N255R, K281R, S72R, Q64R, A237R, T350R, V254R, K310RSame as aboveS57R, A237R, T350R, V254R, K310RQ128R, N255R, K281R, S72R, Q64R, S57R, A237R, T350R, V254R, K310R,

[1179] In the above, each variant was introduced based on the amino acid sequence of SEQ ID NO. 2, and if multiple variants are listed, it means that all variants were introduced. The above TnpB variant protein and reRNA were introduced into cells according to the method of Experimental Example 2.1, and the frequency of introducing indels into target nucleic acids was measured.

[1180] The composition of the reRNA used in the above experiment is shown in the following table:

[1181] LabelScaffoldGuide Domain (RNA)Site9SEQ ID NO: 7AGATGATGTTTCCACACATA (SEQ ID NO: 100)EMX1-1SEQ ID NO: 7GTGATGGGAGCCCTTCTTCT (SEQ ID NO: 83)AGBL1SEQ ID NO: 7TGTTGGGCTCAAACACCAGAT (SEQ ID NO: 81)HEK1-1SEQ ID NO: 7TCAAGTACCACCAGTTTTAT (SEQ ID NO: 86)EMX1-2SEQ ID NO: 7GCATTTCTGTTTTTAATTTAT (SEQ ID NO: 84)

[1182] The experimental results are shown in Figures 6 to 11.

[1183] Experimental results showed that the indel introduction efficiency of TnpB variant proteins with up to 6 or 7 mutations introduced for some target genes was high, but overall, enTnpB with up to 5 mutations introduced showed the highest indel introduction efficiency (Fig. 11).

[1184] Experimental Example 2.4. Comparison of overall gene editing efficiency between wild-type TnpB and engineered TnpB (enTnpB).

[1185] Figure 12 shows the results of comparing the editing efficiency for various target genes in HEK293T according to Experimental Examples 2.2 and 2.3. As a result of the experiment, it was confirmed that the gene editing efficiency of enTnpB, which is specific to the stomach, was much higher across various target genes compared to wild-type TnpB.

[1186] Experimental Example 2.5. Cell Line Change Experiment

[1187] Experiments were conducted to verify whether the trend observed in Experimental Example 2.4 was reproduced even when the cell line was changed. Specifically, the experiment was conducted by changing the cells delivering the TnpB system according to Experimental Example 2.4 to K562 cells and iPSC cells. The composition of the reRNA used in the experiment was as shown in Table 4.

[1188] The experimental results are shown in Figures 13 and 14. The experimental results confirmed that enTnpB exhibited higher gene editing efficiency compared to wild-type TnpB, regardless of cell type. Therefore, it is believed that enTnpB exhibits improved gene editing efficiency compared to wild-type TnpB, regardless of cell type.

[1189] Experimental Example 2.6. Efficiency Comparison Experiment with SpCas9

[1190] Experiments were conducted to compare the gene editing efficiency of the specific enTnpB mentioned above with that of the conventionally most widely used Streptococcus pyogenes-derived Cas9 (SpCas9). However, since TnpB operates by recognizing the TAM sequence of 5'-TTGAT-3', whereas SpCas9 operates by recognizing the PAM of 5'-NGG-3', it is difficult to directly compare editing efficiencies for completely identical target sequences. To create conditions as similar as possible, experiments were conducted by selecting gene targets that could target adjacent regions of the enTnpB system and the SpCas9 system.

[1191] The composition of the guide domain of the enTnpB reRNA and the guide domain of the SpCas9 sgRNA is shown in the following table:

[1192] LabelEffectorScaffoldGuide Domain (RNA)HEK1-1enTnpB (SEQ ID NO: 2)SEQ ID NO: 7TCAAGTACCACCAGTTTTAT (SEQ ID NO: 86)AGBL1Same as aboveSame as aboveTGTTGGCTCAAACACCAGAT (SEQ ID NO: 81)Site1Same as aboveSame as aboveTTCCAGTTAAGGAGAGGAAT (SEQ ID NO: 90)Site9Same as aboveSame as aboveAGATGATGTTTCCACACATA (SEQ ID NO: 100)RUNX1-1Same as aboveSame as aboveATTGATGGCTACATATCAGA (SEQ ID NO: 88)Site4Same as aboveSame as aboveGACCCAAAGAAATGTATTCC (SEQ ID NO: 95)EMX1-1Same as aboveSame as aboveGTGATGGGAGCCCTTCTTCT (SEQ ID NO: 83)Site2Same as aboveSame as aboveTTTACACATCATCATATACA (SEQ ID NO: 93)HEK1-1SpCas9 (SEQ ID NO: 518)SEQ ID NO: 519ATTTTTCTTCATAAAACTGG (SEQ ID NO: 520)AGBL1Same as aboveSame as aboveTTGGCTCAAACACCAGATTT (SEQ ID NO: 521)Site1Same as aboveSame as aboveTTTCCAGTTAAGGAGAGGAA (SEQ ID NO: 522)Site9Same as aboveSame as aboveACAAGTTGTACAAATATGTG (SEQ ID NO: 523)RUNX1-1Same as aboveSame as aboveATTGATGGCTACATATCAGA (SEQ ID NO: 524)Site4Same as aboveSame as aboveTAGAAGGAATACATTTCTTT (SEQ IDNO: 525)EMX1-1Same as aboveSame as aboveAGGGCTCCCATCACATCAAC (SEQ ID NO: 526)Site2Same as aboveSame as aboveGATGATGTGTAAAATCAAAC (SEQ ID NO: 527)

[1193] The positional relationship between reRNA and sgRNA in the table above is schematically shown in Fig. 17. Using the above configuration, the gene editing frequency was measured according to Experimental Example 1.4.

[1194] The experimental results are shown in Figures 15 and 16.

[1195] Experimental results confirmed that the gene editing efficiency of the enTnpB system was higher than that of the SpCas9 system. Furthermore, while the SpCas9 system exhibited relatively large variability in gene editing efficiency, the enTnpB system was confirmed to stably introduce indels with high efficiency. This suggests that, at least for target genes where the enTnpB system can be utilized, reliable gene editing with high efficiency is possible using the enTnpB system, even when compared to the SpCas9 system.

[1196] Experimental Example 3. Confirmation of gene editing activity of length-reduced TnpB

[1197] Experimental Example 3.1. Experiment Overview

[1198] A portion of the 3'-terminal region of the gene encoding the TnpB protein partially overlaps with the gene region encoding reRNA. Accordingly, the inventors of this application predicted that a portion of the C-Terminal Domain (CTD) of the TnpB protein may be a non-biologically functional part and sought to determine whether the gene editing function of the TnpB protein is maintained even when this part is removed.

[1199] The composition of the TnpB proteins used in the experiment is shown in the following table:

[1200] LabelTnpB ProteinTnpB-WT-408aa(WT-TnpB)SEQ ID NO: 1TnpB-ΔCTD-405SEQ ID NO: 485TnpB-ΔCTD-402SEQ ID NO: 488TnpB-ΔCTD-399SEQ ID NO: 491TnpB-ΔCTD-396SEQ ID NO: 494TnpB-ΔCTD-393SEQ ID NO: 497TnpB-ΔCTD-390SEQ ID NO: 500TnpB-ΔCTD-387SEQ ID NO: 503TnpB-ΔCTD-384SEQ ID NO: 506TnpB-ΔCTD-381SEQ ID NO: 509TnpB-ΔCTD-378SEQ ID NO: 512TnpB-ΔCTD-375(aa)SEQ ID NO: 515TnpB-ΔCTD-374aaSEQ ID NO: 516TnpB-ΔCTD-373aaSEQ ID NO: 517TnpB-ΔCTD-372(aa)SEQ ID NO: 518TnpB-ΔCTD-369SEQ ID NO: 519TnpB-ΔCTD-366SEQ ID NO: 520

[1201] The composition of the reRNA used in the experimental example is shown in the following table:

[1202] LabelScaffoldGuide Domain (RNA)AGBL1SEQ ID NO: 7TGTTGGCTCAAACACCAGAT (SEQ ID NO: 81)Site1SEQ ID NO: 7TTCCAGTTAAGGAGAGGAAT (SEQ ID NO: 90)Site9SEQ ID NO: 7AGATGATGTTTCCACACATA (SEQ ID NO: 100)

[1203] Using the TnpB protein of Table 8 and the reRNA of Table 9, the indel editing frequency was determined according to Experimental Example 1.4.

[1204] Experimental Example 3.2. Experimental Results

[1205] The experimental results are shown in Figures 18 and 19. When amino acids were sequentially removed from the C-terminus of the CTD in the wild-type TnpB protein (408aa), the frequency of indel introduction was not statistically significant compared to the wild-type TnpB up to the case of 373aa based on the total length of the TnpB protein, but the frequency of indel introduction decreased to a statistically significant level when it became 372aa. Accordingly, it can be evaluated that the amino acid sequence of approximately 34aa in the C-terminus of the wild-type TnpB protein does not have a significant effect on the target-specific nucleic acid cleavage activity of the TnpB protein.

[1206] Experimental Example 4. Verification of TAM expansion of enTnpB

[1207] Experimental Example 4.1. Confirmation of mutation location in enTnpB

[1208] Cryo-EM structural analysis was performed to analyze how each of the five variant sites introduced into the enTnpB protein selected in Experimental Example 2 interacts with the target nucleic acid or reRNA. The results of the analysis are shown in Figures 20 and 21.

[1209] Analysis results confirmed that the S72, Q64, and S57 positions are located close to the TAM sequence, the Q128 position is located close to a sequence near the TAM, the K281 position is located close to the target site of the target nucleic acid that interacts with the reRNA guide domain, and the N255 position is located close to the pseudonot region of the reRNA, indicating a high probability of interaction with each of these respective sites.

[1210] Experimental Example 4.2. Verification of TAM expansion

[1211] The inventors of this application expected that since a mutation was introduced at a location in the TnpB protein that is highly likely to interact with the TAM sequence when constructing the enTnpB protein, there would also be a change in the TAM sequence of the wild-type TnpB protein, 5'-TTGAT-3'. Accordingly, they designed reRNAs in which gene editing can occur when 5'-TNGAT-3', 5'-TTNAT-3', and 5'-TTGNT-3' are each recognized as TAM sequences, and confirmed the frequency of indel introduction according to Experimental Example 1.5.

[1212] The composition of TnpB used in the experiment is shown in the following table:

[1213] LabelTnpB ProteinTnpBSEQ ID NO: 1enTnpBSEQ ID NO: 2

[1214] The composition of the reRNA used in the experiment is shown in the following table:

[1215] 22-year-old reRNALabelScaffoldGuide Domain (RNA)TAGAT-1SEQ ID NO: 7TGAAGGAAAAGTTACAAAGG (SEQ ID NO: 530)TAGAT-2Same as aboveGGTAGAGTCTATATGAATTG (SEQ ID NO: 530). 531)TCGAT-1Same as aboveATACTGAACACAAGAAACAA (SEQ ID NO: 532)TGGAT-1Same as aboveGAAGGGGACTCAGAGGATGC (SEQ ID NO: 533)TGGAT-2Same as aboveAAGAGTATATGGATAATGTC (SEQ ID NO: 534)TTAAT-1Same as above aboveATCTACTTTCCCAGATATTT (SEQ ID NO: 535)TTAAT-2Same as aboveAAAAAGGGGTTTTAGAATTA (SEQ ID NO: 536)TTCAT-1Same as aboveGGAAAGGATCAAAGTCAAAA (SEQ ID NO: 537)TTCAT-2Same as aboveAGAAGGTCCTAAAGAGATAT (SEQ ID NO: 537) 538)TTTAT-1Same as aboveAATAGAGATACAGATAAAAA (SEQ ID NO: 539)TTTAT-2Same as aboveAAAAGATTGACTCTGATATT (SEQ ID NO: 540)TTGCT-1Same as aboveACCTGGGATCACCTTCCATT (SEQ ID NO: 541)TTGCT-2Same as above aboveACTTTGGGAGATGCTAAGTT (SEQ ID NO: 542)TTGGT-1Same as aboveTGTCACAACTGGGGCTGAGA (SEQ ID NO: 543)TTGGT-2Same as aboveTTGGTCCAGAAAGGTGAGAC (SEQ ID NO: 544)TTGTT-1Same as aboveTTAGTGGGTAACATTATCTT (SEQ ID NO: 543). NO: 545)TTGTT-2Same asaboveTATTAAGACTATTAGATGTT (SEQ ID NO: 546)

[1216] Page 23 of 23 reRNALabelScaffoldGuide Domain (RNA)TAGCT-1SEQ ID NO: 7GCAAACAAGTGCAGAATATC (SEQ ID NO: 547)TAGCT-2Same as aboveTGTGGCAAAAGAGTATATAAT (SEQ ID NO: 547). 548)TAGGT-1Same as aboveAGGAAGGAAAGACAAAATCT (SEQ ID NO: 549)TAGGT-2Same as aboveGAACTGGTAAATGGTAGATA (SEQ ID NO: 550)TAGTT-1Same as aboveGAGCAGATGCTGGATATATC (SEQ ID NO: 551)TCGCT-1Same as above aboveAGTGGGAAGACTAAATGATA (SEQ ID NO: 552)TCGCT-2Same as aboveAACAAGACTAATAAAGAATA (SEQ ID NO: 553)TCGGT-1Same as aboveGTCAACATATGTGTTATTCA (SEQ ID NO: 554)TCGGT-2Same as aboveAATAGCAGGATGCTTACACA (SEQ ID NO: 553). 555)TCGTT-1Same as aboveAACTAAACTCTGGACTTTAC (SEQ ID NO: 556)TGGCT-1Same as aboveAAAAGGGTCCCAGATACATT (SEQ ID NO: 557)TGGCT-2Same as aboveAAGAAAGAATGAGGGCTCAA (SEQ ID NO: 558)TGGGT-1Same as above aboveAAAATAGTGAAGCATTCTCT (SEQ ID NO: 559)TGGGT-2Same as aboveGTGAGTGTGATGAAATAATT (SEQ ID NO: 560)TGGTT-1Same as aboveAAGAAAGAACACCTCTCCTT (SEQ ID NO: 561)TGGTT-2Same as aboveATGGAGGTGAATGCGTTTAC (SEQ ID NO: 560). 562)

[1217] [ PMC free article ] [ PubMed ] 25. reRNALabelScaffoldGuide Domain (RNA)TAGAT-1SEQ ID NO: 7ACTAGACACTTGGAATAATT (SEQ ID NO: 563)TAGAT-2Same as aboveAAAAGAAGGAGTAGGACATA (SEQ ID NO: 563). 564)TAGAT-3Same as aboveAGTAAAGCCATGCTGATATA (SEQ ID NO: 565)TAGAT-4Same as aboveGACAAAGATAATGAACAAAT (SEQ ID NO: 566)TAGAT-5Same as aboveTAAAAGGAATGTAACATATA (SEQ ID NO: 567)TAGAT-6Same as aboveAAAAATATCCTTAAAAAATA (SEQ ID NO: 568)TAGAT-7Same as aboveTGAAGGAAAAGTTACAAAGG (SEQ ID NO: 569)TAGAT-8Same as aboveGGTAGAGTCTATATGAATTG (SEQ ID NO: 570)TAGAT-9Same as aboveAACCACATATCCAGAAAACC (SEQ ID NO: 571)TCGAT-1Same as aboveATAGGTATGTTTATAAATCT (SEQ ID NO: 572)TCGAT-2Same as aboveAAAACCAACAGTTCTTCTTT (SEQ ID NO: 573)TCGAT-3Same as aboveTCTGTTCACTAGTATTTTAT (SEQ ID NO: 574)TCGAT-4Same as aboveATACTGAACACAAGAAACAA (SEQ ID NO: 574) 575)TCGAT-5Same as aboveGTGTATGTCAATGCACTTTT (SEQ ID NO: 576)TCGAT-6Same as aboveACAAGGGTAGACAAATAGAA (SEQ ID NO: 577)TCGAT-7Same as aboveGCTAGGAAGAAACTGCATCA (SEQ ID NO: 578)TCGAT-8Same as aboveaboveCAAGTGGAAGAAAGGATATC (SEQ ID NO: 579)TCGAT-9Same as aboveGTGAAAATCCTCAGTAAAAT (SEQ ID NO: 580)TGGAT-1Same as aboveAAGATGATCTTCTGAAAATC (SEQ ID NO: 581)TGGAT-2Same as aboveAAAGTACTCTATAGGGTGAT (SEQ ID NO: 582)TGGAT-3Same as aboveATGTACAACTTTAGTAGATA (SEQ ID NO: 583)TGGAT-4Same as aboveAATGTCATCTTTGAGAAAAA (SEQ ID NO: 584)TGGAT-5Same as aboveGAAAAGATTAACTGAAAAAA (SEQ ID NO: 585)TGGAT-6Same as aboveTTAACTAACTACCATATGCT (SEQ ID NO: 586)TGGAT-7Same as aboveATAGTGTTTAAGTGTATAAA (SEQ ID NO: 587)TGGAT-8Same as aboveAAGAGTATATGGATAATGTC (SEQ ID NO: 588)TGGAT-9Same as aboveGAACTTGACAAACATAATGC (SEQ ID NO: 589)TGGAT-10Same as aboveGAAGGGGACTCAGAGGATGC (SEQ ID NO: 590)TTGCT-1Same as aboveATTTCCAGCTAGAATATAAA (SEQ ID NO: 591)TTGCT-2Same as aboveAACATGGAAAGTATAGGATG (SEQ ID NO: 592)TTGCT-3Same as aboveAAGGGACAGTCACACTCTAA (SEQ ID NO: 593)TTGCT-4Same as aboveATAAAGATACCTGAAAATGT (SEQ ID NO: 594)TTGCT-5Same as aboveAAAGAAACATGTGTATAAAT (SEQ ID NO: 595)TTGCT-6Same as aboveATAGTCACTCCCTCTTACCA (SEQID NO: 596)TTGCT-7Same as aboveGCAAAGGACATGATCTCATT (SEQ ID NO: 597)TTGCT-8Same as aboveATAAAGATACCTGAGAATAT (SEQ ID NO: 598)TTGCT-9Same as aboveACTTTGGGAGATGCTAAGTT (SEQ ID NO: 599)TTGCT-10Same as aboveACCTGGGATCACCTTCCATT (SEQ ID NO: 600)TTGGT-1Same as aboveAGAATGGCAGTGCAATACGT (SEQ ID NO: 601)TTGGT-2Same as aboveAAGAGTGATACTAGTTTATC (SEQ ID NO: 602)TTGGT-3Same as aboveACACTAGAAAAGGATAAAAC (SEQ ID NO: 603)TTGGT-4Same as aboveAGTGTGACTACACTCAGATT (SEQ ID NO: 604)TTGGT-5Same as aboveAATAAGAAAATATATGAGAA (SEQ ID NO: 605)TTGGT-6Same as aboveTGTCACAACTGGGGCTGAGA (SEQ ID NO: 606)TTGGT-7Same as aboveGTTAGGGTATATCTGTAGTC (SEQ ID NO: 607)TTGGT-8Same as aboveGAAGGCATGTGTGGCAATTG (SEQ ID NO: 608)TTGGT-9Same as aboveAGAAAGACAGATGACCAAAT (SEQ ID NO: 609)TTGGT-10Same as aboveTTGGTCCAGAAAGGTGAGAC (SEQ ID NO: 610)

[1218] The experimental results are shown in FIGS. 22 to 25. The experimental results confirmed that the enTnpB protein recognizes extended TAM sequences compared to the wild-type TnpB protein and exhibits target-specific gene editing activity. Based on the combined experimental results, it is thought that the enTnpB protein recognizes the TAMs of 5'-TNGAT-3' and 5'-TTGVT-3' (Fig. 24).

[1219] Experimental Example 5. Verification of Base Editing System Effectiveness

[1220] Experimental Example 5.1. Verification of Base Editing System Effect #1

[1221] To construct a base editing system using dTnpB proteins prepared by introducing the D191A mutation into wild-type TnpB according to prior literature, and to verify whether base pair correction is possible, a base editing system with the following configuration was constructed:

[1222] LabelConstructSEQID NOTnpB-CBENLS-APOBEC1-dTnpB-UGI-UGI-NLS77enTnpB-CBENLS-APOBEC1-denTnpB-UGI-UGI-NLS78TnpB-ABENLS-TadA8e-dTnpB-NLS74enTnpB-ABENLS-TadA8e-denTnpB-NLS75

[1223] In the table above, dTnpB refers to the TnpB variant protein of SEQ ID NO. 8, and denTnpB refers to the TnpB variant protein of SEQ ID NO. 9. The base editing system of the table above is schematically shown in FIG. 26.

[1224] The composition of the reRNA used in the experimental example is shown in the following table:

[1225] LabelScaffoldGuide Domain (RNA)Site2SEQ ID NO: 7TTTACACATCATCATATACA (SEQ ID NO: 93)Site4SEQ ID NO: 7GACCCAAAGAAATGTATTCC (SEQ ID NO: 95)Site12SEQ ID NO: 7TAAGGAACTAGAATCTAAAA (SEQ ID NO: 103)EMX1-1SEQ ID NO: 5GTGATGGGAGCCCTTCTTCT (SEQ ID NO: 83)Site5SEQ ID NO: 7ATTCAAAAACACGCAAACCC (SEQ ID NO: 96)RUNX1-1SEQ ID NO: 7ATTGATGGCTACATATCAGA (SEQ ID NO: 88)RUNX1-2SEQ ID NO: 7CATAACAGCTAACAGTTTTTT (SEQ ID NO: 89)HEK1-1SEQ ID NO: 7TCAAGTACCACCAGTTTTAT (SEQ ID NO: 86)AGBL1SEQ ID NO: 7TGTTGGCTCAAACACCAGAT (SEQ ID NO: 81)TTRSEQ ID NO: 7GGCAGGACTGCCTCGGACAG (SEQ ID NO: 92)Site9SEQ ID NO: 7AGATGATGTTTCCACACATA (SEQ ID NO: 100) Run ID NO: 98)

[1226] The base editing system shown in the table above was introduced into target cells according to Experimental Example 1.4, and the genome of the target cells was analyzed by deep sequencing to confirm whether the intended base editing effect occurred.

[1227] The experimental results are shown in Figures 28 to 36. The experimental results confirmed that both the base editor protein containing dead-enTnpB and the base editor protein containing dead-TnpB edited nucleotides as intended at various targets. Synthesizing the experimental results, it was confirmed that the cytosine base editor containing dead-enTnpB targeted C located at positions 2 through 13 in the protospacer sequence, and the correction efficiency was particularly high when C was located at positions 2 through 7. Meanwhile, it was confirmed that the adenine base editor containing dead-enTnpB targeted A located at positions 2 through 13 in the protospacer sequence, and the correction efficiency was particularly high when A was located at positions 2 through 7.

[1228] Experimental Example 5.2. Verification of the effect of the base editing system on the extended TAM

[1229] According to Experimental Example 4, an experiment was conducted to determine whether base editing occurs on a target sequence containing an extended TAM sequence that was confirmed to be recognizable by enTnpB.

[1230] Specifically, the base editing protein used in the experimental example is as disclosed in Table 12, and the composition of the reRNA used in the experimental example is shown in the following table:

[1231] LabelScaffoldGuide Domain (RNA)TAGAT-site1SEQ ID NO: 7GACAAAGATAATGAACAAAT (SEQ ID NO: 611)TAGAT-site2Same as aboveACTAGACACTTGGAATAATT (SEQ ID NO: 612)TCGAT-site1Same as aboveATACTGAACACAAGAAACAA (SEQ ID NO: 613)TCGAT-site2Same as aboveAAAACCAACAGTTCTTCTTT (SEQ ID NO: 614)TGGAT-site1Same as aboveAAAGTACTCTATAGGGTGAT (SEQ ID NO: 615)TGGAT-site2Same as aboveATGTACAACTTTAGTAGATA (SEQ ID NO: 616)TTGCT-site1Same as aboveAACATGGAAAGTATAGGATG (SEQ ID NO: 617)TTGCT-site2Same as aboveATTTCCAGCTAGAATATAAA (SEQ ID NO: 618)TTGGT-site1Same as aboveAGAATGGCAGTGCAATACGT (SEQ ID NO: 619)TTGGT-site2Same as aboveACACTAGAAAAGGATAAAAC (SEQ ID NO: 620)

[1232] The experimental results are shown in Figures 37 to 38. As a result of the experiment, it was confirmed that while the base editing system containing dead-TnpB did not operate in the target gene region containing the extended TAM (5'-TNGAT-3' or 5'-TTGVT-3'), the base editing system containing dead-enTnpB exhibited gene editing activity even in the target gene region containing the extended TAM.

[1233] Experimental Example 5.3. Cell Line Change Experiment

[1234] Experiments were conducted to verify whether the base editing system according to Experimental Example 5.1 operates regardless of the type of target cell. Specifically, using K562 cells and iPSC cells, gene editing efficiency was verified according to Experimental Example 1.5 using the base editor protein of Table 12 and the reRNA of the following table.

[1235] LabelScaffoldGuide Domain (RNA)HEK1-1SEQ ID NO: 7TCAAGTACCACCAGTTTTAT (SEQ ID NO: 86)EMX1-2SEQ ID NO: 7GCATTTCTGTTTTTAATTTAT (SEQ ID NO: 84)Site9SEQ ID NO: 7AGATGATGTTTCCACACATA (SEQ ID NO: 100)

[1236] The experimental results are shown in FIGS. 39 to 40. As a result of the experiment, it was confirmed that the base editing system of Experimental Example 5.1 exhibited the intended nucleobase editing effect regardless of the type of target cell.

[1237] Experimental Example 5.4. Verification of Base Editing System Effect #2

[1238] To determine whether base editing systems of various configurations could edit genes, the base editing systems in the following table were configured and the gene editing efficiency was verified according to Experimental Example 1.4.

[1239] LabelConstructSEQID NOTnpB_CBE(APOBEC1)_16bpNLS-APOBEC1-denTnpB-NLS-UGI-NLS623TnpB_CBE(APOBEC1)_20bpNLS-APOBEC1-denTnpB-NLS-UGI -NLS623TnpB_CBE(BE4max)_16bpNLS-BE4max-denTnpB-UGI-UGI-NLS624TnpB_CBE(BE4max)_20bpNLS-BE4max-denTnpB-UGI-UGI -NLS624TnpB_CBE(AncBE4max)_16bpNLS-AncBE4max-denTnpB-UGI-UGI-NLS625TnpB_CBE(AncBE4max)_20bpNLS-AncBE4max-de nTnpB-UGI-UGI-NLS625TnpB_ABE(TadA8e)_16bpNLS-TadA8e-dTnpB-NLS626TnpB_ABE(TadA8e)_20bpNLS-TadA8e-dTnpB-NLS626

[1240] The composition of the reRNA used in the experiment is as follows:

[1241] LabelScaffoldGuide Domain (RNA)EMX1-1SEQ ID NO: 7GTGATGGGAGCCCTTCTTCT (SEQ ID NO: 83)EMX1-2Same as aboveGCATTTCTGTTTTTAATTTAT (SEQ ID NO: 84)EMX1-1(16bp)Same as aboveGTGATGGGAGCCCTTC (SEQ ID NO: 621)EMX1-2(16bp)Same as aboveGCATTTCTGTTTTTAAT (SEQ ID NO: 622)

[1242] The experimental results are shown in Figure 41. The experimental results confirmed that gene editing was successful even when the base editing system was constructed using various deaminases.

[1243] Experimental Example 5.5. Identification of Clinically Significant Targets for Base Editing Systems

[1244] Using pathological SNV data listed in the ClinVar database, targets expected to show therapeutic effects using a dead-enTnpB-based base editing system were selected in-silico.

[1245] Specifically, in the case of enTnpB-CBE:

[1246] Recognize the TAM of 5'-TTGAT-3', 5'-TNGAT-3', or 5'-TTGVT-3', and

[1247] Condition to change C to T in the 2nd to 7th positions relative to the Protoss spacer,

[1248] And for enTnpB-ABE:

[1249] Recognize the TAM of 5'-TTGAT-3', 5'-TNGAT-3', or 5'-TTGVT-3', and

[1250] It was confirmed under the condition that A, included in the 2nd to 7th position relative to the Protoss spacer, is changed to G.

[1251] The experimental results are schematically illustrated in Figure 42. The results confirmed that a total of 305 SNVs could be corrected using the enTnpB-based base editing system, of which 71 were correctable with CBE and 234 with ABE. Additionally, only some of these (25.35% for CBE and 37.36% for ABE) were identified as targets that could also be corrected with the SpCas9-based base editing system. These results suggest that there are a significant number of pathological SNVs for which therapeutic effects can be expected using the enTnpB-based base editing system.

[1252] Experimental Example 6. Introduction of a stabilization motif into reRNA

[1253] Experimental Example 6.1. Verification of Indel Introduction Efficiency

[1254] The inventors of this application investigated whether introducing a stabilization motif into reRNA contributes to increasing the efficiency of indel introduction. The composition used in the experiment is shown in the following table:

[1255] Label5'-end Stable MotifLinker-1ScaffoldGuide DomainLinker-23'-end Stable Motifω*RNA- tGGTGGCTGCGGGAATCTCAGACACCTTAAACGCTCATGGAGGCTATgaaaATGGTCTGCGAAGTGAGAATCACGCGACTTTAGTCGTGTGAGGTTCAA (SEQ ID NO: 7)TCAAGTACCACCAGTTTTAT (SEQ ID NO: 86)--5'evopreQ1-ω*RNATTGACGCGGTTCTATCTAGTTACGCGTTAAACCAACTAGAAA (SEQ ID NO: 21)-SEQ ID NO: 7SEQ ID NO: 86--5'evopreQ1-linker-ω*RNASEQ ID NO: 21CCTCTTCT (SEQ ID NO: 23)SEQ ID NO: 7SEQ ID NO: 86--5'-mpknot-ω*RNAGGGTCAGGAGCCCCCCCCCTGAACCCAGGATAACCCTCAAAGTCGGGGGGCAACCC (SEQ ID NO: 22)-SEQ ID NO: 7SEQ ID NO: 86--5'-mpknot-linker-ω*RNASEQ ID NO: 22SEQ ID NO: 23SEQ ID NO: 7SEQ ID NO: 86--3'evopreQ1-ω*RNA SEQ ID NO: 7SEQ ID NO: 86-SEQ ID NO: 213'evopreQ1-linker-ω*RNA SEQ ID NO: 7SEQ ID NO: 86SEQ ID NO: 23SEQ ID NO: 213'-mpknot-ω*RNA SEQ ID NO: 7SEQ ID NO: 86-SEQ ID NO: 223'-mpknot-linker-ω*RNA SEQ ID NO: 7SEQ ID NO: 86SEQ ID NO: 23SEQ ID NO: 22

[1256] The above configuration is schematically illustrated in Fig. 43. A plasmid encoding the engineered TnpB system of the above configuration and a plasmid encoding reRNA, respectively, were prepared according to Experimental Example 1.1 and transfected into HEK293T cells cultured according to Experimental Example 1.2. Subsequently, targeted deep sequencing was performed according to Experimental Example 1.3 to measure the frequency of indel introduction.

[1257] The experimental results are shown in Fig. 44.

[1258] Experimental results confirmed that applying a stabilization motif to reRNA generally significantly increased the frequency of indel introduction.

[1259] Experimental Example 6.2. Verification of Base Editing Effect

[1260] The inventors of this application sought to verify whether introducing a stabilization motif into reRNA increases base editing efficiency. Specifically, experiments were conducted using a base editing system with the following configuration:

[1261] LabelConstructSEQ ID NOenTnpB-CBENLS-APOBEC1-denTnpB-UGI-UGI-NLS78enTnpB-ABENLS-TadA8e-denTnpB-NLS75

[1262] A plasmid configured to express the above base editing system was constructed according to Experimental Example 1.1, and HEK293T cells were transfected according to Experimental Example 1.2. Subsequently, target deep sequencing was performed according to Experimental Example 1.3 to confirm whether the intended base editing occurred in the TAM and protospacer regions of the target nucleic acid.

[1263] The experimental results are shown in Figures 45 to 48.

[1264] Experimental results confirmed that introducing a stabilization motif into reRNA significantly increased base editing efficiency. Similarly, there was no change in the correction window, and only the correction efficiency increased.

[1265] Experimental Example 7. Verification of Multiplexing Feasibility

[1266] Experimental Example 7.1. Verification of the feasibility of introducing multiple target indels through multiplexing

[1267] The inventors of this application verified whether the engineered TnpB system of this specification could be utilized in a multiplexing gene editing method. Specifically, they designed nucleic acids of the following structures in which the target nucleic acids encode three different types of reRNAs:

[1268] scaffold-1Guide Domain-1scaffold-2Guide Domain-2scaffold-3Guide Domain-3SEQ ID NO: 7SEQ ID NO:86SEQ ID NO: 7SEQ ID NO: 93SEQ ID NO: 7SEQ ID NO: 88

[1269] The process of the above nucleic acid and the resulting reRNA self-maturation is schematically illustrated in Fig. 49.

[1270] The experimental procedure is schematically illustrated in Fig. 50. According to Experimental Example 1.1, a vector containing the nucleic acids of the table above and a plasmid expressing wild-type TnpB protein (SEQ No. 1) or engineered TnpB protein (SEQ No. 2) were constructed, and HEK293T cells were transfected according to Experimental Example 1.2. In addition, wild-type reRNA (SEQ Nos. 5 to 7) or modified reRNA with some sequences removed (SEQ No. 6) were used in the experiment. Subsequently, target deep sequencing was performed according to Experimental Example 1.3 to confirm whether an indel was introduced into the target nucleic acid.

[1271] The experimental results are shown in Fig. 51.

[1272] Experimental results confirmed that multiplexing gene editing is possible through self-cleavage and maturation without the need to introduce separate terminators or spacers between the nucleic acids encoding each reRNA.

[1273] Experimental Example 7.2. Verification of the feasibility of multiple target base editing through multiplexing

[1274] The inventors of this application verified whether the engineered TnpB system of this specification could be utilized in a multiplexing base editing method. A nucleic acid encoding reRNA was designed according to Table 16 of Experimental Example 7.1.

[1275] The composition of the base editor protein used in the experiment is as follows:

[1276] LabelConstructSEQID NOTnpB-CBENLS-APOBEC1-dTnpB-UGI-UGI-NLS77enTnpB-CBENLS-APOBEC1-denTnpB-UGI-UGI-NLS78TnpB-ABENLS-TadA8e-dTnpB-NLS74enTnpB-ABENLS-TadA8e-denTnpB-NLS75

[1277] According to Experimental Example 1.1, a plasmid expressing the nucleic acid encoding the above reRNA and the base editor protein of the table above was constructed, and HEK293T cells were transfected according to Experimental Example 1.2. Subsequently, target deep sequencing was performed according to Experimental Example 1.3 to confirm whether an indel was introduced into the target nucleic acid. In addition, wild-type reRNA (SEQ ID NOs 5 to 7) or modified reRNA with some sequences removed (SEQ ID NO. 6) was used in the experiment. The experimental results are shown in Figures 52 to 53. As a result of the experiment, it was confirmed that multiplexing base editing is possible through self-cleavage and maturation occurring between the nucleic acids encoding each reRNA, even without introducing separate terminators or spacers.

[1278] Experimental Example 7.3. Verification of effects according to composition order

[1279] When constructing an expression vector by linking three different types of reRNAs, experiments were conducted to determine whether the gene editing efficiency varied depending on the order in which each reRNA was positioned.

[1280] Specifically, reRNA vectors containing all possible sequence combinations of reRNAs targeting HEK1-1, Site2, and RUNX1-2 were constructed (a total of 6 types), and these were used in the experiments of Experimental Example 7.1 and Experimental Example 7.2. This is schematically illustrated in the left area of ​​FIGS. 54 to 56.

[1281] The experimental results are shown in Figures 54 to 56. The experimental results showed that the reRNA placed closest to the promoter within the vector tended to perform gene editing most efficiently, but overall, targets placed further back were also edited without issues.

[1282] Experimental Example 8. Confirmation of in-vivo mouse gene editing effect

[1283] Experimental Example 8.1. Experimental Method and Materials

[1284] Experiments were conducted to verify whether the engineered TnpB protein (enTnpB) of this specification exhibits the intended gene editing effect when delivered to mice.

[1285] Specifically, the enTnpB system, which targets the mouse Pcsk9 gene or targets both Pcsk9 and Angptl3 genes simultaneously (multiplex), was loaded onto an AAV vector and injected via the mouse tail vein to verify whether a knockout effect of the corresponding genes occurred in the actual liver. The detailed experimental method is as follows:

[1286] [Prepare mouse]

[1287] 8-week-old male C57BL / 6J mice were prepared.

[1288] [Building AAV Vectors]

[1289] Two single-stranded AAV9 (ssAAV9) vectors were designed to encode the enTnpB system, one to target only Pcsk9 and the other to target Pcsk9 and Angptl3 simultaneously. Each construct included an enTnpB coding sequence driven by a P3 promoter and a reRNA cassette driven by a U6 promoter, said reRNA cassette including a spacer targeting Pcsk9 or a multiplex spacer targeting Pcsk9 (SEQ ID NO: 528) and Angptl3 (SEQ ID NO: 529).

[1290] The AAV vector was commercially synthesized by VectorBuilder (Chicago, IL, USA) and packaged into ssAAV9 particles, and vector genome titers were determined by quantitative PCR.

[1291] The composition of each vector is as follows:

[1292] LabelPromotor for TnpBTnpB proteinreRNA #1 SpacerreRNA #2 SpacerPromotor for reRNAssAAV9-enTnpB-reRNA (Pcsk9)P3SEQ ID NO: 2SEQ ID NO: 528-U6ssAAV9-enTnpB-multi-reRNA (Pcsk9+Angptl3)P3SEQ ID NO: 2SEQ ID NO: 528SEQ ID NO: 529U6

[1293] The structure of each AAV vector is schematically shown in Fig. 57. In addition, the target nucleic acid of the reRNA targeting Pcsk9 is shown in Fig. 59, and the target nucleic acid of the reRNA targeting Angptl3 is shown in Fig. 60.

[1294] [TnpB System Delivery via AAV Vector]

[1295] 5.0 × 10 based on body weight into the mouse tail vein prepared above 13 The AAV vector prepared above was injected at a dose of vg / kg. Three mice were used per experimental group. Six weeks after injection, blood samples were collected from each mouse, and the mice were sacrificed to isolate liver tissue.

[1296] [Sample Analysis]

[1297] Genomic DNA extracted from liver tissue was subjected to targeted deep sequencing to quantify indel frequencies at target sites. Total RNA was isolated from liver tissue, and gene expression levels of PCSK9 and Angptl3 were analyzed by quantitative real-time PCR (qPCR) following reverse transcription. Serum PCSK9 and ANGPTL3 protein levels were measured using commercially available enzyme-linked immunosorbent assay (ELISA) kits according to the manufacturer's instructions. Serum LDL cholesterol, triglycerides, alanine aminotransferase (ALT), and aspartate aminotransferase (AST) levels were analyzed using an automated clinical chemistry analyzer. Data are presented as mean ± standard error (SEM).

[1298] Experimental Example 8.2 Experimental results for Pcsk9 single target vector

[1299] The experimental results are shown in Figures 61 to 66.

[1300] Experimental results confirmed that an indel introduction effect was observed in the Psck9 gene of the liver cells of the experimental group (PCSK9-KO) mice, and as a result, the mice in the experimental group showed a Psck9 expression rate (including both mRNA and protein) of 0.5 times or less compared to the negative control (UNT) mice (Fig. 61).

[1301] In addition, the blood LDL concentration of the experimental group mice was significantly lower than that of the negative control group mice, and there were no significant differences in triglycerides, ALT, and AST (Fig. 65).

[1302] Meanwhile, tissue examination results confirmed that the indel introduction effect of Pcsk9 was specific only in liver tissue (Fig. 66). This is believed to be because 1) the ssAAV9 serotype AAV is characterized by being delivered specifically to liver cells, and 2) the P3 promoter also operates specifically to the liver.

[1303] Synthesizing the above results, it can be seen that when a TnpB system targeting Pcsk9 is loaded into a vector and injected, Pcsk9 in the liver can be edited and knocked out, and this effect can be controlled to occur specifically in liver tissue.

[1304] Experimental Example 8.3 Experimental results for Pcsk9 and Angptl3 multiple target vectors

[1305] The experimental results are shown in Figures 67 to 69.

[1306] Experimental results confirmed that indel introduction effects were observed in the Psck9 and Angptl3 genes of the experimental group (Multi-KO) mouse liver cells, and as a result, the expression rates of Psck9 mRNA, Psck9 protein, Angptl3 mRNA, and Angptl3 protein in the experimental group mice were reduced compared to the negative control (UNT) mice (Fig. 67).

[1307] In addition, LDL and triglyceride concentrations in the blood of the experimental group mice were significantly reduced, while ALT and AST concentrations showed no significant difference (Fig. 68). Meanwhile, histological examination revealed that indels occurred in both Pcsk9 and Angptl3 only in liver tissue. This is attributed to the function of the P3 promoter as previously mentioned.

[1308] Synthesizing the above results, it can be seen that when a TnpB system targeting both Pcsk9 and Angptl3 is loaded into a vector in a multiplex form and injected, both Pcsk9 and Angptl3 genes in the liver can be edited and knocked out, and this effect can be regulated to occur specifically in liver tissue.

[1309] Experimental Example 9. Primer sequences used in the experiment

[1310] The primer sequence information used in the experiment is summarized in the table below.

[1311] LabelDescriptionSEQ ID NOFP-TnpB-V1Forward Primer for Site-Directed Mutage, TnpB Variant (N4R)105FP-TnpB-V2Forward Primer for Site-Directed Mutage, TnpB Variant (K5R)106FP-TnpB-V1Forward Primer for Site-Directed Mutage, Primer for Site-Directed Mutage TnpB Variant (A6R)107FP-TnpB-V4Forward Primer for Site-Directed Mutage, TnpB Variant (V8R)108FP-TnpB-V5Forward Primer for Site-Directed Mutage, TnpB Variant (Y12R)109FPB-Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (A15R)110FP-TnpB-V7Forward Primer for Site-Directed Mutage, TnpB Variant (Y32R)111FP-TnpB-V8Forward Primer for Site-Directed Mutage, TnpB Variant (H34R)112FP-TnpB-V9Forward Primer for Site-Directed Mutage, TnpB Variant (I40R)113FP-TnpB-V10Forward Primer for Site-Directed Mutage, TnpB Variant (Y43R)114FPV-TnpB1 for Site-Directed Mutage Mutage, TnpB Variant (K48R)115FP-TnpB-V12Forward Primer for Site-Directed Mutage, TnpB Variant (G49R)116FP-TnpB-V13Forward Primer for Site-Directed Mutage, TnpB Variant (Y52R)117FP-TnpB-V14Forward Primerfor Site-Directed Mutage, TnpB Variant (S56R)118FP-TnpB-V15Forward Primer for Site-Directed Mutage, TnpB Variant (S57R)119FP-TnpB-V16Forward Primer for Site-Directed Mutage, TnpB Variant (T60R)120FP-TnpB-V17Forward Primer for Site-Directed Mutage, TnpB Variant (K63R)121FP-TnpB-V18Forward Primer for Site-Directed Mutage, TnpB Variant (Q64R)122FP-TnpB-V17Forward Primer for Site-Directed Mutage Mutage, TnpB Variant (A65R)123FP-TnpB-V20Forward Primer for Site-Directed Mutage, TnpB Variant (S72R)124FP-TnpB-V21Forward Primer for Site-Directed Mutage, TnpB Variant (K76R)125FP-TnpB-V22Forward Primer for Site-Directed Mutage, TnpB Variant (F77R)126FP-TnpB-V23Forward Primer for Site-Directed Mutage, TnpB Variant (Q80R)127FP-TnpB-V22Forward Primer for Site-Directed Mutage Mutage, TnpB Variant (K84R)128FP-TnpB-V25Forward Primer for Site-Directed Mutage, TnpB Variant (N85R)129FP-TnpB-V26Forward Primer for Site-Directed Mutage, TnpB Variant (E87R)130FP-TnpB-V27Forward Primer for Site-Directed Mutage, TnpB Variant(T88R)131FP-TnpB-V28Forward Primer for Site-Directed Mutage, TnpB Variant (A89R)132FP-TnpB-V29Forward Primer for Site-Directed Mutage, TnpB Variant (N92R)133FP-TnpB-V28Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (F93R)134FP-TnpB-V31Forward Primer for Site-Directed Mutage, TnpB Variant (F94R)135FP-TnpB-V32Forward Primer for Site-Directed Mutage, TnpB Variant (V97R)136FP-TnpB-V33Forward Primer for Site-Directed Mutage, TnpB Variant (V104R)137FP-TnpB-V34Forward Primer for Site-Directed Mutage, TnpB Variant (G105R)138B-FPV-Forward Primer for 35 Site-Directed Mutage, TnpB Variant (F106R)139FP-TnpB-V36Forward Primer for Site-Directed Mutage, TnpB Variant (P107R)140FP-TnpB-V37Forward Primer for Site-Directed Mutage, TnpB Variant (T114R)141FP-TnpB-V38Forward Primer for Site-Directed Mutage, TnpB Variant (G115R)142FP-TnpB-V39Forward Primer for Site-Directed Mutage, TnpB Variant (S117R)143F-FPVT-Forward Primer for 40B Variant Site-Directed Mutage, TnpB Variant (Q121R)144FP-TnpB-V41ForwardPrimer for Site-Directed Mutage, TnpB Variant (T123R)145FP-TnpB-V42Forward Primer for Site-Directed Mutage, TnpB Variant (N124R)146FP-TnpB-V43Forward Primer for Site-Directed Mutage, TnpB Vageriant (N125R)147FP-TnpB-V44Forward Primer for Site-Directed Mutage, TnpB Variant (N126R)148FP-TnpB-V45Forward Primer for Site-Directed Mutage, TnpB Variant (Q128R)149FPB-V44Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (K135R)150FP-TnpB-V47Forward Primer for Site-Directed Mutage, TnpB Variant (P137R)151FP-TnpB-V48Forward Primer for Site-Directed Mutage, TnpB Variant (K138R)152FP-TnpB-V49Forward Primer for Site-Directed Mutage, TnpB Variant (G140R)153FP-TnpB-V50Forward Primer for Site-Directed Mutage, TnpB Variant (K145R)154B-FVT-Forward Primer for 515 Site-Directed Mutage, TnpB Variant (G146R)155FP-TnpB-V52Forward Primer for Site-Directed Mutage, TnpB Variant (Q148R)156FP-TnpB-V53Forward Primer for Site-Directed Mutage, TnpB Variant (L155R)157FP-TnpB-V54Forward Primer for Site-DirectedMutage, TnpB Variant (N156R)158FP-TnpB-V55Forward Primer for Site-Directed Mutage, TnpB Variant (T158R)159FP-TnpB-V56Forward Primer for Site-Directed Mutage, TnpB Variant (S170R)160FP-TnpB-V57Forward Primer for Site-Directed Mutage, TnpB Variant (L172R)161FP-TnpB-V58Forward Primer for Site-Directed Mutage, TnpB Variant (P178R)162FPB-V57Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (A182R)163FP-TnpB-V60Forward Primer for Site-Directed Mutage, TnpB Variant (V204R)164FP-TnpB-V61Forward Primer for Site-Directed Mutage, TnpB Variant (Y215R)165FP-TnpB-V62Forward Primer for Site-Directed Mutage, TnpB Variant (K224R)166FP-TnpB-V63Forward Primer for Site-Directed Mutage, TnpB Variant (Q226R)167-FPB-V62Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (Q227R)168FP-TnpB-V65Forward Primer for Site-Directed Mutage, TnpB Variant (T228R)169FP-TnpB-V66Forward Primer for Site-Directed Mutage, TnpB Variant (S230R)170FP-TnpB-V67Forward Primer for Site-Directed Mutage, TnpB Variant(K233R)171FP-TnpB-V68Forward Primer for Site-Directed Mutage, TnpB Variant (K234R)172FP-TnpB-V69Forward Primer for Site-Directed Mutage, TnpB Variant (S236R)173-FpB-V68Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (A237R)174FP-TnpB-V71Forward Primer for Site-Directed Mutage, TnpB Variant (Y239R)175FP-TnpB-V72Forward Primer for Site-Directed Mutage, TnpB Variant (G240R)176FP-TnpB-V73Forward Primer for Site-Directed Mutage, TnpB Variant (K241R)177FP-TnpB-V74Forward Primer for Site-Directed Mutage, TnpB Variant (K243R)178FPB-Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (L246R)179FP-TnpB-V76Forward Primer for Site-Directed Mutage, TnpB Variant (A247R)180FP-TnpB-V77Forward Primer for Site-Directed Mutage, TnpB Variant (H250R)181FP-TnpB-V78Forward Primer for Site-Directed Mutage, TnpB Variant (V254R)182FP-TnpB-V79Forward Primer for Site-Directed Mutage, TnpB Variant (N255R)183F-FPB-V78Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant(K256R)184FP-TnpB-V81Forward Primer for Site-Directed Mutage, TnpB Variant (D259R)185FP-TnpB-V82Forward Primer for Site-Directed Mutage, TnpB Variant (H262R)186FB-V81Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (K263R)187FP-TnpB-V84Forward Primer for Site-Directed Mutage, TnpB Variant (L264R)188FP-TnpB-V85Forward Primer for Site-Directed Mutage, TnpB Variant (T266R)189FP-TnpB-V86Forward Primer for Site-Directed Mutage, TnpB Variant (L280R)190FP-TnpB-V87Forward Primer for Site-Directed Mutage, TnpB Variant (K281R)191FB-V86Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (P282R)192FP-TnpB-V89Forward Primer for Site-Directed Mutage, TnpB Variant (D283R)193FP-TnpB-V90Forward Primer for Site-Directed Mutage, TnpB Variant (A292R)194FP-TnpB-V91Forward Primer for Site-Directed Mutage, TnpB Variant (L293R)195FP-TnpB-V92Forward Primer for Site-Directed Mutage, TnpB Variant (S296R)196F-FBVT-Forward Primer for TnpB Variant (L293R) Site-Directed Mutage, TnpB Variant(G299R)197FP-TnpB-V94Forward Primer for Site-Directed Mutage, TnpB Variant (W300R)198FP-TnpB-V95Forward Primer for Site-Directed Mutage, TnpB Variant (G301R)199FB-V94Forward Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (E302R)200FP-TnpB-V97Forward Primer for Site-Directed Mutage, TnpB Variant (Y309R)201FP-TnpB-V98Forward Primer for Site-Directed Mutage, TnpB Variant (K310R)202FP-TnpB-V99Forward Primer for Site-Directed Mutage, TnpB Variant (W313R)203FP-TnpB-V100Forward Primer for Site-Directed Mutage, TnpB Variant (L317R)TnpB-V99Forward Primer for Site-Directed Mutage (L317R) Site-Directed Mutage, TnpB Variant (P323R)205FP-TnpB-V102Forward Primer for Site-Directed Mutage, TnpB Variant (P339R)206FP-TnpB-V103Forward Primer for Site-Directed Mutage, TnpB Variant (T350R)207FP-TnpB-V104Forward Primer for Site-Directed Mutage, TnpB Variant (V374R)208RP-TnpB-V1Reverse Primer for Site-Directed Mutage, TnpB Variant (N4R)209RP-TnpB-V104Forward Primer for Site-Directed Mutage Mutage, TnpB Variant(K5R)210RP-TnpB-V3Reverse Primer for Site-Directed Mutage, TnpB Variant (A6R)211RP-TnpB-V4Reverse Primer for Site-Directed Mutage, TnpB Variant (V8R)212RP-TnpB-V5Reverse Primer for Site-Directed Mutage, TnpB Variant Variant (Y12R)213RP-TnpB-V6Reverse Primer for Site-Directed Mutage, TnpB Variant (A15R)214RP-TnpB-V7Reverse Primer for Site-Directed Mutage, TnpB Variant (Y32R)215RP-TnpB-V6D Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (H34R)216RP-TnpB-V9Reverse Primer for Site-Directed Mutage, TnpB Variant (I40R)217RP-TnpB-V10Reverse Primer for Site-Directed Mutage, TnpB Variant (Y43R)-TnpB-V9Reverse Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (K48R)219RP-TnpB-V12Reverse Primer for Site-Directed Mutage, TnpB Variant (G49R)220RP-TnpB-V13Reverse Primer for Site-Directed Mutage, TnpB Variant (Y52R)221RP-TnpB-V14Reverse Primer for Site-Directed Mutage, TnpB Variant (S56R)222RP-TnpB-V15Reverse Primer for Site-Directed Mutage, TnpB Variant (S57R)223RP-TnpB-V14Reverse Primer for Site-Directed MutageSite-Directed Mutage, TnpB Variant (T60R)224RP-TnpB-V17Reverse Primer for Site-Directed Mutage, TnpB Variant (K63R)225RP-TnpB-V18Reverse Primer for Site-Directed Mutage, TnpB Variant (Q64R)226RP-TnpB-V19Reverse Primer for Site-Directed Mutage, TnpB Variant (A65R)227RP-TnpB-V20Reverse Primer for Site-Directed Mutage, TnpB Variant (S72R)228RP-TnpB-V19Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (K76R)229RP-TnpB-V22Reverse Primer for Site-Directed Mutage, TnpB Variant (F77R)230RP-TnpB-V23Reverse Primer for Site-Directed Mutage, TnpB Variant (Q8RPB-V22R) for Site-Directed Mutage, TnpB Variant (K84R)232RP-TnpB-V25Reverse Primer for Site-Directed Mutage, TnpB Variant (N85R)233RP-TnpB-V26Reverse Primer for Site-Directed Mutage, TnpB Variant (E87R)234RP-TnpB-V27Reverse Primer for Site-Directed Mutage, TnpB Variant (T88R)235RP-TnpB-V28Reverse Primer for Site-Directed Mutage, TnpB Variant (A89R)236RP-TnpB-V29Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant(N92R)237RP-TnpB-V30Reverse Primer for Site-Directed Mutage, TnpB Variant (F93R)238RP-TnpB-V31Reverse Primer for Site-Directed Mutage, TnpB Variant (F94R)239RP-TnpB-V30Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (V97R)240RP-TnpB-V33Reverse Primer for Site-Directed Mutage, TnpB Variant (V104R)241RP-TnpB-V34Reverse Primer for Site-Directed Mutage, TnpB Variant (G15R)240RP-TnpB-V33Reverse Primer Primer for Site-Directed Mutage, TnpB Variant (F106R)243RP-TnpB-V36Reverse Primer for Site-Directed Mutage, TnpB Variant (P107R)244RP-TnpB-V37Reverse Primer for Site-Directed Mutage, TnpB Variant (T114R)245RP-TnpB-V38Reverse Primer for Site-Directed Mutage, TnpB Variant (G115R)246RP-TnpB-V39Reverse Primer for Site-Directed Mutage, TnpB Variant (S117R)247-TnpB-TnpB-V38Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (Q121R)248RP-TnpB-V41Reverse Primer for Site-Directed Mutage, TnpB Variant (T123R)249RP-TnpB-V42Reverse Primer for Site-Directed Mutage, TnpB Variant (N124R)250RP-TnpB-V43ReversePrimer for Site-Directed Mutage, TnpB Variant (N125R)251RP-TnpB-V44Reverse Primer for Site-Directed Mutage, TnpB Variant (N126R)252RP-TnpB-V45Reverse Primer for Site-Directed Mutage, TnpB Variant (Q128R)253RP-TnpB-V46Reverse Primer for Site-Directed Mutage, TnpB Variant (K135R)254RP-TnpB-V47Reverse Primer for Site-Directed Mutage, TnpB Variant (P137R)255-TnpB-TnpB-V46Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (K138R)256RP-TnpB-V49Reverse Primer for Site-Directed Mutage, TnpB Variant (G140R)257RP-TnpB-V50Reverse Primer for Site-Directed Mutage, TnpB Variant (K145R)258RP-TnpB-V51Reverse Primer for Site-Directed Mutage, TnpB Variant (G146R)259RP-TnpB-V52Reverse Primer for Site-Directed Mutage, TnpB Variant (Q148R)260-TnpB-TnpB-V51Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (L155R)261RP-TnpB-V54Reverse Primer for Site-Directed Mutage, TnpB Variant (N156R)262RP-TnpB-V55Reverse Primer for Site-Directed Mutage, TnpB Variant (T158R)263RP-TnpB-V56Reverse Primer for Site-DirectedMutage, TnpB Variant (S170R)264RP-TnpB-V57Reverse Primer for Site-Directed Mutage, TnpB Variant (L172R)265RP-TnpB-V58Reverse Primer for Site-Directed Mutage, TnpB Variant (P178R)266RP-TnpB-V59Reverse Primer for Site-Directed Mutage, TnpB Variant (A182R)267RP-TnpB-V60Reverse Primer for Site-Directed Mutage, TnpB Variant (V204R)268RP-TnpB-TnpB-V60Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (Y215R)269RP-TnpB-V62Reverse Primer for Site-Directed Mutage, TnpB Variant (K224R)270RP-TnpB-V63Reverse Primer for Site-Directed Mutage, TnpB Variant (Q226R)271RP-TnpB-V64Reverse Primer for Site-Directed Mutage, TnpB Variant (Q227R)272RP-TnpB-V65Reverse Primer for Site-Directed Mutage, TnpB Variant (T228R)273-TnpB-V64Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (S230R)274RP-TnpB-V67Reverse Primer for Site-Directed Mutage, TnpB Variant (K233R)275RP-TnpB-V68Reverse Primer for Site-Directed Mutage, TnpB Variant (K234R)276RP-TnpB-V69Reverse Primer for Site-Directed Mutage, TnpB Variant(S236R)277RP-TnpB-V70Reverse Primer for Site-Directed Mutage, TnpB Variant (A237R)278RP-TnpB-V71Reverse Primer for Site-Directed Mutage, TnpB Variant (Y239R)279-TnpB-V70Reverse Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (G240R)280RP-TnpB-V73Reverse Primer for Site-Directed Mutage, TnpB Variant (K241R)281RP-TnpB-V74Reverse Primer for Site-Directed Mutage, TnpB Variant (K243R)282RP-TnpB-V75Reverse Primer for Site-Directed Mutage, TnpB Variant (L246R)283RP-TnpB-V76Reverse Primer for Site-Directed Mutage, TnpB Variant (A247R)284-TnpB-TnpB-V75Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (H250R)285RP-TnpB-V78Reverse Primer for Site-Directed Mutage, TnpB Variant (V254R)286RP-TnpB-V79Reverse Primer for Site-Directed Mutage, TnpB Variant (N255R)287RP-TnpB-V80Reverse Primer for Site-Directed Mutage, TnpB Variant (K256R)288RP-TnpB-V81Reverse Primer for Site-Directed Mutage, TnpB Variant (D259R)289-TnpB-TnpB-V80Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant(H262R)290RP-TnpB-V83Reverse Primer for Site-Directed Mutage, TnpB Variant (K263R)291RP-TnpB-V84Reverse Primer for Site-Directed Mutage, TnpB Variant (L264R)292-TnpB-V83Reverse Primer for Site-Directed Mutage Site-Directed Mutage, TnpB Variant (T266R)293RP-TnpB-V86Reverse Primer for Site-Directed Mutage, TnpB Variant (L280R)294RP-TnpB-V87Reverse Primer for Site-Directed Mutage, TnpB Variant (K281R)295RP-TnpB-V88Reverse Primer for Site-Directed Mutage, TnpB Variant (P282R)296RP-TnpB-V89Reverse Primer for Site-Directed Mutage, TnpB Variant (D283R)297-BRP-TnpB-V88Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant (A292R)298RP-TnpB-V91Reverse Primer for Site-Directed Mutage, TnpB Variant (L293R)299RP-TnpB-V92Reverse Primer for Site-Directed Mutage, TnpB Variant (S296R)300RP-TnpB-V93Reverse Primer for Site-Directed Mutage, TnpB Variant (G299R)301RP-TnpB-V94Reverse Primer for Site-Directed Mutage, TnpB Variant (W300R)302RP-TnpB-TnpB-V93Reverse Primer for Site-Directed Mutage Mutage, TnpB Variant(G301R)303RP-TnpB-V96Reverse Primer for Site-Directed Mutage, TnpB Variant (E302R)304RP-TnpB-V97Reverse Primer for Site-Directed Mutage, TnpB Variant (Y309R)305RP-TnpB-V98Reverse Primer for Site-Directed Mutage, TnpB Variant (K310R)306RP-TnpB-V99Reverse Primer for Site-Directed Mutage, TnpB Variant (W313R)307RP-TnpB-V100Reverse Primer for Site-Directed Mutage, TnpB Variant (L317R)308RP-TnpB-V101Reverse Primer for Site-Directed Mutage, TnpB Variant (P323R)309RP-TnpB-V102Reverse Primer for Site-Directed Mutage, TnpB Variant (P339R)310RP-TnpB-V103Reverse Primer for Site-Directed Mutage, TnpB Variant (T350R)311RP-TnpB-V104Reverse Primer for Site-Directed Mutage, TnpB Variant (V374R)312NGS-AGBL1-1st-FPNGS Forward Primer for AGBL1313NGS-EMX1-1st-FPNGS Forward Primer for EMX1314NGS-HPRT-1st-FPNGS Forward Primer for HPRT315NGS-HEK1-1-1st-FPNGS Forward Primer for HEK1-1316NGS-HEK1-3-1st-FPNGS Forward Primer for HEK1-3317NGS-RUNX1-1st-FPNGS Forward Primer for RUNX1318NGS-TTR-1st-FPNGS ForwardPrimer for TTR319NGS-TnpbB-site1-1st-FPNGS Forward Primer for TnpbB-site1320NGS-TnpB-site2-1st-FPNGS Forward Primer for TnpB-site2321NGS-TnpB-site3-1st-FPNGS Forward Primer for TnpB-site3322NGS-TnpB-site4,5-1st-FPNGS Forward Primer for TnpB-site4,5323NGS-TnpB-site6-1st-FPNGS Forward Primer for TnpB-site6324NGS-TnpB-site7-1st-FPNGS Forward Primer for TnpB-site7325NGS-TnpB-site8-1st-FPNGS Forward Primer for TnpB-site8326NGS-TnpB-site9-1st-FPNGS Forward Primer for TnpB-site9327NGS-TnpB-site10-1st-FPNGS Forward Primer for TnpB-site10328NGS-TnpB-site11,12-1st-FPNGS Forward Primer for TnpB-site11,12329NGS-TnpB-site13-1st-FPNGS Forward Primer for TnpB-site13330NGS-AGBL1-1st-RPNGS Reverse Primer for AGBL1331NGS-EMX1-1st-RPNGS Reverse Primer for EMX1332NGS-HPRT-1st-RPNGS Reverse Primer for HPRT333NGS-HEK1-1-1st-RPNGS Reverse Primer for HEK1-1334NGS-HEK1-3-1st-RPNGS Reverse Primer for HEK1-3335NGS-RUNX1-1st-RPNGS Reverse Primer for RUNX1336NGS-TTR-1st-RPNGS Reverse Primer forTTR337NGS-TnpbB-site1-1st-RPNGS Reverse Primer for TnpbB-site1338NGS-TnpB-site2-1st-RPNGS Reverse Primer for TnpB-site2339NGS-TnpB-site3-1st-RPNGS Reverse Primer for TnpB-site3340NGS-TnpB-site4,5-1st-RPNGS Reverse Primer for TnpB-site4,5341NGS-TnpB-site6-1st-RPNGS Reverse Primer for TnpB-site6342NGS-TnpB-site7-1st-RPNGS Reverse Primer for TnpB-site7343NGS-TnpB-site8-1st-RPNGS Reverse Primer for TnpB-site8344NGS-TnpB-site9-1st-RPNGS Reverse Primer for TnpB-site9345NGS-TnpB-site10-1st-RPNGS Reverse Primer for TnpB-site10346NGS-TnpB-site11,12-1st-RPNGS Reverse Primer for TnpB-site11,12347NGS-TnpB-site13-1st-RPNGS Reverse Primer for TnpB-site13348NGS-AGBL1-2nd-FP2nd NGS Forward Primer for AGBL1349NGS-EMX1-1-2nd-FP2nd NGS Forward Primer for EMX1-1350NGS-EMX1-2-2nd-FP2nd NGS Forward Primer for EMX1-2351NGS-HPRT-2nd-FP2nd NGS Forward Primer for HPRT352NGS-HEK1-1-2nd-FP2nd NGS Forward Primer for HEK1-1353NGS-HEK1-3-2nd-FP2nd NGS Forward Primer for HEK1-3354NGS-RUNX1-1-2nd-FP2nd NGS ForwardPrimer for RUNX1-1355NGS-RUNX1-2-2nd-FP2nd NGS Forward Primer for RUNX1-2356NGS-RUNX1-3-2nd-FP2nd NGS Forward Primer for RUNX1-3357NGS-TTR-2nd-FP2nd NGS Forward Primer for TTR358NGS-TnpbB-site1-2nd-FP2nd NGS Forward Primer for TnpbB-site1359NGS-TnpB-site2-2nd-FP2nd NGS Forward Primer for TnpB-site2360NGS-TnpB-site3-2nd-FP2nd NGS Forward Primer for TnpB-site3361NGS-TnpB-site4-2nd-FP2nd NGS Forward Primer for TnpB-site4362NGS-TnpB-site5-2nd-FP2nd NGS Forward Primer for TnpB-site5363NGS-TnpB-site6-2nd-FP2nd NGS Forward Primer for TnpB-site6364NGS-TnpB-site7-2nd-FP2nd NGS Forward Primer for TnpB-site7365NGS-TnpB-site8-2nd-FP2nd NGS Forward Primer for TnpB-site8366NGS-TnpB-site9-2nd-FP2nd NGS Forward Primer for TnpB-site9367NGS-TnpB-site10-2nd-FP2nd NGS Forward Primer for TnpB-site10368NGS-TnpB-site11-2nd-FP2nd NGS Forward Primer for TnpB-site11369NGS-TnpB-site12-2nd-FP2nd NGS Forward Primer for TnpB-site12370NGS-TnpB-site13-2nd-FP2nd NGS Forward Primer for TnpB-site13371NGS-AGBL1-2nd-RP2ndNGS Reverse Primer for AGBL1372NGS-EMX1-1-2nd-RP2nd NGS Reverse Primer for EMX1-1373NGS-EMX1-2-2nd-RP2nd NGS Reverse Primer for EMX1-2374NGS-HPRT-2nd-RP2nd NGS Reverse Primer for HPRT375NGS-HEK1-1-2nd-RP2nd NGS Reverse Primer for HEK1-1376NGS-HEK1-3-2nd-RP2nd NGS Reverse Primer for HEK1-3377NGS-RUNX1-1-2nd-RP2nd NGS Reverse Primer for RUNX1-1378NGS-RUNX1-2-2nd-RP2nd NGS Reverse Primer for RUNX1-2379NGS-RUNX1-3-2nd-RP2nd NGS Reverse Primer for RUNX1-3380NGS-TTR-2nd-RP2nd NGS Reverse Primer for TTR381NGS-TnpbB-site1-2nd-RP2nd NGS Reverse Primer for TnpbB-site1382NGS-TnpB-site2-2nd-RP2nd NGS Reverse Primer for TnpB-site2383NGS-TnpB-site3-2nd-RP2nd NGS Reverse Primer for TnpB-site3384NGS-TnpB-site4-2nd-RP2nd NGS Reverse Primer for TnpB-site4385NGS-TnpB-site5-2nd-RP2nd NGS Reverse Primer for TnpB-site5386NGS-TnpB-site6-2nd-RP2nd NGS Reverse Primer for TnpB-site6387NGS-TnpB-site7-2nd-RP2nd NGS Reverse Primer for TnpB-site7388NGS-TnpB-site8-2nd-RP2nd NGS Reverse Primer forTnpB-site8389NGS-TnpB-site9-2nd-RP2nd NGS Reverse Primer for TnpB-site9390NGS-TnpB-site10-2nd-RP2nd NGS Reverse Primer for TnpB-site10391NGS-TnpB-site11-2nd-RP2nd NGS Reverse Primer for TnpB-site11392NGS-TnpB-site12-2nd-RP2nd NGS Reverse Primer for TnpB-site12393NGS-TnpB-site13-2nd-RP2nd NGS Reverse Primer for TnpB-site13394NGS-AGBL1-2nd-F-Adaptor2nd NGS Forward Adaptor for AGBL1395NGS-AGBL1-2-2nd-F-Adaptor2nd NGS Forward Adaptor for AGBL1-2396NGS-EMX1-1-2nd-F-Adaptor2nd NGS Forward Adaptor for EMX1-1397NGS-EMX1-2-2nd-F-Adaptor2nd NGS Forward Adaptor for EMX1-2398NGS-HPRT-2nd-F-Adaptor2nd NGS Forward Adaptor for HPRT399NGS-HEK1-1-2nd-F-Adaptor2nd NGS Forward Adaptor for HEK1-1400NGS-HEK1-3-2nd-F-Adaptor2nd NGS Forward Adaptor for HEK1-3401NGS-RUNX1-1-2nd-F-Adaptor2nd NGS Forward Adaptor for RUNX1-1402NGS-RUNX1-2-2nd-F-Adaptor2nd NGS Forward Adaptor for RUNX1-2403NGS-RUNX1-3-2nd-F-Adaptor2nd NGS Forward Adaptor for RUNX1-3404NGS-TTR-2nd-F-Adaptor2nd NGS Forward Adaptor forTTR405NGS-TnpbB-site1-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpbB-site1406NGS-TnpB-site2-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site2407NGS-TnpB-site3-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site3408NGS-TnpB-site4-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site4409NGS-TnpB-site5-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site5410NGS-TnpB-site6-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site6411NGS-TnpB-site7-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site7412NGS-TnpB-site8-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site8413NGS-TnpB-site9-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site9414NGS-TnpB-site10-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site10415NGS-TnpB-site11-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site11416NGS-TnpB-site12-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site12417NGS-TnpB-site13-2nd-F-Adaptor2nd NGS Forward Adaptor for TnpB-site13418NGS-AGBL1-2nd-R-Adaptor2nd NGS Reverse Adaptor for AGBL1419NGS-AGBL1-2-2nd-R-Adaptor2ndNGS Reverse Adaptor for AGBL1-2420NGS-EMX1-1-2nd-R-Adaptor2nd NGS Reverse Adaptor for EMX1-1421NGS-EMX1-2-2nd-R-Adaptor2nd NGS Reverse Adaptor for EMX1-2422NGS-HPRT-2nd-R-Adaptor2nd NGS Reverse Adaptor for HPRT423NGS-HEK1-1-2nd-R-Adaptor2nd NGS Reverse Adaptor for HEK1-1424NGS-HEK1-3-2nd-R-Adaptor2nd NGS Reverse Adaptor for HEK1-3425NGS-RUNX1-1-2nd-R-Adaptor2nd NGS Reverse Adaptor for RUNX1-1426NGS-RUNX1-2-2nd-R-Adaptor2nd NGS Reverse Adaptor for RUNX1-2427NGS-RUNX1-3-2nd-R-Adaptor2nd NGS Reverse Adaptor for RUNX1-3428NGS-TTR-2nd-R-Adaptor2nd NGS Reverse Adaptor for TTR429NGS-TnpbB-site1-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpbB-site1430NGS-TnpB-site2-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site2431NGS-TnpB-site3-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site3432NGS-TnpB-site4-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site4433NGS-TnpB-site5-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site5434NGS-TnpB-site6-2nd-R-Adaptor2nd NGS Reverse Adaptor forTnpB-site6435NGS-TnpB-site7-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site7436NGS-TnpB-site8-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site8437NGS-TnpB-site9-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site9438NGS-TnpB-site10-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site10439NGS-TnpB-site11-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site11440NGS-TnpB-site12-2nd-R-Adaptor2nd NGS Reverse Adaptor for TnpB-site12441NGS-TnpB-site13-2nd-R-Adaptor2nd NGS Reverse Adapter for TnpB-site13442

[1312] By using the engineered TnpB protein disclosed in this specification, the base editor protein containing the engineered TnpB protein, and reRNA, intended gene editing can be induced with high efficiency by targeting the genome of a cell. In addition, since it exhibits gene editing efficiency comparable to the Streptococcus pyogenes-derived Cas9 system currently widely used for gene editing, it can be utilized in various fields of gene editing technology.

Claims

1. Engineered Deinococcus radiodurans ISDra2-derived TnpB protein (ISDra2 TnpB), Herein, the engineered ISDra2 TnpB protein comprises an amino acid sequence in which one to five variations selected from the following are introduced into the amino acid sequence of SEQ ID NO. 1: N4R; V8R; G49R; S57R; Q64R; S72R; K84R; Q128R; G140R; G146R; Q148R; V204R; T228R; K234R; A237R; Y239R; K243R; A247R; V254R; N255R; K256R; K263R; T266R; K281R; K310R; P339R; And T350R.

2. In the ISDra2 TnpB protein of claim 1, The above-mentioned engineered ISDra2 TnpB protein comprises the amino acid sequence of SEQ ID NO. 2 or SEQ ID NO.

3.

3. In any one of the ISDra2 TnpB proteins selected from claims 1 to 2, The above-mentioned engineered ISDra2 TnpB protein further contains one or more Nuclear Localization Signals (NLS) at the N-terminus, C-terminus, or both N-terminus and C-terminus.

4. A gene editing composition comprising the following: An engineered ISDra2 TnpB protein of any one of claims 1 to 3, or a nucleic acid encoding said engineered ISDra2 TnpB protein; and right-end element RNA (reRNA), or nucleic acid encoding said reRNA; Here, the reRNA comprises a scaffold and a spacer, and The above scaffold can form a complex with ISDra2-derived TnpB protein, and The above spacer can bind complementarily to a predetermined target nucleic acid.

5. In the gene editing composition of paragraph 4, The above reRNA includes the following: 5'-[Scaffold]-[Spacer]-[Linker]-[Stabilization Domain]-3' Here, the scaffold comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 5 to 7, and The above linker includes or is absent the nucleic acid sequence of SEQ ID NO. 23, and The above stabilization domain includes a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22, or is absent.

6. In any one of the gene editing compositions selected from paragraphs 4 to 5, The gene editing composition comprises the engineered ISDra2 TnpB protein and an RNA-guided nucleic acid cleavage complex to which the reRNA is bound.

7. In any one of the gene editing compositions selected from paragraphs 4 to 6, The gene editing composition comprises a nucleic acid encoding the engineered ISDra2 TnpB protein and a vector comprising a nucleic acid encoding the reRNA.

8. Intracellular target gene editing methods including the following: A process of delivering a gene editing composition selected from any one of claims 4 to 7 to the cell; Here, the spacer of the reRNA of the gene editing composition is configured to bind complementarily to a predetermined target nucleic acid included in the target gene, and By the above delivery, an RNA-guided nucleic acid cleavage complex bound to the engineered ISDra2 TnpB protein of the composition and reRNA is delivered into the cell, or an RNA-guided nucleic acid cleavage complex is formed within the cell, and Contact between the RNA-guided nucleic acid cleavage complex and the target gene is induced, and By the above contact, the target gene is edited.

9. Engineered ISDra2-derived TnpB protein (ISDra2 TnpB), Herein, the ISDra2 TnpB comprises the amino acid sequence of SEQ ID NO. 8; or The amino acid sequence comprises an amino acid sequence specified by applying one to six variations selected from the following to the amino acid sequence of SEQ ID NO. 8: N4R; V8R; G49R; S57R; Q64R; S72R; K84R; Q128R; G140R; G146R; Q148R; V204R; T228R; K234R; A237R; Y239R; K243R; A247R; V254R; N255R; K256R; K263R; T266R; K281R; K310R; P339R; And the T350R, The above-mentioned engineered ISDra2-derived TnpB protein has its nucleic acid cleavage activity removed.

10. In the engineered ISDra2 TnpB protein of claim 9, The above-mentioned engineered ISDra2 TnpB protein comprises an amino acid sequence selected from the group consisting of SEQ ID NOs 9 to 10.

11. Base editor proteins including the following: An engineered ISDra2 TnpB protein selected from any one of claims 9 to 10; One or more nuclear localization signals (NLS); and One or more base editing domains; Here, the base editing domain is adenosine deaminase or cytidine deaminase.

12. In the base editor protein of paragraph 11, The above base editor protein includes the following: An engineered ISDra2-derived TnpB protein selected from any one of claims 9 to 10; One or more nuclear localization signals (NLS); and One or more adenosine deaminases.

13. In the base editor protein of paragraph 12, The above base editor protein includes the following structure: [NLS1] - [AD] - [Dead TnpB] - [NLS2] Here, the NLS1 is a first nuclear localization signal, and Here, the NLS2 is a second nuclear localization signal, and The above AD is an adenosine deaminase, and The above Dead TnpB is an engineered ISDra2-derived TnpB protein selected from any one of claims 9 to 10.

14. In the base editor protein of paragraph 11, The above base editor protein includes the following: An engineered ISDra2-derived TnpB protein selected from any one of claims 9 to 10; One or more nuclear localization signals (NLS); One or more cytidine deaminases; and One or more uracil glycosylase inhibitors (UGI).

15. In the base editor protein of paragraph 14, The above base editor protein is a protein with the following structure: [NLS1] - [CD] - [Dead TnpB] - [UGI] - [NLS2] Here, the NLS1 is a first nuclear localization signal, and Here, the NLS2 is a second nuclear localization signal, and The above CD is cytidine deaminase, and The above Dead TnpB is an engineered ISDra2-derived TnpB protein selected from any one of claims 9 to 10.

16. A base editor composition comprising the following: The base editor protein of claim 11, or a nucleic acid encoding the base editor protein; and reRNA, or a nucleic acid encoding the reRNA; Here, the reRNA comprises a scaffold and a spacer, and The scaffold can bind to the base editor protein to form a base editor complex, and The above spacer can bind complementarily to a predetermined target nucleic acid.

17. A base editor composition comprising the following: A base editor protein selected from any one of claims 12 to 13, or a nucleic acid encoding said base editor protein; and reRNA, or a nucleic acid encoding the reRNA; Here, the reRNA comprises a scaffold and a spacer, and The scaffold can bind to the base editor protein to form a base editor complex, and The above spacer can bind complementarily to a predetermined target nucleic acid.

18. In the base editor composition of claim 17, The above reRNA includes the following: 5'-[Scaffold]-[Spacer]-[Linker]-[Stabilization Domain]-3' Here, the scaffold comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 5 to 7, and The above linker includes or is absent the nucleic acid sequence of SEQ ID NO. 23, and The above stabilization domain includes a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22, or is absent.

19. A base editor composition comprising the following: A base editor protein selected from any one of claims 14 to 15, or a nucleic acid encoding said base editor protein; and reRNA, or a nucleic acid encoding the reRNA; Here, the reRNA comprises a scaffold and a spacer, and The scaffold can bind to the base editor protein to form a base editor complex, and The above spacer can bind complementarily to a predetermined target nucleic acid.

20. In the base editor composition of claim 19, The above reRNA includes the following: 5'-[Scaffold]-[Spacer]-[Linker]-[Stabilization Domain]-3' Here, the scaffold comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs 5 to 7, and The above linker includes or is absent the nucleic acid sequence of SEQ ID NO. 23, and The above stabilization domain includes a nucleic acid sequence selected from the group consisting of SEQ ID NOs 21 to 22, or is absent.

21. As a method of editing target DNA within a cell, Here: The above target DNA is double-stranded DNA comprising a first strand and a second strand inversely complementary thereto; The above first strand includes a first correction window; The second strand includes a second correction window at a position corresponding to the first correction window; The first correction window and the second correction window are inversely complementary to each other; and The first calibration window and the second calibration window are collectively referred to as calibration windows. The above method includes the following: A process of delivering a base editing composition selected from any one of claims 19 to 20 to the cell, Here: In the above cell, the base editor protein of the composition and reRNA bind to form a base editor complex; and One or more C:G base pairs contained in the correction window are edited into T:A base pairs by the base editor complex.

22. In the method of editing intracellular target DNA of paragraph 21, Here: The length of the above correction window is l-bp; The base editor complex recognizes a sequence consisting of the m-th nucleotide to the (m + 4)-th nucleotide in the upstream direction relative to the 5'-terminus of the first correction window as a transposon-associated motif (TAM); The above reRNA includes a spacer of length n-nt; The above m is an integer between 2 and 10; The above l is an integer between 1 and 9; and The above n is an integer between 18 and 22.

23. In the method of editing intracellular target DNA of paragraph 22, The above m is 2, the above l is 6, and the above n is 20.

24. As a method of editing target DNA within a cell, Here: The above target DNA is double-stranded DNA comprising a first strand and a second strand inversely complementary thereto; The above first strand includes a first correction window; The second strand includes a second correction window at a position corresponding to the first correction window; The first correction window and the second correction window are inversely complementary to each other; and The first calibration window and the second calibration window are collectively referred to as calibration windows. The above method includes the following: The process of delivering a base editing composition selected from any one of claims 17 to 18 to the cell, Here: In the above cell, the base editor protein of the composition and reRNA bind to form a base editor complex; and One or more A:T base pairs included in the correction window are edited into G:C base pairs by the base editor complex.

25. In the method of editing intracellular target DNA of paragraph 24, Here: The length of the above correction window is l-bp; The base editor complex recognizes a sequence consisting of the m-th nucleotide to the (m + 4)-th nucleotide in the upstream direction relative to the 5'-terminus of the first correction window as a transposon-associated motif (TAM); The above reRNA includes a spacer of length n-nt; The above m is an integer between 2 and 11; The above l is an integer between 1 and 10; and The above n is an integer between 18 and 22.

26. In the method of editing intracellular target DNA of paragraph 25, The above m is 2, the above l is 6, and the above n is 20.