Generative ai–based target protein complex formation inhibitor for enhancing prime editing efficiency and uses thereof
A novel MLH1 small binder (MLH1-SB) is designed to inhibit the MutLα complex, addressing inefficiencies in prime editing by enhancing editing efficiency and reducing off-target mutations, suitable for therapeutic applications.
Patent Information
- Application Number
- PCT/KR2025/013019
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-27
- Filing Date
- 2025-08-26
- Publication Date
- 2026-03-05
AI Technical Summary
Existing prime editing technologies face inefficiencies due to interference from the mismatch repair (MMR) pathway, particularly the MutLα complex, limiting editing efficiency and applicability in therapeutic applications, especially for larger genetic modifications.
Development of a novel MLH1 small binder (MLH1-SB) using RFdiffusion and AlphaFold 3-based design to inhibit the MutLα complex, enhancing prime editing efficiency by binding to specific amino acid residues of MLH1, thereby improving the editing process.
The MLH1-SB significantly enhances prime editing efficiency, reducing off-target mutations and cytotoxicity, and supports larger genetic modifications, making it suitable for therapeutic applications.
Smart Images

Figure KR2025013019_05032026_PF_FP_ABST
Abstract
Description
Generative AI-based target protein complex formation inhibitor for improving prime editing efficiency and its use
[0001] The present invention relates to a generative AI-based target protein complex inhibitor for enhancing prime editing and its use, and more particularly, to a novel polypeptide that binds to MLH1 and inhibits its interaction with PMS2 and formation of the MutLα complex, thereby enhancing prime editing efficiency, and its use for prime editing.
[0002]
[0003] Since the advent of CRISPR-Cas9 genome editing technology (Science 337, 816-821; Nat Biotechnol 31, 230-232; Science 339, 819-823 and Science 339, 823-826), various studies have been conducted to expand its application range, resulting in the development of base editors (BEs) and prime editors (ature 533, 420-424; Nature 551, 464-471 and Nature 576, 149-157). While base editors can modify one or a few bases, prime editors have the advantage of being able to introduce not only base substitutions but also small insertions and deletions (indels) (Nat Rev Drug Discov 19, 839-859; Nat Rev Genet 24, 161-177). Due to this flexibility and precision, prime editors have attracted significant attention in therapeutic applications such as cell and gene therapy and disease modeling (Trends Biotechnol 41, 1000-1012; Stem Cell Res Ther 14, 164). Accordingly, various strategies are being explored to improve the efficiency of prime editing (PE).
[0004] PE2 is an optimized prime editing construct that fuses Cas9 nickase (nCas9) and reverse transcriptase (RT) to synthesize a DNA strand containing the desired mutation at the target site (Nature 576, 149-157). PE2 is based on utilizing the 3'-extension of the prime editing guide RNA (pegRNA) as a template for reverse transcription. To enhance prime editing efficiency, PE3 and PE5 employ a strategy of introducing an additional nicking guide RNA (ngRNA) to induce a second nick in the non-edited strand, thereby shifting the flap equilibrium toward the desired editing direction (Nature 576, 149-157; Cell 184, 5635-5652 e5629). However, subsequent studies have shown that the mismatch repair (MMR) pathway interferes with the desired editing at the target site (Cell 184, 5635-5652 e5629). In the MMR pathway, MutS homologs (MutSα, composed of MSH2-MSH6, and MutSβ, composed of MSH2-MSH3) recognize base mismatches and recruit the MutLα complex, composed of MLH1 and PMS2, to induce repair (Nat Rev Mol Cell Biol 7, 335-346). Accordingly, inhibition of key MMR factors such as MSH2, MSH3, MSH6, MLH1, and PMS2 significantly improves PE efficiency (Cell 184, 5635-5652 e5629; Mol Ther Nucleic Acids 32, 914-922; Nat Commun 15, 4002). In particular, the PE4 structure is characterized by co-delivery of MLH1dn (dominant negative MLH1) with PE to inhibit the MutLα complex (Cell 184, 5635-5652 e5629;).However, recent studies have reported that MMR inhibition only significantly improves the efficiency of editing of about 10 base pairs (bp) or less (Mol Ther Nucleic Acids 32, 914-922, Nat Commun 13, 760), and its large size of about 2300 bases limits its use in the development of genetic therapeutics.
[0005] In addition, various strategies have been proposed to improve PE efficiency. For example, inserting an RNA pseudoknot motif at the 3' end of pegRNA to increase stability and prevent 3'-extension degradation was reported (Nat Biotechnol 40, 402-410). Furthermore, various PE6 structures were developed through phage-assisted continuous evolution, improving editing efficiency (Cell 186, 3983-4002 e3926). Another strategy, PE7, utilizes the La protein, an RNA-binding protection factor, to protect the 3' end of pegRNA (Nature 628, 639-647). Despite continuous advancements in PE technology, editing efficiency remains suboptimal, particularly for therapeutic applications. Additionally, improving PE structure without ngRNA, which may induce unintended indel mutations, remains an important challenge.
[0006] Recent advances in artificial intelligence (AI) technology enable not only sequence-based protein structure prediction (Nature 596, 583-589; Nature 630, 493-500; Science 373, 871-876) but also de novo protein and peptide generation (Nature 620, 1089-1100). Furthermore, these AI technologies have been actively applied in the development of CRISPR-related tools (bioRxiv 2024.04.22.590591; BioRxiv 2024.02.27.582234; BioRxiv. 2024.08.03.606485). However, improvements in CRISPR genome editing technology using AI have primarily focused on enhancing the catalytic activity of core enzymes such as Cas9 (bioRxiv 2024.04.22.590591) and TadA (BioRxiv. 2024.08.03.606485). Meanwhile, the application of RFdiffusion, an AI technology for producing binder proteins, is still limited to the development of single-site binders and the design of antibody surrogates (Nature 620, 1089-1100; BioRxiv. 2024.03.14.585103).
[0007] Against this backdrop, the present inventors have made extensive efforts to develop novel inhibitors of the MMR pathway to improve prime editing efficiency. As a result, (i) using RFdiffusion, we designed a de novo protein with multiple binding sites, including a non-competitive anchoring site and an inhibitory site that inhibits the core activity, in addition to a simple competitive binding site, and (ii) using an AlphaFold 3-based filtering system, we developed a novel MLH1 small complex (MLH1-SB) that significantly improves PE efficiency. We confirmed that the MLH1 small complex (MLH1-SB) can be easily integrated into existing PE systems and can significantly increase the PE efficiency of existing prime editing systems, thereby completing the present invention.
[0008]
[0009] The above information described in this background section is solely intended to enhance understanding of the background of the present invention and may not include information that constitutes prior art already known to a person of ordinary skill in the art to which the present invention pertains.
[0010]
[0011] Summary of the invention
[0012] The purpose of the present invention is to provide a novel polypeptide that binds to MLH1 and inhibits the formation of a MutLα complex with PMS2.
[0013] Another object of the present invention is to provide a nucleic acid encoding the polypeptide.
[0014] Another object of the present invention is to provide a vector comprising the nucleic acid.
[0015] Another object of the present invention is to provide a composition or kit for prime editing comprising the polypeptide and / or nucleic acid.
[0016] Another object of the present invention is to provide a prime editing system comprising the polypeptide and / or nucleic acid.
[0017] Another object of the present invention is to provide a use for gene editing of the polypeptide and / or nucleic acid.
[0018] Another object of the present invention is to provide a gene editing method using the polypeptide and / or nucleic acid.
[0019]
[0020] To achieve the above purpose, the present invention provides a novel MLH1 binding polypeptide that binds to one or more amino acid residues of MLH1 and inhibits the formation of the MLH1-PMS2 complex (MutLα complex).
[0021] The present invention also provides a nucleic acid encoding the polypeptide.
[0022] The present invention also provides a prime editing vector comprising the nucleic acid.
[0023] The present invention also provides a composition or kit for prime editing comprising the polypeptide and / or nucleic acid.
[0024] The present invention also provides a prime editing system comprising the polypeptide and / or nucleic acid.
[0025] The present invention also provides uses for prime editing of said polypeptides and / or nucleic acids.
[0026] The present invention also provides a prime editing method using the polypeptide and / or nucleic acid.
[0027]
[0028] Figure 1 illustrates the mechanism and effect of prime editing enhancement of the polypeptide (MLH1-SB) of the present invention. Figure 1a is a schematic diagram of prime editing. DNA and RNA sequences containing the intended mutations are highlighted in red. Generated using BioRender.com. Figure 1b is a schematic diagram of MLH1 SB generation. MLH1-C: C-terminal domain of MLH1; PMS2-C: C-terminal domain of PMS2; MLH1-SB: MLH1 small binder. Figure 1c is the dimer interface structure between MLH1 and PMS2 predicted by AlphaFold 3. Binding hotspot residues designated as “Main binding” (W538, Q542), “Critical activity” (Y750), and “Additional binding” (M682) are circled in red. The MLH1-C domain is shown in green, and the PMS2-C domain is shown in brown. Figure 1d shows dot plots of the interface distance and the interaction pAE calculated via RFdiffusion according to the AlphaFold 3 competition assay: (left) MC binder, (right) MCA binder. The histogram of the interface distance is shown at the top. The red line (9.0157 Å) represents the original interface distance of the MLH1-C and PMS2-C dimeric complex (MutLα complex), and the yellow line (9.7724 Å) represents the threshold. MC, MCA, and failed-MCA binders are highlighted in orange, green, and blue boxes, respectively. Figure 1e is a schematic diagram of the AlphaFold 3 competition assay of the present invention. Figure 1f shows the fold change in prime editing efficiency relative to PE2 for each MLH1-SB in HEK3 (3 bp insertion), RNF2 (3 bp insertion), RNF2 (3 bp deletion), and HEK4 (G→T substitution). The red dotted line indicates a 3-fold improvement, and MLH1 SBs with a 3-fold or greater increase in PE efficiency are indicated by red dots.Bars: mean; error bars: standard deviation (SD); n = 12.
[0029] Figure 2 illustrates the generation and filtering process of MLH1-SB. Figure 2a shows the parameter distribution (left) and filtering result (right) for MC binder proteins. The total number of proteins is 1978, and they were filtered based on the reference threshold of 9.7724 Å (yellow line). The filtered MC binder proteins are indicated by the top 20 numbers. The red line indicates the original interface distance (9.0157 Å) between MLH1-C and PMS2-C, which is the same as Figure 1d. Figure 2b shows the parameter distribution (left) and filtering result (right) for MCA binder proteins. The selected MCA binders are indicated by the top 23 numbers. Figure 2c shows five atom pairs (purple) selected to calculate the interface distance (average atom-pair distance). The average distance between the Cα atoms of the residues is calculated based on the average distance between the Cα atoms of the residues, and includes L540, L743, Q542, C756, and N739 in MLH1-C, and I688, I853, N683, E705, and Q861 in PMS2-C. Three of the selected positions were chosen to span the MLH1-PMS2 major binding site perpendicularly, and the other two were chosen symmetrically at both ends of the C-terminal binding sites of MLH1 and PMS2, comprehensively covering the core region of the binding interface. In particular, the MLH1 C-terminus is an important region for functional activity. Figure 2d shows failed-MCA binder proteins filtered according to the reference threshold. The filtered failed-MCA binders are indicated by the top 17 numbers, and an expanded version of the interface distance distribution for MCA binder proteins is also presented.Figures 2e-2g show the trimeric structure of the MLH1 C-terminal domain (MLH1-C), the PMS2 C-terminal domain (PMS2-C), and the small binder predicted by AlphaFold 3. Figure 2e shows the structure of the MC binder protein, Figure 2f shows the structure of the MCA binder protein, and Figure 2g shows the structure of the failed-MCA binder protein, respectively. MLH1-C: green, PMS2-C: orange, MLH1-small binder: blue.
[0030] Figure 3 shows the characteristics of the selected MLH1-SB. Figures 3a to 3c show the dimeric structure between the MLH1 C-terminal domain (MLH1-C) and the MLH1 small binder predicted by AlphaFold 2 integrated into the RFdiffusion program. Figure 3a shows the structures of the MC binder protein, Figure 3b shows the structures of the MCA binder protein, and Figure 3c shows the structures of the failed-MCA binder protein, respectively. MLH1-C: green, MLH1-small binder: blue. Figure 3d shows the results of a comparison of various parameters between the MCA binder protein and the failed-MCA binder protein. The length unit is amino acid (aa). ns: not significant, p > 0.05; ***, p < 0.001. Figure 3e shows the calculated average TM-scores among the MC, MCA, and failed-MCA binder proteins. TM-score: Mean ± standard deviation (SD) Figure 3f shows the TM-score distribution (left) and sequence identity distribution (right) calculated for the total MC binder proteins and the total MCA binder proteins generated. TM-scores were calculated using Foldseek (Nat Biotechnol 42, 243–246), and sequence identity was calculated using MMseqs2 (Nat Biotechnol 35, 1026–1028).
[0031] Figure 4 shows the effect of MLH1-SB on prime editing efficiency. Figures 4a-d show the prime editing efficiency under the following conditions using the indicated MLH1 small binders: CTT insertion in HEK3 (a), TAC insertion in RNF2 (b), CTG deletion in RNF2 (c), and G to T substitution in HEK4 (d). Bars: mean; error bars: standard deviation (SD); n = 3. Figure 4e shows the overall prime editing fold-change when using the MCA and failed-MCA binders, respectively. Bars: mean; error bars: standard deviation (SD); n = 23 and n = 17, respectively. ****, p < 0.0001. Figure 4f shows the distribution of the mean fold-change in prime editing efficiency using the MC binder (left), the MCA binder (middle), and the failed-MCA binder (right). The red dotted line indicating a 2.5-fold increase is highlighted. Figure 4g shows the sequence identity (left) and E-value distribution (middle) of random amino-acid proteins, random existing proteins, all generated MC binder proteins, and all generated MCA binder proteins. Selected MC, MCA, and failed-MCA binder proteins are additionally indicated as individual dots. The solid line represents the median, and the shaded area represents the interquartile range (IQR, 25th-75th percentile). The right panel is an enlarged graph of the E-value distribution with the y-axis range set to 0.001-10000. Sequence identity and E-value were calculated using MMseqs2.
[0032] Figure 5 shows the compatibility with the PE6 and PE7 architectures of MLH1-SB. Figure 5a is a schematic diagram of the prime editor architecture. nCas9: SpCas9 H840A nickase; RT: human codon-optimized MMLV-RT; MLH1dn: dominant-negative MLH1; MLH1-SB: MCA-23 MLH1 small binder; La: small RNA-binding exonuclease protection factor. Figures 5b and 5c show the prime editing efficiency for the designated targets and mutations when PEmax and PE7 were introduced together with the control vector (Cont), MLH1dn, or MLH1-SB expression vector in HeLa cells. Figure 5b shows the results performed without nicking gRNA, and Figure 5c shows the results when nicking gRNA was introduced together. Bars: mean; error bars: standard deviation (SD); n = 3. Figure 5d shows the normalized prime editing efficiency when PEmax and PE7 were co-transduced with the control vector (Cont), MLH1dn, or MLH1-SB expression vector in HeLa cells. Numbers above the bars represent fold-change values. Bars: mean; error bars: standard deviation (SD); n = 21. ns: non-significant, p > 0.05; *p < 0.05; **p < 0.01; ***p < 0.001. Figure 5e is a schematic diagram of the PE architecture. nCas9: SpCas9 H840A nickase, RT6d: PE6d reverse transcriptase, MLH1-SB: MCA-23 small binder, La: RNA-binding exonuclease protection factor.Figure 5f shows the prime editing efficiency for the indicated targets and mutations when the control vector or MLH1-SB vector was introduced using PE6d and PE6d-La. The bars represent the mean, and the error bars represent the standard deviation (SD) (n = 3). Figure 5g shows the normalized prime editing efficiency and unwanted mutations when PE6d and PE6d-La were co-transduced with the control or MLH1-SB vector in HeLa cells. Samples in which unwanted mutations did not occur due to PE6d were excluded from normalization, and the numbers above the bars represent the fold-change values. Bars: mean; error bars: standard deviation (SD); (prime editing: n = 21, unwanted mutations: n = 18). ns: non-significant, p > 0.05; **p < 0.01; ***p < 0.001; ****p < 0.0001.
[0033] Figure 6 shows the results of confirming unwanted mutations in HeLa cells. Figures 6a and 6b show unwanted mutations and induced mutations for the designated targets when PEmax or PE7 was introduced together with the control vector (Cont) and the MLH1dn or MLH1-SB expression vector in the presence (a) and absence (b) of ngRNA in HeLa cells. Bars: mean; error bars: standard deviation (SD); n=15. ns: not significant, p > 0.05; ***, p < 0.001. Figure 4c shows the normalized unwanted mutations by PEmax and PE7 when HeLa cells were treated with the control vector, MLH1dn, or MLH1-SB expression vector. Cases in which no unwanted mutations were detected by PEmax were excluded from normalization. The numbers above the bars represent fold-change values. Bars: mean; error bars: standard deviation (SD); n=15. ns: not significant, p > 0.05. Figure 6d shows the undesired mutations induced at the indicated targets when PE6d or PE6d-La was introduced together with the control vector, MLH1dn, or MLH1-SB expression vectors. Bars: mean; error bars: standard deviation (SD); n = 3.
[0034] Figure 7 shows that MLH1-SB enhances MLH1 binding and MMR-dependent PE efficiency. Figure 7a shows the Western blot results of co-immunoprecipitation performed after transfection of HeLa cells with a vector containing an HA tag to MLH1-SB (MLH1-SB-HA), alone or together with 3×FLAG-tagged MLH1 (FLAG-MLH1). Figure 7b shows immunofluorescence images of HEK293T cells transfected with the indicated vectors. The cells were stained for nuclei with Hoechst (blue), FLAG-tagged MLH1 was stained with anti-FLAG antibody (magenta), and HA-tagged MLH1-SB was stained with anti-HA antibody (green). Left panels show: Mock (cells transduced with empty vector), MLH1-SB (HA staining appears in the cytoplasm), MLH1 (MLH1-expressing cells), and FLAG-MLH1 + MLH1-SB-HA (co-transduced cells). Right panels show high-magnification images of co-transduced cells. (1, 2) Magnifications of co-localization in the nucleus and cytoplasm are shown. Arrows indicate cells with high MLH1 expression. Scale bar: 20 μm. Figures 7c and 7d show the PE efficiency for the indicated targets when control, MLH1dn, or MLH1-SB was transduced together with PEmax and PE7 in HEK293 and HEK293T cells. Bars: mean; error bars: standard deviation (SD); n = 3. Figures 7e and f show the normalized PE efficiency and unwanted mutations when control (Cont), MLH1dn, or MLH1-SB was introduced together with PEmax and PE7 in HEK293 (e) and HEK293T (f) cells. The numbers above the PE efficiency bars represent fold-change values. Bars: mean; error bars: standard deviation (SD); n = 21. ****p < 0.0001.
[0035] Figure 8 illustrates the enhanced MLH1-dependent prime editing efficiency of MLH1-SB. Figure 8a is a cropped image of the Western blot presented in Figure 7a. The cropped area is indicated by a red box. Figure 8b schematically illustrates an experiment to identify MLH1-SB interacting proteins. HeLa cells were transfected with a puromycin resistance gene expression vector (PuroR) or an MLH1-SB expression vector linked to an HA tag (PuroR-2A-SB-HA). Cells were lysed and immunoprecipitated using anti-HA magnetic beads. The immunoprecipitated proteins were separated by SDS-PAGE, enzymatically digested, and analyzed by LC-MS / MS. Figure 8c is an LC-MS / MS spectrum showing the MLH1-derived peptide sequence. Figures 8d and 8e show the mass point distributions of proteins identified in HeLa cells transfected with PuroR-2A-SB-HA (d) and PuroR (e), respectively. Figure 8f shows the relative expression levels of MLH1 mRNA in HeLa, HEK293, and HEK293T cells. Bars: mean; error bars: standard deviation (SD); n = 3. The left panel of Figure 8g is a representative image taken by a confocal microscope, in which nuclei were labeled from A to E, and the red arrows indicate the distance axis of the fluorescence intensity profile graph. Scale bar = 20 μm. The middle panel shows the fluorescence intensity profiles for three fluorescent markers (green: AF488#-T5, magenta: mCher2#-T3, blue: H3342#-T6), and the dotted red lines indicate the positions of the labeled nuclei. The right panel is a scatter plot showing the normalized area under the curve (AUC) of green and magenta fluorescence in each designated nucleus. The AUC for each nucleus was normalized by dividing it by its length.Figures 8h and 8i show unwanted and induced mutations for the indicated targets when PEmax or PE7 was introduced together with the control vector, MLH1dn, or MLH1-SB expression vectors into HEK293 (h) and HEK293T (i) cells. Bars: mean; error bars: standard deviation (SD); n = 3.
[0036] Figure 9 shows the PE efficiency enhancement effect according to the gene editing size of MLH1-SB. Figure 9a schematically shows the tested prime editor architecture. nCas9: SpCas9 H840A nickase, RT: human codon-optimized MMLV-RT, MLH1dn: dominant-negative MLH1, MLH1-SB: MCA-23 MLH1 small binder, MLH1-SB-NLS: MLH1-SB with nuclear localization sequence (NLS), La represents RNA-binding exonuclease protection factor, 2A represents 2A peptide, and Xt represents SGGS-XTEN16-SGGS linker. Figure 9b shows the normalized PE efficiency and unwanted mutations when PE7, PE7-MLH1dn, PE7-MLH1-SB (SB), PE7-MLH1-SB-NLS (SB-NLS), PE7-MLH1-SB-MLH1-SB-NLS (SB2), and PE7 fused to the N-terminus or C-terminus of MLH1-SB (N-SB, C-SB) were used in HeLa cells. Bars: mean; error bars: standard deviation (SD); n = 9. Figures 9c to 9e show the PE efficiency for substitutions (C), insertions (D), and deletions (E) when PEmax, PE7, and PE7-SB2 were used in the presence or absence of ngRNA in HeLa cells. Bars: mean; error bars: standard deviation (SD); n = 3. Figures 9f to 9g show the normalized PE efficiency and unwanted mutations for edits less than 12 bp (F) and greater than 23 bp (G) when PEmax, PE7, and PE7-SB2 were used in the presence or absence of ngRNA in HeLa cells. Cases of PEmax in which no unwanted mutations were observed were excluded from normalization. Numbers above the bars represent fold-change values.Bars: mean; error bars: standard deviation (SD); n = 54 (edits <12 bp); n = 51 (unwanted <12 bp); n = 9 (edits >23 bp); n = 6 (unwanted >23 bp). ns: not significant, p > 0.05; ***p < 0.001; ****p < 0.0001. Figures 9h to 9i show the normalized PE efficiency in hiPSCs (H) and HEK293 (I) cells when PEmax, PE7, and PE7-SB2 were used in the presence or absence of ngRNA. The numbers above the bars represent fold-change values. Bars: mean; error bars: standard deviation (SD); n = 18 hiPSCs, n = 21 HEK293). ns: not significant, p > 0.05; *p < 0.05; **p < 0.01; ****p < 0.0001.
[0037] Figure 10 shows the PE efficiency enhancement effect of the MLH1-SB fusion prime editor. Figure 10a shows the prime editing efficiency and unwanted mutations performed on designated targets and mutations in HEK3, RNF2, and HEK4 cells using PE7, PE7 linked to MLH1dn (MLH1dn), PE7 linked to MLH1-SB (SB), PE7 linked to MLH1-SB with an added nuclear export sequence (NLS) (SB-NLS), PE7 linked to both MLH1-SB and MLH1-SB-NLS (SB2), and forms in which MLH1-SB is linked to the N-terminus (N-SB) or C-terminus (C-SB) of PE7, respectively. Bars: mean; error bars: standard deviation (SD); n = 3. Figures 10b-10d show the unwanted mutations when PEmax, PE7, and PE7-SB2 were used in the presence or absence of nicking gRNA (ngRNA) for the indicated substitution (b), insertion (c), and deletion (d) mutations in HeLa cells. Bars: mean; error bars: standard deviation (SD); n = 3. Figures 10e and 10f show the prime editing efficiency (e) and unwanted mutation frequency (f) when PEmax, PE7, and PE7-SB2 were used in the presence or absence of ngRNA for the indicated mutations in human induced pluripotent stem cells (hiPSCs). Bars: mean; error bars: standard deviation (SD); n = 3. Figure 10g shows the normalized unwanted mutations when PEmax, PE7, and PE7-SB2 were used in the presence or absence of ngRNA in human hiPSCs. If no unwanted mutations were detected in PEmax, they were excluded from normalization. The numbers above the bars represent the average fold-change. Bars: mean; error bars: standard deviation (SD); n = 15.Figures 10h and 10i show the prime editing efficiency (h) and unwanted mutations (i) for the indicated mutations in HEK293 cells when PEmax, PE7, and PE7-SB2 were used in the presence or absence of ngRNA, respectively. Bars: mean; error bars: standard deviation (SD); n = 3. Figure 10j shows the normalized unwanted mutations when PEmax, PE7, and PE7-SB2 were used in the presence or absence of ngRNA, respectively, in HEK293 cells. Bars: mean; error bars: standard deviation (SD); n = 21.
[0038] Figure 11 shows the results of verifying compatibility with MLH1-SB and improved reverse transcriptase. Figure 11a schematically illustrates the prime editor architecture. nCas9: SpCas9 H840A nickase; RT: human codon-optimized MMLV-RT; La: small RNA-binding exonuclease protection factor; 2A: 2A peptide; RT6b: RT of PE6b; RT6c: RT of PE6c; RT6d: RT of PE6d; RT**: RT of PEmax**; 2A: 2A peptide. Figures 11b and 11c show the prime editing efficiency (b) and unwanted mutation frequency (c). Bars: mean; error bars: standard deviation (SD); n = 3. Figure 11d shows the normalized prime editing efficiency and unwanted mutations by the PE7-SB2 system, which combines PEmax, PE7-SB2, and RT6b (RT6b-SB2), RT6c (RT6c-SB2), RT6d (RT6d-SB2), and RT** (RT**-SB2), respectively. The numbers above the bars represent the average fold-change values. Bars: mean; error bars: standard deviation (SD); n = 21.
[0039] Figure 12 compares the off-target editing of PE7-SB2 and PE7. Figure 12a Off-target sequence information of HEK3 (left), HEK4 (center), and RNF2 (right). Mismatched bases within the off-target spacer sequence are highlighted in red. Figures 12b-12d show the off-target mutation frequencies for HEK3 CTT insertions (b), HEK4 G to A substitutions (c), and RNF2 TAC insertions (d) performed using PEmax, PE7, and PE7-SB2 in the presence and absence of ngRNA (nicking gRNA) in HeLa cells, respectively. Bars: mean; error bars: standard deviation (SD); n = 3. Figure 12e shows the normalized off-target editing when PEmax, PE7, and PE7-SB2 were used in the presence or absence of ngRNA in HeLa cells. Experiments in which no off-target mutations were detected in PEmax were excluded from normalization. Bars: mean; error bars: standard deviation (SD); n = 27.
[0040] Figure 13 shows the results of confirming the cytotoxicity of MLH1-SB and its compatibility with primary cells and cell cycle arrested cells. Figures 13a and 13b show the results of flow cytometry based on Ki67 staining (a) and AnnexinV / propidium iodide staining (b) after introducing puromycin resistance gene (PuroR), dominant-negative MLH1 (MLH1dn), MLH1 small binder (MLH1-SB), MLH1dn and PE7 linked with 2A peptide (PE7-MLH1dn), and PE7-SB2 into HeLa cells, respectively. Figures 13c and 13d show the results of flow cytometry based on Ki67 staining (c) and AnnexinV / propidium iodide staining (d) in fibroblasts introduced with PE7 or PE7-SB2. In Ki67 staining analysis, the right box represents proliferating cells, and the numbers in the boxes represent the proportion of replicating cells. In AnnexinV / propidium iodide staining analysis, the lower left quadrant (Annexin V- / PI-) represents viable cells, the upper left (Annexin V- / PI+) represents necrotic cells, the lower right (Annexin V+ / PI-) represents early apoptotic cells, and the upper right (Annexin V+ / PI+) represents late apoptotic or dead cells, and the numbers in each quadrant represent the proportion of cells in that state. Figures 13e and 13f are bar graphs showing flow cytometric analysis of Ki67 staining (e) and AnnexinV / propidium iodide staining (f) for the above experiment. Bars: mean; error bars: standard deviation (SD); n = 3. Figure 13g shows unwanted mutations and prime editing efficiency at the indicated positions in fibroblasts introduced with PE7 and PE7-SB2. Bars: mean; error bars: standard deviation (SD); n = 3.Figure 13h shows the results of flow cytometry analysis performed using propidium iodide staining after HeLa cells were treated with DMSO or 50 ng / ml Nocodazole for 18 h. The distribution of cells in G0 / G1, S, and G2 / M phases was quantified using histogram-based analysis. Figure 13i shows the prime editing efficiency and unwanted mutations at the indicated positions when PE7 and PE7-SB2 were introduced in HeLa cells treated with Nocodazole. Bars: mean; error bars: standard deviation (SD); n = 3. Figures 13j and 13k show the cumulative (left) and normalized (right) frequency distributions of unique single-nucleotide variants (SNVs, top) and insertions / deletions (indels, bottom) detected by whole-genome sequencing (WGS) of hiPSC cells (j) and HeLa cells (k). Mutations were detected using HaplotypeCaller according to the GATK4 best practices pipeline, and sequencing reads were aligned to the human reference genome (hg38) using BWA-MEM. BAM file alignment was performed using SAMtools, and PCR duplicate removal was performed using Picard.
[0041] Figure 14 shows the cytotoxicity and transcriptome alteration of MLH1-SB. Figure 14a shows the results of principal component analysis (PCA) on gene expression. Each dot represents one condition, and the degree of separation along the principal component axis reflects the difference between gene expression profiles. Figure 14b shows a volcano plot showing differential gene expression for the following pairwise comparison conditions: MLH1dn vs PuroR, MLH1-SB vs PuroR, MLH1-SB vs MLH1dn, PE7 vs PuroR, PE7-SB2 vs PuroR, PE7-SB2 vs PE7. The x-axis represents the log2 fold change in gene expression, and the y-axis represents -log10 (raw p-value). Significantly upregulated genes (fold change ≥ 2, raw p < 0.05) are shown in orange, and significantly downregulated genes (fold change ≤ -2, raw p < 0.05) are shown in blue. Genes commonly altered in both MLH1-SB and MLH1dn are shown as follows: Cg1 (LOC100130744), Cg2 (UBE2F-SCLY), Cg3 (LOC124900953), Cg4 (LOC124904773), Cg5 (LOC124904627), Cg6 (FMC1-LUC7L2). Genes specifically altered in MLH1-SB are indicated as SBg1 (LOC105374748), SBg2 (SENP3-EIF4A1), and genes specifically altered in MLH1dn are indicated as Mdng1 (RIPPLY2-CYB5R4), Mdng2 (LHX4-AS1). Genes specifically altered in PE7 are indicated as PEg1 (HSPA6), PEg2 (HSPA7), PEg3 (SSB), PEg4 (POLR2J3), PEg5 (HBA2).Figure 14C is a bar graph showing the results of flow cytometry analysis for Ki67 staining. Figure 14D shows the results of flow cytometry analysis for Annexin V / propidium iodide staining. Bars: mean; error bars: standard deviation (SD); n = 3. ns: not significant, p > 0.05; *p < 0.05; **p < 0.01.
[0042] Figure 15 shows the effect of enhancing the in vivo prime editing efficiency in a mouse model of MLH1-SB. Figure 15a is a schematic diagram evaluating plasmid delivery and in vivo expression through bioluminescence imaging on days 5, 7, and 10 after hydrodynamic injection of luciferase vector (pCAG-Luc) into 6-week-old ICR mice. Figure 15b is a bioluminescence image showing photon flux measured in the liver of mice (Luc-1, Luc-2) injected with luciferase plasmid on days 5, 7, and 10. Figure 15c shows the results of quantitative analysis of photon flux in mice (Luc-1, Luc-2) injected with PBS or luciferase plasmid. Figure 15d shows an experimental outline in which liver DNA (gDNA) and serum were collected on day 10 after hydrodynamic injection of PE7 or PE7-SB2 together with pegRNA targeting Igf2 into 6-week-old ICR mice. Figure 15e shows serum AST and ALT levels in the non-treated group, PBS-injected group, PE7-SB2-only injection group, PE7 and pegRNA-injected group, and PE7-SB2 and pegRNA-injected group. 60 μg of plasmid was injected in each condition. Bars: mean; error bars: standard deviation (SD); n = 4 (non-treated, PBS, PE7-SB2, PE7-SB2 + pegRNA); n = 5 (PE7 + pegRNA). ns: not significant, p > 0.05. Figure 15f is a bar graph showing the PE efficiency and unwanted mutations at the Igf2 region measured in liver tissue from the untreated group, PBS-injected group, PE7 and pegRNA-injected group, and PE7-SB2 and pegRNA-injected group. A total of 50 or 60 μg of plasmid was used.Bars: mean; error bars: standard deviation (SD); n = 1 (non-treated, PBS); n = 6 (50 μg PE7 / PE7-SB2); n = 5 (60 μg PE7); n = 4 (60 μg PE7-SB2). Figure 15g shows the results comparing the mutation frequencies for seven off-target sites (OT1-OT7) of Igf2 among the untreated group, PBS-injected group, 60 μg PE7 and pegRNA-injected group (PE7), and 60 μg PE7-SB2 and pegRNA-injected group (PE7-SB2). Bars: mean; error bars: standard deviation (SD); n = 1 (non-treated, PBS); n = 5 (PE7); n = 4 (PE7-SB2). Figure 15h shows the sequences of Igf2 off-target sites, with mismatched bases highlighted in red. Figures 15i and 15j are diagrams showing structures predicted by AlphaFold 3. Figure 15i shows the structure and structural conservation of MLH1-SB bound to human or mouse MLH1. Figure 15j shows the predicted structures of human and mouse PMS2 together with the MLH1dn-PMS2 complex (MutLα complex).
[0043]
[0044] Detailed description of the invention and preferred embodiments
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. In general, the nomenclature used herein is well known and commonly used in the art.
[0046]
[0047] Prime editing is a technology that introduces precise genetic modifications into the target genome guided by prime editing RNA (pegRNA), based on a minimal construct consisting of a Cas nickase fused to a reverse transcriptase. However, mismatch repair (MMR) is reported to be a strong limitation of prime editing, and MLH1dn (dominant-negative MLH1), which inhibits the formation of the MLH1-PMS2 complex (MutLα complex), a key component of MMR, is used to enhance prime editing efficiency.
[0048] To develop a novel MutLα complex (MLH1-PMS2 complex) inhibitor with a smaller size than MLH1dn and a superior prime editing efficiency enhancement effect, the present inventors designed a polypeptide that binds to multiple target amino acid residues of MLH1 using RF diffusion, and designed a novel MLH1 small binder (MLH1-SB) using a novel competitive testing strategy based on AlphaFold 3.
[0049] In an embodiment of the present invention, it was confirmed that the designed MLH1 small binder (MLH1-SB) significantly improved gene editing efficiency in general when used in combination with various previously reported prime editing systems, while having very low side effects such as off-target gene editing and cytotoxicity.
[0050]
[0051] MLH1 binding polypeptide
[0052] Accordingly, the present invention relates, in one aspect, to a novel MLH1 binding polypeptide that binds to one or more amino acid residues of MLH1 and inhibits the formation of the MLH1-PMS2 complex (MutLα complex).
[0053] In the present invention, one or more amino acid residues of the MLH1 are
[0054] - Amino acid residues in the region involved in the interaction between MLH1 and PMS2;
[0055] - Amino acid residues in the nuclease active domain of the MutLα complex; and
[0056] - It may be one or more amino acid residues selected from the group consisting of amino acid residues of the non-competitive anchoring site of MLH1.
[0057] The term “MLH1 (MutL Homolog 1)” of the present invention is a protein that plays a key role in the mismatch repair (MMR) pathway, and is a component of a complex that recognizes and repairs base pair mismatches that occur during DNA replication in cells. After the MutSα (MSH2-MSH6) or MutSβ (MSH2-MSH3) complex recognizes a DNA mismatch, MLH1 receives a mismatch repair signal in an ATP-dependent manner and forms a dimeric complex with proteins such as PMS2 (i.e., MutLα complex), thereby participating in the excision and correction of the mismatch site. MLH1 binds to PMS2 and others at the C-terminal region to form a functional MMR complex, and this binding is essential for structural stability and active site formation for MMR activity.
[0058] The term “MutLα” or “MutLα complex” of the present invention refers to a eukaryotic MutL homologous protein complex that functions in the DNA mismatch repair (MMR) pathway, which is formed by a heterodimer of MLH1 protein and PMS2 protein. As used herein, “MutLα” includes functional variants, homologs, isoforms, mutants, or functional fragments thereof of the above proteins, and includes cases where such variants exhibit substantially the same biological activity as MutLα in DNA mismatch repair.
[0059] In the present invention, the MLH1 may be derived from various organisms without limitation. In the present invention, the MLH1 may be, for example, MLH1 derived from an animal, preferably a mammal, and more preferably a human, but is not limited thereto.
[0060] In an embodiment of the present invention, a polypeptide that binds to a plurality of specific amino acid residues based on human MLH1 (UniProt ID: P40692) represented by SEQ ID NO: 1 was prepared, but is not limited thereto. In the present invention, the MLH1 may include a sequence having at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology with SEQ ID NO: 1. In the present invention, the MLH1 may be or include the amino acid sequence of SEQ ID NO: 1.
[0061] In the present invention, the MLH1 may be a human MLH1 comprising the amino acid sequence of SEQ ID NO: 1, and in the amino acid sequence of SEQ ID NO: 1
[0062] The amino acid residues in the region involved in the interaction between the above MLH1 and PMS2 are W538 and / or Q542;
[0063] The amino acid residues in the nuclease active region of the MutLα complex are Y750;
[0064] The amino acid residue of the non-competitive anchoring site of the above MLH1 may be M682.
[0065] In another aspect, the present invention relates to a polypeptide that binds to any one or more amino acid residues selected from the group consisting of W538, Q542, M682 and Y750 of the MLH1 protein.
[0066] In the present invention, the MLH1 protein may be represented by or include SEQ ID NO: 1, but is not limited thereto.
[0067] In the present invention, it can be clearly understood that when the MLH1 protein has a sequence different from SEQ ID NO: 1, a residue corresponding to each amino acid residue can be derived through alignment with SEQ ID NO: 1.
[0068] In the present invention, the MLH1 binding polypeptide is
[0069] One or more amino acid residues selected from the region involved in the interaction between MLH1 and PMS2; and
[0070] It may be characterized by binding to one or more amino acid residues selected from the nuclease active region of the MutLα complex.
[0071] For example, the MLH1 binding polypeptide may be characterized by binding to residues W538 and Y750. For another example, the MLH1 binding polypeptide may be characterized by binding to residues Q542 and Y750. For yet another example, the MLH1 binding polypeptide may be characterized by binding to residues W538, Q542, and Y750.
[0072] In the present invention, the MLH1 binding polypeptide may comprise an amino acid sequence selected from SEQ ID NO: 2 to SEQ ID NO: 21. In the present invention, the MLH1 binding polypeptide may comprise a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology with an amino acid sequence selected from SEQ ID NO: 2 to SEQ ID NO: 21. In the present invention, the MLH1 binding polypeptide may preferably comprise an amino acid sequence of SEQ ID NO: 5.
[0073] In the present invention, the MLH1 binding polypeptide may be an amino acid sequence selected from SEQ ID NO: 2 to SEQ ID NO: 21 or a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology thereto. The MLH1 binding polypeptide may preferably be an amino acid sequence of SEQ ID NO: 5.
[0074]
[0075] In the present invention, the MLH1 binding polypeptide is
[0076] One or more amino acid residues selected from the region involved in the interaction between MLH1 and PMS2;
[0077] One or more amino acid residues selected from the nuclease active domain of the MutLα complex; and
[0078] It may be characterized by binding to one or more amino acid residues selected from the non-competitive anchoring site of MLH1.
[0079] For example, the polypeptide may be characterized by binding to residues W538, M682, and Y750. In the present invention, the polypeptide may be characterized by binding to residues Q542, M682, and Y750. In the present invention, the polypeptide may be characterized by binding to residues W538, Q542, M682, and Y750.
[0080] In the present invention, the MLH1 binding polypeptide may include an amino acid sequence selected from SEQ ID NO: 22 to SEQ ID NO: 61. In the present invention, the MLH1 binding polypeptide may include a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology to an amino acid sequence selected from SEQ ID NO: 22 to SEQ ID NO: 61.
[0081] In the present invention, the MLH1 binding polypeptide may preferably comprise an amino acid sequence selected from SEQ ID NOs: 2, 5, 6, 8, 9, 13, 15, 16, 18, 20, 21, 22, 23, 24, 25, 27, 28, 30, 31, 32, 33, 37, 38, 40, 41, 42, 43, 44, 46, 48 and 51. In the present invention, the MLH1 binding polypeptide may more preferably comprise an amino acid sequence selected from SEQ ID NOs: 5, 22, 23, 32, 33, 38, 40, 41, 42 and 44. In the present invention, the MLH1 binding polypeptide may most preferably comprise an amino acid sequence of SEQ ID NO: 44.
[0082] In the present invention, the MLH1 binding polypeptide may preferably be an amino acid sequence selected from SEQ ID NOs: 2, 5, 6, 8, 9, 13, 15, 16, 18, 20, 21, 22, 23, 24, 25, 27, 28, 30, 31, 32, 33, 37, 38, 40, 41, 42, 43, 44, 46, 48 and 51. In the present invention, the MLH1 binding polypeptide may more preferably be an amino acid sequence selected from 5, 22, 23, 32, 33, 38, 40, 41, 42 and 44. In the present invention, the MLH1 binding polypeptide may most preferably be an amino acid sequence of SEQ ID NO: 44.
[0083]
[0084]
[0085]
[0086]
[0087] In the present invention, the MLH1 binding polypeptide may be characterized by binding to MLH1 and inhibiting the interaction between MLH1 and PMS2 and the formation of a MutLα complex.
[0088] In the present invention, the MLH1 binding polypeptide may be characterized in that, when forming a trimer with MHL1 and PMS2 using AlphaFold 3, the average interface distance between Cα atoms at the MLH1-C / PMS2-C interface is at least 6 Å or more, preferably at least 7 Å or more, more preferably at least 8 Å or more, and most preferably at least 9.7724 Å or more.
[0089] In the present invention, the MLH1 binding polypeptide may be characterized by inhibiting the mismatch repair (MMR) pathway.
[0090] In the present invention, the MLH1 binding polypeptide can significantly improve gene editing efficiency, for example, can improve the efficiency of Cas / CRIPSR gene editing or prime editing using a prime editor.
[0091]
[0092] Nucleic acid encoding MLH1 binding polypeptide
[0093] In another aspect, the present invention relates to a nucleic acid encoding the MLH1 binding polypeptide of the present invention.
[0094] The term "nucleic acid" in the present invention refers to a polynucleotide of any length, which may include DNA, RNA, or a combination of DNA and RNA. In one embodiment, the nucleic acid may be a deoxyribonucleotide, a ribonucleotide, as well as a modified nucleotide or base, and / or an analog thereof, or any substrate capable of being incorporated by a DNA or RNA polymerase.
[0095] The term "coding nucleic acid" in the present invention refers to a nucleic acid sequence encoding a specific protein or polypeptide. In the art, when the sequence of a specific protein or polypeptide is disclosed, methods for designing or deriving a nucleic acid encoding it are well known, and can be further codon optimized.
[0096] Accordingly, the nucleic acid encoding the MLH1 binding polypeptide of the present invention can be readily understood from the description of the MLH1 binding polypeptide herein.
[0097] In the present invention, the nucleic acid encoding the MLH1 binding polypeptide may be or include a nucleic acid sequence selected from SEQ ID NO: 62 to SEQ ID NO: 121. In the present invention, the nucleic acid encoding the MLH1 binding polypeptide may be or include a sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% homology to a nucleic acid sequence selected from SEQ ID NO: 62 to SEQ ID NO: 121. The nucleic acid encoding the above MLH1 binding polypeptide may preferably be or include a nucleic acid sequence selected from SEQ ID NOs: 62, 65, 66, 68, 69, 73, 75, 76, 78, 80, 81, 82, 83, 84, 85, 87, 88, 90, 91, 92, 93, 97, 98, 100, 101, 102, 103, 104, 106, 108 and 111. In the present invention, the nucleic acid encoding the MLH1 binding polypeptide may more preferably be or include a nucleic acid sequence selected from SEQ ID NOs: 65, 82, 83, 92, 93, 98, 100, 101, 102 and 104. In the present invention, the MLH1 binding polypeptide may more preferably be or include a nucleic acid sequence of SEQ ID NO: 104.
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106] A vector comprising a nucleic acid encoding an MLH1 binding polypeptide of the invention
[0107] A nucleic acid encoding the MLH1 binding polypeptide of the present invention can be provided to a target cell via a gene delivery vehicle such as a vector.
[0108] Accordingly, the present invention, from another aspect, relates to a vector comprising a nucleic acid encoding the MLH1 binding polypeptide of the present invention.
[0109] The term “vector” in the present invention means any form of genetic means for delivering and / or expressing a target gene. For example, adenovirus vectors, retrovirus vectors, adeno-associated virus vectors, vaccinia virus (Puhlmann M. et al., Human Gene Therapy, 10:649-657(1999); Ridgeway, 467-492(1988); Baichwal and Sugden, In: Kucherlapati R, ed. Gene transfer. New York: Plenum Press, 117-148(1986) and Coupar et al., Gene, 68:1-10(1988)), lentivirus (Wang G. et al., J. Clin. Invest., 104(11):R55-62(1999)), herpes simplex virus (Chamber R., et al., Proc. Natl. Acad. Sci USA, 92:1411-1415(1995)), poxvirus (GCE, NJL, Krupa M, Esteban M., Curr Gene Ther 8(2):97-120(2008)), viral vectors such as reovirus, measles virus, Semliki Forest virus, poliovirus and cytomegalovirus; cosmid vectors; plasmid vectors (Sambrook et al., 1989) and non-viral vectors such as minicircles (Yew et al. 2000 Mol Ther 1(3), 255-62).
[0110] A vector may generally include, but is not limited to, one or more components selected from a signal sequence, an origin of replication, an enhancer element, a promoter, and a transcription terminator sequence. In the present invention, a nucleic acid encoding an MLH1 binding polypeptide may be operably linked to a promoter and a transcription terminator sequence to express the MLH1 binding polypeptide.
[0111] "Operably linked" means a functional association between a nucleic acid expression regulatory sequence (e.g., a promoter, a signal sequence, or an array of transcription factor binding sites) and another nucleic acid sequence, whereby the regulatory sequence regulates transcription and / or translation of the other nucleic acid sequence.
[0112] In the present invention, the promoter may be a prokaryotic promoter or a eukaryotic promoter, and for example, prokaryotic promoters such as the tac promoter, the lac promoter, the lacUV5 promoter, the lpp promoter, the pLλ promoter, the pRλ promoter, the rac5 promoter, the amp promoter, the recA promoter, the SP6 promoter, the trp promoter, and the T7 promoter; promoters derived from the genome of mammalian cells such as the metallothionine promoter, the β-actin promoter, the human hemoglobin promoter, and the human muscle creatine promoter; It may be a promoter derived from a mammalian virus, such as, but not limited to, the adenovirus late promoter, the vaccinia virus 7.5K promoter, the SV40 promoter, the cytomegalovirus (CMV) promoter, the tk promoter of HSV, the mouse mammary tumor virus (MMTV) promoter, the LTR promoter of HIV, the promoter of Moloney virus, the promoter of Epstein-Barr virus (EBV), and the promoter of Rous sarcoma virus (RSV).
[0113]
[0114] In the present invention, the vector may additionally comprise a nucleic acid encoding a Cas nickase protein.
[0115] In the present invention, the Cas nickase protein may be a Cas protein variant that has lost the endonuclease activity of cleaving DNA double strands in wild-type Cas nuclease and has nickase activity.
[0116] In the present invention, the Cas protein is Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12g, Cas12h, Cas12i, Cas12j, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, CsMT2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, It may be an endonuclease of Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3 or Csf4, preferably an endonuclease of Cas9 or Cas12a, but is not limited thereto.
[0117] In the present invention, the Cas protein is, for example, Corynebacter, Sutterella, Legionella, Treponema, Filifactor, Eubacterium, Streptococcus (Streptococcus pyogenes), Lactobacillus, Mycoplasma, Bacteroides, Flaviflaviivola, Flavobacterium, Azospirillum, Gluconacetobacter, Neisseria, Roseburia, Parvibaculum, Staphylococcus (Staphylococcus aureus), A microbial genus comprising an ortholog of a Cas protein selected from the group consisting of Nitratifractor, Corynebacterium and Campylobacter, which may be simply isolated or recombinant therefrom, but is not limited thereto.
[0118] In the present invention, the Cas nickase can introduce a nick in the strand where the base conversion has occurred or in the opposite strand (e.g., the strand opposite to the strand where the base conversion has occurred) (e.g., a nick is introduced at a position corresponding to between the 3rd and 4th nucleotides in the 5' end direction of the PAM sequence in the opposite strand to the strand where the PAM is located). Such a mutation (e.g., an amino acid substitution, etc.) may occur in the catalytically active domain (e.g., the RuvC catalytic domain in the case of Cas9).
[0119] For Cas9 derived from Streptococcus pyogenes, the mutation may include a mutation in which one or more other amino acids are substituted, such as a catalytic aspartate residue (e.g., aspartic acid at position 10 (D10)), glutamic acid at position 762 (E762), histidine at position 840 (H840), asparagine at position 854 (N854), asparagine at position 863 (N863), aspartic acid at position 986 (D986), etc. In this case, the other amino acid to be substituted may be, but is not limited to, alanine.
[0120] In the present invention, the Cas protein variant having the nickase activity includes, but is not limited to, for example, SpCas9 (H840A), SpCas9 (D10A), SpCas9 (H840A / D10A), SaCas9 (H840A), eSpCas9 (H840A), SpCas9-HF1 (H840A), HypaCas9 (H840A), etc. In the present invention, the Cas protein variant having the nickase activity includes, but is not limited to, for example, AsCas12a, LbCas12a, Cas12f (Cas14), CasX, CasY, CasPhi, etc.
[0121]
[0122] In the present invention, the vector may additionally include a nucleic acid encoding reverse transcriptase (RT).
[0123] The term “reverse transcriptase” of the present invention refers to an enzyme capable of synthesizing DNA using RNA as a template.
[0124] In the present invention, the reverse transcriptase may be a wild-type reverse transcriptase derived from various viruses. In the present invention, the reverse transcriptase may be an artificially synthesized reverse transcriptase or a reverse transcriptase variant. Examples thereof include, but are not limited to, M-MLV reverse transcriptase (Moloney Murine Leukemia Virus RT), HIV reverse transcriptase (e.g., HIV-1 RT), AMV RT (Avian Myeloblastosis Virus), etc.
[0125] In the present invention, the reverse transcriptase may be a recombinant reverse transcriptase, and various engineered reverse transcriptases known in the technical field of the present invention may be selected and used, but are not limited thereto.
[0126]
[0127] In the present invention, the vector may additionally include a nucleic acid encoding a prime editor.
[0128] The term “prime editor” of the present invention comprehensively refers to a fusion protein comprising Cas nickase and reverse transcriptase (RT) as minimum structural units, and various prime editor architectures including additional components such as ngRNA (nick gRNA), MLH1 binder, RNA-binding exonuclease protection factor (e.g., La motif), etc. have been reported. For example, the prime editor architectures include, but are not limited to, PE1, PE2, PE3, PE3b, PEmax, PE4, PE5, PE6, PE6a, PE6b, PE6c, PE6d, PE6e, PE6f, PE7, TwinPE, etc.
[0129] In the present invention, the Cas nickase and reverse transcriptase of the prime editor may be directly connected in series or may be connected by a linker. In the present invention, various linkers may be used as long as the prime editing activity of the prime editor is maintained. In the present invention, the linker may preferably be a flexible linker, and for example, the flexible linker may be an AS linker (e.g., (AnS)m (n and m are each integers from 1 to 10)), a GS linker (e.g., (GS)n, (GGS)n, (GSGGS)n or (GnS)m (n and m are each integers from 1 to 10)), an XTEN linker (e.g., XTEN16, XTEN72, etc.), but is not limited thereto.
[0130] In the present invention, the prime editor may additionally include a single-stranded DNA binding domain.
[0131] The term "single-stranded DNA-binding domain (ssDBD)" as used herein refers to an independently folded protein domain comprising one or more structural motifs that recognize single-stranded DNA. The DNA-binding domain may recognize a specific DNA sequence or have a general affinity for DNA, and some DNA-binding domains may also contain nucleic acids in their folded structures. Examples of the single-stranded DNA-binding domain include, but are not limited to, Rad52, Replication Protein A (RPA), Single-Stranded DNA Binding protein (SSB), and various viral-derived DNA binding proteins (DBPs).
[0132] In the present invention, the single-stranded DNA binding domain may be connected directly or through a linker to the N'-terminus or C'-terminus of Cas nickase, the N'-terminus or C'-terminus of reverse transcriptase, etc., but is not limited thereto.
[0133]
[0134] In the present invention, the nucleic acid encoding the prime editor and the nucleic acid encoding the MLH1 binding polypeptide may be directly linked at the N'-terminus or the C'-terminus or linked via a linker. A nucleic acid sequence encoding an additional domain may be included between the nucleic acid encoding the prime editor and the nucleic acid encoding the MLH1 binding polypeptide.
[0135] In the present invention, the prime editor and MLH1 binding polypeptide may be expressed as a fused protein, expressed as a single fused protein and then separated, or expressed as separate proteins within a single mRNA.
[0136] In the present invention, various linkers may be used without limitation as long as the desired activity of the prime editor and the MLH1 binding polypeptide is maintained. For example, when expressed as a single fusion protein, they may be linked by a flexible linker. In another example, they may be linked by a protease recognition site (e.g., a TEV protease recognition site, an HCV NS3 protease recognition site, etc.) so that they can be separated after expression as a single fusion protein. In another example, they may be linked by a 2A peptide sequence (e.g., P2A, T2A, E2A, F2A, etc.) or an IRES sequence so that they can be expressed as separate proteins within a single mRNA. In the embodiments of the present invention, a vector was constructed using a sequence encoding a 2A peptide so that the prime editor and the MLH1 binding polypeptide are expressed as separate proteins, but this is not limiting.
[0137] In the present invention, the vector may contain a nucleic acid encoding an MLH1 binding polypeptide repeated one or more times. In the present invention, when the vector contains a nucleic acid encoding an MLH1 binding polypeptide repeated one or more times, the nucleic acid encoding each MLH1 binding polypeptide may be linked to a nucleic acid encoding a 2A peptide, but is not limited thereto.
[0138]
[0139] In the present invention, the vector may additionally include a nuclear localization sequence (NLS).
[0140] The term "nuclear localization sequence (NLS)" of the present invention refers to an amino acid sequence that plays a role in transporting a prime editor into the cell nucleus, and generally functions to transport it into the cell nucleus through a nuclear pore (Kalderon D, et al., Cell 39:499-509 (1984); Dingwall C, et al., J CellBiol. 107(3):8419 (1988)).
[0141]
[0142] In the present invention, the vector may include DNA encoding pegRNA (prime editing guide RNA).
[0143] The term "pegRNA (prime editing guide RNA)" as used herein may include a binding region (or guide sequence) that binds to a target genome for editing, a tracrRNA scaffold sequence, a primer binding site (PBS) required for initiating reverse transcription, and a proofreading sequence containing a desired genetic change. The sequence containing the proofreading sequence acts as a template for a reverse transcriptase. The reverse transcriptase template contains the desired proofreading sequence and has homology to a genomic DNA locus. The proofreading sequence is a heterologous sequence and contains a target sequence to be corrected in the genome. The binding region may be arbitrarily positioned in the 5' direction or the 3' direction of the reverse transcriptase template, and specifically, the binding region may be positioned in the 3' direction of the reverse transcriptase template. The binding region may include a sequence complementary to a genomic DNA strand nicked by a nuclease or a variant thereof, such as a nickase, included in a prime editor protein. The above binding region can hybridize to a target site, which can act as an activation initiation point of a reverse transcriptase. The above binding region can include 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 25 or more complementary sequences having 80% or more, for example, 85% or more, 90% or more, 95% or more, 97% or more, 98% or more, 99% or more, or 100% homology with the sequence of the target site.
[0144]
[0145] In a preferred embodiment of the present invention, the vector may be characterized by, but is not limited to, comprising the following structure:
[0146] [Prime Editor]-L1-[MLH1 binding polypeptide];
[0147] [Prime Editor]-L1-[MLH1 binding polypeptide]-[NLS];
[0148] [Prime Editor]-L1-[MLH1 binding polypeptide]-L2-[MLH1 binding polypeptide]; or
[0149] [Prime Editor]-L1-[MLH1 binding polypeptide]-[NLS]-L2-[MLH1 binding polypeptide],
[0150] Here, L1 and L2 are linkers that are each independently selected.
[0151]
[0152] For use in Prime Editing (compositions, kits, methods, etc.)
[0153] In the examples of the present invention, it was confirmed that the MLH1 binding polypeptide of the present invention can be usefully used as an MMR inhibitor of prime editing technology by significantly improving editing efficiency in various prime editing systems while exhibiting low off-target editing and low cytotoxicity.
[0154] Accordingly, the present invention relates, from another aspect, to the use of the MLH1 binding polypeptide of the present invention or a nucleic acid encoding the same for gene editing.
[0155] In another aspect, the present invention relates to a use for the preparation of a composition or kit for gene editing of an MLH1 binding polypeptide of the present invention or a nucleic acid encoding the same.
[0156] In another aspect, the present invention relates to a composition or kit for gene editing comprising the MLH1 binding polypeptide of the present invention or a nucleic acid encoding the same.
[0157] The term "genome editing" of the present invention is interchangeable with "gene correction" and refers to a method of altering a nucleic acid sequence by inducing mutations, such as selective substitutions, insertions, or deletions, in a target sequence. Such specific target sequences include, but are not limited to, a chromosomal region, a gene, a promoter, an open reading frame, or any nucleic acid sequence.
[0158] The term "target sequence" in the present invention refers to a nucleic acid sequence targeted by pegRNA. The target sequence may be a sequence expected to be targeted by the pegRNA. The target sequence may be a portion of a known genomic sequence, or a sequence desired to be edited by a skilled artisan using the system of the present invention.
[0159] In the present invention, the substitution means that an amino acid residue of a gene sequence is replaced with a different amino acid residue.
[0160] In the present invention, the deletion means the loss of a portion (at least one amino acid) of a gene sequence. In the present invention, the deletion may be a deletion of amino acids having a length of less than 24 bp, preferably less than 20 bp, more preferably less than 16 bp, and most preferably less than 13 bp.
[0161] In the present invention, the insertion means the addition of a portion (at least one amino acid) of a gene sequence. In the present invention, the insertion may be an insertion of an amino acid having a length of less than 24 bp, preferably less than 20 bp, more preferably less than 16 bp, and most preferably less than 13 bp.
[0162] In the present invention, the MLH1 binding polypeptide can be directly delivered to cells in the form of a protein. In the present invention, when the MLH1 binding polypeptide is directly delivered in the form of a protein, it can be delivered by fusing with a cell-penetrating peptide (CPP) or using a method known in the art, such as liposomes or nanoparticles.
[0163] In the present invention, more preferably, the nucleic acid encoding the MLH1 binding polypeptide can be directly delivered to a cell for expression. Accordingly, in the present invention, the composition or kit may include a vector comprising a nucleic acid encoding the MLH1 binding polypeptide of the present invention. Examples of such vectors are as described above.
[0164]
[0165] In the present invention, the composition or kit may additionally comprise a Cas nickase or a nucleic acid encoding the same.
[0166] In the present invention, the composition or kit may additionally include a reverse transcriptase enzyme or a nucleic acid encoding the same.
[0167] In the present invention, the composition or kit may additionally comprise a prime editor or a nucleic acid encoding the same.
[0168] In the present invention, the nucleic acid encoding the Cas nickase, reverse transcriptase, or prime editor can be cloned into the same vector as the nucleic acid encoding the MLH1 binding polypeptide. For an embodiment of cloning into the same vector, refer to the above description.
[0169] In the present invention, the nucleic acid encoding the prime editor can be cloned into a different vector from the nucleic acid encoding the MLH1 binding polypeptide and introduced sequentially or co-transfected.
[0170] In the present invention, the composition or kit may additionally comprise pegRNA. In the present invention, the pegRNA may be cloned into the same vector as the nucleic acid encoding the MLH1 binding polypeptide. In the present invention, the pegRNA may be cloned into different vectors and introduced sequentially or jointly.
[0171] In the present invention, the composition or kit may additionally comprise ngRNA. In the present invention, the ngRNA may be cloned into the same vector as the nucleic acid encoding the MLH1 binding polypeptide. In the present invention, the ngRNA may be cloned into different vectors and introduced sequentially or jointly.
[0172]
[0173] In another aspect, the present invention relates to a gene editing method comprising a step of treating or administering the composition to a cell or a subject.
[0174] In the present invention, the cell may be a eukaryotic cell (e.g., a fungus such as yeast, a cell derived from a eukaryotic animal and / or a eukaryotic plant (e.g., an embryonic cell, a stem cell, a somatic cell, a germ cell, etc.), a eukaryotic animal (e.g., a primate such as a human or monkey, a dog, a pig, a cow, a sheep, a goat, a mouse, a rat, etc.), or a eukaryotic plant (e.g., an algae such as green algae, corn, soybeans, wheat, rice, etc.), but is not limited thereto.
[0175] In the present invention, the subject may be a human. In the present invention, the subject may be a living organism other than a human, such as an animal, plant, or fungus, but is not limited thereto.
[0176] In the embodiments of the present invention, cells such as HEK293T, HeLa, HEK293, and hiPSCs were used, but are not limited thereto.
[0177]
[0178] Hereinafter, the present invention will be described in more detail through examples. These examples are intended solely to illustrate the present invention, and it will be apparent to those skilled in the art that the scope of the present invention is not limited by these examples.
[0179]
[0180] Example
[0181] Example 1: Materials and Methods
[0182] cell lines
[0183] HEK293T (ATCC, CRL-11268), HeLa (ATCC, CCL-2), and HEK293 (ATCC, CRL-1573) cells were cultured in DMEM medium (Welgene, LM001-05) supplemented with 10% FBS (Welgene, PK004) and 1% antibiotics (Welgene, LS203-1). Cells were maintained in a humidified incubator at 37°C and 5% CO₂. When subcultured, cells were washed with DPBS (Welgene, LB001-01) and detached using 0.25% Trypsin-EDTA (Welgene, LS015). After detachment, the Trypsin-EDTA reaction was stopped by adding DPBS containing 10% FBS. Cells were then counted and seeded onto cell culture plates containing fresh culture medium.
[0184] hiPSCs (CMC-hiPSC-022, Catholic University of Korea) were cultured in Essential 8™ medium (ThermoScientific, A1517001) on cell culture plates coated with iMatrix-511 (REPROCELL, NP892-012). During passage, hiPSCs were washed with DPBS and detached using ACCUTASE™ (Stem Cell Technologies, 07922). The detached cells were washed twice with DPBS and seeded onto iMatrix-511-coated culture plates containing Essential 8™ medium. To prevent cell death due to cell dissociation, 10 μM ROCK inhibitor solution (Y27632 dihydrochloride, MCE, HY-10583) was added to the culture medium during passage. hiPSCs were maintained at 37°C in an atmosphere with 5% CO₂.
[0185]
[0186] animal
[0187] All animal experiments were approved by the Korea University Institutional Animal Care and Use Committee (KU-IACUC) (KOREA-2024-0179). Six-week-old female ICR mice were housed in a specific pathogen-free (SPF) facility under a 12-h day / night cycle, at a temperature of 20–26°C and a humidity of 40%–60%.
[0188]
[0189] Production of SB (Small Binder) Proteins Using RF Diffusion
[0190] SB proteins were generated using a customized script that integrated the Colab-based RFdiffusion version with ProteinMPNN and AlphaFold 2. The dimeric structures of MLH1-C (residues 486-751) and PMS2-C (residues 606-862) were predicted using the official website of AlphaFold 3. For binder design, the generated MLH1-C structure was used as a target protein, including three designated hotspots (W538, Q542, Y750) for the MC binder and an additional hotspot (M682) for the MCA binder.
[0191] Most adjustable parameters are set to default values as recommended by the RFdiffusion developers.
[0192] - RF diffusion subsection: iterations = 50
[0193] - ProteinMPNN subsection: num_seqs = 2, mpnn_sampling_temp = 0.0001, rm_aa = 'C'
[0194] - AlphaFold 2 subsection: initial_guess = 'TRUE', num_recycles = 3, use_multimer = 'TRUE'
[0195] - Binder length: Set to 50-150 residues for MC binder and 70-150 residues for MCA binder.
[0196] A total of 1,978 MC binders and 821 MCA binders were generated, and filtering was repeated until a sufficient number of candidates were secured.
[0197] The candidate filtering criteria were as follows:
[0198] - interaction pAE ≤ 7.5
[0199] - pLDDT ≥ 0.85
[0200] - RMSD ≤ 1.5 Å
[0201] Candidates that passed the filtering were used in the AlphaFold 3 competition test.
[0202]
[0203] AlphaFold 3 Competitive Testing
[0204] Using the sequences of the selected candidates, the structure of the trimer complex with MLH1-C and PMS2-C was predicted through the official website of AlphaFold 3. To assess structural completeness, the average interface distance between Cα atoms at the MLH1-C / PMS2-C interface was calculated using the following five residue pairs: L540, L743, Q542, C756, and N739 of MLH1-C and I688, I853, N683, E705, and Q861 of PMS2-C. For comparison, the interface distance in the dimeric MLH1-C / PMS2-C structure was calculated to be 9.0157 Å.
[0205] A threshold of 9.7724 Å was set, and candidates with an interface distance exceeding this threshold were distinguished from those without.
[0206] The final selected candidate is
[0207] - Amino acid residues in the region involved in the interaction between MLH1 and PMS2; and a total of 20 candidates (MC) that bind to amino acid residues in the nuclease active region of the MutLα complex;
[0208] - A total of 23 candidates (MCAs) with an interface distance exceeding a threshold among the candidates that additionally bind to amino acid residues of the non-competitive anchoring site of MLH1 along with the above two regions; and
[0209] - In addition to the above two regions, a total of 17 candidates (Failed-MCA) were derived among the candidates that bind additionally to amino acid residues in the non-competitive anchoring site of MLH1 and whose interface distance does not exceed the threshold value (Table 1).
[0210]
[0211] Plasmid construction
[0212] The protein sequence of MLH1-SB generated by RFdiffusion was converted into a nucleic acid sequence using GenSmart Optimization, and the detailed sequence is shown in Table 2.
[0213]
[0214] The codon-optimized MLH1-SB sequence was synthesized as Gene Fragments by Twist Bioscience. These gene fragments were amplified by PCR and then inserted into the PCR-linearized CMV-PuroR-2A backbone vector by Gibson Assembly.
[0215] To construct the MLH1dn expression vector, the MLH1dn sequence was amplified by PCR from the pCMV-PE4stem plasmid (Addgene, #208768). The amplified MLH1dn sequence was inserted into the PCR-linearized CMV-2A-Puro backbone vector by Gibson Assembly.
[0216] To construct pCMV-PE7, total cellular RNA was extracted from HEK293T cells using TRIzol-Chloroform. 1 μg of RNA was then reverse-transcribed using ReverTra Ace® qPCR RT Master Mix (TOYOBO, FSK-101F) according to the manufacturer's instructions. La(1-194) cDNA (also known as 'SSB') was amplified using KOD-Plus-Neo polymerase (TOYOBO, TOKOD-401) and cloned into the pCMV-PEmax vector (Addgene, #174820) with a SGGS×2-XTEN16-SGGS×2 linker using BsrGI-HF (NEB, R3575S) and SapI (NEB, R0569S).
[0217] To generate the MLH1-SB-NLS sequence, the MLH1-SB codon was reoptimized using GenSmart Optimization. The reoptimized MLH1-SB-NLS sequence was synthesized by BIONICS, amplified by PCR, and cloned into the PE7 vector using KpnI-HF (NEB, R3142L) and AgeI-HF (NEB, R3552L) restriction enzymes.
[0218] To construct the PE7-MLH1-SB-NLS vector, site-directed mutagenesis was performed using Platinum™ SuperFi II DNA Polymerase (Invitrogen, 361250). The codon-reoptimized MLH1-SB-NLS sequence was assembled into the PE7-MLH1-SB construct via Gibson Assembly, generating a PE7 vector containing both MLH1-SB and MLH1-SB-NLS.
[0219] To construct PE7-N-term-SB and PE7-C-term-SB vectors, MLH1-SB was inserted into the 5' or 3' end of the PE7 vector, and PCR-linearized to include the SGGS-XTEN16-SGGS linker at the corresponding end. PCR was performed using the KOD Multi & Epi PCR kit (TOYOBO), and Gibson Assembly was performed using Gibson Assembly Master Mix (NEB, E2611).
[0220] Transfection and genomic DNA extraction
[0221] To transfect plasmid DNA, HEK293T, HEK293, and HeLa cells were washed with DPBS (Welgene, LB001-01), detached using 0.25% Trypsin-EDTA (Welgene, LS015), and washed with DMEM containing 10% FBS and 1% Antibiotic Antimycotic solution (Welgene, Ls203-01). After counting, the cells were seeded into 48-well cell culture plates (SPL, 30048) at a density of 3 × 10⁴ per well.
[0222] For the MLH1-SB assay (Fig. 1F), 1.5 × 10⁴ HeLa cells were seeded in a 96-well cell culture plate. After 24 h, the medium was replaced with fresh DMEM containing 10% FBS and 1% Antibiotic Antimycotic solution (Welgene, Ls203-01). MLH1-SB assay (Fig. 1F): 60 ng of pCMV-PE2 vector (Addgene, 132775), 20 ng of pegRNA vector, and 20 ng of MLH1-SB expression vector were mixed with 0.5 μl of Lipofectamine 2000 (ThermoFisher, 11668027) according to the manufacturer's instructions and transfected into HeLa cells.
[0223] PE2 or PE7 and MLH1dn or MLH1-SB experiments (Fig. 5): 120 ng of pCMV-PEmax vector (Addgene, 174820) or PE7 vector, 40 ng of pegRNA vector, control vector (CMV-PuroR vector), and 100 ng of MLH1dn or MLH1-SB expression vector were mixed with 1 μl of Lipofectamine 2000 and transfected into HeLa and HEK293T cells.
[0224] PE7-SB2 screening experiment (Fig. 7a): 180 ng of each prime editor variant (PE7, PE7-dn, PE7-SB, PE7-SB-NLS, PE7-2xSB, PE7-N-term-SB, PE7-C-term-SB) and 60 ng of pegRNA were mixed with 1 μl of Lipofectamine 2000 and transfected into HeLa cells.
[0225] PE7-SB2 test experiment (Fig. 7c-h): HeLa and HEK293 cells were transfected with 1 μl of Lipofectamine 2000 by mixing 180 ng of pCMV-PEmax vector (Addgene, 174820), PE7 vector or PE7-SB2 vector, 60 ng of pegRNA vector, and 20 ng of nicking gRNA (+ngRNA condition).
[0226] For hiPSC transfection, cells were detached using ACCUTASE™ (Stem Cell Technologies, 07922), washed with DPBS, and counted. 3 × 10⁴ cells were seeded into 48-well cell culture plates. 180 ng of pCMV-PEmax vector (Addgene, 174820), PE7 vector, or PE7-SB2 vector, 60 ng of pegRNA vector, and 20 ng of nicking gRNA (+ngRNA condition) were mixed with 1 μl of Lipofectamine Stem (ThermoFisher, STEM00015) and transfected into hiPSCs according to the manufacturer's instructions.
[0227] Seventy-two hours after transfection, the medium was removed, and the cells were washed with DPBS. The washed cells were lysed using lysis buffer (25 μl for 96-well experiments and 60 μl for 48-well experiments). Lysis buffer consisted of the following components: 40 mM Tris (pH 8.0), 1% Tween-20, 0.2 mM EDTA, 0.2% Nonidet P-40, and 2 mg / ml protease K. Genomic DNA extracts were incubated at 60°C for 15 min to activate proteinase K, followed by inactivation at 95°C for 5 min.
[0228] The pegRNA used in the examples is as described in Table 3.
[0229]
[0230] Electroporation and genomic DNA extraction
[0231] Patient-derived fibroblasts (50,000 cells / transfection) were transduced using the Neon transfection system 10 μl kit (Thermo Fisher Scientific, MPK1025) under the following conditions: voltage 1,600 V, pulse width 20 ms, and number of cycles 1.
[0232] Fibroblasts were transduced with 480 ng of a plasmid expressing PE, 120 ng of a plasmid expressing pegRNA, and 120 ng of an overexpression vector such as the SB vector or a control vector (CMV-PuroR vector).
[0233] Seventy-two hours after electroporation, all cells were harvested for next-generation sequencing. The cell pellet was resuspended in lysis buffer containing the following components: 40 mM Tris-HCl (pH 8.0), 1% Tween-20, 0.2 mM EDTA, 0.2% Nonidet P-40 (VWR Life Science, 97064-730), and 2 mg / ml Proteinase K.
[0234] The suspension was incubated at 60°C for 15 minutes to activate Proteinase K, and then heated at 98°C for 5 minutes to inactivate Proteinase K.
[0235]
[0236] Co-immunoprecipitation
[0237] HeLa cells (0.5 million cells) were seeded in 6-well plates 12–16 hours prior to transfection. The following day, cells were transfected with 2 μg of plasmid DNA (MLH1-SB-HA or a combination of MLH1-SB-HA and 3×FLAG-MLH1) using Lipofectamine 2000 (Invitrogen, 11668019) according to the manufacturer's instructions. After transfection, cells were cultured for 48 hours at 37°C and 5% CO₂. After 48 hours, the medium was removed, and cells were washed once with 1X DPBS (Welgene, LB001-01). RIPA buffer supplemented with cOmplete™ Mini Protease Inhibitor Cocktail (Roche, 11836153001) was then added for cell lysis.
[0238] The lysate was passed through a 26G syringe needle several times to ensure complete cell lysis, and then incubated on ice for 30 minutes, vortexing every 10 minutes. The lysate was then incubated overnight at 4°C with Pierce™ Anti-HA Magnetic Beads (Invitrogen, 88836) or anti-FLAG® M2 magnetic beads (Merck, M8823) to allow antigen-antibody binding. The beads used were pre-washed and equilibrated.
[0239] The following day, the beads were separated using a magnetic stand and washed twice with 0.05% TBS-T. Bound proteins were eluted by adding 2X sample buffer and heating at 95°C for 5 minutes. The eluted proteins were electrophoresed on 4%–20% Mini-PROTEAN® TGX™ Precast Protein Gels (Bio-Rad, 4561094) and transferred to 0.45 μm PVDF membranes (Bio-Rad, 1620177).
[0240] 3×FLAG-MLH1 was detected with DYKDDDDK Tag Monoclonal Antibody (Invitrogen, MA1-91878), and MLH1-SB-HA was detected with HA Tag Monoclonal Antibody (Invitrogen, 26183).
[0241]
[0242] LC-MS / MS
[0243] LC-MS / MS analysis was performed at PROTIA Inc. (Seoul, Republic of Korea) using a nano ACQUITY UPLC and LTQ-orbitrap-mass spectrometer (Thermo Electron, San Jose, CA). A BEH C18 1.7 μm, 100 μm × 100 mm column (Waters, Milford, MA, USA) was used for the analysis. Detailed experimental conditions were as follows. Mobile phase A of the LC separation was distilled water containing 0.1% formic acid, and mobile phase B was acetonitrile containing 0.1% formic acid. The chromatographic gradient conditions were a linear increase from 10% B to 40% B over 21 min, a linear increase from 40% B to 95% B over 7 min, and a decrease from 90% B to 10% B over 10 min. The flow rate was 0.5 μL / min. For tandem mass spectrometry, mass spectra were acquired by performing a full mass scan (300–2000 m / z) followed by an MS / MS scan (data-dependent acquisition). Each MS / MS scan was acquired as the average of one microscan on the LTQ. The ion transfer tube temperature was controlled at 275°C, and the spray voltage was 2.3 kV. The normalized collision energy for MS / MS was set to 35%.
[0244] Individual spectra generated from MS / MS were processed using SEQUEST software (Thermo Quest, San Jose, CA, USA), and the generated peak list was requested from an in-house database via the MASCOT program (Matrix Science Ltd., London, UK). Modifications considered during MS analysis were Carbamidomethyl (C), Deamidated (NQ), and Oxidation (M), and the tolerance of the peptide mass was set to 10 ppm. The tolerance of the MS / MS ion mass was 0.8 Da, missed cleavages were allowed up to 2, and the analysis was performed considering charge states (+2, +3). Only significant hits defined by MASCOT probability analysis were used in data analysis.
[0245]
[0246] Targeted deep sequencing
[0247] Each on-target site was amplified through a two-step PCR. In the first step (PCR1), the target region was amplified using primers containing Illumina forward and reverse sequencing adapters. PCR1 was performed using 1 μl of genomic DNA extract with KOD-Multi & Epi polymerase (TOYOBO, KME-101) under the following conditions: initial denaturation at 94°C for 2 min, followed by 32 cycles of 98°C for 10 s, 58°C for 20 s, and 68°C for 12 s.
[0248] The CR1 product was used as a template for the second amplification (PCR2) using Illumina sequencing index primers. PCR2 was performed using 1 μl of the PCR1 product under the following conditions: initial denaturation at 94 °C for 2 min, followed by 20 cycles of 98 °C for 10 s, 60 °C for 20 s, and 68 °C for 12 s.
[0249] PCR products were confirmed by electrophoresis on a 1.5% agarose gel (Condalab, 8100). PCR2 products were then pooled and purified using a PCR purification kit (GeneAll, 103-150).
[0250] The library was sequenced using an Illumina MiniSeq sequencer, and the sequencing results were analyzed using Cas-Analyzer (http: / www.rgenome.net / cas-analyzer / ) and PE-Analyzer (https: / www.rgenome.net / pe-analyzer / ). The list of primers used in this example is as shown in Table 4.
[0251]
[0252]
[0253] RT-qPCR
[0254] Total RNA was extracted from HEK293, HEK293T, and HeLa cells. Cells were first detached (washed with DPBS) and lysed in 500 μl of TRIzol Reagent (Invitrogen, 15596018). RNA extraction was performed according to the manufacturer's instructions.
[0255] For cDNA synthesis, 500 ng of RNA was mixed with 2 μl of PrimeScript™ RT reagent kit (TaKaRa, RR037B) and distilled water (DW) to a final volume of 10 μl. The reaction was performed at 37°C for 15 min, followed by annealing at 85°C for 30 s. The synthesized cDNA was then diluted with 200 μl of DW.
[0256] For qPCR, 10 μl of iTaq Universal SYBR Green Supermix (Bio-Rad, 1725124), 0.5 μl (20 pmol) each of forward and reverse primers, and 9 μl of diluted cDNA were loaded into a 96-well qPCR plate.
[0257] qPCR was performed using a CFX96 Real-Time PCR Detection System (Bio-Rad) under the following cycling conditions: initial denaturation at 95 °C for 3 min, followed by 50 cycles of 95 °C for 10 s, 55 °C for 10 s, and 72 °C for 30 s.
[0258] Relative MLH1 expression levels were calculated using 18s rRNA as an endogenous control. The list of primers used in this example is shown in Table 5.
[0259]
[0260] mRNA sequence analysis
[0261] Total RNA concentration was measured using the Quant-IT RiboGreen assay (Invitrogen, #R11490). RNA integrity was assessed using the TapeStation RNA ScreenTape system (Agilent, #5067-5576). Only high-quality RNA samples with an RNA integrity number (RIN) of 7.0 or higher were used for library construction.
[0262] For each sample, individual libraries were constructed from 10 ng of total RNA using the SMARTer Ultra-Low Input RNA Sample Prep Kit (Clontech Laboratories, Inc., CA, USA). mRNA molecules containing poly-A sequences were purified using poly-T-conjugated magnetic beads. The purified mRNA was fragmented at high temperature using divalent cations. The digested RNA was reverse transcribed into first-strand cDNA using SMARTScribe reverse transcriptase (Clontech) and 3' SMART CDS primers, and second-strand cDNA was then synthesized using DNA Polymerase I, RNase H, and dUTP.
[0263] The generated cDNA fragments underwent end repair, single 'A' base addition, and adapter ligation. The final cDNA library was purified and amplified via PCR. Library quantification was performed using KAPA Library Quantification Kits for Illumina sequencing platforms (KAPA BIOSYSTEMS, #KK4854) according to the qPCR Quantification Protocol Guide. Library quality was assessed using the TapeStation D1000 ScreenTape system (Agilent Technologies, #5067-5582).
[0264] The indexed library was sequenced in paired-end (2 × 100 bp) mode on the Illumina NovaSeq platform (Illumina, Inc., San Diego, CA, USA), and sequencing was performed by Macrogen Incorporated.
[0265]
[0266] mRNA sequencing data processing
[0267] Raw reads obtained from the sequencer were preprocessed to remove low-quality sequences and adapter contamination prior to analysis. The preprocessed reads were aligned to the Homo sapiens (hg38) reference genome using HISAT v2.1.0 (Nat Biotechnol 37, 907–915). The Homo sapiens (hg38) reference genome and annotation data were downloaded from the UCSC Table Browser (http: / genome.ucsc.edu).
[0268] Aligned reads were assembled into transcripts using StringTie v2.1.3b, and expression levels were estimated (Nat Protoc 11, 1650–1667). This provided the relative expression levels of transcripts and genes in each sample as read count values.
[0269]
[0270] Gene expression analysis
[0271] The relative expression level of each gene was measured as read count using StringTie. To identify differentially expressed genes (DEGs), statistical analysis was performed based on within-sample gene expression estimates. Only genes with non-zero read counts across all samples were included. The filtered data were normalized with trimmed mean of M-values (TMM). Differential expression analysis was performed using edgeR's exactTest, and fold-change analysis was performed under the null hypothesis of no difference between groups (Bioinformatics 26, 139-140). The false-positive discovery rate (FDR) was controlled by adjusting the p-value using the Benjamini-Hochberg algorithm. For the DEG set, hierarchical clustering analysis was performed using complete linkage and Euclidean distance as similarity measures. For the list of significant genes, gene enrichment, functional annotation, and pathway analysis were performed using gProfiler (Nucleic Acids Res 47, W191-W198).
[0272]
[0273] Hydrodynamic tail vein injection and genomic DNA extraction
[0274] Female ICR mice (6–8 weeks old) were purchased from Orient Bio (Korea) and acclimatized before the experiment. Plasmids were purified using the NucleoBond Xtra Midi EF, Midi kit (MN, #740420.5). Purified PE7, PE7-SB, and Igf2-pegRNA plasmids were co-injected via the tail vein using a 3 ml syringe in a hydrodynamic manner for 5–8 seconds (Methods Mol Biol 1015, 279–289). Liver tissues were collected for analysis 10 days after injection. Genomic DNA was extracted using the AccuPrep® Genomic DNA Extraction Kit (Bioneer, K-3032).
[0275]
[0276] Liver enzyme analysis
[0277] Serum was isolated from whole blood collected from mice 10 days after injection. Serum ALT and AST levels were measured using FUJI DRI-CHEM slides (GPT / ALT-P III (FUJIFILM, 424911) and GOT / AST-P III (FUJIFILM, 411802)) and a DRI-CHEM NX500i analyzer (FUJIFILM).
[0278]
[0279] Cytotoxicity and cell proliferation assays
[0280] HeLa cells or human-derived primary fibroblasts were harvested and washed once with ice-cold DPBS. Apoptosis was assessed using a Dead Cell Apoptosis Kit (Invitrogen, V13241) containing Annexin V-Alexa Fluor 488 and propidium iodide according to the manufacturer's instructions. Briefly, cells were seeded at a density of 1 × 10 4Cells were suspended in 1X Annexin V binding buffer (BioLegend, 422201) at a concentration of 10 cells / ml, and 100 μl of the suspension was reacted with Annexin V-Alexa Fluor 488 (1:20, v / v) and propidium iodide (final concentration 1 μg / ml) at room temperature in the dark for 20 min. Then, 400 μl of Annexin V binding buffer was added and stored on ice until flow cytometry analysis. For proliferation analysis, cells collected from the same suspension were fixed and permeabilized with 1X FOXP3 Fix / Perm solution (BioLegend, 421403) at room temperature for 20 min, and then washed twice with 1X FOXP3 Perm buffer. Cells were then stained with APC-conjugated anti-human Ki-67 antibody (BioLegend, 350513; 1:200) for 30 min at room temperature in the dark. After two additional washes with 1X FOXP3 Perm buffer, cells were resuspended in FACS buffer (DPBS containing 2% FBS) and immediately analyzed using a BD LSRFortessa™ X-20 flow cytometer.
[0281]
[0282] Data statistical analysis
[0283] Quantitative data were presented as mean ± standard deviation (mean ± SD). Statistical significance of gene editing efficiency and gene expression comparison was assessed using a nonparametric one-tailed Student's t-test, and a p-value < 0.05 was considered statistically significant (*p < 0.05, **p < 0.01, ***p < 0.001, ****p < 0.0001).
[0284]
[0285] Example 2: De novo protein generation targeting multiple MLH1 binding sites using RFdiffusion.
[0286] Integration of a PE-synthesized DNA strand containing the desired genetic modification into the target genome can be hindered by the mismatch repair (MMR) pathway. During this process, MutS homologs, which recognize mismatches, recruit the MutLα complex (Nat Struct Mol Biol 20, 461-468; Proc Natl Acad Sci USA 118), ultimately reducing PE efficiency (Fig. 1a). To inhibit this pathway, we designed de novo small binders (SBs) that can bind to the dimeric interface of MLH1 and PMS2 using RF diffusion, thereby inhibiting the formation of the MutLα complex at the MLH1 C-terminus and PMS2 C-termini (Fig. 1b).
[0287] Although the single structures of MLH1 and PMS2 are readily available from the Protein Data Bank (PDB), their complex structure has not been clearly elucidated. Therefore, we used AlphaFold 3 to predict the interaction between the MLH1 C-terminal and PMS2 C-terminal domains and modeled the dimer binding site (Fig. 1c).
[0288] Existing conjugate design strategies have primarily focused on a single binding site, targeting the same site even when considering multiple binding hotspots for binding optimization (Nature 620, 1089-1100; Nature 605, 551-560; Cell 187, 4305-4317 e4318; Science 378, 49-56; Nat Commun 14, 2625). The present inventors have designed a novel conjugate capable of simultaneously inhibiting various activities by targeting multiple binding sites, each of which is as follows (Fig. 1c):
[0289] - The “Main binding (M)” region is the region that mediates the interaction between MLH1 and PMS2, and targets W538 and Q542 of MLH1.
[0290] - The “Critical activity (C)” region is the region where the nuclease activity of the MutLα complex exists, targeting Y750 of MLH1.
[0291] - The “Additional binding (A)” region is a non-competitive anchoring site, targeting M682 of MLH1.
[0292] Based on the three target sites described above, 1,978 MC binders that simultaneously bind to the "Main binding" and "Critical activity" regions were designed (Figs. 1D and 2A). Additionally, 821 MCA binders that simultaneously bind to all three sites (Figs. 1D and 2B) were designed as candidate MHL1-small binders (designated MHL1-SB).
[0293] Some candidates were excluded by applying three parameters from the RFdiffusion validation output: the predicted local distance difference test (pLDDT), the averaged predicted aligned error of interchain residue pairs (pAE), and the root mean square deviation (RMSD) (Nature 620, 1089-1100). pLDDT reflects the stability of the binding protein, pAE indicates binding affinity, and RMSD evaluates design reliability. Of the three parameters, pAE is known to be the most effective filter for binding protein design (Nat Commun 14, 2625).
[0294] However, the interaction pAE only reflects the overall characteristics of the interaction between MHL1-SB and the target and cannot directly calculate the effectiveness of inhibiting the interaction between MLH1 and PMS2. To address this limitation, we developed the AlphaFold 3 competition assay in three steps (Fig. 1e).
[0295] Specifically, (i) the MHL1-SB candidate sequence was added to the existing MLH1-PMS2 complex sequence to predict the results of the three complexes, and (ii) five representative residue pairs at the MLH1-PMS2 interface were selected to calculate the average atom-pair distance between Cα atoms (Fig. 2c). Then, (iii) using the histogram of the average atom-pair distance, a filtering threshold was set to exclude complexes not affected by the MHL1-SB candidate sequence (Figs. 1d, 2a to 2d).
[0296] AlphaFold 3 competitive results identified three distinct groups in the distance histogram:
[0297] - First peak region: Contains MHL1-SB candidates that did not significantly interfere with the formation of the MutLα complex.
[0298] - Second peak region: Contains MHL1-SB candidates that partially bind and slightly increase the distance between the MLH1-PMS2 interfaces.
[0299] - Wide distribution area: Includes MHL1-SB candidate that completely inhibits the interaction between MLH1 and PMS2.
[0300] The filtering threshold was set to 9.7724 Å, and all ineffective MHL1-SB candidates were removed. As a result, 20 MC binders with an interfacial distance greater than 9.7724 Å and 23 MCA binders were selected (Figures 2e to 2g and Table 1). Among the MCA binders, 17 MCA binders with an interfacial distance less than 9.7724 Å were classified as Failed-MCA (Table 1). Interestingly, no significant differences in structural or parameter-based characteristics were observed between the selected MCA binders and the Failed-MCA binders (Figures 3a to 3e). This suggests that the AlphaFold 3 competition assay designed by the inventors can utilize latent structural knowledge about protein complexes beyond simple structural characteristics. Additionally, sequence identity analysis and TM-scoring results showed that the selected MHL1-SB candidates exhibited structural and sequence diversity, indicating a wide range of potential binding modes and mechanisms (Fig. 3f).
[0301]
[0302] Example 3: Confirmation of the effect of MLH1-SB on improving PE2 prime editing efficiency.
[0303] To evaluate the effect of the MLH1-SB candidate selected in Example 1 on PE efficiency, MLH1-SB, PE2 (Nature 576, 149-157), and pegRNAs designed to induce substitutions (HEK4 G>T), insertions (HEK3 CTT insertion and RNF2 TAC insertion), and deletions (3-bp deletion of RNF2) were co-transfected into HeLa cells (Fig. 1f and Figs. 4a-4d). A 1.5-fold value was set as the threshold to confirm the effect of the MLH1-SB candidate on improving PE efficiency (Fig. 1f). The MCA conjugates generally showed slightly higher Fold change values than the MC conjugates and significantly higher than the Failed-MCA conjugates. However, among the Failed-MCA conjugates, Failed-MCA2, Failed-MCA4, and Failed-MCA7 showed excellent PE efficiency enhancement, contrary to expectations (Fig. 4e).
[0304] These results suggest that the AlphaFold 3 competition assay designed by the inventors can effectively distinguish binders that can interfere with native complex formation. Because existing dimeric complex prediction methods lack information on competitive binding with native binding partners, our ternary complex prediction is more advantageous in identifying MLH1-SB candidates with higher efficacy. Furthermore, no significant correlation was observed between PE efficiency fold changes and the AlphaFold 3 competition test scores (Figure 4f), suggesting that simply passing or failing the test is more predictive than the numerical score.
[0305]
[0306] Example 4: Checking compatibility with other prime editing systems
[0307] Recent studies have reported that protecting the 3'-end of pegRNA with La protein improves PE efficiency, and based on this, the PE7 architecture was developed (Nature 628, 639-647). To confirm the simultaneous application of MMR inhibition by MLH1-SB of the present invention and RNA protection by La protein, MLH1dn or MLH1-SB was administered to HeLa cells together with PEmax (an optimized version of PE2, Cell 184, 5635-5652 e5629) and PE7 (Nature 628, 639-647) (Fig. 5a). By inhibiting MMR, overexpression of MLH1dn or MLH1-SB significantly increased the PE efficiency of insertions (HEK3 CTT and RNF2 TAC), deletions (3-bp deletion in HEK3 and RNF2), and substitutions (HEK3 T->A, HEK4 G->T, RNF2 C->G) using PEmax and PE7, regardless of the presence of ngRNA (Figs. 5b, 5c, 6a, and 6b). Notably, MLH1-SB showed a higher fold change in PE efficiency compared to MLH1dn (Fig. 5d), and the levels of unwanted mutations remained similar between the MLH1dn and MLH1-SB treatment groups (Fig. 6c). Specifically, for PEmax, the average fold increase in PE efficiency was 7.1-fold for MLH1dn and 18.9-fold for MLH1-SB, indicating that MLH1-SB exhibits 2.7-fold higher activity than MLH1dn. Similarly, for PE7, the PE efficiency increased 2.7-fold for MLH1dn and 9.4-fold for MLH1-SB, indicating that MLH1-SB exhibits 3.5-fold higher activity than MLH1dn. These results indicate that MMR inhibition is compatible with pegRNA protection, and that co-expression of MLH1-SB and PE7 exhibits excellent synergistic effects. Furthermore, in the presence of ngRNA, MLH1-SB overexpression increased PE efficiency by 2.8-fold for PEmax and 5.9-fold for PE7.
[0308] Based on PE2, various PE6 architectures, such as PE6b, PE6c, and PE6d, with improved PE efficiency have been developed through phage-assisted evolution of reverse transcriptase (RT) (Cell 186, 3983-4002 e3926). To evaluate whether the PE6d architecture can reconcile increased PE efficiency through pegRNA protection and MMR inhibition, we constructed PE6d fused to the La protein (PE6d-La) (Fig. 5e). Similar to the previous experiment, the PE efficiency was measured in seven different cases in the presence or absence of MLH1-SB overexpression in HeLa cells (Fig. 5f).
[0309] As a result, the PE efficiency significantly increased by 5.4-fold and 11.6-fold, respectively, compared to PE6d alone under both the conditions of PE6d-La alone or PE6d plus MLH1-SB overexpression (Fig. 5g). Furthermore, when PE6d-La and MLH1-SB were co-expressed, the PE efficiency increased by 24.6-fold compared to PE6d alone (Fig. 5g), indicating that pegRNA protection and MMR inhibition exhibited a remarkable synergistic effect on PE efficiency. In particular, PE6d-La increased the frequency of unwanted mutations compared to PE6d, but this was stably maintained under the condition of MLH1-SB overexpression (Fig. 6d).
[0310]
[0311] Example 4: Confirmation of the mechanism of action of MLH1-SB
[0312] To investigate the molecular mechanism by which MLH1-SB enhances intracellular PE efficiency, we constructed a plasmid expressing HA-tagged MLH1-SB (MLH1-SB-HA) and another plasmid expressing wild-type MLH1 tagged with 3×FLAG (FLAG-MLH1). HeLa cells were then transfected with FLAG-MLH1, MLH1-SB-HA, or both plasmids, respectively, and co-immunoprecipitation assays were performed using anti-HA beads, anti-FLAG beads, anti-HA antibodies, and anti-FLAG antibodies (Figs. 7A and 8A). When MLH1-SB-HA was transfected alone, no FLAG signal was detected in the co-sedimentation assay using anti-HA beads. However, when MLH1-SB-HA and FLAG-MLH1 were co-transfected, a FLAG signal was observed in the anti-HA bead pull-down. Similarly, in the co-suppressive assay using anti-FLAG beads, the HA-tagged protein was detected only when MLH1-SB-HA and FLAG-MLH1 were co-expressed. These results confirm that MLH1-SB directly binds to MLH1 in cells. To further verify this binding, liquid chromatography-tandem mass spectrometry (LC-MS / MS) was performed (Fig. 8B). HeLa cells were transfected with the control vector or MLH1-SB-HA, and then co-suppressive assay was performed using anti-HA beads. LC-MS / MS analysis revealed that MLH1 was detected only in cells transfected with MLH1-SB-HA (Fig. 8C-8E), supporting the direct binding of MLH1-SB to MLH1.
[0313] Next, we used immunofluorescence (IF) staining to confirm the subcellular localization of MLH1-SB in relation to MLH1 (Fig. 7b). HEK293T cells, which have low MLH1 expression due to hypermethylation of the MLH1 promoter (Fig. 8f), were transfected with an empty vector (“Mock”), MLH1-SB-HA, FLAG-MLH1, or both plasmids. Cells were then stained with Hoechst (blue), anti-FLAG antibody (magenta), and anti-HA antibody (green). FLAG staining was observed primarily in the nucleus, consistent with the prediction of a nuclear localization of MLH1. In contrast, HA staining was detected primarily in the cytoplasm, suggesting that MLH1-SB is primarily located in the cytoplasm. Interestingly, when both MLH1 and MLH1-SB were expressed, a strong correlation was observed, with green and magenta fluorescent signals overlapping within the nucleus (Fig. 8g). Considering that MLH1 contains a nuclear localization signal (NLS), whereas MLH1-SB does not, these results imply that MLH1-SB binds to MLH1 and cotransports them to the nucleus. To determine whether the enhancement of PE efficiency by MLH1-SB is MMR-dependent, we transfected PEmax and PE7 together with MLH1dn or MLH1-SB into MLH1-present HEK293 cells and MLH1-deficient HEK293T cells (Mol Cancer 13, 11) (Figs. 7c–7f, 8h, and 8i). In HEK293 cells, overexpression of MLH1dn or MLH1-SB significantly increased PE efficiency at both PEmax and PE7 (Fig. 7e). However, in HEK293T cells lacking MLH1, expression of MLH1dn or MLH1-SB did not alter PE efficiency (Fig. 8f). These results imply that the enhancement of PE efficiency mediated by MLH1-SB is MMR (miss match repair) dependent.
[0314]
[0315] Example 5: Confirmation of PE efficiency of MLH1-SB according to editing size.
[0316] Considering that co-expression of MLH1-SB significantly enhanced PE efficiency, we introduced MLH1-SB into the PE construct to develop improved prime editors. Using the 2A peptide, which promotes protein self-cleavage during translation, we constructed constructs fused to PE7: MLH1dn; MLH1-SB; MLH1-SB with a nuclear localization signal (NLS) (designated MLH1-SB-NLS); and a construct containing both MLH1-SB-NLS and MLH1-SB (designated MLH1-SB2) (Fig. 9a). The resulting prime editors were designated PE7-SB (PE7 linked to MLH1-SB via the 2A system), PE7-SB-NLS (PE7 linked to MLH1-SB-NLS), and PE7-SB2 (PE7 linked to MLH1-SB and MLH1-SB-NLS), respectively. In addition, MLH1-SB was fused to the N-terminus or C-terminus of PE7 using an XTEN link (Xt), resulting in PE7-N-SB and PE7-C-SB, respectively. The prime editors were introduced into HeLa cells for PE targeting of three sequences: a G to T mutation in HEK4, a 3-bp deletion in HEK3, and a 3-bp deletion in RNF2. Among the tested prime editors, PE7-SB2 showed a similar level of PE efficiency compared to PE7-SB and PE7-SB-NLS (Figs. 9b and 10a). In contrast, PE7 linked to MLH1dn (PE7-MLH1dn) did not show a significant improvement in PE efficiency compared to PE7. This suggests that the small size of MLH1-SB (82 amino acids) can be efficiently incorporated into the PE structure.
[0317] Next, we examined whether PE7-SB2 could synergize with ngRNA addition. PEmax, PE7, and PE7-SB2 were transfected in the presence or absence of ngRNA to target various mutations in HeLa cells. The targeted mutations included substitutions (HEK3 T→A, T→G, T→C; HEK4 G→A, G→T, G→C; RNF2 C→A, C→T, C→G, A→T, A→G, A→C), insertions of various sizes (TAC, 38 bp insertion in RNF2, 24 bp Flag insertion in HEK3), and deletions of various sizes (3 bp deletion in HEK3, 3 bp, 5 bp, 9 bp, 11 bp, and 45 bp deletions in RNF2) (Figures 9C to 9E and 10B to 10D). As a result, we confirmed that PE7-SB2 improved PE efficiency for all substitutions and deletions or insertions less than 12 bp (Fig. 9f). However, for larger edit sizes, such as a 45 bp deletion in RNF2, a 38 bp insertion, and a 24 bp Flag insertion in HEK3, no improvement in efficiency was observed regardless of the inclusion of ngRNA (Fig. 9g). These results are consistent with previous studies showing that MMR mainly affects mismatches less than 13 bp (Nat Rev Mol Cell Biol 7, 335-346; Biophys J 106, 2483-2492) and that PE edits greater than 15 bp are not associated with MMR (Mol Ther Nucleic Acids 32, 914-922). Overall, for PE edits less than 12 bp, PE7-SB2 exhibited an 18.8-fold improved PE efficiency compared to PEmax and a 2.5-fold improvement compared to PE7 (Fig. 9f). Furthermore, when ngRNA was added to PE7-SB2, it showed a 2.8-fold improvement in PE efficiency compared to PE7-SB2 alone (Fig. 9f), confirming that PE7-SB2 exhibits a strong synergistic effect with ngRNA.
[0318] PE7-SB2 was further applied to various cell lines to determine its PE efficiency. In human induced pluripotent stem cells (hiPSCs), the PE efficiency was evaluated for substitutions (HEK3 T→A and HEK4 G→T), insertions (HEK3 3-bp insertion and RNF2 3-bp insertion), and deletions (HEK3 3-bp deletion and RNF2 3-bp deletion) (Figs. 10e to 10g). Similar to the results in HeLa cells, PE7-SB2 showed a 6.9-fold and 1.9-fold improved PE efficiency compared to PEmax and PE7 in hiPSCs (Fig. 9h). Furthermore, when ngRNA was added, the PE efficiency was further improved by 2.1-fold compared to PE7-SB2 alone.
[0319] In addition, the PE efficiencies of PEmax, PE7, and PE7-SB2 according to the presence or absence of ngRNA were investigated in HEK293 cells targeting substitutions (HEK3 T→A, HEK4 G→T, RNF2 C→G), insertions (HEK3 3-bp insertion and RNF2 3-bp insertion), and deletions (HEK3 3-bp deletion and RNF2 3-bp deletion) (Figs. 10h to 10j). As a result, PE7-SB2 showed a PE efficiency that was 3.0-fold and 2.5-fold improved compared to PEmax and PE7, respectively (Fig. 9i).
[0320] Additionally, we verified whether PE7-SB2 could be used in engineered RT variants. To this end, the RT of PE7-SB2 was replaced with RT of PE6b (RT6b), RT of PE6c (RT6c), RT of PE6d (RT6d), and RT with improved dNTP affinity (RT**) (Fig. 11a) (Nat Biotechnol. 10.1038 / s41587-024-02405-x). Consistently, PE7-SB2 maintained significantly higher efficiency than PEmax when combined with all tested engineered reverse transcriptase variants (Figs. 11b and 11c). Notably, PE7-SB2 containing RT6d exhibited a higher average PE efficiency than the original PE7-SB2 (Fig. 11d).
[0321]
[0322] Example 6: Confirmation of off-target editing by prime editor using MLH1-SB
[0323] Next, we investigated the editing efficiency at clearly defined off-target sites. For HEK3 and HEK4, we used three off-target sites predicted by CIRCLE-seq (Cell 184, 5635-5652 e5629; Nat Methods 14, 607-614), and for RNF2, we used five off-target sites predicted by Cas-OFFinder (Bioinformatics 30, 1473-1475) (Fig. 12a). HeLa cells were transfected with PEmax, PE7, or PE7-SB2 together with pegRNA targeting HEK3, HEK4, and RNF2, respectively. Each target site was amplified and subjected to targeted deep sequencing. As a result, the average off-target editing efficiency of PE7-SB2 was similar to that of PE7, indicating that MLH1-SB did not induce additional significant off-target editing (Figures 12b to 12e).
[0324] Furthermore, we investigated the cytotoxicity and proliferation effects of MLH1-SB and PE7-SB2 on human cells. To assess cell proliferation and apoptosis, HeLa cells were transfected with plasmids expressing the puromycin resistance gene (PuroR), MLH1dn, MLH1-SB, PE7-MLH1dn, and PE7-SB2. Cell proliferation was measured by Ki-67 staining, and cell death was analyzed by annexin V / propidium iodide staining (Figs. 13a and 13b). No significant changes in cell proliferation were observed in cells transfected with MLH1-SB and PE7-SB2 (Fig. 14a). Although MLH1-SB and PE7-SB2-transfected cells showed a tendency toward increased early apoptosis and apoptotic cell population, the levels were similar to or lower than those in PuroR-transfected cells, indicating that the observed cell death was primarily due to damage caused by plasmid transfection (Fig. 14b). Additionally, the effect of PE7-SB2 was evaluated in patient-derived fibroblasts (Figs. 13c and 13d). Similar to the results in HeLa cells, there was no difference in cell proliferation or apoptosis levels between fibroblasts transfected with PE7 and PE7-SB2 (Figs. 13e and 13f). Furthermore, PE7-SB2 increased PE efficiency compared to PE7 in both fibroblasts and cell cycle-arrested HeLa cells (Figs. 13g to 13i). Next, the effects of MLH1-SB and PE7-SB2 on the transcriptome were evaluated. After transfection of HeLa cells with plasmids expressing PuroR, MLH1dn, MLH1-SB, PE7, and PE7-SB2, transcriptome analysis was performed 48 h later (Figs. 14c and 14d). PuroR-transfected HeLa cells were used as a control for DNA plasmid transfection and expression.Interestingly, HeLa cells transfected with MLH1dn, MLH1-SB, and PE7-SB2 formed similar clusters in principal component analysis (PCA), suggesting that they shared MMR inhibitory effects (Fig. 14c). Volcano plot and Gene Ontology (GO) analysis identified shared differentially expressed genes (DEGs) and statistically significant GO terms between MLH1dn and MLH1-SB transfected cells (Fig. 14d and Table 6). In addition, shared DEGs and significant GO terms were observed between cells transfected with PE7 and PE7-SB2, but no statistically significant GO terms were observed in the DEGs of PE7-SB2 alone (Fig. 14d and Table 6).
[0325]
[0326]
[0327]
[0328] These data suggest that PE7-SB2 can enhance PE efficiency without significantly affecting the cellular transcriptome. To investigate the effects of MLH1-SB and PE7-SB2 on the genome-wide level, PuroR, MLH1dn, MLH1-SB, PE7, and PE7-SB2 plasmids were transfected into single-clone HeLa cells and hiPSCs. Ten days after transfection, genomic DNA (gDNA) was collected, and whole genome sequencing (WGS) with 30x coverage was performed. Based on the WGS data from single-clone cells, base substitutions, insertions, and deletions across the genome were analyzed (Figures 13j and 13k). The analysis results showed that the levels of base substitutions, insertions, and deletions across the genome were similar between PuroR-transfected cells and cells transfected with other plasmids in both hiPSCs and HeLa cells.
[0329]
[0330] Example 7: Confirmation of the PE efficiency-enhancing effect of PE7-SB2 in an in vivo mouse model.
[0331] MLH1-SB was designed based on the human MutLα complex, but because human MLH1 shares a high homology of approximately 93.2% with mouse MLH1, we expected that MLH1-SB could inhibit mouse MLH1. To evaluate the possibility of DNA plasmid delivery in vivo, we administered 50 μg of luciferase expression vector via hydrodynamic injection to 6-week-old ICR (Institute of Cancer Research) mice. Bioluminescence imaging was performed using an in vivo imaging system (IVIS) on days 5, 7, and 10 after injection (Fig. 15a). As a result, a marked photon flux was observed in the livers of luciferase-injected mice (Luc-1 and Luc-2) on day 5, which decreased on day 10 (Figs. 15b and 15c). Subsequently, the effect of MLH1-SB on in vivo PE efficiency in the mouse Igf2 gene was evaluated by co-administration of PE7 and PE7-SB2 plasmids with a pegRNA plasmid inducing TA insertion and G to T substitution (in two experiments including injection of a total of 50 μg and 60 μg of plasmid). Serum and liver-derived gDNA were collected 10 days after injection (Fig. 15d). To evaluate the in vivo toxicity associated with PE7 and PE7-SB2, serum concentrations of aspartate aminotransferase (AST) and alanine aminotransferase (ALT) were measured in mice injected with a total of 60 μg of plasmid (Fig. 15e). Untreated mice, mice injected with PBS only, and mice injected with PE7-SB2 only without pegRNA (PE7-SB2 only) were used as controls.AST and ALT levels increased in the PE7-SB2 alone, PE7 + pegRNA, and PE7-SB2 + pegRNA injection groups, but were similar between the PE7 + pegRNA and PE7-SB2 + pegRNA injection groups. This suggests that additional expression of MLH1-SB induces only limited toxicity in vivo. Furthermore, PE7-SB2 increased PE efficiency by 2.3-fold and 3.4-fold compared to PE7 in 50 μg and 60 μg plasmid injection experiments, respectively (Fig. 15f). Notably, the incidence of unwanted mutations was similar to baseline in both PE7 and PE7-SB2. Finally, we investigated pegRNA-dependent off-target editing of the seven known off-target sites of Igf2 pegRNA 41 in the 60 μg injection sample. Across these seven off-target sites, no significant changes were observed compared to baseline levels (Figures 15g and 15h).
[0332]
[0333] While specific aspects of the present invention have been described in detail above, it will be apparent to those skilled in the art that these specific descriptions merely represent preferred embodiments and are not intended to limit the scope of the present invention. Therefore, the substantial scope of the present invention is defined by the appended claims and their equivalents.
[0334]
[0335] The polypeptide of the present invention significantly enhances prime editing efficiency by binding to the MLH1 protein and inhibiting the formation of the MutLα complex with PMS2, while exhibiting significantly lower off-target editing and cytotoxicity. Furthermore, the polypeptide of the present invention has a significantly smaller size than the existing MLH1dn (dominant negative MLH1), and thus can be easily integrated into various existing prime editing systems within a single vector, enabling universal use. Therefore, the novel MLH1-binding polypeptide of the present invention and the nucleic acid encoding it can be usefully applied in the fields of MMR deficiency and gene editing.
[0336]
[0337] Electronic file attached.
Claims
1. An MLH1 binding polypeptide that binds to one or more amino acid residues of MLH1, wherein the one or more amino acid residues of MLH1 are Amino acid residues in the region involved in the interaction between MLH1 and PMS2; Amino acid residues in the nuclease active domain of the MutLα complex; and An MLH1 binding polypeptide selected from the group consisting of one or more amino acid residues selected from the group consisting of amino acid residues of the non-competitive anchoring site of MLH1.
2. An MLH1 binding polypeptide according to claim 1, characterized in that the MLH1 comprises an amino acid sequence of sequence number 1.
3. An MLH1 binding polypeptide characterized in that it binds to at least one amino acid residue selected from the group consisting of W538, Q542, M682, and Y750 of sequence number 1 in the second paragraph.
4. An MLH1 binding polypeptide characterized in that it binds to residues W538, Q542, and Y750 of sequence number 1 in the third paragraph.
5. An MLH1 binding polypeptide characterized in that it binds to residues W538, Q542, M682, and Y750 of sequence number 1 in the third paragraph.
6. An MLH1 binding polypeptide comprising a sequence selected from the group consisting of SEQ ID NO: 2 to SEQ ID NO: 61 in the first paragraph.
7. An MLH1 binding polypeptide comprising a sequence selected from the group consisting of SEQ ID NOs: 2, 5, 6, 8, 9, 13, 15, 16, 18, 20, 21, 22, 23, 24, 25, 27, 28, 30, 31, 32, 33, 37, 38, 40, 41, 42, 43, 44, 46, 48, and 51 in the first paragraph.
8. A nucleic acid encoding the MLH1 binding polypeptide of paragraph 1.
9. A nucleic acid comprising a sequence selected from the group consisting of SEQ ID NO: 62 to SEQ ID NO: 121 in the 8th paragraph.
10. In paragraph 8, a nucleic acid comprising a sequence selected from the group consisting of sequence numbers 62, 65, 66, 68, 69, 73, 75, 76, 78, 80, 81, 82, 83, 84, 85, 87, 88, 90, 91, 92, 93, 97, 98, 100, 101, 102, 103, 104, 106, 108 and 111, 11. A recombinant vector containing the nucleic acid of Article 8.
12. A recombinant vector according to claim 11, further comprising a nucleic acid encoding a Cas nickase protein and a nucleic acid encoding a reverse transcriptase (RT).
13. A recombinant vector comprising the nucleic acid of paragraph 8 and a nucleic acid encoding a prime editor.
14. A recombinant vector according to claim 13, characterized in that the prime editor is selected from the group consisting of PE1, PE2, PE3, PE3b, PEmax, PE4, PE5, PE6, PE6a, PE6b, PE6c, PE6d, PE6e, PE6f, PE7, and TwinPE.
15. A recombinant vector according to claim 13, further comprising a sequence encoding at least one selected from the group consisting of a single-stranded DNA-binding domain (ssDBD), a nuclear localization sequence (NLS), ngRNA (nick gRNA), MLH1 binder, RNA-binding exonuclease protection factor, and pegRNA (prime editing guide RNA).
16. In the 13th paragraph, the recombinant vector characterized in that the prime editor and MLH1 binding polypeptide are cloned into the recombinant vector with the following structure: [Prime Editor]-L1-[MLH1 binding polypeptide]; [Prime Editor]-L1-[MLH1 binding polypeptide]-[NLS]; [Prime Editor]-L1-[MLH1 binding polypeptide]-L2-[MLH1 binding polypeptide]; or [Prime Editor]-L1-[MLH1 binding polypeptide]-[NLS]-L2-[MLH1 binding polypeptide], Here, the above L1 and L2 are direct connections or linkers.
17. A composition for gene editing comprising at least one of the MLH1 binding polypeptide of claim 1, the nucleic acid of claim 8, and the recombinant vector of any one of claims 11 to 16.
18. A composition for gene editing, further comprising a prime editor protein or a nucleic acid encoding the same; and pegRNA, in claim 17.
19. A composition for gene editing further comprising pegRNA in claim 17.
20. A gene editing method comprising a step of treating or administering the composition of Article 17 to a cell or a subject.
Citation Information
Patent Citations
Method for improving base editing efficiency of guided editing system
CN116042573A
Prime editor variants, constructs, and methods for enhancing prime editing efficiency and precision
WO2022150790A2
Methods and compositions for inhibiting mismatch repair
WO2023096847A2
Improved prime editing methods and compositions
WO2023205687A1
Methods and compositions for modulating cellular factors to increase prime editing efficiencies
WO2024138087A2