A micropeptide for promoting HDR-mediated precise editing of a gene and application thereof

By using artificial intelligence-designed micropeptides to inhibit the NHEJ repair pathway and promote HDR repair, the problem of low HDR repair efficiency in existing technologies is solved, thereby improving the precision and safety of gene editing.

CN122103278APending Publication Date: 2026-05-29CHANGZHI MEDICAL COLLEGE

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHANGZHI MEDICAL COLLEGE
Filing Date
2026-02-09
Publication Date
2026-05-29

Smart Images

  • Figure CN122103278A_ABST
    Figure CN122103278A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of gene editing, and particularly relates to a micropeptide for promoting HDR-mediated precise gene editing and application thereof. The amino acid sequence of the micropeptide is obtained through artificial intelligence design and screening, and is completed through a CRISPR / Cas9-mediated HDR precise gene editing report system for functional verification. The micropeptide is specifically combined with Ku70 protein in cells, and then the interaction between Ku70 / Ku80 heterodimer and DNA-PKcs protein is inhibited, so that the directional regulation of the DNA damage repair pathway in cells is realized. The micropeptide enhances the efficiency of the CRISPR / Cas9 system-mediated HDR, and effectively improves the success rate of precise gene editing at the cell level. In addition, the micropeptide can be adapted to various gene editors depending on the DNA strand break repair mechanism, and is compatible with various cell types, has a wide range of applications, and has important industrial application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of gene editing technology, specifically relating to a micropeptide that promotes HDR-mediated precise gene editing and its applications. Background Technology

[0002] CRISPR / Cas9 technology, a transformative tool in the life sciences, can precisely cut specific sequences in an organism's genome to generate DSBs. Its core application value relies on the cell's repair mechanisms for DSBs: while the non-homologous end joining (NHEJ) pathway repairs rapidly, it easily introduces random insertion / deletion mutations (indels); whereas the HDR pathway can utilize donor DNA templates to achieve precise modification of target sites, making it a key approach for achieving precise gene editing in gene therapy, precision crop breeding, and basic medical research. Therefore, how to effectively improve HDR repair efficiency has become a core technical problem urgently needing to be solved in the current gene editing field.

[0003] When Cas9 cuts the target DNA to generate DSBs, the NHEJ repair pathway in the cell becomes absolutely dominant, with more than 90% of DSBs being repaired by NHEJ. This results in unpredictable indels at the target site, making it impossible to achieve the intended precise gene modification. Meanwhile, the HDR repair efficiency is usually less than 10%. This situation seriously restricts the commercial application and technological breakthroughs of precision gene editing technology in fields such as genetic disease correction, tumor targeted therapy, and crop trait targeted improvement.

[0004] To improve HDR efficiency, various technical strategies have been developed in this field, but all have significant drawbacks: First, using small molecule inhibitors (such as DNA-PKcs inhibitors AZD7648 and NU7441, and Lig4 inhibitor SCR7) to block the NHEJ pathway can improve HDR efficiency to some extent, but these inhibitors are prone to causing overall genome instability, leading to large-scale mutations or deletions at non-target sites, and the global inhibitory effect can exacerbate off-target effects; Second, using cell cycle inhibitors to arrest cells in the HDR-active S / G2 phase can improve HDR efficiency, but it can interfere with the normal cell division cycle and significantly increase the apoptosis rate; Third, modifying and optimizing CRISPR / Cas9 editing tools, delivery systems, or donor templates can improve HDR efficiency, but there are problems such as complex processes, high R&D costs, and limited universality of different modification schemes, making it difficult to meet the needs of a wide range of application scenarios.

[0005] In recent years, the application of artificial intelligence (AI) technology in gene editing has gradually expanded, providing new ideas for the design of novel regulatory molecules. Leveraging AI algorithms such as machine learning and deep learning, micropeptides with specific functions can be efficiently designed based on data such as protein structure prediction and molecular interaction simulation. These AI-generated micropeptides have the advantages of small molecular weight and high specificity, enabling them to accurately identify and bind to functional proteins in cells. By regulating protein activity, interactions, or subcellular localization, they can directionally intervene in cellular biological functions. Compared to traditional small molecule drugs, they exhibit lower cytotoxicity and better targeting, providing a potential direction for solving efficiency regulation problems in gene editing. However, currently, there are no relevant technical solutions for using AI to design micropeptides to specifically promote HDR repair efficiency. The field still urgently needs a novel technical solution that can efficiently improve HDR repair efficiency while also possessing advantages such as high safety, strong versatility, and low cytotoxicity. Summary of the Invention

[0006] Based on this, the present invention provides a micropeptide that promotes HDR-mediated precise gene editing and its application. This micropeptide, obtained through AI-assisted design and screening, specifically binds to the key NHEJ pathway protein Ku70. By inhibiting the interaction between the Ku70 / Ku80 heterodimer and DNA-PKcs protein, it blocks the initiation of the NHEJ repair pathway, causing cell damage repair to favor the HDR pathway, thereby significantly improving the efficiency of HDR-mediated precise editing in gene editing.

[0007] To achieve the above objectives, the present invention can adopt the following technical solutions: This invention provides a micropeptide that facilitates HDR-mediated precise gene editing, wherein the amino acid sequence of the micropeptide is selected from any of the following: (1) The amino acid sequence as shown in SEQ ID NO.1; (2) The amino acid sequence shown in SEQ ID NO.1 is obtained by substitution, insertion or deletion of one or more amino acids, and still has the function of promoting HDR-mediated precise gene editing.

[0008] Another aspect of the present invention provides a reagent for improving the efficiency of HDR-mediated precise gene editing, which includes the aforementioned micropeptides.

[0009] Preferably, the reagents described above further include one or more of the following: a CRISPR / Cas9 expression vector, a donor plasmid vector, or an HDR reporter vector.

[0010] In another aspect, the present invention provides a kit for improving the efficiency of HDR-mediated precise gene editing, the kit comprising the reagents described above.

[0011] In another aspect, the present invention provides a method for improving the efficiency of HDR-mediated precise gene editing, the method comprising: co-transfecting the above-mentioned micropeptide with other vectors into target cells to achieve precise gene editing through the HDR pathway; the other vectors include CRISPR / Cas9 expression vectors, donor plasmid vectors and / or HDR reporter vectors.

[0012] In another aspect, the present invention provides the application of the above-mentioned micropeptide in improving the efficiency of HDR-mediated precise gene editing.

[0013] In another aspect, the present invention provides the application of the above-mentioned micropeptide in the preparation of gene editing drugs, wherein the micropeptide inhibits the NHEJ repair pathway and promotes HDR-mediated gene editing.

[0014] In another aspect, the present invention provides a gene-editing drug comprising the aforementioned micropeptide or the aforementioned reagent.

[0015] The beneficial effects of this invention include at least the following: The amino acid sequence of the micropeptide provided by this invention is generated and screened using artificial intelligence, and its function is verified through a CRISPR / Cas9-mediated HDR gene editing reporter system. The micropeptide specifically binds to the Ku70 protein within cells, thereby inhibiting the interaction between the Ku70 / Ku80 heterodimer and DNA-PKcs protein, achieving targeted regulation of intracellular DNA damage repair pathways. Co-transfecting the expression vector of the micropeptide with the cell genome targeting vector (specifically, a CRISPR / Cas expression vector) and the donor vector into target cells can effectively improve the efficiency of precise gene editing at the cellular level. Furthermore, this micropeptide is compatible with various gene editors that rely on DNA strand break repair mechanisms and is compatible with multiple cell types, exhibiting a wide range of applications and significant potential for industrial application. Attached Figure Description

[0016] Figure 1 A schematic diagram of the DNA double-strand break (DSB) repair mechanisms (NHEJ and HDR) mediated by the CRISPR / Cas9 system; Figure 2 A schematic diagram illustrating the design and screening of small peptides based on artificial intelligence; Figure 3 A schematic diagram showing the interaction interface and hotspot residues between Ku70 / 80 heterodimer and DNA-PKcs protein; Figure 4A schematic diagram of the interaction complex between the AlphaFold3-predicted micropeptide (top 25 in IPTM score) (purple) and the Ku70 / 80 heterodimer (green) is shown for PyMOL. Figure 5 To detect the linearization results of the pCAGGS vector using agarose gel electrophoresis; Figure 6 Results of double enzyme digestion identification of pCAG-Puro-P2A vector (a); Schematic diagram of pCAG-Puro-P2A vector (b); Figure 7 The linearization results of the pCAG-Puro-P2A vector were detected by agarose gel electrophoresis. Figure 8 Results of double enzyme digestion identification of pCAG-Puro-P2A-KU-SBs vector (a1,2,3); Schematic diagram of pCAG-Puro-P2A-KU-SB7 vector (b). Figure 9 To detect the linearization results of the pX330-U6-Chimeric_BB-CBh-hSpCas9 vector by agarose gel electrophoresis; Figure 10 Results of Sanger sequencing identification of the CRISPR / Cas9-sgEMX1 vector (a); Schematic diagram of the CRISPR / Cas9-sgEMX1 vector (b); Figure 11 Results of double enzyme digestion identification of pHDR-PMG vector (a) and pcDNA3.1 vector (b); Figure 12 Results of double enzyme digestion identification of HDR reporter vector (a); Schematic diagram of HDR reporter vector (b); Figure 13 Results of nT-EMX1-mCherry-nT-EMX1 sequence detection by agarose gel electrophoresis; Figure 14 A schematic diagram of the EMX1-HDR reporter vector (a); results of double enzyme digestion identification of the EMX1-HDRreporter vector (b). Figure 15 To observe the working status of HDR reporter vectors in cells after transfection with different plasmid vectors using fluorescence microscopy; Figure 16A gating strategy for detecting the efficiency of different micropeptides in HDR reporter vector repair by flow cytometry (a); the efficiency of HDR reporter vector repair by flow cytometry in the control group (pCAG-Puro-P2A) and the experimental group (pCAG-Puro-P2A-KU-SB7) (b); and statistical analysis of the efficiency of different micropeptides in HDR reporter vector repair by flow cytometry (c). Figure 17 This is a schematic diagram of the precise genome editing and detection process of the donor plasmid (pd-EMX1) based on HDR repair at the EMX1 gene locus. Figure 18 To detect the efficiency of KU-SB7 micropeptide in HeLa cells for precise gene editing at the EMX1 gene locus under conditions where no drug screening was used; Figure 19 To determine the efficiency of precise gene editing at the EMX1 gene locus of KU-SB7 micropeptide in HeLa cells using an enzyme digestion method under drug-screened cell conditions; Figure 20 To assess the efficiency of the micropeptide (KU-SB7) in precisely editing the EMX1 gene locus in HeLa cells after drug screening using DeepSeq; Figure 21 The flow cytometry detection of the micropeptide KU-SB7 at different gene loci (AAVS1, CCR5, NUDT5) in HeLa and U20S cells can improve the efficiency of the HDR reporter vector. Figure 22 To detect the expression of KU-SB7 micropeptide and Ku70 protein in cells using immunofluorescence (IF) staining; Figure 23 To detect the effect of KU-SB7 micropeptide on cell proliferation during CRISPR / Cas9-mediated gene editing using CCK8 assay; Figure 24 The effect of KU-SB7 micropeptide on apoptosis during CRISPR / Cas9-mediated gene editing was detected using Annexin V-FITC / PI double staining. Detailed Implementation

[0017] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The embodiments provided are for better illustration of the present invention, but are not intended to limit the scope of the invention to the embodiments described. Therefore, non-essential improvements and adjustments made to the embodiments by those skilled in the art based on the above description are still within the scope of protection of the present invention.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. Singular expressions include plural expressions unless they have a distinct meaning in the context. As used herein, it should be understood that terms such as “comprising,” “having,” “including,” are intended to indicate the presence of features, numbers, operations, components, parts, elements, materials, or combinations thereof. The terminology of the invention is disclosed in the specification and is not intended to exclude the possibility that one or more other features, numbers, operations, components, parts, elements, materials, or combinations thereof may be present or added. As used herein, “ / ” may be interpreted as “and” or “or,” depending on the context.

[0019] In a first aspect, embodiments of the present invention provide a micropeptide that promotes HDR-mediated precise gene editing, wherein the amino acid sequence of the micropeptide is selected from any of the following: (1) The amino acid sequence as shown in SEQ ID NO.1; (2) The amino acid sequence shown in SEQ ID NO.1 is obtained by substitution, insertion or deletion of one or more amino acids, and still has the function of promoting HDR-mediated precise gene editing.

[0020] It should be noted that this invention begins with precise structural biology analysis, utilizes generative AI tools (RFdiffusion and ProteinMPNN) for innovative molecular design, and employs a unique computational competitive simulation combining multiple AI prediction tools (AlphaFold2 and AlphaFold3) for efficient screening. This provides a micropeptide (as shown in SEQ ID NO.1 to SEQ ID NO.25, with the amino acid sequence shown in SEQ ID NO.1 (hereinafter also referred to as KU-SB7) showing the best effect) that can specifically bind to the Ku70 protein and block the interaction between the Ku70 / Ku80 heterodimer and DNA-PKcs protein. This achieves the goal of promoting HDR repair efficiency by inhibiting intracellular NHEJ repair (see [link to invention]). Figure 1 ), and a system design, screening and verification process (see Figure 2This enables the systematic, efficient, and rapid acquisition of powerful micropeptides, providing a solid technical foundation for improving the efficiency of HDR-mediated precision gene editing.

[0021] Secondly, embodiments of the present invention provide a reagent for improving the efficiency of HDR-mediated precise gene editing, which includes the aforementioned micropeptide.

[0022] It should be noted that the micropeptides in this invention can be prepared into reagents to improve the efficiency of HDR-mediated precise gene editing for use in gene editing.

[0023] In some specific examples, the reagents described above further include: a CRISPR / Cas9 expression vector, a donor plasmid vector, and / or an HDR reporter vector. It should be noted that the reagents in this invention for improving the efficiency of HDR-mediated gene editing precision also include a CRISPR / Cas9 expression vector, a donor plasmid vector, and / or an HDR reporter vector, wherein the CRISPR / Cas9 expression vector, the HDR reporter vector, and the donor plasmid vector are well known in the art.

[0024] Thirdly, embodiments of the present invention provide a kit for improving the efficiency of HDR-mediated precise gene editing, the kit comprising the reagents described above.

[0025] It should be noted that the reagents mentioned above for improving the efficiency of HDR-mediated gene editing can be prepared into kits, which is more conducive to transportation and storage.

[0026] Fourthly, embodiments of the present invention provide a method for improving the efficiency of HDR-mediated precise gene editing, the method comprising: co-transfecting the above-mentioned micropeptide with other vectors into target cells to achieve precise gene editing through the HDR pathway; the other vectors include CRISPR / Cas9 expression vectors, HDR reporter vectors and donor plasmid vectors.

[0027] Fifthly, embodiments of the present invention provide an application of the above-mentioned micropeptide in improving the efficiency of HDR-mediated precise gene editing.

[0028] Sixthly, embodiments of the present invention provide an application of the above-mentioned micropeptide in the preparation of gene editing drugs, wherein the micropeptide inhibits the NHEJ repair pathway and promotes HDR-mediated gene editing.

[0029] It should be noted that, based on the fact that the micro peptides in this invention can improve the efficiency of HDR-mediated precise gene editing, they can be applied to the preparation of gene editing drugs.

[0030] In a seventh aspect, embodiments of the present invention provide a gene-editing drug, the gene-editing drug comprising the micropeptide or the reagent described above.

[0031] To better understand the present invention, the following detailed drawings and examples further illustrate the content of the present invention, but the content of the present invention is not limited to the examples below.

[0032] Example 1 This invention provides an artificial intelligence-generated micropeptide that promotes HDR-mediated precise gene editing and its screening process. Details are as follows: (1) Target protein complex structure modeling and identification of key interaction interfaces (Hotspots) Structural basis: The cryo-electron microscopy structures of the human Ku70 / Ku80 / DNA-PKcs complex (PDBID: 7Z6O, 7Z88) from the publicly available protein database (PDB) were used as the initial three-dimensional structural model.

[0033] Interface Analysis: PyMOL was used to perform atomic-level fine analysis of the interaction interface between Ku70 / 80 and DNA-PKcs; key "hotspot residues" (V97, F99, W148, A181) crucial for their binding were identified. These residues are typically hydrophobic cores or amino acids forming critical hydrogen bonds / salt bridges; these hotspot residues will serve as core targets for subsequent AI design (see [link to AI design]). Figure 3 ).

[0034] (2) Design and generation of de novo small peptides based on generative AI This step employs a two-stage generation strategy: first, a structural skeleton is generated, followed by sequence design to ensure structural diversity and sequence optimization. Specifically: (2-1) Backbone Generation: Using the advanced diffusion model RFdiffusion, and constrained by the interaction interface between Ku70 / Ku80 protein and DNA-PKcs protein, de novo protein backbones are generated. This process does not pre-set sequence information and focuses on creating novel protein backbones that can form precise complementarity with the target interface in three dimensions. Through this step, a total of 7,000 candidate peptide backbones with novel structures are generated.

[0035] (2-2) Sequence Design: The backbone generated by RFdiffusion is used as a structural template, and the protein design model ProteinMPNN is used to optimize the sequence design. ProteinMPNN can optimize 3 sequences for each backbone while maintaining the stability of the preset backbone structure, for a total of 21,000 structures.

[0036] Design parameter settings: Target region: The interaction interface between Ku70 / Ku80 and DNA-PKcs identified in the first step is defined as the "main binding site" designed by AI. Multi-site strategy: Use multiple sets of "hotspot residues" in the design to ensure the success rate of the design; Constraints: Set the length range of the generated peptide (50-80 amino acids) to ensure its stability.

[0037] (3) High-throughput virtual screening and computational validation of candidate micropeptides This invention employs an automated, multi-level computational screening strategy based on artificial intelligence; specifically as follows: Preliminary screening: AlphaFold2 (AF2) was used for structure prediction, and screening was performed based on multiple parameters, including: plddt: a score for the prediction accuracy of each amino acid residue; pae_interaction: the prediction error of the relative positions between residues.

[0038] A strict threshold was set for the first round of rapid screening (pae_interaction<10, plddt_target>90, plddt_binder>90), and a total of 105 structures were screened out.

[0039] High-precision structure prediction and validation: AlphaFold3 (AF3) is used to independently predict the structure of candidate micropeptides that have passed the initial screening, and further determine that they can stably fold into the backbone structure expected by RFdiffusion, eliminating those candidates whose sequences and structures do not match.

[0040] This study ranked micropeptides based on an interface-predicted template-modeling score (IPTM) greater than 0.85, identifying a total of 301 micropeptides. The top 25 candidate micropeptides with the highest potential were selected (see Table 1 for details). Figure 4 Then, it will enter the subsequent experimental verification stage.

[0041] Table 1. Micropeptides screened based on interface prediction template modeling and scoring. name IPTM PTM Ranking score sequence KU-SB1 0.94 0.91 0.93 EEEEREKRIEELNKKSIEYAKKSKELEEKAKTASPEDKAKLNAEADEYFDKSWEYLKEAIKLRQE(SEQ ID NO.2) KU-SB2 0.94 0.91 0.93 EEEERKKLIEELKKAIELAKKSVEEAKAETASPEEKAKLAKEADELFDKSWEHLKKALKLKQE(SEQ ID NO.3) KU-SB3 0.94 0.91 0.93 EEEEKEKKRKELKEKALALAKESKALEEQLPNASPEQKPALQAQIDALFDASSEYIAKAIKLTNE(SEQ ID NO. 4) KU-SB4 0.94 0.92 0.94 SLEEEIEKAKKELEEVKEKIKELVEKRKKAEKESSAKAAELEKKIDELEDEGMKLIKKLLELKEE(SEQ ID NO.5) KU-SB5 0.94 0.92 0.93 LSPEDKAKREELKKAIELAKKAKELEEKTKTASAEEKKKLTAEADKYFAESTKYVAEAIEIVQK(SEQ ID NO.6) KU-SB6 0.94 0.91 0.93 EEEEKKKKIEELKKKSIEYAKKSKEYEEKAKTPASEKAELKAKAAEEYWELSNKYLSEALKLQNE(SEQ ID NO.7) KU-SB7 0.94 0.93 0.94 LSPEKLEKQKEYRKKAIEAAKKAKALEEKAKTASPEEKAKLAEEIDKEEKESSEYLKKAIALNQE(SEQ ID NO.1) KU-SB8 0.94 0.92 0.93 EEEEREKLKEEYRKKAIEYAKKAKALKEKMKTASPEEKAKMEKEIEKLEAKSSEYLKKSIKLVEE(SEQ ID NO.8) KU-SB9 0.94 0.93 0.94 EEEEREKLREELKKAIETAKKAKELEKLPELSPEEKAKALEEIEKLEKESSENIAKRIELVES(SEQ ID NO.9) KU-SB10 0.94 0.92 0.93 EEEEREKRRKELKEKALKAAQEAKKLEEAKDLSPEEKAEADKKIDALFDESSKYIAERIKLVEE(SEQ ID NO.10) KU-SB11 0.94 0.91 0.93 DEEEREKIKILEKKIESKAKAAKEAKDKLKTASPEEKKELEAKYDEFAESKYLSESICYKLQDE(SEQ ID NO.11) KU-SB12 0.93 0.91 0.93 EEEEKKEIKELNEXIALAKEAKKLEEKKPELSPEQAKAEKQIDELMAXSEYLAKAIKLREE(SEQ ID NO.12) KU-SB13 0.93 0.91 0.93 EEEERKERQEELRKKSIELAKKAKELKEKEKAKTATPEEKAKLAAEADALEDESWAYLGEAIKLVQE(SEQ ID NO.13) KU-SB14 0.93 0.9 0.92 KEEEIEKKIEENKKSIELAKKAKELEENKPNLTPEEQAKADKEIKELEDKSFKLISENIKLKEE(SEQ ID NO. 14) KU-SB15 0.93 0.91 0.93 ELEKKIKEAKKKIEELKKKVKEKYKEYEKAKSEDSKAKELEKELDELEKENLKAIKELFEAEEK(SEQ ID NO.15) KU-SB16 0.93 0.91 0.92 LPEEVKKKVEELREEAIKLAKKAQELREKAKTASPEEKAKLEAEADKYFDEAFKKVNEAIKLVMS(SEQ ID NO.16) KU-SB17 0.93 0.9 0.93 EKEEKIKELKKKIESKKKVKEKYEKAKSIDSALAAKLQEELDKLESENLKLIGELFKVEEE(SEQ ID NO.17) KU-SB18 0.93 0.91 0.93 EEEKRKKLKKEEYRKKAILELAKKAKELEENKPNLTPEEKAKAEKEIEKLEKESWKLLKKSIELVEE(SEQ ID NO. 18) KU-SB19 0.93 0.91 0.92 SLEEEIKEAEKKIELKKKIKEAVDEAKKYESTDSDAKAAELEKEIDKLESENLEAIKKFELREK(SEQ ID NO.19) KU-SB20 0.93 0.91 0.93 KEEELKELREELRKKAIELGQKAVELKEKMKTASAEEKAELQAKIDELEDEGFKYLAKAIKLNEE(SEQ ID NO.20) KU-SB21 0.93 0.91 0.92 KEKEKEEEKKKLKEKAIEAAAKKAKKLEEKPKLTPEEKAKADKEIDKLESESSKYLKEAIKLVEE(SEQ ID NO. 21) KU-SB22 0.93 0.91 0.92 ETEEKIKEAEKKIAELNKKVTAAYKAVEAAKSGSSAEKAAAEAALDAAMDANMKAIKEKLKLMEE (SEQ ID NO.22) KU-SB23 0.93 0.91 0.93 EEEERRKKQEELRKKSIELAKKAKELLEQTKTASPEEKAVLTSQADKLEAESSKYLAEAIKLVEE (SEQ ID NO.23) KU-SB24 0.93 0.9 0.93 KEEELKKKQKELNKIAIEKAKKAKELKEKAKTASEEEKTKLLSEADKLFDESFKYVAKAIKLVEE (SEQ ID NO.24) KU-SB25 0.93 0.91 0.93 EEEERKKLIEEYRKKSIETAKKAKEVKDNMPKLTPEEQAKAAKEYEALEAESTELLKKAILLKQE (SEQ ID NO.25)

[0042] Example 2 This invention provides an embodiment (Example 1) that verifies the efficiency of micropeptides in promoting HDR in cells using micropeptides generated and screened based on artificial intelligence. Details are as follows: (1) Micro peptide gene synthesis and expression vector construction The selected candidate micropeptide amino acid sequences were codon-optimized (for human cell expression systems), and the corresponding DNA gene fragments were synthesized and integrated into mammalian cell expression vectors.

[0043] (1-1) Codon optimization of the amino acid sequence of miniature peptides Based on the amino acid sequences of the micropeptides shown in Table 1, the corresponding nucleotide sequences were generated using GenSmart™ CodonOptimization (www.genscript.com) and JCAT (www.jcat.de / ) online codon optimization software (see Table 2).

[0044] Table 2. Codonality optimization of nucleotide sequences encoded by mini peptides Name Sequence Length / bp KU-SB1 GACAAGGAGGTGAAGGAGAAGTCCACAGAGCTGCTGAAGAAGGCCGTGACAGGCGATGCCGAGACAGCTAAGGCCGAGGCTAAGAAGCTGGAGGAGCTGGCTAAGAAGGGCAATAAGTGGGCCAAGGACGCCGCCACAATGGCTAAGGACGTGGCTGCTGAGAAGGAGGCC (SEQ ID NO.26) 171 KU-SB2 AGCGCTATCGCTTCCTACATGACCGGCAGGAAGGCCGCTAAGAAGCTGGGAGAGGAGGGCAAGGACTACCTGGCTCAGCTGGACGCTCTGGCTAAGGCTGCTAATGCCGGCAGCCCTGCTGACCAAGCTGCTTATGCCGCCCAGATCGCCGCTCTGACAAAGGAAATCAACGCCAAGGCCGCCGCC (SEQID NO.27) 186 KU-SB3 GCTTCTGAGCTGCTGGCTAAGGCCGACGAGTACCTGGACAAGGCCGATGAGTACTGGAAGAAGGCCCAGAAGGCCCCCGCTTCTGAGAAGGGAGCTCTGCTGGCTGCCGCTAAGGAGTACGCTAAGAAGAGCGCCGAGTACAAGAGGGCCGCCGCTGAG (SEQ IDNO.28) 159 KU-SB4 AGCTCTGCTAAGGAGGAGGCCGCTGAGCTGAATAAGAAGTCCGAGGAGTACTGGGAGAAGTCCACAGAGTACCTGAAGAAGTCCCTGGAGGCCCAGGACGCTGGAGATTTCGAGGCTGCTAAGAAGTACAGAGAGAAGTCCATCGAGTACGCCAAAGTCCAAGAAGTACGAGGAGGAGGAGGAGCAGAGGGAQGGAGCTAGTAAGGAGGAGGAGGAGGAGGAGGGAQGGAQTAGTAAGGAGGAGGAGGAGGAGGGAGCTGAGAGCTGAGTAAGTAAGGAGGAGAAGTCCATCGAGTACGAGAGGAGCTGAAGTAAGGAAGGAGGAGGAGGAGGAGGAGGAGCTGAGAGAGTCCCTGGAGGCCCAGGACGCTGGAGATTTCGAGGCTGCTAAGAAGTACAGAGAGA NO.29) 201 KU-SB5 GCTGCTGCTCTGAAGGAGCAGGCTAAGGCCAACCTGGCCAAGGCTAAGGCCGTGGAAGCCGTTGCTGAAGATCAAGGCCAGCACCGCCAGAGGAGAACAAGGAACTGGGCCTGAAGAACGCCGA 180 KU-SB6 GACTACAAGGCCAAGGCCGAGGAGTACCTGGAGAAGGCTAAGAAGGTGGAGGAGGTCATGAAGAAGCGCCGAGAAGTCCCCTTACAACGCCGATATCCGAGAGGATGGCCAAGTCCAAGACCAAGGAGCTGAACGACCTGGCCAAGGAGTACGAGGAGAAGGCCAAGAAGGCC(31 ID. 177 KU-SB7 GAGGAAGAGGAGAAGAAGAAGCTGAAGGAGGAGTACAAGAAGAAGGCCCTGGAGTACGCCAAGAAGGCCAAGGAGCTGATCGACAAGGCCGAGGCTTCCACCCCTGAGAAGGCTAAGGAGTACCTGGCCAAGGCCGATGAGTACGAGAAGAAGTCCAGCGAGTACCTGAAGAAGGCCATCAAGCTGAGAGGGAGGGAAGCTGAAGAAGGCCATCAAGCTGAAGCTGAGGAGCTGAAGAAGGCCATCAAGCTGAGAGCTGAAGCTGAAGAAGGCCATCAAGCTGAGAGCTGAAGCTGAAGAAGGCCATCAAGCTGAAGCTGAAGCTGAAGAAGGCCATCAAGCTGAGCTGAAGCTGAAGCTGAAGAAGGCCATCAAGCTGAAGCTGAAGAAGGGAAGCTGAAGCTGAAGAAGGGAGGAGGCTGAAGCTGAAG32. 195 KU-SB8 GCTGCTGCTGTTGAGGAGCACAGGAAAGAAGCAGGAGGCCATCGAGAAGTACAAGGAGCTGACCAAGAAGGCGAGGAGGCCAAGGATCCCGCTAAAACAGGCCAAGCTGCTGGCTGAGGCCGAGAAAGCTCTGGATAAGGGCTTTGAGTACGCCGCCGCCTATCAAGGCCAGC(SEQ ID NO.33) 180 KU-SB9 GAGGAAGAGGAGAGAAGGAAGATCAGGGAGGAGGCCAGAAAGGCCCTGCTGGAAGGAGGAAAGCTGTACAAGCTGGGCCTGGAGTACCTGGAGAAGGGCAAGGCTACCGCGACTCTGAGCTGCTGGCTAAGGGAGAGGAGTACCTGAGCAAGGCCGAGGAGCTGAATAAGAAGTACATCGAGCTGAGAGAGAGTACAGGAAG(SEQ ID NO.34) 204 KU-SB10 ACCGCTGCTCTGAAGGCTGCTGCTAAGGCCGCTGCTGATGAGTACGATGCCAAGTCCGACGAGCTGCTGAAGGAGGCTATCCAGGATCCCAGCAATAAGGAGACCAAGCAGGCCGCCATCGAGTACGCTAAGAAGGCCAAGGAGCTGGACGAGCTGGCCGCT(SEQ ID NO.35) 162 KU-SB11 GCTGCTGAAGAGGCTAAGAAGGTGGTGGAAAGCTGCTGGAGAAGGCCAAGGAGAAGGCCAAAAAAGAATATCGAGAACGGCGAGCTGTACGCCAAGCTGGCTGCTTCTTACGCCGAGATCGAGCTGGGCAAGGAGGAGGGAAAGAAGGTGAAGGAGGGAGGCCGAGAAGCTGATCGAGGAGTGAAGAAG (SEQ ID NO.36) 189 KU-SB12 AGCTCTGCTCTGAAGGAGGCCAATAAGTACAACGAGAAGATCGCCAAGATCTACAAGGAGGCGAGAAGGCCAAGGAGGAGGCTGAAAAGACATACGCAAGGCCGCCGCCGATGCTGTTTACAATGAGACATTTGATAAGATCGACGCCCTGGAGAAGGAGAAGAGGAAGGAGCTGGAGGCCATCAGAGAG(SEQ ID NO.37) 195 KU-SB13 GCTGCTGAAAAGGCCAAGGCCGAGGCTGAGAAGGTGAAGAAGGAGGCGAGGACCATCGAGAAGATCAAGAAGGCCGCCGAGCTGTTGCCAAGAAGACCAAGAGCAAGAAGGAGGGAGAGATCTTCAAGAAGAGCGCCGAGAAGGAGCAGGAGCAGATCAAGACAGACACCGAGAAGAAGGTGAGCGAGCTGCTGTCCC(SEQ ID NO.38) 201 KU-SB14 ATGGAGGTGGCTGAGAAGGCCGTGGAGGCTGTTAGGGAGGGAGACTTTGAGAAGGCCCTGGAGTACGCCAGAATCGCCGTGAAGGATCCTGCCACCCTGGACGCTCTGTCTGTCTGGCTATGAGATCCGATAAGGAGACCAGAAAAGAAGATGAACGAGGCCATCAGAAAGGCCAGGGAGGAGCTGGAGAAGGAG(SEQ ID NO.39) 195 KU-SB15 GCTGCTGCTCTGAAGGCTGAGGCTGCTGCTGCAGCTACCTACGACAAGCTGAGCGCCGAGGCTAGCAAGTACCTGAAGCTGTCCATCGCCGCCGTGGAGGCTGGAGACGTGGCTGCTGCTAAGGAGTACAGAGAGAAGGCCTGGCCGCCCTGAAGGAAGCTAAGAAGGCCCTGGACAAGGAGGACGCCGCTAGAGCTGCTGCTGCCGCTGCTGCT (SEQ ID NO.40) 219 KU-SB16 AGCGAAAACAAGGAGAAGGCCAAGGAGCTGAGGAAGAAGTCCGCCGAGCTGTACAAGAAGGCCGAGGAGTACGCCAAGGACCCTAGCAAGGCCGCTGAGGCTAAGAAGCTGGAGGAGCTGAGCGACAAGTACTTTTAAGGAGAGCATGAAGCTGGCCGAC(SEQ IDNO.41). 159 KU-SB17 GACACAACAGAGAAGATCCTGATCACACTGGCCAAGGACGCCCTGGAGGAGGCTAAACAGGCCTAAGGATGGCGATCTGGAGAAGGCCAAGCAGTACCTGAATGCCCTGAAG GTCATGGAGGAGAGAGCCAAGAAGAAGGGCTACAAGAAGGCCGCCGAGGTGGCTGAAAAGGCCGCTAAAGGCCGAGAAGGCCCTGAAGGCCCTGAAAGAGAAGCTGGAG(SEQ ID NO.42) 225 KU-SB18 GCTGCTGCTAAGGCTGCTGCTCTGGCTAAGGCCGACGAGTACTGGAAGAAGGCCAGCGAGCTGCAGAGAGAGGCCCTGAAATCCAGCCCTGAGGAGGCTAAGAAGCTGAACAAGGAGGCCGCCGAGTACCTGAAGAAGTCCGATGCCGCTGGCCGCTGCTGCT(SEQ ID NO.43) 165 KU-SB19 GAGGAAGAGGAGATCAGGGAGGCCTACATCGAGAATGCCAAGATGGCCCAGAGGCTGGCCGAGGACGCTAAAAGAGCCGCTGAGGAGGCCAGCCCTGAGGACGCTGCTCTGTATAAGAAGGCCGCCGAGTACTTTAAAGAAGGCCGAGGAGCTGCTGAAGAAGGCCGCT(SEQ ID NO.44) 171 KU-SB20 AGCGCTGCTAAGGAGGCTGAGGAGAAAGAAGGAGCTGCTGGCCAAGGCCGAGGAGGCTCAGAAAGGCTGACAAGTACTTCGATAAGATGACCAAGCTGCTGAATGAGGCCCTG AAGGGCGATTACGAGACCGCTAAGGAGAAAGAAAGAGATCAAGAAGGTGAAGGAGGAGTACAACAAGGCCAAGGAGGATGCCGAGAAGTACAAAGAAGGCCGAGGAACTG(SEQ IDNO.45) 231 KU-SB21 ATGAACAAGGAGGAGCTGAAGGCCAAGGCCAATGAGTACGCCAAGAAGCGAGGAGTACTACGGCCGCCGAGAAGGCTGCTAAGGAGGATCCTGCCAAGGCCAAGGAGTACCAGGCCAAGGCTGATGAGTACGACGCCAAGTCCCTGGAGCTGCAGAGGAAGGTGCTGTCC(SEQ) ID NO.46. 174 KU-SB22 GAGGAAGAGAAGAAGAAGCTGCTGCTGAAGGCCGAGGGCCTGAAGAAGAATGCCGAGAGCGAAGAAGATCGCCGAGAAGCTGAAGGCCCTGGCCGAAAAGGCCAGC CCTGAAGAAGCCGCTGGCTACAAGAAGGGCGCCGAAGAAGCCGAGGAGAATTACAAGGGGCCCTGAAGGTGGCCGAGGAGTACGAAAAGAAGGCCGAGGAGCTG(SEQ ID NO.47) 213 KU-SB23 AGCAACATCAAGGCCTACATGGAGGCCAGAAAGGCCGCCAAGAAGCTGGGAGAGGAGGAAAGGAGTACCTGAAGAAGATCGACGAGCTGGCCAAGGAGGCCGAGCTGGAAGCCCTGGCTACAATGGAGGCCAAGACAGGAGGATCAAGAAGATCAAGGAGGAGGCCGAGAGAGGAGCTGCTGCT NO.48) 186 KU-SB24 GCTTCTGAGGAGGCTAAGAAGAAGGCCGAGGCCGTGAAGAAGGAGACCGCTGAAACAGTGGCCAAGATCGATAAGGCCACCGCCAGATACGCCGCCAAGACATCTGATCCCGCCGCTGCTACCCACTTTACAGCCACAGGCGCCAAGGAGAAGGCCGAAGTTACCGCTGCCGGCGCTGCTGAGATCGCTAAACTGGAGGCC (SEQ ID NO. 49) 201 KU-SB25 GACTACGAGGAGAAGGCCAAGAGATACAAGGAGGCCTACGAGAAGATGAAGGAGAGGTACGAGAAGGCCAAAAAGAATGACCCCGAGACCGCCAAGATCGCCAAAGGCTAACCTGGAGTCCGTGAAGGAGGAGTACGAGAAGTACAAGGCCAAGGCCGAGGAGGAGAAGAAGAAGAAG (SEQ ID NO. 50) 177

[0045] (1-2) Construction of pCAG-Puro-P2A vector 1) Puro-2A sequence synthesis Based on the pCMV-Puro-2A-MLH1-SB (Addgene #228871) vector sequence information already submitted on Addgene, the Puro-P2A coding sequence can be obtained. To integrate the Puro-P2A coding sequence into the pCAGGS vector (NovoPro, V008798) via homologous recombination, corresponding homologous sequences (25 bp) from both ends of the pCAGGS backbone vector were added to the 5' and 3' ends of the Puro-P2A coding sequence. The above DNA sequence was sent to Jiutian Gene Technology (Tianjin) Co., Ltd. for sequence synthesis. The sequence is as follows: The Puro-P2A coding sequence is as follows (SEQ ID NO.51): ; The homologous sequence added to the 5' end (SEQ ID NO.52): 5'- GGCAAAGAATTCGAGCTCG -3', The homologous sequence added at the 3' end (SEQ ID NO.53): 5'- GATCTGCTAGCTCGAGTCG-3'.

[0046] 2) Linearized pCAGGS vector The pCAGGS vector was linearized using the KpnI restriction endonuclease (NEB, R3142). The enzyme digestion system (see Table 3) was mixed and then incubated in a 37°C water bath for 30 min.

[0047] Table 3 Enzyme digestion reaction system

[0048] 3) Recombination of linearized pCAGGS vector with Puro-P2A sequence The linearized pCAGGS vector fragment after enzyme digestion was detected by agarose gel electrophoresis, and the target band after enzyme digestion was recovered using a DNA recovery kit (Omega, D2500) (see [link to kit]). Figure 5 The recovered linearized pCAGGS vector and the synthesized Puro-P2A sequence were recombined using a homologous recombination kit (Vazyme, C115-01). The recombination reaction system (see Table 4) was mixed and placed in a PCR instrument at 50°C for 10 min, and then stored at 4°C.

[0049] Table 4 Recombination Reaction System

[0050] After the recombinant reaction was complete, 10 μL of the recombinant product was added to 100 μL of competent cells (DH5-α, Vazyme#C502), and the mixture was gently tapped against the tube wall to mix. The mixture was then incubated on ice for 30 min. After heat shock at 42°C for 45 sec, the cells were immediately cooled on ice for 3 min. 900 μL of ampicillin-free (AMP)-free LB broth was added, and the cells were incubated at 37°C for 1 h (220 rpm). The cells were centrifuged at 5,000 rpm for 5 min, and the 900 μL supernatant was discarded. The cells were resuspended in the remaining broth and gently spread onto AMP-resistant plates using a sterile spreader. The cells were incubated upside down at 37°C for 12 h.

[0051] Plasmids (Omega, D6943) were extracted from bacterial cultures of different monoclonal strains and identified using double enzyme digestion (see [link to relevant documentation]). Figure 6 a) Sanger sequencing. The plasmid that is correctly sequenced is the pCAG-Puro-P2A vector. This vector contains the CAG promoter, the Puro sequence, and the P2A peptide to form an expression cassette (see [link]). Figure 6 b).

[0052] (1-3) Construction of miniature peptide expression vectors (pCAG-Puro-P2A-KU-SBs) 1) Micro peptide sequence synthesis To integrate the optimized micropeptide-encoding DNA sequences into the backbone vector (pCAG-Puro-P2A), homologous sequences corresponding to the ends of the pCAG-Puro-P2A vector were added to the 5' and 3' ends of all micropeptide-encoding DNA sequences. These DNA sequences were then sent to Jiutian Gene Technology (Tianjin) Co., Ltd. for sequence synthesis. The sequences are as follows: The 5' homologous sequence (SEQ ID NO.54) is: 5'- GGAGAACCCTGGACCTGGTACC -3'. The 3' homologous sequence (SEQ ID NO.55) is: 5'-GATCTGCTAGCTCGAGTCGTTA-3'.

[0053] 2) Linearization of pCAG-Puro-P2A vector The pCAG-Puro-P2A vector was linearized using the XhoI restriction endonuclease (NEB, R0146). The enzyme digestion system (see Table 5) was mixed and then incubated in a 37°C water bath for 30 min.

[0054] Table 5 Enzyme digestion reaction system

[0055] 3) Recombination of linearized pCAG-Puro-P2A vector with micropeptide coding sequences and vector identification Agarose gel electrophoresis was performed, and the linearized pCAG-Puro-P2A vector band was recovered using a DNA recovery kit (Omega, D2500) (see [link to kit]). Figure 7 The recovered linearized pCAG-Puro-P2A vector fragment was recombined with the gene-synthesized micropeptide-encoded DNA fragment using a homologous recombination kit (Vazyme, C115-01). The recombination reaction system (see Table 6) was mixed and placed in a PCR instrument at 50°C for 10 min, and then stored at 4°C.

[0056] Table 6 Recombination Reaction System

[0057] After the recombinant reaction was complete, 10 μL of the recombinant product was added to 100 μL of competent cells (DH5-α, Vazyme#C502), and the mixture was gently tapped against the tube wall to mix. The mixture was then incubated on ice for 30 min. After heat shock at 42°C for 45 sec, the cells were immediately cooled on ice for 3 min. 900 μL of ampicillin-free (AMP)-free LB broth was added, and the cells were incubated at 37°C for 1 h (220 rpm). The cells were centrifuged at 5,000 rpm for 5 min, and the 900 μL supernatant was discarded. The cells were resuspended in the remaining broth and gently spread onto AMP-resistant plates using a sterile spreader. The cells were incubated upside down at 37°C for 12 h.

[0058] Plasmids (Omega, D6943) were extracted from bacterial cultures of different monoclonal strains and identified using a double enzyme digestion method (see [link to relevant documentation]). Figure 8 a) Sanger sequencing. The plasmid that is correctly sequenced is the pCAG-Puro-P2A-KU-SBs vector. This vector contains the CAG promoter, the Puro sequence, the P2A cleavage peptide, and a mini peptide sequence forming an expression cassette (see [link to documentation]). Figure 8 b).

[0059] (2) Constructing CRISPR / Cas9 expression vectors for the target gene (2-1) Design of sgRNA for target gene loci Taking EMX1 as the target gene as an example, the EMX1 genome sequence was first searched in NCBI (www.ncbi.nlm.nih.gov). Using the CHOPCHOP (https: / / chopchop.cbu.uib.no / ) online sgRNA design website, target sequences suitable for CRISPR / Cas9 editor recognition were screened within the EMX1 genome sequence. Based on the efficiency, GC content, and number of off-targets of each sgRNA sequence in the prediction results, the following sequence was selected as the nuclease target site of the EMX1 gene (nT-EMX1) (SEQ ID NO.56): 5'-GAGTCCGAGCAGAAGAAGAAGGG-3'. The portion of this site preceding the PAM (e.g., GGG within the black box) is cleaved by Cas9 proteins (e.g., spCas9).

[0060] (2-2) sgRNA Fitting Sequence Synthesis and Annealing To integrate the sgRNA sequence into the CRISPR / Cas9 expression vector (pX330-U6-Chimeric_BB-CBh-hSpCas9, Addgene #42230) via enzyme digestion and ligation, BsaI restriction endonuclease recognition sequences need to be added to both ends of the sgRNA sequence. The fitted sgRNA sequence of the EMX1 gene is as follows: EMX1-F (BsaI) (SEQ ID NO.57): 5'-caccGAGTCCGAGCAGAAGAAGAA-3' EMX1-R (BsaI) (SEQ ID NO.58): 5'-aaacTTCTTCTTCTGCTCGGACTC-3' Primers EMX1-F and EMX1-R were mixed according to the components in Table 7 and placed in a PCR instrument for annealing and fitting (primer annealing program: 95℃, 5 min; -0.1℃ / cycles; store at 4°C) to obtain a DNA double-stranded sequence with sticky ends after being digested by BsaI restriction endonuclease.

[0061] Table 7 Primer Annealing System

[0062] (2-3) Linearization of CRISPR / Cas9 expression vectors The CRISPR / Cas9 expression vector was linearized using the BsaI restriction endonuclease (NEB, R3733). The enzyme digestion system (see Table 8) was mixed and then incubated in a 37°C water bath for 30 min.

[0063] Table 8 Enzyme digestion reaction system

[0064] (2-4) Ligation of linearized CRISPR / Cas9 expression vector with sgRNA annealing product and vector identification Agarose gel electrophoresis was performed, and the linearized CRISPR / Cas9 expression vector bands were recovered using a DNA recovery kit (Omega, D2500) (see [link to kit]). Figure 9 The recovered linearized CRISPR / Cas9 expression vector fragment and sgRNA annealing product were ligated using T4 DNA ligase (NEB, M0202). The ligation reaction system (see Table 9) was mixed and placed in a PCR instrument at 16°C for 30 min, and then stored at 4°C.

[0065] Table 9 Connection Reaction System

[0066] After the ligation reaction, 10 μL of the recombinant product was added to 100 μL of competent cells (DH5-α, Vazyme#C502), and the mixture was gently tapped against the tube wall to mix. The mixture was then incubated on ice for 30 min. After heat shock at 42°C for 45 sec, the cells were immediately cooled on ice for 3 min. 900 μL of ampicillin-free (AMP)-free LB broth was added, and the cells were incubated at 37°C for 1 h (220 rpm). The cells were centrifuged at 5,000 rpm for 5 min, and the 900 μL supernatant was discarded. The cells were resuspended in the remaining broth and gently spread onto AMP-resistant plates using a sterile spreader. The cells were incubated upside down at 37°C for 12 h.

[0067] Plasmids (Omega, D6943) were extracted from bacterial cultures of different monoclonal strains, and the plasmids were sequenced for identification (see [link]). Figure 10 a) The plasmid that was correctly sequenced is the CRISPR / Cas9 expression vector of the EMX1 gene (CRISPR / Cas9-sgEMX1). This vector contains the U6 promoter, the sgRNA sequence of the EMX1 gene (EMX1-sgRNA), the CMV promoter, and the Cas9 coding sequence (see [link to CRISPR / Cas9 expression vector]). Figure 10 b).

[0068] (3) Constructing a target gene-specific HDR reporter vector (3-1) Construction and identification of HDR reporter Using the mammalian expression vector pcDNA3.1 (Invitrogen, V79020) as the backbone vector, we selected PuroL from the pHDR-PMG (Addgene, #208858) vector we constructed previously. 1-225 -mCherry-PuroR 330-597 The -T2A-EGFP sequence was used to construct an HDR reporter vector through enzyme digestion and ligation.

[0069] The pHDR-PMG vector and pcDNA3.1 vector were double-digested using HindIII (NEB, R3104) and XhoI (NEB, R0146) restriction endonucleases. The digestion system (see Table 10) was mixed and incubated in a 37°C water bath for 30 min.

[0070] Table 10 Enzyme digestion reaction system

[0071] Agarose gel electrophoresis was performed, and the double-digested pHDR-PMG vector and pcDNA3.1 vector were recovered using a DNA recovery kit (Omega, D2500) (see [link to kit]). Figure 11 ), will recycle PuroL 1-225 -mCherry-PuroR 330-597 The -T2A-EGFP fragment and the pcDNA3.1 vector fragment were ligated using T4 DNA ligase (NEB, M0202). The ligation reaction system (see Table 11) was mixed and placed in a PCR instrument at 16 ℃ for 30 min, and then stored at 4°C.

[0072] Table 11 Connection Reaction System

[0073] After the ligation reaction, 10 μL of the recombinant product was added to 100 μL of competent cells (DH5-α, Vazyme#C502), and the mixture was gently tapped against the tube wall to mix. The mixture was then incubated on ice for 30 min. After heat shock at 42°C for 45 sec, the cells were immediately cooled on ice for 3 min. 900 μL of ampicillin-free (AMP)-free LB broth was added, and the cells were incubated at 37°C for 1 h (220 rpm). The cells were centrifuged at 5,000 rpm for 5 min, and the 900 μL supernatant was discarded. The cells were resuspended in the remaining broth and gently spread onto AMP-resistant plates using a sterile spreader. The cells were incubated upside down at 37°C for 12 h.

[0074] Plasmids (Omega, D6943) were extracted from bacterial cultures of different monoclonal strains, followed by plasmid digestion and identification (see [link to relevant documentation]). Figure 12 a) The correctly identified plasmid is the HDR reporter vector, which contains the CMV promoter and the PuroL homologous sequence of the resistance gene coding region. 1-225 The homologous sequence PuroR in the coding region of mCherry and the resistance gene 330-597 Stop codon (TAA), T2A cleavage peptide, EGFP, polyA (see...) Figure 12 b).

[0075] (3-2) Amplification of the nT-EMX1-mCherry-nT-EMX1 sequence Taking the EMX1 gene as an example, in order for the HDR reporter vector to be specifically recognized and cleaved by the CRISPR / Cas9 expression vector of the EMX1 gene (i.e., CRISPR / Cas9-sgEMX1), the recognition sequence of the CRISPR / Cas9-sg EMX1 vector (i.e., nT-EMX1) needs to be integrated into the HDR reporter vector. A forward primer is obtained by adding a BamHI recognition site to the 5' end of the target sequence and a partially complementary mCherry sequence (see a portion of the 5' end of the mCherry sequence) to the 3' end. A reverse primer is obtained by adding a partially inverse complementary mCherry sequence (see a portion of the 3' end of the mCherry sequence) to the 5' end of the target sequence and a stop codon and a NotⅠ recognition site to the 3' end. The forward and reverse primer sequences are as follows: nT-EMX1-F (BamHI) (SEQ ID NO.59): 5'-cgcGGATCCGAGTCCGAGCAGAAGAAGAAGGGATGGTGAGCAAGGGCG-3'; nT-EMX1-R (NotⅠ) (SEQ ID NO.60): 5'-tatGCGGCCGC TTA CCCTTCTTCTTCTGCTCGGACTCCTTG TACAGCTCG-3'.

[0076] In the above reverse primer sequences, the underlined part is the reverse complementary base of the additional stop codon; the lowercase letters are protective bases, and the uppercase bold italic letters are restriction endonuclease recognition sites. The designed primers were sent to Qingke Biotechnology Co., Ltd. for synthesis (the same applies below).

[0077] After primer synthesis, PCR amplification was performed using the pAAV-minCMV-mCherry vector (Addgene #27970) as a template and DNA polymerase (TaKaRa, R045A) (see Table 12) to obtain the target fragment (nT-EMX1-mCherry-nT-EMX1) containing the EMX1 gene target site at both ends (nT-EMX1) and the mCherry gene coding sequence in the middle.

[0078] Table 12 nT-EMX1-mCherry-nT-EMX1 Sequence Cloning and Amplification System

[0079] The PCR reaction conditions were: (98 ℃, 10 sec; 58 ℃, 10 sec; 72 ℃, 30 sec) × 33 cycles.

[0080] After the PCR reaction was completed, agarose gel electrophoresis was used for detection (see [link]). Figure 13 The target band (778 bp) was recovered using a DNA recovery kit (Omega, D2500), thus completing the cloning of the nT-EMX1-mCherry-nT-EMX1 sequence and introducing restriction endonuclease (i.e., BamHI, NotHI) recognition sites at both ends of the sequence. The nT-EMX1-mCherry-nT-EMX1 sequence includes the following elements arranged in sequence: nT-EMX1, mCherry, nT-EMX1, stopcodon.

[0081] (3-3) Constructing a specific HDR reporter vector for the EMX1 gene (EMX1-HDR reporter) Using an HDR reporter vector as a backbone, the backbone vector and the target fragment containing the nT-EMX1-mCherry-nT-EMX1 sequence obtained by PCR amplification were respectively integrated into the backbone vector by BamHI / NotI double digestion and ligation, resulting in the EMX1-HDR reporter vector (see [link to documentation]). Figure 14 a) Identification was performed using HindIII (NEB, R0104V) / XhoI (R0146V) enzyme digestion (see [link]). Figure 14 b) Sequencing verification. The EMX1-HDRreporter vector includes the following elements in sequence: CMV promoter, PuroL homologous sequence of the resistance gene coding region. 1 -225 The nT-EMX1-mCherry-nT-EMX1 sequence and the PuroR homologous sequence in the coding region of the resistance gene. 330-597 Stop codon, T2A cleavage peptide, EGFP, polyA. The stop codon (specifically TAA) in the nT-EMX1-mCherry-nT-EMX1 sequence is located in PuroR 330-597 Previously, therefore: 1) Prevents T2A-EGFP in the vector from functioning when not targeted; 2) When firing at the target carrier, PuroL 1-225 With PuroR 330-597The sequence can serve as a homologous arm, using the Puro sequence in the miniature peptide expression vector (pCAG-Puro-P2A-KU-SBs) as a template for HDR repair. Therefore, in PuroL 1-225 With PuroR 330-597 Intermediate components, such as mCherry, can be expressed normally when the vector is not targeted.

[0082] (2) Verification of the efficiency of micropeptide-mediated intracellular HDR repair: Different micropeptide expression vectors (pCAG-Puro-P2A-KU-SBs), CRISPR / Cas9 expression vectors, and target gene-specific HDR reporter vectors were transfected into cells. The efficiency of micropeptides in promoting HDR repair was verified by flow cytometry, and the micropeptides with the highest efficiency in promoting HDR repair were further screened.

[0083] Taking the EMX1 gene locus as an example, cells were seeded in 24-well culture plates (Corning, 3738) 12 h in advance. Using cell transfection reagent (jetPRIME, 101000046), different micro peptide expression vectors (pCAG-Puro-P2A-KU-SBs), CRISPR / Cas9-EMX1 expression vector, and EMX1-HDRreporter reporter vector were transfected into cells (e.g., HeLa cells) at a molar ratio of 1:3:2. At the same time, the pCAG-Puro-P2A vector CRISPR / Cas9-EMX1 expression vector and EMX1-HDR reporter vector were selected as the control group for this experiment.

[0084] Forty-eight hours after cell transfection, the working status of the HDR reporter vector in different treatment groups was observed using a fluorescence microscope (see [link to relevant documentation]). Figure 15 Cells were digested into 1.5 mL EP tubes and centrifuged at 1500 rpm for 2 min. The cell pellet was resuspended in PBS (Procell, PB180327) containing 2% FBS and placed on ice. After incubation with DAPI (1 µg / mL) for 2 min, the cell suspension was filtered through a 200-mesh cell strainer into flow cytometry tubes. The efficiency of different micropeptides in repairing the HDR reporter vector in cells was detected using a flow cytometer (Beckman CytoFLEX LX) (see [link to relevant documentation]). Figure 16 a). FACS: Cells are analyzed by flow cytometry using red fluorescence (mCherry). + ) / Green fluorescence (EGFP) +Channel sorting and counting were performed to calculate the efficiency (%) of micropeptides in HDR repair in cells = 100% × HDR repair positive cells (i.e., EGFP-expressing positive cells) / live cells (i.e., DAPI-expressing negative cells).

[0085] Flow cytometry analysis showed that the reporter vector repair efficiency using the pCAG-Puro-P2A control group (co-transfected with CRISPR / Cas9-EMX1 and EMX1-HDR reporter) was 18.36% (17.5% ± 0.86%), while the reporter vector repair efficiency using the pCAG-Puro-P2A-KU-SBs experimental group (co-transfected with CRISPR / Cas9-EMX1 and EMX1-HDR reporter) ranged from 19.16% to 30.94% (see [link to relevant documentation]). Figure 16 b). That is, co-transfection with different micropeptide expression vectors can increase the repair efficiency of the HDR reporter system in cells by 1.04 to 1.68 times (see [link]). Figure 16 c).

[0086] Example 3 This invention provides an example of the application and validation of the micropeptide (KU-SB7) obtained in Example 2, which has the highest efficiency in promoting HDR repair, in precise cell editing. Details are as follows: Taking the EMX1 gene as an example, HeLa cells co-transfected with the pCAG-Puro-P2A-KU-SB7 vector, the CRISPR / Cas9-EMX1 vector, and the donor vector (pd-EMX1) were used as the experimental group for micropeptide-mediated HDR precise editing. The pd-EMX1 donor vector, using pMD19-T (Novopro, V017510) as its backbone, sequentially included approximately 1,000 bp sequences upstream and downstream of the EMX1 gene target site. Simultaneously, the PAM and nearby sequences at the EMX1 gene target site were replaced with EcoRI restriction endonuclease recognition sites (see [link to documentation]). Figure 17 HeLa cells co-transfected with pCAG-Puro-P2A vector, CRISPR / Cas9-EMX1 vector, and donor vector served as control groups.

[0087] 48 h after transfection, cells from each group were digested with trypsin (Gibco, 25200056) and collected. Genomic DNA was extracted from both groups of cells using a cell genomic DNA extraction kit (TIANGEN, DP304) to detect the efficiency of genome editing.

[0088] Forty-eight hours after transfection, the cells in each group were screened for drug use by adding the drug (Puromycin, 3 μg / mL) to the cell culture medium of both the experimental and positive control groups. Forty-eight hours after drug screening, cells from each group were collected. Genomic DNA was extracted from both groups of cells for genome editing efficiency assay.

[0089] Cells from each group were directly digested and collected 48 h after transfection, and cells from each group obtained after drug screening were used to detect gene editing efficiency by either enzyme digestion or DeepSeq deep sequencing.

[0090] Cells were collected 48 h after transfection using two different methods: direct digestion and drug screening. Genomic DNA was extracted from each group of cells. Using the genomic sequence of the target gene submitted to the NCBI website (https: / / www.ncbi.nlm.nih.gov / ) and the donor template, HDR efficiency detection primers were designed. Using the genomic DNA from each group of cells as a template (NCBI, NC_000002.12), the target sequence (2172 bp) was amplified using DNA polymerase (TaKaRa, R045A). Detection was performed by agarose gel electrophoresis, and the genomic amplification products from each group were recovered using a DNA recovery kit. An equal mass (1 μg) of the recovered product from each group was digested and identified using EcoRI (NEB, R0101) restriction endonuclease. The digested products were detected by agarose gel electrophoresis. The DNA electrophoresis bands were analyzed for grayscale using ImageJ (https: / / imagej.net / ij / ) to calculate the efficiency of precise genome editing via HDR repair. HDR efficiency (%) = 100% × sum of grayscale values ​​of the cleaved target band / (grayscale value of the uncleaved target band + sum of grayscale values ​​of the cleaved target band).

[0091] The primer sequences for HDR efficiency detection are as follows:

[0092] EMX1-F (detection) (SEQ ID NO.61): 5'-CACTGCCAGACACAGAATAGG-3' EMX1-R (detection) (SEQ ID NO.62): 5'- GCACGTTGCTCTTTCTTGG -3' Cells were collected 48 h after transfection using a puromycin selection method. Genomic DNA was extracted from each group of cells as templates, and detection primers with different sequence barcodes (see Table 13) were used to amplify the region near the target site. After gel extraction, the amplified cells were sent to DeepSeq sequencing at Rocks Biotechnology Co., Ltd. The sequencing results were analyzed using CRISPResso (http: / / www.crispresso.rocks) online software to determine the HDR precision editing efficiency of each cell group; GraphPadPrism 8.0 was used for statistical analysis. If HDR precision editing did not occur, the target gene amplified from the cell genomic DNA was inconsistent with the donor template sequence, and the sequencing results showed that the gene sequence was identical to the target genome sequence (WT) or had random mutations. If HDR precision editing occurred, the target gene sequence amplified from the cell genomic DNA was identical to the donor template sequence.

[0093] Table 13 Deepseq sequencing primers

[0094] Enzyme digestion analysis showed that, in cells collected by direct digestion 48 h after transfection, the experimental group using the pCAG-Puro-P2A-KU-SB7 vector (combined with CRISPR / Cas9-EMX1 and pd-EMX1 vectors) achieved 8.8% HDR-accurate editing; the control group using the pCAG-Puro-P2A vector (combined with CRISPR / Cas9-EMX1 and donor vectors) achieved 2.8% HDR-accurate editing (see [link to relevant documentation]). Figure 18 That is, by expressing the KU-SB7 micropeptide, the efficiency of HDR-mediated precision editing can be improved by 3.14 times under conditions where no drug screening cells are used.

[0095] In cells collected 48 h after transfection and subsequent drug screening, the experimental group using the pCAG-Puro-P2A-KU-SB7 vector (combined with CRISPR / Cas9-EMX1 and pd-EMX1 vectors) achieved 33.2% HDR-accurate editing; the control group using the pCAG-Puro-P2A vector (combined with CRISPR / Cas9-EMX1 and pd-EMX1 vectors) achieved 16.7% HDR-accurate editing (see [link to relevant documentation]). Figure 19 Specifically, expressing the KU-SB7 micropeptide can improve the efficiency of HDR-mediated precision editing by 1.98 times under drug-screened cell conditions.

[0096] Deepeseq analysis showed that, in cells collected 48 h after transfection via drug screening, the HDR-precise editing efficiency of cells using the pCAG-Puro-P2A-KU-SB7 vector (combined with CRISPR / Cas9-EMX1 and pd-EMX1 vectors) was 18.37%, while the HDR-precise editing efficiency of cells using the pCAG-Puro-P2A vector (combined with CRISPR / Cas9-EMX1 and donor vectors) was 8.80% (see [link to Deepeseq analysis]). Figure 20 Specifically, expressing the micropeptide (KU-SB7) can improve HDR precision editing by 2.09 times in cells screened using flow cytometry.

[0097] The above results indicate that overexpression of the aforementioned KU-SB7 micropeptide can improve the efficiency of CRISPR / Cas9-mediated precise HDR editing in cells.

[0098] In addition, flow cytometry was used to detect the specific HDR reporter vector repair efficiency of the micropeptide KU-SB7 at different gene loci (AAVS1, CCR5, and NUDT5) in HeLa and U20S cells. Vector construction, cell culture, transfection, and detection methods were the same as described above. Flow cytometry results showed that the micropeptide KU-SB7 could improve HDR repair efficiency at different gene loci in various cell types (see [link to relevant documentation]). Figure 21 ).

[0099] In addition, immunofluorescence (IF) staining was used to detect the subcellular localization of the KU-SB7 micropeptide relative to the Ku70 protein. The specific procedures were as follows: HeLa cells were seeded 12 h in advance in 6-well culture plates (Corning, 3335). The HA-KU-SB7 vector and the 3×FLAG-Ku70 vector were transfected into the cells using a cell transfection reagent (jetPRIME, 101000046). The vector backbones were pcDNA3.1-HA (addgene, 128034) and pcDNA3.1-Flag (addgene, 208051), respectively, and the sequence synthesis and vector construction methods were the same as described above. After culturing for another 48 h, the cells were washed with PBS (Procell, PB180327) and fixed with 4% neutral formaldehyde fixative (Solarbio, P1110) at room temperature for 15 min; the fixative was removed, and the cells were washed three times with PBS. Permeabilize with 0.2% Triton X-100 (Solarbio, T8205) at room temperature for 5 min; remove the permeabilization buffer and wash 3 times with PBS. Block with blocking buffer (Beyotime, P0260) for 60 min; remove the blocking buffer and incubate overnight at 4°C with HA-tagged monoclonal antibody (Invitrogen, 26183) and DYKDDDDK-tagged monoclonal antibody (Invitrogen, MA1-91878). Remove the primary antibody and wash 3 times with PBS. Add fluorescent secondary antibody (Beyotime, A0516, A0521) and incubate at room temperature in the dark for 40 min; remove the secondary antibody and wash 3 times with PBS. Add DAPI (Invitrogen, D1306) and incubate in the dark for 3 min, then wash 3 times with PBS. Fluorescence microscopy revealed that the KU-SB7 micropeptide was primarily located in the cytoplasm, while the Ku70 protein was mainly located in the nucleus. Furthermore, the expression of the KU-SB7 micropeptide and the Ku70 protein showed a strong correlation (see [link to relevant documentation]). Figure 22 ).

[0100] In addition, the effect of HeLa cell transfection with the KU-SB7 micropeptide on cell proliferation during CRISPR / Cas9-mediated gene editing was detected using CCK8 assay. The specific procedures were as follows: HeLa cells were seeded 12 h in advance in 96-well plates (Corning, 3599) with a cell suspension volume of 100 μL / well. Using cell transfection reagent (jetPRIME, 101000046), pCAG-Puro-P2A-KU-SB7 or pCAG-Puro-P2A vectors were transfected into HeLa cells along with the CRISPR / Cas9-EMX1 vector, respectively. Cells without transfection plasmids served as a blank control (Mock). After culturing for 24 h, 10 μl of CCK8 solution (MCE, HY-K0301) was added to each well, and the cells were incubated for another 2 h. The absorbance at 450 nm was measured using a microplate reader. The results showed that transfection with the KU-SB7 micropeptide had no significant effect on cell proliferation during CRISPR / Cas9-mediated gene editing (see [link to study]). Figure 23 ).

[0101] In addition, the effect of HeLa cell transfection with KU-SB7 micropeptide on apoptosis during CRISPR / Cas9-mediated gene editing was detected using the Annexin V-FITC / PI double staining method. The specific procedures were as follows: HeLa cells were seeded 12 h in advance in 24-well plates (Corning, 3524) with a cell suspension volume of 200 μL / well; the cell transfection and grouping methods were the same as those described for the CCK8 assay. Twenty-four h after transfection, apoptosis was detected using the Annexin V-FITC / PI double staining kit (Invitrogen, V13241). Cells were suspended in 1× Annexin V binding buffer, and Annexin V-Alexa Fluor 488 and PI were added and gently mixed before incubation at room temperature in the dark for 20 min. After resuspending the cells, they were immediately analyzed using a Beckman CytoFLEX LX flow cytometer. The results showed that transfection with the KU-SB7 micropeptide had no significant effect on apoptosis during CRISPR / Cas9-mediated gene editing (see [link to study]). Figure 24 ).

[0102] In summary, this invention provides a micropeptide generated and screened with the assistance of artificial intelligence, which can specifically promote the intracellular CRISPR / Cas9-mediated HDR pathway, thereby significantly improving the efficiency of precise gene editing. Furthermore, the micropeptide provided by this invention has no significant effect on cell proliferation and apoptosis during CRISPR / Cas9-mediated gene editing, demonstrating safety. This invention effectively solves the technical bottleneck of low gene editing efficiency in existing technologies, providing a safe, efficient, and promising new strategy for promoting precise gene editing technology in gene function research, animal genetics and breeding, and human disease treatment.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A micropeptide that promotes HDR-mediated precise gene editing, characterized in that, The amino acid sequence of the micropeptide is selected from any of the following: (1) The amino acid sequence as shown in SEQ ID NO.1; (2) The amino acid sequence shown in SEQ ID NO.1 is obtained by substitution, insertion or deletion of one or more amino acids, and still has the function of promoting HDR-mediated precise gene editing.

2. A reagent for improving the efficiency of HDR-mediated precise gene editing, characterized in that, Includes the micropeptide as described in claim 1.

3. The reagent according to claim 2, characterized in that, The reagents also include one or more of the following: CRISPR / Cas9 expression vectors, donor plasmid vectors, or HDR reporter vectors.

4. A kit for improving the efficiency of HDR-mediated precise gene editing, characterized in that, The kit includes the reagents described in claim 2 or 3.

5. A method for improving the efficiency of HDR-mediated precise gene editing, characterized in that, The method includes: co-transfecting the micropeptide of claim 1 with other vectors into target cells to achieve precise gene editing through the HDR pathway; the other vectors include CRISPR / Cas9 expression vectors, donor plasmid vectors and / or HDR reporter vectors.

6. The application of the micropeptide of claim 1 in improving the efficiency of HDR-mediated precise gene editing.

7. The use of the micropeptide according to claim 1 in the preparation of gene-editing drugs, wherein, Miniature peptides inhibit the NHEJ repair pathway and promote HDR-mediated gene editing.

8. A gene-editing drug, characterized in that, Gene-editing drugs include the micropeptides of claim 1 or 2 or the reagents of claim 2 or 3.