3c-like protease of plant viral origin and uses thereof
Patent Information
- Application Number
- CN202611299377.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]然而,目前针对MMLRaV 3C样蛋白酶的研究存在关键的技术瓶颈:该蛋白酶的完整成熟形式氨基酸序列尚未被确定,其N端与C端自切割位点的精确位置及其底物识别特异性均属未知
本发明首次鉴定得到桑花叶卷叶病相关病毒来源的成熟3C样蛋白酶完整氨基酸序列,明确了其N端与C端的自切割位点及Q/G的核心底物识别特异性,填补了该植物病毒蛋白酶的研究空白,既为桑花叶卷叶病的抗病毒防控靶点开发提供了分子基础,也为该蛋白酶的工具化改造提供了序列依据;通过密码子偏好性优化结合无细胞原核表达体系,显著提升了该自剪切型蛋白酶的可溶性表达水平与催化活性,突破了常规活细胞原核表达体系中该类蛋白酶自剪切效率低、活性蛋白产量不足的技术瓶颈;同时该蛋白酶为植物病毒来源,天然无人畜共患病风险,相较于现有动物源、人源3C样蛋白酶工具酶具备更高的生物安全性,可适配食品级、药用级重组蛋白制备等高安全要求的应用场景。
Smart Images

Figure CN122811160A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of molecular biology and virology, specifically to a 3C-like protease derived from a plant virus and its applications. Background Technology
[0002] 3C-like proteases are core functional enzymes encoded by positive-sense RNA viruses. They are responsible for specifically cleaving viral polyprotein precursors into multiple mature functional proteins, essential components for the viral replication cycle. Due to their high substrate recognition specificity, irreplaceable role in viral replication, and highly conserved sequence, 3C-like proteases are not only classic targets for antiviral drug development but also important tool enzymes in the biotechnology field. They can be applied to various biomedical and biotechnology scenarios, including tumor therapy, recombinant protein tag removal, and protein peptide preparation, demonstrating broad application value.
[0003] Mulberry is an important economic forest tree widely planted in my country, possessing ecological, economic, and medicinal value. Currently, as many as 16 viruses that can infect mulberry trees have been publicly reported. Among them, Mulberry mosaic leaf roll associated virus (MMLRaV) belongs to the genus MMLRV, which is a nematode-transmitted polyhedron. Its genome is a dichotomous linear positive-sense single-stranded RNA. Its RNA1 encodes a replication-associated polyprotein, which contains a 3C-like protease domain and is a key functional protein for viral replication and pathogenesis. As a protease derived from plant viruses, it naturally poses no zoonotic risk and has the potential to be developed into a safe tool enzyme.
[0004] However, current research on the MMLRaV 3C-like protease faces key technical bottlenecks: the complete mature amino acid sequence of this protease has not been determined, and the precise locations of its N-terminal and C-terminal self-cleavage sites and their substrate recognition specificity are unknown. Due to the unclear self-cleavage sites, on the one hand, the mature catalytic activity of this protease cannot be accurately defined, making it difficult to construct reliable enzyme activity screening models, severely limiting its credibility and operability as a molecular target for antiviral drugs; on the other hand, the unclear substrate recognition pattern prevents its rational design and application as a site-specific tool enzyme, greatly restricting the engineering development and safe application of this plant-derived protease in biotechnological scenarios such as the precise processing of recombinant proteins. Summary of the Invention
[0005] To address the current technical bottlenecks in research on MMLRaV 3C-like proteases, this invention provides a 3C-like protease derived from a plant virus and its applications.
[0006] The present invention adopts the following technical solution: This invention provides a 3C-like protease derived from a mulberry leaf curl virus, the amino acid sequence of which is shown in SEQ ID NO.8; the N-terminal and C-terminal self-cleavage sites of the 3C-like protease are both glutamine-glycine dipeptides.
[0007] Based on field samples infected with the disease, this invention screens dominant virus strains through polyclonal sequencing and assembly, and for the first time obtains the complete functional coding sequence of a 3C-like protease. A differentiated dual recombinant vector strategy is designed to perform affinity enrichment and Edman degradation sequencing on the N-terminal and C-terminal cleavage products, respectively, to accurately identify the self-cleavage site and clarify its Q / G core substrate recognition specificity. Finally, a plant-derived 3C-like protease with both high catalytic activity and natural biosafety is obtained, providing a molecular target for the control of mulberry virus diseases and a safe tool enzyme for biotechnology scenarios such as recombinant protein processing.
[0008] Furthermore, the 3C-like protease was prepared using a cell-free prokaryotic expression system, and the preparation process included the following steps: Using pET-30a plasmid carrying a 3C-like protease fusion coding sequence as a template, a linear expression template with a T7 promoter and a T7 terminator was obtained by high-fidelity PCR amplification; the 3C-like protease fusion coding sequence is shown in SEQ ID NO.6; The linear expression template was added to the E. coli cell-free expression system and incubated at 20℃~30℃ for 14h~18h to complete the expression of soluble protein. The supernatant of the expression product was collected and purified by imidazole gradient elution using Ni ion affinity chromatography to obtain a 3C-like protease with self-cleavage catalytic activity.
[0009] Furthermore, the N-terminal self-cleavage site and the C-terminal self-cleavage site are identified by a combination of recombinant vectors, the combination of recombinant vectors including a first recombinant vector and a second recombinant vector; The first recombinant vector contains a fusion protein coding sequence, wherein the fusion protein consists of a viral natural leader peptide, a 3C-like protease, a flexible linker peptide, and an affinity tag from the N-terminus to the C-terminus, for identifying the N-terminal self-cleavage site of the protease. The second recombinant vector contains a fusion protein coding sequence, wherein the fusion protein consists of a 3C-like protease, a C-terminal extension peptide, a flexible linker peptide, and an affinity tag from the N-terminus to the C-terminus, for identifying the C-terminal self-cleavage site of the protease.
[0010] Furthermore, the amino acid sequence of the flexible linker peptide is shown in SEQ ID NO.3.
[0011] Furthermore, the affinity tag is composed of a polypeptide affinity tag and a histidine tag tandemly, and the amino acid sequence of the polypeptide affinity tag is shown in SEQ ID NO.4.
[0012] This invention also provides the application of the aforementioned 3C-like protease as a site-specific protein cleavage tool enzyme.
[0013] Furthermore, the specific applications include: removing affinity tags in recombinant protein preparation, achieving independent release of multiple proteins in a polycistronic co-expression system, or performing in-situ cleavage processing of fusion proteins in a cell-free protein synthesis system.
[0014] This invention also provides the application of the aforementioned 3C-like protease as a drug target in screening agents against mulberry leaf curl-related viruses.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention provides the first complete amino acid sequence of a mature 3C-like protease derived from a virus associated with mulberry leaf curl disease, clarifying its N-terminal and C-terminal self-cleavage sites and Q / G core substrate recognition specificity. This fills a research gap in plant viral proteases, providing a molecular basis for the development of antiviral control targets for mulberry leaf curl disease and a sequence basis for the tool-like modification of this protease. By optimizing codon preference and combining it with a cell-free prokaryotic expression system, the soluble expression level and catalytic activity of this self-cleaving protease were significantly improved, overcoming the technical bottleneck of low self-cleavage efficiency and insufficient yield of active protein in conventional live-cell prokaryotic expression systems. Furthermore, this protease is derived from a plant virus, naturally eliminating zoonotic risks, and possesses higher biosafety compared to existing animal- and human-derived 3C-like protease tools, making it suitable for applications with high safety requirements such as the preparation of food-grade and pharmaceutical-grade recombinant proteins.
[0016] This invention utilizes sequencing and assembling of consistent sequences from 11 cloned samples to screen for the protease-coding sequences of dominant viral strains. This effectively reduces the problems of high sequence heterogeneity and random cloning of functional sequences in plant virus serotypes, ensuring the representativeness and functional stability of the obtained protease sequences. A differential design using dual recombinant vectors was employed, with affinity enrichment and Edman degradation sequencing performed on the N-terminal and C-terminal cleavage products, respectively. The rigorous experimental design and accurate and reliable determination of cleavage sites provide reproducible technical support for the subsequent functional research and engineering applications of this protease.
[0017] Based on its clear substrate cleavage specificity and efficient active expression system, this 3C-like protease can be used as a tool enzyme in multiple biotechnology scenarios such as recombinant protein tag excision, construction of polycistronic co-expression systems, and in-situ processing of cell-free protein synthesis systems. At the same time, it can serve as a core target to support the screening and development of drugs against mulberry leaf curl disease, thus possessing both basic research value and industrial transformation potential. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 Electrophoresis results for double enzyme digestion identification of recombinant plasmids; in the figure, a is pET-30a-4575 and b is pET-30a-0902.
[0020] Figure 2 The results of SDS-PAGE analysis of recombinant plasmid pET-30a-4575 are shown.
[0021] Figure 3 The image shows the SDS-PAGE results of the purified product of recombinant plasmid pET-30a-4575; in the figure, A is the band of the full-length fusion protein that has not undergone autocleavage, and B is the band of the N-terminal autocleavage product.
[0022] Figure 4 The image shows the SDS-PAGE results of the purified product of recombinant plasmid pET-30a-0902; in the figure, 1 is the band of the full-length fusion protein that has not undergone autocleavage, and 2 is the band of the C-terminal autocleavage product.
[0023] Figure 5 This is the N-terminal sequencing result.
[0024] Figure 6 These are the C-terminal sequencing results. Detailed Implementation
[0025] The specific embodiments of the present invention are described in detail below, but it should be understood that the scope of protection of the present invention is not limited to the specific embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Unless otherwise specified, the experimental methods described in the embodiments of the present invention are conventional methods, and the materials and reagents used in the following embodiments are commercially available unless otherwise specified.
[0026] Example 1: A 3C-like protease derived from a plant virus and its application.
[0027] I. Experimental Methods 1. Sequence Acquisition and Optimization 1.1 Sample Collection and RNA Extraction Mulberry leaf samples infected with Mulberry leaf roll associated virus (MMLRaV) were collected and preserved in Hengkou Town, Hanbin District, Ankang City (certificate number 5). After being flash-frozen in liquid nitrogen, total RNA was extracted from the leaves using an RNA extraction kit (Tiangen Biotech, catalog number DP432) and stored at -80℃ for later use.
[0028] 1.2 Reverse Transcription and PCR Amplification Using the extracted total RNA as a template, first-strand cDNA was synthesized by reverse transcription using random hexanucleotide primers (Promega, catalog number C1181) according to the instructions of the M-MLV first-strand cDNA synthesis kit (Promega, catalog number M1701).
[0029] The obtained cDNA was subjected to transcriptome sequencing and sequence assembly to obtain the full-length sequence of MMLRaV genomic RNA1. Homology annotation was used to predict the coding region and flanking sequences of the 3C-like protease from the RNA1-encoded polyprotein, and specific amplification primers were designed based on this. Forward primer F: 5'-AATGAAAGAACACACATTGCAAG-3', SEQ ID NO.1, Reverse primer R: 5'-TCTATGCAAACCAATTCAGGCAC-3', SEQ ID NO.2; PCR amplification system (25 μL): 12.5 μL high-fidelity DNA polymerase, 2.5 μL forward primer F (10 μM), 2.5 μL reverse primer R (10 μM), 5 μL cDNA, and nuclease-free water to a final volume of 25 μL; amplification program: 95℃ pre-denaturation for 3 min; 95℃ denaturation for 30 s, 53℃ annealing for 35 s, 72℃ extension for 150 s, 30 cycles; 72℃ final extension for 8 min.
[0030] 1.3 Cloning and Sequencing After purification by agarose gel electrophoresis, the PCR product was ligated into a blunt-end cloning vector and transformed into *E. coli* DH5α competent cells. Eleven positive clones were selected for Sanger sequencing, and a 2040 bp consensus sequence was obtained. This sequence was translated to the expected amino acid sequence of the 3C-like protease precursor, and the dominant sequence was determined. (In this study, eleven positive clones were selected for Sanger sequencing. Multiple alignments were performed on the original sequences of all clones to correct base mismatches and remove low-frequency mutation sites. A 2040 bp consensus sequence of the viral polyprotein precursor was obtained. This consensus sequence was the one with the highest base frequency among the eleven clones, contained no harmful missense mutations, and completely preserved the 3C-like protease's coding reading frame. This sequence was used as the dominant amino acid sequence for the dominant strain, i.e., the dominant sequence.) Based on the dominant sequence, the start codon Met, the flexible linker peptide (GGGGSGGGGGG, SEQ ID NO.3) and two affinity purification tags (KRRWKKNFIAVSAANRFKKISSSGAL, SEQ ID NO.4; HHHHHHHH) were introduced to construct the complete recombinant protein coding sequence.
[0031] 1.4 Sequence Optimization The codon preference of the recombinant protein coding sequence was optimized using the GenSmart codon optimization tool, adjusting the GC content to 53%, and the optimized DNA sequence was synthesized.
[0032] 2. Carrier Construction 2.1 Vector Construction 1 (for N-terminal cleavage site identification) The optimized target DNA fragment and pET-30a vector plasmid were double-digested with BamHI and NdeI restriction endonucleases (NEB), respectively, and incubated at 37°C for 2 h. The digestion products were purified by gel extraction and ligated overnight at 16°C using T4 DNA ligase (NEB, catalog number M0202), followed by transformation into DH5α competent cells. Single colonies were picked, and the plasmid was extracted. After double digestion (a 954 bp target band was identified as positive), Sanger sequencing was performed to verify the absence of mutations in the open reading frame, yielding the recombinant plasmid pET-30a-4575.
[0033] 2.2 Vector Construction II (for C-terminal cleavage site identification) An extended peptide (GFSFVPKDGVLTDGYLK, SEQ ID NO.5) and a His tag (HHHHHHHHH) were added downstream of the predicted C-terminal cleavage site of the 3C-like protease to give the C-terminal cleavage product an affinity tag, facilitating purification and enrichment. Using the same enzyme digestion, ligation, transformation, and validation procedures as the vector construction, the C-terminal extended recombinant plasmid pET-30a-0902 was constructed.
[0034] 3. No prokaryotic expression Cell-free linear templates for expression were prepared by high-fidelity PCR amplification and then added to the E. coli cell-free expression system for protein synthesis.
[0035] High-fidelity PCR amplification: Using the recombinant plasmid verified by sequencing as a template, amplification was performed using a universal T7 promoter forward primer and a T7 terminator reverse primer. The product contained the T7 promoter, the target gene open reading frame, and the T7 terminator. Reaction system (50 μL): Z1 high-fidelity DNA polymerase (Hangzhou Jin'ao, G-POL-001) 2 μL, 10× reaction buffer 5 μL, 50× dNTP 1 μL, forward / reverse primers (10 μM) 2 μL each, recombinant plasmid (50 ng / μL) 1 μL, nuclease-free water to a final volume of 50 μL. Reaction program: 95℃ pre-denaturation for 3 min; 95℃ denaturation for 30 s, 55℃ annealing for 30 s, 72℃ extension for 90 s, 30 cycles; final extension at 72℃ for 5 min. The product, after being verified by electrophoresis, was directly used for expression.
[0036] Cell-free expression system: Prepare a 50 μL reaction system according to the instructions of Hangzhou Xunyao CFU-EC-1000D kit. The specific components are: 20 μL of E. coli extract, 20 μL of reaction buffer, and 10 μL of high-fidelity PCR product (500 ng / μL).
[0037] To optimize protein solubility, a gradient expression process screening experiment was conducted using recombinant plasmid pET-30a-4575 as a vector; recombinant plasmid pET-30a-0902 was directly expressed and prepared using the optimal process determined through screening with pET-30a-4575. The expression conditions for each group are as follows:
[0038] Optimized condition 1: Cell-free system, incubated at 37°C for 16 h; Optimized condition 2: Cell-free system, incubated at 33℃ for 16 h; Optimized condition 3: Cell-free system, incubated at 29°C for 16 h; Optimized condition 4: Cell-free system, incubated at 25°C for 16 h; Optimized condition 5: In vivo expression of Escherichia coli BL21(DE3), induced at 25℃ for 16 h (control).
[0039] After the reaction was completed, a portion of the reaction solution was taken as the total protein sample; the remaining sample was centrifuged at 12,000 rpm for 10 min, and the supernatant was taken as the soluble protein sample for subsequent detection.
[0040] 4. Protein purification and cleavage activity verification The target protein was purified by Ni ion affinity chromatography using a His Trap HP affinity chromatography column (GE Healthcare, 1 mL column volume). Five column volumes were equilibrated with equilibration buffer (20 mM Tris-HCl, 500 mM NaCl, 20 mM imidazole, pH 8.0) at a flow rate of 1 mL / min. The supernatant of the cell-free expression product (soluble protein sample) was loaded onto the column at a flow rate of 0.5 mL / min. After loading, the sample was washed with equilibration buffer until the baseline stabilized. Gradient elution was then performed with five column volumes of 20 mM, 50 mM, and 300 mM imidazole elution buffer (20 mM Tris-HCl, 500 mM NaCl, corresponding concentrations of imidazole, pH 8.0), and the fractions of each elution peak were collected.
[0041] SDS-PAGE was used to detect protein expression and purification effects: 10 μL of protein sample was loaded into each well (mixed with an equal volume of 2×SDS loading buffer and denatured in a boiling water bath for 5 min), followed by electrophoresis on a 5% stacking gel at 80 V for 30 min, and then on a 12% separating gel at 120 V for 60 min. Coomassie Brilliant Blue R-250 staining was performed for 2 h, followed by destaining with a destaining solution (methanol: glacial acetic acid: water = 4:1:5, v:v:v) until the bands were clear. The presence of specific bands with molecular weights smaller than the full-length protein was observed to determine whether the 3C-like protease underwent autocleavage and its activity level.
[0042] 5. Identification of the cutting site 5.1 Identification of N-terminal cleavage site The cleaved protein fraction eluted with 300 mM imidazole was separated by SDS-PAGE and electroblotting onto a PVDF membrane. Coomassie brilliant blue staining was used to locate the target band, and the corresponding band was excised. The sequence of the first 5 amino acids at the N-terminus of the protein was determined using a Shimadzu PPSQ-53A fully automated protein and peptide sequencer based on the Edman degradation method, and the N-terminal cleavage site was deduced after alignment with the theoretical sequence.
[0043] 5.2 Identification of C-terminal cleavage site The C-terminal cleavage site was identified using indirect N-terminal sequencing. Cell-free expression products of recombinant plasmid pET-30a-0902 were purified by Ni affinity chromatography to enrich C-terminal cleavage products with His tags. After separation by SDS-PAGE and transfer to a membrane, specific bands with smaller molecular weights were excised. The N-terminal amino acid sequence of the C-terminal cleavage products was determined using the same Edman degradation method as N-terminal sequencing, and the C-terminal cleavage site of the 3C-like protease was deduced by reverse alignment.
[0044] II. Test Results 1. Sequence acquisition results Eleven clones were sequenced and assembled to obtain eleven 2040 bp nucleotide sequences encoding 3C-like protease precursors, which were then translated to their corresponding amino acid sequences. Preferred sequences were screened, and their codons were optimized, resulting in a DNA sequence GC content of 53% (SEQ ID NO. 6), making it more suitable for the *E. coli* expression system.
[0045] The full-length theoretical amino acid sequence of the recombinant protein (SEQ ID NO.7, including the leader peptide, protease domain, linker peptide, and tag) is as follows: MVGVLHSMDIQGASNASSSSYYDSTRGHNNRVNHKHMHLQGGNPTLPCHGASLSLYGPNGFFCPATWARGRSFWITRHQAFAIPDRASMALIMSDGTRVTFLWEASRLHEYAESEICRYFSPAIPPLRESLARKWYLNDYEKHINMTTCDIYGVTVRKT GKYYEDREIQWWKVPGSIVYKQLQIDDAYMGGTYVHYVPKYIHYSAQTQLHDCGAYVCALIAGDWRIIGFHISRKNGQCAATLIPDVIEQDVQGFSFVPKDGVLTDGYLKGGGGSGGGGGGKRRWKKNFIAVSAANRFKKISSSGALHHHHHHHH, SEQ ID NO.7.
[0046] SEQ ID NO. 6, ATGGTAGGAGTTCTACACTCAATGGATATACAAGGTGCGAGCAATGCGTCTTCCTCCTCCTATTACGACAGCACCCGTGGTCATAACAACCGCGTGAACCACAAACACATGCATCTGCAAGGTGGTAATCCGACCTTACCGTGTCATGGTGCGTCCCTGAGCCTGTACGGCCCTAATGGCTTCTTCTGCCCGGCGACCTGGGCGCGTGGTCGTTCGTTCTGGATAACCCGCCATCAGGCATTTGCAATTCCGGACCGTGCATCTATGGCATTGATCATGTCTGATGGCACGCGTGTTACGTTTCTGTGGGAAGCATCACGTTTGCACGAGTATGCTGAAAGCGAAATCTGCCGTTATTTCTCCCCGGCGATCCCACCGCTGCGTGAGAGCCTTGCCCGCAAGTGGTACCTGAATGATTATGAGAAACACATCAACATGACCACGTGCGACATTTATGGCGTCACTGTACGCAAGACCGGTAAATACTACGAGGACCGTGAGATCCAATGGTGGAAAGTTCCGGGTAGCATTGTGTATAAACAGCTGCAAATTGATGATGCCTACATGGGTGGAACCTATGTTCATTACGTGCCGAAGTACATTCACTACAGCGCGCAGACCCAGCTGCATGATTGCGGTGCGTACGTCTGTGCCTTGATCGCGGGTGACTGGAGAATTATCGGCTTTCATATCAGCCGGAAGAACGGCCAGTGCGCGGCTACTCTGATTCCGGACGTTATTGAACAGGACGTGCAAGGCTTCTCGTTCGTGCCGAAGGACGGCGTTCTGACCGATGGTTATCTGAAAGGTGGCGGTGGCAGCGGTGGCGGCGGCGGGGGCAAGCGCCGTTGGAAGAAAAACTTTATCGCCGTTAGCGCTGCGAACCGTTTTAAAAAGATCAGCAGCAGTGGTGCTTTGCACCACCACCATCACCACCACCACTAA.
[0047] 2. Vector construction and expression results 2.1 Results of double enzyme digestion identification The results of double enzyme digestion and electrophoresis identification of recombinant plasmid pET-30a-4575 are shown below. Figure 1 The descriptions for each lane are as follows: Lane 1: Undigested pET-30a empty vector. The native plasmid is predominantly in supercoil form, exhibiting the fastest electrophoretic migration rate and appearing as the main band; a small number of open-loop, relaxed plasmids migrate more slowly and are located above the main band.
[0048] Lane 2: Recombinant plasmid after double digestion with BamHI / NdeI. Digestion produced two clear bands: the upper band is the linearized pET-30a vector backbone, and the lower band is the inserted target gene fragment. After digestion, the vector changed from supercoiled to linear, migrating more slowly than the supercoiled empty vector. Therefore, the linear vector band is located above the supercoiled band in lane 1. Combined with the fact that the target fragment size is consistent with expectations, this confirms successful insertion of the target gene into the vector.
[0049] Lane 3: 1 kb DNA Ladder molecular weight standard, including gradient bands of 12000 bp, 8000 bp, 6000 bp, 5000 bp, 4000 bp, 3000 bp, 2500 bp, 2000 bp, 1500 bp, 1000 bp, 750 bp, 500 bp and 250 bp.
[0050] Figure 1 Image b shows the agarose gel electrophoresis diagram of the recombinant plasmid pET-30a-0902 after double enzyme digestion. Lane 1 contains the double enzyme digestion product of pET-30a-0902, and lane 2 contains the molecular weight standard of a 1 kb DNA ladder. After double enzyme digestion, two clear bands are visible. The upper band represents the linearized pET-30a vector backbone, and the lower band corresponds to the 1269 bp target gene insertion fragment. The fragment size is consistent with the theoretical expectation, confirming the successful construction of the recombinant expression vector pET-30a-0902.
[0051] 2.2 Cell-free expression results SDS-PAGE results of recombinant plasmid pET-30a-4575 ( Figure 2 )show: Lanes 2 and 3 are parallel replicates of optimized condition 4; lane 1 is the negative control group, incubated at 37°C without linear PCR template; the negative control and the replicates of condition 4 are arranged adjacent to each other for easy and intuitive band comparison.
[0052] Lanes 4, 6, and 8 (corresponding to total protein samples under basic and optimized conditions 1, 2, and 3) all showed a clear target band at 25 kDa, but no clear band was observed in the corresponding supernatant lanes (3, 5, 7, and 9). This indicates that although the protein was successfully expressed under the conditions of 37℃, 33℃, and 29℃, it mainly existed in the form of insoluble inclusion bodies.
[0053] Under optimized condition 4 (cell-free expression at 25℃), both the total protein and the supernatant samples (lanes 11 and 12) showed clear bands at 25 kDa, which was considered the optimal condition. The anti-His tag Western blot was positive, indicating that soluble target protein could be obtained under 25℃ conditions.
[0054] In optimized condition 5 (in vivo expression of E. coli), only total protein showed a band, and there was no obvious signal in the supernatant, indicating that the in vivo expression activity and solubility were much lower than those in the cell-free system.
[0055] Higher temperatures lead to faster protein synthesis in cell-free systems, making polypeptide chains more prone to misfolding and aggregation to form inclusion bodies. At 25°C, the synthesis rate slows down, which is conducive to correct polypeptide chain folding, thus significantly increasing the yield of soluble proteins. Compared to in vivo expression, cell-free systems avoid the toxicity of proteases to host cells and intracellular degradation, making them more suitable for the active expression of this self-cleaving protease.
[0056] The theoretical molecular weight of the full-length uncut recombinant protein is 28 kDa. The detected 25 kDa band represents the mature protein after self-cleavage, and the molecular weight difference is consistent with the theoretical expectation of N-terminal leader peptide cleavage.
[0057] 3. Purification and self-cleavage activity verification Soluble expression products of two recombinant vectors, pET-30a-4575 and pET-30a-0902, were purified using His Trap HP affinity chromatography columns. The target protein was enriched by gradient imidazole elution, and the occurrence of autocleavage was detected by SDS-PAGE to verify protease activity.
[0058] 3.1 Purification results of recombinant plasmid pET-30a-4575 (for N-terminal cleavage site identification) The fusion protein expressed by the recombinant plasmid pET-30a-4575 carries the native viral leader sequence at the N-terminus and is fused with a dual-parenteral tag at the C-terminus. After elution with imidazole gradients of 20 mM, 50 mM, and 300 mM, SDS-PAGE results were obtained. Figure 3 )show:
[0059] The flow-through and low-concentration imidazole elution fractions mainly consisted of miscellaneous proteins; the 300 mM imidazole elution fraction contained two specific main bands: the high molecular weight band (band A) was a full-length fusion protein that had not undergone autocleavage, containing the complete N-terminal leader peptide, 3C-like protease domain, and C-terminal affinity tag; the low molecular weight band (band B) was a mature protease that had undergone N-terminal autocleavage, in which the N-terminal leader peptide was specifically cleaved, but the C-terminal His tag was completely preserved, and therefore it could be captured by the nickel affinity column.
[0060] The molecular weight difference between the two bands is consistent with the theoretical size of the N-terminal leader peptide, proving that this 3C-like protease can complete N-terminal autocleavage in a cell-free system and possesses catalytic activity. The low molecular weight band has purity sufficient for sequencing and will be used for subsequent identification of the N-terminal cleavage site.
[0061] 3.2 Purification results of recombinant plasmid pET-30a-0902 (for C-terminal cleavage site identification) The recombinant plasmid pET-30a-0902 was specifically constructed for C-terminal cleavage site identification, retaining a longer viral native flanking sequence downstream of the 3C-like protease coding region and fusing a His tag at the end. Following the same purification process, after elution with 20 mM, 50 mM, and 300 mM imidazole gradients, the SDS-PAGE results were obtained (…). Figure 4 )show:
[0062] The flow-through fraction contained a small amount of unbound contaminating proteins; a small amount of the target protein was eluted in the 50 mM imidazole elution buffer; two clear specific bands (1, 2) appeared in the 300 mM imidazole elution fraction: the high molecular weight band (1) was the full-length fusion protein that had not undergone autocleavage, containing the complete protease domain, C-terminal flanking extension sequence and terminal His tag; the low molecular weight band (2) was the C-terminal autocleavage product, which was generated by the C-terminal site of the protease and had a complete His tag at the C-terminus, so it could be purified by nickel column enrichment.
[0063] The molecular weight difference between the two bands matches the theoretical size of the C-terminal flanking extension sequence, demonstrating that this 3C-like protease can simultaneously undergo C-terminal autocleavage, further validating its autocatalytic activation function. The low molecular weight band was used for subsequent sequencing identification of the C-terminal cleavage site.
[0064] 4. Identification results of the cutting site 4.1 N-terminal cleavage site The purified B protein band was sequenced from its N-terminus, and the results are shown in the figure. Figure 5 .
[0065] Blank is a blank control chromatogram, obtained by detecting a blank reaction of a protein-free sample using the same procedure. It is used to subtract background peaks from reagents and the system to ensure the accuracy of the results. PTH-AA: Phenylthiohydantoin-amino acid, is a stable derivative generated from the coupling, cleavage, and transformation of N-terminal amino acids during Edman degradation. Different amino acid PTH derivatives have different retention times in reversed-phase liquid chromatography; matching the retention time with that of the standard allows for the identification of the amino acid type.
[0066] The chromatographic peaks of the five sequencing cycles were analyzed using PPSQ Postrun software. By comparing the retention times with those of the standard, the N-terminal first five amino acid sequences were determined to be Gly-Gly-Asn-Pro-Thr (GGNPT).
[0067] The sequence was compared with the theoretical amino acid sequence and matched to the "HMHLQGGNPT" segment. This determined that the N-terminal cleavage site was between glutamine (Q) and glycine (G), i.e., the cleavage pattern was: HMHLQ↓GGNPT.
[0068] 4.2 C-terminal cleavage site N-terminal sequencing was performed on the C-terminal cleavage product expressed by recombinant plasmid pET-30a-0902. Figure 6 The sequencing principle and peak composition are consistent with N-terminal sequencing: Blank is the system blank control, and PTH-AA is the detection peak for amino acid derivatives. Software analysis determined that the first 5 amino acids at the N-terminus of this product are Gly-Phe-Ser-Phe-Val (GFSFV).
[0069] The sequence was compared with the theoretical sequence and matched to the “QDVQGFSFV” segment. This determined that the C-terminal cleavage site was also between glutamine (Q) and glycine (G), and the cleavage pattern was: QDVQ↓GFSFV.
[0070] 5. Complete sequence and clear cleavage site Based on the N-terminal and C-terminal cleavage sites, the complete amino acid sequence of the mature 3C-like protease derived from MMLRaV was determined as follows (218 amino acids in total): SEQ ID NO.8, GGNPTLPCHGASLSLYGPNGFFCPATWARGRSFWITRHQAFAIPDRASMALIMSDGTRVTFLWEASRLHEYAESEICRYFSPAIPPLRESLARKWYLNDYEK HINMTTCDIYGVTVRKTGKYYEDREIQWWKVPGSIVYKQLQIDDAYMGGTYVHYVPKYIHYSAQTQLHDCGAYVCALIAGDWRIIGFHISRKNGQCAATLIPDVIEQ.
[0071] The core cleavage site recognized by this protease is the Q / G dipeptide, which exhibits strict substrate specificity.
[0072] This invention clarifies for the first time the complete sequence and cleavage site of the 3C-like protease of the mulberry leaf curl virus. Combined with its plant virus origin and cell-free high activity, it can be applied to the following scenarios: Recombinant protein tag-cleavage enzyme: A Q / G cleavage site is introduced between the recombinant protein and the affinity tag, and the tag is specifically cleaved using this protease. Compared with animal-derived 3C proteases, it poses no zoonotic disease risk, has higher safety, and can be used for the preparation of food-grade and pharmaceutical-grade recombinant proteins;
[0073] Multicistronic expression system elements: Utilizing their self-cleavage properties, a multi-protein co-expression system can be constructed to achieve equal and independent release of multiple target proteins, which is suitable for fields such as plant metabolic engineering and compound biological agent research and development; Screening targets for antiviral drugs against plants: This protease is a key functional enzyme for MMLRaV replication. The identification of the cleavage site provides a molecular basis for constructing an enzyme activity screening system and screening safe agents against mulberry mosaic and leaf curl disease. Cell-free protein synthesis tool enzyme: This protease exhibits significantly better activity in cell-free systems than when expressed in vivo, and can serve as an in situ tool enzyme for cell-free synthesis platforms, directly cleaving fusion proteins in the reaction system to improve the efficiency of functional protein preparation.
[0074] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments.
[0075] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A 3C-like protease derived from a virus associated with mulberry leaf curl disease, characterized in that, The amino acid sequence of the 3C-like protease is shown in SEQ ID NO.8; both the N-terminal and C-terminal self-cleavage sites of the 3C-like protease are glutamine-glycine dipeptides.
2. The 3C-like protease according to claim 1, characterized in that, The 3C-like protease was prepared using a cell-free prokaryotic expression system, and the preparation process included the following steps: Using pET-30a plasmid carrying a 3C-like protease fusion coding sequence as a template, a linear expression template with a T7 promoter and a T7 terminator was obtained by high-fidelity PCR amplification; the 3C-like protease fusion coding sequence is shown in SEQ ID NO.6; The linear expression template was added to the E. coli cell-free expression system and incubated at 20℃~30℃ for 14 h~18 h to complete the expression of soluble protein. The supernatant of the expression product was collected and purified by imidazole gradient elution using Ni ion affinity chromatography to obtain a 3C-like protease with self-cleavage catalytic activity.
3. The 3C-like protease according to claim 1, characterized in that, The N-terminal self-cleavage site and the C-terminal self-cleavage site are identified by a combination of recombinant vectors, which includes a first recombinant vector and a second recombinant vector. The first recombinant vector contains a fusion protein coding sequence, wherein the fusion protein consists of a viral natural leader peptide, a 3C-like protease, a flexible linker peptide, and an affinity tag from the N-terminus to the C-terminus, for identifying the N-terminal self-cleavage site of the protease. The second recombinant vector contains a fusion protein coding sequence, wherein the fusion protein consists of a 3C-like protease, a C-terminal extension peptide, a flexible linker peptide, and an affinity tag from the N-terminus to the C-terminus, for identifying the C-terminal self-cleavage site of the protease.
4. The 3C-like protease according to claim 3, characterized in that, The amino acid sequence of the flexible linker peptide is shown in SEQ ID NO.
3.
5. The 3C-like protease according to claim 3, characterized in that, The affinity tag is composed of a polypeptide affinity tag and a histidine tag tandemly, and the amino acid sequence of the polypeptide affinity tag is shown in SEQ ID NO.
4.
6. The application of the 3C-like protease as described in claim 1 as a site-specific protein cleavage tool enzyme.
7. The application according to claim 6, characterized in that, The specific applications are: removing affinity tags in recombinant protein preparation, achieving independent release of multiple proteins in a polycistronic co-expression system, or performing in-situ cleavage processing of fusion proteins in a cell-free protein synthesis system.
8. The application of the 3C-like protease as described in claim 1 as a drug target in screening agents against mulberry leaf curl virus.