A method for identifying polypeptide antigen epitopes and uses thereof
By designing a fusion protein using an S-tag and a target sequence in the E. coli type I secretion system, we achieved highly efficient extracellular secretory expression of the target protein, solving the problems of low efficiency and cumbersome purification in existing technologies, and improving expression level and purification efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI JIAOTONG UNIV
- Filing Date
- 2022-07-20
- Publication Date
- 2026-05-19
AI Technical Summary
Existing Escherichia coli type I secretion systems have low efficiency in secreting and expressing exogenous proteins, resulting in cumbersome extraction and purification processes for target proteins, as well as problems with endotoxin contamination and intracellular degradation.
By adding an S tag to the N-terminus and a targeting sequence to the C-terminus of the target protein, the fusion protein is directly secreted into the extracellular space using the E. coli type I secretion system, and then cleaved by the protease recognition sequence, achieving efficient extracellular expression and purification.
This method achieves efficient extracellular secretory expression of the target protein, simplifies the extraction and purification process, reduces endotoxin contamination and intracellular degradation, and improves expression level and purification efficiency.
Smart Images

Figure CN115992159B_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to the field of pharmaceutical technology, specifically a method for identifying polypeptide antigenic epitopes and its application. Background Technology
[0002] Escherichia coli is one of the most commonly used host bacteria for expressing exogenous proteins, characterized by its short growth cycle, low culture cost, and suitability for industrial production. However, because it often retains exogenous proteins inside the bacterial cell, the proteins may not fold correctly or may be degraded and lose their activity, making the extraction and purification of the target protein quite cumbersome. Secretory expression is one of the effective methods to solve these problems.
[0003] Escherichia coli has five types of secretion systems, among which type I and type II secretion systems are mainly used for secreting recombinant proteins. Compared with the periplasmic expression of the type II system, the direct efflux mediated by the transmembrane pores of the type I secretion system has advantages. It avoids degradation by periplasmic proteases, does not require cell lysis, and is easier to recover and purify. In addition, the culture medium has sufficient space, which is conducive to the overexpression of exogenous proteins. However, the existing type I secretion systems have very low efficiency in the secretion and expression of exogenous proteins, and related research is limited, far from meeting the needs of practical applications.
[0004] In view of the above, this application is hereby submitted. Summary of the Invention
[0005] One objective of this application is to provide a method for identifying polypeptide antigenic epitopes, wherein the method enables E. coli cells to have a high level of extracellular secretory expression of the target protein by selectively selecting a tag located at the N-terminus and a target sequence located at the C-terminus of the fusion protein, thereby enabling rapid and high-yield production of the target protein.
[0006] Another objective of this application is to provide a fusion protein that, through an S-tag at the N-terminus and a targeting sequence at the C-terminus, allows the target protein to be directly and secreted at a high expression level into the extracellular culture medium of Escherichia coli, without disrupting the E. coli cells, thereby reducing endotoxin contamination and intracellular degradation of the target protein.
[0007] The purpose of this application is not limited to the above-mentioned purposes. Other purposes and advantages of this application not mentioned above can be understood from the following description and will become clearer through the embodiments of this application. Furthermore, it is readily understood that the purposes and advantages of this application can be achieved through the features disclosed in the claims and combinations thereof.
[0008] Specifically, this application provides the following technical solutions:
[0009] In one aspect of the application, this application provides a method for identifying polypeptide antigenic epitopes, comprising:
[0010] The steps of transforming E. coli cells with an expression vector containing nucleic acid encoding the target protein and culturing the E. coli cells;
[0011] The target protein is expressed in the form of a fusion protein, which comprises:
[0012] Target protein;
[0013] The S tag at the end of N; and
[0014] The targeting sequence at the C-terminus enables the fusion protein to target the extracellular domain of E. coli cells.
[0015] In one embodiment, the targeting sequence comprises hemolysin A signal peptide.
[0016] In one embodiment, the fusion protein further comprises one or more tag sequences selected from His tag, GST tag, Strep tag and MBP tag, the tag sequence being located between the target protein and the targeting sequence.
[0017] In one embodiment, there are two tag sequences, located between the target protein and the targeting sequence, and between the S tag and the target protein, respectively.
[0018] In one embodiment, the fusion protein further comprises two protease recognition sequences located at the N-terminus and C-terminus of the target protein, respectively.
[0019] In one embodiment, the protease recognition sequence is selected from at least one of the following: TEV protease recognition sequence, thrombin recognition sequence, enterokinase recognition sequence, 3C protease recognition sequence, and integrin.
[0020] In one embodiment, the method further includes:
[0021] The culture supernatant is collected and purified, and then contacted with a protease to obtain a free target protein, wherein the protease is identified based on the protease recognition sequence.
[0022] In another aspect of this application, a fusion protein suitable for extracellular secretory expression is provided, comprising:
[0023] Target protein;
[0024] The S tag at the end of N; and
[0025] The targeting sequence at the C-terminus enables the fusion protein to target the extracellular domain of E. coli cells.
[0026] In one embodiment, the targeting sequence comprises hemolysin A signal peptide.
[0027] In one embodiment, the tag sequence is further included, which is selected from one or more of His tags, GST tags, Strep tags and MBP tags, and the tag sequence is located between the target protein and the targeting sequence.
[0028] In one embodiment, there are two tag sequences, located between the target protein and the targeting sequence, and between the S tag and the target protein, respectively.
[0029] In one embodiment, it also includes two protease recognition sequences located at the N-terminus and C-terminus of the target protein, respectively.
[0030] In one embodiment, the protease recognition sequence is selected from at least one of the following: TEV protease recognition sequence, thrombin recognition sequence, enterokinase recognition sequence, 3C protease recognition sequence, and integrin.
[0031] In another aspect of the application, this application provides an expression vector comprising a nucleic acid encoding a fusion protein as described above.
[0032] This application also provides the application of the above-mentioned method for identifying multiple antigenic epitopes in target protein activity detection, affinity detection, or mutant screening.
[0033] The technical solutions provided by the embodiments of this application may include the following beneficial effects:
[0034] 1. The method for identifying polypeptide antigenic epitopes provided in this application can achieve efficient extracellular secretory expression of target proteins. It has the advantages of simple process, fast and efficient process and low cost. It has broad application prospects in the fusion expression, activity detection, affinity detection and mutant screening of proteins such as active proteins, polypeptides, antibody epitopes and antigenic epitopes.
[0035] 2. The fusion protein expression vector provided in this application can achieve efficient extracellular secretory expression of the target protein. The target protein can be secreted in large quantities in a short period of time, and the protein can be directly extracted or detected from the culture medium. Attached Figure Description
[0036] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0037] Figure 1 This is a schematic diagram of the structure of a fusion protein according to an embodiment of this application. 1 represents the target protein, 2 represents the S tag, 3 represents the target sequence, 4 represents the tag sequence, and 5 represents the protease recognition sequence.
[0038] Figure 2 To screen for N-terminal tags of the fusion protein in this application, different tag peptides were placed at the N-terminus of the fusion protein, including F: Flag tag, H: His tag; M: Myc tag, S: S tag, and T: T7 tag. W represents whole-cell expressed protein, and S represents culture medium supernatant.
[0039] Figure 3 This is a schematic diagram of a screening scheme for constructing genetically engineered Escherichia coli strains according to an embodiment of this application.
[0040] Figure 4 This is a schematic diagram of an expression scheme for an Escherichia coli genetically engineered strain according to an embodiment of this application.
[0041] Figure 5 This application describes the promoting effect of the S tag on the extracellular secretion expression of three pattern tag peptides in a genetically engineered *E. coli* strain according to one embodiment. The three pattern tag peptides are: F: Flag tag; M: Myc tag; and T: T7 tag. In the SDS-PAGE band sample, W represents whole bacterial protein, and S represents protein from the culture medium.
[0042] Figure 6 This application describes the promoting effect of the S tag on the extracellular secretion expression of three modal antimicrobial peptides in a genetically engineered *E. coli* strain according to one embodiment. The three modal antimicrobial peptides are: C: Cecropin A; P3: PEW300, a basic mutant of Cecropin A; L: LL37; and A12: Aurenin 1.2. In the SDS-PAGE band sample, W represents whole bacterial protein, and S represents protein from the culture medium.
[0043] Figure 7 Flow cytometry plot of the supernatant of the fusion protein expression of RGDS and its mutant RGES peptides binding to U87 MG cells expressing integrin.
[0044] Figure 8 The EC50 values of the medium for IL-15 fusion protein expression added to the Mo7e cell culture system are shown in the figure.
[0045] Figure 9 This is a schematic diagram of a purification scheme for preparing the target protein according to one embodiment of this application.
[0046] Figure 10This is a schematic diagram of the single-chain Fc antibody structure described in this application; wherein, RGDS is a polypeptide targeting integrin, which is fused and expressed with a conventional Fc with dimerizing ability or a single-chain Fc (mFc) that has lost dimerizing ability after mutation, as the target protein. STHRGDSFcTHHly can be digested with TEV and purified with Ni column to obtain the dimer RGDSFc, and STHRGDSmFcTHHly can be digested with TEV and purified with Ni column to obtain the single-chain RGDSmFc.
[0047] Figure 11 SDS-PAGE image of the preparation process of single-chain antibody fusion protein RGDSmFc.
[0048] Figure 12 This is a schematic diagram of antigen epitope splitting.
[0049] Figure 13 This is a schematic diagram illustrating the construction of Escherichia coli genetically engineered strains and the application of antigenic epitope screening in one embodiment of this application.
[0050] Figure 14 The first round of identification results of the expression strain library based on specific fragments described in this application.
[0051] Figure 15 The results of the second round of identification of the expression strain library based on specific fragments described in this application. Detailed Implementation
[0052] The present application will now be described in further detail with reference to the embodiments. It is to be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit the invention.
[0053] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0054] It should be noted that the endpoints and any values of the ranges disclosed herein are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoint values of various ranges, the endpoint values of various ranges and individual point values, and individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein. Where specific techniques or conditions are not specified in the embodiments, they are performed in accordance with the techniques or conditions described in the literature in the art or according to the product instructions. Reagents or instruments used without specified manufacturers are all conventional products that can be obtained commercially.
[0055] It should be noted that, unless otherwise stated, the scientific and technical terms used in this article have the meanings commonly understood by those skilled in the art. Furthermore, the laboratory procedures for cell culture, molecular genetics, nucleic acid chemistry, and immunology used in this article are all standard procedures widely used in their respective fields.
[0056] definition
[0057] Fusion protein: As used herein, "fusion protein" refers to a biologically active polypeptide and effector molecule linked (i.e., fused) by genetic recombination, chemical methods, or other suitable methods. Where desired, fusion molecules may be fused at one or more sites via amino acid linkers. Fusion proteins may exist as monomers or polymers (e.g., dimers).
[0058] Target protein: As used herein, "target protein" refers to a foreign protein intended to be obtained through expression of a member of the *E. coli* family. It can be any protein, including oligopeptides, polypeptides, cytokines, single-domain antibodies, single-chain antibodies, Fab fragments, Fc fragments, antibody epitopes, antigenic epitopes, and recombinant proteins. In one embodiment, the target protein is a cytokine, such as IL-15. In one embodiment, the target protein is an antimicrobial peptide, such as Cecropin A, LL37, or Aurenin 1.2.
[0059] Tag protein: As used herein, “tag protein” refers to a short amino acid sequence (preferably 2-20 amino acids, more preferably 4-20 amino acids) fused to the N- or C-terminus of a target protein.
[0060] Target sequence: As used herein, a "target sequence" refers to an amino acid sequence fused to the N- or C-terminus of a target protein that guides the secretion of the fusion protein into the extracellular domain of a member of the *E. coli* family. This includes, but is not limited to, protein fragments that can be directly secreted extracellularly, outer membrane protein F, or osmolarly induced protein Y. In one embodiment, the target sequence comprises a hemolysin A signal peptide. In another embodiment, the target sequence is a sequence of a fusion protein synthesized and transported extracellularly to *E. coli* via a type I secretion system.
[0061] Protease recognition sequence: As used herein, "protease recognition sequence" refers to an amino acid sequence that is recognized and cleaved by an associated protease, allowing the free target protein to be obtained through cleavage at a specific site by the associated protease. Associated proteases include those with a single specific recognition sequence, which cleaves within or near a specific sequence of one or more amino acids.
[0062] Inteptide: As used in this article, an "inteptide" is an amino acid sequence that exists in a precursor protein. During the process of the precursor protein being converted into a mature protein, it is released from the precursor protein by self-cleaving. Inteptides can be artificially cleaved inteptides, naturally cleaved inteptides, or inteptide mutants with cleaving function. Invention Details
[0064] A method for identifying polypeptide antigenic epitopes includes:
[0065] The steps of transforming E. coli cells with an expression vector containing nucleic acid encoding the target protein and culturing the E. coli cells;
[0066] The target protein is expressed in the form of a fusion protein, which comprises:
[0067] Target protein;
[0068] The S tag at the end of N; and
[0069] The targeting sequence at the C-terminus enables the fusion protein to target the extracellular domain of E. coli cells.
[0070] In this embodiment, the targeting sequence enables the fusion protein to be secreted extracellularly into E. coli cells via the type I secretion system. This application unexpectedly discovered that adding an S tag to the N-terminus of the fusion protein significantly increased its extracellular secretory expression level, indicating that the S tag can help improve the expression efficiency of the E. coli type I secretion system.
[0071] The S tag is a short peptide derived from the N-terminus of pancreatic ribonuclease A (RNase A), and its length and amino acid composition can be adjusted according to the expression level of the fusion protein to obtain optimal expression results. In a preferred embodiment, the amino acid sequence of the S tag is shown in SEQ ID No. 1. In another preferred embodiment, the nucleotide sequence encoding the S tag is shown in SEQ ID No. 2.
[0072] Among them, Escherichia coli cells include any bacteria in the species Escherichia coli, including but not limited to E. coli K12, E. coli DH 5α, E. coli BL21(DE3), E. coli BL21(DE3)pLysS, E. coli Rosetta(DE3), E. coli JM109(DE3), E. coli S17(λpir), E. coli CC118(λpir), etc.
[0073] Furthermore, in one embodiment, the targeting sequence comprises a hemolysin A signal peptide.
[0074] The Escherichia coli α-hemolysin (HlyA) secretion system is a typical type I secretion system, composed of three protein components: HlyB (ABC translocase), HlyD (membrane fusion protein, MFP), and TolC (outer membrane protein OMP). HlyB recognizes the signal peptide sequence located at the C-terminus of HlyA, triggering the assembly of the three proteins into a transport complex, forming a continuous soluble hollow conduit spanning the outer membrane, periplasm, and inner membrane. HlyA is then secreted directly into the extracellular culture medium without passing through a periplasmic intermediate. Therefore, fusing a target protein with this signal peptide sequence can also result in extracellular secretion.
[0075] In this embodiment, when obtaining the target protein, E. coli cells are simultaneously transformed with an auxiliary plasmid vector containing nucleic acids encoding HlyB and HlyD and an expression vector containing nucleic acids encoding the target protein.
[0076] In a preferred embodiment, the amino acid sequence of the hemolysin A signal peptide is shown in SEQ ID No. 3. In another preferred embodiment, the nucleotide sequence encoding the hemolysin A signal peptide is shown in SEQ ID No. 4.
[0077] Furthermore, in one embodiment, the fusion protein further comprises one or more tag sequences selected from His tag, GST tag, Strep tag, and MBP tag, wherein the tag sequence is located between the target protein and the targeting sequence. That is, the fusion protein comprises, from the N-terminus to the C-terminus, an S tag, a target protein, a tag sequence, and a targeting sequence, wherein the tag sequence facilitates the purification of the fusion protein, making post-secretion purification and detection easier.
[0078] Specifically, the tag sequence used in this application is for binding to ion exchange resins, particularly cation exchange resins. The tag sequence should appropriately include charged amino acid residues, such as K, R, H, D, and E. The amino acid composition and length of the tag sequence can be adjusted according to the size, amino acid composition, and charge of the fusion protein to optimize binding to the ion exchange resin. In a preferred embodiment, the tag sequence length can be 4-20 or more amino acids. Preferably, the tag sequence length is between 4-12 amino acids.
[0079] Furthermore, in a preferred embodiment, the tag sequence is a His tag, i.e., containing H n Where n≥2. In another preferred embodiment, the tag sequence is a His6 tag, i.e., the amino acid sequence of the tag sequence is HHHHHH. In a preferred embodiment, the nucleotide sequence encoding the His6 tag is catcaccatcatcaccac.
[0080] Furthermore, in a preferred embodiment, the fusion protein comprises an S-tag, a target protein, a tag sequence, and a targeting sequence connected sequentially from the N-terminus to the C-terminus.
[0081] Furthermore, in one embodiment, there are two tag sequences, located respectively between the target protein and the targeting sequence, and between the S tag and the target protein. Having two tag sequences facilitates the isolation and purification of the target protein.
[0082] Furthermore, in one embodiment, the fusion protein further comprises two protease recognition sequences located at the N-terminus and C-terminus of the target protein, respectively.
[0083] In this embodiment, two protease recognition sequences are located between the target protein and the tag sequence, respectively. After the fusion protein is expressed, the protease recognition sequence can be cleaved at a specific site using a protease associated with the recognition sequence to obtain the free target protein. After cleavage, the C-terminus and N-terminus of the target protein contain no extra residues or only a few residues, significantly improving the quality of the target protein and greatly facilitating the subsequent purification of the product.
[0084] Further, in one embodiment, the protease recognition sequence is selected from at least one of a TEV protease recognition sequence, a thrombin recognition sequence, an enterokinase recognition sequence, a 3C protease recognition sequence, and an integrin, preferably the protease recognition sequence is a TEV protease recognition sequence. In a preferred embodiment, the amino acid sequence of the TEV protease recognition sequence is ENLYFQG. In a preferred embodiment, the nucleotide sequence encoding the TEV protease recognition sequence is gagaacctgtacttccaaggg.
[0085] Furthermore, in one embodiment, the fusion protein comprises, sequentially from the N-terminus to the C-terminus, an S-tag, a tag sequence, a protease recognition sequence, a target protein, a protease recognition sequence, a tag sequence, and a targeting sequence (e.g., ...). Figure 1 As shown in the figure, the fusion protein comprises the structure shown in the following formula:
[0086] S tag - tag sequence - protease recognition sequence - target protein - protease recognition sequence - tag sequence - target sequence, where "-" indicates a covalent bond.
[0087] In a preferred embodiment, the fusion protein provided in this application may optionally contain a linker peptide, i.e., the S tag is linked to the tag sequence, the tag sequence to the protease recognition sequence, the protease recognition sequence to the target protein, the protease recognition sequence to the tag sequence, and / or the tag sequence to the target sequence via an optional linker peptide. The linker peptide may be, for example, G, S, GG, GS, SS, SG, or GGSGG. The size and complexity of the linker peptide may affect the activity of the protein. Generally, the linker peptide should have sufficient length and flexibility to ensure that the two connected parts have sufficient spatial freedom to perform their functions. Simultaneously, the formation of α-helices or β-sheets in the linker peptide should be avoided to prevent the formation of α-helices or β-sheets that could affect the stability of the fusion protein.
[0088] In another preferred embodiment, the parts of the fusion protein provided in this application are directly linked by covalent bonds (e.g., peptide bonds).
[0089] Furthermore, in one embodiment, the method further includes:
[0090] The culture supernatant is collected and purified, and then contacted with a protease to obtain a free target protein, wherein the protease is identified based on the protease recognition sequence.
[0091] In this process, *E. coli* cells are cultured under suitable conditions for fusion protein expression. The fusion protein is secreted and expressed, and can be directly recovered from the cells or cell culture supernatant. Suitable conditions for expressing the fusion protein should be known to those skilled in the art, who can select an appropriate culture medium based on experience and culture the cells under conditions suitable for *E. coli* cell growth. Once the *E. coli* cells have grown to an appropriate cell density, the selected promoter is induced using a suitable method (such as temperature change or chemical induction), and the cells are cultured for a further period. The fusion protein from the above method is then secreted extracellularly.
[0092] The purification of the culture supernatant (i.e., the fusion protein) is well known to those skilled in the art. Examples of these methods include, but are not limited to, treatment with protein precipitants, centrifugation, ultratreatment, ultracentrifugation, molecular sieve chromatography, adsorption chromatography, ion exchange chromatography, high-performance liquid chromatography, and various other liquid chromatography techniques, as well as combinations of these methods.
[0093] In this process, when the culture supernatant comes into contact with the protease under appropriate conditions, the protease cleaves the protease recognition sequence in the fusion protein, thereby obtaining the free target protein. The method of protease cleavage of fusion proteins is well known to those skilled in the art.
[0094] The method further includes steps for separating and purifying the target protein. Methods for separation and purification are well known to those skilled in the art. Examples of these methods include, but are not limited to, treatment with protein precipitants, centrifugation, ultratreatment, ultracentrifugation, molecular sieve chromatography, adsorption chromatography, ion exchange chromatography, high-performance liquid chromatography, and various other liquid chromatography techniques, as well as combinations of these methods.
[0095] Furthermore, in one embodiment, the method includes the steps of purifying the culture supernatant by a first ion exchange chromatography and purifying the free target protein by a second ion exchange chromatography.
[0096] In the process of purifying the culture supernatant using the first ion exchange chromatography, the tag sequence in the fusion protein binds to the chromatography column used in the first ion exchange chromatography. After elution, the purified fusion protein can be obtained. The fusion protein after protease cleavage is further purified using the second ion exchange chromatography. The fragment containing the tag sequence binds to the chromatography column used in the second ion exchange chromatography, while the target protein does not bind and remains in the flow-through. The target protein is obtained by collecting the flow-through.
[0097] This application also provides a fusion protein suitable for extracellular secretory expression, comprising:
[0098] Target protein;
[0099] The S tag at the end of N; and
[0100] The targeting sequence at the C-terminus enables the fusion protein to target the extracellular domain of E. coli cells.
[0101] Furthermore, in one embodiment, the targeting sequence comprises a hemolysin A signal peptide.
[0102] Furthermore, in one embodiment, it also includes a tag sequence selected from any one or more of His tags, GST tags, Strep tags, and MBP tags, the tag sequence being located between the target protein and the targeting sequence.
[0103] Furthermore, in one embodiment, there are two tag sequences, located between the target protein and the targeting sequence, and between the S tag and the target protein, respectively.
[0104] Furthermore, in one embodiment, it also includes two protease recognition sequences located at the N-terminus and C-terminus of the target protein, respectively.
[0105] Furthermore, in one embodiment, the protease recognition sequence is selected from at least one of the following: TEV protease recognition sequence, thrombin recognition sequence, enterokinase recognition sequence, 3C protease recognition sequence, and integrins.
[0106] The fusion protein of this application has the same technical effect as the above-mentioned method for identifying polypeptide antigenic epitopes, which will not be elaborated here.
[0107] This application also provides nucleic acids encoding fusion proteins and corresponding expression vectors.
[0108] The nucleic acid in this application can be a DNA molecule or an RNA molecule, or a nucleic acid analogue. The nucleic acid molecule in this application can contain naturally occurring nucleic acid residues or artificially generated nucleic acid residues. The nucleic acid molecule in this application can be single-stranded or double-stranded, linear or circular, natural or synthetic, and unless otherwise specified, there are no size limitations. The nucleic acid molecule may also contain a promoter, which can be homologous or heterologous.
[0109] The nucleic acid molecules of this application can be cloned into vectors. The term "vector" in this application includes plasmids, granules, viruses, bacteriophages, and other vectors commonly used in genetic engineering. In a preferred embodiment, these vectors are suitable for stable transformation of *E. coli* cells, for example, to transcribe the nucleic acid molecules of this application.
[0110] The vector used in this application can be an expression vector. Suitable prokaryotic expression vectors that have been widely described in the literature can be used in this application. In one embodiment, the expression vector may contain a marker gene and a replication origin, promoter, and transcription termination signal to ensure replication in a selected host. Preferably, between the promoter and the termination signal, there is at least one restriction site capable of inserting the desired nucleic acid sequence / molecule. Preferably, the expression vector used in this application is selected from pET series expression vectors, pGEX series expression vectors, pcDNA series expression vectors, and pMF series expression vectors. More preferably, the expression vector used in this application is the pET-28a expression vector, which has a T7 promoter and a kanamycin resistance selection gene.
[0111] This application also provides a genetically engineered strain of *Escherichia coli* containing the nucleic acid or expression vector described above. Preferably, it is obtained by introducing the expression vector described above into *Escherichia coli* BL21(DE3), which is capable of producing high levels of the target protein via extracellular secretion, thus providing a novel prokaryotic extracellular secretion expression system.
[0112] Furthermore, in one embodiment, the engineered E. coli strain also includes an auxiliary plasmid vector containing nucleic acids encoding HlyB and HlyD. This engineered E. coli strain, as a novel type I extracellular secretion expression system, can efficiently generate the target protein via extracellular secretion.
[0113] This application also provides the application of the above-mentioned method for identifying polypeptide antigenic epitopes in the detection of target protein activity, affinity, or mutant screening.
[0114] For example, in the detection of the activity or affinity of the target protein, the supernatant of E. coli culture containing the fusion protein is collected, and the supernatant sample is added to the detection system for activity and affinity analysis.
[0115] For example, when the target protein is an antibody or an antigenic epitope, the detection of the affinity of the aforementioned E. coli genetically engineered strain for a specific antibody or antigenic epitope can be considered as screening for antibodies or antigenic epitopes, which can be carried out through the following steps:
[0116] 1) Obtain the gene sequence library of antibody or antigen epitope library through methods such as PCR or DNA chemical synthesis, insert the gene fragment into the target protein position to construct recombinant expression plasmid, and obtain plasmid library;
[0117] 2) Simultaneously transform the plasmid library and helper plasmid vector into E. coli host cells, plate them on LB agar plates supplemented with IPTG and arabinose inducer, as well as kanamycin and ampicillin, and incubate overnight to obtain single colonies;
[0118] 3) Take an NC film, mark it and the LB plate, cover the plate with the NC film, and let it stand for 5 minutes;
[0119] 4) Remove the NC membrane and perform subsequent antigen and antibody incubation and color development. Then, map the colored clone sites onto the original plate to obtain a single clone with high-affinity antibody or antigen epitope expression.
[0120] 5) Select the selected single-clone colonies, culture them, and then perform expression plasmid sequencing to obtain the selected antibody sequence or antigen epitope sequence.
[0121] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application.
[0122] Example 1
[0123] Study on the effect of S-tag (S tag, hereinafter referred to as S) on the extracellular expression and secretion efficiency of Escherichia coli expression system
[0124] Previous studies have shown that the HlyA signal peptide (hereinafter referred to as Hly) can achieve extracellular secretory expression, but with low efficiency. To improve the secretory expression efficiency of exogenous proteins or peptides, THHly was used as a model protein for this study, where T stands for TEV restriction site (hereinafter referred to as T), H stands for His, and Hly stands for HlyA signal peptide. Different tags, including Flag, His, Myc, S, and T7 tags, were used at the N-terminus of the fusion protein to compare the expression and secretion efficiency of the fusion protein after induction. The specific steps are as follows:
[0125] 1) Construct expression plasmid: Design primers for conventional PCR or overlap extension PCR, ligate the Flag, His, Myc, S, and T7 tag sequences to the THHly fragment, and add homologous recombination fragments to both ends; use Xba I and Xho I. Digest the vector with restriction endonucleases, recover the vector digestion products and PCR products using the kit and measure the concentration of the recovered products, add each fragment according to the instructions for homologous recombination, and insert the obtained target fragment into the pET-28a plasmid.
[0126] 2) Plasmid transformation: The recombinant product was transformed into competent DH5α strains and plated on kanamycin plates to obtain cloned strains containing the expression plasmid pET-NTHHly (N represents different tags); after confirming the correct sequence, the relevant cloned strains and plasmids were stored at -80 degrees Celsius.
[0127] 3) Construction of helper plasmid vector: Primers were designed for conventional PCR to obtain nucleotide fragments encoding HlyB and HlyD. The amino acid sequence of HlyB is shown in SEQ ID No. 5, the nucleotide fragment encoding HlyB is shown in SEQ ID No. 6, the amino acid sequence of HlyD is shown in SEQ ID No. 7, and the nucleotide fragment encoding HlyD is shown in SEQ ID No. 8. Using pBAD-gIII C as the backbone, the aforementioned HlyB and HlyD were linked together with a ribosome-binding region and inserted downstream of the pBAD promoter to obtain the helper plasmid vector. The nucleotide sequence of the ribosome-binding site between HlyB and HlyD is cagaaagaacagaagaat.
[0128] 4) Construction of expression strain: The expression plasmid and helper plasmid vector were co-transformed into the BL21(DE3) strain, plated on ampicillin and kanamycin double antibody plates, and several transformants were picked for culture, plasmid extraction and enzyme digestion identification to obtain the expression strain containing the transfer plasmid and helper plasmid.
[0129] 5) Detection of fusion protein induction: Take the overnight culture of the strain with correct enzyme digestion, transfer it to a shake flask containing LB medium (Amp 100 μg / mL + Kan 10 μg / mL) at a volume fraction of 1%, culture it at 37°C and 220 rpm for 2 h, then add 0.08% arabinose and induce at 37°C and 220 rpm; after 4 h, add 0.1 mM IPTG and continue to induce at 20°C and 150 rpm for 4 h. Take 1 mL of the culture solution and transfer it to a centrifuge tube, centrifuge at 12000 g * 5 min, transfer the supernatant to a new centrifuge tube, resuspend the pellet with 1 mL of H2O, and prepare samples from the bacterial cells and the supernatant respectively for SDS-PAGE detection.
[0130] The results are as Figure 2 shown. It can be seen that the protein secretion yields after the fusion of each tag are Flag < Myc < T7 < His < S, that is, Stag has the highest expression and secretion level. Since the combination of S tag and HlyA has produced new advantages in the extracellular secretion expression of foreign proteins, it can be used as a new secretion expression system for the rapid expression, detection, screening and preparation of foreign proteins.
[0131] Example 2
[0132] Construction and screening of expression strains for promoting the secretion expression of target proteins by S tag
[0133] 1) Construction of expression plasmid: Design primers for conventional PCR or overlap extension PCR, ligate the SHT, target protein (hereinafter briefly denoted as X) and THHly fragments, and add homologous recombination fragments at both ends; use Xba I and Xho I restriction endonucleases to digest the vector, use a kit to recover the vector digestion product and the PCR product and measure the concentration of the recovered product, add each fragment according to the dosage in the instruction manual for homologous recombination, and insert the obtained target fragment into the pET-28a plasmid;
[0134] 2) Plasmid transformation: Transform the recombinant product into competent DH5α strain, coat the kanamycin plate, and obtain the cloned strain containing the expression plasmid pET-SHTXTHHly; after determining the correct sequence, store the relevant cloned strain and plasmid in a -80°C refrigerator;
[0135] 3) Construction of expression strain: Co-transform the expression plasmid and the auxiliary plasmid vector constructed by the same method as in Example 1 into BL21(DE3) strain, coat the ampicillin and kanamycin double-antibody plate, pick several transformants for culture, plasmid extraction and enzyme digestion identification, and obtain the expression strain containing the transfer plasmid and the auxiliary plasmid ( Figure 3 ); Take the culture solution of the strain with correct enzyme digestion, and freeze it in a -80°C refrigerator using a commercial bacterial cryopreservation solution.
[0136] Example 3
[0137] The method for identifying polypeptide epitopes described in this application is used for the secretory expression of exogenous proteins.
[0138] A specific volume fraction of the overnight culture of the expression strain was transferred to a shake flask for induction. At fixed time points, arabinose and IPTG were added to induce induction. Taking the SHTXTHHly system as an example, the specific operation is as follows: Figure 4 As shown, it includes the following steps:
[0139] 1) Transfer the frozen bacteria containing the expression plasmid and helper plasmid to LB medium (Amp 100 μg / mL + Kan 10 μg / mL) and incubate overnight at 37°C and 220 rpm;
[0140] 2) After culturing for 12 h, take the overnight culture medium and transfer it at a volume fraction of 1% to a shake flask containing LB medium (Amp 100 μg / mL + Kan 10 μg / mL), and incubate at 37 degrees and 220 rpm.
[0141] 3) After culturing for 2 hours, add 0.08% arabinose and induce at 37 degrees Celsius and 220 rpm;
[0142] 4) After 4 h of induction culture, add 0.1 mM IPTG and induce at 20 degrees and 150 rpm;
[0143] 5) After induction culture for 4 h, transfer the culture medium to a centrifuge tube, centrifuge at 12000 g * 5 min, and transfer the supernatant to a new centrifuge tube for subsequent detection and purification.
[0144] Using this culture system and induction expression method, the promoting effect of S tag on HlyA-mediated exogenous protein expression and secretion was verified. SHT tags were added to the N-terminus of the fusion proteins FTHHly, MTHHly, and TTHHly, which had low secretion levels in Example 1. The results are as follows: Figure 5 As shown, the expression and secretion levels of the three fusion proteins were significantly increased, indicating that the newly constructed extracellular soluble expression and purification system SHTXTHHly has good production efficiency.
[0145] The pattern tags in the above experiments were replaced with antimicrobial peptides, namely Cecropin A (abbreviated as C), PEW300 (abbreviated as P3, a basic mutant of Cecropin A); LL37 (abbreviated as L); and A12:Aurein 1.2 (abbreviated as A12), to further verify the secretory expression ability of the SHTXTHHly system. Figure 6It is evident that the expression and secretion levels were significantly enhanced after the N-terminus of SHT was fused, indicating that the newly constructed extracellular soluble expression and purification system SHTXTHHly can be applied to proteins and peptides of different properties and has good versatility.
[0146] Example 4
[0147] The method for identifying polypeptide antigenic epitopes described in this application is used for screening mutants of the target protein.
[0148] RGDS can specifically bind to integrin molecules on the surface of tumor cells. Using RGDS as the model peptide, its amino acid sequence is RGDS, and its nucleotide sequence is cgcggtgacagc. RGDS and its mutant RGES (amino acid sequence RGES, nucleotide sequence cgcggtgaaagc) were expressed using the same steps as in Example 3 to obtain the fusion proteins SHTRGDSTHHly and SHTRGESTHHly. The supernatant from *E. coli* induction expression containing the fusion proteins was directly incubated with U87 MG human glioma cells. PE-labeled goat anti-mouse secondary antibody and mouse anti-S-tag antibody were used as secondary antibodies. Flow cytometry analysis was performed after incubation, and the results are as follows: Figure 7 As shown, the RGDS sequence in the fusion protein can bind to integrins on the cell surface, and the fluorescence signal is significantly shifted to the right compared to the LB negative control. However, after mutation, the RGES sequence loses its ability to bind to integrins, resulting in a negative fluorescence signal. This method can be widely used for the rapid detection of target protein activity and for screening mutants that meet expected characteristics.
[0149] Example 5
[0150] The method for identifying polypeptide antigenic epitopes described in this application is used for rapid detection of target protein activity.
[0151] Using recombinant cytokine IL-15 as the model protein, expression was induced using the same method as in Example 3 to obtain the fusion protein SHTIL15THHly. The supernatant was used for Mo7e cell proliferation experiments, with commercially available monomeric IL-15 protein as a control. Specific operational steps are as follows:
[0152] 1) Mo7e cells were cultured in suspension in RPMI 1640 complete medium (Gibco, USA) containing 10 ng / mL of commercial recombinant human granulocyte-macrophage colony-stimulating factor (hGM-CSF); before the experiment, the cells were washed twice with complete medium without hGM-CSF and the cell density was diluted to 4 × 10⁶ cells / mL. 5 / mL, add 50 μL of cell suspension to each well of a 96-well plate, which is 20,000 cells per well.
[0153] 2) After starving the cells for 4 h, add 50 μL of culture medium containing different concentrations of SHTIL15THHly or IL-15 protein to each well;
[0154] 3) After culturing at 37℃ for 4 days, 10 μL / well of CCK-8 (Cell counting kit-8, CCK-8) solution was added, and after incubation at 37℃ for 2-3 hours, the absorbance at 450 nm was measured. The data were used to fit the S-curve to plot the proliferation curve, and the EC50 value of each protein group promoting the proliferation of Mo7e cells was calculated.
[0155] The results are as follows Figure 8 As shown, the EC50 of the SHTIL15THHly fusion protein is 801 pM, while the EC50 of the IL-15 control protein is 294 pM. Their activities are similar, indicating that the LB supernatant from induced expression of SHTIL15THHly can be directly used for activity detection. Because the extracellular secretory expression system constructed in this application has high efficiency, the content of the target protein in the culture medium is high. Furthermore, the proteins in the LB medium are all digested short peptides with a molecular weight of less than 3 kDa, minimizing the possibility of interference with the identification of target protein activity. These characteristics meet the conditions for direct detection of target protein activity using the culture medium. This is an advantage of the extracellular secretory expression system described in this application.
[0156] Example 6
[0157] Purification of exogenous proteins prepared using the method for identifying polypeptide epitopes described in this application
[0158] The supernatant from the culture medium after induction is collected and purified using a nickel column to obtain the target protein fusion protein. The target protein is then obtained after TEV digestion and a second nickel column purification. The purification steps using the RGDS fusion single-chain Fc antibody molecule RGDSmFc are illustrated below, illustrating the extracellular secretory expression system described in this application.
[0159] The purification process is as follows Figure 9 As shown, it specifically includes:
[0160] 1) Filter the culture medium supernatant used to induce expression using a 0.45 μm filter membrane; this is the loading solution.
[0161] 2) Turn on the purification system, set it to 1 mL / min for rinsing with 20% ethanol, and simultaneously connect a Smart His 6FF purification column;
[0162] 3) Replace with water, rinse 10 mL at 1 mL / min, then replace with equilibration buffer (20 mM PBS + 300 mM NaCl), flowing in at 1 mL / min;
[0163] 4) After equilibration, replace with the loading solution, flow in at 1 mL / min, for a total of 20 mL, and collect the flow-through solution at the same time. After collection, replace with the equilibration solution, flow in at 1 mL / min.
[0164] 5) Replace with elution buffer (20 mM PBS + 300 mM NaCl + 200 mM imidazole), flow in at 1 mL / min, and collect the elution buffer;
[0165] 6) Replace with regeneration solution (20 mM PBS + 300 mM NaCl + 400 mM imidazole), infusing at 1 mL / min;
[0166] 7) Replace with water, rinse 10 mL at 1 mL / min. Replace with 20% ethanol, rinse 10 mL at 1 mL / min, remove the purification column, and store at 4°C;
[0167] 8) Save the loading solution, flow-through solution and elution solution. At this time, the fusion protein is present in the elution solution and will be used for further detection or purification.
[0168] 9) Take 2 mL of elution buffer, add 222 μL of TEV buffer and 25 μL of TEV enzyme (coupled with His tag), and incubate at 4 degrees Celsius for enzyme digestion;
[0169] 10) Take the enzyme digestion solution, dilute it to a volume of 5 mL with loading buffer, and transfer it to a centrifuge tube;
[0170] 11) Open the purification system and perform a second Ni column purification following the same steps as the fusion protein purification described above. Collect the loading solution, flow-through solution, and elution solution. At this point, the target protein is present in the flow-through solution.
[0171] Using the above steps, the normal dimerized form of RGDSFc and the single-chain form of RGDSmFc were purified, and their structures are as follows: Figure 10 As shown. SDS-PAGE was used to analyze the products from each step of the purification process, such as... Figure 11 As shown, the first-step Ni column purification eluent E1 contains relatively pure SHTRGDSmFcTHHly fusion protein. After enzymatic digestion, the second-step Ni column loading flow-through contains high-purity target protein RGDSmFc, while eluent E2 contains HHly with a His tag and the SHT fragment. However, the SHT fragment, due to its smaller molecular weight than the protein marker, is not readily absorbed. Figure 11 It is not visible in the middle.
[0172] Example 7
[0173] The method for identifying polypeptide antigenic epitopes described in this application is used for rapid screening of antigenic epitopes.
[0174] In this embodiment, human glyceraldehyde-3-phosphate dehydrogenase (GAPDH) is used as the model antigen, and a mouse monoclonal antibody targeting human GAPDH (Proteintech, Cat No: 60004-1-Ig) is used as the model antibody. The specific epitope of GAPDH targeted by this antibody is unknown. Two methods can be used for rapid screening and identification of antigenic epitopes using the extracellular secretory expression system of this invention:
[0175] Option 1: Plate colony cloning screening, the specific steps are as follows:
[0176] 1) Since the length of a typical antigenic epitope is 5-10 amino acids, the full-length 335 aa GAPDH protein is sequentially broken down into peptides of 20 aa length, with adjacent peptides having a 10 aa overlap (e.g., Figure 12 (as shown)
[0177] 2) Synthesize double-stranded oligo DNA libraries for each polypeptide segment, and add single-stranded header sequences that overlap with S tag and His tag upstream and downstream, respectively. Through homologous recombination and other methods, homologous recombination is performed between the oligo library and the linearized vector containing S tag and His tag-HlyA to construct the expression plasmid library of SXHHly.
[0178] 3) A two-step transformation method was used. First, helper plasmids were transformed into the BL21(DE3) strain. After selecting single clones and confirming correct sequencing, electroporation competent cells were prepared. Second, the expression plasmid library was transformed into these competent cells and plated onto 10 LB agar plates containing ampicillin, kanamycin, arabinose, and IPTG. Ampicillin and kanamycin served as selection pressure to ensure the positive rate of transformants, while arabinose and IPTG induced the helper protein and target protein, respectively, to achieve the secretory expression of the target protein.
[0179] 4) Take plates of the expression strain cultured overnight at 37°C and observe the appearance of numerous transformants. Place an NC membrane or a methanol-activated PVDF membrane onto the plate, mark its corresponding position on the plate, let it stand for several minutes, then wash with PBST, block with skim milk, incubate with GADPH primary antibody, wash with PBST, incubate with secondary antibody, wash with PBST again, and incubate with TMB chromogenic solution. Observe the color development results, such as... Figure 13 As shown;
[0180] 5) Culture the clones corresponding to the positive clones, extract plasmids, and sequence the expression plasmids. Based on the sequence inserted into the expression plasmid vector, deduce the polypeptide sequence encoded by the sequence, which is the epitope recognized by the GAPDH antibody.
[0181] Option 2 involves a multi-round, progressive approach to gradually narrow down the scope. The specific implementation steps are as follows:
[0182] 1) Antigen sequence segmentation: The full-length 335 aa GAPDH protein was divided into six segments of 70 aa each, with an overlap of 17 aa, based on the average length: 1-70 aa, 54-123 aa, 107-176 aa, 160-229 aa, 213-282 aa, and 266-335 aa.
[0183] 2) Expression plasmid construction: Primers were designed and conventional PCR was performed to amplify the gene fragments into 6 segments. The 6 expression plasmids were constructed according to the steps in Example 2.
[0184] 3) Fusion protein induction expression: The expression plasmid and helper plasmid were transformed into competent BL21(DE3) strain. Six tubes of competent cells were streaked onto six regions of LB agar plates containing ampicillin, kanamycin, arabinose, and IPTG, and then incubated in a bacterial incubator. Ampicillin and kanamycin served as selection pressure to ensure the positive rate of transformants, while arabinose and IPTG induced the helper protein and target protein, respectively, to achieve the secretory expression of the target protein.
[0185] 4) Plate colony blot and immunoblot: Plates containing the expression strain cultured overnight at 37°C were used. Numerous transformants were observed in all six regions. NC membranes or methanol-activated PVDF membranes were plated onto the plates. After standing for several minutes, the plates were washed with PBST, blocked with skim milk, incubated with GADPH primary antibody, washed with PBST, incubated with secondary antibody, washed with PBST again, and incubated with TMB chromogenic solution. The chromogenic results were observed to determine the target epitope within the 1-70 aa range. Figure 14 As shown;
[0186] 5) Continue dividing 1-70 aa into 5 segments: 1-14 aa, 1-28 aa, 15-42 aa, 42-70 aa, and 57-70 aa. Repeat steps 2), 3), and 4) above, and observe the final colorimetric results. Figure 15 As shown, 1-28 aa is positive, but 1-14 aa and 15-28 aa do not contain the target epitope on their own.
[0187] Both of the above methods can quickly narrow down the epitope range to 10-20 amino acids. For more precise antigenic epitope information, the peptide segmentation range can be further narrowed by repeating the above screening steps, or amino acid mutations, such as alanine scanning, can be performed. This method can be applied to the detection and screening of peptides and active regions involved in other molecular interactions.
[0188] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0189] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can modify, substitute, and vary the above embodiments within the scope of this application.
Claims
1. A method for identifying polypeptide antigenic epitopes, characterized in that, include: Escherichia coli cells were transformed with an expression vector containing nucleic acid encoding a target protein, and the E. coli cells were cultured to achieve secretory expression of the target protein, wherein the target protein is a polypeptide fragment containing an antigenic epitope. Using E. coli expression supernatant containing the target protein, without purification, the supernatant was directly added to the detection system for affinity testing to screen for antigen epitopes; The target protein is expressed in the form of a fusion protein, which comprises: Target protein; The S tag at the end of N; and A targeting sequence at the C-terminus, which enables the fusion protein to target to the extracellular domain of E. coli cells; The target sequence is the hemolysin A signal peptide; The fusion protein also includes two tag sequences, located between the target protein and the targeting sequence, and between the S tag and the target protein, respectively; The fusion protein also contains two protease recognition sequences, located between the N-terminus and C-terminus of the target protein and the aforementioned tag sequence, respectively.
2. The method according to claim 1, characterized in that, Antibodies are used for incubation and color development, and affinity testing is performed to screen for antigenic epitopes.
3. The method according to claim 1, characterized in that, The tag sequence includes any one or more selected from His tags, GST tags, Strep tags, and MBP tags.
4. The method according to claim 1, characterized in that, The protease recognition sequence is selected from at least one of the following: TEV protease recognition sequence, thrombin recognition sequence, enterokinase recognition sequence, 3C protease recognition sequence, and inteptide.
5. The method according to claim 1, characterized in that, The method further includes: The culture supernatant is collected and purified, and then contacted with a protease to obtain a free target protein, wherein the protease is identified based on the protease recognition sequence.
6. The application of the method for identifying polypeptide antigenic epitopes according to any one of claims 1-5 in affinity testing or mutant screening.