Fusion protein, and preparation method and use therefor

By fusing other amino acid fragments at the nitrogen and/or carbon ends of Brazilian sweet proteins to form fusion proteins, the problem of planting and low content in Brazilian sweet proteins is solved, and the technical effect of increasing sweetness and large-scale production is achieved.

WO2025130287A1PCT designated stage expired Publication Date: 2025-06-26SHENZHEN TAILI BIOTECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/124751
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-04
Filing Date
2024-10-14
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The application of Brazilian sweet protein in the prior art is limited by the problem of difficulty in planting tropical plants and low natural protein content, and the sweetness during heterologous expression is reduced, which limits large-scale commercial production.

Method used

By fusing other amino acid fragments at the nitrogen and/or carbon ends of Brazilian sweet proteins, a fusion protein is formed, which increases its sweetness after heterologous expression and is correctly folded and expressed in CHO cells.

Benefits of technology

The sweetness of the fusion protein is 3.5-16 times higher than that of natural Brazilian sweet protein, shortening the sense of sweetness delay and achieving the technical basis of large-scale production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024124751_26062025_PF_FP_ABST
    Figure CN2024124751_26062025_PF_FP_ABST
Patent Text Reader

Abstract

A fusion protein and a preparation method and use therefor, relating to the field of synthetic biology and recombinant proteins. The amino acid sequence of the fusion protein may comprise a first unit and a second unit, and a carbon end of the first unit is connected to a nitrogen end of the second unit. The amino acid sequence of the fusion protein may comprise a first unit and a third unit, and a nitrogen end of the first unit is connected to a carbon end of the third unit. The amino acid sequence of the fusion protein may comprise a first unit, a second unit and a third unit, a carbon end of the first unit is connected to a nitrogen end of the second unit, and a nitrogen end of the first unit is connected to a carbon end of the third unit. The application further relates to a preparation method of the fusion protein, a polynucleotide, a vector, a cell and a use thereof. For the fusion protein, other fragments are fused at the nitrogen end and / or the carbon end of a natural Brazzein, so that the sweetness is improved, the sweetness delay is shortened, and heterologous expression can also be achieved, such that a technical foundation is laid for subsequent large-scale application in the fields of sweetening agents etc.
Need to check novelty before this filing date? Find Prior Art

Description

A fusion protein and its preparation method and application

[0001] This application claims priority to the Chinese patent application filed on December 19, 2023 (application number: 2023117586160, invention title: Fusion protein, preparation method and application thereof), the Chinese patent application filed on June 4, 2024 (application number: 202410716061.1, invention title: Fusion protein, preparation method and application thereof in the preparation of sweeteners), and the Chinese patent application filed on June 4, 2024 (application number: 2024107180973, invention title: Sweet fusion protein, preparation method and application thereof), and parts of the contents of the Chinese patent application are incorporated into this application in their entirety by reference. Technical Field

[0002] The present invention relates to the field of synthetic biology and recombinant proteins, and in particular to a fusion protein, a method for preparing the fusion protein, a polynucleotide, a vector, a cell, and applications of the fusion protein, polynucleotide, vector or cell. Background Art

[0003] In the modern food industry, sweeteners such as sucrose, maltose, fructose, and glucose remain indispensable food additives. However, due to long-term excessive consumption of these traditional sweeteners, an increasing number of people are suffering from various health conditions, including but not limited to the three highs, obesity, and dental caries. Therefore, there is an urgent need for new sweeteners that combine high sweetness with low calories, and that avoid the potential health risks of artificial sweeteners, to meet people's lifestyle and health needs.

[0004] As we all know, most proteins in nature are tasteless. However, some plant proteins have a sweet taste or can transform sour flavors into sweetness. These proteins, collectively known as plant sweet proteins, are not only natural and non-toxic, but also highly sweet and low in calories. Their digestion and degradation products also produce naturally occurring amino acids essential to the human body. They are highly safe and offer a promising alternative to traditional and artificial sweeteners.

[0005] So far, eight sweet proteins have been found in plants, including Thaumatin, Mabinlin, Pentadin, Curculin, Brazzein, Monellin, Miraculin, and Neoculin. Among them, Brazzein is a sweet protein isolated and purified by Ding et al. from the fruit of the wild plant Pentadiplandra Brazzeana Baillon in West Africa in 1994. It can cause sweetness by itself, has no sweetness-modifying effect, and its sweetness is 500-2000 times that of sucrose. Compared with other sweet proteins, Brazzein has the smallest molecular weight and the best water solubility. Its aqueous solution still retains sweetness after 4h of heat treatment at 80°C. It has good thermal stability and pH stability and is suitable for various food processing technologies. However, the application of Brazzein in the existing technology also has the following defects:

[0006] First, brasiliensis is obtained from a tropical plant unique to Africa, but this tropical plant is difficult to bear fruit in other environments, which limits its development and application. In addition, the natural brasiliensis content in the plant is low, making it difficult to extract and produce on a large scale. Secondly, although many studies have confirmed that the gene sequence of the target protein can be transferred into various expression hosts through protein engineering technology, most of them remain at the laboratory stage and there are still some problems to be solved. For example, brasiliensis has been heterologously expressed in microorganisms (such as Escherichia coli, yeast and lactic acid bacteria) and plants as vectors. However, the sweet protein is not properly folded or has a mismatched disulfide bond, which leads to a reduction or even disappearance of its sweetness during heterologous expression, thereby limiting the large-scale commercial production and application of the sweet protein. Therefore, how to improve brasiliensis so that it can be more easily used as a sweetener is a problem to be solved.

[0007] In view of this, the present invention is proposed.

[0008] Summary of the Invention

[0009] One of the purposes of the present invention is to provide a fusion protein.

[0010] The technical solution of the present invention to solve the above technical problem is as follows: a fusion protein, wherein the amino acid sequence of the fusion protein comprises a first unit and a second unit, wherein the carbon end of the first unit is connected to the nitrogen end of the second unit;

[0011] The amino acid sequence of the first unit is shown in SEQ ID NO.1 or SEQ ID NO.25;

[0012] The second unit comprises a fragment consisting of 2-22 amino acid residues, and each amino acid residue is independently selected from histidine His, lysine Lys or aspartic acid Asp.

[0013] On the basis of the above technical solution, the present invention can also be improved as follows.

[0014] Furthermore, the second unit is selected from a fragment consisting of 2, 4, 6, 8, 10, 12 or 16 histidines.

[0015] Furthermore, the second unit is a fragment consisting of 6 histidines.

[0016] Furthermore, the amino acid sequence of the fusion protein is selected from SEQ ID NO.3, SEQ ID NO.5, SEQ ID NO.7, SEQ ID NO.9, SEQ ID NO.11, SEQ ID NO.13, SEQ ID NO.15, SEQ ID NO.27, SEQ ID NO.29, SEQ ID NO.31, SEQ ID NO.33, SEQ ID NO.35, SEQ ID NO.37, SEQ ID NO.39, SEQ ID NO.60, SEQ ID NO.62, SEQ ID NO.64, SEQ ID NO.66, SEQ ID NO.68 or SEQ ID NO.70.

[0017] A further beneficial effect of the above is that the present invention has found that when the amino acid sequence of the fusion protein is as above, the sweetness is increased by 9.5-16.5 times compared with the natural Brazilian sweet protein before optimization, achieving an unexpected technical effect.

[0018] Another technical solution of the present invention to solve the above technical problem is as follows: a fusion protein, wherein the amino acid sequence of the fusion protein comprises a first unit and a third unit, wherein the nitrogen end of the first unit is connected to the carbon end of the third unit;

[0019] The amino acid sequence of the first unit is shown in SEQ ID NO.1 or SEQ ID NO.25;

[0020] The third unit comprises a fragment consisting of 2-22 amino acid residues, and each amino acid residue is independently selected from histidine His, lysine Lys or aspartic acid Asp.

[0021] On the basis of the above technical solution, the present invention can also be improved as follows.

[0022] Furthermore, the third unit is a fragment consisting of 2, 8 or 16 histidines.

[0023] Furthermore, the amino acid sequence of the fusion protein is selected from SEQ ID NO.17, SEQ ID NO.19, SEQ ID NO.21, SEQ ID NO.41, SEQ ID NO.43 or SEQ ID NO.45.

[0024] A further beneficial effect of the above method is that the present invention has found that when the amino acid sequence of the fusion protein is as above, the sweetness is 3.7-7.1 times higher than that of the natural Brazilian thaumatin before optimization.

[0025] Another technical solution of the present invention to solve the above technical problem is as follows: a fusion protein, wherein the amino acid sequence of the fusion protein comprises a first unit, a second unit, and a third unit, wherein the carbon end of the first unit is connected to the nitrogen end of the second unit, and the nitrogen end of the first unit is connected to the carbon end of the third unit;

[0026] The amino acid sequence of the first unit is shown in SEQ ID NO.1 or SEQ ID NO.25;

[0027] The second unit comprises a fragment consisting of 2-22 amino acid residues, each amino acid residue being independently selected from histidine His, lysine Lys or aspartic acid Asp;

[0028] The third unit comprises a fragment consisting of 2-22 amino acid residues, and each amino acid residue is independently selected from histidine His, lysine Lys or aspartic acid Asp.

[0029] On the basis of the above technical solution, the present invention can also be improved as follows.

[0030] Furthermore, the second unit and the third unit are independently fragments consisting of 2-16 histidines.

[0031] Furthermore, the second unit is a fragment consisting of 2, 4, 6, 8, 10, 12 or 16 histidines, and the third unit is a fragment consisting of 2, 8 or 16 histidines.

[0032] Furthermore, the second unit is a fragment consisting of 6 histidines.

[0033] Furthermore, the second unit and the third unit are both fragments consisting of 8 histidines.

[0034] Furthermore, the amino acid sequence of the fusion protein is selected from SEQ ID NO.23 or SEQ ID NO.47, SEQ ID NO.68 or SEQ ID NO.70.

[0035] A further beneficial effect of the above method is that the present invention has found that when the amino acid sequence of the fusion protein is as above, the sweetness is 10.5 times higher than that of the natural Brazilian thaumatin before optimization.

[0036] The principle of the fusion protein of the present invention is:

[0037] In the present invention, the first unit is the amino acid sequence of natural brasiliensis, and the second and third units are amino acid sequences to be fused, which can enhance the sweetness of natural brasiliensis after heterologous expression.

[0038] After a lot of creative work, the present inventors surprisingly and unexpectedly discovered that by fusing the second unit and / or the third unit to the nitrogen end and / or the carbon end of the first unit, the sweetness of the resulting fusion protein is at least 3.5-16 times higher than that of natural brasiliensis, and the delayed sweetness of natural brasiliensis is shortened. It can also be heterologously expressed, laying a technical foundation for large-scale production.

[0039] The beneficial effects of the fusion protein of the present invention are:

[0040] 1. The fusion protein of the present invention is based on natural brasiliensis, with other fragments fused to the nitrogen-terminus and / or carbon-terminus of natural brasiliensis. This fusion protein has a sweetness at least 3.5-16 times greater than that of natural brasiliensis, reduces the delayed sweetness of natural brasiliensis, and can be expressed heterologously, achieving unexpected technical benefits.

[0041] 2. The fusion protein of the present invention not only significantly increases the sweetness of natural brasilien, but also improves the taste of natural brasilien, shortens the time it takes to perceive sweetness in the mouth, and lays a technical foundation for the industrial production of sweeteners and downstream products containing sweeteners.

[0042] A second object of the present invention is to provide a polynucleotide.

[0043] The technical solution of the present invention to solve the above technical problems is as follows: a polynucleotide encoding the above fusion protein.

[0044] The third object of the present invention is to provide a carrier.

[0045] The technical solution of the present invention to solve the above technical problems is as follows: a vector carrying the above polynucleotide.

[0046] On the basis of the above technical solution, the present invention can also be improved as follows.

[0047] Furthermore, the vector also contains a plasmid.

[0048] A fourth object of the present invention is to provide a cell.

[0049] The technical solution of the present invention to solve the above technical problems is as follows: a cell, wherein the cell expresses the above fusion protein, or carries the above polynucleotide, or contains the above vector.

[0050] On the basis of the above technical solution, the present invention can also be improved as follows.

[0051] Furthermore, the above-mentioned polynucleotide is integrated into the genome of the cell.

[0052] Furthermore, the cells include mammalian cells.

[0053] Furthermore, the cell is a CHO cell, the genome of the CHO cell is integrated with the polynucleotide of claim 8, and the insertion position of the polynucleotide is NW_003616785.1:83044.

[0054] A fifth object of the present invention is to provide a method for preparing the above-mentioned fusion protein.

[0055] The technical solution of the present invention to solve the above technical problem is as follows: a method for preparing a fusion protein, comprising the following steps: expressing the fusion protein in the above-mentioned cells.

[0056] On the basis of the above technical solution, the present invention can also be improved as follows.

[0057] Furthermore, it also includes integrating the above-mentioned polynucleotide into the genome of the cell.

[0058] Furthermore, the cell used to express the fusion protein contains a marker gene when the polynucleotide encoding the fusion protein is not integrated into the cell genome, and the marker gene is knocked out after the polynucleotide is integrated into the cell genome.

[0059] Furthermore, the cell is a CHO cell expressing green fluorescent protein, and the integration site of the green fluorescent protein gene in the CHO cell and the insertion site of the polynucleotide encoding the fusion protein in the CHO cell are both NW_003616785.1:83044.

[0060] A sixth object of the present invention is to provide applications of the above-mentioned fusion protein, the above-mentioned polynucleotide, the above-mentioned vector or the above-mentioned cell.

[0061] The technical solution of the present invention to solve the above technical problem is as follows: use of the above fusion protein, the above polynucleotide, the above vector or the above cell in the preparation of food additives, sweeteners, medicines and / or feeds.

[0062] The beneficial effects of the present invention are: the above-mentioned fusion protein, the above-mentioned polynucleotide, the above-mentioned vector or the above-mentioned cell can be used to prepare food additives, sweeteners, medicines and / or feeds, overcoming the defects of existing natural Brazilian sweet protein and alleviating the problem of lack of protein sweet substances in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0064] FIG1 is a plasmid map of the natural brasilienin expression plasmid Brz in Example 1;

[0065] FIG2 is a plasmid map of the fusion protein expression plasmid Brz-C1 in Example 1;

[0066] FIG3 is a plasmid map of Bxb-1 integrase in Example 2. DETAILED DESCRIPTION

[0067] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0068] The term "polynucleotide" herein refers to a polymeric form of nucleotides of any length, including ribonucleotides and / or deoxyribonucleotides. Examples of polynucleotides include, but are not limited to, single-stranded, double-stranded, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or polymers comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derived nucleotide bases. Coding may alternatively encode a sense strand or an antisense strand. Polynucleotides may be naturally occurring, synthetic, recombinant, or any combination thereof. The terms "polynucleotide" and "nucleic acid" are used interchangeably herein.

[0069] As used herein, the term "vector" refers to a nucleic acid delivery vehicle into which a polynucleotide can be inserted. A vector is called an expression vector when it can express a protein encoded by the inserted polynucleotide. A vector can be introduced into cells through transformation, transduction, or transfection, allowing the genetic material it carries to be expressed in the cells.

[0070] The vectors are well known to those skilled in the art, and include, but are not limited to, plasmids; phagemids; cosmids; artificial chromosomes, such as yeast artificial chromosomes (YACs), bacterial artificial chromosomes (BACs), or P1-derived artificial chromosomes (PACs); bacteriophages such as lambda phage or M13 phage, and animal viruses. Animal viruses that can be used as vectors include, but are not limited to, retroviruses (including lentiviruses), adenoviruses, adeno-associated viruses, herpes viruses (such as herpes simplex viruses), poxviruses, baculoviruses, papillomaviruses, and papillomaviruses. In some embodiments, the vectors of the present invention contain regulatory elements commonly used in genetic engineering, such as enhancers, promoters, internal ribosome entry sites (IRESs), and other expression control elements (such as transcription termination signals, or polyadenylation signals and poly-U sequences, etc.).

[0071] As used herein, the expressions "cell," "cell line," and "cell culture" are used interchangeably, and all such designations include progeny. Progeny may not necessarily be identical to the original cell due to natural, accidental, or deliberate mutation and may differ from the original cell in morphology and / or in genomic DNA.

[0072] As used herein, the term "amino acid" refers to naturally occurring amino acids and synthetic amino acids, as well as amino acid analogs and amino acid mimetics that function in a manner similar to naturally occurring amino acids. Naturally occurring amino acids include those encoded by the genetic code and modified amino acids thereof. Common natural amino acids include: alanine (Ala; A), arginine (Arg; R), asparagine (Asn; N), aspartic acid (Asp; D), cysteine ​​(Cys; C); glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G); histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V). In this article, the Chinese names, three-letter abbreviations and single-letter abbreviations of amino acids are used interchangeably.

[0073] In a first aspect, a fusion protein is provided, wherein the fusion protein enhances the sweetness of brasiliensis after heterologous expression by fusing other fragments to the nitrogen terminus and / or carbon terminus of the brasiliensis protein.

[0074] The present invention provides a technical solution, in an optional embodiment, the amino acid sequence of the fusion protein comprises a first unit and a second unit, and the carbon end of the first unit is connected to the nitrogen end of the second unit;

[0075] The amino acid sequence of the first unit is shown in SEQ ID NO.1 or SEQ ID NO.25;

[0076] The second unit comprises a fragment consisting of 2-22 amino acid residues, and each amino acid residue is independently selected from histidine His, lysine Lys or aspartic acid Asp.

[0077] In an optional embodiment, the second unit is selected from a fragment consisting of 2, 4, 6, 8, 10, 12 or 16 histidines.

[0078] In an optional embodiment, the second unit is a fragment consisting of 6 histidines.

[0079] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the second unit is shown in SEQ ID NO.50, that is, the carbon-terminus of the brazilian melamine sequence of SEQ ID NO.1 is added with a protein sequence HHHHHHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO.3. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.4.

[0080] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, wherein the amino acid sequence of the first unit is as shown in SEQ ID NO.1, and the amino acid sequence of the second unit is HH, i.e., a protein sequence HH is added to the carbon terminus of the brazilin sequence of SEQ ID NO.1. The amino acid sequence of the fusion protein is as shown in SEQ ID NO.5. In an optional embodiment, the nucleotide sequence encoding the fusion protein is as shown in SEQ ID NO.6.

[0081] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the second unit is shown in SEQ ID NO.51, that is, the carbon-terminus of the brazilian melamine sequence of SEQ ID NO.1 is added with a protein sequence HHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO.7. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.8.

[0082] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the second unit is shown in SEQ ID NO.52, that is, the carbon-terminus of the brazilian melamine sequence of SEQ ID NO.1 is added with a protein sequence HHHHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO.9. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.10.

[0083] In an optional embodiment, the sweet fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the second unit is shown in SEQ ID NO.53, that is, the carbon-terminus of the brazilian thaumatin sequence of SEQ ID NO.1 is added with a protein sequence HHHHHHHHDDDDK. The amino acid sequence of the fusion protein is shown in SEQ ID NO.11. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.12.

[0084] In an optional embodiment, the sweet fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the second unit is shown in SEQ ID NO.54, that is, a protein sequence HHHHHHHHDDDDKHHHHHHHHDDDDK is added to the carbon end of the brazilin sequence of SEQ ID NO.1. The amino acid sequence of the fusion protein is shown in SEQ ID NO.13. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.14.

[0085] In an optional embodiment, the sweet fusion protein consists of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the second unit is shown in SEQ ID NO.55, that is, a protein sequence HHHHHHHHHHHHHHHHHDDDDK is added to the carbon end of the brazilin sequence of SEQ ID NO.1. The amino acid sequence of the fusion protein is shown in SEQ ID NO.15. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.16.

[0086] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.25, and the amino acid sequence of the second unit is shown in SEQ ID NO.53, that is, the carbon-terminus of the brazilian melamine sequence of SEQ ID NO.25 is added with a protein sequence HHHHHHHHDDDDK. The amino acid sequence of the fusion protein is shown in SEQ ID NO.35. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.36.

[0087] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.25, and the amino acid sequence of the second unit is shown in SEQ ID NO.54, that is, the carbon-terminus of the brazilian melamine sequence of SEQ ID NO.25 is added with a protein sequence HHHHHHHHDDDDKHHHHHHHHDDDDK. The amino acid sequence of the fusion protein is shown in SEQ ID NO.37. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.38.

[0088] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.25, and the amino acid sequence of the second unit is shown in SEQ ID NO.55, that is, the carbon-terminus of the brazilian melamine sequence of SEQ ID NO.25 is added with a protein sequence HHHHHHHHHHHHHHHHHDDDDK. The amino acid sequence of the fusion protein is shown in SEQ ID NO.39. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.40.

[0089] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the second unit is shown in SEQ ID NO.72, that is, a protein sequence HHHHHHHHHH is added to the carbon end of the brazilin sequence of SEQ ID NO.1. The amino acid sequence of the fusion protein is shown in SEQ ID NO.60. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.61.

[0090] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the second unit is shown in SEQ ID NO.73, that is, a protein sequence HHHHHHHHHHHH is added to the carbon end of the brazilin sequence of SEQ ID NO.1. The amino acid sequence of the fusion protein is shown in SEQ ID NO.62. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.63.

[0091] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the second unit is shown in SEQ ID NO.56, that is, a protein sequence HHHHHHHHHHHHHHHHH is added to the carbon end of the brazilin sequence of SEQ ID NO.1. The amino acid sequence of the fusion protein is shown in SEQ ID NO.64. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.65.

[0092] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO. 25, and the amino acid sequence of the second unit is shown in SEQ ID NO. 50, that is, the carbon-terminus of the brazilian melamine sequence of SEQ ID NO. 25 is added with a protein sequence HHHHHHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO. 27. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO. 28.

[0093] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, wherein the amino acid sequence of the first unit is shown in SEQ ID NO. 25, and the amino acid sequence of the second unit is HH, i.e., a protein sequence HH is added to the carbon terminus of the brazilin sequence of SEQ ID NO. 25. The amino acid sequence of the fusion protein is shown in SEQ ID NO. 29. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO. 30.

[0094] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO. 25, and the amino acid sequence of the second unit is shown in SEQ ID NO. 51, i.e., the carbon-terminus of the brazilian tamarind sequence of SEQ ID NO. 25 is added with a protein sequence HHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO. 31. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO. 32.

[0095] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO. 25, and the amino acid sequence of the second unit is shown in SEQ ID NO. 52, i.e., the carbon-terminus of the brazilian tamarind sequence of SEQ ID NO. 25 is added with a protein sequence HHHHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO. 33. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO. 34.

[0096] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.25, and the amino acid sequence of the second unit is shown in SEQ ID NO.72, that is, the carbon-terminus of the brazilian melamine sequence of SEQ ID NO.1 is added with a protein sequence HHHHHHHHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO.66. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.67.

[0097] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.25, and the amino acid sequence of the second unit is shown in SEQ ID NO.73, that is, the carbon-terminus of the brazilian melamine sequence of SEQ ID NO.1 is added with a protein sequence HHHHHHHHHHHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO.68. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.69.

[0098] In an optional embodiment, the fusion protein is composed of a first unit and a second unit, the amino acid sequence of the first unit is shown in SEQ ID NO.25, and the amino acid sequence of the second unit is shown in SEQ ID NO.56, that is, a protein sequence HHHHHHHHHHHHHHHHH is added to the carbon end of the brazilin sequence of SEQ ID NO.1. The amino acid sequence of the fusion protein is shown in SEQ ID NO.70. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.71.

[0099] Another technical solution provided by the present invention is a fusion protein, wherein the amino acid sequence of the fusion protein comprises a first unit and a third unit, wherein the nitrogen end of the first unit is connected to the carbon end of the third unit;

[0100] The amino acid sequence of the first unit is shown in SEQ ID NO.1 or SEQ ID NO.25;

[0101] The third unit comprises a fragment consisting of 2-22 amino acid residues, and each amino acid residue is independently selected from histidine His, lysine Lys or aspartic acid Asp.

[0102] The third unit is a fragment consisting of 2, 8 or 16 histidines.

[0103] In an optional embodiment, the fusion protein is composed of a first unit and a third unit, wherein the amino acid sequence of the first unit is as shown in SEQ ID NO.1, and the amino acid sequence of the third unit is HH, i.e., a protein sequence HH is added to the nitrogen end of the brazilian tamanin sequence of SEQ ID NO.1. The amino acid sequence of the fusion protein is as shown in SEQ ID NO.17. In an optional embodiment, the nucleotide sequence encoding the fusion protein is as shown in SEQ ID NO.18.

[0104] In an optional embodiment, the fusion protein is composed of a first unit and a third unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the third unit is shown in SEQ ID NO.50, that is, the brazilian tamanin sequence of SEQ ID NO.1 is added to the nitrogen end of the protein sequence HHHHHHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO.19. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.20.

[0105] In an optional embodiment, the fusion protein is composed of a first unit and a third unit, the amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequence of the third unit is shown in SEQ ID NO.56, that is, a protein sequence HHHHHHHHHHHHHHHHH is added to the nitrogen end of the brazilian tamarind sequence of SEQ ID NO.1. The amino acid sequence of the fusion protein is shown in SEQ ID NO.21. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.22.

[0106] In an optional embodiment, the fusion protein is composed of a first unit and a third unit, the amino acid sequence of the first unit is shown in SEQ ID NO. 25, and the amino acid sequence of the third unit is HH, i.e., a protein sequence HH is added to the nitrogen end of the brazilian tamanin sequence of SEQ ID NO. 25. The amino acid sequence of the fusion protein is shown in SEQ ID NO. 41. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO. 42.

[0107] In an optional embodiment, the fusion protein is composed of a first unit and a third unit, the amino acid sequence of the first unit is shown in SEQ ID NO.25, and the amino acid sequence of the third unit is shown in SEQ ID NO.50, i.e., the brazilin sequence of SEQ ID NO.25 is added to the nitrogen end of the protein sequence HHHHHHHH. The amino acid sequence of the fusion protein is shown in SEQ ID NO.43. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.44.

[0108] In an optional embodiment, the fusion protein is composed of a first unit and a third unit, the amino acid sequence of the first unit is shown in SEQ ID NO.25, and the amino acid sequence of the third unit is shown in SEQ ID NO.56, that is, a protein sequence HHHHHHHHHHHHHHHH is added to the nitrogen end of the brazilian tamanin sequence of SEQ ID NO.25. The amino acid sequence of the fusion protein is shown in SEQ ID NO.45. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.46.

[0109] Another technical solution provided by the present invention is a fusion protein, wherein the amino acid sequence of the fusion protein comprises a first unit, a second unit, and a third unit, wherein the carbon end of the first unit is connected to the nitrogen end of the second unit, and the nitrogen end of the first unit is connected to the carbon end of the third unit;

[0110] The amino acid sequence of the first unit is shown in SEQ ID NO.1 or SEQ ID NO.25;

[0111] The second unit comprises a fragment consisting of 2-22 amino acid residues, each amino acid residue being independently selected from histidine His, lysine Lys or aspartic acid Asp;

[0112] The third unit comprises a fragment consisting of 2-22 amino acid residues, and each amino acid residue is independently selected from histidine His, lysine Lys or aspartic acid Asp.

[0113] The second unit and the third unit are independently fragments consisting of 2-16 histidines.

[0114] The second unit is a fragment consisting of 2, 4, 6, 8, 10, 12 or 16 histidines, and the third unit is a fragment consisting of 2, 8 or 16 histidines.

[0115] The second unit is a fragment consisting of 6 histidines.

[0116] The second unit and the third unit are both fragments consisting of 8 histidines.

[0117] In an optional embodiment, the fusion protein is composed of a first unit, a second unit, and a third unit. The amino acid sequence of the first unit is shown in SEQ ID NO.1, and the amino acid sequences of the second and third units are both shown in SEQ ID NO.50, that is, a protein sequence HHHHHHHH is added to the nitrogen and carbon ends of the brazilian tamanin sequence in SEQ ID NO.1. The amino acid sequence of the fusion protein is shown in SEQ ID NO.23. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO.24.

[0118] In an optional embodiment, the fusion protein is composed of a first unit, a second unit, and a third unit. The amino acid sequence of the first unit is shown in SEQ ID NO. 25, and the amino acid sequences of the second and third units are both shown in SEQ ID NO. 50, i.e., a protein sequence HHHHHHHH is added to the nitrogen and carbon ends of the brazilian tamanin sequence of SEQ ID NO. 25, respectively. The amino acid sequence of the fusion protein is shown in SEQ ID NO. 47. In an optional embodiment, the nucleotide sequence encoding the fusion protein is shown in SEQ ID NO. 48.

[0119] In an optional embodiment, the fusion protein in any of the above embodiments is expressed in mammalian cells, preferably CHO cells (Chinese Hamster Ovary cells). CHO cells enable the fusion protein to correctly fold its spatial structure. After expression in CHO cells, the fusion protein has a sweetness that is at least 3.5-16 times greater than that of the natural natron protein not fused to any fragments, and the delayed sweetness of the natural natron is shortened.

[0120] In an optional embodiment, the fusion protein in any of the above embodiments is endogenously expressed by CHO cells.

[0121] In an optional embodiment, the fusion protein is a protein endogenously expressed by CHO cells and has an amino acid sequence as shown in SEQ ID NO.9.

[0122] In a second aspect, a polynucleotide is also provided, wherein the polynucleotide contains a fragment encoding the fusion protein of the first aspect.

[0123] In an optional embodiment, the nucleotide sequence encoding the brazilin shown in SEQ ID NO.1 in the polynucleotide is shown in SEQ ID NO.2.

[0124] In an optional embodiment, the nucleotide sequence encoding the brazilin shown in SEQ ID NO. 25 in the polynucleotide is shown in SEQ ID NO. 26.

[0125] In an optional embodiment, the polynucleotide contains a fragment as shown in at least one of SEQ ID NO.4, SEQ ID NO.6, SEQ ID NO.8, SEQ ID NO.10, SEQ ID NO.18, SEQ ID NO.20, SEQ ID NO.22, SEQ ID NO.24, SEQ ID NO.28, SEQ ID NO.30, SEQ ID NO.32, SEQ ID NO.34, SEQ ID NO.42, SEQ ID NO.44, SEQ ID NO.46, SEQ ID NO.48, SEQ ID NO.61, SEQ ID NO.63, SEQ ID NO.65, SEQ ID NO.67, SEQ ID NO.69 and SEQ ID NO.71.

[0126] In an optional embodiment, the polynucleotide further contains a homology arm sequence for homologous recombination with the genome of the cell expressing the fusion protein.

[0127] In a third aspect, a vector is further provided, wherein the vector carries the polynucleotide described in the second aspect. In an optional embodiment, the vector comprises a plasmid.

[0128] In a fourth aspect, a cell is further provided, wherein the cell expresses the fusion protein in the first aspect, or carries the polynucleotide in the second aspect, or contains the vector in the third aspect.

[0129] In an optional embodiment, the polynucleotide described in the second aspect, ie, the polynucleotide expressing the fusion protein, is integrated into the genome of the cell to achieve endogenous expression of the fusion protein in the cell.

[0130] In alternative embodiments, the cells comprise mammalian cells.

[0131] In an alternative embodiment, the mammalian cells include CHO cells.

[0132] In an optional embodiment, the cell is a CHO cell, a polynucleotide expressing the fusion protein is integrated into the genome of the CHO cell, and the insertion position of the polynucleotide is NW_003616785.1:83044.

[0133] In an optional embodiment, the nucleotide sequence shown in SEQ ID NO. 4 is integrated into the genome of the CHO cell.

[0134] In a fifth aspect, a method for preparing the above-mentioned fusion protein is also provided, which comprises expressing the fusion protein described in the first aspect in a cell.

[0135] In an optional embodiment, the cell includes the cell described in the fourth aspect.

[0136] In an optional embodiment, the fusion protein is expressed endogenously or exogenously in the cell.

[0137] In an optional embodiment, the cells include CHO cells, and the polynucleotide encoding the fusion protein is integrated into the genome of the CHO cells to achieve endogenous expression of heterologous proteins. The fusion protein can be secreted extracellularly in the CHO cells, and there are no abnormalities in the structure and function of the protein. It can be used after simple purification in downstream processes, which is greatly beneficial to the realization of large-scale production.

[0138] In an optional embodiment, the preparation method includes integrating the polynucleotide encoding the fusion protein into the genome of a cell for expressing the fusion protein, thereby achieving endogenous expression of the fusion protein in the cell. The polynucleotide encoding the fusion protein can be integrated into the genome of the cell using any conventional method known in the art. Exemplary methods include, but are not limited to, using gene editing systems, such as the CRISPR / Cas9 system; using a transposon system; or using an integrase system.

[0139] In an alternative embodiment, the polynucleotide encoding the fusion protein is integrated into the genome of the cell using Bxb-1 integrase.

[0140] In an optional embodiment, the cell used to express the fusion protein contains a marker gene when the polynucleotide encoding the fusion protein is not integrated into the cell genome, and the marker gene is knocked out after the polynucleotide is integrated into the cell genome. The marker gene can be any marker gene known in the art. Exemplary marker genes include fluorescent protein genes, such as EGFP, TagGFP, AcGFP, EBFP, ECFP, CyPet, EYFP, mOrange, mRuby, or mApple; exemplary marker genes may also include resistance genes.

[0141] In an optional embodiment, the preparation method further comprises screening cells that positively express the fusion protein by detecting whether a marker gene in the cells is expressed.

[0142] In an optional embodiment, the cells used to construct the expression of the fusion protein are CHO cells expressing green fluorescent protein, and the integration site of the green fluorescent protein gene (EGFP) in the CHO cells and the insertion site of the polynucleotide encoding the fusion protein in the CHO cells are both NW_003616785.1:83044. When the polynucleotide encoding the fusion protein is integrated into the CHO cells expressing green fluorescent protein, EGFP is knocked out, and the CHO cells that do not express fluorescence are CHO cells that positively express the fusion protein.

[0143] The present invention is further described below by way of specific examples. However, it should be understood that these examples are merely provided for more detailed description and are not to be construed as limiting the present invention in any form.

[0144] The main biological materials and reagents used in the following examples are as follows:

[0145] 1. Strains and vectors: The site-directed integration CHO cell line and vector pDonor1.0 were constructed and preserved by the applicant; Escherichia coli DH5α and Hygromycin B were purchased from Thermo Fisher.

[0146] 2. Enzymes and kits: Restriction endonucleases were purchased from NEB, and plasmid extraction kits and gel purification and recovery kits were purchased from TIANGEN.

[0147] 3. Culture medium:

[0148] Escherichia coli culture medium (LB medium): 0.5% yeast extract, 1% peptone, 1% NaCl, pH 7.0; LB+Amp medium: LB medium plus 100 μg / mL ampicillin;

[0149] Ex-cell Advanced CHO Fed-batch medium and glucose solution were purchased from Sigma, Cell Boost7a and Cell Boost7b were purchased from Hyclone, and Glutamax (100X) was purchased from Gibco. D-(+)-Glucose solution was purchased from Sigma.

[0150] 4. Preparation of Stable High-Fluorescence Cells GBB003 Specific consumables include: Neon Resuspension Buffer R (ThermoFisher), a resuspension buffer for cell electroporation; E1 Buffer (ThermoFisher), an electroporation buffer; recovery medium consisting of 80% (v / v) EX-CELL CHO Cloning Medium (Sigma-Aldrich) and 20% (v / v) EX-CELL Advanced CHO Fed-batch Medium (Sigma-Aldrich) supplemented with 1% GlutaMAX (ThermoFisher); expansion medium and subculture medium are both EX-CELL Advanced CHO Fed-batch CHO medium was supplemented with 1% GlutaMAX (ThermoFisher); pressurized medium 1 was passage medium supplemented with G418 at a final concentration of 800 μg / ml; pressurized medium 2 was passage medium supplemented with hygromycin at a final concentration of 250 μg / ml; the conditioned medium was the supernatant obtained by sterile filtration after inoculating the passage medium with CHO-K1 for 1 day; the cloning medium was 75% (v / v) EX-CELL CHO Cloning Medium, 20% (v / v) conditioned medium, and 5% (v / v) ClonaCell-CHO ACF Supplement supplemented with 1% GlutaMAX; the basal medium in the fed-batch medium was EX-CELL Advanced CHO Fed-batch Medium supplemented with 1% GlutaMAX (ThermoFisher), and the feed medium was Cell Boost 7a / 7b (HyClone).

[0151] 5. Preparation of BCA reagent:

[0152] ① Reagent A, 1 L: Weigh 10 g BCA (1%), 20 g Na2CO3·H2O (2%), 1.6 g Na2C4H4O6·H2O (0.16%), 4 g NaOH (0.4%), and 9.5 g NaHCO3 (0.95%), add water to 1 L, and adjust the pH to 11.25 with NaOH or solid NaHCO3.

[0153] ② Reagent B, 50ml: Take 2g CuSO4·5H2O (4%) and add distilled water to 50mL.

[0154] ③BCA reagent: Take 50 parts of reagent A and 1 part of reagent B and mix them evenly. This reagent is stable for one week.

[0155] Example 1

[0156] 1. Gene synthesis of Brazilian sweet:

[0157] A gene sequence search in GenBank revealed a brazin gene in Pentadiplandra brazzeana, with GenBank accession number P56552.1. The entire gene was synthesized by Nanjing GenScript and named Braz-pUC57. The full-length gene is 159 bp.

[0158] 2. Construction of Brazilian sweet Brz fusion protein:

[0159] Using the Braz-pUC57 gene as a template, PCR amplification was performed using primer F, primer R, and primer RC. The primers were synthesized by Suzhou Jinweizhi Biotechnology Co., Ltd.

[0160] Primer F (SEQ ID NO.57):

[0161] agcctccgggatccGCCACCATGAAGACTCTAGTGCTAC;

[0162] Primer R (SEQ ID NO.58):

[0163] aagttTTAgaattcTCAGTACTCGCAGTAGTCGC;

[0164] Primer RC (SEQ ID NO.59):

[0165] aagttTTAgaattcTCAatgatggtgatgatggtgatgatgGTACTCGCAGTAGTCGCAGATG.

[0166] PCR conditions were as follows: denaturation at 98°C for 2 minutes, followed by 30 cycles of denaturation at 98°C for 15 seconds, annealing at 55°C for 30 seconds, and extension at 72°C for 30 seconds, followed by incubation at 72°C for 5 minutes. The product amplified by primers F and R was named Brz, with a gene length of 159 bp. The carbon-terminally modified brazilian sweet fusion protein was named Brz-C1 (SEQ ID NO. 3), the product amplified by primers F and RC. The Brz-C1 gene is 183 bp long.

[0167] The PCR product was recovered from the gel, digested with EcoRI and BamHI, and ligated with the pDonor1.0 vector, which had been digested with the same enzymes. The product was then transformed into Escherichia coli DH5α, and transformants were selected for sequencing verification. Transformants that were confirmed to be correct by sequencing were transferred to LB+Amp liquid medium and cultured overnight at 37°C. Two plasmids were extracted separately: the natural brazilian tamarind protein expression plasmid Brz, with the natural brazilian tamarind protein amino acid sequence as SEQ ID NO.1 and the nucleotide sequence encoding it as SEQ ID NO.2. The plasmid map is shown in Figure 1; and the other is the fusion protein expression plasmid Brz-C1, with the fusion protein Brz-C1 amino acid sequence as SEQ ID NO.3 and the encoding nucleotide sequence as SEQ ID NO.4. The plasmid map is shown in Figure 2.

[0168] Example 2

[0169] Construction of CHO cell line of Brazilian sweet fusion protein:

[0170] 1. Plasmid construction:

[0171] The plasmid prepared in Example 1 was linearized with PvuI-HF restriction endonuclease, and the product was purified by ethanol recovery.

[0172] The linearized plasmids were co-transfected with Bxb-1 integrase into the stable, highly fluorescent GBB003 cell line by electroporation. The Bxb-1 integrase plasmid map is shown in Figure 3. Brasilien and brasilien fusion protein were integrated at position NW_003616785.1:83044 in the GBB003 cell genome. The green fluorescent protein gene (EGFP) was integrated at position NW_003616785.1:83044 in GBB003 cells. When brasilien or brasilien fusion protein was integrated into this position in GBB003 cells, the EGFP gene was knocked out, allowing cells expressing brasilien or brasilien fusion protein to be screened for positive results by fluorescence changes.

[0173] 2. Cell line construction:

[0174] 24-48 hours after electro-transfection, the cells were subjected to pressure screening in 96-well plates. The pressure screening medium was a passage medium supplemented with hygromycin at a final concentration of 200-300 μg / mL.

[0175] After 10-14 days, non-fluorescent cell clusters were selected and enriched using a fluorescence microscope. Cells were then expanded and cultured to obtain multiple cell pools. The cell expansion culture vessels were: 96-well plates → 24-well plates → 6-well plates → 50 mL shake tubes → 125 mL shake flasks. Culture conditions were 37°C, 5% CO2, saturated humidity, and a shake speed of 160 rpm for shake tubes and 120 rpm for shake flasks.

[0176] Select the 4-6 cell pools with the highest cell proliferation rate of 2.1-2.3 times / 24h and no fluorescence for fed-batch culture. The initial inoculation density is 5×10 5 The inoculation day was designated as D0. Cell counts and nutrient supplementation (3% Cell Boost 7a and 0.3% Cell Boost 7b) were performed starting on day 3 (D3). Thereafter, the above-mentioned procedures were performed every other day, with samples taken and counted on D5 (day 5), D7 (day 7), D9 (day 9), and D11 (day 11), and supplemented with 5% Cell Boost 7a and 0.5% Cell Boost 7b. The glucose concentration was maintained at 2-8 g / L throughout the fed-batch culture. Depending on the cell condition (viability ≥ 80%), the culture was terminated between D13 (day 13) and D16 (day 16). The cell supernatant was obtained by stepwise centrifugation, and the supernatant was subjected to downstream purification to obtain an aqueous solution of the brazilian sweet fusion protein.

[0177] The preparation of the stable high-fluorescence cell GBB003 used in this example mainly includes the following steps:

[0178] (1) The green fluorescent protein gene (EGFP) was used as a marker gene for screening, and the attP sequence was used as a homology arm to construct a recombinant plasmid containing RMCE. After linearizing the constructed plasmid, the linear DNA was purified and recovered. CHO-K1 cells were taken, centrifuged, and the supernatant was discarded. The cells were resuspended in 100 μL of R Buffer, a resuspension buffer specifically for cell electroporation, and electroporated three times.

[0179] (2) Pressurization: the pressurization reagent is G418, the pressurization concentration is 800 μg / ml, and the cells are divided into minipools (i.e., 96-well plates) for culture.

[0180] (3) Check the plate after ten days. When a large number of fluorescent cell clusters are observed, they can be enriched. The one-well-to-one-well principle should be followed during enrichment.

[0181] (4) Observe the growth of enriched cells at any time and expand them when the coverage rate reaches more than 50%.

[0182] (5) After expansion to a shake flask, the cells were passaged three times to form a stable cell pool. The formation of monoclonal cells and screening of stable fluorescence were achieved through artificial intelligence. The specific steps are shown in ah.

[0183] a) The cells in the stable cell pool are diluted to a certain concentration and then inoculated into a culture dish containing a semi-solid culture medium. During the inoculation process, the cells are ensured to be approximately evenly distributed in all positions of the culture dish. The cells are allowed to stand for about half an hour to allow the cells to settle to the bottom of the culture dish.

[0184] b) The culture dish is transferred to the electron microscope stage, and the electron microscope performs high-throughput scanning on the cells in the culture dish;

[0185] c) The electron microscope scanned image is uploaded to the server for artificial intelligence image analysis. The analysis process includes monoclonal cell line detection, protein expression level prediction of monoclonal cell lines, protein expression level ranking, and coding and localization of screened high-protein expressing cell lines. The protein expression level prediction of the monoclonal cell line can be based on fluorescence or not. The following step c is protein expression level prediction not based on fluorescence.

[0186] Step c is based on the detection of monoclonal cell line targets using image processing technology. In this embodiment, the target detection algorithm of YOLOv8 is adopted, and the actual detection effect reaches mAP 94.2%. The target detection model for monoclonal cell lines in step c is consistent with the commonly used deep learning target detection model, so it is not described in detail. Specifically, the cell image needs to be annotated first. In order to improve the prediction accuracy of the algorithm, the bounding boxes of all monoclonal cell lines and adhesion cell lines will be marked and used as the real target bounding box (ground truth) for the loss calculation of the model output. After algorithm training, the model learns the ability to extract the border information of monoclonal cell lines. In the actual application scenario step c, the trained model can predict the border of the monoclonal cell line in the image.

[0187] Step c predicts the fluorescent protein expression level based on monoclonal cell images. In this example, the SqueezeNet deep learning network and the MSE loss function are used. Since the predicted protein expression levels in this project are ranked, the ranking result is the ultimate goal of the algorithm, so the NDCG evaluation standard of the ranking algorithm is used here. The specific calculation formula is as follows:

[0188] Where IDCG = best-ranked DCG. Specifically in this example, cell imaging predicted expression levels, sorted by predicted expression levels, and compared with the actual fluorescence value sorting results, NDCG = 0.89. The expression level prediction model in this example is described in detail in CN112037862B (application number CN202311058132.5, invention title "A Cell Transfer Method and Cell Transfer System").

[0189] d) The coding and location information of the screened protein high-expressing cell lines are returned to the robot control software;

[0190] The cell codes and positions returned in step d, in this embodiment, return the top 100 cells in terms of predicted expression levels, with the codes ranging from 1 to 100. The position information includes the coordinates in the plane coordinate system relative to the center of the microscope (the depth of the culture medium is not considered for the time being, because all the cells to be selected are deposited to the bottom of the culture dish. In addition, during the photography process, cells that have not settled to the bottom of the culture dish will be out of focus, resulting in blurred cells, and cells with blurred images will also be excluded).

[0191] e) The robot control software automatically operates (or manually assists) the robotic arm to aspirate the screened monoclonal cell lines and transfer them to the designated well plate. This process is repeated until all high-expressing cell lines are transferred. For the suction and transfer process, please refer to CN113821287B (application number: CN202111040555.5, invention name: "Robot-based cell manipulation task processing method, device, equipment and medium"), CN113403431B (application number: CN202110735660.4, invention name: "Robot-based cell liquid collection control method, device, equipment and storage medium"), CN113771030A (application number: CN202111040562.5, invention name: "Cell manipulation robot control method, device, equipment and storage medium"), CN113733087B (application number: CN202111039293.0, invention name: "Cell manipulation robot control information configuration method, device, equipment and medium")

[0192] f) After the cell clones in the well plate are cultured to a certain number, they are transferred to a larger well plate for culture, and finally expanded to shake flask culture, and the cell expression level is detected to determine whether the cell line selected by artificial intelligence is a high-yield cell line; step f) in this embodiment is to transfer to a 96-well plate for amplification culture, and then transfer to shake flask culture.

[0193] g) The selected high-yield cell lines are subcultured. Before each generation of cell lines is subcultured, a portion of the sample is diluted and placed in a culture dish, and images are taken using an electron microscope. Images of the cell lines are taken continuously for several generations. Currently, the expression of the selected cells can be predicted by the model learned from the cell morphology using photography, and the expression of the selected cells can also be predicted by fluorescent labeling.

[0194] Step g) Photographing Monoclonal Cells In this example, a Thermo Fisher Scientific M7000 electron microscope was used to photograph the cells.

[0195] h) The collected images of several generations of cell lines are input into an artificial intelligence algorithm to predict the cell line's transgenerational stability based on fluorescence intensity and / or protein expression, and cell lines with high protein expression characteristics that can be stably transgenerated are selected as the final candidate cell lines; the prediction of the cell line's transgenerational stability can be based on fluorescence or on cell morphology rather than fluorescence. The following h) step is the prediction of protein expression based on cell morphology rather than fluorescence.

[0196] Step h) predicting the cell line stability, including but not limited to traditional image recognition algorithms using histograms to extract image features and make predictions, using deep learning neural networks to automatically extract image features and make predictions, and other image-based technical predictions. In this embodiment, a deep learning-based method is used to automatically extract image features of stable cell lines and unstable cell lines and predict their stability, with an accuracy rate of 84% for predictions of different cell lines. See CN114417582A (application number: CN202210010493.1, invention name "Cell line stability prediction method, device, computer equipment and storage medium")

[0197] Wherein, step d) and step e) can be completed manually or automatically by a robot arm. In this embodiment, they are completed automatically by a robot arm.

[0198] (6) CHO-K1 cells were used as control and the seeding density was 5×10 5 cells / ml, inoculate 30 ml of the system, count the cells on the day of inoculation, culture day 3, culture day 5, culture day 7, culture day 9, culture day 11, culture day 13, and culture day 14, and screen out monoclones with a growth rate lower than that of CHO-K1.

[0199] (7) Recover the remaining monoclonal clones and perform stable subculture after recovery. Subculture should be performed every 3 or 4 days. The cell density of the 4-day subculture is 3×10 5 cells / ml, and the cell density after 3 days was 5×10 5cells / ml, and the fluorescence of the cell line was regularly monitored during subculture. The stability of the high-fluorescence cell line was studied for approximately 90 days using the subculture medium. GBB003 cells, which exhibited minimal fluorescence fluctuations and stable cell line growth, were confirmed to be stable high-fluorescence cells. NGS sequencing of the integration site of this cell line revealed the highly expressed fragment sequence located at the integration site as SEQ ID NO. 49. Specifically, the annotation information for the integration site on the CHO gene is: NW_003616785.1:83044.

[0200] Example 3

[0201] Determination of the content of brazilian melamine and brazilian fusion protein by BCA method

[0202] 1. Standard protein solution: Weigh 0.5 g of bovine serum albumin and dissolve it in distilled water to a volume of 100 ml to make a 5 mg / mL solution. Dilute tenfold before use.

[0203] 2. Drawing of standard curve:

[0204] Take a 96-well ELISA plate and add reagents according to Table 1.

[0205] Table 1

[0206] 3. Sample Measurement

[0207] After adding the above reagents, accurately pipette 20 μl of sample solution into the microplate well. Add 200 μl of BCA reagent, gently shake, and incubate at 37°C for 30-60 minutes. After cooling to room temperature, plot a standard curve using a blank as a control, comparing color at 562 nm on a microplate reader with bovine serum albumin content as the abscissa and absorbance as the ordinate. Using the standard curve blank as a control, determine the protein content of the sample from the standard curve based on the sample absorbance: 0.841 mg / mL for Brz and 0.624 mg / mL for Brz-C1. See Table 2 for details.

[0208] Table 2

[0209] Example 4

[0210] Sweetness detection method of natural Brazilian sweet protein and Brazilian sweet fusion protein

[0211] The sweetness of sweet proteins was determined using a blind taste test. In the blind taste test, a control group used sucrose aqueous solutions of varying concentrations, each numbered. The test samples of brazilian sweet (natural brazilian sweet protein and brazilian sweet fusion protein prepared in Example 2) were prepared into aqueous solutions of protein at concentrations ranging from 0.08 mg / mL to 0.02 mg / mL. A panel of seven to eight people tasted the sucrose aqueous solutions of varying concentrations and the samples, comparing them pairwise until they found the number of the sucrose aqueous solutions of varying concentrations that had the same or similar sweetness as the brazilian sweet solution (natural brazilian sweet protein or brazilian sweet fusion protein prepared in Example 2). The sweetness of the sample was then calculated.

[0212] Formula: relative sucrose sweetness multiple = sucrose concentration ÷ brasilien (natural brasilien or brasilien fusion protein prepared in Example 2) concentration

[0213] Preparation of sucrose concentration gradient samples: First, weigh 100g of sucrose (CAS No.: 57-50-1, Sigma) and dilute it to 200ml with distilled water to create a 500mg / mL sucrose solution. A certain volume of the mother solution was then measured and further diluted with different volumes of distilled water to create 12 sets of sucrose aqueous solutions of varying concentrations, all with a volume of 30mL. See Table 3 for details. The sensory test results of brazilian tamarind with different configurations are detailed in Table 4.

[0214] Table 3 Preparation of gradient sucrose aqueous solution and relative sucrose sweetness multiples of samples

[0215] Table 4 Test results of sweetness of Brazilian sweet fusion protein with different configurations

[0216] Example 5

[0217] Using the gene of the natural brazilian thauma protein expression plasmid Brz as a template, Nanjing GenScript performed full gene synthesis and constructed six groups of carbon-terminally optimized fusion proteins, named Brz-C2, Brz-C3, Brz-C4, Brz-C5, Brz-C6, and Brz-C7. The fusion proteins were composed of the natural brazilian thauma protein (first unit) with the amino acid sequence of SEQ ID NO. 1 and the corresponding second unit as shown in Table 5, with the carbon terminus of the first unit connected to the nitrogen terminus of the second unit. As shown in Table 5:

[0218] Table 5 Structure of Brazilian sweet fusion protein with optimized carbon terminal configuration

[0219] The six fusion protein CHO cell lines were constructed using the method described in Example 2. After fed-batch culture, each cell line produced a protein aqueous solution of a certain concentration. Sensory testing showed that Brz-C2 had a sweetness approximately 9,700 times that of standard sucrose, Brz-C3 approximately 14,500 times, and Brz-C4 had the highest sweetness, approximately 16,800 times that of an equal mass of sucrose. The remaining Brz-C5, Brz-C6, and Brz-C7 had sweetnesses of approximately 10,500, 10,300, and 9,700 times that of sucrose, respectively (see Table 6 for details). This represents a 9.5-16.5-fold improvement over the pre-optimized natural brassin, achieving unexpected technical results.

[0220] Table 6 Sweetness test results of Brazilian sweet fusion protein with optimized carbon end configuration

[0221] Example 6

[0222] Using the gene of the natural brazilian thauma protein expression plasmid Brz as a template, Nanjing GenScript performed full gene synthesis to construct three groups of nitrogen-terminally optimized fusion proteins, named N1-Brz, N2-Brz, and N3-Brz. The fusion proteins are composed of the natural brazilian thauma protein (first unit) with the amino acid sequence of SEQ ID NO.1 and the corresponding third unit as shown in Table 7, with the nitrogen terminus of the first unit linked to the carbon terminus of the third unit. As shown in Table 7:

[0223] Table 7 Structure of Brazilian sweet fusion protein with optimized nitrogen terminal configuration

[0224] CHO cell lines expressing the three fusion proteins were constructed using the method described in Example 2. Following the completion of fed-batch cell culture, aqueous solutions of thaumatin at varying concentrations were obtained through downstream purification processes. Sensory evaluation revealed that the sweetness of N1-Brz, N2-Brz, and N3-Brz was approximately 4100-, 7100-, and 3700-fold, respectively, that of standard sucrose of equal quality (see Table 8). This represents a 3.7- to 7.1-fold increase compared to natural brassica Brz, but the sweetness enhancement was not as significant as that achieved by the C-terminus-optimized group described in Example 5.

[0225] Table 8 Sweetness test results of Brazilian sweet fusion protein with optimized nitrogen-terminal and double-terminal configurations

[0226] Example 7

[0227] Using the gene of the natural Brazilian sweet protein expression plasmid Brz as a template, Nanjing KingSher Company performed full gene synthesis to construct a fusion protein with optimized ends (nitrogen end and carbon end), named N-Brz-C. The fusion protein N-Brz-C consists of the natural Brazilian sweet protein (first unit) with an amino acid sequence of SEQ ID NO.1 and the corresponding second and third units as shown in Table 9. The carbon end of the first unit is connected to the nitrogen end of the second unit, and the nitrogen end of the first unit is connected to the carbon end of the third unit. As shown in Table 9. According to the method described in Example 2, a CHO cell line of this protein was constructed, and the supernatant obtained by cell culture and the protein aqueous solution after purification and recovery were subjected to sensory testing. The results showed that the sweetness of N-Brz-C was about 10600 times that of the same quality standard sucrose, as shown in Table 10. The sweetness of natural Brazilian sweet Brz was increased by 10.5 times, which is the median sweetness of Brz-C1 and N2-Brz.

[0228] Table 9 Structure of Brazilian sweet fusion protein with double-end optimized configuration

[0229] Table 10 Sweetness test results of Brazilian sweet fusion protein with double-ended optimized structure

[0230] Comparative Example 1

[0231] Using the gene of the natural brazilian thauma protein expression plasmid Brz as a template, Nanjing GenScript Company performed full gene synthesis to construct three groups of fusion proteins, named Brz-K1, Brz-K2, and Brz-K3. The fusion proteins are composed of the natural brazilian thauma protein (first unit) with the amino acid sequence of SEQ ID NO.1 and the corresponding second unit as shown in Table 11. In addition to histidine (His), the second unit also contains aspartic acid (Asp) and lysine (Lys). The carbon end of the first unit is connected to the nitrogen end of the second unit. As shown in Table 11:

[0232] Table 11 Structure of Brazilian sweet fusion protein with optimized carbon terminal configuration

[0233] Following the method described in Example 2, CHO cell lines expressing these three protein groups were constructed. Sensory testing was performed on the supernatants obtained from cell culture and the purified, recovered protein aqueous solutions. The results showed that the sweetness of the three fusion proteins was approximately 600-1000 times that of standard sucrose of equal mass, and was, to varying degrees, lower than the sweetness of natural Bruzet (Brz), as detailed in Table 12. Facts have shown that adding varying lengths of histidine (His) to both ends of the natural Bruzet sequence enhances the sweetness of the sweet proteins to varying degrees; however, the addition of other amino acids (such as Asp and Lys) can negate the enhanced sweetness or even reduce it to a lesser sweetness than the unoptimized version.

[0234] Table 12 Sweetness test results of three kinds of Brazilian sweet fusion protein in Example 8

[0235] In the above examples, Chinese hamster ovary (CHO) cells were used as the protein expression host. The optimized brazzein gene sequence was cloned into the pDonor1.0 vector and 6-8 histidine residues were modified at the C-terminus of the sequence to form Brz-C1 (brazilian fusion protein). Experiments demonstrated that regardless of whether the C-terminus of the sweet protein sequence was optimized, the brazzein secreted by the CHO cells folded correctly, functioned normally, and exhibited a sweet taste. Sweetness sensory evaluation determined that the novel brazzein fusion protein with optimized C-terminus was at least 3.5-16 times sweeter than native brazzein and significantly improved the delayed sweetness sensation upon ingestion.

[0236] Because sensory testers have a threshold for sensory perception in sensory experiments, the sweetness of the brasiliensis fusion protein is too high. Therefore, in this example, the concentrations of natural brasiliensis and sucrose are adjusted to be consistent, and the concentrations of the brasiliensis fusion protein and sucrose are adjusted to be consistent, and the sweetness multiple of brasiliensis relative to sucrose is measured as X. Furthermore, the histidine fragment in the fusion protein can be used for affinity purification, which facilitates large-scale production.

[0237] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A fusion protein, characterized in that The amino acid sequence of the fusion protein comprises a first unit and a second unit, wherein the carbon end of the first unit is connected to the nitrogen end of the second unit; The amino acid sequence of the first unit is shown in SEQ ID NO.1 or SEQ ID NO.25; The second unit comprises a fragment consisting of 2-22 amino acid residues, and each amino acid residue is independently selected from histidine His, lysine Lys or aspartic acid Asp.

2. The fusion protein according to claim 1, characterized in that The second unit is selected from a fragment consisting of 2, 4, 6, 8, 10, 12 or 16 histidines.

3. The fusion protein according to claim 2, characterized in that The second unit is a fragment consisting of 6 histidines.

4. The fusion protein according to claim 1, characterized in that The amino acid sequence of the fusion protein is selected from SEQ ID NO.3, SEQ ID NO.5, SEQ ID NO.7, SEQ ID NO.9, SEQ ID NO.11, SEQ ID NO.13, SEQ ID NO.15, SEQ ID NO.27, SEQ ID NO.29, SEQ ID NO.31, SEQ ID NO.33, SEQ ID NO.35, SEQ ID NO.37, SEQ ID NO.39, SEQ ID NO.60, SEQ ID NO.62, SEQ ID NO.64, SEQ ID NO.66, SEQ ID NO.68 or SEQ ID NO.

70.

5. A fusion protein, characterized in that The amino acid sequence of the fusion protein contains a first unit and a third unit, and the nitrogen end of the first unit is connected to the carbon end of the third unit; The amino acid sequence of the first unit is shown in SEQ ID NO.1 or SEQ ID NO.25; The third unit comprises a fragment consisting of 2-22 amino acid residues, and each amino acid residue is independently selected from histidine His, lysine Lys or aspartic acid Asp.

6. The fusion protein according to claim 5, characterized in that The third unit is a fragment consisting of 2, 8 or 16 histidines.

7. The fusion protein according to claim 5, characterized in that The amino acid sequence of the fusion protein is selected from SEQ ID NO.17, SEQ ID NO.19, SEQ ID NO.21, SEQ ID NO.41, SEQ ID NO.43 or SEQ ID NO.

45.

8. A fusion protein, characterized in that The amino acid sequence of the fusion protein contains a first unit, a second unit and a third unit, the carbon end of the first unit is connected to the nitrogen end of the second unit, and the nitrogen end of the first unit is connected to the carbon end of the third unit; The amino acid sequence of the first unit is shown in SEQ ID NO.1 or SEQ ID NO.25; The second unit comprises a fragment consisting of 2-22 amino acid residues, each of which is independently selected from histidine His, lysine Lys or aspartic acid Asp; The third unit comprises a fragment consisting of 2-22 amino acid residues, and each amino acid residue is independently selected from histidine His, lysine Lys or aspartic acid Asp.

9. The fusion protein according to claim 8, characterized in that The second unit and the third unit are independently fragments consisting of 2-16 histidines.

10. The fusion protein according to claim 9, characterized in that The second unit is a fragment consisting of 2, 4, 6, 8, 10, 12 or 16 histidines, and the third unit is a fragment consisting of 2, 8 or 16 histidines.

11. The fusion protein according to claim 10, characterized in that The second unit is a fragment consisting of 6 histidines; or the second unit and the third unit are both fragments consisting of 8 histidines.

12. The fusion protein according to claim 8, characterized in that The amino acid sequence of the fusion protein is selected from SEQ ID NO.23 or SEQ ID NO.47, SEQ ID NO.68 or SEQ ID NO.

70.

13. A polynucleotide, characterized in that The polynucleotide encodes the fusion protein according to any one of claims 1 to 10.

14. A carrier, characterized in that The vector carries the polynucleotide of claim 13.

15. A cell, characterized in that The cell expresses the fusion protein according to any one of claims 1 to 12, or carries the polynucleotide according to claim 14, or contains the vector according to claim 14.

16. The method for preparing the fusion protein according to any one of claims 1 to 12, characterized in that: The method comprises the following steps: expressing the fusion protein in the cell according to claim 15.

17. Use of the fusion protein according to any one of claims 1 to 12, the polynucleotide according to claim 13, the vector according to claim 14 or the cell according to claim 15 in the preparation of food additives, sweeteners, medicines and / or feeds.

Citation Information

Patent Citations

  • High-sweetness sweet protein gene and synthesis method thereof

    CN101570754A

  • Method of producing a sweet protein

    CN102498126A

  • Taste and flavor-modifier proteins

    CN112313244A