Method for biosynthesizing type IV collagen, a structural material for the human body.
Recombinant human type IV collagen produced in E. coli addresses the limitations of conventional collagen production by ensuring high yields and safety, with enhanced cell adhesion activity and scalability.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SHANXI JINBO BIO PHARMACEUTICAL CO LTD
- Filing Date
- 2023-12-21
- Publication Date
- 2026-05-08
AI Technical Summary
Conventional methods for producing collagen result in materials that have lost their biological activity, exhibit poor water solubility, and pose risks of viral infection and sensitization, while recombinant expression methods are costly and not scalable.
Development of recombinant human type IV collagen expressed in E. coli, with precise purification and high yields, utilizing amino acid sequences with varying degrees of identity and mutations, and production methods involving culture, harvesting, and purification steps.
The recombinant collagen demonstrates higher cell adhesion activity and is suitable for large-scale production, maintaining biological activity and safety.
Smart Images

Figure 0007855803000005 
Figure 0007855803000006 
Figure 0007855803000007
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the priority of a Chinese patent application with an application date of October 25, 2023, an application number of 202311391711.1, and an invention title of "Manufacturing Method for Biosynthesizing Human Structural Material Type IV Collagen".
[0002] The present invention belongs to the technical field of biopharmaceuticals and relates to recombinant type IV humanized collagen, its manufacturing method, and uses.
Background Art
[0003] Collagen (Collagen, COL), abbreviated as "collagen" in Chinese, is a helical fibrous functional protein composed of three peptide chains. It is also a main component of the extracellular matrix, with rich content and a wide coverage range. Collagen in the human body accounts for 25% - 30% of the total protein content, mainly present in the skin, tendon, and bone, playing an important role in protecting and connecting each tissue, and exerting important physiological functions in the body. Type IV collagen is the main protein that constitutes the basement membrane together with laminin, and is composed of trimers with a helical structure composed of three α - chains as the structural units. Although the subunits that make up the trimer vary depending on the tissue, in many tissues such as the liver, the trimer is composed of two α1 - chains and one α2 - chain. The type IV collagen molecule has structures unique to the type IV collagen molecule at both ends of the TH domain that forms the above - mentioned helical structure. The structure at the N - terminal is called the 7S domain, and the structure at the C - terminal is called the NC1 domain.
[0004] Collagen has good biocompatibility, degradability, and low antigenicity. Due to its unique biological structure, it has become the focus of recent research and is widely applied in many fields such as biomedicine, cosmetics, health foods, and foods.
[0005] The human body contains 28 different types of collagen, which can be classified into two types, fibrous and non-fibrous, depending on whether their structure is fibrous or not. Fibrous collagen primarily functions as a cell scaffold, fixing cell position, exerting an anchoring effect, and providing tensile strength and rigidity to tissues. Non-fibrous collagen is further subdivided into basal collagen, short-chain collagen, transmembrane collagen, etc., each performing different functions.
[0006] Conventional methods for producing collagen involve purifying animal-derived tissues using acid, alkali, and enzymatic hydrolysis to extract collagen derivatives. However, collagen obtained by these methods has already lost its original biological activity, has poor water solubility, does not readily bind to the human body, carries the risk of viral infection and sensitization, and cannot exhibit its true function. Some research institutions express human-derived collagen in vitro using conventional recombinant expression methods, but this is costly, has a long production cycle, and cannot be performed on a large scale.
[0007] Therefore, there is an urgent need in the market for a collagen material that possesses excellent biomaterial properties, has an amino acid sequence highly homologous to that of the human body, and can be mass-produced in industrial systems. [Overview of the project]
[0008] The inventors conducted a large-scale functional domain screening of human type IV collagen and discovered 11 recombinant collagens. These recombinant collagens can be expressed in E. coli and purified. Furthermore, the inventors found that these recombinant collagens have high yields, good purity of the target protein after precise purification, and higher activity than the positive control (bovine type I collagen) in the measurement of cell adhesion activity.
[0009] In one embodiment, the present invention provides recombinant collagen comprising one or more repeating units, the repeating units being linked directly or via linkers, the repeating units comprising an amino acid sequence or variant thereof selected from the group consisting of SEQ ID NO: 1 (Gakgdkgskgevgfpglagspgipgskgeq) or 28 (Gptgpagqkgepgsdgipgsagekgepglp), wherein the variant is (1) an amino acid sequence in which one or more amino acid residues are mutated in the amino acid sequence, or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence.
[0010] In one embodiment, the number of repeating units is 2 to 50, for example, 2 to 45, 2 to 40, 2 to 35, 2 to 30, 2 to 25, 2 to 20, 2 to 15, 4 to 10, or 6 to 10. For example, the number of repeating units is 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 16, 18, 20, 30, 40, 50, or a range in between.
[0011] In one embodiment, the linker comprises one or more amino acid residues, for example, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 to 2 amino acid residues.
[0012] In one embodiment, the mutation is selected from substitution, addition, insertion, or deletion.
[0013] In one embodiment, the substitution is a conservative amino acid substitution.
[0014] In one embodiment, the recombinant collagen is recombinant human type IV collagen or recombinant humanized type IV collagen.
[0015] In one embodiment, recombinant collagen has cell adhesion activity.
[0016] In one embodiment, the mutant of SEQ ID NO:1 includes mutations that add Gfpgfp (SEQ ID NO:34) to the N-terminus of the amino acid sequence of SEQ ID NO:1, or shorten a fragment of 1 to 5 amino acid residues to the N-terminus of SEQ ID NO:34, and / or add GFMGPPGPQGQPGLP (SEQ ID NO:35) to the C-terminus of the amino acid sequence of SEQ ID NO:1, or shorten a fragment of 1 to 14 amino acid residues to the C-terminus of SEQ ID NO:35, or shorten 1 to 5 amino acid residues to the C-terminus of the amino acid sequence of SEQ ID NO:1.
[0017] In one embodiment, the variant of SEQ ID NO:28 includes mutations in which Glpgtp (SEQ ID NO:36) is added to the N-terminus of the amino acid sequence of SEQ ID NO:28, or in which a fragment is shortened to a length of 1 to 5 amino acid residues at the N-terminus.
[0018] In one embodiment, the variants of SEQ ID NO:1 are SEQ ID NO:4 (Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq), 7 (GFPGFPGAKGDKGSKGEVGFPGLAGSPGIPGSKGEQGFMGPPGPQGQPGLP), 10 (Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq), and 13 (Gfpgfpgakgdkgskgevgfpglagspgipg Selected from skgeqgfm), 16(Gfpgfpgakgdkgskgevgfpglagspgipgskgeq), 19(Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp), 22(Gakgdkgskgevgfpglagspgipgskgeqgfm), or 31(Gfpgfpgakgdkgskgevgfpglagspgipgsk).
[0019] In one embodiment, the variant of SEQ ID NO:28 is SEQ ID NO:25 (Glpgtpgptgpagqkgepgsdgipgsagekgepglp).
[0020] In one embodiment, recombinant collagen comprises an amino acid sequence or a variant thereof selected from the group consisting of SEQ ID NO: 2, 5, 8, 11, 14, 17, 20, 23, 26, 29, or 32, wherein the variant is (1) an amino acid sequence in which one or more amino acid residues are mutated in the amino acid sequence, or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence.
[0021] In one embodiment, the mutation is selected from substitution, addition, insertion, or deletion. In one embodiment, the substitution is a conservative amino acid substitution.
[0022] In another embodiment, nucleic acids encoding recombinant collagen as described herein are provided. In one embodiment, the nucleic acid comprises a codon-optimized nucleotide sequence. In one embodiment, the nucleotide sequence is a nucleotide sequence that codon-optimizes expression in eukaryotic or prokaryotic host cells, such as yeast or E. coli. In one embodiment, the nucleic acid comprises a nucleotide sequence selected from the group consisting of SEQ ID NO: 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, or 33.
[0023] In another embodiment, a vector comprising the nucleic acid described herein is provided. In one embodiment, the vector comprises an expression regulatory element operably linked to the nucleic acid, a nucleotide of a purified tag, and / or a nucleotide of a leader sequence. In one embodiment, the expression regulatory element is selected from a promoter, a terminator, or an enhancer. In one embodiment, the purified tag is selected from a His tag, a GST tag, an MBP tag, a SUMO tag, or a NusA tag. In one embodiment, the vector is an expression vector or a clone vector, preferably pET-28a(+). pET-28a(+) may include an N-terminal His tag, a Thrombin tag, and a T7 protein tag, as well as a C-terminal His tag. In this specification, recombinant collagen may include an enzymatic cleavage site at the N-terminus to facilitate purification.
[0024] In another embodiment, a host cell comprising the nucleic acid or vector described herein is provided. In one embodiment, the host cell is a eukaryotic cell or a prokaryotic cell. In one embodiment, the eukaryotic cell is a yeast cell, an animal cell and / or an insect cell, and in one embodiment, the prokaryotic cell is an Escherichia coli cell, for example, Escherichia coli BL21.
[0025] In another embodiment, a composition comprising one or more of the recombinant collagen, nucleic acids, vectors, and host cells described herein is provided. In one embodiment, the composition is a kit. In one embodiment, the composition is one or more of the following: biomimetic materials, human biomimetic materials, cosmetic materials, organoid culture materials, cardiovascular stents, coating materials, tissue injection filling materials, ophthalmic materials, obstetric and gynecological biomaterials, nerve repair and regeneration materials, liver tissue materials and vascular repair and regeneration materials, 3D printed artificial organ biomaterials, cosmetic raw materials, medicinal auxiliary materials, and food additives. In one embodiment, the composition is a surface composition, an injection composition, or an oral composition. In one embodiment, the composition is a solution, a lyophilized powder, a gel, a sponge, or a fibrous composition.
[0026] In another aspect, provided is the use of the recombinant collagen, nucleic acid, vector, host cell and / or composition herein in one or more of a biological coating material, a human biomimetic material, a plastic and aesthetic material, an organoid culture material, a cardiovascular stent, a coating material, a tissue injection filling material, an ophthalmic material, a gynaecological and obstetric biological material, a nerve repair and regeneration material, a liver tissue material and a blood vessel repair and regeneration material, a 3D printed artificial organ biological material, a cosmetic raw material, a pharmaceutical adjuvant material and a food additive.
[0027] In another aspect, provided is a method for promoting cell adhesion, which includes the step of contacting a cell with the recombinant collagen, nucleic acid, vector, host cell and / or composition herein. In one embodiment, the cell is an animal cell. The animal cell may be a mammalian cell or a human cell. Also provided is the use of the recombinant collagen, nucleic acid, vector, host cell and / or composition herein in the manufacture of a kit, wherein the drug is for promoting cell adhesion.
[0028] In another aspect, a cosmetic method includes the step of administering the recombinant collagen described herein to a subject. Preferably, the administration is topical administration, oral administration or injection administration. Preferably, the subject is a human.
[0029] In another aspect, step (1) of culturing the host cell described herein under appropriate culture conditions, step (2) of harvesting the host cell and / or medium containing recombinant collagen, step (3) of purifying the recombinant collagen, are provided, which is a method for producing the recombinant collagen described herein.
[0030] In one embodiment, the host cell is an Escherichia coli cell, preferably an Escherichia coli BL21(DE3) cell.
[0031] In one embodiment, step (1) includes culturing the Escherichia coli cell in LB medium and inducing expression with IPTG.
[0032] In one embodiment, step (2) includes harvesting the E. coli cells, resuspending them in an equilibrium working solution, homogenizing the E. coli cells, preferably under high pressure, and separating the supernatant. In one embodiment, the equilibrium working solution contains 100-500 mM sodium chloride, 10-50 mM Tris, 10-50 mM imidazole, and has a pH of 7-9.
[0033] In one embodiment, step (3) includes crude purification, enzymatic cleavage, fine purification and / or reverse-phase nickel column purification.
[0034] In one embodiment, crude purification includes the step of purifying the supernatant using a Ni-agarose gel column to obtain an eluent containing the target protein, wherein the eluent contains 100-500 mM sodium chloride, 10-50 mM Tris, and 100-500 mM imidazole, and preferably has a pH of 7-9.
[0035] In one embodiment, the precision purification includes the step of gradient eluting an eluent containing the target protein using a strong anion exchange chromatography column. In one embodiment, the gradient elution includes eluting with 0-15% solution B for 1-5 minutes and holding three column volumes, eluting with 15-30% solution B for 1-5 minutes and holding three column volumes, eluting with 30-50% solution B for 1-5 minutes and holding three column volumes, and eluting with 50-100% solution B for 1-5 minutes and holding three column volumes, wherein solution B contains 10-50 mM Tris, 0.5-5 M sodium chloride, and has a pH of 7-9. In one embodiment, step (3) includes purifying recombinant collagen with a purification column, e.g., a nickel column, and / or cleaving the recombinant collagen with a collagen tool enzyme.
[0036] The advantages of the present invention include the following: 1. The recombinant collagen of the present invention is derived from human type IV collagen and is recombinant humanized type IV collagen. 2. The recombinant collagen of the present invention is suitable for production using E. coli and can be isolated and purified. 3. The recombinant collagen of the present invention has a higher yield and is suitable for subsequent purification (purification by Ni column or strong anion column). 4. The recombinant collagen of the present invention has cell adhesion activity. The recombinant collagen of the present invention (e.g., C4P7Ch) has higher cell adhesion activity compared to the positive control. [Brief explanation of the drawing]
[0037] [Figure 1] The electrophoretic detection results for C4P7Ca are shown. [Figure 2] The electrophoretic detection results for C4P7Cb are shown. [Figure 3] The electrophoretic detection results for C4P7Cc are shown. [Figure 4] The electrophoretic detection results for C4P7Cd are shown. [Figure 5] The electrophoretic detection results for C4P7Ce are shown. [Figure 6] The electrophoretic detection results for C4P7Cf are shown. [Figure 7] The electrophoretic detection results for C4P7Cg are shown. [Figure 8] The electrophoretic detection results for C4P7Ch are shown. [Figure 9] The electrophoretic detection results for C4P7Ea are shown. [Figure 10] The electrophoretic detection results for C4P7Eb are shown. [Figure 11] The electrophoretic detection results for C4P7Ec are shown. [Figure 12] This demonstrates the effect of collagen C4P7Ch on cell adhesion. [Figure 13] This demonstrates the effect of collagen C4P7Cf on cell adhesion. [Modes for carrying out the invention]
[0038] As used herein, “recombinant collagen” refers to an amino acid sequence or fragment thereof encoded by a specific, designed and modified gene, produced by DNA recombination technology, or a combination of such functional amino acid sequence fragments. The gene coding sequence or amino acid sequence of recombinant collagen may have low homology to the gene coding sequence or amino acid sequence of human collagen. Recombinant humanized collagen is a fragment of a full-length or partial amino acid sequence encoded by a specific type of human collagen gene, produced by DNA recombination technology, or a combination of functional fragments containing human collagen. In this specification, recombinant collagen comprises one or more repeating units, which may be derived from human type IV collagen. Therefore, recombinant collagen may be recombinant type IV humanized collagen. The repeating units may be linked by linkers, which may be native amino acid residues in the human type IV collagen of the repeating units, for example, 1 to 50 amino acid residues. The repeating units may have SEQ ID NO: 1, 4, 7, 10, 13, 16, 19, 22, 25, 28, 31, or 33. Recombinant collagen may also have SEQ ID NO: 2, 5, 8, 11, 14, 17, 20, 23, 26, 29, or 32.
[0039] As used herein, the term “mutant” means recombinant collagen having cell adhesion activity and containing changes / mutations (i.e., substitutions, additions, insertions and / or deletions) at one or more positions. Substitution means replacing an amino acid occupying a position with a different amino acid, deletion means removing an amino acid occupying a position, and insertion means adding an amino acid adjacent to or immediately following an amino acid occupying a position. Addition means adding one or more amino acid residues to the C-terminus and / or N-terminus of an amino acid sequence. Substitutions may be conservative substitutions. A mutant of a repeating unit may be a sequence after one or more amino acid residues have been changed or mutated (i.e., substituted, added, inserted and / or deleted) at SEQ ID NO: 1, 4, 7, 10, 13, 16, 19, 22, 25, 28, 31 or 33. Recombinant collagen variants may be sequences in which one or more amino acid residues have been altered or mutated (i.e., substituted, added, inserted, and / or deleted) in SEQ ID NO: 2, 5, 8, 11, 14, 17, 20, 23, 26, 29, or 32.
[0040] For example, a variant of the repeating unit of SEQ ID NO:1 may include a mutation that adds Gfpgfp (SEQ ID NO:34) to the N-terminus of the amino acid sequence of SEQ ID NO:1 or shortens a fragment at the N-terminus of SEQ ID NO:34 (i.e., shortens 1 to 5 amino acid residues from the N-terminus of SEQ ID NO:34, for example, 1, 2, 3, 4, or 5 amino acid residues corresponding to fpgfp, pgfp, gfp, fp, and p residues, respectively), and / or a variant that adds GFMGPPGPQGQPGLP (SEQ ID NO:35) to the C-terminus of the amino acid sequence of SEQ ID NO:1 or shortens a fragment at the C-terminus of SEQ ID NO:35 (shortens 1 to 14 amino acid residues from the C-terminus of SEQ ID NO:35, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 amino acid residues). Mutants of the repeating unit of SEQ ID NO:1 may include mutations that add Gfpgfp (SEQ ID NO:34) to the N-terminus of the amino acid sequence of SEQ ID NO:1 or shorten a fragment to its N-terminus (i.e., shortening 1, 2, 3, 4, or 5 amino acid residues from the N-terminus of SEQ ID NO:34), and / or mutants that shorten 1 to 5 amino acid residues, for example 1, 2, or 3 amino acid residues, to the C-terminus of the amino acid sequence of SEQ ID NO:1.
[0041] The repeating unit variants of SEQ ID NO:28 may include mutations that add Glpgtp (SEQ ID NO:36) to the N-terminus of the amino acid sequence of SEQ ID NO:28, or mutations that shorten a fragment at its N-terminus (i.e., shortening 1 to 5 amino acid residues from the N-terminus of SEQ ID NO:36, for example, 1, 2, 3, 4, or 5 amino acid residues).
[0042] In the context of the present invention, a conservative substitution may be defined by a substitution within one or more amino acid types as reflected below. Conservative amino acid residues: Acidic residues D and E Basic residues K, R, and H Hydrophilic uncharged residues S, T, N and Q Aliphatic uncharged residues G, A, V, L and I Nonpolar uncharged residues C, M, and P Aromatic residues F, Y, and W. Physical and functional classification of candidate amino acid residues: Alcohol group-containing residues S and T Aliphatic residues I, L, V, and M Cycloalkenyl group-related residues F, H, W, and Y Hydrophobic residues A, C, F, G, H, I, L, M, R, T, V, W, and Y Loaded electrical residues D and E Polar residues C, D, E, H, K, N, Q, R, S, and T Positively charged residues H, K, and R Small residues A, C, D, G, N, P, S, T and V Tiny residues A, G, and S Residues involved in reverse turn formation: A, C, D, E, G, H, K, N, Q, R, S, P, and T Flexible residues Q, T, K, S, G, P, D, E, and R.
[0043] As used herein, “cell adhesion” refers to the adhesion between cells and collagen. Collagen (for example, recombinant collagen as described herein) can promote adhesion between cells and the vessel in which the cells are cultured.
[0044] As used herein, the term “expression” includes, but is not limited to, any steps relating to the production of recombinant collagen, including transcription, post-transcriptional modification, translation, post-translational modification, and secretion.
[0045] As used herein, the term “expression vector” means a linear or circular DNA molecule comprising a polynucleotide encoding recombinant collagen and operably ligated to a control sequence that provides for its expression.
[0046] As used herein, the term “host cell” means any type of cell that readily undergoes transformation, transfection, transduction, etc., by the nucleic acid construct or expression vector comprising the polynucleotide of the present invention. The term “host cell” includes any offspring of a parent cell that are not identical to the parent cell due to mutations occurring during replication.
[0047] As used herein, the term “nucleic acid” means a single-stranded or double-stranded nucleic acid molecule that is isolated from a naturally occurring gene, modified to contain a nucleic acid segment in a manner not originally occurring in nature, or synthesized, and contains one or more regulatory sequences. The nucleic acid may be SEQ ID NO: 3, 6, 9, 12, 15, 18, 21, 24, 27, 30, or 33. The nucleic acid may be a codon-optimized nucleic acid, for example, a nucleic acid codon-optimized for expression in E. coli cells.
[0048] The term "operably linked" refers to a configuration in which a control sequence is positioned appropriately relative to the coding sequence of a polynucleotide to direct the expression of the coding sequence.
[0049] The degree of association between two amino acid sequences or two nucleotide sequences is described by the parameter "sequence identity". For the purposes of this invention, the sequence identity between two amino acid sequences is determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J.Mol.Biol.[Molecular Biology Journal] 48:443-453), which is performed by the Needleman program in the EMBOSS software package (EMBOSS: European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet.[Genetic Trends] 16:276~277) (preferably version 5.0.0 or updated version). The parameters used are a gap-open penalty of 10, a gap-extension penalty of 0.5, and an EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. The Needleman output (obtained using unsimplified options) marked "longest identity" is used as the identity percentage and calculated as follows: (Same residue × 100) / (Alignment length - Total number of gaps in alignment)
[0050] For the purposes of this invention, the sequence identity between two deoxynucleotide sequences is determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970), which is performed by the Needleman program in the EMBOSS software package (EMBOSS: European Molecular Biology Open Software Suite, Rice et al., 2000) (preferably version 5.0.0 or updated version). The parameters used are a gap-open penalty of 10, a gap-extension penalty of 0.5, and an EDNAFULL (EMBOSS version of NCBI NUC4.4) substitution matrix. The Needleman output (obtained using the unsimplified option) marked as "longest identity" is used as the identity percentage and calculated as follows: (Same deoxyribonucleotide × 100) / (Alignment length - Total number of gaps in alignment)
[0051] Recombinant collagen The recombinant collagen according to the present invention comprises one or more repeating units, the repeating units being linked directly or via linkers, and the repeating units comprising an amino acid sequence or variant thereof selected from the group consisting of SEQ ID NO: 1, 4, 7, 10, 13, 16, 19, 22, 25, 28, 31, or 33. The mutant may be (1) an amino acid sequence in which one or more amino acid residues are mutated in the amino acid sequence of SEQ ID NO: 1, 4, 7, 10, 13, 16, 19, 22, 25, 28, 31, or 33, or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of SEQ ID NO: 1, 4, 7, 10, 13, 16, 19, 22, 25, 28, 31, or 33. For recombinant collagen as described herein, the mutation may be selected from substitution, addition, insertion, or deletion. Preferably, the substitution is a conservative amino acid substitution.
[0052] The recombinant collagen described herein may contain multiple repeating units, for example, 2 to 50 repeating units, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 repeating units.
[0053] The linker in the recombinant collagen described herein may contain one or more amino acid residues, for example, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, or 1 to 2 amino acid residues.
[0054] Recombinant collagen is recombinant human type IV collagen or recombinant humanized type IV collagen, and preferably has cell adhesion activity. Since the recombinant collagen described herein is derived from humans, it may also be human recombinant humanized type IV collagen.
[0055] The recombinant collagen described herein may include an amino acid sequence or a variant thereof selected from the group consisting of SEQ ID NO: 2, 5, 8, 11, 14, 17, 20, 23, 26, 29, or 32, wherein the variant is (1) an amino acid sequence in which one or more amino acid residues are mutated in the above amino acid sequence, or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the above amino acid sequence.
[0056] nucleic acid construct The present invention also relates to nucleic acid constructs comprising nucleic acids of the present invention operably ligated to one or more control sequences, which direct the expression of a coding sequence in a suitable host cell under conditions compatible with the control sequences. The vector may also comprise the nucleic acid construct.
[0057] Nucleic acids can be manipulated in various ways to provide recombinant collagen expression. Depending on the expression vector, it may be desirable or necessary to manipulate the nucleic acid before inserting it into the vector. Techniques for modifying nucleic acids using recombinant DNA are well known in this field.
[0058] The regulatory sequence may be a promoter recognized by the host cell for the expression of the polynucleotide encoding the recombinant collagen of the present invention. The promoter includes a transcriptional regulatory sequence that mediates the expression of recombinant collagen. The promoter may be any nucleic acid exhibiting transcriptional activity in the host cell, including mutant promoters, truncated promoters, and hybrid promoters, and may be obtained from a gene encoding recombinant collagen extracellular or intracellular, homogeneous or heterogeneous, that encodes recombinant collagen.
[0059] Examples of promoters suitable for directing the transcription of the vector or nucleic acid construct of the present invention in bacterial host cells include promoters obtained from the Bacillus amyloliquefaciens α-amylase gene (amyQ), Bacillus licheniformis α-amylase gene (amyL), Bacillus licheniformis penicillinase gene (penP), Bacillus stearothermophilus maltogenic amylase gene (amyM), Bacillus subtilis levansucrase gene (sacB), Bacillus subtilis xylA and xylB genes, Bacillus thuringiensis cryIIIA gene, the Escherichia coli lac operon, and the Escherichia coli trc promoter.
[0060] In the yeast host, useful promoters can be obtained from the genes for Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH1, ADH2 / GAP), Saccharomyces cerevisiae triose phosphate isomerase (TPI), Saccharomyces cerevisiae metallothionein (CUP1), and Saccharomyces cerevisiae 3-phosphoglycerate kinase.
[0061] The regulatory sequence may be a transcriptional terminator that is recognized by the host cell and terminates transcription. The terminator is operably ligated to the 3' end of the polynucleotide encoding recombinant collagen. Any terminator that functions in the host cell may be used in this invention.
[0062] Preferred terminators within bacterial host cells are derived from the genes of Bacillus clausii alkaline protease (aprH), Bacillus licheniformis α-amylase (amyL), and Escherichia coli ribosomal RNA (rrnB).
[0063] Preferred terminators in yeast host cells are obtained from the genes for Saccharomyces cerevisiae enolase, Saccharomyces cerevisiae cytochrome C (CYC1), and Saccharomyces cerevisiae glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators in yeast host cells have been described by Romanos et al. (1992).
[0064] The regulatory sequence may be an mRNA stabilizer region located downstream of the promoter and upstream of the gene's coding sequence, which increases gene expression.
[0065] Examples of suitable mRNA stabilizer regions can be found in the Bacillus thuringiensis cryIIIA gene (WO94 / 25612) and the Bacillus subtilis SP82 gene (Hue et al., 1995, Journal of Bacteriology 177:3465-3471).
[0066] The regulatory sequence may be a leader sequence, which is an untranslated region of mRNA that is crucial for translation in the host cell. The leader sequence is operably ligated to the 5' end of a polynucleotide encoding recombinant collagen. Any leader sequence that functions in the host cell may be used.
[0067] Appropriate leader sequences in yeast host cells can be obtained from the genes for Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae α-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).
[0068] The regulatory sequence may also be a polyadenylated sequence, which is operably ligated to the 3' end of a polynucleotide and recognized by the host cell as a signal to add a polyadenosine residue to the transcribed mRNA during transcription. Any polyadenylated sequence that functions within the host cell may be used.
[0069] Useful polyadenylated sequences in yeast host cells are described in Guo and Sherman, 1995, Mol. Cellular Biol. [Molecular Cell Biology] 15:5983-5990.
[0070] The regulatory sequence may be a signal peptide coding region that codes for a signal peptide ligated to the N-terminus of recombinant collagen, thereby directing the recombinant collagen into the cell's secretory pathway. The 5'-terminus of the polynucleotide coding sequence may essentially contain the signal peptide coding sequence originally ligated in the translational reading frame with the coding sequence segment encoding recombinant collagen. Alternatively, the 5'-terminus of the coding sequence may contain an exogenous signal peptide coding sequence relative to the coding sequence. An exogenous signal peptide coding sequence may be required if the coding sequence does not originally contain a signal peptide coding sequence. Alternatively, the exogenous signal peptide coding sequence may simply replace the native signal peptide coding sequence to enhance the secretion of recombinant collagen. However, any signal peptide coding sequence that directs the expressed recombinant collagen into the host cell's secretory pathway may be used.
[0071] Effective signal peptide coding sequences for bacterial host cells are those derived from the genes of Bacillus NCIB 11837-produced maltogenic amylase, Bacillus licheniformis subtilisin, Bacillus licheniformis β-lactamase, Bacillus stearothermophilus α-amylase, Bacillus stearothermophilus neutral proteases (nprT, nprS, nprM), and Bacillus subtilis prsA. Further signal peptides are described in Simonen and Palva, 1993, Microbiological Reviews 57:109-137.
[0072] Useful signal peptides in yeast host cells can be obtained from the genes for Saccharomyces cerevisiae α-factor and Saccharomyces cerevisiae invertase. Other useful signal peptide coding sequences have been described by Romanos et al. (Yeast 8:423-488).
[0073] Expression vector The present invention also relates to a recombinant expression vector comprising the nucleic acid, promoter, and transcription and translation termination signals of the present invention. The nucleic acid and control sequence can be ligated together to produce a recombinant expression vector, which may contain one or more convenient restriction sites, such that a polynucleotide encoding the recombinant collagen is inserted or substituted at such sites. Alternatively, the polynucleotide may be expressed by inserting the nucleic acid or a nucleic acid construct containing the nucleic acid into a suitable vector for expression. When the expression vector is produced, the coding sequence is thus positioned in the vector so that the coding sequence is operably ligated to a suitable control sequence for expression.
[0074] The recombinant expression vector may be any vector (e.g., plasmid or virus) that can readily undergo recombinant DNA programming and induce polynucleotide expression. Vector selection typically depends on the compatibility between the vector and the host cell into which it is introduced. The vector may be a linear or closed circular plasmid.
[0075] The vector is an extrachromosomal entity, and may be an autonomously replicating vector whose replication is independent of chromosome replication, such as a plasmid, extrachromosomal element, microchromosome, or artificial chromosome. The vector may include any means to ensure self-replication. Alternatively, the vector may be a vector that is integrated into the genome upon introduction into a host cell and replicates together with one or more integrated chromosomes. Alternatively, a single vector or plasmid or two or more vectors or plasmids containing the total DNA introduced into the host cell genome may be used, or a transposon may be used.
[0076] The vector preferably includes one or more selective markers that allow for easy selection of cells such as transformed cells, transfection cells, or transduced cells. The selective markers are genes whose products provide biocide resistance or viral resistance, heavy metal resistance, prototrophicity to trophic requirements, etc.
[0077] Examples of bacterial selective markers include the Bacillus licheniformis or Bacillus subtilis dal gene, or markers conferring antibiotic resistance (such as ampicillin resistance, chloramphenicol resistance, kanamycin resistance, neomycin resistance, spectinomycin resistance, and tetracycline resistance). Appropriate markers for yeast host cells include, but are not limited to, ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3.
[0078] The selectivity marker may be a biselectivity marker system as described in WO2010 / 039889. On the other hand, the biselectivity marker is an hph-tk biselectivity marker system.
[0079] The vector may include elements that enable the vector to be integrated into the host cell's genome or to replicate autonomously within the cell independently of the genome.
[0080] When integrated into the host cell genome, the vector may rely on a polynucleotide sequence encoding the recombinant collagen or any other element of the vector for integration into the genome by homologous or non-homologous recombination. Alternatively, the vector may contain additional polynucleotides to direct integration into a precise location on a chromosome in the host cell genome by homologous recombination. To increase the likelihood of precise integration, the integration elements should contain a sufficient number of nucleic acids, e.g., 100–10,000 base pairs, 400–10,000 base pairs, or 800–10,000 base pairs, and these nucleic acids should have high sequence identity with the corresponding target sequence to increase the probability of homologous recombination. The integration elements may be any sequence homologous to the target sequence in the host cell genome. The integration elements may also be non-coding or coding polynucleotides. On the other hand, the vector can be integrated into the host cell genome by non-homologous recombination.
[0081] For autonomous replication, the vector may further include a replication origin that enables autonomous replication in the host cell under consideration. The replication origin may be any plasmid replicon that functions in the cell and mediates autonomous replication. The terms “replication origin” or “plasmid replicon” refer to a polynucleotide that can replicate a plasmid or vector in the body.
[0082] Examples of bacterial replication origins include the origins of plasmids pBR322, pUC19, pACYC177, and pACYC184 that can replicate in E. coli, as well as the origins of plasmids pUB110, pE194, pTA1060, and pAMβ1 that can replicate in Bacillus species.
[0083] Examples of replication origins used in yeast host cells include a 2-micrometer replication origin, ARS1, ARS4, a combination of ARS1 and CEN3, and a combination of ARS4 and CEN6.
[0084] The production of recombinant collagen can be improved by inserting one or more copies of the polynucleotide of the present invention into host cells. An increased copy number of the polynucleotide can be obtained by incorporating at least one other copy of the sequence into the genome of the host cell or by including a selectable marker gene that can be amplified together with the polynucleotide, and cells containing amplified copies of the selectable marker gene and other copies of the polynucleotide can be selected by culturing the cells in the presence of a suitable selective reagent.
[0085] The procedure for constructing the recombinant expression vector of the present invention by linking the above elements is well known to those skilled in the art (see, for example, Sambrook et al., 1989, *Molecular Cloning: A Laboratory Manual* (2nd edition), Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York).
[0086] host cell The present invention also relates to recombinant host cells comprising the polynucleotides of the present invention, operably linked to one or more control sequences that direct the production of recombinant collagen of the present invention. As described previously, by introducing a construct or vector comprising polynucleotides into a host cell, the construct or vector is maintained as a chromosomal integration or as an autonomously replicating extrachromosomal vector. The term “host cell” encompasses any offspring of a parent cell that is not identical to the parent cell due to mutations occurring during replication. The selection of host cells depends largely on the genes encoding recombinant collagen and their sources.
[0087] The host cell may be any cell useful for recombinant production of recombinant collagen according to the present invention, for example, a prokaryote or a eukaryote.
[0088] The prokaryotic host cell may be any Gram-positive or Gram-negative bacterium. Gram-positive bacteria include, but are not limited to, the genera Bacillus, Clostridium, Enterococcus, Geobacillus, Lactobacillus, Lactococcus, Oceanbacillus, Staphylococcus, Streptococcus, and Streptomyces. Gram-negative bacteria include, but are not limited to, the genera Campylobacter, Escherichia coli, Labobacterium, Fusobacterium, Helicobacter, Lysobacter, Neisseria, Pseudomonas, Salmonella, and Ureaplasma.
[0089] The host cell may be a eukaryote, such as a mammal, insect, plant, or fungal cell. Plant cells as used herein do not include plant cells capable of regenerating plant bodies. Animal cells also do not include cells capable of producing animal bodies.
[0090] The host cell may be a fungal cell belonging to the phyla Basidiomycota, Chytridiomycota, Zygomycota, or Oomycota. The fungal host cell may also be a yeast cell, including ascosporogenous yeast (Endomycetales), basidiosporogenous yeast, and yeasts belonging to the imperfect fungi (Blastomycetes). The yeast host cells may be from the genera Candida, Hansenula, Kluyveromyces, Pichia, Saccharomyces, Schizosaccharomyces, or Yarrowia, for example, Kluyveromyces lactis, Saccharomyces carlsbergensis, Saccharomyces cerevisiae, Saccharomyces diastaticus, Saccharomyces douglasii, Saccharomyces crivelli These are cells of Saccharomyces kluyveri, Saccharomyces norbensis, Saccharomyces oviformis, or Yarrowia lipolytica.
[0091] Production method The present invention also relates to a method for producing recombinant collagen as described herein, the method being Step (1) of culturing the host cells described herein under appropriate culture conditions, (2) Harvesting host cells and / or culture medium containing recombinant collagen, The process includes (3) a step of purifying recombinant collagen.
[0092] Host cells are cultured in a suitable nutrient medium for producing recombinant collagen using methods known in the art. For example, cells may be cultured in a shaking flask or by small or large-scale fermentation (including continuous, batch, fed-batch, or solid-state fermentation) in a laboratory or industrial fermenter, under conditions that allow recombinant collagen to be expressed and / or isolated in a suitable medium. The cells are cultured in a suitable nutrient medium containing a carbon source, a nitrogen source, and inorganic salts using procedures known in the art. Suitable media can be obtained from commercial suppliers or manufactured according to disclosed compositions (e.g., the catalog of the U.S. Center for the Preservation of Typical Cultures). If recombinant collagen is secreted into the nutrient medium, it can be recovered directly from the medium. If recombinant collagen is not secreted, it can be recovered from the cell lysate.
[0093] Recombinant collagen may be detected using methods known in the art that are specific to recombinant collagen. These detection methods include, but are not limited to, the use of specific antibodies. For example, the activity of recombinant collagen may be determined using adhesion assays.
[0094] Recombinant collagen may be recovered using methods known in this art. For example, recombinant collagen may be recovered from the nutrient medium by conventional procedures including, but not limited to, collection, centrifugation, filtration, extraction, spray drying, evaporation, or precipitation. Alternatively, the fermentation broth containing recombinant collagen may be recovered.
[0095] To obtain substantially pure recombinant collagen, recombinant collagen may be purified by various procedures known in the art, including, but not limited to, chromatography (e.g., ion exchange chromatography, affinity chromatography, hydrophobic chromatography, focal chromatography, and size exclusion chromatography), electrophoresis (e.g., calibrated isoelectric focusing), differential dissolution (e.g., ammonium sulfate precipitation), SDS-PAGE, or extraction.
[0096] Step (1) may include one or more of the following steps: Construct an expression plasmid and obtain a recombinant expression plasmid by, for example, inserting the coding nucleotide sequence into a pET-28a-Trx-His expression vector. Transform E. coli cells (e.g., E. coli competent cells BL21(DE3)) with the successfully constructed expression plasmid. The specific process may be as follows: (1) Add the plasmid to be transformed to E. coli competent cells BL21(DE3). (2) After ice bathing the mixture on ice (e.g., 10-60 min, e.g., 30 min), heat shock in a water bath (e.g., 40-50°C, e.g., 42°C, 45-90 s), remove, and then ice bathing on ice (e.g., 1-5 min, e.g., 2 min). (3) Add liquid LB medium and then culture (e.g., culture for 40-80 min, e.g., 60 min under conditions of 35-40°C, e.g., 37°C, 150-300 rpm, e.g., 220 rpm). (4) Apply the bacterial suspension and select a single colony. For example, apply the bacterial suspension uniformly to an LB plate containing ampicillin sodium, and incubate the plate in a 37°C incubator for 15-17 hours to grow a colony of uniform size.
[0097] Step (2) may also include culturing a single colony in LB medium containing an antibiotic preservation solution (for example, culturing for 5 to 10 hours, for example, 7 hours, in a constant temperature shaker at 35 to 40°C, for example, 37°C, at 150 to 300 rpm, for example, 220 rpm), and further lowering the temperature of the shaken flask after culturing to 10 to 20°C, for example, 16°C, adding IPTG to induce expression for a certain period of time, and then collecting the bacterial cells (for example, by centrifugation).
[0098] Step (3) may include resuspending the bacterial cells in a activating solution, cooling the bacterial suspension to ≤15°C, homogenizing it (for example, 1 to 5 times, e.g., 2 times under high pressure), and separating the homogenized bacterial suspension to obtain the supernatant. The activating solution may contain 100 to 500 mM sodium chloride, 10 to 50 mM Tris and 10 to 50 mM imidazole, and have a pH of 7 to 9. For example, the concentration of sodium chloride may be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The concentration of Tris may be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The concentration of imidazole may be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The pH may be 7, 7.5, 8, 8.5, or 9.
[0099] Step (3) may include purifying and enzymatically cleaving the recombinant collagen. Purification may be crude purification, which includes purifying the supernatant on a Ni-agarose gel column to obtain an eluent containing the target protein. Crude purification may include washing the column with water for example 2 to 10 column volumes (CV), for example 5 CVs. The column may be equilibrated with equilibrium solution (200 mM sodium chloride, 25 mM Tris, 20 mM imidazole, pH 8.0) for example 2 to 10 CVs, for example 5 CVs. The equilibrium solution may contain 100 to 500 mM sodium chloride, 10 to 50 mM Tris, and 10 to 50 mM imidazole, with a pH of 7 to 9. For example, the concentration of sodium chloride may be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The concentration of Tris may be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The concentration of imidazole may be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The pH may be 7, 7.5, 8, 8.5, or 9.
[0100] Step (3) may include adding the supernatant to the column and washing away contaminating proteins with a scrubbing solution. The scrubbing solution may contain 100–500 mM sodium chloride, 10–50 mM Tris, and 10–50 mM imidazole, and have a pH of 7–9. For example, the concentration of sodium chloride may be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The concentration of Tris may be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The concentration of imidazole may be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The pH may be 7, 7.5, 8, 8.5, or 9. The eluent may then be added, and the flow-through solution may be collected. The eluent may contain 100-500 mM sodium chloride, 10-50 mM Tris, and 100-500 mM imidazole, with a pH of 8.0. For example, the concentration of sodium chloride may be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The concentration of Tris may be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The imidazole concentration may be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The pH may be 7, 7.5, 8, 8.5, or 9.
[0101] Enzymatic cleavage may include adding collagen tool enzymes and performing enzymatic cleavage (with a ratio of 10-100:1 to 50:1 of total protein to total collagen tool enzymes, and performing enzymatic cleavage at 10-20°C, for example 16°C, for 2-8 hours, for example 4 hours). The enzymatically cleaved protein solution is dialyzed, for example, placed in a dialysis bag, dialyzed at 1-6°C, for example 4°C, for 1-8 hours, for example 2 hours, and then transferred to fresh dialysate and dialyzed overnight at 1-6°C, for example 4°C.
[0102] Purification may include fine purification (e.g., protein isoelectric point > 8.0). Preferably, fine purification involves gradient elution of the eluent containing the target protein or the enzymatically cleaved product (e.g., the enzymatically cleaved and dialyzed product) using a strong anion exchange chromatography column (e.g., pH may be 7, 7.5, 8, 8.5, or 9). Gradient elution includes eluting with 0-15% solution B for 1-5 minutes and holding 1-5 columns, for example, 3 column volumes; eluting with 15-30% solution B for 1-5 minutes and holding 1-5 columns, for example, 3 column volumes; eluting with 30-50% solution B for 1-5 minutes and holding 1-5 columns, for example, 3 column volumes; and eluting with 50-100% solution B for 1-5 minutes and holding 1-5 columns, for example, 3 column volumes. Solution B may contain 10-50 mM Tris, 0.5-5 M sodium chloride, and have a pH of 7-9. For example, the concentration of Tris is 15, 20, 25, 30, 35, 40, or 45 mM. The concentration of sodium chloride is 1, 2, 3, or 4 M. The pH may be 7, 7.5, 8, 8.5, or 9. Precision purification may include equilibrating and loading the column with solution A, followed by gradient elution. Solution A may contain 10-50 mM Tris and 10-50 mM sodium chloride, with a pH of 7-9. For example, the concentration of Tris is 15, 20, 25, 30, 35, 40, or 45 mM. For example, the concentration of sodium chloride is 15, 20, 25, 30, 35, 40, or 45 mM. The pH may be 7, 7.5, 8, 8.5, or 9.
[0103] To further clarify the object, technical means, and advantages of the present invention, the technical means in the embodiments of the present invention will be described clearly and completely below with reference to the embodiments of the present invention, and obviously the embodiments described are some embodiments of the present invention, not all embodiments. All other embodiments that can be obtained without creative work by a person skilled in the art based on the embodiments of the present invention are within the scope of the protection of the present invention.
[0104] To further illustrate the present invention, the following embodiments are provided. [Examples]
[0105] The present invention will be further illustrated by the following examples, but none of these examples or any combination thereof should be understood as limiting the scope or embodiments of the present invention. The scope of the present invention is limited by the appended claims, and those skilled in the art will be able to clearly understand the scope limited by the claims by combining this specification and common sense in the art. Without departing from the spirit and scope of the present invention, those skilled in the art can make any modifications or changes to the technical means of the present invention, and these modifications and changes are also included within the scope of the present invention.
[0106] Example 1: Construction, expression, and screening of recombinant type IV humanized collagen fragments 1. A large-scale functional region screening was performed, and the target gene functional regions of the following different recombinant type IV humanized collagens were obtained.
[0107] 1) Amino acid sequence of C4P7Ch (the amino acid sequence of the repeating unit is Gakgdkgskgevgfpglagspgipgskgeq, SEQ ID NO:1, the number of repeating units is 10, and the amino acid sequence of C4P7Ch is SEQ ID NO:2): Gakgdkgskgevgfpglagspgipgskgeq Gakgdkgskgevgfpglagspgipgskgeq Gakgdkgskgevgfpglagspgipgskgeq Gakgdkgskgevgfpglagspgipgskgeq Gakgdkgskgevgfpglagspgipgskgeq Gakgdkgskgevgfpglagspgipgskgeq Gakgdkgskgevgfpglagspgipgskgeq Gakgdkgskgevgfpglagspgipgskgeq Gakgdkgskgevgfpglagspgipgskgeq Gakgdkgskgevgfpglagspgipgskgeq C4P7Ch base sequence (SEQ ID NO:3): GGGGCTAAAGGAGACAAGGGCAGCAAGGGCGAAGTCGGTTTCCCAGGTCTGGCTGGTAGCCCGGGCATCCCGGGTTCAAAGGGTGAACAAGGTGCTAAAGGCGACAAAGGCAGCAAGGGTGAGGTTGGTTTCCCGGGTCTGGCGGGTTCTCCAGGCATCCCGGGTAGCAAAGGAGAACAAGGTGCGAAAGGCGATAAAGGCTCCAAGGGTGAAGTGGGCTTCCCGGGTTTAGCCGGTAGCCCAGGTATTCCGGGTAGCAAAGGCGAACAGGGTGCGAAAGGCGACAAAGGGAGTAAGGGCGAGGTGGGTTTTCCGGGTTTGGCTGGCTCGCCGGGTATTCCGGGTTCAAAGGGCGAACAGGGCGCGAAAGGTGATAAAGGCAGCAAAGGCGAGGTTGGCTTCCCGGGTCTGGCAGGTAGCCCGGGTATCCCGGGTAGCAAGGGTGAGCAGGGTGCCAAAGGCGACAAAGGTAGCAAGGGGGAAGTGGGTTTTCCGGGACTGGCAGGTAGCCCGGGTATCCCGGGTTCTAAGGGCGAGCAGGGTGCGAAAGGTGACAAAGGTAGCAAGGGCGAGGTTGGCTTTCCGGGCTTGGCGGGTAGCCCGGGCATTCCGGGCTCCAAGGGTGAACAAGGTGCGAAAGGTGATAAAGGCTCTAAGGGTGAGGTTGGTTTTCCGGGTCTGGCGGGTTCCCCGGGCATTCCGGGCTCGAAGGGCGAGCAAGGTGCTAAAGGTGATAAGGGCTCCAAGGGCGAGGTGGGTTTCCCGGGCCTGGCAGGCTCTCCGGGCATCCCGGGTTCGAAGGGCGAACAGGGTGCGAAAGGCGATAAAGGTTCCAAGGGCGAAGTCGGATTCCCTGGCCTCGCCGGTAGCCCGGGCATCCCTGGCTCCAAGGGCGAGCAG
[0108] 2) C4P7Cf (The amino acid sequence of the repeating unit is Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq, SEQ ID NO: 4, the number of repeating units is 8, and the amino acid sequence of C4P7Cf is SEQ ID NO: 5): Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpq The base sequence of C4P7Cf (SEQ ID NO:6): GGAGCTAAAGGGGACAAAGGCTCCAAAGGCGAAGTCGGTTTCCCGGGCCTGGCGGGTAGCCCGGGTATCCCGGGTAGCAAGGGTGAGCAGGGGTTCATGGGTCCACCGGGCCCACAGGGTGCCAAAGGTGATAAAGGTTCTAAGGGCGAGGTGGGTTTCCCGGGGCTGGCGGGTTCTCCGGGCATTCCGGGAAGCAAGGGTGAACAGGGCTTTATGGGTCCGCCAGGTCCGCAGGGTGCGAAAGGTGATAAAGGCAGCAAGGGAGAAGTTGGCTTCCCGGGCCTGGCAGGCAGCCCGGGCATTCCGGGGTCGAAGGGCGAACAAGGTTTCATGGGTCCGCCTGGTCCGCAAGGTGCGAAAGGTGATAAGGGTAGCAAGGGTGAAGTGGGTTTTCCGGGATTAGCGGGTTCTCCGGGCATTCCGGGTTCAAAAGGTGAACAAGGCTTTATGGGTCCGCCTGGCCCGCAGGGTGCTAAAGGCGACAAGGGTAGCAAAGGCGAGGTAGGTTTCCCGGGTTTGGCGGGCAGCCCGGGCATTCCGGGTTCCAAGGGCGAGCAGGGTTTTATGGGCCCACCGGGCCCGCAAGGCGCAAAAGGTGATAAGGGCAGCAAAGGCGAGGTGGGCTTCCCGGGACTGGCAGGTTCTCCGGGTATCCCGGGTTCCAAGGGTGAGCAGGGTTTCATGGGCCCACCGGGTCCGCAGGGTGCGAAAGGCGACAAAGGTAGCAAGGGCGAAGTTGGTTTTCCGGGCCTGGCTGGTTCGCCGGGCATCCCGGGCTCCAAGGGCGAGCAAGGCTTCATGGGTCCACCGGGTCCGCAAGGTGCCAAAGGCGACAAAGGTAGCAAGGGCGAGGTTGGTTTTCCGGGCTTGGCTGGTAGCCCTGGCATCCCGGGGTCCAAGGGTGAACAGGGCTTTATGGGTCCGCCGGGCCCTCAA
[0109] 3) Amino acid sequence of C4P7Ca (the amino acid sequence of the repeating unit is GFPGFPGAKGDKGSKGEVGFPGLAGSPGIPGSKGEQGFMGPPGPQGQPGLP, SEQ ID NO: 7, the number of repeating units is 6, and the amino acid sequence of C4P7Ca is SEQ ID NO: 8): GFPGFPGAKGDKGSKGEVGFPGLAGSPGIPGSKGEQGFMGPPGPQGQPGLP GFPGFPGAKGDKGSKGEVGFPGLAGSPGIPGSKGEQGFMGPPGPQGQPGLP GFPGFPGAKGDKGSKGEVGFPGLAGSPGIPGSKGEQGFMGPPGPQGQPGLP GFPGFPGAKGDKGSKGEVGFPGLAGSPGIPGSKGEQGFMGPPGPQGQPGLP GFPGFPGAKGDKGSKGEVGFPGLAGSPGIPGSKGEQGFMGPPGPQGQPGLP GFPGFPGAKGDKGSKGEVGFPGLAGSPGIPGSKGEQGFMGPPGPQGQPGLP C4P7Ca base sequence (SEQ ID NO:9): GGATTTCCCGGGTTCCCGGGTGCCAAAGGGGATAAAGGTTCAAAGGGCGAAGTGGGTTTCCCGGGTTTGGCTGGTAGCCCGGGTATCCCGGGTAGCAAAGGCGAACAGGGCTTTATGGGTCCGCCAGGACCGCAGGGTCAACCGGGACTGCCGGGTTTTCCGGGCTTCCCGGGTGCGAAAGGCGATAAAGGTTCCAAGGGTGAAGTTGGTTTTCCGGGTCTTGCAGGCAGCCCGGGTATTCCGGGTTCCAAGGGTGAACAGGGTTTCATGGGTCCACCGGGCCCACAAGGTCAGCCGGGTCTGCCTGGTTTCCCGGGCTTCCCGGGTGCCAAAGGCGACAAAGGTAGCAAGGGCGAAGTTGGCTTTCCGGGTCTGGCGGGTTCGCCGGGCATTCCGGGCTCGAAGGGCGAGCAGGGTTTCATGGGCCCACCGGGTCCGCAGGGTCAGCCTGGCCTGCCGGGATTCCCAGGTTTTCCGGGAGCGAAAGGCGACAAGGGTAGTAAGGGTGAGGTCGGTTTTCCAGGCTTGGCGGGCTCTCCCGGTATCCCGGGCTCTAAGGGCGAGCAAGGCTTTATGGGTCCACCGGGTCCGCAAGGTCAACCTGGATTACCGGGATTCCCAGGCTTTCCGGGCGCGAAAGGCGATAAAGGCAGCAAGGGTGAGGTGGGCTTCCCGGGCCTCGCGGGTAGCCCGGGCATCCCGGGTAGCAAGGGTGAGCAGGGCTTCATGGGTCCTCCGGGTCCGCAGGGCCAACCGGGCCTGCCGGGATTCCCGGGTTTCCCGGGCGCTAAAGGCGACAAAGGCAGCAAGGGTGAGGTTGGTTTTCCGGGTCTGGCAGGTAGCCCGGGCATTCCGGGCTCCAAGGGCGAACAGGGTTTTATGGGTCCACCGGGCCCTCAAGGTCAGCCGGGCCTGCCG
[0110] 4) Amino acid sequence of C4P7Cb (the amino acid sequence of the repeating unit is Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq, SEQ ID NO: 10, the number of repeating units is 8, and the amino acid sequence of C4P7Cb is SEQ ID NO: 11): Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq The base sequence of C4P7Cb (SEQ ID NO:12):
[0111] 5) Amino acid sequence of C4P7Cc (the amino acid sequence of the repeating unit is Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfm, SEQ ID NO: 13, the number of repeating units is 8, and the amino acid sequence of C4P7Cc is SEQ ID NO: 14): Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfm Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfm Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfm Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfm Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfm Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfm Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfm Gfpgfpgakgdkgskgevgfpglagspgipgskgeqgfm The base sequence of C4P7Cc (SEQ ID NO: 15): GGATTTCCCGGGTTCCCGGGTGCGAAAGGTGATAAAGGCAGCAAGGGTGAAGTCGGTTTTCCGGGTCTGGCAGGCAGCCCGGGTATCCCGGGTAGCAAAGGCGAACAGGGCTTTATGGGTTTCCCGGGCTTCCCAGGTGCGAAGGGCGATAAAGGTTCGAAAGGTGAGGTAGGTTTCCCGGGTTTAGCAGGTTCCCCGGGCATTCCGGGCAGCAAGGGTGAACAGGGTTTCATGGGCTTTCCGGGCTTCCCAGGAGCTAAAGGCGACAAAGGTTCTAAGGGTGAAGTGGGCTTCCCGGGTCTGGCTGGTAGCCCGGGCATCCCGGGCTCCAAGGGTGAGCAGGGTTTCATGGGTTTTCCGGGCTTCCCAGGCGCGAAAGGCGACAAAGGCAGCAAGGGCGAGGTGGGTTTTCCGGGTTTGGCGGGTAGCCCGGGTATTCCGGGTTCGAAGGGTGAACAAGGTTTCATGGGTTTTCCGGGATTCCCAGGCGCGAAAGGCGATAAGGGCAGCAAGGGCGAGGTTGGCTTCCCGGGACTGGCCGGAAGCCCGGGTATCCCGGGATCTAAGGGCGAACAAGGCTTTATGGGTTTCCCGGGTTTTCCTGGTGCGAAAGGCGATAAAGGCTCCAAGGGCGAGGTTGGTTTTCCAGGCCTGGCTGGCTCTCCGGGCATTCCGGGTAGTAAGGGTGAGCAGGGTTTTATGGGTTTTCCGGGCTTCCCGGGTGCAAAGGGTGACAAAGGTAGCAAGGGTGAAGTTGGCTTTCCGGGTCTGGCGGGTTCCCCGGGCATTCCGGGTAGCAAAGGTGAGCAAGGTTTTATGGGTTTTCCGGGCTTCCCGGGTGCCAAAGGCGACAAAGGTAGCAAGGGAGAGGTGGGCTTCCCGGGATTGGCGGGTTCCCCGGGCATCCCGGGCTCAAAGGGTGAACAGGGTTTCATG
[0112] 6) Amino acid sequence of C4P7Cd (the amino acid sequence of the repeating unit is Gfpgfpgakgdkgskgevgfpglagspgipgskgeq, SEQ ID NO: 16, the number of repeating units is 10, and the amino acid sequence of C4P7Cd is SEQ ID NO: 17): Gfpgfpgakgdkgskgevgfpglagspgipgskgeq Gfpgfpgakgdkgskgevgfpglagspgipgskgeq Gfpgfpgakgdkgskgevgfpglagspgipgskgeq Gfpgfpgakgdkgskgevgfpglagspgipgskgeq Gfpgfpgakgdkgskgevgfpglagspgipgskgeq Gfpgfpgakgdkgskgevgfpglagspgipgskgeq Gfpgfpgakgdkgskgevgfpglagspgipgskgeq Gfpgfpgakgdkgskgevgfpglagspgipgskgeq Gfpgfpgakgdkgskgevgfpglagspgipgskgeq Gfpgfpgakgdkgskgevgfpglagspgipgskgeq The base sequence of C4P7Cd (SEQ ID NO: 18):
[0113] 7) Amino acid sequence of C4P7Ce (the amino acid sequence of the repeating unit is Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp, SEQ ID NO: 19, the number of repeating units is 8, and the amino acid sequence of C4P7Ce is SEQ ID NO: 20): Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp Gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgqpglp C4P7Ce base sequence (SEQ ID NO:21):
[0114] 8) Amino acid sequence of C4P7Cg (the amino acid sequence of the repeating unit is Gakgdkgskgevgfpglagspgipgskgeqgfm, SEQ ID NO:22, the number of repeating units is 10, and the amino acid sequence of C4P7Cg is SEQ ID NO:23): Gakgdkgskgevgfpglagspgipgskgeqgfm Gakgdkgskgevgfpglagspgipgskgeqgfm Gakgdkgskgevgfpglagspgipgskgeqgfm Gakgdkgskgevgfpglagspgipgskgeqgfm Gakgdkgskgevgfpglagspgipgskgeqgfm Gakgdkgskgevgfpglagspgipgskgeqgfm Gakgdkgskgevgfpglagspgipgskgeqgfm Gakgdkgskgevgfpglagspgipgskgeqgfm Gakgdkgskgevgfpglagspgipgskgeqgfm Gakgdkgskgevgfpglagspgipgskgeqgfm The base sequence of C4P7Cg (SEQ ID NO:24): GGGGCTAAAGGAGACAAAGGTTCGAAAGGCGAGGTGGGCTTCCCAGGTCTGGCCGGTTCCCCGGGCATTCCGGGTAGCAAAGGCGAACAAGGTTTCATGGGTGCTAAAGGCGATAAAGGTAGCAAGGGTGAGGTTGGCTTCCCAGGCCTGGCTGGTTCGCCGGGCATTCCGGGCTCTAAGGGTGAACAAGGTTTCATGGGTGCAAAAGGTGATAAGGGTAGCAAGGGAGAAGTCGGTTTTCCGGGATTGGCGGGTAGCCCGGGTATCCCGGGCAGCAAGGGCGAGCAGGGTTTTATGGGTGCAAAGGGCGACAAAGGTAGCAAGGGTGAGGTGGGCTTTCCGGGCCTCGCGGGTAGCCCTGGCATCCCGGGTTCCAAAGGTGAGCAAGGCTTCATGGGTGCTAAAGGTGATAAAGGCTCCAAAGGTGAAGTGGGTTTTCCGGGCCTGGCGGGTAGCCCGGGCATTCCGGGAAGCAAGGGCGAACAGGGTTTTATGGGCGCGAAGGGTGATAAAGGTAGTAAGGGCGAAGTTGGTTTCCCGGGCCTGGCTGGCTCTCCGGGTATCCCGGGCTCCAAAGGCGAGCAGGGTTTCATGGGTGCGAAAGGTGACAAGGGTAGCAAGGGTGAGGTGGGTTTCCCAGGTTTGGCGGGTAGCCCGGGCATTCCGGGTAGCAAGGGTGAACAAGGTTTCATGGGTGCGAAAGGTGACAAAGGCAGCAAGGGCGAGGTTGGTTTCCCGGGTCTGGCGGGTAGCCCGGGCATCCCGGGCTCTAAGGGCGAGCAGGGTTTTATGGGTGCCAAAGGCGACAAGGGCTCAAAGGGTGAAGTCGGTTTTCCGGGTTTAGCCGGTTCCCCGGGCATCCCGGGTTCTAAGGGTGAACAGGGCTTCATGGGCGCGAAAGGAGATAAAGGCAGCAAAGGGGAAGTTGGTTTTCCAGGCCTGGCAGGCTCGCCGGGTATCCCGGGTTCCAAGGGCGAGCAGGGTTTTATG
[0115] 9) Amino acid sequence of C4P7Ea (the amino acid sequence of the repeating unit is Glpgtpgptgpagqkgepgsdgipgsagekgepglp, SEQ ID NO: 25, the number of repeating units is 10, and the amino acid sequence of C4P7Ea is SEQ ID NO: 26): Glpgtpgptgpagqkgepgsdgipgsagekgepglp Glpgtpgptgpagqkgepgsdgipgsagekgepglp Glpgtpgptgpagqkgepgsdgipgsagekgepglp Glpgtpgptgpagqkgepgsdgipgsagekgepglp Glpgtpgptgpagqkgepgsdgipgsagekgepglp Glpgtpgptgpagqkgepgsdgipgsagekgepglp Glpgtpgptgpagqkgepgsdgipgsagekgepglp Glpgtpgptgpagqkgepgsdgipgsagekgepglp Glpgtpgptgpagqkgepgsdgipgsagekgepglp Glpgtpgptgpagqkgepgsdgipgsagekgepglp The base sequence of C4P7Ea (SEQ ID NO:27):
[0116] 10) Amino acid sequence of C4P7Eb (the amino acid sequence of the repeating unit is Gptgpagqkgepgsdgipgsagekgepglp, SEQ ID NO: 28, the number of repeating units is 10, and the amino acid sequence of C4P7Eb is SEQ ID NO: 29): Gptgpagqkgepgsdgipgsagekgepglp Gptgpagqkgepgsdgipgsagekgepglp Gptgpagqkgepgsdgipgsagekgepglp Gptgpagqkgepgsdgipgsagekgepglp Gptgpagqkgepgsdgipgsagekgepglp Gptgpagqkgepgsdgipgsagekgepglp Gptgpagqkgepgsdgipgsagekgepglp Gptgpagqkgepgsdgipgsagekgepglp Gptgpagqkgepgsdgipgsagekgepglp Gptgpagqkgepgsdgipgsagekgepglp The base sequence of C4P7Eb (SEQ ID NO:30): GGTCCCACAGGACCGGCAGGCCAGAAAGGTGAGCCGGGTTCCGACGGCATCCCGGGTTCGGCGGGTGAGAAAGGCGAGCCGGGTTTACCGGGTCCGACCGGTCCCGCGGGTCAAAAGGGCGAGCCGGGTAGCGATGGCATTCCGGGTTCTGCGGGTGAAAAGGGCGAACCGGGCCTCCCGGGTCCTACCGGTCCGGCGGGTCAGAAAGGCGAACCGGGCAGCGATGGCATCCCGGGCAGCGCGGGCGAGAAAGGCGAACCGGGCCTGCCGGGCCCGACCGGACCAGCTGGGCAAAAAGGTGAACCGGGCAGCGACGGCATCCCGGGTTCTGCAGGCGAGAAAGGTGAACCAGGCCTGCCGGGACCGACCGGTCCGGCAGGCCAGAAAGGTGAGCCTGGCAGTGATGGTATTCCGGGTTCTGCCGGTGAAAAAGGTGAGCCGGGCCTGCCGGGGCCAACGGGCCCAGCCGGACAAAAAGGTGAGCCGGGTTCCGACGGCATCCCGGGCTCCGCCGGTGAAAAGGGTGAGCCGGGCCTGCCTGGCCCAACGGGTCCGGCTGGCCAAAAGGGCGAGCCGGGTAGCGACGGCATTCCGGGCAGCGCGGGTGAGAAGGGTGAGCCGGGATTGCCGGGTCCGACTGGTCCTGCGGGCCAGAAGGGTGAACCGGGTTCCGACGGCATCCCCGGCTCGGCGGGTGAAAAGGGCGAACCGGGTCTGCCTGGTCCGACCGGCCCAGCGGGTCAGAAGGGTGAACCGGGTAGCGATGGAATCCCGGGTAGCGCTGGTGAAAAGGGCGAGCCGGGCCTGCCGGGTCCGACCGGTCCGGCAGGCCAGAAGGGTGAACCGGGTAGCGATGGTATTCCGGGTAGCGCGGGCGAAAAAGGTGAGCCGGGCTTGCCG
[0117] 11) Amino acid sequence of C4P7Ec (the amino acid sequence of the repeating unit is Gfpgfpgakgdkgskgevgfpglagspgipgsk, SEQ ID NO:31, the number of repeating units is 10, and the amino acid sequence of C4P7Ec is SEQ ID NO:32): Gfpgfpgakgdkgskgevgfpglagspgipgsk Gfpgfpgakgdkgskgevgfpglagspgipgsk Gfpgfpgakgdkgskgevgfpglagspgipgsk Gfpgfpgakgdkgskgevgfpglagspgipgsk Gfpgfpgakgdkgskgevgfpglagspgipgsk Gfpgfpgakgdkgskgevgfpglagspgipgsk Gfpgfpgakgdkgskgevgfpglagspgipgsk Gfpgfpgakgdkgskgevgfpglagspgipgsk Gfpgfpgakgdkgskgevgfpglagspgipgsk Gfpgfpgakgdkgskgevgfpglagspgipgsk The base sequence of C4P7Ec (SEQ ID NO:33): GGATTTCCCGGGTTCCCAGGCGCAAAAGGTGATAAAGGCAGCAAGGGCGAGGTTGGTTTTCCAGGTTTAGCTGGTAGCCCGGGTATCCCGGGTAGCAAGGGCTTCCCGGGTTTTCCGGGTGCTAAAGGCGACAAAGGCTCCAAGGGCGAAGTCGGTTTCCCGGGTTTGGCGGGTAGCCCGGGTATCCCGGGTAGTAAGGGCTTTCCGGGATTCCCAGGCGCGAAAGGTGACAAAGGTAGCAAGGGCGAAGTTGGCTTCCCGGGTTTGGCGGGTTCCCCGGGTATCCCGGGGTCCAAGGGCTTCCCCGGATTCCCGGGCGCGAAAGGCGATAAAGGTAGCAAGGGTGAAGTGGGTTTTCCGGGTCTCGCTGGCAGCCCGGGTATTCCGGGCTCCAAGGGCTTTCCAGGCTTTCCGGGTGCGAAAGGCGATAAAGGTAGCAAGGGTGAGGTGGGTTTTCCGGGTCTGGCAGGTAGCCCTGGCATCCCGGGCTCGAAGGGGTTCCCGGGCTTCCCGGGAGCCAAGGGTGATAAAGGTTCTAAGGGTGAGGTCGGTTTTCCGGGCCTGGCCGGTAGCCCTGGTATCCCGGGGAGCAAGGGTTTCCCGGGTTTTCCGGGTGCCAAAGGCGATAAAGGCTCTAAGGGCGAGGTGGGCTTCCCCGGTCTGGCGGGTAGCCCGGGTATTCCGGGTTCTAAGGGCTTCCCGGGTTTTCCGGGTGCGAAAGGTGACAAGGGCTCCAAGGGTGAAGTTGGTTTTCCGGGTCTGGCTGGTAGCCCGGGTATCCCGGGTAGCAAGGGCTTCCCGGGTTTCCCGGGCGCGAAAGGCGACAAAGGTTCAAAGGGTGAAGTTGGTTTTCCTGGCCTGGCAGGCAGCCCGGGCATTCCGGGTTCCAAAGGTTTTCCGGGCTTCCCGGGTGCGAAAGGTGACAAAGGCTCGAAGGGTGAGGTGGGCTTCCCGGGTCTGGCAGGTTCTCCTGGCATTCCGGGTTCGAAA
[0118] Each of the above coding nucleotide sequences was commercially synthesized. Each of the above coding nucleotide sequences (with a collagen tool enzyme cleavage site added to the 5' end, the amino acid sequence of the collagen tool enzyme cleavage site being ENLYFQ, and the nucleotide sequence being GAAAACCTGTATTTCCAG) was inserted between the KpnI enzyme cleavage site and the XhoI enzyme cleavage site of the pET-28a-Trx-His expression vector to obtain a recombinant expression plasmid.
[0119] 3. The successfully constructed expression plasmid was used to transform E. coli competent cells BL21(DE3). The specific process was as follows: (1) E. coli competent cells BL21(DE3) were removed from an ultracold refrigerator and placed on ice. When partially thawed, 2 μl of the plasmid to be transformed was added to the E. coli competent cells BL21(DE3) and mixed uniformly 2-3 times. (2) The mixture was bathed on ice for 30 minutes, then heat-shocked in a 42°C water bath for 45-90 seconds. After removal, it was bathed on ice for 2 minutes. (3) The cells were transferred to a biological safety cabinet, 700 μl of liquid LB medium was added, and then cultured at 37°C and 220 rpm for 60 minutes. (4) 200 μl of bacterial suspension was uniformly spread onto an LB plate containing ampicillin sodium. (5) The plate was cultured in a 37°C incubator for 15-17 hours to grow colonies of uniform size.
[0120] 4. Five to six single colonies were selected from the transformed LB plates and placed in a shaking flask containing LB medium with antibiotic preservation solution. The cells were incubated for 7 hours at 220 rpm and 37°C using a constant temperature shaker. After incubation, the shaking flask was cooled to 16°C, IPTG was added to induce expression for a certain period of time, and the bacterial suspension was dispensed into a centrifuge flask. The cells were centrifuged at 8000 rpm and 4°C for 10 minutes to collect the cells, record the weight of the cells, and perform electrophoresis detection on the sampled cells (referred to as "bacterial suspension").
[0121] 5. The collected bacterial cells were resuspended in an equilibrium working solution (200 mM sodium chloride, 25 mM Tris, 20 mM imidazole, pH 8.0), the bacterial suspension was cooled to ≤15°C and homogenized, then homogenized twice under high pressure (the two homogenized samples were labeled "homogeneous" and "homogeneous," respectively), and the bacterial suspension was collected after completion. The homogenized bacterial suspension was dispensed into centrifuge flasks and centrifuged at 17000 rpm at 4°C for 30 minutes, the supernatant was collected, and electrophoresis detection was performed on the supernatant (labeled "supernatant") and precipitate.
[0122] 6. Recombinant type IV humanized collagen was purified and enzymatically cleaved. The specific process is as follows (1) to (4). (1) Crude purification was performed as follows a to g. a. The column was washed with water using 5 CVs (Ni6FF, Cytiva). b. The column was equilibrated with an equilibrium solution using 5 CVs (200 mM sodium chloride, 25 mM Tris, 20 mM imidazole, pH 8.0). c. For loading, the supernatant obtained by centrifugation was added to the column until the liquid had finished flowing, and electrophoresis detection was performed on the flow-through solution (referred to as "flow-through"). d. For washing off contaminating proteins, 25 mL of scrubbing solution (200 mM sodium chloride, 25 mM Tris, 20 mM imidazole) was added until the liquid had finished flowing, and electrophoresis detection was performed on the scrubbed flow-through solution (referred to as "scrubbing"). e. For the collection of the target protein, 20 mL of eluent (200 mM sodium chloride, 25 mM Tris, 250 mM imidazole, pH 8.0) was added, the flow-through solution (indicated as "eluent") was collected, the protein concentration was detected, the amount of protein was calculated, and electrophoretic detection was performed. f. The column was washed with 1 M imidazole working solution (indicated as "1 M wash"). g. The column was washed with purified water. (2) Enzymatic digestion was performed as follows: TEV enzyme was added in a ratio of 50:1 between the total amount of protein and the total amount of TEV enzyme, and the enzyme was digested at 16°C for 4 hours. The sample was taken and electrophoretic detection was performed (indicated as "after digestion"). The enzyme-digested protein solution was placed in a dialysis bag, dialyzed at 4°C for 2 hours, then transferred to fresh dialysate and dialyzed overnight at 4°C (indicated as "fluid change"). (3) For the precise purification as described in a-c below, a) for the column (Capto Q, Cytiva), the column was equilibrated with solution A (20 mM Tris, 20 mM sodium chloride, pH 8.0) and the flow rate was set to 10 ml / min. b) For loading, the flow rate was set to 5 ml / min, the column was loaded, the flow-through solution was collected (indicated as "QFL"), and electrophoretic detection was performed. c) For gradient elution, the following settings were used: elution with 0-15% solution B (20 mM Tris, 1 M sodium chloride, pH 8.0) for 2 minutes, followed by holding three CVs; elution with 15-30% solution B for 2 minutes, followed by holding three CVs; elution with 30-50% solution B for 2 minutes, followed by holding three CVs; and elution with 50-100% solution B for 2 minutes, followed by holding three CVs. When a peak appeared, the sample was collected and electrophoretic detection was performed (indicated as "B wash"). d. The column was washed. The protein was stored at 4°C.
[0123] 7. Electrophoresis detection The specific process is as follows: 10 μl of 5× protein loading buffer (250 mM Tris-HCl (pH: 6.8), 10% SDS, 0.5% bromophenol blue, 50% glycerol, 5% β-mercaptoethanol) was added to 40 μl of sample solution, boiled in 100°C water for 10 minutes, then 10 μl of each well was added to an SDS-PAGE protein gel, electrophoresis was performed at 80 V for 2 hours, the proteins were stained with Coomassie Brilliant Blue stain (0.1% Coomassie Brilliant Blue R-250, 25% isopropanol, 10% glacial acetic acid) for 20 minutes, and then destained with protein destaining solution (10% acetic acid, 5% ethanol).
[0124] Figure 1 shows the electrophoretic detection results for C4P7Ca. Figure 2 shows the electrophoretic detection results for C4P7Cb. Figure 3 shows the electrophoretic detection results for C4P7Cc. Figure 4 shows the electrophoretic detection results for C4P7Cd. Figure 5 shows the electrophoretic detection results for C4P7Ce. Figure 6 shows the electrophoretic detection results for C4P7Cf. Figure 7 shows the electrophoretic detection results for C4P7Cg. Figure 8 shows the electrophoretic detection results for C4P7Ch. Figure 9 shows the electrophoretic detection results for C4P7Ea. Figure 10 shows the electrophoretic detection results for C4P7Eb. Figure 11 shows the electrophoretic detection results for C4P7Ec. Figures 1 to 11 show that the actual molecular weights of each separated protein (C4P7Ca, C4P7Cb, C4P7Cc, C4P7Cd, C4P7Ce, C4P7Cf, C4P7Cg, C4P7Ch, C4P7Ea, C4P7Eb, and C4P7Ec) matched their corresponding predicted molecular weights, demonstrating that the proteins are accurately expressed.
[0125] Example 2: Mass spectrometry detection of recombinant type IV humanized collagen Experimental method [Table 1]
[0126] Protein samples (collagen of C4P7Cf and C4P7Ch) were reduced by DTT and alkylated with iodoacetamide, then enzymatically digested overnight with trypsin. The resulting peptide fragments were further desalted using C18 ZipTip, mixed with matrix α-cyano-4-hydroxycinnamic acid (CHCA), and dropped onto a target plate. Finally, the samples were analyzed using a matrix-assisted laser desorption / ionization-time-of-flight mass spectrometer (MALDI-TOF / TOF Ulraflextreme). TM Analysis was performed using Brucker, Germany (see Protein J. 2016;35:212-7 for the peptide mass fingerprinting technique).
[0127] Data retrieval is performed via the MS / MS Ion Search page on the local masco site. Protein identification results are obtained from the primary mass spectra of peptide fragments generated after enzymatic degradation. For detection parameters, two uncleaved sites are defined in Trypsin enzymatic degradation: cysteine alkylation is considered a fixed modification, and methionine oxidation is considered a variable modification. The database used for identification is NCBprot.
[0128] Table 1. Molecular weight and corresponding polypeptide of recombinant type IV humanized collagen C4P7Cf as detected by mass spectrometry. [Table 2]
[0129] The polypeptide fragment has a coverage rate of 99.04% compared to the theoretical sequence. gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgakgdkgskgevgfpglagspgipgskgeqgfmgppgpq gakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgakgdkgskgevgfpglagspgipgskgeqgfmgppgpqgakgdkgskgevgfpglagspgipgskgeqgfmgppgpqGakgdkgskgevgfpglagspgipgskgeqgfmgppgpq The underlined residue is detected as the cover portion, and the detection result is highly reliable.
[0130] Table 2. Molecular weight and corresponding polypeptide of recombinant type IV humanized collagen C4P7Ch as detected by mass spectrometry. [Table 3]
[0131] The polypeptide fragment has a coverage rate of 98% compared to the theoretical sequence (Gakgdk gskgevgfpglagspgipgskgeqgakgdkgskgevgfpglagspgipgskgeqgakgdkgskgevgfpglagspgipgskgeqgakgdkgskgevgfpglagspgipgskgeqgakgdkgskgevgfpglagspgipgskgeqgak gdkgskgevgfpglagspgipgskgeqgakgdkgskgevgfpglagspgipgskgeqgakgdkgskgevgfpglagspgipgskgeqgakgdkgskgevgfpglagspgipgskgeqgakgdkgskgevgfpglagspgipgskgeq The underlined residue is detected as the cover portion, and the detection result is highly reliable.
[0132] Example 3: Detection of the bioactivity of recombinant type IV humanized collagen For information on methods for detecting collagen activity, please refer to the literature: Juming Yao, Satoshi Yanagisawa, Tetsuo Asakura, Design, Expression and Characterization of Collagen-Like Proteins Based on the Cell Adhesive and Crosslinking Sequences Derived from Native Collagens, J Biochem. 136, 643-649 (2004). The specific procedure is as follows.
[0133] (1) Using the ultraviolet absorption method, the concentrations of target protein samples containing bovine type I collagen (China Food and Drug Administration, No.: 380002), recombinant humanized collagen C4P7Cf, and C4P7Ch according to the present invention were detected.
[0134] Specifically, the amount of ultraviolet absorption of the sample at 215 nm and 225 nm was measured, and the protein concentration was calculated according to the empirical formula C(μg / mL) = 144 × (A215 - A225). Note that detection is necessary when A215 < 1.5. The principle of this method is as follows: It measures the characteristic absorption of peptide bonds in far ultraviolet light, is unaffected by the chromophore content, contains few disturbing substances, is easy to operate, and is suitable for detecting human collagen and its analogs that do not fluoresce with Coomassie Brilliant Blue. (Reference: Walker JM. The Protein Protocols Handbook, second edition. Humana Press. 43-45.) After detecting the protein concentration, the concentration of all target proteins was adjusted to 0.5 mg / mL with PBS.
[0135] (2) Sample preparation: Experiments were performed directly using the stock sample solution. Bovine type I collagen (PC) as a positive control was diluted to 1 mg / ml with D-PBS for use, and D-PBS buffer (NC) was used for the negative control.
[0136] (3) Coating: 100 μL each of collagen (C4P7Cf or C4P7Ch) at different concentrations, along with positive and negative controls, were added to each well of the enzyme standard plate. Five overlapping wells were set up for each group, and the plates were incubated overnight at 4°C.
[0137] (4) Blocking: Discard the supernatant, add 100 μL of 1% BSA (heat-inactivated at 56°C for 30 min), and incubate at 37°C for 60 min. Discard the supernatant and wash three times with D-PBS solution.
[0138] (5) Cell inoculation: Each well is inoculated with 10 wells of cultured cells resuspended in D-PBS. 5 Each 3T3 / NIH cell was added and incubated at 37°C for 120 minutes. Each well was washed three times with D-PBS solution.
[0139] (6) Detection: OD using a CCK8 detection kit (Manufacturer: Beyotime, Product Catalog Number: C0038) 450 Absorbance in nm was detected. The degree of cell adhesion was calculated using the following formula. Cell adhesion rate can reflect the cell adhesion capacity of collagen. The higher the cell adhesion capacity, the more quickly a favorable external environment can be provided to the cells, thus aiding cell adhesion.
number
[0140] (7) Statistical analysis: A two-sided t-test was used to statistically analyze the statistical difference between the target recombinant humanized collagen and the negative control, where P<0.05 for *, P<0.01 for **, and P<0.001 for ***.
[0141] The results are shown in Figures 11 and 12. Compared to the D-PBS group, the positive control clearly had a cell adhesion-promoting effect, and recombinant humanized collagen C4P7Cf and C4P7Ch also had a cell adhesion-promoting effect.
[0142] While the present invention has been described with reference to exemplary embodiments, those skilled in the art will understand that various other modifications, omissions, and / or additions may be made and elements of the described embodiments replaced with substantially equivalent ones without departing from the spirit and scope of the invention. Furthermore, many modifications can be made to adapt the teachings of the invention to specific circumstances or materials without departing from the scope of the invention. Accordingly, this specification is not intended to limit the invention to specific embodiments disclosed for the purpose of carrying out the invention, but rather to include all embodiments contained in the appended claims in the invention.
Claims
1. Recombinant collagen comprising 2 to 10 directly linked repeating units, wherein each repeating unit consists of an amino acid sequence of SEQ ID NO: 1, or an amino acid sequence having at least 90% identity with the amino acid sequence of SEQ ID NO:
1. Recombinant collagen with cell adhesion activity.
2. Recombinant collagen according to claim 1, comprising the amino acid sequence of SEQ ID NO: 2, or an amino acid sequence having at least 90% identity with the amino acid sequence of SEQ ID NO:
2.
3. A nucleic acid encoding recombinant collagen as described in claim 1.
4. The nucleic acid according to claim 3, comprising the nucleotide sequence of SEQ ID NO:
3.
5. A vector comprising the nucleic acid described in Claim 3.
6. The vector according to claim 5, comprising an expression control element operably linked to the nucleic acid, a nucleotide sequence of a purified tag and / or a nucleotide sequence of a leader sequence.
7. The expression control element is selected from the group consisting of a promoter, a terminator, or an enhancer. The vector according to claim 6, wherein the purified tag is selected from the group consisting of His tag, GST tag, MBP tag, SUMO tag, or NusA tag.
8. A host cell comprising the nucleic acid described in claim 3.
9. The host cell according to claim 8, which is a eukaryotic cell or a prokaryotic cell.
10. The host cell according to claim 9, wherein the eukaryotic cell is a yeast cell, an animal cell and / or an insect cell, and / or the prokaryotic cell is an Escherichia coli cell.
11. A composition comprising recombinant collagen as described in claim 1.
12. The composition according to claim 11, which is one or more of the following: a bio-coating material, a human biomimetic material, a cosmetic orthopedic material, an organoid culture material, a cardiovascular stent material, a coating material, a tissue injection filling material, an ophthalmic material, an obstetric and gynecological biological material, a nerve repair and regeneration material, a liver tissue material and a blood vessel repair and regeneration material, a 3D printed artificial organ biological material, a cosmetic raw material, a medicinal auxiliary material and a food additive.
13. The composition according to claim 11, which is a surface composition, an injection composition, or an oral composition.
14. Use of recombinant collagen according to claim 1 or 2, nucleic acid according to claim 3 or 4, vector according to any one of claims 5 to 7 and / or host cell according to any one of claims 8 to 10 in the manufacture of one or more of the following: biocoating materials, human biomimetic materials, cosmetic materials, organoid culture materials, cardiovascular stent materials, coating materials, tissue injection filling materials, ophthalmic materials, obstetric and gynecological biological materials, nerve repair and regeneration materials, liver tissue materials and blood vessel repair and regeneration materials, 3D printed artificial organ biological materials, cosmetic raw materials, medicinal auxiliary materials and food additives.
15. An in vitro method for promoting cell adhesion, comprising the step of bringing cells into contact with recombinant collagen according to claim 1 or 2 in vitro.
16. The in vitro method according to claim 15, wherein the cells are animal cells.
17. Recombinant collagen according to claim 1 or 2, for use in a cosmetic method comprising the step of administering the recombinant collagen according to claim 1 or 2 to a subject.
18. (1) A step of culturing host cells containing a vector having nucleic acid encoding recombinant collagen under appropriate culture conditions, (2) Harvesting host cells and / or culture medium containing recombinant collagen, A method for producing recombinant collagen according to claim 1 or 2, comprising the step (3) of purifying recombinant collagen.
19. The method according to claim 18, wherein the host cell is an Escherichia coli cell.
Citation Information
Patent Citations
Preparation method of biosynthetic human body structural material
CN114940712A
Preparation method of biosynthetic human body structural material
CN116478274A