Method for biosynthesizing type XXII collagen, a structural material for the human body
By screening and expressing 21 types of recombinant human type XXII collagen, especially C22s and C22t, the problems of large-scale production and low bioactivity of recombinant human type XXII collagen in existing technologies have been solved, realizing the preparation of highly active and high-purity recombinant human collagen, which is suitable for the biomedical field.
Patent Information
- Application Number
- CN202410728220.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2043-07-18
AI Technical Summary
Existing technologies make it difficult to produce highly bioactive recombinant human type XXII collagen on a large scale, and traditional methods suffer from problems such as loss of biological activity, poor water solubility, and high sensitization.
Through large-scale functional region screening, 21 types of recombinant human type XXII collagen, especially C22s and C22t, were discovered and expressed. They were expressed and purified using E. coli. The peptide and nucleic acid constructs showed high activity in cell adhesion activity assays. Purification was performed using Ni-agarose gel column and strong anion exchange chromatography column.
We have achieved the preparation of recombinant human type XXII collagen with high yield and high purity, which has excellent cell adhesion activity and is suitable for various applications in the biomedical field.
Smart Images

Figure GDA0005579839960000081 
Figure GDA0005579839960000082 
Figure GDA0005579839960000091
Abstract
Description
[0001] This application is a divisional application of Chinese invention patent application filed on July 18, 2023, application number: 202310884333.4, entitled "Method for Biosynthesizing Human Structural Material Type XXII Collagen". Technical Field
[0002] This invention belongs to the field of biomedical technology and relates to a recombinant XXII type humanized collagen, its preparation method and uses. Background Technology
[0003] Collagen (COL) is a helical fibrous functional protein composed of three polypeptide chains. It is also a major component of the extracellular matrix, abundant in quantity and widely distributed. Collagen accounts for 25%–30% of the total protein in the human body, mainly found in the skin, tendons, and bones, playing a vital role in protecting and connecting various tissues and performing important physiological functions within the body.
[0004] Collagen has good biocompatibility, degradability and low antigenicity. Its unique biological structure has made it a research hotspot in recent years and it can be widely used in many fields such as biomedicine, cosmetics, health products and food.
[0005] The human body contains 28 different types of collagen, which can be divided into two main categories based on whether their structure is fibrous or non-fibrous. Fibrous collagen mainly functions as a cellular scaffold, fixing cell position, acting as an anchor, and providing tensile strength and stiffness to tissues. Non-fibrous collagen is further subdivided into basal collagen, short-chain collagen, transmembrane collagen, etc., each performing different functions.
[0006] Type XXII collagen (encoded by the COL22A1 gene) is mainly distributed in the heart and skeletal muscle of human tissues and is an important component for maintaining tissue connections, muscle attachment, and contraction. Furthermore, studies have shown that microinjection of recombinant human type XXII collagen can reverse malnutrition phenotypes, indicating that COL22A1 is a candidate gene for human muscle dystrophy.
[0007] Traditional methods for producing collagen involve purifying animal-derived tissues using acid, alkali, and enzymatic hydrolysis to extract collagen derivatives. However, collagen obtained through these methods loses its original biological activity, has poor water solubility, is difficult to bind to the human body, and carries the risk of viral infection and sensitization, thus failing to perform its true function. Some research institutions have also expressed human collagen in vitro using conventional recombinant expression methods, but this is costly, has a long production cycle, and cannot be implemented on a large scale.
[0008] Therefore, there is an urgent need in the market for a collagen material with excellent biomaterial properties, an amino acid sequence that is highly homologous to that of the human body, and the ability to be mass-produced in an industrial system. Summary of the Invention
[0009] To address the lack of recombinant human type XXII collagen in existing technologies, the inventors conducted large-scale functional region screening of human type XXII collagen, discovering 21 types of recombinant human type XXII collagen. These recombinant human type XXII collagens can be expressed in *E. coli* and can be purified. Furthermore, the inventors found that the C22s and C22t components of these 21 types of recombinant human type XXII collagen had high yields, good purity after purification, and higher activity than the positive control (bovine type I collagen) in cell adhesion activity assays.
[0010] In one aspect, the present invention provides a polypeptide comprising one or more repeating units, said repeating units being directly or via a linker, said repeating units comprising an amino acid sequence selected from the group consisting of or variants thereof: SEQ ID NO:1(gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKghtglmgpqglpgEngpvg ppgppgqpgfpglRgEs), 4(gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKght), 7(gEKgEmgvagpmglpgpKgDigaigpvgapgpKgEKgDvgigpfgqgEKgEKgslglpgppg RDgsKgmRgEpgElgEpglpgEvgmRgpqgppglpgppgRvgapglqgERgEKgtRgEKgE RglDgfpgKpgDtgqqgRpgps), 12(gEKgEKgslglpgppgRDgsKgmRgEpgElgEpglpgEvgmRgpqgppglpgppgRvgapgl qgERgEKgtRgEKgERglDgfpgKpgDtgqqgRpgps), 18(gEqgapgpRghqgapgppgaRgpigpEgRDgppglqgl RgKKgDmgppgip), 21(gppgppgvpgppgpggspglpgEigfpgKpgppgptgppgKDgpngppgppgtKgEpgERgED glpgKpglRgEigEqglagRpgEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEK gEagpagpp), 24(gEigfpgKpgppgptgppgKDgpngppgppgtKgEpgERgEDglpgKpglRgEigEqglagRp gEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEKgEa), 27(gtKgEpgERgEDglpgKpglRgEigEqglagRpgEKgEaglpgapgfp gvRgEKgDqgEKgElglpglKgDRgEKgEa), 30(gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgahgapgaagnpgapghvgapgpsgppgsvgapglRgtpgKDgERgEKgaagEEgspgpvgpRgDpgapglpgppgKg), 33(gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgahgapgaagnpgapghvgapg psgppgsvgapglRgtpgKDgERgEKgaagEE), 36(gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgah), 39(gvagppgpsgppgDKgs pgsRglpgfpgpqgpagRDgapgnpgERgppgKpgls), 42(gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgiagpqgsqgERgaDgEvgq KgDqghpgvpgfmgppgnpgppgaDgiagaagpp), 45(gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgiagpqgsqgERgaDgEvgq KgDq), 48(gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgppglpglpgfKgDKgvpgKpgREgtEg KKgEagppglpgpp), 51(gKpgppgEpgKagEpglpgpEgaRgppgfKghtgDsgapgpRgEsgamglpgqEglpgKD gDtgptgpqgpqgpRgppgKngspgspgEpgpsgtpgqKgsKgEngspglpgflgpRgppgEpgEKgvpgKE), 54(gKpgppgEpgKagEpglpgpEgaRgpp gfKghtgDsgapgpRgEsgamglpgqEglpgKDgDt), 57(gpqgpqgpRgppgKngspgspgEpgpsgtpgqKgsKgEngspglpgflgpRgppgEpgEKgvpg KE), 60(gRpgppgppgKDglpgRagpmgEpgRpgqgglEgpsgpigpKgERgaKgDpgap), wherein the variant is (1) an amino acid sequence in which one or more amino acid residues are mutated or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence.
[0011] In one implementation, the number of repeating units is 2 to 50, such as 2 to 45, 2 to 40, 2 to 35, 2 to 30, 2 to 25, 2 to 20, 2 to 15, or 2 to 10 repeating units. For example, the number of repeating units can be 2, 3, 4, 5, 6, 7, 8, 9, 10, or a range thereof.
[0012] In one embodiment, the linker comprises one or more amino acid residues, such as 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, or 1-2 amino acid residues.
[0013] In one implementation, the mutation is selected from substitution, addition, insertion, or deletion.
[0014] In one implementation, the substitution is a conserved amino acid substitution.
[0015] In one embodiment, the peptide is recombinant collagen. In another embodiment, the peptide is recombinant type XXII collagen, preferably human recombinant type XXII collagen.
[0016] In one embodiment, the polypeptide has cell adhesion activity.
[0017] In one embodiment, the polypeptide comprises an amino acid sequence selected from the group consisting of: SEQ ID NO: 2, 5, 7, 10, 13, 15, 16, 19, 22, 25, 28, 31, 34, 37, 40, 43, 46, 49, 52, 55, 58, or 29, wherein the variant is (1) an amino acid sequence in which one or more amino acid residues are mutated, or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence.
[0018] In one embodiment, the mutation is selected from substitution, addition, insertion, or deletion. In one embodiment, the substitution is a conserved amino acid substitution.
[0019] In another aspect, a nucleic acid is provided that encodes the polypeptide described herein. In one embodiment, the nucleic acid comprises a codon-optimized nucleotide sequence. In one embodiment, the nucleotide sequence is codon-optimized for expression in *E. coli*. In one embodiment, the nucleic acid comprises a nucleotide sequence selected from the group consisting of: SEQ ID NO: 3, 6, 8, 9, 11, 14, 17, 20, 23, 26, 29, 32, 35, 38, 41, 44, 47, 50, 53, 56, or 59.
[0020] In another aspect, a vector is provided that contains the nucleic acid described herein. In one embodiment, the vector contains an expression control element operatively linked to the nucleic acid, a purification tag nucleotide, and / or a leader sequence nucleotide. In one embodiment, the expression control element is selected from promoters, terminators, or enhancers. In one embodiment, the purification tag is selected from His tags, GST tags, MBP tags, SUMO tags, or NusA tags. In one embodiment, the vector is an expression vector or a cloning vector, preferably pET-28a(+). pET-28a(+) may contain N-terminal His, Thrombin, and T7 protein tags, and a C-terminal His tag. In this document, the N-terminus of the peptide may contain an enzyme cleavage site for purification, such as a TEV enzyme cleavage site.
[0021] In another aspect, a host cell is provided, which contains the nucleic acids or vectors described herein. In one embodiment, the host cell is a eukaryotic or prokaryotic cell. In one embodiment, the eukaryotic cell is a yeast cell, animal cell, and / or insect cell, and in one embodiment, the prokaryotic cell is an Escherichia coli cell, such as Escherichia coli BL21.
[0022] In another aspect, compositions are provided that comprise one or more of the peptides, nucleic acids, carriers, and host cells described herein. In one embodiment, the composition is a kit. In one embodiment, the composition is one or more of the following: bio-dressings, biomimetic materials, plastic and cosmetic materials, organoid culture materials, cardiovascular scaffold materials, coating materials, tissue injection fillers, ophthalmic materials, obstetric and gynecological biomaterials, nerve repair and regeneration materials, liver tissue materials and vascular repair and regeneration materials, 3D-printed artificial organ biomaterials, cosmetic ingredients, and pharmaceutical excipients. In one embodiment, the composition is an injectable composition or an oral composition.
[0023] In another aspect, the uses of the peptides, nucleic acids, carriers, host cells and / or compositions described herein are provided in one or more of the following: bio-dressings, biomimetic materials, plastic and cosmetic materials, organoid culture materials, cardiovascular stent materials, coating materials, tissue injection fillers, ophthalmic materials, obstetric and gynecological biomaterials, nerve repair and regeneration materials, liver tissue materials and vascular repair and regeneration materials, 3D printed artificial organ biomaterials, cosmetic ingredients and pharmaceutical excipients.
[0024] In another aspect, methods for promoting cell adhesion are provided, comprising the step of contacting the polypeptides, nucleic acids, carriers, host cells, and / or compositions described herein with cells. In one embodiment, the cell is an animal cell. The animal cell may be a mammalian cell or a human cell.
[0025] In another aspect, methods are provided for performing cosmetic surgery, tissue filler injections, ophthalmic treatments, nerve repair, or vascular repair on subjects in need, comprising administering the polypeptide of this article to the subject. In one embodiment, the administration is oral or injectable. In one embodiment, the subject has a disease or condition associated with type II collagen deficiency, such as muscular dystrophy; preferably, the subject is a human being.
[0026] In another aspect, the production of the polypeptides described herein is provided, comprising:
[0027] (1) Culture the host cells described herein under suitable culture conditions;
[0028] (2) Harvesting host cells and / or culture medium containing polypeptides; and
[0029] (3) Purify the polypeptide.
[0030] In one embodiment, the host cell is an Escherichia coli cell, preferably an Escherichia coli BL21(DE3) cell.
[0031] In one implementation, step (1) includes culturing E. coli cells in LB medium and inducing expression via IPTG.
[0032] In one embodiment, step (2) includes harvesting Escherichia coli cells, resuspending them in a equilibration working solution, homogenizing the Escherichia coli cells, preferably by high-pressure homogenization, and separating the supernatant. In one embodiment, the equilibration working solution contains 100-500 mM sodium chloride, 10-50 mM Tris, 10-50 mM imidazole, and pH 7-9.
[0033] In one implementation, step (3) includes crude purification, enzymatic digestion, fine purification, and / or reverse nickel column purification.
[0034] In one embodiment, the crude purification includes Ni-agarose gel column purification of the supernatant to obtain an eluent containing the target protein, wherein the eluent contains 100-500 mM sodium chloride, 10-50 mM Tris and 100-500 mM imidazole, preferably at pH 7-9.
[0035] In one embodiment, purification includes gradient elution of the eluent containing the target protein using a strong anion exchange chromatography column; preferably, the gradient elution includes 0-15% solution B for 1-5 minutes followed by 3 column volumes, 15-30% solution B for 1-5 minutes followed by 3 column volumes, 30-50% solution B for 1-5 minutes followed by 3 column volumes, and 50-100% solution B for 1-5 minutes followed by 3 column volumes; wherein solution B contains 10-50 mM Tris, 0.5-5 M sodium chloride, and pH 7-9.
[0036] In one embodiment, the reverse nickel column purification includes purification of the enzymatically digested product onto a Ni-agarose gel column; preferably, the eluent contains 10-50 mM Tris, 10-50 mM sodium chloride, 0.5-5 M imidazole, and pH 7-9.
[0037] In this article, enzyme digestion can be performed by TEV enzyme.
[0038] The advantages of this invention include:
[0039] 1. The polypeptide of this invention is derived from type XXII collagen and is recombinant human type XXII collagen;
[0040] 2. The polypeptides of the present invention are suitable for the preparation of Escherichia coli and can be separated and purified;
[0041] 3. Some peptides of the present invention have high yields and are suitable for subsequent purification (Ni column or strong anion exchange column purification); and
[0042] 4. The peptides of the present invention possess cell adhesion activity. The peptides of the present invention (e.g., C22s and C22t) exhibit higher cell adhesion activity compared to positive controls. Attached Figure Description
[0043] Figure 1 The electrophoresis results for C22a are shown.
[0044] Figure 2 The electrophoresis results for C22b are shown.
[0045] Figure 3 The results of C22c electrophoresis are shown.
[0046] Figure 4 The results of C22d electrophoresis are shown.
[0047] Figure 5 The results of C22e electrophoresis are shown.
[0048] Figure 6 The electrophoresis results of C22f are shown.
[0049] Figure 7 The results of C22g electrophoresis are shown.
[0050] Figure 8 The results of C22h electrophoresis are shown.
[0051] Figure 9 The results of C22i electrophoresis are shown.
[0052] Figure 10 The electrophoresis results of C22j are displayed.
[0053] Figure 11 The results of C22k electrophoresis are shown.
[0054] Figure 12 The results of C22l electrophoresis are shown.
[0055] Figure 13 The results of C22m electrophoresis are shown.
[0056] Figure 14 The results of C22n electrophoresis are shown.
[0057] Figure 15 The results of C22o electrophoresis are shown.
[0058] Figure 16 The results of C22p electrophoresis are shown.
[0059] Figure 17 The results of C22q electrophoresis are shown.
[0060] Figure 18 The results of C22r electrophoresis are shown.
[0061] Figure 19 The results of C22s electrophoresis are shown.
[0062] Figure 20 The results of C22t electrophoresis are shown.
[0063] Figure 21 The results of C22u electrophoresis are shown.
[0064] Figure 22 The cell adhesion experiment results of the recombinant XXII type humanized collagen of the present invention are shown. Detailed Implementation
[0065] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below in conjunction with the embodiments of this invention. Obviously, the described embodiments are only some embodiments of this invention, not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0066] As used in this article, type XXII collagen (encoded by the COL22A1 gene) is mainly distributed in the heart and skeletal muscle of human tissues and is an important component for maintaining tissue connections, muscle attachment, and contraction. Furthermore, studies have shown that microinjection of recombinant human type XXII collagen can reverse malnutrition phenotypes, indicating that COL22A1 is a candidate gene for human muscle dystrophy.
[0067] As used herein, a "peptide" refers to a plurality of amino acid residues linked by peptide bonds. In this document, a peptide comprises one or more repeating units. These repeating units may be derived from human type XXII collagen. Therefore, a peptide may be human recombinant type XXII collagen. Multiple repeating units may be linked by a linker, which may be a native amino acid residue of the repeating unit on human type XXII collagen, for example, 1-50 amino acid residues. The repeating unit may be SEQ ID NO: 1, 4, 7, 10, 13, 16, 19, 22, 25, or 28. The peptide may be SEQ ID NO: 2, 5, 8, 11, 14, 17, 20, 23, 26, or 29.
[0068] As used herein, “human recombinant type XXII collagen” refers to a recombinant protein consisting of or substantially consisting of sequences derived from human type XXII collagen. In this context, human recombinant type XXII collagen may consist of or substantially consist of fragments or multiple repeats of fragments derived from human type XXII collagen.
[0069] As used herein, the term "variant" refers to a polypeptide having cell adhesion activity that includes alterations (i.e., substitutions, additions, insertions, and / or deletions) at one or more positions. Substitution means replacing an amino acid occupying a position with a different amino acid; deletion means removing an amino acid occupying a position; and insertion means adding an amino acid adjacent to and immediately following an amino acid occupying a position. Addition means adding one or more amino acid residues to the C-terminus and / or N-terminus of an amino acid sequence. Substitution can be conserved substitution. A variant of a repeating unit can be a sequence that alters or mutates (i.e., substitutes, additions, insertions, and / or deletions) one or more amino acid residues from SEQ ID NO: 1, 4, 7, 12, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60. Variants of the polypeptide can be sequences that are altered or mutated (i.e., substituted, added, inserted and / or deleted) one or more amino acid residues from SEQ ID NO:2, 5, 7, 10, 13, 15, 16, 19, 22, 25, 28, 31, 34, 37, 40, 43, 46, 49, 52, 55, 58 or 29.
[0070] In the context of this invention, conservative substitution may be defined by substitutions within one or more amino acid classes reflected in one or more of the following tables:
[0071] Conserved amino acid residues:
[0072]
[0073] Physical and functional classification of the alternative amino acid residues:
[0074]
[0075]
[0076] As used herein, “cell adhesion” refers to the adhesion between cells and collagen. Collagen (such as the peptides described herein) can promote adhesion between cells and the container in which the cells are cultured.
[0077] As used herein, the term “expression” includes any step involved in peptide production, including but not limited to: transcription, post-transcriptional modification, translation, post-translational modification, and secretion.
[0078] As used herein, the term “expression vector” refers to a straight or circular DNA molecule that contains a polynucleotide encoding a polypeptide and is operatively linked to a control sequence provided for its expression.
[0079] As used herein, the term "host cell" means any cell type that is readily transformed, transfected, transduced, etc., using nucleic acid constructs or expression vectors containing the polynucleotides of the present invention. The term "host cell" also encompasses any parental cell progeny that differs from the parent cell due to mutations occurring during replication.
[0080] As used herein, the term "nucleic acid" refers to a single-stranded or double-stranded nucleic acid molecule isolated from naturally occurring genes, modified to contain nucleic acid segments in a manner not originally present in nature, or synthesized. The nucleic acid molecule may contain one or more control sequences. Nucleic acids may be SEQ ID NO: 3, 6, 8, 9, 11, 14, 17, 20, 23, 26, 29, 32, 35, 38, 41, 44, 47, 50, 53, 56, or 59. Nucleic acids may be codon-optimized nucleic acids, such as those codon-optimized for expression in *E. coli* cells.
[0081] The term "operably linked" refers to a configuration in which a control sequence is placed at an appropriate position relative to the coding sequence of a polynucleotide, such that the control sequence directs the expression of the coding sequence.
[0082] The degree of association between two amino acid sequences or two nucleotide sequences is described by the parameter “sequence identity”. For the purposes of this invention, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48:443-453) implemented by the Needleman program in the EMBOSS software package (EMBOSS: European Molecular Biology Open Software Suite, Rice et al., 2000, Trends Genet. 16:276-277) (preferably version 5.0.0 or later) is used to determine the sequence identity between two amino acid sequences. The parameters used are a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. The output of Needleman labeled “longest identity” (obtained using the non-simplification option) is used as the identity percentage and calculated as follows:
[0083] (identical residues × 100) / (alignment length - total number of vacancies in the alignment)
[0084] For the purposes of this invention, the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, ibid.) implemented by the Needleman program in the EMBOSS software package (EMBOSS: European Molecular Biology Open Software Suite, Rice et al., 2000, ibid.) (preferably version 5.0.0 or later) is used to determine sequence identity between two deoxynucleotide sequences. The parameters used are a vacancy opening penalty of 10, a vacancy extension penalty of 0.5, and an EDNAFULL (EMBOSS version of NCBI NUC4.4) substitution matrix. The output of Needleman labeled "Longest Identity" (obtained using the non-simplified option) is used as the identity percentage and calculated as follows:
[0085] (identical deoxyribonucleotides x 100) / (alignment length - total number of vacancies in the alignment)
[0086] polypeptide
[0087] The present invention provides a polypeptide comprising one or more repeating units connected directly or via a linker, the repeating units comprising an amino acid sequence selected from the group consisting of: SEQ ID NO: 1, 4, 7, 12, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60. The variant may be (1) an amino acid sequence with one or more amino acid residues mutated in the amino acid sequence of SEQ ID NO:1, 4, 7, 12, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60, or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence of SEQ ID NO:1, 4, 7, 12, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60. For the purposes of this description of the polypeptide, the mutation may be selected from substitution, addition, insertion, or deletion. Preferably, the substitution is a conserved amino acid substitution.
[0088] The polypeptides described herein may contain multiple repeating units, for example, 2 to 50 repeating units, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50 repeating units.
[0089] The linkers in the peptides described herein may contain one or more amino acid residues, such as 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, or 1-2 amino acid residues.
[0090] The peptides described herein are recombinant collagens, particularly recombinant type XXII collagen, preferably possessing cell adhesion activity. The peptides described herein may be human-derived, and therefore are human recombinant type XXII collagen.
[0091] The polypeptides described herein may also comprise amino acid sequences selected from the group consisting of: SEQ ID NO: 2, 5, 7, 10, 13, 15, 16, 19, 22, 25, 28, 31, 34, 37, 40, 43, 46, 49, 52, 55, 58 or 29, wherein the variant is (1) an amino acid sequence in which one or more amino acid residues are mutated or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with the amino acid sequence.
[0092] Nucleic acid constructs
[0093] The present invention also relates to nucleic acid constructs comprising the nucleic acid of the present invention operatively linked to one or more control sequences, which, under conditions compatible with the control sequences, direct the expression of a coding sequence in a suitable host cell. Vectors may comprise the nucleic acid constructs.
[0094] Nucleic acids can be manipulated in various ways to provide peptide expression. Depending on the expression vector, manipulating the nucleic acid before insertion into the vector may be desirable or necessary. Techniques for modifying nucleic acids using recombinant DNA methods are well known in the art.
[0095] The control sequence may be a promoter, i.e., a polynucleotide recognized by the host cell for expressing the polypeptide encoding the present invention. The promoter contains a transcriptional control sequence that mediates the expression of the polypeptide. The promoter can be any nucleic acid exhibiting transcriptional activity in the host cell, including variants, truncated and heterozygous promoters, and can be obtained from a gene encoding an extracellular or intracellular polypeptide that is homologous or heterologous to that of the host cell.
[0096] Examples of suitable promoters for guiding the transcription of the vector or nucleic acid construct of the present invention in bacterial host cells are those obtained from the following promoters: Bacillus amyloliquefaciens α-amylase gene (amyQ), Bacillus licheniformis α-amylase gene (amyL), Bacillus licheniformis penicillinase gene (penP), Bacillus thermophilus malt amylase gene (amyM), Bacillus subtilis fructan sucrase gene (sacB), Bacillus subtilis xylA and xylB genes, Bacillus thuringiensis cryIIIA gene, Escherichia coli lac operon, and Escherichia coli trc promoter.
[0097] In the yeast host, useful promoters are obtained from the following genes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae galactokinase (GAL1), Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH1, ADH2 / GAP), Saccharomyces cerevisiae triose phosphate isomerase (TPI), Saccharomyces cerevisiae metallothionein (CUP1), and Saccharomyces cerevisiae 3-phosphate glycerate kinase.
[0098] The control sequence may also be a transcription terminator recognized by the host cell to terminate transcription. The terminator is operatively linked to the 3' end of a polynucleotide encoding a polypeptide. Any terminator that is functional in the host cell can be used in this invention.
[0099] The preferred terminator for bacterial host cells is obtained from the following genes: Bacillus clausti alkaline protease (aprH), Bacillus licheniformis α-amylase (amyL), and Escherichia coli ribosomal RNA (rrnB).
[0100] Preferred terminators for yeast host cells are derived from the following genes: *Saccharomyces cerevisiae* enolase, *Saccharomyces cerevisiae* cytochrome C (CYC1), and *Saccharomyces cerevisiae* glyceraldehyde-3-phosphate dehydrogenase. Other useful terminators for yeast host cells are described by Romanos et al. (1992, ibid.).
[0101] Control sequences can also be mRNA stabilizing regions downstream of the promoter and upstream of the gene's coding sequence, which enhance the expression of the gene.
[0102] Examples of suitable mRNA stable regions were obtained from the following genes: Bacillus thuringiensis cryIIIA gene (WO94 / 25612) and Bacillus subtilis SP82 gene (Hue et al., 1995, Journal of Bacteriology 177:3465-3471).
[0103] The control sequence can also be a leader sequence, which is a non-translated region of mRNA that is important for translation in the host cell. The leader sequence is operatively linked to the 5' end of a polynucleotide encoding a polypeptide. Any leader sequence that is functional in the host cell can be used.
[0104] The appropriate leader sequence for yeast host cells is obtained from the following genes: Saccharomyces cerevisiae enolase (ENO-1), Saccharomyces cerevisiae 3-phosphoglycerate kinase, Saccharomyces cerevisiae α-factor, and Saccharomyces cerevisiae alcohol dehydrogenase / glyceraldehyde-3-phosphate dehydrogenase (ADH2 / GAP).
[0105] The control sequence can also be a polyadenylation sequence, a sequence operatively linked to the 3' end of the polynucleotide and recognized by the host cell during transcription as a signal to add polyadenylation residues to the transcribed mRNA. Any polyadenylation sequence that is functional in the host cell can be used.
[0106] Useful polyadenylation sequences in yeast host cells are described by Guo and Sherman, 1995, Mol. Cellular Biol. [Molecular Cell Biology] 15: 5983-5990.
[0107] The control sequence can also be a signal peptide coding region encoding a signal peptide that is linked to the N-terminus of the polypeptide and guides the polypeptide into the secretory pathway of the cell. The 5' end of the polynucleotide coding sequence itself may contain a signal peptide coding sequence naturally linked to the coding sequence segment of the polypeptide within the translation reading frame. Alternatively, the 5' end of the coding sequence may contain a signal peptide coding sequence that is exogenous to the coding sequence. In cases where the coding sequence does not naturally contain a signal peptide coding sequence, an exogenous signal peptide coding sequence may be required. Alternatively, an exogenous signal peptide coding sequence may simply replace the natural signal peptide coding sequence to enhance polypeptide secretion. However, any signal peptide coding sequence that guides the secretory pathway of the expressed polypeptide into the host cell can be used.
[0108] The effective signal peptide coding sequences for bacterial host cells are derived from the following genes: maltose amylase produced by Bacillus NCIB 11837, subtilisin protease from Bacillus licheniformis, β-lactamase from Bacillus licheniformis, α-amylase from Bacillus thermophilus, neutral proteases (nprT, nprS, nprM) from Bacillus thermophilus, and prsA from Bacillus subtilis. Additional signal peptides are described by Simonen and Palva, 1993, Microbiological Reviews, 57:109-137.
[0109] Useful signal peptides in yeast host cells are obtained from genes such as *Saccharomyces cerevisiae* α-factor and *Saccharomyces cerevisiae* invertase. Sequences encoding other useful signal peptides are described by Romanos et al. (1992, ibid.).
[0110] expression vector
[0111] The present invention also relates to recombinant expression vectors comprising the nucleic acid, promoter, and transcription and translation termination signals of the present invention. The nucleic acid and control sequence can be linked together to produce a recombinant expression vector, which may include one or more convenient restriction sites for the insertion or substitution of a polynucleotide encoding the polypeptide at such sites. Alternatively, the polynucleotide can be expressed by inserting the nucleic acid or a nucleic acid construct containing the nucleic acid into a suitable vector for expression. In producing the expression vector, the coding sequence is located in the vector such that the coding sequence is operatively linked to a suitable control sequence for expression.
[0112] Recombinant expression vectors can be any vector (e.g., plasmids or viruses) that can readily undergo recombinant DNA procedures and induce polynucleotide expression. The choice of vector will typically depend on its compatibility with the host cell to which it will be introduced. Vectors can be linear or closed circular plasmids.
[0113] The vector can be a self-replicating vector, that is, a vector that exists as an extrachromosomal entity and replicates independently of chromosome replication, such as a plasmid, extrachromosomal element, microchromosome, or artificial chromosome. The vector can contain any means to ensure self-replication. Alternatively, the vector can be one that integrates into the genome when introduced into a host cell and replicates along with one or more chromosomes in which it has already been integrated. Furthermore, a single vector or plasmid, or two or more vectors or plasmids collectively containing the total DNA of the host cell genome to be introduced, or transposons can be used.
[0114] The vector preferably contains one or more selective markers that allow for convenient selection of cells such as transformed cells, transfected cells, and transduced cells. A selective marker is a gene whose product provides resistance to biocides or viruses, resistance to heavy metals, or prototrophic auxotrophic traits, etc.
[0115] Examples of selective bacterial markers include the dal gene in Bacillus licheniformis or Bacillus subtilis, or markers that confer antibiotic resistance (such as ampicillin, chloramphenicol, kanamycin, neomycin, spectinomycin, or tetracycline resistance). Suitable markers for yeast host cells include, but are not limited to: ADE2, HIS3, LEU2, LYS2, MET3, TRP1, and URA3.
[0116] Selective labeling can be a biselective labeling system as described in WO 2010 / 039889. In one aspect, biselective labeling is the hph-tk biselective labeling system.
[0117] Vectors may contain elements that allow the vector to integrate into the host cell genome or to replicate autonomously in the cell independently of the genome.
[0118] For integration into the host cell genome, the vector can rely on a polynucleotide sequence encoding the polypeptide or any other element of the vector for integration into the genome via homologous or non-homologous recombination. Alternatively, the vector can contain additional polynucleotides to guide integration into the host cell genome at a precise location on the chromosome via homologous recombination. To increase the likelihood of integration at a precise location, the integrative element should contain a sufficient number of nucleic acids, such as 100 to 10,000 base pairs, 400 to 10,000 base pairs, and 800 to 10,000 base pairs, that have high sequence identity with the corresponding target sequence to enhance the probability of homologous recombination. The integrative element can be any sequence homologous to the target sequence within the host cell genome. Moreover, the integrative element can be a non-coding or coding polynucleotide. On the other hand, the vector can integrate into the host cell genome via non-homologous recombination.
[0119] For autonomous replication, the vector may further include an origin of replication, which enables the vector to replicate autonomously within the host cell discussed in this context. The origin of replication can be any plasmid replicon that mediates autonomous replication and functions within the cell. The terms "origin of replication" or "plasmid replicon" refer to the polynucleotide that enables a plasmid or vector to replicate in vivo.
[0120] Examples of bacterial origins of replication are the origins of replication of plasmids pBR322, pUC19, pACYC177, and pACYC184, which allow replication in Escherichia coli, and the origins of replication of plasmids pUB110, pE194, pTA1060, and pAMβ1, which allow replication in Bacillus.
[0121] Examples of replication origins used in yeast host cells are 2-micron replication origins, ARS1, ARS4, combinations of ARS1 and CEN3, and combinations of ARS4 and CEN6.
[0122] More than one copy of the polynucleotide of the present invention can be inserted into host cells to enhance polypeptide production. An increased copy number of the polynucleotide can be obtained by integrating at least one additional copy of the sequence into the host cell genome or by including an amplifiable selective marker gene along with the polynucleotide, wherein cells containing the amplified copy of the selective marker gene and thus additional copies of the polynucleotide can be selected by culturing cells in the presence of a suitable selective reagent.
[0123] The procedures for connecting the above-described elements to construct the recombinant expression vector of the present invention are well known to those skilled in the art (see, for example, Sambrook et al., 1989).
[0124] host cells
[0125] This invention also relates to recombinant host cells containing polynucleotides of the invention operably linked to one or more control sequences that direct the production of polypeptides of the invention. A construct or vector containing the polynucleotide is introduced into the host cell such that the construct or vector is maintained as a chromosomal integrase or as an autonomously replicating extrachromosomal vector, as previously described. The term "host cell" encompasses any parental cell progeny that is not identical to the parental cell due to mutations occurring during replication. The selection of the host cell will depend largely on the gene encoding the polypeptide and its origin.
[0126] The host cell can be any cell useful in the recombinant production of the polypeptides of the present invention, such as prokaryotes or eukaryotes.
[0127] Prokaryotic host cells can be any Gram-positive or Gram-negative bacteria. Gram-positive bacteria include, but are not limited to: Bacillus, Clostridium, Enterococcus, Bacillus aeruginosa, Lactobacillus, Lactococcus, Bacillus cereus, Staphylococcus, Streptococcus, and Streptomyces. Gram-negative bacteria include, but are not limited to: Campylobacter, Escherichia coli, Flavobacterium, Fusobacterium, Helicobacter, Staphylococcus, Neisseria, Pseudomonas, Salmonella, and Ureaplasma.
[0128] The host cell can also be a eukaryotic cell, such as a mammalian, insect, plant, or fungal cell. Plant cells in this article do not include plant cells capable of regenerating into plants. Animal cells do not include cells capable of producing animal bodies.
[0129] The host cell can be a fungal cell, such as those belonging to the phyla Basidiomycota, Chytridiomycota, Zygomycota, and Oomycota. Fungal host cells can also be yeast cells, including ascosporogenous yeast (Endomycetales), basidiosporogenous yeast, and yeasts belonging to the class Fungi Imperfecti (Blastomycetes). Yeast host cells can be cells from the genera *Candida*, *Hansenula*, *Kluyveromyces*, *Pichia*, *Saccharomyces*, *Schizosaccharomyces*, or *Yarrowia*, such as *Kluyveromyces lactis*, *Saccharomyces carlsbergensis*, *Saccharomyces diastaticus*, *Saccharomyces douglasii*, *Saccharomyces kluyveri*, *Saccharomyces norbensis*, *Saccharomyces oviformis*, or *Yarrowia polytica*.
[0130] Production methods
[0131] This invention also relates to the production of the polypeptides described herein, comprising:
[0132] (1) Culture the host cells described herein under suitable culture conditions;
[0133] (2) Harvesting host cells and / or culture medium containing polypeptides; and
[0134] (3) Purify the polypeptide.
[0135] The host cells are cultured in a suitable nutrient medium for producing the peptide using methods known in the art. For example, cells can be cultured by shake flask culture or by small-scale or large-scale fermentation (including continuous, batch, fed-batch, or solid-state fermentation) in a laboratory or industrial fermenter, in a suitable medium and under conditions that allow for peptide expression and / or isolation. Using procedures known in the art, the culture occurs in a suitable nutrient medium containing carbon and nitrogen sources and inorganic salts. Suitable media are available from commercial suppliers or can be prepared according to publicly available compositions (e.g., in the catalogue of the U.S. Center for Type Culture Collection). If the peptide is secreted into the nutrient medium, it can be recovered directly from the medium. If the peptide is not secreted, it can be recovered from cell lysates.
[0136] Peptides can be detected using methods known in the art that are specific to peptides. These detection methods include, but are not limited to, the use of specific antibodies, the formation of enzyme products, or the disappearance of enzyme substrates. For example, enzyme assays can be used to determine the activity of a peptide.
[0137] Peptides can be recovered using methods known in the art. For example, peptides can be recovered from nutrient media through conventional procedures, including but not limited to collection, centrifugation, filtration, extraction, spray drying, evaporation, or precipitation. In one aspect, fermentation broth containing peptides can be recovered.
[0138] Peptides can be purified using a variety of procedures known in the art, including but not limited to chromatography (e.g., ion exchange chromatography, affinity chromatography, hydrophobic chromatography, focusing chromatography, and size exclusion chromatography), electrophoresis procedures (e.g., preparative isoelectric focusing electrophoresis), differential dissolution (e.g., ammonium sulfate precipitation), SDS-PAGE, or extraction, to obtain substantially pure peptides.
[0139] Step (1) may include one or more of the following steps: constructing an expression plasmid, for example, inserting the encoding nucleotide sequence into the pET-28a-Trx-His expression vector to obtain a recombinant expression plasmid. The successfully constructed expression plasmid can be transformed into Escherichia coli cells (e.g., Escherichia coli competent cells BL21(DE3)). The specific process may be as follows: (1) Take the plasmid to be transformed and add it into Escherichia coli competent cells BL21(DE3); (2) Place the mixture on ice (e.g., 10-60 min, e.g., 30 min), then heat shock in a water bath (e.g., at 40-50℃, e.g., 42℃, 45-90 s), remove it and place it on ice (e.g., 1-5 min, e.g., 2 min); (3) Add liquid LB medium and then culture (e.g., at 35-40℃, e.g., 37℃, 150-300 rpm, e.g., 220 rpm for 40-80 min, e.g., 60 min); (4) Spread the bacterial solution and select single colonies. For example, spread the bacterial suspension evenly on an LB agar plate containing ampicillin sodium, and incubate the plate in a 37°C incubator for 15-17 hours until uniformly sized colonies grow.
[0140] Step (2) may include culturing single colonies in LB medium containing an antibiotic stock solution (e.g., in a constant temperature shaker at 150-300 rpm, e.g., 220 rpm, 35-40°C, e.g., 37°C for 5-10 h, e.g., 7 h). The cultured shake flask is then cooled to 10-20°C, e.g., 16°C, and IPTG is added to induce expression for a period of time before collecting the cells (e.g., by centrifugation).
[0141] Step (3) may include resuspending the bacterial cells in a equilibration working solution, cooling the bacterial solution to ≤15°C, and homogenizing it (e.g., high-pressure homogenization, e.g., 1-5 times, e.g., 2 times), and separating the supernatant from the homogenized bacterial solution. The equilibration working solution may contain 100-500mM sodium chloride, 10-50mM Tris and 10-50mM imidazole, with a pH of 7-9. For example, the concentration of sodium chloride can be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The concentration of Tris can be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The concentration of imidazole can be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The pH can be 7, 7.5, 8, 8.5 or 9.
[0142] Step (3) may include purification and enzymatic digestion of the peptide. Purification may be crude purification, including Ni-agarose gel column purification of the supernatant to obtain an eluent containing the target protein. Crude purification may include washing the column material with water, for example, 2-10 column volumes (CV), such as 5 CV. The column material may be equilibrated with an equilibration buffer (200 mM sodium chloride, 25 mM Tris, 20 mM imidazole, pH 8.0), for example, 2-10 CV, such as 5 CV. The equilibration buffer may contain 100-500 mM sodium chloride, 10-50 mM Tris and 10-50 mM imidazole, pH 7-9. For example, the concentration of sodium chloride can be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The concentration of Tris can be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The concentration of imidazole can be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The pH can be 7, 7.5, 8, 8.5 or 9.
[0143] Step (3) may include adding the supernatant to the column and washing away any contaminating proteins with a washing solution. The washing solution may contain 100-500 mM sodium chloride, 10-50 mM Tris, and 10-50 mM imidazole, with a pH of 7-9. For example, the concentration of sodium chloride may be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The concentration of Tris can be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The concentration of imidazole can be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The pH can be 7, 7.5, 8, 8.5, or 9. Then, an eluent can be added and the flow-through collected. The eluent can contain 100-500 mM sodium chloride, 10-50 mM Tris, 100-500 mM imidazole, and pH 8.0. For example, the concentration of sodium chloride can be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The concentration of Tris can be 10, 15, 20, 25, 30, 35, 40, 45, or 50 nM. The concentration of imidazole can be 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, or 490 nM. The pH can be 7, 7.5, 8, 8.5, or 9.
[0144] Enzymatic digestion may include adding TEV enzyme (at a protein to TEV ratio of 10-100:1, e.g., 50:1, at 10-20°C, e.g., 16°C, for 2-8 hours, e.g., 4 hours). The digested protein solution is then dialyzed, for example, by placing it in a dialysis bag and dialyzing at 1-6°C, e.g., 4°C, for 1-8 hours, e.g., 2 hours, before transferring it to fresh dialysate and dialyzing overnight at 1-6°C, e.g., 4°C.
[0145] Purification can include fine purification (e.g., protein isoelectric point > 8.0). Preferably, fine purification includes gradient elution of the eluent containing the target protein or the enzymatically digested product (e.g., product after enzymatic digestion and dialyzing) using a strong anion exchange chromatography column (e.g., pH 7, 7.5, 8, 8.5, or 9). Gradient elution includes 0-15% solution B for 1-5 minutes followed by 1-5 holds, e.g., 3 column volumes; 15-30% solution B for 1-5 minutes followed by 1-5 holds, e.g., 3 column volumes; 30-50% solution B for 1-5 minutes followed by 1-5 holds, e.g., 3 column volumes; and 50-100% solution B for 1-5 minutes followed by 1-5 holds, e.g., 3 column volumes. Solution B may contain 10-50 mM Tris, 0.5-5 M sodium chloride, and pH 7-9. For example, the concentration of Tris may be 15, 20, 25, 30, 35, 40, or 45 mM. The sodium chloride concentration is 1, 2, 3, or 4 M. The pH can be 7, 7.5, 8, 8.5, or 9. Purification may involve equilibrating the column with solution A and loading the sample, followed by gradient elution. Solution A may contain 10-50 mM Tris, 10-50 mM sodium chloride, and a pH of 7-9. For example, the Tris concentration may be 15, 20, 25, 30, 35, 40, or 45 mM. The sodium chloride concentration may be 15, 20, 25, 30, 35, 40, or 45 mM. The pH can be 7, 7.5, 8, 8.5, or 9.
[0146] Purification may include reverse-coated nickel column purification (e.g., protein isoelectric point < 8.0). Reverse-coated nickel column purification may include purification of the enzyme-digested product (e.g., the product after dialysis) onto a Ni-agarose gel column. The eluent may contain 10-50 mM (e.g., 15, 20, 25, 30, 35, 40, or 45 mM) Tris, 10-50 mM (e.g., 15, 20, 25, 30, 35, 40, or 45 mM) sodium chloride, 0.5-5 M (e.g., 1, 2, 3, or 4 M) imidazole, and pH 7-9 (e.g., 7, 7.5, 8, 8.5, or 9).
[0147] The following embodiments are provided to further illustrate the present invention.
[0148] Example
[0149] The present invention is further illustrated by the following embodiments, but any embodiment or combination thereof should not be construed as limiting the scope or implementation of the present invention. The scope of the present invention is defined by the appended claims, and those skilled in the art will clearly understand the scope defined by the claims in conjunction with this specification and common knowledge in the art. Without departing from the spirit and scope of the present invention, those skilled in the art can make any modifications or changes to the technical solutions of the present invention, and such modifications and changes are also included within the scope of the present invention.
[0150] Example 1: Construction, expression, and screening of recombinant type XXII humanized collagen fragments
[0151] 1. Large-scale functional region screening was conducted to obtain the target gene functional regions of the following different recombinant XXII type humanized collagen.
[0152] 1) C22a amino acid sequence:
[0153] gEKgEmgvagpmglpgpKgDigaigpvgapgpKgEKgDvgigpfgqgEKgEKgslglpg ppgRDgsKgmRgEpgElgEpglpgEvgmRgpqgppglpgppgRvgapglqgERgEKgtRgEKgERglDgfpgKpgDtgqqgRpgps(SEQ ID NO:7)
[0154] Nucleotide sequence:
[0155] GGAGAAAAAGGGGAGATGGGTGTCGCAGGTCCGATGGGTCTGCCGGGCCCGAAAGGGGACATCGGTGCGATCGGCCCGGTGGGTGCGCCTGGCCCGAAAGGCGAAAAGGGTGACGTTGGTATTGGTCCGTTTGGTCAGGGTGAGAAGGGCGAAAAAGGCTCCCTGGGTTTGCCGGGCCCACCGGGTCGCGATGGTAGCAAAGGTATGAGAGGCGAGCCG GGCGAGCTGGGTGAACCGGGTCTGCCGGGCGAGGTGGGCATGCGTGGTCCGGCAGGGTCCGCCTGGCTTGCCGGGCCCACCGGGCCGTGTTGGCGCTCCAGGCCTGCAGGGCGAACGTGGTGAAAAGGGCACGCGTGGTGAAAAGGGCGAGCGCGGTCTTGATGGTTTCCCGGGCAAACCGGGCGACACCGGTCAACAAGGTCGTCCGGGTCCGAGC(SEQ ID NO:8)
[0156] 2) C22b amino acid sequence:
[0157] gEKgEmgvagpmglpgpKgDigaigpvgapgpKgEKgDvgigpfgqgEKgEKgslglpg ppgRDgsKgmRgEpgElgEpglpgEvgmRgpqgppglpgppgRvgapglqgERgEKgtRgEKgERglDgfpgKpgDtgqqgRpgps(SEQ ID NO:7)
[0158] gEKgEmgvagpmglpgpKgDigaigpvgapgpKgEKgDvgigpfgqgEKgEKgslglpg ppgRDgsKgmRgEpgElgEpglpgEvgmRgpqgppglpgppgRvgapglqgERgEKgtRgEKgERglDgfpgKpgDtgqqgRpgps(The amino acid sequence of the repeating unit of C22b is SEQ ID NO:7, the number of repetitions is 2, and the amino acid sequence of C22b is SEQ ID NO:10)
[0159] Nucleotide sequence:
[0160] GGAGAAAAAGGGGAGATGGGCGTCGCAGGTCCGATGGGTTTGCCGGGCCCGAAA
[0161] GGTGATATTGGTGCGATTGGTCCAGTTGGTGCGCCAGGTCCGAAAGGTGAAAAGG
[0162] GCGACGTGGGCATCGGCCCGTTCGGCCAAGGCGAAAAGGGCGAAAAAGGTTCGCT
[0163] GGGTCTGCCGGGGCCGCCGGGCCGCGATGGCAGCAAAGGCATGCGTGGTGAGCCG
[0164] GGCGAGCTGGGTGAGCCGGGCCTGCCGGGCGAGGTGGGCATGCGTGGGCCTCAGG
[0165] GTCCACCGGGTCTGCCGGGTCCTCCGGGTCGTGTTGGCGCGCCAGGCTTACAGGGT
[0166] GAACGTGGTGAGAAAGGTACGCGTGGCGAGAAGGGCGAGCGCGGTCTGGACGGC
[0167] TTCCCGGGTAAACCGGGGGATACCGGTCAACAAGGTCGCCCTGGCCCGAGCGGTG
[0168] AAAAGGGTGAAATGGGCGTGGCTGGTCCGATGGGTCTTCCGGGCCCGAAAGGCGA
[0169] CATCGGTGCGATTGGTCCGGTTGGTGCCCCTGGGCCGAAAGGTGAAAAGGGTGAC
[0170] GTCGGGATCGGTCCGTTTGGTCAGGGTGAGAAGGGTGAGAAGGGCTCTCTGGGCT
[0171] TGCCGGGCCCACCGGGGCGTGATGGTTCCAAAGGTATGCGTGGTGAACCGGGCGA
[0172] ACTGGGTGAGCCGGGCCTGCCGGGTGAAGTTGGTATGAGAGGTCCGCAGGGCCCG
[0173] CCGGGCTTGCCGGGGCCGCCGGGACGCGTGGGTGCTCCGGGCCTCCAGGGTGAGC
[0174] GTGGCGAAAAGGGAACCCGTGGTGAAAAGGGCGAGCGCGGTCTGGATGGTTTTCCGGGAAAACCGGGCGACACCGGTCAGCAAGGCCGTCCGGGCCCAAGC(SEQ ID NO:11)
[0175] 3) C22c amino acid sequence:
[0176] gEKgEKgslglpgppgRDgsKgmRgEpgElgEpglpgEvgmRgpqgppglpgppgRvg apglqgERgEKgtRgEKgERglDgfpgKpgDtgqqgRpgps(SEQ ID NO:12)
[0177] The amino acid sequence of the repeating unit of C22c is SEQ ID NO: 12, the number of repetitions is 2, and the C22c amino acid sequence is SEQ ID NO: 13
[0178] Nucleotide sequence:
[0179] GGAGAAAAAGGGGAGAAAGGCTCTCTGGGCCTGCCGGGCCCACCGGGCCGTGAC
[0180] GGCAGCAAAGGTATGCGCGGTGAACCGGGAGAGCTGGGTGAACCGGGTTTACCGG
[0181] GCGAAGTTGGTATGAGAGGTCCGCAGGGCCCTCCGGGCCTGCCGGGTCCGCCCGG
[0182] CCGTGTTGGTGCTCCGGGCCTTCAGGGTGAACGTGGTGAAAAGGGCACCCGTGGT
[0183] GAAAAGGGTGAGCGTGGTCTCGACGGTTTTCCGGGTAAACCGGGTGATACGGGTC
[0184] AGCAAGGTCGCCCTGGCCCGAGCGGTGAGAAGGGTGAAAAGGGCTCGTTGGGTCT
[0185] GCCGGGCCCGCCGGGCCGTGACGGCTCCAAAGGTATGCGTGGTGAGCCGGGTGAA
[0186] CTGGGCGAGCCGGGCTTGCCGGGCGAGGTGGGCATGCGTGGTCCGCAGGGTCCGC
[0187] CGGGCTTGCCGGGACCACCGGGCCGCGTGGGTGCGCCAGGCCTGCAGGGTGAGCG
[0188] CGGTGAGAAGGGCACCCGTGGCGAAAAAGGTGAGCGCGGTCTGGATGGTTTCCCGGGTAAACCGGGGGATACCGGTCAACAAGGCCGTCCGGGCCCAAGC(SEQ ID NO:14)
[0189] 4) C22d amino acid sequence:
[0190] gEKgEKgslglpgppgRDgsKgmRgEpgElgEpglpgEvgmRgpqgppglpgppgRvg apglqgERgEKgtRgEKgERglDgfpgKpgDtgqqgRpgps(SEQ ID NO:12)
[0191] gEKgEKgslglpgppgRDgsKgmRgEpgElgEpglpgEvgmRgpqgppglpgppgRvg
[0192] apglqgERgEKgtRgEKgERglDgfpgKpgDtgqqgRpgps
[0193] gEKgEKgslglpgppgRDgsKgmRgEpgElgEpglpgEvgmRgpqgppglpgppgRvgapglqgERgEKgtRgEKgERglDgfpgKpgDtgqqgRpgps(The amino acid sequence of the repeating unit of C22d is SEQ ID NO:12, the number of repeats is 3, and the C22d amino acid sequence is SEQ ID NO:16)
[0194] Nucleotide sequence:
[0195] GGAGAAAAAGGGGAGAAGGGCTCCCTGGGGCTGCCGGGTCCGCCGGGACGCGAC
[0196] GGCTCCAAAGGTATGCGTGGTGAACCGGGTGAGCTGGGTGAACCGGGCCTCCCGG
[0197] GCGAGGTTGGTATGCGTGGACCGCAAGGTCCACCGGGATTGCCGGGCCCGCCGGG
[0198] CCGCGTGGGTGCTCCTGGCCTGCAGGGTGAACGTGGTGAAAAGGGCACCCGTGGT
[0199] GAGAAGGGCGAACGTGGCCTGGACGGCTTCCCGGGTAAACCGGGTGACACCGGTC
[0200] AGCAAGGTCGCCCTGGCCCAAGCGGTGAGAAAGGTGAGAAAGGCTCTCTGGGTTT
[0201] ACCGGGCCCACCGGGCCGTGATGGCAGCAAAGGTATGCGTGGCGAGCCGGGTGAA
[0202] CTGGGCGAACCGGGTTTGCCGGGCGAGGTGGGCATGCGTGGTCCGCAGGGTCCGC
[0203] CTGGTTTGCCGGGCCCGCCGGGCAGAGTTGGCGCACCGGGCCTGCAAGGTGAGCG
[0204] CGGCGAAAAGGGCACCCGTGGTGAAAAGGGCGAACGCGGTTTGGACGGTTTCCCG
[0205] GGTAAACCGGGCGACACGGGTCAGCAGGGTCGTCCGGGTCCAAGCGGCGAAAAG
[0206] GGTGAAAAGGGCAGCCTGGGTCTGCCGGGACCGCCCGGCCGGGATGGTTCGAAAG
[0207] GTATGAGAGGTGAGCCGGGCGAGCTGGGTGAACCGGGCCTGCCGGGTGAGGTTGG
[0208] TATGCGTGGTCCGCAGGGTCCGCCGGGTTTGCCGGGCCCACCGGGGCGCGTCGGG
[0209] GCGCCAGGGCTGCAGGGTGAGCGCGGTGAAAAAGGCACTCGTGGCGAGAAGGGT
[0210] GAGCGTGGCCTTGATGGTTTTCCGGGCAAACCGGGCGATACCGGTCAACAAGGTCGTCCGGGTCCGAGC(SEQ ID NO:17)
[0211] 5) C22e amino acid sequence:
[0212] gEqgapgpRghqgapgppgaRgpigpEgRDgppglqglRgKKgDmgppgip(SEQ ID NO:18)
[0213] gEqgapgpRghqgapgppgaRgpigpEgRDgppglqglRgKKgDmgppgip
[0214] gEqgapgpRghqgapgppgaRgpigpEgRDgppglqglRgKKgDmgppgip
[0215] gEqgapgpRghqgapgppgaRgpigpEgRDgppglqglRgKKgDmgppgip
[0216] gEqgapgpRghqgapgppgaRgpigpEgRDgppglqglRgKKgDmgppgip
[0217] gEqgapgpRghqgapgppgaRgpigpEgRDgppglqglRgKKgDmgppgip(The amino acid sequence of the repeating unit of C22e is SEQ ID NO:18, and the number of repeating units is 6; the amino acid sequence of C22e is SEQ ID NO:19)
[0218] Nucleotide sequence:
[0219] GGAGAACAAGGGGCGCCTGGCCCGCGTGGTCACCAGGGTGCGCCGGGTCCGCCGGGCGCTCGCGGTCCGATCGGTCCAGAGGGCCGTGACGGTCCACCGGGCCTGCAGGGCTTGCGCGGTAAAAAGGGTGACATGGGTCCGCCGGGCATTCCGGGTGAACAAGGTGCGCCGGGTCCGAGAGGCCATCAGGGTGCGCCAGGCCCACCGGGCGCGCGTGGTCCGATTGGTCCGGAAGGCCGTGATGGTCCGCCAGGTTTGCAAGGTCTGCGCGGTAAAAAGGGTGATATGGGTCCGCCGGGTATCCCGGGCGAACAGGGCGCACCGGGCCCGCGTGGCCACCAGGGCGCACCTGGTCCGCCGGGCGCTCGCGGTCCGATCGGCCCGGAGGGTCGCGACGGCCCGCCGGGCCTGCAAGGCCTGCGTGGTAAAAAGGGTGACATGGGCCCGCCAGGCATCCCGGGTGAGCAGGGTGCCCCTGGTCCGCGTGGTCACCAAGGTGCCCCTGGTCCGCCGGGCGCGCGTGGCCCGATTGGTCCAGAAGGTCGTGATGGTCCGCCGGGCCTGCAGGGTTTACGCGGTAAAAAGGGTGATATGGGTCCCCCGGGTATTCCGGGCGAGCAGGGCGCACCGGGCCCGCGCGGTCATCAAGGGGCGCCGGGCCCACCGGGTGCGCGTGGCCCCATCGGTCCGGAGGGCAGAGACGGTCCGCCGGGGTTGCAAGGTCTGCGTGGCAAAAAAGGCGACATGGGTCCGCCGGGTATCCCGGGTGAACAGGGTGCTCCGGGTCCGCGTGGCCATCAGGGCGCGCCAGGCCCGCCGGGCGCTCGCGGTCCTATTGGTCCGGAGGGTCGTGATGGTCCGCCGGGCCTTCAAGGTCTGCGTGGTAAGAAGGGTGATATGGGTCCGCCGGGCATCCCG(SEQ ID NO:20)
[0220] 6) C22f Amino acid sequence:
[0221] gppgppgvpgppgpggspglpgEigfpgKpgppgptgppgKDgpngppgppgtKgEpgERgEDglpgKpglRgEigEqglagRpgEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEKgEagpagpp (SEQ ID NO: 21)
[0222] gppgppgvpgppgpggspglpgEigfpgKpgppgptgppgKDgpngppgppgtKgEpgERgEDglpgKpglRgEigEqglagRpgEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEKgEagpagpp(The amino acid sequence of the repeating unit of C22f is SEQ ID NO:21, the number of repeats is 2, and the amino acid sequence of C22f is SEQ ID NO:22)
[0223] Nucleotide sequence:
[0224] GGACCACCCGGGCCGCCGGGCGTGCCGGGCCCGCCAGGCCCAGGTGGTAGCCCGGGCTTGCCGGGCGAGATCGGTTTCCCGGGCAAGCCGGGCCCGCCGGGTCCGACGGGTCCCCCGGGGAAAGATGGTCCAAATGGTCCGCCGGGCCCGCCGGGAACCAAAGGTGAGCCGGGTGAGCGCGGTGAGGATGGTCTGCCGGGCAAGCCGGGTCTGCGTGGCGAAATTGGTGAACAAGGTCTGGCAGGCAGACCGGGCGAAAAGGGCGAAGCGGGTCTCCCGGGCGCGCCTGGCTTCCCGGGTGTGCGTGGTGAGAAAGGCGATCAGGGAGAGAAAGGTGAGCTGGGTCTGCCGGGTTTAAAAGGTGACCGTGGTGAAAAGGGCGAAGCTGGTCCGGCAGGTCCACCGGGCCCGCCGGGCCCGCCAGGCGTTCCGGGCCCGCCGGGTCCGGGTGGTTCCCCGGGCCTGCCGGGTGAAATTGGCTTTCCGGGTAAACCGGGTCCGCCTGGCCCGACCGGTCCGCCGGGTAAGGACGGCCCCAACGGTCCGCCGGGCCCTCCGGGCACCAAAGGCGAACCGGGTGAACGTGGTGAAGATGGCTTGCCGGGGAAACCGGGCCTGCGTGGTGAGATCGGCGAGCAGGGTCTTGCCGGTCGCCCAGGCGAAAAGGGTGAGGCGGGTTTGCCGGGAGCGCCTGGTTTTCCGGGCGTTCGTGGTGAAAAGGGTGACCAAGGTGAGAAGGGTGAACTGGGTCTGCCGGGCCTGAAAGGCGACCGCGGCGAGAAGGGTGAGGCGGGTCCAGCTGGCCCGCCG(SEQ ID NO:23)
[0225] 7) C22g amino acid sequence:
[0226] gEigfpgKpgppgptgppgKDgpngppgppgtKgEpgERgEDglpgKpglRgEigEqglagRpgEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEKgEa (SEQ ID NO: 24)
[0227] gEigfpgKpgppgptgppgKDgpngppgppgtKgEpgERgEDglpgKpglRgEigEqglagRpgEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEKgEa
[0228] gEigfpgKpgppgptgppgKDgpngppgppgtKgEpgERgEDglpgKpglRgEigEqglagRpgEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEKgEa(The repeating unit amino acid sequence of C22g is SEQ ID NO:24, with 3 repeats; the amino acid sequence of C22g is SEQ ID NO:25)
[0229] Nucleotide sequence:
[0230] GGAGAAATAGGGTTCCCGGGCAAGCCGGGCCCGCCGGGCCCAACGGGTCCGCCAGGCAAGGATGGTCCGAACGGTCCGCCGGGTCCGCCGGGCACCAAAGGTGAGCCGGGCGAACGTGGCGAAGATGGCCTGCCGGGTAAGCCGGGCTTGAGAGGTGAGATCGGCGAACAAGGCCTTGCCGGTCGTCCGGGCGAGAAGGGTGAGGCGGGTTTACCGGGTGCCCCTGGTTTTCCGGGCGTGCGCGGTGAGAAGGGTGACCAAGGTGAAAAAGGTGAGCTGGGTCTGCCGGGCCTGAAAGGTGATCGTGGGGAGAAAGGAGAAGCGGGTGAAATTGGTTTTCCGGGTAAGCCGGGTCCACCGGGTCCAACCGGTCCGCCGGGTAAAGATGGTCCGAATGGTCCGCCGGGCCCACCTGGCACCAAGGGTGAGCCGGGCGAACGTGGCGAGGACGGCCTGCCGGGCAAACCGGGCCTGCGCGGCGAAATCGGCGAGCAGGGTTTGGCGGGACGCCCTGGCGAGAAAGGCGAGGCGGGTCTGCCGGGTGCACCGGGATTCCCGGGCGTCCGTGGTGAGAAGGGCGACCAGGGCGAAAAGGGTGAACTGGGCCTGCCGGGCCTGAAAGGCGATCGTGGTGAAAAAGGCGAAGCTGGTGAAATTGGCTTCCCGGGTAAGCCGGGGCCGCCGGGCCCGACTGGTCCGCCGGGTAAGGACGGTCCCAACGGCCCGCCGGGACCTCCGGGAACCAAAGGCGAACCGGGCGAGCGCGGTGAGGATGGTTTGCCTGGCAAACCGGGCTTGCGTGGCGAGATCGGTGAACAAGGTCTGGCAGGTCGTCCGGGTGAGAAGGGTGAGGCTGGTCTGCCGGGTGCGCCAGGCTTTCCGGGTGTTCGTGGTGAGAAAGGCGACCAGGGTGAAAAGGGTGAACTGGGTTTGCCGGGCCTCAAAGGGGACCGCGGCGAAAAGGGCGAGGCG(SEQ IDNO:26)
[0231] 8) C22h amino acid sequence:
[0232] gtKgEpgERgEDglpgKpglRgEigEqglagRpgEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEKgEa (SEQ ID NO: 27)
[0233] gtKgEpgERgEDglpgKpglRgEigEqglagRpgEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEKgEa
[0234] gtKgEpgERgEDglpgKpglRgEigEqglagRpgEKgEaglpgapgfpgvRgEKgDqgEKgElglpglKgDRgEKgEa(The amino acid sequence of the repeating unit of C22h is SEQ ID NO:27, the number of repeats is 3, and the amino acid sequence of C22h is SEQ ID NO:28)
[0235] Nucleotide sequence:
[0236] GGGACAAAAGGAGAGCCGGGCGAACGTGGTGAGGACGGCTTGCCGGGCAAACCGGGTCTGCGCGGTGAGATCGGTGAACAAGGTCTGGCGGGTCGCCCAGGCGAAAAGGGTGAGGCGGGTTTACCGGGAGCCCCCGGCTTTCCGGGTGTCAGAGGCGAGAAAGGTGATCAAGGTGAAAAAGGTGAGCTGGGCCTGCCGGGCCTGAAAGGCGATCGTGGTGAAAAGGGTGAGGCTGGCACGAAAGGCGAGCCGGGCGAGCGTGGTGAGGATGGTTTGCCGGGTAAACCGGGCCTGCGTGGTGAAATTGGTGAACAGGGCTTGGCGGGTCGTCCGGGTGAGAAGGGCGAAGCTGGTTTGCCAGGCGCACCGGGCTTCCCGGGGGTGCGCGGTGAGAAAGGCGACCAGGGCGAAAAGGGCGAACTGGGTCTGCCGGGTCTGAAAGGCGACCGCGGTGAGAAGGGGGAGGCGGGCACCAAAGGCGAACCGGGTGAGCGCGGCGAAGATGGTTTGCCGGGCAAACCGGGCCTGCGTGGCGAGATCGGCGAGCAGGGTCTTGCCGGTCGTCCGGGTGAGAAGGGTGAGGCGGGTCTGCCTGGTGCACCGGGATTCCCGGGTGTTCGTGGTGAAAAGGGTGACCAAGGCGAAAAGGGCGAACTGGGTCTGCCAGGTCTCAAGGGCGACCGTGGTGAAAAGGGCGAAGCG(SEQ ID NO:29)
[0237] 9) C22i Amino acid sequence:
[0238] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgahgapgaagnpgapghv gapgpsgppgsvgapglRgtpgKDgERgEKgaagEEgspgpvgpRgDpgapglpgppgKg(SEQ ID NO:30)
[0239] The amino acid sequence of the repeating unit of C22i is SEQ ID NO:30, the number of repeats is 2, and the amino acid sequence of C22i is SEQ ID NO:31
[0240] Nucleotide sequence:
[0241] GGAGAACAAGGGCCGAAAGGTGAGAAGGGTGACCCGGGTTTGCCGGGTGAGCCGGGTCTGCAGGGCCGTCCGGGCGAACTGGGTCCACAGGGTCCGACCGGCCCACCGGGCGCGAAAGGTCAAGAGGGCGCGCATGGTGCTCCGGGTGCGGCAGGCAATCCGGGTGCGCCCGGTCATGTTGGTGCCCCGGGCCCGAGCGGTCCGCCTGGCTCGGTTGGTGCGCCGGGCCTGAGAGGCACCCCGGGCAAAGATGGTGAGCGCGGTGAAAAAGGTGCAGCAGGTGAGGAGGGTTCTCCGGGCCCGGTCGGTCCGCGTGGTGACCCGGGCGCGCCGGGTTTGCCGGGTCCGCCAGGTAAAGGCGGTGAACAGGGTCCGAAAGGTGAGAAGGGCGATCCGGGCTTACCGGGAGAACCGGGACTGCAAGGTCGTCCGGGTGAGCTGGGCCCGCAGGGTCCAACGGGTCCGCCGGGCGCGAAGGGCCAAGAAGGTGCCCACGGCGCGCCGGGCGCTGCAGGTAACCCGGGTGCGCCGGGCCACGTGGGCGCGCCAGGTCCTAGCGGTCCGCCAGGCTCCGTGGGCGCTCCGGGCCTGCGTGGCACCCCGGGCAAGGACGGCGAACGTGGCGAAAAGGGTGCTGCCGGTGAGGAAGGTAGCCCGGGCCCGGTTGGTCCGCGCGGTGATCCGGGTGCGCCAGGCCTCCCGGGCCCGCCGGGCAAGGGC(SEQ ID NO:32)
[0242] 10) C22j amino acid sequence:
[0243] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgahgapgaagnpgapghv gapgpsgppgsvgapglRgtpgKDgERgEKgaagEE(SEQ ID NO:33)
[0244] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgahgapgaagnpgapghv
[0245] gapgpsgppgsvgapglRgtpgKDgERgEKgaagEE
[0246] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgahgapgaagnpgapghvgapgpsgppgsvgapglRgtpgKDgERgEKgaagEE (The amino acid sequence of the repeating unit of C22j is SEQ ID NO:33, with 3 repeats; the amino acid sequence of C22i is SEQ ID NO:34) Nucleotide sequence:
[0247] GGAGAACAAGGGCCGAAAGGTGAAAAGGGTGACCCGGGCCTGCCTGGCGAGCCGGGCCTGCAGGGTCGTCCGGGCGAGTTGGGTCCGCAGGGTCCGACCGGTCCGCCGGGCGCGAAAGGTCAAGAAGGTGCGCATGGTGCGCCAGGTGCCGCCGGCAACCCGGGTGCTCCGGGACATGTGGGCGCGCCAGGCCCTTCGGGTCCGCCGGGCAGCGTTGGTGCGCCGGGCTTACGTGGCACCCCGGGCAAAGACGGTGAGCGCGGTGAGAAGGGCGCTGCGGGTGAGGAAGGCGAGCAGGGTCCGAAAGGCGAAAAGGGCGATCCGGGCCTGCCGGGTGAGCCGGGACTGCAGGGCCGTCCGGGCGAGCTGGGCCCTCAAGGTCCGACCGGTCCGCCCGGTGCGAAAGGACAAGAAGGTGCCCACGGCGCGCCTGGCGCGGCAGGTAATCCGGGCGCCCCGGGTCACGTGGGTGCACCGGGTCCGTCTGGTCCACCGGGTTCCGTTGGTGCACCGGGCTTGCGCGGCACCCCGGGTAAAGATGGTGAACGTGGTGAGAAGGGCGCGGCAGGTGAAGAAGGTGAACAGGGCCCAAAAGGAGAGAAGGGTGATCCGGGTCTGCCGGGCGAGCCGGGTTTGCAAGGTCGCCCTGGGGAGCTGGGTCCTCAGGGTCCAACGGGCCCGCCGGGTGCCAAAGGCCAAGAGGGCGCGCATGGTGCTCCGGGTGCGGCAGGAAACCCGGGCGCTCCGGGCCACGTTGGTGCGCCGGGTCCGAGCGGTCCGCCGGGCAGCGTCGGTGCGCCGGGCCTCAGAGGTACTCCGGGCAAGGACGGCGAACGTGGAGAGAAGGGTGCTGCTGGCGAAGAA(SEQ ID NO:35)
[0248] 11)C22k amino acid sequence:
[0249] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgah(SEQ ID NO:36)
[0250] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgah
[0251] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgah
[0252] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgah
[0253] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgah
[0254] gEqgpKgEKgDpglpgEpglqgRpgElgpqgptgppgaKgqEgah(The amino acid sequence of the repeating unit of C22k is SEQ ID NO:36, the number of repeats is 6, and the amino acid sequence of C22i is SEQ ID NO:37)
[0255] Nucleotide sequence:
[0256] GGAGAACAAGGGCCCAAGGGCGAGAAGGGTGACCCGGGCCTGCCGGGAGAGCCGGGCCTGCAGGGCCGCCCTGGCGAACTGGGCCCACAGGGTCCGACCGGCCCACCGGGCGCTAAAGGCCAAGAGGGCGCGCACGGCGAGCAGGGTCCGAAGGGCGAGAAAGGCGACCCGGGTCTCCCTGGTGAGCCGGGTCTGCAAGGTCGTCCGGGTGAACTGGGTCCGCAGGGTCCGACGGGCCCACCGGGTGCCAAGGGCCAAGAAGGTGCGCACGGCGAGCAGGGTCCGAAGGGTGAAAAGGGTGATCCGGGCTTGCCAGGTGAGCCAGGTTTGCAGGGCCGTCCGGGAGAGCTGGGTCCGCAGGGCCCGACCGGTCCGCCGGGCGCGAAGGGTCAAGAAGGTGCGCATGGTGAACAGGGTCCGAAGGGCGAAAAAGGTGATCCGGGCCTGCCGGGCGAGCCGGGCCTGCAAGGCCGTCCGGGTGAACTGGGTCCGCAAGGCCCTACTGGTCCGCCGGGCGCTAAAGGCCAAGAAGGCGCGCACGGTGAGCAGGGTCCGAAAGGTGAAAAAGGCGACCCGGGCTTGCCGGGCGAACCGGGCTTGCAAGGTCGTCCGGGTGAGCTTGGTCCGCAGGGGCCGACCGGTCCGCCTGGCGCGAAAGGCCAAGAAGGTGCACATGGCGAGCAGGGTCCGAAGGGTGAAAAAGGTGATCCGGGTCTGCCGGGCGAGCCGGGGTTACAAGGTCGCCCAGGTGAACTGGGTCCGCAGGGTCCGACCGGTCCGCCGGGTGCAAAAGGTCAGGAGGGTGCCCAC(SEQ ID NO:38)
[0257] 12)C22l Amino acid sequence:
[0258] gvagppgpsgppgDKgspgsRglpgfpgpqgpagRDgapgnpgERgppgKpgls(SEQ ID NO:39)
[0259] gvagppgpsgppgDKgspgsRglpgfpgpqgpagRDgapgnpgERgppgKpgls
[0260] gvagppgpsgppgDKgspgsRglpgfpgpqgpagRDgapgnpgERgppgKpgls
[0261] gvagppgpsgppgDKgspgsRglpgfpgpqgpagRDgapgnpgERgppgKpgls
[0262] gvagppgpsgppgDKgspgsRglpgfpgpqgpagRDgapgnpgERgppgKpgls
[0263] gvagppgpsgppgDKgspgsRglpgfpgpqgpagRDgapgnpgERgppgKpgls(The amino acid sequence of the repeating unit of C22l is SEQ ID NO:39, with 6 repeats, and the amino acid sequence of C22i is SEQ ID NO:40)
[0264] Nucleotide sequence:
[0265] GGAGTAGCTGGGCCGCCGGGTCCGAGCGGTCCACCGGGGGATAAAGGTTCGCCGGGCAGCCGTGGTCTTCCGGGCTTTCCGGGCCCGCAGGGTCCGGCGGGTCGTGATGGTGCGCCGGGTAATCCGGGTGAACGCGGTCCGCCGGGCAAGCCGGGCCTGTCCGGTGTTGCAGGCCCGCCCGGCCCGAGCGGTCCGCCGGGTGACAAGGGTAGCCCGGGTTCACGTGGTTTACCGGGCTTCCCGGGTCCGCAGGGTCCGGCTGGCCGCGACGGCGCGCCAGGCAACCCAGGCGAGCGTGGCCCCCCGGGCAAGCCGGGCTTGTCCGGTGTTGCCGGTCCGCCGGGCCCGAGCGGCCCACCAGGCGACAAGGGCTCGCCGGGGTCTAGAGGTCTGCCGGGTTTCCCGGGTCCGCAAGGTCCGGCTGGTCGCGACGGTGCTCCGGGCAACCCGGGCGAGCGCGGTCCGCCTGGGAAACCGGGTTTGTCTGGCGTGGCGGGTCCGCCGGGTCCGAGCGGTCCGCCAGGCGACAAAGGCTCCCCGGGCAGCCGTGGTCTGCCTGGCTTTCCGGGCCCGCAAGGCCCGGCGGGTCGTGATGGTGCGCCTGGCAATCCGGGTGAGCGTGGTCCGCCGGGCAAACCGGGGCTGTCCGGCGTCGCAGGTCCGCCGGGTCCGAGCGGTCCACCGGGTGACAAAGGTAGCCCGGGTAGCCGTGGTCTGCCGGGCTTTCCTGGCCCGCAGGGTCCAGCGGGTCGTGATGGTGCCCCTGGAAACCCAGGCGAACGCGGTCCGCCTGGCAAGCCGGGCCTGAGCGGTGTGGCCGGTCCACCCGGCCCAAGCGGCCCGCCGGGTGATAAGGGTTCTCCAGGTTCCCGCGGTCTGCCTGGATTCCCGGGTCCGCAAGGTCCGGCAGGCCGTGATGGTGCGCCGGGAAACCCGGGCGAACGTGGTCCGCCGGGCAAACCGGGCCTCTCT(SEQ IDNO:41)
[0266] 13) C22m amino acid sequence:
[0267] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgiagpqgsqgERgaDgE vgqKgDqghpgvpgfmgppgnpgppgaDgiagaagpp (SEQ ID NO:42)
[0268] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgiagpqgsqgERgaDgE
[0269] vgqKgDqghpgvpgfmgppgnpgppgaDgiagaagpp
[0270] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgiagpqgsqgERgaDgEvgqKgDqghpgvpgfmgppgnpgppgaDgiagaagpp (The amino acid sequence of the repeating unit of C22m is SEQ ID NO:42, the number of repeats is 3, and the amino acid sequence of C22m is SEQ ID NO:43) Nucleotide sequence:
[0271] GGACCACCCGGGCTTCCGGGCTTGCCGGGCTTTAAAGGTGACAAAGGCGTGCCGGGTAAGCCGGGCCGTGAAGGCACCGAAGGTAAAAAGGGTGAAGCAGGTCCGCCGGGGCTGCCAGGTCCGCCTGGGATCGCTGGTCCGCAAGGCAGCCAGGGTGAGCGTGGTGCGGATGGCGAGGTGGGCCAAAAAGGTGATCAAGGCCACCCGGGCGTGCCGGGATTCATGGGTCCGCCGGGCAACCCGGGTCCGCCGGGTGCGGATGGCATCGCCGGTGCGGCTGGTCCGCCGGGCCCTCCGGGCCTGCCGGGCCTGCCGGGTTTCAAGGGTGATAAGGGTGTGCCGGGTAAACCGGGTCGTGAAGGAACCGAAGGTAAGAAGGGCGAGGCAGGCCCTCCGGGCCTGCCGGGTCCGCCAGGCATTGCGGGCCCACAGGGTAGCCAGGGTGAACGTGGTGCGGACGGCGAAGTTGGTCAGAAAGGCGATCAGGGTCACCCGGGTGTTCCGGGATTTATGGGTCCGCCAGGAAACCCGGGTCCGCCGGGTGCGGATGGTATTGCGGGTGCGGCAGGCCCACCGGGTCCGCCGGGTCTGCCGGGGCTGCCGGGTTTTAAGGGTGACAAAGGTGTTCCGGGCAAGCCGGGACGCGAGGGTACGGAGGGCAAGAAAGGCGAGGCAGGTCCGCCAGGCTTGCCAGGCCCTCCGGGGATCGCCGGTCCCCAGGGTTCCCAAGGTGAGCGCGGTGCGGACGGCGAGGTTGGTCAGAAAGGCGACCAAGGTCATCCGGGTGTCCCGGGCTTCATGGGCCCGCCGGGCAATCCGGGTCCGCCTGGCGCTGACGGTATTGCCGGCGCTGCCGGTCCTCCG(SEQ ID NO:44)
[0272] 14) C22n amino acid sequence:
[0273] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgiagpqgsqgERgaDgE vgqKgDq(SEQ ID NO:45)
[0274] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgiagpqgsqgERgaDgE
[0275] vgqKgDq
[0276] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgiagpqgsqgERgaDgE
[0277] vgqKgDq
[0278] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgiagpqgsqgERgaDgEvgqKgDq(The amino acid sequence of the repeating unit of C22n is SEQ ID NO:45, the number of repeats is 4, and the amino acid sequence of C22n is SEQ ID NO:46)
[0279] Nucleotide sequence:
[0280] GGGCCCCCAGGATTGCCGGGCCTGCCGGGTTTTAAAGGCGACAAAGGCGTCCCGGGCAAGCCGGGACGCGAGGGCACCGAAGGTAAAAAGGGCGAGGCGGGTCCGCCAGGCCTTCCGGGCCCGCCAGGCATTGCGGGTCCGCAGGGTAGCCAGGGTGAGCGCGGTGCAGACGGCGAGGTTGGTCAAAAAGGTGACCAAGGTCCGCCGGGGCTGCCTGGCTTGCCTGGCTTCAAGGGCGACAAAGGGGTGCCGGGTAAGCCGGGTCGTGAGGGCACCGAAGGTAAGAAGGGCGAGGCCGGACCACCGGGACTCCCGGGTCCACCGGGCATCGCGGGTCCGCAGGGTTCCCAAGGTGAACGTGGTGCGGATGGTGAGGTTGGCCAAAAAGGCGATCAGGGCCCACCGGGTTTACCGGGTCTGCCGGGCTTTAAAGGTGATAAAGGCGTGCCGGGGAAGCCGGGTCGTGAAGGCACGGAAGGTAAGAAAGGTGAGGCTGGTCCGCCAGGCCTGCCGGGTCCGCCGGGTATTGCGGGTCCGCAAGGCTCTCAGGGTGAACGTGGTGCGGATGGTGAGGTGGGCCAGAAGGGTGATCAGGGCCCTCCGGGCCTGCCGGGTCTGCCGGGATTCAAAGGTGACAAGGGCGTTCCGGGTAAACCGGGCCGTGAAGGCACCGAGGGCAAGAAGGGCGAAGCAGGTCCGCCGGGTCTGCCGGGTCCGCCGGGCATCGCCGGTCCCCAAGGCAGCCAAGGTGAGCGCGGTGCTGATGGTGAAGTAGGTCAGAAAGGTGACCAG(SEQ ID NO:47)
[0281] 15) C22o Amino acid sequence:
[0282] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgpp(SEQ ID NO:48)
[0283] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgppglpglpgfKgDKgvp
[0284] gKpgREgtEgKKgEagppglpgpp
[0285] gppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgppgppglpglpgfKgDKgvpgKpgREgtEgKKgEagppglpgpp(The amino acid sequence of the repeating unit of C22o is SEQ ID NO:48, the number of repeats is 3, and the amino acid sequence of C22o is SEQ ID NO:49)
[0286] Nucleotide sequence:
[0287] GGACCCCCAGGGTTGCCGGGCCTGCCGGGCTTTAAAGGCGACAAAGGTGTCCCGGGTAAGCCGGGCCGTGAGGGCACCGAAGGTAAGAAAGGTGAAGCTGGCCCTCCGGGTCTGCCGGGTCCGCCGGGTCCGCCGGGCCTGCCGGGCCTGCCGGGCTTTAAAGGTGACAAGGGCGTGCCGGGCAAGCCGGGCCGTGAGGGCACCGAAGGTAAGAAAGGTGAAGCGGGTCCACCGGGCTTACCGGGCCCACCGGGCCCACCGGGTCTGCCGGGTCTGCCTGGCTTCAAAGGTGACAAAGGTGTGCCGGGCAAGCCTGGTCGTGAAGGTACTGAGGGTAAAAAGGGCGAAGCGGGTCCACCGGGTCTTCCGGGTCCGCCTGGCCCACCGGGTTTGCCGGGCCTGCCGGGTTTCAAGGGCGATAAAGGTGTTCCGGGCAAGCCGGGACGCGAGGGCACGGAGGGTAAGAAAGGCGAGGCAGGCCCACCGGGTCTGCCGGGCCCACCGGGTCCTCCGGGGCTGCCGGGCCTGCCGGGTTTTAAGGGCGATAAAGGTGTTCCGGGTAAACCGGGTCGTGAAGGCACCGAAGGCAAGAAGGGCGAAGCGGGTCCGCCTGGCTTGCCGGGTCCGCCAGGTCCGCCGGGTCTGCCGGGGTTGCCGGGATTCAAAGGCGATAAAGGGGTGCCGGGTAAACCGGGTCGCGAGGGAACCGAGGGTAAGAAGGGCGAGGCCGGTCCGCCGGGTCTCCCGGGCCCACCG(SEQ ID NO:50)
[0288] 16)C22p amino acid sequence:
[0289] gKpgppgEpgKagEpglpgpEgaRgppgfKghtgDsgapgpRgEsgamglpgqEglpgK DgDtgptgpqgpqgpRgppgKngspgspgEpgpsgtpgqKgsKgEngspglpgflgpRgppgEpg EKgvpgKE(SEQ ID NO:51)
[0290] gKpgppgEpgKagEpglpgpEgaRgppgfKghtgDsgapgpRgEsgamglpgqEglpgK DgDtgptgpqgpqgpRgppgKngspgspgEpgpsgtpgqKgsKgEngspglpgflgpRgppgEpg EKgvpgKE (The amino acid sequence of the repeating unit of C22p is SEQ ID NO:51, the number of repeats is 2, and the amino acid sequence of C22p is SEQ ID NO:52)
[0291] Nucleotide sequence:
[0292] GGAAAACCCGGGCCACCGGGCGAGCCGGGCAAGGCAGGTGAGCCAGGCCTGCCGGGTCCGGAAGGTGCGCGTGGTCCGCCGGGATTCAAGGGCCACACCGGCGACAGCGGTGCGCCGGGGCCGCGCGGTGAGAGCGGTGCTATGGGTTTGCCTGGTCAAGAAGGCCTGCCGGGCAAGGATGGTGACACCGGTCCAACGGGCCCGCAGGGTCCGCAAGGTCCGCGTGGTCCGCCGGGCAAAAACGGCTCTCCGGGCTCCCCGGGCGAGCCGGGTCCGAGCGGTACTCCGGGCCAGAAAGGCTCCAAAGGCGAGAATGGCTCTCCGGGCTTGCCGGGCTTCCTGGGCCCGCGTGGTCCGCCGGGCGAACCGGGCGAGAAAGGTGTTCCGGGCAAAGAAGGTAAGCCGGGCCCACCGGGTGAGCCGGGCAAGGCCGGTGAACCGGGCCTTCCGGGCCCGGAAGGTGCGCGTGGTCCGCCCGGTTTTAAAGGCCATACGGGTGATAGCGGTGCGCCTGGTCCTCGCGGTGAAAGCGGTGCAATGGGTCTGCCGGGCCAAGAAGGTTTACCGGGCAAGGACGGCGATACCGGTCCGACCGGCCCGCAGGGTCCGCAGGGTCCGCGTGGTCCGCCAGGCAAGAACGGTTCCCCGGGTTCGCCGGGCGAGCCGGGACCTAGCGGCACCCCGGGCCAGAAAGGTTCGAAAGGTGAGAACGGTAGCCCGGGGCTGCCGGGATTTCTGGGTCCGCGCGGTCCACCGGGCGAGCCGGGCGAAAAAGGCGTGCCGGGTAAGGAA(SEQ ID NO:53)
[0293] 17)C22q amino acid sequence:
[0294] gKpgppgEpgKagEpglpgpEgaRgppgfKghtgDsgapgpRgEsgamglpgqEglpgK DgDt(SEQID NO:54)
[0295] gKpgppgEpgKagEpglpgpEgaRgppgfKghtgDsgapgpRgEsgamglpgqEglpgK
[0296] DgDt
[0297] gKpgppgEpgKagEpglpgpEgaRgppgfKghtgDsgapgpRgEsgamglpgqEglpgKDgDt
[0298] gKpgppgEpgKagEpglpgpEgaRgppgfKghtgDsgapgpRgEsgamglpgqEglpgKDgDt(The amino acid sequence of the repeating unit of C22q is SEQ ID NO:54, the number of repeats is 4, and the amino acid sequence of C22q is SEQ ID NO:55)
[0299] Nucleotide sequence:
[0300] GGGAAACCCGGACCGCCAGGAGAGCCGGGCAAGGCAGGCGAGCCGGGTCTGCCGGGGCCGGAAGGTGCGCGTGGTCCGCCGGGTTTCAAAGGTCATACCGGCGACAGCGGTGCGCCAGGTCCGCGTGGCGAGTCCGGCGCGATGGGCTTACCGGGCCAAGAAGGTCTGCCGGGGAAGGACGGCGACACCGGTAAGCCGGGCCCACCGGGTGAACCGGGCAAAGCCGGTGAACCGGGCCTGCCGGGCCCCGAGGGAGCGCGCGGTCCTCCGGGATTTAAAGGCCACACCGGTGATAGCGGTGCTCCGGGCCCTCGTGGTGAATCGGGTGCGATGGGTCTGCCGGGCCAGGAGGGCTTGCCTGGCAAGGACGGCGACACGGGTAAGCCGGGCCCGCCGGGCGAGCCGGGTAAGGCGGGTGAGCCGGGTCTCCCGGGTCCGGAAGGTGCGAGAGGTCCTCCGGGCTTTAAAGGTCATACGGGCGATAGCGGTGCGCCAGGTCCGCGTGGTGAGAGCGGCGCTATGGGTCTGCCGGGCCAGGAGGGCTTGCCGGGTAAAGATGGTGACACCGGTAAACCGGGTCCGCCGGGTGAGCCGGGCAAGGCGGGTGAACCTGGTTTGCCGGGCCCAGAAGGCGCTCGCGGTCCACCGGGTTTCAAAGGTCACACTGGTGATTCTGGCGCACCGGGCCCGCGTGGTGAATCCGGCGCAATGGGTCTGCCGGGTCAAGAAGGCCTGCCGGGCAAAGATGGTGATACC(SEQ ID NO:56)
[0301] 18)C22r Amino acid sequence: gpqgpqgpRgppgKngspgspgEpgpsgtpgqKgsKgEngspglpgflgpRgppgEpgEKgvpgKE(SEQ ID NO:57)
[0302] gpqgpqgpRgppgKngspgspgEpgpsgtpgqKgsKgEngspglpgflgpRgppgEpgEKgvpgKE
[0303] gpqgpqgpRgppgKngspgspgEpgpsgtpgqKgsKgEngspglpgflgpRgppgEpgEKgvpgKE
[0304] gpqgpqgpRgppgKngspgspgEpgpsgtpgqKgsKgEngspglpgflgpRgppgEpgEKgvpgKE(The amino acid sequence of the repeating unit of C22r is SEQ ID NO:57, the number of repeats is 4, and the amino acid sequence of C22r is SEQ ID NO:58)
[0305] Nucleotide sequence:
[0306] GGACCCCAAGGGCCGCAGGGTCCGCGCGGCCCACCGGGTAAGAATGGTAGCCCGGGCTCACCGGGTGAGCCGGGCCCGAGCGGCACCCCGGGTCAGAAGGGCTCCAAGGGTGAGAACGGCAGCCCGGGTCTGCCTGGTTTTCTGGGTCCGCGTGGCCCGCCGGGCGAACCGGGTGAAAAAGGCGTTCCGGGTAAAGAAGGCCCGCAAGGTCCGCAAGGTCCGCGTGGTCCGCCTGGCAAGAACGGCAGCCCGGGGTCTCCGGGCGAACCGGGCCCGAGCGGCACCCCGGGTCAGAAAGGTTCCAAAGGTGAAAATGGTAGCCCGGGCTTGCCGGGCTTCCTGGGTCCGCGTGGCCCGCCAGGCGAGCCTGGCGAGAAGGGTGTGCCGGGCAAAGAAGGTCCACAAGGTCCGCAGGGTCCGCGTGGTCCACCGGGTAAGAATGGCAGCCCGGGATCTCCGGGCGAGCCGGGCCCAAGCGGCACCCCGGGTCAGAAAGGCTCCAAAGGTGAAAACGGTAGCCCGGGTTTGCCGGGCTTTCTGGGTCCGAGAGGGCCGCCGGGCGAGCCGGGCGAAAAAGGGGTGCCGGGCAAGGAGGGTCCGCAAGGTCCGCAAGGTCCGCGTGGTCCTCCGGGCAAGAACGGCTCCCCGGGTTCTCCGGGCGAGCCGGGCCCAAGCGGTACGCCGGGTCAGAAAGGTTCGAAGGGTGAGAACGGCTCGCCGGGTTTACCGGGATTCCTGGGCCCCCGCGGTCCACCGGGTGAGCCGGGTGAAAAGGGTGTTCCGGGCAAAGAA(SEQ ID NO:59)
[0307] 19) Amino acid sequence of C22s:
[0308] gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKghtglmgpqglpgEngpvgppgppgqpgfpglRgEs(SEQ ID NO:1)
[0309] gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKghtglmgpqglpgEng
[0310] pvgppgppgqpgfpglRgEs
[0311] gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKghtglmgpqglpgEngpvgppgppgqpgfpglRgEs (The amino acid sequence of the repeating unit of C22s is SEQ ID NO: 1, and the number of repetitions is 3; the amino acid sequence of C22s is SEQ ID NO: 2)
[0312] Nucleotide sequence:
[0313] AGGAGTACCCGGAAAGCCGGGTGAGCCGGGTTTTAAGGGTGAACGTGGTGATCCG
[0314] GGGATCAAAGGTGACAAGGGCCCACCGGGTGGTAAGGGCCAACCGGGCGACCCA
[0315] GGCATTCCGGGCCACAAAGGCCACACGGGTCTGATGGGTCCTCAGGGTTTGCCGG
[0316] GCGAGAACGGCCCGGTTGGTCCGCCTGGCCCACCGGGCCAGCCGGGTTTCCCGGG
[0317] GCTGCGTGGTGAAAGCGGTGTGCCGGGCAAACCGGGTGAACCGGGCTTTAAAGGT
[0318] GAGCGCGGTGACCCGGGTATTAAAGGTGATAAAGGCCCTCCGGGCGGTAAGGGTC
[0319] AACCGGGCGATCCGGGTATCCCGGGCCACAAGGGTCATACCGGTCTGATGGGCCC
[0320] GCAGGGTCTGCCGGGTGAAAACGGTCCCGTGGGCCCACCGGGCCCACCGGGACAG
[0321] CCGGGCTTTCCGGGTTTGCGTGGTGAGTCCGGCGTTCCGGGTAAGCCGGGTGAACC
[0322] GGGCTTCAAGGGCGAGCGCGGTGACCCGGGCATCAAAGGCGACAAAGGTCCGCC
[0323] AGGAGGCAAGGGCCAGCCTGGCGATCCGGGCATTCCGGGTCACAAAGGCCATACC
[0324] GGTCTGATGGGTCCGCAAGGTCTTCCGGGTGAAAATGGTCCGGTCGGTCCACCGGGCCCACCGGGACAACCGGGATTCCCGGGTTTACGTGGTGAGAGC(SEQ ID NO:3)20)C22t Amino acid sequence:
[0325] gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKght(SEQ ID NO:4)
[0326] gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKght
[0327] gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKght
[0328] gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKght
[0329] gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKght
[0330] gvpgKpgEpgfKgERgDpgiKgDKgppggKgqpgDpgipghKght(The repeating unit amino acid sequence of C22t is SEQ ID NO:4, the repeating number is 6; The amino acid sequence of C22t is SEQ ID NO:5)
[0331] Nucleotide sequence:
[0332] GGAGTACCCGGAAAGCCGGGCGAGCCGGGTTTCAAAGGTGAGCGTGGTGACCCGGGCATTAAGGGCGACAAAGGTCCACCCGGTGGTAAAGGTCAGCCTGGTGATCCGGGCATCCCGGGCCACAAAGGCCACACCGGTGTTCCGGGCAAGCCGGGCGAGCCGGGCTTCAAAGGCGAGCGTGGTGATCCGGGGATTAAAGGTGACAAAGGTCCGCCAGGTGGTAAGGGCCAGCCGGGAGACCCGGGCATTCCGGGCCATAAAGGTCACACCGGCGTGCCGGGCAAACCGGGTGAGCCGGGTTTTAAAGGTGAACGTGGTGACCCAGGCATCAAAGGCGACAAGGGCCCACCGGGCGGTAAAGGGCAACCGGGTGACCCGGGCATTCCGGGCCATAAAGGCCATACGGGTGTTCCGGGCAAGCCGGGAGAACCGGGTTTTAAAGGTGAACGCGGTGACCCGGGCATCAAAGGCGACAAAGGCCCTCCGGGTGGTAAAGGCCAACCGGGCGATCCAGGCATTCCGGGCCACAAGGGTCATACCGGTGTCCCGGGTAAGCCGGGTGAACCGGGATTCAAGGGCGAGCGTGGCGACCCGGGCATCAAGGGTGATAAGGGTCCTCCGGGTGGCAAGGGTCAGCCGGGTGATCCGGGTATCCCGGGTCACAAGGGTCACACTGGTGTGCCAGGCAAGCCGGGTGAACCGGGTTTTAAAGGCGAACGCGGTGATCCGGGTATCAAGGGCGATAAGGGTCCGCCTGGTGGCAAGGGTCAACCGGGTGATCCGGGTATCCCGGGCCATAAGGGTCACACC(SEQ ID NO:6)
[0333] 21)C22u Amino acid sequence:
[0334] gRpgppgppgKDglpgRagpmgEpgRpgqgglEgpsgpigpKgERgaKgDpgap(SEQ ID NO:60)
[0335] gRpgppgppgKDglpgRagpmgEpgRpgqgglEgpsgpigpKgERgaKgDpgap
[0336] gRpgppgppgKDglpgRagpmgEpgRpgqgglEgpsgpigpKgERgaKgDpgap
[0337] gRpgppgppgKDglpgRagpmgEpgRpgqgglEgpsgpigpKgERgaKgDpgap
[0338] gRpgppgppgKDglpgRagpmgEpgRpgqgglEgpsgpigpKgERgaKgDpgap
[0339] gRpgppgppgKDglpgRagpmgEpgRpgqgglEgpsgpigpKgERgaKgDpgap(The repeating unit amino acid sequence of C22u is SEQ ID NO:60, with 6 repeats; the amino acid sequence of C22u is SEQ ID NO:15)
[0340] Nucleotide sequence:
[0341] GGAAGGCCCGGGCCACCGGGCCCACCGGGTAAGGACGGCCTCCCGGGGAGAGCGGGTCCGATGGGTGAGCCGGGCCGTCCGGGCCAAGGTGGTCTGGAAGGTCCAAGCGGTCCGATTGGCCCGAAAGGTGAGCGCGGTGCGAAAGGTGACCCGGGAGCGCCAGGCCGTCCGGGCCCGCCGGGACCGCCAGGTAAAGATGGTCTGCCGGGTCGTGCAGGCCCGATGGGTGAGCCGGGTCGTCCGGGTCAAGGTGGTTTGGAAGGCCCGTCGGGCCCCATCGGTCCGAAGGGCGAACGTGGTGCGAAGGGCGATCCGGGTGCCCCTGGACGCCCGGGCCCACCGGGACCGCCTGGCAAGGACGGCCTGCCGGGTCGCGCAGGTCCGATGGGTGAACCGGGACGTCCGGGTCAGGGCGGTTTGGAGGGTCCGAGCGGTCCGATCGGCCCGAAGGGCGAACGTGGCGCCAAAGGCGATCCGGGGGCTCCGGGCCGTCCGGGTCCACCGGGCCCGCCGGGCAAGGACGGCCTGCCTGGCCGCGCGGGTCCGATGGGTGAGCCGGGCCGTCCGGGTCAGGGTGGTTTGGAAGGTCCGTCCGGTCCGATTGGTCCGAAAGGTGAGCGCGGTGCGAAAGGTGATCCGGGCGCTCCGGGCCGTCCGGGTCCACCGGGTCCTCCGGGCAAGGATGGTCTGCCAGGTAGAGCGGGTCCGATGGGTGAACCGGGCCGTCCGGGACAAGGTGGTCTTGAGGGTCCGTCTGGCCCGATTGGTCCTAAAGGCGAACGCGGTGCGAAAGGCGACCCGGGTGCTCCGGGTCGTCCGGGCCCGCCGGGGCCTCCGGGCAAAGACGGTCTGCCGGGCCGTGCCGGCCCGATGGGCGAACCAGGACGCCCAGGCCAGGGTGGTCTGGAGGGTCCGAGCGGTCCGATCGGCCCGAAGGGCGAGCGTGGTGCGAAGGGCGATCCGGGCGCACCG(SEQ IDNO:9)
[0342] Each of the above-mentioned coding nucleotide sequences was commercially synthesized. Each of the above-mentioned coding nucleotide sequences (with a collagen toolase cleavage site added at the 5' end, the amino acid sequence of the collagen toolase cleavage site being ENLYFQ, and the nucleotide sequence being GAAAACCTGTATTTCCAG) was inserted between the KpnI and XhoI cleavage sites of the pET-28a-Trx-His expression vector to obtain the recombinant expression plasmid.
[0343] 3. Transform the successfully constructed expression plasmid into E. coli competent cells BL21(DE3). The specific process is as follows: (1) Take out the E. coli competent cells BL21(DE3) from the ultra-low temperature freezer and place them on ice. When they are half-thawed, take 2 μl of the plasmid to be transformed and add it to the E. coli competent cells BL21(DE3), and mix slightly 2-3 times. (2) Place the mixture on ice for 30 min, then heat shock it in a 42℃ water bath for 45-90 s, and then place it on ice for 2 min. (3) Transfer it to a biosafety cabinet and add 700 μl of liquid LB medium, and then incubate it at 37℃ and 220 rpm for 60 min. (4) Take 200 μl of bacterial solution and spread it evenly on an LB plate containing ampicillin sodium. (5) Incubate the plate in a 37℃ incubator for 15-17 h until uniformly sized colonies grow.
[0344] 4. Pick 5-6 single colonies from the transformed LB agar plates and place them in a shake flask containing LB medium with antibiotic stock solution. Incubate at 220 rpm and 37°C for 7 hours in a constant temperature shaker. Then, cool the shake flask to 16°C, add IPTG to induce expression for a period of time, aliquot the bacterial solution into centrifuge bottles, centrifuge at 8000 rpm and 4°C for 10 minutes, collect the bacterial cells, record the cell weight, and take a sample (labeled: bacterial solution) for electrophoresis detection.
[0345] 5. Resuspend the collected bacterial cells in equilibration working solution (200mM sodium chloride, 25mM Tris, 20mM imidazole, pH 8.0), cool the bacterial solution to ≤15℃, and homogenize it twice by high-pressure homogenization (label the samples after each homogenization as: Homogenization I and Homogenization II). Collect the bacterial solution after completion. Aliquot the homogenized bacterial solution into centrifuge bottles and centrifuge at 17000 rpm and 4℃ for 30 min. Collect the supernatant, and perform electrophoresis analysis on the supernatant (labeled as: supernatant) and precipitate.
[0346] 6. The recombinant XXII type humanized collagen was purified and enzymatically digested. The specific process was as follows: (1) Crude purification: a. Wash the column material (Ni6FF, Cytiva) with water, 5 CV. b. Equilibrate the column material with equilibration buffer (200mM sodium chloride, 25mM Tris, 20mM imidazole, pH 8.0), 5 CV. c. Sample loading: Add the supernatant after centrifugation to the column material until the liquid is completely flowed out, and take the flow-through for electrophoresis inspection (labeled as flow-through). d. Washing off contaminants: Add 25mL of washing buffer (200mM sodium chloride, 25mM Tris, 20mM imidazole) until the liquid is completely flowed out, and take the washed flow-through for electrophoresis inspection (labeled as washed). e. Collect the target protein: Add 20 mL of elution buffer (200 mM sodium chloride, 25 mM Tris, 250 mM imidazole, pH 8.0), and collect the flow-through (labeled: elution). Detect the protein concentration, calculate the protein amount, and perform electrophoresis. f. Wash the column with 1 M imidazole working solution (labeled: 1 M wash). g. Wash the column with purified water. (2) Enzyme digestion: Add TEV enzyme at a ratio of total protein to total TEV enzyme of 50:1 (if not digested, consider increasing the enzyme concentration, for example, a total ratio of 20:1 or 5:1), digest at 16℃ for 4 h, and take samples for electrophoresis (labeled: digested). Place the digested protein solution into a dialysis bag, dialyze at 4℃ for 2 h, and then transfer to new dialysis buffer for overnight dialysis at 4℃ (labeled: buffer change).
[0347] (3) Purification (protein isoelectric point > 8.0): a. Equilibrate the column (Capto Q, Cytiva): Equilibrate the column with solution A (20 mM Tris, 20 mM sodium chloride, pH 8.0) at a flow rate of 10 ml / min. b. Load the sample: Load the sample at a flow rate of 5 ml / min and collect the flow-through (labeled QFL), then perform electrophoresis. c. Gradient elution: Set up 0-15% solution B (20 mM Tris, 1 M sodium chloride, pH 8.0) for 2 min followed by 3 CVs, 15-30% solution B for 2 min followed by 3 CVs, 30-50% solution B for 2 min followed by 3 CVs, and 50-100% solution B for 2 min followed by 3 CVs. Collect the eluted peaks and perform electrophoresis (labeled B wash). d. Wash the column. Store the protein at 4°C.
[0348] (4) Nickel-on-coated (Ni6FF, Cytiva) (protein isoelectric point < 8.0): a. Column equilibration: Equilibrate the column using solution A (20 mM Tris, 20 mM Sodium Chloride, 20 mM Imidazole, pH 8.0) for 5 CVs. b. Sample loading: Add the protein (after enzyme switching) to the column until the liquid has completely flowed out. Take a flow-through for electrophoresis detection (labeled as nickel-on-coated). c. Wash the column with 1 M imidazole working solution (20 mM Tris, 20 mM Sodium Chloride, 1 M Imidazole, pH 8.0) (labeled as 1 M wash). d. Wash the column with purified water. Store the protein at 4°C.
[0349] 7. Concentration detection
[0350] Accurately measure an appropriate amount of sample, dilute it 10-50 times with elution buffer, and stir thoroughly with a glass rod. Measure the absorbance at 280 nm using a UV-Vis spectrophotometer. Calculate the protein concentration using the formula C(mg / ml) = A280 × absorbance coefficient × dilution factor (Note: the absorbance coefficient can be obtained from the amino acid sequence, and the absorbance value should be between 0.1 and 1).
[0351] The concentration test results are as follows:
[0352] plasmid Absorption coefficient A280 Dilution factor concentration Elution protein volume Protein content C22a 2.01 0.435 5 4.37 mg / ml 20ml 87.44mg C22b 2.91 0.283 5 4.12mg / ml 20ml 82.35mg C22c 2.36 0.245 5 2.89 mg / ml 20ml 57.82mg C22d 2.99 0.231 5 3.45mg / ml 20ml 69.07mg C22e 2.99 0.121 1 0.36mg / ml 20ml 7.24mg C22f 2.74 0.178 5 2.44 mg / ml 20ml 48.77mg C22g 3.13 0.515 2 3.22mg / ml 20ml 64.48mg C22h 2.60 0.755 2 3.93mg / ml 20ml 78.52mg C22i 2.51 0.550 2 2.76 mg / ml 20ml 55.22mg C22j 2.81 0.750 2 3.23 mg / ml 20ml 64.63mg C22k 2.77 0.648 5 8.97mg / ml 20ml 179.50mg C22l 2.99 0.070 1 0.21mg / ml 20ml 4.19mg C22m 2.82 0.720 1 2.03 mg / ml 20ml 40.61mg C22n 2.72 0.344 5 4.68 mg / ml 20ml 93.57mg C22o 2.64 0.914 1 2.41 mg / ml 20ml 48.26mg C22p 2.70 0.250 1 0.68mg / ml 20ml 13.50mg C22q 2.62 0.340 1 0.89mg / ml 20ml 17.82mg C22r 2.70 0.390 1 1.05mg / ml 20ml 21.06mg C22s 2.55 0.401 5 5.11 mg / ml 20ml 102.26mg C22t 2.79 0.395 5 5.51 mg / ml 20ml 110.21mg C22u 3.05 0.62 2 3.78mg / ml 20ml 75.64mg
[0353] Protein expression levels: C22k > C22t > C22s > C22n > C22a > C22b > C22h > C22u > C22d > C22j > C22g > C22c > C22i > C22f > C22o > C22m > C22r > C22q > C22p > C22e > C22l.
[0354] 8. Electrophoresis detection
[0355] The specific process is as follows: Take 40 μl of sample solution, add 10 μl of 5× protein loading buffer (250 mM Tris-HCl (pH: 6.8), 10% SDS, 0.5% bromophenol blue, 50% glycerol, 5% β-mercaptoethanol), boil in 100℃ water for 10 min, then add 10 μl to each well of SDS-PAGE protein gel, run at 80V for 2 h, stain the protein with Coomassie Brilliant Blue staining solution (0.1% Coomassie Brilliant Blue R-250, 25% isopropanol, 10% glacial acetic acid) for 20 min, and then destain with protein destaining solution (10% acetic acid, 5% ethanol).
[0356] Figure 1 The electrophoresis results for C22a are shown. Figure 2 The electrophoresis results for C22b are shown. Figure 3 The results of C22c electrophoresis are shown. Figure 4 The results of C22d electrophoresis are shown. Figure 5 The results of C22e electrophoresis are shown. Figure 6 The electrophoresis results of C22f are shown. Figure 7 The results of C22g electrophoresis are shown. Figure 8 The results of C22h electrophoresis are shown. Figure 9 The results of C22i electrophoresis are shown. Figure 10 The electrophoresis results of C22j are displayed. Figure 11 The results of C22k electrophoresis are shown. Figure 12 The results of C22l electrophoresis are shown. Figure 13 The results of C22m electrophoresis are shown. Figure 14 The results of C22n electrophoresis are shown. Figure 15 The results of C22o electrophoresis are shown. Figure 16 The results of C22p electrophoresis are shown. Figure 17 The results of C22q electrophoresis are shown. Figure 18 The results of C22r electrophoresis are shown. Figure 19 The results of C22s electrophoresis are shown. Figure 20 The results of C22t electrophoresis are shown. Figure 21 The results of C22u electrophoresis are shown.
[0357] Electrophoresis results showed:
[0358] C22e (see) Figure 5 ), C22p (see) Figure 16 C22q (see Figure 17), C22r (see Figure 17) Figure 18 The crude purity yield is relatively low;
[0359] C22u (see) Figure 21 The enzyme digestion ratio needs to be reduced.
[0360] C22f (see) Figure 6 ), C22k (see) Figure 11 ), C22m (see Figure 13 ), C22n (see Figure 14 ), C22o (see) Figure 15 The enzyme content needs to be increased (total protein to TEV enzyme ratio); C22f (see...) Figure 6 ), C22k (see) Figure 11 ), C22m (see Figure 13 ), C22n (see Figure 14 ), C22o (see) Figure 15 The second batch of screening increased the ratio of total protein to TEV enzyme (from 20:1 to 5:1 or 10:1), but the enzyme digestion was still incomplete and there was enzyme residue.
[0361] C22a (see) Figure 1C22b (see) Figure 2 C22g (see) Figure 7 ), C22i (see) Figure 9 ), C22j (see Figure 10 The separation effect of the target protein was poor after reverse nickel plating;
[0362] C22h (see) Figure 8 The purity of the target protein is better after reverse nickel plating;
[0363] C22c (see) Figure 3 C22d (see) Figure 4 ), C22l (see) Figure 12 After purification, the target protein has good purity but low yield.
[0364] C22s (see Figure 19), C22t (see Figure 19) Figure 20 The yield was high, and the purity of the target protein after purification was good. Therefore, further analysis was performed on the purified proteins of C22s and C22t.
[0365] Example 2: Mass spectrometry detection of recombinant type XXII humanized collagen
[0366] Experimental methods
[0367]
[0368] Protein samples were reduced by DTT and alkylated with iodoacetamide, followed by trypsin digestion overnight. The resulting peptides were then desalted using a C18 ZipTip and mixed with the matrix α-cyano-4-hydroxycinnamic acid (CHCA) for TLC. Finally, analysis was performed using a matrix-assisted laser desorption / ionization time-of-flight mass spectrometer (MALDI-TOF / TOF Ulraflextreme™, Brucker, Germany) (peptide fingerprinting techniques can be found in Protein J. 2016; 35:212-7).
[0369] Data retrieval was performed via the MS / MS Ion Search page on the local Masco website. Protein identification results were obtained from primary mass spectrometry of the peptides produced after enzymatic digestion. Detection parameters: Trypsin digestion, with two missed cleavage sites. Cysteine alkylation was set as a fixed modification. Methionine oxidation was set as a variable modification. The NCBprot database was used for identification.
[0370] Table 1: Molecular weights and corresponding peptides of recombinant XXII type humanized collagen detected by C22s mass spectrometry
[0371]
[0372]
[0373] The detected polypeptide fragments showed 100% coverage compared to the theoretical sequence, making the detection results highly reliable.
[0374] Table 2: Molecular weights and corresponding peptides of recombinant XXII type humanized collagen detected by C22t mass spectrometry
[0375]
[0376] The detected polypeptide fragments showed a coverage rate of 93.3% compared to the theoretical sequence, making the detection results highly reliable.
[0377] Example 3: Bioactivity assay of recombinant XXII type humanized collagen
[0378] The method for detecting collagen activity can be found in the literature Juming Yao, Satoshi Yanagisawa, Tetsuo Asakura, Design, Expression and Characterization of Collagen-Like Proteins Based on the Cell Adhesive and Crosslinking Sequences Derived from Native Collagens, J Biochem. 136, 643-649 (2004). The specific implementation method is as follows:
[0379] (1) The concentration of the protein samples to be tested was detected by ultraviolet absorption method, including bovine type I collagen (China National Institutes for Food and Drug Control, No.: 380002), and recombinant type XXII humanized collagen C22s and C22t provided by the present invention.
[0380] Specifically, the ultraviolet light absorption of the sample at 215 nm and 225 nm was measured separately, and the protein concentration was calculated using the empirical formula C(μg / mL) = 144 × (A215 - A225). Note that the detection must be performed when A215 < 1.5. The principle of this method is to measure the characteristic absorption of peptide bonds under far-ultraviolet light, which is not affected by the content of chromophores, has few interfering substances, is simple to operate, and is suitable for detecting human collagen and its analogues that are not colorimetric by Coomassie Brilliant Blue. (Reference: Walker JM. The Protein Protocols Handbook, second edition. Humana Press. 43-45.). After measuring the protein concentration, the concentration of all the proteins to be tested was adjusted to 1 mg / mL with PBS.
[0381] (2) Add different concentrations of collagen, positive control and negative control to the microplate, 100 μL per well, 5 replicates per group, and incubate overnight at 4°C.
[0382] (3) Discard the supernatant, add 100 μL of 1% BSA (heat inactivated at 56℃ for 30 min), and incubate at 37℃ for 60 min. Discard the supernatant and wash three times with D-PBS solution.
[0383] (4) Add 105 well-cultured 3T3 / NIH cells resuspended in D-PBS to each well and incubate at 37°C for 120 min. Wash each well 3 times with D-PBS solution.
[0384] (5) OD was detected using the CCK8 assay kit (manufacturer: Beyotime, product catalog number C0038). 450 The absorbance is measured in nm. Based on the values of the blank control, the cell adhesion rate can be calculated. The calculation formula is as follows: Cell adhesion rate reflects collagen activity. Higher protein activity allows for a better external environment to be provided to cells in a shorter time, thus aiding cell adhesion.
[0385] The results are as follows Figure 22 As shown, compared with the D-PBS group, the positive control group had a significant effect on promoting cell adhesion (***, P<0.001), and recombinant humanized collagen also promoted cell adhesion at the experimental concentration.
[0386] Although the invention has been described with reference to illustrative embodiments, those skilled in the art will understand that various other changes, omissions, and / or additions can be made without departing from the spirit and scope of the invention, and that elements of the described embodiments can be substituted with substantially equivalents. Furthermore, many modifications can be made without departing from the scope of the invention to adapt particular situations or materials to the teachings of the invention. Therefore, this invention is not intended to be limited to the specific embodiments disclosed for carrying out the invention, but rather is intended to encompass all embodiments falling within the scope of the appended claims.
[0387] This invention includes:
[0388] 1. A polypeptide comprising one or more repeating units connected directly or via a linker, the repeating units comprising an amino acid sequence selected from the group consisting of: SEQ ID NO: 1, 4, 7, 12, 18, 21, 24, 27, 30, 33, 36, 39, 42, 45, 48, 51, 54, 57, 60, the variant being (1) an amino acid sequence in which one or more amino acid residues are mutated or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the amino acid sequence;
[0389] Preferably, the plurality of repeating units is 2-50 repeating units, for example 2-45, 2-40, 2-35, 2-30, 2-25, 2-20, 2-15, 2-10, 2-8 or 2-6 repeating units;
[0390] Preferably, the linker comprises one or more amino acid residues, such as 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, or 1-2 amino acid residues;
[0391] Preferably, the mutation is selected from substitution, addition, insertion, or deletion;
[0392] Preferably, the substitution is a conserved amino acid substitution;
[0393] Preferably, the polypeptide is recombinant collagen; more preferably, it is recombinant type XXII collagen; and more preferably, it is human recombinant type XXII collagen.
[0394] Preferably, the polypeptide has cell adhesion activity.
[0395] 2. The polypeptide according to claim 1, comprising an amino acid sequence selected from the group consisting of: SEQ ID NO: 2, 5, 7, 10, 13, 15, 16, 19, 22, 25, 28, 31, 34, 37, 40, 43, 46, 49, 52, 55, 58 or 29, wherein the variant is (1) an amino acid sequence in which one or more amino acid residues are mutated or (2) an amino acid sequence having at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with the amino acid sequence;
[0396] Preferably, the mutation is selected from substitution, addition, insertion, or deletion;
[0397] Preferably, the substitution is a conserved amino acid substitution.
[0398] 3. Nucleic acid encoding a polypeptide as described in item 1 or 2,
[0399] Preferably, the nucleic acid comprises a codon-optimized nucleotide sequence.
[0400] Preferably, the nucleotide sequence is codon-optimized for expression in *E. coli*;
[0401] Preferably, the nucleic acid comprises a nucleotide sequence selected from the group consisting of: SEQ ID NO: 3, 6, 8, 9, 11, 14, 17, 20, 23, 26, 29, 32, 35, 38, 41, 44, 47, 50, 53, 56 or 59.
[0402] 4. A vector comprising the nucleic acid as described in item 3,
[0403] Preferably, the vector comprises an expression control element operatively linked to the nucleic acid, a purified tag nucleotide, and / or a leader sequence nucleotide.
[0404] Preferably, the expression control element is selected from promoters, terminators, or enhancers;
[0405] Preferably, the purification tag is selected from His tag, GST tag, MBP tag, SUMO tag or NusA tag;
[0406] Preferably, the vector is an expression vector or a cloning vector, preferably pET-28a(+).
[0407] 5. A host cell comprising the nucleic acid according to claim 3 or the vector according to claim 4; wherein preferably, the host cell is a eukaryotic cell or a prokaryotic cell; wherein preferably, the eukaryotic cell is a yeast cell, an animal cell and / or an insect cell, and / or the prokaryotic cell is an Escherichia coli cell, such as Escherichia coli BL21.
[0408] 6. A composition comprising one or more of the polypeptide according to claim 1 or 2, the nucleic acid according to claim 3, the carrier according to claim 4, and the host cell according to claim 5; preferably, the composition is a kit; preferably, the composition is one or more of the following: biological dressings, human biomimetic materials, plastic and cosmetic materials, organoid culture materials, cardiovascular stent materials, coating materials, tissue injection filling materials, ophthalmic materials, obstetric and gynecological biomaterials, nerve repair and regeneration materials, liver tissue materials and vascular repair and regeneration materials, 3D printed artificial organ biomaterials, cosmetic raw materials, and pharmaceutical excipients; preferably, the composition is an injectable composition or an oral composition.
[0409] 7. Use of the polypeptide according to claim 1 or 2, the nucleic acid according to claim 3, the carrier according to claim 4, the host cell according to claim 5, and / or the composition according to claim 6 in one or more of the following: bio-dressings, human biomimetic materials, plastic and cosmetic materials, organoid culture materials, cardiovascular stent materials, coating materials, tissue injection filling materials, ophthalmic materials, obstetric and gynecological biomaterials, nerve repair and regeneration materials, liver tissue materials and vascular repair and regeneration materials, 3D printed artificial organ biomaterials, cosmetic raw materials, and pharmaceutical excipients.
[0410] 8. A method for promoting cell adhesion, comprising the step of contacting a polypeptide according to claim 1 or 2, a nucleic acid according to claim 3, a carrier according to claim 4, a host cell according to claim 5, and / or a composition according to claim 6 with a cell, preferably, the cell being an animal cell, more preferably a mammalian cell, and more preferably a human cell.
[0411] 9. A method for performing cosmetic surgery, tissue injection filling, ophthalmic treatment, nerve repair, or vascular repair on a subject in need, comprising administering the polypeptide according to item 1 or 2 to the subject, preferably by oral or injectable administration; preferably, the subject has a disease or condition related to type II collagen deficiency, such as muscular dystrophy; preferably, the subject is a human being.
[0412] 10. The production of the polypeptide according to claim 1 or 2, comprising:
[0413] (1) Culture the host cells as described in item 5 under suitable culture conditions;
[0414] (2) Harvesting host cells and / or culture medium containing polypeptides; and
[0415] (3) Purify the polypeptide;
[0416] Preferably, the host cell is an Escherichia coli cell, and more preferably an Escherichia coli BL21(DE3) cell;
[0417] Preferably, step (1) includes culturing E. coli cells in LB medium and inducing expression by IPTG;
[0418] Preferably, step (2) includes harvesting Escherichia coli cells, resuspending them in a balanced working solution, homogenizing the Escherichia coli cells, preferably by high-pressure homogenization, and separating the supernatant; preferably, the balanced working solution contains 100-500mM sodium chloride, 10-50mM Tris, 10-50mM imidazole, and pH 7-9.
[0419] Preferably, step (3) includes crude purification and one or more of the following: enzymatic digestion, fine purification and reverse nickel column purification;
[0420] Preferably, the crude purification includes Ni-agarose gel column purification of the supernatant to obtain an eluent containing the target protein, wherein the eluent contains 100-500 mM sodium chloride, 10-50 mM Tris and 100-500 mM imidazole, preferably pH 7-9;
[0421] Preferably, the enzyme digestion includes digestion with TEV enzyme, preferably with a protein-to-TEV enzyme ratio of 10-100:1 for 2-8 hours;
[0422] Preferably, the purification process includes gradient elution of the enzyme-digested product using a strong anion exchange chromatography column; preferably, the gradient elution includes 0-15% solution B for 1-5 minutes followed by 3 column volumes, 15-30% solution B for 1-5 minutes followed by 3 column volumes, 30-50% solution B for 1-5 minutes followed by 3 column volumes, and 50-100% solution B for 1-5 minutes followed by 3 column volumes; wherein solution B contains 10-50 mM Tris, 0.5-5 M sodium chloride, and pH 7-9;
[0423] Preferably, the reverse nickel column purification includes purification of the enzyme-digested product onto a Ni-agarose gel column; preferably, the eluent contains 10-50 mM Tris, 10-50 mM sodium chloride, 0.5-5 M imidazole, and pH 7-9.
Claims
1. Recombinant collagen, the amino acid sequence of which is SEQ ID NO:
5.
2. Nucleic acid encoding the recombinant collagen according to claim 1.
3. The nucleic acid according to claim 2, wherein the nucleic acid comprises a codon-optimized nucleotide sequence.
4. The nucleic acid according to claim 3, wherein the nucleic acid comprises a nucleotide sequence codon-optimized for expression in *E. coli*.
5. The nucleic acid according to any one of claims 2-4, wherein the nucleic acid is the nucleotide sequence shown in SEQ ID NO:
6.
6. A vector comprising the nucleic acid according to any one of claims 2-5.
7. The vector of claim 6, wherein the vector comprises an expression control element operatively linked to the nucleic acid, a polynucleotide of a purification tag, and / or a polynucleotide of a leader sequence.
8. The carrier according to claim 7, wherein the expression control element is selected from promoters, terminators, or enhancers.
9. The vector according to claim 7, wherein the purification tag is selected from His tag, GST tag, MBP tag, SUMO tag or NusA tag.
10. The vector according to any one of claims 6-9, wherein the vector is an expression vector or a cloning vector.
11. The carrier according to claim 6, wherein the carrier is pET-28a(+).
12. A host cell comprising a nucleic acid according to any one of claims 2-5 or a vector according to any one of claims 6-11.
13. The host cell according to claim 12, wherein the host cell is a eukaryotic cell or a prokaryotic cell.
14. The host cell according to claim 13, wherein the eukaryotic cell is a yeast cell, an animal cell, and / or an insect cell; and the prokaryotic cell is an Escherichia coli cell.
15. The host cell according to claim 14, wherein the Escherichia coli cell is Escherichia coli BL21.
16. A composition comprising one or more of the recombinant collagen according to claim 1, the nucleic acid according to any one of claims 2-5, the vector according to any one of claims 6-11, and the host cell according to any one of claims 12-15.
17. The composition according to claim 16, wherein the composition is a kit.
18. The composition according to claim 16, wherein the composition is one or more of the following: biological dressings, human biomimetic materials, plastic and cosmetic materials, organoid culture materials, cardiovascular stent materials, coating materials, tissue injection filling materials, ophthalmic materials, obstetric and gynecological biomaterials, nerve repair and regeneration materials, liver tissue materials and vascular repair and regeneration materials, 3D printed artificial organ biomaterials, cosmetic raw materials and pharmaceutical excipients.
19. The composition according to claim 16, wherein the composition is an injectable composition or an oral composition.
20. The use of the recombinant collagen according to claim 1, the nucleic acid according to any one of claims 2-5, the carrier according to any one of claims 6-11, and / or the host cell according to any one of claims 12-15 in the preparation of one or more of the following: bio-dressings, human biomimetic materials, plastic and cosmetic materials, organoid culture materials, cardiovascular stent materials, coating materials, tissue injection filling materials, ophthalmic materials, obstetric and gynecological biomaterials, nerve repair and regeneration materials, liver tissue materials and vascular repair and regeneration materials, 3D printed artificial organ biomaterials, cosmetic raw materials, and pharmaceutical excipients.
21. A method for promoting cell adhesion in vitro, comprising the step of contacting the recombinant collagen of claim 1 with cells in vitro, wherein the cells are 3T3 / NIH cells.
22. A method for producing the recombinant collagen according to claim 1, comprising: (1) Culture the host cells according to any one of claims 12-15 under suitable culture conditions; (2) Harvesting host cells and / or culture medium containing recombinant collagen; and (3) Purify recombinant collagen.
23. The method of claim 22, wherein the host cell is an Escherichia coli cell.
24. The method according to claim 23, wherein the Escherichia coli cells are Escherichia coli BL21(DE3) cells.
25. The method according to any one of claims 22-24, wherein step (1) comprises culturing Escherichia coli cells in LB medium and inducing expression by IPTG.
26. The method according to any one of claims 22-24, wherein step (2) comprises harvesting Escherichia coli cells, resuspending them in a balanced working solution, homogenizing the Escherichia coli cells, and separating the supernatant.
27. The method of claim 26, wherein the homogenization is high-pressure homogenization.
28. The method of claim 26, wherein the equilibration working solution comprises 100-500 mM sodium chloride, 10-50 mM Tris, 10-50 mM imidazole, and pH 7-9.
29. The method according to any one of claims 22-24, wherein step (3) comprises crude purification and one or more of the following: enzymatic digestion, fine purification and reverse nickel column purification.
30. The method of claim 29, wherein the crude purification comprises Ni-agarose gel column purification of the supernatant to obtain an eluent containing the target protein, wherein the eluent comprises 100-500 mM sodium chloride, 10-50 mM Tris and 100-500 mM imidazole.
31. The method of claim 30, wherein the pH of the eluent is 7-9.
32. The method of claim 29, wherein the digestion comprises digestion with a TEV enzyme.
33. The method according to claim 32, wherein the protein to TEV enzyme ratio is 10-100:1 for 2-8 hours.
34. The method of claim 29, wherein purification comprises gradient elution of the enzyme-digested product using a strong anion exchange chromatography column.
35. The method of claim 34, wherein gradient elution comprises 0-15% solution B for 1-5 minutes followed by 3 column volumes, 15%-30% solution B for 1-5 minutes followed by 3 column volumes, 30%-50% solution B for 1-5 minutes followed by 3 column volumes, and 50%-100% solution B for 1-5 minutes followed by 3 column volumes; wherein solution B comprises 10-50 mM Tris, 0.5-5 M sodium chloride, and pH 7-9.
36. The method of claim 29, wherein reverse nickel column purification comprises purifying the enzyme-digested product using a Ni-agarose gel column.
37. Use of the recombinant collagen of claim 1 in the preparation of a kit for promoting cell adhesion, wherein the cells are 3T3 / NIH cells.
Citation Information
Patent Citations
Nucleotide sequences for the control of the expression of DNA sequences in a cellular host
WO1994025612A2
Methods for using positively and negatively selectable genes in a filamentous fungal cell
WO2010039889A2
Cartyrin composition and method for use
CN112601583A
Allogeneic cell compositions and methods of use
CN113383018A