Collagen domain, collagen, recombinant collagen expression bacteria and application
By using thermal stability prediction analysis and sequence design, highly homologous collagen was expressed in Escherichia coli, solving the problem that collagen is difficult to form a triple helix structure in microorganisms. This resulted in highly stable collagen self-assembled fibers, meeting the needs of biomedical applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGNAN UNIV
- Filing Date
- 2023-09-04
- Publication Date
- 2026-08-04
AI Technical Summary
Existing technologies struggle to express the triple helix structure of human collagen with high homology in microorganisms, and the expressed collagen exhibits low stability, failing to self-assemble into higher-order structures. Furthermore, there is a lack of standardized methods for characterizing the triple helix structure.
Through systematic thermal stability prediction analysis, human type II and III collagen fragments were extracted and spliced, and repetitive sequence modules (such as (GPP)n) were designed and introduced. Collagen was expressed using E. coli to form highly thermally stable fragments, which then self-assembled into triple helical structures.
It has achieved high homology expression of human collagen in E. coli, forming a triple helix structure and periodic light and dark stripe fibers similar to natural collagen, meeting the needs of biomedicine and tissue engineering.
Smart Images

Figure CN117186210B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to collagen domains, collagen, recombinant collagen expression bacteria and their applications, specifically to a method for directly expressing recombinant collagen with a triple helix structure in Escherichia coli, and the recombinantly expressed human type I collagen can self-assemble into regular biomimetic fibers, belonging to the field of genetic engineering technology. Background Technology
[0002] Collagen is the most abundant structural protein in the human body, accounting for approximately 30% of total protein. It is widely distributed in tissues such as bones, tendons, cartilage, and skin. Collagen consists of three polypeptide chains forming a right-handed triple helix around a central axis. This triple helix structure can further assemble into higher-order collagen fibers, which perform their functions in the body. Therefore, the triple helix structure of collagen is the basis for its biological functions. Types I, II, and III collagen account for 80-90% of the total collagen in the human body. Type I collagen is the most abundant functional protein in animals. Its self-assembled collagen fibers, when characterized under a transmission electron microscope, exhibit alternating light and dark bands in overlapping and gap regions, commonly known as the D-cycle. The D-cycle is considered a key structural element that endows collagen with various functions and is related to the load-bearing properties of tissues, bone mineralization, and the regulation of cell differentiation and adhesion during tissue development. Type II collagen is found in the cartilage of the ribs, nose, throat, and trachea and can control the symptoms of joint-related diseases such as osteoarthritis. Type III collagen, along with type I, functions in the skin, ligaments, blood vessels, and joints, and is closely related to the skin damage repair process and its quality. In recent years, with the development of intelligent manufacturing in bioengineering, the demand for high-performance biomimetic materials has been increasing. Collagen materials, due to their excellent biocompatibility and low immunogenicity, have enormous application potential in skin damage treatment, vascular scaffold engineering, cartilage and bone defect repair, skin care, hemostatic sponges, and drug delivery, including coatings and medical nanoparticles.
[0003] Currently, the main source of collagen is animal extraction, but its potential immunogenicity limits its application in the field of biomedical materials. While it's possible to obtain polypeptide chains with collagen-like sequences through chemical synthesis, this is costly, and the synthesized polypeptide chains are limited in length, making them unsuitable for large-scale production. The use of genetic engineering to express natural or optimized human collagen sequences in microorganisms to obtain recombinant collagen is gaining increasing attention and has become a research hotspot. This not only solves the viral risks associated with traditional extraction methods but also allows for sequence modification to increase the hydrophilicity of collagen, resulting in stable and safe samples.
[0004] Microbial expression systems, with their advantages of clear genetic background, ease of genetic manipulation, short fermentation cycles, and high expression levels, are widely used for heterologous protein expression. However, due to the current difficulty in achieving humanized post-translational modifications in microbial expression systems, expressed human collagen cannot be modified to fold into a triple helix structure or self-assemble into higher-order structures. Furthermore, due to the unique structure of collagen, the forces driving collagen folding are not yet fully understood, resulting in a lack of sufficient theoretical support for sequence design in heterologous human collagen expression. Therefore, the problem of recombinantly expressed human collagen folding into a triple helix structure and further assembling into regular higher-order collagen structures remains unsolved.
[0005] Currently, some reports have shown that human collagen can be heterologously expressed through microorganisms such as E. coli, but at least one of the following problems exists: (1) It has low homology with natural collagen sequences. The sequences are obtained by random truncation, modification, and repetition / assembly based on experience, and have a triple helix structure verified by circular dichroism spectroscopy. For example, the literature "The self-assembly of a mini-fibril with axial periodicity from a designed collagen-mimetic triple helix" and "To achieve self-assembled collagen mimetic fibrils using The 108-amino acid collagen domain Col108 reported in the book *Designed Peptides* is derived from the splicing of four short sequence fragments from the collagen domain of human type I collagen, with only 45.61% homology to the natural sequence. CN115521373A discloses a triple-helix recombinant humanized type I collagen, its preparation method, and its applications. The expressed recombinant humanized type I collagen has a triple-helix structure and can self-assemble to form collagen fibers. The collagen domain fragment in the above patent is the Col108 fragment reported in the aforementioned literature, with an inserted functional motif, showing low homology to the natural sequence. The dissertation *Preparation, Structural Characterization, and Performance Analysis of Recombinant Human-like Collagen* designed a 38-amino acid sequence recombinant human-like collagen monoclonal antibody. The fragments were repeated 4 or 8 times respectively and expressed using E. coli. The synthesized human-like collagen had a triple helix structure, but a sequence search of this single collagen fragment could not match any human collagen. CN115819557A discloses a triple helix recombinant humanized type II collagen, its preparation method and application. The expressed recombinant humanized type II collagen has a triple helix structure and can self-assemble to form collagen fibers. The sequence matches the human collagen sequence for up to 7 consecutive amino acid residues. CN115521372A discloses a triple helix recombinant humanized type III collagen, its preparation method and application. The sequence matches the natural sequence for up to 9 consecutive amino acids, but a sequence search could not match any human collagen.
[0006] (2) The expressed collagen has low stability and lacks a triple helix structure at room temperature. For example, the literature "Recombinant expression of hydroxylated human collagen in Escherichia coli The expression of prolyl and lysyl hydroxylases of the mimic virus, along with fragments of human type III collagen, promotes its folding into a triple-helix structure. T m With a temperature of only 24.3℃, collagen with low stability is prone to losing its triple helix structure during in vivo and in vitro applications, thus failing to perform its function.
[0007] (3) The expressed collagen was not characterized by a standardized triple helix, and its triple helix structure could not be determined. Based on the "Guidelines for the Evaluation of Recombinant Humanized Collagen Raw Materials" and Nature Protocols, 2006: VOL.1, NO.6, 2527, etc., the literature "Selective expression of nonsecreted triple-helical and secreted single-chain recombinant collagen fragments in the yeast" was cited. Pichia pastoris The study described the recombinant expression of human type I-III collagen fragments in Pichia pastoris, and its subsequent research, "Expression of recombinant human type I-III collagens in the yeast". Pichia pastoris The study co-expressed proline hydroxylase and human type I, II, and III collagen in *Pichia pastoris*, but failed to characterize the triple helix structure in any of these studies. The paper "Production of human type I collagen in yeast reveals unexpected new insights into the molecular assembly of collagen trimers" promoted the folding of chicken proline hydroxylase and human type I collagen into a triple helix structure through co-expression, but the study only measured the triple helix structure using a thermochromic curve at 197 nm. T m The value is 30℃. At this wavelength, the absorption peak is often similar to the protein spectrum in its unfolded state, making it unsuitable as a standard method for characterizing the triple helix of collagen, thus preventing the determination of the triple helix structure. CN114276435A discloses a recombinant human type III collagen and its applications, selecting a 123-amino acid sequence segment, directionally replacing the tripeptide sequence in this segment, and repeating the process. The C-terminus is linked to a specific sequence, and expression is performed using Pichia pastoris without characterization of the triple helix structure; CN114774460A discloses recombinant human type I triple helix collagen from yeast and its preparation method, which selects the α1 chain sequence of human type I collagen and co-expresses it with hydroxylase; patent CN114480471A discloses recombinant human type III triple helix collagen from yeast and its preparation method, which selects the α1 chain sequence of human type III collagen and co-expresses it with hydroxylase; CN111087464B discloses a recombinant human type III collagen with a functional structure and its expression method, which selects a partial sequence fragment of human type III collagen and co-expresses it with hydroxylase; CN Patent 112851797B discloses a recombinant human type III collagen, its preparation method, and its uses. It involves splicing fragments of human type III collagen with cell-binding capabilities and co-expressing them with hydroxylase. Patent CN116555320A discloses a recombinant human type III triple-helix collagen engineered bacteria, its construction method, and its applications. It selects the α1 chain sequence of human type III collagen and co-expresses it with hydroxylase. Patent CN116082494A discloses a recombinant human type III collagen polypeptide, expression vector, expression strain, and its construction method. It selects a 54-amino acid polypeptide fragment with strong hydrophilicity and stability from the human type III collagen sequence and expresses it in Pichia pastoris. None of these seven patents characterize the triple-helix structure; whether a triple-helix structure can actually be formed remains unknown.
[0008] Furthermore, during the preliminary research process, the inventors' team disclosed CN111333715B (a method for preparing type I collagen fibers) using N-terminus and C-terminus (GPP). n Based on the sequence, insert consecutive Gly in the middle. Xaa The collagen sequence of the Yaa triplet forms a band-like fiber with periodic light and dark stripes, and CN111499729B (a method for regulating the stripe period length of type I collagen fibers), with N and C terminals (PPG). n Based on the sequence, insert consecutive Gly molecules with different quantities in the middle. Xaa The collagen sequence of the Yaa triplet forms a band-like fiber with periodic alternating light and dark stripes of different dark stripe lengths. None of the above systematically designed human collagen sequences. In the master's thesis of the inventor team, Yan Haojie, entitled "Hierarchical Self-Assembly of Collagen Peptides Induced by Multiple Non-Covalent Interactions", a human type I collagen sequence fragment was selected and expressed in E. coli, which can form a triple helix structure, but it did not assemble into a fibrous structure similar to that of natural human collagen.
[0009] Therefore, it is necessary to develop collagen sequences that are highly homologous to natural human collagen and capable of exogenous expression of triple-helix structures, based on systematic thermal stability analysis. Summary of the Invention
[0010] To address at least one of the problems associated with recombinant human collagen, such as low homology with natural human collagen, difficulty in heterologous expression to form triple-helix structures, or difficulty in further self-assembling to form higher-order structures, this invention utilizes systematic thermal stability prediction analysis to extract human collagen... Type I, II, and III collagen fragments were sequenced and designed to obtain collagen domains (also known as collagen structural domains or collagen domains) with high homology to natural collagen; further, repeating sequence modules (GPPs) were introduced at both ends of the collagen domains. n The designed collagen sequence was expressed in *E. coli*, and it was found that the designed high-thermal-stability collagen fragments could all fold correctly to form a triple-helix structure, while the low-thermal-stability fragments could not fold correctly. Furthermore, the designed high-thermal-stability recombinant human type I collagen could self-assemble to form periodic light and dark stripes similar to those of natural type I collagen. This invention has developed and achieved the expression of a collagen sequence with high homology to natural human collagen, enabling exogenous expression of a triple-helix structure, thus meeting the demand for recombinant collagen with structural function in the biomedical and tissue engineering fields.
[0011] The first object of the present invention is to provide an amino acid sequence encoding a collagen domain, said amino acid sequence having: (1) The amino acid sequence as shown in SEQ ID NO.1~7, or (2) An amino acid sequence obtained by combining any two sequences of SEQ ID NO. 1~3, or (3) The amino acid sequence obtained by repeating any of the sequences shown in SEQ ID NO.1~7 2~3 times.
[0012] In one embodiment, the amino acid sequences represented by SEQ ID NO. 1-7 are derived from natural sources. It is obtained by sequence extraction or further sequence splicing and design of collagen fragments of type I, II, and III human collagen.
[0013] In one embodiment, the amino acid sequences represented by SEQ ID NO. 1-7 are obtained by analyzing natural... Thermal stability prediction was performed on human collagen types I, II, and III, and high predictive strength was selected. T m The sequence of values is obtained by truncating or splicing. The amino acid sequence is used as a predictor of the collagen triple helix structure of the collagen domain.T m The value is between 38-39℃.
[0014] The amino acid sequence serves as a prediction of the collagen triple helix structure of the collagen domain. T m The specific prediction method is as follows: taking the first triplet unit (XYG) of the triple helix structure as the starting point for continuous numbering, the average relative stability of each XYG triplet is calculated to obtain the thermal stability value of each triplet; then, n consecutive triplets are taken, and the average of the thermal stability values of these n consecutive triplets is calculated, which is the predicted thermal stability value of the collagen domain sequence; where the thermal stability value of a single triplet i refers to the thermal stability value of a window composed of 10 consecutive triplets in the interval [i-5, i+5); the thermal stability value T of the window. windows Based on the window main chain tendency value T bb Interaction value T between window sidechains side Decide, .
[0015] A second objective of this invention is to provide a single-chain protein for expressing collagen, the single-chain protein containing the amino acid sequence encoding the collagen domain described above.
[0016] In one embodiment, the structure of the protein single chain includes: a folding domain, a repeating sequence module, and a collagen domain.
[0017] In one embodiment, the introduction of the folding domain assists collagen in folding to form a triple helix structure. Optionally, the folding domain is a V-domain with the amino acid sequence shown in SEQ ID NO.13; alternatively, the folding domain is a coiled-coil domain with the amino acid sequence shown in SEQ ID NO.14.
[0018] In one embodiment, the introduction of the repetitive sequence modules can assist in the folding of the collagen triple helix and improve its thermal stability. Optionally, there are multiple repetitive sequence modules, located at both ends of the collagen domains or at both ends of multiple collagen domains; for example, when expressing type II collagen, there can be multiple collagen domains, which are connected by repetitive sequence modules. Optionally, the sequences of the repetitive sequence modules can be the same or different.
[0019] In one implementation, the repeat sequence module employs General Process Planning (GPP). n Optionally, when there are multiple repeating sequence modules, each repeating sequence module (GPP) n In this case, n can take the same or different values. Optionally, for the repeating sequence module (GPP)... nBy adjusting the number of 'n' molecules, the molecules can further assemble into fibrous structures. Optionally, for molecules that can assemble into fibrous morphologies (GPP)... n Collagen (GPP) n In this pattern, the two n's are equal (GPP). n The value of n satisfies 5 < n ≤ 30, and the value of n disclosed in CN 111333715 B, published in the inventors' team's previous research, can be referenced. Optionally, for type II and type III triple helix (GPP) n Collagen (GPP) n Collagen (GPP) n In this pattern, the three n's can be different.
[0020] In one embodiment, the folded domain and the repeat sequence module are linked by an enzyme restriction site, such as an LVPRGSP (sequence as SEQ ID NO. 21). Optionally, the folded domain (V-domain) and the repeat sequence module (GPP) are... n They are connected by LVPRGS (sequence shown in SEQ ID NO.22).
[0021] In one embodiment, the structure of the single-chain protein used to express collagen, from the N-terminus to the C-terminus, includes: a folding domain, an enzyme cleavage site, a {repetitive sequence module, collagen domain}m, and a repetitive sequence module; wherein m is greater than or equal to 1. Optionally, m is 1 or 2.
[0022] In one implementation, the front end (N end) of the folded domain has a 6×His tag.
[0023] In one embodiment, the protein single chain used to express collagen has the following structure: Figure 10 As shown; or the structure is as follows Figure 11 As shown.
[0024] A third object of the present invention is to provide a nucleotide sequence encoding the collagen domain, or a nucleotide sequence encoding the single-chain protein for expressing collagen, or a gene encoding the single-chain protein for expressing collagen, and a plasmid or cell expressing the gene.
[0025] Optionally, the plasmid may be a pColdIII series or pET series plasmid. The cell is an *E. coli* cell, including... E. coli BL21 E . coli BL21(DE3), E. coli Rosetta (DE3) E. coli BL21(DE3)pLysS / pLysE or E. coli Origami2 (DE3), etc.
[0026] A fourth objective of this invention is to provide a collagen protein consisting of three of the aforementioned protein single chains coiled around a common central axis, forming a triple helix structure.
[0027] A fifth object of the present invention is to provide collagen fibers formed by the self-assembly of said collagen polymers.
[0028] In one embodiment, the collagen is type I collagen. Optionally, the collagen fibers have periodic alternating light and dark stripes; alternatively, the collagen fibers exhibit bright stripe morphology when negatively stained under TEM.
[0029] In one embodiment, the collagen fibers can be processed using (GPP). n The quantity is obtained through regulation; optionally, the repetitive sequence module in the regulation is (GPP). 10 The length of the corresponding bright stripe is 10 nm.
[0030] In one embodiment, the amino acid sequence of the collagen domain of the present invention is introduced into the collagen domain region of type I collagen, so that the dark stripes in the collagen fiber reach: (number of amino acids in the collagen domain region ÷ 3 × 0.9) ± 1 nm.
[0031] The sixth object of the present invention is to provide a product containing the collagen of the present invention.
[0032] The products mentioned are products in the fields of beauty, chemical, medical / biomedical, cosmetics, and feed, such as beauty cosmetics (masks, serums, creams, etc.), artificial collagen casings, nutritional supplements (collagen powder, oral liquids), medical dressings, hemostatic materials, artificial bone scaffolds, injectable fillers, artificial blood vessels, eye drops, drug sustained-release carriers, etc.
[0033] A seventh object of the present invention is to provide applications in the preparation of collagen-containing products in the fields of biology, chemical engineering, pharmaceuticals, biomaterials, tissue engineering, or cosmetics, said applications including the use of amino acid sequences encoding collagen domains, protein single chains, collagen, collagen fibers, or nucleotide sequences encoding said collagen domains, nucleotide sequences encoding said protein single chains for expressing collagen, genes encoding said protein single chains for expressing collagen, or plasmids or cells expressing said genes.
[0034] The recombinant human collagen provided by this invention can fold into a triple helix structure and controllably self-assemble into a regular, high-order biomimetic fiber structure. This invention utilizes natural... Thermal stability prediction was performed on type I, II, and III collagen, and high / low prediction values were selected. T m A sequence of values forms a collagen domain, optionally directed to... Figure 10 The structure shown or as Figure 11 The collagen domain of the structure shown introduces different types of collagen sequences; the introduction of folding domains (such as the V-domain of SEQ ID NO. 13 or the coiled-coil domain of SEQ ID NO. 14) assists collagen folding to form a triple helix structure; repetitive sequence modules (such as (GPP)) n The introduction of (GPP) can assist in the folding of the collagen triple helix and improve its thermal stability; through (GPP) n By adjusting the quantity, molecules can be further assembled to form fiber structures, which exhibit bright stripe morphological characteristics under negative staining in TEM.
[0035] This invention also analyzed the thermal stability of the obtained recombinant collagen, characterized it using TEM, and identified highly thermally stable fragments for designing type I, II, and III collagen. Although in practice... T m Values and Predictions T m The values (38-39℃) have some deviation, but they can all be correctly folded to form a triple helix structure, while the predicted low thermal stability segments cannot be folded to form a triple helix structure.
[0036] Beneficial effects 1. The collagen domain of this invention is extracted from natural human collagen. The sequence was designed by splicing collagen fragments from type I, II, and III collagen, and it has high homology with natural human collagen. Among them, the fragments directly extracted from natural human collagen have 100% homology with the natural sequence, and the spliced collagen domain sequence has more than 57% homology with the natural sequence.
[0037] 2. Based on the prediction of the thermal stability of human collagen sequences, this invention performs sequence screening and design, and successfully achieves heterologous expression of high thermal stability fragments of different types of human collagen in Escherichia coli.
[0038] 3. The predicted collagen sequence of the present invention. T m The value is between 38-39℃; the thermal change temperature of the collagen domain was determined using circular dichroism spectroscopy. T m The value is also quite close to human body temperature.
[0039] 4. This invention utilizes sequences with high homology to human collagen to achieve the expression of human type I collagen with a triple-helix structure in *E. coli*, capable of self-assembling into regular, high-order biomimetic fiber structures, thus overcoming the current difficulties in recombinant human collagen expression. Furthermore, the synthesized human type I collagen can all self-assemble into fibers with periodic light and dark stripes, similar in morphology to type I collagen. This can meet the demand for recombinant collagen with structural functions in the biomedical and tissue engineering fields.
[0040] 5. This invention introduces / incorporates integrin binding sites into a designed highly stable human type I collagen sequence, enabling it to fold into a stable triple helix structure and self-assemble into a fibrous morphology. This invention provides a reference for the application of recombinant collagen in tissue culture, tooth tissue repair, etc., and also provides a basis for introducing other functional motifs into collagen sequences. Attached Figure Description
[0041] Figure 1 This is a schematic diagram of the interaction between axial and transverse side chains in the triple helix structure of collagen.
[0042] Figure 2 This is a relative stability map of type I collagen.
[0043] Figure 3 A schematic diagram of collagen sequence design.
[0044] Figure 4 This is an SDS-PAGE image of purified collagen; lanes 1-7 represent the purified V-HC1-1, V-HC1-2, V-HC1-3, V-HC1-12, V-HC1-22, V-HC1F, and V-HC1E, respectively, and the arrows represent the target bands; M: protein marker.
[0045] Figure 5 This is an SDS-PAGE image of purified collagen; lanes 1-2 and 4-6 represent purified V-HC2A, V-HC2B, V-HC3A, V-HC3B, and V-HC3C, respectively, and the arrows indicate the target bands; M: protein marker.
[0046] Figure 6 (a) is the circular dichroism chromatogram of the designed type I collagen; (b) is the full wavelength spectrum; (c) is the thermochromatogram.
[0047] Figure 7 The circular dichroism chromatograms of the designed type II and type III collagen are shown in Figure 1; (a) shows the full wavelength spectrum and thermochromic curve of type II collagen; (b) shows the full wavelength spectrum and thermochromic curve of type III collagen.
[0048] Figure 8 (a) shows the morphology of the self-assembled collagen fibers; (b) shows the TEM characterization and spectral density statistics of HC1-1, HC1-2 and HC1-3.
[0049] Figure 9 (c) shows the morphology of the self-assembled collagen fibers; (d) shows the TEM characterization and spectral bandwidth statistics of HC1-12 and HC1-22.
[0050] Figure 10 The structure of the single-chain protein used to express collagen is a folded domain-repetitive sequence-collagen domain-repetitive sequence.
[0051] Figure 11 The structure of the single-chain protein used to express collagen is: folded domain-repetitive sequence-collagen domain-repetitive sequence-collagen domain-repetitive sequence. Detailed Implementation
[0052] Culture medium: LB medium (g / L): tryptone 10, yeast extract 5, NaCl 10, agar powder 15 (solid). TB medium (g / L): tryptone 12, yeast extract 24, glycerol 4 mL, KH2PO4 2.31, K2HPO4 12.54; Culture method: 50 μL of bacterial culture from the glycerol tube containing the target gene was transferred to 20 mL of LB (Amp resistant) and cultured overnight at 37°C and 200 r / min. Then, 1% of the culture was transferred to 100 mL of TB fermentation broth (Amp resistant) and cultured at 37°C and 200 r / min for 24 h. IPTG was then added to a final concentration of 1 mmol / L, and fermented at 25°C and 200 r / min for 10 h, followed by fermentation at 15°C for 14 h.
[0053] Protein purification method: After fermentation, collect the bacterial cells, break them up, centrifuge, collect the supernatant, and filter it through a 0.45 μm aqueous filter membrane. Then use His Trap. TMHP affinity purification was performed using 5 mL of buffer A (20 mmol / L Na₂HPO₄, 20 mmol / L NaH₂PO₄, 500 mmol / L NaCl, 10 mmol / L Iminazole, pH 7.4), followed by loading at a flow rate of 5 mL / min. After loading, a gradient elution was performed using elution buffer B (20 mmol / L Na₂HPO₄, 20 mmol / L NaH₂PO₄, 500 mmol / L NaCl, 500 mmol / L Iminazole, pH 7.4) to obtain the target protein. The purification status was analyzed using SDS-PAGE.
[0054] Trypsin salt removal: Purified collagen was dissolved in water to a concentration of 4 mg / mL. 200 μL samples were taken and trypsin at a concentration of 2.5 g / L was added at molar ratios of 20:1, 200:1, and 2000:1. The mixture was digested in a 16℃ water bath, with samples taken every 3 h. Finally, the mixture was digested in an incubator for 12 h, and purity was verified by SDS-PAGE analysis. After digestion under optimal conditions, HiTrap Desalting was performed, and the peak samples were collected and freeze-dried under vacuum.
[0055] Sample stability assessment: The desalted samples were vacuum freeze-dried, and their full-wavelength and thermal stability were assessed using circular dichroism spectroscopy (CD). The specific steps were as follows: The freeze-dried samples were dissolved in 10 mmol / L, pH 7.0 sodium phosphate buffer to a concentration of 1 mg / mL, equilibrated at 4℃ for 48 h, and then subjected to CD analysis. Full-wavelength CD spectra were measured at 1 nm intervals from 190 to 250 nm at 4℃, with an average scan time of 5 s. The thermal curve was obtained by monitoring the CD signal at 225 nm, increasing the temperature from 4℃ to 70℃ at a rate of 10℃ / h, equilibrating for 8 s at each temperature, and determining the melting temperature (…). T m The stability of the sample is obtained by taking the median absorbance value of the fitted thermal curve at 4°C and 70°C.
[0056] Transmission electron microscopy (TEM) characterization: The lyophilized collagen sample was dissolved in 10 mmol / L, pH 7.0 sodium phosphate buffer to prepare a final concentration of 0.5 mmol / L, and self-assembled at 4°C for 4 days. 5 μL of the assembled sample was dropped onto a copper grid and allowed to adsorb for 30 s. Excess liquid was blotted away with filter paper, and then 5 μL of 0.75% phosphotungstic acid was added for negative staining. After maintaining the stain for 20 s, the stain was removed, and the sample was air-dried. The images were then observed using a Hitachi H-7650 TEM microscope at 80 kV. At least five clear TEM images were selected, and the bandwidths of the bright and dark fringes were measured using ImageJ. At least 200 measurements were taken for each sample, and the average value was calculated.
[0057] Thermal stability analysis methods: The steps include: (1) The encoding structure is as follows Figure 10 As shown or structured Figure 11 The gene for the protein shown is in E. coli. E. coli Expressed in BL21 (DE3); (2) The intracellularly expressed product was purified to obtain the purified protein, and then identified by SDS-PAGE. (3) The purified sample was digested with trypsin, and after the V-domain was completely removed by SDS-PAGE, it was desalted and lyophilized. (4) The lyophilized collagen sample was prepared into a solution with a final concentration of 1 mg / mL using 10 mmol / L sodium phosphate buffer, equilibrated at 4℃ for 48 h, and identified by full-wavelength circular dichroism spectroscopy and thermal temperature scanning. (5) The lyophilized type I collagen sample was prepared into a collagen solution with a final concentration of 0.5 mmol / L using 10 mmol / L sodium phosphate buffer, equilibrated at 4℃ for 4 days, and then characterized by TEM.
[0058] Example 1: Design of collagen domain sequences The full-length sequence of natural human collagen is analyzed using protein computation and thermostability prediction to obtain a sequence fragment with high thermostability. This fragment is then directly extracted or further spliced to obtain the collagen domain sequence. The obtained target collagen domain sequence is used as the predicted thermostability value for the collagen triple helix structure. T m The value is between 38 and 39℃.
[0059] Among them, the predicted thermal stability of the collagen triple helix structure formed by the collagen domain sequence ( T mThe prediction method is as follows: Starting with the first triplet unit (XYG) of the triple helix structure as the starting point for consecutive numbering, the average relative stability is calculated for each XYG triplet to obtain the thermal stability value of each triplet. Then, n consecutive triplets are taken, and the average of these n consecutive triplet values is calculated, which is the predicted thermal stability value of the collagen domain sequence. The target collagen domain sequence of this invention ensures the predicted thermal stability value of the collagen domain sequence while maximizing n. T m The value is between 38 and 39℃.
[0060] Among them, a single triplet i The thermal stability value refers to the thermal stability value of a window consisting of 10 consecutive triplets in the interval [i-5, i+5).
[0061] The thermal stability value T of the window windows Based on the window main chain tendency value T bb Interaction value T between window sidechains side The decision, among which, .
[0062] The T bb It is obtained using the following method: (1) Based on the host-guest system, using the most stable triplet Pro-Hyp-Gly as the host, guests were constructed by single-point mutation of 19 non-Pro residues at the X position of Pro, and the thermal stability value of each guest was measured. T m This refers to the main chain tendency value at different X positions; similarly, guests were constructed by single-point mutation of 20 natural amino acids of Hyp at the Y position in the Pro-Hyp-Gly triplet, and the thermal stability value of each guest was measured. T m This refers to the main chain tendency value at different Y positions; (2) The calculation of the main chain tendency value for any triplet XYG is based on the types of residues at positions X and Y in the triplet. The corresponding main chain tendency value at position X and position Y in (1) is found, and the main chain tendency value at position X is used to calculate the main chain tendency value. T X and the main chain tendency value at position Y T Y Adding them together, we get: T X + T Y For example, the Ala-Ala-Gly triad has a main-chain susceptibility value of [value missing]. T X + T Y ,T X (X=Ala) represents the Ala-Hyp-Gly determination T m value, T Y (Y=Ala) indicates that the Pro-Ala-Gly assay... T m value; (3) Window main chain tendency value T bb It is based on the calculation method of the main chain tendency value of any triplet XYG in (2), and is obtained by summing the main chain tendency values of all triplets in the window; wherein, the window includes 3 chains, and each chain has 10 triplets (i.e., includes 60 triplets). .
[0063] The T side It is generated by the interaction of all sidechains within the window. .
[0064] Where, Δ T Axi Δ represents the axial interaction value between two chains of adjacent triplets. T Lat This represents the lateral interaction value between the two chains of adjacent triplets.
[0065] The triple-helix folding structure constrains the interaction between adjacent chains into two types of geometries: axial and lateral. Figure 1 When the Y and X positions of two adjacent chains interact in a direction parallel to the helical axis, this is called axial interaction; when the Y and X positions of two adjacent chains interact in a direction perpendicular to the helical axis, this is called transverse interaction.
[0066] Δ T Axi and Δ T Lat The values, determined and calculated through double-mutation experiments, represent the differences between the thermal stability of double mutations at the Y and X sites in axial or transverse geometria and the sum of the stability of single-point mutations at the Y or X sites, as shown in the following formula: ; in, T OP This indicates the experimentally measured value when the Y-bit is Hyp and the X-bit is Pro. T m value; T OX This indicates the experimentally measured value when there is a single-point mutation at position X and position Y remains Hyp. Tm value; T YP This indicates the experimentally measured result when there is a single-point mutation at the Y position and the X position remains Pro. T m value; T YX This indicates the experimentally measured result when there is a double mutation at both the Y and X positions, i.e., the Y position is not Hyp and the X position is not Pro. T m value.
[0067] For example, calculate the lateral action value (Δ) when Y-bit is Lys and X-bit is Asp. T Lat When the double mutation was measured, the thermal stability value was... T YX (Y=Lys, X=Asp), and the corresponding X-position single-point mutation determination T m Value T OX (X=Asp, Y=Hyp), the Tm value determined by single-point mutation at position Y is... T YP (Y=Lys, X=Pro). Main thermal stability value T OP If unchanged, it means Y=Hyp and X=Pro. In the transverse interaction, the Y position can mutate into 20 other natural amino acids, and the X position can mutate into 19 other natural amino acids. There are a total of 20 × 19 = 380 combinations of different Y and X positions, corresponding to 380 transverse interaction values (Δ). T Lat Axial interaction (Δ) T Axi Using a similar method, 380 axial interaction values can also be obtained. For details, please refer to the master's thesis of Liu Han from the inventor team, "The Influence of Amino Acid Components on the Thermal Stability of Collagen-like Peptides".
[0068] The window unit contains 3 chains ( a , b , c (Chains), arranged with one residue misaligned, each chain containing 10 triplets (e.g.) Figure 1 (As shown). Chain a and Chain b Between, chain b and Chain c Between them, there are 10 transverse and 9 axial interaction pairs; chains c and Chain aThere are 9 transverse and 8 axial interaction pairs between them. Therefore, within the window of 10 triplet chains, there are a total of 29 transverse interaction values and 26 axial interaction pairs between the three chains. Summing these values separately yields... ∑Δ T Lat and ∑Δ T Axi The sum of the contributions of all axial and lateral sidechain interactions within the window is... T side .
[0069] The above methods involve experimental determination. T m The thermal stability was determined using circular dichroism spectroscopy. Specifically, the lyophilized pure host or guest collagen peptide powder was weighed and dissolved in 10 mM phosphate buffer (pH 7.0) to prepare a high-concentration (1 mM) stock solution. The stock solutions of the host peptide and guest peptide were further diluted to a final concentration of 0.2 mM. a chain, b Chain and c The collagen triple helices were mixed in a 1:1:1 ratio and heated at 80°C for 10 minutes to unwind the folded triple helices into a disordered single-chain state. The mixture was then incubated at 4°C for at least 24 hours to allow for full self-assembly and the formation of well-folded collagen triple helices. Circular dichroism (CD) experiments were performed on a Chirascan instrument (Applied Photophysics Ltd, England). Wavelength scans from 190 nm to 260 nm were performed at 4°C, with 1 nm intervals between each scan. A thermal distortion experiment was conducted at 225 nm, with the temperature increased from 4°C to 80°C at a gradient rate of 1°C / 6 min. The first derivative of the thermal distortion curve was used to determine the... T m value.
[0070] like Figure 2 This is a relative stability map of type I collagen obtained by calculating the thermal stability value of each triplet in the natural type I human collagen sequence and plotting the average relative stability curve according to the above method. It can be obtained from... Figure 2 By extracting continuous triplets with high thermal stability values, a sequence fragment with high thermal stability can be obtained, or the extracted sequences with high thermal stability can be spliced together to obtain a collagen domain sequence.
[0071] Following the above method, thermally stable sequence fragments were extracted from natural type I human collagen α1 chains (NCBI accession number NP_000079.2), type II human collagen α1 chains (NCBI accession number NP_001835.3), and type III human collagen α1 chains (NCBI accession number NP_000081.2), or the extracted thermally stable sequences were further spliced to obtain collagen domain sequences. Predicted sequences were selected. T m Collagen domain sequences with a temperature range of 38–39 °C and a high tendency for triple helix were selected as target sequences. T m Sequences with lower values and lower triple helix tendency were used as controls.
[0072] Among them, different types of collagen prediction T m The following are several sequences with values between 38 and 39℃: (1) The amino acid sequences shown in SEQ ID NO. 1~7; (wherein, SEQ ID NO. 1~3 are fragments or multiple fragments selected from natural type I human collagen and spliced together, named HC1-1, HC1-2, and HC1-3 of type I collagen, and predicted) T m The temperatures were 38.4℃, 38.5℃, and 38.2℃, respectively; SEQ ID NO.4 is a fragment or multiple fragments spliced from natural type II human collagen, named HC2A of type II collagen, and is predicted to be... T m The temperature was 38.3℃; SEQ ID NO. 5~7 were fragments or multiple fragments spliced from natural type III human collagen, named HC3A, HC3B, and HC3C of type III collagen, and were predicted to be... T m (The temperatures were 38.8℃, 38.8℃, and 39.0℃ respectively.) (2) The amino acid sequence obtained by combining any two sequences of SEQ ID NO.1~3, such as SEQ ID NO.8 (named HC1-12, predicted) obtained by combining SEQ ID NO.1 and SEQ ID NO.2. T m (38.4℃). (3) The amino acid sequence obtained by repeating any of the sequences shown in SEQ ID NO.1~7 2 to 3 times, for example, SEQ ID NO.9 obtained by repeating SEQ ID NO.2 2 times (named HC1-22, predicted) T m (38.4℃).
[0073] Predicted T m The following sequences have lower values (36-37℃): SEQ ID NO.10~12 (named HC1E, HC1F, HC2B, predicted). T m The temperatures were 37.1℃, 36.3℃, and 36.5℃ respectively.
[0074] As shown in Table 1, in SEQ ID NO.1~12, all directly selected sequence fragments were not sequence modified (100% homology with natural human collagen sequence), and all spliced collagen domain sequences had greater than 57% homology with natural human collagen sequence.
[0075] Table 1
[0076] Example 2: Collagen Sequence Design Design a single-chain protein containing the collagen domain of Example 1. The structure of the single-chain protein includes: a folded domain, a repeat sequence module, and a collagen domain.
[0077] The introduction of the folding domain assists collagen in folding to form a triple helix structure. Optionally, the folding domain is a V-domain or a coiled-coil domain; optionally, the amino acid sequence of the V-domain is as shown in SEQ ID NO.13; optionally, the amino acid sequence of the coiled-coil domain is as shown in SEQ ID NO.14.
[0078] The introduction of the repetitive sequence modules can assist in the folding of the collagen triple helix and improve its thermal stability. Optionally, there may be multiple repetitive sequence modules, located at both ends of the collagen domains or at both ends of multiple collagen domains; for example, when expressing type II collagen, there may be multiple collagen domains, which are connected by repetitive sequence modules. Optionally, the sequences of the repetitive sequence modules may be the same or different. Optionally, the repetitive sequence modules adopt (GPP). n Optionally, when there are multiple repeating sequence modules, each repeating sequence module (GPP) n The value of n can be the same or different.
[0079] As an example, the design structure of this embodiment is as follows: Figure 10 The amino acid sequence of the single chain of collagen protein is shown. The amino acid sequence of the V-domain is shown in SEQ ID NO.13; the collagen domain sequence uses SEQ ID NO.1~12 of Example 1.
[0080] As one example, such as Figure 3 As shown, for the sequences derived from type I collagen, repeat sequence modules (Gly-Pro-Pro) were inserted at both ends of the sequences HC1-1, HC1-2, HC1-3, HC1-12, HC1-22, HC1E, and HC1F in Example 1. 10 Short peptides, abbreviated as (GPP). 10 (SEQ ID NO.23), and after inserting a V-domain at the N-terminus, the sequences were named V-HC1-1, V-HC1-2, V-HC1-3, V-HC1-12, V-HC1-22, V-HC1E, and V-HC1F, respectively. The HC1-12 sequence is a splice combination of the HC1-1 and HC1-2 sequences, and the HC1-22 sequence is a splice combination of two HC1-2 sequences. The amino acid sequence of V-HC1-1 is shown in SEQ ID NO.15, and the nucleotide sequence encoding V-HC1-1 is shown in SEQ ID NO.16. The amino acid sequences of V-HC1-2, V-HC1-3, V-HC-12, V-HC1-22, V-HC1E, and V-HC1F are obtained by substituting the corresponding collagen domain sequences from Example 1 into the amino acid sequence of V-HC1-1.
[0081] For sequences derived from type II and III collagen, and considering morphological matching with natural collagen, short peptides (Gly-Pro-Pro)5, (Gly-Pro-Pro)4, and (Gly-Pro-Pro)6 were inserted at the N-terminus, middle, and C-terminus of the collagen fragments in sequences HC2A, HC2B, HC3A, HC3B, and HC3C, abbreviated as (GPP)5 (SEQ ID NO.24), (GPP)4 (SEQ ID NO.25), and (GPP)6 (SEQ ID NO.26), respectively. These sequences were named V-HC2A, V-HC2B, V-HC3A, V-HC3B, and V-HC3C, as designed in the following manner. Figure 3 As shown in SEQ ID NO.17, the amino acid sequence of V-HC2A is shown in SEQ ID NO.18. The amino acid sequence of V-HC2B is obtained by substituting the corresponding collagen domain sequence from Example 1 into the amino acid sequence of V-HC2A. The amino acid sequence of V-HC3A is shown in SEQ ID NO.19, and the nucleotide sequence of V-HC3A is shown in SEQ ID NO.20. The amino acid sequences of V-HC3B and V-HC3C are obtained by substituting the corresponding collagen domain sequence from Example 1 into the amino acid sequence of V-HC3A.
[0082] Example 3: Construction of recombinant plasmids and recombinant bacteria When synthesizing the nucleotide sequence of a protein single strand (such as the protein single strand in Example 2), the base GC is introduced at the 5' flanking end, and bases GC are introduced at the 5' and 3' ends, respectively. Nco I and Bam HI restriction site. Subsequently, the synthesized genes were inserted into the pColdIII-M plasmid. Nco I and Bam Between HI, the corresponding recombinant collagen protein particles were obtained, wherein the pColdIII-M plasmid was obtained by transferring the pColdIII plasmid. Nde I. Mutation of the enzyme cleavage site Nco I restriction site. Transform the correctly sequenced recombinant plasmids into... E. coli BL21 (DE3) competent cells were spread on LB plates containing ampicillin, cultured and screened, and preserved in glycerol tubes to obtain recombinant bacteria containing recombinant collagen.
[0083] Example 4: Expression, purification, and enzyme digestion optimization of collagen sequences The recombinant bacteria obtained in Example 3 were cultured in shake flasks. After collecting, lysing, and centrifuging the bacterial cells, the supernatant was collected and treated with HisTrap. TM HP 5 mL was used for affinity purification, and samples with imidazole concentrations of 175 mmol / L and 400 mmol / L were collected. SDS-PAGE identification of the samples is shown below. Figure 4 and Figure 5 The theoretical molecular weights of V-HC1-1, V-HC1-2, V-HC1-3, V-HC1-12, V-HC1-22, V-HC1E, and V-HC1F are 25.13 kDa, 24.81 kDa, 28.15 kDa, 34.38 kDa, 34.33 kDa, 25.15 kDa, and 26.09 kDa, respectively. Their apparent molecular weights on SDS-PAGE are approximately 36 kDa, 35 kDa, 40 kDa, 48 kDa, 48 kDa, 37 kDa, and 38 kDa, respectively (e.g., ...). Figure 4 (As shown); the theoretical molecular weights of V-HC2A, V-HC2B, V-HC3A, V-HC3B, and V-HC3C are 34.32 kDa, 37.15 kDa, 34.70 kDa, 34.28 kDa, and 32.61 kDa, respectively, while their apparent molecular weights on SDS-PAGE are approximately 37 kDa, 44 kDa, 43 kDa, 43 kDa, and 38 kDa, respectively (e.g. Figure 5As shown in the figure, the target protein is about 1.4 times the theoretical molecular weight. This may be due to the presence of more proline in the collagen sequence, which causes the target protein to migrate more slowly on SDS-PAGE than proteins of the same molecular weight, consistent with literature reports.
[0084] Removal of the folded domain is a prerequisite for collagen molecules to self-assemble through a transverse, head-to-tail arrangement, ultimately promoting the formation of striations and fibrils. Therefore, during sequence design, a trypsin cleavage site (LVPRGS) sequence is introduced between the collagen domain and the folded domain. Thus, the folded domain can be removed by adding an appropriate amount of trypsin for digestion, yielding a pure collagen domain structure. Under the action of trypsin, the V-domain is digested into multiple short peptides containing 2-20 amino acid residues. If the collagen domain correctly folds into a rigid triple helix structure under the action of the V-domain, it will not be digested by trypsin in a short time.
[0085] V-HC1-2 was selected as the model protein for optimization of trypsin digestion conditions. Results showed that at a molar ratio of 20:1 and digestion for 3 h, the V-domain and other proteins were almost completely digested, with only one band of approximately 25 kDa, corresponding to 1.4 times the molecular weight of the HC1-2 collagen domain. After 12 h, the band became lighter, possibly due to prolonged digestion in a high-concentration trypsin solution, resulting in the cleavage of a small portion of the triple helix. At a molar ratio of 200:1, some incompletely cleaved bands remained at 3 h, disappearing after 6 h, indicating that the V-domain was largely removed, and there was no significant fading of the band within 12 h. At a molar ratio of 2000:1 and digestion for 9 h, the V-domain was not completely cleaved; the bands gradually disappeared around 12 h before digestion. Based on the digestion results, a molar ratio of 200:1 was selected for digestion, with the digestion time controlled between 6 and 12 h.
[0086] Example 5: SDS-PAGE identification and analysis of collagenase digestion Under the optimal enzyme digestion conditions in Example 4, five types of collagen were digested with enzymes. The results showed that collagen V-HC1-1, V-HC1-2, V-HC1-3 and V-HC1-22 were all single bands after trypsin digestion, and the purity reached electrophoretic purity. The apparent molecular weight corresponded to 1.4 times the theoretical molecular weight after enzyme digestion.
[0087] Example 6: Collagen forming a triple helix structure, and its sequence characterization by circular dichroism spectroscopy. To confirm the secondary structure of the collagen domain, the lyophilized collagen sample after enzymatic digestion and desalting in Example 5 was prepared into a 1 mg / mL solution using 10 mmol / L sodium phosphate buffer and equilibrated at 4°C for 48 h. After equilibration, full-wavelength scanning was performed using circular dichroism spectroscopy.
[0088] For the design of type I human collagen, such as Figure 6 As shown in (a), HC1-1, HC1-2, and HC1-3 all exhibit characteristic positive absorption peaks at 225 nm, indicating that all three collagen proteins correctly folded into a triple helix structure with the assistance of the V-domain. And as... Figure 6 As shown, the low-prediction fragments HC1E and HC1F, used as controls, showed no characteristic positive absorption peak at 225 nm, indicating that they could not correctly fold to form a triple-helix structure. Further analysis using circular dichroism spectroscopy to determine the thermal change temperature of the collagen domain revealed the predicted values for HC1-1, HC1-2, and HC1-3. T m The temperatures were 38.4℃, 38.5℃, and 38.2℃ respectively. Figure 6 As shown in (b), the thermochromatographic curves from 4℃ to 70℃ at 225 nm were detected using circular dichroism spectroscopy, and the thermochromatographic curves were fitted (see Table 1); the results showed that HCl-1, HCl-2, and HCl-3... T m The predicted temperatures were 37.2℃, 38.7℃, and 32.4℃, respectively, while the predicted temperatures of the low-prediction fragments HC1E and HC1F, used as controls, were... T m The thermal distortion curves from 4℃ to 70℃ were detected at 225 nm using circular dichroism spectroscopy at 37.1℃ and 36.3℃ respectively. Figure 6 As shown, the thermal transition of HC1E and HC1F could not be measured, indicating that HC1E and HC1F did not fold correctly to form a triple helix structure.
[0089] Furthermore, collagen HC1-12 and HC1-22, which combine fragments 1 and 2, which have higher thermal stability, can also correctly fold into a triple helix structure; among them, the predicted values for HC1-12 and HC1-22 are... T m The concentrations of HCl-12 and HCl-22 were determined by circular dichroism spectroscopy at 38.4℃ and 38.4℃, respectively. T m The values were 33.0℃ and 33.6℃, respectively, indicating that the elongation of the collagen domain leads to a certain degree of decrease in thermal stability. The reason for this may be that the growth of the collagen sequence and the splicing of two sequences make the force of the V-domain-assisted triple helix folding insufficient to be transmitted from the N end to the more distant C end. This results in insufficient rigidity of the triple helix in some regions, making them relatively loose and unfolding faster, thus reducing thermal stability.
[0090] For the design of type II and III human collagen, full-wavelength scanning was performed using circular dichroism spectroscopy to confirm the secondary structure of the collagen domain. Figure 7As shown, HC2A, HC3A, HC3B, and HC3C all exhibited characteristic positive absorption peaks at 225 nm, indicating that all four collagen proteins correctly folded into triple helical structures with the assistance of the V-domain. In contrast, the low-prediction fragment HC2B, used as a control, showed no characteristic positive absorption peak at 225 nm, indicating that it could not correctly fold into a triple helical structure. Further analysis using circular dichroism spectroscopy to determine the thermal change temperature of the collagen domain revealed the predicted values for HC2A, HC3A, HC3B, and HC3C. T m The predicted temperatures for HC2B, respectively, are 38.3℃, 38.8℃, 38.8℃, and 39.0℃, representing the low stability predictions. T m It is 36.5℃. For example... Figure 7 As shown, the thermochromatographic curves from 4℃ to 70℃ at 225 nm were detected using circular dichroism spectroscopy, and the thermochromatographic curves were fitted (see Table 2). The thermochromatographic curves of HC2A, HC3A, HC3B, and HC3C were analyzed. T m The thermal transition of the low-prediction fragment HC2B could not be measured at 28.2℃, 25.1℃, 28.2℃ and 30.3℃, respectively, indicating that HC2B did not fold correctly to form a triple helix structure.
[0091] The above results demonstrate that the prediction method designed in this invention... T m Collagen fragments with high thermal stability at 38-39℃ can all correctly fold into a triple helix structure, while the predicted... T m Fragments with low thermal stability below 38°C could not fold correctly, indicating that by predicting the thermal stability of human collagen through calculation, collagen fragments with different thermal stability can be effectively designed and heterologously expressed in E. coli.
[0092] Table 2 Prediction and Fitting of Human-Derived Collagen T m
[0093] Example 7: Collagen fibers formed by high-polymer self-assembly of collagen (morphological characterization of collagen sequence self-assembly) To observe whether the collagen domains could self-assemble into higher-order structures in high-concentration solutions, the type I collagen HC1-1, HC1-2, HC1-3, HC1-12 and HC1-22 of the freeze-dried sequence from Example 1 were dissolved in 10 mmol / L sodium phosphate buffer to prepare a 0.5 mmol / L solution. After assembly at 4°C for 4 days, negative staining was performed, and then their morphological characteristics were characterized by TEM.
[0094] like Figure 8 and Figure 9 As shown, band-like fibers with periodic bright and dark stripes can be observed in the field of view, similar to the fibrous morphology of natural type I collagen, indicating that the collagen domains of the designed collagen can self-assemble to form biomimetic microfiber structures. The literature reports that the length of each Gly-Pro-Pro triplet is 1.0 nm, and the length of each XYG triplet is 0.9 nm. The lengths of the bright and dark stripes were measured using ImageJ (see Table 3). The results show that the bright stripe lengths of HC1-1, HC1-2, HC1-3, HC1-12, and HC1-22 are approximately 10.6 nm, 10.3 nm, 11.7 nm, 10.2 nm, and 9.9 nm, respectively, which is consistent with (GPP). 10 The theoretical length of the repeating sequence module is 10 nm, corresponding to the dark stripe lengths of approximately 32.2 nm, 32.3 nm, 42.8 nm, 63.8 nm and 64.5 nm, respectively, all consistent with the theoretical values.
[0095] In addition, it can be seen from Figure 8 and Figure 9 It was observed that HC1-22 assembled more ribbon-like fibers than HC1-12 in the field of view, indicating that HC1-22 had a better self-assembly effect than HC1-12. This may be because the thermal stability of HC1-2 is higher than that of HC1-1, affecting the assembly effect. At the same time, the results also showed that HC1-12 and HC1-22 had fewer observable ribbon-like fibers, and their self-assembly effect was not as good as that of the shorter HC1-1 and HC1-2, both in terms of fiber length and fiber aggregation morphology.
[0096] Table 3. Statistics on the bandwidth of collagen fibers
[0097] Example 8: Products containing collagen A product containing collagen can be found in the beauty, chemical, medical / biomedical, cosmetic, and animal feed industries; for example, beauty cosmetics (masks, serums, creams, etc.), artificial collagen casings, nutritional supplements (collagen powder, oral liquids), medical dressings, hemostatic materials, artificial bone scaffolds, injectable fillers, artificial blood vessels, eye drops, drug sustained-release carriers, etc.
[0098] In the product containing collagen, the collagen has the collagen domain sequence of Example 1 of the present invention, or has the collagen sequence prepared in Example 2.
[0099] Furthermore, the collagen is collagen expressing a triple helix structure.
[0100] Furthermore, the collagen is type I, type II, or type III collagen.
[0101] Furthermore, in the aforementioned collagen-containing products, other components, formulations, and preparation processes can be implemented by any existing method by those skilled in the art.
[0102] The sequence involved in this invention: SEQ ID NO.1: Amino acid sequence of HC1-1 GARGLPGTAGLPGMKGHRGFPGERGLDGAKGDAGPAGPKGEPGSPGENGAPGQMGPRGPQGPPGPPGPKGNSGEPGAPGSKGDTGAKGEPGPVGVQGPPGPAGEEGKR SEQ ID NO.2: Amino acid sequence of HC1-2 GFPGERGVQGPPGPAGPRGANGAPGNDGAKGDAGAPGAPGSQGAPGLQGMPGERGAAGLPPGPKGDRGDAGPKGADGSPGKDGVRGLTGPIGPPGPAGAPGDKGESGPS SEQ ID NO.3: Amino acid sequence of HC1-3 GPAGFAGPPGADGQPGAKGEPGDAGAKGDAGPPGPAGPAGPPGPIGESGREGAPGAEGSPGRDGSPGAKGDRGETGPAGPPGFPGERGAPGPAGPAGPVGPVGARGPAGPQGPRGDKGETGEQGDRGIKGHRGFSGLQ SEQ ID NO.4: Amino acid sequence of HC2A GLTGPAGEPGREGSPGADGPPGRDGAAGVKGDRGETGAVGAPGAPGPPGDRGEAGAQGPMGPSGPAGARGIQGPQGPRGDKGEAGEPGERGLKGHRGFTGLQGLPGPPGPS SEQ ID NO.5: Amino acid sequence of HC3A GFPGMKGHRGFDGRNGEKGETGAPGLKGENGLPGENGAPGPMGPRGAPGERGSPGPKGDKGEPGPPGADGVPGKDGPRGPTGPIGPPGPAGQPGDKGEP SEQ ID NO.6: Amino acid sequence of HC3B GFPGMKGHRGFDGRNGEKGETGAPGLKGENGLPGENGAPGPMGPRGAPGERGAKGEPGPRGERGEAGIPGVPGAKGEDGKPGEPGPKGDAGAPGAPGPKGDAGAPGER SEQ ID NO.7: Amino acid sequence of HC3C GFPGMKGHRGFDGRNGEKGETGAPGLKGENGLPGENGAPGPMGPRGAPGERGAKGEPGPRGERGEAGIPGVPGAKGEDGRDGNPGSDGLPGRDGSPGPKGDRGENGSP SEQ ID NO.8: Amino acid sequence of HC1-12 GARGLPGTAGLPGMKGHRGFPGERGLDGAKGDAGPAGPKGEPGSPGENGAPGQMGPRGPQGPPGPPGPKGNSGEPGAPGSKGDTGAKGEPGPVGVQGPPGPAGEEGKRGFPGERGVQGPPGPAGPRGANGAPGNDGAKGDAGAPGAPGSQGAPGLQGMPGERGAAGLPGPKGDRGDAGPKGADGSPGKDGVRGLTGPIGPPGPAGAPGDKGESGPS SEQ ID NO.9: Amino acid sequence of HC1-22 GFPGERGVQGPPGPAGPRGANGAPGNDGAKGDAGAPGAPGSQGAPGLQGMPGERGAAGLPGPKGDRGDAGPKGADGSPGKDGVRGLTGPIGPPGPAGAPGDKGESGPSGFPGERGVQGPPGPAGPRGANGAPGNDGAKGDAGAPGAPGSQGAPGLQGMPGERGAAGLPGPKGDRGDAGPKGADGSPGKDGVRGLTGPIGPPGPAGAPGDKGESGPS SEQ ID NO.10: Amino acid sequence of HC1E GPMGPSGPRGLPGPPGAPGPQGFQGPPGEPGEPGASGPMGPRGPPGPPGKNGDDGEAGKPGRPGERGPPGPQGARGLPGTAGLPGMKGHRGFSGLDGAKGDAGPAGPK SEQ ID NO.11: Amino acid sequence of HC1F GPRGLPGPPGAPGPQGFQGPPGEPGEPGASGPMGPRGPPGPPGKNGDDGEAGKPGRPGERGPPGPQGARGLPGTAGLPGMKGPAGSPGFQGLPPGPPGEAGKPGEQGVPGDLGAPGPS SEQ ID NO.12: Amino acid sequence of HC2B GANGDPGRPGEPGLPGARGLTGRPGDAGPQGKVGPSGAPGEDGRPGPPGPQGARGQPGVMGFPGPKGANGEPGKAGEKGLPGAPGLRGLPGKDGETGAAGERGSPGAQGLQGPRGLPGTPGTDGPK SEQ ID NO.13: Amino acid sequence of V-domain ADEQEEKAKVRTELIQELAQGLGGIEKKNFPTLGDEDLDHTYMTKLLTYLQEREQAENSWRKRLLKGIQDHALD SEQ ID NO.14: Amino acid sequence of the coiled-coil domain GEIAAIKQEIAAIKKEIAAIKWEIAAIKQGYG SEQ ID NO.15: Amino acid sequence of V-HC1-1 HHHHHHADEQEEKAKVRTELIQELAQGLGGIEKKNFPTLGDEDLDHTYMTKLLTYLQEREQAENSWRKRLLKGIQDHALDLVPRGSPGPPGPPGPPGPPGPPGPPGPPGPPGPPGPPGARGLPGTAGL PGMKGHRGFPGERGLDGAKGDAGPAGPKGEPGSPGENGAPGQMGPRGPQGPPGPPGPKGNSGEPGAPGSKGDTGAKGEPGPVGVQGPPGPAGEEGKRGPPGPPGPPGPPGPPGPPGPPGPPGPPGPPG SEQ ID NO.16: Nucleotide sequence of V-HC1-1 CACCATCACCATCACCACGCCGACGAGCAAGAAGAAAAGGCCAAAGTTCGCACCGAGCTGATTCAAGAACTGGCGCAAGGTCTGGGCGGCATCGAAAAGAAAAACTTCCCGACGCTGGGCGATGAAGATCTGGACCACACCTACATGACGAAGCTGCTGACCTATCTGCAAGAACGTGAACAAGCCGAGAATAGCTGGCGCAAACGTCTGCTGAAAGGCATCCAAGATCATGCGCTGGATCTGGTGCCACGTGGCAGCCCGGGCCCGCCGGGCCCGCCGGGCCCACCGGGTCCACCGGGCCCGCCGGGCCCACCGGGTCCGCCGGGTCCGCCGGGTCCGCCGGGCCCACCGGGCGCCCGTGGTCTGCCGGGCACCGCCGGTCTGCCGGGCATGAAAGGCCATCGCGGTTTCCCGGGTGAACGTGGTCTGGATGGCGCCAAAGGTGATGCGGGTCCAGCCGGTCCGAAAGGCGAACCGGGCAGCCCGGGCGAAAATGGTGCGCCGGGCCAGATGGGTCCGCGTGGTCCACAAGGCCCGCCGGGCCCACCGGGCCCGAAAGGCAATAGCGGTGAACCGGGCGCCCCGGGCAGTAAAGGCGATACCGGTGCGAAAGGTGAACCGGGCCCGGTTGGTGTTCAAGGCCCACCGGGCCCAGCGGGTGAAGAAGGTAAACGTGGTCCGCCGGGTCCACCGGGTCCACCGGGTCCACCGGGCCCACCGGGCCCGCCGGGCCCACCGGGTCCGCCGGGCCCGCCGGGCCCACCGGGCTAA SEQ ID NO.17: Amino acid sequence of V-HC2A HHHHHHADEQEEKAKVRTELIQELAQGLGGIEKKNFPTLGDEDLDHTYMTKLLTYLQEREQAENSWRKRLLKGIQDHALDLVPRGSPGPPGPPGPPGPPGPPGLTGPAGEPGREGSPGADGPPGRDGAAGVKGDRGETGAVGAPGAPGPPGDRGEAGAQGPMGPSGPAGARGIQGPQGPRGDKGEAGEPGERGLKGHRGFTGLQGLPGPPGPSGPPGPPGPPGPPGLTGPAGEPGREGSPGADGPPGRDGAAGVKGDRGETGAVGAPGAPGPPGDRGEAGAQGPMGPSGPAGARGIQGPQGPRGDKGEAGEPGERGLKGHRGFTGLQGLPGPPGPSGPPGPPGPPGPPGPPGPPG SEQ ID NO.18: Nucleotide sequence of V-HC2A SEQ ID NO.19: Amino acid sequence of V-HC3A HHHHHHADEQEEKAKVRTELIQELAQGLGGIEKKNFPTLGDEDLDHTYMTKLLTYLQEREQAENSWRKRLLKGIQDHALDLVPRGSPGPPGPPGPPGPPGPPGPPGFPGMKGHRGFDGRNGEKGETGAPGLKGENGLPGENGAPGPMGPRGAPGERGSPGPKGDKGEPGPPGADGVPGKDGPRGPTGPIGPPGPAGQPGDKGEPGPPGPPGPPGPPGFPGMKGHRGFDGRNGEKGETGAPGLKGENGLPGENGAPGPMGPRGAPGERGSPGPKGDKGEPGPPGADGVPGKDGPRGPTGPIGPPGPAGQPGDKGEPGPPGPPGPPGPPGPPGPPG SEQ ID NO.20: Nucleotide sequence of V-HC3A CATCACCATCACCATCATGCGGATGAACAAGAAGAAAAAGCGAAAGTGCGCACCGAACTGATTCAAGAACTGGCGCAAGGCCTGGGCGGCATTGAAAAAAAAAACTTTCCGACCCTGGGCGATGAAGATCTGGATCATACCTATATGACCAAACTGCTGACCTATCTGCAAGAACGCGAACAAGCGGAAAACAGCTGGCGCAAACGCCTGCTGAAAGGCATTCAAGATCATGCCCTGGATTTAGTGCCGCGCGGCAGCCCGGGTCCACCGGGTCCGCCGGGCCCGCCGGGCCCACCGGGTCCGCCGGGCTTTCCGGGCATGAAGGGCCATCGCGGTTTTGATGGCCGCAACGGCGAAAAAGGCGAAACGGGTGCCCCGGGCCTGAAAGGCGAAAACGGTTTACCGGGCGAGAACGGCGCGCCGGGCCCGATGGGTCCGCGTGGTGCGCCGGGCGAACGCGGCAGCCCGGGCCCAAAAGGTGATAAGGGTGAACCGGGTCCGCCGGGCGCCGACGGTGTGCCGGGCAAAGATGGCCCGCGCGGCCCGACGGGCCCGATTGGCCCGCCGGGCCCGGCGGGCCAACCGGGCGACAAAGGTGAACCGGGCCCGCCGGGCCCGCCGGGCCCACCGGGTCCACCGGGTTTTCCGGGCATGAAGGGCCATCGCGGCTTTGATGGTCGTAACGGCGAGAAGGGCGAAACCGGTGCGCCGGGCTTAAAAGGTGAAAACGGCCTGCCGGGCGAGAACGGCGCGCCGGGTCCGATGGGCCCACGTGGCGCCCCGGGCGAGCGCGGCAGTCCGGGCCCGAAGGGCGATAAAGGCGAACCGGGCCCGCCGGGCGCGGATGGCGTGCCGGGCAAAGATGGCCCACGCGGTCCAACGGGTCCGATCGGCCCGCCGGGCCCGGCGGGTCAGCCGGGCGATAAGGGTGAGCCGGGCCCGCCGGGCCCGCCGGGCCCGCCGGGCCCGCCGGGCCCACCGGGCCCACCGGGTTAA SEQ ID NO.21: LVPRGSP SEQ ID NO.22: LVPRGS SEQ ID NO.23: GPPGPPGPPGPPGPPGPPGPPGPPGPPGPP SEQ ID NO.24: GPPPGPPGPPGPP SEQ ID NO.25: GPPGPPGPPGPP SEQ ID NO.26:GPPGPPGPPGPPGPPGPP Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Anyone skilled in the art can make various modifications and alterations without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the claims.
Claims
1. A protein fragment, characterized in that, The amino acid sequence of the protein fragment is shown in any one of SEQ ID NO.1 to 9.
2. A single-chain protein, characterized in that, The structure of a single protein chain from the N-terminus to the C-terminus includes: {repetitive sequence module, collagen domain}m, repetitive sequence module; where m is 1 or 2; The amino acid sequence of the collagen domain is shown in any one of SEQ ID NO.1~9; The repeat sequence module uses (GPP). n Among them, (GPP) n The value of n in the equation satisfies 5 < n ≤ 30.
3. The protein single chain according to claim 2, characterized in that, The repeat sequence module is preceded by a folding domain, which is introduced to assist collagen in folding to form a triple helix structure.
4. The protein single chain according to claim 3, characterized in that, The folded domain is either a V-domain or a coiled-coil domain; the amino acid sequence of the V-domain is shown in SEQ ID NO.13, and the amino acid sequence of the coiled-coil domain is shown in SEQ ID NO.
14.
5. The protein single chain according to claim 2, characterized in that, There are multiple repetitive sequence modules, located at both ends of the collagen domain or at both ends of multiple collagen domains.
6. The protein single chain according to claim 3, characterized in that, Folded domains and repetitive sequence modules are connected through restriction enzyme sites.
7. The protein single chain according to claim 3, characterized in that, The folded domain also has a histidine tag at its front end.
8. A gene encoding the protein fragment of claim 1.
9. A gene encoding a single strand of the protein described in any one of claims 2-7.
10. A plasmid expressing the gene of the protein fragment of claim 1 or the gene of the single strand of any of the proteins of claims 2-7.
11. The plasmid according to claim 10, characterized in that, The plasmids mentioned are pColdIII series or pET series plasmids.
12. A cell expressing a gene of the protein fragment of claim 1 or a gene of a single strand of any of the proteins of claims 2-7.
13. The cell according to claim 12, characterized in that, The cells mentioned are Escherichia coli cells, including E. coli BL21 E . coli BL21(DE3), E. coli Rosetta (DE3) E. coli BL21(DE3) pLysS / pLysE or E. coli Origami2(DE3).
14. A collagen protein comprising a triple helix structure formed by the single protein chains of any one of claims 2-7 coiled around a common central axis.
15. Collagen fibers formed by the self-assembly of collagen polymers as described in claim 14.
16. A product containing the collagen of claim 14, characterized in that, The products mentioned are products in the fields of chemicals, medical / biomedicine, or cosmetics.
17. The product according to claim 16, characterized in that, Products in the cosmetics field include face masks, serums, or creams.
18. The product according to claim 16, characterized in that, Products in the medical / biomedical field include medical dressings, hemostatic materials, artificial bone scaffolds, injectable fillers, artificial blood vessels, eye drops, or drug sustained-release carriers.