Reconstructed silk protein and preparation method thereof
The neural network model was constructed through material genomics methods, screening and preparing target silk proteins, solving the cost and time-consuming problems in traditional methods, and achieving efficient preparation of functional silk protein materials.
Patent Information
- Application Number
- CN202210450251.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-04-26
AI Technical Summary
Traditional methods are costly and time-consuming when studying the properties of silk protein materials, making it difficult to efficiently screen out silk protein materials with specific properties.
Using material genomics methods, we can establish a silk protein database, build a neural network model, use molecular descriptors and molecular dynamics simulation to screen out the silk protein structure of the target material performance, and prepare the target silk protein through genetic engineering or chemical synthesis.
It greatly shortens the research time and cost of silk protein materials, and can quickly obtain functional silk protein materials with different morphology and performance.
Smart Images

Figure CN115677842B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a reconstructed silk protein and a preparation method thereof, and belongs to the fields of genetic engineering and material genomics. Background Art
[0002] By entering all known material properties into a database and then extracting and integrating the molecular descriptors of the materials through machine learning algorithms, people can simulate the physical and chemical properties of various material combinations on a computer. This is undoubtedly much faster than conducting experimental tests on each sample one by one.
[0003] Providing a new research method for the design of new materials is of great significance for the more efficient and economical development of advanced materials. The material genome method based on calculation, experiment and database search can accelerate the development of advanced materials.
[0004] Materials science essentially studies structure-activity relationships—the relationship between a substance's composition, structure, and properties. The materials genomics approach doesn't break the boundaries of materials science; rather, it uses a cutting-edge system to accelerate and transform materials research. It abandons the traditional slow and inefficient trial-and-error approach and integrates high-throughput materials experimental methods, computational materials theory and algorithms, and materials data mining to create a new approach to materials research.
[0005] Materials data mining is actually Materials Informatics, which uses machine learning and data mining methods to process and analyze material data and predict and design new materials. On the other hand, it also includes the establishment and updating of materials science databases.
[0006] Materials data mining is a core component of materials genome engineering. Using data mining to construct predictive models, we can generate a large number of predictive models for dependent variables (material properties) determined by independent variables (controllable material parameters such as primary structure, physical parameters, aggregate structure, solvent system, etc.). Based on these models, we can predict the properties of candidate materials and screen out the best performing candidates, often finding high-performance materials that meet our requirements.
[0007] Silk, a natural polymer material, boasts excellent biocompatibility and controllable biodegradability, making it a research hotspot in disciplines such as biomedicine and materials science. However, traditional research methods based on scientific experience and trial-and-error experiments are often constrained by high costs and time-consuming processes in the pursuit of improving the properties of silk materials. Summary of the Invention
[0008] In response to the above technical problems, the present invention proposes a research method based on material genomics, which can effectively break through the limitations of traditional methods. To this end, the present invention provides a method for constructing silk protein with target material properties using material genomics, which comprises the following steps:
[0009] S1. Establish a silk protein database;
[0010] S2. Build the model;
[0011] S3. Analyze and retrieve structures related to the material properties of silk protein, and determine the silk protein structure corresponding to the target material properties;
[0012] S4. Validation and screening of silk protein candidates through computational and materials genomic analysis;
[0013] S5. Obtaining target silk protein;
[0014] S6. Verify the target material properties of silk protein;
[0015] S7. Input the material properties and molecular descriptors of the silk protein obtained in step S5 into the neural network obtained in step S2, and optimize and train the neural network in step S2 based on the test data.
[0016] In a preferred embodiment of the present invention, step S1 is performed as follows:
[0017] By defining the characteristic expressions of silk proteins, including SDF, MOL, SMILES codes or images, relevant silk protein characteristic expressions and experimental data were extracted from different patents and literatures, and structured extraction and analysis were performed. All amino acid sequences and structures with specific functional characteristics were searched in the SciFinder, ChemSpider and PubChem databases, and batch downloads were performed through the database's application program interface.
[0018] In a preferred embodiment of the present invention, step S2 is performed as follows:
[0019] a) obtaining training data of multiple molecular descriptors of silk protein and corresponding silk protein material properties;
[0020] b) obtaining a neural network that characterizes the relationship between the molecular descriptors of the silk protein material and the properties of the silk protein material based on the training data;
[0021] c) obtaining test data, testing the neural network according to the test data, and if the test result does not meet the error requirement, returning to step b) until the test result meets the error requirement.
[0022] In a preferred embodiment of the present invention, step S3 is performed as follows:
[0023] a) Select target material properties that match the actual application scenario;
[0024] b) inputting the target material properties into the neural network obtained in step S2 to find a series of molecular descriptors that meet the target material properties;
[0025] c) Based on the obtained molecular descriptors, several candidate protein structures are retrieved and screened from the database in step S1.
[0026] In a preferred embodiment of the present invention, step S4 is performed as follows:
[0027] a) Calculate the properties of candidate materials using molecular dynamics simulations;
[0028] b) Evaluate and validate candidate materials using materials genome analysis.
[0029] The functionally diverse silk proteins constructed by the above method are Figure 1 shown.
[0030] By comprehensively comparing and analyzing the structures and properties of various animal silks and silk proteins, the inventors discovered that for materials based on the primary structural biological macromolecules of silk protein, the molecular chain aggregation structure, which is jointly influenced by the primary structure of the protein, the formation process, and various exogenous factors, will determine the final properties of the material.
[0031] 1. Primary structure
[0032] The primary structure or specific amino acid residues of silk protein are associated with the mechanical properties of natural silk fibers. For example, the proline (Pro) content is negatively correlated with the modulus and breaking stress of silk; for larger reconstructed proteins with n=10, polyalanine segments can form enough helical regions and aggregate with each other, making the reconstructed protein water-insoluble under certain conditions; because glutamic acid has a side carboxyl group, it can help regulate the water solubility of the material (helping to shape) and chelate Ca 2+ (helps mineralization); the RGD (R, G, and D are arginine, glycine, and aspartic acid, respectively) sequence has animal cell adhesion function; TGRGDSPA can increase cell adhesion while promoting the osteogenic expression of osteoblast-like cells bound thereto.
[0033] 2. Chain structure
[0034] Figure 2The relationship between the primary and secondary structures of mulberry silk protein is given. The entire silk protein chain consists of 12 highly repetitive crystalline regions (R01 to R12), 11 non-repetitive random sequences (A01 to A11) interspersed between them, and two termini (N-terminus and C-terminus).
[0035] In the crystalline regions, highly repetitive hexapeptide sequences (GAGAGS) form the building blocks of the β-pleated microcrystalline regions of silk protein. However, fragments of GAGAGY and GAGAGVGY containing tyrosine and with slightly less repetitive sequences may form less regular semi-crystalline or amorphous regions. Furthermore, a small amount of GAAS forms a tetrapeptide β-turn. Within the disordered regions, a high concentration of amino acid residues with hydrophilic side chains (such as serine and glutamic acid) makes the connecting portion much less hydrophobic than the crystalline regions, forming a ring structure. This ring structure is believed to reverse the direction of silk protein chains by 180°, promoting the formation of an antiparallel β-pleated structure.
[0036] 3. Crystallinity
[0037] The crystallinity of protein materials also shows a certain correlation with water solubility and deceleration rate.
[0038] The crystallinity of mulberry silk is generally around 50%, which is because the latter's silk protein contains more and more regular crystallizable GAGAGS and GAGAGAGS segments.
[0039] Practice has also demonstrated that even a small amount of β-sheets (less than 5%) can render silk protein insoluble in water. Because regular β-sheets (crystalline regions) are highly hydrophobic, silk fibers with the highest degree of crystallinity are generally non-biodegradable. By controlling the β-sheet content in a protein material, the degradation rate of the material can be adjusted, significantly facilitating the application of silk protein scaffolds in various tissue engineering applications.
[0040] 4. Exogenous factors
[0041] Through a systematic and comprehensive investigation of exogenous factors, we can determine the different roles played by various exogenous factors, such as solvents, ions, pH and stress, in the transformation of silk protein conformation from random coils to β-folds (i.e., the essential mechanism of silk protein filamentation).
[0042] Therefore, in a further preferred embodiment of the present invention, the molecular descriptors include: the primary structure, aggregated structure, molecular structure and exogenous factors of silk protein.
[0043] In a more preferred embodiment of the present invention, the primary structure molecular descriptors include: amino acid types and contents, amino acid sequences and contents corresponding to the special aggregation structures forming α-helix, β-sheet, β-turn, and random-coil; the aggregation structure molecular descriptors include: α-helix, β-sheet, β-turn, and random-coil; the exogenous factor molecular descriptors include solvents, ions, pH values, and stress.
[0044] In a preferred embodiment of the present invention, the material properties of the silk protein include: hydrophilicity and hydrophobicity, mechanical strength, water absorption, swelling, antibacterial properties, self-assembly properties, fiber / film-forming behavior, aggregate structure, cell adhesion and responsiveness to enzymes.
[0045] In a more preferred embodiment of the present invention, the target material properties of the silk protein are hydrophilicity, self-assembly property, or film-forming and repairing property.
[0046] In the research and development of silk protein materials, the functional properties of the desired protein can be determined by reverse engineering, such as the corresponding amino acid sequence and molecular weight information. This information can then be expressed using genetic recombination techniques, and then produced through molecular self-assembly techniques to create the target peptide chains and proteins. Furthermore, the assembly behavior of molecules such as DNA, proteins, peptides, and polysaccharides can be exploited to create applicable functional materials. With the continuous advancement of nanoscience, it is possible to design peptide molecules at the molecular level and manipulate the structure and shape of their aggregates to form nanomaterials with specific structures and functions.
[0047] Therefore, in a preferred embodiment of the present invention, step S5 uses a genetic engineering method to construct the target silk protein. In addition, step S5 of the present invention can also use chemical synthesis or genetic biology methods to construct the target silk protein.
[0048] In a more preferred embodiment of the present invention, the silk protein structure determined in step S3 is as shown in Sequence 1, Sequence 8, Sequence 12 or Sequence 16.
[0049] Another aspect of the present invention provides a silk protein having target material properties constructed according to the method of the present invention.
[0050] In a preferred embodiment of the present invention, the target material properties are hydrophilicity, self-assembly properties, or film-forming and repairing properties.
[0051] In another aspect, the present invention provides a silk protein with improved target material properties and a silk protein material prepared therefrom.
[0052] In a preferred embodiment of the present invention, the target material properties are hydrophilicity, self-assembly properties, or film-forming and repairing properties.
[0053] In a preferred embodiment of the present invention, the silk protein comprises mulberry silk protein, spider silk protein, or tussah silk protein.
[0054] Beneficial effects
[0055] The present invention uses materials genomics as a research method to screen, analyze and study the material properties of silk protein. Compared with traditional research methods based on scientific experience and trial and error experiments, it saves research costs, shortens research time, and greatly accelerates the preparation and development of silk protein materials.
[0056] Applying the method for constructing silk protein of the present invention to various aspects of silk protein material research can make the functionalization of silk protein more convenient and diversified, and can obtain silk protein materials with different morphologies and properties. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a schematic diagram of the purposeful construction of functionalized diversified proteins;
[0058] Figure 2 This is a schematic diagram of the relationship between the primary structure and secondary structure of mulberry silk protein;
[0059] Figure 3 It is a schematic diagram of the amino acid sequence of part of silk protein;
[0060] Figure 4 Schematic diagram of the relationship between the molecular descriptors used in the present invention and the properties of silk protein materials;
[0061] Figure 5 Schematic diagram of encoding molecular descriptors characterizing silk protein materials into neural networks;
[0062] Figure 6 This is a schematic diagram of a machine learning model for the structure-activity relationship of materials;
[0063] Figure 7 This is a comparison chart showing the improvement in the hydrophilic properties of the hydrophilic reconstructed silk protein membrane prepared in Example 1 of the present invention;
[0064] Figure 8 is a transmission electron microscope image of the sample prepared in Example 2 of the present invention;
[0065] Figure 9 is a transmission electron microscope image of the sample prepared in Example 3 of the present invention;
[0066] Figure 10 This is a diagram showing the film-forming and repairing properties of the sample prepared in Example 4 of the present invention. DETAILED DESCRIPTION
[0067] In order to better illustrate the purpose and advantages of the present method, the specific implementation content of the present invention is further described in detail with reference to the accompanying drawings and specific embodiments.
[0068] Example 1: Preparation of super hydrophilic silk protein film
[0069] Naturally extracted silk protein chains primarily consist of highly repetitive GAGAGS and GAGAGAGS (G, glycine; A, alanine; S, serine) sequences, which readily form hydrophobic β-pleated regions, serving as physical crosslinks and rendering silk fibroin-based materials water-insoluble. This inherent hydrophobicity significantly impacts cell adhesion to silk materials. Therefore, it is desirable to improve the hydrophilicity of these materials' surfaces and thereby enhance cell adhesion.
[0070] Step S1: Database establishment
[0071] By defining characteristic expressions for silk proteins (SDF / MOL / SMILES codes / images), we extracted relevant silk protein characteristic expressions and experimental data from various patents and literature, and used natural language processing techniques to perform structured extraction and analysis of the structural and performance data. We searched the SciFinder, ChemSpider, and PubChem databases for all amino acid sequences and structures with specific functional characteristics, and wrote automated scripts for batch downloading through the databases' application programming interfaces.
[0072] Figure 3 The amino acid sequence of part of the silk protein is shown, corresponding to the primary structure of natural silk protein.
[0073] Step S2. Build model
[0074] a) obtaining training data of multiple molecular descriptors of silk protein and corresponding silk protein material properties;
[0075] b) obtaining a neural network that characterizes the relationship between the molecular descriptors of the silk protein material and the properties of the silk protein material based on the training data;
[0076] c) obtaining test data, testing the neural network according to the test data, and if the test result does not meet the error requirement, returning to step b until the test result meets the error requirement;
[0077] like Figure 4 As shown, molecular descriptors can include the primary structure, aggregation structure, molecular structure and exogenous factors of silk protein.
[0078] The primary structure molecular descriptors of silk protein include: amino acid types and content, amino acid sequences and contents corresponding to special aggregation structures such as α-helix, β-sheet, β-turn, random-coil, and other primary structures.
[0079] The molecular descriptors of the aggregated structure of silk protein include: α-helix, β-sheet, β-turn, random-coil and other aggregated structures.
[0080] The molecular descriptors of exogenous factors that form silk protein materials include solvents, ions, pH, stress, etc.
[0081] like Figure 5 The molecular descriptors characterizing the silk protein material are encoded into the neural network as shown in the following example, and then used as Figure 6 The neural network shown performs machine learning and builds a model.
[0082] Step S3. Analyze and retrieve the structure related to the material properties of silk protein, and determine the silk protein structure corresponding to the target material properties
[0083] Based on the key feature of the material performance, hydrophilicity, reverse analysis was performed, and the chemical structure of the corresponding material feature was retrieved from the above database, which was the GTGSSGFGPYVANGGYSRSDGYEYAWSSDFGT sequence (sequence 1).
[0084] Step S4: Verify the performance of candidate materials and screen them through molecular dynamics simulation and material genome analysis
[0085] a) Calculate the properties of candidate materials using molecular dynamics simulations
[0086] Use molecular modeling tools to build a chemical molecular model corresponding to the material sequence. Use a universal force field to establish a set of appropriate force field parameters for the model.
[0087] Select appropriate molecular dynamics simulation parameters and simulate the constructed molecular model. Based on the resulting molecular dynamics trajectory, calculate the material properties of the molecular model. For more accurate results, run multiple trajectories and calculate the statistical average to eliminate errors.
[0088] The retrieved materials are verified based on the calculated material performance parameters.
[0089] b) Evaluate and validate candidate materials using materials genome analysis
[0090] The structure of the candidate material is evaluated by material genome analysis technology. A sequence set with target performance is screened from the silk protein database obtained in step S1. Based on the Tanimoto similarity algorithm, the structural similarity between the screened sequence set and the candidate material is calculated. A similarity threshold S is set. threshold , only when the average similarity between the candidate material and the sequence set is greater than the threshold S threshold , you can pass the verification and proceed to the next step.
[0091] Step S5: Obtaining recombinant target protein through genetic engineering methods
[0092] a) Obtain target gene fragments by chemical synthesis
[0093] The target fragment (5'-3' single strand) was obtained by chemical synthesis. The restriction sites BamHI and NcoⅠ, as well as protective bases were added to both ends of the target fragment to synthesize the positive strand and negative strand respectively.
[0094] Take the target protein Pro1 as an example, where:
[0095] Protein sequence (SEQ ID NO: 2): (GAGAGSGAGAGSGAGAGSGAGAGSGAGAGS) n ;
[0096] Corresponding gene sequence (sequence 3):
[0097] (ggtgctggtgctggttcaggtgctggtgctggttcaggtgctggtgctggttcaggtgccggtgctggttcaggtgctggtgctggttca) n ;
[0098] By chemical synthesis, n target fragments were obtained respectively:
[0099] Positive chain (sequence 4):
[0100] GATCCggtgctggtgctggttcaggtgctggtgctggttcaggtgctggtgctggttcaggtgccggtgctggttcaggtgctggtgctggttcaTGAC;
[0101] Negative strand (sequence 5):
[0102] CATGGTCAtgaaccagcaccagcacctgaaccagcaccggcacctgaaccagcaccagcacctgaaccagcaccagcacctgaaccagcaccagcaccG;
[0103] Then, a temperature gradient annealing was used to form a double strand, and the PNK enzyme was used to add a phosphate group to the 5' end;
[0104] The vector was linearized by PCR using a pET vector containing the Amp ampicillin resistance gene with a GST-6×His-TEV tag;
[0105] The forward and reverse primers contain corresponding restriction sites and protection bases, and the annealing temperature is guaranteed to be 54-60°C, for example:
[0106] Forward primer (SEQ ID NO: 6):
[0107] GAAAACCTGTACTTCCAGTCAGGATCCGCG;
[0108] Reverse primer (SEQ ID NO: 7):
[0109] GCGCCATGGTCTGACACATGCAGCTCCCG;
[0110] After verifying that the target fragment is correct, use a dedicated kit for cleaning;
[0111] Use T4 ligase to connect the linearized vector and the target fragment. Since the target gene fragment is too long and is a repeating unit, a step-by-step gene construction method is adopted to connect the gene fragment to the linearized vector n times;
[0112] b) Use the recombinant vector pET-GST-6×His-TEV-Pro1 to transform competent cells
[0113] The recombinant vector was transformed into competent cells DH5α (clone), ice-bathed for 30 minutes, heat-shocked at 42°C for 1 minute, ice-bathed for 3 minutes, plated (LB solid medium with 100 μg / ml Amp), and cultured overnight at 37°C; single colonies were picked and placed in LB liquid medium with corresponding antibiotics, shaken and cultured at 37°C, 220 rpm, and a portion was sequenced after the bacterial solution became turbid.
[0114] c) Plasmid transformation into competent cells
[0115] After the sequencing results were correct, the plasmid was extracted using a kit and transformed into competent cells BL21 (expression), ice-bathed for 30 min, heat-shocked at 42°C for 1 min, ice-bathed for 3 min, plated (LB solid medium with corresponding antibiotics), and cultured overnight at 37°C.
[0116] d) Fermentation culture of genetically engineered bacteria
[0117] Pick a single colony into 3 ml of culture medium (containing antibiotics), expand to 1 L of culture medium (containing antibiotics), and culture until OD 600 The value was between 0.5 and 0.8, and then the inducer IPTG at a concentration of 0.1-2 mM was added to induce protein expression. The culture was continued at 37°C, 220 rpm for 4 hours (or 16°C, 220 rpm for overnight induction). The cells were collected by centrifugation at 8000Xg and 4°C for 30 minutes and stored at -80°C.
[0118] e) Protein purification
[0119] Taking His tag as an example, the target protein contains 6×His and is purified using a Ni column:
[0120] (1) Thaw the cells at -80℃;
[0121] (2) Resuspend in buffer (20 g - 150 ml): 30 ml Ecoli buffer + 300 ml PI + 60 μl 4 M imidazole;
[0122] (3) ultrasonic disruption, 3s for 3s 50 times, 2 cycles;
[0123] (4) Centrifugation at 40,000 × g for 30 min;
[0124] (5) Incubate with Ni beads for 1.5-2 hours;
[0125] (6) Flow though;
[0126] (7) Wash, 40 ml Wash (+50 mM imidazole) + 500 μL 4 M imidazole;
[0127] (8) Elution, 400 mM imidazole, 1 h;
[0128] (9) Collecting and obtaining the target protein;
[0129] (10) Gel running / Western Blot verification.
[0130] Step S6: Performance verification of target protein material
[0131] a) The reconstructed silk protein aqueous solution was transferred into a polystyrene mold and air-dried in a constant temperature and humidity chamber (25° C., 30±5% RH) to obtain a protein film with a thickness of approximately 0.2 cm.
[0132] b) Contact angle measurements were performed using a Dataphysics OCA 40 instrument at 25°C using water (surface tension γ = 72.2 mN / m) as the probe liquid. The static contact angles of water on both naturally extracted silk fibroin membranes and hydrophilic reconstructed silk fibroin membranes were measured.
[0133] The results showed that the contact angle of water on the surface of natural extracted silk protein membrane was 80±℃, while the contact angle on the surface of hydrophilic reconstructed silk protein membrane was 29±℃. Figure 7 This result shows that compared with the naturally extracted silk protein membrane, the hydrophilic reconstructed silk protein membrane prepared by the present invention has improved hydrophilicity and significantly improves the adhesion performance of cells on the surface of the silk protein membrane.
[0134] Step S7. Optimize and train the neural network
[0135] The hydrophilic properties of the silk protein obtained in step S5 and the molecular descriptors shown in sequence 1 are encoded and input into the neural network obtained in step S2, and the neural network in step S2 is optimized and trained based on the above data.
[0136] Example 2: Preparation of amyloid-like fibrous protein 1
[0137] Step S1: Database establishment
[0138] The specific method is the same as described in step S1 of Example 1.
[0139] Step S2: Building the model
[0140] The specific method is the same as described in step S2 of Example 1.
[0141] Step S3: Analyze and retrieve the structure related to the material properties of silk protein, and determine the silk protein structure corresponding to the target material properties
[0142] Based on the key feature that the material can self-assemble into fibers, reverse analysis was performed and the chemical structure of the corresponding material feature was retrieved from the above database, which was the GAGSGAGAGSGA sequence (sequence 8).
[0143] Step S4: Verify the performance of candidate materials and screen them through molecular dynamics simulation and material genome analysis
[0144] The specific method is the same as described in step S4 of Example 1.
[0145] Step S5: Obtaining recombinant target protein by genetic engineering methods
[0146] a) Obtain target gene fragments by chemical synthesis
[0147] The target fragment (5'-3' single strand) was obtained by chemical synthesis. The restriction sites BamHI and NcoⅠ, as well as protective bases were added to both ends of the target fragment to synthesize the positive strand and negative strand respectively.
[0148] Protein sequence (SEQ ID NO: 8):
[0149] (GAGSGAGAGSGA) n ;
[0150] Corresponding gene sequence (SEQ ID NO: 9):
[0151] (ggtgctggttcaggtgctggtgctggttcaggtgct) n ;
[0152] By chemical synthesis, the target fragments were obtained respectively:
[0153] Positive chain (sequence 10):
[0154] GATCC ggtgctggttcaggtgctggtgctggttcaggtgct TGAC;
[0155] Negative strand (sequence 11):
[0156] CATGGTCA ggtgctggttcaggtgctggtgctggttcaggtgct G;
[0157] Then, a temperature gradient annealing was used to form a double strand, and the PNK enzyme was used to add a phosphate group to the 5' end;
[0158] The vector was linearized by PCR using a pET vector containing the Amp ampicillin resistance gene with a GST-6×His-TEV tag;
[0159] The forward and reverse primers contain corresponding restriction sites and protection bases, and the annealing temperature is guaranteed to be 54-60°C, for example:
[0160] Forward primer (SEQ ID NO: 6):
[0161] GAAAACCTGTACTTCCAGTCAGGATCCGCG;
[0162] Reverse primer (SEQ ID NO: 7):
[0163] GCGCCATGGTCTGACACATGCAGCTCCCG;
[0164] After verifying that the target fragment is correct, use a dedicated kit for cleaning;
[0165] Use T4 ligase to connect the linearized vector and the target fragment;
[0166] b) Using the recombinant vector pET-GST-6×His-TEV-Pro1 to transform competent cells: the specific method is as described in S5 of Example 1.
[0167] c) Plasmid transformation into competent cells: The specific method is as described in S5 of Example 1.
[0168] d) Fermentation culture of genetically engineered bacteria: The specific method is as described in S5 of Example 1.
[0169] e) Protein purification: The specific method is as described in S5 of Example 1.
[0170] Step S6: Performance verification of target protein material
[0171] The protein solution obtained in step S5 above was dropped onto the carbon film on the surface of a clean copper mesh. The sample was stained with 0.2 wt % uranyl acetate and dried. The sample morphology was observed under a transmission electron microscope.
[0172] like Figure 8 As shown, the silk protein constructed by the above method can self-assemble into fibers, and the fiber network structure can be observed by transmission electron microscopy.
[0173] Step S7. Optimize and train the neural network
[0174] After encoding the material self-assembly properties of the silk protein obtained in the above step S5 and the molecular descriptors shown in sequence 8, they are input into the neural network obtained in step S2, and the neural network in step S2 is optimized and trained based on the above data.
[0175] Example 3: Preparation of amyloid-like fibrous protein 2
[0176] Step S1: Database establishment
[0177] The specific method is the same as described in step S1 of Example 1.
[0178] Step S2: Building the model
[0179] The specific method is the same as described in step S2 of Example 1.
[0180] Step S3: Analyze and retrieve the structure related to the material properties of silk protein, and determine the silk protein structure corresponding to the target material properties
[0181] Based on the key feature that the material can self-assemble into fibers, reverse analysis is performed and the chemical structure of the corresponding material feature retrieved from the above database can also be the GAGAGAGY sequence (sequence 12).
[0182] Step S4: Verify the performance of candidate materials and screen them through molecular dynamics simulation and material genome analysis
[0183] The specific method is the same as described in step S4 of Example 1.
[0184] Step S5: Obtaining recombinant target protein through genetic engineering methods
[0185] a) Obtain target gene fragments by chemical synthesis
[0186] The target fragment (5'-3' single strand) was obtained by chemical synthesis. The restriction sites BamHI and NcoⅠ, as well as protective bases were added to both ends of the target fragment to synthesize the positive strand and negative strand respectively.
[0187] Protein sequence (SEQ ID NO: 12):
[0188] (GAGAGAGY) n ;
[0189] Corresponding gene sequence (SEQ ID NO: 13):
[0190] (ggagcaggagctggtgctggatac) n ;
[0191] By chemical synthesis, the target fragments were obtained respectively:
[0192] Positive chain (sequence 14):
[0193] GATCC ggagcaggagctggtgctggatac TGAC;
[0194] Negative strand (sequence 15):
[0195] CATGGTCAggagcaggagctggtgctggatac G;
[0196] Then, a temperature gradient annealing was used to form a double strand, and the PNK enzyme was used to add a phosphate group to the 5' end;
[0197] The vector was linearized by PCR using a pET vector containing the Amp ampicillin resistance gene with a GST-6×His-TEV tag;
[0198] The forward and reverse primers contain corresponding restriction sites and protection bases, and the annealing temperature is guaranteed to be 54-60°C, for example:
[0199] Forward primer (SEQ ID NO: 6):
[0200] GAAAACCTGTACTTCCAGTCA GGATCC GCG;
[0201] Reverse primer (SEQ ID NO: 7):
[0202] GCGCCATGGTCTGACACATGCAGCTCCCG;
[0203] After verifying that the target fragment is correct, use a dedicated kit for cleaning;
[0204] Use T4 ligase to connect the linearized vector and the target fragment;
[0205] b) Using the recombinant vector pET-GST-6×His-TEV-Pro1 to transform competent cells: the specific method is as described in S5 of Example 1.
[0206] c) Plasmid transformation into competent cells: The specific method is as described in S5 of Example 1.
[0207] d) Fermentation culture of genetically engineered bacteria: The specific method is as described in S5 of Example 1.
[0208] e) Protein purification: The specific method is as described in S5 of Example 1.
[0209] Step S6: Performance verification of target protein material
[0210] The protein solution obtained in step S5 was dropped onto the carbon film on the surface of a clean copper mesh. The sample was stained with 0.2 wt.% uranyl acetate and dried. The sample morphology was observed under a transmission electron microscope.
[0211] like Figure 9 As shown, the silk protein prepared by the above method can self-assemble into fibers, and the fiber network structure can also be observed by transmission electron microscopy.
[0212] Step S7. Optimize and train the neural network
[0213] The material self-assembly properties of the silk protein obtained in step S5 and the molecular descriptors shown in sequence 12 are encoded and input into the neural network obtained in step S2, and the neural network in step S2 is optimized and trained based on the above data.
[0214] Example 4: Preparation of silk protein material with film-forming and repairing properties
[0215] Step S1: Database establishment
[0216] The specific method is the same as described in step S1 of Example 1.
[0217] Step S2: Building the model
[0218] The specific method is the same as described in step S2 of Example 1.
[0219] Step S3: Analyze and retrieve the structure related to the material properties of silk protein, and determine the silk protein structure corresponding to the target material properties
[0220] Based on the key feature of the material, which is its film-forming and repairing properties, reverse analysis was performed and the chemical structure of the corresponding material feature was retrieved from the above database and found to be the GAGVGAGY sequence (sequence 16).
[0221] Step S4: Verify the performance of candidate materials and screen them through molecular dynamics simulation and material genome analysis
[0222] The specific method is the same as described in step S3 of Example 1.
[0223] Step S5: Obtaining recombinant target protein by genetic engineering methods
[0224] a) Obtain target gene fragments by chemical synthesis
[0225] The target fragment (5'-3' single strand) was obtained by chemical synthesis. The restriction sites BamHI and NcoⅠ, as well as protective bases were added to both ends of the target fragment to synthesize the positive strand and negative strand respectively.
[0226] Protein sequence (SEQ ID NO: 16):
[0227] (GAGVGAGY) n ;
[0228] Corresponding gene sequence (SEQ ID NO: 17):
[0229] (ggatacggagct) n ;
[0230] By chemical synthesis, the target fragments were obtained respectively:
[0231] Positive chain (sequence 18):
[0232] GATCC ggatacggagct TGAC;
[0233] Negative strand (sequence 19):
[0234] CATGGTCAggatacggagct G;
[0235] Then, a temperature gradient annealing was used to form a double strand, and the PNK enzyme was used to add a phosphate group to the 5' end;
[0236] The vector was linearized by PCR using a pET vector containing the Amp ampicillin resistance gene with a GST-6×His-TEV tag;
[0237] The forward and reverse primers contain corresponding restriction sites and protection bases, and the annealing temperature is guaranteed to be 54-60°C, for example:
[0238] Forward primer (SEQ ID NO: 6):
[0239] GAAAACCTGTACTTCCAGTCAGGATCCGCG;
[0240] Reverse primer (SEQ ID NO: 7):
[0241] GCGCCATGGTCTGACACATGCAGCTCCCG;
[0242] After verifying that the target fragment is correct, use a dedicated kit for cleaning;
[0243] Use T4 ligase to connect the linearized vector and the target fragment;
[0244] b) Using the recombinant vector pET-GST-6×His-TEV-Pro1 to transform competent cells: the specific method is as described in S5 of Example 1.
[0245] c) Plasmid transformation into competent cells: The specific method is as described in S5 of Example 1.
[0246] d) Fermentation culture of genetically engineered bacteria: The specific method is as described in S5 of Example 1.
[0247] e) Protein purification: The specific method is as described in S5 of Example 1.
[0248] Step S6: Performance verification of target protein material
[0249] a) The silk protein solution obtained in step S5 is mixed with palm oil and homogenized for 3 minutes to prepare an emulsion, and then the chitosan aqueous solution is added to the obtained emulsion and homogenized for 1 minute to obtain a protein emulsion containing the three substances. The mass fractions of the silk protein, palm oil, and chitosan in the emulsion are 1 wt %, 10 wt %, and 0.2 wt %, respectively.
[0250] At the same time, palm oil, chitosan and water were mixed and homogenized for 4 minutes to obtain a blank control emulsion. The mass fractions of palm oil and chitosan were 10 wt% and 0.2 wt% respectively.
[0251] b) treating hair tresses with the two emulsions prepared in the above steps respectively, as follows:
[0252] 5 g of bleached hair was rubbed with 10 g of the emulsion for 1 minute, the hair was rinsed for 30 seconds, and then the surface of the hair was observed using a scanning electron microscope.
[0253] like Figure 10 As shown, the cuticle scales of bleached hair fibers exhibit irregular and lifted scales; hair treated with the blank emulsion exhibits a certain repair effect, but the effect is not obvious; after treatment with the protein emulsion, the bleached hair fibers have a coating and a smoother surface. This indicates that the silk protein prepared through the above steps does have film-forming and repairing properties.
[0254] Step S7. Optimize and train the neural network
[0255] The film-forming and repairing properties of the silk protein obtained in step S5 and the molecular descriptors shown in sequence 16 are encoded and input into the neural network obtained in step S2, and the neural network in step S2 is optimized and trained based on the above data.
[0256] Although the above examples utilize recombinant methods to prepare the target silk protein, these examples are merely preferred embodiments of the present invention. Other methods, such as chemical synthesis and genetic biology, can also be used to construct the target silk protein of the present invention. The present invention should not be limited to the disclosure of these examples and the accompanying drawings. Any equivalent or modified methods that do not depart from the spirit of the present invention fall within the scope of protection of the present invention. SEQUENCE LISTING <110> Fuxiang Sitai Medical Technology (Suzhou) Co., Ltd. <120> Reconstructed silk protein and preparation method thereof <130> 20220426 <160> 19 <170> PatentIn version 3.3 <210> 1 <211> 32 <212> PRT <213> artificial synthesis <400> 1 Gly Thr Gly Ser Ser Gly Phe Gly Pro Tyr Val Ala Asn Gly Gly Tyr 1 5 10 15 Ser Arg Ser Asp Gly Tyr Glu Tyr Ala Trp Ser Ser Asp Phe Gly Thr 20 25 30 <210> 2 <211> 30 <212> PRT <213> Synthetic <400> 2 Gly Ala Gly Ala Gly Ser Gly Ala Gly Ala Gly Ser Gly Ala Gly Ala 1 5 10 15 Gly Ser Gly Ala Gly Ala Gly Ser Gly Ala Gly Ala Gly Ser 20 25 30 <210> 3 <211> 90 <212> DNA <213> Synthetic <400> 3 ggtgctggtg ctggttcagg tgctggtgct ggttcaggtg ctggtgctgg ttcaggtgcc 60 ggtgctggtt caggtgctgg tgctggttca 90 <210> 4 <211> 99 <212> DNA <213> Synthetic <400> 4 gatccggtgc tggtgctggt tcaggtgctg gtgctggttc aggtgctggt gctggttcag 60 gtgccggtgc tggttcaggt gctggtgctg gttcatgac 99 <210> 5 <211> 99 <212> DNA <213> Synthetic <400> 5 catggtcatg aaccagcacc agcacctgaa ccagcaccgg cacctgaacc agcaccagca 60 cctgaaccag caccagcacc tgaaccagca ccagcaccg 99 <210> 6 <211> 30 <212> DNA <213> Synthetic <400> 6 gaaaacctgt acttccagtc aggatccgcg 30 <210> 7 <211> 29 <212> DNA <213> Synthetic <400> 7 gcgccatggt ctgacacatg cagctcccg 29 <210> 8 <211> 12 <212> PRT <213> Synthetic <400> 8 Gly Ala Gly Ser Gly Ala Gly Ala Gly Ser Gly Ala 1 5 10 <210> 9 <211> 36 <212> DNA <213> Synthetic <400> 9 ggtgctggtt caggtgctgg tgctggttca ggtgct 36 <210> 10 <211> 45 <212> DNA <213> Synthetic <400> 10 gatccggtgc tggttcaggt gctggtgctg gttcaggtgc ttgac 45 <210> 11 <211> 45 <212> DNA <213> artificial synthesis <400> 11 catggtcagg tgctggttca ggtgctggtg ctggttcagg tgctg 45 <210> 12 <211> 8 <212> PRT <213> artificial synthesis <400> 12 Gly Ala Gly Ala Gly Ala Gly Tyr 1 5 <210> 13 <211> twenty four <212> DNA <213> artificial synthesis <400> 13 ggagcaggag ctggtgctgg atac 24 <210> 14 <211> 33 <212> DNA <213> artificial synthesis <400> 14 gatccggagc aggagctggt gctggatact gac 33 <210> 15 <211> 33 <212> DNA <213> artificial synthesis <400> 15 catggtcagg agcaggagct ggtgctggat acg 33 <210> 16 <211> 8 <212> PRT <213> artificial synthesis <400> 16 Gly Ala Gly Val Gly Ala Gly Tyr 1 5 <210> 17 <211> 12 <212> DNA <213> artificial synthesis <400> 17 ggatacggag ct 12 <210> 18 <211> twenty one <212> DNA <213> artificial synthesis <400> 18 gatccggata cggagcttga c 21 <210> 19 <211> twenty one <212> DNA <213> artificial synthesis <400> 19 catggtcagg atacggagct g 21
Claims
1. A method for constructing silk protein with target material properties using materials genomics, comprising the following steps: S1. Establish a silk protein database; S2. Build the model; S3. Analyze and retrieve structures related to the material properties of silk proteins, and determine the silk protein structures corresponding to the target material properties; S4. Validation and screening of silk protein candidates through computational and materials genomic analyses; S5. Obtaining target silk protein; S6. Verify the target material properties of silk protein; S7. The material properties and molecular descriptors of the silk protein obtained in step S5 are input into the neural network obtained in step S2, and the neural network is optimized and trained in step S2 according to the test data; in, Step S1 is performed as follows: By defining characteristic expressions of silk proteins, including SDF, MOL, SMILES codes or images, relevant silk protein characteristic expressions and experimental data were extracted from different patents and literature, and structured extraction and analysis were performed. All amino acid sequences and structures with specific functional characteristics were searched in the SciFinder, ChemSpider and PubChem databases, and batch downloads were performed through the database application program interface; Step S2 is performed in the following manner: a) obtaining training data of multiple molecular descriptors of silk protein and corresponding silk protein material properties; b) obtaining a neural network that characterizes the relationship between the molecular descriptors of the silk protein material and the properties of the silk protein material based on the training data; c) obtaining test data, testing the neural network based on the test data, and if the test result does not meet the error requirement, returning to step b) until the test result meets the error requirement; Step S3 is performed in the following manner: a) Select target material properties that match the actual application scenario; b) inputting the target material properties into the neural network obtained in step S2 to find a series of molecular descriptors that meet the target material properties; c) searching and screening a number of candidate protein structures from the database in step S1 based on the obtained molecular descriptors; Step S4 is performed in the following manner: a) Calculate the properties of candidate materials using molecular dynamics simulations; b) Evaluate and validate candidate materials using materials genome analysis; Step S5: constructing the target silk protein by genetic engineering, chemical synthesis, and genetic biology methods; The target material properties are hydrophilicity, self-assembly property, film-forming property or repair property; The molecular descriptors include: primary structure, aggregated structure, molecular structure and exogenous factors of silk protein.
2. The method according to claim 1, wherein the primary structure of the silk protein comprises: amino acid types and contents, and amino acid sequences and contents corresponding to the special aggregation structures forming α-helix, β-sheet, β-turn, and random-coil.
3. The method according to claim 1, wherein the aggregated structure comprises: α-helix, β-sheet, β-turn, and random-coil.
4. The method according to claim 1, wherein the exogenous factors include solvent, ions, pH value, and stress. 5 . The method according to claim 1 , wherein the silk protein structure determined in step S3 is as shown in Sequence 1, Sequence 8, Sequence 12 or Sequence 16.
Citation Information
Patent Citations
Extraction method for hydrophilic recombinant protein
CN104884463A
Silk protein scaffold material and preparation method thereof
CN104971386A