Method and system for screening and constructing small molecule polypeptide simulant based on interaction between proteins

Through multiple rounds of molecular docking and sequence alignment, the polypeptide design is solved, and the problems of poor selectivity and off-target effects when regulating protein interactions in the prior art are solved, and the screening and verification of polypeptide sequences that efficiently bind ADAM10 protein is achieved, providing a new idea for developing targeted drugs.

CN120108484AInactive Publication Date: 2025-06-06DONGGUAN SOUTHEAST CENTRAL HOSPITAL (DONGGUAN SOUTHEAST TRADITIONAL CHINESE MEDICINE MEDICAL SERVICE CENTER DONGGUAN FIRST HOSPITAL AFFILIATED TO GUANGDONG MEDICAL UNIVERSITY)
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510171205.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art faces the problems of poor selectivity, poor cell permeability and off-target effects when regulating protein interactions (PPI), and lacks efficient peptide design methods and the ability to predict in vivo efficacy.

Method used

By obtaining the three-dimensional structural prediction files of RECK protein and ADAM10 protein, multiple rounds of molecular docking and sequence alignment were carried out, polypeptide sequences that could efficiently bind ADAM10 protein were designed and screened, and their biological functions were verified through cellular experiments.

Benefits of technology

The screening and verification of polypeptide sequences with high binding ability of ADAM10 protein has been achieved, providing new ideas for developing drugs targeting ADAM10 protein, and improving the efficiency and accuracy of peptide design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108484A_ABST
    Figure CN120108484A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for screening and constructing a small molecule polypeptide simulant based on interaction between proteins, relates to the technical field of biology, and comprises the following steps: predicting an interaction model of RECK protein and ADAM10 protein through a molecular docking platform, and analyzing interaction sites of the RECK protein and the ADAM10 protein. And then, determining a KAZAL region of the RECK protein according to the interaction site, and predicting a sequence structure of the KAZAL region. And then carrying out molecular docking on the KAZAL region and ADAM10 protein, designing a plurality of peptide fragments with the length of 30 amino acids, and screening the peptide fragment sequence with the highest score as candidate polypeptide. And finally, performing species homology comparison and quality control detection on the candidate polypeptide, and verifying the binding capacity and biological function of the candidate polypeptide to ADAM10 protein through cell experiments. The invention provides a method for screening to obtain a polypeptide simulant specifically bound with ADAM10 protein. The polypeptide simulant can be used for treating ADAM10 related diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to biotechnology, and in particular to a method and system for constructing small molecule polypeptide mimetics based on protein-protein interaction screening. Background Art

[0002] Protein-protein interactions (PPIs) play crucial roles in various biological processes, including signal transduction, metabolism, and disease development. Understanding and regulating PPIs has become a key goal for drug discovery and therapeutic intervention. Given the importance of PPIs, various methods have been developed to study and target these interactions. These methods include yeast two-hybrid, co-immunoprecipitation, and surface plasmon resonance-based techniques. In recent years, computer simulations, especially molecular docking, have become valuable tools for studying PPIs and designing small molecules that target these interactions.

[0003] However, conventional small-molecule-based inhibitors often face several challenges, such as poor selectivity, poor cell permeability, and off-target effects. To overcome these limitations, there has been increasing interest in developing peptide-based inhibitors that mimic protein interaction interfaces. Peptide mimics offer several advantages, including high specificity, good biocompatibility, and relatively easy synthesis. Despite the promise of peptide mimics, designing peptides that can effectively and specifically target PPIs remains challenging.

[0004] There are three main defects and deficiencies in the prior art:

[0005] 1. Reliance on experimental methods: Existing methods for identifying PPI modulators rely heavily on experimental screening, which can be time-consuming and laborious. These methods often require a lot of resources and may not identify peptides with the required specificity and affinity.

[0006] 2. Lack of rational peptide design methods: Most existing peptide design strategies rely on limited experimental data or time-consuming optimization procedures based on structural information. These methods are usually unable to explore the vast peptide sequence space and may not identify peptides with the best binding affinity and specificity.

[0007] 3. Difficulty predicting in vivo efficacy: Even well-designed peptides may have difficult-to-predict efficacy in vivo. Peptides may be degraded by proteases or may not be able to effectively penetrate cell membranes, limiting their therapeutic potential. Summary of the invention

[0008] The embodiments of the present invention provide a method and system for constructing small molecule polypeptide mimetics based on protein-protein interaction screening, which can solve the problems in the prior art.

[0009] A first aspect of an embodiment of the present invention provides a method for constructing a small molecule polypeptide mimetic based on protein-protein interaction screening, comprising:

[0010] Obtaining three-dimensional structure prediction files of RECK protein and ADAM10 protein, wherein the three-dimensional structure prediction file of RECK protein is AF-O95980-F1, and the three-dimensional structure prediction file of ADAM10 protein is AF-O14672-F1;

[0011] The three-dimensional structure prediction files of the RECK protein and the ADAM10 protein were input into the molecular docking platform for the first round of molecular docking to obtain the interaction model between the RECK protein and the ADAM10 protein, and the interaction site between the RECK protein and the ADAM10 protein was analyzed by the PyMOL visualization software; the KAZAL region of the RECK protein was determined according to the interaction site, and the sequence structure of the KAZAL region was predicted using the Swiss-model platform; the KAZAL region was subjected to the second round of molecular docking with the ADAM10 protein, and multiple peptides with a length of 30 amino acids were designed according to the docking results; the multiple peptides were respectively subjected to the third round of molecular docking with the ADAM10 protein, and the peptide sequence with the highest score was selected as the candidate polypeptide based on the docking score;

[0012] The candidate polypeptide is subjected to species homology comparison and quality control detection, wherein the quality control detection includes polypeptide content detection, high performance liquid chromatography detection and mass spectrometry detection; the binding ability and biological function of the candidate polypeptide to ADAM10 protein are verified through cell experiments, wherein the amino acid sequence of the candidate polypeptide is PCNCADQFVPVCGQNGRTYPSACIARCVG.

[0013] The three-dimensional structure prediction files of the RECK protein and the ADAM10 protein were input into the molecular docking platform for the first round of molecular docking to obtain the interaction model between the RECK protein and the ADAM10 protein, and the interaction sites between the RECK protein and the ADAM10 protein were analyzed by the PyMOL visualization software, including:

[0014] Pre-processing of the RECK protein structure and the ADAM10 protein structure included removing crystal water molecules from the protein structure, adding hydrogen atoms to the protein structure, optimizing the protein structure for energy minimization, and verifying the rationality of the protein structure conformation through the Ramachandran plot;

[0015] The pre-processed RECK protein structure and ADAM10 protein structure were input into the molecular docking platform, and the RECK protein structure and the ADAM10 protein structure were rigidly docked based on the fast Fourier transform algorithm and the Monte Carlo local optimization algorithm. The search space range with a radius of 30 angstroms was set with the center of the RECK protein as the center. 1000 conformations were generated each time the docking was performed, and the top 100 conformations were retained according to the shape complementarity score.

[0016] The flexibility optimization was performed on the top 100 retained conformations, including side chain optimization and local energy minimization of the conformations, and the binding free energy was obtained by calculating the van der Waals interaction energy, electrostatic interaction energy, and desolvation free energy;

[0017] The conformation after flexibility optimization was imported into PyMOL visualization software, and the interaction interface was analyzed by calculating the interface contact area, identifying the hydrogen bond network, analyzing the hydrophobic interaction and identifying the salt bridge; a 5 angstrom cutoff value was set in the interaction interface to screen the interface residues, calculate the change in the buried area of ​​each residue, analyze the residue conservation, and identify the residues that contribute more than 2 kcal / mol to the binding free energy as hotspot residues; the conformations were clustered based on the root mean square deviation to identify the main binding mode and evaluate the stability of the binding conformation. The root mean square deviation was calculated by the following formula: the square of the coordinate difference between each atom and the corresponding atom in the reference structure was summed, the average was taken, and the square root was taken; the analysis results were automatically processed by the PyMOL script to generate a standardized analysis report containing interface residue information, hotspot residue information and main binding modes.

[0018] Determining the KAZAL region of the RECK protein according to the interaction site, and using the Swiss-model platform to predict the sequence structure of the KAZAL region includes:

[0019] A multiple sequence alignment of the RECK protein sequences was performed, and a position-specific scoring matrix was constructed by calculating the logarithmic ratio of the observed frequency of the amino acid at each position to the background frequency to obtain the conservation score of the RECK protein sequence;

[0020] Based on the conservation score, the domain analysis of RECK protein was performed, the KAZAL motif features were analyzed using InterProScan, the domain boundaries were verified in combination with the SMART database, the functional regions were annotated using the Pfam database, and the amino acid sequence of the KAZAL region in the RECK protein was determined;

[0021] Inputting the amino acid sequence of the KAZAL region into the Swiss-model platform, screening structures with sequence similarity greater than 30% and structural resolution less than 2.5 angstroms as modeling templates, evaluating the template quality by calculating the weighted sum of the sequence similarity score, resolution score and R factor score of the template, and selecting the optimal template;

[0022] Based on the optimal template, structural modeling of the KAZAL region was performed, including sequence alignment optimization, backbone construction, side chain reconstruction, loop region optimization and energy minimization, to obtain a three-dimensional structural model of the KAZAL region.

[0023] The KAZAL region was subjected to a second round of molecular docking with the ADAM10 protein, and multiple peptide segments with a length of 30 amino acids were designed based on the docking results, including:

[0024] Performing interface feature analysis on the three-dimensional structural model of the KAZAL region, calculating the solvent accessible surface area of ​​each residue, analyzing the distribution of polar residues, hydrophobic residue aggregation and secondary structure elements, identifying β turns, evaluating α helix stability, predicting disulfide bond sites, and obtaining a map of interaction interface features;

[0025] According to the interaction interface characteristic map, the structural units are divided, the boundary positions of the secondary structures are determined, the hydrophobic core regions are identified, the polar interaction network is analyzed, and the optimal peptide length is determined to be 30 amino acids by calculating the peptide folding free energy;

[0026] Designing candidate peptides based on the structural unit division results, performing amino acid tendency analysis, charge distribution optimization and hydrophobicity balance on each structural unit, and designing multiple candidate peptides;

[0027] The candidate peptides are evaluated and optimized for properties, including prediction of secondary structure accuracy, solubility and aggregation tendency, calculation of isoelectric point and half-life, and evaluation of structural stability. On the premise of ensuring sequence diversity and retention of functional motifs, the final polypeptide sequence is screened out through conformational space sampling. The final polypeptide sequence has the same structural features and functional properties as the KAZAL region.

[0028] The third round of molecular docking was performed on multiple peptides with ADAM10 protein, and the peptide sequences with the highest scores were selected as candidate peptides based on the docking scores, including:

[0029] The PEP-FOLD algorithm was used to predict the secondary structure of the peptide sequence, and the initial conformation of the peptide was obtained through energy minimization optimization. The catalytic active site of the ADAM10 protein was analyzed, and the area around the active site was set as the binding pocket through three-dimensional grid point volume calculation to generate the spatial range and position information of the binding pocket.

[0030] Molecular docking parameters are set based on the spatial range and position information of the binding pocket, and a multidimensional scoring function is constructed, wherein the multidimensional scoring function includes shape complementarity energy, electrostatic interaction energy, hydrogen bond energy, and hydrophobic interaction energy;

[0031] The initial conformation of the polypeptide is molecularly docked with the binding pocket of the ADAM10 protein, and the binding energy, residue pair tendency, secondary structure matching degree, and interface area of ​​the polypeptide and the ADAM10 protein are calculated using the multidimensional scoring function to obtain an initial score of the docking complex;

[0032] Normalizing the initial scores of the docked complex, obtaining a comprehensive score by weighted summation, and calculating the interface contact score based on the ratio of the distance between the atomic pairs to the standard contact distance to generate score data for the peptide-protein complex;

[0033] Based on the scoring data of the polypeptide-protein complex, a molecular dynamics simulation is performed on the complex, conformational entropy is calculated through conformational probability distribution, the structural stability of the complex is evaluated, and conformational stability data is obtained;

[0034] The comprehensive score, interface contact score, and conformational stability data are input into a multi-objective optimization model, and the Pareto optimal solution analysis method is used to optimize and screen the polypeptide sequence to obtain the optimal polypeptide sequence.

[0035] The PEP-FOLD algorithm was used to predict the secondary structure of the peptide sequence, and the initial conformation of the peptide was obtained through energy minimization optimization. The catalytic active site of the ADAM10 protein was analyzed, and the area around the active site was set as the binding pocket through three-dimensional grid point volume calculation. The spatial range and position information of the binding pocket were generated, including:

[0036] The peptide sequence is divided into fragments based on the physical and chemical properties of amino acids, and the dihedral angle tendency of each fragment is calculated by the frequency of dihedral angles in the database to obtain the fragment dihedral angle probability distribution;

[0037] According to the probability distribution of dihedral angles of the fragments, the α-helix tendency and β-folding tendency of each amino acid site in the polypeptide sequence are calculated, and the secondary structure tendency score is obtained by weighting by the position weight factor;

[0038] Based on the secondary structure propensity score, an initial conformation of the polypeptide is constructed, and a potential energy function including bond length energy, bond angle energy, torsion energy, van der Waals interaction energy, and electrostatic interaction energy is used to optimize the conformation, and the optimal conformation of the polypeptide is obtained by iterative calculation using the BFGS algorithm until the energy gradient is less than a preset convergence threshold;

[0039] The ADAM10 protein sequence was subjected to conservation analysis, highly conserved sites were identified by calculating sequence entropy, and the catalytically active regions were determined by combining the distribution of secondary structure elements with surface exposure analysis;

[0040] Taking the catalytically active region as the center, the three-dimensional space is divided into a regular grid with a side length of 0.5 angstroms, and the distance between each grid point and the surface atom of the protein is calculated to determine whether the grid point belongs to the binding pocket, and all grid points belonging to the binding pocket are volume-integrated to obtain the volume data of the binding pocket;

[0041] Calculating the pocket depth index based on the volume data of the binding pocket, characterizing the spatial characteristics of the pocket by the ratio of the pocket opening area to the volume, and analyzing the hydrophobicity distribution in the pocket;

[0042] The Poisson-Boltzmann equation is used to calculate the electrostatic potential distribution in the binding pocket area, and the electrostatic potential field is constructed based on the atomic charge and dielectric constant to obtain the potential distribution map of the binding pocket.

[0043] The optimal conformation of the polypeptide, binding pocket volume data, pocket depth index, and electrostatic potential distribution map are integrated into a structural feature descriptor to guide the subsequent molecular docking process. The structural feature descriptor contains spatial matching information and interaction characteristics between the polypeptide and the binding site.

[0044] The candidate polypeptides are subjected to species homology comparison and quality control detection, wherein the quality control detection includes polypeptide content detection, high performance liquid chromatography detection and mass spectrometry detection including:

[0045] Obtain the observed frequency of amino acid pairs and the background frequency of single amino acids of the candidate polypeptides through BLAST sequence alignment, calculate the amino acid pair substitution scoring matrix, analyze the sequence consistency based on the scoring matrix, and obtain the ratio of the number of matching sites to the alignment length;

[0046] Performing interspecies conservation analysis based on the sequence alignment results, calculating site conservation scores by amino acid frequency, constructing a phylogenetic tree, and obtaining genetic distance data for candidate polypeptides;

[0047] The candidate polypeptide is subjected to acid hydrolysis treatment, a working curve is established using a standard amino acid solution, the content of amino acid components in the hydrolyzate is determined, and the polypeptide recovery rate is calculated;

[0048] Based on the amino acid component content results, the absorbance at 280 nanometers of the candidate polypeptide is measured by ultraviolet spectrophotometry, and the polypeptide concentration is calculated by combining the extinction coefficient and the optical path length to obtain polypeptide content data;

[0049] Reversed-phase high performance liquid chromatography was used to separate and analyze the candidate peptides, the chromatographic conditions were optimized by theoretical plate number calculation, the separation degree was calculated based on retention time and peak width, and the chromatographic spectrum of the candidate peptides was obtained;

[0050] The peak area of ​​each chromatographic peak in the chromatogram is compared with the total peak area, the purity of the target peak is calculated, the symmetry of the peak shape is evaluated by the tailing factor, and the purity data of the peptide is obtained;

[0051] Electrospray ionization mass spectrometry was used to determine the mass-to-charge ratio of candidate peptides, analyze isotope distribution patterns, and obtain fragment ion spectra through tandem mass spectrometry to identify peptide sequences and modification sites;

[0052] Establish quality control indicators including main peak content and single impurity content limits, conduct accelerated tests and solution stability tests on candidate peptides, and evaluate the stability of candidate peptides under different pH and temperature conditions;

[0053] Based on the sequence comparison results, polypeptide content data, polypeptide purity data, mass spectrometry analysis results and stability data, a comprehensive polypeptide quality evaluation system is established to screen candidate polypeptides that meet quality control requirements.

[0054] A second aspect of the embodiments of the present invention provides a system for constructing small molecule polypeptide mimetics based on protein-protein interaction screening, comprising:

[0055] The first unit is used to obtain the three-dimensional structure prediction files of RECK protein and ADAM10 protein, wherein the three-dimensional structure prediction file of RECK protein is AF-O95980-F1, and the three-dimensional structure prediction file of ADAM10 protein is AF-O14672-F1;

[0056] The second unit is used to input the three-dimensional structure prediction files of the RECK protein and the ADAM10 protein into the molecular docking platform for the first round of molecular docking, obtain the interaction model of the RECK protein and the ADAM10 protein, and analyze the interaction site of the RECK protein and the ADAM10 protein by PyMOL visualization software; determine the KAZAL region of the RECK protein according to the interaction site, and use the Swiss-model platform to predict the sequence structure of the KAZAL region; perform a second round of molecular docking on the KAZAL region and the ADAM10 protein, and design multiple peptides with a length of 30 amino acids according to the docking results; perform a third round of molecular docking on the multiple peptides and the ADAM10 protein respectively, and select the peptide sequence with the highest score as the candidate polypeptide based on the docking score;

[0057] The third unit is used to perform species homology comparison and quality control detection on the candidate polypeptide, wherein the quality control detection includes polypeptide content detection, high performance liquid chromatography detection and mass spectrometry detection; verify the binding ability and biological function of the candidate polypeptide to ADAM10 protein through cell experiments, wherein the amino acid sequence of the candidate polypeptide is PCNCADQFVPVCGQNGRTYPSACIARCVG.

[0058] A third aspect of the embodiments of the present invention

[0059] An electronic device is provided, comprising:

[0060] processor;

[0061] a memory for storing processor-executable instructions;

[0062] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0063] A fourth aspect of the embodiments of the present invention is:

[0064] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0065] The beneficial effects of this application are as follows:

[0066] 1. A method for screening and constructing small molecule peptide mimetics based on protein-protein interactions is provided, which is used to simulate the interaction between RECK protein and ADAM10 protein, and to screen for peptides that can bind to ADAM10 protein.

[0067] 2. Through multiple rounds of molecular docking and sequence alignment, this method finally screened out a polypeptide sequence LPCNCADQFVPVCGQNGRTYPSACIARCVG with high binding ability to ADAM10 protein, and verified its biological function through cell experiments.

[0068] 3. This method provides new ideas and candidate molecules for the development of drugs targeting ADAM10 protein. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 A schematic diagram of a process for constructing a small molecule polypeptide mimetic based on protein-protein interaction screening according to an embodiment of the present invention;

[0070] Figure 2 Schematic diagram showing that RECK K1-4 polypeptides do not produce cytotoxicity when cultured with human umbilical vein endothelial cells (HUVEC);

[0071] Figure 3 Schematic diagram showing that RECK K1-4 polypeptide can be well taken up by endothelial cells;

[0072] Figure 4 Schematic diagram showing that RECK K1-4 polypeptides affect ADAM10 protein levels in human umbilical vein endothelial cells in a time-dependent manner;

[0073] Figure 5 Schematic diagram showing that intravenous injection of RECK K1-4 peptide can reduce the levels of circulating inflammatory factors in septic mice;

[0074] Figure 6 This is a schematic diagram of the structure of a small molecule polypeptide mimetic system constructed based on protein-protein interaction screening in an embodiment of the present invention. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0076] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0077] Figure 1 FIG. 1 is a flow chart of a method for constructing small molecule polypeptide mimetics based on protein-protein interaction screening according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0078] S101. Obtaining three-dimensional structure prediction files of RECK protein and ADAM10 protein, wherein the three-dimensional structure prediction file of RECK protein is AF-O95980-F1, and the three-dimensional structure prediction file of ADAM10 protein is AF-O14672-F1;

[0079] S102. Input the three-dimensional structure prediction files of the RECK protein and the ADAM10 protein into the molecular docking platform for the first round of molecular docking, obtain the interaction model of the RECK protein and the ADAM10 protein, and analyze the interaction site of the RECK protein and the ADAM10 protein by PyMOL visualization software; determine the KAZAL region of the RECK protein according to the interaction site, and use the Swiss-model platform to predict the sequence structure of the KAZAL region; perform a second round of molecular docking on the KAZAL region and the ADAM10 protein, and design multiple peptides with a length of 30 amino acids according to the docking results; perform a third round of molecular docking on the multiple peptides and the ADAM10 protein, and select the peptide sequence with the highest score as the candidate polypeptide based on the docking score;

[0080] S103. Perform species homology comparison and quality control testing on the candidate polypeptide, wherein the quality control testing includes polypeptide content testing, high performance liquid chromatography testing and mass spectrometry testing; verify the binding ability and biological function of the candidate polypeptide to ADAM10 protein through cell experiments, wherein the amino acid sequence of the candidate polypeptide is LPCNCADQFVPVCGQNGRTYPSACIARCVG.

[0081] In an optional embodiment, the three-dimensional structure prediction files of the RECK protein and the ADAM10 protein are input into a molecular docking platform for a first round of molecular docking to obtain an interaction model between the RECK protein and the ADAM10 protein, and the interaction sites between the RECK protein and the ADAM10 protein are analyzed by PyMOL visualization software, including:

[0082] Pre-processing of the RECK protein structure and the ADAM10 protein structure included removing crystal water molecules from the protein structure, adding hydrogen atoms to the protein structure, optimizing the protein structure for energy minimization, and verifying the rationality of the protein structure conformation through the Ramachandran plot;

[0083] The pre-processed RECK protein structure and ADAM10 protein structure were input into the molecular docking platform, and the RECK protein structure and the ADAM10 protein structure were rigidly docked based on the fast Fourier transform algorithm and the Monte Carlo local optimization algorithm. The search space range with a radius of 30 angstroms was set with the center of the RECK protein as the center. 1000 conformations were generated each time the docking was performed, and the top 100 conformations were retained according to the shape complementarity score.

[0084] The flexibility optimization was performed on the top 100 retained conformations, including side chain optimization and local energy minimization of the conformations, and the binding free energy was obtained by calculating the van der Waals interaction energy, electrostatic interaction energy, and desolvation free energy;

[0085] The conformation after flexibility optimization was imported into PyMOL visualization software, and the interaction interface was analyzed by calculating the interface contact area, identifying the hydrogen bond network, analyzing the hydrophobic interaction and identifying the salt bridge; a 5 angstrom cutoff value was set in the interaction interface to screen the interface residues, calculate the change in the buried area of ​​each residue, analyze the residue conservation, and identify the residues that contribute more than 2 kcal / mol to the binding free energy as hotspot residues; the conformations were clustered based on the root mean square deviation to identify the main binding mode and evaluate the stability of the binding conformation. The root mean square deviation was calculated by the following formula: the square of the coordinate difference between each atom and the corresponding atom in the reference structure was summed, the average was taken, and the square root was taken; the analysis results were automatically processed by the PyMOL script to generate a standardized analysis report containing interface residue information, hotspot residue information and main binding modes.

[0086] The molecular docking and analysis method of the interaction site between RECK protein and ADAM10 protein, the specific implementation steps are as follows:

[0087] First, protein structure preprocessing is performed. Obtain the three-dimensional structure files of RECK protein and ADAM10 protein, for example, download them from the Protein Data Bank (PDB). Use molecular modeling software, such as UCSF Chimera or PyMOL, to remove unnecessary components in the protein structure, such as crystal water molecules, ligands, and other non-protein molecules. Use molecular mechanics software, such as Amber or Gromacs, to add hydrogen atoms to the protein structure to ensure the integrity of the protein structure. Perform energy minimization optimization on the protein structure to eliminate unreasonable atomic contacts and bond lengths and angles to obtain a stable protein conformation. Use the Ramachandran diagram to analyze the conformational rationality of the protein structure to ensure that the main chain dihedral angles of the protein structure are within the allowed range. For example, the proportion of residues in the most favorable area of ​​the main chain dihedral angles in the Ramachandran diagram should be greater than 90%.

[0088] Next, molecular docking is performed. The pre-processed RECK protein and ADAM10 protein structures are input into a molecular docking platform, such as AutoDock Vina or ZDOCK. Rigid docking of RECK protein and ADAM10 protein is performed based on the fast Fourier transform algorithm and the Monte Carlo local optimization algorithm. A search space range with a radius of 30 angstroms is set with the center of the RECK protein as the center to ensure that the search space is large enough to contain all possible binding modes. The docking parameters are set to generate 1000 conformations for each docking, and the top 100 conformations are retained according to the shape complementarity score, for example, the 100 conformations with the highest shape complementarity score are selected.

[0089] The retained conformations are then subjected to flexibility optimization. The top 100 conformations are subjected to side chain optimization and local energy minimization, allowing the protein side chains to rotate and translate to accommodate the binding interface. The binding free energy is obtained by calculating the van der Waals interaction energy, electrostatic interaction energy, and desolvation free energy, for example, using the MM-GBSA or MM-PBSA method.

[0090] Finally, the interaction interface analysis is performed. The conformation after flexibility optimization is imported into the PyMOL visualization software. The interaction interface is analyzed by calculating the interface contact area, identifying the hydrogen bond network, analyzing the hydrophobic interaction, and identifying the salt bridge. A 5 angstrom cutoff value is set in the interaction interface to screen the interface residues, and the buried area change of each residue is calculated to analyze the residue conservation. Residues that contribute more than 2 kcal / mol to the binding free energy are identified as hotspot residues. For example, hotspot residues are identified by calculating the binding free energy contribution of each residue. The conformations are clustered based on the root mean square deviation to identify the main binding mode and evaluate the stability of the binding conformation. The root mean square deviation is calculated by summing the squares of the coordinate differences between each atom and the corresponding atom of the reference structure, taking the average and then taking the square root. For example, the root mean square deviation is calculated using the align command of PyMOL. The analysis results are automatically processed by the PyMOL script to generate a standardized analysis report containing interface residue information, hotspot residue information, and main binding modes. For example, a PyMOL plug-in is written using Python scripts to automate the analysis process.

[0091] For example, in a docking experiment, the binding free energy of RECK protein and ADAM10 protein was -10 kcal / mol. The interaction interface is mainly composed of hydrophobic interactions and hydrogen bonds. Asp185 of RECK protein and Arg212 of ADAM10 protein form a salt bridge. Trp315 of RECK protein and Phe408 of ADAM10 protein form a π-π interaction. The docking results show that the Loop 2 region of RECK protein is the key region for interaction with ADAM10 protein.

[0092] The beneficial effects of this method are reflected in the following three aspects:

[0093] Improve analysis efficiency: Automated processing can significantly improve analysis efficiency, save researchers’ time and effort, and reduce human errors.

[0094] Provide more comprehensive information: Combining multiple analysis methods can provide more comprehensive interaction information, including binding free energy, interface residues, hot spot residues and main binding modes, which helps to deeply understand the protein interaction mechanism.

[0095] Standardization and reproducibility of results: Standardized analysis procedures and reports can improve the reproducibility and reliability of results and facilitate comparison and communication between different studies.

[0096] In an optional embodiment, determining the KAZAL region of the RECK protein according to the interaction site, and using the Swiss-model platform to predict the sequence structure of the KAZAL region comprises:

[0097] A multiple sequence alignment of the RECK protein sequences was performed, and a position-specific scoring matrix was constructed by calculating the logarithmic ratio of the observed frequency of the amino acid at each position to the background frequency to obtain the conservation score of the RECK protein sequence;

[0098] Based on the conservation score, the domain analysis of RECK protein was performed, the KAZAL motif features were analyzed using InterProScan, the domain boundaries were verified in combination with the SMART database, the functional regions were annotated using the Pfam database, and the amino acid sequence of the KAZAL region in the RECK protein was determined;

[0099] Inputting the amino acid sequence of the KAZAL region into the Swiss-model platform, screening structures with sequence similarity greater than 30% and structural resolution less than 2.5 angstroms as modeling templates, evaluating the template quality by calculating the weighted sum of the sequence similarity score, resolution score and R factor score of the template, and selecting the optimal template;

[0100] Based on the optimal template, structural modeling of the KAZAL region was performed, including sequence alignment optimization, backbone construction, side chain reconstruction, loop region optimization and energy minimization, to obtain a three-dimensional structural model of the KAZAL region.

[0101] The KAZAL region structure prediction and model construction method of RECK protein achieved the three-dimensional structure prediction of the KAZAL region of RECK protein through technical means such as multiple sequence alignment, conservation analysis, domain identification and homology modeling.

[0102] First, the amino acid sequence of the RECK protein is obtained. For example, the FASTA format sequence of the RECK protein can be retrieved and downloaded from the NCBI database. Assume that the obtained RECK protein sequence is SEQ ID NO:1.

[0103] Next, perform a multiple sequence alignment. Use the Clustal Omega program to align the RECK protein sequence with homologous sequences from other species. For example, you can select RECK protein sequences from species such as humans, mice, and rats and perform a multiple sequence alignment together. The alignment results can show the conservation of RECK protein sequences from different species.

[0104] Then, a conservation score is performed. According to the results of the multiple sequence alignment, the frequency of occurrence of different amino acids at each site is counted. The higher the degree of conservation of the amino acid at the site, the smaller the difference between different species, and the higher the corresponding frequency. By comparing the observed frequency of amino acids at each site with the background frequency, the conservation score of each site can be calculated. The higher the conservation score, the more conservative the amino acid at the site. For example, a threshold can be set to screen out sites with a conservation score higher than the threshold.

[0105] Based on the conservation score, domain analysis was performed. InterProScan was used to analyze the RECK protein sequence and identify the KAZAL motif features. The domain boundaries were verified in conjunction with the SMART database, and the functional regions were annotated using the Pfam database. Through these analyses, the specific location and amino acid sequence of the KAZAL region in the RECK protein can be determined. For example, the KAZAL region of the RECK protein was determined to be located at amino acids 100 to 150 by analysis. Its sequence is SEQ ID NO: 2.

[0106] After obtaining the amino acid sequence of the KAZAL region, it was input into the Swiss-Model platform for homology modeling. In the Swiss-Model platform, known structures with a sequence similarity greater than 30% and a structural resolution less than 2.5 angstroms with the KAZAL region were searched as candidate templates. For example, the protein structure with PDB ID 1XYZ was searched for a similarity of 50% with the target sequence and a resolution of 2.0 angstroms.

[0107] Perform quality assessment on candidate templates. Consider the sequence similarity score, resolution score and R factor score of the template, and calculate the weighted sum to evaluate the template quality. Select the template with the highest weighted sum score as the optimal template. For example, the sequence similarity score of 1XYZ is 0.8, the resolution score is 0.9, and the R factor score is 0.7. The weighted total score is 0.8*0.5+0.9*0.3+0.7*0.2=0.85. Assuming that the scores of other candidate templates are all lower than 0.85, select 1XYZ as the optimal template.

[0108] Based on the optimal template, the KAZAL region was structurally modeled. The amino acid sequence of the KAZAL region was aligned and optimized with the optimal template to construct the initial skeleton structure. Then, side chain reconstruction, loop region optimization, and energy minimization were performed to finally obtain the three-dimensional structural model of the KAZAL region.

[0109] The beneficial effects of this method are reflected in three aspects:

[0110] First, the accuracy of the structure prediction of the KAZAL region of the RECK protein was improved. Through multiple sequence alignment and conservation analysis, the key amino acid sites in the KAZAL region can be identified, providing more accurate sequence information for structural modeling.

[0111] Second, the process of predicting the structure of the KAZAL region of the RECK protein has been simplified. Using existing bioinformatics tools and databases, steps such as domain identification, template screening, and model building can be completed quickly and efficiently.

[0112] Third, it promotes the study of the functional mechanism of RECK protein. Obtaining the three-dimensional structural model of the KAZAL region of RECK protein can provide an important structural basis for studying its interaction with other molecules and its functional mechanism.

[0113] In an optional embodiment, the KAZAL region is subjected to a second round of molecular docking with the ADAM10 protein, and a plurality of peptide segments with a length of 30 amino acids are designed according to the docking results, including:

[0114] Performing interface feature analysis on the three-dimensional structural model of the KAZAL region, calculating the solvent accessible surface area of ​​each residue, analyzing the distribution of polar residues, hydrophobic residue aggregation and secondary structure elements, identifying β turns, evaluating α helix stability, predicting disulfide bond sites, and obtaining a map of interaction interface features;

[0115] According to the interaction interface characteristic map, the structural units are divided, the boundary positions of the secondary structures are determined, the hydrophobic core regions are identified, the polar interaction network is analyzed, and the optimal peptide length is determined to be 30 amino acids by calculating the peptide folding free energy;

[0116] Designing candidate peptides based on the structural unit division results, performing amino acid tendency analysis, charge distribution optimization and hydrophobicity balance on each structural unit, and designing multiple candidate peptides;

[0117] The candidate peptides are evaluated and optimized for properties, including prediction of secondary structure accuracy, solubility and aggregation tendency, calculation of isoelectric point and half-life, and evaluation of structural stability. On the premise of ensuring sequence diversity and retention of functional motifs, the final polypeptide sequence is screened out through conformational space sampling. The final polypeptide sequence has the same structural features and functional properties as the KAZAL region.

[0118] Aiming at the interaction between the KAZAL region and the ADAM10 protein, a 30 amino acid-long polypeptide was designed to simulate the structure and function of the KAZAL region. The specific implementation is as follows:

[0119] First, molecular docking of the KAZAL region and the ADAM10 protein is performed. The three-dimensional structural model of the KAZAL region is docked with the structural model of the ADAM10 protein using molecular docking software, such as AutoDock Vina, ZDOCK, etc. The key interaction sites and binding modes of the KAZAL region and the ADAM10 protein are determined by analyzing the docking results, such as binding energy, binding sites, etc. For example, suppose the docking results show that the 10th to 20th amino acids of the KAZAL region are tightly bound to the active site of ADAM10, with a binding energy of -8 kcal / mol.

[0120] Next, the interface characteristics of the KAZAL region are analyzed. Molecular dynamics simulation software, such as GROMACS, Amber, etc., is used to simulate the three-dimensional structural model of the KAZAL region and calculate the solvent accessible surface area of ​​each residue. The distribution of polar residues (such as serine and threonine) and hydrophobic residues (such as valine and leucine), as well as the distribution of secondary structural elements such as α helix and β fold are analyzed. β turns are identified, the stability of α helix is ​​evaluated, and possible disulfide bond sites are predicted. For example, the analysis results show that the 15th to 18th amino acids in the KAZAL region form a β turn, the 5th to 12th amino acids form an α helix, and a disulfide bond may be formed between the 22nd and 28th cysteines. These analysis results will constitute the interaction interface characteristic map of the KAZAL region.

[0121] Then, the structural units are divided according to the characteristic map of the interaction interface. The KAZAL region is divided into different structural units according to the position of the secondary structure boundary, the hydrophobic core region, and the polar interaction network. For example, the KAZAL region is divided into three structural units: the N-terminal region (amino acids 1-10), the core region (amino acids 11-20), and the C-terminal region (amino acids 21-30). By calculating the folding free energy of peptides of different lengths, the optimal peptide length is determined to be 30 amino acids. For example, the simulation results show that a peptide with a length of 30 amino acids has the lowest folding free energy, indicating that its structure is the most stable.

[0122] Subsequently, candidate peptides were designed based on the results of the structural unit division. Amino acid tendency analysis, charge distribution optimization, and hydrophobicity balance were performed on each structural unit to design multiple candidate peptides. For example, based on the amino acid composition and hydrophobicity characteristics of the core region, three candidate peptides were designed: peptide 1 (sequence: XXXXXXX), peptide 2 (sequence: YYYYYYY), and peptide 3 (sequence: ZZZZZZZZ).

[0123] Finally, the properties of the candidate peptides were evaluated and optimized. Online tools and software, such as PSIPRED and Aggrescan, were used to predict the secondary structure, solubility, and aggregation tendency of the candidate peptides. The isoelectric point and half-life were calculated, and the structural stability was evaluated. Under the premise of ensuring sequence diversity and functional motif retention, conformational space sampling was performed through molecular dynamics simulation to screen out the final peptide sequence. For example, the simulation results showed that peptide 2 had the best solubility, the lowest aggregation tendency, and a secondary structure similar to the KAZAL region, so it was selected as the final peptide sequence.

[0124] The beneficial effects of this method are reflected in three aspects:

[0125] 1. Improve design efficiency: By analyzing the structural characteristics of the KAZAL region, the peptide design is guided, avoiding blind screening and thus improving the design efficiency.

[0126] 2. Ensure the quality of polypeptides: Through the evaluation and optimization of polypeptide properties, the final polypeptide sequence is ensured to have good solubility, stability and biological activity, thus improving the quality of polypeptides.

[0127] 3. Simulation of KAZAL function: The designed 30-amino acid-length peptide simulates the structural and functional characteristics of the KAZAL region and is expected to be used to inhibit the activity of ADAM10 protein, with potential medicinal value.

[0128] In an optional embodiment, a plurality of peptides are respectively subjected to a third round of molecular docking with the ADAM10 protein, and the peptide sequence with the highest score is selected as a candidate polypeptide based on the docking score, including:

[0129] The PEP-FOLD algorithm was used to predict the secondary structure of the peptide sequence, and the initial conformation of the peptide was obtained through energy minimization optimization. The catalytic active site of the ADAM10 protein was analyzed, and the area around the active site was set as the binding pocket through three-dimensional grid point volume calculation to generate the spatial range and position information of the binding pocket.

[0130] Molecular docking parameters are set based on the spatial range and position information of the binding pocket, and a multidimensional scoring function is constructed, wherein the multidimensional scoring function includes shape complementarity energy, electrostatic interaction energy, hydrogen bond energy, and hydrophobic interaction energy;

[0131] The initial conformation of the polypeptide is molecularly docked with the binding pocket of the ADAM10 protein, and the binding energy, residue pair tendency, secondary structure matching degree, and interface area of ​​the polypeptide and the ADAM10 protein are calculated using the multidimensional scoring function to obtain an initial score of the docking complex;

[0132] Normalizing the initial scores of the docked complex, obtaining a comprehensive score by weighted summation, and calculating the interface contact score based on the ratio of the distance between the atomic pairs to the standard contact distance to generate score data for the peptide-protein complex;

[0133] Based on the scoring data of the polypeptide-protein complex, a molecular dynamics simulation is performed on the complex, conformational entropy is calculated through conformational probability distribution, the structural stability of the complex is evaluated, and conformational stability data is obtained;

[0134] The comprehensive score, interface contact score, and conformational stability data are input into a multi-objective optimization model, and the Pareto optimal solution analysis method is used to optimize and screen the polypeptide sequence to obtain the optimal polypeptide sequence.

[0135] A method for designing high-affinity peptide inhibitors for ADAM10 protein, through technical means such as peptide sequence design, molecular docking, molecular dynamics simulation and multi-objective optimization, to screen out peptide sequences with high affinity and stable binding ability to ADAM10 protein.

[0136] First, peptide sequences that interact with ADAM10 protein in known databases or literature reports are used as starting sequences, or multiple candidate peptide sequences are designed based on the structural characteristics of the active site of ADAM10 protein. For example, known inhibitory peptides against ADAM10 protein in the peptide library can be selected, or new peptide sequences can be designed based on the amino acid composition and spatial structure of the binding pocket of ADAM10 protein, such as sequence A: GYRYSDW, sequence B: VQMMPL, sequence C: IFRGYP, etc.

[0137] Next, the PEP-FOLD algorithm is used to predict the secondary structure of the candidate peptide. The peptide sequence is input into the PEP-FOLD server, and the appropriate solvent environment and pH parameters are selected to predict the secondary structure. For example, for sequence A: GYRYSDW, the prediction results show that it is mainly composed of α-helix and random coil. Then, the predicted secondary structure is energy minimized to obtain the initial conformation of the peptide. Molecular dynamics software packages such as GROMACS can be used to minimize the energy and obtain the conformation with the lowest energy.

[0138] Subsequently, the catalytic active site of the ADAM10 protein was analyzed to determine the spatial range and position information of the binding pocket. The resolved ADAM10 protein crystal structure in the PDB database can be used in combination with the active site information reported in the literature to determine the key amino acid residues of the binding pocket. For example, it can be determined that residues such as His405, Glu384, and Met381 constitute the catalytic active center of the ADAM10 protein. Then, with these key residues as the center, a suitable distance threshold (e.g. ), all atoms within the threshold range are defined as binding pockets. The three-dimensional structure of the binding pocket can be calculated and displayed using molecular visualization software such as PyMOL.

[0139] Based on the spatial range and position information of the binding pocket, set the molecular docking parameters. Select appropriate molecular docking software, such as AutoDock Vina, and set the docking parameters, including the center coordinates, size, and search range of the binding pocket. Construct a multidimensional scoring function that includes multiple energy terms such as shape complementarity energy, electrostatic interaction energy, hydrogen bond energy, and hydrophobic interaction energy to evaluate the binding strength of the peptide to the ADAM10 protein.

[0140] The initial conformation of the peptide was molecularly docked with the binding pocket of the ADAM10 protein. After the docking was completed, the multidimensional scoring function was used to calculate the binding energy, residue pair preference, secondary structure matching degree and interface area between the peptide and the ADAM10 protein to obtain the initial score of the docking complex. For example, after the sequence A was docked with the ADAM10 protein, the binding energy was -8.5 kcal / mol, and the residue pair preference showed that the Tyr and Asp residues had a strong binding effect with the ADAM10 protein.

[0141] The initial scores of the docked complexes are normalized and the comprehensive scores are obtained by weighted summation. For example, parameters such as binding energy, residue pair preference, secondary structure matching, and interface area can be normalized and assigned different weights, and then weighted summed to obtain the comprehensive score. For example, the comprehensive score of sequence A is 0.85. At the same time, the interface contact score is calculated based on the ratio of the distance between the atomic pairs to the standard contact distance to generate the score data of the peptide-protein complex.

[0142] Molecular dynamics simulations are performed on the complex to evaluate the structural stability of the complex. Molecular dynamics simulations of the peptide-protein complex are performed using molecular dynamics software packages such as GROMACS, with a simulation time of typically 100 ns. The conformational entropy is calculated from the conformational probability distribution during the simulation to evaluate the structural stability of the complex. For example, the conformational entropy of sequence A is 0.6.

[0143] The comprehensive score, interface contact score and conformational stability data are input into the multi-objective optimization model. The Pareto optimal solution analysis method is used to optimize and screen the peptide sequence to obtain the optimal peptide sequence. For example, by comparing the comprehensive scores, interface contact scores and conformational stability data of different peptide segments, the peptide sequence with the best comprehensive performance is screened out.

[0144] Beneficial effects:

[0145] 1. Improve the affinity of peptide binding to target protein: Through multi-objective optimization strategy, peptide sequences with higher affinity to ADAM10 protein can be screened to enhance its inhibitory effect.

[0146] 2. Improve the stability of peptide-protein complexes: Molecular dynamics simulation and conformational entropy analysis can evaluate the stability of the complex, thereby screening out peptide sequences with more stable binding ability.

[0147] 3. Optimize the physicochemical properties of peptides: By optimizing the peptide sequence, its physicochemical properties, such as solubility, stability and bioavailability, can be improved, making it more suitable as a candidate molecule for drug development.

[0148] In an optional embodiment, the PEP-FOLD algorithm is used to predict the secondary structure of the polypeptide sequence, and the initial conformation of the polypeptide is obtained by energy minimization optimization; the catalytic active site of the ADAM10 protein is analyzed, and the area around the active site is set as a binding pocket by three-dimensional grid point volume calculation, and the spatial range and position information of the binding pocket are generated, including:

[0149] The peptide sequence is divided into fragments based on the physical and chemical properties of amino acids, and the dihedral angle tendency of each fragment is calculated by the frequency of dihedral angles in the database to obtain the fragment dihedral angle probability distribution;

[0150] According to the probability distribution of dihedral angles of the fragments, the α-helix tendency and β-folding tendency of each amino acid site in the polypeptide sequence are calculated, and the secondary structure tendency score is obtained by weighting by the position weight factor;

[0151] Based on the secondary structure propensity score, an initial conformation of the polypeptide is constructed, and a potential energy function including bond length energy, bond angle energy, torsion energy, van der Waals interaction energy, and electrostatic interaction energy is used to optimize the conformation, and the optimal conformation of the polypeptide is obtained by iterative calculation using the BFGS algorithm until the energy gradient is less than a preset convergence threshold;

[0152] The ADAM10 protein sequence was subjected to conservation analysis, highly conserved sites were identified by calculating sequence entropy, and the catalytically active regions were determined by combining the distribution of secondary structure elements with surface exposure analysis;

[0153] Taking the catalytically active region as the center, the three-dimensional space is divided into a regular grid with a side length of 0.5 angstroms, and the distance between each grid point and the surface atom of the protein is calculated to determine whether the grid point belongs to the binding pocket, and all grid points belonging to the binding pocket are volume-integrated to obtain the volume data of the binding pocket;

[0154] Calculating the pocket depth index based on the volume data of the binding pocket, characterizing the spatial characteristics of the pocket by the ratio of the pocket opening area to the volume, and analyzing the hydrophobicity distribution in the pocket;

[0155] The Poisson-Boltzmann equation is used to calculate the electrostatic potential distribution in the binding pocket area, and the electrostatic potential field is constructed based on the atomic charge and dielectric constant to obtain the potential distribution map of the binding pocket.

[0156] The optimal conformation of the polypeptide, binding pocket volume data, pocket depth index, and electrostatic potential distribution map are integrated into a structural feature descriptor to guide the subsequent molecular docking process. The structural feature descriptor contains spatial matching information and interaction characteristics between the polypeptide and the binding site.

[0157] A method for predicting peptide-protein binding sites and constructing structural feature descriptors to guide the molecular docking process. This method is based on peptide secondary structure prediction and protein binding pocket analysis to provide spatial matching information and interaction characteristics between peptides and binding sites.

[0158] First, the secondary structure of the polypeptide sequence is predicted. The polypeptide sequence is divided into several fragments based on the physicochemical properties of amino acids (e.g., hydrophobicity, hydrophilicity, charge, etc.). For example, a 20-amino acid polypeptide is divided into 5 fragments based on hydrophobicity. Then, a protein database, such as the PDB database, is searched, and the dihedral angle (φ and ψ angle) distribution frequency of fragments in the database that have similar amino acid composition to each fragment is counted. For example, fragment 1 is composed of four amino acids Ala-Gly-Leu-Val. All fragments composed of these four amino acids are searched in the database, and their dihedral angle distribution frequency is counted. In this way, the dihedral angle probability distribution of each fragment can be obtained. Next, according to the dihedral angle probability distribution of each fragment, the α-helix tendency and β-folding tendency of each amino acid in the polypeptide sequence are calculated. For example, if the dihedral angle distribution of a certain amino acid is concentrated in the α-helix region, its α-helix tendency is higher. Then, different weight factors are assigned according to the position of the amino acid in the sequence. For example, the amino acid in the middle of the sequence has a higher weight, and the amino acids at both ends of the sequence have a lower weight. The α-helix tendency, β-fold tendency and position weight factor are multiplied to obtain the secondary structure tendency score of each amino acid. Finally, the initial conformation of the polypeptide is constructed based on the secondary structure tendency score. For example, the secondary structure type with the highest score will be used to construct the initial conformation of the amino acid. The constructed initial conformation is optimized using a potential energy function including bond length energy, bond angle energy, torsion energy, van der Waals interaction energy, and electrostatic interaction energy. The BFGS algorithm is used to iterate the calculation until the energy gradient is less than the preset convergence threshold (for example, ) to obtain the optimal conformation of the polypeptide.

[0159] Secondly, the catalytic active site of the target protein (such as ADAM10 protein) is analyzed and a binding pocket is constructed. The ADAM10 protein sequence is subjected to conservation analysis. By calculating the sequence entropy, highly conserved amino acid sites are identified. For example, sites with an entropy value of less than 0.5 are considered to be highly conserved. The catalytic active region is determined by combining the distribution of secondary structure elements (such as α helix, β fold, etc.) and surface exposure analysis. For example, highly conserved sites located on the surface of the protein and in the loop region may be involved in catalytic activity. With the determined catalytic active region as the center, the three-dimensional space is divided into a regular grid with a side length of 0.5 angstroms. The distance from each grid point to the surface atom of the protein is calculated. If the distance is less than a certain threshold (such as 4 angstroms), the grid point is considered to belong to the binding pocket. Volume integration is performed on all grid points belonging to the binding pocket, for example, the volume of each grid point (0.5 angstroms to the cube) is added to obtain the volume data of the binding pocket, such as 150 cubic angstroms. Based on the volume data of the binding pocket, the pocket depth index is calculated. For example, the ratio of the pocket opening area to the volume is used as the pocket depth index, such as 0.2. Analyze the hydrophobicity distribution in the pocket. For example, count the number and proportion of hydrophobic amino acid residues in the binding pocket. Use the Poisson-Boltzmann equation to calculate the electrostatic potential distribution in the binding pocket region. Construct an electrostatic potential field based on atomic charge and dielectric constant. For example, divide the charge of each atom by its distance from the grid point and consider the dielectric constant of the solvent to calculate the electrostatic potential of each grid point. Finally, obtain the potential distribution map of the binding pocket.

[0160] Finally, the optimal conformation of the peptide, the binding pocket volume data (e.g., 150 cubic angstroms), the pocket depth index (e.g., 0.2), and the electrostatic potential distribution map are integrated into a structural feature descriptor. This descriptor contains the spatial matching information and interaction characteristics between the peptide and the binding site, which can be used to guide the subsequent molecular docking process. For example, preliminary screening can be performed based on the conformation of the peptide and the shape of the binding pocket, and then fine docking can be performed based on the electrostatic potential distribution and hydrophobicity distribution.

[0161] The beneficial effects of this method are reflected in three aspects:

[0162] First, the accuracy of peptide secondary structure prediction is improved. By combining the physicochemical properties of amino acids, fragment dihedral angle preferences, and position weight factors, the secondary structure of peptides can be predicted more accurately, thereby constructing a more reasonable initial conformation.

[0163] Second, it provides more comprehensive information about the binding pocket. By calculating the volume, depth index, hydrophobicity distribution, and electrostatic potential distribution of the binding pocket, the characteristics of the binding pocket can be more comprehensively described, thereby better guiding molecular docking.

[0164] Third, the efficiency and accuracy of molecular docking are improved. By integrating the peptide conformation and binding pocket features as structural feature descriptors, potential binding sites can be screened more effectively and the accuracy of molecular docking can be improved.

[0165] In an optional embodiment, the candidate polypeptide is subjected to species homology comparison and quality control detection, wherein the quality control detection includes polypeptide content detection, high performance liquid chromatography detection and mass spectrometry detection including:

[0166] Obtain the observed frequency of amino acid pairs and the background frequency of single amino acids of the candidate polypeptides through BLAST sequence alignment, calculate the amino acid pair substitution scoring matrix, analyze the sequence consistency based on the scoring matrix, and obtain the ratio of the number of matching sites to the alignment length;

[0167] Performing interspecies conservation analysis based on the sequence alignment results, calculating site conservation scores by amino acid frequency, constructing a phylogenetic tree, and obtaining genetic distance data for candidate polypeptides;

[0168] The candidate polypeptide is subjected to acid hydrolysis treatment, a working curve is established using a standard amino acid solution, the content of amino acid components in the hydrolyzate is determined, and the polypeptide recovery rate is calculated;

[0169] Based on the amino acid component content results, the absorbance at 280 nanometers of the candidate polypeptide is measured by ultraviolet spectrophotometry, and the polypeptide concentration is calculated by combining the extinction coefficient and the optical path length to obtain polypeptide content data;

[0170] Reversed-phase high performance liquid chromatography was used to separate and analyze the candidate peptides, the chromatographic conditions were optimized by theoretical plate number calculation, the separation degree was calculated based on retention time and peak width, and the chromatographic spectrum of the candidate peptides was obtained;

[0171] The peak area of ​​each chromatographic peak in the chromatogram is compared with the total peak area, the purity of the target peak is calculated, the symmetry of the peak shape is evaluated by the tailing factor, and the purity data of the peptide is obtained;

[0172] Electrospray ionization mass spectrometry was used to determine the mass-to-charge ratio of candidate peptides, analyze isotope distribution patterns, and obtain fragment ion spectra through tandem mass spectrometry to identify peptide sequences and modification sites;

[0173] Establish quality control indicators including main peak content and single impurity content limits, conduct accelerated tests and solution stability tests on candidate peptides, and evaluate the stability of candidate peptides under different pH and temperature conditions;

[0174] Based on the sequence comparison results, polypeptide content data, polypeptide purity data, mass spectrometry analysis results and stability data, a comprehensive polypeptide quality evaluation system is established to screen candidate polypeptides that meet quality control requirements.

[0175] The species homology comparison and quality control detection method of the candidate peptides is specifically implemented as follows:

[0176] First, perform a species homology comparison. Use the BLAST tool to compare the candidate polypeptide sequence with the homologous sequences of other species in the database. Count the number of occurrences of each pair of amino acids in the comparison results and compare them with the background frequency of a single amino acid to obtain a scoring matrix for amino acid pair substitutions. For example, if the number of times alanine and valine appear in a certain polypeptide in the comparison results is much higher than the background frequency of their individual occurrences, the scores for alanine and valine substitutions will be higher. Based on this scoring matrix, analyze the similarity between the candidate polypeptide and the homologous sequences of other species, and calculate the ratio of the number of matched amino acid sites to the total length of the comparison as an indicator of sequence consistency. Assuming that the length of the comparison between a candidate polypeptide and a homologous sequence of a species is 100 amino acids, and 80 amino acid sites match, the sequence consistency is 80%. Analyze the conservation between species based on the sequence comparison results. Count the frequency of occurrence of different amino acids at each site and calculate the site conservation score. For example, if 90% of the species at a site are glycine, the conservation score of the site is high. According to the conservation score of each site, a phylogenetic tree was constructed, and the genetic distance between the candidate peptide and other species was calculated.

[0177] Next, the polypeptide content is detected. The candidate polypeptide is subjected to acid hydrolysis to decompose the polypeptide into amino acids. A working curve is established using a standard amino acid solution, and the content of various amino acids in the hydrolyzate is determined by high-performance liquid chromatography. For example, a standard alanine solution of known concentration is gradiently diluted to obtain a series of alanine solutions of different concentrations, and its peak area in high-performance liquid chromatography is measured. The peak area is plotted against the concentration to obtain a working curve for alanine. The candidate polypeptide hydrolyzate is then subjected to high-performance liquid chromatography analysis, and the content of alanine in the hydrolyzate is calculated based on the working curve. The polypeptide recovery rate is calculated by calculating the total content of all amino acids and comparing it with the theoretical molecular weight of the candidate polypeptide. Assuming that the theoretical molecular weight of a candidate polypeptide is 1000Da, and the total amino acid content measured after hydrolysis is equivalent to 900Da, the polypeptide recovery rate is 90%. According to the results of the amino acid component content, the absorbance of the candidate polypeptide at 280 nanometers is determined by ultraviolet spectrophotometry. Combined with the known extinction coefficient and optical path, the polypeptide concentration is calculated to obtain the polypeptide content data. For example, if the absorbance of a candidate polypeptide at 280 nm is measured to be 0.5, and its extinction coefficient is known to be 10000M-1cm-1, and its optical path length is 1 cm, then its concentration is 50 μM.

[0178] Then, high performance liquid chromatography was performed for detection. The candidate peptides were separated and analyzed using reversed phase high performance liquid chromatography. The chromatographic conditions were optimized by adjusting the composition, flow rate, column temperature and other parameters of the mobile phase so that the candidate peptides could be effectively separated from other impurities. The separation efficiency of the chromatographic column was evaluated using the theoretical plate number as an indicator. The separation degree was calculated based on the retention time and peak width of the candidate peptide to evaluate the separation effect of the candidate peptide from other impurities. The chromatogram of the candidate peptide was obtained. The peak area of ​​each chromatographic peak in the chromatogram was compared with the total peak area to calculate the purity of the target peak. For example, if the peak area of ​​the candidate peptide accounts for 98% of the total peak area, its purity is 98%. The symmetry of the peak shape is evaluated by the tailing factor. The closer the tailing factor is to 1, the more symmetrical the peak shape is.

[0179] Next, mass spectrometry is performed. Electrospray ionization mass spectrometry is used to determine the mass-to-charge ratio of the candidate peptide. The isotope distribution spectrum is analyzed to confirm the molecular weight of the candidate peptide. Through tandem mass spectrometry technology, the fragment ion spectrum of the candidate peptide is obtained to identify the peptide sequence and possible modification sites. For example, by analyzing the difference in the mass-to-charge ratio of the fragment ions, the amino acid sequence can be inferred.

[0180] Finally, a stability test is conducted. Quality control indicators including the main peak content and the single impurity content limit are established. Accelerated tests and solution stability tests are conducted on candidate peptides to evaluate the stability of candidate peptides under different pH values ​​and temperature conditions. For example, the candidate peptide solution is placed under a high temperature of 40°C, and samples are taken regularly to detect changes in its content and purity to evaluate its high temperature stability.

[0181] Based on species homology comparison results, peptide content data, peptide purity data, mass spectrometry analysis results and stability data, a comprehensive peptide quality evaluation system was established to screen candidate peptides that met quality control requirements.

[0182] Beneficial effects:

[0183] 1. High accuracy: This method ensures the accuracy and reliability of the quality assessment results of candidate peptides through multiple means such as peptide content detection, high performance liquid chromatography detection and mass spectrometry detection.

[0184] 2. Comprehensiveness: This method covers multiple aspects such as species homology comparison, peptide content determination, purity analysis, sequence identification and stability evaluation, achieving comprehensive quality control of candidate peptides.

[0185] 3. High efficiency: This method can quickly screen out candidate peptides that meet quality requirements by establishing a comprehensive peptide quality evaluation system, thereby improving research and development efficiency.

[0186] This application also provides a variety of specific embodiments:

[0187] [Implementation Case 1] Screening of active fragment small molecule peptides based on target protein sequence;

[0188] 1. Experimental materials;

[0189] Target protein-K1-4, human umbilical vein endothelial cells, DMEM-basic (Gibco), fetal bovine

[0190] Serum (Zeta), Cell Counting Kit-8 (CCK-8, Seven)

[0191] 2. Experimental methods;

[0192] 1. Preparation of target protein-K1-4 peptide solution;

[0193] The peptide lyophilized powder was dissolved in DMSO, the storage concentration was 30 mg / mL, 30 uL was divided into packages, and stored at -80 °C.

[0194] 2. CCK-8 cell activity detection: The cell density is 5x104 / mL and inoculated in a 96-well culture plate, the final density is 5*103 / well, and incubated at 37°C and 5% C02, 95% air. On the second day, the medium was replaced with serum-free medium, and the serum-free medium containing the target protein-K1-4 peptide at concentrations of 0.1, 1, 10, and 100ug / ml was cultured. The incubation was terminated after 12 hours. 10μL CCK-8 solution was added to each well and incubated in the incubator for 1h. The absorbance of each well at 450nm was measured using a multifunctional microplate reader (Agilent / Synergy H1). Six replicate wells were set for each concentration, and the experiment was repeated 3 times.

[0195] 3. Peptide uptake detection: The cells were inoculated in a 12-well culture plate at a density of 1x106 / mL and incubated at 37°C and 5% C02, 95% air. The medium was replaced with serum-free medium on the second day, and the serum-free medium containing the target protein-K1-4 peptide at a concentration of 2ug / ml was added for culture. The FITC fluorescence intensity of the cells was detected by flow cytometry after 0, 0.5, 1, 2, 4, 6, 12, and 18h ​​of culture. Three replicate wells were set for each concentration, and the experiment was repeated 3 times.

[0196] Figure 2 Schematic diagram showing that RECK K1-4 polypeptide does not produce cytotoxicity when cultured with human umbilical vein endothelial cells (HUVEC); Figure 3This is a schematic diagram showing that RECK K1-4 peptides can be well taken up by endothelial cells. Molecular docking and structural prediction were performed using multiple platforms such as AlphaFold and HDOCK, combined with PyMOL visualization analysis, which can accurately locate the key domains of the target protein; the GMQE and QMEAN scoring systems of the swiss-model platform can reliably evaluate the accuracy of structural prediction; a three-round progressive molecular docking strategy was used, from the overall protein to the functional domain to the active peptide segment, to gradually optimize the screening process and improve the accuracy of the active fragment.

[0197] By combining various analytical methods such as elemental analyzer, high performance liquid chromatography and mass spectrometry, the physical and chemical properties of the peptides are fully characterized; the evolutionary conservation of the obtained peptide sequences is verified through multi-species sequence homology analysis; and a standardized peptide synthesis and quality control process is established to ensure product stability and batch consistency.

[0198] CCK-8 cell activity assay confirmed that the polypeptide had no obvious cytotoxicity to human umbilical vein endothelial cells in the concentration range of 0.1-100 μg / ml; flow cytometry monitoring confirmed that the polypeptide could be effectively taken up by endothelial cells and showed time-dependent uptake kinetic characteristics; the experimental design adopted multiple concentration gradients and time points, and set appropriate repetitions to ensure the reliability and scientificity of the experimental results.

[0199] [Implementation Case 2] The target protein K1-4 peptide was cultured with human umbilical vein endothelial cells (HUVEC) to significantly reduce the protein level of the screening protein

[0200] 1. Experimental materials;

[0201] Target protein-K1-4, human umbilical vein endothelial cells (Sciencell, USA), DMEM-basic (Gibco), fetal bovine serum (Zeta), anti-screening protein (Abcam), anti-β-actin (Santa crutz), horseradish peroxidase-labeled goat anti-mouse IgG (Beyotime), horseradish peroxidase-labeled goat anti-rabbit IgG (Beyotime), luminescent liquid (ZETA);

[0202] 2. Experimental methods;

[0203] 1. Preparation of target protein-K1-4 peptide solution;

[0204] The peptide lyophilized powder was dissolved in DMSO, the storage concentration was 30 mg / mL, 30 uL was divided into packages, and stored at -80 °C.

[0205] 2. Screening protein protein level detection: The cells were inoculated in a 12-well culture plate at a density of 1x106 / mL and incubated at 37°C and 5% C02, 95% air. The medium was replaced with serum-free medium on the second day, and ① the serum-free medium containing the target protein-K1-4 polypeptide at a concentration of 0.1, 1, 10, 100ug / ml was used for culture, and the incubation was terminated after 12 hours; ② the serum-free medium containing the target protein-K1-4 polypeptide at a concentration of 2ug / ml was used for culture, and the total cell protein was extracted and collected by lysing HUVE with Western-IP lysis buffer after 0, 1, 2, 4, 6, 12, 18, 24h.

[0206] 3. Western blot protein blotting detection: The above-collected corresponding protein concentration was determined using a BCA protein content assay kit, and an equal amount (50ug) of protein lysate was transferred to a polyvinylidene fluoride (PVDF) membrane after 12% sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE) electrophoresis. Block with 5% skim milk at room temperature for 2 hours, and incubate with the corresponding mouse anti-human / rabbit anti-mouse primary antibody (anti-screening protein, β-actin antibody) at 4°C for overnight hybridization. The next day, rinse three times with TBS buffer containing 0.05% Tween 20, and further incubate with horseradish peroxidase (HRP)-labeled goat anti-mouse / goat anti-rabbit secondary antibody at room temperature for 1 hour. After sufficient rinsing, the immunoblot bands were exposed and developed using a luminescent visualization agent. Image J software was used to perform density determination and quantification of the immunoblot bands. Western blot experiments were performed to detect the protein level of the screened protein. Four replicate wells were set for each concentration, and the experiment was repeated 3 times.

[0207] Figure 4 Schematic diagram of the time-dependent effect of RECK K1-4 polypeptide on ADAM10 protein level in human umbilical vein endothelial cells; Western blot technology combined with Image J quantitative analysis can accurately detect the regulatory effect of target protein K1-4 polypeptide on the expression level of screened proteins; β-actin is used as an internal reference to ensure the accuracy and reliability of protein quantitative analysis; the experimental design includes multiple technical replicates and biological replicates to ensure the scientificity and repeatability of the experimental results.

[0208] By setting a concentration gradient of 0.1-100 μg / ml, the dose-effect relationship of the target protein K1-4 peptide on the expression of the screened protein was systematically evaluated; the experiment adopted a design of multiple concentration points, which can comprehensively reflect the action characteristics of the peptide and provide a reliable concentration reference for subsequent applications; a standardized peptide storage and use plan was established to ensure the stability of the peptide during the experiment.

[0209] [Implementation Case 3] Intravenous injection of K1-4 peptide can improve inflammatory response and organ damage in sepsis;

[0210] 1. Experimental materials;

[0211] Target protein-K1-4 (GenScript Biotechnology), C57BL / 6J mice;

[0212] 2. Experimental methods;

[0213] 1. Preparation of target protein-K1-4 peptide solution

[0214] 0.5% saline is used to prepare a 10 mg / ml solution, which is then equilibrated to room temperature and used immediately.

[0215] 2. Establishment of mouse sepsis model: Wild-type C57BL / 6 mice were randomly divided into 3 groups for CLP and sham surgery. Fasting was performed for 12 hours before surgery, but water was not allowed. Ketamine (100 mg / kg) and xylazine (5 mg / kg) were injected intraperitoneally. After satisfactory anesthesia, a 1.0 cm laparotomy was performed on the midline of the abdomen. The cecum was found in the left lower abdomen and removed to expose the cecum. The cecum was then ligated 0.5 cm from the top just below the ileocecal valve and punctured back and forth once with a No. 23 prism needle (2.6×65 mm). Gently squeeze to allow a drop of feces to flow out. The cecum was carefully placed back into the abdominal cavity and the incision was sutured. Mice in the sham group (Sham group) underwent the same surgical operation, but no cecal ligation and puncture were performed. In order to maintain stable hemodynamic conditions in mice after surgery, each mouse was given resuscitation fluid (0.9% NaCl, 37°C, 50 mL / kg, subcutaneous injection), and food and water were freely available after anesthesia recovery. After surgery, the mice were closely observed for sepsis-related symptoms such as shortness of breath, diarrhea, piloerection, decreased muscle tension, and poor mental responsiveness. The mice were euthanized 12 h after modeling, and lung and kidney tissues were collected. Blood samples were centrifuged (room temperature, 3000 rpm / min, 10 min) to separate serum, and the samples were stored at -80°C for subsequent experiments.

[0216] 3. Tail vein injection of target protein-K1-4 peptide solution: 0.5 hours after surgery, mice were injected with a total amount of 10 mg / kg peptide solution through the tail vein.

[0217] 4. Detection of IL-6 and TNFα levels in mouse serum:

[0218] Figure 5Schematic diagram of intravenous injection of RECK K1-4 peptide to reduce the level of circulating inflammatory factors in septic mice; Enzyme-linked immunosorbent assay (ELISA) test was performed on mouse serum according to the following procedures: (1) Take out the kit from the 4℃ refrigerator, mix all reagents and samples thoroughly, and equilibrate to room temperature; (2) Estimate the sample concentration and dilute it with sample diluent according to the proportion in the instruction manual; prepare biotin-labeled antibody working solution (antibody: antibody diluent = 1:99) and avidin-peroxidase complex (ABC) working solution (ABC: ABC diluent = 1:99) in proportion one hour before use. (3) Prepare standard: Gently centrifuge the standard powder for 1 minute, slowly add the specified volume of sample diluent and vortex to mix, and prepare the corresponding concentration standard solution according to the instruction manual after fully dissolving. (4) Sample addition: The ELISA plate is equipped with standard wells, blank wells, and sample wells, among which 100ul sample diluent is added to the blank well, and equal amounts of standard solution and sample solution are added to the remaining wells, trying to avoid touching the bottom and wall of the well. The ELISA plate was sealed with a membrane and incubated at 37°C for 90 min; (5) the liquid in (4) was shaken off, without washing, 100 μl of the prepared antibody working solution was added to each well, the ELISA plate was covered with a membrane and incubated at 37°C for 60 min; (6) the liquid in (5) was shaken off, 300 μl of 1× buffered washing solution was added to each well, soaked for 1 min, and shaken off, and the washing was repeated 3 times; (7) the liquid in (6) was shaken off, 100 μl of the prepared ABC working solution was added to each well, the ELISA plate was covered with a membrane and incubated at 37°C for 60 min; (8) the liquid in (7) was shaken off, and at least 300 μl of 1× buffered washing solution was added to each well. Soak in 1× buffered washing solution for 2 min and shake off, repeat washing 5 times; (9) Shake off the liquid in (8), add 90 μl of TMB colorimetric solution to each well in the dark, cover the ELISA plate and incubate at 37°C in the dark for 60 min; (10) Take out the plate in the dark, quickly add 100 μl of reaction stop solution to each well, and immediately use a multifunctional microplate reader to detect the OD value of each well at a wavelength of 450 nm; (11) Data calculation: Draw a standard linear regression equation curve, with the ordinate being the OD value of the standard and the abscissa being the concentration of the standard corresponding to the OD value. A correlation coefficient R value greater than or equal to 0.9900 indicates that the curve is available. Then, based on the obtained sample OD value, use the calculated standard linear regression equation and the sample dilution multiple to calculate the actual concentration of the sample.

[0219] 4. Observation of pathological changes in mouse lung tissue by hematoxylin-eosin (HE) staining:

[0220] The mouse tissue samples obtained were fixed with 4% paraformaldehyde, dehydrated with gradient alcohol, embedded in paraffin, sliced ​​(5μm), stained with HE, and sealed. A digital pathology section scanner was used to scan the tissue pathology sections, collect images, and score tissue damage. Lung tissue damage was evaluated by pulmonary edema, hyaline membrane, hemorrhage, atelectasis, inflammatory cell infiltration, etc. Renal tissue damage was evaluated by cellular vacuolar degeneration, tubular epithelial edema, tubular atrophy, necrosis, and shedding. In order to determine the severity of the lesions, all tissue analyses used a scoring system of 1 to 4 points. All scores were added up according to the area percentage of each area to calculate the total score of the tissue pathology.

[0221] A stable and reliable sepsis mouse model was successfully established through a standardized cecal ligation and puncture (CLP) surgical protocol. Multiple clinical performance indicators (respiration, diarrhea, mental state, etc.) were used to comprehensively evaluate the effectiveness of the model establishment. During the experiment, surgical conditions and postoperative care measures were strictly controlled to ensure the repeatability and reliability of the experimental results.

[0222] The standardized ELISA method was used to achieve accurate quantification of inflammatory factors (IL-6, TNFα) in serum; a complete sample processing and detection process was established, including the preparation of standard curves and data analysis methods; and the combined detection of multiple inflammatory indicators fully reflected the regulatory effect of the target protein K1-4 polypeptide on the inflammatory response of sepsis.

[0223] Through HE staining combined with digital pathological analysis, an objective quantitative evaluation of the degree of organ damage was achieved; a multi-dimensional tissue damage scoring standard was established, including multiple pathological indicators such as pulmonary edema, inflammatory cell infiltration, and cell degeneration; a quantitative scoring system of 1-4 points was used to make the evaluation results more scientific and comparable, providing a reliable basis for evaluating the therapeutic effects of peptides.

[0224] Figure 2 FIG. 1 is a schematic diagram of a structure of a small molecule polypeptide mimetic system based on protein-protein interaction screening according to an embodiment of the present invention. Figure 2 As shown, the system comprises:

[0225] The first unit is used to obtain the three-dimensional structure prediction files of RECK protein and ADAM10 protein, wherein the three-dimensional structure prediction file of RECK protein is AF-O95980-F1, and the three-dimensional structure prediction file of ADAM10 protein is AF-O14672-F1;

[0226] The second unit is used to input the three-dimensional structure prediction files of the RECK protein and the ADAM10 protein into the molecular docking platform for the first round of molecular docking, obtain the interaction model of the RECK protein and the ADAM10 protein, and analyze the interaction site of the RECK protein and the ADAM10 protein by PyMOL visualization software; determine the KAZAL region of the RECK protein according to the interaction site, and use the Swiss-model platform to predict the sequence structure of the KAZAL region; perform a second round of molecular docking on the KAZAL region and the ADAM10 protein, and design multiple peptides with a length of 30 amino acids according to the docking results; perform a third round of molecular docking on the multiple peptides and the ADAM10 protein respectively, and select the peptide sequence with the highest score as the candidate polypeptide based on the docking score;

[0227] The third unit is used to perform species homology comparison and quality control detection on the candidate polypeptide, wherein the quality control detection includes polypeptide content detection, high performance liquid chromatography detection and mass spectrometry detection; verify the binding ability and biological function of the candidate polypeptide to ADAM10 protein through cell experiments, wherein the amino acid sequence of the candidate polypeptide is LPCNCADQFVPVCGQNGRTYPSACIARCVG.

[0228] According to a third aspect of the embodiments of the present invention,

[0229] An electronic device is provided, comprising:

[0230] processor;

[0231] a memory for storing processor-executable instructions;

[0232] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0233] A fourth aspect of the embodiments of the present invention is:

[0234] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0235] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0236] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing small molecule polypeptide mimetics based on protein-protein interaction screening, characterized in that: include: Obtaining three-dimensional structure prediction files of RECK protein and ADAM10 protein, wherein the three-dimensional structure prediction file of RECK protein is AF-O95980-F1, and the three-dimensional structure prediction file of ADAM10 protein is AF-O14672-F1; The three-dimensional structure prediction files of the RECK protein and the ADAM10 protein were input into the molecular docking platform for the first round of molecular docking to obtain the interaction model between the RECK protein and the ADAM10 protein, and the interaction site between the RECK protein and the ADAM10 protein was analyzed by the PyMOL visualization software; the KAZAL region of the RECK protein was determined according to the interaction site, and the sequence structure of the KAZAL region was predicted using the Swiss-model platform; the KAZAL region was subjected to the second round of molecular docking with the ADAM10 protein, and multiple peptides with a length of 30 amino acids were designed according to the docking results; the multiple peptides were respectively subjected to the third round of molecular docking with the ADAM10 protein, and the peptide sequence with the highest score was selected as the candidate polypeptide based on the docking score; The candidate polypeptide is subjected to species homology comparison and quality control detection, wherein the quality control detection includes polypeptide content detection, high performance liquid chromatography detection and mass spectrometry detection; the binding ability and biological function of the candidate polypeptide to ADAM10 protein are verified through cell experiments, wherein the amino acid sequence of the candidate polypeptide is LPCNCADQFVPVCGQNGRTYPSACIARCVG.

2. The method according to claim 1, characterized in that The three-dimensional structure prediction files of the RECK protein and the ADAM10 protein were input into the molecular docking platform for the first round of molecular docking to obtain the interaction model between the RECK protein and the ADAM10 protein, and the interaction sites between the RECK protein and the ADAM10 protein were analyzed by the PyMOL visualization software, including: Pre-processing of the RECK protein structure and the ADAM10 protein structure included removing crystal water molecules from the protein structure, adding hydrogen atoms to the protein structure, optimizing the protein structure for energy minimization, and verifying the rationality of the protein structure conformation through the Ramachandran plot; The pre-processed RECK protein structure and ADAM10 protein structure were input into the molecular docking platform, and the RECK protein structure and the ADAM10 protein structure were rigidly docked based on the fast Fourier transform algorithm and the Monte Carlo local optimization algorithm. The search space range with a radius of 30 angstroms was set with the center of the RECK protein as the center. 1000 conformations were generated each time the docking was performed, and the top 100 conformations were retained according to the shape complementarity score. The flexibility optimization was performed on the top 100 retained conformations, including side chain optimization and local energy minimization of the conformations, and the binding free energy was obtained by calculating the van der Waals interaction energy, electrostatic interaction energy, and desolvation free energy; The conformation after flexibility optimization was imported into PyMOL visualization software, and the interaction interface was analyzed by calculating the interface contact area, identifying the hydrogen bond network, analyzing the hydrophobic interaction and identifying the salt bridge; a 5 angstrom cutoff value was set in the interaction interface to screen the interface residues, calculate the change in the buried area of ​​each residue, analyze the residue conservation, and identify the residues that contribute more than 2 kcal / mol to the binding free energy as hotspot residues; the conformations were clustered based on the root mean square deviation to identify the main binding mode and evaluate the stability of the binding conformation. The root mean square deviation was calculated by the following formula: the square of the coordinate difference between each atom and the corresponding atom in the reference structure was summed, the average was taken, and the square root was taken; the analysis results were automatically processed by the PyMOL script to generate a standardized analysis report containing interface residue information, hotspot residue information and main binding modes.

3. The method according to claim 1, characterized in that: Determining the KAZAL region of the RECK protein according to the interaction site, and using the Swiss-model platform to predict the sequence structure of the KAZAL region includes: A multiple sequence alignment of the RECK protein sequences was performed, and a position-specific scoring matrix was constructed by calculating the logarithmic ratio of the observed frequency of the amino acid at each position to the background frequency to obtain the conservation score of the RECK protein sequence; Based on the conservation score, the domain analysis of RECK protein was performed, the KAZAL motif features were analyzed using InterProScan, the domain boundaries were verified in combination with the SMART database, the functional regions were annotated using the Pfam database, and the amino acid sequence of the KAZAL region in the RECK protein was determined; Inputting the amino acid sequence of the KAZAL region into the Swiss-model platform, screening structures with sequence similarity greater than 30% and structural resolution less than 2.5 angstroms as modeling templates, evaluating the template quality by calculating the weighted sum of the sequence similarity score, resolution score and R factor score of the template, and selecting the optimal template; Based on the optimal template, structural modeling of the KAZAL region was performed, including sequence alignment optimization, backbone construction, side chain reconstruction, loop region optimization and energy minimization, to obtain a three-dimensional structural model of the KAZAL region.

4. The method according to claim 3, characterized in that The KAZAL region was subjected to a second round of molecular docking with the ADAM10 protein, and multiple peptide segments with a length of 30 amino acids were designed based on the docking results, including: Performing interface feature analysis on the three-dimensional structural model of the KAZAL region, calculating the solvent accessible surface area of ​​each residue, analyzing the distribution of polar residues, hydrophobic residue aggregation and secondary structure elements, identifying β turns, evaluating α helix stability, predicting disulfide bond sites, and obtaining a map of interaction interface features; According to the interaction interface characteristic map, the structural units are divided, the boundary positions of the secondary structures are determined, the hydrophobic core regions are identified, the polar interaction network is analyzed, and the optimal peptide length is determined to be 30 amino acids by calculating the peptide folding free energy; Designing candidate peptides based on the structural unit division results, performing amino acid tendency analysis, charge distribution optimization and hydrophobicity balance on each structural unit, and designing multiple candidate peptides; The candidate peptides are evaluated and optimized for properties, including prediction of secondary structure accuracy, solubility and aggregation tendency, calculation of isoelectric point and half-life, and evaluation of structural stability. On the premise of ensuring sequence diversity and retention of functional motifs, the final polypeptide sequence is screened out through conformational space sampling. The final polypeptide sequence has the same structural features and functional properties as the KAZAL region.

5. The method according to claim 1, characterized in that The third round of molecular docking was performed on multiple peptides with ADAM10 protein, and the peptide sequences with the highest scores were selected as candidate peptides based on the docking scores, including: The PEP-FOLD algorithm was used to predict the secondary structure of the peptide sequence, and the initial conformation of the peptide was obtained through energy minimization optimization. The catalytic active site of the ADAM10 protein was analyzed, and the area around the active site was set as the binding pocket through three-dimensional grid point volume calculation to generate the spatial range and position information of the binding pocket. Molecular docking parameters are set based on the spatial range and position information of the binding pocket, and a multidimensional scoring function is constructed, wherein the multidimensional scoring function includes shape complementarity energy, electrostatic interaction energy, hydrogen bond energy, and hydrophobic interaction energy; The initial conformation of the polypeptide is molecularly docked with the binding pocket of the ADAM10 protein, and the binding energy, residue pair tendency, secondary structure matching degree, and interface area of ​​the polypeptide and the ADAM10 protein are calculated using the multidimensional scoring function to obtain an initial score of the docking complex; Normalizing the initial scores of the docked complex, obtaining a comprehensive score by weighted summation, and calculating the interface contact score based on the ratio of the distance between the atomic pairs to the standard contact distance to generate score data for the peptide-protein complex; Based on the scoring data of the polypeptide-protein complex, a molecular dynamics simulation is performed on the complex, conformational entropy is calculated through conformational probability distribution, the structural stability of the complex is evaluated, and conformational stability data is obtained; The comprehensive score, interface contact score, and conformational stability data are input into a multi-objective optimization model, and the Pareto optimal solution analysis method is used to optimize and screen the polypeptide sequence to obtain the optimal polypeptide sequence.

6. The method according to claim 5, characterized in that The PEP-FOLD algorithm was used to predict the secondary structure of the peptide sequence, and the initial conformation of the peptide was obtained through energy minimization optimization. The catalytic active site of the ADAM10 protein was analyzed, and the area around the active site was set as the binding pocket through three-dimensional grid point volume calculation. The spatial range and position information of the binding pocket were generated, including: The peptide sequence is divided into fragments based on the physical and chemical properties of amino acids, and the dihedral angle tendency of each fragment is calculated by the frequency of dihedral angles in the database to obtain the fragment dihedral angle probability distribution; According to the probability distribution of dihedral angles of the fragments, the α-helix tendency and β-folding tendency of each amino acid site in the polypeptide sequence are calculated, and the secondary structure tendency score is obtained by weighting by the position weight factor; Based on the secondary structure propensity score, an initial conformation of the polypeptide is constructed, and a potential energy function including bond length energy, bond angle energy, torsion energy, van der Waals interaction energy, and electrostatic interaction energy is used to optimize the conformation, and the optimal conformation of the polypeptide is obtained by iterative calculation using the BFGS algorithm until the energy gradient is less than a preset convergence threshold; The ADAM10 protein sequence was subjected to conservation analysis, highly conserved sites were identified by calculating sequence entropy, and the catalytically active regions were determined by combining the distribution of secondary structure elements with surface exposure analysis; Taking the catalytically active region as the center, the three-dimensional space is divided into a regular grid with a side length of 0.5 angstroms, and the distance between each grid point and the surface atom of the protein is calculated to determine whether the grid point belongs to the binding pocket, and all grid points belonging to the binding pocket are volume-integrated to obtain the volume data of the binding pocket; Calculating the pocket depth index based on the volume data of the binding pocket, characterizing the spatial characteristics of the pocket by the ratio of the pocket opening area to the volume, and analyzing the hydrophobicity distribution in the pocket; The Poisson-Boltzmann equation is used to calculate the electrostatic potential distribution in the binding pocket area, and the electrostatic potential field is constructed based on the atomic charge and dielectric constant to obtain the potential distribution map of the binding pocket. The optimal conformation of the polypeptide, binding pocket volume data, pocket depth index, and electrostatic potential distribution map are integrated into a structural feature descriptor to guide the subsequent molecular docking process. The structural feature descriptor contains spatial matching information and interaction characteristics between the polypeptide and the binding site.

7. The method according to claim 1, characterized in that The candidate polypeptides are subjected to species homology comparison and quality control detection, wherein the quality control detection includes polypeptide content detection, high performance liquid chromatography detection and mass spectrometry detection including: Obtain the observed frequency of amino acid pairs and the background frequency of single amino acids of the candidate polypeptides through BLAST sequence alignment, calculate the amino acid pair substitution scoring matrix, analyze the sequence consistency based on the scoring matrix, and obtain the ratio of the number of matching sites to the alignment length; Performing interspecies conservation analysis based on the sequence alignment results, calculating site conservation scores by amino acid frequency, constructing a phylogenetic tree, and obtaining genetic distance data for candidate polypeptides; The candidate polypeptide is subjected to acid hydrolysis treatment, a working curve is established using a standard amino acid solution, the content of amino acid components in the hydrolyzate is determined, and the polypeptide recovery rate is calculated; Based on the amino acid component content results, the absorbance at 280 nanometers of the candidate polypeptide is measured by ultraviolet spectrophotometry, and the polypeptide concentration is calculated by combining the extinction coefficient and the optical path length to obtain polypeptide content data; Reversed-phase high performance liquid chromatography was used to separate and analyze the candidate peptides, the chromatographic conditions were optimized by theoretical plate number calculation, the separation degree was calculated based on retention time and peak width, and the chromatographic spectrum of the candidate peptides was obtained; The peak area of ​​each chromatographic peak in the chromatogram is compared with the total peak area, the purity of the target peak is calculated, the symmetry of the peak shape is evaluated by the tailing factor, and the purity data of the peptide is obtained; Electrospray ionization mass spectrometry was used to determine the mass-to-charge ratio of candidate peptides, analyze isotope distribution patterns, and obtain fragment ion spectra through tandem mass spectrometry to identify peptide sequences and modification sites; Establish quality control indicators including main peak content and single impurity content limits, conduct accelerated tests and solution stability tests on candidate peptides, and evaluate the stability of candidate peptides under different pH and temperature conditions; Based on the sequence comparison results, polypeptide content data, polypeptide purity data, mass spectrometry analysis results and stability data, a comprehensive polypeptide quality evaluation system is established to screen candidate polypeptides that meet quality control requirements.

8. A small molecule peptide mimetic system is constructed based on protein-protein interaction screening, for implementing the method according to any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain the three-dimensional structure prediction files of RECK protein and ADAM10 protein, wherein the three-dimensional structure prediction file of RECK protein is AF-O95980-F1, and the three-dimensional structure prediction file of ADAM10 protein is AF-O14672-F1; The second unit is used to input the three-dimensional structure prediction files of the RECK protein and the ADAM10 protein into the molecular docking platform for the first round of molecular docking, obtain the interaction model of the RECK protein and the ADAM10 protein, and analyze the interaction site of the RECK protein and the ADAM10 protein by PyMOL visualization software; determine the KAZAL region of the RECK protein according to the interaction site, and use the Swiss-model platform to predict the sequence structure of the KAZAL region; perform a second round of molecular docking on the KAZAL region and the ADAM10 protein, and design multiple peptides with a length of 30 amino acids according to the docking results; perform a third round of molecular docking on the multiple peptides and the ADAM10 protein respectively, and select the peptide sequence with the highest score as the candidate polypeptide based on the docking score; The third unit is used to perform species homology comparison and quality control detection on the candidate polypeptide, wherein the quality control detection includes polypeptide content detection, high performance liquid chromatography detection and mass spectrometry detection; the binding ability and biological function of the candidate polypeptide to ADAM10 protein are verified through cell experiments, and the amino acid sequence of the candidate polypeptide is PCNCADQFVPVCGQNGRTYPSACIARCVG.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 8.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Multi-modal deep learning-based MHC presentation peptide fragment prediction method and system

    CN120452555A

  • Rapid verification system and method for docking and RMSD optimization

    CN121583390A

  • Method and system for efficiently screening high-activity isoenzyme

    CN122201503A

  • Method for screening optical coupling sites in protein

    CN122266441A

  • Strain culture temperature-based thermosensitive UDG rational mutation site screening method

    CN122337320A