A virtual decompression method for a cyclic peptide

Through the virtual decompression method of cyclic peptides, the virtual screening technology is used to decompress the cyclic peptides in the cyclic peptide library, solving the problems of low efficiency and high cost of traditional methods, and achieving efficient peptide screening and drug screening.

CN117831643BActive Publication Date: 2025-06-20HUNAN ZONSEN PEPLIB BIOTECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202311864121.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-06-20
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

The traditional decompression method of cyclic peptides in the prior art cyclic peptide library is inefficient and costly.

Method used

Using the virtual decompression method of cyclic peptides, the fragmentation of cyclic peptides, virtual screening of polypeptide molecular library and result analysis, the better polypeptide molecules were screened out, and the virtual screening was completed using molecular docking technology.

Benefits of technology

It greatly shortens the experimental time, saves experimental costs, avoids interference from artificial errors and uncontrollable factors, realizes high-throughput design of specific identification elements, and improves the efficiency of drug screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117831643B_ABST
    Figure CN117831643B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of bioinformatics technology, and particularly relates to a method for virtual decompression of cyclic peptides. The method disassembles cyclic peptides into polypeptide molecules of 6 to 40 amino acids to create a polypeptide molecule library, and performs virtual screening on the polypeptide molecule library. Using the method provided by the present invention, polypeptide molecules disassembled from cyclic peptides with high affinity can be quickly and accurately screened out, greatly improving the efficiency of drug screening and reducing the corresponding costs, providing more reliable options for the drugability of polypeptides, and having broad application prospects and research significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics, and particularly relates to a method for virtual decompression of cyclic peptides. Background Art

[0002] In recent years, antibody-specific recognition elements (aptamers, nanobodies, molecularly imprinted polymers, etc.) have been a research hotspot. Polypeptides are substances between amino acids and proteins, with the advantages of small molecular weight, high activity, strong specificity, no immunogenicity, and easy synthesis. At the same time, they have structural diversity, specificity, and binding affinity for target analytes.

[0003] The cost of drug screening using animal models is high, the efficiency is low, and it requires a large amount of time. In recent years, with the progress of technology, the new drug discovery mode has gradually changed from mainly in vivo screening to mainly in vitro screening. The in vitro screening mainly focuses on high throughput screening (HTS). However, the positive rate of high throughput screening is extremely low. In order to more efficiently discover new compounds, people have developed a new drug screening method - virtual screening (VS) to increase the probability of discovering active drug molecules while reducing costs.

[0004] Virtual screening can be defined as: aiming at the three-dimensional structure of important disease target biomacromolecules or the quantitative structure-activity relationship model, searching for compounds that bind to the target biomacromolecules or conform to the QSAR model from existing small molecule databases for experimental screening research. The positive rate of virtual screening is generally between 5% and 20%, which is much higher than that of high throughput screening. Based on molecular docking, combining virtual screening and high throughput screening, high throughput virtual screening (HVS) is carried out. It relies on molecular simulation technology and can more quickly and efficiently select potentially effective candidate compounds from hundreds of thousands or even tens of millions of molecules, greatly shortening the research time and saving research funds.

[0005] Molecular docking technology is one of the important methods of virtual screening, mainly studying intermolecular interactions, and it is a theoretical simulation method that continuously locates and finds the best matching state between binding sites using spatial structures and interactions. This technology docks the molecules in the virtual peptide library one by one with specific active sites of the target protein crystal structure, quickly calculates through a computer and continuously adjusts the binding position of the peptide and the target protein, searches for the best conformation of the peptide molecule and the target protein in terms of spatial structure, and predicts the binding mode and affinity between the two. The peptide ligand with the best affinity for the target protein close to the natural conformation is selected through docking score values, the spatial trend of the peptide, and the sense-antisense peptide theory.

[0006] Prior studies have demonstrated a method for constructing partial or complete peptide libraries using cyclic peptides, such as CN111727194A. The use of cyclic peptide libraries to display more peptides reduces the number of peptides required to construct a complete polypeptide library, thereby increasing the capacity of the peptide library. It is easy to express and purify, and can be easily replicated at any time using cloned genes, while improving the screening efficiency. Summary of the Invention

[0007] To solve the problems of low efficiency and high cost in the traditional decompression method of cyclic peptides in the prior art, the object of the present invention is to provide a virtual decompression method for cyclic peptides, which uses virtual decompression technology to decompress cyclic peptides in a cyclic peptide library. This method avoids the complex preparation steps of polypeptides and the cumbersome screening process of aptamers, can greatly shorten the experimental time, save experimental costs, and at the same time avoid the interference of human errors and uncontrollable factors such as the environment, realizes the high-throughput design of specific recognition elements, provides strong technical support for the establishment of rapid screening of polypeptides, and provides an effective development approach for the development and design of drugs.

[0008] To achieve the above object, the present invention adopts the following technical solutions.

[0009] In a first aspect, the present invention provides a virtual decompression method for cyclic peptides, characterized in that the virtual decompression method comprises the following steps:

[0010] (1) Fragmentation of cyclic peptides to construct a polypeptide molecular library;

[0011] (2) Virtual screening of the polypeptide molecular library;

[0012] (3) Result analysis to screen out better polypeptide molecules, preferably, screen out the top 20 polypeptide molecules in terms of scoring;

[0013] The virtual screening of the fragmented polypeptides is completed based on molecular docking technology;

[0014] The length of the cyclic peptide is 50-100 amino acids.

[0015] In certain embodiments, the virtual decompression method is as Figure 1 shown.

[0016] In certain embodiments, the length of the cyclic peptide is 50-80 amino acids.

[0017] In certain embodiments, the length of the cyclic peptide is 50, 60, 70, or 80 amino acids.

[0018] In certain embodiments, the length of the cyclic peptide is 80 amino acids.

[0019] In certain embodiments, the virtual decompression method may be selected from Method 1 and / or Method 2,

[0020] Method 1 includes the following steps,

[0021] (1) Fragmentation of cyclic peptides: Based on the structure of the cyclic peptide, the cyclic peptide sequence is cleaved to obtain a library of polypeptide molecules containing polypeptide molecules with a length of 6 to 40 amino acids;

[0022] (2) Prediction of the polypeptide molecular structure model;

[0023] (3) Virtual screening of the polypeptide molecule library: The virtual screening method of the polypeptide molecule library may be selected from (a) and / or (b),

[0024] Method (a) includes:

[0025] (I) Preprocessing of the polypeptide molecule model;

[0026] (II) Molecular docking and / or molecular dynamics simulation;

[0027] Method (b) includes:

[0028] (i) Affinity calculation;

[0029] (4) Result analysis, screening the top 20 polypeptide molecules with the highest affinity scores in step (3) above;

[0030] Method 2 includes the following steps,

[0031] (A) Fragmentation of cyclic peptides: Based on the structure of the cyclic peptide, the cyclic peptide sequence is cleaved to obtain a library of polypeptide molecules containing polypeptide molecules with a length of 6 to 40 amino acids;

[0032] (B) Calculation of the interaction between the polypeptide molecule and the target protein;

[0033] (C) Result analysis, predicting the binding residues in the polypeptide molecule that bind to the target protein.

[0034] Method 1 and Method 2 are as Figure 2 shown.

[0035] In certain embodiments, the fragmentation of cyclic peptides: Based on the structure of the cyclic peptide, the cyclic peptide sequence is cleaved to obtain a library of polypeptide molecules containing polypeptide molecules with a length of 6 to 30 amino acids;

[0036] In certain embodiments, the fragmentation of cyclic peptides: Based on the structure of the cyclic peptide, the cyclic peptide sequence is cleaved to obtain a library of polypeptide molecules containing polypeptide molecules with a length of 6 to 20 amino acids;

[0037] In certain embodiments, the target protein for polypeptide molecular docking is pre-treated. The specific structure of the target protein is as follows: Obtain the three-dimensional crystal structure information of the target protein from the protein structure database, and download the PDB format file of the target protein crystal structure, with the resolution of this structure meeting the design requirements.

[0038] In certain embodiments, repair the missing amino acid residues and peptide chain ends of the obtained target protein, remove water molecules, and add polar hydrogens to the entire target protein molecule. Conduct structure prediction and kinetic simulation on the target protein to obtain specific conformational information, and then determine the most suitable conformation of the target protein.

[0039] In certain embodiments, the virtual decompression method can be selected from Method 1 and Method 2. When analyzing the results, the results of step (4) in Method 1 and step (C) in Method 2 need to be combined.

[0040] In certain embodiments, draw line charts and trend lines of the ADCP results, and at the same time draw line charts and trend lines according to the affinity results to more intuitively understand the relationship between the complex structure and affinity. Compare and contrast the ADCP results with the affinity results to evaluate the consistency and correlation between the two.

[0041] In certain embodiments, Method 1 further includes further calculating the binding residues of the top 20 polypeptides screened with the target protein to predict their binding to the target protein.

[0042] In certain embodiments, the software for fragmenting cyclic peptides in the virtual decompression method can be selected from: Cut_80.py, fasta.py.

[0043] In certain embodiments, the software for predicting the polypeptide molecular structure model in the virtual decompression method can be selected from: OmegaFold, PyMOL.

[0044] In certain embodiments, the software for pre-processing the polypeptide molecular model in the virtual decompression method can be selected from: Amber.

[0045] In certain embodiments, the software for molecular docking in the virtual decompression method can be selected from: agfrgui, ADCP (AutoDock CrankPep).

[0046] In certain embodiments, the software for molecular dynamics simulation in the virtual decompression method can be selected from: Amber.

[0047] In certain embodiments, the software for calculating affinity in the virtual decompression method can be selected from: prodigy.

[0048] In certain embodiments, the software selected for calculating the interaction between the polypeptide molecule and the target protein in the virtual decompression method may be selected from: MM-PBSA.

[0049] In certain embodiments, the virtual decompression method may be selected from Method 1 and / or Method 2.

[0050] Method 1 includes the following steps.

[0051] (1) Fragmentation of cyclic peptides: Based on the structure of cyclic peptides, the cyclic peptide sequence is cut using the Python script Cut_80.py to obtain a polypeptide molecule library containing polypeptide molecules with 6 to 40 amino acids in length; preferably, a polypeptide molecule library containing polypeptide molecules with 6 to 30 amino acids in length; more preferably, a polypeptide molecule library containing polypeptide molecules with 6 to 20 amino acids in length.

[0052] In certain embodiments, the Python script Cut_80.py is as shown in Figure 4.

[0053] (2) Prediction of polypeptide molecule structure models: Use the software OmegaFold to predict the model structure of the polypeptide molecules with 6 to 15 amino acids in length after cutting; use the software ColabFold to predict the complex structure model of the polypeptide sequence with 16 to 40 amino acids in length after cutting and the target protein. Preferably, ColabFold predicts the complex structure model of the polypeptide sequence with 16 to 30 amino acids in length after cutting and the target protein, and ColabFold predicts the complex structure model of the polypeptide sequence with 16 to 20 amino acids in length after cutting and the target protein.

[0054] In certain embodiments, using the OmegaFold software to predict the polypeptide molecule structure model includes the following steps: After submitting the fasta file of the polypeptide molecule model after cutting to the server, use a for loop statement to batch predict the polypeptide molecule structure model with OmegaFold.

[0055] In certain embodiments, using a for loop statement to predict the polypeptide molecule structure model with OmegaFold is as Figure 5 shown. Figure 5 In the example shown, all files with the file type fasta in the fasta directory are identified through a for loop statement and wildcards *, and then the omegafold${f}PDB command is used to predict the polypeptide molecule structure model using the OmegaFold software.

[0056] In some embodiments, predicting the molecular model structure of a polypeptide molecule using ColabFold software includes the following steps: storing the fasta file of 16 - 40 amino acids after cleavage in the directory specified by the Python script, and an example of the Python script is as Figure 6 shown. Before using the Figure 6 script shown, store the target protein sequence as a file named input.txt, rename the directory where the fasta file is stored to input (or directly modify the script parameters input_file and input_folder), and then run the script to obtain the fasta file of the complex. Subsequently, submit the complex fasta file to the server, and use a for loop statement and ColabFold to batch predict the molecular structure model of the polypeptide. An example of the shell script is as Figure 7 shown. In the above Figure 7 example shown, all files with the fasta file type in the fasta directory are identified through a for loop statement and the wildcard *, and then the complex structure model is predicted using the "colabfold_batch \"${file}\" \"output / ${folder} / \"" command.

[0057] (3) Virtual screening of the polypeptide molecular library: The virtual screening method of the polypeptide molecular library can be selected from (a) and / or (b). Among them, method (a) is applicable to the molecular model structure of polypeptide molecules with a length of 6 - 15 amino acids, and method (b) is applicable to the molecular model structure of polypeptide molecules with a length of 16 - 40 amino acids; preferably, method (b) is applicable to the molecular model structure of polypeptide molecules with a length of 16 - 30 amino acids; further, method (b) is applicable to the molecular model structure of polypeptide molecules with a length of 16 - 20 amino acids;

[0058] The method (a) includes:

[0059] (I) Pretreatment of the polypeptide molecular model: After the above steps (1) and (2) are completed, classify the polypeptide molecular model according to whether it contains disulfide bonds; through a Python script, it can be identified whether there are disulfide bonds in the polypeptide molecular model and classified. This step provides important information for the subsequent molecular docking process.

[0060] (II) Molecular docking and / or molecular dynamics simulation: Select the target protein, preferably select the appropriate conformation of the target protein according to different screening targets to select screening conditions, and perform molecular docking on the polypeptide molecular model using ADCP software; and / or perform molecular dynamics simulation on the polypeptide molecular model using Amber software.

[0061] The method (b) includes:

[0062] (i) Affinity calculation: Use the Python script extract_Conformation.py to extract the energy values after docking;

[0063] (4) Result analysis, screening the top 20 polypeptide molecules in terms of scoring in the above step (3);

[0064] In certain embodiments, step (4) selects the polypeptide molecule that does not contain the His structure and has the highest scoring result.

[0065] In certain embodiments, preprocess the polypeptide molecular model through a Python script. Among them, an example of the Python script is as Figure 8 shown, Figure 8 The Python script shown can be directly run in the PyMol software command line, which can identify whether there is a disulfide bond in the polypeptide molecular model and classify it.

[0066] In certain embodiments, use the ADCP software to perform molecular docking on the polypeptide molecular model. After submitting the preprocessed target protein docking file and ligand file to the server, use a for loop statement combined with the ADCP docking instruction to perform molecular docking on the polypeptide molecular model. An example of the shell script is as Figure 9 shown, the above Figure 9 parameters -N 50 and -n1000000 in the shown command can be modified to adjust the precision.

[0067] In certain embodiments, use the Python script extract_Conformation.py to extract the energy values after docking. An example of part of the script content is as Figure 10 shown. After running the above script, input the extraction file path to extract the energy values after docking.

[0068] The second method includes the following steps,

[0069] (A) Fragmentation of cyclic peptides: Based on the structure of cyclic peptides, cut the cyclic peptide sequence to obtain a polypeptide molecular library containing polypeptide molecules with 6 - 40 amino acids in length;

[0070] (B) Calculation of the interaction between polypeptide molecules and target proteins: Use the MM - PBSA software to calculate the interaction between polypeptide molecules and target proteins;

[0071] (C) Result analysis, predicting the binding residues in the polypeptide molecule that bind to the target protein.

[0072] In certain embodiments, using the MM - PBSA software to calculate the interaction between polypeptide molecules and target proteins includes the following steps:

[0073] a. File Preparation:

[0074] Generate a fasta file of the complex of the short peptide and the target protein:

[0075] Ensure that appropriate tools are used to generate the correct fasta file, and the connection mode and conformation of the protein and peptide may need to be considered.

[0076] AlphaFold Predicted Structure Model:

[0077] Try to use the latest version of AlphaFold or other high-performance protein structure prediction tools.

[0078] Ensure the accuracy of the model, and multiple models can be considered for comparison.

[0079] Amber Molecular Dynamics Simulation:

[0080] Before performing the simulation, perform energy minimization to optimize the initial structure.

[0081] Consider using different force field parameters, especially those related to peptides, to improve the simulation accuracy.

[0082] Increase the simulation time to obtain more reliable results if computing resources permit.

[0083] Extract the dominant conformations after 1 nanosecond of simulation:

[0084] Use appropriate tools and metrics to identify and extract the dominant conformations, such as clustering analysis, RMSD (root mean square deviation), etc.

[0085] b. Data Preparation:

[0086] Calculate MMPBSA:

[0087] Use MMPBSA.py in AmberTools to calculate the binding free energy. An example of the command is as follows: nohup mpirun -np 192 MMPBSA.py.MPI -O -i mmpbsa.in -o FINAL_RESULTS_MMPBSA.dat -do FINAL_DECOMP_MMPBSA.csv -sp com_solvated.prmtop -cp com.prmtop -rp rec.prmtop -lp lig.prmtop -y prod.nc &

[0088] c. Results:

[0089] Screen the top 20 polypeptide molecules by binding free energy:

[0090] Analyze the MMPBSA results and select the top 20 polypeptide molecules with the lowest binding free energy.

[0091] Combined with other information, such as conformational stability, H-bonds, hydrophobic interactions, etc., conduct more detailed screening.

[0092] Result visualization and analysis:

[0093] Use PyMOL to analyze the structure to ensure the biological interpretability of the results.

[0094] In certain embodiments, the virtual decompression method further includes Method 3, which includes the following two steps: constructing a complex structure model of the cyclic peptide and the target protein, and identifying the binding residues in the polypeptide molecule that bind to the target protein.

[0095] In certain embodiments, the software for predicting the complex structure model of the cyclic peptide and the target protein in Method 3 can be selected from: AlphaFold, OmegaFold, ColabFold, Amber, MM-PBSA, ADCP; the software for identifying the binding residues in the polypeptide molecule that bind to the target protein can be selected from PyMOL.

[0096] In certain embodiments, the software for predicting the complex structure model of the cyclic peptide and the target protein in Method 3 can be selected from: AlphaFold, OmegaFold, ColabFold. Preferably, the software is selected from AlphaFold.

[0097] In certain embodiments, the cyclic peptide sequence is a sequence formed by splicing after the original cyclic peptide is cut.

[0098] In certain embodiments, using AlphaFold to predict the complex structure model of the cyclic peptide and the target protein in Method 3 includes the following steps: submitting the complex fasta file to the server, and running AlphaFold to predict the complex structure model through a shell script. An example of the shell script is Figure 11 as shown. Figure 11 In the script shown, export CUDA_VISIBLE_DEVICES=1 and --gpu_devices=1 represent the selected GPU code number. The GPU can be selected by adjustment. The --model_preset=multimer parameter can be adjusted to predict the complex or monomer (here, the complex is used as an example. If predicting the monomer, the parameter needs to be modified to monomer). After modifying the parameters, running the shell script can predict the complex structure model.

[0099] In certain embodiments, the virtual decompression method further includes performing molecular dynamics simulations on the polypeptide molecules screened by the above methods to increase the accuracy of the results.

[0100] In certain embodiments, the virtual decompression method is as Figure 3 described.

[0101] In certain embodiments, the virtual decompression method further includes an experimental verification part. Screen the polypeptide molecules with top scores in the ranking, and perform solid-phase synthesis or protein expression; contact the obtained polypeptide molecules with the receptor in vitro or in vivo to determine whether the candidate compound regulates the activity of the receptor.

[0102] In certain embodiments, the verification methods for the screened polypeptide molecules can be selected from ELISA, TR-FRET or FLIPR screening.

[0103] In certain embodiments, the polypeptide molecules are obtained by solid-phase synthesis. The polypeptide molecules with top scores in the screening are synthesized by the Fmoc method. After cleavage and shedding, centrifugal precipitation, and freeze-drying, a primary polypeptide product containing impurities is formed. After dissolving it, impurities are removed by high-speed centrifugation, and it is detected and purified by HPLC with a purity greater than 95%. After determining that the molecular weight of the synthesized polypeptide molecule is consistent with the designed polypeptide molecule by MS detection and analysis, the purified solution is freeze-dried into a powder and stored at -20 °C for later use.

[0104] In certain embodiments, the virtual decompression method further includes a cyclic peptide screening step.

[0105] In certain embodiments, the cyclic peptides are screened from a cyclic peptide library. The cyclic peptide library is constructed by cyclic peptides with lengths of 50 - 100 amino acids respectively; further, the cyclic peptide library is constructed by cyclic peptides with lengths of 50 - 80 amino acids respectively; further, the cyclic peptide library is constructed by cyclic peptides with lengths of 50, 60, 70, 80 amino acids respectively; further, the cyclic peptide library is constructed by cyclic peptides with a length of 80 amino acids.

[0106] In certain embodiments, the screening method for the cyclic peptides from the cyclic peptide library can be selected from ELISA, TR-FRET or FLIPR screening.

[0107] In certain embodiments, the virtual decompression method further includes a sequence optimization step. After screening out the polypeptide molecules with higher scores through virtual screening, the polypeptide molecules can be further optimized.

[0108] Further, the sequence optimization is selected from the following forms or combinations thereof:

[0109] (1) Cyclic peptide synthesis: head-to-tail cyclization, side-chain cyclization (lactone, lactam, ether bond, etc.), multiple disulfide bonds, monothioether cyclization, etc.;

[0110] (2) Isotope labeling: 13 C, 15 N, 18 Isotope labeling such as O;

[0111] (3) Polyethylene glycol (PEG) modification: modification with PEG2, PEG4, PEG8, PEG12, PEG24, PEG36, PEG2000, PEG5000, PEG3400, PEG20K, PEG40K, etc.;

[0112] (4) Phosphorylation modification: phosphorylation modification of L-type or D-type amino acids (such as threonine T, serine S), phosphorylation modification of single or multiple amino acids;

[0113] (5) Polypeptide-protein conjugation: KLH, BSA, OVA, etc.;

[0114] (6) Modification of N-terminal or side-chain amino acids: acetylation, formylation, biotinylation, trifluoroacetylation, benzoylation, 2-aminobenzoylation, maleimidation, chloroacetylation, bromoacetylation, succinylation, palmitoylation, malic acidification, fatty acidification, formaldehydeation, chelation reaction (such as Hynic, DTPA, DOTA, NOTA), chlorination, fluorination, bromination, nitro or methoxy substitution, fluorescence labeling (such as Cy series, Texas series, Alexa series, rhodamine, Bodipy, Rox, FAM, FITC, MCA, TAMRA, Dnp), etc.;

[0115] (7) C-terminal modification: amidation, esterification, aldehydeation, alcoholization, succinylation, fluorescence labeling (such as Cy series, rhodamine, AMC, AFC, PNA, CMK, FMK), etc.;

[0116] (8) Alkylation modification: N-methylation, side-chain methylation, N-ethylation, N-phenylpropylation, N-allylation, etc.;

[0117] (9) Other special modifications: glycopeptide, sulfonation, MAPS, etc.

[0118] To make the present invention easier to understand, certain technical and scientific terms are specifically defined below. Unless otherwise clearly defined elsewhere in this document, all other technical and scientific terms used herein have the meanings commonly understood by those of ordinary skill in the art to which the present invention pertains.

[0119] In this document, the term "comprising" or "including" is an open expression, that is, it includes the content specified by the present invention, but does not exclude other aspects of the content.

[0120] The terms "polypeptide molecule" and "polypeptide" are used interchangeably and refer to a polymer composed of multiple amino acid residues, or variants, synthetic or naturally occurring analogs thereof. Thus, polypeptides are applicable to amino acid polymers in which one or more amino acid residues are synthetic non-naturally occurring amino acids (such as chemical analogs of the corresponding naturally occurring amino acids), as well as to naturally occurring amino acid polymers and their naturally occurring chemical derivatives. Examples of such derivatives include, but are not limited to, for example, post-translational modifications and degradation products, including pyroglutamyl, isoaspartyl, proteolytic, phosphorylated, glycosylated, oxidized, isomerized, and deaminated variants.

[0121] The term "target protein" refers to proteins and peptides having any biological function or activity (including structural, regulatory, hormonal, enzymatic, genetic, immunological, contractile, storage, transport, and signal transduction). In some embodiments, the target proteins include structural proteins, receptors, enzymes, cell surface proteins, proteins associated with the integrated functions of cells, including proteins involved in the following: catalytic activity, aromatase activity, motility activity, helicase activity, metabolic processes (anabolic and catabolic), antioxidant activity, proteolysis, biosynthesis, proteins having kinase activity, oxidoreductase activity, transferase activity, hydrolase activity, lyase activity, isomerase activity, ligase activity, enzyme regulator activity, signal transducer activity, structural molecule activity, binding activity (protein, lipid, carbohydrate), receptor activity, cell motility, membrane fusion, cell communication, biological process regulation, development, cell differentiation, stimulus response, behavioral proteins, cell adhesion proteins, proteins involved in cell death, proteins involved in transport (including protein transport activity, nuclear transport, ion transport activity, channel transport activity, carrier activity), permease activity, secretory activity, electron transport activity, pathogens, chaperone regulator activity, nucleic acid binding activity, transcriptional regulator activity, extracellular matrix and biogenesis activity, translational regulator activity. The proteins include proteins from eukaryotes and prokaryotes, including microorganisms, viruses, fungi, and parasites, as well as many others, including humans, microorganisms, viruses, fungi, and parasites targeted for drug therapy, other animals (including domestic animals), microorganisms targeted for the determination of antibiotics and other antimicrobial agents, plants, and even viruses, and many others.

[0122] The term "virtual screening" refers to the computer identification and design of chemical compounds that have the potential to bind to and modulate the function of a specific target protein. There are two variations of virtual screening, known as receptor-based virtual screening, when the three-dimensional structure of the receptor is available (RBVS); or ligand-based virtual screening (LBVS), when structural information about known ligands of the target molecule of interest is available, although a combination of the two variations is also often used.

[0123] The term "molecular docking" is an analytical method that uses computers to predict the interaction of macromolecules. Molecular docking simulates the interaction between small molecule ligands and receptor biomacromolecules based on the "lock and key principle" of ligand-receptor interaction. Ligand-receptor interaction is a process of molecular recognition, which mainly includes electrostatic interaction, hydrogen bonding, hydrophobic interaction, van der Waals interaction, etc. The mode and affinity of mutual binding between molecules can be predicted by calculation. Common methods of molecular docking Molecular docking methods can be roughly divided into the following three categories according to different degrees of simplification: (1) rigid docking; (2) semi-flexible docking; (3) flexible docking. Flexible docking means that during the docking process, the conformation of the research system can basically change freely. It is generally used to accurately examine the recognition between molecules. Since the conformation of the system can change during the calculation process, flexible docking improves the docking accuracy but requires a long calculation time.

[0124] Compared with the prior art, the present invention has the following advantages:

[0125] (1) The present invention uses virtual screening technology to carry out theoretical calculation and experimental verification. The present invention gives full play to the advantages of virtual screening. The design method is simple and easy to operate, and can efficiently construct and screen candidate polypeptides. This method can not only improve the screening success rate of affinity polypeptides, but also has the advantages of greatly shortening the experimental cycle, reducing workload and reducing costs, saving manpower, material resources and financial resources. Traditional methods mostly use a one-by-one synthesis method. Compared with traditional synthesis methods, the virtual decompression method of the present invention is specifically:

[0126] Traditional method Virtual decompression Time-consuming 3 - 9 months 3 days per item (3 - 5 items in parallel) Cost 12W <2000 / piece Accuracy rate Low decompression success rate High decompression success rate

[0127] (2) The preferred polypeptides screened by the present invention have high affinity, can specifically bind to the target protein, have high sensitivity, and can be applied to the prevention, screening, diagnosis and treatment of diseases.

[0128] (3) The series of preferred polypeptides screened by the present invention have good targeting effects and can be used as target polypeptides of targeted drug delivery systems or gene drug carriers. They are safe and reliable and have very broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0129] Figure 1 This is the flowchart of the virtual decompression method of the present invention;

[0130] Figure 2 This is the flowchart of the specific steps of Method 1 and Method 2 in the virtual decompression method of the present invention;

[0131] Figure 3 This is the flowchart of the specific steps of Method 1, Method 2, and Method 3 in the virtual decompression method of the present invention;

[0132] Figure 4 is the Python script Cut_80.py;

[0133] Figure 5 It is the for loop statement used in the prediction of the polypeptide molecular structure model by the OmegaFold software;

[0134] Figure 6 It is the Python script used in the prediction of the polypeptide molecular model structure by the ColabFold software;

[0135] Figure 7 It is the shell script used in the prediction of the polypeptide molecular model structure by the ColabFold software;

[0136] Figure 8 It is the Python script used in the preprocessing of the polypeptide molecular model;

[0137] Figure 9 The shell script used in molecular docking;

[0138] Figure 10 It is part of the script content of the Python script extract_Conformation.py;

[0139] Figure 11 It is the shell script used in the prediction of the complex structure model by AlphaFold;

[0140] Figure 12 It is the Python script Energy_extraction.py;

[0141] Figure 13 It is the Python script extract_Affinity.py;

[0142] Figure 14 It is the interfaceresidues.py script;

[0143] Figure 15 It is the structural formula of the polypeptide molecule shown in SEQ ID No.2;

[0144] Figure 16 is the structural formula of the polypeptide molecule shown in SEQ ID No. 9;

[0145] Figure 17 is the structural formula of the polypeptide molecule shown in SEQ ID No. 11;

[0146] Figure 18 is the structural formula of the polypeptide molecule shown in SEQ ID No. 14;

[0147] Figure 19 is the structural formula of the polypeptide molecule shown in SEQ ID No. 17;

[0148] Figure 20 is the structural formula of the polypeptide molecule shown in SEQ ID No. 20. Specific embodiments

[0149] As described above, the first aspect of the present invention provides a virtual decompression method for cyclic peptides, characterized in that the virtual decompression method comprises the following steps:

[0150] (1) Fragmentation of cyclic peptides to construct a polypeptide molecule library;

[0151] (2) Virtual screening of the polypeptide molecule library;

[0152] (3) Result analysis to screen out the top 20 polypeptide molecules in terms of scoring;

[0153] The virtual screening of the polypeptide molecule library is completed based on the molecular docking technology;

[0154] The cyclic peptide has a length of 50 to 100 amino acids.

[0155] Preferably, the virtual decompression method can be selected from Method 1 and / or Method 2.

[0156] More preferably, the virtual decompression method further includes Method 3, and Method 3 comprises the following two steps: constructing a complex structure model of the cyclic peptide and the target protein, and identifying the binding residues in the polypeptide molecule that bind to the target protein.

[0157] The virtual decompression method is as Figures 1 - 3 shown.

[0158] In some embodiments, the specific steps of the virtual decompression method are as follows:

[0159] 1. Sequence-based fragmentation

[0160] Using a Python script, the given cyclic peptide sequence can be cleaved to generate polypeptide molecular sequences of different lengths. In this task, the shortest length is set to 6 amino acids, the longest length is 40 amino acids, and the cleavage interval is 1. Preferably, the longest length is 20 amino acids.

[0161] For example, for the fragmentation of an 80-cyclic peptide, using the Python script cut_80, the given 80-cyclic peptide sequence can be cleaved to generate polypeptide molecular sequences of different lengths. In this task, we set the shortest length to 6 amino acids, the longest length to 40 amino acids, the cleavage interval to 1, and perform 80 cleavages. To confirm the active regions in the cyclic peptide, the 80-cyclic peptide is decompressed into polypeptide molecular sequences of different lengths. Using the cut_80 script in the Python programming language, the given 80-cyclic peptide sequence is cleaved. During the cleavage process, starting from the shortest length (6 amino acids), one amino acid is added each time until the longest length is reached. The cleavage interval is set to 1, meaning that 80 consecutive cleavages will be performed to obtain 80 polypeptide molecular sequences of different lengths. This cleavage method provides a systematic strategy that enables detailed study of the various regional fragments of the 80-cyclic peptide.

[0162] 2. Prediction of the structural model of polypeptide molecules

[0163] Use the software OmegaFold to predict the structural model of polypeptide molecules with 6 - 15 amino acid sequences after cleavage.

[0164] 3. Preprocessing of polypeptide molecular models

[0165] After the above steps are completed, the polypeptide molecules are classified according to whether they contain disulfide bonds. Disulfide bonds are common and important structural features in proteins and play a key role in protein folding and stability. Through a Python script, it is possible to identify whether there are disulfide bonds in the polypeptide molecular model and classify them. This step provides important information for the subsequent molecular docking process and can be used to adjust the parameters of AutoDock CrankPep to better consider the characteristics of disulfide bonds.

[0166] 4. Preprocessing of target proteins and molecular docking

[0167] First, the appropriate conformation of the target protein needs to be selected according to specific screening conditions. First, search in the PDB database or perform structural prediction and molecular dynamics simulation on the target protein to obtain specific conformation information, and then determine the most suitable conformation of the target protein; next, use the GUI interface of the AutoDock CrankPep software to predict the size and position of the docking box during the molecular docking process; finally, use the AutoDock CrankPep software to perform molecular docking on the polypeptide molecular models with 6 - 15 amino acid sequences.

[0168] 5. Structure model prediction of the complex

[0169] To predict the structure model of the complex of the cleaved polypeptide molecule with 16 or more amino acids (preferably 16 - 20 amino acids) and the target protein using the software ColabFold, a Python script is used to generate a complex fasta file suitable for ColabFold.

[0170] The complex fasta file is a commonly used format for storing protein sequence information. When generating the complex fasta file, the sequences of the fragmented polypeptide molecules and the target protein need to be combined into one fasta file. The ColabFold software can make predictions based on the information in this fasta file and generate the corresponding complex structure model. Through the Python script generate_com.py, the fragmented polypeptide molecules can be saved in a list. Then, using the file operation functions of Python, these polypeptide molecules are written into the fasta file one by one according to the requirements of the complex fasta file format.

[0171] 6. Calculation of affinity score

[0172] Use the Python script extract_Conformation.py to extract the energy values after docking.

[0173] 7. Data processing

[0174] Further analyze the results saved in an Excel sheet after extraction using the Python script Energy_extraction.py (as Figure 12 shown), draw line charts and trend lines. First, use this script to extract the required information from the calculation results, including the energy scores of the complexes. These energy scores can be used to evaluate the stability and binding strength of the complexes. By screening and sorting the energy scores, select the top twenty polypeptide molecule sequences that do not contain the His structure and have the highest scoring results. This screening condition helps to exclude the interference of the His structure on the evaluation of the complex structure and affinity, and select the polypeptide molecule sequences with the highest binding ability. Among the selected polypeptide molecule sequences, splice them according to the original cyclic peptide sequence to generate complete sequences for subsequent analysis and research. Use the Python script extract_Affinity.py (as Figure 13Extract the calculated affinity results as shown and save them in an Excel spreadsheet. Meanwhile, draw line charts and trend lines based on the affinity results to more intuitively understand the relationship between the complex structure and affinity. We will also draw line charts and trend lines of the ADCP results for comparison and contrast with the affinity results to evaluate the consistency and correlation between the two.

[0175] 8. Complex prediction analysis

[0176] Extract the sequences of the cyclic peptide and the target protein into the same fasta file, and predict the complex structure model using AlphaFold.

[0177] After cutting and splicing the cyclic peptide sequence, extract the spliced sequence and the target protein sequence into the same fasta file, and predict the complex structure model using AlphaFold.

[0178] 9. Post-prediction analysis

[0179] Use PyMOL to load the predicted complex structure model file. Run the interfaceresidues.py script (as shown Figure 14 to identify the interacting residue interface in the complex structure model.

[0180] 10. Data analysis

[0181] Compare the ADCP-spliced sequence with the interacting residues identified in the predicted complex structure model. Based on the comparison results, the selected polypeptide molecules can be optimized. If it is found that the spliced sequence does not match or is inconsistent with the interacting residues in the complex structure model, the sequence can be re-adjusted or screened to be more consistent with the complex structure. By comparing the ADCP-spliced sequence with the interacting residues in the predicted complex structure model and performing sequence optimization screening, a sequence that matches the target protein complex structure can be selected more accurately.

[0182] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings.

[0183] Example 1 Virtual decompression of GHSR agonist polypeptide

[0184] 1. Materials and methods

[0185] Cyclic peptide library (self-made), CHO-K1 / GHSR / Gα15 cells (GenScript), Calcium5 Assay Kit. 2. Screening of 80 cyclic peptides

[0186] Dissolution of cyclic peptide library: Place the 96-well deep well plate of cyclic peptide library in a centrifuge and centrifuge at 4000 rpm for 2 - 3 minutes. Use an automatic liquid dispenser to add 200 μL / well of ultrapure water to the 96-well deep well plate. Seal it with a silica gel lid and place it in a 95°C water bath for 5 minutes. Note: At this time, the cyclic peptide concentration is approximately: 50 μM. Place the dissolved cyclic peptide in the 96-well deep well plate in a centrifuge and centrifuge at 4000 rpm for 2 - 3 minutes.

[0187] The "cyclic peptide library" used is from Hunan ZhongSheng QuanPeptide Biochemical Co., Ltd. using the PICT (Peptide Information Compression Technology) patented technology. This technology uses biological means to compress polypeptide information, integrating the information of multiple polypeptides into one polypeptide, thus achieving a relatively small library capacity containing a large amount of polypeptide information; a cyclic peptide library containing nearly 73,000 80 - amino - acid cyclic peptides is constructed through the PICT technology. The specific construction method can refer to Patent CN201580081102.3 and Patent CN201780089941.9.

[0188] Dilution of cyclic peptide library: Transfer the dissolved cyclic peptide to a 384 - well plate using a workstation and dilute it to 10 μM with loading buffer. Obtain a CHO - K1 / GHSR cell line stably expressing human GHSR from GenScript and culture it in F - 12K containing 10% FBS and 500 μg / mL G418 at 37°C and 5% CO2 humidified conditions. CHO - K1 cells are used as GHSR negative control cells and cultured in F - 12K containing 10% FBS at 37°C and 5% CO2 humidified conditions.

[0189] The CHO-K1 and CHO-K1 / GHSR stable cell lines were used for the fluorescence imaging plate reader (FLIPR) assay. One day before the assay, CHO-K1 / GHSR cells were seeded into a 384-well plate (Thermo) at a density of 8×10 cells / 3 wells, with 25 μL of F-12K (10% FBS, 500 μg / mL G418) in each well. CHO-K1 cells were seeded into 384-well plate 3 at a density of 8×10 cells / well, with 25 μL of F-12K (10% FBS) in each well. Incubate at 37 °C in 5% CO2 for 20 - 22 h. On the day of the assay, dye was added to each well (50 μL per well in the 384-well plate). The experimental plate was incubated in the dark at room temperature for 120 minutes. The composition of 20 mL of the dye included 1 mL of Calcium 5 (Molecular Devices), 0.2 mL of 250 mM probenecid (Sigma), and 18.8 mL of loading buffer. The plate was placed in the FLIPR Tetra (Molecular Devices), and 12.5 μL of the test compound at 5× working concentration was added to the FLIPR. The fluorescence signal was detected by the FLIPR at room temperature according to the standard settings for 160 seconds, and then the cells were exposed to 15.6 μL of the agonist at 5× working concentration (the compound added by the FLIPR). The fluorescence signal was detected subsequently for 20 s (antagonist mode).

[0190] Through high-throughput screening, a cyclic peptide that activates the GHSR receptor at the cellular level was screened from nearly 73,000 80-membered cyclic peptides, namely SEQ ID No.1. SEQ ID No.1 is a cyclic peptide with its head and tail connected. The amino acid sequence of SEQ ID No.1 is as follows: HHHHYAIITRQLRAHGFPLTWQQLHHAWRYTRFKQNGDPLRGNQRRLFCRFCQHAVTCRQ CGSDLPGENSQREVPYCFHH (SEQ ID No.1).

[0191] 3. Virtual decompression of cyclic peptides

[0192] (1) Sequence-based fragmentation

[0193] Using a Python script, the given 80-membered cyclic peptide (SEQ ID No.1) sequence can be cleaved to generate polypeptide molecular sequences of different lengths. In this task, the shortest length was set to 6 amino acids, the longest length was set to 15 amino acids, and the cleavage interval was 1.

[0194] (2) Prediction of the structural model of polypeptide molecules

[0195] The software OmegaFold was used to predict the structural model of polypeptide molecules with 6 - 15 amino acid sequences after cleavage.

[0196] (3) Pretreatment of polypeptide molecular models

[0197] Add all hydrogens to the predicted polypeptide molecular model in step (2) and classify it into two categories: those containing disulfide bonds and those not containing disulfide bonds.

[0198] (4) Receptor protein pretreatment

[0199] The target GHSR is a GPCR-type protein. Therefore, it is necessary to confirm whether the conformation type of the target protein is activation or inhibition, select the correct and reasonable activation conformation, remove the crystal water, add all hydrogens, and then use the built-in GUI interface of ADCP to generate a docking file.

[0200] (5) Molecular docking

[0201] Submit the prepared ligand file and receptor file to the server and use ADCP for molecular docking.

[0202] (6) Affinity score calculation

[0203] Use the Python script extract_Conformation.py to extract the energy values after docking.

[0204] (7) Data processing

[0205] Further analyze the results saved in the Excel sheet after extraction using the Python script Energy_extraction.py, draw a line chart and a trend line. First, use this script to extract the required information from the calculation results, including the energy scores of the complexes. These energy scores can be used to evaluate the stability and binding strength of the complexes. By screening and sorting the energy scores, select the top twenty polypeptide molecular sequences that do not contain the His structure and have the highest scoring results, as shown in Tables 1 and 2. This screening condition helps to exclude the interference of the His structure on the evaluation of the complex structure and affinity, and select the sequences with the highest binding ability. Among the selected polypeptide molecular sequences, splice them according to the original 80-cyclic peptide sequence to generate a complete sequence for subsequent analysis and research.

[0206] Table 1 Top twenty results of polypeptide molecules with lengths of 6 - 10 amino acids

[0207] Ranking Sequence ADCP score value SEQ ID No. 1 HAWRYTRFKQ -22.0 15 2 HHAWRYTRFK -21.1 16 3 WRYTRFKQN -20.5 17 4 AWRYTRFKQ -20.2 18 5 PLTWQQLH -20.1 19 6 QQLHHAWR -20.0 20 7 NQRRLFCRFC -19.9 5 8 QLHHAWRYTR -19.7 21 9 PLTWQQL -19.5 22 10 LTWQQLHHAW -19.5 23 11 AWRYTRFKQN -19.3 24 12 WQQLHHAWRY -19.2 25 13 WQQLHHAW -19.1 26 14 YAIITRQLRA -19.1 27 15 HHAWRYTRF -18.9 28 16 RGNQRRLFC -18.9 29 17 LHHAWRYTRF -18.9 30 18 NQRRLFCR -18.8 31 19 RQLRAHGFPL -18.7 32 20 CRFCQHAVTC -18.7 33

[0208] Table 2 Top twenty results of polypeptide molecules with lengths of 11 - 15 amino acids

[0209]

[0210]

[0211] 4. Sequence verification

[0212] When detected by a fluorescence imaging plate reader (FLIPR), the auxin (GenScript) and the polypeptide molecule (solid-phase synthesis) screened in step 3 above were dissolved in ultrapure water to prepare a 1 mM stock solution. On the day of the experiment, Ghrelin was diluted to 500 nM using the loading buffer and serially diluted starting from the stock solution.

[0213] Detection was carried out according to the experimental procedure in step 2, and all experiments had at least 4 replicates. Data analysis was performed using ScreenWorks 4.2 and GraphPad Prism 9. The data were expressed as mean ± standard error (SE). Nonlinear regression analysis was used to calculate the half-maximal effective concentration (EC 50 50). The activation rate was the ratio of the change rate of calcium flux in CHO-K1 / GHSR cells induced by the sample to the change rate of calcium flux in CHO-K1 / GHSR cells induced by 100 nM Ghrelin. The experiments of some polypeptide molecules (SEQ ID No.2 to SEQ ID No.6) are shown in Table 3, where SEQ ID No.2 and SEQ ID No.3 are sequences spliced according to the ranking results.

[0214] Table 3 Concentration-response results of polypeptide molecules

[0215]

[0216] Example 2 Virtual decompression of EGFR inhibitory polypeptides

[0217] 1. Materials

[0218] Name Manufacturer Product number Biotinylated Human EGF R,His,Avitag Acro EGR - H82E3 EGF Protein,Human,Recombinant(ECD,hFc Tag) SINO 10605 - H01H Streptavidin - Eu(SA - Eu) ATTBio 16925 Anti - human Fc antibody - Alexa Fluor647 Jackson 109-605-170 Streptavidin,horseradish peroxidase conjugated Thermo 21126

[0219] 2. Screening of 80-membered cyclic peptides

[0220] (1) Dilute the cyclic peptide library as in Example 1.

[0221] (2) ELISA screening and verification method:

[0222] Fix the EGF Protein at 0.2 μg / mL, 25 μL / well on a 384-well plate (greiner 781097). Transfer 12.5 μL / well to the 384-well plate using a workstation, and then add 12.5 μL / well of a premixed solution composed of EGFR (0.12 μg / mL) and SA-HRP (0.2 μg / mL) premixed at a 1:1 volume ratio at room temperature for 1 h. After transient centrifugation to remove air bubbles, place it in an incubator at 37 °C for 1-2 h. After discarding the liquid in the wells, add 80 μl of 1×TBST washing solution with pH = 7.4 to wash the plate 3-5 times, 3 min - 5 min each time. Drain the detection plate, add 25 μL of chromogenic solution to each well, and continue to incubate in the incubator at 37 °C for 30 min. Finally, add 25 μL of termination solution (1 M HCL) to each well to terminate the reaction. Use a microplate reader cytation5 to read the absorbance at 450 nm. After calculating the inhibition rate, calculate the polypeptide IC 50 value, Inhibition%=(1 - (polypeptide signal - blank signal) / (positive signal - blank signal)*100).

[0223] Through high-throughput screening, a cyclic peptide SEQ ID No.7 with EGFR inhibitory effect at the cellular level was screened from an 80 cyclic peptide library. SEQ ID No.7 is a cyclic peptide with its head and tail connected. The amino acid sequence of SEQ ID No.7 is as follows:

[0224] HHHHSTSPFKRACRILLFTLTKNCHDTRKFFTGGDSWIKNASFRNLCPANRSRWRISSRIPARNWRRKSAFRMRAYCFHH (SEQ ID No.7).

[0225] 3. Virtual decompression of cyclic peptides

[0226] (1) Sequence-based fragmentation

[0227] Using a Python script, the given 80 cyclic peptide (SEQ ID No.7) sequence can be cut to generate polypeptide molecular sequences of different lengths. In this task, the shortest length is set to 6 amino acids, the longest length is set to 20 amino acids, and the cutting interval is 1.

[0228] (2) Prediction of the structural model of polypeptide molecules

[0229] Use the software OmegaFold to predict the structural model of polypeptide molecules with 6-15 amino acid sequences after cutting. Synthesize a complex fasta file from the polypeptide molecular sequences of 16-20 amino acids and the receptor protein sequence, and use ColabFold to predict the complex polypeptide molecular model.

[0230] (3) Preprocessing of polypeptide molecular models

[0231] Add full hydrogen to the 6 - 15 predicted polypeptide molecular models in step (2) and classify them into two categories: those containing disulfide bonds and those without disulfide bonds.

[0232] (4) Preprocessing of receptor proteins

[0233] Select the correct and reasonable active conformation of the target EGFR, remove the crystal water, add full hydrogen, and then generate a docking file using the built - in GUI interface of ADCP.

[0234] (5) Molecular docking

[0235] Submit the prepared ligand file and receptor file to the server and perform molecular docking using ADCP.

[0236] (6) Calculation of affinity score values

[0237] Use the Python script extract_Conformation.py to extract the energy values after docking; use prodigy to calculate the affinity score of the Complex.

[0238] (7) Data processing

[0239] Further analyze the results saved in the Excel sheet after extraction using the Python script Energy_extraction.py, draw line charts and trend lines. First, use this script to extract the required information from the calculation results, including the energy scores of the complexes. These energy scores can be used to evaluate the stability and binding strength of the complexes. By screening and sorting the energy scores, select the top twenty polypeptide molecular sequences that do not contain the His structure and have the highest scoring results, as shown in Table 4. This screening condition helps to exclude the interference of the His structure on the evaluation of the complex structure and affinity, and select the sequences with the highest binding ability. Among the selected polypeptide molecular sequences, splice them according to the original 80 - loop peptide sequence to generate complete sequences for subsequent analysis and research. Use the Python script extract_Affinity.py to extract the calculated affinity results and save them in an Excel sheet. At the same time, draw line charts and trend lines based on the affinity results to more intuitively understand the relationship between the complex structure and affinity. We will also draw line charts and trend lines of the ADCP results for comparison and contrast with the affinity results, so as to evaluate the consistency and correlation between the two.

[0240] Table 4 Results of the top twenty polypeptide molecules

[0241] Ranking Sequence ADCP score value SEQ ID No. 1 FKRACRILLFTLTKN -25.4 9 2 ILLFTLTKNCHDTRK -25.1 53 3 WRISSRIPARNWRRK -24.9 54 4 FKRACRILLFTLTK -24.7 55 5 TRKFFTGGDSWIKNA -23.6 56 6 PFKRACRILLFTLTK -23.5 57 7 RACRILLFTLTKNCH -23.5 10 8 FTGGDSWIKNASFRN -23.5 58 9 RISSRIPARNWRRKS -23.5 8 10 CRILLFTLTKNCH -23.3 59 11 RILLFTLTKNCHDT -23.3 60 12 TRKFFTGGDSWIKN -23.2 61 13 LTKNCHDTRKFFT -23.1 62 14 LFTLTKNCHDTRKFF -23.1 63 15 NCHDTRKFFTGGDSW -23.0 64 16 IPARNWRRKSAFRMR -23.0 65 17 NWRRKSAFRMRAYC -22.7 66 18 LLFTLTKNCHDTRKF -22.7 67 19 DSWIKNASFRNLC -22.6 68 20 RKFFTGGDSWIKNA -22.6 69

[0242] 4. Sequence Verification

[0243] Perform IC verification on the decompressed polypeptide molecules screened in step 3 according to the experimental procedures in step 2. The experimental results of some polypeptides (SEQ ID No. 8 to SEQ ID No. 10) are shown in Table 5. 50 Verification, the experimental results of some polypeptides (SEQ ID No. 8~SEQ ID No.10) are shown in Table 5.

[0244] Table 5 Concentration Response Results of Polypeptide Molecules

[0245]

[0246]

[0247] Example 3 Virtual Decompression of Apelin Receptor-Activating Polypeptides

[0248] 1. Materials and Methods

[0249] Cyclic peptide library (self-made), CHO-K1 / AGTRL1 (APJ) / Gα15 cells (GenScript), Calcium5 Assay Kit.

[0250] 2. Screening of 80-Cyclic Peptides

[0251] Dilute the cyclic peptide library as in Example 1.

[0252] Cell resuscitation: Take out the cells from the liquid nitrogen tank and quickly thaw the cells in a 37°C water bath. Transfer the cells to a 15 mL centrifuge tube, slowly add 9 mL of pre-warmed thawing medium, centrifuge at 800 rpm for 5 minutes, and remove the supernatant medium. Resuspend the cells with 5 mL of thawing medium, transfer to a T25 culture flask, and culture in an incubator at 37°C and 5% CO2. Replace the culture medium with growth medium on the second day after cell resuscitation.

[0253] Cell passage: When the cells reach 80-90% confluence in the culture flask, first rinse the cells with DPBS, then add DPBS again and gently tap the cell flask to make the cells detach from the wall; collect the cell suspension into a centrifuge tube, centrifuge at 800 rpm for 5 minutes, and remove the supernatant medium; add 6-8 mL of fresh growth medium, resuspend the cells, and passage them at a ratio of 1:3 - 1:8, and culture in an incubator at 37°C and 5% CO2. Change the medium every 2-3 days after passage.

[0254] Cell cryopreservation: When the cells reach 80-90% confluence in the culture dish, first rinse the cells with DPBS, then add DPBS again and gently tap the cell flask to make the cells detach from the wall; collect the cell suspension into a centrifuge tube, centrifuge at 800 rpm for 5 minutes, and remove the supernatant medium; resuspend the cells with cryopreservation medium, perform cell counting, and dilute the cells to 2-3×10 6 / mL. Aliquot 1 mL of the cell cryopreservation suspension into each cryotube. Place the cryotubes containing the cells into a cryobox, and after storing the cryobox in an -80 °C refrigerator overnight, transfer the cryotubes to a liquid nitrogen tank.

[0255] Twenty-four hours before the test, digest the CHO-K1 / AGTRL1 / Gα15 cells in the culture flask with 0.25% trypsin and suspend them in cell culture medium. Then, use a pipettor to add the cells to a black clear-bottom 384-well plate at a density of 10,000 cells per well for culture, 25 μL per well, and incubate overnight at 37 °C and 5% CO2.

[0256] On the day of the test, dissolve Component A in the Calcium5 kit in calcium dye buffer (Loading buffer) and add 250 mM probenecid to prepare a calcium dye solution containing 5 mM probenecid. After aspirating the culture medium, add 50 μL of the calcium dye solution to each well of the cell culture plate and incubate at room temperature for 2 hours. Two hours after adding the Calcium 5 dye to the cells, take out the cell culture plate and place it in the dark at room temperature for 10 minutes, then put it into a FLIPR instrument together with the polypeptide solution plate for detection.

[0257] Through high-throughput screening, an 80-membered cyclic peptide SEQ ID No.11 with Apelin receptor agonist activity at the cell level was screened. SEQ ID No.11 is a cyclic peptide with its head and tail connected. The sequence of SEQ ID No.11 is as follows:

[0258] HHHHHGTTPFRMSHLQKKWRWRATVRRWPILPIRQPVRLSSRKANNLVRSPNCAMTVQPPVAAGFSPVAGRRKATYCFHH (SEQ ID No.11).

[0259] 3. Virtual decompression of cyclic peptides

[0260] (1) Sequence-based fragmentation

[0261] Using a Python script, the given 80-membered cyclic peptide (SEQ ID No.13) sequence can be cut to generate polypeptide molecular sequences of different lengths. In this task, the shortest length is set to 6 amino acids, the longest length is set to 15 amino acids, and the cutting interval is 1.

[0262] (2) Prediction of the structural model of polypeptide molecules

[0263] Use the software OmegaFold to predict the structural model of polypeptide molecules with 6 - 15 amino acid sequences after cutting.

[0264] (3) Pretreatment of polypeptide molecule models

[0265] Add all hydrogens to the predicted polypeptide molecular model in step (2) and classify it into two categories: those containing disulfide bonds and those not containing disulfide bonds.

[0266] (4) Receptor protein pretreatment

[0267] The target Apelin receptor is a GPCR-type protein. Therefore, it is necessary to confirm whether the conformation type of the target protein is activation or inhibition, select the correct and reasonable activation conformation, remove the crystal water, add all hydrogens, and then use the GUI interface of ADCP to generate a docking file.

[0268] (5) Molecular docking

[0269] Submit the prepared ligand file and receptor file to the server and use ADCP for molecular docking.

[0270] (6) Calculation of affinity score

[0271] Use the Python script extract_Conformation.py to extract the energy values after docking.

[0272] (7) Data processing

[0273] Further analyze the results saved in the Excel table after extraction using the Python script Energy_extraction.py, draw a line chart and a trend line. First, use this script to extract the required information from the calculation results, including the energy scores of the complexes. These energy scores can be used to evaluate the stability and binding strength of the complexes. By screening and sorting the energy scores, select the top twenty polypeptide molecular sequences that do not contain the His structure and have the highest scoring results. The results are shown in Table 6. This screening condition helps to exclude the interference of the His structure on the evaluation of the complex structure and affinity and select the sequences with the highest binding ability. Among the screened polypeptide molecular sequences, splice them according to the original 80-ring peptide sequence to generate a complete sequence for subsequent analysis and research.

[0274] Table 6 Results of the top twenty polypeptide molecules

[0275]

[0276]

[0277] 4. Sequence verification

[0278] Perform EC 50 verification on the decompressed active polypeptides screened in step 3 according to the steps in step 2. The experimental results of some polypeptides (SEQ ID No. 12 to SEQ ID No. 14) are shown in Table 7.

[0279] Table 7 Concentration response results of polypeptide molecules

[0280]

[0281] Based on the results of the above embodiments, it can be seen that there are polypeptide molecules with good activity among the top 20 polypeptide molecules. Through the virtual decompression method of the present invention, polypeptide molecules disassembled from cyclic peptides with high affinity can be efficiently screened out, greatly improving the efficiency of drug screening and reducing the corresponding costs.

[0282] Example 4 Polypeptide synthesis

[0283] The linear precursors of the polypeptide compounds and their derivatives provided by the present invention are synthesized by solid-phase synthesis, and the target crude peptide is obtained after cleavage. The synthesis carrier is 2-Chlotrityl Resin. During the synthesis process, first, the 2-Chlotrityl Resin is fully swollen in N,N-dimethylformamide (DMF), and then the solid-phase carrier is repeatedly subjected to the operations of condensation → washing → deprotection of Fmoc → washing → condensation of the next amino acid to reach the length of the polypeptide chain to be synthesized. Finally, a mixed solution of trifluoroacetic acid: water: triisopropylsilane: benzyl methyl sulfide (90:2.5:2.5:5, v:v:v:v) is reacted with the resin to cleave the polypeptide from the solid-phase carrier, and the target polypeptide crude product is obtained after precipitation with frozen methyl tert-butyl ether. The polypeptide crude product is purified and separated by a C18 reversed-phase preparative chromatographic column in a system of acetonitrile / water with 0.1% trifluoroacetic acid to obtain the pure product of the polypeptide and its derivatives.

[0284] Experimental reagents

[0285]

[0286] (1) Preparation of the compound of SEQ ID No.2

[0287] Step 1: Coupling the first amino acid Fmoc-Lys(Boc)-OH

[0288] 84 mg (0.1 mmol) of 2-Chlorotrityl chloride resin was fully swollen in DCM for 1 h. Weigh Fmoc-Lys(Boc)-OH (0.08 mmol) and diisopropylethylamine (DIEA, 0.32 mmol), dissolve them in 5 mL of DCM and add them to the resin, and react at room temperature for 2 h. After the reaction is completed, add the blocking solution (10 mL) of DCM: methanol: DIEA (85:10:5, v:v:v) and block at room temperature for 10 min. The blocked resin was washed 5 times with DCM and 5 times with DMF.

[0289] Step 2: Synthesis of the linear precursor peptide chain

[0290] The linear precursor peptide chain K-I-H-W-G-N-G-G-A-E-L of SEQ ID No.2.

[0291] The resin obtained in Step 1 was fully swollen in DMF for 1 h, and then synthesized in the order from the second W at the carboxyl terminus to the amino terminus according to the linear precursor sequence. Each coupling cycle was carried out as follows:

[0292] · Deprotect Fmoc twice with 20% piperidine / DMF (20% v / v, 10 mL), 8 min each time.

[0293] · Rinse the resin with DMF 6 - 8 times until neutral pH.

[0294] · Dissolve 0.5 mmol of Fmoc-AA, 0.5 mmol of 6-chlorobenzotriazole-1,1,3,3-tetramethyluronium hexafluorophosphate (HCTU) and 1 mmol of 4-methylmorpholine (NMM) in DMF, add to the resin and react at room temperature for 1 h.

[0295] · Rinse the resin with DMF 4 - 6 times before coupling the next amino acid.

[0296] After the synthesis of the linear polypeptide, rinse the resin with DMF 5 times and with DCM 5 times. The resin was dried in vacuo.

[0297] Step 3: Cleavage of the linear precursor peptide chain

[0298] Add freshly prepared cleavage cocktail (10 mL) trifluoroacetic acid: water: triisopropylsilane: benzyl methyl sulfide (90:2.5:2.5:5, v:v:v:v) to the resin obtained in Step 2, and react with shaking at room temperature for 2 h. After the reaction, filter the reaction solution, wash the resin with trifluoroacetic acid, combine with the reaction solution, and precipitate with 4 volumes of cold MTBE to obtain the crude product. Wash the crude product with MTBE 3 times and dry in vacuo.

[0299] Step 4: Purification and preparation of the polypeptide

[0300] Dissolve the crude polypeptide in 20% aqueous acetonitrile solution, filter through a 0.45 μm membrane, and separate by a reverse-phase high-performance liquid chromatography system. The buffers are A (0.1% trifluoroacetic acid, aqueous solution) and B (0.1% trifluoroacetic acid, acetonitrile). Among them, the chromatographic column is a BR-C18 (Sepax) reverse-phase chromatographic column. During the purification process, the detection wavelength of the chromatograph is set at 230 nm, the flow rate is 15 mL / min, and the gradient is 20 - 50% acetonitrile in 40 min. Collect the relevant fractions of the product, combine the fractions with a purity >95% after HPLC identification, lyophilize to obtain the pure polypeptide.

[0301] (2) Preparation of the compound of SEQ ID No.9

[0302] Step 1: Coupling of the first amino acid Fmoc-Pro-OH

[0303] Swell 84 mg (0.1 mmol) of 2-Chlorotrityl chloride resin in DCM for 1 h. Weigh Fmoc-Pro-OH (0.08 mmol) and diisopropylethylamine (DIEA, 0.32 mmol), dissolve them in 5 mL of DCM, add to the resin, and react at room temperature for 2 h. After the reaction is completed, add the capping solution (10 mL) of DCM: methanol: DIEA (85:10:5, v:v:v) and cap at room temperature for 10 min. Wash the capped resin 5 times with DCM and 5 times with DMF.

[0304] Step 2: Synthesis of the linear precursor peptide chain

[0305] The linear precursor peptide chain of SEQ ID No.9 is L-Q-K-K-W-R-W-R-A-T-V-R-R-W-P.

[0306] Swell the resin obtained in Step 1 in DMF for 1 h, and then synthesize it in the order from the second W at the carboxyl terminus to the amino terminus according to the linear precursor sequence. Each coupling cycle is as follows:

[0307] · Deprotect Fmoc twice with 20% piperidine / DMF (20% v / v, 10 mL), 8 min each time.

[0308] · Wash the resin 6 - 8 times with DMF until the pH is neutral.

[0309] · Dissolve 0.5 mmol of Fmoc-AA, 0.5 mmol of 6-chlorobenzotriazole-1,1,3,3-tetramethyluronium hexafluorophosphate (HCTU), and 1 mmol of 4-methylmorpholine (NMM) in DMF, add to the resin, and react at room temperature for 1 h.

[0310] · Wash the resin 4 - 6 times with DMF before coupling the next amino acid.

[0311] After the synthesis of the linear polypeptide, wash the resin 5 times with DMF and 5 times with DCM. Dry the resin in vacuo.

[0312] Step 3: Cleavage of the linear precursor peptide chain

[0313] Add freshly prepared cleavage cocktail (10 mL) trifluoroacetic acid: water: triisopropylsilane: benzyl methyl sulfide (90:2.5:2.5:5, v:v:v:v) to the resin obtained in Step 2, and react with shaking at room temperature for 2 hours. After the reaction, filter the reaction solution, wash the resin with trifluoroacetic acid, combine it with the reaction solution, and precipitate with 4 volumes of cold MTBE to obtain the crude product. Wash the crude product 3 times with MTBE and dry it in vacuo.

[0314] Step 4: Purification and preparation of polypeptide

[0315] Dissolve the crude polypeptide in 20% aqueous acetonitrile solution, filter it through a 0.45 µm membrane, and separate it using a reverse-phase high-performance liquid chromatography system. The buffers are A (0.1% trifluoroacetic acid, aqueous solution) and B (0.1% trifluoroacetic acid, acetonitrile). Among them, the chromatographic column is a BR-C18 (Sepax) reverse-phase chromatographic column. During the purification process, the detection wavelength of the chromatograph is set at 230 nm, the flow rate is 15 mL / min, and the gradient is 20 - 50% acetonitrile in 40 min. Collect the relevant fractions of the product, combine the fractions with a purity > 95% after HPLC identification, and lyophilize to obtain the pure polypeptide.

[0316] (3) Preparation of the compound of SEQ ID No.11

[0317] Step 1: Coupling of the first amino acid Fmoc-Lys(Boc)-OH

[0318] Swell 84 mg (0.1 mmol) of 2-Chlorotrityl chloride resin in DCM for 1 h. Weigh Fmoc-Lys(Boc)-OH (0.08 mmol) and diisopropylethylamine (DIEA, 0.32 mmol), dissolve them in 5 mL of DCM, add them to the resin, and react at room temperature for 2 h. After the reaction is completed, add the blocking solution (10 mL) DCM: methanol: DIEA (85:10:5, v:v:v) and block at room temperature for 10 min. Wash the blocked resin 5 times with DCM and 5 times with DMF.

[0319] Step 2: Synthesis of the linear precursor peptide chain

[0320] The linear precursor peptide chain of SEQ ID No.11 is I-L-L-F-T-L-T-K-N-C-H-D-T-R-K.

[0321] Swell the resin obtained in Step 1 in DMF for 1 h, and then synthesize it in the order from the second amino acid W at the carboxyl terminus to the amino terminus according to the linear precursor sequence. Each coupling cycle is as follows:

[0322] · Deprotect with 20% piperidine / DMF (20% v / v, 10 mL) twice, 8 min each time.

[0323] · Wash the resin with DMF 6 - 8 times until neutral pH.

[0324] · Dissolve 0.5 mmol of Fmoc - AA, 0.5 mmol of 6 - chlorobenzotriazol - 1,1,3,3 - tetramethyluronium hexafluorophosphate (HCTU), and 1 mmol of 4 - methylmorpholine (NMM) in DMF, add to the resin, and react at room temperature for 1 h.

[0325] · Wash the resin with DMF 4 - 6 times before coupling the next amino acid.

[0326] After synthesizing the linear polypeptide, wash the resin with DMF 5 times and with DCM 5 times. Dry the resin in vacuo.

[0327] Step 3: Cleavage of the linear precursor peptide chain

[0328] Add freshly prepared cleavage cocktail (10 mL) trifluoroacetic acid: water: triisopropylsilane: benzyl methyl sulfide (90:2.5:2.5:5, v:v:v:v) to the resin obtained in Step 2, and react with shaking at room temperature for 2 h. After the reaction, filter the reaction solution, wash the resin with trifluoroacetic acid, combine with the reaction solution, and precipitate with 4 volumes of cold MTBE to obtain the crude product. Wash the crude product with MTBE 3 times and dry in vacuo.

[0329] Step 4: Purification and preparation of the polypeptide

[0330] Dissolve the crude polypeptide in 20% aqueous acetonitrile solution, filter through a 0.45 - um membrane, and separate by a reversed - phase high - performance liquid chromatography system. The buffers are A (0.1% trifluoroacetic acid, aqueous solution) and B (0.1% trifluoroacetic acid, acetonitrile). Among them, the chromatographic column is a BR - C18 (Sepax) reversed - phase chromatographic column. During the purification process, the detection wavelength of the chromatograph is set at 230 nm, the flow rate is 15 mL / min, and the gradient is 20 - 50% acetonitrile in 40 min. Collect the relevant fractions of the product, combine the fractions with a purity > 95% after HPLC identification, and lyophilize to obtain the pure polypeptide.

[0331] (4) Preparation of the compound of SEQ ID No.14

[0332] Step 1: Couple the first amino acid Fmoc - Asn(Trt) - OH

[0333] 84 mg (0.1 mmol) of 2-Chlorotrityl chloride resin was fully swollen in DCM for 1 h. Weighed Fmoc-Asn(Trt)-OH (0.08 mmol) and diisopropylethylamine (DIEA, 0.32 mmol) were dissolved in 5 mL of DCM and added to the resin, and the reaction was carried out at room temperature for 2 h. After the reaction was completed, a blocking solution (10 mL) of DCM: methanol: DIEA (85:10:5, v:v:v) was added at room temperature for 10 min for blocking. The blocked resin was washed 5 times with DCM and 5 times with DMF.

[0334] Step 2: Synthesis of the linear precursor peptide chain

[0335] The linear precursor peptide chain of SEQ ID No.14 is F-K-R-A-C-R-I-L-L-F-T-L-T-K-N.

[0336] The resin obtained in Step 1 was fully swollen in DMF for 1 h, and then synthesized in the order from the second position W at the carboxyl terminus to the amino terminus according to the linear precursor sequence. Each coupling cycle was carried out as follows:

[0337] · 20% piperidine / DMF (20% v / v, 10 mL) was used for Fmoc-deprotection twice, 8 min each time.

[0338] · The resin was rinsed 6 - 8 times with DMF until the neutral pH.

[0339] · 0.5 mmol of Fmoc-AA, 0.5 mmol of 6-chlorobenzotriazole-1,1,3,3-tetramethyluronium hexafluorophosphate (HCTU) and 1 mmol of 4-methylmorpholine (NMM) were dissolved in DMF, added to the resin and reacted at room temperature for 1 h.

[0340] · The resin was rinsed 4 - 6 times with DMF before the next amino acid coupling.

[0341] After the synthesis of the linear polypeptide, the resin was rinsed 5 times with DMF and 5 times with DCM. The resin was dried in vacuo.

[0342] Step 3: Cleavage of the linear precursor peptide chain

[0343] Freshly prepared cleavage cocktail (10 mL) of trifluoroacetic acid: water: triisopropylsilane: benzyl methyl sulfide (90:2.5:2.5:5, v:v:v:v) was added to the resin obtained in Step 2, and the reaction was carried out with shaking at room temperature for 2 h. After the reaction was completed, the reaction solution was filtered, and the resin was washed with trifluoroacetic acid and combined with the reaction solution. The crude product was precipitated with 4-fold volume of cold MTBE. The crude product was washed 3 times with MTBE and dried in vacuo.

[0344] Step 4: Purification and preparation of the polypeptide

[0345] The crude polypeptide was dissolved in 20% aqueous acetonitrile solution, filtered through a 0.45 μm membrane, and then separated by a reverse-phase high-performance liquid chromatography system. The buffers were A (0.1% trifluoroacetic acid, aqueous solution) and B (0.1% trifluoroacetic acid, acetonitrile). Among them, the chromatographic column was a BR-C18 (Sepax) reverse-phase chromatographic column. During the purification process, the detection wavelength of the chromatograph was set at 230 nm, the flow rate was 15 mL / min, and the gradient was 20 - 50% acetonitrile in 40 min. The relevant fractions of the product were collected, and after HPLC identification of the purity, the fractions with a purity > 95% were combined and freeze-dried to obtain the pure polypeptide.

[0346] (5) Preparation of the compound of SEQ ID No.17

[0347] Step 1: Coupling of the first amino acid Fmoc-His(Boc)-OH

[0348] 84 mg (0.1 mmol) of 2-Chlorotrityl chloride resin was fully swollen in DCM for 1 h. Weigh Fmoc-His(Boc)-OH (0.08 mmol) and diisopropylethylamine (DIEA, 0.32 mmol), dissolve them in 5 mL of DCM and add to the resin, and react at room temperature for 2 h. After the reaction was completed, a blocking solution (10 mL) of DCM:methanol:DIEA (85:10:5, v:v:v) was added and blocked at room temperature for 10 min. The blocked resin was washed 5 times with DCM and 5 times with DMF.

[0349] Step 2: Synthesis of the linear precursor peptide chain

[0350] The linear precursor peptide chain of SEQ ID No.17 is Y-A-I-I-T-R-Q-L-R-A-H.

[0351] The resin obtained in Step 1 was fully swollen in DMF for 1 h, and then synthesized in the order from the second amino acid W at the carboxyl terminus to the amino terminus according to the linear precursor sequence. Each coupling cycle was carried out as follows:

[0352] · 20% piperidine / DMF (20% v / v, 10 mL) was used for Fmoc-deprotection twice, 8 min each time.

[0353] · The resin was rinsed with DMF 6 - 8 times until the pH was neutral.

[0354] · Dissolve 0.5 mmol of Fmoc-AA, 0.5 mmol of 6-chlorobenzotriazole-1,1,3,3-tetramethyluronium hexafluorophosphate (HCTU) and 1 mmol of 4-methylmorpholine (NMM) in DMF, add to the resin and react at room temperature for 1 h.

[0355] · Rinse the resin 4 - 6 times with DMF before coupling the next amino acid.

[0356] After the synthesis of the linear polypeptide, rinse the resin 5 times with DMF and 5 times with DCM. Dry the resin under vacuum.

[0357] Step 3: Cleavage of the linear precursor peptide chain

[0358] Add freshly prepared cleavage cocktail (10 mL) trifluoroacetic acid: water: triisopropylsilane: benzyl methyl sulfide (90:2.5:2.5:5, v:v:v:v) to the resin obtained in Step 2, and react with shaking at room temperature for 2 hours. After the reaction, filter the reaction solution, wash the resin with trifluoroacetic acid, combine it with the reaction solution, and precipitate with 4 volumes of cold MTBE to obtain the crude product. Wash the crude product 3 times with MTBE and dry it under vacuum.

[0359] Step 4: Purification and preparation of the polypeptide

[0360] Dissolve the crude polypeptide in 20% aqueous acetonitrile solution, filter it through a 0.45 μm membrane, and separate it using a reverse-phase high-performance liquid chromatography system. The buffers are A (0.1% trifluoroacetic acid, aqueous solution) and B (0.1% trifluoroacetic acid, acetonitrile). Among them, the chromatographic column is a BR-C18 (Sepax) reverse-phase chromatographic column. During the purification process, the detection wavelength of the chromatograph is set at 230 nm, the flow rate is 15 mL / min, and the gradient is 20 - 50% acetonitrile in 40 min. Collect the relevant fractions of the product, combine the fractions with a purity > 95% after HPLC identification, and lyophilize to obtain the pure polypeptide.

[0361] (6) Preparation of the compound of SEQ ID No. 20

[0362] Step 1: Coupling of the first amino acid Fmoc-Cys(Trt)-OH

[0363] Swell 84 mg (0.1 mmol) of 2-Chlorotrityl chloride resin in DCM for 1 h. Weigh Fmoc-Cys(Trt)-OH (0.08 mmol) and diisopropylethylamine (DIEA, 0.32 mmol), dissolve them in 5 mL of DCM and add to the resin, and react at room temperature for 2 h. After the reaction, add the blocking solution (10 mL) DCM: methanol: DIEA (85:10:5, v:v:v) and block at room temperature for 10 min. Wash the blocked resin 5 times with DCM and 5 times with DMF.

[0364] Step 2: Synthesis of the linear precursor peptide chain

[0365] The linear precursor peptide chain of SEQ ID No. 20 is N-Q-R-R-L-F-C-R-F-C.

[0366] The resin obtained in Step 1 was fully swollen in DMF for 1 h, and then synthesized in the order from the second position W at the carboxyl terminus to the amino terminus according to the linear precursor sequence. Each coupling cycle was carried out as follows:

[0367] · Deprotect Fmoc twice with 20% piperidine / DMF (20% v / v, 10 mL), 8 min each time.

[0368] · Rinse the resin with DMF 6 - 8 times until neutral pH.

[0369] · Dissolve 0.5 mmol of Fmoc - AA, 0.5 mmol of 6 - chlorobenzotriazole - 1,1,3,3 - tetramethyluronium hexafluorophosphate (HCTU) and 1 mmol of 4 - methylmorpholine (NMM) in DMF, add to the resin and react at room temperature for 1 h.

[0370] · Rinse the resin with DMF 4 - 6 times before coupling the next amino acid.

[0371] After the synthesis of the linear polypeptide, rinse the resin with DMF 5 times and with DCM 5 times. Dry the resin in vacuo.

[0372] Step 3: Cleavage of the linear precursor peptide chain

[0373] Add freshly prepared cleavage cocktail (10 mL) trifluoroacetic acid: water: triisopropylsilane: benzyl methyl sulfide (90:2.5:2.5:5, v:v:v:v) to the resin obtained in Step 2, and react with shaking at room temperature for 2 h. After the reaction, filter the reaction solution, wash the resin with trifluoroacetic acid, combine with the reaction solution, and precipitate with 4 - fold volume of cold MTBE to obtain the crude product. Wash the crude product with MTBE 3 times and dry in vacuo.

[0374] Step 4: Purification and preparation of the polypeptide

[0375] Dissolve the polypeptide crude product in 20% aqueous acetonitrile solution, filter through a 0.45 - um membrane, and separate by a reverse - phase high - performance liquid chromatography system. The buffers are A (0.1% trifluoroacetic acid, aqueous solution) and B (0.1% trifluoroacetic acid, acetonitrile). Among them, the chromatographic column is a BR - C18 (Sepax) reverse - phase chromatographic column. During the purification process, the detection wavelength of the chromatograph is set at 230 nm, the flow rate is 15 mL / min, and the gradient is 20 - 50% acetonitrile in 40 min. Collect the relevant fractions of the product, combine the fractions with a purity > 95% after HPLC identification, and lyophilize to obtain the pure polypeptide.

[0376] All the polypeptide molecules of the present invention can be synthesized with reference to the above - mentioned synthesis method.

[0377] The embodiments of the present application have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

Claims

1. A virtual decompression method for an 80-ring peptide with EGFR inhibitory activity, characterized in that, The amino acid sequence of the 80-ring peptide is shown in SEQ ID No.

7. The virtual decompression method includes Step 1 and Step 2. The specific content of Step 1 is as follows: (1) Fragmentation of the cyclic peptide: Based on the structure of the cyclic peptide, the cyclic peptide sequence is cut. The shortest length is set to 6 amino acids, the longest length is 20 amino acids, and the cutting interval is 1, obtaining a polypeptide molecule library containing polypeptide molecules with lengths of 6-20 amino acids. (2) Prediction of the polypeptide molecule structure model: Use the software OmegaFold to predict the structure model of the polypeptide molecules with lengths of 6-15 amino acids after cutting; use the software ColabFold to predict the complex structure model of the polypeptide sequences with lengths of 16-20 amino acids after cutting and the target protein. (3) Virtual screening of the polypeptide molecule library: The virtual screening method of the polypeptide molecule library is selected from (a) and (b). Among them, method (a) is applicable to the structure models of polypeptide molecules with lengths of 6-15 amino acids, and method (b) is applicable to the structure models of polypeptide molecules with lengths of 16-20 amino acids. The method (a) includes: (I) Pretreatment of the polypeptide molecule model: After steps (1) and (2) are completed, the polypeptide molecule models are classified according to whether they contain disulfide bonds. (II) Molecular docking and molecular dynamics simulation: Select the target protein, use the ADCP software to perform molecular docking on the polypeptide molecule model; and use the Amber software to perform molecular dynamics simulation on the polypeptide molecule model. The method (b) includes: (i) Calculate the affinity and extract the energy value after docking. (4) Result analysis: Screen the top 20 polypeptide molecules scored in the above step (3). The specific content of Step 2 is as follows: (A) Fragmentation of the cyclic peptide: Based on the structure of the cyclic peptide, the cyclic peptide sequence is cut to obtain a polypeptide molecule library containing polypeptide molecules with lengths of 6-20 amino acids. (B) Calculation of the interaction between the polypeptide molecule and the target protein: Use the MM-PBSA software to calculate the interaction between the polypeptide molecule and the target protein. (C) Result analysis: Predict the binding residues in the polypeptide molecule that bind to the target protein.

2. The virtual decompression method according to claim 1, characterized in that, Step 1 also includes further predicting the binding residues of the top 20 screened polypeptide molecules that bind to the target protein.

3. The virtual decompression method according to claim 1 or 2, characterized in that, The virtual decompression method also includes Step 3. The specific content of Step 3 is as follows: Construct a complex structure model of the cyclic peptide and the target protein, and identify the binding residues in the polypeptide molecule that bind to the target protein. The software for predicting the complex structure model of the cyclic peptide and the target protein in Step 3 is selected from: AlphaFold, OmegaFold, ColabFold, Amber, MM-PBSA, ADCP; the software for identifying the binding residues in the polypeptide molecule that bind to the target protein is PyMOL.

4. The virtual decompression method according to claim 1, characterized in that, The virtual decompression method also includes an experimental verification part.

5. The virtual decompression method according to claim 1, characterized in that, The virtual decompression method also includes a cyclic peptide screening step.

6. The virtual decompression method according to claim 1, characterized in that, The virtual decompression method also includes a sequence optimization step.

Citation Information

Patent Citations

  • Peptide library constructing method and related vectors

    CN107849737A

  • Peptide library constructing method

    CN111727194A

  • Methods for constructing peptide libraries

    CN111727194B

  • Method for constructing target tumor polypeptide screening model by using structural biology

    CN115966248A

  • Target recognizing binding agents

    CN1756761A