Bionic antimicrobial peptide, pharmaceutical composition and its application in the preparation of antibacterial drugs, and design method of bionic antimicrobial peptide
By constructing and modifying the polypeptide sequence dataset and using deep learning models for pre-training and incremental migration training, a new bionic antimicrobial peptide with broad-spectrum antimicrobial properties and good biocompatibility was designed, which solved the problems of low antibiotic design efficiency and poor antimicrobial effect in the existing technology, and achieved efficient and low-cost antimicrobial drug design.
Patent Information
- Application Number
- CN202411293187.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-09-14
AI Technical Summary
The prior art is difficult to design and screen new bionic antibacterial peptides efficiently and at low cost in a short period of time, and traditional antibiotics usually only have activity on one or one type of bacteria, have slow bactericidal speed, weak biofilm removal ability, and are prone to bacterial resistance.
By constructing a data set containing a variety of N-terminal modifications, N-terminal modifications are performed using compounds to generate self-assembled polypeptide sequences with N-terminal modifications, and polypeptide drug sensitivity tests are performed. The deep learning model is used for pre-training and incremental migration training to predict the antibacterial activity of the peptide sequence, and the polypeptide sequence with prediction results above the threshold is selected for synthesis and purification, and in vitro experimental evaluation.
The new bionic antibacterial peptides designed have good biocompatibility, broad-spectrum antibacterial properties, fast bactericidal speed, strong biofilm removal ability, and are not prone to drug resistance. They have potential clinical application prospects, and through the prediction and screening of deep learning models, the time and cost of the design and screening process are significantly shortened.
Smart Images

Figure CN119541658B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of biomedicine. Specifically, it relates to a biomimetic antibacterial peptide, a pharmaceutical composition, their applications in the preparation of antibacterial drugs, and a design method for the biomimetic antibacterial peptide. Background Art
[0002] Bacterial infection is one of the diseases that seriously threaten human life and health. According to The Lancet, in 2019, approximately 7.7 million people died from common bacterial infections, accounting for 13.6% of the total global deaths, becoming the second leading cause of death globally. The treatment of bacterial infections mainly relies on the use of antibiotics. However, the abuse and misuse of antibiotics have led to the gradual emergence of drug resistance and even multi-drug resistance in bacteria, posing a severe challenge to clinical anti-infection treatment. In 2019, approximately 4.95 million deaths globally were associated with bacterial drug resistance, of which 1.27 million were directly attributed to the emergence of bacterial drug resistance. In the same year, the World Health Organization listed bacterial drug resistance as one of the top ten global health threats. Despite the increasingly severe problem of bacterial drug resistance, almost no new classes of antibiotics have been approved in recent years. Facing this situation, countries around the world regard the research and development of new antibacterial drugs as an important task in the field of biomedicine and are all striving to promote scientific and technological research on bacterial drug resistance prevention and control. It is urgent to develop new antibacterial drugs.
[0003] In nature, the self-assembly phenomenon of proteins and polypeptides exists widely, which can enhance their activities and functions. It has been found that a clovibactin can be isolated from soil bacteria. By self-assembling into higher-order fibers and binding to the bacterial cell wall precursor pyrophosphate (PPi), it can then block the biosynthesis of the cell wall. This simple and less affected killing mechanism makes Clovibactin less likely to develop drug resistance and have low cytotoxicity. In the human body, Paneth cells in the small intestine can secrete a defensin, which self-assembles into a nano-network structure in the intestine to resist bacterial invasion. Compared with traditional antibacterial peptides such as Cathelicidin LL-37, Magainin 2, and Melittin, these antibacterial polypeptides with self-assembly functions have better biocompatibility, longer in vivo half-lives, and are less likely to develop acquired drug resistance, making them a new type of antibacterial drug to replace antibiotics.
[0004] However, the market launch of a new type of antibiotic drug first requires the design and screening of hit compounds, the screening of lead compounds, and lead optimization. This process usually screens 1,000 to 10,000 drug molecules, but it takes 1 to 7 years and costs 5 to 10 million euros. That is to say, traditional screening methods can only screen 1,000 to 10,000 hit compound molecules and optimize the lead compounds among them within several years. The whole process consumes a large amount of time, manpower, material resources and financial resources. In addition, the targets of antibiotics obtained by traditional methods are usually inside bacteria or affect a certain biological process of them, resulting in traditional antibiotics usually being active against only one or a class of bacteria. Moreover, traditional antibiotics have a slow bactericidal speed, weak biofilm clearance ability, and bacteria gradually develop resistance to antibiotics, and the therapeutic effect of antibiotics on drug-resistant bacteria is gradually declining.
[0005] It can be seen from this that there is no design method in the prior art that can efficiently and low-costly design and screen new biomimetic antibacterial peptides in a short time, and there is also no new biomimetic antibacterial peptide with good biocompatibility, broad-spectrum antibacterial property, difficult to produce drug resistance, faster bactericidal speed, and stronger biofilm clearance ability. Summary of the Invention
[0006] An object of the present application is to provide a design method for a bionic antibacterial peptide. The design method includes: constructing a first data set containing a plurality of polypeptide sequences without N-terminal modification, wherein the polypeptide sequences in the first data set are composed of only 20 natural amino acids and are marked with antibacterial activity; using a plurality of compounds to perform N-terminal modification on the polypeptide sequences selected from the first data set to generate a plurality of self-assembled polypeptide sequences with N-terminal modification; performing polypeptide drug sensitivity tests on each of the generated self-assembled polypeptide sequences with N-terminal modification and marking antibacterial activity according to the test results, and using each of the self-assembled polypeptide sequences with N-terminal modification marked with antibacterial activity as a second data set; using the first data set to pre-train a deep learning model so that the pre-trained deep learning model can reconstruct the hidden feature representation of the polypeptide sequence and output the antibacterial activity prediction result of the polypeptide sequence; using the second data set to perform incremental transfer training on the pre-trained deep learning model so that the incrementally transfer-trained deep learning model can output the antibacterial activity prediction result of the self-assembled polypeptide sequence with N-terminal modification; using a plurality of compounds to perform N-terminal modification on the polypeptide sequences in the first data set to generate a third data set containing a plurality of self-assembled polypeptide sequences with N-terminal modification, using the incrementally transfer-trained deep learning model to perform antibacterial activity prediction on each of the self-assembled polypeptide sequences with N-terminal modification in the third data set, and selecting the self-assembled polypeptide sequences with N-terminal modification whose antibacterial activity prediction results are higher than a preset threshold for synthesis and purification and conducting polypeptide drug sensitivity tests; sequentially performing in vitro experimental evaluation and treatment effect evaluation on the self-assembled polypeptide sequences with N-terminal modification that pass the polypeptide drug sensitivity test, and using the self-assembled polypeptide sequences with N-terminal modification whose antibacterial activity is verified through evaluation as the designed novel bionic antibacterial peptide.
[0007] In some embodiments, the deep learning model includes a feature representation part, a pre-trained network, and an incremental transfer part. The incremental transfer part includes a noise enhancement module, an N-terminal modification encoding module, and a transfer learning network. Pre-training the deep learning model using the first dataset so that the pre-trained deep learning model can reconstruct the hidden feature representation of the polypeptide sequence and output the prediction result of the antibacterial activity of the polypeptide sequence specifically includes: the feature representation part performs input feature representation on each polypeptide sequence in the first dataset, trains and optimizes the pre-trained network based on the input feature representation of the polypeptide sequence and the antibacterial activity annotation, and the pre-trained network outputs the hidden feature representation of the reconstructed polypeptide sequence and the prediction result of the antibacterial activity of the polypeptide sequence. The prediction result of the antibacterial activity is the probability value that the polypeptide sequence has antibacterial activity. Performing incremental transfer training on the pre-trained deep learning model using the second dataset so that the incrementally transferred and trained deep learning model can output the prediction result of the antibacterial activity of the self-assembled polypeptide sequence with N-terminal modification specifically includes: using the data enhancement module to perform noise enhancement with sequence invariance on the input feature representation of the polypeptide sequence in the second dataset generated by the feature representation part, and inputting the input feature representation of the polypeptide sequence after noise enhancement into the pre-trained pre-trained network to generate the hidden feature representation of the polypeptide sequence after noise enhancement; using the N-terminal modification encoding module to generate the N-terminal modification feature representation of the self-assembled polypeptide sequence in the second dataset; inputting the spliced N-terminal modification feature representation and the hidden feature representation of the polypeptide sequence after noise enhancement into the transfer learning network, and performing incremental transfer training on the deep learning model in combination with the corresponding antibacterial activity annotation. After the incremental transfer training is completed, when the self-assembled polypeptide sequence with N-terminal modification is input into the deep learning model, the transfer learning network outputs the prediction result of the antibacterial activity of the self-assembled polypeptide sequence with N-terminal modification.
[0008] In some embodiments, the data volume in the second dataset is much less than that in the first dataset.
[0009] In some embodiments, the first dataset is a public dataset, and the public dataset includes positive data and negative data. The positive data is from the DBAASP database, and its polypeptide sequence has antibacterial activity; the negative data is from the Uniprot database, and its polypeptide sequence does not have antibacterial activity.
[0010] In some embodiments, the multiple compounds specifically include: octanoic acid (C8-), lauric acid (C12-), palmitic acid (C16-), phenylacetic acid (PHE-), 4-biphenylacetic acid (BIP-), diphenylacetic acid (DIP-), 2-naphthaleneacetic acid (NAP-), 9-anthracenecarboxylic acid (ANT-), 1-pyrenebutanoic acid (PYR-), cyclopropaneacetic acid (C-PRO), and cyclohexylacetic acid (C-HEX).
[0011] In some embodiments, performing a polypeptide drug sensitivity test on each self-assembled polypeptide sequence with an N-terminal modification generated and performing antibacterial activity annotation according to the test results specifically includes: when the polypeptide drug sensitivity test shows that the MIC is less than or equal to 100 μg / mL, annotating that the self-assembled polypeptide sequence with an N-terminal modification has antibacterial activity.
[0012] In some embodiments, the pre-trained network and the transfer learning network are constructed based on a Transformer network with a multi-head self-attention mechanism.
[0013] In some embodiments, the feature representation unit performing input feature representation on each polypeptide sequence in the first dataset specifically includes: performing feature encoding on 20 natural amino acids to generate 20 natural amino acid feature vectors {x j} j=1,…,20 , where j is the number of the natural amino acid; when the length of the polypeptide sequence in the first dataset is n, taking as the input feature representation of the polypeptide sequence, where j1, j2, …, j n are the numbers of the respective natural amino acids in the polypeptide sequence; using the data augmentation module to perform sequence-invariant noise augmentation on the input feature representation of the polypeptide sequences in the second dataset generated by the feature representation unit, and inputting the input feature representation of the polypeptide sequences after noise augmentation into the pre-trained pre-trained network specifically includes:
[0014] Calculating the Euclidean distance d i from each natural amino acid feature vector x i to the remaining natural amino acid feature vectors of the nearest neighbor according to formula (1):
[0015] d i = min({||x i - x j ||2} j=1,…,20 ), formula (1)
[0016] where, ‖x i - x j ||2 represents the Euclidean distance between x i and x j ;
[0017] At each training step in the incremental migration training, uniform sampling is performed again in the high-dimensional feature space centered on x i and with as the radius to obtain the noise-enhanced feature vector x i ′ of the natural amino acid corresponding to x as shown in formula (2): i ′:
[0018] x i ′ = x i + ε i , formula (2)
[0019] where ε i is the noise vector corresponding to x i and represents a uniform distribution between;
[0020] Take X′ in formula (3) as the input feature representation of the noise-enhanced polypeptide sequence corresponding to X:
[0021]
[0022] where is the noise vector corresponding to and is the noise vector corresponding to and is the noise vector corresponding to and respectively represent a uniform distribution between.
[0023] Another object of the present application is to provide a bionic antibacterial peptide having the following structure:
[0024]
[0025] Another object of the present application is to provide a pharmaceutical composition comprising the bionic antibacterial peptide described in each embodiment of the present application, and optionally a pharmaceutically acceptable carrier.
[0026] Another object of the present application is to provide the use of the bionic antibacterial peptide described in each embodiment of the present application in the preparation of antibacterial drugs, or the use of the pharmaceutical composition comprising the bionic antibacterial peptide described in each embodiment of the present application in the preparation of antibacterial drugs.
[0027] This application designs antibacterial peptides with bionic functions based on the interaction between deep learning and experiments. Compared with the traditional drug molecule screening process, the entire process consumes less time, manpower, material resources, and financial resources. Moreover, the newly designed bionic antibacterial peptides are mainly composed of natural amino acids and have good biocompatibility. The functional groups modified at the N-terminus can activate the self-assembly activity of the polypeptide and help the polypeptide target the bacterial membrane. Therefore, they are not easily prone to inducing bacterial drug resistance, can effectively kill drug-resistant bacteria, have a faster bactericidal speed, and have potential and broad application prospects in the treatment of clinical bacterial and drug-resistant bacterial infectious diseases. Brief Description of the Drawings
[0028] In the drawings, which are not necessarily drawn to scale, the same reference numerals may describe similar components in different views. The same reference numerals with alphabetical suffixes or different alphabetical suffixes may represent different instances of similar components. The drawings generally illustrate various embodiments by way of example and not limitation, and are used together with the description and the claims to explain the disclosed embodiments. Where appropriate, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be an exhaustive or exclusive embodiment of the device or method.
[0029] Figure 1 A flowchart showing the design method of the bionic antibacterial peptide according to an embodiment of the present application.
[0030] Figure 2 Eleven functional groups for modifying the N-terminus of the polypeptide sequence according to an embodiment of the present application are shown.
[0031] FIG. 3(a) shows a schematic diagram of the pre-training of the deep learning model according to an embodiment of the present application.
[0032] FIG. 3(b) shows a schematic diagram of the incremental transfer training of the deep learning model according to an embodiment of the present application.
[0033] Figure 4 The chemical structural formula of polypeptide p45 according to an embodiment of the present application is shown.
[0034] Figure 5 It is shown that polypeptide p45 according to an embodiment of the present application has stronger biofilm clearance ability than the traditional antibiotic ciprofloxacin, where: a shows the results of measuring the residual amount of the biofilm after treatment with ciprofloxacin and polypeptide p45 using the crystal violet method; b shows the results of observing the live / dead situation of bacteria in the biofilm by combining SYTO 9 / PI staining and CLSM.
[0035] Figure 6Shows the weight changes (a) and survival rates (b) of mice infected with Salmonella typhimurium in the intestine after different treatments according to the embodiments of the present application. PBS: PBS buffer; p45: polypeptide p45; Cip: ciprofloxacin. 5 mice in each group.
[0036] Figure 7 Shows the Salmonella typhimurium content in the feces (a), small intestine (b), and colon (c) of mice infected with Salmonella typhimurium in the intestine after different treatments according to the embodiments of the present application. 5 mice in each group.
[0037] Figure 8 Shows the HE, PAS, and AB-PAS stained sections of the small intestine of mice infected with Salmonella typhimurium in the intestine after different treatments according to the embodiments of the present application. Black arrows are used to compare the length of the small intestine appearance. The red square is used to magnify and observe the morphology of the small intestinal villi, and the red arrow points to the lesion area of the small intestinal villi. Scale bar, 100 μm.
[0038] Figure 9 Shows the HE, PAS, and AB-PAS stained sections of the colon of mice infected with Salmonella typhimurium in the intestine after different treatments according to the embodiments of the present application. The red arrow points to the area of neutrophil infiltration. Scale bar, 100 μm.
[0039] Figure 10 Shows the blood routine results of mice after treatment in the acute toxicity and side effect study according to the embodiments of the present application.
[0040] Figure 11 Shows the blood biochemical results of mice after treatment in the acute toxicity and side effect study according to the embodiments of the present application.
[0041] Figure 12 Shows the H&E stained sections of the heart, liver, spleen, lung, kidney, small intestine, and colon of mice after treatment in the acute toxicity and side effect study according to the embodiments of the present application.
[0042] Figure 13 Shows the blood routine results of mice after treatment in the chronic toxicity and side effect study according to the embodiments of the present application.
[0043] Figure 14 Shows the blood biochemical results of mice after treatment in the chronic toxicity and side effect study according to the embodiments of the present application.
[0044] Figure 15 Shows the H&E stained sections of the heart, liver, spleen, lung, kidney, small intestine, and colon of mice after treatment in the chronic toxicity and side effect study according to the embodiments of the present application.
[0045] Figure 16The results of detecting and analyzing 16S rRNA in the fecal homogenate of mice to explore the recovery of intestinal flora in the in vivo treatment effect study according to the embodiments of the present application are shown.
[0046] In the above figure, ns indicates no significant difference, * indicates that the p-value is less than 0.05, ** indicates that the p-value is less than 0.01, *** indicates that the p-value is less than 0.001, and the smaller the p-value, the more significant the difference between the two sets of data. Detailed implementation manners
[0047] The terms used herein are only for explaining the detailed implementation manners and do not limit the present application. Singular expressions include their plural expressions unless clearly stated or obviously not intended to be so from the context.
[0048] Although various modifications can be made to the present application and the present application can have various forms, specific examples will be described and explained in detail below. However, it should be understood that these are not intended to limit the present application to the specific disclosure, and the present application includes all modifications, equivalents or substitutions thereof without departing from the spirit and technical scope of the present application.
[0049] The present application will be further described in detail below through embodiments, but the scope of the present application is not limited to the embodiments.
[0050] In a first aspect, the present application provides a method for designing a bionic antibacterial peptide. Figure 1 The schematic flow diagram of the method for designing a bionic antibacterial peptide according to the embodiments of the present application is shown.
[0051] As Figure 1 shown, first in step 101, a first data set containing a plurality of polypeptide sequences without N-terminal modification is constructed, wherein the polypeptide sequences in the first data set are composed of only 20 natural amino acids and are marked with antibacterial activity.
[0052] In the present application, the polypeptide sequences labeled as having antibacterial activity can also become positive data, while the polypeptide sequences labeled as not having antibacterial activity are called negative data. In some embodiments, the above-mentioned first data set includes both positive data and negative data. Only as an example, the first data set is a publicly available data set, where the positive data is from the DBAASP database and the negative data is from the Uniprot database. In other embodiments, the positive data and negative data in the first data set can also come from other databases, as long as the antibacterial activity of the polypeptide sequences composed of 20 natural amino acids without N-terminal modification is clearly verified by experiments. In some embodiments, the antibacterial activity can be labeled based on the test results of the polypeptide drug susceptibility test on the polypeptide sequences. More specifically, for example, when the polypeptide drug susceptibility test shows that the MIC (minimum inhibitory concentration) is less than or equal to 100 μg / mL, the polypeptide sequence is labeled as having antibacterial activity. Conversely, when the test results show that the MIC is greater than 100 μg / mL, the polypeptide sequence is labeled as not having antibacterial activity. In other embodiments, the in vitro antibacterial activity experimental screening method can also use the inhibition zone experiment instead of the MIC experiment to evaluate the antibacterial activity of the biomimetic self-assembled polypeptide, and the present application does not limit this.
[0053] Next, in step 102, for example, polypeptide sequences can be selected from the first data set, and various compounds are used to perform N-terminal modification on the selected polypeptide sequences to generate various self-assembled polypeptide sequences with N-terminal modification (hereinafter also simply referred to as self-assembled polypeptide sequences); the polypeptide drug susceptibility test is performed on each of the generated self-assembled polypeptide sequences with N-terminal modification and the antibacterial activity is labeled according to the test results, and each self-assembled polypeptide sequence with N-terminal modification with an antibacterial activity label is used as the second data set. Among them, the polypeptide sequences selected from the first data set can be originally antibacterial (positive data) or originally non-antibacterial (negative data), which has no influence on the applicability of the design method described in the embodiments of the present application, and the present application does not limit this.
[0054] In the research practice of the present application, it is found that by using specific compounds to perform N-terminal modification on polypeptide sequences composed of natural amino acids, the self-assembly function of the polypeptide sequences can be greatly enhanced, so that such self-assembled polypeptide sequences containing N-terminal modification can not only have good antibacterial properties, but also may have better biocompatibility, longer in vivo half-life and are not prone to acquire drug resistance compared with traditional antibacterial peptides such as Cathelicidin LL-37, Magainin2, Melittin, etc., and are expected to obtain new antibacterial drugs to replace traditional antibiotics.
[0055] Therefore, the present application creatively proposes to use eleven compounds to perform N-terminal modification on several N-terminal unmodified polypeptide sequences in the first dataset, thereby generating a variety of self-assembling polypeptide sequences with N-terminal modification. Only as an example, for each polypeptide sequence selected from the first dataset, eleven self-assembling polypeptide sequences with different N-terminal modifications can be generated, and each self-assembling polypeptide sequence with N-terminal modification is subjected to a polypeptide drug sensitivity test, and antibacterial activity annotation is carried out according to the test results, and each self-assembling polypeptide sequence with N-terminal modification with antibacterial activity annotation is used as the second dataset.
[0056] Figure 2 Show eleven functional groups for modifying the N-terminus of polypeptide sequences according to an embodiment of the present application. As Figure 2 shown, the eleven compounds (functional groups) specifically include: octanoic acid (C8-), lauric acid (C12-), palmitic acid (C16-), phenylacetic acid (PHE-), 4-biphenylacetic acid (BIP-), diphenylacetic acid (DIP-), 2-naphthaleneacetic acid (NAP-), 9-anthracene carboxylic acid (ANT-), 1-pyrenebutanoic acid (PYR-), cyclopropaneacetic acid (C-PRO), and cyclohexylacetic acid (C-HEX).
[0057] More specifically, for example, polypeptide synthesis can be first carried out by the classical Fmoc solid-phase peptide synthesis method (SPPS), and the N-terminus of the synthesized polypeptide sequence is modified using Figure 2 the various compounds shown, then purified by reverse high-pressure liquid chromatography (RP-HPLC), and finally freeze-dried to obtain the final product. The synthesized polypeptide sequence containing N-terminal modification (i.e., the polypeptide molecule) can be detected for molecular weight and purity using mass spectrometry (MS) and liquid chromatography (HPLC). In some embodiments, for example, Escherichia coli (Gram-negative bacterium) and Staphylococcus aureus (Gram-positive bacterium) can be used to carry out the polypeptide drug sensitivity test. Specifically, 100 μL of polypeptide solutions with different concentrations (200 μg / mL, 100 μg / mL, 50 μg / mL, 25 μg / mL, 12.5 μg / mL) are added to a 96-well plate, and then an equal volume of bacterial solution with a concentration of 106 CFU / mL is added. After 18 hours, the 96-well plate is observed with the naked eye, and the well with a clear solution and the highest concentration is the minimum inhibitory concentration (MIC) of the polypeptide. The antibacterial activity of the polypeptide molecule is evaluated according to the results (MIC ≤ 100 μg / mL is positive and has antibacterial activity). Then, each self-assembling polypeptide sequence with N-terminal modification with antibacterial activity annotation is used as the second dataset.
[0058] Then, in step 103, the deep learning model is pre-trained using the first data set, so that the pre-trained deep learning model can reconstruct the hidden feature representation of the polypeptide sequence and output the prediction result of the antibacterial activity of the polypeptide sequence.
[0059] In some embodiments, the pre-training objective can be set to two parts. One part is for accurately reconstructing the feature representation of the polypeptide sequence input to the deep learning model, and the other part is for accurately predicting the antibacterial activity of the polypeptide sequence without N-terminal modification. In this way, after pre-training using tens of thousands of polypeptide sequence data with antibacterial activity annotations in the first data set, the deep learning model can not only accurately predict the antibacterial activity of the polypeptide sequence without N-terminal modification, but also accurately identify and reconstruct the feature representation of the polypeptide sequence input to the model.
[0060] Next, in step 104, the pre-trained deep learning model is further subjected to incremental transfer training using the second data set, so that the incrementally transferred and trained deep learning model can output the prediction result of the antibacterial activity of the self-assembled polypeptide sequence with N-terminal modification.
[0061] In the embodiments according to the present application, through incremental transfer training, the finally trained deep learning model can accurately predict whether the self-assembled polypeptide sequence with various N-terminal modifications corresponding to the polypeptide sequence has antibacterial activity on the basis of accurately reconstructing the feature representation of the polypeptide sequence without N-terminal modification.
[0062] Finally, in step 105, multiple compounds are used to perform N-terminal modification on the polypeptide sequences in the first data set to generate a third data set containing multiple self-assembled polypeptide sequences with N-terminal modification. Then, the incrementally transferred and trained deep learning model is used to predict the antibacterial activity of each self-assembled polypeptide sequence with N-terminal modification in the third data set. For example, when each polypeptide sequence in the first data set is subjected to N-terminal modification using each of the eleven functional groups described above, the number in the third data set will be eleven times the data volume in the first data set. However, since the deep learning model usually runs on a high-performance processor, the execution speed is very fast. Despite the huge data volume, it can still quickly obtain the antibacterial activity of each self-assembled polypeptide sequence with N-terminal modification in the third data set.
[0063] On this basis, in step 105, self-assembled polypeptide sequences with N-terminal modifications and antibacterial activity prediction results higher than a preset threshold can be further selected for synthesis and purification, and polypeptide drug sensitivity tests can be carried out. For self-assembled polypeptide sequences with N-terminal modifications that have passed the polypeptide drug sensitivity test (that is, the test results show antibacterial activity), in vitro experimental evaluation and treatment effect evaluation are carried out in sequence, and the self-assembled polypeptide sequences with N-terminal modifications whose antibacterial activity has been verified through evaluation are used as the designed novel bionic antibacterial peptides.
[0064] In some embodiments, the prediction result of antibacterial activity is the probability value that the polypeptide sequence / self-assembled polypeptide sequence has antibacterial activity (that is, the value range is [0, 1]). Therefore, the above preset threshold can be determined based on this probability value, for example. In other embodiments, it can also be determined after arranging the probability values of the prediction results of the antibacterial activity of each self-assembled polypeptide sequence in order of magnitude, so as to ensure that an appropriate number of self-assembled polypeptide sequences with higher antibacterial activity are selected, so that subsequent processes such as synthesis and purification, drug sensitivity test, in vitro experimental evaluation, and treatment effect evaluation have acceptable workloads and acceptable time costs. In the embodiments of the present application, for example, 140 self-assembled polypeptide sequences with the highest antibacterial activity can be selected for subsequent processes.
[0065] More specifically, the drug sensitivity experiment can be carried out for Gram-negative bacteria (Escherichia coli and Salmonella typhimurium) and Gram-positive bacteria (Staphylococcus aureus and Listeria), and the in vitro experimental evaluation and treatment effect evaluation will be described in detail in the following embodiments.
[0066] According to the design method of the bionic antibacterial peptide of the embodiments of the present application, the designed novel bionic antibacterial peptide is obtained by N-terminal modification on the basis of natural amino acids, has good biocompatibility, and the functional group modified at the N-terminus can activate the self-assembly activity of the polypeptide and help the polypeptide target the bacterial membrane. Therefore, it is not easy to cause bacterial drug resistance, can effectively kill drug-resistant bacteria, has a faster bactericidal speed, and has potential and broad application prospects in the treatment of clinical bacterial and drug-resistant bacterial infectious diseases.
[0067] In addition, due to the adoption of the method of interacting deep learning with test experiments in the design process, compared with the traditional drug molecule screening process, the time, manpower, material resources and financial resources consumed in the whole process are less. Moreover, the deep learning model is divided into two stages: pre-training and incremental transfer training, enabling the deep learning model to always predict the antibacterial activity of self-assembling polypeptide sequences based on the accurate reconstruction of the feature representation of polypeptide sequences. Therefore, the prediction results are more accurate and reasonable, avoiding wasting a large number of subsequent experimental verification efforts on sequences that actually do not have antibacterial activity. The actual situation proves that using the bionic antibacterial peptide design method in this application can quickly screen a polypeptide sequence database with an order of magnitude of billions. Combining experimental verification and screening, it is possible to screen up to 10^11 polypeptide molecules within 4 weeks, and successfully design a bionic antibacterial peptide with actual medicinal value, whose structure is shown as follows:
[0068] Various experimental verifications of this novel bionic antibacterial peptide will be described in detail later.
[0069] Table 1 compares the bionic antibacterial peptide design method of this application with the traditional antibiotic screening method.
[0070] Table 1 Comparison table of the bionic antibacterial peptide design method of this application and the traditional antibiotic screening method
[0071]
[0072] Next, the composition and training process of the deep learning network according to the embodiments of this application will be introduced in detail with reference to FIGS. 3(a) and 3(b). The deep learning model in this application may include three main components: a feature representation part, a pre-training network, and an incremental transfer part. Among them, the incremental transfer part further includes a noise enhancement module, an N-terminal modification encoding module, and a transfer learning network. Only as an example, the pre-training network and the transfer learning network may be constructed based on a Transformer network with a multi-head self-attention mechanism, but it is not limited thereto. The pre-training network and the transfer learning network may also adopt any other neural network structures in common use.
[0073] Figure 3(a) shows a schematic diagram of the pre-training of a deep learning model according to an embodiment of the present application. As shown in Figure 3(a), only the feature representation part and the pre-training network are involved in the pre-training of the deep learning model. In some embodiments, the first data set may contain tens of thousands of polypeptide sequence data with antibacterial activity annotations. For example, the first data set can be divided into a training data set, a validation data set, and a test data set according to a ratio of 8:1:1. Specifically, each polypeptide sequence in the training data set is input into the feature representation part, and the feature representation part performs input feature representation on the input polypeptide sequence according to a unified coding method. Then, the input feature representation of the polypeptide sequence is input into the pre-training network. During the training optimization process of the pre-training, for example, the sum of the L2 sequence reconstruction error of the model and the binary cross-entropy of the antibacterial activity prediction result can be used as the loss function, and the AdaM optimizer is used to update the network weights and completely traverse the training data in the training data set 100 times. Among them, the L2 sequence reconstruction error refers to the error between the hidden feature representation of the reconstructed polypeptide sequence and the input feature representation of the polypeptide sequence in the training data set. The binary cross-entropy of the antibacterial activity prediction result refers to the binary cross-entropy between the antibacterial activity prediction result of the polypeptide sequence without N-terminal modification output by the deep learning model during pre-training and the antibacterial activity annotation of the polypeptide sequence. Since the loss function of the pre-training consists of the above two parts, after the pre-training is completed, the deep learning model can not only accurately output the prediction result of the antibacterial activity of the polypeptide sequence without N-terminal modification by the pre-training network. Among them, the antibacterial activity prediction result can be, for example, the probability value that the polypeptide sequence has antibacterial activity. At the same time, the pre-training network can also accurately identify and reconstruct the feature representation of the polypeptide sequence input into the deep learning model.
[0074] Since the antibacterial peptides intended to be designed in the present application are not polypeptide sequences composed only of natural amino acids, but self-assembled antibacterial peptides with N-terminal modifications, after pre-training the deep learning model using the polypeptide sequences without N-terminal modifications in the first data set, it is necessary to further supplement the training so that it can accurately predict the antibacterial activity of the self-assembled polypeptide sequences with N-terminal modifications.
[0075] Figure 3(b) shows a schematic diagram of incremental transfer training of a deep learning model according to an embodiment of the present application. As shown in Figure 3(b), in incremental transfer training, all components of the deep learning model need to participate in training. It should be noted that since the self-assembled polypeptide sequences containing N-terminal modifications in the second dataset need to be subjected to polypeptide drug sensitivity tests for antibacterial activity annotation according to the test results, the workload is relatively large. Therefore, the data volume of the second dataset is much smaller than that of the first dataset. In this case, the influence of the small-scale second dataset on the deep learning model is easily diluted by the influence of the large-scale first dataset. In other words, if the polypeptide sequence data in the second dataset is not processed and directly used for incremental transfer training of the deep learning model, after incremental transfer training, the deep learning model may only be able to accurately predict whether the self-assembled polypeptide sequences containing N-terminal modifications corresponding to the trained polypeptide sequences have antibacterial activity, but cannot well predict the antibacterial activity of the self-assembled polypeptide sequences containing N-terminal modifications generated based on other polypeptide sequences.
[0076] To solve the above technical problems, the present application creatively proposes that in the incremental transfer training stage, as shown in Figure 3(b), a data augmentation module is added between the feature representation part and the pre-trained network. The data augmentation module performs sequence-invariant noise augmentation on the input feature representation of the polypeptide sequences in the second dataset generated by the feature representation part, and then inputs the input feature representation of the polypeptide sequences after noise augmentation into the pre-trained network. Moreover, the noise augmentation of the input polypeptide sequences by the above data augmentation module is carried out with each training step in the incremental transfer training and the noise augmentation method changes randomly each time. This method, on the one hand, greatly amplifies the data volume used for incremental transfer training, makes up for the possible defects caused by the much smaller data volume of the second dataset than that of the first dataset, and at the same time does not generate additional storage or time requirements due to data augmentation.
[0077] An implementation manner of the feature representation part and the data augmentation module is exemplarily introduced below.
[0078] In some embodiments, the feature representation part is used to perform input feature representation on polypeptide sequences (specifically, polypeptide sequences composed of only 20 natural amino acids without N-terminal modifications). It is easy to understand that in the pre-training stage, the feature representation part performs input feature representation on each polypeptide sequence in the first dataset, and in the incremental transfer training stage, similarly, the feature representation part is used to perform input feature representation on the polypeptide sequence part (i.e., not including the compound used for N-terminal modification) composed of natural amino acids in each self-assembled polypeptide sequence in the second dataset. The specific feature representation method is as follows.
[0079] First, use an applicable encoding method such as One-Hot Encoding to perform feature encoding on 20 natural amino acids to generate 20 natural amino acid feature vectors {x j}, j=1,…,20 where j is the number of the natural amino acid; when the length of the polypeptide sequence in the first dataset (or the second dataset) is n, use as the input feature representation of the polypeptide sequence, where j1, j2, …, j n are the numbers of the respective natural amino acids in the polypeptide sequence. By way of example only, the natural amino acids can be numbered according to the numbering method commonly used in the art. For example, the single-letter abbreviation of alanine is A, and its corresponding number is 1; the single-letter abbreviation of cysteine is C, and since there is no natural amino acid with a single-letter abbreviation of B, its corresponding number is 2, and so on. In some other embodiments, the 20 natural amino acids can also be feature-encoded according to other numbering and encoding methods, and corresponding input feature representations can be generated for polypeptide sequences of different lengths and different compositions.
[0080] Different from the pre-training stage, in the incremental transfer training stage, the input feature representation of the polypeptide sequence output by the feature representation unit is not directly input into the pre-trained network, but into the data augmentation module, and the data augmentation module performs noise addition enhancement with sequence invariance on the input feature representation of the polypeptide sequence. The specific method is as follows:
[0081] Calculate the Euclidean distance d i from each natural amino acid feature vector x i to the remaining natural amino acid feature vectors of the nearest neighbor according to formula (1):
[0082] d i = min({||x i - x j ||2} j=1,…,20 ), formula (1)
[0083] where ||x i - x j ||2 represents the Euclidean distance between x i and x j ;
[0084] In each training step of the incremental transfer training, uniform sampling is performed again in the high-dimensional feature space centered on x i with as the radius to obtain the noise-added enhanced feature vector x i corresponding to x i ' as shown in formula (2):
[0085] x′ i = x i + ε i , formula (2)
[0086] where ε i is the noise vector corresponding to x i and denotes a uniform distribution between;
[0087] Take X′ in formula (3) as the input feature representation of the noise-added and enhanced polypeptide sequence corresponding to X:
[0088]
[0089] where is the noise vector corresponding to and is the noise vector corresponding to and is the noise vector corresponding to and respectively denote a uniform distribution between.
[0090] In the above manner, while the discrete sequence feature space can be made continuous, the representation of the sequence itself can be retained. That is to say, it can be approximately considered that the noise-added sequence feature is an augmented representation of the original sequence. In this application, it means that there is no need to re-verify the antibacterial activity of the self-assembled polypeptide sequence after adding noise. Thus, it can be seen that what is proposed in this application is an innovative method that can significantly enhance the discrete sequence data features without increasing human, material, and time costs, and even without the need to additionally increase the amount of training data. Input the input feature representation of the noise-added and enhanced polypeptide sequence into the pre-trained network, and the pre-trained network can utilize its reconstruction ability of the hidden feature representation of the polypeptide sequence learned during the pre-training stage to generate the corresponding hidden feature representation for the self-assembled polypeptide sequence after adding noise and enhancement.
[0091] Meanwhile, eleven compounds for N-terminal modification of the self-assembling polypeptide sequences in the second dataset are input into the N-terminal modification coding module to generate N-terminal modification feature representations. Then, after splicing the N-terminal modification feature representations with the hidden feature representations of the polypeptide sequences after noise addition and enhancement, they are jointly used as the feature representations of the self-assembling polypeptide sequences with N-terminal modification after noise addition and input into the transfer learning network. Combined with the corresponding antibacterial activity annotations, an incremental transfer training is jointly performed on the deep learning models including the pre-trained network and the transfer learning network. During the incremental transfer training process, similarly, the data in the second dataset can also be divided into a training dataset, a validation dataset, and a test dataset according to a ratio of 8:1:1. Then, only the binary cross-entropy of the antibacterial activity prediction results is used as the loss function, and the noise vectors for enhancing sequence representation are resampled at each training step, and the Adam optimizer is used to update the network weights and the training dataset is completely traversed 50 times. After each incremental transfer learning training, the deep learning model and the hyperparameters used in the training process are adjusted according to the performance parameters of the deep learning model (such as accuracy, precision, and recall, etc.), and finally the best-performing deep learning model is obtained.
[0092] After the incremental transfer training is completed, the deep learning model has the ability to predict the antibacterial activity of the self-assembling polypeptide sequences containing N-terminal modification. Thus, in the case of inputting the self-assembling polypeptide sequences with N-terminal modification into the deep learning model, the transfer learning network will output the prediction results of the antibacterial activity of the self-assembling polypeptide sequences. Consistent with the foregoing, the prediction results of the antibacterial activity here can also be the probability value that the self-assembling polypeptide sequence has antibacterial activity.
[0093] According to the above design method of the bionic antibacterial peptide, a new type of bionic antibacterial peptide can be designed. Therefore, according to an embodiment of the present application, a bionic antibacterial peptide is further provided, which has the following structure:
[0094]
[0095] According to an embodiment of the present application, a pharmaceutical composition is further provided, which comprises the bionic antibacterial peptide according to the embodiment of the present application and an optional pharmaceutically acceptable carrier.
[0096] According to an embodiment of the present application, the application of the bionic antibacterial peptide in the present application or the pharmaceutical composition in the present application in the preparation of antibacterial drugs is further provided.
[0097] Hereinafter, the experimental materials, instruments and experimental methods involved in the design method of the biomimetic antimicrobial peptide of the present application, or the biomimetic antimicrobial peptide of the present application, the pharmaceutical composition containing the biomimetic antimicrobial peptide of the present application, and the biomimetic antimicrobial peptide / pharmaceutical composition of the present application in the process of preparing antimicrobial drugs will be described in detail through some examples. However, the examples provided herein are for illustrative purposes only and are not intended to limit the present application.
[0098] Unless otherwise specified, the experimental methods used in the following examples are conventional methods.
[0099] Unless otherwise specified, the materials and reagents used in the following examples can be obtained from commercial sources.
[0100] Experimental materials and instruments
[0101] Culture medium, RMPI 1640, purchased from ThermoFisher Scientific, sterile;
[0102] Fetal bovine serum was purchased from ThermoFisher Scientific, sterile;
[0103] 2-Cl resin was purchased from Tianjin Nankai Hecheng Technology Co., Ltd., 1.1 mmol / mg;
[0104] N,N-Diisopropylethylamine (DIPEA), purchased from Sigma-Aldrich, purity 99%;
[0105] Trifluoroacetic acid (TFA), purchased from Sigma-Aldrich, purity 99%;
[0106] Triisopropylsilane (TIS) was purchased from Sigma-Aldrich with a purity of 99%;
[0107] Gastric epithelial cells (GES-1) were purchased from Hunan Fenghui Biotechnology Co., Ltd.;
[0108] Octanoic acid (C8-) and cyclohexylacetic acid (C-HEX-) were purchased from Shanghai Haohong Biopharmaceutical Technology Co., Ltd. with a purity of 98%;
[0109] Lauric acid (C12-), purchased from Shanghai Yuanye Biotechnology Co., Ltd., purity 98%;
[0110] Hexadecanoic acid (C16-), purchased from Shanghai Sane Chemical Technology Co., Ltd., purity 98%;
[0111] 4-Biphenylacetic acid (BIP-), 2-Naphthaleneacetic acid (NAP-), 1-Pyrenebutanoic acid (PYR-), were purchased from Shanghai Aladdin Biochemical Technology Co., Ltd., with a purity of 98%;
[0112] Diphenylacetic acid (DIP-), 9-Anthracenecarboxylic acid (ANT-), Cyclopropaneacetic acid (C-PRO-), were purchased from Shanghai Macklin Biochemical Co., Ltd., with a purity of 98%;
[0113] Phenylacetic acid (PHE) was purchased from Shanghai National Pharmaceutical Group, with a purity of 98%;
[0114] L-configuration amino acids were purchased from Gil Biochemical (Shanghai) Co., Ltd., with a purity of 98%;
[0115] Varioskan LUX multi-functional microplate reader was purchased from Thermo Fisher Scientific;
[0116] Escherichia coli (E. coli, ATCC 25922), Staphylococcus aureus (S. aureus, ATCC 25923), Enterococcus faecium (E. faecium, ATCC 51559), Enterococcus faecalis (E. faecalis, ATCC 51575, ATCC 51299), were sourced from the American Type Culture Collection (ATCC);
[0117] Listeria monocytogenes (L. monocytogenes, CMCC 54004), was sourced from the China Medical Culture Collection Center (CMCC);
[0118] Enterobacter cancerogenus (E. cancerogenus, BNCC 363037), Staphylococcus epidermidis (S. epidermidis, BNCC330867), Acinetobacter baumannii (A. baumannii, BNCC 254392), Escherichia coli (E. coli, BNCC 186732), were sourced from Beijing NanoCell Biotech Co., Ltd. (BNCC);
[0119] Salmonella typhimurium (S. typhimurium, SL 1344), Staphylococcus aureus (S. aureus, USA300), were sourced from the laboratory bacterial strain bank;
[0120] SYTO 9 / PI live bacteria / dead bacteria dual staining kit was purchased from Beijing Lamboid Trading Co., Ltd.;
[0121] Hematoxylin and eosin combined staining (H&E) kit was purchased from Wuhan Sevier Biotechnology Co., Ltd.;
[0122] The PAS and AB-PAS staining kit was purchased from Wuhan Sevier Biotechnology Co., Ltd.
[0123] Examples of experimental methods
[0124] The process of efficiently, at low cost, and accurately predicting the antibacterial activity of each self-assembling polypeptide sequence with N-terminal modification in the third dataset using a deep learning model and a method of interacting deep learning with testing experiments has been described in detail in conjunction with Figure 1 、 Figure 2 、Figure 3(a) and Figure 3(b). The sensitivity test method for polypeptide drugs has also been described above. Next, only the Figure 1 synthesis, in vitro experimental evaluation, therapeutic effect evaluation of the self-assembling polypeptide sequence with N-terminal modification involved in Step 102 and Step 105, and finally the experimental process of obtaining the self-assembling polypeptide sequence with N-terminal modification having medicinal value in the present application will be illustrated by examples.
[0125] Examples of synthesizing bionic antibacterial peptides using the Fmoc-solid phase synthesis method
[0126] (1) Weigh 0.5 mmol of 2-Cl resin into a solid-phase synthesis tube, add 10 mL of anhydrous dichloromethane (DCM), and place it on a shaker and shake for 10 min to ensure that the 2-Cl resin is fully swollen;
[0127] (2) Use an ear bulb to press out the DCM from the solid-phase synthesis tube completely;
[0128] (3) Dissolve 1 mmol of Fmoc-protected amino acid in 10 mL of anhydrous DCM, add 2 mmol of N,N-diisopropylethylamine (DIPEA), and then transfer it to the solid-phase synthesis tube and react at room temperature for 1 h. Use an ear bulb to remove the reaction solution in the solid-phase synthesis tube, and then wash it with anhydrous DCM. The amount of DCM used each time is 10 mL, and the washing time is 1 min, for a total of 5 washes. After washing, add 20 mL of the prepared blocking reaction solution (anhydrous DCM:DIPEA:methanol = 17:1:2) and react at room temperature for 15 min;
[0129] (4) Use an ear bulb to remove the reaction solution in the solid-phase synthesis tube, first wash it with anhydrous DCM. The amount of DCM used each time is 10 mL, and the washing time is 1 min, for a total of 5 washes;
[0130] (5) Wash it with N,N-dimethylformamide (DMF). The amount of DMF used each time is 10 mL and the washing time is 1 min, for a total of 5 times. Add 10 mL of a 20% piperidine DMF solution by volume and react for 30 min, and then wash it with DMF. The amount of DMF used each time is 10 mL and the washing time is 1 min, for a total of 5 times;
[0131] (6) Add 3 mmol of the second Fmoc-protected amino acid, 4.5 mmol of HBTU, and 6 mmol of DIPEA to 20 mL of DMF. Add the prepared reaction solution to the solid-phase synthesis tube and react for 2 h.
[0132] (7) Repeat the operations in (5) and (6) to sequentially add the amino acids or N-terminal modification groups in the peptide sequence; then wash 5 times with DMF and then 5 times with DCM.
[0133] (8) Add a solution containing 95% (v / v) TFA, 2.5% TIS, and 2.5% H2O to the solid-phase synthesis tube, react for 2 h, expel and collect the reaction solution with an ear bulb, remove the solvent by rotary evaporation to obtain the crude product, and only purify it by HPLC.
[0134] Examples of in vitro experimental evaluations
[0135] In vitro experimental evaluations of the novel biomimetic antibacterial peptide can include, for example, in vitro cytotoxicity experiments, in vitro hemolysis experiments, in vitro broad-spectrum antibacterial experiments, biofilm clearance experiments, etc.
[0136] (1) In vitro cytotoxicity experiment
[0137] Seed gastric mucosal epithelial cells (GES-1) into a 96-well plate at a density of 10,000 cells per well. After the cells adhere and grow, prepare polypeptide solutions with different concentrations (1000 μg / mL, 500 μg / mL, 250 μg / mL, 125 μg / mL, 63 μg / mL, 31 μg / mL) in RPMI 1640 medium. Set up control groups (medium, medium + cells) for each plate. After co-incubation for 24 h, detect and calculate the 50% cell lethal concentration (EC50) by the MTT method. Calculate the selectivity coefficient for gastric mucosal epithelial cells: SI_E represents the selectivity coefficient (cytotoxicity), EC50 represents the concentration at which the polypeptide causes 50% cell death, and MIC represents the minimum inhibitory concentration.
[0138] (2) In vitro hemolysis experiment
[0139] Dilute rabbit red blood cells (rRBC) to 8% and add 100 μL to each well of a 96-well plate. Then add 100 μL of polypeptide solutions with different concentrations (1000 μg / mL, 500 μg / mL, 250 μg / mL, 125 μg / mL, 63 μg / mL, 31 μg / mL). Set up control groups on each plate (cell suspension + 1% Triton X-100, cell suspension + PBS). After co-incubation for 3 h, read OD570 using an enzyme-linked immunosorbent assay (ELISA) reader and calculate the 50% hemolysis concentration (HC50). Calculate the selectivity coefficient for mouse red blood cells: SI_H represents the selectivity coefficient (for mouse red blood cells), HC50 represents the concentration of the polypeptide causing 50% hemolysis, and MIC represents the minimum inhibitory concentration. SI_E represents the selectivity coefficient (for human gastric mucosal epithelial cells), EC50 represents the concentration of the polypeptide causing 50% cell death, and MIC represents the minimum inhibitory concentration. The experimental results are shown in Table 2.
[0140] Table 2 MIC values of candidate polypeptides against Salmonella typhimurium, 50% hemolysis concentration (HC50) against mouse red blood cells, 50% lethal concentration (EC50) against human gastric mucosal epithelial cells, and selectivity coefficients (SI_H (for mouse red blood cells), SI_E (for human gastric mucosal epithelial cells)), in μg / mL.
[0141]
[0142] As shown in Table 2, in further in vitro experimental screening, polypeptide p45 exhibited the best biocompatibility (highest selectivity coefficient) and the strongest antibacterial activity (lowest MIC value). At the same time, the MIC value of p45 without N-terminal modification was greater than 100 μg / mL, demonstrating that without N-terminal modification, this polypeptide lost its antibacterial activity. The structure of the best self-assembling antibacterial polypeptide p45 screened out is as Figure 4 shown.
[0143] (3) In vitro broad-spectrum antibacterial experiment
[0144] As shown in Table 3, the MIC values of the verified self-assembling antibacterial peptides against various bacteria (E. coli, S. aureus, S. typhimurium, L. monocytogenes, E. faecium, E. faecalis, E. cancerogenus, S. epidemidis, A. baumannii, etc.) were measured through MIC experiments to evaluate their broad-spectrum antibacterial properties.
[0145] Table 3 Broad-spectrum antibacterial properties of the verified self-assembling antibacterial peptides against various bacteria
[0146]
[0147] Finally, by comparing the above results with the bacterial selection coefficients (EC50 / MIC, HC50 / MIC), biomimetic antibacterial peptides with good biocompatibility and broad-spectrum antibacterial properties were screened out.
[0148] (4) Biofilm clearance experiment
[0149] Salmonella typhimurium biofilm was incubated on the surface of a culture dish, and 1 mL of PBS buffer, ciprofloxacin (concentration: 20×MIC, 20 μg / mL), and p45 polypeptide (concentration: 20×MIC, 125 μg / mL) were added respectively and treated for 3 hours. Then, the crystal violet assay was used to detect the content of residual biofilm. SYTO 9 / PI (live / dead) staining combined with confocal microscopy was used to observe the viability of the treated biofilm. As Figure 5 shown, polypeptide p45 has stronger biofilm clearance ability than the traditional antibiotic ciprofloxacin.
[0150] Examples of evaluating the therapeutic effect of novel biomimetic antibacterial peptides
[0151] (1) In vivo therapeutic effect study
[0152] A mouse model of intestinal infection was constructed using Salmonella typhimurium. Female BALB / c mice aged 6 - 8 weeks (purchased from Shanghai Jihui Laboratory Animal Breeding Co., Ltd., ethical batch number: AP#22 - 025 - WHM, breeding environment temperature 20 - 26 °C, relative humidity 50 - 60%, light cycle 12 hours light / 12 hours dark) were selected. After fasting and water deprivation for 4 h, each mouse was intraperitoneally injected with 20 mg of streptomycin sulfate. After 20 h of resuming diet, they were fasted and water-deprived again for 4 h, and then 100 μL of Salmonella typhimurium bacterial solution (10 9 CFU / mL, PBS) was gavaged to each mouse. After 2 days of resuming diet, the mice were randomly grouped (5 mice per group) for subsequent experiments. Each group of infected mice was intraperitoneally injected once with 200 μL of PBS buffer, ciprofloxacin, and polypeptide p45 at a dose of 30 mg / kg, and the daily survival and body weight changes of each group of mice were recorded ( Figure 6 ). On the fifth day after treatment, fresh feces of the mice were collected and the mice were sacrificed. Small intestine and colon tissues were taken, homogenized respectively, plated, cultured, counted, and the content of Salmonella typhimurium in feces, small intestine, and colon was calculated ( Figure 7 ). Then, the H&E and PAS stained sections of the small intestine and colon of the mice were observed and compared ( Figure 8 , Figure 9), investigate the intestinal recovery after treatment (inflammation, small intestinal villi, intestinal cells, etc.). At the same time, Wuhan Sevier Biotechnology Co., Ltd. was commissioned to conduct 16s rRNA detection and analysis on the mouse fecal homogenate to explore the recovery of the intestinal flora. The experimental results are as Figure 16 shown. It can be seen that the flora after p45 treatment and after antibiotic cip treatment are closer. That is to say, the therapeutic effect of the present invention is equivalent to that of the traditional antibiotic cip.
[0153] (2) Toxicity and side effect research
[0154] Female BALB / c mice at 6 - 8 weeks old were selected and randomly grouped (5 mice in each group, a total of 6 groups) for toxicity and side effect research. The healthy mice were injected with the bionic antibacterial peptide and PBS solution using the same treatment protocol in the treatment experiment. Among them, 3 groups had their blood taken 24 hours after treatment for routine blood tests (red blood cells (RBC), mean corpuscular volume (MCV), platelets (PLT), white blood cells (WBC), neutrophils (NEUT), monocytes (MONO)) ( Figure 10 ) and blood biochemistry tests (alanine aminotransferase (ALT), aspartate aminotransferase (AST), blood urea nitrogen (BUN), uric acid (UA), creatinine (CREA)) ( Figure 11 ), and different organs (heart, liver, spleen, lung, kidney, small intestine, colon) ( Figure 12 ) were collected for H&E staining section scanning to analyze the acute toxicity and side effects of the treatment protocol according to the results. The remaining 3 groups recorded the survival rate and body weight changes of each group of mice after treatment, and the same routine blood test ( Figure 13 ), blood biochemistry ( Figure 14 ), and tissue staining section experiments ( Figure 15 ) were conducted on the 7th day to analyze the chronic toxicity and side effects of the treatment protocol.
[0155] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of their embodiments) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. Additionally, in the above - described specific embodiments, various features can be grouped together to simplify the present application. This should not be construed as an intention that the disclosed features not claimed are necessary for any claim. On the contrary, the subject matter of the present invention can be less than all the features of a particular disclosed embodiment. Thus, the claims are incorporated herein as examples or embodiments into the specific embodiments, where each claim independently serves as a separate embodiment, and considering these embodiments, they can be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the appended claims and the full scope of the equivalents to which these claims are entitled.
Claims
1. A biomimetic antimicrobial peptide having the structure shown below: 。 2. A method for designing a biomimetic antimicrobial peptide according to claim 1, characterized in that: The design method comprises: Constructing a first data set comprising a plurality of polypeptide sequences without N-terminal modification, wherein the polypeptide sequences in the first data set consist of only 20 natural amino acids and are annotated with antibacterial activity; Using a plurality of compounds to perform N-terminal modification on the polypeptide sequences selected from the first data set to generate a plurality of self-assembling polypeptide sequences with N-terminal modification; performing a polypeptide drug sensitivity test on each of the generated self-assembling polypeptide sequences with N-terminal modification and annotating the antibacterial activity according to the test results, and using each of the self-assembling polypeptide sequences with N-terminal modification annotated with antibacterial activity as a second data set; Constructing a deep learning model including a feature representation unit, a pre-trained network and an incremental migration unit, wherein the incremental migration unit includes a noise enhancement module, an N-terminal modification encoding module and a transfer learning network; The feature representation unit performs input feature representation on each polypeptide sequence in the first data set, specifically including: feature encoding 20 natural amino acids to generate 20 natural amino acid feature vectors ,in, is the number of natural amino acids; the length of the polypeptide sequence in the first data set is In the case of As the input feature representation of the polypeptide sequence, is the number of each natural amino acid in the polypeptide sequence; the pre-trained network is trained and optimized based on the input feature representation and antibacterial activity annotation of the polypeptide sequence, and the pre-trained network outputs the hidden feature representation of the reconstructed polypeptide sequence and the antibacterial activity prediction result of the polypeptide sequence, wherein the antibacterial activity prediction result is the probability value of the polypeptide sequence having antibacterial activity; The data enhancement module is used to perform sequence-invariant noise enhancement on the input feature representation of the polypeptide sequence in the second data set generated by the feature representation unit, and the input feature representation of the polypeptide sequence after the noise enhancement is input into the pre-trained network to generate the hidden feature representation of the polypeptide sequence after the noise enhancement, specifically including: calculating the feature vector of each natural amino acid according to formula (1): Euclidean distance to the nearest neighbor of the remaining natural amino acid feature vectors : , Formula (1) in, express and The Euclidean distance of Each training step in the incremental transfer training is based on Centered on Resample uniformly in the high-dimensional feature space with radius to obtain the formula (2) The corresponding noise-enhanced feature vector of natural amino acids : , Formula (2) in, For the corresponding The noise vector , express Uniform distribution between In formula (3), As The corresponding input feature representation of the noise-enhanced peptide sequence is input into the pre-trained network to generate the hidden feature representation of the noise-enhanced peptide sequence: , Formula (3) in, For the corresponding The noise vector , For the corresponding The noise vector , For the corresponding The noise vector , , , Respectively , , Uniform distribution between The N-terminal modification encoding module is used to generate the N-terminal modification feature representation of the self-assembling polypeptide sequence in the second data set; the spliced N-terminal modification feature representation and the hidden feature representation of the polypeptide sequence after noise enhancement are input into the transfer learning network, and the deep learning model is incrementally transferred and trained in combination with the corresponding antibacterial activity annotation, and after the incremental transfer training is completed, when the self-assembling polypeptide sequence with the N-terminal modification is input into the deep learning model, the transfer learning network outputs the prediction result of the antibacterial activity of the self-assembling polypeptide sequence with the N-terminal modification; The peptide sequences in the first data set are modified at the N-terminus using a variety of compounds to generate a third data set containing multiple self-assembling peptide sequences with N-terminus modifications, and the antibacterial activity of each self-assembling peptide sequence with N-terminus modifications in the third data set is predicted using the deep learning model trained by incremental migration. The self-assembling peptide sequences with N-terminus modifications whose predicted antibacterial activity results are higher than a preset threshold are selected for synthesis and purification, and peptide drug sensitivity testing is carried out. The self-assembling peptide sequences with N-terminus modifications that pass the peptide drug sensitivity test are sequentially subjected to in vitro experimental evaluation and therapeutic effect evaluation, and the self-assembling peptide sequences with N-terminus modifications whose antibacterial activity has been verified by the evaluation are used as the designed new bionic antimicrobial peptides.
3. The design method according to claim 2, characterized in that: The amount of data in the second data set is much less than the amount of data in the first data set.
4. The design method according to claim 2, characterized in that: The first data set is a public data set, which includes positive data and negative data, wherein: The positive data came from the DBAASP database, and its peptide sequence had antibacterial activity; Negative data came from the Uniprot database, and its peptide sequence had no antibacterial activity.
5. The design method according to claim 2, characterized in that: The various compounds specifically include: octanoic acid (C8-), lauric acid (C12-), hexadecanoic acid (C16-), phenylacetic acid (PHE-), 4-biphenylacetic acid (BIP-), diphenylacetic acid (DIP-), 2-naphthylacetic acid (NAP-), 9-anthracenecarboxylic acid (ANT-), 1-pyrenebutyric acid (PYR-), cyclopropylacetic acid (C-PRO) and cyclohexylacetic acid (C-HEX).
6. The design method according to claim 2, characterized in that: The generated self-assembling polypeptide sequences with N-terminal modifications are subjected to polypeptide drug sensitivity tests and antimicrobial activity annotations are performed based on the test results, specifically including: When the peptide drug sensitivity test shows that the MIC is less than or equal to 100 μg / mL, the self-assembling peptide sequence with N-terminal modification is marked as having antibacterial activity.
7. The design method according to claim 2, characterized in that: The pre-trained network and the transfer learning network are constructed based on a Transformer network with a multi-head self-attention mechanism.
8. A pharmaceutical composition comprising the biomimetic antimicrobial peptide according to claim 1, and optionally a pharmaceutically acceptable carrier.
9. Use of the biomimetic antimicrobial peptide according to claim 1 or the pharmaceutical composition according to claim 8 in the preparation of antimicrobial drugs.
Citation Information
Patent Citations
Antibacterial peptide constructed based on artificial intelligence method and application thereof
CN118165085A
Antibacterial peptide recognition and directed evolution method based on deep learning
CN118298907A