Method for screening high-activity endogenous antibacterial peptide
Through machine learning and bioinformatics methods, the endogenous antimicrobial peptide sequences with lengths between 8 and 50 amino acid residues were screened, which solved the cumbersome and cost-effective problems of traditional antimicrobial peptide screening methods, and achieved efficient and rapid high-active antimicrobial peptide screening.
Patent Information
- Application Number
- CN202510066641.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-06
AI Technical Summary
The traditional antimicrobial peptide screening method is complicated to operate, complicated process, long cycle and high cost, and the antimicrobial peptide activity is easily affected during the isolation and purification process.
Using machine learning and bioinformatic analysis methods, endogenous antimicrobial peptide sequences with lengths of 8 to 50 amino acid residues were screened out, and highly active endogenous antimicrobial peptides were screened out through antimicrobial peptide prediction model and bioinformatic analysis.
It realizes efficient and rapid screening of highly active endogenous antimicrobial peptides, reducing operational complexity and cost, and improving screening efficiency and accuracy.
Smart Images

Figure CN120108501A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of antimicrobial peptide screening, and in particular to a method for screening highly active endogenous antimicrobial peptides. Background Art
[0002] Endogenous antimicrobial peptides (AMPs) are a class of small molecule peptides naturally produced in organisms, which play an important role in the innate immunity of organisms. These endogenous antimicrobial peptides have a wide range of antimicrobial activities and can fight against a variety of pathogens, including bacteria, fungi, viruses and parasites. Endogenous antimicrobial peptides have been found in a variety of organisms, including mammals, amphibians, insects and plants, and exist in the body fluids and tissues of a variety of organisms. With the increasing severity of the problem of antibiotic resistance, antimicrobial peptides, as a new type of antimicrobial agent, are receiving more and more attention due to their unique mechanism of action and low resistance potential. AMPs can usually act through multiple mechanisms, such as destroying cell membrane integrity, interfering with cell wall synthesis, inhibiting protein and nucleic acid metabolism, etc., making them multi-target antimicrobial agents, thereby reducing the possibility of pathogens developing resistance. In addition to direct antimicrobial effects, AMPs can also regulate host immune responses and enhance the body's defense capabilities. In addition, as natural molecules, AMPs generally have good biocompatibility and low toxicity, and some AMPs can quickly kill pathogens, which is crucial for effective infection control.
[0003] Traditional AMPs screening methods rely on separation and purification from biological materials, followed by in vitro activity verification. However, this method has many limitations, such as cumbersome operation, complex process, long cycle, high cost, and the activity of AMPs is easily affected during the separation and purification process. Therefore, there is an urgent need to develop a new method for rationally, efficiently and quickly screening highly active endogenous AMPs. Summary of the invention
[0004] In order to solve the problems of cumbersome operation, complicated process, long cycle and high cost of traditional high-activity endogenous antimicrobial peptide screening methods, the present invention provides a method for efficiently and quickly screening new high-activity endogenous antimicrobial peptides based on machine learning and bioinformatics analysis, as well as new high-activity endogenous antimicrobial peptides screened by this method.
[0005] According to a first aspect of the present invention, there is provided a method for screening highly active endogenous antimicrobial peptides, comprising the following steps:
[0006] (1) Collect known antimicrobial peptide sequences from the antimicrobial peptide database and construct an antimicrobial peptide prediction model based on a machine learning algorithm;
[0007] (2) Obtain porcine proteome data;
[0008] (3) Sequences with lengths ranging from 8 to 50 residues were selected from the porcine proteome data to construct a candidate dataset;
[0009] (4) Screening out non-redundant sequences from the candidate dataset to construct a prediction dataset;
[0010] (5) Using the antimicrobial peptide prediction model to predict endogenous antimicrobial peptides in the prediction data set to obtain candidate antimicrobial peptides;
[0011] (6) Performing bioinformatics analysis on the selected antimicrobial peptides and comparing them with the antimicrobial peptide database, screening out antimicrobial peptides with less than 100% sequence similarity to known antimicrobial peptides and having secretory protein characteristics or exosomal protein characteristics, and obtaining highly active endogenous antimicrobial peptides.
[0012] The present invention is based on all the data of the porcine proteome, applies a machine learning algorithm to screen out sequences with a length range of 8 to 50 residues, and finally screens out endogenous antimicrobial peptide sequences with high activity through sequence comparison in combination with bioinformatics means. First, the antimicrobial peptide prediction model constructed by using a machine learning algorithm based on known antimicrobial peptide sequences in an antimicrobial peptide database can learn the sequence characteristics of antimicrobial peptides and predict the inherent activity and antimicrobial potential of candidate peptide segments by training a data set containing a large number of known antimicrobial peptides. The accuracy of the model directly affects the quality of the candidate antimicrobial peptides, thereby affecting the number of highly active antimicrobial peptides finally screened. The antimicrobial peptide prediction model constructed by the above method has the characteristic of high accuracy, which is conducive to further improving the accuracy of the screening results of highly active endogenous antimicrobial peptides. Second, based on the length distribution law of known antimicrobial peptides, sequences with a length of 8 to 50 amino acid residues are selected for constructing a candidate data set. This length range covers the vast majority of peptide segments with antimicrobial activity and is sufficient to form a secondary structure with antimicrobial activity. For example, Such as α-helix or β-fold, and at the same time small enough to pass through the bacterial cell membrane. By screening sequences within the above length range for the construction of candidate data sets, the number of candidate peptides can be effectively reduced, the screening efficiency can be improved, and the consumption of computing resources can be reduced. It also helps to exclude peptides that are unlikely to have antibacterial activity, thereby improving the accuracy of the final screening results. Peptides that are too short may not be able to form a stable structure, while peptides that are too long may be difficult to synthesize, have poor solubility or have other unfavorable properties; third, in the process of constructing the prediction data set, redundant sequences (i.e., all repeated sequences in the candidate data set) are removed, and the remaining non-redundant sequences constitute the prediction data set for subsequent machine learning model prediction, which can reduce computational redundancy (i.e., the subsequent amount of calculation) and improve the efficiency and accuracy of subsequent steps; fourth, in the process of the antimicrobial peptide prediction model predicting endogenous antimicrobial peptides in the prediction data set, the prediction ability of the machine learning model can be used to quickly enrich potential antimicrobial peptides from a large number of candidate peptides, providing a basis for subsequent experimental verification;Fifth, the premise for antimicrobial peptides to work is that they can reach the site of action (i.e., the location of the bacteria). By screening secretable antimicrobial peptides (i.e., peptides with secretory protein characteristics or exosome protein characteristics) through bioinformatics analysis, it can be ensured that these peptides can be secreted outside the cell or encapsulated in exosomes, thereby more effectively contacting and killing bacteria. This step of screening is crucial to improving the activity of antimicrobial peptides, which is conducive to making the endogenous antimicrobial peptides finally screened show higher activity, because even if the peptides with antimicrobial potential cannot be secreted, it is difficult to exert their antimicrobial activity. Specifically, by analyzing the candidate peptides, Select the subcellular localization information of the protein where the antimicrobial peptide is located, such as the presence or absence of signal peptides, transmembrane domains, etc., to determine whether it is a secretory protein or an exosome protein. Antimicrobial peptides located in the secretory pathway, outside the cell or in the exosome are more likely to reach the site of action, thereby showing higher activity; Sixth, the purpose of comparing the candidate antimicrobial peptides with the known antimicrobial peptide database is to exclude known antimicrobial peptides and screen out candidate antimicrobial peptides with a similarity of less than 100% with any sequence in the database, that is, a completely new antimicrobial peptide sequence, which ensures that the highly active endogenous antimicrobial peptides finally obtained are completely new and have potential novelty and application value. ;
[0013] Therefore, the method provided by the present invention can reasonably, efficiently and quickly screen out highly active endogenous antimicrobial peptides through the combined action of each step, and can solve the problems of cumbersome separation and extraction process, complicated process, long cycle and low purity in the process of activity verification after separation and extraction of antimicrobial peptides from biological materials, and at the same time provide a new direction for mining highly active endogenous antimicrobial peptides. Moreover, compared with the traditional method of extracting antimicrobial peptides, the screening method provided by the present invention is more efficient, has good applicability and safety, and the highly active endogenous antimicrobial peptides screened out by the method provided by the present invention have broader application prospects.
[0014] In addition, in animal husbandry, animals are susceptible to bacterial infections and traditionally rely on antibiotics for treatment. However, long-term use of antibiotics can lead to the emergence of drug-resistant strains, threatening animal health and food safety. The method for screening highly active endogenous antimicrobial peptides provided by the present invention can efficiently and quickly discover new antimicrobial peptides, providing a new solution to the problems of antibiotic abuse and bacterial resistance in animal husbandry. The highly active endogenous antimicrobial peptides screened and obtained using the method provided by the present invention, as a natural antimicrobial agent, have the advantages of broad-spectrum antimicrobial activity and being less likely to develop drug resistance. They are expected to replace or reduce the use of antibiotics, promote healthy breeding, and improve the quality of livestock products, thereby promoting the sustainable development of agriculture and food safety. This has important application prospects in reducing animal diseases and ensuring the safety of livestock products, and helps to achieve healthy development in the "three rural" areas.
[0015] Preferably, in step (1), the machine learning algorithm includes at least one of a classification model based on the physicochemical characteristics of peptides and a classification model based on the language model BERT.
[0016] Preferably, the classification model based on the physicochemical characteristics of peptides is constructed by the following steps: using a support vector machine model and / or a random forest model algorithm to extract the physicochemical characteristics of peptides of known antimicrobial peptide sequences, and randomly dividing the known antimicrobial peptide sequences into a training set and a test set based on an antimicrobial peptide database, using an elastic network to screen the features in the training set and the test set, constructing a first initial model, optimizing the parameters in the first initial model through five-fold cross validation, and using the test set to evaluate the performance of the first initial model to obtain a classification model based on the physicochemical characteristics of peptides.
[0017] Preferably, the classification model based on the language model BERT is constructed by the following steps: using a pre-trained BERT model to randomly divide the known antimicrobial peptide sequences into a training set and a test set, performing parameter optimization, constructing a second initial model, optimizing the parameters in the second initial model through five-fold cross validation, and using the test set to evaluate the performance of the second initial model, thereby obtaining a classification model based on the language model BERT.
[0018] Preferably, in the steps of constructing a classification model based on peptide physicochemical features and a classification model based on the language model BERT, in the process of randomly dividing the known antimicrobial peptide sequences into a training set and a test set, the ratio of the number of antimicrobial peptide sequences in the training set to that in the test set is 0.7:0.3.
[0019] In the construction steps of the classification model based on the physicochemical characteristics of peptides and the classification model based on the language model BERT, the known antimicrobial peptide sequences are randomly divided into training sets and test sets according to the above ratios, and after the initial model is constructed, the parameters in the initial model are tested and optimized based on different machine learning algorithms to screen out parameter combinations with better prediction effects; Elastic Net is a feature screening method that determines whether a peptide has antimicrobial characteristics. Not all features are important, so in the process of building a classification model based on the physicochemical characteristics of peptides, an elastic network is used to screen the features in the training set and the test set. Thus, through the above technical means, the classification model based on the physicochemical characteristics of peptides and the classification model based on the language model BERT are given better prediction performance, thereby making the prediction effect of the obtained antimicrobial peptide prediction model better and the prediction results more accurate, thereby improving the accuracy of the screening results of highly active endogenous antimicrobial peptides, so as to obtain highly active endogenous antimicrobial peptides.
[0020] Preferably, the peptide physicochemical characteristics include at least one of hydrophobicity, net charge, average hydrophobic moment, and secondary structure tendency, and the language model includes a BERT model.
[0021] Preferably, step (3) comprises the following operations: after screening out sequences with a length ranging from 8 to 50 residues from the porcine proteome data, scoring the sequences using a classification model based on peptide physicochemical features and a classification model based on the language model BERT, integrating the scoring results of the scoring model based on peptide physicochemical features and the classification model based on the language model BERT to obtain a final comprehensive score, sorting the sequences in descending order according to the comprehensive score, and constructing a candidate data set from a set of the top 1000 sequences for subsequent screening.
[0022] The above-mentioned comprehensive score is obtained by integrating the scoring results of the classification model based on the physicochemical characteristics of peptides and the classification model based on the language model BERT using a weighted average strategy. The scoring results of the classification model based on the physicochemical characteristics of peptides and the classification model based on the language model BERT are integrated to obtain the final comprehensive score. The integration strategy adopts weighted average, and the weights can be adjusted according to the performance of different models.
[0023] By combining the physicochemical characteristics of peptides with the BERT language model, first, the information of peptide sequences can be captured more comprehensively, thereby more accurately predicting their antimicrobial activity, improving screening efficiency, and ultimately obtaining more active endogenous antimicrobial peptides. Second, the top 1,000 sequences with the highest scores are selected to strike a balance between computing resources and screening efficiency. Third, the use of weighted average or other fusion strategies can adjust the scores according to the performance of different models to obtain the best prediction results.
[0024] Preferably, in step (6), the obtained candidate antimicrobial peptides are compared with sequences in the antimicrobial peptide database using the BLAST tool.
[0025] Preferably, the antimicrobial peptide database includes at least one of APD and CAMP.
[0026] Preferably, in step (6), during the bioinformatics analysis, a secretory protein prediction tool or an exosome protein prediction tool is used to predict the candidate antimicrobial peptides, the secretory protein prediction tool includes at least one of SignalP, TargetP, and SecretomeP, and the exosome protein prediction tool includes at least one of Exocarta, Vesiclepedia, and FunRich.
[0027] Preferably, in step (6), after screening for antimicrobial peptides having a sequence similarity with the known antimicrobial peptide of less than 100% and having secretory protein characteristics or exosomal protein characteristics, the method further includes a step of verifying the antibacterial activity of the antimicrobial peptides.
[0028] Preferably, the antibacterial activity of the antimicrobial peptide is verified according to the following steps: different concentrations of the antimicrobial peptide are mixed with at least one of Escherichia coli, Staphylococcus aureus, and Thermomyces cerevisiae, and cultured at 25-37°C for 12-36 hours, and the antibacterial activity of the antimicrobial peptide is evaluated by antibacterial experimental indicators, and the antibacterial activity indicators include at least one of the minimum inhibitory concentration, the minimum bactericidal concentration, the growth curve, the total number of colonies, and the diameter of the inhibition zone.
[0029] According to a second aspect of the present invention, there is provided a highly active endogenous antimicrobial peptide, which is obtained by screening using the above method for screening highly active endogenous antimicrobial peptides.
[0030] Preferably, the above-mentioned highly active endogenous antimicrobial peptides include at least one of A0A4X1T497_PIG and A0A4X1THB8_PIG.
[0031] The antibacterial characterization experiment proved that the endogenous antimicrobial peptides screened by the method provided by the present invention have the characteristics of high antibacterial activity, which shows that the method for screening high-activity endogenous antimicrobial peptides provided by the present invention is feasible and can reasonably, efficiently and quickly screen out endogenous antimicrobial peptides with high activity. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 The motif logo display result is for some common sequences of pig proteome.
[0033] Figure 2 This is the minimum inhibitory concentration and minimum bactericidal concentration of A0A4X1T497_PIG and lactobacillus streptococcus against Escherichia coli, Staphylococcus aureus and Thermostatsus thermophilus.
[0034] Figure 3 This is a graph showing the bacterial growth curves of A0A4X1T497_PIG and lactobacillus streptococcus against Escherichia coli, Staphylococcus aureus and Bacillus thermophilus.
[0035] Figure 4 This is the experimental result of the total colony count of three bacteria, namely, Escherichia coli, Staphylococcus aureus and Thermolyticus.
[0036] Figure 5 This is the experimental result of the inhibition zone diameter of A0A4X1T497_PIG and lactobacillus streptococcus against Escherichia coli, Staphylococcus aureus and Thermostatsus thermophilus. DETAILED DESCRIPTION
[0037] The following is a further clear and complete description of the technical features in the technical solution provided by the present invention in conjunction with the specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0038] Example 1
[0039] A method for screening highly active endogenous antimicrobial peptides, the method comprising the following steps:
[0040] Step 1: Construct an antimicrobial peptide prediction model
[0041] Known antimicrobial peptide sequence data were collected from the antimicrobial peptide database, and an antimicrobial peptide prediction model was constructed based on machine learning algorithms (including a classification model based on the physicochemical characteristics of peptides and a classification model based on the language model BERT).
[0042] The classification model based on the physicochemical characteristics of peptides is constructed by the following steps: using a support vector machine model and / or a random forest model algorithm to extract the physicochemical characteristics of peptides of known antimicrobial peptide sequences, and based on the antimicrobial peptide database, the known antimicrobial peptide sequences are randomly divided into a training set and a test set at a quantity ratio of 0.7:0.3, and an elastic network is used to screen the features in the training set and the test set to construct a first initial model, and the parameters in the first initial model are optimized by five-fold cross validation, and the performance of the first initial model is evaluated using the test set to obtain a classification model based on the physicochemical characteristics of peptides.
[0043] The classification model based on the language model BERT was constructed through the following steps: using the pre-trained BERT model to randomly divide the known antimicrobial peptide sequences into training sets and test sets according to the quantity ratio of 0.7:0.3, and optimizing the parameters to construct the second initial model, optimizing the parameters in the second initial model through five-fold cross validation, and using the test set to evaluate the performance of the second initial model to obtain the classification model based on the language model BERT.
[0044] The specific construction steps of the classification model based on the physicochemical characteristics of peptides and the classification model based on the language model BERT are further detailed:
[0045] a) Extraction of physicochemical characteristics: Extract a series of physicochemical characteristics from the collected known antimicrobial peptide sequences, such as hydrophobicity, net charge, average hydrophobic moment, secondary structure tendency, etc.;
[0046] b) BERT embedding: Use the pre-trained BERT model to convert the peptide sequence into a vector representation (embedding). The BERT model can capture the contextual and semantic information in the peptide sequence, thereby providing a richer feature representation;
[0047] c) Model training and scoring integration: Use the collected known antimicrobial peptide sequences to train models based on physicochemical features, such as support vector machines, random forests, etc., to extract a series of physicochemical features, such as hydrophobicity, net charge, average hydrophobic moment, secondary structure tendency, etc.; You can choose to use the collected known antimicrobial peptide sequences to train the BERT model to better adapt to the antimicrobial peptide prediction task;
[0048] d) Python implementation: The model construction, training and prediction processes are all implemented using Python.
[0049] Step 2: Obtaining pig proteome data
[0050] The canonical and isoform sequences of all porcine proteins were downloaded from the UniProt database (https: / / www.uniprot.org / uniprotkb?query=pig).
[0051] Step 3: Construct a pig peptide dataset (i.e., candidate dataset)
[0052] A Python script was used to screen out sequences with a length range of 8 to 50 amino acid residues from the porcine protein sequences obtained above, forming a candidate dataset containing approximately 43,000 peptides. Some sequences in the candidate dataset are shown in Table 1.
[0053] The specific steps for constructing the candidate data set are as follows: after screening out sequences with a length range of 8 to 50 residues from the porcine proteome data, the sequences are scored using a classification model based on the physicochemical characteristics of peptides and a classification model based on the language model BERT. The scoring results of the scoring model based on the physicochemical characteristics of peptides and the classification model based on the language model BERT are integrated to obtain the final comprehensive score. The sequences are sorted in descending order according to the comprehensive score, and the set of sequences with the top 1000 scores is constructed to obtain a candidate data set for subsequent screening.
[0054] Table 1 Some porcine protein peptide sequences in the candidate dataset
[0055]
[0056]
[0057] Step 4: Build a prediction dataset
[0058] Redundant sequences were removed from the candidate dataset and non-redundant sequences were screened out to construct the prediction dataset. The specific operations were as follows: A Python script was used to compare the sequences in the candidate dataset and remove duplicate peptides with exactly the same sequences to ensure that the prediction dataset contained only unique non-redundant sequences.
[0059] Step 5: Machine Learning Prediction
[0060] The antimicrobial peptide prediction model is used to predict the endogenous antimicrobial peptides in the prediction data set to obtain candidate antimicrobial peptides. Specifically, the machine learning model constructed in step one (i.e., the antimicrobial peptide prediction model) is used to predict the antimicrobial activity of all non-redundant sequences in the prediction data set obtained in step four. The specific steps are as follows: Each peptide sequence in the prediction data set is input into the trained machine learning model. The model will output a prediction score representing the antimicrobial potential of the peptide based on the learned antimicrobial peptide features. Based on the score, the prediction results are sorted, and peptides with higher scores are selected as candidate antimicrobial peptides for subsequent bioinformatics analysis and new peptide screening.
[0061] Step 6. Bioinformatics analysis and new peptide screening
[0062] Bioinformatics analysis was performed on the candidate antimicrobial peptides obtained in step 5, including the use of tools such as SignalP, TargetP, and SecretomeP to predict the secretion characteristics of the protein in which the candidate peptides were located, and the use of tools such as Exocarta, Vesiclepedia, and FunRich to predict the characteristics of exosome proteins. At the same time, the candidate antimicrobial peptides were compared with the sequences in the antimicrobial peptide database dbAMP (https: / / awi.cuhk.edu.cn / dbAMP / index.php). Candidate antimicrobial peptides with a sequence similarity of less than 100% with known antimicrobial peptides and having secretory protein characteristics or exosome protein characteristics were screened out to obtain highly active endogenous antimicrobial peptides. Some of the comparison results are shown in Table 2. The highly active endogenous antimicrobial peptides finally screened are shown in Table 3. Among them, A0A4X1T497_PIG is a protein containing an OAS1 domain, which may be involved in the innate immune process of the organism. A0A4X1THB8_PIG is a 2'-5'-oligoadenylate synthetase, which may be involved in the immune response to viruses in the organism. Moreover, the length of naturally occurring endogenous antimicrobial peptides (AMPs) is usually between 12 and 50 amino acids, so the A0A4X1T497_PIG fragment with a length of 24 amino acid residues and the A0A4X1THB8_PIG fragment with a length of 25 amino acid residues were initially selected as highly active endogenous antimicrobial peptides for subsequent experiments.
[0063] Table 2 Some secreted protein peptide sequences
[0064]
[0065] Table 3 Highly active endogenous antimicrobial peptides finally screened
[0066]
[0067] Example 2
[0068] This example aims to perform consensus sequence analysis on two highly active endogenous antimicrobial peptides, A0A4X1T497_PIG and A0A4X1THB8_PIG, which were screened using the method for screening highly active endogenous antimicrobial peptides provided in Example 1.
[0069] MEME is a tool for searching similar fragments from a large number of sequences. Studies have shown that the common sequence of proteins can be related to their target sites, so there is reason to believe that similar sequence fragments searched using MEME may be domains with certain functions, which are related to the action site or activity of antimicrobial peptides.
[0070] In order to further verify the functions of the two highly active endogenous antimicrobial peptides screened above, MEME-STREME was first used to perform consensus sequence analysis on the core amino acid motifs (consensus sequence) of the two peptides.
[0071] A0A4X1T497_PIG is:
[0072] [HS]-RP-[TAV]-KLK-[SN]-LLRLVKHWYLK-[YC]-V;
[0073] A0A4X1THB8_PIG is:
[0074] [HN]-RP-[AT]-KLKSL-[LI]-RLVKHWY-[LQ]-[KT]-[YV]-[VW].
[0075] In addition, we selected the A0A4X1W7A2_PIG sequence fragment with a similarity of 100% as a control, which is 22 amino acid residues long. Using MEME-STREME, we obtained a total of 27 consensus sequences. The results are as follows: Figure 1 As shown, a total of 27 motifs were obtained. Figure 1 Only four of them are shown in Figure 1 It can be seen that the amino acid composition rules or patterns of antimicrobial peptides are obtained, wherein the horizontal and vertical axes represent different amino acid residues and their frequencies of occurrence.
[0076] Example 3
[0077] The purpose of this example is to select A0A4X1T497_PIG with lower similarity from the two highly active endogenous antimicrobial peptides A0A4X1T497_PIG and A0A4X1THB8_PIG obtained by the method for screening highly active endogenous antimicrobial peptides provided in Example 1 to conduct an antibacterial characterization experiment. The antibacterial activity of the antimicrobial peptides is verified by conducting an antibacterial characterization experiment on the screened sequence peptides. Different antibacterial experimental indicators are used to reflect that the method for screening highly active endogenous antimicrobial peptides provided by the present invention is feasible.
[0078] The strains used in this example are as follows: Escherichia coli (E.coli) O157:H7 is from Beina Chuanglian Biotechnology Co., Ltd.; Staphylococcus aureus (S.aureus) ATCC25923 is from Beina Chuanglian Biotechnology Co., Ltd.; Brochothrix.thermosphacta (B.thermo) is from the School of Food and Bioengineering, Hefei University of Technology.
[0079] The reagents and materials used in this example are as follows: Nisin (≥1000 IU / mg) and turbidity standards were purchased from Shanghai Aladdin Reagent Co., Ltd.; tryptone soy broth (TSB), plate technique agar (PCA), crystal violet neutral red bile glucose agar (VRBGA), and mannitol sodium chloride agar (MSA) were purchased from Beijing Luqiao Technology Co., Ltd.; the antimicrobial peptide A0A4X1T497_PIG (purity >98%) was synthesized using synthetic biology methods; other chemical analysis reagents were purchased from Sinopharm Chemical Reagent Co., Ltd.; Oxford cups (inner diameter 6 mm, outer diameter 7.9 mm, height 10 mm) were purchased from the local offline market.
[0080] The instruments used in this embodiment are as follows: a multifunctional microplate reader SynergyHTX, purchased from Berton Instrument Co., Ltd. of the United States; a KLCZ-880A clean bench, purchased from Beijing Yatai Cologne Instrument Technology Co., Ltd.; a vertical pressure steam sterilizer, purchased from Shanghai Boxun Medical Biological Instrument Co., Ltd.; a multifunctional microplate reader SynergyHTX, purchased from Berton Instrument Co., Ltd. of the United States; and an ultrapure water machine, purchased from Hefei Hongke Instrument Co., Ltd.
[0081] The minimum inhibitory concentration refers to the concentration of antimicrobial peptides when the bacterial solution becomes visibly clear; the minimum bactericidal concentration refers to the concentration of antimicrobial peptides when the bacterial growth is completely consistent.
[0082] In this example, two highly active endogenous antimicrobial peptides, A0A4X1T497_PIG and A0A4X1THB8_PIG, were synthesized, and antibacterial characterization experiments were performed on Escherichia coli (pathogenic bacteria), Staphylococcus aureus (pathogenic bacteria) and Pseudomonas thermophila (corrupt bacteria). The antibacterial characterization experiment indicators included minimal inhibitory concentration (MIC), minimal bactericidal concentration (MBC), bacterial growth curves, total colony counts (survival bacteria clones) and inhibition zone diameters (Inhibition zone diameters).
[0083] 1. Antibacterial characterization experiment
[0084] The specific operation steps of the antibacterial characterization experiment are as follows:
[0085] (1) Escherichia coli (stored in a 25% glycerol tube at -80°C) was streaked onto a VRBGA plate, cultured at 37°C for 24 h, and then a rose-red single colony was picked up and placed in TSB (containing 3% NaCl) broth and cultured overnight in a constant temperature incubator at 37°C and 200 rpm;
[0086] (2) Staphylococcus aureus (stored in 25% glycerol tubes at -80°C) was streaked on MSA plates, cultured at 37°C for 24 h, and then single yellow colonies were picked and placed in TSB (containing 3% NaCl) broth and cultured overnight in a constant temperature incubator at 37°C and 200 rpm;
[0087] (3) Streak the heat-killed Sossella (stored in 25% glycerol tubes at -80°C) on a PCA plate, incubate at 25°C for 24 h, pick a single white colony into TSB broth and incubate overnight in a constant temperature incubator at 25°C and 200 rpm;
[0088] (4) The concentrations of the bacterial solutions of Escherichia coli, Staphylococcus aureus and Thermococcus thermophilus were adjusted with 0.85% NaCl solution, and the concentrations of the three bacterial solutions were determined to be approximately 8 log (CFU / mL) using a turbidity standard (McFarland turbidity method);
[0089] (5) The three prepared bacterial solutions were diluted in equal proportions with fresh TSB broth in a test tube to 7 log CFU / mL. A0A4X1T497_PIG and nisin in the test were both cationic peptides, which were dissolved in 0.02 mol / L HCl to prepare a mother solution concentration of 2000 μg / mL.
[0090] (6) Preparation of 96-well plate for antimicrobial activity test: Add antimicrobial peptides to a 96-well plate (flat bottom), add 2 μL of compound to the first column, add 198 μL of bacterial solution, add 100 μL of bacterial solution to columns 2 to 9, use a spray gun to thoroughly pipette 3 times in the first column to mix the compound and bacterial solution, then pipette 100 μL into the second column and pipette to mix thoroughly, repeat to the eighth column, the ninth column is the growth control well, each group has 3 parallels, and the final test concentrations of the compound are 200, 100, 50, 25, 12.5, 6.25, 3.13, 1.56, 0 μg / mL;
[0091] (7) The prepared 96-well plates of Escherichia coli and Staphylococcus aureus were placed in a constant temperature incubator at 37°C for 24 hours, and the prepared 96-well plates of heat-killed Sordella were placed in a constant temperature incubator at 25°C for 24 hours. The OD600 was measured by naked eye observation and microplate reader and the results were recorded. All experiments were repeated three times, and 3 parallel samples were set for each group of samples. The average value was taken after 3 independent repeated experiments. Data were processed by Microsoft Excel 2010, and one-way analysis of variance was performed by SPSS21.0. There was a significant difference (P < 0.05), and Origin 2021 was used for graphical analysis.
[0092] In the antibacterial characterization experiment, Nisin was added as a positive control, and the selected sequence peptides were compared with the mature commercial antimicrobial peptides. The concentration used was consistent with the maximum dosage of 0.5 mg / mL of Nisin stipulated in the "National Food Safety Standard-Food Additive Usage Standard". The minimum inhibitory concentration and minimum bactericidal concentration of A0A4X1T497_PIG and Nisin against the above three bacteria in the antibacterial characterization experiment are shown as follows: Figure 2 As shown in the figure, the minimum inhibitory concentration is the antimicrobial peptide concentration corresponding to a column that is relatively clearer than the growth control well after the mixture of the bacterial solution and the antimicrobial peptide is observed by naked eyes, the minimum bactericidal concentration is the concentration that completely inhibits the growth of bacteria, MIC means minimum inhibitory concentration, MBC means minimum bactericidal concentration, Nisin means nisin, and T479 means A0A4X1T497_PIG.
[0093] Depend on Figure 2It can be seen that nisin has no inhibitory effect on Escherichia coli (Gram-negative bacteria), while the minimum inhibitory concentration of A0A4X1T497_PIG on Escherichia coli is 6.25μg / mL and the minimum bactericidal concentration is 12.5μg / mL; nisin has a significant inhibitory effect on Staphylococcus aureus (Gram-positive bacteria), with a minimum inhibitory concentration of 25μg / mL and a minimum bactericidal concentration of 100μg / mL, while A0A4X1 The minimum inhibitory concentration of W7A2_PIG against Staphylococcus aureus is 25μg / mL, and the minimum bactericidal concentration is 100μg / mL; the minimum inhibitory concentration of nisin against thermophilic Sordella (Gram-positive bacteria) is 12.5μg / mL, and the minimum bactericidal concentration is 25μg / mL, and the minimum inhibitory concentration of A0A4X1T497_PIG against thermophilic Sordella is 12.5μg / mL, and the minimum bactericidal concentration is 25μg / mL.
[0094] The results of the above antibacterial characterization experiments show that the endogenous antimicrobial peptides screened out by the method for screening highly active endogenous antimicrobial peptides provided by the present invention have an inhibitory effect on both Gram-negative and Gram-positive bacteria, and the effect is obvious. Compared with nisin, it has a broader spectrum of antibacterial effect.
[0095] 2. Growth Curve Determination
[0096] By measuring the bacterial growth curve, the inhibitory or killing effect of antimicrobial peptides on specific bacteria can be intuitively evaluated, and the changes in the growth curve can reflect the intensity of the antimicrobial peptide's action.
[0097] The maximum addition amount of lactobacillus streptococcus is specified as 0.5 mg / mL in the "National Food Safety Standard for the Use of Food Additives".
[0098] Referring to the above antibacterial characterization experiment operation steps, the three bacterial solutions of Escherichia coli, Staphylococcus aureus and Thermocidal Sothiae were diluted in a test tube with fresh TSB broth in equal proportions to 10 5 CFU / mL, A0A4X1T497_PIG and nisin in the test were dissolved in 0.02 mol / L HCl, and the concentration of the prepared stock solution was 5000 μg / mL; 9 mL of the above three bacterial solutions were added to the test tubes respectively, and 1 mL of the antimicrobial peptide (A0A4X1T497_PIG) stock solution was added to the test tubes to make its concentration become the maximum amount of 500 μg / ml in the national standard, and the OD was measured by an enzyme marker after 0, 2, 4, 6, 8, 10, 12, and 16 hours respectively. 600 All experiments were repeated three times. The growth curves of Escherichia coli, Staphylococcus aureus and Thermostats were determined according to the above experimental steps. The results are shown in Figure 3 As shown, Figure 3A is the growth curve of Staphylococcus aureus (S. aureus), Figure 3 B is the growth curve of Escherichia coli (E.coli), Figure 3 C is the growth curve of B. thermophila (B. thermo), Nisin represents the positive control group with the addition of nisin, T497 represents the experimental group with the addition of antimicrobial peptide A0A4X1T497_PIG, and Control represents the blank control group without the addition of antimicrobial peptide A0A4X1T497_PIG and nisin.
[0099] Depend on Figure 3 A shows that nisin and A0A4X1T497_PIG have obvious inhibitory effects on Staphylococcus aureus (Gram-positive bacteria). After 12 hours, the absorbance OD 600 The antibacterial activities of the two against Staphylococcus aureus were almost parallel. Figure 3 B shows that the growth curve of nisin against Escherichia coli maintains the same trend as the blank control group, while A0A4X1T497_PIG has a significant antibacterial effect, which also verifies that nisin has no inhibitory effect on most Gram-negative bacteria; Figure 3 C shows that both nisin and A0A4X1T497_PIG have an inhibitory effect on heat-killed Sothia (Gram-positive bacteria), and the absorbance OD 600 The value always remains around 0.2.
[0100] The above growth curve measurement results indicate that the endogenous antimicrobial peptides screened out by the method for screening highly active endogenous antimicrobial peptides provided by the present invention have obvious inhibitory effects on Escherichia coli, Staphylococcus aureus and Bacillus thermophila, and are superior to nisin in inhibiting Gram-negative bacteria.
[0101] 3. Determination of total colony count
[0102] The total colony count is measured by the plate counting method to detect the number of live bacteria. This method can count the live bacteria in the plate and visually show the bactericidal efficiency of antimicrobial peptides on bacteria.
[0103] Referring to the operating steps of the antibacterial characterization experiment and the production curve determination, the three bacterial solutions of Escherichia coli, Staphylococcus aureus and heat-killed Sothia bacteria were diluted in equal proportions with fresh TSB broth in a test tube to 6 log (CFU / mL), and the experimental method was consistent with the above-mentioned growth curve determination; after the bacterial solution and the antimicrobial peptide A0A4X1T497_PIG were mixed and reacted for half an hour, 1 mL of the mixture was added to the PCA plate and evenly coated. The coated Escherichia coli plate and Staphylococcus aureus plate were placed in a constant temperature incubator at 37°C for 24 hours, and the heat-killed Sothia bacteria plate was placed in a constant temperature incubator at 25°C for 24 hours, and counted. All experiments were repeated three times. The results of the total colony count determination of Escherichia coli, Staphylococcus aureus and heat-killed Sothia bacteria by antimicrobial peptide A0A4X1T497_PIG and lactic acid bacteria streptococcus are shown in the figure. Figure 4 As shown, Ⅰ represents the blank control group without adding antimicrobial peptide A0A4X1T497_PIG and nisin, Ⅱ represents the experimental group with adding antimicrobial peptide A0A4X1T497_PIG, and Ⅲ represents the positive control group with adding nisin.
[0104] Depend on Figure 4 The survival of three bacterial clones on agar plates after adding antimicrobial peptide A0A4X1T497_PIG or nisin was shown. The colony counts of Escherichia coli, Staphylococcus aureus and Thermocidol sutschott in the blank control group were 6.52log (CFU / mL), 6.76log (CFU / mL) and 6.34log (CFU / mL), respectively. After nisin treatment, only a few single colonies grew on the plates of Staphylococcus aureus and Thermocidol sutschott, while Escherichia coli The colony count on the plate was 6.22log (CFU / mL), indicating that the inhibition rate of nisin against Staphylococcus aureus and thermophilic bacteria was more than 99.99%, and it had no inhibitory effect on Escherichia coli; after treatment with A0A4X1T497_PIG, the colony count on the plate of Staphylococcus aureus was 1.02CFU / mL, while the plates of Escherichia coli and thermophilic bacteria were almost empty, indicating that the inhibition rate of the endogenous antimicrobial peptide A0A4X1T497_PIG against these three bacteria was more than 99.99%.
[0105] The above-mentioned total colony count determination results indicate that the endogenous antimicrobial peptides screened out by the method for screening highly active endogenous antimicrobial peptides provided by the present invention have significant inhibitory effects on Escherichia coli, Staphylococcus aureus and Thermococcus thermophilus, with the inhibition rates all reaching over 99.99%.
[0106] 4. Determination of the diameter of the inhibition zone
[0107] The inhibition zone diameter was determined using the Oxford cup diffusion method, which provides a direct visual indicator for quickly observing the effect of antimicrobial agents on bacterial growth.
[0108] Referring to the operating steps of the antibacterial characterization experiment, production curve and total colony count determination, the three bacterial solutions of Escherichia coli, Staphylococcus aureus and Thermocidol thiocyanate were diluted in equal proportions with fresh TSB broth in a test tube to 6 log (CFU / mL); 100 μL of the above three bacterial solutions were respectively spread on the PCA plate, and the sterilized Oxford cup was placed on the freshly coated plate with tweezers, and 200 μL of antimicrobial peptide A0A4X1T497_PIG was added to the plate. After culturing in a constant temperature incubator for 24 hours according to the temperature of the total colony count determination experiment, the diameter of the inhibition zone was observed and recorded. All experiments were repeated three times. The observation and determination results of the diameter of the inhibition zone are shown in the figure. Figure 5 As shown, Figure 5 A is the observation result of the diameter of the inhibition zone of Escherichia coli, Staphylococcus aureus and thermophilic bacteria. Figure 5 B is the result graph of the diameter of the inhibition zone of Escherichia coli, Staphylococcus aureus and thermophilic bacteria plates. Nisin or N represents the positive control group with the addition of nisin, T497 or T represents the experimental group with the addition of antimicrobial peptide A0A4X1T497_PIG, and Control represents the blank control group without the addition of antimicrobial peptide A0A4X1T497_PIG and nisin.
[0109] Depend on Figure 5 It can be seen that no inhibition zone was generated in the blank control group, and no inhibition zone was generated by nisin in the Escherichia coli plate. The diameter of the inhibition zone in the Staphylococcus aureus plate was 14.2±0.12 mm, and the diameter of the inhibition zone in the heat-killed Sothia bacteria plate was 10.13±0.25 mm; the diameter of the inhibition zone of the antimicrobial peptide A0A4X1T497_PIG in the Escherichia coli plate was 8.7±0.14 mm, the diameter of the inhibition zone in the Staphylococcus aureus plate was 13.5±0.23 mm, and the diameter of the inhibition zone in the heat-killed Sothia bacteria plate was 11.21±0.18 mm.
[0110] The above results can be summarized to show that the method for screening highly active endogenous antimicrobial peptides provided by the present invention cleverly combines peptidomics, machine learning and bioinformatics technologies. First, a machine learning algorithm is used to quickly predict potential antimicrobial peptides, and then the secretion characteristics and exosome localization of peptide segments are predicted by bioinformatics means, and known antimicrobial peptides are excluded, and finally the antibacterial activity of candidate peptides is verified by in vitro experiments. The method of the present invention can reasonably, efficiently and quickly screen out new highly active endogenous antimicrobial peptide sequences, avoid the cumbersome process of separation, extraction and activity verification of traditional methods, significantly improve the screening efficiency, and provide a new direction for the development of new antimicrobial peptides. Compared with traditional methods, the screening method provided by the present invention has the advantages of simple operation, low cost, good applicability and high safety. The new highly active endogenous antimicrobial peptides screened out by the method provided by the present invention have broader application prospects.
[0111] The above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the protection scope of the present invention. Although the present invention is described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the technical solutions of the present invention can be modified or equivalently replaced, but these modifications or replacements are all within the protection scope of the present invention.
Claims
1. A method for screening highly active endogenous antimicrobial peptides, characterized in that: The following steps are involved: (1) Collect known antimicrobial peptide sequences from the antimicrobial peptide database and construct an antimicrobial peptide prediction model based on a machine learning algorithm; (2) Obtain porcine proteome data; (3) selecting sequences with a length ranging from 8 to 50 residues from the porcine proteome data to construct a candidate data set; (4) selecting non-redundant sequences from the candidate data set to construct a prediction data set; (5) predicting endogenous antimicrobial peptides in the prediction data set using the antimicrobial peptide prediction model to obtain candidate antimicrobial peptides; (6) performing bioinformatics analysis on the candidate antimicrobial peptides and comparing them with the antimicrobial peptide database to screen out antimicrobial peptides whose sequence similarity to the known antimicrobial peptides is less than 100% and which have secretory protein characteristics or exosomal protein characteristics, thereby obtaining the highly active endogenous antimicrobial peptides.
2. The method for screening highly active endogenous antimicrobial peptides according to claim 1, characterized in that: In step (1), the machine learning algorithm includes at least one of a classification model based on the physicochemical characteristics of peptides and a classification model based on the language model BERT.
3. The method for screening highly active endogenous antimicrobial peptides according to claim 2, characterized in that: The classification model based on peptide physicochemical characteristics is constructed by the following steps: using the support vector machine model and / or the random forest model algorithm to extract the peptide physicochemical characteristics of the known antimicrobial peptide sequence, and randomly dividing the known antimicrobial peptide sequence into a training set and a test set based on the antimicrobial peptide database, using an elastic network to screen the features in the training set and the test set to construct a first initial model, optimizing the parameters in the first initial model through five-fold cross validation, and using the test set to evaluate the performance of the first initial model to obtain the classification model based on peptide physicochemical characteristics; The classification model based on the language model BERT is constructed by the following steps: using a pre-trained BERT model to randomly divide the known antimicrobial peptide sequence into a training set and a test set, performing parameter optimization, constructing a second initial model, optimizing the parameters in the second initial model through five-fold cross validation, and using the test set to evaluate the performance of the second initial model to obtain the classification model based on the language model BERT.
4. The method for screening highly active endogenous antimicrobial peptides according to claim 3, characterized in that: The peptide physicochemical characteristics include at least one of hydrophobicity, net charge, average hydrophobic moment, and secondary structure tendency, and the language model includes a BERT model.
5. The method for screening highly active endogenous antimicrobial peptides according to claim 3, characterized in that: The step (3) comprises the following operations: after selecting sequences with a length ranging from 8 to 50 residues from the porcine proteome data, scoring the sequences using the classification model based on peptide physicochemical features and the classification model based on the language model BERT, integrating the scoring results of the scoring model based on peptide physicochemical features and the classification model based on the language model BERT to obtain a final comprehensive score, sorting the sequences in descending order according to the comprehensive scores, and constructing a candidate data set from a set of the top 1000 sequences for subsequent screening.
6. The method for screening highly active endogenous antimicrobial peptides according to claim 1, characterized in that: In the step (6), during the bioinformatics analysis, a secretory protein prediction tool or an exosome protein prediction tool is used to predict the candidate antimicrobial peptide, wherein the secretory protein prediction tool includes at least one of SignalP, TargetP, and SecretomeP, and the exosome protein prediction tool includes at least one of Exocarta, Vesiclepedia, and FunRich.
7. The method for screening highly active endogenous antimicrobial peptides according to claim 1, characterized in that: In the step (6), after screening out an antimicrobial peptide having a sequence similarity with the known antimicrobial peptide of less than 100% and having secretory protein characteristics or exosomal protein characteristics, the method further includes a step of verifying the antibacterial activity of the antimicrobial peptide.
8. The method for screening highly active endogenous antimicrobial peptides according to claim 7, characterized in that: The antibacterial activity of the antimicrobial peptide is verified according to the following steps: different concentrations of the antimicrobial peptide are mixed with at least one of Escherichia coli, Staphylococcus aureus, and Thermomyces cerevisiae, and cultured at 25-37° C. for 12-36 hours, and the antibacterial activity of the antimicrobial peptide is evaluated by antibacterial experimental indicators, and the antibacterial activity indicators include at least one of the minimum inhibitory concentration, the minimum bactericidal concentration, the growth curve, the total number of colonies, and the diameter of the inhibition zone.
9. A highly active endogenous antimicrobial peptide, characterized in that: The highly active endogenous antimicrobial peptide is obtained by screening using the method for screening highly active endogenous antimicrobial peptides as described in any one of claims 1 to 8.
10. The highly active endogenous antimicrobial peptide according to claim 9, characterized in that: The highly active endogenous antimicrobial peptides include at least one of A0A4X1T497_PIG and A0A4X1THB8_PIG.
Citation Information
Cited By
DPP-IV inhibitory peptide screening method based on source perception activity sorting
CN122369698A
Dpp-iv inhibiting peptide screening method based on source perception activity ranking
CN122369698B