Antibacterial peptide screening framework based on protein language model and biological information calculation software
Through the antibacterial peptide screening framework based on protein language model and biological information calculation software, the existing antibacterial peptide screening methods are solved, and efficient and accurate antibacterial peptide screening and novel improvement are achieved.
Patent Information
- Application Number
- CN202510211663.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
The existing antimicrobial peptide screening methods are inefficient and insufficiently accurate, and face the problems of data quality and deviation, difficulty in specific prediction, high false positive rate, insufficient novelty and difficulty in transforming the effect from in vitro to in vivo.
An antibacterial peptide screening framework based on protein language model and bioinformatics calculation software is adopted, and through machine learning screening and posterior verification, the protein language model screening results and the biological activity and properties of antibacterial peptides are combined to improve screening efficiency and accuracy.
The efficient and accurate screening of antimicrobial peptide candidate sequences is achieved, the false positive rate is reduced, novelty is improved and the success rate of transformation from in vitro to in vivo effects is provided, and an efficient and accurate antimicrobial peptide screening method is provided.
Smart Images

Figure BDA0005286169530000031 
Figure BDA0005286169530000041 
Figure BDA0005286169530000081
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of bioinformatics and computational biology. Specifically, the present invention belongs to the technical field of antimicrobial peptide screening, and involves using protein language models and bioinformatics calculation methods for antimicrobial peptide screening, for the efficient screening and optimization of antimicrobial peptides. Background Art
[0002] As a novel antibacterial agent, antimicrobial peptide (AMP) has attracted much attention in dealing with the increasingly serious global antimicrobial resistance (AMR) crisis. AMP is a class of short-chain amino acid sequences (usually defined as 12 - 50 amino acid residues), which can inhibit microbial growth by interfering with mechanisms such as the integrity of the microbial cell wall. Compared with traditional antibiotics, AMP has advantages such as broad-spectrum antibacterial activity and lower development of drug resistance.
[0003] Traditional methods for AMP screening mainly include: isolation from natural sources, screening after rational design, and high-throughput screening. These methods have identified a large number of AMPs, but face significant challenges in terms of efficiency and cost.
[0004] With the development of computing hardware devices, computational-based AMP screening methods can overcome the above limitations to a certain extent, and thus computational-based AMP screening methods have gradually emerged. Currently, computational-based AMP screening methods mainly include the following categories: sequence feature analysis, machine learning methods, deep learning methods, integration of multi-omics data, and virtual screening.
[0005] Although progress has been made in AMP screening using computational methods, the following main challenges still remain: (1) Data quality and bias: Existing AMP databases may be biased and cannot fully represent the diversity of AMPs in nature; (2) Difficulty in specific prediction: Accurately predicting the specific antibacterial mechanism or activity spectrum of AMPs is still challenging; (3) High false positive rate: Many computational methods may produce a large number of false positive results during the screening process, increasing the cost of subsequent experimental verification; (4) Lack of novelty: Existing methods may tend to identify sequences similar to known AMPs, while ignoring potential novel AMPs; (5) Translation from in vitro to in vivo effects: Computationally screened AMPs require comprehensive pharmacokinetic and pharmacodynamic studies to verify their potential as viable drug candidates.
[0006] Therefore, there is an urgent need in the art to develop an efficient and highly accurate antimicrobial peptide screening method. Summary of the Invention
[0007] The objective of the present invention is to provide an antibacterial peptide screening framework based on a protein language model and bioinformatics calculation software, so as to efficiently and accurately screen out target sequences of antibacterial peptides with good antibacterial activity from antibacterial peptide candidate sequences.
[0008] In the first aspect of the present invention, a method for constructing an antibacterial peptide screening model is provided, and the method includes the steps:
[0009] (1) Machine learning screening;
[0010] (2) Posterior verification;
[0011] Among them, the antibacterial peptide candidate sequences are subjected to the machine learning screening and the posterior verification to obtain target sequences of antibacterial peptides.
[0012] In another preferred example, the framework further includes the step: (3) High-throughput screening.
[0013] In another preferred example, in step (1), it includes:
[0014] (A1) Using a machine learning model, evaluating the sequence similarity between an antibacterial peptide candidate sequence and a target functional sample through a scoring system;
[0015] (A2) Predicting antibacterial activity indicators.
[0016] In another preferred example, in step (1), it further includes:
[0017] (A3) Evaluating the stability of antibacterial peptides;
[0018] (A4) Evaluating structural similarity.
[0019] In another preferred example, in step (A1), it includes:
[0020] (a1.1) Using the machine learning model to convert the antibacterial peptide candidate sequence and / or the target functional sample into a numerical form;
[0021] (a1.2) Calculating the similarity between the antibacterial peptide candidate sequence and the target functional sample.
[0022] In another preferred example, the machine learning model is selected from the group consisting of: a protein language model, a deep learning model, or a combination thereof.
[0023] In another preferred example, the protein language model includes: the ESM2 model, ProtBERT, ProtXLNet.
[0024] In another preferred example, the protein language model is the ESM2 model.
[0025] In another preferred example, the ESM2 model evaluates the sequence similarity between the bacteriopeptide candidate sequence and the target functional sample through zero-shot learning.
[0026] In another preferred example, the protein language model converts the antimicrobial peptide candidate sequence and / or the target functional sample into the numerical form by a method selected from the group consisting of: natural language processing algorithms, or sequence alignment algorithms.
[0027] In another preferred example, the natural language processing algorithm is a word vector.
[0028] In another preferred example, the word vector includes: one-hot encoding, word2vec.
[0029] In another preferred example, the word vector is one-hot encoding.
[0030] In another preferred example, the sequence alignment algorithm is a protein substitution scoring matrix.
[0031] In another preferred example, the protein substitution matrix is selected from the group consisting of: PAM matrix, or BLOSUM matrix.
[0032] In another preferred example, the protein substitution matrix is the BLOSUM matrix.
[0033] In another preferred example, the numerical form includes: embedding vector, scoring matrix.
[0034] In another preferred example, the numerical form is an embedding vector.
[0035] In another preferred example, the embedding vector is an ESM2 embedding.
[0036] In another preferred example, the ESM2 embedding is an embedding vector formed by converting the candidate antimicrobial peptide sequence and the target functional sample using the ESM2 model.
[0037] In another preferred example, the deep learning model includes: a model based on a convolutional neural network (CNN), or a model based on a recurrent neural network (RNN).
[0038] In another preferred example, the deep learning model is a model based on a recurrent neural network (RNN).
[0039] In another preferred example, the scoring system employs an algorithm selected from the group consisting of: supervised learning algorithms, reinforcement learning algorithms, unsupervised learning algorithms, or semi-supervised learning algorithms.
[0040] In another preferred example, the scoring system employs an algorithm selected from the group consisting of: supervised learning algorithms, or reinforcement learning algorithms.
[0041] In another preferred example, the supervised learning algorithm is a classification algorithm.
[0042] In another preferred example, the classification algorithm is selected from the group consisting of: linear model, K-nearest neighbor (KNN), random forest, support vector machine (SVM), decision tree, neural network, naive Bayes, Boosting, or a combination thereof.
[0043] In another preferred example, the classification algorithm is selected from the group consisting of: linear model, K-nearest neighbor, random forest, or a combination thereof.
[0044] In another preferred example, the classification algorithm is K-nearest neighbor.
[0045] In another preferred example, the calculation for evaluating sequence similarity by the K-nearest neighbor is:
[0046]
[0047] where s_f is the similarity score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric, x is the generated sample, y_i is the center of the i-th functional cluster, k is the number of clusters, and d is the number of distance metrics.
[0048] In another preferred example, the distance metric is selected from the group consisting of: Euclidean distance, Manhattan distance, cosine similarity, Chebyshev distance, Minkowski distance, standard Euclidean distance, Mahalanobis distance, Hamming distance, Jaccard distance, correlation distance, information entropy, or a combination thereof.
[0049] In another preferred example, the distance metrics are Euclidean distance, Manhattan distance, and cosine similarity.
[0050] In another preferred example, the d = 3.
[0051] In another preferred example, the scoring system includes using multiple of the supervised learning algorithms to obtain multiple of the similarity scores.
[0052] In another preferred example, the multiple similarity scores are synthesized using a method selected from the group consisting of: fuzzy logic, Dempster-Shafer theory, or multi-objective optimization algorithm.
[0053] In another preferred example, the multi-objective optimization algorithm is NSGA-II.
[0054] In another preferred example, the unsupervised learning algorithm is a clustering algorithm.
[0055] In another preferred example, the clustering algorithm is selected from the following group: density clustering, K-means clustering, spectral clustering, hierarchical clustering, grid clustering, or model clustering.
[0056] In another preferred example, the clustering algorithm is selected from the following group: density clustering, or K-means clustering.
[0057] In another preferred example, the density clustering is DBSCAN.
[0058] In another preferred example, a computational model is used to evaluate the stability of the antimicrobial peptide.
[0059] In another preferred example, the computational model includes a computational model for simulating protein degradation.
[0060] In another preferred example, a graph neural network is used to evaluate the structural similarity.
[0061] In another preferred example, the antimicrobial activity indicators include: minimum inhibitory concentration, minimum bactericidal concentration, time-kill curve.
[0062] In another preferred example, the antimicrobial activity indicator is the minimum inhibitory concentration.
[0063] In another preferred example, a method selected from the following group is used to predict the minimum inhibitory concentration: supervised learning algorithm, expert system, or deep learning model.
[0064] In another preferred example, the supervised learning algorithm is selected from the following group: classification algorithm, or regression algorithm.
[0065] In another preferred example, the classification algorithm is Naive Bayes.
[0066] In another preferred example, a classifier based on the Naive Bayes is trained on a known data set such that the classifier reaches the training objective, and a pre-trained classifier is obtained, thereby using the pre-trained classifier to predict the minimum inhibitory concentration.
[0067] In another preferred example, the known data set is an antimicrobial peptide data set with experimentally determined minimum inhibitory concentration values.
[0068] In another preferred example, the classifier based on Naive Bayes is a Bayesian classifier.
[0069] In another preferred example, the Bayesian classifier uses the ESM2 embedding.
[0070] In another preferred example, the training objective is to minimize the L1 loss between the predicted minimum inhibitory concentration value and the actual minimum inhibitory concentration value:
[0071]
[0072] Among them, y_i is the actual minimum inhibitory concentration value, f(x_i) is the predicted minimum inhibitory concentration value of the numerical form x_i of the i-th peptide segment, and N is the number of training samples.
[0073] In another preferred example, the regression algorithm is selected from the following group: linear regression, K-nearest neighbor regression, support vector machine regression, decision tree regression, neural network regression, naive Bayes regression, Boosting regression, random forest regression, deep forest regression, or extremely randomized tree regression.
[0074] In another preferred example, the Boosting regression is gradient boosting decision tree regression.
[0075] In another preferred example, the expert system includes a rule-based expert system.
[0076] In another preferred example, the method of deep learning is to use an end-to-end deep learning model.
[0077] In another preferred example, the end-to-end deep learning model directly identifies the minimum inhibitory concentration from the antimicrobial peptide candidate sequence.
[0078] In another preferred example, in step (2), it includes steps selected from the following group: restricting the peptide segment length, predicting structural features, analyzing similarity, or a combination thereof.
[0079] In another preferred example, in step (2), it includes steps: restricting the peptide segment length, predicting structural features, and analyzing similarity.
[0080] In another preferred example, in step (2), it further includes screening the antimicrobial peptide candidate sequence by using a method selected from the following group: integrating multi-omics data, or knowledge graph technology.
[0081] In another preferred example, the multi-omics data is selected from the following group: proteomics, metabolomics, or a combination thereof.
[0082] In another preferred example, the multi-omics data is proteomics.
[0083] In another preferred example, the peptide segment length is ≤ 50 amino acids, preferably ≤ 30 amino acids, more preferably ≤ 25 amino acids.
[0084] In another preferred example, the structural features are foldability and thermal stability.
[0085] In another preferred example, the structural features are predicted by a method selected from the following group: protein structure prediction tools, or molecular simulation.
[0086] In another preferred example, the protein structure prediction tool is selected from the group consisting of: AlphaFold2, ESMFold, I-TASSER, RoseTTAFold, GalaxyTBM, SWISS-MODEL, or a combination thereof.
[0087] In another preferred example, the protein structure prediction tool is a combination of AlphaFold2 and ESMFold.
[0088] In another preferred example, the molecular simulation is molecular dynamics simulation.
[0089] In another preferred example, the similarity includes: sequence similarity, structural similarity.
[0090] In another preferred example, the similarity is analyzed by a method selected from the group consisting of: sequence alignment algorithms, or structural alignment algorithms.
[0091] In another preferred example, the sequence similarity is analyzed by the sequence alignment algorithm.
[0092] In another preferred example, the structural similarity is analyzed by the structural alignment algorithm.
[0093] In another preferred example, the sequence alignment algorithm is selected from the group consisting of: BLAST, Smith-Waterman algorithm, Needleman-Wunsch algorithm, or a combination thereof.
[0094] In another preferred example, the sequence alignment algorithm is BLAST.
[0095] In another preferred example, the structural alignment algorithm is selected from the group consisting of: FoldSeek, TM-align, DALI, SSAP, FLEXPROT, or a combination thereof.
[0096] In another preferred example, the structural alignment algorithm is a combination of FoldSeek and TM-align.
[0097] In another preferred example, the method further includes: using an evaluation model to comprehensively consider multiple indicators to screen candidate sequences of antimicrobial peptides.
[0098] In another preferred example, the evaluation model is selected from the group consisting of: AHP, TOPSIS, grey relational analysis, fuzzy comprehensive evaluation, entropy weight method, or a combination thereof.
[0099] In another preferred example, the evaluation model is AHP.
[0100] In another preferred example, the indicators are selected from the group consisting of: sequence similarity, structural similarity, structural features, minimum inhibitory concentration, or a combination thereof.
[0101] In a second aspect of the present invention, there is provided a method for generating an antimicrobial peptide candidate sequence, the method comprising generating the antimicrobial peptide candidate sequence using machine learning techniques.
[0102] In another preferred embodiment, the machine learning techniques are selected from the group consisting of: language models, deep learning models, machine learning algorithms, or combinations thereof.
[0103] In another preferred embodiment, the language model employs techniques selected from the group consisting of: rule-based methods, Transformer-based methods, large-scale pre-trained model-based methods, statistics-based methods, neural network-based methods, or combinations thereof.
[0104] In another preferred embodiment, the language model employs techniques selected from the group consisting of: rule-based methods, Transformer-based methods, or large-scale pre-trained model-based methods.
[0105] In another preferred embodiment, the rule-based methods are selected from the group consisting of: sequence-guided peptide design methods, structure-guided peptide design methods, or combinations thereof.
[0106] In another preferred embodiment, the language model is selected from the group consisting of: AMPGen model, AMPGen-MIC model, or combinations thereof.
[0107] In another preferred embodiment, the language model is the AMPGen model.
[0108] In another preferred embodiment, the deep learning model is a deep generation model.
[0109] In another preferred embodiment, the deep generation model includes: variational autoencoder (VAE), generative adversarial network (GAN).
[0110] In another preferred embodiment, the machine learning algorithm is an evolutionary algorithm.
[0111] In another preferred embodiment, the evolutionary algorithm is selected from the group consisting of: genetic algorithm, differential evolution algorithm, co-evolution algorithm, or estimation of distribution algorithm.
[0112] In another preferred embodiment, the evolutionary algorithm is a genetic algorithm.
[0113] In a third aspect of the present invention, there is provided an antimicrobial peptide screening framework or system, the framework or system comprising:
[0114] An input unit configured to input data, the data including an input sequence to be screened, the input sequence to be screened including an antimicrobial peptide candidate sequence generated by the method according to the second aspect of the present invention;
[0115] A screening unit for antimicrobial peptides, configured to execute a screening model for antimicrobial peptides to obtain a target sequence of antimicrobial peptides from the candidate sequences of antimicrobial peptides; wherein, the screening model is constructed by the method described in the first aspect of the present invention.
[0116] An output unit configured to output the target sequence of antimicrobial peptides.
[0117] In another preferred example, the screening model is AMPGen-Filtering.
[0118] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be elaborated one by one here. Description of the Drawings
[0119] Figure 1 The flow of the screening framework for antimicrobial peptides based on protein language models and bioinformatics methods (taking AMPGen-MIC as an example). Detailed Embodiments
[0120] Through extensive and in-depth research, the inventors of the present invention have first constructed a screening framework for antimicrobial peptides based on protein language models and bioinformatics methods. Specifically, the present invention combines the screening results of protein language models with the actual biological activities and biological properties of antimicrobial peptides through innovative computational methods, effectively solving the problems of low screening efficiency and insufficient accuracy in the prior art for antimicrobial peptides, improving the prediction mode of the functional characteristics of antimicrobial peptides, and increasing the diversity and novelty of the screening results, providing a new method for screening antimicrobial peptides that is efficient, accurate, and flexible, and accelerating the screening and development process of new antimicrobial drugs. On this basis, the present invention has been completed.
[0121] The following explains some innovative points of the present invention:
[0122] The present invention uses the screening framework to screen candidate samples of antimicrobial peptides. The screening framework is a screening model of AMPGen-Filtering, including two-stage screening. In the first stage, first, the pre-trained protein language model ESM2 is used to evaluate the similarity between the sequence and known antimicrobial peptides in a zero-shot manner, and a classifier based on MIC is used to predict the antimicrobial activity of the sequence. Secondly, the generated antimicrobial peptide samples are screened by using the K-nearest neighbor algorithm, and the screening score is calculated. The calculation formula of the screening score is:
[0123]
[0124] Among them, s_f is the final similarity screening score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric (Euclidean distance, Manhattan distance, and cosine similarity), x is the generated sample, y_i is the center of the i-th functional cluster, and k is the number of clusters.
[0125] After that, a Bayesian classifier based on ESM2 embeddings is used to further refine the selection process. This classifier is trained on a known AMP dataset with MIC values, and the goal is to minimize the L1 loss between the predicted MIC value and the actual MIC value. The calculation formula for the loss value L1 is as follows:
[0126]
[0127] Among them, y_i is the true MIC value, f(x_i) is the predicted MIC value of the ESM2 embedding x_i of the i-th peptide segment, and N is the number of training samples.
[0128] Finally, the sequences obtained by cross-validating the K-nearest neighbor algorithm and MIC prediction screening are used to further refine the selection of candidate peptides.
[0129] In the second stage, the candidate antimicrobial peptide sequences are optimized through length screening, structure prediction, and similarity analysis. In length screening, candidate sequences with a length of at most 50 amino acids are selected. Protein structure prediction tools such as AlphaFold2 and ESMFold are used in structure prediction to evaluate the structural characteristics of candidate sequences. In similarity analysis, BLAST and FoldSeek are used for sequence similarity search.
[0130] It should be understood that the specific methods and experimental conditions of the present invention are described in various levels of detail below to provide an understanding of the essence of the present invention. Definitions of some terms used in this specification are provided below. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention belongs.
[0131] The term
[0132] As used herein, the terms "comprising", "including", and "containing" can be used interchangeably, including not only closed definitions but also semi-closed and open definitions. In other words, the said terms include "consisting of" and "consisting essentially of".
[0133] As used herein, the term "Foldseek" is a method for searching protein structure data, used to find structurally similar proteins in large-scale structure databases. This method reduces the computational time by 4 to 5 orders of magnitude by organizing the description of the tertiary amino acid interactions within a protein into a sequence of structure alphabets and aligning the queried protein structure with the database.
[0134] As used herein, the terms "Sequence Embedding", "Embedding Vector", and "Embedding" are used interchangeably and refer to the conversion of a protein sequence into a vector representation of a fixed dimension, capturing the semantic information of the sequence. In the present invention, generating a sample embedding refers to an embedding vector transformed from a candidate antimicrobial peptide sequence; a target functional sample embedding refers to an embedding vector transformed from an antimicrobial peptide sequence with known function and sequence.
[0135] As used herein, the terms "physicochemical property" and "physicochemical characteristic" are used interchangeably and are quantitative indicators describing molecular characteristics. In the present invention, the physicochemical properties of antimicrobial peptides include hydrophobicity, hydrophobic moment, charge, isoelectric point, etc.
[0136] As used herein, the terms "Zero-shot Learning" and "Zero-shot Approach" are used interchangeably and refer to the ability of a model to recognize or classify categories not seen during the training process.
[0137] As used herein, the terms "TM-Score (Template Modeling Score)" and "TM score" are used interchangeably and are a score used to evaluate the structural similarity of proteins, ranging from 0 to 1, and the closer to 1, the more similar the structures.
[0138] As used herein, the term "RMSD (Root Mean Square Deviation)" refers to the root mean square deviation, which is used to measure the average distance of the atomic spatial positions between two superimposed protein structures, with the unit of angstrom.
[0139] As used herein, the term "Structure-Activity Relationship (SAR)" is a study describing the relationship between the structural characteristics of a molecule and its biological activity.
[0140] As used herein, the term "Molecular Dynamics Simulation" is a computer simulation technique used to study the physical motion of atomic and molecular systems over time.
[0141] As used herein, the term "Multi-objective Optimization Algorithm" refers to an algorithm that simultaneously optimizes multiple objective functions, such as NSGA-II (Non-dominated Sorting Genetic Algorithm II).
[0142] As used herein, the terms "antimicrobial peptide target sequence" and "antimicrobial peptide target sample" are used interchangeably and refer to the antimicrobial peptide sequence obtained after screening the antimicrobial peptide sequence to be screened through the screening framework or system of the present invention, which has properties such as good antibacterial activity.
[0143] AMP Screening Method
[0144] The AMP screening method includes traditional AMP screening methods and computational-based AMP screening methods.
[0145] The traditional AMP screening methods mainly include the following three: (1) Isolation from natural sources: Extracting and isolating potential AMPs from various organisms; (2) Screening after rational design: Designing peptide sequences based on known structure-activity relationships and then screening; (3) High-throughput screening: Using synthetic peptide libraries for large-scale screening.
[0146] The computational-based AMP screening methods mainly include the following five: (1) Sequence feature analysis: Predicting and screening AMPs based on features such as amino acid composition and physicochemical properties; (2) Machine learning methods: Using algorithms such as support vector machines (SVMs) and random forests (RFs) to build classification models to screen potential AMPs from a large number of candidate peptides; (3) Deep learning methods: Using deep learning architectures such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to improve the accuracy of AMP screening; (4) Integrating multi-omics data: Combining multi-dimensional data such as protein structure information and evolutionary conservation to improve the accuracy and interpretability of screening; (5) Virtual screening: Using methods such as molecular docking and molecular dynamics simulations to predict the interactions between AMPs and targets for large-scale virtual screening.
[0147] Language Models and Protein Language Models
[0148] As used herein, the term "language model" is a type of machine learning model belonging to the fields of natural language processing and deep learning, which includes large language models (LLMs), etc. The goal of a language model is to predict the next possible characters based on a given context. The technical bases used in language models include: rule-based methods, statistics-based methods, neural network-based methods, Transformer-based methods, and large-scale pre-trained model-based methods.
[0149] In the present invention, a language model is used to generate candidate sequences of antimicrobial peptides. Among them, "AMPGen" and "AMPGen-MIC" are two language models, and AMPGen-MIC is a model further fine-tuned based on AMPGen. AMPGen and AMPGen-MIC can generate new candidate sequences of antimicrobial peptides.
[0150] As used herein, the term "Protein Language Models (PLMs)" is a deep learning-based model specifically used to process or generate protein sequences. These models learn the potential patterns and regularities of protein sequences through pre-training on a large amount of protein sequence data, so as to predict the properties, structures and other attributes of unknown proteins.
[0151] As used herein, the terms "ESM2", "ProtBERT", "ProtXLNet" and "AlphaFold2" are several language models commonly used in protein research.
[0152] ESM2 (Evolutionary Scale Modeling 2) is a large protein language model developed by Meta AI and pre-trained on hundreds of millions of protein sequences. ESM2 extracts information by encoding these sequences for subsequent prediction. The term "ESMFold" is a model for protein structure prediction using the ESM2 model. It can perform end-to-end three-dimensional structure prediction using only a single sequence as input by utilizing the information and representations learned by ESM2 to quantify the occurrence of protein structures.
[0153] ProtBERT and ProtXLNet are natural language processing models trained on protein sequences. Among them, ProtBERT adopts a self-supervised learning method, combines protein structures with annotations of Gene Ontology, and can make predictions in terms of protein structures, post-translational modifications and biophysical properties; ProtXLNet, on the other hand, adopts an autoregressive model to make predictions on proteins.
[0154] AlphaFold2 is a protein structure prediction algorithm developed by DeepMind, which can predict the three-dimensional structure of proteins with extremely high accuracy.
[0155] Loss function
[0156] As used herein, the terms "loss function" and "error function" are used interchangeably and refer to a class of functions that calculate the difference between predicted values and true values. By calculating the difference between predicted values and true values, the performance of the model is evaluated. In machine learning, by introducing a loss function during the prediction process, the predicted values can be controlled to be infinitely close to the true values, so as to improve the accuracy of model prediction as much as possible. Therefore, the smaller the value output by the loss function, the better the machine learning effect and the higher the accuracy of the model.
[0157] Loss functions are classified into regression losses and classification losses according to the type of task. Among them, regression losses mainly deal with continuous variables, such as Mean Square Error (MSE) and Mean Absolute Error (MAE); while classification losses mainly deal with discrete variables, such as Cross Entropy Loss and Dice Loss.
[0158] The loss function is located between the forward propagation and backward propagation of the machine learning model. The forward propagation refers to the model generating predicted values based on input features; the loss function receives the above predicted values and calculates the difference from the true values; this difference is then used in the backward propagation stage to update the model's parameters and reduce the error of subsequent predictions.
[0159] In the present invention, a loss function based on Naive Bayes is constructed in the screening framework, and the loss function is:
[0160]
[0161] where y_i is the true MIC value, f(x_i) is the predicted MIC value of the ESM2 embedding x_i of the i-th peptide segment, and N is the number of training samples.
[0162] The scoring system of the present invention
[0163] In the present invention, the scoring system is selected from the group consisting of: supervised learning algorithms, reinforcement learning algorithms, unsupervised learning algorithms, or semi-supervised learning algorithms.
[0164] As used herein, the term "supervised learning algorithm" refers to a class of algorithms used in the process of supervised learning. The supervised learning refers to providing input data and its corresponding label data to the model, and after training, the model accurately finds the optimal mapping relationship between the input data and the label data, so as to predict or classify new unlabeled data. Supervised learning algorithms mainly include classification algorithms and regression algorithms. Classification algorithms are mainly used to output discrete data, while regression algorithms are used to output continuous data.
[0165] As used herein, the term "unsupervised learning algorithm" refers to a class of algorithms used in the process of unsupervised learning, which refers to the classification of unlabeled input data. Unsupervised learning algorithms include clustering algorithms.
[0166] In the present invention, a scoring system is used to evaluate sequence similarity. Preferably, the scoring system employs a supervised learning algorithm or an unsupervised learning algorithm. Preferably, the supervised learning algorithm is a classification algorithm selected from the group consisting of: linear model, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), decision tree, neural network, Naive Bayes, Boosting, Random Forest, or a combination thereof. The unsupervised learning algorithm is selected from the group consisting of: spectral clustering, hierarchical clustering, density clustering, K-means clustering, grid clustering, or model clustering. Among them, SVM classifies data points by finding the optimal hyperplane, while Random Forest classifies by constructing multiple decision trees and taking the majority vote.
[0167] Preferably, the scoring system employs the K-Nearest Neighbors algorithm. The K-Nearest Neighbors algorithm makes predictions based on the K nearest neighbors around a sample. In this process, a distance metric is needed to calculate the distance between the sample and its neighbors. The terms "distance metric" and "similarity metric" can be used interchangeably because the commonly used method for evaluating the similarity between samples is to calculate the distance.
[0168] Preferably, the calculation for evaluating sequence similarity by the K-Nearest Neighbors is:
[0169]
[0170] where s_f is the similarity score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric, x is the generated sample, y_i is the center of the i-th functional cluster, and k is the number of clusters.
[0171] Preferably, the distance metrics are Euclidean distance, Manhattan distance, and cosine similarity.
[0172] In the present invention, preferably, a supervised learning algorithm is used to predict the minimum inhibitory concentration. The supervised learning algorithm is a regression algorithm selected from the group consisting of: Naive Bayes, linear regression, K-Nearest Neighbors regression, Support Vector Machine regression, decision tree regression, neural network regression, Naive Bayes regression, gradient boosting decision tree regression, Random Forest regression, deep forest regression, or extremely randomized tree regression.
[0173] Preferably, a Bayesian Classifier is constructed using the Naive Bayes algorithm to predict the minimum inhibitory concentration. The Bayesian Classifier refers to a probability classifier based on Bayes' theorem, which can perform classification prediction according to the conditional probability of features.
[0174] Method for analyzing similarity in the posterior verification of the present invention
[0175] In the present invention, similarity is analyzed by a method selected from the group consisting of: sequence alignment algorithms, or structure alignment algorithms.
[0176] Preferably, the sequence alignment algorithm is selected from the group consisting of: BLAST, Smith-Waterman algorithm, Needleman-Wunsch algorithm, or a combination thereof.
[0177] Preferably, the sequence alignment algorithm is BLAST. As used herein, the term "BLAST" is the Basic Local Alignment Search Tool, which is a sequence similarity search algorithm used to compare biological sequences (such as DNA, RNA, or protein sequences) with a sequence database to identify similar sequences in the database.
[0178] Preferably, the structure alignment algorithm is selected from the group consisting of: FoldSeek, TM-align, DALI, SSAP, FLEXPROT, or a combination thereof.
[0179] Preferably, the structure alignment algorithm is FoldSeek. As used herein, "FoldSeek" is a bioinformatics tool for rapidly and accurately searching for protein structures, which performs alignment based on 3D structure similarity.
[0180] Expert system
[0181] As used herein, the term "expert system" belongs to the field of artificial intelligence and refers to a computer software system that can solve complex problems like a human expert in a specific field. It can effectively utilize the experience and professional knowledge accumulated by experts over the years, and solve problems that require experts by simulating the thinking process of experts. An expert system needs to save expert knowledge in a knowledge base through a certain knowledge acquisition method, and then work with an inference engine in combination with a man-machine interaction interface.
[0182] Expert systems can be classified according to inference rules into: rule-based expert systems, case-based expert systems, artificial neural network-based expert systems, frame-based expert systems, fuzzy logic-based expert systems, genetic algorithm-based expert systems, etc. Among them, a rule-based expert system consists of five parts: a knowledge base, a database, an inference engine, an explanation facility, and a user interface. The knowledge base contains domain knowledge related to problem-solving, and the knowledge is represented by a set of rules with a condition-action structure; the database contains a set of facts for matching the conditions in the knowledge base; the inference engine is associated with the rules in the knowledge base and the facts in the database to perform inference so that the expert system can find a solution; the user interface enables communication between the user and the system.
[0183] In the present invention, an expert system is used to predict the minimum inhibitory concentration. Preferably, a rule-based expert system is used to predict the minimum inhibitory concentration.
[0184] Deep learning model
[0185] As used herein, the term "deep learning model" belongs to machine learning models and uses a multi-layer neural network to learn from a large amount of data.
[0186] Common deep learning models include supervised neural networks such as Recurrent Neural Networks (RNN), Convolutional Neural Networks (CNN), deep neural networks, recurrent neural networks, etc., and unsupervised or semi-supervised deep learning models such as deep generative models, autoencoders, etc. Among them, deep generative models include Generative Adversarial Network (GAN), and autoencoders include Variational Autoencoder (VAE).
[0187] Deep learning models such as RNN and CNN have been widely used in natural language processing and biological sequence analysis. Among them, RNN is a neural network for processing sequence data and can utilize the temporal or spatial dependence of the sequence; CNN is suitable for processing data with a grid topology structure, such as images or sequence data.
[0188] Variational autoencoders are often used to learn the latent representation of data to generate new samples. A generative adversarial network consists of two networks, a generative model and a discriminative model, and generates realistic samples through adversarial learning.
[0189] The term "end-to-end deep learning model" refers to a type of deep learning model that utilizes an end-to-end approach. Here, "end-to-end" is a data transmission method, which means that data is directly transmitted from the sender to the receiver without the need for an intermediate environment to parse and process the data content, ensuring the directness and integrity of the data. The end-to-end deep learning model can be an end-to-end RNN, an end-to-end CNN, or other deep learning models.
[0190] In the present invention, RNN and CNN can be used to replace the ESM2 model for machine learning screening of antimicrobial peptides; GAN or VAE can replace the current AMPGen and AMPGen-MIC models to generate candidate sequences of antimicrobial peptides.
[0191] Antibacterial activity index
[0192] The present invention uses antibacterial activity indexes to characterize the antibacterial activity of antimicrobial peptides. By predicting the antibacterial activity of candidate sequences of antimicrobial peptides, target sequences of antimicrobial peptides with good antibacterial activity can be screened out.
[0193] As used herein, the term "Minimum Inhibitory Concentration (MIC)" refers to the lowest concentration of an antibacterial substance that can inhibit the visible growth of microorganisms and is an important index for evaluating antibacterial efficacy.
[0194] As used herein, the term "Minimum Bactericidal Concentration (MBC)" refers to the lowest concentration required to kill microorganisms under specific conditions and a fixed extended time (18 to 24 hours), at which the viability of the microorganisms is reduced by 99%.
[0195] As used herein, the term "Time-kill Curve" refers to a curve that describes the change in the killing effect of an antibacterial agent on microorganisms over time and is used to evaluate the bactericidal kinetics of the antibacterial agent.
[0196] Antimicrobial peptide screening framework or system of the present invention and its construction method
[0197] The present invention provides an antimicrobial peptide screening framework or system and its construction method for screening the generated candidate antimicrobial peptide sequences to obtain antimicrobial peptides with better antibacterial activity.
[0198] Specifically, the method includes the steps of:
[0199] (1) Machine learning screening, and the machine learning screening includes the steps of: (A1) Using a machine learning model, evaluating the sequence similarity between candidate antimicrobial peptide sequences and target function samples through a scoring system; (A2) Predicting antibacterial activity indexes.
[0200] (2) Posterior verification, which includes: restricting peptide length, predicting structural features, analyzing similarity, or a combination thereof. Preferably, the posterior verification includes restricting peptide length, predicting structural features, and analyzing similarity.
[0201] The framework or system includes: an input unit configured to input data, the input data including a protein model and / or an input sequence to be screened, the input sequence to be screened including an antimicrobial peptide candidate sequence generated by the method described in the second aspect of the present invention; a screening unit for antimicrobial peptides configured to execute a screening model for antimicrobial peptides to obtain an antimicrobial peptide target sequence from the antimicrobial peptide candidate sequence; wherein the screening model is constructed using the method described in the first aspect of the present invention; and an output unit configured to output a screening result. This framework or system is the "AMPGen - Filtering".
[0202] In the method for constructing the antimicrobial peptide screening framework or system of the present invention, the screening indicators can be adjusted according to the type and properties of bioactive peptides (antimicrobial peptides are a type of bioactive peptide), so as to apply this method to construct screening frameworks or systems for more types of bioactive peptides, such as anticancer peptides, antiviral peptides, and cell - penetrating peptides. For example, for anticancer peptides or antiviral peptides, screening can be carried out according to the killing power against tumor cells or viruses.
[0203] The main advantages of the present invention include:
[0204] (1) The antimicrobial peptide screening framework proposed by the present invention integrates advanced protein language model technology with traditional bioinformatics calculation methods. Compared with the prior art, this integration not only improves the screening accuracy but also enhances the flexibility of the framework. Through modular design, the present invention can adapt to the development of new technologies and update or replace specific components at any time without affecting the effectiveness of the overall framework.
[0205] (2) According to the steps in the invention content 1, the present invention adopts a multi - dimensional evaluation method at the sequence and structure levels, significantly improving the reliability of the screening results. Compared with the prior art that only relies on sequence similarity or single - structure prediction, the present invention comprehensively evaluates the potential of candidate antimicrobial peptides through a multi - stage screening strategy, effectively reducing the false - positive rate while increasing the possibility of discovering novel antimicrobial peptides.
[0206] (3) According to step (2) in the invention content 1, the present invention ensures that the screening results have practical biological significance and experimental feasibility by introducing posterior verification stages such as length screening and folding stability prediction, and improves the success rate of the transformation from in vitro to in vivo effects.
[0207] (4) The framework design of the present invention allows for the integration of new data sources and algorithms, enabling the framework to continuously improve and adapt to the changing research needs of antimicrobial peptides, and having scalability compared to the common fixed processes in the prior art.
[0208] (5) By combining advanced machine learning techniques and traditional bioinformatics methods, the present invention significantly improves the efficiency of antimicrobial peptide screening, reduces the R & D cost, and provides a more cost-effective solution for the development of antimicrobial peptide drugs.
[0209] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The experimental methods without specific conditions noted in the following embodiments are usually carried out under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise stated, percentages and parts are by weight percentage and weight parts.
[0210] Methods and steps:
[0211] 1. Generation of antimicrobial peptide candidate sequences
[0212] The present invention uses a protein designer model to generate candidate antimicrobial peptide (cAMPs) sequences. These models are large-scale protein language models, such as AMPGen, and are obtained by fine-tuning on known antimicrobial peptide data sets, such as AMPGen-MIC. The generated candidate sequences form the initial pool for subsequent screening.
[0213] 2. Screening and verification of antimicrobial peptide candidate sequences
[0214] The present invention adopts a two-stage screening process to ensure the quality of the generated candidate peptides.
[0215] 2.1 Machine learning-based basic screening
[0216] Two parallel methods are adopted in the machine learning screening stage: K-nearest neighbor (KNN) search based on ESM2 and minimum inhibitory concentration (MIC) value identification.
[0217] 2.1.1 KNN search based on ESM2
[0218] The present invention uses the ESM2 (Evolutionary Scale Modeling 2) pre-trained protein language model to evaluate sequence similarity. The specific steps are as follows:
[0219] (A1) Use the ESM2 model to convert the antimicrobial peptide candidate sequence into an embedding vector.
[0220] (A2) Calculate the distance between the candidate sequence embedding vector and the center of the target functional sample embedding vector.
[0221] (A3) Use the weighted sum of multiple distance metrics (Euclidean distance, Manhattan distance, and cosine similarity) as the similarity score:
[0222]
[0223] where \(s_f\) is the final similarity screening score, \(w_m\) is the weight of each functional cluster, \(f_m\) represents the \(m\)-th distance metric, \(x\) is the generated sample, \(y_i\) is the center of the \(i\)-th functional cluster, and \(k\) is the number of clusters.
[0224] 2.1.2 MIC value identification
[0225] The present invention uses a Bayesian classifier based on ESM2 embedding to predict antibacterial activity. The specific steps are as follows:
[0226] (B1) Train the classifier on the AMP dataset with experimentally determined minimum inhibitory concentration (MIC) values.
[0227] (B2) The training objective is to minimize the L1 loss between the predicted MIC value and the actual MIC value:
[0228]
[0229] where \(y_i\) is the true MIC value, \(f(x_i)\) is the predicted MIC value of the ESM2 embedding \(x_i\) of the \(i\)-th peptide segment, and \(N\) is the number of training samples.
[0230] 2.1.3 Cross-validation
[0231] Take the intersection of the results of KNN search and MIC value identification to obtain peptide sequences that are both structurally similar to known AMPs and may have satisfactory antibacterial activity.
[0232] 2.2 Posterior validation
[0233] The post-processing step of the present invention optimizes antibacterial peptide candidates through length screening, structure prediction, and similarity analysis. The specific implementation steps are as follows:
[0234] (C1) Length selection and folding stability
[0235] (a) Limit the peptide segment length to 50 amino acids or less to ensure compatibility with efficient laboratory synthesis techniques.
[0236] (b) Use protein structure prediction tools AlphaFold2 and ESMFold to evaluate the structural properties of the selected peptides, providing key feature predictions regarding peptide foldability and potential thermal stability.
[0237] (C2) Similarity analysis
[0238] (a) Conduct sequence-level comparisons using BLAST, aligning against curated peptide and protein domain databases. In the BLAST analysis evaluation, the following metrics are mainly focused on: highest score, total score, query coverage, E-value, percentage identity, and accepted length. Priority is given to the E-value and percentage identity to determine sequence similarity. Candidates with an E-value lower than 1e-5 and a percentage identity higher than 30% are marked for further investigation.
[0239] (b) Perform structure-based similarity searches using FoldSeek. In the FoldSeek analysis evaluation, the following metrics are mainly focused on: check probability, sequence identity, E-value, score, query position, target position, TM-score, and RMSD. Special attention is paid to the TM-score and RMSD. Structures with a TM-score higher than 0.5 and an RMSD lower than are considered to have significant structural similarity.
[0240] Through the above comprehensive screening and analysis methods, the present invention can efficiently identify novel peptide sequences with potential antibacterial activity, while providing a powerful framework for evaluating their potential and guiding future experimental research.
[0241] 3. Detection of antibacterial peptide properties
[0242] 3.1 Determination of molar mass
[0243] Use mass spectrometry analysis techniques, such as matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF MS) or liquid chromatography-mass spectrometry (LC-MS), which can provide precise molecular weight information of antibacterial peptides. In addition, nuclear magnetic resonance (NMR) can also be used for structure elucidation and molecular weight confirmation.
[0244] 3.2 Determination of bacteriostatic rate
[0245] Evaluation was carried out by spectrophotometry (OD600 measurement) and plate counting method. First, three kinds of bacteria, Pseudomonas aeruginosa, Escherichia coli, and Staphylococcus aureus, were cultured. Bacteria were inoculated into LB (Luria - Bertani) liquid medium and cultured with shaking at 37°C and 200 rpm for 12 - 16 hours until the bacterial suspension reached the logarithmic growth phase (OD600≈0.5 - 0.6), and then diluted to OD600≈0.05 with sterile PBS or LB to obtain the bacterial suspension for experiments.
[0246] After that, for yeasts and fungi (Saccharomyces cerevisiae, Candida albicans, and other fungi), they were cultured in YPD (Yeast Extract - Peptone - Dextrose) liquid medium for 16 - 24 hours with shaking at 37°C and 200 rpm, and then diluted to OD600≈0.05 with sterile PBS or YPD.
[0247] In the experiment of measuring the antibacterial rate by spectrophotometry, the experimental group, positive control group, and negative control group were set up in a 96 - well plate. In the experimental group, 100 μL of the diluted bacterial suspension and 100 μL of antibacterial peptide solutions with different concentrations were added to make the final volume 200 μL; in the positive control group, 100 μL of the bacterial suspension and 100 μL of sterile PBS or medium were added; the negative control group contained only sterile PBS and medium. Subsequently, it was incubated at 37°C for 16 - 20 hours (for fungi, it could be appropriately extended to 24 - 48 hours), and the absorbance at 600 nm (OD600) was measured using a spectrophotometer, and the antibacterial rate was calculated. The calculation formula is:
[0248] Antibacterial rate (%)=(1 - OD600 of experimental group / OD600 of control group)×100%
[0249] If the antibacterial peptide may affect the OD600 reading (such as causing bacterial aggregation or sedimentation), the plate counting method can be combined for verification. In the experiment of measuring the antibacterial rate by plate counting method, 100 μL of the bacterial suspension after the action of antibacterial peptides with different concentrations was taken, appropriately diluted with sterile PBS, and evenly spread on LB (for bacteria) or YPD (for yeast / fungi) agar plates, and incubated at 37°C for 16 - 24 hours (for fungi, it could be appropriately extended to 48 hours). Subsequently, the number of colonies (CFU / mL) was counted, and the antibacterial rate was calculated. The calculation formula is:
[0250] Antibacterial rate (%)=(1 - CFU of experimental group / CFU of control group)×100%
[0251] 3.3 Determination of Minimum Inhibitory Concentration
[0252] In the minimum inhibitory concentration (MIC) determination experiment, the microbroth dilution method is adopted, which is applicable to bacteria and fungi. The culture media used in the experiment vary according to the types of strains. For bacteria (P. aeruginosa, E. coli, S. aureus), Mueller-Hinton (MH) broth medium is used, while for yeast and fungi (S. cerevisiae, C. albicans and other fungi), RPMI 1640 medium (recommended to be carried out under the conditions of pH 7.0 and containing MOPS buffer) is used.
[0253] The experiment first dilutes the antimicrobial peptide two-fold. In the 96-well plate, starting from the first column, it is gradually diluted, and the final concentration range is generally set at 256 μg / mL to 0.125 μg / mL (the concentration gradient can be adjusted according to the activity of the antimicrobial peptide), and 100 μL of the antimicrobial peptide solution is added to each well. Subsequently, the bacterial suspension is added, and the concentration of the bacterial or yeast / fungal suspension is adjusted. The final bacterial concentration is adjusted to 1×106 CFU / mL, while the concentration of yeast / fungi is adjusted to 0.5×10 3 ~104 CFU / mL. 100 μL of the bacterial suspension is added to each well to make the final volume 200 μL.
[0254] In the experiment, positive and negative control groups need to be set. The positive control group contains only the bacterial suspension and the culture medium, while the negative control group contains only the culture medium and the antimicrobial peptide, without the bacterial suspension. Subsequently, the 96-well plate is incubated at 37 °C. Bacteria are incubated for 16 - 20 hours, while yeast and fungi are incubated for 24 - 48 hours. The judgment standard of MIC is to determine the concentration of the antimicrobial peptide with the lowest complete aseptic growth by visual observation or measuring OD600 with a spectrophotometer, which is the MIC value. For fungi, 2,3,5-triphenyltetrazolium chloride staining (TTC) or XTT staining method can be used to enhance the visual judgment.
[0255] To improve the accuracy and repeatability of the experiment, the following optimization measures can be taken. First, use the McFarland standard (such as 0.5 McFarland, corresponding to about 1×108 CFU / mL) to adjust the concentration of the bacterial solution to ensure the consistency of the experiment. Second, appropriately extend the incubation time. Since fungi grow slowly, it is recommended to incubate yeast for 24 hours and Candida albicans for 48 hours to improve the accuracy of MIC determination. In addition, each experiment should be repeated at least three times, and the average value is taken to reduce experimental errors. Finally, the minimum bactericidal concentration (MBC) determination can be combined. Samples are taken from the turbid bacterial solution after MIC determination for plate counting to distinguish between antibacterial and bactericidal effects.
[0256] Example 1: Antimicrobial Peptide Screening Framework Based on Protein Language Model and Bioinformatics Method
[0257] This example relates to an antimicrobial peptide screening framework based on a protein language model and bioinformatics methods, and its process is as Figure 1 shown, including the following steps:
[0258] (A1) Use a large-scale protein language model to generate a candidate sample pool of antimicrobial peptides. The large-scale protein language models used are AMPGen and AMPGen-MIC;
[0259] (A2) Input the above-mentioned antimicrobial peptide candidate samples into the AMPGen-Filtering model for screening. The AMPGen-Filtering model screens the candidate samples through two stages: machine learning-based screening and posterior verification. The methods of the basic screening and posterior verification are as shown in the above 2.1 Machine Learning-Based Screening and 2.2 Posterior Verification.
[0260] Example 2: Screening of Antimicrobial Peptides Generated by the AMPGen Model
[0261] This example relates to obtaining the target sequence of antimicrobial peptides using the antimicrobial peptide screening framework and the antimicrobial peptides generated by the AMPGen model.
[0262] By inputting the candidate antimicrobial peptides generated by the AMPGen model into the AMPGen-Filtering model for screening and verification, 4 antimicrobial peptides with expected properties and MIC activity are finally obtained.
[0263] Test the molar mass and MIC of the above 4 antimicrobial peptides. The specific steps are as shown in 3 Detection of Antimicrobial Peptide Properties.
[0264] The results of the molar mass and MIC of the 4 antimicrobial peptides are shown in Table 1-2:
[0265] Table 1. Molar Mass of 4 Antimicrobial Peptides
[0266]
[0267] Table 2. MIC (μg / mL) of 4 Antimicrobial Peptides
[0268]
[0269] It can be seen from the results in Table 2 that:
[0270] ZJAMP006 had the lowest MIC against Staphylococcus aureus (4 μg / mL) and Fusarium graminearum (8 μg / mL), indicating its effectiveness against Gram-positive bacteria and Fusarium graminearum. ZJAMP013 had a certain effect on Staphylococcus aureus (8 μg / mL), but other strains were not tested and further research is needed.
[0271] ZJAMP015 (MIC against Staphylococcus aureus = 32 μg / mL) had poor activity and may require a higher concentration to effectively inhibit bacteria.
[0272] The above results also reflected the need for further optimization for specific strains. For example, for Pseudomonas aeruginosa and Saccharomyces cerevisiae, higher concentrations of antimicrobial peptides may be required to produce good antibacterial effects; for Escherichia coli and Staphylococcus aureus, most antimicrobial peptides showed good performance at 50 μg / mL, indicating their sensitivity to these antimicrobial peptides.
[0273] The amino acid sequences of the 4 antimicrobial peptides are shown in Table 3.
[0274] Table 3. Amino acid sequences of 4 antimicrobial peptides
[0275]
[0276] Example 3: Screening of antimicrobial peptides generated by the AMPGen-MIC model
[0277] This example involves obtaining the target sequence of antimicrobial peptides using the antimicrobial peptide screening framework and the antimicrobial peptides generated by the AMPGen-MIC model.
[0278] By inputting the candidate antimicrobial peptides generated by the AMPGen-MIC model into the AMPGen-Filtering model for screening and verification, 13 antimicrobial peptides with expected properties and MIC activity were finally obtained.
[0279] The molar mass, antibacterial rate, and MIC of the above 13 antimicrobial peptides were tested. The experimental methods were the same as the corresponding experimental parts in Example 2.
[0280] The results of the molar mass, antibacterial rate, and MIC of the 13 antimicrobial peptides are shown in Tables 4 - 6:
[0281] Table 4. Molar mass of 13 antimicrobial peptides
[0282]
[0283]
[0284] Table 5. Antibacterial rate of 13 antimicrobial peptides
[0285]
[0286] In Table 5, the two values 50 and 100 below each bacterial strain represent the inhibition rates (%) measured at two antimicrobial peptide concentrations of 50 μg / mL and 100 μg / mL, respectively. Therefore, the data in the table represent the antibacterial effects of each antimicrobial peptide on different microorganisms (Pseudomonas aeruginosa, Saccharomyces cerevisiae, Candida albicans, Escherichia coli, Staphylococcus aureus) at two concentrations: 50 represents the inhibition rate (%) when the antimicrobial peptide concentration is 50 μg / mL; 100 represents the inhibition rate (%) when the antimicrobial peptide concentration is 100 μg / mL.
[0287] The calculation of the inhibition rate is usually based on the colony growth between the control group (without antimicrobial peptide) and the experimental group (containing antimicrobial peptide), and the formula is:
[0288] Inhibition rate (%) = (1 - CFU or OD600 of the experimental group / CFU or OD600 of the control group) × 100%,
[0289] The higher the inhibition rate, the stronger the antibacterial effect of the antimicrobial peptide.
[0290] From the results of the inhibition rates in Table 5, it can be concluded that ZJAPM051, ZJAPM052, and ZJAPM062 are relatively ideal broad-spectrum antimicrobial peptides in this experiment, showing high inhibition rates on all bacterial strains and worthy of further study. ZJAPM064 is particularly effective against Escherichia coli and Staphylococcus aureus and may be suitable for applications targeting these two bacteria. ZJAPM065 has good antibacterial effects on Escherichia coli and Staphylococcus aureus at high concentrations, but its activity is low at low concentrations.
[0291] Table 6. MIC (μg / mL) of 13 antimicrobial peptides
[0292]
[0293]
[0294] From the results of the inhibition rates in Table 6, it can be concluded that ZJAPM051, ZJAPM052, and ZJAPM065 have low MIC and high inhibition rates on multiple bacterial strains and are worthy of further study. ZJAPM061 and ZJAPM062 have moderate MIC on Pseudomonas aeruginosa and Candida albicans and may have good effects on specific bacterial strains.
[0295] The amino acid sequences of 13 antimicrobial peptides are shown in Table 7.
[0296] Table 7. Amino acid sequences of 13 antimicrobial peptides
[0297]
[0298] Discussion
[0299] In the screening example of the antimicrobial peptides generated by the AMPGen model, more bacterial species can be tested subsequently. For example, the MIC of Escherichia coli, Candida albicans, and Fusarium graminearum against ZJAMP009, ZJAMP013, and ZJAMP015 can be determined to confirm their antibacterial spectra; or antimicrobial peptides with lower MIC values, such as ZJAMP006, can be tested to further explore their minimum effective concentrations.
[0300] In the screening example of the antimicrobial peptides generated by the AMPGen-MIC model, the experimental results can be further improved in the following ways: increasing more refined concentration gradients (such as 25 μg / mL, 75 μg / mL) to determine the optimal concentration range of the antimicrobial peptides; measuring the MIC (minimum inhibitory concentration) and MBC (minimum bactericidal concentration) to further confirm the effective concentration range of the antimicrobial peptides; investigating the effect of the antimicrobial peptides on the bacterial growth curve and analyzing the specific mechanism of their action, such as whether it is bacteriostatic or bactericidal; further optimizing the sequence or modification of the antimicrobial peptides to enhance their effect on fungi (such as Candida albicans, Saccharomyces cerevisiae); and conducting cytotoxicity experiments to evaluate the impact of these antimicrobial peptides on mammalian cells to verify their safety.
[0301] In addition, the antimicrobial peptides can be further optimized for different bacterial species. For example, Staphylococcus aureus and Escherichia coli are sensitive to most antimicrobial peptides, and lower MIC concentrations can be tested in the future to further optimize the dosage; Pseudomonas aeruginosa and fungi have higher MIC values, and higher concentrations or structurally optimized antimicrobial peptides may be required to improve their activity.
[0302] All the documents mentioned in the present invention are cited herein by reference as if each individual document was specifically cited. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
Claims
1. A method for constructing an antimicrobial peptide screening model, characterized in that: The method comprises the steps of: (1) Machine learning screening; (2) a posteriori verification; The machine learning screening and the a posteriori verification are performed on the antimicrobial peptide candidate sequence to obtain the antimicrobial peptide target sequence.
2. The method according to claim 1, characterized in that In step (1), it includes: (A1) Using a machine learning model, the sequence similarity between the candidate antimicrobial peptide sequence and the target functional sample is evaluated through a scoring system; (A2) predictive indicators of antimicrobial activity; Wherein, the scoring system is selected from the following group: supervised learning algorithm, reinforcement learning algorithm, unsupervised learning algorithm, or semi-supervised learning algorithm; the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the following group: linear model, K nearest neighbor, random forest, or a combination thereof; the unsupervised learning algorithm is a clustering algorithm, and the clustering algorithm is selected from the following group: density clustering, or K-means clustering; The antibacterial activity index is the minimum inhibitory concentration.
3. The method according to claim 1 or 2, characterized in that In step (1), it also includes: (A3) Evaluation of antimicrobial peptide stability; (A4) Assessment of structural similarity.
4. The method according to any one of claims 1 to 3, characterized in that: In step (A1), it includes: (a1.1) using the machine learning model to convert the antimicrobial peptide candidate sequence and / or the target function sample into a numerical form; (a1.2) Calculating the similarity between the antimicrobial peptide candidate sequence and the target functional sample.
5. The method according to any one of claims 1 to 4, characterized in that: The machine learning model is selected from the group consisting of a protein language model, a deep learning model, or a combination thereof; Wherein, the protein language model is an ESM2 model; the protein language model converts the antimicrobial peptide candidate sequence and / or the target function sample into the numerical form by a method selected from the following group: a natural language processing algorithm, or a sequence alignment algorithm; the natural language processing algorithm is a word vector, and the word vector is one-hot encoding; the sequence alignment algorithm is a protein substitution scoring matrix, and the protein substitution matrix is a BLOSUM matrix; the numerical form is an embedding vector; The deep learning model is a model based on a recurrent neural network.
6. The method according to any one of claims 1 to 5, characterized in that: The scoring system is K-nearest neighbors, and the calculation for evaluating sequence similarity by the K-nearest neighbors is: Where s_f is the similarity score, w_m is the weight of each functional cluster, f_m represents the mth distance metric, x is the generated sample, y_i is the center of the i-th functional cluster, k is the number of clusters, and d is the number of distance metrics; The distance metrics are Euclidean distance, Manhattan distance and cosine similarity.
7. The method according to any one of claims 1 to 6, characterized in that: In step (2), the method comprises a step selected from the group consisting of: limiting peptide length, predicting structural features, analyzing similarity, or a combination thereof; Wherein, in step (2), it also includes screening the candidate antimicrobial peptide sequences using a method selected from the following group: integrating multi-omics data, or knowledge graph technology; the multi-omics data is proteomics; The structural features are foldability and thermal stability; Predicting structural features by a method selected from the group consisting of a protein structure prediction tool, or molecular simulation; the protein structure prediction tool is a combination of AlphaFold2 and ESMFold; The similarity is analyzed by a method selected from the group consisting of: a sequence alignment algorithm, or a structure alignment algorithm; the sequence alignment algorithm is BLAST; the structure alignment algorithm is a combination of FoldSeek and TM-align.
8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: using an evaluation model to integrate multiple indicators to screen candidate antimicrobial peptide sequences, wherein the evaluation model is AHP, and the indicators are selected from the following group: sequence similarity, structural similarity, structural characteristics, minimum inhibitory concentration, or a combination thereof.
9. A method for generating candidate antimicrobial peptide sequences, characterized in that: The method comprises generating the antimicrobial peptide candidate sequence using machine learning technology; Wherein, the machine learning technology is selected from the group consisting of: a language model, a deep learning model, a machine learning algorithm, or a combination thereof; The language model adopts a technology selected from the following group: a rule-based method, a Transformer-based method, or a method based on a large-scale pre-trained model; the rule-based method is selected from the following group: a sequence-guided peptide design method, a structure-guided peptide design algorithm, or a combination thereof; the language model is selected from the following group: an AMPGen model, an AMPGen-MIC model, or a combination thereof; The deep learning model is a deep generative model; the deep generative model includes: a variational autoencoder (VAE) and a generative adversarial network (GAN); The machine learning algorithm is an evolutionary algorithm; the evolutionary algorithm is selected from the following group: a genetic algorithm, a differential evolution algorithm, a co-evolution algorithm, or a distribution estimation algorithm.
10. An antimicrobial peptide screening framework or system, characterized in that: The framework or system includes: An input unit, wherein the input unit is configured to input data, wherein the data includes an input sequence to be screened, wherein the input sequence to be screened includes an antimicrobial peptide candidate sequence generated by the method of claim 9; An antimicrobial peptide screening unit, the screening unit being configured to execute an antimicrobial peptide screening model, thereby obtaining an antimicrobial peptide target sequence from the antimicrobial peptide candidate sequence; wherein the screening model is constructed using the method of claim 1; An output unit, wherein the output unit is configured to output the antimicrobial peptide target sequence.
Citation Information
Cited By
Antibacterial peptide function interpretable prediction method and system based on graph causal learning
CN120808899A
Multi-target antibacterial peptide sequence optimization method and system
CN120877880A
A multi-target antimicrobial peptide sequence optimization method and system
CN120877880B