This application discloses a method,
system, and application for predicting
antimicrobial peptide activity, relating to the fields of
bioinformatics and
peptide drug design. The method includes: training multiple classifiers using a sample dataset; dividing the activity dataset into several data subgroups; employing a K-fold cross-validation strategy, using each data subgroup as the validation set and the remaining data subgroups as the
training set to
train a classification model and a regression model corresponding to each data subgroup; preprocessing the target
amino acid sequence to generate candidate peptides; inputting the ESM
feature vector of each candidate
peptide into the trained multiple classifiers, the trained classification model corresponding to the data subgroup to which the candidate peptide belongs, and the regression model, respectively, to obtain the predicted probability that the candidate peptide is an
antimicrobial peptide, the probability of having inhibitory activity against the target
pathogen, and the predicted value of the inhibitory
activity intensity against the target
pathogen; and selecting the final
list of candidate
antimicrobial peptides, thereby improving prediction accuracy and
data quality in the case of small samples.