Antibacterial peptide activity prediction method, system and application

By employing K-fold cross-validation and a multi-task learning framework, the problems of model overfitting and multi-task integration in antimicrobial peptide design were solved, ensuring the diversity of candidate peptides and achieving high accuracy and stability in antimicrobial peptide prediction, which was then applied to the prevention and control of citrus Huanglongbing (HLB).

CN122392642APending Publication Date: 2026-07-14GANNAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GANNAN NORMAL UNIV
Filing Date
2026-05-06
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

In antimicrobial peptide design, existing technologies face challenges such as model overfitting due to extremely limited specific activity data, integration of multi-task prediction frameworks, and ensuring diversity of candidate peptide outputs. In particular, the prediction accuracy is insufficient and the data quality is low in small sample cases.

Method used

We employ a K-fold cross-validation strategy and a multi-task learning framework to construct a high-quality training dataset through systematic integration of multi-source positive sample data and scientific negative sample screening. Combined with length partitioning and similarity control, we screen out the final list of candidate antimicrobial peptides.

Benefits of technology

The prediction accuracy and data quality were improved in small sample cases. The screened antimicrobial peptides showed significant inhibitory effects on the pathogen of citrus Huanglongbing, with an in vitro antibacterial rate of over 80% and an in vivo antibacterial rate of 34.55%~86.13%, making them suitable for the preparation of anti-citrus Huanglongbing related products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122392642A_ABST
    Figure CN122392642A_ABST
Patent Text Reader

Abstract

This application discloses a method, system, and application for predicting antimicrobial peptide activity, relating to the fields of bioinformatics and peptide drug design. The method includes: training multiple classifiers using a sample dataset; dividing the activity dataset into several data subgroups; employing a K-fold cross-validation strategy, using each data subgroup as the validation set and the remaining data subgroups as the training set to train a classification model and a regression model corresponding to each data subgroup; preprocessing the target amino acid sequence to generate candidate peptides; inputting the ESM feature vector of each candidate peptide into the trained multiple classifiers, the trained classification model corresponding to the data subgroup to which the candidate peptide belongs, and the regression model, respectively, to obtain the predicted probability that the candidate peptide is an antimicrobial peptide, the probability of having inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity intensity against the target pathogen; and selecting the final list of candidate antimicrobial peptides, thereby improving prediction accuracy and data quality in the case of small samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of bioinformatics and peptide drug design, and in particular to a method, system and application for predicting the activity of antimicrobial peptides. Background Technology

[0002] Currently, there are three interrelated but unresolved technical challenges in the field of antimicrobial peptide prediction: (1) The challenge of effective modeling with extremely limited specific activity data: In the design of antimicrobial peptides targeting specific pathogens (such as the pathogen of citrus Huanglongbing (HLB)), experimentally validated activity data is extremely scarce, with typically fewer than 30 labeled samples available. Under these conditions, traditional machine learning methods face the following challenges: severe overfitting due to insufficient training samples (performance degradation of more than 40% on the test set); poor model generalization ability; significant performance fluctuations under different data partitions (AUC standard deviation > 0.15); and the inability to establish a reliable dose-response relationship model (R²). 2 <0.5).

[0003] (2) System integration challenges of multi-task prediction framework: Ideal antimicrobial peptide design requires simultaneous optimization of multiple activity indicators, including general antimicrobial activity (broad spectrum), specific pathogen inhibitory activity (targeting), and activity intensity prediction (dose effect). Existing single-task prediction methods cannot coordinate multi-objective optimization under a unified framework, resulting in: inconsistent prediction results for each task, making it difficult to make comprehensive decisions; model complexity increases linearly with the number of tasks, leading to high maintenance costs; lack of knowledge transfer between tasks, resulting in low data utilization efficiency.

[0004] (3) The challenge of ensuring the diversity of candidate peptide output: Traditional prediction methods tend to output candidate peptides with similar sequence characteristics, resulting in: a single activity mechanism, which limits the discovery of new modes of action (candidate peptides with similarity >0.8 account for more than 70%); concentrated length distribution, which cannot cover antimicrobial peptides with different structural types (80% of candidate peptides are concentrated in a narrow length range); and low resource utilization efficiency for subsequent experimental verification (redundant verification wastes more than 50% of experimental resources).

[0005] Therefore, there is an urgent need to develop an antimicrobial peptide prediction method that can systematically solve the above three technical problems, and to improve the integrated antimicrobial peptide screening system of "model prediction - in vitro Lcr screening - in vivo HLB efficacy" by combining the evaluation of the actual control effect of CLas in citrus. Summary of the Invention

[0006] The purpose of this application is to provide a method, system, and application for predicting antimicrobial peptide activity, which can improve the prediction accuracy and data quality in small sample situations.

[0007] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a method for predicting antimicrobial peptide activity, comprising: Obtain a sample dataset and an active dataset of the target pathogen; the sample dataset includes several sample amino acid sequences; the sample amino acid sequences in the sample dataset include antimicrobial peptide amino acid sequences and non-antimicrobial peptide amino acid sequences; the active dataset includes several target amino acid sequences of the target pathogen, as well as the probability and intensity of the target pathogen inhibitory activity of each target amino acid sequence. Feature extraction is performed on the sample amino acid sequence and the target amino acid sequence to obtain the ESM feature vector of each sample amino acid sequence and the ESM feature vector of the target amino acid sequence. Using the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output, multiple classifiers are trained to obtain multiple trained classifiers. The active dataset was divided into several data subgroups. A K-fold cross-validation strategy was adopted, with each data subgroup serving as the validation set and the remaining data subgroups serving as the training set. The ESM feature vector of the target amino acid sequence was used as the input, and the probability that the target amino acid sequence has inhibitory activity against the target pathogen and the intensity of inhibitory activity against the target pathogen were used as the outputs to train the classification model and the regression model, respectively, to obtain the trained classification model and regression model corresponding to each data subgroup. The target amino acid sequence in each data subgroup is preprocessed to obtain the candidate peptides corresponding to the target amino acid sequence, and a candidate peptide library is generated for each data subgroup. For each candidate peptide, the ESM feature vector of the candidate peptide is input into multiple pre-trained classifiers, the pre-trained classification model and regression model corresponding to the data subgroup to which the candidate peptide belongs, respectively, to obtain the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen. The final list of candidate antimicrobial peptides is selected from the candidate peptide library based on the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted intensity of the inhibitory activity against the target pathogen.

[0008] Secondly, this application provides an antimicrobial peptide activity prediction system, comprising: The dataset acquisition module is used to acquire sample datasets and active datasets of target pathogens. The sample dataset includes several sample amino acid sequences. The sample amino acid sequences in the sample dataset include antimicrobial peptide amino acid sequences and non-antimicrobial peptide amino acid sequences. The active dataset includes several target amino acid sequences of target pathogens, as well as the probability and intensity of the target pathogen inhibitory activity of each target amino acid sequence. The feature extraction module is used to extract features from the sample amino acid sequence and the target amino acid sequence to obtain the ESM feature vector of each sample amino acid sequence and the ESM feature vector of the target amino acid sequence. The model training module is used to train multiple classifiers by taking the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output, to obtain multiple trained classifiers. The active dataset is divided into several data subgroups. A K-fold cross-validation strategy is adopted, with each data subgroup as the validation set and the remaining data subgroups as the training set. The ESM feature vector of the target amino acid sequence is taken as input, and the probability that the target amino acid sequence has the inhibitory activity of the target pathogen and the intensity of the inhibitory activity of the target pathogen are taken as outputs, respectively, to train the classification model and the regression model corresponding to each data subgroup. The candidate peptide generation module is used to preprocess the target amino acid sequence in each data subgroup to obtain the candidate peptide corresponding to the target amino acid sequence and generate a candidate peptide library for each data subgroup. The prediction module is used to input the ESM feature vector of each candidate peptide into multiple pre-trained classifiers, the pre-trained classification model and regression model corresponding to the data subgroup to which the candidate peptide belongs, and obtain the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity intensity against the target pathogen. The antimicrobial peptide screening module is used to select the final list of candidate antimicrobial peptides from the candidate peptide library based on the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen.

[0009] Thirdly, this application provides an application of the prediction method or prediction system in the prediction of antimicrobial peptide activity or the design of antimicrobial peptides.

[0010] The type of antimicrobial peptide is not particularly limited, but in some embodiments, it is preferred to include antimicrobial peptides targeting the pathogen of citrus Huanglongbing.

[0011] Fourthly, this application provides an antimicrobial peptide, wherein the amino acid sequence of the antimicrobial peptide preferably includes the amino acid sequence shown in any one of SEQ ID NO.1 to SEQ ID NO.18.

[0012] The antimicrobial peptide is preferably obtained by screening using the prediction method or the prediction system.

[0013] Fifthly, this application provides the application of the aforementioned antimicrobial peptide in the preparation of products resistant to citrus Huanglongbing (HLB).

[0014] The product preferably includes a product capable of inhibiting the pathogen of citrus Huanglongbing (HLB); the pathogen of citrus Huanglongbing preferably includes... Liberibacter crescens The type of product is not specifically limited, but preferably includes drugs, antibacterial agents, or control agents for the prevention and treatment of citrus Huanglongbing (HLB).

[0015] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a method, system, and application for predicting antimicrobial peptide activity. A more reliable training dataset is constructed through systematic integration of multi-source positive samples (antimicrobial peptide amino acid sequences) and scientific screening of negative samples (antimicrobial peptide amino acid sequences). For the extremely limited HLB activity data, a K-fold cross-validation strategy is designed to divide the activity dataset into several data subgroups. For each data subgroup, the remaining data subgroups are used as the training set to train a corresponding pre-trained classification model and a pre-trained regression model, effectively solving the model training problem under extremely limited activity data. A multi-task learning framework (including multiple pre-trained classifiers, a pre-trained classification model, and a pre-trained regression model) is used to screen the final list of candidate antimicrobial peptides, considering both the general characteristics of antimicrobial peptides and the specificity of the target pathogen. Multi-task learning and multi-model integration improve the accuracy and stability of the prediction.

[0016] The 18 antimicrobial peptides (SEQ ID NO.1~SEQ ID NO.18) screened by the predictive method of this application showed an in vitro inhibition rate of over 80% against the pathogen of citrus Huanglongbing (HLB). At the same time, they also showed excellent HLB control effects in citrus, with an in vivo inhibition rate of 34.55%~86.13%. This indicates that the antimicrobial peptides screened in this application have a significant inhibitory effect on HLB-related pathogens and can be used in the preparation of products related to anti-HLB. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an application environment diagram of an antimicrobial peptide activity prediction method in Example 1 of this application.

[0019] Figure 2 This is a flowchart illustrating an antimicrobial peptide activity prediction method provided in Embodiment 1 of this application.

[0020] Figure 3This is a detailed flowchart illustrating the method for predicting antimicrobial peptide activity provided in Example 1 of this application.

[0021] Figure 4 This is a schematic diagram of the functional modules of an antimicrobial peptide activity prediction system provided in Embodiment 2 of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] While existing technologies address general antimicrobial peptide prediction on large datasets, they do not address model stability in small-sample scenarios. Related technologies propose multi-task learning frameworks, but these cannot guarantee the reliability of predictions for each task when data is scarce. They also consider sequence diversity but lack systematic length partitioning and similarity control strategies. To address these issues, this application provides an antimicrobial peptide activity prediction method that overcomes the shortcomings of current methods in terms of insufficient prediction accuracy and low data quality in small-sample scenarios. Specifically, it involves an antimicrobial peptide activity prediction method based on a high-quality training dataset, a K-fold cross-validation strategy, and multi-task machine learning, and its application in antimicrobial peptide design, thus overcoming the shortcomings of current methods in terms of insufficient prediction accuracy and low data quality in small-sample scenarios.

[0024] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] Example 1 The antimicrobial peptide activity prediction method provided in this application embodiment can be applied to, for example... Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on other servers. Terminal 102 can send a dataset to server 104. After receiving the dataset, server 104, for that dataset, uses multiple pre-trained classifiers obtained from sample datasets acquired from public databases. It divides the active dataset into several data subgroups and uses a K-fold cross-validation strategy, sequentially using each data subgroup as the validation set and the remaining data subgroups as the training set, to train the classification model and regression model corresponding to each data subgroup. It preprocesses the target amino acid sequence to generate candidate peptides. The ESM feature vector of each candidate peptide is input into the multiple pre-trained classifiers, the pre-trained classification model corresponding to the data subgroup to which the candidate peptide belongs, and the predicted probability that the candidate peptide is an antimicrobial peptide, the probability of having inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity intensity against the target pathogen, thus selecting the final list of candidate antimicrobial peptides. Server 104 can feed back the final list of candidate antimicrobial peptides to terminal 102. Furthermore, in some embodiments, the antimicrobial peptide activity prediction method can be implemented by either server 104 or terminal 102 independently. For example, terminal 102 can directly predict antimicrobial peptide activity based on the dataset, or server 104 can retrieve the dataset from the data storage system and perform antimicrobial peptide activity prediction based on the dataset.

[0026] The terminal 102 can be, but is not limited to, various desktop computers and laptops. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers, or it can be a cloud server.

[0027] In one exemplary embodiment, such as Figure 2 As shown, a method for predicting antimicrobial peptide activity is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 206.

[0028] Step 201: Obtain the sample dataset and the active dataset of the target pathogen; the sample dataset includes several sample amino acid sequences; the sample amino acid sequences in the sample dataset include antimicrobial peptide amino acid sequences and non-antimicrobial peptide amino acid sequences; the active dataset includes several target amino acid sequences of the target pathogen and the probability and intensity of the target pathogen inhibitory activity of each target amino acid sequence.

[0029] Step 202: Extract features from the sample amino acid sequence and the target amino acid sequence to obtain the ESM feature vector of each sample amino acid sequence and the ESM feature vector of the target amino acid sequence.

[0030] Step 203: Using the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output, train multiple classifiers to obtain multiple trained classifiers; divide the active dataset into several data subgroups; adopt a K-fold cross-validation strategy, using each data subgroup as the validation set and the remaining data subgroups as the training set, using the ESM feature vector of the target amino acid sequence as input, and using the probability that the target amino acid sequence has the inhibitory activity of the target pathogen and the intensity of the inhibitory activity of the target pathogen as outputs to train the classification model and regression model, respectively, to obtain the trained classification model and regression model corresponding to each data subgroup.

[0031] Step 204: Preprocess the target amino acid sequence in each data subgroup to obtain the candidate peptides corresponding to the target amino acid sequence, and generate a candidate peptide library for each data subgroup.

[0032] Step 205: For each candidate peptide, the ESM feature vector of the candidate peptide is input into multiple trained classifiers, the trained classification model and regression model corresponding to the data subgroup to which the candidate peptide belongs, respectively, to obtain the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen.

[0033] Step 206: Select the final list of candidate antimicrobial peptides from the candidate peptide library based on the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen.

[0034] By implementing steps 201 to 206 above, a more reliable training dataset was constructed through systematic integration of multi-source positive samples (antimicrobial peptide amino acid sequences) and scientific screening of negative samples (antimicrobial peptide amino acid sequences). For the extremely limited amount of target pathogen activity data, a K-fold cross-validation strategy was designed to divide the activity dataset into several data subgroups. For each data subgroup, the remaining data subgroups were used as the training set to train the corresponding pre-trained classification and regression models, effectively solving the model training problem under extremely limited activity data. A multi-task learning framework (including multiple pre-trained classifiers, a pre-trained classification model, and a pre-trained regression model) was used to screen the final list of candidate antimicrobial peptides, considering both the general characteristics of antimicrobial peptides and the specificity of the target pathogen. Multi-task learning and multi-model integration improved the accuracy and stability of the predictions. Furthermore, this application also ensures the diversity of candidate peptides through length partitioning and similarity control.

[0035] The steps of the method for predicting antimicrobial peptide activity are as follows: Figure 3 As shown, samples are obtained from public databases to construct a high-quality sample dataset, including positive sample data integration and scientific screening of negative samples.

[0036] 1. Positive Sample Data Integration: Antimicrobial peptide amino acid sequences (i.e., positive antimicrobial peptide samples) were collected from three main data source systems and integrated in the following priority order: (1) First priority: sequences in the APD6 database that have been experimentally verified to have antibacterial activity; (2) Second priority: common antimicrobial peptide sequences in the APD6 database; (3) Third priority: data on Gram-negative bacteria from other professional websites.

[0037] A priority deduplication strategy is used during the integration process: when the same amino acid sequence appears in multiple data sources, the source with the highest priority is retained.

[0038] 2. Scientific Screening of Negative Samples: Non-antimicrobial peptide amino acid sequences (i.e., negative samples) are downloaded from the UniProt database. Sequences that may have antimicrobial activity are excluded based on multiple dimensions, including charge, hydrophobicity, length, cysteine ​​content, and pattern matching. The strict filtering conditions are as follows: (1) Charge filtering: Eliminate sequences with a net charge > 5; (2) Hydrophobic filtration: sequences with hydrophobicity <-0.2 or >0.85 are excluded; (3) Length filtering: Exclude sequences with a length < 40 amino acids; (4) Cysteine ​​content filtering: sequences with cysteine ​​content >15% are excluded; (5) Pattern matching filter: exclude sequences containing multiple antimicrobial peptide feature patterns.

[0039] The collected antimicrobial peptide amino acid sequences and non-antimicrobial peptide amino acid sequences are subjected to data balancing: positive and negative samples are balanced to a 1:1 ratio through random sampling, thereby obtaining the sample dataset in step 201.

[0040] In step 202 above, the ESM feature vector extraction is specifically implemented as follows: using the ESM (Evolutionary Scale Modeling) model variant esm2_t6_8M_UR50D, features are extracted from the sample amino acid sequence and the target amino acid sequence respectively to obtain the corresponding ESM feature vectors.

[0041] The esm2_t6_8M_UR50D model parameters include a 6-layer Transformer encoder with 8 million parameters. Input dimension: The input is an amino acid sequence, resulting in a 320-dimensional ESM feature vector.

[0042] The feature extraction process includes: for each amino acid sequence Perform the following processing: (1) Sequence preprocessing: convert to uppercase and remove leading and trailing spaces; (2) Tokenization processing: The sequence is converted into a numerical token using the ESM alphabet; (3) Forward propagation: Obtain the hidden state through a 6-layer Transformer encoder; (4) Global average pooling: This yields sequence-level feature vectors, i.e., ESM feature vectors. .

[0043] Batch processing optimization: A batch processing strategy is adopted to extract features from several sample amino acid sequences or several target amino acid sequences, with batch_size = 32, and CPU acceleration is used for computation.

[0044] Multiple classifiers, including a first classifier, a second classifier, and a third classifier, are trained separately using the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output. Specifically, the first classifier, the second classifier, and the third classifier are trained separately using the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output, resulting in a trained first classifier, a trained second classifier, and a trained third classifier.

[0045] The first classifier, the second classifier, and the third classifier are LightGBM classifier, Random Forest classifier, and XGBoost classifier, respectively.

[0046] Taking the target pathogen as the citrus Huanglongbing (HLB) pathogen as an example, i.e., the target pathogen is the HLB pathogen, the data processing and K-fold cross-validation strategy specifically includes: To address the issue of limited HLB live data, a K-fold cross-validation strategy is employed to obtain the trained classification and regression models for each data subset. (1) The limited HLB activity data were randomly divided into 4 HLB activity data subgroups, named group1, group2, group3 and group4 respectively.

[0047] (2) For each data subgroup, the complete prediction model is trained using the data from the other 3 data subgroups. That is, the active dataset is divided into 4 data subgroups, which are mutually exclusive subsets. There are 4 groups, and each group contains 6-7 active samples. For each data subgroup, the current data subgroup is used as the validation set, and the other 3 data subgroups are used as the training set. The classification model is trained with the ESM feature vector of the target amino acid sequence as the input and the probability that the target amino acid sequence has the inhibitory activity against the target pathogen as the output. The trained classification model is obtained. The regression model is trained with the ESM feature vector of the target amino acid sequence as the input and the intensity of the inhibitory activity against the target pathogen as the output.

[0048] (3) Each data subgroup independently generates a candidate peptide library and performs activity prediction; each group independently generates a candidate peptide library based on the original candidate amino acid sequence through sliding window slicing and single point mutation; the model of the corresponding group is used to predict the target pathogen inhibitory activity of the candidate peptides in that group.

[0049] The target amino acid sequence in each data subgroup is preprocessed to obtain the candidate peptide corresponding to the target amino acid sequence, and a candidate peptide library corresponding to each data subgroup is generated. Specifically, this includes: performing sliding window slicing and single-point mutation processing on the target amino acid sequence in each data subgroup to obtain the candidate peptide corresponding to the target amino acid sequence. All candidate peptides corresponding to the target amino acid sequence in each data subgroup constitute the candidate peptide library corresponding to each data subgroup.

[0050] (4) Highly active candidate peptides were screened from the four sets of prediction results.

[0051] The technical advantages of this K-fold cross-validation strategy are: 1) Maximizing data utilization: Each sample has the opportunity to be used as both a training set and a test set; 2) Avoiding overfitting: Cross-validation ensures the model's generalization ability; 3) Result reliability: The consistency of multiple independent prediction results improves the screening credibility.

[0052] The classification model can be a logistic regression classifier, and the regression model can be a Bayesian ridge regression model.

[0053] This embodiment uses a multi-task machine learning framework for training, which consists of three core prediction tasks: 1. General Antimicrobial Peptide Classification Task: The input is the ESM feature vector of the amino acid sequence, and the output is the probability P_amp that the sequence is an antimicrobial peptide. The model consists of three classifiers: LightGBM, Random Forest, and XGBoost. The predicted outputs of these three models are weighted to obtain the combined probability that the sequence is an antimicrobial peptide. Training process: A balanced dataset is used, with the training and test sets split in an 8:2 ratio. Hyperparameters are tuned using five-fold cross-validation.

[0054] 2. HLB Classification Task: The input is an ESM feature vector of an amino acid sequence, and the output is the probability P_hlb_class that the sequence has inhibitory activity against the target pathogen. Model: Logistic Regression classifier, using hierarchical cross-validation to handle small sample sizes. Decision function: ,in, Let be the decision function. For the sigmoid function, and For decision parameters, For input.

[0055] 3. HLB Regression Task: The input is the ESM feature vector of the amino acid sequence, and the output is the predicted value of the inhibitory activity intensity of the target pathogen, y_reg. and forecast uncertainty _reg. Model: Bayesian Ridge Regression Model, providing uncertainty estimates based on Bayesian Ridge Regression Model inferences.

[0056] 1. Algorithm complementary combination strategy: (1) Gradient boosting decision tree combination: LightGBM classifier: gradient boosting based on histogram algorithm; XGBoost classifier: extreme gradient boosting, with regularization to control model complexity; (2) Ensemble learning algorithm: Random Forest classifier: Bagging ensemble based on Bootstrap sampling; (3) Combination of linear models: including logistic regression (linear classification model) classifier and Bayesian ridge regression (Bayesian linear regression) model.

[0057] The following section focuses on hyperparameter tuning for the LightGBM classifier, Random Forest classifier, XGBoost classifier, logistic regression classifier, and Bayesian ridge regression model. Hyperparameter tuning methods include: (1) Grid search strategy: Five-fold cross-validation combined with grid search is used for hyperparameter optimization. The core hyperparameter search range of each algorithm includes: a) LightGBM classifier: Number of trees: 100-200; Learning rate: 0.05-0.1; Maximum depth: 5-7 layers; Automatic class weight balancing; b) Random Forest classifier: Number of decision trees: 100-200; Tree depth control: 10-20 layers or unlimited; Minimum number of samples for node splitting: 2-5; Automatic class weight balancing; c) XGBoost classifier: Number of trees: 100-200; Learning rate: 0.05-0.1; Maximum depth: 5-7 layers; Class weights adjusted based on sample proportion.

[0058] (2) Optimization objective: Using ROC-AUC score as the main optimization indicator, the optimal parameter combination is selected through five-fold cross-validation.

[0059] 3. Model Evaluation Criteria: (1) Comprehensive evaluation of multiple indicators: Classification task: The trained classification model is evaluated using metrics such as ROC-AUC, accuracy, precision, recall, and F1 score. The model that meets the validation requirements is considered the trained classification model.

[0060] Regression task: using R 2 The trained regression model is evaluated using metrics such as score, RMSE, MAE, and explained variance. The model that meets the validation requirements is used as the trained regression model.

[0061] (2) Model selection strategy: The hierarchical selection criteria are adopted, with the main indicators ranking in the top 3 and the secondary indicators all being better than the benchmark value.

[0062] Through the above process, we obtain the trained first classifier, the trained second classifier, the trained third classifier, the trained classification model, and the trained regression model.

[0063] For each candidate peptide, the ESM feature vector of the candidate peptide is input into multiple pre-trained classifiers to obtain the predicted probability P_amp that the candidate peptide is an antimicrobial peptide. The process of determining the predicted probability that the candidate peptide is an antimicrobial peptide includes: inputting the ESM feature vector of the candidate peptide into a first pre-trained classifier, a second pre-trained classifier, and a third pre-trained classifier to obtain a first predicted probability, a second predicted probability, and a third predicted probability; and then performing a weighted summation of the first predicted probability, the second predicted probability, and the third predicted probability to obtain the predicted probability that the candidate peptide is an antimicrobial peptide.

[0064] The ESM feature vector of the candidate peptide is input into the trained classification model corresponding to the data subgroup to which the candidate peptide belongs, to obtain the probability P_hlb_class of having inhibitory activity against the target pathogen.

[0065] The ESM feature vector of the candidate peptide is input into the trained regression model corresponding to the data subgroup to which the candidate peptide belongs, to obtain the predicted value y_reg of the inhibitory activity intensity of the target pathogen.

[0066] In step 206 above, the final list of candidate antimicrobial peptides is selected from the candidate peptide library based on the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen. This specifically includes the following steps 301 to 306.

[0067] Step 301: The predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted intensity of inhibitory activity against the target pathogen are weighted and summed to obtain the comprehensive priority score of the candidate peptide.

[0068] Weighted integration strategy: Overall priority score S group = Antimicrobial peptide classification score + Specific classification score + Specific regression score, where , , The weights are respectively the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted intensity of the inhibitory activity against the target pathogen. The antimicrobial peptide classification score is the predicted probability that the candidate peptide is an antimicrobial peptide, the specificity classification score is the probability that the candidate peptide has inhibitory activity against the target pathogen, and the specificity regression score is the predicted intensity of the inhibitory activity against the target pathogen of the candidate peptide.

[0069] The value range for each weight is: 0.3 ≤ ≤0.5, 0.2≤ ≤0.4, 0.2≤ ≤0.4, and + + =1. This embodiment uses a predefined weight allocation scheme: 0.4 0.3 It is 0.3.

[0070] Step 302: For each candidate peptide, calculate the consistency score of the candidate peptide based on the first predicted probability, the second predicted probability, the third predicted probability, the probability of having inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen.

[0071] Step 303: For each pair of candidate peptides, calculate the sequence similarity between them based on the Levenshtein distance.

[0072] Step 304: Sort the candidate peptides according to their comprehensive priority score, and select candidate peptides that meet the similarity requirements based on the sequence similarity between pairs of candidate peptides. The similarity threshold is dynamically adjusted based on the number of candidate peptides (range: 0.6-0.8). In one example, a similarity threshold of 0.7 is set, meaning the sequence similarity between any two candidate peptides does not exceed 0.7. Candidate peptides that meet the similarity requirements are selected to ensure sufficient sequence diversity in the selected peptides.

[0073] Step 305: Select a quota based on the preset peptide length range to obtain candidate peptides for each peptide length range.

[0074] Length partitioning strategy: Quota selection is based on three intervals divided according to peptide length. Short peptide range: 10-12 amino acids, quota 30%-40%, specifically up to 12 amino acids; Medium peptide range: 13-16 amino acids, quota 40%-50%, specifically up to 15 amino acids; Long peptide range: 17-20 amino acids, quota 20%-30%, specifically up to 8 amino acids.

[0075] Step 306: Screen candidate peptides for each peptide length range based on their consistency scores to obtain the final candidate antimicrobial peptide list. Specifically, if a candidate peptide's consistency score meets a set standard, it is added to the final candidate antimicrobial peptide list. The standard is that any three of the trained classifiers (first, second, and third classes) must predict the candidate peptide's performance by a value greater than a set threshold (0.7), and the predicted performance of the trained regression model must be greater than the top 20% quantile. The top 20% quantile is determined by ranking all candidate peptides in the candidate peptide library by their predicted target pathogen inhibitory activity intensity from largest to smallest.

[0076] A consistency score is calculated based on the consistency of high activity predictions across multiple models (the five trained models mentioned above: the trained first classifier, the trained second classifier, the trained third classifier, the trained classification model, and the trained regression model).

[0077] Consistency score = (Number of highly active models) / (Total number of models), where the criteria for judging high activity are: Classification model (including trained first classifier, trained second classifier, trained third classifier, and trained classification model): Prediction probability > 0.7, used in classification model, indicates that the confidence level of predicting it as an active peptide exceeds 70%; Regression model: Predicted value > top 20 percentile, used in regression model, indicates that the predicted activity intensity ranks in the top 20% of all candidate peptides.

[0078] For the three classification models (first classifier, second classifier, and third classifier), if the predicted probability of the candidate peptide output by the classifier is greater than 0.7, the classifier is considered to have high activity. For the regression model, the predicted value of the target pathogen inhibitory activity intensity output by the regression model needs to be compared with the predicted values ​​of the target pathogen inhibitory activity intensity of all peptides in the current candidate peptide library. If the predicted value of the target pathogen inhibitory activity intensity of the candidate peptide is higher than the top 20% percentile (i.e., the 80th percentile) of the predicted values ​​of the target pathogen inhibitory activity intensity of all candidate peptides sorted by size, the regression model is considered to have high activity.

[0079] For example, suppose a candidate peptide receives predictions from five trained models—the first classifier, the second classifier, the third classifier, the classification model, and the regression model—with predictions of 0.85 (>0.7), 0.65 (≤0.7), 0.78 (>0.7), and 0.72 (>0.7), respectively, and a predicted value (regression value) of 1.05 for the inhibitory activity against the target pathogen. If the top 20% quantile of the predicted inhibitory activity against the target pathogen in the entire candidate peptide library is 0.9, then the candidate peptide is classified as highly active by four models (three classification models meet the probability threshold, and the regression value 1.05 > 0.9). Therefore, the consistency score for this candidate peptide = (number of highly active models) / (total number of models) = 4 / 5 = 0.8.

[0080] This scoring mechanism quantifies the consistency of prediction results by integrating the independent judgments of multiple heterogeneous models, effectively enhancing the reliability and robustness of the final candidate peptide screening and providing a high-confidence priority list for subsequent experimental verification.

[0081] Comprehensive screening process: (1) Sort in descending order according to comprehensive priority score; (2) Select peptides that meet the sequence similarity requirements in sequence; (3) Balance the number distribution of each length range; (4) Combine consistency score for final confirmation.

[0082] In a specific example, a candidate peptide may exist in 4 data subgroups. In this case, the candidate peptide needs to be screened by average priority score and predicted stability score.

[0083] The average priority score is a quantitative indicator of the overall activity potential of candidate antimicrobial peptides. Its calculation is based on a weighted ensemble of multi-task prediction results and averaging across data subgroups. The specific calculation steps are as follows: Average priority score across data subgroups: The arithmetic mean of the overall priority scores obtained by the candidate peptide in four independent data subgroups is used to obtain the average priority score S. 平均值 : S 平均值 =(S group1 +S group2 +S group3 +S group4 ) / 4; where S group1 S group2 S group3 S group4 These are the overall priority scores obtained by the candidate peptides in the four independent data subgroups.

[0084] Set average priority score S 平均值 A value of ≥0.8 was used as a screening criterion to ensure that candidate peptides maintained high overall activity prediction across different data divisions.

[0085] The prediction stability score is used to evaluate the consistency of high activity predictions for candidate peptides across data subgroups, measuring the robustness of the model's predictions to the training data partitioning. The specific calculation steps are as follows: (1) High activity determination of a single data subgroup: In the independent prediction model of each data subgroup, the candidate peptide is determined according to a unified high activity judgment standard: For the classification model (including the first classifier, the second classifier, the third classifier and the classification model), if the prediction probability of the candidate peptide is greater than 0.7, it is determined that the model predicts it to be highly active.

[0086] For regression models, if the predicted value of the candidate peptide output is higher than the top 20 percentile (i.e., the 80th percentile) of all predicted values ​​in the candidate peptide library of the current data subgroup, then the model is judged to predict high activity.

[0087] (2) Consistency statistics across data subgroups: The number of times a candidate peptide is judged as highly active in the four data subgroups is counted and denoted as the number of highly active groups N. A candidate peptide is judged as highly active if its consistency score is greater than the set value.

[0088] (3) Calculation of predicted stability score: R = N / Q, where R is the predicted stability score of the candidate peptide and Q is the total number of data subgroups. The predicted stability score ranges from 0 to 1. A higher score indicates better consistency in predicting high activity of the candidate peptide in different data subgroups.

[0089] A predicted stability score of R ≥ 0.8 was set as the screening criterion. Due to the four data subgroups, the candidate peptide was required to be predicted as highly active in all four data subgroups before being added to the final list of candidate antimicrobial peptides. This ensured that the activity prediction of the candidate peptides did not depend on specific data partitions and had high robustness.

[0090] The selection criteria for high-activity candidate peptides were integrated from the prediction results of four data subgroups. These criteria included: an average priority score ≥ the average priority score threshold (0.8) and being predicted as highly active in all four subgroups; and a prediction stability score ≥ the prediction stability score threshold (0.8) and consistently being predicted as highly active in all four independently trained data subgroup models. This ensured the robustness and reliability of the prediction results and met the requirements for length partitioning and diversity. By simultaneously requiring an average priority score ≥ 0.8 and a prediction stability score ≥ 0.8, this application rigorously screened highly active candidate peptides while minimizing the risk of misselection due to model overfitting or data bias, thereby improving the success rate of subsequent experimental validation.

[0091] Based on the above method, this application provides a class of 33 highly active antimicrobial peptides obtained through the prediction method, of which 19 antimicrobial peptides were chemically synthesized, and their antimicrobial activity was experimentally verified. Liberibacter crescens BT-1 is the type strain, with accession number ATCC BAA-2481. GFP-labeled [cell / organization] was constructed. Liberibacter crescens BT-1 strain ( Liberibacter crescens BT-1 / pBBR1-S-GFP, denoted as Lcr -GFP. Resuspended in 0.85% NaCl solution. Lcr -GFP, to obtain Lcr -GFP bacterial suspension; the synthesized antimicrobial peptides were dissolved in DMSO to obtain antimicrobial peptide solutions of different concentrations (16 mg / mL). 1 μL of the antimicrobial peptide solution and 199 μL of the solution were added to the ELISA plate, respectively. Lcr -GFP bacterial suspension, in 1 μL DMSO and 199 μL Lcr -GFP bacterial suspension was used as a blank control. The fluorescence intensity of each group was detected using a full-wavelength multi-mode microplate reader. The excitation wavelength was 488 nm and the emission wavelength was 535 nm. The inhibition rate was calculated based on the fluorescence values ​​of each group according to the following formula: Antibacterial rate (%) = [(Fluorescence value of blank group - Fluorescence value of group with added antimicrobial peptide solution) / Fluorescence value of blank group] × 100% Of the 19 synthesized antimicrobial peptides, 13 were identified as highly active (inhibition rate ≥ 80%), with inhibition rates ranging from 80.0% to 100.0% and an average inhibition rate of 91.3%. The amino acid sequences and inhibition rates of these 13 peptides are shown in Table 1. The remaining peptides exhibited inhibition rates between 26.8% and 66.4%, specifically including: 3 moderately active peptides (inhibition rate 50%–79%), with inhibition rates ranging from 50.1% to 66.4%; and 3 low-activity peptides (inhibition rate < 50%), with inhibition rates ranging from 26.8% to 39.6%.

[0092] Table 1. Amino acid sequences and antibacterial rates of 13 highly active antimicrobial peptides.

[0093] Multiple peptides showed significant inhibitory effects on citrus Huanglongbing-related pathogens, demonstrating the effectiveness of the proposed method in practical applications.

[0094] Based on the above methods, five highly active antimicrobial peptides obtained through the prediction method were further provided. Their in vitro antimicrobial activity was verified using the same method described above, i.e., in vitro efficacy. Furthermore, in vivo efficacy verification was conducted on the leaf injection of candidate antimicrobial peptides into citrus HLB-infected plants: the antimicrobial peptides were first prepared into a 10 mM stock solution using DMSO, then diluted with sterile water to a 20 μM working solution. 400 μL of this solution was injected into the leaves of citrus HLB-infected plants, with 200 μL injected into each side of the main vein. A solution without added antimicrobial peptides was used as a control (referred to as the Mock control). Changes in CLas titers (copy number / μg DNA) at 0, 1, 3, 5, and 7 dpi were measured, and the 7 dpi efficacy, i.e., in vivo efficacy, was calculated. The results are shown in Table 2.

[0095] Table 2. In vitro and in vivo efficacy of antimicrobial peptides

[0096] The antimicrobial peptides of this application all exhibited high antimicrobial activity in in vitro Lcr antimicrobial tests, with the five peptides showing in vitro antimicrobial rates ranging from 81.1% to 86.9%, all at a high level. In in vivo efficacy tests, the antimicrobial rates ranged from 34.55% to 86.13%. Although the antimicrobial rates of each peptide fluctuated to some extent, most peptides still showed significant antimicrobial effects. The in vitro and in vivo antimicrobial effects did not completely correspond, possibly due to factors such as the targeted delivery efficiency of the antimicrobial peptides in vivo, in vivo stability, and local concentration of action. Even with the interference of the in vivo environment, the antimicrobial peptides still maintained considerable antimicrobial activity and did not completely lose their effect due to the complex in vivo environment. This indicates that the method of this application can screen antimicrobial peptides with excellent in vitro antimicrobial activity and effective antimicrobial effects in the in vivo environment.

[0097] To fully verify the significant advancements of the technical solution presented in this application compared to existing technologies, a systematic comparative analysis was designed. The method provided in this application was compared with traditional single-model methods, ensemble methods without diversity control, traditional QSAR methods, and ensemble methods without group training. The results are as follows: Comparative Example 1: Traditional single-model method. Technical solution: A single XGBoost model is used for training with all 27 HLB samples, without grouped cross-validation.

[0098] Feature engineering: same ESM feature extraction; screening strategy: sorted only by prediction score, with no diversity control.

[0099] Detailed experimental results: Model training process: AUC=0.98 for training set and AUC=0.62 for test set, indicating significant overfitting; Prediction stability: The standard deviation of AUC after 5 repeated training iterations was 0.15, indicating poor stability; Candidate peptide output analysis: 28 out of 35 peptide segments had a sequence similarity >0.8, with lengths concentrated between 12-14 amino acids.

[0100] Experimental results showed that only 3 out of 12 synthetic peptides exhibited high activity (sterilization rate >90%), with a recognition rate of 25%.

[0101] Comparative Example 2: Integration method without diversity control. Technical solution: The same multi-model integration and group training as in this application are adopted, but the diversity screening module is removed.

[0102] Training strategy: same group cross-training; Selection strategy: sorted only based on priority score, without length partitioning and similarity control.

[0103] Detailed experimental results: Model performance: AUC=0.85 on the test set, with good predictive stability (standard deviation 0.06); Candidate peptide distribution: 31 out of 35 peptides are concentrated in the 13-15 amino acid length range; Sequence similarity: Average similarity 0.65, with obvious redundancy.

[0104] Experimental results: 8 out of 18 synthetic peptides showed high activity, with a recognition rate of 45%, but their mechanisms of activity were singular. Comparative Example 3: Traditional QSAR method, technical solution: quantitative structure-activity relationship model based on traditional physicochemical descriptors.

[0105] Feature engineering: 15 descriptors including hydrophobicity, net charge, and amino acid composition; Training strategy: linear regression and partial least squares analysis.

[0106] Detailed experimental results: Model performance: Best R 2 =0.35, unable to establish an effective structure-activity relationship; Feature importance analysis: The contributions of each descriptor are scattered, with no significant dominant factor; Predictive ability: The prediction error for the test set exceeds 50%.

[0107] Experimental results: No meaningful candidate peptide ranking can be provided.

[0108] Comparative Example 4: Ensemble method without grouping training. Technical solution: Employs multi-model ensemble training without grouping or cross-training. Training strategy: Trains a single model set using all 27 samples; Selection strategy: Includes diversity control.

[0109] Detailed experimental results: Model stability: Performance fluctuates greatly under different data partitions (AUC range 0.68-0.82); Overfitting: Significant performance gap between training and test sets; Practical application value: Difficult to apply reliably due to model instability.

[0110] The results of the systematic performance comparison analysis of the above schemes are shown in Table 3.

[0111] Table 3 Systematic Performance Comparison Analysis

[0112] Statistical significance analysis was performed using paired t-tests to compare the performance differences between this application and each comparative example: Compared with Comparative Example 1: t=8.32, p<0.001, the difference was extremely significant; Compared with Comparative Example 2: t=4.15, p<0.01, the difference was significant; Compared with Comparative Example 4: t=5.67, p<0.001, the difference was extremely significant.

[0113] Comparative analysis results: Through systematic comparative experiments, it is demonstrated that this application achieves the following through the synergistic effect of three technical features: k-fold cross-validation strategy, multi-task integration, and diversity screening: (1) Establish a stable and reliable prediction model under small sample conditions (AUC standard deviation decreased from 0.15 to 0.04). (2) Maintain high prediction accuracy for each subtask within the multi-task framework (average AUC increased from 0.62 to 0.87). (3) While ensuring activity, sufficient diversity of output peptides is achieved (similarity decreased from 0.82 to 0.32).

[0114] The method provided in this application is applicable to the functional prediction of antimicrobial peptides, antiviral peptides, and cell-penetrating peptides.

[0115] Compared with traditional methods for predicting antimicrobial peptides, this application has the following significant advantages: 1. Data quality advantage: Through systematic integration and scientific screening of multi-source positive sample data, a more reliable training dataset has been constructed.

[0116] 2. Advantages of small sample processing: A K-fold cross-validation strategy was designed for extremely small amounts of HLB active data, which effectively solves the model training problem under extremely small amounts of active data.

[0117] 3. Advantages in prediction accuracy: The multi-task learning framework is adopted, taking into account both the general characteristics of antimicrobial peptides and HLB specificity. Multi-task learning and multi-model integration improve the accuracy and stability of prediction. The accuracy and stability of prediction are further enhanced through complementary algorithm combinations and integration strategies.

[0118] 4. Diversity Guarantee Advantage: By controlling sequence similarity and allocating peptide lengths, the diversity of candidate peptides is ensured. The short peptides / mutants such as AMPs-3-13 and AMPs-3-8 screened out not only showed Lcr inhibitory activity in vitro, but also demonstrated excellent HLB control effects in citrus fruits. This breaks the traditional understanding that antimicrobial peptides are mostly 10-20 amino acids, and provides a new direction for the development of short peptide anti-HLB drugs.

[0119] 5. Practical advantages: Experimental verification results show that the antimicrobial peptides screened by the method of this application exhibit excellent bactericidal activity in experimental verification, and the method of this application can effectively guide the discovery of highly active antimicrobial peptides.

[0120] 6. Advantages of integrated in vitro and in vivo screening. The method of this application can accurately predict the in vitro Lcr antibacterial activity in small samples, and simultaneously screen out highly effective antimicrobial peptides with practical agricultural application value in vivo, solving the problem of the disconnect between the in vitro morphology and in vivo efficacy in existing technologies; the six highly effective short peptides screened have the characteristics of small molecular weight, easy synthesis, and easy transport in plants, making them more suitable for practical field applications and significantly reducing the R&D and transformation costs of antimicrobial peptides from the laboratory to agricultural production.

[0121] The above embodiments illustrate that this application effectively solves the problem of predicting antimicrobial peptide activity under small sample data by constructing high-quality training data, group cross-training, and multi-task learning, providing a reliable computational tool for the development of novel antimicrobial drugs.

[0122] This application also provides an application scenario in which the above-described antimicrobial peptide activity prediction method is applied. Specifically, the antimicrobial peptide activity prediction method provided in this embodiment can be applied to the antimicrobial peptide activity prediction scenario for citrus Huanglongbing (HLB). The antimicrobial peptide activity prediction scenario for citrus HLB includes a data acquisition stage and an antimicrobial peptide activity prediction stage; the dataset enters the antimicrobial peptide activity prediction stage from the data acquisition stage to obtain the final candidate antimicrobial peptide list. The antimicrobial peptide activity prediction method provided in this embodiment belongs to the antimicrobial peptide activity prediction stage. Specifically, in the content processing chain of the dataset, multiple pre-trained classifiers can be obtained from sample datasets acquired from public databases. The active dataset is divided into several data subgroups, and a K-fold cross-validation strategy is adopted. Each data subgroup is used as the validation set, and the remaining data subgroups are used as the training set to train the classification model and regression model corresponding to each data subgroup. The target amino acid sequence is pre-processed to generate candidate peptides. The ESM feature vector of each candidate peptide is input into the multiple pre-trained classifiers, the pre-trained classification model and regression model corresponding to the data subgroup to which the candidate peptide belongs, respectively, to obtain the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity intensity against the target pathogen, thus selecting the final list of candidate antimicrobial peptides.

[0123] Example 2 Based on the same inventive concept, this application also provides an antimicrobial peptide activity prediction system for implementing the aforementioned antimicrobial peptide activity prediction method. The solution provided by this system is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the antimicrobial peptide activity prediction system provided below can be found in the limitations of the antimicrobial peptide activity prediction method described above, and will not be repeated here.

[0124] In one exemplary embodiment, such as Figure 4 As shown, an antimicrobial peptide activity prediction system is provided, comprising: The dataset acquisition module T1 is used to acquire the sample dataset and the active dataset of the target pathogen. The sample dataset includes several sample amino acid sequences. The sample amino acid sequences in the sample dataset include antimicrobial peptide amino acid sequences and non-antimicrobial peptide amino acid sequences. The active dataset includes several target amino acid sequences of the target pathogen, as well as the probability and intensity of the target pathogen inhibitory activity of each target amino acid sequence. The feature extraction module T2 is used to extract features from the sample amino acid sequence and the target amino acid sequence to obtain the ESM feature vector of each sample amino acid sequence and the ESM feature vector of the target amino acid sequence. The model training module T3 is used to train multiple classifiers by taking the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output, to obtain multiple trained classifiers. The active dataset is divided into several data subgroups. A K-fold cross-validation strategy is adopted, with each data subgroup as the validation set and the remaining data subgroups as the training set. The ESM feature vector of the target amino acid sequence is taken as input, and the probability that the target amino acid sequence has the inhibitory activity of the target pathogen and the intensity of the inhibitory activity of the target pathogen are taken as outputs to train the classification model and the regression model, respectively, to obtain the trained classification model and regression model corresponding to each data subgroup. The candidate peptide generation module T4 is used to preprocess the target amino acid sequence in each data subgroup to obtain the candidate peptide corresponding to the target amino acid sequence and generate the candidate peptide library corresponding to each data subgroup. The prediction module T5 is used to input the ESM feature vector of each candidate peptide into multiple trained classifiers, the trained classification model and regression model corresponding to the data subgroup to which the candidate peptide belongs, and obtain the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the intensity of inhibitory activity against the target pathogen. The antimicrobial peptide screening module T6 is used to screen the final list of candidate antimicrobial peptides from the candidate peptide library based on the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen.

[0125] This application also provides the use of the above-mentioned antimicrobial peptides in the preparation of antimicrobial drugs, especially drugs for treating citrus Huanglongbing (HLB).

[0126] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0127] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for predicting antimicrobial peptide activity, characterized in that, The method for predicting antimicrobial peptide activity includes: Obtain a sample dataset and an active dataset of the target pathogen; the sample dataset includes several sample amino acid sequences; the sample amino acid sequences in the sample dataset include antimicrobial peptide amino acid sequences and non-antimicrobial peptide amino acid sequences; the active dataset includes several target amino acid sequences of the target pathogen, as well as the probability and intensity of the target pathogen inhibitory activity of each target amino acid sequence. Feature extraction is performed on the sample amino acid sequence and the target amino acid sequence to obtain the ESM feature vector of each sample amino acid sequence and the ESM feature vector of the target amino acid sequence. Using the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output, multiple classifiers are trained to obtain multiple trained classifiers. The active dataset was divided into several data subgroups. A K-fold cross-validation strategy was adopted, with each data subgroup serving as the validation set and the remaining data subgroups serving as the training set. The ESM feature vector of the target amino acid sequence was used as the input, and the probability that the target amino acid sequence has inhibitory activity against the target pathogen and the intensity of inhibitory activity against the target pathogen were used as the outputs to train the classification model and the regression model, respectively, to obtain the trained classification model and regression model corresponding to each data subgroup. The target amino acid sequence in each data subgroup is preprocessed to obtain the candidate peptides corresponding to the target amino acid sequence, and a candidate peptide library is generated for each data subgroup. For each candidate peptide, the ESM feature vector of the candidate peptide is input into multiple pre-trained classifiers, the pre-trained classification model and regression model corresponding to the data subgroup to which the candidate peptide belongs, respectively, to obtain the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen. The final list of candidate antimicrobial peptides is selected from the candidate peptide library based on the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted intensity of the inhibitory activity against the target pathogen.

2. The method for predicting antimicrobial peptide activity according to claim 1, characterized in that, Multiple classifiers include a first classifier, a second classifier, and a third classifier; Using the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output, multiple classifiers are trained separately to obtain multiple trained classifiers, specifically including: Using the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output, the first classifier, the second classifier, and the third classifier are trained respectively to obtain the trained first classifier, the trained second classifier, and the trained third classifier.

3. The method for predicting antimicrobial peptide activity according to claim 2, characterized in that, The process of determining the predicted probability that a candidate peptide is an antimicrobial peptide includes: The ESM feature vectors of the candidate peptides are input into the trained first classifier, the trained second classifier, and the trained third classifier, respectively, to obtain the first prediction probability, the second prediction probability, and the third prediction probability. The first, second, and third prediction probabilities are weighted and summed to obtain the prediction probability that the candidate peptide is an antimicrobial peptide.

4. The method for predicting antimicrobial peptide activity according to claim 2, characterized in that, The first classifier, the second classifier, and the third classifier are LightGBM classifier, Random Forest classifier, and XGBoost classifier, respectively.

5. The method for predicting antimicrobial peptide activity according to claim 1, characterized in that, The target amino acid sequence in each data subgroup is preprocessed to obtain the candidate peptides corresponding to the target amino acid sequence, generating a candidate peptide library for each data subgroup, specifically including: The target amino acid sequence in each data subgroup is subjected to sliding window slicing and single-point mutation processing to obtain candidate peptides corresponding to the target amino acid sequence. All candidate peptides corresponding to the target amino acid sequence in each data subgroup constitute the candidate peptide library for each data subgroup.

6. The method for predicting antimicrobial peptide activity according to claim 3, characterized in that, The final list of candidate antimicrobial peptides was selected from the candidate peptide library based on the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted intensity of the inhibitory activity against the target pathogen. Specifically, this list includes: The predicted probability of a candidate peptide being an antimicrobial peptide, the probability of it having inhibitory activity against the target pathogen, and the predicted intensity of its inhibitory activity against the target pathogen are weighted and summed to obtain a comprehensive priority score for the candidate peptide. For each candidate peptide, a consistency score is calculated based on the first predicted probability, the second predicted probability, the third predicted probability, the probability of having inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen. For each pair of candidate peptides, the sequence similarity between the two candidate peptides is calculated based on the Levenshtein distance; Candidate peptides are ranked according to their comprehensive priority score, and candidate peptides that meet the similarity requirements are selected based on the sequence similarity between pairs of candidate peptides. Quota selection is performed based on preset peptide length ranges to obtain candidate peptides for each peptide length range. Candidate peptides in each peptide length range are screened based on their consistency scores to obtain the final list of candidate antimicrobial peptides.

7. An antimicrobial peptide activity prediction system, characterized in that, The antimicrobial peptide activity prediction system includes: The dataset acquisition module is used to acquire sample datasets and active datasets of target pathogens. The sample dataset includes several sample amino acid sequences. The sample amino acid sequences in the sample dataset include antimicrobial peptide amino acid sequences and non-antimicrobial peptide amino acid sequences. The active dataset includes several target amino acid sequences of target pathogens, as well as the probability and intensity of the target pathogen inhibitory activity of each target amino acid sequence. The feature extraction module is used to extract features from the sample amino acid sequence and the target amino acid sequence to obtain the ESM feature vector of each sample amino acid sequence and the ESM feature vector of the target amino acid sequence. The model training module is used to train multiple classifiers by taking the ESM feature vector of the sample amino acid sequence as input and the probability that the sample amino acid sequence is an antimicrobial peptide as output, to obtain multiple trained classifiers. The active dataset is divided into several data subgroups. A K-fold cross-validation strategy is adopted, with each data subgroup as the validation set and the remaining data subgroups as the training set. The ESM feature vector of the target amino acid sequence is taken as input, and the probability that the target amino acid sequence has the inhibitory activity of the target pathogen and the intensity of the inhibitory activity of the target pathogen are taken as outputs, respectively, to train the classification model and the regression model corresponding to each data subgroup. The candidate peptide generation module is used to preprocess the target amino acid sequence in each data subgroup to obtain the candidate peptide corresponding to the target amino acid sequence and generate a candidate peptide library for each data subgroup. The prediction module is used to input the ESM feature vector of each candidate peptide into multiple pre-trained classifiers, the pre-trained classification model and regression model corresponding to the data subgroup to which the candidate peptide belongs, and obtain the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity intensity against the target pathogen. The antimicrobial peptide screening module is used to select the final list of candidate antimicrobial peptides from the candidate peptide library based on the predicted probability that the candidate peptide is an antimicrobial peptide, the probability that it has inhibitory activity against the target pathogen, and the predicted value of the inhibitory activity against the target pathogen.

8. The application of the prediction method according to any one of claims 1 to 6 or the prediction system according to claim 7 in the prediction of antimicrobial peptide activity or the design of antimicrobial peptides.

9. An antimicrobial peptide, characterized in that, The amino acid sequence of the antimicrobial peptide includes any one of the amino acid sequences shown in SEQ ID NO.1 to SEQ ID NO.

18.

10. The use of the antimicrobial peptide according to claim 9 in the preparation of products resistant to citrus Huanglongbing (HLB).