Protein language model supervised fine tuning-based antibacterial peptide design framework and generation sequence thereof

Through supervised fine-tuning technology based on protein language model, an antimicrobial peptide generation and screening framework was constructed, which solved the problems of low efficiency and insufficient accuracy of antimicrobial peptide design and screening in the existing technology, and achieved efficient and accurate antimicrobial peptide sequence generation and screening.

CN120048361APending Publication Date: 2025-05-27杭州天衍星旅科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510211662.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing antimicrobial peptide design and screening methods are inefficient and insufficiently accurate, making it difficult to efficiently generate antimicrobial peptides with target characteristics, and may ignore peptides with unconventional mechanisms of action.

Method used

Using supervised fine-tuning technology based on protein language model, an antimicrobial peptide generation and screening framework is constructed, and antimicrobial peptide candidate sequences are generated and screened by training and fine-tuning of protein base models.

Benefits of technology

The efficiency and accuracy of antimicrobial peptide design and screening are significantly improved, and the generated antimicrobial peptide sequences are diverse, novel and good antimicrobial activity, which improves its practical application potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005286169360000021
    Figure BDA0005286169360000021
  • Figure BDA0005286169360000031
    Figure BDA0005286169360000031
  • Figure BDA0005286169360000061
    Figure BDA0005286169360000061
Patent Text Reader

Abstract

The invention provides an antibacterial peptide design framework based on supervised fine tuning of a protein language model, in particular, an antibacterial peptide generation framework is developed by utilizing the protein language model through a supervised fine tuning technology, and an antibacterial peptide screening framework is designed based on various properties of antibacterial peptide. The two frameworks jointly form a novel antibacterial peptide design framework, and the problems that in the prior art, antibacterial peptide design is low in efficiency, insufficient in accuracy and the like are effectively solved. Besides, when diversified and novel antibacterial peptide sequences are generated, the key antibacterial activity of the antibacterial peptide sequences is maintained, the practical application potential of generating the antibacterial peptide sequences is remarkably improved, and an efficient, accurate and flexible new method is provided for research and development of the antibacterial peptide. Based on the antibacterial peptide design and screening model, four brand-new high-activity antibacterial peptide sequences are also found.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of bioinformatics, artificial intelligence, and biomedicine, and relates to a technology for designing, optimizing, and screening antimicrobial peptide sequences using artificial intelligence technology, especially a deep learning method based on a protein language model, and also relates to a research and development and design method for novel antimicrobial drugs. Background Art

[0002] Antimicrobial Peptides (AMP) are a class of short-chain amino acid sequences (usually defined as 12 - 50 amino acid residues), which can inhibit microbial growth by interfering with cell wall integrity. Antimicrobial peptides play an important role in coping with the increasingly serious Antimicrobial Resistance (AMR) crisis. According to the prediction of the World Health Organization, by 2050, AMR may cause 10 million deaths annually and result in a cumulative loss of $100 trillion to the global economy.

[0003] Traditional methods for discovering antimicrobial peptides include isolation from natural sources, rational design, and high-throughput screening (HTS) of synthetic peptide libraries. These methods have successfully identified many antimicrobial peptides, but also face many challenges. For example, isolation from natural sources is time-consuming and limited by the available biodiversity; rational design requires a large amount of resources for verification; HTS is limited by the initial library design and infrastructure requirements. These methods often have difficulty efficiently generating AMPs with target properties and may overlook peptides with unconventional mechanisms of action.

[0004] To overcome the above limitations, researchers have begun to explore artificial intelligence (AI)-based methods for discovering and designing antimicrobial peptides. However, existing AI-based computational methods still face two major challenges in antimicrobial peptide discovery and design: (1) Transformation from in vitro to in vivo: The transformation of computationally designed antimicrobial peptides from in vitro to in vivo efficacy requires comprehensive pharmacokinetic and pharmacodynamic studies. Key factors such as bioavailability, metabolic stability, and potential off-target effects need to be thoroughly studied in a complex biological environment; (2) Inherent limitations of computational methods, such as biases and limitations in the AMP database, difficulties in accurately predicting functional properties, and challenges in balancing diversity, novelty, and antimicrobial properties.

[0005] In summary, there is an urgent need in this field to develop an efficient and highly accurate method for designing and screening antimicrobial peptides. Summary of the Invention

[0006] The purpose of the present invention is to provide an antimicrobial peptide design framework based on supervised fine-tuning of a protein language model and a batch of antimicrobial peptide sequences obtained through this design framework.

[0007] In the first aspect of the present invention, a method for constructing an antibacterial peptide generation model is provided, and the method includes the steps of:

[0008] (1) Training a protein base model on a dataset;

[0009] (2) Performing supervised fine-tuning on the protein base model to obtain a fine-tuned protein base model and antibacterial peptide candidate sequences.

[0010] In another preferred example, in step (2), the method further includes the step of: performing reinforcement learning on the fine-tuned protein base model to obtain an optimized protein model and antibacterial peptide candidate sequences.

[0011] In another preferred example, in step (1), the dataset is a publicly available antibacterial peptide dataset.

[0012] In another preferred example, in step (2.1), the dataset is a publicly available antibacterial peptide dataset with a low minimum inhibitory concentration and a publicly available non-active antibacterial peptide dataset.

[0013] In another preferred example, the dataset is obtained from the original dataset by a data augmentation method.

[0014] In another preferred example, the data augmentation method is selected from the group consisting of:

[0015] (i) Performing sequence variation on the original dataset;

[0016] (ii) Performing structural perturbation on the original dataset;

[0017] (iii) Performing conditional generation on the original dataset;

[0018] or a combination thereof.

[0019] In another preferred example, the protein base model is a protein language model.

[0020] In another preferred example, the protein language model includes: ProGen2 model, ESM-1b, ProtBERT, ProtXLNet, TAPE.

[0021] In another preferred example, the protein language model is the ProGen2 model.

[0022] In another preferred example, the methods for performing the supervised fine-tuning include: Low-Rank Adaptation (LoRA), Prefix-tuning, Prompt-tuning, P-tuning, P-tuning v1, P-tuning v2, Adapter-tuning.

[0023] In another preferred example, the method for performing the supervised fine-tuning is low-rank adaptation.

[0024] In another preferred example, perplexity is used to guide the supervised fine-tuning.

[0025] In another preferred example, the loss function of the low-rank adaptation is:

[0026]

[0027] where N is the number of sequences in the training batch, T is the length of each sequence, \(x_{i,t}\) represents the \(t\)-th token of the \(i\)-th sequence, \(x_{i,<t}\) represents all tokens before \(t\) in the \(i\)-th sequence, and \(\theta_{LoRA}\) represents the low-rank adaptation parameters.

[0028] In another preferred example, the fine-tuned protein base model is AMPGen.

[0029] In a second aspect of the present invention, there is provided an antimicrobial peptide generation framework or system, the framework or system comprising:

[0030] (A1) An input unit configured to input data, the input data including a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence;

[0031] (A2) A fine-tuning unit configured to execute a fine-tuning model on the input data to obtain fine-tuned input data; wherein the fine-tuning model includes a supervised fine-tuning model, and the supervised fine-tuning model includes the steps of: performing supervised fine-tuning on the input data using a fine-tuning method to obtain fine-tuned input data;

[0032] (A3) An output unit configured to output the result of the fine-tuning unit.

[0033] In another preferred example, the protein base model is a protein language model.

[0034] In another preferred example, the protein language model includes: ProGen2 model, ESM-1b, ProtBERT, ProtXLNet, TAPE.

[0035] In another preferred example, the protein language model is the ProGen2 model.

[0036] In another preferred example, the fine-tuning method includes: Low-Rank Adaptation (LoRA), Adapter-tuning, Prefix-tuning, Prompt-tuning, P-tuning, P-tuning v1, P-tuning v2.

[0037] In another preferred example, the fine-tuning method is Low-Rank Adaptation.

[0038] In another preferred example, perplexity is used to guide the supervised fine-tuning model.

[0039] In another preferred example, the loss function of the Low-Rank Adaptation is:

[0040]

[0041] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the Low-Rank Adaptation parameters.

[0042] In a third aspect of the present invention, a method for constructing an antimicrobial peptide screening model is provided, and the method includes the steps:

[0043] (M1) Machine learning screening, and the machine learning screening includes the steps:

[0044] (B1) Using a machine learning model, evaluating the sequence similarity between antimicrobial peptide candidate sequences and target functional samples through a scoring system;

[0045] (B2) Predicting antimicrobial activity indicators;

[0046] (M2) Posterior verification, and the posterior verification includes steps selected from the following group:

[0047] (C1) Limiting the peptide segment length;

[0048] (C2) Predicting structural features;

[0049] (C3) Analyzing similarity;

[0050] or a combination thereof;

[0051] wherein, the scoring system is selected from the following group: supervised learning algorithm, reinforcement learning algorithm, unsupervised learning algorithm, or semi-supervised learning algorithm; the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the following group: linear model, K-nearest neighbor, or random forest; the unsupervised learning algorithm is a clustering algorithm, and the clustering algorithm is selected from the following group: density clustering, or K-means clustering;

[0052] The antibacterial activity index is the minimum inhibitory concentration;

[0053] The structural features are foldability and thermal stability;

[0054] The structural features are predicted by a method selected from the group consisting of: protein structure prediction tools, or molecular simulations; the protein structure prediction tools are a combination of AlphaFold2 and ESMFold;

[0055] The similarity is analyzed by a method selected from the group consisting of: sequence alignment algorithms, or structure alignment algorithms; the sequence alignment algorithm is BLAST; the structure alignment algorithms are a combination of FoldSeek and TM-align.

[0056] In another preferred embodiment, the method further comprises the step: (M3) high-throughput screening.

[0057] In another preferred embodiment, in step (M1), it further comprises:

[0058] (B3) Evaluating the stability of the antimicrobial peptide;

[0059] (B4) Evaluating the structural similarity.

[0060] In another preferred embodiment, in step (B1), it includes:

[0061] (b1.1) Using the machine learning model, converting the antimicrobial peptide candidate sequence and / or the target functional sample into a numerical form;

[0062] (b1.2) Calculating the similarity between the antimicrobial peptide candidate sequence and the target functional sample;

[0063] Wherein,

[0064] The machine learning model is selected from the group consisting of: protein language models, deep learning models, or combinations thereof; the protein language model is the ESM2 model;

[0065] The protein language model converts the antimicrobial peptide candidate sequence and / or the target functional sample into the numerical form by a method selected from the group consisting of: natural language processing algorithms, or sequence alignment algorithms; the natural language processing algorithm is word vectors, the word vectors are one-hot encoding; the sequence alignment algorithm is a protein substitution scoring matrix, the protein substitution matrix is the BLOSUM matrix; the numerical form is an embedding vector;

[0066] The deep learning model is a model based on a recurrent neural network.

[0067] In another preferred embodiment, the machine learning model is a combination of a protein language model and a deep learning model.

[0068] In another preferred example, the protein language model includes: the ESM2 model, ProtBERT, and ProtXLNet.

[0069] In another preferred example, the protein language model is the ESM2 model.

[0070] In another preferred example, the protein language model converts the antimicrobial peptide candidate sequence and / or the target functional sample into the numerical form by a method selected from the following group: natural language processing algorithms, or sequence alignment algorithms.

[0071] In another preferred example, the natural language processing algorithm is a word vector.

[0072] In another preferred example, the ESM2 model evaluates the sequence similarity between the bacteriocin peptide candidate sequence and the target functional sample through zero-shot learning.

[0073] In another preferred example, the word vector is selected from the following group: one-hot encoding, or word2vec.

[0074] In another preferred example, the word vector is one-hot encoding.

[0075] In another preferred example, the sequence alignment algorithm is a protein substitution scoring matrix.

[0076] In another preferred example, the protein substitution scoring matrix is selected from the following group: PAM matrix, or BLOSUM matrix.

[0077] In another preferred example, the protein substitution scoring matrix is the BLOSUM matrix.

[0078] In another preferred example, the numerical form includes: embedding vectors, scoring matrices.

[0079] In another preferred example, the numerical form is an embedding vector.

[0080] In another preferred example, the embedding vector is an ESM2 embedding.

[0081] In another preferred example, the ESM2 embedding is an embedding vector formed by using the ESM2 model to convert the candidate antimicrobial peptide sequence and the target functional sample.

[0082] In another preferred example, the deep learning model includes: a model based on a convolutional neural network (CNN), a model based on a recurrent neural network (RNN).

[0083] In another preferred example, the deep learning model is a model based on a recurrent neural network (RNN).

[0084] In another preferred example, the scoring system employs an algorithm selected from the group consisting of: supervised learning algorithms, reinforcement learning algorithms, unsupervised learning algorithms, or semi-supervised learning algorithms.

[0085] In another preferred example, the scoring system employs an algorithm selected from the group consisting of: supervised learning algorithms, or reinforcement learning algorithms.

[0086] In another preferred example, the supervised learning algorithm is a classification algorithm.

[0087] In another preferred example, the classification algorithm is selected from the group consisting of: linear models, K-nearest neighbors (KNN), random forests, support vector machines (SVM), decision trees, neural networks, naive Bayes, Boosting, or combinations thereof.

[0088] In another preferred example, the classification algorithm is selected from the group consisting of: linear models, K-nearest neighbors, random forests, or combinations thereof.

[0089] In another preferred example, the classification algorithm is K-nearest neighbors.

[0090] In another preferred example, the calculation for evaluating sequence similarity by the K-nearest neighbors is as follows:

[0091]

[0092] where s_f is the similarity score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric, x is the generated sample, y_i is the center of the i-th functional cluster, k is the number of clusters, and d is the number of distance metrics.

[0093] In another preferred example, the distance metric is selected from the group consisting of: Euclidean distance, Manhattan distance, cosine similarity, Chebyshev distance, Minkowski distance, standard Euclidean distance, Mahalanobis distance, Hamming distance, Jaccard distance, correlation distance, information entropy, or combinations thereof.

[0094] In another preferred example, the distance metrics are Euclidean distance, Manhattan distance, and cosine similarity.

[0095] In another preferred example, d = 3.

[0096] In another preferred example, the scoring system includes obtaining multiple similarity scores by using multiple of the supervised learning algorithms.

[0097] In another preferred example, the multiple similarity scores are combined by a method selected from the group consisting of: fuzzy logic, Dempster-Shafer theory, or multi-objective optimization algorithms.

[0098] In another preferred example, the multi-objective optimization algorithm is NSGA-II.

[0099] In another preferred example, the unsupervised learning algorithm is a clustering algorithm.

[0100] In another preferred example, the clustering algorithm is selected from the group consisting of: density clustering, K-means clustering, spectral clustering, hierarchical clustering, grid clustering, or model clustering.

[0101] In another preferred example, the density clustering is DBSCAN.

[0102] In another preferred example, a computational model is used to evaluate the stability of antimicrobial peptides.

[0103] In another preferred example, the computational model includes a computational model for simulating protein degradation.

[0104] In another preferred example, a graph neural network is used to evaluate structural similarity.

[0105] In another preferred example, the antimicrobial activity indicators include: minimum inhibitory concentration, minimum bactericidal concentration, time-kill curve.

[0106] In another preferred example, the antimicrobial activity indicator is the minimum inhibitory concentration.

[0107] In another preferred example, the minimum inhibitory concentration is predicted using a method selected from the group consisting of: supervised learning algorithms, expert systems, or deep learning models.

[0108] In another preferred example, the supervised learning algorithm is selected from the group consisting of: classification algorithms, or regression algorithms.

[0109] In another preferred example, the classification algorithm is Naive Bayes.

[0110] In another preferred example, the Naive Bayes-based classifier is trained on a known dataset so that the classifier reaches the training objective, obtaining a pre-trained classifier, and thus using the pre-trained classifier to predict the minimum inhibitory concentration.

[0111] In another preferred example, the known dataset is an antimicrobial peptide dataset with experimentally determined minimum inhibitory concentration values.

[0112] In another preferred example, the Naive Bayes-based classifier is a Bayesian classifier.

[0113] In another preferred example, the Bayesian classifier uses the ESM2 embedding.

[0114] In another preferred example, the training objective is to minimize the L1 loss between the predicted minimum inhibitory concentration value and the actual minimum inhibitory concentration value, where the calculation of the L1 loss is as follows:

[0115]

[0116] where y_i is the actual minimum inhibitory concentration value, f(x_i) is the predicted minimum inhibitory concentration value of the numerical form x_i of the i-th peptide segment, and N is the number of training samples.

[0117] In another preferred example, the regression algorithm is selected from the following group: linear regression, K-nearest neighbor regression, support vector machine regression, decision tree regression, neural network regression, naive Bayes regression, Boosting regression, random forest regression, deep forest regression, or extremely randomized tree regression.

[0118] In another preferred example, the Boosting regression is gradient boosting decision tree regression.

[0119] In another preferred example, the expert system includes a rule-based expert system.

[0120] In another preferred example, the deep learning model is an end-to-end deep learning model.

[0121] In another preferred example, the end-to-end deep learning model directly identifies the minimum inhibitory concentration from the antimicrobial peptide candidate sequence.

[0122] In another preferred example, the posterior verification screens the antimicrobial peptide candidate sequences by methods selected from the following group:

[0123] (C1) Limiting the peptide segment length;

[0124] (C2) Predicting structural features; and

[0125] (C3) Analyzing similarity.

[0126] In another preferred example, the peptide segment length is ≤50 amino acids, preferably ≤30 amino acids, more preferably ≤25 amino acids.

[0127] In another preferred example, the structural features are foldability and thermal stability.

[0128] In another preferred example, the structural features are predicted by methods selected from the following group: protein structure prediction tools, or molecular simulation.

[0129] In another preferred example, the protein structure prediction tool is selected from the group consisting of: AlphaFold2, ESMFold, I-TASSER, RoseTTAFold, GalaxyTBM, SWISS-MODEL, or a combination thereof.

[0130] In another preferred example, the protein structure prediction tool is a combination of AlphaFold2 and ESMFold.

[0131] In another preferred example, the molecular simulation is molecular dynamics simulation.

[0132] In another preferred example, the similarity includes: sequence similarity, structural similarity.

[0133] In another preferred example, the similarity is analyzed by a method selected from the group consisting of: sequence alignment algorithms, or structural alignment algorithms.

[0134] In another preferred example, the sequence similarity is analyzed by the sequence alignment algorithm.

[0135] In another preferred example, the structural similarity is analyzed by the structural alignment algorithm.

[0136] In another preferred example, the sequence alignment algorithm is selected from the group consisting of: BLAST, Smith-Waterman algorithm, Needleman-Wunsch algorithm, or a combination thereof.

[0137] In another preferred example, the sequence alignment algorithm is BLAST.

[0138] In another preferred example, the structural alignment algorithm is selected from the group consisting of: FoldSeek, TM-align, DALI, SSAP, FLEXPROT, or a combination thereof.

[0139] In another preferred example, the structural alignment algorithm is a combination of FoldSeek and TM-align.

[0140] In another preferred example, the posterior verification further includes screening the antimicrobial peptide candidate sequences by a method selected from the group consisting of: integrating multi-omics data, or knowledge graph technology.

[0141] In another preferred example, the multi-omics data is selected from the group consisting of: proteomics, metabolomics, or a combination thereof.

[0142] In another preferred example, the multi-omics data is proteomics.

[0143] In another preferred example, the method further includes: using an evaluation model to comprehensively screen antimicrobial peptide candidate sequences based on multiple indicators.

[0144] In another preferred example, the evaluation model is selected from the following group: AHP, TOPSIS, grey relational analysis, fuzzy comprehensive evaluation, entropy weight method, or a combination thereof.

[0145] In another preferred example, the evaluation model is AHP.

[0146] In another preferred example, the indicators are selected from the following group: sequence similarity, structural similarity, structural features, minimum inhibitory concentration, or a combination thereof.

[0147] In a fourth aspect of the present invention, there is provided an antimicrobial peptide screening framework or system, which comprises:

[0148] (D1) An input unit configured to input data, the input data including an input sequence to be screened, and the input sequence to be screened including antimicrobial peptide candidate sequences generated by the antimicrobial peptide generation framework or system according to the second aspect of the present invention;

[0149] (D2) A screening unit for antimicrobial peptides, configured to execute a screening model for antimicrobial peptides, so as to obtain antimicrobial peptide target sequences from the antimicrobial peptide candidate sequences; wherein, the screening model is constructed by the method according to the third aspect of the present invention;

[0150] (D3) An output unit configured to output a screening result.

[0151] In another preferred example, the antimicrobial peptide generation framework or system is a fine-tuned antimicrobial peptide generation framework or system.

[0152] In another preferred example, the fine-tuned antimicrobial peptide generation framework or system is AMPGen.

[0153] In another preferred example, the screening model is AMPGen-Filtering.

[0154] In a fifth aspect of the present invention, there is provided an antimicrobial peptide design and screening framework or system, which comprises:

[0155] (W) An input unit configured to input data, the input data including a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence;

[0156] (X) A fine-tuning unit configured to fine-tune a model, and the fine-tuning model performs a predetermined fine-tuning on the input data to obtain a fine-tuning result; wherein, the fine-tuning model is a supervised fine-tuning model, and the supervised fine-tuning model includes the steps of: performing supervised fine-tuning on the input data by using a fine-tuning method to obtain fine-tuned input data;

[0157] (Y) Screening unit, the screening unit is configured to execute a screening model for antimicrobial peptides, and the screening model obtains an antimicrobial peptide target sequence from the fine-tuned input data; wherein, the screening model is constructed by the method described in the third aspect of the present invention;

[0158] (Z) Output unit, the output unit is configured to output the result of the screening module.

[0159] In another preferred example, the fine-tuned input data includes: a fine-tuned input sequence and / or a fine-tuned protein base model.

[0160] In another preferred example, the fine-tuned input sequence includes an amino acid sequence shown in any one of SEQ ID NO: 1-4.

[0161] In another preferred example, the antimicrobial peptide target sequence includes an amino acid sequence shown in any one of SEQ ID NO: 1-4.

[0162] In another preferred example, the antimicrobial peptide target sequence has an amino acid sequence shown in any one of SEQ ID NO: 1-4.

[0163] In another preferred example, the molar mass range of the antimicrobial peptide target sequence is 1400-2800 g / mol.

[0164] In another preferred example, the MIC of the antimicrobial peptide target sequence is < 35 μg / mL, preferably < 20 μg / mL, more preferably < 8 μg / mL.

[0165] In the sixth aspect of the present invention, there is provided a drug or a pharmaceutical composition, the drug or the pharmaceutical composition includes an antimicrobial peptide sequence, and the antimicrobial peptide sequence is obtained by the framework or system described in the fifth aspect of the present invention.

[0166] In another preferred example, the drug or the pharmaceutical composition includes an amino acid sequence shown in any one of SEQ ID NO: 1-4.

[0167] In another preferred example, the drug or the pharmaceutical composition further includes: the protein family to which the antimicrobial peptide sequence belongs, and antimicrobial peptide sequences in the protein family to which the antimicrobial peptide sequence belongs that have a similar structure or function to the antimicrobial peptide sequence.

[0168] In another preferred embodiment, the drug or pharmaceutical composition further comprises: a protein family to which the amino acid sequence shown in any one of SEQ ID NO: 1-4 belongs, and an antimicrobial peptide sequence having a similar structure or function to the amino acid sequence shown in any one of SEQ ID NO: 1-4 in the protein family to which the amino acid sequence shown in any one of SEQ ID NO: 1-4 belongs.

[0169] In another preferred embodiment, the drug or pharmaceutical composition further comprises a pharmaceutically acceptable carrier, diluent or excipient.

[0170] In another preferred embodiment, the antimicrobial peptide sequence is prepared by direct synthesis.

[0171] In the seventh aspect of the present invention, there is provided a use of the drug or pharmaceutical composition according to the sixth aspect of the present invention, and the use includes:

[0172] (E1) Preparing a preparation for inhibiting the growth of microorganisms;

[0173] (E2) Preparing a kit, and the kit further comprises a label or an instruction manual, and the instruction manual indicates that the kit is used for inhibiting the growth of microorganisms.

[0174] In another preferred embodiment, the microorganisms include: bacteria, fungi.

[0175] In another preferred embodiment, the bacteria include: Pseudomonas aeruginosa, Escherichia coli, Staphylococcus aureus.

[0176] In another preferred embodiment, the fungi include: Saccharomyces cerevisiae, Candida albicans, Fusarium graminearum.

[0177] In another preferred embodiment, the preparation includes: a liquid preparation.

[0178] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features specifically described below (such as in the examples) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be repeated one by one here. BRIEF DESCRIPTION OF THE DRAWINGS

[0179] Figure 1 Shows the process of the antimicrobial peptide sequence design framework based on supervised fine-tuning.

[0180] Figure 2 Shows the process of the antimicrobial peptide screening framework based on protein language models and bioinformatics methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0181] After extensive and in-depth research, the present inventor has provided for the first time an antimicrobial peptide design framework based on supervised fine-tuning of protein language models. Specifically, the present invention uses protein language models through supervised fine-tuning technology to develop an antimicrobial peptide generation framework, and designs an antimicrobial peptide screening framework based on various properties of antimicrobial peptides. The two frameworks together constitute a novel antimicrobial peptide design framework, effectively solving the problems of low design efficiency and insufficient accuracy in the prior art. In addition, while generating diverse and novel antimicrobial peptide sequences, maintaining their key antimicrobial activities significantly improves the practical application potential of the generated antimicrobial peptide sequences, providing an efficient, accurate, and flexible new method for the research and development of antimicrobial peptides. Based on the above antimicrobial peptide design and screening models, the present invention has also discovered 4 novel highly active antimicrobial peptide sequences. On this basis, the present invention has been completed.

[0182] The following explains some of the innovative points of the present invention:

[0183] First, the present invention uses the antimicrobial peptide data in the public dataset as input data, inputs it into the protein base model, and obtains the training model AMPGen through supervised fine-tuning. Among them, ProGen2 is used as the base model, low-rank adaptation technology is used for supervised fine-tuning, and perplexity is used to guide model fine-tuning. Low-rank adaptation can minimize the number of parameters and reduce the risk of overfitting; perplexity can evaluate the quality of model fine-tuning results and guide model optimization. In the supervised fine-tuning using low-rank adaptation technology, the loss function of supervised fine-tuning is defined as follows:

[0184]

[0185] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the LoRA parameters.

[0186] After the above model training is completed, the present invention uses the screening framework to screen antimicrobial peptide candidate samples. The screening framework is AMPGen-Filtering, which includes two-stage screening. In the first stage, first use the pre-trained protein language model ESM2 to evaluate the similarity between the sequence and known antimicrobial peptides in a zero-shot manner, and use a MIC-based classifier to predict the antimicrobial activity of the sequence. Secondly, use the K-nearest neighbor algorithm to screen the generated antimicrobial peptide samples, calculate the screening score, and the calculation formula of the screening score is:

[0187]

[0188] Among them, s_f is the final similarity screening score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric (Euclidean distance, Manhattan distance, and cosine similarity), x is the generated sample, y_i is the center of the i-th functional cluster, and k is the number of clusters.

[0189] After that, a Bayesian classifier based on ESM2 embeddings is used to further refine the selection process. This classifier is trained on a known AMP dataset with MIC values, and the goal is to minimize the L1 loss between the predicted MIC value and the actual MIC value. The formula for calculating the loss value L1 is:

[0190]

[0191] Among them, y_i is the true MIC value, f(x_i) is the predicted MIC value of the ESM2 embedding x_i of the i-th peptide, and N is the number of training samples.

[0192] Finally, the sequences obtained by MIC prediction and K-nearest neighbor algorithm screening are cross-validated to further refine the selection of candidate peptides.

[0193] In the second stage, the antimicrobial peptide candidates are optimized through length screening, structure prediction, and similarity analysis. Among them, AlphaFold2 and ESMFold are used to evaluate the sequence structure characteristics in structure prediction; BLAST and FoldSeek are used for sequence similarity search in similarity analysis.

[0194] It should be understood that the specific methods and experimental conditions of the present invention are described in various levels of detail below to provide an understanding of the essence of the present invention. Definitions of some terms used in this specification are provided below. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention belongs.

[0195] The term

[0196] As used herein, the terms "comprising", "including", and "containing" can be used interchangeably and include not only closed definitions but also semi-closed and open definitions. In other words, the said terms include "consisting of" and "consisting essentially of".

[0197] As used herein, the terms "baseline generator", "base model", "protein baseline generator", and "protein base model" can be used interchangeably and are all Baseline generator, which refers to an original machine learning model that can be used for protein research without training or fine-tuning and can have good target protein generation ability after training and fine-tuning.

[0198] As used herein, the term "Perplexity" is an evaluation metric for the performance of a language model, used to measure the prediction accuracy of the model for a given sequence. The lower the perplexity, the better the model performance.

[0199] As used herein, the terms "physicochemical property" and "physicochemical characteristic" can be used interchangeably and are quantitative indicators describing molecular characteristics. In the present invention, the physicochemical properties of the antimicrobial peptide include hydrophobicity, hydrophobic moment, charge, isoelectric point, etc.

[0200] As used herein, the terms "Zero-shot Learning" and "zero-shot manner" can be used interchangeably, referring to the ability of a model to recognize or classify categories not seen during the training process.

[0201] As used herein, the term "Protein Family" is a group of evolutionarily related proteins with similar sequences, structures, or functions.

[0202] As used herein, the terms "Sequence Embedding", "embedding vector", and "embedding" can be used interchangeably, referring to the conversion of a protein sequence into a vector representation of a fixed dimension to capture the semantic information of the sequence. In the present invention, generating a sample embedding refers to the embedding vector transformed from the candidate antimicrobial peptide sequence; the target functional sample embedding refers to the embedding vector transformed from the antimicrobial peptide sequence with known function and sequence.

[0203] As used herein, the term "Molecular Dynamics Simulation" is a computational method for simulating the evolution of a molecular system over time by computer, used to study the motion and interactions of molecules.

[0204] As used herein, the terms "antimicrobial peptide MIC value predictor", "minimum inhibitory concentration predictor", and "MIC identifier" can be used interchangeably, referring to a classifier that can predict the MIC obtained by training on a known data set containing MIC data during the activity-based feedback fine-tuning stage.

[0205] As used herein, the terms "TM-Score (Template Modeling Score)" and "TM score" can be used interchangeably, which is a score used to evaluate the similarity of protein structures, ranging from 0 to 1, and the closer to 1, the more similar the structures.

[0206] As used herein, the term "RMSD (Root Mean Square Deviation)" refers to the root mean square deviation, which is used to measure the average distance of the atomic spatial positions between two superimposed protein structures, with the unit of angstrom.

[0207] As used herein, the term "Multi-objective Optimization Algorithm" refers to an algorithm that aims to optimize multiple objective functions simultaneously, such as NSGA-II (Non-dominated Sorting Genetic Algorithm II).

[0208] As used herein, the term "K-Nearest Neighbors (KNN)" is a machine learning algorithm for classification and regression, which makes predictions based on the K nearest neighbors around the sample.

[0209] As used herein, the term "Bayesian Classifier" refers to a probability classifier based on Bayes' theorem, which can make classification predictions according to the conditional probabilities of features.

[0210] As used herein, the term "BLAST" is the Basic Local Alignment Search Tool, which is a sequence similarity search algorithm used to compare biological sequences (such as DNA, RNA, or protein sequences) with a sequence database to identify similar sequences in the database.

[0211] As used herein, the terms "antimicrobial peptide target sequence" and "antimicrobial peptide target sample" can be used interchangeably, and refer to the antimicrobial peptide sequence obtained after screening the antimicrobial peptide sequence to be screened through the screening framework or system of the present invention, which has properties such as good antimicrobial activity.

[0212] Language Model and Protein Language Model

[0213] As used herein, the term "language model" is a type of machine learning model in the fields of natural language processing and deep learning, which includes large language models (LLMs), etc. The goal of a language model is to predict the next possible characters based on the given context. The technical bases used in language models include: rule-based methods, statistic-based methods, neural network-based methods, Transformer-based methods, and large-scale pre-trained model-based methods.

[0214] As used herein, the term "Protein Language Models (PLMs)" refers to language models specifically designed to process or generate protein sequences. These models learn the underlying patterns and regularities of protein sequences through pre-training on a large amount of protein sequence data, enabling predictions of properties such as the nature and structure of unknown proteins.

[0215] As used herein, the terms "ESM2", "ESM-1b", "ProtBERT", "ProtXLNet", "TAPE", "ProGen2", and "AlphaFold2" are several language models commonly used in protein research.

[0216] Both ESM2 and ESM-1b belong to the ESM large biological models. Among them, ESM2 is a large protein language model based on the Transformer framework developed by Meta AI, which has been pre-trained on hundreds of millions of protein sequences. It consists of multiple layers of self-attention mechanisms and feed-forward neural networks. The self-attention mechanism allows the model to consider the relationships between different positions in the sequence when processing the sequence, while the feed-forward neural network further processes and integrates the above information. The term "ESMFold" is a model for protein structure prediction using the ESM2 model. It can perform end-to-end three-dimensional structure prediction using only a single sequence as input by leveraging the information and representations learned by ESM2 to quantify the occurrence of protein structures. The input of ESM2 is an amino acid sequence, which is converted into numerical vectors and then input into the model for learning and prediction. The output of ESM2 is the three-dimensional structure prediction of the protein, usually represented in the form of atomic coordinates, which describe the spatial positions of each atom in the protein molecule and can be used for further biophysical analysis and molecular simulations. ESM-1b is also a large protein language model based on the Transformer framework, which contains multiple attention layers and is trained using a self-supervised learning method through a masked language model. The input of ESM-1b is an amino acid sequence, and the output is the feature representation of each amino acid position for downstream analysis.

[0217] ProtBERT and ProtXLNet are natural language processing models trained on protein sequences. Among them, ProtBERT uses a self-supervised learning method, combining protein structures with annotations of the Gene Ontology, and can make predictions on protein structures, post-translational modifications, and biophysical properties. ProtXLNet, on the other hand, uses an autoregressive model to predict proteins.

[0218] ProGen2 is a GPT-like protein language model developed by Salesforce Research. This model is trained on a diverse sequence dataset of over one billion proteins extracted from genomic, metagenomic, and immunoglobulin databases, learning the evolutionary distribution of protein sequences, generating new viable sequences, and predicting protein fitness. Its autoregressive mode can enhance the diversity and novelty of protein sequence generation. The term "GPT-like" refers to a neural network-based autoregressive language model that uses a novel sequence-to-sequence model capable of avoiding the vanishing gradient problem present in traditional recurrent neural networks when processing long sequence data. ProGen2 contains an attention module, which is a mechanism in deep learning for locating key tokens.

[0219] AlphaFold2 is a protein structure prediction algorithm developed by DeepMind, capable of predicting the three-dimensional structure of proteins with extremely high accuracy.

[0220] Supervised Fine-tuning

[0221] In the present invention, a fine-tuned protein base model and antimicrobial peptide candidate sequences are obtained by performing supervised fine-tuning on the protein base model.

[0222] As used herein, the terms "Supervised Fine-tuning (SFT)" and "supervised fine-tuning" are used interchangeably and refer to the process of fine-tuning a pre-trained model using labeled data on the basis of the pre-trained model, enabling the model to adapt to a specific task or domain. Generally, supervised fine-tuning includes the following steps: pre-training, data collection and annotation, supervised fine-tuning, evaluation, and optimization.

[0223] In the present invention, the methods for performing the supervised fine-tuning include: Low-Rank Adaptation (LoRA), Prefix-tuning, Prompt-tuning, P-tuning, P-tuning v1, P-tuning v2, Adapter-tuning.

[0224] Preferably, low-rank adaptation is used for supervised fine-tuning. Low-rank adaptation is a parameter-efficient model fine-tuning technique that achieves model adaptation by adding a low-rank matrix to the weight matrix of the pre-trained model, significantly reducing the number of parameters that need to be updated, which can significantly reduce the debugging cost and can reduce the risk of overfitting.

[0225] Deep learning model

[0226] As used herein, the term "deep learning model" belongs to machine learning models that use multi-layer neural networks to learn from large amounts of data.

[0227] Common deep learning models include supervised neural networks such as Recurrent Neural Networks (RNN), Convolutional Neural Networks (CNN), deep neural networks, recurrent neural networks, etc., and unsupervised or semi-supervised deep learning models such as deep generative models, autoencoders, etc. Among them, deep generative models include Generative Adversarial Network (GAN), and autoencoders include Variational Autoencoder (VAE).

[0228] Deep learning models such as RNN and CNN have been widely used in natural language processing and biological sequence analysis. Among them, RNN is a neural network for processing sequence data that can utilize the temporal or spatial dependence of the sequence; CNN is suitable for processing data with a grid topology structure, such as images or sequence data.

[0229] Deep learning models such as variational autoencoders are often used in data compression and generation tasks. The generative adversarial network consists of two networks, a generative model and a discriminative model, and generates realistic samples through adversarial learning.

[0230] The term "end-to-end deep learning model" refers to a type of deep learning model that uses an end-to-end approach. Among them, "end-to-end" is a data transmission method, which means that data is directly transmitted from the sender to the receiver without the need for an intermediate environment to parse and process the data content, ensuring the directness and integrity of the data. The end-to-end deep learning model can be an end-to-end RNN, an end-to-end CNN, or other deep learning models.

[0231] In the present invention, the end-to-end deep learning model is trained on a data set containing antimicrobial peptides and their minimum inhibitory concentrations, so that the corresponding minimum inhibitory concentration can be directly obtained from the antimicrobial peptide candidate sequence.

[0232] In the present invention, RNN and CNN can be used to replace the ESM2 model for machine learning screening of antimicrobial peptides.

[0233] The scoring system of the present invention

[0234] In the present invention, the scoring system is selected from the following group: supervised learning algorithms, reinforcement learning algorithms, unsupervised learning algorithms, or semi-supervised learning algorithms.

[0235] As used herein, the term "supervised learning algorithm" refers to a class of algorithms used in the process of supervised learning, where supervised learning means that by providing input data and its corresponding labeled data to a model, the model is trained to accurately find the optimal mapping relationship between the input data and the labeled data, so as to predict or classify new unlabeled data. Supervised learning algorithms mainly include classification algorithms and regression algorithms. Classification algorithms are mainly used to output discrete data, while regression algorithms are used to output continuous data.

[0236] As used herein, the term "unsupervised learning algorithm" refers to a class of algorithms used in the process of unsupervised learning, where unsupervised learning means classifying unlabeled input data. Unsupervised learning algorithms include clustering algorithms.

[0237] In the present invention, a scoring system is used to evaluate sequence similarity. Preferably, the scoring system employs a supervised learning algorithm or an unsupervised learning algorithm. Preferably, the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the group consisting of: linear models, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), decision trees, neural networks, Naive Bayes, Boosting, Random Forest, or combinations thereof. The unsupervised learning algorithm is selected from the group consisting of: spectral clustering, hierarchical clustering, density clustering, K-means clustering, grid clustering, or model clustering. Among them, SVM classifies data points by finding the optimal hyperplane, while Random Forest classifies by constructing multiple decision trees and taking the majority vote.

[0238] Preferably, the scoring system employs the K-Nearest Neighbors algorithm. The K-Nearest Neighbors algorithm makes predictions based on the K nearest neighbors around a sample. In this process, a distance metric is needed to calculate the distance between the sample and its neighbors. The terms "distance metric" and "similarity metric" can be used interchangeably because the commonly used method for evaluating the similarity between samples is to calculate the distance.

[0239] Preferably, the calculation for evaluating sequence similarity through the K-Nearest Neighbors is as follows:

[0240]

[0241] where s_f is the similarity score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric, x is the generated sample, y_i is the center of the i-th functional cluster, and k is the number of clusters.

[0242] Preferably, the distance metrics are Euclidean distance, Manhattan distance, and cosine similarity.

[0243] In the present invention, preferably, a supervised learning algorithm is used to predict the minimum inhibitory concentration. The supervised learning algorithm is a regression algorithm, and the regression algorithm is selected from the following group: Naive Bayes, linear regression, K-nearest neighbor regression, support vector machine regression, decision tree regression, neural network regression, Naive Bayes regression, gradient boosting decision tree regression, random forest regression, deep forest regression, or extremely randomized tree regression.

[0244] Preferably, the Naive Bayes algorithm is used to construct a Bayesian Classifier to predict the minimum inhibitory concentration. The Bayesian Classifier refers to a probability classifier based on Bayes' theorem, which can perform classification prediction according to the conditional probability of features.

[0245] Antibacterial activity index

[0246] The present invention uses antibacterial activity indexes to characterize the antibacterial activity of antibacterial peptides. By predicting the antibacterial activity of antibacterial peptide candidate sequences, target sequences of antibacterial peptides with good antibacterial activity can be screened out.

[0247] As used herein, the term "Minimum Inhibitory Concentration (MIC)" refers to the lowest concentration of an antibacterial substance that can inhibit the visible growth of microorganisms and is an important indicator for evaluating antibacterial efficacy.

[0248] As used herein, the term "Minimum Bactericidal Concentration (MBC)" refers to the lowest concentration required to kill microorganisms under specific conditions and a fixed extended time (18 to 24 hours), at which the viability of the microorganisms is reduced by 99%.

[0249] As used herein, the term "Time-kill Curve" refers to a curve that describes the change in the killing effect of an antibacterial agent on microorganisms over time and is used to evaluate the bactericidal kinetics of the antibacterial agent.

[0250] Antibacterial peptide generation framework or system of the present invention and its construction method

[0251] The present invention provides an antibacterial peptide generation framework or system and its construction method for generating candidate antibacterial peptide sequences.

[0252] Specifically, the method includes the steps of: (1) training a protein base model on a data set; (2) performing supervised fine-tuning on the protein base model to obtain a fine-tuned protein base model and antibacterial peptide candidate sequences.

[0253] Preferably, step (2) further includes the steps of: performing reinforcement learning on the fine-tuned protein base model to obtain an optimized protein model and candidate antimicrobial peptide sequences.

[0254] The framework or system includes: (D1) an input unit configured to input data, the input data including a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence; (D2) a fine-tuning unit configured to perform a fine-tuning model on the input data to obtain fine-tuned input data; wherein the fine-tuning model is a supervised fine-tuning model; (D3) an output unit configured to output the result of the fine-tuning unit. This framework or system is "AMPGen".

[0255] In the method for constructing the antimicrobial peptide generation framework or system of the present invention, the dataset for training can be adjusted according to the type of bioactive peptides (antimicrobial peptides are a type of bioactive peptides), so as to apply the method to construct more types of bioactive peptide generation frameworks or systems, such as anticancer peptides, antiviral peptides, and cell-penetrating peptides. For example, generating new anticancer peptides based on an anticancer peptide dataset.

[0256] The antimicrobial peptide screening framework or system of the present invention and its construction method

[0257] The present invention provides an antimicrobial peptide screening framework or system and its construction method for screening the generated candidate antimicrobial peptide sequences to obtain antimicrobial peptides with good antimicrobial activity.

[0258] Specifically, the method includes the steps:

[0259] (M1) Machine learning screening, which includes the steps of: (B1) using a machine learning model to evaluate the sequence similarity between candidate antimicrobial peptide sequences and target function samples through a scoring system; (B2) predicting antimicrobial activity indicators.

[0260] (M2) Posterior verification, which includes (C1) restricting the peptide segment length; (C2) predicting structural features; (C3) analyzing similarity; or a combination thereof. Preferably, the posterior verification includes restricting the peptide segment length, predicting structural features, and analyzing similarity.

[0261] The described framework or system includes: (D1) an input unit configured to input data, the input data including an input sequence to be screened, the input sequence to be screened including an antimicrobial peptide candidate sequence generated using the antimicrobial peptide generation framework or system according to the second aspect of the present invention; (D2) a screening unit for antimicrobial peptides, the screening unit being configured to execute a screening model for antimicrobial peptides so as to obtain an antimicrobial peptide target sequence from the antimicrobial peptide candidate sequences; wherein, the screening model is constructed using the method according to the third aspect of the present invention; (D3) an output unit configured to output a screening result. This framework or system is the "AMPGen-Filtering".

[0262] Preferably, the input sequence to be screened is an antimicrobial peptide candidate sequence generated by AMPGen.

[0263] In the method for constructing the antimicrobial peptide screening framework or system of the present invention, the screening indicators can be adjusted according to the type and properties of bioactive peptides (antimicrobial peptides are a type of bioactive peptide), so as to apply the method to construct screening frameworks or systems for more types of bioactive peptides, such as anticancer peptides, antiviral peptides, and cell-penetrating peptides. For example, for anticancer peptides or antiviral peptides, screening can be carried out according to the killing power against tumor cells or viruses.

[0264] The main advantages of the present invention include:

[0265] (1) According to the steps in Invention Content 1, the present invention effectively reduces data-driven bias and captures a wider range of protein sequence knowledge by using a large protein language model and fine-tuning techniques.

[0266] (2) According to the steps in Invention Content 3, the present invention proposes a comprehensive screening and analysis method that effectively balances the diversity, novelty, and key antibacterial activity of the generated antimicrobial peptide sequences, addressing the deficiencies in existing methods.

[0267] (3) According to the steps in Invention Content 3, the comprehensive evaluation system established by the present invention provides a comprehensive quality control and screening mechanism for the generated antimicrobial peptide candidates, improving the accuracy and efficiency of screening.

[0268] (4) The present invention significantly improves the practical application potential of the generated antimicrobial peptide sequences by integrating multiple computational methods and experimental verification, narrowing the gap between in vitro prediction and in vivo effects.

[0269] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The experimental methods without specific conditions noted in the following embodiments are generally carried out under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise specified, percentages and parts are by weight percentage and weight parts.

[0270] Methods and steps:

[0271] 1. Supervised Fine-tuning (SFT) of Protein Language Model

[0272] The present invention uses ProGen2 as the baseline generator (base model) and performs supervised fine-tuning using Low-Rank Adaptation (LoRA) technology. As an efficient adapter-based method, Low-Rank Adaptation can minimize the number of parameters and reduce the risk of overfitting. In the present invention, rank-based adapters are introduced into the attention module of ProGen2 to regulate residue interactions for conditional generation. Perplexity is used to guide the fine-tuning of the model to optimize the generation of peptides similar in function to the training data.

[0273] The loss function for supervised fine-tuning is defined as follows:

[0274]

[0275] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the LoRA parameters.

[0276] 2. Screening and Validation of Antimicrobial Peptide Candidate Samples

[0277] The present invention adopts a two-stage screening process to ensure the quality of the generated candidate peptide segments.

[0278] 2.1 Machine Learning-based Basic Screening

[0279] The specific implementation steps of machine learning-based basic screening are as follows:

[0280] (S1) Use the pre-trained protein language model ESM2 to evaluate the similarity between the sequence and known antimicrobial peptides in a zero-shot manner.

[0281] (S2) Predict the antimicrobial activity using a classifier based on the minimum inhibitory concentration.

[0282] (S3) Implement the K-nearest neighbor algorithm to screen the generated antimicrobial peptide samples, and use the distance between ESM2 protein sequence embeddings to evaluate similarity.

[0283] (S4) Calculate the screening score. The screening criterion is based on the weighted sum of multiple normalized distance metrics between the generated sample embedding and the center of the target functional sample embedding. The specific calculation formula is as follows:

[0284]

[0285] where \(s_f\) is the final similarity screening score, \(w_m\) is the weight of each functional cluster, \(f_m\) represents the \(m\)-th distance metric (Euclidean distance, Manhattan distance, and cosine similarity), \(x\) is the generated sample, \(y_i\) is the center of the \(i\)-th functional cluster, and \(k\) is the number of clusters.

[0286] (S5) Use a Bayesian classifier based on ESM2 embeddings to further refine the selection process. The classifier is trained on an AMPs dataset with experimentally determined minimum inhibitory concentration (MIC) values. The training objective is to minimize the L1 loss between the predicted MIC value and the actual MIC value:

[0287]

[0288] where \(y_i\) is the actual MIC value, \(f(x_i)\) is the predicted MIC value of the ESM2 embedding \(x_i\) of the \(i\)-th peptide segment, and \(N\) is the number of training samples.

[0289] (S6) Further refine the selection of candidate peptides through a cross-validation method. Take the intersection of the two sets of KNN-based screening and MIC prediction to identify peptides that are not only structurally similar to known AMPs but also likely to have good antibacterial activity.

[0290] 2.2 Posterior verification

[0291] The post-processing steps of the present invention optimize antimicrobial peptide candidates through length screening, structure prediction, and similarity analysis. The specific implementation steps are as follows:

[0292] (S1) Length selection and folding stability

[0293] (a) Limit the peptide segment length to 50 amino acids or less to ensure compatibility with efficient laboratory synthesis techniques.

[0294] (b) Use protein structure prediction tools AlphaFold2 and ESMFold to evaluate the structural characteristics of the selected peptide segments, providing key feature predictions regarding the foldability and potential thermal stability of the peptide segments.

[0295] (S2) Similarity analysis

[0296] (a) Perform sequence-level comparison using BLAST, and align against the curated peptide and protein domain databases. In the BLAST analysis and evaluation, the following indicators are mainly concerned: the highest score, the total score, the query coverage rate, the E-value, the percentage identity, and the accepted length. Priority is given to the E-value and the percentage identity to determine sequence similarity. Candidates with an E-value lower than 1e-5 and a percentage identity higher than 30% are marked for further investigation.

[0297] (b) Perform structure-based similarity search using FoldSeek. In the FoldSeek analysis and evaluation, the following indicators are mainly concerned: the inspection probability, the sequence identity, the E-value, the score, the query position, the target position, the TM score, and the RMSD. Special emphasis is placed on the TM score and the RMSD for structure similarity evaluation. Structures with a TM score higher than 0.5 and an RMSD lower than are considered to have significant structural similarity.

[0298] Through the above steps, the present invention provides a comprehensive method for screening and evaluating antibacterial peptide candidates, balancing synthetic feasibility and potential structural stability, and at the same time providing a framework for explaining experimental results.

[0299] 3. Detection of antibacterial peptide properties

[0300] 3.1 Determination of molar mass

[0301] Use mass spectrometry analysis techniques, such as matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS) or liquid chromatography-mass spectrometry (LC-MS), which can provide accurate molecular weight information of antibacterial peptides. In addition, nuclear magnetic resonance (NMR) can also be used for structure analysis and molecular weight confirmation.

[0302] 3.2 Determination of bacteriostatic rate

[0303] Spectrophotometry (OD600 measurement) and plate counting method are used for evaluation. First, three kinds of bacteria, Pseudomonas aeruginosa, Escherichia coli, and Staphylococcus aureus, are cultured. Bacteria are inoculated in LB (Luria-Bertani) liquid medium and cultured at 37°C with shaking at 200 rpm for 12 - 16 hours until the bacterial solution reaches the logarithmic growth phase (OD600≈0.5 - 0.6), and then diluted to OD600≈0.05 with sterile PBS or LB as the experimental bacterial suspension.

[0304] Afterwards, for yeasts and fungi (Saccharomyces cerevisiae, Candida albicans and other fungi), they were cultured in YPD (Yeast Extract-Peptone-Dextrose) liquid medium for 16 - 24 hours, with shaking at 37°C and 200 rpm, and then diluted to OD600 ≈ 0.05 with sterile PBS or YPD.

[0305] In the experiment of measuring the antibacterial rate by spectrophotometry, the experimental group, positive control group and negative control group were set up in a 96-well plate. In the experimental group, 100 μL of the diluted bacterial suspension and 100 μL of antibacterial peptide solutions with different concentrations were added to make the final volume 200 μL; in the positive control group, 100 μL of the bacterial suspension and 100 μL of sterile PBS or medium were added; in the negative control group, only sterile PBS and medium were contained. Subsequently, they were incubated at 37°C for 16 - 20 hours (for fungi, it could be appropriately extended to 24 - 48 hours), and the absorbance at 600 nm (OD600) was measured using a spectrophotometer, and the antibacterial rate was calculated. The calculation formula is:

[0306] Antibacterial rate (%) = (1 - OD600 of the experimental group / OD600 of the control group) × 100%

[0307] If the antibacterial peptide may affect the OD600 reading (such as causing bacterial aggregation or sedimentation), the plate counting method can be combined for verification. In the experiment of measuring the antibacterial rate by the plate counting method, 100 μL of the bacterial suspension after the action of antibacterial peptides with different concentrations was taken, appropriately diluted with sterile PBS, and evenly spread on LB (for bacteria) or YPD (for yeast / fungi) agar plates, and incubated at 37°C for 16 - 24 hours (for fungi, it could be appropriately extended to 48 hours). Subsequently, the number of colonies (CFU / mL) was counted, and the antibacterial rate was calculated. The calculation formula is:

[0308] Antibacterial rate (%) = (1 - CFU of the experimental group / CFU of the control group) × 100%

[0309] 3.3 Determination of the minimum inhibitory concentration

[0310] In the minimum inhibitory concentration (MIC) determination experiment, the microbroth dilution method is adopted, which is applicable to bacteria and fungi. The culture media used in the experiment vary according to the types of strains. For bacteria (P. aeruginosa, E. coli, S. aureus), Mueller-Hinton (MH) broth medium is used, while for yeast and fungi (S. cerevisiae, C. albicans and other fungi), RPMI 1640 medium is used (it is recommended to carry out under the conditions of pH 7.0 and containing MOPS buffer).

[0311] First, the antimicrobial peptides are serially diluted two-fold. In the 96-well plate, starting from the first column, they are gradually diluted, and the final concentration range is generally set at 256 μg / mL to 0.125 μg / mL (the concentration gradient can be adjusted according to the activity of the antimicrobial peptide), and 100 μL of the antimicrobial peptide solution is added to each well. Subsequently, the bacterial suspension is added, and the concentration of the bacterial or yeast / fungal suspension is adjusted. The final bacterial concentration is adjusted to 1×106 CFU / mL, while the concentration of yeast / fungi is adjusted to 0.5×10 3 ~104 CFU / mL. 100 μL of the bacterial suspension is added to each well to make the final volume 200 μL.

[0312] In the experiment, positive and negative control groups need to be set. The positive control group contains only the bacterial suspension and the culture medium, while the negative control group contains only the culture medium and the antimicrobial peptide, without the bacterial suspension. Subsequently, the 96-well plate is incubated at 37°C. Bacteria are incubated for 16 - 20 hours, while yeast and fungi are incubated for 24 - 48 hours. The judgment standard of MIC is to determine the concentration of the antimicrobial peptide with the lowest complete aseptic growth by visual observation or measuring OD600 with a spectrophotometer, which is the MIC value. For fungi, 2,3,5-triphenyltetrazolium chloride staining (TTC) or XTT staining method can be used to enhance the visual judgment.

[0313] To improve the accuracy and repeatability of the experiment, the following optimization measures can be taken. First, use the McFarland standard (such as 0.5 McFarland, corresponding to approximately 1×108 CFU / mL) to adjust the concentration of the bacterial suspension to ensure the consistency of the experiment. Second, appropriately extend the incubation time. Since fungi grow slowly, it is recommended to incubate yeast for 24 hours and Candida albicans for 48 hours to improve the accuracy of MIC determination. In addition, each experiment should be repeated at least three times and the average value should be taken to reduce experimental errors. Finally, the minimum bactericidal concentration (MBC) determination can be combined. Samples are taken from the turbid bacterial suspension after MIC determination for plate counting to distinguish the inhibitory effect and the bactericidal effect.

[0314] Example 1: Antimicrobial Peptide Sequence Design Framework Based on Supervised Fine-Tuning

[0315] This embodiment relates to a framework for designing antimicrobial peptide sequences based on supervised fine-tuning, and its process is as follows Figure 1 shown, including the following steps:

[0316] (S1) Collect antimicrobial peptide data from public datasets;

[0317] (S2) Input the above antimicrobial peptide data into a protein backbone model and train the model. In this embodiment, the protein backbone model is ProGen2;

[0318] (S3) Fine-tune the model using the methods and steps in supervised fine-tuning in 1 to obtain the AMPGen model.

[0319] Embodiment 2: Antimicrobial Peptide Screening Framework Based on Protein Language Model and Bioinformatics Method

[0320] This embodiment relates to an antimicrobial peptide screening framework based on protein language model and bioinformatics method, and its process is as follows Figure 2 shown, including the following steps:

[0321] (C1) Use the AMPGen model obtained in Embodiment 1 to generate amino acid sequences as antimicrobial peptide candidate samples;

[0322] (C2) Input the above antimicrobial peptide candidate samples into the AMPGen-Filtering model for screening, where the screening methods and steps involved in the AMPGen-Filtering model are as shown in the screening and verification of 2 antimicrobial peptide candidate samples.

[0323] Embodiment 3: Antimicrobial Peptides Obtained from the Antimicrobial Peptide Sequence Design Framework Based on Supervised Fine-Tuning and the Antimicrobial Peptide Screening Framework

[0324] This embodiment relates to antimicrobial peptides obtained from the antimicrobial peptide sequence design framework based on supervised fine-tuning and the antimicrobial peptide screening framework. Specifically, a batch of candidate antimicrobial peptides can be obtained through the design framework in Embodiment 1, and the candidate antimicrobial peptides are input into the antimicrobial peptide screening framework for screening and verification, and finally 4 antimicrobial peptides with expected properties and MIC activities are obtained.

[0325] Test the molar mass and MIC of the above 4 antimicrobial peptides. The specific steps are as shown in the detection of 3 antimicrobial peptide properties.

[0326] The results of the molar mass and MIC of the 4 antimicrobial peptides are shown in Table 1-2:

[0327] Table 1. Molar Mass of 4 Antimicrobial Peptides

[0328]

[0329] Table 2. MIC (μg / mL) of 4 antimicrobial peptides

[0330]

[0331] As can be seen from the results in Table 2:

[0332] The MIC of ZJAMP006 is the lowest in Staphylococcus aureus (4 μg / mL) and Fusarium graminearum (8 μg / mL), indicating that it is more effective against Gram-positive bacteria and Fusarium graminearum. ZJAMP013 has a certain effect on Staphylococcus aureus (8 μg / mL), but other strains were not tested and further research is needed.

[0333] ZJAMP015 (MIC against Staphylococcus aureus = 32 μg / mL) has poor activity and may require a higher concentration to effectively inhibit bacteria.

[0334] The above results also reflect the need for further optimization for specific strains. For example, for Pseudomonas aeruginosa and Saccharomyces cerevisiae, higher concentrations of antimicrobial peptides may be required to produce good antibacterial effects; for Escherichia coli and Staphylococcus aureus, most antimicrobial peptides have shown good performance at 50 μg / mL, indicating that they are more sensitive to these antimicrobial peptides.

[0335] The amino acid sequences of the 4 antimicrobial peptides are shown in Table 3.

[0336] Table 3. Amino acid sequences of 4 antimicrobial peptides

[0337]

[0338] Discussion

[0339] In the examples of the design and screening of the antimicrobial peptides of the present invention, more strain tests can be added in the future. For example, determine the MIC of Escherichia coli, Candida albicans and Fusarium graminearum for ZJAMP009, ZJAMP013, ZJAMP015 to confirm their antibacterial spectra; or test antimicrobial peptides with lower MIC values, such as ZJAMP006, to further explore their minimum effective concentrations.

[0340] All documents mentioned in the present invention are cited herein as references, as if each document was individually cited as a reference. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

Claims

1. A method for constructing an antimicrobial peptide production model, characterized in that: The method comprises the steps of: (1) Train a protein pedestal model on the dataset; (2) Performing supervised fine-tuning on the protein base model to obtain a fine-tuned protein base model and an antimicrobial peptide candidate sequence.

2. The method according to claim 1, characterized in that The method for performing the supervised fine-tuning is low-rank adaptation.

3. A framework or system for producing antimicrobial peptides, characterized in that: The framework or system includes: (A1) an input unit, wherein the input unit is configured to input data, wherein the input data includes a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence; (A2) a fine-tuning unit, the fine-tuning unit being configured to execute a fine-tuning model on the input data, thereby obtaining fine-tuned input data; wherein the fine-tuning model comprises a supervised fine-tuning model, the supervised fine-tuning model comprising the steps of: performing supervised fine-tuning on the input data using a fine-tuning method, thereby obtaining fine-tuned input data; (A3) An output unit, wherein the output unit is configured to output the result of the fine-tuning unit.

4. A method for constructing an antimicrobial peptide screening model, characterized in that: The method comprises the steps of: (M1) machine learning screening, the machine learning screening comprising the steps of: (B1) Using a machine learning model, the sequence similarity between the candidate antimicrobial peptide sequence and the target functional sample is evaluated through a scoring system; (B2) predictive indicators of antimicrobial activity; (M2) a posteriori verification, the a posteriori verification comprising a step selected from the group consisting of: (C1) limit peptide length; (C2) predicting structural features; (C3) Analyze similarities; or a combination thereof; Wherein, the scoring system is selected from the following group: supervised learning algorithm, reinforcement learning algorithm, unsupervised learning algorithm, or semi-supervised learning algorithm; the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the following group: linear model, K nearest neighbor, random forest, or a combination thereof; the unsupervised learning algorithm is a clustering algorithm, and the clustering algorithm is selected from the following group: density clustering, or K-means clustering; The antibacterial activity index is the minimum inhibitory concentration; The structural features are foldability and thermal stability; Predicting structural features by a method selected from the group consisting of a protein structure prediction tool, or molecular simulation; the protein structure prediction tool is a combination of AlphaFold2 and ESMFold; The similarity is analyzed by a method selected from the group consisting of: a sequence alignment algorithm, or a structure alignment algorithm; the sequence alignment algorithm is BLAST; the structure alignment algorithm is a combination of FoldSeek and TM-align.

5. The method according to claim 4, characterized in that In step (B1), comprising: (b1.1) using the machine learning model to convert the antimicrobial peptide candidate sequence and / or the target function sample into a numerical form; (b1.2) calculating the similarity between the antimicrobial peptide candidate sequence and the target functional sample; in, The machine learning model is selected from the group consisting of a protein language model, a deep learning model, or a combination thereof; the protein language model is an ESM2 model; The protein language model converts the antimicrobial peptide candidate sequence and / or the target function sample into the numerical form by a method selected from the following group: a natural language processing algorithm or a sequence alignment algorithm; the natural language processing algorithm is a word vector, and the word vector is one-hot encoding; the sequence alignment algorithm is a protein substitution scoring matrix, and the protein substitution matrix is ​​a BLOSUM matrix; the numerical form is an embedding vector; The deep learning model is a model based on a recurrent neural network.

6. An antimicrobial peptide screening framework or system, characterized in that: The framework or system includes: (D1) an input unit, wherein the input unit is configured to input data, wherein the input data includes an input sequence to be screened, wherein the input sequence to be screened includes an antimicrobial peptide candidate sequence generated using the antimicrobial peptide generation framework or system according to claim 3; (D2) an antimicrobial peptide screening unit, the screening unit being configured to execute an antimicrobial peptide screening model to obtain an antimicrobial peptide target sequence from the antimicrobial peptide candidate sequence; wherein the screening model is constructed using the method of claim 4; (D3) An output unit, wherein the output unit is configured to output the screening result.

7. A framework or system for designing and screening antimicrobial peptides, characterized in that: The framework or system includes: (W) an input unit, wherein the input unit is configured to input data, wherein the input data includes a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence; (X) a fine-tuning unit, the fine-tuning unit being configured as a fine-tuning model, the fine-tuning model performing a predetermined fine-tuning on the input data, thereby obtaining a fine-tuning result; wherein the fine-tuning model is a supervised fine-tuning model, the supervised fine-tuning model comprising the steps of: performing supervised fine-tuning on the input data, thereby obtaining fine-tuned input data; (Y) a screening unit, the screening unit being configured to execute a screening model for antimicrobial peptides, the screening model obtaining an antimicrobial peptide target sequence from the fine-tuned input data; wherein the screening model is constructed by the method of claim 3; (Z) an output unit, wherein the output unit is configured to output the result of the screening module.

8. The framework or system of claim 7, wherein: The antimicrobial peptide target sequence has an amino acid sequence as shown in any one of SEQ ID NOs: 1-4.

9. A drug or a pharmaceutical composition, characterized in that The drug or pharmaceutical composition comprises an antimicrobial peptide sequence, and the antimicrobial peptide sequence is obtained by the framework or system of claim 7.

10. Use of the medicine or pharmaceutical composition according to claim 9, characterized in that: The uses include: (F1) preparing a medicament for inhibiting the growth of microorganisms; (F2) preparing a kit, the kit further comprising a label or instructions indicating that the kit is used to inhibit the growth of microorganisms.