Batch of antibacterial peptide sequences based on protein language model supervised fine tuning and screening design
The supervised fine-tuning and screening design of antimicrobial peptides is solved through deep learning methods based on protein language model, and the problem of low antimicrobial peptide discovery and design efficiency in the prior art was successfully discovered, and a variety of highly active antimicrobial peptides were improved, which has improved its application potential.
Patent Information
- Application Number
- CN202510211660.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
Existing antimicrobial peptide discovery and design methods face problems such as low efficiency, high resource consumption, difficulty in generating antimicrobial peptides of target properties, and ignoring unconventional mechanisms of action.
Using a deep learning method based on protein language model, antimicrobial peptide sequences are generated through supervised fine-tuning and screening design, and the antimicrobial peptide generation and screening framework is used to screen and verify candidate sequences.
The generation efficiency and accuracy of antimicrobial peptide sequences have been improved, and a variety of new high-active antimicrobial peptides have been discovered, which can effectively inhibit the growth of a variety of microorganisms and significantly improve the practical application potential of antimicrobial peptides.
Smart Images

Figure BDA0005286168850000021 
Figure BDA0005286168850000051 
Figure BDA0005286168850000061
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of antimicrobial peptide design and biomedicine. Specifically, it involves the development of a method for the research and design of a novel antimicrobial drug by using bioinformatics and artificial intelligence technologies, especially deep learning methods based on protein language models, and a batch of novel antimicrobial peptides have been discovered. Background Art
[0002] Antimicrobial Peptides (AMP) are a class of short-chain amino acid sequences (usually defined as 12 - 50 amino acid residues) that can inhibit microbial growth by interfering with cell wall integrity. Antimicrobial peptides play an important role in coping with the increasingly serious Antimicrobial Resistance (AMR) crisis. According to the prediction of the World Health Organization, by 2050, AMR may cause 10 million deaths annually and impose a cumulative loss of $100 trillion on the global economy.
[0003] Traditional methods for discovering antimicrobial peptides include isolation from natural sources, rational design, and high-throughput screening (HTS) of synthetic peptide libraries. These methods have successfully identified many antimicrobial peptides, but also face many challenges. For example, isolation from natural sources is time-consuming and limited by the available biodiversity; rational design requires a large amount of resources for verification; HTS is limited by the initial library design and infrastructure requirements. These methods often have difficulty efficiently generating AMPs with target properties and may overlook peptides with unconventional mechanisms of action.
[0004] To overcome the above limitations, researchers have begun to explore artificial intelligence (AI)-based methods for discovering and designing antimicrobial peptides. However, existing AI-based computational methods still face two main challenges in antimicrobial peptide discovery and design: (1) Transformation from in vitro to in vivo: The transformation of computationally designed antimicrobial peptides from in vitro to in vivo efficacy requires comprehensive pharmacokinetic and pharmacodynamic studies. Key factors such as bioavailability, metabolic stability, and potential off-target effects need to be thoroughly studied in a complex biological environment; (2) Inherent limitations of computational methods, such as biases and limitations in the AMP database, difficulties in accurately predicting functional properties, and the challenge of balancing diversity, novelty, and antimicrobial properties.
[0005] In summary, there is an urgent need in this field to develop a method for the research and design of a novel antimicrobial drug with high efficiency and high accuracy. Summary of the Invention
[0006] The object of the present invention is to provide a batch of antimicrobial peptide sequences designed based on supervised fine-tuning and screening of protein language models, as well as an antimicrobial peptide generation and screening framework and its construction method used in the process of generating the antimicrobial peptide sequences.
[0007] In the first aspect of the present invention, an antimicrobial peptide sequence is provided, and the antimicrobial peptide is obtained through an antimicrobial peptide generation framework and an antimicrobial peptide screening framework.
[0008] In another preferred embodiment, the antimicrobial peptide sequence comprises the amino acid sequence shown in any one of SEQ ID NO: 1-4.
[0009] In another preferred embodiment, the molar mass range of the antimicrobial peptide sequence is 1400-2800 g / mol.
[0010] In another preferred embodiment, the minimum inhibitory concentration of the antimicrobial peptide sequence is < 35 μg / mL, preferably < 20 μg / mL, more preferably < 8 μg / mL.
[0011] In another preferred embodiment, the antimicrobial peptide generation framework comprises:
[0012] (A1) An input unit configured to input data, the input data including a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence;
[0013] (A2) A fine-tuning unit configured to perform a fine-tuning model on the input data to obtain fine-tuned input data; wherein the fine-tuning model includes a supervised fine-tuning model, and the supervised fine-tuning model includes the steps of: performing supervised fine-tuning on the input data by using a fine-tuning method to obtain fine-tuned input data;
[0014] (A3) An output unit configured to output the result of the fine-tuning unit.
[0015] In another preferred embodiment, the protein base model is a protein language model.
[0016] In another preferred embodiment, the protein language model includes: ProGen2 model, ESM-1b, ProtBERT, ProtXLNet, TAPE.
[0017] In another preferred embodiment, the protein language model is the ProGen2 model.
[0018] In another preferred embodiment, the input sequence is from a publicly available antimicrobial peptide dataset.
[0019] In another preferred embodiment, the fine-tuning methods include: Low-Rank Adaptation (LoRA), Adapter-tuning, Prefix-tuning, Prompt-tuning, P-tuning, P-tuning v1, P-tuning v2.
[0020] In another preferred example, the fine-tuning method is low-rank adaptation.
[0021] In another preferred example, perplexity is used to guide the supervised fine-tuning model.
[0022] In another preferred example, the loss function of the low-rank adaptation is:
[0023]
[0024] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the low-rank adaptation parameters.
[0025] In another preferred example, the antimicrobial peptide screening framework includes:
[0026] (B1) An input unit configured to input data, where the input data includes an input sequence to be screened;
[0027] (B2) A screening unit for antimicrobial peptides, configured to execute a screening model for antimicrobial peptides to obtain an antimicrobial peptide target sequence from the input sequence to be screened;
[0028] (B3) An output unit configured to output a screening result.
[0029] In another preferred example, the input sequence to be screened includes an antimicrobial peptide candidate sequence generated using the antimicrobial peptide generation framework.
[0030] In another preferred example, the antimicrobial peptide generation framework is AMPGen.
[0031] In another preferred example, the screening model includes the steps of:
[0032] (I) Machine learning screening, which includes the steps of:
[0033] (C1) Using a machine learning model, evaluating the sequence similarity between an antimicrobial peptide candidate sequence and a target function sample through a scoring system;
[0034] (C2) Predicting an antimicrobial activity index;
[0035] (II) Posterior verification, which includes steps selected from the following group:
[0036] (D1) Limiting the peptide segment length;
[0037] (D2) Predicting structural features;
[0038] (D3) Analyze similarity;
[0039] or a combination thereof;
[0040] wherein, the scoring system is selected from the group consisting of: supervised learning algorithms, reinforcement learning algorithms, unsupervised learning algorithms, or semi-supervised learning algorithms; the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the group consisting of: linear models, K-nearest neighbors, random forests, or a combination thereof; the unsupervised learning algorithm is a clustering algorithm, and the clustering algorithm is selected from the group consisting of: density clustering, or K-means clustering;
[0041] The antibacterial activity index is the minimum inhibitory concentration;
[0042] The structural features are foldability and thermal stability;
[0043] The structural features are predicted by a method selected from the group consisting of: protein structure prediction tools, or molecular simulation; the protein structure prediction tool is a combination of AlphaFold2 and ESMFold;
[0044] The similarity is analyzed by a method selected from the group consisting of: sequence alignment algorithms, or structure alignment algorithms; the sequence alignment algorithm is BLAST; the structure alignment algorithm is a combination of FoldSeek and TM-align.
[0045] In another preferred example, the method further includes the step: (III) High-throughput screening.
[0046] In another preferred example, in step (I), it further includes:
[0047] (C3) Evaluate the stability of the antimicrobial peptide;
[0048] (C4) Evaluate the structural similarity.
[0049] In another preferred example, in step (C1), it includes:
[0050] (c1.1) Using the machine learning model, convert the antimicrobial peptide candidate sequence and / or the target functional sample into a numerical form;
[0051] (c1.2) Calculate the similarity between the antimicrobial peptide candidate sequence and the target functional sample;
[0052] wherein,
[0053] The machine learning model is selected from the group consisting of: protein language models, deep learning models, or a combination thereof; the protein language model is the ESM2 model;
[0054] The protein language model converts the antimicrobial peptide candidate sequence and / or the target functional sample into the numerical form by a method selected from the group consisting of: natural language processing algorithms, or sequence alignment algorithms; the natural language processing algorithm is a word vector, and the word vector is a one-hot encoding; the sequence alignment algorithm is a protein substitution scoring matrix, and the protein substitution matrix is a BLOSUM matrix; the numerical form is an embedding vector.
[0055] The deep learning model is a model based on a recurrent neural network.
[0056] In another preferred example, the machine learning model is a combination of a protein language model and a deep learning model.
[0057] In another preferred example, the protein language model includes: ESM2 model, ProtBERT, ProtXLNet.
[0058] In another preferred example, the protein language model is the ESM2 model.
[0059] In another preferred example, the protein language model converts the antimicrobial peptide candidate sequence and / or the target functional sample into the numerical form by a method selected from the group consisting of: natural language processing algorithms, or sequence alignment algorithms.
[0060] In another preferred example, the natural language processing algorithm is a word vector.
[0061] In another preferred example, the ESM2 model evaluates the sequence similarity between the antimicrobial peptide candidate sequence and the target functional sample through zero-shot learning.
[0062] In another preferred example, the word vector is selected from the group consisting of: one-hot encoding, or word2vec.
[0063] In another preferred example, the word vector is a one-hot encoding.
[0064] In another preferred example, the sequence alignment algorithm is a protein substitution scoring matrix.
[0065] In another preferred example, the protein substitution scoring matrix is selected from the group consisting of: PAM matrix, or BLOSUM matrix.
[0066] In another preferred example, the protein substitution scoring matrix is a BLOSUM matrix.
[0067] In another preferred example, the numerical form includes: embedding vector, scoring matrix.
[0068] In another preferred example, the numerical form is an embedding vector.
[0069] In another preferred example, the embedding vector is an ESM2 embedding.
[0070] In another preferred example, the ESM2 embedding is an embedding vector formed by converting the candidate antimicrobial peptide sequence and the target functional sample using the ESM2 model.
[0071] In another preferred example, the deep learning model includes: a model based on a convolutional neural network (CNN), a model based on a recurrent neural network (RNN).
[0072] In another preferred example, the deep learning model is a model based on a recurrent neural network (RNN).
[0073] In another preferred example, the scoring system employs an algorithm selected from the group consisting of: supervised learning algorithms, reinforcement learning algorithms, unsupervised learning algorithms, or semi-supervised learning algorithms.
[0074] In another preferred example, the scoring system employs an algorithm selected from the group consisting of: supervised learning algorithms, or reinforcement learning algorithms.
[0075] In another preferred example, the supervised learning algorithm is a classification algorithm.
[0076] In another preferred example, the classification algorithm is selected from the group consisting of: linear models, K-nearest neighbors (KNN), random forests, support vector machines (SVM), decision trees, neural networks, naive Bayes, Boosting, or combinations thereof.
[0077] In another preferred example, the classification algorithm is selected from the group consisting of: linear models, K-nearest neighbors, random forests, or combinations thereof.
[0078] In another preferred example, the classification algorithm is K-nearest neighbors.
[0079] In another preferred example, the calculation of evaluating sequence similarity by the K-nearest neighbors is as follows:
[0080]
[0081] where s_f is the similarity score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric, x is the generated sample, y_i is the center of the i-th functional cluster, k is the number of clusters, and d is the number of distance metrics.
[0082] In another preferred example, the distance metric is selected from the group consisting of: Euclidean distance, Manhattan distance, cosine similarity, Chebyshev distance, Minkowski distance, standard Euclidean distance, Mahalanobis distance, Hamming distance, Jaccard distance, correlation distance, information entropy, or combinations thereof.
[0083] In another preferred example, the distance metric is Euclidean distance, Manhattan distance, and cosine similarity.
[0084] In another preferred example, d = 3.
[0085] In another preferred example, the scoring system includes obtaining multiple similarity scores by using multiple supervised learning algorithms.
[0086] In another preferred example, the multiple similarity scores are synthesized by using a method selected from the following group: fuzzy logic, Dempster-Shafer theory, or multi-objective optimization algorithm.
[0087] In another preferred example, the multi-objective optimization algorithm is NSGA-II.
[0088] In another preferred example, the unsupervised learning algorithm is a clustering algorithm.
[0089] In another preferred example, the clustering algorithm is selected from the following group: density clustering, K-means clustering, spectral clustering, hierarchical clustering, grid clustering, or model clustering.
[0090] In another preferred example, the density clustering is DBSCAN.
[0091] In another preferred example, a computational model is used to evaluate the stability of antimicrobial peptides.
[0092] In another preferred example, the computational model includes a computational model for simulating protein degradation.
[0093] In another preferred example, a graph neural network is used to evaluate structural similarity.
[0094] In another preferred example, the antimicrobial activity indicators include: minimum inhibitory concentration, minimum bactericidal concentration, time-kill curve.
[0095] In another preferred example, the antimicrobial activity indicator is the minimum inhibitory concentration.
[0096] In another preferred example, the minimum inhibitory concentration is predicted by using a method selected from the following group: supervised learning algorithm, expert system, or deep learning model.
[0097] In another preferred example, the supervised learning algorithm is selected from the following group: classification algorithm, or regression algorithm.
[0098] In another preferred example, the classification algorithm is Naive Bayes.
[0099] In another preferred example, the Naive Bayes-based classifier is trained on a known data set so that the classifier reaches the training objective, and a pre-trained classifier is obtained, thereby using the pre-trained classifier to predict the minimum inhibitory concentration.
[0100] In another preferred example, the known data set is an antimicrobial peptide data set with experimentally determined minimum inhibitory concentration values.
[0101] In another preferred example, the Naive Bayes-based classifier is a Bayesian classifier.
[0102] In another preferred example, the Bayesian classifier uses the ESM2 embedding.
[0103] In another preferred example, the training objective is to minimize the L1 loss between the predicted minimum inhibitory concentration value and the actual minimum inhibitory concentration value, where the calculation of the L1 loss is:
[0104]
[0105] where \(y_i\) is the actual minimum inhibitory concentration value, \(f(x_i)\) is the predicted minimum inhibitory concentration value of the numerical form \(x_i\) of the \(i\)-th peptide segment, and \(N\) is the number of training samples.
[0106] In another preferred example, the regression algorithm is selected from the group consisting of: linear regression, K-nearest neighbor regression, support vector machine regression, decision tree regression, neural network regression, Naive Bayes regression, Boosting regression, random forest regression, deep forest regression, or extremely randomized tree regression.
[0107] In another preferred example, the Boosting regression is gradient boosting decision tree regression.
[0108] In another preferred example, the expert system includes a rule-based expert system.
[0109] In another preferred example, the deep learning model is an end-to-end deep learning model.
[0110] In another preferred example, the end-to-end deep learning model directly identifies the minimum inhibitory concentration from the antimicrobial peptide candidate sequence.
[0111] In another preferred example, the posterior verification screens the antimicrobial peptide candidate sequences by a method selected from the group consisting of:
[0112] (D1) Limiting the peptide segment length;
[0113] (D2) Predicting structural features; and
[0114] (D3) Analyzing similarity.
[0115] In another preferred embodiment, the length of the peptide segment is ≤ 50 amino acids, preferably ≤ 30 amino acids, more preferably ≤ 25 amino acids.
[0116] In another preferred embodiment, the structural feature is foldability and thermal stability.
[0117] In another preferred embodiment, the structural feature is predicted by a method selected from the group consisting of: protein structure prediction tools, or molecular simulations.
[0118] In another preferred embodiment, the protein structure prediction tool is selected from the group consisting of: AlphaFold2, ESMFold, I-TASSER, RoseTTAFold, GalaxyTBM, SWISS-MODEL, or a combination thereof.
[0119] In another preferred embodiment, the protein structure prediction tool is a combination of AlphaFold2 and ESMFold.
[0120] In another preferred embodiment, the molecular simulation is molecular dynamics simulation.
[0121] In another preferred embodiment, the similarity includes: sequence similarity, structural similarity.
[0122] In another preferred embodiment, the similarity is analyzed by a method selected from the group consisting of: sequence alignment algorithms, or structural alignment algorithms.
[0123] In another preferred embodiment, the sequence similarity is analyzed by the sequence alignment algorithm.
[0124] In another preferred embodiment, the structural similarity is analyzed by the structural alignment algorithm.
[0125] In another preferred embodiment, the sequence alignment algorithm is selected from the group consisting of: BLAST, Smith-Waterman algorithm, Needleman-Wunsch algorithm, or a combination thereof.
[0126] In another preferred embodiment, the sequence alignment algorithm is BLAST.
[0127] In another preferred embodiment, the structural alignment algorithm is selected from the group consisting of: FoldSeek, TM-align, DALI, SSAP, FLEXPROT, or a combination thereof.
[0128] In another preferred embodiment, the structural alignment algorithm is a combination of FoldSeek and TM-align.
[0129] In another preferred embodiment, the post hoc verification further includes screening antibacterial peptide candidate sequences by a method selected from the group consisting of: integrating multi-omics data, or knowledge graph technology.
[0130] In another preferred example, the multi-omics data is selected from the group consisting of: proteomics, metabolomics, or a combination thereof.
[0131] In another preferred example, the multi-omics data is proteomics.
[0132] In another preferred example, the screening model further includes: using an evaluation model to comprehensively screen antibacterial peptide candidate sequences based on multiple indicators.
[0133] In another preferred example, the evaluation model is selected from the group consisting of: AHP, TOPSIS, grey relational analysis, fuzzy comprehensive evaluation, entropy weight method, or a combination thereof.
[0134] In another preferred example, the evaluation model is AHP.
[0135] In another preferred example, the indicators are selected from the group consisting of: sequence similarity, structural similarity, structural features, minimum inhibitory concentration, or a combination thereof.
[0136] In another preferred example, the screening model is AMPGen-Filtering.
[0137] In a second aspect of the present invention, there is provided a drug or a pharmaceutical composition, which contains one or more antibacterial peptide sequences as described in the first aspect of the present invention.
[0138] In another preferred example, the drug or the pharmaceutical combination includes the amino acid sequence shown in any one of SEQ ID NO: 1-4.
[0139] In another preferred example, the drug or the pharmaceutical composition further includes: the protein family to which the antibacterial peptide sequence belongs, and antibacterial peptide sequences in the protein family to which the antibacterial peptide sequence belongs and having a similar structure or function to the antibacterial peptide sequence.
[0140] In another preferred example, the drug or the pharmaceutical composition further includes: the protein family to which the amino acid sequence shown in any one of SEQ ID NO: 1-4 belongs, and antibacterial peptide sequences in the protein family to which the amino acid sequence shown in any one of SEQ ID NO: 1-4 belongs and having a similar structure or function to the amino acid sequence shown in any one of SEQ ID NO: 1-4.
[0141] In another preferred example, the drug or the pharmaceutical composition further contains a pharmaceutically acceptable carrier, diluent or excipient.
[0142] In another preferred example, the antibacterial peptide sequence is prepared by direct synthesis.
[0143] The third aspect of the present invention provides a use of the drug or pharmaceutical composition described in the second aspect of the present invention, and the use includes:
[0144] (E1) preparing a preparation for inhibiting the growth of microorganisms;
[0145] (E2) preparing a kit, which further includes a label or an instruction manual, and the instruction manual indicates that the kit is used for inhibiting the growth of microorganisms.
[0146] In another preferred example, the microorganisms include bacteria and fungi.
[0147] In another preferred example, the bacteria include Pseudomonas aeruginosa, Escherichia coli, and Staphylococcus aureus.
[0148] In another preferred example, the fungi include Saccharomyces cerevisiae, Candida albicans, and Fusarium graminearum.
[0149] In another preferred example, the preparation includes a liquid preparation.
[0150] The fourth aspect of the present invention provides a method for constructing an antibacterial peptide generation model, and the method includes the steps:
[0151] (1) training a protein base model on a data set;
[0152] (2) performing supervised fine-tuning on the protein base model to obtain a fine-tuned protein base model and antibacterial peptide candidate sequences.
[0153] In another preferred example, in step (2), it further includes the step of performing reinforcement learning on the fine-tuned protein base model to obtain an optimized protein model and antibacterial peptide candidate sequences.
[0154] In another preferred example, in step (1), the data set is a publicly available antibacterial peptide data set.
[0155] In another preferred example, in step (2.1), the data set is a publicly available antibacterial peptide data set with a low minimum inhibitory concentration and a publicly available non-active antibacterial peptide data set.
[0156] In another preferred example, the data set is obtained from the original data set by a data augmentation method.
[0157] In another preferred example, the data augmentation method is selected from the following group:
[0158] (i) performing sequence variation on the original data set;
[0159] (ii) performing structural perturbation on the original data set;
[0160] (iii) Perform conditional generation on the original dataset;
[0161] Or a combination thereof.
[0162] In another preferred example, the protein base model is a protein language model.
[0163] In another preferred example, the protein language model includes: ProGen2 model, ESM-1b, ProtBERT, ProtXLNet, TAPE.
[0164] In another preferred example, the protein language model is the ProGen2 model.
[0165] In another preferred example, the method for performing the supervised fine-tuning includes: Low-Rank Adaptation (LoRA), Prefix-tuning, Prompt-tuning, P-tuning, P-tuning v1, P-tuning v2, Adapter-tuning.
[0166] In another preferred example, the method for performing the supervised fine-tuning is Low-Rank Adaptation.
[0167] In another preferred example, perplexity is used to guide the supervised fine-tuning.
[0168] In another preferred example, the loss function of the Low-Rank Adaptation is:
[0169]
[0170] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the Low-Rank Adaptation parameters.
[0171] In another preferred example, the fine-tuned protein base model is AMPGen.
[0172] In the fifth aspect of the present invention, a method for constructing an antimicrobial peptide screening model is provided, and the method includes the steps:
[0173] (I) Machine learning screening, and the machine learning screening includes the steps:
[0174] (C1) Using a machine learning model, evaluating the sequence similarity between an antimicrobial peptide candidate sequence and a target function sample through a scoring system;
[0175] (C2) Predicting antimicrobial activity indicators;
[0176] (II) Posterior verification, and the posterior verification includes steps selected from the following group:
[0177] (D1) Restricting the peptide length;
[0178] (D2) Predicting structural features;
[0179] (D3) Analyzing similarity;
[0180] Or a combination thereof;
[0181] Wherein, the scoring system is selected from the following group: supervised learning algorithm, reinforcement learning algorithm, unsupervised learning algorithm, or semi-supervised learning algorithm; the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the following group: linear model, K-nearest neighbor, or random forest; the unsupervised learning algorithm is a clustering algorithm, and the clustering algorithm is selected from the following group: density clustering, or K-means clustering;
[0182] The antibacterial activity index is the minimum inhibitory concentration;
[0183] The structural features are foldability and thermal stability;
[0184] The structural features are predicted by a method selected from the following group: protein structure prediction tool, or molecular simulation; the protein structure prediction tool is a combination of AlphaFold2 and ESMFold;
[0185] The similarity is analyzed by a method selected from the following group: sequence alignment algorithm, or structure alignment algorithm; the sequence alignment algorithm is BLAST; the structure alignment algorithm is a combination of FoldSeek and TM-align.
[0186] In another preferred example, the method further includes the step: (III) High-throughput screening.
[0187] In another preferred example, in step (I), it further includes:
[0188] (C3) Evaluating the stability of the antibacterial peptide;
[0189] (C4) Evaluating the structural similarity.
[0190] In another preferred example, in step (C1), it includes:
[0191] (c1.1) Using the machine learning model to convert the antibacterial peptide candidate sequence and / or the target functional sample into a numerical form;
[0192] (c1.2) Calculating the similarity between the antibacterial peptide candidate sequence and the target functional sample;
[0193] Wherein,
[0194] The machine learning model is selected from the group consisting of: a protein language model, a deep learning model, or a combination thereof; the protein language model is the ESM2 model;
[0195] The protein language model converts the antimicrobial peptide candidate sequence and / or the target functional sample into the numerical form by a method selected from the group consisting of: a natural language processing algorithm, or a sequence alignment algorithm; the natural language processing algorithm is a word vector, and the word vector is a one-hot encoding; the sequence alignment algorithm is a protein substitution scoring matrix, and the protein substitution matrix is a BLOSUM matrix; the numerical form is an embedding vector;
[0196] The deep learning model is a model based on a recurrent neural network.
[0197] In a sixth aspect of the present invention, there is provided an antimicrobial peptide design and screening framework or system, the framework or system comprising:
[0198] (W) An input unit configured to input data, the input data including a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence;
[0199] (X) A fine-tuning unit configured to fine-tune a model, the fine-tuned model performing a predetermined fine-tuning on the input data to obtain a fine-tuning result; wherein the fine-tuned model is a supervised fine-tuning model, and the supervised fine-tuning model includes the steps of: performing supervised fine-tuning on the input data by a fine-tuning method to obtain fine-tuned input data;
[0200] (Y) A screening unit configured to execute an antimicrobial peptide screening model, the screening model obtaining an antimicrobial peptide target sequence from the fine-tuned input data; wherein the screening model is constructed by the method of the fifth aspect of the present invention;
[0201] (Z) An output unit configured to output the result of the screening module.
[0202] In another preferred example, the fine-tuned input data includes: a fine-tuned input sequence and / or a fine-tuned protein base model.
[0203] In another preferred example, the fine-tuned input sequence includes an amino acid sequence as shown in any one of SEQ ID NO: 1-4.
[0204] In another preferred example, the antimicrobial peptide target sequence has an amino acid sequence as shown in any one of SEQ ID NO: 1-4.
[0205] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features specifically described hereinafter (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be elaborated one by one here. BRIEF DESCRIPTION OF THE DRAWINGS
[0206] Figure 1 Shows the process of the antimicrobial peptide sequence design framework based on supervised fine-tuning.
[0207] Figure 2 Shows the process of the antimicrobial peptide screening framework based on protein language models and bioinformatics methods. DETAILED DESCRIPTION OF THE INVENTION
[0208] Through a large number of in-depth studies, the inventors of the present invention for the first time discovered 4 novel highly active antimicrobial peptide sequences based on the supervised fine-tuning of protein language models and multiple screening processes. These antimicrobial peptides have inhibitory effects on the growth of various microorganisms. In addition, the present invention also for the first time developed an antimicrobial peptide generation framework using protein language models through supervised fine-tuning technology, and designed an antimicrobial peptide screening framework based on various properties of antimicrobial peptides for the first time. The two frameworks together constitute a novel antimicrobial peptide design framework, which can generate diverse and novel antimicrobial peptide sequences while maintaining their key antimicrobial activities, significantly improving the practical application potential of the generated antimicrobial peptide sequences, providing an efficient, accurate and flexible new method for the research and development of antimicrobial peptides, and is expected to accelerate the discovery and development process of novel antimicrobial drugs. On this basis, the present invention was completed.
[0209] The following explains some of the innovative points of the present invention:
[0210] First, the present invention uses the antimicrobial peptide data in the public dataset as input data, inputs it into the protein base model, and obtains the training model AMPGen through supervised fine-tuning. Among them, ProGen2 is used as the base model, low-rank adaptation technology is used for supervised fine-tuning, and perplexity is used to guide the model fine-tuning. Low-rank adaptation can minimize the number of parameters and reduce the risk of overfitting; perplexity can evaluate the quality of the model fine-tuning results and guide the model optimization. In the supervised fine-tuning using low-rank adaptation technology, the loss function of the supervised fine-tuning is defined as follows:
[0211]
[0212] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the LoRA parameters.
[0213] After the above model training is completed, the present invention uses a screening framework to screen candidate samples of antimicrobial peptides. The screening framework is AMPGen-Filtering, which includes two-stage screening. In the first stage, first, the pre-trained protein language model ESM2 is used to evaluate the similarity between the sequence and known antimicrobial peptides in a zero-shot manner, and a MIC-based classifier is used to predict the antimicrobial activity of the sequence. Secondly, the generated antimicrobial peptide samples are screened by using the k-nearest neighbor algorithm, and the screening score is calculated. The calculation formula of the screening score is:
[0214]
[0215] where s_f is the final similarity screening score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric (Euclidean distance, Manhattan distance, and cosine similarity), x is the generated sample, y_i is the center of the i-th functional cluster, and k is the number of clusters.
[0216] After that, a Bayesian classifier based on ESM2 embedding is used to further refine the selection process. This classifier is trained on the known AMP dataset with MIC values, and the goal is to minimize the L1 loss between the predicted MIC value and the actual MIC value. The calculation formula of the loss value L1 is:
[0217]
[0218] where y_i is the true MIC value, f(x_i) is the predicted MIC value of the ESM2 embedding x_i of the i-th peptide segment, and N is the number of training samples.
[0219] Finally, the sequences obtained by MIC prediction and k-nearest neighbor algorithm screening are cross-validated to further refine the selection of candidate peptide segments.
[0220] In the second stage, the antimicrobial peptide candidates are optimized by length screening, structure prediction, and similarity analysis. Among them, AlphaFold2 and ESMFold are used to evaluate the sequence structure characteristics in structure prediction; BLAST and FoldSeek are used for sequence similarity search in similarity analysis.
[0221] Through the above antimicrobial peptide generation framework and screening framework, the present invention has discovered 4 novel highly active antimicrobial peptide sequences. The above 4 sequences have the potential to be applied in fields such as agriculture and food preservation.
[0222] It should be understood that the specific methods and experimental conditions of the present invention described below in various levels of detail are used to provide an essential understanding of the present invention. Definitions of certain terms used in this specification are provided below. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention pertains.
[0223] The term
[0224] As used herein, the terms "comprising", "including", and "containing" can be used interchangeably, and include not only closed definitions but also semi-closed and open definitions. In other words, the said terms include "consisting of" and "consisting essentially of".
[0225] As used herein, the terms "baseline generator", "base model", "protein baseline generator", and "protein base model" can be used interchangeably, and all refer to the original machine learning model that can be used for protein research without training or fine-tuning, and can have good target protein generation ability after training and fine-tuning.
[0226] As used herein, the term "Perplexity" is an evaluation metric for the performance of a language model, used to measure the prediction accuracy of the model for a given sequence. The lower the perplexity, the better the model performance.
[0227] As used herein, the terms "physicochemical property" and "physicochemical property" can be used interchangeably, and are quantitative indicators describing molecular characteristics. In the present invention, the physicochemical properties of antimicrobial peptides include hydrophobicity, hydrophobic moment, charge, isoelectric point, etc.
[0228] As used herein, the terms "Zero-shot Learning" and "zero-shot manner" can be used interchangeably, and refer to the ability of a model to recognize or classify categories not seen during the training process.
[0229] As used herein, the term "Protein Family" is a group of evolutionarily related proteins with similar sequences, structures, or functions.
[0230] As used herein, the terms "Sequence Embedding", "embedding vector", and "embedding" can be used interchangeably, and refer to the conversion of a protein sequence into a vector representation of a fixed dimension to capture the semantic information of the sequence. In the present invention, generating a sample embedding refers to the embedding vector transformed from a candidate antimicrobial peptide sequence; the target functional sample embedding refers to the embedding vector transformed from an antimicrobial peptide sequence with known function and sequence.
[0231] As used herein, the term "Molecular Dynamics Simulation" is a computational method for simulating the evolution of a molecular system over time by computer, and is used to study the motion and interactions of molecules.
[0232] As used herein, the terms "antibacterial peptide MIC value predictor", "minimum inhibitory concentration predictor", and "MIC identifier" are used interchangeably, and refer to a classifier that can predict MIC, which is obtained by training on a known dataset containing MIC data during the activity-based feedback fine-tuning stage.
[0233] As used herein, the terms "TM-Score (Template Modeling Score)" and "TM score" are used interchangeably, and are a score used to evaluate the similarity of protein structures, ranging from 0 to 1, and the closer to 1, the more similar the structures.
[0234] As used herein, the term "RMSD (Root Mean Square Deviation)" refers to the root mean square deviation, which is used to measure the average distance of the atomic spatial positions between two superimposed protein structures, with the unit of angstrom.
[0235] As used herein, the term "Multi-objective Optimization Algorithm" is an algorithm designed to optimize multiple objective functions simultaneously, such as NSGA-II (Non-dominated Sorting Genetic Algorithm II).
[0236] As used herein, the term "K-Nearest Neighbors (KNN)" is a machine learning algorithm for classification and regression, which makes predictions based on the K nearest neighbors around a sample.
[0237] As used herein, the term "Bayesian Classifier" refers to a probabilistic classifier based on Bayes' theorem, which can make classification predictions based on the conditional probabilities of features.
[0238] As used herein, the term "BLAST" is the Basic Local Alignment Search Tool, which is a sequence similarity search algorithm used to compare biological sequences (such as DNA, RNA, or protein sequences) with a sequence database to identify similar sequences in the database.
[0239] As used herein, the terms "antimicrobial peptide target sequence" and "antimicrobial peptide target sample" are used interchangeably and refer to an antimicrobial peptide sequence obtained after screening an antimicrobial peptide sequence to be screened through the screening framework or system of the present invention, and having properties such as good antimicrobial activity.
[0240] Language Models and Protein Language Models
[0241] As used herein, the term "language model" is a type of machine learning model belonging to the fields of natural language processing and deep learning, and includes large language models (LLMs), etc. The goal of a language model is to predict the next possible characters based on a given context. The technical bases applied by language models include: rule-based methods, statistic-based methods, neural network-based methods, Transformer-based methods, and large-scale pre-trained model-based methods.
[0242] As used herein, the term "Protein Language Models (PLMs)" refers to language models specifically used to process or generate protein sequences. These models learn the latent patterns and rules of protein sequences through pre-training on a large amount of protein sequence data, so as to predict properties, structures, and other attributes of unknown proteins.
[0243] As used herein, the terms "ESM2", "ESM-1b", "ProtBERT", "ProtXLNet", "TAPE", "ProGen2", and "AlphaFold2" are several language models commonly used in protein research.
[0244] Both ESM2 and ESM-1b belong to the ESM large biological models. Among them, ESM2 is a large protein language model based on the Transformer framework developed by Meta AI, which has been pre-trained on hundreds of millions of protein sequences. It consists of multiple layers of self-attention mechanisms and feed-forward neural networks. The self-attention mechanism allows the model to consider the relationships between different positions in the sequence when processing the sequence, while the feed-forward neural network further processes and integrates the above information. The term "ESMFold" is a model for protein structure prediction using the ESM2 model. It can perform end-to-end three-dimensional structure prediction using only a single sequence as input by leveraging the information and representations learned by ESM2, so as to quantify the occurrence of protein structures. The input of ESM2 is an amino acid sequence. By converting the amino acid sequence into a numerical vector and then inputting it into the model for learning and prediction; the output of ESM2 is the three-dimensional structure prediction of the protein, usually represented in the form of atomic coordinates, which describe the spatial positions of each atom in the protein molecule, and thus can be used for further biophysical analysis and molecular simulation. ESM-1b is also a large protein language model based on the Transformer framework. It contains multiple attention layers and is trained using a self-supervised learning method through masked language models. The input of ESM-1b is an amino acid sequence, and the output is the feature representation of each amino acid position, which is used for downstream analysis.
[0245] ProtBERT and ProtXLNet are natural language processing models trained on protein sequences. Among them, ProtBERT uses a self-supervised learning method that combines protein structures with annotations of Gene Ontology, and can make predictions on protein structures, post-translational modifications, and biophysical properties; ProtXLNet uses an autoregressive model to make predictions on proteins.
[0246] ProGen2 is a GPT-like protein language model developed by Salesforce Research. This model is trained on a diverse sequence dataset of over one billion proteins extracted from genomic, metagenomic, and immunological databases, learns the evolutionary distribution of protein sequences, generates new viable sequences, and predicts protein fitness. Its autoregressive mode can enhance the diversity and novelty of protein sequence generation. The term "GPT-like" refers to a neural network-based autoregressive language model that uses a novel sequence-to-sequence model and can avoid the vanishing gradient problem existing in traditional recurrent neural networks when processing long sequence data. ProGen2 contains an attention module, and the attention module is a mechanism in deep learning for locating key tokens.
[0247] AlphaFold2 is a protein structure prediction algorithm developed by DeepMind, which can predict the three-dimensional structure of proteins with extremely high accuracy.
[0248] Supervised fine-tuning
[0249] In the present invention, a fine-tuned protein base model and an antimicrobial peptide candidate sequence are obtained by performing supervised fine-tuning on the protein base model.
[0250] As used herein, the terms "Supervised Fine-tuning (SFT)" and "supervised fine-tuning" can be used interchangeably, and refer to the process of fine-tuning a pre-trained model using labeled data on the basis of the pre-trained model, so that the model can adapt to a specific task or field. Generally, supervised fine-tuning includes the following steps: pre-training, data collection and annotation, supervised fine-tuning, evaluation and optimization.
[0251] In the present invention, the methods for performing the supervised fine-tuning include: Low-Rank Adaptation (LoRA), Prefix-tuning, Prompt-tuning, P-tuning, P-tuning v1, P-tuning v2, Adapter-tuning.
[0252] Preferably, low-rank adaptation is used for supervised fine-tuning. The low-rank adaptation is a parameter-efficient model fine-tuning technique, which realizes model adaptation by adding a low-rank matrix to the weight matrix of the pre-trained model, significantly reducing the number of parameters to be updated, can significantly reduce the debugging cost, and can reduce the risk of overfitting.
[0253] Deep learning model
[0254] As used herein, the term "deep learning model" belongs to the machine learning model, which uses a multi-layer neural network to learn from a large amount of data.
[0255] Common deep learning models include supervised neural networks, such as Recurrent Neural Networks (RNN), Convolutional Neural Networks (CNN), deep neural networks, recursive neural networks, etc., and unsupervised or semi-supervised deep learning models, such as deep generative models, autoencoders, etc. Among them, deep generative models include Generative Adversarial Network (GAN), and autoencoders include Variational Autoencoder (VAE).
[0256] Deep learning models such as RNN and CNN have been widely used in natural language processing and biological sequence analysis. Among them, RNN is a neural network for processing sequence data, which can utilize the temporal or spatial dependence of the sequence; CNN is suitable for processing data with a grid topology structure, such as images or sequence data.
[0257] Deep learning models such as variational autoencoders are often used in data compression and generation tasks. Generative adversarial networks consist of two networks, a generative model and a discriminative model, and generate realistic samples through adversarial learning.
[0258] The term "end-to-end deep learning model" refers to a class of deep learning models that utilize an end-to-end approach. Among them, "end-to-end" is a data transmission method, which means that data is directly transmitted from the sender to the receiver without the need for an intermediate environment to parse and process the data content, ensuring the directness and integrity of the data. The end-to-end deep learning model can be an end-to-end RNN, an end-to-end CNN, or other deep learning models.
[0259] In the present invention, the end-to-end deep learning model is trained on a dataset containing antimicrobial peptides and their minimum inhibitory concentrations, so that the corresponding minimum inhibitory concentration can be directly obtained from the antimicrobial peptide candidate sequence.
[0260] In the present invention, RNN and CNN can be used to replace the ESM2 model for machine learning screening of antimicrobial peptides.
[0261] The scoring system of the present invention
[0262] In the present invention, the scoring system is selected from the following group: supervised learning algorithms, reinforcement learning algorithms, unsupervised learning algorithms, or semi-supervised learning algorithms.
[0263] As used herein, the term "supervised learning algorithm" refers to a class of algorithms used in the process of supervised learning. The supervised learning refers to providing input data and its corresponding label data to the model, and after training, the model accurately finds the optimal mapping relationship between the input data and the label data, so as to predict or classify new unlabeled data. Supervised learning algorithms mainly include classification algorithms and regression algorithms. Classification algorithms are mainly used to output discrete data, while regression algorithms are used to output continuous data.
[0264] As used herein, the term "unsupervised learning algorithm" refers to a class of algorithms used in the process of unsupervised learning. The unsupervised learning refers to classifying unlabeled input data. Unsupervised learning algorithms include clustering algorithms.
[0265] In the present invention, a scoring system is used to evaluate sequence similarity. Preferably, the scoring system employs a supervised learning algorithm or an unsupervised learning algorithm. Preferably, the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the following group: linear model, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), decision tree, neural network, Naive Bayes, Boosting, Random Forest, or a combination thereof. The unsupervised learning algorithm is selected from the following group: spectral clustering, hierarchical clustering, density clustering, K-means clustering, grid clustering, or model clustering. Among them, SVM classifies data points by finding the optimal hyperplane, while Random Forest classifies by constructing multiple decision trees and taking the majority vote.
[0266] Preferably, the scoring system employs the KNN algorithm. The KNN algorithm makes predictions based on the K nearest neighbors around a sample. In this process, a distance metric is needed to calculate the distance between the sample and its neighbors. The terms "distance metric" and "similarity metric" can be used interchangeably because the commonly used method for evaluating the similarity between samples is to calculate the distance.
[0267] Preferably, the calculation for evaluating sequence similarity by the KNN is as follows:
[0268]
[0269] where s_f is the similarity score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric, x is the generated sample, y_i is the center of the i-th functional cluster, and k is the number of clusters.
[0270] Preferably, the distance metric is Euclidean distance, Manhattan distance, and cosine similarity.
[0271] In the present invention, preferably, a supervised learning algorithm is used to predict the minimum inhibitory concentration. The supervised learning algorithm is a regression algorithm, and the regression algorithm is selected from the following group: Naive Bayes, linear regression, KNN regression, Support Vector Machine regression, decision tree regression, neural network regression, Naive Bayes regression, gradient boosting decision tree regression, Random Forest regression, deep forest regression, or extremely randomized tree regression.
[0272] Preferably, the Naive Bayes algorithm is used to construct a Bayesian Classifier to predict the minimum inhibitory concentration. The Bayesian Classifier refers to a probability classifier based on Bayes' theorem, which can perform classification prediction according to the conditional probability of features.
[0273] Antibacterial activity index
[0274] The present invention uses antibacterial activity indicators to characterize the antibacterial activity of antibacterial peptides. By predicting the antibacterial activity of candidate sequences of antibacterial peptides, target sequences of antibacterial peptides with good antibacterial activity are screened out.
[0275] As used herein, the term "Minimum Inhibitory Concentration (MIC)" refers to the lowest concentration of an antibacterial substance that can inhibit the visible growth of microorganisms and is an important indicator for evaluating antibacterial efficacy.
[0276] As used herein, the term "Minimum Bactericidal Concentration (MBC)" refers to the lowest concentration required to kill microorganisms under specific conditions and a fixed extended time (18 to 24 hours), at which the viability of the microorganisms is reduced by 99%.
[0277] As used herein, the term "Time-kill Curve" refers to a curve that describes the change in the killing effect of an antibacterial agent on microorganisms over time and is used to evaluate the bactericidal kinetics of the antibacterial agent.
[0278] The antibacterial peptide generation framework of the present invention and its construction method
[0279] The present invention provides a method for constructing an antibacterial peptide generation framework for forming an antibacterial peptide generation framework, thereby generating candidate antibacterial peptide sequences.
[0280] Specifically, the method includes the steps of: (1) training a protein base model on a data set; (2) performing supervised fine-tuning on the protein base model to obtain a fine-tuned protein base model and candidate antibacterial peptide sequences.
[0281] Preferably, in step (2), it further includes the step of performing reinforcement learning on the fine-tuned protein base model to obtain an optimized protein model and candidate antibacterial peptide sequences.
[0282] The antibacterial peptide generation framework constructed by using the above method includes: (A1) an input unit configured to input data, where the input data includes a protein base model and / or an input sequence, and the protein base model is trained on the input sequence; (A2) a fine-tuning unit configured to execute a fine-tuning model on the input data to obtain fine-tuned input data, where the fine-tuning model is a supervised fine-tuning model; (A3) an output unit configured to output the result of the fine-tuning unit. This framework is "AMPGen".
[0283] In the method for constructing the antibacterial peptide generation framework of the present invention, the dataset for training can be adjusted according to the type of bioactive peptides (antibacterial peptides are a type of bioactive peptides), so as to apply the method to construct more types of bioactive peptide generation frameworks or systems, such as anticancer peptides, antiviral peptides, and cell-penetrating peptides. For example, new anticancer peptides are generated based on the anticancer peptide dataset.
[0284] Antibacterial Peptide Screening Framework and Its Construction Method of the Present Invention
[0285] The present invention provides a method for constructing an antibacterial peptide screening framework for forming an antibacterial peptide screening framework, so as to screen the generated candidate antibacterial peptide sequences to obtain antibacterial peptides with better antibacterial activity.
[0286] Specifically, the method includes the steps:
[0287] (I) Machine learning screening, the machine learning screening includes the steps: (C1) Using a machine learning model, evaluating the sequence similarity between the antibacterial peptide candidate sequence and the target function sample through a scoring system; (C2) Predicting antibacterial activity indicators;
[0288] (II) Posterior verification, the posterior verification includes (D1) restricting the peptide segment length; (D2) predicting structural features; (D3) analyzing similarity; or a combination thereof. Preferably, the posterior verification includes restricting the peptide segment length, predicting structural features, and analyzing similarity.
[0289] The antibacterial peptide screening framework constructed by using the above method includes: (B1) An input unit, the input unit is configured to input data, the input data includes an input sequence to be screened, and the input sequence to be screened includes antibacterial peptide candidate sequences generated by the antibacterial peptide generation framework constructed by the method described in the fourth aspect of the present invention; (B2) A screening unit for antibacterial peptides, the screening unit is configured to execute a screening model for antibacterial peptides, so as to obtain an antibacterial peptide target sequence from the antibacterial peptide candidate sequences; wherein, the screening model is constructed by using the method described in the fifth aspect of the present invention; (B3) An output unit, the output unit is configured to output the screening result. This framework is the "AMPGen-Filtering".
[0290] Preferably, the input sequence to be screened is an antibacterial peptide candidate sequence generated by AMPGen.
[0291] In the method for constructing the antibacterial peptide screening framework or system of the present invention, the screening indicators can be adjusted according to the types and properties of bioactive peptides (antibacterial peptides are a type of bioactive peptide), so as to apply the method to construct screening frameworks or systems for more types of bioactive peptides, such as anticancer peptides, antiviral peptides, and cell-penetrating peptides. For example, for anticancer peptides or antiviral peptides, screening can be carried out according to the killing power against tumor cells or viruses.
[0292] The main advantages of the present invention include:
[0293] (1) According to the content of Invention 1, the present invention discovered 4 novel and highly active antibacterial peptides through integrating multiple computational methods and experimental verification. The above antibacterial peptide sequences have practical application potential, narrowing the gap between prediction and actual effect.
[0294] (2) According to the content of Invention 1, the present invention effectively reduced data-driven bias and captured a wider range of protein sequence knowledge by using large protein language models and fine-tuning techniques.
[0295] (3) According to the content of Invention 1, the present invention proposed a comprehensive screening and analysis method that effectively balances the diversity, novelty, and key antibacterial activity of generated antibacterial peptide sequences, addressing the deficiencies in existing methods.
[0296] (4) According to the steps in the content of Invention 5, the comprehensive evaluation system established by the present invention provides a comprehensive quality control and screening mechanism for generated antibacterial peptide candidates, improving the accuracy and efficiency of screening.
[0297] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The experimental methods without specific conditions noted in the following embodiments are usually carried out under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise specified, percentages and parts are weight percentages and weight parts.
[0298] Methods and steps:
[0299] 1. Supervised Fine-tuning (SFT) of protein language models
[0300] The present invention uses ProGen2 as the baseline generator (base model) and performs supervised fine-tuning using the Low-Rank Adaptation (LoRA) technique. As an efficient adapter-based method, Low-Rank Adaptation can minimize the number of parameters and reduce the risk of overfitting. The present invention introduces rank-based adapters into the attention module of ProGen2 to regulate residue interactions for conditional generation. Perplexity is used to guide the fine-tuning of the model to optimize the generation of peptides that are functionally similar to the training data.
[0301] The loss function for supervised fine-tuning is defined as follows:
[0302]
[0303] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the LoRA parameters.
[0304] 2. Screening and Validation of Antimicrobial Peptide Candidate Samples
[0305] The present invention adopts a two-stage screening process to ensure the quality of the generated candidate peptide segments.
[0306] 2.1 Machine Learning-Based Basic Screening
[0307] The specific implementation steps of the machine learning-based basic screening are as follows:
[0308] (S1) Use the pre-trained protein language model ESM2 to evaluate the similarity of sequences to known antimicrobial peptides in a zero-shot manner.
[0309] (S2) Use a classifier based on the minimum inhibitory concentration to predict antimicrobial activity.
[0310] (S3) Implement the K-nearest neighbor algorithm to screen the generated antimicrobial peptide samples, and use the distance between ESM2 protein sequence embeddings to evaluate similarity.
[0311] (S4) Calculate the screening score. The screening criterion is based on the weighted sum of multiple normalized distance metrics between the generated sample embedding and the center of the target functional sample embedding. The specific calculation formula is as follows:
[0312]
[0313] where s_f is the final similarity screening score, w_m is the weight of each functional cluster, f_m represents the m-th distance metric (Euclidean distance, Manhattan distance, and cosine similarity), x is the generated sample, y_i is the center of the i-th functional cluster, and k is the number of clusters.
[0314] (S5) Further refine the selection process using a Bayesian classifier based on ESM2 embeddings. This classifier is trained on a dataset of AMPs with experimentally determined minimum inhibitory concentration (MIC) values. The training objective is to minimize the L1 loss between the predicted MIC value and the actual MIC value:
[0315]
[0316] where \(y_i\) is the actual MIC value, \(f(x_i)\) is the predicted MIC value of the ESM2 embedding \(x_i\) of the \(i\)-th peptide, and \(N\) is the number of training samples.
[0317] (S6) Further refine the selection of candidate peptides through a cross-validation method. Take the intersection of the two sets of KNN-based screening and MIC prediction to identify peptides that are not only structurally similar to known AMPs but also likely to have good antibacterial activity.
[0318] 2.2 Posterior Verification
[0319] The post-processing steps of the present invention optimize antibacterial peptide candidates through length screening, structure prediction, and similarity analysis. The specific implementation steps are as follows:
[0320] (S1) Length Selection and Folding Stability
[0321] (a) Limit the peptide length to 50 amino acids or less to ensure compatibility with efficient laboratory synthesis techniques.
[0322] (b) Use the protein structure prediction tools AlphaFold2 and ESMFold to evaluate the structural properties of the selected peptides, providing key feature predictions regarding peptide foldability and potential thermal stability.
[0323] (S2) Similarity Analysis
[0324] (a) Perform sequence-level comparisons using BLAST against curated peptide and protein domain databases. In the BLAST analysis evaluation, the following metrics are mainly concerned: highest score, total score, query coverage, E-value, percentage identity, and acceptance length. Priority is given to the E-value and percentage identity to determine sequence similarity. Candidates with an E-value lower than 1e-5 and a percentage identity higher than 30% are marked for further investigation.
[0325] (b) Use FoldSeek for structure-based similarity search. In the FoldSeek analysis and evaluation, the following metrics are mainly concerned: check probability, sequence identity, E-value, score, query position, target position, TM-score, and RMSD. Special emphasis is placed on TM-score and RMSD for structure similarity assessment. Structures with a TM-score higher than 0.5 and an RMSD lower than are considered to have significant structural similarity.
[0326] Through the above steps, the present invention provides a comprehensive method for screening and evaluating antimicrobial peptide candidates, balancing synthetic feasibility and potential structural stability, while providing a framework for interpreting experimental results.
[0327] 3. Detection of Antimicrobial Peptide Properties
[0328] 3.1 Determination of Molar Mass
[0329] Mass spectrometry analysis techniques such as matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS) or liquid chromatography-mass spectrometry (LC-MS) are used. These methods can provide accurate molecular weight information of antimicrobial peptides. In addition, nuclear magnetic resonance (NMR) can also be used for structure analysis and molecular weight confirmation.
[0330] 3.2 Determination of Bacteriostatic Rate
[0331] Spectrophotometry (OD600 measurement) and plate counting method are used for evaluation. First, three kinds of bacteria, Pseudomonas aeruginosa, Escherichia coli, and Staphylococcus aureus, are cultured. Bacteria are inoculated in LB (Luria-Bertani) liquid medium and cultured with shaking at 37 °C and 200 rpm for 12 - 16 hours until the bacterial suspension reaches the logarithmic growth phase (OD600 ≈ 0.5 - 0.6), and then diluted to OD600 ≈ 0.05 with sterile PBS or LB as the experimental bacterial suspension.
[0332] After that, for yeasts and fungi (Saccharomyces cerevisiae, Candida albicans, and other fungi), they are cultured in YPD (Yeast Extract-Peptone-Dextrose) liquid medium for 16 - 24 hours with shaking at 37 °C, and then diluted to OD600 ≈ 0.05 with sterile PBS or YPD.
[0333] In the experiment of measuring the antibacterial rate by spectrophotometry, an experimental group, a positive control group, and a negative control group were established in a 96-well plate. In the experimental group, 100 μL of the diluted bacterial suspension and 100 μL of antibacterial peptide solutions with different concentrations were added to make the final volume 200 μL; in the positive control group, 100 μL of the bacterial suspension and 100 μL of sterile PBS or culture medium were added; in the negative control group, only sterile PBS and culture medium were contained. Subsequently, it was incubated at 37 °C for 16 - 20 hours (for fungi, it can be appropriately extended to 24 - 48 hours). The absorbance at 600 nm (OD600) was measured using a spectrophotometer, and the antibacterial rate was calculated. The calculation formula is:
[0334] Antibacterial rate (%) = (1 - OD600 of the experimental group / OD600 of the control group) × 100%
[0335] If the antibacterial peptide may affect the OD600 reading (such as causing bacterial aggregation or sedimentation), the plate counting method can be combined for verification. In the experiment of measuring the antibacterial rate by the plate counting method, 100 μL of the bacterial suspension after the action of different concentrations of antibacterial peptide was taken, appropriately diluted with sterile PBS, and evenly spread on LB (for bacteria) or YPD (for yeast / fungi) agar plates, and incubated at 37 °C for 16 - 24 hours (for fungi, it can be appropriately extended to 48 hours). Subsequently, the number of colonies (CFU / mL) was counted, and the antibacterial rate was calculated. The calculation formula is:
[0336] Antibacterial rate (%) = (1 - CFU of the experimental group / CFU of the control group) × 100%
[0337] 3.3 Determination of the minimum inhibitory concentration
[0338] In the experiment of determining the minimum inhibitory concentration (MIC), the microbroth dilution method was adopted, which is applicable to both bacteria and fungi. The culture media used in the experiment vary depending on the strains. For bacteria (P. aeruginosa, E. coli, S. aureus), MH (Mueller - Hinton) broth medium was used, while for yeast and fungi (S. cerevisiae, C. albicans and other fungi), RPMI 1640 medium was used (it is recommended to carry out under the conditions of pH 7.0 and containing MOPS buffer).
[0339] The experiment first performs a two-fold dilution of the antimicrobial peptide. In a 96-well plate, starting from the first column, it is gradually diluted step by step. The final concentration range is generally set from 256 μg / mL to 0.125 μg / mL (the concentration gradient can be adjusted according to the activity of the antimicrobial peptide), and 100 μL of the antimicrobial peptide solution is added to each well. Subsequently, the bacterial suspension is added, and the concentration of the bacterial or yeast / fungal suspension is adjusted. The final bacterial concentration is adjusted to 1×106 CFU / mL, while the concentration of yeast / fungi is adjusted to 0.5×10 3 ~104 CFU / mL. 100 μL of the bacterial suspension is added to each well to make the final volume 200 μL.
[0340] In the experiment, a positive control group and a negative control group need to be set up. The positive control group contains only the bacterial suspension and the culture medium, while the negative control group contains only the culture medium and the antimicrobial peptide, without the bacterial suspension. Subsequently, the 96-well plate is incubated at 37 °C. Bacteria are incubated for 16 - 20 hours, while yeast and fungi are incubated for 24 - 48 hours. The judgment standard for MIC is to determine the concentration of the antimicrobial peptide with the lowest complete aseptic growth by visual observation or measuring OD600 with a spectrophotometer, which is the MIC value. For fungi, 2,3,5-triphenyltetrazolium chloride staining (TTC) or XTT staining can be used to enhance the visual judgment.
[0341] To improve the accuracy and repeatability of the experiment, the following optimization measures can be taken. First, use the McFarland standard (such as 0.5 McFarland, corresponding to approximately 1×108 CFU / mL) to adjust the concentration of the bacterial solution to ensure experimental consistency. Second, appropriately extend the incubation time. Since fungi grow slowly, it is recommended to incubate yeast for 24 hours and Candida albicans for 48 hours to improve the accuracy of MIC determination. In addition, each experiment should be repeated at least three times and the average value should be taken to reduce experimental errors. Finally, the minimum bactericidal concentration (MBC) determination can be combined. After the MIC determination, samples are taken from the turbid bacterial solution for plate counting to distinguish between bacteriostatic and bactericidal effects.
[0342] Example 1: Antimicrobial Peptide Sequence Design Framework Based on Supervised Fine-Tuning
[0343] This example relates to an antimicrobial peptide sequence design framework based on supervised fine-tuning, and its process is as Figure 1 shown, including the following steps:
[0344] (S1) Collect antimicrobial peptide data from public datasets;
[0345] (S2) Input the above antimicrobial peptide data into the protein base model and train the model. In this example, the protein base model is ProGen2;
[0346] (S3) Fine-tune the model using the methods and steps in the supervised fine-tuning in 1 to obtain the AMPGen model.
[0347] Example 2: Antimicrobial Peptide Screening Framework Based on Protein Language Model and Bioinformatics Method
[0348] This example relates to an antimicrobial peptide screening framework based on a protein language model and bioinformatics methods, and its process is as Figure 2 shown, including the following steps:
[0349] (C1) Use the AMPGen model obtained in Example 1 to generate amino acid sequences as antimicrobial peptide candidate samples;
[0350] (C2) Input the above antimicrobial peptide candidate samples into the AMPGen-Filtering model for screening, where the screening methods and steps involved in the AMPGen-Filtering model are as shown in the screening and verification of antimicrobial peptide candidate samples in 2.
[0351] Example 3: Antimicrobial Peptide Sequences Designed Based on Supervised Fine-Tuning and Antimicrobial Peptides Obtained from the Antimicrobial Peptide Screening Framework
[0352] This example relates to antimicrobial peptides obtained from an antimicrobial peptide sequence design framework based on supervised fine-tuning and an antimicrobial peptide screening framework. Specifically, a batch of candidate antimicrobial peptides can be obtained through the design framework in Example 1, and the candidate antimicrobial peptides are input into the antimicrobial peptide screening framework for screening and verification, and finally 4 antimicrobial peptides with expected properties and MIC activities are obtained.
[0353] Test the molar mass and MIC of the above 4 antimicrobial peptides. The specific steps are as shown in the detection of antimicrobial peptide properties in 3.
[0354] The results of the molar mass and MIC of the 4 antimicrobial peptides are shown in Table 1-2:
[0355] Table 1. Molar Mass of 4 Antimicrobial Peptides
[0356]
[0357] Table 2. MIC (μg / mL) of 4 Antimicrobial Peptides
[0358]
[0359] It can be seen from the results in Table 2 that:
[0360] ZJAMP006 has the lowest MIC against Staphylococcus aureus (4 μg / mL) and Fusarium graminearum (8 μg / mL), indicating its effectiveness against Gram-positive bacteria and Fusarium graminearum. ZJAMP013 has a certain effect against Staphylococcus aureus (8 μg / mL), but other strains were not tested and further research is needed.
[0361] ZJAMP015 (MIC against Staphylococcus aureus = 32 μg / mL) has poor activity and may require higher concentrations to effectively inhibit bacteria.
[0362] The above results also reflect the need for further optimization for specific strains. For example, for Pseudomonas aeruginosa and Saccharomyces cerevisiae, higher concentrations of antimicrobial peptides may be required to produce good antibacterial effects; for Escherichia coli and Staphylococcus aureus, most antimicrobial peptides have shown good performance at 50 μg / mL, indicating their sensitivity to these antimicrobial peptides.
[0363] The amino acid sequences of the 4 antimicrobial peptides are shown in Table 3.
[0364] Table 3. Amino acid sequences of 4 antimicrobial peptides
[0365]
[0366] Discussion
[0367] In the examples of the design and screening of the antimicrobial peptides of the present invention, more strain tests can be added in the future. For example, determine the MIC of ZJAMP009, ZJAMP013, and ZJAMP015 against Escherichia coli, Candida albicans, and Fusarium graminearum to confirm their antibacterial spectra; or test antimicrobial peptides with lower MIC values, such as ZJAMP006, to further explore their minimum effective concentrations.
[0368] All documents mentioned in the present invention are incorporated herein by reference as if each document was individually incorporated by reference. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of the present application.
Claims
1. An antimicrobial peptide sequence, characterized in that: The antimicrobial peptide is obtained through an antimicrobial peptide generation framework and an antimicrobial peptide screening framework.
2. The antimicrobial peptide sequence according to claim 1, characterized in that The antimicrobial peptide production framework includes: (A1) an input unit, wherein the input unit is configured to input data, wherein the input data includes a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence; (A2) a fine-tuning unit, the fine-tuning unit being configured to execute a fine-tuning model on the input data, thereby obtaining fine-tuned input data; wherein the fine-tuning model comprises a supervised fine-tuning model, the supervised fine-tuning model comprising the steps of: performing supervised fine-tuning on the input data using a fine-tuning method, thereby obtaining fine-tuned input data; (A3) An output unit, wherein the output unit is configured to output the result of the fine-tuning unit.
3. The antimicrobial peptide sequence according to claim 1, characterized in that The antimicrobial peptide screening framework includes: (B1) an input unit, the input unit being configured to input data, the input data comprising an input sequence to be screened; (B2) an antimicrobial peptide screening unit, the screening unit being configured to execute an antimicrobial peptide screening model, thereby obtaining an antimicrobial peptide target sequence from the input sequence to be screened; (B3) An output unit, wherein the output unit is configured to output the screening result.
4. The antimicrobial peptide sequence according to claim 3, characterized in that The screening model comprises the steps of: (I) machine learning screening, the machine learning screening comprising the steps of: (C1) Using a machine learning model, the sequence similarity between the candidate antimicrobial peptide sequence and the target functional sample is evaluated through a scoring system; (C2) predictive indicators of antimicrobial activity; (II) a posteriori validation, the a posteriori validation comprising a step selected from the group consisting of: (D1) limit peptide length; (D2) predicting structural features; (D3) Analyze similarities; or a combination thereof; Wherein, the scoring system is selected from the following group: supervised learning algorithm, reinforcement learning algorithm, unsupervised learning algorithm, or semi-supervised learning algorithm; the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the following group: linear model, K nearest neighbor, random forest, or a combination thereof; the unsupervised learning algorithm is a clustering algorithm, and the clustering algorithm is selected from the following group: density clustering, or K-means clustering; The antibacterial activity index is the minimum inhibitory concentration; The structural features are foldability and thermal stability; Predicting structural features by a method selected from the group consisting of a protein structure prediction tool, or molecular simulation; the protein structure prediction tool is a combination of AlphaFold2 and ESMFold; The similarity is analyzed by a method selected from the group consisting of: a sequence alignment algorithm, or a structure alignment algorithm; the sequence alignment algorithm is BLAST; the structure alignment algorithm is a combination of FoldSeek and TM-align.
5. A drug or a pharmaceutical composition, characterized in that The medicine or pharmaceutical composition contains one or more antimicrobial peptide sequences according to claim 1.
6. Use of the medicine or pharmaceutical composition according to claim 5, characterized in that: The uses include: (E1) preparing a drug for inhibiting the growth of microorganisms; (E2) preparing a kit, the kit further comprising a label or instructions indicating that the kit is used to inhibit the growth of microorganisms.
7. A method for constructing an antimicrobial peptide production framework, characterized in that: The method comprises the steps of: (1) Train a protein pedestal model on the dataset; (2) Performing supervised fine-tuning on the protein base model to obtain a fine-tuned protein base model and an antimicrobial peptide candidate sequence.
8. The method according to claim 7, characterized in that The method for performing the supervised fine-tuning is low-rank adaptation.
9. A method for constructing an antimicrobial peptide screening framework, characterized in that: The method comprises the steps of: (I) machine learning screening, the machine learning screening comprising the steps of: (C1) Using a machine learning model, the sequence similarity between the candidate antimicrobial peptide sequence and the target functional sample is evaluated through a scoring system; (C2) predictive indicators of antimicrobial activity; (II) a posteriori validation, the a posteriori validation comprising a step selected from the group consisting of: (D1) limit peptide length; (D2) predicting structural features; (D3) Analyze similarities; or a combination thereof; Wherein, the scoring system is selected from the following group: supervised learning algorithm, reinforcement learning algorithm, unsupervised learning algorithm, or semi-supervised learning algorithm; the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the following group: linear model, K nearest neighbor, random forest, or a combination thereof; the unsupervised learning algorithm is a clustering algorithm, and the clustering algorithm is selected from the following group: density clustering, or K-means clustering; The antibacterial activity index is the minimum inhibitory concentration; The structural features are foldability and thermal stability; Predicting structural features by a method selected from the group consisting of a protein structure prediction tool, or molecular simulation; the protein structure prediction tool is a combination of AlphaFold2 and ESMFold; The similarity is analyzed by a method selected from the group consisting of: a sequence alignment algorithm, or a structure alignment algorithm; the sequence alignment algorithm is BLAST; the structure alignment algorithm is a combination of FoldSeek and TM-align.
10. An antimicrobial peptide design and screening framework or system, characterized in that: The framework or system includes: (W) an input unit, wherein the input unit is configured to input data, wherein the input data includes a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence; (X) a fine-tuning unit, the fine-tuning unit being configured as a fine-tuning model, the fine-tuning model performing a predetermined fine-tuning on the input data, thereby obtaining a fine-tuning result; wherein the fine-tuning model is a supervised fine-tuning model, the supervised fine-tuning model comprising the steps of: performing supervised fine-tuning on the input data, thereby obtaining fine-tuned input data; (Y) a screening unit, the screening unit being configured to execute a screening model for antimicrobial peptides, the screening model obtaining an antimicrobial peptide target sequence from the fine-tuned input data; wherein the screening model is constructed by the method of claim 9; (Z) an output unit, wherein the output unit is configured to output the result of the screening module.