Antibacterial peptide sequence design framework based on protein language model supervised fine tuning and human feedback reinforcement learning
Through the method of supervised fine-tuning based on protein language model and human feedback reinforcement learning, the design and screening of antimicrobial peptide candidate sequences is optimized, and the problems of low efficiency and insufficient accuracy of antimicrobial peptide discovery and screening in the existing technology are solved, and efficient and accurate antimicrobial peptide design and screening are achieved.
Patent Information
- Application Number
- CN202510211661.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
AI Technical Summary
The existing antimicrobial peptide discovery and screening methods are inefficient and insufficiently accurate, making it difficult to effectively solve the antimicrobial resistance crisis.
An antimicrobial peptide sequence design framework based on supervised fine-tuning and human feedback reinforcement learning is adopted to optimize the design and screening of antimicrobial peptide candidate sequences by training the protein base model and performing supervised fine-tuning and human feedback reinforcement learning.
The efficiency and accuracy of antimicrobial peptide design are improved, the transformation cycle from calculation prediction to practical application is shortened, and 13 new high-active antimicrobial peptide sequences have been discovered, enhancing the efficiency of antimicrobial drug discovery and development.
Smart Images

Figure BDA0005286169020000031 
Figure BDA0005286169020000051 
Figure BDA0005286169020000071
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of bioinformatics, artificial intelligence and drug design. Specifically, the present invention belongs to the technical field of antimicrobial peptide sequence design and screening, and relates to a computational method and system for antimicrobial peptide sequence design and screening using a protein language model, a supervised learning algorithm and a human feedback mechanism. Background Art
[0002] Antimicrobial peptides (AMPs), as a new type of antimicrobial agent, have attracted much attention in response to the increasingly serious antimicrobial resistance (AMR) crisis worldwide. AMPs are a class of short-chain amino acid sequences (usually defined as 12-50 amino acid residues) that are widely present in various organisms and can inhibit microbial growth by interfering with the integrity of microbial cell walls. Compared with traditional antibiotics, AMPs have advantages such as broad-spectrum antimicrobial activity and lower resistance development.
[0003] Traditional AMP discovery methods include: isolation from natural sources, rational design, and high-throughput screening (HTS) of synthetic peptide libraries. However, these methods have significant limitations: natural source isolation is time-consuming and limited by available biodiversity; rational design requires a lot of resources for verification; and HTS is limited by initial peptide library design and infrastructure requirements. These methods often have difficulty in efficiently generating AMPs with target properties and may overlook peptides with unconventional mechanisms of action.
[0004] Traditional AMP screening methods mainly include: isolation from natural sources, rational design and post-screening, and high-throughput screening. These methods have identified a large number of AMPs, but face significant challenges in terms of efficiency and cost.
[0005] With the development of artificial intelligence (AI) technology and computing hardware facilities, AMP discovery and screening methods based on AI and computing can overcome the above limitations to a certain extent, so AI-based AMP discovery and computing-based AMP screening are gradually emerging. At present, AI-based AMP discovery methods mainly include the following categories: AMP identification based on machine learning, AMP discovery based on metagenomics, and de novo AMP generation. At present, computing-based AMP screening methods mainly include the following categories: sequence feature analysis, machine learning methods, deep learning methods, integration of multi-omics data, and virtual screening.
[0006] Although AI-based methods have made some progress in AMP discovery and design, they still face the following major challenges: (1) Conversion from in vitro to in vivo: AMPs predicted by computational methods may face problems such as low bioavailability and poor metabolic stability in practical applications, requiring comprehensive pharmacokinetic and pharmacodynamic studies; (2) Database limitations: Existing AMP databases may not fully capture the diversity of natural antimicrobial sequences, potentially introducing bias; (3) Difficulty in function prediction: Accurately predicting functional properties associated with specific antimicrobial mechanisms or activity spectra remains challenging; (4) Balance issues in generative models: Existing generative models have difficulty balancing sequence diversity and novelty while maintaining key antimicrobial properties; (5) Insufficient understanding of structure-activity relationships (SAR): Existing methods provide limited SAR insights, hindering the systematic exploration of sequence-function correlations, which is critical for optimizing AMP design; (6) Environmental adaptability: In vitro screening may not accurately predict in vivo effects, and many promising AMPs exhibit poor stability under physiological conditions; (7) Rapid microbial evolution: The rapid evolution of microorganisms requires AMP design methods to be able to continuously innovate to cope with emerging drug resistance.
[0007] AMP screening based on computational methods has the following shortcomings: (1) Data quality and bias: Existing AMP databases may be biased and cannot fully represent the diversity of AMPs in nature; (2) Difficulty in specificity prediction: Accurately predicting the specific antimicrobial mechanism or activity spectrum of AMPs remains challenging; (3) High false positive rate: Many computational methods may produce a large number of false positive results during the screening process, increasing the cost of subsequent experimental verification; (4) Lack of novelty: Existing methods may tend to identify sequences similar to known AMPs and ignore potential new AMPs; (5) Conversion from in vitro to in vivo effects: Computationally screened AMPs require comprehensive pharmacokinetic and pharmacodynamic studies to verify their potential as viable drug candidates.
[0008] In summary, there is an urgent need in this field to develop an efficient and highly accurate method for discovering, designing and screening antimicrobial peptides. Summary of the invention
[0009] The purpose of the present invention is to provide an antimicrobial peptide sequence design framework based on supervised fine-tuning of a protein language model and human feedback reinforcement learning.
[0010] The first aspect of the present invention provides a method for constructing an antimicrobial peptide production model, the method comprising the steps of:
[0011] (1) Train the protein pedestal model on the dataset;
[0012] (2) Performing supervised fine-tuning on the protein base model to obtain a fine-tuned protein base model and an antimicrobial peptide candidate sequence.
[0013] In another preferred embodiment, step (2) further comprises a step selected from the following group:
[0014] (2.1) training a predictor on the data set to obtain a pre-trained predictor, wherein the pre-trained predictor outputs a prediction value, inputs the prediction value into a reward function, and uses the reward function to perform human feedback reinforcement learning on the fine-tuned protein base model to obtain an optimized protein model and an antimicrobial peptide candidate sequence;
[0015] (2.2) designing a reward function, and using the reward function to perform human feedback reinforcement learning on the fine-tuned protein base model to obtain an optimized protein model and an antimicrobial peptide candidate sequence;
[0016] or a combination thereof.
[0017] In another preferred embodiment, step (2) further includes the following steps:
[0018] (2.3) Reinforcement learning is performed on the fine-tuned protein base model to obtain an optimized protein model and antimicrobial peptide candidate sequences.
[0019] In another preferred embodiment, in step (1), the dataset is a public antimicrobial peptide dataset.
[0020] In another preferred embodiment, in step (2.1), the data set is a publicly available data set of antimicrobial peptides with low minimum inhibitory concentrations and a publicly available data set of inactive antimicrobial peptides.
[0021] In another preferred embodiment, the data set is obtained from the original data set by a data enhancement method.
[0022] In another preferred embodiment, the data enhancement method is selected from the following group:
[0023] (i) performing sequence variation on the original data set;
[0024] (ii) performing structural perturbation on the original data set;
[0025] (iii) performing conditional generation on the original data set;
[0026] or a combination thereof.
[0027] In another preferred embodiment, the protein base model is a protein language model.
[0028] In another preferred example, the protein language model includes: ProGen2 model, ESM-1b, ProtBERT, ProtXLNet, TAPE.
[0029] In another preferred example, the protein language model is the ProGen2 model.
[0030] In another preferred example, the methods for performing the supervised fine-tuning include: Low-Rank Adaptation (LoRA), Prefix-tuning, Prompt-tuning, P-tuning, P-tuning v1, P-tuning v2, Adapter-tuning.
[0031] In another preferred example, the method for performing the supervised fine-tuning is Low-Rank Adaptation.
[0032] In another preferred example, perplexity is used to guide the supervised fine-tuning.
[0033] In another preferred example, the loss function of the Low-Rank Adaptation is:
[0034]
[0035] where N is the number of sequences in the training batch, T is the length of each sequence, \(x_{i,t}\) represents the \(t\)-th token of the \(i\)-th sequence, \(x_{i,<t}\) represents all tokens before \(t\) in the \(i\)-th sequence, and \(\theta_{LoRA}\) represents the Low-Rank Adaptation parameters.
[0036] In another preferred example, the following methods selected from the group are used for the human feedback reinforcement learning: reinforcement learning algorithms, reinforcement learning frameworks, active learning strategies, or real-time interactions.
[0037] In another preferred example, the reinforcement learning framework is a multi-agent reinforcement learning framework.
[0038] In another preferred example, the fine-tuned protein base model is AMPGen.
[0039] In another preferred example, in step (2.1), the optimized protein model is AMPGen-MIC.
[0040] In another preferred example, in step (2.2), the optimized protein model is AMPGen-property.
[0041] In another preferred example, in the human feedback reinforcement learning, a reinforcement learning algorithm is used to adjust the fine-tuned protein base model.
[0042] In another preferred example, the reinforcement learning algorithms include: proximal policy optimization (PPO), trust region policy optimization (TRPO), dominant actor-critic (A2C), asynchronous actor-critic (A3C), soft actor-critic (SAC), and deep deterministic policy gradient (DDPG).
[0043] In another preferred embodiment, the reinforcement learning algorithm is proximal strategy optimization.
[0044] In another preferred embodiment, the loss function of the proximal strategy optimization is:
[0045] L_PPO-AMP=L_policy+c1·L_value-c2·H(π_θ)
[0046] Among them, L_policy is the policy loss, L_value is the value function loss, H(π_θ) is the policy entropy, and c1 and c2 are weight coefficients.
[0047] In another preferred embodiment, the pre-trained predictor is a minimum inhibitory concentration predictor.
[0048] In another preferred example, the minimum inhibitory concentration predictor is trained on a public antimicrobial peptide dataset with low minimum inhibitory concentration and a public inactive antimicrobial peptide dataset.
[0049] In another preferred embodiment, the reward function is designed using the predicted value of the minimum inhibitory concentration predictor.
[0050] In another preferred embodiment, in (2.1), the reward function is:
[0051] R_h={
[0052] (s-γ)*β, if s<0.5
[0053] 1.0, if s ≥ 0.5
[0054] }
[0055] Wherein, s is the predicted value output by the minimum inhibitory concentration predictor; 0.5 is the threshold used when training the minimum inhibitory concentration predictor; γ and β are thresholds.
[0056] In another preferred embodiment, the γ and β are set to 0.35 and 4 respectively.
[0057] In another preferred embodiment, the reward function is designed using the properties of antimicrobial peptides.
[0058] In another preferred embodiment, the antimicrobial peptide properties are selected from the following group: physicochemical properties, stability, antimicrobial activity, or a combination thereof.
[0059] In another preferred embodiment, the physicochemical property is selected from the following group: hydrophobicity, hydrophobic moment, charge, isoelectric point, toxicity, activity, energy of interaction between antimicrobial peptide and bacterial membrane, or a combination thereof.
[0060] In another preferred embodiment, the physicochemical property is a combination of hydrophobicity, hydrophobic moment, charge and isoelectric point.
[0061] In another preferred embodiment, the physicochemical properties are calculated by a method selected from the following group: quantum chemical calculation method, molecular dynamics simulation, or a combination thereof.
[0062] In another preferred embodiment, the energy of the interaction between the antimicrobial peptide and the bacterial membrane is calculated by the molecular dynamics simulation.
[0063] In another preferred embodiment, the antibacterial activity is predicted by a machine learning model.
[0064] In another preferred embodiment, in (2.2), the reward function is:
[0065] R_p=d1·clamp(p_h,-0.5,0.8)+d2·clamp(p_hm,0.0,0.6)+d3·clamp(p_q,-5.0,9.0)
[0066] +d4 clamp(p_i,8.0,11.0)+d5
[0067] Among them, p_h, p_hm, p_q and p_i represent hydrophobicity, hydrophobic moment, charge and isoelectric point respectively; d1 to d4 are weights; d5 is the normalization factor; clamp is the clamp function, which receives three parameters, namely minimum value, preferred value and maximum value.
[0068] In another preferred embodiment, the reward function integrates multiple properties of the antimicrobial peptides through a multi-objective optimization algorithm.
[0069] In another preferred example, the multi-objective optimization algorithm includes Pareto optimization.
[0070] The second aspect of the present invention provides a framework or system for producing antimicrobial peptides, the framework or system comprising:
[0071] (D1) an input unit, wherein the input unit is configured to input data, wherein the input data includes a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence;
[0072] (D2) Fine-tuning unit, which is configured to execute a fine-tuning model on the input data to obtain fine-tuned input data; wherein, the fine-tuning model includes a supervised fine-tuning model, and the supervised fine-tuning model includes the steps of: performing supervised fine-tuning on the input data by using a fine-tuning method to obtain fine-tuned input data;
[0073] (D3) Output unit, which is configured to output the result of the fine-tuning unit.
[0074] In another preferred example, the protein base model is a protein language model.
[0075] In another preferred example, the protein language model includes: ProGen2 model, ESM-1b, ProtBERT, ProtXLNet, TAPE.
[0076] In another preferred example, the protein language model is the ProGen2 model.
[0077] In another preferred example, the fine-tuning method includes: Low-Rank Adaptation (LoRA), Adapter-tuning, Prefix-tuning, Prompt-tuning, P-tuning, P-tuning v1, P-tuning v2.
[0078] In another preferred example, the fine-tuning method is Low-Rank Adaptation.
[0079] In another preferred example, perplexity is used to guide the supervised fine-tuning model.
[0080] In another preferred example, the loss function of the Low-Rank Adaptation is:
[0081]
[0082] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the Low-Rank Adaptation parameter.
[0083] In another preferred example, the fine-tuning model further includes a model selected from the following group:
[0084] (E1) Active feedback fine-tuning model, which includes the steps of: training a predictor on a dataset to obtain a pre-trained predictor, the pre-trained predictor outputs a prediction value, inputting the prediction value into a reward function, and using the reward function to perform human feedback reinforcement learning on the input data to obtain fine-tuned input data; or
[0085] (E2) An attribute feedback fine-tuning model, wherein the attribute feedback fine-tuning model comprises the steps of: designing a reward function, and using the reward function to perform human feedback reinforcement learning on the input data, thereby obtaining fine-tuned input data.
[0086] In another preferred example, the framework or system includes: an input unit as described in (D1); a fine-tuning unit as described in (D2), wherein the fine-tuning unit is configured to execute a fine-tuning model on the input data to obtain fine-tuned input data; wherein the fine-tuning model includes a supervised fine-tuning model and an active feedback fine-tuning model as described in (E1); and an output unit as described in (D3).
[0087] In another preferred example, the framework or system includes: an input unit as described in (D1); a fine-tuning unit as described in (D2), wherein the fine-tuning unit is configured to execute a fine-tuning model on the input data to obtain fine-tuned input data; wherein the fine-tuning model includes a supervised fine-tuning model and an attribute feedback fine-tuning model as described in (E2); and an output unit as described in (D3).
[0088] In another preferred embodiment, the data set is a publicly available low minimum inhibitory concentration antimicrobial peptide data set and a publicly available inactive antimicrobial peptide data set.
[0089] In another preferred example, in the human feedback reinforcement learning, a reinforcement learning algorithm is used to adjust the micro-input data.
[0090] In another preferred example, the reinforcement learning algorithms include: proximal policy optimization (PPO), trust region policy optimization (TRPO), dominant actor-critic (A2C), asynchronous actor-critic (A3C), soft actor-critic (SAC), and deep deterministic policy gradient (DDPG).
[0091] In another preferred embodiment, the reinforcement learning algorithm is proximal strategy optimization.
[0092] In another preferred embodiment, the loss function of the proximal strategy optimization is:
[0093] L_PPO-AMP=L_policy+c1·L_value-c2·H(π_θ)
[0094] Among them, L_policy is the policy loss, L_value is the value function loss, H(π_θ) is the policy entropy, and c1 and c2 are weight coefficients.
[0095] In another preferred embodiment, the pre-trained predictor is a minimum inhibitory concentration predictor.
[0096] In another preferred example, the minimum inhibitory concentration predictor is trained on a public antimicrobial peptide dataset with low minimum inhibitory concentration and a public inactive antimicrobial peptide dataset.
[0097] In another preferred embodiment, the reward function is designed using the predicted value of the minimum inhibitory concentration predictor.
[0098] In another preferred embodiment, in (E1), the reward function is:
[0099] R_h={
[0100] (s-γ)*β, if s<0.5
[0101] 1.0, if s ≥ 0.5
[0102] }
[0103] Wherein, s is the predicted value output by the minimum inhibitory concentration predictor; 0.5 is the threshold used when training the minimum inhibitory concentration predictor; γ and β are thresholds.
[0104] In another preferred embodiment, the γ and β are set to 0.35 and 4 respectively.
[0105] In another preferred embodiment, the reward function is designed using the properties of antimicrobial peptides.
[0106] In another preferred embodiment, the antimicrobial peptide properties are selected from the following group: physicochemical properties, stability, antimicrobial activity, or a combination thereof.
[0107] In another preferred embodiment, the physicochemical property is a combination of hydrophobicity, hydrophobic moment, charge and isoelectric point.
[0108] In another preferred embodiment, in (E2), the reward function is:
[0109] R_p=d1·clamp(p_h,-0.5,0.8)+d2·clamp(p_hm,0.0,0.6)+d3·clamp(p_q,-5.0,9.0)
[0110] +d4 clamp(p_i,8.0,11.0)+d5
[0111] Among them, p_h, p_hm, p_q and p_i represent hydrophobicity, hydrophobic moment, charge and isoelectric point respectively; d1 to d4 are weights; d5 is the normalization factor; clamp is the clamp function, which receives three parameters, namely minimum value, preferred value and maximum value.
[0112] The third aspect of the present invention provides a method for constructing an antimicrobial peptide screening model, the method comprising the steps of:
[0113] (M1) machine learning screening, the machine learning screening comprising the steps of:
[0114] (A1) Using a machine learning model, the sequence similarity between the candidate antimicrobial peptide sequence and the target functional sample is evaluated through a scoring system;
[0115] (A2) predictive indicators of antimicrobial activity;
[0116] (M2) a posteriori verification, the a posteriori verification comprising a step selected from the group consisting of:
[0117] (C1) limit peptide length;
[0118] (C2) predicting structural features;
[0119] (C3) Analyze similarities;
[0120] or a combination thereof.
[0121] In another preferred embodiment, the method further comprises the step of: (M3) high throughput screening.
[0122] In another preferred embodiment, in step (M1), further comprising:
[0123] (A3) Evaluation of antimicrobial peptide stability;
[0124] (A4) Assessment of structural similarity.
[0125] In another preferred embodiment, in step (A1), the process comprises:
[0126] (a1.1) using the machine learning model to convert the antimicrobial peptide candidate sequence and / or the target function sample into a numerical form;
[0127] (a1.2) Calculating the similarity between the antimicrobial peptide candidate sequence and the target functional sample.
[0128] In another preferred embodiment, the machine learning model is selected from the following group: a protein language model, a deep learning model, or a combination thereof.
[0129] In another preferred example, the machine learning model is a combination of a protein language model and a deep learning model.
[0130] In another preferred example, the protein language model includes: ESM2 model, ProtBERT, and ProtXLNet.
[0131] In another preferred example, the protein language model is the ESM2 model.
[0132] In another preferred example, the ESM2 model evaluates the sequence similarity between the bacterial peptide candidate sequence and the target functional sample through zero-sample learning.
[0133] In another preferred embodiment, the protein language model converts the antimicrobial peptide candidate sequence and / or the target function sample into the numerical form by a method selected from the following group: a natural language processing algorithm, or a sequence alignment algorithm.
[0134] In another preferred example, the natural language processing algorithm is a word vector.
[0135] In another preferred embodiment, the word vector is selected from the following group: one-hot encoding, or word2vec.
[0136] In another preferred embodiment, the word vector is one-hot encoded.
[0137] In another preferred embodiment, the sequence alignment algorithm is a protein substitution scoring matrix.
[0138] In another preferred example, the protein substitution scoring matrix is selected from the following group: PAM matrix, or BLOSUM matrix.
[0139] In another preferred example, the protein substitution scoring matrix is a BLOSUM matrix.
[0140] In another preferred example, the numerical form includes: an embedding vector and a scoring matrix.
[0141] In another preferred embodiment, the numerical form is an embedded vector.
[0142] In another preferred example, the embedding vector is ESM2 embedding.
[0143] In another preferred example, the ESM2 embedding is an embedding vector formed by converting the candidate antimicrobial peptide sequence and the target functional sample using the ESM2 model.
[0144] In another preferred example, the deep learning model includes: a model based on a convolutional neural network (CNN) and a model based on a recurrent neural network (RNN).
[0145] In another preferred example, the deep learning model is a model based on a recurrent neural network (RNN).
[0146] In another preferred embodiment, the scoring system adopts an algorithm selected from the following group: a supervised learning algorithm, a reinforcement learning algorithm, an unsupervised learning algorithm, or a semi-supervised learning algorithm.
[0147] In another preferred embodiment, the scoring system adopts an algorithm selected from the following group: a supervised learning algorithm, or a reinforcement learning algorithm.
[0148] In another preferred example, the supervised learning algorithm is a classification algorithm.
[0149] In another preferred embodiment, the classification algorithm is selected from the following group: linear model, K nearest neighbor (KNN), random forest, support vector machine (SVM), decision tree, neural network, naive Bayes, Boosting, or a combination thereof.
[0150] In another preferred embodiment, the classification algorithm is selected from the following group: linear model, K nearest neighbor, random forest, or a combination thereof.
[0151] In another preferred embodiment, the classification algorithm is K nearest neighbor.
[0152] In another preferred embodiment, the calculation of sequence similarity evaluated by the K nearest neighbors is:
[0153]
[0154] Among them, s_f is the similarity score, w_m is the weight of each functional cluster, f_m represents the mth distance metric, x is the generated sample, y_i is the center of the ith functional cluster, k is the number of clusters, and d is the number of distance metrics.
[0155] In another preferred example, the distance metric is selected from the following group: Euclidean distance, Manhattan distance, cosine similarity, Chebyshev distance, Minkowski distance, standard Euclidean distance, Mahalanobis distance, Hamming distance, Jaccard distance, correlation distance, information entropy, or a combination thereof.
[0156] In another preferred embodiment, the distance metric is Euclidean distance, Manhattan distance and cosine similarity.
[0157] In another preferred embodiment, d=3.
[0158] In another preferred embodiment, the scoring system includes utilizing multiple supervised learning algorithms to obtain multiple similarity scores.
[0159] In another preferred embodiment, the multiple similarity scores are integrated using a method selected from the following group: fuzzy logic, Dempster-Shafer theory, or a multi-objective optimization algorithm.
[0160] In another preferred example, the multi-objective optimization algorithm is NSGA-II.
[0161] In another preferred example, the unsupervised learning algorithm is a clustering algorithm.
[0162] In another preferred embodiment, the clustering algorithm is selected from the following group: density clustering, K-means clustering, spectral clustering, hierarchical clustering, grid clustering, or model clustering.
[0163] In another preferred example, the density clustering is DBSCAN.
[0164] In another preferred embodiment, the stability of antimicrobial peptides is evaluated using a computational model.
[0165] In another preferred embodiment, the computational model includes a computational model for simulating protein degradation.
[0166] In another preferred embodiment, a graph neural network is used to evaluate structural similarity.
[0167] In another preferred embodiment, the antibacterial activity indicators include: minimum inhibitory concentration, minimum bactericidal concentration, and time-killing curve.
[0168] In another preferred embodiment, the antibacterial activity index is the minimum inhibitory concentration.
[0169] In another preferred embodiment, the minimum inhibitory concentration is predicted using a method selected from the following group: a supervised learning algorithm, an expert system, or a deep learning model.
[0170] In another preferred embodiment, the supervised learning algorithm is selected from the following group: a classification algorithm, or a regression algorithm.
[0171] In another preferred example, the classification algorithm is Naive Bayes.
[0172] In another preferred embodiment, the naive Bayes-based classifier is trained on a known data set so that the classifier reaches the training target, and a pre-trained classifier is obtained, so as to predict the minimum inhibitory concentration using the pre-trained classifier.
[0173] In another preferred embodiment, the known data set is an antimicrobial peptide data set with experimentally determined minimum inhibitory concentration values.
[0174] In another preferred example, the naive Bayes-based classifier is a Bayesian classifier.
[0175] In another preferred example, the Bayesian classifier is embedded using the ESM2.
[0176] In another preferred embodiment, the training objective is to minimize the L1 loss between the predicted minimum inhibitory concentration value and the actual minimum inhibitory concentration value, wherein the L1 loss is calculated as:
[0177]
[0178] Among them, y_i is the actual minimum inhibitory concentration value, f(x_i) is the predicted minimum inhibitory concentration value of the numerical form of x_i of the i-th peptide segment, and N is the number of training samples.
[0179] In another preferred embodiment, the regression algorithm is selected from the following group: linear regression, K nearest neighbor regression, support vector machine regression, decision tree regression, neural network regression, naive Bayes regression, boosting regression, random forest regression, deep forest regression, or extreme random tree regression.
[0180] In another preferred embodiment, the Boosting regression is a gradient boosting decision tree regression.
[0181] In another preferred embodiment, the expert system comprises a rule-based expert system.
[0182] In another preferred example, the deep learning model is an end-to-end deep learning model.
[0183] In another preferred embodiment, the end-to-end deep learning model directly identifies the minimum inhibitory concentration from the antimicrobial peptide candidate sequence.
[0184] In another preferred embodiment, the a posteriori verification is to screen the candidate antimicrobial peptide sequences by a method selected from the following group:
[0185] (C1) limit peptide length;
[0186] (C2) predicting structural features; and
[0187] (C3) Analyze similarities.
[0188] In another preferred embodiment, the length of the peptide segment is ≤50 amino acids, preferably ≤30 amino acids, and more preferably ≤25 amino acids.
[0189] In another preferred embodiment, the structural features are foldability and thermal stability.
[0190] In another preferred embodiment, the structural features are predicted by a method selected from the following group: protein structure prediction tools, or molecular simulations.
[0191] In another preferred embodiment, the protein structure prediction tool is selected from the following group: AlphaFold2, ESMFold, I-TASSER, RoseTTAFold, GalaxyTBM, SWISS-MODEL, or a combination thereof.
[0192] In another preferred example, the protein structure prediction tool is a combination of AlphaFold2 and ESMFold.
[0193] In another preferred embodiment, the molecular simulation is molecular dynamics simulation.
[0194] In another preferred embodiment, the similarity includes: sequence similarity and structural similarity.
[0195] In another preferred embodiment, the similarity is analyzed by a method selected from the group consisting of a sequence alignment algorithm, or a structure alignment algorithm.
[0196] In another preferred embodiment, the sequence similarity is analyzed by the sequence alignment algorithm.
[0197] In another preferred embodiment, the structural similarity is analyzed by the structural comparison algorithm.
[0198] In another preferred embodiment, the sequence alignment algorithm is selected from the following group: BLAST, Smith-Waterman algorithm, Needleman-Wunsch algorithm, or a combination thereof.
[0199] In another preferred embodiment, the sequence alignment algorithm is BLAST.
[0200] In another preferred embodiment, the structure alignment algorithm is selected from the following group: FoldSeek, TM-align, DALI, SSAP, FLEXPROT, or a combination thereof.
[0201] In another preferred embodiment, the structure alignment algorithm is a combination of FoldSeek and TM-align.
[0202] In another preferred embodiment, the a posteriori validation further comprises screening antimicrobial peptide candidate sequences by a method selected from the following group: integrating multi-omics data, or knowledge graph technology.
[0203] In another preferred embodiment, the multi-omics data is selected from the group consisting of proteomics, metabolomics, or a combination thereof.
[0204] In another preferred embodiment, the multi-omics data is proteomics.
[0205] In another preferred embodiment, the method further comprises: using an evaluation model to comprehensively analyze multiple indicators to screen candidate antimicrobial peptide sequences.
[0206] In another preferred embodiment, the evaluation model is selected from the following group: AHP, TOPSIS, grey correlation, fuzzy synthesis, entropy weight method, or a combination thereof.
[0207] In another preferred example, the evaluation model is AHP.
[0208] In another preferred embodiment, the indicator is selected from the following group: sequence similarity, structural similarity, structural feature, minimum inhibitory concentration, or a combination thereof.
[0209] In a fourth aspect of the present invention, there is provided an antimicrobial peptide screening framework or system, the framework or system comprising:
[0210] (B1) an input unit, wherein the input unit is configured to input data, wherein the input data includes an input sequence to be screened, wherein the input sequence to be screened includes an antimicrobial peptide candidate sequence generated using the antimicrobial peptide generation framework or system according to the second aspect of the present invention;
[0211] (B2) an antimicrobial peptide screening unit, the screening unit being configured to execute an antimicrobial peptide screening model, thereby obtaining an antimicrobial peptide target sequence from the antimicrobial peptide candidate sequence; wherein the screening model is constructed using the method described in the third aspect of the present invention;
[0212] (B3) An output unit, wherein the output unit is configured to output the screening result.
[0213] In another preferred embodiment, the antimicrobial peptide production framework or system includes: AMPGen, AMPGen-MIC.
[0214] In another preferred embodiment, the antimicrobial peptide production framework or system is AMPGen-MIC.
[0215] In another preferred embodiment, the screening model is AMPGen-Filtering.
[0216] In a fifth aspect, the present invention provides an antimicrobial peptide design and screening framework or system, the framework or system comprising:
[0217] (W) an input unit, wherein the input unit is configured to input data, wherein the input data includes a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence;
[0218] (X) a fine-tuning unit, wherein the fine-tuning unit is configured as a fine-tuning model, wherein the fine-tuning model performs a predetermined fine-tuning on the input data to obtain a fine-tuning result; wherein the fine-tuning model includes:
[0219] (x1) a supervised fine-tuning model, the supervised fine-tuning model comprising the steps of: performing supervised fine-tuning on the input data using a fine-tuning method, thereby obtaining fine-tuned input data;
[0220] (x2) an activity feedback fine-tuning model, the activity feedback fine-tuning model comprising the steps of: training a predictor on a data set to obtain a pre-trained predictor, the pre-trained predictor outputting a prediction value, inputting the prediction value into a reward function, and performing human feedback reinforcement learning on the input data using the reward function, thereby obtaining fine-tuned input data;
[0221] (x3) an attribute feedback fine-tuning model, the attribute feedback fine-tuning model comprising the steps of: designing a reward function, and using the reward function to perform human feedback reinforcement learning on the input data, thereby obtaining fine-tuned input data;
[0222] or a combination thereof;
[0223] The fine-tuning unit is configured to fine-tune a model selected from the group consisting of:
[0224] (X1) supervised fine-tuning model as described in (x1);
[0225] (X2) A combination of the supervised fine-tuning model described in (x1) and the activity feedback fine-tuning model described in (x2); or
[0226] (X3) A combination of the supervised fine-tuning model described in (x1) and the attribute feedback fine-tuning model described in (x3);
[0227] (Y) a screening unit, the screening unit being configured to execute a screening model for antimicrobial peptides, the screening model obtaining an antimicrobial peptide target sequence from the fine-tuned input data; wherein the screening model is constructed using the method described in the third aspect of the present invention;
[0228] (Z) an output unit, wherein the output unit is configured to output the result of the screening module.
[0229] In another preferred embodiment, the fine-tuned input data includes: a fine-tuned input sequence and / or a fine-tuned protein base model.
[0230] In another preferred example, the fine-tuned input sequence includes an amino acid sequence as shown in any one of SEQ ID NOs: 1-13.
[0231] In another preferred embodiment, the antimicrobial peptide target sequence includes an amino acid sequence as shown in any one of SEQ ID NOs: 1-13.
[0232] In another preferred embodiment, the antimicrobial peptide target sequence has an amino acid sequence as shown in any one of SEQ ID NOs: 1-13.
[0233] In another preferred embodiment, the molar mass range of the antimicrobial peptide target sequence is 1400-2800 g / mol.
[0234] In another preferred embodiment, the MIC of the antimicrobial peptide target sequence is less than 32 μg / mL, preferably less than 16 μg / mL, and more preferably less than 4 μg / mL.
[0235] The sixth aspect of the present invention provides a use of a drug or a pharmaceutical composition, wherein the drug or the pharmaceutical composition comprises an antimicrobial peptide sequence, and the antimicrobial peptide sequence is obtained by the framework or system described in the fifth aspect of the present invention, and the use comprises:
[0236] (F1) preparing a formulation for inhibiting the growth of microorganisms;
[0237] (F2) preparing a kit, the kit further comprising a label or instructions indicating that the kit is used to inhibit the growth of microorganisms.
[0238] In another preferred embodiment, the drug or pharmaceutical composition further comprises: the protein family to which the antimicrobial peptide target sequence belongs, and an antimicrobial peptide sequence in the protein family to which the antimicrobial peptide target sequence belongs that has a similar structure or function to the antimicrobial peptide target sequence.
[0239] In another preferred embodiment, the drug or pharmaceutical composition comprises an amino acid sequence as shown in any one of SEQ ID NOs: 1-13.
[0240] In another preferred embodiment, the drug or pharmaceutical composition further comprises: a protein family to which the amino acid sequence shown in any one of SEQ ID NOs: 1-13 belongs, and an antimicrobial peptide sequence in the protein family to which the amino acid sequence shown in any one of SEQ ID NOs: 1-13 belongs that has a similar structure or function to the amino acid sequence shown in any one of SEQ ID NOs: 1-13.
[0241] In another preferred embodiment, the drug or pharmaceutical composition further comprises a pharmaceutically acceptable carrier, diluent or excipient.
[0242] In another preferred embodiment, the antimicrobial peptide target sequence is prepared by direct synthesis.
[0243] In another preferred embodiment, the microorganisms include: bacteria and fungi.
[0244] In another preferred embodiment, the bacteria include: Pseudomonas aeruginosa, Escherichia coli, and Staphylococcus aureus.
[0245] In another preferred embodiment, the fungi include: Saccharomyces cerevisiae and Candida albicans.
[0246] In another preferred embodiment, the preparation comprises a liquid preparation.
[0247] In another preferred embodiment, the drug or pharmaceutical composition comprises an amino acid sequence selected from the following group: an amino acid sequence as shown in SEQ IN NO: 1, an amino acid sequence as shown in SEQ IN NO: 2, an amino acid sequence as shown in SEQ IN NO: 13, or a combination thereof.
[0248] In another preferred embodiment, the antimicrobial peptide sequence is prepared by direct synthesis.
[0249] It should be understood that within the scope of the present invention, the above-mentioned technical features of the present invention and the technical features specifically described below (such as embodiments) can be combined with each other to form a new or preferred technical solution. Due to space limitations, they will not be described one by one here. BRIEF DESCRIPTION OF THE DRAWINGS
[0250] Figure 1 The pipeline of the supervised fine-tuning-based antimicrobial peptide sequence design framework is shown.
[0251] Figure 2 The pipeline of the antimicrobial peptide sequence design framework based on supervised fine-tuning and reinforcement learning with human feedback is shown.
[0252] Figure 3 The process of the antimicrobial peptide screening framework based on protein language model and bioinformatics methods is shown. DETAILED DESCRIPTION
[0253] After extensive and in-depth research, the inventors have designed an antimicrobial peptide sequence design and screening framework for the first time based on supervised fine-tuning of a protein language model and human feedback reinforcement learning. Specifically, the present invention has developed an antimicrobial peptide design and screening model and a fine-tuning method for the model through innovative computational methods and comprehensive analysis processes, effectively solving the problems of low efficiency, insufficient accuracy, single optimization dimension, and lack of flexibility and scalability of the antimicrobial peptide discovery and screening process based on artificial intelligence and computing in the prior art. Thirteen new highly active antimicrobial peptide sequences have been discovered, providing an efficient, accurate, and flexible new method for the research and development of antimicrobial peptides, which is expected to accelerate the discovery and development process of new antimicrobial drugs. On this basis, the present invention has been completed.
[0254] It should be understood that the specific methods and experimental conditions of the present invention are described below in various degrees of detail to provide a substantial understanding of the present invention. The definitions of certain terms used in this specification are provided below. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention belongs.
[0255] the term
[0256] As used herein, the terms "include", "comprising", and "containing" are used interchangeably and include not only closed definitions, but also semi-closed and open definitions. In other words, the terms include "consisting of", "consisting essentially of".
[0257] As used in this article, the term “Reinforcement Learning from Human Feedback (RLHF)” is a machine learning method that guides the model to optimize in a direction that is more consistent with human expectations by integrating feedback from human experts into the reinforcement learning process.
[0258] As used in this article, the terms "baseline generator", "base model", "protein baseline generator" and "protein base model" can be used interchangeably, all of which are Baseline generators, which refer to original machine learning models that can be used for protein research without training or fine-tuning, and can have good target protein generation capabilities after training and fine-tuning.
[0259] As used herein, the term "perplexity" is an evaluation metric for the performance of a language model, which is used to measure the model's prediction accuracy for a given sequence. The lower the perplexity, the better the model performance.
[0260] As used herein, the term "physicochemical property" and "physicochemical property" are interchangeable and are quantitative indicators that describe molecular characteristics. In the present invention, the physicochemical properties of antimicrobial peptides include hydrophobicity, hydrophobic moment, charge and isoelectric point.
[0261] As used in this article, the terms “zero-shot learning” and “zero-shot approach” are used interchangeably and refer to the ability of a model to recognize or classify categories that have not been seen during training.
[0262] As used herein, the term "protein family" is a group of evolutionarily related proteins with similar sequences, structures or functions.
[0263] As used herein, the terms "sequence embedding", "embedding vector", and "embedding" are used interchangeably and refer to converting a protein sequence into a vector representation of a fixed dimension to capture the semantic information of the sequence. In the present invention, the generated sample embedding refers to the embedding vector converted from the candidate antimicrobial peptide sequence; the target function sample embedding refers to the embedding vector converted from the antimicrobial peptide sequence of known function and sequence.
[0264] As used herein, the term "Structure-Activity Relationship (SAR)" is the study that describes the relationship between the structural features of a molecule and its biological activity.
[0265] As used herein, the term "Molecular Dynamics Simulation" is a computational method for studying the motion and interaction of molecules by simulating the evolution of a molecular system over time through a computer.
[0266] As used herein, the terms "antimicrobial peptide MIC value predictor", "minimum inhibitory concentration predictor" and "MIC identifier" are used interchangeably and refer to a classifier that can predict MIC by training based on a known data set containing MIC data during the activity-based feedback fine-tuning stage.
[0267] As used herein, the term "TM-Score (Template Modeling Score)" and "TM score" are used interchangeably and are a score for evaluating protein structural similarity, ranging from 0 to 1, where the closer to 1, the more similar the structures.
[0268] As used herein, the term "RMSD (Root Mean Square Deviation)" refers to the root mean square deviation, which is used to measure the average distance between the atomic spatial positions of two superimposed protein structures, and the unit is angstrom.
[0269] As used herein, the term "multi-objective optimization algorithm" refers to an algorithm that optimizes multiple objective functions simultaneously, such as NSGA-II (Non-dominated Sorting Genetic Algorithm II).
[0270] As used herein, the terms "antimicrobial peptide target sequence" and "antimicrobial peptide target sample" can be used interchangeably, and refer to the antimicrobial peptide sequence to be screened, which is obtained after screening by the screening framework or system of the present invention, and has properties such as good antimicrobial activity.
[0271] AI-based AMP discovery approach
[0272] There are three main AI-based AMP discovery methods: (1) Machine learning-based AMP identification: using complex machine learning algorithms to identify sequences with antimicrobial activity from protein databases. These methods usually integrate multidimensional feature vectors (including physicochemical properties, amino acid composition, and sequence patterns) to build predictive models to distinguish AMPs from non-antimicrobial peptides; (2) Metagenomics-based AMP discovery: using computational algorithms to reveal potential antimicrobial sequences from complex microbial communities. These methods usually combine genome fragment reconstruction and machine learning models to identify potential antimicrobial candidates based on sequence features and structural patterns; (3) De novo AMP generation: using advanced deep learning architectures, mainly generative models, to explore vast sequence spaces and design novel peptides with desired antimicrobial properties. These methods usually use recurrent neural networks or variational autoencoders to generate new candidates by learning potential sequence patterns from existing AMP databases.
[0273] Language Model and Protein Language Model
[0274] As used herein, the term "language model" is a type of machine learning model that belongs to the field of natural language processing and deep learning, including large-scale language models (LLMs). The goal of a language model is to predict the next character that may appear based on a given context. The technical foundations used by language models include: rule-based methods, statistical methods, neural network-based methods, Transformer-based methods, and methods based on large-scale pre-trained models.
[0275] As used herein, the term "Protein Language Models (PLMs)" refers to language models specifically used to process or generate protein sequences. These models are pre-trained on a large amount of protein sequence data to learn the underlying patterns and regularities of protein sequences, thereby predicting properties, structures, and other attributes of unknown proteins.
[0276] As used herein, the terms “ESM2,” “ESM-1b,” “ProtBERT,” “ProtXLNet,” “TAPE,” “ProGen2,” and “AlphaFold2” are several language models commonly used in protein research.
[0277] Both ESM2 and ESM-1b belong to the ESM biological large model. Among them, ESM2 is a large protein language model based on the Transformer framework developed by Meta AI, which has been pre-trained on hundreds of millions of protein sequences. It consists of multiple levels of self-attention mechanisms and feedforward neural networks. The self-attention mechanism allows the model to consider the relationship between each position in the sequence when processing the sequence, while the feedforward neural network further processes and integrates the above information. The term "ESMFold" refers to a model that uses the ESM2 model for protein structure prediction. It uses the information and representation learned by ESM2 and only uses a single sequence as input to perform end-to-end three-dimensional structure prediction to quantify the appearance of protein structure. The input of ESM2 is an amino acid sequence, which is converted into a numerical vector and then input into the model for learning and prediction; the output of ESM2 is a three-dimensional structure prediction of the protein, usually expressed in the form of atomic coordinates, which describe the spatial position of each atom in the protein molecule, so that it can be used for further biophysical analysis and molecular simulation. ESM-1b is also a large protein language model based on the Transformer framework, which contains multiple attention layers and uses a self-supervised learning method to train through a masked language model. The input of ESM-1b is an amino acid sequence, and the output is a feature representation of each amino acid position for downstream analysis.
[0278] ProtBERT and ProtXLNet are natural language processing models trained on protein sequences. ProtBERT uses a self-supervised learning method to combine protein structure with Gene Ontology annotations, and can predict protein structure, post-translational modifications, and biophysical properties; ProtXLNet uses an autoregressive model to predict proteins.
[0279] ProGen2 is a set of GPT-like protein language models developed by Salesforce Research. The model is trained on different sequence data sets of more than 1 billion proteins extracted from genome, metagenome and immune library databases to learn the evolutionary distribution of protein sequences, generate new feasible sequences, and predict protein adaptability. Its autoregressive mode can enhance the diversity and novelty of protein sequence generation. The term "GPT-like" refers to an autoregressive language model based on a neural network, which uses a new sequence-to-sequence model that can avoid the gradient vanishing problem in traditional recurrent neural networks when processing long sequence data. ProGen2 contains an attention module, which is a mechanism for locating key tokens in deep learning.
[0280] AlphaFold2 is a protein structure prediction algorithm developed by DeepMind that can predict the three-dimensional structure of proteins with extremely high accuracy.
[0281] As used in this article, the terms "scaling" and "whitening" are natural language processing methods. Scaling can significantly increase the model capacity of large language models; while whitening is a data preprocessing technique used to reduce the correlation between data and improve its interpretability.
[0282] Supervised fine-tuning
[0283] In the present invention, the protein base model is fine-tuned in a supervised manner to obtain a fine-tuned protein base model and an antimicrobial peptide candidate sequence.
[0284] As used in this article, the terms "supervised fine-tuning (SFT)" and "supervised fine-tuning" are used interchangeably and refer to the process of fine-tuning a pre-trained model using labeled data based on the pre-trained model to adapt the model to a specific task or domain. Typically, supervised fine-tuning includes the following steps: pre-training, data collection and annotation, supervised fine-tuning, evaluation and optimization.
[0285] In the present invention, the method for performing the supervised fine-tuning includes: Low-Rank Adaptation (LoRA), Prefix-tuning, Prompt-tuning, P-tuning, P-tuning v1, P-tuning v2, and Adapter-tuning.
[0286] Preferably, supervised fine-tuning is performed using low-rank adaptation, which is a parameter-efficient model fine-tuning technique that implements model adaptation by adding a low-rank matrix to the weight matrix of the pre-trained model, significantly reducing the number of parameters that need to be updated, significantly reducing debugging costs, and reducing the risk of overfitting.
[0287] Reinforcement Learning Algorithms
[0288] In the present invention, the fine-tuned protein pedestal model is adjusted by a reinforcement learning algorithm.
[0289] Reinforcement learning refers to the process of the environment changing from one state to the next through the interaction between the agent and the environment, that is, taking actions and then obtaining rewards; the agent continuously optimizes its own action strategy to maximize its long-term benefits. Among them, the reward obtained by the agent is quantified by the "reward function", which is a function used to evaluate the quality of actions and guide the model to learn the optimal strategy. The reward function contains weight parameters or weight functions. By adjusting the weight parameters or weight functions, the proportion of each reward component in the reward function can be adjusted to personalize or optimize the reward function, thereby effectively guiding the model to explore; the environment changes from one state to the next through the "transition probability", which can be a deterministic transfer process or a random transfer process.
[0290] Reinforcement learning algorithms include policy gradient, temporal-difference, etc., where the policy gradient refers to learning the optimal policy by directly optimizing the policy function. On the basis of the traditional policy gradient algorithm, KL divergence is introduced to form a natural policy gradient algorithm; conjugate gradient method and line search are further introduced to form a trust region policy optimization algorithm; target divergence or clip function is further set to form a proximal policy optimization algorithm. Policy gradient is combined with temporal difference to form an actor-critic algorithm.
[0291] In the present invention, the reinforcement learning algorithms include: Proximal Policy Optimization (PPO), Trust Region Policy Optimization (TRPO), Deep Deterministic Policy Gradient (DDPG), Advantage Actor-Critic (A2C), Asynchronous Actor-Critic (A3C), and Soft Actor-Critic (SAC).
[0292] Preferably, the fine-tuned protein base model is adjusted using proximal strategy optimization. The proximal strategy optimization is a widely applicable reinforcement learning algorithm developed through algorithm iteration based on the traditional policy gradient algorithm. Proximal strategy optimization stabilizes the training process by limiting the difference between the new and old strategies, thereby improving learning efficiency and robustness.
[0293] Preferably, the present invention constructs a reward function based on the physicochemical properties of antimicrobial peptides:
[0294] R_p=d1·clamp(p_h,-0.5,0.8)+d2·clamp(p_hm,0.0,0.6)+d3·clamp(p_q,-5.0,9.0)
[0295] +d4 clamp(p_i,8.0,11.0)+d5
[0296] Among them, p_h, p_hm, p_q and p_i represent hydrophobicity, hydrophobic moment, charge and isoelectric point respectively; d1 to d4 are weights; d5 is the normalization factor; clamp is the clamp function, which receives three parameters, namely minimum value, preferred value and maximum value.
[0297] Preferably, the present invention constructs a reward function based on the minimum inhibitory concentration:
[0298] R_h={
[0299] (s-γ)*β, if s<0.5
[0300] 1.0, if s ≥ 0.5
[0301] }
[0302] Wherein, s is the predicted value output by the minimum inhibitory concentration predictor; 0.5 is the threshold used when training the minimum inhibitory concentration predictor; and the thresholds γ and β are set to 0.35 and 4, respectively.
[0303] Loss Function
[0304] As used herein, the terms "loss function" and "error function" are used interchangeably and refer to a type of function that calculates the difference between the predicted value and the true value. The performance of the model is evaluated by calculating the difference between the predicted value and the true value. In machine learning, by introducing a loss function in the prediction process, the predicted value can be controlled to be infinitely close to the true value to maximize the accuracy of the model prediction. Therefore, the smaller the value output by the loss function, the better the machine learning effect and the higher the accuracy of the model.
[0305] The loss function is divided into regression loss and classification loss according to the task type. Among them, regression loss mainly deals with continuous variables, such as mean square error (MSE) and mean absolute error (MSE); while classification loss mainly deals with discrete variables, such as cross entropy loss (Cross Entropy Loss) and Dice Loss.
[0306] The loss function is located between the forward propagation and backward propagation of the machine learning model. The forward propagation is when the model generates a prediction based on the input features; the loss function takes the prediction and calculates the difference from the true value; this difference is then used in the backward propagation phase to update the model's parameters and reduce the error of subsequent predictions.
[0307] In the present invention, a loss function for low-rank adaptation is constructed during the supervised fine-tuning process, and the loss function is as follows:
[0308]
[0309] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the low-rank adaptation parameter.
[0310] In the present invention, a loss function for proximal policy optimization is constructed during the reinforcement learning process, and the loss function is as follows:
[0311] L_PPO-AMP = L_policy + c1·L_value - c2·H(π_θ)
[0312] where L_policy is the policy loss, L_value is the value function loss, H(π_θ) is the policy entropy, and c1 and c2 are weight coefficients.
[0313] In the present invention, a loss function based on Naive Bayes is constructed in the screening framework, and the loss function is as follows:
[0314]
[0315] where y_i is the true MIC value, f(x_i) is the predicted MIC value of the ESM2 embedding x_i of the i-th peptide, and N is the number of training samples.
[0316] The scoring system of the present invention
[0317] In the present invention, the scoring system is selected from the group consisting of: supervised learning algorithms, reinforcement learning algorithms, unsupervised learning algorithms, or semi-supervised learning algorithms.
[0318] As used herein, the term "supervised learning algorithm" refers to a class of algorithms used in the process of supervised learning, which refers to providing input data and its corresponding label data to the model, and after training, the model accurately finds the optimal mapping relationship between the input data and the label data, so as to predict or classify new unlabeled data. Supervised learning algorithms mainly include classification algorithms and regression algorithms. Classification algorithms are mainly used to output discrete data, while regression algorithms are used to output continuous data.
[0319] As used herein, the term "unsupervised learning algorithm" refers to a class of algorithms used in the process of unsupervised learning, which refers to classifying unlabeled input data. Unsupervised learning algorithms include clustering algorithms.
[0320] In the present invention, a scoring system is used to evaluate sequence similarity. Preferably, the scoring system uses a supervised learning algorithm or an unsupervised learning algorithm. Preferably, the supervised learning algorithm is a classification algorithm, and the classification algorithm is selected from the following group: linear model, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), decision tree, neural network, naive Bayes, Boosting, random forest, or a combination thereof. The unsupervised learning algorithm is selected from the following group: spectral clustering, hierarchical clustering, density clustering, K-means clustering, grid clustering, or model clustering. Among them, SVM classifies data points by finding the best hyperplane, while random forest classifies by constructing multiple decision trees and taking majority votes.
[0321] Preferably, the scoring system uses a K-nearest neighbor algorithm. The K-nearest neighbor algorithm makes predictions based on the K nearest neighbors around the sample. In this process, a distance metric is needed to calculate the distance between the sample and the neighbors. The terms "distance metric" and "similarity metric" can be used interchangeably because the method usually used to evaluate the similarity between samples is to calculate the distance.
[0322] Preferably, the calculation of sequence similarity assessed by the K nearest neighbors is:
[0323]
[0324] Among them, s_f is the similarity score, w_m is the weight of each functional cluster, f_m represents the mth distance metric, x is the generated sample, y_i is the center of the i-th functional cluster, and k is the number of clusters.
[0325] Preferably, the distance metric is Euclidean distance, Manhattan distance and cosine similarity.
[0326] In the present invention, preferably, a supervised learning algorithm is used to predict the minimum inhibitory concentration, and the supervised learning algorithm is a regression algorithm, and the regression algorithm is selected from the following group: naive Bayes, linear regression, K nearest neighbor regression, support vector machine regression, decision tree regression, neural network regression, naive Bayes regression, gradient boosting decision tree regression, random forest regression, deep forest regression, or extremely random tree regression.
[0327] Preferably, a naive Bayesian algorithm is used to construct a Bayesian classifier to predict the minimum inhibitory concentration. The Bayesian classifier refers to a probabilistic classifier based on the Bayesian theorem, which can perform classification prediction based on the conditional probability of the feature.
[0328] Similarity analysis method in a posteriori verification of the present invention
[0329] In the present invention, similarity is analyzed by a method selected from the group consisting of a sequence alignment algorithm, or a structure alignment algorithm.
[0330] Preferably, the sequence alignment algorithm is selected from the group consisting of BLAST, Smith-Waterman algorithm, Needleman-Wunsch algorithm, or a combination thereof.
[0331] Preferably, the sequence alignment algorithm is BLAST. As used herein, the term "BLAST" is a Basic Local Alignment Search Tool, which is a sequence similarity search algorithm used to compare biological sequences (such as DNA, RNA or protein sequences) with sequence databases to identify similar sequences in the database.
[0332] Preferably, the structural alignment algorithm is selected from the group consisting of FoldSeek, TM-align, DALI, SSAP, FLEXPROT, or a combination thereof.
[0333] Preferably, the structural alignment algorithm is FoldSeek. As used herein, "FoldSeek" is a bioinformatics tool for fast and accurate protein structure search, based on 3D structural similarity for alignment.
[0334] Expert System
[0335] As used in this article, the term "expert system" belongs to the field of artificial intelligence and refers to a computer software system that can solve complex problems in a specific field like a human expert. It can effectively use the experience and professional knowledge accumulated by experts over the years to solve problems that require experts by simulating the expert's thinking process. The expert system needs to use certain knowledge acquisition methods to store expert knowledge in a knowledge base, and then use the inference engine and combine it with the human-computer interaction interface to work.
[0336] Expert systems can be divided into rule-based expert systems, case-based expert systems, artificial neural network-based expert systems, framework-based expert systems, fuzzy logic-based expert systems, fuzzy logic-based expert systems, genetic algorithm-based expert systems, etc. According to the reasoning rules, the expert system based on rules is divided into: expert system based on rules, expert system based on case-based expert systems, expert system based on artificial neural networks, expert system based on framework-based expert systems, expert system based on fuzzy logic, expert system based on genetic algorithms, etc. Among them, the expert system based on rules is composed of five parts: knowledge base, database, reasoning engine, interpretation device and user interface. The knowledge base contains domain knowledge related to problem solving, which is represented by a set of rules and has a structure of conditions and behaviors; the database contains a set of facts for matching the conditions in the knowledge base; the reasoning engine is associated with the rules in the knowledge base and the facts in the database, thereby performing reasoning and enabling the expert system to find a solution; the user interface realizes the communication between the user and the system.
[0337] In the present invention, the minimum inhibitory concentration can be predicted using an expert system. Preferably, the minimum inhibitory concentration is predicted using a rule-based expert system.
[0338] Deep Learning Models
[0339] As used herein, the term “deep learning model” pertains to machine learning models that use multi-layered neural networks to learn from large amounts of data.
[0340] Common deep learning models include supervised neural networks, such as recurrent neural networks (RNN), convolutional neural networks (CNN), deep neural networks, recursive neural networks, etc., and unsupervised or semi-supervised deep learning models, such as deep generative models, autoencoders, etc., among which deep generative models include generative adversarial networks (GAN), and autoencoders include variational autoencoders (VAE).
[0341] Deep learning models such as RNN and CNN have been widely used in natural language processing and biological sequence analysis. Among them, RNN is a neural network that processes sequence data and can utilize the temporal or spatial dependencies of sequences; CNN is suitable for processing data with grid topology structures, such as images or sequence data.
[0342] Deep learning models such as variational autoencoders are often used for data compression and generation tasks. Generative adversarial networks consist of two networks: a generative model and a discriminative model, and generate realistic samples through adversarial learning.
[0343] The term "end-to-end deep learning model" refers to a type of deep learning model that uses an end-to-end approach, where "end-to-end" is a data transmission method that means that data is directly transmitted from the sender to the receiver without the need for an intermediate environment to parse and process the data content, thereby ensuring data directness and integrity. The end-to-end deep learning model can be an end-to-end RNN, end-to-end CNN or other deep learning model.
[0344] In the present invention, the end-to-end deep learning model is trained on a data set containing antimicrobial peptides and their minimum inhibitory concentrations, so that the corresponding minimum inhibitory concentration can be directly obtained from the antimicrobial peptide candidate sequences.
[0345] In the present invention, RNN and CNN can replace the ESM2 model for machine learning screening of antimicrobial peptides.
[0346] Antimicrobial activity index
[0347] The present invention uses antimicrobial activity indicators to characterize the antimicrobial activity of antimicrobial peptides, and screens out antimicrobial peptide target sequences with good antimicrobial activity by predicting the antimicrobial activity of antimicrobial peptide candidate sequences.
[0348] As used herein, the term "minimum inhibitory concentration (MIC)" refers to the lowest concentration of an antibacterial substance that can inhibit the visible growth of microorganisms and is an important indicator for evaluating antibacterial efficacy. The smaller the minimum inhibitory concentration, the stronger the antibacterial ability.
[0349] As used herein, the term "Minimum Bactericidal Concentration (MBC)" refers to the lowest concentration required to kill microorganisms under specific conditions and for a fixed extended period of time (18 to 24 hours), at which concentration the viability of the microorganisms is reduced by 99%.
[0350] As used herein, the term "time-kill curve" refers to a curve that describes the change in the killing effect of an antimicrobial agent on microorganisms over time, and is used to evaluate the killing kinetics of the antimicrobial agent.
[0351] Antimicrobial peptide production framework or system of the present invention and construction method thereof
[0352] The present invention provides an antimicrobial peptide generation framework or system and a construction method thereof, which are used to generate candidate antimicrobial peptide sequences.
[0353] Specifically, the method comprises the steps of: (1) training a protein base model on a data set; and (2) fine-tuning the protein base model in a supervised manner to obtain a fine-tuned protein base model and an antimicrobial peptide candidate sequence.
[0354] Preferably, step (2) also includes the following steps: (2.1) training a predictor on a data set to obtain a pre-trained predictor, the pre-trained predictor outputs a prediction value, the prediction value is input into a reward function, and the fine-tuned protein base model is subjected to human feedback reinforcement learning using the reward function to obtain an optimized protein model and an antimicrobial peptide candidate sequence; (2.2) designing a reward function, and the fine-tuned protein base model is subjected to human feedback reinforcement learning using the reward function to obtain an optimized protein model and an antimicrobial peptide candidate sequence.
[0355] The framework or system includes: (D1) an input unit, the input unit is configured to input data, the input data includes a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence; (D2) a fine-tuning unit, the fine-tuning unit is configured to execute a fine-tuning model on the input data to obtain fine-tuned input data; wherein the fine-tuning model includes a supervised fine-tuning model; (D3) an output unit, the output unit is configured to output the result of the fine-tuning unit. The framework or system is "AMPGen".
[0356] Preferably, the unit (D2) also includes the following models: (E1) an activity feedback fine-tuning model, the activity feedback fine-tuning model comprising the steps of: training a predictor on a data set to obtain a pre-trained predictor, the pre-trained predictor outputting a predicted value, inputting the predicted value into a reward function, and using the reward function to perform human feedback reinforcement learning on the input data to obtain fine-tuned input data; (E2) an attribute feedback fine-tuning model, the attribute feedback fine-tuning model comprising the steps of: designing a reward function, and using the reward function to perform human feedback reinforcement learning on the input data to obtain fine-tuned input data.
[0357] Preferably, the framework or system comprises: an input unit as described in (D1); a fine-tuning unit as described in (D2), wherein the fine-tuning unit is configured to execute a fine-tuning model on the input data to obtain fine-tuned input data; wherein the fine-tuning model comprises a supervised fine-tuning model and an active feedback fine-tuning model as described in (E1); and an output unit as described in (D3). The framework or system is "AMPGen-MIC".
[0358] Preferably, the framework or system comprises: an input unit as described in (D1); a fine-tuning unit as described in (D2), wherein the fine-tuning unit is configured to execute a fine-tuning model on the input data to obtain fine-tuned input data; wherein the fine-tuning model comprises a supervised fine-tuning model and a property feedback fine-tuning model as described in (E2); and an output unit as described in (D3). The framework or system is "AMPGen-property".
[0359] In the method for constructing the antimicrobial peptide generation framework or system of the present invention, the data set used for training, and the reward function in feedback fine-tuning, etc. can be adjusted according to the type and properties of the bioactive peptide (antimicrobial peptides are a type of bioactive peptides), so that the method can be applied to construct more types of bioactive peptide generation frameworks or systems, such as anticancer peptides, antiviral peptides, and cell penetrating peptides. For example, for anticancer peptides or antiviral peptides, the reward function can be designed using the predicted value of the lethality predictor for tumor cells or viruses.
[0360] Antimicrobial peptide screening framework or system of the present invention and construction method thereof
[0361] The present invention provides an antimicrobial peptide screening framework or system and a construction method thereof, which are used to screen generated candidate antimicrobial peptide sequences to obtain antimicrobial peptides with good antimicrobial activity.
[0362] Specifically, the method comprises the steps of:
[0363] (M1) machine learning screening, the machine learning screening comprising the steps of: (A1) using a machine learning model to evaluate the sequence similarity between the candidate antimicrobial peptide sequence and the target functional sample through a scoring system; (A2) predicting antimicrobial activity indicators;
[0364] (M2) a posteriori validation, the a posteriori validation comprising (C1) limiting peptide length; (C2) predicting structural features; (C3) analyzing similarity; or a combination thereof. Preferably, the a posteriori validation comprises limiting peptide length, predicting structural features, and analyzing similarity.
[0365] The framework or system comprises: (B1) an input unit, the input unit is configured to input data, the input data comprises an input sequence to be screened, the input sequence to be screened comprises an antimicrobial peptide candidate sequence generated by the antimicrobial peptide generation framework or system according to the second aspect of the present invention; (B2) an antimicrobial peptide screening unit, the screening unit is configured to execute an antimicrobial peptide screening model, thereby obtaining an antimicrobial peptide target sequence from the antimicrobial peptide candidate sequence; wherein the screening model is constructed using the method according to the third aspect of the present invention; (B3) an output unit, the output unit is configured to output a screening result. The framework or system is "AMPGen-Filtering".
[0366] The input sequence for screening is an antimicrobial peptide candidate sequence generated by a protein model. Preferably, the protein model includes AMPGen and AMPGen-MIC. Preferably, the protein model is AMPGen-MIC.
[0367] In the construction method of the antimicrobial peptide screening framework or system of the present invention, the index used for screening can be adjusted according to the type and property of the bioactive peptide (antimicrobial peptide is a type of bioactive peptide), so as to be applied to the generation of more types of bioactive peptides, such as anticancer peptides, antiviral peptides, and cell penetrating peptides. For example, for anticancer peptides or antiviral peptides, screening can be performed based on the lethality to tumor cells or viruses. Accordingly, the framework or system constructed using the adjusted screening method can also be used for the screening of bioactive peptides.
[0368] In the method for constructing an antimicrobial peptide screening framework or system of the present invention, the index used for screening can be adjusted according to the type and property of the bioactive peptide (antimicrobial peptide is a type of bioactive peptide), so that the method can be applied to construct more types of bioactive peptide screening frameworks or systems, such as anticancer peptides, antiviral peptides, and cell penetrating peptides. For example, for anticancer peptides or antiviral peptides, screening can be performed based on the lethality to tumor cells or viruses.
[0369] The main advantages of the present invention include:
[0370] (1) According to the steps in Invention Content 1, the present invention improves the efficiency and accuracy of AMP design and shortens the conversion cycle from computational prediction to practical application by combining a large protein language model and human feedback reinforcement learning technology.
[0371] (2) According to step (2) in the invention content 1, the antimicrobial peptide design sequence framework of the present invention adopts low-rank adaptation technology to fine-tune the model, which significantly reduces the computing resource requirements, so that antimicrobial peptide design can be efficiently performed on ordinary hardware, greatly reducing research costs and improving computing efficiency.
[0372] (3) According to steps (2.1) and (2.2) in Invention Content 1, the present invention proposes an attribute-based feedback fine-tuning method, which overcomes the limitation of the prior art that only focuses on a single or a few attributes. By comprehensively considering multiple key physicochemical properties of antimicrobial peptides, multi-dimensional optimization of antimicrobial peptide sequences is achieved, thereby enabling a more comprehensive design of antimicrobial peptides that meet actual application requirements.
[0373] (4) According to steps (2.1) and (2.2) in the content of the invention 1, the antimicrobial peptide design sequence framework of the present invention has high flexibility and scalability, and can adapt to the design requirements of different types of antimicrobial peptides by adjusting the weight coefficients in the reward function. At the same time, the framework can integrate new attributes or evaluation criteria, so that it can be continuously updated with the advancement of scientific cognition.
[0374] (5) According to the steps in Invention Content 3, the present invention provides a more comprehensive method to understand the structure-function relationship of antimicrobial peptides by combining sequence-level and structure-level similarity analysis, thereby deepening the understanding of the structure-function relationship.
[0375] (6) Through the antimicrobial peptide design sequence framework and screening framework of the present invention, 13 new highly active antimicrobial peptide sequences were discovered for the first time. After experimental verification, the above antimicrobial peptides have good minimum inhibitory concentration values.
[0376] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The experimental methods without specific conditions noted in the following embodiments are generally carried out under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or according to the conditions recommended by the manufacturer. Unless otherwise specified, percentages and parts are weight percentages and weight parts.
[0377] Methods and steps:
[0378] 1. Fine-tuning of protein language models
[0379] The present invention adopts a three-stage fine-tuning process: supervised fine-tuning, attribute-based feedback fine-tuning, and activity-based feedback fine-tuning. This method makes full use of the advantages of large language models, and at the same time guides through human feedback to output antibacterial peptides that meet strict functional requirements, which can optimize computational efficiency and model performance.
[0380] 1.1 Supervised Fine-tuning (SFT)
[0381] The present invention uses ProGen2 as the baseline generator (base model). In the first stage, the Low-Rank Adaptation (LoRA) technique is used. As an efficient adapter-based method, it can minimize the number of parameters and reduce the risk of overfitting. The present invention introduces a rank-based adapter into the attention module of ProGen2 to regulate residue interactions for conditional generation. Perplexity is used to guide the fine-tuning of the model to optimize the generation of peptides similar to the training data in function.
[0382] The loss function of supervised fine-tuning is defined as follows:
[0383]
[0384] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the LoRA parameters.
[0385] 1.2 Fine-tuning based on human feedback reinforcement learning
[0386] The second and third phases optimized the key properties of the antimicrobial peptides using an online reinforcement learning technique inspired by reinforcement learning with human feedback (RLHF). The present invention adapted this approach to use peptide prediction scores based on identification software as a reward function, thereby achieving property-based optimization without the need for extensive manual evaluation.
[0387] 1.2.1 Attribute-based feedback fine-tuning
[0388] In property-based feedback tuning, the optimization objective is formulated as a weighted combination of key physicochemical properties:
[0389] R_p=d1·clamp(p_h,-0.5,0.8)+d2·clamp(p_hm,0.0,0.6)+d3·clamp(p_q,-5.0,9.0)
[0390] +d4 clamp(p_i,8.0,11.0)+d5
[0391] Among them, p_h, p_hm, p_q, and p_i represent hydrophobicity, hydrophobic moment, charge, and isoelectric point, respectively. The clamp function restricts these properties to a specified range to prevent the generation of antimicrobial peptides with undesirable properties.
[0392] Coefficients d1 to d4 act as weights, defining the relative importance of each physicochemical property in the reward function. By adjusting these weights, the model can be effectively guided to explore the antimicrobial peptide sequence space and balance the importance of different properties. The d5 term acts as a normalization factor to shift the baseline of the reward close to zero to improve training stability and increase the interpretability of the reward signal between different training stages and model comparisons.
[0393] 1.2.2 Activity-based Feedback Fine-tuning
[0394] The third stage implements policy-based fine-tuning, using a mixed reward design to guide the protein language model from the first stage to align with the functional clusters. The reward function is based on the output of the minimum inhibitory concentration (MIC) predictor:
[0395] R_h={
[0396] (s-γ)*β, if s<0.5
[0397] 1.0, if s ≥ 0.5
[0398] }
[0399] Among them, s is the predicted value of the antimicrobial peptide MIC value predictor f_mic. The antimicrobial peptide MIC value predictor is trained on a general antimicrobial peptide cluster and a low MIC cluster to distinguish low MIC antimicrobial peptides from general antimicrobial peptides. The 0.5 threshold is consistent with the training of f_mic, while the thresholds γ and β are set to 0.35 and 4 respectively to ensure the stability of the policy gradient.
[0400] This piecewise function provides a progressive learning signal for sequences with poor performance while maintaining a constant maximum reward for high-performance sequences. It promotes exploration in the lower performance range and prevents reward inflation for the best sequences, effectively balancing exploration and exploitation during the fine-tuning process.
[0401] 1.3 Training details
[0402] In the first stage, the present invention adopts the low-rank adaptation (LoRA) technique to efficiently update the model while maintaining its general knowledge of protein sequences.
[0403] The loss function for supervised fine-tuning is defined as:
[0404]
[0405] where N is the number of sequences in the training batch, T is the length of each sequence, x_i,t represents the t-th token of the i-th sequence, x_i,<t represents all tokens before t in the i-th sequence, and θ_LoRA represents the LoRA parameters.
[0406] In the second and third stages, the present invention adopts the Proximal Policy Optimization (PPO) algorithm to fine-tune the protein language model to generate peptide sequences with potential antibacterial activity. The PPO-AMP loss function is as follows:
[0407] L_PPO-AMP = L_policy + c1·L_value - c2·H(π_θ)
[0408] where L_policy is the policy loss, L_value is the value function loss, H(π_θ) is the policy entropy, and c1 and c2 are weight coefficients.
[0409] To improve the training stability and efficiency, the present invention performs a series of processing steps on the rewards, including scaling, normalization, and whitening.
[0410] 2. Screening and verification of antimicrobial peptide candidate samples
[0411] The present invention adopts a two-stage screening process to ensure the quality of the generated candidate peptide segments.
[0412] 2.1 Machine learning-based basic screening
[0413] The specific implementation steps of machine learning basic screening are as follows:
[0414] (S1) The pre-trained protein language model ESM2 was used to evaluate the similarity of sequences with known antimicrobial peptides in a zero-shot manner.
[0415] (S2) Predict antibacterial activity using a minimum inhibitory concentration-based classifier.
[0416] (S3) K nearest neighbor algorithm was implemented to screen the generated antimicrobial peptide samples, and the distance between ESM2 protein sequence embeddings was used to evaluate the similarity.
[0417] (S4) Calculate the screening score. The screening criterion is based on the weighted sum of multiple normalized distance metrics between the generated sample embedding and the target function sample embedding center. The specific calculation formula is as follows:
[0418]
[0419] Among them, s_f is the final similarity screening score, w_m is the weight of each functional cluster, f_m represents the mth distance metric (Euclidean distance, Manhattan distance and cosine similarity), x is the generated sample, y_i is the center of the i-th functional cluster, and k is the number of clusters.
[0420] (S5) The selection process was further refined using a Bayesian classifier based on ESM2 embedding. The classifier was trained on a dataset of AMPs with experimentally determined minimum inhibitory concentration (MIC) values. The training objective was to minimize the L1 loss between the predicted MIC value and the actual MIC value:
[0421]
[0422] Where y_i is the actual MIC value, f(x_i) is the predicted MIC value of the ESM2 embedding x_i of the ith peptide, and N is the number of training samples.
[0423] (S6) The candidate peptide selection was further refined by a cross-validation approach. The intersection of the two sets of KNN-based screening and MIC prediction was taken to identify peptides that are not only structurally similar to known AMPs but also likely to have good antimicrobial activity.
[0424] 2.2 A posteriori verification
[0425] The post-processing step of the present invention optimizes the antimicrobial peptide candidates through length screening, structure prediction and similarity analysis. The specific implementation steps are as follows:
[0426] (S1) Length selection and folding stability
[0427] (a) Limit peptide length to 50 amino acids or less to ensure compatibility with efficient laboratory synthesis techniques.
[0428] (b) The protein structure prediction tools AlphaFold2 and ESMFold were used to evaluate the structural properties of the selected peptides and provide predictions of key features regarding the foldability and potential thermal stability of the peptides.
[0429] (S2) Similarity analysis
[0430] (a) Sequence-level comparisons were performed using BLAST against a curated database of peptides and protein domains. In the evaluation of BLAST analysis, the following metrics were of interest: highest score, total score, query coverage, E-value, percent identity, and acceptance length. E-value and percent identity were prioritized for determining sequence similarity. Candidates with an E-value lower than 1e-5 and a percent identity higher than 30% were flagged for further investigation.
[0431] (b) Use FoldSeek to perform structure-based similarity search. In the FoldSeek analysis evaluation, the following indicators are mainly focused on: inspection probability, sequence identity, E value, score, query position, target position, TM score and RMSD. Special emphasis is placed on TM score and RMSD for structural similarity evaluation. TM score higher than 0.5 and RMSD lower than The structures of the two species are considered to have significant structural similarities.
[0432] Through the above steps, the present invention provides a comprehensive method to screen and evaluate antimicrobial peptide candidates, balancing synthetic feasibility and potential structural stability, while providing a framework for interpreting experimental results.
[0433] 3. Detection of antimicrobial peptide properties
[0434] 3.1 Determination of molar mass
[0435] Using mass spectrometry techniques, such as matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS) or liquid chromatography-mass spectrometry (LC-MS), these methods can provide accurate molecular weight information of antimicrobial peptides. In addition, nuclear magnetic resonance (NMR) can also be used for structural elucidation and molecular weight confirmation.
[0436] 3.2 Determination of antibacterial rate
[0437] The evaluation was performed using spectrophotometry (OD600 determination) and plate count methods. First, three bacteria, Pseudomonas aeruginosa, Escherichia coli, and Staphylococcus aureus, were cultured. The bacteria were inoculated in LB (Luria-Bertani) liquid medium and cultured at 37°C and 200 rpm for 12-16 hours to allow the bacterial solution to reach the logarithmic growth phase (OD600≈0.5-0.6), and then diluted with sterile PBS or LB to OD600≈0.05 as the experimental bacterial suspension.
[0438] Afterwards, for yeast and fungi (Saccharomyces cerevisiae, Candida albicans and other fungi), culture them in YPD (Yeast Extract-Peptone-Dextrose) liquid medium for 16-24 hours, shaking at 37°C and 200 rpm, and then dilute them with sterile PBS or YPD to OD600≈0.05.
[0439] In the experiment of determining the inhibition rate by spectrophotometry, the experimental group, positive control group and negative control group were set up in 96-well plates. 100 μL of diluted bacterial suspension and 100 μL of antimicrobial peptide solution of different concentrations were added to the experimental group to make the final volume 200 μL; 100 μL of bacterial suspension and 100 μL of sterile PBS or culture medium were added to the positive control group; the negative control group only contained sterile PBS and culture medium. Then incubate at 37°C for 16-20 hours (fungi can be appropriately extended to 24-48 hours), use a spectrophotometer to measure the absorbance at 600nm (OD600), and calculate the inhibition rate. The calculation formula is:
[0440] Antibacterial rate (%) = (1- OD600 of experimental group / OD600 of control group) × 100%
[0441] If the antimicrobial peptide may affect the OD600 reading (such as causing bacterial aggregation or sedimentation), it can be verified in combination with the plate count method. In the experiment of determining the inhibition rate by the plate count method, 100 μL of the bacterial suspension after the action of different concentrations of antimicrobial peptides is taken, appropriately diluted with sterile PBS, and evenly spread on LB (bacteria) or YPD (yeast / fungus) agar plates, and incubated at 37°C for 16-24 hours (fungi can be appropriately extended to 48 hours). Then count the number of colonies (CFU / mL) and calculate the inhibition rate. The calculation formula is:
[0442] Antibacterial rate (%) = (1-experimental group CFU / control group CFU) × 100%
[0443] 3.3 Determination of minimum inhibitory concentration
[0444] In the minimum inhibitory concentration (MIC) determination experiment, the Broth Microdilution Method was used, which is applicable to bacteria and fungi. The culture gene species used in the experiment are different. Bacteria (P. aeruginosa, E. coli, S. aureus) use MH (Mueller-Hinton) broth medium, while yeast and fungi (S. cerevisiae, C. albicans and other fungi) use RPMI 1640 medium (recommended under pH 7.0, MOPS buffer).
[0445] The experiment first diluted the antimicrobial peptide twice, and in a 96-well plate, it was gradually diluted from the first column. The final concentration range was generally set at 256 μg / mL to 0.125 μg / mL (the concentration gradient can be adjusted according to the activity of the antimicrobial peptide), and 100 μL of antimicrobial peptide solution was added to each well. Then, the bacterial suspension was added to adjust the concentration of the bacterial or yeast / fungal suspension. The final bacterial concentration of the bacteria was adjusted to 1×106 CFU / mL, and the concentration of the yeast / fungus was adjusted to 0.5×10 3 ~104 CFU / mL. Add 100 μL of bacterial suspension to each well to make the final volume 200 μL.
[0446] In the experiment, a positive control group and a negative control group need to be set up. The positive control group contains only bacterial suspension and culture medium, while the negative control group contains only culture medium and antimicrobial peptides without bacterial suspension. The 96-well plate is then incubated at 37°C for 16-20 hours for bacteria and 24-48 hours for yeast and fungi. The MIC judgment standard is to determine the minimum antimicrobial peptide concentration for complete sterile growth by visual observation or spectrophotometer measurement of OD600, which is the MIC value. For fungi, 2,3,5-triphenyltetrazolium staining (TTC) or XTT staining can be used to enhance visual judgment.
[0447] In order to improve the accuracy and repeatability of the experiment, the following optimization measures can be taken. First, the concentration of the bacterial solution is adjusted using the McFarland standard (such as 0.5McFarland, corresponding to approximately 1×108CFU / mL) to ensure experimental consistency. Secondly, the incubation time should be appropriately extended. Due to the slow growth of fungi, it is recommended to incubate yeast for 24 hours and Candida albicans for 48 hours to improve the accuracy of MIC determination. In addition, each experiment should be repeated at least three times, and the average value should be taken to reduce experimental errors. Finally, the minimum bactericidal concentration (MBC) determination can be combined with the plate count to distinguish between antibacterial and bactericidal effects by taking samples from the turbid bacterial solution after the MIC determination.
[0448] Example 1: Antimicrobial peptide sequence design framework based on supervised fine-tuning
[0449] This embodiment relates to an antimicrobial peptide sequence design framework based on supervised fine-tuning, and its process is as follows: Figure 1 As shown, the following steps are included:
[0450] (S1) Collect antimicrobial peptide data from public datasets;
[0451] (S2) inputting the antimicrobial peptide data into a protein base model to train the model. In this embodiment, the protein base model is ProGen2;
[0452] (S3) Fine-tune the model using the methods and steps in 1.1 supervised fine-tuning to obtain the AMPGen model.
[0453] Example 2: Antimicrobial peptide sequence design framework based on supervised fine-tuning and human feedback reinforcement learning
[0454] This embodiment relates to an antimicrobial peptide sequence design framework based on supervised fine-tuning and human feedback reinforcement learning, and its process is as follows: Figure 2 As shown, the fine-tuning includes two stages, wherein the fine-tuning of the first stage includes the following steps:
[0455] (A1) Collect low MIC value antimicrobial peptide data from public datasets;
[0456] (A2) inputting the above antimicrobial peptide data into the AMPGen model obtained in Example 1 and adjusting the model;
[0457] (A3) Fine-tune the model using the method and steps in 1.2.1 attribute-based feedback fine-tuning, wherein the attributes include four physicochemical attributes of the antimicrobial peptide: hydrophobicity, hydrophobic moment, charge, and isoelectric point. After fine-tuning, the AMPGen-property model is obtained.
[0458] The second stage of fine-tuning includes the following steps:
[0459] (B1) Collect low MIC value antimicrobial peptide data and inactive antimicrobial peptide data from public datasets;
[0460] (B2) inputting the above low MIC value antimicrobial peptide data and inactive antimicrobial peptide data into an antimicrobial peptide MIC value predictor to train the predictor;
[0461] (B3) Using the method and steps in 1.2.2 Activity-based feedback fine-tuning, the AMPGen model obtained in Example 1 was aligned with the MIC functional cluster using the antimicrobial peptide MIC value predictor, thereby fine-tuning the model. After fine-tuning, the AMPGen-MIC model was obtained.
[0462] Example 3: Antimicrobial peptide screening framework based on protein language model and bioinformatics method
[0463] This embodiment relates to an antimicrobial peptide screening framework based on a protein language model and a bioinformatics method, and its process is as follows: Figure 3 As shown, the following steps are included:
[0464] (C1) using the AMPGen-MIC model obtained in Example 2 to generate amino acid sequences as candidate antimicrobial peptide samples;
[0465] (C2) The above antimicrobial peptide candidate samples are input into the AMPGen-Filtering model for screening, wherein the screening method and steps involved in the AMPGen-Filtering model are as shown in 2 Screening and verification of antimicrobial peptide candidate samples.
[0466] Example 4: Antimicrobial peptide sequence design framework based on supervised fine-tuning and human feedback reinforcement learning and antimicrobial peptide screening framework obtained antimicrobial peptides
[0467] This embodiment relates to an antimicrobial peptide sequence design framework based on supervised fine-tuning and human feedback reinforcement learning and antimicrobial peptides obtained by an antimicrobial peptide screening framework. Specifically, a batch of candidate antimicrobial peptides can be obtained through the design framework in Example 2, and the candidate antimicrobial peptides are input into the antimicrobial peptide screening framework for screening and verification, and finally 13 antimicrobial peptides with expected properties and MIC activity are obtained.
[0468] The molar mass, inhibition rate, and MIC of the 13 antimicrobial peptides were tested. The specific steps are shown in 3. Testing of antimicrobial peptide properties.
[0469] The molar mass, inhibition rate, and MIC results of the 13 antimicrobial peptides are shown in Tables 1-3.
[0470] Table 1. Molar masses of 13 antimicrobial peptides
[0471]
[0472]
[0473] Table 2. Inhibition rate of 13 antimicrobial peptides
[0474]
[0475] In Table 2, the two values 50 and 100 below each bacterial species respectively indicate that the inhibition rate (%) was measured at two antimicrobial peptide concentrations of 50 μg / mL and 100 μg / mL. Therefore, the data in the table represent the antibacterial effects of each antimicrobial peptide on different microorganisms (Pseudomonas aeruginosa, Saccharomyces cerevisiae, Candida albicans, Escherichia coli, Staphylococcus aureus) at two concentrations: 50 represents the inhibition rate (%) when the antimicrobial peptide concentration is 50 μg / mL; 100 represents the inhibition rate (%) when the antimicrobial peptide concentration is 100 μg / mL.
[0476] The calculation of the inhibition rate is usually based on the colony growth between the control group (without antimicrobial peptide) and the experimental group (with antimicrobial peptide), and the formula is:
[0477] Inhibition rate (%) = (1-CFU or OD600 of experimental group / CFU or OD600 of control group) × 100%,
[0478] The higher the inhibition rate, the stronger the antibacterial effect of the antimicrobial peptide.
[0479] From the results of the antibacterial rate in Table 2, it can be concluded that:
[0480] (1) Overall trend:
[0481] For most bacterial species, when the concentration of antimicrobial peptides increased from 50 μg / mL to 100 μg / mL, the inhibition rate increased significantly, indicating that the antibacterial effect of antimicrobial peptides was concentration-dependent.
[0482] The inhibition rates of some antimicrobial peptides were low or even negative at 50 μg / mL (such as ZJAPM055, ZJAPM056, ZJAPM059, and ZJAPM065), but significantly increased at 100 μg / mL, indicating that these antimicrobial peptides may require higher concentrations to effectively inhibit bacteria.
[0483] The inhibition rate of some antimicrobial peptides on certain bacterial species is close to 100% (such as ZJAPM051, ZJAPM052, and ZJAPM062), indicating that they may be highly effective antimicrobial peptides on these bacterial species.
[0484] (2) Performance of different antimicrobial peptides:
[0485] ZJAPM051, ZJAPM052, and ZJAPM062: They showed high inhibition rates (>90%) against all bacterial species, whether at 50 μg / mL or 100 μg / mL, indicating that these antimicrobial peptides have broad-spectrum and highly effective antimicrobial activity and may be the most promising candidate antimicrobial peptides in the study.
[0486] ZJAPM055, ZJAPM056, and ZJAPM059: At a concentration of 50 μg / mL, they performed poorly against some bacterial species (such as Pseudomonas aeruginosa, Saccharomyces cerevisiae, and Candida albicans), and even showed negative values, which may indicate that they are ineffective against these bacterial species at low concentrations and may even promote bacterial growth. However, at 100 μg / mL, some of the inhibition rates were restored, indicating that they may require higher concentrations to work.
[0487] ZJAPM064: It exhibited relatively stable antibacterial activity at both 50μg / mL and 100μg / mL, especially against Escherichia coli and Staphylococcus aureus, with a high inhibition rate (>95%), which may be more targeted to these bacteria.
[0488] ZJAPM065: The inhibition rate was low at low concentration (50 μg / mL), but the inhibition rate against Escherichia coli and Staphylococcus aureus increased significantly at 100 μg / mL, indicating that its antibacterial effect was weak but still had certain potential.
[0489] (3) Differences in resistance among different strains:
[0490] Escherichia coli and Staphylococcus aureus showed high inhibition rates (usually >90%) to most antimicrobial peptides, indicating that these bacteria may be more sensitive to these antimicrobial peptides.
[0491] The inhibition rates of Pseudomonas aeruginosa and Saccharomyces cerevisiae were quite different. Some antimicrobial peptides (such as ZJAPM05 and ZJAPM056) had lower inhibition rates against them at 50 μg / mL, indicating that these strains may have a certain tolerance and require higher concentrations to inhibit their growth.
[0492] The overall inhibition rate of Candida albicans was high, but the activity of some antimicrobial peptides was weak at 50 μg / mL, which may indicate that it has a certain tolerance to low concentrations of antimicrobial peptides.
[0493] Table 3. MIC of 13 antimicrobial peptides (μg / mL)
[0494]
[0495] From the MIC results in Table 3, it can be concluded that:
[0496] (1) Overall analysis of MIC results:
[0497] The MIC value range is approximately 1 to 32 μg / mL, where:
[0498] Lowest MIC value (1 μg / mL): ZJAPM065 had the lowest MIC against Staphylococcus aureus (1 μg / mL), indicating that it had the strongest antibacterial ability against this strain;
[0499] The highest MIC value (32 μg / mL): ZJAPM052 had the highest MIC against Candida albicans (32 μg / mL), indicating that its antibacterial effect against this species was weak and a higher concentration was required to completely inhibit its growth.
[0500] Most antimicrobial peptides were effective against Escherichia coli and Staphylococcus aureus at concentrations between 2 and 8 μg / mL, indicating that these bacteria were relatively susceptible to the effects of antimicrobial peptides.
[0501] The MIC values of Candida albicans and Saccharomyces cerevisiae were generally high (8-32 μg / mL), indicating that fungi were less sensitive to these antimicrobial peptides and that higher concentrations might be required to effectively inhibit their growth. This may be related to the complex structure of the fungal cell wall.
[0502] (2) Analysis of the performance of each antimicrobial peptide:
[0503] The antimicrobial peptides that performed better (low MIC and strong activity) are as follows:
[0504] (I) ZJAPM065: The MIC values were the lowest in Staphylococcus aureus (1 μg / mL) and Escherichia coli (4 μg / mL), indicating that it had good antibacterial effects against both Gram-positive bacteria (S. aureus) and Gram-negative bacteria (E. coli).
[0505] (II) ZJAPM051 and ZJAPM052: They showed low MIC against Staphylococcus aureus (2 μg / mL), Escherichia coli (4 μg / mL), and Saccharomyces cerevisiae (4 μg / mL), indicating that they are relatively broad-spectrum antimicrobial peptides.
[0506] (III) ZJAPM061 and ZJAPM062: They had lower MICs against Pseudomonas aeruginosa (4 μg / mL), Candida albicans (8 μg / mL), and Escherichia coli (8 μg / mL), indicating that they may be effective against these bacterial species.
[0507] The antimicrobial peptides that performed poorly (high MIC, requiring higher concentrations to inhibit bacteria) are as follows:
[0508] (I) ZJAPM052 had the highest MIC against Candida albicans (32 μg / mL), indicating that its inhibitory ability against Candida albicans was weak.
[0509] (II) The MIC values of ZJAPM056 against various bacterial species were relatively high (e.g., 16 μg / mL for Pseudomonas aeruginosa and 16 μg / mL for Candida albicans), indicating that its antibacterial activity was weak and a higher dose might be required for effective inhibition of bacteria.
[0510] (III) The MIC values of ZJAPM057 and ZJAPM059 were high (8-16 μg / mL) against multiple bacterial species, indicating that their antibacterial activities were relatively poor.
[0511] (3) Analysis of tolerance of different bacterial species:
[0512] (a) Staphylococcus aureus (S. aureus): The MIC values of most antimicrobial peptides were low (1-4 μg / mL), indicating that it was relatively sensitive to these antimicrobial peptides.
[0513] (b) Escherichia coli (E. coli): The MIC was mainly concentrated in 4-8 μg / mL, indicating that the antimicrobial peptides had good antibacterial activity against it.
[0514] (c) Pseudomonas aeruginosa: MIC values ranged from 4 to 16 μg / mL. Some antimicrobial peptides (such as ZJAPM056) showed weak antibacterial activity (MIC = 16 μg / mL), indicating that the bacterium had a high tolerance to these antimicrobial peptides.
[0515] (d) Saccharomyces cerevisiae and Candida albicans: The MIC values were generally high (8-32 μg / mL), indicating that these fungi have strong tolerance to antimicrobial peptides and may require higher concentrations or more specific antimicrobial peptides to effectively inhibit bacteria.
[0516] (e) Fungi (unspecified species): MIC values were between 4 and 8 μg / mL, indicating that some antimicrobial peptides still had a certain inhibitory effect on them.
[0517] (4) Comprehensive analysis combined with the antibacterial rate experiment:
[0518] ZJAPM051, ZJAPM052, and ZJAPM065 had high inhibition rates (>90%) against multiple bacterial species in the inhibition rate experiment, and showed low MIC values (1-8 μg / mL) in the MIC test, indicating that these antimicrobial peptides may be more effective antimicrobial candidates.
[0519] ZJAPM056, ZJAPM057, and ZJAPM059 performed poorly in the MIC test (8-32 μg / mL), and the inhibition rates of some bacterial species were low in the inhibition rate experiment, indicating that these antimicrobial peptides may have low activity within the existing concentration range, or their mechanisms of action are different.
[0520] The MICs of Candida albicans and Saccharomyces cerevisiae were generally high, which was consistent with the phenomenon that some antimicrobial peptides were less effective at low concentrations in the inhibition rate experiment, indicating that fungi may have stronger tolerance and may require higher concentrations or optimization of antimicrobial peptide sequences.
[0521] The amino acid sequences of the 13 antimicrobial peptides are shown in Table 4.
[0522] Table 4. Amino acid sequences of 13 antimicrobial peptides
[0523]
[0524] All documents mentioned in the present invention are cited as references in this application, just as each document is cited as reference individually. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the claims attached to this application.
Claims
1. A method for constructing an antimicrobial peptide production model, characterized in that: The method comprises the steps of: (1) Train the protein pedestal model on the dataset; (2) Performing supervised fine-tuning on the protein base model to obtain a fine-tuned protein base model and an antimicrobial peptide candidate sequence.
2. The method according to claim 1, characterized in that In step (2), the step further includes a step selected from the following group: (2.1) training a predictor on the data set to obtain a pre-trained predictor, wherein the pre-trained predictor outputs a prediction value, inputs the prediction value into a reward function, and uses the reward function to perform human feedback reinforcement learning on the fine-tuned protein base model to obtain an optimized protein model and an antimicrobial peptide candidate sequence; (2.2) designing a reward function, and using the reward function to perform human feedback reinforcement learning on the fine-tuned protein base model to obtain an optimized protein model and an antimicrobial peptide candidate sequence; or a combination thereof.
3. The method according to claim 1 or 2, characterized in that In the human feedback reinforcement learning, a reinforcement learning algorithm is used to adjust the fine-tuned protein pedestal model.
4. A framework or system for producing antimicrobial peptides, characterized in that: The framework or system includes: (D1) an input unit, wherein the input unit is configured to input data, wherein the input data includes a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence; (D2) a fine-tuning unit, the fine-tuning unit being configured to execute a fine-tuning model on the input data, thereby obtaining fine-tuned input data; wherein the fine-tuning model comprises a supervised fine-tuning model, the supervised fine-tuning model comprising the steps of: performing supervised fine-tuning on the input data using a fine-tuning method, thereby obtaining fine-tuned input data; (D3) An output unit, wherein the output unit is configured to output the result of the fine-tuning unit.
5. The framework or system according to claim 4, characterized in that The fine-tuning model also includes a model selected from the following group: (E1) an activity feedback fine-tuning model, the activity feedback fine-tuning model comprising the steps of: training a predictor on a data set to obtain a pre-trained predictor, the pre-trained predictor outputting a prediction value, inputting the prediction value into a reward function, and performing human feedback reinforcement learning on the input data using the reward function, thereby obtaining fine-tuned input data; or (E2) An attribute feedback fine-tuning model, wherein the attribute feedback fine-tuning model comprises the steps of: designing a reward function, and using the reward function to perform human feedback reinforcement learning on the input data, thereby obtaining fine-tuned input data.
6. A method for constructing an antimicrobial peptide screening model, characterized in that: The method comprises the steps of: (M1) machine learning screening, the machine learning screening comprising the steps of: (A1) Using a machine learning model, the sequence similarity between the candidate antimicrobial peptide sequence and the target functional sample is evaluated through a scoring system; (A2) predictive indicators of antimicrobial activity; (M2) a posteriori verification, the a posteriori verification comprising a step selected from the group consisting of: (C1) limit peptide length; (C2) predicting structural features; (C3) Analyze similarities; or a combination thereof.
7. An antimicrobial peptide screening framework or system, characterized in that: The framework or system includes: (B1) an input unit, wherein the input unit is configured to input data, wherein the input data includes an input sequence to be screened, wherein the input sequence to be screened includes an antimicrobial peptide candidate sequence generated using the antimicrobial peptide generation framework or system according to claim 4; (B2) an antimicrobial peptide screening unit, the screening unit being configured to execute an antimicrobial peptide screening model, thereby obtaining an antimicrobial peptide target sequence from the antimicrobial peptide candidate sequence; wherein the screening model is constructed using the method of claim 6; (B3) An output unit, wherein the output unit is configured to output the screening result.
8. A framework or system for designing and screening antimicrobial peptides, characterized in that: The framework or system includes: (W) an input unit, wherein the input unit is configured to input data, wherein the input data includes a protein base model and / or an input sequence, wherein the protein base model is trained on the input sequence; (X) a fine-tuning unit, wherein the fine-tuning unit is configured as a fine-tuning model, wherein the fine-tuning model performs a predetermined fine-tuning on the input data to obtain a fine-tuning result; wherein the fine-tuning model includes: (x1) a supervised fine-tuning model, the supervised fine-tuning model comprising the steps of: performing supervised fine-tuning on the input data using a fine-tuning method, thereby obtaining fine-tuned input data; (x2) an activity feedback fine-tuning model, the activity feedback fine-tuning model comprising the steps of: training a predictor on a data set to obtain a pre-trained predictor, the pre-trained predictor outputting a prediction value, inputting the prediction value into a reward function, and performing human feedback reinforcement learning on the input data using the reward function, thereby obtaining fine-tuned input data; (x3) an attribute feedback fine-tuning model, the attribute feedback fine-tuning model comprising the steps of: designing a reward function, and using the reward function to perform human feedback reinforcement learning on the input data, thereby obtaining fine-tuned input data; or a combination thereof; The fine-tuning unit is configured to fine-tune a model selected from the group consisting of: (X1) supervised fine-tuning model as described in (x1); (X2) A combination of the supervised fine-tuning model described in (x1) and the activity feedback fine-tuning model described in (x2); or (X3) A combination of the supervised fine-tuning model described in (x1) and the attribute feedback fine-tuning model described in (x3); (Y) a screening unit, the screening unit being configured to execute a screening model for antimicrobial peptides, the screening model obtaining an antimicrobial peptide target sequence from the fine-tuned input data; wherein the screening model is constructed using the method of claim 6; (Z) an output unit, wherein the output unit is configured to output the result of the screening module.
9. The framework or system of claim 8, wherein: The antimicrobial peptide target sequence includes an amino acid sequence as shown in any one of SEQ ID NOs: 1-13.
10. Use of a drug or a pharmaceutical composition, wherein the drug or the pharmaceutical composition comprises an antimicrobial peptide sequence, wherein the antimicrobial peptide sequence is obtained by the framework or the system according to claim 8, characterized in that: The uses include: (F1) preparing a formulation for inhibiting the growth of microorganisms; (F2) preparing a kit, the kit further comprising a label or instructions indicating that the kit is used to inhibit the growth of microorganisms.
Citation Information
Cited By
Antibody multi-site mutation evolution method based on multi-agent reinforcement learning
CN120260676A
An antibody multi-site mutation evolution method based on multi-agent reinforcement learning
CN120260676B
Antibacterial peptide function interpretable prediction method and system based on graph causal learning
CN120808899A
Calculation design method and system of antioxidant peptide
CN120932743A