Antibacterial peptide function recognition optimization method and system based on deep learning and LLM

Through the fusion method of deep learning and large language model, structured prompt word templates and natural language reasoning are constructed, and antimicrobial peptide functional prediction is optimized, which solves the problem of feature dependence and insufficient samples in the antimicrobial peptide functional design, and achieves efficient and accurate functional recognition and generation.

CN120472985APending Publication Date: 2025-08-12SHANDONG UNIV QILU HOSPITAL
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510557301.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing antimicrobial peptide functional design and prediction technology have strong feature dependence, insufficient samples, and long-tail distribution problems, resulting in low prediction accuracy in low-frequency functional categories, limited generation ability, and high experimental verification costs.

Method used

The fusion method of deep learning and large language model (LLM) is adopted to build structured prompt word templates and natural language inference, combine the sequence information and physical and chemical properties of antimicrobial peptides, optimize the prediction results, and use the knowledge verification mechanism of large language models to improve prediction accuracy and robustness.

Benefits of technology

It significantly improves the accuracy and stability of antimicrobial peptide function prediction, reduces the risk of misjudgment by traditional deep learning methods under conditions of insufficient samples or incomplete characteristics, and improves the efficiency and accuracy of antimicrobial peptide function recognition and generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472985A_ABST
    Figure CN120472985A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer-aided drug research and development, and provides an antibacterial peptide function recognition optimization method and system based on deep learning and LLM. Systematic innovation is performed in multiple links such as feature extraction, data generation and reasoning verification by introducing a fusion mechanism of a large language model and a deep learning model; and the accuracy and robustness of antibacterial peptide function prediction are obviously improved. On one hand, a strict feature extraction process and a structured cue word template are constructed, sequence information and physicochemical property indexes of the antibacterial peptide and a prediction result of a deep learning model can be effectively integrated, and the prediction result is reasonably corrected and optimized through the reasoning ability and knowledge verification mechanism of a large language model. The misjudgment risk of a traditional deep learning method under the condition of insufficient samples or incomplete features is effectively reduced, and the prediction accuracy and stability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field related to computer-aided drug development, and more specifically, to a method and system for antimicrobial peptide function identification and optimization based on deep learning and LLM. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] Antimicrobial peptides (polypeptides) are a class of biomacromolecules composed of amino acids linked by peptide bonds. They perform a variety of key functions in life processes, including regulation, signal transduction, immune response, and metabolic control. They have been widely used in a variety of biomedical fields, including antibacterial, antiviral, anti-tumor, cytokine mimicry, and cardiovascular therapy. Compared to traditional small molecule drugs, antimicrobial peptides offer significant advantages such as structural flexibility, strong targeting, high safety, and good biocompatibility, and are increasingly attracting attention in life science and bioengineering research.

[0004] However, the number of natural antimicrobial peptides in nature is limited, and their functional distribution is highly uneven, which limits the application of antimicrobial peptides in specific functional areas. For example, antimicrobial peptides account for more than 40% of the total data in public databases, while the number of antimicrobial peptides with functions such as antiviral and antifungal is significantly smaller. In addition, the functional design of antimicrobial peptides requires sequence optimization based on specific needs, but the cost and cycle of experimental verification are high, which seriously restricts the efficiency of new antimicrobial peptide screening based on trial and error. Therefore, how to construct new antimicrobial peptide sequences with diverse functions has become a key issue that urgently needs to be broken through in current antimicrobial peptide engineering.

[0005] At present, the mainstream technical routes for antimicrobial peptide design and function prediction mainly include the following two categories:

[0006] One type involves numerical simulation and energy optimization methods based on physical and chemical rules, such as molecular dynamics simulation (MD), Monte Carlo sampling, and molecular docking. These methods emphasize precise modeling of the relationship between sequence and structure and can provide a certain degree of physical interpretation, but they often have high computational overhead and low search efficiency in high-dimensional sequence space, making them difficult to apply to high-throughput antimicrobial peptide design tasks.

[0007] Another category, deep learning methods, has emerged in recent years. These methods predict or generate antimicrobial peptide functions by training on large datasets of functionally labeled antimicrobial peptides. Common models include convolutional neural networks (CNNs), long short-term memory networks (LSTMs), Transformers, and their variants. While these methods have achieved significant progress in overall prediction performance, they still face the following challenges: First, they are highly feature-dependent. The performance of deep models is highly dependent on the selection and processing of input features. Improper feature engineering can lead to poor model generalization. Second, antimicrobial peptide function databases exhibit a severe long-tail distribution, with over 80% of samples concentrated in three to five high-frequency functional categories. This significantly reduces the model's prediction accuracy for low-frequency functional categories (such as antiviral and antifungal). Furthermore, training data is insufficient. In particular, high-quality training data is severely lacking in sequence generation tasks, limiting generative capabilities. Existing generative models struggle to balance multiple functional optimization objectives, such as enhancing antimicrobial activity while simultaneously reducing toxicity. Summary of the Invention

[0008] To address the above-mentioned issues, the present disclosure proposes a method and system for antimicrobial peptide function identification and optimization based on deep learning and LLM. The prediction results of a small-scale deep learning model and sequence information are used as prompts to prompt the LLM for reasoning. The prediction results of the small-scale deep learning model are modified using the LLM's vast knowledge base, thereby significantly improving the accuracy of the recognition task and achieving more accurate computer-assisted drug development.

[0009] In order to achieve the above objectives, the present disclosure adopts the following technical solutions:

[0010] One or more embodiments provide an antimicrobial peptide function identification and optimization method based on deep learning and LLM, including an antimicrobial peptide function identification method and an antimicrobial peptide optimization method. The antimicrobial peptide function identification method comprises the following steps:

[0011] Obtaining the sequence information of the antimicrobial peptide to be identified and inputting it into a trained antimicrobial peptide prediction model based on a deep learning model to obtain a first identification result of the antimicrobial peptide function;

[0012] Constructing a structured first prompt word template, and filling the antimicrobial peptide sequence information, the first recognition result, and the analysis requirement information into the first prompt word template;

[0013] The filled prompt word template information is input into the large language model, and the natural language reasoning of the large language model is used to re-judge and correct the first prediction result in combination with context information to obtain the antimicrobial peptide function recognition result.

[0014] One or more embodiments provide an antimicrobial peptide function identification and optimization system based on deep learning and LLM, including an antimicrobial peptide function identification unit and an antimicrobial peptide optimization unit. The antimicrobial peptide function identification unit includes:

[0015] A function prediction module is configured to obtain sequence information of the antimicrobial peptide to be identified and input it into a trained antimicrobial peptide prediction model based on a deep learning model to obtain a first identification result of the antimicrobial peptide function;

[0016] A first prompt word template construction module is configured to construct a structured first prompt word template and fill the antimicrobial peptide sequence information, the first recognition result, and the analysis requirement information into the first prompt word template;

[0017] The first reasoning module is configured to input the filled prompt word template information into the large language model, use the natural language reasoning of the large language model, and combine the context information to re-judge and correct the first prediction result to obtain the antimicrobial peptide function recognition result.

[0018] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the above-mentioned antimicrobial peptide function identification and optimization method based on deep learning and LLM are completed.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] The disclosed method for identifying antimicrobial peptide functions, by integrating a large language model with a deep learning model, has made systematic innovations in multiple aspects, including feature extraction, data generation, and reasoning verification, significantly improving the accuracy and robustness of antimicrobial peptide function prediction. On the one hand, a rigorous feature extraction process and structured prompt word templates are constructed to effectively integrate the sequence information, physicochemical properties, and prediction results of the antimicrobial peptide with the deep learning model. The prediction results are then rationally corrected and optimized through the reasoning power and knowledge verification mechanism of the large language model. This effectively reduces the risk of misjudgment in traditional deep learning methods when samples are insufficient or features are incomplete, significantly improving the accuracy and stability of predictions.

[0021] The advantages of the present disclosure and additional advantages will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings, which constitute a part of the present disclosure, are used to provide a further understanding of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure but do not constitute a limitation of the present disclosure.

[0023] Figure 1 is a flow chart of the antimicrobial peptide function identification method of Example 1 of the present disclosure;

[0024] Figure 2 is a flow chart of the antimicrobial peptide optimization method of Example 1 of the present disclosure;

[0025] Figure 3 is the ROC curve obtained by testing on an independent test set in the simulation experiment of Example 1 of the present disclosure;

[0026] Figure 4 is the PR curve obtained by testing on an independent test set in the simulation experiment of Example 1 of the present disclosure;

[0027] Figure 5 is the MIC regression score obtained by the antimicrobial peptide prediction method in the simulation experiment of Example 1 of the present disclosure;

[0028] Figure 6 3. This is a comparison diagram of the predicted distribution values of the minimum inhibitory concentration (MIC) and the half-maximal hemolytic lethal concentration (HC50) of the antimicrobial peptide generation method and the existing method in the simulation experiment of Example 1 of the present disclosure; DETAILED DESCRIPTION

[0029] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0030] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs.

[0031] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof. It should be noted that, in the absence of conflict, the various embodiments in the present disclosure and the features in the embodiments can be combined with each other. The embodiments will be described in detail below with reference to the accompanying drawings.

[0032] Example 1

[0033] In the technical solutions disclosed in one or more embodiments, Figures 1 to 2 As shown, the antimicrobial peptide function identification and optimization method based on deep learning and LLM includes an antimicrobial peptide function identification method and an antimicrobial peptide optimization method, wherein the antimicrobial peptide function identification method includes the following steps:

[0034] Step 1: Obtain the sequence information of the antimicrobial peptide to be identified and input it into a trained antimicrobial peptide prediction model based on a deep learning model to obtain a first identification result of the antimicrobial peptide function;

[0035] Step 2: constructing a structured first prompt word template, and filling the antimicrobial peptide sequence information, the first recognition result, and the analysis requirement information into the first prompt word template;

[0036] Step 3: Input the filled prompt word template information into the large language model, use the natural language reasoning of the large language model, and combine the context information to re-judge and correct the first prediction result to obtain the antimicrobial peptide function recognition result;

[0037] The antimicrobial peptide function identification method of this embodiment, by introducing a fusion mechanism of a large language model and a deep learning model, has made systematic innovations in multiple aspects, including feature extraction, data generation, and reasoning verification, significantly improving the accuracy and robustness of antimicrobial peptide function prediction. On the one hand, a rigorous feature extraction process and structured prompt word templates are constructed to effectively integrate the sequence information, physical and chemical properties of antimicrobial peptides, and the prediction results of the deep learning model. The reasoning ability and knowledge verification mechanism of the large language model are then used to reasonably correct and optimize the prediction results. This effectively reduces the risk of misjudgment of traditional deep learning methods under conditions of insufficient samples or incomplete features, significantly improving the accuracy and stability of predictions.

[0038] In step 1, the deep learning model can adopt an existing deep learning model. The model selection and training strategy of the deep learning model are as follows:

[0039] Step 11: Set the training data threshold and compare it with the number of training data samples;

[0040] Step 12: When the number of training data samples is not less than the set training data threshold, select multiple sequence models for training. The sequence models can be multi-layer LSTM, Mamba sequence modeling network, or Transformer model, that is, multi-sequence model training;

[0041] Step 13: When the number of training data samples is less than the set training data threshold, a lightweight network structure model is selected for training. The lightweight network structure can be a single-layer GRU, a small CNN-LSTM hybrid network, etc., to reduce the risk of overfitting, i.e., lightweight model training;

[0042] Step 14: Evaluate the stability and generalization ability of the trained model through k-fold cross validation.

[0043] Specifically, the process of training the antimicrobial peptide prediction model includes the following:

[0044] Step 101: Acquire antimicrobial peptide data with target functions from an antimicrobial peptide database;

[0045] Specifically, antimicrobial peptide data containing target functions are obtained from public antimicrobial peptide databases (such as APD, DBAASP, CAMPR3, etc.);

[0046] Antimicrobial peptide data may include: antimicrobial peptide amino acid sequence information, corresponding functional tags, and experimentally measured quantitative indicators (such as MIC value, CC50 value).

[0047] Among them, antimicrobial peptide functional tags include antibacterial, antiviral, cytotoxicity, etc., and the quantitative indicators measured by experiments can include MIC value, CC50 value, etc.

[0048] Step 102: Construct an antimicrobial peptide template with target functional characteristics. Utilize a large language model (LLM) to generate antimicrobial peptide sequences based on the input set functional requirements. The generated antimicrobial peptide sequences are screened using existing prediction models or physicochemical rules. The acquired antimicrobial peptide data are supplemented to construct a synthetic training dataset. This step involves LLM generating synthetic sequences to enhance the training set when data is insufficient.

[0049] Among them, the large language model can be DeepSeek R1 or V3;

[0050] To address the problem of scarce samples (n<500) in certain functional categories (such as antifungal and antiviral), this embodiment proposes a completion mechanism based on a large language model to effectively expand the amount of training data.

[0051] This approach addresses the uneven distribution of natural antimicrobial peptide data and the scarcity of samples for low-frequency functional categories by leveraging a large language model to generate synthetic sequences with target functional characteristics, thereby expanding the training sample library and improving the model's recognition capabilities for long-tail categories. This approach not only enhances the generalization performance of deep learning models in small sample sizes but also effectively reduces the risk of overfitting caused by insufficient training samples.

[0052] Step 103: Based on the antimicrobial peptide data of the synthetic training dataset, feature encoding is performed and input into a deep learning model for model training to obtain an antimicrobial peptide function prediction result of the antimicrobial peptide data;

[0053] Optionally, feature encoding can use one-hot encoding, amino acid index (AAIndex) or embedded representation (such as ProtVec, ESM embedding);

[0054] Among them, for the prediction results of antimicrobial peptide functions, for deep learning models of classification tasks, Sigmoid or Softmax activation functions can be used to output category probabilities, such as antimicrobial prediction; for deep learning models of regression tasks, linear output layers can be used to predict quantitative indicators, such as MIC values;

[0055] Step 104: Calculate the loss function value based on the prediction result, iteratively train the antimicrobial peptide prediction model until a cutoff condition is met, and obtain a trained antimicrobial peptide prediction model;

[0056] Among them, the loss function can be set according to the task, using Binary Crossentropy (classification) or MSE / MAE (regression);

[0057] In step 2, the constructed structured first prompt word template may include the following elements in sequence:

[0058] 1) Antimicrobial peptide sequence information: the complete amino acid sequence of the target antimicrobial peptide;

[0059] 2) Physical and chemical properties of antimicrobial peptides: including isoelectric point, hydrophobicity index, molecular weight, secondary structure ratio, etc.;

[0060] 3) The first prediction result obtained based on the antimicrobial peptide prediction model: for example, "antimicrobial probability: 85%", "MIC predicted value: 64 μg / mL";

[0061] 4) Specific analysis requirements;

[0062] In a specific example, the prompt word structure can be constructed as follows:

[0063] "Sequence: XXX...X, isoelectric point: 6.8, hydrophobicity: 0.43, α-helix ratio: 40%, model-predicted antibacterial probability: 85%, predicted MIC: 64 μg / mL. Please determine whether this antimicrobial peptide possesses antibacterial activity and identify possible optimization options."

[0064] In step 3, the natural language reasoning capability of the large language model is used to re-judge and correct the first prediction result in combination with contextual information to obtain the antimicrobial peptide function recognition result. Then, the output format is fine-tuned through instructions to finally convert the LLM output text into structured prediction data.

[0065] The structured prediction data may be in JSON format, table form, or structured records in a database;

[0066] The data format of JSON format is {"field name": value}. For example, the output natural language text is: "This antimicrobial peptide has good antibacterial activity, with an antibacterial probability of about 91% and a predicted MIC of 48 μg / mL, which is a moderate to strong antibacterial level." Converted to JSON format is {"corrected antibacterial probability": 91, "corrected MIC value": 48, "function judgment": "has moderate to strong antibacterial activity"};

[0067] To verify the effectiveness of the antimicrobial peptide function identification method proposed in this example, which incorporates a large language model (LLM) inference correction mechanism, a comparative experiment was designed and conducted for antimicrobial peptide function prediction. First, using the antimicrobial performance prediction task with sufficient training samples as an example, 151 experimentally verified antimicrobial peptide samples were selected from a public experimental database as an independent test set to determine whether they possessed antimicrobial function. When a deep learning model trained on a basic dataset containing approximately 12,000 real-world annotated data points performed predictions, its AUROC (area under the receiver operating characteristic curve) was only 0.778. The model exhibited some misjudgment when processing samples with blurred functional boundaries and complex sequence representations. However, the antimicrobial peptide function identification method proposed in this example, which incorporates the LLM inference correction module, was able to perform a semantically comprehensive judgment based on sequence, physicochemical characteristics, and preliminary prediction results, effectively compensating for the basic model's shortcomings in understanding local information. This significantly improved the final prediction AUROC to 0.851, demonstrating excellent classification performance and stability.

[0068] Furthermore, in order to verify the practicality and generalization ability of the framework under small sample conditions, this embodiment selected the antifungal performance task, a task with scarce training data, as an evaluation scenario, and extracted 300 antimicrobial peptide samples with antifungal function annotations from the dbAMP database for experiments. Since there were very few high-quality samples with antifungal labels in the original training set, the basic model was trained only with a small amount of data synthesized by LLM, resulting in a prediction accuracy of only 58% on the independent test set. However, after integrating the LLM reasoning correction module proposed in this embodiment, the model's ability to understand the deep underlying laws between samples was significantly enhanced, ultimately increasing the detection success rate to 84%, verifying the adaptability and robustness of this method in small sample and data imbalance scenarios. Experimental results show that this method can effectively overcome the problems faced by existing models in function prediction, such as strong feature dependence and high sample sparsity, and significantly improve prediction accuracy, especially showing superior performance in long-tail categories and boundary samples.

[0069] As a further technical solution, this embodiment provides a method for generating antimicrobial peptides with specific functions based on the fusion optimization of prior-guided reinforcement learning (PRL) and a large language model (LLM). This method aims to address the high dependence of traditional deep learning models on large-scale annotated data. It is particularly suitable for generating and screening antimicrobial peptide sequences with clear functional objectives (such as high antimicrobial activity and low toxicity). By incorporating the design concept of a reward mechanism from reinforcement learning, this method combines the basic model, domain knowledge, and the reasoning capabilities of the LLM to continuously optimize the target properties of the generated antimicrobial peptide sequences.

[0070] The overall process is as follows Figure 2 As shown, the antimicrobial peptide optimization method comprises the following steps:

[0071] Step S1: setting an optimization goal according to the functional requirements of the target antimicrobial peptide to be optimized, and constructing a reward function based on the optimization goal and prior knowledge embedding;

[0072] To ensure that the generated antimicrobial peptide sequences possess biological properties consistent with their intended functions, this example introduces a reward function design mechanism based on customized optimization objectives. Specifically, the reward function serves as feedback during the reinforcement learning training process. Based on the set antimicrobial peptide optimization objective, the reward function ensures that the generated model's behavior aligns with the functional requirements.

[0073] In some embodiments, optimization goals include improving antimicrobial activity, reducing toxicity risk, and controlling hydrophobicity levels;

[0074] Prior knowledge can include biological physical and chemical property laws, literature and database statistical knowledge, etc.

[0075] Biophysical and chemical property patterns, including the correlation between the isoelectric point, hydrophobicity, amino acid composition and functional performance of antimicrobial peptides;

[0076] Statistical knowledge of literature and databases, including common structural fragments, functional distribution, and ranges of physical and chemical parameters of antimicrobial peptides;

[0077] Optionally, these abstract optimization objectives are converted into quantifiable learning information, and the antimicrobial peptide function recognition results predicted by the antimicrobial peptide prediction model are combined with prior knowledge to construct a reward function, including:

[0078] Based on the physical and chemical properties, the first part of the reward function is constructed as follows:

[0079] R1 = λ1×f1 (antibacterial probability) - λ2×f2 (toxicity probability) - λ3×f3 (hydrophobicity deviation) ...

[0080] Among them, f1, f2, and f3 represent the quantitative functions of the corresponding functions or physicochemical parameters, and λ1, λ2, and λ3 are adjustable weight parameters used to reflect the importance trade-offs between the optimization objectives;

[0081] Based on the predicted antimicrobial peptide properties:

[0082]

[0083] Among them, p anti and p toxic are the predicted probability that the peptide is an antimicrobial peptide and the probability that the peptide has hemolytic toxicity, respectively, MIC is an adjustable weight parameter, MIC max is the maximum value that the MIC predictor can output, and MIC is the actual output value of the MIC predictor.

[0084] Finally, add up the overall reward function:

[0085] R=R1+R2

[0086] Through the aforementioned reward function construction method, this embodiment enables the reinforcement learning model to dynamically receive "the performance of the current generated sequence in each optimization dimension" during training, and adjust the antimicrobial peptide generator parameters based on this feedback, gradually improving the overall performance of the generated antimicrobial peptides in terms of target performance. This mechanism significantly improves the learning efficiency and result reliability of the generative model in data-scarce scenarios, effectively avoiding the uncertainty in target control associated with traditional methods that rely on large-scale labeled data training. This reward function design strategy gives this embodiment a clear goal orientation, ensuring the functional biological validity and engineering feasibility of the generated sequence, and providing a scalable and controllable generation optimization path for antimicrobial peptide drug development.

[0087] The reward function will be directly used to train the target optimization strategy in the subsequent reinforcement training phase and is the core constraint mechanism in the entire generation process.

[0088] Step S2: obtaining the antimicrobial peptide template or starting fragment information to be optimized, performing reinforcement learning based on the deep learning-based antimicrobial peptide generator to generate an optimized antimicrobial peptide sequence structure, and iteratively optimizing the antimicrobial peptide sequence structure based on the antimicrobial peptide generator based on the constructed reward function to obtain the reinforcement learning-optimized sequence structure of the target antimicrobial peptide;

[0089] An antimicrobial peptide generator is constructed based on a deep learning sequence generation model. Specifically, a sequence generation model based on structures such as Transformer and GRU is used as an antimicrobial peptide generator (policy network). The antimicrobial peptide generator generates a complete sequence based on an antimicrobial peptide template or starting fragment.

[0090] Furthermore, based on the constructed reward function, the sequence structure of the antimicrobial peptide is iteratively optimized based on the antimicrobial peptide generator. The process is as follows:

[0091] Step S21: For each candidate antimicrobial peptide sequence structure generated by the antimicrobial peptide generator, the antimicrobial peptide function recognition method of steps 1 to 3 is used to identify each optimized antimicrobial peptide function recognition result, calculate the value of the reward function, and score the antimicrobial peptide sequence structure.

[0092] Calculate the dependent variable of the reward function based on prior knowledge, and calculate the value of the reward function based on the dependent variable, including:

[0093] 21.1) Obtain prior knowledge, domain expert rules, and algorithms based on physical and chemical properties from a biological algorithm library to calculate the first dependent variable of the reward function, which may include α-helix, isoelectric point, net charge, positive charge ratio, hydrophobicity, and stability;

[0094] 21.2) Using the antimicrobial peptide function identification method of steps 1 to 3 above, the predicted second dependent variable of the antimicrobial peptide function (MIC predicted value, AMP probability predicted value, hemolytic toxicity probability predicted value);

[0095] 21.3) calculating a reward value based on the first dependent variable and the second dependent variable;

[0096] Step S22: The reward value is used as a feedback signal to input into the antimicrobial peptide generator for reinforcement learning. The antimicrobial peptide generator adjusts parameters based on the feedback signal and outputs the next candidate antimicrobial peptide sequence structure.

[0097] The antimicrobial peptide generator adjusts its parameters based on the feedback signal, gradually optimizing the generated sample distribution towards "higher rewards";

[0098] Step S23, iteratively executing steps 21 to 22 until the reward function value meets the set requirements, thereby obtaining the optimized sequence structure of the target antimicrobial peptide;

[0099] The above process goes through multiple rounds of iterations, enabling the generator to quickly approach the optimization target on an initial unsupervised basis.

[0100] Unlike traditional data-driven sequence generation, the reinforcement learning mechanism can keep the optimization direction of the generative model highly consistent with the target attributes even under conditions of extremely small or unbalanced data.

[0101] After completing the preliminary PRL sequence generation, in order to further improve the structural rationality and functional accuracy of the candidate antimicrobial peptides, this embodiment introduces a large language model (such as DeepSeek R1 / V3) for high-level semantic reasoning optimization.

[0102] Step S3: construct a second prompt word template, and use the large language model to supplement and modify the semantics and implicit patterns of the sequence structure optimized by reinforcement learning to obtain the final optimized antimicrobial peptide sequence structure;

[0103] Construct a structured second prompt word template that can include the following information:

[0104] 3.1) Sequence structure generated by the antimicrobial peptide generator after reinforcement learning optimization;

[0105] 3.2) Physicochemical properties calculated based on prior knowledge, including isoelectric point, hydrophobicity, α-helix ratio, etc.

[0106] 3.3) The optimization goal set, such as “improving antibacterial activity while reducing toxicity”;

[0107] 3.4) Specific optimization analysis requirements;

[0108] In a specific example, the prompt word example is: "Sequence: KKLLKLLKLLKK, isoelectric point 8.6, hydrophobicity 0.49, toxicity probability: 0.31, goal: improve antibacterial activity while avoiding cytotoxicity, please optimize this sequence."

[0109] The prompt words are input into the Large Language Model (LLM), which uses its ability to integrate language and life science knowledge to output optimization suggestions or revised antimicrobial peptide sequences.

[0110] Furthermore, the output structure of the specification is fine-tuned through instructions and converted into a structured format, wherein the structured prediction data can be in JSON format, table form or structured records in a database;

[0111] This example innovatively integrates a deep learning model with a large language model to construct an efficient and accurate method for antimicrobial peptide function identification and optimization. This method leverages the advantages of deep semantic mining and a massive knowledge base, combining the strengths of deep learning in extracting antimicrobial peptide sequence features with the global logical reasoning and knowledge integration of a large language model. This method achieves significant performance improvements in the identification and generation of antimicrobial peptides and other functional antimicrobial peptides.

[0112] This embodiment provides prompt engineering design and structured output of results, further improving the interpretability and output availability of the deep learning model. While ensuring prediction accuracy, it reduces the number of manual verifications and experimental costs, significantly improving overall R&D efficiency. It proposes a new computational path that breaks through the limitations of traditional feature engineering, providing important technical support for antimicrobial peptide function prediction and related computer-aided drug development, and also offers broad application prospects for the field of drug design in processing high-dimensional features and low-resource tasks.

[0113] In order to illustrate the effect of the above identification optimization method, an experiment was conducted, which is described as follows;

[0114] (1) Antimicrobial peptide recognition performance;

[0115] The receiver operating characteristic curve (ROC curve) and precision-recall curve (PR curve) were plotted to compare the performance of various mainstream deep learning methods and the method used in this example (i.e., LLM+LLM-GME in the figure) in the antimicrobial peptide classification task. As shown in the figure, the evaluation index uses the area under the curve (AUC) as a quantitative standard, and its value directly reflects the classification efficiency of the model. The experimental results show that Figure 3 and Figure 4 As shown in the figure, the method adopted in this embodiment achieved excellent performance of 0.851 area under the ROC curve and 0.874 area under the PR curve on the independent test set, which showed significant advantages compared with other methods, fully verifying the advancement and reliability of the method adopted in this embodiment in the field of antimicrobial peptide classification.

[0116] AMP-CLIP: a method proposed in the literature (DOI:10.1162 / neco.1997.9.8.1735);

[0117] LSTM: The model proposed in the literature (DOI:10.1162 / neco.1997.9.8.1735);

[0118] Mamba: a model proposed in the literature (DOI:10.48550 / arXiv.2312.00752);

[0119] MLA: Model proposed in the literature (DOI:10.48550 / arXiv.2405.04434);

[0120] Ensemble all: average results of LSTM, Mamba, MHA, and MLA predictions;

[0121] LLM-GME: Use LLM to judge the results of selective fusion of LSTM, Mamba, MHA, and MLA

[0122] LLM+LLM-GME: The antimicrobial peptide function identification method of this embodiment uses LLM to correct the results of LLM-GME;

[0123] like Figure 5As shown, the performance of the mainstream deep learning model and the method adopted in this embodiment (i.e., LLM+LLM-GME in the figure) in the regression prediction task of antimicrobial peptide minimum inhibitory concentration (MIC) was evaluated by two core indicators: mean absolute error (MAE) and relative square error (RSE). As evaluation indicators, MAE and RSE both show a negative correlation - the closer their values are to zero, the higher the model prediction accuracy. The method adopted in this embodiment achieved a MAE of 0.969 and an RSE of 0.495 on an independent test set, showing a significant performance advantage over other comparison methods. This result fully verifies the advancement and reliability of the method adopted in this embodiment in the field of antimicrobial peptide MIC prediction.

[0124] (2) Antimicrobial peptide production performance;

[0125] like Figure 6 The results of comparative analysis of key pharmacodynamic indicators of antimicrobial peptides using various advanced antimicrobial peptide production methods and the method used in this example (i.e., PRL and PRL+DS-R1 shown in the figure) are presented;

[0126] PRL refers to the result generated by the reinforcement learning method proposed in this example, and PRL+DS-R1 refers to the result optimized by DeepSeek-R1 based on PRL.

[0127] HydrAMP: a method proposed in the literature (DOI:10.1038 / s41467-023-36994-z);

[0128] DeepAMP: a method proposed in the literature (DOI:10.1038 / s41467-024-51933-2);

[0129] ProGen2-s: a method proposed in the literature (DOI:10.48550 / arXiv.2004.03497);

[0130] DS-V3: DeepSeek V3 Large Language Model;

[0131] GPT-o1: OpenAI GPT-o1 large language model;

[0132] Specifically, Figure 6 The distribution of predicted minimum inhibitory concentrations (MICs) and hemolytic lethal concentrations (HC50s) of the antimicrobial peptides generated by each method is presented. The results demonstrate that the method employed in this example exhibits significant advantages in antimicrobial activity, with the generated antimicrobial peptides exhibiting lower MICs (i.e., more pronounced antimicrobial potency) and higher HC50s (i.e., significantly reduced cytotoxicity), fully demonstrating the innovative nature and application value of the method employed in this example in the field of antimicrobial peptide research and development.

[0133] Example 2

[0134] Based on Example 1, this embodiment provides an antimicrobial peptide function identification and optimization system based on deep learning and LLM, including an antimicrobial peptide function identification unit and an antimicrobial peptide optimization unit. The antimicrobial peptide function identification unit includes:

[0135] A function prediction module is configured to obtain sequence information of the antimicrobial peptide to be identified and input it into a trained antimicrobial peptide prediction model based on a deep learning model to obtain a first identification result of the antimicrobial peptide function;

[0136] A first prompt word template construction module is configured to construct a structured first prompt word template and fill the antimicrobial peptide sequence information, the first recognition result, and the analysis requirement information into the first prompt word template;

[0137] The first reasoning module is configured to input the filled prompt word template information into the large language model, use the natural language reasoning of the large language model, and combine the context information to re-judge and correct the first prediction result to obtain the antimicrobial peptide function recognition result.

[0138] Furthermore, the antimicrobial peptide optimization unit includes:

[0139] A reward function construction module is configured to set an optimization goal according to the functional requirements of the target antimicrobial peptide to be optimized, and to embed and construct a reward function based on the optimization goal and prior knowledge;

[0140] a reinforcement learning module configured to obtain information about an antimicrobial peptide template or starting fragment to be optimized, perform reinforcement learning based on a deep learning-based antimicrobial peptide generator to generate an optimized antimicrobial peptide sequence structure, and iteratively optimize the antimicrobial peptide sequence structure based on the antimicrobial peptide generator based on a constructed reward function to obtain a reinforcement learning-optimized sequence structure of a target antimicrobial peptide;

[0141] The second reasoning module is configured to construct a second prompt word template, and use the large language model to supplement and correct the semantics and implicit patterns of the sequence structure optimized by reinforcement learning to obtain the final optimized antimicrobial peptide sequence structure.

[0142] It should be noted here that the various modules in this embodiment correspond one-to-one to the various steps in Example 1, and the specific implementation processes are the same, which will not be repeated here.

[0143] Example 3

[0144] This embodiment provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the antimicrobial peptide function identification and optimization method based on deep learning and LLM in Example 1 are completed.

[0145] Example 4

[0146] This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the antimicrobial peptide function identification and optimization method based on deep learning and LLM in Example 1 are completed.

[0147] The foregoing description is merely a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

[0148] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Those skilled in the art should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. Antimicrobial peptide function identification and optimization method based on deep learning and LLM, characterized by: The invention includes an antimicrobial peptide function identification method and an antimicrobial peptide optimization method. The antimicrobial peptide function identification method includes the following steps: Obtaining the sequence information of the antimicrobial peptide to be identified and inputting it into a trained antimicrobial peptide prediction model based on a deep learning model to obtain a first identification result of the antimicrobial peptide function; Constructing a structured first prompt word template, and filling the antimicrobial peptide sequence information, the first recognition result, and the analysis requirement information into the first prompt word template; The filled prompt word template information is input into the large language model, and the natural language reasoning of the large language model is used to re-judge and correct the first prediction result in combination with context information to obtain the antimicrobial peptide function recognition result.

2. The method for antimicrobial peptide function identification and optimization based on deep learning and LLM according to claim 1, characterized in that: It also includes the model selection and training strategy of the deep learning model for the antimicrobial peptide prediction model, specifically: Set the training data threshold and compare it with the number of training data samples; When the number of training data samples is not less than the set training data threshold, multiple sequence models are selected for training; When the number of training data samples is less than the set training data threshold, a lightweight network structure model is selected for training; The stability and generalization ability of the trained model were evaluated through k-fold cross validation.

3. The antimicrobial peptide function identification and optimization method based on deep learning and LLM according to claim 1, characterized in that: The process of training the antimicrobial peptide prediction model includes the following: Obtain antimicrobial peptide data with target functions from the antimicrobial peptide database; Construct antimicrobial peptide templates with target functional characteristics, use a large language model to generate antimicrobial peptide sequences based on the input set functional requirements, screen the generated antimicrobial peptide sequences using physical and chemical rules, supplement the acquired antimicrobial peptide data, and construct a synthetic training dataset; Based on the antimicrobial peptide data of the synthetic training dataset, feature encoding is input into the deep learning model for model training to obtain the antimicrobial peptide function prediction results of the antimicrobial peptide data; The loss function value is calculated according to the prediction results, and the antimicrobial peptide prediction model is iteratively trained until the cutoff condition is met to obtain the trained antimicrobial peptide prediction model.

4. The method for antimicrobial peptide function identification and optimization based on deep learning and LLM according to claim 1, characterized in that: The first prompt word template includes the following elements in sequence: antimicrobial peptide sequence information, antimicrobial peptide physicochemical properties, a first prediction result obtained based on an antimicrobial peptide prediction model, and analysis requirements.

5. The antimicrobial peptide function identification and optimization method based on deep learning and LLM according to claim 1, characterized in that: The antimicrobial peptide optimization method comprises the following steps: According to the functional requirements of the target antimicrobial peptide to be optimized, the optimization goal is set, and the reward function is constructed based on the optimization goal and prior knowledge embedding; Obtaining the antimicrobial peptide template or starting fragment information to be optimized, performing reinforcement learning based on a deep learning-based antimicrobial peptide generator to generate an optimized antimicrobial peptide sequence structure, and iteratively optimizing the antimicrobial peptide sequence structure based on the antimicrobial peptide generator based on the constructed reward function to obtain the reinforcement learning-optimized sequence structure of the target antimicrobial peptide; A second prompt word template was constructed, and the sequence structure optimized by reinforcement learning was supplemented and corrected in terms of semantics and implicit patterns using a large language model to obtain the final optimized antimicrobial peptide sequence structure.

6. The method for antimicrobial peptide function identification and optimization based on deep learning and LLM according to claim 1, characterized in that: Based on the constructed reward function, the sequence structure of the antimicrobial peptide is iteratively optimized based on the antimicrobial peptide generator. The process is as follows: Step S21: The antimicrobial peptide generator generates a candidate antimicrobial peptide sequence structure each time, identifies each optimized antimicrobial peptide function recognition result based on the antimicrobial peptide function recognition method, and calculates the value of the reward function; Step S22: The reward value is used as a feedback signal to input into the antimicrobial peptide generator for reinforcement learning. The antimicrobial peptide generator adjusts parameters based on the feedback signal and outputs the next candidate antimicrobial peptide sequence structure. Step S23: iteratively execute steps 21 to 22 until the reward function value meets the set requirements, thereby obtaining the optimized sequence structure of the target antimicrobial peptide.

7. The method for antimicrobial peptide function identification and optimization based on deep learning and LLM according to claim 1, characterized in that: The second prompt word template includes the following information: the sequence structure generated by the antimicrobial peptide generator after reinforcement learning optimization; the physical and chemical property parameters calculated based on prior knowledge; the set optimization goals; and the optimization analysis requirements.

8. An antimicrobial peptide function identification and optimization system based on deep learning and LLM, comprising an antimicrobial peptide function identification unit and an antimicrobial peptide optimization unit, characterized in that: Antimicrobial peptide functional identification unit, including: A function prediction module is configured to obtain sequence information of the antimicrobial peptide to be identified and input it into a trained antimicrobial peptide prediction model based on a deep learning model to obtain a first identification result of the antimicrobial peptide function; A first prompt word template construction module is configured to construct a structured first prompt word template and fill the antimicrobial peptide sequence information, the first recognition result, and the analysis requirement information into the first prompt word template; The first reasoning module is configured to input the filled prompt word template information into the large language model, use the natural language reasoning of the large language model, and combine the context information to re-judge and correct the first prediction result to obtain the antimicrobial peptide function recognition result.

9. The antimicrobial peptide function identification and optimization system based on deep learning and LLM according to claim 8, characterized in that: Antimicrobial Peptide Optimization Unit, including: A reward function construction module is configured to set an optimization goal according to the functional requirements of the target antimicrobial peptide to be optimized, and to embed and construct a reward function based on the optimization goal and prior knowledge; a reinforcement learning module configured to obtain information about an antimicrobial peptide template or starting fragment to be optimized, perform reinforcement learning based on a deep learning-based antimicrobial peptide generator to generate an optimized antimicrobial peptide sequence structure, and iteratively optimize the antimicrobial peptide sequence structure based on the antimicrobial peptide generator based on a constructed reward function to obtain a reinforcement learning-optimized sequence structure of a target antimicrobial peptide; The second reasoning module is configured to construct a second prompt word template, and use the large language model to supplement and correct the semantics and implicit patterns of the sequence structure optimized by reinforcement learning to obtain the final optimized antimicrobial peptide sequence structure.

10. An electronic device, characterized in that: The invention comprises a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the steps of the antimicrobial peptide function identification and optimization method based on deep learning and LLM according to any one of claims 1 to 7 are completed.

Citation Information

Cited By

  • Large-model-driven intelligent biological research method and system

    CN120805532A

  • Calculation design method and system of antioxidant peptide

    CN120932743A

  • Antibody sequence optimization method based on large language model

    CN121191600A

  • Antibacterial peptide recognition method and device based on secondary structure characteristics and robust statistics

    CN122245452A