A multi-target antimicrobial peptide sequence optimization method and system
By constructing a sequence search tree and a co-evolutionary framework, combined with a multi-attribute prediction model, the antimicrobial peptide sequence was optimized, solving the problem of balancing antimicrobial activity, toxicity, and half-life in existing technologies. This resulted in the generation of highly efficient and safe antimicrobial peptide sequences, providing high-quality candidates for the development of antibiotic alternatives.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-03-24
AI Technical Summary
Existing methods for optimizing and designing antimicrobial peptides have failed to effectively balance key indicators such as antimicrobial activity, toxicity, and half-life, resulting in poor overall performance of the generated peptide sequences. Furthermore, the optimization process is prone to getting stuck in local optima and lacks global search capabilities.
A multi-objective antimicrobial peptide sequence optimization method was adopted. By constructing a sequence search tree and a co-evolutionary framework, combined with a multi-attribute prediction model and a Pareto optimal screening strategy, the balanced optimization of attributes such as antimicrobial activity, hemolysis, and half-life was achieved.
A novel antimicrobial peptide sequence with high activity, low toxicity, and long half-life was generated, which is suitable for the development of antibiotic alternatives and improves the drug development process of antimicrobial peptides.
Smart Images

Figure CN120877880B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of antimicrobial peptides, and more specifically, to a method and system for multi-target antimicrobial peptide sequence optimization. Background Technology
[0002] Antimicrobial peptides are a class of short peptides naturally found in plants, animals, and microorganisms. They possess important functions such as broad-spectrum antibacterial, antifungal, and antiviral activity. Because they primarily exert their effects by disrupting microbial membrane structures or regulating immune responses, they are less likely to induce bacterial resistance, thus holding significant strategic importance in the development of anti-infective drugs. However, the in vivo application of most natural or artificially designed antimicrobial peptides currently faces significant limitations, mainly due to high toxicity, poor stability, and insufficient pharmacokinetic properties. For example, many peptides with high antimicrobial activity also exhibit strong hemolytic activity or toxic side effects on mammalian cells, posing significant safety risks when used in vivo. Conversely, some peptides, while having lower toxicity, lack sufficient activity to achieve effective antibacterial inhibition. Furthermore, antimicrobial peptides are readily degraded by proteases in vivo, have short half-lives, and struggle to maintain effective concentrations, further impacting their therapeutic efficacy. These factors significantly limit the drug development process of antimicrobial peptides.
[0003] Existing antimicrobial peptide optimization design methods often focus solely on enhancing antimicrobial activity, lacking systematic consideration of key indicators such as toxicity and half-life. This results in poor overall performance of the generated candidate peptides. Some studies have attempted to use machine learning or deep learning models to predict peptide activity or toxicity, but their optimization processes still fail to achieve a balance and synergy among multiple indicators, and they remain insufficient in terms of screening efficiency and quality. Current antimicrobial peptide optimization methods based on rule-based design or single-point mutation strategies often fail to effectively model the nonlinear coupling relationships between the high-dimensional physicochemical features inherent in the sequence. The optimization process is prone to getting trapped in local optima and lacks global search capabilities, leading to performance bottlenecks in many aspects of the generated peptide sequences.
[0004] Therefore, a solution is needed to optimize multi-target antimicrobial peptide sequences. Summary of the Invention
[0005] To address the problems existing in the prior art, this application provides a method and system for multi-target antimicrobial peptide sequence optimization. The specific solution is as follows:
[0006] A multi-target antimicrobial peptide sequence optimization method includes the following:
[0007] Determine the initial peptide sequence to be processed and its improvement requirements. Construct a sequence search tree using the initial peptide sequence as nodes. Iteratively perform the following operations based on the sequence search tree:
[0008] The nodes in the sequence search tree are scored according to the preset scoring rules, and a node is selected as the current node based on the score.
[0009] Based on the aforementioned improvement requirements, point mutation operations are performed on the current node according to preset mutation rules to generate multiple candidate peptides, and the candidate peptides are added as new nodes to the sequence search tree.
[0010] Based on the aforementioned improvement requirements, each candidate peptide is copied to its corresponding subpopulation. Based on the co-evolutionary framework, the candidate peptides are locally optimized using the subpopulations to obtain new candidate peptides. Each new candidate peptide is then evaluated using a preset multi-attribute prediction model.
[0011] The evaluation results of each new candidate peptide are fed back to the corresponding current node in the sequence search tree. Each node is re-scored based on the evaluation results and the current node is reselected.
[0012] When the iteration stopping condition is met, a peptide sequence is selected from the sequence search tree based on the score as the peptide sequence that meets the improvement requirements and output.
[0013] In some specific embodiments, the scoring rules include: scoring based on the number of times a node is selected, the number of times its corresponding parent node is selected, and the evaluation result of the node.
[0014] In some specific embodiments, the expression for the scoring rule is:
[0015]
[0016] Where C1 is the cumulative reward derived from all evaluation results for this node. This represents the number of times a node is accessed. To explore the weighting factors, This represents the number of times the parent node of this node has been visited.
[0017] In some specific embodiments, each subpopulation corresponds to the optimization of an attribute objective. The corresponding attribute objective is determined according to the improvement requirements, and then the corresponding subpopulation is selected for optimization.
[0018] In some specific embodiments, the multi-attribute prediction model uses a weighted summation method to quantify and integrate the multi-target attributes of new candidate peptides, and on this basis, introduces a Pareto optimal screening strategy to ensure a reasonable balance among multiple target attributes.
[0019] In some specific embodiments, the target properties include antibacterial activity, hemolysis, half-life, thermal stability, serum stability, hydrophobicity, immunogenicity, biomembrane penetration ability, and anti-endotoxin ability.
[0020] In some specific embodiments, the mutation rules include:
[0021] When the sequence search tree contains only the initial peptide sequence node, multiple peptide sequences are generated by introducing random points to guide mutation of the initial peptide sequence, thereby forming a diverse initial population. The number of peptide sequences in the initial population is set proportionally according to the length of the initial peptide sequence.
[0022] When the sequence search tree has multiple nodes, the peptide sequences corresponding to two parent nodes are randomly exchanged and recombinated according to a preset crossover rate to generate new child nodes, or the peptide sequences corresponding to child nodes generated by mutations are guided to mutate again based on a preset mutation rate to generate new child nodes.
[0023] In some specific embodiments, the multi-attribute prediction model specifically includes:
[0024] Features of new candidate peptides are extracted by a preset feature extraction part and multi-source feature fusion is performed by a multi-scale fusion strategy to generate a fused feature vector.
[0025] By performing deep encoding on the fused feature vector through the pre-defined shared representation modeling part, global dependency information between residues is extracted, and a dynamic representation with context-aware capability is constructed to obtain the core semantic features that simultaneously associate antibacterial activity and hemolytic activity prediction.
[0026] Regression prediction is performed based on the feature subspace related to the antibacterial mechanism in the core semantic features through a preset antibacterial activity prediction branch, and classification prediction is performed based on the feature subspace related to the toxicity in the core semantic features through a preset hemolyticity scoring branch.
[0027] The prediction errors of the two branches are backpropagated to the feature extraction part and the shared representation modeling part through a joint loss function, thereby dynamically optimizing the fusion ratio of multi-source feature fusion and the encoding weight of deep encoding.
[0028] In some specific embodiments, for peptide sequences with known chemical structures, atoms in the molecule are regarded as graph nodes and chemical bonds as edges. The topological features and functional group distribution information at the molecular level are extracted through a message passing mechanism and finally mapped into a fusionable low-dimensional vector to achieve modeling of the molecular structure.
[0029] A multi-objective antimicrobial peptide sequence optimization system, used to implement the multi-objective antimicrobial peptide sequence optimization method described in any one of the above claims, includes the following:
[0030] The input unit is used to determine the initial peptide sequence to be processed and its improvement requirements. A sequence search tree is constructed using the initial peptide sequence as nodes, and the following operations are iteratively performed based on the sequence search tree:
[0031] The iterative optimization unit is used to score the nodes in the sequence search tree according to the preset scoring rules, and select a node as the current node based on the score.
[0032] Based on the aforementioned improvement requirements, point mutation operations are performed on the current node according to preset mutation rules to generate multiple candidate peptides, and the candidate peptides are added as new nodes to the sequence search tree.
[0033] Based on the aforementioned improvement requirements, each candidate peptide is copied to its corresponding subpopulation. Based on the co-evolutionary framework, the candidate peptides are locally optimized using the subpopulations to obtain new candidate peptides. Each new candidate peptide is then evaluated using a preset multi-attribute prediction model.
[0034] The evaluation results of each new candidate peptide are fed back to the corresponding current node in the sequence search tree. Each node is re-scored based on the evaluation results and the current node is reselected.
[0035] The output unit is used to select a peptide sequence from the sequence search tree based on the score as the peptide sequence that meets the improvement requirements and output it when the iteration stopping condition is met.
[0036] Beneficial Effects: This application proposes a multi-objective antimicrobial peptide sequence optimization method and system. It models antimicrobial peptide sequence mutations as a tree-structured search path and introduces multi-objective co-evolution in the simulation stage. By combining screening strategies and weighted scoring mechanisms, it achieves efficient coupling between global search and attribute balance. It can simultaneously perform joint modeling of pharmacological properties such as antimicrobial activity, host toxicity, and half-life, and complete efficient and refined optimization of antimicrobial peptide sequences. It is especially suitable for performance reconstruction and safety improvement starting from existing peptides, and has significant technical advantages and application prospects in the field of antibiotic alternative development.
[0037] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0038] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a schematic diagram of the multi-target antimicrobial peptide sequence optimization method of this application;
[0040] Figure 2This is a flowchart illustrating the optimization strategy based on Monte Carlo tree search.
[0041] Figure 3 This is a schematic diagram of the overall structure of the co-evolutionary framework of this application;
[0042] Figure 4 This is a schematic diagram of the multi-task model;
[0043] Figure 5 This is a schematic diagram of the verification test for this application;
[0044] Figure 6 This is a diagram of the hemolysis test results of this application;
[0045] Figure 7 This is a schematic diagram of the multi-target antimicrobial peptide sequence optimization system module of this application.
[0046] Figure labels: 1-Input unit; 2-Iterative optimization unit; 3-Output unit. Detailed Implementation
[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0048] This application proposes a multi-objective method for optimizing antimicrobial peptide sequences, the process of which is attached. Figure 1 As shown, the specific solution is as follows:
[0049] A multi-target antimicrobial peptide sequence optimization method includes the following:
[0050] 101. Determine the initial peptide sequence to be processed and its improvement requirements. Construct a sequence search tree using the initial peptide sequence as nodes. Iterate through steps 102-105 based on the sequence search tree:
[0051] 102. Score the nodes in the sequence search tree according to the preset scoring rules, and select a node as the current node based on the score;
[0052] 103. Based on the improvement requirements, perform point mutation operations on the current node according to the preset mutation rules to generate multiple candidate peptides, and add the candidate peptides as new nodes to the sequence search tree;
[0053] 104. Based on the improvement requirements, each candidate peptide is copied to the corresponding subpopulation. Based on the co-evolution framework, the candidate peptides are locally optimized using the subpopulation to obtain new candidate peptides. Each new candidate peptide is evaluated by a preset multi-attribute prediction model.
[0054] 105. The evaluation results of each new candidate peptide are fed back to the corresponding current node in the sequence search tree. Each node is re-scored based on the evaluation results and the current node is reselected.
[0055] 106. When the iteration stopping condition is met, select peptide sequences from the sequence search tree based on the score as peptide sequences that meet the improvement requirements and output them.
[0056] This invention relates to a multi-objective computational optimization method for antimicrobial peptides, aiming to enhance antimicrobial activity while significantly reducing host toxicity and prolonging in vivo half-life, thereby obtaining candidate antimicrobial peptide sequences with greater drug potential. By constructing a multi-attribute prediction and evaluation model, the method comprehensively evaluates the target attributes of antimicrobial peptide sequences, such as antimicrobial activity, hemolytic activity, and half-life. Using existing antimicrobial peptides as initial seed sequences, new candidate sequences are generated through operations such as fragment substitution, mutation, and recombination, and Monte Carlo sampling is used to improve global search capabilities. The optimization process is conducted in multiple iterations guided by a multi-objective scoring function, dynamically balancing the relationship between activity, toxicity, and stability to screen sequences with excellent overall performance. Ultimately, this invention can obtain a batch of novel antimicrobial peptides with high activity, low toxicity, and long half-life, suitable for subsequent drug development and experimental validation, and has significant advantages such as clear optimization objectives, strong search capabilities, and high sequence adaptability.
[0057] In this application, the initial peptide sequence to be processed in step 101 serves as the input peptide sequence and is the starting point for optimization. It typically exhibits characteristics of high toxicity or low activity requiring improvement. It should be noted that this invention is applicable not only to the optimization of existing antimicrobial peptide sequences but also to scenarios such as local modification and functional reconstruction of artificially constructed or natural peptides. For example, for functional segments such as signal peptides, transport peptides, and cell-penetrating peptides extracted from natural proteins, the optimization method provided by this invention can enhance their antimicrobial function and inhibit toxicity, expanding their applicability in the field of novel functional peptide design. This invention focuses on the multi-objective optimization of antimicrobial peptides, but its multi-task evaluation-optimization feedback framework can be extended to other short peptide drugs (such as antiviral peptides, anticancer peptides, and cell-penetrating peptides). Only the input structural form and objective function design need to be adjusted to reuse the technical route and optimization process of this invention.
[0058] Improvement requirements specify the target attributes to be optimized, such as a combination of multiple objectives, such as enhancing antibacterial activity while reducing hemolysis and extending half-life. Improvement requirements need to be set based on the initial peptide sequence and are selected manually. The sequence search tree is a tree-like structure with peptide sequences as nodes and sequence relationships formed by mutation / optimization operations as edges, used to record and trace the optimization path.
[0059] In some embodiments, the sequence search tree is a Monte Carlo Tree Search (MCTS) tree structure. The Monte Carlo Tree Search (MCTS), as the core search engine in this invention, is responsible for efficient exploration and decision-making in the peptide sequence conformation space. The process based on the Monte Carlo Tree Search (MCTS) optimization strategy is attached. Figure 2 As shown, this method models the antimicrobial peptide sequence optimization problem as a state transition process. Through four stages in a tree structure—Selection, Expansion, Simulation, and Backpropagation—it progressively explores the optimal region in the peptide sequence space. In each simulation, a new peptide is generated based on the current state, and its antimicrobial activity and toxicity levels are evaluated using a multi-task prediction model. The evaluation results guide the backpropagation of node scores, thus achieving a heuristic search oriented towards multiple objectives. This strategy can effectively escape local optima, improving optimization efficiency and result diversity. The algorithm constructs a state-action search tree, modeling each mutation or modification of the peptide as an action on a path in the tree, thereby guiding the optimization process to search in a more purposeful way in the complex space. In practical applications, without changing the optimization objective, it can also replace other algorithm systems with global search capabilities, such as Genetic Algorithm (GA), Particle Swarm Optimization (PSO), Simulated Annealing (SA), and Deep Reinforcement Learning (DRL). These alternative algorithms can be adjusted in terms of structure search strategies, candidate peptide mutation methods, and evaluation feedback mechanisms, but they are still essentially based on model-driven multi-objective evolutionary optimization.
[0060] By optimizing existing peptide sequences in a targeted manner, the blind approach of designing peptides from scratch is avoided. A tree structure organizes discrete peptide sequences into an ordered search space, making the optimization process traceable and controllable. The iterative mechanism simulates the trial-and-error selection process of natural evolution, gradually accumulating high-quality sequence features. In this application, the sequence search tree needs to be iteratively executed. By repeating steps 102-105, the search tree is continuously expanded, gradually approaching the optimal peptide that meets the improvement requirements. Specifically, MCTS sequentially performs four steps during the optimization process: selection, expansion, simulation, and backpropagation, corresponding to steps 102-105 respectively.
[0061] A sequence search tree is constructed using an initial peptide sequence. The initial peptide provides the structural and functional basis, focusing the search on potential sequence neighborhoods. Initially, the sequence search tree contains only one node, corresponding to the initial peptide sequence. Subsequent iterative execution of steps 102-105 continuously expands the sequence search tree by adding more peptide sequences based on the initial sequence. When the iteration stops, steps 102-105 are stopped, and a suitable peptide sequence is selected from the sequence search tree. In this application, the potential of existing peptides in the search tree is evaluated using quantification rules, and the peptide sequence is ultimately selected based on its potential ranking.
[0062] Specifically, a priority score is generated by combining the historical performance of a node, the sufficiency of its exploration, and the performance of its parent nodes. The current node is the peptide to be mutated and optimized in this iteration, serving as the parent for generating new candidate peptides. The scoring rules quantify potential through a mathematical model, tilting resources towards nodes with high scores and less exploration, ensuring that the search process both deeply explores the optimization potential of high-quality sequences and does not overlook potential high-value new sequences. In some specific embodiments, the scoring rules include: a score based on the number of times a node is selected, the number of times its corresponding parent node is selected, and the node's evaluation result. The number of times a node is selected is the number of times the peptide sequence has been chosen as the current node for mutation optimization, i.e., its exploration frequency; the number of times a parent node is selected is the number of times its parent peptide (the node that generated it) has been selected, reflecting the overall potential of the parent sequence; the node's evaluation result is the comprehensive score of the peptide evaluated by a multi-attribute prediction model, such as a weighted total score of antibacterial activity, hemolytic activity, and half-life. This combination of three factors ensures that the scoring not only focuses on current performance but also on exploration value, ensuring that low-frequency but high-potential nodes in the search tree have a chance to be optimized.
[0063] Furthermore, the expression for the scoring rule is:
[0064]
[0065] Where C1 is the cumulative reward derived from all evaluation results for this node. This represents the number of times a node is accessed. To explore the weighting factors, This represents the number of visits to the parent node of the given node. During the selection phase, this application will calculate this evaluation value for each candidate child node. Each evaluation result for a node will be included in the node's cumulative reward. Based on these evaluation values, the system can adopt a deterministic strategy (selecting the node with the highest evaluation value) or a probabilistic strategy (converting the evaluation value into a selection probability and then sampling) to iteratively traverse the tree until a preset termination condition is reached, thereby completing an efficient and balanced path selection.
[0066] Step 103 is the expansion phase, where rule-based point mutations are performed on the target peptide starting from the selected node to generate new candidate peptides, which are then added to the tree as child nodes. By performing controlled mutations on the current node, new sequences are generated and added to the search tree, expanding the optimization scope. In this application, mutation methods are designed according to improvement requirements, with controlled mutation as the core. For example, if the improvement requirement is to increase activity, the activity-related structures are adjusted, or a subpopulation that enhances activity is introduced for optimization, thereby achieving peptide sequence mutations in the direction of activity. The new sequence generated after mutation is added to the search tree as a child node, forming a parent-child relationship with the current node. This simulates gene mutations in the natural evolution of proteins, but reduces harmful mutations (such as mutations that lead to peptide inactivation) through rule constraints, ensuring that the new sequence retains the basic function of the parent while introducing mutations to explore better performance.
[0067] In some specific embodiments, the mutation rules include: when there is only one node, the initial peptide sequence, in the sequence search tree, multiple peptide sequences are generated by introducing random points to guide the mutation of the initial peptide sequence, thereby forming a diverse initial population, and the number of peptide sequences in the initial population is proportionally set according to the length of the initial peptide sequence; when there are multiple nodes in the sequence search tree, the peptide sequences corresponding to two parent nodes are randomly exchanged and recombinated according to a preset crossover rate to generate new child nodes, or the peptide sequences corresponding to the child nodes generated by mutation are guided to mutate again based on a preset mutation rate to generate new child nodes.
[0068] Step 104 is the simulation phase, which is the core of the entire optimization strategy. In this phase, this application introduces a co-evolution mechanism to replicate the current candidate peptide into multiple subpopulations, each independently performing local optimization for different attribute targets. Each subpopulation employs a carefully designed mutation strategy and selection mechanism to generate several new sequences, which are then evaluated in conjunction with a prediction model.
[0069] In the evaluation process, this invention employs a weighted summation method to quantitatively integrate multiple objective attributes, and introduces a Pareto optimal screening strategy to ensure that the selected solution achieves a reasonable balance among multiple performance indicators. This reward mechanism retains the operability of weighted aggregation while avoiding a single objective dominating the optimization direction through Pareto front constraints, thereby achieving synergistic optimization among enhanced activity, reduced toxicity, and improved stability. Candidate peptides are assigned to specialized optimization populations (subpopulations) for fine-tuning, and their multi-objective attributes are quickly evaluated through a model, solving the core problem of balancing multiple objectives. In some specific embodiments, each subpopulation corresponds to the optimization of one attribute objective. The corresponding attribute objective is determined according to the improvement requirements, and then the corresponding subpopulation is selected for optimization. Subpopulations are divided according to the target attributes of the improvement requirements, and each subpopulation focuses on only one attribute objective. For example, the antibacterial activity subpopulation focuses on improving antibacterial ability, the hemolytic subpopulation focuses on reducing host toxicity, and the half-life subpopulation focuses on prolonging metabolic stability. Each subpopulation independently executes a mutation-evaluation-selection cycle, achieving a balanced improvement of multiple attributes. A single subpopulation can optimize the target attribute through targeted mutation depth, avoiding the inefficiency caused by multi-objective mixing. If the improvement requirement does not involve a certain attribute, the corresponding subpopulation can be shut down, and resources can be concentrated on optimizing the core objective. The feedback during the optimization process can be replaced by numerical prediction results generated by multi-task neural networks, expert knowledge scoring, experimental measurement results (such as MIC, hemolytic test), or rule engine evaluation (such as whether it contains a certain type of toxic structural motif), to support an experiment- / knowledge-driven optimization closed loop, thereby forming a hybrid optimization process of computational prediction + experimental feedback, which is also within the scope of protection of this invention.
[0070] Step 105 is the backpropagation phase. The optimization results from step 104 are propagated back along the search path to the ancestor nodes of the tree structure, gradually adjusting the value assessment and strategy distribution of the entire tree to improve the effectiveness and diversity of subsequent search directions. The multi-attribute scores of new candidate peptides are associated with their parent nodes, updating the historical performance records of the parent nodes. Based on the backpropagation results, the priorities of all nodes are recalculated using the scoring rules from step 102. A closed-loop feedback mechanism is introduced to give the optimization process learning capabilities—using historical data to determine which nodes are more likely to evolve into high-quality offspring, dynamically allocating resources, and improving iteration efficiency.
[0071] Step 106 involves iteration termination and outputting the optimal peptide. When optimization reaches the preset target or has no further potential for improvement, the iteration terminates and the optimal result is output. Iteration termination conditions include reaching the preset number of iterations, performance meeting the target, and convergence stabilization. For example, if a certain number of new candidate peptides satisfy all improvement requirements, the performance is considered to have met the target; if the score improvement of the optimal peptide is <0.5% in 50 consecutive iterations, meaning the peptide has no further potential for improvement, convergence stabilization can be considered. The peptide with the highest score is selected from the search tree as the final optimization result.
[0072] In some specific embodiments, the target attributes include antimicrobial activity, hemolytic activity, and half-life. Antimicrobial activity indicates the peptide's ability to inhibit / kill bacteria, usually measured by the MIC (minimum inhibitory concentration) (the smaller the value, the stronger the activity); hemolytic activity indicates the peptide's ability to damage human erythrocytes (a toxicity indicator), expressed as erythrocyte lysis rate (the lower the value, the higher the safety); half-life indicates the stability of the peptide's metabolism in vivo, i.e., the time it takes for the concentration to drop to half of its initial value (the longer the half-life, the longer the duration of efficacy). As potential antibiotic alternatives, antimicrobial peptides must simultaneously meet the three core requirements of highly effective bactericidal action, low toxicity, and long-lasting effect; the balance of these three is key to their drug development. Target attributes provide specific benchmarks for subpopulation specialization and multi-attribute evaluation, giving clear quantitative basis for co-evolution and Pareto screening, and are directly related to the clinical application needs of antimicrobial peptides, ensuring that the optimization results are not only effective in the laboratory but also have practical translational potential. While existing activity and toxicity predictions provide important references for antimicrobial peptide screening, their accuracy is still affected by factors such as data quality, coverage, and ambiguous sequence boundaries, making it difficult to fully reflect the true performance of candidate peptides in certain situations. Furthermore, high-throughput experimental validation has high requirements in terms of cost and time, and cannot quickly provide feedback on the massive amounts of sequences generated through computation. Therefore, this invention, based on fully utilizing existing prediction results, employs the aforementioned multi-objective synergistic optimization strategy to improve the overall quality of generated sequences from the source, reduce experimental validation costs and risks, and accelerate the development of efficient, safe, and stable antimicrobial peptides.
[0073] The overall structure of the co-evolutionary framework is shown in the attached diagram. Figure 3 As shown, the multi-objective optimization task of antimicrobial peptides is decomposed into three cooperative sub-evolutionary processes, corresponding to key pharmacological properties such as antimicrobial activity, hemolytic activity (toxicity), and half-life. The initial population generates progeny peptides through mutation or recombination strategies, which then enter three parallel attribute evaluation paths. Each path employs its own independent evaluation metrics and elite selection mechanism to update the population. After multiple iterations, when all paths meet the termination conditions, candidate peptide sequences with superior multi-attribute performance are output.
[0074] Therefore, the present application proposes a sequence optimization method that takes into account the triple objectives of antibacterial activity, host toxicity, and half-life. Starting from peptides with high activity and high toxicity or low activity and low toxicity, it optimizes the generation of novel antimicrobial peptide molecules with high activity, low toxicity, and long half-life, providing high-quality, convertible lead compounds for the development of anti-infective drugs and promoting the practical transformation of antimicrobial peptides from in vitro research to in vivo application.
[0075] In some embodiments, the target attributes also include thermal stability, serum stability, hydrophobicity, immunogenicity, biofilm penetration ability, and anti-endotoxin ability. Thermal stability is the ability of a peptide to maintain its structure and function under high-temperature conditions (such as storage or processing), measured by the heat denaturation temperature (Tm) or the percentage of activity retained after high-temperature treatment. Serum stability is the ability of a peptide to resist degradation by serum proteases (such as trypsin and pepsin), expressed as the percentage of remaining activity after serum incubation (e.g., >80% retention after 2 hours indicates stability). Hydrophobicity is the interaction characteristic of a peptide with water molecules, quantified by the GRAVY value (mean hydrophobicity index), affecting its binding ability to bacterial membranes (excessive hydrophobicity may increase hemolysis). Immunogenicity is the likelihood that a peptide will trigger an immune response (such as antibody production), assessed by predicting T-cell epitopes or in vitro immunogenicity experiments; low immunogenicity reduces the risk of allergies. Biofilm penetration ability is the ability of a peptide to penetrate bacterial biofilms (such as Pseudomonas aeruginosa biofilms), measured by the penetration depth or the reduction rate of viable bacteria within the biofilm, and is crucial for the treatment of chronic infections. Anti-endotoxin ability refers to the peptide's capacity to neutralize bacterial endotoxins (such as lipopolysaccharide LPS), expressed as the endotoxin neutralization rate, which can reduce the risk of complications such as sepsis. The expanded nine attributes cover the entire chain of requirements for antimicrobial peptides, from in vitro activity to in vivo drug formulation (such as thermal stability related to storage conditions, and immunogenicity related to clinical safety), overcoming the limitations of traditional optimization that focuses on activity rather than drug formulation. Each attribute can correspond to an independent subpopulation (such as a biomembrane penetration subpopulation and an anti-endotoxin subpopulation), which can be specifically optimized through targeted mutation (such as enhancing positive charge to improve membrane permeability), and then balanced by Pareto screening.
[0076] In some specific embodiments, the multi-attribute prediction model uses a weighted summation method to quantify and integrate the multiple target attributes of new candidate peptides. Based on this, a Pareto optimal screening strategy is introduced to ensure a reasonable balance among multiple target attributes. Attributes such as antibacterial activity, hemolytic activity, and half-life are converted into a unified score, with weights adjusted according to improvement needs. Pareto optimality compensates for the limitations of weighted summation, retaining sequences without weaknesses or with prominent attributes, ensuring a balance among multiple objectives.
[0077] In some specific embodiments, the multi-attribute prediction model specifically includes: extracting features of new candidate peptides through a preset feature extraction part and performing multi-source feature fusion using a multi-scale fusion strategy to generate a fused feature vector; performing deep encoding on the fused feature vector through a preset shared representation modeling part to extract global dependency information between residues, constructing a context-aware dynamic representation, and obtaining core semantic features that simultaneously associate antibacterial activity and hemolytic activity predictions; performing regression prediction based on the feature subspace related to the antibacterial mechanism in the core semantic features through a preset antibacterial activity prediction branch, and performing classification prediction based on the feature subspace related to toxicity in the core semantic features through a preset hemolytic activity scoring branch; and backpropagating the prediction errors of the two branches to the feature extraction part and the shared representation modeling part through a joint loss function to dynamically optimize the fusion ratio of multi-source feature fusion and the encoding weight of deep encoding. The multi-attribute prediction model uses a multi-task neural network architecture with a shared Transformer encoder and a branch prediction structure for attribute prediction.
[0078] Multi-attribute prediction models, as a type of multi-task learning model, are used to simultaneously evaluate the antibacterial activity and hemolytic activity of peptides, providing quantitative feedback for subsequent evolutionary optimization. The structure of this multi-task model is as follows: Figure 4 As shown, the whole system consists of three core parts: input feature extraction, shared representation modeling, and multi-task prediction.
[0079] First, the input amino acid sequence is processed by the feature extraction module and transformed into a high-dimensional vector representation. To enhance representation capabilities, this feature extraction module employs a multi-scale fusion strategy, integrating three feature sources: sequence-level semantic modeling, traditional physicochemical property encoding, and molecular structure diagram modeling. This module combines multiple representation methods for antimicrobial peptide sequences, including one-dimensional sequence representation, various traditional physicochemical property features, and graph neural network feature modeling based on molecular structure diagrams. Information is extracted in parallel through three paths: Transformer, manual feature encoding, and graph convolutional networks, and normalized, projected, and fused in a unified dimensional space, providing a comprehensive and rich expression foundation for subsequent multi-attribute prediction and optimization. In the sequence path, the peptide segment first transforms the amino acid symbols into a continuous vector representation through an embedding layer, and positional encoding information is added to maintain the sequence order. Subsequently, a multi-layer Transformer structure (including a multi-head attention mechanism, a feedforward network, and a normalization module) is used to extract global dependency information between residues, constructing a dynamic representation with context-aware capabilities. In the physicochemical feature path, various validated descriptors are used for encoding, such as DPC (dipeptide composition), QSO (sequence ordering coupling feature), PAAC (pseudo-amino acid composition), CKSAAGP (K-spacer component pairs), and PHYC (physicochemical property feature), to capture fundamental properties that significantly influence antibacterial activity and toxicity, such as hydrophobicity and charge distribution. These static regular features can supplement deep characterization, helping to improve the interpretability and robustness of the model. In addition to using amino acid physicochemical features, one-dimensional sequence features, and molecular graph features, the feature extraction module can also alternatively introduce protein language models (such as ESM, ProtXLNet, ProtBert), positional encoding, and structure prediction models (such as AlphaFold2 predicted structure maps) as input feature sources to further improve the ability to characterize the functional-structural coupling relationship.
[0080] In the structural path, for peptides with known or predicted chemical structures, a graph neural network (GNN) is introduced to model their molecular structure. Specifically, structures such as graph attention networks (GAT) and graph isomorphic networks (GIN) are used, treating atoms in the molecule as graph nodes and chemical bonds as edges. A message-passing mechanism is used to extract topological features and functional group distribution information at the molecular level, which are then mapped into a fusionable low-dimensional vector. These multi-source features are normalized and fused in a unified space and input to... Figure 1 The shared Transformer coding layer shown is used to uniformly extract core semantic information that is relevant to both antibacterial activity and toxicity prediction, reducing redundancy and enhancing generalization ability.
[0081] Finally, the model structure is divided into two independent prediction branches, corresponding to the antibacterial activity prediction and hemolytic activity scoring tasks, respectively. Each branch consists of an independent fully connected network, trained with the supervision signals of its respective task, supporting parallel inference. The model as a whole adopts a joint loss function optimization strategy, minimizing the prediction errors of both tasks simultaneously, thereby improving the accuracy and robustness of multi-attribute prediction. In some specific embodiments, for peptide sequences with known chemical structures, atoms in the molecule are treated as graph nodes and chemical bonds as edges. A message passing mechanism is used to extract molecular-level topological features and functional group distribution information, which are then mapped into a fusionable low-dimensional vector to model the molecular structure. This module can be called multiple times in the co-evolutionary optimization process to evaluate the attribute scores of candidate sequences from different generations, supporting a rapid feedback mechanism and dynamic adjustment of the optimization path, thereby significantly improving the overall performance and practicality of this invention in the multi-objective optimization task of antibacterial peptide sequences.
[0082] In practical applications, it can also be replaced by other mainstream deep model structures, such as shared BiLSTM networks, CNN-BiLSTM hybrid structures, fine-tuning frameworks based on pre-trained models (such as ESM, ProtT5), and even combined with ensemble learning methods (such as XGBoost, RandomForest, etc.) to achieve model architecture replacement. As long as it can achieve joint prediction of multiple peptide attributes and provide scores that can be used to optimize feedback, it can be considered an equivalent replacement of the present invention.
[0083] Experimental Validation: Using melivesin as the starting sequence, 10 potential candidate peptides were obtained through optimization using the method described in this application. To verify their antibacterial efficacy and drug-likeness, in vitro experiments were conducted to evaluate their antibacterial activity and hemolytic activity. The experimental results are attached. Figure 5 As shown in the figure. Experimental results indicate that candidate peptides T2, T4, T5, T6, T7, T10, and T11 exhibit varying degrees of antibacterial activity against *Escherichia coli*. Among them, T7 showed a significant antibacterial effect, with a minimum inhibitory concentration (MIC) close to 15 μM. Notably, none of the optimized peptides exhibited hemolytic activity against human erythrocytes within the tested concentration range, suggesting good biocompatibility and potential safety.
[0084] Original polypeptide sequence: GIGAVLKVLTTGLPALISWIKRKRQQ
[0085] Peptide sequence: PIGALMLKLHTGFPCMICAIKRKRVI
[0086] The minimum inhibitory concentration (MIC) of antimicrobial peptides against *Escherichia coli* ATCC25922 was determined using the broth microdilution method. In the experiment, the antimicrobial peptide sample was dissolved in a sterile solution (e.g., dH₂O) to prepare an initial concentration (e.g., 1000 μM), and then serially diluted in 96-well plates at a 1:2 ratio. The logarithmic growth phase *E. coli* bacterial culture (OD₂O₃) was then... 600 ≈0.5) diluted to approximately 5×10 4 After inoculating with CFU / mL, the solution was added to each well, with a final volume of 200 μL per well. Blank, negative, positive, and growth control samples were included. After incubation at 37°C for 18 hours, OD values were recorded. 600 The minimum concentration required to completely inhibit visible bacterial growth was determined as the MIC value. The experiment was repeated three times, and the average result was used to analyze antibacterial activity.
[0087] The hemolytic toxicity of antimicrobial peptides to mouse erythrocytes (MRBCs) was evaluated using a erythrocyte hemolysis assay. Blood was collected from healthy mice, anticoagulated, washed with PBS, and resuspended to a 5% MRBC suspension. The antimicrobial peptide sample was dissolved in PBS at a maximum concentration of 1000 μM, and then serially diluted 1:2 in 96-well plates. An equal volume of MRBC suspension and antimicrobial peptide solution was added to each well. PBS was used as a negative control, and ddH2O or Triton X-100 was used as a positive control. After incubation at 37°C for 60 minutes, the supernatant was collected by centrifugation, and the absorbance at 540 nm was measured. The hemolysis rate was calculated using the formula. A concentration-hemolysis rate curve was plotted, and the HC ratio was determined. 50 The experiment was repeated three times to ensure the reliability of the results. The results of the hemolysis experiment are shown in the attached figure. Figure 6 As shown.
[0088] A multi-target antimicrobial peptide sequence optimization system is provided to implement the multi-target antimicrobial peptide sequence optimization method described in any of the above embodiments. The system modules are as follows: Figure 7 As shown, it includes the following:
[0089] Input unit 1 is used to determine the initial peptide sequence to be processed and its improvement requirements. A sequence search tree is constructed using the initial peptide sequence as a node. The following operations are performed iteratively based on the sequence search tree.
[0090] Iterative optimization unit 2 is used to score the nodes in the sequence search tree according to the preset scoring rules, and select a node as the current node based on the score;
[0091] Based on the improvement requirements, point mutation operations are performed on the current node according to the preset mutation rules to generate multiple candidate peptides, and the candidate peptides are added as new nodes to the sequence search tree.
[0092] Based on the improvement requirements, each candidate peptide is copied to the corresponding subpopulation. Based on the co-evolutionary framework, the candidate peptides are locally optimized using the subpopulation to obtain new candidate peptides. Each new candidate peptide is evaluated by a preset multi-attribute prediction model.
[0093] The evaluation results of each new candidate peptide are fed back to the corresponding current node in the sequence search tree. Each node is re-scored based on the evaluation results and the current node is reselected.
[0094] Output unit 3 is used to select peptide sequences from the sequence search tree based on the score when the iteration stopping condition is met, and output the peptide sequences that meet the improvement requirements.
[0095] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a multi-target antimicrobial peptide sequence optimization method. Applying a multi-target antimicrobial peptide sequence optimization method to a computer program product facilitates execution.
[0096] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the multi-target antimicrobial peptide sequence optimization method as described above.
[0097] The computer storage medium of this application can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. Computer-readable storage media can be, for example, but not limited to: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. This application applies a multi-target antimicrobial peptide sequence optimization method to a computer-readable storage medium storing a computer program. When executed by a processor, this program implements the method steps provided in this application, which is simple, fast, easy to store, and not easily lost.
[0098] Those skilled in the art will understand that the modules described above can be implemented using general-purpose computing systems. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Optionally, they can be implemented using computer-executable program code, allowing them to be stored in a storage system for execution by the computing system. Alternatively, they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0099] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.
[0100] The above disclosures are only a few specific implementation scenarios of this application. However, this application is not limited to these. Any variations that can be conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. A method for optimizing multi-target antimicrobial peptide sequences, characterized in that, Including the following: Determine the initial peptide sequence to be processed and its improvement requirements. Construct a sequence search tree using the initial peptide sequence as nodes. Iteratively perform the following operations based on the sequence search tree: The nodes in the sequence search tree are scored according to the preset scoring rules, and a node is selected as the current node based on the score. Based on the aforementioned improvement requirements, point mutation operations are performed on the current node according to preset mutation rules to generate multiple candidate peptides, and the candidate peptides are added as new nodes to the sequence search tree. Based on the aforementioned improvement requirements, each candidate peptide is copied to its corresponding subpopulation. Based on the co-evolutionary framework, the candidate peptides are locally optimized using the subpopulations to obtain new candidate peptides. Each new candidate peptide is then evaluated using a preset multi-attribute prediction model. The evaluation results of each new candidate peptide are fed back to the corresponding current node in the sequence search tree. Each node is re-scored based on the evaluation results and the current node is reselected. When the iteration stopping condition is met, a peptide sequence is selected from the sequence search tree based on the score as the peptide sequence that meets the improvement requirements and output. The scoring rules include: a comprehensive score based on the number of times a node is selected, the number of times its corresponding parent node is selected, and the node's evaluation result; the expression for the scoring rules is: Where C1 is the cumulative reward derived from all evaluation results for this node. This represents the number of times a node is accessed. To explore the weighting factors, This represents the number of times the parent node of this node has been visited.
2. The multi-target antimicrobial peptide sequence optimization method according to claim 1, characterized in that, Each subpopulation corresponds to the optimization of a certain attribute objective. Based on the improvement requirements, the corresponding attribute objective is determined, and then the appropriate subpopulation is selected for optimization.
3. The multi-target antimicrobial peptide sequence optimization method according to claim 1, characterized in that, The multi-attribute prediction model uses a weighted summation method to quantify and integrate the multi-target attributes of new candidate peptides, and introduces a Pareto optimal screening strategy to ensure a reasonable balance among multiple target attributes.
4. The multi-target antimicrobial peptide sequence optimization method according to claim 3, characterized in that, The target properties include antibacterial activity, hemolysis, half-life, thermal stability, serum stability, hydrophobicity, immunogenicity, biomembrane penetration ability, and anti-endotoxin ability.
5. The multi-target antimicrobial peptide sequence optimization method according to claim 1, characterized in that, The mutation rules include: When the sequence search tree contains only the initial peptide sequence node, multiple peptide sequences are generated by introducing random points to guide mutation of the initial peptide sequence, thereby forming a diverse initial population. The number of peptide sequences in the initial population is set proportionally according to the length of the initial peptide sequence. When the sequence search tree has multiple nodes, the peptide sequences corresponding to two parent nodes are randomly exchanged and recombinated according to a preset crossover rate to generate new child nodes, or the peptide sequences corresponding to child nodes generated by mutations are guided to mutate again based on a preset mutation rate to generate new child nodes.
6. The multi-target antimicrobial peptide sequence optimization method according to claim 1, characterized in that, The multi-attribute prediction model specifically includes: Features of new candidate peptides are extracted by a preset feature extraction part and multi-source feature fusion is performed by a multi-scale fusion strategy to generate a fused feature vector. By performing deep encoding on the fused feature vector through the pre-defined shared representation modeling part, global dependency information between residues is extracted, and a dynamic representation with context-aware capability is constructed to obtain the core semantic features that simultaneously associate antibacterial activity and hemolytic activity prediction. Regression prediction is performed based on the feature subspace related to the antibacterial mechanism in the core semantic features through a preset antibacterial activity prediction branch, and classification prediction is performed based on the feature subspace related to the toxicity in the core semantic features through a preset hemolyticity scoring branch. The prediction errors of the two branches are backpropagated to the feature extraction part and the shared representation modeling part through a joint loss function, thereby dynamically optimizing the fusion ratio of multi-source feature fusion and the encoding weight of deep encoding.
7. The multi-target antimicrobial peptide sequence optimization method according to claim 1, characterized in that, For peptide sequences with known chemical structures, atoms in the molecule are treated as graph nodes and chemical bonds as edges. The topological features and functional group distribution information at the molecular level are extracted through a message passing mechanism and finally mapped into a fusionable low-dimensional vector to achieve molecular structure modeling.
8. A multi-target antimicrobial peptide sequence optimization system, characterized in that, A method for implementing a multi-target antimicrobial peptide sequence optimization method according to any one of claims 1-7 includes the following: The input unit is used to determine the initial peptide sequence to be processed and its improvement requirements. A sequence search tree is constructed using the initial peptide sequence as nodes, and the following operations are iteratively performed based on the sequence search tree: The iterative optimization unit is used to score the nodes in the sequence search tree according to the preset scoring rules, and select a node as the current node based on the score. Based on the aforementioned improvement requirements, point mutation operations are performed on the current node according to preset mutation rules to generate multiple candidate peptides, and the candidate peptides are added as new nodes to the sequence search tree. Based on the aforementioned improvement requirements, each candidate peptide is copied to its corresponding subpopulation. Based on the co-evolutionary framework, the candidate peptides are locally optimized using the subpopulations to obtain new candidate peptides. Each new candidate peptide is then evaluated using a preset multi-attribute prediction model. The evaluation results of each new candidate peptide are fed back to the corresponding current node in the sequence search tree. Each node is re-scored based on the evaluation results and the current node is reselected. The output unit is used to select a peptide sequence from the sequence search tree based on the score as the peptide sequence that meets the improvement requirements and output it when the iteration stopping condition is met. The scoring rules include: a comprehensive score based on the number of times a node is selected, the number of times its corresponding parent node is selected, and the node's evaluation result; the expression for the scoring rules is: Where C1 is the cumulative reward derived from all evaluation results for this node. This represents the number of times a node is accessed. To explore the weighting factors, This represents the number of times the parent node of this node has been visited.
Citation Information
Patent Citations
Method for predicting antibacterial activity of antibacterial peptide
CN116884472A
Antibacterial peptide screening framework based on protein language model and biological information calculation software
CN120048357A