Antibacterial peptide generation and screening method based on conditional potential diffusion model

By employing a conditional potential diffusion model and a multi-stage screening process, the challenges of sequence diversity and stability in antimicrobial peptide generation and screening methods were addressed, enabling efficient and reliable antimicrobial peptide discovery and generating candidate sequences with broad-spectrum activity.

CN121789781APending Publication Date: 2026-04-03SOUTHWEST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for generating and screening antimicrobial peptides struggle to maintain the stability and controllability of the results while expanding sequence diversity, and lack systematic screening strategies, thus limiting the efficiency of antimicrobial peptide discovery.

Method used

An antimicrobial peptide generation and screening method based on a conditional latent diffusion model is adopted. The peptide sequence is mapped to a continuous latent space through a variational autoencoder. Combined with the conditional diffusion generation mechanism and a multi-level screening process, including data preprocessing, multi-model prediction and experimental verification, the automatic generation and efficient screening of antimicrobial peptide sequences are achieved.

Benefits of technology

The generated antimicrobial peptide sequences were structurally sound and uniformly distributed, which improved the efficiency and reliability of antimicrobial peptide discovery, reduced the redundancy of candidate sequences, and enhanced broad-spectrum antimicrobial activity and diversity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789781A_ABST
    Figure CN121789781A_ABST
Patent Text Reader

Abstract

The invention discloses an antibacterial peptide generating and screening method based on a conditional potential diffusion model, which comprises the following steps: modeling antibacterial peptide sequence distribution in a continuous submerged space, realizing controllable generation of antibacterial peptide sequences through conditional constraints, and improving the efficiency and reliability of an antibacterial peptide discovery process by combining a multi-stage calculation screening and experimental verification strategy; therefore, the problem that in an existing antibacterial peptide design method, due to the factors that the sequence space scale is huge, the generation process is difficult to control, the candidate sequence redundancy is high, and the screening cost is high, the antibacterial peptide discovery efficiency is limited is solved. According to the method, a conditional diffusion generation mechanism is introduced into a continuous submerged space, and a multi-stage screening process formed by calculation screening and experimental verification is combined, so that systematic design and optimization of an antibacterial peptide sequence from data driven generation to performance evaluation are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, and more specifically to a method for screening the generation of antimicrobial peptides based on a conditional potential diffusion model. Background Technology

[0002] Combating microbial infections has long been a significant challenge in clinical medicine and public health. With the continued use of antibiotics, pathogen resistance is constantly increasing, and multidrug-resistant strains are emerging and spreading, leading to a decline in the therapeutic efficacy of some traditional small-molecule antibiotics and increasing the difficulty of preventing and controlling related infectious diseases. Against this backdrop, novel anti-infective strategies with different mechanisms of action have attracted widespread attention.

[0003] Antimicrobial peptides, a class of functional polypeptides with natural or artificial sources, typically exert their antibacterial effects by disrupting the cell membrane structure of microorganisms or interfering with their physiological functions. They are characterized by diverse mechanisms of action, rapid onset of action, and relatively low risk of drug resistance, and can cover a variety of pathogen types, including bacteria and fungi. However, the sequence design space for antimicrobial peptides is vast, with a complex search space composed of amino acid types, sequence lengths, and structural combinations. This makes it difficult for traditional methods relying on empirical rules or experimental screening to efficiently discover antimicrobial peptide sequences with superior performance.

[0004] To improve the efficiency of antimicrobial peptide design, some existing technologies have introduced computational models to assist in the generation or prediction of peptide sequences. However, these methods still have shortcomings in practical applications. On the one hand, existing methods often struggle to ensure the stability and controllability of the generated results while expanding sequence diversity during the antimicrobial peptide sequence generation process. The generated sequences are easily limited by the distribution of training samples, leading to high repetition rates or inconsistent quality of candidate sequences. On the other hand, the screening process for generated sequences is usually fragmented, relying heavily on a single activity prediction indicator and lacking a systematic and hierarchical screening strategy. This makes it difficult to simultaneously consider antimicrobial activity, broad spectrum, and potential application performance.

[0005] Furthermore, while some existing technologies attempt to combine generative models with screening methods, the synergy between the generation and screening processes is limited, failing to form a complete closed loop from sequence generation and computational evaluation to experimental verification. This results in room for improvement in the overall efficiency and reliability of the antimicrobial peptide discovery process. Therefore, how to construct a unified technical solution that balances antimicrobial peptide generation and screening, ensuring sequence diversity while efficiently locating candidate antimicrobial peptide sequences with application potential through a systematic screening process, remains a pressing technical problem to be solved in this field. Summary of the Invention

[0006] In view of this, one of the objectives of this invention is to provide an antimicrobial peptide generation and screening method based on a conditional latent diffusion model. The aim is to achieve automated generation and efficient screening of antimicrobial peptide sequences, thereby addressing the limitations in antimicrobial peptide discovery efficiency caused by factors such as the large sequence space, difficulty in controlling the generation process, high candidate sequence redundancy, and high screening costs in existing antimicrobial peptide design methods. This invention introduces a conditional diffusion generation mechanism into a continuous latent space and combines it with a multi-level screening process consisting of computational screening and experimental verification, achieving a systematic design and optimization of antimicrobial peptide sequences from data-driven generation to performance evaluation.

[0007] To achieve the above objectives, the present invention provides the following technical solution: 1. A method for screening antimicrobial peptides based on a conditional potential diffusion model, comprising: (1) Data acquisition: Obtain experimentally validated antimicrobial peptide sequences from publicly available antimicrobial peptide databases, and construct non-antimicrobial peptide sequence data with consistent length distribution to form training and validation datasets; (2) Data preprocessing: The peptide sequences are standardized and encoded, including amino acid letter mapping, length alignment and invalid sequence filtering; (3) Latent space coding: using variational autoencoder (VAE) through encoder network Discrete peptide sequences are mapped to a continuous latent space to obtain the probability distribution of latent variables; this is then processed through a decoder network. Reconstructing the original sequence from latent variables; model training is achieved by maximizing the lower bound of evidence, ELBO. (4) Conditional diffusion modeling: A conditional diffusion model is constructed in the latent space. Noise is gradually added to the latent variable of antimicrobial peptide through a forward diffusion process, and a reverse denoising process is learned under the constraint of antimicrobial property condition c. To model the distribution of potential representations of antimicrobial peptides; optionally, a fine-tuning process based on minimum inhibitory concentration (MIC) data is introduced to adjust model parameters using MIC information; (5) Antimicrobial peptide generation: Starting from random noise, backsampling is performed using a trained conditional diffusion model to obtain a latent representation, and then candidate antimicrobial peptide sequences are generated through a VAE decoder. (6) Screening of candidate antimicrobial peptides: First, multiple antimicrobial activity prediction models based on convolutional neural networks, recurrent neural networks and self-attention mechanisms are used to independently predict and jointly screen the generated candidate peptide sequences; then, similarity calculation and cluster analysis are performed on the sequences that pass the activity screening to remove highly similar or repetitive sequences and obtain a set of candidate sequences with redundancy removed. (7) Minimum inhibitory concentration prediction screening: Construct a minimum inhibitory concentration prediction model based on attention mechanism, perform MIC prediction on candidate sequences after redundancy removal, and screen out sequences with predicted MIC below a set threshold. (8) Screening of broad-spectrum antimicrobial peptides: Using the minimum inhibitory concentration prediction sub-models for Gram-positive and Gram-negative bacteria respectively, candidate sequences are evaluated and sequences with low prediction MICs against multiple pathogens are retained to obtain a candidate set of broad-spectrum antimicrobial peptides.

[0008] Preferably, in step (4) of this invention, the forward process of the conditional diffusion model is defined as: in , The preset noise control coefficient, These are initial latent variables used to characterize the potential features of antimicrobial peptides. The variables after adding noise at step t are used to train the model to learn the denoising process of generating potential representations of antimicrobial peptides from noisy states.

[0009] Preferably, in step (4) of this invention, the objective function of the fine-tuning process of the MIC data is: , in The basic training loss for the diffusion model, is a weighting function constructed from the minimum inhibitory concentration, used to enhance the contribution of highly active samples, and λ is a balance coefficient used to adjust the relative weight between the MIC weighted loss term and the basic training loss of the diffusion model.

[0010] Preferably, in step (6) of this invention, the antibacterial activity screening rule is as follows: candidate sequence Retained if and only if in Indicates the first Candidate antimicrobial peptide sequences, Indicates the first The antibacterial activity threshold corresponding to each prediction model This represents the minimum number of models required to satisfy the criteria after filtering. For indicator functions, Let represent the predicted antimicrobial activity value of the i-th antimicrobial activity prediction model for the i-th candidate antimicrobial peptide sequence. This represents the set of antimicrobial peptide candidate sequences obtained after screening by multiple prediction models with consistency constraints.

[0011] Preferably, in step (6) of this invention, the sequence diversity screening includes: calculating the pairwise similarity between candidate sequences. Clustering is performed based on similarity, and only one or more representative sequences are retained for each cluster. This represents the sequence similarity metric function, using cosine similarity.

[0012] Preferably, in step (7) of this invention, the minimum inhibitory concentration prediction model adopts a low-rank self-attention mechanism to adaptively focus on amino acid fragments in the sequence that contribute significantly to antibacterial activity.

[0013] Preferably, in step (8) of this invention, the screening of broad-spectrum antimicrobial peptides specifically involves: inputting candidate sequences into the MIC prediction sub-models for Gram-positive and Gram-negative bacteria, respectively, to obtain predicted values. in Indicates candidate antimicrobial peptide sequence Predicted minimum inhibitory concentration for Gram-positive bacteria Indicates candidate antimicrobial peptide sequence Predicted minimum inhibitory concentration for Gram-negative bacteria.

[0014] The beneficial effects of this invention are as follows: This invention provides a method for antimicrobial peptide generation and screening based on a conditional latent diffusion model, which can obtain antimicrobial peptide sequences with reasonable structure and consistent distribution. Unsupervised distribution analysis based on protein language model features shows that the model-generated sequences exhibit a high degree of overlap with natural antimicrobial peptides in the high-dimensional feature space, while maintaining a clear distinction from non-antimicrobial peptides. This indicates that the model can effectively capture the latent sequence characteristics and statistical regularities of antimicrobial peptides without explicitly introducing classification supervision. Further physicochemical property analysis results show that the generated sequences are consistent with the distribution of known antimicrobial peptides in terms of amino acid composition, net charge, isoelectric point, and secondary structure-related indicators, especially showing a reasonable bias in key antimicrobial-related properties such as positive charge and hydrophobicity. This phenomenon indicates that diffusion modeling based on latent space can constrain the generated results to fall within a biologically reasonable antimicrobial peptide subspace while preserving sequence diversity, laying a reliable foundation for subsequent functional screening. Attached Figure Description

[0015] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration: Figure 1 A flowchart illustrating the steps of an antimicrobial peptide generation and screening method. Figure 2 Distribution characteristics and comparative analysis of antimicrobial peptide sequences; Figure 3 The amino acid composition distribution of peptide sequences generated by different methods; Figure 4 For aromaticity analysis; Figure 5 For α-spiral ratio analysis; Figure 6 For isoelectric point analysis; Figure 7 Net charge analysis; Figure 8Sequence consistency and similarity analysis. Detailed Implementation

[0016] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.

[0017] Example 1 (1) Data acquisition: We collected experimentally validated antimicrobial peptide sequences from multiple publicly available antimicrobial peptide databases (APD3, DRAMP, and DBAASP, etc.) and constructed a dataset of non-antimicrobial peptide sequences with consistent length distribution to form a dataset for model training and validation.

[0018] (2) Data preprocessing: The peptide sequences are standardized through encoding processes, including amino acid letter mapping, length alignment, and invalid sequence filtering, to meet the model input requirements.

[0019] (3) Latent space coding: The preprocessed peptide sequences are feature extracted using a sequence coding network, and the discrete amino acid sequences are mapped to a continuous latent space representation.

[0020] Let a peptide sequence of length L be represented as: in This represents a finite alphabet consisting of 20 standard amino acids and necessary special symbols. To eliminate the adverse effects of symbol discreteness on neural network modeling, each amino acid symbol is first converted into a continuous vector representation using an embedding mapping function: This transforms the original sequence into an embedding vector sequence. .

[0021] Based on this, the encoder network The embedded sequences are modeled to learn the probabilistic mapping from the sequence space to the latent space. The encoder aggregates features based on the internal order relationships and contextual dependencies of the sequence, outputting the conditional distribution parameters of the latent variables, typically modeled as a multidimensional Gaussian distribution. in These represent the mean and standard deviation vectors of the latent variables predicted by the encoding network, respectively. These are the encoder parameters. This design ensures that each peptide sequence corresponds to a continuous probability distribution in the latent space, rather than a single point representation, thereby enhancing the robustness and smoothness of the latent space representation. To achieve a differentiable random sampling process, a reparameterization technique is used to sample the latent variables during the model training phase: This operation ensures that randomness originates only from independent noise variables. This ensures that the model parameters can be effectively optimized during backpropagation. Corresponding to the encoder is the decoder network. Responsible for latent variables Reconstructing the original peptide sequence is essentially a conditional sequence generation model that progressively predicts the probability distribution of amino acid symbols given a potential representation. The reconstruction probability output by the decoder can be expressed as: in Indicates decoder parameters, Indicates position The previous amino acid sequence. This autoregressive form can fully capture the sequence dependence characteristics in the peptide sequence.

[0022] The training objective of the VAE model is achieved by maximizing the Evidence Lower Bound (ELBO), and its optimization objective function can be expressed as: The first term is the reconstruction term, which measures the decoder's ability to reconstruct the original sequence under latent space conditions; the second term is the KL divergence term, which constrains the difference between the latent variable distribution obtained by encoding and the prior distribution p(z) (usually set as a standard normal distribution), thereby guiding the latent space to form a structured, continuous and sampleable distribution.

[0023] (4) Conditional diffusion modeling: A conditional diffusion model is constructed in the latent space. By progressively injecting Gaussian noise and learning an inverse denoising process under the constraint of antibacterial properties, the latent distribution model is completed. The specific method is as follows: The preprocessed antimicrobial peptide sequences were mapped to a continuous latent space to obtain the corresponding latent representations. The corresponding conditional label is ,in Indicates antimicrobial peptide, This represents a non-antimicrobial peptide. The diffusion model gradually injects Gaussian noise into the latent representation through a forward diffusion process, constructing a series of intermediate variables. Its forward process is defined as: in , The noise scheduling coefficient is preset. As the number of diffusion steps increases, the latent representation gradually evolves towards a standard Gaussian distribution, thereby achieving unified perturbation modeling for antibacterial and non-antibacterial latent space representations.

[0024] And by learning the corresponding reverse denoising process In condition information The distribution characteristics of potential representations of antimicrobial peptides are modeled under constraints, thereby generating potential representations of candidate antimicrobial peptides.

[0025] After model training is completed or before model generation, a fine-tuning process based on minimum inhibitory concentration (MIC) data is introduced into the conditional latent diffusion model. Using known antimicrobial peptide sequences and their corresponding MIC information, the model parameters are adjusted. The fine-tuning objective function can be expressed as follows: in The basic training loss for the diffusion model, The weighting function is constructed based on the minimum inhibitory concentration (MIC) and is used to enhance the contribution of low MIC (high inhibitory activity) antimicrobial peptide samples in the model optimization process.

[0026] (5) Antimicrobial peptide generation: In the generation phase, starting from a random noise vector, backsampling is performed under the guidance of a conditional latent diffusion model and its fine-tuning parameters to obtain a representation of the potential antimicrobial peptide. This representation is then restored to a candidate antimicrobial peptide sequence via a decoding network. The fine-tuning method based on minimum inhibitory concentration (MIC) data is not limited to a specific implementation form and can employ parameter updates, weight adjustments, or conditional guidance enhancements to ensure the flexibility and replaceability of the generation process.

[0027] (6) Screening of candidate antimicrobial peptides: The generated candidate peptide sequences were initially evaluated using an antimicrobial peptide activity prediction model, and sequences with predicted antimicrobial activity below a preset threshold were removed. Then, based on sequence similarity calculation and cluster analysis, the candidate antimicrobial peptides that passed the initial screening were subjected to diversity screening to remove highly similar or repetitive sequences, thereby improving the coverage and innovation of the candidate set.

[0028] The antimicrobial peptide activity screening utilizes sequence prediction models based on convolutional neural networks, recurrent neural networks, and self-attention mechanisms to independently predict the antimicrobial activity of candidate antimicrobial peptide sequences, and then progressively filters candidate sequences based on the prediction results of multiple models. After generating candidate antimicrobial peptide sequences, to improve the reliability and robustness of the screening results, this invention employs multiple antimicrobial activity prediction models with different structures to independently evaluate candidate sequences, and progressively filters them based on multiple prediction results.

[0029] Let the set of candidate antimicrobial peptide sequences be... in Indicates the first A number of candidate antimicrobial peptide sequences were identified. First, each candidate sequence was represented using sequence embedding: in This represents a sequence coding function used to map an amino acid sequence into a numerical feature representation.

[0030] Subsequently, the antimicrobial activity of the candidate antimicrobial peptide sequences was independently predicted using a convolutional neural network model, a recurrent neural network model, and a Transformer model based on a self-attention mechanism, respectively. The corresponding prediction results are expressed as follows: in , and These represent antimicrobial activity prediction models based on convolutional neural networks, recurrent neural networks, and Transformers, respectively.

[0031] Based on this, candidate antimicrobial peptide sequences are filtered multiple times according to preset antimicrobial activity screening rules. Only when a candidate sequence meets the antimicrobial activity threshold requirement in at least two prediction models is the sequence retained for the next stage of screening. The screening rules can be expressed as follows: in Indicates the first The antibacterial activity threshold corresponding to each prediction model This represents the minimum number of models required to satisfy the criteria after filtering. This is an indicator function.

[0032] By employing the aforementioned multi-model independent prediction and multiple filtering mechanisms, the risk of misjudgment can be reduced and the stability and generalization of screening results can be improved without relying on a single prediction model. At the same time, it provides a higher quality set of candidate antimicrobial peptide sequences for subsequent sequence diversity screening and minimum inhibitory concentration prediction.

[0033] Sequence diversity screening includes the following steps: After completing the antimicrobial activity prediction screening, a set of candidate antimicrobial peptide sequences that passed the initial screening was obtained. Each candidate sequence meets the preset antibacterial activity requirements.

[0034] First, the similarity of the candidate antimicrobial peptide sequences is calculated. The similarity between any two candidate antimicrobial peptide sequences is calculated using sequence alignment or feature representation methods. in This represents a sequence similarity metric function used to reflect the degree of similarity between two antimicrobial peptide sequences in terms of amino acid composition, sequence, or feature representation.

[0035] Subsequently, based on the similarity calculation results, cluster analysis is performed on the candidate antimicrobial peptide sequences, grouping sequences with similarity higher than a preset threshold into the same category. For each category of candidate sequences, preferably only one or more representative sequences are retained, while other highly similar or repetitive sequences are removed, thus obtaining a deredundant set of candidate antimicrobial peptide sequences. Through the above sequence diversity screening process, the redundancy between candidate antimicrobial peptide sequences can be effectively reduced without relying on specific similarity calculation methods or clustering algorithms. This improves the distribution coverage and diversity of the screened antimicrobial peptide candidate set in the sequence space, providing more representative candidate sequences for subsequent minimum inhibitory concentration prediction and experimental verification.

[0036] (7) Screening based on prediction of minimum inhibitory concentration: After sequence diversity screening, a minimum inhibitory concentration (MIC) prediction model based on a low-rank self-attention mechanism was constructed to further screen candidate antimicrobial peptides. The prediction model was trained based on MIC data of Gram-negative and Gram-positive bacteria to evaluate the inhibitory ability of candidate antimicrobial peptides against different types of bacteria.

[0037] After completing the sequence diversity screening, a set of candidate antimicrobial peptide sequences after redundancy removal was obtained. The candidate antimicrobial peptide sequences showed high differences at the sequence level.

[0038] First, the candidate antimicrobial peptide sequences are characterized by mapping the amino acid sequences to sequence features suitable for minimum inhibitory concentration (MIC) prediction: in This represents a sequence feature extraction function used to generate sequence representations that include amino acid composition and contextual information.

[0039] Subsequently, a minimum inhibitory concentration (MIC) prediction model based on an attention mechanism was constructed to evaluate candidate antimicrobial peptide sequences. This prediction model adaptively focuses on amino acid positions or fragments in the sequence that contribute significantly to antimicrobial activity through an attention weight allocation mechanism, thereby outputting the corresponding MIC prediction value. in Indicates candidate antimicrobial peptide sequence Predicted minimum inhibitory concentration for the target pathogen.

[0040] After obtaining the prediction results, the predicted minimum inhibitory concentration is compared with a preset screening threshold, and candidate antimicrobial peptide sequences are further screened according to the following rules: in This represents the screening threshold for the minimum inhibitory concentration. This represents the set of candidate antimicrobial peptide sequences selected based on the minimum inhibitory concentration prediction.

[0041] By using the minimum inhibitory concentration (MIC) prediction screening steps described above, the antibacterial ability of candidate antimicrobial peptide sequences can be quantitatively evaluated without relying on specific model structures or training methods. This allows for the screening of antimicrobial peptide sequences with high potential antimicrobial effects, providing a reliable candidate set for subsequent screening and experimental verification of broad-spectrum antimicrobial peptides.

[0042] (8) Screening of broad-spectrum antimicrobial peptides: Using the minimum inhibitory concentration (MIC) prediction model, candidate antimicrobial peptides with lengths within a preset range are screened, and sequences with low predicted MICs against multiple pathogens are retained, thereby obtaining a candidate set of antimicrobial peptides with potential broad-spectrum antimicrobial activity.

[0043] The minimum inhibitory concentration (MIC) prediction model includes multiple prediction sub-models constructed for different types of pathogens, and specifically includes the following steps: Based on experimentally determined minimum inhibitory concentration (MIC) data, MIC prediction models were constructed for Gram-positive and Gram-negative bacteria, respectively. The prediction model for Gram-positive bacteria was trained using MIC data for the corresponding bacterial species to evaluate the inhibitory ability of candidate antimicrobial peptide sequences against Gram-positive pathogens; the prediction model for Gram-negative bacteria was also trained using MIC data for the corresponding bacterial species to evaluate the inhibitory ability of candidate antimicrobial peptide sequences against Gram-negative pathogens.

[0044] During the prediction phase, the candidate antimicrobial peptide sequences obtained through the aforementioned screening steps are input into the Gram-positive bacteria prediction model and the Gram-negative bacteria prediction model, respectively, to obtain the corresponding minimum inhibitory concentration (MIC) prediction results: in Indicates candidate antimicrobial peptide sequence Predicted minimum inhibitory concentration for Gram-positive bacteria Indicates candidate antimicrobial peptide sequence Predicted minimum inhibitory concentration for Gram-negative bacteria.

[0045] Based on the prediction results, candidate antimicrobial peptide sequences are comprehensively evaluated according to preset screening rules. Antimicrobial peptide sequences that exhibit low predicted minimum inhibitory concentrations in both Gram-positive and Gram-negative bacteria prediction models, or that meet preset inhibition thresholds in at least one type of pathogen, are retained, thereby obtaining a candidate set of antimicrobial peptides with potential inhibitory effects against multiple pathogens.

[0046] By constructing prediction models based on minimum inhibitory concentration (MIC) data of different types of pathogens and conducting joint screening, the reliability and generalizability of the screened antimicrobial peptides in terms of antimicrobial spectrum coverage can be improved without relying on the prediction results of a single bacterial species, thus providing effective support for the discovery of broad-spectrum antimicrobial peptides.

[0047] Example 2 After the computational screening was completed, representative sequences were selected from the selected antimicrobial peptide candidate set for experimental verification. Through minimum inhibitory concentration determination and related performance evaluation, their antimicrobial activity, biological characteristics and application potential were comprehensively evaluated to obtain antimicrobial peptide candidate sequences with practical application value.

[0048] 1) Distribution characteristics and comparative analysis of generated antimicrobial peptide sequences The t-SNE spatial distribution of antimicrobial peptides, non-antimicrobial peptides, and generated antimicrobial peptides was analyzed, and the results are as follows: Figure 2As shown in the figure. In the analysis of experimental results, different categories of sequences and methods are distinguished by a unified identifier to facilitate a direct comparison of their distribution characteristics and performance differences. Representative antimicrobial peptide sequences and the results generated by the proposed method are the focus of analysis, used to characterize the model's ability to model in the antimicrobial peptide feature space. Various existing generation models and baseline methods are used as comparison objects to evaluate the differences in sequence rationality and diversity among different generation strategies. Natural non-antimicrobial peptide sequences are used as a reference distribution to characterize the background boundaries of non-antimicrobial sequences at the feature space and physicochemical property levels. The experimental results show that the antimicrobial peptide sequences generated by this invention are closer to the actual antimicrobial peptides in overall distribution, while maintaining a clear distinction from the non-antimicrobial peptide reference distribution, indicating that the model can effectively capture the key statistical and structural features of antimicrobial peptides in the unsupervised feature space. In contrast, some comparative methods generate sequences that are more dispersed in distribution or biased towards the reference distribution to some extent, reflecting their shortcomings in antimicrobial feature modeling or generation stability. The above results, from the perspective of distribution consistency and comparative analysis, verify the effectiveness of the proposed method in the task of antimicrobial peptide generation, and provide a reasonable basis for subsequent physicochemical property analysis and multi-level screening results.

[0049] 2) Distribution characteristics and comparative analysis of generated antimicrobial peptide sequences The amino acid distribution of peptide sequences generated by different methods was analyzed, and the results are as follows: Figure 3 As shown in the figure, the horizontal axis represents 20 standard amino acid types, and the vertical axis represents the frequency of occurrence of the corresponding amino acid in the generated peptide sequence. It can be observed that different generation methods show significant differences in amino acid preferences. The methods based on the latent diffusion model and variational autoencoder generally maintain amino acid distribution characteristics similar to the real positive samples, while some unconstrained generation methods show significant shifts in the distribution of charged or hydrophobic amino acids (such as K, L, R, etc.). In contrast, the UniProt database and randomly generated sequences exhibit a relatively smooth and nearly uniform distribution trend, reflecting the statistical characteristics of its role as a reference distribution. Overall, these results indicate that the proposed generation model can effectively learn the intrinsic distribution patterns of antimicrobial peptide sequences while maintaining the overall rationality of the amino acid composition.

[0050] 3) Aromaticity analysis Aromaticity is an important physicochemical indicator reflecting the ability of antimicrobial peptides to interact with bacterial membranes. The enrichment and distribution of aromatic amino acid residues (F, Y, W) in the sequence have a key impact on membrane binding stability and insertion behavior. Figure 4This study compared the aromaticity distribution of different generation strategies with that of real antimicrobial peptides. It was observed that real antimicrobial peptides (Positive) exhibit a relatively concentrated aromaticity distribution range, reflecting clear structural constraints and functional orientation in their sequence design. In contrast, the aromaticity distribution of random sequences, UniProt sequences, and negative samples was more dispersed or generally lower, indicating a lack of the directional structural features required for antimicrobial peptides. Notably, generative model-based methods, especially the latent diffusion model and some comparative models, closely approximate the aromaticity distribution of real antimicrobial peptides, maintaining consistency not only at the median level but also showing good matching in distribution width. This demonstrates that such models can effectively learn and reproduce the statistical regularity of aromatic residues along the sequence direction in antimicrobial peptide sequences, rather than simply replicating the overall amino acid proportions. This result further validates the effectiveness of the proposed model in preserving the key directional physicochemical characteristics of antimicrobial peptides.

[0051] 4) α-spiral ratio analysis The α-helix ratio is a key structural indicator characterizing the conformational stability of antimicrobial peptides and their ability to interact with cell membranes. Numerous studies have shown that antimicrobial peptides with continuous or semi-continuous α-helix structures are more prone to conformational changes in hydrophobic environments, thereby promoting membrane insertion and disruption. Figure 5 It is evident that positive antimicrobial peptides exhibit a high and relatively concentrated distribution of α-helix ratios, reflecting their tendency to form stable secondary structures during evolution. In contrast, random sequences, UniProt sequences, and negative samples generally show low and dispersed α-helix ratios, indicating their difficulty in forming stable helical conformations. Notably, various generative model-based methods show high consistency with positive antimicrobial peptides on this metric, especially the latent diffusion model and some deep generative models, whose median α-helix ratios and distribution ranges highly overlap with those of positive samples. This suggests that the models not only maintain amino acid composition characteristics during generation but also effectively learn the intrinsic rules governing the formation of α-helix structures along the sequence direction of antimicrobial peptides. This result further demonstrates that the model possesses good biophysical rationality at the structural level.

[0052] 5) Isoelectric point analysis The isoelectric point is an important physicochemical parameter reflecting the overall charge state of antimicrobial peptides and their ability to interact with negatively charged bacterial membranes. Under physiological conditions, a higher isoelectric point usually indicates a sequence rich in basic residues, which facilitates the initial binding of antimicrobial peptides to the bacterial membrane surface via electrostatic adsorption. Figure 6As can be seen, the isoelectric point distribution of real antimicrobial peptides (Positive) is concentrated in the higher range with relatively small overall dispersion, indicating a stable and clear positive charge tendency at the sequence level. In contrast, the isoelectric points of random sequences, UniProt, and negative samples are generally lower and more dispersed, reflecting their lack of the positive charge enrichment characteristics required for antimicrobial peptides. Notably, the generative model-based method shows good consistency with the isoelectric point distribution of real antimicrobial peptides, especially the latent diffusion model, whose median and distribution range of isoelectric points highly overlap with those of positive samples. This indicates that the model can effectively capture the statistical regularity of basic residues along the sequence direction in antimicrobial peptide sequences, thus maintaining good biophysical rationality at the overall charge level. This result further validates the effectiveness of the proposed method in reconstructing the key charge-related characteristics of antimicrobial peptides.

[0053] 6) Net charge analysis Net charge is an important indicator for measuring the overall electrical properties of antimicrobial peptides under physiological conditions. Its size and distribution directly affect the strength of the electrostatic interaction between the antimicrobial peptide and the negatively charged bacterial membrane, as shown in the following figures. Figure 7 As shown in the figure, the net charge distribution of the real antimicrobial peptide (Positive) is significantly biased towards the positive range and relatively concentrated, reflecting a stable positive charge enrichment pattern at the sequence level. In contrast, the net charge of random sequences, UniProt, and negative samples is generally lower, with some even approaching neutral or negative values, indicating a lack of charge-driven characteristics required for antimicrobial peptides. Notably, the generative model-based method shows high consistency with the real antimicrobial peptide in terms of net charge distribution, especially the latent diffusion model, whose median net charge level and distribution range are very close to those of the positive samples. This indicates that the model can effectively learn the combination rules of positive and negative charge residues in the antimicrobial peptide sequence during generation and maintain good directional characteristics at the overall charge level. This result further illustrates that the model has strong biophysical rationality in modeling charge-related physicochemical properties.

[0054] 7) Sequence consistency and similarity analysis To evaluate the novelty of sequences generated by different methods and their relationship with existing sequences, statistical analyses were performed on the sequence consistency and similarity between the generated sequences and the training set and other generated sequences. For example... Figure 7As shown, traditional methods and some machine learning-based generation strategies exhibit high levels of sequence consistency and similarity with the training set, indicating that their generation results rely to some extent on the training data and pose a significant risk of sequence redundancy. In contrast, the proposed method is significantly lower than the comparative methods in both sequence consistency and similarity with the training set, indicating that the generated antimicrobial peptide sequences maintain a higher degree of difference from the training samples, which is beneficial for expanding the sequence space and improving the ability to explore new sequences. Further analysis of the internal similarity between generated sequences reveals that most comparative methods still maintain a high level of "sequence consistency / similarity with other generated sequences," indicating a certain degree of homogenization among their generated results. In contrast, the proposed method significantly reduces this metric, with greater differences between generated sequences, indicating that the model effectively suppresses pattern collapse while maintaining functionally relevant features, thus improving the diversity and coverage of the generated results. In summary, the proposed method has significant advantages in avoiding overfitting to training data, reducing generated sequence redundancy, and improving novelty, providing a more promising set of candidate sequences for subsequent screening and experimental verification of high-quality antimicrobial peptides.

[0055] In summary, the proposed conditional latent diffusion model can stably generate structurally sound and uniformly distributed antimicrobial peptide sequences within a continuous latent space, based on the overall experimental results. Unsupervised distribution analysis based on protein language model features shows that the generated sequences exhibit a high degree of overlap with natural antimicrobial peptides in the high-dimensional feature space, while maintaining clear distinction from non-antimicrobial peptides. This indicates that the model can effectively capture the latent sequence characteristics and statistical regularities of antimicrobial peptides without explicit classification supervision. Further physicochemical property analysis reveals that the generated sequences maintain consistency with the distribution of known antimicrobial peptides in terms of amino acid composition, net charge, isoelectric point, and secondary structure-related indices, showing a reasonable bias, especially in key antimicrobial properties such as positive charge and hydrophobicity. This phenomenon demonstrates that latent space-based diffusion modeling can preserve sequence diversity while constraining the generated results to fall within a biologically plausible antimicrobial peptide subspace, laying a reliable foundation for subsequent functional screening.

[0056] The multi-level computational screening process built upon the generated results further improves the overall quality and usability of candidate antimicrobial peptides. By sequentially introducing antimicrobial activity prediction, sequence similarity and diversity constraints, and a multi-pathogen MIC prediction model, the screened sequences outperform the unscreened set in predicting antimicrobial performance, while significantly reducing redundancy between candidate sequences and between candidate sequences and the training set. Novelty and diversity evaluation results demonstrate that this method has significant advantages in avoiding model memory of training data and expanding the sequence chemical space. In particular, the multi-pathogen joint prediction mechanism built for broad-spectrum antimicrobial potential eliminates reliance on prediction results from a single species during the screening process, contributing to improved robustness and generalization ability of candidate sequences in real-world applications. In summary, the above experimental results validate the feasibility of the proposed antimicrobial peptide generation and screening framework at both the methodological and engineering levels, and provide a set of candidate sequences with guiding significance for subsequent in vitro experimental validation and practical antimicrobial peptide design.

[0057] The above-described embodiments are merely preferred embodiments provided to fully illustrate the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are all within the scope of protection of the present invention. The scope of protection of the present invention is defined by the claims.

Claims

1. A method for screening antimicrobial peptides based on a conditional latent diffusion model, characterized in that, include: (1) Data acquisition: Obtain experimentally validated antimicrobial peptide sequences from publicly available antimicrobial peptide databases, and construct non-antimicrobial peptide sequence data with consistent length distribution to form training and validation datasets; (2) Data preprocessing: The peptide sequences are standardized and encoded, including amino acid letter mapping, length alignment and invalid sequence filtering; (3) Latent space coding: using variational autoencoder (VAE) through encoder network Discrete peptide sequences are mapped to a continuous latent space to obtain the probability distribution of latent variables; this is then processed through a decoder network. Reconstructing the original sequence from latent variables; model training is achieved by maximizing the lower bound of evidence, ELBO. (4) Conditional diffusion modeling: A conditional diffusion model is constructed in the latent space. Noise is gradually added to the latent variable of antimicrobial peptide through a forward diffusion process, and a reverse denoising process is learned under the constraint of antimicrobial property condition c. To model the distribution of potential representations of antimicrobial peptides; optionally, a fine-tuning process based on minimum inhibitory concentration (MIC) data is introduced to adjust model parameters using MIC information; (5) Antimicrobial peptide generation: Starting from random noise, backsampling is performed using a trained conditional diffusion model to obtain a latent representation, and then candidate antimicrobial peptide sequences are generated through a VAE decoder. (6) Screening of candidate antimicrobial peptides: First, multiple antimicrobial activity prediction models based on convolutional neural networks, recurrent neural networks and self-attention mechanisms are used to independently predict and jointly screen the generated candidate peptide sequences; then, similarity calculation and cluster analysis are performed on the sequences that pass the activity screening to remove highly similar or repetitive sequences and obtain a set of candidate sequences with redundancy removed; (7) Minimum inhibitory concentration prediction screening: A minimum inhibitory concentration prediction model based on attention mechanism is constructed to predict the MIC of the candidate sequences after redundancy removal and to screen out sequences with predicted MICs lower than a set threshold; (8) Screening of broad-spectrum antimicrobial peptides: Using the minimum inhibitory concentration prediction sub-models for Gram-positive and Gram-negative bacteria respectively, candidate sequences are evaluated and sequences with low prediction MICs against multiple pathogens are retained to obtain a candidate set of broad-spectrum antimicrobial peptides.

2. The method for screening and generating antimicrobial peptides according to claim 1, characterized in that, In step (4), the forward process of the conditional diffusion model is defined as: in , The preset noise control coefficient, These are initial latent variables used to characterize the potential features of antimicrobial peptides. The variables after adding noise at step t are used to train the model to learn the denoising process of generating potential representations of antimicrobial peptides from noisy states.

3. The method for screening and generating antimicrobial peptides according to claim 1 or 2, characterized in that, In step (4), the objective function for the fine-tuning process of the MIC data is: , in The basic training loss for the diffusion model, is a weighting function constructed from the minimum inhibitory concentration, used to enhance the contribution of highly active samples, and λ is a balance coefficient used to adjust the relative weight between the MIC weighted loss term and the basic training loss of the diffusion model.

4. The method for screening and generating antimicrobial peptides according to claim 1, characterized in that, In step (6), the antibacterial activity screening rule is as follows: candidate sequence Retained if and only if in Indicates the first Candidate antimicrobial peptide sequences, Indicates the first The antibacterial activity threshold corresponding to each prediction model This represents the minimum number of models required to satisfy the criteria after filtering. For indicator functions, Let represent the predicted antimicrobial activity value of the i-th antimicrobial activity prediction model for the i-th candidate antimicrobial peptide sequence. This represents the set of antimicrobial peptide candidate sequences obtained after screening by multiple prediction models with consistency constraints.

5. The method according to claim 1, characterized in that, In step (6), the sequence diversity screening includes: calculating the pairwise similarity between candidate sequences. Clustering is performed based on similarity, and only one or more representative sequences are retained for each cluster. This represents the sequence similarity metric function, using cosine similarity.

6. The method according to claim 1, characterized in that, In step (7), the minimum inhibitory concentration prediction model adopts a low-rank self-attention mechanism to adaptively focus on amino acid fragments in the sequence that contribute significantly to antibacterial activity.

7. The method according to claim 1, characterized in that, In step (8), the screening of broad-spectrum antimicrobial peptides specifically involves: inputting candidate sequences into the MIC prediction sub-models for Gram-positive and Gram-negative bacteria, respectively, to obtain predicted values. in Indicates candidate antimicrobial peptide sequence Predicted minimum inhibitory concentration for Gram-positive bacteria Indicates candidate antimicrobial peptide sequence Predicted minimum inhibitory concentration for Gram-negative bacteria.

Citation Information

Cited By

  • Multi-objective antimicrobial peptide generation method based on constrained monte carlo tree search and diffusion model

    CN122551897A