Method for rapidly screening high-efficiency alpha helix antibacterial peptide and application thereof

By extracting characteristic parameters of α-helical antimicrobial peptides and employing a hierarchical screening strategy, the problem of scarce data on plant-derived antimicrobial peptides has been solved. This enables rapid and efficient screening of high-efficiency α-helical antimicrobial peptides on ordinary computers, reducing screening time and costs, and making it suitable for various databases and library design scenarios.

CN122369672APending Publication Date: 2026-07-10NINGBO UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NINGBO UNIV
Filing Date
2026-06-09
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently screen for highly effective α-helical antimicrobial peptides when data on plant-derived antimicrobial peptides are scarce. Furthermore, traditional methods are computationally intensive, time-consuming, and inefficient, failing to meet the demands for rapid screening and reduced research costs.

Method used

By extracting characteristic parameters such as the hydrophobic cluster scores HCS3_HF and HCS4_HF of α-helical antimicrobial peptides, and combining them with hydrophobicity Hyd, hydrophobic moment HMom, net charge z, hydrophobic surface HoF, and hydrophilic surface HiF, a programmatic AMPs key feature calculation system written in Python was established for screening. A hierarchical screening strategy was adopted, including physicochemical constraints, confirmation of amphiphilicity, and screening based on hydrophobic cluster score thresholds. The screening process was optimized by using multi-centroid distance sorting.

Benefits of technology

It achieves efficient screening of billions of sequence spaces on ordinary computers, shortening the screening time to 1-2 hours. The MIC of all screened antimicrobial peptides is controlled within 32μg/mL, which significantly improves screening efficiency and reduces costs. It is suitable for various database and library design scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369672A_ABST
    Figure CN122369672A_ABST
Patent Text Reader

Abstract

This invention discloses a method for rapidly screening highly efficient α-helical antimicrobial peptides and its application. The method includes extracting hydrophobic cluster scores HCS3_HF and / or HCS4_HF from the peptide sequences to be screened, and screening the peptide sequences based on the values ​​of these features. Using known low-MIC α-helical antimicrobial peptides as multi-centroid references, precise selection of candidate peptides is achieved through feature value ranges. This invention constructs a screening system based on the structure-activity relationship of the synergistic effect of hydrophobic clusters at positions i+3 and i+4 of the hydrophobic surface. This system can efficiently identify α-helical antimicrobial peptides with low minimum inhibitory activity from a vast sequence space without relying on large-scale training data, significantly shortening the antimicrobial peptide development cycle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of peptide antibiotic technology, specifically relating to a method for rapidly screening highly efficient α-helical antimicrobial peptides and their applications. Background Technology

[0002] Pectinobacter is the main pathogen causing bacterial soft rot in solanaceous and cruciferous vegetables, resulting in severe losses to their production. Although strict field management and the breeding of disease-resistant varieties can mitigate the losses caused by soft rot, effective control measures are lacking. Antimicrobial peptides, as the first line of defense in the host's immune system, possess broad-spectrum antibacterial properties and are less prone to inducing drug resistance, making them ideal candidates for antibiotic alternatives.

[0003] Antimicrobial peptides (AMPs) can be classified into α-helical, β-sheet, α / β mixed, and cyclic peptides based on their structural characteristics, with α-helical peptides being the most common. Because parameters such as net charge, hydrophobicity, and amphiphilicity of α-helical AMPs can be quantified and modeled, and their structure-activity relationships are relatively clear, α-helical AMPs have become a core object in current machine learning prediction systems. Currently, α-helical AMP learning models are mostly trained based on antimicrobial peptide databases, achieving high-throughput screening and targeted optimization of candidate peptides through the mapping relationship between antimicrobial peptide sequence composition and antimicrobial activity. However, this data-driven strategy is highly dependent on the scale and quality of the training data, but currently available training data for plant-derived AMPs is insufficient. For example, the APD database, as of January 2026, contains 6309 AMPs, of which 3379 are natural AMPs and only 272 are plant-derived antimicrobial peptides. Due to the relative insufficiency of experimental validation data and standardized functional annotations for plant-derived AMPs, problems such as class imbalance and feature bias in the training set have emerged, limiting the generalization ability of the model in predicting plant-derived antimicrobial peptides. How to achieve the discovery and functional optimization of plant-derived antimicrobial peptides under limited data conditions has become a critical scientific problem that urgently needs to be solved. Due to data scarcity, enumeration methods for screening suffer from the problem of large data volumes and low screening efficiency; for example, for an antimicrobial peptide with 18 amino acids, there are up to 20 possible permutations and combinations. 18 Power bar sequence.

[0004] Although there are reports on predicting antimicrobial peptides using bioinformatics, these methods typically require massive computing power and tens of hours of computation to complete the initial screening. Furthermore, the antimicrobial peptides obtained in the final screening tend to have low activity.

[0005] Therefore, given the scarcity of plant antimicrobial peptide data, there is a need for a technology that can screen out efficient antimicrobial peptide sequences from a large database, while saving computing power, improving screening efficiency, and shortening the screening time and research costs for new antimicrobial peptides. Summary of the Invention

[0006] To address the aforementioned needs, this invention discloses a technique based on the quantification of α-helical antimicrobial peptide characteristics, enabling rapid screening of highly efficient α-helical antimicrobial peptides. The method provided by this invention can complete the screening process in just 1-2 hours using a typical home computer, effectively controlling the minimum inhibitory concentration (MIC) of all screened α-helical antimicrobial peptides to within 32 μg / mL. Furthermore, this technique has been demonstrated to be suitable for screening data from site-directed mutagenesis, plant protein libraries, and custom enumeration libraries, thus holding promise as a novel method for screening control products against *Pectinobacterium soft rot*.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: On one hand, the present invention provides a method for screening α-helical antimicrobial peptides, comprising the following steps: (i) Extract at least one α-helix structure-related feature of the peptide sequence to be screened, the feature including hydrophobic cluster score HCS3_HF (hydrophobic cluster score on the i+3 hydrophobic surface) and / or HCS4_HF (hydrophobic cluster score on the i+4 hydrophobic surface). (ii) Screen the peptide sequences according to the value of the aforementioned characteristics to obtain α-helical antimicrobial peptides with low MIC (minimum inhibitory concentration).

[0008] In our previous research, we discovered that the antimicrobial activity of antimicrobial peptides mainly depends on the electrostatic interaction between positively charged amino acids and negatively charged cell membranes, as well as the penetration and disruption of cell membranes by hydrophobic amino acids. We designed a method using features such as helix number, Predicted Local Distance Difference Test (pLDDT), net charge, and hydrophobicity percentage to successfully screen for highly active Pth-St1 antimicrobial derivative peptides (see domestic invention patent CN2025114079962). However, this method is based on the large language model ESM-3, requiring a massive GPU configuration and computational power, with a screening time exceeding 16 hours. Furthermore, this feature-generated screening method cannot achieve large-scale enumeration. Additionally, among the thousands of candidate sequences obtained, many antimicrobial peptides with high MICs still exist, requiring extensive validation experiments to identify truly effective antimicrobial peptides.

[0009] To effectively shorten the computational screening time, save computing power, and improve screening efficiency, this invention systematically studies the relationship between α-helical antimicrobial peptide characteristics and their minimum MICs (micro-internal ratios). A programmatic AMPs key feature calculation system written in Python was established, including average hydrophobicity (Hyd), hydrophobic moment (HMom), the ratio of polar to nonpolar amino acids, hydrophobic surface (HoF), and hydrophilic surface (HiF). A total of 54 α-helical antimicrobial peptide characteristics were obtained using these characteristics (as shown in Table 3 of the examples), including 28 basic characteristics and 26 amphiphilic characteristics. It was found that the MIC values ​​of α-helical antimicrobial peptide characteristics are directly related to the hydrophobic amino acid structure characteristics of the α-helical antimicrobial peptide, especially HCS3_HF and HCS4_HF.

[0010] Specifically, the method provided by this invention includes extracting at least one characteristic parameter related to the α-helix structure and hydrophobic amphiphilicity of the peptide sequence to be screened, and screening the peptide sequence according to the value of the characteristic parameter to obtain candidate antimicrobial peptides. The characteristic parameter includes at least the hydrophobic cluster scores HCS3_HF and / or HCS4_HF defined based on the spatial arrangement of hydrophobic residues.

[0011] HCS3_HF and HCS4_HF are two key features used to quantify the degree of spatial aggregation between amino acid residues on the hydrophobic surface of α-helical antimicrobial peptides. HCS3_HF characterizes the strength of clustering forces between hydrophobic amino acid side chains spaced three residues apart (i and i+3) along the peptide chain direction on the assumed α-helical conformation; HCS4_HF characterizes the strength of clustering forces between hydrophobic amino acid side chains spaced four residues apart (i and i+4).

[0012] In the α-helix structure, each amino acid residue rotates approximately 100° around the helical axis. Therefore, i+3: starting from the i-th residue, count forward 3 residues along the helix (i.e., the i+3 position). Since 3 × 100° = 300°, equivalent to a 300° rotation from position i (i.e., 60° in the opposite direction), these two residues are not exactly on the same side of the helical cylinder, but slightly offset. i+4: counting forward 4 residues from the i-th residue (i+4 position). 4 × 100° = 400°, equivalent to 360° + 40°, that is, a full rotation plus 40°, so i and i+4 are almost on the same side of the helix, very close in space. That is to say, i and i+3 are on the same helix, and i+4 is the position corresponding to i in the next helix immediately following it. i+3 and i+4 are both spatially adjacent positions of i.

[0013] In some embodiments, the HCS3_HF (hydrophobic cluster score on the i+3 hydrophobic surface) mentioned in this invention refers to the average hydrophobic distance between two adjacent i and i+3 positions in the hydrophobic surface of an α-helix structure; HCS4_HF (hydrophobic cluster score on the i+4 hydrophobic surface) refers to the average hydrophobic distance between two adjacent i and i+4 positions in the hydrophobic surface of an α-helix structure. For example, the hydrophobic amino acid is composed of amino acids at positions 1, 4, 8, 11, and 15 in each helix (usually, the hydrophobic surface in an α-helix structure consists of amino acids at these five positions), where 1 and 4, 8 and 11 are the i+3 cases, and 4 and 8, 11 and 15 are the i+4 cases. Therefore, HCS4_HF represents the average hydrophobic distance between the combinations 4 and 8, and 11 and 15, and HCS3_HF represents the average hydrophobic distance between the combinations 1 and 4, and 8 and 11. The calculation of HCS3_HF and HCS4_HF has been integrated into the AMPLLM feature calculation suite developed in this invention (such as the amp_amphiphilic_features.py script), which can perform batch calculations on hundreds or thousands of peptide sequences with one click and output the specific values ​​of HCS3_HF and HCS4_HF for subsequent threshold screening or multi-centroid distance sorting.

[0014] Through extensive mutant analysis and experimental verification, this invention has found that HCS4_HF is significantly negatively correlated with the minimum inhibitory concentration (MIC) of antimicrobial peptides, and there is a dynamic synergistic balance between HCS3_HF and HCS4_HF. For example, when HCS4_HF is greater than zero and the specific directional hydrophobic moment component meets the preset conditions, the antimicrobial peptide exhibits an extremely low MIC value. This provides a clear physicochemical criterion for rational design in the absence of training samples of plant antimicrobial peptides.

[0015] In some embodiments, the features include HCS3_HF and HCS4_HF.

[0016] Studies have shown that HCS4_HF plays a leading role in regulating the antibacterial activity of α-helical antimicrobial peptides, while HCS3 plays a synergistic role. A dynamic synergistic balance exists between HCS3_HF and HCS4_HF, jointly regulating the antibacterial activity of α-helical antimicrobial peptides.

[0017] Furthermore, the features described in step (a) also include one or more of the following: Hyd (average hydrophobicity), HMom (hydrophobic moment), z (net charge), HoF (hydrophobic surface), and HiF (hydrophilic surface).

[0018] Since approximately 3.6 amino acids form a helix, and after 5 helices, the amino acids return from 0 degrees to 0 degrees, 18-19 amino acids can rotate once to form a helical wheel, which is sufficient to cross the hydrophobic core of the bacterial cell membrane and form a stable amphiphilic structure, enough for the α-helix to form at least 4-5 turns. Helical wheel analysis can clearly distinguish between the hydrophobic and hydrophilic surfaces.

[0019] This invention selects peptide combinations of 18-19 amino acids (including 5 helices) as the antimicrobial peptide database to be screened. For 20 amino acids, any combination of 18-19 amino acids can be selected, resulting in billions of possible combinations. If only HCS3_HF and HCS4_HF are used for screening, the computational power would be insufficient to complete the screening smoothly. Therefore, preliminary screening through a series of physicochemical and structural constraints is required to reduce the number of peptide sequences to be screened. This improves screening efficiency, saves computational power, and reduces costs. The screening of α-helical antimicrobial peptides can be completed using only a regular computer.

[0020] The ordinary computer mentioned in this invention refers to a computer with a Windows 10 / 11 operating system, a modern 6-core or higher processor, and at least 16 GB of memory, which can ensure smooth operation of the screening process.

[0021] To complete the preliminary screening, this invention selected average hydrophobicity (Hyd), hydrophobic moment (HMom), net charge (z), hydrophobic surface (HoF), and hydrophilic surface (HiF) as preliminary screening characteristics for α-helical antimicrobial peptides from the aforementioned 54 α-helical antimicrobial peptide characteristics.

[0022] Alpha-helical antimicrobial peptides must be able to form a stable amphiphilic alpha-helical structure and interact effectively with the bacterial cell membrane to exert their bactericidal function. This invention demonstrates that hydrophobicity (Hyd), hydrophobic moment (HMom), net charge (z), hydrophobic surface (HoF), and hydrophilic surface (HiF) are preliminary screening characteristics for alpha-helical antimicrobial peptides. These characteristics are directly related to whether the alpha-helical antimicrobial peptide can form an amphiphilic alpha-helical structure and whether it can interact effectively with the bacterial cell membrane. They can serve as basic physicochemical and structural constraints, enabling the rapid elimination of over 99% of inactive sequences from a massive pool of candidate sequences without experimental synthesis. This is the core technological foundation for achieving single-machine, billion-level high-efficiency screening in this invention.

[0023] Hyd (mean hydrophobicity) refers to the average value that measures the overall hydrophobicity of a peptide sequence, reflecting its ability to insert into the hydrophobic core of a bacterial cell membrane.

[0024] HMom (hydrophobic moment) refers to the degree of spatial aggregation of hydrophobic amino acid side chains on one side of the helix under the assumed α-helix conformation. The larger the HMom, the more complete the hydrophobic surface and the stronger the amphiphilicity.

[0025] z (net charge) refers to the net charge carried by a peptide sequence under physiological pH conditions (approximately 7.0), which determines its electrostatic attraction to negatively charged bacterial cell membranes (such as lipopolysaccharides).

[0026] HoF (hydrophobic surface) refers to one side surface formed by the concentrated distribution of continuous hydrophobic amino acid residues (such as L, W, F, I, V, etc.) in the α-helix projection, which is responsible for inserting the hydrophobic core of the membrane.

[0027] The HiF (hydrophilic surface) refers to the region on the opposite side of the hydrophobic surface in the projection of the α-helix wheel. It is usually rich in positively charged hydrophilic residues (such as K and R) and is responsible for electrostatic interactions with the polar heads and lipopolysaccharides on the membrane surface.

[0028] Furthermore, the screening described in step (ii) includes the following three steps: (1) Hydrophobic moment HMom≥0.5, hydrophobicity Hyd≤0.5, net charge 7≤z≤11; (2) The hydrophobic surface HoF ≠ none, and the hydrophilic surface HiF ≠ none; (3) The hydrophobic cluster score HCS3_HF on the i+3 hydrophobic surface is greater than 0, and the hydrophobic cluster score HCS4_HF on the i+4 hydrophobic surface is greater than 0.

[0029] First, in the initial screening in step (1), the thresholds HMom ≥ 0.5, Hyd ≤ 0.5 and 7 ≤ z ≤ 11 are set based on quantitative experience of the relationship between the structure and activity of α-helix antimicrobial peptides. The purpose is to quickly eliminate a large number of sequences that are unlikely to have high antimicrobial activity and retain candidate peptides with an amphiphilic cationic helical basic backbone.

[0030] Cationic antimicrobial peptides anchor to bacterial surfaces via electrostatic attraction, which is the first step in their action. This invention, through site-directed mutagenesis experiments, discovered that the net charge of highly effective antimicrobial peptides (MIC ≤ 7.8 μg / mL) almost entirely falls within the +7 to +11 range; therefore, this is set as a rigid constraint. If the net charge is too low (< +7), the binding force with LPS (lipopolysaccharide) is insufficient, making it difficult to effectively aggregate on the bacterial surface; if the net charge is too high (> +11), it may non-specifically bind to host cell membranes (such as negatively charged glycoproteins on the surface of erythrocytes), increasing the risk of hemolytic toxicity.

[0031] Hydrophobicity (on the Fauchère-Pliska scale) determines a peptide's ability to insert into the lipid bilayer. This invention, through analysis of a large number of known amphiphilic cationic AMPs, shows that the Hyd of highly efficient peptides is almost always ≤ 0.5. A Hyd that is too high (> 0.5) can cause peptides to self-aggregate in aqueous solutions, even forming irreversible precipitation, and can easily damage eukaryotic cell membranes, leading to hemolysis; a Hyd that is too low (< 0 or even negative) cannot penetrate the hydrophobic core, resulting in the loss of membrane lysis activity. Therefore, setting Hyd ≤ 0.5 as the initial screening upper limit ensures sufficient membrane insertion capacity while avoiding the toxicity caused by excessive hydrophobicity.

[0032] Hydrophobic moment (HMom) measures the degree of aggregation of hydrophobic amino acids on one side of an α-helix. A larger HMom indicates a more concentrated spatial orientation of the hydrophobic side chains, resulting in a more complete hydrophobic surface. Our study using the Q1_19 mutant revealed that all peptides with low MIC (≤ 15.6 μg / mL) have an HMom ≥ 0.5. Therefore, HMom ≥ 0.5 is used as a mandatory initial screening threshold. If HMom < 0.5, the distribution of hydrophobic residues tends to be uniform, making it difficult to form a clear amphiphilic helix. Even with suitable charge and hydrophobicity, these peptides are unlikely to form pores or disrupt membrane integrity.

[0033] The three thresholds—HMom ≥ 0.5, Hyd ≤ 0.5, and 7 ≤ z ≤ 11—form a first-level screening funnel. Within minutes of computation, sequences that either cannot bind to bacteria (insufficient charge), are too toxic or have poor solubility (too hydrophobic), or are not amphiphilic (too low hydrophobic moment) can be eliminated from billions of random sequences. Experiments show that among sequences meeting these three conditions, subsequent screening with HCS3 / 4 yields a significantly higher proportion of active candidate peptides than unconstrained enumeration, thus achieving highly efficient single-machine screening at the billion-level.

[0034] Secondly, merely satisfying the three quantitative values ​​of net charge, hydrophobicity, and hydrophobic moment can indicate that the screened antimicrobial peptide has amphiphilic potential, but it cannot guarantee that the sequence will actually form an effective amphiphilic α-helix structure in three-dimensional space. It is also necessary to use the condition that the hydrophobic surface (HoF) ≠ absent and the hydrophilic surface (HiF) ≠ absent as the screening condition for step (2), to confirm that there is indeed a continuous region composed of multiple adjacent hydrophobic residues on one side of the helix, and that positively charged amino acids (such as lysine K and arginine R) are concentrated on the other side of the helix to form a hydrophilic surface. The hydrophilic surface should preferably contain positively charged amino acids K or R, which can increase the adsorption capacity of the peptide. The hydrophobic surface can anchor to the membrane surface and play a role in disrupting the membrane integrity, thereby performing antimicrobial activity.

[0035] Furthermore, step (1) (Hyd, HMom, z) is extremely fast (requiring only sequence composition information) and can instantly eliminate over 99% of invalid sequences from billions of sequences. Step (2) (HoF / HiF verification) requires spiral wheel projection or calling structure prediction tools, which has a slightly higher computational cost and is not suitable for direct operations on the original 20^18 space. Therefore, it is placed after the initial screening and only verifies tens of thousands to hundreds of thousands of candidate sequences that meet the basic physicochemical conditions.

[0036] Experimental data from this invention show that sequences that meet the threshold in step one but lack HoF or HiF almost entirely exhibit extremely poor antibacterial activity (MIC > 64 μg / mL). Therefore, step (2) can effectively eliminate pseudo-amphiphilic sequences.

[0037] Subsequently, in step (c) screening, HCS3_HF > 0 and HCS4_HF > 0 are required because the first two steps (physicochemical parameters, presence of hydrophobic / hydrophilic surfaces) can only ensure that the antimicrobial peptide has a basic amphiphilic helical backbone, but cannot quantify the more refined spatial aggregation pattern between amino acids on the hydrophobic surface. It is this aggregation pattern that determines whether the helix can efficiently insert into and destroy the bacterial cell membrane and whether it has efficient antimicrobial activity.

[0038] HCS4_HF > 0 is the dominant condition for highly efficient antibacterial activity. The residues spaced at i+4 intervals lie on the same side of the α-helix (nearly longitudinally aligned), and their hydrophobic interactions are crucial for forming a continuous, dense hydrophobic surface. Correlation analysis showed a significant negative correlation between HCS4_HF and MIC (higher HCS4_HF corresponds to lower MIC), and HCS4_HF was greater than 0 in all low MIC groups. Therefore, requiring HCS4_HF > 0 can directly eliminate sequences with a loose, discontinuous hydrophobic surface (e.g., excessively large residue spacing).

[0039] HCS3_HF > 0 provides a synergistic enhancement effect. Although the residues spaced at i+3 are not completely on the same side of the helix, they can interleave and complement the backbone formed by i+4 in three-dimensional space, enhancing the overall lateral stability of the hydrophobic surface. Through mutation experiments, this invention found that when HCS3_HF = 0 (even if HCS4_HF > 0), the antibacterial activity often does not reach the optimal level (MIC increases by 24 times). Therefore, requiring HCS3_HF > 0 is a necessary condition to ensure that the hydrophobic surface possesses a synergistic aggregation network and achieves the lowest MIC.

[0040] Step (3) screening is the final confirmation of the structural quality of α-helical antimicrobial peptides. After passing the screening in steps (1) and (2), two problems may still exist: 1) Although the hydrophobic surface exists, the hydrophobic residues are arranged randomly (e.g., F, W, L are randomly spaced), resulting in a hydrophobic cluster score close to 0; 2) Some sequences have high HMom values, but the high contribution comes from a few hydrophobic residues, and the remaining positions are filled with small hydrophobic residues (e.g., A, V), which cannot form efficient clusters. By directly calculating the density of hydrophobic residues and the interaction strength at positions i+3 and i+4 using HCS3_HF and HCS4_HF, the hydrophobic surface quality of candidate peptides can be improved from the presence or absence to the quantitative level of strength at a very low computational cost (based only on sequence and helical wheel mapping). Experiments show that the proportion of sequences that satisfy HCS3_HF > 0 and HCS4_HF > 0 with a final MIC ≤ 15.6 μg / mL is more than 6 times higher than that of sequences that only satisfy the first two steps.

[0041] Furthermore, the screening in step (ii) also includes step (4): using at least one α-helical antimicrobial peptide with a known low MIC as the centroid, using the characteristic value of the centroid as a reference to determine the screening interval, and performing screening according to the screening interval.

[0042] After completing the first three steps (steps (1) to (3)) of screening (physicochemical property constraints, confirmation of the existence of amphiphilicity, and basic hydrophobic cluster scoring threshold), the candidate antimicrobial peptide database has been compressed to the level of tens of thousands, and the basic screening has been completed. At this time, it is necessary to further screen out antimicrobial peptides with higher activity, so that their MIC is lower, achieve the expected goal, and further reduce the number of candidate antimicrobial peptides to the point where wet experiment verification can be carried out.

[0043] The wet experiment refers to the actual physical operations in the laboratory, such as synthesizing peptides, culturing bacteria, and determining the minimum inhibitory concentration, as opposed to the dry experiment simulated by computer.

[0044] This invention uses at least one experimentally verified α-helical antimicrobial peptide with a known low molecular weight enzyme (MIC) as a reference centroid. It examines various characteristic values ​​of this low-MIC α-helical antimicrobial peptide as screening reference values. For example, if there is one centroid, its characteristic values ​​are used to determine the range of characteristic values ​​to be screened; if there are multiple centroids, the characteristic values ​​of all centroids are used to determine the range of characteristic values ​​to be screened for further screening. This strategy, by constructing a metric similar to the characteristic fingerprint of highly active antimicrobial peptides, effectively overcomes the difficulty of capturing activity fluctuations caused by small variations in peptide sequences using fixed cutoff values.

[0045] Furthermore, the centroid preferably includes at least two of (Q90, Q10); the eigenvalues ​​include the eigenvalues ​​of HCS4_HF, HCS3_HF, and HMom HF_mag.

[0046] Q10 and Q90 are statistical quantiles calculated based on α-helical antimicrobial peptide samples with known low minimum inhibitory concentrations (MICs). Specifically, a key characteristic of these highly active peptides (e.g., HCS4_HF, which is used for sorting because it has the most significant impact on MIC) is ranked from smallest to largest: the value at the 10th percentile is called Q10, representing the top level, with only 10% of the best peptides reaching or exceeding this value; the value at the 90th percentile is called Q90, representing the general level of these highly active peptides. The range formed by these two quantiles is used as a reference, rather than relying solely on a single optimal value, to accurately identify highly active peptides while avoiding overfitting and ensuring sequence diversity of candidate peptides.

[0047] The HMom_HF_mag (magnitude of the hydrophobic moment vector) is calculated for the hydrophobic surface of the α-helical antimicrobial peptide (i.e., the side where hydrophobic amino acid residues are concentrated in the helical projection). It projects the hydrophobic value (e.g., Fauchère-Pliska scale) of each hydrophobic residue into a two-dimensional vector according to its spatial rotation angle within the α-helix. The magnitude of the sum of these vectors is then given as HMom_HF_mag. The previously mentioned HMom (hydrophobic moment) calculates the sum of the vector magnitudes of all hydrophobic amino acid residues in the helical projection for the entire α-helical peptide chain, reflecting the overall amphiphilicity. HMom_HF_mag goes further, focusing only on the residues on the hydrophobic surface (i.e., the side where hydrophobic amino acids are concentrated in the helical projection).

[0048] In step (4), based on the existing HCS4_HF and HCS3_HF, the HMom_HF_mag (magnitude of the hydrophobic moment vector of the hydrophobic surface) feature is added to supplement the quantification of the overall vector intensity of the hydrophobic surface. HCS3_HF and HCS4_HF only evaluate the average hydrophobic distance between several pairs of residues at specific intervals (i+3, i+4) on the hydrophobic surface. HMom_HF_mag, on the other hand, synthesizes the vectors of all hydrophobic residues. The larger the modulus, the more consistent the spatial orientation of the hydrophobic residues, the more concentrated the resultant force, and the higher the physical destruction efficiency of the bacterial cell membrane. Conversely, if HMom_HF_mag is too low, even if the local pairing score is qualified, the overall synergy of the hydrophobic surface will be insufficient.

[0049] In some methods, such as targeting 1867_Q1_19 and its 53 mutants, mutant antimicrobial peptides with a known Exp_Log2_MIC <= 3 (equivalent to MIC ≤ 8 μg / mL) are selected. Peptide sequences ranking in the 10%–90% range are then screened based on HCS4_HF values. The interval values ​​of HCS4_HF, HCS3_HF, and HMom HF_mag are examined, yielding characteristic interval values ​​including: 3.3 ≤ HCS4_HF ≤ 4.5, 2.8 ≤ HCS3_HF ≤ 3.6, and 7.6 ≤ HMom HF_mag ≤ 8.7. These characteristic interval values ​​can be used for efficient screening of antimicrobial peptides from *Pectinobacterium soft rot*.

[0050] Specifically, researchers used the low MIC subgroup (i.e., the batch of highly effective peptides with the lowest minimum inhibitory concentration) from 53 Q1_19 mutants determined by wet assays as a learning sample. For each key structural feature (HCS4_HF, HCS3_HF, HMom_HF_mag), the values ​​of this batch of highly effective peptides were sorted from smallest to largest, forming intervals representing the core characteristic ranges common to the top-performing highly active peptides: 3.3 ≤ HCS4_HF ≤ 4.5, 2.8 ≤ HCS3_HF ≤ 3.6, 7.6 ≤ HMom_HF_mag ≤ 8.7. Candidate peptides screened according to this interval were validated by wet assays, and the MICs of the screened antimicrobial peptides were no higher than 32 μg / mL, proving the scientific validity and effectiveness of this numerical range.

[0051] It is understandable that the screening ranges for HCS4_HF, HCS3_HF, and HMom HF_mag may differ for different antimicrobial peptides. Therefore, it is necessary to select the centroid (Q90, Q10) based on the antimicrobial peptide to redetermine the screening ranges for HCS4_HF, HCS3_HF, and HMom HF_mag, so as to screen for antimicrobial peptides with low MIC values ​​more efficiently.

[0052] The method provided by this invention possesses strong generalization ability and scalability. When applied to mining natural protein databases, a sliding window technique can be used to cut complete protein sequences into continuous peptides of a preset length (e.g., 18 amino acid residues). The resulting peptide library is then subjected to a first layer of physicochemical constraints based on average hydrophobicity, hydrophobic moment, net charge range, and hydrophobic clustering score. A second layer of precise screening is then performed using multi-centroid distance sorting, allowing for the rapid identification of over a hundred high-potential candidate antimicrobial peptides from plant proteomes containing hundreds of millions of peptides. Similarly, when applied to the design of combinatorial mutant libraries based on a fixed backbone, by limiting the types of amino acids at variable sites (e.g., preferably arginine, lysine, tryptophan, and leucine positively correlated with low MIC) and enumerating all combinations, combined with the same hierarchical screening framework, it is possible to efficiently converge to the optimal sequence region with low MIC potential in a sequence space of billions of sequences. All of the above screening processes can be automated by computer programs, with processing time completed within tens of hours, fully demonstrating the overwhelming efficiency advantage of this invention's method compared to traditional wet experimental enumeration methods.

[0053] In some approaches, the number of peptides in the antimicrobial peptide database to be screened is preferably controlled to no more than 1×10^8. If the preset limit of 1×10^8 per task is exceeded, the number of combinations needs to be split to keep it below this threshold, thereby reducing the computational burden on the CPU and memory.

[0054] For example, based on a fixed backbone (M****************GG), a combinatorial mutant library was constructed, with * indicating variable residues extracted from K, A, W, and L. The total number of combinations is 4^16 ≈ 4.3 × 10^9, exceeding the preset limit of 1 × 10^8 per task. The system automatically calculates the number of split combinations to keep each enumeration batch below this threshold: 4^15 ≈ 1.07 × 10^9 (> 1 × 10^8), 4^14 = 2.68 × 10^8 (> 1 × 10^8), 4^13 = 6.71 × 10^7 (< 1 × 10^8). Therefore, the original 16 variable positions are divided: the first three positions are fixed (e.g., KKK, KKA…), resulting in 4^3 = 64 sub-models. Each sub-model has 13 remaining variable positions, which can be exhaustively enumerated within the 1 × 10^8 limit. Subsequently, a gradient filtering process is performed on these sub-models. (This step is based on the computer's CPU computing power and memory pressure. It is characterized by enumeration and generation, and the number of combinations calculated at one time is controlled to 100,000,000 peptides. The filtering is carried out in order of increasing computing power requirements.)

[0055] In actual computation, the system employs a step-by-step enumeration, filtering, and screening approach for each sub-model, rather than generating all sequences first and then screening. Following the order of computational cost from lowest to highest, physicochemical constraints such as net charge, hydrophobicity, and hydrophobic moment are applied sequentially, the existence of hydrophobic / hydrophilic surfaces is verified, and finally, the thresholds for HCS3_HF and HCS4_HF are calculated. This method discards a large number of unqualified sequences early in the enumeration process, resulting in a significantly lower actual number of candidate sequences stored in memory compared to the theoretical total. After the 64 sub-models have been computed independently, all screening results are merged and then uniformly sorted by multi-centroid distance, ultimately completing the efficient screening of a mutation library of 4.3 billion sequences within 2 hours. This strategy fully demonstrates the core technological advantage of this invention: enabling rapid screening of a billion-sequence space on a standard single machine.

[0056] On the other hand, the present invention provides an α-helical antimicrobial peptide, obtained by screening using the method described above.

[0057] Furthermore, the α-helical antimicrobial peptide has an amino acid sequence as shown in any one or more of SEQ ID NO.1 to SEQ ID NO.26.

[0058] The most preferred option is SEQ ID NO.3, which has a minimum MIC of 2.8125 μg / mL.

[0059] Experimental data show that the α-helical antimicrobial peptides obtained by the screening method of the present invention have extremely low inhibitory concentrations against Pectinobacter softrot. The minimum inhibitory concentration of all sequences is no higher than 32 μg / mL, and the lowest can reach 2.8 μg / mL, which verifies the high accuracy and practicality of the screening system of the present invention.

[0060] In summary, this invention overcomes the reliance on large-scale, high-quality plant antimicrobial peptide training datasets. For the first time, it uses the i+3 and i+4 hydrophobic clustering interactions on the hydrophobic surface of α-helical antimicrobial peptides as the core screening indicator. Combined with multidimensional amphiphilic feature quantification and a multi-centroid distance approximation strategy, it establishes a screening method with clear physical meaning and extremely high computational efficiency. The antimicrobial peptides screened using this method not only exhibit excellent in vitro antimicrobial activity, but the screening process also covers two major scenarios: natural protein database mining and artificially combined mutant library enumeration, fully demonstrating its enormous potential as a general-purpose α-helical antimicrobial peptide discovery engine.

[0061] In another aspect, the present invention provides the application of the α-helical antimicrobial peptide described above in the preparation of antibacterial drugs or agricultural fungicides.

[0062] Furthermore, the bacteria are *Pectinobacterium soft rot*, or Gram-negative or Gram-positive bacteria.

[0063] In another aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described above.

[0064] The present invention has the following beneficial effects: Without relying on training data, it starts directly from physicochemical and structural features, achieving efficient and rapid screening in the context of data scarcity.

[0065] For the first time, this study quantitatively revealed the synergistic regulatory relationship between the hydrophobic cluster scores (HCS3_HF and HCS4_HF) at the i+3 and i+4 positions on the hydrophobic surface of α-helical antimicrobial peptides and the minimum inhibitory concentration (MIC). Furthermore, an activity prediction model with HCS4 as the dominant factor and HCS3 as the synergist was established, providing clear physicochemical criteria for rational design.

[0066] By adopting a hierarchical screening strategy (physicochemical constraints → amphiphilicity existence → hydrophobic cluster scoring → multi-centroid distance sorting), high-throughput screening of a billion-level sequence space can be achieved on a regular single machine, with the entire process taking only 1 to 2 hours, which is several orders of magnitude more efficient than the traditional wet experimental enumeration method.

[0067] Wide range of applications: This method is not only applicable to sliding window mining based on plant protein databases, but also to de novo design of fixed backbone combined mutant libraries, and can also be used for site-directed mutant activity optimization, demonstrating strong versatility.

[0068] By introducing a low MIC centroid and determining the screening range using three features—HCS4_HF, HCS3_HF, and HMom_HF_mag—the screening efficiency was significantly improved, and the minimum inhibitory concentration of all the antimicrobial peptide sequences of *Pectinobacter softrot* obtained through screening was no higher than 32 μg / mL.

[0069] The entire screening process can be completed automatically on a computer. Only a few dozen peptides selected at the end need to be verified by wet experiments, which greatly reduces the workload of synthesis and biological testing, shortens the research and development cycle of novel antimicrobial peptides and reduces research costs.

[0070] It is easy to extend to the fields of bioinformatics, agricultural fungicides and antibacterial drug research and development, and has good industrialization prospects. Attached Figure Description

[0071] Figure 1 The overall flowchart of the programmatic computation model of amphipathic characteristics based on LLM established in this invention is shown. Figure 2 This describes the process of programmatically calculating and validating an amphiphilic trait model based on LLM. Figure 3A dedicated feature analysis suite was developed for the programmatic computation of amphipathic feature models based on LLM. Figure 4 The study showed the impact of the Q1_19 hydrophobic surface mutation on P. MIC, and Spearman analyzed the association between AMPs characteristics and MIC; Figure 5 The study showed the impact of the Q1_19 hydrophobic surface mutation on P. MIC, and Pearson analyzed the association between AMPs characteristics and MIC; Figure 6 The study shows the effect of the Q1_19 hydrophobic surface mutation on P. MIC, including the helical wheel structure, hydrophobic surface, and Log2MIC value of the Q1_19 mutant antimicrobial peptide; Figure 7 The effect of the synergistic effect of hydrophobic surfaces HCS4 and HCS3 on MIC is shown, along with the Q1_19 helical wheel structure and schematic diagrams of i+3 and i+4. Figure 8 The effect of the synergistic effect of hydrophobic surfaces HCS4 and HCS3 on MIC is shown. 3D structure diagram of Q1_19 predicted by ESMFold and schematic diagram of i+3 and i+4 are also shown. Figure 9 The study showed the synergistic effect of hydrophobic HCS4 and HCS3 on MIC, and Pearson analysis was used to analyze the correlation between AMPs characteristics and MIC. Figure 10 The effect of the synergistic effect of hydrophobic HCS4 and HCS3 on MIC is shown, and the comparative analysis of Q1_19_3 and Q1_19_2 is presented. Figure 11 The study shows the effect of the synergistic effect of hydrophobic surfaces HCS4 and HCS3 on MIC, and the correlation between the hydrophobic force formed by the number of hydrophobic surfaces W and MIC. Figure 12 The correlation between HCS4_HF and Log2MIC at different thresholds is shown; Figure 13 The correlation between HCS3_HF and Log2MIC at different thresholds is shown; Figure 14 The diagram shows the helical wheel representation of HCS4, which is beneficial for obtaining low MIC of antimicrobial peptides. Figure 15 The diagram shows the projection vector of the hydrophobic distance direction of the hydrophobic surface, which is beneficial for HCS4 to obtain antimicrobial peptides with low MIC. Figure 16 HCS4 was shown to be beneficial for antimicrobial peptides to achieve low MICs. Spearman analysis showed the association between AMP characteristics and MICs. Figure 17The results showed that HCS4 is beneficial for obtaining low MIC of antimicrobial peptides, and the correlation analysis between different threshold Log2MIC and HCS3 and HCS4 was performed. Figure 18 The results showed that HCS4 is beneficial for obtaining low MIC of antimicrobial peptides, and the correlation analysis between different threshold Log2MIC and HCS4 was performed. Figure 19 The results showed that HCS4 is beneficial for obtaining low MIC of antimicrobial peptides, and the correlation analysis between different threshold Log2MIC and HCS3 was performed. Figure 20 The results showed that HCS4 is beneficial for obtaining low MICs of antimicrobial peptides, and the correlation between HCS3_HF and different threshold MICs was analyzed in the full library analysis of α-helical antimicrobial peptides. Figure 21 The results showed that HCS4 is beneficial for obtaining low MICs of antimicrobial peptides, and the correlation between HCS4_HF and different threshold MICs was analyzed in the full library analysis of α-helical antimicrobial peptides. Figure 22 This demonstrates the large-scale low-MIC structural feature prediction model established in this invention. Detailed Implementation

[0072] The present invention will be further described in detail below with reference to specific embodiments. The scope of protection of the present invention is not limited to the following embodiments. All equivalent transformations made based on the technical concept of the present invention are within the scope of protection of the present invention.

[0073] Unless otherwise specified, all raw materials used in this embodiment are commercially available conventional raw materials; and all testing methods used are existing conventional testing methods unless otherwise specified.

[0074] Example 1: Characterization and Activity Correlation Analysis of α-Helical Antimicrobial Peptides (1) Overview of the screening process of the present invention This embodiment employs the self-developed AMP-LLM (Antimicrobial Peptide Large Language Model) computational framework, using the Python programming language to implement the entire process of feature analysis, differential analysis, design, and validation of α-helical antimicrobial peptides. The overall process is as follows: Figure 1 As shown, the process includes AMPs Feature Analysis (using a feature analysis suite to extract multi-dimensional physicochemical and structural features of peptide sequences), Delta-Based Analysis (analyzing feature changes based on the differences between mutant and parent peptides), AMPs Design (designing novel candidate peptides based on feature-activity association rules), and AMP-MIC (verifying the minimum inhibitory concentration of candidate peptides through wet experiments). Finally, the experimental data is fed back to the feature analysis module to iteratively optimize the screening model.

[0075] like Figure 3 As shown, this embodiment developed a dedicated feature analysis suite. The core scripts include PeptidePipeline_Core_V3.py (the main flow control script, coordinating the operation of each module), feature_schema.py (feature pattern definition and validation), amp_features.py (basic peptide feature calculation module), amp_amphiphilic_features.py (amphiphilic feature calculation module), and Run.py (batch processing and parallel computing entry point). This suite can calculate two main categories of features: basic peptide descriptors and amphiphilic characteristics.

[0076] (2) Mutant construction and differential characterization This embodiment uses the patented α-helical antimicrobial peptide Design_1867 (1867) as a template. Studies of α-helical antimicrobial peptides have shown that 18 amino acids can form a helical loop. To facilitate the characterization of the antimicrobial peptide, Design_1867 was truncated to 16 amino acids, forming the antimicrobial peptide derivative 1867_Q. To facilitate subsequent biosynthesis of the antimicrobial peptide, the start codon methionine (Met, M) was introduced at the N-terminus; and the variable amino acid glycine (Gly, G) was introduced at the C-terminus, forming a series of mutant derivative peptides (Table 1). Antimicrobial activity verification showed that the antimicrobial activity of 1867_Q_19, 1867_Q1, and 1867_Q1_19 was 7.81 µg / mL.

[0077] Table 1. Correlation between antimicrobial peptide characteristics and antibacterial activity of Design_1867 derived peptides

[0078] To facilitate subsequent mutation studies, this embodiment used 1867_Q1_19 (MRKLLKKLHRFKAKLVRGG) as a template to construct 53 site-directed mutants (Table 2) via solid-phase synthesis. The mutation sites were mainly located on the hydrophobic surface, replacing original residues with different amino acids (such as F, W, L, etc.). The effects of each mutant on *Pectinobacter softrot* (a type of bacteria) were determined using the double dilution method. Pectobacterium carotovorum MIC: Dilute the overnight culture to 1×10⁻⁶. 6 CFU / mL was mixed with serially diluted antimicrobial peptides (starting at 2000 μg / mL, then 2-fold dilution), and incubated at 37°C for 12 h. The OD value was then calculated. 600The MIC (minimum inhibitory concentration) was determined by measuring turbidity and identifying the lowest concentration that completely inhibited bacterial growth. Kanamycin (50 mg / mL) was used as a positive control, and sterile water as a negative control. Each sample was tested in triplicate.

[0079] Table 2. Sequences and antimicrobial activities of Q1_19 and its 53 mutant antimicrobial peptides.

[0080]

[0081] Then, using the scripts Differential_Feature_Analysis.py and PeptideAnalyzer.py, differential feature analysis was performed on the 53 mutants and the parent peptide Q1_19. The change in each feature (Delta value) of each mutant relative to the parent peptide was calculated, and the Delta value was correlated with the change in MIC (ΔLog2_MIC). After correlation analysis, 54 features with certain descriptive functions were selected (Table 3), which can describe the changes in mutant antimicrobial peptides to a certain extent. Specific steps included: a) Calculate the 54 eigenvalues ​​for each mutant; b) Calculate the difference between each feature and Q1_19: ΔFeature_i = Feature_i(mutant) - Feature_i(Q1_19); c) Calculate the change in MIC: ΔLog2_MIC = Log2_MIC(mutant) - Log2_MIC(Q1_19); d) Spearman rank correlation and Pearson linear correlation were used to analyze the correlation between ΔFeature and ΔLog2_MIC.

[0082] Table 3. Characteristics of α-helical antimicrobial peptides

[0083]

[0084] The hydrophobic amino acids are composed of amino acids at positions 1, 4, 8, 11, and 15 in each helix. Positions 1 and 4, 8 and 11 represent the i+3 case, while positions 4 and 8, 11 and 15 represent the i+4 case. HCS4_HF represents the average hydrophobic distance for the combinations 4 and 8, and 11 and 15, while HCS3_HF represents the average hydrophobic distance for the combinations 1 and 4, and 8 and 11.

[0085] Hyd refers to the average hydrophobicity of the entire peptide sequence, reflecting its ability to insert into the hydrophobic core of the bacterial cell membrane. Calculation method: The Fauchère-Pliska hydrophobicity scale is used, summing the hydrophobicity values ​​of each amino acid in the sequence and then dividing by the peptide chain length (number of residues). The calculation process can be automated using the amp_features.py script.

[0086] HMom characterizes the spatial aggregation of hydrophobic amino acid side chains on one side of the helix under the assumed α-helix conformation. A larger HMom indicates a more complete hydrophobic surface and stronger amphiphilicity. Calculation method: The hydrophobicity value of each residue is treated as a vector, projected onto a plane perpendicular to the helical axis according to its rotation angle within the α-helix (approximately 3.6 residues / turn), and the magnitude of the sum of all residue vectors is calculated. This calculation is implemented using the Python feature suite (amp_amphiphilic_features.py).

[0087] z refers to the net charge carried by the peptide sequence under physiological pH conditions (approximately 7.0), which determines its electrostatic attraction to negatively charged bacterial cell membranes (such as lipopolysaccharides). Calculation method: Count the positive charges of basic amino acids (arginine R, lysine K, histidine H) minus the negative charges of acidic amino acids (aspartic acid D, glutamic acid E). This method does not involve prediction of complex structures and is calculated directly from the sequence composition.

[0088] HoF refers to a side surface formed by a concentrated distribution of continuous hydrophobic amino acid residues (such as L, W, F, I, V, etc.) in the α-helix wheel projection, responsible for inserting into the membrane's hydrophobic core. Calculation method: Based on the predicted α-helix structure or helix wheel model (e.g., obtaining the helix wheel projection through HeliQuest or ESMFold), the residues are grouped according to spatial angles, identifying continuous fan-shaped regions with a significantly higher proportion of hydrophobic residues than the average. For example, it is required that "HoF ≠ None," meaning that such a clearly identifiable hydrophobic fan must exist.

[0089] HiF refers to the region on the opposite side of the hydrophobic surface in the α-helix projection, typically enriched with positively charged hydrophilic residues (such as K and R), responsible for electrostatic interactions with polar heads and lipopolysaccharides on the membrane surface. Calculation method: Also based on the helix projection or three-dimensional structure, identify the region where hydrophilic residues (especially polar and positively charged residues) are concentrated. For example, it is required that "HiF ≠ None," meaning the hydrophilic surface must be clearly present and should typically be distributed at approximately 180° to the hydrophobic surface.

[0090] HMom_HF_mag (the magnitude of the hydrophobic moment vector of the hydrophobic surface) is calculated by projecting the hydrophobic value (such as the Fauchère-Pliska scale) of each hydrophobic residue onto a two-dimensional vector according to its spatial rotation angle in the α-helix of the hydrophobic surface of the α-helix antimicrobial peptide. Then, the magnitude of the sum vector is obtained by summing the vectors of all hydrophobic residues. The detection method is as follows: (a) Based on the primary sequence of the peptide, the spatial angular position of each residue in the α-helix conformation is determined using the HeliQuest helical wheel projection algorithm or structure prediction tools such as ESMFold; (b) Each amino acid is assigned a hydrophobic value according to a predefined hydrophobic scale (such as the Fauchère Pliska scale); (c) Residues on the hydrophobic surface (usually residues within a continuous hydrophobic sector identified by helical wheel analysis) are extracted, and their respective hydrophobic vectors are calculated (the magnitude of which is the hydrophobic value, and the direction is the angle of the residue in the helical wheel); (d) The vectors of all hydrophobic residues are summed to obtain the resultant vector, and then the magnitude of the resultant vector (i.e., the square root of the sum of squares) is calculated. This calculation process has been integrated into the AMPLLM feature calculation suite developed in this invention (such as the amp_amphiphilic_features.py script), which can realize high-throughput batch calculation.

[0091] (3) Correlation analysis and key feature screening The MIC values ​​of the 53 mutants were log2 transformed (Log2_MIC), and Spearman and Pearson correlation analyses were performed with the original eigenvalues ​​(using the Correlation Analysis module). The results are as follows: Figure 4 and Figure 5 As shown.

[0092] The results showed that the proportion of phenylalanine (F) in the hydrophobic surface was significantly positively correlated with Log2_MIC. P <0.001), meaning that the higher the F ratio, the higher the MIC and the worse the activity; the ratio of leucine (L) and tryptophan (W) in the hydrophobic surface, the average hydrophobicity Hyd, the hydrophobic moment HMom, HCS4_HF and Log2_MIC are significantly negatively correlated, that is, the higher these characteristic values, the lower the MIC and the stronger the activity.

[0093] Further analysis revealed that when L was replaced with W, replacing one or three W molecules reduced the MIC to its lowest level (Log2_MIC: 1.97 μg / mL), but replacing two or all W molecules with W did not show a significant change. Figure 6 This indicates that the amino acid composition of the hydrophobic surface exists in dynamic equilibrium, which plays a key regulatory role in antibacterial activity.

[0094] (4) Synergistic effect analysis of HCS3 and HCS4 The continuous hydrophobic regions of α-helical AMPs are mainly formed by the interactions of amino acids at positions i, i+3, and i+4, with i+4 located on the same side of the helix. The three-dimensional structures of Q1_19 and its mutants were predicted using ESMFold, and combined with HeliQuest helical wheel analysis, the spatial interaction sites of i+3 and i+4 were determined. Figure 7 and Figure 8 Hydrophobic cluster scores HCS3_HF and HCS4_HF were calculated and correlated with Exp_Log2_MIC. The results showed that HCS4_HF was moderately negatively correlated with MIC. Figure 9 Three-dimensional structural analysis revealed that a more "flattened" hydrophobic surface can reduce the MIC (micro-microstructure). Figure 10 and Figure 11 Group analysis showed that the HCS4_HF distribution was more concentrated in the low MIC group. Figure 12 and Figure 13 This indicates that HCS4 plays a leading role in regulating antibacterial activity, while HCS3 plays a synergistic role.

[0095] Further calculations of the directional hydrophobic moments (HMom_x, HMom_y, and HMom_mag) of HCS3 and HCS4 showed that HMom_HF_y was positively correlated with MIC, while HMom_HF_mag, HMom_HF_x, and HCS4_HF were negatively correlated with MIC. Figures 14-16 Force variation analysis showed that when Paris1 HCS4 ≥ Paris2 HCS4, the MIC decreased significantly; when Paris2HCS3 ≥ Paris1 HCS3, it was beneficial to reduce the MIC. Figures 17-19 Validation of 733 amphiphilic cationic AMPs with an amino acid content of 18-24 aa in the public AMP database revealed that both HCS3_HF and HCS4_HF were greater than 0 in the low MIC group, with HCS4 contributing more significantly. Figure 20 and Figure 21 This indicates a dynamic balance between HCS3 and HCS4, which jointly regulate antimicrobial activity. The constructed HCS4_Code system can be used to explain the relationship between the characteristics of α-helical antimicrobial peptides and their MICs.

[0096] Using the key features identified in the above correlation analysis as screening indicators, several candidate peptides (such as the W-enriched mutant of Q1_19) were designed, synthesized, and their Log2_MIC against *Pectinobacter softrot* was measured. Experimental results showed that mutants screened according to HCS4_HF > 0 and HCS3_HF > 0 had significantly lower MICs than random mutants, validating the effectiveness of this feature analysis model.

[0097] (5) Preliminary screening of a series of physicochemical and structural constraint characteristics Because the database of antimicrobial peptides to be screened contains a very large number, typically billions, using only HCS3_HF and HCS4_HF for screening would be computationally unsustainable and would not be able to complete the screening smoothly. Therefore, preliminary screening through a series of physicochemical and structural constraints is required to narrow down the number of peptide sequences to be screened.

[0098] In this embodiment, the average hydrophobicity (Hyd), hydrophobic moment (HMom), net charge (z), hydrophobic surface (HoF), and hydrophilic surface (HiF) were selected from the 54 α-helical antimicrobial peptide characteristics in Table 3 as preliminary screening characteristics for α-helical antimicrobial peptides.

[0099] Cationic antimicrobial peptides anchor to the bacterial surface via electrostatic attraction, which is the first step in their action. This embodiment, through experiments with 53 mutants, found that the net charge of highly effective antimicrobial peptides (MIC ≤ 7.8 μg / mL) almost entirely fell within the +7 to +11 range; therefore, this was set as a rigid constraint. If the net charge is too low (< +7), the binding force with LPS (lipopolysaccharide) is insufficient, making it difficult to effectively aggregate on the bacterial surface; if the net charge is too high (> +11), it may non-specifically bind to the host cell membrane (such as negatively charged glycoproteins on the surface of erythrocytes), increasing the risk of hemolytic toxicity.

[0100] Hydrophobicity (on the Fauchère-Pliska scale) determines a peptide's ability to insert into the lipid bilayer. This example, through analysis of 733 known amphiphilic cationic AMPs, shows that the Hyd of highly efficient peptides is almost always ≤ 0.5. Too high a Hyd (> 0.5) can cause peptides to self-aggregate in aqueous solutions, even forming irreversible precipitation, and can easily damage eukaryotic cell membranes, leading to hemolysis; too low a Hyd (< 0 or even negative) prevents penetration into the hydrophobic core, resulting in the loss of membrane lysis activity. Therefore, setting Hyd ≤ 0.5 as the initial screening upper limit ensures sufficient membrane insertion capacity while avoiding the toxicity caused by excessive hydrophobicity.

[0101] Hydrophobic moment (HMom) measures the degree of aggregation of hydrophobic amino acids on one side of an α-helix. A larger HMom indicates a more concentrated spatial orientation of the hydrophobic side chains, resulting in a more complete hydrophobic surface. In this study of 53 Q1_19 mutants, it was found that all peptides with low MIC (≤ 15.6 μg / mL) had an HMom ≥ 0.5. Therefore, HMom ≥ 0.5 was used as a mandatory initial screening threshold. If HMom < 0.5, the distribution of hydrophobic residues tends to be uniform, making it difficult to form a clear amphiphilic helix. Even with suitable charge and hydrophobicity, these peptides are unlikely to form pores or disrupt membrane integrity.

[0102] The three thresholds HMom ≥ 0.5, Hyd ≤ 0.5, and 7 ≤ z ≤ 11 together form a first-level screening funnel. In just a few minutes of calculation, it can eliminate sequences from billions of random sequences that either cannot bind to bacteria (insufficient charge), are too toxic or have poor solubility (too hydrophobic), or are not amphiphilic (too low hydrophobic moment).

[0103] Secondly, merely satisfying the three quantitative values ​​of net charge, hydrophobicity, and hydrophobic moment can indicate that the screened antimicrobial peptide has amphiphilic potential, but it cannot guarantee that the sequence will actually form an effective amphiphilic α-helix structure in three-dimensional space. It is also necessary to use the condition that the hydrophobic surface (HoF) ≠ absent and the hydrophilic surface (HiF) ≠ absent as the screening condition for step (2), to confirm that there is indeed a continuous region composed of multiple adjacent hydrophobic residues on one side of the helix, and that positively charged amino acids (such as lysine K and arginine R) are concentrated on the other side of the helix to form a hydrophilic surface. The hydrophilic surface should preferably contain positively charged amino acids K or R, which can increase the adsorption capacity of the peptide. The hydrophobic surface can anchor to the membrane surface and play a role in disrupting the membrane integrity, thereby performing antimicrobial activity.

[0104] This embodiment analyzes 733 known amphiphilic cationic AMPs. Experimental data show that sequences that satisfy HMom ≥ 0.5, Hyd ≤ 0.5, and 7 ≤ z ≤ 11, but lack HoF or HiF, almost all have extremely poor antibacterial activity (MIC > 64 μg / mL).

[0105] Furthermore, Hyd, HMom, and z are calculated extremely quickly (requiring only sequence composition information), instantly eliminating over 99% of invalid sequences from billions of sequences. In contrast, HoF / HiF validation requires spiral projection or the use of structure prediction tools, resulting in higher computational costs and making it unsuitable for direct operation on billions of sequences. Therefore, it is performed after the initial screening, validating only tens to hundreds of thousands of candidate sequences that meet basic physicochemical conditions.

[0106] Through the above process, and combined with HCS4_HF > 0 and HCS3_HF > 0 screening, this embodiment successfully established a preliminary screening and prediction model based on Python programmatic α-helical antimicrobial peptide characteristic calculation and activity, providing a theoretical basis and algorithmic support for subsequent large-scale screening.

[0107] (6) Screening of α-helical antimicrobial peptides with high antibacterial activity In this embodiment, among 53 Q1_19 mutants, the batch of highly efficient peptides with the lowest minimum inhibitory concentration (MIC) (Exp_Log2_MIC <= 3, equivalent to MIC < 8 μg / mL) was selected as the learning sample. The sequences with HCS4_HF values ​​in the 10-90% range were selected as centroids (Q10, Q90). For each key structural feature (HCS4_HF, HCS3_HF, HMom_HF_mag), the characteristic numerical range of this batch of highly efficient peptides was determined, forming the intervals 3.3 ≤ HCS4_HF ≤ 4.5, 2.8 ≤ HCS3_HF ≤ 3.6, and 7.6 ≤ HMom_HF_mag ≤ 8.7, which represent the characteristic core range common to the highly active peptides (Exp_Log2_MIC <= 3) in the 53 Q1_19 mutants.

[0108] Example 2: Screening based on a fixed-backbone combined mutant library This embodiment aims to verify the ability of the method of the present invention to design novel antimicrobial peptides from an artificially combined mutant library. Unlike the natural database in Example 2, the sequence space in this embodiment is completely artificially enumerated and is much larger (4.3 billion entries) to simulate a de novo design scenario.

[0109] I. Constructing a combined mutation library The fixed backbone sequence is: M***************GG (length 19, with methionine M and glycine G fixed at the beginning and end, and 16 variable sites in the middle). Amino acids shown in Example 1 to be positively correlated with low MIC (L and W) or to help maintain cationicity (R or K); and amino acids that increase peptide stability (A). The mutant amino acid library is limited to four amino acids: K, A, W, and L, because these four amino acids have a theoretical total sequence count of 4. 16 Approximately 4.3 billion entries.

[0110] The bottleneck of enumeration-based filtering of antimicrobial peptide data lies in hard drive storage and file retrieval. These two steps place enormous pressure on hard drive space and memory capacity, and are also a major reason for the time-consuming data filtering process. To reduce the computational burden on the CPU and memory, a method is proposed for 4... 16 Approximately 4.3 billion records were split into two parts, limiting the number of operations per cycle to 10. 8 The system automatically calculates the number of splits and combinations based on the total number of theoretical sequences, calculating in the following order: 4 15 ≈ 1.07×10 9 (>1×10) 8 ), 4 14 = 2.68×10 8 (>1×10) 8 ), 4 13 =6.71×10 7 (<1×10)8 Therefore, the original 16 variable positions are further divided: based on the mutant amino acid library, the first three positions are fixed (e.g., KKK, KKA...), resulting in 4... 3 = 64 sub-models (MKKK*************GG, MKKA*************GG...), each sub-model has 13 remaining variable positions, which can be used in 1×10 8 Enumeration filtering is performed under constraints.

[0111] II. Screening based on physicochemical constraints (first-level screening) Using the same feature calculation kit established in Example 1, each candidate peptide was calculated according to the following screening conditions: Step (1) Hydrophobic moment HMom≥0.5, hydrophobicity Hyd≤0.5, net charge 7≤z≤11.

[0112] Step (2) The hydrophobic surface HoF ≠ none, the hydrophilic surface HiF ≠ none, and the amino acid positions on the hydrophobic surface are (11, 4, 8, 15, 1).

[0113] Step (3) The hydrophobic cluster score HCS3_HF on the i+3 hydrophobic surface is greater than 0, and the hydrophobic cluster score HCS4_HF on the i+4 hydrophobic surface is greater than 0.

[0114] Steps (1)-(3) are dynamic filtering, which saves hard disk space and access speed by generating and filtering at the same time.

[0115] III. Multi-centroid distance sorting (second-level filtering) Using the exact same settings as in Example 2, the screening thresholds (Q90, Q10) were selected based on the characteristic values ​​of low-MIC antimicrobial peptides, ranked from 10% to 90%. The thresholds were set as follows: 3.3 ≤ HCS4_HF ≤ 4.5, 2.8 ≤ HCS3_HF ≤ 3.6, 7.6 ≤ HMom_HF_mag ≤ 8.7 (Step 4). 4312 high-quality antimicrobial peptides were obtained from 51.34 million candidate sequences. The entire process (construction of a combinatorial mutant library, combinatorial enumeration, and hierarchical screening) was significantly shortened. The original template screening method required 48 hours, while the hierarchical screening process was completed in 1.5 hours. The screening and filtering results showed no difference, verifying the scalability of this method in a sequence space of billions.

[0116] Example 3: Screening for highly efficient candidate peptides through site-directed mutagenesis of existing antimicrobial peptides Example 2 screened a batch of candidate peptides predicted to have high activity from an artificial combinatorial library. This example performed site-directed mutagenesis on existing antimicrobial peptides and verified the actual antimicrobial effects of these candidate peptides through experiments to prove the effectiveness of the screening method of the present invention. The fixed backbone sequence is: MRK*LKK*HR*KAK*VRGG, and the mutant amino acid library is hydrophobic amino acids (ALIWMVF). P is not conducive to the continuity of the hydrophobic surface, and Y is an aromatic amino acid that is not conducive to peptide stability. The theoretical total number of mutants is 2,401. Steps (1)-(3) filter out the remaining 304 mutants, and step (4) filters out the remaining 124 mutants.

[0117] (1) Selection of candidate peptides Five peptides (SEQ ID NO. 1 to SEQ ID NO. 5) were randomly selected from 124 peptides, and their sequences are shown in Table 4. All peptides were synthesized by solid-phase synthesis and purified by high-performance liquid chromatography (HPLC) to a purity >95%.

[0118] (2) MIC determination method The effects of each candidate peptide on *Pectinobacterium soft rot* (Pectinobacterium tumefaciens) were determined using a double dilution method. Pectobacterium carotovorum The minimum inhibitory concentration (MIC) of [the substance / organism] is determined. The specific steps are as follows: a) Inoculate *Pectinobacterium soft rot* into sterile NB culture medium and culture at 37°C with shaking for 6 hours until the logarithmic growth phase; b) Dilute the bacterial culture to 1×10⁻⁶ using sterile NB culture medium. 6 CFU / mL; c) Dissolve and dilute each candidate peptide with sterile water to a starting concentration of 2000 μg / mL; d) In a 96-well plate, add 50 μL of bacterial culture and 50 μL of serially diluted peptide solution (2-fold dilution series) to each well to make the final peptide concentration range from 2000 μg / mL to 0.98 μg / mL; e) Set up a positive control (50 mg / mL kanamycin) and a negative control (sterile water); f) Incubate at 37℃ for 12 hours until obvious turbidity appears in the negative control wells; g) Measure OD using an enzyme-linked immunosorbent assay (ELISA) reader. 600 Value, to completely inhibit bacterial growth (OD) 600 The lowest peptide concentration (with no significant difference from the negative control) was used as the MIC, and three replicates were set up for each experimental group.

[0119] (3) Experimental results The results are shown in Table 4. The MICs of the five candidate peptides ranged from 2.8125 to 15.625 μg / mL, with sequence_3179 showing a low MIC of 2.8125 μg / mL, indicating strong antibacterial activity. This result confirms that the candidate peptides screened from the mutant amino acid library by the method of this invention all possess genuine antibacterial activity, and the activity levels are highly consistent with the predicted sequence. Therefore, this invention can rapidly and efficiently screen α-helical antimicrobial peptides with high antibacterial activity against *Pectinobacterium softrot* without relying on large-scale training data.

[0120] Table 4. MIC determination results of candidate antimicrobial peptides

[0121] The two best sequences are sequence_2864 and sequence_3179, with MICs as low as 7.8125 μg / mL, especially sequence_3179, which has the lowest MIC (2.8125 μg / mL).

[0122] This embodiment also demonstrates that the method can obtain candidate antimicrobial peptides with a maximum MIC of no more than 15.625 μg / mL by site-directed mutation of hydrophobic surfaces.

[0123] Example 4: Comparison of different screening methods I. Design of the screening process This embodiment follows the method provided in Embodiment 2 to screen for a fixed-backbone combined mutant library. The first and second layer screening methods employ the following methods, which are then compared with existing screening methods: The basic screening conditions remain unchanged: (1) Hydrophobic moment HMom≥0.5, hydrophobicity Hyd≤0.5, net charge 7≤z≤11; Method a The fixed backbone sequence is: M****************G, and the mutant amino acid library contains I, K, and W; (the antimicrobial peptides obtained through screening are shown in Table 6, SEQ ID NO. 6-11) (2) The hydrophobic surface HoF ≠ none, and the hydrophilic surface HiF ≠ none; Method b The fixed backbone sequence is: M****************GG, and the mutant amino acid library contains K, A, W, and L; (the antimicrobial peptides obtained through screening are shown in Table 6, SEQ ID NO.12-21) (2) The hydrophobic surface HoF ≠ none, the hydrophilic surface HiF ≠ none, and the amino acid positions on the hydrophobic surface are (11, 4, 8, 15, 1). (3) HCS3_HF > 0 and HCS4_HF > 0; (4) The hydrophobic surfaces form a FLLLM; Method c Existing screening methods (CN2025114079962) require a Predicted Local Distance Difference Test (pLDDT) greater than 0.8, a helix number greater than 5, a net charge greater than +5, and a hydrophobicity percentage greater than 21%.

[0124] Method d The fixed backbone sequence is: M****************GG, and the mutant amino acid library contains K, A, W, and L; (the antimicrobial peptides obtained through screening are listed in Table 6 as SEQ ID NO.22-26) (2) The hydrophobic surface HoF ≠ none, and the hydrophilic surface HiF ≠ none; (3) HCS3_HF > 0 and HCS4_HF > 0; (4) The thresholds are set as follows: 3.3 ≤ HCS4_HF ≤ 4.5, 2.8 ≤ HCS3_HF ≤ 3.6, 7.6 ≤ HMom_HF_mag ≤ 8.7 All other screening steps remained unchanged. The screening computation time required for different methods was examined, and the highest and lowest MICs of the candidate peptides obtained were determined. The screening results are shown in Table 5.

[0125] Table 5. Comparison of screening results using different methods

[0126] Table 6. MIC determination results of candidate antimicrobial peptides

[0127] As shown in Tables 5-6, comparing methods a to c, since all methods limited the hydrophobic moment HMom ≥ 0.5, hydrophobicity Hyd ≤ 0.5, and net charge 7 ≤ z ≤ 11; and the hydrophobic surface HoF ≠ none, the hydrophilic surface HiF ≠ none, the MIC values ​​of the antimicrobial peptides were all between 15.625 and 31.25 μg / mL. In this case, step (4) (3.3 ≤ HCS4_HF ≤ 4.5, 2.8 ≤ HCS3_HF ≤ 3.6, 7.6 ≤ HMom_HF_mag ≤ 8.7) and limited the amino acid positions on the hydrophobic surface (11, 4, 8, 15, 1) for screening, the MIC values ​​of the finally screened antimicrobial peptides were closer to the preset antimicrobial activity, and the antimicrobial activity was higher.

[0128] Methods a and b both use enumeration to set basic screening conditions. Compared with the existing method c, the minimum MIC of the antimicrobial peptides obtained by screening is lower, and the computation time is significantly shorter (where method a uses 3 to the power of 16 enumerations, and b and d use 4 to the power of 16 enumerations, so method a has the shortest time). Using the method provided by this invention, through the screening steps of (1) to (3), computing power can be greatly saved, and the computation time can be shortened from 26 hours to about 1.5 hours. Method d adds step (4) (3.3 ≤ HCS4_HF ≤ 4.5, 2.8 ≤ HCS3_HF ≤ 3.6, 7.6 ≤ HMom_HF_mag ≤ 8.7) and hydrophobic amino acid positions (11,4,8,15,1). Compared with existing screening methods, it not only has higher screening efficiency, but also ensures that the MIC of the final screened antimicrobial peptide sequence is lower. The proportion of active candidate peptides obtained is much higher than that of unconstrained enumeration, truly achieving a single-machine billion-level high-efficiency screening.

[0129] II. Feature Filtering in Step (4) Next, in this embodiment, screening is performed according to the above method d, with the fixed backbone sequence as: M****************GG, and the mutant amino acid library as K, A, W, L, wherein step (4) is performed using the following three methods respectively: Method f: 3.3 ≤ HCS4_HF ≤ 4.5; 2.8 ≤ HCS3_HF ≤ 3.6; Method c, 3.3 ≤ HCS4_HF ≤ 4.5, 2.8 ≤ HCS3_HF ≤ 3.6, 7.6 ≤ HMom_HF_mag ≤ 8.7 (Example 3); Method g, 3.3 ≤ HCS4_HF ≤ 4.5, 2.8 ≤ HCS3_HF ≤ 3.6, 7.6 ≤ HMom_HF_mag ≤ 8.7, 5.41 ≤ HMom_HF_x ≤ 6.28; All other screening steps remained unchanged. The screening computation time required for different methods, the number of candidate peptides obtained, and the highest and lowest MICs among the obtained candidate peptides were examined. The screening results are shown in Table 6.

[0130] Table 7, Feature Filtering in Step (4)

[0131] As can be seen from Table 7, the comparison methods f, c and g all require 4 steps of screening and the required computation time is relatively consistent. However, when different features are used in step (4), the screening results are quite different. If the feature HMom_HF_mag is not added on the basis of the two features HCS4_HF and HCS3_HF, the number of antimicrobial peptides with low quality will increase significantly. Among the antimicrobial peptide sequences obtained by screening, some sequences with high MIC values ​​will also be selected, resulting in high MIC values.

[0132] Furthermore, if step (4) adds HMom_HF_x (method g) to the three features HCS4_HF, HCS3_HF, and HMom_HF_mag, it does not improve the final screening; instead, it increases the lowest MIC in the antimicrobial peptide sequences obtained through screening. This embodiment also tried adding other features (such as Hyd) to HCS4_HF and HCS3_HF in step (4), but the screening effect was still not as good as method g (Example 3). Therefore, method c is adopted, and including the three features HCS4_HF, HCS3_HF, and HMom_HF_mag in step (4) is the optimal screening method.

[0133] The above embodiments are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art can select different preparation methods and process parameters according to actual needs. Any modifications, equivalent substitutions, improvements, etc., made within the technical concept of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for screening α-helical antimicrobial peptides, characterized in that, Includes the following steps: (i) Extract at least one α-helix structure-related feature of the peptide sequence to be screened, the feature including HCS3_HF and / or HCS4_HF; (ii) Screen the peptide sequences according to the value of the aforementioned characteristics to obtain α-helical antimicrobial peptides with low MIC.

2. The method as described in claim 1, characterized in that, The features described in step (a) also include one or more of the following: Hyd, HMom, z, HoF, and HiF.

3. The method as described in claim 2, characterized in that, Step (II) of the screening process includes the following three steps: (1) HMom≥0.5, Hyd≤0.5, 7≤z≤11; (2) HoF ≠ none, HiF ≠ none; (3) HCS3_HF>0, HCS4_HF>0.

4. The method as described in claim 3, characterized in that, The screening in step (2) also includes step (4): using at least one α-helical antimicrobial peptide with a known low minimum inhibitory concentration (MIC) as the centroid, using the characteristic value of the centroid as a reference to determine the screening interval, and screening is performed according to the screening interval.

5. The method as described in claim 4, characterized in that, The centroid preferably includes at least two of (Q90, Q10); the eigenvalues ​​include the eigenvalues ​​HCS4_HF, HCS3_HF, and HMom HF_mag.

6. An α-helical antimicrobial peptide, characterized in that, Obtained by screening using the method described in any one of claims 1-5.

7. The α-helical antimicrobial peptide as described in claim 6, characterized in that, It has any one or more amino acid sequences as in SEQ ID NO. 1 to SEQ ID NO.

26.

8. The use of the α-helical antimicrobial peptide as described in claim 6 or 7 in the preparation of antibacterial drugs or agricultural fungicides.

9. The application as described in claim 8, characterized in that, The bacteria are either soft-rot pectinobacterium, Gram-negative bacteria, or Gram-positive bacteria.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.