Antimicrobial peptides

Antimicrobial peptides (AMPs) designed using bioinformatics tools offer a promising solution to the growing problem of antibiotic resistance. These peptides exhibit strong antimicrobial activity against resistant bacterial strains, providing a potential alternative to conventional antibiotics.

WO2025102148A1PCT designated stage expired Publication Date: 2025-05-22PROVINCIAL HEALTH SERVICES AUTHORITY +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/CA2024/050914
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-15
Filing Date
2024-07-08
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

The increasing resistance of microorganisms to conventional antibiotics poses a significant challenge in treating microbial infections, both in humans and animals. The overuse of antibiotics in livestock has contributed to the emergence of superbugs resistant to common antibiotics, and there is a need for new antimicrobial approaches to mitigate this issue.

Method used

The development of antimicrobial peptides (AMPs) with specific amino acid sequences that exhibit antimicrobial properties. These peptides can be used alone or in combination with other antimicrobial agents to treat or prevent infectious diseases. A bioinformatics approach, including the use of tools like AMPd-Up, is employed to design and generate novel AMP sequences with desired antimicrobial activities.

Benefits of technology

The AMPs described demonstrate significant antimicrobial activity against a range of bacterial strains, including Escherichia coli and Staphylococcus aureus, with some peptides showing no hemolytic activity. This suggests that AMPs can be effective alternatives to conventional antibiotics, reducing the risk of drug resistance and environmental impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2024050914_22052025_PF_FP_ABST
    Figure CA2024050914_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Antimicrobial peptides (AMPs) exhibiting broad spectrum antimicrobial activity are described. Such peptides are useful in treating or preventing infections and other conditions, and are of special interest for treating antibiotic-resistant bacterial pathogens, viruses, fungi, and other pathogens, or for anti-cancer applications. Treatment or prevention of infection in humans or animals is described as well as uses and methods for disinfecting or prevention of growth of microbes on a surface, a material, or in an environment, such as for environmental remediation. Development of new AMPs is arduous due to the practical limitations of classical protein-based discovery approaches. Two high throughput bioinformatics approaches leading to (1) de novo design of numerous antimicrobial peptides, and (2) identification of numerous antimicrobial peptides from known genomic sequences are described.
Need to check novelty before this filing date? Find Prior Art

Description

ANTIMICROBIAL PEPTIDESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to U.S. Provisional Patent Application No. 63 / 599,227 filed November 15, 2023 entitled “ANTIMICROBIAL PEPTIDES”, the entirety of which is hereby incorporated by reference.BACKGROUND OF THE INVENTION

[0002] The present disclosure relates generally to antimicrobial peptides for the treatment or mitigation of disease.

[0003] There is a need for peptides and pharmaceutical compositions thereof which are useful as therapies for microbial infections or as chemopreventative agents to slow or arrest the progression of microbial infections.

[0004] Use of antibiotics in livestock may have direct and indirect impact on medical use in addressing human disease. The ubiquitous use of antibiotics in all industries has contributed to the emergence of superbugs which have become resistant to the most common antibiotics. Some strains illustrate multi-drug resistance, which is a global concern. Although the search for new antibiotic approaches continues in earnest to address challenges in both human and animal health.

[0005] Consumers have concerns about the use of prophylactic antibiotics due to the potential environmental impact, increasing drug resistance, and the possible consumption of antibiotic lace meat, egg, or dairy products. Restrictions on prophylactic antibiotic use in livestock that have been implemented to address these concerns, but have downstream consequences such as increased rates of animal infections, leading to productivity loss due to the increase disease burden. Sick animals that are then treated with antibiotics will continue to contribute to potential drug resistance. Poultry and swine raised in close quarters are particularly susceptible to the rapid spread of disease. Different approaches to reducing infections disease in livestock animals are under development, including investigation of new antibiotic approaches, and development of vaccines. While small molecule drugs have conventionally been used, antimicrobial peptide and polypeptide therapeutic approaches are also under consideration.

[0006] International Patent Publication Nos. WO 2020 / 118427 (Bird et al.) and describes antimicrobial peptides.

[0007] It is desirable to find new antimicrobial approaches to reduce the onset and spread of disease in humans and animals.SUMMARY

[0008] Peptides and / or amino acid sequences with antimicrobial properties are described herein. A process for use as an antimicrobial peptide (AMP) sequence generation tool is described, for de novo AMP design.

[0009] A bioinformatics approach, starting with sequences exhibiting effect, and making strategic modifications thereto, has led to the discovery of antimicrobial peptides. In a bioinformatics approach, sufficient similarity among sequences can be maintained so as to permit functional equivalency. Sequences similar to isolated sequences from which a consensus is derived are also described. Such similar sequences contain conserved amino acid substitutions and a limited number of non-conserved modifications.

[0010] It is an object of the present disclosure to provide antimicrobial peptides, which may obviate or mitigate at least one disadvantage of previous antimicrobial approaches.

[0011] There is described herein an antimicrobial peptide comprising an amino acid sequence according to any one of SEQ ID NO:1 to SEQ ID NO:8066, or SEQ ID NO:8076 to 8433, or a variant thereof having at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or having 100% amino acid sequence identity thereto.

[0012] Further, there is described herein a composition comprising the described antimicrobial peptide together with a suitable carrier or excipient.

[0013] The composition comprising the described antimicrobial peptide may be used in treatment or prevention of a disease or condition, such as infectious disease or a tumour.

[0014] The composition comprising the described antimicrobial peptide may be used for application to a surface for disinfecting or prevention of growth of microbes. For example, such a use may be for disinfecting of a surface, a material, or an environment such as for use in environmental remediation. A method is described herein for cleaning or disinfecting of a surface, a material, or an environment comprising application of the described composition thereto.

[0015] A use for the antimicrobial peptide is provided, for treatment or prevention of a disease or condition in a subject in need thereof. Further, the use of the antimicrobial peptide for preparation of a medicament for treatment or prevention of a disease or condition in a subject in need thereof is also described herein. Additionally, a method of treating or preventing a disease or condition is described, comprising administering to a subject in needthereof an effective amount of the antimicrobial peptide or composition thereof. The disease may be, for example, an infectious disease or a tumour. The subject may be a human or an animal, such as a livestock animal, for example poultry, or a companion animal, such as a domestic animal or pet. The infectious agent may be a Gram-negative bacteria, Grampositive bacteria, acid fast bacteria, bacteria resistant to other drugs, a virus, a fungi, or a parasite.

[0016] A lipid vesicle comprising the antimicrobial peptide is described. A nucleic acid molecule encoding the antimicrobial peptide is also provided, as is a vector comprising such a nucleic acid molecule.

[0017] A method of identifying a target molecule associated with an infectious agent is described, in which the target molecule targets, or is affected by, the antimicrobial peptide. The method comprises the step of screening a library of candidate target molecules associated with the infectious agent, for a molecule that targets or affects the antimicrobial peptide. A kit for conducting such a method for identifying a target molecule associated with an infectious agent is also described, in which the kit comprises the antimicrobial peptide described herein together with instructions.

[0018] Other aspects and features of the present disclosure will become apparent to those ordinarily skilled in the art upon review of the following description of specific embodiments in conjunction with the accompanying figures.BRIEF DESCRIPTION OF THE FIGURES

[0019] Embodiments of the present disclosure will now be described, by way of example only, with reference to the attached Figures.

[0020] Figure 1 shows the architecture of the recurrent neural network (RNN) language model of Example 1.

[0021] Figure 2 illustrates length and net charge distributions of the generated sequences of Example 1.

[0022] Figure 3 shows sequence similarity distribution of the generated sequences to the training set of Example 1.

[0023] Figure 4 shows sequence similarity distribution of the generated sequences to known AMPs of Example 1. The known AMP sequence set comprises 4,538 distinct sequences downloaded from APD3 and DADP databases.

[0024] Figure 5 shows distributions of pairwise sequence similarities between different sequence sets in Example 1.

[0025] Figure 6 shows in vitro validation results of the 58 selected AMPs. (Panel a) Antimicrobial and hemolytic activities of the 40 peptides that were active against at least one bacterial strain of Escherichia coli ATCC 25922 and Staphylococcus aureus ATCC 29213, for Example 1 are grouped into List A, List B, and List C. (Panel b) shows proportions of peptides displaying antimicrobial activity. (Panel c) plots antimicrobial activity of tested peptide AMPlify scores and AMPd-Up scores.

[0026] Figure 7 shows AMP mining workflow. The AMP mining workflow utilizes the rAMPage (Lin et al., 2022) cleavage module with ProP (Duckert et al., 2004) to cleave precursor sequences and AMPlify (Li C. et al., 2022) (Li C. et al., 2023) to predict AMP sequences in Example 2.

[0027] Figure 8 shows length and net charge distributions of the AMPs predicted by AMPlify from the UniProtKB / Swiss-Prot database.

[0028] Figure 9 shows Sequence similarity distributions of the predicted AMPs from the UniProtKB / Swiss-Prot database to known AMPs for Example 2.

[0029] Figure 10 shows distribution for the number of predicted mature AMP sequences found within each AMP precursor sequence mined from the UniProtKB / Swiss-Prot database for Example 2.

[0030] Figure 11 shows categorization of the AMP entries identified from the UniProtKB / Swiss-Prot database based on source organisms for Example 2.

[0031] Figure 12 shows structural similarity distributions of the short cationic AMPs mined from the UniProtKB / Swiss-Prot database to known chicken AMPs for Example 2.

[0032] Figure 13 shows antimicrobial and hemolytic activities of the 13 novel AMPs mined from the UniProtKB / Swiss-Prot database that were active against at least one bacterial strain of Escherichia coli ATCC 25922 and Staphylococcus aureus ATCC 29213 for Example 2.

[0033] Figure 14 shows a visualization of antimicrobial activity of the 38 tested AMPs with respect to AMPlify score and associated structural similarity to known chicken AMPs in Example 2.DETAILED DESCRIPTION

[0034] Peptides and / or amino acid sequences with antimicrobial properties are described herein. A bioinformatics approach, starting with sequences exhibiting effect, and making strategic modifications thereto, has led to the discovery of antimicrobial peptides. In a bioinformatics approach, sufficient similarity among sequences can be maintained so as to permit functional equivalency. Sequences similar to isolated sequences from which aconsensus is derived are also described. Such similar sequences may contain conserved amino acid substitutions together with a limited number of non-conserved substitutions, such as modifications or deletions, but while still maintaining functionality.

[0035] An AMP sequence generation method and tool (AMPd-Up) is described. AMPd-Up adapts recurrent neural network (RNN), and samples candidate AMP sequences from multiple model instances trained with different random initializations. With AMPd-Up, a user is able to generate novel AMP sequences in a de novo manner without any pre-designed features specified. 58 AMP sequences were generated by AMP-Up for validation, and 40 showed antimicrobial activity against Escherichia coli and / or Staphylococcus aureus. The AMP sequence generation tool AMPd-Up is described, and the following AMPs (SEQ ID NO:1 - SEQ ID NO: 58) are described in Table 1.Table 1SEQ ID NOs 1 - 58SEQ ID NO:1. DLLSGLGKAAKKVAKTVLKNLLKCSEQ ID NO:2. NLLDTLKNLAKKLAKKLLKKLLKKLSEQ ID NO:3. NLLSTLLDAAKKAAKGAAKSAAKKLAKKLAKKLSEQ ID NO:4. HLLSGLLSAAKKAAKKAAKKALKKLLKKLLKKLSEQ ID NO:5. GLFSLLKKLLKKLLKKLLKKLLKKLLKKLSEQ ID NO:6. NLLDTLKKKAKKVAKKVLKKLLKKLLKKLSEQ ID NO:7. FLPSI I KGAAKKLPKI FCKI LKKCSEQ ID NO:8. GLLSLLKKLLKKLLKKLLKKLSEQ ID NO:9. DLLKTLGKAAKKAAKTALKAALKGLLKKLAKKLSEQ I D NO: 10. VLGGLLKKLLKKLLKKLSEQ ID NO:11. HLLSLLKKAAKKLLKKLLKKLAKKLSEQ ID NO:12. CLLDTLKCVAKGVAGTLLDTLKCKITGKCSEQ ID NO:13. KIFGKILKKLLKKLLKKLLKKLSEQ ID NO:14. ALPSLLKKLAKKLAKKLLKKLLKKLLKKLLKKLSEQ ID NO:15. N LLDTLKN VAKN VAKN VLDTLKCKITCKCSEQ ID NO:16. FLPIIAGLAAKFLPKIFCKITKKCSEQ ID NO:17. FLPIIAGLAAKLLPKLFCKITKKCSEQ ID NO:18. WLPKIAGKIAGKLLKKLLKKIKKKSEQ ID N0:19. FLPKIAGKAAKKLPKIFCKITKKCSEQ ID NQ:20. TLPDVAKNVAKNVAKTVLDTLKCKITGKCSEQ ID N0:21. KLFGKLLKKLKKILKKIAKKIKKKLSEQ ID NO:22. GLLSLLKKIGKKIGKLLSEQ ID NO:23. DLLKTLKKIAKKLLKTLLKKLLKKLLKKLSEQ ID NO:24. KLFGKI LGKI AKKI LGKI LGALLSKLLSALSEQ ID NO:25. DLLSCLKKKGKCVLKNLSEQ ID NO:26. RLPSLFKKLFKKIAKWGKIAKKILKKSEQ ID NO:27. RLPSIIPGIAGKLGGLLGGLLKGLSEQ ID NO:28. CLPSLLPSLFKKLSEQ ID NO:29. SLPSILSGIAGKLSEQ ID NQ:30. RLPRIFRGIRGKLSEQ ID N0:31. PLPPI I PGI AGKLLGGLLGLLKKLSEQ ID NO:32. YLPSVLPSVLKPLSEQ ID NO:33. PLPPIIPGLASGLLSGLCSEQ ID NO:34. KLPSIIKAAAKALPKLFSEQ ID NO:35. QLPRIAGKIAKKLSEQ ID NO:36. QLPSVLPAIAKALSEQ ID NO:37. CLPSILCSEQ ID NO:38. MLPSIAGAAAKGLPKLFCKITKKCSEQ ID NO:39. M LPKI FGKI FKKI LKKI LKKI LKKI LKKLLKKLSEQ ID NQ:40. MLPSILGALLKLLSEQ ID N0:41. MLPKIAGKIAKKLSEQ ID NO:42. MLPKIAGAIAKLLSEQ ID NO:43. WLPKIAGKIAGKLSEQ ID NO:44. CLPSILCKITKKCSEQ ID NO:45. FLPKIFKKIAKKLSEQ ID NO:46. VLGSLLKGLLKKLSEQ ID NO:47. ALPSIIKGLLKKLSEQ ID NO:48. LLPSLLKGLLKKLSEQ ID NO:49. ALLSLLKKLLKKLSEQ ID NQ:50. FLPKIAGKIAGKLSEQ ID N0:51. ALPSLLKKLLKKLSEQ ID NO:52. YLPSVLKGLLKKLSEQ ID NO:53. LLPSLLKGLAKKLSEQ ID NO:54. QLPKIAGKIAKKLSEQ ID NO:55. FLPKIFKKIAKKISEQ ID NO:56. GLLSLLKKLLKKLSEQ ID NO:57. ILGKLLKKLLKKLSEQ ID NO:58. FLPKIAGKIAKKL

[0036] Further, antimicrobial peptides were mined from the UniProt / Swiss-Prot database (2022_02 release) on the basis of selection parameters described herein, from which 8008 eukaryotic peptides were identified. Based on original entries of their parent sequences, certain sequence were discovered herein to be antimicrobial peptides (AMPs). A cleavage module for precursor sequence cleavage was adapted for antimicrobial activity prediction. None of the original entries of the parent sequences identified in the database was associated or annotated with antimicrobial activities therein. Herein is described the identification and synthesis, as well as in vitro testing of the identified peptides, many of which exhibit striking antimicrobial activity against at least Escherichia coli and / or Staphylococcus aureus.

[0037] The AMP sequences were not located in APD3 (aps. unmc.edu / AP) or DADP (split4.pmfst.hr / dadp) databases, and the original entries of their parent sequences were not annotated with antimicrobial effect or keywords in the UniProt / Swiss-Prot database.

[0038] The antimicrobial peptides SEQ ID NO:59 - SEQ ID NQ:8066 were identified, and at least the following peptides were synthesized and / or evaluated further herein:

[0039] SEQ ID NQ:2208. WLLVLQRGHRLASIKHVCQLSERKR

[0040] SEQ ID NQ:2075. AWDWAKNWVFWTCSYLLNFLYHHHCSRDLIRR

[0041] SEQ ID NO:3751. DLWFGLKTEGELAFVKRTIFDVIYKSAKRKR

[0042] SEQ ID NO:1978. GLQKLINKIKSQMSRFSTKTNKICGP

[0043] SEQ ID NO: 1159. TRWEAMKAKATELRVCCARRKR

[0044] SEQ ID NO:3162. VWEFFEAGAGLRASNASKKIYGVAKRFRR

[0045] SEQ ID NO:2425. FWNLNKKKKFFYKTVKNSIGQVILRDMSNN

[0046] SEQ ID NQ:460. GVLKIPLAFVQMIAISIALICLLIP

[0047] SEQ ID NO:613. FMRFGKRFMRFGRFGKSAEVENNIQIAAKQS

[0048] SEQ ID NO:1114. RWSQWAYAGRAQFCAVRRSVFGFSVRSGMVCRPRR

[0049] SEQ ID NO:3941. SYWWWGFHKNVIDRREAFYADLAEKKKAEN

[0050] SEQ ID NO:3534. AKETLEKKKLLKELWESSKKVH

[0051] SEQ ID NO:2254. AWSSLHSGWAKRAWQDMSSAWGKR

[0052] SEQ ID NO:3319. FMIHNTLGLFYRSVVRNEIKKR

[0053] SEQ ID NO:1851. FTLFFPIFMIVVCVISFFNLHKR

[0054] SEQ ID NQ:2950. YFAFYNLFHCLKKDSNNVEMYLKLLKCRLIRSKC

[0055] SEQ ID NO:285. WIIQCWWRQVLEKLLAKRRR

[0056] SEQ ID NO:1291. GLIKRIIRQKR

[0057] SEQ ID NO:3881. ILTPVSTVLALLLIAALILLKR

[0058] SEQ ID NO:1478. ASWEYLVHHVMAMGAFFSGIFWKR

[0059] SEQ ID NO: 1847. HTSRLKARKHSKRRVRYICEFTIPQ

[0060] SEQ ID NQ:3901. ALWKLGEPSDHLLQWLVLHLASLHLRLLFKR

[0061] SEQ ID NO:483. ILDKVRLWLARIRLNLLKR

[0062] SEQ ID NO:3598. YFAFHNLFHCLKKDSSHVEMYLKLLKCRLIQSNC

[0063] SEQ ID NQ:4032. ARLDVAAEFRKKWNKWALSRGKR

[0064] SEQ ID NO:2797. RSIIRNGIRTLTWRKETKRKKK

[0065] SEQ ID NO:2444. VLKKSEIRIDFISRTILHI

[0066] SEQ ID NO:416. GVLKKVIRHKR

[0067] SEQ ID NO:3895. QFFLPMILRLYVSRLFISKL

[0068] SEQ ID NQ:2048. LYVFCCRTRAKTPSVIYTINLVVTDLLVGLSLPTR

[0069] SEQ ID NO:3732. ARSKKSPWPGWAPLAAPHSH

[0070] SEQ ID NO:2452. VMRSAGSRSSKR

[0071] SEQ ID NO:833. GRGRPKGRGRGRPRGRPRGSKR

[0072] SEQ ID NO:2787. QVWKRR

[0073] SEQ ID NO: 1994. FQWHGRKPGPETGVPQSRPPIPR

[0074] SEQ ID NO:3682. RIRGHTIKG

[0075] SEQ ID NO:2913. FLEKPPGPLGARPLGGK

[0076] SEQ ID NQ:3700. TPKFVGK

[0077] SEQ ID NO:2343. VRCRFACC

[0078] SEQ ID NQ:2074. SPLHKR

[0079] Additional peptides described and tested herein include SEQ ID NOs: 8076-8281 as listed in Table 13, and the peptides listed in Table 15, including SEQ ID NOs: 8282-8433.

[0080] These peptides and their pharmaceutical compositions and modifications thereof are also useful as therapies for microbial infections or as chemopreventative agents to slow or arrest the progression of microbial infections. Uses are also envisaged in agriculture and livestock applications, in addition to human health applications. Antimicrobial peptides maybe used as disinfection or cleaning agents, on surfaces, or integrated into environments where environmental remediation may be desired. Modifications of peptides described herein may include but are not limited to incorporation of the peptides or their modifications in lipid vesicles for enhanced therapeutic delivery and the modulation of other ADMET properties (absorption, distribution, metabolism, excretion, toxicity) as well.

[0081] Chemical modifications of the peptides are described, which are known to individuals skilled in the art of peptide chemistry to be useful to enhance stability and otherwise make the peptides more drug-like and useful for the desired applications. Such modifications include peptide cyclization and the use of amino acids of opposite chirality - so-called D- amino acids. Such modifications also include alternative backbone chemistries and novel side chains that retain the binding specificity.

[0082] Also described is the application of the peptides, and modifications of the peptides obvious to those skilled in the art, to other microbial targets. Antimicrobial therapies useful and effective in one type of infection may be useful and effective in other diseases.

[0083] Also described are vector constructs incorporating the disclosed peptides and / or their amino acid sequences and coding nucleic acid sequences for the purposes of the production of antimicrobial peptides.

[0084] The peptides described herein, and the modifications thereof are also useful in combination with other antimicrobial agents for the treatment or prevention of disease, such as an infectious disease or a cancer.

[0085] Uses of the AMPs either alone or as part of a kit to isolate or identify target molecules, or molecules upon which an effect is observed, that may be associated with the infectious agent, are also described herein.

[0086] The peptides and / or amino acid sequences described herein have selective antimicrobial properties. Further aspects and advantages will become apparent from consideration of the ensuing description of various embodiments. A person skilled in the art will realize that other embodiments, combinations and variations are possible, and that the details described herein can be modified in a number of respects, all without departing from the overall concept. Thus, the following drawings, descriptions and examples are to be regarded as illustrative in nature and not restrictive.

[0087] Treatment or prevention of a disease or condition encompasses treatment before and after outward signs or symptoms of the disease or condition are present in the subject. For example, a subject exposed an infectious agent may or may not exhibit symptoms. Further, the prevention or prophylaxis of a disease or condition may encompass partial prevention,lessening of severity when onset occurs, decreasing likelihood of outward signs or symptoms, or preventing the spread of infection by keeping severity so low as to be undetectable or negligible. Treatment and prevention may involve modulating the immune system of the subject to the extent that the subject’s own defenses ward off the disease or condition, such as infection. An inflammatory or anti-inflammatory effect of the peptides described herein may modulate the outward signs or symptoms of a disease or condition.

[0088] Anti-cancer activity, such as against solid tumours or liquid tumours, may be modulated by peptides as described herein. Indirect or direct attack on cancer cells by the peptides described herein through effects on the immune system by the peptides may alleviate cancerous cell growth.

[0089] An antimicrobial peptide comprising: an amino acid sequence according to any one of SEQ ID NO:1 to SEQ ID NO:8066, or SEQ ID NO:8076 - 8433, or a variant thereof, having at least 95% amino acid sequence identity thereto. The threshold of amino acid sequence identity for the variant may optionally be at least 96%, at least 97%, at least 98%, at least 99%, or may be 100% amino acid sequence identity to any one of SEQ ID NO:1 to SEQ ID NQ:8066 or SEQ ID NQ:8076 to SEQ ID NO:8433.

[0090] The antimicrobial peptide may be modified, or may be a variant which comprises a modification that is a conservative amino acid substitution. Such amino acid sequences as are known in the art may include the following candidates, with the substitutable options shown in parentheses: Ala (Gly, Ser); Arg (Gly, Gin); Asn (Gin, His); Asp (Glu); Cys (Ser); Gin (Asn, Lys); Glu (Asp); Gly (Ala, Pro); His (Asn, Gin); lie (Leu, Vai); Leu (lie, Vai); Lys (Arg, Gin); Met (Leu, lie); Phe (Met, Leu, Tyr); Ser (Thr, Gly); Thr (Ser; Vai); Trp (Tyr); Tyr (Trp, Phe); and Vai (lie, Leu). Furthermore, ‘functional’ variants, mutations, insertions, or deletions encompass sequences in which the activity or function is substantially the same as that of the reference sequence from which the altered sequence is derived. Activity or function may be tested according to such parameters as described herein, such as minimum inhibitory concentration (MIC) or minimum bactericidal concentration (MBC). Further, it may be desirable to reduce the antigenicity of a peptide, for example by PEGylated, or the peptide may comprise a D-amino acid. The peptide may be cyclized.

[0091] A composition is described herein which comprises the antimicrobial peptide as described herein, together with a suitable excipient, such as a pharmaceutically acceptable carrier. The composition may be one that is suitable for use in treatment or prevention of a disease or condition, such as an infectious disease, or a cancer, such as may be attributable to a solid tumour or a liquid tumour.

[0092] The composition may be formulated for oral, injectable, rectal, topical, transdermal, nasal, or ocular delivery. Such compositions can thus be administered to subjects in need thereof through any acceptable route, such as topically (as by powders, ointments, or drops); oral tablets, capsules, gels or liquids; or rectal suppositories. Further modes of delivery include mucosally, sublingually, parenterally, intravaginally, intraperitoneally, bucally, ocularly, or intranasally. The composition may be formulated for application to a surface for prevention of growth of bacteria or other microbes, wherein the antimicrobial peptide would be formulated with a suitable carrier. Such a composition may be used in a method of cleaning or disinfecting of a surface, a material, or an environment, such as an environment in need of environmental remeciation.

[0093] When formulated for oral use or administration in a liquid formulation, the excipients or ingredients may include but are not limited to those accepted in the art of pharmaceutical formulations, for example emulsions, microemulsions, solutions, suspensions, syrups and elixirs. Liquid dosage forms may contain inert diluents such as water or other solvents, solubilizing agents, emulsifiers, ethyl alcohol, isopropyl alcohol, ethyl carbonate, ethyl acetate, benzyl alcohol, benzyl benzoate, propylene glycol, 1 ,3-butylene glycol, or dimethylformamide. Further, a liquid formulation may comprise oils such as cottonseed, groundnut, corn, germ, olive, castor, and sesame oils; glycerol, tetrahydrofurfuryl alcohol, polyethylene glycols and fatty acid esters of sorbitan; and mixtures thereof. Besides inert diluents, such oral compositions can also include adjuvants such as wetting agents, emulsifying and suspending agents, sweetening, flavoring, and perfuming agents. In livestock, such as poultry and in particular turkey or chicken farming, application in feed or liquid supplements may be considered.

[0094] The composition may be one that is lyophilized. The composition may comprise a suitable preservative.

[0095] The composition may be one that is distributed evenly in a diet intended for livestock, such as swine or poultry. Such a composition may be sprayed or mixed into a ground or powdered ingredient, and then mixed evenly into a coarser animal feed to ensure even distribution.

[0096] A composition for application to a surface is described herein, for use in the prevention of growth of bacteria, said composition comprising the antimicrobial peptide according to any one of claims 1 to 5 and a suitable carrier.

[0097] A use of the antimicrobial peptide is provided herein, for treatment or prevention of a disease or condition in a subject in need thereof, such as an infectious disease. The disease or condition may also be a cancer, such as a solid tumour or a liquid tumour.

[0098] Further, a use is provided for preparation of a medicament for treatment or prevention of such a disease or condition in a subject in need thereof. A method of treating or preventing such a disease or condition is also described herein, which comprises administering to a subject in need thereof an effective amount of the peptide or the composition described herein.

[0099] The disease or condition may be one attributable to Gram-negative bacteria, or it may be a disease or condition attributable to Gram-positive bacteria. The disease or condition may be one that is attributable to acid fast bacteria, or one that is attributable to bacteria that has become resistant to other drugs. Such diseases or conditions may be ones attributable to E. coli, S. enterica, S. aureus, P. aeruginosa, S. pyogenes, M. smegmatis, MRSA, S. enterica serovar Enteritidis, S. enterica serovar Heidelberg, A. baumannii, K. pneumoniae, or E. faecalis bacteria, for example. The disease or condition may be one attributable to a virus, a fungi, or a parasite.

[0100] Further, the disease or condition may be a cancer, such as a solid tumour or a liquid tumour.

[0101] A lipid vesicle may be used to deliver the antimicrobial peptide described herein. A nucleic acid molecule encoding the antimicrobial peptide described is also envisioned. A vector comprising the nucleic acid molecule is also encompassed.

[0102] A method of identifying a target molecule associated with an infectious agent is described, wherein the target molecule targets the antimicrobial peptide described herein. Such a method involves the step of screening a library of candidate target molecules associated with the infectious agent, for a molecule that is affected by the antimicrobial peptide. The infectious agent may be Gram-negative bacteria, or may be Gram-positive bacteria. Further, the infections agent may be acid fast bacteria, bacteria that has become resistant to other drugs, a virus, a fungi, or a parasite. Exemplary infectious agents include but are not limited to E. coli, S. enterica, S. aureus, P. aeruginosa, S. pyogenes, M. smegmatis, MRSA, S. enterica serovar Enteritidis, S. enterica serovar Heidelberg, A. baumannii, K. pneumoniae, or E. faecalis bacteria. Further, a method of identifying a target molecule for modulating biological activity is described, wherein the target molecule targets a peptide as described herein. The method comprising the step of screening a library of candidate target molecules for a molecule that targets the peptide. Modulating of biologicalactivity may comprise anti-tumour action, anti-inflammatory action, or inflammatory action. In such methods of target identification, the screening of a library of candidate target molecules may comprise in silico screening.

[0103] A kit is encompassed herein for identifying a target molecule associated with an infectious agent. Such a kit comprises an antimicrobial peptide as described herein together with instructions for conducting the method described herein for identifying a target molecule associated with the infectious agent. Optionally, additional reagents may be provided with the kit. A kit for identifying a target molecule for modulating biological activity, is also described. Such a kit comprises a peptide, as described herein, together with instructions for conducting a screening method for molecules that bind to the peptide.

[0104] There are instances herein where the term “putative” precedes the term “antimicrobial peptide”. The term makes no implication regarding antimicrobial action of the peptide but acknowledges difference between the establishment of antimicrobial effect versus use as an approved drug, given the years of downstream efforts after antimicrobial activity is established. Thus the term putative as used herein acknowledges the long term efforts required in establishing commercial viability and bringing such a product to market.

[0105] A naming strategy utilized herein is indicative of the parent peptide (Organism base) and its AMP# name / numbering nomenclature. Organism base: indicated by the first two letters of both parts of the organism's latin name; AMP #: Next highest number based on total number of AMPs discovered in that organism. Mutants are listed following the AMP name where applicable following the format: original amino acid, position, desired amino acid. Modified peptides described herein follow the same naming conventions but include Dap, Dab, Orn or lower case letters to indicate D form (such as “k”) as the new amino acid substitutions. See below for codes and explanations. D-Lysine may be represented as lower case “k”, indicating D- instead of L-form. Dab is represented by “II”, indicative of diaminobutyric acid; Dap is represented by J, indicative of diaminopropionic acid; Hor is indicative of X**; which indicates homoarginine instead of L-Lysine; Orn is represented by O, indicative of ornithine; and lower case is indicative of the D-form of the amino acid at the stated position. The use of “X” in the peptide naming convention is not to be confused with the meaning of “any amino acid”.

[0106] For Example, “RaCa” represents Rana Catesbeiana, for example a parent name of: RaCa2 FFPIIARLAAKVIPSLVCAVTKKC (described as SEQ ID NO:43 of W02020 / 118427 A1 to Bird et al.) may have progeny named RaCa2A19K FFPIIARLAAKVIPSLVCKVTKKC (herein SEQ ID NO:8269), indicative of mutation A19K(Lysine (K) amino acid substitution for Alanine (A) at position 19,). Naming with regard to truncations is indicated by the peptide name followed by “tX_Y”, meaning the residues from X to Y are present in the newly truncated sequence, for example “t15_23” indicates a truncated version of the originally named peptide from positions 15 to 23. C-terminal amidation is indicated as “_a”, meaning “-CONH2”, for example “RaCa2_a” indicates the peptide is amidated.

[0107] Examples

[0108] The following Examples outline exemplary embodiments and / or studies conducted pertaining thereto. While the Examples are illustrative, they should not be viewed as limiting.

[0109] Example 1

[0110] Recurrent Neural Network Allows For Potent And Novel Antimicrobial Peptide Sequence Generation

[0111] Summary. Antibiotic resistance is recognized as an imminent and growing global health threat. New antimicrobial drugs are urgently needed due to the decreasing effectiveness of conventional small-molecule antibiotics. Antimicrobial peptides (AMPs), a class of host defence peptides, are emerging as promising candidates to address this need. The potential sequence space of amino acids is combinatorially vast, making it possible to extend the current arsenal of antimicrobial agents with a practically infinite number of new peptide-based candidates. However, mining naturally-occurring AMPs, whether directly by wet lab screening methods or aided by bioinformatics prediction tools, has its theoretical limit regarding the number of samples or genomic / transcriptomic resources researchers have access to. Further, manually designing novel synthetic AMPs requires prior field knowledge, restricting its throughput. In silico sequence generation methods are gaining interest as a high-throughput solution to the problem. Here, AMPd-Up was introduced, a recurrent neural network (RNN) based tool for AMP sequence generation, and its utility was demonstrated over existing methods. Validation of candidates designed by AMPd-Up through antimicrobial susceptibility testing revealed that 40 of the 58 generated sequences possessed antimicrobial activity against Escherichia coli and / or Staphylococcus aureus. These results illustrate that AMPd-Up can be used to design synthetic novel AMPs.

[0112] 1.1. Introduction

[0113] The worldwide overuse of antibiotics has created an alarming number of bacteria reported to possess antibiotic resistance, resulting in conventional antibiotics to beless effective (Reardon, 2014) It is estimated that 1.27 million people died due to antibiotic resistance in 2019 (Murray et al, 2022), and the fast speed of the bacterial evolution of antibiotic resistance is expected to greatly increase this death toll in the next few decades (O’Neill, 2014; Laxminarayan et al., 2013) Moreover, the sluggish pace of discovery and development of new therapeutics is exacerbating this public health crisis (Koo et al., 2019). As a result, novel effective substitutes for conventional antibiotics are urgently needed as weapons to fight against the multi-drug resistant bacteria referred to as “superbugs”.

[0114] Antimicrobial peptides (AMPs), a diverse class of short and often cationic peptides, are considered to be one viable alternative to conventional antibiotics (van der Does et al., 2019). Naturally-occurring AMPs are observed amongst all forms of life (Zhang and Gallo, 2016). In higher eukaryotic organisms, AMPs have co-evolved with environmental microbes as part of the innate immunity of the respective hosts (Zhang and Gallo, 2016). Microbes can also produce AMPs for inter-competition purposes against the growth of other microbes (Zhang and Gallo, 2016). Most of the known AMPs reported in public databases are antibacterial, with some AMPs active or additionally active against other microbes (e.g. fungi, viruses) (Wang et al., 2016). Unlike most conventional antibiotics, which have specific functional or structural targets, the majority of AMPs act directly on the bacterial membranes or cell walls leading to non-enzymatic disruption, with some eukaryotic AMPs performing additional modulation of the host immune system (Zhang and Gallo, 2016) (Nguyen et al., 2011). As a result, it may be more difficult for bacteria to develop resistance to AMPs compared with conventional antibiotics (Boman, 2003). However, resistance to AMPs can still be observed if bacteria are exposed to AMPs for sufficient periods of time (Boman, 2003), highlighting that antibiotic resistance is a problem that investigators will have to solve over and over again. Thus, high-throughput methods for the rapid discovery and design of novel AMPs would be instrumental in the fight against superbugs (Lin et al., 2022).

[0115] Recently, a number of in silico AMP prediction tools have been developed (Li C. et al., 2022) (Veltri et al., 2018) (Meher et al., 2017) (Xiao et al., 2013) to reduce the labor and costs associated with large-scale wet lab screening for AMP discovery. State-of-the-art AMP prediction tools include AMPlify, (Li C. et al., 2022) AMP Scanner Vr.2 (Veltri et al., 2018), iAMPpred (Meher et al., 2017), and iAMP-2L (Xiao et al., 2013). Each of these tools utilizes machine learning methods, with AMPlify outperforming the latter three tools by adapting a deep learning model with attention mechanisms (Li C. et al., 2022) (Yang Z., et al., 2016) (Vaswani et al., 2017). These in silico tools have successfully been applied in identifying novel, naturally-occurring AMPs from genomic or transcriptomic resources (Lin etal., 2022) (Li C. et al., 2022) (Richter et al., 2022). However, the discovery of naturally- occurring AMPs requires organism sources, from tissue samples for direct wet-lab screening to sequencing data for in silico mining. Even though in silico mining methods are high- throughput, they require massive amounts of upstream work for careful data preparation, which further limits the pace of development and the number of novel AMPs that can be discovered.

[0116] Furthermore, the potential sequence space of amino acids is combinatorially large, allowing for peptide sequences that may not exist in nature but with desirable antimicrobial properties. Traditional approaches for AMP design include 1) modification of known AMP sequences to generate their congeners, fragments, or hybrids; 2) minimalist approaches by which AMPs are designed de novo purely based on structural requirements (e.g. amphipathic alpha-helical structures) but with limited types (e.g. physicochemical properties) of residues used; 3) creating sequence templates by comparing structurally homologous fragments from known AMPs for conserved patterns in terms of residue types and using the templates as guides for novel AMP design; and 4) utilizing peptide libraries (Huan et al., 2020)(Tossi, 2011). All of these methods could be efficient if combined with high-throughput screening methods, though currently they require prior AMP work expertise for more accurate designs.

[0117] Recently, a series of deep learning based models have been proposed for the automatic de novo design of AMP sequences (Nagarajan et al, 2018) (Dean et al., 2021) (Das et al., 2021) (Szymczak et al., 2022) (Gupta et al., 2019) (Tues et al., 2020) (Van Oort et al., 2021). They make it possible for the users to sample novel AMP sequences directly from the models, without any artificial design. Deep learning, a class of machine learning models, has outperformed traditional methods in many bioinformatics tasks with its strong ability to capture high-level features of the training data (Li Y. et al., 2019). Popular sequence generation models include recurrent neural network (RNN) language models (Mikolov et al., 2010), variational autoencoders (VAE), and generative adversarial networks (GAN). Nagarajan et al. developed a long short-term memory (LSTM) (Nagarajan et al, 2018) RNN language model, and embedded it into a framework with multiple filtering steps for the generation of novel AMPs with strong antibacterial activity (Nagarajan et al, 2018). Dean et al. proposed a VAE based AMP generation framework, named PepVAE, for generation of highly active AMPs (Dean et al., 2021). Das et al. further adapted VAE and introduced CLaSS for controlled AMP sequence generation with attributes of interest (Das et al., 2021). HydrAMP is another VAE based model (Szymczak et al., 2022), which incorporates two pre-trained classifies monitoring the quality of the generated peptides during training, improving upon a conditional VAE (cVAE). Gupta et al. 2019 proposed Feedback GAN for generating DNA sequences encoding proteins with optimized properties, and applied it to AMP sequence generation as an example. Tues et al. (2020) adapted an activity-aware LeakGAN to generate highly active AMPs, while Van Oort et al. (2021) introduced AMPGAN v.2 based on a bidirectional conditional GAN (BiCGAN) to generate AMP sequences of different types and properties. The flurry of activities represented by these methods explore expertise-free approaches and illustrate a strong interest in the field for de novo AMP design.

[0118] In this work, AMPd-Up is introduced, a novel AMP sequence generation tool that implements a standard RNN language model (Mikolov et al., 2010). AMPd-Up samples candidate AMP sequences from multiple model instances trained with different random initializations. For de novo AMP sequence generation, the RNN language model learns the “grammar” - the arrangement of the amino acids - of the training AMP sequences and estimates the probabilities of amino acid occurrence at each position recurrently starting from the N-terminus. Thus, the model generates a putative AMP sequence, residue by residue, based on the probability distribution estimated at each residue position (or each time step of the process). It is expected that different model instances would capture the complicated underlying features of AMP sequences from slightly different aspects, thus exploring various localities in the state space represented by a rich repertoire of natural AMPs. With this approach, 40 AMPs were generated that have never been reported in public databases but were proven to be active against laboratory strains of Escherichia coli and / or Staphylococcus aureus. The generated AMP sequences were limited to only have standard amino acids with a maximum length of 50 amino acids (aa), reflecting the fact that most documented AMPs are relatively short (Zhang and Gallo, 2016). This validation work primarily focused on AMPs with direct antibacterial activity, a major function of most known AMPs. These results illustrate the power of AMPd-Up in contributing to an expanding arsenal of synthetic antimicrobial agents.

[0119] 1.2. Materials and Methods

[0120] Training set. To get the RNN language model well trained, a curated set of known AMP sequences are required to comprise the training set. All antibacterial peptide sequences were downloaded from the Antimicrobial Peptide Database (APD3, aps. unmc.edu / AP) on March 20, 2019, a manually curated and annotated database for AMPs. This set of sequences contained 2,571 AMP records with antibacterial activity, 2,276 of which were < 50 aa in length. After removing duplicates and sequences with non-standardamino acids, there remained a non-redundant set of 2,253 antibacterial sequences < 50 aa in length, forming the training set for the RNN language model.

[0121] Model architecture and implementation. The implementation of the RNN language model was adapted from the PyTorch online tutorial by Sean Robertson (pytorch.org / tutorials / intermediate / char_rnn_generation_tutorial.html, accessed March 15, 2021), with PyTorch library 1.7.1 in Python 3.6.7. During the training process, cross-entropy was used as the loss function, and stochastic gradient descent (Robbins et al., 1951) was applied to optimize the model weights. A dropout technique was adopted to prevent overfitting. The hyperparameters, which cannot be learnt directly from training, were tuned through stratified 5-fold cross-validation on the training set. The set of hyperparameters for model architecture and training settings with the lowest average cross-validation loss was determined to be the optimal one to train the final model.

[0122] Figure 1 shows the architecture of the RNN language model, represented as a chain of repeating RNN cells. Given the first N-terminal amino acid, the RNN language model generates a peptide sequence residue by residue until reaching the end-of-sequence (EOS) signal. In this specific task of AMP sequence generation, the maximum length was set to be 50 and only the 20 standard amino acids are considered. Amino acids, together with the EOS signal, are encoded as twenty-one distinct one-hot vectors, with xte R31representing the t-th residue of a generated sequence. In this task, a time step t is defined as the process of an RNN cell predicting the (t 4- i)-th residue xM1of a sequence. At each time step t of the generation process, the RNN cell takes the hidden state ht-1from the previous time step and the predicted amino acid for the t-th residue xtas input, and outputs a set of probabilities pfof amino acid and EOS occurrence at the next position, from which xMican be sampled. The hidden state ht E RSftand probability vectorR31at each time step are calculates as:

[0123] (formula I)

[0124] (formula II)

[0125] and ^ £ R31^ft 31>are weight matrices, and bA£ RSft, bee R31, and bp£ R31are bias vectors. Here, denotes theconcatenation of two vectors vtand va, and the softmax function ensures that the probabilities sum up to 1. It was noted that the initial hidden state hflis set to a zero vector. It was found that the best tunedto be 128. A dropout rate of 0.1 was applied before the softmax function during training, and the training process was conducted with 100,000 iterations and a learning rate of 0.0005.

[0126] Predictions can be made by sampling from the output probabilities of the RNN cells. Sequence generation process stops if an EOS signal is predicted or if the maximum length is reached without EOS signal predicted. Sequences generated in the former case are annotated as “complete”, while those in the latter case as “incomplete”. The geometric mean of probabilities of all predicted symbols in a sequence (EOS signal included if the sequence is “complete”) was defined as the “AMPd-Up score”, measuring the confidence of the RNN language model in generating the sequence. In AMPd-Up, the model is trained multiple times with different random initializations, yielding multiple model instances.

[0127] Given one of the 20 possible starting amino acids, the symbol with the highest probability estimated at each time step is taken as the next amino acid prediction (including the EOS signal), resulting in a maximum of 20 candidate AMP sequences generated by a single model instance. In a practical use case, the model will be trained k times and the users would get a candidate AMP list of up to 2&k sequences. Assuming a non-convex loss function like most deep learning tasks, different initializations may result in different trained models (Fort et al., 2019), allowing different model instances of AMPd-Up to capture slightly different aspects of the complex but unknown features of AMPs.

[0128] Model evaluation. In order to measure the performance of AMPd-Up, the predictions from three state-of-the-art in silico AMP prediction tools were used: AMPlify (Li C. et al., 2022), AMP Scanner Vr.2 (Veltri et al., 2018), and iAMPpred (Meher et al., 2017), as a proxy for AMP sequence generation accuracy. Here, the AMP sequence generation accuracy measured by a selected prediction tool was calculated based on the proportion of peptide sequences predicted as AMPs among a generated sequence set. A default setting of balanced model was chosen for AMPlify (v1.1.0) as described in a data note (Li C. et al., 2023), while the “original production model” was chosen for AMP Scanner Vr.2 on its online server (Veltri et al., 2018). Predictions by iAMPpred were obtained through its online server with its trained model as described in the publication (Meher et al., 2017).

[0129] AMPd-Up was compared with three other AMP sequence generation methods with publicly available models or generated sequences: the LSTM language model (Nagarajan et al, 2018), AMPGAN v2 (Van Oort et al., 2021), and HydrAMP (Szymczak et al., 2022). For each method, a total of 2,000 sequences were generated for comparison in five batches without any further downstream filtering. This resulted in five generated sequence sets of 400 sequences for each method. Sequences for the LSTM language model was sampled from the dataset the authors provided, while those for HydrAMP were obtained through their online server (hydramp.mimuw.edu.pl, accessed November 7, 2022). While all other methods focus on the generation of antibacterial peptides, AMPGAN v2 additionally allows for generating AMPs of other function types (e.g. antifungal, antiviral) and the generated sequences are annotated with their predicted functions in the results (Van Oort et al., 2021). For a fairer comparison, only AMPs targeting bacteria were selected for AMPGAN v2. For each AMP sequence generation method measured by each AMP prediction tool, the average accuracy value of the five generated sets was reported, along with the corresponding standard deviation value.

[0130] In addition to the sequence generation accuracy, the quality of sequences generated was evaluated by AMPd-Up based on their physicochemical properties as well as their sequence similarities to the training set and all publicly available known AMPs.

[0131] The properties that cause a peptide sequence to have antimicrobial activity is complex and the mechamisms are still not well understood (Teimouri et al., 2021). Considering the fact that most known AMPs share common characteristics of short lengths and net positive charges (Zhang and Gallo, 2016), the focus was placed on these two important and easy-to-calculate physicochemical properties.

[0132] Moreover, sequence similarities of the generated sequences to the training set were calculated to evaluate whether the model instances capture high-level features of AMPs rather than only genarating the same or highly similar sequences to the training set. A similar comparison between the generated sequences and all publicly available known AMPs was done to evaluate the novelty of the generated sequences to those known AMP sequences. The training AMPs are antibacterial, while the known AMP sequence set additionally includes those targeting microbes other than bacteria. The known AMP sequence set comprises 4,538 distinct sequences that were downloaded from APD3 (see Wang et al.) and Database of Anuran Defense Peptides (DADP, split4.pmfst.hr / dadp) (Novkovic et al., 2012) on July 11 , 2022 and December 6, 2018, respectively. The similaritybetween two sequences was calculated as is the edit distanceare lengths of the sequences. The similarity of a sequence to a set of sequences was defined as the maximum of all similarity values calculated between that sequence and the sequences in the target set for comparison (i.e. the similarity of that sequence to its most similar sequence in the target set).

[0133] Selecting AMPs for validation. Table 2 shows AMP sequences generated by AMPd-Up that have been prioritized for synthesis. Lists A, B, and C include 38, 4, and 16 sequences, respectively. All sequences in Lists A and C were predicted as AMPs by AMPlify (4, 5), while those in List B were predicted as non-AMPs. Sequences were sampled from all candidate peptide sequences generated by 1 ,000 model instances, with incomplete sequences removed. Sequences in Lists A and B were sampled through AMPd-Up scores, while List C comprises a set of sequences that appear with high frequency (> 40 in sequence counts) in all candidate peptides. The numbering of peptide names for Lists A and B was byAMPlify score, while List C was by sequence count, both in descending order.

[0134] To demonstrate the utility of the tool, the model was trained 1 ,000 times, yielding 1,000 model instances and 20,000 sequences, 14,188 of which were complete, and 8,737 of the complete sequences were distinct. The “count” of a sequence was defined to be the number of times it appears in the entire generated set. Short sequences were filtered for, with lengths < 35 and obatained 7,434 peptide sequences, since shorter peptides are more cost-effective for synthesis (Lin et al., 2022). 58 of these peptides were selected using different strategies (forming Lists A, B, and C), and their bioactivity validated through in vitro experiments (Table 2).

[0135] The peptides comprising Lists A and B were chosen following a strategy that stratifies the AMP-Up score range of 7,434 sequences into same-length score invervals. For n intervals, each interval can be written as a range fromwith k = i,z n and a,b being the mimimum and maximum AMPd-Up scores investigated in the generated set. In the present case, « = ai462 and & = Q.3&79. All intervals are of left-open and right-closed, except the first one (ft = i) that is closed. If multiple model instances generated the same sequence, the AMPd-Up score from the first model that generated this sequence was used for stratification. Peptides for List A were sampled by splitting the AMPd- Up score range of [0.1462, 0.3579] into 40 intervals, and the sequence with top AMPd-Up score within each interval was chosen. List B peptides were chosen by splitting the same AMPd-Up score range into 5 intervals, and selecting one predicted non-AMP (as assessed by AMPlify) in each interval with the highest count, or with top AMPd-Up score if all sequences have the same count, within each interval was chosen. Some intervals did nothave any sequences, resulting in 38 sequences in List A and 4 sequences in List B. Additionally, 16 more peptide sequences that appeared with high frequency (> 40 in sequence counts) in the generated set were selected to be List C. All sequences in Lists A and C were predicted as AMPs by AMPlify. Table 2 presents the sequence similarity of each sequence to the known AMPs, showing the novelty of those sequences to the known AMP sequences.

[0136] Antimicrobial susceptibility testing (AST). The antimicrobial activity of the selected peptides was measured in the laboratory by broth microdilution assays, to determine the minimum inhibitory and minimum bactericidal concentrations (MICs and MBCs, respectively), as outlined by the Clinical and Laboratory Standards Institute (CLSI), with some adaptations for testing cationic AMPs as described previously Wiegand et al., 2008. Laboratory isolates of E. coli 25922 and S. aureus 29213 were purchased from the American Type Culture Collection (ATCC; Manassas, VA, USA) and were used to validate the 58 selected AMPs. Bacteria from frozen stocks were streaked onto nonselective Columbia blood agar with 5% sheep blood (Oxoid) and incubated for 18-24 h at 37°C. The following day, 2-4 colonies were streaked onto a new agar plate and incubated for 18-24 h at 37°C to ensure uniform colony health prior to the assay. The standardized bacterial inoculum was prepared by suspending isolated colonies in Mueller-Hinton Broth (MHB; Sigma-Aldrich, St. Louis. MO, USA). The suspension was adjusted to an optical density of 0.08-0.1 at 600nm, equivalent to a 0.5 McFarland standard of approximately 1-2 x 108CFU / mL (CFU: colony forming units). The inoculum was then diluted 1:250 to achieve a final concentration of 5±3 x 105CFU / mL. The target bacterial density was confirmed by routinely examing the total viability counts from the final inoculum.

[0137] Candidate AMPs were purchased from and synthesized by GenScript (Piscataway, NJ, USA). These were received in lyophilized format and stored at -20°C, and were suspended in sterile ultrapure water prior to testing. A two-fold serial dilution of 1280 down to 2.5 pg / mL was prepared in sterile 96-well polypropylene microtitre plates (Greiner Bio-One #650261, Kremsmunster, Austria) before the addition of 100 pL of the standardized bacterial inoculum, providing a final AMP testing range of 128 down to 0.25 pg / mL. The MIC values were reported as the lowest peptide concentration where no visible bacterial growth were observed following a 20-24 h incubation at 37°C. For determination of MBC, well contents of the MIC and the two adjacent wells containing the two- and four-fold MIC were plated onto nonselective nutrient agar. The concentration in which 99.9% of the inoculum were killed after incubation for 24 hours at 37°C was reported as the MBC.

[0138] In the present tests, a known AMP Ranatuerin-4 (Goraya et al., 1998) from the American bullfrog and an in-house peptide [TKPKG]s (OT15, SEQ ID NO:8074) were used as the positive and negative control peptides, respectively. OT15 was truncated and derived from a negative control peptide [TKPKG]4 (OT20, SEQ ID NO:8075), not antimicrobial but with similar characteristics to AMPs, used in previous studies (Horvati et al., 2017).

[0139] Hemolysis assay. The toxicity to red blood cells (RBCs) of the select peptides was evaluated by hemolysis experiments. Whole blood from healthy donor pigs was purchased from Lampire Biological Laboratories (Pipersville, PA, USA). RBCs were washed and isolated by centrifugation, using Roswell Park Memorial Institute medium (RPMI) (Life Technologies, Grand Island, NY, USA; Gibco cat# 11835-030). AMPs were suspended and serially diluted from 1280 down to 10 pg / mL using RPMI in a 96-well polypropylene microtitre plate, and then they were combined with 100 pL of the 1% RBC solution. This resulted in a final AMP testing range of 128 down to 1 pg / mL. Following an incubation at 37°C for 30-45 minutes, plates were centrifuged and 1 / 2 volume from each supernatant was transferred to a new 96-well plate. The absorbance of the wells was measured at 415 nm; the AMP concentration that lysed > 50% of the RBCs (HCso) was used to report the hemolytic activity. Absorbance readings from wells containing RBCs treated with 11 pL of a 2% Triton-X100 solution or RPMI (AMP solvent-only) were used to define 100% and 0% hemolysis, respectively. All centrifugation steps were performed at 500* g for five minutes in an Allegra- 6R centrifuge (Beckman Coulter, CA, USA).

[0140] 1.3. Results

[0141] Performance comparison with state-of-the-art methods. The performance of AMPd-Up was measured by assessing the generated sequences using three state-of-the- art AMP prediction tools: AMPlify (Li C. et al., 2022), AMP Scanner Vr.2 (Veltri et al., 2018), and iAMPpred (Meher et al., 2017). The sequence generation accuracy values, reported as the percentages of sequences predicted as AMPs evaluated by each AMP prediction tool, are reported in Table 3, with three other AMP sequence generation methods listed for comparison. Although none of the in silico prediction tools are perfect in identifying AMPs, their reported performance (Li C. et al., 2022) (Veltri et al., 2018) (Meher et al., 2017) would be suitable for evaluating the AMP sequence generation methods. Details of how the accuracy values were calculated can be found in the Materials And Methods section.

[0142] Table 3 shows a performance comparison of different antimicrobial peptide (AMP) sequence generation methods. Different methods were evaluated using three in silico AMP prediction tools: AMPlify (Li C. et al., 2022), AMP Scanner Vr.2 (Veltri et al., 2018), and iAMPpred (Meher et al., 2017), based on sequences generated by each of the methods. The AMP sequence generation accuracy measured by a selected prediction tool was defined as the proportion of peptide sequences predicted as AMPs among a generated sequence set. For each sequence generation method, five sets of sequences were generated, with 400 in each set. For each AMP sequence generation method, an average accuracy value of the five generated sets was reported when measured by a specific AMP prediction tool, along with the corresponding standard deviation value.

[0143] As measured by AMPlify, AMPd-Up obtains the highest accuracy with 95.50% of the generated sequences predicted as AMPs on average, which outperforms the best comparator AMPGAN v2 by 4.60%, followed by HydrAMP (by 8.00%) and then the LSTM language model (by 10.65%). When evaluating using AMP Scanner Vr.2 and iAMPpred, AMPd-Up generates AMP sequences with accuracies of 100.00% and 99.30%, surpassing the best comparator HydrAMP by 5.40% and 1.60%, respectively. Although the rankings of the AMP sequence generation methods evaluated by the three AMP prediction tools are slightly different from each other, AMPd-Up always performs the best compared with its comparators.

[0144] Quality analyses of de novo generated sequences. Besides using the outputs of in silico AMP prediction tools as proxy for performance, the length and net charge distributions of the generated sequences were analysed, as well as their sequence similarity levels to the training set and all known AMP sequences.

[0145] Figure 1 shows the architecture of the recurrent neural network (RNN) language model. Given a starting amino acid, the RNN language model predicts the next amino acids residue by residue until reaching the end-of-sequence (EOS) signal. Amino acids, including the EOS signal, are one-hot encoded. The output of RNN at each time step is a probability vector of amino acid and EOS occurrence at the next position, to which sampling strategies can be applied.

[0146] Short lengths and net positive charges are common characteristics for most previously discovered AMPs (Zhang and Gallo, 2016), therefore many AMP studies investigate these key properties (Gagnon et al., 2017). Shorter peptides are also cheaper to sythesize (Lin et al., 2022), making translating shorter sequences for clinical application potentially more cost-effective. Further, the net positive charges of cationic AMPs are responsible for the electrostatic interation with the negatively charged bacterial membranes or cell walls (Zhang and Gallo, 2016), with studies illustrating that the antimicrobial activity of some AMPs can be improved by increasing their net charges (Zelezetsky et al., 2006).

[0147] Figure 2 illustrates length and net charge distributions of the generated sequences. Length and net charge distributions were calculated based on 2,000 sequences generated by AMPd-Up, along with training sequences for comparison. 1,484 of the 2,000 generated sequences in “complete” status were chosen for an additional comparison. Mean (p) and standard deviation (o) of each distribution are as follows: training sequences (length: p = 26.21 aa, o = 10.34 aa; net charge: p= 3.30, o = 2.74), all generated sequences (length: p = 28.90 aa, o = 15.07 aa; net charge: p = 9.08, o = 7.33), and complete generated sequences (length: p = 21.56 aa, o = 9.87 aa; net charge: p = 6.45, o = 4.73).

[0148] The top section of Figure 2 compares the length distributions of the generated sequences with those constituting the training set. The model may fail to reach the EOS signals when generating some sequences (referred to as incomplete sequences), and thus an additional comparison was made of the generated sequence set with those incomplete sequences removed. The average generated sequence length is 28.90 aa, but is reduced to 21.56 aa after incomplete sequences were removed. The incomplete sequences are 50 aa, by default. The complete sequences were 4.65 aa shorter than the training sequences, on average. The bottom section of Figure 2 shows a similar comparision for net charge distributions. The average generated sequence net charge is 9.08, but is reduced to 6.45 after incomplete sequence removal. However, the net charge of the complete sequences is still 3.15 greater than the training sequences, on average.

[0149] Figure 3 shows sequence similarity distribution of the generated sequences to the training set. The sequence similarity distribution, with a mean of 49.97% and a standard deviation of 9.83%, was calculated based on the 2,000 sequences generated by AMPd-llp. The sequence similarity of each generated sequence to the training set was considered as the similarity of that sequence to its most similar sequence in the training set, based on which the distribution was plotted.

[0150] The sequence similarity of each generated sequence to the training set, which includes all available antibacterial peptides, was calculated for analysis, with details described in the Materials and Methods section. Figure 3 shows the sequence similarity distribution of the generated sequences to the training set, with a peak between 50.00% and 55.00%. The generated sequences possess a similarity level of 49.97% compared with the training sequences, on average, indicating that the present sequence generation method generates novel AMPs less similar to the training sequences. This implies that AMP-lip may be capturing high-level features of AMPs, rather than only memorizing sequence-level information during training.

[0151] Figure 4 shows sequence similarity distribution of the generated sequences to known AMPs. The known AMP sequence set comprises 4,538 distinct sequences downloaded from APD3 (aps. unmc.edu / AP see Wang et al., 2016) and DADP (split4.pmfst.hr / dadp see Novkovic et al., 2012) databases. The sequence similarity distribution, with a mean of 51.03% and a standard deviation of 9.38%, was calculated based on the 2,000 sequences generated by AMPd-llp. The sequence similarity of each generated sequence to known AMPs was considered as the similarity of that sequence to its most similar known AMP sequence, based on which the distribution was plotted.

[0152] Figure 5 shows distributions of pairwise sequence similarities between different sequence sets in Example 1. The pairwise sequence similarities between two different sets of sequences were calculated as the similarities of all sequence pairs between the two sets, while the pairwise sequence similarities of the same set of sequences (i.e. sequence diversity measurements) were defined as the similarities of sequences to each other in the set. The intra-model sequence similarities were calculated as the similarities of sequences generated by the same model instance to each other, while inter-model sequence similarities were calculated as pairwise sequence similarities between sets of sequences generated by different model instances. A set of 2,000 random sequences matching the length distribution of the generated sequences were added for comparison, in addition to the training and generated sequence sets. Mean (p) and standard deviation (o) values of eachdistribution are as follows: Random vs. Random (p = 13.71%, o = 5.09%), Random vs. Generated (p = 10.34%, o = 5.28%), Random vs. Training (p = 13.76%, o = 4.91%), Training vs. Training (p = 18.06%, o = 7.92%), Generated vs. Training (p = 18.80%, o = 8.94%), Generated vs. Generated (p = 33.61%, o = 16.18%), Intra-model (p = 39.14%, o = 19.79%), and Inter-model (p = 33.56%, o = 16.14%). Two-sided Kolmogorov-Smirnov tests reveal that the difference between any two of the distributions is significant (p < 0.0018), except that between Generated vs. Generated and Inter-model (p = 0.0519). Notably, 0.0018 is an adjusted alpha level calculated with Sidak correction (Sidak, 1967) from a family-wise alpha level of 0.05 for the multiple comparisons.

[0153] An additional test on the sequence similarity of each generated sequence to all available known AMPs was done, with an average sequence similarity level of 51.03%, indicating the novelty of the generated sequences as compared with AMPs that have already be discovered or designed (Figure 4). To supplement the sequence similarity analysis, the pairwise sequence similarities between different sequence sets was visualized (Figure 5). A lower generated sequence similarity level between different model instances (33.56%) than within the same model instance (39.14%) indicates that different model instances tend to capture features of AMPs from slightly differently aspects. The AMP sequences revealed as described herein will to add diversity to the current AMP sequence databases.

[0154] In vitro validation results. A total of 58 sequences generated by 1,000 AMPd-Up model instances were selected for in vitro validation and bioactivity assessment. Peptides were ordered into three lists: List A (DeNo1001 to DeNo1038) and List B (DeNo1039 to DeNo1042) were sampled through AMPd-Up scores as described in the Materials And Methods section, and a few more sequences that appear with high frequency (> 40 in sequence counts) in the generated set were selected to make List C (DeNo1043 to DeNo1058). Table 2 summarizes the sequence specifications of the 58 selected AMPs. All sequences in Lists A and C were predicted as AMPs by AMPlify, while all sequences in List B were predicted as non-AMPs.

[0155] Figure 6 shows in vitro validation results of the 58 selected AMPs. (Panel a) Antimicrobial and hemolytic activities of the 40 peptides that were active against at least one bacterial strain of Escherichia coli ATCC 25922 and Staphylococcus aureus ATCC 29213. Antimicrobial and hemolytic activities were measured by minimum inhibitory concentration (MIC) and concentration that lyses 50% (HC50) of the red blood cells (RBCs), respectively. HC50 was determined using porcine RBCs. Data was obtained to determine the lowest effective peptide concentration range (pg / mL) observed in three independent experimentsperformed in duplicate, with one maximum data point and one minimum data point dropped for each measurement. Three sections from left to right correspond to peptides with observable antimicrobial activity from List A (n = 28), List B (n = 1), and List C (n = 11), respectively. Full data is shown in Table 4. Activity of the peptides was split into four levels: high (< 4 pg / mL), moderate (8-16 pg / mL), low (32-128 pg / mL), and without observable activity (> 128 pg / mL), as separated by different background shades in the plot. (Panel b) Stacked bar chart showing proportions of peptides that displayed antimicrobial activity with different sequence similarity levels to known AMPs from APD3 (Wang et al., 2016) and DADP (Novkovic et al., 2012) databases. All similarity ranges are of left-open and right- closed, and the sequence similarity of each candidate peptide to known AMPs was considered as the sequence similarity of that sequence to its most similar known AMP sequence. (Panel c) Visualization of antimicrobial activity of the 58 tested peptides regarding AMPlify scores and AMPd-Up scores. AMPd-Up scores of the same peptide sequences generated by multiple model instances were averaged. Peptides without any observable antimicrobial activity are presented as grey crosses, and the active peptides are presented in shaded dots. Dots with darker shading indicate stronger antimicrobial activity against Escherichia coli ATCC 25922, determined by the minimum MIC value of each peptide against the strain.

[0156] The 58 candidate peptides were tested against two bacterial isolates: the Gram-negative E. coli ATCC 25922, and the Gram-positive S. aureus ATCC 29213. Porcine RBCs were used to assess the hemolytic activity of the peptides. Out of the 58 peptides selected for in vitro validation, 40 peptides displayed antimicrobial activity against at least one bacterial strain tested. All 15 peptides that were observed active against S. aureus ATCC 29213 also showed antimicrobial activity against E. coli ATCC 25922. Figure 6, Panel A visualizes the antimicrobial and hemolytic activities (in MIC and HCso, respectively) of the 40 peptides, with the entire in vitro validation results of the 58 peptides shown in Table 4. For a better interpretation of the results, the activity of the tested peptides was split into four levels according to the MIC / HC50 ranges: high (< 4 pg / mL), moderate (8-16 pg / mL), low (32- 128 pg / mL), and without observable activity (> 128 pg / mL).

[0157] Table 4 shows the results of antimicrobial susceptibility testing and hemolysis experiments of the 58 selected peptides in vitro. Peptides were tested for their antimicrobial activity against Escherichia coli ATCC 25922 and Staphylococcus aureus ATCC 29213 for their minimum inhibitory concentration (MIC) and minimum bactericidal concentration (MBC) values. Porcine red blood cells (RBCs) were used to test the hemolytic activity of theselected peptides for their hemolytic concentration (HC50) values. Data is presented as the lowest effective peptide concentration range (pg / mL) observed in three independent experiments performed in duplicate, with one maximum data point and one minimum data point dropped for each measurement. Ranateurin-4 and OT15 (SEQ ID NO:8074) were used as the positive and negative control peptides, respectively.

[0158] Among the 38 List A peptides tested, a total of 28 peptides displayed antimicrobial activity in the tests, 12 of which were active against both of the strains tested. Nine of the List A peptides were highly active against E. coli ATCC 25922, with DeNo1018 being the most active with an MIC of 1-2 pg / mL. All of these nine peptides were also active against S. aureus ATCC 29213. Four of the nine peptides were highly active against S. aureus ATCC 29213 (MIC = 2-4 pg / mL for DeNo1016 and DeNo1017; MIC = 4 pg / mL for DeNo1007 and DeNo1022), with one (DeNo1018) moderately active (MIC = 8 pg / mL). Six peptides from List A were moderately antibacterial against E. coli ATCC 25922, and another two showed low to moderate activity against the strain. Three of these eight peptides displayed some antimicrobial activity against S. aureus ATCC 29213, one of which (DeNo1031) was moderately active (MIC = 16 pg / mL) with the other two (DeNo1021 andDeNo1026) showed low (MIC = 32-64 pg / mL) or minimal activity (MIC > 128 pg / mL), respectively. Among all 28 List A peptides with proven antimicrobial activity, three were minimally hemolytic (HCso s 128 pg / mL) and 17 showed no hemolytic activity (HCso > 128 pg / mL) in the tests. DeNo1007 was the only AMP with high antimicrobial activity against both of the bacterial strains tested (MIC = 4 pg / mL) and no observable hemolytic activity (HCso > 128 pg / mL).

[0159] Among the four peptides from List B tested, only DeNo1040 displayed some low-level activity against the two bacterial strains tested. Specifically, this peptide inhibited the growth of E. coli ATCC 25922 and S. aureus ATCC 29213 providing MICs of 64-128 pg / mL and 64 pg / mL, respectively. DeNo1040 also did not show any hemolytic activity in the tests (HCso > 128 pg / mL). Peptides in List B were predicted as non-AMPs by AMPlify.

[0160] Among the 16 List C peptides tested, a total of 11 peptides showed antimicrobial activity against E. coli ATCC 25922 in the tests, with two of them additionally active against S. aureus ATCC 29213. DeNo1049 displayed moderate to high activity against E. coli ATCC 25922 (MIC = 4-8 pg / mL), which was the strongest in List C. DeNo1057 was moderately antibacterial against E. coli ATCC 25922 (MIC = 8-16 pg / mL), followed by DeNo1051 (MIC = 16-32 pg / mL) and DeNo1046 (MIC = 16-64 pg / mL). DeNo1057 and DeNo1046 were the only two List C peptides with antibacterial activity against S. aureus ATCC 29213, though with low activity (MIC = 32 pg / mL and 128 pg / mL, respectively). None of the List C peptides displayed hemolytic activity in the tests (HCso > 128 pg / mL).

[0161] Among the peptides that did not show any antimicrobial activity against the bacterial strains tested, most of them were also not hemolytic to the porcine RBCs except DeNo1008 (HCso = 16-32 pg / mL) and DeNo1039 (HCso = 32-64 pg / mL).

[0162] Figure 6, Panel b presents the proportions of peptides that were active against at least one of the bacterial strains tested under different sequence similarity levels to the known AMPs from APD3 (Wang et al., 2016) and DADP (Novkovic et al., 2012) databases. All six peptides between sequence similarities of 70% and 90% to known AMPs showed antimicrobial activity in the tests. The largest proportion of the tested peptides fall between sequence similarities of 60% and 70% to known AMPs, with nine out of 20 sequences displaying antimicrobial activity. Interestingly, lower similarity intervals of 50%- 60% and 40%-50% possess relatively high proportions of antimicrobially active peptides with rates of 75.00% (12 / 16) and 81.25% (13 / 16), respectively. More than half (62.50%) of the peptides with antimicrobial activity from these experiments fall into these intervals, implying there is much to be explored in the sequence space for novel AMPs.

[0163] Figure 6, Panel c visualizes the distribution of the 58 tested AMPs with regard to AMPlify scores and AMPd-Up scores. AMPlify score, ranging from 0 to 80, is a prediction score reported by AMPlify, which is a log transformation of the AMPlify probability scoreAMPd-Up score, reported by AMPd-Up, ranges from0 to 1 and is a measure of the confidence level of the model when generating the sequence. Considering the fact that multiple model instances may generate the same sequence but with different AMPd-Up scores, the average was taken in the visualization for a more comprehensive analysis. As can been seen from Figure 6, Panel c, most of the peptides that did not show any antimicrobial activity in the tests are located at the bottom left of the figure, suggesting that it is ideal to prioritize generative sequences with both high AMPlify and AMPd-Up scores for in vitro validation assays.

[0164] DISCUSSION

[0165] In the presented work, AMPd-Up was used as a tool for de novo AMP sequence generation. AMPd-Up adopts an RNN language model, sampling from multiple model instances trained with different random initializations. Although the architecture of the model is simple compared with existing methods, it was shown that simple models like AMPd-Up can work well if properly trained, as illustrated. Moreover, the sequences generated by AMPd-Up are of high novelty compared with existing AMP sequences in public databases, demonstrating the ability of the model to learn high-level AMP features. While AMPd-Up shows great promise and favorable performance, the size of its training set is still relatively small (2,253 sequences) compared with that of many traditional deep learning tasks for broader sequence data analysis, such as sentiment analysis or machine translation, which both typically enjoy hundreds of thousands to millions of data points for training available through public databases (Khurana et al., 2023). This is a critical problem that all current in silico AMP prediction and generation tools are facing. This limitation will gradually become resolved as more AMPs are being discovered and validated, leading to further improvement in de novo AMP sequence generation tools like AMPd-Up.

[0166] Although the AMPd-Up-generated AMPs have a considerable level of sequence diversity (Figure 5), it was still noticed that some patterns at the sequence level. Analyzing the generated sequences, it was observed that “LLKK” and “LKKL” were the two most frequently occurring 4-mer motifs, appearing in 44.09% and 40.73% of the 20,000 sequences, respectively. Previous studies have shown that synthetic amphipathic alphahelical peptides made up of repeat units [LLKK]nor [LKKL]nhave antimicrobial properties(Wiradharma et al., 2011) (Khara et al., 2017), which explain these findings to some extent. In fact, it is suggested that repeats of 4-mer units such as these are responsible for the formation of cationic amphipathic alpha-helical structures, a key initiating step to the bioactivity and membrane-disrupting properties of many AMPs (Wiradharma et al., 2011) (Khara et al., 2017).

[0167] Among the 58 novel AMP sequences generated by AMPd-Up, 40 showed antimicrobial activity in the tests, 15 of which were broadly antibacterial against both the Gram-positive and Gram-negative isolates. Promisingly, one of the most active peptides, DeNo1007, not only possessed high antimicrobial activity against the two bacterial strains tested, but also displayed no observable hemolytic activity. The AMP candidates generated by AMPd-Up should increase the diversity of known peptide-derived antibiotics, currently populated by mostly naturally-occurring sequences, and to augment the candidate set of potential alternatives to conventional antibiotics. Although some of the identified AMPs did not show any antimicrobial activity against the two bacterial strains tested in vitro, they may still be active against other bacterial species and / or possess some immunomodulatory functions. Also, the structures of some AMPs may vary based on their microenvironment (Candido et al., 2019). Further experimentation could be done to test candidate sequences on a wider panel of bacterial species, or to interrogate in vivo biological interactions.

[0168] Results from work like ours also have broader potential impact. Resistance to last-line peptide-based therapeutics, such as colistin and other polymyxins, is increasingly being reported (Aghapour et al., 2019). Concerningly, this is sometimes presented with cross-resistance to multiple AMPs (Fleitas et al., 2016), highlighting the need for multiple and diverse classes of peptide-based antimicrobials. De novo AMP sequence generation allows for a rational solution to this problem, as one would theoretically expect that pathogens would be naive to many of the diverse cfe-novo-generated AMPs. Even though there may be natural AMPs similar to some of the cte-novo-generated ones, the vast sequence space of amino acids (e.g. 1020or one hundred quintillion for a 10-residue peptide sequence) virtually ensures that there would be a practically infinite number of them out there that are “new” to most common pathogens. Thus, it is expected that high-throughput in silico AMP sequence design tools like AMPd-Up to play a vital role in the fight against antibiotic resistance and the imminent rise of anti bi otic- resista nt bacteria.

[0169] Example 2

[0170] Mining the UniProtKB / Swiss-Prot Database for Antimicrobial Peptides

[0171] Summary. The ever-growing global health threat of antibiotic resistance is compelling researchers to explore alternatives to conventional antibiotics. Antimicrobial peptides (AMPs), a class of short and often cationic biological molecules, is emerging as a promising solution to fill this need. Naturally-occurring AMPs are produced by all forms of life as part of the innate immune system. High-throughput bioinformatics tools have enabled fast and large-scale discovery of AMPs from genomic, transcriptomic, and proteomic resources of selected organisms. With over 200 million records and counting, public protein sequence databases represent a comprehensive compendium of sequences from a broad range of source organisms, and large-scale in silico probing of those databases for novel AMPs has never been done. In this study, it was proposed that an AMP mining workflow to predict novel AMPs from the UniProtKB / Swiss-Prot database using the AMP prediction tool AMPlify as the core module, with 8,008 novel AMPs (SEQ ID NO:59 - SEQ ID NO:8066) identified from all eukaryotic sequences in the database. Focusing on the practical use of AMPs as suitable antimicrobial agents with application to use in the poultry industry, 40 of those AMPs were prioritized based on structural similarities to known chicken AMPs. In the tests conducted, 13 of the 38 successfully synthesized peptides showed antimicrobial activity against Escherichia coli and / or Staphylococcus aureus.

[0172] 2.1. Introduction.

[0173] As a consequence of the worldwide overuse of antibiotics, the world is under the threat of entering a “post-antibiotic era” (Reardon, 2014). The decreasing effectiveness of conventional antibiotics is posing great challenges in the treatment of many infectious diseases (Reardon, 2014), with an estimated 1.27 million people dying due to antibiotic resistance in 2019 (Murray et al., 2022) Despite the extensive use of conventional antibiotics in clinical settings, antibiotics are also widely used in the agriculture industry (Laxminarayan et al., 2013) It has been reported that certain multidrug-resistant (MDR) bacteria can be transmitted between humans and other animals, which further exacerbates the problem (Laxminarayan et al., 2013). While the occurrence of antibiotic resistance is increasing, there has been a significant decline in the discovery of new antimicrobial agents starting from the 1990s (Koo et al., 2019; Terreni et al., 2021) Consequently, there is an urgent need for novel and effective substitutes for conventional antibiotics.

[0174] Antimicrobial peptides (AMPs), a family of short and often cationic peptides, are regarded as one promising substitute for conventional antibiotics (van der Does et al., 2019). Naturally-occurring AMPs are produced by all forms of life as part of the innate immunity (Zhang and Gallo, 2016), usually in an inactive precursor form with a signalpeptide, an acidic pro-sequence, and the bioactive mature peptide (Beckloff et al., 2008). The bioactive mature AMPs are released by proteolytic cleavage of their precursors (Zhang and Gallo, 2016). While the majority of known AMPs recorded in public AMP databases have been shown to possess antibacterial activity (Wang et al., 2016), many of them have also been reported with other types of antimicrobial activities, including antifungal (De Lucca et al., 1999) and antiviral (Klotman et al., 2006). Most AMPs exert their effects by directly interacting with bacterial membranes or cell walls, causing non-enzymatic disruption (Zhang and Gallo, 2016). Additionally, some eukaryotic AMPs also perform modulation of immune responses (Zhang and Gallo, 2016) (Nguyen et al., 2011). In comparison with conventional small-molecule antibiotics, which have specific functional or structural targets, the distinct modes of action of AMPs may hold an advantage in being able to better overcome bacterial resistance (Boman, 2003). Nevertheless, it is still possible to observe resistance to AMPs if bacteria are exposed to AMPs for extended periods of time (Boman, 2003), which highlights a pressing need to augment the peptide-based therapeutics arsenal.

[0175] Traditional approaches of discovering naturally-occurring AMPs through wet lab screening are time-consuming, labor-intensive and costly (Wu Q., 2019). In the past few decades, a series of machine learning based high-throughput in silico AMP prediction tools have been developed to overcome this problem, including iAMP-2L (Xiao et al., 2013), iAMPpred (Meher et al., 2017), AMP Scanner Vr.2 (Veltri et al., 2018), and AMPlify (Li C. et al., 2022) (Li C. et al., 2023). Among those tools, AMPlify, which uses a deep learning model with attention mechanisms (Vaswani et al., 2017) (Yang Z., et al., 2016), has become the state-of-the-art method for AMP prediction. It has also successfully identified over a hundred of lab-validated AMPs from amphibian and insect genomes and transcriptomes (Lin et al., 2022) (Li C. et al., 2022) (Richter et al., 2022).

[0176] Current large-scale AMP mining studies mostly focus on genomic (Li C. et al., 2022) (Amaral et al., 2012) (Prichula et al., 2021), transcriptomic (Lin et al., 2022) (Richter et al., 2022), or proteomic (Tomazou et al, 2019) resources. However, in most of these studies, specific organisms were chosen for the mining. AMP mining by large-scale scanning of public protein databases, such as UniProt (UniProt Consortium, 2019) (uniprot.org), in which the majority of the sequences have already been annotated with their functions, has never been done. However, conducting AMP mining using comprehensive protein databases provides another valuable way to unearth novel AMPs. First, these large protein databases allow researchers to discover novel AMPs from a vast number of available organisms, rather than biasing on popular source organisms such as amphibians (Li C. et al., 2022) , providingthe community with a broad source of AMPs. Some AMP prediction studies do mine protein sequences from databases like UniProt, but they are limited in scope, and only focus on a few selected organisms (Tomazou et al., 2019). Second, there are many uncharacterized sequences awaiting annotation in public databases like UniProt (UniProt Consortium, 2019), which serves as an untapped source of AMPs. Further, it would be of interest to uncover possible antimicrobial properties of characterized proteins or peptides, where only nonantimicrobial functions have been reported to date. Studies have found that some proteins or peptides with other biological functions exhibit antimicrobial properties (Burdukiewicz et al., 2020) (Khurshid et al., 2017). Histatins, a series of peptides found in saliva, fall in this category. Histatins harbor multiple functions beneficial to oral health besides antimicrobial activity, including oral hemostasis, development of acquired tooth pellicle, and assistance in bonding of some metal ions (Khurshid et al., 2017). Protein databases represent a valuable resource for novel AMP mining, which is demonstrated herein.

[0177] Here is presented an AMP mining workflow with AMPlify as the core AMP prediction module, to predict novel AMPs from all eukaryotic sequences in the UniProtKB / Swiss-Prot database (UniProt Consortium, 2019), which contains only manually annotated and reviewed records from the entire UniProt database. With the AMP mining workflow, 8,008 novel putative AMPs were identified, and some sequences are released in the Zenodo repository at doi.org / 10.5281 / zenodo.8133088 (Li and Bird 2023(a)) to provide the community with candidate peptides for lab validation. Inspired by the growing global concern of animals passing antibiotic-resistant bacteria to humans as a direct consequence of antibiotics overuse in animal agriculture including poultry farming (Apata 2009), a total of 38 predicted peptides were tested in vitro based on their structural similarities to known chicken AMPs. Preliminarily tests were conducted of the synthesized peptides against two lab strains of Escherichia coli ATCC 25922 and Staphylococcus aureus ATCC 29213, and 13 demonstrated antimicrobial activity against at least one bacterial strain. While chicken AMPs have theoretically been evolved to fight against pathogens in chickens’ living environment (Zhang and Gallo, 2016) the utility of some of these newly discovered AMPs if foreseeable, with high structural similarities to known chicken AMPs, in chicken farming as novel substitutes for conventional antibiotics.

[0178] 2.2. Materials and Methods

[0179] Data Sets. This study involves four main datasets: 1) all eukaryotic entries from UniProtKB / Swiss-Prot (UniProt Consortium, 2019), 2) entries annotated with AMP- related keywords in UniProtKB / Swiss-Prot (UniProt Consortium, 2019), 3) known AMPsequences in AMP databases, and 4) known chicken AMPs. All UniProtKB / Swiss-Prot sequences were downloaded from the 2022_02 release.

[0180] All eukaryotic entries from the UniProtKB / Swiss-Prot database (UniProt Consortium, 2019) were downloaded by searching with query “(taxonomy_id:2759) AND (reviewed:true)”. This set of sequences include 186,302 distinct sequences from 195,188 UniProt entries.

[0181] The set of entries annotated with AMP-related keywords in their annotations in the UniProtKB / Swiss-Prot database (UniProt Consortium, 2019) includes 18,470 distinct sequences from 19,845 UniProt entries. UniProt entries with their annotations containing any of the following 16 AMP-related keywords were downloaded: {antimicrobial, antibiotic, antibacterial, antiviral, antifungal, antimalarial, antiparasitic, anti-protist, anticancer, defense, defensin, cathelicidin, histatin, bacteriocin, microbicidal, fungicide} (Li C. et al., 2022). There were 11,925 sequences from 12,229 UniProt entries in this set belong the above eukaryotic sequence set.

[0182] The AMP sequence set was taken from a previous considered set, comprising 4,538 distinct AMP sequences. All AMP sequences were downloaded from two curated AMP databases: Antimicrobial Peptide Database (APD3, aps. unmc.edu / AP) on July 11 , 2022 and Database of Anuran Defense Peptides (DADP, split4.pmfst.hr / dadp) on December 6, 2018.

[0183] The known chicken AMP sequence set include 22 sequences in total. Sequences that were downloaded on Oct 14, 2022 from the APD3 database were located by searching for source organism “Gallus gallus".

[0184] AMP mining workflow. The AMP mining workflow proposed in this study includes five steps, as shown in Figure 7.

[0185] Figure 7 shows AMP mining workflow. The AMP mining workflow utilizes the rAMPage (Lin et al., 2022) cleavage module with ProP (Duckert et al., 2004) to cleave putative precursor sequences and AMPlify (Li C. et al., 2022) (Li C., Warren, et al., 2023) to predict AMP sequences. The numbers of candidate mature sequences, candidate mature sequence records, parent sequences, and UniProt entries, that remained at each step are reported. Two putative mature sequence records are considered distinct from each other if any of these three attributes are different: UniProt entry ID, position(s) in the parent sequence, and candidate mature sequence. All numbers presented in the workflow are non- redundant.

[0186] In Step 1, all eukaryotic protein sequences from the UniProtKB / Swiss-Prot database (UniProt Consortium, 2019) were cleaved to generate mature sequences using thecleavage module of the AMP discovery pipeline rAMPage v1.0.1 (Lin et al., 2022). The cleavage module of rAMPage is based on ProP vl .Oc (Duckert et al., 2004), a machine learning based tool for cleavage site prediction for eukaryotic protein sequences. The rAMPage cleavage module also provides post-processing of the ProP output (Lin et al., 2022). With exception of the signal peptide, ProP only predicts the cleavage sites rather than identifying whether each cleaved sequence is a bioactive peptide or a pro-sequence (Lin et al., 2022) (Duckert et al., 2004). As a result, for each putative precursor sequence, the rAMPage cleavage module takes all cleaved pieces, excluding the predicted signal peptide, as well as all possible recombinations of non-adjacent cleaved pieces (with a maximum of three pieces within each recombination) as candidate mature sequences (Lin et al., 2022). Additionally, the rAMPage cleavage module removes candidate mature sequences with lengths < 2 or > 200 amino acids (aa), as AMPs are typically short sequences (van der Does et al., 2019). The redundancy removal function of the rAMPage cleavage module was turned off in this study, as the same AMP sequence from different UniProt entries were kept as different records (Figure 7) for convenience of future investigations into their source organisms. However, these were not considered as the same AMP sequence when calculating the sequence counts.

[0187] In Step 2, sequences without signal peptides predicted were removed, as AMPs are mostly secreted peptides (Bals 2000).

[0188] In Step 3, the remained candidate mature sequences from Step 2 were passed through the imbalanced model of AMPlify v1.1.0 (Li C., Warren, et al., 2023) for a preliminary filtering. Candidate mature sequences predicted as AMPs by the AMPlify imbalanced model were passed on to the next step. Sequences with non-standard amino acids were also filtered out by AMPlify, as AMPlify does not assign those sequences with any predictions.

[0189] In Step 4, the predicted AMPs by the AMPlify imbalanced model from Step 3 were then passed through the balanced model of AMPlify v1.1.0 (Li C. et al., 2022) for a more precise filtering. Sequences that were predicated as AMPs by the AMPlify imbalanced model from Step 3 as well as the AMPlify balanced model were collected as the final prediction results of the AMPlify balanced and imbalanced models.

[0190] In Step 5, all the predicted AMPs from Step 4 were checked against the known AMP sequence set as well as the annotations of the corresponding UniProt entries for novel putative AMPs. Predicted AMPs were found to be heretofore undescribed if they werenot in the known AMP sequence set and if their corresponding UniProt entry annotation did not have AMP-related keywords.

[0191] Sequence and structural similarities between peptides. This work involves two different types of similarity measurements between peptides: sequence similarity and structural similarity.

[0192] The sequence similarity between two peptides was calculated asare lengths of the peptide sequences. The structural similarity between two peptides was measured with TM-scores normalized by the average peptide sequence length (Zhang et al., 2004) using TM-align (Zhang et al., 2005). TM-scores are between 0 and 1 , with higher TM-scores indicating more similarity between the two structures (Zhang et al., 2004). The three-dimensional (3D) structures of peptides were predicted using ColabFold (Mirdita et al, 2022), which provides a faster prediction speed by combining MMseqs2 (Steinegger et al., 2017) homology search with AlphaFold2 (Jumper et al., 2021).

[0193] The sequence / structural similarity of a single peptide to a set of peptides was defined as the maximum of all sequence / structural similarity values calculated between that peptide and the peptides in the target set for comparison (i.e. the sequence / structural similarity of that peptide to its most similar target set peptide in their sequences / structures).

[0194] Selecting putative AMPs for validation. Considering the large number of putative AMPs identified by the AMP mining workflow, a few of the sequences were prioritized for synthesis and in vitro validation. This work primarily focused on the application of AMPs in poultry farming. Considering the facts that 3D structures of proteins determine their biological functions and eukaryotic AMPs have evolved to fight against the pathogens that infect the respective hosts (Zhang and Gallo, 2016), putative AMPs were prioritized by their structural similarities to known chicken AMPs, foreseeing their potentials in chicken farming against pathogens that affect chickens.

[0195] Among all putative AMPs unearthed by the AMP mining workflow, 397 short cationic peptides were prioritized with AMPlify scores > 10 (i.e. AMPlify probability scores > 0.9) for selection, as shorter peptides are more cost-effective to synthesize (Lin et al., 2022). Short cationic peptides were defined to be those with lengths between 5 and 35 aa (amino acids) and net charge greater than 0. The AMPlify score is a prediction score reported byAMPlify, which is a log transformation of the AMPlify probability scoreas— IQlOg’igCl “ fiijfft y)■ , and here the AMPlify scores from the balanced model were taken for analysis.

[0196] First, the 3D structures of the 397 putative AMPs, together with the 22 known chicken AMPs, were predicted by ColabFold (Mirdita et al, 2022). Next, the structural similarity (TM-score) of each AMP to the known chicken AMPs (i.e. structural similarity of the putative AMP to the most structurally similar known chicken AMP) was calculated using TM- align (Zhang et al., 2005). The most structurally similar known chicken AMP to a specific putative AMP was defined as the reference chicken AMP of that putative AMP. Finally, were ranked the peptides by their structural similarities to known chicken AMPs. The top 30 AMPs with highest TM-scores were prioritized for synthesis, as well as the bottom 10 with lowest TM-scores for comparison (Table 5, Table 6). Furthermore, the three reference chicken AMPs for the top 30 AMPs: Chicken CATH-2 (van Dijk et al., 2005), Chicken CATH-3 (Xiao et al, 2006), and Ovipin (Santos et al., 2022), were also sent for synthesis for comparison (Table 7). Two of the top 30 AMPs could not be successfully synthesized, resulting in a final number of 38 peptides in total proceeding to further in vitro validation.

[0197] Table 5 shows characteristics of the 40 AMP sequences mined from the UniProtKB / Swiss-Prot database that have been prioritized for synthesis. A total number of 397 short cationic novel putative AMPs with AMPlify scores > 10 (i.e. AMPlify probability scores > 0.9) were compared with known chicken AMPs for their similarities in three- dimensional (3D) structures. The top section shows the 30 most structurally similar AMPs to the known chicken AMPs as measured by TM-scores, while the bottom section shows the 10 least structurally similar ones to the known chicken AMPs.

[0198] Table 6 shows an overview of the 40 AMP sequences mined from the UniProtKB / Swiss-Prot database that have been prioritized for synthesis regarding their corresponding parent sequence information in the database. This table supplements Table 5, by providing information about the 40 AMP sequences regarding their corresponding parent sequences. The top section shows the 30 most structurally similar AMPs to the known chicken AMPs as measured by TM-scores, while the bottom section shows the 10 least structurally similar ones to the known chicken AMPs.

[0199] Table 7 shows seven reference chicken AMP sequences for the 40 AMPs mined from the UniProtKB / Swiss-Prot database and prioritized for synthesis. The AMP sequences listed in this table correspond to the seven reference chicken AMPs for the 40 AMPs listed in Table 5. Only the three reference chicken AMPs for the top 30 AMPs which share highest structural similarities to known chicken AMPs were prioritized for synthesis and further tests (i.e. Chicken CATH-2, Chicken CATH-3, and Ovipin).

[0200] Antimicrobial susceptibility testing (AST). To measure the antimicrobial activity of the selected peptides in vitro, broth microdilution assays were conducted to determine the minimum inhibitory and minimum bactericidal concentrations (MICs and MBCs, respectively), as outlined by the Clinical and Laboratory Standards Institute (CLSI). Some adaptations for testing cationic AMPs as described previously (Wiegand et al, 2008) were incorporated. Laboratory isolates of Escherichia coli 25922 and Staphylococcus aureus 29213, purchased from the American Type Culture Collection (ATCC; Manassas, VA, USA), were used to validate the selected AMPs. Bacteria from frozen stocks were streaked onto nonselective Columbia blood agar with 5% sheep blood (Oxoid) and incubated for 18-24 h at 37°C. On the next day, 2-4 colonies were streaked onto a new agar plate and incubated for 18-24 h at 37°C to ensure uniform colony health before the assay. To create a standardized bacterial inoculum, isolated colonies were suspended in Mueller-Hinton Broth (MHB; Sigma- Aldrich, St. Louis. MO, USA). The suspension was adjusted to an optical density of 0.08-0.1 at 600nm, which corresponds to a 0.5 McFarland standard of approximately 1-2 x 108CFU / mL (CFU: colony forming units). The inoculum was then diluted 1 :250, resulting in a final concentration of 5±3 x 105CFU / mL. The target bacterial density was confirmed by routinely examining the total viability counts from the final inoculum.

[0201] The selected AMPs were purchased from and synthesized by GenScript (Piscataway, NJ, USA). The peptides were received in lyophilized format and stored at -20°C, and were suspended in sterile ultrapure water prior to testing. A two-fold serial dilution of 1280 to 2.5 pg / mL was prepared in sterile 96-well polypropylene microtitre plates (Greiner Bio-One #650261, Kremsmunster, Austria). Then, 100 pL of the standardized bacterial inoculum was added to each well, resulting in a final AMP testing range of 128 to 0.25 pg / mL. The MIC values were determined as the lowest peptide concentration in which no visible bacterial growth were observed after a 20-24 h incubation at 37°C. To determine the MBC values, well contents of the MIC and the two adjacent wells containing the two- and four-fold MIC were plated onto nonselective nutrient agar. The concentration in which 99.9% of the inoculum was killed after incubation for 24 hours at 37°C was reported as the MBC.

[0202] In the tests conducted, a known AMP Ranatuerin-4 (Goraya et al., 1998) from the American bullfrog and an in-house peptide [TKPKGk (OT15, SEQ ID NO:8074) were used as the positive and negative control peptides, respectively. The OT15 was truncated and derived from a negative control peptide [TKPKG^ (OT20, SEQ ID NO:8075), not antimicrobial but with similar characteristics to AMPs, used in previous studies (Horvati et al., 2017).

[0203] Hemolysis assay. Hemolysis experiments were performed to evaluate the toxicity of the select peptides to red blood cells (RBCs). Whole blood from healthy donor pigs was purchased from Lampire Biological Laboratories (Pipersville, PA, USA). RBCs were washed and isolated by centrifugation, using Roswell Park Memorial Institute medium (RPMI) (Life Technologies, Grand Island, NY, USA; Gibco cat# 11835-030). AMPs were suspended and serially diluted from 1280 to 10 pg / mL using RPMI in a 96-well polypropylene microtitre plate before being combined with 100 pL of the 1% RBC solution, resulting in a final AMP testing range of 128 to 1 pg / mL. Following an incubation at 37°C for 30-45 minutes, plates were centrifuged and 1 / 2 volume from each supernatant was transferred to a new 96-well plate. The absorbance of the wells was measured at 415 nm, and the AMP concentration that lysed > 50% of the RBCs (HCso) was used to determine the hemolytic activity. Absorbance readings from wells containing RBCs treated with 11 pL of a 2% Triton- Xi 00 solution or RPMI (AMP solvent-only) were used to define 100% and 0% hemolysis, respectively. All centrifugation steps were performed at 500* g for five minutes in an Allegra- 6R centrifuge (Beckman Coulter, CA, USA).

[0204] 2.3. Results

[0205] Integration of AMPlify balanced and imbalanced models. The AMP prediction module of the AMP mining workflow utilized both AMPlify imbalanced and balanced models (Li C. et al., 2022) (Li C., Warren, et al., 2023) to obtain a curated set of candidate AMP sequences. The imbalanced model of AMPlify, trained on an imbalanced training set and added in v1.1.0, has been proved to perform well on large, highly imbalanced candidate sequence sets with far more non-AMPs than AMPs (Li C., Warren, et al., 2023). On the other hand, the balanced model of AMPlify, trained on a balanced training set, has demonstrated stronger performance on more curated candidate sequence sets, such as what has been applied in the mining of the bullfrog genome (Li C. et al., 2022) (Li C., Warren, et al., 2023). Based on the performance of the two models in their respective advantageous application scenarios, both of the two models were incorporated in the AMP mining workflow. As shown in Figure 7, the prediction module (Step 3 and Step 4) performed a two-stage filtering by applying the imbalanced and balanced models in turn. The integration of the two models can also be considered as a filtering scheme of only taking sequences that are predicted by both models as AMPs to be AMPs.

[0206] In previous work, two test sets were built to evaluate the performance of the balanced and imbalanced models (Li C., Warren, et al., 2023). The balanced test set comprises 835 AMPs and 835 non-AMPs, while the imbalanced test set comprises 835 AMPs and 25,689 non-AMPs (Li C., Warren, et al., 2023). Table 8 shows the performance comparison of the balanced and imbalanced models of AMPlify as well as the integration of the two (noted as “balanced + imbalanced”) on the two test sets with regard to accuracy, sensitivity, specificity, F1 score and area under the receiver operating characteristic curve (AUROC). The integration of the two models output predicted labels instead of probabilities, hence the AUROC values were not reported for it.

[0207] Table 8 shows a performance comparison of the balanced / imbalanced models of AMPlify, as well as the integration of the two, on the balanced and imbalanced test sets. Values of accuracy (acc), sensitivity (sens), specificity (spec), F1 score (F1) and area under the receiver operating characteristic curve (AUROC) are presented in percentage.

[0208] On the balanced test set, where the AMPlify balanced model performs better than the imbalanced model except for sensitivity (Li C., Warren, et al., 2023), the integration of the two AMPlify models does not show much improvement upon the two models when used alone, with only an increase in specificity by 1.2% (95.69% vs. 94.49%).

[0209] On the imbalanced test set, where the AMPlify imbalanced model performs better than the balanced model in all five metrics (Li C., Warren, et al., 2023), the integration of the two models achieves improvement in accuracy (99.31% vs. 98.94%), specificity (99.56% vs. 99.09%), and F1 score (89.25% vs. 84.87%). For highly imbalanced candidate sequence sets where non-AMP sequences far outnumber AMP sequences, it is crucial to minimize the number of false positives to prevent an excessive number of non-AMP sequences from being forwarded to downstream in vitro validation. Although there was a slight decrease in sensitivity (91.50% vs. 94.37%), the improvement in other metrics, particularly specificity, suggests that the integration of the two models is better than utilizing the imbalanced model alone when facing prediction scenarios where AMPs are only a small portion of a large candidate sequence set.

[0210] Additionally, all AMPlify models (balanced, imbalanced, and their integration) outperforms other existing AMP prediction methods for comparison on the two test sets (Table 9, Table 10).

[0211] Table 9 shows a performance comparison among different tools on the balanced test set. Values of accuracy (acc), sensitivity (sens), specificity (spec), F1 score(F1) and area under the receiver operating characteristic curve (AUROC) are presented in percentage.

[0212] Table 10 shows a performance comparison among different tools on the imbalanced test set. Values of accuracy (acc), sensitivity (sens), specificity (spec), F1 score (F1) and area under the receiver operating characteristic curve (AUROC) are presented in percentage.

[0213] Considering that the source for AMP mining is a large protein database (UniProt Consortium, 2019), in which most of the sequences do not belong to the AMP family, filtering for candidate AMP sequences by applying the integration of imbalanced and balanced models is an optimal way to reduce the number of false positives.

[0214] Predicted AMPs

[0215] By applying the AMP mining workflow to all eukaryotic sequences in UniProtKB / Swiss-Prot (27), 10,720 distinct candidate mature sequences were predicted as AMPs, of which 8,008 (74.70%) (SEQ ID NO:59 - SEQ ID NO: 8066) are novel AMPs (Figure 7). All these predicted AMPs were derived from parent sequences with signal peptides predicted. Although, there are mature or partial precursor sequences reported without their complete precursor forms presented in the UniProtKB / Swiss-Prot database (UniProt Consortium, 2019), all sequences without signal peptides predicted were filtered out to ensure a cleaner candidate set.

[0216] Figure 8 shows length and net charge distributions of the AMPs predicted by AMPlify from the UniProtKB / Swiss-Prot database. Length and net charge distributions were calculated for all 10,720 predicted AMPs as well as the 8,008 AMPs (SEQ ID NO:59 to SEQ ID NQ:8066) described herein, among them. The distributions of the 4,538 known AMP sequences from APD3 (Wang et al., 2016) and DADP (Novkovic et al., 2012) databases were plotted alongside for comparison. Mean (p) and standard deviation (o) values of each distribution are as follows: all predicted AMPs (length: p = 52.83 aa, o = 26.59 aa; net charge: p= 3.02, o = 3.90), novel putative AMPs (length: p = 52.49 aa, o = 26.49 aa; net charge: p = 3.04, o = 3.88), and known AMPs (length: p = 30.21 aa, o = 20.28 aa; net charge: p = 3.05, o = 3.10).

[0217] The 10,720 predicted AMPs have an average length of 52.83 aa and an average net charge of 3.02, while the 8,008 AMPs (SEQ ID NO:59 to SEQ ID NQ:8066) described herein, have an average length of 52.49 aa and an average net charge of 3.04 (Figure 8). Compared with the 4,538 known AMP sequences, which have an average length of 30.21 aa and an average net charge of 3.05 (Figure 8), the average lengths of both the predicted AMP set and the novel putative AMP set mined from UniProtKB / Swiss-Prot are at least 20 aa larger. However, researchers tend to prioritize shorter sequences for validation due to their lower synthesis costs for in vitro lab validation (Lin et al., 2022), which could explain the difference in average lengths. Furthermore, the novel putative AMP sequences exhibit a remarkable level of novelty, sharing a low sequence similarity level of 32.71% on average to the known AMP sequences as shown in Figure 9.

[0218] Figure 9 shows Sequence similarity distributions of the predicted AMPs from the UniProtKB / Swiss-Prot database to known AMPs. The sequence similarity distribution of all 10,720 predicted AMPs to known AMPs from APD3 and DADP databases was visualized, along with that of the 8,008 AMPs (SEQ ID NO:59 to SEQ ID NO:8066) described herein from all predicted AMPs. The former distribution holds a mean of 38.27% and a standard deviation of 17.28%, while the latter holds a mean of 32.71% and a standard deviation of 7.21%. The sequence similarity of each predicted AMP to known AMPs was considered as the sequence similarity of that predicted AMP sequence to its most similar sequence in the known AMP set, based on which the distributions were plotted.

[0219] Tracing back to the original source entries in the protein database, the 10,720 predicted AMPs corresponded to 8,481 parent sequences from a total of 8,862 UniProt entries, while the 8,008 novel AMPs (SEQ ID NO:59 to SEQ ID NQ:8066) described herein, corresponded to 6,349 parent sequences from 6,654 UniProt entries (Figure 7). Upon analysis, parent sequences with AMPs predicted from their cleaved mature sequences were defined to be putative AMP precursor sequences, and UniProt entries of those putative AMP precursor sequences to be putative AMP entries.

[0220] Figure 10 shows distribution for the number of predicted mature AMP sequences found within each putative AMP precursor sequence mined from the UniProtKB / Swiss-Prot database. The bar chart was plotted based on 8,481 distinct putative AMP precursor sequences from the 8,862 putative AMP entries. Each bar shows the number of putative AMP precursor sequences that were predicted with the corresponding number of distinct mature AMPs by AMPlify.

[0221] Among all 8,481 putative AMP precursor sequences identified, 86.84% (7,365) only had one distinct predicted AMP sequence, as shown in Figure 10. However, it was noticed that cases where multiple AMP sequences were predicted from a single putative precursor sequence. Specifically, 698 putative precursor sequences were predicted with two distinct AMPs from each, 207 predicted with three, 74 predicted with four, and 57 predicted with five. There were 80 putative precursor sequences that had more than five distinct AMPs predicted from each of them.

[0222] Figure 11 shows categorization of the AMP entries identified from the UniProtKB / Swiss-Prot database based on source organisms. The 8,862 putative AMP entries with mature AMPs predicted by AMPlify were classified into five categories of mammal, plant, amphibian, insect, and others, based on their source organisms. The predicted AMPs were further checked against the known AMP sequences in APD3 andDADP databases as well as the annotations of the corresponding UniProt entries. Those not found in those two AMP databases and not annotated with the AMP-related keywords were labelled as novel AMP sequences.

[0223] Figure 11 categorizes the 8,862 putative AMP entries identified by the AMP mining workflow by source organisms. According to the statistics made by the APD3 web server on January 2023, out of all 3,569 sequences in records, amphibians are the largest organism source of AMPs (1,196 sequences), followed by bacteria (380 sequences), plants (371 sequences), insects (367 sequences), and mammals (363 sequences) (Wang et al., 2016). Based on this, all the AMP entries were classified into the following five categories: amphibian, plant, insect, mammal, and others. The category of bacterial AMPs was not included, as only eukaryotic AMPs were considered in this work.

[0224] As seen in Figure 11 , mammalian entries constitute the largest source organism category, comprising 37.93% (3,361 / 8,862) of all AMP entries. Out of all 3,361 putative AMP entries from the mammalian category, 2,703 (80.42%) of them had novel AMPs predicted from their cleaved candidate mature sequences, making it the largest source organism category for novel putative AMP entries as well. Putative AMP entries not belonging to any of the four specified source organism categories also make up a large portion of all putative AMP entries, following the mammalian category, and represent 33.78% (2,994 / 8,862) of all putative AMP entries. In this category, 79.93% (2,393 / 2,994) had novel AMPs predicted from their cleaved candidate mature sequences. Plants are the third largest group among putative AMP entries, contributing 17.42% (1,544 / 8,862) to the total. Overall, 1,170 out of 1,544 putative plant AMP entries (75.78%) were considered as novel putative AMP entries. Amphibian and insect are the two smallest categories, with 509 and 454 putative AMP entries in each, respectively. While 59.25% (269 / 454) of the insect category were determined to be novel AMP entries, amphibian is the only source organism category with more than half of the entries (76.62%) already identified as known AMP entries and / or annotated with AMP-related keywords. The main reason for this is that amphibians are considered as a rich source of AMPs (Helbing et al., 2019), hence has become a popular source for AMP mining (Lin et al., 2022) (Li C. et al., 2022) (Richter et al., 2022), resulting in more of the amphibian AMPs having been discovered than other organisms. There are no cases where a single putative AMP precursor has novel putative AMP(s), but at the same time, has known AMP(s) in its cleaved candidate mature peptides and / or has been annotated with AMP-related keywords. UniProt entries annotated with AMP-related keywords can be proteins or peptides that assist in the defense against microbes rather than possessingantimicrobial activities themselves. However, to ensure a cleaner set of novel putative AMPs, all of the predicted AMPs with their parent sequence entries annotated with AMP-related keywords were not considered as novel discovery.

[0225] The information regarding AMPs mined in this study were deposited to Zenodo repository at doi.org / 10.5281 / zenodo.8133088 (Li and Bird 2023(a)), providing the community with a long list of potential AMP sequences for future validation.

[0226] In vitro validation results

[0227] Figure 12 shows structural similarity distributions of the short cationic AMPs mined from the UniProtKB / Swiss-Prot database to known chicken AMPs. The structural similarity distribution of 1 ,596 short cationic putative AMPs to known chicken AMPs holds a mean of 0.4266 and a standard deviation of 0.1690. Among all short cationic putative AMPs, 397 of them are with AMPlify scores > 10 (i.e. AMPlify probability scores > 0.9). The structural similarity distribution of these 397 putative AMPs to known chicken AMPs holds a mean of 0.4434 and a standard deviation of 0.1611. The structural similarity (TM-score) of each predicted AMP to known chicken AMPs was considered as the structural similarity of the predicted AMP to the most structurally similar known chicken AMP (i.e. reference chicken AMP) from the APD3 database, based on which the distributions were plotted.

[0228] A total of 40 short cationic novel AMP sequences were selected with AMPlify scores > 10 for synthesis based on their structural similarities, as measured by TM-scores, to known chicken AMPs (Figure 12), with detailed methods shown in the Materials and Methods section. Of those selected, 30 had the highest TM-scores, while the remaining 10 had the lowest TM-scores. Table 5 and Table 6 list the characteristics of the 40 peptides and their original UniProt entry information, respectively. Two of the top 30 sequences were not successfully synthesized, resulting in a final list of 38 AMPs for in vitro validation. The three reference chicken AMPs, Chicken CATH-2 (van Dijk et al., 2005), Chicken CATH-3 (Xiao et al, 2006), and Ovipin (Santos et al., 2022), that matched the top 30 sequences were tested for comparison, with their information shown in Table 7.

[0229] Figure 13 shows antimicrobial and hemolytic activities of the 13 novel AMPs mined from the UniProtKB / Swiss-Prot database that were active against at least one bacterial strain of Escherichia coli ATCC 25922 and Staphylococcus aureus ATCC 29213. Antimicrobial and hemolytic activities were measured by minimum inhibitory concentration (MIC) and concentration that lyses 50% (HCso) of the red blood cells (RBCs), respectively. HC50 was determined using porcine RBCs. Data is presented as the lowest effective peptide concentration range (pg / mL) observed in three independent experiments performed induplicate, with one maximum data point and one minimum data point dropped for each measurement. The results are divided into two sections, as separated by the solid vertical line. The left section shows results of the 11 active peptides from the 28 successfully synthesized peptides among the 30 most structurally similar peptides to known chicken AMPs as measured by TM-scores, while the right section shows those of the two active peptides from the 10 least structurally similar ones to known chicken AMPs. These 11 active peptides in the left section are categorized into three sub-sections, as separated by the dashed vertical lines, according to their reference chicken AMPs (i.e. the most similar known chicken AMP in 3D structure). The results of the three reference chicken AMPs: Chicken CATH-2, Chicken CATH-3, and Ovipin, are listed for comparison in each sub-section. Hemolysis experiments were not performed for AMPs showed no antimicrobial activity (MIC > 128) in at least two repeats for each bacterial strain tested.

[0230] The 38 AMPs were tested against two bacterial isolates: the Gram-negative Escherichia coli ATCC 25922, and the Gram-positive Staphylococcus aureus ATCC 29213. Porcine RBCs were used to assess the hemolytic activity of the peptides. Out of the 38 AMPs for in vitro validation, 13 peptides displayed antimicrobial activity against Escherichia coli ATCC 25922, with three of them additionally active against Staphylococcus aureus ATCC 29213. Figure 13 summarizes the antimicrobial and hemolytic activities (in MIC and HCso, respectively) of the 13 active peptides, with the entire in vitro validation results of all peptides shown in Table 11. The left section of Figure 13 shows the results for the 11 active peptides from the 30 most structurally similar AMPs to known chicken AMPs, with their reference chicken AMPs for comparison. Results of the two active peptides from 10 least structurally similar AMPs to known chicken AMPs are shown in the right section of Figure 13.

[0231] Table 11 shows antimicrobial susceptibility testing and hemolysis experiment results of the 38 successfully synthesized AMPs mined from the UniProtKB / Swiss-Prot database and tested in vitro. Peptides were tested for their antimicrobial activity against Escherichia coli ATCC 25922 and Staphylococcus aureus ATCC 29213 for their minimum inhibitory concentration (MIC) and minimum bactericidal concentration (MBC) values. Porcine red blood cells (RBCs) were used to test the hemolytic activity of the selected peptides for their hemolytic concentration (HC50) values. Data is presented as the lowest effective peptide concentration range (pg / mL) observed in three independent experiments performed in duplicate, with one maximum data point and one minimum data point dropped for each measurement. The top section shows results of the 28 successfully synthesizedpeptides among the top 30 peptides in structural similarity to known chicken AMPs as measured by TM-scores, while the second section shows results of the 10 least structurally similar ones to known chicken AMPs. Results of the three reference chicken AMPs (Chicken CATH-2, Chicken CATH-3, and Ovipin) for the top 30 peptides are listed in the third section for comparison. The control peptides in the bottom section includes: a positive control peptide Ranatuerin-4 (Goraya et al., 1998) and a negative control peptide OT15 (SEQ ID NO:8074).

[0232] Figure 14 shows a visualization of antimicrobial activity of the 38 tested AMPs with respect to AMPlify score and associated structural similarity to known chicken AMPs. Peptides without any observable antimicrobial activity are presented as grey crosses, and the active peptides are presented in shaded dots with their names annotated. Dots with darker shades indicate stronger antimicrobial activity against Escherichia coli ATCC 25922, determined by the minimum MIC value of each peptide against the strain. AMPlify scores from the balanced model were used for visualization. The structural similarity (TM-score) of each tested peptide to known chicken AMPs was considered as the structural similarity of that peptide to the most structurally similar known chicken AMP (i.e. reference chicken AMP) from the APD3 database. Figure 14 illustrates the activity of all tested AMPs with regard to their AMPlify scores and sequence similarities to known chicken AMPs.

[0233] Among 10 tested AMPs with Chicken CATH-2 as their reference chicken AMP, five showed antimicrobial activity in the tests conducted. Within this group of five peptides, HoSal , with its original UniProt entry annotated as an olfactory receptor from human (UniProt Consortium, 2019), possessed the strongest antimicrobial activity against E. coli ATCC 25922 (MIC = 16 pg / mL), and was the only peptide that showed additional activity against S. aureus ATCC 29213 (MIC = 128 pg / mL). DiDi2 showed the same antibacterial activity as HoSal against E. coli ATCC 25922 with an MIC of 16 pg / mL, with its original UniProt entry annotated as a putative uncharacterized transmembrane protein from social amoeba (UniProt Consortium, 2019). DaRel, of which the parent sequence is annotated as a GTP-binding protein from zebrafish, inhibited the growth of E. coli ATCC 25922 at MIC of 32-64 pg / mL. Molnl and MuMul presented weakest antibacterial activity against E. coli ATCC 25922 this group, with MICs of 64->128 pg / mL and >128 pg / mL. The parent sequence of Molnl is annotated as part of an intermediate translocation complex at the inner envelope membrane of chloroplasts from mulberry, while the parent sequence of MuMul is annotated as a mouse surfeit locus protein, a component of the MITRAC (mitochondrial translation regulation assembly intermediate of cytochrome c oxidase complex) (UniProt Consortium, 2019). The reference chicken AMP, Chicken CATH-2, possessed stronger antimicrobial activity against E. coli ATCC 25922 (MIC = 8-16 pg / mL) and S. aureus ATCC 29213 (MIC = 32 pg / mL) than all five peptides in this group.

[0234] Among 16 tested AMPs with Chicken CATH-3 as their reference chicken AMP, five showed antimicrobial activity in the tests conducted. Among the five peptides in the group, PIVil (SEQ ID NO:483) had the strongest antimicrobial activity against both E. coli ATCC 25922 (MIC = 4-8 pg / mL) and S. aureus ATCC 29213 (MIC = 32 pg / mL), and it wasalso the most active peptides among all 38 AMPs tested. The parent sequence of PIVil (SEQ ID NO:483) is annotated as the precursor of a RxLR effector protein that completely suppresses the host cell death induced by cell death-inducing proteins, which is derived from the a type of oomycete that causes the downy mildew disease of grapevines (UniProt Consortium, 2019). HoSa4 (SEQ ID NO: 285), derived from an IQ domain-containing protein from human without detailed functions annotated in UniProtKB / Swiss-Prot (UniProt Consortium, 2019), was the second most active AMP in this group, with antibacterial activity against E. coli ATCC 25922 and S. aureus ATCC 29213 with MICs of 32 pg / mL and 64-128 pg / mL. The PIVil (SEQ ID NO:483) and HoSa4 (SEQ ID NO: 285) peptides are the only two peptides that were active against both bacterial strains tested in this group. ArTh11 (SEQ ID NO:2797) inhibited the growth of E. coli ATCC 25922 at MIC of 64-128 pg / mL. The parent sequence of ArTh11 (SEQ ID NO:2797) is annotated as an F-box protein from mouse-ear cress in UniProtKB / Swiss-Prot (UniProt Consortium, 2019) but without detailed functions specified. SuSc1 (SEQ ID NQ:4032) and HoSa7 (SEQ ID NO: 2048) were the two least active peptides in the group, with antibacterial activity against E. coli ATCC 25922 MICs of 128 pg / mL and 64->128 pg / mL, respectively. SuSc1 (SEQ ID NQ:4032) was derived from an precursor sequence of adrenomedullin in pigs and cattle, while HoSa7 (SEQ ID NO: 2048) was cleaved from a human G-protein coupled receptor, as annotated in UniProtKB / Swiss- Prot (UniProt Consortium, 2019). The reference chicken AMP, Chicken CATH-3, possessed stronger antimicrobial activity against E. coli ATCC 25922 (MIC = 2-4 pg / mL) and S. aureus ATCC 29213 (MIC = 2 pg / mL) than all five peptides in this group.

[0235] Two peptides in the top 30 AMP list matched Ovipin as their reference chicken AMP, with only one of them (GaGa2, SEQ ID NO:1291) showing antimicrobial activity in the tests conducted. However, it only showed minimal antibacterial activity against E. coli ATCC 25922 (MIC = 128 pg / mL), with no activity against S. aureus ATCC 29213. This peptide was derived from a chicken tenascin (UniProt Consortium, 2019). The reference chicken AMP, Ovipin, did not show any activity against the two bacterial strains tested in this study, but was proved with activity against a Gram-positive Micrococcus luteus strain in the previous study (Santos et al., 2022).

[0236] Among the 10 least structurally similar AMPs to known chicken AMPs, only two (GoGol of SEQ ID NO:833; and UnBil of SEQ ID NO:2343) showed activity in the tests conducted. GoGol (SEQ ID NO:833) was part of the sequence of a transcriptional repressor found in five different primate species of western lowland gorilla, bonobo, Bornean orangutan, chimpanzee, as well as human (UniProt Consortium, 2019). It inhibited thegrowth of E. coli ATCC 25922 with MIC of 64 pg / mL. The parent peptide of UnBil was annotated as a neurotoxin from sea snail in UniProtKB / Swiss-Prot (UniProt Consortium, 2019). It only showed minimal activity against E. coli ATCC 25922 (MIC = 128 pg / mL). Neither of these two peptides were active against S. aureus ATCC 29213.

[0237] Except Molnl (SEQ ID NO:2444), MuMul (SEQ ID NO: 1114), and HoSa7 (SEQ ID NO:2048) that were not sent for hemolysis assays (Table 11), all the other 10 novel AMPs proven with antimicrobial activity shown in this work were not hemolytic to the porcine RBCs (HC5O > 128).

[0238] 2.4. DISCUSSION

[0239] In this work, an AMP mining workflow was introduced for novel AMP discovery from eukaryotic protein databases, and applied the workflow to unearth novel eukaryotic AMPs from the UniProtKB / Swiss-Prot database (UniProt Consortium, 2019). The AMP mining workflow utilizes the state-of-the-art AMP prediction tool AMPlify (Li C. et al., 2022) (Li C., Warren, et al., 2023), as well as the rAMPage cleavage module (Lin et al., 2022) with ProP (Duckert et al., 2004) for precursor sequence cleavage. Following the AMP mining workflow, 8,008 distinct novel AMP sequences (SEQ ID NO:59 - SEQ ID NQ:8066) were identified from all eukaryotic sequences in the UniProtKB / Swiss-Prot database. Sequences were submitted to Zenodo repository at doi.org / 10.5281 / zenodo.8133088 (Li and Bird 2023(a)).

[0240] Although the AMP mining workflow identified a considerable number of putative AMPs, there do exist some limitations due to the machine learning based tools utilized. Neither AMPlify nor ProP is 100% accurate at respective tasks. Like many other machine learning based bioinformatics tools, the performance of such tools can be limited by the degree of curation and size of the available data used for training (Lin et al., 2022) (Li C. et al., 2022). Nevertheless, their reported performance (Li C. et al., 2022) (Li C., Warren, et al., 2023) (Duckert et al., 2004) still establishes them as state-of-the-art tools. Consequently, they are still highly suitable for integration into any AMP mining workflow. It is expected that he limitations of these tools will gradually diminish as more data becomes available for training (Li C. et al., 2022) and deep learning techniques continue to advance (Li Y. et al., 2019).

[0241] In this study, the focus was placed on AMPs that have potential utility in poultry farming for further validation. As a first step towards the field application, the selected AMPs were initially tested against E. coli and S. aureus, and found 13 to be active against atleast one of the bacterial strains tested. A total of 11 of the active peptides were relatively similar to chicken AMPs in their 3D structures (TM scores ranging from 0.6934 to 0.8352), 10 of which were from organisms other than chickens. Based on the fact that the biological function of a protein is strongly influenced by the 3D structure it adopts and chicken AMPs have evolved to fight the pathogens that infect them (Zhang and Gallo, 2016), these newly discovered AMPs sharing high structural similarities to chicken AMPs may target similar pathogens as chicken AMPs do. On the other hand, the 10 “exogenous” AMPs may have different evolutionary background than chicken AMPs, suggesting that many pathogens in the chickens’ living environments, especially those mostly found in chickens, may not have had enough opportunity to develop resistance against these newly discovered AMPs. As a result, the use of the aforementioned 10 AMPs as substitutes for conventional antibiotics to fight against pathogens in chicken farming can be foreseen. Although some peptides were not active against the bacterial strains tested, they may still be active against other pathogens, especially those that infect chickens. Further tests on a broader range of bacterial species could be done to validate their application as novel antibiotics for chickens.

[0242] Although it is hypothesized that AMPs may not induce antibiotic resistance to the extent of conventional antibiotics (Boman, 2003), cases of bacterial cross-resistance to multiple AMPs have been reported (Fleitas et al., 2016), raising a concern of the usage of these newly discovered AMPs. The similarity between a newly discovered AMP and a chicken AMP in 3D structures infers a potential similarity in their modes of action, indicating a potential risk that a newly discovered AMP may lose its effect if some pathogens have already developed resistance to its reference chicken AMP (Kintses et al., 2019). However, it is unclear how similar their modes of action are. Properties other than the peptide structure itself, such as the distribution of the hydrophobic and positively charged residues, can all affect how an AMP works. None of the described AMPs have a 100% match in 3D structures (TM-scores = 1) to the known chicken AMPs. Additionally, the structures of some AMPs may vary according to the surrounding microenvironment (Candido et al., 2019). All of these cannot be observed through the current set of tests, and warrant further experimentation. In the worst case, if cross-resistance to both of the newly discovered AMPs and chicken AMPs does exist in some pathogens, then this work sets up a warning to the scientific community in the use of those AMPs in chicken farming, though they can still be applied to other clinical scenarios.

[0243] Although the AMPs discovered in this study were derived from putative precursor sequences from natural organisms, it still requires further investigation whetherthose AMPs really exist in nature. Due to the limitations in the in silica cleavage tool, it is likely that some sequences, which may or may not be a real precursor, were incorrectly cleaved, even though the resulted fragments displayed antimicrobial properties against one or two bacterial strains tested in vitro. It has been reported that the degradation products of some non-antimicrobial proteins with other biological functions exhibit antimicrobial activities (Papareddy et all, 2010), and some of the aforementioned antimicrobial protein fragments may belong to this category of AMPs. Further, cases were noticed in which the predicted AMPs were derived from a non-secreted protein (e.g. GoGol (SEQ ID NO: 833) as the fragment of a transcriptional repressor), while those AMPs are less likely to exist in nature. However, regardless of whether a discovered AMP exists in nature or not, it can be considered as a valuable discovery because it provides more potential alternatives to conventional antibiotics. Those that do not exist in nature could have unique advantages in that they may be “new” to most extant pathogens and where, theoretically, no resistance has been developed to them. Lastly, it is also of interest to study the relationship between the newly discovered antimicrobial activities of those peptides and the originally annotated functions of their parent sequences in the database, which can help get a better understanding of the full life cycle of those proteins or peptides.

[0244] As the UniProt database frequently updates with more protein sequence data from additional organisms, it is expected that the AMP mining workflow presented herein will be of utility for some time. Moreover, future work can be extended to additional peptide sequence databases (e.g. UniProtKB / TrEMBL, the unreviewed section of the UniProt database (UniProt Consortium, 2019)). It is expected that in silico bioinformatics workflows like that described herein, will continuously unearth novel AMPs and to help win the battle against MDR bacteria.

[0245] The discovered AMPs have the potential to reduce the morbidity and mortality caused by bacterial infections (especially by antibiotic resistant strains or multidrug resistant strains), viral diseases in humans or non-human animals, and cancers that may be triggered by or relating to such bacterial infections. This direct relevance to public health makes the present bioinformatics approach and the resulting discoveries relevant to industrial players. The future applications of the discovery pipeline may expand into the analysis of genomic resources from other studies, including on species from environmental, agricultural, and forestry genomics studies, as well as from metagenomics surveys.

[0246] Developing AMPs into therapeutics can involve a variety of delivery modes and forms such as oral, injectable, rectal, topical, transdermal, nasal, or ocular delivery.Using a high throughput methodology to identify a large number of AMPs with therapeutic potential may permit the development of combination compositions. The antimicrobial peptides can be used in combination with more than one providing antimicrobial activity, or in a composition that provides a type of “cocktail” approach, together with other active ingredients.

[0247] Example 3

[0248] Predicting toxicity of antimicrobial peptides for drug design using toxicity prediction of antimicrobial peptides with representation learning (tAMPer).

[0249] Antimicrobial peptides (AMPs) are naturally occurring peptides produced by all multicellular organisms. These host defense peptides are actively being considered as alternative therapeutics in the fight against microbes due to the current antimicrobial resistance crisis. One of the biggest challenges in the path toward AMP based drugs is related to the potential toxicity of natural or synthetic peptides against mammalian cells. This motivates development of in silico methods that can inform the candidate prioritization and downstream peptide optimization, thus reducing costs and effort by filtering out the peptides that are potentially toxic.

[0250] In this example, a machine learning based pipeline that informs filtering decisions of AMPs by toxicity predictions based on peptide sequence only. We have curated publicly available data on natural and synthetic AMP sequences with annotations of hemolytic or cytotoxic activities. Using the dataset collected, a deep learning model (ELMo) is trained, as adapted by the SeqVec method, to obtain latent vector embeddings for our data. SeqVec is an unsupervised deep learning method that trains on a large corpus of proteins (UniRef50) and learns internal representations of them. Vector representations or latent features of a AMP dataset is input on a 3-hidden-layer feed-forward neural network with tanh activation.

[0251] These embeddings contain information that encodes for peptide toxicity. This classification model was able to correctly predict toxicity on the test data with an accuracy of 95%. When projecting the embeddings to a lower dimensional space using Principal Component Analysis, specific components which encode the toxicity of the sequence embeddings can be pinpointed. It was found that the second principal component correlates with the hemolysis or cytotoxicity values of the AMPs in these data with a Pearson correlation coefficient of 0.46. This suggests that the pipeline can be used to predict thetoxicity of antimicrobial peptides, as well as providing a design platform for nontoxic antimicrobial peptides.

[0252] The typical solution to determine the level of toxicity for AMPs is through wet lab testing. The strategy used herein provides a method to prioritize AMPs for downstream applications in drug development.

[0253] Structure-aware deep learning models for toxicity prediction of peptides are described herein. Antimicrobial resistance is a critical public health concern, necessitating the exploration of alternative treatments to conventional antibiotics. Antimicrobial peptides (AMPs) have emerged as a promising avenue to explore such alternatives. However, assessing their toxicity through wet lab methods is time-consuming and costly.

[0254] Computational tools that accurately predict peptide toxicity may offer a solution by enabling the rapid screening of candidate AMPs. In response to this need, we introduce tAMPer, a deep learning model that predicts AMP toxicity by integrating the underlying sequence composition and its predicted three-dimensional (3D) structure. tAMPer adopts a graph-based representation for peptides, modeling them as graphs encoding AlphaFold2-predicted structures. In these graphs, nodes correspond to amino acids, and edges represent spatial interactions. The model extracts structural features using graph neural networks, and employs recurrent neural networks to capture sequential dependencies. tAM Per's performance was assessed on both publicly available protein toxicity benchmark dataset and in-house AMP hemolysis data. The method achieved a 68.7% F1-score, surpassing the second-best method's score of 45.3%. On the protein benchmark dataset, tAMPer exhibited over 3.0% improvement in F1 -score compared to current state-of-the-art methods. This work highlights the potential of 3D peptide structure predictions and graph neural networks in developing safer AMP therapeutics to combat antimicrobial resistance.

[0255] In recent years, there has been a growing focus on scaling up in silico AMP discovery pipelines using computational methods for prediction and de novo design of AMPs. These computational approaches generate a large pool of potential AMP candidates, which must be evaluated further. One crucial aspect of this evaluation is assessing their toxicity towards host cells. While there is limited understanding of the underlying mechanisms of AMP toxicity, several studies have suggested exploring membrane interactions as an explanatory factor, underscoring the importance of understanding AMP structures in this context. Traditionally, the first step in evaluating AMP toxicity is to measure their hemolytic (toxicity against red blood cells) or broader cytotoxic activity. Computationally identifyingpotentially toxic AMP candidates prior to conducting such wet-lab experiments would greatly streamline this process by filtering out potentially harmful AMPs. Such computational tools would enable prioritization of the most promising candidates for further experimental validation.

[0256] Computational methods for predicting toxic peptides can be broadly categorized into two groups: conventional bioinformatics tools and machine learning (ML) models. Conventional tools rely on similarity or homology-based searches. For instance, BLAST and BLAST-score classify a protein sequence as toxic if it shares similarity with known toxic sequences, determined by a threshold applied to the E-value. InterProScan and HmmSearch, on the other hand, detect toxic protein domains and categorize a sequence as toxic if it possesses or is associated with these domains. In contrast, ML models are trained to predict toxic peptides based on their specific characteristics. ML models have demonstrated notably higher predictive performance in toxicity assessment compared to conventional bioinformatics tools.

[0257] ML-based methods are optimized to learn discriminative features extracted from peptide sequences to distinguish between toxic and non-toxic peptides. While various methods, such as ToxinPred, HemoPl, and HAPPENN, have been in use for the task, they typically rely on hand-crafted sequence-derived or physicochemical features as inputs to their models. This limits their ability explore the input data for potentially more informative features for toxicity prediction. Furthermore, these methods often overlook the inherent order present in input peptide sequences. Some existing methods, such as Toxify, ToxDL, and ToxinPred2 have primarily been trained on long protein sequences and are not specifically designed for peptide sequences. ATSE incorporates position-specific scoring matrices (PSSMs) and molecular graphs for toxicity prediction. However, for generating PSSMs it relies on the PSI-BLAST algorithm, which is both database-dependent and time-consuming to run. ToxlBTL utilizes transfer learning by leveraging knowledge learned from protein toxicity for improved peptide toxicity prediction. Importantly, none of the cited methods above exploit the potential of 3D structures of peptides, which can be a valuable source of information for toxicity prediction.

[0258] Here, a multi-modal structure-aware deep learning model designed for predicting peptide toxicity. Initial embeddings are generated for peptide sequences, capturing higher-level amino acid representations. 3D peptide structure information represented as graphs is then incorporated, where residues are nodes and interactions are edges. This model architecture jointly learns from both sequential and structural features for in silicotoxicity prediction. A curated training and validation sets of peptides is thus created, with wetlab validated toxicity. The described method holds promise for advancing the field of peptide toxicity prediction and may contribute to the development of safer and more effective AMPs.

[0259] MATERIAL AND METHODS

[0260] Data collection

[0261] Training and validation sets. A dataset was compiled by aggregating relevant peptide sequences from several manually curated databases, including DBAASP v3 (Pirtskhalava et al., 2021), hemolytik (Gautam et al., 2014), APD3 (Wang G. et al., 2016) and UniProtKB / Swiss-Prot (accessed on Jan 2023). To ensure the quality and relevance of the dataset, peptide sequences were included that were 5 to 50 residues in length containing only natural amino acids, and redundant entries were filtered out. Notably, peptide sequences were incorporated with C-terminal amidation, as this post-translational modification can impact toxic activity of peptides. This modification was accommodated in the model.

[0262] From sequences obtained from the DBAASP and hemolytik databases, which provided experimental data on hemolytic activities, rigorous thresholds were utilized to classify them as either hemolytic (toxic) or non-hemolytic (non-toxic), as described in Table 12. Sequences that failed these criteria were excluded, resulting in 1 ,604 toxic and 4,042 non-toxic peptides. To supplement this dataset, 104 non-redundant sequences from the APD3 database with validated hemolytic activity were included as positive samples. From the Swiss-Prot database, sequences associated with the keyword “hemolysis” were downloaded. Mature peptide sub-sequences were retained if the search results contained peptide annotations. If these annotations were not available, the sub-sequences that had been annotated as the main chain in the sequence were considered. In cases where neither of these options was present, the entire sequence was included. Collected sequences were labelled as non-toxic if the function description included following keywords: “No hemolysis”, “low hemolysis”, or “weak hemolysis”, and as toxic otherwise. These steps yielded 221 hemolytic and 5 non-hemolytic new sequences from the Swiss-Prot database.

[0263] Table 13 shows the determination of toxicity for 206 peptides.

[0264] The final non-redundant dataset comprised 1,929 hemolytic and 4,047 nonhemolytic peptide sequences. This dataset was split into training and validation sets with a ratio of 4:1 for each class, respectively. To ensure that the validation set was representative of the overall distribution of peptide sequences in the dataset and to remove potential bias of learning sequence similarity, CD-HIT was used to select validation samples that share less than 80% sequence similarity with the training sequences. This approach helps to ensure that the model will generalize well to new peptide sequences.

[0265] AMP hemolysis dataset. To comprehensively evaluate the performance of tAMPer and compare it with other existing methods, an independent set of 340 peptide sequences was created the hemolytic activity of which was assessed in vitro. Whole blood was obtained from healthy donor pigs (Lampire Biological Laboratories; Pipersville, PA, USA) and the red blood cells (RBCs) isolated through centrifugation and washing with Roswell Park Memorial Institute medium (RPMI; Thermo Fisher Scientific, MA, USA), to prepare a 1% solution (v / v) in RPMI. Lyophilized AMPs (GenScript; Piscataway, NJ, USA) were suspendedand serially diluted in RPMI from 1280 down to 10 pg / mL in 96-well polypropylene plates (Greiner Bio-One; Kremsmunster, Austria) and combined with 100pL of the 1% RBC solution. Following an incubation at 37°C for 30-45 minutes, plates underwent centrifugation, and half of the supernatants were transferred to new 96-well plates. Absorbance was measured at 415 nm utilizing the Cytation 5 Cell Imaging Multimode Reader (BioTek, CA, USA), with the AMP concentration causing 50% hemolysis of RBCs (HC50) serving as the hemolytic activity indicator. Absorbance readings from wells containing RBCs treated with 11 pL of a 2% Triton-X100 solution and RPMI (AMP solvent-only), established the baseline of 100% and 0% hemolysis, respectively. All centrifugation steps were done with the Allegra-6R centrifuge (Beckman Coulter, CA, USA) at 500 x g for five minutes. To remove biases and confounding factors in experiments, each hemolysis assessment was performed in technical duplicates and repeated three times (N=3). A peptide was labelled as toxic if it showed an HC50 value of less than or equal to 128 pg / mL in at least two experiments. This approach helped to ensure that the hemolytic peptides were consistently active over multiple trials and reduced the likelihood of false positives due to experimental variability. An independent test set comprising 56 hemolytic and 284 non-hemolytic peptides was tested for HC50 (ug / ML) values, forming an AMP hemolysis dataset with peptide sequence and HC50 information. Each hemolysis assessment was performed in triplicate (N=3) to determine the concentration required for 50% hemolysis of red blood cells (HC50). A peptide was deemed toxic if it exhibited an HC50 value of less than or equal to 128 pg / mL in at least two of the three experiments.

[0266] Three-dimensional structure prediction. Version 1.3.1 of local ColabFold was utilized to predict the 3D structures of peptides. ColabFold uses AlphaFold2 (AF2) models (Jumper et al., 2021) with Mmseqs2 homology search for faster structure predictions (Mirdita et al., 2022). To run ColabFold locally, AF2 model inputs were obtained, including multiple sequence alignments and structural templates, from ColabFold’s server. Five structures were generated, corresponding to the five AF2 sub-models, each initialized using a different random seed. The output structures are ranked based on their average pLDDT score, which represents the per-residue confidence of AF2 in its predictions.

[0267] Data augmentation. To augment the collected data, the structure with the highest average pLDDT score was selected, as is customary, and also the other four structures were considered. This inclusion of additional structures functions as a form of data augmentation, a technique known to enhance the performance and robustness of deep learning models. Specifically, in tAMPer, it makes the model more robust against variationsin predicted structures, simulating different possible conformations of peptide structures. For a given peptide at test time, the toxicity probability for each structure is predicted, and the final output is the average probability across the five predicted structures.

[0268] tAMPer model

[0269] The tAMPer model is comprised of three main components: (i) Bi-directional Gated Recurrent Units (Bi-GRUs) for processing sequences, (ii) Graph Neural Networks (GNNs) (Scarselli et al., 2009) for extracting structural patterns from graph-encoded structures, and (iii) a self-attention layer for integrating sequential and structural feature vectors for each amino acid residue. Each component plays a specific role within the tAMPer model. The Bi-GRUs capture the sequential dependencies in the peptide sequences. The GNNs extract structural features by performing message passing between nodes, aggregating information from their neighboring nodes and edges. The self-attention layer combines the sequential and structural feature vectors for each amino acid, enhancing their representation and capturing potential relationships between residues. This also allows the model to focus on important features and weight the contributions of residues within each peptide sequence. By jointly learning from both sequential and structural data, tAMPer minimizes the error in predictions of toxicity and secondary structure during the training phase.

[0270] To represent peptide sequences numerically as vectors, sequence embeddings were obtained from protein language models (PLMs) trained on millions of raw protein sequences. In this particular work, the ESM2 (evolutionary-scale prediction of atomic- level) protein structure PLM was utilized to generate initial embeddings for amino acids in peptides.

[0271] A sequence processing module was used. Bi-GRUs in tAMPer capture the sequential dependencies between the amino acid residues in peptide sequences. Bi-GRUs are able to do this by maintaining a hidden state that represents the "memory" of the network, allowing it to encode and retain information about previous residues as it processes the current residue. The bi-directional nature of the Bi-GRUs allows for processing the sequence both forwards and backwards, enabling it to capture dependencies in both directions. Bi-GRUs transform the input embeddings to represent the extracted sequential features for each amino acid in model’s dimensionality.

[0272] In tAMPer, the 3D structure of a peptide are represented as a graph encoding the interactions between residues. To determine whether two residues should be connected by an edge, the distance between their atoms in 3D space is evaluated, and if this distance isless than a certain threshold, an edge is established between the residues. This criterion indicates that the two amino acids are in contact or in close proximity to each other. One of the advantages of this graph construction approach is its flexibility in incorporating additional features for nodes and edges, including backbone structural characteristics, physicochemical features, or specific attributes related to the interactions between residues. This would result in a graph representation that is more comprehensive and informative.

[0273] Other aspects of the tAMPer model are generally described herein. In the structure processing module, GNNs are utilized to capture structural patterns in graph- encoded peptides. GNNs operate based on a message passing paradigm which includes the three steps of: featurization, aggregation, and updating of node representations. During featurization, the relevant numerical features are encoded on the nodes and edges of the graphs. In the aggregation step, each node collects information from its neighboring edges and nodes. Finally, each node updates its feature vector based on received messages from aggregation step and its previous state. This process of aggregation and updating can be repeated multiple times through layers of GNNs to iteratively refine the node representations.

[0274] The GNNs component was pre-trained on a reverse-folding task with the predicted structures as well as structures of sequences with less 100 amino acids length available from AlphaFold2 database (Varadi et al., 2022). The objective is to predict the corresponding amino acid type for each node within featurized graphs. Each node is assigned a label from the set, representing the specific natural amino acid it corresponds to. Throughout the training process, the model utilizes the initial scalar and vector features present in the graphs as inputs to predict these labels. Given that the features do not explicitly convey information about the amino acids themselves, GNNs rely on the structural features and the conformation of the graph to predict the amino acids.

[0275] Integration of sequential and structural features. The initial representation of peptides is formed by concatenating the extracted sequential and structural features for each amino acid residue obtained from Bi-GRUs and GNNs. To avoid overreliance on a single source of information during training, a two-dimensional dropout with p=0.05 probability was applied separately to the sequential and structural features before their concatenation.

[0276] tAMPer utilizes an 8-headed self-attention layer to identify potential relationships between amino acids contributing to toxicity, and to combine the sequential and structural features in a same space. Attention weights are used to provide interpretability fortAMPer predictions, enabling an assessment of which residues or combinations of them play substantial role in making a peptide toxic.

[0277] For toxicity prediction, the final representations of peptides is obtained, and a single binary feature is appended to indicate whether the peptide is amidated or not. The augmented vector obtained is then fed into a fully connected layer, followed by a softmax activation function. The softmax function outputs a probability representing the predicted likelihood of toxicity. The input peptide was classified as “toxic” if the output probability is greater than 0.5, and as “non-toxic” otherwise. The Binary Cross Entropy loss function is employed to calculate the error between the predicted probability and the true label of toxicity.

[0278] An additional loss term is introduced to utilize the combined features to predict the secondary structure of each amino acid individually. To predict the secondary structure, one layer of fully connected network is used to map the combined feature to a dimension in which classes of secondary structures are represented. The cross-entropy loss function is then applied to calculate the error between the predicted secondary structure and the true labels. The final loss for the peptide is a linear combination of the toxicity loss and the secondary structure loss, with a hyperparameter controlling the weighting between the two losses. Hyperparameter tuning was conducted on the validation set.

[0279] RESULTS

[0280] AMP hemolysis test set

[0281] The performance of tAMPer and other tools on a real-world data set of peptides was evaluated. While the Toxify method (Cole et al., 2019) allowed us to re-train their model on our dataset, none of the other methods provided this option. Therefore, we used the original models of these methods to predict toxicity and compared the results against tAMPer. For the ToxlBTL model, predictions were obtained from the online server (server.wei-group.net). Similarly, predictions for the ToxinPred (crdd.osdd.net), ToxDL (csbio.sjtu.edu.cn), HAPPENN (research. timmons.eu), and HemoPl (webs. iiitd.edu. in) models were obtained using the default configurations from their online servers. For Toxify, the model was downloaded from the Toxify GitHub repository (github.com) and re-trained on our collected data using the default hyperparameters provided by the authors.

[0282] Table 14 shows that tAMPer outperforms the other competing methods, achieving an F1 score of 68.66%, MCC of 62.65%, auROC of 91.69%, and auPRC of 68.98%. Specifically, with tAMPer, a 51.6% higher F1 score was observed (i.e. , a 23.4% net increase), 23.30% in MCC, 7.79% in auROC, and 10.75% in auPRC over the second-bestmethod in each metric. Additionally, the performance of all tools was analysed using Receiver Operating Characteristic (ROC) curves, which revealed the true positive rate versus the false positive rate at different classification thresholds. A ROC curve comparing tAMPer and other toxicity prediction methods on the AMP hemolysis dataset revealed that when tAMPer is compared to original Toxify method; the re-trained model of Toxify using the present collected data; HAPPEN N; HemoPl; ToxlBTL; and ToxDL, true positive rate (sensitivity) against the false positive rate (1 -specificity) at different classification thresholds found that tAMPer achieved the highest area under the ROC curve (auROC) 91.69%, indicating a higher discriminatory ability in distinguishing between toxic and non-toxic peptides.

[0283] Correlation between tAMPer’s toxicity probability and HC50. We examined the correlation between the toxicity probabilities predicted by tAMPer and logarithmic transformation (base 2) of the HC50 value measured in pg / ml of our test peptides. The relationship between the toxicity probabilities predicted by tAMPer to the HC50 of the test peptides was observed. An inverse correlation was found between the HC50 values and the predicted toxicity probabilities. As the HC50 values increased, indicating a higher concentration required for hemolysis, the predicted toxicity probabilities decreased, indicating a lower likelihood of peptide toxicity. The fitted curve exhibited a gradual decline intoxicity probability for peptides within the toxic HC50 range of 8 to 128 pg / ml, while mostly remaining above a probability of 0.5. Meanwhile, peptides with HC50 values exceeding 128 pg / ml (non-toxic) displayed a probability of less than 0.3. notably, the model was not provided with HC50 values during the training process, highlighting the ability of tAMPer to infer the relationship between peptide toxicity and HC50 values without direct knowledge of these concentrations. This correlation aligns with the desired behavior of a reliable toxicity prediction model.

[0284] Protein benchmark

[0285] To ensure an unbiased evaluation of tAMPer and to avoid any potential biases from relying solely on curated data, additional training and testing was conducted using the dataset established by the toxDL method (Pan et al., 2021). The use of this benchmark dataset is widespread in the literature, and thus permits impartial and objective evaluation of tAMPer's performance. The toxDL training set consists of 4,472 neurotoxic and 6,341 nontoxic animal protein sequences, respectively. In the test set, each sequence has less than 40% sequence identity to any sequence in the training set. Moreover, to ensure diversity, none of the sequences in the training and test sets belong to the same Pfam clans.

[0286] In this study, and in line with other competitors, the F1 score, MCC, auROC, and auPRC were evaluated. To maintain fairness in comparison with other competitors, there was no augmentation of the data with supplementary structures. Instead, the structure with the highest average PLDDT score for each sequence was utilized. To accommodate longer sequences and larger structures in this dataset, specific modifications were made to the tAMPer model, utilizing t33 embeddings from the ESM2 model, providing a more suitable representation for the longer sequences. Additionally, the number of Graph Neural Networks (GNNs) layers was increased to 3, in order to capture more complex structural patterns. Furthermore, the dimensionality of the model was augmented to 128 to enhance its capacity to learn from the larger structures.

[0287] The performance of tAMPer was compared with other methods that were tested on the protein benchmark dataset, including ToxDL. The methods applied include BLAST and BLAST-score (Altschul, 1997), InterProScan (Quevillon et al., 2005), HmmSearch, (Potter et al. 2018), ClanTox (Naamati et al., 2009), ToxinPred (Gupta et al., 2013), Toxify (Cole et al., 2019), ToxDL (Pan et al., 2021), and ToxlBTL (Wei et al., 2022). A comparison of tAMPer with these existing methods revealed that tAMPer outperforms across all measured metrics, achieving an F1 score of 86.0%, MCC of 85.0%, auROC of 99.2%, and auPRC of 91.6%. More specifically, tAMPer achieves a net 3.0% higher F1 score, 3.4%higher MCC, 0.3% higher auROC, and a 0.3% higher auPRC over ToxlBTL, which was the next best method for protein toxicity prediction.

[0288] DISCUSSION

[0289] tAMPer uses both sequential and structural data to predict the toxicity of peptides, as the functional attributes of AMPs are assumed to originate not only from their amino acid composition but also from their structural characteristics (Fjell et al., 2012; Hollmann et al., 2016; DeGrado et al., 1982). The model is built upon amino acid embeddings generated by ESM2 and AlphaFold2 predicted structures.

[0290] To ensure a balanced integration of sequential and structural features of proteins, particularly considering the PLM-generated nature of sequential features, deliberate steps were taken to ensure the contribution of structural features to predictions. tAM Per’s performance on the independent AMP hemolysis test set, which is representative of real- world experimental data, highlights its robustness. In the evaluation of sensitivity and specificity, certain methods exhibited a trade-off, leaning towards either excessively high sensitivity or specificity in their predictions. In contrast, tAMPer achieved the most balanced results overall, with 82.1% and 88.7% prediction sensitivity and specificity, respectively. Further, the inverse relationship between tAMPer's predicted probability of toxicity and the HC50 values of peptides facilitates the prioritization and selection of AMP candidates for subsequent experimental testing.

[0291] The performance of deep learning models often improves with an increase in the amount of available training data. However, in the field of peptide toxicity prediction, it is important to acknowledge that the scale of the training data used is relatively limited compared to domains such as computer vision or natural language processing. There are several factors that contribute to the limited availability of training data for peptide toxicity prediction tasks. Experimental determination of peptide toxicity is a time-consuming, costly, and labor-intensive process. This leads to a scarcity of well-annotated peptide datasets with reliable toxicity labels. Gathering comprehensive and high-quality toxicity data requires significant resources and expertise. Furthermore, the mechanisms of toxicity and the complete toxicology profiles of peptides are still not fully understood. Peptides can exhibit diverse modes of toxicity, and their toxicological properties can vary depending on the target organism or cell type. The lack of comprehensive knowledge about the mechanisms underlying peptide toxicity poses challenges in data collection and model development. As the field progresses and more research is conducted, more well-annotated peptide sequences datasets will become available, and researchers will gain a deeper understandingof peptide toxicity mechanisms. This will contribute to improved model performance and a better understanding of peptide toxicity in various biological contexts.

[0292] This model offers effective protein and peptide toxicity prediction, including evaluation of antimicrobial peptides, offering valuable insights for the development of more effective peptide-based therapeutics to address the pressing challenge of antimicrobial resistance. Its ability to predict toxicity profiles with high accuracy can streamline the screening and design of antimicrobial peptides, facilitating identification of candidates with improved safety profiles. By reducing the reliance on labor-intensive and costly wet lab experiments, tAMPer not only accelerates the discovery process but also contributes to substantial cost savings.

[0293] Example 4

[0294] Efficacy and toxicity testing of priority antimicrobial peptides.

[0295] In this example, 918 peptides were tested against 24 bacterial strains. Those peptide (515) showing low or no activity in the bacterial strains tested were not prioritized for further testing. Of the 403 remaining peptide that showed activity in at least some bacterial strains, 59 showed some toxicity and were not prioritized for further testing. There were 314 active peptides remaining showing no toxicity, while 30 required further toxicity testing. Of these, 334 are presented below as priority antimicrobial peptides. These peptides are listed in Table 15, which indicates the AMP by name (according to the specified naming strategy), sequence, length and charge.

[0296] The toxicity testing was conducted against Porcine RBC HC50 and HEK293 CC50. The testing with pRBC for toxicity was evaluated, and it passed, the question of whether both parameters of eukaryotic toxicity were tested and passed. Evaluation of minimum inhibitory concentration (MIC) for antimicrobial efficacy was deemed adequate in all peptides tested, in Table 15. Eukaryotic toxicity and antimicrobial efficacy (MIC) testing was conducted as outlined in Example 3, and results are provided in Tables 16-27 below, with a summary provided in Table 28.

[0297] The microbes against which the peptides were tested were as follows: Escherichia coli ATCC 25922; Avian Pathogenic Escherichia coli 317; Escherichia coli ATCC BAA-2469; Escherichia coli ATCC BAA-2340; Escherichia coli ATCC BAA-2471; Escherichia coli CPO-NDM (BCCDC clinical isolate); Escherichia coli CPO-NDM (BCCDC clinical isolate); Escherichia coli ESBL (BCCDC clinical isolate); Escherichia coli ESBL (BCCDC clinical isolate); Escherichia coli MCR-1 (BCCDC clinical isolate); Escherichia coli MCR-2 (BCCDC clinical isolate); Escherichia coli MCR / NDM (BCCDC clinical isolate);Escherichia coli MCR / NDM (BCCDC clinical isolate); Salmonella enterica serovar Enteritidis ATCC 4931; Salmonella enterica serovar Enteritidis LS101 ; Salmonella enterica serovar Enteritidis (BCCDC clinical isolate); Salmonella enterica serovar Heidelberg (BCCDC clinical isolate); Pseudomonas aeruginosa ATCC 27853; Pseudomonas aeruginosa NDM / VIM (BCCDC clinical isolate); Acinetobacter baumannii ATCC 2208; Acinetobacter baumannii ATCC 19606; Acinetobacter baumannii NDM / OXA-51 (BCCDC clinical isolate); Klebsiella pneumoniae ATCC 10031 ; Klebsiella pneumoniae KPC / NDM (BCCDC clinical isolate);Enterococcus faecalis ATCC 29212; and Enterococcus faecalis vanA (BCCDC clinical isolate).

[0298] RESULTS

[0299] Tables 16-27 show results of toxicology testing (consistent with Example 3) and MIC analysis for various microorganism, consistent with above-noted methodologies. In all tables, median (MED) values are presented, together with median absolute deviation (MAD) as a measure of variability. Table 28 provides a summary of results of toxicity and antimicrobial activity (with MIC provided at a 32 g / L cut off for the MIC “pass” parameter.

[0300] Table 16 shows the antimicrobial peptides, presented consistently with the order and SEQ ID NOs of Table 15 tested for toxicity against to porcine RBC, and providing the parameter of halfmaximal concentrations to cause hemolysis (H50) together with n of tests and median absolute deviation (MAD) as a measure of variability.

[0301] The data of Table 17 illustrate toxicity testing of select antimicrobial peptides using HEK293 cells to assess the half maximal cytotoxic concentration at which each peptide is cytotoxic to 50% of a the population of HEK293 cells (CC50).

[0302] The data presented in Table 18 illustrate efficacy against Escherichia coli ATCC 25922 based on minimum inhibitory concentrations (MIC) as described above.

[0303] The data presented in Table 19 illustrate efficacy against Avian Pathogenic Escherichia coli 317 based on minimum inhibitory concentrations (MIC) as described above.

[0304] The data presented in Table 20 illustrate efficacy against three different Escherichia coli ATCC BAA-2469, ATCC BAA-2340, and ATCC BAA-2471 , showing data based on minimum inhibitory concentrations (MIC) as described above for select peptides of Table 15.

[0305] The data presented in Table 21 illustrate efficacy against Escherichia coli Escherichia coli CPO-NDM (BCCDC clinical isolate), Escherichia coli CPO-KPC (BCCDC clinical isolate), and Escherichia coli ESBL (BCCDC clinical isolate) showing data based on minimum inhibitory concentrations (MIC) as described above for select peptides of Table 15.

[0306] The data presented in Table 22 illustrate efficacy against different Escherichia coli clinical isolates, namely: Escherichia coli MCR-1 (BCCDC clinical isolate), Escherichiacoli MCR-2 (BCCDC clinical isolate), and Escherichia coli MCR / NDM (BCCDC clinical isolate), showing data based on minimum inhibitory concentrations (MIC) as described above for select peptides of Table 15.

[0307] The data presented in Table 23 illustrate efficacy against Salmonella enterica, specifically: Salmonella enterica serovar Enteritidis ATCC 4931 , Salmonella enterica serovar Enteritidis LS101, and Salmonella enterica serovar Enteritidis (BCCDC clinical isolate), showing data based on minimum inhibitory concentrations (MIC) as described above for select peptides of Table 15.

[0308] The data presented in Table 24 illustrate efficacy against different microbes, specifically: Salmonella enterica serovar Heidelberg (BCCDC clinical isolate), Pseudomonas aeruginosa ATCC 27853, Pseudomonas aeruginosa NDM / VIM (BCCDC clinical isolate), showing data based on minimum inhibitory concentrations (MIC) as described above for select peptides of Table 15.

[0309] The data presented in Table 25 illustrate efficacy against different Acinetobacter microbes, specifically: Acinetobacter baumannii ATCC 2208, Acinetobacter baumannii ATCC 19606, and Acinetobacter baumannii NDM / OXA-51 (BCCDC clinical isolate), showing data based on minimum inhibitory concentrations (MIC) as described above for select peptides of Table 15.

[0310] The data presented in Table 26 illustrate efficacy against Klebsiella pneumoniae, in particular: Klebsiella pneumoniae ATCC 10031 and Klebsiella pneumoniae KPC / NDM (BCCDC clinical isolate), showing data based on minimum inhibitory concentrations (MIC) as described above for select peptides of Table 15.

[0311] The data presented in Table 27 illustrate efficacy against various microbes, specifically: Enterococcus faecalis ATCC 29212, Enterococcus faecalis vanA (BCCDC clinical isolate), and Staphylococcus aureus ATCC 29213, showing data based on minimum inhibitory concentrations (MIC) as described above for select peptides of Table 15.

[0312] The data presented in Table 28 summarize the toxicity parameters and the efficacy parameters evaluated for each of peptides of Table 15, as presented in Tables 16 - 27, herein.

[0313] These data illustrate 334 priority antimicrobial peptides having efficacy against a variety of microbes as tested, as well as promising non-toxic results: SEQ ID NOs: 1-3, 6-7, 9-10, 12, 18-19, 21-22, 26-27, 31 , 34, 49, 51 , 91 , 187, 285, 483, 896, 1218, 1238, 1399, 1555, 1802, 1978, 2046, 2425, 2806, 3363, 3568, 4668, 8078, 8082-8084, 8087-8089, 8091- 8092, 8093-8096, 8097- 8100, 8102-8103, 8106-8113, 8115-8117, 8118-8125, 8126-8132, 8141-8143, 8148-8150, 8153, 8155, 8157-8158, 8162-8165, 8172, 8175-8199, 8201-8206, 8208, 8209-8211, 8213-8220, 8222-8226, 8228, 8230-8242, 8244, 8263-8265, 8245-8254, 8256, 8258-8262, 8266, 8270-8272, 8275, 8281 and 8282-8433. For example, peptides according to SEQ ID NOs: 1-3, 6-7, 9-10, 12, 18-19, 21-22, 26-27, 31, 34, 49, 51 , or 8282- 8433 are effective antimicrobial peptides for applications as described herein or for applications to which new antimicrobial peptides can be beneficially applied. Peptides having SEQ ID NOs:1-3, 7, 10, 18, 22, 8301 , 8304-8306, 8312, 8319, 8323-8324, 8336, 8343, 8366-8369, 8371, 8380-8381 ,8387-8389, 8392, 8397, 8398, 8400-8402, 8406, 8416, and 8420-8422 (among others) illustrated both low toxicity and high efficacy. Such peptideshaving an agreeable safety profile of low toxicity can be used in antimicrobial applications as described herein.

[0314] In the preceding description, for purposes of explanation, numerous details are set forth in order to provide a thorough understanding of the embodiments. However, it will be apparent to one skilled in the art that these specific details are not required.

[0315] The above-described embodiments are intended to be examples only. Alterations, modifications and variations can be effected to the particular embodiments by those of skill in the art. The scope of the claims should not be limited by the particular embodiments set forth herein, but should be construed in a manner consistent with the specification as a whole.

[0316] References

[0317] The following publications are incorporated by reference herein.WO 2020 / 118427 A1 - Antimicrobial Peptides (Bird et al.)Adamczak R., Porollo A., Meller J. Combining Prediction of Secondary Structure and Solvent Accessibility in Proteins. Proteins. 2005;59:467-475. doi: 10.1002 / prot.20441.Aghapour.Z., Gholizadeh.P., Ganbarov.K., Bialvaei.A.Z., Mahmood, S.S., Tanomand.A., Yousefi.M., Asgharzadeh.M., Yousefi.B. and Samadi Kafil.H. (2019) Molecular mechanisms related to colistin resistance in Enterobacteriaceae. Infect. Drug Resist., 12, 965-975.Altschul.S. (1997) Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res., 25, 3389-3402.Amaral A.C., Silva O.N., Mundim N.C.C.R., de Carvalho M.J.A., Migliolo L., Leite J.R.S.A., Prates M.V., Bocca A.L., Franco O.L., Felipe M.S.S. Predicting Antimicrobial Peptides from Eukaryotic Genomes: In Silico Strategies to Develop Antibiotics. Peptides. 2012;37:301-308. doi: 10.1016 / j. peptides.2012.07.021.Andersson D.I., Hughes D., Kubicek-Sutherland J.Z. Mechanisms and Consequences of Bacterial Resistance to Antimicrobial Peptides. Drug Resist. Updat. 2016;26:43-57. doi: 10.1016 / j. drup.2016.04.002.Apata.D.F. (2009) Antibiotic Resistance in Poultry. Int. J. Poult. Sci., 8, 404-408.Aronica P.G.A., Reid L.M., Desai N., Li J., Fox S.J., Yadahalli S., Essex J.W., Verma C.S. Computational Methods and Tools in Antimicrobial Peptide Research. J. Chem. Inf. Model. 2021 ;61:3172-3196. doi: 10.1021 / acs.jcim.1c00175.Arvidson R., Kaiser M., Lee S.S., llrenda J.-P., Dail C., Mohammed H., Nolan C., Pan S., Stajich J.E., Libersat F., et al. Parasitoid Jewel Wasp Mounts Multipronged Neurochemical Attack to Hijack a Host Brain. Mol. Cell. Proteom. 2019;18:99-114. doi: 10.1074 / mcp.RA118.000908.Bals.R. (2000) Epithelial antimicrobial peptides in host defense against infection. Respir. Res., 1 , 5.Barbosa-Morais N.L., Irimia M., Pan Q., Xiong H.Y., Gueroussov S., Lee L.J., Slobodeniuc V., Kutter C., Watt S., Colak R., et al. The Evolutionary Landscape of Alternative Splicing in Vertebrate Species. Science. 2012;338:1587-1593. doi: 10.1126 / science.1230612.Becchimanzi A., Avolio M., Bostan H., Colantuono C., Cozzolino F., Mancini D., Chiusano M.L., Pucci P., Caccia S., Pennacchio F. Venomics of the Ectoparasitoid Wasp Bracon Nigricans. BMC Genom. 2020;21 :34. doi: 10.1186 / s12864-019-6396-4.Beckloff.N. and Diamond, G. (2008) Computational Analysis Suggests Beta-Defensins Are Processed to Mature Peptides By Signal Peptidase. Protein Pept. Lett., 15, 536-540.Bird I., Behsaz B., Hammond S.A., Kucuk E., Veldhoen N., Helbing C.C. De Novo Transcriptome Assemblies of Rana (Lithobates) Catesbeiana and Xenopus Laevis Tadpole Livers for Comparative Genomics without Reference Genomes. PLoS ONE. 2015;10:e0130720. doi: 10.1371 / journal.pone.0130720.Boman, H.G. (2003) Antibacterial peptides: basic facts and emerging concepts. J. Intern. Med., 254, 197-215.Bossuyt F., Schulte L.M., Maex M., Janssenswillen S., Novikova P.Y., Biju S.D., Van de Peer Y., Matthijs S., Roelants K., Martel A., et al. Multiple Independent Recruitment of Sodefrin Precursor-Like Factors in Anuran Sexually Dimorphic Glands. Mol. Biol.Evol. 2019;36:1921-1930. doi: 10.1093 / molbev / msz115.Bouzid W., Verdenaud M., Klopp C., Ducancel F., Noirot C., Vetillard A. De Novo Sequencing and Transcriptome Analysis for Tetramorium Bicarinatum: A Comprehensive Venom Gland Transcriptome Analysis from an Ant Species. BMC Genom. 2014;15:987. doi: 10.1186 / 1471-2164-15-987.Brandenburg K., Heinbockel L., Correa W., Lohner K. Peptides with Dual Mode of Action: Killing Bacteria and Preventing Endotoxin-Induced Sepsis. Biochim. Biophys. Acta BBA-Biomembr. 2016;1858:971-979. doi: 10.1016 / j.bbamem.2016.01.011.Burdukiewicz.M., Sidorczuk.K., Rafacz.D., Pietluch.F., Chilimoniuk.J., Rddiger.S. and Gagat.P. (2020) Proteomic Screening for Prediction and Design of Antimicrobial Peptides with AmpGram. Int. J. Mol. Sci., 21 , 4310.Burke G.R., Strand M.R. Systematic Analysis of a Wasp Parasitism Arsenal. Mol. Ecol. 2014;23:890-901. doi: 10.1111 / mec.12648.Candido, E.S., Cardoso.M.H., Chan.L.Y., Torres, M.D.T., Oshiro, K.G.N., Porto, W.F., Ribeiro, S.M., Haney, E.F., Hancock, R.E.W., Lu,T.K., et al. (2019) Short CationicPeptide Derived from Archaea with Dual Antibacterial Properties and Anti- Infective Potential. ACS Infect. Dis., 5, 1081-1086.Cao J., de la Fuente-Nunez C., Ou R.W., Torres M.D.T., Pande S.G., Sinskey A. J., Lu T.K. Yeast-Based Synthetic Biology Platform for Antimicrobial Peptide Production. ACS Synth. Biol. 2018;7:896-902. doi: 10.1021 / acssynbio.7b00396.Caty S.N., Alvarez-Buylla A., Byrd G.D., Vidoudez C., Roland A.B., Tapia E.E., Budnik B., Trauger S.A., Coloma L.A., O’Connell L.A. Molecular Physiology of Chemical Defenses in a Poison Frog. J. Exp. Biol. 2019;222:jeb.204149. doi: 10.1242 / jeb.204149.Chang L., Zhu W., Shi S., Zhang M., Jiang J., Li C., Xie F., Wang B. Plateau Grass and Greenhouse Flower? Distinct Genetic Basis of Closely Related Toad Tadpoles Respectively Adapted to High Altitude and Karst Caves. Genes. 2020; 11:123. doi: 10.3390 / genesl 1020123.Chen C.H., Bepler T., Pepper K., Fu D., Lu T.K. Synthetic Molecular Evolution of Antimicrobial Peptides. Curr. Opin. Biotechnol. 2022;75:102718. doi: 10.1016 / j. copbio.2022.102718.Chen S., Zhou Y., Chen Y., Gu J. Fastp: An Ultra-Fast All-in-One FASTQ Preprocessor. Bioinformatics. 2018;34:i884-i890. doi: 10.1093 / bioinformatics / bty560.Chen W., Hwang Y.Y., Gleaton J.W., Titus J.K., Hamlin N.J. Optimization of a Peptide Extraction and LC-MS Protocol for Quantitative Analysis of Antimicrobial Peptides. Future Sci. OA. 2019;5:FSO348. doi: 10.4155 / fsoa-2018-0073.Cho J.H., Sung B.H., Kim S.C. Buforins: Histone H2A-Derived Antimicrobial Peptides from Toad Stomach. Biochim. Biophys. Acta BBA-Biomembr. 2009;1788:1564-1569. doi: 10.1016 / j. bbamem.2008.10.025.Chowdhury T., Mandal S.M., Kumari R., Ghosh A.K. Purification and Characterization of a Novel Antimicrobial Peptide (QAK) from the Hemolymph of Antheraea Mylitta. Biochem. Biophys. Res. Commun. 2020;527:411-417. doi: 10.1016 / j.bbrc.2020.04.050.Christenson M.K., Trease A.J., Potluri L.-P., Jezewski A. J., Davis V.M., Knight L.A., Kolok A.S., Davis P.H. De Novo Assembly and Analysis of the Northern Leopard Frog Rana Pipiens Transcriptome. J. Genom. 2014;2:141-149. doi: 10.7150 / jgen.9760.Coffman K.A., Harrell T.C., Burke G.R. A Mutualistic Poxvirus Exhibits Convergent Evolution with Other Heritable Viruses in Parasitoid Wasps. J. Virol. 2020;94:e02059-19. doi: 10.1128 / JVI.02059-19.Cole, T. J. and Brewer, M.S. (2019) TOXIFY: a deep learning approach to classify animal venom proteins. PeerJ, 7, e7200.Conlon J.M. Reflections on a Systematic Nomenclature for Antimicrobial Peptides from the Skins of Frogs of the Family Ranidae. Peptides. 2008;29:1815-1819. doi: 10.1016 / j. peptides.2008.05.029.Conlon J.M., Ahmed E., Condamine E. Antimicrobial Properties of Brevinin-2-Related Peptide and Its Analogs: Efficacy Against Multidrug-Resistant AcinetobacterBaumannii. Chem. Biol. Drug Des. 2009;74:488-493. doi: 10.1111 / j.1747- 0285.2009.00882.x.Conlon J.M., Mechkarska M. Host-Defense Peptides with Therapeutic Potential from Skin Secretions of Frogs from the Family Pipidae. Pharmaceuticals. 2014;7:58-77. doi: 10.3390 / ph7010058.Cook N., Boulton R.A., Green J., Trivedi II., Tauber E., Pannebakker B.A., Ritchie M.G., Shuker D.M. Differential Gene Expression Is Not Required for Facultative Sex Allocation: A Transcriptome Analysis of Brain Tissue in the Parasitoid Wasp Nasonia vitripennis. R. Soc. Open sci. 2018;5:171718. doi: 10.1098 / rsos.171718.Cruz J., Ortiz C., Guzman F., Fernandez-Lafuente R., Torres R. Antimicrobial Peptides: Promising Compounds against Pathogenic Microorganisms. CMC. 2014;21 :2299- 2321. doi: 10.2174 / 0929867321666140217110155. da Cunha N.B., Cobacho N.B., Viana J.F.C., Lima L.A., Sampaio K.B.O., Dohms S.S.M., Ferreira A.C.R., de la Fuente-Nunez C., Costa F.F., Franco O.L., et al. The next Generation of Antimicrobial Peptides (AMPs) as Molecular Therapeutic Tools for the Treatment of Diseases with Social and Economic Impacts. Drug Discov.Today. 2017;22:234-248. doi: 10.1016 / j.drudis.2016.10.017.Das P., Wadhawan K., Chang O., Sercu T., Santos C.D., Riemer M., Chenthamarakshan V., Padhi I., Mojsilovic A. PepCVAE: Semi-Supervised Targeted Design of Antimicrobial Peptide Sequences. arXiv. 2018 doi: 10.48550 / ARXIV.1810.07743.1810.07743.Das,P., Sercu, T., Wadhawan, K., Padhi, I., Gehrmann.S., Cipcigan.F., Chenthamarakshan, V., Strobelt.H., dos Santos, C., Chen,P.-Y., et al. (2021) Accelerated antimicrobial discovery via deep generative models and molecular dynamics simulations. Nat. Bio med. Eng., 5, 613-623.Dean S.N., Alvarez J.A.E., Zabetakis D., Walper S.A., Malanoski A.P. PepVAE: Variational Autoencoder Framework for Antimicrobial Peptide Generation and Activity Prediction. Front. Microbiol. 2021; 12:725727. doi: 10.3389 / fmicb.2021.725727. de Bekker C., Ohm R.A., Loreto R.G., Sebastian A., Albert I., Merrow M., Brachmann A., Hughes D.P. Gene Expression during Zombie Ant Biting Behavior Reflects the Complexity Underlying Fungal Parasitic Behavioral Manipulation. BMC Genom. 2015;16:620. doi: 10.1186 / s12864-015-1812-x.DeGrado.W.F., Musso.G.F., Lieber, M., Kaiser, E.T. and Kezdy.F.J. (1982) Kinetics and mechanism of hemolysis induced by melittin and by a synthetic melittin analogue. Biophys. J., 37, 329-338De Gregorio E., Spellman P.T., Tzou P., Rubin G.M., Lemaitre B. The Toll and Imd Pathways Are the Major Regulators of the Immune Response in Drosophila. EMBO J. 2002;21 :2568-2579. doi: 10.1093 / emboj / 21.11.2568.De la Lastra J.M.P., Garrido-Orduha C., Borges A. A., Jimenez-Arias D., Garcia-Machado F.J., Hernandez M., Gonzalez C., Boto A. Bioinformatics discovery of vertebratecathelicidins from the mining of available genomes. In: Bobbarala V., editor. Drug Discovery — Concepts to Market. InTech; London, UK: 2018.De Lucca A. J., Walsh T.J. Antifungal Peptides: Novel Therapeutic Compounds against Emerging Pathogens. Antimicrob. Agents Chemother 1999;43:1-11. doi: 10.1128 / AAC.43.1.1.Dong R., Peng Z., Zhang Y., Yang J. MTM-Align: An Algorithm for Fast and Accurate Multiple Protein Structure Alignment. Bioinformatics. 2018;34:1719-1725. doi: 10.1093 / bioinformatics / btx828.Duckert P., Brunak S., Blom N. Prediction of Proprotein Convertase Cleavage Sites. Protein Eng. Des. Sei. 2004;17:107-112. doi: 10.1093 / protein / gzh013.Durand G.A., Raoult D., Dubourg G. Antibiotic Discovery: History, Methods and Perspectives. Int. J. Antimicrob. Agents. 2019;53:371-382. doi: 10.1016 / j.ijantimicag.2018.11.010.Eskew E.A., Shock B.C., LaDouceur E.E.B., Keel K., Miller M.R., Foley J.E., Todd B.D. Gene Expression Differs in Susceptible and Resistant Amphibians Exposed to Batrachochytrium Dendrobatidis. R. Soc. Open sci. 2018;5:170910. doi: 10.1098 / rsos.170910.Evans, E.W., Beach, G.G., Wunderlich, J. and Harmon, B.G. (1994) Isolation of antimicrobial peptides from avian heterophils. J. Leukoc. Biol., 56, 661-665.Fan W., Jiang Y., Zhang M., Yang D., Chen Z., Sun H., Lan X., Yan F., Xu J., Yuan W. Comparative Transcriptome Analyses Reveal the Genetic Basis Underlying the Immune Function of Three Amphibians’ Skin. PLoS ONE. 2017;12:e0190023. doi: 10.1371 / journal. pone.0190023.Finn R.D., Clements J., Eddy S.R. HMMER Web Server: Interactive Sequence Similarity Searching. Nucleic Acids Res. 2011;39:W29-W37. doi: 10.1093 / nar / gkr367.Finn R.D., Bateman A., Clements J., Coggill P., Eberhardt R.Y., Eddy S.R., Heger A., Hetherington K., Holm L., Mistry J., et al. Pfam: The Protein Families Database. Nucleic Acids Res. 2014;42:D222-D230. doi: 10.1093 / nar / gkt1223.Fjell.C. D. , Hiss, J. A., Hancock, R.E.W. and Schneider, G. (2012) Designing antimicrobial peptides: form follows function. Nat. Rev. Drug Discov., 11 , 37-51Fleitas.O. and Franco, O.L. (2016) Induced Bacterial Cross- Resistance toward Host Antimicrobial Peptides: A Worrying Phenomenon. Front. Microbiol., 7:381. doi: 10.3389 / fmicb.2016.00381.Fort,S., Hu,H. and Lakshminarayanan.B. (2019) Deep Ensembles: A Loss Landscape Perspective. arXiv, doi. org / 10.48550 / arXiv.1912.02757.Frishman D., Argos P. Knowledge- Based Protein Secondary Structure Assignment. Proteins. 1995;23:566-579. doi: 10.1002 / prot.340230412.Fu L., Niu B., Zhu Z., Wu S., Li W. CD-HIT: Accelerated for Clustering the next-Generation Sequencing Data. Bioinformatics. 2012;28:3150-3152. doi: 10.1093 / bioinformatics / bts565.Furman B.L.S., Evans B.J. Sequential Turnovers of Sex Chromosomes in African Clawed Frogs (Xenopus) Suggest Some Genomic Regions Are Good at Sex Determination. G3 (Bethesda) 2016;6:3625-3633. doi: 10.1534 / g3.116.033423.Gagnon, M.-C., Strandberg.E., Grau-Campistany,A., Wadhwani.P., Reichert, J., Burck.J., Rabanal.F., Auger, M., Paquin, J. -F. and Ulrich, A. S. (2017) Influence of the Length and Charge on the Activity of a-Helical Amphipathic Antimicrobial Peptides. Biochemistry, 56, 1680-1695.Gautam, A., Chaudhary.K., Singh, S., Joshi, A., Anand, P., Tuknait.A., Mathur.D., Varshney.G.C. and Raghava.G.P.S. (2014) Hemolytik: a database of experimentally determined hemolytic and non-hemolytic peptides. Nucleic Acids Res., 42, D444-D449.Gillings M.R., Paulsen I.T., Tetu S.G. Genomics and the Evolution of Antibiotic Resistance: Genomics and Antibiotic Resistance. Ann. N.Y. Acad. Sci. 2017;1388:92-107. doi: 10.1111 / nyas.13268.Goraya.J., Knoop, F.C. and Conlon, J. M. (1998) Ranatuerins: Antimicrobial Peptides Isolated from the Skin of the American Bullfrog, Rana catesbeiana. Biochem. Biophys. Res. Common., 250, 589-592.Greco I., Molchanova N., Holmedal E., Jenssen H., Hummel B.D., Watts J.L., Hakansson J., Hansen P.R., Svenson J. Correlation between Hemolytic Activity, Cytotoxicity and Systemic in Vivo Toxicity of Synthetic Antimicrobial Peptides. Sci. Rep. 2020; 10: 13206. doi: 10.1038 / S41598-020-69995-9.Grogan L.F., Mulvenna J., Gummer J.P.A., Scheele B.C., Berger L., Cashins S.D., McFadden M.S., Harlow P., Hunter D.A., Trengove R.D., et al. Survival, Gene and Metabolite Responses of Litoria Verreauxii Alpina Frogs to Fungal Disease Chytridiomycosis. Sci. Data. 2018;5: 180033. doi: 10.1038 / sdata.2018.33.Guilhelmelli F., Vilela N., Albuquerque P., Derengowski L.D.S., Silva-Pereira I., Kyaw C.M. Antibiotic Development Challenges: The Various Mechanisms of Action of Antimicrobial Peptides and of Bacterial Resistance. Front. Microbiol. 2013;4:353. doi: 10.3389 / fmicb.2013.00353.Guo R., Chen D., Diao Q., Xiong C., Zheng Y., Hou C. Transcriptomic Investigation of Immune Responses of the Apis Cerana Cerana Larval Gut Infected by Ascosphaera Apis. J. Invertebr. Pathol. 2019;166:107210. doi: 10.1016 / j.jip.2019.107210.Gupta, A. and Zou,J. (2019) Feedback GAN for DNA optimizes protein functions. Nat. Mach. Intell., 1, 105-111.Gupta, S., Kapoor, P., Chaudhary.K., Gautam, A., Kumar, R., Open Source Drug Discovery Consortium and Raghava.G.P.S. (2013) In Silico Approach for Predicting Toxicity of Peptides and Proteins. PLoS ONE, 8, e73957.Haas B.J., Papanicolaou A., Yassour M., Grabherr M., Blood P.D., Bowden J., Couger M.B., Eccles D., Li B., Lieber M., et al. De Novo Transcript Sequence Reconstruction from RNA-Seq Using the Trinity Platform for Reference Generation and Analysis. Nat. Protoc. 2013;8:1494-1512. doi: 10.1038 / nprot.2013.084.Hammond, S.A., Warren, R.L, et al., North American bullfrog draft genome provides insight into hormonal regulation of long noncoding RNA. Nature Communications, 2017(8):1433. doi: 10.1038 / s41467-017-01316-7.Hancock R.E.W., Sahl H.-G. Antimicrobial and Host-Defense Peptides as New Anti-Infective Therapeutic Strategies. Nat. Biotechnol. 2006;24:1551-1557. doi: 10.1038 / nbt1267.Hart A. J., Ginzburg S., Xu M., Fisher C.R., Rahmatpour N., Mitton J.B., Paul R., Wegrzyn J.L. EnTAP: Bringing Faster and Smarter Functional Annotation to Non-model Eukaryotic Transcriptomes. Mol. Ecol. Resour. 2020;20:591-604. doi: 10.1111 / 1755- 0998.13106.Hazam P.K., Goyal R., Ramakrishnan V. Peptide Based Antimicrobials: Design Strategies and Therapeutic Potential. Prog. Biophys. Mol. Biol. 2019;142:10-22. doi: 10.1016 / j.pbiomolbio.2018.08.006.Hede K. Antibiotic Resistance: An Infectious Arms Race. Nature. 2014;509:S2-S3. doi: 10.1038 / 509S2a.Helbing C.C., Hammond S.A., Jackman S.H., Houston S., Warren R.L., Cameron C.E., Birol I. Antimicrobial Peptides from Rana [Lithobates] Catesbeiana: Gene Structure and Bioinformatic Identification of Novel Forms from Tadpoles. Sci. Rep. 2019;9: 1529. doi: 10.1038 / S41598-018-38442-1.Hirano M., Saito C., Goto C., Yokoo H., Kawano R., Misawa T., Demizu Y. Rational Design of Helix-Stabilized Antimicrobial Peptide Foldamers Containing a,a-Disubstituted AAs or Side-Chain Stapling. ChemPlusChem. 2020;85:2731-2736. doi: 10.1002 / cplu.202000749.Hochreiter, S. and Schmidhuber.J. (1997) Long Short-Term Memory. Neural Comput., 9, 1735-1780.Hollmann.A., Martinez, M., Noguera.M.E., Augusto, M.T., Disalvo, A., Santos, N.C., Semorile.L. and Maffia.P.C. (2016) Role of amphipathicity and hydrophobicity in the balance between hemolysis and peptide-membrane interactions of three related antimicrobial peptides. Colloids Surf. B Biointerfaces, 141 , 528-536Horvati.K., Bacsa.B., Mlinko.T., Szabo, N., Hudecz.F., Zsila.F. and Bosze.S. (2017) Comparative analysis of internalisation, haemolytic, cytotoxic and antibacterial effect of membrane-active cationic peptides: aspects of experimental setup. Amino Acids, 49, 1053-1067.Huan,Y., Kong,Q., Mou,H. and Yi,H. (2020) Antimicrobial Peptides: Classification, Design, Application and Research Progress in Multiple Fields. Front. Microbiol., 11 , 978-992.llic N., Novkovic M., Guida F., Xhindoli D., Benincasa M., Tossi A., Juretic D. Selective Antimicrobial Activity and Mode of Action of Adepantins, Glycine-Rich Peptide Antibiotics Based on Anuran Antimicrobial Peptide Sequences. Biochim. Biophys. Acta (BBA)-Biomembr. 2013;1828:1004-1012. doi: 10.1016 / j.bbamem.2012.11.017.Jiang Z., Vasil A. I., Hale J.D., Hancock R.E.W., Vasil M.L., Hodges R.S. Effects of Net Charge and the Number of Positively Charged Residues on the Biological Activity of Amphipathic Alpha-Helical Cationic Antimicrobial Peptides. Biopolymers. 2008;90:369- 383. doi: 10.1002 / bip.20911.Johnson L.S., Eddy S.R., Portugaly E. Hidden Markov Model Speed Heuristic and Iterative HMM Search Procedure. BMC Bioinform. 2010;11 :431. doi: 10.1186 / 1471-2105-11- 431.Johnson, M.; Zaretskaya, I.; Raytselis, Y.; Merezhuk, Y.; McGinnis, S.; Madden, T.L. NCBI BLAST : A Better Web Interface. Nucleic Acids Research 2008, 36, W5-W9, doi:10.1093 / nar / gkn201.Jones P., Binns D., Chang H.-Y., Fraser M., Li W., McAnulla C., McWilliam H., Maslen J., Mitchell A., Nuka G., et al. InterProScan 5: Genome-Scale Protein Function Classification. Bioinformatics. 2014;30:1236-1240. doi: 10.1093 / bioinformatics / btu031.Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Zidek A., Potapenko A., et al. Highly Accurate Protein Structure Prediction with AlphaFold. Nature. 2021 ;596:583-589. doi: 10.1038 / s41586-021-03819-2.Kaur R., Stoldt M., Jongepier E., Feldmeyer B., Menzel F., Bornberg-Bauer E., Foitzik S. Ant Behaviour and Brain Gene Expression of Defending Hosts Depend on the Ecological Success of the Intruding Social Parasite. Philos. Trans. R. Soc. Lond. B Biol.Sci. 2019;374:20180192. doi: 10.1098 / rstb.2018.0192.Kazuma K., Masuko K., Konno K., Inagaki H. Combined Venom Gland Transcriptomic and Venom Peptidomic Analysis of the Predatory Ant Odontomachus Monticola. Toxins. 2017;9:323. doi: 10.3390 / toxins9100323.Khara.J.S., Obuobi.S., Wang,Y., Hamilton, M.S., Robertson, B.D., Newton, S.M., Yang.Y.Y., Langford, P.R. and Ee.P.L.R. (2017) Disruption of drug-resistant biofilms using de novo designed short a-helical antimicrobial peptides with idealized facial amphiphilicity. Acta Biomater., 57, 103-114.Khurana.D., Koli,A., Khatter.K. and Singh, S. (2023) Natural language processing: state of the art, current trends and challenges. Multimed. Tools Appl., 82, 3713-3744.Khurshid.Z., Najeeb, S., Mali,M., Moin.S.F., Raza.S.Q., Zohaib.S., Sefat.F. and Zafar, M.S. (2017) Histatin peptides: Pharmacological functions and their applications in dentistry. Saudi Pharm. J., 25, 25-31.Kintses.B., Jangir.P.K., Fekete.G., Szamel.M., Mehi,O., Spohn, R., Daruka.L., Martins, A., Hosseinnia.A., Gagarinova.A., et al. (2019) Chemical-genetic profiling reveals limitedcross-resistance between antimicrobial peptides with different modes of action. Nat. Commun., 10, 5731.Klotman M.E., Chang T.L. Defensins in Innate Antiviral Immunity. Nat. Rev.Immunol. 2006;6:447-456. doi: 10.1038 / nri1860.Koehbach J., Craik D.J. The Vast Structural Diversity of Antimicrobial Peptides. Trends Pharmacol. Sci. 2019;40:517-528. doi: 10.1016 / j.tips.2019.04.012.Koo H.B., Seo J. Antimicrobial Peptides under Clinical Investigation. Pept.Sci. 2019;111 :e24122. doi: 10.1002 / pep2.24122.Laxminarayan.R., Duse, A., Wattal.C., Zaidi.A.K.M., Wertheim, H.F.L., Sumpradit.N., Vlieghe.E., Hara.G.L., Gould, I. M., Goossens.H., et al. (2013) Antibiotic resistance — the need for global solutions. Lancet Infect. Dis., 13, 1057-1098.Leinonen R., Sugawara H., Shumway M., On behalf of the International Nucleotide Sequence Database Collaboration Sequence Read Archive. Nucleic Acids Res. 2011;39:D19-D21. doi: 10.1093 / nar / gkq1019.Li C., Sutherland D., Hammond S.A., Yang C., Taho F., Bergman L., Houston S., Warren R.L., Wong T., Hoang L.M.N., et al. AMPlify: Attentive Deep Learning Model for Discovery of Novel Antimicrobial Peptides Effective against WHO Priority Pathogens. BMC Genom. 2022;23:77. doi: 10.1186 / s12864-022-08310-4.Li,C., Warren, R.L. and Bird, I. (2023) Models and data of AMPlify: a deep learning tool for antimicrobial peptide prediction. BMC Res. Notes, 16, 11.Li,C. and Bird, I. (2023)(a) Candidate antimicrobial peptide sequences mined from the UniProtKB / Swiss-Prot database using AMPlify. Zenodo, July 10, 2023 doi.org / 10.5281 / zenodo.8133088.Li,C. and Bird, I. (2023)(b) Model files of AMPd-Up: a tool for antimicrobial peptide sequence generation. Zenodo, doi.org / 10.5281 / zenodo.7905591.Li J., Xu X., Xu C., Zhou W., Zhang K., Yu H., Zhang Y., Zheng Y., Rees H.H., Lai R., et al. Anti-Infection Peptidomics of Amphibian Skin. Mol. Cell. Proteom. 2007;6:882-894. doi: 10.1074 / mcp.M600334-MCP200.Li,Y., Huang, C., Ding,L., Li,Z., Pan,Y. and Gao,X. (2019) Deep learning in bioinformatics: Introduction, application, and perspective in the big data era. Methods, 166, 4-21.Li W.-F., Ma G.-X., Zhou X.-X. Apidaecin-Type Peptides: Biodiversity, Structure-Function Relationships and Mode of Action. Peptides. 2006;27:2350-2359. doi: 10.1016 / j. peptides.2006.03.016.Lin, D. (Published: October 31, 2022) M.Sc. Thesis High throughput in silico discovery of antimicrobial peptides in amphibian and insect transcriptomes (T). University of British Columbia, open. library. ubc.ca / collections / ubctheses / 24 / items / 1.0402476.Lin D., Sutherland D., Aninta S.I., Louie N., Nip K.M., Li C., Yanai A., Coombe L., Warren R.L., Helbing C.C., et al., Mining Amphibian and Insect Transcriptomes forAntimicrobial Peptide Sequences with rAMPage, Antibiotics (Basel), 2022 Jul 15;11(7):952. doi: 10.3390 / antibioticsl 1070952.Liscano Martinez Y., Arenas Gomez C.M., Smith J., Delgado J.P. A Tree Frog (Boana Pugnax) Dataset of Skin Transcriptome for the Identification of Biomolecules with Potential Antimicrobial Activities. Data Brief. 2020;32: 106084. doi: 10.1016 / j. dib.2020.106084.Llor C., Bjerrum L. Antimicrobial Resistance: Risk Associated with Antibiotic Overuse and Initiatives to Reduce the Problem. Ther. Adv. Drug Saf. 2014;5:229-241. doi: 10.1177 / 2042098614554919.Maher S., McClean S. Investigation of the Cytotoxicity of Eukaryotic and Prokaryotic Antimicrobial Peptides in Intestinal Epithelial Cells in Vitro. Biochem.Pharmacol. 2006;71 :1289-1298. doi: 10.1016 / j.bcp.2006.01.012.Mahlapuu M., Hakansson J., Ringstad L., Bjorn C. Antimicrobial Peptides: An Emerging Category of Therapeutic Agents. Front. Cell. Infect. Microbiol. 2016;6:194. doi: 10.3389 / fcimb.2016.00194.Mangoni M.L., Papo N., Mignogna G., Andreu D., Shai Y., Barra D., Simmaco M. Ranacyclins, a New Family of Short Cyclic Antimicrobial Peptides: Biological Function, Mode of Action, and Parameters Involved in Target Specificity. Biochemistry. 2003;42:14023-14035. doi: 10.1021 / bi0345211.Martinson E.O., Mrinalini, Kelkar Y.D., Chang C.-H., Werren J.H. The Evolution of Venom by Co-Option of Single-Copy Genes. Curr. Biol. 2017;27:2007-2013. doi: 10.1016 / j. cub.2017.05.032.Maturana P., Martinez M., Noguera M.E., Santos N.C., Disalvo E.A., Semorile L., Maffia P.C., Hollmann A. Lipid Selectivity in Novel Antimicrobial Peptides: Implication on Antimicrobial and Hemolytic Activity. Colloids Surf. B Biointerfaces. 2017;153:152-159. doi: 10.1016 / j.colsurfb.2017.02.003.McNamara-Bordewick N.K., McKinstry M., Snow J. W. Robust Transcriptional Response to Heat Shock Impacting Diverse Cellular Processes despite Lack of Heat Shock Factor in Microsporidia. mSphere. 2019;4:e00219-19. doi: 10.1128 / mSphere.00219-19.Meher P.K., Sahu T.K., Saini V., Rao A.R. Predicting Antimicrobial Peptides with Improved Accuracy by Incorporating the Compositional, Physico-Chemical and Structural Features into Chou’s General PseAAC. Sci. Rep. 2017;7:42362. doi: 10.1038 / srep42362.Meylan S., Andrews I.W., Collins J. J. Targeting Antibiotic Tolerance, Pathogen by Pathogen. Cell. 2018;172:1228-1238. doi: 10.1016 / j.cell.2018.01.037.Mhade S., Panse S., Tendulkar G., Awate R., Kadam S., Kaushik K.S. AMPing Up the Search: A Structural and Functional Repository of Antimicrobial Peptides for Biofilm Studies, and a Case Study of Its Application to Corynebacterium striatum, an EmergingPathogen. Front. Cell. Infect. Microbiol. 2021 ;11:803774. doi: 10.3389 / fcimb.2021.803774.Mikolov.T., Karafiat.M., Burget.L., Cernocky.J. and Khudanpur.S. (2010) Recurrent neural network based language model. In Proceedings of the 11th Annual Conference of the International Speech Communication Association, INTERSPEECH 2010. ISCA, pp. 1045-1048.Mirdita M., Schutze K., Moriwaki Y., Heo L., Ovchinnikov S., Steinegger M. ColabFold: Making Protein Folding Accessible to All. Nat. Methods. 2022;19:679-682. doi: 10.1038 / S41592-022-01488-1.Moravej H., Moravej Z., Yazdanparast M., Heiat M., Mirhosseini A., Moosazadeh Moghaddam M., Mirnejad R. Antimicrobial Peptides: Features, Action, and Their Resistance Mechanisms in Bacteria. Microb. Drug Resist. 2018;24:747-767. doi: 10.1089 / mdr.2017.0392.Muir P., Li S., Lou S., Wang D., Spakowicz D.J., Salichos L., Zhang J., Weinstock G.M., Isaacs F., Rozowsky J., et al. The Real Cost of Sequencing: Scaling Computation to Keep Pace with Data Generation. Genome Biol. 2016;17:53. doi: 10.1186 / s 13059-016- 0917-0.Murray C. J., Ikuta K.S., Sharara F., Swetschinski L., Robles Aguilar G., Gray A., Han C., Bisignano C., Rao P., Wool E., et al. Global Burden of Bacterial Antimicrobial Resistance in 2019: A Systematic Analysis. Lancet. 2022;399:629-655. doi: 10.1016 / 50140-6736(21)02724-0.Naamati.G., Askenazi.M. and Linial.M. (2009) ClanTox: a classifier of short animal toxins. Nucleic Acids Res., 37, W363-W368.Nagarajan.D., Nagarajan.T., Roy,N., Kulkarni.O., Ravichandran.S., Mishra, M., Chakravortty.D. and Chandra, N. (2018) Computational antimicrobial peptide design and evaluation against multidrug-resistant clinical isolates of bacteria. J. Biol. Chem., 293, 3492-3509.NCBI Resource Coordinators Database Resources of the National Center for Biotechnology Information. Nucleic Acids Res. 2016;44:D7-D19. doi: 10.1093 / nar / gkv1290.Negroni M.A., Foitzik S., Feldmeyer B. Long-Lived Temnothorax Ant Queens Switch from Investment in Immunity to Antioxidant Production with Age. Sci. Rep. 2019;9:7270. doi: 10.1038 / S41598-019-43796-1.Nguyen, L.T., Haney, E.F. and Vogel, H. J. (2011) The expanding scope of antimicrobial peptide structures and their modes of action. Trends Biotechnol., 29, 464-472.Nip K.M., Chiu R., Yang C., Chu J., Mohamadi H., Warren R.L., Bird I. RNA-Bloom Enables Reference-Free and Reference-Guided Sequence Assembly for Single-Cell Transcriptomes. Genome Res. 2020;30:1191-1200. doi: 10.1101 / gr.260174.119.Novkovic, M., Simunic.J., Bojovic.V., Tossi.A. and Juretic.D. (2012) DADP: the database of anuran defense peptides. Bioinformatics, 28, 1406-1407.O’Leary N.A., Wright M.W., Brister J. R., Ciufo S., Haddad D., McVeigh R., Rajput B., Robbertse B., Smith-White B., Ako-Adjei D., et al. Reference Sequence (RefSeq) Database at NCBI: Current Status, Taxonomic Expansion, and Functional Annotation. Nucleic Acids Res. 2016;44:D733-D745. doi: 10.1093 / nar / gkv1189.O’Neill, J. (2014) Antimicrobial Resistance: Tackling a crisis for the health and wealth of nations. The Review on Antimicrobial Resistance.Ozbek R., Wielsch N., Vogel H., Lochnit G., Foerster F., Vilcinskas A., von Reumont B.M. Proteo-Transcriptomic Characterization of the Venom from the Endoparasitoid Wasp Pimpla Turionellae with Aspects on Its Biology and Evolution. Toxins. 2019; 11:721. doi: 10.3390 / toxins11120721.Palme J., Hochreiter S., Bodenhofer U. KeBABS: An R Package for Kernel-Based Analysis of Biological Sequences. Bioinformatics. 2015;31:2574-2576. doi: 10.1093 / bioinformatics / btv176.Pan,X., Zuallaert.J., Wang,X., Shen,H.-B., Campos, E.P., Marushchak.D.O. and De Neve,W. (2021) ToxDL: deep learning using primary structure and domain embeddings for assessing protein toxicity. Bioinformatics, 36, 5159-5168.Papareddy.P., Rydengard.V., Pasupuleti.M., Walse.B., Mdrgelin.M., Chalupka.A., Malmsten.M. and Schmidtchen.A. (2010) Proteolysis of Human Thrombin Generates Novel Host Defense Peptides. PLoS Pathog., 6, e1000857.Paradis E., Schliep K. Ape 5.0: An Environment for Modern Phylogenetics and Evolutionary Analyses in R. Bioinformatics. 2019;35:526-528. doi: 10.1093 / bioinformatics / bty633.Park S., Park S.-H., Ahn H.-C., Kim S., Kim S.S., Lee B.J., Lee B.-J. Structural Study of Novel Antimicrobial Peptides, Nigrocins, Isolated from Rana Nigromaculata. FEBS Lett. 2001;507:95-100. doi: 10.1016 / 50014-5793(01)02956-8.Patro R., Duggal G., Love M.I., Irizarry R.A., Kingsford C. Salmon Provides Fast and Bias- Aware Quantification of Transcript Expression. Nat. Methods. 2017;14:417-419. doi: 10.1038 / nmeth.4197.Pei J., Feng Z., Ren T., Sun H., Han H., Jin W., Dang J., Tao Y. Purification, Characterization and Application of a Novel Antimicrobial Peptide from Andrias Davidianus Blood. Lett. Appl. Microbiol. 2018;66:38-43. doi: 10.1111 / lam.12823.Petchiappan A., Chatterji D. Antibiotic Resistance: Current Perspectives. ACS Omega. 2017;2:7400-7409. doi: 10.1021 / acsomega.7b01368.Pirtskhalava.M., Armstrong, A. A., Grigolava.M., Chubinidze.M., Alimbarashvili.E., Vishnepolsky.B., Gabrielian, A., Rosenthal, A., Hurt.D.E. and Tartakovsky.M. (2021) DBAASP v3: database of antimicrobial / cytotoxic activity and structure of peptides as a resource for development of new therapeutics. Nucleic Acids Res., 49, D288-D297.Porto W.F., Pires A.S., Franco O.L. Computational Tools for Exploring Sequence Databases as a Resource for Antimicrobial Peptides. Biotechnol. Adv. 2017;35:337-349. doi: 10.1016 / j. biotechadv.2017.02.001.Potter, S.C., Luciani.A., Eddy.S.R., Park.Y., Lopez, R. and Finn.R.D. (2018) HMMER web server: 2018 update. Nucleic Acids Res., 46, W200-W204Price S.J., Garner T.W.J., Balloux F., Ruis C., Paszkiewicz K.H., Moore K., Griffiths A.G.F. A de Novo Assembly of the Common Frog (Rana Temporaria) Transcriptome and Comparison of Transcription Following Exposure to Ranavirus and Batrachochytrium Dendrobatidis. PLoS ONE. 2015;10:e0130500. doi: 10.1371 / journal. pone.0130500.Prichula J., Primon-Barros M., Luz R.C.Z., Castro I. M.S., Paim T.G.S., Tavares M., Ligabue- Braun R., d’Azevedo P.A., Frazzon J., Frazzon A.P.G., et al. Genome Mining for Antimicrobial Compounds in Wild Marine Animals-Associated Enterococci. Mar. Drugs. 2021;19:328. doi: 10.3390 / md19060328.Qiao L., Yang W., Fu J., Song Z. Transcriptome Profile of the Green Odorous Frog (Odorrana Margaretae) PLoS ONE. 2013;8:e75211. doi: 10.1371 / journal.pone.0075211.Quevillon.E., Silventoinen.V., Pillai, S., Harte, N., Mulder, N., Apweiler.R. and Lopez, R. (2005) InterProScan: protein domains identifier. Nucleic Acids Res., 33, W116-W120.Ramazi S., Mohammadi N., Allahverdi A., Khalili E., Abdolmaleki P. A Review on Antimicrobial Peptides Databases and the Computational Tools. Database. 2022;2022:baac011. doi: 10.1093 / database / baac011.Reardon, S. (2014) Antibiotic resistance sweeping developing world. Nature, 509, 141-142.Reilly B.D., Schlipalius D.I., Cramp R.L., Ebert P.R., Franklin C.E. Frogs and Estivation: Transcriptional Insights into Metabolism and Cell Survival in a Natural Model of Extended Muscle Disuse. Physiol. Genom. 2013;45:377-388. doi: 10.1152 / physiolgenomics.00163.2012.Richter, A., Sutherland, D., Ebrahimikondori, H., Babcock, A., Louie, N., Li, C., Coombe, L., Lin, D., Warren, R. L., Yanai, A., Kotkoff, M., Helbing, C. C., Hof, F., Hoang, L. M. N., & Bird, I. (2022). Associating Biological Activity and Predicted Structure of Antimicrobial Peptides from Amphibians and Insects. Antibiotics (Basel), Nov. 27, 2022: 11(12), 1710. doi.org / 10.3390 / antibioticsl 1121710.Rifflet A., Gavalda S., Tene N., Orivel J., Leprince J., Guilhaudis L., Genin E., Vetillard A., Treilhou M. Identification and Characterization of a Novel Antimicrobial Peptide from the Venom of the Ant Tetramorium Bicarinatum. Peptides. 2012;38:363-370. doi: 10.1016 / j. peptides.2012.08.018.Rima M., Rima M., Fajloun Z., Sabatier J. -M., Bechinger B., Naas T. Antimicrobial Peptides: A Potent Alternative to Antibiotics. Antibiotics. 2021 ;10:1095. doi: 10.3390 / antibiotics10091095.Robbins, H. and Monro, S. (1951) A Stochastic Approximation Method. Ann. Math. Stat., 22, 400-407.Robinson S.D., Mueller A., Clayton D., Starobova H., Hamilton B.R., Payne R.J., Vetter I., King G.F., Undheim E.A.B. A Comprehensive Portrait of the Venom of the Giant Red Bull Ant, Myrmecia Gulosa, Reveals a Hyperdiverse Hymenopteran Toxin Gene Family. Sci. Adv. 2018;4:eaau4640. doi: 10.1126 / sciadv.aau4640.Rodriguez-Rojas A., Baeder D.Y., Johnston P., Regoes R.R., Rolff J. Bacteria Primed by Antimicrobial Peptides Develop Tolerance and Persist. PLoS Pathog. 2021;17:e1009443. doi: 10.1371 / journal.ppat.1009443.Sanchez E., Rodriguez A., Grau J.H., Letters S., Kunzel S., Saporito R.A., Ringler E., Schulz S., Wollenberg Valero K.C., Vences M. Transcriptomic Signatures of Experimental Alkaloid Consumption in a Poison Frog. Genes. 2019; 10:733. doi: 10.3390 / genesl 0100733.Santos, S.R. dos, Miranda, A. and Silva Junior, P. I. da (2022) Ovipin: a new antimicrobial peptide from chicken eggs Gallus gallus. bioRxiv, doi. org / 10.1101 / 2021.09.28.462162.Sarkar N. Polyadenylation of MRNA in Prokaryotes. Annu. Rev. Biochem. 1997;66:173-197. doi : 10.1146 / annurev. biochem.66.1.173.Scarselli.F., Gori,M., Ah Chung Tsoi, Hagenbuchner.M. and Monfardini.G. (2009) The Graph Neural Network Model. IEEE Trans. Neural Netw., 20, 61-80.Sheehan G., Farrell G., Kavanagh K. Immune Priming: The Secret Weapon of the Insect World. Virulence. 2020;11 :238-246. doi: 10.1080 / 21505594.2020.1731137.Shen W., Chen Y., Yao H., Du C., Luan N., Yan X. A Novel Defensin-like Antimicrobial Peptide from the Skin Secretions of the Tree Frog, Theloderma Kwangsiensis. Gene. 2016;576:136-140. doi: 10.1016 / j.gene.2015.09.086.Shu Y., Xia J., Yu Q., Wang G., Zhang J., He J., Wang H., Zhang L., Wu H. Integrated Analysis of MRNA and MiRNA Expression Profiles Reveals Muscle Growth Differences between Adult Female and Male Chinese Concave-Eared Frogs (Odorrana Tormota) Gene. 2018;678:241-251. doi: 10.1016 / j.gene.2018.08.007.Sievers F., Wilm A., Dineen D., Gibson T.J., Karplus K., Li W., Lopez R., McWilliam H., Remmert M., Sbding J., et al. Fast, Scalable Generation of High-Quality Protein Multiple Sequence Alignments Using Clustal Omega. Mol. Syst. Biol. 2011;7:539. doi: 10.1038 / msb.2011.75.Sim A.D., Wheeler D. The Venom Gland Transcriptome of the Parasitoid Wasp Nasonia Vitripennis Highlights the Importance of Novel Genes in Venom Function. BMC Genom. 2016;17:571. doi: 10.1186 / s12864-016-2924-7.Siu-Ting K., Torres-Sanchez M., San Mauro D., Wilcockson D., Wilkinson M., Pisani D., O’Connell M.J., Creevey C.J. Inadvertent Paralog Inclusion Drives Artifactual Topologies and Timetree Estimates in Phylogenomics. Mol. Biol. Evol. 2019;36: 1344- 1356. doi: 10.1093 / molbev / msz067.Slater G., Birney E. Automated Generation of Heuristics for Biological Sequence Comparison. BMC Bioinform. 2005;6:31. doi: 10.1186 / 1471-2105-6-31.Smith C.R., Helms Cahan S., Kemena C., Brady S.G., Yang W., Bornberg-Bauer E., Eriksson T., Gadau J., Helmkampf M., Gotzek D., et al. How Do Genomes Create Novel Phenotypes? Insights from the Loss of the Worker Caste in Ant Social Parasites. Mol. Biol. Evol. 2015;32:2919-2931. doi: 10.1093 / molbev / msv165.Song Y., Ji S., Liu W., Yu X., Meng Q., Lai R. Different Expression Profiles of Bioactive Peptides in Pelophylax Nigromaculatus from Distinct Regions. Biosci. Biotechnol. Biochem. 2013;77:1075-1079. doi: 10.1271 / bbb.130044.Strandberg E., Tiltak D., leronimo M., Kanithasen N., Wadhwani P., Ulrich A.S. Influence of C-Terminal Amidation on the Antimicrobial and Hemolytic Activities of Cationic a-Helical Peptides. Pure Appl. Chem. 2007;79:717-728. doi: 10.1351 / pac200779040717.Steinegger.M. and Sdding.J. (2017) MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol., 35, 1026-1028.Stuckert A.M.M., Chouteau M., McClure M., LaPolice T.M., Linderoth T., Nielsen R., Summers K., MacManes M.D. The Genomics of Mimicry: Gene Expression throughout Development Provides Insights into Convergent and Divergent Phenotypes in a Mullerian Mimicry System. Mol. Ecol. 2021 ;30:4039-4061. doi: 10.1111 / mec.16024.Szymczak P., Mozejko M., Grzegorzek T., Bauer M., Neubauer D., Michalski M., Sroka J., Setny P., Kamysz W., Szczurek E. HydrAMP: A Deep Generative Model for Antimicrobial Peptide Discovery. bioRxiv. 2022 doi: 10.1101 / 2022.01.27.478054.Taho, F. (2020). Antimicrobial peptide host toxicity prediction with transfer learning for proteins (Thesis). University of British Columbia. open, library. ubc.ca / collections / ubctheses / 24 / items / 1.0394496Teimouri.H., Nguyen, T.N. and Kolomeisky.A.B. (2021) Single-cell stochastic modelling of the action of antimicrobial peptides on bacteria. J. R. Soc. Interface, 18, 20210392.Tene N., Bonnafe E., Berger F., Rifflet A., Guilhaudis L., Segalas-Milazzo I., Pipy B., Coste A., Leprince J., Treilhou M. Biochemical and Biophysical Combined Study of Bicarinalin, an Ant Venom Antimicrobial Peptide. Peptides. 2016;79:103-113. doi: 10.1016 / j. peptides.2016.04.001.Terreni.M., Taccani.M. and Pregnolato.M. (2021) New Antibiotics for Multidrug-Resistant Bacterial Strains: Latest Research Developments and Future Perspectives. Molecules, 26, 2671.Tomazou M., Oulas A., Anagnostopoulos A.K., Tsangaris G.T., Spyrou G.M. In Silico Identification of Antimicrobial Peptides in the Proteomes of Goat and Sheep Milk and Feta Cheese. Proteomes. 2019;7:32. doi: 10.3390 / proteomes7040032.Tossi A., Sandri L., Giangaspero A. Amphipathic, a-Helical AntimicrobialPeptides. Biopolymers. 2000;55:4-30. doi: 10.1002 / 1097-0282(2000)55:1 <4:: Al D- BIP30>3.0.CG;2-M.Tossi.A. (2011) Design and Engineering Strategies for Synthetic Antimicrobial Peptides. In Drider.D., Rebuffat.S. (eds), Prokaryotic Antimicrobial Peptides. Springer, New York, NY, pp. 81-98.Tues, A., Tran.D.P., Yumoto.A., lto,Y. , Uzawa.T. and Tsuda.K. (2020) Generating Ampicillin- Level Antimicrobial Peptides with Activity-Aware Generative Adversarial Networks. ACS Omega, 5, 22847-22851.UniProt Consortium (The), (2019) UniProt: a worldwide hub of protein knowledge. Nucleic Acids Res., 47, D506-D515.UniProt Consortium (The). Bateman A., Martin M.-J., Orchard S., Magrane M., Agivetova R., Ahmad S., Alpi E., Bowler-Barnett E.H., Britto R., et al. UniProt: The Universal Protein Knowledgebase in 2021. Nucleic Acids Res. 2021 ;49:D480-D489. doi: 10.1093 / nar / gkaa1100. van der Does, A. M., Hiemstra.P.S. and Mookherjee.N. (2019) Antimicrobial Host Defence Peptides: Immunomodulatory Functions and Translational Prospects. In Matsuzaki.K. (ed), Antimicrobial Peptides. Advances in Experimental Medicine and Biology. Springer, Singapore, pp. 149-171 van Dijk, A., Veldhuizen.E.J.A., van Asten.A.J.A.M. and Haagsman.H.P. (2005) CMAP27, a novel chicken cathelicidin-like antimicrobial protein. Vet. Immunol. Immunopathol., 106, 321-327.Vanhoye D., Bruston F., Nicolas P., Amiche M. Antimicrobial Peptides from Hylid and Ranin Frogs Originated from a 150-Million-Year-Old Ancestral Precursor with a Conserved Signal Peptide but a Hypermutable Antimicrobial Domain. Eur. J.Biochem. 2003;270:2068-2081. doi: 10.1046 / j.1432-1033.2003.03584.x.Van Oort.C.M., Ferrell, J. B. , Remington, J. M., Wshah.S. and Li, J. (2021) AMPGAN v2: Machine Learning-Guided Design of Antimicrobial Peptides. J. Chem. Inf. Model., 61 , 2198-2207.Varadi.M., Anyango.S., Deshpande, M., Nair,S., Natassia.C., Yordanova.G., Yuan,D., Stroe.O., Wood,G., Laydon.A., et al. (2022) AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high- accuracy models. Nucleic Acids Res., 50, D439-D444.Vasilchenko.A.S., Rogozhin, E. A., Vasilchenko, A. ., Kartashova, O.L. and Sycheva, M.V. (2016) Novel haemoglobin-derived antimicrobial peptides from chicken (Gallus gallus) blood: purification, structural aspects and biological activity. J. Appl. Microbiol., 121 , 1546-1557.Vaswani.A., Shazeer.N., Parmar.N., Uszkoreit.J., Jones, L., Gomez.A.N., Kaiser, L. and Polosukhin.l. (2017) Attention Is All You Need. In Advances in Neural Information Processing Systems. pp. 6000-6010.Veltri D., Kamath U., Shehu A. Deep Learning Improves Antimicrobial PeptideRecognition. Bioinformatics. 2018;34:2740-2747. doi: 10.1093 / bioinformatics / bty179.Virtanen P., Gommers R., Oliphant T.E., Haberland M., Reddy T., Cournapeau D., Burovski E., Peterson P., Weckesser W., Bright J., et al. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python. Nat. Methods. 2020;17:261-272. doi: 10.1038 / s41592- 019-0686-2. von Wyschetzki K., Lowack H., Heinze J. Transcriptomic Response to Injury Sheds Light on the Physiological Costs of Reproduction in Ant Queens. Mol. Ecol. 2016;25:1972-1985. doi: 10.1111 / mec.13588.Wang G. Post-Translational Modifications of Natural Antimicrobial Peptides and Strategies for Peptide Engineering. CBIOT. 2012;1:72-79. doi: 10.2174 / 2211550111201010072.Wang.G., Li,X. and Wang.Z. (2016) APD3: the antimicrobial peptide database as a tool for research and education. Nucleic Acids Res., 44, D1087-D1093.Wang X., Song Y., Li J., Liu H., Xu X., Lai R., Zhang K. A New Family of Antimicrobial Peptides from Skin Secretions of Rana Pleuraden. Peptides. 2007;28:2069-2074. doi: 10.1016 / j. peptides.2007.07.020.Wang X., Ren S., Guo C., Zhang W., Zhang X., Zhang B., Li S., Ren J., Hu Y., Wang H. Identification and Functional Analyses of Novel Antioxidant Peptides and Antimicrobial Peptides from Skin Secretions of Four East Asian Frog Species. Acta Biochim. Biophys. Sin. 2017;49:550-559. doi: 10.1093 / abbs / gmx032.Wangsanuwat C., Heom K.A., Liu E., O’Malley M.A., Dey S.S. Efficient and Cost-Effective Bacterial MRNA Sequencing from Low Input Samples through Ribosomal RNA Depletion. BMC Genom. 2020;21 :717. doi: 10.1186 / s12864-020-07134-4.Wei,L., Ye,X., Sakurai.T., Mu,Z. and Wei,L. (2022) ToxlBTL: prediction of peptide toxicity based on information bottleneck and transfer learning. Bioinformatics, 38, 1514-1524.Wiegand I., Hilpert K., Hancock R.E.W. Agar and Broth Dilution Methods to Determine the Minimal Inhibitory Concentration (MIC) of Antimicrobial Substances. Nat.Protoc. 2008;3:163-175. doi: 10.1038 / nprot.2007.521.Wiradharma.N., Khoe,U., Hauser, C. A. E., Seow.S.V., Zhang, S. and Yang,Y.-Y. (2011) Synthetic cationic amphiphilic a-helical peptides as antimicrobial agents. Biomaterials, 32, 2204-2212.Wu Q., Patocka J., Kuca K. Insect Antimicrobial Peptides, a Mini Review. Toxins. 2018; 10:461. doi: 10.3390 / toxinsl 0110461.Wu,Q., Ke,H., Li,D., Wang,Q., Fang, J. and Zhou, J. (2019) Recent Progress in Machine Learning-based Prediction of Peptide Activity for Drug Discovery. Curr. Top. Med. Chem., 19, 4-16.Xia Y., Luo W., Yuan S., Zheng Y., Zeng X. Microsatellite Development from Genome Skimming and Transcriptome Sequencing: Comparison of Strategies and Lessons from Frog Species. BMC Genom. 2018;19:886. doi: 10.1186 / s12864-018-5329-y.Xiao X., Wang P., Lin W.-Z., Jia J.-H., Chou K.-C. IAMP-2L: A Two-Level Multi- Label Classifier for Identifying Antimicrobial Peptides and Their Functional Types. Anal. Biochem. 2013;436:168-177. doi: 10.1016 / j.ab.2013.01.019.Xiao.Y., Cai,Y., Bommineni.Y.R., Fernando, S.C., Prakash, O., Gilliland, S.E. and Zhang, G. (2006) Identification and Functional Characterization of Three Chicken Cathelicidins with Potent Antimicrobial Activity. J. Biol. Chem., 281 , 2858-2867.Yang L., Yang Y., Liu M.-M., Yan Z.-C., Qiu L.-M., Fang Q., Wang F., Werren J.H., Ye G.-Y. Identification and Comparative Analysis of Venom Proteins in a Pupal Ectoparasitoid, Pachycrepoideus Vindemmiae. Front. Physiol. 2020;11:9. doi: 10.3389 / fphys.2020.00009.Yang Z., Yang,D., Dyer,C., He,X., Smola.A. and Hovy,E. (2016) Hierarchical Attention Networks for Document Classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Association for Computational Linguistics, Stroudsburg, PA, USA, pp. 1480-1489.Yek S.H., Boomsma J. J., Schiott M. Differential Gene Expression in Acromyrmex Leaf- Cutting Ants after Challenges with Two Fungal Pathogens. Mol. Ecol. 2013;22:2173- 2187. doi: 10.1111 / mec.12255.Yi H.-Y., Chowdhury M., Huang Y.-D., Yu X.-Q. Insect Antimicrobial Peptides and Their Applications. Appl. Microbiol. Biotechnol. 2014;98:5807-5822. doi: 10.1007 / s00253- 014-5792-6.Yoon K.A., Kim K., Kim W.-J., Bang W.Y., Ahn N.-H., Bae C.-H., Yeo J.-H., Lee S.H. Characterization of Venom Components and Their Phylogenetic Properties in Some Aculeate Bumblebees and Wasps. Toxins. 2020;12:47. doi: 10.3390 / toxinsl 2010047.Yoshida N., Kaito C. Dataset for de Novo Transcriptome Assembly of the African Bullfrog Pyxicephalus Adspersus. Data Brief. 2020;30: 105388. doi: 10.1016 / j. dib.2020.105388.Yu G., Smith D.K., Zhu H., Guan Y., Lam T.T. ggtree: An r Package for Visualization and Annotation of Phylogenetic Trees with Their Covariates and Other Associated Data. Methods Ecol. Evol. 2017;8:28-36. doi: 10.1111 / 2041-210X.12628.Yu G., Baeder D.Y., Regoes R.R., Rolff J. Predicting Drug Resistance Evolution: Insights from Antimicrobial Peptides and Antibiotics. Proc. R. Soc. B. 2018;285:20172687. doi: 10.1098 / rspb.2017.2687.Yu G., Lam T.T.-Y., Zhu H., Guan Y. Two Methods for Mapping and Visualizing Associated Data on Phylogeny Using Ggtree. Mol. Biol. Evol. 2018;35:3041-3043. doi: 10.1093 / molbev / msy194.Yu G. Using Ggtree to Visualize Data on Tree-Like Structures. Curr. Protoc.Bioinform. 2020;69:e96. doi: 10.1002 / cpbi.96.Yu,L., Xiao,Y.-P., Li, J.-J., Ran,J.-S., Yin,L.-Q., Liu,Y.-P. and Zhang, L. (2018) Molecular characterization of a novel ovodefensin gene in chickens. Gene, 678, 233-240.Zasloff M. Antimicrobial Peptides of Multicellular Organisms. Nature. 2002;415:389-395. doi: 10.1038 / 415389a.Zelezetsky.l. and Tossi.A. (2006) Alpha-helical antimicrobial peptides — Using a sequence template to guide structure-activity relationship studies. Biochim. Biophys. Acta - Biomembr., 1758, 1436-1449.Zhang L., Gallo R.L. Antimicrobial Peptides. Curr. Biol. 2016;26:R14-R19. doi: 10.1016 / j.cub.2015.11.017.Zhang R.-W., Liu W.-T., Geng L.-L., Chen X.-H., Bi K.-S. Quantitative Analysis of a Novel Antimicrobial Peptide in Rat Plasma by Ultra Performance Liquid Chromatography- Tandem Mass Spectrometry. J. Pharm. Anal. 2011;1 :191-196. doi: 10.1016 / j.jpha.2011.04.001.Zhang, Y. and Skolnick.J. (2004) Scoring function for automated assessment of protein structure template quality. Proteins Struct. Fund. Bioinforma., 57, 702-710.Zhang, Y. and Skolnick.J. (2005) TM-align: a protein structure alignment algorithm based on the TM-score. Nucleic Acids Res., 33, 2302-2309.Zhang Y., Li Y., Qin Z., Wang H., Li J. A Screening Assay for Thyroid Hormone Signaling Disruption Based on Thyroid Hormone-Response Gene Expression Analysis in the Frog Pelophylax Nigromaculatus. J. Environ. Sci. 2015;34:143-154. doi: 10.1016 / j.jes.2015.01.028.Zhao W., Shi M., Ye X., Li F., Wang X., Chen X. Comparative Transcriptome Analysis of Venom Glands from Cotesia Vestalis and Diadromus Collaris, Two Endoparasitoids of the Host Plutella Xylostella. Sci. Rep. 2017;7:1298. doi: 10.1038 / s41598-017-01383-2.

Claims

CLAIMS1. An antimicrobial peptide comprising: an amino acid sequence according to any one of SEQ ID NO:1 to SEQ ID NO:8066, SEQ ID NO:8076 to SEQ ID NO: 8433, or a variant thereof, having at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or having 100% amino acid sequence identity thereto.

2. The antimicrobial peptide of claim 1, wherein the amino acid sequence, fragment or variant thereof, has at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or having 100%amino acid sequence identity toDeNo1001 - DeNo1058 (SEQ ID NO:1- SEQ ID NO:58),SEQ ID NO:91 , SEQ ID NO:187,HoSa4 (SEQ ID NO:285),HoSa6 (SEQ ID NO:416),DiDi3 (SEQ ID NQ:460),PIVil (SEQ ID NO:483),LySt2 (SEQ ID NO:613),GoGol (SEQ ID NO:833),SEQ ID NO:896,MuMul (SEQ ID NO:1114),HoSa2 (SEQ ID NO: 1159),SEQ ID NO: 1218,SEQ ID NO: 1238,GaGa2 (SEQ ID NO:1291),SEQ ID NO:1399,RaNo2 (SEQ ID NO:1478),SEQ ID NO:1555,SEQ ID NQ:1802,HoSa5 (SEQ ID NO:1847),DiDil (SEQ ID NO:1851),HoSal (SEQ ID NO:1978),HoSa9 (SEQ ID NO:1994),SEQ ID NO: 2046,HoSa7 (SEQ ID NQ:2048),ScPo3 (SEQ ID NQ:2074),ScPol (SEQ ID NQ:2075),SEQ ID NO: 2806,DaRel (SEQ ID NQ:2208),BoMo2 (SEQ ID NO:2254),UnBI1 (SEQ ID NO:2343),DiDi2 (SEQ ID NO:2425),Molnl (SEQ ID NO:2444),OrSa6 (SEQ ID NO:2452),HoSa8 (SEQ ID NO:2787),ArTh11 (SEQ ID NO:2797),MuMu4 (SEQ ID NO:2913),RaNol (SEQ ID NQ:2950),DIVI1 (SEQ ID NO:3162),ScPo2 (SEQ ID NO:3319),SEQ ID NO: 3363,HoSa3 (SEQ ID NO:3534),SEQ ID NO:3568,MuMu3 (SEQ ID NO:3598),DrMe6 (SEQ ID NO:3682),HoSalO (SEQ ID NQ:3700),GaGa4 (SEQ ID NO:3732),ArThlO (SEQ ID NO:3751),MuMu2 (SEQ ID NO:3881),CaGI2 (SEQ ID NO:3895),EiHel (SEQ ID NQ:3901),CaGI1 (SEQ ID NO:3941),SuSc1 (SEQ ID NQ:4032),SEQ ID NOs: 4668, 8078, 8082-8084, 8087-8089, 8091-8092, 8093-8096, 8097- 8100, 8102-8103, 8106-8113, 8115-8117, 8118-8125, 8126-8132, 8141-8143, 8148- 8150, 8153, 8155, 8157-8158, 8162-8165, 8172, 8175-8199, 8201-8206, 8208, 8209- 8211 , 8213-8220, 8222-8226, 8228, 8230-8242, 8244, 8263-8265, 8245-8254, 8256, 8258-8262, 8266, 8270-8272, 8275, 8281 or 8282-8433.

3. The antimicrobial peptide of claim 1 or 2, wherein the variant comprises a modification that is a conservative amino acid substitution.

4. The antimicrobial peptide of claim 1, wherein the peptide comprises or consists of an amino acid sequence according to SEQ ID NO:1-SEQ ID NO:8066 or according to SEQ ID NO:8076 to SEQ ID NO: 8433.

5. The antimicrobial peptide of claim 1, wherein the peptide comprises an amino acid sequence according to:Table 2 List A / Figure 6 (28):DeNo1018 (SEQ ID NO:18),DeNo1016 (SEQ ID NO:16),DeNo1017 (SEQ ID NO:17),DeNo1007 (SEQ ID NO:7),DeNo1022 (SEQ ID NO:22),DeNo1031 (SEQ ID NO:31),DeNo1021 (SEQ ID NO:21),DeNo1026 (SEQ ID NO:26),Table 2 List B / Figure 6 (1):DeNo1040 (SEQ ID NQ:40);Table 2 List C / Figure 6 (11):DeNo1049 (SEQ ID NO:49),DeNo1057 (SEQ ID NO:57),DeNo1051 (SEQ ID NO:51), andDeNo1046 (SEQ ID NO:46);Table 11:SEQ ID NQ:2208, DaRelSEQ ID NO:1978, HoSalSEQ ID NO:3162, DiVilSEQ ID NO:2425, DiDi2SEQ ID NO:1114, MuMulSEQ ID NO:285, HoSa4SEQ ID NO:483, PIVilSEQ ID NO:2797, ArTh11SEQ ID NO:2444, MolnlSEQ ID NO:833, GoGolSEQ ID NO:2343, UnBH ;Table 13; orTable 15.

6. A composition comprising the antimicrobial peptide according to any one of claims 1 to 5, and a suitable carrier or excipient.

7. The composition of claim 6, for use in treatment or prevention of an infectious disease or condition.

8. The composition of claim 7, wherein the disease or condition is attributable to Gramnegative bacteria, Gram-positive bacteria, acid fast bacteria, bacteria resistant to other drugs, a virus, a fungi, or a parasites; or wherein the disease or condition is prevention or treatment of a tumour, such as a solid tumour or a liquid tumour.

9. The composition of any one of claims 6 to 8, formulated for oral, injectable, rectal, topical, transdermal, nasal, or ocular delivery.

10. The composition of any one of claims 6 to 19, wherein the composition is lyophilized.

11. A composition for application to a surface for disinfecting or prevention of growth of microbes, said composition comprising the antimicrobial peptide according to any one of claims 1 to 5 and a suitable carrier.

12. A method of cleaning or disinfecting of a surface, a material, or an environment comprising application of the composition of claim 11 thereto.

13. Use of the antimicrobial peptide as defined in any one of claims 1 to 5 or the composition of any one of claims 6 to 10 for treatment or prevention, or for preparation of a medicament for treatment or prevention, of a disease or condition in a subject in need thereof.

14. The use of claim 15, wherein said disease or condition is infectious, and optionally is attributable to Gram-negative bacteria, attributable to Gram-positive bacteria, attributable to acid fast bacteria, attributable to bacteria resistant to other drugs, or attributable to a virus, a fungi, or a parasite.

15. The use of claim 13, wherein said disease or condition is attributable to E. coli, S. enterica, S. aureus, P. aeruginosa, S. pyogenes, M. smegmatis, MRSA, S. enterica serovar Enteritidis, S. Heidelberg, A. baumannii, K. pneumoniae, or E. faecalis bacteria.

16. The use of claim 13, wherein said disease or condition is attributable to a tumour, such as a solid tumour or a liquid tumour.

17. The use of any one of claims 13 to 16, wherein the subject is a human or an animal, such as a livestock animal, or a pet.

18. A method of treating or preventing a disease or condition comprising administering to a subject in need thereof an effective amount of the peptide according to any one of claims 1 to 5, or the composition according to any one of claims 6 to 10.

19. The method of claim 18, wherein said disease or condition is infectious, and optionally is attributable to Gram-negative bacteria, attributable to Gram-positive bacteria, attributable to acid fast bacteria, attributable to bacteria resistant to other drugs, or attributable to a virus, a fungi, or a parasite.

20. The method of claim 18, wherein said disease or condition is attributable to E. coli, S. enterica, S. aureus, P. aeruginosa, S. pyogenes, M. smegmatis, MRSA, S. enterica serovar Enteritidis, S. Heidelberg, A. baumannii, K. pneumoniae, or E. faecalis bacteria.

21. The method of claim 18, wherein said disease or condition is attributable to a tumour, such as a solid tumour or a liquid tumour.

22. The method of any one of claims 18 to 21, wherein the subject is a human or an animal, such as a livestock animal or a pet.

23. A lipid vesicle comprising the antimicrobial peptide of any one of claims 1 to 5.

24. A nucleic acid molecule encoding the antimicrobial peptide of any one of claims 1 to 5.

25. A vector comprising the nucleic acid molecule of claim 24.

26. A method of identifying a target molecule associated with an infectious agent, wherein said target molecule targets the antimicrobial peptide of any one of claims 1 to 5, said method comprising the step of screening a library of candidate target molecules associated with the infectious agent, for a molecule that targets the antimicrobial peptide.

27. The method of claim 26, wherein said infectious agent is Gram-negative bacteria, Gram-positive bacteria, acid fast bacteria, bacteria resistant to other drugs, a virus, a fungi, or a parasite.

28. The method of claim 27, wherein said infectious agent is E. coli, S. enterica, S. aureus, P. aeruginosa, S. pyogenes, M. smegmatis, MRSA, S. enterica serovar Enteritidis, S. enterica serovar Heidelberg, A. baumannii, K. pneumoniae, or E. faecalis bacteria.

29. A method of identifying a target molecule for modulating biological activity, wherein said target molecule targets the peptide of any one of claims 1 to 5, said method comprising the step of screening a library of candidate target molecules for a molecule that targets the peptide.

30. The method of claim 29, wherein modulating biological activity comprises anti-tumour action, anti-inflammatory action, or inflammatory action.

31. The method of claim 29 or 30, wherein the screening of a library of candidate target molecules comprises in silico screening.

32. A kit for identifying a target molecule associated with an infectious agent, said kit comprising the antimicrobial peptide of any one of claims 1 to 5, together with instructions for conducting the method of any one of claims 26 to 28.

33. A kit for identifying a target molecule for modulating biological activity, said kit comprising the peptide of any one of claims 1 to 5, together with instructions for conducting the method of any one of claims 29 to 31.

Citation Information

Cited By

  • Antibacterial peptide prediction method and system based on quantum peep hole LSTM neural network

    CN120526854A