Machine learning models to generate rule sets to discriminate sequence data
ML-guided AMPs selectively target keystone pathogens in peri-implantitis, addressing the limitations of current treatments by effectively inhibiting Aggregatibacter actinomycetemcomitans while preserving Streptococcus gordonii, thus mitigating peri-implantitis progression and restoring bacterial balance.
Patent Information
- Application Number
- PCT/US2025/021560
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2025-03-26
- Publication Date
- 2025-10-02
Smart Images

Figure US2025021560_02102025_PF_FP_ABST
Abstract
Description
[0001]Atty. Dkt. No.: 104434-0340 MACHINE LEARNING MODELS TO GENERATE RULE SETS TO DISCRIMINATE SEQUENCE DATA CROSS-REFERENCE TO RELATED PATENT APPLICATIONS The present application claims priority to U.S. Provisional Patent Application No. 63 / 570,116, filed March 26, 2024, the entirety of which is incorporated by reference in its entirety. STATEMENT OF GOVERNMENT SUPPORT This invention was made with government support under R01DE025476 and R56DE032903 awarded by the National Institutes of Health. The government has certain rights in the invention BACKGROUND Peri-implantitis is a complex infectious disease that manifests as progressive loss of alveolar bone around the dental implants and hyper-inflammation associated with microbial dysbiosis. Despite high success rates for dental implants, their bacterial plaque- associated inflammatory lesions, known as peri-implant diseases, still occur. These lesions continue to degrade the stability of peri-implant soft and hard tissues, which can result in loss of the implant. While peri-implant mucositis is a reversible inflammatory condition, peri- implantitis is an irreversible pathological condition leading to loss of supporting alveolar bone. The reported prevalence of peri-implant mucositis and peri-implantitis shows a substantial increase over time following implant placement. Meta-analysis for patient-based peri-implant mucositis and peri-implantitis was reported as 46.83% and 19.83%. In a separate study, meta-analyses estimated the peri-implant mucositis and peri-implantitis as 43% and 22%, respectively. Peri-implantitis is also reported to be in the range of 11-47% among dental implants 10 years after their placement. These numbers further increase in periodontally compromised patients. Current treatments for peri-implantitis and periodontitis include mechanical debridement, disinfection of exposed implant surfaces, and antibiotic or antiseptic -1- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 prescription to suppress the associated bacteria. The use of adjunctive antibiotics for treating peri-implantitis or periodontitis is debated mainly due to the concern of microbial antibiotic resistance, the non-selective suppression of both pathogenic and commensal species, and the adverse systemic reactions. Notably, these conventional treatment modalities may not prevent relapses, as the pathogens may either remain unaffected or quickly re-emerge after treatment. The poor efficacy of antibiotic treatment in peri-implantitis may be explained by their non-specific suppression of dysbiotic biofilms. The adaptability and resiliency of pathogenic bacteria in biofilms is well documented. The unique structure and inter-species relationships within a biofilm enhance the individual strengths of the bacteria present, creating an unbalanced community organized to promote communal success at the expense of the host.(9) Notably, keystone pathogens play an outsized role in shaping the community structure. Therefore, targeting keystone pathogens may be the most effective approach to reverse microbial dysbiosis and return to health-compatible eubiosis. In peri-implantitis, Porphyromonas gingivalis (P. gingivalis) is widely acknowledged as a keystone pathogen. P. gingivalis is associated with increased levels of inflammation and subsequent alveolar bone loss. Once P. gingivalis has initiated biofilm growth, other pathogens are free to flourish and further contribute to the dysbiotic community. The interdependent-relationships among pathogens in a biofilm are a defining factor in their treatment difficulty. In peri-implantitis, this is evident by the coexistence of Aggregatibacter actinomycetemcomitans, another keystone pathogen associated with aggressive periodontitis, and Streptococcus gordonii, a commensal and accessory pathogen. Microbial communities exhibiting both P. gingivalis and S. gordonii are linked to more severe cases of peri-implantitis, resulting in increased infection and bone loss as compared to others. Undoubtedly, in peri-implantitic biofilms, pathogens grow synergistically to promote each other’s survival. Keystone pathogens, such as P. gingivalis and A. actinomycetemcomitans, play a pivotal role in shifting the oral microbiome to induce the host into a disease-oriented state. The presence of these pathogens is widely associated with intensified inflammation levels, prolonged infection, and enhanced alveolar bone loss in patients. -2- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 SUMMARY Successful mitigation of disease progression in peri-implantitis requires a specific mode of treatment capable of targeting keystone pathogens and restoring bacterial community balance toward commensal species. Broad spectrum approaches have difficulty in preventing oral dysbiosis. This addresses this need by offering ML enabled design features that can target keystone pathogens without destroying the commensal species and restore bacterial community balance. Presented herein are systems and methods to provide machine learning (ML) enabled peptide design features to selectively target keystone pathogens that plays a critical role in peri-implantitis disease progression and restore bacterial community balance. Current treatment options may not prevent relapses, as the pathogens either remain unaffected or quickly re-emerge after treatment. Successful mitigation of disease progression in peri- implantitis requires a specific mode of treatment capable of targeting keystone pathogens and restoring bacterial community balance toward commensal species. The interdependent relationships among the pathogens relate to their treatment difficulty aligned with their role as the disease progress. A confirmed peptide can selectively target the keystone pathogen Aggregatibacter actinomycetemcomitans without compromising the accessory commensal species, Streptococcus gordonii. A transparent ML model was developed, combined with a genetic algorithm, that empowers AMP design targeted to a keystone pathogen. Predicted activity of the generated AMPs were classified using rough sets and the rules are improved using empirical growth inhibition data for specific pathogens. The validation tests were run with the peptides having the highest inhibition predictability score against the keystone pathogen, A. actinomycetemcomitans. A novel peptide VL-13 was confirmed to be active against the selected keystone pathogen without compromising the accessory-commensal species, S. gordonii. A path for developing an engineering approach to iteratively discover targeted antimicrobials may be provided as a robust method for potential therapeutic treatment for peri-implant biofilm infections to reduce the resulting host response of hyper-inflammation during disease progression. -3- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Aspects of the present disclosure are directed to systems and methods for using machine learning (ML) models to generate rule sets to discriminate sequence data. One or more processors coupled with memory may retrieve a training dataset comprising a plurality of AMP sequences and a plurality of non-AMP sequences. The plurality of AMP sequences may target the select microbial populations. The one or more processors may generate, for each of the plurality of AMP sequences and of the plurality of non-AMP sequences of the training dataset, a respective plurality of properties comprising at least one of (i) a hydropathy or (ii) a microbial inhibitory activity for at least one microbial species of the select microbial populations. The one or more processors may provide as input to a ML model, the plurality of AMP sequences, the plurality of non-AMP sequences, and the respective plurality of properties for each of the plurality of AMP sequences and the plurality of non-AMP sequences. The one or more processors may determine, based on providing the input to the ML model, a set of rules defining values for the respective plurality of properties to discriminate between the plurality of AMP sequences and the plurality of non-AMP sequences. The one or more processors may apply the set of rules to a plurality of candidate AMP sequences to identify a subset of AMP sequences that satisfy the values defined by the set of rules for the plurality of AMP sequences. The one or more processors may store, using one or more data structures, the subset of candidate AMPs. In some embodiments, an AMP corresponding to a candidate AMP sequence of the subset of AMP sequences may be synthesized using a peptide synthesizer. The AMP may be tested in a microbial model including the select microbial population to determine a score indicating degree of efficacy of the AMP against the select microbial populations. The one or more processors may identify the AMP as effective or ineffective in targeting the select microbial populations based on a comparison of the score with a threshold. In some embodiments, the microbial model may be one of a single-species microbial model or a poly- species microbial model. A concentration of the AMP may be at one of a minimum inhibitory concentration (MIC) or minimum bactericidal concentration (MBC). In some embodiments, the one or more processors may add the candidate AMP sequence to the plurality of AMP sequences to generate a second plurality of AMP sequences for the training dataset, responsive to identifying the candidate AMP as effective. -4- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 The one or more processors may generate, for the candidate AMP sequence of the training dataset, a second plurality of properties. The one or more processors may provide, as input to the ML model: (i) the second plurality of AMP sequences including the candidate AMP sequence, (ii) the plurality of non-AMP sequences, (iii) the respective plurality of properties for each of the plurality of AMP sequences and the plurality of non-AMP sequences, and (iv) the second plurality of properties. The one or more processors may determine, based on providing the input to the ML model, a second set of rules defining values for the respective plurality of properties to discriminate between the second plurality of AMP sequences and the plurality of non-AMP sequences. In some embodiments, the one or more processors may receive, for a subject at risk of or diagnosed with peri-implantitis or periodontitis associated with dental tissue about a dental implant, an indication of a presence of the select microbial populations in the tissue of the subject. The one or more processors may provide an output identifying the AMP sequence to target the select microbial populations, responsive to identifying the AMP as effective. The dental tissue about the dental implant of the subject may be administered with a therapy including the AMP to address the peri-implantitis or periodontitis. In some embodiments, the one or more processors may refrain from adding the candidate AMP sequence to the plurality of AMP sequences of the training dataset, responsive to identifying the AMP as ineffective. In some embodiments, the one or more processors may generate the plurality of candidate AMP sequences using a second plurality of AMP sequences identified as targeting the select microbial populations. In some embodiments, the one or more processors may train, using the set of rules, a second ML model to generate the second plurality of AMP sequences that satisfy the values defined by the set of rules. In some embodiments, the one or more processors may determine, for each of the second plurality of AMP sequences, a respective metric indicating a degree of similarity of a respective AMP sequence with at least one of a third plurality of AMP sequences, wherein the third plurality of AMP sequences targets the select microbial populations. The one or more processors may select, from the second plurality of AMP sequences, the plurality of candidate AMP sequences based on the respective metric. -5- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 In some embodiments, the one or more processors may update the ML model based on checking the set of rules on a second plurality of AMP sequences corresponding to a plurality of AMPs present in a commensal sample. In some embodiments, the one or more processors may determine, from at least one layer of the ML model, a plurality of embeddings defining the set of rules to discriminate between the plurality of AMP sequences and the plurality of non-AMP sequences. In some embodiments, the one or more processors may reduce, using an approximator, a number of boundary conditions corresponding to the values of the set of rules to discriminate between the plurality of AMP sequences and the plurality of non-AMP sequences. In some embodiments, the one or more processors may generate the respective plurality of properties further comprises generating, for each of the plurality of AMP sequences and of the plurality of non-AMP sequences of the training dataset, the respective plurality of properties to include at least one of: (i) a matrix defining a plurality of pair distances for a respective peptide, each of the plurality of pair distances between a respective pair of residues in the respective peptide; (ii) a Fourier score determined based on the plurality of pair distances of the matrix; or (iii) a secondary structure identifying a spatial arrangement of a polypeptide backbone of the respective peptide. In some embodiments, the select microbial populations may be associated with peri-implantitis or periodontitis. The select microbial populations may include at least one of Porphyromonas gingivalis, Aggregatibacter actinomycetemcomitans, or Streptococcus gordonii. At least one aspect of the present disclosure is directed to an antimicrobial peptide. The antimicrobial peptide may be any peptide disclosed herein, or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. In some embodiments, the antimicrobial peptide may be of an amino acid sequence of any one of SEQ ID NOs: 10-11, 13-15, 17-46, 51-168, and 173-747, or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. In some embodiments, the antimicrobial peptide may be of an amino acid sequence of any one of SEQ ID NOs: 17-46, 51-168, and 173-747, or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. -6- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 In some embodiments, the antimicrobial peptide may be of an amino acid sequence that is: KWKLFKTTAKFLHLAK (SEQ ID NO: 14) (“KK-15”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof, FLHWVPLRRVV (SEQ ID NO: 15) (“FV-11”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof, VDWKKVFGKLLKL (SEQ ID NO: 16) (“VL-13”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof, or LGKLLKKIPKFLHLVNK (SEQ ID NO: 387) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. In some embodiments, the antimicrobial peptide may be of an amino acid sequence that is VDWKKVFGKLLKL (SEQ ID NO: 16) (“VL-13”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. At least one aspect of the present disclosure is directed to a composition. The composition may include an antimicrobial peptide of any one of those disclosed herein and a pharmaceutically acceptable carrier. At least one aspect of the present disclosure is directed to a method of treating peri-implant disease in a subject in need thereof. The method may include administering an effective amount of an antimicrobial peptide of any of those disclosed herein to a dental implant in the subject. In some embodiments, the peri-implant disease may be peri-implantitis. At least one aspect of the present disclosure is directed to a method of controlling bacterial colonization on a dental implant in a subject. The method may include administering to the dental implant an effective amount of an antimicrobial peptide of any of those disclosed herein. At least one aspect of the present disclosure is directed to a method to control biofilm formation on a dental implant in a subject. The method may include administering to the dental implant an effective amount of the peptide of an antimicrobial peptide of any of those disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS The foregoing and other objects, aspects, features, and advantages of the disclosure will become more apparent and better understood by referring to the following description taken in conjunction with the accompanying drawings, in which: -7- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 FIG. 1: Peptide targeting design scheme for antimicrobial peptides. The training is based on peptide sequences with identified growth inhibition. The expanded sequences are selected for consistency with identified training relationships. FIG. 2: Peptide signals for iAMP-2l antibacterial sequences and training sequences (2,347 peptides). The signals are stacked columns of four different signals: log P, peptide charge, peptide length and net inhibition rules. The height of the stacked signals is limited to 0.5 using the inverse logit of the signal value, normalized as percentage of the overall observed range for the value. The database is divided into quartiles (Q1–Q4) of decreasing fitness. This view indicates immediately how many peptides have positive log P (hydrophobic), and which have negative log P (hydrophilic), in the order of the fitness criteria. The fitness criteria prioritize sequences having positive net inhibition rules for targeting. Priority is also given to shorter, more hydrophilic peptides for easier synthesis. Signal strength λ is calculated to maximize the sensitivity at the lower percentile ranges of property values. See Supplementary Materials 1.1. FIG. 3: Peptide fitness signal quartiles (Q1–Q4) for 10th and 25th generation sequences. The initial generation seen in Figure 2, has only 20 sequences in the first quartile which meet the targeted inhibition rules. From this small set, the genetic algorithm generated peptides which have >200 sequences meeting the inhibition rules by Generation 10. However, the variation of the net inhibition rules for Generation 10 is high among these peptides. By Generation 25, maturation is achieved with >300 sequences with low variation of net inhibition rule conditions. Signal strength λ is calculated to maximize the sensitivity at the lower percentile ranges of property values. See Supplementary Materials 1.1. FIG. 4: Peptide property values for Top-25 scoring sequences in the 25th generation of 2,261 sequences. The candidates selected for further experimental validation are highlighted in purple: KK-15, FV-11, and VL-13. FIG. 5: The in vitro analyses of ML generated top scoring peptides: VL-13, KK-15, FV-11 against keystone pathogen A. actinomycetemcomitans. Peptides without 6 h columns had no countable CFU / ml. The other included peptides are antimicrobial peptides which have known activity outside of the oral environment. Keystone targeting score is the -8- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 difference in change in CFU / ml of the keystone pathogen and the minimum inhibition for either commensal strain. The keystone targeting scores are in Table 3. FIG. 6: The in vitro analyses of ML generated top scoring peptides: VL-13, KK-15, FV-11 against accessory pathogen S. gordonii. Peptides without 6 h columns had no countable CFU / ml. The other peptides are antimicrobial peptides which have known activity outside of the oral environment. Keystone targeting score is the difference in change in CFU / ml of the keystone pathogen and the minimum inhibition for either commensal strain. The keystone targeting scores are in Table 3. FIG. 7: The in vitro analyses of ML generated top scoring peptides: VL-13, KK-15, FV-11 against commensal strain, S. sanguinis. Peptides without 6 h columns had no countable CFU / ml. The other peptides are antimicrobial peptides which have known activity outside of the oral environment. Keystone targeting score is the difference in change in CFU / ml of the keystone pathogen and the minimum inhibition for either commensal strain. The keystone targeting scores are in Table 3. FIG. 8: Ensemble structure description of VL-13 (VDWKKVFGKLLKL (SEQ ID NO: 16)). The hydrophobic feature of VFG (see Figure 10) has very high solvent accessibility. This indicates high accessibility for potential binding partners. The locations of the side chains are stable through the ensemble. Only the Phenylalanine 7 seems to have multiple distinct clusters of orientations, which may indicate a local rotation between two smaller vibrations of the aromatic ring. The conserved backbone structure simplifies structure-function hypotheses for peptides. This lower folding entropy simplifies structure- function studies compared to using AMPa or AMP1. Superposition of 192 PEP-FOLD3 (66) structures aligned with MatchAlign (64) in UCSF Chimera (65). FIG. 9: JACR890101 hydrophobicity property trend as both a single amino acid property (blue) and as the sum of three consecutive amino acids (orange). In (A), VL-13 meets Rule 1 Condition 2 in Table 2 at residue F7 where the 3-aa window falls within the rule bounds (dashed lines). In (B), KF-18 does not have a hydropathy feature of the 3-aa sum that meets this condition. The dashed lines indicate the Rule 1 Condition 2 upper and lower -9- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 boundaries in Table 2. Figure 16 compares Rule 2 and Rule 3 conditions in Table 2 with Rule 1. FIG. 10: High folding entropy of AMPa (A) and AMP1 (B) compared to VL- 13. The orange ribbon is the first half of each sequence and green is the second half of the corresponding sequence. The divergent ribbons indicate a folding ensemble with many different orientations of backbone structure having similar energy values. In Figure 9, the conserved backbone shape among superposition structures indicates a relatively large folding energy barrier to other backbone shapes. While the high structural entropy complicates the structure-function analysis of these peptides, surface interactions are still facilitated through electrostatic means. The electrostatic surface indicates highly charged surface which is mostly electropositive (blue) for the first half of AMPa with most electronegative (red) surface for the second half of AMPa. The distribution of electronegativity for AMP1 is more heterogeneous compared to AMPa. Superposition of 100 PyRosetta structures (63) aligned with MatchAlign (64) in UCSF Chimera (65). FIG. 11: Antimicrobial activity against keystone periodontal pathogen A. actinomycetemcomitans for selected antimicrobial peptides. This data was used to supplement the training data of the CLN-MLEM2 model. AMPa: KWKLWKKIEKWGQGIGAVLKWLTTW-NH2 (SEQ ID NO: 1), AMP1: LKLLKKLLKLLKKL (SEQ ID NO: 2), AMP2-NH2: KWKRWWWWR-NH2(SEQ ID NO: 3), AMP7: ESYKKML (SEQ ID NO: 4), AMP10: GILGKLWEGVKSTF (SEQ ID NO: 5). FIG. 12: Antimicrobial activity against keystone periodontal pathogen A. actinomycetemcomitans for selected antimicrobial peptides. This data was used to supplement the training data of the CLN-MLEM2 model. TIBP-S5-AMPA: RPRENRGRERGLGSGGGKWKLWKKIEKWGQGIGAVLKWLTTW-NH2 (SEQ ID NO: 6), TIBP-AH-AMP1: RPRENRGRERGLKGSVLSADLKLLKKLLKLLKKL (SEQ ID NO: 7), TiBP-AH-GL13K: RPRENRGRERGLKGSVLSADGKIIKLKASLKLL (SEQ ID NO: 8). -10- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 FIG. 13: Antimicrobial activity against periodontal accessory pathogen S. gordonii for selected antimicrobial peptides. This data was used to supplement the training data of the CLN-MLEM2 model. AMPa: KWKLWKKIEKWGQGIGAVLKWLTTW-NH2 (SEQ ID NO: 1), AMP1: LKLLKKLLKLLKKL (SEQ ID NO: 2), AMP2-NH2: KWKRWWWWR-NH2 (SEQ ID NO: 3), AMP7: ESYKKML (SEQ ID NO: 4), AMP10: GILGKLWEGVKSTF (SEQ ID NO: 5). FIG. 14: Antimicrobial activity against periodontal accessory pathogen S. gordonii for selected antimicrobial peptides. This data was used to supplement the training data of the CLN-MLEM2 model. TIBP-S5-AMPA: RPRENRGRERGLGSGGGKWKLWKKIEKWGQGIGAVLKWLTTW-NH2(SEQ ID NO: 6), TIBP-AH-AMP1: RPRENRGRERGLKGSVLSADLKLLKKLLKLLKKL (SEQ ID NO: 7), TiBP-AH-GL13K: RPRENRGRERGLKGSVLSADGKIIKLKASLKLL (SEQ ID NO: 8). FIG. 15: Antimicrobial activity against periodontal accessory pathogen P. gingivalis for selected antimicrobial peptides. This data was used to supplement the training data of the CLN-MLEM2 model. TIBP-AH-AMP1: RPRENRGRERGLKGSVLSADLKLLKKLLKLLKKL (SEQ ID NO: 7), TiBP-AH-GL13K: RPRENRGRERGLKGSVLSADGKIIKLKASLKLL (SEQ ID NO: 8), AMP10: GILGKLWEGVKSTF (SEQ ID NO: 5). FIG. 16: Detailed Rough Set Theory Rule Evaluation for JACR890101 in the AAindex1 of VL-13. In Figure _, the peak hydrophobicity features are compared for VL-13 and KF-18. While VL-13 met the first rule in Table 2, KF-18 did not. This plot shows the ranges of Rules 1 and 3 from Table 2, which correspond to A. actinomycetemcomitans (Aa) growth inhibition and no growth inhibition respectively. Rule 1 has two conditions for JACR890101: Condition 2 for peak hydrophobicity, indicating by a three amino acid (3-aa) hydrophobic feature, and Condition 4 for the mean hydropathy of the overall sequence to be hydrophilic. This description is a part of a quantitative boundary for amphiphilic sequence structure which is discriminating between inhibition and non-inhibition. Rule 3 Condition 4 also contains a hydrophobic 3-aa window. Rules are the conjunction of conditions, usually -11- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 across multiple key physicochemical properties. Rule 1 (+ Aa inhibition) contains 5 conditions, Rule 2 (+ Aa inhibition) contains 4 conditions and Rule 3 (- Aa inhibition) contains 4 conditions. FIG. 17: Detailed Rough Set Theory Rule Evaluation for ZIMJ680103 in the AAindex1 of VL-13. VL-13 meets a total sequence characteristic of Rule 1 Condition 1 for this property. Rule 1 relates the overall mean of ZIMJ680103 (mean value: 13.6) to the probability of the amino acid components occurring in a sequence which is, on average, slightly hydrophilic (ranging from 17.1 to 23.0). Higher scores for this property are more hydrophilic. Rules are the conjunction of conditions, usually across multiple key physicochemical properties. Rule 1 (+ Aa inhibition) contains 5 conditions, Rule 2 (+ Aa inhibition) contains 4 conditions and Rule 3 (- Aa inhibition) contains 4 conditions. FIG. 18: Detailed Rough Set Theory Rule Evaluation for COWR900101 in the AAindex1 of VL-13. This property is contained in both Rule 1, in Condition 3 and in Condition 5, and Rule 3 Condition 2, partially discriminating between inhibiting the growth of Aa and its non-inhibition. This property relates to the changing polarity in acidic conditions. At pH 3, aspartic and glutamic acid side chain residues are uncharged and are counted as polar residues instead of charged residues. When sequences contain these two residues, this estimate will be less hydrophilic than for other hydropathy scales. For Rule 1 the sequence must average slightly hydrophilic average in acidic conditions as well have a sum of hydrophilic index between -27.82 and 0.27, indirectly constraining the length of peptides and hydropathy variation which meet Rule 1. Rule 3 requires hydrophobic peak under acidic conditions. Rules are the conjunction of conditions, usually across multiple key physicochemical properties. Rule 1 (+ Aa inhibition) contains 5 conditions, Rule 2 (+ Aa inhibition) contains 4 conditions and Rule 3 (- Aa inhibition) contains 4 conditions. FIG. 19: Detailed Rough Set Theory Rule Evaluation for WARP780101 in the AAindex1 of VL-13. Peptides meeting Rule 2 Condition 1 include residues which average between 6.4 – 7.1 interactions per side chain, where the mean of the property is 7.4. Higher values tend to be less polar amino acids. Rule 2 contains a condition of being slightly non-polar for the total sequence. Therefore, VL-13 is slightly less polar by property -12- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 WARP780101 and slightly polar by JACR890101 property because it also meets Rule 1 conditions. Being less polar and slightly polar is not a contradiction when different properties are considered. Rule 3 Condition 3 indicates that peptides in its set must as a highly non-polar 3-amino acid peptide window. Rules are the conjunction of conditions, usually across multiple key physicochemical properties. Rule 1 (+ Aa inhibition) contains 5 conditions, Rule 2 (+ Aa inhibition) contains 4 conditions and Rule 3 (- Aa inhibition) contains 4 conditions. FIG. 20: Detailed Rough Set Theory Rule Evaluation for MEEJ810102 in the AAindex1 of VL-13. The key property is retention in NaH2PO4, with a pKa of 7.2. The mean of all of the amino acids is 2.6. The condition of Rule 2 Condition 2 for this property is an overall mean between 3.2 and 7.3, which is relatively hydrophobic for this hydropathy property. VL-13 is slightly overall non-polar by the Rule 2 description, but slightly overall polar by the Rule 1 description. Different polarity properties are used for the two rules. Rules are the conjunction of conditions, usually across multiple key physicochemical properties. Rule 1 (+ Aa inhibition) contains 5 conditions, Rule 2 (+ Aa inhibition) contains 4 conditions and Rule 3 (- Aa inhibition) contains 4 conditions. FIG. 21: Detailed Rough Set Theory Rule Evaluation for FAUJ880110 in the AAindex1 of VL-13. The number of full non-bonding orbitals is high for polar amino acids and tyrosine. Generally, hydrophobic amino acids have none. Having highly polar 3-aa window is Rule 3 Condition 1 that corresponds to non-growth inhibition. VL-13 does not meet this condition, but Rule 2 Condition 3 has a moderate polarity condition of a 3-aa feature for this property that VL-13 meets via residues 2-4 (DWK). A second condition for Rule 2 (Condition 4) is a moderately non-polar average between 0.4 and 0.69. The average for the property is 1.25 and lower values are less polar. Rules are the conjunction of conditions, usually across multiple key physicochemical properties. Rule 1 (+ Aa inhibition) contains 5 conditions, Rule 2 (+ Aa inhibition) contains 4 conditions and Rule 3 (- Aa inhibition) contains 4 conditions. FIG. 22: Flow diagram for ML design approach which iteratively discovers new antimicrobial peptides. -13- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 FIG. 23: Increased aggregation potential range of generated candidates using the codon representation (-1 to 2) compared to without the representation (-0.8 to 1.3). FIGs. 24A–D: Codon representation results in candidates with similar convergence in this study. FIG. 25: The genetic algorithm produced candidates with a high probability near the global optimum of a score of zero with the codon representation. FIG. 26: Zone of inhibition for antibiotic control (Ampicillin), two natural AMP controls (GIHD…) and (DYHH…), and five designed antibacterial candidates designed from generic antibacterial rules. FIG. 27: Flow Chart 1: Targeted AMP Generation. An autoencoder (AE) is provided with sequences encoded by AAindex1 hydropathy properties, inhibition activity, transformed distances and secondary structures. The rules may be outputted by the AE. Scan commensal genomes for AMPs with selected rules and retrain both AE and a multi-class CNN The multi-class CNN model can generate sequences from desired inhibition patterns. The newly generated AMP sequences can be tested in a single-species model to perform targeted AMP generation. The selected AMPs may be tested in polymicrobial dysbiosis model to evaluate for dysbiosis and keystone pathogen collapse. Based on the results, targeted AMP generation can be iteratively performed. FIG. 28: Flow Chart 2: ML guided SB-AMP design for modifying microbial dysbiosis. Steps: (a) Pre-training a model (e.g., ESM-2) to calculate transformed distances and secondary structures of various peptide structures. (b) Preprocess training set into pair distances from PLM. Calculate Fourier distances and secondary structure features. (c) Train AMP autoencoder (AE) with sequences encoded by AAindex1 hydropathy properties, inhibition activity, transformed distances and secondary structures. (d) Once trained, the AMP AE can be used to generate AMP sequences to target certain microbial populations. CLN-MLEM2 selects inhibition rules from AE-embedded coordinate intervals. Scan commensal genomes for AMPs with selected rules and retrain both AE and CLN- MLEM2. Train semi-supervised multi-class CNN model and generate sequences from -14- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 desired inhibition patterns. (e) Select AMP sequences based on target inhibition criterion and sequence similarity. (f) Test single-species MIC and MBC of 1st Gen AMPs in both the planktonic state and in biofilms. Retrain ML model to generate SB-AMPs. FIG. 29: Flow Chart 3: ML guided SB-AMP design for modifying microbial dysbiosis. Steps: (a) Preprocess training set into pair distances from PLM. Initial set: 4,134 AMPs and 4,134 non-AMPs18. Calculate Fourier distances and secondary structure features (b) Train AMP autoencoder (AE) with sequences encoded by AAindex1 hydropathy properties, inhibition activity, transformed distances and secondary structures. CLN-MLEM2 selects inhibition rules from AE-embedded coordinate intervals. Scan commensal genomes for AMPs with selected rules and retrain both AE and CLN-MLEM2. Train semi-supervised multi-class CNN model and generate sequences from desired inhibition patterns. (c) Filter these sequences by the retrained rules and enrich with genetic algorithm by sequence similarity from prior patterns yielding 1st Gen AMP set. (d) Test single-species MIC and MBC of 1st Gen AMPs in both the planktonic state and in biofilms. Retrain ML model to generate SB-AMPs (a-c). Test SB-AMPs in polymicrobial model and iterative ML retraining, validate SB-AMPs for induced changes. FIG. 30: Codon-Based Genetic Algorithm Flow Diagram: The flow can include: select predicted active sequences by CLN-MLEM 2 rules; select top-scoring amino acid sequences; encode sequence to nucleic acid using codon table; and find new peptide sequences using single-point mutations and crossover. FIG. 31: Annotated flow diagram: The flow can further include: select predicted active sequences by CLN-MLEM 2 rules; select top-scoring amino acid sequences; encode sequence to nucleic acid using codon table; find new peptide sequences using generation operators (e.g., single-point mutations and crossover); determine unique codon representation, escaping local minima; rank the codon representations; and apply specific rules, and repeat flow. FIGs. 32A–C depict graphs showing percentages of colony-forming units (%CFU) of various microbial populations with new candidate AMPs, over time. -15- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 FIGs. 33A and 33B depict graphs showing percentages of colony-forming units (%CFU) of various microbial populations with new candidate AMPs, over time FIG. 34 depicts a graph of an inhibition zone of various candidate novel antimicrobial peptides in targeting S. epidermidis FIGs. 35A and 35B. Inhibition of planktonic bacteria (upper) and biofilms (lower) by 1st Gen ML-AMPs designed with relevance rules and genetic algorithm (VL-13, KK-15 and FV-11) compared to peptide (AMP1). A. actinomycetemcomitans D7S-1, S. gordonii ATCC 35105 and S. sanguinis ATCC 10556 were grown for 6 hrs (upper panel) or 4 hours (lower panel) in Shi with peptides (100 µM for planktonic cells, and 200 µM for biofilms) at 37°C in an atmosphere supplemented with 5% CO2. The CFUs were enumerated at T=0 (blue) or after incubation (orange). FIG. 36 depicts a block diagram of a system for using machine learning (ML) models to identify species-biased antimicrobial peptides (SB-AMPs) that target select microbial populations, in accordance with an illustrative embodiment. FIG. 37 depicts a block diagram of a process to determine properties of antimicrobial peptide (AMP) and non-AMP sequences in the system for using ML models to identify species-biased antimicrobial peptides (SB-AMPs), in accordance with an illustrative embodiment. FIG. 38 depicts a block diagram of a process to generate rule sets to discriminate antimicrobial peptide (AMP) and non-AMP sequences and new candidate AMP sequences in the system for using ML models to identify species-biased antimicrobial peptides (SB-AMPs), in accordance with an illustrative embodiment. FIG. 39 depicts a block diagram of a process to validate inhibition levels of candidate antimicrobial peptides (AMPs) in the system for using ML models to identify species-biased antimicrobial peptides (SB-AMPs), in accordance with an illustrative embodiment. -16- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 FIG. 40 depicts a block diagram of a process to provide therapy to dental tissues in the system for using ML models to identify species-biased antimicrobial peptides (SB-AMPs), in accordance with an illustrative embodiment. FIG. 41 depicts a flow diagram of a method of using machine learning (ML) models to identify species-biased antimicrobial peptides (SB-AMPs) that target select microbial populations, in accordance with an illustrative embodiment. FIG. 42 depicts a block diagram of a server system and a client computer system, in accordance with one or more implementations DETAILED DESCRIPTION Following below are more detailed descriptions of various concepts related to, and embodiments of, systems and methods for using machine learning (ML) models to identify species-biased antimicrobial peptides (SB-AMPs) that target select microbial populations. It should be appreciated that various concepts introduced above and discussed in greater detail below may be implemented in any of numerous ways, as the disclosed concepts are not limited to any particular manner of implementation. Examples of specific implementations and applications are provided primarily for illustrative purposes. Section A describes machine learning-enabled design features of antimicrobial peptides selectively targeting peri-implant disease progression. Section B describes interpretable and transparent machine learning guided antimicrobial peptide design empowers targeted inhibition of a keystone pathogen. Section C describes example antimicrobial peptides (AMPs) for targeting select microbial populations. Section D describes systems and methods for using machine learning (ML) models to identify species-biased antimicrobial peptides (SB-AMPs) that target select microbial populations. -17- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Section E describes a network environment and computing environment which may be useful for practicing various computing related embodiments described herein Section F describes antimicrobial peptides of the present technology as well as compositions thereof and methods of use. Various embodiments are described hereinafter. It should be noted that the specific embodiments are not intended as an exhaustive description or as a limitation to the broader aspects discussed herein. One aspect described in conjunction with a particular embodiment is not necessarily limited to that embodiment and can be practiced with any other embodiment(s). As used herein and in the appended claims, singular articles such as “a” and “an” and “the” and similar referents in the context of describing the elements (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the embodiments and does not pose a limitation on the scope of the claims unless otherwise stated. No language in the specification should be construed as indicating any non-claimed element as essential. As used herein, “about” will be understood by persons of ordinary skill in the art and will vary to some extent depending upon the context in which it is used. If there are uses of the term which are not clear to persons of ordinary skill in the art, given the context in which it is used, “about” will mean up to plus or minus 10% of the particular term – for example, “about 10 wt.%” would be understood to mean “9 wt.% to 11 wt.%.” It is to be understood that when “about” precedes a term, the term is to be construed as disclosing -18- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 “about” the term as well as the term without modification by “about”^for example, “about 10 wt.%” discloses “9 wt.% to 11 wt.%” as well as disclosing “10 wt.%.” The phrase “and / or” as used in the present disclosure will be understood to mean any one of the recited members individually or a combination of any two or more thereof^for example, “A, B, and / or C” would mean “A, B, C, A and B, A and C, B and C, or the combination of A, B, and C.” As will be understood by one skilled in the art, for any and all purposes, particularly in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to,” “at least,” “greater than,” “less than,” and the like include the number recited and refer to ranges which can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 atoms refers to groups having 1, 2, or 3 atoms. Similarly, a group having 1-5 atoms refers to groups having 1, 2, 3, 4, or 5 atoms, and so forth. As used herein, the term “peptide” refers to a polymer of amino acid residues joined by amide linkages, which may optionally be chemically modified to achieve desired characteristics. The term “amino acid residue,” includes but is not limited to amino acid residues contained in the group consisting of alanine (Ala or A), cysteine (Cys or C), aspartic acid (Asp or D), glutamic acid (Glu or E), phenylalanine (Phe or F), glycine (Gly or G), histidine (His or H), isoleucine (Ile or I), lysine (Lys or K), leucine (Leu or L), methionine (Met or M), asparagine (Asn or N), proline (Pro or P), glutamine (Gln or Q), arginine (Arg or R), serine (Ser or S), threonine (Thr or T), valine (Val or V), tryptophan (Trp or W), and tyrosine (Tyr or Y) residues. The term “amino acid residue” also may include unnatural amino acids or residues contained in the group consisting of homocysteine, 2-Aminoadipic -19- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 acid, N-Ethylasparagine, 3-Aminoadipic acid, Hydroxylysine, β-alanine, β-Amino-propionic acid, allo-Hydroxylysine acid, 2-Aminobutyric acid, 3-Hydroxyproline, 4-Aminobutyric acid, 4-Hydroxyproline, piperidinic acid, 6-Aminocaproic acid, Isodesmosine, 2-Aminoheptanoic acid, allo-Isoleucine, 2-Aminoisobutyric acid, N-Methylglycine, sarcosine, 3- Aminoisobutyric acid, N-Methylisoleucine, 2-Aminopimelic acid, 6-N-Methyllysine, 2,4- Diaminobutyric acid, N-Methylvaline, Desmosine, Norvaline, 2,2′-Diaminopimelic acid, Norleucine, 2,3-Diaminopropionic acid, Ornithine, and N-Ethylglycine. Typically, the amide linkages of the peptides are formed from an amino group of the backbone of one amino acid and a carboxyl group of the backbone of another amino acid. By “pharmaceutically acceptable” is meant a material that is not biologically or otherwise undesirable, e.g., the material may be incorporated into a pharmaceutical composition administered to a patient without causing any undesirable biological effects or interacting in a deleterious manner with any of the other components of the composition in which it is contained. When the term “pharmaceutically acceptable” is used to refer to a pharmaceutical carrier or excipient, it is implied that the carrier or excipient has met the required standards of toxicological and manufacturing testing or that it is included on the Inactive Ingredient Guide prepared by the U.S. Food and Drug administration. Pharmaceutically acceptable salts of peptides described herein are within the scope of the present technology and include acid or base addition salts which retain the desired pharmacological activity and is not biologically undesirable (e.g., the salt is not unduly toxic, allergenic, or irritating, and is bioavailable). When the compound of the present technology has a basic group, such as, for example, an amino group, pharmaceutically acceptable salts can be formed with inorganic acids (such as hydrochloric acid, hydroboric acid, nitric acid, sulfuric acid, and phosphoric acid), organic acids (e.g., alginate, formic acid, acetic acid, benzoic acid, gluconic acid, fumaric acid, oxalic acid, tartaric acid, lactic acid, maleic acid, citric acid, succinic acid, malic acid, methanesulfonic acid, benzenesulfonic acid, naphthalene sulfonic acid, and p-toluenesulfonic acid) or acidic amino acids (such as aspartic acid and glutamic acid). When the compound of the present technology has an acidic group, such as for example, a carboxylic acid group, it can form salts with metals, such as alkali and earth alkali metals (e.g., Na+, Li+, K+, Ca2+, Mg2+, Zn2+), -20- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 ammonia or organic amines (e.g., dicyclohexylamine, trimethylamine, triethylamine, pyridine, picoline, ethanolamine, diethanolamine, triethanolamine) or basic amino acids (e.g., arginine, lysine and ornithine). Such salts can be prepared in situ during isolation and purification of the compounds or by separately reacting the purified compound in its free base or free acid form with a suitable acid or base, respectively, and isolating the salt thus formed. The peptides of the present technology may exist as solvates, especially hydrates. Hydrates may form during manufacture of the compounds or compositions comprising the compounds, or hydrates may form over time due to the hygroscopic nature of the compounds. Compounds of the present technology may exist as organic solvates as well, including DMF, ether, and alcohol solvates among others. The identification and preparation of any particular solvate is within the skill of the ordinary artisan of synthetic organic or medicinal chemistry. As used herein, “subject” refers to an animal, such as a mammal (including a human), that has been or will be the object of treatment, observation or experiment. “Subject” and “patient” may be used interchangeably, unless otherwise indicated. Mammals include, but are not limited to, mice, rodents, rats, simians, humans, farm animals, dogs, cats, sport animals, and pets. The methods described herein may be useful in human therapy and / or veterinary applications. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. The term “treatment” or “treating” means administering a compound disclosed herein for the purpose of: (i) delaying the onset of a disease, that is, causing the clinical symptoms of the disease not to develop or delaying the development thereof; (ii) inhibiting the disease, that is, arresting the development of clinical symptoms; and / or (iii) relieving the disease, that is, causing the regression of clinical symptoms or the severity thereof. The term “dental implant” refers to a titanium-containing post surgically placed (or to be surgically placed) in the upper or lower jaw of a subject, which functions as an anchor for one or more replacement teeth. The replacement tooth, commonly referred to as a dental crown, may be attached to the implant via an abutment, a connector that supports and holds the tooth. -21- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this present technology belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present technology, representative illustrative methods and materials are described herein. Throughout this disclosure, various publications, patents and published patent specifications are referenced by an identifying citation. Also within this disclosure are Arabic numerals referring to referenced citations, the full bibliographic details of which are provided subsequent to the Examples section. The disclosures of these publications, patents and published patent specifications are hereby incorporated by reference into the present disclosure to more fully describe the present technology. A. Machine Learning-Enabled Design Features of Antimicrobial Peptides Selectively Targeting Peri-Implant Disease Progression Peri-implantitis is a complex infectious disease that manifests as progressive loss of alveolar bone around the dental implants and hyper-inflammation associated with microbial dysbiosis. Using antibiotics in treating peri-implantitis is controversial because of antibiotic resistance threats, the non-selective suppression of pathogens and commensals within the microbial community, and potentially serious systemic sequelae. Therefore, conventional treatment for peri-implantitis comprises mechanical debridement by nonsurgical or surgical approaches with adjunct local microbicidal agents. Consequently, current treatment options may not prevent relapses, as the pathogens either remain unaffected or quickly re-emerge after treatment. Successful mitigation of disease progression in peri- implantitis requires a specific mode of treatment capable of targeting keystone pathogens and restoring bacterial community balance toward commensal species. Antimicrobial peptides (AMPs) hold promise as alternative therapeutics through their bacterial specificity and targeted inhibitory activity. However, peptide sequence space exhibits complex relationships such as sparse vector encoding of sequences, including combinatorial and discrete functions describing peptide antimicrobial activity. -22- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Presented herein isa transparent machine learning (ML) model that identifies sequence-function relationships based on rough set theory using simple summaries of the hydropathic features of AMPs. Comparing the hydropathic features of peptides according to their differential activity for different classes of bacteria empowered the predictability of antimicrobial targeting. Enriching the sequence diversity by a genetic algorithm, numerous candidate AMPs designed for selectively targeting pathogens were generated and predicted their activity using classifying rough sets. Empirical growth inhibition data are iteratively fed back into the ML training to generate new peptides, resulting in increasingly more rigorous rules for which peptides match targeted inhibition levels for specific bacterial strains. The subsequent top scoring candidates were empirically tested for their inhibition against keystone and accessory peri-implantitis pathogens as well as an oral commensal bacterium. A novel peptide, VL-13, was confirmed to be selectively active against a keystone pathogen. Considering the continually increasing number of oral implants placed each year and the complexity of the disease progression, the prevalence of peri-implant diseases continues to rise. The approach offers transparent ML-enabled paths towards developing antimicrobial peptide-based therapies targeting the changes in the microbial communities that can beneficially impact disease progression. 1 Introduction Despite high success rates for dental implants, their bacterial plaque- associated inflammatory lesions, known as peri-implant diseases, still occur (1, 2). These lesions continue to degrade the stability of peri-implant soft and hard tissues, which can result in loss of the implant. While peri-implant mucositis is a reversible inflammatory condition, peri-implantitis is an irreversible pathological condition leading to loss of supporting alveolar bone (2). The reported prevalence of peri-implant mucositis and peri-implantitis shows a substantial increase over time following implant placement. Meta-analysis for patient-based peri-implant mucositis and peri-implantitis was reported as 46.83% and 19.83% by Lee et al. (3). In a separate study, meta-analyses estimated the peri-implant mucositis and peri- implantitis as 43% and 22%, respectively. Peri-implantitis is also reported to be in the range of 11%–47% among dental implants 10 years after their placement (4, 5). These numbers further increase in periodontally compromised patients (1, 6, 7). -23- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Current treatments for peri-implantitis and periodontitis include mechanical debridement, disinfection of exposed implant surfaces, and antibiotic or antiseptic prescriptions to suppress the associated bacteria (5). The use of adjunctive antibiotics for treating peri-implantitis or periodontitis is debated mainly because of concerns about microbial antibiotic resistance, the non-selective suppression of both pathogenic and commensal species, and the adverse systemic reactions. Notably, these conventional treatment modalities may not prevent relapses, as the pathogens may either remain unaffected or quickly re-emerge after treatment. The poor efficacy of antibiotic treatment in peri-implantitis may be explained by the non-specific suppression of dysbiotic biofilms. The adaptability and resiliency of pathogenic bacteria in biofilms is well documented (8). The unique structure and inter- species relationships within a biofilm enhance the individual strengths of the bacteria present, creating an unbalanced community organized to promote communal success at the expense of the host (9). Notably, keystone pathogens play an outsized role in shaping the community structure. Therefore, targeting keystone pathogens may be the most effective approach to reverse microbial dysbiosis and return to health- compatible eubiosis. In peri-implantitis, Porphyromonas gingivalis (P. gingivalis) is widely acknowledged as a keystone pathogen (10). P. gingivalis is associated with increased levels of inflammation and subsequent alveolar bone loss (11). Once P. gingivalis has initiated biofilm growth, other pathogens are free toflourish and further contribute to the dysbiotic community (11, 12). The interdependent-relationships among pathogens in a biofilm are a defining factor in their treatment difficulty. In peri-implantitis, this is evident by the coexistence of Aggregatibacter actinomycetemcomitans, another keystone pathogen associated with aggressive periodontitis, and Streptococcus gordonii, a commensal and accessory pathogen (13, 14). Microbial communities exhibiting both P. gingivalis and S. gordonii are linked to more severe cases of peri-implantitis, resulting in increased infection and bone loss as compared to others (15–17). Undoubtedly, in peri-implantitic biofilms, pathogens grow synergistically to promote each other’s survival. Keystone pathogens, such as P. gingivalis and A. actinomycetemcomitans, play a pivotal role in shifting the oral microbiome to induce the host into a disease-oriented state. The presence of these pathogens -24- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 is widely associated with intensified inflammation levels, prolonged infection, and enhanced alveolar bone loss in patients (18, 19). Successful mitigation of disease progression in peri- implantitis requires a specific mode of treatment capable of targeting keystone pathogens and restoring bacterial community balance toward commensal species. Broad-spectrum approaches have difficulty in preventing oral dysbiosis (11, 18, 20). Antimicrobial peptides (AMPs) have been receiving increasing attention as promising therapeutic candidates since their use leads to no or low antibiotic resistance (21, 22). Moreover, AMPs with a short sequence domain offer straight-forward manufacturing, and relatively low-cost production (23). Within AMPs, antibacterial peptides account for the largest proportion of peptides with inhibitory activities ranging from broad- to specific- species (21). Complexity in structure-function relationships in AMPs is increasing with the increased number of peptide sequences that are isolated from a wide range of organisms, designed using computational search methods, or developed as peptide-mimics as potential candidates (24–27). Despite such progress, and the promise of AMPs as alternative treatments to antibiotics and antiseptics, still only a handful of AMPs have been applied to oral-craniofacial applications (28–33). The AMPs and other groups were investigated that may serve to reduce biofilm load and / or to target emergent keystone pathogens on dental implants, and mitigation of bacterial-induced peri-implantitis has been demonstrated by rationally designed chimeric AMPs and peptides (30, 34–40). As a result of growing interest in AMPs, large databases on their sequence and known functions are now readily available (41–46). With the increasing number of AMPs discovered at the lab bench and via computational methods, determining the boundaries of similar AMPs and identifying their bacteria- specific function remains challenging. Machine learning models offer unique tools tofind AMPs with targeted functionality (42). Through bioinformatics similarity tools, ML models are effective in identifying possible antimicrobial peptides among the large number of nucleic acid sequences (41). Recurrent neural networks (RNNs), long-short term models (LSTMs) and other deep learning methods have demonstrated success in peptide related-prediction models and these methods are now being used in constructive model approaches to design AMPs (47–51). However, when these methods are applied in generative approaches, they severely lack training adaptability. This -25- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 is mainly due to their requirement for a relatively large number of training sets needed to learn specific paths in high-dimensional decision space. This makes them at risk for re- enforcing errors when the models incorrectly classify cases by correlated features. Customizing decisions for the sample distribution may optimize prediction performance, but this makes the ability of the model to adjust to the trends in unbiased sampling data extremely difficult to achieve. Gradient descent or backpropagation methods can help recognize the contribution from previously under-represented subgroups that are negatively impacting the models’ applicability and overall efficacy. However, since retraining the whole model comes with a large computational cost, alternative methods are sought to address the bias. Still, no strategies are currently known that solve this issue. In the previous work, the use of rough set theory was pioneered for the classification of peptide sequences and demonstrated how to achieve training adaptability by bypassing neural networks in the context of AMP identification (52). In this approach, different descriptors in accordance with antibacterial activity are analyzed by rough set theory (RST) which is used as a heuristic method to discover the rules distinguishing different outcomes. By combining the rough set theory approach with the algorithm of Modified Learning from Examples Module, Version 2 (MLEM2) and the Interesting Rule Induction Module (IRIM) algorithm, high-specificity performance was achieved. The method provides a transparent selection approach to define explicit boundaries that distinguish between classes of AMPs by their activity. The model can adapt by using the explicit decision components and the related rules that are introduced by new hypotheses and labelled data. Non-linear categories distinguishing between active and inactive peptides reduce training time of the model, while preserving the structure of the explicit model choices. This improves training flexibility and avoids the cost barrier associated with re-training or the creation of new models. This method also guards against irrational decision relationships by maintaining transparency for each decision step throughout the decision process. In a separate study, the RST based ML approach (CLN-MLEM2) was combined with a codon-based genetic algorithm (CB-GA) and increased the variations of peptide sequences generated by RST ML search (53). Using the CB-GA combined ML approach, an AMP sequence effective against Staphylococcus epidermidis was identified. The training false discovery rate, i.e., probability of false positives, was approximately 5% (53). -26- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 In this study, a transparent ML model was developed and was combined with a genetic algorithm, that empowers AMP design targeted to a keystone pathogen. The predicted activity of the generated AMPs was classified using rough sets and the rules were improved using empirical growth inhibition data for specific pathogens. The validation tests were run with the peptides having the highest inhibition predictability score against the keystone pathogen, A. actinomycetemcomitans. A novel peptide VL-13 was confirmed to be active against the selected keystone pathogen without compromising the accessory- commensal species, S. gordonii. 2 Materials and methods 2.1 Materials A. actinomycetemcomitans strain D7S-1, S. gordonii strain Challis, and Streptococcus sanguinis strain ATCC10556 were cultured using either modified Trypticase Soy Broth (mTSB) containing 3% trypticase soy broth and 0.6% yeast extract or on mTSB agar (mTSB with 1.5% agar Becton Dickinson and Company). In some experiments, the bacteria were cultured in SHI medium supplemented with hemin (5 µg / ml) (Sigma- Aldrich, St. Louis, MO, USA), menadione (1 µg / ml) (Sigma- Aldrich), human serum (10%) (Sigma- Aldrich), and sucrose (0.25%). Bacteria were cultured at 37 °C in a humidified atmosphere supplemented with 5% CO2 (54). P. gingivalis strain ATCC33277 was cultured in the brain- heart infusion (BHI) broth at 37 °C under anaerobic conditions. 2.2 Bacterial viability tests Bacterial viability tests were done byfirst adjusting the optical density of bacterial cultures to 0.2 (equivalent to approximately 107CFU per ml) at 600 nm. The bacterial cultures were then diluted 1:20 to get to 5 × 105CFU / ml and incubated with 100 µM AMPs. Bacterial viability was then determined at 0 and 6 h by CFU counts and at an additional 24-h time point for P. gingivalis. 2.3 Machine learning model 2.3.1 Initial datasets -27- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 The complete listing of the antimicrobial peptides used to generate the targeted rough set theory rules is given in Supplementary Table S1. In addition to literature- derived peptides, additional antimicrobial peptides was included previously studied for different applications shown in Table 1. The initial datasets for peptide generation were taken from the iAMP-2l database (55). This database wasfiltered to only include examples of anti-bacterial peptides, resulting in 1,274 unique peptides from the database. Of the 21 peptide sequences provided in Table 1 and Supplementary Table S1, 15 of them were already included in the database. Therefore, a total of 1,280 AMP sequences are included in the initial set. TABLE 1 Rough set theory identities for in vitro inhibition against A. actinomycetemcomitans (Aa), S. gordonii (Sg), and P. gingivalis (Pg). 2.3.2 Peptide customization by rough set theory The rough set theory rules were generated as described in an earlier publication on the CLN-MLEM2 method developed for classification of antimicrobial peptides with two enhancements (52). Previously, the rough set rules were designed to establish the boundaries simply between active and inactive antibacterial peptides, using non- -28- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 correlated AAindex1 properties. Thefirst enhancement that was made is to combine the rough set rules from the previous paper with the targeting rough set theory rules and generate a multiple-dimensional view of the predicted activity. To generate these targeted activity rules, a set of AMPs with confirmed antibacterial activity against any of the three pathogens of interest, i.e., P. gingivalis, S. gordonii, A. actinomycetemcomitans, were used as positive data set (see Table 1, Supplementary Table S1 and Figures 11–15). The second enhancement for targeting was introduced by focusing on the key physicochemical property features. Eight indices were integrated proposed by a recent study as reduced AAindex (rAAindex) obtained from a subset of original 544 indices in the amino acid index database (56). Kibinge et al., applied a random forest (RF) algorithm for property reduction and maximizing metadata capturing. With the two enhancements introduced, this method creates collections of discriminating attributes separating targeted AMP activity trends focused on hydropathy variations for peptide sequences. The resulting rules are each characterized by a hydropathy property of importance determined by the updated CLN-MLEM2 method, and a simple summary characteristic that portrays the features to be used to predict an AMP`s targeted ability. Sequence features that are most relevant for the observed peptide activity are collected as simple arithmetic summaries. The summary characteristics used included the sum of the property across the amino acids of the sequence, the mean of the property across the amino acids of the sequence and the maximum value of the property across three consecutive amino acids within a sequence (50). These summary characteristics correspond to non- linear boundaries between activity classes. Each time a rule set is generated, these boundaries and properties are chosen to separate the sequences into the desired classification groups. The heuristic goal of the method is to create definitions of activity classification with the minimum number of rules and conditions per rule possible. Rule sets contain descriptions of active and inactive peptides, but likely do not contain the set of all peptides in the union of the rule sets. Peptides that do not meet any of the rule sets are not identified directly by the classification method. Since specific peptide activity is not likely to be present for randomly selected sequences, peptides were imputed as inactive if they do not belong to any rule set positive for activity. -29- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 2.3.3 Sequence expansion by codon-based genetic algorithm The codon-based genetic algorithm was used for sequence expansion reported in one of the earlier publications (53). To begin this process, a total of 1,280 AMP sequences was used in the initial set, which contained a large variety of sequences to recombine and mutate through artificial genetic operations. These sequences were subsequently ranked by which generated sequences met the rough set theory rules for targeted antimicrobial activity. The rule sets have two separate descriptions of antimicrobial activity in this study. Thefirst description is the rough sets previously used as a measure for broad-spectrum activity estimation. The second description is the newly generated rules targeting the chosen periodontal keystone pathogens. This second level of activity distinguishes between keystone-only activity and other antimicrobial activity with the goal of avoiding impacting commensal species. These two descriptions were weighed as components of thefitness function. There is a large disparity between the number of previously trained sequences, i.e., 2,347 sequences, and the empirically tested targeted sequences, i.e., 21 sequences (Table 1 and Supplementary Table S1). The rule counts were therefore independently calculated and normalized before being combined as separate terms in thefitness objective function. Fitness objective function were defined for the codon-based genetic algorithm by the following equation in which AB is referred as antibacterial: ^^^^^^ ൌ min^10 ∗ ^^^^^^^ ^^^^^^^^^^^^ ^^^^^^^^ ^^^^^^^^^^ െ ^^^^^^^^^^^^^^^^ ^^^^^^^^^^^^ ^^^^^^^^ ^^^^^^^^^^^^0.01 ∗ ^^^^^^^ ^^^^ ^^^^^^^^ ^^^^^^^^^^ െ ^^^^^^^^^^^^^^^^ ^^^^ ^^^^^^^^ ^^^^^^^^^^^ (1)^5 ∗ min ^^10, 30^ െ ^^^^^^^^^^^^^^^^ ^^^^^^^^^^ ^^^^^^^^ ^^^^^^^^^^ℎ^^The mutation rate for sequences to go to the next generation is 25%. Mutation changes a codon in the sequence, which may not result in an amino acid change or could result in a stop codon. The cross- over rate was 50%. Crossing over was completed with codon representation, often resulting in frameshifts for new candidate sequences compared to the parent codon sequences. The generations were monitored for convergence, both with the bestfitness between generations and the consistency of the targeting rules to generate the top-scoring sequences. Generations are deemed mature for identifying top candidates when the majority -30- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 of sequences meeting the targeting rules are within one half of the maximum number of targeting rules. 2.3.4 Peptide synthesis Peptides were synthesized using Wang resin following a standard Fmoc chemistry method using an Aapptec Focus XC peptide synthesizer. Dimethylformamide (DMF) and 20%–40% piperidine in DMF were used for Fmoc deprotection with two repetitions. The peptide-resins were then washed with DMF. Activation of 0.2 M amino acids / DMF (2 equivalents) was performed by addition of 0.2M 2-(1H-benzotriazol-1-yl)- 1,1,3,3- tetramethyluronium hexafluorophosphate (HBTU) / DMF. The coupling step was completed twice, and the procedure was repeated until the complete peptide was assembled on the solid resin support. Following synthesis, the peptide-resin was removed from the reaction vessel using DMF. Following the removal of DMF from the peptide-resin by washing with ethanol, a cleavage cocktail (15 ml / 1 gram of resin) was added to the dried resin for 2 h with gentle stirring to remove the peptide from the solid support and remove the side chain protecting groups. The standard cleavage cocktail is composed of trifluoroacetic acid (TFA) / triisopropylsilane (TIS) / water (95:2.5:2.5, % vol / vol / vol). To remove side chain protecting groups from peptides containing histidine or cysteine, 2.5% thioanisole and 2.5% 1,2 ethanedithiol were added to the cocktail and for peptides containing methionine, tyrosine, or arginine, 5% phenol was added. The cleavage products werefiltered, and crude peptide product was isolated by precipitation in cold ether. The crude peptide was pelleted by centrifugation (2,000 rpm for 2 min), the supernatant was removed, and the process was repeated for two to four times prior to lyophilization of the peptide products. 3 Results In this study, a transparent ML model was developed that allowed for iterative training sets to be incorporated and enrich the sequence space by a genetic algorithm to design antimicrobial peptides with inhibitory activity targeted to periodontal keystone and accessory pathogens. To identify antimicrobial peptides specific to oral pathogens, rough set boundaries were first established for training rules, building upon the initial iteration of known AMPs and their antimicrobial activity (Figures 11–15, Table 1, and Supplementary -31- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Table S1). The rough set theory classifier was trained to identify possible antimicrobial peptides specific to oral keystone and supporting pathogens, with separate rule sets for each strain. Sequence diversity was expanded using a genetic algorithm and selected novel candidates consistent with their predicted inhibitory activities using the identified training relationships. These candidates are generated by the second iteration. While the rough set rules generated apply to many other sequences than the sequences that were trained on, the rules will not, in general, cover all sequences. Non-conforming sequences are imputed as non-targeted. Therefore, the second iteration is focused onfinding the sequences which are like the sequences that were identified as active against a targeted species in thefirst iteration considering the feature properties found to be discriminating between active and inactive against that single species. Training sequences do not need to have specific activity to generate sequences with targeted activity; multiple examples of non-specific peptides can still provide direction for what features are needed in generating rules. Figure 1 provides the selective targeting antimicrobial peptide design scheme which includes training the next iteration of the model with known AMPs, establishing training rules, expansion of sequence diversity, candidate selection and verification of their predicted inhibition properties targeted to specific organisms. The targeting rules are heuristically made to be a minimal set that covers all the training sequences. 3.1 Establishing boundaries for growth inhibition rules The machine learning approach builds upon the previously developed CLN- MLEM2 method that utilizes rough set theory principles (52). The CLN-MLEM2 method separates sequences by their amino acid properties to establish functional classification. This method simultaneouslyfilters which properties are key properties and provides boundaries for classification. The key properties identified by the CLN-MLEM2 method are provided in Supplementary Table S2. These properties were selected from the rAAindex (56), a reduced subset of the AAindex focused on hydropathy. The machine learning training was initiated by providing literature data to set up the initial inhibition descriptions for rough set participation to address selected pathogens (Supplementary Table S1). The rough sets are key property summary descriptions (e.g., property sum, property mean, property peak window) that identify peptides to have a certain activity level. The key property summary -32- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 descriptions are used as features in the ML method (57). Both the full sequence length properties and short sequence segments are included as features. The short segments are summarized for a sequence by selecting the property peak window features. The full- sequence length features are included as the property sum and the property mean values. The CLN-MLEM2 method used 6 of the 8 properties in the rAAindex to build the rule sets. The selected properties and their short descriptions are included in Supplementary Table S2. The inhibition activity data captured from the literature was also supplemented by incorporating the inhibitory activity from additional antimicrobial peptide sequences into the model. These sequences have been shown to be active in different contexts; therefore their activity was evaluated against selected oral pathogens, P. gingivalis, A. actinomycetemcomitans, and S. gordonii (see Figures 11–15). The level of minimum inhibitory activities is categorized from low to high and provided in Table 1. The ratio of correctly identified cases to the total number of applicable cases for a set is identified as α (0 ≤ α ≤ 1) (58–60). The CLN-MLEM2 method selects rules based on α (0 ≤ α ≤ 1). Using higher values of α generates fewer rules with higher probability (Pr) values of training accuracy if generated rules do not meet the accuracy specification. Using lower values of α generates more rules with lower Pr values of training accuracy when rules do not meet the accuracy specification. For all the training peptide sets, every rule with α = 1.0 was found. Therefore, all generated rules meet the maximum accuracy specification of 1.0. The current descriptions allow for enough discrimination to uniquely describe all peptide sequences when they differ in activity. Hydropathy feature boundaries which explicitly classified the training examples were found. The CLN-MLEM2 rules for predicting the activity against keystone pathogen members A. actinomycetemcomitans and P. gingivalis were generated next, and accessory-commensal class member S. gordonii. These rules provide design criteria for either increasing or decreasing the probability of a peptide sequence having an antimicrobial activity against each strain. The initial design strategy involvedfinding an antimicrobial peptide which heuristically has as many features as possible to be active through the genetic algorithm to satisfy the maximum count of non-linear boundary rules simultaneously through computational search. -33- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Table 2 shows the selected sequence property rules that are associated with the inhibition of A. actinomycetemcomitans and S. gordonii. The sequence summary features in this table provide conditions for classifying peptides with potential inhibitory properties. A detailed analysis of the A. actinomycetemcomitans inhibition rules in Figures 16–21 were provided. The rules positive for A. actinomycetemcomitans inhibition were found to be consistent with thefirst iteration identities and apply to either slightly polar mean polarity or to slightly non-polar mean polarity sequence descriptions, depending on the scale used. Interestingly, instead of the descriptions being mutually exclusive, VL-13 combined membership of both rules into a single peptide sequence. TABLE 2 A set of CLN-MLEM2 rules for A. actinomycetemcomitans (Aa) and S. gordonii (Sg). -34- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 3.2 Ranking antimicrobial peptides by rough set theory relevance -35- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Once “rough set” boundaries are established from thefirst iteration, antimicrobial peptides in the database can be compared in relation to the selected physical properties of their amino acids. The selectivity of the training methods were explored for the iAMP-2l database, which included 2,347 peptides (55). Using the CLN-MLEM2 rules, this database was ranked and generated a relatively small number of sequences of interest for further exploration (Figure 2). After combining the peptides in the iAMP-2l database and the initial testing set, only 20 peptide sequences (0.9%) met conditions for CLN- MLEM2 rules, the net of which were rules for inhibition instead of non-inhibition. Overall, only 20 peptides in the iAMP-2l database (∼2,200 peptides) had similar hydropathy features to the confirmed active AMPs and thus met the CLN-MLEM2 rules derived from thefirst iteration in antimicrobial activity prediction. Selecting peptides with cysteine to enable rational cyclization studies using disulfide bonds in the future was avoided. Therefore, the third- ranked peptide KF-18 (KWKLFKKIPKFLHLAKKF (SEQ ID NO: 13)) was selected directly from the database to synthesize and to evaluate its in vitro activity. To the knowledge, inhibitory activity for this peptide is not reported for oral bacteria. 3.3 Identifying candidate antimicrobial peptides with enhanced relevance to existing rough sets A codon-based genetic algorithm (53) to identify antimicrobial peptides relevant to inhibition-related rough sets was recently developed, as well as no growth inhibition rough sets. The codon based- genetic algorithm uses reading frameshifts probabilistically to generate new amino acid sequences that have low sequence similarity to previously generated sequences. The codon-based operations supplement the recombination and mutation operators. Peptides identified as possessing antimicrobial properties for at least one of three target bacterial strains were targeted. The peptides were further ranked by length and solubility estimates. Next, the later-generation peptide sequences were enriched with sequences that met the rough set criteria. In Figure 3, this enrichment can be seen in the net-positive inhibition rule sequences where the initial generation has increased from 20 sequences -36- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 (Figure 2) to 269 sequences by Generation 10. Generation 25 has a tighter distribution of net inhibition rules compared to Generation 10, which resulted in 329 peptide sequences. These sequences also contained the maximum number of net inhibition rules observed with the top scoring sequences transferred between generations. The end of 25th generation run resulted in 2,261 sequences, from these three peptides were selected to evaluate their inhibition activity: KK-15 (KWKLFKTTAKFLHLAK (SEQ ID NO: 14)), FV-11 (FLHWVPLRRVV (SEQ ID NO: 15)) and VL-13 (VDWKKVFGKLLKL (SEQ ID NO: 16)) (Figure 4). Many of the top scoring sequences contained cysteine and they were avoided to enable future cyclization of candidates through rational placement of disulfide bonds. A high number of rules were chosen (>7), which were satisfied by KK-15 and VL-13. FLFAFFRALRHVGK (FK-14 (SEQ ID NO: 17)), LKLLKRLLKLLKK (LK-13-1 (SEQ ID NO: 18)) and LKLLKKLLKLLKK (LK-13-2 (SEQ ID NO: 19)) peptides were not selected. The LK-13 and LK-13-2 peptides were observed as close analogues of AMP1, only missing the last lysine residue or also having an arginine-lysine substitution. Since AMP1 is already included in the study, other candidates were selected. Peptide FK-14 was not selected because it has a GRAVY score of +0.72, indicating a solubility risk. However, the shorter peptide was selected, FV-11, because it has less of a solubility risk with a high score of 7. In summary, the attention on potential solubility and possessing the most applicable rules were focused on. These candidates were among the top 25 sequences with 13 amino acids or less that fell within the 98th percentile of net inhibition rules. The maturation of generating new candidates with similar net inhibition rule counts was completed by the 25th generation. Noteworthy is thefinding in Figure 3 that the 11th generation shows large variations of net inhibition rules among the candidates. Experimental evaluation of ML-generated sequences, KK-15, FV-11, and VL- 13 was performed against keystone pathogen A. actinomycetemcomitans, accessory / commensal S. gordonii and commensal S. sanguinis (Figures 5–7, respectively). These second iteration peptides were compared with twofirst iteration peptides, AMP1 and AMPa. Table 3 shows the scores for targeting keystone pathogens. VL-13 showed strong inhibition activity selectively against A. actinomycetemcomitans compared to the other peptides FV-11, KK-15, and KF-18. The score for targeting keystone pathogen is the difference between the number of CFU logs reduced after 6 h. VL-13 has the largest -37- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 targeting score, having +7 more log reduction for the keystone pathogen A. actinomycetemcomitans than for the commensal S. gordonii. VL-13 also had a high score for targeting keystone pathogen of +4.5 when compared to the other commensal strain S. sanguinis. The other predicted peptides had positive targeting scores, but less than AMP1. Peptides used in the training sets AMPa resulted in the lowest targeting score of only a 0.5 log reduction difference between the two groups; AMP1 also resulted in activity against S. sanguinis. VL-13 was validated for predicted targeted activity against A. actinomycetemcomitans. The superimposed structures generated for VL-13 are given in Figure 8. From the structures, a peptide structural feature with a high amount of hydrophobicity is indicated as a low-energy rotation barrier compared to the more hydrophilic structural features in the antimicrobial peptide. This hydrophobicity feature is also discussed in the rough set theory sequence feature example with the JACR8901013-amino acid window in Figure 9 and in Figure 16. As a parallel method to assess the progress of the CLN- MLEM2 model, the performance of additional peptides reported as potential therapeutics for oral bacteria and oral biofilms as the test set was predicted (61). In Supplementary Tables S3 and S4, the performance of the rules specific to two keystone pathogens, A. actinomycetemcomitans and P. gingivalis, respectively were evaluated. In Supplementary Table S5, the performance of the rules specific to the commensal / accessory pathogen S. gordonii were evaluated. The false discovery rate was found to be low for A. actinomycetemcomitans and the commensal pathogen, S. gordonii (Supplementary Tables S3 and S5) using prediction performance. Overall, the model performance shows good prediction for positive activity of the sequences for selected keystone pathogen and the commensal / accessory pathogen. 4 Discussion 4.1 Shifting focus toward pathogen A. actinomycetemcomitans Treatment of periodontal and peri-implant diseases necessitates a targeted, polymicrobial approach that sufficiently inhibits progression of the disease state without -38- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 detrimentally impacting commensal and otherwise opportunistic species. To address this, a ML-enabled targeting AMP prediction approach was attempted. To train the CLN-MLEM2 model, an initial round of AMPs was tested against keystone pathogens P. gingivalis and A. actinomycetemcomitans, and opportunistic pathogen S. gordonii. Both P. gingivalis and A. actinomycetemcomitans display pathogenic characteristics that contribute significantly to biofilm prevalence and virulence (11). Although these species are referred as key stone pathogens, their involvement appears in different stages of the disease progression. The ML- based tunability in this paper presents an opportunity to control the disease progression by targeting different keystones and other microbiome components. A. actinomycetemcomitans is unique for its association with localized, aggressive cases of peri-implantitis and periodontitis—specifically those in younger individuals under the age of 35 (13, 62). This is extremely problematic not only for the livelihood of impacted individuals, but also for the whole of human health. Disease occurrence in younger individuals entails longer timelines of recurrent infection and treatment cycles. This makes these individuals, and thus bacterial species present in them, prime candidates for emergent bacterial resistance and innate microbiome depletion. It also makes them at heightened risk for implant installation failure and loss of oral functionality as the tissue and bone surrounding their dental implants degrade. In the ML-predicted novel peptide generation, therefore focused on assessment of the antimicrobial activity against A. actinomycetemcomitans, S. gordonii, and the commensal S. sanguinis. TABLE 3 Net inhibition rule counts by bacterial species. -39- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Positive scores in the “NC” columns indicate predicted antibacterial activity for the strain. Zero or negative scores in these columns indicate no activity predicted. The “LR” columns indicate the decimal log change after 6 h of incubation with the peptide indicated in Figures 5–7. Keystone targeting is the difference between the A. actinomycetemcomitans (Aa) log reduction and the minimum of the S. gordonii (Sg) log reduction or the S. sanguinis (Ss) log reduction. Higher log reduction scores indicate more growth inhibition. Negative log reduction indicates growth during incubation instead of inhibition. Higher targeting scores indicate better targeting performance. Thefirst two rows arefirst iteration sequences with known activity used to compare the targeting performance of the second iteration peptides generated by the codon-based genetic algorithm. The Aa inhibition rules describe transferred activity for VL-13 but not for KK-15, while the Sg inhibition rules did not describe transferred activity in any of the second iteration peptides. NC, net count of inhibition rules; LR, log reduction; and ND, not determined. 4.2 Rule application and predicted vs. experimental antimicrobial activity in VL- 13 The rules in Table 2 show amphipathic descriptions of targeting growth inhibition of A. actinomycetemcomitans, both for descriptions of which peptides inhibit and which peptides do not inhibit growth. In Figures 16–21, the feature characteristics for sequences that also have rule membership for rules which apply to VL-13 in Table 2 were -40- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 described. The second-iteration decision system made incorrect predictions for KK-15 and FV-11. The next iteration of the decision system will avoid these incorrect predictions and generalize with these new cases to determine new decision boundaries. The next decision system draws more accurate boundaries for negative results but does not move boundaries when all cases within a sub-domain are accurate. This attribute of the transparent decision system shows the system development value of having test cases. This way the decision system challenges the rough set membership boundaries of the previous iteration rather than having test cases which are very likely to be accurate or peptides which have no rough set theory membership. Testing truly random peptides which do not belong to either rules for or against activity may not build on the knowledge gained from previously tested peptides. However, any peptide test results would start a new knowledge base to build on in future studies. A nested-rough set rules approach was used by evaluating the rules generated when classifying peptides for having any antibacterial activity with rough set rules for peptides having targeted activity. This nested methodology is an example of transfer learning in the machine learning method. Further targeting of the activity of the peptides incorporating different design goals can also adapt this nesting approach in future studies. In thefirst iteration of the codon-based genetic algorithm, generated a novel antimicrobial peptide against S. epidermidis (53). In the second iteration of this system in this work, a nested version of the decision system was applied to the in vitro oral environment for the mitigation of the progression of peri-implantitis. In this study, an antimicrobial peptide, VL-13, was demonstrated targeted growth inhibition against an oral keystone pathogen A. actinomycetemcomitans without inhibiting accessory / commensal species, S. gordonii (Figure 5). VL-13 is both a positive result for this decision system iteration for activity against the keystone pathogen and a negative result for activity against the accessory / commensal pathogen S. gordonii. Peptides were designed to meet multiple targeting classes, inferring that if the targeting criteria are relatively difficult to describe, finding one working targeting description would likely come before two working targeting -41- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 descriptions. The result does not limit building on the targeting criteria which could be introduced as new design characteristics. In future studies, both experimental results can be used to draw boundaries which include VL-13 for A. actinomycetemcomitans inhibition and exclude VL-13 for S. gordonii inhibition. The method learns boundaries from inhibition and non-inhibition results for each of the tested peptides. Superimposed structures of AMPa and AMP1 were evaluated to gain further insight, as both sequences were used in the training set based on their demonstrated activity (Figure 10). The folding dynamics of these structures have relatively high folding entropy compared to VL-13, shown in Figure 8. To discuss what the rough set theory rules imply about inhibition activity for the keystone pathogens, begin with one AAindex1 property. The three-amino acid window JACR890101 is the amino-acid wise component of hydrophobic interactions of amino acids at the bilayer (Figure 9, Figure 16). While further study can investigate if this feature is necessary for the peptide’s activity, also note that this feature may be related to some motion allowing the peptide to efficiently attack / bypass the membrane of A. actinomycetemcomitans, which is a gram-negative pathogen. In Rules 1 and 3 in Table 2, this feature was selected as a tripeptide window or a mean (see Figure 10). The rules both select tripeptide windows for this feature, in which Rule 1 has a left-shifted range (from 0.89–1.32 to 0.615– 0.935). This window is very hydrophobic under the conditions of the bilayer (see Figures 16–21). The overall mean of the peptide for this property in Rule 1 was between −1.2 and −1.8, indicating that these tripeptide windows of greater than 0 are not likely to be common among peptides which are active against A. actinomycetemcomitans. The combination of these rules describes a preliminary description of amphipathy that is useful to inhibit the pathogen’s growth. Having varying hydrophobicity has long been studied for antimicrobial peptides (57, 67, 68). Further non-linear boundary ranges are shown in Figures 11–18. The CLN-MLEM2 rough set theory method has added a process to identifying which hydrophobicity features relate to inhibiting the growth for a specific pathogen. The target that the antimicrobial peptide is affecting with the hydrophobicity feature is unknown. 4.3 Testing performance, limitations and future perspectives -42- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 This paper used data from a review article on peptides as therapeutics targeting oral bacteria by Sztukowska et al. (61) as a test set. Previous reviews of antimicrobial peptides in the oral environment with peptide sequences that were not included in the training set exist (69–72). The reasons to not include all experimentally tested peptides in the model training phase is to continue to develop the model through testing performance beyond training performance. Having peptides in the literature that are not included in the training set allows for testing the performance evaluation of the model. Indeed, the ML model should be validated with experimental results. The literature peptide activity (61) was used as a test set for the rule sets describing targeted activity for the two keystone pathogens and the commensal / accessory pathogen. Thefirst keystone pathogen rules for A. actinomycetemcomitans had higher accuracy than the commensal pathogen rules for S. gordonii, indicating a closer relationship between trained sequences and tested sequences by the rule set descriptions related to Supplementary Table S3 than to the rule set descriptions in Supplementary Table S5. The rule set for the second keystone pathogen had a high false discovery rate, indicating the rule set related to Supplementary Table S4 for P. gingivalis has a higher chance of leading to unexpected negative inhibitory results. These results also confirm the critical iterations of the enhanced rule sets for different pathogens including keystones such as P. gingivalis. The future studies focus on developing rule sets leading to low false discovery rates (<10%), then incorporate this bacterial strain in the ML-candidate evaluation process. The ML-models are re-trained between iterations to avoid carrying over recognized performance errors. Hydrophobicity trends among peptides are discovered during training and applied to the selection of new peptides. The method is identifying database peptides that have similar hydropathy features to the peptides that are confirmed for their antimicrobial activity using in vitro evaluation against the targeted bacterial strains. These activities build better information for peptides with similar hydropathy features, rather than using brute-force testing on all related peptides to see if their inhibition activity changes in a useful way between the targeted bacterial strains. Incorporation of new information from in vitro results strengthens the approach for either positive or negative results. The fact that the peptides studied here need to be further evaluated for their toxicity is recognized. The future work -43- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 will include experimental evaluation of the potential peptide toxicity as well as other clinically relevant properties of these peptides. Plans to incorporate multiple factors into the ML design of targeted peptides is also planned. Broader types of data, such as toxicity and stability, can be tightly integrated together with inhibitory activity because each new factor will have its own rule sets to simultaneously constrain the sequences the genetic algorithm will target. These distinct descriptions can be integrated with the inhibitory activity descriptions by building separate rule sets and using the genetic algorithm tofind examples that combine multiple rule sets for distinct descriptions. The transparent approach of rough set theory allows us to re- classify sequences based on new results in a short time without using GPU-parallelized computation. Therefore, re-training of the entire dataset is feasible when integrating data between sources and extending data sources. During the training, hydrophobicity trends among peptides were discovered and applied to the selection of new peptides. The validation tests were run on three top candidates, one of which had the highest inhibition predictability score against the keystone pathogen, A. actinomycetemcomitans. The in vitro test results demonstrated that peptide VL- 13 is an antimicrobial peptide with the largest change of inhibition compared between the keystone pathogen and the accessory-commensal species, S. gordonii. Further, VL-13 has an added advantage in possessing the largest change of inhibition between the keystone pathogen and the commensal species, S. sanguinis. 5 Conclusion A transparent machine learning model was developed with an iterative in vitro experimental validation approach to design antimicrobial peptides that selectively target keystone pathogens believed to play critical roles during biofilm dysbiogenesis leading to progression of peri-implant disease. The transparency of the machine learning method allows us to compare the discovered relationships with trends in literature. The non-linear nature of the boundaries also provides for rapid learning between iterations of new sequences to explore. Through the transparent machine learning methods, thefinding that antimicrobial peptide VL-13 inhibits the keystone pathogen A. actinomycetemcomitans is demonstrated, while the ML model also provided better learning descriptions forfinding antimicrobial -44- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 peptides inhibitory against the accessory pathogen S. gordonii. VL-13 was the most targeted antimicrobial peptide sequence tested, with high levels of inhibition against A. actinomycetemcomitans and minimal impact on S. sanguinis and S. gordonii. The descriptions used to select the VL-13 sequence from the candidate sequence generation method were learned from tested peptides in this study combined with known sequences in the literature. From antimicrobial peptide sequences available in iAMP-2l antimicrobial peptide database, KF-18 was chosen to test for activity. KF-18 has strong homology with AMPa, which demonstrated is inhibitory toward A. actinomycetemcomitans. Since this peptide and a second generated candidate with homology to AMPa (KK-15) did not test as active, there is further insight into the features of AMPa which relate to inhibition activity for A. actinomycetemcomitans and S. gordonii. Future work aims to generate non-inhibitory sequences for commensal and accessory species, such as S. gordonii, while retaining inhibitory activity against keystone pathogens. Developing an engineering approach to iteratively discover targeted antimicrobials is a robust method for potential therapeutic treatment for peri-implant biofilm infections and thus to reduce the resulting host response of hyper-inflammation during disease progression. With increasing use of implants to replace missing teeth and support oral function, the number of patients suffering from peri-implantitis will continue to increase. It is critical tofind targeted approaches to address this complex infectious disease. Antimicrobial peptides with selective bioactivity against keystone pathogens, while preserving commensals could be the next generation therapy that will respond to this urgent healthcare need. References 1. Iacono VJ, Bassir SH, Wang HH, Myneni SR. Peri-implantitis: effects of periodontitis and its risk factors—a narrative review. Front Oral Maxillofac Med. (2022) 5:27. doi: 10.21037 / fomm-21-63 2. Berglundh T, Wennström JL, Lindhe J. Long-term outcome of surgical treatment of peri-implantitis. A 2–11-year retrospective study. Clin Oral Implants Res. (2018) 29 (4):404–10. doi: 10.1111 / clr.13138 -45- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 3. Lee CT, Huang YW, Zhu L, Weltman R. Prevalences of peri- implantitis and peri- implant mucositis: systematic review and meta-analysis. J Dent. (2017) 62:1–12. doi: 10.1016 / j.jdent.2017.04.011 4. Ardila CM, Vivares-Builes AM. Antibiotic resistance in patients with peri- implantitis: a systematic scoping review. Int J Environ Res Public Health. (2022) 19 (23):15609. doi: 10.3390 / ijerph192315609 5. Roccuzzo A, Stähli A, Monje A, Sculean A, Salvi GE. Peri- implantitis: a clinical update on prevalence and surgical treatment outcomes. J Clin Med. (2021) 10(5):1107. doi: 10.3390 / jcm10051107 6. Schwarz F, Derks J, Monje A, Wang HL. Peri-implantitis. J Clin Periodontol. (2018) 45:S246–66. doi: 10.1111 / jcpe.12954 7. Romandini M, Shin HS, Romandini P, Laforí A, Cordaro M. Hormone-related events and periodontitis in women. J Clin Periodontol. (2020) 47(4):429– 41. doi: 10. 1111 / jcpe.13248 8. Wicaksono WA, Erschen S, Krause R, Müller H, Cernava T, Berg G. Enhanced survival of multi-species biofilms under stress is promoted by low-abundant but antimicrobial-resistant keystone species. J Hazard Mater. (2022) 422:126836. doi: 10.1016 / j.jhazmat.2021.126836 9. Costerton JW. Introduction to biofilm. Int J Antimicrob Agents. (1999) 11(3– 4):217–21; discussion 37–9. doi: 10.1016 / s0924-8579(99)00018-7 10. Hajishengallis G, Darveau RP, Curtis MA. The keystone-pathogen hypothesis. Nat Rev Microbiol. (2012) 10(10):717–25. doi: 10.1038 / nrmicro2873 11. Hajishengallis G, Lamont RJ. Beyond the red complex and into more complexity: the polymicrobial synergy and dysbiosis (psd) model of periodontal disease etiology. Mol Oral Microbiol. (2012) 27(6):409–19. doi: 10.1111 / j.2041-1014.2012.00663.x -46- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 12. Costalonga M, Herzberg MC. The oral microbiome and the immunobiology of periodontal disease and caries. Immunol Lett. (2014) 162(2 Pt A):22–38. doi: 10.1016 / j.imlet.2014.08.017 13. Fine DH, Markowitz K, Furgang D, Fairlie K, Ferrandiz J, Nasri C, et al. Aggregatibacter actinomycetemcomitans and its relationship to initiation of localized aggressive periodontitis: longitudinal cohort study of initially healthy adolescents. J Clin Microbiol. (2007) 45(12):3859–69. doi: 10.1128 / jcm.00653-07 14. Zhu B, Macleod LC, Newsome E, Liu J, Xu P. Aggregatibacter actinomycetemcomitans mediates protection of Porphyromonas Gingivalis from Streptococcus Sanguinis hydrogen peroxide production in multi-Species biofilms. Sci Rep. (2019) 9(1):4944. doi: 10.1038 / s41598-019-41467-9 15. Pollanen MT, Paino A, Ihalin R. Environmental stimuli shape biofilm formation and the virulence of periodontal pathogens. Int J Mol Sci. (2013) 14(8):17221–37. doi: 10.3390 / ijms140817221 16. Lamont RJ, Koo H, Hajishengallis G. The oral Microbiota: dynamic communities and host interactions. Nat Rev Microbiol. (2018) 16(12):745–59. doi: 10.1038 / s41579-018-0089-x 17. Brown JL, Yates EA, Bielecki M, Olczak T, Smalley JW. Potential role for Streptococcus Gordonii-derived hydrogen peroxide in heme acquisition by Porphyromonas Gingivalis. Mol Oral Microbiol. (2018) 33(4):322–35. doi: 10.1111 / omi.12229 18. Yu XL, Chan Y, Zhuang L, Lai HC, Lang NP, Keung Leung W, et al. Intra-oral single-site comparisons of periodontal and peri-implant microbiota in health and disease. Clin Oral Implants Res. (2019) 30(8):760–76. doi: 10.1111 / clr.13459 19. Cheng WC, van Asten SD, Burns LA, Evans HG, Walter GJ, Hashim A, et al. Periodontitis-associated pathogens P. Gingivalis and A. Actinomycetemcomitans -47- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 activate human Cd14(+) monocytes leading to enhanced Th17 / il-17 responses. Eur J Immunol. (2016) 46(9):2211–21. doi: 10.1002 / eji.201545871 20. Mishra B, Reiling S, Zarena D, Wang G. Host defense antimicrobial peptides as antibiotics: design and application strategies. Curr Opin Chem Biol. (2017) 38:87–96. doi: 10.1016 / j.cbpa.2017.03.014 21. Huan Y, Kong Q, Mou H, Yi H. Antimicrobial peptides: classification, design, application and research progress in multiplefields. Front Microbiol. (2020) 11:582779. doi: 10.3389 / fmicb.2020.582779 22. Dini I, De Biasi MG, Mancusi A. An overview of the potentialities of antimicrobial peptides derived from natural sources. Antibiotics (Basel). (2022) 11(11):1483. doi: 10.3390 / antibiotics11111483 23. Wibowo D, Zhao CX. Recent achievements and perspectives for large-scale recombinant production of antimicrobial peptides. Appl Microbiol Biotechnol. (2019) 103(2):659–71. doi: 10.1007 / s00253-018-9524-1 24. Renaud S, Mansbach RA. Latent spaces for antimicrobial peptide design. Digital Discovery. (2023) 2(2):441–58. doi: 10.1039 / D2DD00091A 25. Azmat M, Ghalandari B, Jessica J, Xu Y, Li X, Su W, et al. Pepdred: De Novo peptide design with strong binding affinity for target protein. Anal Chem. (2023) 95 (33):12264–72. doi: 10.1021 / acs.analchem.3c01057 26. Martínez OF, Duque HM, Franco OL. Peptidomimetics as potential anti- virulence drugs against resistant bacterial pathogens. Front Microbiol. (2022) 13:831037. doi: 10.3389 / fmicb.2022.831037 27. Mahlapuu M, Håkansson J, Ringstad L, Björn C. Antimicrobial peptides: an emerging category of therapeutic agents. Front Cell Infect Microbiol. (2016) 6:194. doi: 10.3389 / fcimb.2016.00194 -48- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 28. Griffith A, Mateen A, Markowitz K, Singer SR, Cugini C, Shimizu E, et al. Alternative antibiotics in dentistry: antimicrobial peptides. Pharmaceutics. (2022) 14(8):1679. doi: 10.3390 / pharmaceutics14081679 29. Fischer NG, Münchow EA, Tamerler C, Bottino MC, Aparicio C. Harnessing biomolecules for bioinspired dental biomaterials. J Mater Chem B. (2020) 8 (38):8713–47. doi: 10.1039 / D0TB01456G 30. Holmberg KV, Abdolhosseini M, Li Y, Chen X, Gorr SU, Aparicio C. Bio- inspired stable antimicrobial peptide coatings for dental applications. Acta Biomater. (2013) 9(9):8224–31. doi: 10.1016 / j.actbio.2013.06.017 31. Xie S-X, Song L, Yuca E, Boone K, Sarikaya R, VanOosten SK, et al. Antimicrobial peptide–polymer conjugates for dentistry. ACS Appl Polym Mater. (2020) 2(3):1134–44. doi: 10.1021 / acsapm.9b00921 32. Spencer P, Ye Q, Misra A, Chandler JR, Cobb CM, Tamerler C. Engineering peptide-polymer hybrids for targeted repair and protection of cervical lesions. Front Dent Med. (2022) 3:1007753. doi: 10.3389 / fdmed.2022.1007753 33. Yuca E, Xie SX, Song L, Boone K, Kamathewatta N, Woolfolk SK, et al. Reconfigurable dual peptide tethered polymer system offers a synergistic solution for next generation dental adhesives. Int J Mol Sci. (2021) 22(12):6552. doi: 10. 3390 / ijms22126552 34. Wisdom C, Chen C, Yuca E, Zhou Y, Tamerler C, Snead ML. Repeatedly applied peptidefilm kills Bacteria on dental implants. Jom (1989). (2019) 71(4):1271–80. doi: 10.1007 / s11837-019-03334-w 35. Wisdom EC, Zhou Y, Chen C, Tamerler C, Snead ML. Mitigation of peri-implantitis by rational design of bifunctional peptides with antimicrobial properties. ACS Biomater Sci Eng. (2020) 6(5):2682–95. doi: 10.1021 / acsbiomaterials.9b01213 36. Zhang X, Geng H, Gong L, Zhang Q, Li H, Zhang X, et al. Modification of the surface of titanium with multifunctional chimeric peptides to prevent -49- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 biofilm formation via inhibition of initial colonizers. Int J Nanomedicine. (2018) 13:5361– 75. doi: 10.2147 / IJN.S170819 37. Godoy-Gallardo M, Mas-Moruno C, Yu K, Manero JM, Gil FJ, Kizhakkedathu JN, et al. Antibacterial properties of Hlf1-11 peptide onto Titanium surfaces: a comparison study between silanization and surface initiated polymerization. Biomacromolecules. (2015) 16(2):483–96. doi: 10.1021 / bm501528x 38. Godoy-Gallardo M, Mas-Moruno C, Fernández-Calderón MC, Pérez- Giraldo C, Manero JM, Albericio F, et al. Covalent immobilization of Hlf1-11 peptide on a Titanium surface reduces bacterial adhesion and biofilm formation. Acta Biomater. (2014) 10(8):3522–34. doi: 10.1016 / j.actbio.2014.03.026 39. Yazici H, Fong H, Wilson B, Oren EE, Amos FA, Zhang H, et al. Biological response on a Titanium implant-grade surface functionalized with modular peptides. Acta Biomater. (2013) 9(2):5341–52. doi: 10.1016 / j.actbio.2012.11.004 40. Zhou Y, Snead ML, Tamerler C. Bio-inspired hard-to-soft interface for implant integration to bone. Nanomedicine. (2015) 11(2):431–4. doi: 10.1016 / j.nano.2014.10. 003 41. Porto WF, Pires AS, Franco OL. Computational tools for exploring sequence databases as a resource for antimicrobial peptides. Biotechnol Adv. (2017) 35 (3):337–49. doi: 10.1016 / j.biotechadv.2017.02.001 42. Wang G, Li X, Wang Z. Apd3: the antimicrobial peptide database as a tool for research and education. Nucleic Acids Res. (2016) 44(D1):D1087–93. doi: 10.1093 / nar / gkv1278 43. Waghu FH, Idicula-Thomas S. Collection of antimicrobial peptides database and its derivatives: applications and beyond. Protein Sci. (2020) 29(1):36–42. doi: 10.1002 / pro.3714 -50- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 44. Ye G, Wu H, Huang J, Wang W, Ge K, Li G, et al. Lamp2: a major update of the database linking antimicrobial peptides. Database (Oxford). (2020) 2020:baaa061. doi: 10.1093 / database / baaa061 45. Fan L, Sun J, Zhou M, Zhou J, Lao X, Zheng H, et al. Dramp: a comprehensive data repository of antimicrobial peptides. Sci Rep. (2016) 6:24482. doi: 10.1038 / srep24482 46. Azam MW, Kumar A, Khan AU. Acd: antimicrobial chemotherapeutics database. PLoS One. (2020) 15(6):e0235193. doi: 10.1371 / journal.pone.0235193 47. Gupta R, Srivastava D, Sahu M, Tiwari S, Ambasta RK, Kumar P. Artificial intelligence to deep learning: machine intelligence approach for drug discovery. Mol Divers. (2021) 25(3):1315–60. doi: 10.1007 / s11030-021-10217-3 48. Liu Z, Jin J, Cui Y, Xiong Z, Nasiri A, Zhao Y, et al. Deepseqpanii: an interpretable recurrent neural network model with attention mechanism for peptide-hla class ii binding prediction. IEEE / ACM Trans Comput Biol Bioinform. (2022) 19(4):2188–96. doi: 10.1109 / TCBB.2021.3074927 49. Pertseva M, Gao B, Neumeier D, Yermanos A, Reddy ST. Applications of machine and deep learning in adaptive immunity. Annu Rev Chem Biomol Eng. (2021) 12:39–62. doi: 10.1146 / annurev-chembioeng-101420-125021 50. Wang C, Garlick S, Zloh M. Deep learning for novel antimicrobial peptide design. Biomolecules. (2021) 11(3):471. doi: 10.3390 / biom11030471 51. Khabbaz H, Karimi-Jafari MH, Saboury AA, BabaAli B. Prediction of antimicrobial peptides toxicity based on their physico-chemical properties using machine learning techniques. BMC Bioinform. (2021) 22(1):549. doi: 10.1186 / s12859- 021-04468-y 52. Boone K, Camarda K, Spencer P, Tamerler C. Antimicrobial peptide similarity and classification through rough set theory using physicochemical boundaries. BMC Bioinform. (2018) 19(1):469. doi: 10.1186 / s12859-018-2514-6 -51- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 53. Boone K, Wisdom C, Camarda K, Spencer P, Tamerler C. Combining genetic algorithm with machine learning strategies for designing potent antimicrobial peptides. BMC Bioinform. (2021) 22(1):239. doi: 10.1186 / s12859-021-04156-x 54. Edlund A, Yang Y, Hall AP, Guo L, Lux R, He X, et al. An in vitro biofilm model system maintaining a highly reproducible species and metabolic diversity approaching that of the human oral microbiome. Microbiome. (2013) 1(1):25. doi: 10.1186 / 2049-2618-1-25 55. Xiao X, Wang P, Lin WZ, Jia JH, Chou KC. Iamp-2l: a two-level multi-label classifier for identifying antimicrobial peptides and their functional types. Anal Biochem. (2013) 436(2):168–77. doi: 10.1016 / j.ab.2013.01.019 56. Kibinge N, Ikeda S, Ono N, Altaf-Ul-Amin M, Kanaya S. Integration of residue attributes for sequence diversity characterization of terpenoid enzymes. BioMed Res Int. (2014) 2014:753428. doi: 10.1155 / 2014 / 753428 57. Kawashima S, Pokarowski P, Pokarowska M, Kolinski A, Katayama T, Kanehisa M. Aaindex: amino acid index database, progress report 2008. Nucleic Acids Res. (2008) 36(Database issue):D202–5. doi: 10.1093 / nar / gkm998 58. Grzymala-Busse JW. Mining numerical data—a rough set approach. In: Kryszkiewicz M, Peters JF, Rybinski H, Skowron A, editors. Proceedings of the International Conference on Rough Sets and Intelligent Systems Paradigms. Berlin, Heidelberg: Springer-Verlag (2007). p. 12–21. doi: 10.1007 / 978-3-642-11479-3_1 59. Yao YY, Wong SKM, Lin TY. A review of rough set models. In: Lin TY, Cercone N, editors. Rough Sets and Data Mining: Analysis of Imprecise Data. Boston, MA: Springer US (1997). p. 47–75. doi: 10.1007 / 978-1-4613-1461-5_3 60. Pawlak Z. Rough sets: theoretical aspects of reasoning about data. Theory and Decision Library Series D, System Theory, Knowledge Engineering, and Problem Solving. Vol. 9. Netherlands: Springer Netherlands (1991). p. 1–19. doi: 10.1007 / 978-94- 011-3534-4 -52- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 61. Sztukowska MN, Roky M, Demuth DR. Peptide and non-peptide mimetics as potential therapeutics targeting oral Bacteria and oral biofilms. Mol Oral Microbiol. (2019) 34(5):169–82. doi: 10.1111 / omi.12267 62. Claesson R, Höglund-Åberg C, Haubek D, Johansson A. Age-related prevalence and characteristics of Aggregatibacter Actinomycetemcomitans in periodontitis patients living in Sweden. J Oral Microbiol. (2017) 9(1):1334504. doi: 10.1080 / 20002297.2017. 1334504 63. Chaudhury S, Lyskov S, Gray JJ. Pyrosetta: a script-based interface for implementing molecular modeling algorithms using Rosetta. Bioinform. (2010) 26 (5):689–91. doi: 10.1093 / bioinformatics / btq007 64. Meng EC, Pettersen EF, Couch GS, Huang CC, Ferrin TE. Tools for integrated sequence-structure analysis with ucsf chimera. BMC Bioinform. (2006) 7(1):339. doi: 10.1186 / 1471-2105-7-339 65. Pettersen EF, Goddard TD, Huang CC, Couch GS, Greenblatt DM, Meng EC, et al. UCSF chimera—a visualization system for exploratory research and analysis. J Comput Chem. (2004) 25(13):1605–12. doi: 10.1002 / jcc.20084 66. Lamiable A, Thevenet P, Rey J, Vavrusa M, Derreumaux P, Tuffery P. Pep-Fold3: faster De Novo structure prediction for linear peptides in solution and in complex. Nucleic Acids Res. (2016) 44(W1):W449–54. doi: 10.1093 / nar / gkw329 67. Chen CH, Starr CG, Troendle E, Wiedman G, Wimley WC, Ulmschneider JP, et al. Simulation-guided rational De Novo design of a small pore-forming antimicrobial peptide. J Am Chem Soc. (2019) 141(12):4839–48. doi: 10.1021 / jacs.8b11939 68. Kauffman WB, Fuselier T, He J, Wimley WC. Mechanism matters: a taxonomy of cell penetrating peptides. Trends Biochem Sci. (2015) 40(12):749–64. doi: 10.1016 / j. tibs.2015.10.004 -53- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 69. Lin B, Li R, Handley TNG, Wade JD, Li W, O’Brien-Simpson NM. Cationic antimicrobial peptides are leading the way to combat oropathogenic infections. ACS Infect Dis. (2021) 7(11):2959–70. doi: 10.1021 / acsinfecdis.1c00424 70. da Silva BRD, Freitas VAAD, Nascimento-Neto LG, Carneiro VA, Arruda FVS, Aguiar ASWD, et al. Antimicrobial peptide control of pathogenic microorganisms of the oral cavity: a review of the literature. Peptides. (2012) 36(2):315–21. doi: 10.1016 / j. peptides.2012.05.015 71. Niu JY, Yin IX, Mei ML, Mei ML, Wu WKK, Li Q-L, et al. The multifaceted roles of antimicrobial peptides in oral diseases. Mol Oral Microbiol. (2021) 36(3):159–71. doi: 10.1111 / omi.12333 72. Gorr SU. Antimicrobial peptides of the oral cavity. Periodontol 2000. (2009) 51:152–80. doi: 10.1111 / j.1600-0757.2009.00310.x Supplementary Material 1 Supplementary Data This document contains materials that supplement Machine learning enabled design features of antimicrobial peptides selectively targeting peri-implant disease progression. Within the following pages, additional figures and tables are provided to create a more comprehensive picture of the development and design of novel, targeted antimicrobial peptides (AMPs) via the unique machine learning models (CLN-MLEM2, CB-GA). Multiple peptides against P. gingivalis, A. actinomycetemcomitans, and S. gordonii. This data was used in the training and tailoring of the CLN-MLEM2 method, as discussed in the main paper. Figures 16–20 contains a series of graphical depictions of the targeted AMP VL-13 across six hydropathy properties of interest. The figures delineate how VL-13’s summary characteristics in these properties correspond to its predicted growth inhibition for A. actinomycetemcomitans. An additional summary of the rough set theory follows these figures. -54- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Supplementary table, i.e., Supplementary Table 1, presents a compilation of literature-based AMPs known for their activity against several periodontal pathogens. This table was also used in training the CLN-MLEM2 model of the current study. Supplementary Table 2 provides short descriptions of the hydropathy properties identified by the CLN- MLEM2 method as defining physiochemical traits for targeted antimicrobial ability. The performance of the CLN-MLEM2 model was also predicted on additional peptides reported as potential therapeutics for oral bacteria and oral biofilms as the test set (Supplementary Tables 3-5). The model did not train on this test set. The rules sets showed good predictability for the selected keystone pathogen and the commensal / accessory pathogen. The false discovery rate was found to be low for the keystone pathogen A. actinomycetemcomitans and the commensal pathogen S. gordonii (Supplementary Table 3 and Supplementary Table 5). 1.1 Peptide Sequence Signals In Figure 15 and Figure 16, the attributes of many peptide sequences, counted on the horizontal axis, are displayed together. One stacked bar of peptide signals is one peptide sequence summary. Each signal is a normalization of the value of the property over the range of the peptides in the database. The equation for the signal is: ^^ ൌ 0.5 The factor of 50 increases the sensitivity of the signal values for small values in the range. A factor of 1 has maximum sensitivity at the 50th-percentile. Adjusting the factor to 50 focuses the maximum sensitivity at the first percentile of the range. This characteristic makes this chart tailored to show the difference between differences in properties at low percentiles and less sensitive to show differences at larger percentiles. This is a desired property for summarizing which peptides have any magnitude for each of the properties, such as the net rule counts. Using either a linear scale or a logarithmic scale with a factor of one, would make small magnitudes difficult to distinguish from no magnitudes. In addition to the magnitude information this provides, the property value sign information is -55- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 included by moving the stack above the x-axis when the property is positive and to below the x-axis when the property is negative. Rough Set Theory Rule Summary for VL-13 Both Rule 1 and Rule 2 have descriptions of amphipathy which are quantitative for targeting Aa. Using machine learning (ML) is important in this domain because a critical review of the features by CLN-MLEM2 has made distinctions which have hydropathy meaning in different formal senses, without losing technical rigor. ML can help make progress when the technical descriptions of properties change between these scales. Instead of being limited to selecting one scale to use, we use multiple scales with the traceability of knowing which scale is referenced at any given time. As an overview, targeting Aa seems to require amphipathic features, such as being generally hydrophilic, possessing one key hydrophobic 3-aa window, and also having a second 3-aa window which is moderately polar. For other properties, the rule is possessing a hydrophobic 3-aa window and being moderately non-polar. The lack of contradiction comes from which properties are chosen. The rule for a peptide not inhibiting Aa growth seems to be the possession of a 3-aa window of highly polar amino acids and a window which is hydrophobic by three different hydropathy descriptions. The machine learning model arranged the details of the known data, allowing for more precise predictions than discussed in the previous paragraph. This precision is not a safe guard against the discovered rules being inaccurate for future cases, but rather a guard against missing information due to vague, overlapping terminology. Because we have precise, conjunctive descriptions of when activity is predicted to occur, we benefit from automation for selection of sequences which meet these specific descriptions for activity. Our genetic algorithm makes finding sequences which meet multiple rules and conditions reliable. 2 Supplementary Figures and Tables 2.1 Supplementary Tables -56- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Supplementary Table 1. Antimicrobial peptides with known activity against periodontal pathogens. ‘NR’ indicates not reported (1). Pathogens are abbreviated as follows: A. actinomycetemcomitans (Aa), S. gordonii (Sg), P. gingivalis (Pg). MIC values higher than 100 μg / mL are not considered active for peptides. The MIC values for the proteins (nigrescinB and nigrescinC) are considered active below 100 mM. The activities of low and none are not distinguished in the targeting rule sets used for this study. Rules are distinguished between high activity and either low-activity or no-activity for each of the three bacterial species. -57- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -58- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -59- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Supplementary Table 2. Amino acid hydropathy properties selected in the CLN-MLEM2 rules for describing targeted antibacterial activity. These properties depict non-linear boundaries between activity types. LIFS790102 (Conformational preference for parallel beta-strands) and PONP800108 (Average number of surrounding residues) were in referenced study, but not selected for rules. -60- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Supplementary Table 3. Testing performance of the model targeting against A. actinomycetemcomitans (Aa). The rough set theory (RST) rules generated by CLN- MLEM2 model was tested on peptides reported as targeting oral bacteria and oral biofilms in a review by Sztukowska et al. (17). The match column describes if the sign of RST support, either positive or negative, matched the activity of the sequence. A “1” in “Match” indicates a match. A “0” indicates a mismatch. “0.5” was given for the case when the activity is dependent on an additional condition (Sequence #9: addition of an acetyl group at the C- terminal). RST Support is the number of training peptides which are identified by the rule set. False positive RST support is “FP”. If the rule is for activity, then the RST support is positive. If the rule is for inactivity, the RST support is negative. The weight is the magnitude of RST support. The area under the curve (AUC) weight is the product of the weight and the match values. Accuracy is the proportion of matches, and AUC is the ratio of the sum of the AUC weight column over the sum of the weight column. The AUC indicates that 76% of the RST support was classified correctly. Accuracy indicates 46% of test cases were classified correctly. FDR indicates 0% of weight relates to false positive predictions. -61- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Accuracy 46% FDR 0.0% AUC 0.74 -62- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Supplementary Table 4. Testing performance of the model targeting against P. gingivalis (Pg). The rules generated by CLN-MLEM2 model was tested on peptides reported as targeting oral bacteria and oral biofilms in a review by Sztukowska et al. (17). The match column describes if the sign of RST support, either positive or negative, matched the activity of the sequence. A “1” in “Match” indicates a match. A “0” indicates a mismatch. “0.5” was given for the case when the activity is dependent on an additional condition (Sequence #9: addition of an acetyl group at the C-terminal). RST Support is the number of training peptides which are identified by the rule set. False positive RST support is “FP”. If the rule is for activity, then the RST support is positive. If the rule is for inactivity, the RST support is negative. The weight is the magnitude of RST support. The area under the curve (AUC) weight is the product of the weight and the match values. Accuracy is the proportion of matches, and AUC is the ratio of the sum of the AUC weight column over the sum of the weight column. The AUC indicates that 66% of the RST support was classified correctly. Accuracy indicates 68% of test cases were classified correctly. FDR indicates that 27% of weight relates to false positive predictions. -63- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Accuracy 68% FD 27% AUC 0.66 R Supplementary Table 5. Testing performance of the model targeting against S. gordonii (Sg). The rules generated by CLN-MLEM2 model was tested on peptides reported as targeting oral bacteria and oral biofilms in a review by Sztukowska et al. (17). The match column describes if the sign of RST support, either positive or negative, matched the activity of the sequence. A “1” in “Match” indicates a match. A “0” indicates a mismatch. “0.5” was given for the case when the activity is dependent on an additional condition (Sequence #9: addition of an acetyl group at the C-terminal). RST Support is the number of training peptides which are identified by the rule set. False positive RST support is “FP”. If the rule is for activity, then the RST support is positive. If the rule is for inactivity, the RST support is negative. The weight is the magnitude of RST support. The area under the curve (AUC) weight is the product of the weight and the match values. Accuracy is the proportion of matches, and AUC is the ratio of the sum of the AUC weight column over the sum of the weight column. The AUC of 0.49 indicates that 49% of the RST -64- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 support was classified correctly. Accuracy of 39% indicates 39% of test cases were classified correctly. FDR indicates that 1.1% of weight relates to false positive predictions. -65- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Accuracy 39% FD 1.1% AUC 0.49 R Supplementary References 1. Grzymala-Busse JW. Mlem2—Discretization During Rule Induction. In: Kłopotek MA, Wierzchoń, S.T., Trojanowski, K., editor. Intelligent Information Processing and Web Mining. Berlin, Heidelberg: Springer (2003). p. 499-508. doi: 10.1007 / 978-3-540-36562-4_53. 2. Suwandecha T, Srichana T, Balekar N, Nakpheng T, Pangsomboon K. Novel Antimicrobial Peptide Specifically Active against Porphyromonas Gingivalis. Archives of microbiology (2015) 197(7):899-909. doi: https: / / doi.org / 10.1007 / s00203-015- 1126-z. 3. Joly S, Maze C, McCray PB, Jr., Guthmiller JM. Human Beta- Defensins 2 and 3 Demonstrate Strain-Selective Activity against Oral Microorganisms. Journal of clinical microbiology (2004) 42(3):1024-9. doi: https: / / doi.org / 10.1128 / jcm.42.3.1024-1029.2004. 4. Guthmiller JM, Vargas KG, Srikantha R, Schomberg LL, Weistroffer PL, McCray Jr PB, et al. Susceptibilities of Oral Bacteria and Yeast to Mammalian Cathelicidins. Antimicrobial agents and chemotherapy (2001) 45(11):3216-9. doi: https: / / doi.org / 10.1128 / aac.45.11.3216-3219.2001. 5. Teanpaisan R, Narawatthana S, Utarabhand P. The Gene Coding for Nigrescin Produced by Prevotella Nigrescens Atcc 25261. Letters in Applied Microbiology (2009) 49(3):293-8. doi: https: / / doi.org / 10.1111 / j.1472-765X.2009.02657.x. 6. Dale BA, Fredericks LP. Antimicrobial Peptides in the Oral Environment: Expression and Function in Health and Disease. Current issues in molecular biology (2005) 7(2):119-34. https: / / doi.org / 10.21775 / cimb.007.119 -66- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 7. Hirt H, Hall JW, Larson E, Gorr S-U. A D-Enantiomer of the Antimicrobial Peptide Gl13k Evades Antimicrobial Resistance in the Gram Positive Bacteria Enterococcus Faecalis and Streptococcus Gordonii. PLOS ONE (2018) 13(3):e0194900. doi: 10.1371 / journal.pone.0194900. 8. Concannon SP, Crowe TD, Abercrombie JJ, Molina CM, Hou P, Sukumaran DK, et al. Susceptibility of Oral Bacteria to an Antimicrobial Decapeptide. Journal of Medical Microbiology (2003) 52(12):1083-93. doi: https: / / doi.org / 10.1099 / jmm.0.05286-0. 9. Wang W, Tao R, Tong Z, Ding Y, Kuang R, Zhai S, et al. Effect of a Novel Antimicrobial Peptide Chrysophsin-1 on Oral Pathogens and Streptococcus Mutans Biofilms. Peptides (2012) 33(2):212-9. doi: https: / / doi.org / 10.1016 / j.peptides.2012.01.006. 10. Chen L, Jia L, Zhang Q, Zhou X, Liu Z, Li B, et al. A Novel Antimicrobial Peptide against Dental-Caries-Associated Bacteria. Anaerobe (2017) 47:165- 72. doi: https: / / doi.org / 10.1016 / j.anaerobe.2017.05.016. 11. Cowan R, Whittaker RG. Hydrophobicity Indices for Amino Acid Residues as Determined by High-Performance Liquid Chromatography. Pept Res (1990) 3(2):75-80. PMID: 2134053 12. Fauchere JL, Charton M, Kier LB, Verloop A, Pliska V. Amino Acid Side Chain Parameters for Correlation Studies in Biology and Pharmacology. Int J Pept Protein Res (1988) 32(4):269-78. doi: https: / / doi.org / 10.1111 / j.1399-3011.1988.tb01261.x 13. Jacobs RE, White SH. The Nature of the Hydrophobic Binding of Small Peptides at the Bilayer Interface: Implications for the Insertion of Transbilayer Helices. Biochemistry (1989) 28(8):3421-37. doi: 10.1021 / bi00434a042. 14. Meek JL. Prediction of Peptide Retention Times in High-Pressure Liquid Chromatography on the Basis of Amino Acid Composition. Proc Natl Acad Sci U S A (1980) 77(3):1632-6. doi: 10.1073 / pnas.77.3.1632. -67- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 15. Warme PK, Morgan RS. A Survey of Amino Acid Side-Chain Interactions in 21 Proteins. J Mol Biol (1978) 118(3):289-304. doi: 10.1016 / 0022- 2836(78)90229-2. 16. Zimmerman JM, Eliezer N, Simha R. The Characterization of Amino Acid Sequences in Proteins by Statistical Methods. J Theor Biol (1968) 21(2):170-201. doi: https: / / doi.org / 10.1016 / 0022-5193(68)90069-6. 17. Sztukowska MN, Roky M, Demuth DR. Peptide and Non-Peptide Mimetics as Potential Therapeutics Targeting Oral Bacteria and Oral Biofilms. Mol Oral Microbiol (2019) 34(5):169-82. Epub 20190815. doi: 10.1111 / omi.12267. B. Interpretable and Transparent Machine Learning Guided Antimicrobial Peptide Design Empowers Targeted Inhibition of a Keystone Pathogen As a proxy for antimicrobial activity, rough set theory rules were used for activity. The development is to change the rough set theory rules from generic antibacterial activity to narrow species activity for identified oral bacterial species. Negative data is the peptides generated for the generic antibacterial AMPs which did not incorporate all of the specific inhibitory rules as the AMPs, which led to the identification of VL13. Collected data: planktonic and biofilm specificity of single species cultures; uncollected data of effect of our designed targeted AMPs in microbiota community; however, bacteriocins are natural, highly selective AMPs produced within the microbiota community by member bacteria which help to create observed compositions of bacteria in the community. The structural and physical properties of the training sets are explicitly the source of the rough set rules selected to evaluate the outputted predicted peptides. The rough set rules are the conjunction of multiple structural / physical property ranges which describes the set of peptides which are within all those ranges that make up the rule. More specifically of how the rules impact the output peptides, the number of training peptides which meet the rough set rule and the condition (e.g., inhibition of the specific strain) is added to the rule count for the peptide. The rule count for the peptides has a high target, which means that the more counts a peptide has for inhibition of that strain, the lower the addition of that weighed -68- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 component to the fitness score. The approach had evolved from scoring of a sequence when it was predicted to be antibacterial, ad expanded with using the counts of the supporting training sequences. Lower fitness scores are better. Training data used for machine learning model are not routine and non- conventional. Generic approach related negative data set was trained on the iAMP-2L data set. In the field of antimicrobial peptide training, this is a routine data set. The training dataset for the outcome were selected antimicrobial peptides with annotated inhibition of oral bacteria. These literature results were supplemented with AMP activity tests against oral bacteria within our paper. This is non-conventional because (1) no training set had previously targeted only oral bacteria and (2) the training set size was orders of magnitude smaller compared to conventional data sets (~20 peptides for our training set while conventional data sets are ~1,000 – 8,000 peptides). Inputs: Data table of peptide sequences labeled with inhibition activity with one row per peptide sequence and one column per labeled inhibition value. If structural features are considered, data table is extended by an inner join of structural features by peptide sequence, carrying over inhibition activity columns using the primary keys of (peptide sequence, structure id) in the joined data table. While no structural features were explicitly included for the negative data set or positive data set, these may be implemented. Inputs may include calculation functions for peptide properties (either structural or physicochemical properties) add to the data table with a column per row (also joined by peptide sequence), such as sum of property for sequence, mean of property for sequence, and maximum property value for a window of 3 amino acids, among others. For the negative data set, the 74 non-correlated properties were selected from the AAindex1 property set. For the positive data set, the properties were focused to 8 hydropathy properties in the AAindex1 property set. Outputs: Rough set theory rules with conjunctive conditions over data table columns, conditioned on inhibition outcomes. Genetic Algorithm starts with an initial pool (APD sequences for negative data set and iAMP-2L antibacterial peptides for positive data set. Genetic algorithm selects the best fitness scores (based on rule score counts, peptide -69- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 length and peptide hydrophobicity) for applying the genetic operators to get the next generation. Sequence collection is then scored and ranked. The sequences with the worst fitness scores are not represented in the next generation of sequences. Initial set of peptides predicted by the model: Validation that the predicted output peptides have the desired functionality on the microbiota community. Total Score is a priority metric in which lower scores receive more priority. The range of scores in a generation tends stayed consistent for our positive data set, while the generation of peptides was continued until the range of scores collapsed in our negative data set. Previous versions were limited by only coming up with repeated convergent peptide sequences when targeting shorter sequences (7aa – 12 amino acids) within 12 runs (2 sets of duplicates). Sequences with high homology to the duplicate sequences were not antibacterial when tested. Targeting longer sequences (7 aa to 15 aa) has reduced the number of duplicate sequences. Overview: ML Design Prediction Improvement This describes the interpretable and transparent machine learning approach enabling the species-specific targeting inhibition of keystone pathogen. Ideally, a successful host immune response should balance sufficient inflammation to reduce / eradicate pathogens, while maintaining beneficial bacteria and preventing damage to the commensal organisms, which potentially protects host health by coordinating cooperative roles for colonization resistance. The approach utilizes decision trees and enable the extraction of rules that relates to the inhibition activities of the pathogen, that are related to keystone pathogens while sparing the commensal and accessary pathogens. The approach focuses on identifying and expanding on design features that play a critical role in guiding the predictions in an iterative manner supported by the wetlab experimental testing. Machine learning approaches allow to solve multiple complex relationship and address the vulnerability of deep learning methods in the training processes. Transparency in decision process eliminates the vulnerability due to making illogical connections where the correlated relations lack causation. While these methods may still allow to model several -70- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 different kinds of complex relationships, the lack of transparency about the classification becomes especially difficult in determining boundaries of similar antimicrobial peptide clusters. The rise of antibiotic-resistant bacteria necessitates immediate and effective interventions. To address this critical issue, antimicrobial peptides have been receiving increasing interest as an alternative antimicrobial agent. While antimicrobial peptides can target several different range of pathogens, they exhibit differential antimicrobial activity ranging from broad spectrum to specific influenced by several different factors including environment. Furthermore, they have a wide range of antimicrobial mechanisms reflected within the structural and functional diversity. With the increased number of antimicrobial peptides that are discovered from natural sources, several new approaches have been introduced to their search to find more candidates. While their laboratory bench discovery, introducing antimicrobial peptide-mimetics to the existing peptide libraries continues, integration of computational methods has increased the number of antimicrobial peptide candidates tremendously. Computer aided molecular design using quantitative sequence activity relationship, encrypted antimicrobial peptides for DNA repositories, grammar- and regular expression-based match sequence patterns have been used to predict sequence activity relationship and identify similarities. The transparent ML antimicrobial peptide design approach starts with increasing comprehension of relationship between specific design solutions in a design space while broadening the understanding of the structure of the design space beyond a single cluster of design iterations. The earlier design approach shown in Fig. 22 use rough set theory to find highly selective boundaries, in the form of peptide sequence property rules, for differentiating antibacterial peptides from other peptides. Then new peptide sequences were generated from the diversity of known AMPs and filter them with these rules. The method goes beyond the known AMP sequences to continue to evolve them in silico and identify new candidates. A custom algorithm was developed using rough set theory with prioritization of low false discovery rates and maintained interpretability and generalizability through the -71- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 condition limited number – modified learning form experience module 2 (CLN-MLEM2) in Python. Rough set theory has been used in data mining as a heuristic method for discovering rules that distinguished between outcomes. Order sensitivity was addressed by calculating the physicochemical properties of sub-sequences in addition to using descriptors of physicochemical properties for length independence. With this approach, order-sensitivity and length independence was combined and select combinations of both kind of descriptors into a single rule defining its own cluster. Among the two main approaches in designing antimicrobial peptides, one uses curated insights into antimicrobial activity (rational design), and the other (opaque machine learning) leverages sequence data trends such as deep neural networks to identify patterns and trends using large datasets. Rational design involves utilizing existing knowledge of the structure-activity relationships and inserts key patterns to generate new candidates (like the Joker algorithm). Opaque machine learning identifies patterns and predict sequences, but, often lack interpretability, and makes it harder to incorporate the outcome from these models for further refinement of antimicrobial peptides. A custom genetic algorithm in R was further developed which uses a codon-based representation to escape local minima for finding more globally optimal peptide sequences fitting the targets (Table 1). While codon representation allows to uncover the diversity for customizing the design and direct the selection of sequences related in DNA-space, transparent machine learning guides the sequence selection process for targeted properties. Table 1. Antimicrobial peptide targets using generic antibacterial CLN- MLEM2 rules. -72- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 How the addition of the codon-representation aided in generating a wider representation of aggregation potential among sequence candidates than without the codon representation was evaluated. (Fig 23). The number of distinct sequences was less with the codon-based representation for the wider property representation (Fig 24a). This increased design flexibility did not stop the codon-based method from finding sequences that converged the target properties as well as the non-codon genetic algorithm converged. (Fig 24c). The convergence of the codon and non-codon variations were asymptotically similar. (Figs. 24b and d). Some of the best scoring sequences were selected from the codon-based genetic algorithm to evaluate their antibacterial activity against a commensal pathogen, e.g., S. epidermidis. While it is a common bacterium found in the oral cavity, it has an important biofilm forming capacity and has been strongly associated with peri-implantitis, an infection of the oral tissues surrounding an implant. Peri-implantitis has been significantly increasing and raising a concern for increased failures of the implant. In the search, the closest sequence similarity neighbor for the designed sequences in the Antimicrobial Peptide Database (APD) was a natural AMP from scorpion venom. The initial results were promising to be able to describe the boundary between antibacterial peptides and non-antibacterial peptides well enough to find a neighbor of a natural AMP in peptide design space (Table 2). Table 2. Antibacterial activity of designed AMPs with a natural AMP control and an antibiotic control of Ampicillin by zone of inhibition. -73- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Evaluating other groupings of the designed antimicrobial candidates produced through the codon-based genetic algorithm was continued. In Fig 25, the previous concept of finding the best sequence similarity scoring peptide in the APD (Myxinidin: DYHHGRVL) was repeated. A secondary peptide control Serracin-P which had the highest CLN-MLEM2 rule score in the APD was also added. While discovered that Serracin-P is active against S. epidermidis at relevant concentrations, it was discovered that the designed peptides didn’t have activity until extremely large concentrations were attempted. Further, the APD-selected peptide didn’t have any relevant activity against S. epidermidis. Concluding that while promising initial results were had; to make more progress, the method from “antibacterial” to targeted for a specific bacterial taxon was focused on. Literature sequences were selected with annotated activity against oral bacteria strains, and relevant peptides were tested this time with the oral bacteria strains. Then CLN-MLEM2 rough set theory rules were developed specific to each strain. The method generated sequences from evolving known AMPs to find the active peptides, when applying rules specific for Aa. The method further evolved as two peptides KK-15 and FV- 11 were predicted for activity but showed no activity for the oral species that have been tested. The development trajectory is not described as a method which always predicts the correct AMPs. The trajectory instead is developed as a method with the transparency and interpretability. This approach allows us both to succeed in finding some working sequences, and to learn from the prediction misses by considering the most relevant connections in the data tables aided by rough set theory, which helps us to guide the ML training process. With relatively few sequences, there is informed hypotheses about when hydropathy sequence features lead to activity and when they do not for specific bacteria. The negative results become more valuable when there is a more interpretable and transparent representation of the problem space. -74- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Machine Learning Method An antimicrobial peptides (AMP) classification approach was developed for antibacterial peptides using non-correlated AAindex1 properties. The classification approach is based on rough set theory, which is used in genome analysis in bioinformatics to determine the most relevant expression differences between disease and non-disease states.1-3The use of rough set theory for the classification of peptide function is pioneered. The customized algorithm, condition-limited number modified from experience learning module 2 (CLN- MLEM2) constrains rules to have a specified maximum number of constraints to adjust the simplicity of the generated rules based on the application. Appendix A further explains a brief history of the CLN-MLEM2 algorithm procedure noting where some of the distinct features originated. The training set had 1,274 antibacterial sequences and 1,440 non-antibacterial sequences. One rule for antibacterial activity which applies to 446 of the 1,274 antibacterial sequences (35%) and 10 out of 1,440 non-antibacterial sequences (0.7%) is in Table 3. This rule collects peptides with a strong features: transmembrane motif, helix termination motif, thermophilic motif, secondary structure C feature, and coil feature. In addition, the sequence can have at most 3 negative charges and two conjunctive constraints: overall sum of solution free energy between -0.61 kcal / mol and 19.51 kcal / mol and the total sum of the mean frequency of the amino acid being in an alpha helix to be between 12.68 and 39.9. The exclusively positive frequency sum constraints means that all peptides meeting this rule must at least be 10 amino acids long and must be at most 48 amino acids long. Since the free energy of solution is either positive or negative, no similar length constraint exists. Table 3. CLN-MLEM2 rule for bacterial growth inhibition. The accuracy of this rule is 97.8% (446 / 456) for the peptides that met the conditions from the iAMP-2L dataset. Percentages are the interpolated positions on a window range from 0% to 100% for possible values. -75- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 As shown that the rough set theory methods have a low false positive rate, a genetic algorithm was next built to design new antimicrobial peptides with the classification. Appendix B shows the unique and standard parts of the genetic algorithm protocol. were able to target the sequences which have generic antibacterial rules. About 97% (20,980 of 21,672) of the unique sequences generated met one of the rules for antibacterial inhibition -76- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 identified as a boundary between antibacterial and non-antibacterial peptides. verified that a sequence designed showed inhibitory activity against S. epidermidis. The design effort was continued for increasingly targeting AMP inhibition beyond generic antibacterial activity by focusing on fewer features of the peptide sequences. Feature generation was narrowed down to just 8 hydropathy properties of the 544 AAindex1 property set, down from the 74 non-correlated features used for generic antibacterial activity. In parallel, the training sequence scope was further narrowed down to find rules specific to bacterial strains with recognized roles in the oral microbiome. start to gather the known examples of AMPs inhibiting oral bacteria. With the initial set of known antimicrobial peptides which have been limited in number, were able to generate rules targeting inhibition of Aggregatibacter actinomycetemcomitans, a keystone pathogen (Table 4). Table 4 Top 25 sequences meeting A. actinomycetemcomitans inhibition rules at the 25th Generation from 2,261 sequences -77- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -78- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Table 5 provides the rules selected guiding the identification of the antimicrobial peptides with the species biased inhibition property. A key difference between peptides meeting inhibitory Rule 1 and peptides meeting non-inhibitory Rule 3 is the hydrophobicity at pH 3. For Rule 1, three conditions indicate that these values must be generally negative for inhibitory peptides, but Rule 3 has a condition in which this hydrophobicity feature must be high (87%-99%). This difference may indicate it is necessary in the inhibition process for peptides to remain highly soluble under acidic conditions, while less soluble peptides under acidic conditions are less bioactive for this strain. The inhibitory Rule #2 indicates that peptides with a high number of full non-bonding orbitals may be less bioactive against this strain. Rule #2 has a limit of 46% of the maximum number of such orbitals in a 3-amino acid subsequence. Rule #3 has a requirement of one 3-amino acid subsequence with at least 87% of the maximum number of orbitals. This distinction may indicate some mechanistic role of full non-bonding orbitals making peptides more susceptible to the AMP defenses for this strain. Table 5. A set of CLN-MLEM2 rules for A. actinomycetemcomitans. Percentages for 3-amino acid windows are interpolated for the possible sums on a range from 0% to 100%. -79- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -80- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Table 4 lists the top 25 sequences which were selected at the 25 Generation of the genetic algorithm. meeting some A. actinomycetemcomitans rules. To understand the necessity of using the A. actinomycetemcomitans inhibition rules, sequences meeting CLN-MLEM2 rules for A were also generated. actinomycetemcomitans non-inhibition. Appendix B has the top 25 sequences for sequences meeting at least one A. actinomycetemcomitans non-inhibition rule. The motifs in the sequences with non-inhibition rules show low homology to sequences meeting the A. actinomycetemcomitans inhibition rules. Table 6 Top 25 sequences meeting A. actinomycetemcomitans non-inhibition rules at the 25th Generation from 989 sequences -81- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 The necessity of the A. actinomycetemcomitans rules in the fitness function for the genetic algorithm were further explored. Sequences generated for S. epidermidis with general antimicrobial peptide rules (nonspecific) were next checked if they meet these specific A. actinomycetemcomitans related rules. Of 21,672 sequences generated without the specific rules, only 74 (0.342%) met the criteria for one of the rules for Aa inhibition, specifically Rule 1 (solubility in acidic conditions, critical for oral environment) in the recent oral antimicrobial targeting study.5This membership percentage is low compared to the membership for the rules used with in the scoring for the genetic algorithm of 96.8% (20,980 of 21,672). This membership percentage is also low compared to the membership percentage for the genetic algorithm sequences in the generation pool using the specific rules in the recent paper of 14.6% (329 of 2,261). None of the sequences generated in the previous study met the conditions for Rule #2 (low full non-bonding orbitals). Sequences from the generic antimicrobial activity rules were provided that met one of two specific rules in Table 3. None -82- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 of the peptides selected for testing in the earlier study of inhibition against S. epidermidis met the rule for specificity to A. actinomycetemcomitans. Table 7 Sequences meeting A. actinomycetemcomitans inhibition rule #1 which did not use this rule in the scoring of the genetic algorithm. 74 out of 21,672 sequences generated for finding S. epidermidis AMP met this rule. No other A. actinomycetemcomitans inhibition rule was met by these sequences. -83- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -84- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 The method shifted generation sequence characteristics based on membership within boundaries selected by the CLN-MLEM2 algorithm with increased targeting compared to the generic antibacterial peptide rules previously studied, and expanded with the training sets that are evolved with antimicrobial properties that are known for oral community. Combination of design features and expanding on training sets allows the identification of antimicrobial peptides with the specific inhibitory activity against a keystone pathogen strain while having reduced inhibition for the growth of commensal and support polymicrobial community structure, maintaining microbial homeostasis. -85- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 ML Method Expansion Continues with additional design features and training sets ML guided species targeting AMP design (Flowcharts (1-3) evolving with different training sets and design features- These are ongoing studies to reflect the interpretable and transparent strength of the ML model. References Cited (1) Peters, G.; Crespo, F.; Lingras, P.; Weber, R. Soft clustering – Fuzzy and rough approaches and their extensions and derivatives. International Journal of Approximate Reasoning 2013, 54 (2), 307-322. DOI: https: / / doi.org / 10.1016 / j.ijar.2012.10.003. (2) Chen, Y.; Zhang, Z.; Zheng, J.; Ma, Y.; Xue, Y. Gene selection for tumor classification using neighborhood rough sets and entropy measures. Journal of Biomedical Informatics 2017, 67, 59-68. DOI: https: / / doi.org / 10.1016 / j.jbi.2017.02.007. (3) Sun, L.; Zhang, X.; Qian, Y.; Xu, J.; Zhang, S. Feature selection using neighborhood entropy-based uncertainty measures for gene expression data classification. Information Sciences 2019, 502, 18-41. DOI: https: / / doi.org / 10.1016 / j.ins.2019.05.072. (4) Hajishengallis, G.; Darveau, R. P.; Curtis, M. A. The keystone-pathogen hypothesis. Nat Rev Microbiol 2012, 10 (10), 717-725. DOI: 10.1038 / nrmicro2873 From NLM. (5) Boone, K.; Tjokro, N.; Chu, K. N.; Chen, C.; Snead, M. L.; Tamerler, C. Machine learning-enabled design features of antimicrobial peptides selectively targeting peri- implant disease progression. Frontiers in Dental Medicine 2024, 5, Original Research. DOI: 10.3389 / fdmed.2024.1372534. (6) Boone, K.; Camarda, K.; Spencer, P.; Tamerler, C. Antimicrobial peptide similarity and classification through rough set theory using physicochemical boundaries. BMC Bioinformatics 2018, 19 (1), 469. DOI: 10.1186 / s12859-018-2514-6. (7) Boone, K.; Wisdom, C.; Camarda, K.; Spencer, P.; Tamerler, C. Combining genetic algorithm with machine learning strategies for designing potent -86- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 antimicrobial peptides. BMC Bioinformatics 2021, 22 (1), 239. DOI: 10.1186 / s12859-021- 04156-x. (8) Grzymala-Busse, J. W. Rule induction. In Machine Learning for Data Science Handbook: Data Mining and Knowledge Discovery Handbook, Springer, 2023; pp 55-74. (9) Grzymala-Busse, J. W.; Grzymala-Busse, W. J. Handling missing attribute values. Data mining and knowledge discovery handbook 2010, 33-51. Appendix A The condition limited number – modified learning from experience module 2 (CLN-MLEM2) algorithm is a member of the rough set theory algorithms using discretization schemes to handle continuous numerical data with a heuristic method invented to maintain the interpretability and generalizability of the rules by limiting their conditions. The algorithm is descended from Jerzy W. Grzymala-Busse’s MLEM2 Algorithm.8LEM1, part of the Learning from Experience using Rough Sets (LERS) system, uses the indiscernibility property of a subset of attribute columns in a data table as an equivalence relation to group together similar rows conditioned on an outcome of interest, called a decision D, which is recorded in the data table. U is the “universe” set of all rows in the data table. The partition of groups r1, r2, … for a decision form global covering R if: Statement 1 ^∀ ^^, ^^ ∈ ^^,^^^^^^ ൌ ^^^^^^^^^^ ^^^^^^ ^^^^^^^^ ^^^^ ^^^^^^ ൌ ^^^^^^^where X(e) = group x which contains element e This covering R is heuristically minimal in the sense that if any attribute can be removed and the discernability relationship holds, it is removed from the set of column attributes. LEM1 Algorithm8-87- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Algorithm to compute a single global covering (input: the set A of all attributes, partition [d]* on U; output: a single global covering R); begin global covering compute partition A*; P := A; R := ∅; if A* ⊆ [d]* then begin attribute removal attempts for each attribute a in A do begin Q := P.remove[a]; compute partition Q*; if Q* ⊆ [d]* then P := Q end for R := P end then end algorithm The LEM2 algorithm shifts the grouping strategy from picking attributes to form partitions where each attribute value describes a single set within a partition, to grouping attribute-value pairs to form single local coverings to approximate a global covering. The algorithm heuristically searches for the smallest set of single local coverings whose union is U. To accomplish this, attribute-value pairs are selected based on how many rows with the specified outcome they cover, and then by how selective they are. The conjunction of these pairings can rapidly increase the selectivity of a local covering. A problem may arise when local coverings are too specific and do not generalize when more data is added to the data table. To mitigate this, a maximum condition limit was developed for rules to reduce their specificity to maintain generalization. -88- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 The modified LEM2 algorithm contains two enhancements.9The first is support for handing different types of non-discretized data: numerical values and missing values. The second is the addition of the concept of the addition of probable rules beyond certain rules which the quantification of the training probability for the rule sets generated. CLN-MLEM2 Algorithm G: goal set, the rows which have specified outcome B: set of all rows being described T: target subgoal which is addressed for a single rule iteration T(G): rows in G with no applicable rules t: conjunctive attribute-value pairs, each rule is one set of these pairs α: minimum acceptable training probability (input: a set B, output: a single local covering L of set B); begin local covering G := B; L := ∅; while G ് ∅;begin single rule T := ∅; T(G) := [t | t ∩ G ≠∅^ ; while (T = ∅ or [T] ⊈ B) and |pairs| < condition-limit number ; begin adding a condition to the rule select a pair t ∈ T(G) such that |[t] ∩ G| is maximum; if a tie occurs, select min |t| for tied sets if another tie occurs, select by min column number; T := T ∪ [t] ; G:= |t] ∩ G; T(G) = T(G).remove(T) ; end while P(T) := |T ∩ G| / |T| -89- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 if P(T) < α then rule is below the minimal acceptable training probability continue else L := L ∪ T LEM1 algorithm removes attribute-value pairs which are not required for the rule to stay valid end algorithm Appendix B Codon-Based Genetic Algorithm Flow Diagram Definitions: Support = number of rows rule correctly classified in training set Coverage = number of rows rule applies in training set Confidence Estimate = Support / Coverage If peptide sequence has multiple applicable rules, then support is difference of the sum of the support for activity and the sum of the support against activity. Then, the confidence estimate is the arithmetic mean difference between confidence of rules for activity and confidence for rules against activity. Annotated flow diagram for Figs. 30 and 31 Codon Representation Procedure: Input: P, List of peptide sequences Output: D, list of DNA codon sequences begin output for each peptide sequence p in P do -90- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 begin DNA codon sequence d for each amino acid in p do Select codon by uniform probability the reverse mapping in NCBI Codon Table 1 Add selected codon to d end amino acid end DNA codon sequence d end peptide sequence p end procedure While converting peptide sequences into codon sequences is non- deterministic, the reverse of converting codons to peptide sequences is deterministic. The codon representation is operated on by single point mutation or cross over operators. A benefit of the conversion is that since the stop codon may be generated, the use of this representation can dynamically change the length of the corresponding peptide sequence. Another benefit of the conversion is that single point mutations are more likely to change to a more similar amino acid than if amino acids are selected with uniform probability. Equation 1. Objective function for genetic algorithm (standard): ^^^^^^^^^^ ൌ ^^^^^^^^^^^^ ^^^^^^^^ℎ^^ ^^^^^^^^^^^^ ∗ ∑௧^^^^௧ ^^^^^^^^^^^^^^^^^ െ ^^^^^^^^^^^^^A standard convex scoring function is used with a known optimal value of zero. The targets, and the product of their user-specified weights, indicate the most probable directions for future genetic algorithm candidates to migrate. Given enough generations, the genetic algorithm will converge on similar sequences. Restarting the generations from an intermediate pool will most likely change the convergence sequences unless the scoring function has steep gradients for sequences outside of the original path. Targets: -91- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Maximum Support of Rules Peptide length > 7 amino acids Peptide length < 15 amino acids No cysteine residues The unique aspects of the peptide sequence generative method include the capability of changing design direction based on highly-selective rough set theory approximations of sequence properties and that use a codon translation to increase the ability to escape local minima. C. Example Antimicrobial Peptides (AMPs) for Targeting Select Microbial Populations 1. Preliminary Analysis FIGs. 32A–C depict graphs showing percentages of colony-forming units (%CFU) of various microbial populations with new candidate AMPs, over time. A legend for the candidate AMP sequences is listed below: TiBP-S5-AMPA: RPRENGRERGLGSGGGKKWKLWKKIEKWGQGIGAVLKWLTTW (SEQ ID NO: 168) TiBP-AH-AMP1: RPRENRGRERGLKGSVLSALKLLKKLLKLLKKL (SEQ ID NO: 10) FIG. 33A and 33B depict graphs showing percentages of colony-forming units (%CFU) of various microbial populations with new candidate AMPs, over time. A legend for the candidate AMP sequences is listed below: AMP1: LKLLKKLLKLLKKL(SEQ ID NO: 2) TiBP-S5-AMPA: RPRENGRERGLGSGGGKKWKLWKKIEKWGQGIGAVLKWLTTW (SEQ ID NO: 168) -92- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 TiBP-AH-AMP1:RPRENRGRERGLKGSVLSALKLLKKLLKLLKKL (SEQ ID NO: 10) TiBP-AH-GL13K: RPRENRGRERGLKGSVLSAGKIIKLKASLKLL (SEQ ID NO: 11) AMP7: ESYKKML (SEQ ID NO: 4) AMP10: GILGKLWEGVKSTF (SEQ ID NO: 5) AMP1: LKLLKKLLKLLKKL(SEQ ID NO: 2) AMP2-NH2: KWKRWWWWR-NH2 (SEQ ID NO: 3) AMPA: KWKLWKKIEKWGQGIGAVLKWLTTW (SEQ ID NO: 9) 2. S. Epidermidis Data FIG. 34 depicts a graph of an inhibition zone of various candidate novel antimicrobial peptides in targeting S. epidermidis. Table 1. Inhibition zone of S. epidermidis for candidate novel antimicrobial peptides 3. Species-Specific AMPs Single Biofilm -93- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Figs. 35A and B. Inhibition of planktonic bacteria (upper) and biofilms (lower) by 1st Gen ML-AMPs designed with relevance rules and genetic algorithm (VL-13, KK-15 and FV-11) compared to peptide (AMP1). A. actinomycetemcomitans D7S-1, S. gordonii ATCC 35105 and S. sanguinis ATCC 10556 were grown for 6 hrs (upper panel) or 4 hours (lower panel) in Shi with peptides (100 µM for planktonic cells, and 200 µM for biofilms) at 37°C in an atmosphere supplemented with 5% CO2. The CFUs were enumerated at T=0 (blue) or after incubation (orange). Fig. 35A showed three replicates with standard deviations. FIG. 35B showed the average of two replicates. VL-13 demonstrated >4 logs of inhibition of planktonic and biofilm Aa D7S-1 to below the detection limit (50 CFU / ml) with minimal or no inhibition of the other two species. AMP1 also showed species-biased inhibition of planktonic Aa D7S-1. All bacteria grew in Shi without peptides (data not shown) in 4 or 6 hours. 4. Generic Antibacterial Generated Sequences High Priority (Lowest 100 Scores) -94- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -95- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Low Priority (Highest 100 Scores) Sequence Score -96- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -97- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -98- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -99- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 5. Oral Species-Specific AMPs #1 of 2 High Inhibition for Aggregatibacter actinomycetemcomitans High Priority (Lowest 100 Scores) -100- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -101- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -102- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Low Priority (Highest 100 Scores) -103- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -104- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -105- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 6. Oral Species-Specific AMPs #2 of 2 Low Inhibition of Aggregatibacter actinomycetemcomitans High Priority (Lowest 100 Scores) -106- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -107- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -108- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Low Priority (Highest 100 Scores) -109- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -110- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 -111- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 D. Systems and Methods for Using Machine Learning (ML) Models to Identify Species-Biased Antimicrobial Peptides (SB-AMPs) to Target Select Microbial Populations Referring now to FIG. 36, depicted is a block diagram of a system 100 for using machine learning (ML) models to identify species-biased antimicrobial peptides (SB- AMPs) that target select microbial populations. In brief overview, the system 100 may include at least one data processing system 105, at least one gene sequencer 110, and at least one client device 115, communicatively coupled with one another via at least one network 120. The data processing system 105 may include at least one dataset manager 125, at least one property calculator 130, at least one model trainer 135, at least one rule checker 140, at least one sequence expander 145, at least one validation handler 150, at least one output evaluator 155, at least one rule generation model 165, at least one sequence creation model 170, and at least one database 175, among others. The gene sequencer 110 or the client device 115 may be associated with at least one test environment 180, among others. Each of the components in the system 100 (such as the data processing system 105 and its subcomponents and the client device 115) as detailed herein may be implemented using hardware (e.g., one or more processors coupled with memory), or a combination of hardware and software as detailed herein in Section E. The system 100 may be used to implement the functionalities described herein in Sections A and B. In further detail, the data processing system 105 may be any computing device including one or more processors coupled with memory and software and capable of performing the various processes and tasks described herein. The data processing system 105 may be associated with an entity to process peptide sequence data. The data processing system 105 can be in communication with the gene sequencer 110, the client device 115, the database 175, and other devices, via the network 120. The data processing system 105 may be situated, located, or otherwise associated with at least one server group. The server group may correspond to a data center, a branch office, or a site at which one or more servers corresponding to the data processing system 105 is situated. In some embodiments, the functionalities ascribed to the data processing system 105 may be performed on the client -112- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 device 115 (e.g., as a software installed thereon and executed using one or more processors coupled with memory). The data processing system 105 may include one or more modules, components, or subsystems to perform the various processes and tasks described herein. On the data processing system 105, the dataset manager 125 may retrieve a training dataset including antimicrobial peptides (AMP) sequences and non-AMP sequences. The property calculator 130 may generate a set of properties for each of the AMP and non-AMP sequences. The model trainer 135 may use the training dataset to train the rule generation model 165 to output rule sets to discriminate between AMP and non-AMP sequences. The rule checker 140 may compare the rule sets outputted by the rule generation model 165 against a sample AMPs. The sequence expander 145 may use the rule set to generate new, additional candidate AMP sequences. The validation handler 150 may manage testing of candidate AMPs corresponding to the candidate AMP sequences to target select microbial populations. The output evaluator 155 may generate output including information on the AMP sequences. The rule generation model 165 (sometimes herein referred to generally as a machine learning model) may include any type of machine learning (ML) or artificial intelligence (AI) architecture to generate rule sets to discriminate the AMP and non-AMP sequences. The architecture for the rule generation model 165 may include, for example, a deep learning artificial neural network (ANN) (e.g., an autoencoder, a convolutional neural network (CNN), a recurrent neural network (RNN), or a transformer), a Markov chain, a support vector machine (SVM), a clustering algorithm, a Bayesian classifier, or a decision tree, among others. In general, the rule generation model 165 may include a set of inputs, at least one output, and a set of weights arranged across a set of layers relating the inputs to the output. The set of inputs may include the AMP or non-AMP sequences and the properties for each sequence. The output may include a rule set defining values for the properties to discriminate between the AMP and non-AMP sequences. The set of weights may be in accordance with the ML or AI architecture. The rule generation model 165 may be trained by the model trainer 135 and the rule checker 140. -113- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 The sequence creation model 170 (sometimes herein referred to generally as a machine learning model) may include any type of machine learning (ML) or artificial intelligence (AI) architecture to create a set of candidate AMP sequences. The architecture for the rule generation model 165 may include, for example, a deep learning artificial neural network (ANN) (e.g., an autoencoder, a convolutional neural network (CNN), a recurrent neural network (RNN), or a transformer), a Markov chain, a support vector machine (SVM), a clustering algorithm, a Bayesian classifier, or a decision tree, among others. In general, the sequence creation model 170 may include a set of inputs, at least one output, and a set of weights arranged across a set of layers relating the inputs to the output. The set of inputs may include seed AMP sequences. The output may include new, candidate AMP sequences. The set of weights may be in accordance with the ML or AI architecture. The sequence creation model 170 may be trained using the set of rules generated by the rule generation model 165. The gene sequencer 110 may be any device to perform peptide sequencing on peptides. The peptide sequencing may be performed via Edman degradation or mass spectroscopy to yield the sequence data. For example, using mass spectroscopy, a protein may be extracted from a microbial sample. The protein may be broken down into peptides using an enzyme (e.g., trypsin). The peptides may be ionized (e.g., via electrospray ionization or matrix-assisted laser desorption or ionization) using an electromagnetic field. For each peptide, the mass-to-charge ratio may be measured to estimate precursor mass. The peptide ions (precursors) may be fragmented (e.g., at the peptide backbone) using collision- induced dissociation (CID) to create fragment ions (e.g., b-ions and y-ions). The results may be analyzed via mass spectrometer to generate a fragmentation spectrum. The spectrum may be used to search a database of peptides or may be used to reconstruct the peptide sequence. The sequence data may be stored and maintained as one or more files, such as comma- separate values (CSV) files. The gene sequencer 110 may be in communication with the data processing system 105, the client device 115, and the database 175 via the network 120. The gene sequencer 110 may be associated with the client device 115 or the test environment 180. The client device 115 (sometimes herein referred to as an end user computing device) may be any computing device comprising one or more processors coupled with memory and software and capable of performing the various processes and tasks described -114- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 herein. The client device 115 may be in communication with the data processing system 105, the gene sequencer 110, and the database 175 via the network 120. The client device 115 may have at least one display. The client device 115 may be associated with an entity (e.g., a clinician) examining the subject. The client device 115 may also be associated with the test environment 180. The display may present information about the subject provided by the data processing system 105. The test environment 180 may correspond to or may include any one or more components or settings used to synthesize and test peptides for targeting select microbial populations. The test environment 180 may include testing equipment, such as media, agents, assays, incubators, petri dishes, loops, pipettes, rules, microplate reader, counter, or data logger devices, among others. The test environment 180 may include one or more components to synthesize peptides, such as solid support resin (e.g., Wang resin), protection group (e.g., Fmoc), activating agents, purifiers, solvents (e.g., Dimethylformamide (DMF)), or cleavage reagents, among others. The test environment 180 may be administered and managed by the same entity as the gene sequencer 110 or the client device 115. For example, the test environment 180 may correspond to a laboratory room, containing the gene sequencer 110, the client device 115, and the equipment for testing and synthesizing peptides. The database 175 may store and maintain various resources and data associated with the data processing system 105, the gene sequencer 110, and the client device 115, among others. The database 175 may include a database management system (DBMS) to arrange and organize the data maintained thereon. The database 175 may be in communication with the data processing system 105, the gene sequencer 110, and the client device 115, via the network 120. While running various operations, the data processing system 105, the gene sequencer 110, and the client device 115 may access the database 175 to retrieve identified data therefrom. The data processing system 105, the gene sequencer 110, and the client device 115 may also write data onto the database 175 from running such operations. As used herein, “antimicrobial peptides” or “AMPs” refer to a class of small peptides that have a wide range of inhibitory effects against bacteria, fungi, and / or -115- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 protozoans. In contrast, “non-AMPs” refer to small peptides that lack any detectable or clinically relevant inhibitory activity against bacteria, fungi, and / or protozoans. Referring now to FIG. 37, depicted is a block diagram of a process 200 to determine properties of antimicrobial peptide (AMP) and non-AMP sequences in the system for using ML models to identify species-biased antimicrobial peptides (SB-AMPs). Under the process 200, the gene sequencer 110 may execute, carry out, or otherwise perform sequencing on a set of AMPs 205A–N (hereinafter generally referred to as AMPs 205). The gene sequencer 110 may also perform sequencing on a set of non-AMPs 210A–N (hereinafter referred to as non-AMPs 210). The AMPs 205 may include naturally occurring peptides (e.g., 10–200 amino acids long) that targets select microbial populations. In some embodiments, the microbial populations may comprise one or more species of bacteria. For instance, the AMPs 205 may exhibit inhibitive activity against the certain microbial populations. The select microbial populations may be associated with inflammatory disease of oral region, such as peri- implantitis or periodontitis, among others. The select microbial populations may include, for example, at least one of Porphyromonas gingivalis, Aggregatibacter actinomycetemcomitans, or Streptococcus gordonii, among others, including those listed herein. In addition, the non- AMPs 210 may include peptides that do not possess inhibitive activity against the select microbial populations. In some embodiments, the non-AMPs 210 may include peptides that are non-naturally occurring. From performing the sequencing on the AMPs 205, the gene sequencer 110 may produce, create, or otherwise generate a set of AMP sequences 215A–N (hereinafter generally referred to AMP sequences 215). Each AMP sequence 215 may identify or include a sequence of alphanumeric characters corresponding to amino acid code of a respective AMP 205. The AMP sequence 215 may be used as a structural identifier to the respective AMP 205, with each alphanumeric character defining an amino acid and a position of the amino acid in the respective AMP 205. The sequence of the alphanumeric characters may correspond to the sequence of amino acids from one terminus (e.g., N-terminus) to another terminus (e.g., C-terminus) of the respective AMP 205. The alphanumeric characters may include the standard one-letter codes for the 20 canonical amino acids (e.g., in accordance with International Union of Biochemistry and Molecular Biology (IUBMB) system). The AMP sequence 215 may have a length of 10–200 alphanumeric characters. -116- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 In addition, from performing the sequencing on the AMPs 210, the gene sequencer 110 may produce, create, or otherwise generate a set of non-AMP sequences 220A–N (hereinafter generally referred to non-AMP sequences 220). The non-AMP sequences 220 may be similar in format and structure as the AMP sequences 215. Each non- AMP sequence 220 may identify or include a sequence of alphanumeric characters corresponding to amino acid code of a respective non-AMP 210. The AMP sequence 220 may be used as a structural identifier to the respective non-AMP 210, with each alphanumeric character defining an amino acid and a position of the amino acid in the respective non-AMP 210. The alphanumeric characters may include the standard one-letter codes for the 20 canonical amino acids. The sequence of the alphanumeric characters may correspond to sequence of amino acids from one terminus (e.g., N-terminus) to another terminus (e.g., C- terminus) of the respective non-AMP 210. The alphanumeric characters may include the standard one-letter codes for the 20 canonical amino acids (e.g., in accordance with International Union of Biochemistry and Molecular Biology (IUBMB) system). The non- AMP sequence 220 may have a length of 10–200 alphanumeric characters corresponding to the respective non-AMP 210. With the generation of sequences, the gene sequencer 110 may store and maintain the set of AMP sequences 215 and the set of non-AMP sequences 220 on the database 175. In some embodiments, the gene sequencer 110 may send, transmit, or otherwise provide the set of AMP sequences 215 and the set of non-AMP sequences 220 to the data processing system 105. The dataset manager 125 may identify at least one training dataset 225 to include the set of AMP sequences 215 and the set of non-AMP sequences 220. The training dataset 225 may be used to initialize, train, and establish the rule generation model 165 or the sequence creation model 170, or both. The dataset manager 125 may receive, identify, or otherwise retrieve the training dataset 225 (e.g., from the database 175 or the gene sequencer 110). The training dataset 225 may be an initial training dataset 225, when used for the first iteration of training or when there are no additional AMP sequences 215 generated by the sequence creation model 170 included in the training dataset 225. For each of the AMP sequences 215 and the non-AMP sequences 220 of the training dataset 225, the property calculator 130 may compute, determine, or otherwise -117- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 generate a set of properties 230A–N (hereinafter generally referred to as properties 230). The set of properties 230 may define, specify, or identify various physical or functional characteristics of the respective AMP 205 or non-AMP 210. The set of properties 230 may identify or include one or more physical characteristics, such as hydropathy (e.g., hydrophobicity or hydrophilicity from polar or non-polar residues), amphipathicity (e.g., ability to possess both hydrophobicity and hydrophilicity), solubility (e.g., ability to dissolve in solvents), or secondary structure propensity (e.g., ability to form secondary structures, such as α-helices, β-sheet structure, or turns and loops upon contact), among others. The set of properties 230 may include, for example, microbial inhibitory activity for at least one microbial species of the select microbial populations, among others. For instance, the microbial inhibitory activity may correspond to an ability to inhibit or kill one or more microbial species of the select microbial population. In some embodiments, the property calculator 130 may calculate, generate, or otherwise determine at least one matrix for a respective peptide (e.g., AMP 205 or non-AMP 210) corresponding each of the AMP sequences 215 and the non-AMP sequences 220 of the training dataset 225. The matrix may identify or define a set of pair distances, with each distance defining a respective distance between a corresponding pair of amino acid residues in the peptide. For example, the property calculator 130 may use Evolutionary Scale Modeling 2 (ESM-2) to estimate or determine a protein structure defining the amino acid residues of the peptide. For each pair of amino acid residues in the estimated protein structure, the property calculator 130 may generate a respective pair distance. The property calculator 130 may iterate through the pairs of amino acid residues in the proteins structure to generate the set of pair distances. Using the set of pair distances, the property calculator 130 may generate the matrix to include the set of pair distances. Based on the set of pair distances in the matrix, the property calculator 130 may generate or determine a Fourier score for the peptide (e.g., corresponding to the AMP 205 or the non-AMP 210). The Fourier score may define, identify, or characterize the frequency domain of the matrix, such as a spectral energy, entropy, or frequency modes, among others. To determine, the property calculator 130 may convert or transform the matrix to the frequency domain, and use the frequency domain representation of the matrix to determine the Fourier score. -118- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 In some embodiments, the property calculator 130 may determine or identify at least one secondary structure of the peptide (e.g., corresponding to the AMP 205 or the non-AMP 210). The secondary structure may define or identify a spatial arrangement of a polypeptide backbone of the peptide. For instance, the property calculator 130 may scan the protein structure determined using ESM-2 to identify one or more secondary structures (e.g., α-helices, β-sheet structure, or turns and loops) on the polypeptide backbone of the peptide. In some embodiments, the set of properties 230 may identify or include one or more of the physical characteristics, the microbial inhibitory activity, the matrix, the Fourier score, or secondary structure, among others. With the generation of the set of properties 230, the property calculator 130 may add, insert, or otherwise include the set of properties 230 to the training dataset 225. The property calculator 130 may iterative or traverse through the set of AMP sequences 215 and the non-AMP sequences 220 to generate the corresponding sets of properties 230 to include in the training dataset 225. Referring now to FIG. 38, depicted is a block diagram of a process 300 to generate rule sets to discriminate antimicrobial peptide (AMP) and non-AMP sequences and new candidate AMP sequences in the system for using ML models to identify species-biased antimicrobial peptides (SB-AMPs). Under the process 300, the model trainer 135 may apply, feed, or otherwise provide the training dataset 225 toe the rule generation model 165. The input may include the set of AMP sequences 215, the set of non-AMP sequences 220, and the set of properties 230 for each of the set of AMP sequences 215 and the set of non-AMP sequences 220, among others. For the initial iteration, the input may lack any new AMP sequences 215". In feeding, the model trainer 135 may process the input in accordance with the set of weights arranged across the set of layers of the rule generation model 165. Based on providing the training dataset 225 as the input to the rule generation model 165, the model trainer 135 may produce, construct, or otherwise generate at least one rule set 305. The rule set 305 may specify, define, or otherwise identify values for the set of properties 230 to distinguish, differentiate, or discriminate the AMP sequences 215 and the non-AMP sequences 220. The values of the rule set 305 may define a range (e.g., a lower and upper limits), a statistical measure (e.g., a mean, median, or variance), or another metric (e.g., sum or a window) to use as boundary conditions to differentiate the set of properties -119- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 230 associated with the AMP sequences 215 versus the set of properties 230 associated with the non-AMP sequences 220. By extension, the values of the rule set 305 may define the boundary conditions to discriminate between the physical or functional characteristics of the AMPs 205 versus the physical or functional characteristics of the non-AMP 210. For example, the rule set 305 may include boundary conditions in the form of upper and lower bounds for: physical characteristics, such as mean polarity, interfacial hydrophobicity, and a hydrophobicity at a defined pH level; structural characteristics, such as side chain interactions, retention coefficient in a particular substance (e.g., NaH2PO4), full non-bonding orbitals; and microbial inhibitory activity (e.g., growth inhibition) for the select microbial population, among others. The rule set 305 may be stored and maintained by the model trainer 135 using one or more data structures or files, such as a matrix, a table, extensible markup (XML) file, or comma- separated values (CSV) file, among others. In some embodiments, the model trainer 135 may produce, generate, or otherwise determine a set of embeddings from at least one layer of the rule generation model 165 based on providing the input to the rule generation model 165. The set of embeddings may be obtained from an output a layer within the rule generation model 165. For example, the model trainer 135 may identify the set of embeddings outputted by an intermediary layer between an encoder and a decoder in an autoencoder model used to implement the rule generation model 165. The model trainer 135 may identify the set of embeddings from a penultimate layer (e.g., prior to the last activation layer) in the autoencoder model. The set of embeddings may identify or include features associated with the boundary condition to differentiate the set of properties 230 associated with the AMP sequences 215 versus the set of properties 230 associated with the non-AMP sequences 220. In some embodiments, the model trainer 135 may apply or use an approximator on the rule set 305 to decrease, lower, or otherwise reduce a number of the boundary conditions in the rule set 305. The approximator may include a rough set theory algorithm (e.g., the LEM2 or CLN-MLEM2 as detailed herein) to generalize the boundary conditions of the rule set 305 to reduce the number of boundary conditions. The resultant boundary condition generated using the approximator may define values for the set of properties 230 that are wider in range than the original boundary conditions defined in the -120- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 rule set 305. In applying the approximator, for each boundary condition defined in the initial rule set 305, the model trainer 135 may determine or identify whether another boundary condition of the initial rule set 305 at least partially overlaps. If there are no other boundary conditions that overlap with the boundary condition, the model trainer 135 may traverse over the boundary conditions of the rule set 305 to identify another boundary condition. If the two boundary conditions are determined to be at least partially overlapping, the model trainer 135 may determine or generate an approximate boundary condition to cover or correspond to the values of the two boundary conditions. The approximate boundary condition may include conjunctive values for the set of properties 230 as defined in the original two boundary conditions. With the generation, the model trainer 135 may substitute or replace the two boundary conditions with the approximate boundary condition in the rule set 305. In this manner, the new approximate boundary condition may be more generalizable than the two more specified boundary conditions. The model trainer 135 may repeat this process over the boundary conditions specified by the rule set 305. In some embodiments, the model trainer 135 may calculate or determine a training probability for each boundary condition. In some embodiments, the training probability may identify a likelihood that input peptide sequences and properties, against which the boundary condition is to be checked, belongs to the values defined by the boundary condition. In some embodiments, the training probability may define a confidence level that input peptide sequences and properties support or fall under the boundary condition. The model trainer 135 may compare the training probability of the boundary condition with a threshold. If the training probability satisfies (e.g., greater than or equal to) the threshold, the model trainer 135 may maintain the boundary condition in the rule set 305. Otherwise, if the training probability does not satisfy (e.g., less than) the threshold, the model trainer 135 may remove the boundary condition from the rule set 305. The model trainer 135 may use the rule set 305 with the reduced number of boundary conditions for further processing. The rule checker 140 may identify or determine whether the rule set 305 defines values for the set of properties 230 that satisfy values of properties of at least one sample dataset 310. The sample dataset 310 may include a set of AMP sequences derived -121- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 from a microbial sample (e.g., a commensal sample containing AMPs or non-pathogenic microbes). As used herein, a “commensal sample” refers to a biological sample comprising microbes that reside in or on another organism (the host) without causing harm to the host. In some embodiments, the sample dataset 310 may include a set of non-AMP sequences isolated from the microbial sample or a reference sample that is substantially free of microbes. The sample dataset 310 may also include a set of properties for each of the set of AMP sequences derived from the microbial sample. To determine, the rule checker 140 may compare values defined by the rule set 305 against the values of the properties of the sample dataset 310. The comparison may be to determine whether the ranges defined by the boundary conditions of the rule set 305 encompass or cover the values of the properties in the sample dataset 310. In some embodiments, the rule checker 140 may determine or identify a number of correct classifications and a number of incorrect classifications using the rule set 305. The number of correct classifications may correspond to a number of AMP sequences of the sample dataset 310 that were correctly identified as an AMP sequence (or vice-versa) using the rule set 305. Conversely, the number of incorrect classifications may correspond to a number of AMP sequences of the sample dataset 310 that were incorrectly identified as a non-AMP sequence (or vice-versa) using the rule set 305. The rule checker 140 may calculate, generate, or otherwise determine at least one loss metric 315 based on the comparison. The loss metric 315 may define or identify a degree of deviation of the rule set 305 outputted by the rule generation model 165 versus the sample dataset 310. The loss metric 315 may be calculated with any number of loss functions, such as any number of loss functions, such as a norm loss (e.g., L1 or L2), mean absolute error (MAE), mean squared error (MSE), a quadratic loss, a cross-entropy loss, and a Huber loss, among others. In some embodiments, the loss metric 315 may be determined as a function of the number of correct or incorrect classifications, such as classification error rate, approximation loss, or membership loss metric, among others. In general, the more than the ranges defined by the boundary conditions of the rule set 305 do not cover the values of the properties in the sample dataset 310, the higher the loss metric 315 may be. Conversely, the more than the ranges defined by the boundary conditions of the rule set 305 cover the values of the properties in the sample dataset 310, the lower the loss metric 315 may be. -122- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 The rule checker 140 may modify, change, or otherwise update the rule generation model 165 using the loss metric 315. The updating of the weights may be in accordance with a back propagation and optimization function (sometimes referred to herein as an objective function) with one or more parameters (e.g., learning rate, momentum, weight decay, and number of iterations). The optimization function may define one or more parameters at which the weights of the rule generation model 165 are to be updated. The optimization function may be in accordance with stochastic gradient descent, and may include, for example, an adaptive moment estimation (Adam), implicit update (ISGD), and adaptive gradient algorithm (AdaGrad), among others. The rule checker 140 can iteratively train the rule generation model 165 until convergence. Upon convergence, the rule checker 140 can store and maintain the set of weights of the rule generation model 165 for use in inference. In conjunction, the sequence expander 145 may use the rule set 305 to initialize, train, and establish the sequence creation model 170. In some embodiments, the sequence expander 145 may initiate training of the sequence creation model 170, subsequent to completion of the training of the rule generation model 165. The rule set 305 used to train the sequence creation model 170 may be generated by the rule generation model 165 upon completion of training of the rule generation model 165. The sequence creation model 170 may be trained to generate AMP sequences with values of properties that satisfy at least some of the values of the properties defined by the rule set 305. The AMP sequences may include one or more modifications relative to the input seed AMP sequence, such as nucleotide mutations (e.g., substitutions, insertions, or deletions), translocation, or codon frame shifts, among others. In some embodiments, the sequence creation model 170 may be trained by the sequence expander 145 to generate AMP sequences with a microbial inhibitory activity that satisfies a threshold level. To train, the sequence expander 145 may feed, apply, or otherwise provide one or more seed AMP sequences as input to the sequence creation model 170. The one or more AMP sequences may correspond to AMPs identified as targeting (e.g., exhibiting inhibitory activity) the select microbial populations. In some embodiments, the set of seed AMP sequences may include at least a portion of the original AMP sequences 215 generated -123- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 from the AMPs 205. In some embodiments, the set of seed AMP sequences may include sequences derived from AMPs besides the AMPs 205 used to derive the original set of AMP sequences 215. In feeding, the sequence expander 145 may process the input in accordance with the set of weights arranged across the set of layers of the sequence creation model 170. Based on providing the input to the sequence creation model 170, the sequence expander 145 may generate one or more corresponding AMP sequences. The output AMP sequences may include one or more modifications relative to the input seed AMP sequences. The sequence expander 145 may identify or determine whether the output AMP sequence have values of properties that satisfy the values of the properties (e.g., target antimicrobial activity) defined by the rule set 305. When the values of the properties of the output AMP sequence do not satisfy the rule set 305, the sequence expander 145 may identify the output AMP sequence as an incorrect generation. When the values of the properties of the output AMP sequence the rule set 305, the sequence expander 145 may identify the output AMP sequence as a correct generation. Based on the determination, the sequence expander 145 may calculate, generate, or otherwise determine at least one loss metric based on the comparison. The loss metric may define or identify a degree of deviation of the values of the properties of the AMP sequences generated by the sequence creation model 170, relative to the values of the properties as defined by the rule set 305 or the target antimicrobial activity level. The loss metric may be calculated with any number of loss functions, such as any number of loss functions, such as a norm loss (e.g., L1 or L2), mean absolute error (MAE), mean squared error (MSE), a quadratic loss, a cross-entropy loss, and a Huber loss, among others. The sequence expander 145 may modify, change, or otherwise update the sequence creation model 170 using the loss metric. The updating of the weights may be in accordance with a back propagation and optimization function (sometimes referred to herein as an objective function) with one or more parameters (e.g., learning rate, momentum, weight decay, and number of iterations). The optimization function may define one or more parameters at which the weights of the sequence creation model 170 are to be updated. The optimization function may be in accordance with stochastic gradient descent, and may include, for example, an adaptive moment estimation (Adam), implicit update (ISGD), and -124- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 adaptive gradient algorithm (AdaGrad), among others. The sequence expander 145 can iteratively train the sequence creation model 170 until convergence. Upon convergence, the sequence expander 145 can store and maintain the set of weights of the sequence creation model 170 for inference. With the establishment of the sequence creation model 170, the sequence expander 145 may create, produce, or otherwise generate a set of candidate AMP sequences 215’A–N (hereinafter generally referred to as candidate AMP sequences 215’). To generate, the sequence expander 145 may provide one or more seed AMP sequences 320A–N (hereinafter generally referred to as seed AMP sequences 320) as input to the sequence creation model 170. The one or more AMP sequences 320 may correspond to AMPs identified as targeting (e.g., exhibiting inhibitory activity) the select microbial populations. In some embodiments, the set of seed AMP sequences 320 may include at least a portion of the original AMP sequences 215 generated from the AMPs 205. In some embodiments, the set of seed AMP sequences 320 may include sequences derived from AMPs besides the AMPs 205 used to derive the original set of AMP sequences 215. In feeding, the sequence expander 145 may process the input in accordance with the set of weights arranged across the set of layers of the sequence creation model 170. Based on providing the input to the sequence creation model 170, the sequence expander 145 may generate the candidate AMP sequences 215’. The candidate AMP sequences 215’ may include one or more modifications relative to the input seed AMP sequences 320. From the initial set of candidate AMP sequences 215’, the sequence expander 145 may use or apply the rule set 305 to select or identify a subset of candidate AMP sequences 215"A–N (hereinafter generally referred to as candidate AMP sequences 215"). The subset of candidate AMP sequences 215" may have properties that satisfy the values of properties as defined by the rule set 305 to discriminate the AMP sequences 215 and the non- AMP sequences 220. For each candidate AMP sequence 215’, the sequence expander 145 may determine a set of properties (e.g., similar to the set of properties 230), such as one or more of the physical characteristics, the microbial inhibitory activity, the matrix, the Fourier score, or secondary structure, among others. The sequence expander 145 may identify or determine whether the values of the properties of the candidate AMP sequence 215’ satisfy -125- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 the values of the properties defined by the rule set 305. When the values of the candidate AMP sequence 215’ satisfy the rule set 305, the sequence expander 145 may include the candidate AMP sequence 215’ as part of the subset of candidate AMP sequences 215". In contrast, when the values of the candidate AMP sequence 215’ do not satisfy the rule set 305, the sequence expander 145 may exclude the candidate AMP sequence 215’ from the subset of candidate AMP sequences 215". The sequence expander 145 may repeat the determination of the set of candidate AMP sequences 215’. In some embodiments, the sequence expander 145 may calculate, generate, or otherwise determine at least one similarity metric 325A–N (hereinafter generally referred to the similarity metric 325) for each candidate AMP sequence 215’. The similarity metric 325 may identify or indicate a degree of similarity of the candidate AMP sequence 215’ with one or more of a set of reference AMP sequences identified as targeting the select microbial populations. The degree of similarity may be in terms of the sequence of alphanumeric characters forming the candidate AMP sequence 215’ itself relative to the sequence of alphanumeric characters of the one or more of a set of reference AMP sequences. The sequence expander 145 may identify or determine whether the similarity metric 325 satisfies a threshold metric. When the similarity metric 325 satisfies (e.g., greater than or equal to) the threshold metric, the sequence expander 145 may include the candidate AMP sequence 215’ as part of the subset of candidate AMP sequences 215". In contrast, when the similarity metric 325 do not satisfy (e.g., less than) the threshold metric, the sequence expander 145 may exclude the candidate AMP sequence 215’ from the subset of candidate AMP sequences 215". The sequence expander 145 may repeat the determination of the set of candidate AMP sequences 215’. With the generation, the sequence expander 145 may store and maintain the subset of AMP sequences 215" using one or more data structures or files, such as a matrix, a table, extensible markup (XML) file, or comma-separated values (CSV) file, among others. Referring now to FIG. 39, depicted is a block diagram of a process 400 to validate inhibition levels of candidate antimicrobial peptides (AMPs) in the system using ML models to identify species-biased antimicrobial peptides (SB-AMPs). Under the process 400, the validation handler 150 may create or generate at least one test package 405 using the candidate AMP sequences 215". The test package 405 may identify or include at least one of -126- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 the subset of the candidate AMP sequences 215" to be validated for inhibition activity against the select microbial populations. The test package 405 may also identify or include information (e.g., in text format) indicating testing procedures for the candidate AMP sequences 215". The validation handler 150 may send, transmit, or otherwise provide the test package 405 to the client device 115. The client device 115 may retrieve, obtain, or otherwise receive the test package 405 from the data processing system 105. With receipt, the client device 115 may render, display, or otherwise present the information of the test package 405, such as at least one of the subset of candidate AMP sequence 215" and testing procedures. A user (e.g., a clinician or laboratory staff) may use the testing environment 180 to run or carry out experiments to validate the efficacy of an AMP corresponding to the subset of candidate AMP sequence 215". In some embodiments, the testing may be in accordance as detailed herein in Sections A–C. In running the experiment, at least one microbial model 410 may be created and prepared in the testing environment 180. The microbial model 410 may include a single-species microbial model (e.g., with a single species of the select microbial population) or a poly-species microbial model (e.g., with multiple species of the select microbial population). For example, the microbial model 410 may include Porphyromonas gingivalis, Aggregatibacter actinomycetemcomitans, or Streptococcus gordonii, among others, including those listed herein. To the test environment 180, a synthesized AMP 415 corresponding to at least one of the subset of candidate AMP sequences 215" may be inserted or added to the microbial model 410. The AMP 415 corresponding to a respective candidate AMP sequences 215" may be synthesized, for example, using solid-phase peptide synthesis (SPPS). For example, the synthesized AMP 415 may be created as detailed herein in Sections A–C. The amount or concentration of the synthesized AMP 415 may encompass at least one of a minimum inhibitory concentration (MIC) or minimum bactericidal concentration (MBC), among others. Once added to the microbial model 410, the synthesized AMP 415 may be evaluated and assessed for level of inhibitory activity against the select microbial population in the microbial model 410. For example, assessing the level of inhibitory activity may include measuring the amount of the select microbial population (e.g., measured in terms of -127- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 colony-forming units per milliliter (CFU / ml)) in the microbial model 410 upon exposure to the synthesized AMP 415. The test may be run over a defined period of time, ranging from 1 hour, 2 hours, 4 hours, 6 hours, 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, or 1 week, among others, relative to the addition of the synthesized AMP 415. The client device 115 may produce, create, or otherwise generate at least one result package 425 using the assessment of the synthesized AMP 415 with respect to the inhibitory activity against the select microbial population in the microbial model 410 of the test environment 180. For each candidate AMP sequence 215" for which the corresponding synthesized AMP 415 is tested, the client device 115 may identify or determine at least one efficacy score 420A–N (hereinafter generally referred to as efficacy score 420). The efficacy scores 420 may be inputted or entered by the user of the client device 115 that is managing the test in the test environment 180. The efficacy score 420 may identify or indicate degree of efficacy (e.g., inhibitory antimicrobial activity) of the AMP 415 corresponding to the candidate AMP sequence 215" against the select microbial populations. For example, the efficacy score 420 may correspond to the amount of the select microbial population in the microbial model 410, after the defined of time subsequent to the addition of the synthesized AMP 415. The client device 115 may create, produce, or otherwise generate at least one result package 425 to include the set of efficacy scores 425 for the subset of candidate AMP sequences 215". With the generation, the client device 115 may transmit, send, or otherwise provide the result package 425 to the data processing system 105. The validation handler 150 may retrieve, identify, or otherwise receive the result package 425 from the client device 115. For each candidate AMP sequence 215", the validation handler 150 may determine or identify whether the respective AMP 415 is effective or ineffective in targeting the select microbial population. To identify, the validation handler 150 may compare the efficacy score 420 of the candidate AMP sequence 215" (and by extension, the respective AMP 415) to a threshold score. If the efficacy score 420 satisfies (e.g., is greater than or equal to) the threshold, the validation handler 150 may identify the AMP 415 as effective in targeting the select microbial population. When the AMP 415 is identified as effective, the validation handler 150 may include, insert, or add the corresponding AMP sequence 215" (e.g., new AMP sequence 215"X) to the set of AMP sequence 215 of the training dataset 225. -128- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 In some embodiments, the validation handler 150 may may store and maintain the new AMP sequence 215"X with an indication as effective using one or more data structures or files, such as a matrix, a table, extensible markup (XML) file, or comma-separated values (CSV) file, among others. On the other hand, if the efficacy score 420 does not satisfy (e.g., is less than) threshold, the validation handler 150 may identify the AMP 415 as ineffective in targeting the select microbial population. When the AMP 415 is identified as ineffective, the validation handler 150 may refrain from adding the corresponding AMP sequence 215" to the training dataset 225. In some embodiments, the validation handler 150 may may store and maintain the new AMP sequence 215" with an indication as ineffective using one or more data structures or files Using the new training dataset 225 with one or more new AMP sequences 215"X, the training of the rule generation model 165 and the sequence creation model 170 as detailed herein may be repeated. For example, for each new AMP sequences 215"X, the property calculator 130 may determine a set of properties 230 for the new AMP sequence 215"X, such as one or more of the physical characteristics, the microbial inhibitory activity, the matrix, the Fourier score, or secondary structure, among others. The property calculator may add the set of properties 230 for each new AMP sequence 215"X to the training dataset 225. The model trainer 135 may provide the training dataset 225 with the additional the set of properties 230 for each new AMP sequence 215"X as input to the rule generation model 165. The input may include the original set of AMP sequences 215, the original set of non- AMP sequences 220, the original set of properties 230 for each of the original set of AMP sequences 215, the original set of non-AMP sequences 220, the one or more new AMP sequences 215"X, and the additional the set of properties 230 for each new AMP sequence 215"X, among others. Based on providing the input, the model trainer 135 may generate a new rule set 305 to define the values for the set of properties 230 to distinguish, differentiate, or discriminate the AMP sequences 215 (and the one or more AMP sequences 215"X) versus the non-AMP sequences 220. The rule checker 140 may determine whether the new rule set 305 defines values for the set of properties 230 that satisfy values of properties of at least one sample dataset 310. Based on the comparison, the rule checker 140 may generate a new loss -129- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 metric 315 and update the rule generation model 165 using the new loss metric 315. In addition, the sequence expander 145 may generate a new set of candidate AMP sequences 215’ using the sequence creation model 170 and the one or more seed AMP sequences 320. The seed sequences 320 in this iteration may differ from the seed sequences 320 used in the previous iteration. The sequence expander 145 may apply the new rule set 305 to the set of candidate AMP sequences 215’ to select a subset of candidate AMP sequences 215". The process of the training and using of the rule generation model 165 and the sequence creation model 170 to create new candidate AMP sequences may be repeated any number of times, for example, for at least a set number of times (e.g., 2–100 times). Referring now to FIG. 40, depicted is a block diagram of a process 500 to provide therapy to dental tissues in the system for using ML models to identify species-biased antimicrobial peptides (SB-AMPs). Under the process 500, the user of the client device 115 may be a clinician (e.g., dentist or oral surgeon) examining at least one subject 505, in particular an oral region 510. The oral region 510 may correspond to a mouth or oral cavity within the subject 505 and the surrounding tissues and anatomical structures. The oral region 510 may include a dental tissue 515 with at least one dental implant 520. The subject 505 may be at risk of or diagnosed with inflammatory disease of the oral region 510, such as peri- implantitis or periodontitis in the dental tissue 515, such as a portion of the dental tissue 515 about the dental implant 520. The client device 115 may create, produce, or otherwise generate at least one indication 525. The indication 525 may be generated in response to the user inputting or entering interactions via the client device 115 (e.g., on a user interface) to indicate potential risk or diagnosis of the inflammatory disease of the oral region 510. The user may have also examined the dental tissue 515 for various microbial populations, such as Porphyromonas gingivalis, Aggregatibacter actinomycetemcomitans, or Streptococcus gordonii, among others, including those listed herein. The indication 525 may identify of the select microbial populations in the dental tissue 515 of the subject 505. With the generation, the client device 115 may transmit, send, or otherwise provide at least one indication 525 to the data processing system 105. -130- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 The output evaluator 155 may retrieve, identify, or otherwise receive the indication 525 from the client device 115. With receipt, the output evaluator 155 may process or parse the indication 525 to extract or identify the select microbial populations marked as present in the dental tissue 515 of the subject 505. The output evaluator 155 may select or identify at least one AMP sequence 530 from the set of AMP sequences 215 or the one or more new AMP sequences 215" that target the identified microbial population. The AMP sequence 530 may have been identified as effective in targeting the identified microbial population. With the identification, the output evaluator 155 may send, transmit, or otherwise provide at least one output 535 to the client device 115. The output 535 may identify or include the AMP sequence 530. In some embodiments, the output 535 may indicate that a therapy 540 including AMP corresponding to the AMP sequence 530 is (e.g., as a recommendation) to be administered to the oral region 510 of the subject 505. The client device 115 may retrieve, identify, or otherwise receive the output 535 from the output evaluator 155. With receipt, the client device 115 may display, render, or present information from the output 535, such as the AMP sequence 530. The user may use the information presented via the client device 115 to decide whether to administer the therapy 540. At least a portion of the therapy 540 may include administering an effective amount (e.g., minimum inhibitory concentration (MIC) or minimum bactericidal concentration (MBC)) of an AMP corresponding to the AMP sequence 530. The therapy 540 including the AMP corresponding to the AMP sequence 530 may be delivered, provided, or otherwise administered to the dental tissue 515 of the subject 505 to address the oral inflammatory disease (e.g., peri-implantitis or periodontitis). The delivery of the therapy 540 may be in the form of a topical form (e.g., gel or mouth rinse). For instance, the user may apply the gel (an example of the therapy 540) containing the AMP corresponding to the AMP sequence 530 to the dental tissue 515 about the dental implant 520 to address the oral inflammatory disease. In this manner, the data processing system 105 may use the rule generation model 165 and the sequence creation model 170 to generate new AMP sequences 215"X that can effectively target select microbial populations. The rule generation model 165 may allow for the creation of rule sets 305 that are interpretable and relatable to various physical and -131- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 functional characteristics of AMPs 205 and non-AMPs 210. The sequence expander 145 may allow for the generation of new AMP sequences 215" from modifications of previous AMP sequences 215, while satisfying target boundary conditions or inhibitory activity levels. The rule generation model 165 and the sequence creation model 170 may be leveraged to find complex relationship among AMP and non-AMP sequence data and property data. By generating new AMP sequences 215" predicted to satisfy target inhibitory activity levels, the data processing system 105 may significantly shorten the amount of time to find potential AMPs that target microbial populations associated with oral inflammatory disease. As the time is greatly shortened, the consumption of computing resources, such as processor, memory, and power, as well as network bandwidth used by the data processing system 105 may be significantly reduced, thereby freeing up such resources for other uses. Referring now to FIG. 41, depicted is a flow diagram of a method 600 of using machine learning (ML) models to identify species-biased antimicrobial peptides (SB-AMPs) that target select microbial populations. The method 600 may be implemented or performed by any of the components detailed herein, such as the system 100 or the system 700. Under the method 600, a computing system may retrieve a training dataset including antimicrobial peptides (AMP) sequences and non-AMP sequences (605). The computing system may determine a set of properties for each of the AMP and non-AMP sequences (610). The computing system may provide the set of AMP and non-AMP sequences along with the set of properties for each sequence to a machine learning (ML) model (615). The computing system may generate a rule set to discriminate the AMP sequences versus the non-AMP sequences (620). The computing system may generate a set of candidate AMP sequences (625). The computing system may apply the rule set to select a candidate AMP sequence from the set of candidate AMP sequences using the rule set (630). An inhibition of an AMP corresponding to the candidate AMP sequence may be validated against a select microbial population (635). Based on the validation, an inhibition score for the AMP may be determined (640). The computing system may determine whether the AMP corresponding to the candidate AMP sequence is effective based on the inhibition score (645). If the AMP is determined to be effective, the computing system may include the new AMP sequence to the training dataset (650). Otherwise, if the AMP is -132- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 determined to be ineffective, the computing system may exclude the new AMP sequence from the training dataset (655). The computing system may repeat the method 600 from step (605) for a set number of iterations. E. Computing and Network Environment Various operations described herein can be implemented on computer systems. FIG. 42 shows a simplified block diagram of a representative server system 700, client computing system 714, and network 726 usable to implement certain embodiments of the present disclosure. In various embodiments, server system 700 or similar systems can implement services or servers described herein or portions thereof. Client computing system 714 or similar systems can implement clients described herein. The system 100 described herein can be similar to the server system 700. Server system 700 can have a modular design that incorporates a number of modules 702 (e.g., blades in a blade server embodiment); while two modules 702 are shown, any number can be provided. Each module 702 can include processing unit(s) 704 and local storage 706. Processing unit(s) 704 can include a single processor, which can have one or more cores, or multiple processors. In some embodiments, processing unit(s) 704 can include a general-purpose primary processor as well as one or more special-purpose co-processors such as graphics processors, digital signal processors, or the like. In some embodiments, some or all processing units 704 can be implemented using customized circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In other embodiments, processing unit(s) 704 can execute instructions stored in local storage 706. Any type of processors in any combination can be included in processing unit(s) 704. Local storage 706 can include volatile storage media (e.g., DRAM, SRAM, SDRAM, or the like) and / or non-volatile storage media (e.g., magnetic or optical disk, flash memory, or the like). Storage media incorporated in local storage 706 can be fixed, removable or upgradeable as desired. Local storage 706 can be physically or logically divided into various subunits such as a system memory, a read-only memory (ROM), and a -133- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 permanent storage device. The system memory can be a read-and-write memory device or a volatile read-and-write memory, such as dynamic random-access memory. The system memory can store some or all of the instructions and data that processing unit(s) 704 need at runtime. The ROM can store static data and instructions that are needed by processing unit(s) 704. The permanent storage device can be a non-volatile read-and-write memory device that can store instructions and data even when module 702 is powered down. The term “storage medium” as used herein includes any medium in which data can be stored indefinitely (subject to overwriting, electrical disturbance, power loss, or the like) and does not include carrier waves and transitory electronic signals propagating wirelessly or over wired connections. In some embodiments, local storage 706 can store one or more software programs to be executed by processing unit(s) 704, such as an operating system and / or programs implementing various server functions such as functions of the system 100 of FIG. 36 or any other system described herein, or any other server(s) associated with system 100 or any other system described herein. “Software” refers generally to sequences of instructions that, when executed by processing unit(s) 704 cause server system 700 (or portions thereof) to perform various operations, thus defining one or more specific machine embodiments that execute and perform the operations of the software programs. The instructions can be stored as firmware residing in read-only memory and / or program code stored in non-volatile storage media that can be read into volatile working memory for execution by processing unit(s) 704. Software can be implemented as a single program or a collection of separate programs or program modules that interact as desired. From local storage 706 (or non-local storage described below), processing unit(s) 704 can retrieve program instructions to execute and data to process in order to execute various operations described above. In some server systems 700, multiple modules 702 can be interconnected via a bus or other interconnect 708, forming a local area network that supports communication between modules 702 and other components of server system 700. Interconnect 708 can be implemented using various technologies including server racks, hubs, routers, etc. -134- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 A wide area network (WAN) interface 710 can provide data communication capability between the local area network (interconnect 708) and the network 726, such as the Internet. Technologies can be used, including wired (e.g., Ethernet, IEEE 702.3 standards) and / or wireless technologies (e.g., Wi-Fi, IEEE 702.11 standards). In some embodiments, local storage 706 is intended to provide working memory for processing unit(s) 704, providing fast access to programs and / or data to be processed while reducing traffic on interconnect 708. Storage for larger quantities of data can be provided on the local area network by one or more mass storage subsystems 712 that can be connected to interconnect 708. Mass storage subsystem 712 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network-attached storage, and the like can be used. Any data stores or other collections of data described herein as being produced, consumed, or maintained by a service or server can be stored in mass storage subsystem 712. In some embodiments, additional data storage resources may be accessible via WAN interface 710 (potentially with increased latency). Server system 700 can operate in response to requests received via WAN interface 710. For example, one of modules 702 can implement a supervisory function and assign discrete tasks to other modules 702 in response to received requests. Work allocation techniques can be used. As requests are processed, results can be returned to the requester via WAN interface 710. Such operation can generally be automated. Further, in some embodiments, WAN interface 710 can connect multiple server systems 700 to each other, providing scalable systems capable of managing high volumes of activity. Other techniques for managing server systems and server farms (collections of server systems that cooperate) can be used, including dynamic resource allocation and reallocation. Server system 700 can interact with various user-owned or user-operated devices via a wide-area network such as the Internet. An example of a user-operated device is shown in FIG. 12 as client computing system 714. Client computing system 714 can be implemented, for example, as a consumer device such as a smartphone, other mobile phone, -135- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 tablet computer, wearable computing device (e.g., smart watch, eyeglasses), desktop computer, laptop computer, and so on. For example, client computing system 714 can communicate via WAN interface 710. Client computing system 714 can include computer components such as processing unit(s) 716, storage device 718, network interface 720, user input device 722, and user output device 724. Client computing system 714 can be a computing device implemented in a variety of form factors, such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, or the like. Processing unit(s) 716 and storage device 718 can be similar to processing unit(s) 704 and local storage 706 described above. Suitable devices can be selected based on the demands to be placed on client computing system 714; for example, client computing system 714 can be implemented as a “thin” client with limited processing capability or as a high-powered computing device. Client computing system 714 can be provisioned with program code executable by processing unit(s) 716 to enable various interactions with server system 700. Network interface 720 can provide a connection to the network 726, such as a wide area network (e.g., the Internet) to which WAN interface 710 of server system 700 is also connected. In various embodiments, network interface 720 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various RF data communication standards such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, LTE, etc.). User input device 722 can include any device (or devices) via which a user can provide signals to client computing system 714. The client computing system 714 can interpret the signals as indicative of particular user requests or information. In various embodiments, user input device 722 can include any or all of a keyboard, touch pad, touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, and so on. -136- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 User output device 724 can include any device via which client computing system 714 can provide information to a user. For example, user output device 724 can include a display to display images generated by or delivered to client computing system 714. The display can incorporate various image generation technologies, e.g., a liquid crystal display (LCD), light-emitting diode (LED) including organic light-emitting diodes (OLED), projection system, cathode ray tube (CRT), or the like, together with supporting electronics (e.g., digital-to-analog or analog-to-digital converters, signal processors, or the like). Some embodiments can include a device such as a touchscreen that function as both input and output device. In some embodiments, other user output devices 724 can be provided in addition to or instead of a display. Examples include indicator lights, speakers, tactile “display” devices, printers, and so on. Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a computer-readable storage medium. Many of the features described in this specification can be implemented as processes that are specified as a set of program instructions encoded on a computer-readable storage medium. When these program instructions are executed by one or more processing units, they cause the processing unit(s) to perform various operation indicated in the program instructions. Examples of program instructions or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter. Through suitable programming, processing unit(s) 704 and 716 can provide various functionality for server system 700 and client computing system 714, including any of the functionality described herein as being performed by a server or client, or other functionality. It will be appreciated that server system 700 and client computing system 714 are illustrative and that variations and modifications are possible. Computer systems used in connection with embodiments of the present disclosure can have other capabilities not specifically described here. Further, while server system 700 and client computing system 714 are described with reference to particular blocks, it is to be understood that these blocks are defined for convenience of description and are not intended to imply a particular physical arrangement of component parts. For instance, different blocks can be but need not be -137- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 located in the same facility, in the same server rack, or on the same motherboard. Further, the blocks need not correspond to physically distinct components. Blocks can be configured to perform various operations, e.g., by programming a processor or providing appropriate control circuitry, and various blocks might or might not be reconfigurable depending on how the initial configuration is obtained. Embodiments of the present disclosure can be realized in a variety of apparatus including electronic devices implemented using any combination of circuitry and software. F. Antimicrobial Peptides of The Present Technology, Compositions Thereof, and Methods of Use In an aspect, provided herein is an antimicrobial peptide consisting of an amino acid sequence of any peptide disclosed herein, or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. In any embodiment herein, the antimicrobial peptide may consist of an amino acid sequence of any one of SEQ ID NOs: 10-11, 13-15, 17- 46, 51-168, and 173-747, or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. In any embodiment herein, the antimicrobial peptide may consist of an amino acid sequence of any one of SEQ ID NOs: 17-46, 51-168, and 173-747, or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. In any embodiment herein, the antimicrobial peptide may consist of an amino acid sequence of KWKLFKTTAKFLHLAK (SEQ ID NO: 14) (“KK-15”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof, FLHWVPLRRVV (SEQ ID NO: 15) (“FV-11”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof, VDWKKVFGKLLKL (SEQ ID NO: 16) (“VL-13”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof, or LGKLLKKIPKFLHLVNK (SEQ ID NO: 387) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. In any embodiment herein, the antimicrobial peptide may consist of an amino acid sequence of VDWKKVFGKLLKL (SEQ ID NO: 16) (“VL-13”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof. The antimicrobial peptide of the present technology is also alternatively referred to herein as “a peptide of the present technology,” -138- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 “the peptide of the present technology,” “the peptide,” and the like. The antimicrobial peptide of the present technology may include one or more D-amino acids as well as one or more L-amino acids. In any embodiment herein, the antimicrobial peptide may consist of only D-amino acids, or alternatively in any embodiment herein the antimicrobial peptide may consist only of L-amino acids. An antimicrobial peptide of the present technology may be synthesized by any technique known to those of skill in the art and by methods as disclosed herein. Methods for synthesizing the disclosed peptides may include chemical synthesis of proteins or peptides, the expression of peptides through standard molecular biological techniques, and / or the isolation of proteins or peptides from natural sources. The disclosed antimicrobial peptide thus synthesized may be subject to further chemical and / or enzymatic modification. Various methods for commercial preparations of peptides and polypeptides are known to those of skill in the art. An antimicrobial peptide of the present technology may alternatively be made by recombinant means or by cleavage from a longer polypeptide. The particular composition of an antimicrobial peptide may be confirmed by amino acid analysis or sequencing. Compositions In an aspect, a composition is provided that includes an antimicrobial peptide of any embodiment disclosed herein, a pharmaceutically acceptable carrier or one or more excipients, fillers or agents (collectively referred to hereafter as “pharmaceutically acceptable carrier” unless otherwise indicated and / or specified). In a related aspect, a medicament for treating peri-implant disease is provided that includes an antimicrobial peptide of any embodiment disclosed herein and optionally a pharmaceutically acceptable carrier. In a related aspect, a medicament for controlling bacterial colonization on a dental implant is provided that includes an antimicrobial peptide of any embodiment disclosed herein and optionally a pharmaceutically acceptable carrier. In a related aspect, a medicament for controlling biofilm formation on a dental implant is provided that includes an antimicrobial peptide of any embodiment disclosed herein and optionally a pharmaceutically acceptable carrier. In a related aspect, a pharmaceutical composition is provided that includes an -139- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 effective amount of an antimicrobial peptide of any embodiment disclosed herein as well as a pharmaceutically acceptable carrier. For ease of reference, the compositions, medicaments, and pharmaceutical compositions of the present technology may collectively be referred to herein as “compositions.” In further related aspects, the present technology provides methods and uses that include an antimicrobial peptide of any aspect or embodiment disclosed herein and / or a composition of any embodiment disclosed herein as well as uses thereof. “Effective amount” refers to the amount of a compound (e.g., an antimicrobial peptide of the present technology) or composition required to produce a desired effect. One example of an effective amount includes amounts or dosages that yield acceptable toxicity and bioavailability levels for therapeutic (pharmaceutical) use including, but not limited to, treating peri-implant disease, controlling bacterial colonization on a dental implant, and / or controlling biofilm formation on a dental implant. In any aspect or embodiment disclosed herein (collectively referred to herein as “any embodiment herein,” “any embodiment disclosed herein,” or the like) of the compositions, pharmaceutical compositions, and methods including an antimicrobial peptide of the present technology, the effective amount may be an amount effective in treating peri-implant disease, controlling bacterial colonization on a dental implant, and / or controlling biofilm formation on a dental implant. By way of example, the effective amount of any embodiment herein including an antimicrobial peptide of the present technology may be from about 0.01 μg to about 200 mg of the peptide (such as from about 0.1 μg to about 50 mg of the peptide or about 10 μg to about 20 mg of the peptide). The methods and uses according to the present technology may include an effective amount of an antimicrobial peptide of any embodiment disclosed herein. In any aspect or embodiment disclosed herein, the effective amount may be determined in relation to a subject and / or in relation to a dental implant. The term “subject” and “patient” can be used interchangeably. Thus, the present technology provides pharmaceutical compositions and medicaments including an antimicrobial peptide of any embodiment disclosed herein (or a composition of any embodiment disclosed herein) and a pharmaceutically acceptable carrier. The compositions may be used in the methods and treatments described herein. The pharmaceutical composition may be packaged in unit dosage form. The unit dosage form is -140- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 effective in treating peri-implant disease, controlling bacterial colonization on a dental implant, and / or controlling biofilm formation on a dental implant when administered to a subject in need thereof and / or administered to a dental implant. Generally, a unit dosage including a peptide of the present technology will vary depending on patient considerations. Such considerations include, for example, age, protocol, condition, sex, extent of disease, contraindications, concomitant therapies and the like. Further, a unit dosage including a peptide of the present technology may vary depending on the dental implant considerations, such as the titanium surface area of the dental implant. An exemplary unit dosage based on these considerations may also be adjusted or modified by a physician skilled in the art. Suitable unit dosage forms, include, but are not limited to oral solutions, powders, lozenges, topical varnishes, lipid complexes, liquids, etc. The pharmaceutical compositions and medicaments may be prepared by mixing an antimicrobial peptide of the present technology with one or more pharmaceutically acceptable carriers, excipients, binders, diluents or the like. Such compositions can be in the form of, for example, powders, syrup, emulsions, suspensions or solutions. The instant compositions can be formulated for various routes of administration, for example, by intraoral administration or via administration (e.g., application) to a dental implant external to a patient. The following dosage forms are given by way of example and should not be construed as limiting the instant present technology. For intraoral administration, powders and suspensions are acceptable as solid dosage forms. These can be prepared, for example, by mixing an antimicrobial peptide of the instant present technology with at least one additive such as a starch or other additive. Suitable additives are sucrose, lactose, cellulose sugar, mannitol, maltitol, dextran, starch, agar, alginates, chitins, chitosans, pectins, tragacanth gum, gum arabic, gelatins, collagens, casein, albumin, synthetic or semi-synthetic polymers or glycerides. Optionally, oral dosage forms can contain other ingredients to aid in administration, such as an inactive diluent, or lubricants such as magnesium stearate, or preservatives such as paraben or sorbic acid, or anti-oxidants such as ascorbic acid, tocopherol or cysteine, a disintegrating agent, binders, thickeners, buffers, sweeteners, flavoring agents and / or perfuming agents. -141- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Liquid dosage forms for oral administration (e.g., intraoral administration) may be in the form of pharmaceutically acceptable emulsions, syrups, suspensions, or solutions, which may contain an inactive diluent, such as water. Pharmaceutical formulations and medicaments may be prepared as liquid suspensions or solutions using a sterile liquid, such as, but not limited to, an oil, water, an alcohol, and combinations of these. Pharmaceutically suitable surfactants, suspending agents, emulsifying agents, may be added for oral administration. As noted above, suspensions may include oils. Such oils include, but are not limited to, peanut oil, sesame oil, cottonseed oil, corn oil and olive oil. Suspension preparation may also contain esters of fatty acids such as ethyl oleate, isopropyl myristate, fatty acid glycerides and acetylated fatty acid glycerides. Suspension formulations may include alcohols, such as, but not limited to, ethanol, isopropyl alcohol, hexadecyl alcohol, glycerol and propylene glycol. Ethers, such as but not limited to, poly(ethyleneglycol), petroleum hydrocarbons such as mineral oil and petrolatum; and water may also be used in suspension formulations. The pharmaceutical formulation and / or medicament may be a powder suitable for reconstitution with an appropriate solution as described above. Examples of these include, but are not limited to, freeze dried, rotary dried or spray dried powders, amorphous powders, granules, precipitates, or particulates. The formulations may optionally contain stabilizers, antimicrobial agents, antioxidants, pH modifiers, surfactants, bioavailability modifiers and combinations of these. The carriers and stabilizers vary with the requirements of the particular composition, but typically include nonionic surfactants (Tweens, Pluronics, or polyethylene glycol), innocuous proteins like serum albumin, sorbitan esters, oleic acid, lecithin, amino acids such as glycine, buffers, salts, sugars, or sugar alcohols. Powders and sprays can be prepared, for example, with excipients such as lactose, talc, silicic acid, aluminum hydroxide, calcium silicates and polyamide powder, or mixtures of these substances. Ointments, pastes, creams and gels may also contain excipients such as animal and vegetable fats, oils, waxes, paraffins, starch, tragacanth, cellulose derivatives, polyethylene glycols, silicones, bentonites, silicic acid, talc and zinc oxide, or mixtures thereof. -142- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 Besides those representative dosage forms described above, pharmaceutically acceptable excipients and carriers are generally known to those skilled in the art and are thus included in the instant present technology. Such excipients and carriers are described, for example, in “Remingtons Pharmaceutical Sciences” Mack Pub. Co., New Jersey (1991), and “Remington: The Science and Practice of Pharmacy,” 20thEdition, Editor: Alfonso R Gennaro, Lippincott, Williams & Wilkins, Baltimore (2000), each of which is incorporated herein by reference. Methods Disclosed herein, in one aspect, is an method for slowing or halting the progression of peri-implant disease by applying an antimicrobial peptide of the present technology to a dental implant, e.g., to produce a film on the dental implant. This film can be applied in two minutes and can be repeated at follow up appointments. The renewable effects of the antimicrobial peptide upon successive reapplication was evaluated on bacteria-fouled and -cleaned dental implant surfaces, mimicking the re-treatment of implants affected by peri-implant disease in a dental office. This non-surgical approach can improve oral health by controlling microbial dysbiogenesis and reducing peri-implant disease progression. In another aspect, provided herein are methods of treating peri-implant disease in a subject in need thereof, the methods comprising, consisting essentially of, or consisting of administering an effective amount of an antimicrobial peptide of the present technology or a composition of the present technology to a dental implant in the subject. In another aspect, provided herein are methods of treating peri-implantitis in a subject in need thereof, the methods comprising, consisting essentially of, or consisting of administering an effective amount of an antimicrobial peptide of the present technology or a composition of the present technology to a dental implant in the subject. In another aspect, provided herein are methods of controlling bacterial colonization on a dental implant in a subject in need thereof, the methods comprising, consisting essentially of, or consisting of administering an effective amount of an -143- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 antimicrobial peptide of the present technology or a composition of the present technology to a dental implant in the subject. In another aspect, provided herein are methods of controlling biofilm formation on a dental implant in a subject in need thereof, the methods comprising, consisting essentially of, or consisting of administering an effective amount of an antimicrobial peptide of the present technology or a composition of the present technology to a dental implant in the subject. While the disclosure has been described with respect to specific embodiments, one skilled in the art will recognize that numerous modifications are possible. Embodiments of the disclosure can be realized using a variety of computer systems and communication technologies including but not limited to the specific examples described herein. Embodiments of the present disclosure can be realized using any combination of dedicated components and / or programmable processors and / or other programmable devices. The various processes described herein can be implemented on the same processor or different processors in any combination. Where components are described as being configured to perform certain operations, such configuration can be accomplished, e.g., by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, or any combination thereof. Further, while the embodiments described above may make reference to specific hardware and software components, those skilled in the art will appreciate that different combinations of hardware and / or software components may also be used and that particular operations described as being implemented in hardware might also be implemented in software or vice versa. Computer programs incorporating various features of the present disclosure may be encoded and stored on various computer-readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or DVD (digital versatile disk), flash memory, and other non-transitory media. Computer-readable media encoded with the program code may be packaged with a compatible electronic device, -144- 4917-2850-9996.7 Atty. Dkt. No.: 104434-0340 or the program code may be provided separately from electronic devices (e.g., via Internet download or as a separately packaged computer-readable storage medium). Thus, although the disclosure has been described with respect to specific embodiments, it will be appreciated that the disclosure is intended to cover all modifications and equivalents within the scope of the following claims. . -145- 4917-2850-9996.7
Claims
Atty. Dkt. No.: 104434-0340 WHAT IS CLAIMED IS:
1. A method of using machine learning (ML) models to generate rule sets to discriminate sequence data, comprising: retrieving, by one or more processors, a training dataset comprising a plurality of AMP sequences and a plurality of non-AMP sequences, wherein the plurality of AMP sequences targets select microbial populations; generating, by the one or more processors, for each of the plurality of AMP sequences and of the plurality of non-AMP sequences of the training dataset, a respective plurality of properties comprising at least one of (i) a hydropathy or (ii) a microbial inhibitory activity for at least one microbial species of the select microbial populations; providing, by the one or more processors, as input to a ML model, the plurality of AMP sequences, the plurality of non-AMP sequences, and the respective plurality of properties for each of the plurality of AMP sequences and the plurality of non-AMP sequences; determining, by the one or more processors, based on providing the input to the ML model, a set of rules defining values for the respective plurality of properties to discriminate between the plurality of AMP sequences and the plurality of non-AMP sequences; applying, by the one or more processors, the set of rules to a plurality of candidate AMP sequences to identify a subset of AMP sequences that satisfy the values defined by the set of rules for the plurality of AMP sequences; storing, by the one or more processors, using one or more data structures, the subset of candidate AMP sequences.
2. The method of claim 1, further comprising: synthesizing, using a peptide synthesizer, an AMP corresponding to a candidate AMP sequence of the subset of AMP sequences; testing the AMP in a microbial model including the select microbial population to determine a score indicating degree of efficacy of the AMP against the select microbial populations; and -146- 4917-2850-9996.7Atty. Dkt. No.: 104434-0340 identifying, by the one or more processors, the AMP as effective or ineffective in targeting the select microbial populations based on a comparison of the score with a threshold.
3. The method of claim 2, wherein the microbial model is one of a single-species microbial model or a poly-species microbial model, or wherein a concentration of the AMP is at one of a minimum inhibitory concentration (MIC) or minimum bactericidal concentration (MBC).
4. The method of claim 2, further comprising: adding, by the one or more processors, the candidate AMP sequence to the plurality of AMP sequences to generate a second plurality of AMP sequences for the training dataset, responsive to identifying the candidate AMP as effective; generating, by the one or more processors, for the candidate AMP sequence of the training dataset, a second plurality of properties; providing, by the one or more processors, as input to the ML model: (i) the second plurality of AMP sequences including the candidate AMP sequence, (ii) the plurality of non- AMP sequences, (iii) the respective plurality of properties for each of the plurality of AMP sequences and the plurality of non-AMP sequences, and (iv) the second plurality of properties; determining, by the one or more processors, based on providing the input to the ML model, a second set of rules defining values for the respective plurality of properties to discriminate between the second plurality of AMP sequences and the plurality of non-AMP sequences.
5. The method of claim 2, further comprising: receiving, by the one or more processors, for a subject at risk of or diagnosed with peri-implantitis or periodontitis associated with dental tissue about a dental implant, an indication of a presence of the select microbial populations in the tissue of the subject; and providing, by the one or more processors, an output identifying the AMP sequence to target the select microbial populations, responsive to identifying the AMP as effective, -147- 4917-2850-9996.7Atty. Dkt. No.: 104434-0340 wherein the dental tissue about the dental implant of the subject is administered with a therapy including the AMP to address the peri-implantitis or periodontitis.
6. The method of claim 2, further comprising refraining, by the one or more processors, from adding the candidate AMP sequence to the plurality of AMP sequences of the training dataset, responsive to identifying the AMP as ineffective.
7. The method of claim 1, further comprising generating, by the one or more processors, the plurality of candidate AMP sequences using a second plurality of AMP sequences identified as targeting the select microbial populations.
8. The method of claim 7, further comprising training, by the one or more processors, using the set of rules, a second ML model to generate the second plurality of AMP sequences that satisfy the values defined by the set of rules.
9. The method of claim 7, wherein generating the plurality of candidate AMP sequences further comprises: determining, for each of the second plurality of AMP sequences, a respective metric indicating a degree of similarity of a respective AMP sequence with at least one of a third plurality of AMP sequences, wherein the third plurality of AMP sequences targets the select microbial populations; and selecting, from the second plurality of AMP sequences, the plurality of candidate AMP sequences based on the respective metric.
10. The method of claim 1, further comprising updating, by the one or more processors, the ML model based on checking the set of rules on a second plurality of AMP sequences corresponding to a plurality of AMPs present in a commensal sample.
11. The method of claim 1, wherein determining the set of rules further comprises reducing, using an approximator, a number of boundary conditions corresponding to the values of the -148- 4917-2850-9996.7Atty. Dkt. No.: 104434-0340 set of rules to discriminate between the plurality of AMP sequences and the plurality of non- AMP sequences.
12. The method of claim 1, wherein generating the respective plurality of properties further comprises wherein generating the respective plurality of properties further comprises generating, for each of the plurality of AMP sequences and of the plurality of non-AMP sequences of the training dataset, the respective plurality of properties to include at least one of: (i) a matrix defining a plurality of pair distances for a respective peptide, each of the plurality of pair distances between a respective pair of residues in the respective peptide; (ii) a Fourier score determined based on the plurality of pair distances of the matrix; or (iii) a secondary structure identifying a spatial arrangement of a polypeptide backbone of the respective peptide.
13. The method of claim 1, wherein the select microbial populations are associated with peri- implantitis or periodontitis, and further comprise at least one of Porphyromonas gingivalis, Aggregatibacter actinomycetemcomitans, or Streptococcus gordonii.
14. A system for using machine learning (ML) models to generate rule sets to discriminate sequence data, comprising: one or more processors coupled with memory, configured to: retrieve a training dataset comprising a plurality of AMP sequences and a plurality of non-AMP sequences, wherein the plurality of AMP sequences targets select microbial populations; generate, for each of the plurality of AMP sequences and of the plurality of non-AMP sequences of the training dataset, a respective plurality of properties comprising at least one of (i) a hydropathy or (ii) a microbial inhibitory activity for at least one microbial species of the select microbial populations; -149- 4917-2850-9996.7Atty. Dkt. No.: 104434-0340 provide as input to a ML model, the plurality of AMP sequences, the plurality of non-AMP sequences, and the respective plurality of properties for each of the plurality of AMP sequences and the plurality of non-AMP sequences; determine, based on providing the input to the ML model, a set of rules defining values for the respective plurality of properties to discriminate between the plurality of AMP sequences and the plurality of non-AMP sequences; apply the set of rules to a plurality of candidate AMP sequences to identify a subset of AMP sequences that satisfy the values defined by the set of rules for the plurality of AMP sequences; store, using one or more data structures, the subset of candidate AMP sequences.
15. The system of claim 14, wherein an AMP corresponding to a candidate AMP sequence of the subset of AMP sequences is synthesized using a peptide synthesizer, and wherein the AMP is tested in a microbial model including the select microbial population to determine a score indicating degree of efficacy of the AMP against the select microbial populations, and wherein the one or more processors are further configured to identify the AMP as effective or ineffective in targeting the select microbial populations based on a comparison of the score with a threshold.
16. The system of claim 15, wherein the microbial model is one of a single-species microbial model or a poly-species microbial model, or wherein a concentration of the AMP is at one of a minimum inhibitory concentration (MIC) or minimum bactericidal concentration (MBC).
17. The system of claim 15, wherein the one or more processors are further configured to: add the candidate AMP sequence to the plurality of AMP sequences to generate a second plurality of AMP sequences for the training dataset, responsive to identifying the candidate AMP as effective; generate, for the candidate AMP sequence of the training dataset, a second plurality of properties; -150- 4917-2850-9996.7Atty. Dkt. No.: 104434-0340 provide, as input to the ML model: (i) the second plurality of AMP sequences including the candidate AMP sequence, (ii) the plurality of non-AMP sequences, (iii) the respective plurality of properties for each of the plurality of AMP sequences and the plurality of non-AMP sequences, and (iv) the second plurality of properties; determine, based on providing the input to the ML model, a second set of rules defining values for the respective plurality of properties to discriminate between the second plurality of AMP sequences and the plurality of non-AMP sequences.
18. The system of claim 15, wherein the one or more processors are further configured to: receive, for a subject at risk of or diagnosed with peri-implantitis or periodontitis associated with dental tissue about a dental implant, an indication of a presence of the select microbial populations in the tissue of the subject; and provide an output identifying the AMP sequence to target the select microbial populations, responsive to identifying the AMP as effective, wherein the dental tissue about the dental implant of the subject is administered with a therapy including the AMP to address the peri-implantitis or periodontitis.
19. The system of claim 15, wherein the one or more processors are further configured to refrain from adding the candidate AMP sequence to the plurality of AMP sequences of the training dataset, responsive to identifying the AMP as ineffective.
20. The system of claim 14, wherein the one or more processors are further configured to generate the plurality of candidate AMP sequences using a second plurality of AMP sequences identified as targeting the select microbial populations.
21. The system of claim 20, wherein the one or more processors are further configured to train, using the set of rules, a second ML model to generate the second plurality of AMP sequences that satisfy the values defined by the set of rules.
22. The system of claim 20, wherein the one or more processors are further configured to: -151- 4917-2850-9996.7Atty. Dkt. No.: 104434-0340 determine, for each of the second plurality of AMP sequences, a respective metric indicating a degree of similarity of a respective AMP sequence with at least one of a third plurality of AMP sequences, wherein the third plurality of AMP sequences targets the select microbial populations; and select, from the second plurality of AMP sequences, the plurality of candidate AMP sequences based on the respective metric.
23. The system of claim 14, wherein the one or more processors are further configured to update the ML model based on checking the set of rules on a second plurality of AMP sequences corresponding to a plurality of AMPs present in a commensal sample.
24. The system of claim 14, wherein the one or more processors are further configured to reduce, in accordance with an approximator, a number of boundary conditions corresponding to the values of the set of rules to discriminate between the plurality of AMP sequences and the plurality of non-AMP sequences.
25. The system of claim 14, wherein the one or more processors are further configured to generate the respective plurality of properties further comprises generating, for each of the plurality of AMP sequences and of the plurality of non-AMP sequences of the training dataset, the respective plurality of properties to include at least one of: (i) a matrix defining a plurality of pair distances for a respective peptide, each of the plurality of pair distances between a respective pair of residues in the respective peptide; (ii) a Fourier score determined based on the plurality of pair distances of the matrix; or (iii) a secondary structure identifying a spatial arrangement of a polypeptide backbone of the respective peptide.
26. The system of claim 14, wherein the select microbial populations are associated with peri- implantitis or periodontitis, and further comprise at least one of Porphyromonas gingivalis, Aggregatibacter actinomycetemcomitans, or Streptococcus gordonii. -152- 4917-2850-9996.7Atty. Dkt. No.: 104434-0340 27. An antimicrobial peptide of any peptide disclosed herein, or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof.
28. The antimicrobial peptide of claim 27, wherein the antimicrobial peptide is of an amino acid sequence of any one of SEQ ID NOs: 10-11, 13-15, 17-46, 51-168, and 173-747, or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof.
29. The antimicrobial peptide of claim 27 or claim 28, wherein the antimicrobial peptide is of an amino acid sequence of any one of SEQ ID NOs: 17-46, 51-168, and 173-747, or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof.
30. The antimicrobial peptide of any one of claims 27-29, wherein the antimicrobial peptide is of an amino acid sequence that is: KWKLFKTTAKFLHLAK (SEQ ID NO: 14) (“KK-15”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof, FLHWVPLRRVV (SEQ ID NO: 15) (“FV-11”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof, VDWKKVFGKLLKL (SEQ ID NO: 16) (“VL-13”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof, or LGKLLKKIPKFLHLVNK (SEQ ID NO: 387) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof.
31. The antimicrobial peptide of any one of claims 27-29, wherein the antimicrobial peptide is of an amino acid sequence that is VDWKKVFGKLLKL (SEQ ID NO: 16) (“VL-13”) or one or both of a pharmaceutically acceptable salt thereof and a solvate thereof.
32. A composition comprising an antimicrobial peptide of any one of claims 27-30 and a pharmaceutically acceptable carrier. -153- 4917-2850-9996.7Atty. Dkt. No.: 104434-0340 33. A method of treating peri-implant disease in a subject in need thereof, the method comprising administering an effective amount of an antimicrobial peptide of any one of claims 27-30 to a dental implant in the subject.
34. The method of claim 32, wherein the peri-implant disease is peri-implantitis.
35. A method of controlling bacterial colonization on a dental implant in a subject, the method comprising administering to the dental implant an effective amount of an antimicrobial peptide of any one of claims 27-30.
36. A method to control biofilm formation on a dental implant in a subject, the method comprising administering to the dental implant an effective amount of the peptide of an antimicrobial peptide of any one of claims 27-30. -154- 4917-2850-9996.7
Citation Information
Patent Citations
Filtering artificial intelligence designed molecules for laboratory testing
US20210366580A1
Artificial intelligence designed antimicrobial peptides
US20220009966A1
Cited By
Industrial energy-saving monitoring method based on Copula graph model
CN121980461A