A system, method, and device to evaluate drug potency in the treatment of animal pathogens

A system using advanced computational methods rapidly identifies effective antibiotics and predicts drug potency against pathogens, addressing the inefficiencies of current MIC determination methods and antimicrobial resistance challenges.

WO2025179225A1PCT designated stage Publication Date: 2025-08-28MACHINE TRANSLATION LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/016915
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-21
Filing Date
2025-02-21
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Current methods for determining Minimum Inhibitory Concentrations (MICs) of antibiotics take too long, leading to ineffective empirical therapies and increased morbidity and mortality due to antimicrobial resistance, with existing systems failing to provide prompt, accurate, and generalizable predictions for effective antimicrobial agents against pathogens, including those with incipient resistance.

Method used

A system utilizing a processor and non-transitory computer readable medium with executable instructions for determining drug potency, incorporating a sampler interface, pathogen characterizer, drug characterizer, and potency analyzer, leveraging large language models and genetic algorithms to rapidly identify effective drugs against pathogens, including those with predicted mutational derivatives.

Benefits of technology

The system significantly reduces the time to determine MICs, enabling clinicians to select effective antibiotics and drug developers to create new compounds that counter emerging resistance, providing rapid and generalizable predictions for various pathogens and potential mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025016915_28082025_PF_FP_ABST
    Figure US2025016915_28082025_PF_FP_ABST
Patent Text Reader

Abstract

A system includes a processor and a non-transitory computer readable medium that stores executable instructions for determining an expected potency of a drug in treating a non-human animal pathogen. The executable instructions include a sampler interface that receives sequence data representing the animal pathogen and generates a dataset comprising one of raw nucleic acid sequence data, ribonucleic acid sequence data, or amino acid sequence data. A pathogen characterizer generates a representation of the animal pathogen from the dataset. A potency analyzer receives the representation of the animal pathogen from the pathogen characterizer and a representation of the drug and outputs a value representing a potency of the drug as applied to the animal pathogen. The potency analyzer can be used with a pathogen simulator that generates a representation of a pathogen and a drug simulator that generates a representation of a drug to iteratively discover new pathogens and drugs.
Need to check novelty before this filing date? Find Prior Art

Description

A SYSTEM, METHOD, AND DEVICE TO EVALUATE DRUG POTENCY IN THE TREATMENT OF ANIMAL PATHOGENS RELATED APPLICATION

[0001] This application claims priority from U.S. Provisional Patent 63 / 556,114, filed February 21, 2024 and entitled “A SYSTEM, METHOD, AND DEVICE TO PREDICT AND COUNTER EVOLUTION OF ANTIMICROBIAL RESISTANCE.” This reference is hereby incorporated by reference in its entirety. FIELD OF THE INVENTION

[0002] This invention provides a method, a system, and a device useful in reducing the time for a clinician to select, or for a pharmaceutical developer to choose to produce, an effective drug to counter infection in a non-human animal caused by a pathogen. BACKGROUND OF THE INVENTION

[0003] When a patient enters a hospital suffering from a bacterial infection, a microbial sample is taken from the patient in order to assist a physician in prescribing an empirical antimicrobial therapy, which is prescribed based on site of infection (e.g. nasal, lungs, urinary tract, etc.) and probable species causing the infection. However, antimicrobial resistance is increasingly found to arise, meaning, these empirical therapies are becoming less and less effective over time against known pathogens. A measure of resistance of a bacterial isolate to a given antimicrobial is termed the Minimum Inhibitory Concentrations (MIC), that is, the lowest concentration of an antibiotic candidate that prevents growth of the microbes in the sample taken from the patient. Unfortunately, a major problem in the art is that MICs typically take between 24 to 72 hours to determine using known culturing methods, see Figures 1 and 2.

[0004] Due to the decreasing effectiveness of the empirical therapies, and the time to determine MICs, the patient is at risk of morbidity and mortality until a suitably effective antibiotic is identified. The degree of this problem can be understood from predictions in the literature that bacterial infections will be the leading cause of death by 2050 due to antibiotic-resistant microbial infections.

[0005] An attempt to resist resistance is to decrease the creation and spread of antimicrobial-resistant mechanisms. Stopping the spread of resistance through “antibiotic stewardship” is already the focus of physicians worldwide. However, this only slows the spread of antibiotic resistant strains. Bacteria are always evolving, leading to increased resistance, and modern medicine has found itself reacting to, rather than getting ahead of the incipient resistance. Morbidity and mortality could be reduced by determining MICs more quickly, and by implementing predictive technology, as described and claimed herein below, to get ahead of incipient resistance, by rapidly identifying compounds with sufficiently low MIC’s to be effective against even multidrug resistant pathogens, and against their predicted mutational derivatives. SUMMARY OF THE INVENTION

[0006] In one example, a system includes a processor and a non-transitory computer readable medium that stores executable instructions for determining an expected potency of a drug in treating a non-human animal pathogen. The executable instructions include a sampler interface that receives sequence data representing the non-human animal pathogen and generates a dataset comprising one of raw nucleic acid sequence data, ribonucleic acid sequence data, or amino acid sequence data. A pathogen characterizer generates a representation of the non-human animal pathogen from the dataset. A potency analyzer receives the representation of the non-human animal pathogen from the pathogen characterizer and a representation of the drug and outputs a value representing a potency of the drug as applied to the non-human animal pathogen.

[0007] The system can further include a drug characterizer that receives one of an identity or different representation of the drug and generates the representation of the drug provided to the potency analyzer. In one example, the drug characterizer is implemented using a large language model, such as the MAMBA architecture, that is trained on a sets of sample data each including a SMILES representation of the drug and a representation of the drug appropriate for analysis at the potency analyzer. In another example, the drug characterizer is implemented using a graph neural network that is trained on a sets of sample data each including an input representing a molecular structure of the therapeutic as a graph and a representation of the drug appropriate for analysis at the potency analyzer.

[0008] In one example, the non-human animal pathogen is a bacterium and the drug is an antibiotic. In another example, the non-human animal pathogen is a fungus and the drug is an antifungal. In a further example, the non-human animal pathogen is a virus and the drug is an antiviral.

[0009] In one implementation, the sampler interface receives sequences representing raw DNA and generates DNA contigs from the sequences representing raw DNA. In one example, the sampler interface generates the DNA contigs using reference-guided construction. Additionally or alternatively, the sampler interface can include a quality control component, implemented as a machine learning model trained on a set of samples each including a set of raw DNA and a label representing the quality of DNA contigs generated from the set of raw DNA, that determines if the raw DNA meets a threshold quality for generating DNA contigs. In one example, the quality control component is implemented using a large language model, such as the MAMBA architecture.

[0010] In one implementation, the pathogen characterizer generates a count of unique K- mers from the dataset, normalizes the K-mer counts according to a total count, and generates a matrix including the K-mer counts. In another implementation, the pathogen characterizer is implemented using a large language model, such as the MAMBA architecture, that is trained on a sets of sample data each including a set of DNA or protein data as well as an appropriate representation of the non-human animal pathogen based on the set of DNA or protein data.

[0011] The system can further include a species predictor determines a species of the non-human animal pathogen according to the data provided from the sampler interface. In one example, the system further includes a pathogen predictor implemented as a machine learning model trained on a set of training samples each including a set of raw DNA sequences from a bodily fluid or tissue of an animal, a type of the bodily fluid or tissue, a location of an infection, and one of a plurality of pathogen classes, wherein the pathogen classes include a first class representing bacteria, a second class representing viruses, a third class representing fungi, and a fourth class representing no detected pathogen, the species predictor determining the species of the non-human animal pathogen according to the data provided from the sampler interface and a determine class of the non-human animal pathogen at the pathogen predictor.

[0012] In one implementation, the potency analyzer includes a pathogen validity model that determines if the pathogen represented by the representation of the non-human animal pathogen could represent a living organism and a drug validity model that determines if the drug represented by the representation of the drug could represent aviable molecule, the potency analyzer outputting the value representing a potency of the drug, a value representing a validity of the representation of the pathogen, and a value representing a validity of the drug. In one example, the potency analyzer includes at least one attention layer that provides a set of importance values indicating regions of the representation of the non-human animal pathogen and the representation of the drug that are important in determining the potency of the drug as applied to the non-human animal pathogen.

[0013] In another example, a system includes a processor and a non-transitory computer readable medium that stores executable instructions for generating novel drugs and pathogens. The executable instructions include a non-human animal pathogen simulator that generates a representation of a pathogen for analysis. A drug simulator that generates a representation of a drug for analysis. A potency analyzer that determines a potency for the drug as applied to the non-human animal pathogen.

[0014] In one example, each of the pathogen simulator and the drug simulator are implemented as genetic algorithms that operate adversarially, such that the pathogen simulator iteratively selects for a pathogen that minimizes the potency, and the therapeutic simulator iteratively selects for a therapeutic that maximizes the potency. In one example, the potency analyzer includes a pathogen validity model that determines if the pathogen represented by the representation of the non-human animal pathogen could represent a living organism and a drug validity model that determines if the drug represented by the representation of the drug could represent a viable molecule, the potency analyzer outputting the value representing a potency of the drug, a value representing a validity of the representation of the pathogen, and a value representing a validity of the drug. Additionally or alternatively, the potency analyzer can include at least one attention layer that provides a set of importance values indicating regions of the representation of the non-human animal pathogen and the representation of the drug that are important in determining the potency of the drug as applied to the non-human animal pathogen.

[0015] In one example, the pathogen simulator is implemented as a mutation learning model that provides a pathogen that provides target potency for a given drug, the pathogen simulator receiving a dataset representing an initial pathogen, a species of the initial pathogen, a representation of the therapeutic, and a target potency and providing the pathogen as a mutation dataset, comprising one of raw DNA, RNA, or protein sequences, representing of the pathogen that provides the target potency. Additionally or alternatively, the drug simulator can be implemented as a drug tuning model that providesa drug that provides target potency for a given pathogen, the drug simulator receives a representation of an initial drug, a representation of the pathogen, and a target potency and provides the drug as a modified version of the initial drug that provides the target potency.

[0016] In one example, the system includes a mutation evaluator that estimates the time necessary for a pathogen to mutate from a first DNA sequence to a second, mutated DNA sequence.

[0017] In one example, the non-human animal pathogen is a bacterium and the drug is an antibiotic. In another example, the non-human animal pathogen is a fungus and the drug is an antifungal. In a further example, the non-human animal pathogen is a virus and the drug is an antiviral. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 provides a representation of Broth Microdilution Panel Schematics - Increasing dilutions of the antimicrobial agent are displayed at the top of the panel. The bacterial growth is represented as darker wells while lack of growth is represented as lighter wells.

[0019] Figure 2 shows determination by standard methods, the Minimum Inhibitory Concentration, MIC, which is the lowest concentration of an antimicrobial agent that can inhibit bacterial growth.

[0020] Figure 3 provides a diagrammatic representation of elements which form one embodiment of the invention described and enabled herein, comprising functional modules, system components, devices, subsystems, algorithms, methods, and routines which, in integrated combination as described herein, provides A SYSTEM, METHOD AND DEVICE TO PREDICT AND COUNTER EVOLUTION OF ANTIMICROBIAL RESISTANCE, as claimed.

[0021] Figure 4 shows the origins of a dataset, which may be considered an input to, and therefore a part of the invention, when sequence data as shown is used according to the system, method and device shown in Figure 3.

[0022] Figure 5 shows a diagrammatic representation of a BioInformatics Module (BIM) utilized according to the to the system, method and device shown in Figure 3.

[0023] Figure 6 shows a diagrammatic representation of a Machine Learning Bacterial Processing (MLBP) module utilized according to the to the system, method and device shown in Figure 3.

[0024] Figure 7 shows a diagrammatic representation of a MIC Prediction Pipeline (MPP) module utilized according to the to the system, method and device shown in Figure 3.

[0025] Figure 8 shows a diagrammatic representation of a Machine Learning Antibiotic Processing (MLAP) module utilized according to the to the system, method and device shown in Figure 3.

[0026] Figure 9 shows a detailed diagrammatic representation of a Bacterial Adversarial Genetic Algorithm (BAAGA) module utilized according to the to the system, method and device according to this invention, as shown in Figure 3. This figure provides visualization of what happens for each iteration of the AGA and BGA components within BAAGA. Seeding the simulation consists of an isolate’s contigs and an antimicrobial agent’s SMILES string. From there, AGA runs 150 iterations (1 segment) with each iteration going through Crossover, Mutation, having MICs predicted for new individuals using the RNN, then Selection. After 1 segment, the best SMILES string goes to the BGA which then runs its own 150 iterations. After the BGA finishes 1 segment, the best fingerprint passes to the AGA to start a new round (1 segment run by each GA).

[0027] Figure 10 Provides four random fingerprints from the dataset as examples of what isolate fingerprints are.

[0028] Figure 11 Visualization of two antimicrobial agent structures and their respective canonical SMILES strings. The Aztreonam and Cefepime structures and canonical SMILES were collected from PubChem.

[0029] Figure 12 Data flow diagram through processing and prediction. The diagram also includes the representation of the RNN.

[0030] Figure 13 Architecture diagram of the RNN used.

[0031] Figure 14 Visualizing F1-micro scores for all cross-validation folds.

[0032] Figure 15 Log scale histogram showing the number of samples per new antimicrobial agent per MIC bin for the four generalized antibiotics. This MIC dataset is only for testing the generalization performance, and is not used during training.

[0033] Figure 16 Comparing all generalization performance metric results for four new antimicrobial agents. The zero random SMILES represent training on thecanonical SMILES string for the antimicrobial agent. The five random SMILES represent having up to six possible SMILES to train on for each antimicrobial agent. One canonical SMILES and one to five unique, random SMILES. Ten random SMILES have the same definition. One canonical SMILES and one to ten unique, random SMILES.

[0034] Figure 17 Simulation results for Aztreonam. Each plot in the figure represents a segment which is 150 iterations. The top left, segment 0, is the first segment run of the simulation. The bottom right, segment 5, is the last run of the simulation.

[0035] Figure 18 Simulation results for Ceftriaxone. Each plot in the figure represents a segment which is 150 iterations. The top left, segment 0, is the first segment run of the simulation. The bottom right, segment 5, is the last run of the simulation.

[0036] Figure 19 Simulation results for Meropenem. Each plot in the figure represents a segment which is 150 iterations. The top left, segment 0, is the first segment run of the simulation. The bottom right, segment 5, is the last run of the simulation. These results are from the second simulation run.

[0037] Figure 20 shows algorithms 1-3 in Figures 20A-20C respectively.

[0038] Figure 21 shows algorithms 4 and 5 in Figures 21A-B respectively.

[0039] Figure 22 shows, in Figure 22A, a box plot that shows timing results for all 1,000 runs for each test of the different algorithms and the respective language(s) on which they ran; Figure 22B shows, in tabular form, average timing and memory results for running 1,000 runs for each algorithm and the language(s) the algorithm was run in.

[0040] Figure 23 illustrates a system for determining a metric representing an expected potency of a drug in treating a pathogen.

[0041] Figure 24 illustrates a system for identifying novel drugs and pathogens DETAILED DISCLOSURE OF THE PREFERRED EMBODIMENTS ACCORDING TO THE INVENTION

[0042] Definitions

[0043] A nucleotide consists of a phosphate group linked by a phosphoester bond to a pentose (ribose in RNA, and deoxyribose in DNA) that is linked in turn to an organic base. The monomeric units of a nucleic acid are nucleotides. Naturally occurring DNA and RNA each contain four different nucleotides: nucleotides having adenine,guanine, cytosine and thymine bases are found in naturally occurring DNA, and nucleotides having adenine, guanine, cytosine and uracil bases are found in naturally occurring RNA. The bases adenine, guanine, cytosine, thymine, and uracil often are abbreviated A, G, C, T and U, respectively.

[0044] A polynucleotide, as used herein, may mean any molecule including a plurality of nucleotides, including but not limited to DNA or RNA. Preferably, the polynucleotide includes at least 5 nucleotides, and more preferably it includes 10 or more nucleotides. The depiction of a single strand also defines the sequence of the complementary strand. Thus, a nucleic acid also encompasses the complementary strand of a depicted single strand. A polynucleotide may be single stranded or double stranded, or may contain portions of both double stranded and single stranded sequence. Double stranded polynucleotides are a sequence and its complementary sequence that are associated with one another, as understood by those skilled in the art. The polynucleotide may be DNA, both genomic and cDNA, RNA, or a hybrid, where the nucleic acid may contain combinations of deoxyribo- and ribo-nucleotides, and combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine and isoguanine. Polynucleotides may be obtained by chemical synthesis methods or by recombinant methods. When a polynucleotide has been defined as consisting of either DNA or RNA, it may be referred to as a DNA strand, or RNA strand, respectively.

[0045] k-mers are substrings of length k contained within a nucleotide sequence. Primarily used within the context of computational genomics and sequence analysis, in which k-mers are composed of nucleotides (i.e. A, T, G, and C). The term k-mer refers to all of a sequence's subsequences of length k, such that the sequence AGAT would have four monomers (A, G, A, and T), three 2-mers (AG, GA, AT), two 3-mers (AGA and GAT) and one 4-mer (AGAT).

[0046] A “pathogen”, as used herein, refers to an organism or multicellular structure that causes disease, and is explicitly intended to include disease-causing bacteria, viruses, and fungi.

[0047] A “drug” as used herein, refers to a substance introduced to a human, non- human animal, or plant to treat a pathogen causing disease in the human, non-human animal, or plant. The term “drug” is explicitly intended to include at least antibiotics, antivirals, bacteriophages, antifungals, and anticancer agents.

[0048] Antibiotics, as defined herein, are bactericidal or bacteriostatic compounds already known in the art. Examples of known antibiotics include agents that target the bacterial cell wall, such as penicillins, cephalosporins, agents that target the cell membrane such as polymixins, agents that interfere with essential bacterial enzymes, such as quinolones and sulfonamides, and agents that that target protein synthesis such as the aminoglycosides, macrolides and tetracyclines. Additional known antibiotics include cyclic lipopeptides, glycylcyclines, and oxazolidinones. Antibiotic resistance represents the ability of intracellular pathogens to decrease (i.e., resist) the cytotoxic and cytostatic effects of antibiotics.

[0049] Pathogenic bacteria are harmful bacteria, typically as a result of their ability to cause an infection having harmful symptoms in a subject. Examples of pathogenic bacteria include Mycobacterium tuberculosis, Escherichia coli, Vibrio cholerae, Strepthococcus pneumoniae, and Staphylococcus aureus.

[0050] Those skilled in the art are aware of a corpus of published information and attempts at producing various methods, systems, and devices with the intent of minimizing the time required to obtain reliable MIC information for a given antibiotic, also referred to herein as an antimicrobial agent, and potentially antibiotic- resistant pathogens, also referred to as antimicrobial resistant microbes or pathogens.

[0051] To date, no known system, method or device has been developed which adequately meets the clinical needs for prompt, accurate, reliable and add generalizable presentation to clinicians, drug development researchers, or the like, of a set of prescribing or drug development options in real time or as close to real time as possible, where the prescribing and development options are weighted by likelihood of being effective against a given microbe, including those with incipient antimicrobial or multi-drug resistance. Attempts at solving this need have not been generalizable, and so have been applicable only to a few antibiotics or species. The present technology is generalizable to essentially any species or compound. That is, an antimicrobial agent which is predicted to have a low MIC, notwithstanding evolution of mutants predicted to evolve antimicrobial resistance. That is, the method, system and device of this invention further predicts antimicrobial agents, known or generated by the method, system or device, with low predicted MIC’s for such mutants predicted to evolve antimicrobial resistance.

[0052] Accordingly, this invention provides a method, system, and device useful in reducing the time for a clinician to select an antibiotic to contain an infection in apatient. Properly executed, the present invention further provides additional information, beyond MIC predictions, to allow physicians and drug developers to make a long-term therapy and drug development decisions, based on predictions of incipient or emergent drug resistant microbes and the chemical entities which might best be used to contain such pathogens.

[0053] To be further described herein below, embodiments of the invention include a generalizable RNN which achieves state-of-the-art performance, with an F1 score of 0.82 on trained agents, with respect to antimicrobials and MICs.

[0054] The combination of steps, elements, systems, and components, including simulations, as disclosed herein, is able to identify and predict new resistance mechanisms. The BAAGA module is, and the system, method and mechanism incorporating BAAGA is, furthermore, able to generate new, testable, antimicrobial agents which may be synthesized and confirmed to exhibit the desired MICs. Preferably, these embodiments provide outputs which include all explanations (SMILES strings, MICs, and mutations) for graphic presentation and interpretation by trained clinicians, rather than being used to intake human data or provide clinical decisions. This invention provides the foundation for predicting future resistance and MICs. Experimental results are provided herein for bacterial species and antimicrobial agents experimented with as described herein. The more the RNN system, method and device according to this invention is trained on many more agents and species, the more the robustness, including accuracy, of various embodiments of the underlying invention will increase. More data is anticipated to permit optimization, for example, of fingerprint datasets by including more K-Mer sizes. Furthermore, disclosed herein are a series of examples showing the use and adaptation of the System, Method and Device according to this invention for identifying antimicrobially effective compounds where the microbe is a bacterium, a fungus, a virus, and where the antimicrobial is a standard antibiotic or a bacteriophage. Referring now to the figures, Figure 1 provides a representation of Broth Microdilution Panel Schematics - Increasing dilutions of the antimicrobial agent are displayed at the top of the panel. The bacterial growth is represented is darker while lack of growth is represented as lighter wells.

[0055] Figure 2 shows an MIC determination as the lowest concentration of an antimicrobial agent that can inhibit bacterial growth. These methods known in the artare slow to produce actionable information for clinicians, drug developers and the like, and such methods only provide a limited amount of information.

[0056] Referring now to 3, a simplified, diagrammatic representation of the multiple interconnected elements which form one embodiment 100 according to the present invention is described, enabled and claimed herein. This embodiment comprises functional modules, system components, devices, hardware, software, subsystems, algorithms, methods, and routines which, in integrated combination as described herein, provides the claimed systems. The key provided in this figure 3 is applicable to figures 3-9:

[0057] The overall architecture of one embodiment of the system, method, and device 100 according to this invention, which includes modular sub-systems, apparatuses, components, methods, and routines, wherein each module is described herein below. In this figure, a workflow starting 101 is shown in which, in the typical lab, hospital, researcher, pharmaceutical company or the like 102 using the present invention, isolates and produces a nucleic acid sequence for a pathogen 103. This workflow includes known procedures used in the art, including Whole Genome Sequencing (WGS), now taking a few minutes to a few hours using such advanced sequencing systems such as those available from Illumina. Absent this invention, the standard workflow includes and awaits culturing over several days, by which time the patient may have been misdiagnosed, or treated with an inefficient antimicrobial for the given infection, which may exhibit antimicrobial resistance or incipient antimicrobial resistance. Figure 4 provides further details on this workflow, including the annotation provided in Figure 3.

[0058] The workflow shown in Figures 3 and 4, 101-103, is conventional in the art of molecular biology, which includes standard nucleic acid sequencing, protein sequencing, transcription, translation, and the like. The workflow is depicted as occurring in a hospital, with a patient, but this workflow may begin anywhere in the world or, in outer space, provided the correct equipment is available. The important thing is for some form of workflow to be implemented to achieve the objectives of this invention, that is, to provide the starting information, namely the sequence data 103 which is the raw starting material of this invention.

[0059] The embodiment, referred to as MT herein, 100 according to Figure 3 initiates at 104, using the microbial sequence 103. From here, the workflow bifurcates via a first arm 105, wherein the sequence data is uploaded to the device, system, or utilizedin the method of the present invention, and from there to Bioinformatics Pipeline 200. The other arm 106 of the bifurcation utilizes user inputs of at least one or multiple antibiotics, which input is utilized by Machine Learning Antibiotic Processing module 400, for each antibiotic input.

[0060] Bioinformatics Pipeline 200, see Figure 3 and detail in Figure 5, utilizes its raw material input 103, via 104 and bifurcation at 105, to produce, as further described herein below, outputs 201 quality control to ensure use of sample or not, and if used, 202 the identification of the strain of the microbe, and 203 the taxonomy, including species, of the microbe. Further products of this module 200 include the microbe’s list of contigs which is input 204 to BAAGA module 500 (see Figure 3 and detail in Figure 9) and also 205 to a Machine Learning Bacterial processing module 300 (see Figure 6) which produces a bacterial fingerprint which, (see Figure 7 for detail), in a first arm, 301, is input to Species ML Model (RNN) 302 which outputs the species 303 of the microbe. In a second arm 304, the same bacterial fingerprint is input to MIC Prediction Pipeline 305, which identifies the MIC 306, by virtue of also receiving input 401 of sequentially provided antibiotic data from Machine Learning Antibiotic Processing module 400, see Figure 8 for detail.

[0061] With respect to the contigs, for example, unique 3-mers, 307, 4-mers, 308, 5- mers, etc, up to and including 8-mers, 309, are identified, counted to produce a count of k-mers 310, which, reconstructed in a matrix 311 provides the bacterial fingerprint output 304. Those skilled in the art will appreciate that, theoretically, the present system, method and device may utilize k-mers of, theoretically any length; computation power and need for speed of processing will guide those skilled in the art to balance accuracy, precision, costs, and timeline to obtain results with ever finer confidence limits. It is anticipated that there is a point of diminishing returns where the k-mer sizes and counts thereof slows computation time without providing much additional accuracy, precision, or predictive power, but, theoretically, there is no definitive limit to the k-mer size or counts used.

[0062] In this embodiment, the species 303 is identified, as noted above in describing Figure 3, via system components 301, and 302. The bacterial fingerprint 304 is also fed to MIC Prediction pipeline, 305, see Figures 3, 6 and 7 for detail, for each potential antibiotic. At 310 a Recursive Neural Network module, “RNN”, uses the fingerprint 304 to compare against antibiotic list 401 from Machine Learning Antibiotic processing module 400, such that if the entire list has not yet beencompared, 312, the process reiterates, until the list is completed 313 at which point the loop is halted 314 and the MICs have been output 306.

[0063] Within module 400, see Figure 8 for detail, a user 403 inputs list of “1+ Antibiotics” 404, which is, through further processing produces antibiotic SMILES 405, which are sent 406 to a SMILES one-hot matrix engine, 407 produces a SMILES one-hot matrix, which, when all recursions are complete, produces SMILES one-hot matrix 401 which is fed to MIC Prediction Pipeline 305, as described above.

[0064] It should be recognized that module 200 utilizes datasets which include data from microbial, including fungal, bacterial, viral, and other pathogens, preferably worldwide, according to the infection type (see Examples). Preferably, isolates are tested for antimicrobial susceptibility by standard reference broth microdilution methods (Figures 1 and 2) in response to the many antimicrobial agents available for clinical use. Preferably, the dataset includes the identification of species and taxonomy, with quality controls to ensure exclusion of spurious data inputs.

[0065] Product 204 from the Bioinformatics Pipeline 200, and antibiotic SMILES strings 402 from Machine Learning Antibiotic Processing module 400 are fed into the BAAGA engine 500 of this invention, see Figure 3 and detail in Figure 9 which produces the following non-limiting products: 501, a list of mutations and made; 502, a list of antimicrobial compounds made, which, when passed 503 to a ChemProp ML model 504 providers chemical properties for new compounds 505; a list or MIC values predicted 506; and a list of generated bacteria 507, which, using a phylogenetic tool provides a tree 508 from which is derived a phylogenetic tree of all microbes which appeared during the simulation 509.

[0066] Methods of carrying out mutation analysis of DNA are known by those skilled in the art. Mutations on germline or somatic DNA include mis-sense and nonsense mutations, SNPs (single nucleotide polymorphisms) deletions and insertions. In some embodiments, the mutation analysis comprises denaturing gradient gel electrophoresis, high-resolution melting curve analysis, and directed Sanger sequencing. Other types of mutation analysis include single-strand conformation polymorphism (SSCP) analysis, SSCP / heteroduplex analysis, enzyme mismatch cleavage, allele-specific hybridization, and restriction analysis of the genomic DNA.

[0067] Chemical entities, their predicted MICs, and potential New Chemical Entities (NCE’s), and their predicted MIC’s for the microbe under investigation and potentially antimicrobially resistant mutants thereof are presented as products of thedevice, method and system 100 according to this invention. The clinician and drug developers are thereby empowered to, preferably in real time or as near to real time as possible, make as informed as possible prescribing decisions. This invention incorporates various embodiments of the following elements such that a novel, inventive and useful device, system and method is disclosed herein. “Machine Learning” (“ML”) or “Artificial Intelligence” (“AI”) when referred to herein, is used synonymously, is a tool to enhance the speed and thus the practicality of carrying out the inventive embodiments of this invention.

[0068] Those skilled in the art will appreciate, from the present disclosure, that equivalents to various components, elements, methods, routines, subroutines, modules, models, simulations, and the like disclosed herein, may be substituted for the specifics taught in the present disclosure and exemplary support provided herewith. Technology is developing extremely quickly in this art area, and as methods such as Whole Genome Sequencing, WGS, Whole Transcriptome Sequencing, WTS, and the like progress, and as various methods and hardware are developed to optimize the speed at which the present system, apparatus and method operates, the more quickly and more accurately the present invention may operate, permitting higher iterations of predictions and greater data sets to be transformed, as described herein, into useful, actionable, clinical and drug development decision support. The novelty, inventiveness, utility and enablement, including written description, is to be determined by reference to the appended claims.

[0069] For prediction of future antimicrobial resistance emergence and prediction of compounds likely to be effective notwithstanding emergence of antimicrobial resistance, see Figures 3 and 9, which depicts an MIC Machine Learning (“ML”) module 500 comprising, preferably, an MIC (RNN) module 501, which runs simulations to provides outputs, 501-509, as described above in connection with Figure 3.

[0070] The detail provided for component 500 provides a depiction of a simulation engine which, when provided with the product, 402 from MIC / antibiotic module 400, and product 204 from the Bioinformatics Pipeline module 200, initiates a recursion between a first Antibiotic Genetic Algorithm (AGA) module, 510, and a second Bacterial Genetic Algorithm (BGA) module 520. AGA 510 operates through cyclic iterations, 511, comprising the steps and components for the system and device to execute the steps of selection, 512, crossover, 513, mutation, 514, and prediction ofnew MIC’s 515. The best SMILES string 527 is saved and passed to a Bacterial Genetic Algorithm (BGA) module 520, which through cyclic iterations, 521, comprising the steps and components for the system and device to execute the steps of selection, 522, crossover, 523, mutation, 524, and prediction of new MIC’s 525. At each iteration of the simulation, the best fingerprint 526 is passed back to the AGA engine 510 and best SMILES string 527 is passed to BGA, 520. Upon completion of a defined number of iterations 530, say 150 iterations, the module 500 generates Individuals SMILES string and saved. Best fingerprint, on the AGA 510 side of the system, back to 501, while, on the BGA side of the module 520, likewise, provides 532 the saved best SMILES string and the saved best individual’s fingerprint, back to 501 where the MIC ML Model (RNN), provides predictions on MICs.

[0071] Fitness (MIC prediction) 533 and 534 is built into the model as noted above in connection with Figure 3, the output 501-509 from the BAAGA device 500 includes a list of: mutations and when they occurred 502; antibiotic compounds made 503 which at 504 is identified, e.g. via ChemProp, to provide predicted chemical properties 505 for new compound; MIC values 506 predicted throughout the simulation; a list of generated bacteria 507 which, via a phylogenetic tree production system 508 produces a phylogenetic tree 509 of all bacteria that appeared during a given simulation.

[0072] The functioning of this system, the method, and the device according to this invention is evidenced in Figures 10-19 and the Examples provided herein below. Figure 10 provides four random fingerprints from the dataset as examples of what isolate fingerprints are, while Figure 11 provides visualization of two antimicrobial agent structures and their respective canonical SMILES strings. The Aztreonam and Cefepime structures and canonical SMILES were collected from PubChem. Figure 12 provides a data flow diagram through processing and prediction. The diagram also includes a representation of the RNN as a box, while figure 13 provides an enabling architecture diagram of the RNN used in the Figure 12 process, which those skilled in the art are able to practice based on this teaching. Figure 14 provides visualization of the F1-micro scores for all cross-validation folds, and Figure 15 provides a log scale histogram showing the number of samples per new antimicrobial agent per MIC bin for the four generalized antibiotics. This MIC dataset is only for testing the generalization performance, and is not used during training.

[0073] Figure 16 provides a comparison of all generalization performance metric results for four new antimicrobial agents. The zero random SMILES represent training on the canonical SMILES string for the antimicrobial agent. The five random SMILES represent having up to six possible SMILES to train on for each antimicrobial agent. One canonical SMILES and one to five unique, random SMILES. Ten random SMILES have the same definition. One canonical SMILES and one to ten unique, random SMILES.

[0074] In light of the foregoing disclosure, one of ordinary skill in the art is enabled to make and use various embodiments of the present invention. Applications for this invention include: addressing the need by clinicians to promptly select an antimicrobially effective antimicrobial agent, as well as addressing the need for clinicians to be able to determine long-term therapy or even prophylactic treatments, based on predictions mutations in bacterial species, and predictions of chemical entities, known or New Chemical Entities (NCE’s), to contain further anti-microbial resistance before or as it emerges. These pro-active capabilities are anticipated to become ever more important in the arts of medical treatment, and in the arts of development of NCE’s by pharmaceutical companies, chemists and the like to supplement Structure Activity Relationship (SAR) research.

[0075] In the present case, one of ordinary skill in the art of Machine Learning alone, or in the art of bacterial biology, alone, potentially experience challenges in comprehending and making practical use of the invention disclosed and claimed herein. Accordingly, as a starting point, one of ordinary skill in art would need to have such skills in at least two quite divergent disciplines. Further, as described herein the parameters used in the architecture, i.e. values used in the layers, e.g. in a Dropout layer, as well as for BAAGA, absent the present disclosure, would require significant experimentation.

[0076] Accordingly, what is taught herein goes well beyond ordinary skill in the art, in that it is taught herein how to build and use a two-core system, device, and method, comprising both an MIC ML module and BAAGA module. The operative architecture for ML module is taught herein and it is also taught how to collect datasets for use in the two-core system. It is further taught herein how to use such datasets to fine-tune the operative parameters of the two-core technology described herein. Furthermore, it is taught herein how to train the two-core invention to achieve precise, reliable, and generalizable predictive MIC data and the active antimicrobials associated with suchMIC predictions. It is further taught herein an engine, BAAGA, along with probabilities used in the bacterial side of the simulation. It is further taught herein how to implement these elements in a form amenable to machine learning, requiring significant understanding, again, of diverse fields of human endeavor. It is further taught herein how to build two Genetic Algorithms, (GA’s), one for bacteria and one for antibiotics, as described herein. Further, it is taught herein the processing portions of the system, method, and device, which includes a workflow of: SMILES -> one- hot matrix, Contigs -> K-Mer frequency matrix / Fingerprint, as disclosed and claimed herein. The fingerprint processing as disclosed herein has the potential to produce results five thousand times (5,000 x) faster than conventional processing known in the art because the system takes out a lot of work needed in both the middle and end of fingerprint processing. It also streamlines processing in a way that allows for a computer to take advantage of parallel computation on a GPU. To make that more explicit, conventional means to generate a fingerprint is to slowly count K-Mers through each size. For example, one would first count all 3-Mers, then restart and count all 4-Mers from scratch, then restart and count all 5-Mers from scratch, etc. According to the present system, method and device, a specific 3-Mer is first counted, then a nucleotide character is added to count the 4-Mer, then again for 5-Mer, etc. Essentially, all K-Mer sizes are counted at the same time rather than one-by-one. The end portion speedup is due to the fact that the present system is built to slowly generate the fingerprint while running. Conventional code will just be counting K- Mers for each k size. Then, at the end, the conventional means would have to build the fingerprint with all the counts. This would be somewhat timely due to the data structure that conventional means uses to count K-Mers. By parallel computation, counting all K-Mer sizes at once, and not needing to build the fingerprint at the end, a speedup factor of up to 5,000X is achieved. JellyFish, for example, is one of the conventional K-Mer counting tools used in the art and, as compared to the present method, system and device. For further details and support of this element of the invention, those skilled in the art may refer to Kromer-Edwards, Cory, "Optimizing K-Mer Fingerprint Generation for Machine Learning." Proceedings of the 14th ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics, September 2023, Article No.: 101, Pages 1–5 , where this speed-up factor is confirmed as follows: “With the increasing availability of genomic data obtained through Whole-Genome Sequencing (WGS), Machine Learning (ML) algorithms arebeing used to analyze this data. However, processing large datasets or files poses challenges. One approach is to count K-Mers, which has been used in ML studies. However, larger K-Mer sizes may lead to decreased accuracy and training difficulties. Alternatively, combining multiple K-Mers of smaller sizes into fingerprints has shown promise in predicting species and antibiotic resistance. This study compares existing fingerprint generation techniques with a new algorithm called GPU K-Mer Fingerprinting (GKF), which utilizes a GPU for parallel processing. GKF demonstrates similar memory utilization compared to other approaches but achieves a speedup of 5,546X.”

[0077] One skilled in the art is taught by the present disclosure how to understand chemical compounds enough to know how to work with SMILES. It is further taught herein how to fine-tune the Genetic Algorithms to produce meaningful outputs.

[0078] Those skilled in the art know how to assemble bacterial sequences, build general pipelines, use bioinformatics tools, and read and interpret QC values, as disclosed herein. In reviewing the foregoing disclosure and the experiments disclosed herein below, it should be borne in mind the non-trivial accomplishment disclosed herein of a generalizable system, method and device comprising processing from Contigs -> K-Mer frequency matrix / Fingerprint, which requires undue experimentation and inventive activity to develop the operative invention as disclosed herein.

[0079] Having described the system, method and device according to this invention herein above, those skilled in the art will appreciate that the present invention includes a device comprising (a)at least one memory module; and (b) at least one data processing module. The at least one memory module (a) is selected to be adequate to receive at least one dataset comprising nucleic acid sequence data for at least one microbial isolate. The data processing module comprises at least one Recursive Neural Network (RNN) adequate to utilize said nucleic acid sequence data to produce (i) contiguous sequences, taxonomy, and species; (ii) K-mers, counts thereof, or both, which, in matrix format, provides a fingerprint for said at least one microbial isolate adequate to provide a definitive identification of said at least one microbial species; and (iii) an MIC prediction module adequate to provide, for any given antibiotic and microbial isolate, predicted MICs. Preferably, this device includes at least one simulation module the output of which predicts either or both potential future antibiotic-resistant phenotypes, and potential future antibiotics along with theirpredicted MICs for control of potential future antibiotic resistant phenotypes, or both. This is achieved herein by providing a Bacterial Antibiotic Adversarial Genetic Algorithm (BAAGA) module comprising (a) at least one Bacterial Genetic Algorithm (BGA) unit; and (b) at least one an Antibiotic Genetic Algorithm (AGA). As described herein, these modules are set to interface with each other such that, recursively, it may be used as a tool to minimize the time to determine Minimum Inhibitory Concentrations (MICs) for potential selection by a clinician as an antimicrobial agent to control a pathogen in a subject, or to predict an antimicrobial agent adequate to control a variant of a pathogen. The BAAGA module recursively selects antibiotics and microbes, provides crossover and mutations to predict MICs for likely future antimicrobial resistance and compounds likely to be able to counter said future antimicrobial resistance. In certain embodiments a cording to this invention, tools such as the known Kraken tool is used to determine taxonomy, including species or the known MLST tool is employed to determine strain, including species, or both such tools or their equivalents may be employed.

[0080] The system according to this invention includes, whether assembled in a single location or in distributed locations or whether controlled by a single entity or not, adequate machine hardware, software, and interconnections between the various modules described immediately above for the device. Similarly, those skilled in the art will appreciate that to practice this invention, one skilled in the art is tasked with providing a device or system as described herein above, including providing and connecting each of its modules and components into a functioning whole, as shown in Figure 3, for example, for one embodiment according to this invention. In one preferred embodiment of such a system, device and method according to this invention, includes at least one bioinformatics module, at least one Machine Learning for bacterial processing module, at least one machine learning for antibiotic processing module, needed for each candidate antibiotic, and a Bacterial Antibiotic Adversarial Genetic Algorithm (BAAGA) engine, which is adapted to produce time- sensitive, accurate, and informative outputs from the system, subsystems, devices, subroutines, methods, integrated with machine learning, upon which clinicians and drug developers may rely to make more informed, prompt and safe decisions when faced with antimicrobial resistant infectious microbes, and emergent antimicrobially resistant infections, epidemics, or pandemics.

[0081] In further preferred embodiments, the bioinformatic module includes, utilization of microbial nucleic acid sequence data to assemble non-overlapping, contiguous nucleic acid sequences, such that, via further processing the Taxonomy (species) and strain are identified, optionally with inbuilt Quality Control confirming reliability of data upon which these identifications are made; contigs are introduced to at least one fingerprinting block wherein unique k-mers are identified, counted to produce a count of k-mers, which, reconstructed in a matrix provides a bacterial fingerprint. Similarly, for MIC prediction for each potential antibiotic, the system, method or device includes at least one RNN module, which uses bacterial fingerprints to produce a definitive species identification.

[0082] In further preferred embodiments, the bacterial fingerprint is fed to at least one MIC prediction module which includes a RNN module, which compares against an antibiotic list until the list is completed to provide MIC predictions.

[0083] In further preferred embodiments, the MIC predictions module referred to herein as the BAAGA module, is included, an MIC machine learning module comprising, an RNN which provides outputs including: a list of: mutations and when they occurred, antibiotic compounds made, to provide predicted chemical properties for new compounds, MIC values predicted throughout the simulation, a list of generated bacteria which, via a phylogenetic tree production systems produces a phylogenetic tree of all bacteria that appeared during a given simulation.

[0084] EXAMPLES

[0085] The present invention has been described in both broad terms and in some detail herein above. To ensure that an adequate written description is provided herein to enable those skilled in the art to make and use embodiments according to this invention, the following detailed exemplary support is provided. It will be appreciated, however, that the specific details provided in this exemplary support are not to be interpreted as limiting on the present invention. Rather, for a scope of the present invention, reference should be had to the claims as herein appended, including equivalents thereof.

[0086] In the Examples that follow, various considerations are described so that one skilled in the art may better understand what is herein disclosed and claimed. Based on the following examples, those skilled in the art receive further guidance as to therelevant considerations for properly implementing and obtaining the benefits of the present invention.

[0087] The Exemplary support which follows may be best understood with the following introductory materials exploring relevant considerations for each aspect of the present invention. This introductory material is provided here to ensure that one of ordinary skill in the art reading the present disclosure is made familiar with various distinct developments in various and divergent art areas, and that absent the present teach, one of ordinary skill in the art would first have to become an expert in all of the art areas discussed herein below, and then would need to understand how to assemble all of the elements described herein to obtain the operative invention disclosed and claimed herein. This is a complex art area, and the present teachings should not be used in seeking to construct the present invention as though it existed independently from the present disclosure and claims. Furthermore, what follows may not be utilized as an admission that any of the art discussed herein below is effective “prior art” with respect to the subject matter to which the present claims are directed. Further, it is the very “threading of the needle” through all of the art and experiments described herein below, which elevates the present disclosure to an advancement in the art of antimicrobial resistance containment and treatment.

[0088] Bacterial infections are first treated with an empirical agent. Hospitals choose an empirical therapy based on the site of infection and probable bacterial species causing the infection. This is because physicians have no laboratory results to guide definitive therapy. Empirical agents are usually broad-spectrum and can lead to Antimicrobial Resistance (AMR). Broad-spectrum refers to an agent that is effective against multiple bacterial species. Physicians rely on Antimicrobial Susceptibility Testing (AST) results to choose the definitive therapy. The chosen therapy may be a more potent agent. This would be the case if the infection is resistant to the empirical antibiotics. The chosen therapy could also be a narrow-spectrum agent. This would be the case if the broad-spectrum empirical antibiotic is not necessary. A narrow- spectrum antibiotic is an agent that targets a specific, or few, species. Thus, the chosen agent would decrease the impact on AMR while still being effective. Individual hospitals create guidelines on what empirical therapy to use for each scenario

[0042] . Studies have demonstrated that delaying the application of appropriate antibiotics increases morbidity and mortality [42, 70]. For these reasons, AST results are a cornerstone of the treatment of bacterial infections. Lab and hospital staff useMinimum Inhibitory Concentration (MIC) values to determine AST results. The longer it takes to choose the proper antimicrobial agent, the higher chance a patient will not have a successful outcome

[0070] .

[0089] Authors in Bassetti et al.

[0018] predict that, by 2050, bacterial infections will be the leading cause of death, overtaking cancer and other diseases. Timing is critical for individual patients. Every hour of inadequate antimicrobial therapy can mean higher mortality rates. For these reasons, researchers have been trying to identify a faster, more efficient means of determining MICs.

[0090] One possibility is using Machine Learning (ML) to predict MICs based on genotypic data from the bacterial isolates. DNA and RNA sequencing methods are becoming faster and cheaper. It is now possible to have sequencing results within a few hours

[0091] . A computer system can feed these sequences through an assembly pipeline and get results in less than an hour. Finally, the computer system can send the results from the pipeline through a ML model to predict MICs. If the model is generalizable, the model could predict many antimicrobial agents within minutes. One step after creating a generalizable model could be to create a simulation. The simulation could generate new resistance mechanisms and new antimicrobial agents. At that point, physicians and researchers could act proactive rather than reactive to antibiotic resistance. Physicians in hospitals could plan for future resistance mechanisms and adjust empirical antibiotics. Researchers could plan new antimicrobial agents based on future resistance patterns. The following will discuss the challenges of predicting MICs in more depth. Alongside that, many bacterial datasets and the ML algorithms that could use those datasets are discussed.

[0091] Discussion of antimicrobial resistance must be complete before moving too far into the data or Machine Learning (ML). This chapter will discuss what antimicrobial resistance is and how to measure it through Minimum Inhibitory Concentrations (MIC). MICs are cumbersome to determine and how they are interpreted to categorize isolates as susceptible or resistant is subject to error. Antimicrobial mechanisms and how they impact resistance are also discussed.

[0092] Measuring resistance by MIC

[0093] The reference method for MIC determination is the broth microdilution procedure described by the Clinical and Laboratory Standards Institute (CLSI)

[0112] .This procedure is a multistep process performed in a 96-well round bottom panel as shown in Figure 2. The broth microdilution method is an in vitro assay where a bacterial isolate is grown in increasing concentrations of an antimicrobial agent. The MIC is the lowest concentration of an antimicrobial agent that prevents visible growth of an isolate

[0113] after incubation for 24 to 72 hours at established temperatures. The antimicrobial concentrations are tested in 2-fold dilutions and are expressed in mg / L

[0080] . Figure 3 shows the schematics on how to determine an MIC. The process of making the panels with multiple antimicrobial agents is demanding and clinical laboratories are not well equipped to perform this task. Setting up and reading the MIC results requires time and trained personnel.

[0094] Variability in determining MIC

[0095] Furthermore, Antimicrobial Susceptibility Testing (AST) has reproducibility issues and variations can be noticed within the same laboratory. The reasons for these variations range from differences, in preparing the panels, media used in the panels, the temperature that panels are left to allow growth, and incubation time. There are also variations in the reading of the results by the staff in the laboratories and hospitals as some staff may have different levels of training / technical knowledge

[0080] .

[0096] Clinical microbiology laboratories inside hospitals do not often perform the broth microdilution procedure [98, 80]. Rather, they utilize commercial systems that are validated against the reference broth microdilution method. This choice is mainly due to the easy-to-use nature of commercial systems that do not require visual reading and interpretation. These systems have also evolved to, in certain cases, be slightly faster to generate results than the reference method with incubation periods for MIC determination of 6 to 10 hours [98, 80].

[0097] Acceptance criteria of MICs

[0098] Due to the difficulty and variability in determining MICs, the FDA

[0039] and CLSI

[0112] have advised that determining an MIC within a ±12-fold dilution is considered acceptable. This recommendation means that if the MIC should be 1.0 mg / L, a human / machine could determine 0.5 mg / L or 2.0 mg / L, and the MIC would still be accepted.

[0099] Interpretation criteria for MIC values

[0100] AST results require interpretation criteria to indicate to the physician if the bacterial isolate causing the infection will be able to be treated with a given agent. The main interpretation criteria used for AST are named breakpoints.

[0101] Resistance to antimicrobials

[0102] AMR can be caused by the acquisition of foreign genes or the selection of isolates that are advantageous in the environment where an antimicrobial agent is present

[0075] . The acquisition of foreign genes can occur by the transfer of mobile genetic elements such as plasmids or transposons or in certain cases by the accumulation of foreign single-stranded DNA in a process called natural transformation.

[0103] Plasmids are small, independent, circular pieces of DNA that can be replicated and expressed separately from the main chromosome in a bacterial isolate

[0035] . These plasmids can be passed between bacteria, through conjugation, which is how ARGs can be spread

[0035] . Another type of gene is intrinsic. Intrinsic genes are found on the bacterial chromosome and mutations in these genes, or their regulatory genes, can lead to traits that allow the bacteria to survive in environments containing antimicrobial agents.

[0104] Mutations in intrinsic genes and acquired genes can be detected using Whole Genome Sequencing (WGS) using DNA samples; however, mutations in regulatory genes that cause differential expressions of genes leading to AMR can be detected by measuring the RNA levels in the cell using a transcriptome analysis or RNA- sequencing (RNA-seq)

[0037] .

[0105] Phenotypes of Antimicrobial Mechanisms

[0106] There are three types of phenotypes for ARGs which consist of target alteration of the antimicrobial, enzyme inactivation of the antimicrobial, and permeability of outer-membrane in Gram-negative bacteria

[0022] .

[0107] Antimicrobials generally seek to target specific molecules within bacterial pathogens. Because of this specific targeting, minor mutations in those molecules can confer resistance to the respective antimicrobial. An example of target modification asa mechanism of resistance is alterations in cell wall precursors that confer resistance to glycopeptide antibiotics

[0022] . Resistance through target alteration is variable, though, as it requires that the modified cell can still perform its normal function. An example of a low level of conferred resistance is with mutations in penicillin-binding proteins (PBP) of S. pneumoniae

[0044] . However, an example of conferring a very high level of resistance is VanA-type vancomycin resistance for vancomycin in enterococci [9].

[0108] Enzyme inactivation mechanisms seek to make the antimicrobial compound inactive by modification. Many of these enzymes are acquired, for example, β- lactamases, and Aminoglycoside modifying enzymes (AME). However, some enzymes are intrinsic to certain genera. An example of one of a genus with intrinsic antibiotic-modifying enzymes is Enterobacterales [9]. When these enzymes are modified, high levels of resistance can be conferred to antibiotics. One such example is the expression of TEM-1 β-lactamase in E. coli. This TEM-1 enzyme can increase the ampicillin MIC from 4 mg / L to >128 mg / L

[0021] .

[0109] As time passes, bacteria gain more resistance mechanisms to various antimicrobial classes, allowing them to become Multi-Drug Resistant (MDR)

[0031] . One common way a bacterium can become MDR is through a mutation in an efflux system gene, that can expel out multiple antibiotic classes, or a gene encoding for an outer-membrane porin protein (OMP), which can prevent penetration of multiple antibiotic classes

[0094] . More specifically, the function of an efflux pump is responsible for extruding toxins, dyes, and antibiotics from the cell

[0022] . Efflux pumps are found in both Gram-positive and Gram-negative bacteria. OMPs are found in Gramnegative bacteria and are responsible for allowing material inside the outer cell wall

[0022] .

[0110] One final note on resistance is that, occasionally, resistance cannot be attributed to a single mechanism

[0022] . One such example is P. aeruginosa conferring resistance to imipenem through the downgrading of an OMP gene named OprD and production of AmpC β-lactamase

[0074] . In this case, both mechanisms must be present for clinically significant levels of resistance.

[0111] MOLECULAR DATASETS AND DATA COLLECTION

[0112] To establish a reliable Machine Learning (ML) algorithm, the input data is of utmost importance. Collection of genotypic data used for ML and data collection are discussed.

[0113] WGS

[0114] Whole-Genome Sequencing (WGS) is the basis for many ML algorithm datasets. Reference guided assembly

[0073] can guide the collection of genes of interest. Chosen genes for this dataset are often genes that confer resistance to specific antimicrobial agents. An example is β-lactamase genes that confer resistance to β- lactam agents

[0024] .

[0115] Short-Read

[0116] For short-read sequences, Illumina sequencing techniques and equivalents thereof are commercially available

[0102] . To start, a biologist preps a DNA library by splitting up the DNA into many denatured fragments. These fragments are then attached to a flow-cell panel. The DNA fragments are then amplified using PCR to fill the panel. Then, the biologist adds special terminator nucleotides with a fluorescent dye and blocking group to the cell. Once the proper nucleotide is added, light emits and other nucleotides are blocked. The Illumina machine picks up the light by an optical device and the wavelength of the light emitted determines which base was added. Finally, the blocking agent and fluorescent dye are removed. This allows the next nucleotide to be added

[0023] .

[0117] Illumina sequencing for short-reads can produce read lengths of at most 300 bp long. Each run can contain at most 2,000 Mb due to the clustering nature of Illumina sequencing. Output from sequencing is many fragmented DNA sequences with quality control scores in FASTQ file format. The equation below determines the quality control score, given by Q. In the equation, e represents the estimated probability of the base call being wrong.

[0118] Q = −10 log10(e)

[0119] K-Mers Dataset

[0120] K-Mers are unique substrings of length k made up of either nucleotides or amino acids. For example, taking a small piece of DNA, ATCGATT, and finding all K-Mers of size 2 would lead to a list of [AT, TC, CG, TT]. Another step in this dataset is to go back through and count the number of each K-Mer in the set of sequences. In the example, that would give [AT: 2, TC: 1, CG: 1, TT: 1]. Each K-Mer then gets a unique ID: [0: 2, 1: 1, 2: 1, 3: 1]. That final list of ids and counts is the input to ML algorithms.

[0121] Single Nucleotide Polymorphisms Dataset

[0122] The gene sequence-based dataset mentioned above can also form another dataset. When performing reference-guided assembly a list of Single Nucleotide Polymorphisms (SNPs) can be exported. A SNP shows a change in a single position of a sequence based on what was in that position in the reference sequence. These SNPs can then merge into a list for each gene and passed into the ML algorithm. These lists are the input rather than the DNA or protein sequences themselves.

[0123] Expressions Dataset

[0124] Another way resistance can emerge is if benign genes can cause resistance

[0024] . An example is a gene dictating how much material to push out of the cell. This gene is on the outer-membrane and is an efflux pump. For the efflux pump, over- expression may expel the antimicrobial compounds before a reaction can occur. This conclusion led to the gene expression dataset, which may be utilized in certain circumstances utilizing the present invention. When creating the gene expression dataset, the amount of RNA produced by the gene in question is measured. One way to check if these expression levels confer resistance is to compare the expression for one isolate against the expression from a known resistant isolate. Comparing expression levels gives a fold increase or decrease in expression. However, if resistant expression measurements are not known, ML may predict resistance.

[0125] After discussing the different forms of data, how to use that data in Machine Learning (ML) algorithms is next addressed. The main takeaway is that some of the algorithms discussed below have multiple applications for Antimicrobial Resistance(AMR). The specific approach to apply to the algorithms is based on the specific problem that needs to be answered.

[0126] There are two ways to use K-Mers as input. One approach would be to take the list of unique KMers and feed them one by one into the model to get a list of output predictions for each K-Mer. The predictions would then combine into one final prediction. This approach is not used in practice as it would be difficult to determine how to combine the predictions. The difficulty is that each K-Mer could be the cause of resistance. A second approach is to have a distinct number of K-Mers. Each isolate has all K-Mers counted. Then, there would be an input neuron for each K-Mer with the input being the counts for those K-Mers.

[0127] Recursive Neural Networks (RNN)

[0128] A RNN is a type of NN that holds a state or memory at each layer

[0050] . This memory allows the RNN to be capable of learning time-series datasets. In the case of AMR, the only datasets that can use an RNN are DNA or protein sequences. Each sequence represents a time series as order matters. Each nucleotide or amino acid is a point in that series. A language, where each word is a single character, is another way to describe these sequences. In this way, Natural Language Processing (NLP) applies these sequences using RNNs.

[0129] XGBoost

[0130] XGBoost is another ensemble approach to the Decision Tree (“DT”) algorithm

[0027] . Unlike the random generation of DTs in RF, XGBoost creates each DT sequentially. After the generation of a DT, the XGBoost algorithm calculates a new loss using all previous DTs and the new one. Gradient descent uses that loss to help create the next DT. In this sense, all DTs are built in a way that the next DT helps the shortcomings of the previous DTs.

[0131] EXAMPLE 1

[0132] COMPOUND RNN TO PREDICT MICS USING K-MER FINGERPRINTS AND ANTIBIOTIC SMILES

[0133] Training a generalizable model for any antimicrobial agent is a significant contribution to the art. The system according to this invention achieved a 10-fold cross-validation test F1 score of 0.82 on 10 antimicrobial agents, and achieved a notable F1 score of 0.63 average on four new, unique antimicrobial agents.

[0134] To effectively use Machine Learning (ML) to predict Minimum Inhibitory Concentration (MIC), multiple antimicrobial agents need to be predictable. Currently, there have been two methods for predicting multiple agents. One method is to create multiple models with each model predicting a single agent [63, 57, 56, 15]. The other method creates a single model that also takes in an agent id to distinguish between agents [59, 85]. While both methods can predict MICs for a small list of antibiotics, neither can generalize to new antibiotics.

[0135] Whole Genome Sequencing (WGS)

[0037] data can be utilized as one input to ML. WGS can generate short-read sequences within one hour

[0091] . These sequences can then assemble into contigs within another hour. However, to generalize antimicrobial agents, the agent’s chemical structure must be passed in as another input. To accomplish this, each agent compound is converted into Simplified Molecular Input Line Entry System (SMILES) representations [6].

[0136] A novel way to generally predict MIC values by fingerprinting isolates and antimicrobial agent SMILES strings is disclosed herein. These fingerprints and SMILES were passed into a Recursive Neural Network (RNN) to predict MIC values for 10 different antimicrobial agents. Results indicate that the models described herein can achieve an average F1-micro score of 0.82 using 10-fold cross-validation. That performance is compared against other models that also use K-Mers as input. Additionally, generalization performance is shown by having the model predict MICs for four antimicrobial agents that were not used during training or testing. Generalization F1-macro scores were, on average, 0.63 for the four agents.

[0137] Background

[0138] Fingerprints were generated for each isolate as input. The process for creating each fingerprint is known in the art. SMILES strings encode structural information of a chemical into a text representation. SMILES strings have a set alphabet and grammar associated with them. This alphabet allows chemists the ability to look at aSMILES string and infer the chemical structure. Each character in the SMILES alphabet corresponds to either an atom, type of atom, bond, ring, or side chain. Most SMILES used in research are canonical SMILES

[0110] . A canonical SMILES string is a unique representation of a particular molecule. A canonicalization algorithm generates canonical SMILES strings. This algorithm utilizes rules to ensure every canonical SMILES string output is unique, given the same structure. Figure 5 shows two, of the twelve, antimicrobial agent structures. For comparison with the 2D structures, Equation 10.1 shows the canonical SMILES string for Aztreonam, and Equation 10.2 shows the canonical SMILES string for Cefepime. Both canonical SMILES strings were collected from PubChem [4, 2]. SMILES strings have been used in other ML problems in the past [10, 66, 51, 117, 41]. With the onset of Graph Neural Networks (GNN), studies have investigated if the graph representation of a molecule achieves higher performance and efficiency, rather than SMILES representations

[0051] . Results revealed that GNN models achieve the same, or worse, performance on low-complexity tasks. The lower threshold for performance was fewer than 50 classes

[0051] . Here, 19 possible MICs are predicted, which results in 19 classes. Because of this, canonical SMILES were used as input to the RNN architecture rather than a chemical structure into a GNN.

[0139] Figure 5: Visualization of two antimicrobial agent structures and their respective canonical SMILES strings. The Aztreonam and Cefepime structures and canonical SMILES were collected from PubChem [3, 1]

[0140] Isolate Data

[0141] Isolate data was collected. For the validation of this system, data includes contigs of DNA, held in FASTA files, for Klebsiella pneumoniae and Escherichia coli. These isolates were selected for WGS as the isolates obtained resistant MIC results to β-lactam and / or aminoglycoside. This led to 10 antimicrobial agents being utilized. In total, 19,860 isolates over years 2016, 2017, 2018, 2019, 2020, and 2021 were selected for WGS. One isolate per patient infection episode was collected to avoid duplication of isolates.173 hospitals worldwide used a common protocol to collect the isolates.

[0142] Data Processing

[0143] Contigs are processed similarly to [76, 53, 59]. Each contig had K-Mers of size three, four, five, six, seven, and eight counted. All contigs for an isolate had their K-Mer counts summed together to get an overall K-Mer count for the six sizes. Then,a universal scalar was applied to all K-Mer counts. This normalized all counts for each isolate to be between [-1.0, 1.0]. Finally, each isolate had the normalized counts combined into a matrix of size 6x65,536. Here, the six rows stood for each k size, and each column was a count. K-Mer sizes three, four, five, six, and seven were padded to the right with 0's to reach 65,536. The RNN used these final matrices as input. These matrices are named fingerprints moving forward. Figure 6 shows the steps to process the contigs before input into the RNN, providing a data flow diagram through processing and prediction. The diagram also includes the representation of the RNN. K-Mer normalization is not shown here for simplification. For simplification, the diagram shows processing fingerprints utilizing k sizes of three, four, and five.

[0144] By way of example, one approach that those skilled in the art may implement to address this aspect of the present system is create code using binary operations to slowly count the number of K-Mers in order rather than one at a time. By distinction, those skilled in the art have adopted a workflow of generating a fingerprint by counting all 3-Mers, then counting all 4-Mers, etc. That approach is slow because it requires rebuilding the K-Mer string and searching through all of the DNA for every K-Mer. Rather than doing that, an improvement implemented according to this invention includes the following workflow:

[0145] Start, for example, at a given position, say nucleotide 42 in a given DNA sequence;

[0146] Check if the 3-Mer starting at position 42 is AAA (first 3-Mer) and count it if so;

[0147] Check if the 4-Mer starting at position 42 is AAAA (first 4-Mer) and count if so;

[0148] Check if the 5-Mer starting at position 42 is AAAAA (first 5-Mer) and count if so etc. for 6, 7, and 8;

[0149] Check the 3-Mer starting at position 42 is AAC (second 3-Mer) and count it if so;

[0150] Check the 4-Mer starting at position 42 is AAAC (second 4-Mer) and count it if so etc. for all K-Mers of 3-8…

[0151] After all K-Mers for 3-8 are counted at position 42, move to position 43, and so forth.

[0152] Moving between an X-Mer and a Y-Mer (where X+1=Y) is 3 bit operations. So, rather than counting through all the DNA thousands of times, this improvement counts through the entire sequence once, while carrying out hundred-bit operations for each position in the sequence. Those skilled in the art will appreciate that, in this way, the workflow described above becomes amenable to GPU parallelization, to achieve counting of K-mers, per the above described workflow, on multiple GPU cores. This combination of data processing structure, that is the workflow, with hardware structure, results in approximately a 5000X speedup in processing and development of the relevant fingerprint(s). A further improvement that those skilled in the art may implement to achieve this level of computation acceleration is to require the concurrent construction of the fingerprint as the parallelized hardware configuration conducts the above described workflow. This process confers significant operational gains to known methods, configurations, systems or devices which built up lists of counts and then had to build the fingerprint , separately, and in series, rather than in parallel. The table shown in Figure 22 B, demonstrates actual results achieved in operating this element of the system, method and device according to this invention as described herein, for average timing results for running 1,000 runs for each algorithm, and the language(s) in which the algorithm was run.

[0153] FIG.23 illustrates a system 2300 for determining a metric representing an expected potency of a drug in treating a pathogen. The system 2300 includes a processor 2302 and a non-transitory computer readable medium 2310 that stores executable instructions for determining an expected potency of a drug in treating a pathogen. The system further includes a sample evaluator 2304 that receives a sample representing a pathogen and generates a representation of one of the DNA, RNA, and protein sequence associated with the pathogen. The executable instructions can include a sampler interface 2312 that includes appropriate software components for communicating with the sample evaluator 2304 or a repository of stored pathogen representations (not shown) over a network via a network interface (not shown) or via a bus connection.

[0154] In some implementations, the sample evaluator 2304 can extract a sequence representing raw fragments of DNA and provide it to the sampler interface 2312, for example, as a Fasta file. In one example, this raw DNA data provided directly to a sample characterizer 2314. In another example, the sample evaluator 2304 can include tools for generating sequences of contiguous DNA from the raw DNA data,referred to as contigs. In one example, the sampler interface 2312 can include a tool for trimming low quality data from the edges of the raw DNA samples, an assembler for joining the DNA, and one or more error correction programs that align the raw DNA reads with the generated contigs, sort and index the alignments, and correct the reads using those alignments. It will be appreciated that a quality of the contigs can be checked after the initial assembly and after error correction to ensure that the assembly is performed correctly and that the assembled contigs provide sufficient coverage of the raw DNA fragments. In another example, the sampler interface performs a reference-guided construction of DNA contigs. In the reference-guided process, the raw DNA is provided along with an identification of the specific species of the pathogen, which is generally determined via lab work prior to the process. Appropriate reference sequences can be retrieved from a database representing chromosome DNA of the species as well as fully assembled sequences for plasmids for the determined species. These can be retrieved from a database or stored locally, for example, after multiple high-quality sequences have been generated for a specific species of pathogen. The raw reads are then aligned to the reference to assemble the main chromosome and any plasmids into a continuous sequence. The assembled contigs can then be checked for missing portions, or holes, with the user warned if any holes are present. Additionally or alternatively, the assembled contigs, generated by either process, can be translated into series of proteins for analysis.

[0155] In one implementation, the sampler interface 2312 performs an initial quality control estimate from the raw DNA before any contigs are assembled. The estimate is performed by a quality control model 2314 that receives the raw DNA as an input and provides one or more quality control metrics representing the expected quality of any contigs that can be formed from the raw DNA. The quality control model 2314 can be implemented as a supervised learning model that is trained on samples including a set of raw data paired with quality metrics assigned to a set of contigs generated from the set of raw data. In one example, the quality control model 2314 is implemented using the MAMBA architecture. The output of the quality control model 2314 can be used to avoid the time and cost of generating contigs or performing further analysis on the raw DNA if the raw DNA is not of sufficient quality to continue analysis. When low quality inputs are detected, the user can be provided with an alert to allow for collection of a higher-quality sample.

[0156] The output from the sampler interface 2314 is provided to a pathogen characterizer 2316 that generates a representation of the pathogen from the provided data. In one implementation, as described herein previously, the pathogen characterizer 2316 generates a count of unique K-mers from the provided DNA or protein data, normalizes the K-mer counts according to a total count, and generates a matrix including the K-mer counts. In another implementation, the pathogen characterizer 2316 can be implemented a machine learning model that generates a representation of the pathogen according to the provided data. In one implementation, the pathogen characterizer 2316 is implemented using a MAMBA architecture that is trained on a sets of sample data each including a set of DNA or protein data as well as an appropriate representation of the pathogen based on the set of DNA or protein data.

[0157] A species predictor 2318 can determine a species of the pathogen according to the data provided from the sampler interface 2314. The species predictor 2318 can be implemented as a machine learning model that classifies the input data into an output class representing a particular species of pathogen. The species predictor 2318 can be trained on a set of training examples each including a set of DNA or protein data and a determined pathogen species. In one example, the species predictor 2318 includes a model based on the MAMBA architecture to determine the species of the pathogen represented by the sample. In one implementation, the species predictor 2318 can include a pathogen predictor 2320 that determines a general category of the pathogen (e.g., virus, bacteria, fungus, or no pathogen) from the data provided by the sampler interface 2314. This determination can then be used to either supplement the input data provided to the model used to determine the specific species of pathogen or to select upon multiple models comprising the species predictor, each trained to determine specific species for a general category of pathogens. Where no pathogen is detected, the user can be alerted and given the option to end the analysis or submit another sample. In one implementation, the input to the pathogen predictor 2320 includes not only the DNA or protein data from the sampler interface 2314, but also a location of an infection, a type of bodily fluid or tissue used for the sample. In this implementation, the pathogen predictor can be trained on a set of samples each including a set of DNA or protein data, a location, a sample type, and a label indicating the class of pathogen.

[0158] A drug characterizer 2322 can receive an identity or representation of a drug and generates a representation of the drug that is suitable for analysis at a potencyanalyzer 2324. It will be appreciated that the term “potency” is explicitly intended to include a minimum inhibitory concentration for an antibiotic applied to a bacterium, as wells as concentrations or dosages of antifungals, antivirals, and other drugs that can be applied to fungi and virus. In one example, the drug characterizer 2322 generates a canonical SMILES representation of a drug from an input SMILES string or an identity of a drug. Additionally or alternatively, the drug characterizer 2322 can include a machine learning model that receives a SMILEs string and generates a representation of the drug for analysis at the potency analyzer 2324 from the SMILES string. In one implementation, the drug characterizer 2322 is implemented using a MAMBA architecture that is trained on a sets of sample data each including a SMILES sequence as well as an appropriate representation of the drug. In another implementation, the drug characterizer 2322 comprises a graph neural network that receives an input representing a molecular structure of the drug as a graph and provides a representation of the drug for analysis at the potency analyzer 2324.

[0159] The potency analyzer 2324 receives the representation of the pathogen from the pathogen characterizer 2316 and the representation of the drug from the drug characterizer 2322 and determines a potency for the drug as applied to the pathogen. It will be appreciated that the potency analyzer 2324 can be implemented as a machine learning model trained on training samples each comprising representations of a pathogen and a drug as well as a label presenting a potency value. In one example, the potency analyzer 2324 is implemented as a recurrent neural network. In one implementation, the potency analyzer 2324 can receive both the output of the pathogen characterizer 2316 and the dataset output by the sampler interface 2314. In this implementation, the potency analyzer 2324 can be implemented as a machine learning model trained on training samples each comprising representations of a pathogen, a dataset representing the DNA, RNA, or protein sequence of the pathogen, and a drug as well as a label presenting a potency value.

[0160] In one implementation, the potency analyzer 2324 can also review the representation of the pathogen and the representation of the drug to determine both a validity of the representation, that is, if the representation of the pathogen could represent a live organism or the representation of the drug represents a valid chepotencyal structure, and importance values for various portions of the representations. In one implementation, the importance values can be determined via attention layers in the machine learning model or models comprising the potencyanalyzer 2324. In this implementation, the potency analyzer 2324 can contain one or more models trained to produce these values from the representations of the drug and the representation of the pathogen, such that the output of the potency analyzer 2324 includes a value representing the validity of each of the drug and the pathogen, outputs representing the importance of locations, such as portions of the DNA sequence representing the pathogen or bonds and atoms within the representation of the drug, and the potency value for the drug given the pathogen.

[0161] FIG.24 illustrates a system 2400 for identifying novel drugs and pathogens. The system 2400 includes a processor 2402 and a non-transitory computer readable medium 2410 that stores executable instructions for simulating both novel pathogens and novel drugs and determining the potency of the drug in treating the pathogen. To this end, the executable instructions can include a pathogen simulator 2412 that generates a novel pathogen for analysis, a drug simulator 2414 that generates a novel drug for analysis, and a potency analyzer 2416 that determines a potency for the drug given the pathogen. In one implementation, each of the pathogen simulator 2412 and the drug simulator are each implemented as genetic algorithms that operate adversarially. The pathogen simulator 2412 iteratively selects for a pathogen that minimizes the potency, and the drug simulator 2414 iteratively selects for a drug that maximizes the potency, and, in one example, the system 2400 can operate for a predetermined time or number of iterations. It will be appreciated that potency analyzer can be implemented in the same manner as the potency analyzer 2324 of FIG.23, and can include the additional models for determining the validity of each of the drug and the pathogen and the outputs representing the importance of locations within the representations of the drug and the pathogen.

[0162] In another implementation, the pathogen simulator 2412 can be implemented as a mutation learning model that provides a pathogen that provides target potency for a given drug. For example, the pathogen simulator 2412 can operate to simulate mutations in a bacterium species. In this implementation, the pathogen simulator 2412 receives DNA (e.g., raw DNA) representing a pathogen, a species of the pathogen, a SMILES sequence representing a drug, and a target potency and provides a mutation of the pathogen that provides the target potency. In one example, the pathogen simulator 2412 can be implemented using the MAMBA architecture and a Stable Diffusion architecture including a U-Net neural network with two encoders, one decoder, and a time stage that iterates over the U-Net neural network.

[0163] Additionally or alternatively, the drug simulator 2414 can be implemented as a drug tuning model that provides a drug that provides target potency for a given pathogen. For example, the drug simulator 2414 can operate to determine an antibiotic to provide a desired potency for a bacterium species. In this implementation, the drug simulator 2414 receives a representation of a base drug as a SMILES sequence, a representation of the pathogen, and a target potency and provides a modified version of the drug that provides the target potency. In one example, the drug simulator 2414 can be implemented using the MAMBA architecture and a Stable Diffusion architecture including a U-Net neural network with two encoders, one decoder, and a time stage that iterates over the U-Net neural network.

[0164] In one implementation, the system 2400 includes a mutation evaluator 2420 that estimates the time necessary to go from a first DNA sequence to a second, mutated DNA sequence. The mutation evaluator 2420 receives the two DNA sequences and provides a value representing the time necessary for a pathogen, such as a bacteria, to mutate from the first sequence to the second sequence. It will be appreciated that the mutation evaluator 2420 can be implemented as a supervised learning model that is trained on a set of samples, each including two sequences of DNA and a label indicating a time necessary to mutate from the first sequence to the second sequence. In one implementation, the mutation evaluator 2420 can be implemented as two MAMBA models providing inputs to a recurrent neural network.

[0165] Further details regarding this element of the system, method and device according to this invention are provided below

[0166] AKF a baseline algorithm, and GKF is the GPU parallel version of AKF, as further described in the Appendix below and the associated Figures. Jellyfish and KMC are well-known K-Mer counting programs within the Microbiology community. Jellyfish was used as a comparator for this element of the invention. Such known K- Mer counters, such as Jellyfish

[0012] or KMC [7], may miss K-Mers that do not appear in a training dataset due to the efficiencies required to process large sizes of ^. This workflow for fingerprint optimization allows the Machine Learning models and BAAGA to operate in near real-time, as compared to the hours or days required for those models to interface with BAAGA, which we have found can take over three weeks to properly tune. The improvement described herein shortens tuning time to a few hours maximum.

[0167] SMILES strings were processed per character in the string. A SMILES string had each character converted into a one-hot-encoding. The one-hot-encodings were based on all possible options that the character could be. The number of possible options a character could be was 27, so each one-hot-encoding was a vector of size 27 with the index of the respective character set to 1. For example, characters could include molecules such as Oxygen or could be simple characters like “=”, which represents a double bond. Once characters were one-hot- each encodingwas then concatenated to form a matrix. The processing of these strings and input to the RNN was similar to processing and input from

[0041] .

[0168] MICs collected contained off-edge results, meaning some values were above or below measurable MIC values and needed to be normalized. To normalize these MIC values over multiple years, the following steps were performed. Any MIC having “≤” had the “≤” removed and the MIC was not changed. An example of this normalization would turn the MIC “≤ 1.0” into “1.0”. Any MIC having “>” had the “>” removed and the MIC raised one dilution. An example of this normalization would turn the MIC “> 1.0” into “2.0”. These MICs were then made into one-hot vectors based on the dilution range [0.008, 2048.0]. This led to each MIC becoming a vector of 19 zeros with the dilution of the specified MIC value being set to 1.

[0169] RNN Model

[0170] At the beginning of the architecture, a SMILES string became a second input. Both fingerprint and SMILES representation inputs were first run through a single embedding layer before being concatenated and sent through the rest of the architecture. Figure 13 shows the basis for the architecture of the RNN and its workflow which those skilled in the art may utilize as a component in setting up the present system, method and device.

[0171] Tuning was performed using KerasTuner version 1.1.3

[0089] . The hyperparameters that were tuned were the number of units in the GRU layer, the Dropout percent in the first dropout layer, the number of units in the hidden dense layer, the Dropout percent in the second Dropout layer, how many convolutions to be used for the compound input, how many filters to use in each convolution layer used in the compound input, learning rate, and epsilon used in the Adam optimizer. Another parameter that was tuned separately was the number of epochs to perform.

[0172] In contrast to Kromer-Edwards et al.

[0056] , which trained a model for each antimicrobial agent, this study used one single model trained on all agents. Keras

[0043] in Tensorflow-GPU version 2.4.0

[0077] created and trained the model. Figure 7 provides an Architecture diagram of the RNN used.

[0173] Training and Testing Results

[0174] Training results showed that overfitting was not a concern like other studies [63, 57, 59, 56]. Overfitting was especially an issue when adding the antimicrobial agents as IDs to the model [59, 85]. These results revealed that the poor performance shown in other studies was due to low complexity in the input data, rather than small input datasets.

[0175] The Food and Drug Administration (FDA)

[0039] and Clinical and Laboratory Standards Institute (CLSI)

[0112] have advised that determining an MIC within a ±12- fold dilution is acceptable. Because of these recommendations, all test predictions that were ±12-fold dilution away from the actual MIC were considered correct.

[0176] Figure 14 displays the F1 score for each of the 10 cross-validation folds used. These cross-validation results show that the model can effectively predict as the variance between folds is minimal. The average F1 score for all folds is 0.82. To compare this model against other models, average F1-micro scores were compared against other average F1-micro scores given in those model manuscripts. Figure 14: Visualizing F1-micro scores for all cross-validation folds

[0177] With that, the below Table shows all average F1-micro scores.

[0178] Table: Compare F1 score averages for the models against other F1 scores

[0179] Generalizing Antibiotics

[0180] To test how well the model can generalize to new antimicrobial agents, four new agents were tested that the model had not seen before. Doxycycline and Minocycline had similar targeting mechanisms or structures to the agents in the utilized dataset. Trimethoprim-sulfamethoxazole and Gentamicin either had a different targeting mechanism or a different structure.

[0181] Testing of the model with these four new antimicrobial agents gave an unremarkable performance with an F1-macro score ranging from 0.41 to 0.5. Here, the F1-macro is used instead of F1-micro due to the large class imbalance shown in Figure 15. Log scale histogram showing the number of samples per new antimicrobial agent per MIC bin for the four generalized antibiotics. This MIC dataset is only for testing the generalization performance and is not used during training. Looking at zero extra random SMILES models in Figure 15 shows low performance. Comparing all generalization performance metric results for four new antimicrobial agents. The zero random SMILES represents training on the canonical SMILES string for the antimicrobial agent. The five random SMILES represent having up to six possible SMILES to train on for each antimicrobial agent. One canonical SMILES and one to five unique, random SMILES. Ten random SMILES have the same definition. One canonical SMILES and one to ten unique, random SMILES. Because of this low performance, more data had to be collected.

[0182] Antimicrobial agent SMILES strings were augmented by randomly generating new SMILES strings. These new strings are based on the canonical SMILES string for each agent. Randomizing of the canonical SMILES strings was performed in a restricted manner based on the findings from Arús-Pous et al.

[0010] . This restrictive form of randomization randomized the ordering of atoms in the molecule formed with RDKit

[0067] . Then, the new SMILES strings were generated based on the new ordering of atoms, while adhering to all RDKit restrictions. An example of such a restriction is prioritizing side chains when traversing a ring, rather than completing the ring first. Two new datasets were formed from the maximum number of randomly generated SMILES each molecule could make. These two datasets were made with five and ten max random SMILES, respectively. Each antimicrobial agent had the respective number of extra, random SMILES strings generated based on the canonical SMILES string. Then, the list of SMILES had duplicates removed. For example, Meropenem in the five random SMILES dataset had the canonical SMILES string and five other random SMILES strings. Once duplicates were removed, Meropenem could be left with anywhere from one to six SMILES in the dataset depending on the randomness of the SMILES generation. Figure 10 shows the test results against the four new antimicrobial agents with augmentation of either five random SMILES or ten random SMILES. F1-macro scores for the four antimicrobial agents show improvement between no extra SMILES and five extra random SMILES. However,there was little performance improvement from five extra SMILES to ten extra SMILES. Average results show that F1-macro scores for five and ten random SMILES were 0.623 and 0.614 respectively.

[0183] Conclusion

[0184] The model initially showed promise with a test average F1 score of 0.82. However, it failed to predict when given new antimicrobial agents. After augmenting training data with extra random SMILES, prediction performance for the four new agents increased significantly on the very skewed new agent datasets used to test. F1- macro scores increased anywhere from 0.1 to 0.2 over the four agents. The difference between using five or ten extra SMILES per antimicrobial agent in the training dataset used was minimal. This shows that adding more random SMILES would not yield improvement. Five random SMILES should be what is used to train the model in the future.

[0185] Comparison of this model against the best models known shows that this could outperform the XGBoost model and DNN model in Kromer-Edwards et al.

[0059] , while still only needing to train one model for multiple antimicrobial agents. This model also has the training efficiency of the RNN models that were trained in Kromer-Edwards et al.

[0056] . In addition to these benefits, the model disclosed herein was able to predict MICs for new antimicrobial agents that the model had not trained on.

[0186] EXAMPLE 2

[0187] FINDING, AND COUNTERING, FUTURE RESISTANCE USING BACTERIAL ANTIBIOTIC ADVERSARIAL GENETIC ALGORITHM (BAAGA)

[0188] Introduction

[0189] Previous studies have delved into predicting Minimum Inhibitory Concentrations (MIC) [63, 15, 57, 13, 62, 59, 56]. Other studies have investigated predicting new antimicrobial agents, given resistance mechanisms [116, 29]. A study by Campos et al.

[0025] did create a simulation algorithm that tried to simulate antibiotic resistance. However, this algorithm requires background knowledge of plasmids, resistance, bacteria, and hospital patient inflow. Background knowledge may not exist in cases of unknown or new resistance mechanisms. The study also is only able to predict counts of plasmids and bacteria rather than specific MIC valuesand mutations that may cause them. This paper presents an algorithm for predicting hypothetical MIC results as bacteria and antibiotics change randomly, but in an adversarial manner. Accordingly, the present system, method, and device includes a method, module, component and integrated algorithms for predicting hypothetical MIC results as bacteria and antibiotics change randomly, but in an adversarial manner.

[0190] Disclosed herein is new device, system, and method called Bacterial Antibiotic Adversarial Genetic Algorithm (BAAGA) that simulates generating new antibiotics and new bacterial mutations. This simulation works in an adversarial manner using two Genetic Algorithms (GA). A RNN

[0058] (for predicting MICs) is the fitness function that links together two GAs. Two GAs were implemented: one for generating bacteria (BGA) and another for antibiotic development (AGA). The first (BGA) is part of a model trying to maximize the fitness function (MIC). The second (AGA) is trying to minimize this same fitness function. The fitness function here is the MIC prediction from an RNN. Restricting the randomness in the generation of bacteria in the BGA and antibiotics in the AGA lead to viable individuals. GAs have been used in the past in an adversarial manner. An example of adversarial GAs is Bock et al.

[0020] to overcome country censorship. Here, these GAs, together in the BAAGA simulation, find what could be the next resistance mechanisms and the antimicrobial agents that counter those mechanisms. BAAGA is demonstrated to simulate the cycling of new resistance mechanisms forming and new antimicrobial agents which counter those new mechanisms.

[0191] Data Processing and Background

[0192] An individual bacterium is represented by a genetic fingerprint, and an antimicrobial agent is represented by the Simplified Molecular Input Line Entry System (SMILES) [6] string that encodes its chemical structure.

[0193] RNN Model

[0194] The MIC RNN predicts the MIC for a given bacterium and antibiotic combination. The architecture of the RNN consists of two inputs, one for a fingerprint and one for a SMILES string. The fingerprint input goes through a bidirectional Gated Recurrent Unit (GRU) layer with attention. Then, that goes through a dropout layer and a dense layer. The SMILES string input goes through two 1-D convolutional layers before flattening and going through a dense layer. At this point, both inputs are concatenated together and go through a dense layer, a dropout layer, and, finally, a dense layer again. Building and training the RNN model uses Keras

[0043] inTensorflow-GPU version 2.4.0

[0077] . A training dataset built with five additional randomized SMILES strings provides the highest generalized performance

[0056] . Therefore, the RNN in BAAGA utilizes this training dataset.

[0195] GA Setup

[0196] There are two GAs in BAAGA: an Antimicrobial agent GA (AGA); a Bacterial agent GA (BGA). Helper functions from the DEAP v1.3.1 library

[0040] build both GAs. Both AGA and BGA use DEAP’s tournament selection function selTournament for the Selection step. The Crossover and Mutation steps are distinct, and custom-made, for AGA and BGA.

[0197] The first task when working on GAs is to determine how specific or vague the search spaces should be. For example, the BGA search space could be large, modeling real life, as other nonGA models have tried [54, 118]. A small search space could have a K-Mer count fingerprint be an individual and mutations be the changes to the K-Mer counts of that fingerprint. Both search space sizes have their pros and cons. A medium-sized search space must be capable of finding new individuals, while also being capable enough to report possible resistance mechanisms. An example search space proposed by Thomas et al.

[0105] used K-Mers as individuals, and successfully- selected K-Mers were then combined into a final individual. With these options in mind, BGA utilizes a list of contigs made of nucleotides as the search space for each individual. Contigs can have Single Nucleotide Polymorphisms (SNPs) added, generating resistance mechanisms. Processing contigs into fingerprints is also easy. Therefore, the search space for BGA consists of contigs from an isolate.

[0198] The Crossover step randomly swaps contigs between two isolates. The Mutation step iterates every position in all contigs for an isolate. Each position has a chance to pseudo-randomly create a SNP. Furthermore, the Mutation step has a chance to create an Insertion or Deletion (INDEL) in each contig. When adding a SNP, the initial isolate’s nucleotide transition matrix decides which nucleotide to add or swap. Counting and determining averages for all K-Mers of size two from the seed isolate’s contigs generates this transition matrix, and this matrix forms a list of conditional probabilities. These conditional probabilities are defined as the probability of nucleotide y following nucleotide z in a sequence, or P(y|z). An example of such a probability is P(A|T) being the probability that nucleotide T follows nucleotide A in a DNA sequence. When creating a SNP, a position where the SNP occurred is first chosen, here labeled x. The nucleotide at position x, the nucleotide at position x + 1,and the transition matrix, generate four probabilities. Each probability corresponds to each of the four nucleotides (A, T, C, and G) that can replace the nucleotide at position x. For example, if the nucleotide at position x + 1 is A, then the four probabilities are P(A|A), P(T|A), P(C|A), and P(G|A).

[0199] A transition change converts nucleotides A and G or T and C. A transversion change is one of any of the other possible changes, such as A to T, A to C, C to G, etc. If the SNP creates a transversion change, the respective probability is multiplied by 0.3 based on the probabilities found in [34, 92]. The four probabilities are, then, sorted by nucleotide and converted to a Cumulative Distribution Function (CDF) to randomly choose a nucleotide for the SNP.

[0200] The process for selecting the next nucleotide in an insertion is similar to the process for SNPs. For each new nucleotide, there is a probability to stop inserting and finalize the insertion. Placement of the insertion into the contig occurs at this point. If the insertion is not stopped, the previously inserted nucleotide, along with the transition matrix, randomly chose the next nucleotide. However, the difference is that the type of change (transition or transversion) is not taken into effect for the four probabilities before forming the CDF.

[0201] For the AGA, the Crossover step from MolFinder

[0066] combines both SMILES strings into two new SMILES strings. For the Mutation step, the respective functions in MolFinder add, delete, and replace atoms. Every time a SMILES string mutates, one of these options, or nothing at all, occurs.

[0202] Simulation Modeling

[0203] The simulation model defined in this study is BAAGA which consists of the single RNN model, AGA, and BGA. The RNN model, using MICs, is the fitness function that joins the two GAs together. The AGA uses the fingerprint of the previous best isolate found from the BGA as a baseline through all iterations of a segment. The goal of the AGA is to find the antimicrobial SMILES string that gives the lowest MIC prediction from the RNN. The BGA uses the previous best SMILES string from the AGA as a baseline to find the isolate that produces the highest MIC from the RNN for a segment. Therefore, the AGA is minimizing the fitness function and BGA is maximizing the fitness function, and that fitness function is the predicted MIC from the RNN. The AGA and BGA do not run in parallel. The AGA runs for one segment, then the BGA runs for one segment, then AGA, etc. At the end of each segment, the respective GA sends the best individual to the other GA for that GA’ssegment.A round of simulation consists of the AGA and BGA running for 150 iterations / generations each. Starting the simulation consists of seeding an isolate’s contigs and an antimicrobial agent’s SMILES string. The first 150 iterations (1 segment) of the AGA seeds the population with the SMILES string provided. Once that segment completes, the best SMILES string is then passed to the BGA to have a segment run of the BGA. This first segment run of the BGA seeds the population with the input isolate. After that segment is complete, the best isolate is then passed to the AGA, and the previous best SMILES string seeds the next population. In all, each segment of a GA has its population reseeded using that respective GA’s previous segment best. When predicting an MIC, the GA will use the respective individual along with the best output from the other GA in that GA’s previous segment run.

[0204] In total, three rounds of the simulation, three segments of AGA and three segments of BGA, run. This leads to six total segments of a GA run in a single simulation. Each segment consists of 150 iterations for the respective GA, so 900 total iterations run for each complete simulation.

[0205] Tuning Parameters

[0206] Both GAs have specific parameters to tune, along with parameters set up for the simulation as a whole. Overall GA parameters that need to tune for both GAs are population size, number of generations / iterations to run, tournament size, and crossover probability. The AGA specific parameters that need to tune are the probability to add an atom, the probability to replace an atom, the probability to delete an atom, and the probability to keep ring structures when crossing over. The BGA specific parameters are SNP probability, INDEL probability, and probability to mutate a contig when traversing contigs.

[0207] Tuning uses 2k factorial design

[0115] which refers to testing three points, rather than all possible points like GridSearch. The three points are a low value, a default value, and a high value. Here, the selected points are as is, rather than performing multiple short rounds of 2k factorial design, due to timing restrictions. In total, the simulation tunes 12 parameters. Each parameter had either three or four options to test (all three AGA atom change probabilities combined into one parameter with four options since the probabilities had to add up to one). Altogether, the simulation has to run 708,588 times to traverse all the possible parameter combinations. However, to get a better understanding of how these parameters behave over different inputs, each set of parameters run three simulations. The same isolatecontigs, but different antimicrobial agent SMILES strings (Aztreonam, Meropenem, or Ceftriaxone), seed each of the three simulations. An overall result for a set of parameters comes from combining results across the three runs. This means that, in total, the tuning has to run 2,125,764 times to obtain the proper set of parameters to use. To that end, the tuning process split into four separate tuning runs. Each run tests three parameters while the other parameters do not change. These four tuning runs result in 297 (81+108+81+27) total simulation runs. While splitting the tuning process up into multiple tuning runs reduces the parameter search space, it also makes tuning possible, as each run of BAAGA takes, on average, three hours to perform. Therefore, splitting the tuning up into multiple runs drops the runtime from approximately 728 years to 37 days.

[0208] Three attributes determine which value to select for a parameter during tuning. The first is the Average Overall Standard Deviation (AOSD) for all segments of a simulation. The standard deviations of the last iteration for each segment of each simulation run are added together to give the AOSD for a single antimicrobial agent. Total AOSD is the average AOSD of the three antimicrobial agents used.

[0209] The second and third attributes collected were the Overall Segment MIC Mean (OSMM) and Overall Segment MIC Standard Deviation (OSMSD). To calculate these, first, collect the best “performers” for each segment. The best performer for AGA is a SMILES string that has the lowest MIC, and the best performer for BGA is the isolate contigs that give the highest MIC. MICs extracted from these performers form the basis for OSSMM and OSMSD. For a given simulation run, OSMM is the mean of the MICs and OSMSD is the standard deviation of the MICs. The total OSMSD and OSMM are the average of the final results for both attributes for the three antimicrobial simulations. The goal is to have a set of parameters that has the highest AOSD and OSMSD while also having a low OSMM. These goals yield simulations that have the highest chance to find new SMILES or contigs in the shortest time possible.

[0210] By wanting high standard deviations, the GAs have more diverse populations by the end of the simulations. Having lower OSMM means that the AGA has a SMILES string that the BGA cannot find better isolates for, thus leading to new antimicrobial agents. The final parameter set ends with an AOSD of 12.187, an OSMM of 7.722, and an OSMSD of 3.526. Note here that the OSMM is the MIC index coming from the RNN prediction. Rounding that up gives an MIC index of 8,which corresponds to an MIC of 2.0. The simulations all start with MICs of 16.0 or 32.0, which supports that BAAGA can produce new antimicrobial agents to lower the MIC and struggles to find new resistance mechanisms to raise the MIC.

[0211] Results

[0212] As previously mentioned, the search space utilized for BGA was large enough that the algorithm may not find novel mutations to impact the predicted MIC in the 150 iterations, or segment, it can run. Because of this, each antimicrobial agent runs through the simulation three times with the most diverse simulation selected as the final result. Figure 17 for Aztreonam, Figure 18 for Ceftriaxone, and Figure 19 for Meropenem shows these chosen results, each displaying the best and average individuals for each iteration of the respective GA in log2(MIC) format. MIC values are on a log scale, so a plot visualizing these values is difficult to interpret. Going forward, plots will show MICs as log2(MIC) to better visualize the MICs that BAAGA produces. The best individual for AGA would have the lowest log2(MIC) while the best individual for BGA would have the highest. For example, an MIC of two is log2(2) = 1 and an MIC of one is log2(1) = 0. Once a GA finds the best individual, that stays the best individual until the GA finds a better individual. This means that straight lines, denoting “Best”, are straight until the GAs find a better individual. The longer a line is straight, the more stagnate the GA is, due to the large search space. Lines denoting “Average” represent the average log2(MIC) for all individuals in a given iteration / generation. The dashed “Seed MIC” lines represent the starting log2(MIC) that the RNN found, given the seed isolate’s contigs and antimicrobial agent’s SMILES string. These dashed lines are a reference to see how the GAs progress. The orange side of the graph (left) represents segments that the AGA ran, while the blue side of the graph (right) represents segments that the BGA ran. Segments are in increasing order so that the top left plot in each is segment zero, which is the first segment of the simulation. The bottom right plot in the figures displays the last segment to run in the simulation.

[0213] Figure 17: Simulation results for Aztreonam. Each plot in the figure represents a segment which is 150 iterations. The top left, segment 0, is the first segment run of the simulation. The bottom right, segment 5, is the last run of the simulation. These results are from the third simulation run.

[0214] Figure 18: Simulation results for Ceftriaxone. Each plot in the figure represents a segment which is 150 iterations. The top left, segment 0, is the firstsegment run of the simulation. The bottom right, segment 5, is the last run of the simulation. These results are from the first simulation run.

[0215] Figure 19: Simulation results for Meropenem. Each plot in the figure represents a segment which is 150 iterations. The top left, segment 0, is the first segment run of the simulation. The bottom right, segment 5, is the last run of the simulation. These results are from the second simulation run.

[0216] Overall results demonstrate that the AGA, with its smaller search space, is capable of finding SMILES strings that lower the MIC more often. The larger search space of the BGA causes the algorithm to struggle to find better contigs. However, while the BGA struggles to find new, useful SNPs and INDELs, the new SNPs and INDELs the algorithm does find have an extreme impact on not only the best individual but also the population. This result persists throughout the figures as the average AGA population struggles to catch up to the best SMILES string, while the populations in the BGA quickly catch up (within a few iterations). Unintentionally, the Crossover step appears to behave like horizontal gene transfer, which allows for the population to quickly obtain which SNPs or INDELs are most beneficial.

[0217] Along with the population changes, another outcome from Figures 17, 18, and 19 is that the simulation is able to capture the cycling nature described earlier, but on a smaller scale. In real life, there is an exponentially larger number of bacteria compared to the small 25 individuals used in BGA. Though the search space for BGA is large, these results support that, even on a small scale, resistance mechanisms can generate quickly to counter new antimicrobial agents. The cycling shown here can determine how bacteria may react to a given antimicrobial agent and what steps to take to prepare for that reaction.

[0218] Finally, Figure 18 shows two separate times when the BGA found new resistance mechanisms that made the MIC go higher than the baseline. Figure 19 also shows this, but the AGA was able to find a new compound to overcome the new resistance mechanism. The result in Figure 19 from the AGA is a candidate for further review. The new compound shows, through simulation, that it is able to lower MICs more than Meropenem on its own. The new compound also can reduce MICs after a new resistance mechanism emerges.

[0219] Conclusion

[0220] A system component includes a simulation algorithm, named BAAGA, that successfully simulates the vicious cycle that researchers and physicians encounterwith bacteria. This cycle uses antimicrobial agents and antimicrobial resistance to counter one another. Using BAAGA to simulate three antimicrobial agents against a known resistant isolate demonstrates the ease at which bacteria can gain resistance mechanisms and how slow the process is to find new antimicrobial agents to counter those new mechanisms.

[0221] How BAAGA can produce possible new antimicrobial agents to counter mechanisms of resistance that currently do not exist is also demonstrated This demonstration may be the key to understanding exactly how bacteria may react to new antimicrobial agents. The results also may provide insight into what antimicrobial agents researchers and physicians study next, once new mechanisms of resistance emerge.

[0222] EXAMPLE 3

[0223] USE OF AMINO ACID DATASETS IN THE PRESENT INVENTION

[0224] Protein sequences can be utilized rather than DNA sequences within the system. The DNA sequences are translated into protein sequences, giving a list of amino acid strings. This translation step occurs at the end of the Bioinformatics Pipeline on the contigs. Each contig is translated into the respective protein sequence from DNA. The new list of protein contigs is sent to the Machine Learning Bacterial processing module and BAAGA module. The Machine Learning Bacterial processing module operates without modification as K-Mers are agnostic to the input sequence type. The MIC prediction pipeline module is retrained with a new MIC ML Model. This new model is trained by taking the original bacterial training data and translating all isolate contigs into protein sequences before training. The same dataset is also used to train a new Species ML Model. BAAGA utilizes the new MIC ML Model trained on the bacterial protein sequence dataset to predict MICs for fitness. Also, BAAGA modifies the Bacterial Genetic Algorithm by converting the nucleotide probabilities into amino acid probabilities. This modification allows the Bacterial Genetic Algorithm to mutate the amino acids within the contigs. Another optional way the Bacterial Genetic Algorithm could be modified, rather than converting the input and probabilities, is to translate the DNA contigs into protein contigs just before MIC prediction. This would leave DNA contigs as input and DNA-based probabilities but still predict MICs based on the amino acids through translation.

[0225] EXAMPLE 4 USE OF BACTERIOPHAGE AS THE ANTIBIOTIC

[0226] Antibiotic compounds can be replaced within the system with Bacteriophage viruses. This replacement allows for the development of new Bacteriophage strains to combat bacteria rather than relying on new antibiotic compounds. The following changes are required to utilize Bacteriophages. First, the antibiotic input at the start of the system is replaced with Bacteriophage DNA sequences. These sequences are processed via a Bioinformatics Pipeline that assembles the virus DNA, provides taxonomy, and strain information. This Bioinformatics Pipeline would be specific for the Bacteriophage and replace the Machine Learning Antibiotic processing module. The assembled Bacteriophage contigs are sent to BAAGA and a Machine Learning Bacteriophage processing module. The Machine Learning Bacteriophage processing module performs the same processing as the Machine Learning Bacterial processing module. The two modules could be the same in code but may be kept separate in planning for clarity. BAAGA’s Antibiotic Genetic Algorithm is converted into a Bacteriophage Genetic Algorithm. This Bacteriophage Genetic Algorithm behaves like the Bactria Genetic Algorithm, but the probabilities are specific to virus DNA probabilities. Finally, the MIC prediction pipeline requires the MIC ML Model to be changed by removing the antibiotic branch of the model and replacing it with another fingerprint input used for Bacteriophage. This means that the MIC ML Model has the bacteria fingerprint and the Bacteriophage fingerprint as the two inputs. The output remains the MIC of a given Bacteriophage for that bacterium, and thus, bacteriophage, for the purposes of this invention, are included as another antimicrobial available for development by researchers and use by Clinicians to treat pathogenic microbial infections, according to this invention.

[0227] EXAMPLE 5 MIC’S FOR ANTIFUNGALS

[0228] In the system, method and device according to this invention, when utilized for fungal isolates rather than bacterial isolates, the Bioinformatics Pipeline module’s Kraken database files are modified, MLST database files are changed, and QC modified to support fungal isolates rather than bacterial isolates. A new training dataset is used to contain fungal DNA contigs, antifungal compounds, and the MICs that join the two data types together. This new dataset is used to train the MIC ML Model and Species ML Model. BAAGA is provided with the Bacterial GeneticAlgorithm converted into a Fungal Genetic Algorithm. The modifications required to perform this conversion require changing the nucleotide, insertion, deletion, and substitution probabilities that make up the mutation step of the Genetic Algorithm.

[0229] EXAMPLE 6

[0230] ANTIVIRAL MIC PREDICTION

[0231] The system, device and method according to this invention is modified to utilize virus isolates rather than bacterial isolates. To do this, the Bioinformatics Pipeline module’s Kraken database files are modified, MLST database files changed, and QC modified to support virus isolates rather than bacterial isolates. Training datasets containing both virus DNA contigs, antiviral compounds, and a data points to join the two data types are provided to define reduction in plaque counts of 50% from a Plaque Reduction Assays (PRAs). This dataset is used to train a new PRA ML Model and Species ML Model. BAAGA’s Bacterial Genetic Algorithm is converted into a Viral Genetic Algorithm. The modifications required to perform this conversion include changing the nucleotide, insertion, deletion, and substitution probabilities that make up the mutation step of the Genetic Algorithm.

[0232] Appendix

[0233] This appendix is being included in this patent disclosure to ensure that the general disclosure and specific examples provided herein above, and the claims, herein below, are fully described, supported and enabled. In this regard, while not considered essential for either written description or enablement purposes, because 37 CFR 1.57 only permits incorporation by reference to patent documents for material which, retrospectively, may be considered “essential”, the following text is included, absent intermediate section headers, from Kromer-Edwards, Cory, "Optimizing K- Mer Fingerprint Generation for Machine Learning." Proceedings of the 14th ACM International Conference on Bioinformatics, Computational Biology, and Health Informatics, September 2023, Article No.: 101, Pages 1–5, which is herein incorporated by reference in its entirety for this purpose, including explicitly as follows:

[0234] With the increasing availability of genomic data obtained through Whole- Genome Sequencing (WGS), Machine Learning (ML) algorithms are being used to analyze this data. However, processing large datasets or files poses challenges. One approach is to count K-Mers, which has been used in ML studies. However, larger K- Mer sizes may lead to decreased accuracy and training difficulties. Alternatively, combining multiple K-Mers of smaller sizes into fingerprints has shown promise in predicting species and antibiotic resistance. This study compares existing fingerprint generation techniques with a new algorithm, called GPU K-Mer Fingerprinting (GKF), which utilizes a GPU for parallel processing. GKF demonstrates similar memory utilization compared to other approaches but achieves a speedup of 5,546X.

[0235] As Whole-Genome Sequencing (WGS) becomes cheaper and more readily available, genomic data becomes more plentiful as input to Machine Learning (ML) algorithms. Processing of WGS reads through de novo assembly creates contigs. There have been studies using genomic data from bacterial isolates to predict antibiotic resistance [8, 9], species

[0011] , and predict gene phenotypes [1, 3] as a few examples. However, bacterial isolates are relatively small for the amount of genomic data that is sequenced. Contig files range from 20 MB to 200 MB in size.

[0236] The issues come into play with large datasets, larger data files, or requiring real-time inference. While the files may be small, processing them is extremely timeintensive for ML. There must be ways to process these large amounts of data in a memory and time-efficient manner.

[0237] One approach to processing these contigs is to count the number of K-Mers of a particular size ^. Counting K-Mers has been utilized in ML studies [3, 8, 9, 11]. Studies have delved into making K-Mer counting efficient [7, 12] and have it parallelized using a GPU [4, 5]. However, when training ML models, large K-Mer sizes have been found to lead to diminished returns on the accuracy, but also may not be trainable at all [9, 14]. At a K-Mer size 10, Kromer-Edwards et al. [9] showed that training an XGBoost model can only occur on large, expensive compute instances. Along with this, standard K-Mer counters, such as Jellyfish

[0012] or KMC [7], may miss K-Mers that do not appear in a training dataset due to the efficiencies required to process large sizes of ^.

[0238] A more straightforward approach is to count multiple K-Mers of small ^ sizes and combine them to form a data representation called a fingerprint. Studies have shown that fingerprints built from ^ sizes of three, four, five, six, seven, and eight are enough to train ML models to accurately predict species

[0011] and antibiotic resistance [8]. However, little work has been done to make fingerprints efficiently. This paper compares current Fingerprint generation techniques with new approaches, including using a GPU to parallelize the process of generating fingerprints. A new algorithm named GPU K-Mer Fingerprinting (GKF) has memory utilization on par with other approaches while obtaining a speed up of 5,546X compared to using a HashMap with other KMer counters.

[0239] WGS sequencing can generate many short-read sequences that can be filtered and de novo assembled into contigs. Contigs are non-overlapping, contiguous sequences of DNA. These contigs can then be further processed depending on the problem that is trying to be solved. For example, one problem is finding whether genes from bacteria could confer resistance to antimicrobial agents [2, 18]. For this problem, contigs can be processed by searching for the genes of interest and collecting mutations from the contigs in the gene regions. For larger problems, such as determining if bacterial isolates confer resistance to antimicrobials, the input must be the entire list of contigs for a single isolate. With these problems, the contigs may be processed into K-Mers. A K-Mer is a unique substring of size ^ of either DNA or protein sequences.

[0240] There are two main approaches to using K-Mers to tackle these problems. The first approach is to take a large ^ value, for example, 27, and only count K-Mers of that size. Many studies have delved into making this task both time efficient, using GPU parallelization [5], and memory efficient, using the disk when the data becomes too large to fit into memory [7, 12]. However, this approach is not suitable for ML as the K-Mers that are present may sometimes be different from sample to sample. These large ^ sizes have too many possible K-Mers to fit into memory to train a model, and a dataset to train a model may be overwhelmingly large, for example, 300 GB [9]. The other approach uses multiple, small K-Mers to form a representation, or fingerprint, of a sample. The fingerprinting method has a small dataset size and can easily be trainable for multiple ML algorithms [8, 10, 11]. Little work has been done to optimize the fingerprint approach for generating ML datasets. Optimizing fingerprint timing is crucial as ML algorithms for clinical data become more widespread in clinical settings and simulations. An example of such a simulation is finding new resistance mechanisms that may appear over time as antibiotic use increases

[0010] . In both cases, fingerprints must be generated near real-time to allow research to complete faster and decrease computation time and cost.

[0241] There are two ways to create fingerprints for ML algorithms. One way is to have a matrix where each row is a K-Mer size, and the number of columns is 4^^^_^ . For example, if a fingerprint is generated from K-Mer sizes [3, 4, and 5], there would be three rows and 45= 1, 024 columns. Since ^ sizes three and four only have 64 and 256 possible K-Mers, respectively, the first 4^ columns contain the K-Mer counts for that row, and the rest are filled with 0’s. This matrix can then be used in algorithms such as Convolutional Neural Networks. Another way of creating fingerprints is to take the matrix previously mentioned and unroll it to form a list. From here, the extra 0 paddings may optionally be removed. This list can then be input for Recurrent or Dense Neural Networks.

[0242] There are two ways to generate a fingerprint. One approach is to use embeddings for K-Mers and concatenate those embeddings together

[0013] . While this approach may be efficient since it uses a single embedding layer to create the encodings, a ML model must be kept and maintained and cannot be parallelized easily.

[0243] The other approach to generating a fingerprint is similar to creating large K- Mer sizes. This approach first creates a Hashmap data structure to hold all possible K- Mers of sizes three through five, as done in Lugo et al.

[0011] or three through eight, as done in Kim et al. [6]. The algorithm associated with this approach is named Hashmap K-Mer Fingerprint (HKF). The keys in the Hashmap are the K-Mers; the values are where the K-Mers would appear in their respective K-Mer size’s ordered list. For example, 3-Mers has 64 possible options. The 3-Mer “AAA” would have a value of zero as this K-Mer is the first in that list, and the K-Mer “GGG” may have a value of 63 as it would be the last K-Mer in the list. The sequences of interest then have all the respective K-Mer sizes counted. Those counts are placed into arrays based on the KMer’s Hashmap value. While this approach does not rely on prior knowledge, such as a ML model, it could be more efficient as multiple look-ups in the Hashmap are required per K-Mer per round of counting.

[0244] This study improves the HKF algorithm by removing the Hashmap requirement. Once the Hashmap is removed, a GPU can make the algorithm parallel. This new, parallelized algorithm is named GKF. This study compares the GKF algorithm with a CPU-only version, called Array KMer Fingerprint (AKF), and the original HKF. All three versions of the algorithms are then compared in the C language and Python programming language as extensions. Having GKF as a Python extension allows for tight integration with ML programs.

[0245] When discussing the possible algorithms, the HKF should be addressed first to grasp what is currently being utilized. The algorithm’s core is shown in Algorithm 1. In Algorithm 1, the input parametergenerated, and the input parameter^^^^^^^^^ contains strings of all DNA sequences of interest that need a fingerprint ^^^^^^^^_^^ is a list of integers being either [3, 4, 5], [3, 4, 5, 6, 7, 8], or any list of wanted K-Mer counts to generate a fingerprint. For the purposes here, all the function code referencing this function will use a K-Mer list of [3, 4, 5, 6, 7, 8]. Lastly, in Algorithm 1, designated for a specific ^^^^^^^^^^_h^^ size and has the keys set to K-Mers of that ^h^^^() builds a list of Hashmaps. Each Hashmap in the list is ^ size and values being the index of that K-Mer in that size. The values in this Hashmap are incremented each time a respective K-Mer appears in one of the given sequences. Other tools have augmented this HashMap to work better in high K-Mer size / parallel settings and on GPUs [5, 12]. However, studies generating fingerprints in the past have not used thesetools to aid in the generation of the fingerprints [6, 8, 10, 11]. These studies used the HKF algorithm in Algorithm 1, see Figure 20 A.

[0246] Removing the Hashmap requires a new means of tracking K-Mer counts. Instead of a HashMap, a single array can be used to track counts. The new algorithm that uses this array is denoted AKF and is given in Algorithm 3. The crucial function of AKF, represented with the function name of AKF iteratively makes the index as it traverses K-Mers by repeatedly calling ^^^^^ in Algorithm 2, is where bit logic computes a K-Mer index through every function call. ^^^^^^. There are index three main components to ^^^^^^: the left shift by two bits, the inserting of the new nucleotide through the bitwise OR operator, and finally, the creation of the index for the input nucleotide. Shifting the bits to the left by two allows the new nucleotide bits to go in the first and second bits.

[0247] The index for the new nucleotide is created using the binary representation of the nucleotide’s capital ASCII code. For the capital letters that make up nucleotides (A, T, C, and G), the second and third bits can create a means of ordering them. For example, the letter “A”s ASCII code is 01000001. Looking at bits two and three show 00. Nucleotides, then, have the following indices: “A” has an index of “00”, “C” has an index of “01”, “T” has an index of “10”, and “G” has an index of “11”. This two- bit index allows the ordering of the nucleotide letters in binary. The third part of ^^^^^^ selects bits two and three through a bitwise AND operation of the letters with seven (00000111), then bit-shifting to the right by one. The integer created from these bitwise operations of the nucleotide’s ASCII code is then the index of the K-Mer built to that point. See Figure 20 B.

[0248] Along with creating an index through bitwise operations, the other main difference between HKF and AKF is that AKF unrolls the processing of the K-Mers and does them all at once. Since the ^^^^^^ function can dynamically generate an index for a K-Mer, there is no need to keep track of an index for each K-Mer separately. Iteratively generating an index leads to only needing to process one nucleotide for each successive K-Mer size. For example, going from ^ = 3 to ^ = 4 only requires doing bitwise operations on one more nucleotide rather than processing a whole KMer. Because the index is then shared between all K-Mers, it must be further augmented by a bitwise OR operation with the number signifying which ^ size is currently being processed. In

[0249] Algorithm 3, this index augmentation is the size of the fingerprint array multiplied by the index of the ^ size being checked against.

[0250] Looping over each sequence was made parallel to allow AKF to benefit from GPU parallel processing. Looping over the list of sequences was kept sequential on the CPU. AKF runs each DNA sequence (contig here) parallel across a GPU. Each thread was assigned a single nucleotide position in the sequence to count all K-Mer sizes. The fingerprint array was global with atomic adding operations to increment each K-Mer count. These code changes are shown in the ^^^^^^^^ function in Algorithm 4. The complete GKF algorithm is in Algorithm 5.

[0251] To be useful for ML, the algorithm should be callable in Python to generate new data on-the-fly. To that end, the C and C++ code was built into Python extensions. The extensions were created using the Python C-API

[0015] . These extensions are also tested alongside standalone C and C++ code for the respective algorithms.

[0252] Every tested algorithm, language, and device combination ran 1,000 times on the same Fasta file. The Fasta file held 134 contigs from a bacterial isolate with 5,770,170 bases over all contigs. The total file size was 5.59 MB. Loading of the file into memory to be processed was tracked in the timing calculations. The Fasta file reader in C was utilized from Turro et al.

[0017] to read in the Fasta files, contributing to timing.

[0253] To validate the algorithms’ correctness, the last run for each combination had the K-Mer counts manually checked against the K-Mer counts from the Fasta file. To manually generate the K-Mer counts, a set of regex strings representing each K-Mer was used to count each available K-Mer and check that all outputs from the algorithms matched. The regex string template utilized was (? =X) where ^ is the K- Mer of interest to be counted. For example, if the K-Mer was “AAA”, then the regex would be (?=AAA), and the number of matches of that regex would be the count of

[0254] “AAA” in the sequence of interest. All algorithms in this study are deterministic, so the ordering of the K-Mers will be the same for the respective algorithm. However, something to note is that the nucleotide indexing in AKF and GKF did result in a different ordering of K-Mers than Jellyfish and HKF. Due to the ordering difference, the same algorithm should be used to generate training and inference input for ML algorithms. After verifying the K-Mer order for eachalgorithm, counts of the K-Mers were manually checked for each algorithm, and all algorithms had correct counts for all K-Mers.

[0255] All testing and validations were run on a desktop computer with 64 GB of RAM, an AMD Ryzen 95900X 12-Core CPU, and an NVIDIA 3060TI GPU with 8 GB of memory. All testing was run single-threaded on the CPU. The time to run the algorithms, RAM utilization, GPU utilization, and GPU memory were tracked over all runs.

[0256] Memory and timing results are recorded for each algorithm. As a comparison, Jellyfish is also included. Jellyfish was compiled and run with Python bindings. To that end, Jellyfish results will be recorded using Python and C as the bindings called C functions. Jellyfish is the only K-Mer counter algorithm comparison made because KMC is built to utilize the disk for larger K-Mers and files rather than focus on small K-Mers, and Gerbil mentions in the original manuscript that any K-Mer counts less than 32 will have comparable results to Jellyfish even though a GPU is utilized [5]. The average timing results for all 1,000 runs of each combination, in seconds, is recorded in the table in Figure 22B.

[0257] The C extensions can add timing variance due to data moving between Python and C. To view this variance, the Box Plot in Figure 22 A shows all running combinations that used C extensions and running in base Python. Along with timing results, memory usage was also captured. Table 1 shows memory usage in MiB across all algorithm / language combinations. Tracking of memory usage was accomplished through Procpath

[0016] .

[0258] The average GPU utilization for GKF was 28%, and the average GPU memory used was 128 MiB. During testing, the number of threads per block was set to 512. Different threads per block values, such as 2,432, were tested, but the GPU utilization and memory did not change. The bottleneck for higher utilization, thus faster running time, is the size of the processed contigs from a Fasta file. Because the bottleneck is in the data, further speedup can be achieved by processing multiple sequences or files in parallel on multiple CPU threads.

[0259] As expected, creating the C extensions did impact timing and memory utilization. However, the timing differences largely resulted from the Hashmap requiring generation every time a Fasta file was processed. Generating the Hashmap each time was due to the C code being unable to store the Hashmap data structure easily while processing multiple Fasta files. There were two possible options toovercome this. One option would be to create and store the Hashmap in Python, then send that Hashmap to C for every Fasta file. The other option, and one applied here, was to generate the Hashmap for every Fasta file in C. The reason for choosing this option was because of the processing time that would have occurred when moving an object from Python to C. When moving the Hashmap over from Python, a PyObject variable temporarily stored the Hashmap. The C code would then create a new data structure to store the data from the temporary Python object. Rather than slowly creating a Python object and slowly transferring that object to C for every Fasta file, C code was made to create the Hashmap every time. Jellyfish provides a way to generate the HashMap once and use it on each successive run, with a fine-tuned HashMap algorithm dedicated to performance optimization.

[0260] The optimizations made in Jellyfish demonstrated half the time of HKF in Python and Python with C extensions. Studies, such as Kromer-Edwards et. al

[0010] and Lugo et al.

[0011] that use HKF, thus, could benefit by switching to Jellyfish for generating fingerprints. However, AKF and GKF improved the timing of Jellyfish by having a speedup of 437X and 2,822X, respectively. While Jellyfish has an optimized HashMap-based algorithm, it was designed for computing a single, large K-Mer count at a time. AKF and GKF simultaneously count all K-Mers three through eight due to the incremental K-Mer index. If the average Jellyfish time, shown in Table 1, only had K-Mer counts of size eight counted and did not need to generate a fingerprint, then Jellyfish would have a time of around four to five seconds. This distinction is why AKF, and subsequently GKF, can have such a drastic speedup compared to Jellyfish.

[0261] The timing results in the table shown in Figure 22B show that the average speedup between HKF and AKF is 1,029X, and the average speedup between AKF and GKF is 5X. Along with these speedups, moving from HKF to either AKF or GKF also decreases the variance in timing, as shown in Figure 22A. While the speedup between AKF and GKF is 5X, both work in milliseconds to process a Fasta file. If many Fasta files do not need to be processed in real-time, or if working on larger Fasta files where memory is a concern, then AKF is suitable. However, in the case of ML, where it can be assumed that one or multiple GPUs are available with larger memory capacities, GKF can increase dataset processing speeds significantly when hundreds or thousands of Fasta files are required to be processed. Utilizing CPU threading as another means of parallelization can further decrease processing time forAKF and GKF, as both algorithms had low memory and CPU utilization throughout all testing. References cited in the Appendix: [1] Ekaterina Avershina, Priyanka Sharma, Arne M. Taxt, Harpreet Singh, Stephan A. Frye, Kolin Paul, genotype-to-phenotype prediction of resistance towards ^-lactams in Escherichia coli and Klebsiella Arti Kapil, Umaer Naseer, Punit Kaur, and Rafi Ahmad. 2021. AMR-Diag: Neural network based pneumoniae. Computational and Structural Biotechnology Journal 19 (2021), 1896–1906. https: / / doi.org / 10.1016 / j.csbj.2021.03.027 [2] Ekaterina Avershina, Priyanka Sharma, Arne M. Taxt, Harpreet Singh, Stephan A. Frye, Kolin Paul, genotype-to-phenotype prediction of resistance towards ^-lactams in Escherichia coli and Klebsiella Arti Kapil, Umaer Naseer, Punit Kaur, and Rafi Ahmad. 2021. AMR-Diag: Neural network based pneumoniae. Computational and Structural Biotechnology Journal 19 (2021), 1896–1906. https: / / doi.org / 10.1016 / j.csbj.2021.03.027[3]Cristian C. Barros.2021. Neural network-based predictions of antimicrobial resistance in Salmonella spp. using k-mers counting from whole-genome sequences. bioRxiv (2021). https: / / doi.org / 10.1101 / 2021 .08 .10 .455825 a r X i v : h t t p s : / / w w w. b i o r x i v. o rg / c o n t e n t / e a r l y / 2021 / 08 / 11 / 2021.08.10.455825.full.pdf [4] Nicola Cadenelli, Jordà Polo, and David Carrera.2017. Accelerating K-mer Frequency Counting with GPU and Non-Volatile Memory. In 2017 IEEE 19th International Conference on High Performance Computing and Communications; IEEE 15th International Conference on Smart City; IEEE 3rd International Conference on Data Science and Systems(HPCC / SmartCity / DSS).434–441. https: / / doi.org / 10.1109 / HPCC- SmartCity-DSS.2017.57 [5] Marius Erbert, Steffen Rechner, and Matthias Müller-Hannemann.2017. Gerbil: A fast and memoryefficient K-mer counter with GPU-support. Algorithms for Molecular Biology 12, 1 (2017). https: / / doi.org / 10.1186 / s13015-017-0097-9[6]Sunkyu Kim, Heewon Lee, Keonwoo Kim, and Jaewoo Kang.2018. Mut2Vec: Distributed representation of cancerous mutations. BMC Medical Genomics 11, S2 (2018). https: / / doi.org / 10.1186 / s12920-018-0349-7 [7] Marek Kokot, Maciej Długosz, and Sebastian Deorowicz.2017. KMC 3: counting and manipulating k-mer statistics. Bioinformatics 33, 17 (052017), 2759–2761. https: / / doi.org / 10.1093 / bioinformatics / btx304 arXiv:https: / / academic.oup.com / bioinformatics / article-pdf / 33 / 17 / 2759 / 25163905 / btx304_supplementary_kmc_tools.pdf [8] C. Kromer-Edwards, M. Castanheira, and S. Oliveira.2022. K-Mer Fingerprinting with RNN to predict MICs for K. pneumoniae. In 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE Computer Society, Los Alamitos, CA, USA, 1927–1932. https: / / doi.org / 10.1109 / BIBM55620.2022.9995374[9]Cory Kromer-Edwards, Mariana Castanheira, and Suely Oliveira.2023. Using Feature Selection from XGBoost to Predict MIC Values with Neural Networks. In 2023International Joint Conference on Neural Networks (IJCNN) (IJCNN 2023). Queensland, Australia.

[0010] Cory Kromer-Edwards and Suely Oliveira.2023. Finding, and Countering, Future Resistance Using Bacterial Antibiotic Adversarial Genetic Algorithm (BAAGA). In 2023 International Joint Conference on Neural Networks (IJCNN) (IJCNN 2023). Queensland, Australia.

[0011] Luis Lugo and Emiliano Barreto-Hernández.2021. A Recurrent Neural Network approach for whole genome bacteria identification. Applied Artificial Intelligence 35, 9 (2021), 642–656. https: / / doi.org / 10.1080 / 08839514.2021.1922842 arXiv:https: / / doi.org / 10.1080 / 08839514.2021.1922842

[0012] Guillaume Marçais and Carl Kingsford.2011. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27, 6 (012011), 764–770. https: / / doi.org / 10.1093 / bioinformatics / btr011 arXiv:https: / / academic.oup.com / bioinformatics / article-pdf / 27 / 6 / 764 / 48866141 / bioinformatics_27_6_764.pdf

[0013] Patrick Ng.2017. dna2vec: Consistent vector representations of variable-length k-mers. (012017).

[0014] Marcus Nguyen, Thomas Brettin, S. Wesley Long, and et al.2017. Developing an in silico minimum inhibitory concentration panel test for Klebsiella pneumoniae. bioRxiv (2017). https: / / doi.org / 10.1101 / 193797 arXiv:https: / / www.biorxiv.org / content / early / 2017 / 09 / 25 / 193797.full.pdf

[0015] Python.2023. Python documentation. https: / / docs.python.org / 3 / c-api / index.html

[0016] Saaj.2022. Procpath documentation. https: / / procpath.readthedocs.io /

[0017] Ernest Turroand et al.2011. Haplotype and isoform specific expression estimation using multimapping RNA-seq reads.GenomeBiology12,2(Feb2011). https: / / doi.org / 10.1186 / gb-2011-12-2-r13

[0018] Ziye Wang, Shuo Li, Ronghui You, Shanfeng Zhu, Xianghong Jasmine v Zhou, and Fengzhu Sun.2021. ARG-SHINE: improve antibiotic resistance class prediction by integrating sequence homology, functional information and deep convolutional neural network. NAR Genomics and Bioinformatics 3, 3 (082021). https: / / doi.org / 10.1093 / nargab / lqab066 arXiv:https: / / academic.oup.com / nargab / article-pdf / 3 / 3 / lqab066 / 39584656 / lqab066.pdf lqab066. REFERENCES CITED IN THE TEXT OF THE DISCLOSURE [1] https: / / pubchem.ncbi.nlm.nih.gov / compound / Cefepime#section= Canonical-SMILES. Accessed: 2023-1-13. [2] https: / / pubchem.ncbi.nlm.nih.gov / compound / Cefepime#section= 2D-Structure. Accessed: 2023-1-13. [3] https: / / pubchem.ncbi.nlm.nih.gov / compound / Azactam#section= Canonical-SMILES. Accessed: 2023-1-13. [4] https: / / pubchem.ncbi.nlm.nih.gov / compound / Azactam#section= 2D-Structure. Accessed: 2023-1-13.[5] Ahmad Alsahaf et al. “A framework for feature selection through boost-ing”. In: Expert Systems with Applications 187 (2022), p.115895. issn: 0957-4174. doi: https: / / doi.org / 10.1016 / j.eswa.2021.115895. url: https: / / www.sciencedirect.com / science / article / pii / S0957417421012513. [6] E. Anderson, G. D. Veith, and David Weininger. “SMILES: a line notation and computerized interpreter for chemical structures.” In: (1987). [7] G. A. Arango-Argoty et al. “DeepARG: A deep learning approach for predicting antibiotic resistance genes from metagenomic data”. In: bioRxiv (2017). doi: 10.1101 / 149328. eprint: https: / / www.biorxiv. org / content / early / 2017 / 06 / 12 / 149328.full.pdf. url: https: / / www.biorxiv.org / content / early / 2017 / 06 / 12 / 149328.[8]Gustavo Arango-Argoty et al. “Deeparg: A deep learning approach for predicting antibiotic resistance genes from metagenomic data”. In: Microbiome 6.1 (2018). doi: 10.1186 / s40168-018-0401-z. [9] M ARTHUR, P REYNOLDS, and P COURVALIN. “Glycopeptide resistance in Enterococci”. In: Trends in Microbiology 4.10 (1996), pp.401–407. doi: 10.1016 / 0966- 842x(96)10063-9.

[0010] Josep Arús-Pous et al. “Randomized smiles strings improve the quality of molecular generative models”. In: Journal of Cheminformatics.71st ser.11.1 (2019). doi: 10.26434 / chemrxiv.8639942.v1.

[0011] Mostafa Atlam et al. “A New Feature Selection Method for Enhancing Cancer Diagnosis Based on DNA Microarray”. In: 202037th National Radio Science Conference (NRSC). 2020, pp.285–295. doi: 10.1109 / NRSC49500.2020.9235095.

[0012] Ekaterina Avershina et al. “AMR-Diag: Neural network based genotype-to-phenotype prediction of resistance towards β-lactams in Escherichia coli and Klebsiella pneumoniae”. In: Computational and Structural Biotechnology Journal 19 (2021), pp. 1896–1906. issn: 2001-0370. doi: https: / / doi.org / 10.1016 / j.csbj.2021.03.027. url: https: / / www.sciencedirect.com / science / article / pii / S2001037021000994.

[0013] Ekaterina Avershina et al. “AMR-Diag: Neural network based genotype-to-phenotype prediction of resistance towards β-lactams in Escherichia coli and Klebsiella pneumoniae”. In: Computational and Structural Biotechnology Journal 19 (2021), pp. 1896–1906. issn: 2001-0370. doi: https: / / doi.org / 10.1016 / j.csbj.2021.03.027. url: https: / / www.sciencedirect.com / science / article / pii / S2001037021000994.

[0014] Junchen Bao. “Multi-features Based Arrhythmia Diagnosis Algorithm Using Xgboost”. In: 2020 International Conference on Computing and Data Science (CDS).2020, pp. 454–457. doi: 10.1109 / CDS49703.2020.00095.

[0015] Cristian C. Barros. “Neural network-based predictions of antimicrobial resistance in Salmonella spp. using kmers counting from whole-genome sequences”. In: bioRxiv (2021). doi: 10.1101 / 2021.08.10.455825. eprint: https: / / www.biorxiv.org / content / early / 2021 / 08 / 11 / 2021.08.10.455825.full.pdf. url: https: / / www.biorxiv.org / content / early / 2021 / 08 / 11 / 2021.08.10.455825.

[0016] Cristian C. Barros. “Neural network-based predictions of antimicrobial resistance in Salmonella spp. using kmers counting from whole-genome sequences”. In: bioRxiv (2021). doi: 10.1101 / 2021.08.10.455825. eprint: https: / / www.biorxiv.org / content / early / 2021 / 08 / 11 / 2021.08.10.455825.full.pdf. url: https: / / www.biorxiv.org / content / early / 2021 / 08 / 11 / 2021.08.10.455825.

[0017] Syed K. Bashar, Abdullah Al Fahim, and Ki H. Chon. “Smartphone Based Human Activity Recognition with Feature Selection and Dense Neural Network”. In: 202042nd Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC).2020, pp.5888–5891. doi: 10.1109 / EMBC44109.2020.9176239.

[0018] Matteo Bassetti et al. “Antimicrobial resistance in the next 30 years, humankind, bugs and drugs: a visionary approach.” In: Intensive Care Medicine 43.10 (2017), pp.1464– 1475. issn: 0340-0964. doi: 10.1007 / s00134-017-4878-x.

[0019] Marlon Bayot and Bradley Bragg. Antimicrobial Susceptibility Testing. StatPearls. StatPearls Publishing, 2021. url: https: / / www.ncbi.nlm. nih.gov / books / NBK539714 / .

[0020] Kevin Bock et al. “Geneva: Evolving Censorship Evasion Strategies”. In: Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security. CCS ’19. London, United Kingdom: Association for Computing Machinery, 2019, pp.2199– 2214. isbn: 9781450367479. doi: 10.1145 / 3319535.3363189. url: https: / / doi.org / 10.1145 / 3319535.3363189.

[0021] P A Bradford and C C Sanders. “Development of test panel of beta-lactamases expressed in a common Escherichia coli host background for evaluation of new beta-lactam antibiotics”. In: Antimicrobial Agents and Chemotherapy 39.2 (1995), pp.308–313. doi: 10.1128 / AAC.39.2.308. eprint: https: / / journals.asm.org / doi / pdf / 10.1128 / AAC.39. 2.308. url: https: / / journals.asm.org / doi / abs / 10.1128 / AAC.39.2.308.

[0022] Patricia A. Bradford and Mariana Castanheira. Mechanisms of Resistance to Antibacterial Agents.12th Edition. American Society of Microbiology, 2019. doi: 10.1128 / 9781683670438.MCM.ch71. eprint: https : / / clinmicronow . org / doi / pdf / 10 .1128 / 9781683670438 . MCM.ch71. url: https: / / clinmicronow.org / doi / abs / 10.1128 / 9781683670438.MCM.ch71.

[0023] T. A. Brown. Genomes 4. Garland Science, 2018.

[0024] Wanda C Reygaert. “An overview of the antimicrobial resistance mechanisms of bacteria”. In: AIMS Microbiology 4.3 (June 2018), pp.482– 501. doi: 10.3934 / microbiol.2018.3.482.

[0025] Marcelino Campos et al. “Simulating Multilevel Dynamics of Antimicrobial Resistance in a Membrane Computing Model”. In: mBio 10.1 (2019), e02460–18. doi: 10.1128 / mBio.02460-18. eprint: https: / / journals.asm.org / doi / pdf / 10.1128 / mBio.02460- 18. url: https: / / journals.asm.org / doi / abs / 10.1128 / mBio.02460-18.

[0026] Nitesh V. Chawla et al. “SMOTE: Synthetic Minority over-Sampling Technique”. In: J. Artif. Int. Res.16.1 (June 2002), pp.321–357. issn: 1076-9757.

[0027] Tianqi Chen and Carlos Guestrin. “XGBoost: A Scalable Tree Boosting System”. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. KDD ’16. San Francisco, California, USA: ACM, 2016, pp. 785–794. isbn: 978-1-4503-4232-2. doi: 10.1145 / 2939672.2939785. url: http: / / doi.acm.org / 10.1145 / 2939672.2939785.

[0028] Ying Chen et al. “LDNNET: Towards Robust Classification of Lung Nodule and Cancer Using Lung Dense Neural Network”. In: IEEE Access 9 (2021), pp.50301–50320. doi: 10.1109 / ACCESS.2021.3068896.

[0029] Artem Cherkasov et al. “Use of Artificial Intelligence in the Design of Small Peptide Antibiotics Effective against a Broad Spectrum of Highly Antibiotic-Resistant Superbugs”. In: ACS Chemical Biology 4.1 (2009). PMID: 19055425, pp.65–74. doi: 10.1021 / cb800240j. eprint: https: / / doi.org / 10.1021 / cb800240j. url: https: / / doi.org / 10. 1021 / cb800240j.

[0030] Nancy Chinchor. “MUC-4 Evaluation Metrics”. In: Proceedings of the 4th Conference on Message Understanding. MUC4 ’92. McLean, Virginia: Association for Computational Linguistics, 1992, pp.22–29. isbn: 1-55860-273-9. doi: 10.3115 / 1072064.1072067. url: https: / / doi. org / 10.3115 / 1072064.1072067.

[0031] Joana Rosado Coelho et al. “The Use of Machine Learning Methodologies to Analyse Antibiotic and Biocide Susceptibility in Staphylococcus aureus”. In: PLOS ONE 8.2 (Feb.2013), pp.1–10. doi: 10.1371 / journal.pone.0055582. url: https: / / doi.org / 10.1371 / journal. pone.0055582.

[0032] Joana Rosado Coelho et al. “The Use of Machine Learning Methodologies to Analyse Antibiotic and Biocide Susceptibility in Staphylococcus aureus”. In: PLOS ONE 8.2 (Feb.2013), pp.1–10. doi: 10.1371 / journal.pone.0055582. url: https: / / doi.org / 10.1371 / journal. pone.0055582.

[0033] Joana Rosado Coelho et al. “The Use of Machine Learning Methodologies to Analyse Antibiotic and Biocide Susceptibility in Staphylococcus aureus”. In: PLOS ONE 8.2 (Feb.2013), pp.1–10. doi: 10.1371 / journal.pone.0055582. url: https: / / doi.org / 10.1371 / journal. pone.0055582.

[0034] David W. Collins and Thomas H. Jukes. “Rates of Transition and Transversion in Coding Sequences since the Human-Rodent Divergence”. In: Genomics 20.3 (1994), pp.386– 396. issn: 0888-7543. doi: https: / / doi.org / 10.1006 / geno.1994.1192. url: https: / / www.sciencedirect. com / science / article / pii / S088875438471192X.

[0035] M Couturier et al. “Identification and classification of bacterial plasmids”. In: Microbiological Reviews 52.3 (1988), pp.375–395. doi: 10.1128 / mr.52.3.375-395.1988.

[0036] Preethi Devan and Neelu Khare. “An efficient xgboost–DNN-based classification model for network Intrusion Detection System”. In: Neural Computing and Applications 32.16 (2020), pp.12499–12514. doi: 10.1007 / s00521-020-04708-x.

[0037] Robert Ekblom and Jochen B. Wolf. “A field guide to whole-genome sequencing, assembly and Annotation”. In: Evolutionary Applications 7.9 (2014), pp.1026–1042. doi: 10.1111 / eva.12178.

[0038] Christopher D. Fjell et al. “Identification of Novel Antibacterial Peptides by Chemoinformatics and Machine Learning”. In: Journal of Medicinal Chemistry 52.7 (2009). PMID: 19296598, pp.2006–2015. doi: 10.1021 / jm8015365. eprint: https: / / doi.org / 10.1021 / jm8015365. url: https: / / doi.org / 10.1021 / jm8015365.

[0039] U.S. Food and Drug Administration. Class II Special Controls Guidance Document: Antimicrobial Susceptibility Test (AST) Systems. U.S. Food and Drug Administration, 2009.

[0040] F ́elix-Antoine Fortin et al. “DEAP: Evolutionary Algorithms Made Easy”. In: Journal of Machine Learning Research 13 (July 2012), pp.2171– 2175.

[0041] Garrett B. Goh et al. SMILES2VEC: An interpretable general-purpose deep neural network for predicting chemical properties. Mar.2018. url: https: / / arxiv.org / abs / 1712.02034.

[0042] N. C. Gordon et al. “Prediction of Staphylococcus aureus Antimicrobial Resistance by Whole-Genome Sequencing”. In: Journal of Clinical Microbiology 52.4 (2014), pp. 1182–1191. doi: 10.1128 / JCM.03117-13. eprint: https: / / journals.asm.org / doi / pdf / 10.1128 / JCM.03117- 13.url:https: / / journals.asm.org / doi / abs / 10.1128 / JCM.03117-13.

[0043] Antonio Gulli and Sujit Pal. Deep learning with Keras. Packt Publishing Ltd, 2017.

[0044] R. Hakenbeck and J. Coyette. “Resistant penicillin-binding proteins”. In: Cellular and Molecular Life Sciences CMLS 54.4 (1998), pp.332– 340. doi: 10.1007 / s000180050160.

[0045] Md-Nafiz Hamid and Iddo Friedberg. “Transfer learning improves antibiotic resistance class prediction”. In: bioRxiv (2020). doi: 10.1101 / 2020.04.17.047316. eprint: https: / / www.biorxiv.org / content / early / 2020 / 04 / 18 / 2020.04.17.047316.full.pdf. url: https: / / www.biorxiv.org / content / early / 2020 / 04 / 18 / 2020.04.17.047316.

[0046] Stephan Harbarth and Matthew Samore. “Antimicrobial Resistance Determinants and Future Control”. In: Emerging infectious diseases 11 (July 2005), pp.794–801. doi: 10.3201 / eid1106.050167.

[0047] Tin Kam Ho. “Random decision forests”. In: Proceedings of 3rd international conference on document analysis and recognition. Vol.1. IEEE.1995, pp.278–282.

[0048] Cheng-Ping Hsieh et al. “Feature Selection Framework for XGBoost Based on Electrodermal Activity in Stress Detection”. In: 2019 IEEE International Workshop on Signal Processing Systems (SiPS).2019, pp.330–335. doi: 10.1109 / SiPS47522.2019.9020321.

[0049] Mesayu Elida Irawati and Hasballah Zakaria. “Classification Model for Covid-19 Detection Through Recording of Cough Using XGboost Classifier Algorithm”. In: 2021 International Symposium on Electronics and Smart Devices (ISESD).2021, pp.1–5. doi: 10.1109 / ISESD53023.2021.9501695.

[0050] L. C. Jain and L. R. Medsker. Recurrent Neural Networks: Design and Applications.1st. USA: CRC Press, Inc., 1999. isbn: 0849371813.

[0051] Dejun Jiang et al. “Could graph neural networks learn better molecular representation for drug discovery? A comparison study of descriptor-based and graph-based models”. In: Journal of Cheminformatics 13.1 (2021). doi: 10.1186 / s13321-020-00479-8.

[0052] J H Jorgensen. “Selection criteria for an antimicrobial susceptibility testing system.” In: Journal of Clinical Microbiology 31.11 (1993), pp.2841–2844. issn: 0095-1137. eprint: https: / / jcm.asm.org / content / 31 / 11 / 2841.full.pdf. url: https: / / jcm.asm.org / content / 31 / 11 / 2841.

[0053] Sunkyu Kim et al. “Mut2Vec: Distributed representation of cancerous mutations”. In: BMC Medical Genomics 11.S2 (2018). doi: 10.1186 / s12920-018-0349-7.

[0054] Alexandra Klimenko et al. “Modeling evolution of spatially distributed bacterial communities: A simulation with the haploid evolutionary constructor”. In: BMC Evolutionary Biology 15.Suppl 1 (2015). doi: 10.1186 / 1471-2148-15-s1-s3.

[0055] Marek Kokot, Maciej Dlugosz, and Sebastian Deorowicz. “KMC 3: counting and manipulating k-mer statistics”. In: Bioinformatics 33.17 (May 2017), pp.2759–2761. issn: 1367-4803. doi: 10.1093 / bioinformatics / btx304. eprint: https: / / academic.oup.com / bioinformatics / article-pdf / 33 / 17 / 2759 / 25163905 / btx304 \ _supplementary \ _kmc\_tools.pdf. url: https: / / doi.org / 10.1093 / bioinformatics / btx304.

[0056] C. Kromer-Edwards, M. Castanheira, and S. Oliveira. “K-Mer Fingerprinting with RNN to predict MICs for K. pneumoniae”. In: 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). Los Alamitos, CA, USA: IEEE Computer Society, Dec.2022, pp.1927– 1932. doi: 10.1109 / BIBM55620.2022.9995374. url: https: / / doi. ieeecomputersociety.org / 10.1109 / BIBM55620.2022.9995374.

[0057] C. Kromer-Edwards, M. Castanheira, and S. Oliveira. “Year, Location, and Species Information In Predicting MIC Values with Beta-Lactamase Genes”. In: 2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). Los Alamitos, CA, USA: IEEE Computer Society, Dec.2020, pp.1383–1390. doi: 10.1109 / BIBM49941.2020.9313331. url: https: / / doi.ieeecomputersociety.org / 10. 1109 / BIBM49941.2020.9313331.

[0058] C. Kromer-Edwards and S. Oliveira. “Compound RNN to Predict MICs Using K-Mer Fingerprints and Antibiotic SMILES”. Submitted.2023.

[0059] Cory Kromer-Edwards, Mariana Castanheira, and Suely Oliveira. “Using Feature Selection from XGBoost to Predict MIC Values with Neural Networks”. In: 2023 International Joint Conference on Neural Networks (IJCNN) (IJCNN 2023). Queensland, Australia, June 2023.

[0060] Cory Kromer-Edwards and Suely Oliveira. “Compound RNN to Predict MICs Using K- Mer Fingerprints and Antibiotic SMILES”. In: 2023 ACM Conference on Bioinformatics, Computational Biology, and Health Informatics (BCB 2023). Houston, Texas, Sept.2023.

[0061] Cory Kromer-Edwards and Suely Oliveira. “Finding, and Countering, Future Resistance Using Bacterial Antibiotic Adversarial Genetic Algorithm (BAAGA)”. In: 2023 International Joint Conference on Neural Networks (IJCNN) (IJCNN 2023). Queensland, Australia, June 2023.

[0062] Cory Kromer-Edwards, Suely Oliveira, and Mariana Castanheira. Minimum Inhibitory Concentration (mic) Predictions For Four B-lactam Agents For Escherichia Coli And Klebsiella Pneumoniae From A Large Surveillance Program Using Genomic Data And A Machine Learning Model. Abstract presented at ASM Microbe, Washington, DC. June 2022.

[0063] Cory Kromer-Edwards et al. “Identifying Beta-Lactam Resistance with Neural Networks”. In: 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM).2019, pp.1324–1330. doi: 10.1109 / BIBM47256.2019.8983058.

[0064] J. Kulski. “Next-Generation Sequencing — An Overview of the History, Tools, and “Omic” Applications”. In: Next Generation Sequencing Advances, Applications and Challenges.2016.

[0065] A. Kumar et al. “Duration of hypotension before initiation of effective antimicrobial therapy is the critical determinant of survival in human septic shock”. In: Crit. Care Med. 34.6 (June 2006), pp.1589–1596.

[0066] Yongbeom Kwon and Juyong Lee. “Molfinder: An evolutionary algorithm for the global optimization of molecular properties and the extensive exploration of chemical space using smiles”. In: Journal of Chem-informatics 13.1 (2021). doi: 10.1186 / s13321-021- 00501-7.

[0067] Greg Landrum. “RDKit: Open-Source Cheminformatics Software”. In: (). url: https: / / github.com / rdkit / rdkit.

[0068] Steven M LaValle, Michael S Branicky, and Stephen R Lindemann. “On the relationship between classical grid search and probabilistic roadmaps”. In: The International Journal of Robotics Research 23.7-8 (2004), pp.673–692.

[0069] Y. Lecun et al. “Gradient-based learning applied to document recognition”. In: Proceedings of the IEEE 86.11 (1998), pp.2278–2324. doi: 10.1109 / 5.726791.

[0070] Surbhi Leekha, Christine L. Terrell, and Randall S. Edson. “General principles of antimicrobial therapy”. In: Mayo Clinic Proceedings 86.2 (2011), pp.156–167. doi: 10.4065 / mcp.2010.0639.

[0071] Guillaume Lemaˆıtre, Fernando Nogueira, and Christos K. Aridas. “Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning”. In: Journal of Machine Learning Research 18.17 (2017), pp.1–5. url: http: / / jmlr.org / papers / v18 / 16-365.

[0072] Edo Liberty et al. “Elastic Machine Learning Algorithms in Amazon SageMaker”. In: Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data. SIGMOD ’20. Portland, OR, USA: Association for Computing Machinery, 2020, pp.731–737. isbn: 9781450367356. doi: 10.1145 / 3318464.3386126. url: https: / / doi.org / 10.1145 / 3318464.3386126.

[0073] Heidi E. Lischer and Kentaro K. Shimizu. “Reference-guided de novo assembly approach improves genome reconstruction for related species”. In: BMC Bioinformatics 18.1 (2017). doi: 10.1186 / s12859-017-1911-6.

[0074] D M Livermore. “Interplay of impermeability and chromosomal beta-lactamase activity in imipenem-resistant Pseudomonas aeruginosa”. In: Antimicrobial Agents and Chemotherapy 36.9 (1992), pp.2046–2048. doi: 10.1128 / AAC.36.9.2046. eprint: https: / / journals.asm.org / doi / pdf / 10.1128 / AAC.36.9.2046. url: https: / / journals.asm. org / doi / abs / 10.1128 / AAC.36.9.2046.

[0075] Carl Llor and Lars Bjerrum. “Antimicrobial resistance: risk associated with antibiotic overuse and initiatives to reduce the problem”. In: Therapeutic Advances in Drug Safety 5.6 (2014). PMID: 25436105, pp.229– 241. doi: 10.1177 / 2042098614554919. eprint: https: / / doi.org / 10.1177 / 2042098614554919. url: https: / / doi.org / 10.1177 / 2042098614554919.

[0076] Luis Lugo and Emiliano Barreto-Hernandez. “A Recurrent Neural Network approach for whole genome bacteria identification”. In: Applied Artificial Intelligence 35.9 (2021), pp.642–656. doi: 10.1080 / 08839514.2021.1922842. eprint:https: / / doi.org / 10.1080 / 08839514.2021.1922842. url: https: / / doi.org / 10.1080 / 08839514.2021.1922842.

[0077] Mart ́ın Abadi et al. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. Software available from tensorflow.org.2015. url: https: / / www.tensorflow.org / .

[0078] Warren S McCulloch and Walter Pitts. “A logical calculus of the ideas immanent in nervous activity”. In: The bulletin of mathematical biophysics 5.4 (1943), pp.115–133.

[0079] Danesh Moradigaravand et al. “Prediction of antibiotic resistance in Escherichia coli from large-scale pangenome data”. In: PLOS Computational Biology 14.12 (Dec. 2018), pp.1–17. doi: 10.1371 / journal. pcbi.1006258. url: https: / / doi.org / 10.1371 / journal.pcbi.1006258.

[0080] Johan W Mouton et al. “MIC-based dose adjustment: facts and fables”. In: Journal of Antimicrobial Chemotherapy 73.3 (Dec.2017), pp.564– 568. issn: 0305-7453. doi: 10.1093 / jac / dkx427. eprint: https: / / academic.oup.com / jac / article- pdf / 73 / 3 / 564 / 23823576 / dkx427. pdf. url: https: / / doi.org / 10.1093 / jac / dkx427.

[0081] Antonio Mucherino, Petraq J. Papajorgji, and Panos M. Pardalos. “k-Nearest Neighbor Classification”. In: Data Mining in Agriculture. New York, NY: Springer New York, 2009, pp.83–106. isbn: 978-0-387-88615-2. doi: 10.1007 / 978-0-387-88615-2_4. url: https: / / doi.org / 10.1007 / 978-0-387-88615-2_4.

[0082] Bryan Naidenov et al. “Pan-Genomic and Polymorphic Driven Prediction of Antibiotic Resistance in Elizabethkingia”. In: Frontiers in Microbiology 10 (2019), p.1446. issn: 1664-302X. doi: 10.3389 / fmicb. 2019.01446. url: https: / / www.frontiersin.org / article / 10.3389 / fmicb.2019.01446.

[0083] H. C. Neu. “Contribution of beta-lactamases to bacterial resistance and mechanisms to inhibit betalactamases”. In: Am. J. Med.79.5B (Nov.1985), pp.2–12.

[0084] Marcus Nguyen et al. “Developing an in silico minimum inhibitory concentration panel test for Klebsiella pneumoniae”. In: bioRxiv (2017). doi: 10.1101 / 193797. eprint: https: / / www.biorxiv.org / content / early / 2017 / 09 / 25 / 193797.full.pdf. url: https: / / www.biorxiv. org / content / early / 2017 / 09 / 25 / 193797.

[0085] Marcus Nguyen et al. “Developing an in silico minimum inhibitory concentration panel test for Klebsiella pneumoniae”. In: Scientific Reports 8.1 (Jan.2018). doi: 10.1038 / s41598-017-18972-w.

[0086] Marcus Nguyen et al. “Developing an in silico minimum inhibitory concentration panel test for Klebsiella pneumoniae”. In: Scientific Reports 8.1 (Jan.2018). doi: 10.1038 / s41598-017-18972-w.

[0087] Marcus Nguyen et al. “Using Machine Learning To Predict Antimicrobial MICs and Associated Genomic Features for Nontyphoidal Salmonella”. In: Journal of Clinical Microbiology 57.2 (2019). Ed. by Daniel J. Diekema. issn: 0095-1137. doi: 10.1128 / JCM.01260-18. eprint: https: / / jcm.asm.org / content / 57 / 2 / e01260-18.full.pdf. url: https: / / jcm.asm.org / content / 57 / 2 / e01260-18.

[0088] Marcus Nguyen et al. “Using Machine Learning To Predict Antimicrobial MICs and Associated Genomic Features for Nontyphoidal Salmonella”. In: Journal of Clinical Microbiology 57.2 (2019). issn: 0095-1137. doi: 10.1128 / JCM.01260-18. eprint:https: / / jcm.asm.org / content / 57 / 2 / e01260-18.full.pdf. url: https: / / jcm.asm.org / content / 57 / 2 / e01260-18.

[0089] Tom O’Malley et al. KerasTuner. https: / / github.com / keras-team / keras-tuner.2019.

[0090] Adeola Ogunleye and Qing-Guo Wang. “XGBoost Model for Chronic Kidney Disease Diagnosis”. In: IEEE / ACM Transactions on Computational Biology and Bioinformatics 17.6 (2020), pp.2131–2140. doi: 10.1109 / TCBB.2019.2911071.

[0091] Chandra Shekhar Pareek, Rafal Smoczynski, and Andrzej Tretyn. “Sequencing technologies and genome sequencing”. In: Journal of Applied Genetics 52.4 (2011), pp. 413–435. doi: 10.1007 / s13353-011-0057-x. url: https: / / doi.org / 10.1007 / s13353-011- 0057-x.

[0092] Joshua L. Payne et al. “Transition bias influences the evolution of antibiotic resistance in Mycobacterium tuberculosis”. In: PLOS Biology 17.5 (May 2019), pp.1–23. doi: 10.1371 / journal.pbio.3000265. url: https: / / doi.org / 10.1371 / journal.pbio.3000265.

[0093] F. Pedregosa et al. “Scikit-learn: Machine Learning in Python”. In: Journal of Machine Learning Research 12 (2011), pp.2825–2830.

[0094] Keith Poole. “Outer Membranes and efflux: The path to multidrug resistance in gram- negative bacteria”. In: Current Pharmaceutical Biotechnology 3.2 (2002), pp.77–98. doi: 10.2174 / 1389201023378454.

[0095] Didi Liliana Popa et al. “Comparing two different types of deep neural networks to improve accuracy in ECG interpretation”. In: 202125th International Conference on System Theory, Control and Computing (ICSTCC).2021, pp.260–265. doi: 10.1109 / ICSTCC52150.2021.9607110.

[0096] Andrey Prjibelski et al. “Using SPAdes De Novo Assembler”. In: Current Protocols in Bioinformatics 70.1 (2020), e102. doi: https: / / doi.org / 10.1002 / cpbi.102. eprint: https: / / currentprotocols. onlinelibrary.wiley.com / doi / pdf / 10.1002 / cpbi.102. url: https: / / currentprotocols.onlinelibrary.wiley.com / doi / abs / 10.1002 / cpbi.102.

[0097] J. R. Quinlan. “Induction of Decision Trees”. In: Mach. Learn.1.1 (Mar.1986), pp.81– 106. issn: 0885-6125. doi: 10.1023 / A:1022643204877. url: https: / / doi.org / 10.1023 / A:1022643204877.

[0098] L. Barth Reller et al. “Antimicrobial Susceptibility Testing: A Review of General Principles and Contemporary Practices”. In: Clinical Infectious Diseases 49.11 (Dec. 2009), pp.1749–1755. issn: 1058-4838. doi: 10.1086 / 647952. eprint: https: / / academic.oup.com / cid / article-pdf / 49 / 11 / 1749 / 984933 / 49-11-1749.pdf. url: https: / / doi.org / 10.1086 / 647952.

[0099] Xiuzhi Sang et al. “HMMPred: Accurate Prediction of DNA-Binding Proteins Based on HMM Profiles and XGBoost Feature Selection”. In: Computational and Mathematical Methods in Medicine 2020 (2020).

[0100] Johannes Schmidt-Hieber. “Nonparametric regression using deep neural networks with ReLU activation function”. In: The Annals of Statistics 48.4 (2020), pp.1875–1897. doi: 10.1214 / 19-AOS1875. url: https: / / doi.org / 10.1214 / 19-AOS1875.

[0101] Nitish Srivastava et al. “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”. In: J. Mach. Learn. Res.15.1 (Jan.2014), pp.1929–1958. issn: 1532-4435.

[0102] Nicholas Stoler and Anton Nekrutenko. “Sequencing error profiles of Illumina sequencing instruments”. In: NAR Genomics and Bioinformatics 3.1 (Mar.2021). lqab019. issn: 2631-9268. doi: 10.1093 / nargab / lqab019. eprint: https: / / academic.oup.com / nargab / article-pdf / 3 / 1 / lqab019 / 36763060 / lqab019.pdf. url: https: / / doi.org / 10.1093 / nargab / lqab019.

[0103] Michelle Su et al. “Genome-Based Prediction of Bacterial Antibiotic Resistance”. In: Journal of Clinical Microbiology 57.3 (2019), e01405– 18. doi: 10.1128 / JCM.01405-18. eprint: https: / / journals.asm. org / doi / pdf / 10.1128 / JCM.01405-18. url: https: / / journals. asm.org / doi / abs / 10.1128 / JCM.01405-18.

[0104] V Sugumaran, V Muralidharan, and Ramachandran K I. “Feature selection using Decision Tree and classification through Proximal Support Vector Machine for fault diagnostics of roller bearing”. In: Mechanical Systems and Signal Processing 21 (Feb. 2007), pp.930–942. doi: 10.1016 / j.ymssp.2006.05.004.

[0105] Aubin Thomas et al. “Gecko is a genetic algorithm to classify and explore high throughput sequencing data”. In: Communications Biology 2.1 (2019). doi: 10.1038 / s42003-019-0456-9.

[0106] J. Turnidge, G. Kahlmeter, and G. Kronvall. “Statistical characterisation of bacterial wild-type mic value distributions and the determination of epidemiological cut-off values”. In: Clinical Microbiology and Infection 12.5 (2006), pp.418–425. doi: 10.1111 / j.1469-0691.2006.01377.x.

[0107] Pieter-Jan Van Camp, David B. Haslam, and Aleksey Porollo. “Prediction of Antimicrobial Resistance in Gram-Negative Bacteria From Whole-Genome Sequencing Data”. In: Frontiers in Microbiology 11 (2020), p.1013. issn: 1664-302X. doi: 10.3389 / fmicb.2020.01013. url: https: / / www.frontiersin.org / article / 10.3389 / fmicb. 2020.01013.

[0108] Pieter-Jan Van Camp, David B. Haslam, and Aleksey Porollo. “Prediction of Antimicrobial Resistance in Gram-Negative Bacteria From Whole-Genome Sequencing Data”. In: Frontiers in Microbiology 11 (2020), p.1013. issn: 1664-302X. doi: 10.3389 / fmicb.2020.01013. url: https: / / www.frontiersin.org / article / 10.3389 / fmicb. 2020.01013.

[0109] Ziye Wang et al. “ARG-SHINE: improve antibiotic resistance class prediction by integrating sequence homology, functional information and deep convolutional neural network”. In: NAR Genomics and Bioinformatics 3.3 (Aug.2021). lqab066. issn: 2631- 9268. doi: 10.1093 / nargab / lqab066. eprint: https: / / academic.oup.com / nargab / article- pdf / 3 / 3 / lqab066 / 39584656 / lqab066.pdf. url: https: / / doi.org / 10.1093 / nargab / lqab066.

[0110] David Weininger. “SMILES, a chemical language and information system.1. Introduction to methodology and encoding rules”. In: Journal of Chemical Information and Computer Sciences 28.1 (1988), pp.31– 36. doi: 10.1021 / ci00057a005. eprint: https: / / doi.org / 10.1021 / ci00057a005. url: https: / / doi.org / 10.1021 / ci00057a005.

[0111] Melvin P. Weinstein. Methods for dilution antimicrobial susceptibility tests for bacteria that grow aerobically. Clinical Laboratory Standards Institute, 2018.

[0112] Melvin P. Weinstein. Performance standards for antimicrobial susceptibility testing. Clinical and Laboratory Standards Institute, 2021.

[0113] Matthew A. Wikler. Development of in vitro susceptibility testing criteria and quality control parameters. NCCLS, 2018.

[0114] Dennis L. Wilson. “Asymptotic Properties of Nearest Neighbor Rules Using Edited Data.” In: IEEE Trans. Systems, Man, and Cybernetics 2.3 (1972), pp.408–421. url: http: / / dblp.uni-trier.de / db / journals / tsmc / tsmc2.html / Wilson72.

[0115] B. J. Winer, D. R. Brown, and K. M Michels. Statistical principles in experimental design. Vol.2. McGrawHill New York, 1971.

[0116] Jason H. Yang et al. “A White-Box Machine Learning Approach for Revealing Antibiotic Mechanisms of Action”. In: Cell 177.6 (2019), 1649–1661.e9. issn: 0092-8674. doi: https: / / doi.org / 10.1016 / j.cell.2019.04.016. url: https: / / www.sciencedirect.com / science / article / pii / S0092867419304027.

[0117] Qianqian Yao, Simone Ludwig, and Mingao Yuan. “COMPARISON OF NON- LEARNED AND LEARNED MOLECULE REPRESENTATIONS FOR CATALYST DISCOVERY”. In: 2022.

[0118] Qinglong Zeng et al. “Models of microbiome evolution incorporating host and Microbial Selection”. In: Microbiome 5.1 (2017). doi: 10.1186 / s40168-017-0343-x.

[0119] Yudong Zhang et al. “Binary PSO with Mutation Operator for Feature Selection Using Decision Tree Applied to Spam Detection”. In: Know.-Based Syst.64.1 (July 2014), pp. 22–31. issn: 0950-7051. doi: 10.1016 / j.knosys.2014.03.015. url: https: / / doi.org / 10.1016 / j. knosys.2014.03.015.

Claims

WHAT IS CLAIMED IS:

1. A system comprising: a processor; and a non-transitory computer readable medium that stores executable instructions for determining an expected potency of a drug in treating a non-human animal pathogen, the executable instructions comprising: a sampler interface that receives sequence data representing the non-human animal pathogen and generates a dataset comprising one of raw nucleic acid sequence data, ribonucleic acid sequence data, or amino acid sequence data; a pathogen characterizer that generates a representation of the non-human animal pathogen from the dataset; and a potency analyzer receives the representation of the non-human animal pathogen from the pathogen characterizer and a representation of the drug and outputs a value representing a potency of the drug as applied to the non-human animal pathogen.

2. The system of claim 1, wherein the representation of the drug is a first representation and the system further comprises a drug characterizer that receives one of an identity or a second representation of the drug and generates the first representation of the drug.

3. The system of claim 2, wherein the drug characterizer is implemented using a large language model that is trained on a sets of sample data each including a SMILES representation of the drug and a representation of the drug appropriate for analysis at the potency analyzer.

4. The system of claim 2, wherein the wherein the drug characterizer is implemented using a graph neural network that is trained on a sets of sample data each including an input representing a molecular structure of the therapeutic as a graph and a representation of the drug appropriate for analysis at the potency analyzer 5. The system of claim 1, wherein the non-human animal pathogen is a bacterium and the drug is an antibiotic.

6. The system of claim 1, wherein the non-human animal pathogen is a fungus and the drug is an antifungal.

7. The system of claim 1, wherein the non-human animal pathogen is a virus and the drug is an antiviral.

8. The system of claim 1, wherein the sampler interface receives sequences representing raw DNA and generates DNA contigs from the sequences representing raw DNA.

9. The system of claim 8, wherein the sampler interface generates the DNA contigs using reference-guided construction.

10. The system of claim 8, wherein the sampler interface comprises a quality control component, implemented as a machine learning model trained on a set of samples each including a set of raw DNA and a label representing the quality of DNA contigs generated from the set of raw DNA, that determines if the raw DNA meets a threshold quality for generating DNA contigs.

11. The system of claim 10, wherein the quality control component is implemented using a large language model.

12. The system of claim 1, wherein the pathogen characterizer generates a count of unique K- mers from the dataset, normalizes the K-mer counts according to a total count, and generates a matrix including the K-mer counts.

13. The system of claim 1, wherein the pathogen characterizer is implemented using a large language model that is trained on a sets of sample data each including a set of DNA or protein data as well as an appropriate representation of the non-human animal pathogen based on the set of DNA or protein data.

14. The system of claim 1, further comprising a species predictor determines a species of the non-human animal pathogen according to the data provided from the sampler interface.

15. The system of claim 14, further comprising a pathogen predictor implemented as a machine learning model trained on a set of training samples each including a set of raw DNA sequences from a bodily fluid or tissue of an animal, a type of the bodily fluid or tissue, a location of an infection, and one of a plurality of pathogen classes, wherein the pathogen classes include a first class representing bacteria, a second class representing viruses, a third class representing fungi, and a fourth class representing no detected pathogen, the species predictor determining the speciesof the non-human animal pathogen according to the data provided from the sampler interface and a determine class of the non-human animal pathogen at the pathogen predictor.

16. The system of claim 1, wherein the potency analyzer includes a pathogen validity model that determines if the pathogen represented by the representation of the non-human animal pathogen could represent a living organism and a drug validity model that determines if the drug represented by the representation of the drug could represent a viable molecule, the potency analyzer outputting the value representing a potency of the drug, a value representing a validity of the representation of the pathogen, and a value representing a validity of the drug.

17. The system of claim 15, wherein the potency analyzer includes at least one attention layer that provides a set of importance values indicating regions of the representation of the non-human animal pathogen and the representation of the drug that are important in determining the potency of the drug as applied to the non-human animal pathogen.

18. A system comprising: a processor; and a non-transitory computer readable medium that stores executable instructions for generating novel drugs and pathogens, the executable instructions comprising: a non-human animal pathogen simulator that generates a representation of a pathogen for analysis; a drug simulator that generates a representation of a drug for analysis; and a potency analyzer that determines a potency for the drug as applied to the non- human animal pathogen.

19. The system of claim 18, wherein each of the pathogen simulator and the drug simulator are implemented as genetic algorithms that operate adversarially, such that the pathogen simulator iteratively selects for a pathogen that minimizes the potency, and the therapeutic simulator iteratively selects for a therapeutic that maximizes the potency.

20. The system of claim 18, wherein the potency analyzer includes a pathogen validity model that determines if the pathogen represented by the representation of the non-human animal pathogen could represent a living organism and a drug validity model that determines if the drug represented by the representation of the drug could represent a viable molecule, the potencyanalyzer outputting the value representing a potency of the drug, a value representing a validity of the representation of the pathogen, and a value representing a validity of the drug.

21. The system of claim 20, wherein the potency analyzer includes at least one attention layer that provides a set of importance values indicating regions of the representation of the non-human animal pathogen and the representation of the drug that are important in determining the potency of the drug as applied to the non-human animal pathogen.

22. The system of claim 18, wherein the pathogen simulator is implemented as a mutation learning model that provides a pathogen that provides target potency for a given drug, the pathogen simulator receiving a dataset representing an initial pathogen, a species of the initial pathogen, a representation of the therapeutic, and a target potency and providing the pathogen as a mutation dataset, comprising one of raw DNA, RNA, or protein sequences, representing a pathogen that provides the target potency.

23. The system of claim 18, wherein the drug simulator is implemented as a drug tuning model that provides a drug that provides target potency for a given pathogen, the drug simulator receives a representation of an initial drug, a representation of the pathogen, and a target potency and provides the drug as a modified version of the initial drug that provides the target potency.

24. The system of claim 18, further comprising a mutation evaluator that estimates the time necessary for a pathogen to mutate from a first DNA sequence to a second, mutated DNA sequence.

25. The system of claim 18, wherein the non-human animal pathogen is a bacterium and the drug is an antibiotic.

26. The system of claim 18, wherein the non-human animal pathogen is a fungus and the drug is an antifungal.

27. The system of claim 18, wherein the non-human animal pathogen is a virus and the drug is an antiviral.

Citation Information

Patent Citations

  • Determining drug effectiveness ranking for a patient using machine learning

    US20200227176A1

  • Systems and methods for multi-label cancer classification

    US20210142904A1

  • Hybrid computational system of classical and quantum computing for drug discovery and methods

    US20210287773A1

  • Computer representations of peptides for efficient design of drug candidates

    WO2023200866A1