Systems and methods for predicting antibiotic resistance from bacterial genomes

The PARP model uses a machine learning system trained on genomic data to predict antibiotic resistance, offering rapid and accurate predictions across bacterial species, improving healthcare outcomes and reducing economic burdens.

JP2025538115APending Publication Date: 2025-11-26BOARD OF RGT THE UNIV OF TEXAS SYST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025524777
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-26
Filing Date
2023-10-25
Publication Date
2025-11-26

AI Technical Summary

Technical Problem

Traditional diagnostic methods for antibiotic resistance, such as culture followed by antibiotic susceptibility testing, are time-consuming and inefficient, leading to prolonged hospital stays, increased mortality, and significant economic burdens due to the development of antibiotic-resistant pathogens.

Method used

A machine learning system, such as the Pan-Antibiotic Resistance Prediction (PARP) model, trained on genomic information and antibiotic signatures, uses advanced deep learning algorithms to predict antibiotic resistance or susceptibility across various bacterial species, even for pathogens not included in the training dataset, leveraging cloud-based infrastructure for scalability and affordability.

Benefits of technology

The PARP model provides high predictive accuracy, enabling rapid and accurate antibiotic resistance predictions, reducing the time and cost associated with traditional methods and addressing the public health threat of antibiotic-resistant infections.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025538115000001_ABST
    Figure 2025538115000001_ABST
Patent Text Reader

Abstract

The systems, methods, and devices disclosed herein provide antibiotic resistance prediction using a pan-antibiotic resistance prediction (PARP) model. The PARP model includes a machine learning system trained with a training dataset of genetic information associated with multiple bacterial species and / or antibiotic signature information associated with multiple antibiotics. The PARP model is deployed to a cloud-based service for scalability, providing access to the PARP model for clinic, hospital, and / or laboratory devices. For example, a web-based portal of the cloud-based service receives genomic sequences associated with specific bacterial isolates uploaded via a remote device. The PARP model outputs a predictive index of antibiotic resistance for the specific bacterial isolate. The predictive index may include a bar graph (e.g., displayed in a graphical user interface) showing susceptibility / resistance predictions for multiple antibiotics for the specific bacterial isolate.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 381,086, filed October 26, 2022, entitled "SYSTEMS AND METHODS FOR PREDICTION OF ANTIBIOTIC RESISTANCE FROM BACTERIAL GENOMES," the entire contents of which are incorporated herein by reference.

[0002] Government support approval This invention was made with government support under Grant No. AI169298 awarded by the National Institutes of Health and Grant Nos. W81XWH-20-1-0149, PR192594 awarded by the U.S. Department of Defense. The government has certain rights in this invention.

[0003] 1.Technical Field Aspects of the technology disclosed herein relate generally to systems and methods for predicting antibiotic resistance, and more particularly to systems and methods for predicting antibiotic resistance based on genomic information. [Background technology]

[0004] 2. Description of Related Technology Antibiotic resistance (AR) is a public health threat. Each year in the United States, at least 2.8 million people become infected with antibiotic-resistant bacteria or fungi, and more than 35,000 die as a result. The development of resistance among pathogens complicates the delivery of effective treatment, leading to longer hospital stays, costly alternative therapies, and increased mortality. Economically, the total economic burden of AR infections could reach up to $20 billion annually in healthcare costs and $35 billion annually in lost productivity. Traditional diagnostic methods rely on culture followed by antibiotic susceptibility testing (AST), which can take days to weeks to complete. Summary of the Invention

[0005] The systems, methods, and devices disclosed herein can address the aforementioned problems. For example, a method for antibiotic resistance prediction may include receiving, in a machine learning system, genetic information associated with bacteria, where the species of the bacteria is one of a plurality of bacterial species; receiving, in the machine learning system, an indication of one of a plurality of antibiotics, where the machine learning system is trained for the plurality of antibiotics using the genetic information associated with the plurality of bacterial species; and / or outputting, in the machine learning system, an indication of antibiotic resistance or susceptibility associated with the bacteria and antibiotic received in the machine learning system.

[0006] In some examples, the machine learning system can be trained using protein sequences associated with multiple bacterial species. The machine learning system can include a feature-based linear modulation (FiLM) machine learning system. Additionally, the machine learning system can jointly model antibiotics and bacterial variants. Additionally, the machine learning system can include a rectified linear activation function (ReLU) layer, a batch normalization layer, and / or a dropout layer.

[0007] In some examples, a method for antibiotic resistance prediction includes training a pan-antibiotic resistance prediction (PARP) model by providing one or more training datasets to a machine learning system, the training dataset including genetic information associated with a plurality of bacterial species and / or antibiotic signature information associated with a plurality of antibiotics. The method also includes receiving, in the machine learning system, a genomic sequence associated with a particular bacterial isolate, and / or outputting, in the machine learning system, a predictive indicator of antibiotic resistance associated with the one or more antibiotics for the particular bacterial isolate.

[0008] In some examples, the method further includes performing a data preparation procedure on one or more training datasets by one-hot encoding the antibiotic feature information. Furthermore, the one or more training datasets may include at least one of an isolate-variant matrix, an antibiotic indicator matrix, or an isolate resistance signature vector. The method may also include performing a nested cross-validation procedure on the PARP model using a validation dataset including multiple bacteria-antibiotic combinations. Furthermore, the method may include determining, by the PARP model, one or more classes associated with the multiple antibiotics or multiple bacterial species using weights from one or more dense layers of a feature-based linear modulation (FiLM) generator to form clusters, and the machine learning system may output a predictive indicator of antibiotic resistance using the one or more classes. The PARP model may also be deployed on a container orchestration service such that the PARP model provides a cloud-based antibiotic resistance prediction service.

[0009] In some examples, receiving the genome sequence can include receiving an upload from a remote device in a cloud-based antibiotic resistance prediction service. The genome sequence can correspond to a bacterial species not included in the one or more training datasets. Furthermore, the antibiotic resistance predictive indicator can include a bar graph for presentation on a graphical user interface (GUI) of the computing device that provided the genome sequence. Furthermore, the x-axis of the bar graph can represent different antibiotics, and the y-axis of the bar graph can represent predicted values ​​of resistance or susceptibility to the different antibiotics. The PARP model can generate shared and unique variant data indicating one or more variants shared between different bacterial species and one or more variants unique to the different bacterial species. Training the PARP model can include generating paired antibiotic susceptibility data based on testing the isolates against antibiotic pairs that indicate a shared pathway for the antibiotic pair. The method can also include performing a predictive accuracy assessment for the predictive indicator, where the predictive accuracy assessment outputs one or more predictive accuracy values ​​corresponding to the one or more antibiotics. Furthermore, the method can include tuning multiple hyperparameters of the PARP model, where the multiple hyperparameters include the number of dense blocks, the number of layers or feature stacking blocks, and / or the geometric size of the dense layers.

[0010] In some examples, a system for antibiotic resistance prediction includes a Pan-Antibiotic Resistance Prediction (PARP) model deployed on a cloud-based service, the PARP model having a machine learning system trained with one or more training datasets including genetic information associated with a plurality of bacterial species and / or antibiotic signature information associated with a plurality of antibiotics. The system also includes a web-based portal for receiving genomic sequences associated with particular bacterial isolates and providing the genomic sequences to the PARP model, and / or a predictive indicator of antibiotic resistance for the particular bacterial isolate output by the PARP model and configured for display in a graphical user interface (GUI) of a computing device.

[0011] Other embodiments are described and referenced herein. Moreover, while multiple embodiments are disclosed, still other embodiments of the disclosed technology will become apparent to those skilled in the art from the following detailed description, which shows and describes illustrative embodiments of the disclosed technology. As will be recognized, the disclosed technology can be modified in various aspects without departing from the spirit and scope of the disclosed technology. Accordingly, the drawings and detailed description are to be regarded as illustrative in nature and not restrictive. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 illustrates an example of a computing device in accordance with certain aspects of the disclosed techniques. [Figure 2] FIG. 1 illustrates an example of an antibiotic resistance prediction system implemented using a machine learning model, according to certain aspects of the present disclosure. [Figure 3] FIG. 1 is a block diagram illustrating an example of a machine learning system in accordance with certain aspects of the present disclosure. [Figure 4] FIG. 1 is a flow diagram illustrating example operations for antibiotic resistance prediction in accordance with certain aspects of the disclosed technology. [Figure 5A]FIG. 1 illustrates an example workflow of an antibiotic resistance prediction system in accordance with certain aspects of the disclosed technology. [Figure 5B] FIG. 1 illustrates an example antibiotic resistance heat map of an antibiotic resistance prediction system, in accordance with certain aspects of the disclosed technology. [Figure 5C] FIG. 1 illustrates an example of a mutation graph for an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 6A] FIG. 1 illustrates an example of an evaluation of the predictive accuracy of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 6B] FIG. 1 illustrates an example of an evaluation of the predictive accuracy of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 7A] FIG. 1 illustrates an example of a classification graph of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 7B] FIG. 1 shows an example of an agglomerative clustering dendrogram of an antibiotic resistance prediction system, in accordance with certain aspects of the disclosed technology. [Figure 7C] FIG. 1 illustrates an example of a classification graph of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 7D] FIG. 1 shows an example of an agglomerative clustering dendrogram of an antibiotic resistance prediction system, in accordance with certain aspects of the disclosed technology. [Figure 8A] FIG. 1 illustrates an example of a cloud-based deployment of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 8B] FIG. 10 illustrates an example bar graph of predicted output results of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 9A] FIG. 1 shows an example of a unique / shared mutation bar graph of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 9B] FIG. 1 shows an example of a unique / shared mutation bar graph of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 10A]FIG. 1 shows an example of a unique / shared mutation bar graph of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 10B] FIG. 1 shows an example of a multiple unique / shared mutations bar graph of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 10C] FIG. 1 shows an example of a multiple unique / shared mutations bar graph of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 10D] FIG. 1 shows an example of a multiple unique / shared mutations bar graph of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 11] FIG. 1 illustrates an antibiotic evaluation engine of an antibiotic resistance prediction system, in accordance with certain aspects of the disclosed technology. [Figure 12A] FIG. 1 illustrates an example of a nested cross-validation procedure for an antibiotic resistance prediction system, in accordance with certain aspects of the disclosed technology. [Figure 12B] FIG. 1 illustrates an example of a nested cross-validation procedure for an antibiotic resistance prediction system, in accordance with certain aspects of the disclosed technology. [Figure 13] FIG. 1 illustrates an example of a comparison of the prediction output of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 14] FIG. 1 illustrates an example of an evaluation of the predictive accuracy of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 15] FIG. 1 illustrates an example of an evaluation of the predictive accuracy of an antibiotic resistance prediction system, according to certain aspects of the disclosed technology. [Figure 16] FIG. 1 illustrates an example of an antibiotic resistance prediction system with a feature-based linear modulation (FiLM) generator, in accordance with certain aspects of the disclosed technology. DETAILED DESCRIPTION OF THE INVENTION

[0013] Upon review of the entire disclosure, it will be apparent to one of ordinary skill in the art that the steps shown in the above figures may be performed in a different order than described, and that one or more of the steps shown in these figures may be optional.

[0014] Certain aspects of the disclosed technology are directed to methods and systems for predicting antibiotic resistance based on genomic information. The antibiotic resistance prediction system described herein may include a machine learning system for in-silico antibiotic resistance determination. The development of the machine learning system can be based on the curation of bacterial isolates (e.g., over 3,000 bacterial isolates) and various antibiotics (e.g., 29 antibiotics). The machine learning system can provide high predictive performance by using advanced deep learning algorithms. As a cloud-native solution, the antibiotic resistance prediction system offers scalability and affordability. Furthermore, the antibiotic resistance prediction system is a pathogen-independent prediction algorithm and can predict antibiotic resistance against any pathogen whose genome has been sequenced, even if the pathogen is not included in the training dataset.

[0015] Further benefits and advantages of the techniques of this disclosure will become apparent from the detailed description that follows.

[0016] 1 illustrates an example of a computing device 100 in accordance with certain aspects of the disclosed techniques. Computing device 100 may include a processor 103 that controls the overall operation of computing device 100 and associated components, including input / output devices 109, a communication interface 111, and / or a memory 115. A data bus may interconnect processor 103, memory 115, I / O devices 109, and / or communication interface 111.

[0017] The input / output (I / O) devices 109 may include a microphone, keypad, touchscreen, and / or stylus through which a user of the computing device 100 can provide input, and may also include one or more of a speaker for providing audio output and a video display device for providing textual, audiovisual, and / or graphical output. Software may be stored within the memory 115 to provide instructions to the processor 103 to enable the computing device 100 to perform various operations. For example, the memory 115 may store software used by the computing device 100, such as an operating system 117, application programs 119, and / or associated internal databases 121. The various hardware memory units within the memory 115 may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The memory 115 may include one or more physical persistent memory devices and / or one or more non-persistent memory devices. Memory 115 may include, but is not limited to, random access memory (RAM), read-only memory (ROM), electronically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or any other medium that can be used to store desired information and that can be accessed by processor 103.

[0018] Communications interface 111 may include one or more transceivers, digital signal processors, and / or additional circuitry and software for communicating over any network, wired or wireless, using any of the protocols described herein. Processor 103 may include a single central processing unit (CPU), which may be a single-core or multi-core processor (e.g., dual-core, quad-core, etc.), or may include multiple CPUs. Processor 103 and associated components may enable computing device 100 to execute a sequence of computer-readable instructions and perform some or all of the processes described herein. Although not shown in FIG. 1 , various elements within memory 115 or other components within computing device 100 may include one or more caches, such as a CPU cache used by processor 103, a page cache used by operating system 117, a disk cache on a hard drive, and / or a database cache used to cache the contents of database 121. For embodiments that include a CPU cache, the CPU cache may be used by one or more processors 103 to reduce memory latency and access time. The processor 103 can retrieve data from and write data to the CPU cache rather than reading from and writing to memory 115, thereby increasing the speed of these operations. In some examples, a database cache can be created in which certain data from the database 121 is cached in a separate, smaller database in a memory separate from the database, such as RAM or another computing device. For example, in a multi-tier application, a database cache on the application server can reduce the time for data retrieval and data manipulation by not requiring communication over the network with a back-end database server.These types of caches and others can be included in various embodiments and can provide potential benefits in certain embodiments of the software deployment system, such as faster response times and reduced dependency on network conditions when sending and receiving data.

[0019] In certain aspects of the present disclosure, computing device 100 may include machine learning model 120. Machine learning model 120 may, in some embodiments, be implemented as part of processor 103. Machine learning model 120 may be trained using training circuitry 122. Machine learning model 120 may be trained to predict antibiotic resistance based on genomic information across various bacterial species, thereby forming a pan-antibiotic resistance prediction (PARP) model 504, described in more detail below. In some aspects, computing device 100 may be implemented on a network (e.g., on a server) to perform antibiotic resistance prediction on the cloud.

[0020] 2 illustrates an example of an antibiotic resistance prediction system 200 implemented using the machine learning model 120 according to certain embodiments of the present disclosure. The antibiotic resistance prediction system 200 may be implemented on the cloud and provides an interface for a user to interact with the antibiotic resistance prediction system 200 to predict antibiotic resistance.

[0021] Genomic information for any bacterium (e.g., across various bacterial species) can be provided to antibiotic resistance prediction system 200. In some embodiments, a specific antibiotic can be provided to antibiotic resistance prediction system 200. Using a trained machine learning model (machine learning system), antibiotic resistance prediction system 200 can predict the resistance level of the bacterium to the antibiotic provided to antibiotic resistance prediction system 200. In some embodiments, input to the machine learning model can be the consensus protein sequence of the translated DNA of the bacterium and characterization / definition of protein variants.

[0022] To train a machine learning model, entire bacterial genomes of various bacterial species can be provided for training. In some cases, specific antibiotic resistance genes can be used to train the model. Thus, the trained machine learning model can provide resistance predictions based on input of any genome of any pathogen (e.g., the model is not limited to a specific bacterial species or pathogen, but can provide resistance predictions for any bacteria of any bacterial species input to the model). The machine learning model can predict resistance to multiple antibiotics and identify new potential resistance genes and / or mutations in genes important for resistance. The machine learning system can be implemented using a feature-based linear modulation (FiLM) generator deep learning technique to generate multiple layers and blocks containing specific optimization parameters, as described in more detail below. The systems disclosed herein use the FiLM machine learning system. Additionally or alternatively, other modeling systems, such as conditional batch normalization, gated layers, cross-modal fusion, and / or attention layers, can be included.

[0023] FIG. 3 is a block diagram illustrating an example of a machine learning system 300 according to certain embodiments of the present disclosure. The machine learning system 300 can form at least a portion of any of the antibiotic resistance prediction systems 200 and 500-1600 discussed herein. The machine learning system 300 can be used to implement the antibiotic resistance prediction system 200. The machine learning system 300 can be a FiLM machine learning model. The machine learning system 300 can use a deep learning model to jointly model variants and antibiotics, as shown. The machine learning system 300 includes dense block dimensions. Multiple FiLM blocks are optimized using nested cross-validation. The model includes dense / fully connected layers (e.g., linear operations on the layer's input vectors). For activation, the machine learning system 300 can include a rectified linear unit (ReLU) layer. An activation function can transform the total weighted input from a node into the node's activation or output for that input. The ReLU layer can be a piecewise linear function that outputs the input as is if the input is positive, and outputs 0 otherwise.

[0024] The machine learning system 300 may also include a batch normalization layer and a dropout layer, as shown. Batch normalization may be used to make model training faster and more stable by normalizing the layer's inputs through recentering and rescaling. A dropout layer may be used to ignore units (e.g., neurons) during the training phase for a specific set of neurons, which may be randomly selected. For example, these units may not be considered during a particular forward or backward pass. Dropout may reduce interdependent learning between neurons.

[0025] As described, the machine learning system 300 may be trained across two or more species of bacteria. The machine learning system 300 may be trained using protein variants or amino acid variants, as described. For example, in some embodiments, instead of using DNA sequences, the machine learning system 300 may be trained using translated protein variants.

[0026] 4 is a flow diagram illustrating an example of operations 400 for antibiotic resistance prediction according to certain aspects of the disclosed technology. Operations 400 may be performed, for example, by computing system 100, machine learning system 300, and / or any of antibiotic resistance prediction systems 200 and 500-1600.

[0027] At block 402, the computing system can receive, at the machine learning system, genetic information associated with a bacterium, where the species of the bacterium is one of a plurality of bacterial species. At block 404, the computing system can receive, at the machine learning system, an antibiotic instruction for one of a plurality of antibiotics, where the machine learning system is trained for the plurality of antibiotics using the genetic information associated with the plurality of bacterial species. At block 406, the computing system can output, at the machine learning system, an antibiotic resistance instruction associated with the bacterium and antibiotic received at the machine learning system.

[0028] In some embodiments, the machine learning system is trained using protein sequences associated with multiple bacterial species. In some embodiments, the machine learning system includes a feature-based linear modulation (FiLM) machine learning system. The machine learning system may jointly model antibiotics and bacterial variants. The machine learning system may include a rectified linear activation function (ReLU) layer, a batch normalization layer, and a dropout layer.

[0029] These and various other configurations are described in more detail herein. As will be appreciated by those skilled in the art upon reading the following disclosure, various aspects described herein can be methods, computer systems, or computer program products. Accordingly, these aspects can take the form of entirely hardware embodiments, entirely software embodiments, or at least one embodiment combining software and hardware aspects. Furthermore, such aspects can take the form of a computer program product stored on one or more computer-readable storage media (e.g., non-transitory computer-readable media) having computer-readable program code or instructions contained therein or thereon. Any suitable computer-readable storage medium can be used, including hard disks, CD-ROMs, optical storage devices, magnetic storage devices, and / or any combination thereof. Furthermore, various signals representing data or events described herein can be transmitted between sources and destinations in the form of electromagnetic waves traveling over signal-conducting media, such as metal wires, optical fibers, and / or wireless transmission media (e.g., air and / or outer space).

[0030] As mentioned above, embodiments of the disclosed technology include various steps described herein. These steps may be performed by hardware components or may be embodied in machine-executable instructions that can be used to cause a special-purpose processor programmed with the instructions to perform the steps. Alternatively, the steps may be performed by a combination of hardware, software, and / or firmware.

[0031] 5A-5C illustrate an example of an antibiotic resistance prediction system 500 according to certain embodiments of the disclosed technology.

[0032] 5A, the antibiotic resistance prediction system 500 can include a workflow 502 for determining shared genetic features between different pathogen species. The antibiotic resistance prediction system 500 can be a pan-antibiotic resistance prediction model, or PARP model 504, for predicting resistance across a wide variety of pathogens, including pathogens not previously analyzed by the PARP model 504.

[0033] In some examples, the antibiotic resistance prediction system 500 includes a curated training dataset comprising paired bacterial genomes and antibiotic resistance phenotypes, which can be used to train the PARP model 504. The PARP model 504 can also include an independent test dataset used to evaluate the model's performance. For example, the data source can include a publicly available training data source that can include 3,393 different isolates belonging to nine bacterial species with 29 different antibiotics. The data source can also include an external test data source that includes 1,970 different isolates belonging to four bacterial species with 10 antibiotics. The PARP model 504 can also include data preparation procedures for converting the training data and / or test data into usable training and / or test datasets. The data preparation procedures can include one-hot encoding of antibiotic signature information, combinations for sequencing gene variants, and / or validation splits for nested cross-validation. Thus, the antibiotic resistance prediction system 500 can generate a training dataset having a first isolate-variant matrix, a first antibiotic indicator matrix, and / or a first isolate resistance signature vector. Additionally, the antibiotic resistance prediction system 500 can generate a test dataset including a second isolate-variant matrix, a second antibiotic indicator matrix, and / or a second isolate resistance signature vector. The PARP model 504 can undergo a data training procedure, in which hyperparameters are tuned by nested cross-validation and model fit assessment is performed on the training dataset. Additionally, the PARP model 504 can undergo a data validation procedure, in which predictions generated by the PARP model 504 are evaluated using the test dataset. The data training procedure can also include generating weights for bacteria and antibiotics and / or visualizing the weights. Figure 5B shows a heatmap 506 including rows representing samples, columns representing antibiotic resistance genes (ARGs), and shading representing the presence of ARGs.Figure 5C shows a graph 508 depicting the number of variants for different isolates sorted in descending order, with darker shaded bars representing the first number of variants shared with other isolates and lighter shaded bars representing the second number of variants unique to that isolate.

[0034] 6A and 6B show an example of an antibiotic resistance prediction system 600 including a prediction accuracy assessment 602 according to certain embodiments of the disclosed technology. For example, the first prediction accuracy assessment 604 shown in FIG. 6A can be based on a first test dataset (e.g., a National Center for Biotechnology Information (NCBI) dataset), and the second prediction accuracy assessment 606 shown in FIG. 6B can be based on a second test dataset (e.g., an MD Anderson Cancer Center dataset). According to the prediction accuracy assessment 602, the PARP model 504 can have improved prediction accuracy over other prediction methods, such as a support vector machine (SVM) model, a logistic regression model with L2 regularization, and / or a random forest (RF) model.

[0035] 7A-7D show an example of an antibiotic resistance prediction system 700 according to the disclosed technology, including results output by the PARP model 504. For example, FIG. 7A shows a classification graph 702 illustrating that antibiotics of the same class can be clustered based on antibiotics embedded in the first and second principal components. This output can use weights from four dense layers from the FiLM generator component of the PARP model 504. For example, carbapenems can form the first cluster and / or quinolone antibiotics / fluoroquinolone antibiotics can form the second cluster. FIG. 7B shows an agglomerative clustering dendrogram 704 of antibiotics corresponding to the classification graph 702 of FIG. 7A. This dendrogram can use a Euclidean distance measure and Ward linkage criterion, with a cluster threshold of 70% of the maximum linkage value. Furthermore, FIG. 7C shows a second classification graph 706 for classifying bacterial isolates. For example, the second classification graph 706 may show a third cluster of Acinetobacter baumannii, a fourth cluster of Streptococcus pneumoniae, a fifth cluster of Pseudomonas aeruginosa, a sixth cluster of Klebsiella pneumoniae, a seventh cluster of Escherichia coli, and / or an eighth cluster of Salmonella enterica. Figure 7D shows an agglomerative clustering dendrogram 708 of bacterial isolates using a Euclidean distance measure and Ward linkage criterion, corresponding to the classification graph 706 of Figure 7. The cluster threshold for the agglomerative clustering dendrogram 708 may be 70% of maximum linkage.

[0036] FIG. 8A illustrates an example of an antibiotic resistance prediction system 800 according to the disclosed technology, including a cloud-based deployment 802. The cloud-based deployment 802 can include a web service portal, such as a content delivery network (CDN) accelerated website (e.g., Cloudfront), which users can access to upload data for the PARP model 504. The data can be uploaded to a storage service (e.g., S3), which can trigger an event-driven platform, such as a serverless platform like Lambda. Triggering the event-driven platform can initiate a computation process in a container orchestration service (e.g., Elastic Container Service), which can execute the machine learning model 120 (e.g., the PARP model 504) disclosed herein. Prediction results from the PARP model 504 can be sent back from the container orchestration service to the storage service for viewing, display, download, etc.

[0037] FIG. 8B illustrates an example of an antibiotic resistance prediction system 800 according to the disclosed technology, showing predicted output results 804 of the PARP model 504 (e.g., via cloud-based deployment 802). These predicted results can correspond to multiple different uploaded genomes of different pathogens represented by the x-axis. Lower y-values ​​can correspond to antibiotic susceptibility, and higher y-values ​​can correspond to antibiotic resistance. For example, the predicted output results 804 of the PARP model 504 can be based on multiple different antibiotics (e.g., 10 to 30 antibiotics, or more than 30 antibiotics). The cloud-based deployment 802 can form a distributed diagnostic test, where any remote device can upload any genome sequence via the cloud-based deployment 802 infrastructure, and predicted results can be generated and provided to the remote device. In other words, the cloud-based deployment 802 of the PARP model 504 can provide a scalable antibiotic resistance prediction platform over a wide area network (WAN), such as the Internet.

[0038] 8B or throughout this disclosure may be presented on one or more graphical user interfaces (GUIs) of one or more user devices. For example, a computing device associated with a clinic, hospital, laboratory, etc. may receive and / or present the output results in its GUI. In some scenarios, the GUI presenting the predicted output results 804 may be the same GUI that provided the upload of the genome sequence for analysis by cloud-based deployment 802, or the device presenting the output results may be a different device than the one that provided the genome sequence.

[0039] 9A and 9B illustrate an example antibiotic resistance prediction system 900 according to the disclosed technology, including one or more bar graphs 902 representing unique and / or shared mutation data. For example, a first bar graph 904 shown in FIG. 9A represents the number of antibiotic resistance genes shared between species as determined by the PARP model 504. A second bar graph 906 shown in FIG. 9B represents the shared and unique variants of isolates from different pathogen species as determined in the training dataset. Unique variants represented by lighter shaded bars are shared only by that particular species represented on the x-axis. Darker shaded bars represent shared variants shared by at least two species.

[0040] 10A-10D illustrate an antibiotic resistance prediction system 1000 according to the disclosed technology, including one or more bar graphs 1002 representing unique and / or shared mutation data that may be determined by the PARP model 504. The one or more bar graphs 1002 in FIGS. 10A-10D may represent shared and / or unique variants for a particular pathogen species. The x-axis may represent isolates, with lighter shading y-values ​​representing the number of unique variants for that isolate, and darker shading y-values ​​representing the number of shared variants for that isolate. For example, FIG. 10A illustrates a first bar graph 1004 representing shared and unique variants per isolate for Enterobacter cloacae. Figure 10B shows a second bar graph 1006 representing shared and unique variants per isolate for Acinetobacter baumannii, a third bar graph 1008 representing shared and unique variants per isolate for Klebsiella aerogenes, and a fourth bar graph 1010 representing shared and unique variants per isolate for Salmonella enterica. Additionally, Figure 10C shows a fifth bar graph 1012 representing shared and unique variants per isolate for Enterobacter cloacae, a sixth bar graph 1014 representing shared and unique variants per isolate for Klebsiella pneumoniae, and a seventh bar graph 1016 representing shared and unique variants per isolate for Staphylococcus aureus. Additionally, Figure 10D shows an eighth bar graph 1018 representing shared and unique variants per isolate for E. coli, a ninth bar graph 1020 representing shared and unique variants per isolate for P. aeruginosa, and a tenth bar graph 1022 representing shared and unique variants per isolate for Streptococcus pneumoniae.

[0041] 11 illustrates an antibiotic resistance prediction system 1100 according to the disclosed technology, including the output of an antibiotic evaluation engine 1102, which can form part of a PARP model 504. In some scenarios, the PARP model 504 can include a two-stage approach, including a first stage in which pathogen genomes are analyzed to determine similarities and differences that may affect their antibiotic resistance. A second stage can include analysis of the antibiotics themselves using the antibiotic evaluation engine 1102 to determine phenotypic similarities and differences between antibiotics that may affect whether a pathogen is susceptible or resistant to an antibiotic.

[0042] For example, the antibiotic evaluation engine 1102 can generate antibiotic susceptibility data represented by a paired antibiotic susceptibility heat map 1104. For each pair of antibiotics in the paired antibiotic susceptibility heat map 1104, a square represents the percentage of isolates with identical phenotypes tested on both the antibiotic on the x-axis and the antibiotic on the y-axis. A blank square indicates no isolates were tested on that particular antibiotic combination. The paired antibiotic susceptibility heat map 1104 can show how different classes of antibiotics target different pathways. The PARP model 504 can integrate the results of the paired antibiotic susceptibility heat map 1104 into a determination of antibiotic resistance for different mutants through extrapolation of identical phenotypes. In this way, the PARP model 504 can make predictions for a particular antibiotic by recognizing similarities with other antibiotics that have been tested, even if that antibiotic has not been specifically tested.

[0043] 12A and 12B show an example of an antibiotic resistance prediction system 1200 according to the disclosed technology, including a nested cross-validation procedure 1202. For example, FIG. 12A shows an outer loop 1204 of the nested cross-validation procedure 1202, in which a training dataset can be divided into three outer folders, with one-third of the data treated as test data and two-thirds of the data treated as training data. The training dataset can be shuffled before this division to ensure that the selected dataset is representative of the entire data. The first outer folder 1206 can use the first one-third of the dataset as test data and the last two-thirds of the dataset as training data. The second outer folder 1208 can use the first and last third of the dataset as training data and the middle one-third as test data. The third outer folder 1210 can use the first two-thirds of the dataset as training data and the last one-third of the dataset as test data. In some cases, the results of the first outer folder 1206, the second outer folder 1208, and / or the third outer folder 1210 can be combined. FIG. 12B illustrates an inner loop 1212 of the nested cross-validation procedure 1202. While the inner loop 1212 illustrated in FIG. 12B corresponds to the first outer folder 1206, a similar or identical inner loop 1212 can be used for the second outer folder 1208 and / or the third outer folder 1210. The inner loop 1212 can include dividing the outer folder into three inner folds, with 20% of the dataset treated as a validation dataset. These validation datasets can be used to select hyperparameters, and predictive metrics on the test subsets can be used as training metrics. The three inner folds can report predictive accuracy on the validation set, and the nested cross-validation procedure 1202 can select the dense size with the greatest average accuracy among the three inner validation sets.The model can then be retrained on the training set of the outer fold using the selected density size. Finally, the inner loop 1212 can report the accuracy on the test set from the outer fold. Various operations of the nested cross-validation procedure 1202 disclosed herein can, in some cases, reduce bias in the PARP model 504.

[0044] FIG. 13 illustrates an example of an antibiotic resistance prediction system 1300 according to the disclosed technology, including a prediction output comparison 1302. For example, the prediction output comparison 1302 can include a box plot of the prediction accuracy in a training set using the PARP model 504, which can be compared to other prediction models. For example, the prediction output comparison 1302 can generate a prediction comparison between the PARP model 504 and a logistic regression model with L2 regularization, an RF model, and / or an SVM model. In some examples, pairwise p-values ​​between the PARP model 504 and the other models can be obtained from a Wilcoxon signed-rank test and / or a Kruskal-Wallis test, which can be used to compare results between the four models. In some scenarios, the prediction output comparison 1302 can indicate higher accuracy (e.g., via tighter box clusters) for the PARP model 504 compared to the other models.

[0045] In some examples, the overall accuracy of the PARP model 504 on the training data can be 94.8%, while 64 of the 93 bacteria-antibiotic pairs can have an accuracy greater than 90%, and 83 pairs have an accuracy greater than 80%. Compared to existing machine methods, the PARP model 504 can have the best performance (e.g., PARP 94.8%, Elastic Net L2 93.2%, RF 91.1%, and SVM 93.2%).

[0046] FIG. 14 illustrates an example of an antibiotic resistance prediction system 1400 according to the disclosed technology, including a prediction accuracy assessment 1402 for unknown bacteria and antibiotic combinations. The prediction accuracy assessment 1402 can predict the accuracy of various bacteria and antibiotic pairs based on a PARP model 504 trained on a subset of the dataset, as described above. The prediction accuracy assessment 1402 can include bar graphs 1404 corresponding to specific bacterial groups, such as the Enterobacter cloacae group. Crosses in different bar graphs represent the predicted training accuracy of this pair of isolates in the original PARP model 504. The threshold can be set to 0.5. While FIG. 14 shows an assessment of the Enterobacter cloacae group, it should be understood that the multiple predictive accuracy assessments 1402 generating the multiple bar graphs 1404 can be used for multiple different bacterial groups, such as Acinetobacter baumannii, Escherichia coli, Klebsiella aerogenes, Klebsiella pneumoniae, Pseudomonas aeruginosa, Salmonella enterica, Staphylococcus aureus, and / or Streptococcus pneumoniae.

[0047] FIG. 15 illustrates an example of an antibiotic resistance prediction system 1500 according to the disclosed technology, including a predicted area under the receiver operating characteristic curve (AUROC) 1502 for an unknown bacteria-antibiotic combination. The predicted AUROC 1502 can predict the accuracy of various bacteria-antibiotic pairs based on a PARP model 504 trained on a subset of the dataset that excludes isolates from the pair. The predicted AUROC 1502 can include a bar graph 1504 corresponding to a specific bacterial group, such as Enterobacter cloacae. The crosses in the various bar graphs represent the predicted training AUROC for isolates from this pair under the original PARP model 504. While Figure 15 shows an assessment of the Enterobacter cloacae group, it is understood that multiple AUROCs 1502 generating multiple bar graphs 1504 can be used for multiple different bacterial groups, such as Acinetobacter baumannii, Escherichia coli, Klebsiella aerogenes, Klebsiella pneumoniae, Pseudomonas aeruginosa, Salmonella enterica, Staphylococcus aureus, and / or Streptococcus pneumoniae.

[0048] 16 illustrates an example of an antibiotic resistance prediction system 1600 according to the disclosed technology, including a feature-based linear modulation (FiLM) generation block 1602. Block 1602 can include the FiLM machine learning model described above with respect to FIG. 3. In some scenarios, the antibiotic resistance prediction system 1600 includes a conditional affine transformation structure. After the transformation from the FiLM generation block 1602, information from the antibiotic one-hot matrix can be integrated into the deep learning model using two interactions, such as multiplicative and additive interactions.

[0049] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 shown in Figures 1-16 can address the emerging public health threat of antibiotic resistance (AR). Additional details of the antibiotic resistance prediction systems 200 and 500-1600 are provided below.

[0050] The antibiotic resistance prediction systems 200 and 500-1600 can leverage advances in whole-genome bacterial sequencing technology and machine learning by providing in silico antibiotic resistance prediction results in a timely and accurate format. The bioinformatics techniques disclosed herein can profile bacterial sequences for machine learning features. These predictive models generated by the antibiotic resistance prediction systems 200 and 500-1600 can use both genomic features and antimicrobial susceptibility testing (AST) data to facilitate AR prediction. The antibiotic resistance prediction systems 200 and 500-1600 disclosed herein can address issues related to the limited availability of paired bacterial genome and their AST phenotype data, which can make building accurate predictive models difficult.

[0051] For example, antibiotic resistance prediction systems 200 and 500-1600 can include a deep learning model, such as machine learning model 120, to uncover relationships between antibiotic resistance genes (ARGs) and a wide range of antibiotics. Antibiotic resistance prediction systems 200 and 500-1600 can use orthologous gene variants as input machine learning features to identify their association with antibiotic resistance. The approach disclosed herein can offer at least two advantages. Ortholog-based features can have identifiable and / or explainable relationships between variants and antibiotics, which may not be the case with other approaches that simply check for the presence of ARGs or the number of short DNA fragments (e.g., k-mers). Furthermore, some other models rely on the availability of paired bacterial genomic features and their phenotypes (e.g., resistance or susceptibility phenotypes to each antibiotic). This can make the prediction task difficult when applied to less representative bacteria, which are constrained by a lack of data availability.

[0052] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 can include a machine learning framework for investigating protein variants across at least 9 bacterial species (e.g., or more or fewer bacterial species) and / or at least 29 antibiotics (e.g., or more or fewer antibiotics). The antibiotic resistance prediction systems 200 and 500-1600 can identify similar protein variants with similar antibiotic function across different bacterial species or antibiotic classes. In this way, the developed deep learning prediction models of the antibiotic resistance prediction systems 200 and 500-1600 may be suitable for predicting antibiotic resistance across a wide range of bacterial species. The antibiotic resistance prediction systems 200 and 500-1600 can predict resistance even for bacterial species with a limited number of available genomes.

[0053] In some examples, the antibiotic resistance prediction system 200, 500, 1600 workflow, such as the workflow 502 described above with reference to Figures 5A-5C, can begin with the curation of paired bacterial genomes and their AR phenotypes. After quality control procedures, a large dataset of 3,393 isolates with paired AST results can be curated for the antibiotic resistance prediction system 200, 500, 1600. Next, the antibiotic resistance prediction system 200, 500, 1600 can determine the sharing of genetic features associated with antibiotic resistance between species, followed by the fusion of shared genetic features with antibiotic features. The PARP model 504 is optimized through the nested cross-validation procedure 1202 described above with reference to Figures 12A and 12B, and its performance is evaluated unbiasedly. Compared to other prediction models, the PARP model 504 can have high accuracy (e.g., 94.8%). By examining the model parameters, the antibiotic resistance prediction systems 200 and 500-1600 can determine that the PARP model 504 is interpretable as exhibiting a unique relationship between AR and bacterial classification. Additionally, the antibiotic resistance prediction systems 200 and 500-1600 can validate the performance of the PARP model 504 using an independent dataset (e.g., including 197 isolates of four bacterial species).

[0054] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 can determine antibiotic resistance characteristics shared between species.

[0055] For example, by accessing the NCBI Resistance Record Database and / or the NCBI Short Read Archive, a large collection of paired bacterial genomes and AST data can be curated for the antibiotic resistance prediction systems 200 and 500-1600. As a quality control step, the antibiotic resistance prediction systems 200 and 500-1600 can filter out bacterial genomes with unclear species identity and / or limited sequencing coverage and classify minimum inhibitory concentration (MIC) test results using Clinical and Laboratory Standards Institute (CLSI) breakpoints. Furthermore, the final cohort can include nine bacterial species, 3,393 isolates, and a total of 29,187 binary AR test results.

[0056] In some scenarios, the next step may involve using the antibiotic resistance prediction systems 200 and 500-1600 to derive genetic features associated with ARGs from some or all species. Information between ARGs and antibiotics can be learned across multiple species and antibiotics. This contrasts with other machine learning models, in which each model is suitable for one bacterial species-antibiotic combination and may be limited by the sample size of available paired bacterial genomic features and resistance phenotypes. To develop an integrated model capable of studying all bacterial genomic features and their pan-antibiotic resistance profiles, the antibiotic resistance prediction systems 200 and 500-1600 can determine amino acid changes occurring in homologous genes. For different bacterial genomes, the antibiotic resistance prediction systems 200 and 500-1600 can characterize these amino acid changes using specific bioinformatics techniques, as discussed herein. The antibiotic resistance prediction systems 200 and 500-1600 can determine that 406 variants are shared across multiple species and / or that 174.7 (±42.6) variants are shared across bacterial isolates. The antibiotic resistance prediction systems 200 and 500-1600 can also determine that of the 402 AR genes, 250 genes are shared with three or more species. Furthermore, the antibiotic resistance prediction systems 200 and 500-1600 can determine that, on average, any given bacterial species is likely to possess 82.3% to 100% of the genes observed in other species. Furthermore, the antibiotic resistance prediction systems 200 and 500-1600 can also determine that bacterial isolates are likely to frequently exhibit resistance between antibiotics in similar classes. For example, the antibiotic resistance prediction systems 200 and 500-1600 can determine that AST results for doripenem and imipenem are similar. The antibiotic resistance prediction systems 200 and 500-1600 can utilize the extensive sharing and concordance of genetic features between antibiotic phenotypes to provide an integrated predictive machine learning framework for broad-based prediction, forming the PARP model 504.

[0057] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 can include a PARP model 504 with a deep learning model to predict antibiotic resistance across multiple combinations of bacterial species and antibiotics.

[0058] For example, the PARP model 504 can use paired genetic features and antibiotics as inputs and AST results as outputs. For example, the inputs for the PARP model 504 can include at least one of protein-level orthologous gene features (e.g., those listed in the Kyoto Encyclopedia of Genes and Genomes (KEGG)), orthologous gene names, orthologous gene variants, and / or AST data. One-hot encoding can be used for the antibiotic set (e.g., 29 antibiotics) included in the dataset. The model architecture design can be optimized by the antibiotic resistance prediction systems 200 and 500-1600, embedding and / or blending information from both genetic features and antibiotic features through nested cross-validation (e.g., nested cross-validation step 1202 in Figure 12). The optimal model for the PARP model 504 can be determined with one dense block, two FiLM generators, and / or a dense size that can be 1024.

[0059] In some examples, the PARP model 504 can predict resistance to both pathogens in the training set and pathogens not included in the training set by determining frequently shared genetic characteristics possessed by the pathogens. To evaluate the performance of the PARP model 504 for unknown bacteria-antibiotic combinations, the antibiotic resistance prediction systems 200 and 500-1600 can perform a leave-one-combination-out (LOCO) procedure, in which, for a given bacteria-antibiotic combination, the PARP model 504 is trained on samples not included in this combination and can predict resistance to samples included in this combination. This LOCO procedure can be repeated for multiple bacteria-antibiotic combinations (e.g., all 93 combinations).

[0060] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 can be externally validated using an independent dataset.

[0061] For example, the performance of PARP model 504 can be evaluated using an independent test dataset collected independently at MD Anderson (e.g., as described above with respect to Figure 6B). This test can use 197 unique isolates from four different bacterial species: Enterobacter cloacae (N = 13), Escherichia coli (N = 31), Klebsiella pneumoniae (N = 24), and Pseudomonas aeruginosa (N = 129). These isolates can be tested against 10 different antibiotics, resulting in 1,203 samples from 21 pathogen-drug pairs included in the test dataset. PARP model 504 can be evaluated using all 3,393 bacterial isolates included in the NCBI Resistance Record dataset. PARP model 504 can have the highest accuracy (accuracy = 93.55%) for the combination of E. coli and meropenem. The overall accuracy of the PARP model 504 is 76.56%, which may be better than other methods (e.g., SVM = 53.11%, Elastic Net L2 = 54.42%, RF = 56.28%).

[0062] The PARP model 504 can then be evaluated to determine whether it can predict novel bacteria-antibiotic combinations not seen in the training dataset. For example, the combination of E. coli and meropenem can be omitted from the NCBI Resistance Record dataset. The PARP model 504 can be retrained using the reduced dataset and can predict resistance in the MD Anderson dataset with 93.55% prediction accuracy.

[0063] In some scenarios, the antibiotic resistance prediction systems 200 and 500-1600 can determine, via embedding, explainable genetic features.

[0064] For example, after the final predictive model of the PARP model 504 is developed, the antibiotic resistance prediction systems 200 and 500-1600 can explore whether the network parameters from the trained model reflect hidden relationships between genetic features and antibiotics. For unsupervised classification of bacteria and antibiotics, principal component analysis (PCA) and / or hierarchical clustering analysis (HCA) can be performed based on the feature maps in the dense layer. The PARP model 504 can represent relationships between different antibiotics or bacterial species without prior information about these relationships. Figure 7A shows antibiotics embedded in the first two principal components. Although the first two dimensions can explain only 31.01% of the total variance in some scenarios, the antibiotic resistance prediction systems 200 and 500-1600 can determine that the hidden representations of antibiotic features tend to cluster by their classes, such as carbapenems and quinolones, as highlighted in the dashed box in Figure 7A. This indicates that the PARP model 504 automatically captures shared features within the same antibiotic class. Figure 7B further shows antibiotics clustered by class. Figure 7C further includes a visualization of the nine bacterial species included in the training dataset with respect to the first two principal components. Samples can be clustered with other samples belonging to the same bacterium. This allows the PARP model 504 to learn information to distinguish between different bacteria.

[0065] Furthermore, techniques for integrating bacterial genetic features and AST data can be explored by the antibiotic resistance prediction systems 200 and 500-1600. Model-independent methods can be used by artificially inducing different genetic features and quantifying the predicted change in resistance. Higher values ​​can reflect that the variant is more important for inducing resistance. For example, the antibiotic resistance prediction systems 200 and 500-1600 can determine that the K18768 variant of blaKPC (beta-lactamase class A KPC) contributes to the greatest resistance to doripenem, imipenem, and meropenem. Similarly, the K19096 variant of aph3-I (aminoglycoside 3'-phosphotransferase I), a gene that modifies aminoglycoside antibiotics (such as amikacin), can contribute to the greatest resistance.

[0066] In some scenarios, the antibiotic resistance prediction systems 200 and 500-1600 can facilitate the use of the PARP model 504 by the broader research community by providing a website for the developed PARP model 504 (e.g., as shown above with respect to Figure 8A). This website can provide a portal for users to upload genome sequences and / or obtain in silico predicted resistance profiles within minutes. Users can receive reports containing resistance probabilities for multiple antibiotics, such as 35 antibiotics, as shown in Figure 8B. Furthermore, this cloud-based deployment 802 can improve the scalability, affordability, and availability of the PARP model 504 in at least three ways. First, by packaging the software pipeline to extract orthologous gene features and compute resistance predictions within a software container, the cloud-based deployment 802 can be scaled up and online analyses can be performed simultaneously and reproducibly. Second, the website can use a serverless architecture, so minimal costs are incurred only when computations are performed. Additionally, the cloud-based deployment 802 includes built-in backup and replication mechanisms, allowing the server to provide uninterrupted service to researchers around the world.

[0067] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 can use a variety of data acquisition techniques.

[0068] For example, the training dataset can be processed based on the NCBI BioSample Antibiograms database. Resistance record tabular data can be used to validate bacterial isolates and can hold reported minimum inhibitory concentration (MIC) values. Corresponding sequence data can be from the NCBI Sequence Read Archive (SRA). In some scenarios, the PARP model 504 can hold bacterial isolates with matching species from resistance records and sequence data analysis, and resistance and susceptibility phenotypes can be based on CLSI standards.

[0069] The training dataset can include 3,393 unique bacterial isolates representing nine species: Acinetobacter baumannii (772), Enterobacter cloacae (79), Escherichia coli (350), Klebsiella aerogenes (68), Klebsiella pneumoniae (344), Pseudomonas aeruginosa (83), Salmonella enterica (1349), Staphylococcus aureus (31), and Streptococcus pneumoniae (317). Additionally, resistance phenotypes for 29 different antibiotics can be curated. Thus, the antibiotic resistance prediction systems 200 and 500–1600 can obtain 29,187 paired pathogen–antibiotic samples covering 93 different species and antibiotic combinations.

[0070] An external dataset from MD Anderson, the antibiotic resistance prediction systems 200 and 500–1600, can sequence 197 unique isolates. This dataset can serve as a validation cohort for the developed models and includes four bacteria: Enterobacter cloacae (13 isolates), Escherichia coli (31 isolates), Klebsiella pneumoniae (24 isolates), and Pseudomonas aeruginosa (129 isolates). These isolates can have paired antibiotic test phenotypes, resulting in a total of 1,203 paired pathogen-antibiotic combinations covering 21 species and antibiotic combinations.

[0071] In some embodiments, the training dataset can include multiple bacterial species (e.g., two or more), where the multiple bacteria can be Yersinia, Vibrio, Treponema, Streptococcus, Staphylococcus, Shigella, Salmonella, Rickettsia, Orientia, Pseudomonas, Neisseria, Mycoplasma, Mycobacterium, Listeria, Leptospira, Legionella, Klebsiella, Helicobacter, Haemophilus, Francisella, Escherichia, or Escherichia. The bacterial strains are from one or more genera including, but not limited to, Jerrichia, Ehrlichia, Enterococcus, Coxiella, Corynebacterium, Clostridium, Chlamydia, Chlamydophila, Campylobacter, Burkholderia, Brucella, Borrelia, Bordetella, Bifidobacterium, Bacillus, Proteus, Morganella, Sphingobium, Sphingomonas, Zymomonas, Cupriavidus, or any combination thereof.

[0072] In some embodiments, the plurality of bacterial species are selected from the group consisting of Achromobacter spp., Acidaminococcus fermentans, Acinetobacter calcoaceticus, Actinomyces spp., Actinomyces viscosus, Actinomyces naeslundii, Aeromonas spp., Aggregatibacter actinomycetemcomitans, Anaerobiospirillum spp., Alcaligenes faecalis, Arachnia propionica, Bacillus spp., Bacteroides spp., Bacteroides gingivalis, Bacteroides fragilis, Bacteroides intermedius, Bacteroides spp. Ides melaninogenicus, Bacteroides pneumocyntes, Bacterionema matrcotti, Bifidobacterium spp., Buchnera aphidicola, Butyriviverio fibrosolvens, Bordetella pertussis, Campylobacter spp., Campylobacter coli, Campylobacter sputum, Campylobacter upsaliensis, Capnocytophaga spp., Chlamydophila pneumoniae, Clostridium spp., Citrobacter freundii, Clostridium difficile, Clostridium sordellii, Corynebacterium spp., Eikenella Genus Corodens, Enterobacter cloacae, Enterococcus spp., Enterococcus faecalis, Enterococcus faecium, Escherichia coli, Eubacterium spp., Flavobacterium spp., Fusobacterium spp., Fusobacterium nucleatum, Gordonia spp., Haemophilus parainfluenzae, Helicobacter pylori, Haemophilus paraphlophilus, Klebsiella, Lactobacillus spp., Listeria monocytogenes, Leptotrichia buccalis, Methanobrevibacter smithii, Micrococcus fiavus, Moraxella spp. catarrhalis, Mycobacterium tuberculosis, Mycobacterium paratuberculosis, Mycoplasma pneumoniae, Morganella morganii, Mycobacteria spp., Mycoplasma spp., Micrococcus spp., Mycobacterium chelonae, Neisseria spp., Neisseria sicca, Pasteurella multocida, Peptococcus spp., Peptostreptococcus spp., Plesiomonas shigelloides, Porphyromonas gingivalis, Proteus spp., Proteus mirabilis, Proteus vulgaris, Propionibacterium spp., Propionibacterium acnes, Providencia spp., Pseudomonas aeruginosa, Orientia,The bacterial strain is selected from Ruminococcus bromii, Rothia dentocariosa, Ruminococcus spp., Sarcinaltea spp., Serratia marcescens, Shigella boydii, Shigella fiexneri, Shigella sonnei, Sarcina spp., Staphylococcus aureus, Staphylococcus epidermidis, Streptococcus anginosus, Streptococcus faecalis, Streptococcus mutans, Streptococcus oxalis, Streptococcus pneumoniae, Streptococcus sobrinus, Streptococcus viridans, Streptococcus pyogenes, Salmonella enterica, Salmonella typhi, Salmonella paratyphi, Rickettsiae, Torulopsis glabrata, Treponema denticola, Treponema refringens, Veillonella spp., Vibrio spp., Vibrio sputum, Wolinella succinogenes, Yersinia enterocolitica, or any combination thereof.

[0073] In some embodiments, one or more of the plurality of bacterial species is a pathogenic bacteria. In some embodiments, the pathogenic bacteria is selected from the group consisting of Clostridium difficile, Salmonella, enteropathogenic Escherichia coli, multidrug-resistant bacteria such as Klebsiella and Escherichia coli, carbapenem-resistant Enterobacteriaceae (CRE), extended-spectrum beta-lactam-resistant Enterococci (ESBL), fluoroquinolone-resistant Enterobacteriaceae, and vancomycin-resistant Enterococci (VRE), multidrug-resistant bacteria, extended-spectrum beta-lactam-resistant Enterococci (ESBL), carbapenem-resistant Enterobacteriaceae (CRE), fluoroquinolone-resistant Enterobacteriaceae, and vancomycin-resistant Enterococci (VRE), Aeromonas hydrophila, Campylobacter fetus, Plesiomonas shigelloides, Bacillus cereus, Campylobacter jejuni, Clostridium botulinum, Clostridium difficile, Clostridium perfringens, and the like. The bacterial strain can be selected from the group consisting of: Enterococcus mutans, Enteroaggregative E. coli, Enterohemorrhagic E. coli, Enteroinvasive E. coli, Enterotoxigenic E. coli (such as, but not limited to, LT and / or ST), E. coli O157:H7, Helicobacter pylori, Klebsiella pneumoniae, Listeria monocytogenes, Plesiomonas shigelloides, Salmonella spp., Salmonella typhi, Salmonella paratyphi, Shigella dysenteriae, Staphylococcus aureus, Staphylococcus aureus, Vancomycin-resistant Enterococci, Vibrio spp., Vibrio cholerae, Vibrio parahaemolyticus, Vibrio vulnificus, and Yersinia enterocolitica, antibiotic-resistant Proteobacteria, Vancomycin-resistant Enterococci (VRE), Carbapenem-resistant Enterobacteriaceae (CRE), Fluoroquinolone-resistant Enterobacteriaceae, Extended-spectrum beta-lactamase-producing Enterobacteriaceae (ESBL-E), or any combination thereof.

[0074] In some embodiments, one or more of the plurality of bacterial species can be antibiotic-resistant bacteria, such as Acinetobacter baumannii, Enterobacter cloacae, Escherichia coli, Klebsiella aerogenes, Klebsiella pneumoniae, Pseudomonas aeruginosa, Salmonella enterica, Staphylococcus aureus, Streptococcus pneumoniae, Klebsiella oxytoca, Serratia marcescens, Enterobacter aerogenes, Proteus mirabilis, Acinetobacter baumannii, Stenotrophomonas maltophilia, Staphylococcus epidermidis, Staphylococcus haemolyticus, Staphylococcus saprophyticus, Streptococcus pyogenes, Streptococcus agalactiae, Streptococcus mitis, Enterococcus faecium, Enterococcus faecalis, Candida albicans ... The bacterial infection may include, but is not limited to, Candida tropicalis, Candida parapsilosis, Candida krusei, Candida glabrata, Mycobacterium tuberculosis, Neisseria meningitidis, Listeria monocytogenes, Citrobacter friendly, Salmonella enteritidis, Serratia marcescens, Proteus mirabilis, Hafnia alvei, Enterobacter spp., Serratia marcescens, Pseudomonas putida, Enterobacter cloacae, Proteus vulgaris, Providencia rettgeri, Shigella dysenteriae, Shewanella algae, Acinobacter junis, Ralstonia pickettii, Pandoraea pnomenaeus, Pasteurella multocida, Bordetella bronchiseptica, Listeria monocytogenes, Bacillus cereus, or any combination thereof.

[0075] In some embodiments, one or more of the plurality of bacterial species can be multidrug-resistant. The multidrug-resistant bacteria can be Acinetobacter baumannii, such as ATCC isolate #2894233-696-101-1, ATCC isolate #2894257-696-101-1, ATCC isolate #2894255-696-101-1, ATCC isolate #2894253-696-101-1, or ATCC #2894254-696-101-1; ATCC isolate #33128, ATCC isolate #2894218-696-101-1, ATCC isolate #2894219-696-101-1, ATCC Citrobacter freundii, such as isolate #2894224-696-101-1, ATCC isolate #2894218-632-101-1, or ATCC isolate #2894218-659-101-1; ATCC isolate #22894251-659-101-1, ATCC isolate #22894264-659-101-1, ATCC isolate #22894246-659-101-1, ATCC isolate #22894243-659-101-1, or ATCC isolate #22894245-659 Enterobacter cloacae, such as ATCC isolate #22894228-659-101-1, ATCC isolate #22894222-659-101-1, ATCC isolate #22894221-659-101-1, ATCC isolate #22894225-659-101-1, or ATCC isolate #22894245-659-101-1; Enterococcus faecalis, such as ATCC isolate #51858, ATCC isolate #35667, ATCC isolate #2954833_26 Enterococcus faecium, such as ATCC isolate #94008, ATCC isolate #2954833_2692765, or ATCC isolate #2954836_2694361; Escherichia coli, such as ATCC isolate #CGUC11332, CGUC11350, CGUC11371, CGUC11378, or CGUC11393; Kiemann's disease, such as ATTC isolate #27736, ATTC isolate #29011, ATTC isolate #20013, ATTC isolate #33495, or ATTC isolate #35657;Serratia marcescens, such as ATCC isolate #43862, ATCC isolate #2338870, ATCC isolate #2426026, ATCC isolate #SIID2895511, or ATCC isolate #SIID2895538; or Staphylococcus aureus, such as ATCC isolate #JHH02, ATCC isolate #JHH02, ATCC isolate #JHH03, ATCC isolate #JHH04, ATCC isolate #JHH05, or ATCC isolate #JHH06;

[0076] In some embodiments, the machine learning algorithm is trained using genetic information of bacteria associated with antibiotic-resistant bacteria. In some embodiments, the machine learning algorithm is trained using genetic information associated with Acinetobacter baumannii, Enterobacter cloacae, Escherichia coli, Klebsiella aerogenes, Klebsiella pneumoniae, Pseudomonas aeruginosa, Salmonella enterica, Staphylococcus aureus, Streptococcus pneumoniae, or any combination thereof. In some embodiments, the machine learning algorithm is trained using genetic information associated with Acinetobacter baumannii. In some embodiments, the machine learning algorithm is trained using genetic information associated with Enterobacter cloacae. In some embodiments, the machine learning algorithm is trained using genetic information associated with Escherichia coli. In some embodiments, the machine learning algorithm is trained using genetic information associated with Klebsiella aerogenes. In some embodiments, the machine learning algorithm is trained using genetic information associated with Klebsiella pneumoniae. In some embodiments, the machine learning algorithm is trained using genetic information associated with Pseudomonas aeruginosa. In some embodiments, the machine learning algorithm is trained using genetic information associated with Salmonella enterica. In some embodiments, the machine learning algorithm is trained using genetic information associated with Staphylococcus aureus. In some embodiments, the machine learning algorithm is trained using genetic information associated with Acinetobacter baumannii, Enterobacter cloacae, Escherichia coli, Klebsiella aerogenes, Klebsiella pneumoniae, Pseudomonas aeruginosa, Salmonella enterica, Staphylococcus aureus, and Streptococcus pneumoniae.

[0077] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 can perform various operations to derive genetic signatures.

[0078] For example, as described above, antibiotic resistance prediction systems 200 and 500-1600 can derive interpretable KEGG orthologous gene-based sequence variants. Sequence reads can be aggregated to obtain a consensus reference genome, gene sequences can be matched to detected and clustered UniRef20 protein sequences, and different reference gene clusters can be associated with KEGG orthologous (KO) genes. For example, variant K18768.0 can indicate blaKPC, a K. pneumoniae carbapenemase. Here, K18768 is the KEGG KO gene name, and 0 represents the UniRef cluster. Amino acid signatures can lead to good predictive performance for a single bacterial species-antibiotic combination.

[0079] In some examples, the PARP model 504 can have different model architectures and parameters.

[0080] For example, the PARP model 504 can be constructed from at least four types of blocks, as shown in Figure 3. These blocks include a variance block (Var block) for computing embeddings of bacterial variants; a dense block for representing variant-level features using a deep neural network; a FiLM block consisting of a FiLM generator that transforms antibiotic features and blends bacterial variant features via a conditional affine transformation, thereby effectively blending features from both domains; and / or a classifier block for calculating the probability of resistance or susceptibility. The model hyperparameters of these blocks can be tuned using grid search. Specifically, the geometric size of the dense layer can be adjusted (e.g., 64, 128, 256, 512, or 1024), the number of dense blocks can be adjusted (e.g., 1, 2, or 3), and / or the number of FiLM generators can be adjusted (e.g., 1, 2, 3, 4, 5, 6, or 7).

[0081] In some scenarios, the antibiotic resistance prediction systems 200 and 500-1600 can determine optimized parameter sets using the nested cross-validation procedure 1202 of Figures 12A and 12B to report unbiased prediction accuracy. For example, as described above with respect to Figures 12A and 12B, the entire dataset (e.g., N = 29,187) can first be divided into three outer folds, with two outer folds serving as training and the remainder as testing. The training set can then be divided into three inner folds, with two inner folds serving as sub-training and the remainder serving as validation. Each hyperparameter combination can be used to train a neural network for 10 epochs on the sub-training set. The hyperparameters can be retrained with the maximum validation accuracy and used to retrain the neural network on the training samples for each outer fold. The average accuracy across the three outer folds as the overall prediction accuracy can be reported and defined as follows:

number

[0082] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 can generate one or more model descriptions.

[0083] For example, to understand the inner structure of the PARP model 504, the antibiotic resistance prediction systems 200 and 500-1600 can analyze estimated parameters from "hidden" blocks (e.g., neural network layers) within the model. The antibiotic resistance prediction systems 200 and 500-1600 can employ at least two unsupervised learning algorithms, such as principal component analysis (PCA) and / or hierarchical cluster analysis (HCA). PCA can project the original data into a principal component space, which can be a low-dimensional feature set that best-efforts to preserve the variability of the original data. Similar observations can be clustered together in the low-dimensional space. Additionally or alternatively, HCA can seek homogeneous subgroups among the original observations by iteratively fusing two clusters that share the most similarity. To interpret the 29 antibiotics, the antibiotic resistance prediction systems 200 and 500-1600 can average the feature maps of the dense layers from all FiLM generators to generate a weight matrix.

number

number

[0084] Furthermore, through model-independent explanations, we can estimate the contribution of each genetic mutation to resistance, conditional on a wild-type baseline. We can manually create an indicator vector of length 14,615 as the bacterial genetic feature input, and use one-hot encoded antibiotics as the antibiotic feature input. The PARP model 504 then calculates the resistance probability relative to the baseline p 0 resistance Then, the antibiotic resistance prediction system 200 and 500-1600 can mutate the i-th element to 1 to mimic a bacterial isolate carrying the corresponding mutation. Using this mutated genetic feature vector, the PARP model 504 calculates a new resistance probability p i resistance The PARP model 504 can calculate the Effect i =p i resistance -p 0 resistance The difference between the two can be used to represent the effect of genetic variant i on a particular antibiotic.

[0085] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 can generate one or more predictions for unknown bacteria-antibiotic combinations.

[0086] For example, the PARP model 504 may be an integrated AR prediction model for multiple bacteria-antibiotic combinations, and thus may have the potential to predict resistance probabilities for novel bacteria-antibiotic combinations by leveraging information learned from existing combinations. To quantitatively evaluate its performance, the antibiotic resistance prediction systems 200, 500, and 1600 may conduct a leave-one-combination-out (LOCO) experiment by excluding isolates from one bacteria-antibiotic combination, reconstructing the PARP model 504 based on the remaining training data, and predicting AR for the excluded isolates. The prediction accuracy of each LOCO experiment may also be reported based on the architecture of the PARP model.

[0087] Next, in some scenarios, the LOCO model can be evaluated on an external dataset containing, for example, 1,075 samples from 18 different bacteria-antibiotic pairs. For each specific bacteria-antibiotic pair in the external dataset, the antibiotic resistance prediction systems 200 and 500-1600 can select the model trained in the LOCO experiment described above, excluding samples from this pair. The model's performance on samples belonging to this bacteria-antibiotic pair can then be evaluated. For amikacin, a high accuracy of 99.26% can be achieved for Pseudomonas aeruginosa, indicating that the PARP model 504 was able to predict some bacteria-antibiotic pairs not present in the training dataset.

[0088] In some examples, the antibiotic resistance prediction systems 200 and 500-1600 perform various external validation procedures.

[0089] For example, to validate the predictive performance of the PARP model 504, the antibiotic resistance prediction systems 200 and 500-1600 can use an external dataset collected at MD Anderson. As previously described, the PARP model 504 can be trained using a dataset (N=29,187) with an optimal hyperparameter set (e.g., one dense block, two FiLM generators, and / or a dense layer size of 1,024), using 30 epochs and a batch size of 32. Performance metrics can include overall prediction accuracy, receiver operating characteristic (ROC) curves, and / or AUROC values ​​for individual bacteria and antibiotic combinations. The overall prediction accuracy is calculated by multiplying the weighted average prediction accuracy of individual bacteria-antibiotic pairs by the ROC curve.

number

[0090] In some examples, the PARP model 504 can be trained end-to-end from scratch with a batch size of 32, an RMSprop optimizer with a learning rate of 0.001, ReLU activation, 10 epochs, and / or a dropout rate of 0.5. During the nested cross-validation procedure 1202, the proportion of test size in outer folds can be 33% and the proportion of validation size in inner folds can be 20%.

[0091] In some scenarios, the PARP model 504 for pan-antibiotic resistance prediction is based on deep learning modeling techniques, which are sometimes criticized for being black-box. To enhance interpretability, the PARP model 504 can be explicitly designed, with each network block having its own purpose. The FiLM structure allows efficient blending of bacterial genetic features and antibiotic features. The PARP model 504 can also be visualized by optimizing network parameters. Therefore, providing a model explanation allows users to better understand the output. Furthermore, the PARP model 504 can predict untrained bacteria-antibiotic combinations. The PARP model 504 can be used when underrepresentative combinations or sample size are a concern.

[0092] While proteins are in some cases the primary functional units contributing to common resistance mechanisms in prokaryotes, there may be other mechanisms associated with resistance (e.g., metabolic genes) that are not captured by genomics alone, and antibiotic resistance prediction systems 200 and 500-1600 can use these to expand the gene features based on the PARP model 504. Thus, the PARP model 504 can be a useful tool for predicting antibiotic resistance across the pathogen space using a variety of different mechanisms.

[0093] As described above, the present disclosure provides a method for predicting antibiotic resistance using multiple antibiotics. The multiple antibiotics (e.g., two or more) can include antibiotics known in the art. Similarly, the antibiotic resistance predicted by the disclosed method can be an antibiotic known in the art.

[0094] In some embodiments, the antibiotic may be a macrolide antibiotic, a sulfa antibiotic, a carbostyril antibiotic, a nitrofuran antibiotic, a cephalosporin analog, or any combination thereof.

[0095] In some embodiments, the antibiotic of the present disclosure can be from an antibiotic class, non-limiting examples of antibiotic classes include aminoglycosides, carbapenems and monobactams, cephalosporins, chloramphenicol, lincosamides, macrolides, pleuromutilins, glycopeptides, polypeptides, penicillins, polymyxins, quinolones, sulfonamides, and tetracyclines, among others. In some embodiments, the antibiotic can include a penicillin (e.g., ampicillin, piperacillin, benzylpenicillin, methicillin, and cloxacillin), a cephalosporin (e.g., cefotaxime and ceftazidime, cephaloridine), a carbapenem (e.g., iminipenem, meropenem, etrapenem, doripenem), a monobactam (e.g., aztreonam), or any combination thereof.

[0096] In some embodiments, the antibiotic is gentamicin, kanamycin, streptomycin, neomycin, tetracycline, terramycin, aureomycin, doxycycline, erythromycin, roxithromycin, sulfadiazine, sulfadimidine, sulfadimethoxine, sulfamethoxazole, sulfadoxine, norfloxacin, ciprofloxacin, ofloxacin, gatifloxacin, sparfloxacin, moxifloxacin, furazolidone, furaltadone, furantoin, nitrofuran, Zon, chloromycetin, thiamphenicol, clindamycin, lincomycin, ampicillin, gentamicin, kanamycin, streptomycin, erythromycin, clindamycin, tetracycline, chloramphenicol, balofloxacin, ceftiofur, cinoxacin, ciprofloxacin, clinafloxacin, enoxacin, fleroxacin, gemifloxacin, levofloxacin, lomefloxacin, nadifloxacin, nalidixic acid, oxolinic acid, pazufloxacin, pefloxacin, piperazine Mimidate, piromidate, prulifloxacin, losoxacin, rufloxacin, sitafloxacin, sparfloxacin, tosufloxacin, chlortetracycline, demeclocycline, doxycycline, lymecycline, meclocycline, methacycline, minocycline, omadacycline, oxytetracycline, rolitetracycline, sarecycline, amikacin, cefepime, imipenem, amoxicillin, amoxicillin / clavulanic acid, ampicillin / sulbactam, azithromycin, cephalosporin The active ingredient may be cephalosporin, cefazolin, cefepime, cefotaxime, cefoxitin, ceftriaxone, cefuroxime, daptomycin, ertapenem, fosfomycin, fusidic acid, linezolid, meropenem, methicillin, mupirocin, nitrofurantoin, oxacillin, penicillin, piperacillin / tazobactam, quinupristin / dalfopristin, rifampicin, teicoplanin, teigecycline, tobramycin, trimethoprim / sulfamethoxazole, vancomycin, or any combination thereof.

[0097] In some embodiments, the antibiotic comprises amoxicillin, meropenem, amoxicillin / clavulanic acid, cefoxitin, chloramphenicol, kanamycin, trimethoprim / sulfamethoxazole, ceftiofur, ciprofloxacin, ceftazidime, ampicillin, cefotaxime, ampicillin / sulbactam, aztreonam, ceftriaxone, tetracycline, ertapenem, erythromycin, tobramycin, amikacin, clindamycin, cefazolin, levofloxacin, doripenem, impipenem, gentamicin, cefepime, cefuroxime, piperacillin / tazobactam, or any combination thereof. In some embodiments, the antibiotic is amoxicillin, meropenem, amoxicillin / clavulanic acid, cefoxitin, chloramphenicol, kanamycin, trimethoprim / sulfamethoxazole, ceftiofur, ciprofloxacin, ceftazidime, ampicillin, cefotaxime, ampicillin / sulbactam, aztreonam, ceftriaxone, tetracycline, ertapenem, erythromycin, tobramycin, These include amikacin, clindamycin, cefazolin, levofloxacin, doripenem, impipenem, gentamicin, cefepime, cefuroxime, piperacillin / tazobactam, linezolid, tidezolid, ceftazidime-avibactam, ceftolozane-tazobactam, cefiderocol, imipenem-relebactam, durrobactam-sulbactam, fidaxomicin, eravacycline, dalbavancin, and ceftaroline.

[0098] In some embodiments, the antibiotic comprises amikacin, ampicillin, cefepime, linezolid, tidezolid, ceftazidime-avibactam, ceftolozane-tazobactam, cefiderocol, imipenem-relebactam, durrobactam-sulbactam, fidaxomicin, eravacycline, dalbavancin, ceftaroline, or any combination thereof. In some embodiments, the antibiotic is amikacin. In some embodiments, the antibiotic is ampicillin. In some embodiments, the antibiotic is cefepime. In some embodiments, the antibiotic is linezolid. In some embodiments, the antibiotic is tidezolid. In some embodiments, the antibiotic is ceftazidime-avibactam. In some embodiments, the antibiotic is ceftolozane-tazobactam. In some embodiments, the antibiotic is cefiderocol. In some embodiments, the antibiotic is imipenem-relebactam. In some embodiments, the antibiotic is durrobactam-sulbactam. In some embodiments, the antibiotic is fidaxomicin. In some embodiments, the antibiotic is eravacycline. In some embodiments, the antibiotic is dalbavancin. In some embodiments, the antibiotic is ceftaroline.

[0099] While specific embodiments are described, it should be understood that this is done for illustrative purposes only. Those skilled in the relevant art will recognize that other components and configurations may be used without departing from the spirit and scope of the technology of the present disclosure. Therefore, the following description and drawings are illustrative and should not be construed as limiting. Numerous specific details are set forth to provide a thorough understanding of the technology of the present disclosure. However, in some cases, well-known or conventional details are not set forth to avoid obscuring the description. A reference to one embodiment or an embodiment in the technology of the present disclosure may be a reference to the same embodiment or any embodiment, and such a reference means at least one of the embodiments.

[0100] Reference to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the disclosed technology. The appearances of the phrase "in one embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, or to separate or alternative embodiments that are mutually exclusive of other embodiments. Furthermore, various features are described that may be exhibited in some embodiments and not in other embodiments.

[0101] The terms used herein generally have their ordinary meanings in the art, in the context of the technology of this disclosure, and in the specific context in which each term is used. Alternative terms and synonyms may be used for any one or more of the terms discussed herein, and no special significance should be attached to whether a term is specifically described or discussed herein. In some cases, synonyms for particular terms are provided. The description of one or more synonyms does not exclude the use of other synonyms. The use of examples anywhere in this specification, including examples of terms discussed herein, is for illustrative purposes only and is not intended to further limit the scope and meaning of the technology of this disclosure or any exemplified term. Similarly, the technology of this disclosure is not limited to the various embodiments described herein.

[0102] Although not intended to limit the scope of the technology of the present disclosure, examples of devices, apparatuses, methods, and their related results according to embodiments of the technology of the present disclosure are provided below. Note that for the convenience of the reader, titles or subtitles may be used in the examples, but these in no way limit the scope of the technology of the present disclosure. Unless otherwise defined, technical and scientific terms used herein have the meanings commonly understood by those skilled in the art to which the technology of the present disclosure pertains. In the event of any conflict, the present specification, including definitions, shall prevail.

[0103] Additional features and advantages of the techniques of the present disclosure will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by practice of the principles disclosed herein. The features and advantages of the techniques of the present disclosure may be realized and obtained by means of the instruments and combinations particularly pointed out in the appended claims. These and other features of the techniques of the present disclosure will become more fully apparent from the following description and the appended claims, or may be learned by practice of the principles described herein.

Claims

1. 1. A method for antibiotic resistance prediction, comprising: receiving, in a machine learning system, genetic information associated with a bacterium, the species of the bacterium being one of a plurality of bacterial species; receiving, at the machine learning system, an indication of one antibiotic of a plurality of antibiotics, the machine learning system being trained for the plurality of antibiotics using genetic information associated with the plurality of bacterial species; outputting, at the machine learning system, an indication of antibiotic resistance associated with the bacteria and the antibiotic received at the machine learning system; A method comprising:

2. The method of claim 1 , wherein the machine learning system is trained using protein sequences associated with the plurality of bacterial species.

3. The method of claim 1 , wherein the machine learning system comprises a feature-based linear modulation (FiLM) machine learning system.

4. The method of claim 1 , wherein the machine learning system jointly models antibiotics and bacterial mutants.

5. The method of claim 1 , wherein the machine learning system includes a rectified linear activation function (ReLU) layer, a batch normalization layer, and a dropout layer.

6. 1. A method for antibiotic resistance prediction, comprising: Machine learning systems, genetic information associated with multiple bacterial species; and Antibiotic characteristics associated with multiple antibiotics training a pan-antibiotic resistance prediction (PARP) model by providing one or more training datasets comprising: receiving, in the machine learning system, a genome sequence associated with a particular bacterial isolate; outputting, in the machine learning system, a predictive indicator of antibiotic resistance associated with one or more antibiotics for the particular bacterial isolate; A method comprising:

7. 7. The method of claim 6, further comprising performing a data preparation procedure on the one or more training datasets by one-hot encoding the antibiotic feature information.

8. 7. The method of claim 6, wherein the one or more training datasets comprise at least one of an isolate-variant matrix, an antibiotic indicator matrix, or an isolate resistance signature vector.

9. The method of claim 6 , further comprising performing a nested cross-validation procedure on the PARP model using a validation dataset comprising multiple bacteria-antibiotic combinations.

10. determining, by the PARP model, one or more classes associated with the plurality of antibiotics or the plurality of bacterial species using weights of one or more dense layers of a Feature Linear Modulation (FiLM) generator to form clusters; The method of claim 6 , wherein the machine learning system uses the one or more classes to output a predictive indicator of antibiotic resistance.

11. 7. The method of claim 6, wherein the PARP model is deployed on a container orchestration service such that the PARP model provides a cloud-based antibiotic resistance prediction service.

12. 12. The method of claim 11, wherein receiving the genome sequence comprises receiving an upload from a remote device at the cloud-based antibiotic resistance prediction service.

13. 7. The method of claim 6, wherein the genome sequence corresponds to a bacterial species not present in the one or more training datasets.

14. 7. The method of claim 6, wherein the predictive indicator of antibiotic resistance comprises a bar graph for display in a graphical user interface (GUI) of a computing device that provided the genome sequence.

15. 15. The method of claim 14, wherein the x-axis of the bar graph represents different antibiotics and the y-axis of the bar graph represents predicted values ​​of resistance or susceptibility to the different antibiotics.

16. 7. The method of claim 6, wherein the PARP model generates shared and unique variant data indicative of one or more variants shared among different bacterial species and one or more variants unique to the different bacterial species.

17. 7. The method of claim 6, wherein training the PARP model comprises generating paired antibiotic susceptibility data based on testing isolates in the antibiotic pair that indicate a shared pathway for the antibiotic pair.

18. 7. The method of claim 6, further comprising performing a predictive accuracy assessment for the predictive index, the predictive accuracy assessment outputting one or more predictive accuracy values ​​corresponding to the one or more antibiotics.

19. and further comprising adjusting a plurality of hyperparameters of the PARP model, the plurality of hyperparameters comprising: Number of dense blocks, the number of FiLM layers or feature stacking blocks, and Geometric size of the dense layer The method of claim 6, comprising:

20. 1. A system for antibiotic resistance prediction, comprising: A pan-antibiotic resistance prediction (PARP) model deployed into a cloud-based service, the PARP model comprising: genetic information associated with multiple bacterial species; and Antibiotic characteristic information associated with the plurality of antibiotics a pan-antibiotic resistance prediction (PARP) model, the machine learning system being trained with one or more training datasets comprising: a web-based portal for receiving genomic sequences associated with particular bacterial isolates and providing the genomic sequences to the PARP model; a predictive indicator of antibiotic resistance for the particular bacterial isolate output by the PARP model and configured for display in a graphical user interface (GUI) of a computing device; and Including, the system.