Pharmacology autoencoder machine learning models, systems and methods of using same
An autoencoder machine learning model trained on pharmacologic signatures addresses data quality and transparency issues in pharmacology, enabling efficient identification of candidate ingredients and targets for microbiome-based treatments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CASE WESTERN RESERVE UNIV
- Filing Date
- 2026-01-21
- Publication Date
- 2026-07-30
AI Technical Summary
Machine learning models in pharmacology face challenges due to the need for high-quality data, complexity of biological systems, and the 'black box' nature, leading to inaccurate predictions and a lack of transparency.
A processor-implemented autoencoder machine learning model is trained on pharmacologic signatures to generate a latent space representation, allowing for the identification of candidate ingredients and targets, and provides a visualization of this representation to enhance understanding and accuracy.
The model enables computationally efficient derivation of novel microbiome-based treatments with high success likelihood, reducing costs and accelerating drug development by providing transparent and accurate predictions.
Smart Images

Figure US2026011961_30072026_PF_FP_ABST
Abstract
Description
PHARMACOLOGY AUTOENCODER MACHINE LEARNING MODELS, SYSTEMS AND METHODS OF USING SAME CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of and priority to U.S. provisional patent application no. 63 / 747418, filed on January 21, 2025, which is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0002] This disclosure relates to machine learning models and, more specifically, to pharmacology autoencoder machine learning models as well as to systems and methods using such models.BACKGROUND
[0003] Machine learning (ML) models are computer programs that recognize patterns in data to make predictions or classifications. For example, machine learning models can analyze vast amounts of biological data to predict how therapeutics interact with the body, identifying potential benefits and adverse effects. However, challenges include the need for high-quality data and the complexity of biological systems, which can lead to inaccurate predictions. The effectiveness of ML models heavily relies on the quality and quantity of data, and incomplete or biased datasets can lead to inaccurate results. Additionally, many ML models operate as "black boxes," making it difficult for users to understand how they arrive at specific predictions, which can hinder trust in their results.SUMMARY
[0004] One example relates to a processor-implemented method. The processor-implemented method includes a processor-implemented method includes receiving input data to an autoencoder machine learning model to provide a latent space representation of feature vectors characterizing pharmacologic factors. The input data can include one or more data items representing one or more ingredients and / or one or more targets. The autoencoder machine learning model can be trained based on training data representing pharmacologic signatures for known ingredients and / or known targets. The method can also include analyzing the latent space representation of feature vectors. The method can also include identifying one or more candidate ingredients and / or one or more candidate targets based on the analysis.
[0005] Another example relate to a system that includes one or more processors and one or more non-transitory computer-readable media having instructions. The instructions, when executed by the one or more processors, cause the one or more processors to:apply input data to an autoencoder machine learning model to generate a latent space representation of feature vectors that encode pharmacologic factors based on the input data, in which the input data includes one or more data items representing one or more ingredients and / or one or more targets, and the autoencoder machine learning model is trained based on training data representing pharmacologic signatures for known ingredients and / or known targets;analyze the latent space representation of the feature vectors; andidentify one or more candidate ingredients and / or one or more candidate targets based on the analysis.
[0006] Yet another example relates to a processor-implemented method. The processor-implemented method includes generating a training dataset from microbiomc -based signatures of individual and combinations of microbiome factors and signatures of drug treatment from a cohort of subjects based on comparison of a quantitative similarity of the signatures of the microbiome factors and drug signatures. An autoencoder machine learning model can be trained using the training dataset.
[0007] Another example relate to a computer-implemented method of using an autoencoder machine learning model to determine at least one formulation of ingredients to achieve a desired biological and / or pharmacological effect or action in response to input data representing one or more ingredients and / or one or more targets.
[0008] Further examples relate to one or more therapeutic agents produced based on formulations derived using system and / or methods described hereinBRIEF DESCRIPTION OF THE DRAWINGS
[0009] FIG. 1 is a block diagram of an example of an autoencoder neural network.
[0010] FIG. 2 depicts an example visualization of a latent space representation that can be generated by the autoencoder neural network of FIG. 1.
[0011] FIG. 3 depicts another example visualization of a latent space representation that can be generated by the autoencoder neural network of FIG. 1 showing a given class and factors in the given class.
[0012] FIG. 4 depicts another example visualization of a latent space representation that can be generated by the autoencoder neural network of FIG. 1 showing feature vectors with a radius of a selected target.
[0013] FIG. 5 depicts another example visualization of a latent space representation that can be generated by the autoencoder neural network of FIG. 1 showing microbiome factors and products relative to a selected target pharmacologic function.
[0014] FIG. 6 depicts another example visualization of a latent space representation that can be generated by the autoencoder neural network of FIG. 1 showing synthetic microbial communities relative to pharmacologic functions, such as drug functions.
[0015] FIG. 7 depicts another example of an autoencoder neural network having an adaptive bottleneck.
[0016] FIG. 8 depicts an example visualization of a latent space representation that can be generated by the autoencoder neural network of FIG. 1 or 7 for identifying metabolites for a target treatment.
[0017] FIG. 9 depicts an example visualization of a latent space representation that can be generated by the autoencoder neural network of FIG. 1 or 7 for identifying metabolites for a target pharmacologic activity.
[0018] FIG. 10 depicts an example visualization of a latent space representation that can be generated by the autoencoder neural network of FIG. 1 or 7 for identifying ingredients and / or targets that would be expected to influence a particular phenotype / biological function.
[0019] FIG. 11 depicts another example visualization of a latent space representation that can be generated by the autoencoder neural network of FIG. 1 or 7 for identifying a synthetic community of factors for a target phenotype, patient condition, or biological state (e.g. successful pregnancy after IVF treatment).
[0020] FIG. 12 is a flow diagram illustrating an example of a computer-implemented method for using a pharmacology autoencoder neural network.
[0021] FIG. 13 is a block diagram illustrating an example operating environment for using a pharmacology autoencoder neural network.
[0022] FIG. 14 is a flow diagram illustrating an example method for training an autoencoder neural network.
[0023] FIGS. 15 through 20 are diagrams and graphs illustrating an autoencoder neural network and its evaluation at different stages of the training method of FIG. 14.
[0024] FIGS. 21 and 22 are example latent space representations that can be generated by the autoencoder neural network trained by the method of FIG. 14 for validation purposes.
[0025] FIGS. 23, 24, and 25 are example latent space representations that can be generated by the autoencoder neural network trained by the method of FIG. 14 for identifying metabolites and microbiota near metronidazole, clindamycin, and clotrimazole.
[0026] FIG. 26 is a workflow diagram illustrating creation of a database that can be used for training the autoencoder neural network of FIGS. 1, 7, and / or 12.DETAILED DESCRIPTION
[0027] This disclosure relates to machine learning models and, more specifically, to pharmacology machine learning models as well as to systems and methods using such models.
[0028] As an example, systems and methods described herein can use an autoencoder machine learning model to identify one or more candidates, such as one or more candidate ingredients and / or one or more candidate targets, in response to input data. The input data can include one or more data items representing one or more ingredients and / or one or more targets depending on the desired output from the model. As one example, the autoencoder machine learning model can identify one or more candidate ingredients in response to input data representing one or more targets or ingredients. As another example, the autoencoder machine learning model can identify one or more targets, corresponding to pharmacologic actions and / or therapeutic effects, in response to input data representing one or more ingredients. As described herein, the autoencoder machine learning model can be trained to be microbiome agnostic to facilitate development of treatments across one or more microbiomes. Alternatively, the model can be trained for one or a combination of specific microbiomes.
[0029] As used herein, the term microbiome can refer to a community of microorganisms that occur in a physiological context. For example, the term microbiomc can encompass bacteria, viruses, fungi, phages as well as their products and / or genes. Examples of microbiomes for a human body include gut microbiome, skin microbiome, oral microbiome, vaginal microbiome, nasal microbiome, lung microbiome, and urothelial microbiome.
[0030] Various autoencoder architectures may be used depending on the training data and application requirements. For example, the autoencoder machine learning model can be designed as an autoencoder neural network (also referred to herein as a neural network or, simply, a network). For example, the network can include a plurality of processing nodes arranged in multiple layers, including an encoder portion, one or more bottleneck portions, and a decoder portion, in which nodes of one layer are connected to nodes of one or more other layers. The network includes a bottleneck portion configured to provide a latent feature space representation that includes feature vectors characterizing one or more pharmacologic factorsfor the input data. The term pharmacologic factors can refer to the properties and effects of therapeutic agents or other ingredients, including their origin, composition, therapeutic uses, and how they interact with biological systems, particularly microbiomes. For example, the latent space representation can encode pharmacologic signature similarity of ingredients and / or targets by embedding host signatures (e.g., proteins and / or gene expressions) of individual and combinations of ingredients along with signatures (e.g., proteins and / or gene expressions) of individual and combinations of targets based on quantitative analysis of the signatures. Thus, the trained autoencoder machine learning model further projects input data, which represents signatures for at least one ingredient and / or at least one target, into the latent feature space representation. The systems and methods described herein can analyze the latent feature space representation to identify one or more candidates, which can include candidate ingredients and / or candidate targets. Systems and methods can include a dashboard or other user platform configured to generate an output to visualize the latent space representation and / or the candidate ingredients and / or candidate targets. For example, the dashboard further can provide an interactive user interface (e.g., graphical user interface) to enable a user to select input data and interact with the latent space representation in response to user input instructions.
[0031] As used herein, the term ingredients can refer to any substance, agent, or combination of substances or agents, including naturally occurring, derived from organisms, or synthetically derived, which can be used to treat, manage, or otherwise influence or modulate a biological and / or pharmacological effect or action. For example, ingredients include therapeutic agents (e.g., drugs), small molecules, biologies, cell, proteins, peptides, nucleic acids, genes or gene products / fragments, microbiome factors, metabolites, other model-designed products, or any combination thereof. Microbiome factors further include microbiome members (e.g., bacteria, viruses, and / or fungi), microbiome products (e.g., proteins, peptides, and / or metabolites).
[0032] As used herein, the term target can refer to any desired therapeutic effect, modulation of a biomolecular or cellular entity, and / or pharmacologic activity, action, or biological entity whose modulation produces a desired therapeutic effect and / or pharmacologic activity. For example, targets include specific cellular (e.g., host or microbiome) receptors, pathways, proteins (e.g. kinase, phosphatase, transcription factor, enzyme), other biological processes or factors a user wants to position ingredients or other model-designed product toward, or any combination thereof.
[0033] In examples herein, the autoencoder machine learning model can be trained based on training data representing pharmacologic signatures, defined by respective similarity scoresof proteins and / or gene expressions (e.g., gene expression perturbations) for known ingredients and / or known targets. In examples described herein, the training data for the autoencoder machine learning model can be derived from a pharmacology database representing a similarity between signatures of known ingredients and / or known targets. For example, the database can be a microbiome pharmacology database representing a quantitative similarity of host signatures between microbiome factors and therapeutic agents (e.g., drugs), such as defining a similarity matrix (e.g., a signature-correlation matrix) that includes a functional similarity score for signatures of each pair of therapeutic agent and microbiome factor. The similarity scores can be determined by quantifying (e.g., by a statistical metric) the similarity of the therapeutic agent signatures with the microbiome signatures.
[0034] As used herein, a database can refer to a table or a set of tables. In still other examples, the term database may refer to a set of data stores and methods for accessing and / or manipulating those data stores. In one embodiment, a database may be stored, for example, at a disk, data store, and / or a memory. A database may be stored locally or remotely and accessed via a network. Additionally, the term memory can refer to volatile memory and / or nonvolatile memory. Non-volatile memory may include, for example, ROM (read only memory), PROM (programmable read only memory), EPROM (erasable PROM), and EEPROM (electrically erasable PROM). Volatile memory may include, for example, RAM (random access memory), synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), and direct RAM bus RAM (DRRAM). The memory may store an operating system that controls or allocates resources of a computing device.
[0035] Advantageously, the systems and methods described herein can help derive novel microbiome-based treatments, in a computationally efficient manner, with a high likelihood of success without the need to do expensive screening experiments, thereby reducing costs and accelerating drug development. For example, the systems and methods described herein can receive model input data, including or derived from existing or proposed formulations of ingredients, and use the model to create new formulations of probiotics, prebiotics, vitamins, minerals, and / or other supplements. The systems and methods described herein can also be used to create microbiome-derived molecules as therapeutics.
[0036] FIG. 1 is a block diagram of an example architecture of an autoencoder neural network 100 that can be used in the systems and methods described herein. The autoencoder neural network 100 includes an encoder portion 102, a bottleneck portion 104, and a decoder portion 106, in which nodes of one layer are connected to nodes of one or more other layers. The encoder portion 102 includes an input layer 108 and one or more hidden layers 110, eachhaving a number of nodes representing corresponding features. The input layer 108 of the encoder has nodes to receive respective entries of input data 112 representing one or more targets and / or one or more ingredients. The input data 112 includes one or more targets and / or one or more ingredients, which has been prepared to include a set of input features mapping to respective nodes of the input layer 110 and having input feature values (e.g., signatures) scaled according to scaling of the training data set. For example, a database including training data of signature similarity scores can prepare user input data into the input data 112 having features mapping to input layer 110 to enable the autoencoder neural network to process the data. The nodes in the input layer 108 are connected to provide a set of features to respective nodes in the hidden layer 110. The encoder portion can include one or more feature encoders between input layer and an output layer thereof, shown as layer 110. The encoder portion is configured to encode or compress the input data 112 into an encoded set of features provided at the output layer 110. The output layer of the encoder portion 110 has nodes connected to respective nodes of the bottleneck portion 104.
[0037] The bottleneck portion 104 is configured to provide a latent feature space representation 114 that includes feature vectors characterizing one or more pharmacologic factors for the input data 112 in latent space. As an example, the bottleneck portion 104 includes three nodes representing three features in three-dimensional latent space 114. The bottleneck portion 104 can include other numbers of nodes (e.g., 4 nodes, 5 nodes, 6 nodes, 7 nodes, etc.) to provide other dimensional latent spaces into which encoded input data is projected. The bottleneck portion 104 can project the input data 112, as corresponding feature vectors, into the latent space representation 114. In some examples, the latent space representation 114, into which the input data is projected, defines a universal latent space that embeds encoded host signatures (e.g., proteins, gene expression) of individual and combinations of microbiome factors (bacteria, viruses, fungi and their products) along with host signatures (proteins, gene expression) of drug treatment based on quantitative similarity of the signatures between the drugs and microbiome factors. As a result, the latent space representation can provide a side-by-side feature comparison in latent space between encoded features of the input data and known data (e.g., data for ingredients and / or targets having known and / or desired pharmacologic properties). The known data can include a portion or all pertinent training data. Also, or alternatively, the input data 112 can be selected in response to a user input.
[0038] Additionally, the decoder portion 106 includes an input layer 116 and one or more hidden layers 118, and an output layer, each having a number of nodes representingcorresponding features. The bottleneck layer 104 also can include an output layer having nodes connected to respective nodes of the input layer 116 of the decoder portion 106. The output layer 118 of the decoder portion 106 can include the same number of nodes as the input layer 108 of the encoder portion 102. The decoder portion 106 is configured to reconstruct the input data based on the latent space representation 114 provided by the bottleneck portion 104. In some examples, the decoder portion 106 can be utilized to reconstruct at least some of the input data from the latent space representation to identify similar known targets and / or ingredients in the input / training space.
[0039] As described herein, a graphical representation of the latent space representation 114 can be generated to provide visualization of the organization of latent space for the input data and known data. For example, the graphical representation can enable a user to visualize how data points are organized based on their features, such as by showing relationships (e.g., clusters of features) in a lower-dimensional space. FIGS. 2-6 depict example visualizations of latent space representations for feature vectors that can be generated by an autocncodcr neural network as described herein (e.g., network 100 of FIG. 1). In the examples of FIGS. 2-6, a visualization scale 154 (e.g., a color scale, a pattern-coded scale, or the like) can be rendered in the visualization based on metadata associated with the input data and / or feature vectors. The metadata can include other information associated with the data in the visualization, such as a class of the data and / or an identifier specifying the data being projected in the latent space. In some examples, graphical and / or textual features can be rendered in the visualization based on metadata for one or more feature vectors in response to a user input (e.g., hovering a pointer of a GUI over parts of the visualization).
[0040] FIG. 2 depicts an example visualization 150 of a latent space representation of input data projected as feature vectors 152 in three dimensions defined by variables VI, V2, and V3. Each of the feature vectors 152 embedded in the latent space can include metadata characterizing the type (e.g., class) of input data being projected therein. In the example of FIG. 2, the projected data includes is labeled according to its class, including drugs, metabolites, microbiota, and bacteria, which is indicated by a visualization scale 154 (e.g., a color scale, a pattern-coded scale, or the like).
[0041] FIGS. 3, 4 and 5 depict example visualizations 160, 170, and 180 of respective latent space representations, in which the spatial proximity of microbiome factors to drugs in the latent space can be used to identify mechanistic, pharmacologic functions of the microbiome members or products. The example visualization 160, 170, and 180 as well asother similar visualizations can provide useful tools to identify microbiome factors and products that perform similar functions to the drug of interest.
[0042] In the example of FIG. 3, the visualization 160 includes feature vectors 162 that can be generated by the autoencoder neural network of FIG. 1 showing a given class of data and factors in the given class, as indicated by scale 164. Specifically, the visualization 160 shows mapping of estrogen receptor modulators and microbiome factors in that class embedded in the latent space. The visualization 170 of FIG. 4 includes feature vectors 172 showing a set of ingredients mapping to specific targets indicated by scale 174, namely, a latent space representation of lactobacillus crispatus drugs and metabolites. In the example FIG. 4, a cluster of the bacteria is identified based on relative proximity to the target(s), shown by a radius 176 (e.g., a radius of 0.25 rad). The visualization 180 of FIG. 5 includes feature vectors 182 showing a mapping of potential antibiotic functions of microbiome factors and products (e.g., corresponding to metronidazole), as indicated by scale 184. Additionally, a cluster of microbiome factors (e.g., products and members) is identified in the latent space based on relative proximity to one or more targets, shown by a radius 186 (e.g., a radius of 0.5rad). The specified target(s) and / or radius can be set as default values or be programmable in response to a user input to enable further customization of targets and ingredients.
[0043] FIG. 6 depicts another example visualization 190 that can be generated by the autoencoder neural network of FIG. 1 showing a latent space representation of feature vectors 192 for a number of classes, as indicated by scale 194. Specifically, the visualization 190 includes feature vectors 192 for synthetic microbial communities (e.g., one or more simulated combinations of microbiome factors, such as microbiome products and / or microbiome members) embedded in the latent space relative to pharmacologic functions, such as drug functions. The latent representation thus may be analyzed (e.g., through decoder portion 106 or another mapping function) to extract a predicted pharmacologic profile of one or more simulated combinations of microbiome factors.
[0044] FIG. 7 depicts an architecture of another example of an autoencoder neural network 200 (also referred to as a network or model) that can be used to identify individual or combinations of candidate ingredients and / or targets by projecting input data representing the ingredients and / or targets into a latent space of the network 200. The network 200 includes an encoder portion 202, a bottleneck portion 204, and a decoder portion 206. The encoder portion 202 includes an input layer 208 and one or more other layers 210, such as hidden layers. The input layer 208 includes a number of neurons (or nodes) based on the number of data units (e.g., entries) in the input data to be received by the network, such as defined by respectivesignature scores of training data. Examples of numbers of neurons for each of the layers of the network are shown above each layer, and can vary based on training data and / or training of the network. The encoder portion 202 is trained to encode the input data and provide encoded data (e.g., an encoded features set) representing encoded features of input data. The encoder portion 202 includes an output layer 212 that maps to the bottleneck portion 204.
[0045] The bottleneck portion 204 includes a layer having neurons configured to reduce the dimensionality of the encoded data and to provide feature vectors in the latent space representation for each unit of the input data. The bottleneck portion 204 can include three or other numbers of neurons, which defines the dimensionality of the latent space into which the input data is projected. As an example, the latent space representation can be a universal latent space that embeds encoded host signatures (e.g., proteins, gene expression perturbations) of individual and combinations of microbiome factors (bacteria, viruses, fungi and their products) along with host signatures (proteins, gene expression perturbations) of drug treatment based on quantitative similarity of the signatures between the drugs and microbiomc factors. In some examples, the bottleneck portion 204 of autoencoder neural network 200 has an architecture that is adapted during training. Examples of training that can be implemented to train the network 200 are described herein with respect to FIGS. 14-21.
[0046] The decoder portion 206 includes an input layer 214, an output layer 216, and one or more hidden layers 218 between the input and output layers, and an output layer, each having a number of nodes representing corresponding features. The nodes of the bottleneck layer 204 are connected to respective nodes of the input layer 214 of the decoder portion 206. The output layer 216 of the decoder portion 206 can include the same number of nodes as the input layer 208 of the encoder portion 202. The decoder portion 206 is configured to reconstruct the input data based on compressed data in the latent space representation of the bottleneck portion 204. In some examples, the decoder portion 206 can be utilized to reconstruct at least some of the projected data from the latent space representation to identify similar known targets and / or ingredients in the input / training space.
[0047] FIGS. 8-11 depict examples of visualizations of latent space representations that can be generated by the autoencoder neural network 200 of FIG. 7. Similar to FIGS. 2-6, the visualizations can include respective scales for differentiating types (or classes) of ingredients and / or targets projected in the latent space. In the examples of FIGS. 8-11, the input data provided to the network 200 can include ingredients and / or targets (e.g., drugs, patients, microbiota, metabolites, microbiome factors, etc.), represented by signatures of scaled scores for proteins or gene expression perturbations. The projection of the input data in the latentspace representation can define output data for the network 200, which can specify one or more pharmacologic properties, novel combinations, predicted targets, and / or patient-specific mappings. The type of output and / or input data can be set for the network 200 in response to control instructions, such as in response to a user input instruction. The examples of FIGS. 8-11 demonstrate examples of some different uses of the autoencoder neural network described herein. Other uses of the autoencoder neural network are possible. Additionally, while the examples of FIGS. 8-11 demonstrate a three-dimensional latent space, other dimensions (e.g., higher dimensions) can be used in other examples.
[0048] FIG. 8 depicts an example visualization 250 of feature vectors 252 projected in a latent space representation for identifying ingredients to achieve a target treatment. Specifically, the visualization 250 of the latent space representation produced by the autoencoder neural network can be used for identifying vaginal metabolites for endometrial cancer treatment and / or prognosis. For example, the latent space representation can be analyzed to understand how the production and consumption of metabolites by vaginal taxa affect tumor viability. As shown in FIG. 8, a radius of compounds, shown at 254, identifies compounds around anti-cancer agent in latent space (e.g. cisplatin). The compounds within the radius 254 further be analyzed to rank metabolites (or other ingredients) for enabling a rank-based selection of metabolites based on proximity to one or more desired drugs to mimic.
[0049] FIG. 9 depicts an example visualization 260 of a latent space representation of feature vectors 262 that can be generated by the autoencoder neural network for identifying metabolites for a target pharmacologic activity. Specifically, the visualization 260 of the latent space representation produced by the autoencoder neural network can be used to identify vaginal metabolites with antibiotic potential. For example, loss of vaginal homeostasis and overgrowth of anaerobes is called bacterial vaginosis and it has been shown that 50% of cases are recurrent and many plagued by antibiotic resistance. The visualization can identify a radius of compounds around a known antibiotic agent in latent space (e.g. clindamycin). Other antibiotic agents can also be identified in the visualization (e.g., hydroxyisocaproate, taurine). Associated tools can rank the proximity of metabolites (or other ingredients) to enable a rankbased selection of metabolites based on proximity to the desired drug (or desired drug combination) to mimic.
[0050] FIG. 10 depicts an example visualization 270 of a latent space representation of feature vectors 272 that can be generated by the autoencoder neural network for identifying pharmacologic factors for a target phenotype (e.g., ingredients and / or targets that would be expected to influence a particular phenotype / biological function, demonstrated here assuccessful pregnancy after IVF in a VMFI cohort). Specifically, the visualization 270 of the latent space representation produced by the autoencoder neural network can be used for mechanism-based design of eubiotics, which are formulations of supplements and therapies containing probiotics, prebiotics, and / or postbiotics. For example, an input data set can be provided for individuals having a desired phenotype (e.g., microbiome, RNA-seq, metabolomics, etc.). The input data set can be mapped into the autoencoder neural network of FIG. 1 or 7 to provide the visualization 270. The latent space representation and / or visualization 270 thereof can be analyzed (e.g., automatically and / or manually) to identify a candidate factor (or factors) in a radius 274 of the target phenotype in latent space. The candidate factors can be utilized to simulate and extract combinations of additional candidate factors that map one or more formulations to the target phenotype. The process can be repeated until a formulation(s) is within a threshold distance / proximity of the target.
[0051] FIG. 11 depicts another example visualization 280 of a latent space representation feature vectors 282 that can be generated by the autocncodcr neural network, as described herein, for identifying a synthetic community of factors for a target phenotype. The visualization 280 can be used for mechanism-based design of eubiotics based on an input data set, such as can be provided for individuals having a desired phenotype (e.g., microbiome, RNA-seq, metabolomics, etc.). The input data set can be mapped into the autoencoder neural network of FIG. 1 or 7 to provide the visualization 280. The latent space representation and / or visualization 280 thereof can be analyzed (e.g., automatically and / or manually) to identify a plurality of candidate factors in a radius 284 of the target phenotype in latent space. The candidate factors can be utilized to simulate and extract combinations of additional candidate factors that map one or more formulations to the target phenotype. The process can be repeated until a formulation of factors is within a threshold distance / proximity of the target phenotype in the latent space.
[0052] FIG. 12 is a block diagram illustrating an example system 300 implementing an autoencoder neural network 302. One or more instances of the network 302 can be used in the system 300 according to examples described herein, including FIGS. 1 and 7. Also, or alternatively, the system 300 can be configured to implement the method 400 of FIG. 13. Accordingly, the description of FIG. 12 may refer to certain aspects of FIGS. 1, 7, and 13. Further, the system 300 can include one or more computing devices, each having one or more processors and memory (e.g., one or more non-transitory machine-readable media). The memory can store instructions and data, in which the instructions, when executed by one or more processors, cause the one or more processors to perform functions and methods describedherein, including one or more instances of the network 302. Those skilled in the art will understand various computing architectures that may be used to implement the system. For instance, the system 300 can include distributed computing systems, cloud computing platforms, and edge devices. Additionally, the memory used to store instructions and data can include local memory, shared memory, distributed memory, and hybrid memory systems.
[0053] As used herein, the term computer-readable medium can refer to a non-transitory medium that stores instructions and / or data. For example, a computer-readable medium may take forms, including, but not limited to, non-volatile media, and volatile media. Non-volatile media may include, for example, optical disks, magnetic disks, and so on. Volatile media may include, for example, semiconductor memories, dynamic memory, and so on. Common forms of a computer-readable medium may include a floppy disk, a flexible disk, a hard disk, a magnetic tape, other magnetic medium, an ASIC, a CD, other optical medium, a RAM, a ROM, a memory chip or card, a memory stick, and other media from which a computer, a processor or other electronic device may read, directly or indirectly through a physical or logical connection.
[0054] In the example of FIG. 12, the 300 system includes an autoencoder interface and analysis module 304 (e.g., instructions) that is coupled to the autoencoder neural network 302. The autoencoder interface and analysis module 304 includes input data preparation module 306 and a latent space analyzer module 308. The data preparation module 306 is operative to prepare input data for receipt by the autoencoder neural network 302 and apply the prepared input data to the respective nodes of the input layer for processing by the network 302, such as described herein. The latent space analyzer module 308 is operative to analyze a latent space representation provided by the autoencoder neural network in response to the input data. As used herein, the term module can refer to non-transitory computer readable medium that stores instructions, instructions in execution on a machine, hardware, firmware, software in execution on a machine, and / or combinations of each to perform a function(s) or an action(s), and / or to cause a function or action from another module, method, and / or system. Multiple modules may be combined into one module and single modules may be distributed among multiple modules.
[0055] As a further example, the input data preparation module 306 receives user data 310 from a user computing device 312, which can be coupled to the computing device implementing the autoencoder interface and analysis module 304 through one or more networks 314. The network(s) 314 can include one or more local area networks (LAN), widearea networks (WAN), the internet, and the like. The user data 310 can include or describe one or more ingredients and / or one or more targets, such as described herein.
[0056] As an example, the user data 310 can be in the form of table or other data structure, including a table with rows for records identifying each ingredient (e.g., by name or another identifier), columns for candidate formulations, and entries as percentages / proportions where the columns sum to 100%. The table can be user-defined or can be generated by the input set generator 318 (e.g., via random sampling / simulation). The table may be generated automatically or in response to a user input specifying a desired target and / or one or more ingredients. An example ingredient / formulation table is shown in Table 1 (where I and J are positive integers representing the number of formulations and ingredients, respectively and entries are percentages of the ingredient in a given formulation such that the set of percentages arranged in each column sums to 100%).TABLE 1
[0057] The input data preparation module 306 can include a database (DB) query module 316 and an input data set generator module 318. For example, the DB query module is programmed to query a database 320 based on the user data 310. The database 320 can be a microbiome pharmacology database representing a quantitative similarity of host signatures between microbiome factors and therapeutic agents (e.g., drugs), such described herein. The DB query module 316 can provide ingredients and / or targets in the user data 310 to query the database 320. The input data set generator 318 is operative to generate one more files, such as the form of tables or other data structures, including a respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets represented in the user data based on pharmacological profile scores stored in the database 320. As described herein, the pharmacologic signature for each of the one or more ingredients, formulations, and / or the one or more targets in the input data set can include a score for respective data entries thereof that map to (are readable by) respective neurons of the input layer of the autoencoder neural network 302. In some examples, the scores in the input data can be scaled (e.g., by scalingmodule 322) according to the scale defined by respective entries of records in the database 320 (e.g., correlated similarity scores for proteins and gene expression values on a normalized scale).
[0058] In some examples, where an ingredient or target is not in the database, a new database entry can be constructed for that factor, such as using either user-provided data or analysis of public resources to put the new factor on the scale of the microbiome pharmacology database. The model may be retrained to enable processing of such data. The retraining can be initiated automatically or in response to a user input. Alternatively, a message can be sent to alert an authorized user to initiate retraining of the model.
[0059] The input set generator 318 is operative to construct one or more input files as tables that includes data values configured and arranged to be applied to respective neurons of the input layer of the autoencoder neural network 302. As an example, the following tables (TABLE 2, TABLE 3, and TABLE 4) demonstrate some examples of tables, defining input data, which can be fed into the trained autocncodcr neural network 302, individually or collectively, for generating one or more latent space representations for each input data record. While shown as separate tables, the tables can be aggregated into a single table for inputting into the network 302 or may have other arrangements. In the following tables N represents the number of neurons in the input layer of the network 302, S represents a score value, K represents the number of ingredients, L represents the number of formulations, and M represents the number of formulations. It is to be understood that the score values, “S” for each entry can be different even though represented the same in one or more tables (i.e., in different tables Sl-1 AS1-1 or S 1 -1=S-1 -1).
[0060] Table 2 is an example of a table, defining input data, in which each record represents a different ingredient (e.g., substance, agent, or combination thereof) and entries of each record include respective scores (e.g., pharmacological profile similarity scores) that define a signature for a respective ingredient. In this context, different ingredients listed of Table 2 may represent the same item but in different concentrations or may represent different items (e.g., Ingredient l=metabolite A in concentration X, Ingredient 2=metabolite A in concentration Y, Ingredient 3=metabolite B in concentration Z, Ingredient 4=small molecule D, etc.). Thus, the name or other identifier in the list is used to uniquely identify each ingredient.TABLE 2
[0061] Table 3 is an example of a table, defining input data, in which each record represents a different formulation (e.g., a combination of multiple ingredients at respective concentrations) and entries of each record include respective scores (e.g., pharmacological profile similarity scores) that define a signature for a respective formulation.TABLE 3
[0062] Table 4 is an example of a table, defining input data, in which each record represents a different target (e.g., therapeutic effect and / or pharmacologic activity, action, or biological entity whose modulation produces a desired therapeutic effect and / or pharmacologic activity) and entries of each record include respective scores (e.g., pharmacological profile similarity scores) that define a signature for a respective formulation.TABLE 4
[0063] The input data set generator 318 can also include a metadata generator 324 and / or an expansion module 326. The expansion module 326 can cause a processor to expand one or more records of the input data returned by the database 320, corresponding to an individual ingredient or combination of ingredients, to include additional combinations and / or concentration of ingredients, additional formulations and / or additional targets. For example, the expansion module 326 can generate additional combinations in response to a user input instruction from the user computing device 312. The expansion module 326 can derive the additional combinations of a set of ingredients, which may include the same or different numbers of ingredients and have concentrations thereof that are the same or different for the respective ingredients. The metadata generator 324 further can cause a processor to add (e.g., append) metadata to each record of input data, such as including a unique identifier for the respective records and / or other descriptive information (e.g., class type) for ingredients or targets including additional ingredients, targets, and / or formulations created by the expansion module 326. The resulting input data set (e.g., one or more tables of ingredients, formulations, and / or targets) can be stored in memory and provided to the autoencoder neural network 302 for processing. The resulting data set can thus include ingredients and / or targets explicitly specified in the user data as well as additional combinations derived (e.g., by expansion module) based on the user data, in which data items for each such record of input data has a respective pharmacologic signature.
[0064] In some examples, the input data preparation module 306 can control respective functions (e.g., DB query 316 and / or input data set generator 318 and component modules thereof) in response to user input instructions provided at the user computing device 312. As an example, the user computing device 312 includes one or more processors 330 and memory 332. The memory 332 can store data and instructions, in which the instructions are operative to cause the processor to perform functions and methods, shown as including an autoencoder dashboard 334. The autoencoder dashboard 334 can include an autoencoder API 336 and a GUI 338. The autoencoder API 336 includes a set of protocols that enables the dashboard to communicate with the autoencoder interface and analysis module 304, enabling the dashboard to access and utilize functions and methods of the autoencoder interface and analysis module 304. The GUI 338 can include GUI elements (e.g., buttons, menus and the like) operative to enable users to interact with the autoencoder dashboard 334 for accessing and communicating data (e.g., user data 310) and instructions associated with using and / or controlling the autoencoder neural network 302. For example, the GUI 338 can select the user data 310 from the memory 332 in response to a user input enter through a user input device(s) 340 and sendthe user data (e.g., as a request via the autoencoder API 336) to the autoencoder interface and analysis module 304. The autoencoder dashboard 334 further can receive responses from the autoencoder interface and analysis module 304, which can include output data representing a latent space representation generated by the autoencoder neural network 302 and / or analysis data derived from the latent space representation, such as described herein.
[0065] In an example, to enable communication, display, and / or manipulation of data, the user computing device 312 further can include a network interface 344, an input interface 346, and a display interface 346, which are each operably connected for communication via a bus and / or other wired and wireless communication technologies. The network interface 344 provides software and hardware to facilitate data input and output between the components of the computing device 312 and other components, the network(s) 314, and data sources, such as described herein. The memory 332 may also store an operating system that controls or allocates resources for implementing the autoencoder dashboard 334. The network 104 serves as a communication medium to various remote devices (e.g., databases, web servers, remote servers, application servers, intermediary servers, client machines, and other portable devices).
[0066] The display interface 116 provides software and hardware to facilitate data input and output between the autoencoder dashboard 334 and a display 348. The display 348 is a device for outputting information and may be a light-emitting diode (TED) display panel, liquid crystal display (ECD) panel, a plasma display panel, and touch screen displays, among others. The display 348 can include graphical input controls for the GUI 338, which can include software and hardware-based controls, interfaces, touch screens, or touch pads or plug and play devices for a user.
[0067] As described herein, a bottleneck layer (e.g., layer 104, 204) of the autoencoder neural network 302 constructs a latent space representation of feature vectors responsive to the input data received from the autoencoder interface and analysis module 304. The latent space analyzer 308 can cause a processor to analyze the latent space representation for identifying one or candidate ingredients and / or targets. The analysis and identification performed by the latent space analyzer 308 can be an automated process responsive to the latent space representation that is generated. Also, or alternatively, the functions performed by the latent space analyzer 308 may be controlled in response to user input instructions (entered at the user input device 340) communicated through the autoencoder API 336.
[0068] In the example of FIG. 12, the latent space analyzer 308 includes one or more visualization tools 350, a proximity calculator 352, a clustering module 354, and a decoder 356, each of which can cause a processor to perform respective functions described herein.
[0069] The visualization tool 350 and / or other functions of the latent space analyzer can cause a processor to communicate output data to the user computing device 312 via the autoencoder API 336. The output data can include an identification of one or more candidate targets and / or candidate ingredients. Also, or alternatively, the output data can include a graphical representation of the latent space representation generated by the autoencoder neural network 302. For example, the user computing device 312 can generate, based on the output data, a visualization of the latent space representation on the display 348. In some examples, the GUI 338 can provide the visualization as an interactive graphical representation, which a user can interact with through the input device 340, such as to implement controls to manipulate the visualization (e.g., rotate the visualization of the latent space, change viewing angle, zoom in or out, etc.), in response to user commands communicated to the visualization tool 350 via the autoencoder API 336. As another example, a user can control selection of one or more feature vectors in response to user selection input(s) entered through the input device 340, which can be selected to visualize additional information associated with the selected feature vector(s) (e.g., metadata identifying the corresponding ingredient or target) and / or to trigger further analysis with respect to the corresponding ingredient or target for the selected feature vector(s). Examples of visualizations that can be generated (by visualization tool 350) are shown in FIGS. 2-6 and 8-11.
[0070] The proximity calculator 352 can be operative to determine a proximity of generated feature vectors (generated in response to the input data) relative to one or more selected feature vectors in latent space. For example, the one or more selected feature vectors can correspond to one or more known ingredients and / or targets forming the universal latent space representation, which can be provided with the input data or selected in an interactive visualization through the GUI 338 in response to a user input. In an example, the proximity calculator 352 can be programmed to compute a Euclidean distance between each generated feature vector and the selected feature vector(s). Other distance metrics may be used in other examples. The computed distance may be stored in memory. The latent space analyzer 308 can analyze the computed distance values to identify one or more candidate ingredients and / or the one or more candidate targets, which candidates can be stored in memory. Also, or alternatively, the computed distance values can be provided to the user computing device 312 (via the autoencoder API) to provide the proximity information on the visualization of the latent space representation and allow a user to evaluate and manually identify one or more candidate ingredients and / or the one or more candidate targets through the GUI 338.
[0071] In some examples, the clustering module 354 can cluster one or more feature vectors in the latent space representation. For example, the clustering module 354 can be operative to select at least one feature vector or location in the latent space representation corresponding to one or more known ingredients and / or targets. The selected feature vector can be specified as one or more ingredients or targets, which can be provided with the input data to the autoencoder neural network 302 or other data. The clustering module 354 can further be operative to cluster feature vectors in the latent space representation within a distance threshold of the selected at least one feature vector. The distance threshold (e.g., a radius in latent space) can be a fixed value or be a variable value, which can be adjusted based on the relative spacing of feature vectors and / or in response to a user input. Alternatively, the distance threshold can define a number (P) of feature vectors, and the clustering modules can be operative (e.g., by implementing a nearest neighbor algorithm) to identify a set of P feature vectors nearest to the selected feature vector or location. For example, a user can adjust the distance threshold or number P (through the GUI 338 in response to a user input entered at input device 340) to increase or decrease the feature vectors being identified. The clustering module 354 further can be operative to generating a candidate list based on the clustered features. The candidate list can specify one or more candidate ingredients and / or the one or more candidate targets based on the identified feature vectors and / or identify the feature vectors in latent space. The clustering module 354 further can be operative to rank candidate ingredients and / or targets in order of proximity to the selected feature vector or location (e.g., determined by proximity calculator 352). In some examples, the decoder module 356 is operative to map one or more selected feature vectors from the latent space back into input space. The decoder module 356 can be implemented by the decoder portion (e.g., decoder portion 106, 206) of the autoencoder neural network 302 or by a separately training decoder for reconstructing ingredients and targets (in the input data set) according to feature vector location in latent space. The candidate list can be stored in memory and sent (e.g., through the autoencoder API 336) to the autoencoder dashboard 334 for visualizing on the display 348 with the GUI 338.
[0072] The autoencoder dashboard 334 can be used to enable the user to determine whether further analysis is needed or if the candidate list of ingredients and / or targets is satisfactory. For example, a user can interact with GUI 338 to select some or all the ingredients and / or targets in the candidate list for further analysis. In one example, the further analysis includes providing the selected ingredients and / or targets to the input set generator (shown byarrow 358) for generating an updated input data set based on at least some of the ingredients and / or targets identified in the candidate list.
[0073] In some examples, the input data set generator 318 can invoke the expansion module 326 to expand the candidate ingredients and / or targets (e.g., as mapped into the input space) to include additional combinations of ingredients and / or targets. The additional combinations of ingredients and / or targets further may be generated by adjusting a concentration of one or more of the candidate ingredients or adjusting features of candidate targets. Also, or alternatively, the latent space analyzer 308 can be operative to command the input set generator 318 to generate the update input data set in response to determine that further analysis is needed (or desired). Thus, the updated input data can be generated automatically (e.g., by latent space analyzer 308 and input set generator 318) or in response to a user input provided (via the autoencoder dashboard) at the user computing device 312.
[0074] FIG. 13 is a flow diagram illustrating an example of a computer-implemented method 400 for using an autocncodcr machine learning model, such as one or example models (e.g., model 100, 200, 302) described herein, including FIGS. 1, 7, and 12. Accordingly, the description of FIG. 13 may refer to certain aspects of FIGS. 1, 7, and 12. The method 400 can be a computer-implemented method, which can be stored in memory as machine-readable instructions that when executed cause one or more processors to perform the method. The method 400 can be performed by an individual computing device or multiple computing devices, such as in a distributed computing system, cloud computing platform, or other configuration.
[0075] At 402, the method 400 includes receiving input data. For example, the input data includes user data provided by a user describing one or more candidate ingredients and / or targets. At 404, a database (e.g., database 320) can be queried based on the received input data. As described herein, the database can be a microbiome pharmacology similarity matrix including pharmacologic signatures for predetermined targets and ingredients. At 406, one or more input data sets can be generated based on the querying. The input data set(s) can be generated (at 406, e.g., by input data set generator 318) to include a respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets. The targets in the input data set can include substances, agents, formulations, or combinations thereof, such as described herein. For example, the input data set(s) can include data records for respective ingredients and / or targets, and each record includes data entries having respective scores (e.g., scaled pharmacological similarity scores). The scores for a given ingredient or target record define the signature for the given ingredient or target. As described herein, the pharmacologicsignature for each of the one or more ingredients and / or the one or more targets in the input data set(s) can include scores for respective data entries (see, e.g., Table 2, Table 3, Table 4). The data entry for each record (ingredient or target) in the input data set can map to a respective neuron of an input layer of the autoencoder machine learning model.
[0076] In some examples, the input data set(s) are generated at 406 (e.g., by expansion module 326) to further include combinations of the one or more ingredients received in the user data, each having a respective pharmacologic signature. For example, the user data may include a number x of ingredients, formulations, or targets, and the number may be expanded by increasing the number of ingredients, formulations, or targets with additional combinations thereof, which further may include different concentrations. Each of the combinations defines a unique combination having a respective pharmacologic based on pharmacological profile scores stored in the database.
[0077] At 408, the input data set is provided to an input layer to the trained autoencoder machine learning model (e.g., model 100, 200, 302) and, at 410, a feature representation is generated in latent space of the autoencoder machine learning model based on the input data set. As described herein, the autoencoder machine learning model can be trained based on training data representing pharmacologic signatures for known ingredients and / or known targets. For example, the autoencoder machine learning model is implemented as an autoencoder neural network having an architecture that includes an encoder portion having an input layer and one or more other layers, in which the encoder portion is trained to encode the input data and provide encoded data representing features of the input data set generated at 406. The autoencoder neural network also includes a bottleneck portion operative to generate (at 410) feature vectors in the latent space representation (e.g., a universal latent space) based on the encoded data. The autoencoder neural network can also include a decoder portion operative to reconstruct the input data set based on latent space representation. In some examples, data representing the feature vectors in the latent space representation can be provided to the user computing device (e.g., via autoencoder API) for generating a user-perceptible output (e.g., a visualization on display 348). The visualization may be interactive in response to a user input, as described herein.
[0078] The method 400 further includes analyzing (e.g., by latent space analyzer 308) at least a portion of the feature vectors in the latent space representation and identifying one or more candidate ingredients and / or one or more candidate targets based on the analysis (e.g., by decoding or otherwise mapping selected feature vectors from latent space into input space). The targets can include substances, agents, or formulations, such as described herein. In theexample of FIG. 13, at 412, the method 400 includes evaluating the proximity of feature vectors to a location of one or more selected items (e.g., feature vectors for selected ingredients and / or targets) in latent space. For example, a proximity of feature vectors generated based on the input data set can be determined (e.g., by proximity calculator 352, such as using a distance metric in latent space) relative to the selected item(s). The selected item(s) can be selected (e.g., tagged in metadata) for the at least one known ingredient and / or target having desired pharmacologic properties. The selected item(s) further can be projected into the latent space representation as part of the input data set have been projected separately from the input data set. The one or more candidate ingredients and / or the one or more candidate targets can be identified based on the determined proximity.
[0079] At 414, the method includes clustering (e.g., by clustering module 354) feature vectors in the latent space representation. For example, the clustering (at 414) includes performing a nearest neighbor algorithm to identify a set of feature vectors within a distance threshold of the selected itcm(s) in latent space (e.g., a universal latent space). The clustered features can include targets and / or features that were in the input dataset and / or entries in the training data. At 416, a candidate list of the one or more candidate ingredients and / or the one or more candidate targets can be generated (e.g., by clustering module) based on the clustered features. The candidate list can be generated to identify one or more feature vectors in latent space and / or to identify the corresponding ingredients and / or targets in input space, such as by mapping (e.g., by decoder module 356) feature vectors from the latent space to the input space. The candidate list can be provided to the user computing device (e.g., via autoencoder API) for generating a user-perceptible output (e.g., a visualization on display 348).
[0080] At 418, a determination is made whether further analysis is needed or the candidate target(s) and / or ingredient(s) are satisfactory. In response to determining (e.g., automatically or responsive to a user input) that no more analysis is needed, the method can proceed to 420. At 420, the results can be output (e.g., by the autoencoder interface and analysis module 304). For example, output data representing the candidate target(s) and / or ingredient(s) can be stored in memory. Also, or alternatively, the output data can be reported to a user, such as to the autoencoder dashboard 334 via the autoencoder API 336 or by other messaging technologies. As described herein, the autoencoder dashboard 334 can generate a graphical, textual, and or hybrid representation based on the output data, such to visualize the candidate list on the display 348.
[0081] In response to determining (e.g., automatically or responsive to a user input) that more analysis is needed, the method can proceed to 422 for performing additional processing.At 422, a processing request (e.g., a command) is generated. The request can include the list of candidate ingredients and / or targets. In some examples, the request can further include user-selected ingredients and / or targets in response to a user input specifying additional ingredients and / or targets that were not in the original input data set. From 422, the method can return to 406. At 406, an updated input data set can be generated (e.g., by input data set generator 318) based on the ingredients and / or targets in the processing request. The updated input data set can include simulated permutation data derived based on the list of candidates (provided at 416). For example, the simulated permutation data can include additional combinations and / or adjusted concentrations of ingredients in the list of candidates and / or permutations of targets provided in the list of candidates (e.g., generated by input data set generator 318), such as by performing interpolation, random sampling, and / or simulation based on the candidates. Each newly added ingredient or target can include an identifier (e.g., name) to uniquely identify each respective ingredient and target in the updated input data set. The steps at 408 to 418 can be repeated (one or more times) based on the updated input data set to generate an updated set of feature vectors in the latent space representation and provide an updated version of the candidate list.
[0082] FIG. 14 is a flow diagram illustrating an example method 450 for training an autoencoder neural network. The method can be used to train the autoencoder neural network of FIG. 1, 7, or 12. The method 450 can be a computer-implemented method, which can be stored in memory as machine-readable instructions that when executed cause one or more processors to perform the method. The method 450 can be performed by an individual computing device or multiple computing devices, such as in a distributed computing system, cloud computing platform, or other computing environments. FIGS. 15 through 20 are diagrams and graphs illustrating an autoencoder neural network and its evaluation at different stages of the training method of FIG. 14. Accordingly, the description of FIG. 14 refers to certain aspects of FIGS. 1, 7, and 12 as well as to FIGS. 15 through 20. Hyperparameters for training can be defined for training, such as including the number training epochs, a batch size, a learning rate, the tanh activation function, and an error metric (e.g., mean-square error loss function). Other hyperparameters can be used in other examples. The autoencoder model (e.g., autoencoder neural network model) can be trained on GPUs using the CUDA library. Other training environments are possible.
[0083] At 452, the method includes generating a database. For example, the database (e.g., database 320) can be generated, based on microbiome multi-omics data, as a microbiome pharmacology database (e.g., a similarity matrix) representing a quantitative similarity of hostsignatures between ingredients and targets, such described herein. As a further example, the database can include pharmacologic signatures, defined as similarity of host-perturbed gene expression / protein signature associated with microbiome factors (e.g., bacteria, virus, fungi, metabolites, peptides, gene products, secreted factors, etc.), small molecules, and drugs. The pharmacologic signatures can be constructed via statistical analysis of clinical samples, drug screening, and perturbation biology datasets. The database can be implemented as a table or other database format (e.g., relational or other non-relational database formats). As described herein, the pharmacologic signature for each of the one or more ingredients, formulations, and / or the one or more targets in the database can include a score for respective data entries thereof that will map to (are readable by) respective neurons of the input layer of the autoencoder neural network being trained by the method 450. An example process for generating the database is described with respect to FIG. 26.
[0084] At 454, the method includes defining an autoencoder model architecture. For example, the autocncodcr model architecture can be an autocncodcr neural network (e.g., network 100, 200, 302) that includes an encoder portion, a bottleneck portion, and a decoder portion, rhe neurons of each layer can include an activation function (e.g., tanh or other function). Additionally, an input layer of the autoencoder neural network defined at 454 includes a neuron for each drug / perturbation signature type, which can be derived from the database generated at 452. In some examples, the autoencoder model architecture includes an adaptively trained bottleneck architecture, which can include an initial bottleneck, and framework for adapting the bottleneck layer to data types during the training method 450 and ultimately providing a desired dimensionality for the bottleneck layer.
[0085] At 456, the method includes training the autoencoder model using the database (at 452). As mentioned, the training will vary depending on whether the architecture defined at 454 includes a static or adaptive bottleneck portion. As an example, the autoencoder neural network model can be trained at 456 using the python library PyTorch and the microbiome pharmacology database (provided at 452). The encoder portion includes an input layer having features, one for each drug similarity score entry of the database, a hidden layer neurons, and a bottleneck. The decoder portion maps bottleneck features to an input layer of the decoder portion, and decoder mapped the bottleneck to a hidden layer, and an output layer of drug similarity features commensurate with the input layer of the encoder portion. As an example, the training data from the database is applied to the input layer, encoded by the encoder portion and projected into the latent space, which is reconstructed by the decoder portion to provide reconstructed outputs at the output layer of the decoder portion.
[0086] As an example, FIG. 15 depicts an initial stage of an autoencoder model 500 during training with a certain type of training (e.g., for reconstructing drug training data). The model 500 includes an encoder portion 502 having an input layer 504 and one or more other layers, shown as hidden layer 506. Input data 508, corresponding to a training set from the database at 452, is provided to the input layer to enncode the input data and provide encoded data representing features of the input data. The training data 508 set can include the full set in the database or different subsets (e.g., segments of the database) can be added at different training stages in an order according to the type of pharmacologic factors. For example, respective data segments can be applied in certain order, such as of drugs and different microbiome factors (e.g., microbiota, metabolites, and the like) until all factors for a defined use case have been included. Other orders and segmenting of training data are possible. A bottleneck portion 510 generates feature vectors in the latent space representation based on the encoded data. A decoder portion 512 has an input layer 514 mapping to an output layer 516 for providing output data 518 as a reconstruction of the input data 508.
[0087] At 458, performance and error metrics of the autoencoder model are evaluated. The performance and error metrics can evaluate the performance and error based on the input data 508 and reconstructed data 518 using a variety of error and performance metrics. For example, as shown in FIG. 16, the reconstructed data 518 provided by the model 500 can be evaluated with respect to the input training data 508 during a series of training epochs by computing and evaluating error metrics over epochs, including loss, shown at 530, mean square error (MSE), shown at 532, R-squared error, shown at 534, and Pearson correlation, shown at 536. Other combinations of metrics may be used at 458in other examples for evaluating the model.
[0088] At 460, a determination is made whether the model performance is satisfactory based on the evaluation at 458. If the performance is determined (at 460) to be not satisfactory (NO), the method proceeds to 462. At 462, the method can include adjusting model parameters (e.g., weights and biases). In examples, where the model architecture includes an adaptive bottleneck, the method can also include tuning the model architecture and adjusting the adaptive bottleneck. From 462, the method can return to 456 for a next stage of training of the autoencoder model based on the updating and tuning of parameters and / or architecture at 462.
[0089] For example, FIG. 17 depicts an encoder model 550 during a next stage of training with reconstructing a second type of training data (e.g., for reconstructing microbiota training data) and responsive to the adjustments and tuning at 462. The model 550 includes an encoder portion 552 having an input layer and hidden layers, a bottleneck portion 554, and a decoderportion 556. Input data 508', corresponding to a microbiota segment of training data from the database at 452, is provided to the input layer to encode the input data and provide encoded data representing features of the input data. The bottleneck portion 554 projects the encoded data as feature vectors in a latent space representation (e.g., having a dimensionality according to the number of neurons in the bottleneck portion). The decoder portion 556 is operative to decode the features from the latent space to provide reconstructed data 560.
[0090] As a further example, as shown in FIG. 18, the reconstructed data 560 provided by the updated model 600 can be evaluated (at 458) relative to the input data 508’ by computing and evaluating error metrics over a number of epochs (e.g., about 145 epochs) in response to applying the training data from the database. The metrics can include loss, shown at 570, mean square error (MSE), shown at 572, R-squared error, shown at 574, and Pearson correlation, shown at 576. Other combinations of metrics may be used at 458 in other examples for evaluating the model.
[0091] At 460, a determination is again made whether performance (e.g., FIG. 18) of the model 550 is satisfactory based on the evaluation at 458. If the performance is determined (at 460) to be not satisfactory (NO), the method proceeds to 462, which can include adjusting model parameters (e.g., weights and biases). In some examples, the tuning and adjustments at 462 can also include tuning the model architecture and adjusting the adaptive bottleneck based on the evaluation. From 462, the method can return to 456 for performing a next stage of training of the autoencoder model based on the updating and tuning of parameters and / or architecture at 462.
[0092] For example, FIG. 19 depicts an encoder model 600 during a next stage of training with reconstructing a third type of training (e.g., for reconstructing metabolite data) and responsive to the preceding adjustments and tuning at 462, which can include deep layer adaptation and decoder layer adaptations. The model 600 includes an encoder portion 602 having an input layer and hidden layers, a bottleneck portion 604, and a decoder portion 606. Input data 508, corresponding to a training set from the database at 452, is provided to the input layer to encode the input data and provide encoded data representing features of the input data. The bottleneck portion 604 projects the encoded data as feature vectors in a latent space representation (e.g., having a dimensionality according to the number of neurons in the bottleneck portion). The decoder portion 606 decodes the features from the latent space to provide reconstructed data 610.
[0093] As a further example, as shown in FIG. 20, the reconstructed data provided by the updated model 600 can be evaluated (at 458) relative to the input training data 508” bycomputing and evaluating error metrics over a number of epochs (e.g., about 60 epochs) in response to applying the training data from the database. The metrics can include loss, shown at 620, mean square error (MSE), shown at 622, R-squared error, shown at 624, and Pearson correlation, shown at 626. Other combinations of metrics may be used at 458 in other examples for evaluating the model.
[0094] At 460, if the performance criteria evaluated at 458 is determined to be satisfactory (YES), the method proceeds to 464 and the trained autoencoder machine learning model is stored in memory. After training is completed, the training data or a select portion thereof can be applied to the training autoencoder machine learning model for further validation. FIGS. 21 and 22 are visualizations of latent space representations, shown at 640 and 650, that can be generated by a trained autoencoder neural network (e.g., and rendered on display 348). The visualization 640 demonstrates selected feature vectors for certain lactobacillus species, and the visualization 650 demonstrates selected feature vectors for certain facultative and obligate anaerobes, confirming the accuracy of the autocncodcr model.
[0095] As a further example, FIGS. 23, 24, and 25 depict visualizations 660, 662, and 664 for example latent space representations that can be generated by the autoencoder neural network trained by the method of FIG. 14 (e.g., and rendered on display 348). The visualizations 660, 662, and 664 demonstrate clusters of identified metabolites and microbiota near metronidazole, clindamycin, and clotrimazole, respectively, as residing within respective radius GUI elements.
[0096] FIG. 26 is a workflow diagram 700 illustrating creation of a database that can be used for training an autoencoder neural network of FIGS. 1, 7, and / or 12 and / or for generating input data that can be applied to a trained autoencoder neural network. In the example of FIG 26, the database is created in the context of vaginal microbiome-drug similarities. In other examples, the concepts are equally applicable to and can be extended to one or more other ingredients and targets, such as for one or more microbiomes.
[0097] As shown in FIG. 26, microbiome multi-omics data can be integrated with matched host transcriptomics via Spearman correlation analysis (or other statistical measure of relationships) to generate microbe-gene, metabolite-gene, and bacterial functions-gene lists. Additionally, microbiome-host gene lists are combined with gene expression or protein response signatures after treatment with a drug. The microbiome -host gene lists in the present example are characteristic direction signatures obtained from the library of integrated networkbased cellular signatures (LINCS) dataset), for example, via Spearman correlation with False Discovery Rate correction to identify vaginal microbiome-drug mimicry associations.
[0098] As a further example, in the context of vaginal microbiome, the analysis for creating the database can begin by extracting (1) drug-gene and (2) vaginal microbiome factor-gene signatures for comparison of candidate microbiome factor-drug mimicry relationships. To extract the drug-gene signatures, post-drug treatment perturbation gene expression data can be obtained, in one example, from the LINCS database at the level of the Characteristic Direction (CD) score for the “Chemical Perturbations” category for each druggene pair for metronidazole and clindamycin. CD profiles aggregate the post-treatment gene expression data across all cell lines, drug doses, and time points to obtain a consensus gene signature for the drug. As an alternative approach, a specific cell line may be used for generating drug signatures. It is to be understood that the database can be generated using drug perturbation gene expression / protein responses obtained from one or more other sources and / or using other methods to provide a list of genes and / or proteins, in which the expression or activity of a particular gene or protein increases or decreases (e.g., up / down-rcgulatcd) due to treatment with a drug, small molecule, or other factor that is desirable to incorporate into the training data and database. To extract the vaginal microbiome factor-gene signatures, vaginal microbiome multi-omics data can be obtained, for example, from a subset of HIV-negative participants enrolled in a clinical trial. The data for a plurality of participants can be analyzed, including vaginal microbiome composition, vaginal metabolomics, vaginal metaproteomics inferred via mass spectrometry analysis, and vaginal epithelial transcriptomics data obtained from vaginal biopsies. Spearman correlation analysis can be used to calculate a per-gene correlation coefficient in the host vaginal transcriptomics data for each microbe, metabolite, and bacterial function variable in the vaginal microbiome composition, metabolomics, or metaproteomics data. The result of such analysis can provide a correlation matrix for each microbiome data type quantifying how the abundance of a microbe, metabolite, or bacterial pathway activity is related to the expression of genes in vaginal epithelium, with the coefficients approximating a magnitude and directional change of the gene in response to the microbiome factors. Having extracted the sets of ding-gene relationships and vaginal microbiome factor-gene relationships, these relationships can be connected to infer similarities between the vaginal microbiome factors and drugs. For example, a matrix of potential vaginal microbiome factor-drug mimicry inferred associations can be generated by calculating a Spearman correlation coefficient for each vector of microbiome factor-gene coefficients and drug-CD gene coefficients from LINCS. At this stage, a Benjamini Hochberg FDR correction was applied to the microbiome-drug correlation coefficient p-values to identify statistically significant linkages for downstreamvisualization and analysis, such as shown in FIG. 26. The ingredients and / or targets (e.g., metabolites, taxa, bacterial functions, etc.) associated with respective host gene signatures, further indicating corresponding candidate mimicry associations. After filtering out missing values and technical dropouts, the resulting database can be generated and stored in memory and used, as described herein, such as for generating input data and training an autoencoder machine learning model.
[0099] Further information associated with the diagram 700, which can be used as a basis for generating the database, as disclosed herein, are set forth in Identifying a Vaginal Microbiome-Derived Selective Antibiotic Metabolite via Microbiome Pharmacology Analysis, Smrutiti Jena et al., available online at https: / / doi.org / 10.1101 / 2025.08.28.672927, which is incorporated herein by reference. Another example approach, in the context anti-cancer targets, which can be used to generate the database, which that can be used for training an autoencoder neural network and / or for generating input data that can be applied to a trained autocncodcr neural network, is disclosed in Computational microbiome pharmacology analysis elucidates the anti-cancer potential of vaginal microbes and metabolites, Front. Microbiol., 15 September 2025, Sec. Systems Microbiology, Volume 16 - 2025, which is available online at https: / / doi.org / 10.3389 / fmicb.2025.1602217, which is also incorporated herein by reference.EXAMPLE EMBODIMENTS:
[0100] Several aspects of the present technology are set forth in the following numbered examples.Example 1. A computer-implemented method comprising:receiving input data to an autocncodcr machine learning model to provide a latent space representation of feature vectors characterizing pharmacologic factors, in which the input data includes one or more data items representing one or more ingredients and / or one or more targets, and the autoencoder machine learning model is trained based on training data representing pharmacologic signatures for known ingredients and / or known targets;analyzing the latent space representation of the feature vectors; and identifying one or more candidate ingredients and / or one or more candidate targets based on the analysis.Example 2. The computer-implemented method of example 1, further comprising:receiving user data representing one or more ingredients and / or one or more targets;querying a database based on the user data; andgenerating, based on the querying, the input data to include a respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets represented in the user data.Example 3. The computer-implemented method of example 2, wherein the respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets includes a score for respective data entries, an input layer of the autoencoder machine learning model includes an input layer having neurons for each score, and each score is scaled according to a scale defined by respective entries of records in the database. Example 4. The computer-implemented method of example 2 or 3, further comprising generating combinations of the one or more ingredients received in the user data, the input data including a respective pharmacologic signature for each of the combinations.Example 5. The computer-implemented method of example 4, wherein each of the combinations defines a unique combination including the same or different combination of ingredients having respective concentrations thereof that are the same or different from each other.Example 6. The computer-implemented method of any ofexamples 2, 3, or 4, wherein the respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets represented in the user data represents pharmacological profile scores for each ingredient and / or target based on pharmacological profile scores stored in the database.Example 7. The computer-implemented method of any of examples 2, 3, 4, or 5, wherein the database comprises a microbiome pharmacology similarity matrix including pharmacologic signatures for predetermined targets and ingredients.Example s. The computer-implemented method of example 1, wherein the autoencoder machine learning model comprises an autoencoder neural network.Example 9. The computer-implemented method of example 8, wherein the autoencoder neural network comprises:an encoder portion, including an input layer and one or more other layers, in which the input layer has neurons for each score defined by the training data, and the encoder portion is trained to encode the input data and provide encoded data representing features of the input data;a bottleneck portion trained to generate the feature vectors in the latent space representation based on the encoded data; anda decoder portion, including an output layer and one or more other layers, in which the output layer has neurons for each score defined by the training data, and the decoder portion is trained to reconstruct the input data based on latent space representation.Example 10. The computer-implemented method of example 9, wherein the bottleneck portion includes an adaptively trained bottleneck architecture.Example 11. The computer-implemented method of example 1 , wherein analyzing the latent space representation comprises:determining a proximity of the feature vectors generated based on the input data relative to selected feature vectors, corresponding to at least one of the known ingredients and / or targets projected into the latent space representation; andevaluating the proximity to identify the one or more candidate ingredients and / or the one or more candidate targets.Example 12. The computer-implemented method of example 11, wherein determining the proximity comprises calculating a distance metric in the latent space representation.Example 13. The computer-implemented method of example 11 or 12, wherein analyzing the latent space representation comprises:selecting at least one feature vector or location in the latent space representation for the at least one of the known ingredients and / or targets;clustering feature vectors in the latent space representation within a distance threshold of the selected at least one feature vector or location; andgenerating a candidate list of the one or more candidate ingredients and / or the one or more candidate targets based on the clustered feature vectors.Example 14. The computer-implemented method of example 13, wherein the clustered feature vectors comprise a set of feature vectors generated in the latent space representation responsive to the input data and / or responsive to entries in the training data. Example 15. The computer-implemented method of any of examples 13, or 14, further comprising:generating updated input data based on at least some of the ingredients and / or targets in the candidate list;applying the updated input data to the autoencoder machine learning model to generate an updated set of feature vectors in the latent space representation; and evaluating the updated set of feature vectors to provide an updated version of the candidate list.Example 16. The computer-implemented method of example 15, wherein generating updated input data comprises:mapping at least some of the clustered feature vectors from the latent space representation into ingredients and / or targets in an input space representation, based on a database that includes the training data, to provide at least a portion of the updated input data.Example 17. The computer-implemented method of example 16, wherein generating updated input data comprises:expanding the ingredients and / or targets mapped into the input space representation to include simulated combinations of at least some of the candidate ingredients and / or the candidate targets; and / oradjusting a concentration of one or more of the ingredients mapped into the input space representation thereof.Example 18. The computer-implemented method of example 16, wherein the updated input data is generated automatically or in response to a user input.Example 19. The computer-implemented method of any of the preceding examples, further comprising:generating metadata for each of the feature vectors in the latent space representation to specify the ingredient or target; andmapping at least some of the feature vectors from the latent space to an input space of the input data.Example 20. The computer-implemented method of any of examples 2-7 or 16, wherein the database comprises a patient- specific microbiome profile for an individual such that respective locations of the feature vectors in the latent space representation provides a pharmacologic assessment of the ingredients and / or targets in the input data for the individual.Example 21. The computer-implemented method of any of examples 2-7 or 16, wherein the database comprises a microbiome-drug similarity matrix, defining correlations between drugs and microbiome profiles in drug space, and the latent space representation characterizes a predicted correlation for the ingredients in the drug space.Example 22. The computer-implemented method of any of examples 1-19, wherein the input data includes ingredients representing at least one simulated synthetic community of ingredients, and the latent space representation characterizes a predicted correlation for each of the at least one simulated synthetic community of ingredients.Example 23. The computer-implemented method of any of the preceding examples, further comprising generating the training data, as a similarity matrix, from microbiome based signatures of individual and combinations of microbiome factors and signatures of drug treatment from a cohort of subjects via comparison of a quantitative similarity of the signatures of the microbiome factors and drug signatures.Example 24. The computer-implemented method of example 23, further comprising predicting pharmacological actions or profile of one or more of the microbiome factors or the combinations of the microbiome factors based on spatial proximity of the microbiome factors and drags projected by the autoencoder machine learning model in the latent space representation.Example 25. The computer-implemented method of any of the preceding examples, further comprising providing output data to a user computing device, in which the output data includes the one or more candidate targets and / or candidate ingredients and / or a graphical representation of the latent space representation.Example 26 The computer-implemented method of example 25, further comprising generating a visualization at the user computing device based on the output data.Example 27. The computer-implemented method of any of the preceding examples, further comprising:applying at least a portion of the training data into the trained autoencoder machine learning model such that respective features of the portion of the training data and the input data are encoded and mapped, as the feature vectors, in the latent space representation. Example 28. The computer-implemented method of any of the preceding examples, wherein the one or more ingredients comprise one or more therapeutic agents, small molecules, biologies, cell, proteins, peptides, nucleic acids, genes or gene products / fragments, microbiome factors, metabolites, other model-designed products, or any combination thereof.Example 29. The computer-implemented method of any of the preceding examples, wherein the one or more targets comprise one or more specific cellular receptors, pathways, proteins, other biological processes or factors to position the ingredients or other model-designed product toward, or any combination thereof.Example 30. A therapeutic agent produced based on a formulation determined according to the computer-implemented method of any one or more of the preceding examples.Example 31. One or more targets determined for one or more simulated communities of ingredients, defining the input data, according to the computer-implemented method of any of examples 1-28.Example 32. A system, comprising:one or more processors; andone or more non-transitory computer-readable media having instructions that, when executed by the one or more processors, cause the one or more processors to:apply input data to an autoencoder machine learning model to generate a latent space representation of feature vectors that encode pharmacologic factors based on the input data, in which the input data includes one or more data items representing one or more ingredients and / or one or more targets, and the autoencoder machine learning model is trained based on training data representing pharmacologic signatures for known ingredients and / or known targets;analyze the latent space representation of the feature vectors; andidentify one or more candidate ingredients and / or one or more candidate targets based on the analysis.Example 33. The system of example 32, wherein the instructions further cause the one or more processors to:receive user data representing one or more ingredients and / or one or more targets; query a database in response to user data; andgenerating, based on the query, the input data to include a respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets represented in the user data.Example 34. The system of example 33, wherein the respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets includes a score for respective data entries, an input layer of the autoencoder machine learning model includes an input layer having neurons for each score, and each score is scaled according to a scale defined by respective entries of records in the database.Example 35. The system of example 33 or 34, wherein the instructions further cause the one or more processors to:generate combinations of the one or more ingredients and / or one or more targets received in the user data;add, to the input data, a respective combination data record for each of the combinations that is generated, in which each respective combination data record has a respective pharmacologic signature.Example 36. The system of example 35, wherein each respective combination data record in the input data defines a unique combination of ingredients and / or targets including the same or different combination of ingredients and / or targets, and each unique combination has a respective concentration for of ingredients and / or targets thereof that is the same or different from each other combination.Example 37. The system of any of examples 33, 34, 35, or 36, wherein the respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets represented in the input data set represents pharmacological profile scores for each ingredient and / or target based on pharmacological profile scores stored in the database. Example 38. The system of any of examples 33, 34, 35, or 36, wherein the database comprises a microbiomc pharmacology similarity matrix including pharmacologic signatures for predetermined targets and ingredients.Example 39. The system of any of examples 32, 33, 34, 35, 36, 37, or 38, wherein the autoencoder machine learning model comprises an autoencoder neural network.Example 40. The system of example 39, wherein the autoencoder neural network comprises:an encoder portion, including an input layer and one or more other layers, in which the input layer has neurons for each score defined by the training data, and the encoder portion is trained to encode the input data and provide encoded data representing features of the input data;a bottleneck portion trained to generate the feature vectors in the latent space representation based on the encoded data; anda decoder portion, including an output layer and one or more other layers, in which the output layer has neurons for each score defined by the training data, and the decoder portion is trained to reconstruct the input data based on latent space representation.Example 41. The system of example 40, wherein the bottleneck portion includes an adaptively trained bottleneck architecture.Example 42. The system of any of examples 32 through 41, wherein the instructions to analyze the latent space representation further cause the one or more processors to:determine a proximity of the feature vectors generated based on the input data relative to selected feature vectors, corresponding to at least one of the known ingredients and / or targets projected into the latent space representation; andevaluate the proximity to identify the one or more candidate ingredients and / or the one or more candidate targets.Example 43. The system of example 42, wherein the proximity of the feature vectors is determined by calculating a distance metric in the latent space representation.Example 44. The system of example 42, wherein the instructions to analyze the latent space representation further cause the one or more processors to:select at least one feature vector or location in the latent space representation for the at least one of the known ingredients and / or targets;cluster feature vectors in the latent space representation within a distance threshold of the selected at least one feature vector or location; andgenerate a candidate list of the one or more candidate ingredients and / or the one or more candidate targets based on the clustered feature vectors.Example 45. The system of example 44, wherein the clustered feature vectors comprise a set of feature vectors generated in the latent space representation responsive to the input data and / or responsive to entries in the training data.Example 46. The system of example 44 or 45, wherein the instructions further cause the one or more processors to:generate updated input data based on at least some of the ingredients and / or targets in the candidate list;apply the updated input data to the autoencoder machine learning model to generate an updated set of feature vectors in the latent space representation; andevaluate the updated set of feature vectors to provide an updated version of the candidate list.Example 47. The system of example 46, wherein the instructions to generate the updated input data further cause the one or more processors to:mapping at least some of the clustered feature vectors from the latent space representation into ingredients and / or targets in an input space representation, based on a database that includes the training data, to provide at least a portion of the updated input data.Example 48. The system of example 47, wherein the instructions to generate the updated input data further cause the one or more processors to:expand the ingredients and / or targets mapped into the input space representation to include simulated combinations of at least some of the candidate ingredients and / or the candidate targets; and / oradjust a concentration of one or more of the ingredients mapped into the input space representation thereof.Example 49. The system of example 47, wherein the updated input data is generated automatically or in response to a user input.Example 50. The system according to any of examples 32 through 49, wherein the instructions further cause the one or more processors to:generate metadata for each of the feature vectors in the latent space representation to specify the ingredient or target; andmap at least some of the feature vectors from the latent space to an input space of the input data.Example 51. The system according to any of examples 33 through 38 or 47, wherein the database comprises a patient-specific microbiome profile for an individual such that respective locations of the feature vectors in the latent space representation provides a pharmacologic assessment of the ingredients and / or targets in the input data for the individual.Example 52. The system according to any of examples 33 through 38 or 47, wherein the database comprises a microbiome-drug similarity matrix, defining correlations between drugs and microbiome profiles in drug space, and the latent space representation characterizes a predicted correlation for the ingredients in the drag space.53. The system according to any of examples 32 through 49, wherein the input data includes ingredients representing at least one simulated synthetic community of ingredients, and the latent space representation characterizes a predicted correlation for each of the at least one simulated synthetic community of ingredients.Example 54. The system according to any of examples 32 through 53, wherein the instructions further cause the one or more processors to:generate the training data, as a similarity matrix, from microbiome based signatures of individual and combinations of microbiome factors and signatures of drag treatment from a cohort of subjects via comparison of a quantitative similarity of the signatures of the microbiome factors and drag signatures.Example 55. The system of example 54, wherein the instructions further cause the one or more processors to:predict pharmacological actions or profile of one or more of the microbiome factors or the combinations of the microbiome factors based on spatial proximity of the microbiome factors and drugs projected by the autoencoder machine learning model in the latent space representation.Example 56. The system according to any of the examples 32 through 55, wherein the one or more processors are one or more first processors, the one or more non-transitory computer-readable media are one or more first non-transitory computer-readable media, and the system further comprises:a first computing device that includes at least one of the one or more first processors and the one or more first non-transitory computer-readable media; anda second computing device in communication with the first computing device, wherein the first computing device provides output data to the second computing device, in which the output data includes the one or more candidate targets and / or candidate ingredients and / or a graphical representation of the latent space representation Example 57. The system of example 56, wherein the second computing device comprising:one or more second processors; andone or more second non-transitory computer-readable media including instructions that, when executed by the one or more processors of the second computing device, cause the one or more processors to generate a visualization at the second computing device based on the output data.Example 58. The system of example 56 or 57, wherein the instructions of the one or more second non-transitory computer-readable media further comprise an application programming interface operative to use at least some of the instructions of the one or more first non-transitory computer-readable media.Example 59. The system according to any of the examples 32 through 55, wherein the instructions further cause the one or more processors to:apply at least a portion of the training data into the trained autoencoder machine learning model such that respective features of the portion of the training data and the input data are encoded collectively, as corresponding feature vectors, in the latent space representation.Example 60. The system according to any of the examples 32 through 59, wherein the one or more ingredients comprise one or more therapeutic agents, small molecules, biologies, cell, proteins, peptides, nucleic acids, genes or gene products / fragments.microbiome factors, metabolites, other model-designed products, or any combination thereof.Example 61. The system according to any of the examples 32 through 60, wherein the one or more targets comprise one or more specific cellular receptors, pathways, proteins, other biological processes or factors to position the ingredients or other model-designed product toward, or any combination thereof.Example 62. A computer-implemented method, comprising:generating a training dataset from microbiome-based signatures of individual and combinations of microbiome factors and signatures of drag treatment from a cohort of subjects based on comparison of a quantitative similarity of the signatures of the microbiome factors and drag signatures.Example 63. The computer-implemented method of example 62, further comprising deriving, from the training data set, pharmacological actions of microbiome members, microbiomc products, and combinations of the microbiomc members and products.Example 64. The computer-implemented of example 62 or 63, wherein the dataset includes obtained and inputted drag-response gene and protein expression signatures and microbiome-to-host gene or protein signatures of the cohort of subjects, in which each of the signatures includes scores for respective data entries in the training data.Example 65. The computer-implemented method of any of examples 62 through 64, wherein the microbiome factors include microbiome members and / or microbiome products.Example 66. The computer-implemented method of any of examples 62 through 65, wherein the quantitative similarity comprises a functional similarity based on a statistical metric.Example 67. The computer-implemented method of any of examples 62 through 66, further comprising training an autoencoder machine learning model using the training dataset.Example 68. The computer-implemented method of any of examples 62 through 67, further comprising training a neural network using the training dataset to determine combinations of microbiome factors that target biological mechanisms of the human body.Example 69. A computer-implemented method of using an autoencoder machine learning model to determine at least one formulation of ingredients to achieve a desiredbiological and / or pharmacological effect or action in response to input data representing one or more ingredients and / or one or more targets.Example 70. A therapeutic agent produced according to the at least one formulation of ingredients determined according to the method of example 69.
[0101] It should be understood that various aspects described herein may be combined in different combinations than the combinations specifically presented in the description and accompanying drawings. It should also be understood that, depending on the example, certain acts or events of any of the processes or methods described herein may be performed in a different sequence, may be added, merged, or left out altogether (e.g., all described acts or events may not be necessary to carry out the techniques). In addition, while certain aspects of this description are described as being performed by a single module or unit for purposes of clarity, it should be understood that the techniques of this description may be performed by a combination of units or modules.
[0102] In one or more examples, the described techniques may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media may include non-transitory computer-readable media, which corresponds to a tangible medium such as data storage media (e.g., RAM, ROM, EEPROM, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a processor). For example, instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term “processor” as used herein may refer to any of the foregoing structure(s) or any other physical structure suitable for implementation of the described techniques. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0103] In this description, numerical designations “first,” “second,” etc. are not necessarily consistent with same designations in the claims herein and these numerical designations are used to simply distinguish one element from another. Also, the term “based on” means based at least in part on.
[0104] Additionally, the term "couple" or variants thereof may cover connections, communications, or signal paths that enable a functional relationship consistent with this description. For example, if device A generates a signal to control device B to perform anaction, then: (a) in a first example, device A is directly coupled to device B; or (b) in a second example, device A is indirectly coupled to device B through intervening component C if intervening component C does not alter the functional relationship between device A and device B, so device B is controlled by device A via the control signal generated by device A.
[0105] In this description, the term “based on” means based at least in part on. Also, as used herein, the term “includes” means includes but not limited to, and the term “including” means including but not limited to.
[0106] Also, in this description, a device that is “configured to” perform a task or function may be configured (e.g., programmed and / or hardwired) at a time of manufacturing by a manufacturer to perform the function and / or may be configurable (or reconfigurable) by a user after manufacturing to perform the function and / or other additional or alternative functions. The configuring may be through firmware and / or software programming of the device, through a construction and / or layout of hardware components and interconnections of the device, or a combination thereof.
[0107] In this description, unless otherwise stated, “about,” “approximately” or “substantially” preceding a parameter means being within + / - 10 percent of that parameter. Modifications are possible in the described embodiments and other embodiments are possible within the scope of the claims.
[0108] What have been described above are examples. It is, of course, not possible to describe every conceivable combination of components or methods, but one of ordinary skill in the art will recognize that many further combinations and permutations are possible. Accordingly, the invention is intended to embrace all such alterations, modifications, and variations that fall within the scope of this application, including the appended claims. Where the description or claims recite “a,” “an,” “a first,” or “another” element, or the equivalent thereof, it should be interpreted to include one or more than one such element, neither requiring nor excluding two or more such elements.
[0109] Furthermore, a circuit or device that is said to include certain components may instead be configured to couple to those components to form the described circuitry, device, or system. For example, a structure described as including one or more elements A, B and C may instead include only the A elements within a single physical device and may be configured to couple to at least some of the elements B and / or C to form the described circuitry, device, or system, either at a time of manufacture or after a time of manufacture, for example, by an enduser and / or a third-party.
[0110] All references, publications, and patents cited in the present application are herein incorporated by reference in their entirety.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method comprising :receiving input data to an autoencoder machine learning model to provide a latent space representation of feature vectors characterizing pharmacologic factors, in which the input data includes one or more data items representing one or more ingredients and / or one or more targets, and the autoencoder machine learning model is trained based on training data representing pharmacologic signatures for known ingredients and / or known targets;analyzing the latent space representation of the feature vectors; and identifying one or more candidate ingredients and / or one or more candidate targets based on the analysis.
2. The computer-implemented method of claim 1, further comprising:receiving user data representing one or more ingredients and / or one or more targets;querying a database based on the user data; andgenerating, based on the querying, the input data to include a respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets represented in the user data.
3. The computer-implemented method of claim 2, wherein the respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets includes a score for respective data entries, an input layer of the autoencoder machine learning model includes an input layer having neurons for each score, and each score is scaled according to a scale defined by respective entries of records in the database.
4. The computer-implemented method of claim 2 or 3 , further comprising generating combinations of the one or more ingredients received in the user data, the input data including a respective pharmacologic signature for each of the combinations.
5. The computer-implemented method of claim 4, wherein each of the combinations defines a unique combination including the same or different combination of ingredients having respective concentrations thereof that are the same or different from each other.
6. The computer-implemented method of any one of claims 2, 3, or 4, wherein the respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets represented in the user data represents pharmacological profile scores for each ingredient and / or target based on pharmacological profile scores stored in the database.
7. The computer-implemented method of any one of claims 2, 3, 4, or 5, wherein the database comprises a microbiome pharmacology similarity matrix including pharmacologic signatures for predetermined targets and ingredients.
8. The computer-implemented method of claim 1, wherein the autoencoder machine learning model comprises an autoencoder neural network.
9. The computer-implemented method of claim 8, wherein the autoencoder neural network comprises:an encoder portion, including an input layer and one or more other layers, in which the input layer has neurons for each score defined by the training data, and the encoder portion is trained to encode the input data and provide encoded data representing features of the input data;a bottleneck portion trained to generate the feature vectors in the latent space representation based on the encoded data; anda decoder portion, including an output layer and one or more other layers, in which the output layer has neurons for each score defined by the training data, and the decoder portion is trained to reconstruct the input data based on latent space representation.
10. The computer-implemented method of claim 9, wherein the bottleneck portion includes an adaptively trained bottleneck architecture.
11. The computer-implemented method of claim 1 , wherein analyzing the latent space representation comprises:determining a proximity of the feature vectors generated based on the input data relative to selected feature vectors, corresponding to at least one of the known ingredients and / or targets projected into the latent space representation; andevaluating the proximity to identify the one or more candidate ingredients and / or the one or more candidate targets.
12. The computer-implemented method of claim 11, wherein determining the proximity comprises calculating a distance metric in the latent space representation.
13. The computer-implemented method of claim 11 or 12, wherein analyzing the latent space representation comprises:selecting at least one feature vector or location in the latent space representation for the at least one of the known ingredients and / or targets;clustering feature vectors in the latent space representation within a distance threshold of the selected at least one feature vector or location; andgenerating a candidate list of the one or more candidate ingredients and / or the one or more candidate targets based on the clustered feature vectors.
14. The computer-implemented method of claim 13, wherein the clustered feature vectors comprise a set of feature vectors generated in the latent space representation responsive to the input data and / or responsive to entries in the training data.
15. The computer-implemented method of any one of claims 13, or 14, further comprising:generating updated input data based on at least some of the ingredients and / or targets in the candidate list;applying the updated input data to the autoencoder machine learning model to generate an updated set of feature vectors in the latent space representation; andevaluating the updated set of feature vectors to provide an updated version of the candidate list.
16. The computer-implemented method of claim 15, wherein generating updated input data comprises:mapping at least some of the clustered feature vectors from the latent space representation into ingredients and / or targets in an input space representation, based on a database that includes the training data, to provide at least a portion of the updated input data.
17. The computer-implemented method of claim 16, wherein generating updated input data comprises:expanding the ingredients and / or targets mapped into the input space representation to include simulated combinations of at least some of the candidate ingredients and / or the candidate targets; and / oradjusting a concentration of one or more of the ingredients mapped into the input space representation thereof.
18. The computer-implemented method of claim 16, wherein the updated input data is generated automatically or in response to a user input.
19. The computer-implemented method of any one of the preceding claims, further comprising:generating metadata for each of the feature vectors in the latent space representation to specify the ingredient or target; andmapping at least some of the feature vectors from the latent space to an input space of the input data.
20. The computer-implemented method of any one of claims 2-7 or 16, wherein the database comprises a patient- specific microbiome profile for an individual such that respective locations of the feature vectors in the latent space representation provides a pharmacologic assessment of the ingredients and / or targets in the input data for the individual.
21. The computer-implemented method of any one of claims 2-7 or 16, wherein the database comprises a microbiome-drug similarity matrix, defining correlations between drugs and microbiome profiles in drug space, and the latent space representation characterizes a predicted correlation for the ingredients in the drag space.
22. The computer-implemented method of any one of claims 1-19, wherein the input data includes ingredients representing at least one simulated synthetic community of ingredients, and the latent space representation characterizes a predicted correlation for each of the at least one simulated synthetic community of ingredients.
23. The computer-implemented method of any one of the preceding claims, further comprising generating the training data, as a similarity matrix, from microbiome based signatures of individual and combinations of microbiome factors and signatures of drug treatment from a cohort of subjects via comparison of a quantitative similarity of the signatures of the microbiome factors and drug signatures.
24. The computer-implemented method of claim 23, further comprising predicting pharmacological actions or profile of one or more of the microbiome factors or the combinations of the microbiome factors based on spatial proximity of the microbiome factors and drugs projected by the autoencoder machine learning model in the latent space representation.
25. The computer-implemented method of any one of the preceding claims, further comprising providing output data to a user computing device, in which the output data includes the one or more candidate targets and / or candidate ingredients and / or a graphical representation of the latent space representation.26 The computer-implemented method of claim 25, further comprising generating a visualization at the user computing device based on the output data.
27. The computer-implemented method of any one of the preceding claims, further comprising:applying at least a portion of the training data into the trained autoencoder machine learning model such that respective features of the portion of the training data and the input data are encoded and mapped, as the feature vectors, in the latent space representation.
28. The computer-implemented method of any one of the preceding claims, wherein the one or more ingredients comprise one or more therapeutic agents, small molecules, biologies, cell, proteins, peptides, nucleic acids, genes or gene products / fragments, microbiome factors, metabolites, other model-designed products, or any combination thereof.
29. The computer-implemented method of any one of the preceding claims, wherein the one or more targets comprise one or more specific cellular receptors, pathways, proteins, other biological processes or factors to position the ingredients or other model-designed product toward, or any combination thereof.
30. A therapeutic agent produced based on a formulation determined according to the computer-implemented method of any one or more of the preceding claims.
31. One or more targets determined for one or more simulated communities of ingredients, defining the input data, according to the computer-implemented method of any one of claims 1-28.
32. A system, comprising:one or more processors; andone or more non-transitory computer-readable media having instructions that, when executed by the one or more processors, cause the one or more processors to:apply input data to an autoencoder machine learning model to generate a latent space representation of feature vectors that encode pharmacologic factors based on the input data, in which the input data includes one or more data items representing one or more ingredients and / or one or more targets, and the autoencoder machine learning model is trained based on training data representing pharmacologic signatures for known ingredients and / or known targets;analyze the latent space representation of the feature vectors; andidentify one or more candidate ingredients and / or one or more candidate targets based on the analysis.
33. The system of claim 32, wherein the instructions further cause the one or more processors to:receive user data representing one or more ingredients and / or one or more targets; query a database in response to user data; andgenerating, based on the query, the input data to include a respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets represented in the user data.
34. The system of claim 33, wherein the respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets includes a score for respective data entries, an input layer of the autoencoder machine learning model includes an input layer having neurons for each score, and each score is scaled according to a scale defined by respective entries of records in the database.
35. The system of claim 33 or 34, wherein the instructions further cause the one or more processors to:generate combinations of the one or more ingredients and / or one or more targets received in the user data;add, to the input data, a respective combination data record for each of the combinations that is generated, in which each respective combination data record has a respective pharmacologic signature.
36. The system of claim 35, wherein each respective combination data record in the input data defines a unique combination of ingredients and / or targets including the same or different combination of ingredients and / or targets, and each unique combination has a respective concentration for of ingredients and / or targets thereof that is the same or different from each other combination.
37. The system of any one of claims 33, 34, 35, or 36, wherein the respective pharmacologic signature for each of the one or more ingredients and / or the one or more targets represented in the input data set represents pharmacological profile scores for each ingredient and / or target based on pharmacological profile scores stored in the database.
38. The system of any one of claims 33, 34, 35, or 36, wherein the database comprises a microbiome pharmacology similarity matrix including pharmacologic signatures for predetermined targets and ingredients.
39. The system of any one of claims 32, 33, 34, 35, 36, 37, or 38, wherein the autoencoder machine learning model comprises an autoencoder neural network.
40. The system of claim 39, wherein the autoencoder neural network comprises: an encoder portion, including an input layer and one or more other layers, in which the input layer has neurons for each score defined by the training data, and the encoder portion is trained to encode the input data and provide encoded data representing features of the input data;a bottleneck portion trained to generate the feature vectors in the latent space representation based on the encoded data; anda decoder portion, including an output layer and one or more other layers, in which the output layer has neurons for each score defined by the training data, and the decoder portion is trained to reconstruct the input data based on latent space representation.
41. The system of claim 40, wherein the bottleneck portion includes an adaptively trained bottleneck architecture.
42. The system of any one of claims 32 through 41, wherein the instructions to analyze the latent space representation further cause the one or more processors to:determine a proximity of the feature vectors generated based on the input data relative to selected feature vectors, corresponding to at least one of the known ingredients and / or targets projected into the latent space representation; andevaluate the proximity to identify the one or more candidate ingredients and / or the one or more candidate targets.
43. The system of claim 42, wherein the proximity of the feature vectors is determined by calculating a distance metric in the latent space representation.
44. The system of claim 42, wherein the instructions to analyze the latent space representation further cause the one or more processors to:select at least one feature vector or location in the latent space representation for the at least one of the known ingredients and / or targets;cluster feature vectors in the latent space representation within a distance threshold of the selected at least one feature vector or location; andgenerate a candidate list of the one or more candidate ingredients and / or the one or more candidate targets based on the clustered feature vectors.
45. The system of claim 44, wherein the clustered feature vectors comprise a set of feature vectors generated in the latent space representation responsive to the input data and / or responsive to entries in the training data.
46. The system of claim 44 or 45, wherein the instructions further cause the one or more processors to:generate updated input data based on at least some of the ingredients and / or targets in the candidate list;apply the updated input data to the autoencoder machine learning model to generate an updated set of feature vectors in the latent space representation; andevaluate the updated set of feature vectors to provide an updated version of the candidate list.
47. The system of claim 46, wherein the instructions to generate the updated input data further cause the one or more processors to:mapping at least some of the clustered feature vectors from the latent space representation into ingredients and / or targets in an input space representation, based on a database that includes the training data, to provide at least a portion of the updated input data.
48. The system of claim 47, wherein the instructions to generate the updated input data further cause the one or more processors to:expand the ingredients and / or targets mapped into the input space representation to include simulated combinations of at least some of the candidate ingredients and / or the candidate targets; and / oradjust a concentration of one or more of the ingredients mapped into the input space representation thereof.
49. The system of claim 47, wherein the updated input data is generated automatically or in response to a user input.
50. The system according to any one of claims 32 through 49, wherein the instructions further cause the one or more processors to:generate metadata for each of the feature vectors in the latent space representation to specify the ingredient or target; andmap at least some of the feature vectors from the latent space to an input space of the input data.
51. The system according to any one of claims 33 through 38 or 47, wherein the database comprises a patient-specific microbiome profile for an individual such that respective locations of the feature vectors in the latent space representation provides a pharmacologic assessment of the ingredients and / or targets in the input data for the individual.
52. The system according to any one of claims 33 through 38 or 47, wherein the database comprises a microbiome-drug similarity matrix, defining correlations between drugs and microbiome profiles in drug space, and the latent space representation characterizes a predicted correlation for the ingredients in the drug space.
53. The system according to any one of claims 32 through 49, wherein the input data includes ingredients representing at least one simulated synthetic community of ingredients, and the latent space representation characterizes a predicted correlation for each of the at least one simulated synthetic community of ingredients.
54. The system according to any one of claims 32 through 53, wherein the instructions further cause the one or more processors to:generate the training data, as a similarity matrix, from microbiome based signatures of individual and combinations of microbiome factors and signatures of drug treatment from a cohort of subjects via comparison of a quantitative similarity of the signatures of the microbiome factors and drug signatures.
55. The system of claim 54, wherein the instructions further cause the one or more processors to:predict pharmacological actions or profile of one or more of the microbiome factors or the combinations of the microbiome factors based on spatial proximity of the microbiome factors and drugs projected by the autoencoder machine learning model in the latent space representation.
56. The system according to any one of the claims 32 through 55, wherein the one or more processors are one or more first processors, the one or more non-transitory computer-readable media are one or more first non-transitory computer-readable media, and the system further comprises:a first computing device that includes at least one of the one or more first processors and the one or more first non-transitory computer-readable media; anda second computing device in communication with the first computing device, wherein the first computing device provides output data to the second computing device, in which the output data includes the one or more candidate targets and / or candidate ingredients and / or a graphical representation of the latent space representation57. The system of claim 56, wherein the second computing device comprising:one or more second processors; andone or more second non-transitory computer-readable media including instructions that, when executed by the one or more processors of the second computing device, cause the one or more processors to generate a visualization at the second computing device based on the output data.
58. The system of claim 56 or 57, wherein the instructions of the one or more second non-transitory computer-readable media further comprise an application programming interface operative to use at least some of the instructions of the one or more first non-transitory computer-readable media.
59. The system according to any one of the claims 32 through 55, wherein the instructions further cause the one or more processors to:apply at least a portion of the training data into the trained autoencoder machine learning model such that respective features of the portion of the training data and the input data are encoded collectively, as corresponding feature vectors, in the latent space representation.
60. The system according to any one of the claims 32 through 59, wherein the one or more ingredients comprise one or more therapeutic agents, small molecules, biologies, cell, proteins, peptides, nucleic acids, genes or gene products / fragments, microbiome factors, metabolites, other model-designed products, or any combination thereof.
61. The system according to any one of the claims 32 through 60, wherein the one or more targets comprise one or more specific cellular receptors, pathways, proteins, other biological processes or factors to position the ingredients or other model-designed product toward, or any combination thereof.
62. A computer-implemented method, comprising:generating a training dataset from microbiome-based signatures of individual and combinations of microbiome factors and signatures of drug treatment from a cohort of subjects based on comparison of a quantitative similarity of the signatures of the microbiome factors and drug signatures.
63. The computer-implemented method of claim 62, further comprising deriving, from the training data set, pharmacological actions of microbiome members, microbiome products, and combinations of the microbiomc members and products.
64. The computer-implemented of claim 62 or 63, wherein the dataset includes obtained and inputted drug-response gene and protein expression signatures and microbiome-to-host gene or protein signatures of the cohort of subjects, in which each of the signatures includes scores for respective data entries in the training data.
65. The computer-implemented method of any one of claims 62 through 64, wherein the microbiome factors include microbiome members and / or microbiome products.
66. The computer-implemented method of any of claims 62 through 65, wherein the quantitative similarity comprises a functional similarity based on a statistical metric.
67. The computer-implemented method of any of claims 62 through 66, further comprising training an autoencoder machine learning model using the training dataset.
68. The computer-implemented method of any of claims 62 through 67, further comprising training a neural network using the training dataset to determine combinations of microbiome factors that target biological mechanisms of the human body.
69. A computer-implemented method of using an autoencoder machine learning model to determine at least one formulation of ingredients to achieve a desired biological and / or pharmacological effect or action in response to input data representing one or more ingredients and / or one or more targets.
70. A therapeutic agent produced according to the at least one formulation of ingredients determined according to the method of claim 69.