Systems and methods for predicting biological activity of chemical / biological agents
A two-step process using contrastive learning on multimodal datasets trains a predictive model to accurately predict the biological impact of chemical agents on targets, addressing inefficiencies in drug discovery by enhancing computational processing and accelerating the development of safe medicines.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-03-26
AI Technical Summary
Current methods lack the ability to effectively leverage prior data presentations to predict the biological activity of novel chemical/biological agents, leading to inefficiencies in the drug discovery process regarding time, cost, and success rate.
A two-step process involving a foundation model trained using contrastive learning on multimodal datasets to learn abstract representations of chemical agents, followed by a predictive model that predicts the biological impact of these agents on targets like biological pathways, transcription factors, or genes, without the need for empirical data collection.
Enhances the efficiency and accuracy of drug discovery by enabling predictive modeling of agent modulation on biological targets, improving computational processing of diverse datasets and accelerating the development of efficacious and safe medicines.
Smart Images

Figure US2025047367_26032026_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR PREDICTING BIOLOGICAL ACTIVITY OF CHEMICAL / BIOLOGICAL AGENTSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 698,025, filed on September 23, 2024, the entire contents of which are incorporated by reference herein.TECHNICAL FIELD
[0002] The subject matter of this disclosure relates to systems and methods for predicting the biological activity of chemical / biological agents. In certain examples, the process of selecting, curating and processing multiple data modalities to train a machine learning model to perform multi-task multi-class classification of an agent is described.BACKGROUND
[0003] Drug discovery efforts often rely on multiple empirical data modalities to determine the clinical viability and safety profile of experimental therapeutics intended for medical use. For example, the process of developing a therapeutic for a given medical indication can involve identification of a biological target, generation of chemical / biological agents able to modulate the biological target, and biological assessment through various assays to determine the biological target engagement and safety profile (e.g., off-target activity) of the agents.
[0004] While this process generates multimodal, multidimensional datasets for specific agents, current methods are lacking in their ability to leverage prior data presentations to inform predictors of the biological activity of novel agents for therapeutic use. Furthermore, there is an unmet need to improve the efficiency of drug discovery process such that there is an improvement (e.g., length of time, cost, and success rate) in the discovery of efficacious and safe medicines.SUMMARY
[0005] In various examples, the subject matter of this disclosure relates to developing and implementing a predictive model to predict the biological impact of chemical agents by applying contrastive learning to a multimodal dataset. Upon application of an agent (e.g., an agent structure) to the predictive model, the predictive model predicts modulation (e.g., activity or expression) of a biological target, such as a biological pathway, a transcriptionfactor or a gene. In one embodiment, the predictive model predicts the agent modulation of the biological target as comprising stimulation, inhibition, or non-modulation (e.g., no change) of the biological target. In another embodiment, the predictive model predicts the agent modulation of the biological target as yes / no (binary scoring system) or a probability score.
[0006] In an embodiment of the invention, a two-step process for developing and training machine learning models, with the resultant predictive model predicting the biological impact of chemical agents. In some embodiments, step one comprises a foundation model. The foundation model is trained using multiple dataset types (e.g., distinct dataset modalities) to learn abstract representations of agents that are informed by the multiple dataset types. The foundation model is suited to processing a variety of dataset types (e.g., multimodal) for heterogeneous collections of compounds. Through a novel algorithm that innovates on contrastive learning methods, which allow for feature mapping of seemingly disparate data modalities, the output of the foundation model is a learned representation of chemical agents, e.g., the chemical structure dataset of agents. The mathematical representation of chemical structures learned by the foundational model is useful for tasks of interest, including multiclass classification.
[0007] In some embodiments, step two comprises a predictive model, derived from the foundation model. The predictive model characterizes and predicts the biological impact of an agent. The predictive model, can, in some instances, predicts the biological impact(s) of a chemical / molecular structure prior to and without the need for collection of empirical data. In an aspect of the invention, upon application of an agent (e.g., an agent structure) to the predictive model, the predictive model predicts modulation (e.g., activity or expression) of a biological target, such as a biological pathway, a transcription factor or a gene. In one embodiment, the predictive model predicts the agent modulation of the biological target as comprising stimulation, inhibition, or non-modulation (e.g., no change) of the biological target. In another embodiment, the predictive model predicts the agent modulation of the biological target as yes / no (binary scoring system) or a probability score.
[0008] In one aspect, the method includes the steps of generating a foundation model and a predictive model based on one or more biological or chemical datasets for the purposes of multi-task multi-class classification of biological target modulation. The process comprises training the foundation model, using an encoder function, by applying contrastive learning to a variety of datasets (e.g., a multimodal dataset), resulting in representational feature maps ofrelationships among the datasets, training a predictive machine learning model using the representational feature map learned by the foundation model as an input, wherein the predictive machine learning model predicts modulatory or non-modulatory effects of an agent on one or more biological targets. In another aspect, applying contrastive learning to the multimodal dataset comprises embedding each dataset modality in a shared representation space and creating a learned representation of each dataset modality, wherein training the predictive model based on the learned representation of the chemical structure dataset modality and a classifier function that assigns modulation status for the biological target results in the trained predictive model. In other instances, an agent structure is presented to the predictive model to predict biological activity of the agent.
[0009] In some examples, each dataset modality of the plurality of dataset modalities is treated independently with a modality-specific machine learning model. Each modalityspecific machine learning model projects its respective dataset modality into a shared representation space. The representations from each dataset modality are then aligned by contrastive learning by the foundation model.
[0010] In some examples, the dataset of each data modality can be derived from public and / or private and / or curated datasets / sources. In instances where curated datasets are utilized for training the machine learning models and / or generating the multimodal dataset, these datasets are, in some instances, publicly and / or privately available and / or a mixture of the two, generated datasets that have undergone modification. Furthermore, the dataset of each data modality, in some instances, include at least one of the following categories of data: chemical data, pharmacological data, cell imaging data, bioinformatics data, computational data, or cell biological data. In certain examples, a combination of datasets / sources are utilized to train the machine learning models and / or generate the multimodal dataset. In some examples, the combined datasets / sources or multimodal dataset can include, but are not limited to, any one or more of molecular structures, bioactivity assays, computationally- derived protei ligand docking scores, differential gene expression following gene knockdown, differential gene expression following compound treatment, scores for stimulation or inhibition of multiple biological pathways, transcription factors or genes derived from external computational methods applied to gene expression data, or cell morphological images following compound treatment datasets.
[0011] In certain implementations, a foundation model is trained on various datasets (e.g., the multimodal dataset) using an innovation on a contrastive learning approach. A non-exhaustive list of example contrastive learning approaches further innovated on includes: Contrastive Multiview Coding, geometric multimodal contrastive learning, InfoNCE, and multimodal contrastive learning. In some instances, the trained foundation model can output a representation for an agent which can be used as the input for the subsequent predictive model.
[0012] In some instances, the foundation model comprises a multimodal datasetembedding system comprising dataset modalities generated based on multiple and heterogeneous network architectures, including graph neural networks, convolutional neural networks, and feed-forward neural networks. The predictive model can be based on various machine learning architectures, which can include: a neural network, support vector machine learning, a random forest model, a decision tree model, linear regression, and a logistic regression model. Moreover, in some instances the predictive model in certain examples is a graph neural network.
[0013] In some implementations of the predictive model, an agent structure is presented to the predictive model to predict the modulatory or non-modulatory effect (e.g., activity or expression) of the agent on one or more biological targets, for example biological pathways, transcription factors or genes. In certain examples, the agent structure presented to the predictive model is a chemical structure or a coded representation of the chemical structure. An example of a chemical structure utilized in various examples of the present disclosure includes a small molecule. An example of a code representing a chemical structure is a SMILES string (“Simplified Molecular Input Line Entry System”) of the chemical structure. An example output of certain implementations of the present disclosure includes a measuring of the effect of introducing the agent on one or more biological targets, for example biological pathways, transcription factors or genes. In certain instances, the use of the predictive model allows for estimation of an agent’s probability of stimulating, inhibiting, or having no effect (e.g., activity or expression) on one or more biological targets, for example biological pathways, transcription factors or genes.
[0014] These and other objects, along with advantages and features of embodiments of the present invention herein disclosed, will become more apparent through reference to the following description, the figures, and the claims. Furthermore, it is to be understood that the features of the various embodiments described herein are not mutually exclusive and can exist in various combinations and permutations.BRIEF DESCRIPTION OF THE DRAWINGS
[0015] FIG. 1A is a schematic representation of a foundation / representational model based on training data featuring a plurality of dataset modalities. FIG. IB is a schematic representation of a predictive model that predicts the modulation of biological pathway(s), transcription factor(s) and / or gene(s) as a result of presentation of an agent.
[0016] FIG. 2 is a block diagram depicting steps of training a method to train a foundation model.
[0017] FIG. 3 is a block diagram depicting steps of estimating the biological pathway(s) activity of an agent.
[0018] FIG. 4A is a diagram of the pretraining of a foundation model. Several data modalities are used during pretraining: 1) the molecular structure of a small molecule; 2) the differential gene expression profile upon gene knockdown of the gene targeted by the small molecule; 3) the differential gene expression profile upon small molecule treatment; 4) blind docking scores and bioactivity assay results; 5) morphological features from cell imaging. A separate, independent encoder for each data modality is trained during pretraining to project its respective data modality into a single shared representation space. A main objective during pretraining is the contrastive loss, depicted by dotted purple arrows. The pretrained foundation model does not predict effects on pathways, transcription factors or genes. The ground truth data labels for effects on pathways are not used during pretraining.
[0019] FIG. 4B is a diagram of the training of a predictive model to predict effects on pathway activity via fine-tuning. The pretrained foundation model from FIG. 4A is shown, but the pretrained encoders for all data modalities are discarded upon completion of pretraining (here: greyed-out) except for the encoder for molecular structure. A projection head (e.g., a multilayer perceptron)(red, far right) then produces classification scores for effects on each pathway of interest. The molecular structure encoder from the foundation model and the new MLP project (red, far right) can be fine-tuned via supervised learning using the ground truth data labels for effects on pathway activity.
[0020] FIG. 4C is a diagram depicting the application of the fine-tuned model of FIG. 4B. The sole input to the fine-tuned model is the molecular structure of an agent. The output of the fine-tuned model is a separate classification score for each pathway of interest. The pretrained foundation model features the graph neural networks (GNN) over molecularstructure, readout, and MLP model components. Further specified are components (e.g., MLP) added during model fine tuning.
[0021] FIG. 4D is a diagram depicting a roadmap for implementation of the predictive model. The foundation model from FIG. 4A together with the appended MLP (red, top right) for the fine-tuned model to predict effects on pathways from FIG. 4B are shown. The numbers within the red circles indicate the order of implementation of the subcomponents of the model, la) and lb) would be implemented simultaneously, in parallel.
[0022] FIG. 5A is a graphical depiction of an exemplary foundation model utilized in the present disclosure, where four dataset modalities are utilized to construct a feature map of the shared representation space between molecular structures.
[0023] FIG. 5B is a graphical overview of the benchmarking approach utilized to determine the performance of the method of the present disclosure vs. baseline approaches (e.g., Classic ML and GNN baseline approaches).
[0024] FIG. 5C is a plot of a t-SNE embedding where 1 point represents 1 compound. Datapoints are colored and / or shaded by labels predicted from NP-Classifier.
[0025] FIG. 5D is a plot showing results from a non-parametric permutation test of distance between compounds with same labels predicted from NP-Classifier. Compounds with similar properties were found to have more similar representations than expected by chance (p < le-4).DETAILED DESCRIPTION
[0026] Methods, processes, and supporting systems described in the present disclosure use novel computational techniques and analysis of chemical and biological datasets for drug discovery purposes. It is, however, understood that the present disclosure can also be utilized for other scientific purposes-e.g., the study of physiological processes and / or phenomena where one or more biological pathway regulates, affects, or predicts said physiological process. In both instances, biological pathways are viewed as processes where biological elements (e.g., DNA, RNA, and / or proteins) interact with one another to drive one or more biological events. The boundaries of the biological pathways are understood to be flexible and non-limiting, as there is crosstalk among biological pathways and that multiple biological pathways can converge on the same biological event.
[0027] In some examples, the biological pathways can include, and are not limited to: metabolic processes, genetic information processing, environmental information processing, cellular processes, organismal systems, human diseases, and / or drug development pathways. It is understood that each of these processes can be composed of multiple biological and / or chemical elements that work together to impact a cellular / molecular process, which in turn can affect cellular events. These cellular events may have an impact on organ systems which can, in turn, modulate physiological processes that the organ system is responsible for. Often, these changes in organ systems cascade into organism-level activity changes and can result in various disease states.
[0028] As an illustrative non-limiting example, the process of external signal transduction can include a cell surface receptor (such as a G-protein Coupled Receptor, of which there are 100+ such in mammals) to recognize a specific substrate. This example receptor system can then use one or more signal transduction proteins / elements (in the example of G-protein Coupled Receptors, alpha, beta, and gamma subunits, each of which have multiple subtypes) to modulate the activity of a number of effector proteins and molecules. Non-exhaustive examples of such effector proteins and molecules include alpha, beta, and gamma subunits effector molecules / proteins include adenylyl cyclase, axin, cAMP, PKA, phosphodiesterases, phospholipases, PLCbeta, Ibc, Calcium2+, PKC, Rho, P115-RhoGEF, LARG, PDZ-RhoGEF, AKAP-lbc, PI3K, and ion channels.
[0029] Effector molecules can in turn interact with other biological molecules / elements such as transcription factors, which can drive gene expression that underlies biological responses such as cell proliferation, communication, survival, differentiation, migration, and extracellular matrix construction / degradation. These biological responses can impact organ system and organism function. As this example illustrates, the dimensions of biological pathway activity measurement can be large, as there is a complex network of both distinct and overlapping biological elements that interact with one another to modulate similar and / or different biological responses from the cellular to organism level. As such, there is a need for effective and efficient computational approaches for predicting biological pathway activity, including if desired a specific point on the pathway. The present disclosure offers an innovative approach to meet this need by, in part, performing dimensionality reduction of the vast number of biological features that underlie biological pathways in a manner that produces accurate and actionable predictions that can be leveraged for a range of applications, including drug discovery. Additionally, the present disclosure offers novelinsights and methods for generating, curating, and deploying seemingly disparate datasets for the classification of the activity of a biological pathway, the activity of a transcription factor or the expression of a gene. Such an approach vastly improves the computational efficiency of the systems used to process the input datasets, accelerating the drug discovery process.
[0030] In some embodiments, the methods of performing contrastive learning on multiple datasets to estimate the biological impact of an agent is demonstrated by FIGs. 1A-1B. FIG.1 A illustrates a method 100A where one or more agent structures 101A and one or more dataset modalities 101B are integrated using a contrastive learning approach 102 A applied to generate a representational map of the similarity and / or dissimilarity between each data point 103 A. In some embodiments, the method 100A is referred to as the foundation model. FIG. IB illustrates a method of 100B that is utilized to predict biological pathway (s) activity, transcription factor(s) activity, or gene(s) expression 103B as a result of exposing a biological system (e.g., DNA, RNA, and / or proteins in solution, cell(s), an organ system, or organism) to an agent 101 B through a machine learning model 102B based on the output 103 A of the foundation model of 100 A. In some embodiments, the approach of FIG. IB is utilized in the context of estimating the effect of an agent on biological pathways, transcription factors or genes and referred to as the predictive model. In some embodiments, FIG. 2 is an exemplary flow diagram that outlines one, of multiple, implementations of FIG. 1 A. In brief, to be elaborated in additional depth in subsequent sections, the foundation model within the system 200 is pretrained on multiple biological and / or chemical datasets 201 where an agent is presented, and the predictive model within the system is derived from the pretrained foundation model by training on a dataset and task of interest. A foundation model is pretrained with a contrastive learning loss function 202 on datasets to learn a representation through machine learning of an agent based on the various data modalities 203. Optionally, a representational map 204 of similarity and / or dissimilarity between other data from multiple data modalities can be generated.Datasets
[0031] Further detail is now provided concerning the datasets used for the present disclosure. In some instances, the biological and / or chemical datasets 201 are derived from public 201 A or private datasets 20 IB. Example public biological and / or chemical datasets 201 A can be sourced from public repositories featuring datasets or computational models such as PUBMED, PUBCHEM, CONNECTIVITY MAP, GITHUB, NIH GEO, DIFFDOCK, SMINA, CMAP, UNIDOCK, and CELL PAINTING. In some instances, thesedatasets are computationally derived. Example private biological and / or chemical datasets 20 IB can be generated or collected from non-public sources where the data generally concerns the molecular structure of small molecules, differential gene expression, binding affinity scores, computational data that estimates biological pathway activity, cell biological analysis, and / or bioactivity analysis curated or generated by a specific entity, which may consider these private datasets 20 IB as confidential and proprietary. Additionally, the biological and or chemical datasets 201, whether publicly sourced 201 A and / or privately sourced 201 B, can generally include the broad categories of chemical data, pharmacological data, bioinformatics data, computational data, or cell biological data 201 C.
[0032] In some instances, the datasets of 201C used for contrastive learning 202 are both public 201 A and privately 201B generated biological and chemical datasets composed of: molecular structures , gene expression (e.g., Connectivity Map, N > 28,000 small molecules), bioactivity assays (e.g., PubChem and privately generated datasets, N = 8.2M small molecules assays readout) and computationally-generated proteimligand docking scores (e.g., DiffDock+UniDock, N= 200,000 small molecules, Feature space = 60 proteins), differential gene expression following compound (e.g, agent) treatment (e.g., Connectivity Map, N> 28,000 small molecules, 978 genes), and cell morphological images following compound treatment datasets (e.g., Cell Painting, N= 117,000 small molecules, Feature space =1,800 precomputed features (3M images)).
[0033] In some examples, the docking fingerprint is computed from a model that scores binding affinity for a pose that was computed from another model. Docking fingerprints are mathematical vectors representing the three-dimensional structural characteristics of molecules based on computational estimates of their binding affinity to various proteins. They can be estimated, for example, by first performing docking simulations with a blind docking model, which identify one or more optimal poses of a molecule, then scoring that affinity with a scoring model which is in general distinct from the docking model, and then in some examples the embedding the affinity scores across proteins in a lower dimensional space.
[0034] In some examples, the gene expression data is determined by RNA-sequencing, RT-qPCR, or microarray data. Methods of generating RNA-sequencing can include any known method of collecting coding and non-coding RNA from a biological sample, performing library preparation, and sequencing steps. Example RNA-sequencing protocols and procedures can include: single-cell RNA-sequencing, single-nuclei RNA-sequencing, andbulk RNA-sequencing. RT-qPCR data can be generated using established or custom-built probes for genes of interest using any known protocol (e.g., Everaert, C., Luypaert, M., Maag, J.L.V. et al. Benchmarking of RNA-sequencing analysis workflows using whole- transcriptome RT-qPCR expression data. Sci Rep 7, 1559 (2017). incorporated by reference). Lastly, establishedcommercial or custom gene microarrays can be utilized for gene expression analysis. Each of these methods of gene expression analysis would follow conventional bioinformatic analysis practices and can additionally feature preprocessing steps.
[0035] In some embodiments, computational data is generated by computational methods / systems that estimate of biological pathway activity (e.g., stimulation, inhibition, and / or no change). Estimates from any known computational method / system that utilizes computer-aided estimation of biological activity profiles can be utilized. In some examples, computational data is sourced from algorithms that estimate the biological activity profiles of drug-like compounds (e.g., Filimonov DA, Rudik AV, Dmitriev AV, Poroikov VV. Computer-Aided Estimation of Biological Activity Profiles of Drug-Like Compounds Taking into Account Their Metabolism in Human Body. Int J Mol Sci. 2020 Oct 11 ;21(20):7492. doi: 10.3390 / ijms21207492. PMID: 33050610 (incorporated by reference)). In other examples, the computational data set is computationally-derived proteimligand docking scores.
[0036] In some embodiments, cell biological data are generated by any known methods and compositions directed towards characterization of biological components (e.g., molecules, organelles, proteins, DNA, RNA, etc.) within a biological system (e.g., cell(s), tissue, organs, organism, etc.). In some instances, the methods and compositions for characterization of cell biological processes includes microscopy (brightfield, fluorescence, atomic force, electron, etc.). These methods can optionally be combined with imaging approaches that utilize in situ hybridization (RNA), antibodies (protein detection), dyes (pH sensing, organelle labeling, etc.), and / or heavy metals (electron microscopy) to detect and quantify cellular features.
[0037] In some embodiments, bioactivity data is generated by any known biochemical or biophysical assay that measures the ability of a compound (e.g., agent) to bind and / or modulate a protein. In some instances, bioactivity data is deriving from radioligand assays, affinity chromatography, surface plasmon resonance, isothermal titration calorimetry and light-based methods (absorbance, fluorescence, or luminescence readouts).
[0038] In some instances, biological assays can also include cell-based bioactivity, biophysical, and biochemical assays. The results of these assays can be utilized in downstream applications such as predicting the biological activity of a chemical / molecular structure.
[0039] In other embodiments, methods of assessing genome architecture and epigenetic status can be used in conjunction or independently to the previously described gene expression methods. Such methods can include ATAC-sequencing or methylation assays.
[0040] In certain embodiments, a foundation model with data from multiple data modalities is trained. However, many compounds do not have data from all data modalities. For example, compound A may have gene expression data, but not cell imaging data, and vice versa for compound B. This is a challenge for multimodal contrastive learning approaches. Typical multimodal contrastive learning approaches cannot be run with incomplete data modalities for samples. They generally require all samples to have data for all data modalities. Overcoming this limitation, referred to as the problem of “incomplete data modalities”, is a critical challenge addressed in the present disclosure. This issue is common in biological application areas, and potentially beyond as well.
[0041] As such, in some embodiments, a new contrastive loss function enables one to perform multimodal contrastive learning even when there are incomplete data modalities. In some embodiments, the approach is derived from one or more contrastive loss functions for performing multimodal contrastive learning with complete data modalities.Contrastive Learning
[0042] As discussed above and in other sections, a biological pathway is a high-level, abstract concept of a process that can occur, generally, within a cell. It represents complex, heterogeneous interactions between a multitude of molecules, cellular actors and other biological pathways and systems. Identifying, inferring, or predicting the stimulation or inhibition of a pathway may, in some cases, require high-dimensional datasets to accurately characterize. However, the encompassed biological complexity present intrinsic challenges for computational systems directed towards the estimation of biological activity. In certain embodiments, the present disclosure teaches systems and methods of increasingly the tractability of estimating biological activity of agents on biological targets (e.g., the activity of a biological pathway or the activity of a transcription factor or the expression of a gene)through a combination of supervised and unsupervised machine learning methods, through contrastive learning.
[0043] Further detail is now provided concerning the application of contrastive learning to the datasets used for the present disclosure, 202 (FIG. 2). In some instances, contrastive learning refers to a method for training a machine learning model which produces a representation of a datapoint in a latent space based on similarity and / or dissimilarity between other data from data modalities. For example, a contrastive learning method may seek to train a model to produce representations that are similar for datapoints from data modalities that correspond to the same semantic entity, such as the same chemical compound, while also producing representations that are relatively less similar for datapoints from data modalities that do not correspond to the same semantic entity. In some instances, similarity between representations is computed using the cosine similarity. It is understood that contrastive loss functions are often implemented in unsupervised learning scenarios. In some cases, a contrastive loss function is utilized to perform the contrastive learning step 202. In some instances, the contrastive loss function 202 is one known in the art, such as SimCLR, max margin contrastive loss, a Triplet loss, N-pair loss, InfoNCE, and / or NT-Xent Loss. In other instances, the contrastive loss function 202 is one or a combination of the equations as described in Tian, Y. et. al., 2020 Dec 18 v. 5, Contrastive Multiview Coding, hEFA / Lirxjy orgA;bs / i906.05g.i1;9, incorporated by reference. Additionally, multiple approaches can be leveraged for the implementation of contrastive learning 202 to the biological and / or chemical datasets 201 including, but not limited to, Contrastive Multi view Coding, geometric multimodal contrastive learning, InfoNCE 202A. In some instances, the contrastive learning process includes a data augmentation module, such as transforming a data sample to create a positive or negative pair. For images, this can include a random crop, resizing, random flips, color distortions, and / or gaussian blur. Subsequently, a neural network-based encoder can be utilized to learn representations. While it is understood any network backbone can be utilized, for visual data (e.g., images) one can use a ResNet backbone, where the feature extraction occurs after the final averaging pooling layer. The computed representations are then mapped to a latent space through a neural network projection head. Lastly, the defined contrastive loss function is applied to the representations. Where, the output is a representation that is amenable for downstream analysis through one or more machine learning model(s). When applied to multimodal data, the disclosed method of use of contrastive learning can enrich the representation of one of the data modalities,which can serve as a more useful input for a subsequent machine learning model that is trained for a downstream task such as classification.Machine Learning Models and Classification
[0044] Further detail is now provided concerning the training of a machine learning foundation model with contrastive learning results / representations used in the present disclosure, 203 (FIG. 2). Machine learning is a collection of methods of for fitting parameters of a model from data. Machine learning can involve providing data into a computer program, which then uses computational, mathematical, and / or statistical analysis to identify patterns and relationships within the data. In supervised learning, the goal is to enable the program to make predictions or decisions based on these patterns and relationships learned from training data, without being explicitly told how to do so. Some examples of machine learning models include, but is not limited to, neural networks, graph neural networks (GNNs), support vector machines, random forest, decision tree, and logistic regression, 203 A.
[0045] Specifically, neural networks are a type of machine learning algorithm / model that are inspired by the structure and function of the human brain. They consist of layers of interconnected “neurons,” sometimes called nodes, which process and transmit information. Each neuron receives input from other neurons, processes it, and passes it on to other neurons in the next layer.
[0046] The layers in a neural network refer to the layers of interconnected neurons. There are typically multiple layers in a neural network, with the input layer receiving the raw data and the output layer producing the final prediction or decision. Between the input and output layers, there are one or more hidden layers, which process the data and pass it on to the next layer.
[0047] By training a neural network on a large dataset, the connections between neurons (called “weights41) can be adjusted to reduce the loss function. To train a neural network, the data is fed through the network and the output is evaluated by a loss function. In supervised learning, the loss function compares the output to a desired result, and if the output is not accurate, the weights are adjusted to reduce the error. This process is repeated multiple times, with the network continually adjusting the weights to reduce the loss. Once the network has been trained, it can be applied to new data.
[0048] In various examples, “machine learning” can refer to the application of certain techniques (e.g., pattern recognition and / or statistical inference techniques) by computersystems to perform specific tasks. Machine learning techniques (automated or otherwise) may be used to build data analytics models based on sample data (e.g., “training data”) and to validate the models using validation data (e.g., “testing data”). The sample and validation data may be organized as sets of records (e.g., “observations” or “data samples”), with each record indicating values of specified data fields (e.g., “independent variables,” “inputs,” “features,” or “predictors”) and corresponding values of other data fields (e.g., “dependent variables,” “outputs,” or “targets”). Machine learning techniques may be used to train models to infer the values of the outputs based on the values of the inputs. When presented with other data (e.g., “inference data”) similar or related to the sample data, such models may accurately infer the unknown values of the targets of the inference data set. Such models can be referred to herein as “machine learning models,” predictive models,” or “computer-implemented models.”
[0049] A feature of a data sample may be a measurable property of an entity (e.g., a cell, a biological sample, a person, thing, event, activity, etc.) represented by or associated with the data sample. For example, a feature can be a characteristic of a cell of an organism. As a further example, a feature can be a gene expression level associated with the cell. In some cases, a feature of a data sample is a description of (or other information regarding) an entity represented by or associated with the data sample. A value of a feature may be a measurement of the corresponding property of an entity or an instance of information regarding an entity.
[0050] For instance, if a feature of a cell is the expression level of a gene, a value of the feature can be an integer, which is the number of messenger RNA fragments in that cell that map to the region of gene on the human reference genome. In some cases, a value of a feature can indicate a missing value (e.g., no value). For instance, in the above example in which a feature is the gene expression level, the value of the feature may be ‘NULL,’ indicating that the gene level was not measured by a given technology.
[0051] Features can also have data types. For instance, a feature can have an image data type, a numerical data type, a text data type (e.g., a structured text data type or an unstructured (“free”) text data type), a categorical data type, or any other suitable data type. In the above example, the feature of a shape extracted from an image of a cell can be of an image data type. In general, a feature’ s data type is categorical if the set of values that can be assigned to the feature is finite.
[0052] As used herein, the “development” of a machine learning model may refer to construction of the machine learning model. Machine learning models may be constructed by computers using training datasets. Thus, “development” of a machine learning model may include the training of the machine learning model using a training data set. In some cases (generally referred to as “supervised learning”), a training data set used to train a machine learning model can include known outcomes (e.g., labels or target values) for individual data samples in the training data set. For example, when training a supervised computer vision model to detect images of cats, a target value for a data sample in the training data set may indicate whether the data sample includes an image of a cat. In other cases (generally referred to as “unsupervised learning”), a training data set does not include known outcomes for individual data samples in the training data set.
[0053] Following development, a machine learning model may be used to generate inferences with respect to “inference” datasets. For example, following development, a computer vision model may be configured to distinguish data samples including images of cats from data samples that do not include images of cats. As used herein, the “deployment” of a machine learning model may refer to the use of a developed machine learning model to generate inferences / predictions about data other than the training data.
[0054] In some instances, the output of the foundation model 200 is a representational map of the similarity and / or dissimilarity between each data point 204. In some embodiments, this output 204 is utilized for a classification task.
[0055] During deployment / implementation, 300, the classification task can include the prediction of modulation / non-modulation effects of an agent on a biological pathway. During classification, the machine learning model predicts whether a biological pathway is stimulated or inhibited 304A or there is no change in the biological pathway 304A. These predictions / estimations can be represented through known frameworks, including binary (yes / no) and / or probability scores assigned to each of the classes for each biological pathway 304. The machine learning model 303 can perform this classification through multiple methods through a projection head. The multiple methods of performing the classification can include k-nearest neighbors (KNNs), decision trees, support vector machines (SVMs), and / or GNNs 303A. In the case of GNNs, a SoftMax function (assignment of decimal probabilities to each class) approach can be utilized. Additionally, a one-vs-all approach can also be used for the GNN-based classification, where a binary classification (yes or no) is assigned to each of the classes.Deployment / Implementation of the Method to Characterize Effects on Biological Pathways
[0056] Further detail is now provided concerning the deployment of the system 300 (FIG. 3) for estimating effects of an agent structure 301 on biological pathways 304. In some embodiments, an agent structure 301 is provided as an input to the system 300. The agent structure 301 can be a chemical structure 301 A that can represent multiple types of compounds, including a small molecule 301 B, amino acid peptide, antibody, RNA / DNA structure, etc. Additionally, the system 300 includes a trained machine learning model (“foundation model”, 302) which received ground truth data from multiple data modalities for computing representations of a chemical compound. The subsequent, predictive machine learning model / proj ection head, derived from the foundation model, estimates the impact an agent structure can have on biological pathways (e.g., activate, inhibit, or no change, 304 / 304A). The multiple datasets used to train this foundation machine learning model 302 can include biological and / or chemical datasets, as described in further detail in other sections. The source of this data can include both public and / or private sources which can provide, either individually or together, chemical, pharmacological, bioinformatics, and / or cell biological data concerning the impact of an agent structure on biological pathway activity. These data sources serve as training data for which the system 300 is then suitable to estimate the biological activity of an agent structure 301 that has not been part of the training dataset.
[0057] In some embodiments, the foundation model is trained by contrastive learning. In some instances, the contrastive learning process implemented is derived from or related to Contrastive Multiview Coding, geometric multimodal contrastive learning, TuplelnfoNCE. Upon completion of training of the foundation model, an encoder from the foundation model can produce a representation of a chemical compound. In some instances, the encoders for the data modalities within the foundation model are neural networks 203A. A predictive machine learning model 300 is then trained by supervised learning to perform multi-class classification by predicting modulatory / non-modulatory effects of an agent on biological pathways 304. Methods for performing this classification are described in other sections. Generally, the model performs a classification based on the presented agent structure to determine stimulation, inhibition, or no effect on one or more of biological pathway 305A. The results of this classification can be calculated as a binary or probability score, as described in previous sections.
[0058] In some embodiments, the method is referenced as “Pathway Scoring” involving two distinct steps. In step 1, a multimodal foundation model is trained. This is called “pretraining”. This is where the multiple data modalities and contrastive learning are utilized. A foundation model is a collection of several machine learning models: there is one machine learning model, called an “encoder”, for each data modality. The encoder for each data modality is distinct. As such, for example, if the data modalities are molecular structure, gene expression, and cell imaging, then the foundation model consists of 3 encoders (machine learning models): one for molecular structure, one for gene expression, and one for cell imaging. For example, the input for the encoder for gene expression is the gene expression data, the input for the cell imaging encoder is the cell imaging data, and so on. The foundation model does not make predictions. The encoders do not make predictions. Rather, after pretraining is completed, each encoder from within the foundation model outputs a “learned representation” for its respective data modality, e.g., a long vector of numbers that is putatively a “good” representation of its respective data modality (i.e. “good” as in representations ability to be used to supervise a variety of downstream prediction task with high performance). Contrastive learning is the procedure by which the encoders within the foundation model are trained on the multimodal data. Step 2 comprises training a predictive model, derived from the encoder for molecular structure from the pretrained foundation model. The encoder for the molecular structure data modality is extracted from the foundation model. Next, a new machine learning model is defined. It is a predictive machine learning model. It is the concatenation of two machine learning models. The first is the encoder for molecular structure from the foundation model. The second is any machine learning model (ML), called the “projection head”. The input to this new predictive ML model is a molecular structure (no other data modality). The output is a prediction, e.g., classification for pathway activity / transcription factor activity / gene expression profiles. This predictive model is trained by supervised learning. In some examples, it would be trained on pathway activity / transcription factor activity / gene expression scores 304. The input is the molecular structure of a compound 301. The encoder from the foundation model outputs an abstract “representation” 302. The representation is input to the projection head 303. The output of the projection head is the e.g., probability score for classification 304 / 304A.Computer Implementation
[0059] In some examples, some or all of the processing described above can be carried out on a personal computing device, on one or more centralized computing devices, or viacloud-based processing by one or more servers. Some types of processing can occur on one device and other types of processing can occur on another device. Some or all of the data described above can be stored on a personal computing device, in data storage hosted on one or more centralized computing devices, and / or via cloud-based storage. Some data can be stored in one location and other data can be stored in another location. In some examples, quantum computing can be used and / or functional programming languages can be used. Electrical memory, such as flash-based memory, can be used.
[0060] General-purpose computers, network appliances, mobile devices, or other electronic systems may also include at least portions of the system. The system includes a processor, a memory, a storage device, and an input / output device. Additionally, one or more graphical processor unit (GPUs) may be also included in the aforementioned system. Each of the components may be interconnected, for example, using a system bus. The processor is capable of processing instructions for execution within the system. In some implementations, the processor is a single-threaded processor. In some implementations, the processor is a multi-threaded processor. The processor is capable of processing instructions stored in the memory or on the storage device.
[0061] The memory stores information within the system. In some implementations, the memory is a non- transitory computer-readable medium. In some implementations, the memory is a volatile memory unit. In some implementations, the memory is a non-volatile memory unit.
[0062] The storage device is capable of providing mass storage for the system. In some implementations, the storage device can be a non-transitory computer-readable medium. In various different implementations, the storage device may include, for example, a hard disk device, an optical disk device, a solid-date drive, a flash drive, or some other large capacity storage device. For example, the storage device may store long-term data (e.g., database data, file system data, etc.). A input / output device provides input / output operations for the system. In some implementations, the input / output device may include one or more of a network interface device, e.g., an Ethernet card, a serial communication device, e.g., an RS-232 port, and / or a wireless interface device, e.g., an 802.11 card, a wireless modem. In some implementations, the input / output device may include driver devices configured to receive input data and send output data to other input / output devices, e.g., keyboard, printer, and display devices. In some examples, mobile computing devices, mobile communication devices, and other devices may be used.
[0063] In some implementations, at least a portion of the approaches described above may be realized by instructions that upon execution cause one or more processing devices to carry out the processes and functions described above. Such instructions may include, for example, interpreted instructions such as script instructions, or executable code, or other instructions stored in a non-transitory computer readable medium. A storage device may be implemented in a distributed way over a network, for example as a server farm or a set of widely distributed servers or may be implemented in a single computing device.
[0064] In some embodiments of the subject matter, functional operations and processes described in this specification can be implemented in other types of digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible nonvolatile program carrier for execution by, or to control the operation of, data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them.
[0065] The term “system” may encompass all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. A processing system may include special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). A processing system may include, in addition to hardware, code that creates an execution environment for the computer program in question, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0066] A computer program (which may also be referred to or described as a program, software, a software application, an engine, a pipeline, a module, a software module, a script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in anyform, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, sub programs, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0067] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit).
[0068] Computers suitable for the execution of a computer program can include, by way of example, general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random-access memory or both. A computer generally includes a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or he operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name just a few.
[0069] Computer readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0070] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can he received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’ s user device in response to requests received from the web browser.
[0071] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (“LAN”) and a wide area network (“WAN”), e.g., the Internet.
[0072] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0073] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, oneor more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0074] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0075] It is contemplated that systems, methods, and processes of the claimed invention encompass variations and adaptations developed using information from the embodiments described herein. Adaptation and / or modification of the systems, methods, and processes described herein may be performed by those of ordinary skill in the relevant art.
[0076] It should be understood that the order of steps or order for performing certain actions is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.ExamplesSystem and Task Overview
[0077] To illustrate the tractability of training the system of 200 (FIG. 2, the representational / foundation model) and its implementation / deployment 300 (FIG. 3, the predictive model), the following examples are provided (FIGs. 4A-4D and 5A-5B). The modeling task was predicting the effects of an agent (e.g., chemical entity) on the activity of biological pathways, transcription factors and / or genes of interest. The initial input to the trained model was a molecular structure (SMIEES string) of an agent. The output of the trained model can be a distinct three-way continuous simplex score for each pathway, transcription factor and / or gene of interest. The outcome categories for each pathway are: 1) stimulates, 2) inhibits, or 3) no change. These categories are interpreted relative to the basal activity level of the respective pathway, transcription factor or gene in the cell. Thus, thetrained model performs multi-task multi-class classification given the molecular structure of an agent.
[0078] The trained model (the predictive model) for predicting effects on pathways can be obtained by finetuning a pretrained foundation model in a supervised setting. The foundation model includes additional data modalities and is pretrained through a selfsupervised contrastive learning method. The additional data modalities involved in pretraining the foundation model are not required for input to the fine-tuned model at inference time. Therefore, the fine-tuned model can predict the effect of any agent on pathways, transcription factors and genes of interest from its molecular structure alone.
[0079] The model for predicting the effects of agent on pathway activity from their molecular structures alone can be fine-tuned from the pretrained foundation model via a supervised learning task. Specifically, the supervised learning task is to predict the effect on pathway activity using ground truth classification labels (stimulates, inhibits, or no change).Datasets and Biological Pathway Ground Truth
[0080] The ground truth classification labels for effects on pathway activity, transcription factor and gene activity for the training data can be derived from gene expression profiles of small molecule treatment experiments. In one example, the Connectivity Map includes >720k prepared gene expression profiles spanning 39k small molecules and >80 cell lines. In this example, a subset of molecules with a history of safe human consumption “subset molecules”, comprising 345 of the 39k small molecules interrogated in the Connectivity Map was utilized. An initial strategy used to estimate model performance for the subset molecules chemical space specifically was to holdout all 345 subset molecules as part of the hold-out test set, in addition to many other natural products in the Connectivity Map. Subsequently, a private dataset for additional subset molecules to improve the training of the model and improve the reliability of this estimate across the chemical space, though data generation is not required to develop the model proposed here was generated. The LI 000 microarray platform used in the Connectivity Map was an efficient starting point because it had an established protocol, had an open-source data processing pipeline, and was lowcost. RNA-seq data was generated by developing an analogous protocol and data processing pipeline and can be readily incorporated. By any approach, one could strategically select additional subset molecules to capture diversity in the subset molecules chemical space to improve reliability of the estimate of model performance.
[0081] The derivation of ground truth labels for effects on pathways hinges on key insights, including: 1. Stimulation or inhibition of a pathway upon small molecule treatment can be transient or short-lived [Gilmore TD. Introduction to NF-kappaB: players, pathways, perspectives. Oncogene. 2006 Oct 30;25(51):6680-4. doi: 10. 1038 / sj. one.1209954. PMID: 17072321 ; Bliithgen N, Legewie S. Robustness of signal transduction pathways. Cell Mol Life Sci. 2013 Jul;70(13):2259-69. doi: 10.1007 / s00018-012-l 162-7. Epub 2012 Sep 25. PMID: 23007845; Bliithgen N. Signaling output: it's all about timing and feedbacks. Mol Syst Biol. 2015 Nov 27;1 1(11 ):843. doi: 10.15252 / msb.20156642. PMID: 26613962], whereas samples are often collected at a later timepoint in which effects of a downstream cascade have manifested [Parikh JR, Klinger B, Xia Y, Marto JA, Bliithgen N. Discovering causal signaling pathways through gene-expression patterns. Nucleic Acids Res. 2010 Jul;38(Web Server issue):W109-17. doi: 10.1093 / nar / gkq424. Epub 2010 May 21. PMID: 20494976, incorporated by reference]. 2. The induced gene expression profile can be used to infer that a particular pathway was stimulated, inhibited, or unaffected. 3. The genes whose expression levels provide evidence that a particular pathway was stimulated or inhibited and the genes whose protein products are involved in the signaling of the pathway can be quite different [Parikh JR, Klinger B, Xia Y, Marto JA, Bliithgen N. Discovering causal signaling pathways through gene-expression patterns. Nucleic Acids Res. 2010 Jul;38(Web Server issue):W109- 17. doi: 10.1093 / nar / gkq424. Epub 2010 May 21. PMID: 20494976; Schubert, M., Klinger, B., Klunemann, M. et al. Perturbation-response genes reveal signaling footprints in cancer gene expression. Nat Commun 9, 20 (2018). btins: / / doi.org / 10.1038 / s4l 467-017-02391 -6;Liu A, Trairatphisan P, Gjerga E, Didangelos A, Barratt J, Saez-Rodriguez J. From expression footprints to causal pathways: contextualizing large signaling networks with CARNIVAL. NPJ Syst Biol Appl. 2019 Nov 11;5:40. doi: 10.1038 / s41540-019-0118-z. PMID: 31728204, all incorporated by reference].
[0082] An established bioinformatics method which addresses these properties was used to derive the ground truth label for pathway effects from the differential gene expression profiles [Schubert, M., Klinger, B., Klunemann, M. et al. Perturbation-response genes reveal signaling footprints in cancer gene expression. Nat Commun 9, 20 (2018).» Pau Badia-i-Mompel, Jesus Velez Santiago, Jana Braunger, Celina Geiss, Daniel Dimitrov, Sophia Miiller-Dott, Petr Taus, Aurelien Dugourd, Christian H Holland, Ricardo O Ramirez Flores, Julio Saez-Rodriguez, decoupleR: ensemble of computational methods to infer biological activities from omicsdata, Bioinformatics Advances, Volume 2, Issue 1, 2022, both incorporated by reference]. Themethod is based on the PROGENy set of pathway-specific footprints; that was, weighted sets of genes whose expression was regulated as a direct consequence of activation of a specific pathway [Schubert 2018]. PROGENy was built from a curated set of >560 controlled pathway-perturbation experiments over >2600 microarrays. PROGENy includes a multivariate linear model fit to differential gene expression profiles of pathway-perturbation experiments from which p-values for direction-aware pathway-specific effects (i.e., stimulated vs inhibited vs no change) can be obtained for a new differential gene expression profile of interest. This method has been validated with phosphoproteomics and holdout pathway-perturbation experiments [Schubert 2018; Badia-I-Mompel 2022], The companion software package decoupleR can be used to compute the direction-aware scores and p values from the multivariate linear model [Badia-I-Mompel 2022]. The direction aware scores from the statistical hypothesis tests can be converted directly into ground truth labels for pathway effects (stimulates vs inhibits vs no change). Ground truth labels for the following pathways can be obtained using the PROGENy method: Androgen, EGFR, Estrogen, Hypoxia, JAK- STAT, MAPK, NFkB, p53, PI3K, TGFb, TNFa, Trail, VEGF, WNT. Ground truth labels for the following additional pathways, H2O2, Hippo, IL-1, Insulin, TLR, Notch, and PPAR, could be derived through a related bioinformatics method, SPEED2 [Rydenfelt M, Klinger B, Klunemann M, Bliithgen N. SPEED2: inferring upstream pathway activity from differential gene expression. Nucleic Acids Res. 2020 Jul 2;48(Wl):W307-W312. doi: 10.1093 / nar / gkaa236. PMID: 32313938, incorporated by reference]. SPEED2 was also built on footprints from a curated set of pathway-specific perturbation experiments [Parikh 2010; Rydenfelt 2020], analogous to PROGENy. SPEED2 applies the Bates Test to differential expression profiles to compute direction aware scores and p-values for effects on pathways. The senior authors of PROGENy and SPEED2 overlap, and the methods share the footprint idea and curated-experiments approach. SPEED2 was not validated with phosphoproteomics, but rather with gene expression profiles from holdout single pathway perturbation experiments.
[0083] (Phospho)proteomics and gene expression profiling are techniques orthogonal to bioassays for inferring modulation of pathways. Phosphoproteomics can be challenging data source for building models because the experiments are low throughput, challenging to perform, and rare [Staudt DE, Murray HC, Skerrett-Byme DA, Smith ND, Jamaluddin MFB,Kahl RGS, Duchatel RJ, Germon ZP, McLachlan T, Jackson ER, Findlay IJ, Kearney PS, Mannan A, McEwen HP, Douglas AM, Nixon B, Verrills NM, Dun MD. Phospho-heavy- labeled- spiketide FAIMS stepped-CV DDA (pHASED) provides real-time phosphoproteomics data to aid in cancer drug selection. Clin Proteomics. 2022 Dec 19;19(1):48. doi: 10.1186 / s 12014-022-09385 -7. Erratum in: Clin Proteomics. 2023 Apr 8;20(l):16. doi: 10.1186 / sl2014-023-09406-z. PMID: 36536316 (incorporated by reference), Liu 2019, Schubert 2018]. In contrast, gene expression profiling experiments were more abundant and easily performed, thus opening up the possibility of obtaining ground truth data labels for pathway modulation from public and private datasets. Further, gene expression profiling was not susceptible to the same sources of systematic false positives in bioassay designs which necessitate counter-screens and secondary screens.
[0084] Ground truth labels of pathway activity for a pathway not listed above can be obtained by following the procedure of Schubert 2018 for building PROGENy, upon two conditions. The two conditions were: 1. It was possible to perform a controlled perturbation experiment that stimulates or inhibits that pathway in particular, and ideally no other pathway (though crosstalk was inevitable for certain pathways and accounted for in the procedure below). 2. Gene expression data from such controlled perturbation experiments are available, or the experiments can be performed. If these two conditions are satisfied, then one can follow the procedure described in Schubert 2018 to obtain a footprint for the pathway, transcription factor or gene. That procedure was summarized as follows: 1. Curate gene expression profiles from controlled perturbation experiments for the pathway. RNA-seq or microarray data can be used. If there was insufficient suitable public data, then singlepathway-perturbation experiments must be performed. Schubert 2018 found that 20-50 experiments were needed to build a footprint for a single pathway that can be applied to infer pathway activity in any cell line reliably. If the provided scope was restricted to a set of cell lines, then the number of experiments necessary may be less. 2. Perform a differential expression analysis in each controlled perturbation experiment. 3. Derive footprints from fitting a multivariate linear model between pathways and genes. Details in Schubert 2018.
[0085] The output of this procedure was a footprint for the pathway; that was, a set of genes with signed weights. Subsequently, the footprint can be used to infer the activity of the pathway (stimulated vs inhibited vs no change) from any gene expression profile. In other words, one could use the footprint to generate a ground truth label for activity of that pathway from any gene expression profile. One could extend PROGENy to include the footprint forthis pathway. One could then expand the ground truth labels to include the pathway by applying the updated PROGENy method to the gene expression profiles upon small molecule treatment from Connectivity Map. There are >720k gene expression profiles upon small molecule treatment in the Connectivity Map, and 21 pathways already with footprints via PROGENy and SPEED2 (above), yielding >15. IM = 720k * 21 ground truth labels. Each additional pathway built by this procedure would yield an additional >720k ground truth labels.Pretraining a Foundation Model
[0086] A foundation model can be pretrained using contrastive learning and additional data modalities (FIG. 4A). Pretraining does not necessarily involve ground truth labels for pathway activity. Upon completion of pretraining, the foundation model does not predict anything per se. Rather, the purpose of the foundation model was to learn an encoder that produces an embedding of a molecule from its molecular structure alone, where the embedding was enriched because it was informed by additional data modalities via contrastive learning. Subsequently, the model for pathway activity will be derived by transfer learning, starting from an encoder from the pretrained foundation model. Example additional data modalities and their models that comprise the foundation model are as follows.Molecular Structure
[0087] A GNN over molecular structure can be pretrained offline (”pre-pretrained”) using an auxiliary chemical knowledge graph and contrastive learning, following Pengcheng Jiang, Cao Xiao, Tianfan Fu, Jimeng Sun, Bi-level Contrastive Learning for Knowledge-Enhanced Molecule Representations, arXiv, June 2, 2023, https: / / doi.org / lO.4855O / arXiv.23O6.016 1, incorporated by reference. The chemical knowledge graph MolKG [Jiang 2023] can be built from PrimeKG [Chandak P, Huang K, Zitnik M. Building a knowledge graph to enable precision medicine. Sci Data. 2023 Feb 2;10(l):67. doi: 10.1038 / s41597-023-01960-3.PMID: 36732524, incorporated by reference] and PubChemRDF (PubChem Substance and Compound information) [Fu 2015]. This “pre-pretraining” step was independent of the pretraining strategy described below and does not involve pathways or any of the data modalities described here. Subsequently, during the pretraining of the foundation model, a graph embedding will be readout from the pretrained GNN. A shallow MLP that projects the graph embedding into the shared representation space will be trained during pretraining via contrastive learning (below). Importantly, pretrained molecular structure representations haveboosted the performance of molecular property prediction tasks because they incorporate additional prior knowledge (e.g., chemical knowledge graphs) [Wang, Y., Wang, J., Cao, Z. et al. Molecular contrastive learning of representations via graph neural networks. Nat Mach Intell 4, T19- S (2022). https: / / doi.org / 10.1038 / s42256-022-00447-x; Fang, Y„ Zhang, Q., Zhang, N. et al. Knowledge graph-enhanced molecular contrastive learning with functional prompt. Nat Mach Intell 5, 542-553 (2023). https: / / doi.org / 10.1038 / s42256-023- 00654-0; Jiang 2023, all incorporated by reference].Gene Expression Profiles from Gene Knockdown Perturbations
[0088] A GNN can be trained over a graph whose nodes represent genes and edges indicate co-membership in Biological Processes from the Gene Ontology (GO). The node features was the prepared differential gene expression profiles from experiments involving the gene knockdowns (no small molecules involved) from the Connectivity Map. A graph embedding can be a readout from the GNN. A shallow MLP can then project the graph embedding into the shared representation space. The GNN and MLP projector was trained during pretraining via contrastive learning (below). One source defines the graph from GO Biological Processes based on the biological intuition that, “. . . genes that are involved in similar pathways [i.e. GO biological processes] should impact the expression of similar genes after perturbation.” [Roohani Y, Huang K, Leskovec J. Predicting transcriptional outcomes of novel multigene perturbations with GEARS. Nat Biotechnol. 2024 Jun;42(6):927-935. doi: 10. 1038 / s41587 -023-01905-6. Epub 2023 Aug 17. PMID: 37592036, incorporated by reference]. Gene expression profiles observed upon small molecule treatment or knockdown of the target of the small molecule treatment are highly correlated [Pabon NA, Xia Y, Estabrooks SK, Ye Z, Herbrand AK, et al. (2018) Predicting protein targets for drug-like compounds using transcriptomics. PLOS Computational Biology 14(12): incorporated by reference]. Proteintargets of small molecules of interest have been identified from such pairs of gene expression profiles, as have small molecules which can target a protein of interest [Pabon 2018;Feisheng Zhong, Xiaolong Wu, Ruirui Yang, Xutong Li, Dingyan Wang, Zunyun Fu, Xiaohong Liu, XiaoZhe Wan, Tianbiao Yang, Zisheng Fan, Yinghui Zhang, Xiaomin Luo, Kaixian Chen, Sulin Zhang, Hualiang Jiang, Mingyue Zheng, Drug target inference by mining transcriptional data using a novel graph convolutional network framework, Protein & Cell, Volume 13, Issue 4, April 2022, Pages 281-301, hups : / / doi .org / 10. 1007 / s 021 ■■00885-0, incorporated by reference]. Therefore, gene expression profiles from geneknockdown perturbation experiments provided an alternative “view” of the small molecule to facilitate contrastive learning.Gene Expression Profiles from Small Molecules Treatment
[0089] A GNN can be trained over a graph whose nodes represent genes and edges indicated co-membership in Biological Processes from the GO. The node features can be the prepared differential gene expression profiles from experiments involving small molecule treatments from the Connectivity Map. A graph embedding can be readout from the GNN. A shallow MLP will then project the graph embedding into the shared representation space. The GNN and MLP projector will be trained during pretraining via contrastive learning (below). The justifications for the use of GO Biological Processes and gene expression profiles was analogous to the justifications given above for gene expression profiles from gene knockdown experiments.Cell Morphology Images from Small Molecule Treatments
[0090] An MLP was trained to encode precomputed handcrafted features from morphological profiles into the shared representation space. The morphological profiles are from the Cell Painting high-throughput imaging database. The handcrafted features are computed from the CellProfiler imaging analysis software and provided with the Cell Painting dataset. Recent studies integrated cell morphology from Cell Painting with chemical structure [Trapotsi MA, Mervin LH, Afzal AM, Sturm N, Engkvist O, Barrett IP, Bender A. Comparison of Chemical Structure and Cell Morphology Information for Multitask Bioactivity Predictions. J Chem Inf Model. 2021 Mar 22;61(3): 1444-1456. doi: 10.1021 / acs.jcim.0c00864. Epub 2021 Mar 4. PMID: 33661004 (incorporated by reference)] and gene expression profiles [Moshkov, N., Becker, T., Yang, K. et al. Predicting compound activity from phenotypic profiles and chemical structures. Nat Commiin 14, 1967 (2023). htps: / / doi. org / 10.1038 / s41467 -023 ■ 37570- 1 (incorporated by reference)] to predict bioactivity assay results for small molecules. The studies showed that each data modality provided complementary information that contributed to improved performance. The small molecules examined in the Cell Painting dataset substantially overlap with those examined in the Connectivity Map dataset [Moshkov 2023].Bioactivity Assay Results and Blind Docking Scores
[0091] A GNN over a protein-protein interaction (PPI) network can be trained to project bioactivity assay results and blind docking scores into the shared representation space. ThePPI network can be from the PrimeKG knowledge graph because it integrates over several PPI databases [Chandak 2023]. The blind docking scores can be computed from the pretrained DiffDock model [Gabriele Corso, Hannes Stark, Bowen Jing, Regina Barzilay, Tommi Jaakkola, DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking, October 4, 2022, b t ps : / / d» . or g / 10.48550 / arXi y .2210.01776 (incorporated by reference) ]. The bioactivity assay results can be from the ExCape database [Sun J, Jeliazkova N, Chupakin V, Golib-Dzib JF, Engkvist O, Carlsson L, Wegner J, Ceulemans H, Georgiev I, Jeliazkov V, Kochev N, Ashby TJ, Chen H. ExCAPE-DB: an integrated large scale dataset facilitating Big Data analysis in chemogenomics. J Cheminform. 2017 Mar 7;9: 17. doi: 10.1186 / s 13321 -017-0203-5. Erratum in: J Cheminform. 2017 Jun 14;9(1):41. doi: 10.1186 / s 13321 -017-0222-2. PMID: 28316655 (incorporated by reference)], which was drawn from PubChem and ChEMBL. Privately generated bioactivity assay results can be included here as well. Models for predicting bioactivity assay results can have higher performance when trained on historical bioactivity assay results as compared to models trained on chemical fingerprints alone [Riniker S, Wang Y, Jenkins JL, Landrum GA. Using information from historical high-throughput screens to predict active compounds. J Chem Inf Model. 2014 Jul 28;54(7): 1880-91. doi: 10.1021 / ci500190p. Epub 2014 Jun 26. PMID: 24933016 (incorporated by reference)Similarly, the inclusion of docking scores improved the performance of a model for predicting compound- protein interactions [Pabon 2018]. As such, bioactivity assays and blind docking scores can be analyzed together in a single GNN because they both report on physical interactions and direct binding targets of the small molecule and because each alone was sparse within the full subset molecules x Protein space. Lastly, A PPI network was used because it likely represents the effects of physical interactions better than other networks, such as gene co-expression networks. Further, a PPI network subsumes pathways and naturally relates pathways to each other through interactions and shared proteins in the PPI network. Using pathways alone would not capture these interpathway interactions. Further, pathway knowledge was limited because pathways are conceptual constructions, whereas PPI networks are less constrained because PPIs are measured at-scale in experiments. Also, compound-protein interactions assessed by blind docking scores and bioactivity assays do not in general manifest in the expression levels of the genes involved in the interactions, making gene co-expression networks less appropriate for contextualizing this information.Contrastive Learning
[0092] The representations from each data modality can be aligned by contrastive learning. Successively more complex contrastive learning strategies will be pursued. Initial testing can include Contrastive Multiview Coding [Tian, Y., Krishnan, D., Isola, P. (2020). Contrastive Multiview Coding. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, JM. (eds) Computer Vision - ECCV 2020. ECCV 2020. Lecture Notes in Computer ScienceQ, vol 12356. Springer, Cham. I•' (incorporated by reference)] because of its relative simplicity and impact. Then, a geometric multimodal contrastive (GMC) learning strategy can be tried, which aligns the representation of each separate modality to their more informed multimodal fusion representation [Petra Poklukar, Miguel Vasco, Hang Yin, Francisco S. Melo, Ana Paiva, Danica Kragic, Geometric Multimodal Contrastive Representation Learning, Proceedings of the 39th International Conference on Machine Learning, PMLR 162: 17782-17800, 2022 (incorporated by reference)]. This strategy can be explored secondarily because of its limited complexity. Lastly, the TuplelnfoNCE method seeks to ensure that “complementary synergies between modalities that might be useful for downstream tasks” are leveraged and that information from weaker data modalities was not ignored during contrastive learning [Yunze Liu, Qingnan Fan, Shanghang Zhang, Hao Dong, Thomas Funkhouser, Li Yi, Contrastive Multimodal Fusion With TuplelnfoNCE, Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV), 2021, pp. 754-763 (incorporated by reference)]. A simplified baseline approximation of the TuplelnfoNCE loss can be explored next, followed by the full optimization method described by the authors. A method for extending the contrastive loss to include incomplete data modalities (e.g., unmatched molecules and targets) was also explored [Ryumei Nakada, Halil Ibrahim Gulluk, Zhun Deng, Wenlong Ji, James Zou, Linjun Zhang, Understanding Multimodal Contrastive Learning and Incorporating Unpaired Data, Proceedings of The 26th International Conference on Artificial Intelligence and Statistics, PMLR 206:4348-4380, 2023. (incorporated by reference) ].Training a Model to Predict Effects on PathwaysFine-Tuning from the Pretrained Foundation Model
[0093] A model for the ultimate task of predicting the effect of a subset molecule on each pathway of interest was obtained by fine-tuning the pretrained foundation model (FIG. 4B). This model was obtained from the pretrained foundation model by appending a shallow MLPthat generates the multi-task multi-class classification scores (stimulates, inhibits, no change) for each pathway of interest from the shared representation space. The model was fine-tuned through a supervised learning task, where the ground truth data labels for the effects on pathway activity, in some instances, are derived from PROGENy, as described above. This fine-tuning stage was the only stage in which the ground truth data labels for the effects on pathway activity are used.Architecture of the Fine-Tuned Model
[0094] Description of how the model obtained by fine-tuning the pretrained foundation model operates at inference time to predict the effects of the subset molecules on pathways of interest was shown in the above text and FIG. 4C. The input molecular structure (SMILES string) of a subset molecule was processed by a GNN over the molecular structure graph to obtain a graph embedding. The graph embedding was then projected into what was once the shared representation space by an MLP. Then, another MLP computes the classification scores for each biological pathway of interest simultaneously from the embedding in what was once the shared representation space. Where, the classification scores for each are represented as the probability (0-1.0) of activating, inhibiting, or not changing a given biological pathway.Adaptation
[0095] The modularity of the disclosed approach enables implementation and testing of the end-to-end model without inclusion of all data modalities (molecular structure, gene expression, cell morphology, and bioactivity). In particular, one can implement the proposed model incrementally with frequent checkpoints (FIG. 4D). Each checkpoint can include a full end-to-end performance evaluation of the model (pretrain a foundation model — > fine-tune on pathway activation labels — > evaluate on a holdout test set — > make predictions for any subset molecule). The first checkpoint represents a minimal model which serves as a low-cost, low effort baseline. Namely, it effectively replaces the ground truth labels used to train the current ACN model with ground truth labels for effects on pathway derived from PROGENy.
[0096] The sequence of incremental checkpoints for implementing the full model can be as follows (FIG. 4D):1. A minimal directed message passing neural network (D / WJAW)-like model: a. Implement the GNN encoder for molecular structure, b. Use PROGENy to derive ground truth labels for effects on pathway effects from the Connectivity Map.2. Implement the GNN encoder for differential gene expression profiles upon small molecule treatment and train with data from the Connectivity Map.3. Implement the GNN encoder for differential gene expression profiles upon gene knockdown perturbations and train with data from the Connectivity Map.4. Implement the MLP encoder for precomputed handcrafted CellProfiler features of cell morphology and train with the Cell Painting dataset.5. Implement the GNN encoder for bioactivity assay results.6. Include docking scores into the GNN encoder together with bioactivity assay results.Foundation and Predictive Model Implementation and Performance
[0097] To evaluate the performance of the aforementioned method, a non-limiting specific implementation of the Pathway Scoring model was utilized to predict bioactivity assay outcomes for two target-specific bioassays (e.g., target 1 and target 2). This method was then compared to benchmark bioactivity models as described below. The method comprises the two-step formulation of the Pathway Scoring approach (FIG. 5 A):1. Pretraining a multimodal foundation model on four data modalities. a. molecular structure b. gene expression from compound treatment experiments c. pre-featurized cell imaging data from compound treatment experiments d. pathway and transcription factor activity scores from compound treatment experiments2. Transfer learning to predict bioactivity assay outcomes of target 1 and target 1 from molecular structure alone, starting from the encoder for molecular structure from the pretrained multimodal foundation model.
[0098] To benchmark the Pathway Scoring Model, two gradient boosted trees models were trained (one for each target) to predict bioactivity assay outcomes from a molecule structure (FIG. 5B Classic ML baseline). For another benchmark, a single GNN was trained to predict bioactivity assay outcomes from molecular structure (DMPNN model) in a multitask fashion (FIG. 5B GNN baseline).
[0099] The Pathway Scoring outperformed all baseline models as shown in Table 1 below.Table 1
[0100] Table 1 shows the Area Under the Receiver Operating Characteristic curve (AUROC) for the three models mentioned in FIG. 5B. The best scores for each metric are bolded. Evaluation was performed separately for two target- specific bioactivity assays, target 1 and target 2, as binary classification tasks. The same single train / test split was used to train and evaluate all models.
[0101] Additionally, properties of natural products are emergent in learned representations of compounds (FIG. 5C) and compounds with similar properties have more similar representations than expected by chance (p < le-4; FIG. 5D).Utilization
[0102] Together, the aforementioned examples of identifying, curating, and modifying seemingly disparate datasets, pretraining a foundation model by a contrastive learning method, and finally training a machine learning model for multi-task multi-class classification of biological pathway activity illustrate one embodiment of the present disclosure. The implementation of this exemplary system can occur across the fields of chemistry, biology, and computer science. Where, the deployment of the exemplary system can be utilized to expedite the discovery of therapeutic molecules. For example, the exemplary system can be executed on a computer or web-based platform where a user can input one or more agent structures. The user then would receive a list of biological pathways and their modulation / non-modulation status (0-1.0 probability assignment or yes / no per classification category (activate, inhibit, or no change)). The user can then utilize this output to prioritize agent structures with a desired biological activity profile. For example, a desirable biological activity profile can include modulation (activation / inhibition) of a specific biological pathway while having no modulatory effects on other biological pathways. In this instance, the agent structure could be classified as a specific modulator with low to nooff-target effects. The user could then, optionally, perform a series of empirical validation experiments to generate data to show the validity of such a prediction that would inform downstream drug discovery efforts.
[0103] One unique feature of the present disclosure relative to previous automation attempts to estimate the impact of an agent structure on biological pathways was the innovative use of seemingly disparate datasets. As discussed in previous sections, previously disclosed automation attempts have relied on narrower and more homogeneous datasets and data modalities for training a machine learning model for this estimation / classification task, due, in part, to the computational challenges of unifying complex and high dimensional datasets for a classification task. As such, the exemplary system offers a tractable, modular solution that integrates multiple datasets into a modulatory status classifier based on a provided agent structure. The potential benefits of this approach may include improved accuracy of a downstream classification task, which in turn can expedite the processes of evaluating investigational compounds. In particular, when the examples and teachings of the present disclosure are combined with computational generation of chemical structures, a user is enabled to identify promising chemical structures for the desired biological / therapeutic applications entirely through in silico methods, which improves the throughput and efficiency of the early stages of drug discovery.
[0104] It also is understood that various other embodiments may be practiced, given the genera] description provided herein. Additionally, the above are examples of specific embodiments for carrying out the claimed subject matter of the present disclosure. The examples are offered for illustrative purposes only, and are not intended to limit the scope of the present disclosure in any way. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, measurements, etc.), but some experimental error and deviation should, of course, be allowed for.
[0105] What is claimed is:
Claims
CLAIMS1. A method of generating a predictive model to predict modulation of a biological target, the method comprising: providing a multimodal dataset comprising a plurality of distinct dataset modalities wherein at least one dataset modality comprises a chemical structure dataset; training, using an encoder function, a foundation model by applying contrastive learning to the multimodal dataset, thereby embedding each dataset modality in a shared representation space and creating a learned representation of each dataset modality; and training the predictive model based on the learned representation of the chemical structure dataset modality and a classifier function that assigns modulation status for the biological target.
2. The method of claim 1 , wherein the biological target is a biological pathway, a transcription factor or a gene.
3. The method of claim 1, wherein the plurality of distinct dataset modalities comprises at least one of pharmacological data, bioinformatics data, computational data, or cell biological data.
4. The method of claim 3, wherein the plurality of distinct dataset modalities comprises molecular structures, bioactivity assays, compound docking scores, differential gene expression following gene knockdown, differential gene expression following compound treatment, scores for simulation or inhibition of multiple biological pathways, transcription factors or genes derived from external computational methods applied to gene expression data, or cell morphological images following compound treatment datasets.
5. The method of claim 1, wherein contrastive learning comprises one or more of Contrastive Multiview Coding, geometric multimodal contrastive learning, InfoNCE, or multimodal contrastive learning.
6. The method of claim 1 , wherein the predictive model comprises one of a neural network, support vector machine learning, a random forest model, a decision tree model, or a logistic regression model.
7. The model of claim 6, wherein the predictive model comprises a graph neural network.
8. The method of claim 1, wherein the shared representation space denotes a degree of similarity among the dataset modalities.
9. The method of claim 1 , wherein the modulation status for the biological target comprises stimulation, inhibition, or no change of a biological target.
10. The method of claim 1, wherein the modulation status for the biological target is represented by yes / no or a probability score.
11. A method of predicting modulation of a biological target as a result of an application of an agent to a predictive model, the method comprising: providing a multimodal dataset comprising a plurality of distinct dataset modalities wherein at least one dataset modality comprises a chemical structure dataset; training, using an encoder function, a foundation model by applying contrastive learning to the multimodal dataset, thereby embedding each dataset modality in a shared representation space and creating a learned representation of each dataset modality; training a predictive model based on the learned representation of the chemical structure dataset modality and a classifier function that assigns modulation status for the biological target; and introducing the agent to the predictive model to predict modulation of the biological target.
12. The method of claim 11, wherein the biological target is a biological pathway, a transcription factor or a gene.
13. The method of claim 11, wherein the plurality of distinct dataset modalities comprises at least one of pharmacological data, bioinformatics data, computational data, or cell biological data.
14. The method of claim 13, wherein the plurality of distinct dataset modalities comprises molecular structures, bioactivity assays, compound docking scores, differential gene expression following gene knockdown, differential gene expression following compound treatment, scores for simulation or inhibition of multiple biologicalpathways or transcription factors derived from external computational methods applied to gene expression data, or cell morphological images following compound treatment datasets.
15. The method of claim 11, wherein contrastive learning comprises one or more of Contrastive Multiview Coding, geometric multimodal contrastive learning, InfoNCE, or multimodal contrastive learning.
16. The method of claim 11, wherein the predictive model comprises one of a neural network, support vector machine learning, a random forest model, a decision tree model, or a logistic regression model.
17. The model of claim 16, wherein the predictive model comprises a graph neural network.
18. The method of claim 11, wherein the shared representation space denotes a degree of similarity among the dataset modalities.
19. The method of claim 11, wherein the modulation of a biological target comprises stimulation, inhibition, or no change of a biological target.
20. The method of claim 11 , wherein the modulation of a biological target is represented by yes / no or a probability score.
21. The method of claim 11, wherein the agent is comprised of a chemical structure.
22. The method of claims 21, wherein the agent is comprised of a small molecule.