Manifold learning of chemical perturbations and genetic responsome for de novo drug discovery

Manifold learning with reduced-rank Sparse Gaussian Process and Spectral Mixture kernel functions addresses the challenges of drug discovery by accurately predicting molecular and genetic interactions, enhancing drug candidate prediction and reducing computational complexity.

WO2025226823A1PCT designated stage Publication Date: 2025-10-30PURDUE RES FOUND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/025981
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-23
Filing Date
2025-04-23
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Current drug discovery methods are costly and time-consuming, particularly for complex diseases like Alzheimer's, due to the inability to consider multiplexed protein-protein interactions, and existing machine learning approaches fail to reliably predict molecular interactions and gene expressions, leading to late-stage failures and side effects.

Method used

A method involving manifold learning to encode molecular and genetic interactions using reduced-rank Sparse Gaussian Process with Spectral Mixture kernel functions, transforming molecular and gene interaction networks into two-dimensional representations for use in deep learning models to predict elicited differential gene expression.

Benefits of technology

This approach allows for accurate prediction of drug candidates by capturing non-linear molecular interactions, reducing computational complexity and ensuring manifold topology is respected, thereby improving drug discovery efficiency and reducing late-stage failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025025981_30102025_PF_FP_ABST
    Figure US2025025981_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are computational methods, systems, and devices for encoding electronic quantities from a molecular surface to a molecular representation for ML / DL-leveraged prediction of elicited differential gene expression in a particular cellular context and de novo design in deep learning applications.
Need to check novelty before this filing date? Find Prior Art

Description

Manifold Learning of Chemical Perturbations and Genetic Responsome for De Novo Drug DiscoveryCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. provisional patent application no. 63 / 637,747, filed April 23, 2024.TECHNICAL FIELD

[0002] The present disclosure relates to computational methods, systems, and devices for encoding electronic quantities from a molecular surface to a molecular representation for the prediction of elicited differential gene expression in a particular cellular context and de novo design in deep learning applications.BACKGROUND

[0003] The current paradigm of developing a medicine is a costly (> SI bn) and agonizingly long process (> 10 yr), which becomes even more dismal when finding a cure for complex diseases such as, e.g., Alzheimer's Disease and Alzheimer's Disease Related Dementias (AD / ADRD). The practice lies in a linear, progressive approach primarily centering on identifying a single protein target and subsequently conducting ligand screening against the target before a lead compound is tested in pre- and clinical studies. The inability to fully consider the disease’s multiplexed protein-protein interactions (PPIs) at the beginning of drug discovery- likely results in late-stage failure because of off-target binding and side effects. Tire recent rise in artificial intelligence (Al) is poised to reshape the “one disease, one target, and one drug” convention.

[0004] In treating disease, despite numerous efforts in developing and using machine / deep learning (ML / DL) for drug design, few drug candidates have been successfully predicted and tested. With respect to AD / ADRD, for example, what contributes to the difficulty lies in (1) most of the current AD / ADRD-related databases contain little chemical information on drug treatment that is collected consistently under identical experimental conditions; (2) current approaches to representing molecules and gene attributes for ML / DL are incompetent to reliably predict the causality between a molecule and its elicited differential gene expressions (DGEs) and other pathological / neurological behaviors.

[0005] The latter is mainly because most chemical information of molecular interactions resides in a non-linear manifold space. Moreover, conventional molecular representations are based on the structural description of a molecule, making it challenging to infer molecular interactions and properties that are determined by the molecule's electronic structure and attributes. The structural representations are also noisy — meaning that only a tiny fraction of such (mathematic) representations are chemically correct — causing the generative prediction of molecules based on desired biological / therapeutic outcomes to be ''hallucinating."

[0006] Moreover, a major technical hurdle to advancing ML / DL in drag research is that most chemical information, manifested through molecular interactions, resides on high- dimensional manifolds rather than in a (low-dimensional) Euclidean space. At least two challenges arise: (1) capturing chemical information saliently by enabling mathematical representations to effectively fend off the “Curse of Dimensionality” (COD), and (2) ensuring manifold topography is properly respected by ML. Traditional molecular descriptors are prone to COD, and structure-based fingerprints carry little chemistry of molecular interactions. It is even more challenging to featurize networked interactions such as gene regulatory networks (GRNs). Because molecular or chemical representation resides on a manifold, it is imperative to utilize manifold metrics and capture multiplexed, non-linear relationships with molecular properties or interactions. Importantly for generative Al, a molecule or networked interactions must be encoded in a mathematically and chemically differentiable form.

[0007] As but one example, innovative manifold learning methods can be used to capture the underlying chemistry of molecular interactions and develop generative Al models for predicting AD / ADRD drag candidates using the experimental training data collected on a zebrafish AD model. These efforts are driven by the perturbation theory, that is, by introducing a small chemical perturbation to a biological system, the response allows the exploration of deviations from the homeostatic state and conceivably design a chemical (i.e., perturbagen) to alter the diseased state.SUMMARY

[0008] Presently described are methods, systems and devices for predicting a perturbagen’s elicited cellular differential gene expression. The present approach can be accomplished through a variety of steps, including, for example: (1) the two-dimensional embedding or kemelization of a molecule’s electronic patterns; (2) the two-dimensionalkernelization of a molecule’s elicited differential gene expression (DGE) pattern over a gene interaction network and corresponding biomarkers; and (3) the manifold regression between the two manifolds.

[0009] The present disclosure involves, in one embodiment, a method for creating a representation of a molecule as a chemically authentic and dimensionally reduced feature for computing molecular interactions and pertinent properties. A plurality of molecules are measured by observation of electronic patterns on a three-dimensional molecular surface. According to methods herein, a manifold kernelization of the observed electronic patterns can be created, the kernelization having a two-dimensional representation of the observed electronic patterns. The two-dimensional representation can be associated with an elicited differential gene expression (DGE) in a particular cellular context and the corresponding biomarkers.

[0010] In one embodiment, a database of molecular representations includes a plurality of data structures. The data structures can include manifold kernelization of the observed electronic patterns on a molecular surface, where at least one of the observed electronic patterns being dimensionally reduced via manifold kernelization, with an association of the manifold kernelization with a particular cellular context’s DGE and corresponding biomarkers.

[0011] In yet a further embodiment, a method of determining the elicited DGE in a particular cellular context of an observed molecule’s electronic patterns using a database is disclosed, according to which a plurality of data structures representing molecular surfaces and associated DGE can be provided. A quantum calculation of the observed molecule electronic patterns can be created, along with a manifold kernelization of molecular surface (MKMS) or other similar representation of the observed molecule using dimensionality reduction. The kernel representation of the observed molecule can be used as an input with a neural network utilizing the database to identify gene expression changes elicited by the observed molecule.

[0012] In example embodiments, a non-transitory computer-readable medium CRM) can encode instructions that, when executed by at least one processor, cause the at least one processor to: (1) receive a dataset of electronic quantities across a manifold topology of a molecule of interest; (2) perform manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; (3) utilize a multi-variate regression to predict an elicited DGE covariance matrix in a particularcellular context; and (4) export the covariance matrix as a symmetric semi-positive definite matrix. By way of example, multi-variate regression analysis can include partial least square or NIP AL. In certain examples, the elicited DGE manifold can be determined through kernelization of DGE data in a particular cellular context across a gene interaction network, and an array of biomarkers of the same particular cellular context. In the same or other examples, the gene interaction network can include cellular protein-protein interactions.

[0013] The CRM herein can perform the kernelization process using a reduced-rank Sparse Gaussian Process (SGP) with Spectral Mixture (SM) as the kernel function. In particular, kernel function can be a covariance function of the spectral solution of a manifold identified through an eigen decomposition of the Graph Laplacian of the gene interaction network. Moreover, the SGP can utilize a set of landmark points to approximate the shape of the underlying manifold. In one example, the landmark points can be selected as those points that exhibit the highest amount of uncertainty or variance during the SPG regression. In various embodiments, the encoded quantum information is determined through MKMS of a dataset of molecular surfaces.

[0014] In another embodiment, a method for creating a representation of the genetic changes in a population of cells in a particular context is disclosed for a biologically authentic and dimensionally-reduced feature for computing perturbagen elicited gene expression alterations. Genes, or their various products, interact with other molecules to perform cellular functions. These interactions can be measured through a variety' of different methods and can be expressed as a graph, where the nodes represent different genes, and the edges represent the presence of a molecular interaction. A manifold kernelization of the observed gene interactions is created, the kernelization having a two-dimensional representation of the observed genetic interactions.

[0015] To that end, in one aspect, a one non- transitory computer-readable medium is disclosed that includes instructions that, when executed by at least one processor, cause the at least one processor to: (1) receive a dataset consisting of a gene interaction network and DGE data of a cellular context of interest; (2) perform manifold regression on the gene interaction network and the DGE data to encode genetic alterations of a particular cellular context in a covariant matrix (kernel), wherein the covariant matrix capture mutual relationships among the genetic alterations and the manifold topology; and (3) export the covariance matrix as asymmetric semi-positive definite matrix for use as an input to a neural network model configured to utilize the matrix for predicting molecular properties.

[0016] A cell’s gene regulatory network (GRN) manifold according to embodiments herein can be transformed using for example, reduced-rank SGP with SM as the kernel function. According to examples, the SGP utilizes a set of landmark points chosen, manually or numerically, as the points with the highest variance observed during the SPG. In the same or other examples, the covariance function of the spectral solution of a manifold can be identified through an eigen decomposition of the Graph Laplacian of the underlying gene interaction network. In the same or yet other examples, the hyperparameters used in the reduced-rank SPG, particularly those associated with the means of the Gaussian functions can be set to 0 to improve convergence. The kernel matrix can also be modified via eigen value reduction to a number near 0, generally between 0.001 and 0.000001, to allow for matrix or other value optimization.

[0017] According to embodiments, dimensionality reduction can be used to project the data from a Euclidean tangent space or other similar space to a Riemannian manifold or other similar space. Projection can also be used to project the data from a Riemannian manifold or other similar space to a Euclidean tangent or other similar space.

[0018] Any neural networks] s) (or other machine or deep learning models) configured and / or trained to receive a single positive definite (SPD) matrix or derivates thereof as input to predict an embedded molecule’s elicited genetic alterations can be used consistent with the present disclosure, including in one example, a trained SPDNet Attention model. According to the examples, the model of embodiments of the invention includes a Deep Sets-based Graph and Self -Attention Network for evaluating differential gene expression in a particular cellular context. In one embodiment the algorithm can use -21,000 molecules and their elicited DGE in U20S cells from National Institute of Health (NIH) LINCS (Library of Integrated Network- Based Cellular Signatures) program.

[0019] The present method of generating unique elicited DGE and biomarkers in a particular cellular context from the quantum surfaces of an input molecule can be performed using a trained ML / DL model. In one example, this method includes the steps of: (1) receiving a dataset of electronic quantities across a manifold topology of a molecule of interest; (2) performing manifold regression on the electronic quantities to encode quantum information of the molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutualrelationships among the electronic quantities and the manifold topology; (3) utilizes a multi- variate regression analysis from the encoded quantum information to produce it’s elicited DGE manifold; and (4) exports a covariant matrix (kernel) encoding of the elicited DGE manifold.

[0020] The present method of training a ML / DL model to produce unique elicited DGE and biomarkers in a particular cellular context form quantum sur faces of an input molecule. In one example, this method includes the steps of: (1 ) receiving a dataset of electronic quantities across a manifold topology of a molecule of interest; (2) performing manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; (3) receiving a dataset of a molecule’s elicited DGE and corresponding biomarkers in a particular cellular context and a gene interaction network; (4) performing manifold regression on the DGE data across the gene interaction network to encode molecular interaction information from tire cellular context in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the genetic alterations and the manifold topology; (5) performs a multi-variate regression analysis between the molecular electronic manifold and the DGE manifold.

[0021] In a number of embodiments, systems are described for performing a method of manifold learning of chemical perturbations and genetic responsome via a trained ML / DL model, the method comprising: (1) receiving a dataset of molecular surfaces; (2) performing manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; (3) utilizing a multi-variate regression analysis from the encoded quantum information to produce it’s elicited DGE manifold; (4) exporting a covariant matrix (kernel) encoding of the elicited DGE manifold.

[0022] In one particular example, the method can include: (1) receiving a dataset of electronic quantities across a manifold topology of a molecule of interest; (20 performing manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; (3) receiving a dataset of a molecule’s elicited DGE and corresponding biomarkers in a particular cellular' context and a gene interaction network; (4) performing manifold regression on the DGE data across the gene interactionnetwork to encode molecular interaction information from the cellular context in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the genetic alterations and the manifold topology; and (5) performing a multi-variate regression analysis between the molecular electronic manifold and the DGE manifold.

[0023] In examples, multi-variate regression analysis can include partial least square or NIPAL. In the same or other examples, the elicited DGE manifold is determined through kernelization of DGE data in a particular cellular context across a gene interaction network, and an array of biomarkers of the same particular cellular context. In the same or still other examples, the gene interaction network where the gene interaction network can include cellular protein- protein interactions.BRIEF DESCRIPTION OF THE FIGURES

[0024] Fig. 1 illustrates a schematic diagram of a system environment implementing a cheminformatics subsystem.

[0025] Fig. 2 is a flow diagram of an example cheminformatics subsystem process for perturbagen manifold learning.

[0026] Fig. 3. is an illustration of the experimental procedure for generating data for perturbagen manifold learning.

[0027] Fig. 4. demonstrates how experimental chemical and biological data is transformed into a ML / DL compatible format.

[0028] Fig 5. presents ESP and f2MKMS kernels of astemizol and aconitine being mapped onto a chemical perturbagen space.

[0029] Fig 6. presents different DGE networks, represented by then kernels being mapped onto a transcriptomic response space.

[0030] Fig 7. shows the mathematical foundations by which a perturbagen’s manifold mapping can be used to predict that perturbagen’s transcriptomic effects.

[0031] Fig. 8. shows a diagram of a computer device.DETAILED DESCRIPTION

[0032] The following discussion is presented to enable any person skilled in the art to make and use the technology disclosed and is provided in the context of a particular applicationand its requirements as well as a particular system environment and its requirements. Various modifications to the disclosed implementations will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other implementations and applications without departing from the spirit and scope of the technology disclosed. Thus, the technology disclosed is not intended to be limited to the implementation shown but it should be accorded the widest scope consistent with the principles and features disclosed herein.

[0033] Examples of programming languages, codes and / or data libraries, and / or operating environments for use with the present technology include Python, Numpy, R, Java, Javascript, C#, C++, Julia, Shell, Go, TypeScript, and Scala.

[0034] Notably, while the disclosure provides certain examples in the context of AD / ADRD diseases and drug candidates, any disease and drug candidate for which a drug molecule’s electronic patterns and / or the molecule’s elicited DGE is obtainable may form the basis of analysis as described herein.

[0035] Fig. 1 illustrates a schematic diagram of a system environment 100 (or “environment”) to (a) receive a dataset of electronic quantities across a manifold topology of a molecule of interest; (b) create a two-dimensional embedding or kernelization of a molecule’s electronic patterns; (c) utilize a multi-variate regression to predict an elicited DGE covariance matrix in a particular cellular context; and (d) export the covariance matrix (kernel) as a symmetric semi-positive definite matrix, which may be connected to one or more server device(s) 102 and / or one or more client (or user) device(s) 108 via a network 114. As shown in Fig. 1, the sen- er device(s) 102 and the client device(s) 108 may communicate with each other via the network 114 or directly. The network 114 may comprise any suitable network over which computing devices may communicate. The network 114 may include a wired and / or wireless communication network. Example wireless communication signals using one or more wireless communication protocols, such as a cellular communication protocol, a wireless local area network (WLAN) or WIFI communication protocol, and / or another wireless communication protocol. In addition, or in the alternative to may bypass the network 114 and may communicate directly with one another.

[0036] As further illustrated in Fig. 1, the environment 100 may include storage 106. The storage 106 may store information for being accessed by the devices in the environment 100. The serve device(s) 102 and the client device(s) 108 may communicate with the storage 106 (e.g.,directly or via the network 1 14) to store and / or access information, including, e.g., any of the various libraries.

[0037] As further indicated by Fig. 1, the sen- er device(s) 102 may generate, receive, analyze, store, and / or transmit digital data, such as datasets of electronic quantities across a manifold topology of a molecule of interest. The server device(s) 102 may communicate with the client device(s) 108. In particular, the server device(s) 102 may send data to the client device(s) 108, including molecules, and DGE kernels, and the server device(s) 102 may receive input from users via client device(s) 108.

[0038] The server device(s) 102 may comprise a distributed collection of servers where the server device(s) 102 include a number of server devices distributed across the network 114 and located in the same or different physical locations. Further, the server device(s) 102 may comprise a content server, an application server, a communications server, a web-hosting server, or another type of server.

[0039] As further shown in Fig. 1, the server device(s) 102 and / or the sequencing device(s) 214 may include a cheminformatics subsystem 104.

[0040] The cheminformatics subsystem 104 may also include software and / or hardware utilized by the server device(s) 102 for receiving datasets of electronic quantities across a manifold topology of a molecule of interest; creating a two-dimensional embedding or kernelization of a molecule’s electronic patterns, utilizing a multi-variate regression to predict an elicited DGE kernel in a particular cellular context; and export the covariance matrix as a symmetric semi-positive definite matrix.

[0041] The cheminformatics subsystem 104 may be included in a single server device 102 or may be distributed across multiple devices. The cheminformatics subsystem 104 may include a sequencing systems that spans multiple layers of software and / of hardware for servicing requests at cheminformatics subsystem 104. The cheminformatics subsystem 104 may support drug optimization and filtering sendees for clients or other users.

[0042] Each client device 108 may generate, store, receive, and / or send digital data. For example, the client device 108 may communicate with the server device(s) 102 to receive one or more datafile comprising output from the cheminformatics subsystem 104, including datafiles comprising datasets of parent derivative molecules and / or drug candidate molecules.Furthermore, the client device 108 may present or display information pertaining to a dataset ofelectronic quantities across a manifold topology of a molecule of interest and / or DGE kernels within a graphical user interface to a human user associated with the client device 108. The client device(s) 108 may comprise various types of client devices. In examples, the client device 108 may include non-mobile devices, such as desktop computers or servers, or other types of client devices. In other examples, the client device 108 may include mobile devices, such as laptops, tablets, mobile telephones, or smartphones.

[0043] As further illustrated in Fig. 1, each client device 108 may include a client subsystem 110. The client subsystem 110 may include software and / or hardware utilized by the client device 108 for processing datasets of electronic quantities across a manifold topology of a molecule of interest. The client subsystem 110 may span multiple layers of software and / or hardware. The client subsystem 110 may be included in a single client device 108 or may be distributed across multiple client devices 108.

[0044] Client processes may be operated on one or more the client device 108 and / or the server device 102 for requesting drug optimization and filtering services from the cheminformatics subsystem 104. For example, client processes executing on any of the client device(s) 108 and / or server device(s) 102 may transmit requests for performing drug optimization and / or filtering services for particular biologically active molecules of interest. The cheminformatics subsystem 104 may load and / or execute different bitstreams to perform various types of analysis to support requests from the client processes.

[0045] Referring to Fig. 2, in one aspect, a cheminformatic subsystem 104 may include one or more engines for implementing one or more methods or procedures. For example, the cheminformatics subsystem 104 may include a receiving engine 202, a molecular kemelization engine 204, a multi-variate regression 206, and an exporting engine 208.

[0046] In examples, the cheminformatic subsystem 104 may cause receiving engine 202 to receive of electronic quantities across a manifold topology of a molecule of interest, e.g. molecules from the LINCS database. The chemoinformatic subsystem 104 may then cause the kemelization engine 204 to produce a unique two-dimensional representation of the observed electronic patterns. Next, the cheminformatic subsystem 104 may cause the regression engine 206 to perform multi-variate regression to predict an elicited DGE kernel. Next, the cheminformatic subsystem 104 may cause the exporting engine 208 to export the kernel encoding of the elicited DGE manifold.

[0047] In one example, the receiving engine 202 of cheminformatic subsystem 104 may receive electronic quantities for a molecule of interest such as electrostatic potential (ESP) and / or Fukui function quantities. In one example, the kernelization engine 204 of cheminformatic subsystem 104 may perform kernelization utilizing manifold regression on the electronic quantities to generate the molecular kernel, such as MKMS. MKMS is preferred for this process as it allows dimensional reduction of the data from 3 to 2, significantly reducing the computational complexity of the data for DL / ML applications, while still maintaining the molecule’s structural quantum and electronic information. In certain embodiments MEMS could be used by the kernelization engine 204 of the cheminformatics subsystem 104. In one example, the multi- variate regression engine 206 of cheminformatic subsystem 104 may utilize partial least squares or NIP AL to perform the multi-variate regression to predict the elicited DGE kernel. In one example, the exporting engine 208 of cheminformatic subsystem 104 may export an inputted perturbagen’s predicted elicited DGE kernel in a particular cellular context, as a two- dimensional covariance matrix (kernel) or other representation thereof.

[0048] In example embodiments, the cheminformatic subsystem 104 may comprise one or more non-transitory computer-readable media (CRM) comprising instruction that, when executed by at least one processor, cause the at least one processor to: (1), via receiving engine 202, receive a dataset of electronic quantities across a manifold topology of a molecule of interest; (2), via kernelization engine 204, perform manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; (3), multi-variate regression engine 206, utilize a multi-variate regression to predict an elicited DGE covariance matrix in a particular cellular context; and (4), via exporting engine 208, export the covariance matrix as a symmetric semi-positive definite matrix. In examples, multi-variate regression analysis includes partial least square or NIP AL. In the same or other examples, the elicited DGE manifold is determined through kernelization of DGE data in a particular cellular context across a gene interaction network, and an array of biomarkers of the same particular cellular context. In the same or still other examples, the gene interaction network where the gene interaction network is comprised of cellular protein-protein interactions.

[0049] The CRM may perform the kernelization process using a reduced-rank Sparse Gaussian Process (SGP) with Spectral Mixture (SM) as the kernel function. In particular, kernelfunction may be a covariance function of the spectral solution of a manifold identified through an eigen decomposition of the Graph Laplacian of the gene interaction network. Moreover, for example, the SGP may utilize a set of landmark points to approximate the shape of the underlying manifold. In one example, the landmark points are selected as those points that exhibit the highest amount of uncertainty or variance during the SPG regression. In various embodiments, the encoded quantum information is determined through MKMS of a dataset of molecular surfaces.

[0050] In a number of embodiments, the cheminformatic subsystem 104 performs a method (or directs performance of a method) by which a trained ML / DL model: (I), via receiving engine 202, receives a dataset of molecular surfaces; (2), via kemelization engine 204, performs manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; (3), via multi-variate regression engine 206, utilizes a multi-variate regression analysis from the encoded quantum information produce it’s elicited DGE manifold; (4), via exporting engine 208, exports a covariant matrix (kernel) encoding of the elicited DGE manifold.

[0051] In one particular example, the method may include: (1), via receiving engine 202, receiving a dataset of electronic quantities across a manifold topology of a molecule of interest;(2), via kemelization engine 204, performing manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology;(3), via receiving engine 202, receiving a dataset of a molecule’s elicited DGE and corresponding biomarkers in a particular cellular context and a gene interaction network; (4), via kemelization engine 204, performing manifold regression on the DGE data across the gene interaction network to encode molecular interaction information from the cellular context in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the genetic alterations and the manifold topology; (5), via multi-variate regression engine 206, performing a multi-variate regression analysis between the molecular electronic manifold and the DGE manifold. In certain embodiments the covariant matrices of steps (2) and (4) are generated in a non ML / DL processing steps as dimensionally reduced inputs into a trained ML / DL model,in the trained ML / DL model causes the multi-variate regression engine 206 to perform multi- variate regression analysis between the molecular electronic manifold and the DGE manifold.

[0052] Referring back to Fig. 1, in certain embodiments, a cloud computing platform 112 may include one or more engines for implementing one or more methods or procedures herein. In the illustrated example, the cloud computing platform 112 may host an artificial neural network (ANN) for implementing one or more procedures herein, e.g., on a trained ML / DL model. By way of example, an ML / DL model programed on the ANN may include the multi- variate regression engine 206 for performing a multi-variate regression analysis between the molecular electronic manifold and the DGE manifold.

[0053] In examples, multi-variate regression analysis includes partial least square or NIP AL. In the same or other examples, the elicited DGE manifold is determined through kernelization of DGE data in a particular cellular context across a gene interaction network, and an array of biomarkers of the same particular cellular context. In the same or still other examples, the gene interaction network where the gene interaction network is comprised of cellular' protein- protein interactions.

[0054] The kernelization process according to certain examples may be performed using a reduced-rank Sparse Gaussian Process (SGP) with Spectral Mixture (SM) as the kernel function. In particular, kernel function may be a covariance function of the spectral solution of a manifold identified through an eigen decomposition of the Graph Laplacian of the gene interaction network. Moreover, for example, the SGP may utilize a set of landmark points to approximate the shape of the underlying manifold. In one example, the landmark points are selected as those points that exhibit the highest amount of uncertainty or variance during the SPG regression. In various embodiments, the encoded quantum information is determined through MKMS of a dataset of molecular surfaces.

[0055] Multi-omics has become a defining force in studying disease etiology and therapeutic mechanisms of a molecular' agent, resulting in large amounts of omics data to be exploited as illustrated in Fig. 3. By fully considering the networked protein-protein interactions (PPIs) in a targeted cell at the very beginning of drug discovery, it is possible to design a molecule that concertedly effects regulation only on diseased targets while avoiding the homeostatic machinery, to lessen late-stage failures in human trials. This approach is currently greatly stymied because molecular interactions, like those between a perturbagen and protein andprotein and other proteins, reside on high-dimensional manifolds rather than a (low-dimensional) Euclidean space as illustrated in Fig. 4. In mathematics, a manifold signifies a space with unique topology in which the local neighborhood of a manifold point is approximated Euclidean; importantly the, the local Euclidean metrics are collectively associated and defined globally by the topological metrics. For instance, a gene interaction network mathematically is a 2-D manifold in 3-D space, and a surface point bears a 2-D tangent plane that characterizes its neighborhood metrics.

[0056] Elicited differential gene expression (DGE) may be regarded as cellular' states distributed on a manifold connected to a condition of interest. Such conditions may include nominal (healthy), diseased, or disordered. Moreover, most of the current cellular representations for ML / DL stem from individual genetic values, but they carry very little information about the molecular interactions, making data-driven drug research highly challenging. It is thus critical for developing ML / DL models to directly utilize genetic information, in particular, a cell’s molecular interactions, to fend off the “curse of dimensionality” (COD) and to ensure manifold topology is correctly and fully observed for ML methods.

[0057] The present disclose utilizes kernel learning to directly capture the information of molecular' interactions contained in a gene interaction network without resorting to a manifold embedding process. The essence of the present approach is to conduct Gaussian Process (GP) regression on a manifold. Because of its non-parametric nature, GP utilizes the covar iances among training or existing data points to predict the distribution function at a new point (both mean and variance). The covariance matrix, or kernel, signifies the mutual relationships among the (training) data points. Various kernel functions are devised to define how two data points are related typically based on their distance and (trainable) hyperparameters. Sparse GP (SGP) utilizes a fraction of training or inducing points to fit all data points optimizing the hyperparameters of the covariance matrix of the inducing points. The resultant covariance or kernel matrix thus encodes data relationships and their mutual influences among the inducing points, as well as the connections with all the training points.

[0058] Testing disclosed herein demonstrated that using dozens of spectral mixture (SM) kernel in SGP resulted in highly expressive kernels for elicited DGE activity across a gene interaction network. Notably, if applying SGP directly on elicited DGE activity across a gene interaction network, it would render the covariance matrix not semi-positive definite even ifgeodesic distances are used in kernel calculation. In that regard, a reduced-rank GP approach was adopted along with calculated covariances with eigensolutions of the graph Lapalacian, resulting in kernels representing both DGE and the topology of the gene interaction network.

[0059] An artificial neural network (ANN) model was developed to predict elicited DGE and utilize MKMS as a molecular representation in DE, like those seen in Fig. 5. As both MKMS and elicited DGE kernels are symmetric positive definite (SPD), SPDNet was adopted to maintain the underlying Riemannian topology by SPD matrices in data training. Self-attention in the ANN architecture was further utilized to ensure permutation-invariance in processing the kernel input. The supervised manifold learning model, dubbed SPDNet Attention, performed well using data from NIH LINCS. The overall technical framework herein includes three components: (1) Unsupervised manifold kernelization of molecular surface (which we have filed) and DGE; (2) supervised deep learning to handle MKMS or DGE kernels as ANN input; and (3) supervised manifold learning (NIPAL) to link MKMS kernels with DGE kernels

[0060] The manifold learning methodology as described herein is advanced by directly running kernel learning on a DGE data across a gene interaction network. According to this approach, SGP is performed with proper kernel functions on the DGE data across a gene interaction network or a surface manifold. SGP is an approximate GP approach given by Eq. 1. where k is a kernel function between data points X (e.g., radial basis function or RBF).

[0061] In Eq. 1, the complex 3-dimensional surface can be mathematically defined as a joint distribution (p(f(X)|X)), of a Gaussian function f(X) and X, where X are data points obtained from the surface. This covariance function can also be modeled as a Gaussian function (N( u, K)) of the mean function of the surface (p), and a covariance function of the surface (K). With sufficiently complex surfaces these functions cannot be exactly determined but can be estimated with a SGP, involving X*, a set of new data points derived from the surface. The mean function can be approximated as a product of K*TK-lf(X), where K* is K(X,X*). The covariance function can be approximated as the difference of K** and K*T K-l K*, where K** is K(X*, X*).

[0062] The hyperparameters used in the kernel functions are optimized iteratively as the SGP regression is performed. Spectral Mixture (SM) is used as the kernel function, and is run directly on the embedding points in Euclidean space. The SM kernel is critical to this process, asit models the spectral density of the covariance function as a mixture of Gaussians and contains more trainable hyperparameters than alternative kernel functions like RBF or Mateni, as shown in Eq. 2. where T is the Euclidean distance between the true points and the selected points (xi - xj). T p is the projected distance on dimension p of with P, maintaining the dimensionality of x. q is an index of Q, the number of SM mixtures used in the analysis. Wq is the weight of the qth SM. Uq and mq are hyperparameters defining the variances and means of the Gaussian function of the SM mixtures.

[0063] Utilizing a recent development of reduced-rank GP, the covariance function of a spectral solution manifold can be solved using SM as the spectral density to generate GP kernels without going through the embedding process. The Graph Laplacian is calculated from a triangular mesh of the molecular surface and then is eigen-decomposed. By selecting the eigenvectors of 4 of the 5 smallest eigenvalues (excluding the smallest), an SM kernel can be created, as shown in Eq. 3.Where ki is the ith eigen value of the surface graph’s Laplacian, M is the number of SM mixtures, and wm is the weight hyperparameter of a respective SM Gaussian. Em and mm are hyperparameters of variances and means of the Gaussian functions. These values may be kept to 0 to improve convergence to singular values.

[0064] A kernel or covariance matrix is symmetric semi-positive definite; in practice, such a matrix is regularized by adding a small number (e.g., 1.0A-7) to its eigenvalues, becoming SPD. Importantly, SPD matrices reside in a Riemannian manifold, and the topology of the Riemannian manifold needs to be respected by ML / DL models. SPDNet was therefore adopted, where three neural network operators, BiMap, ReEig, and LogEig, aim to preserve the Riemannian manifold during training. In particular, BiMap layer achieves dimensionality reduction of an input SPD, ReEig regulates the learning (by replacing the smallest eigenvalues of an SPD with a predetermined cutoff parameter, generally, 0.00001 to 0.001), and logEig projects an SPD from the Riemannian manifold to its Euclidean tangent space.

[0065] SPDNet Attention as described herein was developed to ensure permutation invariance of gene expression vectors. In other words, exchanging i and j rows (and corresponding columns) of a DGE kernel should not affect the prediction outcome. The attention algorithm follows the essence of self-attention or Transformer. Two SPDNets may be utilized to generate the “query'” and “key” of the input gene expression vector, (i.e., DGE values of individual genes used as inducing points of the GRN). The query and key then multiply to form the weighting matrix and then mask the input vector by matrix multiplication. The weighted electronic vectors are stacked together, forming a feature matrix further processed by DeepSets layers.

[0066] Kernelization was performed on DGE data from NIH LINCS for astemizol and aconitine in U2OS cells. Fig. 6 highlights the manifold kernelizations as well as the results of over representation analysis (ORA) of the landmark genes from the kernel. The illustrated DGE kernel of astemizol was calculated with 24 SM functions. When directly conducted on DGE data across a gene interaction network, manifold kernelization was initiated by spectrum analysis of the manifold graph Laplacian followed by SGP regression with SM kernel functions. The exemplified aconitine kernel was calculated with the top 2400 spectral eigenvectors and 24 SM functions. The two DGE kernels share very little in common, despite both utilizing the same gene interaction network. The recovery' of distinct ORA results further demonstrates the robustness and performance of this technique.

[0067] Fig. 7 shows the mathematical basis by which a multi-variate regression could be performed to map a molecular kernel to its elicited DGE kernel in a particular cellular context.

[0068] A one non-transitory computer-readable medium (CRM) is disclosed that includes instruction that, when executed by at least one processor, cause the at least one processor to generate a unique elicited DGE and biomarkers in a particular context from the quantum surfaces of input molecule. In one example the CRM may cause the processor to: (1) receive a dataset of electronic quantities across a manifold topology of a molecule of interest; (2) perform manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; (3) utilize a multi-variate regression to predict an elicited DGE covariance matrix in a particular cellular context; and (4) export the covariance matrix as a symmetric semi-positive definite matrix.

[0069] A particular cellular context’s GRN according to embodiments herein can be transformed using, for example reduced-rank Sparse Gaussian Process (SGP) and Spectral Mixture (SM) as the kernel function. According to examples, the SPG utilizes a set of inducing points chosen manually or numerically as the points which exhibited the highest uncertainty, variance, during the regression process. In the same or other examples, the covariance function of the spectral solution of a manifold can be identified through an eigen decomposition of the Graph Laplacian of a gene interaction network. In the same or other examples, the hyperparameters used in the reduced-rank SGP, particularly those associated with the means of the Gaussian functions can be set to 0, to improve convergence. The kernel matrix can also be modified via eigen value reduction to a number near 0, generally between 0.001 and 0.000001, to allow for matrix or other value optimization.

[0070] According to embodiment, dimensionality reduction may be used to project the data from a Euclidean tangent space or other similar space to a Riemannian manifold or other similar space. Projection can also be used to project the data from a Riemannian manifold or other similar space to a Euclidean tangent or other similar space.

[0071] Any neural network(s) (or other machine or deep learning models) configured and / or trained to receive a single positive definite (SPD) matrix or derivates thereof as input to predict an embedded molecule’s chemical properties can be used consistent with the present disclosure, including, in one example, a trained SPDNet Attention model. According to examples, the model of embodiments of the invention involves a Deep Sets-based Graph and Self-Attention Network with input from a biological approach to predict a particular cellular context’s elicited DGE.

[0072] The present method of generating unique elicited DGE and biomarkers in a particular cellular context from the quantum surfaces of an input molecule can be performed using a trained ML / DL model. In one example, this method includes the steps of: (1) receiving a dataset of electronic quantities across a manifold topology of a molecule of interest; (2) performing manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; (3) utilizes a multi- variate regression analysis from the encoded quantum information to produce it’s elicited DGE manifold; and (4) exports a covariant matrix (kernel) encoding of the elicited DGE manifold.

[0073] The present method of training a ML / DL model to produce unique elicited DGE and biomarkers in a particular cellular context form quantum surfaces of an input molecule. In one example, this method includes the steps of: (1) receiving a dataset of electronic quantities across a manifold topology of a molecule of interest; (2 ) performing manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology': (3 ) receiving a dataset of a molecule’s elicited DGE and corresponding biomarkers in a particular cellular context and a gene interaction network; (4) performing manifold regression on the DGE data across the gene interaction network to encode molecular interaction information from the cellular context in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the genetic alterations and the manifold topology; (5) performs a multi-variate regression analysis between the molecular electronic manifold and the DGE manifold.

[0074] An example of Artificial Intelligence (Al) is intelligence as manifested by computer systems. Various examples of ANNs include Convolutional Neural Networks (CNNs) generally (e.g., any ANN having one or more layers performing convolution), as well as NNs having elements that include one or more CNNs and / or CNN-related elements (e.g., one or more convolutional layers), such as various implementations of Generative Adversarial Networks (GANs) generally, as well as various implementations of Conditional Generative Adversarial Networks (CGANs), cycle-consistent Generative Adversarial Networks (CycleGANs), and autoencoders. In various scenarios, any ANN having at least one convolutional layer is referred to as a CNN. The various examples of ANNs further include transformer-based ANNs generally (e.g., any ANN having one or more layers performing an attention operation such as a self- attention operation or any other type of attention operation), as well as ANNs having elements that include one or more transformers and / or transformer-related elements. The various examples of ANNs further include Recurrent Neural Networks (RNNs) generally (e.g., any ANN in which output from a previous step is provided as input to a current step and / or having hidden state), as well as ANNs having one or more elements related to recurrence. The various examples of ANNs further include graph neural networks and diffusion neural networks. The various examples of NNs further include MultiLayer Perceptron (MLP) neural networks. In some implementations, a GAN is implemented at least in part via one or more MLP elements.

[0075] Examples of elements of ANNs include layers, such as processing, activation, and pooling layers, as well as loss functions and objective functions. According to implementation, functionality corresponding to one or more processing layers, one or more pooling layers, and / or one or more an activation layers is included in layers of a ANN. According to implementation, some layers are organized as a processing layer followed by an activation layer and optionally the activation layer is followed by a pooling layer. For example, a processing layer produces layer results via an activation function followed by pooling. Additional example elements of NNs include batch normalization layers, regularization layers, and layers that implement dropout, as well as recurrent connections, residual connections, highway connections, peephole connections, and skip connections. Further additional example elements of ANNs include gates and gated memory units, such as Long Short-Term Memory (LSTM) blocks or Gated Recurrent Unit (GRU) blocks, as well as residual and / or attention blocks.

[0076] Examples of processing layers include convolutional layers generally, upsampling layers, downsampling layers, averaging layers, and padding layers. Examples of convolutional layers include ID convolutional layers, 2D convolutional layers, 3D convolutional layers, 4D convolutional layers, 5D convolutional layers, multi-dimensional convolutional layers, single channel convolutional layers, multi-channel convolutional layers, 1 x 1 convolutional layers, atrous convolutional layers, dilated convolutional layers, transpose convolutional layers, depthwise separable convolutional layers, pointwise convolutional layers, 1 x 1 convolutional layers, group convolutional layers, flattened convolutional layers, spatial convolutional layers, spatially separable convolutional layers, cross-channel convolutional layers, shuffled grouped convolutional layers, and pointwise grouped convolutional layers. Convolutional layers vary according to various convolutional layer parameters, for example, kernel size (e.g., field of view of the convolution), stride (e.g., step size of the kernel when traversing an image), padding (e.g., how sample borders are processed), and input and output channels. An example kernel size is 3x3 pixels for a 2D image. An example default stride is one. In various implementations, strides of one or more convolutional layers are larger than unity (e.g., two). A stride larger than unity is usable, for example, to reduce sizes of non-channel dimensions and / or downsampling. A first example of padding (sometimes referred to as ‘padded’) pads zero values around input boundaries of a convolution so that spatial input and output sizes are equal (e.g., a 5x5 2D input image is processed to a 5x5 2D output image). A second example of padding (sometimesreferred to as ‘unpadded’) includes no padding in a convolution so that spatial output size is smaller than input size (e.g., a 6x6 2D input image is processed to a 4x42D output image).

[0077] Example activation layers implement, e.g., non-linear functions, such as a rectifying linear unit function (sometimes referred to as ReLU), a leaky rectifying linear unit function (sometimes referred to as a leaky-ReLU), a parametric rectified linear unit (sometimes referred to as a PreLU), a Gaussian Error Linear Unit (GELU) function, a sigmoid linear unit function, a sigmoid shrinkage function, an SiL function, a Swish- 1 function, a Mish function, a Gaussian function, a softplus function, a maxout function, an Exponential Linear Unit (ELU) function, a Scaled Exponential Linear Unit (SELU) function, a logistic function, a sigmoid function, a soft step function, a softmax function, a Tangens hyperbolicus function, a tanh function, an arctan function, an ElliotSig / Softsign function, an Inverse Square Root Unit (ISRU) function, an Inverse Square Root Linear Unit (ISRLU) function, and a Square Nonlinearity (SQNL) function.

[0078] Examples of pooling layers include maximum pooling layers, minimum pooling layers, average pooling layers, and adaptive pooling layers.

[0079] Example techniques to train ANNs, such as to determine and / or update parameters of the NNs, include backpropagation-based gradient update and / or gradient descent techniques, such as Stochastic Gradient Descent (SGD), synchronous SGD, asynchronous SGD, batch gradient descent, and mini-batch gradient descent. The backpropagation-based gradient techniques are usable alone or in any combination, e.g., stochastic gradient descent is usable in a mini-batch context. Example optimization techniques usable with, e.g., backpropagation-based gradient techniques (such as gradient update and / or gradient descent techniques) include Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.

[0080] According to implementation, elements of ANNs, such as layers, loss functions, and / or objective functions, variously correspond to one or more hardware elements, one or more software elements, and / or various combinations of hardware elements and software elements. For a first example, a convolution layer, such as a N x M x D convolutional layer, is implemented as hardware logic circuitry comprised in an Application Specific Integrated Circuit (ASIC). For a second example, a plurality of convolutional, activation, and pooling layers ar e implemented in a TensorFlow machine learning framework on a collection of Internet-connectedservers. For a third example, a first one or more portions an ANN, such as one or more convolution layers, are respectively implemented in hardware logic circuitry according to the first example, and a second one or more portions of the ANN, such as one or more convolutional, activation, and pooling layers, are implemented on a collection of Internet-connected servers according to the second example. Various implementations are contemplated that use various combinations of hardware and software elements to provide corresponding price and performance points.

[0081] Example characterizations of an ANN architecture include any one or more of topology, interconnection, number, arrangement, dimensionality', size, value, dimensions and / or number of hyperparameters, and dimensions and / or number of parameters of and / or relating to various elements of a NN (e.g., any one or more of layers, loss functions, and / or objective functions of the NN).

[0082] Example implementations of an ANN architecture include various collections of software and / or hardware elements that collectively perform operations according to the ANN architecture. Various ANN implementations vary according to machine learning framework, programming language, runtime system, operating system, and underlying hardware resources. The underlying hardware resources variously include one or more computer systems, such as having any combination of Central Processing Units (CPUs), Graphics Processing Units (GPUs), Field Programmable Gate Arrays (FPGAs), Coarse-Grained Reconfigurable Architectures (CGRAs), Application- Specific Integrated Circuits (ASICs), Application Specific Instruction-set Processors (ASIPs), and Digital Signal Processors (DSPs), as well as computing systems generally, e.g., elements enabled to execute programmed instructions specified via programming languages. Various ANN implementations are enabled to store programming information (such as code and data) on non-transitory computer readable media and are further enabled to execute the code and reference the data according to programs that implement ANN architectures.

[0083] Examples of machine learning frameworks, platforms, runtime environments, and / or libraries, such as enabling investigation, development, implementation, and / or deployment of ANNs and / or ANN-related elements, include TensorFlow, Theano, Torch, PyTorch, Keras, MLpack, MATLAB, IBM Watson Studio, Google Cloud Al Platform, Amazon SageMaker, Google Cloud AutoML, RapidMiner, Azure Machine Learning Studio, Jupyter Notebook, and Oracle Machine Learning.

[0084] Fig. 8. is a block diagram illustrating an example computing device 800. One or more computing devices such as the computing device 800 may implement one or more processes described herein. For example, the computing device 800 may comprise one or more of the engines of the cheminformatics subsystem 104. As shown by Fig. 8, the computing device 800 may comprise a processor 802, a memory 804, a storage device 806, an I / O interface 808, and a communication interface 810, which may be communicatively coupled by way of a communication infrastructure 812. The computing device 800 may include fewer or more components than those shown in Fig. 8.

[0085] The processor 802 may include hardware for executing instructions, such as those making up a computer application or system. In examples, to execute instructions from an internal register, an internal cache, the memory 804, or the storage device 806 and decode and execute the instructions. The memory 804 may be a volatile or non-volatile memory used for storing data, metadata, computer-readable or machine-readable instructions, and / or programs for execution by the processor! s) for operating as described herein. The storage device 806 may include storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods.

[0086] The I / O interface 808 may allow a user to provide input to, receive output from, and / or otherwise transfer data to and receive data from the computing device 800. The I / O interface 808 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 808 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g. display drivers ), one or more audio speakers, and one or more audio drivers. The I / O interface 808 may be configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content.

[0087] The communication interface 810 may include hardware, software, or both. In any event, the communication interface 810 may provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 800 and one or more other computing devices and / or networks. The communication may be a wired or wireless communication. As an example, and not by way of limitation, the communicationinterface 810 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as WI-FI.

[0088] Additionally, the communication interface 810 may facilitate communications with various types of wire or wireless networks. The communication infrastructure 812 may also include hardware, software, or both that couples components of the computing device 800 to each other. Fore example, the communication interface 810 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes.

[0089] In addition to what has been described herein, the methods and systems may also be implemented in a computer program(s), software, or firmware incorporated in one or more computer-readable media include electronic signals (transmitted over wired or wireless connections) and tangible / non-transitory computer-readable storage media. Examples of tangible / non-transitory computer-readable storage media include, but are not limited to, a read only memory (ROM), a random-access memory' (RAM), removable disks, and optical media such as CD-ROM disks, and digital versatile disks (DVDs).

Claims

WHAT IS CLAIMED IS:

1. A non-transitory computer-readable medium (CRM) comprising instruction that, when executed by at least one processor, cause the at least one processor to: receive a dataset of electronic quantities across a manifold topology of a molecule of interest; perform manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; utilize a multi-variate regression to predict an elicited DGE covariance matrix in a particular cellular context; and export the covariance matrix as a symmetric semi-positive definite matrix.

2. The CRM of claim 1, wherein the multi-variate regression analysis method comprises a partial least square or NIPAL.

3. The CRM of claim 2, where the elicited DGE manifold is determined through kernelization of DGE data in a particular cellular context across a gene interaction network, and an array of biomarkers of the same particular cellular context.

4. The CRM of claim 3, where the gene interaction network where the gene interaction network is comprised of cellular protein-protein interactions.

5. The CRM of claim 4, where the kernelization process is performed using a reduced-rank Sparse Gaussian Process (SGP) with Spectral Mixture (SM) as the kernel function.

6. The CRM of claim 5, wherein the SGP utilizes a set of landmark points to approximate the shape of the underlying manifold.

7. The CRM of claim 6, where in the landmark points are selected as those points that exhibit the highest amount of uncertainty or variance during the SPG regression.

8. The CRN! of claim 7, wherein the kernel function is a covariance function of the spectral solution of the manifold identified through an eigen decomposition of the Graph Laplacian of the gene interaction network.

9. The CRM of claim 2, where the manifold is determined through MKMS of a dataset of molecular surfaces.

10. A method for manifold learning of chemical perturbations and genetic responsome via a trained ML / DL model, the method comprising: receiving a dataset of molecular surfaces; performing manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and a manifold topology; utilizing a multi-variate regression analysis from the encoded quantum information to produce the molecule’s elicited DGE manifold; exporting a covariant matrix (kernel) encoding of the elicited DGE manifold.

11. A method for manifold learning of chemical perturbations and genetic responsome via a trained ML / DL model, the method comprising: receiving a dataset of electronic quantities across a manifold topology of a molecule of interest; performing manifold regression on the electronic quantities to encode quantum information of a molecule in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the electronic quantities and the manifold topology; receiving a dataset of a molecule’s elicited DGE and corresponding biomarkers in a particular cellular context and a gene interaction network; performing manifold regression on the DGE data across the gene interaction network to encode molecular interaction information from the cellular context in a covariant matrix (kernel), wherein the covariant matrix captures mutual relationships among the genetic alterations and the manifold topology;performing a multi-variate regression analysis between the encoded quantum information and the DGE manifold.

12. The method of claim 11, where the multi- variate regression analysis method comprises a partial least squares regression or NIP AL.

13. The method of claim 12, where the elicited DGE manifold is determined through kernelization of DGE data in a particular cellular context across a gene interaction network, and an array of biomarkers of the same particular cellular context.

14. The method of claim 13, where the gene interaction network where the gene interaction network is comprised of cellular protein-protein interactions.

15. The method of claim 14, where the kernelization process is performed using a reduced- rank Sparse Gaussian Process (SGP) with Spectral Mixture (SM) as the kernel function.

16. The method of claim 15, wherein the SGP utilizes a set of landmark points to approximate the shape of the underlying manifold.

17. The method of claim 16, where in the landmark points are selected as those points that exhibit the highest amount of uncertainty or variance during the SPG regression.

18. The method of claim 17, wherein the kernel function is a covariance function of the spectral solution of a manifold identified through an eigen decomposition of the Graph Laplacian of the gene interaction network.

19. The method of claim 11, where the molecular encoding manifold is determined throughMKMS of a dataset of molecular surfaces.

Citation Information

Patent Citations

  • Method and systems for phytomedicine analytics for research optimization at scale

    US20220130493A1

  • Effects of a Molecule

    US20220277813A1