Systems and methods to identify candidate drugs for repurposing

A multi-modal approach integrating gene signatures, drug perturbation, and clinical data efficiently identifies drug candidates for repurposing, addressing the challenges of novel and rare diseases by leveraging existing drug data for rapid and cost-effective treatment solutions.

WO2026161415A1PCT designated stage Publication Date: 2026-07-30THE RGT UNIV OF MICHIGAN
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
THE RGT UNIV OF MICHIGAN
Filing Date
2026-01-21
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Conventional drug discovery methods face challenges in swiftly identifying effective treatments for novel or rare diseases and diseases refractory to standard-of-care medical therapy, necessitating a need for new techniques that facilitate drug repurposing.

Method used

A multi-modal approach integrating disease-associated gene signatures, drug perturbation data, and clinical data to identify candidate drugs for repurposing, utilizing systems and methods that include transcriptomic analysis, drug perturbation databases, and electronic health records to evaluate drug efficacy.

Benefits of technology

Facilitates rapid identification of drug candidates with reduced risk, integrating diverse data sources for comprehensive analysis and leveraging previously approved drugs, thereby reducing development time and cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000027_0001
    Figure IMGF000027_0001
  • Figure IMGF000027_0002
    Figure IMGF000027_0002
  • Figure IMGF000027_0003
    Figure IMGF000027_0003
Patent Text Reader

Abstract

Provided herein are systems and methods to identify candidate drugs for repurposing. In particular, provided herein are systems and methods employing a multi-modal approach to identify candidate drugs for repurposing by integrating disease-associated gene signature, drug perturbation, and clinical data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Attorney Docket Number: UM-43517.601

[0002] SYSTEMS AND METHODS TO IDENTIFY CANDIDATE DRUGS FOR REPURPOSING

[0003] The present application claims priority to U.S. Provisional application serial number 63 / 747,508, filed January 21, 2025, which is herein incorporated by reference in its entirety.

[0004] FIELD

[0005] Provided herein are systems and methods to identify candidate drugs for repurposing. In particular, provided herein are systems and methods employing a multi-modal approach to identify candidate drugs for repurposing by integrating disease-associated gene signature, drag perturbation, and clinical data.

[0006] BACKGROUND

[0007] Conventional drag discovery methods often face challenges in swiftly identifying effective treatments, particularly for novel or rare diseases and diseases refractory to standard-of-care medical therapy. Drug repurposing, or the use of drags outside of the scope of their original medical indication, may offer a potential solution for many medical conditions without known therapies or refractory to existing standards of care. Potential benefits of drug repurposing include greater likelihood of safety given that the drag has undergone prior clinical trials, shorter time frame for drag development, and potentially less cost investment with regards to getting the drag to the market. For rare diseases, drag repurposing may be the only economically viable way to bring drags that comply with regulatory standards to market. There is a need for new techniques that facilitate drag repurposing efforts.

[0008] SUMMARY

[0009] Provided herein are systems and methods to identify candidate drags for repurposing. In particular, provided herein are systems and methods employing a multi-modal approach toAttomey Docket Number: UM-43517.601

[0010] identify candidate drugs for repurposing by integrating disease-associated gene signature, drag perturbation, and clinical data.

[0011] In some embodiments, the systems and methods for drag repurposing for a given disease state or condition involve three steps. The first is to utilize one or more disease-associated gene signatures (e.g., scRNA-seq data) to pinpoint transcriptomic signatures associated with the disease state or condition. The second is to mine drag perturbation data from a predefined database (e.g., the Library of Integrated Network-Based Cellular Signatures (LINCS) database), to identify potential drug candidates capable of reversing the identified transcriptomic signatures. These drag candidates are then pooled into drag classes based on standard mechanism of action and the efficacy of these candidate drag classes is evaluated using clinical data on drug exposure and disease-associated outcome (e.g., remission as obtained from an electronic health record (EHR)). According to the technology provided herein, the amalgamation of gene signature findings with drug perturbation and clinical data provides a comprehensive framework for scrutinizing specific drugs and drug classes' effects on a given disease state.

[0012] Advantages of these systems and methods include rapid identification of drug candidates for a given disease state, integration of diverse data sources for a comprehensive analysis, and reduced risk given utilization of drags that were previously approved (e.g., FDA- approved).

[0013] In some embodiments, provided herein are methods for characterizing drugs, comprising one or more or all of the steps of: (a) identifying a disease-associated gene signature for a disease of interest; (b) selecting a plurality of candidate drugs, using drug perturbation data, that alter the disease-associated gene signatures; and (c) characterizing the candidate drags related to the disease of interest by analyzing data contained in electronic health records. In some embodiments, the characterizing comprises testing the efficacy of the candidate drags. In some embodiments, the characterizing comprises identifying a new drug for treating the disease of interest. In some embodiments, selecting comprises generating a ranked list of candidate drags. In some embodiments, electronic health records are obtained from or contain data from an observational cohort study or clinical trial dataset. In some embodiments, further comprising step (d) of testing a candidate drag in a laboratory disease model.

[0014] In certain embodiments, provided herein are methods of characterizing drags comprising: a) identifying a disease-associated gene signature for a disease of interest, wherein the disease-Attorney Docket Number: UM-43517.601

[0015] associated gene signature comprises a plurality of differentially expressed genes (DEGs) in a first cell type, wherein the plurality of DEGs comprise a plurality of upregulated and / or downregulated genes, compared to non-disease wild type status, in the first cell type; b) selecting a plurality of drugs by inputting the plurality of DEGs into a drug perturbation dataset (LINCs) which outputs a plurality of candidate drugs that reverses upregulation and / or downregulation of at least some, or all, of the plurality of upregulated and / or downregulated genes in the first cell type, wherein the drug perturbation dataset comprises data that comprises a plurality of individual gene expression results of a particular candidate drug interacting with the particular cell type; c) characterizing said at last one of said plurality of candidate drugs related to said disease of interest by analyzing data contained in electronic health records.

[0016] In particular embodiments, the first cell type is selected from the group consisting of: T cells (cd4+, cd8+, regulatory t cells / tregs, exhausted t cells), b cells (naive, memory, plasma cells), nk cells, monocytes (classical / non-classical), macrophages (tissue-resident and inflammatory), dendritic cells (conventional des, plasmacytoid des), neutrophils, mast cells, eosinophils I basophils, innate lymphoid cells (ilcs), hematopoietic stem cells (hscs), progenitors (myeloid and lymphoid progenitors), erythroid lineage cells, megakaryocytes / platelet precursors, epithelial cells (basal, luminal, secretory, proliferative), goblet cells, paneth cells, enterocytes / colonocytes, tuft cells, ciliated cells, club cells, alveolar type i / type ii cells (at1 / at2), keratinocytes, fibroblasts (including activated / myofibroblast states), myofibroblasts, endothelial cells (vascular, lymphatic), pericytes, smooth muscle cells, mesothelial cells, neurons (excitatory, inhibitory, region-specific subtypes), astrocytes, microglia, oligodendrocytes, oligodendrocyte precursor cells (opes), ependymal cells, schwann cells, hepatocytes, cholangiocytes, kupffer cells, hepatic stellate cells, adipocytes, adipose stromal cells, pancreatic islet cells (β, α, δ, pp, ε cells), cardiomyocytes, cardiac fibroblasts, vascular endothelial cells, vascular smooth muscle cells, podocytes, proximal tubule cells, loop of henle cells, distal tubule cells, collecting duct cells (principal / intercalated), mesangial cells, malignant tumor cells, tumorinfiltrating t cells, tumor- associated macrophages (tarns), cancer-associated fibroblasts (cafs), myeloid-derived suppressor-like cells (mdsc-like), stem / progenitor cells (tissue-specific), activated fibroblasts, inflammatory macrophages, endothelial progenitors, and cycling / proliferating cells.Attorney Docket Number: UM-43517.601

[0017] In some embodiments, the assessment for characterizing comprises weighing: (a) two or more factors associated with a candidate drug selected from the group consisting of: (i) drug potency; (ii) drug selectivity; (iii) gene perturbation score based on connectivity score of altering target genes; and (iv) class score based on a number of drugs that have negative connectivity scores that belong to the same class as the candidate drug; and (b) the data from the electronic health records.

[0018] In some embodiments, the identifying of disease-associated gene signatures comprises transcriptomic analysis of cells from disease and non-disease samples. In some embodiments, transcriptomic analysis comprises single-cell RNA-seq analysis.

[0019] In some embodiments, provided herein is a system comprising a computer processor configured to conduct characterization of candidate drugs related to the disease of interest. In some embodiments, a computer processor is further configured to conduct the selection of a plurality of candidate drugs, using drug perturbation data. In some embodiments, a computer processor is further configured to conduct the identification of a disease-associated gene signature for a disease of interest.

[0020] DEFINITIONS

[0021] The terms “comprise(s),” “include(s),” “having,” “has,” “can,” “contain(s),” and variants thereof, as used herein, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of additional acts or structures. The singular forms “a,” “and” and “the” include plural references unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments “comprising,” “consisting of’ and “consisting essentially of,” the embodiments or elements presented herein, whether explicitly set forth or not.

[0022] The term “medical data,” such as data that may be found in an electronic health record, as used herein refers to certain parameters / conditions of a specific individual. Such medical data include one or more of the following: medical history, physical examination (e.g. age, weight, height, gender, etc.). In some embodiments, the medical data of the specific individual also include the pretreatment clinical data of the specific individual, which may also include the disease-related clinical data, e.g. pathology review, histologic subtype; imaging data; blood counts (CBC); biochemistry profile; hormone profile and markers of inflammation; tumorAttorney Docket Number: UM-43517.601

[0023] markers; molecular diagnostic tests; Immunohistochemical Staining (IHC); gene status, such as mutation in one or more genes, one or more amplification in one or more copies, genetic recombination, partial or complete genetic sequencing; and death indicator. One or more anticipated or completed treatment regimens may also be included, e.g. chemotherapy drugs, immunotherapy drugs, biological drugs, combination of two or more drugs, including treatment outcomes.

[0024] As used herein, the term “sample” is used in its broadest sense. In one sense, it is meant to include a specimen or culture obtained from any source, as well as biological and environmental samples. Biological samples may be obtained from animals (including humans) and encompass fluids, solids, tissues, and gases. Biological samples include blood products, such as plasma, serum and the like. Environmental samples include environmental material such as surface matter, soil, water, and industrial samples. Such examples are not however to be construed as limiting the sample types applicable to the present disclosure.

[0025] A “subject” or “patient” may be human or non-human (e.g., dogs, cats, cows, horses, sheep, poultry, fish, crustaceans, etc.) and may include, for example, animal strains or species used as “model systems” for research purposes, such a mouse model as described herein.

[0026] Likewise, subject or patient may include either adults or juveniles (e.g.. children). Moreover, subject or patient may mean any living organism, preferably a mammal (e.g., human or non-human). As used herein, the term “patient” typically refers to a subject that is being treated for a disease or condition. In one embodiment of the methods and systems provided herein, the mammal is a human.

[0027] The terms “candidate drugs” and “candidate drug compounds” may be used interchangeably herein.

[0028] The terms “transmit,” “receive,” and “communicate,” as well as derivatives thereof, encompass both direct and indirect communication. The terms “transmit,” “receive,” and “communicate.” as well as derivatives thereof, encompass both communication with remote systems and communication within a system, including reading and writing to different portions of a memory device.Attorney Docket Number: UM-43517.601

[0029] The term “translate” may refer to any operation performed wherein data is input in one format, representation, language (computer, purpose-specific, such as drug design or integrated circuit design), structure, appearance or other written, oral or representable instantiation and data is output in a different format, representation, language (computer, purpose-specific, such as drug design or integrated circuit design), structure, appearance or other written, oral or representable instantiation, wherein the data output has a similar or identical meaning, semantically or otherwise, to the data input. Translation as a process includes but is not limited to substitution (including macro substitution), encryption, hashing, encoding, decoding or other mathematical or other operations performed on the input data. The same means of translation performed on the same input data will consistently yield the same output data, while a different means of translation performed on the same input data may yield different output data which nevertheless preserves all or part of the meaning or function of the input data, for a given purpose. Notwithstanding the foregoing, in a mathematically degenerate case, a translation can output data identical to the input data.

[0030] The term “controller” means any device, system or part thereof that controls at least one operation. Such a controller may be implemented in hardware or a combination of hardware and software and / or firmware. The functionality associated with any particular controller may be centralized or distributed, whether locally or remotely.

[0031] Various functions described below can be implemented or supported by one or more computer programs, each of which is formed from computer readable program code and embodied in a computer readable storage medium. The terms “application” and “program” refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or a portion thereof adapted for implementation in a suitable computer readable program code. The phrase “computer readable program code” includes any type of computer code, including source code, object code, and executable code. The phrase “computer readable storage medium” includes any type of medium capable of being accessed by a computer, such as read only memory (ROM), random access memory (RAM), a hard disk drive, a compact disc (CD), a digital video disc (DVD), solid state drive (SSD), or any other type of memory. A “non-transitory” computer readable storage medium excludes wired, wireless, optical, or other communication links that transport transitoryAttorney Docket Number: UM-43517.601

[0032] electrical or other signals. A non-transitory computer readable storage medium includes media where data can be permanently stored and media where data can be stored and later overwritten, such as a rewritable optical disc or an erasable memory device.

[0033] As used herein, the term “database” shall generally mean a digital collection of data or information.

[0034] It will be understood that when an element (such as a first element) is referred to as being (operatively or communicatively) “coupled / coupled” or “connected / connected” to another element (such as a second element), it can be directly coupled or connected / coupled or connected to the other element (such as the second element) or via a third element. Conversely, it will be understood that when an element (such as a first element) is referred to as being “directly coupled” / ”directly coupled to” or “directly connected” / ”directly connected” to another element (such as a second element), there is no other element (such as a third element) intervening between the element and the other element.

[0035] As used herein, the phrase “configured (or set) to” may be used interchangeably with the phrases “adapted to”, “having... capability”, “designed to”, “adapted to”, “made to”, or “capable”, as the case may be. The phrase “configured (or set) to” does not substantially mean “specially designed in hardware”. Rather, the phrase “configured to” may indicate that a device is capable of performing an operation with another device or component. For example, the phrase “a processor configured (or arranged) to perform A, B and C” may refer to a general-purpose processor (such as a CPU or an application processor) or a special-purpose processor (such as an embedded processor) that may perform operations by executing one or more software programs stored in a memory device.

[0036] The terms and phrases used herein are used only to describe some embodiments of the present disclosure and do not limit the scope of other embodiments of the present disclosure. It is to be understood that the singular includes plural referents unless the context clearly dictates otherwise. All terms and phrases used herein (including technical and scientific terms and phrases) have the same meaning as commonly understood by one of ordinary skill in the art to which embodiments of the present disclosure belong. It will be further understood that terms and phrases, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not beAttorney Docket Number: UM-43517.601

[0037] interpreted in an idealized or overly formal sense unless expressly so defined herein. In some instances, the terms and phrases defined herein may be construed to exclude embodiments of the disclosure.

[0038] Definitions for other specific words and phrases may be provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most instances, such definitions apply to prior, as well as future uses of such defined words and phrases.

[0039] None of the description in this application should be read as implying that any particular element, step, or function is an essential element which must be included in the claim scope. The scope of patented subject matter is defined only by the claims. Any other term used in the claims, including, but not limited to, “mechanism,” “module,” “device,” “unit,” “assembly,” “element,” “member,” “apparatus,” “machine,” “system,” “processor,” or “controller,” is understood by the applicants to refer to structures known to those of ordinary skill in the relevant art.

[0040] DESCRIPTION OF THE FIGURES FIG. 1 is a heterogeneous network encompassing n drug classes, drugs, and genes and demonstrating their relationship to a target disease state in accordance with some embodiments. The ultimate goal is how to identify a drug that could reverse the genes signature associated with the disease.

[0041] FIG. 2 is a flowchart demonstrating a drug repurposing workflow. “DEGs” in this flowchart refers to a table comprising differentially expressed genes (DEGs), with dimensions encompassing the number of genes in rows and number of cells derived from samples in columns.

[0042] FIG. 3 is a flowchart of a process for creating a sorted list of differentially-expressed genes in accordance with some embodiments. A, B, and C represent cell types derived from patient samples among which gene expression is assessed.Attorney Docket Number: UM-43517.601

[0043] FIG. 4 is a flowchart of a process for creating datasets composed of synergistic and antagonistic drugs within a specific cell line C derived from drug perturbation data (e.g., LINCS), in accordance with some embodiments.

[0044] FIG. 5 is a flowchart of a process for creating a Cx2 ranked list of drugs that reverse the disease-associated gene signature, pooled into drug classes by mechanism of action in accordance with some embodiments, μ, η, λ, and τ refer to mathematical scoring of drugs described further in Detailed Description that are used to rank drugs by efficacy and C refers to the number of drug classes present in the final dataset.

[0045] FIG. 6 is a flowchart of a process for identifying exposure status to one or more of C candidate drug classes among patients with a target disease state, derived from a clinical dataset in accordance with some embodiments.

[0046] FIG. 7 is a flowchart of a process for creating a dataset composed of patients with a target disease state and whether or not they recovered / improved from this target disease state in response to one or more of C candidate drug classes in accordance with some embodiments.

[0047] FIG. 8 is a flowchart of a process for creating a dataset encompassing the odds of achieving the disease-related outcome among patients exposed to one or more of C candidate drug classes, derived from generalized linear models (GLM), in accordance with some embodiments.

[0048] Figure 9: Workflow from sCRNA-seq data on patients with COVID-AKI to detect COVID-AKI-associated gene signatures, identify candidate drugs for COVID-AKI gene signature reversal, and validate the utility of candidate drugs using N3C.

[0049] Figure 10. Figure 10: N3C pre-processing steps used to generate our final cohort of COVID-AKI patients.

[0050] DETAILED DESCRIPTION

[0051] Conventional approaches to drug development involve testing a single drug candidate through rigorous clinical trials designed to assess the safety and efficacy of the drug candidate on a disease-related outcome, such as improvement in associated laboratory measures or reductionAttorney Docket Number: UM-43517.601

[0052] in mortality. These approaches also do not tend to integrate multiple sources of biological information, such as transcriptomic data, in the evaluation of drug candidates. Drug repurposing, or the use of drugs outside of the scope of their original medical indication, can provide therapeutic options for many rare and common diseases with no existing treatment. These include a shorter time frame for drug development, greater likelihood of safety given that the drug candidates in drug repurposing have undergone prior clinical trials, and potentially less cost investment with regards to getting the drug to the market. Notable examples of drug repurposing include the use of trazodone for insomnia or sildenafil for erectile dysfunction. Hypothesis generation for drug repurposing can be performed systematically using computational approaches on big datasets, including molecular and electronic health record (EHR) data.

[0053] Computational methods can be utilized to identify genes that are associated with a disease state, illuminate associations between drugs and potential drug targets, and systematically analyze EHR and existing observational studies and clinical trial data to determine the efficacy of a drug for a different indication. A schematic example of a heterogeneous network that represents the inter-relatedness of genes, drugs, and a disease is shown in FIG. 1. Drug classes are defined for the purposes of this patent as drugs with a similar mechanism of action, such as glucocorticoid receptor agonism.

[0054] While exemplified herein using scRNA-seq data, the disease-associated gene signature data can be obtained using any of a variety of different technologies and approaches. In some embodiments, the disease-associated gene signature data comprises one or more of gene expression levels (e.g., RNA levels, protein levels, etc.), epigenetic status, mutational status, etc. determined at the single cell, tissue, organ, sample, or organism level. In some embodiments, gene expression is assessed using a technique, including but not limited to, scRNA-seq, real-time reverse transcription polymerase chain reaction (RT-PCR), fluorescent in situ hybridization (FISH), serial analysis of gene expression (SAGE), and microarray analysis. In some embodiments, the disease-associated gene signatures comprises transcriptomic analysis of cells from disease and non-disease samples.

[0055] Disease-associated gene signature data may be generated from bulk disease samples. For example, a bulk sample (e.g., biopsy sample) may be used to identify disease-associated gene signatures as compared to control samples, of similar type for non-disease subjects. Conversely,Attorney Docket Number: UM-43517.601

[0056] the methods may categorize or separate cells or tissue types from the sample prior to or as part of the workflow for generating gene signatures. For example, the cells or tissue types may be differentially tagged or separated (e.g.. single cell tagging) during analysis such that the gene signature data includes embedded information regarding the gene signatures in different cell or tissue types. Alternatively, cell or tissue types may be first separated prior to generation of the gene signature data, such that the final data is compiled across cell and tissues or interest.

[0057] The technology finds use for any number of disease and conditions, including, but not limited to, infectious diseases (e.g., bacterial infections (e.g., tuberculosis, strep throat, urinary tract infections (UTIs), Lyme disease, syphilis, pneumonia(bacterial)), viral infections (e.g., influenza, COVID-19, hepatitis A, hepatitis B, hepatitis C, human papillomavirus (HPV) infections, herpes simplex virus (HSV) infections, HIV / AIDS (manageable with antiretroviral therapy)), parasitic infections (e.g., malaria, giardiasis, toxoplasmosis, schistosomiasis), fungal infections (e.g., candidiasis, aspergillosis, athlete’s foot, histoplasmosis)), chronic diseases (e.g., cardiovascular conditions (e.g., hypertension, coronary artery disease, heart failure, arrhythmias), metabolic disorders (e.g., type 1 and 2 diabetes, hyperlipidemia, thyroid disorders (e.g., hyperthyroidism, hypothyroidism), polycystic ovary syndrome (PCOS)), respiratory diseases (e.g., asthma. Chronic obstructive pulmonary disease (COPD), sleep apnea), gastrointestinal disorders (e.g., acid reflux / GERD, irritable bowel syndrome (IBS), Crohn’s disease, ulcerative colitis)), cancers (e.g., breast cancer, prostate cancer, cervical cancer, colorectal cancer, melanoma, non-Hodgkin’s lymphoma), neurological and mental health conditions (e.g., depression, anxiety, bipolar disorder, schizophrenia (symptom management), epilepsy, Parkinson’s disease (symptom management), multiple sclerosis (disease-modifying treatments)), autoimmune disorders (e.g., rheumatoid arthritis, psoriasis, lupus, celiac disease, ankylosing spondylitis), musculoskeletal conditions (e.g., osteoporosis, osteoarthritis, tendinitis, rotator cuff injuries), and genetic disorders (with some level of treatment or management) (e.g., cystic fibrosis, sickle cell disease, hemophilia, thalassemia).

[0058] The technology can be applied to any number of agents, drugs, and medications, including, but not limited to, analgesics (e.g., acetaminophen, aspirin, ibuprofen, naproxen, codeine), antacids (e.g., calcium carbonate, aluminum hydroxide, magnesium hydroxide, sodium bicarbonate, simethicone), antianxiety drugs / sedatives (e.g., benzodiazepines (e.g., clonazepam,Attorney Docket Number: UM-43517.601

[0059] alprazolam, lorazepam, diazepam, bromazepam), azapirones (e.g. buspirone, gepirone)), antiarrhythmics (e.g., hyperpolarization-activated cyclic nucleotide-gated (HCN) channel blockers (e.g. ivabradine), sodium channel blockers (e.g., quinidine, procainamide, disopyramide, lidocaine, mexiletine, flecainide, propafenone), beta-blockers, potassium channel blockers (e.g., amiodarone, dronedarone, dofetilide, sotalol, ibutilide), non-dihydropyridine calcium channel blockers (e.g.. diltiazem, verapamil), others (e.g.. adenosine, digoxin)), antibacterials / antibiotics (e.g., glycylcylines (e.g. tigecycline), tetracyclines (e.g., doxycycline, minocycline), lincosamides (e.g., clindamycin), macrolides (e.g., azithromycin, clarithromycin, erythromycin), oxazolidinones (e.g., linezolid), sulfonamides (e.g., sulfamethoxazole), aminoglycosides (e.g., tobramycin, gentamicin, amikacin), beta-lactams (penacillins, cephalosporins, carbapenems) (e.g., amoxicillin, cefazolin, meropenem), fluoroquinolones (e.g., ciprofloxacin, levofloxacin), glycopeptides (e.g., vancomycin), cyclic lipopeptides (e.g., daptomycin), nitroimidazoles (e.g., metronidazole)), anticoagulants (e.g., warfarin, apixaban, dabigatran, heparin), thrombolytics (e.g., Tenecteplase. alteplase, streptokinase), anticonvulsants (commonly regarded to as antiseizure medications) (e.g., brivaracetam, cannabidiol, carbamazepine, cenobamate, clobazam, clonazepam, eslicarbazepine, ethosuximide, felbamate, fosphenytoin, gabapentin, lacosamide, levetiracetam, oxcarbazepine, perampanel, phenobarbital, phenytoin, pregabalin, primidone, rufinamide, stiripentol, tiagabine, topiramate, valproate products (e.g., valproate sodium, divalproex sodium, valproic acid), vigabatrin, zonisamide), antidepressants (e.g., selective serotonin reuptake inhibitors (SSRIs) (e.g., fluoxetine, paroxetine, sertraline, citalopram, escitalopram), serotonin and norepinephrine reuptake inhibitors (SNRIs) (e.g., duloxetine. venlafaxine, desvenlafaxine. levomilnacipran), atypical antidepressants (e.g., trazodone, mirtazapine, vortioxetine, vilazodone, bupropion), tricyclic antidepressants (e.g., imipramine, nortriptyline, amitriptyline, doxepin, desipramine), monoamine oxidase inhibitors (MAOIs) (e.g.. tranylcypromine, phenelzine, isocarboxazid, selegiline), other / an tidepressant augmentation medications (e.g., aripiprazole, quetiapine, lithium)), antidiarrheals (e.g., bismuth subsalicylate, octreotide, Crofelemer, loperamide, difenoxin HCl / atropine, psyllium, loperamide / simethicone, diphenoxylate / atropine, rifaximin), antiemetics (e.g., serotonin-receptor antagonists (e.g., ondansetron, granisetron, dolasetron, palonosetron), glucocorticoids (e.g., dexamethasone), anticholinergics (e.g., scopolamine), NK receptor antagonists (e.g., aprepitant, fosaprepitant), dopamine receptor antagonists (e.g., prochlorperazine, chlorpromazine),Attorney Docket Number: UM-43517.601

[0060] butyrophenones (e.g., droperidol, haloperidol), benzamides (e.g., metoclopramide), cannabinoid therapy (e.g., nabilone, dronabinol), antihistamines (e.g., diphenhydramine, meclizine, promethazine)), antifungals (e.g.. clotrimazole, econazole, miconazole, terbinafine, fluconazole, ketoconazole, nystatin, amphotericin), antihistamines (e.g., fexofenadine, hydroxyzine, diphenhydramine, loratadine, clarinex, cetirizine, xyzal), anti-inflammatories (non-steroidal antiinflammatory drugs) (e.g., diclofenac, celecoxib, mefenamic acid, etoricoxib, indomethacin), antineoplastics (e.g., alkylating agents (e.g., altretamine, bendamustine, busulfan, carmustine, chlorambucil), antimetabolites (e.g., antifolates (e.g., methotrexate, pemetrexed, pralatrexate, trimetrexate), purine analogues (e.g., azathioprine, cladribine, fludarabine, mercaptopurine, thioguanine), pyrimidine analogues (e.g., azacitidine, capecitabine, cytarabine, decitabine, floxuridine, fluorouracil, gemcitabine, trifluridine / tipracil))). antipsychotics (e.g., chlorpromazine, fluphenazine, haloperidol, loxapine, molindone, perphenazine, pimozide, prochlorperazine, thiothixene, thoridazine, trifluoperazine, aripiprazole, asenapine, brexpiprazole, cariprazine, clozapine, iloperidone, lumateperone, lurasidone, olanzapine, quetiapine, paliperidone, pimavanserin, risperidone, ziprasidone, xanomeline, trospium chloride), antipyretics (e.g., acetaminophen, aspirin, sodium salicylate, salicylic acid, piroxicam, meloxicam, indomethacin, naproxen, ketoprofen, phenylbutazone), antivirals (e.g., acyclovir, adefovir, amantadine, ampligen, amprenavir, umifenovir, atazanavir, atripla, baloxavir marboxil, biktarvy, boceprevir, bulevirtide, cidofovir, cobicistat, combivir, daclatasvir, darunavir, delavirdine, descovy, didanosine, docosanol, dolutegravir, doravirine, edoxudine, efavirenz, elvitegravir, emtricitabine, enfuvirtide), beta-blockers (e.g., atenolol, bisoprolol, carvedilol, labetalol, metoprolol, propranolol, sotalol), bronchodilators (e.g., drenergic bronchodilators (e.g.. albuterol, epinephrine, levalbuterol, formoterol, terbutaline, racepinephrine, arformoterol, olodaterol, salmeterol, isoproterenol), anticholinergic bronchodilators (e.g., tiotropium, umeclidinium, aclidinium, ipratropium, glycopyrrolate, revenfenacin), methylxanthines (e.g.. theophylline)), cold cures (e.g., ibuprofen, naproxen, diphenhydramine, pseudoephedrine, phenylephrine, acetaminophen, dextromethorphan, doxylamine), corticosteroids (e.g., cortisone, prednisolone, methylprednisolone, dexamethasone, betamethasone, hydrocortisone), cough suppressants (e.g., benzonatate systemic, dextromethorphan systemic, hydrocodone systemic, codeine systemic, chlophedianol systemic), cytotoxics (e.g., azathioprine, cyclophosphamide, methotrexate), decongestants (e.g., pseudoephedrine systemic, phenylephrine systemic,Attorney Docket Number: UM-43517.601

[0061] ephedrine systemic), diuretics (e.g., carbonic anhydrase inhibitors (e.g., acetazolamide systemic, methazolamide systemic, dichlorphenamide systemic), loop diuretics (e.g., furosemide systemic, bumetanide systemic, torsemide systemic, ethacrynic acid systemic), miscellaneous diuretics (e.g., pamabrom systemic, urea systemic, mannitol systemic), potassium-sparing diuretics (e.g., spironolactone systemic, eplerenone systemic, triamterene systemic, amiloride systemic), thiazide diuretics (e.g., hydrochlorothiazide systemic, chlorthalidone systemic, indapamide systemic, metolazone systemic, Bendroflumethiazide systemic, chlorothiazide systemic)), expectorants (e.g., guaifenesin systemic, potassium iodide systemic), hormones (e.g., 5-alpha-reductase inhibitors (e.g., finasteride systemic, dutasteride systemic, dutasteride / tamsulosin systemic), adrenal cortical steroids (e.g., corticotropin (e.g., corticotropin systemic, cosyntropin systemic, corticorelin systemic), glucocorticoids (e.g., prednisone systemic, methylprednisolone systemic, triamcinolone systemic, dexamethasone systemic, budesonide systemic, prednisolone systemic, hydrocortisone systemic, betamethasone systemic, cortisone systemic, vamorolone systemic, deflazacort systemic), mineralocorticoids (e.g.. fludrocortisone systemic)), adrenal corticosteroid inhibitors (e.g., osilodrostat systemic, metyrapone systemic, levoketoconazole systemic), antiandrogens (e.g., apalutamide systemic, enzalutamide systemic, bicalutamide systemic, darolutamide systemic, nilutamide systemic, flutamide systemic), antidiuretic hormones (e.g., desmopressin systemic, vasopressin systemic, terlipressin systemic), antigonadotropic agents (e.g.. danazol systemic), antithyroid agents (e.g., methimazole systemic, propylthiouracil systemic, potassium iodide systemic, sodium iodide-i- 131 systemic), aromatase inhibitors (e.g., letrozole systemic, anastrozole systemic, exemestane systemic), calcimimetics (e.g., cinacalcet systemic, etelcalcetide systemic), calcitonin, estrogen receptor antagonists (e.g., felvestrant systemic, elacestrant systemic), gonadotropin-releasing hormone antagonists (e.g., elagolix systemic, relugolix systemic, degarelix systemic, ganirelix systemic, cetrorelix systemic), growth hormone receptor blockers (e.g., teprotumumab systemic, pegvisomant systemic), growth hormones (e.g., somatropin systemic, tesamorelin systemic, somatrogon systemic, somapacitan-beco systemic, macimorelin systemic, lonapegsomatropin systemic), insulin-like growth factors (e.g., mecasermin systemic), melanocortin receptor agonists (e.g., bremelanotide systemic, setmelanotide systemic, afamelanotide systemic), miscellaneous hormones (e.g., vosoritide systemic), parathyroid hormone and analogs (e.g., teriparatide systemic, abaloparatide systemic, parathyroid hormone systemic, palopegteriparatide systemic),Attorney Docket Number: UM-43517.601

[0062] progesterone receptor modulators (e.g., ulipristal systemic, mifepristone systemic), prolactin inhibitors (e.g., cabergoline systemic, bromocriptine systemic), selective estrogen receptor modulators (e.g.. tamoxifen systemic, ospemifene systemic, raloxifene systemic, toremifene systemic), somatostatin and somatostatin analogs (e.g., octreotide systemic, lanreotide systemic, pasireotide systemic), synthetic ovulation stimulants (e.g., clomiphene systemic), thyroid drugs (e.g., levothyroxine systemic, thyroid desiccated systemic, liothyronine systemic, thyrotropin alpha systemic), hypoglycemics (oral) (e.g., sulfonylureas (e.g., glipizide, glyburide, gliclazide, glimepiride), meglitinides (e.g., repaglinide, nateglinide), biguanides (e.g., metformin), thiazolidinediones (e.g., rosiglitazone, pioglitazone), a-glucosidase inhibitors (e.g., acarbose, miglitol, voglibose), DPP-4 inhibitors (e.g., sitagliptin, saxagliptin, vildagliptin, linagliptin, alogliptin), SGLT2 inhibitors (e.g.. dapagliflozin, canagliflozin), cycloset (e.g., bromocriptine)), immunosuppressives (e.g., glucocorticoids (e.g., calcineurin inhibitors (e.g., cyclosporine, tacrolimus, everolimus), nucleotide synthesis inhibitors (e.g., mycophenolate mofetil, mizoribine. leflunomide, azathioprine)), protein drugs (e.g., muromonab CD3, alemtuzumab, rituximab, basiliximab, belatacept)), laxatives (e.g., bulk-forming laxatives (e.g., fybogel, psyllium, polycarbophil, methylcellulose), osmotic laxatives (e.g.. polyethylene glycol, magnesium hydroxide solution, glycerin), stool softener laxatives (e.g., docusate), stimulant laxatives (e.g., bisacodyl, senna), prescription-only laxatives (e.g., lactulose, linaclotide, lubiprostone, prucalopride, plecanatide, lactulose, lactitol, methylnaltrexone, naloxegol, naldemedine)), muscle relaxants (e.g., antispasmodics (centrally acting skeletal muscle relaxants) (e.g., carisoprodol, chlorzoxazone, cyclobenzaprine, metaxalone, methocarbamol, orphenadrine, tizanidine). antispastics (e.g., baclofen, dantrolene, diazepam)), sex hormones (female) (e.g., estradiol, norethindrone, ethinyl estradiol, drospirenone, norgestimate, conjugated estrogens, medroxyprogesterone), sex hormones (male) (e.g., bicalutamide, goserelin, degarelix, abiraterone, testosterone, abiraterone, enzalutamide, apalutamide, darolutamide), sleeping drugs (e.g., trazodone, zolpidem, eszopiclone, temazepam, daridorexant, amitriptyline, quetiapine, mirtazapine, clonazepam, gabapentin, lorazepam, estazolam, flurazepam, doxepin, suvorexant, triazolam, diphenhydramine, ramelteon, quazepam), tranquilizer (e.g., minor tranquilizers (e.g., benzodiazepines (e.g., diazepam, chlordiazepoxide, alprazolam), buspirone, lorazepam), major tranquilizers (e.g., aripiprazole, chlorpromazine, haloperidol, lithium carbonate, olanzapine, risperidone, thioridazine, trifluoperazine)), and vitamins (e.g., vitamin A, vitamin C, vitamin D,Attorney Docket Number: UM-43517.601

[0063] vitamin E, vitamin K, choline, B vitamins (e.g., thiamin, riboflavin, niacin, pantothenic acid, biotin, vitamin B6, vitamin B12, folate / folic acid)).

[0064] While exemplified herein with LINCS, the drug perturbation data can be obtained using any of a variety of different databases, including newly created databases and disease-specific databases (e.g., LINPS for cancer). Such databases include, but are not limited to, LINCS, PerturBase, Connectivity Map (CMap), and ChemPert. The drug perturbation data may also be generated in silico. For example, the methods may include applying in silico perturbagens to gene signatures using a pre-existing drug-expression profile database(s) to identify drugs candidates associated with reversal of the gene signatures.

[0065] The methods described herein may generate a drug potency score; a drug selectivity score; a genetic perturbation score based on connectivity score of altering target genes; a drug class score based on a number of drugs that have negative connectivity scores that belong to the same class as the candidate drug; and / or a summary or total score of a given drug(s) or drug class(es) based on drug perturbation data for one or more candidate drugs. Accordingly, the candidate drugs or drag classes may be characterized, classified, and / or ranked by one or more or all of drug potency; drug selectivity; gene perturbation; and drug class directed to a target disease(s).

[0066] The methods described herein further use observational studies or other clinical data, e.g., from electronic health records, to weigh the efficacy of the high scoring drug candidates and drug classes. The clinical data considers real-world usage of the drugs with metrics of disease remission or improvement for the candidate drug(s) or drug class(es) being characterized. In some embodiments, the electronic health records are masked to the identity of the patient in compliance with HIPAA and other health record confidentiality and disclosure requirements. In some embodiments, the electronic health records are obtained from or contain data from an observational cohort study or clinical trial dataset. In some embodiments, electronic health record comprises data from one or more observational cohorts or clinical trials. Drug classes which meet or exceed a desired disease-related outcome are used to generate a table of drug classes that may be applied in a real-world scenario off-label to treat the target disease.

[0067] In some embodiments, the methods may further comprise testing efficacy candidate drug(s). The testing may involve testing of candidate drug(s) in vitro, in cell or tissue models ofAttorney Docket Number: UM-43517.601

[0068] the target disease(s), animal disease models of the target disease(s), or in a clinical trial of patients.

[0069] Non-limiting, exemplary configurations of the systems and methods are provided below to illustrate structures and functions of the technology.

[0070] FIG. 1 is an exemplary network of drug classes, drugs, and genes including a plurality of nodes, which are associated with biological data for a different modality such as a disease, drug or drag class, or gene. Both inter and intra-modality connections are shown between nodes, describing the relationship between the nodes with the disease at the center. Pointed arrows represent synergistic effects, as a drag resulting in upregulation of a gene, and flattened arrows represent antagonistic effects, such as a drug resulting in downregulation of a gene. The inclusion of several modalities and possible interactions in the schematic reflects the complexity of drug therapy. Drugs, drug classes, and genes may have interactions within a given modality, such as drug-drug interactions, or across modalities, such as drugs decreasing expression of specific genes or genetic variation reducing the efficacy of a drug. Information regarding genegene interactions may be derived from a wide variety of data sources, such as transcriptomics analyses. Gene-gene interactions may reflect many different biological pathways including shared inheritance, epigenetic expression changes such as the expression of one gene increasing the expression of another, or protein-protein interactions. Similarly, drug-drug interactions may reflect direct antagonistic effects of one drug on the ability of another drug to bind to a target site, allosteric inhibition or activation, or off-target effects. Gene-drug interactions, in turn, may be derived from publicly available drag perturbation databases such as CMAP-LINCS-L1000, the Connectivity Map (CMAP) Library of Integrated Network-based Cellular Signatures (LINCS) L1000 dataset. As an example in FIG. 1, drug 1 belonging to drug class 1 may, for example, upregulate the expression of gene 1 which then increases the severity of a disease state. On the other hand, gene n, which may increase the severity of a disease state, may be downregulated by drug n which is therefore therapeutic with respect to the disease state.

[0071] The approach described herein integrating data from, for example, single cell transcriptomics, with drug perturbation and ultimately clinical data allows for the elucidation of possible inter-modality interactions between drugs, genes, and a target disease state.Attorney Docket Number: UM-43517.601

[0072] FIG. 2 demonstrates an exemplary drug repurposing workflow in accordance with some embodiments. In this workflow, the inputs are tissue samples (200), or samples derived from non-invasive approaches (210) such as urine from patients with and without a target disease. These samples are then processed by a sequencer (220). Cells are first dissociated into a singlecell suspension, typically using enzymatic digestion or mechanical disruption. Next, single cells, containing a unique barcode or molecular tag. are encapsulated into tiny droplets or microwells. This step, known as cell barcoding, allows for the identification of individual cells and their respective RNA transcripts throughout the sequencing process. Once encapsulated, the RNA molecules within the cells are reverse transcribed into complementary DNA (cDNA).

[0073] Importantly, this step preserves information about the abundance and identity of RNA transcripts within the cell. Subsequently, the cDNA molecules are amplified using polymerase chain reaction (PCR) to generate sufficient material for sequencing. To ensure that cDNA molecules retain their unique cellular origin, the PCR amplification is performed in a manner that preserves the cell barcodes. Following amplification, the cDNA libraries are sequenced using high-throughput sequencing platforms. During sequencing, the cell barcodes and unique molecular identifiers (UMIs) are used to assign sequenced reads to their respective cell and transcript, allowing for the quantification of gene expression at the single-cell level.

[0074] The output of the sequencer is a count matrix representing the degree of expression of thousands of genes within a given cell, which is then matched with whether or not the patients from which the sample is derived have a given disease state. The limma statistical test is then used to identify differentially expressed genes with respect to the target health condition in an applicable cell type. In this fashion, a list of differentially expressed genes are generated representing genes that are up- or down -regulated within a given cell type among patients with the target disease relative to those without the target disease. “DEGs” refers here to a table containing differentially expressed genes in the rows and columns showing the log-fold change of gene expression associated with the disease, p-value from the appropriate statistical test, and the adjusted p-value based on user preference for methodology of p-value correction. The list of differentially expressed genes is then filtered to those expressed in an appropriate cell type associated with the disease state, such as kidney epithelial cells for kidney disease as well as to the columns for log-fold change, unadjusted p-value, and adjusted p-value to produce a final gene list table. The genes in this list are then inputted to a drug perturbation platform, such asAttorney Docket Number: UM-43517.601

[0075] LINCS, which is used to generate a list of drugs by drug mechanism that are associated with reversal of disease-associated gene signatures serving as a proxy for the disease-related outcome. The drug classes identified in this step as being associated with reversal of disease-associated gene signatures may serve as promising drugs.

[0076] While conventional approaches to drug discovery typically apply the identified drug candidates in clinical trials, the use of existing clinical data in which patients have already been exposed to the FDA-approved drug candidate may expedite drug repurposing for the target disease. In that vein, the next step of the workflow involves the use of an existing clinical dataset such as an electronic health record (EHR), containing patients with the target disease and who may have been exposed to one or more of the drug candidates. This cohort is then filtered to criteria set per the user, including filtering to patients with target disease state, adequate duration of drag therapy, and sufficient length of exposure to a given drag class. Lastly, generalized linear models comparing patients with and without exposure for the odds of disease-related outcome are used to generate a table of drag classes with columns odds ratio, unadjusted p-value, and adjusted p-value. Drag classes with adjusted p-values below the user’s set threshold, such as 0.05, therefore reflect drug classes that may be applied in a real-world scenario off-label to treat the target disease in lieu of clinical trials data for this specific indication.

[0077] FIG. 3 schematically illustrates an exemplary process for generating a sorted list of genes that are either up or downregulated in a target disease. The first step in this schematic is node linking the transcriptomic data (300) to two individual cell types represented by rectangle, hexagon, and circle for a disease state and a non-disease state. Transcriptomics is utilized to generate gene expression data using tissue or non-invasive samples derived from patients with and without a target disease to gene expression data. Initially, gene expression data is not limited to individual cell types and incorporates a large majority of or all of cell types present in the sample. There are several possible cell types for which transcriptomics data can be subset to. These are represented as squares for cell type A, hexagons for cell type B, and circles for cell type C. Moreover, gene expression is compared between patients with and without the disease to elucidate which genes are associated with the disease state specifically, embodied in the schematic as 310 and 320 corresponding to patients with or without the disease, respectively. The approach shown in the figure for evaluating a specific cell type can be scaled to includeAttorney Docket Number: UM-43517.601

[0078] other cell types or comparisons. For example, one may include nested disease states, such as a complication of an existing disease, and therefore include three comparison groups encompassing patients with the disease and complication, patients with the disease but not the complication, and patients without the disease. However, this data can be filtered to columns representing cell types that are specific to the target disease, such as kidney epithelial cells in kidney disease. Accordingly, the next step in the schematic is to subset 310 and 320 to cell type C, represented by nodes 330 and 340, respectively.

[0079] Comparison of the two study groups at the level of a single cell type C, as shown in this schematic, produces a data table of differentially expressed genes between groups as shown in 330, containing gene identifiers in the rows and cell type identifiers in the columns. DEGs refers here to genes whose expression levels significantly vary between different experimental conditions, such as healthy versus diseased tissues, treated versus untreated samples, or different developmental stages. Identifying DEGs is helpful in elucidating the genetic basis of complex traits and diseases. Techniques like single cell RNA sequencing are used to measure gene expression levels across different conditions. Once data is obtained, statistical methods are applied to detect genes with significant expression changes. These DEGs can then be further analyzed to infer their functional roles and regulatory networks. As illustrated in this schematic, the list of DEGs in the target cell type is filtered to a subsetted data table, termed “Gene List,” containing the same genes themselves in rows and log fold change in gene expression for the genes with respect to the disease state, unadjusted p-values, and adjusted p-values for this log fold change in the columns.

[0080] The “Gene List” generated using the workflow shown in FIG. 3 may be inputted to a drug perturbation dataset to generate a list of drugs that amplify and reverse the overall disease-associated gene signature as shown in FIG. 4. One example of a drug perturbation tool is Connectivity Map Linked User Environment (CLUE), hosted by the Library of Integrated Network-based Cellular Signatures, an NIH funded project that measures changes in gene expression that occur in the setting of various diseases and drugs. The drug perturbation tool then compares the expression patterns of input DEGs with the gene expression signatures of compounds or drugs stored in its database. This comparison is often based on gene expression profiles obtained from drug-treated cell lines. Through computational algorithms and statisticalAttorney Docket Number: UM-43517.601

[0081] methods, the tool identifies drugs or compounds that may reverse the expression patterns of the input DEGs. These candidate drugs may exert their effects through various mechanisms, such as activating or inhibiting specific pathways or molecular targets associated with the disease phenotype. While drug perturbation tools can also prioritize potential therapeutic agents based on their predicted efficacy and safety profiles, for the purposes of this example, the reversal of expression patterns was assessed specifically. As shown in process 400, this drug perturbation dataset can be applied selectively to a specific cell line, embodied by a circle containing the letter “C” in this schematic. Drug candidates that are revealed through this drug perturbation approach to amplify disease-associated gene signatures are stored in a table titled “Synergistic Drugs” with rows referring to the drug name, while drug candidates that reverse disease-associated gene signatures are stored in a table titled “Antagonistic Drugs” with rows referring to the drug name.

[0082] As shown in FIG. 5, the drug classes may be ranked by their efficacy with respect to reversal of disease-associated gene signatures. With the “Antagonistic Drugs” data table as the input, four metrics are calculated using a data input utility. The metrics include (1) the potency score of a given drug in the target cell line generated from the CLUE server, (2) the selectivity score of the drug, (3) the genetic perturbation score of the drug which represents the averaged connectivity score of knocking down the target genes of the drug, and (4) the class score, or the number of drugs that have negative connectivity scores that belong to the same class as a given drug. These four metrics are summed in order to determine the total score of the drug, and this process is repeated for antagonistic drugs to create a “Total Drug Scores” table with rows as drugs and columns including drug class and total score. The drugs in the “Total Drug Scores” table are then sorted in descending order by total drug score and grouped into drug classes by mechanism of action, creating a Cx2 data table “Ranked Drug Classes” where C is the number of drug classes.

[0083] FIG. 6 is an exemplary framework for identifying an applicable cohort of patients with a target disease state and who may have been exposed to one or more of the target drug classes. In this framework, a clinical data set first filtered to patients with the target disease state. This creates a data table titled “Patients with Disease State” that has patient IDs in the rows and covariates, such as demographics and other diagnoses, in the columns along with the clinical outcome of interest. In the second step, patients are then filtered to those patients with adequateAttorney Docket Number: UM-43517.601

[0084] duration of drug therapy to assess for the outcome of interest. Ding therapy may be completed either in the inpatient or outpatient setting. In the next step of this process, patients are split into those who were exposed to one or more ranked drug classes from the “Ranked Drug Classes” table. Exposure criteria may represent, for example, one third of the patient's length of stay in the hospital. This process is repeated for C drug classes in the “Ranked Drug Classes” table. This process ultimately creates C individual data tables composed of patients who are either exposed or unexposed to a given drug class during their follow-up period. Additional filtration criteria for the cohort definition may be pursued by the user, however this patent represents a general, scalable approach with some embodiments.

[0085] FIG. 7 is a framework for creating a final “Cohort Definition” table composed of patients, their exposure status to one or more of the C drug classes, and their outcome, such as whether or not they recovered from the target disease during the duration of drug therapy. Specifically, the analysis begins with C data tables composed of patients and their exposure status with respect to one or more of the C drug classes. Subsequently these data tables are parsed to identify the patient outcome - that is, whether or not they experienced the disease -related outcome during the period of drug therapy. Disease-related outcomes are specific to a target disease as described in this patent.

[0086] Examples include reduction of serum creatinine to baseline levels after an acute kidney injury, reduction in blast counts below the threshold in an acute hematologic malignancy, or reduction in A1C among patients with diabetes. Subsequently, for an individual drug class a “Cohort Definition” table is created consisting of patient IDs in rows and exposure and disease-related outcome in the columns. These individual C “Cohort Definition” tables are then merged to create a final “Cohort Definition” table containing patient IDs and rows and exposure status, outcome status, and covariates in the columns. A data table containing drug classes and the associated odds ratios, unadjusted p-values, and adjusted p-values derived from generalized linear mixed models may be generated as shown in FIG. 8 in accordance with some embodiments. First, the “Cohort Definition” table created using the schematic shown in FIG. 7 parsed into individual cohort definition tables for drug class one through C. For the cohort definition tables, an individual, generalized linear mixed model is generated. The outcome modeled by the generalized linear mixed model is outcome status, as in whether a patientAttorney Docket Number: UM-43517.601

[0087] experienced the disease-related outcome, and the exposures include exposure to the drug class and any applicable literature- guided covariates. For example, age or sex may be utilized as covariates in a generalized linear mixed model to assess the efficacy of a drug class on disease-related outcome. Subsequently, the odds ratios and unadjusted P values derived from the generalized linear mixed model are extracted and input into a new data table where a row corresponds to an individual drug class. Once this is repeated for C drug classes, adjusted p-values are determined using the adjustment method of the user's choice, such as Bonferroni correction. The final output of this workflow is a data table containing drug classes and their associated odds ratios, unadjusted p-values, and adjusted P values with respect to experiencing a disease-related outcome within a given duration of drug therapy that is user-determined. Drugs belonging to the drug classes

[0088] Some embodiments are directed to a computer system comprising at least one computer processor and at least one storage device encoded with a plurality of instructions that, when executed by the at least one computer processor, perform a method of predicting associations between data in a plurality of modalities (e.g., disease-associated gene signature, drug perturbation, and clinical data) using a statistical model trained to represent links between data having a plurality of modalities, the statistical model comprising a plurality of encoders and decoders, each of which is trained to process data for one of the plurality of modalities, and a joint-modality representation coupling the plurality of encoders and decoders.

[0089] In some embodiments, the various embodiments of the present disclosure are associated with a plurality of computers or computer systems that operate in concert to perform a method as described herein. For example, in some embodiments, a plurality of computers (e.g., connected by a network) may work in parallel to collect and process data, e.g., in an implementation of cluster computing or grid computing or some other distributed computer architecture that relies on complete computers (with onboard CPUs, storage, power supplies, network interfaces, etc.) connected to a network (private, public, or the internet) by a conventional network interface, such as Ethernet, fiber optic, or by a wireless network technology.

[0090] In some embodiments, the computer system comprises a data receiving component configured to receive input transcriptomic data. For example, the computer system may receive transcriptomic data from patients with the target disease, without the target disease, or aAttorney Docket Number: UM-43517.601

[0091] combination thereof. Alternatively or additionally, the data receiving component may be configured to receive input other gene expression signatures associated with the target disease. In some embodiments, the data receiving component may be configured to receive clinical data. In some embodiments, the data receiving component comprises data storage and management capabilities.

[0092] In some embodiments, the computer system comprises an analysis module. In some embodiments, the analysis module identifies disease-associated gene signatures for a target disease from the input transcriptomic data and gene expression signatures as received by the data receiving component. In some embodiments, the analysis module selects candidate drugs for treatment of a target disease, e.g., by using drug perturbation data or in silico perturbation data. The analysis module may apply in silico perturbagens to gene signatures using a pre-existing drug-expression profile database to identify drags candidates associated with reversal of gene signatures. The analysis module may calculate one or more of: a potency score; a selectivity score; genetic perturbation score; a class score; and a summary score of a given drag(s) or drag class(es) based on drug perturbation data or in silico perturbation data. Exemplary calculations for the above listed scores is provided in the examples.

[0093] In some embodiments, the analysis module characterizes candidate drugs related to a target disease by analyzing data contained in electronic health records. The analysis module may identify a group of individuals with the specified disease state, e.g., from a clinical database of health records, who were and were not exposed to a given drug class for a specified period of time and calculates the odds ratio for a defined disease-related outcome among individuals with and without exposure to the drag class and adjusting for other medical data on the group of individuals such as demographics. In some embodiments, the computer system or analysis module comprises a modifier module that groups candidate drags into drag classes based on standard mechanisms of action of each candidate drug.

[0094] The analysis module may also generate results and / or reports based on the analysis. The results or reports may include any or all of: lists of genes having differential expression for the target disease as categorized by significance of differential expression, implicated biological pathways, cell type, disease state, type of regulations each differentially expressed gene; results of in silico perturbation data; summary tables or spreadsheets including scores (e.g., potencyAttorney Docket Number: UM-43517.601

[0095] scores, selectivity scores, genetic perturbation scores, class scores, summary scores) of a given drug(s) or drug class(es); lists of candidate drugs ranked by the above listed scores; and disease remission or improvement metrics following treatments of target disease(s) with candidate drugs.

[0096] The computer systems and methods may find use with a variety of artificial intelligence or machine learning (AI / ML) systems including, but not limited to, neural networks (e.g., with reinforcement learning through human feedback), large language models (LLMs) or small language models (SLMs), large sequence models, fine-tuned large language models, an ensemble of machine learning models of any permutation of the above, and causal inference models, retrieval augmented generation (RAG) techniques (e.g., where prompts and information generated by LLMs is combined with data retrieved directly from databases and other data sources), agentic Al (e.g., where a system comprised of multiple components work together to autonomously set goals, create action plans, and perform workflows towards achieving them), and Mixture of Experts (MOE) techniques (e.g., where multiple smaller models trained to have expertise in specific fields are employed over a single monolithic model, a gating network is used to route inputs to the expert models and weigh their responses, and a collaboration / answer model is used to combine outputs into a single cohesive response).

[0097] In some embodiments, the computer system comprises a communication component. In some embodiments, the communication component communicates information from any one component of the system (e.g., the data receiving component) to any other component of the system (e.g., the analysis component), between a component and a sub-component of the system (e.g., analysis module and modifier module), or between a component outside of the system (e.g., instrument(s) generating gene signature data). In some embodiments, the communication component communicates information to or from the data receiving component. In some embodiments, the communication component communicates information from the data receiving component to the analysis component.

[0098] In some embodiments, a portion or all of the communication component is wired. In some embodiments, a portion of or the entire communication component is wireless. Any desired wireless communication technology may be employed, including but not limited to, electromagnetic wireless telecommunications (e.g., wireless networking, cellular, satellite), and electromagnetic induction (such as light, magnetic, or electric fields or the use of sound). WhereAttorney Docket Number: UM-43517.601

[0099] wireless networks are employed, any desired protocol can be used (e.g., ZigBee, EnOcean, Personal area networks, Bluetooth, TransferJet, ultra-wideband).

[0100] Any of a variety of computing devices may be used in the computer systems. Examples of computing devices are personal computers, digital assistants, personal digital assistants, cellular phones, mobile phones, smart phones, digital tablets, laptop computers, and other processor-based devices. In general, the computing devices related to aspects of the technology provided herein may be any type of processor-based platform that operates on any operating system, such as Microsoft Windows, Linux, UNIX, Mac OS X, etc., capable of supporting one or more programs or applications capable of carrying out one or more parts of the methods described herein. Some embodiments comprise a personal computer executing other application programs (e.g.. applications). All such components (e.g., computing devices and systems) described herein as associated with the technology may be logical or virtual.

[0101] EXAMPLES EXAMPLE 1

[0102] The above described methods integrate data from scRNA-seq, drug perturbation datasets (like LINCS), and EHR data to identify candidate drugs for repurposing. Transcriptomic signatures from single-cell RNA-seq data identify disease-associated gene signatures and drug perturbation data identifies candidate drugs that reverse the disease-associated gene signatures. Finally, efficacy of these drug candidates is evaluated by analyzing clinical outcomes using clinical data (e.g., EHR data). Clinical data are used to weigh the efficacy of the drug classes in real-world settings, with outcomes like disease remission or improvement used as key metrics. This integration creates a comprehensive framework for identifying potential treatments.

[0103] To find the drugs that target selected differentially expressed genes (DEGs), the cloudbased Connectivity Map Linked User Environment (CLUE) platform was used for the analysis of perturbational datasets (Subramanian A, et al., Cell, 171(6), 1437-1452. M7, incorporated herein by reference). CLUE outputs a list of drugs that reverse the up and down regulated genes, essentially measuring the changes in gene expression that occurs when cells are exposed to manyAttorney Docket Number: UM-43517.601

[0104] perturbing agents (e.g. drugs). CLUE is hosted by the Library of Integrated Network-based Cellular Signatures (LINCS) project.

[0105] Several factors were considered in identifying drug repurposing candidates. The potency score (Sx di) of a given drug z in the target cell line generated from the CLUE server was calculated using equation 1.

[0106]

[0107] (1)

[0108] CSCLUE is the connectivity score of drug i in the target cell line generated from the CLUE server. Lor example, the closer the CS is to -100, the greater the chance the drug has of reversing upregulated differentially expression genes.

[0109] The input into the CLUE query system was a list of up (qup) and down (qdown) regulated genes that were compared to the reference database (Touchstone, r) in LINCS using the weighted Kolmogorov-Smirnov enrichment statistic as shown in equation 2.

[0110]

[0111] wq,r is the weighted connectivity score which represents the similarity measure between a query q (qup, qa own ) and a reference signature r. ESup is the enrichment of qupin r and ESdown is the enrichment of qdown in r.

[0112] The selectivity score (S2 di) of drag i was calculated using equation 3.

[0113] (3)

[0114]

[0115] TASi is the average transcriptional impact score of drug i which describes its activity, relative to all other drugs, as derived from its replicate reproducibility and magnitude of differential gene expression.

[0116] The genetic perturbation score (S3 di) of drug i, which represents the averaged connectivity score of knocking down the target genes,j~), was calculated using equation 4.

[0117]

[0118] ii&(4)Attorney Docket Number: UM-43517.601

[0119] CS(TGj)i is the connectivity score of knocking down target gene j by drug i.

[0120] The class score (S4 di) of drug i, or the number of drugs that have negative connectivity scores that belong to the same class as drug i, was calculated using equation 5.

[0121] < ■■■ >

[0122]

[0123] (5) negRdass of drug i is the ratio between the number of drugs that have a negative score and the total number of drugs belonging to the class of drug i, as shown in equation 6.

[0124] ,<>■. ■ ~...

[0125]

[0126] HNEG is the number of drugs that have a negative connectivity score and belong to class of drug i. npos is the number of drugs that have positive connectivity score and belong to class of drug i.

[0127] Each of the factors were quantitatively scored and summed to rank the most promising candidates, as in equation 7.

[0128] '>,:$x (7)

[0129]

[0130] Clinical outcomes from EHR data were used to weigh the efficacy of the high scoring drug candidates and overall drug class in real-world settings, with outcomes like disease remission or improvement used as key metrics.

[0131] Initially, a clinical dataset is filtered to include patients with the target disease state, creating a foundational dataset titled “Patients with Disease State” that includes patient IDs, covariates (e.g., demographics, diagnoses), and clinical outcomes. The cohort is further refined by selecting patients with adequate duration of drug therapy, such as exposure to a drug class for at least one-third of the patient’s length of stay in the hospital or another user-specified threshold. Patients are grouped into those exposed and unexposed to each drug class from the “Ranked Drug Classes” table, with this process repeated for all drug classes being evaluated. Each group is analyzed for the occurrence of disease-related outcomes, such as reductions in serum creatinine levels for acute kidney injury, reductions in blast counts for hematologic malignancies, or decreases in HbAlc for diabetes. These outcomes are modeled using generalized linear models (GLMs) of the form:Attorney Docket Number: UM-43517.601

[0132] log i —

[0133]

[0134] : — | 4-xDrug Expos ure ■+■ & x Age - #3 x Sex 4- >,. - x Covanates \ l - p /

[0135] where p represents the probability of achieving the desired clinical outcome, P 1 quantifies the effect of drug exposure, P2, P3,.. and Pk account for the influence of age, sex, and other relevant covariates. The GLM outputs odds ratios epithat describe the likelihood of achieving the desired outcome for patients exposed to a given drug class compared to those unexposed, adjusting for other covariates. For each drug class, odds ratios, unadjusted p-values, and adjusted p-values are computed and summarized. Results from individual analyses are aggregated into a final data table, where drug classes are ranked based on their efficacy, as reflected by statistical significance and effect sizes. This comprehensive process provides a robust framework for identifying drug candidates for repurposing in real-world clinical settings.

[0136] EXAMPLE 2

[0137] Drug Repurposing for Acute Kidney Injury in COVID-19 Patients Using scRNA-seq and Electronic Health Record Data

[0138] The COVID-19 pandemic caused by the Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2) has had devastating impacts across the world, contributing to nearly 7 million deaths worldwide (WHO). Patients hospitalized with COVID-19 can develop acute kidney injury (AKI), which is a predictor of further complications and mortality (Rubin, Chao, Yoo, Ostermann). Studies have demonstrated that anywhere from approximately 30% to nearly half of patients hospitalized with COVID- 19 develop AKI (henceforth referred to as COVID-AKI), with a mortality rate of up to 50% and need for dialysis of 19% (Chan, Hilton, Zhang). In addition, 35% of patients do not experience renal function recovery by the time of discharge (Chan). Possible mechanisms of COVID-AKI include hypotension, vascular thrombosis and endothelial injury, tubular necrosis, secondary infection, glomerular jury, and exposure to nephrotoxic drugs (Hilton, Chan). Covid-associated nephropathy has also been recognized as a novel form of collapsing glomerulopathy following COVID- 19, particularly among patients with high-risk APOL1 alleles (Velez). Current treatment modalities for COVID-AKI rely heavily on reducing further renal damage and kidney replacement therapies (Hilton).Attorney Docket Number: UM-43517.601

[0139] Given the high rate of morbidity and mortality associated with COVID-AKI and the absence of existing targeted drug therapies, it is therefore imperative to identify therapeutic options to treat this devastating complication.

[0140] In this example, we report a novel multi-modal approach to drug repurposing in the treatment of COVID-AKI. The first step of our approach is to utilize single cell RNA-seq (scRNA-seq) data to identify transcriptomic signatures associated with COVID-AKI. ScRNA-seq allows us to study cellular changes in various disease states through the quantification of gene expression within individual cells (Xu). As a high-resolution molecular technology, scRNA-seq has been used or is currently being applied in many studies attempting to elucidate molecular mechanisms associated with disease and in pursuit of drug repurposing in areas such as Alzheimer’s disease, various cancers, multiple sclerosis, and Crohn’s disease (Gupta.

[0141] Shevtsov, Mo, Kwak). The use of urine transcriptomics in the case of kidney research offers the additional benefit of a non-invasive method in the field of personalized medicine. The second step of our approach is to identify drug candidates that reverse the COVID- AKI-associated gene signature through the use of drug perturbation data available in Library of Integrated Networkbased Cellular Signatures (LINCS), an NIH funded project that measures changes in gene expression that occur in the setting of various diseases and drugs (LINCS). LINCS provides the opportunity to examine drug perturbation data on target cell types, including kidney epithelial cells, to understand and rank drugs by their ability to reverse specific gene signals (LINCS, Alakwaa). Lastly, we assess the efficacy of these candidate drug classes in the treatment of COVID-AKI using real- world electronic health record data. With respect to COVID-AKI, the National COVID Cohort Collaborative (N3C) EHR dataset offers an enticing opportunity to understand the effects of various drugs on AKI therapy. N3C is an NIH National Center for Advancing Translational Sciences (NCATS)-sponsored program that collects electronic health record data, including on hospital admissions and laboratory results, from over 80 sites across the United States and over 20 million patients dating back to January 1, 2018 (N3C). Institutions contributed EHR data to the N3C consortium using several different data models including OMOP, PCORnet, and PEDSnet, and mapped all data to the OMOP Common Data Model (V5.3.1) (N3C).Attorney Docket Number: UM-43517.601

[0142] METHODS:

[0143] LINCS

[0144] Once we identified a list of up and down-regulated genes among patients with AKI, we then sought to find candidate drugs that reverse the selected differentially expressed genes (DEGs) using a bioinformatics pipeline previously trailed on a study that sought to repurpose drugs for tamoxifen-resistant breast cancer (Alakwaa). This pipeline required the use of the cloud-based platform CLUE version 1.1.1.1. CLUE which is hosted by the Library of Integrated Network-based Cellular Signatures, an NIH funded project that measures changes in gene expression that occur in the setting of various diseases and drugs (LINCS). For our reference dataset, we used the L1000 LINCS dataset containing 1,328.098 gene expression profiles secondary to applying 42,553 perturbagens for a total of 476,251 signatures (Subramanian). We also utilized CLUE API RESTflu web services, which offer programmatic access to annotations and perturbational signatures in the LI 000 dataset (LINCS). For our cell line, we selected the HA1E line which in CLUE incorporates 8,668 gene signatures [LINCS],

[0145] We first provided our list of up and down-regulated genes as the input into CLUE, which were compared to the reference database in order to generate a connectivity score. We then applied a drug prioritization system as described in Alakwaa (Alakwaa, herein incorporated by reference). This entailed calculating 4 metrics, (1) the potency score of a given drug in the HA IE cell line generated from the CLUE server, (2) the selectivity score of the drug, (3) the genetic perturbation score of the drug which represents the averaged connectivity score of knocking down the target genes of the drug, and (4) the class score, or the number of drugs that have negative connectivity scores that belong to the same class as a given drug. We then summed these four metrics in order to determine the rank of the drug, or the degree to which the drug reverses the gene signals associated with AKI, and our output is shown in Table 1.

[0146] TABLE 1

[0147] no id name description target Drug_Tar Class Drug_ total_ Dif ferenc genes getGene_ Score Score score e in score score due to different class

[0148]

[0149] scoreAttorney Docket Number: UM-43517.601

[0150] 1 BRD- XMD-1150 leucine LRRK2 0. 9 / 27 1 0. 97 8 0. 985 b 0

[0151] K0143 63 66 rich repeat 3 66667

[0152] kinase

[0153] inhibit or ( L

[0154] RRK2

[0155] inh ibit © i

[0156] BRD- pyrimidine 0. 90 4 0. 962 7

[0157] 2 K8214371 6 f lucyt o s ine ana l og DNMT 1 0. 9842 1 1 66667 0

[0158] FAAH RRD- reupt a k e 0. 997 0. 0 833333 K37865504 LY-218324 0 inh ibit or FAAH 0. 9386 0. 7 5 9 0. 8 955

[0159] 0. 666

[0160] BRD- p 1 f 11 h r i n - HSP HSPA1A, 6 6666 0. 994 0. 8 801

[0161] 4 K96799727 mu mh ibit or TP 53 0. 97 975 1 72222 Ti BRD- XMD-385 leucine LRRK2, 0. 9727 0. 675 0. 991 0. 3 7 98 0. 1033333 K64857 S 43 r i c h repeat MARK 7 666 67 J 3

[0162] kinase

[0163] inhibit or

[0164] ( LRRK2

[0165] inhibit o )

[0166] SSTR1,

[0167] SSTR2,

[0168] SSTR3,

[0169] SSTR4,

[0170] SSTR5,

[0171] GF1,

[0172] s omat o st at i GFR,

[0173] BRD- s omat o st ac i n recept or OPRD 1, 0. 927 0. 8743

[0174] 6 K14 6818 67 agoni st OPRM1 0. 6957 1 4 66667 0

[0175] 0. 823

[0176] BRD- mTOR 5294 1 0. 997 0. 83 60

[0177] 7 K67 566344 KU-00 637 94 inh ibit or MTOP 0. 6873 2 2 0 9804 0

[0178] 0. 823

[0179] BRD- mTOR 5294 1 0. 994 0. 8350

[0180] S K699324 63 AZD-8055 inh ibit or MTOR 0. 6873 2 43137 l'l capi l lary AKR1C3,

[0181] BRD- st ab 111 z mg AKR1B1, 0. 7184 66 0. 982 0. 333 6 C1. C1666666 9 K204820 99 r u L m agen t F 10 0. 8 4 22222

[0182] 0. 823

[0183] BRD- ml OR 52941 0. 98 9 0. 8335

[0184] 10 K94294671 OS I-027 inhibit or MTOR 0. 6873 7 0 9804 0

[0185] 0. 823

[0186] BRD- mTOR 52941 0. 971 0. 8274

[0187] 1 1 K68174511 t onn-2 inhibit or MTOR 0. 687 3 2 7 6471 0

[0188] avramvi l la nucleopho sm 0. 571

[0189] BRD- mide- in 42857 0. 992 0. 82 69

[0190] 12 K3 9569857 analoq-3 inhibit or NFM1 0. 917 1 3 0 9524 0

[0191] serotonin 0. 487

[0192] BRD- recept or MAO A, 17948 0. 999 0. 8181

[0193] 1 3 K1 65514 C 1 PNU-223 94 agon i st MAOB 0. 9679 7 4 5 982 9 0

[0194] imi daz o l ine

[0195] BRD- recept or MAO A, 0. 978 0. 8155

[0196] 1 4 K55344148 BU-224 l icand MAOB 0. 9679 0. 5 7 33333 0

[0197] calc ium 0. 559

[0198] BRD- channe l MAOB, 4104 3 0. 904 0. 810 6 0. 146863 1 1 5 K84 63 9753 s a t i n am i de mode 1 at o r SRC 6 A3 0. 9679 1 5 03477 9

[0199] neur opept id 0. 428

[0200] BRD- 57142 0. 975 0. 8 (J 13

[0201] 1 6 K1207 98 98 PD-1 60170 ant agon i st NPY1R 0. 9998 7 57143

[0202] BRD- NAMP T 0. 7 97 6

[0203] 1 7 K8328 913 1 CAY- 10 618 inh ibit or NAME’ T 0. 4 1 0. 993 66667 0

[0204] 0. 736

[0205] BRD- gluc okinase 84210 0. 97 9 0. 75 37 0. 08771 92 1 8 K 96670504 lonidamine inh ibit or GCK 0. 5453 c 1 473 68 98

[0206] PDGFR alpha 0. 4 92

[0207] BRD- and c-Kit 0 634 9 0. 928 0. 7515 0. 1 693121

[0208]

[0209] 1 9 K47150025 KI-8751 inh ibit or KDR 0. 8 344 3 373 31 69Attorney Docket Number: UM-43517.601

[0210] CSNK1E,

[0211] CSNK1A1

[0212] 0. 678

[0213] BRD- tubul in CSNK1D, 57442 0. 978 0. 7438 0. 0595238 20 K0 96383 61 SA- 63133 inhibit or CSNK1G2 0. 57455 9 5 7384 1

[0214] Brut on1s

[0215] tyro s ine

[0216] k inase

[0217] BRD- terre ic- ( BTK ) 0. 999 0. 7408

[0218] 2 1 A6422845 1 inh ibit or BTK 0. 2233 1 2 33333 0 BRD- PKA PKIA, 0. 980 0. 7355

[0219] 97 K1 8742343 H-8 inh ibit or PRKACA 0. 22 65 1 1 33333 o SCNN1A,

[0220] c odrum SCNN1B, 0. 41 6

[0221] BRD- channe l SCNN1G, 66666 0. 985 0. 7 347

[0222] 23 K9204 95 97 t r iamt erene blocker 3 CNN ID 0. 3022 3838 9

[0223] recept or

[0224] tyro s ine

[0225] p r o t e i n 0. 333

[0226] BRD- tyrpho st m- kinase 0. 98 9 0. 7248

[0227] 24 K87 91 973 9 AG-825 inhibit or ERBB2 0. 8546 11111 0

[0228] DNA 0. 785

[0229] BRD- repl icat ion 74428 0. 7242 0. 0714285 2 5 K35 9605 C 2 niclo s amide inhibit or STAT 3 0. 383 0. 974 380 95 71

[0230] PGR,

[0231] CYP 17A1

[0232] NR3C2,

[0233] CATSPER

[0234] 1,

[0235] CATSPER

[0236] 2,

[0237] CATSPER CATSPER

[0238] 4,

[0239] CYP 2C 1 9

[0240] progest eron, ESRI, 0. 692

[0241] BRD- proge s teron e recept or OFRK1, 3 07 69 0. 91 6 0. 7 127

[0242] 2 6 K64 994 963 agoni st TRPC5 0. 52 95 69231

[0243] CDC25A, 0. 714

[0244] BRD- CDC C D C 25 B, 0. 423133 28571 0. 98 6 0. 7 C' 80

[0245] 27 K0310 94 92 NSC- 663284 inhibit or CDC25C 4 7 3 9683 0 BRD- BMX BMX, 0. 998 0. 6990 0. 041 6 b 66 28 U8 69221 68 QL-XI I-47 inhibit or BTK 0. 2233 0. 875 9 66667 67

[0246] VKORC T,

[0247] CYP 2C 4 9

[0248] 0. 7 93

[0249] BRD- vitamin CYP 2C 8, 65079 0. 997 0. 6965 0. 0 687830 2 9 A245145 65 warf arin inh ibit or CYP 4F2 0. 2 986 4 5 835 98 69

[0250] aryl

[0251] hyd r oca rbon 0. 8 1 1

[0252] BRD- recepto r A ER, 1 1 1 1 1 0. 98 1 0. 6740 0. 0 62 962 9 30 K01815685 rndc le acTon i st IDO1 0. 22 97 1 4 70 37 63

[0253] NFKB BRD- pathway NFKB1, 0. 999 0. 6685

[0254] 3 1 K074035 93 CAY 10470 inh ibit or TNF 0. 00 6 1 6 33333

[0255] KCNJ11,

[0256] ABCC 8,

[0257] CFTR,

[0258] KCNJ5,

[0259] KCNJ8,

[0260] ABCA1,

[0261] ABCB11,

[0262] ABCC 9,

[0263] CP T 1A,

[0264] KCNJ1,

[0265] BRD- gl ihenclami sul f onylure SLCO2B1 0. 7 65 0. 98 6 0. 6 661

[0266]

[0267] 32 K3 692723 6 de, TRPA1 0. 24585 625 9 25 0. 078425Attorney Docket Number: UM-43517.601

[0268] 0. 811

[0269] BRD- mTOR MTOR, 7 6470 0. 998 0. 6634 0. 003 9215 33 A454983 68 WYE-125132 inhibit or P IK3CA 0. 1801 4 215 69 69

[0270] A0X1, 0. 666

[0271] BRD- HSR HSP 90AA o 6 o 66 0. 982 0. 6602

[0272] 34 K51 967704 B I IB021 inh ibit or 1 0. 331 5 7 7 8888 9 0

[0273] ABCC 8,

[0274] KCNJ1 0,

[0275] BRD- sul f ony lure KCNJ1 1, 0. 756 0. 964 0. 65 60

[0276] 35 K80396088 g 1 i ci ] i d one KGNJ8 0. 2471 25 9 83333 0. 081 25

[0277] 0. 666

[0278] BRD- HSR HSP 90AA 66666 0. 968 0. 655 6

[0279] 3 6 K3 652 9613 PU-H71 inh ibit or 1 0. 3 315 7 7 fi BRD thro mb 1 n 0. 991 0. 654 1 K73824 630 skatole inh ibit or F2 0. 471 I-,, 5 4 33333 I l 0. 811

[0280] BRD- mT OR MTOR, 7 6470 0. 957 0. 64 97 0. 003 9215 38 K40175214 t orm-1 inhibit or P IK3CA 0. 1601 6 88235 69

[0281] ErbB2

[0282] tyro s ine 0. 132

[0283] BRD- kina se EGFR, 73809 0. 975 0. 6400 0. 122420 6 3 9 K7 99301 C 1 GW-583340 inhibit or ERBB2 0. 31155 5 8 2 93 65 35

[0284] ATP 1A1,

[0285] ATP 1A2,

[0286] ATP 1A3,

[0287] ATP 1A4,

[0288] ATP 1B1,

[0289] ATP 1B2,

[0290] a lpha ATP 1B3, 0. 8 92

[0291] BRD- subun it ATP 1B4, 95714 0. 633 1 0. 035 7142 4 0 A68 930007 ouaba in binder FXYD2 0. 010625 0. 9 96 60714 8 6

[0292] e s t rogen

[0293] BRD- 17-bet a- recept o r 0. 993 0. 6325

[0294] 4 1 K 66766661 est radio l agoni st ESRI 0. 52 95 0. 375 6 6667 1'1

[0295] e s t rogen

[0296] BRD- e st radiolrecept or 0. 987 0. 6305

[0297] 42 A3 9747742 valerate agoni st ESRI 0. 52 95 0. 375 1 33333 1'1

[0298] MAPK7,

[0299] DCLK2,

[0300] LRRK2, 0. 535

[0301] BRD- BMK PLK4, 71428 0. 996 0. 6281 0. 154761 9 43 K50387473 XMD-8 92 inhibit or TNK1 0. 35225 6 4 2142 9 05

[0302] ATP 1A1,

[0303] ATP 1A2,

[0304] ATP 1A3,

[0305] ATP 1A4,

[0306] ATP 1B1,

[0307] ATP 1B2,

[0308] ch l oride ATP 1B3, 0. 8 92

[0309] BRD- channe l ATP 1B4, 85714 0. 980 0. 6280 0. 0357142 4 4 A80502530 c i nobu f agi n a o.t i vat o r FXYD2 0. 0 1 0625 3 6 2738 1 8 6

[0310] valo s in

[0311] cont a m mg 0. 722

[0312] BRD- p r o t e i n CYP 1 9A1 0. 62 6 6 0. 0 925925 4 5 K77390737 xanthohumo 1 inh ibit or, VCR 0. 1 976 2 0. 9 6 07407 93

[0313] MTOR,

[0314] P IK3CA,

[0315] P IK3CG,

[0316] P IK3CD, 0. 721

[0317] BRD- mTOR ATR, 73202 0. 991 0. 61' 4 9 0. 033 9324 4 6 K1218491 6 NVB-BEZ 235 inhibit or P 1K3CB 0. 10165 6 60 675 62

[0318] e st rogen 0. 520

[0319] BRD- recept or ESRI, 83333 0. 6025 0. 048 6111 47 A18 6209 C 0 e st rio l antagoni st ESR2 0. 3 5075 3 0. 92 6 27778 11

[0320] progesteron PGR, 0. 666

[0321] BRD- e re cept or AR, 06666 0. 998 0. 5 67 6 0. 1111 11 1 48 A9479305 1 gest rmone ant agoni st ESRI 0. 03755 7 6 0555 6 11

[0322] PPAR BRD- pt ero st i lbe recept o r FAAH, 0. 24 67 66 0. 487 0. 953 0. 5 625

[0323]

[0324] 4 9 K 92870997 a gon i st PP AP A, 67 5 888 8 9 0. 0375Attorney Docket Number: UM-43517.601

[0325] P TGS 1,

[0326] P TGS2

[0327] P TGS2,

[0328] pro st ano id RTF, 0. 523

[0329] BRD- recept or PLA2G2E 80952 0. 911 0. 5 624 0. 0158730 50 K7 6775527 mmesul ide inh ibit or, P TGS 1 0. 252 4 5 3 6508 1 6

[0330] cyc l ooxygen

[0331] BRD- P TGS2, 0. 90 6 0. 5527

[0332] 5 1 K8256263 1 t o lmet m inh ibit or P TGS 1 0. 252 0. 5 2 33333 0

[0333] 0. 272

[0334] BRD- RAF 72727 0. 985 0. 541 8

[0335] 52 K01253243 SB-590895 inh ibit or BRA. F 0. 3 671 g 7 42424 f'l

[0336] AD RAIA,

[0337] ADRA2A,

[0338] ADRA1B,

[0339] ADRA1D,

[0340] ADRA2B,

[0341] ADRA2C,

[0342] unident i f ie CYP 2C 1 9

[0343] d

[0344] pharroaco log HTR1B, 0. 532

[0345] BRD- oxyrnet az o l i ical HTR1D, 62411 0. 950 0. 5 032 0. 0 0023 64 53 K1 6195444 ne act ivity HTR2C 0. 02 615 9 247 04 07

[0346] ADRA1A,

[0347] ADRA2A,

[0348] AD RAI B,

[0349] adrenergic ADRA1D,

[0350] BRD- phent o lamm recept o r ADRA2B, 0. 421 0. 999 0. 4 82 6

[0351] 54 K2131733 6 ant agoni st ADRA2C 0. 02 615 8 75 8 08333 0

[0352] 0. 333

[0353] BRD- ATM k ina se 33333 0. 93 9 0. 48 13

[0354]

[0355] 55 K15592317 CP 4 66722 inh ibit or ATM 0. 1 715 1 11 1 1 1 0

[0356] In Table 1, the drags are ranked based on total score, with the highest (best at the top, closest to "1"). Rows 1-39 are considered positive hits. When the total score is close to 1, it demonstrates higher reversal potential of the disease gene signature through scoring of the database (e.g„ gene score, class score, etc.).

[0357] One example from the above table is the drag XMD-1150 in the first row. This drag has the following values:

[0358] • Target gene(s): LRRK2

[0359] • Gene Score: 0.9727

[0360] • Class Score: 1

[0361] • Drug Score: 0.9783

[0362] • Total Score: 0.9837

[0363] • Difference in score due to different class score: 0

[0364] How these values are derived is in Table 1.5 below:Attorney Docket Number: UM-43517.601

[0365] TABLE 1.5

[0366] Metric Definition Formula / Calculation Average genetic perturbation score

[0367] Gene Mean of CLUE-based CS scores (connectivity score of knocking down target

[0368] Score for target genes (e.g., LRRK2) genes)

[0369] Class Ratio of drugs in the same class that have

[0370] n_NEG / (n_NEG + n_POS ) Score negative connectivity scores

[0371] Drug Connectivity score of the drug itself (from Score from CLUE API (closer to - Score LINCS / CLUE) 100 = stronger reversal) Total Average of Gene Score, Class Score, and Drug (Gene Score + Class Score + Drug Score Score Score) / 3

[0372] Change in Total Score if Class Score were Shown explicitly in table to Difference

[0373] different (used for sensitivity) highlight Class Score impact

[0374]

[0375] The total score closer to 1.0 indicate higher predicted ability to reverse the disease signature. There is no strict threshold, but drugs with Total Scores > 0.85 (e.g., top 5-10 drugs) would typically be prioritized for downstream steps such as: i) validation in model systems (in vitro / in vivo); ii) clinical association testing via EHR (e.g., N3C); iii) investigating compound availability / known side effects. This scoring system helps triage candidates and justify prioritization, especially when combined with downstream EHR-based validation.

[0376] N3C Cohort

[0377] We utilized data from N3C for validation of candidate drug classes in COVID-AKI therapy. Each N3C contributor site maintains a data transfer agreement approved by its IRB. The analyses reported in this study were separately approved by the IRB of each participating institution. The IRB reviews included a waiver of informed consent. Within the N3C enclave, we utilized PySpark and R version 3.5.0 for all data pre-processing and analysis. We froze the available de-identified N3C data on September 23, 2023 to allow for replication.

[0378] Patients with a history of COVID- 19 were defined as those patients documented in the “Covid positive persons" catalog table curated directly by N3C, all of whom had positive PCR or antigen test or a COVID- 19 diagnostic code within 7 days of hospitalization [N3C], Patients with AKI, CKD, history of kidney transplant, hypertension, type 2 diabetes, or cardiovascular disease were identified using the codeset IDs 692484006, 889975596, 2103567, 343580462, 54654297, and 904688609 respectively, in the Observational Medical Outcomes PartnershipAttorney Docket Number: UM-43517.601

[0379] (OMOP) concept set tool [N3C], Patients with a history of kidney transplant were identified using the concept ID 42539502 in the OMOP concept set tool [N3C], Length of stay for each patient was defined as the difference between the start and end dates of admission following COVID positivity using the “microvisits_to_macrovisits” catalog table provided by N3C. We followed the standard OMOP canonical unit standard for serum creatinine, in which we excluded measurements greater than 30 mg / dL or less than 0 mg / dL [N3C],

[0380] For each drug class, we defined exposure to a drug class as having been treated with a member of the drag class for at least one third of the patient’s length of stay. Patients who were not treated with the drag class during their hospitalization or who were treated for less than one third of their length of stay were labeled as unexposed. We defined the outcome of AKI recovery as a decrease in serum creatinine prior to discharge from the baseline serum creatinine of at least 33% within 7 days, using KDIGO criteria for definition of AKI [Neyra], For each drug class, we matched patients who were exposed to patients who were unexposed 1:1 using the Matchit package version 4.5.5 [Ho] on age, sex, race, ethnicity, and diagnosis of hypertension, type 2 diabetes, or cardiovascular disease present in the “condition_occurrence” table in N3C [N3C], Lastly, we built generalized linear models in R version 3.5.0 using the glm function present in base R modeling the outcome, or whether or not they experienced AKI recovery, with the exposure as the predictor and a binomial distribution. We then examined the odds ratio derived from each generalized linear model for each drug class along with the p-value to identify whether or not exposure to the drag class was associated with AKI recovery.

[0381] RESULTS:

[0382] As shown in Figure 9, we utilized a multilevel approach to detect gene signatures associated with COVID-AKI, identify candidate drugs that reverse the gene signatures associated with COVID-AKI based on perturbation studies, and validate the utility of these candidate drags for COVID-AKI therapy in a large retrospective cohort.

[0383] Validation in N3C Cohort

[0384] Next, we sought to determine if drugs associated with reversal of COVID AKLassociated gene signals can successfully be used to treat COVID-AKI in a clinical setting. We utilizedAttorney Docket Number: UM-43517.601

[0385] retrospective data from N3C which included hospital admissions and laboratory results from over 20 million patients. A summary of our pre-processing steps is shown in Figure 10.

[0386] We first selected patients in N3C who were documented to have had both COVID and an AKI. Given our focus on patients with AKI secondary to COVID, we filtered out patients with a history of prior AKIs, chronic kidney disease, or kidney transplant who may be predisposed to AKI and proceeded with the remaining 224,994 patients [Liu, Dudreuilh], These patients may also be on prior therapies that preclude accurate calculation of drug exposures [Neyra, Yoon, Liu, Ostermann], In order to identify patients with potentially a COVID-associated AKI (hereby referred to as COVID- AKI), we then filtered to patients with an AKI diagnosis within 30 days following COVID diagnosis (n=101,756). The median number of days from COVID diagnosis to AKI diagnosis was 34 (IQR=145 days).

[0387] In our next pre-processing step, we included only patients with a hospital stay following COVID diagnosis (n=34,665), as inpatient stay allows monitoring of AKI resolution and accurate drug exposure, as well as that patients requiring hospital stay for COVID-19 were more likely to have developed complications such as AKI [Matsuomoto], Subsequently, we filtered to only those patients with a length-of-stay, or time from admission to discharge, within 5 to 90 days (n=25,458). Length of stay fewer than 5 days may be unlikely to provide sufficient time to monitor drug response and patients with length-of-stay over 90 days may either reflect severe complications of COVID-19 infection that contribute to prolonged AKI or clerical errors. These thresholds were selected using a density plot which shows that the median length of stay was 9 days (IQR=13 days).

[0388] Subsequently, we filtered our cohort to those patients with at least 2 serum creatinines recorded during their hospital stay. A minimum of 2 creatinines are required to assess for trend in creatinine during hospital admission, including determining whether AKI resolved prior to discharge. The median number of serum creatinine recorded during admission for all patients with COVID-AKI was 22 (IQR=28).

[0389] We then filtered this cohort to only include patients with at least 1 serum creatinine within the first 7 days from admission and 1 in the last 7 days prior to discharge (n=20,864). The presence of a serum creatinine recorded close to admission would allow us to determine baseline creatinine associated with COVID, and a serum creatinine recorded close to discharge would allow us to assess for AKI recovery prior to discharge. For example, patients with their firstAttorney Docket Number: UM-43517.601

[0390] serum creatinine recorded several days after admission may have other factors during hospitalization beyond COVID that contribute to AKI. A The density plot showed that, on average, patients had a serum creatinine recorded on the day of admission and the day of discharge.

[0391] For each patient, we then defined baseline serum creatinine as either the serum creatinine on admission or the serum creatinine on date of AKI diagnosis, whichever was higher. For patients who had more than 2 serum creatinines recorded during admission, we filtered to include only the baseline serum creatinine and the serum creatinine closest to discharge when assessing for AKI recovery. For our last data pre-processing step, we retained only those 10,926 patients whose AKI persisted at least 7 days after admission, as shown in Figure 10, as these patients did not recover within 7 days per KDIGO criteria and are likely to benefit from targeted drug therapy [Neyra], The characteristics of our final cohort of patients including demographics and rates of hypertension and type 2 diabetes diagnoses are shown below in Table 2.

[0392] TABLE 2

[0393] N or Mean (% or SD)

[0394] Age (Years) 64.4 (15.7)

[0395] Gender

[0396] Male 6741 (61.7%)

[0397] Female 4182 (38.3%)

[0398] Not Reported 3 (0.03%)

[0399] Race

[0400] White 6857 (62.8%)

[0401] African American 2681 (24.5%)

[0402] Asian or Pacific Islander 76 (0.7%)

[0403] American Indian or Alaskan Native 61 (0.6%)

[0404] Multiracial 4 (0.04%)

[0405] Unknown 1247 (11.4%)

[0406] Ethnicity

[0407] Not Hispanic or Latino 9350 (85.6%)

[0408] Hispanic or Latino 910 (8.3%)

[0409]

[0410] Attorney Docket Number: UM-43517.601

[0411] Other / Un known 666 (6.1%) Hypertension Diagnosis

[0412] Yes 8904 (81.5%)

[0413] No 2022 (18.5%)

[0414] Type 2 Diabetes Diagnosis

[0415] Yes 5813 (53.2%)

[0416] No 5113 (46.8%) Cardiovascular Disease Diagnosis

[0417] Yes 8950 (81.9%)

[0418] No 1976 (18.1%)

[0419]

[0420] In order to focus on identifying and evaluating drug candidates that may reverse COVID-associated AKI, we first pooled all of the individual drugs from our scRNA-seq analysis into drug classes based on mechanism of action (Table 1). We then defined exposure to a drug class as if the start and end dates of the drug exposure class were greater than or equal to one-third of the patient’s LOS. Patients in N3C met exposure criteria to 13 of 18 drug classes.

[0421] Glucocorticoid receptor agonists had the greatest number of patients who were treated (n=6,782), followed by acetylcholine receptor antagonists (n=5863) and ARBs (n=2138). PI3K inhibitors had the fewest number of patients who were treated (n=2) followed by tubulin polymerization inhibitors (n=37) and PPAR agonists (n=209). Conversely. PI3K inhibitors had the greatest number of patients who did not have any exposure and were therefore labeled as untreated (n=9,793), while glucocorticoid receptor agonists again had the fewest number of patients without any exposure (n=l,198).

[0422] For each drug class, we performed 1:1 propensity score matching of patients who were treated with a member of the drug class for at least one third of their length of stay to those who were not. The matching variables included age. gender, race, ethnicity, and diagnosis of hypertension, type 2 diabetes, or cardiovascular disease. We then built generalized linear mixed models comparing the odds of recovery from AKI as defined in Methods by exposure status to each drug class. As shown in Table 3, we found that patients exposed to acetylcholine receptor antagonists or glucocorticoid receptor agonists during their hospitalization were significantly more likely to recover from COVID-AKI compared to patients who were not exposed to these drug classes.Attorney Docket Number: UM-43517.601

[0423] TABLE 3

[0424] Acetylcholine 109 1673 1673 1.459 [1.201, <o.oor Receptor Antagonists 1.776] Glucocorticoid 86 1198 1198 1.323 [1.046, 0.020* Receptor Agonists 1.678]

[0425] Opioid Receptor 7 1043 1043 1.054 [0.829, 0.667 Antagonists 1.342] Peroxisome 1 209 209 1.000 [0.590, 1 Proliferator-Activated 1.694] Receptor (PPAR)

[0426] Agonists

[0427] Protein Synthesis 7 682 682 0.964 [0.708, 0.814 Inhibitors 1.311]

[0428] DNA Inhibitors 10 1504 1504 0.924 [0.750, 0.457

[0429] 1.140] Angiotensin Receptor 7 2138 2138 0.857 [0.723, 0.076 Blockers (ARBs) 1.016] Dehydropeptidase 1 588 588 0.845 [0.615, 0.295 Inhibitors 1.158] Sodium-Glucose 4 551 551 0.821 [0.590, 0.240 Cotransporter-2 1.141] (SGLT2) Inhibitors

[0430] Glutamate Receptor 8 1195 1195 0.796 [0.631, 0.052 Antagonists 1.002] Remdesivir 1 1912 1912 0.779 [0.647, 0.008*

[0431] 0.938] Angiotensin 7 774 774 0.619 [0.456, 0.002* Converting Enzyme 0.837]

[0432] (ACE) Inhibitors

[0433]

[0434] Table 3 shows the number of COVID-AKI patients treated with each drug class and odds ratios, confidence intervals, and p-values from comparison of odds of AKI recovery by treatment with each class. Patients with and without exposure to therapy are matched 1:1 by age, sex, race, ethnicity, and diabetes, hypertension, or cardiovascular disease. Drug classes are sorted by descending odds of recovery from COVID-AKI. *p < 0.05.Attorney Docket Number: UM-43517.601

[0435] Specifically, acetylcholine receptor antagonists were associated with 1.459 times greater odds of recovery from COVID-AKI compared to matched controls (95% CI [1.201, 1.776], p < 0.001) and glucocorticoid receptor agonists were associated with 1.323 times greater odds of recovery from COVID-AKI compared to matched controls (95% CI [1.046, 1.678], p = 0.02). On the other hand, patients exposed to ACE inhibitors during their hospitalization were significantly less likely to recover from COVID-AKI compared to matched controls (OR 0.619, 95% CI [0.456, 0.837], p = 0.002). We also modeled COVID-AKI recovery among patients exposed to remdesivir, an antiviral agent used in the treatment of COVID-19 among hospitalized patients, to serve as a control measuring the effect of COVID- 19 recovery independent of COVID-AKI, although this drug was not one that reversed the gene signature associated with COVID-AKI in LINCS. Patients exposed to remdesivir were significantly less likely to recover from COVID-AKI compared to matched controls (OR 0.779, 95% CI [0.647, 0.938], p = 0.008).

[0436] In this example, we utilized a multi-modal drug repurposing pipeline incorporating scRNA-seq, LINCS, and EHR data from N3C to identify candidate drug classes for COVID-AKI therapy. Our use of scRNA-seq identified several genes associated with COVID-AKI, with our subsequent drug perturbation analysis in LINCS identifying 39 drugs that reverse the gene signature associated with COVID-AKI. Ultimately, these drugs were pooled into 13 drug classes that were evaluated using real-world EHR data on patients with AKI-associated with COVID- 19 found in N3C. Two candidate drug classes, acetylcholine receptor antagonists and glucocorticoid receptor agonists, were found to be significantly associated with COVID-AKI recovery. Our results indicate that patients hospitalized with COVID-AKI who have not recovered within 7 days of admission may benefit from therapy with members of these two drug classes. This example applies a multi-modal approach encompassing genomics, drug perturbation, and clinical data for drug repurposing and the first to identify candidate drug classes in the therapy of COVID-AKI specifically. More broadly, while clinical trials are crucial to determine drug efficacy, they can be resource and time intensive, making it challenging in situations such as future pandemics, where rapid repurposing may be required, or for rare diseases in which robust sample sizes for clinical trials may not be achieved. Therefore, application of our pipeline supports further drug repurposing initiatives beyond COVID-AKI therapy.Attorney Docket Number: UM-43517.601

[0437] One major strength of our study is the use of scRNA-seq data to identify a gene expression signature specific to COVID-AKI. scRNA-seq data provides us with a high-resolution approach beyond individual cells to examine gene signatures across the entire transcriptome associated with COVID-AKI.

[0438] A strength of this study was our evaluation of candidate drug classes for COVID-AKI therapy using a large independent electronic health record in the form of N3C. This is the largest observational cohort examining drug therapies for COVID-AKI, consisting of over 10,000 patients with COVID-AKI and many with exposure to at least 1 of 13 drug classes.

[0439] There are three categories of drug classes that we found to be associated with COVID-AKI using scRNA-seq, LINCS, and N3C. These include drug classes that have been demonstrated to be beneficial among patients with COVID-AKI in the literature, drug classes without existing evidence in the literature regarding their efficacy in COVID-AKI therapy, and drug classes demonstrated to be associated with worsening AKI in the literature. Evidence in the literature supports our findings that glucocorticoid receptor agonists may be beneficial in COVID-AKI therapy. Prior studies have shown that dexamethasone, a member of the glucocorticoid receptor agonist class which we found to be associated with COVID-AKI recovery, is associated with lower odds of developing AKI and complications such as COVID-associated nephropathy (COVAN) [Kadariya, Oieux, Iglesias, Oweis], For example, Orieux et al. conducted a study of 126 patients with COVID admitted to the ICU and found that the odds of developing AKI was 3.23 times lower among patients exposed to dexamethasone, although exposure to dexamethasone did not lower the risk of needing kidney replacement therapy [Orieux], Another study of 249 patients conducted by Iglesias et al. demonstrated that use of tocilizumab and corticosteroids together was associated with 2.29 times reduced risk of developing AKI among CO VID patients admitted to the ICU [Iglesias], While several mechanisms for COVID-AKI, including direct viral injury through ACE2 receptor activation in the kidneys, imbalanced RAAS activation, inflammatory glomerular and tubular damage, and vascular injury, one theory suggests a common pathway between COVID-associated acute respiratory distress syndrome (ARDS) and AKI in which the release of pro-inflammatory cytokines such as IL-1 and IL-6, and consequently increased vascular damage, renal interstitial pressure, and edema contribute to AKI in the setting of ARDS [Legrand], Thus, through theAttorney Docket Number: UM-43517.601

[0440] inhibition of pro-inflammatory cytokine production, glucocorticoids could play a role in the treatment of COVID-AKI.

[0441] The potential role of acetylcholine receptor antagonists in the treatment of COVID-AKI is less clear. While acetylcholine receptor antagonists are known to cause urinary retention, particularly in older adults, which can contribute to AKI, other studies on COVID-associated myocarditis suggest that ACE-2 expression, a primary target of SARS-CoV-2, can be altered by acetylcholine receptor antagonists [Liu], Moreover, the inclusion of inhaled acetylcholine receptor antagonists such as tiotropium and ipratropium in our study used to treat the respiratory manifestations of COVID-19 may indicate that the association between this drug class and recovery of COVID-AKI is a proxy for recovery from overall COVID-19 infection. However, similar to the case with glucocorticoids, the association between remdesivir and lower odds of AKI recovery suggests a potential effect of acetylcholine receptor antagonists on AKI recovery independent of COVID-19 recovery.

[0442] We also sought to assess the effectiveness of drugs already used in the therapy of COVID- 19 such as remdesivir to determine whether the observed COVID-AKI recovery among patients exposed to glucocorticoids and acetylcholine receptor antagonists was due to overall COVID- 19 recovery. We found that the use of remdesivir was associated with decreased odds of COVID-AKI recovery, despite evidence in the literature highlighting the efficacy of remdesivir in the COVID-19 therapy with regards to mortality reduction and reduction in time to clinical recovery [Chen, Wu], Remdesivir is also associated with increased odds of AKI in the literature, with the mechanism of nephrotoxicity currently unknown [Wu], Although glucocorticoids and inhaled acetylcholine receptor antagonists are often given in combination with remdesivir for inpatient COVID therapy, the difference in the direction of association between these two drug classes and remdesivir with respect to COVID-AKI suggests that glucocorticoids and acetylcholine receptor antagonists may have a beneficial effect on kidney function independent of the treatment of COVID-19 and strengthens our conclusions.

[0443] We also observed that ACE inhibitors and ARBs were associated with lower odds of AKI recovery. Several studies have demonstrated the detrimental effect of ACE inhibitor use in patients hospitalized with COVID with respect to the development and severity of COVID-AKI, possibly due to RAAS blockade and impaired renal blood flow autoregulation [Lee, Oussalah], Indeed, ACE inhibitors and ARBs are typically held on hospital admission due to risk ofAttorney Docket Number: UM-43517.601

[0444] precipitating AKI [Lee, Oussalah]. ARBs were also associated with lower odds of AKI recovery, albeit not significant. ARBs and ACE inhibitors have a similar mechanism of action, however the association between ARBs and COVID-19 complications is heterogeneous [Soleimani, Chambergo-Michilot], One study found that discontinuation of ARBs during hospitalization is associated with greater risk of AKI in COVID-19 patients [Soleimani], However, a larger meta analysis did not find a significant difference in mortality, ICU admission, or notably AKI requiring renal replacement therapy among patients who had their ARB discontinued on hospital admission compared to patients without discontinuation [Chambergo-Michilot],

[0445] With regard to SGLT2 inhibitors, there are several potential reasons why we did not observe a significant association with COVID-AKI recovery. Relative to other drug classes, fewer patients were on SGLT2 inhibitors in our cohort (n=551) and therefore our example may have been underpowered with regard to showing a treatment effect. Furthermore, studies on diabetic patients hospitalized with COVID-19 have shown associations between the use of SGLT2 inhibitors and diabetic ketoacidosis (DKA) due to poor oral intake, potentially contributing to AKI. However, other studies have not demonstrated a significant association between SGLT2 inhibitor use and COVID-AKI and in fact have shown cardiorenal benefits of this drug class [Milder, Salvatore, Heerspink], Furthermore, SGLT2 inhibitors have only been recently approved for a variety of indications including heart failure and there is thus far limited data regarding their safety and efficacy when continued during a hospitalization [Singh].

[0446] REFERENCES:

[0447] 1. WHO coronavirus (COVID-19) dashboard. World Health Organization. Accessed December 7, 2023. https: / / covidl9.who.int / .

[0448] 2. 6. Rubin S. Orieux A, Prevel R, et al. Characterization of acute kidney injury in critically ill patients with severe coronavirus disease 2019. Clin Kidney J. 2020;13(3):354-361.

[0449] 3. 7. Chao CT, Tsai HB, Wu CY, et al. The severity of initial acute kidney injury at admission of geriatric patients significantly correlates with subsequent in-hospital complications. S ci Rep. 2015;5:13925.

[0450] 4. Yoo YJ, Wilkins KJ, Alakwaa F, et al. COVID-19-associated AKI in hospitalized US patients: incidence, temporal trends, geographical distribution, risk factors and mortality.Attorney Docket Number: UM-43517.601

[0451] Preprint. medRxiv. 2022;2022.09.02.22279398. Published 2022 Sep 2.

[0452] doi: 10.1101 / 2022.09.02.22279398

[0453] 5. Xu Y, Kong J, Hu P. Computational Drug Repurposing for Alzheimer's Disease Using Risk Genes From GWAS and Single-Cell RNA Sequencing Studies. Front Pharmacol.

[0454] 2021;12:617537. Published 2021 Jun 30.

[0455] 6. Kwak MS. Hwang CI, Cha JM, Jeon JW, Yoon JY, Park SB. Single-Cell Network-Based Drug Repositioning for Discovery of Therapies against Anti-Tumour Necrosis Factor-Resistant Crohn's Disease. Int J Mol Sci. 2023;24(18): 14099. Published 2023 Sep 14.

[0456] 7. Gupta C, Xu J, Jin T, et al. Single-cell network biology characterizes cell type gene regulation for drug repurposing and phenotype prediction in Alzheimer's disease. PLoS Comput Biol. 2022;18(7):el010287. Published 2022 Jul 18.

[0457] 8. Shevtsov A, Raevskiy M, Stupnikov A, Medvedeva Y. In Silico Drug Repurposing in Multiple Sclerosis Using scRNA-Seq Data. Int J Mol Sci. 2023;24(2):985. Published 2023 Jan 4.

[0458] 9. Mo Z, Liu D, Chen Y, et al. Single-cell transcriptomics reveals the role of Macrophage-Naive CD4 + T cell interaction in the immunosuppressive microenvironment of primary liver carcinoma. J Transl Med. 2022;20(l):466. Published 2022 Oct 11.

[0459] 10. Lui G, Guaraldi G. Drug treatment of COVID- 19 infection. Curr Opin Pulm Med.

[0460] 2023;29(3): 174-183.

[0461] 11. Pezoulas VC, Kourou KD, Mylona E, et al. ICU admission and mortality classifiers for COVID-19 patients based on subgroups of dynamically associated profiles across multiple timepoints. Comput Biol Med. 2022;141: 105176.

[0462] 12. Neyra JA, Chawla LS. Acute Kidney Disease to Chronic Kidney Disease. Crit Care Clin.

[0463] 2021;37(2):453-474.

[0464] 13. Yoon SY, Kim JS, Jeong KH, Kim SK. Acute Kidney Injury: Biomarker-Guided Diagnosis and Management. Medicina (Kaunas). 2022;58(3):340. Published 2022 Feb 23.

[0465] 14. Ostermann M, Karsten E, Lumlertgul N. Biomarker-Based Management of AKI: Fact or Fantasy? [published correction appears in Nephron. 2022;146(3):324-326], Nephron.

[0466] 2022;146(3):295-301.

[0467] 15. Alakwaa FA, Benny P, Chaudhary K et al. A data driven approach to identify ZM-447439 as a potential repurposed drug to overcome tamoxifen-resistance in breast cancer, 21Attorney Docket Number: UM-43517.601

[0468] August 2020, PREPRINT (Version 1) available at Research Square [https: / / doi. Org / 10.21203 / rs.3.rs-56103 / vl]

[0469] 16. The LINCS project, clue.io. Accessed November 19, 2023. https: / / clue.io / lincs#access.

[0470] 17. Subramanian A. Query DataSets for GSE92742. National Center for Biotechnology Information. September 8, 2021. Accessed November 19, 2023. https: / / www.ncbi.nlm.nih. gov / geo / query / acc.cgi?acc=GSE92742.

[0471] 18. National Institutes of Health (NIH). National Center for Advancing Translational Sciences (NCATS). National COVID Cohort Collaborative Data Enclave Repository. Bethesda, Maryland: U. S. Department of Health and Human Services, National Institutes of Health, 2023. https: / / covid.cd2h.org / .

[0472] 19. Liu KD, Yang J. Tan TC, et al. Risk Factors for Recurrent Acute Kidney Injury in a Large Population-Based Cohort. Am J Kidney Dis. 2019;73(2): 163-173.

[0473] 20. Dudreuilh C, Aguiar R, Ostermann M. Acute kidney injury in kidney transplant patients. Acute Med. 2018; 17(1):31-35.

[0474] 21. Matsumoto K, Prowle JR. COVID-19-associated AKI. Curr Opin Crit Care.

[0475] 2022;28(6):630-637.

[0476] 22. Ho, D., Imai, K., King, G., & Stuart, E. A. (2011). Matchit: Nonparametric Preprocessing for Parametric Causal Inference. Journal of Statistical Software, 42(8), 1-28.

[0477] 23. Pushpakom S, Iorio F, Eyers PA, et al. Drug repurposing: progress, challenges and recommendations. Nat Rev Drug Discov. 2019; 18( l):41-58.

[0478] 24. Sanseau P, Agarwal P, Barnes MR, et al. Use of genome-wide association studies for drug repositioning. Nat Biotechnol. 2012;30(4):317-320. Published 2012 Apr 10.

[0479] 25. Khafipour A, Eissa N, Munyaka PM, et al. Denosumab Regulates Gut Microbiota Composition and Cytokines in Dinitrobenzene Sulfonic Acid (DNBS) -Experimental Colitis. Front Microbiol. 2020; 11:1405. Published 2020 Jun 25.

[0480] 26. Crandall CJ, Manson JE, Hohensee C, et al. Association of genetic variation in the tachykinin receptor 3 locus with hot flashes and night sweats in the Women's Health Initiative Study. Menopause. 2017;24(3):252-261.

[0481] 27. Griebel G, Beeské S. Is there still a future for neurokinin 3 receptor antagonists as potential drugs for the treatment of psychiatric diseases?. Pharmacol Ther. 2012;133(l): 116-123.Attorney Docket Number: UM-43517.601

[0482] 28. Prague JK, Roberts RE, Comninos AN, et al. Neurokinin 3 receptor antagonism as a novel treatment for menopausal hot flushes: a phase 2, randomised, double-blind, placebo-controlled trial. Lancet. 2017:389(10081): 1809-1820.

[0483] 29. Panchapakesan U, Pollock C. Drug repurposing in kidney disease. Kidney Int.

[0484] 2018;94(l):40-48.

[0485] 30. Tholen M, Ricksten SE, Lannemyr L. Effects of levosimendan on renal blood flow and glomerular filtration in patients with acute kidney injury after cardiac surgery: a double blind, randomized placebo-controlled study. Crit Care. 2021;25(l):207. Published 2021 Jun 12.

[0486] 31. Mao Z, Valluru MK, Ong ACM. Drug repurposing in autosomal dominant polycystic kidney disease: back to the future with pioglitazone. Clin Kidney J. 2021; 14(7): 1715-1718. Published 2021 Mar 26.

[0487] 32. Keller SA, Chen Z, Gaponova A, Korzinkin M, Berquez M, Luciani A. Drug discovery and therapeutic perspectives for proximal tubulopathies. Kidney Int. 2023; 104(6): 1103-1112. 33. Benmerah A, Briseno-Roa L, Annereau JP, Saunier S. Repurposing small molecules for nephronophthisis and related renal ciliopathies. Kidney Int. 2023;104(2):245-253.

[0488] 34. Chan L, Chaudhary K, Saha A, et al. AKI in Hospitalized Patients with COVID-19. J Am Soc Nephrol. 2021;32(1): 151- 160.

[0489] 35. Hilton J, Boyer N, Nadim MK, Fomi LG, Kellum JA. COVID-19 and Acute Kidney Injury. Crit Care Clin. 2022;38(3):473-489.

[0490] 36. Velez ICQ, Caza T, Larsen CP. COVAN is the new HIV AN: the re-emergence of collapsing glomerulopathy with COVID- 19 [published correction appears in Nat Rev Nephrol.

[0491] 2020 Aug 11;:]. Nat Rev Nephrol. 2020;16(10):565-567.

[0492] 37. Zhang J, Pang Q, Zhou T, et al. Risk factors for acute kidney injury in COVID- 19 patients: an updated systematic review and meta-analysis. Ren Fail. 2023;45(l):2170809.

[0493] 38. Orieux A, Khan P, Prevel R, Gruson D, Rubin S, Boyer A. Impact of dexamethasone use to prevent from severe COVID-19-induced acute kidney injury. Crit Care. 2021;25(1):249. Published 2021 Jul 16.

[0494] 39. Kadariya K, Soji-Ayoade D, Isaac S, Paul S. Potential Role of High-Dose Steroids in the Treatment of COVID-19-Associated Nephropathy. Cureus. 2023;15(2):e35372. Published 2023 Feb 23.Attorney Docket Number: UM-43517.601

[0495] 40. Iglesias J, Vassallo A, Ilagan J, et al. Acute Kidney Injury Associated with Severe SARS-CoV-2 Infection: Risk Factors for Morbidity and Mortality and a Potential Benefit of Combined Therapy with Tocilizumab and Corticosteroids. Biomedicines. 2023; 11(3):845. Published 2023 Mar 10.

[0496] 41. Gabarre P, Dumas G, Dupont T, Darmon M, Azoulay E, Zafrani L. Acute kidney injury in critically ill patients with COVID-19. Intensive Care Med. 2020:46(7): 1339-1348.

[0497] 42. Yildiz S, Heybeli C, Soysal P, Smith L, Veronese N, Kazancioglu R. Frequency and Clinical Impact of Anticholinergic Burden in older patients: Comparing older patients with and without chronic kidney disease. Arch Gerontol Geriatr. 2023; 112: 105041.

[0498] 43. Liu W, Liu Z, Li YC. COVID-19-related myocarditis and cholinergic anti-inflammatory pathways. Hellenic J Cardiol. 2021;62(4):265-269.

[0499] 44. Lee SA, Park R, Yang JH, et al. Increased risk of acute kidney injury in coronavirus disease patients with renin-angiotensin-aldosterone-system blockade use: a systematic review and meta- analysis. Sci Rep. 2021; 11(1): 13588. Published 2021 Jun 30.

[0500] 45. Oussalah A, Gleye S, Clerc Urmes I, et al. Long-term ACE Inhibitor / ARB Use Is Associated With Severe Renal Dysfunction and Acute Kidney Injury in Patients With Severe COVID- 19: Results From a Referral Center Cohort in the Northeast of France. Clin Infect Dis.

[0501] 2020;71(9):2447-2456.

[0502] 46. Jaffer KY, Chang T, Vanle B, et al. Trazodone for Insomnia: A Systematic Review. Innov Clin Neurosci. 2017; 14(7-8):24-34. Published 2017 Aug 1.

[0503] 47. https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC9064179 /

[0504] 48. Oweis AO. Alshelleh SA, Hawasly L, Alsabbagh G, Alzoubi KH. Acute Kidney Injury among Hospital-Admitted COVID-19 Patients: A Study from Jordan. Int J Gen Med.

[0505] 2022;15:4475-4482. Published 2022 Apr 29

[0506] 49. Chambergo-Michilot D, Runzer-Colmenares FM, Segura-Saldana PA. Discontinuation of Antihypertensive Drug Use Compared to Continuation in COVID-19 Patients: A Systematic Review with Meta-analysis and Trial Sequential Analysis. High Blood Press Cardiovasc Prev.

[0507] 2023;30(3):265-279.

[0508] 50. Soleimani A, Kazemian S, Karbalai Saleh S, et al. Effects of Angiotensin Receptor Blockers (ARBs) on In-Hospital Outcomes of Patients With Hypertension and Confirmed or Clinically Suspected COVID-19. Am J Hypertens. 2020;33( 12): 1102-1111.Attorney Docket Number: UM-43517.601

[0509] 51. Zhang Y, He D, Zhang W, et al. ACE Inhibitor Benefit to Kidney and Cardiovascular Outcomes for Patients with Non-Dialysis Chronic Kidney Disease Stages 3-5: A Network MetaAnalysis of Randomised Clinical Trials. Drugs. 2020;80(8):797-811.

[0510] 52. Milder TY, Stocker SL, Day RO, Greenfield JR. Potential Safety Issues with Use of Sodium-Glucose Cotransporter 2 Inhibitors, Particularly in People with Type 2 Diabetes and Chronic Kidney Disease. Drug Saf. 2020;43(12): 1211-1221.

[0511] 53. Heerspink HJL, Furtado RHM, Berwanger O, et al. Dapagliflozin and Kidney Outcomes in Hospitalized Patients with COVID-19 Infection: An Analysis of the DARE- 19 Randomized Controlled Trial. Clin J Am Soc Nephrol. 2022;17(5):643-654.

[0512] 54. Salvatore T, Galiero R, Caturano A, et al. An Overview of the Cardiorenal Protective Mechanisms of SGLT2 Inhibitors. Int J Mol Sci. 2022;23(7):3651. Published 2022 Mar 26. 55. Matsushita K, Mori K, Saritas T, et al. Cilastatin Ameliorates Rhabdomyolysis-induced AKI in Mice. J Am Soc Nephrol. 2021;32(10):2579-2594.

[0513] 56. Jado JC, Humanes B, Gonzalez-Nicolas MA, et al. Nephroprotective Effect of Cilastatin against Gentamicin-Induced Renal Injury In Vitro and In Vivo without Altering Its Bactericidal Efficiency. Antioxidants (Basel). 2020;9(9):821. Published 2020 Sep 3. doi:10.3390 / antiox9090821

[0514] 57. Kale A, Shelke V, Dagar N, Anders HJ, Gaikwad AB. How to use COVID- 19 antiviral drugs in patients with chronic kidney disease. Front Pharmacol. 2023:14:1053814. Published 2023 Feb 9.

[0515] 58. Behiry S, Rabie A, Kora M, Ismail W, Sabry D, Zahran A. Effect of combination sildenafil and gemfibrozil on cisplatin-induced nephrotoxicity: role of heme oxygenase- 1. Ren Fail. 2018;40(l):371-378.

[0516] 59. Wan H. Hu Z, Wang J, Zhang S, Yang X, Peng T. Clindamycin-induced Kidney Diseases: A Retrospective Analysis of 50 Patients. Intern Med. 2016:55(11): 1433- 1437.

[0517] 60. Wu B, Luo M, Wu F, He Z, Li Y, Xu T. Acute Kidney Injury Associated With Remdesivir: A Comprehensive Pharmacovigilance Analysis of COVID-19 Reports in FAERS. Front Pharmacol. 2022;13:692828. Published 2022 Mar 25.

[0518] 61. Singh LG, Ntelis S, Siddiqui T, Seliger SL, Sorkin JD, Spanakis EK. Association of Continued Use of SGLT2 Inhibitors From the Ambulatory to Inpatient Setting With HospitalAttorney Docket Number: UM-43517.601

[0519] Outcomes in Patients With Diabetes: A Nationwide Cohort Study. Diabetes Care. Published online December 5, 2023. doi:10.2337 / dc23-1129.

[0520] 62. Dann E, Teeple E, Elmentaite R, Meyer KB, Gaglia G, Nestle F, et al. Single-cell RNA sequencing of human tissue supports successful drug targets. medRxiv. [Preprint], 2024.04.04.24305313: doi: 10.1101 / 2024.04.04.24305313.

[0521] 63. Jagadeesh KA, Dey KK, Montoro DT, Mohan R, Gazal S, Engreitz JM, et al. Identifying disease-critical cell types and cellular processes by integrating single-cell RNA-sequencing and human genetics. Nat Genet. 2022;54: 1479-1492.

[0522] 64. Jackson HW, Fischer JR, Zanotelli VRT, Ali HR, Mechera R, Soysal SD, et al. The single-cell pathology landscape of breast cancer. Nature. 2020;578: 615-620.

[0523] 65. Van Galen P, Hovestadt V, Wadsworth M II, Hughes T, Griffin GK, Verga JA, et al. Single-cell RNA-seq reveals AML cellular hierarchies relevant to clinical outcomes and immunity. Blood. 2018; 132: 542-542.

[0524] 66. Dominguez CX, Muller S, Keerthivasan S, Koeppen H, Hung J, Gierke S, et al. Singlecell RNA sequencing reveals stromal evolution into LRRC15+ myofibroblasts as a determinant of patient response to cancer immunotherapy. Cancer Discov. 2020;10: 232-253.

[0525] 67. Imai Y, Kusakabe M, Nagai M, Yasuda K, Yamanishi K. Dupilumab effects on innate lymphoid cell and helper T cell populations in patients with atopic dermatitis. JID Innov. 2021;1: 100003.

[0526] 68. Sun D, Guan X, Moran AE, Wu L-Y, Qian DZ, Schedin P, et al. Identifying phenotype-associated subpopulations by integrating bulk and single-cell sequencing data. Nat Biotechnol. 2022;40: 527-538.

[0527] 69. Chen C, Fang J, Chen S, et al. The efficacy and safety of remdesivir alone and in combination with other drugs for the treatment of COVID-19: a systematic review and metaanalysis. BMC Infect Dis. 2023;23(1):672. Published 2023 Oct 9.

[0528] All publications and patents mentioned in the specification and / or listed below are herein incorporated by reference. Various modifications and variations of the described method and system of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been described in connection withAttorney Docket Number: UM-43517.601

[0529] specific embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in the relevant fields are intended to be within the scope described herein.

Claims

Attorney Docket Number: UM-43517.601CLAIMSWe claim:

1. A method for characterizing drugs, comprising:a) identifying a disease-associated gene signature for a disease of interest;b) selecting a plurality of candidate drugs, using drug perturbation data, that alter said disease-associated gene signatures; andc) characterizing said candidate drugs related to said disease of interest by analyzing data contained in electronic health records.

2. The method of claim 1, wherein said characterizing comprises testing efficacy said candidate drugs.

3. The method of claim 1, wherein said characterizing comprises identifying a new drug for treating the disease of interest.

4. The method of claim 1, wherein said characterizing comprises weighing: a) two or more factors associated with a candidate drug selected from the group consisting of: i) drug potency; ii) drug selectivity; iii) gene perturbation score based on connectivity score of altering target genes; and iv) class score based on a number of drugs that have negative connectivity scores that belong to the same class as the candidate drug; and b) said data.

5. The method of claim 1, wherein said identifying disease-associated gene signatures comprises transcriptomic analysis of cells from disease and non-disease samples.

6. The method of claim 5, wherein said transcriptomic analysis comprises single-cell RNA-seq analysis.

7. The method of claim 1, further comprising step d) testing a candidate drug in a laboratory disease model.Attorney Docket Number: UM-43517.6018. The method of claim 1, wherein said selecting comprising generating a ranked list of candidate drugs.

9. The method of claim 1, wherein said electronic health records are obtained from or contain data from an observational cohort study or clinical trial dataset.

10. The method of claim 1, wherein the electronic health record comprises data from one or more clinical trials.

11. A system comprising a computer processor configured to conduct step c) of claim 1 or 14.

12. The system of claim 11, wherein said computer processor is further configured to conduct step b) of claim 1 or 14.

13. The system of claim 2, wherein said computer processor is further configured to conduct step a) of claim 1 or 14.

14. A method of characterizing drags comprising:a) identifying a disease-associated gene signature for a disease of interest, wherein said disease-associated gene signature comprises a plurality of differentially expressed genes (DEGs) in a first cell type,wherein said plurality of DEGs comprise a plurality of upregulated and / or downregulated genes, compared to non-disease wild type status, in said first cell type;b) selecting a plurality of drugs by inputting said plurality of DEGs into a drug perturbation dataset which outputs a plurality of candidate drags that reverses upregulation and / or downregulation of at least some, or all, of said plurality of upregulated and / or downregulated genes in said first cell type,Attorney Docket Number: UM-43517.601wherein said drug perturbation dataset comprises data that comprises a plurality of individual gene expression results of a particular candidate drug interacting with said particular cell type;c) characterizing said at last one of said plurality of candidate drugs related to said disease of interest by analyzing data contained in electronic health records.

15. The method of claim 14, wherein said characterizing comprises testing efficacy said candidate drugs.

16. The method of claim 14, wherein said characterizing comprises identifying a new drug for treating the disease of interest.

17. The method of claim 14, wherein said characterizing comprises weighing: a) two or more factors associated with a candidate drug selected from the group consisting of: i) drug potency; ii) drug selectivity; iii) gene perturbation score based on connectivity score of altering target genes; and iv) class score based on a number of drugs that have negative connectivity scores that belong to the same class as the candidate drug; and b) said data.

16. The method of claim 14, wherein said identifying disease-associated gene signatures comprises transcriptomic analysis of cells from disease and non-disease samples.

17. The method of claim 16, wherein said transcriptomic analysis comprises single-cell RNA-seq analysis.

18. The method of claim 14, further comprising step d) testing a candidate drug in a laboratory disease model.

19. The method of claim 14, wherein said selecting comprising generating a ranked list of candidate drugs.Attorney Docket Number: UM-43517.60120. The method of claim 14, wherein said electronic health records are obtained from or contain data from an observational cohort study or clinical trial dataset.

21. The method of claim 14, wherein the electronic health record comprises data from one or more clinical trials.