In silico drug development and drug activity determination
In silico drug development using polygenic risk scores predicts drug efficacy and safety in specific patient populations, addressing the inefficiencies of traditional drug development for common chronic diseases by reducing the need for extensive clinical trials.
Patent Information
- Application Number
- JP2025528187
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-15
- Filing Date
- 2023-11-14
- Publication Date
- 2025-12-17
AI Technical Summary
Existing drug development processes are costly and time-consuming, particularly for common chronic diseases, due to the need for large-scale clinical trials to assess efficacy and safety, while new approaches are often focused on oncology and rare diseases with defined molecular mechanisms.
Utilizing polygenic risk scores (PRS) in conjunction with computer models to predict drug activity and efficacy without empirical data, enabling in silico drug development and clinical trial design, including patient-specific models for virtual cohorts to test safety and efficacy.
Enables efficient and cost-effective drug development by predicting drug efficacy and safety in specific patient populations, reducing the need for extensive clinical trials and improving treatment response through PRS-informed methodologies.
Smart Images

Figure 2025540934000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority under 35 U.S.C. §119(e) to U.S. Provisional Application No. 63 / 425,579, filed November 15, 2022, the entire contents of which are incorporated herein by reference.
[0002] Field of Disclosure The present disclosure relates to systems, methods and computer program products for drug development using in silico techniques. [Background technology]
[0003] Modeling and simulation is a rapidly developing area in both technology and applications of life sciences. The use of modeling and simulation has expanded beyond the depiction of drug exposure to the dynamic depiction of complex drug effects and disease subtypes and progression. In recent years, new approaches in modeling and simulation have begun to provide important insights in biomedicine, paving the way for their potential use to reduce, refine, and partially replace both animal and human testing.
[0004] Overview of an exemplary embodiment A polygenic risk score (PGS), used interchangeably herein as a polygenic risk score (PRS), combines genome-wide genotype data with gene-phenotype association data, such as from genome-wide association studies (GWAS), to summarize genetic burden for a trait into a metric. This disclosure provides in silico systems, methods, and computer program products using PRS for various aspects of drug development and patient treatment optimization, including drug indication selection and clinical trial replication. Specifically, this disclosure provides a PRS-informed drug development system to facilitate the development of novel therapeutic products and the repurposing of existing therapeutics for new populations and / or indications. PRS-informed drug development has numerous applications, ranging from discovering novel drug targets for the treatment of certain indications, to repurposing existing therapeutics for new indications, to novel and efficient designs of clinical trials, to incorporating novel endpoints in ongoing clinical trials, to providing evidence of efficacy to power clinical trials through improved data analysis, to improving the use of existing therapeutics in patient populations. These systems mimic multiple aspects of the drug development process by utilizing existing patient variant data in combination with computer models. Importantly, the disclosed systems of this approach enable such development without the need for empirical data, which has been required to date for various steps in the drug development process, from drug discovery through clinical trials to regulatory approval, although empirical data may also be incorporated in certain embodiments.
[0005] In addition to patient variant data, the computer models of the disclosed systems can optionally utilize mechanistic knowledge, such as physical, chemical, and / or environmental data relevant to the patient population, as well as available biological and physiological knowledge, such as patient treatment (e.g., medical intervention) or activity (e.g., exercise or lack thereof).
[0006] In some aspects, the present disclosure provides a preclinical drug development system for identifying and / or developing treatments (e.g., drugs) for various therapeutic areas and indications. The preclinical drug discovery system is configured to use mammalian polygenic scores in conjunction with computer models to provide a PRS-informed drug discovery methodology. A potential application of such a method is illustrated in Figure 1.
[0007] In some aspects, the present disclosure provides a drug discovery system for repurposing existing therapeutics for new therapeutic areas and indications, configured to use mammalian polygenic scores in conjunction with computer models to provide a PRS-informed methodology for identifying potential new uses for existing therapeutics using known drug and patient information.
[0008] In some embodiments, the present disclosure provides in silico systems and methods for determining drug activity. In particular, the present disclosure provides systems and methods for determining drug activity for a particular patient population, phenotype, or for use in a therapeutic indication / area. This drug activity may be determined for known or de novo agents.
[0009] Thus, in some aspects, the present disclosure provides an in silico method for determining drug activity of a drug target against a plurality of intermediate phenotypes, the method comprising, alternatively consisting essentially of, or further consisting of the steps of: obtaining genetic variant and phenotypic data from each of a plurality of subjects; determining a variant effect value for each intermediate phenotype based on external data; calculating a polygenic score for each intermediate phenotype in a disease process; calculating a pharmacomimetic genetic score for each drug target for each intermediate phenotype; and identifying predicted drug activity of the drug target against one or more phenotypes based on a statistical interaction between the polygenic score and the pharmacomimetic genetic score for each drug target in an association analysis between the intermediate phenotype and an outcome phenotype. In some aspects, the present disclosure provides a system for providing in silico clinical trials of novel therapeutics using a mammalian polygenic score in conjunction with a computer model. Historically, safety and efficacy data provided to regulatory agencies in support of marketing approval of new pharmaceuticals has been generated experimentally, either in vitro or in vivo. More recently, regulatory agencies have begun to receive and accept in silico data generated using modeling and simulation, and the data provided by the disclosed system can be data supporting regulatory applications without the need for expensive and time-consuming experimentation.
[0010] The disclosed system enables the development of patient-specific models that form virtual cohorts for testing the safety and / or efficacy of new drugs and new medical devices.
[0011] In some embodiments, the present disclosure provides methods for determining drug activity. determining a value for the effect of each variant on each phenotype based on external data from a non-overlapping set of subjects; constructing a matrix where variants are columns and phenotypes are rows, with the values being the externally obtained variant effects; performing a truncated principal component analysis on this matrix to calculate a user-specified number of principal components, typically between 2 and 5; calculating a polygenic score for each calculated principal component by reweighting the existing polygenic score for a user-specified "reference" phenotype so that the new weight for each variant included in the polygenic score is equal to its old weight multiplied by the square of the loading for the principal component in question, divided by the sum of the squared loadings for all calculated principal components; and finally, identifying a predicted drug activity for the "reference" phenotype based on each of the reweighted polygenic scores, which may often be interpreted as corresponding to a biological pathway. This example method is visually illustrated in Figure 2.
[0012] In some embodiments, the disclosure provides an in silico system for drug development, the system including, alternatively consisting essentially of, or additionally consisting of, at least one hardware processor and a non-transitory computer-readable storage medium having program code stored thereon, the program code being executable by the at least one hardware processor to: obtain data regarding genetic variants and phenotypes from each of a plurality of subjects; determine a value of a variant effect for each phenotype based on external data; construct a matrix based on the variants, phenotypes, and values; perform principal component analysis on the variant x phenotype association phenotype-variant effect matrix; calculate a polygenic score for each principal component; and determine a predicted drug effect for one or more phenotypes based on the polygenic score calculated for each principal component.
[0013] In some embodiments, the present disclosure provides in silico systems for determining intermediate phenotypes, such as risk factors or genetic variations, that distinguish responders from non-responders to a particular therapeutic intervention. Accordingly, in some aspects, the present disclosure provides in silico systems and methods that include, alternatively consisting essentially of, or further consisting of, obtaining data regarding genetic variants and phenotypes from each of a plurality of subjects, determining a variant effect value for each phenotype based on external data, calculating a polygenic score for selected intermediate phenotypes (e.g., risk factors, biomarkers, or genetic indications of response) in a disease process, calculating a pharmacomimetic gene score for each drug target, and identifying predicted drug activity for one or more phenotypes based on a statistical interaction between the polygenic score and the pharmacomimetic gene score for each drug target in an association analysis with an outcome phenotype.
[0014] In particular aspects, the present disclosure provides methods in which the PRS information is provided in whole or in part from the associated disease phenotype.
[0015] Thus, the present disclosure provides a method for determining drug activity of a drug target, the method comprising, alternatively consisting essentially of, or further comprising the steps of: obtaining data regarding genetic variants and phenotypes from each of a plurality of subjects; determining a value of the variant effect for each phenotype based on external data; calculating a polygenic score for selected outcome phenotypes related to disease outcomes; calculating a pharmacomimetic genetic score for the drug target; and identifying predicted drug activity of the drug target for one or more phenotypes based on a statistical interaction between the polygenic score and the pharmacomimetic genetic score for the drug target in an association analysis with the outcome phenotype. The drug target may be, for example, a known drug for the phenotype, an approved drug, a drug under investigation for an indication different from the phenotype being analyzed, or a de novo agent.
[0016] One aspect of the present disclosure relates to an in silico method for determining drug activity of multiple drug targets, the method comprising, alternatively consisting essentially of, or further consisting of: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the molecular biomarker stratification factor data including, alternatively consisting essentially of, or further consisting of, a plurality of biomarker stratification factors and at least one disease phenotype; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; and identifying a predicted drug activity of each drug target for at least one disease phenotype in a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with the disease phenotype.
[0017] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0018] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0019] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0020] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0021] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; (viii) a polygenic score for biological pathway activity. (xiii) somatic mutation score for drug target expression; (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof.
[0022] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
[0023] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis or weighted principal component analysis based on a matrix of one or more biomarker stratification factor effects.
[0024] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of constructing a matrix based on the one or more biomarker stratification factors, the phenotype, and the plurality of values.
[0025] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
[0026] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratification factor scores for the selected disease phenotype.
[0027] In some embodiments, the method further comprises, alternatively consists of, or further consists of, weighted principal component analysis identifying predicted drug activity of multiple drug targets against at least one disease phenotype.
[0028] In some embodiments, the multiple drug targets are associated with a disease selected from Table 1. In some embodiments, the multiple drug targets are associated with a disease selected from non-alcoholic steatohepatitis (NASH), coronary heart disease, dry age-related macular degeneration, rheumatoid arthritis, atopic disease, or obesity.
[0029] In some embodiments, the multiple drug targets are associated with non-alcoholic steatohepatitis (NASH), and the multiple drug targets include, alternatively consist essentially of, or further consist of HSD17B13.
[0030] In some embodiments, the multiple drug targets are associated with coronary heart disease and the multiple drug targets include, or alternatively consist essentially of, or further consist of LPL, ANGPTL4 and / or ANGPTL3.
[0031] In some embodiments, the multiple drug targets are associated with dry age-related macular degeneration, and the multiple drug targets include, alternatively consist essentially of, or further consist of C3, CFB, CFH, and / or HTRA1.
[0032] In some embodiments, the multiple drug targets are associated with rheumatoid arthritis and the multiple drug targets include, alternatively consist essentially of, or further consist of TYK2.
[0033] In some embodiments, the multiple drug targets are associated with atopic disease and the multiple drug targets include, or alternatively consist essentially of, or further consist of, IL33, TSLP and / or IL4R.
[0034] In some embodiments, the multiple drug targets are associated with obesity and the multiple drug targets include, or alternatively consist essentially of, or further consist of GLP1R and / or GIPR.
[0035] Another aspect of the present disclosure relates to an in silico method for determining drug activity of multiple drug targets, said in silico method comprising: obtaining molecular biomarker stratification factor data and at least one disease phenotype from each of a plurality of subjects; determining the value of the biomarker stratification factor effect for each disease phenotype based on external data; constructing a matrix based on the biomarker stratification factors, disease phenotypes and values; performing a principal component analysis on the biomarker stratification factor x phenotype-associated phenotype biomarker stratification factor effect matrix; calculating a biomarker stratification factor score for each principal component; calculating a pharmacomimetic gene score for each drug target; identifying predicted drug activity of each drug target against one or more phenotypes of a subset of the biomarker stratification factor distributions based on a statistical correlation between the polygenic score calculated for each principal component and the pharmacomimetic gene score for each drug target in an association analysis with outcome phenotypes; Includes.
[0036] Another aspect of the present disclosure relates to an in silico method for determining drug activity of a drug target against a plurality of intermediate phenotypes, said in silico method comprising: obtaining molecular biomarker stratification factor data and a disease phenotype from each of a plurality of subjects; determining a biomarker stratification factor effect value for each intermediate phenotype based on external data; calculating a biomarker stratification factor score for each intermediate phenotype in the disease process; calculating a pharmacomimetic gene score for each drug target for each intermediate phenotype; identifying a predicted drug activity of the drug target for at least one phenotype in a subset of the biomarker stratification factor distribution based on a statistical correlation between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with intermediate and outcome phenotypes; Includes.
[0037] Another aspect of the present disclosure relates to a method for determining drug activity of a drug target, the method comprising: obtaining molecular biomarker stratification factor data and a disease phenotype from each of a plurality of subjects; determining the value of the biomarker stratification factor effect for each phenotype based on external data; calculating biomarker stratification factor scores for selected outcome phenotypes for disease outcomes; calculating a pharmacomimetic gene score for the drug target; identifying a predicted drug activity of the drug target for one or more phenotypes in a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for the drug target in an association analysis with the outcome phenotype; Includes.
[0038] Another aspect of the present disclosure is: generating a plurality of principal components (PCs) corresponding to the genetic ancestry data of the subjects in the study cohort; generating a biomarker stratification factor score for each subject in the study cohort based on at least (i) the PCs and (ii) the biomarker stratification factor weights for the disease of interest; determining which of a plurality of disease-associated variants is a pharmacomimetic for a disease of interest, wherein a disease-associated variant is a pharmacomimetic if it modulates the function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to predict a second effect of the drug on the phenotype; determining a statistical interaction between the pharmacomimetic measures and biomarker stratification factor scores for one or more drug targets, wherein the statistical interaction is predictive of drug target-specific differential treatment response for the disease of interest; taking one or more prediction-based actions based on the determined interactions; The present invention relates to an in silico method, including:
[0039] In some embodiments, the one or more prediction-based actions include at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics.
[0040] In some embodiments, the genetic ancestry data indicates whether subjects in a study cohort are cases or controls for a disease of interest.
[0041] In some embodiments, the genetic ancestry data is based on genotyping arrays or whole genome sequencing.
[0042] In some embodiments, the genetic ancestry data is obtained from publicly available databases.
[0043] In some embodiments, the plurality of PCs comprises at least five PCs.
[0044] In some embodiments, the study cohort includes at least 200 control subjects.
[0045] In some embodiments, each disease-associated variant in the plurality of disease-associated variants meets a genome-wide significance threshold.
[0046] In some embodiments, the plurality of disease-associated variants is determined based on a genome-wide association study (GWAS) of the first disease.
[0047] In some embodiments, the biomarker stratification factor score for each subject is a scaled biomarker stratification factor score.
[0048] In some embodiments, the scaled biomarker stratification factor score is based on at least the raw PGS and the ancestry-normalized PGS.
[0049] In some embodiments, the PGS variant weights are calculated independently of the genetic ancestry data corresponding to the study cohort.
[0050] In some embodiments, determining which of the plurality of disease-associated variants are pharmacomimetics for the disease of interest comprises generating an allele score.
[0051] Another aspect of the present disclosure relates to an in silico system for drug development, the system including: at least one hardware processor; and a non-transitory computer-readable storage medium having program code stored thereon, the program code being executable by the at least one hardware processor such that: obtaining molecular biomarker stratification factor data and a disease phenotype from each of a plurality of subjects; determining the value of the biomarker stratification factor effect for each phenotype based on external data; constructing a matrix based on said biomarker stratification factors, phenotypes and values; Conduct principal component analysis on the biomarker stratification factor × phenotype-related phenotype biomarker stratification factor effect matrix; Calculate the biomarker stratification factor score for each principal component; Calculate a pharmacomimetic gene score for each drug target; Association analysis with outcome phenotypes identifies predicted drug effects for one or more phenotypes in a subset of the biomarker stratification factor distribution based on statistical interactions between the polygenic scores calculated for each principal component and the pharmacomimetic gene scores for each drug target.
[0052] These and other embodiments, features and advantages will be set forth in this disclosure. [Brief explanation of the drawings]
[0053] The accompanying drawings, which are incorporated herein and form a part of this specification, depict one or more embodiments and, together with the description, explain these embodiments. The accompanying drawings are not necessarily drawn to scale. Any values or dimensions shown in the accompanying graphs and drawings are for illustration purposes only and may or may not represent actual or preferred values or dimensions. Where appropriate, some or all features may not be shown to help illustrate essential features. [Figure 1] FIG. 1 is a schematic diagram illustrating the potential application of polygenic risk score (or "polygenic score")-mediated drug development using the in silico system of the present disclosure. [Figure 2] FIG. 2 is a flow diagram illustrating a method according to one embodiment of the present disclosure. [Figure 3] FIG. 3 shows a block diagram of a system for in silico drug development and drug activity determination in one embodiment of the present disclosure. [Figure 4] FIG. 4 shows a flow diagram illustrating a method for in silico drug development and determination of drug activity in one embodiment of the present disclosure. [Figure 5] FIG. 5 shows a flow diagram illustrating a method for in silico drug development and determination of drug activity in one embodiment of the present disclosure. [Figure 6] FIG. 6 shows a flow diagram illustrating a method for in silico drug development and determination of drug activity in one embodiment of the present disclosure. [Figure 7] FIG. 7 shows a sequence for in silico drug development and drug activity determination in one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0054] Detailed Description The following detailed description of the preferred embodiments of the present disclosure will be better understood when read in conjunction with the accompanying drawings.
[0055] The systems and methods provided are illustrated for coronary artery disease, but are generalizable for use in PRS-informed drug development in any polygenic therapeutic area or indication; therefore, one skilled in the art will recognize that the present disclosure is exemplary and applicable to other in silico systems and methods.
[0056] All publications mentioned in this application, including patent documents, scientific articles, and databases, are incorporated herein by reference in their entirety for all purposes to the same extent as if each individual publication were individually incorporated by reference. To the extent that a definition set forth herein contradicts or is inconsistent with a definition set forth in a patent, application, published application, or other publication incorporated herein by reference, the definition set forth herein shall take precedence over the definition incorporated herein by reference.
[0057] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.
[0058] In some embodiments, the present disclosure describes an in silico system for replicating empirical clinical trial data. Disclosures are provided herein demonstrating that genetic analysis using pharmacomimetic variants can replicate drug-PRS interactions observed in retrospective analyses of clinical trials. Furthermore, it is also demonstrated that selectively enrolling patients with high PRS in clinical trials can improve the average treatment response of enrolled patients, as shown by comparing empirical evidence with results obtained with the disclosed system. For example, as shown, the average treatment response of a known anti-PCSK9 therapy can be improved by up to two-fold in patients with high PRS.
[0059] Over the past decade, genome-wide association studies (GWAS) have uncovered the contribution of inherited variants to a wide range of common and complex disorders. Many non-communicable disorders with significant public health impact have highly polygenic genetic backgrounds that contain, or alternatively consist essentially of, hundreds or even thousands of genetic variants (or polymorphisms), each of which has only a small effect on disease risk. While each genetic variant associated with a given disease is valuable in implicating genes or biological pathways involved in the disorder, there is also hope that this genetic data can be used to predict disease risk, potentially with clinical applications.
[0060] For example, substantial success has been achieved in discovering and developing new health interventions, including therapeutic drugs, for common chronic human diseases, defined here as diseases with a prevalence greater than 0.1% in the general population. Despite these advances in managing common chronic diseases, significant unmet needs remain in clinical areas that contribute significantly to the population burden of disease incidence, premature death, and associated societal and insurance costs. It is estimated that six out of ten Americans have at least one common chronic disease, and these diseases collectively account for $2.7 trillion in U.S. health care costs. Because disease progression and response to treatment are heterogeneous, drug development in these areas requires large-scale clinical trials with substantial clinical development costs to assess clinical efficacy. Thus, despite substantial market opportunity and unmet clinical need, drug development has shifted over the past two decades to oncology and rare diseases, where defined molecular mechanisms and convergence among subpopulations have increased the likelihood of demonstrating sufficient efficacy and safety for regulatory approval of novel chemical entities.
[0061] Large-scale genome-wide association studies conducted over the past few decades have revealed genetic contributions to variation in common disease risk and underlying risk factors that play a causal role in their pathology. These studies have found that the architecture of disease risk is genetically complex, reflecting the contributions of tens to millions of individual common genetic variants, each with a small effect, and a smaller number of less frequent and rare genetic variants with larger effects. Collectively, these risk factors explain up to 40% to 80% of the observable variation in disease risk at the population level. Larger-scale studies of the effects of these genetic variants on disease risk have enabled more precise estimates of the contribution of genetic variants to disease risk across a range of effect sizes. This has prompted the development of a scaled summary of these effects, a continuous polygenic genetic score (PGS), that estimates a measurable aggregate of genetic risk for any individual. PGS can be isotropic, reflecting the aggregate effect of genes on disease risk, agnostic to effects on underlying risk factors or biological pathways, or they can reflect underlying promoting pathways in subsets of common chronic diseases among individuals with specific risk factors. Both isotropic and pathway- or risk factor-specific PGS may resolve general disease heterogeneity and identify population subgroups that respond exceptionally well to specific therapeutic mechanisms. Indeed, previous clinical trials have shown that patients with high PGS benefited more from drugs than patients with low PGS. This phenomenon may represent a PGS-stratified drug efficacy trial. PGS-stratified drug efficacy trials have been observed in clinical trials of PCSK9 antibodies for the secondary prevention of cardiovascular disease (Marston et al. 2019, Damask et al. 2019) and statins for the primary prevention of cardiovascular disease (Natarajan 2017). In both cases, clinical benefit was increased in individuals with high PGS.
[0062] Systems implementing the systems and methods described herein may use analytical methods and knowledge bases to identify certain individuals in subpopulations of common chronic human diseases who are predicted to derive greater benefit from specific therapeutic mechanisms. Specifically, the methods identify drug target combinations or disease indications where patients with a high PGS for a disease, for associated risk factors, or for a collection of multiple genetically driven biological pathways are likely to derive greater benefit from a given drug compared to patients with a low PGS.
[0063] To perform the method, a computer can access genetic and phenotypic data from a large group of humans (e.g., a "study cohort"). The genetic data can be derived from genotyping arrays or whole-genome sequencing. The phenotypic data need not be in a particular format. However, in some embodiments, the phenotypic data must be sufficient to determine whether each subject is a case, control, or neither of the disease of interest. The algorithm for assigning case-control status to subjects can be manually defined by the user. The algorithm can take into account, for example, (i) ICD-10 or ICD-9 diagnosis codes from hospital or primary care records; (ii) self-reported diagnoses; (iii) medication prescriptions; and / or (iv) OPCS-4 or OPCS-3 operation codes from hospital records. The computer can also use each subject's age and sex. Examples of suitable research cohorts include UK Biobank, the FinnGen research project, and the All of Us research program. There are no strict requirements for the number of subjects in a study cohort, but typically the computer may use at least up to 5000 cases and up to 5000 controls for a particular disease of interest.
[0064] The computer may also access PGS variant weights for the disease of interest. "PGS variant weights" can refer to a table of genetic variants in which each genetic variant is assigned a weight (e.g., a numerical value). The weights can be an estimate of how much a given variant contributes to a disease's risk. The computer can calculate PGS variant weights using publicly available programs such as PRS-CS (Ge et al. 2019) or LDPred2 (Prive et al. 2020). In some cases, PGS variant weights may be published by academic researchers. The computer can download such weights from web resources, such as the PGS catalog. The calculation of PGS variant weights does not need to include any data from the study cohort, but in some embodiments, they must be derived from an entirely independent dataset.
[0065] The computer may access results from genome-wide association studies (GWAS), typically referred to as "summary statistics," for the disease of interest. To do so, the computer may retrieve GWAS summary statistics from web resources, such as the EBI GWAS catalog. Alternatively, the data processing system may perform GWAS within a study cohort, such as by using publicly available software such as Plink (Purcell et al. 2007, Chang et al. 2015), SAIGE (Zhou et al. 2018), or REGENIE (Mbatchou et al. 2021).
[0066] Computers have access to computational pipelines for mapping GWAS variants to causal genes, such as those published by Mountjoy et al. 2021 and Gazal et al. 2022. Drug-by-PGS interaction discovery methods are agnostic to the specific pipeline used.
[0067] The method can utilize the following paradigm operations: (1) modeling the predicted effect of drug target modulation from genotype-phenotype association analysis between human genetic variants in drug target-encoding genes and observed clinical phenotypes. The genetic variants used for modeling can be individual variants or sets of statistically independent variants in the same drug target gene identified as an "allele score." These statistical tools are "pharmaco-mimetic tools." (2) segmenting human populations according to isotropic PGS, risk factor PGS, or biological pathway PGS; and (3) applying a method to identify statistical interactions of pharmacomimetic tools and PGS with respect to drug targets to predict drug target-specific differential treatment responses. In some embodiments, the method can include using the predicted drug target-specific differential treatment responses to select patients for treatment and / or administer treatment. In some embodiments, the method can include sending the predicted drug target-specific differential treatment response and / or any data used to determine the predicted drug target-specific differential treatment response (e.g., statistical interactions, determined PGS, etc.) to another computer system (e.g., a downstream pipeline) for an entity connected to the computer system used for treatment.
[0068] In carrying out the aforementioned method, the computer system can use human genetic data to predict in advance whether any drug mechanism will show PG-stratified efficacy without the need for clinical trial data. For example, the computer system can apply the method to model the effect of anti-PCSK9 mechanisms on cardiovascular disease risk according to a PGS, thereby exploiting the predictive capabilities of the method, such as accurately reproducing the results of clinical trial data. The computer system can use the method for at least 35 common chronic diseases to identify drug targets for which PGS stratification is likely to produce differential treatment responses and / or implement novel indications for the same drug targets. The computer system can use the method to identify or use PGS-stratified effective LDL-lowering therapies (e.g., statins and PCSK9 antibodies) for cardiovascular disease and drugs for a wide range of indications.
[0069] FIG. 1 is a schematic diagram illustrating the potential use of polygenic risk scores (or "polygenic scores") mediated drug development in one embodiment of the present disclosure.
[0070] 2 is a flow diagram illustrating a method 200 according to one embodiment of the present disclosure. Method 200 may be performed by a data processing system (e.g., a client device or data processing system 302 shown and described with reference to FIG. 3, a server system, etc.). Method 200 may include more or fewer operations, and the operations may be performed in any order. Execution of method 200 enables the data processing system to automatically identify predicted drug effects on phenotypes based on polygenic risk scores.
[0071] In method 200, in act 202, a data processing system obtains genetic variant data from a plurality of phenotyped subjects. In act 204, the data processing system determines a variant effect value for each phenotype based on external data. In act 206, the data processing system constructs a matrix based on the variants, phenotypes, and variant effect values. In act 208, the data processing system performs principal component analysis on the biomarker variant x phenotype association phenotype variant effect matrix. In act 210, the data processing system calculates a polygenic risk score for each principal component. In act 212, the data processing system identifies predicted drug effects for one or more phenotypes based on the polygenic risk scores calculated for each principal component.
[0072] FIG. 3 shows a block diagram of a system 300 for in silico drug development and drug activity determination in one embodiment of the present disclosure. Briefly, system 300 can include a data processing system 302 and data sources 304, 306, and 308. Data processing system 302 can receive or retrieve molecular biomarker stratification factor data for a plurality of subjects. Data processing system 302 can determine values representing biomarker stratification factor effects based on external data from individuals separate from the plurality of subjects. Data processing system 302 can calculate biomarker stratification factor scores, selected disease phenotypes, and pharmacomimetic gene scores for different drug targets. Data processing system 302 can identify predicted drug activity for each drug target through association analysis with disease phenotypes based on the biomarker stratification factor scores and the pharmacomimetic gene scores for the drug target. System 300 may include more, fewer, or different components than those shown in FIG. 3 .
[0073] The data processing system 302 may include one or more processors configured to identify predicted drug activity of drug targets against disease phenotypes. The data processing system 302 may include a network interface 312, a processor 314, and / or a memory 316. The data processing system 302 may communicate with data sources 304, 306, and / or 308 via the network interface 312, which may include an antenna or other networking device that enables communication across a network and / or to other devices. The processor 314 may be or include an ASIC, one or more FPGAs, a DSP, a circuit including one or more processing components, a microprocessor-supporting circuit, a group of processing components, or other suitable electronic processing components. In some embodiments, the processor 314 may execute computer code or modules (e.g., executable code, object code, source code, script code, machine code, etc.) stored in the memory 316 to facilitate the activities described herein. The memory 316 may be any computer-readable, volatile or non-volatile storage medium capable of storing data or computer code.
[0074] Memory 316, in some embodiments, may include a data collector 318, a principal component (PC) generator 320, a pharmacomimetic measure identifier, a statistical interaction calculator 324, and / or an action performer 326. Components 318-326 may operate to identify predicted drug activity of drug targets for different disease phenotypes.
[0075] For example, the data collection component 318 may include programmable instructions that, when executed, cause the processor 314 to communicate with the data sources 304, 306, 308, and / or any other data sources. The data collection component 318 may be or include an application programming interface (API) that facilitates communication between the data processing system 302 and other computing devices. The communication component 318 may communicate with the data sources 304, 306, 308, and / or any other computing devices via the network 310.
[0076] Data sources 304, 306, and 308 may be data sources that store molecular biomarker stratification factor data and / or external data. For example, data sources 304, 306, and 308 may be or include databases (e.g., relational or graph databases) that store molecular biomarker stratification factor data, including biomarker stratification factors from or for a plurality of subjects or individuals and at least one disease phenotype. Data sources 304, 306, and 308 may additionally or alternatively include external data for individuals separate from the plurality of subjects or individuals for which molecular biomarker stratification factor data is provided.
[0077] The data collector 318 can establish a connection with one or more of the data sources 304, 306, and / or 308. The data collector 318 can establish this connection over the network 310. To do so, the data collector 318 can communicate with the data sources 304, 306, and / or 308 over the network 310. In one example, the data collector 318 can send a syn packet to each data source 304, 306, and / or 308 to establish this connection using the TLS handshake protocol. The data collector 318 can use any handshake protocol to establish a connection with the data sources 304, 306, and / or 308.
[0078] The data collector 318 can obtain the molecular biomarker stratification factor data and the external data. In some embodiments, the data collector 318 can obtain the molecular biomarker stratification factor data and the external data through an established connection from one or more of the data sources 304, 306, or 308. The data collector 318 can do so by issuing queries to the data sources 304, 306, or 308. In some embodiments, the data collector 318 can receive the molecular biomarker stratification factor data and / or the external data as input (e.g., manual input) by a user accessing the data processing system. The data collector 318 can obtain the molecular biomarker stratification factor data and portions of the external data using any method or any combination of methods.
[0079] The PC generator 320 may include programmable instructions that, when executed, cause the processor 314 to generate principal components. The PC generator 320 can generate principal components that correspond to (or are based on) genetic ancestry data (e.g., different genes) for subjects in a study cohort. The genetic ancestry data can be based on genotyping arrays or whole-genome sequencing. The PC generator 320 can obtain genetic ancestry data from publicly available databases, as described herein. The PC generator 320 can be, for example, Flash PCA. In some embodiments, the principal components can be used to control subsequent statistical analyses for population stratification.
[0080] For example, the PC generator 320 can generate multiple PCs corresponding to the genetic ancestry data of the subjects in the study cohort. The PC generator 320 can do so by generating principal components on a genetic similarity matrix from the genetic ancestry data of the subjects in the study cohort and using principal component analysis on the genetic similarity matrix. The PC generator 320 can generate any number of PCs using principal component analysis.
[0081] The PC generator 320 can generate a biomarker stratification factor score for each subject in the study cohort. The PC generator 320 can generate the biomarker stratification factor score based on at least (i) the PC and (ii) the weight of the biomarker stratification factor for the disease of interest. For example, the "raw PGS" (e.g., biomarker stratification factor score) for subject S can be defined as the sum of the weights for each variant in a PGS variant weight table: ((the weight of the variant) × (the number of copies S has of that variant)). The weights of PGS variants in the PGS variant weight table can be calculated independently from the genetic ancestry data corresponding to the study cohort. The PC generator 320 can generate biomarker stratification factor scores for any number of diseases, such as nonalcoholic steatohepatitis (NASH), coronary heart disease, dry age-related macular degeneration, rheumatoid arthritis, atopic disease, and / or obesity.
[0082] The PC generator 320 can determine an "ancestry-normalized" PGS. The PC generator 320 can determine this ancestry-normalized PGS as the residual of a linear regression model where the outcome is the raw PGS and the predictors are the PCs of the genetic ancestry data. This normalization can correct for differences in mean PGS values between populations. Additionally, the ancestry-normalized PGS can also be referred to as an ancestry-normalized biomarker stratification factor score.
[0083] The PC generator 320 can determine a "scaled PGS." The PC generator 320 can do so using the following equation: scaled PGS = ((ancestry-normalized PGS - ancestry-normalized PGS across all subjects in the cohort) / (deviation of ancestry-normalized PGS across all subjects in the cohort)). The scaled PGS values can therefore be approximately normally distributed. The scaled PGS can be used for all subsequent analyses. The scaled PGS can also be referred to as a scaled biomarker stratification factor score.
[0084] The pharmacomimetic identifier 322 may include programmable instructions that, when executed, cause the processor 314 to determine which of a plurality of disease-associated variants is a pharmacomimetic for a disease of interest. A disease-associated variant may be a pharmacomimetic if it modulates the function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to predict a second effect of the drug on that phenotype.
[0085] To determine which of multiple disease-associated variants are pharmacomimetic for a disease of interest, the pharmacomimetic identifier 322 may identify "conditionally independent disease-associated variants" by running a program for the disease of interest (e.g., GCTA COJO-SLCT). To do so, for example, the pharmacomimetic identifier 322 may first perform a GWAS (e.g., a first GWAS) in the normal manner. The pharmacomimetic identifier 322 may identify the variant with the most significant association (e.g., smallest p-value) with the disease outcome across the entire genome. The pharmacomimetic identifier 322 may perform the GWAS again, but this time conditioning all variant-disease association tests on the genotype of the most significant variant. The pharmacomimetic identifier 322 may identify the variant with the next most significant association with the disease outcome. The pharmacomimetic identifier 322 may perform the GWAS again, but this time conditioning all variant-disease association tests on the genotype of the most significant variant and the second most significant variant. Disease-associated variants below a predefined p-value threshold (e.g., 5 × 10 -8The pharmacomimetic identifier 322 can repeat this process until it is determined that none of the variants meet a genome-wide significance threshold (e.g., a threshold above 100). Using a lower threshold reduces the likelihood that the analysis is a fluke, but increases the likelihood that the data processing will fail to identify important variants, and vice versa. The pharmacomimetic identifier 322 can use either threshold. Because each variant is associated with disease even after conditioning on the genotypes of all previous variants, the set of variants identified by the pharmacomimetic identifier can be referred to as "conditionally independent disease-associated variants."
[0086] The pharmacomimetic identifier 322 can generate a table, where each row corresponds to one subject from the study cohort and can include demographic and other determined data about the subject (e.g., PGS, PCS of genetic ancestry, number of alternative alleles of each independent disease-associated variant, etc.).
[0087] The pharmacomimetic identifier 322 can determine which conditionally independent disease-associated variants are "pharmacomimetics" (e.g., variants that modulate the function or expression of a drug target gene such that the effect of the variant on a human phenotype is likely to be predictive of the drug's effect on the human phenotype). Depending on the configuration, the pharmacomimetic identifier 322 can use different criteria and / or priorities to identify pharmacomimetic variants from the identified conditionally independent disease-associated variants. In one example, a user may be interested in clinical-stage drugs for a disease. In this case, the pharmacomimetic identifier 322 can compile a table of clinical-stage drugs and their targets using, for example, public sources such as clinicaltrials.gov and / or proprietary databases such as Cortellis. The pharmacomimetic identifier 322 can determine which variants are pharmacomimetic if they are close to a drug target gene (e.g., within 150 kb of the gene's transcription start site) or map to the drug target. In another example, a user may be interested in known and novel drug targets for the antibody modality. In this case, the pharmacomimetic identifier 322 can compile a list of genes encoding proteins that can be drugged with the antibody modality (e.g., proteins that are secreted or localized at the cell surface, etc.). The pharmacomimetic identifier 322 can determine, based on the compiled table, that variants mapped to antibody-druggable proteins are pharmacomimetics (e.g., based on determined variants recorded as associated with antibody-druggable proteins).
[0088] In some embodiments, there are two or more suitable variants for a given drug mechanism. In this case, combining allele scores by aggregating variants (e.g., all variants) identified as pharmacomimetics associated with a given drug mechanism may improve statistical power for detecting drug-PGS interactions. To do so, for example, the pharmacomimetics identification unit 322 can calculate an allele score in the same or similar manner as a PGS (e.g., by assigning a weight to each variant and calculating, for each subject in the cohort, the sum of (weight of the variant) × (number of subjects for that variant) across each variant). The difference between an allele score and a PGS is that an allele score consists of variants that modulate the function or expression of a single drug target gene, while a PGS can consist of genome-wide variants affecting many different genes. Variants in an allele score may be weighted using a disease GWAS (e.g., a second GWAS) that did not include any subjects from the study cohort. Alternatively, variants can be weighted using a GWAS (e.g., a third GWAS) on a biomarker that causally mediates the effect of the drug target on the disease (e.g., LDL cholesterol mediates the effect of PCSK9 on coronary artery disease), yet the biomarker GWAS must not overlap the study cohort.
[0089] The statistical interaction calculator 324 may include programmable instructions that, when executed, cause the processor 314 to determine statistical interactions between pharmacomimetic measures and biomarker stratification factor scores for one or more drug targets. The statistical interactions may predict drug target-specific differential treatment responses for a disease of interest. To determine statistical interactions, the statistical interaction calculator 324 may generate a data table of phenotypes, covariates, and a set of pharmacomimetic measures determined using the systems and methods described above, which may include allele scores for single variants and multiple variants. In response to doing so, the statistical interaction calculator 324 may test for interactions between the pharmacomimetic measures and PGS in association analyses with disease or risk factor traits of interest.
[0090] For example, for each pharmacomimetic measure, the statistical interaction calculator 324 can use statistical software, such as the R programming language, to fit a logistic regression model using the following exemplary formula: disease case-control status ~ age + sex + ancestry PC + (pharmacomimetic measure) + (scaled PGS) + (pharmacomimetic measure):(scaled PGS). Doing so, the statistical interaction calculator 324 can determine the scaled PGS. The model can be implemented using the R programming language, such as using the command glm(model_formula, family="binomial", data=our_genetic_and_phenotypic_data). In the example formula above, the variable to the left of the "~" symbol is the dependent variable (e.g., the variable to be predicted). Each variable to the right of the "~" symbol with a "+" symbol between them is an independent variable (e.g., a variable observed in the data and used to predict the dependent variable). The ":" symbol indicates a variable that is the product (the result of multiplication) of the variables to the left and right of the ":" symbol. This product is the "interaction term." If the pharmacomimetic:scaled polygenic score (PGS) interaction term for one of the independent variables has a statistically significant association with the dependent variable (disease status in this case) when all of the other independent variables, such as age and sex, are considered and controlled for, then the statistical interaction calculator 324 can predict or determine that the drug modeled by the pharmacomimetic will have a different effect between subjects with high vs. low PGS. The statistical interaction calculator 324 can fit a logistic regression model for each pharmacomimetic determined or identified by the pharmacomimetic identifier 322.
[0091] The statistical interaction calculator 324 can use the same program to calculate p-values for interaction terms. If the p-value for the interaction between the pharmacomimetic measure and the PGS is <0.05 / (number of measures tested), the statistical interaction calculator 324 can determine that the interaction is "Bonferroni significant." Alternatively, the statistical interaction calculator 324 can be configured to use a more lenient p-value threshold, at the risk of false positives.
[0092] Such agent-PGS interactions can also be "drug-PGS interactions," since each pharmacomimetic agent is a model of a specific drug mechanism.
[0093] The statistical interaction calculator 324 can generate tables that provide statistically significant drug-PGS interactions in a manner that is easier to interpret than looking at the raw regression model output. For example, the statistical interaction calculator 324 can group subjects in a study cohort by PGS quantiles (e.g., 0-33 percentile, 34-66 percentile, and 67-100 percentile, etc.). Then, for each drug-PGS interaction, the statistical interaction calculator 324 can fit a separate logistic regression model within each PGS quantile using the formula: disease case-control status ~ age + sex + ancestry PC + pharmacomimetic measure of drug. Using these regression outputs, the statistical interaction calculator 324 can construct a table showing the effect of the pharmacomimetic measure on disease risk (and a 95% confidence interval for that effect) within each PGS quantile.
[0094] For example, the statistical interaction calculator 324 can divide study subjects into a finite number of groups (e.g., several subgroups) based on the subjects' PGS scores. Within each group, the statistical interaction calculator 324 can predict disease outcomes based on the patient's genetic factors, specifically genetic measures that mimic the effects of a drug ("pharmaco-mimetic measures"). The statistical interaction calculator 324 can then determine the difference in how strongly the pharmacomimetic measures predict disease outcomes in individuals with high PGS vs. moderate PGS vs. low PGS. This difference would indicate that a drug may have different efficacy in people with high PGS vs. moderate PGS vs. low PGS.
[0095] The statistical interaction calculator 324 can also calculate a statistic called the "treatment effect multiplier," which is the effect of a pharmacomimetic measure on disease risk in the highest quartile divided by the effect for all subjects. The "treatment effect multiplier" can represent how much the average treatment effect in a drug clinical trial can be increased by enrolling only patients in the highest quartile of the PGS rather than all eligible patients. The treatment effect multiplier is one way of comparing drug-PGS interactions and distinguishing "strong" from "weak" interactions.
[0096] The action performer 326 may include programmable instructions that, when executed, cause the processor 314 to take one or more prediction-based actions based on the determined interactions. The one or more prediction-based actions may include at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics. To perform a prediction-based action, the action performer 326 may select patients for treatment. The action performer 326 may select patients based on the likelihood that the patient will benefit from the treatment. For example, the action performer 326 may determine the patient's PGS using the systems and methods described herein and select the patient by determining that the patient's PGS is above a threshold (e.g., a PGS threshold) or within a range (e.g., a quartile, quintile, defined range, etc.). In another example, the action performer 326 may select patients for novel therapeutic mechanisms from a polygenic high-risk subset of patients with a common disease for whom benefit is most or only apparent. This allows for the identification of novel targets and novel therapeutic discoveries. In another example, action execution 326 can select patients from large, commonly treated disease populations who will benefit most from existing drug therapies, and conversely, identify patients who will not derive clinically meaningful benefit. This can improve the pharmaco-economic profile of existing mechanisms, allowing for more cost-effective implementation. In another example, action execution 326 can select candidate clinical trial mechanisms and designs for investigational drugs that are predicted to produce greater benefit in individuals with high PGS. These trials could be conducted over a fraction of the observation period required to demonstrate clinical benefit, because increased event rates and treatment response rates can be expected.
[0097] In some embodiments, treatments can be administered to the selected patients. For example, the action taker 326 can select patients with PGS scores that meet criteria for a particular treatment to treat the disease. Examples of such treatments would be treatments for NASH, coronary heart disease, dry age-related macular degeneration, rheumatoid arthritis, atopic disease, and / or obesity. The action taker 326 can generate a record containing a list of the selected patients. An entity (e.g., a user or clinician) can view the list and administer disease treatments to the patients identified in the record. Examples of such treatments can include injections, therapy, or other treatment techniques. In some embodiments, the action taker 326 can connect to another computer and send the record containing the list of patients or the PGS scores determined by the PC generator 320 to the other computer. The computer can display the scores or list of patients, and an entity or user accessing the computer can use the list and / or scores to direct treatments.
[0098] FIG. 4 shows a flow diagram illustrating a method 400 for in silico drug development and drug activity determination in one embodiment of the present disclosure. Method 400 can be performed on a data processing system (e.g., a client device, or data processing system 302 described with reference to FIG. 3, a server system, etc.). Method 400 may include more or fewer operations, and operations may be performed in any order. Execution of method 400 may enable the data processing system to automatically identify predicted drug activities of different drug targets against disease phenotypes and / or intermediate phenotypes.
[0099] In method 400, at operation 402, a data processing system obtains molecular biomarker stratification factor data. The molecular biomarker stratification factor data can include a plurality of biomarker stratification factors from each of a plurality of subjects and at least one disease phenotype or endophenotype. The molecular stratification biomarker data can include at least one of the following types of data: genomics, translipomics, metabolomics, or proteomics. These data can be obtained by assays specific to each biomarker type. The at least one disease phenotype or endophenotype can include or be associated with non-alcoholic steatohepatitis (NASH), coronary heart disease, dry age-related macular degeneration, rheumatoid arthritis, atopic disease, and / or obesity.
[0100] The molecular biomarker stratification factor data can include data regarding at least one disease phenotype. The data can include patient status at a time point (e.g., whether CAD has been diagnosed, yes / no). The data can also include quantitative disease traits (e.g., LDL, HDL, and Tg measurements at a time point relative to disease diagnosis), medication information (e.g., whether the patient received or is receiving statin or other lipid-lowering therapy at a time point relative to disease outcome), and demographic information (e.g., age at diagnosis, biological sex, etc.).
[0101] In operation 404, the data processing system determines a plurality of values (e.g., numerical values) representing biomarker stratification factor effects. Each value may individually represent how each of the plurality of biomarker stratification factors affects each of at least one disease phenotype or endophenotype. Each biomarker stratification factor effect may be an estimate from a statistical model of the biomarker's effect on disease risk or a continuous disease trait. The data processing system may determine the plurality of values based on external data (e.g., data relating to an individual other than the subject among the plurality of subjects). The data processing system may determine the plurality of values using a logistic regression model for disease (yes / or) versus biomarker (e.g., protein measurement) plus covariates (age, sex, etc.) on the data.
[0102] In operation 406, the data processing system calculates biomarker stratification factor scores for the selected disease phenotypes or endophenotypes. The data processing system can calculate a biomarker stratification factor score for each of the at least one disease phenotype or endophenotype. The data processing system can determine the biomarker stratification factor scores as polygenic risk scores for each disease phenotype. The data processing system can calculate a biomarker stratification factor score for the biomarker stratification factor. For example, the data processing system can thereby determine the biomarker stratification factor score as one of the following: a polygenic risk score for a disease phenotype; a polygenic score for a disease risk factor; a polygenic score for biological pathway activity; a polygenic score for drug target expression; a proteomic risk score for a disease phenotype; a proteomic score for a disease risk factor; a proteomic score for biological pathway activity; a proteomic score for drug target expression; a transcriptomic risk score for a disease phenotype; a transcriptomic score for a disease risk factor; a transcriptomic score for biological pathway activity; a somatic mutation score for drug target expression; a somatic mutation score for a disease phenotype; a somatic mutation score for a disease risk factor; a somatic mutation score for biological pathway activity; or a somatic mutation score for drug target expression. For example, if the stratification factor is a polygenic risk score (PRS), a statistical model may be trained using summary statistics from published studies of large cohorts or from meta-analyses of multiple studies based on large and diverse cohorts (none of which include UK Biobank). The data processing system can then apply the weights estimated from the model to datasets (e.g., UK Biobank) that are orthogonal to the one used for training.
[0103] In operation 408, the data processing system calculates a pharmacomimetic gene score for each drug target for which the data processing system is attempting to determine drug activity. The pharmacomimetic gene score can represent genes for which two or more pharmacomimetic variants have been identified (e.g., functional variants of a gene that are strongly and significantly associated with a given disease). The data processing system can determine the pharmacomimetic gene score by weighting each variant by its estimated effect on the phenotype and summing the weighted values.
[0104] In operation 410, the data processing system identifies a predicted drug activity for each drug target for at least one disease phenotype or endophenotype of the subset of biomarker stratification factor distributions. The data processing system can identify the predicted drug activity based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with the disease phenotype or endophenotype. For example, the data processing system can utilize a regression model (e.g., a linear regression model, a logistic regression model, a Cox proportional hazards model, etc.) including a predictor defined by a multiplicative interaction between the pharmacomimetic gene score and the biomarker score to perform the association analysis with each drug target and disease outcome phenotype. The predictor can be the predicted drug activity.
[0105] The data processing system can identify subsets of biomarker stratification factor distributions using one or more thresholds. The one or more thresholds can be chosen to define different subsets depending on the context of use. In one example, it may be optimal to use the 25th percentile of the CAD polygenic risk score to define a subset of patients predicted to experience exceptional clinical benefit from therapy X. In another example, it may be optimal to use the 33% (top tertile) of the CAD PRS to define a subset of patients who will benefit most from therapy Y. For other disease states with different PRSs, thresholds will also vary. The optimal threshold definition will depend on factors such as: the estimated effect size of the subset versus all participants in the therapy; the screening failure rate (if the threshold is set too conservatively in a clinical trial, e.g., 5%, more patients will have to be screened to enroll a minority); etc. The subsets can be from individual levels of cohort data, such as the top 10%, 20%, 30%, etc. of the distribution.
[0106] 5 shows a flow diagram illustrating a method 500 for in silico drug development and drug activity determination in one embodiment of the present disclosure. Method 500 can be performed by a data processing system (e.g., a client device, or data processing system 302 shown and described with reference to FIG. 3, a server system, etc.). Method 500 may include more or fewer operations, and operations may be performed in any order. By performing method 500, the data processing system could automatically identify predicted drug activities of different drug targets against a disease phenotype or intermediate phenotype.
[0107] In method 500, a data processing system obtains molecular biomarker stratification factor data and at least one disease phenotype or endophenotype at operation 502. The molecular biomarker stratification factor data can include multiple biomarker stratification factors. The data processing system can obtain the molecular biomarker stratification factor data from each of a plurality of subjects.
[0108] At operation 504, the data processing system determines biomarker stratification factor effect values for each disease phenotype or endophenotype. Each value may individually represent how each of a plurality of biomarker stratification factors affects at least one disease phenotype or endophenotype. The data processing system may determine the plurality of values based on external data (e.g., data relating to an individual separate from the subjects). At operation 506, the data processing system constructs a matrix based on the biomarker stratification factor data (e.g., biomarker stratification factors), disease phenotypes, or endophenotype values. At operation 508, the data processing system performs principal component analysis on the biomarker stratification factor x phenotype-associated phenotype biomarker stratification factor effect matrix. At operation 510, the data processing system calculates a biomarker stratification factor score for each principal component. At operation 512, the data processing system calculates a pharmacomimetic gene score for each drug target. In operation 514, the data processing system identifies predicted drug activity for each drug target for one or more phenotypes or endophenotypes in the subset of biomarker stratification factor distributions. The data processing system can identify predicted drug activity based on a statistical interaction between the polygenic score calculated for each principal component and the pharmacomimetic gene score for each drug target in an association analysis with the outcome phenotype or outcome endophenotype.
[0109] FIG. 6 shows a flow diagram illustrating a method 600 for in silico drug development and drug activity determination in one embodiment of the present disclosure. Method 600 may be performed by a data processing system (e.g., a client device, or data processing system 302 shown and described with reference to FIG. 3, a server system, etc.). Method 600 may include more or fewer operations, and operations may be performed in any order. Execution of method 600 may enable the data processing system to automatically determine interactions between pharmacomimetics and take predictive-based action based on the determined interactions. One or more of the operations of methods 400, 500, and / or 600 may also be performed during the operation of any other of methods 400, 500, and / or 600.
[0110] In method 600, in operation 602, a data processing system generates a plurality of principal components (PCs) corresponding to genetic ancestry data (e.g., genetic data) of subjects in a study cohort. The data processing system can generate any number of PCs. For example, the data processing system can generate five PCs, at least six PCs for a European ancestry cohort, and / or 10 to 40 PCs for a multi-ancestry cohort (e.g., 20 for UK Biobank, and / or perhaps up to 40 for a more diverse cohort). The study cohort can include at least 200 control subjects. In some embodiments, the study cohort can include at least 950 control subjects or at least 990 control subjects. The data processing system can determine whether a subject in the study cohort is a case or control for a disease of interest, for example, based on a flag on the subject indicating whether the subject is a case or control. The genetic ancestry data can be based on genotyping arrays or whole-genome sequencing. The data processing system can obtain the genetic ancestry data, for example, from a publicly available database or through any other method.
[0111] In operation 604, the data processing system generates a biomarker stratification factor score for each subject in the study cohort. The data processing system can generate the biomarker stratification factor score based on at least (i) the PC, and (ii) the weight of the biomarker stratification factor for the disease of interest. The biomarker stratification factor score can be specific to the disease of interest. For example, the "raw PGS" (e.g., biomarker stratification factor score) for subject S can be calculated as follows: ((variant weights) *It can be defined as the sum over each variant in a PGS variant weight table, where weight ((number of copies of that variant in S)) is the sum of the weights over each variant in the PGS variant weight table. For example, consider three variants A, B, and C with weights of 1, 2, and 3, respectively. If subject S has 2, 0, and 1 copies of variants A, B, and C, their unscaled PGS can be calculated as 2*weight(A)+0*weight(B)+1*weight(C)=2*1+0*2+1*3=5. PGS variant weights can be calculated independently of the genetic ancestry data corresponding to the study cohort.
[0112] The data processing system can determine an "ancestry-normalized PGS," which can be defined as the residual of a linear regression model where the outcome is the raw PGS and the predictors are the PCs of the genetic ancestry selected in act 602. This normalization can correct for differences in mean PGS values between populations.
[0113] The data processing system can determine a "scaled PGS," which can be defined as ((ancestry-normalized PGS - ancestry-normalized PGS of all subjects in the cohort) / (standard deviation of ancestry-normalized PGS of all subjects in the cohort)). The scaled PGS values will therefore be approximately normally distributed. The scaled PGS can be used for all subsequent analyses. The task of generating and scaling a valid and comparable PGS across populations can be accomplished using any method.
[0114] In operation 606, the data processing system determines which of the plurality of disease-associated variants is a pharmacomimetic for the disease of interest. A disease-associated variant is likely to be a pharmacomimetic if it modulates the function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to predict a second effect of the drug on the phenotype.
[0115] To determine which of multiple disease-associated variants are pharmacomimetics for a disease of interest, the data processing system may run a program for the disease of interest (e.g., GCTA COJO-SLCT) to identify "conditionally independent disease-associated variants."
[0116] For example, genetic variants that are close to each other on a DNA molecule (e.g., within roughly 500,000 base pairs) may be inherited together. Variants that are typically inherited together are said to be in "linkage disequilibrium." If a single variant contributes to the risk of a disease, a GWAS will find both a statistical association of that variant (the "causative variant") with the disease and a statistical association of the disease with other variants that are in linkage disequilibrium (co-inherited) with the causative variant. Thus, if a GWAS has a group of variants that are close together in their locations on a DNA molecule, and many of these variants are associated with a disease, the data processing system can distinguish between two scenarios: 1) all of the apparent disease-associated variants are inherited together, meaning there is likely a single "causative variant," or 2) there are two or more distinct groups of variants, each with one or more variants, such that the variants in each group are inherited together, but the inheritance of one group is independent of the inheritance of the other. In the latter case, there is likely one "causative variant" per group of linked variants.
[0117] To distinguish between "single causative variant" and "multiple causative variant" scenarios, the data processing system can implement a regression function. For example, a GWAS (e.g., a first GWAS) can first be performed as usual. The data processing system can identify the variant with the most significant association with the disease outcome across the entire genome (e.g., smallest p-value). The data processing system can then perform the GWAS again, but this time conditioning all of the variant-disease association tests on the genotype of the most significant variant. The data processing system can identify the variant with the next most significant association with the disease outcome. The data processing system can then perform the GWAS again, but this time conditioning all of the variant-disease association tests on the genotypes of the most significant variant and the second most significant variant ... -8 This method can be repeated until it is determined that there are no disease-associated variants below a genome-wide significance threshold (e.g., a threshold). Using a lower threshold reduces the likelihood that the analysis is a fluke, but increases the likelihood that the data processing will fail to identify significant variants, and vice versa. The data processing system can use any threshold. At the end of the method, the set of variants identified by the data processing system can be called "conditionally independent disease-associated variants," because each variant is still associated with disease after conditioning on all genotypes of the previous variant.
[0118] Returning to the "single causative variant" and "multiple causative variant" scenarios, if during stepwise regression the data processing system can identify only a single "conditionally independent disease-associated variant" in any DNA region, then the "single causative variant" scenario is more likely to be true. However, if stepwise regression identifies multiple "conditionally independent disease-associated variants," then the "multiple causative variant" scenario is more likely to be true.
[0119] If a disease (e.g., coronary artery disease) is associated with multiple biomarkers (e.g., LDL cholesterol and blood pressure), the data processing system may use multiple independent acceptance criteria to identify biomarker-associated variants, thereby increasing the number of variants the data processing system identifies. For example, the data processing system may use multiple independent acceptance criteria to identify biomarker-associated variants, such as p < 5 x 10 for the disease. -8 Identify variants with p < 5 x 10 for at least one disease-associated biomarker -8 p < 5 x 10 for diseases that also have -6 The data processing system can use any criteria to identify variants. Variations in the acceptance criteria allow the data processing system to identify a wide range of variants using the COJO-SLCT analysis.
[0120] The data processing system can generate a table, where each row corresponds to one subject in the study cohort. The columns of the table are described as follows: Whether the subject is a case or control for the disease. Cases are assigned a value of "1" and controls a value of "0." Subjects that are neither cases nor controls are excluded from the table. b. Age and sex of the subject. c. The PCs of the genetic ancestors from operations 602 or 604. d. Scaled PGS of disease risk from behavior 604. e. Number of subjects with each alternative allele of an independent disease-associated variant. i. Here, "alternative allele" refers to a reference genome from the Genome Reference Consortium, e.g., GRCh37 or GRCh38. In a reference genome, each variant has a "reference allele." Thus, a biallelic variant is a genomic site where an individual has either the "reference allele" or the "alternative allele" at that site. ii. Some sites are "polyallelic," meaning that there are three or more alleles in the population. For computational purposes, each alternative allele is treated as a separate biallelic variant (e.g., if there are "reference," "alternative-1," and "alternative-2," it would be treated as two biallelic variants: (reference, alternative-1), and (reference-alternative-2)).
[0121] The data processing system can identify which disease-associated variants are "pharmaco-mimetics" (e.g., when a variant modulates the function or expression of a drug target such that the effect of the variant on a human phenotype is likely to be predictive of the effect of the drug on the human phenotype). Depending on the configuration, the data processing system can use different criteria and / or priorities to identify pharmacomimetics from identified disease-associated variants. In one example, a user may be interested in clinical-stage drugs for a disease. In this case, the data processing system can compile a table of clinical-stage drugs and their targets using, for example, public resources such as clinicaltrials.gov and / or proprietary databases such as Cortellis. The data processing system can determine that variants that are close to a drug target gene (e.g., within 150 kb of the transcription start site of the gene) or that map to the drug target are pharmacomimetics. In another example, a user may be interested in known and novel drug targets for antibody modalities. In this case, the data processing system can compile a list of genes encoding druggable proteins (e.g., proteins secreted or localized at the cell surface). The data processing system can determine that variants mapped to antibody-druggable proteins are pharmacomimetics. The data processing system can determine that variants mapped to antibody-druggable proteins are pharmacomimetics based on the compiled table (e.g., based on determined variants recorded in association with antibody-druggable proteins).
[0122] Optionally, genetic evidence may be required to support the hypothesis that target inhibition, rather than target activation, confers benefit. This assessment could consider data from expression quantitative trait loci (eQTL) and protein quantitative trait loci (pQTL) datasets that inform how genetically determined changes in target expression affect disease risk.
[0123] In some embodiments, there are two or more suitable variants for a given drug mechanism. In this case, the data processing system can aggregate variants (e.g., all of the variants) identified as pharmacomimetics associated with a given drug mechanism and combine them into an allele score to increase the statistical power of detecting drug-PGS interactions. To do so, for example, the data processing system can calculate an allele score in the same or similar manner as a PGS (e.g., assigning a weight to each variant and calculating, for each subject in the cohort, the sum of (weight of the variant) x (number of subjects with that variant) across each variant). The difference between an allele score and a PGS is that an allele score consists of variants that modulate the function or expression of a single drug target gene, while a PGS can consist of genome-wide variants affecting many different genes. Variants in an allele score may be weighted using a disease GWAS (e.g., a second GWAS) that did not include subjects from the study cohort. Alternatively, variants can be weighted using a GWAS (e.g., a third GWAS) for a biomarker that causally mediates the effect of the drug target on the disease (e.g., LDL cholesterol mediates the effect of PCSK9 on coronary artery disease). However, the biomarker GWAS should not overlap the study cohort. Allele scores can also be used in place of PGS to implement the systems and methods described herein.
[0124] In operation 608, the data processing system determines statistical interactions between the pharmacomimetic measures and biomarker stratification factor scores for one or more drug targets. This statistical interaction may predict drug target-specific differential treatment responses for the disease of interest. To determine the statistical interactions, the data processing system can generate a data table of phenotypes (e.g., nonalcoholic steatohepatitis (NASH), coronary heart disease, dry age-related macular degeneration, rheumatoid arthritis, atopic disease, obesity, etc.), covariates, and a set of pharmacomimetic measures determined using the above-described systems and methods, which may include single variant or multiple variant allele scores. In response to doing so, the data processing system can test for interactions between the pharmacomimetic measures and PGS (or allele scores) in association analyses with the disease or risk factor trait of interest.
[0125] For example, for each pharmacomimetic measure, the data processing system can use statistical software, such as the R programming language, to fit a logistic regression model using the following exemplary formula: disease case-control status ~ age + sex + ancestry PC + (pharmacomimetic measure) + (scaled PGS) + (pharmacomimetic measure):(scaled PGS). The model can be implemented using the R programming language, such as using the command glm(model_formula, family=binomial, data=our_genetic_and_phenotypic_data). In the example formula above, the variable to the left of the "~" symbol is the dependent variable (e.g., the variable to be predicted). Each of the variables to the right of the "~" symbol with a "+" symbol between them is an independent variable (e.g., a variable observed in the data and used to predict the dependent variable). The ":" symbol indicates a variable that is the product (the result of multiplication) of the variables to the left and right of the ":" symbol. This product is the "interaction term." If the pharmacomimetic:scaled polygenic score (PGS) interaction term—one of the independent variables—has a statistically significant association with the dependent variable (in this case, disease status) when controlling for all of the other independent variables, such as age and sex, then the data processing system can predict or determine that the drug modeled by the pharmacomimetic will have a different effect between subjects with high vs. low PGS. For each pharmacomimetic determined or identified by the data processing system, the data processing system can fit a logistic regression model, e.g., based on the PGS of different subjects.
[0126] The data processing system can also calculate p-values for interaction terms using the same program. If the p-value for the interaction between a pharmacomimetic measure and a PGS is <0.05 / (number of measures tested), the data processing system can determine that this interaction is "Bonferroni significant." Alternatively, the data processing system can be configured to use a more lenient p-value threshold, even at the risk of false positives.
[0127] A tool-PGS interaction can also be a "drug-PGS interaction," because each pharmacomimetic tool is a model of a specific drug mechanism.
[0128] The data processing system can generate tables that provide statistically significant drug-PGS interactions in a manner that is easier to interpret than looking at the raw regression model output. For example, the data processing system can first group the subjects in a study cohort by PGS quantiles (e.g., 0-33rd percentile, 34-66th percentile, and 67-100th percentile). Then, for each drug-PGS interaction, the data processing system can fit a separate logistic regression model within each PGS quantile using the formula: disease case-control status ~ age + sex + ancestry PC + pharmacomimetic measure of drug. Using these regression outputs, the data processing system can construct tables showing the effect of the pharmacomimetic measure on disease risk (and the 95% confidence interval for that effect) within each PGS quantile.
[0129] For example, a data processing system can divide study subjects into a finite number of groups (e.g., several subgroups) based on the subjects' PGS values. Within each group, the data processing system can predict disease outcomes based on the patient's genetic factors, specifically genetic measures that mimic the effects of a drug ("pharmaco-mimetic measures"). In doing so, the data processing system can determine the difference in how strongly the pharmacomimetic measures predict disease outcomes in individuals with high PGS vs. moderate PGS vs. low PGS. This difference would indicate that a drug may have different efficacy in people with high PGS vs. moderate PGS vs. low PGS.
[0130] The data processing system can also calculate a statistic called the "treatment effect multiplier," which is the effect of a pharmacomimetic measure on disease risk in the top quartile divided by the effect in all subjects. The "treatment effect multiplier" can represent how much the average treatment effect in a drug clinical trial could be increased by enrolling only patients in the top quartile of the PGS rather than all eligible patients. The treatment effect multiplier is one way of comparing drug-PGS interactions and distinguishing "strong" from "weak" interactions.
[0131] In operation 610, the data processing system performs one or more prediction-based actions. The data processing system can perform the one or more prediction-based actions based on the determined interaction. The one or more prediction-based actions can include at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics. To perform the prediction-based actions, the data processing system can select patients for treatment. The data processing system can select patients based on the likelihood that the patient will benefit from the treatment. For example, the data processing system can determine the patient's PGS using the systems and methods described herein and select patients by determining that the patient's PGS is above a threshold (e.g., a PGS threshold) or within a range (e.g., a quartile, quintile, defined range, etc.). In another example, the data processing system can select patients for novel therapeutic mechanisms from polygenic high-risk subsets of patients with a common disease for which benefit is most or only evident. This enables novel target identification and novel therapeutic discovery. This application can be called "polygenic therapeutic target ID." In another example, a data processing system can select patients from large populations eligible for common disease treatments who will most likely benefit from existing drug therapies, and conversely, identify patients who will not likely derive clinically meaningful benefit. This can improve the pharmacological-economic profile of existing mechanisms, enabling more cost-effective implementation. This can be called "polygenic pharmacogenomics." In another example, a data processing system can select candidate clinical trial mechanisms and designs for investigational drugs predicted to produce greater benefit in individuals with high PGS. These trials could be conducted within a fraction of the observation years required to demonstrate clinical benefit, because increased event rates and treatment response rates can be expected. This application can be called "polygenic therapeutic development." In some embodiments, injections, therapies, and / or treatment techniques can be administered to selected individuals when one or more prediction-based actions are taken.
[0132] In some examples, a patient can be treated using the systems and methods described herein. For example, the patient may visit a clinic. A clinician can collect a blood sample, saliva sample, or tissue sample from the patient. The clinician can then use the collected sample to perform a genotyping array on the patient. A data processing system implementing the systems and methods described herein can analyze the patient's genetic data to generate principal components (PCs) that reflect the patient's genetic ancestry. The patient can undergo testing to identify relevant biomarkers for a disease of interest (e.g., diabetes, cancer, cardiovascular disease, psychiatric disorders, autoimmune diseases, neurological disorders, infectious diseases, asthma, allergies, etc.). These biomarkers are quantifiable biological parameters, such as blood glucose levels, cholesterol levels, and specific protein markers, that may characterize the patient depending on the disease. Based on the results of the biomarker testing and genetic ancestry data (e.g., PCs), the data processing system can calculate the patient's biomarker stratification factor score for the disease of interest (e.g., a disease of interest from the list above).
[0133] The data processing system can scan the patient's genetic data for disease-associated variants. The data processing system can then identify genes or genetic variants in the patient's genetic data that are associated with a disease for which the data processing system has determined a biomarker stratification factor score for the patient. For example, the data processing system can identify disease-associated variants for the disease using the systems and methods described herein. The data processing system can use the identified disease-associated variants as keys for querying through the patient's genetic array data. Based on this query, the data processing system can identify any disease-associated variants in the patient's genetic data.
[0134] The data processing system can determine which of the identified disease-associated variants are pharmacomimetics (e.g., whether these genetic variants are predictive of how a patient will respond to a particular drug or treatment). For example, before determining a treatment for a patient, the data processing system may have identified a list of pharmacomimetics for the disease of interest using the systems and methods described herein. The data processing system can compare the disease-associated variants identified in the patient's genetic makeup to this list of pharmacomimetics. Based on this comparison, the data processing system can identify a set of pharmacomimetics for the patient. In some embodiments, the data processing system can identify pharmacomimetics for a patient without identifying disease-associated variants, for example, by querying the patient's genetic data for pharmacomimetics related to the disease.
[0135] The data processing system can determine one or more treatments for the patient based on the patient's pharmacomimetic measure and biomarker stratification factor score. The data processing system can do this, for example, based on a statistical interaction the data processing system has previously determined for the pharmacomimetic measure and biomarker stratification factor score. For example, the data processing system can store records (e.g., records based on PGS or logistic regression of the biomarker stratification factor score and the pharmacomimetic measure) indicating that individuals with a high polygenic risk score or biomarker stratification factor score (e.g., above a threshold) for the disease of interest will likely have a high positive response to a therapeutic inhibitor (represented by the pharmacomimetic measure), and individuals with a low polygenic risk score or biomarker stratification factor score (e.g., below a threshold) for the disease of interest will likely have a low positive response to the inhibitor. A patient may have a high polygenic risk score or biomarker stratification factor score. Thus, the data processing system can generate a record (e.g., a file, a notification, an alert, a user interface, a data structure, etc.) ordering or recommending treatment of a patient with a therapeutic inhibitor. The data processing system can display the record on a user interface of a client device accessed by a clinician. The data processing system can recommend either form of treatment based on the interaction between the patient's pharmacomimetic measure and the polygenic risk score biomarker stratification factor score.
[0136] The clinician can treat the patient based on this recommendation. For example, the clinician can view a recommendation to treat the patient with an inhibitor. The clinician can administer the treatment by administering a pill, capsule, giving an injection or set of injections, or using a gene editing tool. The clinician can administer any type of treatment based on the patient's polygenic risk score or biomarker stratification factor score, disease type, and / or pharmaco-mimetic means.
[0137] In some embodiments, the data processing system can automatically perform a treatment (e.g., using a gene editing tool or another treatment device or mechanism) in response to the treatment decision. For example, the data processing system can determine a treatment to inhibit a specific gene to treat a disease in a patient, as described herein. The data processing system can configure a treatment machine (gene editing tool) to automatically implement the treatment for the patient in response to the decision, such as by controlling a machine to perform the treatment (e.g., inhibiting or editing a specific gene, or administering an injection).
[0138] In another example, the systems and methods described herein can be used to conduct clinical trials to determine the efficacy of a drug candidate. For example, the data processing system can identify a disease of interest that the drug candidate is intended to treat. The data processing system can identify a pharmacomimetic for the disease of interest. The data processing system can identify a pharmacomimetic for the drug candidate, e.g., based on user input. The data processing system can receive genetic data for a research cohort. The data processing system can use the systems and methods described herein to determine a polygenic risk score for the research cohort based on the genetic data.
[0139] The data processing system can select participants for the clinical trial based on the polygenic risk score and an interaction between the polygenic risk score and a pharmacomimetic measure of the candidate drug. For example, the data processing system can identify an interaction between the pharmacomimetic measure and the polygenic risk score. From this identified interaction, the data processing system can determine that there is a high positive interaction between a high polygenic risk score (e.g., a polygenic risk score above a threshold) and the pharmacomimetic measure. Accordingly, the data processing system can filter participants from the study cohort and identify participants with high polygenic risk scores for the study cohort. The data processing system can generate a record including a list of patients identified for the clinical trial. The data processing system can present the list of patients to a clinician on a user interface and / or take action on the participants for the clinical trial as described above.
[0140] FIG. 7 illustrates a sequence 700 for in silico drug development and drug activity determination, according to one embodiment of the present disclosure. Sequence 700 includes an illustration of a matrix that can be used to perform the systems and methods described herein. A data processing system (e.g., a client device, data processing system 302 shown and described with reference to FIG. 3, a server system, etc.) can perform the operations to generate the matrix of sequence 700. Sequence 700 may include more or fewer operations, and the operations may be performed in any order. By performing method 700, the data processing system may automatically determine interactions between pharmacomimetics and take predictive-based actions based on the determined interactions.
[0141] For example, a data processing system running sequence 700 can generate variant effect matrix 702. The data processing system can generate variant effect matrix 702 to have column 704a (column 704) and rows 706a-n (rows 706). n can be any number and can be different between column 704 and row 706. Each column 704 can correspond to a different phenotype. In some embodiments, column 704a can correspond to a reference phenotype, and columns 704b-n can be phenotypes biologically related to the reference phenotype in column 704a. Each row 706 can correspond to the variant effect of a different variant on the phenotype in each column 704. The data processing system can determine values for the variant effects and insert the values into variant effect matrix 702.
[0142] The data processing system can perform a truncated principal component analysis (PCA) 708 on the variant effect matrix 702 to generate a variant-PC loading matrix 710. The data processing system can generate the variant-PC loading matrix 710 to have columns 712a-n (columns 712) and rows 714a-n (rows 714). Each column 712 can correspond to a different principal component. Each row 706 can correspond to a different variant. The intersection value between the column 712 and the row 714 can indicate the variant loading on a particular principal component.
[0143] The data processing system can generate or calculate a weight (e.g., a new weight) for each PC and variant combination 716. The data processing system can use the weights to generate a new PGS for the phenotype. The data processing system can do this using the following formula: PC i to new weight of variant = old weight of variant * (PC i (variant loadings on 2 / (total j =[PC j Variant loadings on 21...k). The old weights of the variants may be the weights of the variants used to determine the existing PGS718. The data processing system can use the newly calculated weights to calculate a PGS of the phenotype using the systems and methods described herein. The data processing system can use the PGS to identify patients for treatment and / or an entity can treat patients based on the PGS. Variant data for polygenic scores
[0144] Numerous approaches are available that combine information across loci to assess polygenic risk scores (PRS) for a variety of conditions. Numerous studies have shown that PRS can predict disease status in research-based case-control studies. See, for example, Mavaddat N, et al., Am J Hum Genet. 2019; 104:21-34; Wray N Ret al., Nat Genet. 2018; 50:668-81. More convincingly, this prediction is also valid in population-based cohort studies and electronic health record-based studies, particularly for psychiatric disorders. Musliner KL et al. JAMA Psychiatry. 2019;76:516-25; Lewis CM and Hagenaars SP, JAMA Psychiatry. 2019;76:470-2; Zheutlin AB, et al., Am J Psychiatry. 2019;176(10):846-55. A PRS can be formed from a set of independent risk variants for a disorder based on current evidence from the largest or most informative genome-wide association studies. For each individual, the number of risk alleles (0, 1, or 2) for each variant is summed and weighted by its effect size (i.e., log(OR) for binary traits or beta coefficient for continuous traits). The output is a single score of each individual's genetic burden for a disease or continuous trait.
[0145] Much of the research on polygenic scores comes from research studies in cardiovascular disease, type 2 diabetes, breast and prostate cancer, and Alzheimer's disease (Lambert SA et al. Hum Mol Genet. 2019;28(R2):R133-42). A study using UK Biobank demonstrated that PRS based on variant data can identify a percentage of patients at at least three times higher risk for coronary artery disease, atrial fibrillation, type 2 diabetes, inflammatory bowel disease, and breast cancer, with the proportion of individuals identified ranging from 1.5 to 8%, depending on the disorder (Khera AV et al., Nat Genet. 2018;50:1219-24). While these effects appear modest, PRSs have the potential for greater clinical importance because they can identify a substantially larger proportion of the population at high risk for disease than single-gene mutations. definition
[0146] Unless otherwise defined, all technical terms, notations, and other technical and scientific terms or terminology used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms that have a commonly understood meaning are defined herein for clarity and / or ease of reference, and the inclusion of such definitions herein should not necessarily be construed as representing a substantial departure from what is commonly understood in the art.
[0147] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, "a" or "an" means "at least one" or "one or more." It is understood that the aspects and variations discussed herein include "consisting of" and / or "consisting essentially of" aspects and variations.
[0148] Throughout this disclosure, various aspects of claimed subject matter are presented in range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges and individual numerical values within that range. For example, when a range of values is recited, it is understood that every intermediate value between the upper and lower limits of that range, as well as any other stated or intermediate value within the stated range, is also encompassed within the claimed subject matter. The upper and lower limits of these smaller ranges may also be individually included in the smaller ranges, and are encompassed within the claimed subject matter, subject to any specifically excluded limit in the stated range. When a stated range includes one or both of the limits, ranges excluding either or both of those included limits are also encompassed within the claimed subject matter. This applies regardless of the breadth of the range.
[0149] As used herein, "about" when referring to a measurable value such as an amount, length of time, etc., is intended to mean encompassing a variation of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and even more preferably ±0.1% from the specified value, where such variations are appropriate for performing the disclosed method.
[0150] The terms "coronary artery disease," "coronary heart disease," or "CAD" refer to a cardiac condition caused by reduced blood flow to the heart, resulting in decreased oxygen and nutrients to the myocardium. Medical manifestations of this condition include, but are not limited to, angina pectoris, myocardial infarction, and coronary artery revascularization.
[0151] The term "in silico modeling" refers to the use of computer models to predict drug effects and / or health outcomes under various scenarios.
[0152] "Molecularly stratified biomarker data" includes at least one of the following types of data: genomics, transcriptomics, metabolomics, or proteomics, which are obtained by assays specific to each biomarker type.
[0153] A disease phenotype is based on knowledge of a patient's disease status at any time point (e.g., whether a patient has been diagnosed with CAD, yes / or). Disease phenotypes may also include quantitative disease traits (e.g., LDL, HDL, and Tg measurements at any time point relative to disease diagnosis), medication information (e.g., whether a patient has taken or is taking statins or other lipid-lowering therapy at any time point for disease outcomes), and demographic information (e.g., age at diagnosis, biological sex, etc.).
[0154] As used herein, "biological pathway" refers to either (1) a custom-made set of genes whose protein products are known to interact in a biological pathway (e.g., complement pathway, JAK / STAT pathway, MAPK pathway, etc.), or (2) an unsupervised learning approach (e.g., principal component analysis (PCA)) that derives patterns from data such that gene clusters comprising a pathway can be identified. In some embodiments, approach (1) is based on canonical pathways identified in the literature and external pathway databases, and approach (2) is a data-driven analysis that leads to pathway identification.
[0155] "Biological pathway activity," as used herein, is defined in the form of information from the literature or external databases that provides evidence of gene or protein expression indicative of pathway activity (e.g., "upregulation" or "downregulation" of pathway X in disease Y).
[0156] As used herein, "drug target expression" refers to protein expression measured in participants from a large cohort (e.g., UK Biobank) where the analyzed protein is a known drug target or a potentially druggable target (even if not yet drugged). In some embodiments, a biomarker stratification factor score (e.g., polygenic score, proteomic score, transcriptomic score, or somatic mutation score) for "drug target expression" is calculated in the same way as any other quantitative trait where the model outcome is a quantitative measure.
[0157] As used herein, "disease outcome" refers to a clinical event that is a result of having a disease. For example, coronary artery disease (CAD) or atherosclerotic cardiovascular disease (ASCVD) are diseases, but myocardial infarction or stroke are events that occur in patients with ASCVD / CAD and are therefore "disease outcomes."
[0158] The term "drug target," as used herein, refers to any gene or gene product (e.g., RNA or polypeptide) implicated in a disease or disorder. Non-limiting examples include various proteins, such as enzymes, oncogenes and their polypeptide products, and cell cycle regulatory genes and their polypeptide products.
[0159] As used herein, the term "external data" refers to data from an independent set of individuals or populations that are not used in determining the molecular biomarker stratification factor data score. In some embodiments, the external data is used to assess the predictive power of the methods described herein. In some embodiments, the external data is used to prevent overfitting during the determination of the molecular biomarker stratification factor data score. At a minimum, the external data is from a large cohort of participants from an observational institution (e.g., UK Biobank or eMERGE). In some embodiments, the external data is divided into a training subset and a validation subset. In some embodiments, estimates of biomarker effect are validated in multiple studies, which may include both observational studies and health system patient cohorts.
[0160] As used herein, the term "genetic variant" refers to an alteration, variant, or polymorphism in a subject's nucleic acid sample or genome. Such alteration, variant, or polymorphism may be with respect to a reference genome, which may be the reference genome for the species (e.g., hGI9 or hG38 for humans), or with respect to the subject or other individual. Variations include one or more single nucleotide variations (SNVs), insertions, deletions, repeats, small insertions, small deletions, small repeats, structural variant junctions, variable length tandem repeats, and / or flanking sequences, although copy number variants (CNVs), transversions, gene fusions, and other rearrangements are also forms of genetic variation. Variations may be single nucleotide variations (SNVs), insertions or deletions (indels), repeats, copy number variations (CNVs), transversions, or combinations thereof.
[0161] As used herein, a "genome-wide association study" (GWA study, or GWAS) refers to an observational study of a genome-wide set of genetic variants among different individuals to determine whether any variants are associated with a trait. In some embodiments, GWAS studies focus on the relationship between genetic variants (e.g., single nucleotide polymorphisms (SNPs)) and phenotypic traits (e.g., diseases). In some embodiments, the genetic variants are somatic genetic variants. In some embodiments, the genetic variants are germline genetic variants.
[0162] In some embodiments, GWA studies compare DNA from participants with various phenotypes for a particular trait or disease. These participants may be people with a disease (cases) and similar people without the disease (controls), or they may be people with different phenotypes for a particular trait, such as blood pressure. This approach is known as phenotype-first, in which participants are first classified by their clinical symptoms, as opposed to genotype-first. Each person provides a DNA sample from which millions of genetic variants are determined (e.g., using SNP arrays or sequencing). If there is significant statistical evidence that one type of variant (one allele) is more frequent in people with the disease, that variant is said to be associated with the disease. Associated SNPs are then considered to mark regions in the human genome that may influence disease risk.
[0163] As used herein, the phrase "molecular biomarker stratification factor" or "biomarker stratification factor" refers to a molecular marker that differs between two different states (eg, healthy vs. diseased).
[0164] As used herein, the phrase "molecular biomarker stratification factor distribution" or "biomarker stratification factor distribution" refers to a subset obtained from individual-level cohort data (e.g., data from UK Biobank), which may be defined as the top 10%, 20%, 30%, etc., of the distribution. The threshold chosen to define a subset of a molecular biomarker stratification factor distribution will depend on the application case. For example, to define a subset of patients predicted to have exceptional clinical benefit from a given treatment, it may be optimal to use the 25th percentile of a molecular biomarker stratification factor score (e.g., a polygenic score) for coronary heart disease. Alternatively, the 33rd percentile of a molecular biomarker stratification factor score for coronary heart disease may be used to define a subset of patients likely to benefit most from a second, different, given treatment. In other disease settings with different molecular biomarker stratification factor scores, thresholds will also be different. The definition of the optimal threshold will depend on factors such as: the estimated effect size of the subset receiving the therapy versus all participants in the therapy; and the screening failure rate ("screening failure" means that candidates who are screened, i.e., checked for eligibility, do not meet the clinical inclusion / exclusion criteria in the trial; "screening failure rate" is the number of ineligible candidates divided by the number of candidates screened). If the threshold is set too conservatively, e.g., at 5%, then many more patients will have to be screened to enroll the few eligible candidates.
[0165] As used herein, the phrase "molecular biomarker stratification factor effect" or "biomarker stratification factor effect" refers to the magnitude of the effect of a biomarker on disease risk or a continuous disease trait, as estimated from a statistical model.
[0166] As used herein, the phrase "molecular biomarker stratification factor score" or "biomarker stratification factor score" refers to a score that is an accumulation of data from multiple (hundreds, thousands, tens of thousands, or more) potential biomarker stratification factors that may be used to predict an individual's risk of disease. In some embodiments, a statistical model is trained using summary statistics from published studies of large cohorts or from meta-analyses of multiple studies of large and diverse sets of cohorts. Weights estimated from this model are then applied to a dataset orthogonal to the one used for training (e.g., UK Biobank). For example, a Bayesian high-dimensional linear regression model may be used for a set of populations, with body mass index (BMI) as the dependent variable and a genome-wide set of SNPs and their effect sizes from the published BMI study from the GIANT consortium (Locke et al. 2015) as input. A population-specific reference panel (e.g., 1000 genomes) is used to infer linkage disequilibrium and derive a population-specific PRS. Linear regressions of the normalized PRS for each population can then be fitted and applied to individual-level genotype and phenotype data from UK Biobank.
[0167] In some embodiments, the molecular biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, and the biomarker stratification factor scores comprise, alternatively consist essentially of, or further consist of polygenic scores.
[0168] In some embodiments, the molecular biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of a proteomic variant, and the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of a proteomic score.
[0169] In some embodiments, the molecular biomarker stratification factor comprises, alternatively consists essentially of, or further consists of a transcriptional variant, and the biomarker stratification factor score comprises, alternatively consists essentially of, or further consists of a transcriptomic risk score.
[0170] In some embodiments, the molecular biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, and the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of a somatic mutation risk score.
[0171] As used herein, "pharmacomimetics" refers to exhibiting a similar response or lack of response to a therapeutic agent or class of therapeutic agents.
[0172] As used herein, the phrase "pharmaco-mimetic variant interaction" refers to the correlation of multiple biomarker stratification factors, when combined, to show a similar response or lack of response to a therapeutic agent or class of therapeutic agents.
[0173] The term "pharmaco-mimetic means" as used herein refers to a biomarker stratification factor that modulates the function or expression of a drug target gene such that the effect of the biomarker stratification factor on the human phenotype is likely to be predictive of the effect of the drug on the human phenotype.
[0174] As used herein, a "pharmaco-mimetic gene score" is a score that indicates the likelihood that a patient will respond to a particular drug or class of drugs. In some embodiments, more than one pharmacomimetic variant (defined as a functional variant of a gene that is strongly or significantly associated with a given disease) may be identified for any given gene. Instead of analyzing the variants of the gene separately, their effects are combined into a "pharmaco-mimetic gene score."
[0175] As used herein, the terms "polygenic score" or "polygenic risk score" or "PGS" or "PRS" refer to an index that summarizes the estimated effect of many genetic variants on an individual's phenotype, typically calculated as a weighted sum of trait-associated alleles. In some embodiments, a polygenic score reflects an individual's estimated genetic predisposition for a given trait and can be used as a predictor for that trait. In some embodiments, a polygenic score represents an estimate of how likely an individual is to have a given trait based solely on genetics, without taking into account environmental factors.
[0176] As used herein, the term "proteomic score" refers to an index that summarizes the estimated effect of numerous differences in protein composition (e.g., protein abundance, protein mutational variants, or differences in protein post-translational modifications) on an individual's phenotype, typically calculated as a weighted sum of trait-related differences. In some embodiments, the proteomic score can be used as a predictor of any trait. In some embodiments, the proteomic score is used to identify patients likely to derive exceptional benefit from a particular treatment. In some embodiments, the protein composition of a subject is determined by mass spectrometry. In some embodiments, the "proteomic score" is estimated from a regression model, such as a least absolute shrinkage and selection operator (LASSO) (which implements variable selection and regularization for optimal predictive accuracy), of protein measurements versus outcomes, followed by cross-validation techniques.
[0177] As used herein, the term "transcriptomics score" refers to an index that summarizes the estimated effect of numerous differences in RNA transcript composition on an individual's phenotype, typically calculated as a weighted sum of trait-associated differences. In some embodiments, a transcriptomics score refers to a linear or nonlinear combination of transcript abundance values associated with a disease or clinical phenotype. In some embodiments, a transcriptomics score can be used as a predictor of any trait. In some embodiments, a subject's RNA transcript composition is determined by RNA sequencing or microarray.
[0178] As used herein, the term "somatic mutation risk score" refers to an index that summarizes the estimated effect of somatic mutations on an individual's phenotype, typically calculated as a weighted sum of trait-associated somatic mutations. In some embodiments, the somatic mutation risk score can be used as a predictor of any trait. In some embodiments, the somatic mutation composition of a subject is determined by next-generation (high-throughput) sequencing.
[0179] As used herein, the term "principal component analysis" (PCA) refers to a technique for analyzing large datasets containing multiple observational dimensions / features, enhancing data interpretability while retaining the maximum amount of information and enabling visualization of multidimensional data. In some embodiments, PCA refers to a statistical technique that reduces the dimensionality of a dataset. In some embodiments, dimensionality reduction is achieved by linearly transforming the data into a new coordinate system that describes the variability of the data in fewer dimensions than the original data. As used herein, "principal component analysis" refers to a set of new variables that are combinations of the original variables in the original, large dataset.
[0180] As used herein, "statistical interaction" refers to the coefficient of a predictor defined as the product of a pharmacomimetic gene score and a biomarker score in a regression model for predicting a phenotype related to a drug indication. Specifically, a PGS and a biomarker are considered to have a "statistical interaction" if the p-value of the hypothesis test that the coefficient of this PGS-biomarker product term is non-zero is statistically significant (p < 0.05, or, if multiple hypotheses are tested, p < 0.05 / number of tests performed). If the sign of the coefficient of the PGS-biomarker product term is the same as the sign of the marginal biomarker term, this suggests that the predictive and / or causal effect of the biomarker on the phenotype is amplified in subjects with a high PGS.
[0181] As used herein, "association analysis" refers to the process of searching for hidden associations or patterns in large data sets. In some embodiments, association analysis is performed using a statistical model, such as linear regression, logistic regression, or Cox proportional hazards model, in which the dependent variable is a clinical trait or disease phenotype assessed at a single time point or with censored survival time to disease onset or event, and the independent variables include the weighted effects of biomarker stratification factors, demographic information, and other clinical characteristics.
[0182] As used herein, "intermediate phenotype" refers to the demonstration of partial or incomplete correlation between two or more genes (e.g., one of two alleles does not completely determine the phenotype, but an intermediate phenotype may appear when both are present).
[0183] As used herein, the term "HSD17B13" refers to the "hydroxysteroid 17-beta dehydrogenase 13" enzyme (UniProtKB / Swiss-Prot: Q7Z5P4). A pharmacomimetic gene score "associated with response to drugs targeting HSD17B13" is used to determine whether a subject will respond to an HSD17B13 inhibitor (e.g., BI-3231, described in Thamm, Sven, et al., Journal of Medicinal Chemistry 66.4 (2023): 2832-2850, which is incorporated herein by reference in its entirety).
[0184] As used herein, "LPL" refers to the lipoprotein lipase enzyme (UniProtKB / Swiss-Prot: P06858), "ANGPTL3" refers to the angiopoietin-like 3 protein (UniProtKB / Swiss-Prot: Q9Y5C1), and "ANGPTL4" refers to the angiopoietin-like 4 protein (UniProtKB / Swiss-Prot: Q9BY76). Subjects with a pharmacomimetic gene score associated with response to drugs targeting LPL / ANGPTL4 / ANGPTL3 are expected to respond to an LPL agonist, an ANGPTL4 inhibitor, and / or an ANGPTL3 inhibitor.
[0185] As used herein, "C3" refers to the complement C3 protein (UniProtKB / Swiss-Prot: P01024). The pharmacomimetic gene score "Associated with Response to Drugs Targeting C3" is used to determine whether a subject will respond to a C3 inhibitor.
[0186] As used herein, "CFB" refers to the complement factor B protein (UniProtKB / Swiss-Prot: P00751). A pharmacomimetic gene score "associated with response to drugs targeting CFB" is used to determine whether a subject will respond to a CFB inhibitor.
[0187] As used herein, "CFH" refers to the complement factor H protein (UniProtKB / Swiss-Prot: P08603). The pharmacomimetic gene score "Associated with Response to Drugs Targeting CFH" is used to determine whether a subject will respond to a CFH inhibitor.
[0188] As used herein, "HTRA1" refers to the HtrA serine peptidase 1 protein (UniProtKB / Swiss-Prot: Q92743). A pharmacomimetic gene score "associated with response to drugs targeting HTRA1" is used to determine whether a subject will respond to an HTRA1 inhibitor.
[0189] As used herein, "TYK2" refers to the "tyrosine kinase 2" enzyme (UniProtKB / Swiss-Prot: P29597). A pharmacomimetic gene score "associated with response to drugs targeting TYK2" is used to determine whether a subject will respond to a TYK2 inhibitor.
[0190] As used herein, "IL33" refers to "interleukin 33" (UniProtKB / Swiss-Prot: O95760). As used herein, "TSLP" refers to "thymic stromal lymphopoietin" (UniProtKB / Swiss-Prot: Q969D9). As used herein, "IL4R" refers to "interleukin 4 receptor" (UniProtKB / Swiss-Prot: P24394). A pharmacomimetic gene score "associated with response to drugs targeting IL33 / TSLP / IL4R" is used to determine whether a subject will respond to an IL33 / TSLP / IL4R inhibitor, such as a bispecific antibody targeting IL33 / TSLP / IL4R. As used herein, "GLP1R" refers to "glucagon-like peptide 1 receptor" (UniProtKB / Swiss-Prot: P43220). As used herein, "GIPR" refers to "gastric inhibitory polypeptide receptor" (UniProtKB / Swiss-Prot: P48546). A pharmacomimetic gene score "associated with response to drugs that target GLP1R / GIPR" is used to determine whether a subject will respond to a GLP1R and / or GIPR agonist. The presently disclosed system for risk score-based drug development
[0191] The disclosed system has many applications, ranging from drug development using agnostic analysis of new chemical entities combined with the distribution of genetic variants in patient populations, to novel and efficient designs of clinical trials, to the introduction of novel endpoints in clinical trials, to how data from previous clinical trials are analyzed to provide evidence of efficacy. Particularly valuable applications include combining genetic information from patient populations to further understand pharmacodynamic endpoints and biomarkers and how they relate to clinical outcomes.
[0192] For example, the disclosed system can be used for in silico replication and regulatory evaluation of clinical trials. Historically, safety and efficacy data provided to regulatory agencies to support commercialization has required preclinical and clinical efficacy data using empirical methods. treatment area
[0193] The disclosed system can be used for drug development in a variety of therapeutic areas, including, but not limited to, autoimmune diseases; cardiovascular diseases such as CAD; dental and oral health; dermatology; endocrinology; gastroenterology; ultra-rare diseases; hematology; hepatology; immunology; infectious diseases; metabolic disorders; musculoskeletal disorders; nephrology; neurology, including neurodegenerative diseases; obstetrics and gynecology; oncology; ophthalmology; orthopedics; otolaryngology; psychiatric disorders; pulmonary / respiratory diseases; rheumatology; and urology.
[0194] Therapeutic indications within specific therapeutic areas that can be addressed using the drug development system of the present disclosure include, but are not limited to, the following listed in Table 1:
[0195] [Table 1] JPEG2025540934000003.jpg247170JPEG2025540934000004.jpg34170
[0196] The present approach may be configured as a system or a method, or may be provided as a computer program product, which may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the present disclosure.
[0197] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, or semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanically encoded device having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage medium should not be construed as a transitory signal itself, such as an electric wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or an electrical signal transmitted over a wire. Rather, the computer-readable storage medium is a non-transitory (ie, non-volatile) medium.
[0198] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device over a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface of each computing / processing device receives the computer-readable program instructions from the network and transmits the computer-readable program instructions to a computer-readable storage medium within the respective computing / processing device for storage.
[0199] The computer-readable program instructions for carrying out the operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, etc., and conventional procedural programming languages, e.g., the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the present disclosure.
[0200] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, executed by the processor of the computer or other programmable data processing apparatus, produce means for implementing the functions and / or steps specified in the disclosure. These computer-readable program instructions may also be stored in a computer-readable storage medium capable of directing a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that a computer-readable storage medium having instructions stored therein may comprise, alternatively consist essentially of, or further consist of an article of manufacture including instructions that implement aspects of the functions and / or steps specified in the present disclosure.
[0201] Computer-readable program instructions may be loaded into a computer, other programmable data processing apparatus, or other device and caused to execute a series of operational steps on the computer, programmable apparatus, or other device to produce a computer-implemented process such that the instructions executing on the computer, other programmable apparatus, or other device implement the functions and / or steps set forth in this disclosure. Methods for determining drug activity
[0202] One aspect of the present disclosure relates to an in silico method for determining drug activity of multiple drug targets, the method comprising, alternatively consisting essentially of, or further consisting of: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the data including, alternatively consisting essentially of, or further consisting of, a plurality of biomarker stratification factors and at least one disease phenotype; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; and identifying a predicted drug activity of each drug target for at least one disease phenotype in a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with the disease phenotype.
[0203] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0204] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0205] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0206] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0207] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is selected from the group consisting of: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; and (viii) a biological pathway activity. (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof.
[0208] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
[0209] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis or weighted principal component analysis based on a matrix of one or more biomarker stratification factor effects.
[0210] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of constructing a matrix based on the one or more biomarker stratification factors, the phenotype, and the plurality of values.
[0211] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
[0212] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratification factor scores for the selected disease phenotype.
[0213] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, weighted principal component analysis identifying predicted drug activity of multiple drug targets against at least one disease phenotype.
[0214] Non-alcoholic steatohepatitis (NASH): Another aspect of the present disclosure is to determine drug activity of multiple drug targets associated with non-alcoholic steatohepatitis (NASH). For in silico methods, the method may comprise, alternatively consist essentially of, or further consist of the steps of: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the data including, alternatively consisting essentially of, or further consisting of, a plurality of biomarker stratification factors and at least one disease phenotype; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; and identifying a predicted drug activity of each drug target for at least one disease phenotype in a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with the disease phenotype.
[0215] In some embodiments, the NASH disease phenotype includes, alternatively consists essentially of, or additionally consists of one or more of quantitative disease traits (e.g., blood markers), medication information, demographic information (e.g., age at diagnosis, biological sex), itchy skin, abdominal swelling (also called ascites), shortness of breath, swelling of the legs, spider veins under the skin, enlarged spleen, redness of the palms of the hands, or yellowing of the skin and eyes (jaundice).
[0216] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0217] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0218] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0219] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0220] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is selected from the group consisting of: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; and (viii) a biological pathway activity. (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof.
[0221] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
[0222] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis or weighted principal component analysis on the one or more biomarker stratification factor effect matrices.
[0223] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of constructing a matrix based on the one or more biomarker stratification factors, the phenotype, and the plurality of values.
[0224] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
[0225] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratification factor scores for the selected disease phenotype.
[0226] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, weighted principal component analysis identifying predicted drug activity of multiple drug targets against at least one disease phenotype.
[0227] In some embodiments, the biomarker stratification factor score is a fatty liver polygenic score.
[0228] In some embodiments, the multiple drug targets are associated with non-alcoholic steatohepatitis (NASH), and the multiple drug targets include, alternatively consist essentially of, or further consist of HSD17B13.
[0229] In some embodiments, the pharmacomimetic gene score is associated with response to drugs that target HSD17B13.
[0230] In some embodiments, the drug that targets HSD17B13 is an HSD17B13 inhibitor.
[0231] Coronary Heart Disease: Another aspect of the present disclosure is to determine drug activity of multiple drug targets associated with coronary heart disease. For in silico methods, the method may comprise, alternatively consist essentially of, or further consist of the steps of: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the data including, alternatively consisting essentially of, or further consisting of, a plurality of biomarker stratification factors and at least one disease phenotype; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; and identifying a predicted drug activity of each drug target for at least one disease phenotype in a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with the disease phenotype.
[0232] In some embodiments, the disease phenotype includes, alternatively consists essentially of, or further consists of one or more of quantitative disease traits (e.g., LDL levels, HDL levels, triglyceride levels), medication information, demographic information (e.g., age at diagnosis, biological sex), chest pain (angina), shortness of breath, fatigue, heart attack, nausea, sweating, tiredness, tachycardia, weakness, or dizziness.
[0233] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0234] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0235] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0236] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0237] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; (viii) a polygenic score for biological pathway activity. (xiii) somatic mutation score for drug target expression; (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof.
[0238] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
[0239] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis or weighted principal component analysis based on a matrix of one or more biomarker stratification factor effects.
[0240] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of constructing a matrix based on the one or more biomarker stratification factors, the phenotype, and the plurality of values.
[0241] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
[0242] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratification factor scores of the selected disease phenotype.
[0243] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, weighted principal component analysis identifying predicted drug activity of multiple drug targets against at least one disease phenotype.
[0244] In some embodiments, the multiple drug targets are associated with coronary heart disease and the multiple drug targets include, or alternatively consist essentially of, or further consist of, LPL, ANGPTL4 and / or ANGPTL3.
[0245] In some embodiments, the pharmacomimetic gene score is associated with response to a drug that targets LPL, ANGPTL4, and / or ANGPTL3. In some embodiments, the drug is an LPL, ANGPTL4, and / or ANGPTL3 inhibitor.
[0246] Dry Age-Related Macular Degeneration: Another aspect of the present disclosure is to determine drug activity of multiple drug targets associated with dry age-related macular degeneration. For in silico methods, the method may comprise, alternatively consist essentially of, or further consist of the steps of: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the data including, alternatively consisting essentially of, or further consisting of, a plurality of biomarker stratification factors and at least one disease phenotype; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; and identifying a predicted drug activity of each drug target for at least one disease phenotype in a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with the disease phenotype.
[0247] In some embodiments, the disease phenotype includes, alternatively consists essentially of, or additionally consists of one or more of: blurred or blind spots in central vision in one or both eyes, difficulty recognizing faces or reading letters, reduced light intensity or brightness of colors, distorted vision such as seeing curves on straight objects, or increased difficulty adapting to low light levels.
[0248] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0249] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0250] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0251] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0252] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; (viii) a polygenic score for biological pathway activity. (xiii) somatic mutation score for drug target expression; (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof.
[0253] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
[0254] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis or weighted principal component analysis based on a matrix of one or more biomarker stratification factor effects.
[0255] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of constructing a matrix based on the one or more biomarker stratification factors, the phenotype, and the plurality of values.
[0256] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
[0257] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratification factor scores for the selected disease phenotype.
[0258] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, weighted principal component analysis identifying predicted drug activity of multiple drug targets against at least one disease phenotype.
[0259] In some embodiments, the multiple drug targets are associated with dry age-related macular degeneration and the multiple drug targets include, or alternatively consist essentially of, or further consist of C3, CFB, CFH, and / or HTRA1.
[0260] In some embodiments, the pharmacomimetic gene score is associated with response to a drug that targets C3, CFB, CFH, and / or HTRA1. In some embodiments, the drug is a C3, CFB, CFH, and / or HTRA1 inhibitor.
[0261] Rheumatoid arthritis: Another aspect of the present disclosure is to determine drug activity of multiple drug targets associated with rheumatoid arthritis. For in silico methods, the method may comprise, alternatively consist essentially of, or further consist of the steps of: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the data including, or alternatively consisting essentially of, or further consisting of, a plurality of biomarker stratification factors and at least one disease phenotype; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; and identifying a predicted drug activity of each drug target for at least one disease phenotype in a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with the disease phenotype.
[0262] In some embodiments, the disease phenotype comprises, alternatively consists essentially of, or additionally consists of one or more of: swollen joints, fluid accumulation in the heels, morning stiffness, joint pain, fatigue, joint redness, increased eye sensitivity and dryness, dry mouth, skin nodules, or lung inflammation.
[0263] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of genetic variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0264] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0265] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0266] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0267] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; (viii) a polygenic score for biological pathway activity. (xiii) somatic mutation score for drug target expression; (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof.
[0268] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
[0269] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis or weighted principal component analysis based on a matrix of one or more biomarker stratification factor effects.
[0270] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of constructing a matrix based on the one or more biomarker stratification factors, the phenotype, and the plurality of values.
[0271] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
[0272] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratification factor scores of the selected disease phenotype.
[0273] In some embodiments, the multiple drug targets are associated with rheumatoid arthritis and the multiple drug targets include, alternatively consist essentially of, or further consist of TYK2.
[0274] In some embodiments, the pharmacomimetic gene score is associated with response to a drug that targets TYK2. In some embodiments, the drug is a TYK2 inhibitor.
[0275] Atopic Disease: Another aspect of the present disclosure is to determine drug activity of multiple drug targets associated with atopic disease. For in silico methods, the method may comprise, alternatively consist essentially of, or further consist of the steps of: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the data including, alternatively consisting essentially of, or further consisting of, a plurality of biomarker stratification factors and at least one disease phenotype; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; and identifying a predicted drug activity of each drug target for at least one disease phenotype in a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with the disease phenotype.
[0276] In some embodiments, the disease phenotype comprises, alternatively consists essentially of, or additionally consists of one or more of: dry skin, itching, swelling and inflammation, a red, brown, purple, or gray rash, small fluid-filled bumps or crusts, cracked skin, a rash, wrinkles on the palms or under the eyes, or darkening of the skin around the eyes.
[0277] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0278] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0279] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0280] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0281] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; (viii) a polygenic score for biological pathway activity. (xiii) somatic mutation score for drug target expression; (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof.
[0282] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
[0283] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis or weighted principal component analysis based on a matrix of one or more biomarker stratification factor effects.
[0284] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of constructing a matrix based on the one or more biomarker stratification factors, the phenotype, and the plurality of values.
[0285] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
[0286] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratification factor scores of the selected disease phenotype.
[0287] In some embodiments, the multiple drug targets are associated with atopic dermatitis and the multiple drug targets include, or alternatively consist essentially of, or further consist of, IL33, TSLP and / or IL4R.
[0288] In some embodiments, the pharmacomimetic gene score is associated with response to drugs that target IL33, TSLP, and / or IL4R. In some embodiments, the drugs are IL33, TSLP, and / or IL4R inhibitors.
[0289] Obesity: Another aspect of the present disclosure is to determine drug activity of multiple drug targets associated with obesity. For in silico methods, the method may comprise, alternatively consist essentially of, or further consist of the steps of: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the data including, alternatively consisting essentially of, or further consisting of, a plurality of biomarker stratification factors and at least one disease phenotype; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; and identifying a predicted drug activity of each drug target for at least one disease phenotype in a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with the disease phenotype.
[0290] In some embodiments, the disease phenotype comprises, alternatively consists essentially of, or additionally consists of one or more of: a body mass index (BMI) greater than 30, sleep apnea, varicose veins, skin problems caused by fluid accumulation in skin folds, gallstones, or osteoarthritis.
[0291] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0292] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0293] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0294] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0295] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; (viii) a polygenic score for biological pathway activity. (xiii) somatic mutation score for drug target expression; (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof.
[0296] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
[0297] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis or weighted principal component analysis based on a matrix of one or more biomarker stratification factor effects.
[0298] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of constructing a matrix based on the one or more biomarker stratification factors, the phenotype, and the plurality of values.
[0299] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
[0300] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratification factor scores of the selected disease phenotype.
[0301] In some embodiments, the multiple drug targets are associated with obesity and the multiple drug targets include, or alternatively consist essentially of, or further consist of GLP1R and / or GIPR.
[0302] In some embodiments, the pharmacomimetic gene score is associated with response to drugs that target GLP1R and / or GIPR. In some embodiments, the drugs are GLP1R and / or GIPR agonists.
[0303] In some embodiments, the polygenic score is a body mass index polygenic score.
[0304] Another aspect of the present disclosure relates to an in silico method for determining drug activity of multiple drug targets, the method comprising: obtaining molecular biomarker stratification factor data and at least one disease phenotype from each of a plurality of subjects; determining the value of the biomarker stratification factor effect for each disease phenotype based on external data; constructing a matrix based on the biomarker stratification factors, disease phenotypes and values; performing a principal component analysis on the biomarker stratification factor x phenotype-associated phenotype biomarker stratification factor effect matrix; calculating a biomarker stratification factor score for each principal component; calculating a pharmacomimetic gene score for each drug target; identifying predicted drug activity of each drug target against one or more phenotypes in the subset of biomarker stratification factor distributions based on statistical interactions between the polygenic scores calculated for each principal component and the pharmacomimetic gene scores for each drug target in association analysis with outcome phenotypes; Includes.
[0305] Another aspect of the present disclosure relates to an in silico method for determining drug activity of a drug target against a plurality of intermediate phenotypes, the method comprising: obtaining molecular biomarker stratification factor data and a disease phenotype from each of a plurality of subjects; determining a value of the biomarker stratification factor effect for each intermediate phenotype based on external data; calculating a biomarker stratification factor score for each intermediate phenotype in the disease process; calculating a pharmacomimetic gene score for each drug target for each intermediate phenotype; identifying predicted drug activity of each drug target for one or more phenotypes in the subset of biomarker stratification factor distributions based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with intermediate and outcome phenotypes; Includes.
[0306] Another aspect of the present disclosure relates to a method for determining drug activity of a drug target, the method comprising: obtaining molecular biomarker stratification factor data and a disease phenotype from each of a plurality of subjects; determining the value of the biomarker stratification factor effect for each phenotype based on external data; calculating a biomarker stratification factor score for a selected outcome phenotype related to a disease outcome; calculating a pharmacomimetic gene score for the drug target; identifying a predicted drug activity of the drug target for one or more phenotypes of a subset of the biomarker stratification factor distribution based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for the drug target in an association analysis with the outcome phenotype; Includes.
[0307] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0308] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0309] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0310] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0311] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; (viii) a polygenic score for biological pathway activity. (xiii) somatic mutation score for drug target expression; (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof.
[0312] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
[0313] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, performing principal component analysis or weighted principal component analysis based on a matrix of one or more biomarker stratification factor effects.
[0314] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of constructing a matrix based on the one or more biomarker stratification factors, the phenotype, and the plurality of values.
[0315] In some embodiments, the method further comprises, or alternatively consists essentially of, or further consists of, adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
[0316] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratification factor scores of the selected disease phenotype.
[0317] In some embodiments, the method further comprises, alternatively consists of, or further consists of, the weighted principal component analysis identifying predicted drug activities of multiple drug targets against at least one disease phenotype.
[0318] Another aspect of the present disclosure is a method for manufacturing a semiconductor device comprising: generating a plurality of principal components (PCs) corresponding to the genetic ancestry data of the subjects in the study cohort; generating a biomarker stratification factor score for each subject in the study cohort based on at least (i) the PC and (ii) the weight of the biomarker stratification factor for the disease of interest; determining which of a plurality of disease-associated variants is a pharmacomimetic for a disease of interest, wherein a disease-associated variant is a pharmacomimetic if it modulates the function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to predict a second effect of the drug on the phenotype; determining a statistical interaction between the pharmacomimetic measures and biomarker stratification factor scores for one or more drug targets, said statistical interaction predicting drug target-specific differential treatment response for the disease of interest; taking one or more prediction-based actions based on the determined interactions; The present invention relates to an in silico method, including:
[0319] In some embodiments, the one or more prediction-based actions include at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics.
[0320] In some embodiments, the genetic ancestry data indicates whether subjects in a study cohort are cases or controls for a disease of interest.
[0321] In some embodiments, the genetic ancestry data is based on genotyping arrays or whole genome sequencing.
[0322] In some embodiments, the genetic ancestry data is obtained from publicly available databases.
[0323] In some embodiments, the plurality of PCs comprises at least 5 PCs.
[0324] In some embodiments, the study cohort includes at least 200 control subjects.
[0325] In some embodiments, each disease-associated variant in the plurality of disease-associated variants meets a genome-wide significance threshold.
[0326] In some embodiments, the plurality of disease-associated variants is determined based on a genome-wide association study (GWAS) of the first disease.
[0327] In some embodiments, the biomarker stratification factor score for each subject is a scaled biomarker stratification factor score.
[0328] In some embodiments, the scaled biomarker stratification factor score is based on at least the raw PGS and the ancestry-normalized PGS.
[0329] In some embodiments, the PGS variant weights are calculated independently from genetic ancestry data corresponding to a study cohort.
[0330] In some embodiments, determining which of the plurality of disease-associated variants are pharmacomimetics for the disease of interest comprises generating an allele score.
[0331] In some embodiments, the variants in the allele score are weighted based at least on a second GWAS that did not include any of the subjects in the study cohort.
[0332] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0333] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0334] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0335] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0336] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; (viii) a polygenic score for biological pathway activity. (xiii) somatic mutation score for drug target expression; (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof. system
[0337] Another aspect of the present disclosure relates to an in silico system for drug development, the system including at least one hardware processor and a non-transitory computer-readable storage medium having program code stored thereon, the program code comprising: obtaining molecular biomarker stratification factor data and a disease phenotype from each of a plurality of subjects; Determine the value of the biomarker stratification factor effect for each phenotype based on external data; Construct a matrix based on biomarker stratification factors, phenotypes and values; Conduct a principal component analysis on the biomarker stratification factor x phenotype-related phenotype biomarker stratification factor effect matrix: Calculate the biomarker stratification factor score for each principal component; Calculate a pharmacomimetic gene score for each drug target; Identifying predicted drug effects for one or more phenotypes for a subset of the biomarker stratification factor distribution based on statistical interactions between the polygenic scores calculated for each principal component and the pharmacomimetic gene scores for each drug target in association analyses with outcome phenotypes; and wherein the at least one processor is operable to execute the at least one program.
[0338] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or further consist of genetic variants, wherein the biomarker stratification factor score comprises, alternatively consist essentially of, or further consist of one or more scores selected from: (i) a polygenic score of a disease phenotype; (ii) a polygenic score of disease risk factors; (iii) a polygenic score of biological pathway activity; (iv) a polygenic score of drug target expression, or any linear or non-linear combination thereof.
[0339] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of proteome variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a proteome score of a disease phenotype; (ii) a proteome score of a disease risk factor; (iii) a proteome score of a biological pathway activity; (iv) a proteome score of a drug target expression; or any linear or non-linear combination thereof.
[0340] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of, transcriptional variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a transcriptome score of a disease phenotype; (ii) a transcriptome score of a disease risk factor; (iii) a transcriptome score of a biological pathway activity; or any linear or non-linear combination thereof.
[0341] In some embodiments, the biomarker stratification factor comprises, or alternatively consists essentially of, or further consists of somatic mutation variants, wherein the biomarker stratification factor score comprises, or alternatively consists essentially of, or further consists of one or more scores selected from: (i) a somatic mutation score of drug target expression; (ii) a somatic mutation score of disease phenotype; (iii) a somatic mutation score of disease risk factors; (iv) a somatic mutation score of biological pathway activity; (v) a somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
[0342] In some embodiments, the biomarker stratification factors comprise, alternatively consist essentially of, or additionally consist of genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, wherein the biomarker stratification factor score is: (i) a polygenic score for disease phenotype; (ii) a polygenic score for disease risk factors; (iii) a polygenic score for biological pathway activity; (iv) a polygenic score for drug target expression; (v) a proteomic score for disease phenotype; (vi) a proteomic score for disease risk factors; (vii) a proteomic score for biological pathway activity; (viii) a polygenic score for biological pathway activity. (xiii) somatic mutation score for drug target expression; (xiv) somatic mutation score for disease phenotype; (xv) somatic mutation score for disease risk factor; (xvi) somatic mutation score for biological pathway activity; (xvii) somatic mutation score for drug target expression; or any linear or non-linear combination thereof. example
[0343] The following examples are included for illustrative purposes only and are not intended to limit the scope of the present disclosure. Example 1: Regional CAD PRS
[0344] In clinical trials of PSCK9 inhibitors for secondary prevention of coronary heart disease (CAD), patients with a high PRS for CAD had a substantially greater relative risk reduction than patients with a low PRS, The reduction in LDL cholesterol between the two groups was similar (Marston et al. 2019, Damask et al. 2019). Similarly, a high PRS amplifies the relative risk reduction provided by statins for the primary prevention of CAD (Natarajan 2017).
[0345] The disclosed in silico modeling system was able to reproduce the clinical trial interaction between pharmacological PCSK9 inhibition and PRS observed in clinical trials. The clinical data were replicated in silico using UK Biobank genetic variants (Kheraj AV et al., Nat Genet. 2018; 50:1219-24). Cohort information regarding these subjects and the prevalence of CAD is summarized in Table 2.
[0346] Table 2: Summarizing cohort statistics of patients in the UK Biobank study (ukbiobank.ac.uk).
[0347] [Table 2]
[0348] These "pharmaco-mimetic" variants discovered in silico mimic the effects of known drugs by reducing PCSK9 function. These pharmacomimetic variants in LPL, the target of triglyceride-lowering drugs, interact with PRS in silico to determine CAD risk. rsll591147 (R46L), weight = -0.48 rsll206510 (intergenic), weight = -0.07 rs2479409 (intergenic), weight = -0.047 rs505151 (G670Q), weight = -0.09
[0349] We also calculated two-gene GRSs for six pharmacomimetic variants of LPL and ANGPTL4, a protein that physically binds and inhibits LPL. Variants were weighted by their effect on serum triglycerides in the same study as above (Id.). rsl3702 (LPL 3' UTR), weight = -0.12 rs328 (LPL S474X), weight = -0.18 rs268 (LPL N318S), weight = 0.24 rsl801177 (LPL D36N), weight = 0.17 rsll6843064 (ANGPTL4 E40K), weight = -0.27 rs7255436 (ANGPTL4 intron), weight = -0.019
[0350] Finally, we calculated two-gene GRS from two pharmacomimetic variants of the incretin hormone receptors GLP1R and GIPR. rsl042044 (GLP1R L260F), weight = 0.00895 rsl800437 (GIPR E354Q), weight = -0.0283
[0351] GLP1R and GIPR variants were weighted by their effect in a manually performed meta-analysis of four BMI GWAS (for only two variants), none of which overlapped with the others, nor with the others or with UK BioBank. · Locke et al. 2015 (GIANT consortium) (Locke, Adam E., et al. "Genetic studies of body mass index yield new insights for obesity biology." Nature 518.7538 (2015): 197-206.) · Sakae et al. 2021 (Biobank Japan) (Sakaue, S., Kanai, M., Tanigawa, Y. et al. A cross-population atlas of genetic associations for 220 human phenotypes. Nat Genet 53, 1415-1424 (2021)) · Fernandez-Rhodes et al. 2022 (Hispanic Community Health Study / Study of Latinos) (Fernandez-Rhodes, Lindsay, et al. "Ancestral diversity improves discovery and fine-mapping of genetic loci for anthropometric traits-The Hispanic / Latino Anthropometry Consortium." Human Genetics and Genomics Advances 3.2 (2022)) · Wojcik et al. 2019 (PAGE study) (Wojcik, Genevieve L., et al. "Genetic analyzes of diverse populations improves discovery for complex traits." Nature 570.7762 (2019): 514-518.) Example 2: Genome-wide CAD PRS
[0352] The PRS for CAD was developed by applying the program PRS-CS (Ge et al. 2019) to an inverse-variance weighted meta-analysis of two GWAS: · Nikpay et al. 2015 = CARDioGRAMplusC4D ("A comprehensive 1000 Genomes-based genome-wide association meta-analysis of coronary artery disease." Nature genetics 47, no. 10 (2015): 1121-1130.) · Kukri = Finnegan release 6, using the phenotype I9_CHD (Kurki, Mitja I., et al. "FinnGen provides genetic insights from a well-phenotyped isolated population." Nature 613.7944 (2023): 508-518.) These GWAS do not have any overlap with UK Biobank.
[0353] The PRS was selected from a set of PRSs trained on the PRS-CS using various hyperparameters. The hyperparameter values that provided the best predictive performance in the UK Biobank, the same cohort in which the pharmacomimetic variant × PRS interaction test was performed, were selected. However, confounding effects were expected to be minimal. This is because the only hyperparameter in the PRS-CS, phi, represents the polygenicity of the disease under study, and the optimal value of phi should not depend on the LD structure, which is the main driver of the differential performance of the PRS across cohorts. Example 3: Pathway-specific CAD PRS
[0354] A list of 264 conditionally independent (p < 5e-7) variants associated with CAD was obtained from van der Harts and Verweij 2017 (Van Der Harst, Pim, and Niek Verweij. "Identification of 64 novel genetic loci provides an expanded view on the genetic architecture of coronary artery disease." Circulation research 122.3 (2018): 433-443). A matrix was constructed in which rows are CAD variants, columns are CAD-related phenotypes (including CAD itself), and values are the effect of the variant on each phenotype, sourced from multiple GWAS studies listed below.
[0355] To select CAD-associated phenotypes, the genetic correlation (R2) between CAD and a number of candidate phenotypes was calculated using 264 CAD variants. All UK Biobank biomarkers and blood counts, as well as various other CAD-associated phenotypes, were included in the matrix of nine phenotypes that had a genetic correlation to CAD of R2 >= 0.25. · CAD itself, R A2 = 1, (van der Harts P and Verweij N. Circ Res. 2018 Feb 2;122(3):433-443) · Abdominal aortic aneurysm, RA 2 = 0.49, (Zhou Z, et al., Medicine (Baltimore). 2021 Oct 1;100(39): e27306.) Lipoprotein A, RA2 = 0.46, UK Biobank Non-HDL cholesterol, RA2 = 0.31, Graham SE et al. Nature. 2021 Dec;600(7890):675-679. Systolic blood pressure, RA 2 = 0.31, UI < Biobank Total cholesterol, RA 2 = 0.30, Graham et al. LDL cholesterol, RA 2 = 0.29, Graham et al. · Pulse pressure, RA 2 = 0.28, UK Biobank Ischemic stroke, RA2 = 0.25, Malik et al. 2018
[0356] Phenotypes that had an RA 2 < 0.25 and were therefore excluded from the matrix include: HDL cholesterol Triglycerides Diastolic blood pressure Mean arterial pressure Body Mass Index · type 2 diabetes HbA1c C-reactive protein
[0357] Applicant performed principal component analysis (PCA) on a CAD variant x CAD-related phenotype variant effect matrix (e.g., a matrix in which rows represent the CAD risk phenotypes listed above [abdominal aortic aneurysm, lipoprotein A, non-HDL cholesterol, systolic blood pressure, total cholesterol, LDL cholesterol, pulse pressure, and ischemic stroke] and columns represent the genetic association weights for each gene variant position). The estimated percentage of each variant's effect on CAD mediated by principal component 1 was defined as (variant loading on PC1) / (sum of squared variant loadings on all PCs). This was repeated for principal components 2-9.
[0358] Based on the trait associations of each PC and the candidate genes of the most strongly associated variants for each PC, PC1-3 were inferred to represent specific and distinct CAD-related pathways according to manual inspection of the biological functions of the major genetic loadings of each principal component.
[0359] Next, pathway-specific PRSs were constructed for each of PC1-3. Each variant in the original genome-wide PRS for CAD was assigned to the closest variant included in the PCA. PCA variants whose unscaled PC loadings were inconsistent with their effect on CAD were excluded. For example, if variant X had a loading of 0.5 on PC1 and PC1 was positively correlated with CAD, then variant X would be predicted and required to be associated with a high risk of CAD. Variants located more than 500 kb away from any other PCA variant were also excluded.
[0360] Next, the weight of each variant in the PRS, originally calculated by PRS-CS, was multiplied by the scaled PC loading of the PCA variant closest to the PC of interest. In other words, loci were reweighted based on whether their effect was predicted to be mediated or not by the pathway represented by a given PC. Finally, the PRS was recalculated using these new weights. This resulted in three new pathway-specific PRSs for each of PC1-3. Phenotype
[0361] CAD cases were defined as subjects with at least one of the following codes: ICD-10: 120-125; ICD-9: 410, 412, 414; UK Biobank self-report: 1074, 1075; OPCS-4: K40-K46, K49-K50, K75; or any record in the Myocardial Infarction Registry (UKB data field 42000). dosage
[0362] Subjects who self-reported taking lipid-lowering medications at the initial UKB assessment visit were excluded from all analyses. Because lipid-lowering medications have been shown to reduce polygenic risk, given that the pharmacomimetic variants used explained a smaller proportion of the variability in lipid traits in the population compared with medication, medication-PRS interactions may mask potential pharmacomimetic variant-PRS interactions.
[0363] Using PCA on variant effects on CAD and CAD-related phenotypes, we were able to construct pathway-specific PRSs. Two of these PRSs (PC1 and PC2) appeared to have no interactions with LPL pharmacomimetic variants. Adjusting the genome-wide CAD PRS to remove the contributions of PC1 and PC2 resulted in a small but significant (p < 0.01) improvement in the magnitude of the LPL pharmacomimetic variant × PRS interaction, but also increased the sparsity of the PRS. PCSK9 GRS x CAD PRS
[0364] PCSK9 GRS had a nominally significant (p < 0.05) interaction with CAD PRS, consistent with the results of clinical trials: the mean effect of PCSK9 gene loss of function in subjects in the top tertile of PRS was nearly three times (2.9x) the mean effect in subjects in the bottom or middle tertiles of PRS.
[0365] The CAD PRS was adjusted to eliminate the contribution of PCSK9 to the PRS CAD (Table 3).
[0366] Table 3 summarizes the interaction between PCSK9 GRS and CAD PRS in determining CAD risk.
[0367] [Table 3]
[0368] Subjects receiving lipid-lowering medications were excluded from the analysis. The PCSK9 GRS is scaled so that 1 unit corresponds to a lifetime LDL reduction equivalent to that of evolocumab in a clinical trial setting, or -1.56 standard deviations. The CAD PRS is in standard deviations.
[0369] The CAD PRS was adjusted to eliminate the contribution of PCSK9 to the PRS (Table 4).
[0370] Table 4 summarizes the effect of PCSK9 GRS on CAD risk stratified by quantiles of CAD PRS.
[0371] [Table 4]
[0372] Subjects receiving lipid-lowering medications were excluded from the analysis. The PCSK9 GRS is scaled so that 1 unit corresponds to a lifetime LDL reduction equivalent to that of evolocumab in a clinical trial setting, or -1.56 standard deviations. The CAD PRS is in standard deviations.
[0373] The CAD PRS was adjusted to remove the contribution of PCSK9 to the PRS (Table 5). Subjects receiving lipid- or blood pressure-lowering medications were excluded from the analysis. The PCSK9 GRS was scaled so that 1 unit corresponds to a lifetime LDL reduction equivalent to that of evolocumab in a clinical trial setting, or -1.56 standard deviations. The units of the CAD PRS are standard deviations.
[0374] Table 5: Summary of the interaction between PCSK9 GRS and CAD PRS in determining CAD risk.
[0375] [Table 5]
[0376] The CAD PRS was adjusted to remove the contribution of PCSK9 to the PRS (Table 6). Subjects receiving lipid- or blood pressure-lowering medications were excluded from the analysis. The PCSK9 GRS was scaled so that 1 unit corresponds to the lifetime LDL reduction equivalent to that of evolocumab in a clinical trial setting, i.e., -1.56. The units of the CAD PRS are standard deviations.
[0377] Table 6 summarizes the effect of PCSK9 GRS on CAD risk stratified by quantiles of CAD PRS.
[0378] [Table 6]
[0379] LPL / ANGPTL4 GRS x CAD PRS
[0380] The LPL / ANGPTL4 GRS also had a significant (p < 0.01) interaction with CAD PRS: the effect of LPL gene activation in subjects in the highest PRS tertile appeared to be approximately twice (2.2×) the average effect in subjects in the lowest and middle PRS tertiles.
[0381] The CAD PRS was adjusted to remove the contributions of LPL and ANGPTL4 to the PRS (Figure 8). Subjects receiving lipid-lowering medications were excluded from the analysis. The LPL / ANGPTL4 GRS was scaled so that one unit corresponds to a lifetime TG reduction equivalent to 0.5 standard deviations = 1.67 times the effect of Vascepa (ethyl eicosapentaenoic acid) in a clinical trial setting. The units of the PRS are standard deviations.
[0382] Table 7: Summary of the interaction between LPL / ANGPTL4 GRS and CAD PRS in determining the risk of CAD.
[0383] [Table 7]
[0384] The CAD PRS was adjusted to remove the contributions of LPL and ANGPTL4 to the PRS (Table 8). Subjects taking lipid-lowering medications were excluded from the analysis. The LPL / ANGPTL4 GRS was scaled so that 1 unit corresponds to a lifetime TG reduction equivalent to 0.5 standard deviations = 1.67 times the effect of Vascepa (ethyl eicosapentaenoic acid) in a clinical trial setting. The units of the PRS are standard deviations.
[0385] Table 8 summarizes the effect of LPL / ANGPTL4 GRS on CAD risk stratified by quantiles of CAD PRS.
[0386] [Table 8]
[0387] The CAD PRS was adjusted to remove the contributions of LPL and ANGPTL4 to the PRS (Table 9). Subjects taking lipid- or blood pressure-lowering medications were excluded from the analysis. The LPL / ANGPTL4 GRS was scaled so that 1 unit corresponds to a lifetime TG reduction equivalent to 0.5 standard deviations = 1.67 times the effect of Vascepa (ethyl eicosapentaenoic acid) in a clinical trial setting. The units of the PRS are standard deviations.
[0388] Table 9: Summary of the interaction between LPL / ANGPTL4 GRS and CAD PRS in determining the risk of CAD.
[0389] [Table 9]
[0390] The CAD PRS was adjusted to remove the contributions of LPL and ANGPTL4 to the PRS (Table 10). Subjects taking lipid- or blood pressure-lowering medications were excluded from the analysis. The LPL / ANGPTL4 GRS was scaled so that 1 unit corresponds to a lifetime TG reduction equivalent to 0.5 standard deviations = 1.67 times the effect of Vascepa (ethyl eicosapentaenoic acid) in a clinical trial setting. The units of the PRS are standard deviations.
[0391] Table 10 summarizes the effect of LPL / ANGPTL4 GRS on CAD risk stratified by quantiles of CAD PRS.
[0392] [Table 10]
[0393] LPL / ANGPTL4 GRS × TG PRS (exclude these loci)
[0394] The LPL / ANGPTL4 GRS does not interact with the triglyceride PRS, suggesting that the GRS:CAD PRS is driven by pathways other than triglyceride metabolism.
[0395] The TG PRS was adjusted to remove the contributions of LPL and ANGPTL4 to the PRS (Table 11). The LPL / ANGPTL4 GRS was scaled so that 1 unit corresponds to a lifetime TG reduction equivalent to 0.5 standard deviations = 1.67 times the effect of Vascepa (ethyl eicosapentaenoic acid) in a clinical trial setting. The units of the PRS are standard deviations.
[0396] Table 11: Summary of the interaction of LPL / ANGPTL4 GRS with TG PRS in determining the risk of CAD in subjects not treated with lipid-lowering drugs.
[0397] [Table 11]
[0398] The TG PRS was adjusted to eliminate the contribution of LPL and ANGPTL4 to the PRS (Table 12).
[0399] Table 12: summarizes the effect of LPL / ANGPTL4 GRS (normalized to -0.5 SD TG) on risk of CAD in subjects not treated with lipid-lowering drugs, stratified by percentile of TG PRS.
[0400] [Table 12] LDL PRS × Residual CAD PRS
[0401] The LDL PRS has a significant interaction with the "residual CAD PRS" after the correlation with the LDL PRS is removed from the regression data. The ratio of this interaction effect to the marginal effect is less than half that observed when the PCSK9 GRS was used instead of the LDL PRS. This tentatively suggests that the PCSK9 GRS:CAD PRS interaction may not be explained by a class-wide effect, where all of the hypothesized LDL-lowering treatments interact with the CAD PRS. PCSK9 (as well as HMGCR, based on the observed statin-PRS interaction) appears to have something unique about it. However, because we excluded subjects taking statins from the analysis, our confidence intervals are wide, making it difficult to be confident.
[0402] The LDL PRS is scaled so that 1 unit corresponds to a lifetime LDL reduction equivalent to the LDL reduction of evolocumab in the clinical trial setting, or −1.56 standard deviations (Table 13). The outcome phenotype is CAD in subjects not treated with lipid-lowering drugs.
[0403] Table 13: A table summarizing the interaction of LDL PRS with CAD PRS conditioned on LDL PRS using a simple linear regression model.
[0404] [Table 13]
[0405] The effect of LDL PRS on the risk of CAD in subjects not treated with lipid-lowering drugs (standardized to the effect of evolocumab, −1.56 SD LDL) stratified by quantiles of CAD PRS conditioned on LDL PRS is shown in Table 14.
[0406] Table 14: summarizes the effect of LDL PRS on CAD risk in subjects not treated with lipid-lowering drugs.
[0407] [Table 14] TG PRS × Residual CAD PRS
[0408] Triglyceride PRS does not interact with the "residual CAD PRS" after the correlation with triglyceride PRS is removed from the regression data, suggesting that the LPL / ANGPTL4 GRS:CAD PRS interaction cannot be explained by a class-wide effect in which all of the hypothesized TG-lowering treatments interact with CAD PRS.
[0409] The TG PRS is scaled so that 1 unit corresponds to a lifetime TG reduction equal to 0.5 standard deviations = 1.67 times the effect of Vascepa (ethyl eicosapentaenoic acid) in a clinical trial setting (Table 15).
[0410] Table 15: is a table summarizing that the outcome phenotype is CAD in subjects not treated with lipid-lowering drugs.
[0411] [Table 15]
[0412] The effect of TG PRS (standardized to −0.5 SD TG) on CAD risk in subjects not treated with lipid-lowering drugs, stratified by quantiles of CAD PRS conditioned on TG PRS, is shown in Table 16.
[0413] Table 16: summarizes the effect of TG PRS on CAD risk in subjects not treated with lipid-lowering drugs.
[0414] [Table 16] GLP1R / GIPR GRS × CAD PRS
[0415] No association of GLP1R / GIPR GRS with CAD was observed either by itself or in interaction with CAD PRS.
[0416] Subjects receiving lipid-lowering medications were excluded from the analysis. The interaction between GLP1R / GIPR GRS and CAD PRS (with the contribution of GLP1R / GIPR GRS removed) in determining the risk of CAD is shown in Table 17. The units of GLP1R / GIPR GRS and CAD PRS are standard deviations.
[0417] Table 17 summarizes the interaction between GLP1R / GIPR GRS and CAD PRS in determining the risk of CAD.
[0418] [Table 17] PCSK9, LPL GRS × PCA-derived pathway-specific CAD PRS
[0419] The LPL / ANGPTL4 pharmacomimetic GRS did not appear to interact at all with the PRS derived from PC1 or PC2 of our PCA of variant effects on CAD-related phenotypes, but it did interact strongly with the PRS derived from PC3 of the same PCA and with the "residual PRS" obtained after removing the contributions of PC1, PC2, and PC3 from the genome-wide CAD PRS. To assess the significance of this result, we performed a likelihood ratio test to assess whether a model of CAD risk with separate GRS:PC1 PRS, GRS:PC2 PRS, GRS:PC3 PRS, and GRS:residual CAD PRS interaction terms provided a better fit than a model in which all four GRS:PRS interaction terms were constrained to be equal. This likelihood ratio test was significant (p < 0.01).
[0420] All of the variants with the highest positive loads in PC1 were located in the LPA locus, suggesting that this PC is related to lipoprotein(a) levels. High-positive load variants in PC2 increased LDL cholesterol and CAD risk and were located in loci containing many canonical regulators of LDL, such as HMGCR, PCSK9, and LDLR. High-positive load variants in PC3 increased the risk of CAD and abdominal aortic aneurysm without significantly affecting lipids. Manual review of loci containing these variants suggests that PC3 is related to vascular smooth muscle cell and / or endothelial cell function.
[0421] No specificity was observed in the interaction between PCSK9 GRS and PCA-derived PRS.
[0422] All PRSs were adjusted by linear regression to remove correlations between GRSs and PRSs. The "residual CAD PRSs" were further adjusted to remove correlations between PC1 / 2 / 3 PRSs and genome-wide CAD PRSs (Table 18). Units for each PRS are standard deviations. GRSs were scaled to match the expected drug effect (-1.56 SD LDL for PCSK9 GRS and -0.5 SD TG for LPL / ANGPTL4 GRS). Likelihood ratio tests. Table 18 summarizes the interactions between PCSK9 GRS and several PCA-derived CAD PRSs in the joint model of CAD risk in subjects not treated with lipid-lowering drugs.
[0423] [Table 18]
[0424] Model 1: cad ~ age_at_right_censoring + age_at_right_censoring_sq + sex + uk_bileve + PC1 + pc2 + pc3 + pc4 + pc5 + pc6 + grs * (cad_PC1_lpa_prs + cad_pc2_ldl_prs + cad_pc3_vascular_prs + resid-ual_cad_prs) Model 2: cad~ age_at_right_censoring + age_at_right_censoring_sq + sex + uk_bileve + PC1 + pc2 + pc3 + pc4 + pc5 + pc6 + grs + cad_PC1_lpa_prs + cad_pc2_1dl_prs + cad_pc3_ vascular_prs + residual_cad_prs + grs:I(cad_PC1_lpa_prs + cad_pc2_ldl_prs + cad_pc3_vascular_prs + resid-ual_cad_prs) #Df LogLik Df Chisq Pr(>Chisq)
[0425] All PRSs were adjusted by linear regression to remove correlations between GRSs and PRSs. The "residual CAD PRSs" were further adjusted to remove correlations between PC1 / 2 / 3 PRSs and genome-wide CAD PRSs (Table 19). The units for each PRS are standard deviations. GRSs were scaled to match the expected drug effect (-1.56 SD for LDL for PCSK9 GRS and -0.5 SD for TG for LPL / ANGPTL4 GRS). Likelihood ratio test
[0426] Table 19: Summary of interactions between LPL / ANGPTL4 GRS and several PCA-derived CAD PRS in the joint model of CAD risk in subjects not treated with lipid-lowering drugs.
[0427] [Table 19]
[0428] Model 1: cad~ age_at_right_censoring + age_at_right_censoring_sq +sex+ uk_bileve + pc1 + pc2 + pc3 + pc4 + pc5 + pc6 + grs * (cad_pc1_lpa_prs + cad_pc2_ldl_prs + cad_pc3_vascular_prs + resid-ual_cad_prs) Model 2: cad~ age_at_right_censoring + age_at_right_censoring_sq +sex+ uk_bileve + pc1 + pc2 + pc3 + pc4 + pc5 + pc6 + grs + cad_pc1_lpa_prs + cad_pc2_ldl_prs + cad_pc3_ vascular_prs + residua.l_cad_prs + grs:I( cad_pc1_lpa_prs + cad_pc2_ldl_prs + cad_pc3_ vascular_prs + residual_cad_prs) 2 17 -50193 -3 13.924 0.003011 ** - Signif. codes: 0 '' O. 001 " 0. 01 " 0.05 '.' 0.1 ' ' 1
[0429] When the contributions of PC1 and PS2 are removed from the CAD PRS, an adjusted CAD PRS is obtained. This adjusted CAD PRS is slightly better than the unadjusted genome-wide CAD PRS at predicting subject response to the LPL / ANGPTL4 pharmacomimetic GRS, even though the adjusted PRS has a marginal effect on CAD that is only 73% of the marginal effect of the genome-wide PRS on CAD. Essentially, a more parsimonious PRS can be constructed that performs at least as well as the default PRS when applied to studies of this drug mechanism.
[0430] The CAD PRS was adjusted to remove the contributions of PC1, PC2, LPL, and ANGPTL4 to the PRS (Table 20).
[0431] Table 20: Summary of the interaction between LPL / ANGPTL4 GRS and CAD PRS in determining the risk of CAD.
[0432] [Table 20]
[0433] Subjects taking lipid- or blood pressure-lowering medications were excluded from the analysis. The LPL / ANGPTL4 GRS is scaled so that 1 unit corresponds to a lifetime TG reduction equivalent to 0.5 standard deviations = 1.67 times the effect of Vascepa (ethyl eicosapentaenoate) in a clinical trial setting. The units of the PRS are standard deviations.
[0434] The CAD PRS was adjusted to remove the contributions of PC1, PC2, LPL, and ANGPTL4 to the PRS (Table 21). Subjects taking lipid- or blood pressure-lowering medications were excluded from the analysis. The LPL / ANGPTL4 GRS was scaled so that 1 unit corresponds to a lifetime TG reduction equivalent to 0.5 standard deviations = 1.67 times the effect of Vascepa (ethyl eicosapentaenoate) in a clinical trial setting. The units of the PRS are standard deviations.
[0435] Table 21 summarizes the effect of LPL / ANGPTL4 GRS on CAD risk stratified by quantiles of CAD PRS.
[0436] [Table 21]
[0437] Using PCA on variant effects on CAD and CAD-related phenotypes, we were able to construct pathway-specific PRSs. These two PRSs (PC1 and PC2) did not appear to interact with LPL pharmacomimetic variants. Adjusting the genome-wide CAD PRS to remove contributions from PC1 and PC2 resulted in a small but significant (p < 0.01) improvement in the magnitude of the LPL pharmacomimetic variant × PRS interaction, while also increasing PRS sparsity. Example 4: Materials and Methods Phenotype:
[0438] Coronary artery disease (CAD) - ICD-10 I20-I25 ICD-9 410, 412, 414; excluding controls with 411, 413 - Self-reported myocardial infarction or angina - OPC-4 K40-K46, K49-K50, K75 - Note: This phenotype definition is broader than that used in much of the published literature because stable angina was included in the case definition. We found that using a broad phenotype definition improved the power to detect genetic association compared with narrower phenotypes. - Note: In analyses using CAD as the outcome, the applicant excluded subjects who self-reported taking lipid-lowering or antihypertensive medications.
[0439] Hypertension (HTN) - Applicant used medication-adjusted systolic blood pressure as a proxy for hypertension. - Subjects who self-reported taking blood pressure medication had 15mmHg added to their blood pressure readings. Blood pressure medications included ACE inhibitors, angiotensin II receptor blockers, calcium channel blockers, diuretics, alpha-blockers, beta-blockers, alpha-2 receptor agonists, central agonists, peripheral adrenergic agonists, and vasodilators.
[0440] Venous thromboembolism (VTE) - ICD-10 I26, I80, I82.9, O22.3, O87.1; excluding controls with I74, I81-I82, O22, O87, O88.2. ICD-9 4151, 451, 453.9, 671.3, 671.4; excluding controls with 415.0, 444, 452-453, 593.81, 593.82, 671, 673.2. - Self-reported venous thromboembolism, pulmonary embolism, or deep vein thrombosis (1068, 1093, 1094); - Excluding controls with peripheral vascular disease or leg claudication (1067, 1087).
[0441] Atrial fibrillation (AFib) - ICD-10 I48; excluding controls with I47, I49. - ICD-9 427.3; excluding controls with 427. - Self-reported atrial fibrillation or flutter (1471, 1483); excluding controls with other arrhythmias (1077, 1484, 1485, 1486, 1487). - This phenotype includes atrial flutter. Applicants have discovered that using a broad atrial fibrillation phenotype that includes atrial flutter improves the power to detect genetic associations compared to narrower phenotypes.
[0442] ·Chronic liver disease (CLD) - The applicant used alanine aminotransferase (ALT) levels as a surrogate for CLD, and ALT levels above 84 U / L were set at 84 U / L, i.e., the phenotype was winsorized.
[0443] ·Type 2 diabetes (T2D) - This phenotype excludes subjects with type 1 diabetes from both case and control groups, but allows for cases of unspecified diabetes without a T1D code. - ICD-10 E11, E14; exclude cases and controls with E10; exclude controls with E12-E13. - ICD-9 250 - Self-reported diabetes or type 2 diabetes (1220, 1223); exclude cases and controls with type 1 diabetes (1222); exclude controls with gestational diabetes (1221) - Doctor-diagnosed diabetes (UKB data field 2443) - Any of multiple self-reported medications, including metformin, PPARG agonists, sulfonylureas, non-sulfonylurea secretagogues, and alpha-glucosidase inhibitors.
[0444] Obesity (BMI) Body mass index was used as a proxy for obesity. - In analyses using BMI, Applicant focused specifically on obese (BMI >= 30) subjects, as this is the population most likely to receive weight loss medications. Some pharmacomimetic variants appear to have a significantly amplified effect on BMI in obese subjects, even when expressed in relative terms (i.e., percentage change in weight). Applicant suspects this is due to gene x environment interactions.
[0445] · Gout - ICD-10 M10 - ICD-9 274 - Self-reported gout (1466)
[0446] Chronic kidney disease (CKD) - Estimated glomerular filtration rate (eGFR) was used as a surrogate for CKD. - eGFR was estimated using the 2021 CKD-EPI(crea-cys) equation (Inker et al. 2021).
[0447] Age-related macular degeneration (AMD) - ICD-10 H35.3; exclude controls with E10.3, E11.3, E12.3, E13.3, E14.3, H35, H36.0. - ICD-9 362.5; exclude controls with 362. - OPCS-4: Exclude controls with C79-C85 and X93. - Self-reported macular degeneration (1528); exclude controls with diabetic eye disease.
[0448] Glaucoma - ICD-10: H40, 42 - ICD-9 365 - OPCS-4, C60-C61, C65 - Self-reported glaucoma - Any one of several glaucoma medications, including prostaglandin analogues (timolol, dorzolamide, brinzolamide, brimonidine tartrate, levobunolol). For medications with both systemic and ocular uses, only eye drops were used.
[0449] Alzheimer's disease (AD) - Due to the limited sample size of Alzheimer's disease cases ascertained in UK Biobank, the applicant used a broad phenotype definition that included the code for "unspecified dementia." Excluding controls with ICD-10 G30, F00, F03; F01-F02, F04-F05, F06.7, G31, G32, R32, R41, R54. ICD-9 331.0; excluding controls with 290, 331 - Exclude controls with self-reported dementia or cognitive impairment (1263).
[0450] Alzheimer's Disease (AD) (parent history score) - Applicant used the method of Janssen et al. 2019 to define a score combining the observed AD in the proband and self-reported parental history of AD. The phenotype was defined as the number of affected parents ranging from 0 to 2, with a value of 2 for a proband with AD. Affected parents contribute a value of 1 to the score. Unaffected parents contribute a fractional value reflecting their likelihood of developing AD in the future or the likelihood they would have developed the disease had they not died of another cause. The fractional value was calculated as min(0.32, (100 - current age or age at death, in years) / 100). 0.32 is an estimate of the maximum population prevalence of AD in people who survive to very old age.
[0451] Parkinson's disease - ICD-10:F02.3, G20; excluding controls with G21-G22. - ICD-9: 332.0; excluding controls with 332.1. - Self-reported Parkinson's disease (1262) Parkinson's disease (parent history score) - This was constructed in the same way as the AD parent history score above, except that the fractional value for the unaffected parent was set to min(0.029, (85 - current age or age at death) / 85).
[0452] ·migraine ICD-10 G43; excluding controls with G44 - ICD-9 346 - Self-reported migraine - Migraine medication
[0453] ·asthma - ICD-10 J45-46; excluding controls with J30, J33, and L20 - ICD-9 493; excluding controls with 471, 477, 691.80 - OPCS-4: Excluding controls with E08.1 - Self-reported asthma (1111); excludes controls with allergic rhinitis, nasal polyps, or atopic dermatitis
[0454] Nasal polyps - ICD-10 J33; excluding controls with J30, J45, J46, L20. - ICD-9 471; excluding controls with 477, 493, 691.80 - OPCS-4 E08.1 - Self-reported nasal polyps; excludes controls with asthma, allergic rhinitis, or atopic dermatitis.
[0455] ·Chronic obstructive pulmonary disease (COPD) - ICD-10 J41-J44; excluding controls with J40. - ICD-9 491-492; excluding controls with 490. - Self-reported COPD, chronic bronchitis, or emphysema (1112, 1113, 1472); exclude controls with unspecified bronchitis.
[0456] Idiopathic pulmonary fibrosis (IPF) - ICD-10 J84.1; excluding controls with D86.0, D86.2, J60-J68, J70, J84.0, J84.2, J84.8, J84.9, J99, M05.1. - ICD-9 516.3; excluding controls with 495, 500-506, 508, 515-517. - Exclude controls with self-reported interstitial lung disease, asbestosis, pulmonary fibrosis, or fibrosing alveolitis (1115, 1120, 1121, 1122).
[0457] Osteoporosis - The applicant used heel bone mineral density (BMD) as a surrogate for osteoporosis.
[0458] Osteoarthritis of the knee (OAofKnee) - ICD-10 M17; excluding controls with M05-M16, M18-M19, M25.5, M79.0. - Excluding controls with ICD-9 712-716, 719.4, 729.0. - Excludes controls with self-reported rheumatoid arthritis (1464), osteoarthritis (1465), gout (1466), other joint disorders (1467), psoriatic arthropathy (1477), joint pain (1537), or "arthritis (NOS)" (1538).
[0459] Rheumatoid arthritis (RA) - ICD-10 M05-M06; excluding controls with M08.0. - ICD-9 714.0-714.2 - Self-reported rheumatoid arthritis (1464)
[0460] Crohn's disease (Crohn's disease) - Self-reported Crohn's disease (1462) - Excludes controls with ICD-10 K50 or ICD-9 555. - : The applicant found that not including ICD codes in the case criteria significantly improved the consistency of the UKB Crohn's disease GWAS with published Crohn's disease GWAS from high-quality case-control cohorts. This may be related to the fact that the ICD-10 K50 code is called "Crohn's disease [regional enterocolitis]." Perhaps cases of enterocolitis other than Crohn's disease are erroneously recorded using this code.
[0461] ·Ulcerative colitis (UC) - ICD-10 K51 - ICD-9 556 - Self-reported ulcerative colitis (1463)
[0462] ·Inflammatory bowel disease (IBD) - Combination of the Crohn's disease and UC phenotypes listed above
[0463] ·Atopic dermatitis (AtoDerm) - ICD-10 L20; excluding controls with J30, J33, J45-J46. - ICD-9 691.80; excluding controls with 471, 477, 493. - OPCS-4; except for controls with E08.1. - Self-reported eczema / dermatitis (1452); excluding controls with asthma, allergic rhinitis, or nasal polyps (1111, 1387, 1417).
[0464] ·psoriasis - ICD-10 L40.0-L40.3, L40.8-L40.9; excluding controls with L40.4-L40.5 - Self-reported psoriasis (1453)
[0465] ·Type 1 diabetes (T1D) - ICD-10 E10 - Self-reported type 1 diabetes (1222)
[0466] ·Ankylosing spondylitis - ICD-10 M45; excluding controls with M46. - ICD-9 720.0; excluding controls with 720. - Self-reported ankylosing spondylitis (1313), excluding controls with self-reported "spondyloarthritis / spondylitis" (1311).
[0467] Celiac disease - ICD-10 K90.0 - ICD-9 579.0 -Excluding controls with self-reported "malabsorption / celiac disease" (1456).
[0468] ·Multiple sclerosis (MS) - ICD-10 G35 - ICD-9 340.9 - Self-reported multiple sclerosis (1261)
[0469] ·Hypothyroidism - ICD-10 E03 - ICD-9 243-244 - Self-reported hypothyroidism (1226) - Self-reported levothyroxine prescription (1141191044)
[0470] Diverticulosis - ICD-10 K57 - ICD-9 562 - Self-reported "diverticulosis / diverticulitis" (1458)
[0471] ·Uterine fibroids - Exclude men from cases and controls - ICD-10 D25-D26; excluding controls with C50-C58, D24, D27-D28, N93.8, N93.9, N94, N97, Z90.7. - ICD-9 218-219; excluding controls with 174, 179-184, 217, 220-221, 625, 626, 628. - OPCS-4 Q09.2; Q07-Q09, excluding controls with Y75.2. - OPCS-3 700.1; excluding controls with 691-698. - Self-reported uterine fibroids (1351); excluding controls with self-reported female infertility (1403), menorrhagia (1556), or dysmenorrhea (1664).
[0472] ·breast cancer - ICD-10 C50 - ICD-9 174 - Self-reported breast cancer (1002)
[0473] Prostate cancer - ICD-10 C61 - ICD-9 185 - Self-reported prostate cancer (1044)
[0474] ·Non-melanoma skin cancer - ICD-10 C44; excluding controls with C43. - ICD-9 173; excluding controls with 172. - Self-reported non-melanoma skin cancer (1060), basal cell carcinoma (1061), or squamous cell carcinoma (1062); exclude controls with self-reported skin cancer (1003) or melanoma (1059). Variant Selection
[0475] For each disease, Applicants performed a variant:PRS interaction test using a set of variants that met all of the following criteria: The variant is associated with either the disease or a disease-related trait with genome-wide significance (p < 5e-8). The variant is disease-associated with p < 5e-5. If a variant was not associated with disease with genome-wide significance, it had to be uniquely mapped to a gene using eQTL / pQTL colocalization analysis or fine mapping of coding variants. In QTL analysis, Applicant only used results from QTL datasets that Applicant determined to be biologically relevant to disease.
[0476] When multiple independent variants meeting the selection criteria mapped to the same gene for the same disease, we used only the variant with the most significant association with the binary disease phenotype. For this filtering, genome-wide significant variants lacking formal v2g mapping were mapped to the nearest protein-coding gene.
[0477] For computational convenience, a maximum of 250 variants were tested per disease. If there were additional variants meeting the above criteria, Applicants tested only the top 250, ranked by p-value for association with disease. Computer Modeling
[0478] For each (variant, disease) pair, applicants fit a logistic regression model to predict disease, or, for diseases where applicants use a biomarker (e.g., BMI for obesity) as a surrogate for disease, a linear regression model. In each model, predictors were the number of variants, the PRS of the disease, an interaction term between the number of variants and the PRS, and the following covariates: age, age squared, sex, genotyping chip, and PCs 1-6 of genetic ancestry.
[0479] In theory, the PRS can include any specific variant being tested. Applicant performed several experiments and found that this does not seem to make a significant difference, as each variant makes a small contribution to the overall PRS. However, to be safe, before fitting the model, Applicant adjusted the PRS to correct for the number of variants. Specifically, Applicant fitted a linear regression model: PRS ~ number of tested variants, and defined the corrected PRS as the original PRS minus the estimate from this model.
[0480] For quality control, we performed full GWAS of variant × PRS interactions for each disease phenotype. These are not exactly the same models we used in our main analyses because we do not remove the effect of each variant from the PRS before association testing, but they are informative as to whether interaction tests are well calibrated or produce inflated test statistics. By referencing the LD score regression intercept (Bulik-Sullivan et al. 2015) for each GWAS, we found no, minimal, or moderate inflated test statistics, depending on the phenotype. The decay rate for GWASs with inflated test statistics, i.e., (LDSC intercept - 1) / (average chi-square statistic - 1) (Loh et al. 2018), is usually greater than 0.5, indicating that the inflation is real and not an artifact of the high statistical power of GWASs. (Because the scaling of the LD score regression intercept depends on the sample size and trait heritability of the GWAS, even very small / negligible miscalibration in the GWAS model can increase the intercept rate for powerful traits such as BMI and height; however, this was not the case for our GWAS.)
[0481] Therefore, to control our type I error (false positive) rate, we divided the chi-squared test statistic for each variant × PRS interaction test by the LD score regression intercept of all variant × PRS interaction GWAS for a particular disease phenotype.
[0482] Applicants then recalculated p-values from the adjusted test statistics. If the LD score regression intercept was less than 1, Applicants did not apply any correction.
[0483] Finally, to account for multiple testing, we used a Bonferroni significance threshold of 0.05 / number of variants tested for a particular disease. We treated each disease analysis separately and did not correct for multiple testing across diseases. equivalent
[0484] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this technology belongs.
[0485] The technology illustratively described herein may suitably be practiced in the absence of one or more elements, limitation, or limitations not specifically disclosed herein. Thus, for example, terms such as "comprising," "including," and "containing" should be interpreted expansively and without limitation. Furthermore, the terms and expressions employed herein are used as descriptive terms and not limiting terms. The use of such terms and expressions is not intended to exclude equivalents of the functions shown and described or portions thereof, and it is recognized that various modifications are possible within the scope of the claims of the technology.
[0486] Accordingly, it should be understood that the materials, methods, and examples provided herein are representative of preferred embodiments, are exemplary, and are not intended as limitations on the scope of the technology.
[0487] While the present invention has been specifically disclosed by certain aspects, embodiments, and optional features, it is to be understood that improvements, modifications, and variations of such aspects, embodiments, and optional features may be made by those skilled in the art, and such improvements, modifications, and variations are considered to be within the scope of the present disclosure.
[0488] The present technology has been described broadly and generically herein. Each of the narrower species and subgeneric groupings falling within the generic disclosure also form part of the present technology. This includes any general description of the technology with a provisory or negative limitation that removes any subject matter from the genus, regardless of whether the excised material is specifically described herein.
[0489] Furthermore, where features or aspects of the technology are described in terms of a Markush group, those skilled in the art will recognize that the technology is also thereby described in terms of any individual member or subgroup of members of the Markush group.
[0490] All publications, patent applications, patents, and other references mentioned herein are expressly incorporated by reference in their entirety to the same extent as if each were individually incorporated by reference. In the case of conflict, the present specification, including definitions, will control.
[0491] Other aspects are set forth in the following claims.
Claims
1. 1. An in silico method for determining drug activity of multiple drug targets, comprising: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the molecular biomarker stratification factor data including a plurality of biomarker stratification factors and at least one disease phenotype; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; identifying a predicted drug activity of each drug target for said at least one disease phenotype in the subset of biomarker stratification factor distributions based on a statistical interaction between said biomarker stratification factor score and said pharmacomimetic gene score for each drug target in an association analysis with said disease phenotype; A method comprising:
2. (a) the biomarker stratification factor comprises a genetic variant, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) Polygenic score of drug target expression; or any linear or non-linear combination thereof, or (b) the biomarker stratification factor comprises a proteome variant, and wherein the biomarker stratification factor score is: (i) Proteomic score of disease phenotype; (ii) proteomic scores of disease risk factors; (iii) proteomic scores of biological pathway activity; (iv) Proteome score of drug target expression: or any linear or non-linear combination thereof, or (c) said biomarker stratification factor comprises a transcription variant, and wherein said biomarker stratification factor score is: (i) Transcriptome score of disease phenotype; (ii) transcriptome scores for disease risk factors; (iii) transcriptome scores of biological pathway activity; or any linear or non-linear combination thereof, or (d) the biomarker stratification factor comprises a somatic mutation variant, and wherein the biomarker stratification factor score is: (i) somatic mutation score of drug target expression; (ii) somatic mutation score of the disease phenotype; (iii) somatic mutation scores for disease risk factors; (iv) somatic mutation scores for biological pathway activity; (v) somatic mutation score of drug target expression; or any linear or non-linear combination thereof, or (e) the biomarker stratification factors include genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) polygenic score of drug target expression; (v) proteomic score of the disease phenotype; (vi) Proteome scores of disease risk factors: (vii) proteomic scores of biological pathway activity; (viii) proteomic scores of biological pathway activity; (ix) proteomic score of drug target expression; (x) a transcriptome score for the disease phenotype; (xi) Transcriptome score of disease risk factors; (xii) transcriptome scores of biological pathway activity; (xiii) somatic mutation score of drug target expression; (xiv) somatic mutation score of the disease phenotype; (xv) disease risk factor somatic mutation score; (xvi) somatic mutation score of biological pathway activity; (xvii) somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
3. 3. The method of claim 1, further comprising performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
4. The method of claim 3 , further comprising the step of performing the principal component analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratification factor effects.
5. 5. The method of claim 4, further comprising constructing a matrix based on one or more of the biomarker stratification factors, the phenotype, and the plurality of values.
6. The method of any of claims 1-5, further comprising adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
7. The method of claims 3-6, wherein the weighted principal component analysis is performed according to biomarker stratification factor scores of the selected disease phenotypes.
8. The method of claims 3-7, wherein the weighted principal component analysis identifies predicted drug activity of multiple drug targets against at least one disease phenotype.
9. 1. An in silico method for determining drug activity of multiple drug targets associated with nonalcoholic steatohepatitis (NASH), comprising: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the molecular biomarker stratification factor data including a plurality of biomarker stratification factors and at least one disease phenotype associated with nonalcoholic steatohepatitis (NASH); determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for the selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; identifying a predicted drug activity of each drug target for said at least one disease phenotype in the subset of biomarker stratification factor distributions based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score of each drug target in an association analysis with the disease phenotype; A method comprising:
10. (a) the biomarker stratification factor comprises a genetic variant, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) Polygenic score of drug target expression; or any linear or non-linear combination thereof, or (b) the biomarker stratification factor comprises a proteome variant, and wherein the biomarker stratification factor score is: (i) Proteomic score of disease phenotype; (ii) proteomic scores of disease risk factors; (iii) proteomic scores of biological pathway activity; (iv) Proteome score of drug target expression: or any linear or non-linear combination thereof, or (c) said biomarker stratification factor comprises a transcription variant, and wherein said biomarker stratification factor score is: (i) Transcriptome score of disease phenotype; (ii) transcriptome scores for disease risk factors; (iii) transcriptome scores of biological pathway activity; or any linear or non-linear combination thereof, or (d) the biomarker stratification factor comprises a somatic mutation variant, and wherein the biomarker stratification factor score is: (i) somatic mutation score of drug target expression; (ii) somatic mutation score of the disease phenotype; (iii) somatic mutation scores for disease risk factors; (iv) somatic mutation scores for biological pathway activity; (v) somatic mutation score of drug target expression; or any linear or non-linear combination thereof, or (e) the biomarker stratification factors include genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) polygenic score of drug target expression; (v) proteomic score of the disease phenotype; (vi) Proteome scores of disease risk factors: (vii) proteomic scores of biological pathway activity; (viii) proteomic scores of biological pathway activity; (ix) proteomic score of drug target expression; (x) a transcriptome score for the disease phenotype; (xi) Transcriptome score of disease risk factors; (xii) transcriptome scores of biological pathway activity; (xiii) somatic mutation score of drug target expression; (xiv) somatic mutation score of the disease phenotype; (xv) disease risk factor somatic mutation score; (xvi) somatic mutation score of biological pathway activity; (xvii) somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
11. 11. The method of claim 9, further comprising the step of performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
12. The method of claim 11 , further comprising performing the principal component analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratification factor effects.
13. 13. The method of claim 12, further comprising constructing the matrix based on one or more of the biomarker stratification factors, the phenotype, and the plurality of values.
14. The method of any of claims 9-13, further comprising adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
15. The method of any of claims 9-14, wherein the weighted principal component analysis is performed according to the biomarker stratification factor scores of the selected disease phenotypes.
16. The method of any of claims 9-15, wherein the weighted principal component analysis identifies predicted drug activity of multiple drug targets against at least one disease phenotype.
17. The method of any of claims 11-16, wherein the biomarker stratification factor score is a fatty liver polygenic score.
18. The method of any of claims 11-17, wherein the pharmacomimetic gene score is associated with response to a drug that targets HSD17B13.
19. The method of claim 18, wherein the drug that targets HSD17B13 is an HSD17B13 inhibitor.
20. 1. An in silico method for determining drug activity of multiple drug targets associated with coronary heart disease, comprising: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the molecular biomarker stratification factor data including a plurality of biomarker stratification factors and at least one disease phenotype associated with coronary heart disease; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; identifying a predicted drug activity of each drug target for said at least one disease phenotype in the subset of biomarker stratification factor distributions based on a statistical interaction between said biomarker stratification factor score and said pharmacomimetic gene score of each drug target in an association analysis with said disease phenotype; A method comprising:
21. (a) the biomarker stratification factor comprises a genetic variant, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) Polygenic score of drug target expression; or any linear or non-linear combination thereof, or (b) the biomarker stratification factor comprises a proteome variant, and wherein the biomarker stratification factor score is: (i) Proteomic score of disease phenotype; (ii) proteomic scores of disease risk factors; (iii) proteomic scores of biological pathway activity; (iv) Proteome score of drug target expression: or any linear or non-linear combination thereof, or (c) said biomarker stratification factor comprises a transcription variant, and wherein said biomarker stratification factor score is: (i) Transcriptome score of disease phenotype; (ii) transcriptome scores for disease risk factors; (iii) transcriptome scores of biological pathway activity; or any linear or non-linear combination thereof, or (d) the biomarker stratification factor comprises a somatic mutation variant, and wherein the biomarker stratification factor score is: (i) somatic mutation score of drug target expression; (ii) somatic mutation score of the disease phenotype; (iii) somatic mutation scores for disease risk factors; (iv) somatic mutation scores for biological pathway activity; (v) somatic mutation score of drug target expression; or any linear or non-linear combination thereof, or (e) the biomarker stratification factors include genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) polygenic score of drug target expression; (v) proteomic score of the disease phenotype; (vi) Proteome scores of disease risk factors: (vii) proteomic scores of biological pathway activity; (viii) proteomic scores of biological pathway activity; (ix) proteomic score of drug target expression; (x) a transcriptome score for the disease phenotype; (xi) Transcriptome score of disease risk factors; (xii) transcriptome scores of biological pathway activity; (xiii) somatic mutation score of drug target expression; (xiv) somatic mutation score of the disease phenotype; (xv) disease risk factor somatic mutation score; (xvi) somatic mutation score of biological pathway activity; (xvii) somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
22. The method of any of claims 20-21, further comprising the step of performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
23. 23. The method of claim 22, further comprising performing said principal component analysis or said weighted principal component analysis based on a matrix of one or more of said biomarker stratification factor effects.
24. 24. The method of claim 23, further comprising constructing the matrix based on one or more of the biomarker stratification factors, the phenotype, and the plurality of values.
25. The method of claims 20-24, further comprising adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
26. The method of claims 22-25, wherein the weighted principal component analysis is performed according to biomarker stratification factor scores of the selected disease phenotypes.
27. 27. The method of any of claims 22-26, wherein said weighted principal component analysis identifies predicted drug activities of multiple drug targets against said at least one disease phenotype.
28. The method of any one of claims 20-27, wherein the pharmacomimetic gene score is associated with response to drugs that target LPL, ANGPTL4 and / or ANGPTL3.
29. The method of claim 28 , wherein the drug comprises an LPL agonist, an ANGPTL4 inhibitor and / or an ANGPTL3 inhibitor.
30. 1. An in silico method for determining drug activity of multiple drug targets associated with dry age-related macular degeneration, comprising: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the molecular biomarker stratification factor data including a plurality of biomarker stratification factors and at least one disease phenotype associated with dry age-related macular degeneration; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; identifying a predicted drug activity of each drug target for said at least one disease phenotype in the subset of biomarker stratification factor distributions based on a statistical interaction between said biomarker stratification factor score and said pharmacomimetic gene score of each drug target in an association analysis with said disease phenotype; A method comprising:
31. (a) the biomarker stratification factor comprises a genetic variant, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) Polygenic score of drug target expression; or any linear or non-linear combination thereof, or (b) the biomarker stratification factor comprises a proteome variant, and wherein the biomarker stratification factor score is: (i) Proteomic score of disease phenotype; (ii) proteomic scores of disease risk factors; (iii) proteomic scores of biological pathway activity; (iv) Proteome score of drug target expression: or any linear or non-linear combination thereof, or (c) said biomarker stratification factor comprises a transcription variant, and wherein said biomarker stratification factor score is: (i) Transcriptome score of disease phenotype; (ii) transcriptome scores for disease risk factors; (iii) transcriptome scores of biological pathway activity; or any linear or non-linear combination thereof, or (d) the biomarker stratification factor comprises a somatic mutation variant, and wherein the biomarker stratification factor score is: (i) somatic mutation score of drug target expression; (ii) somatic mutation score of the disease phenotype; (iii) somatic mutation scores for disease risk factors; (iv) somatic mutation scores for biological pathway activity; (v) somatic mutation score of drug target expression; or any linear or non-linear combination thereof, or (e) the biomarker stratification factors include genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) polygenic score of drug target expression; (v) proteomic score of the disease phenotype; (vi) Proteome scores of disease risk factors: (vii) proteomic scores of biological pathway activity; (viii) proteomic scores of biological pathway activity; (ix) proteomic score of drug target expression; (x) a transcriptome score for the disease phenotype; (xi) Transcriptome score of disease risk factors; (xii) transcriptome scores of biological pathway activity; (xiii) somatic mutation score of drug target expression; (xiv) somatic mutation score of the disease phenotype; (xv) disease risk factor somatic mutation score; (xvi) somatic mutation score of biological pathway activity; (xvii) somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
32. 32. The method of any of claims 30-31, further comprising the step of performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
33. 33. The method of claim 32, further comprising performing said principal component analysis or said weighted principal component analysis based on a matrix of one or more of said biomarker stratification factor effects.
34. 34. The method of claim 33, further comprising constructing the matrix based on one or more of the biomarker stratification factors, the phenotype, and the plurality of values.
35. The method of any one of claims 30-34, further comprising adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
36. The method of claims 32-35, wherein the weighted principal component analysis is performed according to the biomarker stratification factor scores of the selected disease phenotypes.
37. 37. The method of claims 32-36, wherein said weighted principal component analysis identifies predicted drug activities of multiple drug targets against said at least one disease phenotype.
38. 38. The method of any one of claims 30-37, wherein the pharmacomimetic gene score is associated with response to drugs that target C3, CFB, CFH, and / or HTRA1.
39. 39. The method of claim 38, wherein the drug comprises a C3, CFB, CFH, and / or HTRA1 inhibitor.
40. 1. An in silico method for determining drug activity of multiple drug targets associated with rheumatoid arthritis, comprising: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the molecular biomarker stratification factor data comprising a plurality of biomarker stratification factors and at least one disease phenotype associated with rheumatoid arthritis; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating the biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; identifying a predicted drug activity of each drug target for said at least one disease phenotype in the subset of biomarker stratification factor distributions based on a statistical interaction between said biomarker stratification factor score and said pharmacomimetic gene score of each drug target in an association analysis with said disease phenotype; A method comprising:
41. (a) the biomarker stratification factor comprises a genetic variant, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) Polygenic score of drug target expression; or any linear or non-linear combination thereof, or (b) the biomarker stratification factor comprises a proteome variant, and wherein the biomarker stratification factor score is: (i) Proteomic score of disease phenotype; (ii) proteomic scores of disease risk factors; (iii) proteomic scores of biological pathway activity; (iv) Proteome score of drug target expression: or any linear or non-linear combination thereof, or (c) said biomarker stratification factor comprises a transcription variant, and wherein said biomarker stratification factor score is: (i) Transcriptome score of disease phenotype; (ii) transcriptome scores for disease risk factors; (iii) transcriptome scores of biological pathway activity; or any linear or non-linear combination thereof, or (d) the biomarker stratification factor comprises a somatic mutation variant, and wherein the biomarker stratification factor score is: (i) somatic mutation score of drug target expression; (ii) somatic mutation score of the disease phenotype; (iii) somatic mutation scores for disease risk factors; (iv) somatic mutation scores for biological pathway activity; (v) somatic mutation score of drug target expression; or any linear or non-linear combination thereof, or (e) the biomarker stratification factors include genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) polygenic score of drug target expression; (v) proteomic score of the disease phenotype; (vi) Proteome scores of disease risk factors: (vii) proteomic scores of biological pathway activity; (viii) proteomic scores of biological pathway activity; (ix) proteomic score of drug target expression; (x) a transcriptome score for the disease phenotype; (xi) Transcriptome score of disease risk factors; (xii) transcriptome scores of biological pathway activity; (xiii) somatic mutation score of drug target expression; (xiv) somatic mutation score of the disease phenotype; (xv) disease risk factor somatic mutation score; (xvi) somatic mutation score of biological pathway activity; (xvii) somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
42. 42. The method of any of claims 40-41, further comprising the step of performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of said one or more biomarker stratification factor effects.
43. 43. The method of claim 42, further comprising the step of performing said principal component analysis or said weighted principal component analysis based on a matrix of one or more said biomarker stratification factor effects.
44. 44. The method of claim 43, further comprising constructing the matrix based on one or more of the biomarker stratification factors, the phenotype, and the plurality of values.
45. 45. The method of claims 40-44, further comprising adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
46. 46. The method of claims 42-45, wherein said weighted principal component analysis is performed according to said biomarker stratification factor scores of said selected disease phenotypes.
47. 47. The method of any of claims 42-46, wherein said weighted principal component analysis identifies predicted drug activities of multiple drug targets against said at least one disease phenotype.
48. 48. The method of any one of claims 40-47, wherein the pharmacomimetic gene score is associated with response to a drug that targets TYK2.
49. 49. The method of claim 48, wherein the drug comprises a TYK2 inhibitor.
50. 1. An in silico method for determining drug activity of multiple drug targets associated with atopic disease, comprising: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the molecular biomarker stratification factor data including a plurality of biomarker stratification factors and at least one disease phenotype associated with an atopic disease; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; identifying a predicted drug activity of each drug target for said at least one disease phenotype in the subset of biomarker stratification factor distributions based on a statistical interaction between said biomarker stratification factor score and said pharmacomimetic gene score of each drug target in an association analysis with said disease phenotype; A method comprising:
51. (a) the biomarker stratification factor comprises a genetic variant, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) Polygenic score of drug target expression; or any linear or non-linear combination thereof, or (b) the biomarker stratification factor comprises a proteome variant, and wherein the biomarker stratification factor score is: (i) Proteomic score of disease phenotype; (ii) proteomic scores of disease risk factors; (iii) proteomic scores of biological pathway activity; (iv) Proteome score of drug target expression: or any linear or non-linear combination thereof, or (c) said biomarker stratification factor comprises a transcription variant, and wherein said biomarker stratification factor score is: (i) Transcriptome score of disease phenotype; (ii) transcriptome scores for disease risk factors; (iii) transcriptome scores of biological pathway activity; or any linear or non-linear combination thereof, or (d) the biomarker stratification factor comprises a somatic mutation variant, and wherein the biomarker stratification factor score is: (i) somatic mutation score of drug target expression; (ii) somatic mutation score of the disease phenotype; (iii) somatic mutation scores for disease risk factors; (iv) somatic mutation scores for biological pathway activity; (v) somatic mutation score of drug target expression; or any linear or non-linear combination thereof, or (e) the biomarker stratification factors include genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) polygenic score of drug target expression; (v) proteomic score of the disease phenotype; (vi) Proteome scores of disease risk factors: (vii) proteomic scores of biological pathway activity; (viii) proteomic scores of biological pathway activity; (ix) proteomic score of drug target expression; (x) a transcriptome score for the disease phenotype; (xi) Transcriptome score of disease risk factors; (xii) transcriptome scores of biological pathway activity; (xiii) somatic mutation score of drug target expression; (xiv) somatic mutation score of the disease phenotype; (xv) disease risk factor somatic mutation score; (xvi) somatic mutation score of biological pathway activity; (xvii) somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
52. 52. The method of any of claims 50-51, further comprising the step of performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of said one or more biomarker stratification factor effects.
53. 53. The method of claim 52, further comprising the step of performing said principal component analysis or said weighted principal component analysis based on a matrix of one or more said biomarker stratification factor effects.
54. 54. The method of claim 53, further comprising constructing the matrix based on one or more of the biomarker stratification factors, the phenotype, and the plurality of values.
55. 55. The method of claims 50-54, further comprising adjusting the gene score for each drug target and the biomarker stratification factor score for each principal component.
56. 56. The method of claims 52-55, wherein said weighted principal component analysis is performed according to said biomarker stratification factor scores of said selected disease phenotypes.
57. 57. The method of claims 52-56, wherein said weighted principal component analysis identifies predicted drug activities of multiple drug targets against said at least one disease phenotype.
58. 58. The method of any one of claims 50-57, wherein the pharmacomimetic gene score is associated with response to drugs that target IL33, TSLP and / or IL4R.
59. 59. The method of claim 58, wherein the drug comprises an IL33, TSLP and / or IL4R inhibitor.
60. 1. An in silico method for determining drug activity of multiple drug targets associated with obesity, comprising: obtaining molecular biomarker stratification factor data from each of a plurality of subjects, the molecular biomarker stratification factor data including a plurality of biomarker stratification factors and at least one disease phenotype associated with obesity; determining a plurality of values representing biomarker stratification factor effects, each value separately representing how each of the plurality of biomarker stratification factors affects each of the at least one disease phenotype based on external data; calculating a biomarker stratification factor score for a selected disease phenotype; calculating a pharmacomimetic gene score for each drug target; identifying a predicted drug activity of each drug target for said at least one disease phenotype in the subset of biomarker stratification factor distributions based on a statistical interaction between said biomarker stratification factor score and said pharmacomimetic gene score of each drug target in an association analysis with said disease phenotype; A method comprising:
61. (a) the biomarker stratification factor comprises a genetic variant, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) Polygenic score of drug target expression; or any linear or non-linear combination thereof, or (b) the biomarker stratification factor comprises a proteome variant, and wherein the biomarker stratification factor score is: (i) Proteomic score of disease phenotype; (ii) proteomic scores of disease risk factors; (iii) proteomic scores of biological pathway activity; (iv) Proteome score of drug target expression: or any linear or non-linear combination thereof, or (c) said biomarker stratification factor comprises a transcription variant, and wherein said biomarker stratification factor score is: (i) Transcriptome score of disease phenotype; (ii) transcriptome scores for disease risk factors; (iii) transcriptome scores of biological pathway activity; or any linear or non-linear combination thereof, or (d) the biomarker stratification factor comprises a somatic mutation variant, and wherein the biomarker stratification factor score is: (i) somatic mutation score of drug target expression; (ii) somatic mutation score of the disease phenotype; (iii) somatic mutation scores for disease risk factors; (iv) somatic mutation scores for biological pathway activity; (v) somatic mutation score of drug target expression; or any linear or non-linear combination thereof, or (e) the biomarker stratification factors include genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) polygenic score of drug target expression; (v) proteomic score of the disease phenotype; (vi) Proteome scores of disease risk factors: (vii) proteomic scores of biological pathway activity; (viii) proteomic scores of biological pathway activity; (ix) proteomic score of drug target expression; (x) a transcriptome score for the disease phenotype; (xi) Transcriptome score of disease risk factors; (xii) transcriptome scores of biological pathway activity; (xiii) somatic mutation score of drug target expression; (xiv) somatic mutation score of the disease phenotype; (xv) disease risk factor somatic mutation score; (xvi) somatic mutation score of biological pathway activity; (xvii) somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
62. 62. The method of any of claims 60-61, further comprising the step of performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of the one or more biomarker stratification factor effects.
63. 63. The method of claim 62, further comprising the step of performing said principal component analysis or said weighted principal component analysis based on a matrix of one or more said biomarker stratification factor effects.
64. 64. The method of claim 63, further comprising constructing the matrix based on one or more of the biomarker stratification factors, the phenotype, and the plurality of values.
65. 65. The method of claims 60-64, further comprising adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
66. 66. The method of claims 62-65, wherein said weighted principal component analysis is performed according to said biomarker stratification factor scores of said selected disease phenotypes.
67. 67. The method of claims 62-66, wherein said weighted principal component analysis identifies predicted drug activities of multiple drug targets against said at least one disease phenotype.
68. 68. The method of any one of claims 60-67, wherein the pharmacomimetic gene score is associated with response to drugs that target GLP1R and / or GIPR.
69. 69. The method of claim 68, wherein the drug comprises a GLP1R and / or GIPR agonist.
70. 70. The method of any one of claims 61-69, wherein said polygenic score is a body mass index polygenic score.
71. 1. An in silico method for determining drug activity of multiple drug targets, comprising: obtaining molecular biomarker stratification factor data and at least one disease phenotype from each of a plurality of subjects; determining a value of the biomarker stratification factor effect for each disease phenotype based on external data; constructing a matrix based on the biomarker stratification factors, the disease phenotypes, and the values; performing a principal component analysis on the biomarker stratification factor x phenotype-associated phenotype biomarker stratification factor effect matrix; calculating a biomarker stratification factor score for each principal component; calculating a pharmacomimetic gene score for each drug target; identifying a predicted drug activity of each drug target for at least one phenotype in the subset of biomarker stratification factor distributions based on a statistical interaction between the polygenic score calculated for each principal component and the pharmacomimetic gene score for each drug target in an association analysis with an outcome phenotype; A method comprising:
72. 1. An in silico method for determining drug activity of a drug target against multiple intermediate phenotypes, comprising: obtaining molecular biomarker stratification factor data and a disease phenotype from each of a plurality of subjects; determining a value of the biomarker stratification factor effect for each intermediate phenotype based on external data; calculating a biomarker stratification factor score for each endophenotype in the disease process; calculating a pharmacomimetic gene score for each drug target for each intermediate phenotype; identifying predicted drug activity of each drug target for one or more phenotypes in the subset of biomarker stratification factor distributions based on a statistical interaction between the biomarker stratification factor scores and the pharmacomimetic gene scores for each drug target in an association analysis with intermediate and outcome phenotypes; A method comprising:
73. 1. A method for determining drug activity of a drug target, comprising: obtaining molecular biomarker stratification factor data and a disease phenotype from each of a plurality of subjects; determining the value of the biomarker stratification factor effect for each phenotype based on external data; calculating a biomarker stratification factor score for a selected outcome phenotype related to a disease outcome; calculating a pharmacomimetic gene score for the drug target; identifying predicted drug activity of the drug target for one or more phenotypes in the subset of biomarker stratification factor distributions based on a statistical interaction between the biomarker stratification factor score and the pharmacomimetic gene score for each drug target in an association analysis with an outcome phenotype; A method comprising:
74. (a) the biomarker stratification factor comprises a genetic variant, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) Polygenic score of drug target expression; or any linear or non-linear combination thereof, or (b) the biomarker stratification factor comprises a proteome variant, and wherein the biomarker stratification factor score is: (i) Proteomic score of disease phenotype; (ii) proteomic scores of disease risk factors; (iii) proteomic scores of biological pathway activity; (iv) Proteome score of drug target expression: or any linear or non-linear combination thereof, or (c) said biomarker stratification factor comprises a transcription variant, and wherein said biomarker stratification factor score is: (i) Transcriptome score of disease phenotype; (ii) transcriptome scores for disease risk factors; (iii) transcriptome scores of biological pathway activity; or any linear or non-linear combination thereof, or (d) the biomarker stratification factor comprises a somatic mutation variant, and wherein the biomarker stratification factor score is: (i) somatic mutation score of drug target expression; (ii) somatic mutation score of the disease phenotype; (iii) somatic mutation scores for disease risk factors; (iv) somatic mutation scores for biological pathway activity; (v) somatic mutation score of drug target expression; or any linear or non-linear combination thereof, or (e) the biomarker stratification factors include genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) polygenic score of drug target expression; (v) proteomic score of the disease phenotype; (vi) Proteome scores of disease risk factors: (vii) proteomic scores of biological pathway activity; (viii) proteomic scores of biological pathway activity; (ix) proteomic score of drug target expression; (x) a transcriptome score for the disease phenotype; (xi) Transcriptome score of disease risk factors; (xii) transcriptome scores of biological pathway activity; (xiii) somatic mutation score of drug target expression; (xiv) somatic mutation score of the disease phenotype; (xv) disease risk factor somatic mutation score; (xvi) somatic mutation score of biological pathway activity; (xvii) somatic mutation score of drug target expression; 74. The method of any one of claims 71-73, comprising one or more scores selected from:
75. 75. The method of any one of claims 71-74, further comprising the step of performing principal component analysis (PCA) or weighted principal component analysis (wPCA) to identify one or more principal components of said one or more biomarker stratification factor effects.
76. 76. The method of claim 75, further comprising the step of performing said principal component analysis or said weighted principal component analysis based on a matrix of one or more said biomarker stratification factor effects.
77. 77. The method of claim 76, further comprising constructing the matrix based on one or more of the biomarker stratification factors, the phenotype, and the plurality of values.
78. 78. The method of claims 71-77, further comprising adjusting the gene score for each drug target and the biomarker stratification factor score for each of the principal components.
79. 79. The method of claims 75-78, wherein said weighted principal component analysis is performed according to said biomarker stratification factor scores of said selected disease phenotypes.
80. 80. The method of claims 75-79, wherein said weighted principal component analysis identifies predicted drug activities of multiple drug targets against said at least one disease phenotype.
81. An in silico system for drug development, comprising: at least one hardware processor; a non-transitory computer-readable storage medium having program code stored thereon; The program code comprises: Obtain molecular biomarker stratification factor data and disease phenotypes from each of a plurality of subjects: determining a value of the biomarker stratification factor effect for each phenotype based on external data; constructing a matrix based on said biomarker stratification factors, said phenotypes and said values; A principal component analysis was performed on the biomarker stratification factor × phenotype-related phenotype biomarker stratification factor effect matrix; Calculate the biomarker stratification factor score for each principal component; Calculate a pharmacomimetic gene score for each drug target; and identifying a predicted drug effect for at least one or more phenotypes in a subset of biomarker stratification factor distributions based on a statistical interaction between the polygenic score calculated for each principal component and the pharmacomimetic gene score for each drug target in an association analysis with outcome phenotypes. a system executable by said at least one hardware processor to:
82. (a) the biomarker stratification factor comprises a genetic variant, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) Polygenic score of drug target expression; or any linear or non-linear combination thereof, or (b) the biomarker stratification factor comprises a proteome variant, and wherein the biomarker stratification factor score is: (i) Proteomic score of disease phenotype; (ii) proteomic scores of disease risk factors; (iii) proteomic scores of biological pathway activity; (iv) Proteome score of drug target expression: or any linear or non-linear combination thereof, or (c) said biomarker stratification factor comprises a transcription variant, and wherein said biomarker stratification factor score is: (i) Transcriptome score of disease phenotype; (ii) transcriptome scores for disease risk factors; (iii) transcriptome scores of biological pathway activity; or any linear or non-linear combination thereof, or (d) the biomarker stratification factor comprises a somatic mutation variant, and wherein the biomarker stratification factor score is: (i) somatic mutation score of drug target expression; (ii) somatic mutation score of the disease phenotype; (iii) somatic mutation scores for disease risk factors; (iv) somatic mutation scores for biological pathway activity; (v) somatic mutation score of drug target expression; or any linear or non-linear combination thereof, or (e) the biomarker stratification factors include genetic variants, proteomic variants, transcriptional variants, and / or somatic mutation variants, and wherein the biomarker stratification factor score is: (i) Polygenic score of disease phenotype; (ii) polygenic scores of disease risk factors; (iii) polygenic scores of biological pathway activities; (iv) polygenic score of drug target expression; (v) proteomic score of the disease phenotype; (vi) Proteome scores of disease risk factors: (vii) proteomic scores of biological pathway activity; (viii) proteomic scores of biological pathway activity; (ix) proteomic score of drug target expression; (x) a transcriptome score for the disease phenotype; (xi) Transcriptome score of disease risk factors; (xii) transcriptome scores of biological pathway activity; (xiii) somatic mutation score of drug target expression; (xiv) somatic mutation score of the disease phenotype; (xv) disease risk factor somatic mutation score; (xvi) somatic mutation score of biological pathway activity; (xvii) somatic mutation score of drug target expression; or any linear or non-linear combination thereof.
83. generating a plurality of principal components (PCs) corresponding to the genetic ancestry data of the subjects in the study cohort; generating a biomarker stratification factor score for each subject in the study cohort based on at least (i) the PCs and (ii) biomarker stratification factor weights for the disease of interest; determining which of a plurality of disease-associated variants is a pharmacomimetic for a disease of interest, wherein a disease-associated variant is a pharmacomimetic if it modulates the function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to predict a second effect of the drug on the phenotype; determining a statistical interaction between said pharmacomimetic measure and said biomarker stratification factor score for one or more drug targets, said statistical interaction predicting drug target-specific differential treatment response for a disease of interest; taking one or more prediction-based actions based on the determined interactions; In silico methods, including:
84. 84. The method of claim 83, wherein the one or more prediction-based actions comprise at least one of: (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics.
85. 84. The method of claim 83, wherein the genetic ancestry data indicates whether subjects in a study cohort are cases or controls for a disease of interest.
86. 86. The method of claim 85, wherein the genetic ancestry data is based on genotyping arrays or whole genome sequencing.
87. 86. The method of claim 85, wherein the genetic ancestry data is obtained from a publicly available database.
88. 84. The method of claim 83, wherein the plurality of PCs comprises at least five PCs.
89. 84. The method of claim 83, wherein the study cohort comprises at least 200 control subjects.
90. 84. The method of Claim 83, wherein each disease-associated variant in said plurality of disease-associated variants meets a genome-wide significance threshold.
91. 84. The method of claim 83, wherein the plurality of disease-associated variants is determined based on a genome-wide association study (GWAS) of the first disease.
92. 84. The method of claim 83, wherein the biomarker stratification factor score for each subject is a scaled biomarker stratification factor score.
93. 93. The method of claim 92, wherein the scaled biomarker stratification factor is based on at least raw PGS and ancestry-normalized PGS.
94. 84. The method of claim 83, wherein the PGS variant weights are calculated independently from the genetic ancestry data corresponding to a study cohort.
95. 84. The method of Claim 83, wherein determining which of the plurality of disease-associated variants are the pharmacomimetics for the disease of interest comprises generating an allele score.
96. 96. The method of claim 95, wherein the variants in the allele score are weighted using a second GWAS that did not include any of the subjects in the study cohort.
97. 96. The method of claim 95, wherein the variants in the allele score are weighted using at least a third GWAS for biomarkers that causally mediate the effect of one or more drug targets on a disease of interest, wherein the third GWAS is independent of the study cohort.
98. 84. The method of claim 83, wherein determining statistical interactions comprises fitting a logistic regression model to each pharmacomimetic measure.
99. 84. The method of claim 83, further comprising generating one or more tables comprising grouping subjects by a set of percentiles of the PGS and fitting a separate logistic regression model for each percentile in the set to show the effect of each pharmacomimetic measure on disease risk.
100. 84. The method of claim 83, wherein taking one or more prediction-based actions comprises providing a PGS threshold used to select subjects for a prospective clinical trial.
101. 84. The method of claim 83, wherein taking one or more prediction-based actions comprises identifying therapeutic mechanisms that demonstrate benefit in a subset of patients with a common disease who are at high polygenic risk.
102. 84. The method of claim 83, wherein taking one or more prediction-based actions comprises selecting patients for treatment based on their likelihood of benefiting from the treatment.