Drug development and drug activity determination for chronic liver disease, NASH, adiposity, and diabetes

The use of polygenic scores and computational models simulates drug development for CLD, adiposity, and diabetes, addressing inefficiencies in current processes by predicting drug activity and optimizing therapeutic use in specific patient populations, thus accelerating drug development and reducing costs.

WO2025240664A1PCT designated stage Publication Date: 2025-11-20FORESITE LABS LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/029437
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-15
Filing Date
2025-05-14
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Current drug development processes for chronic liver disease (CLD), adiposity, and diabetes are inefficient and require extensive empirical data, leading to high costs and time consumption, especially in addressing the heterogeneity of patient responses and disease progression.

Method used

Utilizing polygenic scores (PGS) and computational models to simulate drug development and clinical trials, allowing for the prediction of drug activity and patient response without the need for extensive empirical data, by integrating genetic and phenotypic data to identify drug targets and optimize therapeutic use.

Benefits of technology

Accelerates drug development and repurposing of existing therapies by predicting drug efficacy and safety in specific patient populations, reducing the need for costly and time-consuming clinical trials and regulatory approvals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025029437_20112025_PF_FP_ABST
    Figure US2025029437_20112025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems, methods and computer program products for drug development for chronic liver disease (CLD) or non-alcoholic steatohepatitis (NASH) using in silico techniques. An aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets by determining biomarker stratifier effects, calculating a biomarker stratifier score for a chosen disease phenotype; and calculating a pharmacomimetic genetic score using molecular biomarker stratifier data.
Need to check novelty before this filing date? Find Prior Art

Description

DRUG DEVELOPMENT AND DRUG ACTIVITY DETERMINATION FOR CHRONIC LIVER DISEASE, NASH, ADIPOSITY, AND DIABETESCROSS-REFERENCE TO RELATED PATENT APPLICATION

[0001] This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No. 63 / 648,109, filed May 15, 2024, the entire contents of which are incorporated herein by reference.FIELD OF THE DISCLOSURE

[0002] The present disclosure relates to systems, methods and computer program products for drug development for chronic liver disease (CLD), adiposity, diabetes, or non-alcoholic steatohepatitis (NASH) using in silico techniques.BACKGROUND OF THE DISCLOSURE

[0003] Modelling and simulation are rapidly evolving areas in terms of both technologies and application fields in the life sciences. The use of modeling and simulation has expanded beyond the description of drug exposure, towards the dynamic description of complex drug effects and disease subtypes and progressions. In recent years, new approaches in modelling and simulation have started to provide important insights in biomedicine, opening the way for their potential use in the reduction, refinement and partial substitution of both animal and human experimentation.

[0004] There is a need for determining how a patient will respond to a given or potential chronic liver disease (CLD), adiposity, diabetes, or non-alcoholic steatohepatitis (NASH) drug. This disclosure satisfies this need and provides related advantages as well.SUMMARY OF ILLUSTRATIVE EMBODIMENTS

[0005] Polygenic scores (PGS), herein used interchangeably with polygenic risk scores (PRS), summarize genome-wide genotype data in combination with gene-phenotype association data, such as from genome-wide association studies (GWAS), into metrics that represent genetic liability to a trait. The present disclosure provides in silico systems, methods and computer program products using PRS for various aspects of drug development and optimization of patient treatment, including drug indication selection and clinical trial recapitulation, and in several aspects, for CLD adiposity, diabetes, or NASH drug.Specifically, the disclosure provides PRS-informed drug development systems to acceleratethe development of new therapeutic products as well as repurposing existing therapies for new populations and / or indications. There are many applications of PRS-informed drug development, ranging from discovery of new drug targets for the treatment of a given indication, to repurposing of existing therapeutics for new indications, to new and efficient designs of clinical trials, to incorporation of new end points in ongoing clinical trials, to modifying how the data are analyzed to provide evidence of effectiveness in powering clinical trials, to modifying how existing therapeutics are used in patient populations. These systems utilize existing patient variant data in conjunction with computational models to mimic aspects of the drug development process. Importantly, the systems of the disclosed approach allow such development without the requirement of empirical data that has to date been required for various steps in the drug development process, from drug discovery through clinical trials and regulatory approval, although empirical data is optionally incorporated in certain embodiments. In some embodiments, the disease or indication is chronic liver disease (CLD), adiposity, diabetes, or non-alcoholic steatohepatitis (NASH).

[0006] In addition to the patient variant data, the computational models of the systems of the disclosure can optionally use mechanistic knowledge of physical, chemical, and / or environmental data related to patient populations and available biological and physiological knowledge of patient treatment (such as medical intervention) or activity (such as exercise or lack thereof) for the treatment of CLD, adiposity, diabetes, adiposity, diabetes, or NASH.

[0007] In some aspects, the disclosure provides a preclinical drug discovery system for identifying and / or developing treatments (e.g., drugs) for various therapeutic areas and indications to treat CLD, adiposity, diabetes, or NASH. The preclinical drug discovery system is configured to utilize mammalian polygenic scores with computational models to provide PRS-informed drug discovery methodologies.

[0008] In some aspects, the disclosure provides a drug discovery system for repurposing existing therapeutics for new therapeutic areas for the treatment of CLD, adiposity, diabetes, or NASH, and associated indications. The system is configured to utilize mammalian polygenic scores with computational models to provide PRS-informed methodologies with known drug and patient information to find potential new uses of existing therapies.

[0009] In some embodiments, the disclosure provides in silico systems and methods for determining drug activity for the treatment of CLD adiposity, diabetes, or NASH. In particular, the disclosure provides systems and methods for determining drug activity forparticular patient populations, phenotypes, or for use in therapeutic indications / areas for the treatment of CLD adiposity, diabetes, or NASH. This drug activity may be determined for known agents or for de novo agents.

[0010] Accordingly, in some aspects, the disclosure provides an in silico method for determining drug activity of drug targets on a plurality of phenotype intermediates, comprising, or alternatively consisting essentially of, or yet further consisting of obtaining data on genetic variants and phenotypes from each of a plurality of subjects, determining the value of the variant effects for each phenotype intermediate based on external data, calculating a polygenic score for each phenotype intermediate in the disease process, calculating a pharmacomimetic genetic score for each drug target on each phenotype intermediate, and identifying predicted drug activity of the drug targets on one or more phenotypes based on the statistical interaction of the polygenic score with the pharmacomimetic genetic score for each drug target in association analysis with the phenotype intermediate and the outcome phenotype for the treatment of CLD, adiposity, diabetes, or NASH. In some aspects, the disclosure provides a system for providing in silico clinical trials for new therapeutics using mammalian polygenic scores with computational models for the treatment of CLD, adiposity, diabetes, or NASH. Historically, safety and efficacy data provided to regulatory agencies in support of marketing authorization of a new medical product have been produced experimentally, either in vitro or in vivo. More recently, regulatory agencies started receiving and accepting in silico data produced using modelling and simulation, and the data provided by the systems of the disclosure can provide data supporting regulatory filings without the need for expensive and time-consuming experimentation.

[0011] The systems of the disclosure allow development of patient-specific models to form virtual cohorts for testing the safety and / or efficacy of new drugs and of new medical devices.

[0012] In certain embodiments, the disclosure provides in silico methods for determining drug activity, the methods comprising, or alternatively consisting essentially of, or yet further consisting of obtaining data on genetic variants and phenotypes from each of a plurality of subjects; determining the value of each variants’ effect on each phenotype based on external data from a non-overlapping set of subjects; constructing a matrix in which the variants are rows, the phenotypes are columns, and the values are the externally-derived variant effects; running a truncated principal components analysis on this matrix, computing a user-definednumber of principal components, typically two to five; calculating a polygenic score for each computed principal component by reweighting an existing polygenic score for a user-defined “reference” phenotype such that the new weight for each variant included the polygenic score is equal to its old weight times its squared loading for the principal component in question divided by the sum of its squared loadings across all computed principal components; and lastly identifying predicted drug activity for the “reference” phenotype based on each of the reweighted polygenic scores, which often can be interpreted as corresponding to biological pathways. This example method is set forth visually in Figure 1.

[0013] In certain embodiments, the disclosure provides in silico systems for drug development for the treatment of CLD adiposity, diabetes, or NASH, such systems comprising, or alternatively consisting essentially of, or yet further consisting of at least one hardware processor and a non-transitory computer-readable storage medium having stored thereon program code, the program code executable by the at least one hardware processor to obtain data on genetic variants and phenotypes from each of a plurality of subjects, determine the value of the variant effects for each phenotype based on external data, construct a matrix based on the variants, phenotypes and values, run a principal components analysis on the variant x phenotype-related phenotype variant effects matrix, calculate a polygenic score for each principal component, and determine predicted drug effects on one or more phenotype based on the polygenic score calculated for each principal component.

[0014] In certain embodiments, the disclosure provides in silico systems for determination of intermediate phenotypes, such as risk factors or genetic variations that identify responders versus non-responders for a specific therapeutic intervention. Accordingly, in certain aspects, the disclosure provides in silico systems and methods for determining drug activity comprising, or alternatively consisting essentially of, or yet further consisting of obtaining data on genetic variants and phenotypes from each of a plurality of subjects, determining the value of the variant effects for each phenotype based on external data, calculating a polygenic score for a chosen phenotype intermediate in the disease process (e.g., a risk factor, biomarker, or genetic indicator of response), calculating a pharmacomimetic genetic score for each drug target; and identifying predicted drug activity on one or more phenotypes based on the statistical interaction of the polygenic score with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0015] In specific aspects, the disclosure provides methods in which the PRS information is provided in whole or in part from related disease phenotypes.

[0016] Accordingly, the disclosure provides a method of determining the drug activity of a drug target for the treatment of CLD adiposity, diabetes, or NASH, comprising, or alternatively consisting essentially of, or yet further consisting of obtaining data on genetic variants and phenotypes from each of a plurality of subjects, determining the value of the variant effects for each phenotype based on external data, calculating a polygenic score for a chosen outcome phenotype related to the disease outcome, calculating a pharmacomimetic genetic score for the drug target, and identifying predicted drug activity of the drug target on one or more phenotypes based on the statistical interaction of the polygenic score with the pharmacomimetic genetic score for the drug target in association analysis with the outcome phenotype. The drug target may be, e.g., a known drug for the phenotype, a drug that has been approved or in trials for a different indication than the phenotype for which the analysis is being performed, or a de novo agent.

[0017] An aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets for the treatment of CLD adiposity, diabetes, or NASH, the method comprising, or alternatively consisting essentially of, or yet further consisting of obtaining molecular biomarker stratifier data comprising, or alternatively consisting essentially of, or yet further consisting of a plurality of biomarker stratifiers and at least one disease phenotype from each of a plurality of subjects; determining a plurality of values representing biomarker stratifier effects, wherein each value separately represents how each of the plurality of biomarker stratifiers affects each of the at least one disease phenotype based on external data; calculating a biomarker stratifier score for a chosen disease phenotype; calculating a pharmacomimetic genetic score for each drug target; and identifying the predicted drug activity of each drug target the at least one disease phenotype in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the disease phenotype.

[0018] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for adisease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0019] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0020] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0021] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity;(v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0022] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype;(vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomicsscore for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0023] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

[0024] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0025] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and the plurality of values.

[0026] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0027] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0028] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of the weighted principal component analysis identifies predicted drug activity of the plurality of drug targets for the at least one disease phenotype.

[0029] Another aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets, the method comprising: obtaining molecular biomarker stratifier data and at least one disease phenotype from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each disease phenotype based on external data;constructing a matrix based on the biomarker stratifiers, disease phenotypes and values; running a principal components analysis biomarker stratifier x phenotype-related phenotype biomarker stratifier effects matrix; calculating a biomarker stratifier score for each principal component; calculating a pharmacomimetic genetic score for each drug target; and identifying predicted drug activity of each drug target on one or more based on the statistical interaction of the polygenic score calculated for each principal component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0030] In some embodiments, the drug targets comprise, consist essentially of, or consist of PNPLA3 and / or HSD 17B 13.

[0031] In some embodiments, the biomarker stratifier data comprises, consists essentially of, or consists of data on one or more of PNPLA3 genotype, HSD17B13 genotype, body mass index (BMI), BMI-adjusted waist to hip ratio, and diabetes status of the patient.

[0032] Another aspect of the disclosure is directed to an in silico method for determining the drug activity of drug targets on a plurality of phenotype intermediates, comprising: obtaining molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each phenotype intermediate based on external data; calculating a biomarker stratifier score for each phenotype intermediate in the disease process; calculating a pharmacomimetic genetic score for each drug target on each phenotype intermediate; and identifying predicted drug activity of the drug targets on one or more phenotypes based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the phenotype intermediate and the outcome phenotype.

[0033] Another aspect of the disclosure is directed to method of determining the drug activity of a drug target, comprising: obtaining molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each phenotype based on external data; calculating a biomarker stratifier score for a chosen outcome phenotype related to the disease outcome; calculating a pharmacomimetic genetic score for the drug target; and identifying predicted drug activity of the drug target on one or more phenotypes based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for the drug target in association analysis with the outcome phenotype.

[0034] Another aspect of the disclosure is directed to an in silico method comprising: generating a plurality of principal components (PCs) corresponding to genetic ancestry data for subjects in a study cohort; generating a biomarker stratifier score for each subject in the study cohort based at least on (i) the PCs and on (ii) biomarker stratifier weights for a disease of interest; determining which of a plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest, where a disease-associated variant is a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype; determining statistical interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores, wherein the statistical interactions are predictive of drug target-specific differential treatment response for the disease of interest; and performing one or more prediction-based actions based on the determined interactions.

[0035] In some embodiments, the drug targets comprise, consist essentially of, or consist of PNPLA3 and / or HSD 17B 13.

[0036] In some embodiments, the biomarker stratifier data comprises, consists essentially of, or consists of data on one or more of PNPLA3 genotype, HSD17B13 genotype, body mass index (BMI), BMI-adjusted waist to hip ratio, and diabetes status of the patients.

[0037] In some embodiments, the one or more prediction-based actions comprises at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics.

[0038] In some embodiments, the genetic ancestry data indicates whether subjects in the study cohort are a case or a control for the disease of interest.

[0039] In some embodiments, the genetic ancestry data is based on genotyping arrays or whole-genome sequencing.

[0040] In some embodiments, the genetic ancestry data is obtained from a publicly available database.

[0041] In some embodiments, the plurality of PCs comprises at least 5 PCs.

[0042] In some embodiments, the study cohort comprises at least 200 control subjects.

[0043] In some embodiments, each disease-associated variant in the plurality of disease- associated variants satisfies a genome-wide significance threshold.

[0044] In some embodiments, the plurality of disease-associated variants is determined based on a first disease genome-wide association study (GWAS).

[0045] In some embodiments, the biomarker stratifier score for each subject is a scaled biomarker stratifier score.

[0046] In some embodiments, the scaled biomarker stratifier score is based at least on a raw PGS and an ancestry-normalized PGS.

[0047] In some embodiments, the PGS variant weights are computed independent of the genetic ancestry data corresponding to the study cohort.

[0048] In some embodiments, the PGS variant weights are computed using a method for ancestry normalization that estimates the joint likelihood of the mean and variance of the PRS scores as a function of ancestry PCs, and calculates an adjustment score for every individual in the dataset.

[0049] In some embodiments, determining which of the plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest comprises generating an allelic score.

[0050] Another aspect of the disclosure is directed to an in silico system for drug development, such system comprising: at least one hardware processor; and a non-transitory computer-readable storage medium having stored thereon program code, the program code executable by the at least one hardware processor to: obtain molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determine the value of the biomarker stratifier effects for each phenotype based on external data; construct a matrix based on the biomarker stratifiers, phenotypes and values; run a principal components analysis biomarker stratifier x phenotype-related phenotype biomarker stratifier effects matrix; calculate a biomarker stratifier score for each principal component; calculate a pharmacomimetic genetic score for each drug target; determine predicted drug effects on one or more phenotypes based on the statistical interaction of the polygenic score calculated for each principal component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0051] An aspect of the disclosure is directed to a method for selecting a patient for an HSD17B13 inhibitor trial comprising, or consisting essentially of, or consisting of determining the genotype for HSD17B13 and PNPLA3 genes in the patient; wherein the patient having two functional HSD17B13 alleles and at least one PNPLA3 148M allele is selected for the trial. In some embodiments, the patient suffers from CLD or NASH. In a further aspect, the HSD17B13 inhibitor is administered to the patient.

[0052] Another aspect of the disclosure is directed to a method for selecting a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH) patient for an HSD17B13 inhibitor therapy comprising, or consisting essentially of, or consisting of determining the genotype for HSD17B13 and PNPLA3 genes in the patient; and selecting the patient having two functionalHSD17B13 alleles and at least one PNPLA3 148M allele. In a further aspect, the HSD17B13 inhibitor is administered to the patient.

[0053] Another aspect of the disclosure is directed to a method for treating a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH), the method comprising, or consisting essentially of, or yet further consisting of determining the genotype for HSD17B 13 and PNPLA3 genes in the patient; and treating the patient having two functional HSD17B13 alleles and at least one PNPLA3 148M allele with an HSD17B13 inhibitor.

[0054] An aspect of the disclosure is directed to a method for determining whether a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH) patient is at risk of progressing to cirrhosis or hepatocellular carcinoma comprising, or consisting essentially of, or consisting of: (i) sequencing the PNPLA3 gene in the patient; and (ii) determining that the patient is at risk of progressing to cirrhosis or hepatocellular carcinoma when a PNPLA3 I148M mutation is detected in the patient.

[0055] In some embodiments, the patient is heterozygous for the PNPLA3 I148M mutation. In some embodiments, the patient is homozygous for the PNPLA3 I148M mutation.

[0056] In some embodiments, the method further comprises, or consists essentially of, or consists of determining a higher risk of progressing to cirrhosis or hepatocellular carcinoma if the patient is diabetic. In some embodiments, the method further comprises, or consists essentially of, or consists of determining a higher risk of progressing to cirrhosis or hepatocellular carcinoma when the patient has a body mass index (BMI) over 30.

[0057] In some embodiments, the method further comprises, or consists essentially of, or consists of administering to the patient an HSD17B13 inhibitor who is at higher risk for disease progression.

[0058] An aspect of the disclosure is directed to a method for determining whether a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH) patient will benefit from a PNPLA3 degrader, the method comprising, or consisting essentially of, or consisting of: (i) sequencing the PNPLA3 gene in a sample isolated from the patient; and (ii) determining that the patient will benefit from a PNPLA3 degrader when a PNPLA3 I148M mutation is detected in the patient sample.

[0059] In some embodiments, the patient is heterozygous for the PNPLA3 I148M mutation. In some embodiments, the patient is homozygous for the PNPLA3 I148M mutation.

[0060] In some embodiments, the method further comprises, or consists essentially of, or consists of determining a higher benefit from a PNPLA3 degrader if the patient is diabetic. In some embodiments, the method further comprises, or consists essentially of, or consists of determining the patient’s BMI and determining a higher benefit from a PNPLA3 degrader if the patient has a body mass index (BMI) over 30.

[0061] In some embodiments, the method further comprises, or consists essentially of, or consists of administering to the patient determined to receive a greater benefit a PNPLA3 degrader.

[0062] An aspect of the disclosure is directed to a method for determining whether a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH) patient will benefit from an HSD17B13 inhibitor, the method comprising, or consisting essentially of, or consisting of generating a combined score by combining the body mass index (BMI) and BMI-adjusted waist-to-hip ratio of the patient; and determining that the patient will benefit from the HSD17B13 inhibitor when the combined score is above a predetermined threshold.

[0063] In some embodiments, the method further comprises, or consists essentially of, or consists of administering to the patient determined to benefit from the therapy an HSD17B13 inhibitor.

[0064] These and other embodiments, features, and advantages will be set forth in the present disclosure.BRIEF DESCRIPTION OF THE DRAWINGS

[0065] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate one or more embodiments and, together with the description, explain these embodiments. The accompanying drawings have not necessarily been drawn to scale. Any values or dimensions illustrated in the accompanying graphs and figures are for illustration purposes only and may or may not represent actual or preferred values or dimensions. Where applicable, some or all features may not be illustrated to assist in the description of underlying features.

[0066] Figure 1 is a flow diagram showing the method of one embodiment of the disclosure.

[0067] Figure 2 illustrates a block diagram of a system for in silico drug development and determination of drug activity in one embodiment of the disclosure.

[0068] Figure 3 illustrates a flow diagram showing a method for in silico drug development and determination of drug activity in one embodiment of the disclosure.

[0069] Figure 4 illustrates a flow diagram showing a method for in silico drug development and determination of drug activity in one embodiment of the disclosure.

[0070] Figure 5 illustrates a flow diagram showing a method for in silico drug development and determination of drug activity in one embodiment of the disclosure.

[0071] Figure 6 illustrates a sequence for in silico drug development and determination of drug activity in one embodiment of the disclosure.

[0072] Figures 7A-7B illustrate the effect of HSD17B13 loss of function on liver damage from I148M plotted for (A) log ALT, (B) winsorized ALT.

[0073] Figure 8 illustrates the distribution of cTl values.

[0074] Figures 9A-9B illustrate the effect of diabetes on effects of PNPLA3 and HSD17B13 genetic variation, plotted for (A) log ALT, (B) winsorized ALT.

[0075] Figures 10A-10C illustrate (A) the effect of PNPLA3 I148M on CLD progression to cirrhosis; (B) the predictive power of 3-SNP score made from PNPLA3, TM6SF2 and HSD17B13; and (C) the effect of PNPLA3 I148M on CLD progression in subjects with prior diabetes.

[0076] Figure 11 illustrates the cutoffs for medium and high-risk groups using the 3-SNP score made from PNPLA3, TM6SF2 and HSD17B13.

[0077] Figure 12 illustrates CLD progression risk for 3 bins of patients as defined using a 3- SNP score made from PNPLA3, TM6SF2 and HSD17B13.

[0078] Figure 13 illustrates the marginal effects of HSD17B13 and PNPLA3 on ALT in NVH cases, providing justification for ALT winsorization.

[0079] Figure 14 illustrates the log-transformed ALT in NVH cases.

[0080] Figure 15 illustrates the winsorized ALT in NVH cases.

[0081] Figures 16A and 16B illustrate the effect of HSD17B13 loss of function on liver damage in patients with the PNPLA3 genotype. In Figure 16A, log(ALT) was used as the disease outcome and in Figure 16B winsorized ALT was used as the disease outcome.

[0082] Figure 17 illustrates the distribution of cTl values in the UK Biobank cohort.DETAILED DESCRIPTION

[0083] The following detailed description of preferred embodiments of the disclosure will be better understood when read in conjunction with the appended drawings.

[0084] All publications, including patent documents, scientific articles and databases, referred to in this application are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication were individually incorporated by reference. If a definition set forth herein is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein incorporated by reference, the definition set forth herein prevails over the definition that is incorporated herein by reference.

[0085] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0086] HSD17B13 is an enzyme of unknown biological function that also localizes to lipid droplets. A splice variant in HSD17B13 (rs72613567) causes loss of enzymatic function (LoF), reduces liver injury biomarkers, reduces the risk of progression from hepatic steatosis to steatohepatitis (but does not affect the incidence of steatosis), and reduces risk of all-cause cirrhosis (Abul-Husn, Noura S., et al. New England Journal of Medicine 378.12 (2018): 1096-1106; Emdin, Connor A., et al. Gastroenterology 160.5 (2021): 1620-1633). The protective effect of HSD17B13 LoF is greatest in carriers of PNPLA3 I148M (Abul-Husn, Noura S., et al. New England Journal of Medicine 378.12 (2018)), one of the few robust gene-gene interactions that have been described to date in the statistical genetics literature. Several biotech companies are developing HSD17B13 inhibitors, with some positive early clinical data (Arrowhead Pharmaceuticals, Inc., 23 Jun. 2021, ir.arrowheadpharma.com / news-releases / news-release-details / arrowhead-presents-positive- interim-clinical-data-aro-hsd. Press release; Alnylam Pharmaceuticals, Inc., 03 Aug. 2020, investors. alnylam.com / press-release?id=25046. Press release; Inipharm, Inc., 05 Oct. 2021, businesswire.com / news / home / 20211005005420 / en / Inipharm-to-Present-Data-Showing- Potential-of-Small-Molecule-Inhibitors-of-HSD17B13-to-Combat-Liver-Fibrosis-at- AASLD%E2%80%99s-The-Liver-Meeting. Press release).

[0087] PNPLA3 is an enzyme that localizes to lipid droplets and is involved in lipid metabolism. A missense variant in PNPLA3, 1148M, is associated with substantially increased liver fat (steatosis), liver injury biomarkers, and risk of non-alcoholicsteatohepatitis (NASH), alcoholic liver disease, hepatocellular carcinoma, and all-cause cirrhosis (Romeo, Stefano, et al. Nature genetics 40.12 (2008): 1461-1465; Sookoian, Silvia, et al. Journal of lipid research 50.10 (2009): 2111-2116; Tian, Chao, et al. Nature genetics42.1 (2010): 21-23; Speliotes, Elizabeth K., et al. Hepatology 52.3 (2010): 904-912;Sookoian, Silvia, and Carlos J. Pirola. Hepatology 53.6 (2011): 1883-1894; Chambers, John C., et al. Nature genetics 43.11 (2011): 1131-1138; Zain, Shamsul Mohd, et al. Human genetics 131 (2012): 1145-1152; Valenti, Luca, et al. Digestive and Liver Disease 45.8 (2013): 619-624; Singal, Amit G., et al. Official journal of the American College of Gastroenterology | ACG 109.3 (2014): 325-334; Abul-Husn, Noura S., et al. New England Journal of Medicine 378.12 (2018): 1096-1106; Emdin, Connor A., et al. Gastroenterology 160.5 (2021): 1620-1633). It has been shown that, in mice, the effect of PNPLA3 I148M on hepatic steatosis is driven by evasion of ubiquitin-mediated protein degradation and consequent accumulation of mutant PNPLA3 protein on lipid droplets, with some evidence to suggest that PNPLA3 accumulation interferes with the interaction of adipose triglyceride lipase (ATGL = PNPLA2 ) with its cofactor, CGL58 (= ABHD5 ) (Li, John Zhong, et al. The Journal of clinical investigation 122.11 (2012): 4130-4144; Smagris, Eriks, et al. Hepatology61.1 (2015): 108-118; BasuRay, Soumik, et al. Hepatology 66.4 (2017): 1111-1124; Tsang, Felice Ho-Ching, et al. Hepatology 69.6 (2019): 2502-2517; BasuRay, Soumik, et al.Proceedings of the National Academy of Sciences 116.19 (2019): 9521-9526).

[0088] In some aspects, the present disclosure describes in silico systems for recapitulation of empirical clinical trial data. Disclosure is provided herein demonstrating that genetic analysis using pharmacomimetic variants can recapitulate drug-PRS interactions observed in retrospective analyses of clinical trials. It also demonstrates that selectively enrolling patients with high PRS for clinical trials can increase average treatment response by enrolled patients as shown in the comparison of the empirical evidence and the results obtained by the systems of the disclosure.

[0089] Over the last decade, genome-wide association studies (GWAS) have uncovered the contribution of inherited variants to common complex disorders. Many non-communicable disorders with a major public health impact have a genetic underpinning that is highly polygenic, comprising, or alternatively consisting essentially of, or yet further consisting of hundreds or thousands of genetic variants (or polymorphisms), each having a small effect on disease risk. Each genetic variant associated with a disease is valuable in indicating a gene orpathway of biological relevance to the disorder, but there are also expectations that the genetic data could be used to predict disease risk, with potential clinical utility.

[0090] For example, there have been substantial successes in discovering and developing new health interventions, including therapeutics, for common chronic human diseases, here defined as diseases that have a prevalence in the general population of >0.1%. Despite these advances in management of common chronic diseases, there remains substantial unmet need in clinical areas that contribute significantly to the population burden of disease morbidity, premature mortality, and associated societal and healthcare costs. It is estimated that 6 in 10 Americans have at least one common chronic disease, and that these diseases collectively are responsible for ~$2.7T in US healthcare costs. Owing to substantial heterogeneity in disease progression and treatment response, drug development in these areas has required large clinical trials, at substantial clinical development cost, to evaluate clinical efficacy. Thus, despite substantial market opportunity and clinical unmet need, drug development has shifted over the last two decades to oncology and rare disease, where molecularly defined mechanisms and subpopulations have converged to enable a higher likelihood of demonstrating efficacy and safety sufficient for regulatory approval of new chemical entities.

[0091] Large scale genome-wide association studies performed over the last several decades have identified genetic contributions to variation in common disease risk, as well as the underlying risk factors that play a causal role in their pathobiology. These studies have found that the architecture of disease risk is genetically complex, reflecting the contributions of tens to millions of individual common genetic variants, each of small effect, as well as a smaller number of low frequency and rare genetic variants of greater effect. Collectively, these risk factors explain between -40% and -80% of the variation in disease risk observable at a population level. With larger studies of the effects of these genetic variants on disease risk has come an ability to more precisely estimate the contributions to disease risk of genetic variants at a range of effect sizes. This has facilitated the scaled summation of these effects into continuous polygenic genetic scores (PGS) that estimate, for any individual, the aggregate of measurable genetic risk. PGS may be isotropic, reflecting the sum total of genetic effects on disease risk agnostic to effects on underlying risk factors or biological pathways, or may reflect fundamental driver pathways in subsets of common chronic disease or in individuals with specific risk factors. Both isotropic as well as pathway- or risk-factor- specific PGS may resolve common disease heterogeneity and identify subgroups of the population who respond exceedingly well to certain therapeutic mechanisms. Indeed, there isclinical trial precedent for patients with high PGS receiving greater benefit from a drug than patients with low PGS.

[0092] A system implementing the systems and methods described herein may perform a method of analysis and use a knowledgebase for identifying individuals with subsets of common chronic human diseases who are predicted to enjoy greater benefit from specific therapeutic mechanisms. Specifically, the method identifies combinations of drug targets and disease indications where it is predicted that patients with a high PGS for the disease, its associated risk factors, or an aggregate of several genetically-driven biological pathways, will receive greater benefit from a drug compared to patients with a low PGS.

[0093] To perform the method, a computer may access genetic and phenotypic data for a large group of people (e.g., the “study cohort”). The genetic data can originate from genotyping arrays or whole-genome sequencing. The phenotype data does not have to come in a specific format. But, in some embodiments, phenotype data must be sufficient to determine whether each subject is a case, a control, or neither for the disease of interest. The algorithm for assigning case-control status to subjects may be manually defined by a user. The algorithm may consider, for example: (i) ICD-10 or ICD-9 diagnosis codes from hospital records or primary care records; (ii) self-reported diagnoses; (iii) medication prescriptions; and / or (iv) OPCS-4 or OPCS-3 operation codes from hospital records. The computer may also use the age and sex of each subject. Examples of suitable study cohorts can include the UK Biobank(pan.ukbb. broadinstitute.org, last accessed May 13, 2025), the FinnGen Research Project (world wide web fmngen.fi / en, last accessed May 13, 2025), and the All of Us Research Program (allofus.nih.gov, last accessed May 13, 2025). There is no exact requirement for the number of subjects in the study cohort, but typically the computer may use at least ~5k cases and ~5k controls for a particular disease of interest.

[0094] The computer may also access PGS variant weights for the disease of interest. “PGS variant weights” can refer to a table of genetic variants in which each genetic variant is assigned a weight (e.g., a numerical value). The weight can be an estimate of how much a given variant contributes to the risk of a disease. The computer can compute PGS variant weights using publicly available programs such as PRS-CS (Ge et al. (2019) Polygenic prediction via Bayesian regression and continuous shrinkage priors. Nat Commun 10: 1776) or LDPred2 (Prive et al. (2020) LDpred2: better, faster, stronger, Bioinfomiatics, Volume 36, Issue 22-23 : 5424 -5431 )). In some cases, PGS variant weights can be published byacademic researchers. The computer can download such weights from a web resource called the PGS Catalog, for example. Computing the PGS variant weights may not involve any data from the study cohort and must be derived from a completely independent dataset, in some embodiments.

[0095] The computer may access results from a genome-wide association study (GWAS) for the disease of interest, commonly referred to as “summary statistics.” To do so, the computer may retrieve GWAS summary statistics from a web resource, such as the EBI GWAS Catalog, for example. Alternatively, the data processing system may perform the GWAS within the study cohort, such as by using publicly available software such as Plink (Purcell et al. (2007) Am J Hum Genet. Sep;81(3):559-75), Chang et al. (2015) Gigascience Feb 25;4:7), SAIGE (Zhou et al. (2022)) Nat Genet 54: 1466- 1469, or REGENIE (Mbatchou et al.(2021) Nat Genet 53:1097-1103).

[0096] The computer can access a computational pipeline for mapping GWAS variants to causal genes, such as the pipelines published by Mountjoy et al. (2021) Nat Genet. Vol. 53(11): 1527-1533. The drug-by-PGS interaction discovery method is agnostic to the particular pipeline that is used.

[0097] The method can apply the following operations of a framework: (1) modeling the predicted effects of drug target modulation from genotype-phenotype association analysis of human genetic variants in drug target-encoding genes and observed clinical phenotypes. The genetic variants used for modeling may be individual variants or sets of statistically- independent variants in the same drug target gene identified as “allelic scores.” These statistical instruments are “pharmacomimetic instruments;” (2) partitioning a human population according to isotropic PGS, risk factor PGS, or biological pathway-PGS; and (3) application of a method for identifying statistical interactions between pharmacomimetic instruments for drug targets and PGS that predict drug target-specific differential treatment response. In some embodiments, the method can include using the predicted drug targetspecific differential treatment response to select patients for treatment and / or applying the treatment. In some embodiments, the method can include transmitting the predicted drug target-specific differential treatment response and / or any data used to determine the predicted drug target-specific differential treatment response (e.g., the statistical interactions, the determined PGS, etc.) to another computing system (e.g., into a downstream pipeline) for an entity associated with the computing system to use for treatment.

[0098] In performing the aforementioned method, a computing system can use human genetic data to predict a-priori whether a given drug mechanism will exhibit PGS-stratified efficacy without needing to have clinical trial data. For example, the computing system can use the predictive ability of the method by applying the method to modeling the effects of the anti-PCSK9 mechanism on the risk of cardiovascular disease according to PGS, accurately recapitulating the results of clinical trial data. The computing system can use the method on at least 35 common chronic diseases to identify and / or implement novel indications for drug targets in which PGS stratification is likely to yield differential treatment response. The computing system can use the method to identify or use PGS-stratified efficacy LDL- lowering therapies (e.g., statins and PCSK9 antibodies) in cardiovascular disease as well as drugs for a wide variety of indications.

[0099] One attempt for modelling and simulation involves implementing machine learning techniques, such as by using a linear regression model or a neural network to generate the simulations. However, such machine learning models have inherent technical problems that limit their capabilities when processing large amounts of data. For instance, neural networks are limited by the number of nodes that are included at the input layer and linear regression models are limited by the number of axes they may have. These issues with machine learning models may cause problems with biomedical simulations, because biomedical simulations can invol ve mi llions of datapoints of varying types. The large amount of data points can make it difficult to use machine learning models for the simulations or otherwise cause such simulations to be impractical or impossible given the large amount of computing resources they would require (e.g., a neural network with a large amount of input nodes or a linear regression model with a large number of axes can be computationally expensive to execute, particularly for a large number of datapoints).

[0100] A computer implementing the systems and methods described herein can overcome these technical deficiencies of machine learning models. For instance, the computer can first use a principal component analysis (PCA) technique on genetic ancestry data to generate principal components (PCs) for the genetic ancestry data. In doing so, the computer can reduce a size of the genetic ancestry data to reduce the number of inputs for subsequent processing using machine learning techniques. After generating the PCs and / or generating a matrix from the PCs, the computer can generate a feature vector with the PCs, matrices, and / or biomarker stratifier weights for a particular disease of interest. The computer can use a linear regression model or a neural network on the feature vector to generate biomarkerstratifi er scores for individuals of the genetic ancestry data. The computer can determine pharmacomimetic instruments for the disease of interest, where a disease-associated variant is a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype. The computer can subsequently use a neural network or linear regression model to determine interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores. The statistical interactions can be predictive of drug targetspecific differential treatment response for the disease of interest. The computer can perform one or more prediction-based actions based on the determined interactions. By using PCA in this way, the computer can format a feature vector from genetic ancestry data to substantially reduce the number of axes and / or nodes of a linear regression model and / or neural network to enable the linear regression model and / or neural network to operate and / or reduce the amount of computational resources that are required for the modeling or simulations.

[0101] Figure l is a flow diagram showing a method 200 of one embodiment of the disclosure. The method 200 can be performed by a data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG. 3, a server system, etc.). The method 200 may include more or fewer operations and the operations may be performed in any order. Performance of the method 200 may enable the data processing system to automatically identify predicted drug effect on phenotypes based on polygenic risk scores.

[0102] In the method 200, at operation 202, the data processing system obtains genetic variant data from a plurality of subjects with identified phenotypes. At operation 204, the data processing system determines the value of the variant effects for each phenotype based on external data. At operation 206, the data processing system constructs a matrix based on the variants, phenotypes, and values of the variant effects. At operation 208, the data processing system runs a principal components analysis biomarker variant x phenotype- related phenotype variant effects matrix. At operation 210, the data processing system calculates a polygenic risk score for each principal component. At operation 212 the data processing system identifies predicted drug effects on one or more phenotypes based on the polygenic risk scores calculated for each principal component.

[0103] Figure 2 illustrates a block diagram of a system 300 for in silico drug development and determination of drug activity in one embodiment of the disclosure. In brief overview, the system 300 can include a data processing system 302 and data sources 304, 306, and 308. The data processing system 302 can receive or retrieve molecular biomarker stratifier data of a plurality of subjects. The data processing system 302 can determine values representing biomarker stratifier effects based on external data of individuals separate from the plurality of subjects. The data processing system 302 can calculate a biomarker stratifier score, a chosen disease phenotype and a pharmacomimetic genetic score for different drug targets. The data processing system 302 can identify a predicted drug activity for each drug target based on the biomarker stratifier score and the pharmacomimetic genetic score for the drug target in association analysis with the disease phenotype. The system 300 may include more, fewer, or different components than shown in FIG. 3.

[0104] The data processing system 302 may comprise one or more processors that are configured to identify predicted drug activity for drug targets for disease phenotypes. The data processing system 302 may comprise a network interface 312, a processor 314, and / or memory 316. The data processing system 302 may communicate with the data sources 304, 306, and / or 308 via the network interface 312, which may be or include an antenna or other network device that enables communication across a network and / or with other devices. The processor 314 may be or include an ASIC, one or more FPGAs, a DSP, circuits containing one or more processing components, circuitry for supporting a microprocessor, a group of processing components, or other suitable electronic processing components. In some embodiments, the processor 314 may execute computer code or modules (e.g., executable code, object code, source code, script code, machine code, etc.) stored in memory 316 to facilitate the activities described herein. The memory 316 may be any volatile or non-volatile computer-readable storage medium capable of storing data or computer code.

[0105] The memory 316 may include a data collector 318, a principal component (PC) generator 320, a pharmacomimetic instrument identifier, a statistical interaction calculator 324, and / or an action performer 326, in some embodiments. The components 318-326 may operate to identify predicted drug activity of drug targets for different disease phenotypes.

[0106] For example, the data collector 318 may comprise programmable instructions that, upon execution, cause the processor 314 to communicate with the data sources 304, 306, 308, and / or any other data sources. The data collector 318 may be or include an applicationprogramming interface (API) that facilitates communication between the data processing system 302 and other computing devices. The communicator 318 may communicate with the data sources 304, 306, 308, and / or any other computing device across the network 310.

[0107] The data sources 304, 306, and 308 can be data sources that store molecular biomarker stratifier data and / or external data. For example, the data sources 304, 306, and / or 308 can be or include a database (e.g., a relational or graph database) that stores molecular biomarker stratifier data that includes biomarker stratifiers and at least one disease phenotype from or of a plurality of subjects or individuals. The data sources 304, 306 and / or 308 can additionally or instead include external data of individuals separate from the plurality of subjects or individuals of molecular biomarker stratifier data.

[0108] The data collector 318 can establish connections with one or more of the data sources 304, 306, and / or 308. The data collector 318 can establish the connections over the network 310. To do so, the data collector 318 can communicate with the data sources 304, 306, and / or 308 across the network 310. In one example, the data collector 318 can transmit syn packets to the respective data sources 304, 306, and / or 308 and establish the connections using a TLS handshaking protocol. The data collector 318 can use any handshaking protocol to establish a connection with the data sources 304, 306, and / or 308.

[0109] The data collector 318 can obtain the molecular biomarker stratifier data and the external data. In some embodiments, the data collector 318 can obtain the molecular biomarker stratifier data and the external data over the established connections from one or more of the data sources 304, 306, or 308. The data collector 318 may do so by querying the data sources 304, 306, or 308. In some embodiments, the data collector 318 can receive the molecular biomarker stratifier data and / or the external data as input (e.g., a manual input) by a user accessing the data processing system. The data collector 318 can obtain portions of the molecular biomarker stratifier data and the external data using any method or any combination of methods.

[0110] The PC generator 320 may comprise programmable instructions that, upon execution, cause the processor 314 to generate principal components. The PC generator 320 can generate principal components that correspond to (or based on) genetic ancestry data (e.g., different genes) for subjects in a study cohort. The genetic ancestry data can be based on genotyping arrays or whole-genome sequencing. The PC generator 320 can obtain the genetic ancestry data from a publicly available database, as described herein. The PCgenerator 320 can be FlashPCA, for example. The principal components can be used to control subsequent statistical analyses for population stratification, in some embodiments.[OHl] For example, the PC generator 320 can generate a plurality of PCs corresponding to genetic ancestry data for subjects in a study cohort. The PC generator 320 can do so by generating a genetic similarity matrix from the genetic ancestry data of the subjects in the study cohort and using principal component analysis on the genetic similarity matrix. The PC generator 320 can generate any number of PCs using principal components analysis.

[0112] The PC generator 320 can generate a biomarker stratifier score for each subject in the study cohort. The PC generator 320 can generate the biomarker stratifier scores based at least on (i) the PCs and on (ii) biomarker stratifier weights for a disease of interest. For example, a “raw PGS” (e.g., a biomarker stratifier score) for a subject S can be defined as the sum over each variant in a PGS variant weights table of ((the variant’s weight) * (the # of copies S has of that variant)). The PGS variant weights of the PGS variant weights table can be computed independent of the genetic ancestry data corresponding to the study cohort. The PC generator 320 can generate biomarker stratifier scores for any number of diseases, such as nonalcoholic steatohepatitis (NASH), chronic liver disease, diabetes, and / or obesity.

[0113] In some embodiments, the PC generator 320 can use machine learning techniques (e.g., a neural network or a linear regression model) to generate the biomarker stratifier scores. For instance, the PC generator 320 can generate a feature vector from the PCs for the individuals of the study cohort and execute a neural network or a linear regression model using the feature vector as input to generate the biomarker stratifier scores. In doing so, the PC generator 320 can reduce the size of the dataset that is used to generate biomarker stratifiers scores for individuals. The reduced size can enable the neural network or linear regression model to operate or reduce the computational resources that are required to do so compared with systems that may input genetic ancestry data directly into a neural network or a linear regression model.

[0114] The PC generator 320 can determine “ancestry-normalized” PGS. The PC generator 320 can determine the ancestry-normalized PGS to be the residuals of a linear regression model where the outcome can be the raw PGS and the predictors are the PCs of genetic ancestry data. This normalization can correct for differences in mean and / or variance of PGS values between populations. The ancestry-normalized PGS can also be called an ancestry- normalized biomarker stratifier score.

[0115] The PC generator 320 can determine “Scaled PGS.” The PC generator 320 can do so using the following equation: scaled PGS = ((the ancestry-normalized PGS - the ancestry- normalized PGS across all subjects in the cohort) / (the standard deviation of the ancestry- normalized PGS across all subjects in the cohort)). Scaled PGS values can thus be approximately normally distributed. The scaled PGS can be used for all subsequent analyses. Scaled PGSs are also called scaled biomarker stratifier scores.

[0116] The pharmacomimetic instrument identifier 322 may comprise programmable instructions that, upon execution, cause the processor 314 to determine which of a plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest. A disease-associated variant can be a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype.

[0117] To determine which of the plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest, the pharmacomimetic instrument identifier 322 may execute a program (e.g., GCTA COJO-SLCT) for the disease of interest to identify “conditionally independent disease-associated variants.” To do so, for example, the pharmacomimetic instrument identifier 322 may first run a GWAS (e.g., a first GWAS) as normal. The pharmacomimetic instrument identifier 322 can identify the variant that has the most significant association with the disease outcome (e.g., the smallest p-value) across the whole genome. The pharmacomimetic instrument identifier 322 can rerun the GWAS, this time conditioning all of the variant-disease association tests on the genotype of the most- significant variant. The pharmacomimetic instrument identifier 322 can identify the variant that has the next-most significant association with the disease outcome. The pharmacomimetic instrument identifier 322 can rerun the GWAS, this time conditioning all of the variant-disease association tests on the genotype of the most-significant and the 2nd- most-significant variants. The pharmacomimetic instrument identifier 322 can iteratively repeat this procedure until determining no variant is associated with the disease below a defined p-value threshold (e.g., a genome-wide significance threshold, such as 5 x 1 O’8). The pharmacomimetic instrument identifier 322 can use any threshold. The set of variants identified by the pharmacomimetic instrument identifier can be called “conditionally- independent, disease-associated variants” because each variant is associated with the disease even after conditioning on the genotypes of all of the previous variants.

[0118] The pharmacomimetic instrument identifier 322 can create a table. Each row of this table can correspond to one subject from the study cohort and include data regarding the subject as demographic data and other determined data (e.g., PGS, PCS of genetic ancestry, dosage for the alternate allege of each of the independent disease associated variants, etc.).

[0119] The pharmacomimetic instrument identifier 322 can identify which of the conditionally-independent, disease-associated variants are “pharmacomimetic instruments” (e.g., variants that modulate the function or expression of one or more drug target genes so that the variants’ effect on human phenotypes is likely to be predictive of the drug’s effect on human phenotypes).

[0120] Depending on the configuration, the pharmacomimetic instrument identifier 322 can use different criteria and / or priorities to identify pharmacomimetic instruments from the identified conditionally-independent, disease-associated variants. In one example, a user may be interested in clinical-stage drugs for the disease. In this case, the pharmacomimetic instrument identifier 322 can compile a table of clinical-stage drugs and their targets using public resources, such as clinicaltrials.gov and / or proprietary databases, such as Cortellis.The pharmacomimetic instrument identifier 322 can determine that variants that are close to a drug target gene (e.g., within 150 kb of the gene’s transcription start site) or are mapped to the drug target are pharmacomimetic instruments. In another example, a user may be interested in known and novel drug targets for the antibody modality. In this case, the pharmacomimetic instrument identifier 322 can compile a list of genes that encode proteins that are druggable with the antibody modality (e.g., proteins that are secreted or localized to the cell surface). The pharmacomimetic instrument identifier 322 can determine variants that are mapped to an antibody-druggable protein are pharmacomimetic instruments based on the compiled table (e.g., based on the determined variants having stored associations with the antibody-druggable protein).

[0121] In some embodiments, there is more than one suitable variant for a drug mechanism. In this case, it will increase statistical power to detect drug-PGS interactions by aggregating variants (e.g., all of the variants) identified as pharmacomimetic instruments that were associated with a given drug mechanism into a combined allelic score. To do so, for example, the pharmacomimetic instrument identifier 322 can compute an allelic score in the same or a similar manner to a PGS (e.g., assign each variant a weight and compute, for each subject in the cohort, the sum over each variant of (the variant’s weight) * (the subject’sdosage for the variant)). The difference between an allelic score and a PGS is that the allelic score is composed of variants that modulate the function or expression of a single drug target gene, while a PGS can be composed of variants across the genome that affect many different genes, e.g., ANGPTL3, ANGPTL4, and / or LPL.

[0122] The statistical interaction calculator 324 may comprise programmable instructions that, upon execution, cause the processor 314 to determine statistical interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores. The statistical interactions can be predictive of drug target-specific differential treatment responses for the disease of interest. To determine the statistical interactions, the statistical interaction calculator 324 can generate a data table with phenotypes, covariates, and a set of pharmacomimetic instruments determined using the aforementioned systems and methods, which may include single variants and multi-variant allelic scores. Responsive to doing so, the statistical interaction calculator 324 can test for interactions between the pharmacomimetic instruments and the PGS in association analysis with the disease or risk factor trait of interest.

[0123] For example, for each pharmacomimetic instrument, the statistical interaction calculator 324 can use statistical software, such as the R programming language, to fit a logistic regression model using the following example formula: disease case-control status ~ age + sex + ancestry PCs + (pharmacomimetic instrument) + (scaled PGS) + (pharmacomimetic instrument) : (scaled PGS). In doing so, the statistical interaction calculator 324 can determine the scaled PGS. The model can be implemented using the R programming language, such as by using the command glm(model_formula, family = “binomial”, data = our_genetic_and_phenotypic_data). In the example formula, the variable to the left of thesymbol is the dependent variable (e.g., the variable that is being predicted). Each of the variables to the right of the symbol, separated by “+” symbols, are independent variables, (e.g., variables that are observed in the data and used to predict the dependent variable). The symbol indicates a variable that is the product (result of multiplication) of the variables to the left and right of the symbol. This product is an “interaction term.” If the pharmacomimetic instrument : scaled polygenic score (PGS) interaction term — one of the independent variables — has a statistically-significant association with the dependent variable (in this case, disease status), controlling for all of the other independent variables — such as age and sex — then the statistical interaction calculator 324 can predict or determine that the drug that is modeled by the pharmacomimeticinstrument will have different effects in subjects with high PGS vs. low PGS. The statistical interaction calculator 324 can fit a logistic regression model for each pharmacomimetic instrument that the pharmacomimetic instrument identifier 322 determines or identifies.

[0124] The statistical interaction calculator 324 can use the same program to compute a p- value for the interaction term. If the p-value for the interaction between a pharmacomimetic instrument and the PGS is < 0.05 / (the # of instruments tested), the statistical interaction calculator 324 can determine the interaction to be “Bonferroni significant”. Alternatively, the statistical interaction calculator 324 can be configured to use a more lenient p-value threshold, at the risk of false positives.

[0125] The instrument-PGS interactions can be “drug-PGS interactions” because each pharmacomimetic instrument is a model for a particular drug mechanism.

[0126] The statistical interaction calculator 324 can generate tables that present the statistically significant drug-PGS interactions in a way that is easier to interpret than looking at the raw regression model outputs. For example, first, the statistical interaction calculator 324 can group the subjects of the study cohort by quantiles of the PGS, (e.g., the 0-33rd percentile, the 34th-66th percentile, and the 67th-100th percentile). Then, for each drug-PGS interaction, the statistical interaction calculator 324 can fit separate logistic regression models within each PGS quantile using the formula: disease case-control status ~ age + sex + ancestry PCs + the pharmacomimetic instrument for the drug. Using these regression outputs, the statistical interaction calculator 324 can construct a table that shows the effect of the pharmacomimetic instrument on disease risk (as well as a 95% confidence interval for that effect) within each PGS quantile.

[0127] For instance, the statistical interaction calculator 324 can divide the subjects of the study into a finite number of groups (e.g., the groups) based on their PGS values. Within each group, the statistical interaction calculator 324 can predict a disease outcome based on the patients’ genetics, specifically a genetic instrument that mimics the effects of a drug (the “pharmacomimetic instrument”). In doing so, the statistical interaction calculator 324 can determine a difference in how strongly the pharmacomimetic instrument predicts the disease outcome in individuals with high PGS vs. middle PGS vs. low PGS. The difference may indicate that a drug might have different efficacy in people with high PGS vs. middle PGS vs. low PGS.

[0128] The statistical interaction calculator 324 can also compute a statistic called the “treatment effect multiplier.” This is the effect of the pharmacomimetic instrument on disease risk in the top quartile divided by the effect in all subjects. The “treatment effect multiplier” can represent how much one could increase the average treatment effect in a clinical trial of the drug if one enrolled only patients in the top quantile of the PGS as opposed to enrolling all qualified patients. The treatment effect multiplier is a way to compare drug-PGS interactions to distinguish “strong” from “weak” interactions.

[0129] The action performer 326 may comprise programmable instructions that, upon execution, cause the processor 314 to perform one or more prediction-based actions based on the determined interactions. The one or more prediction-based actions can include at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics. To perform a prediction-based action, the action performer 326 can select a patient for treatment. The action performer 326 can select the patient based on a likelihood that the patient will benefit from the treatment. For example, the action performer 326 can select the patient by determining a PGS for the patient using the systems and methods described herein and determining the PGS for the patient is above a threshold (e.g., a PGS threshold) or within a range (e.g., a quartile, a quintile, a defined range, etc.). In another example, the action performer 326 can select patients for novel therapeutic mechanisms whose benefit is mostly or only apparent in subsets of common disease patients who have elevated polygenic risk. This enables novel target identification and novel therapeutic discovery. In another example, the action performer 326 can select patients from large common disease treatment-eligible populations who will benefit most from existing drug therapies, and, conversely, identify patients who are not likely to receive clinically meaningful benefit. This can enhance the pharmaco-economic profile of existing therapeutic mechanisms, yielding more cost-effective utilization. In another example, the action performer 326 can select mechanisms and design of prospective clinical trials of investigational drugs that are predicted to yield greater benefit in individuals with elevated PGS. These trials may be run with many-fold fewer patient years required to demonstrate clinical benefit, given the expectation of magnified event rate and treatment response rates.

[0130] In some embodiments, treatment can be performed on the selected patients. For example, the action performer 326 can select patients with PGS scores that satisfy a criteria for a particular treatment to treat a disease. The action performed 326 can generate a record that includes a list of selected patients. An entity (e.g., a user or a clinician) may view the listand apply a treatment for the disease to the patients identified in the record. Examples of such treatment can include injections, therapy, or other treatment techniques. In some embodiments, the action performer 326 can connect with another computer and transmit the record containing the list of patients or the PGS scores determined by the PC Generator 320 to another computer. The computer may display the scores or list of patients and the entity or user accessing the computer may use the list and / or scores for treatment.

[0131] Figure 3 illustrates a flow diagram showing a method 400 for in silico drug development and determination of drug activity in one embodiment of the disclosure. The method 400 can be performed by a data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG. 3, a server system, etc.). The method 400 may include more or fewer operations and the operations may be performed in any order. Performance of the method 400 may enable the data processing system to automatically identify predicted drug activity of different drug targets for disease phenotypes and / or phenotype intermediates.

[0132] In the method 400, at operation 402, the data processing system obtains molecular biomarker stratifier data. The molecular biomarker stratifier data can include a plurality of biomarker stratifiers and at least one disease phenotype or phenotype intermediate from each of a plurality of subjects. The molecular stratified biomarker data can include at least one of the following types of data: genomic, transcriptomic, metabolomic, or proteomic. These data are obtained by assays specific to each biomarker type. The at least one disease phenotype or phenotype intermediate can include or be associated with nonalcoholic steatohepatitis (NASH), chronic liver disease, diabetes, and / or obesity.

[0133] The molecular biomarker stratifier data can include data on the at least one disease phenotype. The data can include a patient’s disease status (e.g., has been diagnosed with NASH yes / no) at a given point in time. The data can also include quantitative disease traits (e.g., BMI measurements at a given time with regard to disease diagnosis), medication information (e.g., patient has taken or is taking a statin or other lipid-lowering therapy at a given time with regard to disease outcomes), and demographic information (e.g., age at diagnosis, biological sex, etc.).

[0134] At operation 404, the data processing system determines a plurality of values (e.g., numeric values) representing biomarker stratifier effects. Each value can separately represent how each of the plurality of biomarker stratifiers affects each of the at least one diseasephenotype or phenotype intermediate. The biomarker stratifier effects can respectively be the estimated magnitude of effect of the biomarker on disease risk or continuous disease trait from a statistical model. The data processing system can determine the plurality of values based on external data (e.g., data regarding individuals separate from the subjects of the plurality of subjects). The data processing system can determine the plurality of values using a logistic regression model for disease (yes / no) vs. biomarker (e.g., protein measurement) + covariates (age, sex, etc.) on the data.

[0135] At operation 406, the data processing system calculates a biomarker stratifier score for a chosen disease phenotype or phenotype intermediate. The data processing system can calculate a biomarker stratifier score for each of the at least one disease phenotype or phenotype intermediate. The data processing system can determine biomarker stratifier scores as polygenic risk scores for individual disease phenotypes. The data processing system can calculate biomarker stratifier scores for biomarker stratifiers. In doing so, for example, the data processing system can determine a biomarker stratifier score as one of the following: a polygenic risk score for the disease phenotypes; a polygenic score for a disease risk factor; a polygenic score for biological pathway activity; a polygenic score for drug target expression; a proteomics risk score for the disease phenotypes; a proteomics score for a disease risk factor; a proteomics score for biological pathway activity; a proteomics score for drug target expression; a transcriptomics risk score for the disease phenotypes; a transcriptomics score for a disease risk factor; a transcriptomics score for a biological pathway activity; a somatic mutation score for drug target expression; a somatic mutation score for the disease phenotypes; a somatic mutation score for a disease risk factor; a somatic mutation score for a biological pathway activity; or a somatic mutation score for drug target expression. For instance, if the stratifier is a polygenic risk score (PRS), a statistical model may be trained using summary statistics from a published study on large-scale cohorts or from a meta-analysis of multiple studies based on large and diverse set of cohorts (none of which include UK Biobank). The data processing system can apply the estimated weights from that model to a dataset that is orthogonal to the ones used for training (e.g., UK Biobank).

[0136] At operation 408, the data processing system calculates a pharmacomimetic genetic score for each drug target for which the data processing system is determining a drug activity. The pharmacomimetic genetic score can represent loci for which more than one pharmacomimetic variant (e.g., a functional variant in a locus that is strongly andsignificantly associated with a given disease) is identified. The data processing system can determine the pharmacomimetic genetic score by weighting each variant by its estimated effect on the phenotype and summing the weighted values.

[0137] At operation 410, the data processing system identifies the predicted drug activity of each drug target for the at least one disease phenotype or phenotype intermediate in subsets of a biomarker stratifier distribution. The data processing system can identify the predicted drug activity based on a statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the disease phenotype or phenotype intermediate. For example, the data processing system can use a model (e.g., a linear regression model, a logistic regression model, a Cox Proportional Hazards model, etc.) that includes a predictor defined by the multiplicative interaction between the pharmacomimetic genetic score and biomarker score to perform the association analysis for each drug target and disease outcome phenotype. The predictor can be the predicted drug activity.

[0138] The data processing system can identify the subsets of the biomarker stratifier distribution using one or more thresholds. The one or more thresholds can be chosen to define different subsets depending on the use case. In one example, it may be optimal to use a 25th percentile of the polygenic risk score to define a subset of patients who are predicted to have an outsized clinical benefit from therapy X. In another example, it may be optimal to use a 33% (top tertile) of the PRS to define a subset of patients who would benefit most from therapy Y. In other disease settings with different PRS, the thresholds may also vary. Defining the optimal threshold will depend on factors such as: estimated effect size for the subset vs. all-comers on therapy, failure screening rate (in a trial, if the threshold is set too conservatively, e.g., 5%, a lot more patients will have to be screened to enroll a few), etc. Subsets can be from individual-level cohort data as the top 10%, 20%, 30%, etc. of the distribution.

[0139] Figure 4 illustrates a flow diagram showing a method 500 for in silico drug development and determination of drug activity in one embodiment of the disclosure. The method 500 can be performed by a data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG. 3, a server system, etc.). The method 500 may include more or fewer operations and the operations may be performed in any order. Performance of the method 500 may enable the data processing system toautomatically identify predicted drug activity of different drug targets for disease phenotypes or phenotype intermediates.

[0140] In the method 500, at operation 502, the data processing system obtains molecular biomarker stratifier data and at least one disease phenotype or phenotype intermediate. The molecular biomarker stratifier data can include a plurality of biomarker stratifiers. The data processing system can obtain the molecular biomarker stratifier data from each of a plurality of subjects.

[0141] At operation 504, the data processing system determines the value of the biomarker stratifier effects for each disease phenotype or phenotype intermediate. Each value can separately represent how each of the plurality of biomarker stratifiers affects each of the at least one disease phenotype or phenotype intermediate. The data processing system can determine the plurality of values based on external data (e.g., data regarding individuals separate from the subjects of the plurality of subjects). At operation 506, the data processing system constructs a matrix based on the biomarker stratifier data (e.g., biomarker stratifiers), disease phenotypes or phenotype intermediate values. At operation 508, the data processing system runs a principal components analysis biomarker stratifier phenotype-related x phenotype biomarker stratifier effects matrix. At operation 510, the data processing system calculates a biomarker stratifier score for each principal component. At operation 512, the data processing system calculates a pharmacomimetic genetic score for each drug target. At operation 514, the data processing system identifies predicted drug activity of each drug target on one or more phenotypes or phenotype intermediates in subsets of the biomarker stratifier distribution. The data processing system can identify the predicted drug activity based on a statistical interaction of the polygenic score calculated for each principal component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype or outcome phenotype intermediate.

[0142] Figure 5 illustrates a flow diagram showing a method 600 for in silico drug development and determination of drug activity in one embodiment of the disclosure. The method 600 can be performed by a data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG. 3, a server system, etc.). The method 600 may include more or fewer operations and the operations may be performed in any order. Performance of the method 600 may enable the data processing system to automatically determine interactions between pharmacomimetic instruments and performprediction-based actions based on the determined interactions. One or more of the operations in the methods 400, 500, and / or 600 can be performed during operations in any other of the methods 400, 500, and / or 600.

[0143] In the method 600, at operation 602, the data processing system generates a plurality of principal components (PCs) corresponding to genetic ancestry data (e.g., genetic data) for subjects in a study cohort. The data processing system can generate any number of PCs. For example, the data processing system can generate five PCs, at least six PCs for a European ancestry cohort, and / or 10-to-40 PCs for a multi-ancestry cohort (e.g., 20 for UK Biobank and / or possibly up to 40 for a more diverse cohort). The study cohort can include at least 200 control subjects. In some embodiments, the study cohort can include at least 950 control subjects or at least 990 control subjects. The data processing system can determine whether subjects in the study cohort are a case or a control for a disease of interest, such as based on flags for the subjects that indicate whether the subjects are cases or controls. The genetic ancestry data can be based on genotyping arrays or whole-genome sequencing. The data processing system can obtain the genetic ancestry data from a publicly available database, for example, or through any other method.

[0144] At operation 604, the data processing system generates a biomarker stratifier score for each subject in the study cohort. The data processing system can generate the biomarker stratifier scores based at least on (i) the PCs and on (ii) biomarker stratifier weights for a disease of interest. The biomarker stratifier score can be for the particular disease of interest. For example, a “raw PGS” (e.g., a biomarker stratifier score) for a subject S can be defined as the sum over each variant in a PGS variant weights table of ((the variant’s weight) * (the # of copies S has of that variant)). For example, consider three variants A, B, and C, with weights 1, 2, and 3 respectively. If subject S has 2, 0, and 1 copies of variants A, B, C, then their unsealed PGS is computed as 2*weight(A) + 0*weight(B) + l*weight(C) = 2*1 + 0*2 + 1*3 = 5. The PGS variant weights can be computed independent of the genetic ancestry data corresponding to the study cohort.

[0145] The data processing system can determine “ancestry-normalized PGSs.” Ancestry- normalized PGS can be defined as the residuals of a linear regression model where the outcome can be the raw PGS and the predictors are the PCs of genetic ancestry that were selected in operation 602. This normalization can correct for differences in mean and / or variance of PGS values between populations.

[0146] The data processing system can determine “scaled PGSs.” Scaled PGS can be defined as ((the ancestry-normalized PGS - the ancestry -normalized PGS across all subjects in the cohort) / (the standard deviation of the ancestry-normalized PGS across all subjects in the cohort)). Scaled PGS values can thus be approximately normally distributed. The scaled PGS can be used for all subsequent analyses. The task of creating and scaling PGS that is effective and comparable across populations may be performed using any method.

[0147] At operation 606, the data processing system determines which of a plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest. A disease-associated variant can be a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype.To determine which of the plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest, the data processing system may execute a program (e.g., GCTA COJO-SLCT) for the disease of interest to identify “conditionally independent disease-associated variants.”

[0148] For example, groups of genetic variants that are close together on a DNA molecule (e.g., roughly within 500,000 base pairs, although this distance may vary depending upon where in the genome the variants are located) can be inherited together. Variants that are usually inherited together are said to be in “linkage disequilibrium”, and this distance between loci that are inherited together may differ depending upon where the loci and how they are distributed in the genome. When a single variant contributes to the risk of a disease, a GWAS will reveal statistical associations of that variant (the “causal variant”) with the disease, but also statistical associations with the disease for other variants that are in linkage disequilibrium (inherited together) with the causal variant. Therefore, in GWAS, when there is a group of variants that are close together in their position on a DNA molecule, and many of those variants are associated with the disease, the data processing system can distinguish between two scenarios: 1) a scenario in which all of the apparent disease-associated variants are inherited together, which would imply that there is likely to be a single “causal variant”, or 2) a scenario in which there are two or more distinct groups of variants, with one or more variants in each group, such that variants within each group are inherited together, but theinheritance of one group vs. another is independent. In this latter case, there is likely to be one “causal variant” per group of linked variants.

[0149] To distinguish between the “single causal variant” and the “multiple causal variants” scenario, the data processing system can implement a regression function. For example, first, a GWAS (e.g., a first GWAS) can be performed as normal. The data processing system can identify the variant that has the most significant association with the disease outcome (e.g., the smallest p-value) across the whole genome. The data processing system can rerun the GWAS, this time conditioning all of the variant-disease association tests on the genotype of the most-significant variant. The data processing system can identify the variant that has the next-most significant association with the disease outcome. The data processing system can rerun the GWAS, this time conditioning all of the variant-disease association tests on the genotype of the most-significant and the 2nd-most-significant variants. The data processing system can iteratively repeat this procedure until determining no variant is associated with the disease below a defined p-value threshold (e.g., a genome-wide significance threshold, such as 5 x 10'8). The data processing system can use any threshold. Responsive to stopping the procedure, the set of variants the data processing system identified can be called “conditionally-independent, disease-associated variants” because each variant is associated with the disease even after conditioning on the genotypes of all of the previous variants.

[0150] Returning to the “single causal variant” and “multiple causal variants” scenarios, if the data processing system identified only a single “conditionally-independent disease- associated variant” in a given region of DNA during the stepwise regression, then the “single causal variant” scenario is more likely to be true. However, if the stepwise regression identified several “conditionally-independent disease-associated variants”, then the “multiple causal variants” scenario is more likely to be true.

[0151] In instances in which a disease (e.g., coronary artery disease) is associated with multiple biomarkers (e.g., LDL cholesterol and blood pressure), the data processing system may use multiple independent acceptance criteria to identify biomarker-associated variants to increase the number of variants that the data processing system identifies. For example, the data processing system can identify variants with p < 5 x 10'8for the disease and variants with p < 5 x 10'6for the disease that also have p < 5 x 10'8for at least one disease-relevant biomarker. The data processing system can use any criteria to identify variants. Thevariations in acceptance criteria can enable the data processing system to identify a wide range of variants using a COJO-SLCT analysis.

[0152] The data processing system can create a table. Each row of this table can correspond to one subject from the study cohort. The columns of the table are described as follows: a. Whether the subject is a case or control for the disease. Cases are assigned a value of “1” and controls are assigned a value of “0”. Subjects who are neither cases nor controls are excluded from the table. b. Age and sex of the subject. c. The PCs of genetic ancestry from the operation 602 or 604. d. The scaled PGS for disease risk from the operation 604. e. The subject’s dosage for the alternate allele of each of the independent disease- associated variants. i. Here, “alternate allele” is in reference to a reference genome such as the Genome Reference Consortium’s GRCh37 or GRCh38. In a reference genome, each variant has a “reference allele”. A bi-allelic variant is therefore a site in the genome where an individual has either the “reference allele” or an “alternate allele” at the site. ii. Some sites are “multi-allelic”, meaning that in the population, three or more alleles exist. Computationally, each alternate allele is treated as a separate bi-allelic variant, (e.g., if you have alleles “reference”, “altemate-1”, and “alternate-2”, then it would be treated as two bi-allelic variants (reference, alternate-1) and (reference, alternate-2)).

[0153] The data processing system can identify which of the disease-associated variants are “pharmacomimetic instruments” (e.g., variants that modulate the function or expression of a drug target gene so that the variants’ effect on human phenotypes is likely to be predictive of the drug’s effect on human phenotypes). Depending on the configuration, the data processing system can use different criteria and / or priorities to identify pharmacomimetic instruments from the identified disease-associated variants. In one example, a user may be interested in clinical-stage drugs for the disease. In this case, the data processing system can compile a table of clinical-stage drugs and their targets using public resources, such as clinicaltrials.gov, and / or proprietary databases, such as Cortellis. The data processing system can determine the variants that are close to a drug target gene (e.g., within 150 kb of the gene’s transcription start site) or that are mapped to the drug target are pharmacomimetic instruments. In another example, a user may be interested in known and novel drug targets for the antibody modality. In this case, the data processing system can compile a list of genes that encode proteins that are druggable with the antibody modality (e.g., proteins that are secreted or localized to the cell surface). The data processing system can determine variants that are mapped to an antibody-druggable protein are pharmacomimetic instruments. The data processing systemcan determine variants that are mapped to an antibody-druggable protein are pharmacomimetic instruments based on the compiled table (e.g., based on the determined variants having stored associations with the antibody-druggable protein).

[0154] Optionally, one might require that genetic evidence supports the hypothesis that inhibition of the target would be beneficial, as opposed to activation of the target. This evaluation could consider data from expression quantitative trait locus (eQTL) and protein quantitative trait locus (pQTL) datasets that inform on how genetically-determined changes in expression of the target affect disease risk.

[0155] In some embodiments, there is more than one suitable variant for a drug mechanism. In this case, the data processing system can increase statistical power to detect drug-PGS interactions by aggregating variants (e.g., all of the variants) identified as pharmacomimetic instruments that were associated with a given drug mechanism into a combined allelic score. To do so, for example, the data processing system can compute an allelic score in the same or a similar manner to a PGS (e.g., assign each variant a weight and compute, for each subject in the cohort, the sum over each variant of (the variant’s weight) * (the subject’s dosage for the variant)). The difference between an allelic score and a PGS is that the allelic score is composed of variants that modulate the function or expression of a single drug target gene, while a PGS can be composed of variants across the genome that affect many different genes. The variants in the allelic score may be weighted using a GWAS (e.g., a second GWAS) for the disease that did not include any subjects from the study cohort. Alternatively, the variants can be weighted using a GWAS (e.g., a third GWAS) for a biomarker that causally mediates the effect of the drug target on the disease (e.g., LDL cholesterol mediates the effect of PCSK9 on coronary artery disease). A biomarker GWAS still must not overlap the study cohort. The allelic score can be used in place of PGS to perform the systems and methods described herein.

[0156] At operation 608, the data processing system determines statistical interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores. The statistical interactions can be predictive of drug target-specific differential treatment responses for the disease of interest. To determine the statistical interactions, the data processing system can generate a data table with phenotypes (e.g., nonalcoholic steatohepatitis (NASH), chronic liver disease, diabetes, obesity, etc.) covariates, and a set of pharmacomimetic instruments determined using the aforementioned systems andmethods, which may include single variants and multi-variant allelic scores. Responsive to doing so, the data processing system can test for interactions between the pharmacomimetic instruments and the PGS (or allelic scores) in association analysis with the disease or risk factor trait of interest.

[0157] For example, for each pharmacomimetic instrument, the data processing system can use statistical software, such as the R programming language, to fit a logistic regression model using the following example formula: disease case-control status ~ age + sex + ancestry PCs + (pharmacomimetic instrument) + (scaled PGS) + (pharmacomimetic instrument) : (scaled PGS). The model can be implemented using the R programming language, such as by using the command glm(model_formula, family = “binomial”, data = our_genetic_and_phenotypic_data). In the example formula, the variable to the left of the symbol is the dependent variable (e.g., the variable that is being predicted). Each of the variables to the right of the symbol, separated by “+” symbols, are independent variables, (e.g., variables that are observed in the data and used to predict the dependent variable). The symbol indicates a variable that is the product (result of multiplication) of the variables to the left and right of the symbol. This product is an “interaction term”. If the pharmacomimetic instrument : scaled polygenic score (PGS) interaction term — one of the independent variables — has a statistically-significant association with the dependent variable (in this case, disease status), controlling for all of the other independent variables — such as age and sex — then the data processing system can predict or determine that the drug that is modeled by the pharmacomimetic instrument will have different effects in subjects with high PGS vs. low PGS. The data processing system can fit a logistic regression model for each pharmacomimetic instrument that the data processing system determines or identifies, such as based on the PGS of the different subjects.

[0158] The data processing system can use the same program to compute a p-value for the interaction term. If the p-value for the interaction between a pharmacomimetic instrument and the PGS is < 0.05 / (the # of instruments tested), the data processing system can determine the interaction to be “Bonferroni significant.” Alternatively, the data processing system can be configured to use a more lenient p-value threshold, at the risk of false positives.

[0159] The instrument-PGS interactions can be “drug-PGS interactions” because each pharmacomimetic instrument is a model for a particular drug mechanism.

[0160] The data processing system can generate tables that present the statistically significant drug-PGS interactions in a way that is easier to interpret than looking at the raw regression model outputs. For example, first, the data processing system can group the subjects of the study cohort by quantiles of the PGS, (e.g., the 0-33rd percentile, the 34th-66th percentile, and the 67th- 100th percentile). Then, for each drug-PGS interaction, the data processing system can fit separate logistic regression models within each PGS quantile using the formula: disease case-control status ~ age + sex + ancestry PCs + the pharmacomimetic instrument for the drug. Using these regression outputs, the data processing system can construct a table that shows the effect of the pharmacomimetic instrument on disease risk (as well as a 95% confidence interval for that effect) within each PGS quantile.

[0161] For instance, the data processing system can divide the subjects of the study into a finite number of groups (e.g., the groups) based on their PGS values. Within each group, the data processing system can predict a disease outcome based on the patients’ genetics, specifically a genetic instrument that mimics the effects of a drug (the “pharmacomimetic instrument”). In doing so, the data processing system can determine a difference in how strongly the pharmacomimetic instrument predicts the disease outcome in individuals with high PGS vs. middle PGS vs. low PGS. The difference may indicate that a drug might have different efficacy in people with high PGS vs. middle PGS vs. low PGS.

[0162] The data processing system can also compute a statistic called the “treatment effect multiplier.” This is the effect of the pharmacomimetic instrument on disease risk in the top quartile divided by the effect in all subjects. The “treatment effect multiplier” can represent how much one could increase the average treatment effect in a clinical trial of the drug if one enrolled only patients in the top quantile of the PGS as opposed to enrolling all qualified patients. The treatment effect multiplier is a way to compare drug-PGS interactions to distinguish “strong” from “weak” interactions.

[0163] At operation 610, the data processing system performs one or more prediction-based actions. The data processing system can perform the one or more prediction-based actions based on the determined interactions. The one or more prediction-based actions can include at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics. To perform a prediction-based action, the data processing system can select a patient for treatment. The data processing system can select the patient based on a likelihood that the patient will benefit from the treatment. For example, the data processingsystem can select the patient by determining a PGS for the patient using the systems and methods described herein and determining the PGS for the patient is above a threshold (e.g., a PGS threshold) or within a range (e.g., a quartile, a quintile, a defined range, etc.). In another example, the data processing system can select patients for novel therapeutic mechanisms whose benefit is mostly or only apparent in subsets of common disease patients who have elevated polygenic risk. This enables novel target identification and novel therapeutic discovery. This application can be called “polygenic therapeutic target ID.” In another example, the data processing system can select patients from large common disease treatment-eligible populations who will benefit most from existing drug therapies, and, conversely, identify patients who are not likely to receive clinically meaningful benefit. This can enhance the pharmaco-economic profile of existing therapeutic mechanisms, yielding more cost-effective utilization. This can be called “polygenic pharmacogenomics.” In another example, the data processing system can select mechanisms and design of prospective clinical trials of investigational drugs that are predicted to yield greater benefit in individuals with elevated PGS. These trials may be run with many-fold fewer patient years required to demonstrate clinical benefit, given the expectation of magnified event rate and treatment response rates. This application can be called “polygenic therapeutic development.” In some embodiments, the injections, therapy, and / or treatment techniques can be performed on the selected individuals when performing the one or more predictionbased actions.

[0164] In one example, the systems and methods described herein can be used to treat a patient. For instance, a patient can visit a clinic. The clinician can collect a blood sample, a saliva sample, or a tissue sample from the patient. The clinician can use the collected sample to perform a genotyping array on the patient. A data processing system implementing the systems and methods described herein can analyze the patient’s genetic data to generate principal components (PCs) that reflect the patient’s genetic ancestry. The patient can undergo tests to identify relevant biomarkers for a disease of interest (e.g., diabetes, cancer, cardiovascular diseases, psychiatric disorders, autoimmune diseases, neurological disorders, infectious diseases, asthma and allergies, etc.). These biomarkers are quantifiable biological parameters that could include patient characteristics such as blood sugar levels, cholesterol levels, specific protein markers, etc., depending on the disease. The data processing system can calculate a biomarker stratifier score for the patient for a disease of interest (e.g., adisease of interest from the list above) based on the results from the biomarker tests and the genetic ancestry data (e.g., the PCs).

[0165] The data processing system can scan the patient’s genetic data for disease-associated variants. In doing so, the data processing system can identify genes or gene variants in the patient’s genetic data that are associated with the disease for which the data processing system determined the biomarker stratifier score for the patient. For instance, the data processing system can identify disease-associated variants for the disease using the systems and methods described herein. The data processing system can use the identified disease- associated variants as a key in a query through the genetic array data of the patient. Based on the query, the data processing system can identify any disease-associated variants in the patient’s genetic data.

[0166] The data processing system can determine which of the identified disease-associated variants are pharmacomimetic instruments (e.g., whether these genetic variants can predict how the patient might respond to certain drugs or methods of treatment). For example, prior to determining treatment for the patient, the data processing system may have identified a list of pharmacomimetic instruments for the disease of interest using the systems and methods described herein. The data processing system can compare the identified disease-associated variants in the patient’s genetic makeup with the list of pharmacomimetic instruments. Based on the comparison, the data processing system can identify a set of pharmacomimetic instruments of the patient. In some embodiments, the data processing system can identify the pharmacomimetic instruments of the patient by querying the patient’s genetic data for instances of the pharmacomimetic instruments for the disease without identifying the disease- associated variants.

[0167] The data processing system can determine one or more treatments for the patient based on the pharmacomimetic instruments and the biomarker stratifier score of the patient. The data processing system can do so, for example, based on statistical interactions that the data processing system has previously determined for the pharmacomimetic instruments and biomarker stratifier scores. For instance, the data processing system can store a record (e.g., a record based on logistic regressions of PGS or biomarker stratifier scores and the pharmacomimetic instruments) indicating that individuals with a high polygenic risk score or biomarker stratifier score (e.g., above a threshold) for the disease of interest may have a high positive response to a therapeutic inhibitor (represented by a pharmacomimetic instrument)and individuals with a low polygenic risk score or biomarker stratifier score (e.g., below the threshold) for the disease of interest may have a low positive response to the inhibitor. The patient may have a high polygenic risk score or biomarker stratifier score. Accordingly, the data processing system can generate a record (e.g., a file, notification, alert, user interface, data structure, etc.) indicating or recommending treating the patient with a therapeutic inhibitor. The data processing system can display the record on a user interface of a client device accessed by the clinician. The data processing system can recommend any form of treatment based on interactions between pharmacomimetic instruments and polygenic risk scores biomarker stratifier scores of patients.

[0168] The clinician can treat the patient based on the recommendation. For example, the clinician can view the recommendation to use the inhibitor to treat the patient. The clinician can administer the treatment, such as by administering pills or capsules, administering an injection or set of injections, or using gene editing tools. The clinician can perform any type of treatment based on the polygenic risk scores or biomarker stratifier scores, type of disease, and / or pharmacomimetic instruments of the patient.

[0169] In some embodiments, the data processing system can automatically perform a treatment (e.g., using a gene editing tool or another treatment device or mechanism) responsive to determining the treatment. For example, the data processing system can determine a treatment to inhibit a particular gene to treat a disease for a patient as described herein. The data processing system can configure a treatment machine (e.g., a gene editing tool) to automatically implement the treatment on the patient responsive to the determination, such as by controlling the machine to perform the treatment (e.g., inhibiting or otherwise editing a specific gene or by performing an injection).

[0170] In another example, the systems and methods described herein can be used to perform a clinical trial to determine the efficacy of a candidate drug. For instance, the data processing system can identify a disease of interest that the candidate drug is intended to treat. The data processing system can identify pharmacomimetic instruments for the disease of interest. The data processing system can identify a pharmacomimetic instrument for the candidate drug, such as based on a user input. The data processing system can receive genetic data of a study cohort. The data processing system can determine polygenic risk scores for the study cohort using the systems and methods described herein based on the genetic data.

[0171] The data processing system can select participants for the trial based on the polygenic risk scores and interaction data between the polygenic risk scores and the pharmacomimetic instrument for the candidate drug. For instance, the data processing system can identify interactions between the pharmacomimetic instrument and the polygenic risk scores. From the identified interactions, the data processing system can determine that there is a high positive interaction between high polygenic risk scores (e.g., polygenic risk scores that exceed a threshold) and the pharmacomimetic instrument. Accordingly, the data processing system may filter participants from the study cohort and identify participants for the study cohort that have a high polygenic risk score. The data processing system can generate a record including the list of patients identified for the trial. The data processing system can present the list of patients on a user interface to a clinician and / or perform treatment on the participants in the trial as described above.

[0172] Figure 7 illustrates a sequence 700 for in silico drug development and determination of drug activity in one embodiment of the disclosure. The sequence 700 includes an illustration of matrices that can be used in performing the systems and methods described herein. A data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG. 3, a server system, etc.) can perform the operations and generate the matrices of the sequence 700. The sequence 700 may include more or fewer operations and the operations may be performed in any order. Performance of the method 700 may enable the data processing system to automatically determine interactions between pharmacomimetic instruments and perform prediction-based actions based on the determined interactions.

[0173] For example, the data processing system performing the sequence 700 can generate a variant effects matrix 702. The data processing system can generate the variant effects matrix 702 to have columns 704a-n (columns 704) and rows 706a-n (rows 706). n can be any number and can vary between the columns 704 and the rows 706. The columns 704 can each correspond to a different phenotype. In some embodiments, the column 704a can correspond to a reference phenotype and the columns 704b-n can be phenotypes that are biologically relevant to the reference phenotype of the column 704a. The rows 706 can each correspond to variant effects of separate variants on the phenotypes of the respective columns 704. The data processing system can determine the values for the variant effects and insert the values into the variant effects matrix 702.

[0174] The data processing system can perform a truncated principal components analysis (PC A) 708 on the variant effects matrix 702 to generate a variant-PC loadings matrix 710. The data processing system can generate the variant-PC loadings matrix 710 to have columns 712a-n (columns 712) and rows 714a-n (rows 714). The columns 712 can each correspond to a different principal component. The rows 706 can each correspond to a different variant. The values in the intersections between the columns 712 and the rows 714 can indicate variant loads for particular principal components.

[0175] The data processing system can generate or calculate a weight (e.g., a new weight) for each PC and variant combination 716. The data processing system can use the weight to generate a new PGS for the phenotypes. The data processing system can do so using the following equation: A variant’s new weight for PCi = the variant’s old weight * (the variant’ s loading for PCi)2 / (sumj= 1 ... k of [the variant’ s loading for PCj]2). The variant’ s old weight can be the weight for the variant that was used to determine an existing PGS 718. The data processing system can use the newly calculated weights to calculate PGSs for phenotypes using the systems and methods described herein. In one example, the data processing system can include weights and / or the variant-PC loadings matrix 710 or PCs in a feature vector, in some cases with the PCs themselves. The data processing system can execute a neural network or linear regression model based on the feature vector to generate the PGSs for different individuals. The data processing system can use the PGSs to identify patients for treatment and / or entities may treat patients based on the PGSs.Variant data for polygenic scores

[0176] There are many possible approaches to combine information across loci for assessment of polygenic risk scores (PRSs) for various conditions. Many studies have shown that PRSs can predict disease status in research-based case-control studies. See, e.g., Mavaddat N, et al., (2019) Am J Hum Genet. Vol. 104:21-34; Wray N Ret al., (2018) Nat Genet. Vol. 50:668-81. More convincingly, the prediction is also valid in population-based cohort studies and in electronic health record-based studies, especially for psychiatric disorders. Musliner KL et al. (2019) JAMA Psychiatry. Vol. 76:516-25; Lewis CM and Hagenaars SP. (2019) JAMA Psychiatry. Vol. 76:470-472, Zheutlin AB, et al., (2019) Am J Psychiatry. Vol. 176(10):846-55. The PRS can be formed from a set of independent risk variants associated with a disorder, based on the current evidence from the largest or most informative genome-wide association studies. For each individual, the number of risk allelescarried at each variant (0, 1, or 2) is summed, weighted by its effect size (i.e., log (OR) for binary traits or beta coefficient for continuous traits). The outcome is a single score of each individual’s genetic loading for a disease or for a continuous trait.

[0177] Much of the research on polygenic scores comes from research studies in cardiovascular disease, type 2 diabetes, breast and prostate cancers, and Alzheimer’s disease (Lambert SA et al. (2019) Hum Mol Genet. Vol. 28( R2): R133-42). Studies using the UK Biobank have demonstrated that PRS based on variant data can identify which percentage of patients have at least 3-fold increased risk for coronary artery disease, atrial fibrillation, type 2 diabetes, inflammatory bowel disease, and breast cancer, with the proportion of individuals identified varying between 1 .5 and 8% depending on the disorder. Khera AV et al. (2018) Nat Genet. Vol. 50:1219-24. Although these effects appear modest Applicant’s disclosure can identify substantially larger fractions of the population at high disease risk than monogenic mutations, making PRS potentially more clinically relevant.Definitions

[0178] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.

[0179] Throughout this disclosure, various publications, patents and published patent specifications are referenced by an identifying citation. The disclosures of these publications, patents and published patent specifications are hereby incorporated by reference into the present disclosure to more fully describe the state of the art to which this disclosure pertains.

[0180] As used herein, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. For example, “a” or “an” means “at least one” or “one or more.” It is understood that aspects and variations described herein include “consisting of and / or “consisting essentially of aspects and variations.

[0181] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Accordingly, the description of a range should be considered tohave specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For example, where a range of values is provided, it is understood that each intervening value, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the claimed subject matter. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the claimed subject matter, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the claimed subject matter. This applies regardless of the breadth of the range.

[0182] “About” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±20% or± 10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% from the specified value, as such variations are appropriate to perform the disclosed methods.

[0183] As used herein, the term “comprising” is intended to mean that the compositions and methods include the recited elements, but not excluding others. “Consisting essentially of’ when used to define compositions and methods, shall mean excluding other elements of any essential significance to the composition or method. “Consisting of’ shall mean excluding more than trace elements of other ingredients for claimed compositions and substantial method steps. Embodiments defined by each of these transition terms are within the scope of this disclosure. Accordingly, it is intended that the methods and compositions can include additional steps and components (comprising) or alternatively including steps and compositions of no significance (consisting essentially of) or alternatively, intending only the stated method steps or compositions (consisting of).

[0184] The drugs or treatments disclosed herein may be first line, second line or third line therapies. The phrase “first line” or “second line” or “third line” refers to the order of treatment received by a patient. First line therapy regimens are treatments given first, whereas second or third line therapy are given after the first line therapy or after the second line therapy, respectively. The National Cancer Institute defines first line therapy as “the first treatment for a disease or condition.

[0185] The term “allele,” which is used interchangeably herein with “allelic variant” refers to alternative forms of a gene or portions thereof. Alleles occupy the same locus or position on homologous chromosomes. When a subject has two identical alleles of a gene, the subject issaid to be homozygous for the gene or allele. When a subject has two different alleles of a gene, the subject is said to be heterozygous for the gene. Alleles of a specific gene can differ from each other in a single nucleotide, or several nucleotides, and can include substitutions, deletions and insertions of nucleotides. An allele of a gene can also be a form of a gene containing a mutation.

[0186] As used herein, the term “determining the genotype of a cell or tissue sample” intends to identify the genotypes of polymorphic loci of interest in the cell or tissue sample. In one aspect, a polymorphic locus is a single nucleotide polymorphic (SNP) locus. If the allelic composition of a SNP locus is heterozygous, the genotype of the SNP locus will be identified as “X / Y” wherein X and Y are two different nucleotides, e.g., A / G for the rsl 1068551 A / G SNP. If the allelic composition of an SNP locus is heterozygous, the genotype of the SNP locus will be identified as “X / X” wherein X identifies the nucleotide that is present at both alleles, e.g., G / G for the rsl 1068551 A / G SNP.

[0187] The term “genetic marker” refers to an allelic variant of a polymorphic region of a gene of interest and / or the expression level of a gene of interest.

[0188] The term “wild-type allele” refers to an allele of a gene which, when present in two copies in a subject results in a wild-type phenotype. There can be several different wild-type alleles of a specific gene, since certain nucleotide changes in a gene may not affect the phenotype of a subject having two copies of the gene with the nucleotide changes.

[0189] The term “polymorphism” refers to the coexistence of more than one form of a gene or portion thereof. A portion of a gene of which there are at least two different forms, i.e., two different nucleotide sequences, is referred to as a “polymorphic region of a gene.” A polymorphic region can be a single nucleotide, the identity of which differs in different alleles.

[0190] A “polymorphic gene” refers to a gene having at least one polymorphic region.

[0191] The term “genotype” refers to the specific allelic composition of an entire cell or a certain gene and in some aspects a specific polymorphism associated with that gene, whereas the term “phenotype” refers to the detectable outward manifestations of a specific genotype.

[0192] The phrase “amplification of polynucleotides” includes methods such as PCR, ligation amplification (or ligase chain reaction, LCR) and amplification methods. These methods are known and widely practiced in the art. See, e.g., U.S. Pat. Nos. 4,683,195 and4,683,202 and Innis et al., 1990 (for PCR); and Wu, D.Y. et al. (1989) Genomics 4:560-569 (for LCR). In general, the PCR procedure describes a method of gene amplification which is comprised of (i) sequence-specific hybridization of primers to specific genes within a DNA sample (or library), (ii) subsequent amplification involving multiple rounds of annealing, elongation, and denaturation using a DNA polymerase, and (iii) screening the PCR products for a band of the correct size. The primers used are oligonucleotides of sufficient length and appropriate sequence to provide initiation of polymerization, i.e., each primer is specifically designed to be complementary to each strand of the genomic locus to be amplified.

[0193] Reagents and hardware for conducting PCR are commercially available. Primers useful to amplify sequences from a particular gene region are preferably complementary to, and hybridize specifically to sequences in the target region or in its flanking regions. Nucleic acid sequences generated by amplification may be sequenced directly. Alternatively the amplified sequence(s) may be cloned prior to sequence analysis. A method for the direct cloning and sequence analysis of enzymatically amplified genomic segments is known in the art.

[0194] The term “encode” as it is applied to polynucleotides refers to a polynucleotide which is said to “encode” a polypeptide if, in its native state or when manipulated by methods well known to those skilled in the art, it can be transcribed and / or translated to produce the mRNA for the polypeptide and / or a fragment thereof. The antisense strand is the complement of such a nucleic acid, and the encoding sequence can be deduced therefrom.

[0195] The term “isolated” as used herein refers to molecules or biological or cellular materials being substantially free from other materials. In one aspect, the term “isolated” refers to nucleic acid, such as DNA or RNA, or protein or polypeptide, or cell or cellular organelle, or tissue or organ, separated from other DNAs or RNAs, or proteins or polypeptides, or cells or cellular organelles, or tissues or organs, respectively, that are present in the natural source. The term “isolated” also refers to a nucleic acid or peptide that is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA techniques, or chemical precursors or other chemicals when chemically synthesized. Moreover, an “isolated nucleic acid” is meant to include nucleic acid fragments which are not naturally occurring as fragments and would not be found in the natural state. The term “isolated” is also used herein to refer to polypeptides which are isolated from other cellular proteins and is meant to encompass both purified and recombinant polypeptides. The term “isolated” is also used herein to refer to cellsor tissues that are isolated from other cells or tissues and is meant to encompass both cultured and engineered cells or tissues.

[0196] The term “treating” as used herein is intended to encompass curing as well as ameliorating at least one symptom of the condition or disease. For example, in the case of cancer, a response to treatment includes a reduction in cachexia, increase in survival time, elongation in time to tumor progression, reduction in tumor mass, reduction in tumor burden and / or a prolongation in time to tumor metastasis, time to tumor recurrence, tumor response, complete response, partial response, stable disease, progressive disease, progression free survival, overall survival, each as measured by standards set by the National Cancer Institute and the U.S. Food and Drug Administration for the approval of new drugs. See Johnson et al. (2003) J. Clin. Oncol. 21(7): 1404-1411.

[0197] The term “suitable for a therapy” or “suitably treated with a therapy” shall mean that the patient is likely to exhibit one or more desirable clinical outcome as compared to patients having the same disease and receiving the same therapy but possessing a different characteristic that is under consideration for the purpose of the comparison. In one aspect, the characteristic under consideration is a genetic polymorphism or a somatic mutation. In another aspect, the characteristic under consideration is expression level of a gene or a polypeptide. When the disease is chronic liver disease or nonalcoholic steatohepatitis, a more desirable clinical outcome is improved liver function. In another aspect, a more desirable clinical outcome is having normal levels of liver function markers such as alanine aminotransferase (ALT), aspartate aminotransferase (AST), alkaline phosphatase (ALP), gamma-glutamyl transferase (GGT), 5'nucleotidase, total bilirubin, conjugated (direct) bilirubin, unconjugated (indirect)bilirubin, prothrombin time (PT), the international normalized ratio (INR), lactate dehydrogenase, total protein, globulins, and albumin. In another aspect, a more desirable clinical outcome is relatively lower relative risk. In yet another aspect, a more desirable clinical outcome is relatively reduced toxicity or side effects. In some embodiments, more than one clinical outcomes are considered simultaneously. In one such aspect, a patient possessing a characteristic, such as a genotype of a genetic polymorphism, may exhibit more than one more desirable clinical outcomes as compared to patients having the same disease and receiving the same therapy but not possessing the characteristic. As defined herein, the patient is considered suitable for the therapy. In another such aspect, a patient possessing a characteristic may exhibit one or more desirable clinical outcome but simultaneously exhibit one or more less desirable clinical outcome. Theclinical outcomes will then be considered collectively, and a decision as to whether the patient is suitable for the therapy will be made accordingly, taking into account the patient’s specific situation and the relevance of the clinical outcomes.

[0198] The term “blood” refers to blood which includes all components of blood circulating in a subject including, but not limited to, red blood cells, white blood cells, plasma, clotting factors, small proteins, platelets and / or cryoprecipitate. This is typically the type of blood which is donated when a human patient gives blood.

[0199] A “patient” as used herein intends an animal patient, a mammal patient or yet further a human patient. For the purpose of illustration only, a mammal includes but is not limited to a simian, a murine, a bovine, an equine, a porcine or an ovine.

[0200] The term “ / / / silico modeling” refers to use of computational models to predict drug effects and / or health outcomes in different scenarios.

[0201] “Molecular stratified biomarker data” includes at a minimum one of the following types of data: genomic, transcriptomic, metabolomic, or proteomic. These data are obtained by assays specific to each biomarker type.

[0202] Disease phenotypes are based on knowledge of a patient’s disease status (e.g., whether a patient has been diagnosed with diabetes or adiposity , yes / no) at a given point in time. Disease phenotypes may also include quantitative disease traits (e.g., LDL, HDL, and Tg measurements at a given time with respect to disease diagnosis), medication information (e.g., patient has taken or is taking a diabetes drug at a given time with respect to disease outcomes), and demographic information (e.g., age at diagnosis, biological sex, etc.).

[0203] As used herein, a “biological pathway” refers to (1) a bespoke set of genes for which their protein products are known to interact in a biological pathway (e.g., complement pathway, JAK / STAT pathway, MAPK pathway, etc.), or (2) an unsupervised learning approach (e.g., principal component analysis (PCA)) that yields patterns in the data such that clusters of genes comprising a pathway may be identified. In some embodiments, approach (1) is based on canonical pathways identified by literature and external pathway databases, while approach (2) is a data-driven analysis that results in identification of pathways.

[0204] As used herein, “biological pathway activity” is defined by information form the literature or external databases that provide evidence for gene or protein expression indicative of pathway activity (e.g., ‘up regulation’ or ‘down regulation’ of pathway X in disease Y).

[0205] As used herein, a “drug target expression” refers to a protein expression as measured in participants of a large cohort (e.g., UK Biobank), where the protein analyzed is a known drug target or is a target that may be druggable (even if not already drugged). In some embodiments, a biomarker stratifier score (e.g., a polygenic score, a proteomics score, a transcriptomics score, or a somatic mutation score) for “a drug target expression” is computed same as it would be for any other quantitative trait, where the outcome of the model is a quantitative measurement.

[0206] As used herein, a “disease outcome” refers to a clinical event that is the result of having a disease.

[0207] As used herein, the term “drug target” refers to any gene or gene product (e.g., RNA or polypeptide) with implications in an associated disease or disorder. Non-limiting examples include various proteins such as enzymes, oncogenes and their polypeptide products, and cell cycle regulatory genes and their polypeptide products.

[0208] As used herein, the phrase “external data” refers to data from an independent set of individuals or population, unused in molecular biomarker stratifier data score determination. In some embodiments, external data is used to assess the predictive power of the methods described herein. In some embodiments, external data is used to prevent overfitting during molecular biomarker stratifier data score determination. At a minimum, the external data is from a large cohort of participants of an observational study (e.g., UK Biobank or eMERGE). In some embodiments, the external data is divided into training and validation subsets. In some embodiments, estimates of biomarker effect are validated in multiple studies, which may include both observational studies and health system patient cohorts.

[0209] As used herein, the phrase “genetic variant” refers to an alteration, variant or polymorphism in a nucleic acid sample or genome of a subject. Such alteration, variant or polymorphism can be with respect to a reference genome, which may be a reference genome of the species (e.g., for human, hG19 or hG38), the subject or other individual. Variations include one or more single nucleotide variations (SNVs), insertions, deletions, repeats, small insertions, small deletions, small repeats, structural variant junctions, variable length tandem repeats, and / or flanking sequences, copy number variants (CNVs), transversions, gene fusions and other rearrangements are also forms of genetic variation. A variation can be a single nucleotide variation (SNV), insertion or deletion (indel), repeat, copy number variation (CNV), transversion, or a combination thereof.

[0210] As used herein, a “biomarker” or “biological marker” refers to a measurable indicator of some biological state or condition. Biomarkers are often measured and evaluated using biological samples (e.g., blood, urine, stool, sputum, swabbed cells or soft tissues) to examine normal biological processes, pathogenic processes, or pharmacologic responses to a therapeutic intervention. In some embodiments, a biomarker is a genetic biomarker, such as a genotype, a mutation or a single nucleotide polymorphism.

[0211] As used herein, “a genome-wide association study” (GWA study, or GWAS), refers to an observational study of a genome-wide set of genetic variants in different individuals to see if any variant is associated with a trait. In some embodiments, a GWAS study focuses on associations between gene variants (e.g., single-nucleotide polymorphisms (SNPs)) and phenotypic traits (e.g., diseases). In some embodiments, a genetic variant is a somatic genetic variant. In some embodiments, a genetic variant is a germline genetic variant.

[0212] In some embodiments, GWA studies compare the DNA of participants having varying phenotypes for a particular trait or disease. These participants may be people with a disease (cases) and similar people without the disease (controls), or they may be people with different phenotypes for a particular trait, for example blood pressure. This approach is known as phenotype-first, in which the participants are classified first by their clinical manifestation(s), as opposed to genotype-first. Each person gives a sample of DNA, from which millions of genetic variants are determined (e.g., using SNP arrays or sequencing). If there is significant statistical evidence that one type of the variant (one allele) is more frequent in people with the disease, the variant is said to be associated with the disease. The associated SNPs are then considered to mark a region of the human genome that may influence the risk of disease.

[0213] As used herein, the phrase “molecular biomarker stratifier” or “biomarker stratifier” refers to a molecular marker that is different between two distinct states (e.g., healthy versus diseased).

[0214] As used herein, the phrase “molecular biomarker stratifier distribution” or “biomarker stratifier distribution” refers to a subset of data obtained from individual-level cohort data (e.g., data from UK Biobank), where subsets may be defined as the top 10%, 20%, 30%, etc. of the distribution. The threshold that is chosen to define a subset of a molecular biomarker stratifier distribution will depend on the use case. For example, it may be optimal to use a 25th percentile of the NASH molecular biomarker stratifier score (e.g., a polygenic score) to define a subset of patients who are predicted to have an outsized clinical benefit from a giventherapy. Alternatively, a 33rd percentile of the NASH molecular biomarker stratifier score may be used to define a subset of patients who would benefit most from a second, different given therapy. In other disease settings with different molecular biomarker stratifier scores, the thresholds may also vary. Defining the optimal threshold will depend on factors such as: estimated effect size for the subset vs. all-comers on therapy, and screen failure rate (a “screen failure” is candidate who undergoes screening - meaning they are checked for eligibility - but who does not meet the clinical trial’s inclusion / exclusion criteria in a trial. A “screen failure rate” refers to the number of ineligible candidates divided by the number of screened candidates). If the threshold is set too conservatively, e.g., 5%, a lot more patients will have to be screened to enroll a few eligible candidates.

[0215] As used herein, the phrase “molecular biomarker stratifier effect” or “biomarker stratifier effect” refers to an estimated magnitude of the effect of a biomarker on disease risk or continuous disease trait from a statistical model.

[0216] As used herein, the phrase “molecular biomarker stratifier score” or “biomarker stratifier score” refers to a score that is cumulative of data from across a plurality (hundreds, thousands, tens of thousands or more) of possible biomarker stratifiers that could be used to predict an individual’s risk for an illness. In some embodiments, a statistical model is trained using summary statistics from a published study on large-scale cohorts or from a metaanalysis of multiple studies based on large and diverse set of cohorts. Then, the estimated weights from that model are applied to a dataset that is orthogonal to the ones used for training (e.g., UK Biobank). For example, a Bayesian high-dimensional linear regression model for a set of populations may be used, where body mass index (BMI) is the dependent variable, and a genome-wide set of SNPs and their effect sizes from a published study of BMI from the GIANT Consortium (Locke et al. 2015) serve as inputs. Population-specific reference panels (e.g., 1000 Genomes) are used to infer linkage disequilibrium, and population-specific PRS are derived. A linear regression of the normalized PRS for each population is then fit, and can be applied to the individual-level genotype and phenotype data in the UK Biobank.

[0217] In some embodiments, molecular biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of a polygenic score.

[0218] In some embodiments, molecular biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant and the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of a proteomics score.

[0219] In some embodiments, molecular biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant and the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of a transcriptomics risk score.

[0220] In some embodiments, molecular biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant and the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational risk score.

[0221] As used herein, the term “pharmacomimetic” refers to showing a similar response or a similar lack of a response to a therapeutic agent or a class of therapeutic agents.

[0222] As used herein, the phrase “pharmacomimetic variant interaction” refers to a plurality of biomarker stratifiers that in combination correlate with showing a similar response or a similar lack of a response to a therapeutic agent or a class of therapeutic agents.

[0223] As used herein, the phrase “pharmacomimetic instruments” refers to biomarker stratifiers that modulate the function or expression of a drug target gene so that the biomarker stratifiers’ effect on human phenotypes is likely to be predictive of the drug’s effect on human phenotypes.

[0224] As used herein, a “pharmacomimetic genetic score” is a score that indicates the likelihood of a patient to respond to a particular drug or a particular class of drugs. In some embodiments, for a given gene, more than one pharmacomimetic variant (defined as a functional variant in gene that is strongly and significantly associated with a given disease) may be identified. Instead of analyzing the variants of this gene separately, their effects are combined as a “pharmacomimetic genetic score.”

[0225] As used herein, the phrase “polygenic score” or “polygenic risk score” or “PGS” or “PRS” refers to a metric that summarizes the estimated effect of many genetic variants on an individual’s phenotype, typically calculated as a weighted sum of trait-associated alleles. In some embodiments, a polygenic score reflects an individual’s estimated geneticpredisposition for a given trait and can be used as a predictor for that trait. In some embodiments, a polygenic score gives an estimate of how likely an individual is to have a given trait only based on genetics, without taking environmental factors into account.

[0226] As used herein, the phrase “proteomics score” refers to a metric that summarizes the estimated effect of many differences in the protein composition (e.g., differences in the amount of proteins, mutational variants of proteins; posttranslational modification of proteins) on an individual’s phenotype, typically calculated as a weighted sum of trait- associated differences. In some embodiments, a proteomics score can be used as a predictor for a given trait. In some embodiments, a proteomics score is used to identify patients most likely to receive an outsized benefit from a particular therapy. In some embodiments, a protein composition of a subject is determined by mass spectroscopy. In some embodiments, a “proteomics score” is estimated from a regression model, e.g., least absolute shrinkage and selection operator (LASSO) (which implements variable selection and regularization for optimal prediction accuracy) for protein measurements vs. outcome, followed by cross- validation procedures.

[0227] As used herein, the phrase “transcriptomics score” refers to a metric that summarizes the estimated effect of many differences in the RNA transcript composition on an individual’s phenotype, typically calculated as a weighted sum of trait-associated differences. In some embodiments, a transcriptomics score refers to a linear or non-linear combination of transcript abundance values that associate with a disease or clinical phenotype. In some embodiments, a transcriptomics score can be used as a predictor for a given trait. In some embodiments, an RNA transcript composition of a subject is determined by RNA-seq, or microarray.

[0228] As used herein, the phrase “somatic mutational risk score” refers to a metric that summarizes the estimated effect of somatic mutations on an individual’s phenotype, typically calculated as a weighted sum of trait-associated somatic mutations. In some embodiments, a somatic mutational risk score can be used as a predictor for a given trait. In some embodiments, a somatic mutational composition of a subject is determined by next generation (high-throughput) sequencing.

[0229] As used herein, the phrase “Principal Component Analysis” (PCA) refers to a technique for analyzing large datasets containing a high number of dimensions / features per observation, increasing the interpretability of data while preserving the maximum amount ofinformation, and enabling the visualization of multidimensional data. In some embodiments, PCA refers to a statistical technique for reducing the dimensionality of a dataset. In some embodiments, reduction in the dimensionality is accomplished by linearly transforming the data into a new coordinate system where the variation in the data can be described with fewer dimensions than the initial data. As used herein, “principal components” are a set of new variables that are combinations of the original variables in the original large datasets.

[0230] As used herein, a “statistical interaction” refers to the coefficient of a predictor defined by the product of the pharmacomimetic genetic score and biomarker score in a regression model to predict a phenotype relevant to the drug indication. Specifically, if the p- value for a test of the hypothesis that the coefficient for this PGS-biomarker product term is non-zero is statistically significant (p < 0.05 or p < 0.05 / # of tests performed if multiple hypotheses are tested), then the PGS and biomarker are considered to have a “statistical interaction”. If the sign of the coefficient for the PGS-biomarker product term is the same as the sign of the marginal biomarker term, that indicates that the predictive and / or causal effects of the biomarker on the phenotype are amplified in subjects with high PGS.

[0231] As used herein an “association analysis” refers to a process of searching for hidden association or pattern in a large dataset. In some embodiments, an association analysis is carried out using a statistical model such as linear regression, logistic regression, or Cox Proportional Hazards model, e.g., where the dependent variable is a clinical trait or disease phenotype assessed at a single time point, or censored survival time to disease onset or event, and the independent variables include weighted effects of biomarker stratifiers, demographic information, and other clinical features.

[0232] As used herein, a “phenotype intermediate” refers to a demonstration of partial or incomplete dominance between two or more genes (e.g., two alleles may produce an intermediate phenotype when both are present, rather than one fully determining the phenotype).

[0233] As used herein, the term “HSD17B13” refers to “Hydroxysteroid 17-Beta Dehydrogenase 13” enzyme (UniProtKB / Swiss-Prot: Q7Z5P4). A pharmacomimetic genetic score “associated with a response to a drug that targets HSD17B13” is used to determine whether the subject will respond to an HSD17B13 inhibitor (e.g., BI-3231 as described in Thamm, Sven, et al., Journal of Medicinal Chemistry 66.4 (2023): 2832-2850, which is incorporated herein in its entirety).

[0234] As used herein, “body mass index” or “BMI” refers to a person’s weight in kilograms divided by the square of height in meters. In some embodiments, a BMI less than 18.5 means underweight, a BMI between 18.5 and <25 means healthy weight, a BMI between 25.0 to <30 means overweight, and a BMI 30.0 or higher means obese.

[0235] As used herein, “diabetes” refers to a group of diseases diagnosed by two independent measurements of fasting blood plasma sugar levels of 126 mg / dL or higher in a patient.

[0236] As used herein, the term “adiposity” refers to body fat. Different fat depots can be found in the body independently of central obesity. Subcutaneous adiposity refers to the accumulation of fat underneath the skin, in the adipose tissue layer. Visceral adiposity refers to the accumulation of fat in the abdominal cavity, specifically around the organs such as the liver, pancreas, and intestines. Ectopic fat refers to the accumulation of fat in areas where it is not normally found, such as the liver, muscle, and pancreas. Both types of fat are associated with increased risk of metabolic disorders such as insulin resistance, type 2 diabetes, and cardiovascular disease.

[0237] As used herein, “Nonalcoholic steatohepatitis” or “NASH” refers to the ectopic accumulation of fat in the liver (hepatic steatosis) when no other causes of secondary liver fat accumulation are present and there is evidence of inflammatory activity and hepatocyte injury in a steatotic liver tissue.

[0238] As used herein, “chronic liver disease” or “CLD” refers to a progressive deterioration of liver functions for more than six months. This is a continuous process of inflammation, destruction, and regeneration of liver parenchyma leading to fibrosis and cirrhosis. Cirrhosis is a final stage of chronic liver disease that results in disruption of liver architecture, the formation of widespread nodules, vascular reorganization, neo-angiogenesis, and deposition of an extracellular matrix.

[0239] As used herein, an “HSD17B13 inhibitor” refers to any agent capable of specifically inhibiting the expression or activity of HSD17B13 RNA and / or HSD17B13 protein at the molecular level. In some embodiments, inhibitors of HSD17B13 include nucleic acids (including antisense compounds), peptides, antibodies, small molecules and other agents capable of inhibiting the expression of HSD17B13 RNA and / or HSD17B13 protein. In some embodiments, an HSD17B13 inhibitor comprises BI-3231. In some embodiments, an HSD17B13 inhibitor comprises an oligonucleotide inhibitor selected from antisense- oligonucleotide, siRNA, shRNA, antisense oligonucleotide targeting a mRNA, non-codingRNA (ncRNA), miRNA and long non-coding RNA (IncRNA); or a protein or nucleic acid aptamer. In some embodiments, the oligonucleotide inhibitor is as described in US11702700B2, which is incorporated herein in its entirety.

[0240] As used herein “PNPLA3” refers to “Patatin-Like Phospholipase Domain Containing 3” gene. (Genecards: PNPLA3 HGNC: 18590 NCBI Gene: 80339 Ensembl: ENSG00000100344 OMIM®: 609567 UniProtKB / Swiss-Prot: Q9NST1).

[0241] As used herein, a “PNPLA3 degrader” or “PNPLA3 inhibitor” refers to any agent capable of specifically inhibiting the expression or activity of PNPLA3 RNA and / or PNPLA3 protein at the molecular level. In some embodiments, inhibitors of PNPLA3 include nucleic acids (including antisense compounds), peptides, antibodies, small molecules and other agents capable of inhibiting the expression of PNPLA3 RNA and / or PNPLA3 protein. In some embodiments, a PNPLA3 inhibitor comprises an oligonucleotide inhibitor selected from antisense-oligonucleotide, siRNA, shRNA, antisense oligonucleotide targeting a mRNA, non-coding RNA (ncRNA), miRNA and long non-coding RNA (IncRNA); or a protein or nucleic acid aptamer. In some embodiments, the oligonucleotide inhibitor is as described in US 10597661B2, which is incorporated herein in its entirety.Modes For Carrying Out the Disclosure

[0242] Applicant has developed a method of analysis and a knowledge base for identifying individuals with subsets of chronic liver disease who are predicted to enjoy greater benefit and / or lower risk from specific therapeutic mechanisms. Specifically, the method identifies combinations of drug targets and disease indications where it is predicted that patients with subsets of polygenic scores (PGS) for the disease, its associated risk factors, or an aggregate of several genetically-driven biological pathways, will receive greater benefit and lower risk from a drug compared to patients with a low PGS.

[0243] Potential general applications of the method and knowledgebase are:(1) Selection of mechanisms and design of prospective clinical trials of investigational drugs that are predicted to yield greater benefit in individuals with elevated PGS. These trials may be run with many-fold fewer patient years required to demonstrate clinical benefit, given the expectation of magnified event rate and treatment response rates. This application is called “polygenic therapeutic development.”(2) Identification of novel therapeutic mechanisms whose benefit is mostly or only apparent in subsets of common disease patients who have elevated polygenic risk. This enables novel target identification and novel therapeutic discovery. This application is called “polygenic therapeutic target ID.”(3) Selection of patients from large common disease treatment-eligible populations who will benefit most from existing drug therapies, and, conversely, identifying patients who are not likely to receive clinically meaningful benefit. This enhances the pharmaco-economic profile of existing therapeutic mechanisms, yielding more cost-effective utilization. This application is called “polygenic pharmacogenomics.”

[0244] The disclosure further provides diagnostic, prognostic and therapeutic methods, which are based, at least in part, on determination of the identify of a genotype of interest identified herein.

[0245] For example, information obtained using the diagnostic assays described herein is useful for determining if a subject is suitable for cancer treatment of a given type. Based on the prognostic information, a doctor can recommend a therapeutic protocol, useful for improving liver function and treating CLD or NASH in the individual.

[0246] A patient’s likely clinical outcome following a clinical procedure such as a therapy or surgery can be expressed in relative terms. For example, a patient having a particular genotype or expression level may experience relatively longer overall survival than a patient or patients not having the genotype or expression level. The patient having the particular genotype or expression level, alternatively, can be considered as likely to survive. Similarly, a patient having a particular genotype or expression level may experience relatively longer progression free survival, than a patient or patients not having the genotype or expression level. The patient having the particular genotype or expression level, alternatively, can be considered as not likely to suffer liver disease progression. Yet in another example, a patient having a particular genotype or expression level may experience relatively more complete response or partial response than a patient or patients not having the genotype or expression level. The patient having the particular genotype or expression level, alternatively, can be considered as likely to respond.

[0247] It is to be understood that information obtained using the diagnostic assays described herein may be used alone or in combination with other information, such as, but not limited to, genotypes or expression levels of other genes, clinical chemical parameters, histopathologicalparameters, or age, gender and weight of the subject. When used alone, the information obtained using the diagnostic assays described herein is useful in determining or identifying the clinical outcome of a treatment, selecting a patient for a treatment, or treating a patient, etc. When used in combination with other information, on the other hand, the information obtained using the diagnostic assays described herein is useful in aiding in the determination or identification of clinical outcome of a treatment, aiding in the selection of a patient for a treatment, or aiding in the treatment of a patient and etc. In a particular aspect, the genotypes or expression levels of one or more genes as disclosed herein are used in a panel of genes, each of which contributes to the final diagnosis, prognosis or treatment.

[0248] The methods of this disclosure are useful for the diagnosis, prognosis and treatment of patients suffering from a liver disease, e.g., CLD or NASH.

[0249] The methods are useful in the assistance of an animal, a mammal or yet further a human patient. For the purpose of illustration only, a mammal includes but is not limited to a human, a simian, a murine, a bovine, an equine, a porcine or an ovine.Diagnostic Methods

[0250] Provided in one aspect is a method for aiding in the selection of or selecting or not selecting a CLD or NASH patient for a therapy.

[0251] Further provided is a method for aiding in the determination of or determining whether or not a CLD or NASH patient is suitable for a therapy comprising an HSD17B13 inhibitor, a PNPLA3 degrader, or a combination thereof.

[0252] Still further provided is a method for aiding in the determination of or determining whether a CLD or NASH patient is likely to experience a longer or shorter progression free survival following a therapy comprising an HSD17B13 inhibitor, a PNPLA3 degrader, or a combination thereof.

[0253] In one aspect of any of the above methods, the therapy comprises, or alternatively consists essentially of, or consists of, administration of an HSD17B13 inhibitor, a PNPLA3 degrader, or a combination thereof.

[0254] In one aspect of any of the above methods, the patient suffers from CLD or NASHc.

[0255] Suitable patient samples in the methods include, but are not limited to a sample comprises, or alternatively consisting essentially of, or yet further consisting of, at least one of blood, plasma, a blood cell, a peripheral blood lymphocyte, a liver cell, or combinationsthereof. The samples can be at least one of an original sample recently isolated from the patient, a fixed tissue, a frozen tissue, a biopsy tissue, a resection tissue, a microdissected tissue, or combinations thereof.

[0256] Any suitable method for identifying the genotype in the patient sample can be used and the disclosures described herein are not to be limited to these methods. For the purpose of illustration only, the genotype is determined by a method comprising, or alternatively consisting essentially of, or yet further consisting of, Sanger sequencing, next-gen sequencing, hybridization, PCR or more specifically, PCR-RFLP or microarray. These methods as well as equivalents or alternatives thereto are described herein.

[0257] The methods are useful in the assistance of an animal, a mammal or yet further a human patient. For the purpose of illustration only, a mammal includes but is not limited to a human, a simian, a murine, a bovine, an equine, a porcine or an ovine.Polymorphic Region

[0258] For example, information obtained using the diagnostic assays described herein is useful for determining if a subject will likely, more likely, or less likely to respond to a treatment of a given type. Based on the prognostic information, a doctor can recommend a therapeutic protocol, useful for treating improving liver function in the individual.

[0259] In addition, knowledge of the identity of a particular allele in an individual (the gene profile) allows customization of therapy for a particular disease to the individual’s genetic profile, the goal of “pharmacogenomics”. For example, an individual’s genetic profile can enable a doctor: 1) to more effectively prescribe a drug that will address the molecular basis of the disease or condition; 2) to better determine the appropriate dosage of a particular drug and 3) to identify novel targets for drug development. The identity of the genotype or expression patterns of individual patients can then be compared to the genotype or expression profile of the disease to determine the appropriate drug and dose to administer to the patient.

[0260] The ability to target populations expected to show the highest clinical benefit, based on the normal or disease genetic profile, can enable: 1) the repositioning of marketed drugs with disappointing market results; 2) the rescue of drug candidates whose clinical development has been discontinued as a result of safety or efficacy limitations, which are patient subgroup-specific; and 3) an accelerated and less costly development for drug candidates and more optimal drug labeling.

[0261] Detection of point mutations or additional base pair repeats can be accomplished by molecular cloning of the specified allele and subsequent sequencing of that allele using techniques known in the art, in some aspects, after isolation of a suitable nucleic acid sample using methods known in the art. Alternatively, the gene sequences can be amplified directly from a genomic DNA preparation from the tumor tissue using PCR, and the sequence composition is determined from the amplified product. As described more fully below, numerous methods are available for isolating and analyzing a subject’s DNA for mutations at a given genetic locus such as the gene of interest.

[0262] A detection method is allele specific hybridization using probes overlapping the polymorphic site and having about 5, or alternatively 10, or alternatively 20, or alternatively 25, or alternatively 30 nucleotides around the polymorphic region. In another embodiment of the disclosure, several probes capable of hybridizing specifically to the allelic variant are attached to a solid phase support, e.g., a “chip”. Oligonucleotides can be bound to a solid support by a variety of processes, including lithography. For example, a chip can hold up to 250,000 oligonucleotides (GeneChip, Affymetrix). Mutation detection analysis using these chips comprising oligonucleotides, also termed “DNA probe arrays” is described e.g., in Cronin et al. (1996) Human Mutation 7:244.

[0263] In other detection methods, it is necessary to first amplify at least a portion of the gene of interest prior to identifying the allelic variant. Amplification can be performed, e.g., by PCR and / or LCR, according to methods known in the art. In one embodiment, genomic DNA of a cell is exposed to two PCR primers and amplification for a number of cycles sufficient to produce the required amount of amplified DNA.

[0264] Alternative amplification methods include: self-sustained sequence replication (Guatelli et al. (1990) Proc. Natl. Acad. Sci. USA 87: 1874-1878), transcriptional amplification system (Kwoh et al. (1989) Proc. Natl. Acad. Sci. USA 86: 1173-1177), Q-Beta Replicase (Lizardi et al. (1988) Bio / Technology 6: 1197), or any other nucleic acid amplification method, followed by the detection of the amplified molecules using techniques known to those of skill in the art. These detection schemes are useful for the detection of nucleic acid molecules if such molecules are present in very low numbers.

[0265] In one embodiment, any of a variety of sequencing reactions known in the art can be used to directly sequence at least a portion of the gene of interest and detect allelic variants, e.g., mutations, by comparing the sequence of the sample sequence with the correspondingwild-type (control) sequence. Exemplary sequencing reactions include those based on techniques developed by Maxam and Gilbert (1997) Proc. Natl. Acad. Sci, USA 74:560) or Sanger et al. (1977) Proc. Nat. Acad. Sci, 74:5463). It is also contemplated that any of a variety of automated sequencing procedures can be utilized when performing the subject assays (Biotechniques (1995) 19:448), including sequencing by mass spectrometry (see, for example, U.S. Patent No. 5,547,835 and International Patent Application Publication Number WO 94 / 16101, entitled DNA Sequencing by Mass Spectrometry by Koster; U.S. Patent No. 5,547,835 and international patent application Publication Number WO 94 / 21822 entitled “DNA Sequencing by Mass Spectrometry Via Exonuclease Degradation” by Koster; U.S. Patent No. 5,605,798 and International Patent Application No. PCT / US96 / 03651 entitled DNA Diagnostics Based on Mass Spectrometry by Koster; Cohen et al. (1996) Adv.Chromat. 36: 127-162; and Griffin et al. (1993) Appl. Biochem. Bio. 38: 147-159). It will be evident to one skilled in the art that, for certain embodiments, the occurrence of only one, two or three of the nucleic acid bases need be determined in the sequencing reaction. For instance, A-track or the like, e.g., where only one nucleotide is detected, can be carried out.

[0266] Yet other sequencing methods are disclosed, e.g., in U.S. Patent No. 5,580,732 entitled “Method of DNA Sequencing Employing A Mixed DNA-Polymer Chain Probe” and U.S. Patent No. 5,571,676 entitled “Method For Mismatch-Directed In Vitro DNA Sequencing.”

[0267] In some cases, the presence of the specific allele in DNA from a subject can be shown by restriction enzyme analysis. For example, the specific nucleotide polymorphism can result in a nucleotide sequence comprising a restriction site which is absent from the nucleotide sequence of another allelic variant.

[0268] In a further embodiment, protection from cleavage agents (such as a nuclease, hydroxylamine or osmium tetroxide and with piperidine) can be used to detect mismatched bases in RNA / RNA DNA / DNA, or RNA / DNA heteroduplexes (see, e.g., Myers et al. (1985) Science 230: 1242). In general, the technique of “mismatch cleavage” starts by providing heteroduplexes formed by hybridizing a control nucleic acid, which is optionally labeled, e.g., RNA or DNA, comprising a nucleotide sequence of the allelic variant of the gene of interest with a sample nucleic acid, e.g., RNA or DNA, obtained from a tissue sample. The double-stranded duplexes are treated with an agent which cleaves single-stranded regions of the duplex such as duplexes formed based on base pair mismatches between the control andsample strands. For instance, RNA / DNA duplexes can be treated with RNase and DNA / DNA hybrids treated with SI nuclease to enzymatically digest the mismatched regions. In other embodiments, either DNA / DNA or RNA / DNA duplexes can be treated with hydroxylamine or osmium tetroxide and with piperidine in order to digest mismatched regions. After digestion of the mismatched regions, the resulting material is then separated by size on denaturing polyacrylamide gels to determine whether the control and sample nucleic acids have an identical nucleotide sequence or in which nucleotides they are different. See, for example, U.S. Patent No. 6,455,249, Cotton et al. (1988) Proc. Natl. Acad. Sci. USA 85:4397; Saleeba et al. (1992) Methods Enzy. 217:286-295. In another embodiment, the control or sample nucleic acid is labeled for detection.

[0269] In other embodiments, alterations in electrophoretic mobility are used to identify the particular allelic variant. For example, single strand conformation polymorphism (SSCP) may be used to detect differences in electrophoretic mobility between mutant and wild type nucleic acids (Orita et al. (1989) Proc. Natl. Acad. Sci USA 86:2766; Cotton (1993) Mutat. Res. 285: 125-144 and Hayashi (1992) Genet Anal Tech. Appl. 9:73-79). Single-stranded DNA fragments of sample and control nucleic acids are denatured and allowed to renature. The secondary structure of single-stranded nucleic acids varies according to sequence, the resulting alteration in electrophoretic mobility enables the detection of even a single base change. The DNA fragments may be labeled or detected with labeled probes. The sensitivity of the assay may be enhanced by using RNA (rather than DNA), in which the secondary structure is more sensitive to a change in sequence. In another preferred embodiment, the subject method utilizes heteroduplex analysis to separate double stranded heteroduplex molecules on the basis of changes in electrophoretic mobility (Keen et al. (1991) Trends Genet. 7:5).

[0270] In yet another embodiment, the identity of the allelic variant is obtained by analyzing the movement of a nucleic acid comprising the polymorphic region in polyacrylamide gels containing a gradient of denaturant, which is assayed using denaturing gradient gel electrophoresis (DGGE) (Myers et al. (1985) Nature 313 :495). When DGGE is used as the method of analysis, DNA will be modified to ensure that it does not completely denature, for example by adding a GC clamp of approximately 40 bp of high-melting GC-rich DNA by PCR. In a further embodiment, a temperature gradient is used in place of a denaturing agent gradient to identify differences in the mobility of control and sample DNA (Rosenbaum and Reissner (1987) Biophys. Chem. 265: 1275).

[0271] Examples of techniques for detecting differences of at least one nucleotide between 2 nucleic acids include, but are not limited to, selective oligonucleotide hybridization, selective amplification, or selective primer extension. For example, oligonucleotide probes may be prepared in which the known polymorphic nucleotide is placed centrally (allele-specific probes) and then hybridized to target DNA under conditions which permit hybridization only if a perfect match is found (Saiki et al. (1986) Nature 324: 163); Saiki et al. (1989) Proc. Natl. Acad. Sci. USA 86:6230 and Wallace et al. (1979) Nucl. Acids Res. 6:3543). Such allele specific oligonucleotide hybridization techniques may be used for the detection of the nucleotide changes in the polymorphic region of the gene of interest. For example, oligonucleotides having the nucleotide sequence of the specific allelic variant are attached to a hybridizing membrane and this membrane is then hybridized with labeled sample nucleic acid. Analysis of the hybridization signal will then reveal the identity of the nucleotides of the sample nucleic acid.

[0272] Alternatively, allele specific amplification technology which depends on selective PCR amplification may be used in conjunction with the instant disclosure. Oligonucleotides used as primers for specific amplification may carry the allelic variant of interest in the center of the molecule (so that amplification depends on differential hybridization) (Gibbs et al. (1989) Nucleic Acids Res. 17:2437-2448) or at the extreme 3’ end of one primer where, under appropriate conditions, mismatch can prevent, or reduce polymerase extension (Prossner (1993) Tibtech 11 :238 and Newton et al. (1989) Nucl. Acids Res. 17:2503). This technique is also termed “PROBE” for Probe Oligo Base Extension. In addition, it may be desirable to introduce a novel restriction site in the region of the mutation to create cleavagebased detection (Gasparini et al. (1992) Mol. Cell Probes 6: 1).

[0273] In another embodiment, identification of the allelic variant is carried out using an oligonucleotide ligation assay (OLA), as described, e.g., in U.S. Patent No. 4,998,617 and in Landegren et al. (1988) Science 241 : 1077-1080. The OLA protocol uses two oligonucleotides which are designed to be capable of hybridizing to abutting sequences of a single strand of a target. One of the oligonucleotides is linked to a separation marker, e.g., biotinylated, and the other is detectably labeled. If the precise complementary sequence is found in a target molecule, the oligonucleotides will hybridize such that their termini abut, and create a ligation substrate. Ligation then permits the labeled oligonucleotide to be recovered using avidin, or another biotin ligand. Nickerson et al. have described a nucleic acid detection assay that combines attributes of PCR and OLA (Nickerson et al. (1990) Proc.Natl. Acad. Sci. (U.S.A.) 87:8923-8927). In this method, PCR is used to achieve the exponential amplification of target DNA, which is then detected using OLA.

[0274] Several techniques based on this OLA method have been developed and can be used to detect the specific allelic variant of the polymorphic region of the gene of interest. For example, U.S. Patent No. 5,593,826 discloses an OLA using an oligonucleotide having 3’- amino group and a 5 ’-phosphorylated oligonucleotide to form a conjugate having a phosphoramidate linkage. In another variation of OLA described in Tobe et al. (1996) Nucleic Acids Res. 24: 3728, OLA combined with PCR permits typing of two alleles in a single microtiter well. By marking each of the allele-specific primers with a unique hapten, i.e. digoxigenin and fluorescein, each OLA reaction can be detected by using hapten specific antibodies that are labeled with different enzyme reporters, alkaline phosphatase or horseradish peroxidase. This system permits the detection of the two alleles using a high throughput format that leads to the production of two different colors.

[0275] In one embodiment, the single base polymorphism can be detected by using a specialized exonuclease-resistant nucleotide, as disclosed, e.g., in Mundy, C. R. (U.S. Patent No. 4,656,127). According to the method, a primer complementary to the allelic sequence immediately 3’ to the polymorphic site is permitted to hybridize to a target molecule obtained from a particular animal or human. If the polymorphic site on the target molecule contains a nucleotide that is complementary to the particular exonuclease-resistant nucleotide derivative present, then that derivative will be incorporated onto the end of the hybridized primer. Such incorporation renders the primer resistant to exonuclease, and thereby permits its detection. Since the identity of the exonuclease-resistant derivative of the sample is known, a finding that the primer has become resistant to exonucleases reveals that the nucleotide present in the polymorphic site of the target molecule was complementary to that of the nucleotide derivative used in the reaction. This method has the advantage that it does not require the determination of large amounts of extraneous sequence data.

[0276] In another embodiment of the disclosure, a solution-based method is used for determining the identity of the nucleotide of the polymorphic site. Cohen, D. et al. (French Patent 2,650,840; PCT Appln. No. W091 / 02087). As in the Mundy method of U.S. Patent No. 4,656,127, a primer is employed that is complementary to allelic sequences immediately 3’ to a polymorphic site. The method determines the identity of the nucleotide of that siteusing labeled dideoxynucleotide derivatives, which, if complementary to the nucleotide of the polymorphic site will become incorporated onto the terminus of the primer.

[0277] An alternative method, known as Genetic Bit Analysis or GBA™ is described by Goelet, P. et al. (PCT Appln. No. 92 / 15712). This method uses mixtures of labeled terminators and a primer that is complementary to the sequence 3’ to a polymorphic site. The labeled terminator that is incorporated is thus determined by, and complementary to, the nucleotide present in the polymorphic site of the target molecule being evaluated. In contrast to the method of Cohen et al. (French Patent 2,650,840; PCT Appln. No. W091 / 02087) the method of Goelet, P. et al. supra, is preferably a heterogeneous phase assay, in which the primer or the target molecule is immobilized to a solid phase.

[0278] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, J. S. et al. (1989) Nucl. Acids. Res. 17:7779- 7784; Sokolov, B. P. (1990) Nucl. Acids Res. 18:3671; Syvanen, A.-C. et al. (1990) Genomics 8:684-692; Kuppuswamy, M. N. et al. (1991) Proc. Natl. Acad. Sci. (U.S.A.) 88: 1143-1147; Prezant, T. R. et al. (1992) Hum. Mutat. 1 : 159-164; Ugozzoli, L. et al. (1992) GATA 9: 107-112; Nyren, P. et al. (1993) Anal. Biochem. 208: 171-175). These methods differ from GBA™ in that they all rely on the incorporation of labeled deoxynucleotides to discriminate between bases at a polymorphic site. In such a format, since the signal is proportional to the number of deoxynucleotides incorporated, polymorphisms that occur in runs of the same nucleotide can result in signals that are proportional to the length of the run (Syvanen, A.-C. et al. (1993) Amer. J. Hum. Genet. 52:46-59).

[0279] If the polymorphic region is located in the coding region of the gene of interest, yet other methods than those described above can be used for determining the identity of the allelic variant. For example, identification of the allelic variant, which encodes a mutated signal peptide, can be performed by using an antibody specifically recognizing the mutant protein in, e.g., immunohistochemistry or immunoprecipitation. Antibodies to the wild-type or signal peptide mutated forms of the signal peptide proteins can be prepared according to methods known in the art.

[0280] Often a solid phase support is used as a support capable of binding of a primer, probe, polynucleotide, an antigen or an antibody. Well-known supports include glass, polystyrene, polypropylene, polyethylene, dextran, nylon, amylases, natural and modified celluloses, polyacrylamides, gabbros, and magnetite. The nature of the support can be either soluble tosome extent or insoluble for the purposes of the present disclosure. The support material may have virtually any possible structural configuration so long as the coupled molecule is capable of binding to an antigen or antibody. Thus, the support configuration may be spherical, as in a bead, or cylindrical, as in the inside surface of a test tube, or the external surface of a rod. Alternatively, the surface may be flat such as a sheet, test strip, etc. or alternatively polystyrene beads. Those skilled in the art will know many other suitable supports for binding antibody or antigen, or will be able to ascertain the same by use of routine experimentation.

[0281] Moreover, it will be understood that any of the above methods for detecting alterations in a gene or gene product or polymorphic variants can be used to monitor the course of treatment or therapy.

[0282] The methods described herein may be performed, for example, by utilizing prepackaged diagnostic kits, such as those described below, comprising at least one probe or primer nucleic acid described herein, which may be conveniently used, e.g., to determine whether a subject is likely to benefit from a therapy as described herein or has or is at risk of developing liver failure or cirrhosis.

[0283] Sample nucleic acid for use in the above-described diagnostic and prognostic methods can be obtained from any suitable cell type or tissue of a subject. For example, a subject’s bodily fluid (e.g., blood) can be obtained by known techniques (e.g., venipuncture).Alternatively, nucleic acid tests can be performed on dry samples (e.g., hair or skin). Diagnostic procedures can also be performed in situ directly upon tissue sections (fixed and / or frozen) of patient tissue obtained from biopsies or resections, such that no nucleic acid purification is necessary. Nucleic acid reagents can be used as probes and / or primers for such in situ procedures (see, for example, Nuovo, G. J. (1992) PCR IN SITU HYBRIDIZATION: PROTOCOLS AND APPLICATIONS, Raven Press, NY).

[0284] In addition to methods which focus primarily on the detection of one nucleic acid sequence, profiles can also be assessed in such detection schemes. Fingerprint profiles can be generated, for example, by utilizing a differential display procedure, Northern analysis and / or RT-PCR.

[0285] Antibodies directed against wild type or mutant peptides encoded by the allelic variants of the gene of interest may also be used in disease diagnostics and prognostics. Such diagnostic methods, may be used to detect abnormalities in the level of expression of thepeptide, or abnormalities in the structure and / or tissue, cellular, or subcellular location of the peptide. Protein from the tissue or cell type to be analyzed may easily be detected or isolated using techniques which are well known to one of skill in the art, including but not limited to Western blot analysis. For a detailed explanation of methods for carrying out Western blot analysis, see Sambrook and Russell (2001) supra. The protein detection and isolation methods employed herein can also be such as those described in Harlow and Lane, (1999) supra. This can be accomplished, for example, by immunofluorescence techniques employing a fluorescently labeled antibody (see below) coupled with light microscopic, flow cytometric, or fluorimetric detection. The antibodies (or fragments thereof) useful in the present disclosure may, additionally, be employed histologically, as in immunofluorescence or immunoelectron microscopy, for in situ detection of the peptides or their allelic variants. In situ detection may be accomplished by removing a histological specimen from a patient, and applying thereto a labeled antibody of the present disclosure. The antibody (or fragment) is preferably applied by overlaying the labeled antibody (or fragment) onto a biological sample. Through the use of such a procedure, it is possible to determine not only the presence of the subject polypeptide, but also its distribution in the examined tissue. Using the present disclosure, one of ordinary skill will readily perceive that any of a wide variety of histological methods (such as staining procedures) can be modified in order to achieve such in situ detection.

[0286] In one embodiment, it is necessary to first amplify at least a portion of the gene of interest prior to identifying the polymorphic region of the gene of interest in a sample. Amplification can be performed, e.g., by PCR and / or LCR, according to methods known in the art. Various non-limiting examples of PCR include the herein described methods.

[0287] Allele-specific PCR is a diagnostic or cloning technique is used to identify or utilize single-nucleotide polymorphisms (SNPs). It requires prior knowledge of a DNA sequence, including differences between alleles, and uses primers whose 3’ ends encompass the SNP. PCR amplification under stringent conditions is much less efficient in the presence of a mismatch between template and primer, so successful amplification with an SNP-specific primer signals presence of the specific SNP in a sequence (See, Saiki et al. (1986) Nature 324(6093): 163-166 and U.S. Patent Nos.: 5,821,062; 7,052,845 or 7,250,258).

[0288] Assembly PCR or Polymerase Cycling Assembly (PCA) is the artificial synthesis of long DNA sequences by performing PCR on a pool of long oligonucleotides with shortoverlapping segments. The oligonucleotides alternate between sense and antisense directions, and the overlapping segments determine the order of the PCR fragments thereby selectively producing the final long DNA product (See, Stemmer et al. (1995) Gene 164(l):49-53 and U.S. Patent Nos.: 6,335,160; 7,058,504 or 7,323,336)

[0289] Asymmetric PCR is used to preferentially amplify one strand of the original DNA more than the other. It finds use in some types of sequencing and hybridization probing where having only one of the two complementary stands is required. PCR is carried out as usual, but with a great excess of the primers for the chosen strand. Due to the slow amplification later in the reaction after the limiting primer has been used up, extra cycles of PCR are required (See, Innis et al. (1988) Proc Natl Acad Sci U.S.A. 85(24):9436-9440 and U.S. Patent Nos.: 5,576,180; 6,106,777 or 7,179,600) A recent modification on this process, known as Linear- After- The-Exponential-PCR (LATE-PCR), uses a limiting primer with a higher melting temperature (Tm) than the excess primer to maintain reaction efficiency as the limiting primer concentration decreases mid-reaction (Pierce et al. (2007) Methods Mol. Med. 132:65-85).

[0290] Colony PCR uses bacterial colonies, for example E. coli, which can be rapidly screened by PCR for correct DNA vector constructs. Selected bacterial colonies are picked with a sterile toothpick and dabbed into the PCR master mix or sterile water. The PCR is started with an extended time at 95 °C when standard polymerase is used or with a shortened denaturation step at 100°C and special chimeric DNA polymerase (Pavlov et al. (2006) “Thermostable DNA Polymerases for a Wide Spectrum of Applications: Comparison of a Robust Hybrid TopoTaq to other enzymes”, in Kieleczawa J: DNA Sequencing II: Optimizing Preparation and Cleanup. Jones and Bartlett, pp. 241-257)

[0291] Helicase-dependent amplification is similar to traditional PCR, but uses a constant temperature rather than cycling through denaturation and annealing / extension cycles. DNA Helicase, an enzyme that unwinds DNA, is used in place of thermal denaturation (See, Myriam et al. (2004) EMBO reports 5(8):795-800 and U.S. Patent No. 7,282,328).

[0292] Hot-start PCR is a technique that reduces non-specific amplification during the initial set up stages of the PCR. The technique may be performed manually by heating the reaction components to the melting temperature (e.g., 95°C) before adding the polymerase (Chou et al. (1992) Nucleic Acids Research 20: 1717-1723 and U.S. Patent Nos.: 5,576,197 and 6,265,169). Specialized enzyme systems have been developed that inhibit the polymerase’sactivity at ambient temperature, either by the binding of an antibody (Sharkey et al. (1994) Bio / Technology 12:506-509) or by the presence of covalently bound inhibitors that only dissociate after a high-temperature activation step. Hot-start / cold-finish PCR is achieved with new hybrid polymerases that are inactive at ambient temperature and are instantly activated at elongation temperature.

[0293] Intersequence-specific (ISSR) PCR method for DNA fingerprinting that amplifies regions between some simple sequence repeats to produce a unique fingerprint of amplified fragment lengths (Zietkiewicz et al. (1994) Genomics 20(2): 176-83).

[0294] Inverse PCR is a method used to allow PCR when only one internal sequence is known. This is especially useful in identifying flanking sequences to various genomic inserts. This involves a series of DNA digestions and self-ligation, resulting in known sequences at either end of the unknown sequence (Ochman et al. (1988) Genetics 120:621-623 and U.S. Patent Nos.: 6,013,486; 6,106,843 or 7,132,587).

[0295] Ligation-mediated PCR uses small DNA linkers ligated to the DNA of interest and multiple primers annealing to the DNA linkers; it has been used for DNA sequencing, genome walking, and DNA footprinting (Mueller et al. (1988) Science 246:780-786).

[0296] Methylation-specific PCR (MSP) is used to detect methylation of CpG islands in genomic DNA (Herman et al. (1996) Proc Natl Acad Sci U.S.A. 93(13):9821-9826 and U.S. Patent Nos.: 6,811,982; 6,835,541 or 7,125,673). DNA is first treated with sodium bisulfite, which converts unmethylated cytosine bases to uracil, which is recognized by PCR primers as thymine. Two PCRs are then carried out on the modified DNA, using primer sets identical except at any CpG islands within the primer sequences. At these points, one primer set recognizes DNA with cytosines to amplify methylated DNA, and one set recognizes DNA with uracil or thymine to amplify unmethylated DNA. MSP using qPCR can also be performed to obtain quantitative rather than qualitative information about methylation.

[0297] Multiplex Ligation-dependent Probe Amplification (MLP A) permits multiple targets to be amplified with only a single primer pair, thus avoiding the resolution limitations of multiplex PCR (see below).

[0298] Multiplex-PCR uses of multiple, unique primer sets within a single PCR mixture to produce amplicons of varying sizes specific to different DNA sequences (See, U.S. Patent Nos.: 5,882,856; 6,531,282 or 7,118,867). By targeting multiple genes at once, additional information may be gained from a single test run that otherwise would require several timesthe reagents and more time to perform. Annealing temperatures for each of the primer sets must be optimized to work correctly within a single reaction, and amplicon sizes, i.e., their base pair length, should be different enough to form distinct bands when visualized by gel electrophoresis.

[0299] Nested PCR increases the specificity of DNA amplification, by reducing background due to non-specific amplification of DNA. Two sets of primers are being used in two successive PCRs. In the first reaction, one pair of primers is used to generate DNA products, which besides the intended target, may still consist of non-specifically amplified DNA fragments. The product(s) are then used in a second PCR with a set of primers whose binding sites are completely or partially different from and located 3’ of each of the primers used in the first reaction (See, U.S. Patent Nos.: 5,994,006; 7,262,030 or 7,329,493). Nested PCR is often more successful in specifically amplifying long DNA fragments than conventional PCR, but it requires more detailed knowledge of the target sequences.

[0300] Overlap-extension PCR is a genetic engineering technique allowing the construction of a DNA sequence with an alteration inserted beyond the limit of the longest practical primer length.

[0301] Quantitative PCR (Q-PCR), also known as RQ-PCR, QRT-PCR and RTQ-PCR, is used to measure the quantity of a PCR product following the reaction or in real-time. See, U.S. Patent Nos.: 6,258,540; 7,101,663 or 7,188,030. Q-PCR is the method of choice to quantitatively measure starting amounts of DNA, cDNA or RNA. Q-PCR is commonly used to determine whether a DNA sequence is present in a sample and the number of its copies in the sample. The method with currently the highest level of accuracy is digital PCR as described in U.S. Patent No. 6,440,705; U.S. Publication No. 2007 / 0202525; Dressman et al. (2003) Proc. Natl. Acad. Sci USA 100(15):8817-8822 and Vogelstein et al. (1999) Proc. Natl. Acad. Sci. USA. 96(16):9236-9241. More commonly, RT-PCR refers to reverse transcription PCR (see below), which is often used in conjunction with Q-PCR. QRT-PCR methods use fluorescent dyes, such as Sybr Green, or fluorophore-containing DNA probes, such as TaqMan, to measure the amount of amplified product in real time.

[0302] Reverse Transcription PCR (RT-PCR) is a method used to amplify, isolate or identify a known sequence from a cellular or tissue RNA (See, U.S. Patent Nos.: 6,759,195;7, 179,600 or 7,317, 111). The PCR is preceded by a reaction using reverse transcriptase to convert RNA to cDNA. RT-PCR is widely used in expression profiling, to determine theexpression of a gene or to identify the sequence of an RNA transcript, including transcription start and termination sites and, if the genomic DNA sequence of a gene is known, to map the location of exons and introns in the gene. The 5’ end of a gene (corresponding to the transcription start site) is typically identified by an RT-PCR method, named Rapid Amplification of cDNA Ends (RACE-PCR).

[0303] Thermal asymmetric interlaced PCR (TAIL-PCR) is used to isolate unknown sequence flanking a known sequence. Within the known sequence TAIL-PCR uses a nested pair of primers with differing annealing temperatures; a degenerate primer is used to amplify in the other direction from the unknown sequence (Liu et al. (1995) Genomics 25(3):674-81).

[0304] Touchdown PCR is a variant of PCR that aims to reduce nonspecific background by gradually lowering the annealing temperature as PCR cycling progresses. The annealing temperature at the initial cycles is usually a few degrees (3-5 °C) above the Tmof the primers used, while at the later cycles, it is a few degrees (3-5 °C) below the primer Tm. The higher temperatures give greater specificity for primer binding, and the lower temperatures permit more efficient amplification from the specific products formed during the initial cycles (Don et al. (1991) Nucl Acids Res 19:4008 and U.S. Patent No. 6,232,063).

[0305] In one embodiment of the disclosure, probes are labeled with two fluorescent dye molecules to form so-called “molecular beacons” (Tyagi, S. and Kramer, F.R. (1996) Nat. Biotechnol. 14:303-8). Such molecular beacons signal binding to a complementary nucleic acid sequence through relief of intramolecular fluorescence quenching between dyes bound to opposing ends on an oligonucleotide probe. The use of molecular beacons for genotyping has been described (Kostrikis, L.G. (1998) Science 279: 1228-9) as has the use of multiple beacons simultaneously (Marras, S.A. (1999) Genet. Anal. 14: 151-6). A quenching molecule is useful with a particular fluorophore if it has sufficient spectral overlap to substantially inhibit fluorescence of the fluorophore when the two are held proximal to one another, such as in a molecular beacon, or when attached to the ends of an oligonucleotide probe from about 1 to about 25 nucleotides.

[0306] Labeled probes also can be used in conjunction with amplification of a gene of interest. (Holland et al. (1991) Proc. Natl. Acad. Sci. 88:7276-7280). U.S. Patent No. 5,210,015 by Gelfand et al. describe fluorescence-based approaches to provide real time measurements of amplification products during PCR. Such approaches have either employed intercalating dyes (such as ethidium bromide) to indicate the amount of double-strandedDNA present, or they have employed probes containing fluorescence-quencher pairs (also referred to as the “Taq-Man” approach) where the probe is cleaved during amplification to release a fluorescent molecule whose concentration is proportional to the amount of doublestranded DNA present. During amplification, the probe is digested by the nuclease activity of a polymerase when hybridized to the target sequence to cause the fluorescent molecule to be separated from the quencher molecule, thereby causing fluorescence from the reporter molecule to appear. The Taq-Man approach uses a probe containing a reporter moleculequencher molecule pair that specifically anneals to a region of a target polynucleotide containing the polymorphism.

[0307] Probes can be affixed to surfaces for use as “gene chips.” Such gene chips can be used to detect genetic variations by a number of techniques known to one of skill in the art. In one technique, oligonucleotides are arrayed on a gene chip for determining the DNA sequence of a by the sequencing by hybridization approach, such as that outlined in U.S. Patent Nos. 6,025,136 and 6,018,041. The probes of the disclosure also can be used for fluorescent detection of a genetic sequence. Such techniques have been described, for example, in U.S. Patent Nos. 5,968,740 and 5,858,659. A probe also can be affixed to an electrode surface for the electrochemical detection of nucleic acid sequences such as described by Kayem et al. U.S. Patent No. 5,952,172 and by Kelley, S.O. et al. (1999) Nucleic Acids Res. 27:4830-4837.

[0308] This disclosure also provides for a prognostic panel of genetic markers selected from, but not limited to the genetic polymorphisms identified herein. The prognostic panel comprises probes or primers or microarrays that can be used to amplify and / or for determining the molecular structure of the polymorphisms identified herein. The probes or primers can be attached or supported by a solid phase support such as, but not limited to a gene chip or microarray. The probes or primers can be detectably labeled. This aspect of the disclosure is a means to identify the genotype of a patient sample for the genes of interest identified above.

[0309] In one aspect, the panel contains the herein identified probes or primers as wells as other probes or primers. In a alternative aspect, the panel includes one or more of the above noted probes or primers and others. In a further aspect, the panel consist only of the abovenoted probes or primers.

[0310] Primers or probes can be affixed to surfaces for use as “gene chips” or “microarray.” Such gene chips or microarrays can be used to detect genetic variations by a number of techniques known to one of skill in the art. In one technique, oligonucleotides are arrayed on a gene chip for determining the DNA sequence of a by the sequencing by hybridization approach, such as that outlined in U.S. Patent Nos. 6,025,136 and 6,018,041. The probes of the disclosure also can be used for fluorescent detection of a genetic sequence. Such techniques have been described, for example, in U.S. Patent Nos. 5,968,740 and 5,858,659. A probe also can be affixed to an electrode surface for the electrochemical detection of nucleic acid sequences such as described by Kayem et al. U.S. Patent No. 5,952,172 and by Kelley et al. (1999) Nucleic Acids Res. 27:4830-4837.

[0311] Various “gene chips” or “microarray” and similar technologies are known in the art. Examples of such include, but are not limited to LabCard (ACLARA Bio Sciences Inc.); GeneChip (Affymetrix, Inc); LabChip (Caliper Technologies Corp); a low-density array with electrochemical sensing (Clinical Micro Sensors); LabCD System (Gamera Bioscience Corp.); Omni Grid (Gene Machines); Q Array (Genetix Ltd.); a high-throughput, automated mass spectrometry systems with liquid-phase expression technology (Gene Trace Systems, Inc.); a thermal jet spotting system (Hewlett Packard Company); Hyseq HyChip (Hyseq, Inc.); BeadArray (Illumina, Inc.); GEM (Incyte Microarray Systems); a high-throughput microarraying system that can dispense from 12 to 64 spots onto multiple glass slides (Intelligent Bio-Instruments); Molecular Biology Workstation and NanoChip (Nanogen, Inc.); a microfluidic glass chip (Orchid biosciences, Inc.); BioChip Arrayer with four PiezoTip piezoelectric drop-on-demand tips (Packard Instruments, Inc.); FlexJet (Rosetta Inpharmatic, Inc.); MALDLTOF mass spectrometer (Sequnom, Inc.); ChipMaker 2 and ChipMaker 3 (TeleChem International, Inc.); and GenoSensor (Vysis, Inc.) as identified and described in Heller (2002) Annu. Rev. Biomed. Eng. 4: 129-153. Examples of “Gene chips” or a “microarray” are also described in U.S. Patent Publ. Nos.: 2007 / 0111322, 2007 / 0099198, 2007 / 0084997, 2007 / 0059769 and 2007 / 0059765 and US Patent 7,138,506, 7,070,740, and 6,989,267.

[0312] In one aspect, “gene chips” or “microarrays” containing probes or primers for the gene of interest are provided alone or in combination with other probes and / or primers. A suitable sample is obtained from the patient extraction of genomic DNA, RNA, or any combination thereof and amplified if necessary. The DNA or RNA sample is contacted to the gene chip or microarray panel under conditions suitable for hybridization of the gene(s) ofinterest to the probe(s) or primer(s) contained on the gene chip or microarray. The probes or primers may be detectably labeled thereby identifying the polymorphism in the gene(s) of interest. Alternatively, a chemical or biological reaction may be used to identify the probes or primers which hybridized with the DNA or RNA of the gene(s) of interest. The genetic profile of the patient is then determined with the aid of the aforementioned apparatus and methods.

[0313] HSD17B13 pharmaconumetic variant interaction with fatty liver polygenic score in association with NASH. Applicant has found that 7AS7J / 77> / 3 loss of function (equivalent to HSD17B13 inhibition) by itself protects against CLD progression to cirrhosis. Applicant has also found that patients with high polygenic risk for fatty liver disease (patients having predictors of liver fat: polygenic risk scores (PRS) for chronically elevated ALT and MRI- quantified liver fat; anthropometric measurements that are correlated with liver fat, such as body mass index (BMI) and BMI-adjusted waist-to-hip ratio (BMIadjWHR); liver fat directly measured by MRI; and more general adipose tissue volume measurements derived from MRI) are expected to have the greatest response to HSD17B13 inhibitor therapies with respect to NASH progression, and that patients with low polygenic risk for fatty liver disease will have more limited response to these agents.Selecting patients for clinical trials or treatment

[0314] Applicant also noted that a trial of asx HSD17B13 inhibitor could meaningfully improve its expected treatment effect by selecting patients with two functional HSD17B13 alleles and at least one PNPLA3 148M allele, compared to selecting on only PNPLA3 or taking all-comers.

[0315] An aspect of the disclosure is directed to a method for selecting a patient for an HSD17B13 inhibitor trial comprising, or consisting essentially of, or consisting of determining the genotype for HSD17B 13 and PNPLA3 genes in the patient; and selecting the patient having two functional HSD17B13 alleles and at least one PNPLA3 148M allele for the trial. In some embodiments, the patient suffers from CLD or NASH. In a further aspect, the HSD17B13 inhibitor is administered to the patient.

[0316] Another aspect of the disclosure is directed to a method for selecting a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH) patient for an HSD17B13 inhibitor therapy comprising , or consisting essentially of, or consisting of determining the genotype for HSD17B 13 and PNPLA3 genes in the patient; and selecting the patient having twofunctional HSD17B13 alleles and at least one PNPLA3 148M allele . In a further aspect, the HSD17B13 inhibitor is administered to the patient.

[0317] Another aspect of the disclosure is directed to a method for treating a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH) patient, the method comprising, or consisting essentially of, or yet further consisting of determining the genotype for HSD17B 13 and PNPLA3 genes in the patient; and treating the patient having two functional HSD17B13 alleles and at least one PNPLA3 148M allele with an HSD17B13 inhibitor.

[0318] An aspect of the disclosure is directed to a method for determining whether a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH) patient is at risk of progressing to cirrhosis or hepatocellular carcinoma comprising, or consisting essentially of, or consisting of: (i) sequencing the PNPLA3 gene in the patient; and (ii) determining that the patient is at risk of progressing to cirrhosis or hepatocellular carcinoma when a PNPLA3 I148M mutation is detected in the patient.

[0319] In some embodiments, the patient is heterozygous for the PNPLA3 I148M mutation. In some embodiments, the patient is homozygous for the PNPLA3 I148M mutation.

[0320] In some embodiments, the method further comprises, or consists essentially of, or consists of determining a higher risk of progressing to cirrhosis or hepatocellular carcinoma if the patient is diabetic. In some embodiments, the method further comprises, or consists essentially of, or consists of determining a higher risk of progressing to cirrhosis or hepatocellular carcinoma when the patient has a body mass index (BMI) over 30.

[0321] In some embodiments, the method further comprises, or consists essentially of, or consists of administering to the patient an HSD17B13 inhibitor.

[0322] An aspect of the disclosure is directed to a method for determining whether a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH) patient will benefit from a PNPLA3 degrader comprising, or consisting essentially of, or consisting of: (i) sequencing the PNPLA3 gene in the patient; and (ii) determining that the patient will benefit from a PNPLA3 degrader when a PNPLA3 I148M mutation is detected in the patient.

[0323] In some embodiments, the patient is heterozygous for the PNPLA3 I148M mutation. In some embodiments, the patient is homozygous for the PNPLA3 I148M mutation.

[0324] In some embodiments, the method further comprises, or consists essentially of, or consists of determining a higher benefit from a PNPLA3 degrader if the patient is diabetic. Insome embodiments, the method further comprises, or consists essentially of, or consists of determining a higher benefit from a PNPLA3 degrader if the patient has a body mass index (BMI) over 30.

[0325] In some embodiments, the method further comprises, or consists essentially of, or consists of administering to the patient a PNPLA3 degrader.

[0326] An aspect of the disclosure is directed to a method for determining whether a chronic liver disease (CLD) or a non-alcoholic steatohepatitis (NASH) patient will benefit from an HSD17B13 inhibitor comprising, or consisting essentially of, or consisting of combining the body mass index (BMI) and BMI-adjusted waist-to-hip ratio of the patient, thereby generating a combined score; and determining that the patient will benefit from the HSD17B13 inhibitor when the combined score is above a predetermined threshold.

[0327] In some embodiments, the method further comprises, or consists essentially of, or consists of administering to the patient an HSD17B13 inhibitor.Systems of the Disclosure for Risk Score-Informed Drug Development for chronic liver disease (CLD) or non-alcoholic steatohepatitis (NASH)

[0328] There are many applications of the systems of the present disclosure, ranging from drug discovery using agnostic analysis of new chemical entities in combination with genetic variant distribution in patient populations, to new and efficient designs of clinical trials, to incorporation of new end points in clinical trials, to how the data from former clinical studies are analyzed to provide evidence of effectiveness. Applications of particular value include combining genetic information on patient populations with a better understanding of pharmacodynamic end points and biomarkers and how they relate to clinical outcomes.

[0329] For example, the systems of the disclosure can be used for in silico clinical trial recapitulation and regulatory evaluation. Historically, safety and efficacy data provided to regulatory agencies in support of marketing has required preclinical and clinical efficacy data using empirical methods.Computer Program Products

[0330] The present approach may be configured as a system or method, but also may be provided as a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable programinstructions thereon for causing a processor to carry out aspects of the present disclosure. The systems and methods disclosed herein have the superior benefit of reducing memory processing resources.

[0331] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD- ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire. Rather, the computer readable storage medium is a non-transient (i.e., not- volatile) medium.

[0332] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0333] Computer readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions,machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’s computer, partly on the user’ s computer, as a stand-alone software package, partly on the user’ s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0334] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions and / or steps specified in the disclosure. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises, or alternatively consists essentially of, or yet further consists of an article of manufacture including instructions which implement aspects of the functions and / or steps specified in the disclosure.

[0335] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions and / or steps specified in the disclosure.Methods for Determining Drug Activity for Treating Chronic Liver Disease

[0336] An aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets for treating chronic liver disease, the method comprising, or alternatively consisting essentially of, or yet further consisting of: obtaining molecular biomarker stratifier data comprising, or alternatively consisting essentially of, or yet further consisting of a plurality of biomarker stratifiers and at least one disease phenotype from each of a plurality of subjects; determining a plurality of values representing biomarker stratifier effects, wherein each value separately represents how each of the plurality of biomarker stratifiers affects each of the at least one disease phenotype based on external data; calculating a biomarker stratifier score for a chosen disease phenotype; calculating a pharmacomimetic genetic score for each drug target; and identifying the predicted drug activity of each drug target the at least one disease phenotype in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the disease phenotype.

[0337] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof. In some embodiments, the biomarker stratifier comprises polygenic risk scores for liver fat or chronically elevated ALT.

[0338] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0339] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarkerstratifi er score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0340] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity;(v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0341] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype;(vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0342] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

[0343] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0344] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and the plurality of values.

[0345] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0346] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0347] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of the weighted principal component analysis identifies predicted drug activity of the plurality of drug targets for the at least one disease phenotype.

[0348] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0349] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0350] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of oneor more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0351] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity;(v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0352] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype;(vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0353] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

[0354] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or theweighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0355] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and the plurality of values.

[0356] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0357] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0358] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of the weighted principal component analysis identifies predicted drug activity of the plurality of drug targets for the at least one disease phenotype.

[0359] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0360] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0361] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) atranscriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0362] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity;(v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0363] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype;(vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0364] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

[0365] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or theweighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0366] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and the plurality of values.

[0367] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0368] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0369] In some embodiments, the plurality of drug targets is associated with obesity, and the plurality of drug targets comprises, or alternatively consists essentially of, or yet further consists of GLP1R and / or GIPR.

[0370] In some embodiments, the pharmacomimetic genetic score is associated with a response to a drug that targets GLP1R and / or GIPR. In some embodiments, the drug is a GLP1R and / or GIPR agonist.

[0371] In some embodiments, the polygenic score is a body mass index polygenic score.

[0372] Another aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets, the method comprising: obtaining molecular biomarker stratifier data and at least one disease phenotype from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each disease phenotype based on external data; constructing a matrix based on the biomarker stratifiers, disease phenotypes and values; running a principal components analysis biomarker stratifier x phenotype-related phenotype biomarker stratifier effects matrix; calculating a biomarker stratifier score for each principal component; calculating a pharmacomimetic genetic score for each drug target; and identifying predicted drug activity of each drug target on one or more phenotypes based onthe statistical interaction of the polygenic score calculated for each principal component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0373] Another aspect of the disclosure is directed to an in silico method for determining the drug activity of drug targets on a plurality of phenotype intermediates, comprising: obtaining molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each phenotype intermediate based on external data; calculating a biomarker stratifier score for each phenotype intermediate in the disease process; calculating a pharmacomimetic genetic score for each drug target on each phenotype intermediate; and identifying predicted drug activity of the drug targets on one or more phenotypes based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the phenotype intermediate and the outcome phenotype.

[0374] Another aspect of the disclosure is directed to method of determining the drug activity of a drug target, comprising: obtaining molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each phenotype based on external data; calculating a biomarker stratifier score for a chosen outcome phenotype related to the disease outcome; calculating a pharmacomimetic genetic score for the drug target; and identifying predicted drug activity of the drug target on one or more phenotypes based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for the drug target in association analysis with the outcome phenotype.

[0375] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0376] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0377] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0378] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity; (v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0379] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenicscore for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0380] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

[0381] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0382] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and the plurality of values.

[0383] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0384] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0385] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of the weighted principal component analysis identifies predicted drug activity of the plurality of drug targets for the at least one disease phenotype.

[0386] Another aspect of the disclosure is directed to an in silico method comprising: generating a plurality of principal components (PCs) corresponding to geneticancestry data for subjects in a study cohort; generating a biomarker stratifier score for each subject in the study cohort based at least on (i) the PCs and on (ii) biomarker stratifier weights for a disease of interest; determining which of a plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest, where a disease-associated variant is a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype; determining statistical interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores, wherein the statistical interactions are predictive of drug target-specific differential treatment response for the disease of interest; and performing one or more prediction-based actions based on the determined interactions.

[0387] In some embodiments, the one or more prediction-based actions comprises at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics.

[0388] In some embodiments, the genetic ancestry data indicates whether subjects in the study cohort are a case or a control for the disease of interest.

[0389] In some embodiments, the genetic ancestry data is based on genotyping arrays or whole-genome sequencing.

[0390] In some embodiments, the genetic ancestry data is obtained from a publicly available database.

[0391] In some embodiments, the plurality of PCs comprises at least 5 PCs.

[0392] In some embodiments, the study cohort comprises at least 200 control subjects.

[0393] In some embodiments, each disease-associated variant in the plurality of disease- associated variants satisfies a genome-wide significance threshold.

[0394] In some embodiments, the plurality of disease-associated variants is determined based on a first disease genome-wide association study (GWAS).

[0395] In some embodiments, the biomarker stratifier score for each subject is a scaled biomarker stratifier score.

[0396] In some embodiments, the scaled biomarker stratifier score is based at least on a raw PGS and an ancestry-normalized PGS.

[0397] In some embodiments, the PGS variant weights are computed independent of the genetic ancestry data corresponding to the study cohort.

[0398] In some embodiments, determining which of the plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest comprises generating an allelic score.

[0399] In some embodiments, variants in the allelic score are weighted based at least on a second GWAS that did not include any of the subjects in the study cohort.

[0400] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0401] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0402] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0403] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet furtherconsists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity;(v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0404] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype;(vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.Systems

[0405] Another aspect of the disclosure is directed to an in silico system for drug development, such system comprising: at least one hardware processor; and a non-transitory computer-readable storage medium having stored thereon program code, the program code executable by the at least one hardware processor to: obtain molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determine the value of the biomarker stratifier effects for each phenotype based on external data; construct a matrix based on the biomarker stratifiers, phenotypes and values; run a principal components analysis biomarker stratifier x phenotype-related phenotypebiomarker stratifier effects matrix; calculate a biomarker stratifier score for each principal component; calculate a pharmacomimetic genetic score for each drug target; determine predicted drug effects on one or more phenotypes based on the statistical interaction of the polygenic score calculated for each principal component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0406] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0407] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0408] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0409] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutationscore for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity;(v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0410] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.EXAMPLES

[0411] The following examples are included for illustrative purposes only and are not intended to limit the scope of the disclosure.Example 1: HSD17B13 is a drug target for chronic liver disease (CLD).In this disclosure, Applicant:• consolidated and replicated analyses showing a genetic interaction between HSD17B13 and PNPLA3 and an effect of PNPLA3 on CLD progression to cirrhosis.• investigated interactions between PISD17B13, PNPLA3, and several predictors of liver fat: polygenic risk scores (PRS) for chronically elevated ALT and MRI-quantified liver fat; anthropometric measurements that are correlated with liver fat, such as body mass index (BMI) and BMI-adjusted waist-to-hip ratio (BMIadjWHR); liver fat directly measured by MRI; and more general adipose tissue volume measurements derived from MRI.• investigated interactions between HSD17B13, PNPLA3, and diabetes.

[0412] Applicant investigated genetic and environmental predictors of serum alanine aminotransferase (ALT) levels, a biomarker of liver function. ALT is a quantitative trait, so Applicant used simple linear regression models fit to data from unrelated, European-ancestry subjects from the UK Biobank (UK Biobank).

[0413] The distribution of ALT was roughly log-normal, so GWAS of ALT have typically log-transformed it (e.g., Ward, Lucas D., et al. Nature Communications 12.1 (2021): 4571). This assumed that genetic variants had multiplicative effects. In contrast, GWAS of estimated glomerular filtration rate (eGFR), a biomarker of kidney function, did not logtransform the phenotype, instead handling outliers though “winsorization”, a procedure in which measurements greater than a threshold are replaced with the threshold value (Wuttke, Matthias, et al. Nature Genetics 51.6 (2019): 957-972). This assumed that genetic variants have additive effects.

[0414] Applicant compared loglO-transformed ALT to ALT winsorized with values greater than 84 U / L (twice the upper reference limit recommended for males in Valenti, Luca, et al. Hepatology Communications 5.11 (2021): 1824-1832; the upper reference limit recommended for females is lower, so Applicant used the higher value). The values performed similarly in predicting prevalent CLD and cirrhosis, but winsorized ALT gave more significant associations for positive control variants in PNPLA3 and HSD17B13. So, winsorized ALT was selected as the outcome for the present linear regression models.

[0415] All of the present linear regression models included age, age squared, sex, genotyping chip, and the first through sixth principal components (PCs) of genetic ancestry as covariates. Other predictors, e.g., PNPLA3 and HSD17B13 genotype, are described in line in the results. In some models, Applicant used BMLadjusted waist-to-hip ratio as a measure of adiposity. This statistic was actually adjusted for sex in addition to BMI.

[0416] The following phenotype definitions were used to consider CLD, cirrhosis, or diabetes as binary phenotypes:CLD• Cases must have at least one of: ICD-10 B18, C22.0, K70, K72.1, K73, K74.0-K74.2, K74.6, K75.8, K76.0, K76.6-K76.7; ICD-9 155.0, 456.2, 571.0, 571.2, 571.4, 571.50, 571.51, 571.58, 571.59, 571.8, 572.3, 572.4; UK Biobank self-report (non-cancer) 1158, 1579-1580, 1604.• Controls must not be cases and also must not have any of: ICD-10 185, K76.9; ICD-9 456.0, 456.1; UK Biobank self-report (non-cancer) 1141, 1155-1157.• In addition to the requirements above, cases and controls must not have any of: ICD- 10 K74.3, K75.4; ICD-9 571.6; UK Biobank self-report (non-cancer) 1475, 1506.Non-viral hepatitis (NVH)• Same as CLD except does not use case codes: ICD-10 B18, C22.0, K76.0; ICD-9 155.0, 571.0; UK Biobank self-report (non-cancer) 1579-1580.Cirrhosis• Cases must have at least one of: ICD-10 185, K70.2-K70.4, K74.1-K74.2, K74.6, K76.6-K76.7; ICD-9 456.2, 571.2, 571.50-571.51, 571.59; UK Biobank self-report (noncancer) 1158.• Control must not be CLD cases.• In addition to the requirements above, cases and controls must not have any of: ICD- 10 K74.3, K75.4; ICD-9 571.6; UK Biobank self-report (non-cancer) 1475, 1506.Diabetes: cases must have at least one of the codes below; controls must have none.ICD-10: E10-E14• ICD-9: 250• UK Biobank self-report (non-cancer): 1220-1223• Doctor-diagnosed diabetes (UK Biobank data field 2443, code = 1)• UK Biobank medication: 1140883066, 1140884600, 1140874686, 1141171646,1141171652, 1141177600, 1141189090, 1141189094, 1141177606, 1141153254, 1141153262, 1140874744, 1141152590, 1141156984, 1140874718, 1140874736,1140874724, 1140874726, 1140874728, 1140874740, 1140910566, 1140874746,1140874646, 1141157284, 1140874652, 1140874674, 1140874690, 1140874706,1140874712, 1140874716, 1140874664, 1140874666, 1141168660, 1141168668,1141173882, 1141173786, 1140868902, 1140868908

[0417] Applicant investigated polygenic risk scores (PRS) for liver fat, BMI, and type 2 diabetes (T2D). The PRS were trained using PRS-CS (Ge, Tian, et al. Naturecommunications 10.1 (2019): 1776) applied to summary statistics from the GWAS listed below.• Liver fat: Haas, Mary E., et al. Cell genomics 1.3 (2021). This GWAS was performed on MRI data from UK Biobank subjects, so for all downstream analyses using the liver fat PRS, Applicant excluded all UK Biobank subjects with MRI data.• BMI: Locke, Adam E., et al. Nature 518.7538 (2015): 197-206. This GWAS did not include UK Biobank subjects.• T2D: an in-house meta-analysis of an in-house meta-analysis of Scott, Robert A., et al. Diabetes 66.11 (2017): 2888-2902, FinnGen r6 phenocode E4 DM2 (Kurki, Mitja I., et al. MedRxiv (2022): 2022-03.), and Spracklen, Cassandra N., et al. Nature 582.7811 (2020): 240-245. None of these GWAS include UK Biobank subjects.)

[0418] PRS-CS had a hyperparameter, phi, which related to the polygenicity of the phenotype. Applicant selected the value for phi for each PRS that gave the best power to predict the corresponding disease in the UK Biobank, or in the case of the liver fat PRS, Applicant selected the value for phi that gave the highest correlation between the PRS and ALT levels. As downstream analysis used the same UK Biobank subjects Applicant used to select phi, Applicant is potentially overfitting. However, any confounding effects are expected to be minimal, as phi should not be dependent on LD structure, which is the main driver of differential performance of PRS across cohorts.

[0419] Applicant also investigated a PRS for chronically elevated ALT (cALT) suggestive of non-alcoholic fatty liver disease (NAFLD), trained using data from Vujkovic, Marijana, et al. Nature Genetics 54.6 (2022): 761-771. This PRS was constructed with a manually curated set of variants and weights instead of using PRS-CS, as Applicant did not have access to the genome-wide summary statistics from this study. Specifically, Applicant used a set of 41 variants reported in the supplementary tables of the Vujkovic study that 1) had a genomewide significant association with cALT, and 2) were replicated at p< 0.05 in either a casecontrol study of biopsy-confirmed NAFLD or a study of hepatic fat as a quantitative trait. One SNP, rsl 1601507, near TRIM22, met these criteria but was excluded as it was not present in the UK Biobank imputed genotypes.

[0420] Finally, Applicant examined a 2-SNP version of the above cALT PRS, including just PNPLA3 I148M and TM6SF2 E167K (excluding HSD17B13 as Applicant are using this score to test for an interaction with HSD17B13 ).

[0421] Before use, all PRS were adjusted via linear regression to remove correlation with ancestry PCs 1-6.

[0422] Applicant performed an analysis of progression from a subject’s first inpatient CLD diagnosis to a composite endpoint comprised of inpatient cirrhosis (using the codes listed above) or hepatocellular carcinoma (ICD-10 C22, ICD-9 155.0), procedures for drainage of ascites (OPCS-4 T46.1, T46.2), liver transplant (OPCS-4 JOI), or liver-related mortality (ICD-10 K70-K77). This progression analysis used a standard Cox proportional hazards model using age at CLD diagnosis, sex, genotyping chip, and ancestry PCs 1-6 as covariates.ResultsCohort statistics

[0423] Table 1: Cohort statistics, stratified by chronic liver disease (CLD) case-control status.Statistic CLD cases CLD controls# subjects 8,006 347,527% of cohort 2.3% 97.7%Median age at end of data (or death, 72 (55, 83) 72 (56, 83) whichever comes first) (95% interval)% male 56.3% 46.1%% with diabetes 29.3% 7.1%Median ALT (U / L) (95% interval) 28 (11, 108) 20 (10, 56)Median BMI (95% interval) 30 (21, 43) 27 (20, 39)Median waist-to-hip ratio (95% interval) 0.93 (0.75, 1.09) 0.87 (0.71, 1.04)

[0424] Table 2: Cohort statistics, stratified by non-viral hepatitis (NVH) case-control status.Statistic NVH cases NVH controls# subjects 3,119 352,141% of cohort 0.9% 99.1%Median age at end of data (or death, 71 (54, 83) 72 (56, 83) whichever comes first) (95% interval)% male 68.0% 46.1%% with di ab ete s 31.5 % 7.4%Median ALT (U / L) (95% interval) 30 (11, 131) 20 (10, 56)Median BMI (95% interval) 29 (20, 43) 27 (20, 39)Median waist-to-hip ratio (95% interval) 0.94 (0.76, 1.10) 0.87 (0.71, 1.04)

[0425] Table 3: Cohort statistics, stratified by cirrhosis case-control status.Statistic Cirrhosis cases Cirrhosis controls# subjects 2,412 347,122% of cohort 0.7% 99.3%Median age at end of data (or death, 71 (54, 83) 72 (56, 83) whichever comes first) (95% interval)% male 69.1% 46.1%% with di ab ete s 32.9% 7.1 %Median ALT (U / L) (95% interval) 30 (11, 131) 20 (10, 56)Median BMI (95% interval) 29 (20, 43) 27 (20, 39)Median waist-to-hip ratio (95% interval) 0.95 (0.76, 1.10) 0.87 (0.71, 1.04)

[0426] Table 4: Statistics for CLD cases with an inpatient diagnosis.Statistic Value# subjects 7,830# subjects, excluding those in the liver fat PRS 7,505 training datasetMedian age at diagnosis 66.1Median years of follow up 3.8Reached endpoint (cirrhosis, HCC, ascites 30.4% operation, liver transplant, or liver-related death)PNPLA3 148M carrier 46.1%Prior diabetes 18.2%Prior diabetes and PNPLA3 148M carrier 8.2%

[0427] Table 5: Statistics for NVH cases with an inpatient diagnosis.Statistic Value# subjects 3,024# subjects, excluding those in the liver fat PRS 2,950 training datasetMedian age at diagnosis 64.9Median years of follow up 4.3Reached endpoint (cirrhosis, HCC, ascites 70.9% operation, liver transplant, or liver-related death)PNPLA3 148M carrier 48.8%Prior diabetes 19.0%Prior diabetes and PNPLA3 148M carrier 9.9%

[0428] Table 6: Genotypes of the HSD17B13 splice variant (rs72613567) and PNPLA3 I148M (rs738409) within UK Biobank NVH cases. Allele frequencies in the general UK Biobank European population are 0.28 for HSD17B13 rs72613567 and 0.22 for PNPLA3 rs738409.% NVH cases inHSD17B13 PNPLA3 1148M # NVH % NVH HSD17B 13 genotype rs72613567 genotype genotype cases cases groupT / T FI 868 27.8 50.5T / T I / M 717 23.0 41.7T / T M / M 135 4.3 7.8T / TA FI 620 19.9 52.1T / TA I / M 458 14.7 38.5T / TA M / M 113 3.6 9.5TA / TA FI 117 3.8 56.2TA / TA I / M 72 2.3 34.6TA / TA M / M 19 0.6 9.1

[0429] Marginal effects of HSD17B13 and PNPLA3 on ALT + justification for ALT winsorization

[0430] Winsorized ALT explained slightly more variance in NVH and cirrhosis liability than log-transformed ALT, though they were basically equivalent. The reason that the odds ratio of NVH per SD of log-transformed ALT was greater than the OR / SD of winsorized ALT was because log-transformed was more normal / less right-skewed, so the standard deviation was smaller. See, e.g., Figures 13-15.

[0431] The ALT measurements were taken at each subject’s initial UK Biobank assessment visit, while most NVH and cirrhosis diagnoses were based on hospital inpatient data accrued subsequent to the initial assessment visits.

[0432] Table 7: Comparison of untransformed ALT, log transformed ALT, and winsorized ALT in their ability to predict CLD, NVH, and cirrhosis case-control status.Outcome ALT OR per SD of ALT Test statistic McFadden’s R 2 transformationCirrhosis None 1.39 (1.36, 1.41) 34.5 0.089Cirrhosis LoglO 2.00 (1.93, 2.07) 39.7 0.102Cirrhosis Winsorized 1.74 (1.70, 1.78) 43.6 0.105CLD None 1.45 (1.43, 1.47) 53.6 0.091CLD LoglO 1.90 (1.86, 1.94) 63.1 0.103CLD Winsorized 1.67 (1.65, 1.70) 67.9 0.103NVH None 1.39 (1.37, 1.42) 38.3 0.088NVH LoglO 1.97 (1.91, 2.03) 44.1 0.100NVH Winsorized 1.71 (1.67, 1.75) 47.9 0.102

[0433] Winsorized ALT gave a stronger association for PNPLA3 I148M and the HSD17B13 splice variant than either log-transformed ALT or untransformed ALT. Applicant concludedthat winsorized ALT was an equivalent biomarker of liver function compared to ALT with a log or no transformation, and may in fact be a modest improvement.

[0434] Table 8: Comparison of the significance of associations between positive control variants in PNPLA3 and HSD17B13 with untransformed ALT, log transformed ALT, and winsorized ALT. The analysis included BMI as a covariate.Variant ALT Beta Test statistic P-value transformationHSD17B 13 splice variant None -0.65 (-0.72, -0.58) -18.2 4.0e-74(rs72613567)HSD17B 13 splice variant LoglO -0.01 (-0.01, -0.01) -16.8 1.6e-63(rs72613567)HSD17B 13 splice variant Winsorized -0.58 (-0.64, -0.52) -19.0 9.0e-81(rs72613567)PNPLA3 1148M None 1.69 (1.62, 1.77) 43.6 0.0e+00(rs738409)PNPLA3 1148M LoglO 0.02 (0.02, 0.02) 42.4 0.0e+00(rs738409)PNPLA3 1148M Winsorized 1.57 (1.51, 1.63) 47.7 0.0e+00(rs738409)

[0435] Marginal effects of HSD17B13 and PNPLA3 on CLD, NVH, and cirrhosis

[0436] Table 9: Associations of HSD17B13 and PNPLA3 variants with liver disease. The analysis included BMI as a covariate.Variant Phenotype OR P-value # cases # controlsHSD17B 13 splice Cirrhosis 0.90 (0.84, 0.96) 1.6e-03 2,412 349,534 variant (rs72613567)HSD17B13 splice Cirrhosis 0.91 (0.86, 0.98) 7.3e-03 2,382 349,909 variant (rs72613567) (Emdin 2021)HSD17B 13 splice CLD 0.95 (0.92, 0.98) 4.1e-03 8,006 355,533 variant (rs72613567)HSD17B 13 splice NVH 0.90 (0.85, 0.95) 3.3e-04 3,119 355,260 variant (rs72613567)PNPLA3 I148M Cirrhosis 1.56 (1.47, 1.67) 1.3e-44 2,412 349,534(rs738409)PNPLA3 I148M Cirrhosis 1.58 (1.48, 1.68) 3.0e-46 2,382 349,909(rs738409) (Emdin 2021)PNPLA3 I148M CLD 1.37 (1.32, 1.42) 6.5e-67 8,006 355,533(rs738409)PNPLA3 I148M NVH 1.49 (1.41, 1.57) 8.6e-45 3,119 355,260

[0437] Comparison of three liver fat phenotypes available in UK Biobank

[0438] Three companies have derived liver fat fraction measurements from UK Biobank subjects’ abdominal MRI data: AMRA, Calico, and Perspectum. Applicant compared these phenotypes using positive control variants that were known to be associated with liver fat, PNPLA3 alleles I148M and TM6SF2 E167K. The variants’ associations with the AMRA and Perspectum phenotypes were about equivalent, while the associations with the Calico phenotype were attenuated. Perspectum had data for more subjects than AMRA, so Applicant used the Perspectum liver fat phenotype in its further analysis.

[0439] Consistent with Abul-Husn, Noura S., et al. (New England Journal of Medicine 378.12 (2018): 1096-1106), HSD17B13 was not found to associate with liver fat.

[0440] Table 10: Associations of HSD17B13 and PNPLA3 variants with three liver fat phenotypes available in the UK Biobank data. The units of the Beta column are percentage points of liver fat. This analysis included BMI as a covariate.# subjectsVariant Phenotype Beta P-value with dataPNPLA3 I148M Liver fat fraction 1.16 (1.08, 1.25) LOe-147 29,989(rs738409) (Perspectum)PNPLA3 I148M Liver fat fraction 0.95 (0.86, 1.03) 5.5e-114 23,098(rs738409) (AMRA)PNPLA3 I148M Liver fat fraction 1.15 (1.05, 1.24) l.le-110 22,441(rs738409) (Calico)TM6SF2 E167K Liver fat fraction 1.63 (1.50, 1.77) 2.9e-118 29,989(rs58542926) (Perspectum)TM6SF2 E167K Liver fat fraction 1.41 (1.28, 1.53) 2.8e-102 23,098(rs58542926) (AMRA)TM6SF2 E167K Liver fat fraction 1.67 (1.51, 1.83) 1.4e-95 22,441(rs58542926) (Calico)HSD17B 13 splice variant Liver fat fraction 0.06 (-0.02, 0.15) 0.124 29,989(rs72613567) (Perspectum)HSD17B 13 splice variant Liver fat fraction 0.05 (-0.02, 0.13) 0.185 23,098(rs72613567) (AMRA)HSD17B 13 splice variant Liver fat fraction 0.04 (-0.05, 0.14) 0.361 22,441(rs72613567) (Calico)

[0441] Replication: HSD17B13 x PNPLA3 interaction

[0442] Applicant replicated the HSD17B13-PNPLA3 interaction in determining serum levels of alanine aminotransferase (ALT, a biomarker of liver injury), that was reported in Abul-Husn, Noura S., et al. New England Journal of Medicine 378.12 (2018): 1096-1106.

[0443] Table 11: Interaction between the HSD17B13 splice variant (rs72613567) and PNPLA3 I148M (rs738409) in determining winsorized ALT. The analysis includes BMI as a covariate.Term Beta P-valueHSD17B13 splice variant -0.36 (-0.44, -0.29) 4.9e-22PNPLA3 1148M 1.85 (1.76, 1.93) 0.0e+00HSD17B13 splice variant : -0.50 (-0.60, -0.39) 1.5e-21PNPLA3I148M

[0444] Table 12: Effect of the HSD17B13 splice variant (rs72613567) on winsorizedALT, stratified by PNPLA3 I148M genotype. The analysis included BMI as a covariate.PNPLA3 genotype Beta P-valueI / I -0.38 (-0.46, -0.31) 5.5e-26I / M -0.78 (-0.89, -0.67) 1.6e-45M / M -1.64 (-1.98, -1.30) 4.0e-21PNPLA3 genotype Beta P-value

[0445] Table 13: Effect of the HSD17B13 splice variant (rs72613567) on winsorized ALT, stratified by PNPLA3 I148M carrier status. The analysis included BMI as a covariate.PNPLA3 genotype Beta P-valueI / I -0.38 (-0.46, -0.31) 5.5e-26I / M or M / M -0.89 (-0.99, -0.78) l. le-62

[0446] There is an apparent non-linearity in the effect of HSD17B13 loss of function on ALT that was consistent across all PNPLA3 genotypes. Heterozygous HSD17B13 loss of function had substantially >50% of the effect of homozygous HSD17B13 loss of function.

[0447] In Figure 7A, Applicant used log(ALT) as the present outcome instead of winsorized ALT. The choice in outcome made only a small difference to the results (see Figure 7B).

[0448] Table 14: Likelihood ratio tests comparing genotypic to allelic models of HSD17B13 loss of function on winsorized ALT, stratified by PNPLA3 genotype. The genotypic model included separate terms for HSD17B13 rs72613567-TA heterozygous and homozygous status. The allelic model had a single term for the dosage of rs72613567-TA (0, 1, or 2 copies). This analysis included BMI as a covariate.PNPL A3 genotype HSD17B 13 heterozygote LR test p-value effect, as % of homozygote effectVI 75% 6.3e-03I / M 79% 3.0e-05M / M 77% 0.010

[0449] Looking specifically within subjects who had CLD or NVH at the time of their UK Biobank initial assessment visit, when the ALT measurements were taken, the HSD17B13 interaction term with PNPLA3 trended accordingly, though nominally significant only for CLD and not NVH.

[0450] Table 15: Interaction between the HSD17B13 splice variant (rs72613567) and PNPLA3 I148M (rs738409) in determining winsorized ALT in subjects who had CLD at the time the ALT measurements were taken. The analysis includes BMI as a covariate.Term Beta P-valueHSD17B 13 splice variant 0.61 (-1.77, 2.98) 0.616PNPLA3 H48M 5.45 (3.09, 7.80) 6.2e-06HSD17B13 splice variant : -3.71 (-6.66, -0.75) 0.014PNPLA3 I148M

[0451] Table 16: Interaction between the HSD17B13 splice variant (rs72613567) and PNPLA3 I148M (rs738409) in determining winsorized ALT in subjects who had NVH at the time the ALT measurements were taken. The analysis included BMI as a covariate.Term Beta P-valueHSD17B13 splice variant 0.43 (-2.67, 3.53) 0.784PNPLA3 1148M 5.33 (2.31, 8.35) 5.8e-04HSD 17B 13 splice variant : -3.12 (-7.05, 0.80) 0.119PNPLA3 I148M

[0452] The HSD17B13: PNPLA3 interaction was nominally significant with binary CLD as the endpoint, but has no signal with NVH or cirrhosis as the endpoint.

[0453] Table 17: Interaction between the HSD17B13 splice variant (rs72613567) and PNPLA3 I148M (rs738409) in determining risk of CLD. The analysis included BMI as a covariate.Term OR P-valueHSD 17B 13 splice variant 0.98 (0.94, 1.03) 0.444PNPLA3 1148M 1.42 (1.35, 1.49) 3.6e-48HSD17B13 splice variant : 0.94 (0.89, 0.99) 0.030PNPLA3 I148M

[0454] Table 18: Effect of the HSD17B13 splice variant (rs72613567) on risk of CLD, stratified by PNPLA3 I148M carrier status. The analysis included BMI as a covariate.PNPLA3 genotype OR P-valueI / I 0.99 (0.95, 1.04) 0.729I / M or M / M 0.90 (0.85, 0.95) 1.5e-04

[0455] Table 19: Interaction between the HSD17B13 splice variant (rs72613567) and PNPLA3 I148M (rs738409) in determining risk of NVH. The analysis included BMI as a covariate.Term OR P-valueHSD17B 13 splice variant 0.92 (0.85, 0.99) 0.034PNPLA3 1148M 1.52 (1.41, 1.63) 8.2e-30HSD 17B 13 splice variant : 0.96 (0.88, 1.05) 0.405PNPLA3 I148M

[0456] Table 20: Effect of the HSD17B13 splice variant (rs72613567) on risk of NVH, stratified by PNPLA3 I148M carrier status. The analysis includes BMI as a covariate.PNPLA3 genotype OR P-valueI / I 0.95 (0.87, 1.02) 0.177I / M or M / M 0.85 (0.79, 0.93) 2.2e-04

[0457] Table 21: Interaction between the HSD17B13 splice variant (rs72613567) and PNPLA3 I148M (rs738409) in determining risk of cirrhosis. The analysis included BMI as a covariate.Term OR P-valueHSD 17B 13 splice variant 0.89 (0.81, 0.97) 7.7e-03PNPLA3 1148M 1.54 (1.42, 1.67) 2.9e-25HSD 17B 13 splice variant : 1.03 (0.93, 1.14) 0.597PNPLA3 I148M

[0458] Table 22: Effect of the HSD17B13 splice variant (rs72613567) on risk of cirrhosis, stratified by PNPLA3 I148M carrier status. The analysis includes BMI as a covariate.PNPLA3 genotype OR P-valueI / I 0.92 (0.84, 1.00) 0.058I / M or M / M 0.89 (0.81, 0.97) 0.012

[0459] HSD17B13 effect stratified by both PNPLA3 and TM6SF2

[0460] Table 23: Effect of the HSD17B13 splice variant (rs72613567) on CLD-related phenotypes, in subgroups defined by the number of PNPLA3 / TM6SF2 risk alleles that a subject carries.# ofPNPLA3 / TM6SF2 #Endpoint risk alleles Beta / OR P-value # cases # controls #subjectsALT 0 -0.37 (-0.44, - 3.5e-21 NA NA 1795640.29)ALT 1 -0.70 (-0.80, - 2.4e-42 NA NA 1280210.60)ALT 2 -1.16 (-1.39, - 1.7e-22 NA NA 311890.92)ALT 3 or 4 -1.90 (-2.74, - l.le-05 NA NA 29181.05)Cirrhosis 0 0.90 (0.82, 1.00) 0.049 1015 183195 184210Cirrhosis 1 0.85 (0.77, 0.95) 2.6e-03 956 129785 130741Cirrhosis 2 1.02 (0.87, 1.21) 0.782 370 31255 31625Cirrhosis 3 or 4 1.00 (0.68, 1.48) 0.995 71 2887 2958CLD 0 1.00 (0.95, 1.05) 0.899 3532 183401 186933CLD 1 0.92 (0.87, 0.97) 2.2e-03 3266 129921 133187CLD 2 0.92 (0.83, 1.01) 0.084 1060 31314 32374CLD 3 or 4 0.90 (0.68, 1.19) 0.455 148 2891 3039NVH 1 0.85 (0.77, 0.93) 4.2e-04 1278 131806 133084NVH 2 0.93 (0.80, 1.08) 0.333 450 31893 32343NVH 3 or 4 0.99 (0.67, 1.47) 0.976 70 2963 3033

[0461] HSD17B13, PNPLA3, and liver cTl

[0462] Liver iron-corrected cTl is a biomarker derived from MRI that has been used as a proxy for liver inflammation and fibrosis. It is also a proposed endpoint for clinical trials of NASH. The distribution of cTl values in the UK Biobank was roughly normal, with a slightly heavy right tail (Figure 17).

[0463] PNPLA3 I148M and TM6SF2 E167K increase liver cTl, liver fat, ALT, and cirrhosis risk. The HSD17B13 splice variant did not have any effect on cTl or liver fat, despite associating with reduced ALT and cirrhosis risk.

[0464] Table 24: Effects of CLD / cirrhosis-associated variants on liver iron- corrected cTl (units = s.d.). This analysis included BMI as a covariate.Variant Beta P-valuePNPLA3 1148M 0.12 (0.10, 0.14) 2.7e-35(rs738409)TM6SF2 E167K 0.16 (0.13, 0.19) 2.5e-23HSD17B 13 splice variant -0.00 (-0.02, 0.02) 0.863(rs72613567) _

[0465] cTl was better at distinguishing cirrhosis cases from CLD cases without cirrhosis than liver fat or ALT.

[0466] Table 25: Comparison of the ability of biomarkers to distinguish subjects with cirrhosis from subjects with CLD but not cirrhosis.

[0467] Subjects with liver cTl data and subjects without liver cTl data did not have any systematic differences in HSD17B13 or PNPLA3 variant dosage, or in ALT.

[0468] Table 26: Comparison of subjects with vs. without liver cTl data available.Has liver cTl # subjects Mean # HSD17B13 Mean # Mean ALT data? splice variant PNPLA3 148M (winsorized) alleles allelesFALSE 333,415 0.56 0.43 23.4TRUE 25,101 0.55 0.43 22.8

[0469] HSD17B13 x liver fat, cALT PRS compared to HSD17B13 x PNPLA3

[0470] The present three CLD-related PRS (liver fat, 2-SNP cALT, and 41-SNP cALT) modified the protective effect of HSD17B13 loss of function on ALT levels, but the PRS- HSD17B13 interactions were not as strong as the HSD17B13 -PNPLA3 interaction. This suggested that the statistical HSD17B13 -PNPLA3 interaction was not wholly attributable to PNPLA3 genotype acting as a proxy for liver fat as an aggregate measure. Without being bound by a particular theory, PNPLA3 may affect regional liver adiposity or subcellular structure of lipid droplets (conditional on overall liver fat) that are uniquely relevant to liver health - more-so than other loci that modulate aggregate liver fat.

[0471] Table 27: Comparison of variables that interact with HSD17B13 splice variant (rs72613567) genotype in determining winsorized ALT levels. The units of all of the candidate interactors were standard deviations. UK Biobank subjects who had abdominal MRI data and therefore could have been included in the liver fat PRS training dataset wereexcluded from this analysis. The analysis included BMI as covariate. Each PRS was adjusted to remove the effect of HSD17B13 rs72613567 (and PNPLA3 I148M, if specified).

[0472] Excluding PNPLA3 from the liver fat and cALT PRS, there was a statistically significant but relatively weak interaction with HSD17B13 in PNPLA3 I148M non-carriers. The already-weak interaction was almost completely attenuated in PNPLA3 I148M carriers.

[0473] Table 28: Interactions of the liver fat and cALT PRS with the HSD17B13 splice variant (rs72613567) in determining winsorized ALT levels, stratified by PNPLA3 I148M genotype. The units of the PRS are standard deviations. The PRS were adjusted to remove the contributions of HSD17B13 and PNPLA3 to the PRS. UK Biobank subjects who had abdominal MRI data and therefore could have been included in the liver fat PRS training dataset were excluded from the analysis. The analysis included BMI as covariate.Variable interacting Interaction effect / P-PNPLA3 Carrier with HSD17B13 marginal HSD17B 13 effect valueNo Genome-wide liver fat 29.8% 2.9e-PRS 41-SNP cALT PRS 03(excludingNo 41-SNP cALT PRS 28.6 2.9e-(excluding PNPLA3) 03Yes 41-SNP cALT PRS 12.4% 0.040(excluding PNPL A3)Yes Genome-wide liver fat 10.4% 0.094PRS (excluding PNPLA3)

[0474] Subjects who were not PNPLA3 148M carriers, but who were in the top 20% of the liver fat PRS (excluding PNPLA3 ), were predicted to gain approximately the same benefit from the HSD17B13 LoF variant as the all-comer population. By contrast, PNPLA3 148M carriers were predicted to gain 1.5x the benefit from HSD17B13 LOF variant as the all-comer population. This suggested that the PNPLA3 genotype alone may suffice for enrollment criteria of a clinical trial for HSI) 17 13 inhibitor.

[0475] Table 29: Selecting subjects who are PNPLA3 I148M non-carriers but were in the top N% of the liver fat PRS: what % of the CLD patient population did this correspond to, and what was the effect of the HSD17B13 splice variant (rs72613567) relative to its effect in PNPLA3 carriers and to its effect in all-comers.

[0476] The analysis included BMI as a covariate. For comparison, 46.1% of CLD patients in UK Biobank are PNPLA3 carriers, and the HSD17B13 effect in PNPLA3 carriers was 1.53x the HSD17B13 effect in all-comers.Quantile of liver fat % of CLD cases HSD17B13 effect HSD17B13 effect PRS (excluding who are in the PRS in PNPLA3 nonin PNPLA3 nonPNPLA3) quantile but are not carriers / effect in carriers / effect in PNPLA3 carriers un-PRS-selected all-comers PNPLA3 carriers

[0477] HSD17B13 x PNPLA3 x BMI and HSD17B13 x PNPLA3 x diabetes

[0478] BMI interacts with both HSD17B13 and PNPLA3 in determining ALT. The effect of one PNPLA3 148M allele in an obese subject with functional HSD17B13 was expected to be -70% greater than in an overweight-but- not-obese subject with functional HSD17B13. The effect of one HSD17B13 loss of function allele was expected to be -60% greater in an obese PNPLA3 148M heterozygote compared to an overweight-but-not-obese PNPLA3 148M heterozygote.

[0479] Table 30: Interactions between HSD17B13 rs72613567 genotype, PNPLA3 rs738409 genotype, and BMI (centered so that mean = 0 and scaled in units of 5, i.e. 1 unit = difference between borderline overweight BMI = 25 and borderline obese BMI = 30), in determining winsorized ALT.Term Beta P-valueBMI (units of 5) 2.96 (2.89, 3.02) 0.0e+00PNPLA3 1.86 (1.77, 1.94) 0.0e+00PNPLA3 : BMI (units of 1.32 (1.23, 1.42) 2.7e-1795)HSD17B13 : PNPLA3 -0.51 (-0.61, -0.40) 2.1e-22I l lHSD17B13 -0.36 (-0.43, -0.29) 1.3e-21HSD17B13 : PNPLA3 : -0.37 (-0.48, -0.26) 1.9e-l lBMI (units of 5)HSD17B13 : BMI (units of -0.19 (-0.27, -0.12) 9.4e-075) _

[0480] If a BMI PRS was considered instead of observed BMI, the PNPLA3 -BMI interaction remained significant, indicating that BMI was a causal modifier of PNPLA3 effects on liver function. The HSD17B13 -BMI interaction however lost significance.

[0481] Table 31: Interactions between HSD17B13 rs72613567 genotype, PNPLA3 rs738409 genotype, and a BMI PRS (units = standard deviations), in determining winsorized ALT.

[0482] The BMI PRS was adjusted by linear regression to remove correlation with HSD17B13 and PNPLA3.Term Beta P-valuePNPLA3 1.83 (1.74, 1.91) 0.0e+00BMI PRS 0.56 (0.49, 0.62) 9.1e-64HSD17B13 : PNPLA3 -0.50 (-0.61, -0.40) 7.6e-21HSD17B13 -0.35 (-0.42, -0.27) 8.9e-19PNPLA3 : BMI PRS 0.16 (0.07, 0.25) 5.1e-04HSD17B 13 : BMI PRS -0.03 (-0.11, 0.04) 0.410HSD17B13 : PNPLA3 : -0.00 (-0.11, 0.10) 0.951BMI PRS

[0483] Applicant repeated the above analyses using diabetes case-control status (including both type 1 and type 2 diabetes) instead of BMI. The effect of one PNPLA3 148M allele in a diabetic subject with functional HSD17B13 was expected to be more than double the effect in a non-diabetic subject with functional HSD17B13. The effect of one HSD17B13 loss of function allele in a diabetic PNPLA3 148M heterozygote was also expected to more than double the effect in a non-diabetic PNPLA3 148M heterozygote.

[0484] Table 32: Interactions between HSD17B13 rs72613567 genotype, PNPLA3 rs738409 genotype, and diabetes in determining winsorized ALT.Term Beta P-valueDiabetes 5.11 (4.87, 5.35) 0.0e+00PNPLA3 1.65 (1.56, 1.74) 2.7e-271PNPLA3 : Diabetes 2.15 (1.82, 2.48) 5.2e-37HSD17B13 : PNPLA3 -0.45 (-0.56, -0.34) 4.6e-16HSD17B13 -0.33 (-0.40, -0.25) 7.5e-16HSD17B13 : PNPLA3 : -0.65 (-1.04, -0.25) 1.2e-03DiabetesHSD17B13 : Diabetes -0,34 (-0,62, -0,05) 0,020

[0485] Using a type 2 diabetes PRS in the analysis instead of observed diabetes status strongly supported the hypothesis that diabetes is a causal modifier of PNPLA3 effects on liver function. For PISD17B13, the analysis was inconclusive due to limited power.

[0486] Table 33: Interactions between HSD17B13 rs72613567 genotype, PNPLA3 rs738409 genotype, and a T2D PRS (units = standard deviations), in determining winsorized ALT.

[0487] The T2D PRS was adjusted by linear regression to remove correlation with HSD17B13 and PNPLA3 genotype.Term Beta P-valuePNPLA3 1.83 (1.74, 1.91) 0.0e+00T2D PRS 0.88 (0.82, 0.94) 1.9e-157HSD17B13 : PNPLA3 -0.50 (-0.61, -0.40) 5.9e-21HSD17B13 -0.35 (-0.42, -0.27) 7.8e-19PNPLA3 : T2D PRS 0.23 (0.14, 0.32) 2.6e-07HSD17B13 : T2D PRS -0.05 (-0.13, 0.03) 0.192HSD17B13 : PNPLA3 : -0.05 (-0.16, 0.05) 0.315T2D PRS

[0488] Including both BMI and diabetes, or both the BMI and T2D PRS, in a single model showed that these liver disease risk factors are each a causal modifier of PNPLA3 effects on liver function, independent of each other. Observed BMI and diabetes were also independent modifiers of HSD17B13 effects on ALT, but the PRS analysis was too underpowered to confirm the causality of these interactions.

[0489] Table 34: Interactions between HSD17B13 rs72613567 genotype, PNPLA3 rs738409 genotype, BMI, and diabetes in determining winsorized ALT.Term Beta P-valueBMI (units of 5) 2.78 (2.72, 2.85) 0.0e+00PNPLA3 1.78 (1.69, 1.87) 0.0e+00PNPLA3 : BMI (units of 1.24 (1.15, 1.34) 1.3e-1495)Diabetes 2.84 (2.60, 3.09) 8.0e-117HSD17B13 : PNPLA3 -0.48 (-0.59, -0.38) 4.0e-19HSD17B13 -0.34 (-0.42, -0.27) 2.3e-18HSD17B13 : PNPLA3 : -0.34 (-0.46, -0.23) 1.2e-09BMI (units of 5)PNPLA3 : Diabetes 0.96 (0.63, 1.29) 1.5e-08HSD17B13 : BMI (units of -0.18 (-0.26, -0.11) 5.4e-065)Term Beta P-valueHSD17B13 : Diabetes -0.26 (-0.54, 0.03) 0.079HSD17B13 : PNPLA3 : -0.25 (-0.64, 0.14) 0.212Diabetes

[0490] Table 35: Interactions between HSD17B13 rs72613567 genotype, PNPLA3 rs738409 genotype, a BMI PRS, and a T2D PRS in determining winsorized ALT.

[0491] The units of the PRS are standard deviations. The PRS were adjusted by linear regression to remove correlation with HSD17B13 and PNPLA3 genotype.Term Beta P-valuePNPLA3 1.82 (1.74, 1.91) 0.0e+00T2D PRS 0.80 (0.73, 0.87) 1.0e-125BMI PRS 0.39 (0.33, 0.46) 5.7e-32HSD17B13 : PNPLA3 -0.50 (-0.61, -0.40) 8.4e-21HSD17B13 -0.35 (-0.42, -0.27) 6.7e-19PNPLA3 : T2D PRS 0.21 (0.12, 0.30) 4.5e-06PNPLA3 : BMI PRS 0.12 (0.03, 0.21) 0.011HSD17B13 : T2D PRS -0.05 (-0.13, 0.03) 0.237HSD17B13 : PNPLA3 : -0.06 (-0.16, 0.05) 0.304T2D PRSHSD17B13 : BMI PRS -0.02 (-0.10, 0.05) 0.531HSD17B13 : PNPLA3 : 0.01 (-0.10, 0.12) 0.879BMI PRS

[0492] Looking at cirrhosis as a binary phenotype, despite very limited power, Applicant observed nominal significance for the diabetes- / A7J / A3 interaction.

[0493] Table 36: Interaction between PNPLA3 rs738409 genotype and Diabetes in determining risk of cirrhosis.Term OR P-valueDiabetes 5.38 (4.77, 6.06) 5.1e-168PNPLA3 1.47 (1.36, 1.59) 5.3e-23PNPLA3 : Diabetes 1.17 (1.03, 1.34) 0.020

[0494] Table 37: Interaction between PNPLA3 rs738409 genotype and HSD17B13 rs72613567 genotype in determining risk of cirrhosis.Term OR P-valuePNPLA3 1.52 (1.40, 1.65) 4.3e-24HSD17B13 rs72613567 0.89 (0.81, 0.97) 6.7e-03 genotypePNPLA3 : HSD17B13 1.04 (0.94, 1.14) 0.494 rs72613567 genotype

[0495] Table 38: Comparison of BMI, WHR, and BMI-adjusted WHR as correlates of liver function.

[0496] Separate linear regression models were fit to predict winsorized ALT, including the present standard covariates, PNPLA3 rs738409 genotype, HSD17B13 rs72613567 genotype, and one of the three adiposity metrics, with 2-and 3 -way interaction terms between the genotypes and the adiposity metric. Betas and p-values were presented for selected terms in each model.Model TestTerm Beta statistics P- valueBMI BMI (units of 5) 2.96 (2.90, 88.5 0.0e+003.03)BMI PNPLA3 : BMI 1.33 (1.24, 28.7 2.5e-180(units of 5) 1.42)BMI HSD17B13 : BMI -0.19 (-0.27, -4.9 8.7e-07(units of 5) -0.12)BMI HSD17B13 : -0.37 (-0.48, -6.8 1.0e-l lPNPLA3 :BMI (units -0.27) of 5)WHR WHR (units of 0.1) 4.06 (3.98, 100.3 0.0e+004.14)WHR PNPLA3 : WHR 1.18 (1.08, 24.1 3.5e-128(units of 0.1) 1.28)WHR HSD17B13 : WHR -0.17 (-0.25, -4.0 6.6e-05(units of 0.1) -0.09)WHR HSD17B13 : -0.38 (-0.49, -6.5 l. le-10PNPLA3 WHR -0.26)(units of 0.1)BMIadjWHR BMI-adjusted WHR 2.31 (2.19, 40.5 0.0e+00(units of 0.1) 2.42)BMIadjWHR PNPLA3 : BMI- 1.00 (0.85, 12.8 1.6e-37 adjusted WHR (units 1.15) of 0.1)BMIadjWHR HSD17B13 : BMI- -0.07 (-0.20, -1.1 0.272 adjusted WHR (units 0.06) of 0.1)BMIadjWHR HSD17B13 : -0.36 (-0.55, -3.9 8.9e-05PNPLA3 : BMI- -0.18) adjusted WHR (units of 0.1)(BMI, (BMI, BMIadjWHR) 3.11 (3.04, 98.9 0.0e+00BMIadjWHR) combined statistic 3.17) combined (s.d.) statistic (BMI, PNPLA3 : (BMI, 1.36 (1.28, 31.3 2.3e-214BMIadjWHR) BMIadjWHR) 1.45) combined combined statistic statistic (s.d.)(BMI, HSD17B13 : (BMI, -0.17 (-0.24, -4.5 5.5e-06BMIadjWHR) BMIadjWHR) -0.10) combined combined statistic statistic(BMI, HSD17B13 : -0.44 (-0.54, -8.5 2.7e-17BMIadjWHR) PNPLA3 : (BMI, -0.34) combined BMIadjWHR)combin statistic ed statistic (s.d.)

[0497] Applicant repeated the above analysis using MRI-derived adiposity statistics: liver proton density fat fraction (PDFF), abdominal subcutaneous adipose tissue volume (ASAT), and visceral adipose tissue volume (VAT). Compared to BMI, liver PDFF is a better overall predictor of ALT, has a stronger interaction with HSD17B13, but a weaker interaction with PNPLA3. VAT is similar to WHR in that it has weaker genetic interactions than BMI despite being a better overall predictor of ALT (though not as good as liver PDFF). ASAT does worse than BMI across the board.

[0498] The weak interaction of liver PDFF with PNPLA3 compared to the strong interaction between BMI and PNPLA3 was intriguing. PNPLA3 may regulate the proportion of total fat in the body that is stored in the liver; thus, it interacted with BMI in determining ALT and liver fat (see the table after the next table), but given a liver fat measurement, the interaction between liver fat and PNPLA3 was weak.

[0499] Log-transformed liver PDFF gives a -67% stronger interaction with HSD17B13 than BMI. This supported the idea that patients with rare diseases like familial partial lipodystrophy that have a predisposition to highly elevated liver fat are likely to respond well to an HSD17 13 inhibitor.

[0500] Table 39: Comparison of BMI, VAT, ASAT, and raw and log- transformed liver PDFF as they correlate to liver function.

[0501] Separate linear regression models were fit to predict winsorized ALT, including the present standard covariates, PNPLA3 rs738409 genotype, HSD17B13 rs72613567 genotype, and one of the five adiposity metrics, with interaction terms between the genotypes and the adiposity metric. Only subjects with measurements for all of the adiposity variables were included in the models. Betas and p-values were presented for selected terms in each model.Model Term Beta Test P-valuestatisticBMI BMI (s.d.) 2.91 (2.71, 3.11) 28.9 1.6e-180BMI PNPLA3 : BMI (s.d.) 1.10 (0.87, 1.32) 9.6 8.9e-22BMI HSD17B13 : BMI -0.21 (-0.42, - -2.0 0.043(s.d.) 0.01)VAT VAT (s.d.) 3.51 (3.29, 3.72) 32.5 9.6e-228VAT PNPLA3 : VAT (s.d.) 1.01 (0.79, 1.23) 9.0 2.3e-19VAT HSD17B13 : VAT -0.17 (-0.38, -1.6 0.103(s.d.) 0.03)ASAT ASAT (s.d.) 2.02 (1.82, 2.23) 19.0 2.9e-80ASAT PNPLA3 : ASAT 0.73 (0.50, 0.96) 6.3 3.2e-10(s.d.)ASAT HSD17B13 : ASAT -0.03 (-0.24, -0.2 0.812(s.d.) 0.19)Liver PDFF Liver PDFF (s.d.) 3.74 (3.54, 3.95) 35.2 1.4e-265Liver PDFF PNPL A3 : Liver PDFF 0.23 (0.03, 0.43) 2.3 0.021(s.d.)Liver PDFF HSD17B 13 : Liver -0.25 (-0.45, - -2.4 0.017PDFF (s.d.) 0.04)Log liver Log liver PDFF (s.d.) 3.70 (3.50, 3.91) 35.5 7.9e-270PDFFLog liver PNPLA3 : Log liver 0.65 (0.44, 0.86) 6.1 9.0e-10PDFF PDFF (s.d.)Log liver HSD17B 13 : Log liver -0.35 (-0.55, - -3.3 9.0e-04PDFF PDFF (s.d.) 0.14)

[0502] Table 40: Interaction between PNPLA3 and BMI in determining liver fat fraction. The units of the Beta column are percentage points of liver fat.Term Beta P-valueBMI (units of 5) 1.98 (1.91, 2.06) 0.0e+00PNPLA3 I148M (rs738409) 1.27 (1.18, 1.36) 2.1e-168BMI : PNPLA3 0.59 (0.48, 0.69) 1.5e-27

[0503] Table 41: Effect of HSD17B13 splice variant (rs72613567) dosage on winsorizedALT, stratified by quantiles of liver PDFF. The liver PDFF data was adjusted to remove the effect of PNPLA3 I148M (rs738409).Percentile of liver PDFF Effect on ALT (U / L) P-value N0-19 -0.06 (-0.42, 0.29) 0.733 572320-39 -0.08 (-0.45, 0.29) 0.664 617240-59 -0.50 (-0.88, -0.12) 0.010 597760-79 -0.36 (-0.81, 0.08) 0.110 609380-100 -0.97 (-1.56, -0.39) l. le-03 6024

[0504] A model of HSD17B13 interacting with PNPLA3, liver PDFF, and BMI provided tentative evidence that the HSD17B13 -PNPLA3 and HSD17B13 -BMI interactions Applicantdescribed previously were mediated by effects of PNPLA3 and BMI on liver fat. This was consistent with the inability to confirm the causal effect of BMI on HSD17B13 response with BMI PRS analysis. Although the PNPLA3 effect on HSD17B13 response may be completely mediated by liver fat, in certain applications there was limited scope for improvement by incorporating other genetic loci beyond PNPLA3 and thus PNPLA3 may have predictive value as a single marker.

[0505] Table 42: Interaction of HSD17B13 rs72613567 with log-transformed liver PDFF, BMI, and PNPLA3 rs738409 in determining winsorized ALT. term Beta P-valueLog liver PDFF (s.d.) 3.01 (2.83, 3.20) 8.4e-215BMI (units of 5) 2.43 (2.21, 2.66) 5.0e-103PNPLA3 0.62 (0.34, 0.90) 1.2e-05HSD17B13 -0.40 (-0.64, -0.15) 1.3e-03HSD17B 13 : Log liver PDFF (s.d.) -0.23 (-0.45, -0.01) 0.039HSD17B13 : BMI (units of 5) -0.13 (-0.39, 0.13) 0.334HSD17B13 : PNPLA3 -0.09 (-0.43, 0.24) 0.590Example 2: CLD to cirrhosis progression analysis using risk analysis of Individual Loci

[0506] PNPLA3 I148M had a modest but statistically significant effect on progression from an inpatient chronic liver disease (CLD) diagnosis to a composite endpoint comprised of inpatient cirrhosis or hepatocellular carcinoma diagnosis, procedures for drainage of ascites, liver transplant, or liver-related mortality. This effect was greatly amplified (>100%) by a subject having a prior diabetes diagnosis. A PNPLA3 148M heterozygote with diabetes had ~1.8x the progression risk of PNPLA3 1481 homozygote without diabetes.

[0507] Subjects who have at least one PNPLA3 148M allele and prior diabetes have a very high absolute risk of CLD progression, with 20% progress within 2 years and 42% progress within 5 years.

[0508] BMI also modified the effect of PNPLA3 on CLD progression, but independent of PNPLA3, it “protects” against progression-which is almost certainly an artifact of collider bias. Regardless, Applicant estimates that a borderline-obese (BMI = 30) PNPLA3 148M heterozygote does not have a greater risk of CLD progression than a borderline-overweight (BMI = 25) PNPLA3 heterozygote.

[0509] HSD17B13 by itself has a nominally significant (p < 0.05) protective effect against CLD progression. There does not appear to be any significant or even suggestive interaction between HSD17B13 and PNPLA3 in determining CLD progression, and so these may be of independent predictive value.

[0510] Table 43: Evaluation of candidate predictors of CLD progression risk. Separate Cox proportional hazards models were fit including covariates (age at CLD diagnosis, sex, genotyping chip, ancestry PCs 1-6) plus one of the candidate predictors at a time.Predictor HR P-valuePNPLA3 I148M 1.16 (1.09, 1.23) 4.0e-06TM6SF2 E167K 1.12 (1.02, 1.23) 0.019HSD17B 13 splice variant 0.93 (0.87, 0.99) 0.034Prior diabetes 1.52 (1.37, 1.68) 1.6e-15BMI (units of 5) 0.95 (0.91, 0.99) 7.0e-03

[0511] Table 44: Comparison of prior diabetes, BMI, and HSD17B13 as modifiers of PNPLA3 contribution to CLD progression risk.

[0512] Separate Cox proportional hazards models were fit including covariates (age at CLD diagnosis, sex, genotyping chip, ancestry PCs 1-6), PNPLA3 rs738409 genotype, one of the three candidate PNPLA3 modifiers at a time, and a PNPLA3 -modifier interaction term. Hazard ratios and p-values are shown for selected terms in each model.Model Term HR P-valuePrior diabetes PNPLA3 1.11 (1.03, 1.19) 3.7e-03Prior diabetes Prior diabetes 1.35 (1.18, 1.55) 2.2e-05Prior diabetes PNPLA3 : prior diabetes 1.21 (1.05, 1.40) 0.011BMI (units of 5) PNPLA3 1.08 (1.01, 1.16) 0.030BMI (units of 5) BMI (units of 5) 0.88 (0.83, 0.93) 1.5e-06BMI (units of 5) PNPLA3 : BMI (units of 5) 1.14 (1.07, 1.21) 2.1e-05HSD17B13 PNPLA3 1.13 (1.04, 1.22) 3.0e-03HSD17B13 HSD17B13 0.91 (0.84, 0.99) 0.038HSD17B13 PNPLA3 : HSD17B13 1.04 (0.95, 1.15) 0.395

[0513] Table 45: UK Biobank CLD cases were scored using 3 SNPs in PNPLA3, TM6SF2, and HSD17B13, weighted by the SNPs’ effects on risk of having chronically elevated ALT in a GWAS from the Million Veterans Program.

[0514] Based on the distribution of the score, Applicant manually defined high-risk subjects as those with a score >= +1.25 s.d. from the mean, medium-risk subjects as those with a score >= -0.35 s.d. from the mean, and low-risk subjects as the remainder. This table shows the hazards of progression of CLD to a composite endpoint of cirrhosis, hepatocellularcarcinoma, drainage of ascites, liver transplant, or liver-related death, stratified by genetic risk group.Risk bin from 3 -SNP score HR P-value # CLD casesLow 1.00 (reference) NA 3,443Medium 1.07 (0.98, 1.17) 0.122 3,243High 1.34 (1.20, 1.50) 4.7e-07 1,144

[0515] Using models of ALT as a single-timepoint snapshot of liver function, and models of CLD progression to cirrhosis, Applicant showed that CLD patients with prior diabetes were likely to derive substantially greater benefit from a PNPLA3 degrader than subjects without prior diabetes. The present analysis using a T2D PRS in place of observed diabetes status confirmed that this interaction is likely to be causal.

[0516] Body mass index was also a causal modifier PNPLA3 response, independent of diabetes. Due to the comorbidity of diabetes with obesity, selecting patients based on diabetes alone would probably be a sufficient and more straightforward way to improve PNPLA3 degrader response rates compared to considering both diabetes and BMI. The PNPLA3 -BMI interaction is unchanged by conditioning on liver fat measurements from MRI data, indicating that the interaction is likely driven by systemic aspects of metabolism rather than local hepatic metabolism.

[0517] Applicant successfully replicated the report of Abul-Husn, Noura S., et al. New England Journal of Medicine 378.12 (2018): 1096-1106, A iP[SD17B13 loss of function mitigates ALT elevation from PNPLA3 I148M. Applicant found that HSD17B13 loss of function by itself protected against CLD progression to cirrhosis (p < 0.05), a novel result.

[0518] Applicant found that across all PNPLA3 genotypes (I / I, I / M, M / M), one HSD17B13 loss of function allele appears to have -75% of the effect of two HSD17B13 loss of function alleles, a non-additive effect. This result implied that a trial of an HSD17B13 inhibitor could meaningfully improve its expected treatment effect by selecting patients with two functional HSD17B13 alleles and at least one PNPLA3 148M allele, compared to selecting on only PNPLA3 or taking all-comers.

[0519] Applicant showed that HSD17B13 has a greater protective effect on ALT in subjects with elevated liver fat. The HSD17B13 -PNPLA3 interaction seems to be mediated by PNPI 3A effect on liver fat. Combining BMI, a measure of total adiposity, and BML adjusted waist-to-hip ratio, a measure of adipose tissue distribution, could potentially enable rough prediction of liver fat in the general population that is accurate enough to be useful forHSD17B13 inhibitor patient selection. More immediately actionable is the idea that HSD17B13 inhibition could be effective in patients with familial partial lipodystrophy, who are predisposed to liver fat accumulation.

[0520] Genome-wide polygenic risk scores for liver fat or chronically elevated ALT did not outperform the observational measures of adiposity or PNPLA3 genotype alone in predicting HSD17B13 effects on ALT. Excluding PNPLA3 from the PRS, the residual PRS did interact with HSD17B13, but only in PNPLA3 I148M non-carriers; and even in this subgroup, the magnitude of the interaction was modest. In certain applications, it may be beneficial to weight the PRS so that it represents a biological pathway or quantity that is more directly relevant to HSD17B13 function than liver fat in...

Claims

WHAT IS CLAIMED IS:

1. An in silico method for determining a predicted drug activity of a plurality of drug targets associated with at least one disease phenotype associated with non-alcoholic steatotic hepatitis (NASH), the method comprising: obtaining molecular biomarker stratifier data comprising a plurality of biomarker stratifiers and at least one disease phenotype associated with NASH from each of a plurality of subjects; determining a plurality of values representing one or more biomarker stratifier effects, wherein each value of the plurality separately represents how each of the plurality of biomarker stratifiers affects a disease phenotype associated with NASH based on external data; calculating a biomarker stratifier score for a chosen disease phenotype; calculating a pharmacomimetic genetic score for each drug target; and identifying the predicted drug activity of each drug target for at least one disease phenotype associated with NASH in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the at least one disease phenotype associated with NASH.

2. The method of claim 1, wherein the molecular biomarker stratifier data comprises a genetic variant and wherein the biomarker stratifier score comprises one or more score selected from:(i) a polygenic score for a disease phenotype;(ii) a polygenic score for a disease risk factor;(iii) a polygenic score for a biological pathway activity;(iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

3. The method of claim 1 or 2, wherein the biomarker stratifier data comprises a proteomic variant, and wherein the biomarker stratifier score comprises one or more score selected from:(i) a proteomics score for a disease phenotype;(ii) a proteomics score for a disease risk factor;(iii) a proteomics score for a biological pathway activity;(iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

4. The method of claim 1 or claim 2, wherein the biomarker stratifier data comprises a transcriptional variant, and wherein the biomarker stratifier score comprises one or more scores selected from:(i) a transcriptomics score for a disease phenotype;(ii) a transcriptomics score for a disease risk factor;(iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

5. The method of claim 1 or claim 2, wherein the biomarker stratifier data comprises a polymorphic variant, and wherein the biomarker stratifier score comprises one or more scores selected from:(i) a polymorphism score for a drug target expression;(ii) a polymorphism score for a disease phenotype;(iii) a polymorphism score for a disease risk factor;(iv) a polymorphism score for a biological pathway activity;(v) a polymorphism score for a drug target expression; or any linear or non-linear combination thereof.

6. The method of claim 1, wherein the biomarker stratifier data comprises a polymorphic variant, wherein the biomarker stratifier data comprises a genetic variant, a proteomic variant, a transcriptional variant and / or a polymorphic variant, and wherein the biomarker stratifier score comprises one or more score selected from:(i) a polygenic score for a disease phenotype;(ii) a polygenic score for a disease risk factor;(iii) a polygenic score for a biological pathway activity;(iv) a polygenic score for a drug target expression;(v) a proteomics score for a disease phenotype;(vi) a proteomics score for a disease risk factor;(vii) a proteomics score for a biological pathway activity;(viii) a proteomics score for a biological pathway activity;(ix) a proteomics score for a drug target expression;(x) a transcriptomics score for a disease phenotype;(xi) a transcriptomics score for a disease risk factor;(xii) a transcriptomics score for a biological pathway activity;(xiii) a polymorphism score for a drug target expression;(xiv) a polymorphism score for a disease phenotype;(xv) a polymorphism score for a disease risk factor;(xvi) a polymorphism score for a biological pathway activity;(xvii) a polymorphism score for a drug target expression; or any linear or non-linear combination thereof.

7. The method of claim 1 or claim 2, further comprising running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for the one or more of the biomarker stratifier effects.

8. The method of claim 7, further comprising running the principal components analysis or the weighted principal component analysis based on a matrix of the one or more of the biomarker stratifier effects.

9. The method of claim 1 or claim 2, wherein the pharmacomimetic genetic score is associated with a response to a drug that targets HSD17B13.

10. The method of claim 2 or 6, wherein the polygenic score is a body mass index polygenic score.

11. A system for drug development for non-alcoholic steatotic hepatitis (NASH), comprising one or more processors configured by computer-readable instructions to: generate a plurality of principal components (PCs) corresponding to genetic variant data associated with non-alcoholic steatotic hepatitis (NASH) for entities in a dataset; generate a biomarker stratifier score for each entity in the dataset to reduce a size of the dataset based at least on the PCs and biomarker stratifier weights for NASH;determine one or more genetic variant that arepharmacomimetic based on the datasets; and determine one or more interaction between the pharmacomimetic instruments and the biomarker stratifier scores, wherein the interactions are indicative of a drug response for one or more drugs.

12. The system of claim 11, further comprising performing one or more actions based on the determined interaction between the pharmacomimetic instruments and the biomarker stratifier scores.

13. The system of claims 11 or 12, wherein the PCs analysis corresponding to genetic variant data is based on a matrix of the one or more of the biomarker stratifier score.

14. The system of claim 13, further comprising constructing the matrix based on the biomarker stratifier score, the phenotypes, and the external data from a plurality of entities.

15. An in silico method for determining a predicted drug activity of one or more drug targets on a plurality of non-alcoholic steatotic hepatitis (NASH) phenotype intermediates, comprising: obtaining molecular biomarker stratifier data and NASH disease phenotypes from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each NASH phenotype intermediate based on external data; calculating a biomarker stratifier score for each NASH phenotype intermediate in a disease process; calculating a pharmacomimetic genetic score for each drug target on each NASH phenotype intermediate; and identifying predicted drug activity of the one or more drug targets on the one or more NASH phenotype intermediates in one or more subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the NASH phenotype intermediate and a NASH outcome phenotype.

16. A method of determining predicted drug activity of an agent on reducing NASH progression and / or risk in a subject, comprising: obtaining molecular biomarker stratifier data and NASH progression data from a plurality of subjects; determining a value of one or more biomarker stratifier effects for NASH progression based on external data; calculating a biomarker stratifier score for a chosen drug outcome related to subjects’ NASH progression data; calculating a pharmacomimetic genetic score for the agent; and identifying the predicted drug activity of the agent on reducing NASH progression in one or more subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for the agent in association analysis with the NASH progression and / or risk in the subject.

17. The method of claim 16, wherein the biomarker stratifier data comprises a genetic variant, a proteomic variant, a transcriptional variant and / or a polymorphic variant.

18. The method of claim 17, further comprising running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

19. The method of claim 18, further comprising running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the one or more biomarker stratifier effects.

20. The method of claim 19, further comprising constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and external data from a plurality of subjects.

Citation Information

Patent Citations

  • Identification of genetic components of drug response

    US20070072232A1

  • Multiparameter analysis for drug response and related methods

    US20080131905A1

  • Machine learning platform for generating risk models

    US20220044761A1

  • Polygenic Score for Cardiac Heart Failure

    US20220333169A1

  • A method of precision treatment

    WO2023168499A1