Pharmacomimetic variant and PRS interactions in AMD and cnv

In silico systems utilizing PRS and pharmacomimetic variants address the challenges of AMD and CNV treatment heterogeneity, enabling efficient drug development and personalized treatment strategies.

WO2025240665A1PCT designated stage Publication Date: 2025-11-20FORESITE LABS LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/029438
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-15
Filing Date
2025-05-14
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

There is a need for improved systems and methods to treat age-related macular degeneration (AMD) and choroidal neovascularization (CNV), particularly in addressing the heterogeneity in disease progression and treatment response, which has hindered efficient drug development and clinical trial design.

Method used

The use of in silico systems and methods that incorporate polygenic risk scores (PRS) and pharmacomimetic variants to predict drug activity and treatment response, allowing for the development and repurposing of therapies without empirical data, and enabling patient-specific models for virtual clinical trials.

Benefits of technology

This approach accelerates drug development by identifying optimal treatments for specific patient populations, enhances treatment response prediction, and reduces the need for costly and time-consuming clinical trials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025029438_20112025_PF_FP_ABST
    Figure US2025029438_20112025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems, methods and computer program products for drug development using in silico techniques. An aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets by determining biomarker stratifier effects, calculating a biomarker stratifier score for a chosen disease phenotype; and calculating a pharmacomimetic genetic score using molecular biomarker stratifier data.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.136622-1010 PHARMACOMIMETIC VARIANT AND PRS INTERACTIONS IN AMD AND CNV CROSS-REFERENCE TO RELATED PATENT APPLICATION

[0001] This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No.63 / 648,067, filed May 15, 2024, and U.S. Provisional Application No. 63 / 648,133, filed May 15, 2024, the entire contents of each of which is incorporated herein by reference. FIELD OF THE DISCLOSURE

[0002] The present disclosure relates to systems, methods and computer program products to aid in the treatment of age related macular degeneration (AMD) and choroidal neovascularization (CNV) using in silico techniques. BACKGROUND OF THE DISCLOSURE

[0003] Modelling and simulation are rapidly evolving areas in terms of both technologies and application fields in the life sciences. The use of modeling and simulation has expanded beyond the description of drug exposure, towards the dynamic description of complex drug effects and disease subtypes and progressions. In recent years, new approaches in modelling and simulation have started to provide important insights in biomedicine, opening the way for their potential use in the reduction, refinement and partial substitution of both animal and human experimentation.

[0004] Age-related macular degeneration (AMD) is a leading cause of blindness. Advanced AMD has two main subtypes: choroidal neovascularization (CNV), or “wet” AMD, and geographic atrophy (GA), or “dry” AMD. Early and intermediate stage AMD are typically included in the “dry” category. In total, dry AMD accounts for 85-90% of AMD cases. There is a need for improved systems and methods for improving the treatment of AMD and CNV. SUMMARY OF ILLUSTRATIVE EMBODIMENTS

[0005] In this disclosure, methods and systems as described herein are applied to identify pharmacomimetic variant interactions with a complement polygenic score in association with dry age-related macular degeneration (AMD) and choroidal neovascularization. In one aspect, any interactions between a polygenic risk score (PRS) for AMD were tested with pharmacomimetic variants for several drug targets that are being developed for GA / dry AMD. A PRS x pharmacomimetic variant interaction, if statistically significant and ofAttorney Docket No.136622-1010 meaningful effect size, would indicate that AMD patients with high PRS might receive the most benefit from the drug that the variant is mimicking.

[0006] Polygenic scores (PGS), herein used interchangeably with polygenic risk scores (PRS), summarize genome-wide genotype data in combination with gene-phenotype association data, such as from genome-wide association studies (GWAS), into metrics that represent genetic liability to a trait. The present disclosure provides in silico systems, methods and computer program products using PRS for various aspects of drug development and optimization of patient treatment, including drug indication selection and clinical trial recapitulation. Specifically, the disclosure provides PRS-informed drug development systems to accelerate the development of new therapeutic products as well as repurposing existing therapies for new populations and / or indications. There are many applications of PRS-informed drug development, ranging from discovery of new drug targets for the treatment of a given indication, to repurposing of existing therapeutics for new indications, to new and efficient designs of clinical trials, to incorporation of new end points in ongoing clinical trials, to modifying how the data are analyzed to provide evidence of effectiveness in powering clinical trials, to modifying how existing therapeutics are used in patient populations. These systems utilize existing patient variant data in conjunction with computational models to mimic aspects of the drug development process. Importantly, the systems of the disclosed approach allow such development without the requirement of empirical data that has to date been required for various steps in the drug development process, from drug discovery through clinical trials and regulatory approval, although empirical data is optionally incorporated in certain embodiments.

[0007] In addition to the patient variant data, the computational models of the systems of the disclosure can optionally use mechanistic knowledge of physical, chemical, and / or environmental data related to patient populations and available biological and physiological knowledge of patient treatment (such as medical intervention) or activity (such as exercise or lack thereof).

[0008] In some aspects, the disclosure provides a preclinical drug discovery system for identifying and / or developing treatments (e.g., drugs) for various therapeutic areas and indications. The preclinical drug discovery system is configured to utilize mammalian polygenic scores with computational models to provide PRS-informed drug discovery methodologies.Attorney Docket No.136622-1010

[0009] In some aspects, the disclosure provides a drug discovery system for repurposing existing therapeutics for new therapeutic areas and indications. The system is configured to utilize mammalian polygenic scores with computational models to provide PRS-informed methodologies with known drug and patient information to find potential new uses of existing therapies.

[0010] In some embodiments, the disclosure provides in silico systems and methods for determining drug activity. In particular, the disclosure provides systems and methods for determining drug activity for particular patient populations, phenotypes, or for use in therapeutic indications / areas. This drug activity may be determined for known agents or for de novo agents.

[0011] Accordingly, in some aspects, the disclosure provides an in silico method for determining drug activity of drug targets on a plurality of phenotype intermediates, comprising, or alternatively consisting essentially of, or yet further consisting of obtaining data on genetic variants and phenotypes from each of a plurality of subjects, determining the value of the variant effects for each phenotype intermediate based on external data, calculating a polygenic score for each phenotype intermediate in the disease process, calculating a pharmacomimetic genetic score for each drug target on each phenotype intermediate, and identifying predicted drug activity of the drug targets on one or more phenotypes based on the statistical interaction of the polygenic score with the pharmacomimetic genetic score for each drug target in association analysis with the phenotype intermediate and the outcome phenotype. In some aspects, the disclosure provides a system for providing in silico clinical trials for new therapeutics using mammalian polygenic scores with computational models. Historically, safety and efficacy data provided to regulatory agencies in support of marketing authorization of a new medical product have been produced experimentally, either in vitro or in vivo. More recently, regulatory agencies started receiving and accepting in silico data produced using modelling and simulation, and the data provided by the systems of the disclosure can provide data supporting regulatory filings without the need for expensive and time-consuming experimentation.

[0012] The systems of the disclosure allow development of patient-specific models to form virtual cohorts for testing the safety and / or efficacy of new drugs and of new medical devices.

[0013] In certain embodiments, the disclosure provides in silico methods for determining drug activity, the methods comprising, or alternatively consisting essentially of, or yet furtherAttorney Docket No.136622-1010 consisting of obtaining data on genetic variants and phenotypes from each of a plurality of subjects; determining the value of each variants’ effect on each phenotype based on external data from a non-overlapping set of subjects; constructing a matrix in which the variants are rows, the phenotypes are columns, and the values are the externally-derived variant effects; running a truncated principal components analysis on this matrix, computing a user-defined number of principal components, typically two to five; calculating a polygenic score for each computed principal component by reweighting an existing polygenic score for a user-defined “reference” phenotype such that the new weight for each variant included the polygenic score is equal to its old weight times its squared loading for the principal component in question divided by the sum of its squared loadings across all computed principal components; and lastly identifying predicted drug activity for the “reference” phenotype based on each of the reweighted polygenic scores, which often can be interpreted as corresponding to biological pathways . This example method is set forth visually in Figure 1.

[0014] In certain embodiments, the disclosure provides in silico systems for drug development, such systems comprising, or alternatively consisting essentially of, or yet further consisting of at least one hardware processor and a non-transitory computer-readable storage medium having stored thereon program code, the program code executable by the at least one hardware processor to obtain data on genetic variants and phenotypes from each of a plurality of subjects, determine the value of the variant effects for each phenotype based on external data, construct a matrix based on the variants, phenotypes and values, run a principal components analysis on the variant x phenotype-related phenotype variant effects matrix, calculate a polygenic score for each principal component, and identify predicted drug effects on one or more phenotype based on the polygenic score calculated for each principal component.

[0015] In certain embodiments, the disclosure provides in silico systems for determination of intermediate phenotypes, such as risk factors or genetic variations that identify responders versus non-responders for a specific therapeutic intervention. Accordingly, in certain aspects, the disclosure provides in silico systems and methods for determining drug activity comprising, or alternatively consisting essentially of, or yet further consisting of obtaining data on genetic variants and phenotypes from each of a plurality of subjects, determining the value of the variant effects for each phenotype based on external data, calculating a polygenic score for a chosen phenotype intermediate in the disease process (e.g., a risk factor, biomarker, or genetic indicator of response), calculating a pharmacomimetic genetic score forAttorney Docket No.136622-1010 each drug target; and identifying predicted drug activity on one or more phenotypes based on the statistical interaction of the polygenic score with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0016] In specific aspects, the disclosure provides methods in which the PRS information is provided in whole or in part from related disease phenotypes.

[0017] Accordingly, the disclosure provides a method of determining the drug activity of a drug target, comprising, or alternatively consisting essentially of, or yet further consisting of obtaining data on genetic variants and phenotypes from each of a plurality of subjects, determining the value of the variant effects for each phenotype based on external data, calculating a polygenic score for a chosen outcome phenotype related to the disease outcome, calculating a pharmacomimetic genetic score for the drug target, and identifying predicted drug activity of the drug target on one or more phenotypes based on the statistical interaction of the polygenic score with the pharmacomimetic genetic score for the drug target in association analysis with the outcome phenotype. The drug target may be, e.g., a known drug for the phenotype, a drug that has been approved or in trials for a different indication than the phenotype for which the analysis is being performed, or a de novo agent.

[0018] An aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets, the method comprising, or alternatively consisting essentially of, or yet further consisting of: obtaining molecular biomarker stratifier data comprising, or alternatively consisting essentially of, or yet further consisting of a plurality of biomarker stratifiers and at least one disease phenotype from each of a plurality of subjects; determining a plurality of values representing biomarker stratifier effects, wherein each value separately represents how each of the plurality of biomarker stratifiers affects each of the at least one disease phenotype based on external data; calculating a biomarker stratifier score for a chosen disease phenotype; calculating a pharmacomimetic genetic score for each drug target; and determining the predicted drug activity of each drug target the at least one disease phenotype in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the disease phenotype.

[0019] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or moreAttorney Docket No.136622-1010 score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0020] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0021] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0022] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity; (v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0023] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biologicalAttorney Docket No.136622-1010 pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0024] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

[0025] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0026] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and external data from a plurality of subjects.

[0027] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0028] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0029] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of the weighted principal component analysis identifies predicted drug activity of the plurality of drug targets for the at least one disease phenotype.

[0030] In some embodiments, the plurality of drug targets is associated with AMD or CNV.

[0031] Another aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets, the method comprising: obtaining molecular biomarker stratifier data and at least one disease phenotype from each of a plurality of subjects;Attorney Docket No.136622-1010 determining the value of the biomarker stratifier effects for each disease phenotype based on external data; constructing a matrix based on the biomarker stratifiers, disease phenotypes and values; running a principal components analysis biomarker stratifier x phenotype-related phenotype biomarker stratifier effects matrix; calculating a biomarker stratifier score for each principal component; calculating a pharmacomimetic genetic score for each drug target; and identifying predicted drug activity of each drug target on one or more phenotypes in subsets of the biomarker stratifier distribution based on the statistical interaction of the polygenic score calculated for each principal component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0032] Another aspect of the disclosure is directed to an in silico method for determining the drug activity of drug targets on a plurality of phenotype intermediates, comprising: obtaining molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each phenotype intermediate based on external data; calculating a biomarker stratifier score for each phenotype intermediate in the disease process; calculating a pharmacomimetic genetic score for each drug target on each phenotype intermediate; and identifying predicted drug activity of the drug targets on one or more phenotypes in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the phenotype intermediate and the outcome phenotype.

[0033] Another aspect of the disclosure is directed to method of determining the drug activity of a drug target, comprising: obtaining molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects;Attorney Docket No.136622-1010 determining the value of the biomarker stratifier effects for each phenotype based on external data; calculating a biomarker stratifier score for a chosen outcome phenotype related to the disease outcome; calculating a pharmacomimetic genetic score for the drug target; and identifying predicted drug activity of the drug target on one or more phenotypes in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for the drug target in association analysis with the outcome phenotype.

[0034] Another aspect of the disclosure is directed to an in silico method comprising: generating a plurality of principal components (PCs) corresponding to genetic ancestry data for subjects in a study cohort; generating a biomarker stratifier score for each subject in the study cohort based at least on (i) the PCs and on (ii) biomarker stratifier weights for a disease of interest; determining which of a plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest, where a disease-associated variant is a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype; determining statistical interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores, wherein the statistical interactions are predictive of drug target-specific differential treatment response for the disease of interest; and performing one or more prediction-based actions based on the determined interactions.

[0035] In some embodiments, the one or more prediction-based actions comprises at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics.

[0036] In some embodiments, the genetic ancestry data indicates whether subjects in the study cohort are a case or a control for the disease of interest.Attorney Docket No.136622-1010

[0037] In some embodiments, the genetic ancestry data is based on genotyping arrays or whole-genome sequencing.

[0038] In some embodiments, the genetic ancestry data is obtained from a publicly-available database.

[0039] In some embodiments, the plurality of PCs comprises at least 5 PCs.

[0040] In some embodiments, the study cohort comprises at least 200 control subjects.

[0041] In some embodiments, each disease-associated variant in the plurality of disease- associated variants satisfies a genome-wide significance threshold.

[0042] In some embodiments, the plurality of disease-associated variants is determined based on a first disease genome-wide association study (GWAS).

[0043] In some embodiments, the biomarker stratifier score for each subject is a scaled biomarker stratifier score.

[0044] In some embodiments, the scaled biomarker stratifier score is based at least on a raw PGS and an ancestry-normalized PGS.

[0045] In some embodiments, the PGS variant weights are computed independent of the genetic ancestry data corresponding to the study cohort.

[0046] In some embodiments, the PGS variant weights are computed using a method for ancestry normalization that estimates the joint likelihood of the mean and variance of the PRS scores as a function of ancestry PCs, and calculates an adjustment score for every individual in the dataset.

[0047] In some embodiments, determining which of the plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest comprises generating an allelic score.

[0048] Another aspect of the disclosure is directed to an in silico system for drug development, such system comprising: at least one hardware processor; and a non-transitory computer-readable storage medium having stored thereon program code, the program code executable by the at least one hardware processor to: obtain molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects;Attorney Docket No.136622-1010 determine the value of the biomarker stratifier effects for each phenotype based on external data; construct a matrix based on the biomarker stratifiers, phenotypes and values; run a principal components analysis biomarker stratifier x phenotype-related phenotype biomarker stratifier effects matrix; calculate a biomarker stratifier score for each principal component; calculate a pharmacomimetic genetic score for each drug target; identify predicted drug effects on one or more phenotypes in subsets of the biomarker stratifier distribution based on the statistical interaction of the polygenic score calculated for each principal component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0049] These and other embodiments, features, and advantages will be set forth in the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate one or more embodiments and, together with the description, explain these embodiments. The accompanying drawings have not necessarily been drawn to scale. Any values or dimensions illustrated in the accompanying graphs and figures are for illustration purposes only and may or may not represent actual or preferred values or dimensions. Where applicable, some or all features may not be illustrated to assist in the description of underlying features.

[0051] Figure 1 is a flow diagram showing the method of one embodiment of the disclosure.

[0052] Figure 2 illustrates a block diagram of a system for in silico drug development and determination of drug activity in one embodiment of the disclosure.

[0053] Figure 3 illustrates a flow diagram showing a method for in silico drug development and determination of drug activity in one embodiment of the disclosure.

[0054] Figure 4 illustrates a flow diagram showing a method for in silico drug development and determination of drug activity in one embodiment of the disclosure.

[0055] Figure 5 illustrates a flow diagram showing a method for in silico drug development and determination of drug activity in one embodiment of the disclosure.Attorney Docket No.136622-1010

[0056] Figure 6 illustrates a sequence for in silico drug development and determination of drug activity in one embodiment of the disclosure.

[0057] Figure 7 graphically depicts the number of patients (individuals, y-axis) vs. the CFB allelic scaled score that falls into three buckets. The CFB scaled score falls broadly into 3 bins. Low is less than or equal to 0. Medium is less than or equal to 1, but greater than 0. High is greater than 1.

[0058] Figure 8 is a histogram showing the number of patients (individuals, y-axis), vs. the complement score (excluding CFB).

[0059] Figure 9 is a histogram showing the number of patients (individuals, y-axis), vs. the complement score (full). The distribution is similar to that shown in FIG.8.

[0060] Figure 10 graphically depicts density (y-axis) vs. complement score (excluding CFB). The impact on the score for dry AMD is shown.

[0061] Figure 11 shows the mean ISOS RPE thickness central subfield (in micrometer, y axis) vs. complement genetic score. CFB LoF is depicted as low (red), medium (green) and high (blue). For the RPE ISOS Central Subfield Thickness, a clear interaction between the Complement Genetic Score and the CFB LoF Burden is shown. High CFB LoF burden essentially rescues the reduction in RPE thickness observed in individuals with a high complement factor score.

[0062] Figure 12 shows distribution of ISOS RPE thickness by age, fitted with a smooth line. The non-linear reduction in ISOS RPE thickness, beginning around age 60 is shown.

[0063] Figures 13A-13B graphically depict the (A) comparison of variant allele frequencies in the Fritsche 2016 dataset (disclosed in Fritsche et al. (2016) Nature Genetics Vol.41:688- 695, which is incorporated herein in its entirety) versus the UK Biobank European-ancestry (pan.ukbb.broadinstitute.org, last accessed May 13, 2025) subset. (B) Comparison of GWAS betas for 48 variants reported in Fritsche 2016 versus the betas for the same variants in the present analysis of the Fritsche 2016 data. DETAILED DESCRIPTION

[0064] The following detailed description of preferred embodiments of the disclosure will be better understood when read in conjunction with the appended drawings.

[0065] All publications, including patent documents, scientific articles and databases, referred to in this application are incorporated by reference in their entirety for all purposes toAttorney Docket No.136622-1010 the same extent as if each individual publication were individually incorporated by reference. If a definition set forth herein is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein incorporated by reference, the definition set forth herein prevails over the definition that is incorporated herein by reference.

[0066] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0067] In some aspects, the present disclosure describes in silico systems for recapitulation of empirical clinical trial data. Disclosure is provided herein demonstrating that genetic analysis using pharmacomimetic variants can recapitulate drug-PRS interactions observed in retrospective analyses of clinical trials. It also demonstrates that selectively enrolling patients with high PRS for clinical trials can increase average treatment response by enrolled patients as shown in the comparison of the empirical evidence and the results obtained by the systems of the disclosure. For example, as demonstrated, the average treatment response of a known anti-PCSK9 therapy can increase by a factor of ~2 in patients with high PRS.

[0068] Over the last decade, genome-wide association studies (GWAS) have uncovered the contribution of inherited variants to common complex disorders. Many non-communicable disorders with a major public health impact have a genetic underpinning that is highly polygenic, comprising, or alternatively consisting essentially of, or yet further consisting of hundreds or thousands of genetic variants (or polymorphisms), each having a small effect on disease risk. Each genetic variant associated with a disease is valuable in indicating a gene or pathway of biological relevance to the disorder, but and such genetic data can be used to predict disease risk, with clinical utility.

[0069] For example, there have been substantial successes in discovering and developing new health interventions, including therapeutics, for common chronic human diseases, here defined as diseases that have a prevalence in the general population of >0.1%. Despite these advances in management of common chronic diseases, there remains substantial unmet need in clinical areas that contribute significantly to the population burden of disease morbidity, premature mortality, and associated societal and healthcare costs. It is estimated that 6 in 10 Americans have at least one common chronic disease, and that these diseases collectively are responsible for ~$2.7T in US healthcare costs. Owing to substantial heterogeneity in disease progression and treatment response, drug development in these areas has required largeAttorney Docket No.136622-1010 clinical trials, at substantial clinical development cost, to evaluate clinical efficacy. Thus, despite substantial market opportunity and clinical unmet need, drug development has shifted over the last two decades to oncology and rare disease, where molecularly defined mechanisms and subpopulations have converged to enable a higher likelihood of demonstrating efficacy and safety sufficient for regulatory approval of new chemical entities.

[0070] Large scale genome-wide association studies performed over the last several decades have identified genetic contributions to variation in common disease risk, as well as the underlying risk factors that play a causal role in their pathobiology. These studies have found that the architecture of disease risk is genetically complex, reflecting the contributions of tens to millions of individual common genetic variants, each of small effect, as well as a smaller number of low frequency and rare genetic variants of greater effect. Collectively, these risk factors explain between ~40% and ~80% of the variation in disease risk observable at a population level. With larger studies of the effects of these genetic variants on disease risk has come an ability to more precisely estimate the contributions to disease risk of genetic variants at a range of effect sizes. This has facilitated the scaled summation of these effects into continuous polygenic genetic scores (PGS) that estimate, for any individual, the aggregate of measurable genetic risk. PGS may be isotropic, reflecting the sum total of genetic effects on disease risk agnostic to effects on underlying risk factors or biological pathways, or may reflect fundamental driver pathways in subsets of common chronic disease or in individuals with specific risk factors. Both isotropic as well as pathway- or risk-factor- specific PGS may resolve common disease heterogeneity and identify subgroups of the population who respond exceedingly well to certain therapeutic mechanisms. Indeed, there is clinical trial precedent for patients with high PGS receiving greater benefit from a drug than patients with low PGS.

[0071] A system implementing the systems and methods described herein may perform a method of analysis and use a knowledgebase for identifying individuals with subsets of common chronic human diseases who are predicted to enjoy greater benefit from specific therapeutic mechanisms. Specifically, the method identifies combinations of drug targets and disease indications where it is predicted that patients with a high PGS for the disease, its associated risk factors, or an aggregate of several genetically-driven biological pathways, will receive greater benefit from a drug compared to patients with a low PGS.

[0072] To perform the method, a computer may access genetic and phenotypic data for a large group of people (e.g., the “study cohort”). The genetic data can originate fromAttorney Docket No.136622-1010 genotyping arrays or whole-genome sequencing. The phenotype data does not have to come in a specific format. But, in some embodiments, phenotype data must be sufficient to determine whether each subject is a case, a control, or neither for the disease of interest. The algorithm for assigning case-control status to subjects may be manually defined by a user. The algorithm may consider, for example: (i) ICD-10 or ICD-9 diagnosis codes from hospital records or primary care records; (ii) self-reported diagnoses; (iii) medication prescriptions; and / or (iv) OPCS-4 or OPCS-3 operation codes from hospital records. The computer may also use the age and sex of each subject. Examples of suitable study cohorts can include the UK Biobank(pan.ukbb.broadinstitute.org, last accessed May 13, 2025), the FinnGen Research Project (world wide web finngen.fi / en, last accessed May 13, 2025), and the All of Us Research Program (allofus.nih.gov, last accessed May 13, 2025). There is no exact requirement for the number of subjects in the study cohort, but typically the computer may use at least ~5k cases and ~5k controls for a particular disease of interest.

[0073] The computer may also access PGS variant weights for the disease of interest. “PGS variant weights” can refer to a table of genetic variants in which each genetic variant is assigned a weight (e.g., a numerical value). The weight can be an estimate of how much a given variant contributes to the risk of a disease. The computer can compute PGS variant weights using publicly available programs such as PRS-CS (Ge et al. (2019) Polygenic prediction via Bayesian regression and continuous shrinkage priors. Nat Commun 10:1776) or LDPred2 (Prive et al. (2020) LDpred2: better, faster, stronger, Bioinformatics, Volume 36, Issue 22-23: 5424–5431)). In some cases, PGS variant weights can be published by academic researchers. The computer can download such weights from a web resource called the PGS Catalog, for example. Computing the PGS variant weights may not involve any data from the study cohort and must be derived from a completely independent dataset, in some embodiments.

[0074] The computer may access results from a genome-wide association study (GWAS) for the disease of interest, commonly referred to as “summary statistics.” To do so, the computer may retrieve GWAS summary statistics from a web resource, such as the EBI GWAS Catalog, for example. Alternatively, the data processing system may perform the GWAS within the study cohort, such as by using publicly available software such as Plink (Purcell et al. (2007) Am J Hum Genet. Sep;81(3):559-75), Chang et al. (2015) Gigascience Feb 25;4:7), SAIGE (Zhou et al. (2022)) Nat Genet 54:1466–1469, or REGENIE (Mbatchou et al. (2021) Nat Genet 53:1097–1103).Attorney Docket No.136622-1010

[0075] The computer can access a computational pipeline for mapping GWAS variants to causal genes, such as the pipelines published by Mountjoy et al. (2021) Nat Genet. Nov;53(11):1527-1533. The drug-by-PGS interaction discovery method is agnostic to the particular pipeline that is used.

[0076] The method can apply the following operations of a framework: (1) modeling the predicted effects of drug target modulation from genotype-phenotype association analysis of human genetic variants in drug target-encoding genes and observed clinical phenotypes. The genetic variants used for modeling may be individual variants or sets of statistically- independent variants in the same drug target gene identified as “allelic scores.” These statistical instruments are “pharmacomimetic instruments;” (2) partitioning a human population according to isotropic PGS, risk factor PGS, or biological pathway-PGS; and (3) application of a method for identifying statistical interactions between pharmacomimetic instruments for drug targets and PGS that predict drug target-specific differential treatment response. In some embodiments, the method can include using the predicted drug target- specific differential treatment response to select patients for treatment and / or applying the treatment. In some embodiments, the method can include transmitting the predicted drug target-specific differential treatment response and / or any data used to determine the predicted drug target-specific differential treatment response (e.g., the statistical interactions, the determined PGS, etc.) to another computing system (e.g., into a downstream pipeline) for an entity associated with the computing system to use for treatment.

[0077] In performing the aforementioned method, a computing system can use human genetic data to predict a-priori whether a given drug mechanism will exhibit PGS-stratified efficacy without needing to have clinical trial data.

[0078] One attempt for modelling and simulation involves implementing machine learning techniques, such as by using a linear regression model or a neural network to generate the simulations. However, such machine learning models have inherent technical problems that limit their capabilities when processing large amounts of data. For instance, neural networks are limited by the number of nodes that are included at the input layer and linear regression models are limited by the number of axes they may have. These issues with machine learning models may cause problems with biomedical simulations, because biomedical simulations can involve millions of datapoints of varying types. The large amount of data points can make it difficult to use machine learning models for the simulations or otherwise cause such simulations to be impractical or impossible given the large amount of computingAttorney Docket No.136622-1010 resources they would require (e.g., a neural network with a large amount of input nodes or a linear regression model with a large number of axes can be computationally expensive to execute, particularly for a large number of datapoints).

[0079] A computer implementing the systems and methods described herein can overcome these technical deficiencies of machine learning models. For instance, the computer can first use a principal component analysis (PCA) technique on genetic ancestry data to generate principal components (PCs) for the genetic ancestry data. In doing so, the computer can reduce a size of the genetic ancestry data to reduce the number of inputs for subsequent processing using machine learning techniques. After generating the PCs and / or generating a matrix from the PCs, the computer can generate a feature vector with the PCs, matrices, and / or biomarker stratifier weights for a particular disease of interest. The computer can use a linear regression model or a neural network on the feature vector to generate biomarker stratifier scores for individuals of the genetic ancestry data. The computer can determine pharmacomimetic instruments for the disease of interest, where a disease-associated variant is a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype. The computer can subsequently use a neural network or linear regression model to determine interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores. The statistical interactions can be predictive of drug target- specific differential treatment response for the disease of interest. The computer can perform one or more prediction-based actions based on the determined interactions. By using PCA in this way, the computer can format a feature vector from genetic ancestry data to substantially reduce the number of axes and / or nodes of a linear regression model and / or neural network to enable the linear regression model and / or neural network to operate and / or reduce the amount of computational resources that are required for the modeling or simulations.

[0080] Figure 1 is a flow diagram showing a method 200 of one embodiment of the disclosure. The method 200 can be performed by a data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG.2, a server system, etc.). The method 200 may include more or fewer operations and the operations may be performed in any order. Performance of the method 200 may enable the data processing system to automatically identify predicted drug effect on phenotypes based on polygenic risk scores.Attorney Docket No.136622-1010

[0081] In the method 200, at operation 202, the data processing system obtains genetic variant data from a plurality of subjects with identified phenotypes. At operation 204, the data processing system determines the value of the variant effects for each phenotype based on external data. At operation 206, the data processing system constructs a matrix based on the variants, phenotypes, and values of the variant effects. At operation 208, the data processing system runs a principal components analysis biomarker variant x phenotype- related phenotype variant effects matrix. At operation 210, the data processing system calculates a polygenic risk score for each principal component. At operation 212 the data processing system identifies predicted drug effects on one or more phenotypes based on the polygenic risk scores calculated for each principal component.

[0082] Figure 2 illustrates a block diagram of a system 300 for in silico drug development and determination of drug activity in one embodiment of the disclosure. In brief overview, the system 300 can include a data processing system 302 and data sources 304, 306, and 308. The data processing system 302 can receive or retrieve molecular biomarker stratifier data of a plurality of subjects. The data processing system 302 can determine values representing biomarker stratifier effects based on external data of individuals separate from the plurality of subjects. The data processing system 302 can calculate a biomarker stratifier score, a chosen disease phenotype and a pharmacomimetic genetic score for different drug targets. The data processing system 302 can identify a predicted drug activity for each drug target based on the biomarker stratifier score and the pharmacomimetic genetic score for the drug target in association analysis with the disease phenotype. The system 300 may include more, fewer, or different components than shown in FIG.2.

[0083] The data processing system 302 may comprise one or more processors that are configured to identify predicted drug activity for drug targets for disease phenotypes. The data processing system 302 may comprise a network interface 312, a processor 314, and / or memory 316. The data processing system 302 may communicate with the data sources 304, 306, and / or 308 via the network interface 312, which may be or include an antenna or other network device that enables communication across a network and / or with other devices. The processor 314 may be or include an ASIC, one or more FPGAs, a DSP, circuits containing one or more processing components, circuitry for supporting a microprocessor, a group of processing components, or other suitable electronic processing components. In some embodiments, the processor 314 may execute computer code or modules (e.g., executable code, object code, source code, script code, machine code, etc.) stored in memory 316 toAttorney Docket No.136622-1010 facilitate the activities described herein. The memory 316 may be any volatile or non-volatile computer-readable storage medium capable of storing data or computer code.

[0084] The memory 316 may include a data collector 318, a principal component (PC) generator 320, a pharmacomimetic instrument identifier, a statistical interaction calculator 324, and / or an action performer 326, in some embodiments. The components 318-326 may operate to identify predicted drug activity of drug targets for different disease phenotypes.

[0085] For example, the data collector 318 may comprise programmable instructions that, upon execution, cause the processor 314 to communicate with the data sources 304, 306, 308, and / or any other data sources. The data collector 318 may be or include an application programming interface (API) that facilitates communication between the data processing system 302 and other computing devices. The communicator 318 may communicate with the data sources 304, 306, 308, and / or any other computing device across the network 310.

[0086] The data sources 304, 306, and 308 can be data sources that store molecular biomarker stratifier data and / or external data. For example, the data sources 304, 306, and / or 308 can be or include a database (e.g., a relational or graph database) that stores molecular biomarker stratifier data that includes biomarker stratifiers and at least one disease phenotype from or of a plurality of subjects or individuals. The data sources 304, 306 and / or 308 can additionally or instead include external data of individuals separate from the plurality of subjects or individuals of molecular biomarker stratifier data.

[0087] The data collector 318 can establish connections with one or more of the data sources 304, 306, and / or 308. The data collector 318 can establish the connections over the network 310. To do so, the data collector 318 can communicate with the data sources 304, 306, and / or 308 across the network 310. In one example, the data collector 318 can transmit syn packets to the respective data sources 304, 306, and / or 308 and establish the connections using a TLS handshaking protocol. The data collector 318 can use any handshaking protocol to establish a connection with the data sources 304, 306, and / or 308.

[0088] The data collector 318 can obtain the molecular biomarker stratifier data and the external data. In some embodiments, the data collector 318 can obtain the molecular biomarker stratifier data and the external data over the established connections from one or more of the data sources 304, 306, or 308. The data collector 318 may do so by querying the data sources 304, 306, or 308. In some embodiments, the data collector 318 can receive the molecular biomarker stratifier data and / or the external data as input (e.g., a manual input) byAttorney Docket No.136622-1010 a user accessing the data processing system. The data collector 318 can obtain portions of the molecular biomarker stratifier data and the external data using any method or any combination of methods.

[0089] The PC generator 320 may comprise programmable instructions that, upon execution, cause the processor 314 to generate principal components. The PC generator 320 can generate principal components that correspond to (or based on) genetic ancestry data (e.g., different genes) for subjects in a study cohort. The genetic ancestry data can be based on genotyping arrays or whole-genome sequencing. The PC generator 320 can obtain the genetic ancestry data from a publicly available database, as described herein. The PC generator 320 can be FlashPCA, for example. The principal components can be used to control subsequent statistical analyses for population stratification, in some embodiments.

[0090] For example, the PC generator 320 can generate a plurality of PCs corresponding to genetic ancestry data for subjects in a study cohort. The PC generator 320 can do so by generating a genetic similarity matrix from the genetic ancestry data of the subjects in the study cohort and using principal component analysis on the genetic similarity matrix. The PC generator 320 can generate any number of PCs using principal components analysis.

[0091] The PC generator 320 can generate a biomarker stratifier score for each subject in the study cohort. The PC generator 320 can generate the biomarker stratifier scores based at least on (i) the PCs and on (ii) biomarker stratifier weights for a disease of interest. For example, a “raw PGS” (e.g., a biomarker stratifier score) for a subject S can be defined as the sum over each variant in a PGS variant weights table of ((the variant’s weight) * (the # of copies S has of that variant)). The PGS variant weights of the PGS variant weights table can be computed independent of the genetic ancestry data corresponding to the study cohort. The PC generator 320 can generate biomarker stratifier scores for any number of diseases, such as AMD or CNV.

[0092] In some embodiments, the PC generator 320 can use machine learning techniques (e.g., a neural network or a linear regression model) to generate the biomarker stratifier scores. For instance, the PC generator 320 can generate a feature vector from the PCs for the individuals of the study cohort and execute a neural network or a linear regression model using the feature vector as input to generate the biomarker stratifier scores. In doing so, the PC generator 320 can reduce the size of the dataset that is used to generate biomarker stratifiers scores for individuals. The reduced size can enable the neural network or linearAttorney Docket No.136622-1010 regression model to operate or reduce the computational resources that are required to do so compared with systems that may input genetic ancestry data directly into a neural network or a linear regression model.

[0093] The PC generator 320 can determine “ancestry-normalized” PGS. The PC generator 320 can determine the ancestry-normalized PGS to be the residuals of a linear regression model or another type of machine learning model (e.g., a neural network or a support vector machine) where the outcome can be the raw PGS and the predictors are the PCs of genetic ancestry data. This normalization can correct for differences in mean and / or variance of PGS values between populations. The ancestry-normalized PGS can also be called an ancestry- normalized biomarker stratifier score.

[0094] The PC generator 320 can determine “Scaled PGS.” The PC generator 320 can do so using the following equation: scaled PGS = ((the ancestry-normalized PGS - the ancestry- normalized PGS across all subjects in the cohort) / (the standard deviation of the ancestry- normalized PGS across all subjects in the cohort)). Scaled PGS values can thus be approximately normally distributed. The scaled PGS can be used for all subsequent analyses. Scaled PGSs are also called scaled biomarker stratifier scores.

[0095] The pharmacomimetic instrument identifier 322 may comprise programmable instructions that, upon execution, cause the processor 314 to determine which of a plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest. A disease-associated variant can be a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype.

[0096] To determine which of the plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest, the pharmacomimetic instrument identifier 322 may execute a program (e.g., GCTA COJO-SLCT) for the disease of interest to identify “conditionally independent disease-associated variants.” To do so, for example, the pharmacomimetic instrument identifier 322 may first run a GWAS (e.g., a first GWAS) as normal. The pharmacomimetic instrument identifier 322 can identify the variant that has the most significant association with the disease outcome (e.g., the smallest p-value) across the whole genome. The pharmacomimetic instrument identifier 322 can rerun the GWAS, this time conditioning all of the variant-disease association tests on the genotype of the most-Attorney Docket No.136622-1010 significant variant. The pharmacomimetic instrument identifier 322 can identify the variant that has the next-most significant association with the disease outcome. The pharmacomimetic instrument identifier 322 can rerun the GWAS, this time conditioning all of the variant-disease association tests on the genotype of the most-significant and the 2nd- most-significant variants. The pharmacomimetic instrument identifier 322 can iteratively repeat this procedure until determining no variant is associated with the disease below a defined p-value threshold (e.g., a genome-wide significance threshold, such as 5 x 10-8). Using lower thresholds can reduce the chances that an analytical result is a fluke but increases the chance that the data processing will not identify an important variant, and vice versa. The pharmacomimetic instrument identifier 322 can use any threshold. The set of variants identified by the pharmacomimetic instrument identifier can be called “conditionally-independent disease-associated variants” because each variant is associated with the disease even after conditioning on the genotypes of all of the previous variants.

[0097] The pharmacomimetic instrument identifier 322 can create a table. Each row of this table can correspond to one subject from the study cohort and include data regarding the subject as demographic data and other determined data (e.g., PGS, PCS of genetic ancestry, dosage for the alternate allege of each of the independent disease associated variants, etc.).

[0098] The pharmacomimetic instrument identifier 322 can identify which of the conditionally-independent disease-associated variants are “pharmacomimetic instruments” (e.g., variants that modulate the function or expression of one or more drug target genes so that the variants’ effect on human phenotypes is likely to be predictive of the drug’s effect on human phenotypes). Depending on the configuration, the pharmacomimetic instrument identifier 322 can use different criteria and / or priorities to identify pharmacomimetic instruments from the identified conditionally-independent disease-associated variants. In one example, a user may be interested in clinical-stage drugs for the disease. In this case, the pharmacomimetic instrument identifier 322 can compile a table of clinical-stage drugs and their targets using public resources, such as clinicaltrials.gov and / or proprietary databases, such as Cortellis. The pharmacomimetic instrument identifier 322 can determine that variants that are close to a drug target gene (e.g., within 150 kb of the gene’s transcription start site) or are mapped to the drug target are pharmacomimetic instruments. In another example, a user may be interested in known and novel drug targets for the antibody modality. In this case, the pharmacomimetic instrument identifier 322 can compile a list of genes that encode proteins that are druggable with the antibody modality (e.g., proteins that are secretedAttorney Docket No.136622-1010 or localized to the cell surface). The pharmacomimetic instrument identifier 322 can determine variants that are mapped to an antibody-druggable protein are pharmacomimetic instruments based on the compiled table (e.g., based on the determined variants having stored associations with the antibody-druggable protein).

[0099] In some embodiments, there is more than one suitable variant for a drug mechanism. In this case, it will increase statistical power to detect drug-PGS interactions by aggregating variants (e.g., all of the variants) identified as pharmacomimetic instruments that were associated with a given drug mechanism into a combined allelic score. To do so, for example, the pharmacomimetic instrument identifier 322 can compute an allelic score in the same or a similar manner to a PGS (e.g., assign each variant a weight and compute, for each subject in the cohort, the sum over each variant of (the variant’s weight) * (the subject’s dosage for the variant)). The difference between an allelic score and a PGS is that the allelic score is composed of variants that modulate the function or expression of a single drug target gene, while a PGS can be composed of variants across the genome that affect many different genes. The variants in the allelic score may be weighted using a GWAS (e.g., a second GWAS) for the disease that did not include any subjects from the study cohort. Alternatively, the variants can be weighted using a GWAS (e.g., a third GWAS) for a biomarker that causally mediates the effect of the drug target on the disease. A biomarker GWAS still must not overlap the study cohort.

[0100] The statistical interaction calculator 324 may comprise programmable instructions that, upon execution, cause the processor 314 to determine statistical interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores. The statistical interactions can be predictive of drug target-specific differential treatment responses for the disease of interest. To determine the statistical interactions, the statistical interaction calculator 324 can generate a data table with phenotypes, covariates, and a set of pharmacomimetic instruments determined using the aforementioned systems and methods, which may include single variants and multi-variant allelic scores. Responsive to doing so, the statistical interaction calculator 324 can test for interactions between the pharmacomimetic instruments and the PGS in association analysis with the disease or risk factor trait of interest.

[0101] For example, for each pharmacomimetic instrument, the statistical interaction calculator 324 can use statistical software, such as the R programming language, to fit a logistic regression model or another type of machine learning model (e.g., a neural networkAttorney Docket No.136622-1010 or a support vector machine) using the following example formula: disease case-control status ~ age + sex + ancestry PCs + (pharmacomimetic instrument) + (scaled PGS) + (pharmacomimetic instrument) : (scaled PGS). In doing so, the statistical interaction calculator 324 can determine the scaled PGS. The model can be implemented using the R programming language, such as by using the command glm(model_formula, family = “binomial”, data = our_genetic_and_phenotypic_data). In the example formula, the variable to the left of thesymbol is the dependent variable (e.g., the variable that is being predicted). Each of the variables to the right of thesymbol, separated by “+” symbols, are independent variables, (e.g., variables that are observed in the data and used to predict the dependent variable). The “:” symbol indicates a variable that is the product (result of multiplication) of the variables to the left and right of the “:” symbol. This product is an “interaction term.” If the pharmacomimetic instrument : scaled polygenic score (PGS) interaction term -- one of the independent variables -- has a statistically-significant association with the dependent variable (in this case, disease status), controlling for all of the other independent variables -- such as age and sex -- then the statistical interaction calculator 324 can predict or determine that the drug that is modeled by the pharmacomimetic instrument will have different effects in subjects with high PGS vs. low PGS. The statistical interaction calculator 324 can fit a logistic regression model for each pharmacomimetic instrument that the pharmacomimetic instrument identifier 322 determines or identifies.

[0102] The statistical interaction calculator 324 can use the same program to compute a p- value for the interaction term. If the p-value for the interaction between a pharmacomimetic instrument and the PGS is < 0.05 / (the # of instruments tested), the statistical interaction calculator 324 can determine the interaction to be “Bonferroni significant”. Alternatively, the statistical interaction calculator 324 can be configured to use a more lenient p-value threshold, at the risk of false positives.

[0103] The instrument-PGS interactions can be “drug-PGS interactions” because each pharmacomimetic instrument is a model for a particular drug mechanism.

[0104] The statistical interaction calculator 324 can generate tables that present the statistically significant drug-PGS interactions in a way that is easier to interpret than looking at the raw regression model outputs. For example, first, the statistical interaction calculator 324 can group the subjects of the study cohort by quantiles of the PGS, (e.g., the 0-33rd percentile, the 34th-66th percentile, and the 67th-100th percentile). Then, for each drug-PGS interaction, the statistical interaction calculator 324 can fit separate logistic regression modelsAttorney Docket No.136622-1010 within each PGS quantile using the formula: disease case-control status ~ age + sex + ancestry PCs + the pharmacomimetic instrument for the drug. Using these regression outputs, the statistical interaction calculator 324 can construct a table that shows the effect of the pharmacomimetic instrument on disease risk (as well as a 95% confidence interval for that effect) within each PGS quantile.

[0105] For instance, the statistical interaction calculator 324 can divide the subjects of the study into a finite number of groups (e.g., the groups) based on their PGS values. Within each group, the statistical interaction calculator 324 can predict a disease outcome based on the patients’ genetics, specifically a genetic instrument that mimics the effects of a drug (the “pharmacomimetic instrument”). In doing so, the statistical interaction calculator 324 can determine a difference in how strongly the pharmacomimetic instrument predicts the disease outcome in individuals with high PGS vs. middle PGS vs. low PGS. The difference may indicate that a drug might have different efficacy in people with high PGS vs. middle PGS vs. low PGS.

[0106] The statistical interaction calculator 324 can also compute a statistic called the “treatment effect multiplier” or “TEM”. This is the effect of the pharmacomimetic instrument on disease risk in the top quartile divided by the effect in all subjects. The “treatment effect multiplier” can, e.g., represent how much one increases the average treatment effect in a clinical trial of the drug if one enrolled only patients in the top quantile of the PGS as opposed to enrolling all qualified patients. The TEM is a way to compare drug-PGS interactions to distinguish “strong” from “weak” interactions.

[0107] The action performer 326 may comprise programmable instructions that, upon execution, cause the processor 314 to perform one or more prediction-based actions based on the determined interactions. The one or more prediction-based actions can include at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics. To perform a prediction-based action, the action performer 326 can select a patient for treatment. The action performer 326 can select the patient based on a likelihood that the patient will benefit from the treatment. For example, the action performer 326 can select the patient by determining a PGS for the patient using the systems and methods described herein and determining the PGS for the patient is above a threshold (e.g., a PGS threshold) or within a range (e.g., a quartile, a quintile, a defined range, etc.). In another example, the action performer 326 can select patients for novel therapeutic mechanisms whose benefit is mostly or only apparent in subsets of common disease patients who haveAttorney Docket No.136622-1010 elevated polygenic risk. This enables novel target identification and novel therapeutic discovery. In another example, the action performer 326 can select patients from large common disease treatment-eligible populations who will benefit most from existing drug therapies, and, conversely, identify patients who are not likely to receive clinically meaningful benefit. This can enhance the pharmaco-economic profile of existing therapeutic mechanisms, yielding more cost-effective utilization. In another example, the action performer 326 can select mechanisms and design of prospective clinical trials of investigational drugs that are predicted to yield greater benefit in individuals with elevated PGS. These trials may be run with many-fold fewer patient years required to demonstrate clinical benefit, given the expectation of magnified event rate and treatment response rates.

[0108] In some embodiments, treatment can be performed on the selected patients. For example, the action performer 326 can select patients with PGS scores that satisfy a criteria for a particular treatment to treat a disease. The action performed 326 can generate a record that includes a list of selected patients. An entity (e.g., a user or a clinician) may view the list and apply a treatment for the disease to the patients identified in the record. Examples of such treatment can include injections, therapy, or other treatment techniques. In some embodiments, the action performer 326 can connect with another computer and transmit the record containing the list of patients or the PGS scores determined by the PC Generator 320 to another computer. The computer may display the scores or list of patients and the entity or user accessing the computer may use the list and / or scores for treatment.

[0109] Figure 3 illustrates a flow diagram showing a method 400 for in silico drug development and determination of drug activity in one embodiment of the disclosure. The method 400 can be performed by a data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG.2, a server system, etc.). The method 400 may include more or fewer operations and the operations may be performed in any order. Performance of the method 400 may enable the data processing system to automatically identify predicted drug activity of different drug targets for disease phenotypes and / or phenotype intermediates.

[0110] In the method 400, at operation 402, the data processing system obtains molecular biomarker stratifier data. The molecular biomarker stratifier data can include a plurality of biomarker stratifiers and at least one disease phenotype or phenotype intermediate from each of a plurality of subjects. The molecular stratified biomarker data can include at least one of the following types of data: genomic, transcriptomic, metabolomic, or proteomic. TheseAttorney Docket No.136622-1010 data are obtained by assays specific to each biomarker type.

[0111] The molecular biomarker stratifier data can include data on the at least one disease phenotype. The data can include a patient’s disease status (e.g., has been diagnosed with AMD yes / no) at a given point in time. The data can also include quantitative disease traits (e.g., ISOS RPE thickness, AMD lesion size), medication information (e.g., patient has taken or is taking a statin or other lipid-lowering therapy at a given time with regard to disease outcomes), and demographic information (e.g., age at diagnosis, biological sex, etc.).

[0112] At operation 404, the data processing system determines a plurality of values (e.g., numeric values) representing biomarker stratifier effects. Each value can separately represent how each of the plurality of biomarker stratifiers affects each of the at least one disease phenotype or phenotype intermediate. The biomarker stratifier effects can respectively be the estimated magnitude of effect of the biomarker on disease risk or continuous disease trait from a statistical model. The data processing system can determine the plurality of values based on external data (e.g., data regarding individuals separate from the subjects of the plurality of subjects). The data processing system can determine the plurality of values using a logistic regression model for disease (yes / no) vs. biomarker (e.g., protein measurement) + covariates (age, sex, etc.) on the data.

[0113] At operation 406, the data processing system calculates a biomarker stratifier score for a chosen disease phenotype or phenotype intermediate. The data processing system can calculate a biomarker stratifier score for each of the at least one disease phenotype or phenotype intermediate. The data processing system can determine biomarker stratifier scores as polygenic risk scores for individual disease phenotypes. The data processing system can calculate biomarker stratifier scores for biomarker stratifiers. In doing so, for example, the data processing system can determine a biomarker stratifier score as one of the following: a polygenic risk score for the disease phenotypes; a polygenic score for a disease risk factor; a polygenic score for biological pathway activity; a polygenic score for drug target expression; a proteomics risk score for the disease phenotypes; a proteomics score for a disease risk factor; a proteomics score for biological pathway activity; a proteomics score for drug target expression; a transcriptomics risk score for the disease phenotypes; a transcriptomics score for a disease risk factor; a transcriptomics score for a biological pathway activity; a somatic mutation score for drug target expression; a somatic mutation score for the disease phenotypes; a somatic mutation score for a disease risk factor; a somatic mutation score for a biological pathway activity; or a somatic mutation score for drug targetAttorney Docket No.136622-1010 expression. For instance, if the stratifier is a polygenic risk score (PRS), a statistical model may be trained using summary statistics from a published study on large-scale cohorts or from a meta-analysis of multiple studies based on large and diverse set of cohorts (none of which include UK Biobank). The data processing system can apply the estimated weights from that model to a dataset that is orthogonal to the ones used for training (e.g., UK Biobank).

[0114] At operation 408, the data processing system calculates a pharmacomimetic genetic score for each drug target for which the data processing system is determining a drug activity. The pharmacomimetic genetic score can represent loci for which more than one pharmacomimetic variant (e.g., a functional variant in a locus that is strongly and significantly associated with a given disease) is identified. The data processing system can determine the pharmacomimetic genetic score by weighing each variant by its estimated effect on the phenotype and summing the weighted values.

[0115] At operation 410, the data processing system identifies the predicted drug activity of each drug target for the at least one disease phenotype or phenotype intermediate in subsets of a biomarker stratifier distribution. The data processing system can identify the predicted drug activity based on a statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the disease phenotype or phenotype intermediate. For example, the data processing system can use a model (e.g., a linear regression model, a logistic regression model, a Cox Proportional Hazards model, etc.) that includes a predictor defined by the multiplicative interaction between the pharmacomimetic genetic score and biomarker score to perform the association analysis for each drug target and disease outcome phenotype. The predictor can be the predicted drug activity.

[0116] The data processing system can identify the subsets of the biomarker stratifier distribution using one or more thresholds. The one or more thresholds can be chosen to define different subsets depending on the use case. In one example, it may be optimal to use a 25th percentile of the AMD or choroidal neovascularization polygenic risk score to define a subset of patients who are predicted to have an outsized clinical benefit from therapy X. In another example, it may be optimal to use a 33% (top tertile) of the AMD or choroidal neovascularization PRS to define a subset of patients who would benefit most from therapy Y. In other disease settings with different PRS, the thresholds may also vary. Defining the optimal threshold will depend on factors such as: estimated effect size for the subset vs. all-Attorney Docket No.136622-1010 comers on therapy, failure screening rate (in a trial, if the threshold is set too conservatively, e.g., 5%, a lot more patients will have to be screened to enroll a few), etc. Subsets can be from individual-level cohort data as the top 10%, 20%, 30%, etc. of the distribution. Thus, the exact values of “high” and “low” PGS will vary with these factors.

[0117] Figure 4 illustrates a flow diagram showing a method 500 for in silico drug development and determination of drug activity in one embodiment of the disclosure. The method 500 can be performed by a data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG.2, a server system, etc.). The method 500 may include more or fewer operations and the operations may be performed in any order. Performance of the method 500 may enable the data processing system to automatically identify predicted drug activity of different drug targets for disease phenotypes or phenotype intermediates.

[0118] In the method 500, at operation 502, the data processing system obtains molecular biomarker stratifier data and at least one disease phenotype or phenotype intermediate. The molecular biomarker stratifier data can include a plurality of biomarker stratifiers. The data processing system can obtain the molecular biomarker stratifier data from each of a plurality of subjects.

[0119] At operation 504, the data processing system determines the value of the biomarker stratifier effects for each disease phenotype or phenotype intermediate. Each value can separately represent how each of the plurality of biomarker stratifiers affects each of the at least one disease phenotype or phenotype intermediate. The data processing system can determine the plurality of values based on external data (e.g., data regarding individuals separate from the subjects of the plurality of subjects). At operation 506, the data processing system constructs a matrix based on the biomarker stratifier data (e.g., biomarker stratifiers), disease phenotypes or phenotype intermediate values. At operation 508, the data processing system runs a principal components analysis biomarker stratifier phenotype-related x phenotype biomarker stratifier effects matrix. At operation 510, the data processing system calculates a biomarker stratifier score for each principal component. At operation 512, the data processing system calculates a pharmacomimetic genetic score for each drug target. At operation 514, the data processing system identifies predicted drug activity of each drug target on one or more phenotypes or phenotype intermediates in subsets of the biomarker stratifier distribution. The data processing system can identify the predicted drug activity based on a statistical interaction of the polygenic score calculated for each principalAttorney Docket No.136622-1010 component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype or outcome phenotype intermediate.

[0120] Figure 5 illustrates a flow diagram showing a method 600 for in silico drug development and determination of drug activity in one embodiment of the disclosure. The method 600 can be performed by a data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG.2, a server system, etc.). The method 600 may include more or fewer operations and the operations may be performed in any order. Performance of the method 600 may enable the data processing system to automatically determine interactions between pharmacomimetic instruments and perform prediction-based actions based on the determined interactions. One or more of the operations in the methods 400, 500, and / or 600 can be performed during operations in any other of the methods 400, 500, and / or 600.

[0121] In the method 600, at operation 602, the data processing system generates a plurality of principal components (PCs) corresponding to genetic ancestry data (e.g., genetic data) for subjects in a study cohort. The data processing system can generate any number of PCs. For example, the data processing system can generate five PCs, at least six PCs for a European ancestry cohort, and / or 10-to-40 PCs for a multi-ancestry cohort (e.g., 20 for UK Biobank and / or up to 40 for a more diverse cohort). The study cohort can include at least 200 control subjects. In some embodiments, the study cohort can include at least 950 control subjects or at least 990 control subjects. The data processing system can determine whether subjects in the study cohort are a case or a control for a disease of interest, such as based on flags for the subjects that indicate whether the subjects are cases or controls. The genetic ancestry data can be based on genotyping arrays or whole-genome sequencing. The data processing system can obtain the genetic ancestry data from a publicly available database, for example, or through any other method, e.g., direct sequencing of samples of a chosen patient population.

[0122] At operation 604, the data processing system generates a biomarker stratifier score for each subject in the study cohort. The data processing system can generate the biomarker stratifier scores based at least on (i) the PCs and on (ii) biomarker stratifier weights for a disease of interest. The biomarker stratifier score can be for a particular disease of interest, e.g., AMD. For example, a “raw PGS” (e.g., a biomarker stratifier score) for a subject S can be defined as the sum over each variant in a PGS variant weights table of ((the variant’s weight) * (the # of copies S has of that variant)). For example, consider three variants A, B, and C, with weights 1, 2, and 3 respectively. If subject S has 2, 0, and 1 copies of variants A,Attorney Docket No.136622-1010 B, C, then their unscaled PGS is computed as 2*weight(A) + 0*weight(B) + 1*weight(C) = 2*1 + 0*2 + 1*3 = 5. The PGS variant weights can be computed independent of the genetic ancestry data corresponding to the study cohort.

[0123] The data processing system can determine “ancestry-normalized PGSs.” Ancestry- normalized PGS can be defined as the residuals of a linear regression model where the outcome can be the raw PGS and the predictors are the PCs of genetic ancestry that were selected in operation 602. This normalization can correct for differences in mean and / or variance of PGS values between populations.

[0124] The data processing system can determine “scaled PGSs.” Scaled PGS can be defined as ((the ancestry-normalized PGS - the ancestry-normalized PGS across all subjects in the cohort) / (the standard deviation of the ancestry-normalized PGS across all subjects in the cohort)). Scaled PGS values can thus be approximately normally distributed. The scaled PGS can be used for all subsequent analyses. The task of creating and scaling PGS that is effective and comparable across populations may be performed using any method.

[0125] At operation 606, the data processing system determines which of a plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest. A disease-associated variant can be a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype.

[0126] To determine which of the plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest, the data processing system may execute a program (e.g., GCTA COJO-SLCT) for the disease of interest to identify “conditionally independent disease-associated variants.”

[0127] For example, groups of genetic variants that are close together on a DNA molecule (e.g., roughly within 500,000 base pairs, although this distance may vary depending upon where in the genome the variants are located) can be inherited together. Variants that are usually inherited together are said to be in “linkage disequilibrium”, and this distance between loci that are inherited together may differ depending upon where the loci and how they are distributed in the genome.

[0128] When a single variant contributes to the risk of a disease, a GWAS will reveal statistical associations of that variant (the “causal variant”) with the disease, but alsoAttorney Docket No.136622-1010 statistical associations with the disease for other variants that are in linkage disequilibrium (inherited together) with the causal variant. Therefore, in GWAS, when there is a group of variants that are close together in their position on a DNA molecule, and many of those variants are associated with the disease, the data processing system can distinguish between two scenarios: 1) a scenario in which all of the apparent disease-associated variants are inherited together, which would imply that there is likely to be a single “causal variant”, or 2) a scenario in which there are two or more distinct groups of variants, with one or more variants in each group, such that variants within each group are inherited together, but the inheritance of one group vs. another is independent. In this latter case, there is likely to be one “causal variant” per group of linked variants.

[0129] To distinguish between the “single causal variant” and the “multiple causal variants” scenario, the data processing system can implement a regression function. For example, first, a GWAS (e.g., a first GWAS) can be performed as normal. The data processing system can identify the variant that has the most significant association with the disease outcome (e.g., the smallest p-value) across the whole genome. The data processing system can rerun the GWAS, this time conditioning all of the variant-disease association tests on the genotype of the most-significant variant. The data processing system can identify the variant that has the next-most significant association with the disease outcome. The data processing system can rerun the GWAS, this time conditioning all of the variant-disease association tests on the genotype of the most-significant and the 2nd-most-significant variants. The data processing system can iteratively repeat this procedure until determining no variant is associated with the disease below a defined p-value threshold (e.g., a genome-wide significance threshold, such as 5 x 10-8). Using lower thresholds can reduce the chances that an analytical result is a fluke but increases the chance that the data processing will not identify an important variant, and vice versa. The data processing system can use any threshold. Responsive to stopping the procedure, the set of variants the data processing system identified can be called “conditionally-independent disease-associated variants” because each variant is associated with the disease even after conditioning on the genotypes of all of the previous variants.

[0130] Returning to the “single causal variant” and “multiple causal variants” scenarios, if the data processing system identified only a single “conditionally-independent disease- associated variant” in a given region of DNA during the stepwise regression, then the “single causal variant” scenario is more likely to be true. However, if the stepwise regression identified several “conditionally-independent disease-associated variants”, then the “multipleAttorney Docket No.136622-1010 causal variants” scenario is more likely to be true.

[0131] In instances in which a disease (e.g., AMD) is associated with multiple biomarkers (e.g., ISOS RPE thickness, AMD lesion size), the data processing system may use multiple independent acceptance criteria to identify biomarker-associated variants to increase the number of variants that the data processing system identifies. For example, the data processing system can identify variants with p < 5 x 10-8for the disease and variants with p < 5 x 10-6for the disease that also have p < 5 x 10-8for at least one disease-relevant biomarker. The data processing system can use any criteria to identify variants. The variations in acceptance criteria can enable the data processing system to identify a wide range of variants using a COJO-SLCT analysis.

[0132] The data processing system can create a table. Each row of this table can correspond to one subject from the study cohort. The columns of the table are described as follows: a. Whether the subject is a case or control for the disease. Cases are assigned a value of “1” and controls are assigned a value of “0”. Subjects who are neither cases nor controls are excluded from the table. b. Age and sex of the subject. c. The PCs of genetic ancestry from the operation 602 or 604. d. The scaled PGS for disease risk from the operation 604. e. The subject’s dosage for the alternate allele of each of the independent disease-associated variants. i. Here, “alternate allele” is in reference to a reference genome such as the Genome Reference Consortium’s GRCh37 or GRCh38. In a reference genome, each variant has a “reference allele”. A bi-allelic variant is therefore a site in the genome where an individual has either the “reference allele” or an “alternate allele” at the site. ii. Some sites are “multi-allelic”, meaning that in the population, three or more alleles exist. Computationally, each alternate allele is treated as a separate bi-allelic variant, (e.g., if one had alleles “reference”, “alternate-1”, and “alternate-2”, then it would be treated as two bi- allelic variants (reference, alternate-1) and (reference, alternate-2)).

[0133] The data processing system can identify which of the disease-associated variants are “pharmacomimetic instruments” (e.g., variants that modulate the function or expression of one or more drug target genes so that the variants’ effect on human phenotypes is likely to be predictive of the drug’s effect on human phenotypes). Depending on the configuration, the data processing system can use different criteria and / or priorities to identify pharmacomimetic instruments from the identified disease-associated variants. In oneAttorney Docket No.136622-1010 example, a user may be interested in clinical-stage drugs for the disease. In this case, the data processing system can compile a table of clinical-stage drugs and their targets using public resources, such as clinicaltrials.gov, and / or proprietary databases, such as Cortellis. The data processing system can determine the variants that are close to a drug target gene (e.g., within 150 kb of the gene’s transcription start site) or that are mapped to the drug target are pharmacomimetic instruments. In another example, a user may be interested in known and novel drug targets for the antibody modality. In this case, the data processing system can compile a list of genes that encode proteins that are druggable with the antibody modality (e.g., proteins that are secreted or localized to the cell surface). The data processing system can determine variants that are mapped to an antibody-druggable protein are pharmacomimetic instruments. The data processing system can determine variants that are mapped to an antibody-druggable protein are pharmacomimetic instruments based on the compiled table (e.g., based on the determined variants having stored associations with the antibody-druggable protein).

[0134] Optionally, one might require that genetic evidence supports the hypothesis that inhibition of the target would be beneficial, as opposed to activation of the target. This evaluation can consider data from expression quantitative trait locus (eQTL) and protein quantitative trait locus (pQTL) datasets that inform on how genetically-determined changes in expression of the target affect disease risk.

[0135] In some embodiments, there is more than one suitable variant for a drug mechanism. In this case, the data processing system can increase statistical power to detect drug-PGS interactions by aggregating variants (e.g., all of the variants) identified as pharmacomimetic instruments that were associated with a given drug mechanism into a combined allelic score. To do so, for example, the data processing system can compute an allelic score in the same or a similar manner to a PGS (e.g., assign each variant a weight and compute, for each subject in the cohort, the sum over each variant of (the variant’s weight) * (the subject’s dosage for the variant)).

[0136] The difference between an allelic score and a PGS is that the allelic score is composed of variants that modulate the function or expression of a single drug target gene, while a PGS can be composed of variants across the genome that affect many different genes, e.g., ANGPTL3, ANGPTL4, and / or LPL. The variants in the allelic score may be weighted using a GWAS (e.g., a second GWAS) for the disease that did not include any subjects from the study cohort. Alternatively, the variants can be weighted using a GWAS (e.g., a thirdAttorney Docket No.136622-1010 GWAS) for a biomarker that causally mediates the effect of the drug target on the disease. A biomarker GWAS still must not overlap the study cohort. The allelic score can be used in place of PGS to perform the systems and methods described herein.

[0137] At operation 608, the data processing system determines statistical interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores. The statistical interactions can be predictive of drug target-specific differential treatment responses for the disease of interest. To determine the statistical interactions, the data processing system can generate a data table with phenotypes (e.g., wet and dry age-related macular degeneration, and choroidal neovascularization, etc.), covariates, and a set of pharmacomimetic instruments determined using the aforementioned systems and methods, which may include single variants and multi-variant allelic scores. Responsive to doing so, the data processing system can test for interactions between the pharmacomimetic instruments and the PGS (or allelic scores) in association analysis with the disease or risk factor trait of interest.

[0138] For example, for each pharmacomimetic instrument, the data processing system can use statistical software, such as the R programming language, to fit a logistic regression model using the following example formula: disease case-control status ~ age + sex + ancestry PCs + (pharmacomimetic instrument) + (scaled PGS) + (pharmacomimetic instrument) : (scaled PGS). The model can be implemented using the R programming language, such as by using the command glm(model_formula, family = “binomial”, data = our_genetic_and_phenotypic_data). In the example formula, the variable to the left of the symbol is the dependent variable (e.g., the variable that is being predicted). Each of the variables to the right of the “~” symbol, separated by “+” symbols, are independent variables, (e.g., variables that are observed in the data and used to predict the dependent variable). The “:” symbol indicates a variable that is the product (result of multiplication) of the variables to the left and right of the “:” symbol. This product is an “interaction term”. If the pharmacomimetic instrument : scaled polygenic score (PGS) interaction term -- one of the independent variables -- has a statistically-significant association with the dependent variable (in this case, disease status), controlling for all of the other independent variables -- such as age and sex -- then the data processing system can predict or determine that the drug that is modeled by the pharmacomimetic instrument will have different effects in subjects with high PGS vs. low PGS. The data processing system can fit a logistic regression model for each pharmacomimetic instrument that the data processing system determines or identifies, such asAttorney Docket No.136622-1010 based on the PGS of the different subjects.

[0139] The data processing system can use the same program to compute a p-value for the interaction term. If the p-value for the interaction between a pharmacomimetic instrument and the PGS is < 0.05 / (the # of instruments tested), the data processing system can determine the interaction to be “Bonferroni significant.” Alternatively, the data processing system can be configured to use a more lenient p-value threshold, at the risk of false positives.

[0140] The instrument-PGS interactions can be “drug-PGS interactions” because each pharmacomimetic instrument is a model for a particular drug mechanism.

[0141] The data processing system can generate tables that present the statistically significant drug-PGS interactions in a way that is easier to interpret than looking at the raw regression model outputs. For example, first, the data processing system can group the subjects of the study cohort by quantiles of the PGS, (e.g., the 0-33rd percentile, the 34th-66th percentile, and the 67th-100th percentile). Then, for each drug-PGS interaction, the data processing system can fit separate logistic regression models within each PGS quantile using the formula: disease case-control status ~ age + sex + ancestry PCs + the pharmacomimetic instrument for the drug. Using these regression outputs, the data processing system can construct a table that shows the effect of the pharmacomimetic instrument on disease risk (as well as a 95% confidence interval for that effect) within each PGS quantile.

[0142] For instance, the data processing system can divide the subjects of the study into a finite number of groups (e.g., the groups) based on their PGS values. Within each group, the data processing system can predict a disease outcome based on the patients’ genetics, specifically a genetic instrument that mimics the effects of a drug (the “pharmacomimetic instrument”). In doing so, the data processing system can determine a difference in how strongly the pharmacomimetic instrument predicts the disease outcome in individuals with high PGS vs. middle PGS vs. low PGS. The difference may indicate that a drug might have different efficacy in people with high PGS vs. middle PGS vs. low PGS.

[0143] The data processing system can also compute TEM, e.g., an effect of the pharmacomimetic instrument on disease risk in the top quartile divided by the effect in all subjects. The TEM can represent how much one can increase the average treatment effect in a clinical trial of the drug if one enrolled only patients in the top quantile of the PGS as opposed to enrolling all qualified patients. The TEM is a way to compare drug-PGSAttorney Docket No.136622-1010 interactions to distinguish “strong” from “weak” interactions.

[0144] At operation 610, the data processing system performs one or more prediction-based actions. The data processing system can perform the one or more prediction-based actions based on the determined interactions. The one or more prediction-based actions can include at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics. To perform a prediction-based action, the data processing system can select a patient for treatment. The data processing system can select the patient based on a likelihood that the patient will benefit from the treatment. For example, the data processing system can select the patient by determining a PGS for the patient using the systems and methods described herein and determining the PGS for the patient is above a threshold (e.g., a PGS threshold) or within a range (e.g., a quartile, a quintile, a defined range, etc.). In another example, the data processing system can select patients for novel therapeutic mechanisms whose benefit is mostly or only apparent in subsets of common disease patients who have elevated polygenic risk. This enables novel target identification and novel therapeutic discovery. This application can be called “polygenic therapeutic target ID.” In another example, the data processing system can select patients from large common disease treatment-eligible populations who will benefit most from existing drug therapies, and, conversely, identify patients who are not likely to receive clinically meaningful benefit. This can enhance the pharmaco-economic profile of existing therapeutic mechanisms, yielding more cost-effective utilization. This can be called “polygenic pharmacogenomics.” In another example, the data processing system can select mechanisms and design of prospective clinical trials of investigational drugs that are predicted to yield greater benefit in individuals with elevated PGS. These trials may be run with many-fold fewer patient years required to demonstrate clinical benefit, given the expectation of magnified event rate and treatment response rates. This application can be called “polygenic therapeutic development.” In some embodiments, the injections, therapy, and / or treatment techniques can be performed on the selected individuals when performing the one or more prediction- based actions.

[0145] In one example, the systems and methods described herein can be used to treat a patient. For instance, a patient can visit a clinic. The clinician can collect a blood sample, a saliva sample, or a tissue sample from the patient. The clinician can use the collected sample to perform a genotyping array on the patient. A data processing system implementing the systems and methods described herein can analyze the patient’s genetic data to generateAttorney Docket No.136622-1010 principal components (PCs) that reflect the patient’s genetic ancestry. The patient can undergo tests to identify relevant biomarkers for a disease of interest (e.g., AMD, CNV, etc.). These biomarkers are quantifiable biological parameters that include patient characteristics such as blood sugar levels, cholesterol levels, specific protein markers, etc., depending on the disease. The data processing system can calculate a biomarker stratifier score for the patient for a disease of interest (e.g., a disease of interest from the list above) based on the results from the biomarker tests and the genetic ancestry data (e.g., the PCs).

[0146] The data processing system can scan the patient’s genetic data for disease-associated variants. In doing so, the data processing system can identify genes or gene variants in the patient’s genetic data that are associated with the disease for which the data processing system determined the biomarker stratifier score for the patient. For instance, the data processing system can identify disease-associated variants for the disease using the systems and methods described herein. The data processing system can use the identified disease- associated variants as a key in a query through the genetic array data of the patient. Based on the query, the data processing system can identify any disease-associated variants in the patient’s genetic data.

[0147] The data processing system can determine which of the identified disease-associated variants are pharmacomimetic instruments (e.g., whether these genetic variants can predict how the patient might respond to certain drugs or methods of treatment). For example, prior to determining treatment for the patient, the data processing system may have identified a list of pharmacomimetic instruments for the disease of interest using the systems and methods described herein. The data processing system can compare the identified disease-associated variants in the patient’s genetic makeup with the list of pharmacomimetic instruments. Based on the comparison, the data processing system can identify a set of pharmacomimetic instruments of the patient. In some embodiments, the data processing system can identify the pharmacomimetic instruments of the patient by querying the patient’s genetic data for instances of the pharmacomimetic instruments for the disease without identifying the disease- associated variants.

[0148] As demonstrated herein, the data processing system can determine one or more treatments for the patient, risk for disease progression or disease development based on the pharmacomimetic instruments and the biomarker stratifier score of the patient. The data processing system can do so, for example, based on statistical interactions that the data processing system has previously determined for the pharmacomimetic instruments andAttorney Docket No.136622-1010 biomarker stratifier scores. For instance, the data processing system can store a record (e.g., a record based on logistic regressions of PGS or biomarker stratifier scores and the pharmacomimetic instruments) indicating that individuals with a high polygenic risk score or biomarker stratifier score (e.g., above a threshold) for the disease of interest e.g., AMD, may have a high positive response to a therapeutic inhibitor (represented by a pharmacomimetic instrument) and individuals with a low polygenic risk score or biomarker stratifier score (e.g., below the threshold) for the disease of interest may have a low positive response to the inhibitor. The patient may have a high polygenic risk score or biomarker stratifier score. Accordingly, the data processing system can generate a record (e.g., a file, notification, alert, user interface, data structure, etc.) indicating or recommending to treat the patient with a therapeutic inhibitor. The data processing system can display the record on a user interface of a client device accessed by the clinician. The data processing system can recommend any form of treatment based on interactions between pharmacomimetic instruments and polygenic risk scores biomarker stratifier scores of patients.

[0149] The clinician can treat the patient based on the recommendation. For example, the clinician can view the recommendation to use the inhibitor to treat the patient. The clinician can administer the treatment, such as by administering pills or capsules, administering an injection or set of injections, or using gene editing tools. The clinician can perform any type of treatment based on the polygenic risk scores or biomarker stratifier scores, type of disease, and / or pharmacomimetic instruments of the patient.

[0150] In some embodiments, the data processing system can automatically perform a treatment (e.g., using a gene editing tool or another treatment device or mechanism) responsive to determining the treatment. For example, the data processing system can determine a treatment to inhibit a particular gene to treat a disease for a patient as described herein. The data processing system can configure a treatment machine (e.g., a gene editing tool) to automatically implement the treatment on the patient responsive to the determination, such as by controlling the machine to perform the treatment (e.g., inhibiting or otherwise editing a specific gene or by performing an injection).

[0151] In another example, the systems and methods described herein can be used to perform a clinical trial to determine the efficacy of a candidate drug. For instance, the data processing system can identify a disease of interest that the candidate drug is intended to treat. The data processing system can identify pharmacomimetic instruments for the disease of interest. The data processing system can identify a pharmacomimetic instrument for the candidate drug,Attorney Docket No.136622-1010 such as based on a user input. The data processing system can receive genetic data of a study cohort. The data processing system can determine polygenic risk scores for the study cohort using the systems and methods described herein based on the genetic data.

[0152] The data processing system can select participants for the trial based on the polygenic risk scores and interaction data between the polygenic risk scores and the pharmacomimetic instrument for the candidate drug. For instance, the data processing system can identify interactions between the pharmacomimetic instrument and the polygenic risk scores. From the identified interactions, the data processing system can determine that there is a high positive interaction between high polygenic risk scores (e.g., polygenic risk scores that exceed a threshold) and the pharmacomimetic instrument. Accordingly, the data processing system may filter participants from the study cohort and identify participants for the study cohort that have a high polygenic risk score. The data processing system can generate a record including the list of patients identified for the trial. The data processing system can present the list of patients on a user interface to a clinician and / or perform treatment on the participants in the trial as described above.

[0153] Figure 6 illustrates a sequence 700 for in silico drug development and determination of drug activity in one embodiment of the disclosure. The sequence 700 includes an illustration of matrices that can be used in performing the systems and methods described herein. A data processing system (e.g., a client device or the data processing system 302, shown and described with reference to FIG.2, a server system, etc.) can perform the operations and generate the matrices of the sequence 700. The sequence 700 may include more or fewer operations and the operations may be performed in any order. Performance of the method 700 may enable the data processing system to automatically determine interactions between pharmacomimetic instruments and perform prediction-based actions based on the determined interactions.

[0154] For example, the data processing system performing the sequence 700 can generate a variant effects matrix 702. The data processing system can generate the variant effects matrix 702 to have columns 704a-n (columns 704) and rows 706a-n (rows 706). n can be any number and can vary between the columns 704 and the rows 706. The columns 704 can each correspond to a different phenotype. In some embodiments, the column 704a can correspond to a reference phenotype and the columns 704b-n can be phenotypes that are biologically relevant to the reference phenotype of the column 704a. The rows 706 can each correspond to variant effects of separate variants on the phenotypes of the respective columns 704. TheAttorney Docket No.136622-1010 data processing system can determine the values for the variant effects and insert the values into the variant effects matrix 702.

[0155] The data processing system can perform a truncated principal components analysis (PCA) 708 on the variant effects matrix 702 to generate a variant-PC loadings matrix 710. The data processing system can generate the variant-PC loadings matrix 710 to have columns 712a-n (columns 712) and rows 714a-n (rows 714). The columns 712 can each correspond to a different principal component. The rows 706 can each correspond to a different variant. The values in the intersections between the columns 712 and the rows 714 can indicate variant loads for particular principal components.

[0156] The data processing system can generate or calculate a weight (e.g., a new weight) for each PC and variant combination 716. The data processing system can use the weight to generate a new PGS for the phenotypes. The data processing system can do so using the following equation: A variant’s new weight for PCi= the variant’s old weight * (the variant’s loading for PCi)2 / (sumj= 1…k of [the variant’s loading for PCj]2). The variant’s old weight can be the weight for the variant that was used to determine an existing PGS 718. The data processing system can use the newly calculated weights to calculate PGSs for phenotypes using the systems and methods described herein. In one example, the data processing system can include weights and / or the variant-PC loadings matrix 710 or PCs in a feature vector, in some cases with the PCs themselves. The data processing system can execute a neural network or linear regression model based on the feature vector to generate the PGSs for different individuals. The data processing system can use the PGSs to identify patients for treatment and / or entities may treat patients based on the PGSs.

[0157] The data processing system can perform the PCA to reduce the memory processing resources that are required to determine biomarker stratifier scores for individuals and / or identify interactions between the pharmacomimetic instruments and biomarker stratifier scores that are indicative of a drug target-specific differential treatment response for a human phenotype. For example, genetic ancestry data can include millions of different data points for different individuals. Processing each data point using machine learning techniques, such as using a neural network or a regression model, can require a large amount of memory because machine learning models and other computer models have limits on the size of the data that can be used as input (e.g., a neural network may only be configured to receive a limited amount of data into input nodes of the input layer of the neural network or a regression model may be limited in the number of dimensions that the regression model isAttorney Docket No.136622-1010 trained to implement). By employing principal component analysis and generating variant- PC loadings matrices, the data processing system can substantially compress the dimensionality of complex genetic, proteomic, transcriptional, and / or somatic mutational data while preserving essential variance information. The data processing system can generate one or more feature vectors from the output PCs and / or the variant PC loadings matrices that can be used as input into a machine learning model (e.g., a machine learning model trained to configured to generate biomarker stratifier scores for individuals) to generate biomarker stratifier scores and / or interactions between pharmacomimetic instruments and biomarker stratifier scores. The biomarker stratifier scores can be used to generate interactions that are indicative of a drug target-specific differential treatment response for a human phenotype.

[0158] In doing so, the data processing system can format and reduce the size of each feature vector that the data processing system generates to use as input into a machine learning model (e.g., a regression model or a neural network) to facilitate the machine learning model’s ability to process the large amount of data and reduce the processing requirements of executing the machine learning model. Accordingly, the data processing system can address the limitations of how machine learning models are configured and additionally address computational bottlenecks inherent to machine learning models when processing extremely large datasets that would otherwise overwhelm available memory resources of a computer attempting to do so. Variant data for polygenic scores

[0159] There are many possible approaches to combine information across loci for assessment of polygenic risk scores (PRSs) for various conditions. Many studies have shown that PRSs can predict disease status in research-based case-control studies. See, e.g., Mavaddat N, et al. (2019) Am J Hum Genet. Vol.104:21–34; Wray N Ret al. (2018) Nat Genet. Vol.50:668–81. More convincingly, the prediction is also valid in population-based cohort studies and in electronic health record-based studies, especially for psychiatric disorders. Musliner KL et al. (2019) JAMA Psychiatry. Vol.76:516–25; Lewis CM and Hagenaars SP. (2019) JAMA Psychiatry. Vol.76:470-472; Zheutlin AB, et al. (2019) Am J Psychiatry. Vol.176(10):846–55. The PRS can be formed from a set of independent risk variants associated with a disorder, based on the current evidence from the largest or most informative genome-wide association studies. For each individual, the number of risk alleles carried at each variant (0, 1, or 2) is summed, weighted by its effect size (i.e., log (OR) for binary traits or beta coefficient for continuous traits). The outcome is a single score of eachAttorney Docket No.136622-1010 individual’s genetic loading for a disease or for a continuous trait.

[0160] Much of the research on polygenic scores comes from research studies in cardiovascular disease, type 2 diabetes, breast and prostate cancers, and Alzheimer’s disease (Lambert SA et al. (2019) Hum Mol Genet. Vol.28(R2): R133-42). Studies using the UK Biobank have demonstrated that PRS based on variant data can identify which percentage of patients have at least 3-fold increased risk for coronary artery disease, atrial fibrillation, type 2 diabetes, inflammatory bowel disease, and breast cancer, with the proportion of individuals identified varying between 1.5 and 8% depending on the disorder. Khera AV et al. (2018) Nat Genet. Vol.50:1219–24. Although these effects appear modest, PRS can identify substantial larger fractions of the population at high disease risk than monogenic mutations, making PRS potentially more clinically relevant. Definitions

[0161] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.

[0162] As used herein, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. For example, “a” or “an” means “at least one” or “one or more.” It is understood that aspects and variations described herein include “consisting of’ and / or “consisting essentially of’ aspects and variations.

[0163] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For example, where a range of values is provided, it is understood that each intervening value, between the upper and lower limit of that range and any other stated or intervening value in that stated range is encompassed within the claimed subject matter. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the claimed subject matter,Attorney Docket No.136622-1010 subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the claimed subject matter. This applies regardless of the breadth of the range.

[0164] “About” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ±20% or± 10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% from the specified value, as such variations are appropriate to perform the disclosed methods.

[0165] As used herein, the term “comprising” is intended to mean that the compositions and methods include the recited elements, but not excluding others. “Consisting essentially of” when used to define compositions and methods, shall mean excluding other elements of any essential significance to the composition or method. “Consisting of” shall mean excluding more than trace elements of other ingredients for claimed compositions and substantial method steps. Embodiments defined by each of these transition terms are within the scope of this disclosure. Accordingly, it is intended that the methods and compositions can include additional steps and components (comprising) or alternatively including steps and compositions of no significance (consisting essentially of) or alternatively, intending only the stated method steps or compositions (consisting of).

[0166] The drugs or treatments disclosed herein may be first line, second line or third line therapies. The phrase “first line” or “second line” or “third line” refers to the order of treatment received by a patient. First line therapy regimens are treatments given first, whereas second or third line therapy are given after the first line therapy or after the second line therapy, respectively. The National Cancer Institute defines first line therapy as “the first treatment for a disease or condition.

[0167] The term “ age related macular degeneration” or “AMD” is the most common cause of severe loss of eyesight among people aged 50 and older. It specifically affects the center of vision, impairing the ability to see fine details. Although it rarely leads to complete blindness, AMD can significantly impact daily activities such as driving, recognizing faces, and reading smaller print. There are two types of AMD. Dry AMD is the most common type, accounting for about 80% of AMD cases. The exact cause is unknown, but it involves the gradual breakdown of light-sensitive cells in the macula. Vision loss in dry AMD typically progresses slowly over time. Wet AMD wet AMD often leads to more severe vision loss. It occurs when abnormal blood vessels grow beneath the retina, leaking fluid and blood. This canAttorney Docket No.136622-1010 create large blind spots in the center of the visual field. (See hopkinsmedicine.org / health / conditions-and-diseases / agerelated-macular-degeneration-amd, last accessed on May 10, 2024).

[0168] “Choroidal neovascularization” or “CNV” as used herein describes the growth of new blood vessels that originate from the choroid through a break in the Bruch membrane into the sub–retinal pigment epithelium (sub-RPE) or subretinal space. This is a major cause of vision loss.

[0169] A complement blood test measures the amount or activity of complement proteins in the blood. Complement proteins are part of the complement system. This system is made up of a group of proteins that work with the immune system to identify and fight disease-causing substances like viruses and bacteria. There are nine major complement proteins. They are labeled C1 through C9. Complement proteins may be measured individually or together. C3 and C4 proteins are the most commonly tested individual complement proteins. A CH50 test (sometimes called CH100) measures the amount and activity of all the major complement proteins.

[0170] Complement factor B (CFB) (GenBank AQY76745.1, May 10, 2024) is a protein in humans encoded by the CFB gene. CFB is a component of the alternative pathway of complement activation. Factor B circulates in the blood as a single chain polypeptide. Upon activation of the alternative pathway, it is cleaved by complement factor D yielding the noncatalytic chain Ba and the catalytic subunit Bb. The active subunit Bb is a serine protease which associates with C3b to form the alternative pathway C3 convertase. Bb is involved in the proliferation of preactivated B lymphocytes, while Ba inhibits their proliferation. This gene localizes to the major histocompatibility complex (MHC) class III region on chromosome 6. This cluster includes several genes involved in regulation of the immune reaction. Polymorphisms in this gene are associated with a reduced risk of age-related macular degeneration.

[0171] Complement factor H (CFH) gene is a member of the Regulator of Complement Activation (RCA) gene cluster and encodes a protein with twenty short consensus repeat (SCR) domains. This protein is secreted into the bloodstream and has an essential role in the regulation of complement activation, restricting this innate defense mechanism to microbial infections. Mutations in this gene have been associated with hemolytic-uremic syndrome (HUS) and chronic hypocomplementemic nephropathy.Attorney Docket No.136622-1010 See,ncbi.nlm.nih.gov / gene?Db=gene&Cmd=DetailsSearch&Term=3075, last accessed on May 10, 2024.

[0172] Complement factor I (CFI) gene encodes a serine proteinase that is essential for regulating the complement cascade. The encoded preproprotein is cleaved to produce both heavy and light chains, which are linked by disulfide bonds to form a heterodimeric glycoprotein. This heterodimer can cleave and inactivate the complement components C4b and C3b, and it prevents the assembly of the C3 and C5 convertase enzymes. Defects in this gene cause complement factor I deficiency, an autosomal recessive disease associated with a susceptibility to pyogenic infections. Mutations in this gene have been associated with a predisposition to atypical hemolytic uremic syndrome, a disease characterized by acute renal failure, microangiopathic hemolytic anemia and thrombocytopenia. Primary glomerulonephritis with immune deposits and age-related macular degeneration are other conditions associated with mutations of this gene. See, ncbi.nlm.nih.gov / gene / 3426, last accessed May 10, 2024.

[0173] The term “allele” refers to alternative forms of a gene or portions thereof. Alleles occupy the same locus or position on homologous chromosomes. When a subject has two identical alleles of a gene, the subject is said to be homozygous for the gene or allele. When a subject has two different alleles of a gene, the subject is said to be heterozygous for the gene. Alleles of a specific gene can differ from each other in a single nucleotide, or several nucleotides, and can include substitutions, deletions and insertions of nucleotides. An allele of a gene can also be a form of a gene containing a mutation.

[0174] The term “allelic variant” means a specific allele of determined sequence.

[0175] As used herein, the term “determining the genotype of a cell or tissue sample” intends to identify the genotypes of polymorphic loci of interest in the cell or tissue sample. In one aspect, a polymorphic locus is a single nucleotide polymorphic (SNP) locus. If the allelic composition of a SNP locus is heterozygous, the genotype of the SNP locus will be identified as “X / Y” wherein X and Y are two different nucleotides, e.g., A / G for the rs204993 A / G SNP. If the allelic composition of a SNP locus is heterozygous, the genotype of the SNP locus will be identified as “X / X” wherein X identifies the nucleotide that is present at both alleles, e.g., G / G for the rs204993 A / G SNP.Attorney Docket No.136622-1010

[0176] Some SNPs that are useful for the methods of the present disclosure are summarized in the table below:

[0177] The term “genetic marker” refers to an allelic variant of a polymorphic region of a gene of interest and / or the expression level of a gene of interest.

[0178] The term “wild-type allele” refers to an allele of a gene which, when present in two copies in a subject result in a wild-type phenotype. There can be several different wild-type alleles of a specific gene, since certain nucleotide changes in a gene may not affect the phenotype of a subject having two copies of the gene with the nucleotide changes.Attorney Docket No.136622-1010

[0179] The term “polymorphism” refers to the coexistence of more than one form of a gene or portion thereof. A portion of a gene of which there are at least two different forms, i.e., two different nucleotide sequences, is referred to as a “polymorphic region of a gene.” A polymorphic region can be a single nucleotide, the identity of which differs in different alleles.

[0180] A “polymorphic gene” refers to a gene having at least one polymorphic region.

[0181] The term “genotype” refers to the specific allelic composition of an entire cell or a certain gene and in some aspects a specific polymorphism associated with that gene, whereas the term “phenotype” refers to the detectable outward manifestations of a specific genotype.

[0182] The phrase “amplification of polynucleotides” includes methods such as PCR, ligation amplification (or ligase chain reaction, LCR) and amplification methods. These methods are known and widely practiced in the art. See, e.g., U.S. Pat. Nos.4,683,195 and 4,683,202 and Innis et al., 1990 (for PCR); and Wu, D.Y. et al. (1989) Genomics 4:560-569 (for LCR). In general, the PCR procedure describes a method of gene amplification which is comprised of (i) sequence-specific hybridization of primers to specific genes within a DNA sample (or library), (ii) subsequent amplification involving multiple rounds of annealing, elongation, and denaturation using a DNA polymerase, and (iii) screening the PCR products for a band of the correct size. The primers used are oligonucleotides of sufficient length and appropriate sequence to provide initiation of polymerization, i.e., each primer is specifically designed to be complementary to each strand of the genomic locus to be amplified.

[0183] Reagents and hardware for conducting PCR are commercially available. Primers useful to amplify sequences from a particular gene region are preferably complementary to, and hybridize specifically to sequences in the target region or in its flanking regions. Nucleic acid sequences generated by amplification may be sequenced directly. Alternatively the amplified sequence(s) may be cloned prior to sequence analysis. A method for the direct cloning and sequence analysis of enzymatically amplified genomic segments is known in the art.

[0184] The term “encode” as it is applied to polynucleotides refers to a polynucleotide which is said to “encode” a polypeptide if, in its native state or when manipulated by methods well known to those skilled in the art, it can be transcribed and / or translated to produce the mRNA for the polypeptide and / or a fragment thereof. The antisense strand is the complement of such a nucleic acid, and the encoding sequence can be deduced therefrom.Attorney Docket No.136622-1010

[0185] The term “isolated” as used herein refers to molecules or biological or cellular materials being substantially free from other materials. In one aspect, the term “isolated” refers to nucleic acid, such as DNA or RNA, or protein or polypeptide, or cell or cellular organelle, or tissue or organ, separated from other DNAs or RNAs, or proteins or polypeptides, or cells or cellular organelles, or tissues or organs, respectively, that are present in the natural source. The term “isolated” also refers to a nucleic acid or peptide that is substantially free of cellular material, viral material, or culture medium when produced by recombinant DNA techniques, or chemical precursors or other chemicals when chemically synthesized. Moreover, an “isolated nucleic acid” is meant to include nucleic acid fragments which are not naturally occurring as fragments and would not be found in the natural state. The term “isolated” is also used herein to refer to polypeptides which are isolated from other cellular proteins and is meant to encompass both purified and recombinant polypeptides. The term “isolated” is also used herein to refer to cells or tissues that are isolated from other cells or tissues and is meant to encompass both cultured and engineered cells or tissues.

[0186] The term “treating” as used herein is intended to encompass curing as well as ameliorating at least one symptom of the condition or disease. For example, in the case of cancer, a response to treatment includes a reduction in cachexia, increase in survival time, elongation in time to tumor progression, reduction in tumor mass, reduction in tumor burden and / or a prolongation in time to tumor metastasis, time to tumor recurrence, tumor response, complete response, partial response, stable disease, progressive disease, progression free survival, overall survival, each as measured by standards set by the National Cancer Institute and the U.S. Food and Drug Administration for the approval of new drugs. See Johnson et al. (2003) J. Clin. Oncol.21(7):1404-1411.

[0187] The term “suitable for a therapy” or “suitably treated with a therapy” shall mean that the patient is likely to exhibit one or more desirable clinical outcome as compared to patients having the same disease and receiving the same therapy but possessing a different characteristic that is under consideration for the purpose of the comparison. In one aspect, the characteristic under consideration is a genetic polymorphism or a somatic mutation. In another aspect, the characteristic under consideration is expression level of a gene or a polypeptide. In another aspect, a more desirable clinical outcome is relatively lower relative risk. In yet another aspect, a more desirable clinical outcome is relatively reduced toxicity or side effects. In some embodiments, more than one clinical outcomes are considered simultaneously. In one such aspect, a patient possessing a characteristic, such as a genotypeAttorney Docket No.136622-1010 of a genetic polymorphism, may exhibit more than one more desirable clinical outcomes as compared to patients having the same disease and receiving the same therapy but not possessing the characteristic. As defined herein, the patient is considered suitable for the therapy. In another such aspect, a patient possessing a characteristic may exhibit one or more desirable clinical outcome but simultaneously exhibit one or more less desirable clinical outcome. The clinical outcomes will then be considered collectively, and a decision as to whether the patient is suitable for the therapy will be made accordingly, taking into account the patient’s specific situation and the relevance of the clinical outcomes.

[0188] The term “blood” refers to blood which includes all components of blood circulating in a subject including, but not limited to, red blood cells, white blood cells, plasma, clotting factors, small proteins, platelets and / or cryoprecipitate. This is typically the type of blood which is donated when a human patient gives blood.

[0189] A “patient” as used herein intends an animal patient, a mammal patient or yet further a human patient. For the purpose of illustration only, a mammal includes but is not limited to a simian, a murine, a bovine, an equine, a porcine or an ovine. Polymorphic Region

[0190] In addition, knowledge of the identity of a particular allele in an individual (the gene profile) allows customization of therapy for a particular disease to the individual’s genetic profile, the goal of “pharmacogenomics”. For example, an individual’s genetic profile can enable a doctor: 1) to more effectively prescribe a drug that will address the molecular basis of the disease or condition; 2) to better determine the appropriate dosage of a particular drug and 3) to identify novel targets for drug development. The identity of the genotype or expression patterns of individual patients can then be compared to the genotype or expression profile of the disease to determine the appropriate drug and dose to administer to the patient.

[0191] The ability to target populations expected to show the highest clinical benefit, based on the normal or disease genetic profile, can enable: 1) the repositioning of marketed drugs with disappointing market results; 2) the rescue of drug candidates whose clinical development has been discontinued as a result of safety or efficacy limitations, which are patient subgroup-specific; and 3) an accelerated and less costly development for drug candidates and more optimal drug labeling.

[0192] Detection of point mutations or additional base pair repeats can be accomplished by molecular cloning of the specified allele and subsequent sequencing of that allele usingAttorney Docket No.136622-1010 techniques known in the art, in some aspects, after isolation of a suitable nucleic acid sample using methods known in the art. Alternatively, the gene sequences can be amplified directly from a genomic DNA preparation from the tumor tissue using PCR, and the sequence composition is determined from the amplified product. As described more fully below, numerous methods are available for isolating and analyzing a subject’s DNA for mutations at a given genetic locus such as the gene of interest.

[0193] A detection method is allele specific hybridization using probes overlapping the polymorphic site and having about 5, or alternatively 10, or alternatively 20, or alternatively 25, or alternatively 30 nucleotides around the polymorphic region. In another embodiment of the disclosure, several probes capable of hybridizing specifically to the allelic variant are attached to a solid phase support, e.g., a “chip”. Oligonucleotides can be bound to a solid support by a variety of processes, including lithography. For example, a chip can hold up to 250,000 oligonucleotides (GeneChip, Affymetrix). Mutation detection analysis using these chips comprising oligonucleotides, also termed “DNA probe arrays” is described e.g., in Cronin et al. (1996) Human Mutation 7:244.

[0194] In other detection methods, it is necessary to first amplify at least a portion of the gene of interest prior to identifying the allelic variant. Amplification can be performed, e.g., by PCR and / or LCR, according to methods known in the art. In one embodiment, genomic DNA of a cell is exposed to two PCR primers and amplification for a number of cycles sufficient to produce the required amount of amplified DNA.

[0195] Alternative amplification methods include: self-sustained sequence replication (Guatelli et al. (1990) Proc. Natl. Acad. Sci. USA 87:1874-1878), transcriptional amplification system (Kwoh et al. (1989) Proc. Natl. Acad. Sci. USA 86:1173-1177), Q-Beta Replicase (Lizardi et al. (1988) Bio / Technology 6:1197), or any other nucleic acid amplification method, followed by the detection of the amplified molecules using techniques known to those of skill in the art. These detection schemes are useful for the detection of nucleic acid molecules if such molecules are present in very low numbers.

[0196] In one embodiment, any of a variety of sequencing reactions known in the art can be used to directly sequence at least a portion of the gene of interest and detect allelic variants, e.g., mutations, by comparing the sequence of the sample sequence with the corresponding wild-type (control) sequence. Exemplary sequencing reactions include those based on techniques developed by Sanger et al. (1977) Proc. Nat. Acad. Sci, 74:5463). It is alsoAttorney Docket No.136622-1010 contemplated that any of a variety of automated sequencing procedures can be utilized when performing the subject assays (Biotechniques (1995) 19:448), including sequencing by mass spectrometry (see, for example, U.S. Patent No.5,547,835 and International Patent Application Publication Number WO 94 / 16101, entitled DNA Sequencing by Mass Spectrometry by Koster; U.S. Patent No.5,547,835 and international patent application Publication Number WO 94 / 21822 entitled “DNA Sequencing by Mass Spectrometry Via Exonuclease Degradation” by Koster; U.S. Patent No.5,605,798 and International Patent Application No. PCT / US96 / 03651 entitled DNA Diagnostics Based on Mass Spectrometry by Koster; Cohen et al. (1996) Adv. Chromat.36:127-162; and Griffin et al. (1993) Appl. Biochem. Bio.38:147-159). It will be evident to one skilled in the art upon reading this disclosure that, for certain embodiments, the occurrence of only one, two or three of the nucleic acid bases need be determined in the sequencing reaction. For instance, A-track or the like, e.g., where only one nucleotide is detected, can be carried out.

[0197] Yet other sequencing methods are disclosed, e.g., in U.S. Patent No.5,580,732 entitled “Method of DNA Sequencing Employing A Mixed DNA-Polymer Chain Probe” and U.S. Patent No.5,571,676 entitled “Method For Mismatch-Directed In Vitro DNA Sequencing.”

[0198] In some cases, the presence of the specific allele in DNA from a subject can be shown by restriction enzyme analysis. For example, the specific nucleotide polymorphism can result in a nucleotide sequence comprising a restriction site which is absent from the nucleotide sequence of another allelic variant.

[0199] In a further embodiment, protection from cleavage agents (such as a nuclease, hydroxylamine or osmium tetroxide and with piperidine) can be used to detect mismatched bases in RNA / RNA DNA / DNA, or RNA / DNA heteroduplexes (see, e.g., Myers et al. (1985) Science 230:1242). In general, the technique of “mismatch cleavage” starts by providing heteroduplexes formed by hybridizing a control nucleic acid, which is optionally labeled, e.g., RNA or DNA, comprising a nucleotide sequence of the allelic variant of the gene of interest with a sample nucleic acid, e.g., RNA or DNA, obtained from a tissue sample. The double-stranded duplexes are treated with an agent which cleaves single-stranded regions of the duplex such as duplexes formed based on base pair mismatches between the control and sample strands. For instance, RNA / DNA duplexes can be treated with RNase and DNA / DNA hybrids treated with S1 nuclease to enzymatically digest the mismatched regions. In other embodiments, either DNA / DNA or RNA / DNA duplexes can be treated withAttorney Docket No.136622-1010 hydroxylamine or osmium tetroxide and with piperidine in order to digest mismatched regions. After digestion of the mismatched regions, the resulting material is then separated by size on denaturing polyacrylamide gels to determine whether the control and sample nucleic acids have an identical nucleotide sequence or in which nucleotides they are different. See, for example, U.S. Patent No.6,455,249, Cotton et al. (1988) Proc. Natl. Acad. Sci. USA 85:4397; Saleeba et al. (1992) Methods Enzy.217:286-295. In another embodiment, the control or sample nucleic acid is labeled for detection.

[0200] In other embodiments, alterations in electrophoretic mobility are used to identify the particular allelic variant. For example, single strand conformation polymorphism (SSCP) may be used to detect differences in electrophoretic mobility between mutant and wild type nucleic acids (Orita et al. (1989) Proc. Natl. Acad. Sci USA 86:2766; Cotton (1993) Mutat. Res.285:125-144 and Hayashi (1992) Genet Anal Tech. Appl.9:73-79). Single-stranded DNA fragments of sample and control nucleic acids are denatured and allowed to renature. The secondary structure of single-stranded nucleic acids varies according to sequence, the resulting alteration in electrophoretic mobility enables the detection of even a single base change. The DNA fragments may be labeled or detected with labeled probes. The sensitivity of the assay may be enhanced by using RNA (rather than DNA), in which the secondary structure is more sensitive to a change in sequence. In another preferred embodiment, the subject method utilizes heteroduplex analysis to separate double stranded heteroduplex molecules on the basis of changes in electrophoretic mobility (Keen et al. (1991) Trends Genet.7:5).

[0201] In yet another embodiment, the identity of the allelic variant is obtained by analyzing the movement of a nucleic acid comprising the polymorphic region in polyacrylamide gels containing a gradient of denaturant, which is assayed using denaturing gradient gel electrophoresis (DGGE) (Myers et al. (1985) Nature 313:495). When DGGE is used as the method of analysis, DNA will be modified to ensure that it does not completely denature, for example by adding a GC clamp of approximately 40 bp of high-melting GC-rich DNA by PCR. In a further embodiment, a temperature gradient is used in place of a denaturing agent gradient to identify differences in the mobility of control and sample DNA (Rosenbaum and Reissner (1987) Biophys. Chem.265:1275).

[0202] Examples of techniques for detecting differences of at least one nucleotide between 2 nucleic acids include, but are not limited to, selective oligonucleotide hybridization, selective amplification, or selective primer extension. For example, oligonucleotide probes may beAttorney Docket No.136622-1010 prepared in which the known polymorphic nucleotide is placed centrally (allele-specific probes) and then hybridized to target DNA under conditions which permit hybridization only if a perfect match is found (Saiki et al. (1986) Nature 324:163); Saiki et al. (1989) Proc. Natl. Acad. Sci. USA 86:6230 and Wallace et al. (1979) Nucl. Acids Res.6:3543). Such allele specific oligonucleotide hybridization techniques may be used for the detection of the nucleotide changes in the polymorphic region of the gene of interest. For example, oligonucleotides having the nucleotide sequence of the specific allelic variant are attached to a hybridizing membrane and this membrane is then hybridized with labeled sample nucleic acid. Analysis of the hybridization signal will then reveal the identity of the nucleotides of the sample nucleic acid.

[0203] Alternatively, allele specific amplification technology which depends on selective PCR amplification may be used in conjunction with the instant disclosure. Oligonucleotides used as primers for specific amplification may carry the allelic variant of interest in the center of the molecule (so that amplification depends on differential hybridization) (Gibbs et al. (1989) Nucleic Acids Res.17:2437-2448) or at the extreme 3’ end of one primer where, under appropriate conditions, mismatch can prevent, or reduce polymerase extension (Prossner (1993) Tibtech 11:238 and Newton et al. (1989) Nucl. Acids Res.17:2503). This technique is also termed “PROBE” for Probe Oligo Base Extension. In addition, it may be desirable to introduce a novel restriction site in the region of the mutation to create cleavage- based detection (Gasparini et al. (1992) Mol. Cell Probes 6:1).

[0204] In another embodiment, identification of the allelic variant is carried out using an oligonucleotide ligation assay (OLA), as described, e.g., in U.S. Patent No.4,998,617 and in Landegren et al. (1988) Science 241:1077-1080. The OLA protocol uses two oligonucleotides which are designed to be capable of hybridizing to abutting sequences of a single strand of a target. One of the oligonucleotides is linked to a separation marker, e.g., biotinylated, and the other is detectably labeled. If the precise complementary sequence is found in a target molecule, the oligonucleotides will hybridize such that their termini abut, and create a ligation substrate. Ligation then permits the labeled oligonucleotide to be recovered using avidin, or another biotin ligand. Nickerson et al. have described a nucleic acid detection assay that combines attributes of PCR and OLA (Nickerson et al. (1990) Proc. Natl. Acad. Sci. (U.S.A.) 87:8923-8927). In this method, PCR is used to achieve the exponential amplification of target DNA, which is then detected using OLA.Attorney Docket No.136622-1010

[0205] Several techniques based on this OLA method have been developed and can be used to detect the specific allelic variant of the polymorphic region of the gene of interest. For example, U.S. Patent No.5,593,826 discloses an OLA using an oligonucleotide having 3’- amino group and a 5’-phosphorylated oligonucleotide to form a conjugate having a phosphoramidate linkage. In another variation of OLA described in Tobe et al. (1996) Nucleic Acids Res.24: 3728, OLA combined with PCR permits typing of two alleles in a single microtiter well. By marking each of the allele-specific primers with a unique hapten, i.e. digoxigenin and fluorescein, each OLA reaction can be detected by using hapten specific antibodies that are labeled with different enzyme reporters, alkaline phosphatase or horseradish peroxidase. This system permits the detection of the two alleles using a high throughput format that leads to the production of two different colors.

[0206] In one embodiment, the single base polymorphism can be detected by using a specialized exonuclease-resistant nucleotide, as disclosed, e.g., in Mundy, C. R. (U.S. Patent No.4,656,127). According to the method, a primer complementary to the allelic sequence immediately 3’ to the polymorphic site is permitted to hybridize to a target molecule obtained from a particular animal or human. If the polymorphic site on the target molecule contains a nucleotide that is complementary to the particular exonuclease-resistant nucleotide derivative present, then that derivative will be incorporated onto the end of the hybridized primer. Such incorporation renders the primer resistant to exonuclease, and thereby permits its detection. Since the identity of the exonuclease-resistant derivative of the sample is known, a finding that the primer has become resistant to exonucleases reveals that the nucleotide present in the polymorphic site of the target molecule was complementary to that of the nucleotide derivative used in the reaction. This method has the advantage that it does not require the determination of large amounts of extraneous sequence data.

[0207] In another embodiment of the disclosure, a solution-based method is used for determining the identity of the nucleotide of the polymorphic site. Cohen, D. et al. (French Patent 2,650,840; PCT Appln. No. WO91 / 02087). As in the Mundy method of U.S. Patent No.4,656,127, a primer is employed that is complementary to allelic sequences immediately 3’ to a polymorphic site. The method determines the identity of the nucleotide of that site using labeled dideoxynucleotide derivatives, which, if complementary to the nucleotide of the polymorphic site will become incorporated onto the terminus of the primer.

[0208] An alternative method, known as Genetic Bit Analysis or GBA™is described by Goelet, P. et al. (PCT Appln. No.92 / 15712). This method uses mixtures of labeledAttorney Docket No.136622-1010 terminators and a primer that is complementary to the sequence 3’ to a polymorphic site. The labeled terminator that is incorporated is thus determined by, and complementary to, the nucleotide present in the polymorphic site of the target molecule being evaluated. In contrast to the method of Cohen et al. (French Patent 2,650,840; PCT Appln. No. WO91 / 02087) the method of Goelet, P. et al. supra, is preferably a heterogeneous phase assay, in which the primer or the target molecule is immobilized to a solid phase.

[0209] Several primer-guided nucleotide incorporation procedures for assaying polymorphic sites in DNA have been described (Komher, J. S. et al. (1989) Nucl. Acids. Res.17:7779- 7784; Sokolov, B. P. (1990) Nucl. Acids Res.18:3671; Syvanen, A.-C. et al. (1990) Genomics 8:684-692; Kuppuswamy, M. N. et al. (1991) Proc. Natl. Acad. Sci. (U.S.A.) 88:1143-1147; Prezant, T. R. et al. (1992) Hum. Mutat.1:159-164; Ugozzoli, L. et al. (1992) GATA 9:107-112; Nyren, P. et al. (1993) Anal. Biochem.208:171-175). These methods differ from GBA™ in that they all rely on the incorporation of labeled deoxynucleotides to discriminate between bases at a polymorphic site. In such a format, since the signal is proportional to the number of deoxynucleotides incorporated, polymorphisms that occur in runs of the same nucleotide can result in signals that are proportional to the length of the run (Syvanen, A.-C. et al. (1993) Amer. J. Hum. Genet.52:46-59).

[0210] If the polymorphic region is located in the coding region of the gene of interest, yet other methods than those described above can be used for determining the identity of the allelic variant. For example, identification of the allelic variant, which encodes a mutated signal peptide, can be performed by using an antibody specifically recognizing the mutant protein in, e.g., immunohistochemistry or immunoprecipitation. Antibodies to the wild-type or signal peptide mutated forms of the signal peptide proteins can be prepared according to methods known in the art.

[0211] Often a solid phase support is used as a support capable of binding of a primer, probe, polynucleotide, an antigen or an antibody. Well-known supports include glass, polystyrene, polypropylene, polyethylene, dextran, nylon, amylases, natural and modified celluloses, polyacrylamides, gabbros, and magnetite. The nature of the support can be either soluble to some extent or insoluble for the purposes of the present disclosure. The support material may have virtually any possible structural configuration so long as the coupled molecule is capable of binding to an antigen or antibody. Thus, the support configuration may be spherical, as in a bead, or cylindrical, as in the inside surface of a test tube, or the external surface of a rod. Alternatively, the surface may be flat such as a sheet, test strip, etc. orAttorney Docket No.136622-1010 alternatively polystyrene beads. Those skilled in the art will know many other suitable supports for binding antibody or antigen, or will be able to ascertain the same by use of routine experimentation.

[0212] Moreover, it will be understood that any of the above methods for detecting alterations in a gene or gene product or polymorphic variants can be used to monitor the course of treatment or therapy.

[0213] The methods described herein may be performed, for example, by utilizing pre- packaged diagnostic kits, such as those described below, comprising at least one probe or primer nucleic acid described herein, which may be conveniently used, e.g., to determine whether a subject is likely to benefit from a therapy as described herein or has or is at risk of developing AMD or CNV.

[0214] Sample nucleic acid for use in the above-described diagnostic and prognostic methods can be obtained from any suitable cell type or tissue of a subject. For example, a subject’s bodily fluid (e.g., blood) can be obtained by known techniques (e.g., venipuncture). Alternatively, nucleic acid tests can be performed on dry samples (e.g., hair or skin). Diagnostic procedures can also be performed in situ directly upon tissue sections (fixed and / or frozen) of patient tissue obtained from biopsies or resections, such that no nucleic acid purification is necessary. Nucleic acid reagents can be used as probes and / or primers for such in situ procedures (see, for example, Nuovo, G. J. (1992) PCR IN SITU HYBRIDIZATION: PROTOCOLS AND APPLICATIONS, Raven Press, NY).

[0215] In addition to methods which focus primarily on the detection of one nucleic acid sequence, profiles can also be assessed in such detection schemes. Fingerprint profiles can be generated, for example, by utilizing a differential display procedure, Northern analysis and / or RT-PCR.

[0216] Antibodies directed against wild type or mutant peptides encoded by the allelic variants of the gene of interest may also be used in disease diagnostics and prognostics. Such diagnostic methods, may be used to detect abnormalities in the level of expression of the peptide, or abnormalities in the structure and / or tissue, cellular, or subcellular location of the peptide. Protein from the tissue or cell type to be analyzed may easily be detected or isolated using techniques which are well known to one of skill in the art, including but not limited to Western blot analysis. For a detailed explanation of methods for carrying out Western blot analysis, see Sambrook and Russell (2001) supra. The protein detection and isolationAttorney Docket No.136622-1010 methods employed herein can also be such as those described in Harlow and Lane, (1999) supra. This can be accomplished, for example, by immunofluorescence techniques employing a fluorescently labeled antibody (see below) coupled with light microscopic, flow cytometric, or fluorimetric detection. The antibodies (or fragments thereof) useful in the present disclosure may, additionally, be employed histologically, as in immunofluorescence or immunoelectron microscopy, for in situ detection of the peptides or their allelic variants. In situ detection may be accomplished by removing a histological specimen from a patient, and applying thereto a labeled antibody of the present disclosure. The antibody (or fragment) is preferably applied by overlaying the labeled antibody (or fragment) onto a biological sample. Through the use of such a procedure, it is possible to determine not only the presence of the subject polypeptide, but also its distribution in the examined tissue. Using the present disclosure, one of ordinary skill will readily perceive that any of a wide variety of histological methods (such as staining procedures) can be modified in order to achieve such in situ detection.

[0217] In one embodiment, it is necessary to first amplify at least a portion of the gene of interest prior to identifying the polymorphic region of the gene of interest in a sample. Amplification can be performed, e.g., by PCR and / or LCR, according to methods known in the art. Various non-limiting examples of PCR include the herein described methods.

[0218] Allele-specific PCR is a diagnostic or cloning technique is used to identify or utilize single-nucleotide polymorphisms (SNPs). It requires prior knowledge of a DNA sequence, including differences between alleles, and uses primers whose 3’ ends encompass the SNP. PCR amplification under stringent conditions is much less efficient in the presence of a mismatch between template and primer, so successful amplification with an SNP-specific primer signals presence of the specific SNP in a sequence (See, Saiki et al. (1986) Nature 324(6093):163-166 and U.S. Patent Nos.: 5,821,062; 7,052,845 or 7,250,258).

[0219] Assembly PCR or Polymerase Cycling Assembly (PCA) is the artificial synthesis of long DNA sequences by performing PCR on a pool of long oligonucleotides with short overlapping segments. The oligonucleotides alternate between sense and antisense directions, and the overlapping segments determine the order of the PCR fragments thereby selectively producing the final long DNA product (See, Stemmer et al. (1995) Gene 164(1):49-53 and U.S. Patent Nos.: 6,335,160; 7,058,504 or 7,323,336).Attorney Docket No.136622-1010

[0220] Asymmetric PCR is used to preferentially amplify one strand of the original DNA more than the other. It finds use in some types of sequencing and hybridization probing where having only one of the two complementary stands is required. PCR is carried out as usual, but with a great excess of the primers for the chosen strand. Due to the slow amplification later in the reaction after the limiting primer has been used up, extra cycles of PCR are required (See, Innis et al. (1988) Proc Natl Acad Sci U.S.A.85(24):9436-9440 and U.S. Patent Nos.: 5,576,180; 6,106,777 or 7,179,600) A recent modification on this process, known as Linear-After-The-Exponential-PCR (LATE-PCR), uses a limiting primer with a higher melting temperature (Tm) than the excess primer to maintain reaction efficiency as the limiting primer concentration decreases mid-reaction (Pierce et al. (2007) Methods Mol. Med.132:65-85).

[0221] Colony PCR uses bacterial colonies, for example E. coli, which can be rapidly screened by PCR for correct DNA vector constructs. Selected bacterial colonies are picked with a sterile toothpick and dabbed into the PCR master mix or sterile water. The PCR is started with an extended time at 95˚C when standard polymerase is used or with a shortened denaturation step at 100˚C and special chimeric DNA polymerase (Pavlov et al. (2006) “Thermostable DNA Polymerases for a Wide Spectrum of Applications: Comparison of a Robust Hybrid TopoTaq to other enzymes”, in Kieleczawa J: DNA Sequencing II: Optimizing Preparation and Cleanup. Jones and Bartlett, pp.241-257)

[0222] Helicase-dependent amplification is similar to traditional PCR but uses a constant temperature rather than cycling through denaturation and annealing / extension cycles. DNA Helicase, an enzyme that unwinds DNA, is used in place of thermal denaturation (See, Myriam et al. (2004) EMBO reports 5(8):795–800 and U.S. Patent No.7,282,328).

[0223] Hot-start PCR is a technique that reduces non-specific amplification during the initial set up stages of the PCR. The technique may be performed manually by heating the reaction components to the melting temperature (e.g., 95˚C) before adding the polymerase (Chou et al. (1992) Nucleic Acids Research 20:1717-1723 and U.S. Patent Nos.: 5,576,197 and 6,265,169). Specialized enzyme systems have been developed that inhibit the polymerase’s activity at ambient temperature, either by the binding of an antibody (Sharkey et al. (1994) Bio / Technology 12:506-509) or by the presence of covalently bound inhibitors that only dissociate after a high-temperature activation step. Hot-start / cold-finish PCR is achieved with new hybrid polymerases that are inactive at ambient temperature and are instantly activated at elongation temperature.Attorney Docket No.136622-1010

[0224] Intersequence-specific (ISSR) PCR method for DNA fingerprinting that amplifies regions between some simple sequence repeats to produce a unique fingerprint of amplified fragment lengths (Zietkiewicz et al. (1994) Genomics 20(2):176-83).

[0225] Inverse PCR is a method used to allow PCR when only one internal sequence is known. This is especially useful in identifying flanking sequences to various genomic inserts. This involves a series of DNA digestions and self-ligation, resulting in known sequences at either end of the unknown sequence (Ochman et al. (1988) Genetics 120:621-623 and U.S. Patent Nos.: 6,013,486; 6,106,843 or 7,132,587).

[0226] Ligation-mediated PCR uses small DNA linkers ligated to the DNA of interest and multiple primers annealing to the DNA linkers; it has been used for DNA sequencing, genome walking, and DNA footprinting (Mueller et al. (1988) Science 246:780-786).

[0227] Methylation-specific PCR (MSP) is used to detect methylation of CpG islands in genomic DNA (Herman et al. (1996) Proc Natl Acad Sci U.S.A.93(13):9821-9826 and U.S. Patent Nos.: 6,811,982; 6,835,541 or 7,125,673). DNA is first treated with sodium bisulfite, which converts unmethylated cytosine bases to uracil, which is recognized by PCR primers as thymine. Two PCRs are then carried out on the modified DNA, using primer sets identical except at any CpG islands within the primer sequences. At these points, one primer set recognizes DNA with cytosines to amplify methylated DNA, and one set recognizes DNA with uracil or thymine to amplify unmethylated DNA. MSP using qPCR can also be performed to obtain quantitative rather than qualitative information about methylation.

[0228] Multiplex Ligation-dependent Probe Amplification (MLPA) permits multiple targets to be amplified with only a single primer pair, thus avoiding the resolution limitations of multiplex PCR (see below).

[0229] Multiplex-PCR uses of multiple, unique primer sets within a single PCR mixture to produce amplicons of varying sizes specific to different DNA sequences (See, U.S. Patent Nos.: 5,882,856; 6,531,282 or 7,118,867). By targeting multiple genes at once, additional information may be gained from a single test run that otherwise would require several times the reagents and more time to perform. Annealing temperatures for each of the primer sets must be optimized to work correctly within a single reaction, and amplicon sizes, i.e., their base pair length, should be different enough to form distinct bands when visualized by gel electrophoresis.Attorney Docket No.136622-1010

[0230] Nested PCR increases the specificity of DNA amplification, by reducing background due to non-specific amplification of DNA. Two sets of primers are being used in two successive PCRs. In the first reaction, one pair of primers is used to generate DNA products, which besides the intended target, may still consist of non-specifically amplified DNA fragments. The product(s) are then used in a second PCR with a set of primers whose binding sites are completely or partially different from and located 3’ of each of the primers used in the first reaction (See, U.S. Patent Nos.: 5,994,006; 7,262,030 or 7,329,493). Nested PCR is often more successful in specifically amplifying long DNA fragments than conventional PCR, but it requires more detailed knowledge of the target sequences.

[0231] Overlap-extension PCR is a genetic engineering technique allowing the construction of a DNA sequence with an alteration inserted beyond the limit of the longest practical primer length.

[0232] Quantitative PCR (Q-PCR), also known as RQ-PCR, QRT-PCR and RTQ-PCR, is used to measure the quantity of a PCR product following the reaction or in real-time. See, U.S. Patent Nos.: 6,258,540; 7,101,663 or 7,188,030. Q-PCR is the method of choice to quantitatively measure starting amounts of DNA, cDNA or RNA. Q-PCR is commonly used to determine whether a DNA sequence is present in a sample and the number of its copies in the sample. The method with currently the highest level of accuracy is digital PCR as described in U.S. Patent No.6,440,705; U.S. Publication No.2007 / 0202525; Dressman et al. (2003) Proc. Natl. Acad. Sci USA 100(15):8817-8822 and Vogelstein et al. (1999) Proc. Natl. Acad. Sci. USA.96(16):9236-9241. More commonly, RT-PCR refers to reverse transcription PCR (see below), which is often used in conjunction with Q-PCR. QRT-PCR methods use fluorescent dyes, such as Sybr Green, or fluorophore-containing DNA probes, such as TaqMan, to measure the amount of amplified product in real time.

[0233] Reverse Transcription PCR (RT-PCR) is a method used to amplify, isolate or identify a known sequence from a cellular or tissue RNA (See, U.S. Patent Nos.: 6,759,195; 7,179,600 or 7,317,111). The PCR is preceded by a reaction using reverse transcriptase to convert RNA to cDNA. RT-PCR is widely used in expression profiling, to determine the expression of a gene or to identify the sequence of an RNA transcript, including transcription start and termination sites and, if the genomic DNA sequence of a gene is known, to map the location of exons and introns in the gene. The 5’ end of a gene (corresponding to the transcription start site) is typically identified by an RT-PCR method, named Rapid Amplification of cDNA Ends (RACE-PCR).Attorney Docket No.136622-1010

[0234] Thermal asymmetric interlaced PCR (TAIL-PCR) is used to isolate unknown sequence flanking a known sequence. Within the known sequence TAIL-PCR uses a nested pair of primers with differing annealing temperatures; a degenerate primer is used to amplify in the other direction from the unknown sequence (Liu et al. (1995) Genomics 25(3):674-81).

[0235] Touchdown PCR is a variant of PCR that aims to reduce nonspecific background by gradually lowering the annealing temperature as PCR cycling progresses. The annealing temperature at the initial cycles is usually a few degrees (3-5˚C) above the Tmof the primers used, while at the later cycles, it is a few degrees (3-5˚C) below the primer Tm. The higher temperatures give greater specificity for primer binding, and the lower temperatures permit more efficient amplification from the specific products formed during the initial cycles (Don et al. (1991) Nucl Acids Res 19:4008 and U.S. Patent No.6,232,063).

[0236] In one embodiment of the disclosure, probes are labeled with two fluorescent dye molecules to form so-called “molecular beacons” (Tyagi, S. and Kramer, F.R. (1996) Nat. Biotechnol.14:303-8). Such molecular beacons signal binding to a complementary nucleic acid sequence through relief of intramolecular fluorescence quenching between dyes bound to opposing ends on an oligonucleotide probe. The use of molecular beacons for genotyping has been described (Kostrikis, L.G. (1998) Science 279:1228-9) as has the use of multiple beacons simultaneously (Marras, S.A. (1999) Genet. Anal.14:151-6). A quenching molecule is useful with a particular fluorophore if it has sufficient spectral overlap to substantially inhibit fluorescence of the fluorophore when the two are held proximal to one another, such as in a molecular beacon, or when attached to the ends of an oligonucleotide probe from about 1 to about 25 nucleotides.

[0237] Labeled probes also can be used in conjunction with amplification of a gene of interest. (Holland et al. (1991) Proc. Natl. Acad. Sci.88:7276-7280). U.S. Patent No. 5,210,015 by Gelfand et al. describe fluorescence-based approaches to provide real time measurements of amplification products during PCR. Such approaches have either employed intercalating dyes (such as ethidium bromide) to indicate the amount of double-stranded DNA present, or they have employed probes containing fluorescence-quencher pairs (also referred to as the “Taq-Man” approach) where the probe is cleaved during amplification to release a fluorescent molecule whose concentration is proportional to the amount of double- stranded DNA present. During amplification, the probe is digested by the nuclease activity of a polymerase when hybridized to the target sequence to cause the fluorescent molecule to be separated from the quencher molecule, thereby causing fluorescence from the reporterAttorney Docket No.136622-1010 molecule to appear. The Taq-Man approach uses a probe containing a reporter molecule-- quencher molecule pair that specifically anneals to a region of a target polynucleotide containing the polymorphism.

[0238] This disclosure also provides for a prognostic panel of genetic markers selected from, but not limited to the genetic polymorphisms identified herein. The prognostic panel comprises probes or primers or microarrays that can be used to amplify and / or for determining the molecular structure of the polymorphisms identified herein. The probes or primers can be attached or supported by a solid phase support such as but not limited to a gene chip or microarray. The probes or primers can be detectably labeled. This aspect of the disclosure is a means to identify the genotype of a patient sample for the genes of interest identified above.

[0239] In one aspect, the panel contains the herein identified probes or primers as wells as other probes or primers. In a alternative aspect, the panel includes one or more of the above noted probes or primers and others. In a further aspect, the panel consist only of the above- noted probes or primers.

[0240] Primers or probes can be affixed to surfaces for use as microarray. Such gene chips or microarrays can be used to detect genetic variations by a number of techniques known to one of skill in the art. In one technique, oligonucleotides are arrayed on a gene chip for determining the DNA sequence of a by the sequencing by hybridization approach, such as that outlined in U.S. Patent Nos.6,025,136 and 6,018,041. The probes of the disclosure also can be used for fluorescent detection of a genetic sequence. Such techniques have been described, for example, in U.S. Patent Nos.5,968,740 and 5,858,659. A probe also can be affixed to an electrode surface for the electrochemical detection of nucleic acid sequences such as described by Kayem et al. U.S. Patent No.5,952,172 and by Kelley et al. (1999) Nucleic Acids Res.27:4830-4837.

[0241] Various microarrays and similar technologies are known in the art. Examples of such include, but are not limited to LabCard (ACLARA Bio Sciences Inc.); GeneChip (Affymetrix, Inc); LabChip (Caliper Technologies Corp); a low-density array with electrochemical sensing (Clinical Micro Sensors); LabCD System (Gamera Bioscience Corp.); Omni Grid (Gene Machines); Q Array (Genetix Ltd.); a high-throughput, automated mass spectrometry systems with liquid-phase expression technology (Gene Trace Systems, Inc.); a thermal jet spotting system (Hewlett Packard Company); Hyseq HyChip (Hyseq,Attorney Docket No.136622-1010 Inc.); BeadArray (Illumina, Inc.); GEM (Incyte Microarray Systems); a high-throughput microarraying system that can dispense from 12 to 64 spots onto multiple glass slides (Intelligent Bio-Instruments); Molecular Biology Workstation and NanoChip (Nanogen, Inc.); a microfluidic glass chip (Orchid biosciences, Inc.); BioChip Arrayer with four PiezoTip piezoelectric drop-on-demand tips (Packard Instruments, Inc.); FlexJet (Rosetta Inpharmatic, Inc.); MALDI-TOF mass spectrometer (Sequnom, Inc.); ChipMaker 2 and ChipMaker 3 (TeleChem International, Inc.); and GenoSensor (Vysis, Inc.) as identified and described in Heller (2002) Annu. Rev. Biomed. Eng.4:129-153. Examples of microarrays are also described in U.S. Patent Publ. Nos.: 2007 / 0111322, 2007 / 0099198, 2007 / 0084997, 2007 / 0059769 and 2007 / 0059765 and US Patent 7,138,506, 7,070,740, and 6,989,267.

[0242] In one aspect, microarrays containing probes or primers for the gene of interest are provided alone or in combination with other probes and / or primers. A suitable sample is obtained from the patient extraction of genomic DNA, RNA, or any combination thereof and amplified if necessary. The DNA or RNA sample is contacted to the gene chip or microarray panel under conditions suitable for hybridization of the gene(s) of interest to the probe(s) or primer(s) contained on the gene chip or microarray. The probes or primers may be detectably labeled thereby identifying the polymorphism in the gene(s) of interest. Alternatively, a chemical or biological reaction may be used to identify the probes or primers which hybridized with the DNA or RNA of the gene(s) of interest. The genetic profile of the patient is then determined with the aid of the aforementioned apparatus and methods.

[0243] As used herein, a “predetermined threshold” refers to a preset value determined based on a polygenic score distribution of a cohort of patients, and is useful for determining whether a given patient will give a response to a drug or a treatment.

[0244] The term “in silico modeling” refers to use of computational models to predict drug effects and / or health outcomes in different scenarios.

[0245] “Molecular stratified biomarker data” includes at a minimum one of the following types of data: genomic, transcriptomic, metabolomic, or proteomic. These data are obtained by assays specific to each biomarker type.

[0246] Disease phenotypes are based on knowledge of a patient’s disease status (e.g., whether a patient has been diagnosed with AMD or CNV, yes / no) at a given point in time. Disease phenotypes may also include quantitative disease traits (e.g., measurements at a given time with respect to disease diagnosis), medication information (e.g., patient has takenAttorney Docket No.136622-1010 or is taking a drug to treat AMD or CNV at a given time with respect to disease outcomes), and demographic information (e.g., age at diagnosis, biological sex, etc.).

[0247] “Disease outcome” as used herein, means the phenotypes that are associated with a particular disease, e.g., loss of eyesight in patients with AMD.

[0248] As used herein, a “biological pathway” refers to (1) a bespoke set of genes for which their protein products are known to interact in a biological pathway (e.g., complement pathway, JAK / STAT pathway, MAPK pathway, etc.), or (2) an unsupervised learning approach (e.g., principal component analysis (PCA)) that yields patterns in the data such that clusters of genes comprising a pathway may be identified. In some embodiments, approach (1) is based on canonical pathways identified by literature and external pathway databases, while approach (2) is a data-driven analysis that results in identification of pathways.

[0249] As used herein, “biological pathway activity” is defined by information form the literature or external databases that provide evidence for gene or protein expression indicative of pathway activity (e.g., ‘up regulation’ or ‘down regulation’ of pathway X in disease Y).

[0250] As used herein, a “drug target expression” refers to a protein expression as measured in participants of a large cohort (e.g., UK Biobank), where the protein analyzed is a known drug target or is a target that may be druggable (even if not already drugged). In some embodiments, a biomarker stratifier score (e.g., a polygenic score, a proteomics score, a transcriptomics score, or a somatic mutation score) for “a drug target expression” is computed same as it would be for any other quantitative trait, where the outcome of the model is a quantitative measurement.

[0251] As used herein, the term “drug target” refers to any gene or gene product (e.g., RNA or polypeptide) with implications in an associated disease or disorder. Non-limiting examples include various proteins such as enzymes, oncogenes and their polypeptide products, and cell cycle regulatory genes and their polypeptide products.

[0252] As used herein, the phrase “external data” refers to data from an independent set of individuals or population, unused in molecular biomarker stratifier data score determination. In some embodiments, external data is used to assess the predictive power of the methods described herein. In some embodiments, external data is used to prevent overfitting during molecular biomarker stratifier data score determination. At a minimum, the external data is from a large cohort of participants of an observational study (e.g., UK Biobank or eMERGE). In some embodiments, the external data is divided into training and validation subsets. InAttorney Docket No.136622-1010 some embodiments, estimates of biomarker effect are validated in multiple studies, which may include both observational studies and health system patient cohorts.

[0253] As used herein, the phrase “genetic variant” refers to an alteration, variant or polymorphism in a nucleic acid sample or genome of a subject. Such alteration, variant or polymorphism can be with respect to a reference genome, which may be a reference genome of the species (e.g., for human, hGl9 or hG38), the subject or other individual. Variations include one or more single nucleotide variations (SNVs), insertions, deletions, repeats, small insertions, small deletions, small repeats, structural variant junctions, variable length tandem repeats, and / or flanking sequences, copy number variants (CNVs), transversions, gene fusions and other rearrangements are also forms of genetic variation. A variation can be a single nucleotide variation (SNV), insertion or deletion (indel), repeat, copy number variation (CNV), transversion, or a combination thereof.

[0254] As used herein, “a genome-wide association study” (GWA study, or GWAS), refers to an observational study of a genome-wide set of genetic variants in different individuals to see if any variant is associated with a trait. In some embodiments, a GWAS study focuses on associations between gene variants (e.g., single-nucleotide polymorphisms (SNPs)) and phenotypic traits (e.g., diseases). In some embodiments, a genetic variant is a somatic genetic variant. In some embodiments, a genetic variant is a germline genetic variant.

[0255] In some embodiments, GWA studies compare the DNA of participants having varying phenotypes for a particular trait or disease. These participants may be people with a disease (cases) and similar people without the disease (controls), or they may be people with different phenotypes for a particular trait, for example blood pressure. This approach is known as phenotype-first, in which the participants are classified first by their clinical manifestation(s), as opposed to genotype-first. Each person gives a sample of DNA, from which millions of genetic variants are determined (e.g., using SNP arrays or sequencing). If there is significant statistical evidence that one type of the variant (one allele) is more frequent in people with the disease, the variant is said to be associated with the disease. The associated SNPs are then considered to mark a region of the human genome that may influence the risk of disease.

[0256] As used herein, the phrase “molecular biomarker stratifier” or “biomarker stratifier” refers to a molecular marker that is different between two distinct states (e.g., healthy versus diseased).

[0257] As used herein, the phrase “molecular biomarker stratifier distribution” or “biomarkerAttorney Docket No.136622-1010 stratifier distribution” refers to a subset of data obtained from individual-level cohort data (e.g., data from UK Biobank), where subsets may be defined as the top 10%, 20%, 30%, etc. of the distribution. The threshold that is chosen to define a subset of a molecular biomarker stratifier distribution will depend on the use case. For example, it may be optimal to use a 25th percentile of the molecular biomarker stratifier score (e.g., a polygenic score) to define a subset of patients who are predicted to have an outsized clinical benefit from a given therapy. Alternatively, a 33rd percentile of the molecular biomarker stratifier score may be used to define a subset of patients who would benefit most from a second, different given therapy. In other disease settings with different molecular biomarker stratifier scores, the thresholds may also vary. Defining the optimal threshold will depend on factors such as: estimated effect size for the subset vs. all-comers on therapy, and screen failure rate (a “screen failure” is candidate who undergoes screening - meaning they are checked for eligibility - but who does not meet the clinical trial’s inclusion / exclusion criteria in a trial. A “screen failure rate” refers to the number of ineligible candidates divided by the number of screened candidates). If the threshold is set too conservatively, e.g., 5%, a lot more patients will have to be screened to enroll a few eligible candidates.

[0258] As used herein, the phrase “molecular biomarker stratifier effect” or “biomarker stratifier effect” refers to an estimated magnitude of the effect of a biomarker on disease risk or continuous disease trait from a statistical model.

[0259] As used herein, the phrase “molecular biomarker stratifier score” or “biomarker stratifier score” refers to a score that is cumulative of data from across a plurality (hundreds, thousands, tens of thousands or more) of possible biomarker stratifiers that can be used to predict an individual’s risk for an illness. In some embodiments, a statistical model is trained using summary statistics from a published study on large-scale cohorts or from a meta- analysis of multiple studies based on large and diverse set of cohorts. Then, the estimated weights from that model are applied to a dataset that is orthogonal to the ones used for training (e.g., UK Biobank). For example, a Bayesian high-dimensional linear regression model for a set of populations may be used, where body mass index (BMI) is the dependent variable, and a genome-wide set of SNPs and their effect sizes from a published study of BMI from the GIANT Consortium (Locke et al.2015) serve as inputs. Population-specific reference panels (e.g., 1000 Genomes) are used to infer linkage disequilibrium, and population-specific PRS are derived. A linear regression of the normalized PRS for each population is then fit and can be applied to the individual-level genotype and phenotype dataAttorney Docket No.136622-1010 in the UK Biobank.

[0260] In some embodiments, molecular biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of a polygenic score.

[0261] In some embodiments, molecular biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant and the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of a proteomics score.

[0262] In some embodiments, molecular biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant and the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of a transcriptomics risk score.

[0263] In some embodiments, molecular biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant and the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational risk score.

[0264] As used herein, the term “pharmacomimetic” refers to showing a similar response or a similar lack of a response to a therapeutic agent or a class of therapeutic agents.

[0265] As used herein, the phrase “pharmacomimetic variant interaction” refers to a plurality of biomarker stratifiers that in combination correlate with showing a similar response or a similar lack of a response to a therapeutic agent or a class of therapeutic agents.

[0266] As used herein, the phrase “pharmacomimetic instruments” refers to biomarker stratifiers that modulate the function or expression of a drug target gene so that the biomarker stratifiers’ effect on human phenotypes is likely to be predictive of the drug’s effect on human phenotypes.

[0267] As used herein, a “pharmacomimetic genetic score” is a score that indicates the likelihood of a patient to respond to a particular drug or a particular class of drugs. In some embodiments, for a given gene, more than one pharmacomimetic variant (defined as a functional variant in gene that is strongly and significantly associated with a given disease) may be identified. Instead of analyzing the variants of this gene separately, their effects areAttorney Docket No.136622-1010 combined as a “pharmacomimetic genetic score.”

[0268] As used herein, the phrase “polygenic score” or “polygenic risk score” or “PGS” or “PRS” refers to a metric that summarizes the estimated effect of many genetic variants on an individual’s phenotype, typically calculated as a weighted sum of trait-associated alleles. In some embodiments, a polygenic score reflects an individual’s estimated genetic predisposition for a given trait and can be used as a predictor for that trait. In some embodiments, a polygenic score gives an estimate of how likely an individual is to have a given trait only based on genetics, without taking environmental factors into account.

[0269] As used herein, the phrase “proteomics score” refers to a metric that summarizes the estimated effect of many differences in the protein composition (e.g., differences in the amount of proteins, mutational variants of proteins; posttranslational modification of proteins) on an individual’s phenotype, typically calculated as a weighted sum of trait- associated differences. In some embodiments, a proteomics score can be used as a predictor for a given trait. In some embodiments, a proteomics score is used to identify patients most likely to receive an outsized benefit from a particular therapy. In some embodiments, a protein composition of a subject is determined by mass spectroscopy. In some embodiments, a “proteomics score” is estimated from a regression model, e.g., least absolute shrinkage and selection operator (LASSO) (which implements variable selection and regularization for optimal prediction accuracy) for protein measurements vs. outcome, followed by cross- validation procedures.

[0270] As used herein, the phrase “transcriptomics score” refers to a metric that summarizes the estimated effect of many differences in the RNA transcript composition on an individual’s phenotype, typically calculated as a weighted sum of trait-associated differences. In some embodiments, a transcriptomics score refers to a linear or non-linear combination of transcript abundance values that associate with a disease or clinical phenotype. In some embodiments, a transcriptomics score can be used as a predictor for a given trait. In some embodiments, an RNA transcript composition of a subject is determined by RNA-seq, or microarray.

[0271] As used herein, the phrase “somatic mutational risk score” refers to a metric that summarizes the estimated effect of somatic mutations on an individual’s phenotype, typically calculated as a weighted sum of trait-associated somatic mutations. In some embodiments, a somatic mutational risk score can be used as a predictor for a given trait. In someAttorney Docket No.136622-1010 embodiments, a somatic mutational composition of a subject is determined by next generation (high-throughput) sequencing.

[0272] As used herein, the phrase “Principal Component Analysis” (PCA) refers to a technique for analyzing large datasets containing a high number of dimensions / features per observation, increasing the interpretability of data while preserving the maximum amount of information, and enabling the visualization of multidimensional data. In some embodiments, PCA refers to a statistical technique for reducing the dimensionality of a dataset. In some embodiments, reduction the dimensionality is accomplished by linearly transforming the data into a new coordinate system where the variation in the data can be described with fewer dimensions than the initial data. As used herein, “principal components” are a set of new variables that are combinations of the original variables in the original large datasets.

[0273] As used herein, a “statistical interaction” refers to the coefficient of a predictor defined by the product of the pharmacomimetic genetic score and biomarker score in a regression model to predict a phenotype relevant to the drug indication. Specifically, if the p- value for a test of the hypothesis that the coefficient for this PGS-biomarker product term is non-zero is statistically significant (p < 0.05 or p < 0.05 / # of tests performed if multiple hypotheses are tested), then the PGS and biomarker are considered to have a “statistical interaction”. If the sign of the coefficient for the PGS-biomarker product term is the same as the sign of the marginal biomarker term, that indicates that the predictive and / or causal effects of the biomarker on the phenotype are amplified in subjects with high PGS.

[0274] As used herein an “association analysis” refers to a process of searching for hidden association or pattern in a large dataset. In some embodiments, an association analysis is carried out using a statistical model such as linear regression, logistic regression, or Cox Proportional Hazards model, e.g. where the dependent variable is a clinical trait or disease phenotype assessed at a single time point, or censored survival time to disease onset or event, and the independent variables include weighted effects of biomarker stratifiers, demographic information, and other clinical features.

[0275] As used herein, a “phenotype intermediate” refers to a demonstration of partial or incomplete dominance between two or more genes (e.g., two alleles may produce an intermediate phenotype when both are present, rather than one fully determining the phenotype).

[0276] As used herein, “HTRA1” refers to HtrA Serine Peptidase 1 proteinAttorney Docket No.136622-1010 (UniProtKB / Swiss-Prot: Q92743). A pharmacomimetic genetic score “associated with a response to a drug that targets HTRA1” is used to determine whether the subject will respond to a HTRA1 inhibitor. Such are known in the art, e.g., Edgar M. et al. (2024) Inhibition of the serine protease HtrA1 by SerpinE2 suggests an extracellular proteolytic pathway in the control of neural crest migration eLife 12:RP91864 and Gerhardy S. et al. (2022) Allosteric inhibition of HTRA1 activity by a conformational lock mechanism to treat age-related macular degeneration, Nat Commun 13:5222.

[0277] As used herein, the term a “predetermined threshold” refers to a preset value determined based on a polygenic score distribution of a cohort of patients and is useful for determining whether a given patient will give a response to a drug or a treatment. A non- limiting example of such is provided in FIG.7. Modes For Carrying Out The Disclosure

[0278] Analytical methods were developed, and a knowledgebase for identifying patients or individuals more likely to develop AMD or choroidal neovascularization and patients who are likely to experience disease progression. Physicians can utilize this information to determine the best course of treatment based and identify treatments for patients, based on the knowledge which patients are predicted to enjoy greater benefit and / or lower risk from specific therapeutic mechanisms. In one aspect, methods are provided to identify patients that are at risk of developing AMD or choroidal neovascularization. The methods identify combinations of drug targets and disease indications where it is predicted that patients with subsets of polygenic risk scores (PRS) for the disease, its associated risk factors, or an aggregate of several genetically-driven biological pathways, will receive greater benefit and lower risk from a drug compared to patients with a low PRS. In one aspect, the AMD is wet AMD. In another aspect it is dry AMD. In a further aspect, it is both wet and dry AMD (“all AMD”).

[0279] Potential general applications of the method and knowledgebase are: (1) Selection of mechanisms and design of prospective clinical trials of investigational drugs that are predicted to yield greater benefit in individuals with elevated PRS. These trials may be run with many-fold fewer patient years required to demonstrate clinical benefit, given the expectation of magnified event rate and treatment response rates. This application is called “polygenic therapeutic development.”Attorney Docket No.136622-1010 (2) Identification of novel therapeutic mechanisms whose benefit is mostly or only apparent in subsets of common disease patients who have elevated polygenic risk. This enables novel target identification and novel therapeutic discovery. This application is called “polygenic therapeutic target ID.” (3) Selection of patients from large common disease treatment-eligible populations who will benefit most from existing drug therapies, and, conversely, identifying patients who are not likely to receive clinically meaningful benefit. This enhances the pharmaco-economic profile of existing therapeutic mechanisms, yielding more cost- effective utilization. This application is called “polygenic pharmacogenomics.”

[0280] For example, information obtained using the diagnostic assays described herein is useful for determining if a subject is suitable for treatment or a higher risk for disease. Based on the prognostic information, a doctor can recommend a therapeutic protocol, useful for improving treating or preventing AMD or CNV in the individual.

[0281] A patient’s likely clinical outcome following a clinical procedure such as a therapy or surgery can be expressed in relative terms. The patient having a particular PRS can be considered as likely to respond to a particular therapy or likely to develop AMD or choroidal neovascularization.

[0282] It is to be understood that information obtained using the diagnostic assays described herein may be used alone or in combination with other information, such as, but not limited to, genotypes or expression levels of other genes, clinical chemical parameters, histopathological parameters, or age, gender and weight of the subject. When used alone, the information obtained using the diagnostic assays described herein is useful in determining or identifying the clinical outcome of a treatment, selecting a patient for a treatment, or treating a patient, etc. When used in combination with other information, on the other hand, the information obtained using the diagnostic assays described herein is useful in aiding in the determination or identification of clinical outcome of a treatment, aiding in the selection of a patient for a treatment, or aiding in the treatment of a patient and etc. In a particular aspect, the genotypes or expression levels of one or more genes as disclosed herein are used in a panel of genes, each of which contributes to the final diagnosis, prognosis or treatment.

[0283] The methods of this disclosure are useful for the diagnosis, prognosis and treatment of patients who may be at risk for AMD or choroidal neovascularization.Attorney Docket No.136622-1010

[0284] The methods are useful in the assistance of an animal, a mammal or yet further a human patient. For the purpose of illustration only, a mammal includes but is not limited to a human, a simian, a murine, a bovine, an equine, a porcine or an ovine. Diagnostic and Therapeutic Methods

[0285] Provided in one aspect is a method for determining the likelihood of developing all AMD in a patient, wherein the patient’s combined PRS for C3, CFB, CFH, and CFI of as compared to a predetermined threshold for greater likelihood identifies the patient as having a greater likelihood of developing all AMD, and the patient’s combined PRS for C3, CFB, CFH, and CFI as compared to a predetermined threshold for lower likelihood identifies the patient as having a lower likelihood of developing AMD. See Table 2.

[0286] Provided in one aspect is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s C3 PRS to a predetermined threshold for greater likelihood identifies the patient as having a greater likelihood of developing dry AMD, and the patient’s C3 PRS as compared to a predetermined threshold for lower likelihood identifies the patient as having a lower likelihood of developing dry AMD. See Table 3.

[0287] Provided in one aspect is a method for determining the likelihood of developing choroidal neovascularization in a patient, wherein the patient’s C3 PRS as compared to a predetermined value for greater likelihood identifies the patient as having a greater likelihood of developing choroidal neovascularization, and the patient’s C3 PRS as compared to a predetermined value for lower likelihood identifies the patient as having a lower likelihood of developing choroidal neovascularization. See Table 4.

[0288] Provided in one aspect is a method for determining the likelihood of developing all AMD in a patient, wherein the patient’s CFB as compared to a predetermined value for greater likelihood identifies the patient as having a greater likelihood of developing all AMD, and the patient’s CFB PRS as compared to a predetermined value for lower likelihood identifies the patient as having a lower likelihood of developing all AMD. See Table 5.

[0289] Provided in one aspect is a method for determining the likelihood of developing all AMD in a patient, wherein a patient’s CFB PRS as compared to a predetermined value for greater likelihood identifies the patient as having a greater likelihood of developing all AMD, and the patient’s CFR PRS as compared to a predetermined value for lower likelihood identifies the patient as having a lower likelihood of developing AMD. See Table 6.Attorney Docket No.136622-1010

[0290] Provided in one aspect is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFB PRS as compared to a predetermined value for greater likelihood developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD, and the patient’s CFB PRS as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Tables 7 and 8.

[0291] Provided in one aspect is a method for determining the likelihood of developing choroidal neovascularization in a patient, wherein the patient’s CFB PRS as compared to a predetermined value for greater likelihood of developing CNV identifies the patient as having a greater likelihood of developing choroidal neovascularization, and the patient’s CFB PRS as compared to a predetermined value for lower likelihood of developing CNV identifies the patient as having a lower likelihood of developing choroidal neovascularization. See Tables 9 and 10.

[0292] Provided in one aspect is a method for determining the likelihood of developing all AMD in a patient, wherein the patient’s CFH PRS as compared to a predetermined value for greater likelihood of developing all AMD identifies the patient as having a greater likelihood of developing all AMD, and the patient’s CFH PRS as compared to a predetermined value for lower likelihood of developing all AMD identifies the patient as having a lower likelihood of developing AMD. See Table 11.

[0293] Provided in one aspect is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFH PRS as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD, and the patient’s CFH PRS as compared to a predetermined value for lower likelihood of developing CNV identifies the patient as having a lower likelihood of developing dry AMD. See Table 12.

[0294] Provided in one aspect is a method for determining the likelihood of developing choroidal neovascularization in a patient, wherein the patient’s CFH PRS as compared to a predetermined value for greater likelihood of developing CNV identifies the patient as having a greater likelihood of developing choroidal neovascularization, and the patient’s CFH PRS as compared to a predetermined value for lower likelihood of developing CNV identifies the patient as having a lower likelihood of developing choroidal neovascularization. See Table 13.Attorney Docket No.136622-1010

[0295] Provided in one aspect is a method for determining the likelihood of developing all AMD in a patient, wherein the patient’s CFI PRS all AMD as compared to a predetermined value for greater likelihood of developing identifies the patient as having a greater likelihood of developing all AMD, and the patient’s CFI PRS as compared to a predetermined value for lower likelihood of developing all AMD identifies the patient as having a lower likelihood of developing AMD. See Table 14.

[0296] Provided in one aspect is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFI PRS as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD, and the patient’s CFI PRS as compared to a predetermined value for lower of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 15.

[0297] Provided in one aspect is a method for determining the likelihood of developing choroidal neovascularization in a patient, wherein the patient’s CFI PRS as compared to a predetermined value for greater likelihood of developing CNV identifies the patient as having a greater likelihood of developing choroidal neovascularization, and the patient’s CFI PRS as compared to a predetermined value for lower likelihood of developing CNV identifies the patient as having a lower likelihood of developing choroidal neovascularization. See Table 16.

[0298] Also provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s allelic score for C3 and CFH as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s allelic score for C3 and CFH as compared to a predetermined value for lower likelihood of developing CNV identifies the patient as having a lower likelihood of developing dry AMD. See Table 18.

[0299] Also provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s allelic score for C3 and HTRA1 as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s allelic score for C3 and HTRA1 as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 19.Attorney Docket No.136622-1010

[0300] Also provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s allelic score for CFB and CFH as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s allelic score for CFB and CFH as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 20.

[0301] Also provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s allelic score for CFB and HTRA1 as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s allelic score for CFB and HTRA1 as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 21.

[0302] Also provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s allelic score for CFH and HTRA1 as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s allelic score for CFH and HTRA1 as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 22.

[0303] Also provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s allelic score for CFI and CFH as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s allelic score for CFI and CFH as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 23.

[0304] Also provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s allelic score for CFI and HTRA1 as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s allelic score for CFI and HTRA1 as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 24.Attorney Docket No.136622-1010

[0305] Further provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s C3 score and default PRS (all AMD loci) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s C3 score and default PRS (all AMD loci) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 25.

[0306] Further provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s C3 score and complement pathway PRS (including HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s C3 score and complement pathway PRS (including HTRA1) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 26.

[0307] In another aspect, provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s C3 score and complement pathway PRS (excluding CFH and HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s C3 score and complement pathway PRS as compared to a predetermined value for lower likelihood of developing dry AMD (excluding CFH and HTRA1) identifies the patient as having a lower likelihood of developing dry AMD. See Table 27.

[0308] In another aspect, provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s C3 score and lipid metabolism PRS as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and the patient’s C3 score and lipid metabolism as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 28.

[0309] In another aspect, provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s C3 score and extracellular matrix degradation PRS (excluding HTRA1) as compared to a predetermined value for greater likelihood ofAttorney Docket No.136622-1010 developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s C3 score and extracellular matrix degradation PRS (excluding HTRA1) as compared to a predetermined value for lower likelihood of developing identifies dry AMD patient as having a lower likelihood of developing dry AMD. See Table 29.

[0310] In another aspect, provided is a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s C3 score and non-complement pathway PRS as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s C3 score and non-complement pathway PRS as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 30.

[0311] The disclosure also provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFB score and default PRS (all AMD loci) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFB score and default PRS (all AMD loci) as compared to a predetermined value for low likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 32.

[0312] The disclosure also provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFB score and complement pathway PRS (including HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFB score and complement pathway PRS (including HTRA1) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 33.

[0313] The disclosure also provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFB score and complement pathway PRS (excluding CFH and HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFB score and complement pathway PRS (excluding CFH and HTRA1) as compared to a predetermined value for lower likelihood of developing dryAttorney Docket No.136622-1010 AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 34.

[0314] The disclosure further provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFB score and lipid metabolism PRS as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFB score and lipid metabolism PRS as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 35.

[0315] The disclosure further provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFB score and extracellular matrix degradation PRS (excluding HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFB score and extracellular matrix degradation PRS (excluding HTRA1) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 36.

[0316] The disclosure further provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFB score and non-complement PRS (the default PRS with the complement PRS regressed out) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFB score and a non-complement PRS (the default PS with the complement PRS regressed out) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 37.

[0317] The disclosure further provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFB score and default PRS (all AMD loci) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFB score and default PRS (all AMD loci) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 38.Attorney Docket No.136622-1010

[0318] The disclosure further provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFH score and default PRS (all AMD loci) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFH score and default PRS (all AMD loci) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 39.

[0319] The disclosure further provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFH score and complement pathway PRS (including HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFH score and complement pathway PRS (including HTRA1) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 40.

[0320] The disclosure further provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFH score and complement pathway PRS (excluding CFH and HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFH score and complement pathway PRS (excluding CFH and HTRA1) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 41.

[0321] The disclosure further provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFH score and lipid metabolism PRS as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFH score and lipid metabolism PRS as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 42.

[0322] The disclosure further provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFH score and extracellular matrix degradation PRS (that excludes HTRA1) as compared to a predetermined value for greater likelihood ofAttorney Docket No.136622-1010 developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFH score and extracellular matrix degradation PRS as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 43.

[0323] In another aspect, the disclosure provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFH score and non-complement pathway PRS (the default PRS with the complement PRS regressed out) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFH score and non-complement pathway PRS (the default PRS with the complement PRS regressed out) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 44.

[0324] In another aspect, the disclosure provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFI score and default PRS (all AMD loci) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFI score and default PRS (all AMD loci) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 46.

[0325] In another aspect, the disclosure provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFI score and complement pathway PRS (including HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFI score and complement pathway PRS (including HTRA1) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 47.

[0326] In another aspect, the disclosure provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFI score and complement pathway PRS (that excludes CFH and HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFI score and complement pathway PRS (that excludes CFH and HTRA1) as compared to a predetermined value for lower likelihoodAttorney Docket No.136622-1010 of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 48.

[0327] In another aspect, the disclosure provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFI score and lipid metabolism PRS as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFI score and lipid metabolism PRS as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 49.

[0328] In another aspect, the disclosure provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFI score and extracellular matrix PRS (that excludes HTRA1) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFI score and extracellular matrix PRS (that excludes HTRA1) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 50.

[0329] In a further aspect, the disclosure provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFI score and non-complement PRS (the default PRS and the complement PRS regressed out) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFI score and non-complement PRS (the default PRS and the complement PRS regressed out) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 51.

[0330] In a further aspect, the disclosure provides a method for determining the likelihood of developing dry AMD in a patient, wherein the patient’s CFI score and non-complement PRS (the default PRS and the complement PRS regressed out) as compared to a predetermined value for greater likelihood of developing dry AMD identifies the patient as having a greater likelihood of developing dry AMD and wherein the patient’s CFI score and non-complement PRS (the default PRS and the complement PRS regressed out) as compared to a predetermined value for lower likelihood of developing dry AMD identifies the patient as having a lower likelihood of developing dry AMD. See Table 52.Attorney Docket No.136622-1010

[0331] The above methods can further comprise evaluation of a patient’s retinal optical coherence tomography (OCT) images. See Table 56. The OCT-derived phenotype that is most predictive of dry AMD was determined to be the photoreceptor layer thickness. See Tables 57 and 58. In a further aspect, the above methods can further include an evaluation and incorporation of the patient’s age, gender, and disease stage (see Table 60).

[0332] In a further aspect, the method further comprises, or alternatively consists essentially of, or consists of, administration of an HSD17B13 inhibitor, thereby treating the patient.

[0333] Further provided are methods for treating one or more of AMD or CNV in a patient in need thereof wherein the patient is identified for the treatment as at high risk for developing AMD or CNV as respectively determined.

[0334] Suitable patient samples in the methods include, but are not limited to a sample comprises, or alternatively consisting essentially of, or yet further consisting of, at least one of blood, plasma, a blood cell, a peripheral blood lymphocyte, a liver cell, or combinations thereof. The samples can be at least one of an original sample recently isolated from the patient, a fixed tissue, a frozen tissue, a biopsy tissue, a resection tissue, a microdissected tissue, or combinations thereof.

[0335] Any suitable method for identifying the genotype in the patient sample can be used and the disclosures described herein are not to be limited to these methods. For the purpose of illustration only, the genotype is determined by a method comprising, or alternatively consisting essentially of, or yet further consisting of, Sanger sequencing, next-gen sequencing, hybridization, PCR or more specifically, PCR-RFLP or microarray. These methods as well as equivalents or alternatives thereto are described herein.

[0336] The methods are useful in the assistance of an animal, a mammal or yet further a human patient. For the purpose of illustration only, a mammal includes but is not limited to a human, a simian, a murine, a bovine, an equine, a porcine or an ovine. Systems of the Disclosure for Risk Score-Informed Drug Development

[0337] There are many applications of the systems of the present disclosure, ranging from drug discovery using agnostic analysis of new chemical entities in combination with genetic variant distribution in patient populations, to new and efficient designs of clinical trials, to incorporation of new end points in clinical trials, to how the data from former clinical studies are analyzed to provide evidence of effectiveness. Applications of particular value include combining genetic information on patient populations with a better understanding ofAttorney Docket No.136622-1010 pharmacodynamic end points and biomarkers and how they relate to clinical outcomes.

[0338] For example, the systems of the disclosure can be used for in silico clinical trial recapitulation and regulatory evaluation. Historically, safety and efficacy data provided to regulatory agencies in support of marketing has required preclinical and clinical efficacy data using empirical methods. Therapeutic Areas

[0339] The present approach may be configured as a system or method, but also may be provided as a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure. The systems and methods disclosed herein have the superior benefit of reducing memory processing resources.

[0340] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD- ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire. Rather, the computer readable storage medium is a non-transient (i.e., not- volatile) medium.

[0341] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a localAttorney Docket No.136622-1010 area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0342] Computer readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’s computer, partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0343] These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions and / or steps specified in the disclosure. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium havingAttorney Docket No.136622-1010 instructions stored therein comprises, or alternatively consists essentially of, or yet further consists of an article of manufacture including instructions which implement aspects of the functions and / or steps specified in the disclosure.

[0344] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions and / or steps specified in the disclosure. Methods for Determining Drug Activity

[0345] An aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets for the treatment of AMD and CNV, the method comprising, or alternatively consisting essentially of, or yet further consisting of: obtaining molecular biomarker stratifier data comprising, or alternatively consisting essentially of, or yet further consisting of a plurality of biomarker stratifiers and at least one disease phenotype from each of a plurality of subjects; determining a plurality of values representing biomarker stratifier effects, wherein each value separately represents how each of the plurality of biomarker stratifiers affects each of the at least one disease phenotype based on external data; calculating a biomarker stratifier score for a chosen disease phenotype; calculating a pharmacomimetic genetic score for each drug target; and determining the predicted drug activity of each drug target the at least one disease phenotype in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the disease phenotype.

[0346] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof. Examples of such for the diagnosis and treatment of for the treatment of AMD and CNV are provided herein.Attorney Docket No.136622-1010

[0347] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0348] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0349] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity; (v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0350] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somaticAttorney Docket No.136622-1010 mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0351] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

[0352] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0353] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and external data from a plurality of subjects.

[0354] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0355] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0356] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of the weighted principal component analysis identifies predicted drug activity of the plurality of drug targets for the at least one disease phenotype.

[0357] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0358] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of oneAttorney Docket No.136622-1010 or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0359] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0360] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity; (v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0361] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.Attorney Docket No.136622-1010

[0362] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

[0363] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0364] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and the external data from a plurality of subjects.

[0365] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0366] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0367] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of the weighted principal component analysis identifies predicted drug activity of the plurality of drug targets for the at least one disease phenotype.

[0368] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0369] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.Attorney Docket No.136622-1010

[0370] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0371] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity; (v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0372] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0373] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.Attorney Docket No.136622-1010

[0374] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0375] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and external data from a plurality of subjects.

[0376] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0377] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0378] In some embodiments, the polygenic score is a body mass index polygenic score.

[0379] Another aspect of the disclosure is directed to an in silico method for determining drug activity of a plurality of drug targets, the method comprising: obtaining molecular biomarker stratifier data and at least one disease phenotype from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each disease phenotype based on external data; constructing a matrix based on the biomarker stratifiers, disease phenotypes and values; running a principal components analysis biomarker stratifier x phenotype-related phenotype biomarker stratifier effects matrix; calculating a biomarker stratifier score for each principal component; calculating a pharmacomimetic genetic score for each drug target; and identifying predicted drug activity of each drug target on one or more phenotypes in subsets of the biomarker stratifier distribution based on the statistical interaction of the polygenic score calculated for each principal component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0380] Another aspect of the disclosure is directed to an in silico method for determining theAttorney Docket No.136622-1010 drug activity of drug targets on a plurality of phenotype intermediates, comprising: obtaining molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each phenotype intermediate based on external data; calculating a biomarker stratifier score for each phenotype intermediate in the disease process; calculating a pharmacomimetic genetic score for each drug target on each phenotype intermediate; and identifying predicted drug activity of the drug targets on one or more phenotypes in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the phenotype intermediate and the outcome phenotype.

[0381] Another aspect of the disclosure is directed to method of determining the drug activity of a drug target, comprising: obtaining molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each phenotype based on external data; calculating a biomarker stratifier score for a chosen outcome phenotype related to the disease outcome; calculating a pharmacomimetic genetic score for the drug target; and identifying predicted drug activity of the drug target on one or more phenotypes in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for the drug target in association analysis with the outcome phenotype.

[0382] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for aAttorney Docket No.136622-1010 disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0383] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0384] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0385] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity; (v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0386] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomicsAttorney Docket No.136622-1010 score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0387] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

[0388] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

[0389] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and external data from a plurality of subjects.

[0390] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of adjusting the genetic score for each drug target and the biomarker stratifier score for each of the principal components.

[0391] In some embodiments, the weighted principal component analysis is performed according to the biomarker stratifier score of the chosen disease phenotype.

[0392] In some embodiments, the method further comprises, or alternatively consists essentially of, or yet further consists of the weighted principal component analysis identifies predicted drug activity of the plurality of drug targets for the at least one disease phenotype.

[0393] Another aspect of the disclosure is directed to an in silico method comprising: generating a plurality of principal components (PCs) corresponding to genetic ancestry data for subjects in a study cohort; generating a biomarker stratifier score for each subject in the study cohort based at least on (i) the PCs and on (ii) biomarker stratifier weights for a disease of interest; determining which of a plurality of disease-associated variants are pharmacomimeticAttorney Docket No.136622-1010 instruments for the disease of interest, where a disease-associated variant is a pharmacomimetic instrument if the disease-associated variant modulates a function or expression of a target gene of a drug such that a first effect of the disease-associated variant on a phenotype is likely to be predictive of a second effect of the drug on the phenotype; determining statistical interactions between the pharmacomimetic instruments for one or more drug targets and the biomarker stratifier scores, wherein the statistical interactions are predictive of drug target-specific differential treatment response for the disease of interest; and performing one or more prediction-based actions based on the determined interactions.

[0394] In some embodiments, the one or more prediction-based actions comprises at least one of (i) therapeutic development, (ii) therapeutic target identification, or (iii) pharmacogenomics.

[0395] In some embodiments, the genetic ancestry data indicates whether subjects in the study cohort are a case or a control for the disease of interest.

[0396] In some embodiments, the genetic ancestry data is based on genotyping arrays or whole-genome sequencing.

[0397] In some embodiments, the genetic ancestry data is obtained from a publicly-available database.

[0398] In some embodiments, the plurality of PCs comprises at least 5 PCs.

[0399] In some embodiments, the study cohort comprises at least 200 control subjects.

[0400] In some embodiments, each disease-associated variant in the plurality of disease- associated variants satisfies a genome-wide significance threshold.

[0401] In some embodiments, the plurality of disease-associated variants is determined based on a first disease genome-wide association study (GWAS).

[0402] In some embodiments, the biomarker stratifier score for each subject is a scaled biomarker stratifier score.

[0403] In some embodiments, the scaled biomarker stratifier score is based at least on a raw PGS and an ancestry-normalized PGS.

[0404] In some embodiments, the PGS variant weights are computed independent of theAttorney Docket No.136622-1010 genetic ancestry data corresponding to the study cohort.

[0405] In some embodiments, determining which of the plurality of disease-associated variants are pharmacomimetic instruments for the disease of interest comprises generating an allelic score.

[0406] In some embodiments, variants in the allelic score are weighted based at least on a second GWAS that did not include any of the subjects in the study cohort.

[0407] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0408] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0409] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0410] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity; (v) a somatic mutation score for a drug target expression; or any linear or non-linearAttorney Docket No.136622-1010 combination thereof.

[0411] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof. Systems

[0412] Another aspect of the disclosure is directed to an in silico system for drug development, such system comprising: at least one hardware processor; and a non-transitory computer-readable storage medium having stored thereon program code, the program code executable by the at least one hardware processor to: obtain molecular biomarker stratifier data and disease phenotypes from each of a plurality of subjects; determine the value of the biomarker stratifier effects for each phenotype based on external data; construct a matrix based on the biomarker stratifiers, phenotypes and values; run a principal components analysis biomarker stratifier x phenotype-related phenotype biomarker stratifier effects matrix; calculate a biomarker stratifier score for each principal component; calculate a pharmacomimetic genetic score for each drug target;Attorney Docket No.136622-1010 identify predicted drug effects on one or more phenotypes in subsets of the biomarker stratifier distribution based on the statistical interaction of the polygenic score calculated for each principal component with the pharmacomimetic genetic score for each drug target in association analysis with the outcome phenotype.

[0413] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a genetic variant and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

[0414] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a proteomic variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

[0415] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a transcriptional variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

[0416] In some embodiments, the biomarker stratifier comprises, or alternatively consists essentially of, or yet further consists of a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a somatic mutation score for a drug target expression; (ii) a somatic mutation score for a disease phenotype; (iii) a somatic mutation score for a disease risk factor; (iv) a somatic mutation score for a biological pathway activity; (v) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof.

[0417] In some embodiments, the biomarker stratifier comprises, or alternatively consistsAttorney Docket No.136622-1010 essentially of, or yet further consists of a genetic variant, a proteomic variant, a transcriptional variant and / or a somatic mutational variant, and wherein the biomarker stratifier score comprises, or alternatively consists essentially of, or yet further consists of one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a somatic mutation score for a drug target expression; (xiv) a somatic mutation score for a disease phenotype; (xv) a somatic mutation score for a disease risk factor; (xvi) a somatic mutation score for a biological pathway activity; (xvii) a somatic mutation score for a drug target expression; or any linear or non-linear combination thereof. EXAMPLES

[0418] The following examples are included for illustrative purposes only and are not intended to limit the scope of the disclosure. Example 1: Analysis of European Ancestry Subjects from UK Biobank

[0419] For each of four AMD drug targets in the complement pathway (C3, CFB, CFH, CFI) an allelic score was generated that aggregated the effects of two or more conditionally independent variants in or near the target gene. Each score was defined as a weighted sum of an individual’s dosage for included variants (examples below). A variant’s weight was set equal to its regression coefficients in a joint model of the effect of 52 variants on AMD risk, published in Supplementary Table 4 of Fritsche et al. (2016) Nature Genetics Vol.41:688- 695. The Fritsche et al. publication did not have any variants for CFD or C5, two additional drug targets.

[0420] The variants included in each allelic score were: • C3 : rs2230199 (R102G,) and rs147859257 (K155Q); • CFB: rs429608, rs2746394, rs204993 and rs181705462; • CFH : rs10922109, rs570618 (also known in the literature as Y402H), rs148553336, rs187328863, 196815450L, and rs35292876; and • CFI : rs10033900 and rs141853578 (G119R).Attorney Docket No.136622-1010

[0421] The drugs relevant to the complement pathway drug targets include, but are not limited to: • pegcetacoplan / Syfovre, an FDA-approved, intravitreally-delivered, peptide inhibitor of C3 developed by Apellis Pharmaceuticals. • AMY-106, an intravitreally-delivered peptide inhibitor of C3 being developed by Amyndas Pharmaceuticals. • iptacopan, an orally-delivered small molecule inhibitor of CFB being developed by Novartis. • IONIS-FB-LRx (also known as RG6299), a subcutaneously-delivered antisense oligonucleotide for CFB being developed by Ionis and Roche / Genentech. • GEM103, a recombinant CFH protein developed by Gemini Therapeutics. failed in clinical trials. • GT005 (also known as PPY998), an intravitreally-delivered gene therapy to increase expression of CFI being developed by Gyroscope Therapeutics, which was acquired by Novartis. Polygenic risk score

[0422] An AMD polygenic risk score (PRS) was generated using an approach similar to that used for the drug target allelic scores, but also included variants at all AMD GWAS loci reported in Fritsche et al. (2016) rather than one locus at a time. Versions of the PRS were also constructed that excluded each of the drug target genes, i.e. all genes except C3, all genes except CFB, etc.

[0423] 4 out of 52 variants reported in the Fritsche et al. (2016) study were excluded from this reported PRS: • rs3138141, rs121913059, and rs142450006 are not present in the UK Biobank genotypes file (after QC). • rs191281603 is not even nominally significant in the Fritsche et al. (2016) unconditioned analysis and is sub-genome-wide significant in the fully conditioned analysis.

[0424] In total, the PRS included 48 SNPs. Complement pathway-specific PRSAttorney Docket No.136622-1010

[0425] A parsimonious AMD PRS was created using only loci where the candidate gene reported by Fritsche et al. (2016) is in the complement pathway: • the complement components C3 and C9 • the complement regulatory factors CFB, CFH, and CFI • VTN (vitronectin), an extracellular matrix protein that inhibits the formation of the terminal complement complex (Milis et al.1993, Sheehan et al.1995, Schvartz et al.1999). • ARMS2 / HTRA1, two nearby genes that are both candidate genes for the same AMD GWAS locus. ARMS2 is a poorly-characterized protein that has been reported to regulate complement-mediated opsonization in the retina (Micklisch et al.2017). HTRA1 is serine protease that degrades extracellular matrix proteins, including complement-regulating ECM proteins such as vitronectin and clusterin (Tschopp et al.1993). Regardless of which is the causal gene, the lead SNP at this GWAS locus correlates with increased complement activation (Smailhodzic et al.2012).

[0426] The annotation of the Fritsche et al. (2016) GWAS candidate genes as part of / not part of the complement pathway was done by manual literature review. Post-hoc, the selected annotations to the HGNC gene groups “Complement system activation components” and “Complement system regulators and receptors“ were compared.. All of the genes in these HGNC gene groups that are also Fritsche et al. (2016) GWAS candidate genes were present in the complement gene list.

[0427] All of the genes in the complement gene list were present in those HGNC gene groups, except for ARMS2 / HTRA1. As detailed below, it was shown that the ARMS2 / HTRA1 GWAS locus lead SNP, by itself, amplifies the effect of C3 pharmacomimetic variants, justifying its inclusion post-hoc. The history of the Github repository that contains the code for this disclosure and confirmed that the decision to include ARMS2 / HTRA1 in the complement pathway PRS was made prior to doing the single gene-gene interaction test.

[0428] The lead variant at the PILRA / PILRB locus was originally included in the complement PRS, but subsequently removed. As far as is known, only Rathore et al. (2018) identified complement component C4A as a non-exclusive ligand of PILRA.

[0429] The contribution of PILRA / PILRB rs7803454 to the complement PRS, when it was included, was tiny compared to genes like CFH and HTRA1. In Fritsche et al. (2016), theAttorney Docket No.136622-1010 GWAS p-value for PILRA / PILRB rs7803454 was 5e-9, compared to 1e-617 for the lead CFH SNP and 6e-735 for the lead ARMS2 / HTRA1 SNP.

[0430] The final complement-pathway PRS has a total of 18 SNPs. Phenotypes

[0431] AMD (“Codes” refer to diagnosis and reimbursement codes): • Cases must have ICD-10 code H35.3, ICD-9 code 362.5, or UK Biobank self- report code 1528. Cases must not have ICD-10 code H36.0 or ICD-9 code 362.0 (diabetic retinopathy). • Controls must not be cases, and additionally must not have ICD-10 codes E10.3, E11.3, E12.3, E13.3, E14.3 (diabetic eye disease), H35 (macular degeneration), or H36.0 (diabetic retinopathy); ICD-9 code 362 (other retinal disorders); UK Biobank self-report code 1276 (diabetic eye disease); or OPCS-4 codes C79-C85 (eye surgery) or X93 (high-cost ophthalmology drugs).

[0432] Choroidal neovascularization (“Codes” refer to diagnosis and reimbursement codes): • Cases must be AMD cases and have OPCS-4 code X931 (drugs for CNV) or C82 (laser photocoagulation and related procedures which are used for CNV but not GA). • Controls must be AMD controls.

[0433] Dry AMD: • Cases must be AMD cases and not CNV cases. • Controls must be AMD controls.

[0434] Diabetic retinopathy (“Codes” refer to diagnosis and reimbursement codes): • Cases must have ICD-10 code H36.0 or ICD-9 code 362.0 (the specific codes for diabetic retinopathy), or they must have a code for diabetes in general (ICD-10 E10-E14, ICD-9250, self-report 1220-1223, medication codes listed below) AND EITHER (a code for diabetic eye disease in general (ICD-10 E1*3, ICD-9250.4, self-report 1276) or a code for a DR-related surgical procedure (OPCS-4 C80, C82-C84, X93)). • Cases and controls must not have AMD. • If a case does not have specific code for diabetic retinopathy (ICD-10 H36.0,Attorney Docket No.136622-1010 ICD-9362.0), they must not have a code for diabetic cataract (ICD-10 H28.0). If they do, the subject will be excluded from both the cases and the controls.

[0435] The UK Biobank medication codes used in the diabetes phenotype definition described above were: 1140883066, 1140884600, 1140874686, 1141171646, 1141171652, 1141177600, 1141189090, 1141189094, 1141177606, 1141153254, 1141153262, 1140874744, 1141152590, 1141156984, 1140874718, 1140874736, 1140874724, 1140874726, 1140874728, 1140874740, 1140910566, 1140874746, 1140874646, 1141157284, 1140874652, 1140874674, 1140874690, 1140874706, 1140874712, 1140874716, 1140874664, 1140874666, 1141168660, 1141168668, 1141173882, 1141173786, 1140868902, and 1140868908. Statistical models

[0436] Tests were performed for drug-target-score interactions with leave-one-gene-out AMD PRS in the UK Biobank. Briefly, the PRS was performed excluding the effects of individual genes A simple logistic regression model was used with a binary AMD phenotype as the outcome and the following covariates: age, age squared, sex, genotyping chip, and the first through twentieth principal components of genetic ancestry. Additionally, PCs 1-20 were regressed out of the PRS prior to fitting the main regression model.

[0437] The drug-target-score x PRS interaction tests were repeated using several refinements of the AMD phenotypes.

[0438] Table 1: Cohort statistics (unrelated European-ancestry subjects only). Dry AMD (GA or Statistic CNV intermediate AMD) Control Neither case nor control # subjects 736 6,813 332,874 12,943% of cohort 0.2% 1.9% 94.2% 3.7%% female 53% 63% 54% 44%Median age (5th, 77 (65, 78 (67, 84) 72 (57, 82) 75 (59, 83) 95th percentiles) 83) AMD case-control analysis using the genome-wide AMD PRS

[0439] C3, CFB, CFH, and CFI scores all have strong interactions with the AMD PRS in determining risk of any AMD. These interactions remain significant if wet AMD cases are excluded.

[0440] Table 2: Interaction between the C3 score and the leave-C3-out PRS in determining the risk of all AMD.Attorney Docket No.136622-1010 Term OR P-value (raw) P-value (GC) C3 score 0.81 (0.76, 0.87) 6.0e-10 2.1e-08 PRS 1.37 (1.33, 1.41) 5.3e-1013.9e-83 C3 score : PRS 0.79 (0.75, 0.84) 2.9e-14 5.8e-12 interaction

[0441] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The C3 score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al. (2016)). The units of the PRS are standard deviations.

[0442] Table 3: Interaction between the C3 score and the leave-C3-out PRS in determining the risk of dry AMD. Term OR P-value P-value (GC) (raw) C3 score 0.80 (0.75, 5.7e-10 2.0e-08 0.86) PRS 1.37 (1.33, 1.41) 4.0e-92 7.4e-76 C3 score : PRS interaction 0.80 (0.75, 0.85) 8.5e-12 6.3e-10

[0443] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The C3 score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0444] Table 4: Interaction between the C3 score and the leave-C3-out PRS in determining the risk of choroidal neovascularization. Term OR P-value (raw) P-value (GC) C3 score 0.89 (0.72, 0.317 0.365 1.11) PRS 1.37 (1.25, 1.49) 1.1e-11 7.9e-10 C3 score : PRS 0.69 (0.58, 0.83) 5.8e-05 2.7e-04 interaction

[0445] . P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The C3 score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0446] Table 5: Interaction between the CFB score and the leave-CFB-out PRS in determining the risk of all AMD.Attorney Docket No.136622-1010

[0447] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFB score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0448] Table 6: Association of the CFB score with all AMD, stratified by quantiles of the leave-CFB-out PRS.60-79 0.85 (0.76, 0.94) 2.2e-03 5.5e-03 1,536 66,569 80-100 0.72 (0.66, 0.78) 9.7e-14 1.6e-11 2,540 65,377

[0449] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFB score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016).

[0450] Table 7: Interaction between the CFB score and the leave-CFB-out PRS in determining the risk of dry AMD.

[0451] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFB score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0452] Table 8: Association of the CFB score with dry AMD, stratified by quantiles of the leave-CFB-out PRS.Attorney Docket No.136622-1010 40-59 0.94 (0.84, 1.06) 0.30360-79 0.82 (0.74, 0.92) 5.4e-04 1.7e-03 1,391 66,569 80-100 0.71 (0.65, 0.78) 3.1e-13 4.1e-11 2,280 65,377

[0453] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFB score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016).

[0454] Table 9: Interaction between the CFB score and the leave-CFB-out PRS in determining the risk of choroidal neovascularization.

[0455] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFB score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0456] Table 10: Association of the CFB score with choroidal neovascularization, stratified by quantiles of the leave-CFB-out PRS.

[0457] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFB score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016).

[0458] Table 11: Interaction between the CFH score and the leave-CFH-out PRS in determining the risk of all AMDAttorney Docket No.136622-1010 interaction

[0459] . P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFH score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0460] Table 12: Interaction between the CFH score and the leave-CFH-out PRS in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) CFH score 0.82 (0.80, 0.84) 2.3e-66 9.9e-55 PRS 1.61 (1.54, 1.69) 6.4e-94 2.5e-77 CFH score : PRS interaction 0.91 (0.89, 0.93) 2.4e-18 2.6e-15

[0461] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFH score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0462] Table 13: Interaction between the CFH score and the leave-CFH-out PRS in determining the risk of choroidal neovascularization. Term OR P-value (raw) P-value (GC) CFH score 0.84 (0.78, 0.90) 1.4e-06 1.3e-05 PRS 1.84 (1.61, 2.10) 1.7e-19 3.0e-16 CFH score : PRS interaction 0.87 (0.82, 0.93) 1.5e-05 8.7e-05

[0463] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFH score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0464] Table 14: Interaction between the CFI score and the leave-CFI-out PRS in determining the risk of all AMD.

[0465] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFI score isAttorney Docket No.136622-1010 scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0466] Table 15: Interaction between the CFI score and the leave-CFI-out PRS in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) CFI score 0.77 (0.67, 0.88) 1.2e-04 5.0e-04 PRS 1.52 (1.47, 1.58) 5.1e-133 2.2e-109 CFI score : PRS interaction 0.86 (0.76, 0.97) 0.013 0.024

[0467] P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFI score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations.

[0468] Table 16: Interaction between the CFI score and the leave-CFI-out PRS in determining the risk of choroidal neovascularization. Term OR P-value (raw) P-value (GC) CFI score 0.58 (0.41, 0.82) 2.3e-03 5.8e-03 PRS 1.54 (1.41, 1.69) 1.2e-20 3.3e-17 CFI score : PRS interaction 0.94 (0.69, 1.28) 0.707 0.734 P-value (GC) refers to a p-value that is corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome (see Methods). The CFI score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016). The units of the PRS are standard deviations. AMD case-control analysis: other variants from Fritsche et al.2016

[0469] Individual AMD-associated variants from Fritsche et al.2016 were tested for interaction with the AMD PRS, other than those already tested above as part of the C3, CFB, CFH, and CFI complement pathway loci.

[0470] Table 17: Interaction between individual AMD variants reported in Fritsche et al.2016 with the AMD PRS in determining the risk of AMD. Variant Marginal OR Marginal Interaction OR Interaction P-value P-value HTRA1_rs3750846 1.44 (1.38, 1.50) 4.6e-63 1.16 (1.12, 1.21) 5.2e-13 SLC16A8_rs8135665 0.98 (0.94, 1.02) 0.427 1.08 (1.04, 1.12) 1.0e-03 LIPC_rs2070895 0.97 (0.93, 1.02) 0.270 0.93 (0.90, 0.97) 2.3e-03 ABCA1_rs2740488 0.99 (0.95, 1.03) 0.619 0.95 (0.91, 0.98) 0.012Attorney Docket No.136622-1010 Variant Marginal OR Marginal Interaction OR Interaction P-value P-value C9_rs62358361 1.29 (1.10, 1.51) 4.2e-03 1.19 (1.03, 1.37) 0.030 TIMP3_rs5754227 0.94 (0.89, 0.99) 0.039 0.95 (0.90, 0.99) 0.046 COL8A1_rs140647181 1.00 (0.88, 1.15) 0.959 1.11 (0.98, 1.26) 0.130 LIPC_rs2043085 1.00 (0.96, 1.03) 0.913 1.03 (0.99, 1.06) 0.138 C20orf85_rs201459901 1.00 (0.93, 1.07) 0.953 0.94 (0.88, 1.01) 0.138 VEGFA_rs943080 1.03 (0.99, 1.06) 0.153 1.03 (0.99, 1.06) 0.174 RORB_rs10781182 1.05 (1.01, 1.09) 0.039 0.98 (0.95, 1.01) 0.300 COL8A1_rs55975637 1.02 (0.96, 1.07) 0.585 1.03 (0.98, 1.08) 0.316 TNFRSF10A_rs13278062 1.05 (1.02, 1.09) 0.011 1.02 (0.99, 1.05) 0.322 COL4A3_rs11884770 1.02 (0.98, 1.06) 0.341 1.02 (0.98, 1.06) 0.325 CNN2_rs67538026 0.99 (0.96, 1.03) 0.621 1.02 (0.98, 1.05) 0.340 PILRA / PILRB_rs7803454 1.02 (0.98, 1.07) 0.417 1.02 (0.98, 1.06) 0.386 KMT2E_rs1142 1.01 (0.98, 1.05) 0.491 1.02 (0.98, 1.05) 0.412 RAD51B_rs61985136 1.02 (0.98, 1.06) 0.341 1.01 (0.98, 1.05) 0.467 CETP_rs17231506 1.08 (1.04, 1.12) 1.7e-04 1.01 (0.98, 1.05) 0.502 TGFBR1_rs1626340 0.95 (0.91, 1.00) 0.050 0.99 (0.95, 1.03) 0.512 APOE_rs429358 0.89 (0.84, 0.93) 2.2e-05 0.98 (0.94, 1.03) 0.538 TSPAN10_rs6565597 1.00 (0.97, 1.04) 0.947 0.99 (0.95, 1.02) 0.543 TRPM3_rs71507014 0.97 (0.94, 1.01) 0.180 1.01 (0.98, 1.04) 0.586 CETP_rs5817082 0.95 (0.91, 0.99) 0.019 0.99 (0.95, 1.03) 0.608 APOE_rs73036519 1.02 (0.98, 1.06) 0.469 0.99 (0.96, 1.03) 0.640 ADAMTS9_rs62247658 0.98 (0.94, 1.01) 0.231 0.99 (0.96, 1.02) 0.646 B3GALTL_rs9564692 0.98 (0.94, 1.02) 0.298 0.99 (0.96, 1.03) 0.710 ACAD10_rs61941274 1.02 (0.91, 1.14) 0.805 1.02 (0.92, 1.14) 0.712 RAD51B_rs2842339 0.99 (0.93, 1.05) 0.756 1.01 (0.96, 1.07) 0.742 VTN_rs11080055 1.05 (1.02, 1.09) 8.0e-03 1.00 (0.97, 1.04) 0.797 ARHGAP21_rs12357257 1.02 (0.98, 1.07) 0.370 1.00 (0.96, 1.04) 0.836 PRLR_rs114092250 0.91 (0.82, 1.02) 0.138 0.99 (0.90, 1.10) 0.908 CTRB1 / CTRB2_rs72802340 2.92 (0.86, 0.98) 0.029 1.00 (0.94, 1.07) 0.997

[0471] The PRS was adjusted to remove the effect of the variant being tested. The units of the PRS are standard deviations. The units of the term for the individual variant are number of alleles. Interaction p-values are corrected for global test statistic inflation in variant x PRS interaction tests using AMD as the outcome. AMD case-control analysis: single gene-gene interactions

[0472] The CFH and HTRA1 loci have extremely strong associations with AMD, and are sufficiently powered to examine interactions of allelic scores constructed from those genes specifically, with allelic scores (drug-target-scores) for the other AMD drug targets (C3, CFB, CFI ). Significant C3 -CFH, CFB-CFH, CFI -CFH, and C3 -HTRA1 interactions were observed. CFB and CFI do not appear to interact with HTRA1.

[0473] Table 18: Interaction between single-gene AMD allelic scores for C3 and CFH inAttorney Docket No.136622-1010 determining the risk of dry AMD.

[0474] The C3 score is scaled so that 1 unit corresponds to 2-fold decreased AMD risk in the Fritsche et al. training dataset, mimicking a drug. The CFH score is scaled so that 1 unit corresponds to 2-fold increased risk in the Fritsche et al. training dataset.

[0475] Table 19: Interaction between single-gene AMD allelic scores for C3 and HTRA1 in determining the risk of dry AMD Term OR P-value (raw) P-value (GC) C3 score 0.84 (0.77, 0.92) 9.9e-05 4.2e-04 HTRA1 score 1.23 (1.19, 1.27) 1.6e-38 6.9e-32 C3 score : HTRA1 score 0.86 (0.81, 0.92) 5.1e-06 3.6e-05

[0476] The C3 score is scaled so that 1 unit corresponds to 2-fold decreased AMD risk in the Fritsche et al. training dataset, mimicking a drug. The HTRA1 score is scaled so that 1 unit corresponds to 2-fold increased risk in the Fritsche et al. training dataset.

[0477] Table 20: Interaction between single-gene AMD allelic scores for CFB and CFH in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) CFB score 0.72 (0.65, 0.79) 4.1e-11 2.3e-09 CFH score 1.27 (1.24, 1.30) 1.2e-91 1.9e-75 CFB score : CFH score 0.93 (0.89, 0.97) 1.3e-03 3.7e-03

[0478] The CFB score is scaled so that 1 unit corresponds to 2-fold decreased AMD risk in the Fritsche et al. training dataset, mimicking a drug. The CFH score is scaled so that 1 unit corresponds to 2-fold increased risk in the Fritsche et al. training dataset

[0479] Table 21: Interaction between single-gene AMD allelic scores for CFB and HTRA1 in determining the risk of dry AMD. The CFB score is scaled so that 1 unit corresponds to 2-fold decreased AMD risk in the Fritsche et al. training dataset, mimicking a drug. The HTRA1 score is scaled so that 1 unit corresponds to 2-fold increased risk in the Fritsche et al. training dataset. Term OR P-value (raw) P-value (GC)Attorney Docket No.136622-1010 CFB score 0.83 (0.78, 0.89) 7.7e-08 1.1e-06 HTRA1 score 1.29 (1.26, 1.32) 1.7e-87 4.7e-72 CFB score : HTRA1 score 0.99 (0.94, 1.04) 0.571 0.608

[0480] Table 22: Interaction between single-gene AMD allelic scores for CFH and HTRA1 in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) CFH score 0.84 (0.82, 0.87) 1.4e-29 1.6e-24 HTRA1 score 1.46 (1.39, 1.53) 1.1e-54 3.9e-45 CFH score : HTRA1 score 0.93 (0.91, 0.96) 4.5e-09 1.1e-07

[0481] The CFH score is scaled so that 1 unit corresponds to 2-fold decreased AMD risk in the Fritsche et al. training dataset, mimicking a drug. The HTRA1 score is scaled so that 1 unit corresponds to 2-fold increased risk in the Fritsche et al. training dataset.

[0482] Table 23: Interaction between single-gene AMD allelic scores for CFI and CFH in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) CFI score 0.54 (0.43, 0.68) 9.2e-08 1.3e-06 CFH score 1.30 (1.26, 1.34) 5.8e-58 7.9e-48 CFI score : CFH score 0.86 (0.76, 0.96) 7.8e-03 0.016

[0483] The CFI score is scaled so that 1 unit corresponds to 2-fold decreased AMD risk in the Fritsche et al. training dataset, mimicking a drug. The CFH score is scaled so that 1 unit corresponds to 2-fold increased risk in the Fritsche et al. training dataset, similar to how PRS was used in the previous section.

[0484] Table 24: Interaction between single-gene AMD allelic scores for CFI and HTRA1 in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) CFI score 0.70 (0.59, 0.82) 1.5e-05 8.9e-05 HTRA1 score 1.28 (1.24, 1.33) 3.7e-45 2.5e-37 CFI score : HTRA1 score 1.02 (0.90, 1.15) 0.780 0.800

[0485] The CFI score is scaled so that 1 unit corresponds to 2-fold decreased AMD risk in the Fritsche et al. training dataset, mimicking a drug. The HTRA1 score is scaled so that 1 unit corresponds to 2-fold increased risk in the Fritsche et al. training dataset AMD case-control analysis: pathway specific PRS

[0486] A parsimonious AMD PRS was created using only loci where the candidate gene reported by Fritsche et al. is in the complement pathway.Attorney Docket No.136622-1010

[0487] The complement PRS as-strong or nearly-as-strong interactions with each of the drug- target-scores in determining risk of dry AMD compared to the full PRS. As before, variants at the drug targets’ locus were excluded from the complement PRS before running each interaction test.

[0488] The full complement PRS (not excluding any of the drug target loci) has a total of 18 SNPs.

[0489] The CFH and HTRA1 loci have very strong associations with AMD, much stronger than any other single locus. To test whether the drug-target-score x complement PRS interaction was a general phenomenon not specific to these high-leverage loci, the above analysis was repeated excluding CFH and HTRA1 from the complement PRS. The interactions remained nominally significant.

[0490] No evidence of interaction between the drug-target-scores and PRS were observed for two other AMD-related pathways, lipid metabolism (included loci = ABCA1, APOE, CETP, LIPC ) and extracellular matrix degradation (excluding HTRA1 ; included loci = ADAMTS9, CTRB1 / CTRB2, TIMP3; the MMP9 locus was excluded because the lead variant in Fritsche et al. is not present in the UK Biobank genotypes file). An exception was the CFB score, which had a nominally-significant interaction with the lipid-metabolism PRS.

[0491] Table 25: Interaction between the C3 score and the default PRS (all AMD loci) in determining the risk of dry AMD.

[0492] 1 unit of the C3 score corresponds to a 2-fold decreased risk of AMD in the Fritsche et al. training dataset. The units of the PRS are standard deviations. The effect of the C3 score was regressed out from the PRS prior to performing the interaction test. Term OR P-value (raw) P-value (GC) C3 score 0.80 (0.75, 0.86) 6.9e-10 2.3e-08 PRS 1.37 (1.33, 1.41) 3.8e-92 7.1e-76 C3 score : PRS 0.80 (0.75, 0.85) 8.6e-12 6.3e-10

[0493] Table 26: Interaction between the C3 score and a complement pathway PRS (including HTRA1) in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) C3 score 0.81 (0.75, 0.86) 7.8e-10 2.6e-08 PRS 1.35 (1.31, 1.40) 1.1e-85 1.4e-70 C3 score : PRS 0.80 (0.75, 0.85) 1.9e-12 1.8e-10Attorney Docket No.136622-1010

[0494] 1 unit of the C3 score corresponds to a 2-fold decreased risk of AMD in the Fritsche et al. training dataset. The units of the PRS are standard deviations. The effect of the C3 score was regressed out from the PRS prior to performing the interaction test.

[0495] Table 27: Interaction between the C3 score and a complement pathway PRS (that excludes CFH and HTRA1 to evaluate whether interactions are a pan-complement gene phenomenon or are specific to these high-leverage AMD risk genes) in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) C3 score 0.75 (0.71, 0.80) 6.3e-18 5.7e-15 PRS 1.11 (1.08, 1.15) 4.0e-11 2.2e-09 C3 score : PRS 0.90 (0.85, 0.96) 2.2e-03 5.5e-03

[0496] 1 unit of the C3 score corresponds to a 2-fold decreased risk of AMD in the Fritsche et al. training dataset. The units of the PRS are standard deviations. The effect of the C3 score was regressed out from the PRS prior to performing the interaction test.

[0497] Table 28: Interaction between the C3 score and a lipid metabolism PRS in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) C3 score 0.75 (0.70, 0.80) 2.0e-19 3.3e-16 PRS 1.07 (1.04, 1.11) 5.9e-06 4.1e-05 C3 score : PRS 0.95 (0.89, 1.02) 0.157 0.201

[0498] 1 unit of the C3 score corresponds to a 2-fold decreased risk of AMD in the Fritsche et al. training dataset. The units of the PRS are standard deviations. The effect of the C3 score was regressed out from the PRS prior to performing the interaction test.

[0499] Table 29: Interaction between the C3 score and an extracellular matrix degradation PRS (that excludes HTRA1 to evaluate whether interactions are a pan- ECM-degrader gene phenomenon or are specific to HTRA1) in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) C3 score 0.74 (0.70, 0.79) 3.0e-20 7.0e-17 PRS 1.05 (1.02, 1.08) 1.8e-03 4.7e-03 C3 score : PRS 1.00 (0.94, 1.07) 0.918 0.926

[0500] 1 unit of the C3 score corresponds to a 2-fold decreased risk of AMD in the Fritsche et al. training dataset. The units of the PRS are standard deviations. The effect of the C3 score was regressed out from the PRS prior to performing the interaction test.Attorney Docket No.136622-1010

[0501] Table 30: Interaction between the C3 score and a non-complement pathway PRS (the default PRS with the complement PRS regressed out) in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) C3 score 0.75 (0.70, 0.79) 1.3e-19 2.4e-16 PRS 1.11 (1.07, 1.14) 1.8e-10 7.9e-09 C3 score : PRS 0.97 (0.91, 1.04) 0.429 0.474

[0502] 1 unit of the C3 score corresponds to a 2-fold decreased risk of AMD in the Fritsche et al. training dataset. The units of the PRS are standard deviations. The effect of the C3 score was regressed out from the PRS prior to performing the interaction test.

[0503] Table 31: Association of the C3 score with dry AMD, stratified by quantiles of the complement PRS (excluding C3). PRS percentile OR P-value (raw) P-value (GC) # cases # controls 0-19 0.94 (0.79, 1.13) 0.532 0.572 919 67,045 20-39 0.86 (0.73, 1.01) 0.062 0.091 1,075 66,983 40-59 0.92 (0.79, 1.08) 0.320 0.368 1,192 66,803 60-79 0.69 (0.60, 0.79) 2.1e-07 2.6e-06 1,377 66,590 80-100 0.58 (0.52, 0.64) 1.8e-23 1.6e-19 2,250 65,453

[0504] The C3 score is scaled such that 1 unit corresponds to 2-fold decreased AMD risk in the training GWAS (Fritsche et al.2016).

[0505] Table 32: Interaction between the CFB score and the default PRS (all AMD loci) in determining the risk of dry AMD. Term OR P-value (raw) P-value (GC) CFB score 0.85 (0.81, 0.89) 1.0e-09 3.2e-08 PRS 1.48 (1.44, 1.52) 8.4e-218 6.7e-179 CFB score : PRS 0.90 (0.86, 0.95) 4.8e-05 2.3e-04

[0506] 1 unit of the CFB score corresponds to a 2-fold decreased risk of AMD in the Fritsche et al. training dataset. The units of the PRS are standard deviations. The effect of the CFB score was regressed out from the PRS prior to performing the interaction test.

[0507] Table 33: Interaction between the CFB score and a complement pathway PRS (including HTRA1) in determining the risk of dry AMD. Term OR P-va...

Claims

Attorney Docket No.136622-1010 WHAT IS CLAIMED IS:

1. An in silico method for determining a predicted drug activity of a plurality of drug targets associated with at least one disease phenotype associated with age-related macular degeneration (AMD), the method comprising: obtaining molecular biomarker stratifier data comprising a plurality of biomarker stratifiers and at least one disease phenotype associated with AMD from each of a plurality of subjects; determining a plurality of values representing one or more biomarker stratifier effects, wherein each value of the plurality separately represents how each of the plurality of biomarker stratifiers affects a disease phenotype associated with AMD based on external data; calculating a biomarker stratifier score for a chosen disease phenotype; calculating a pharmacomimetic genetic score for each drug target; and determining the predicted drug activity of each drug target for at least one disease phenotype associated with AMD in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the at least one disease phenotype associated with AMD.

2. The method of claim 1, wherein the molecular biomarker stratifier data comprises a genetic variant and wherein the biomarker stratifier score comprises one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity; (iv) a polygenic score for a drug target expression, or any linear or non-linear combination thereof.

3. The method of claim 1 or 2, wherein the biomarker stratifier data comprises a proteomic variant, and wherein the biomarker stratifier score comprises one or more score selected from:Attorney Docket No.136622-1010 (i) a proteomics score for a disease phenotype; (ii) a proteomics score for a disease risk factor; (iii) a proteomics score for a biological pathway activity; (iv) a proteomics score for a drug target expression; or any linear or non-linear combination thereof.

4. The method of claim 1 or 2, wherein the biomarker stratifier data comprises a transcriptional variant, and wherein the biomarker stratifier score comprises one or more score selected from: (i) a transcriptomics score for a disease phenotype; (ii) a transcriptomics score for a disease risk factor; (iii) a transcriptomics score for a biological pathway activity; or any linear or non-linear combination thereof.

5. The method of claim 4, wherein the biomarker stratifier data comprises a polymorphic variant, and wherein the biomarker stratifier score comprises one or more score selected from: (i) a polymorphism score for a drug target expression; (ii) a polymorphism score for a disease phenotype; (iii) a polymorphism score for a disease risk factor; (iv) a polymorphism score for a biological pathway activity; (v) a polymorphism score for a drug target expression; or any linear or non-linear combination thereof.

6. The method of claim 1, wherein the biomarker stratifier data comprises a polymorphic variant, wherein the biomarker stratifier data comprises a genetic variant, a proteomic variant, a transcriptional variant and / or a polymorphic variant, and wherein the biomarker stratifier score comprises one or more score selected from: (i) a polygenic score for a disease phenotype; (ii) a polygenic score for a disease risk factor; (iii) a polygenic score for a biological pathway activity;Attorney Docket No.136622-1010 (iv) a polygenic score for a drug target expression; (v) a proteomics score for a disease phenotype; (vi) a proteomics score for a disease risk factor; (vii) a proteomics score for a biological pathway activity; (viii) a proteomics score for a biological pathway activity; (ix) a proteomics score for a drug target expression; (x) a transcriptomics score for a disease phenotype; (xi) a transcriptomics score for a disease risk factor; (xii) a transcriptomics score for a biological pathway activity; (xiii) a polymorphism score for a drug target expression; (xiv) a polymorphism score for a disease phenotype; (xv) a polymorphism score for a disease risk factor; (xvi) a polymorphism score for a biological pathway activity; (xvii) a polymorphism score for a drug target expression; or any linear or non-linear combination thereof.

7. The method of claim 1 or 2, further comprising running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for the one or more of the biomarker stratifier effects.

8. The method of claim 7, further comprising running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the biomarker stratifier effects.

9. The method of claim 1 or 2, wherein the pharmacomimetic genetic score is associated with a response to a drug that targets a complement protein.

10. The method of claim 9, wherein the complement protein is C3, complement factor B (CFB), complement factor H (CFH), and / or complement factor I (CFI).

11. A system for drug development for age-related macular degeneration (AMD), comprising one or more processors configured by computer-readable instructions to: generate a plurality of principal components (PCs) corresponding to genetic variantAttorney Docket No.136622-1010 data associated with AMD for entities in a dataset to reduce a size of the dataset; generate a biomarker stratifier score for each entity in the dataset to reduce a size of the dataset based at least on the PCs and biomarker stratifier weights for AMD; determine one or more genetic variants that are pharmacomimetic based on the datasets; and determine one or more interactions between the pharmacomimetic instruments and the biomarker stratifier scores, wherein the interactions are indicative of a drug response for one or more drugs.

12. The system of claim 11, further comprising performing one or more actions based on the determined interactions between the pharmacomimetic instruments and the biomarker stratifier scores.

13. The system of claims 11 or 12, wherein the PCs analysis corresponding to genetic variant data is based on a matrix of the one or more of the biomarker stratifier effects.

14. The system of claim 13, further comprising constructing the matrix based on the biomarker stratifier scores, the phenotypes, and the plurality of values.

15. An in silico method for determining a predicted drug activity of one or more drug targets on a plurality of age-related macular degeneration (AMD) phenotype intermediates, comprising: obtaining molecular biomarker stratifier data and AMD disease phenotypes from each of a plurality of subjects; determining the value of the biomarker stratifier effects for each AMD phenotype intermediate based on external data; calculating a biomarker stratifier score for each AMD phenotype intermediate in a disease process; calculating a pharmacomimetic genetic score for each drug target on each AMD phenotype intermediate; and identifying predicted drug activity of the drug targets on one or more AMD phenotypes in subsets of the biomarker stratifier distribution based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for each drug target in association analysis with the AMD phenotype intermediate and the AMDAttorney Docket No.136622-1010 outcome phenotype.

16. A method of determining predicted drug activity of an agent on reducing AMD progression and / or risk in a subject, comprising: obtaining molecular biomarker stratifier data and AMD progression data from a plurality of subjects; determining a value of one or more biomarker stratifier effects for AMD progression based on external data; calculating a biomarker stratifier score for a chosen drug outcome related to subjects’ AMD progression data; calculating a pharmacomimetic genetic score for the agent; and determining the predicted drug activity of the agent on reducing AMD progression based on the statistical interaction of the biomarker stratifier score with the pharmacomimetic genetic score for the agent in association analysis with the AMD progression and / or risk in the subject.

17. The method of claim 16, wherein the biomarker stratifier data comprises a genetic variant, a proteomic variant, a transcriptional variant and / or a polymorphic variant.

18. The method of claim 17, further comprising running a principal components analysis (PCA) or a weighted principal components analysis (wPCA) to identify one or more principal components for one or more of the biomarker stratifier effects.

19. The method of claim 18, further comprising running the principal components analysis or the weighted principal component analysis based on a matrix of one or more of the one or more biomarker stratifier effects.

20. The method of claim 19, further comprising constructing the matrix based on one or more of the biomarker stratifiers, the phenotypes, and external data from a plurality of subjects.

Citation Information

Patent Citations

  • Method and System for Diagnosing Disease and Generating Treatment Recommendations

    US20170137968A1

  • Method of predicting treatment response

    WO2020072004A2