Personalized multi-factor genetic risk prediction system and method of using same

The personalized multi-factor genetic risk prediction system addresses the limitations of static genetic reports by analyzing DNA sequences to generate polygenic scores and disease-specific models, offering interactive risk assessment and tailored testing schedules.

WO2026101884A1PCT designated stage Publication Date: 2026-05-15RES INST AT NATIONWIDE CHILDRENS HOSPITAL
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
RES INST AT NATIONWIDE CHILDRENS HOSPITAL
Filing Date
2025-11-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing genetic risk prediction systems provide static reports lacking interactive capabilities and fail to utilize polygenic scores effectively for personalized risk estimation, with users unable to interact with background population genetic data.

Method used

A personalized multi-factor genetic risk prediction system that utilizes a processor to analyze DNA sequences, identify genetic variants, generate polygenic scores, and create disease-specific models for predicting lifetime disease likelihood and generating tailored testing timelines.

Benefits of technology

Enables interactive disease risk assessment, stratifies individuals based on genetic susceptibility, and provides personalized disease prediction and testing schedules, enhancing the accuracy and applicability of genetic risk reports.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025053927_15052026_PF_FP_ABST
    Figure US2025053927_15052026_PF_FP_ABST
Patent Text Reader

Abstract

A personalized multi-factor genetic risk system for disease prediction and method of use is described herein. The personalized multi-factor genetic risk system is generated by generating a disease specific polygenic score model. Responsive to receiving one or more inputs, wherein the one or more inputs include genomes or partial genomes of an individual, the disease specific polygenic score model generates a lifetime disease likelihood for the individual.
Need to check novelty before this filing date? Find Prior Art

Description

PERSONALIZED MULTI-FACTOR GENETIC RISK PREDICTION SYSTEM ANDMETHOD OF USING SAMECROSS REFERENCES TO RELATED APPLICATIONS

[0001] The following application claims priority under 35 U.S.C. § 119 (e) to U.S. Provisional Patent Application Serial No. 63 / 716,306 filed November 5, 2024 entitled PERSONALIZED MULTI-FACTOR GENETIC RISK PREDICTION SYSTEM AND METHOD OF USING SAME. The above-identified application is incorporated herein by reference in its entirety for all purposes.TECHNICAL FIELD

[0002] The present disclosure generally relates to an analytical system for personalized multi-factor genetic risk stratification and disease risk prediction and method of using same, and more particularly to provide an interactive interfaces and functionality to utilize polygenic scores to generate time-to-event predictions, assess individual life-time disease risks based on a genetic profile, stratify groups of individuals based on genetic susceptibility to particular traits, generate generic risk reports and / or method of using same.BACKGROUND

[0003] Testing of genetic risks across multiple diseases is often offered by commercial genetic companies (e.g., 23andMe). A genetic report is a static form in which results of genetic testing are returned to an individual. These genetic reports give a basic genetic risk for certain diseases. However, background population genetic data that serves as a reference for personal risk estimation for these genetic testing companies remains a “black box” and the individual has no ability to interact with the reference. The static reports also lack any analytical capabilities, including interactive exploration of reference polygenic score (PGS) distributions used for risk estimation.SUMMARY

[0004] One aspect of the present disclosure comprises a non-transitory computer readable medium storing instructions executable by an associated processor to perform a method ofPage 1 of 28NCH-032902.F WO ORDutilizing a multifactor genetic risk model comprising receiving a deoxyribonucleic acid (DNA) sequence from a user, receiving one or more phenotypes of the user, and identifying genetic variants in the DNA sequence including one or more single nucleotide polymorphisms (SNPs) associated with a disease of interest of the user. The method includes generating a plurality of polygenic scores, wherein each polygenic score of the plurality of polygenic scores corresponds to each SNP identified, inputting the plurality of polygenic scores into a disease specific polygenic score model, generating a lifetime disease likelihood prediction from the disease specific polygenic score model based upon the plurality of polygenic scores, and generating an altered testing timeline for the disease of interest based upon the lifetime disease likelihood.

[0005] Another aspect of the present disclosure comprises a personalized multi-factor genetic risk system for disease prediction. The system includes a processing device having a processor configured to perform a predefined set of operations in response to receiving a corresponding input from a secondary device, the processing device comprising memorw wherein a disease specific polygenic score model is stored. In this system the processor intakes a deoxyribonucleic acid (DNA) sequence from a user, the processor intakes a designation of a disease of interest of the user, and the processor intakes one or more phenotypes of the user. Further, the processor identifies genetic variants from the DNA sequence including one or more single nucleotide polymorphisms (SNPs) associated with the disease of interest, and the processor generates a plurality of polygenic scores, wherein each polygenic score of the plurality of polygenic scores is associated with one SNP. The processor generates a lifetime disease likelihood of the user acquiring the disease based upon an output from the disease specific polygenic score model generated based upon the plurality of polygenic score and the processor generates and displays on the secondary device the lifetime disease likelihood and generate an altered testing timeline for the disease of interest.

[0006] Yet another aspect of the present invention includes a non-transitory computer readable medium storing instructions executable by an associated processor to perform a method of generating a multifactor genetic risk model. The method includes receiving a deoxyribonucleic acid (DNA) sequence from a person having a known disease, receiving one or more phenotypes of the person having the known disease, and performing variant matching and endpoint matching on the DNA sequence to generate known disease variants and potential disease variants including single nucleotide polymorphisms (SNPs). The method further includes performing a phenome wide association study (PheWAS) on the potential diseasePage 2 of 28NCH-032902.F WO ORDvariant, generating a polygenic score for the potential disease variant based upon the phenome wide association study, assembling a plurality of generated polygenic scores for potential disease variants linked to the known disease, and utilizing the plurality of generated scores to create a disease specific polygenic score model for the known disease.BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The foregoing and other features and advantages of the present disclosure will become apparent to one skilled in the art to which the present disclosure relates upon consideration of the following description of the disclosure with reference to the accompanying drawings, wherein like reference numerals, unless otherwise described refer to like parts throughout the drawings and in which:

[0008] FIG. 1 illustrates a schematic diagram of a personalized multi-factor genetic risk prediction system in accordance with a second example embodiment of the present disclosure;

[0009] FIG. 2 illustrates a method of training a personalized multi-factor genetic risk prediction system utilizing a personalized multi-factor genetic risk prediction system tool in accordance with a second example embodiment of the present disclosure;

[0010] FIG. 3 illustrates a method of utilizing a personalized multi-factor genetic risk prediction system utilizing a personalized multi-factor genetic risk prediction system tool in accordance with a second example embodiment of the present disclosure;

[0011] FIG. 4 illustrates a method of utilizing a polygenic score (PGS) browser system in accordance with a second example embodiment of the present disclosure;

[0012] FIG. 5A illustrates a known disease variant in accordance with a second example embodiment of the present disclosure;

[0013] FIG. 5B illustrates a known disease variant and other potential variants in accordance with a second example embodiment of the present disclosure;

[0014] FIG. 5C illustrates an endpoint listing related to a disease in accordance with a second example embodiment of the present disclosure;

[0015] FIG. 6 illustrates single nucleotide polymorphisms as related to disease, in accordance with one example embodiment of the present disclosure;

[0016] FIG. 7 illustrates a phenome wide association study (PheWAS) as related to aPage 3 of 28NCH-032902.F WO ORDdisease, in accordance with one example embodiment of the present disclosure;

[0017] FIG. 8 illustrates a correlation matrix as related to a disease, in accordance with one example embodiment of the present disclosure;

[0018] FIG. 9 illustrates method of using a personalized multi-factor genetic risk prediction system . in accordance with one example embodiment of the present disclosure;

[0019] FIG. 10 illustrates method of using a personalized multi-factor genetic risk prediction system to identify an absolute risk of one or more diseases, in accordance with one example embodiment of the present disclosure;

[0020] FIG. 11 illustrates method of using a personalized multi-factor genetic risk prediction system to identify7a time based model of disease acquisition, in accordance with one example embodiment of the present disclosure;

[0021] FIG. 12A illustrates area under the curve and receiver operating characteristic (ROC AUC) distributions for particular diseases, in accordance with one example embodiment of the present disclosure;

[0022] FIG. 12B illustrates area under the curve and receiver operating charactenstic (ROC AUC) distributions for particular diseases, in accordance with one example embodiment of the present disclosure;

[0023] FIG. 12C illustrates area under the curve and receiver operating characteristic (ROC AUC) distributions for particular polygenic scores, in accordance with one example embodiment of the present disclosure;

[0024] FIG. 12D illustrates a highest ranked polygenic scores across categories ranked by unadjusted area under the curve and receiver operating characteristic (ROC AUC), in accordance with one example embodiment of the present disclosure;

[0025] FIG. 12E illustrates a highest ranked polygenic scores across categories ranked by a percentage of explained variability on a logistic liability- scale, in accordance with one example embodiment of the present disclosure;

[0026] FIG. 12F illustrates a plurality of performance metrics for a highest ranked polygenic score, in accordance with one example embodiment of the present disclosure;

[0027] FIG. 13 A illustrates an overview of four distinct phenome wide association study (PheWAS) designs, in accordance with one example embodiment of the present disclosure;

[0028] FIG. 13B illustrates a number of significant phenome-wide associations across aPage 4 of 28NCH-032902.F WO ORDphenome wide association study (PheWAS) atlas, in accordance with one example embodiment of the present disclosure;

[0029] FIG. 13C illustrates phenome wide association study (PheWAS) results for ranked polygenic scores and depression, in accordance with one example embodiment of the present disclosure;

[0030] FIG. 13D illustrates phenome wide association study (PheWAS) results for polygenic scores for depression, in accordance with one example embodiment of the present disclosure;

[0031] FIG. 13E illustrates combined phenome wide association study (PheWAS) results for a depression endpoint and illustrating associated non-target polygenic scores, in accordance with one example embodiment of the present disclosure;

[0032] FIG. 14A illustrates publications that were utilized for polygenic score analysis, in accordance with one example embodiment of the present disclosure;

[0033] FIG. 14B illustrates a scheme of training and testing process for three predictive model types, in accordance with one example embodiment of the present disclosure;

[0034] FIG. 14C illustrates a representation of relative feature importances visualized in panels, in accordance with one example embodiment of the present disclosure;

[0035] FIG. 14D illustrates highest ranked elastic net models for binary disease-status classification, in accordance with one example embodiment of the present disclosure;

[0036] FIG. 14E illustrates relative feature importances for optimal combination of PGSs for a given disease, in accordance with one example embodiment of the present disclosure;

[0037] FIG. 14F illustrates Cox Net models for time-to-event prediction, in accordance with one example embodiment of the present disclosure; and

[0038] FIG. 14G illustrates relative feature importances for optimal combination of polygenic scores for substance abuse, in accordance with one example embodiment of the present disclosure.

[0039] Skilled artisans will appreciate that elements in the figures are illustrated for simplicity and clarity and have not necessarily been drawn to scale. For example, the dimensions of some of the elements in the figures may be exaggerated relative to other elements to help to improve understanding of embodiments of the present disclosure.

[0040] The apparatus and method components have been represented wherePage 5 of 28NCH-032902.F WO ORDappropriate by conventional symbols in the drawings, showing only those specific details that are pertinent to understanding the embodiments of the present disclosure so as not to obscure the disclosure with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.DETAILED DESCRIPTION

[0041] Referring now to the figures generally wherein like numbered features shown therein refer to like elements throughout unless otherwise noted. The present disclosure generally relates to an analytical system for personalized multi-factor genetic risk stratification and disease risk prediction and method of using same, and more particularly to provide an interactive interfaces and functionality to utilize polygenic scores to generate time-to-event predictions, assess individual life-time disease risks based on a genetic profile, stratify groups of individuals based on genetic susceptibility to particular traits, generate genetic risk reports and / or method of using same.

[0042] FIG. 1 illustrates a schematic diagram of a personalized multi-factor genetic risk prediction system 100, in accordance with one of the exemplary7embodiments of the disclosure. The risk prediction system 100 includes a processing device 12, which includes a computing device (e.g. a database server, a file server, an application server, a computer, or the like) with computing capability and / or a processor 14. The processor 14 comprises central processing units (CPU), such as a programmable general purpose or special purpose microprocessor, and / or other similar device or a combination thereof.

[0043] The processing device 12 would generate outputs based upon inputs received from a secondary device 16, cloud storage, a local input form a user, etc. It would be appreciated by having ordinary' skill in the art that the processing device 12 would include a data storage device 17 in various forms of non-transitory. volatile, and non-volatile memories which would store buffered or permanent data as well as compiled programming codes used to execute functions of the processing device 12. In another example embodiment, the data storage device 17 can be external to and accessible by the processing device 12, the data storage device 17 may comprise an external hard drive, cloud storage, and / or other external recording devices 19.

[0044] In one example embodiment, the processing device 12 comprises one of a remote or local computer system 21. The computer system includes desktop, laptop, tablet hand-heldPage 6 of 28NCH-032902.F WO ORDpersonal computing device, IAN, WAN, WWW, and the like, running on any number of known operating systems and are accessible for communication with remote data storage, such as a cloud, host operating computer, via a world-wide-web or Internet.

[0045] In another example embodiment, the processing device 12 comprises a processor, a data storage, computer system memory that includes random-access-memory ("RAM”), read-only-memory (“ROM”) and / or an input / output interface. The processing device 12 executes instructions by non-transitory computer readable medium either internal or external through the processor that communicates to the processor via input interface and / or electrical communications, such as from the secondary device 16 (e.g., smart phone, tablet, personal computer, or other device). In yet another example embodiment, the processing device 12 communicates with the Internet, a network such as a LAN, WAN, and / or a cloud, input / output devices such as flash drives, remote devices such as a smart phone or tablet, and displays. The secondary device 16 includes a display 18, the display having visual, audio, etc. output. In one example embodiment, the risk prediction system 100 is a web-based tool (e.g, no download or installation is needed to utilize the binding affinity prediction system 100). In another example embodiment, the risk prediction system 100 is partially and / or completely downloadable.

[0046] Illustrated in FIG. 2, a personalized multi-factor genetic risk system tool 200 of the personalized multi-factor genetic risk system 100 ingests a deoxyribonucleic acid (DNA) sequence (e.g, FASTA, FASTQ, BAM, SAM files) of a person having a known disease 202. The DNA sequence 202 is sequenced, wherein for a particular disease the DNA sequence 202 will include known disease variants 210 and other variants 214 for a particular disease. In one example embodiment, the DNA sequence 202 includes known disease variants 210 and other variants 214 for a particular disease are known based upon stored data (e.g., such as data stored on FinnGEN). In one example embodiment, a first phenotype 204a, a second phenoty pe 204b, and / or a n number of phenotypes 204c is included as information related to the person having the disease 202. In one example embodiment, the phenotype 214 is one of age, gender, weight, height, existing disease diagnosed, or the like.

[0047] The genetic risk system tool 200 instructs the DNA sequence 202 to undergo a variant or end point analysis 216 to identify the known disease variants 210 and other variants 214. In this example embodiment, the variant analysis identifies specific geneticPage 7 of 28NCH-032902.F WO ORDdifferences in an individual's DNA sequence compared to a reference genome. In this example embodiment, the end point analysis confirms a simple presence or absence of a known, targeted DNA sequence in a sample. As illustrated in the example embodiment of FIG. 5A, responsive to the known disease being cystic fibrosis, a distinct known variant, the cystic fibrosis transmembrane conductance regulator (CFTR) gene is identified, and other potential variants are cataloged for potential disease linkage. As illustrated in the example embodiment of FIG. 5B, responsive to the known disease being coronary artery disease, known variants 210, identified as having a potential linkage to coronary artery disease, are identified, and other potential variants 214 are cataloged for potential disease linkage. As illustrated in FIG. 5C, an example endpoint analysis is illustrated for a selected disease (e.g, dementia). In one example embodiment, endpoints represent binary features indicating the presence or absence of specific conditions / disease, with diagnoses made on the basis of the International Classification of Diseases (e.g, ICD-8, ICD-9, and ICD-10) criteria. The other potential variants 214 include identification of one or more single nucleotide polymorphisms (SNP) 212. In one example embodiment, a first, a second, or a nth SNP are identified, 212a, 212b, 212c.

[0048] In one example embodiment, the DNA sequence 202 undergoes an elastic net analysis 216 to identify the SNPs 212, copy number variations (CNV) 218, and / or DNA methylations 220. when present. In one example embodiment, an optional step of harmonization 221 of the SNP 212, CNV 218, and / or DNA methylations 220 is performed by the processing device 12. In this example embodiment, data (including DNA sequences 202, linked to disease and other phonotypes 204) from multiple databases is utilized to identify SNP 212, CNV 218, and / or DNA methylations 220 related to a condition / disease.

[0049] In one example embodiment, at least one of the variant and endpoint matching 208 or the elastic net analysis 216 utilize phenome-w ide association studies (PheWAS) 222 as illustrated in FIG. 6. For example, FIG. 6 illustrates a first graph 602, wherein a P-value for a particular disease being acquired is on a y-axis. and a genomic position in base pairs of a particular SNP is on an x-axis. In one example embodiment, such as illustrated in FIG. 6, the genetic risk system tool 200 analyzes genotype data from multiple individuals 606a-606d having the known disease, wherein the plurality of SNPs 608 identified by the variant or end point analysis 616 are scored according to equation 1. below:Page 8 of 28NCH-032902.F WO ORD

[0051] Wherein Bi is the P-value for a particular SNP (e.g., effect size of the variant i), Xi is a number of copies of an at risk allele for a variant i that an individual possesses, m is the number of analyzed SNPs, and PRS is the polygenic score, which is output be a polygenic score model 220 created by the genetic risk system tool 200. In this example embodiment, the genetic risk system tool 200 generates a polygenic score 222 for a particular SNP 212 based upon multiple iterations of the SNP identification being linked to a specific disease, and identifying the variation (SNP) as a cause of a specific disease. In one example embodiment, individual PGSs were calculated using whole genome association analysis toolset (e.g., such as, for example, PLINK 2.0).

[0052] In another example embodiment, such as illustrated in FIG. 7, the identified SNPs 212, CNVs 218, and / or DNA methylations 220, undergo a phenome-wide association study (PheWAS) to pair single genetic variants (e.g.. SNPs) against multiple phenotypes simultaneously. In this example embodiment, each PheWAS 222 is fitted to a logistic- regression model with lower-triangular matrix, diagonal matrix, and a transpose of a lower- triangular matrix (LDLT) decomposition using a fast and stable fitting of generalized linear models (fastglni) R package to increase computation speeds of the processing device 12. In one example embodiment, the genetic risk system tool 200 adjusts a relationship between the PGS 222 and each endpoint (to generate PGS-endpoint pairs) for sex, age at the end of follow-up, genetic principal components (PCs) 1-6, and six genotyping-array dummy variables, using Equation (2) below:

[0053] (2) logit (P (E

[0054] In this example embodiment, PCs 1-6 represent a small set (e.g., 6) of new, uncorrelated variables derived from an individual's genetic data using Principal Component Analysis and six genotyping-array dummy variables that represent six distinct categories or factors related to the genotyping process, used as binary7predictors in a statistical model. InPage 9 of 28NCH-032902.F WO ORDthis example embodiment, i is a variant derived from the PheWAS 222 for individual j. In this example embodiment. B represents the effect size of a genetic variant on a particular trait or phenotype, wherein Bo is a statistical notation for an intercept in a regression model. The intercept represents the expected mean value of the phenotype when all other variables (including the genetic variant, age, sex, and principal components) are zero. In this example embodiment, BPgsis the effect size of a given polygenic score of a genotype or SNP 212 of the person having the identified condition or disease. In this example embodiment, Bageis the effect size of an age of the person having the identified condition or disease. In this example embodiment, Bsexis the effect size of a sex of the person having the identified condition or disease. In this example embodiment, Bpciis an effect size of principal components for the variant. In this example embodiment. Barrayj is an effect size of a genotyping array for the individual.

[0055] Further, in this example embodiment, the genetic risk system tool 200 computes an unadjusted area under a receiver operating characteristic curve (ROC AUC) for each PGS-endpoint pair using a SNP array analysis. One example SNP array analysis is a bigsnpr R package. For gender-specific traits, the genetic risk system tool 200 excludes Bsex. In one example embodiment, the genetic risk system tool 200 assesses multiple comparisons for assessment of associations for a single PGS by setting a Bonferroni threshold at p-value indicating statistical significance. In one example embedment, the Bonferroni threshold at p- value that is less than 1.06*10A

[0056] The genetic risk system tool 200 generates three additional PheWAS 222 designs: (1) an exclusion design in which DNA of individuals with a target phenotype were removed to mitigate confounding in secondary associations; (2) noMHC design, in which variants in the human major histocompatibility complex (MHC) region (located at chromosome 6 between 28.5-33.5 megabase pairs (MB)) were excluded and PGS were recalculated before conducting a second PheWAS was performed on the altered scores; and (3) a survival design, in which the logistic-regression model is replaced with Cox proportional-hazards models (survival R package) to use time-to-event as the response variable. In this example embodiment, comparing the exclusion design results and / or the noMHC design results with the original design (unaltered scores and binary endpoints) allows the genetic risk system tool 200 to identify associations driven solely by the MHC locus or phenotypic hitchhiking.Page 10 of 28NCH-032902.F WO ORD

[0057] The genetic risk system tool 200, based upon the PheWAS 222, generates a polygenic score (PGS) 224 for a specific SNP 212, CNV 218, or DNA methylation 220 for a particular disease or phenotype. The genetic risk system tool 200 outputs a plurality' of PGSs 224 for each SNP 212, CNV 218, or DNA methylation 220 associated with the particular disease or phenotype. In one example embodiment, the plurality of polygenic scores are entered into a correlation matrix 800, such as illustrated in FIG. 8. In the example embodiment, FIG. 8 illustrates the correlation matrix 800 for a plurality' of polygenic scores 222 and a likelihood of having combinations of two polygenic scores for a given disease increases overall likelihood of having the disease. In this example embodiment, the correlation matrix 800 illustrates a linear relationship between different PGSs and between PGSs and various traits / disease / conditions and / or outcomes. The correlation matrix 800 illustrates how closely related different polygenic scores are to each other and how well a specific PGS predicts a particular outcome, such as a disease or a quantitative trait.

[0058] The genetic risk system tool 200 generates a disease or condition specific polygenic score model 226 based upon the plurality of PGSs 224. In one example embodiment, the genetic risk system tool 200 utilizes a binary classification using logistic regression with an elastic-net penalty to transform the plurality of PGSs into the disease or condition specific polygenic score model 226. In one example embodiment, the genetic risk system tool 200, implements logistic regression utilizing, for example, scikit-learn Python package. In this example embodiment, the genetic risk system tool 200 uses optimized regularization strength (C) and the elastic-net mixing parameter (e.g, at a 11 ratio) through grid-search cross-validation (e.g, wherein the cross-validation is 5), using ROC AUC as a scoring metric.

[0059] The disease or condition specific polygenic score model 226 includes a time- to-disease / time-to-event or absolute risk of disease manifestation prediction. Regarding the time-to-disease / time-to-event prediction, the polygenic score model 226 presents a likely event age (e.g, when a disease or condition will manifest in an individual). In this example embodiment, the event age is the age the disease or condition is most likely to manifest based on the disease or condition specific polygenic score model 226. To generate the time-to- disease or condition manifestation prediction, the genetic risk system tool 200 utilizes a Cox’s proportional-hazards model with an elastic-net penalty', for example, scikit-learnPage 11 of 28NCH-032902.F WO ORDPython package. Years after the date of the event age are used as a time scale, calculated by subtracting a baseline age from either event age or an end of a follow up (e.g, when the DNA and phenotype data underlying the PGS ceased being collected). In this example embodiment, the end of the follow-up is a predefined time point or date at which data collection for a study subject or the entire research study is concluded, wherein the study subject is the source of the DNA sequence 202 and phenotype data 204. For each disease or condition, DNA sequences of individuals with baseline ages greater than event ages are excluded. The regularization strength (C) and the elastic-net mixing parameter are optimized using grid-search cross-validation (e.g., wherein the cross-validation is 5), with a mean timedependent ROC AUC (1-10 years) as a scoring metric. The genetic risk system tool 200 ranks importance for disease or condition outcomes of PGSs 222 for various SNP 212. CNV 218, or DNA methylation 220 based upon the weights of the elastic-net models. The PGSs 222 are standardized before model fitting, to directly compare importances of each PGSs. Advantageously, the disease specific polygenic score model 226 is optimized per disease, and can be used to quickly and accurately predict disease in users based upon said users DNA.

[0060] Illustrated in FIGS. 3, 9-11, the personalized multi-factor genetic risk system tool 200 of the personalized multi-factor genetic risk system 100 ingests a deoxyribonucleic acid (DNA) sequence 302 of a user. In one example embodiment, the DNA sequence 302 is sequenced and stored as a DNA file (e.g, FASTA, FASTQ, BAM, SAM files). In one example embodiment, a first phenotype 314a, a second phenotype 314b, and / or a n number of phenotypes 314c are received by to the genetic risk system tool 200. As illustrated in the example embodiment of FIG. 9, the user data 304 includes the first phenoty pe 314a as sex, wherein the user is female, the second phenotype 314b as age. wherein the user is forty’ (40) years of age, and an SNP 212. As illustrated in FIGS. 3 and 9, the DNA sequence 302 and the phenotypes 314, collectively the user data 304, are provided to a polygenic score model 318, which calculates a PGS 324 for each SNP 312, CNV 316, and / or DNA methylation 318 identified in the DNA sequence 302 to generate a plurality of PGSs. In one example embodiment, the PGSs 324 are calculated for SNP 312, CNV 316, and / or DNA methylation 318 known to be associated with a specific disease or condition. In this example embodiment, the specificity’ of the disease selection reduces computing pow er, as only SNPs, CNVs and / or DNA methylations that are known to be associated with the disease are used to calculate PGSs. In the example embodiment of FIG. 10, five (5) disease or conditions arePage 12 of 28NCH-032902.F WO ORDselected: breast cancer, Alzheimer, schizophrenia, glaucoma, and ulcerative colitis. The disease specific polygenic score model 226 generates a graph showing density on a y-axis and a standardized PGS score on an x-axis, wherein a bell curve showing an absolute likelihood of disease is shown on each graph for each disease. Based upon the plurality of PGSs entered into the disease specific polygenic score model 226, a standard PGS score is determined and shown on the bell curve for each disease, showing a likelihood specific to the user of acquiring said disease or condition.

[0061] As illustrated in FIGS. 3 and 11, the plurality of PGSs are entered into the disease specific polygenic score model 226 generated by the personalized multi-factor genetic risk system tool 200 in FIG. 2, to generate a lifetime disease likelihood 320. In this example embodiment, an absolute risk graph 1100 is generated. As illustrated in the example embodiment of FIG. 11, an absolute risk 1102 from 0 to 1 is illustrated on a y-axis, wherein 0 means there is no risk and 1 means the user will acquire the disease or condition or die of the disease or condition, and age in years 1104 on an x-axis. A baseline survival rate 1106 of the particular disease or condition is illustrated next to a user specific survivor rate 1 108 based on the plurality of PGSs from the user. The personalized multi-factor genetic risk system 100 removes a requirement for explicit principal-component covariates, allowing the disease specific polygenic score models to be applied externally with only minimal inputs: sex, current age, and the individual’s PGS percentile, which can shift a computational load off of the processing device 12 and onto the secondary device 16. To reduce identifiability risk and reduce computing power, the personalized multi-factor genetic risk system tool 200 utilizes percentiles (an aggregated measure) rather than raw individual-level PGS values. The lifetime disease likelihood 320 is utilized to alter testing schedules or lifestyle factors for users. For example, a user who has a predicted early onset of breast cancer will be assigned breast cancer screening beginning five to ten years prior to predicted onset of breast cancer, rather than waiting for the standard age to begin mammograms. Likewise, for a user who has a predicted early onset of colon cancer will be assigned colon cancer screening beginning five to ten years prior to predicted onset of colon cancer, rather than waiting for the standard age to begin colonoscopies. States another way, common medical testing schedules will be altered based upon the lifetime disease likelihood 320.

[0062] Illustrated in the example embodiment of FIG. 4, a method of using a polygenic score (PGS) browser 400 is illustrated. At 402, data from multiple sources, such asPage 13 of 28NCH-032902.F WO ORDFinnGen project and / or PGS catalog models, are harmonized. At 404, a plurality of models are generated from the harmonized data, including performance evaluation of PGS models, disease prediction models and / or phenome-wide association models. At 406, utilizing the plurality of models to generate PGS browser 400. At 408, the PGS browser 100 receives an input. In one example embodiment, the input includes one or more genomes or partial genomes of one or more individuals. At 410. responsive to receiving inputs, generating a graphical interface including PGS browser results using inputs and plurality of models.

[0063] In one example embodiment, the PGS browser 400 as describe herein is a web-based application that have several functionalities: 1) access to a comprehensive, well- annotated database containing 3,168 polygenic scores (PGS) models; 2) access to a comprehensive database of results of phenome-wide associations for each PGS, including a total of 10,531 studies; 3) enables users to upload PGS values for multiple individuals in the form of raw-values (individual-level data) and / or in the form of pre-calculated percentiles (non-individual level data). Further, such group of individuals could be filtered based on their placement in the desired PGS distributions (i.e. percentile); 4) enables the estimation of lifelong disease risk and survival functions (e.g., a probability of the disease onset at every given age) for individual patients based on their biological sex, current age, and / or percentile in the PGS distribution for the studied disease.

[0064] Advantageously, the PGS browser 400 allows for significantly improved accessible and clinically integrated tools, setting a convenient standard for dissemination of PGS-based genetic analyses. Additionally, the PGS browser 400 provides access to the cutting-edge research results from FinnGen to the broader research community. Further, the analysis and novel interface provided by the PGS browser 400 bridges the gap between the development of PGS models and their practical application in clinical settings.

[0065] The personalized multi-factor genetic risk system tool 200 was tested for efficacy and accuracy by generating lifetime disease likelihood 320 for 3,025 disease specific polygenic score models 226. One model used for testing is FinnGen, which includes genome (DNA sequences) and health data from over 500,000 Finish person. The personalized multifactor genetic risk system 100 efficacy tests included genetic and phenotypic data for 473,681 participants in FinnGen, representing approximately 10% of Finland’s population, including 19,947 of non-Finmsh ancestry. A PGS catalog, which outputs a per mutation (SNP, CNV,Page 14 of 28NCH-032902.F WO ORDor DNA methylation) score was utilized to calculate polygenic scores during testing of the 3,688 models in the PGS Catalog, 3.025 (82%) achieved at least 75% variant overlap with FinnGen and were retained for downstream analyses. The median number of variants per model was 7,372 (range: 1-10,318,272).

[0066] Phenotype definitions were harmonized by matching trait descriptions from existing PGS models to FinnGen disease endpoints, defined using nationwide registries and International Classification of Diseases (TCD) codes. Direct endpoint matches were available for 1,308 models, while 1,717 lacked a match, often because they targeted quantitative traits not available in FinnGen at the time (e.g., brain volume, cystatin C level, etc.). These models were nonetheless retained for PGS calculation and subsequent analyses, as the disease specific polygenic score model 226 is iterative, additions of such quantitative traits can be added as they become available.

[0067] To avoid inflated performance estimates, multiple rounds of manual review were performed to flag and annotate models with sample overlap between FinnGen and PGS development datasets. 460 models (15%) were identified with overlap, which were flagged in a results database and excluded from benchmarking and predictive model analyses. The remaining 2.565 models (85%), free of overlap, were mapped to 227 unique FinnGen endpoints.

[0068] FIG. 12A illustrates ROC AUC distributions for eight (8) PGS diseases and / conditions. FinnGen endpoints were preprocessed to reflect disease prevalence in the Finnish population. For each model, a discriminative ability was quantified using the area under the receiver-operating characteristic curve (ROC AUC) and an association with the corresponding endpoint was tested using logistic regression, adjusting for sex, age, the first six principal components, and genotyping array. In FIG. 12A, a dashed box highlights the ROC AUC distribution for Cancer PGSs.

[0069] FIG. 12B illustrates ROC AUC distributions for the ten (10) types of cancer PGSs. Multiple models were tested for each disease, with varying ROC AUCs. In the graph of FIG. 12B, box plots show a distributions of the multiple models, and each dot indicates the model achieving a highest ROC AUC for a selected disease. To identity’ the most predictive PGSs for each endpoint, a significantly associated model (having a p value <1.06* 10'5) with the highest ROC AUC was selected. Endpoints without significantPage 15 of 28NCH-032902.F WO ORDassociations (as determined by the P-value) were excluded. For example, among 13 models for testicular cancer, PGS000796 was significantly associated and achieved the best ROC AUC of 0.69.

[0070] FIG. 12C illustrates ROC AUC distribution for a 157 highest ranked PGSs. Dots represent in the graph of FIG. 12C indicate a percentage of PGSs exceeding the corresponding ROC AUC thresholds. Only six (3.8%) PGSs achieved a ROC AUC of 0.70 or higher. FIG. 12D illustrates a highest ranked fifteen (15) PGSs across all categories ranked by unadjusted ROC AUC. The highest-performing scores included coeliac disease, ankylosing spondylitis and disorders of iron metabolism. Despite similar overall performance in some cases, substantial individual-level variability among models for the same disease were observed. For example, two hypertension scores yielded comparable AUCs of 0.60 and 0.59 but identified largely distinct individuals in the top 2.5% of the distribution, with only 30% overlap. This discordance persisted even among top performing PGSs. For instance, four coeliac disease models with AUCs of 0.83-0.84 shared only 18-45% overlap in the bottom 2.5% quantile, which can be partly explained by the multimodal shape of the distributions. The primary drivers of discordance were methodological differences in summary statistics from genome-wide association studies and PGS construction methods selection, even when models were built from the same genome-wide association studies. These findings highlight the importance of considering individual-level variability when selecting or applying a PGS for risk stratification.

[0071] FIG. 12E illustrates a highest ranked fifteen (15) PGSs across categories ranked by the percentage of explained variability on a logistic liability scale. To facilitate comparisons individual-level variability, a Spearman rank-correlation matrix was for all 3,025 scores. For each best-performing model, an estimated variance was explained on the logistic liability7scale. The top highest ranked scores were coeliac disease (23.5%), bilirubin metabolism (15%) and congenital vitamin K-dependent coagulation-factor deficiency (12%).

[0072] FIG. 12F illustrates a plurality of performance metrics for the highest ranked PGS, including standardized positive predictive value (set at a 5% standard prevalence) and relative risk (using the 50th percentile as a reference). Follow ing PGS reporting standards, additional quality measures were calculated, including odds ratios (OR), true-positive rates (TPR), standardized positive predictive value (assuming 5% prevalence), and relative andPage 16 of 28NCH-032902.F WO ORDabsolute risks across all percentiles. For example, coeliac disease demonstrated the highest TPR of 0.59 (90%), the relative risk of 11.09 (90 vs 50%), the OR of 2.85, and the absolute risk of 3.32%. These measures provide absolute-risk estimates and enable threshold selection based on desired error rates. Additionally, performance within each major ancestry group in FinnGen was assessed. Of 473,681 participants, 19,947 were non-Finnish: 15,101 nonFinnish Europeans, 1,620 Admixed Americans, 690 East Asians. 433 Africans, 352 South Asians, and 1,751 classified as "Others ’. Endpoints with fewer than eight cases were excluded, reducing the analysis to 740 PGSs and 200 endpoints. In total, 122 best-performing PGS for non-Finnish Europeans, 54 for Admixed Americans, 16 for East Asians, 9 for Africans, 14 for South Asians, and 55 for “Others” were identified. Immune and metabolic traits consistently achieved the highest predictive performance across ancestries. Notably, PGSs for type 2 diabetes achieved statistically significant discrimination across all ancestry groups.

[0073] FIG. 13A illustrates an overview of the four distinct PheWAS designs used, along with the corresponding numbers of studies and phenome-wide associations. In the exclusion design, cases of a target phenotype were excluded to mitigate confounding effects on secondary' associations. In the noMHC design, variants in the MHC region were removed from the PGS model to evaluate PGS-endpoint associations based solely on non-MHC variants. FIG. 13 A illustrates a performance of a phenome-wide association study (PheWAS) for 3,168 polygenic scores across 4,739 FinnGen endpoints, using four complementary designs. The first design is an intact design, wherein the number of studies was 3,168 (n=3,168). The second design is a survival design wherein the number of studies was 3,168 (n=3.168). The survival study employed Cox proportional hazards instead of logistic regression. The third design is an exclusion design wherein the number of studies was 1,172 (n=l,172). The exclusion design excludes individuals with the target phenotype to mitigate phenoty pic hitchhiking. The fourth design is a noMHC design, wherein the number of studies was 3,023 (n=3,023). The noMHC design removes all variants from the MHC locus. The exclusion and noMHC designs provide validation and interpretative support for the primary intact design.

[0074] FIG. 13B illustrates a number of significant phenome-wide associations across a PheWAS Atlas 1300. In the intact design, 3,001 of 3,168 PGSs had at least one phenome-wide significant association. The median number of associations per score was 81Page 17 of 28NCH-032902.F WO ORD(mean=206, max= 1,632). Endpoints were grouped using ICD- 10-based tags. Illustrated in x- y graph of FIG. 13B. the dot indicates the median number of associations.

[0075] FIG. 13C illustrates a PheWAS results for a fifteen (15) highest ranked PGSs and depression. Numbers represent the number of significant phenome-wide associations per endpoint category. Most PGSs showed strongest associations within the target trait category but frequently revealed additional associations in other domains. Collectively, these analyses produced the PGS-PheWAS atlas 1300, cataloging PGS-endpoint associations in FinnGen. The PGS-PheWAS atlas 1300 is usable, for instance, to investigate endpoints associated with a specific PGS. For example, the highest ranking depression PGS was strongly associated with the depression endpoint and showed 1,256 additional significant associations

[0076] FIG. 13D illustrates a PheWAS results for the best-performing PGS for depression. Solid lines indicates results from the intact design, while dashed lines represent results from the exclusion design. The PheWAS results for the best-performing PGS for depression include one hundred and sixteen (116) digestive-system endpoints, one hundred and ten (110) musculoskeletal and connective-tissue endpoints, and eighty t o (82) mentalhealth endpoints. The most associated traits included bipolar disorder, schizophrenia or delusion and panic disorder. Significant, and biologically relevant, associations were also observed with hypothyroidism, irritable bowel syndrome, and cardiovascular disease.

[0077] The exclusion design and noMHC design were compared with the intact design to assess whether secondary7associations w ere driven by the primary trait or the MHC region. For depression, MHC removal had no effect, as the PGS did not include MHC variants. Excluding depression reduced effect sizes but did not eliminate associations, confirming the assocaitions robustness.

[0078] FIG. 13E illustrates combined PheWAS results for the depression endpoint, showing associated non-target PGSs. Alternatively, the PGS-PheWAS atlas 1300 is usable to identify PGSs, beyond a targeted disease, associated with a given endpoint. For example, with depression, significant associations were found for PGSs of schizophrenia, esophagitis, and hypertension. Conversely, negative associations were observed for PGSs of income, lifespan, and height. The PGS-PheWAS atlas 1300 includes 10,531 PheWASs, covering the four study designs, is incorporated into the multi-factor genetic risk prediction system tool 200 to rank the best PGS for a particular diseases, to incorporate the best PGSs into thePage 18 of 28NCH-032902.F WO ORDdisease specific polygenic score model 226..

[0079] FIG. 14A illustrates six publications taht were utilized for further analysis, covering approximately 50% of the retained PGS Catalog models and reliably lacking FinnGen sample overlap. To determine which traits / diseases / conditions benefit from integrating multiple PGSs, sample overlap was excluded between score-development cohorts and FinnGen. During efficacy and accuracy testing, the PGS analysis was restricted to six publications, covering nearly half of the PGS Catalog, and manually confirmed that none of them included FinnGen samples. From these studies, 1,514 models spanning 110 matched FinnGen disease endpoints were extracted. It would be understood by one of ordinary skill in the art that during non-testing phases, any publication could be integrated into the multifactor genetic risk prediction system 100.

[0080] FIG. 14B illustrates a scheme of training and testing process for three predictive model types. FinnGen data was divided into training and testing sets. In the training set, three models were fit, namely, a null model (sex, age, and six genetic PCs), a best-PGS model (null + single highest ranked PGS), and a multi-PGS model (null + an optimal combination of PGSs).

[0081] FIG. 14C illustrates a representation of the relative feature importances visualized in the illustrated panels. Wherein, elastic-net logistic regression was used for binary outcomes and elastic-net Cox proportional hazards for time-to-event prediction. It was found that PGSs with non-zero elastic-net coefficients were informative as to disease or condition acquisition.

[0082] FIG. 14D illustrates highest ranked elastic net models for binary disease-status classification. A left panel of FIG. 14D illustrates incremental ROC AUC improvements (percentile scale) between the best-PGS model, the multi-PGS model, and the null models. The right panel of FIG. 14D illustrates relative feature importances for four feature types: checked pattern (all PGSs combined), slashed pattern (age at the end of follow-up), dotted pattern (sex), and block pattern (first six principal components). The black dashed box highlights PGS feature importances for inflammatory bowel disease (IBD).

[0083] FIG. 14E illustrates relative feature importances for an optimal combination of PGSs for IBD. For disease-status prediction in a left panel of FIG. 14E, 80 of 110 endpoints (73%) showed significant ROC AUC improvement when the best single PGS was added toPage 19 of 28NCH-032902.F WO ORDthe null model (mean AAUC=0.027, p<2.2* 10'16). The largest gains were observed for coeliac disease (AAUC=0.27) and type 1 diabetes (AAUC=0. 17). Multi-PGS models further improved performance for 101 endpoints (mean AAUC=0.037, p<2.2* 10‘16), reaching 0.29 and 0.19 gains for coeliac disease and type 1 diabetes, respectively. Compared with the best- PGS models, the multi-PGS models showed additional significant gains for 87 endpoints (mean AAUC=0.017). with the largest for obesity and hyperalimentation (AAUC=0.08), rheumatoid arthritis (AAUC=0.05), and Crohn’s disease (AAUC=0.04). Many of these improvements were driven by both target and non-target scores. For instance, in inflammatory bowel disease (IBD), key contributors included scores for IBD and ulcerative colitis (target), as well as rheumatoid arthritis, polycythemia vera, and eosinophil count (nontarget).

[0084] FIG. 14F illustrates highest ranked Cox Net models for time-to-event prediction. A left panel of FIG. 14F shows incremental time-dependent ROCAUC (1-10 years) improvements between the best-PGS model, the multi-PGS model and the null models. In this case, the slashed pattern represents baseline age. The black dashed box highlights PGS feature importances for substance abuse. FIG. 14G illustrates relative feature importances for the optimal combination of PGSs for substance abuse. Asterisks in (e) and (g) indicate scores with a negative effect on the disease endpoint.

[0085] To test time-to-event predictions generated by the disease specific polgenic score model 226, time-to-event predictions were compared using elastic-net Cox models. Compared with the null model, the best PGS delivered a significant increase in timedependent (td) ROC AUC for 29 of 110 endpoints (26%; mean AAUC=0.048, p<2.2*10'16). The multi-PGS models extended this to 42 endpoints (38%; mean AAUC=0.056, p<2.2* 10‘16), with the largest gains seen for coeliac disease (AAUC=0.29), obesity and related hyperalimentation disorders (AAUC=0.15), and type 1 diabetes (A AUC A). I I ). Direct comparison of the multi-PGS with the best-PGS models revealed improvements for 15 endpoints (13%; mean AAUC=0.037, p<2.2*10’16), including obesity (AAUC=0.08), sleep disorders (AAUC=0.05), and diabetic retinopathy (AAUC=0.04). As with binary outcomes, these gains were often driven by non-target scores. For example, prediction of substanceabuse onset benefited from integrating PGSs for smoking status, age at first sexual intercourse and neuroticism score.Page 20 of 28NCH-032902.F WO ORD

[0086] To prevent population-structure bias in percentile interpretation during testing, the multi-factor genetic risk prediction system 100 adjusted all 3,168 scores based on ancestry. In this example embodiment, each score was residualized on genetic principal components (PCs) to remove population-structure effects, making distributions comparable across ancestral groups. For each PGS, the predicted scores based on genetic PCs alone were calculated. Several scores, such as those for height (r=0.62; PGS002989), intracranial aneurysm (r=0.59; PGS003407), and atopic dermatitis (r=0.54; PGS002755), showed a substantial correlation with population structure. Subtracting this predicted component from the raw scores eliminated ancestry-driven variation, leaving adjusted scores uncorrelated with PCs.

[0087] To make the disease specific polygenic score models 226 applicable across major continental populations, the PGSs are ancestry-adjusted and refit to time-to-event models using the adjusted scores. In this example embodiment, twenty -two best-PGS models met performance criteria - each achieved a time-dependent ROC AUC above 0.65 in internal FinnGen testing, showed excellent calibration (D-calibration p>0.99), and improved AAUC by at least 0.01 over a null model. These models were then validated in UK Biobank nonWhite British participants (n=78,336).

[0088] It was determined whether 22 PGSs used in time-to-event models showed minimal correlation with population structure (R2<0.05) and evaluated their predictive performance across ancestry groups. Eleven models achieved tdROC AUC > 0.6 in at least three of six ancestries highest ranked performers included palmar fascial fibromatosis (tdROC AUC=0.84), malignant neoplasm of prostate (tdROC AUC=0.81), and atrial fibrillation / flutter (tdROC AUC=0.8).

[0089] Advantageously, combining target-trait PGSs with scores for genetically related, non-target traits to generate the disease specific polygenic score model improved predictive performance for the diseases examined. The phenome-wide association analyses identified numerous cross-trait associations, allowing that shared genetic architecture to be leveraged to improve disease prediction. Across 110 diseases, the disease specific polygenic score model 226 generated by the genetic risk prediction system 100 consistently outperformed corresponding single-score models, with notable AUC gains for obesity, rheumatoid arthritis, Crohn’s disease, and several other conditions. The lifetime diseasePage 21 of 28NCH-032902.F WO ORDlikelihood, whether time-to-event or absolute, will improve testing for given diseases for the users identified as at risk, and will allow the implementation of early medical intervention.

[0090] In the foregoing specification, specific embodiments have been described. However, one of ordinary skill in the art appreciates that various modifications and changes can be made without departing from the scope of the disclosure as set forth in the claims below. Accordingly, the specification and figures are to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of present teachings.

[0091] The benefits, advantages, solutions to problems, and any element(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential features or elements of any or all the claims. The disclosure is defined solely by the appended claims including any amendments made during the pendency of this application and all equivalents of those claims as issued.

[0092] Moreover in this document, relational terms such as first and second, top and bottom, and the like may be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms "comprises," "comprising." “has”, “having,” “includes”, “including,” “contains”, “containing” or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises, has, includes, contains a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by “comprises . . . a”, “has . . . a”, “includes . . . a”, “contains ... a” does not, without more constraints, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises, has, includes, contains the element. The terms “a” and “an” are defined as one or more unless explicitly stated otherwise herein. The terms "substantially”, “essentially”, “approximately”, “about” or any other version thereof, are defined as being close to as understood by one of ordinary skill in the art. In one non-limiting embodiment the terms are defined to be within for example 100%, in another possible embodiment within 5%, in another possible embodiment within 1%. and in another possible embodiment within 0.5%. The term “coupled” as used herein is defined as connected or in contact either temporarily or permanently, although notPage 22 of 28NCH-032902.F WO ORDnecessarily directly and not necessarily mechanically. A device or structure that is “configured” in a certain way is configured in at least that way. but may also be configured in ways that are not listed.

[0093] To the extent that the materials for any of the foregoing embodiments or components thereof are not specified, it is to be appreciated that suitable materials would be known by one of ordinary skill in the art for the intended purposes.

[0094] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Thus the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.Page 23 of 28NCH-032902.F WO ORD

Claims

CLAIMSWhat is claimed is:

1. A non-transitory computer readable medium storing instructions executable by an associated processor to perform a method of utilizing a multifactor genetic risk model comprising: receiving a deoxyribonucleic acid (DNA) sequence from a user; receiving one or more phenotypes of the user; identifying genetic variants in the DNA sequence including one or more single nucleotide polymorphisms (SNPs) associated with a disease of interest of the user; generating a plurality of polygenic scores, wherein each polygenic score of the plurality of polygenic scores corresponds to each SNP identified; inputting the plurality of polygenic scores into a disease specific polygenic score model; generating a lifetime disease likelihood prediction from the disease specific polygenic score model based upon the plurality of polygenic scores; and generating an altered testing timeline for the disease of interest based upon the lifetime disease likelihood.

2. The method of claim 1, wherein each polygenic score of the plurality of polygenic scores is weighted based upon a determined association with the disease of interest.

3. The method of claim 1, wherein a polygenic score weight is quantified using the area under the receiver-operating characteristic curve (ROC AUC), wherein the polygenic score is ranked highest based on a determined P-value of the ROC AUC.

4. The method of claim 1, wherein generating a lifetime disease likelihood prediction comprises generating an absolute likelihood of acquiring the disease of interest.

5. The method of claim 1, wherein generating a lifetime disease likelihood predictionPage 24 of 28NCH-032902.F WO ORDcomprises generating a time-of-event prediction, wherein the event is one of acquiring the disease of interest or dying of the disease of interest.

6. The method of claim 5, wherein the generating an altered testing timeline for the disease of interest is based upon the time-of-event prediction.

7. The method of claim 6, wherein the generating an altered testing timeline for the disease of interest comprises setting a time in years that predates the time-of-event prediction by one of one (1), two (2) or five (5) years.

8. A personalized multi-factor genetic risk system for disease prediction, the system comprising: a processing device having a processor configured to perform a predefined set of operations in response to receiving a corresponding input from a secondary device, the processing device comprising memory, wherein a disease specific polygenic score model is stored; the processor intakes a deoxyribonucleic acid (DNA) sequence from a user; the processor intakes a designation of a disease of interest of the user; the processor intakes one or more phenotypes of the user; the processor identifies genetic variants from the DNA sequence including one or more single nucleotide polymorphisms (SNPs) associated with the disease of interest = the processor generates a plurality of polygenic scores, wherein each polygenic score of the plurality of polygenic scores is associated with one SNP; the processor generates a lifetime disease likelihood of the user acquiring the disease based upon an output from the disease specific polygenic score model generated based upon the plurality of polygenic score; and the processor generates and displays on the secondary device the lifetime disease likelihood and generate an altered testing timeline for the disease of interest.

9. The system of claim 8, wherein the processor weighs each polygenic score based upon an impact of the polygenic score on disease likelihood.Page 25 of 28NCH-032902.F WO ORD10. The system of claim 8, wherein the processor generates a correlation matrix to measure of likelihood of having combinations of two polygenic scores for a given disease increasing an overall likelihood of having the disease, and weighting the polygenic scores impact based upon the correlation matrix.

11. The system of claim 8, wherein the processor weighs each of the plurality of polygenic score by quantifying the area under the receiver-operating characteristic curve (ROC AUC), wherein the processor ranks a polygenic score based on a determined P-value of the ROC AUC.

12. The system of claim 8. the processor generates the lifetime disease likelihood to comprise an absolute likelihood of acquiring the disease of interest.

13. The system of claim 8, the processor generates the lifetime disease likelihood to comprise a time-of-event prediction, wherein the event is one of acquiring the disease of interest or dying of the disease of interest.

14. The system of claim 13, the processor generates the altered testing timeline for the disease of interest based upon the time-of-event prediction, wherein the processor sets a time for testing in years that predates the time-of-event prediction by one of one (1). two (2) or five (5) years.

15. The system of claim 8, wherein the phenotypes comprise at least one of age, sex, height, weight, or other disease of the person having the know n disease.

16. A non-transitory computer readable medium storing instructions executable by an associated processor to perform a method of generating a multifactor genetic risk model comprising: receiving a deoxyribonucleic acid (DNA) sequence from a person having a known disease; receiving one or more phenotypes of the person having the know n disease;Page 26 of 28NCH-032902.F WO ORDperforming variant matching and endpoint matching on the DNA sequence to generate known disease variants and potential disease variants including single nucleotide polymorphisms (SNPs); performing a phenome wide association study (PheWAS) on the potential disease variant; generating a polygenic score for the potential disease variant based upon the phenome wide association study; assembling a plurality of generated polygenic scores for potential disease variants linked to the known disease; and utilizing the plurality of generated scores to create a disease specific polygenic score model for the known disease.

17. The method of claim 16, wherein the generating a polygenic score for the potential disease variant is generated by a selected model, wherein the selected model has a highest score quantified using the area under the receiver-operating characteristic curve (ROC AUC).

18. The method of claim 17, wherein performance metrics for the selected model comprise calculated odds ratios (OR), true-positive rates (TPR), and a standardized positive predictive value, wherein a highest value correlates to a highest score.

19. The method of claim 1 , further comprising performing an elastic net analysis on the DNA sequence to identify SNPs, copy number variations, or DNA methylation in the DNA sequence associated with the known disease.

20. The method of claim 16, wherein the phenotypes comprise at least one of age, sex, height, w eight, or other disease of the person having the known disease.Page 27 of 28NCH-032902.F WO ORD