A method for early disease detection that combines multiple data sources
Combining polygenic risk scores with biomarkers and clinical variables enhances early cancer detection accuracy, enabling targeted screening and improving survival rates by identifying high-risk individuals.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-26
- Publication Date
- 2026-03-12
AI Technical Summary
Existing early cancer detection methods face challenges in maximizing accuracy and reducing false-positive diagnoses, which limits treatment options and survival rates.
Combining polygenic risk scores (PRS) with biomarkers, clinical variables, and genetic correlations to identify individual cancer risk, allowing for targeted screening and intervention.
Improves the accuracy of early cancer detection by identifying high-risk subpopulations, leading to earlier detection and increased survival rates.
Smart Images

Figure 2026508719000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 381,198, filed October 27, 2022, which is incorporated herein by reference in its entirety.
[0002] Technical Field FIELD OF THE DISCLOSURE The present disclosure relates generally to identifying disease risk, and more particularly to methods for identifying disease risk to enable early detection of disease. [Background technology]
[0003] Early cancer detection (ECD) aims to identify cancer or precancerous lesions in patients when the disease is most treatable. Approximately 50% of cancers reach an advanced stage before diagnosis, which limits treatment options and reduces survival rates. Early cancer detection can significantly improve survival rates, and recent advances in the ability to detect biomarkers associated with cancerous or precancerous tissue hold considerable promise. However, maximizing the accuracy of these methods is essential, as false-positive diagnoses can lead to potentially harmful and unnecessary treatment.
[0004] Polygenic risk scores (PRS) have been shown to effectively stratify disease risk in several cancers. Herein, we describe how a patient's PRS, along with ECD biomarkers, can be combined with other relevant information, such as a family history of relevant cancers, to improve the accuracy of early cancer detection and reduce false-positive rates. Similar approaches can be used to predict other common diseases, including coronary artery disease (CAD). Summary of the Invention
[0005] The present invention relates to a method for improving the accuracy of early cancer detection by a priori identification of PRS and then combining it with accuracy-improving biomarkers. The method comprises the following steps: i) Methods for calculating an individual's genetic risk, including a polygenic risk score (PRS), which can be identified either at screening or as early as birth using low-coverage whole-genome analysis, WGS, or microarray genotyping. ii) Methods for measuring biomarkers (including proteins and metabolites). iii) updating the calculated risk (if necessary) to take into account additional clinical variables, including age, sex, and history of infectious diseases or environmental exposures (e.g., smoking);
[0006] Ultimately, this information can identify an individual's current risk of cancer (both solid and liquid tumors), heart disease, or other common diseases, allowing for more effective intervention. The above approach can be improved by also considering genetic correlations between various cancer types and correlations between different cancer types and bioanalytical profiles in the blood. Using the methods of the present disclosure, subpopulations at higher than average risk of cancer can be identified, which can then inform more routine and / or additional testing, leading to earlier detection and ultimately increased recovery / survival rates.
[0007] In practice, patients whose early cancer detection (ECD) biomarkers are borderline negative or negative but have a high polygenic risk for a particular cancer type may be recommended for further testing. In more complex iterations, individual risks for multiple cancer types (e.g., breast, colon, pancreas), each affecting multiple analytes used in early cancer detection, may be leveraged to improve cancer prediction (in some cases, a biomarker, like CEA, may represent risk for multiple cancers, including colon, prostate, lung, thyroid, and others). In further iterations, an individual's PRS may be used to assign weights to tumor origin. This information can then guide additional screening and management of the individual's cancer risk. For example, if an individual has a strong genetic predisposition to breast cancer and also has an ECD signal associated with cancer, a positive ECD result may lead to more focused imaging / testing, such as a breast MRI.
[0008] This approach can also be applied in other contexts, such as screening for primary prevention of coronary artery disease. In addition to serum cholesterol and weight measurements, polygenic risk scores for body mass index (e.g., PRS-BMI) and cholesterol (e.g., PRS-cholesterol) can be used to communicate predicted risk for coronary artery disease. Genetic risk for unrelated diseases (e.g., BRCA1 pathogenic variants) associated with CAD-related diabetes and increased risk of coronary complications can also be used. This information can be used to recommend dietary and lifestyle modifications, additional laboratory tests, and imaging and procedures (e.g., CT scans or coronary catheterization).
[0009] Dietary recommendations may be similar to those for other at-risk patients, but identifying additional genetic risk increases the likelihood of compliance. Laboratory testing may include more sophisticated (but more expensive) tests, which are typically reserved for high-risk patients. Ongoing screening may consist of cancer screening tests recommended for the general population, but these may begin at an earlier age and be performed more frequently. The recommended frequency of screening may be determined by matching the patient's combined risk (age plus genetic risk) to the risk associated with the average person of a particular age. Examples of screening tests that may be recommended include mammography for breast cancer and colonoscopy for colon cancer.
[0010] Genetic predisposition to cancer is often associated with multiple different types of cancer. This is true for specific monogenic cancer genes, such as BRCA1, as well as polygenic risk scores that can capture genome-wide genetic correlations between cancer types. In addition, bioanalytical blood profiles and clinical data on cancer risk factors can be correlated with each other and with polygenic cancer risk scores. Modeling the joint probability distribution of these different modalities can improve the accuracy of early cancer detection.
[0011] Having described above in general terms certain exemplary embodiments, reference is now made to the accompanying drawings, which are not necessarily drawn to scale. Some embodiments may include fewer or more components than those shown. [Brief explanation of the drawings]
[0012] [Figure 1] 1 shows an exemplary relationship between the predictive value (PPV) of a test and cancer prevalence in a test population, according to exemplary embodiments described herein.
[0013] [Figure 2A] 1 illustrates an example simulation architecture using PRS and biomarkers to identify early disease detection for various diseases, according to certain example embodiments described herein. [Figure 2B] 1 illustrates an example simulation architecture using PRS and biomarkers to identify early disease detection for various diseases, according to certain example embodiments described herein. [Figure 2C] 1 illustrates an example simulation architecture using PRS and biomarkers to identify early disease detection for various diseases, according to certain example embodiments described herein. [Figure 2D] 1 illustrates an example simulation architecture using PRS and biomarkers to identify early disease detection for various diseases, according to certain example embodiments described herein. [Figure 2E] 1 illustrates an example simulation architecture using PRS and biomarkers to identify early disease detection for various diseases, according to certain example embodiments described herein.
[0014] [Figure 3A]10 illustrates exemplary distributions of PRS and biomarkers generated by simulation and used for early detection of disease, according to certain exemplary embodiments described herein. [Figure 3B] 10 illustrates exemplary distributions of PRS and biomarkers generated by simulation and used for early detection of disease, according to certain exemplary embodiments described herein. [Figure 3C] 10 illustrates exemplary distributions of PRS and biomarkers generated by simulation and used for early detection of disease, according to certain exemplary embodiments described herein.
[0015] [Figure 4] 1 shows a schematic block diagram of an example circuit embodying a device that may perform various operations according to example embodiments described herein. DETAILED DESCRIPTION OF THE INVENTION
[0016] Certain illustrative embodiments will now be described in more detail below with reference to the accompanying drawings, in which some, but not necessarily all, embodiments are shown. Because the invention described herein may be embodied in many different forms, the invention should not be limited to only the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements.
[0017] Cancer can be treated more effectively if detected early. Subpopulations with higher-than-average risk can be identified, and the cost-benefit tradeoff for these subpopulations justifies more frequent testing compared to the general population. This can ultimately result in higher survival rates and more effective resource allocation.
[0018] From an individual's clinical whole-genome sequencing (WGS), a bank of polygenic models can be generated covering many serious diseases, including breast cancer (BC), lung cancer, prostate cancer, and colorectal cancer (CRC), cardiovascular disease, type 2 diabetes, stroke, Alzheimer's disease, liver disease, and kidney disease. These models include the actions of tens of thousands of variants across the genome, rather than just rare variants in a few genes. Polygenic models can be combined with age, family history, and clinical data, as well as other available analytes, to create an integrated risk score (IRS). Clinical reporting of the IRS allows screening and intervention planning based on the IRS, thereby reducing healthcare costs, improving outcomes, and optimizing the utility of multi-cancer early detection (MCED) tests. The IRS can be continuously optimized across diverse populations with additional clinical data and analytes such as methylation, as well as new modeling methods such as deep neural networks (DNNs) that capture disease-causing signaling pathways.
[0019] Scalable solutions can improve access to appropriate screening and interventional treatments, allowing individuals of all ethnicities to access their own WGS and its interpretation to proactively manage their health. This can potentially improve compliance as individuals understand their personal risks incorporating their genetic makeup, and can help clinicians address various phenotypes, tailor interventions, and better characterize disease risk for targeted treatments, behaviors, and diets.
[0020] IRS can be improved in at least four ways: 1) enhanced analysis, 2) incorporating rare variants from WGS, 3) incorporating clinical data and additional analytes, and 4) expanded datasets.
[0021] In some cases, models can be extended with "opaque" machine learning methods such as neural networks and additional analytes to improve the AUC and odds ratio per standard deviation (OR / SD). For example, in BC, integrating the Tyrer-Cuzick (TC) model with the PRS can improve the area under the receiver operating characteristic curve (AUC) of the IRS for remaining lifetime risk by integrating the TC. This can involve, for example, a fixed stratification method that co-estimates correlates such as age and family history, calibrating subgroup risks, confirming this calibration using the Hosmer-Lemeshow test, and creating an integrated Cox proportional hazards model using the subpopulation calibration.
[0022] In some cases, OR / SD can be improved over standard PRS. This can include, for example, decomposing the genome into ethnic subcomponents, finding the optimal PRS for each subcomponent, weighting SNPs using HapMap data and SNP ORs for multiple ethnicities and functional genomics, and ensemble methods that combine PRSs for the same phenotype with PRSs for related phenotypes. In some cases, the AUC of linear PRS in BC can also be improved by incorporating neural networks (NNs) and deep neural networks (DNNs) to model nonlinear gene interactions and predict gene expression based on gene motifs. These methods can be applied to many diseases.
[0023] In some cases, WGS can detect small copy number variants (CNVs) and structural variants (SVs) along with single nucleotide variants (SNVs) by emulating the performance of multiprotocol gene panel tests and validating with orthogonal long-read sequencing methods. In some cases, rare disease-associated loss-of-function (LoF) or missense variants can be identified using ensemble methods and disease genes weighted using public disease association data. In some cases, reports can be generated that include sets of pharmacogenomic-associated genes to improve compliance with personalized medication and dosing recommendations.
[0024] Additional analytes that influence risk may include methylation, mRNA, miRNA, protein, and clinical phenotypes such as blood counts and metabolomics. Blood methylation, in particular, is a promising robust analyte that captures many of the epigenetic effects. Furthermore, methylation may enhance the AUC of the CVD IRS. Similar multi-ethnic curation has been employed for multiple diseases. In some cases, multi-ethnic datasets may be curated to improve polygenic performance. Similar multi-ethnic curation can also be employed for other diseases. In some cases, a WBS-based "wellness test" may be offered. In some cases, patients may enter personal clinical and family history data to improve the performance and feasibility of their test.
[0025] For each disease, data can be pooled to examine primary outcomes of cases versus suppression, including reduction in incidence, early detection, changes following screening, and intervention. Data can be pooled across enriched and non-enriched cohorts. If necessary, these measures can be combined with renal and CVD to demonstrate efficacy. Sensitivity and specificity can also be assessed for each cancer separately and for annual MCED testing of high-risk subjects. Health economic models of IRS testing can be generated both with and without MCED. The goal is to demonstrate sufficient utility and economic benefit for guideline changes and reimbursement.
[0026] Definitions of Certain Terms Unless otherwise defined, technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. Materials referenced in the following description and examples are available from commercial sources unless otherwise noted.
[0027] The terms "computer-readable medium" and "memory" refer to non-transitory storage hardware, non-transitory storage devices, or non-transitory computer system memory capable of storing computer-executable instructions or software programs that can be accessed by a controller, microcontroller, computing system, or module of a computing system. The non-transitory computer-readable medium can be accessed by the computing system or module of a computing system to retrieve and / or execute the computer-executable instructions or software programs stored on the non-transitory computer-readable medium. Exemplary non-transitory computer-readable media can include, but are not limited to, one or more types of hardware memory, non-transitory tangible media (e.g., one or more magnetic storage disks, one or more optical disks, one or more USB flash drives), and computer system memory or random access memory (e.g., DRAM, SRAM, EDO RAM).
[0028] The term "computing device" may refer to any computer embodied in hardware, software, firmware, and / or any combination thereof. Non-limiting examples of computing devices include personal computers, servers, laptops, mobile devices, smartphones, fixed terminals, personal digital assistants ("PDAs"), kiosks, custom hardware devices, wearable devices, smart home devices, Internet of Things ("IoT") enabled devices, and network-connected computing devices.
[0029] Exemplary Implementation Device FIG. 4 illustrates an apparatus 400 that may include an example system that may implement example embodiments described herein. The apparatus may include a processor 402, a memory 404, communications circuitry 406, and input / output circuitry 408, each of which is described in more detail below, along with any number of additional hardware components not explicitly shown in FIG. 4. While FIG. 4 merely shows various components connected to the processor 402, it will be understood that the apparatus 400 may further include a bus (not explicitly shown in FIG. 4) for passing information between any combination of the various components of the apparatus 400. The apparatus 400 may be configured to perform various operations described above, as well as various operations described below in connection with FIG. 4.
[0030] The processor 402 (and / or coprocessors, or any other processors assisting or otherwise associated with the processor) may communicate with the memory 404 via a bus to pass information between components of the apparatus. The processor 402 may be embodied in several different ways, for example, may include one or more processing devices configured to execute independently. Additionally, the processor may include one or more processors configured in series via a bus to enable independent execution of software instructions, pipeline processing, and / or multithreading. Use of the term "processor" may be understood to include a single-core processor, a multi-core processor, multiple processors in the apparatus 400, a remote processor or a "cloud" processor, or any combination thereof.
[0031] Processor 402 may be configured to execute software instructions stored in memory 404 or otherwise accessible to the processor (e.g., software instructions stored on a separate storage device). In some cases, the processor may be configured to perform hard-coded functions. Thus, whether configured in a hardware or software manner, or a combination of hardware and software, processor 402 represents an entity (e.g., an entity physically embodied in circuitry) that, when configured appropriately, can perform operations in accordance with various embodiments of the present invention. Alternatively, as another example, if processor 402 is embodied as an executor of software instructions, the software instructions, when executed, may tangibly configure processor 402 to perform the algorithms and / or operations described herein.
[0032] The memory 404 may be non-transitory and may include, for example, one or more volatile and / or non-volatile memories. In other words, for example, the memory 404 may be an electronic storage device (e.g., a computer-readable storage medium). The memory 404 may be configured to store information, data, content, applications, software instructions, or the like to enable the device to perform various functions in accordance with the exemplary embodiments contemplated herein.
[0033] The communications circuitry 406 may be any means, such as a device or circuit, embodied in either hardware or a combination of hardware and software, configured to receive and / or transmit data from a network and / or any other device, circuit, or module in communication with the apparatus 400. In this regard, the communications circuitry 406 may include, for example, a network interface for enabling communication with a wired communications network or a wireless communications network. For example, the communications circuitry 406 may include one or more network interface cards, antennas, buses, switches, routers, modems, and supporting hardware and / or software, or any other devices suitable for enabling communication over a network. Additionally, the communications circuitry 406 may include processing circuitry for causing the transmission of such signals to a network or for processing the reception of received signals from the network.
[0034] Device 400 may include input / output circuitry 408 configured to provide output to a user and, in some embodiments, to receive indications of user input. Note that some embodiments do not include input / output circuitry 408, in which case user input may be received via a separate device. Input / output circuitry 408 may comprise a user interface, such as a display, and may further comprise components that govern use of the user interface, such as a web browser, a mobile application, or a dedicated client device. In some embodiments, input / output circuitry 408 may include a keyboard, a mouse, a touchscreen, a touch area, soft keys, a microphone, a speaker, and / or other input / output mechanisms. Input / output circuitry 408 may utilize processor 402 to control one or more functions of one or more of these user interface elements via software instructions (e.g., application software and / or system software such as firmware) stored in memory accessible to processor 402 (e.g., memory 404).
[0035] In some embodiments, various components of device 400 may be hosted remotely (e.g., by one or more cloud servers), so that not all components need be present in one physical location. Additionally, some of the functionality described herein may be provided by third-party circuitry. For example, device 400 may access one or more third-party circuitry via any type of network connection that facilitates the transmission of data and electronic information between device 400 and the third-party circuitry. Device 400 may then remotely communicate with one or more of the components previously described as comprising device 400.
[0036] As will be appreciated based on this disclosure, some exemplary embodiments may take the form of a computer program product that includes software instructions stored on at least one non-transitory computer-readable storage medium (e.g., memory 404). In such embodiments, any suitable non-transitory computer-readable storage medium may be utilized, some examples of which include non-transitory hard disks, CD-ROMs, flash memory, optical storage devices, and magnetic storage devices. With respect to the particular device embodied by apparatus 400 illustrated in FIG. 4, it will be understood that loading the software instructions into a computing device or apparatus creates a special-purpose machine with means for performing the various functions described herein.
[0037] Having described specific components of the device 400, exemplary embodiments are described below.
[0038] Example Operation Relationship between IRS threshold and sensitivity, PPV, and specificity Assume that the IRS is normalized to a mean of 0 and a standard deviation (SD) of 1, and the odds ratio (OR) per SD is r. The probability of being positive for the disease at an IRS value x is p(x) = Cr x where C is some constant. Integrating over all possible values of x and the associated probabilities should reproduce the population incidence p.
number
[0039] If a subject is flagged as being above some PRS threshold t, then the probability P(t) that the subject is above some PRS threshold t and positive for the disease can be:
number
[0040] where normcdf is the normal cumulative distribution. Similarly, the probability N(t) that a subject is above some PRS threshold t and negative for the disease is:
number
[0041] The sensitivity at the threshold t is given by the following formula:
number
[0042] Thus, given a population incidence rate p and a PRS with OR / SD r, to achieve a particular sensitivity, the threshold t can be set as follows:
number
[0043] This gives the following positive predictive value (PV):
number
[0044] In addition, the following specificity is obtained:
number
[0045] Approximation of changes in sensitivity and specificity due to changes in AUC Although the shape of the ROC curve is unknown, when it is necessary to approximate the change in screening performance based on the change in AUC, further approximation is necessary. Assume that the ROC curve is symmetric around the x=1-y line representing the average case or the specificity=sensitivity line, and consists of two lines, as shown in Figure 2. Based on these assumed shapes, it can be shown that A0=Sens0=Spec0 and A1=Sens1=Spec1. When the AUC is A0, the operating point at sensitivity Sens2 can be estimated as follows:
number
[0046] If the AUC is improved to A1 and the sensitivity is improved while keeping the specificity the same, the new sensitivity achievable is:
number
[0047] Work by PRS A PRS can be used to select individuals most at risk and recommend ongoing screening. In this scenario, a risk level is pre-assigned to patients using the PRS, and then individuals deemed to be at high risk for the disease of interest can be recommended ongoing monitoring. By testing individuals who are more likely to develop the disease, the accuracy of the test increases. The following equation illustrates this relationship using a cancer screening test as an example:
number
[0048] The left side shows the accuracy of the test, i.e., the positive predictive value (PPV) of the test. As the cancer prevalence in the test population, as identified by the PRS (P(cancer)), increases, so does the PPV of the test. This phenomenon is further illustrated in Figure 1, which shows this relationship for tests with various false positive rates (FPR). For tests with low FPR, the PPV can be significantly improved by using the PRS to focus on patients at highest risk of disease.
[0049] This approach allows prioritization of individuals for early cancer detection, thus providing a more accurate estimate of each individual's cancer risk and increasing the accuracy of the test by applying the ECD only to high-risk individuals.
[0050] Potential benefits: ECD biomarker testing can be more costly than identifying cancer PRS. To maximize the benefit of ECD biomarker testing, it is useful to identify patients with ECD biomarker test results that are most likely to be clinically treatable. Patients at high polygenic risk for cancer may benefit more from ECD biomarker testing than patients at low polygenic risk for cancer.
[0051] Combining PRS and ECD signals to reduce false positives In addition to risk stratification, PRS can be combined with existing ECD tests to reduce the false positive rate of the test. Expected false positive rates for cancer detection were derived using ECD biomarkers alone and then using the reduction in false positive rate resulting from combining ECD biomarkers with PRS. In this example, a population of control samples (D) and a population of cancer patients (
number
number
number
[0052] The overall probability that a patient has cancer was assumed to be equal to the overall probability that the patient does not have cancer (i.e.,
number
number
number
[0053] The probability of a signal X1 corresponding to a false positive (i.e., mischaracterizing a control sample as a cancer sample) was then calculated from the cumulative distribution function using X1 as follows:
number
[0054] False positives in cancer detection using ECD biomarkers and PRS A method for making a cancer / control call using two signals together (X1, the ECD biomarker signal and an orthogonal signal, X2, which may be a PRS) was computationally simulated according to the calling scheme shown in Table 1 below. [Table 1]
[0055] As mentioned above, the same assumptions were made for the distribution of signal X2 as for the distribution of signal X1. Based on the use of both distributions according to Table 1, the probability of calling a false positive and the probability of making no call are specified as in Table 2 below, where "normcdf" is the normal cumulative distribution function (e.g., in MATLAB®). [Table 2]
[0056] Assuming m1=6 and m2=6 / sqrt(3), the probability value is P FPX1 =0.0013, P FPX2 = 0.0416, and P FPX1X2 =0.000056, as calculated above.
[0057] The measurements of the control population (D) and the cancer population were assumed to have the same distribution as in the previous example. The method for making a cancer / control call by mathematically combining the two signals, X1 and X2, to produce a single product (X1*X2 or "X1X2") was calculated as follows:
number
number
[0058] Assume again that the overall cancer-free probability is equal to the overall cancer probability (i.e.,
number
number
number
number
[0059] Next, to estimate the false positive rate, the joint probability function was integrated as follows:
number
number
number
[0060] Next, X2 was solved as follows:
number
[0061] Therefore, the false positive rate was determined as follows:
number
[0062] The false positive rate can then be empirically calculated using the following MATLAB® code, where “sum” is the false positive rate for the different signal means m1 and m2: % variables n = 2000; m1 = 6; m2 = 6 / sqrt(3); lim = 20; delta = 2*lim / (n-1); x1_vec = [-lim:delta:lim]; x2_vec = [-lim:delta:lim]; sum = 0; for x1 = x1_vec ind = find( x2_vec > (m1^2 + m2^2 -2*m1*x1) / (2*m2) ); for x2 = x2_vec(ind) sum = sum + exp(-0.5*(x1^2+x2^2))*delta^2 / (2*pi); end end sum
[0063] Here, the "sum" corresponds to the probability of observing a false positive in this joint probability scenario combining signal mean m1 with some weaker signal mean m2. The probability of a false positive is determined to be P(False Positive) = sum = 0.00026, while the individual probabilities (as assessed in the previous example) are P FPX1 =0.0013 and P FPX2 =0.0416, which was higher.
[0064] Simulations demonstrate that combining two independent signals, one of which has three times higher variance than the other, can reduce the false positive rate by at least five-fold compared to using either signal alone.
[0065] Simulation results showing the increased specificity of correlated traits are shown in Table 3 below. [Table 3]
[0066] Simulations involving PRS and a single cancer biomarker Figure 2A shows an exemplary simulation of a single cancer biomarker. The utility of PRS is further demonstrated by comparing early detection models simulating data with and without PRS as a component. The following plate illustrates the simulation scenario.
[0067] Here, the PRS predicts cancer risk and cancer status is related to biomarkers used as indicators of cancer in early detection tests.
[0068] The outline of the simulation was as follows. 1. Assume that the PRS is normalized to N~(0, 1). 2. Simulate cancer status using a PRS log odds ratio of 1.5–3, consistent with results seen for other diseases (4.5 for T1 diabetes, 1.9 for prostate cancer, and 1.74 for breast cancer), and an intercept based on average population risk. 3. Simulate the ECD biomarker with Poisson (λ+k(t)). a. λ is the baseline value of the biomarker in an uninfected population. bk(t) simulates the growth of biomarkers with cancer progression at time t, which is simulated using a logistic function. 4. Time since cancer onset, t, is modeled as a negative binomial distribution with mean 10 days. 5. Simulate cancer conditions in 500,000 individuals. 6. Compare a biomarker-only model with a model incorporating both biomarkers and PRS.
[0069] Adding a PRS is most effective when the power of the conventional ECD test is low; as shown in Table 4, when combined with a PRS, recall improves by up to 15 points, and even a low-powered PRS results in an increase in recall. [Table 4]
[0070] As shown in Figures 3A-B, when the PRS log odds is 3 and the ECD biomarker has moderate power (see Figure 3A), the recall in the simulation increases by 15%. This PRS is well within the range found empirically and could be made more powerful by including additional covariates. The simulated biomarker test has moderate power, in which case early disease status is difficult to distinguish. This is realistic given that most early cancer tests are still in development.
[0071] The inclusion of PRS may allow for earlier case identification than biomarker-only testing. As shown in Figure 3C, the median cases identified by the full model were 6 days earlier than the biomarker-only model, and the cases missed by the biomarker-only model and identified by the full model were earlier-stage cancers.
[0072] In this relatively simple scenario, incorporating a PRS into an ECD test provided substantial improvement. More complex scenarios, including additional biomarkers and correlated PRSs, can provide even greater improvements.
[0073] Additional Example 1: Screening for Multiple Cancers As shown in Figure 2B, the simple scenario of a single cancer / biomarker can be expanded to include multiple cancers or diseases, where the PRSs of different cancers may not be correlated, but the biomarkers are. This allows us to use the correlation between biomarkers to improve the detection power for both Cancer #1 and Cancer #2. Examples of cancer biomarkers shared between different cancer types include HER2 / neu, alpha-fetoprotein, and carcinoembryonic antigen. Predicting the PRSs of cancers that share biomarkers can not only improve the detection power for each cancer type, but also help distinguish between cancer types.
[0074] Additional Example 2: As shown in Figure 2C, the model is further extended to multiple PRSs correlated with multiple cancers. Each cancer has a specific biomarker profile (occasionally sharing biomarkers). The correlation of PRSs confers additional power for cancer detection. The genetic correlation between cancer types shown in this example can increase the power of detection for both cancer types in addition to the increased power provided by the shared biomarkers.
[0075] Additional Example 3: As shown in Figure 2D, the model can be extended to cases where the PRS predicts biomarkers that contribute to disease risk itself, rather than just response. Examples 2 and 3 show cases where tumors cause the secretion of biomarkers that are later used for diagnosis. However, some biomarkers are present in causal early stages of tumor development, which affects the combined distribution and polygenic score of biomarkers, as well as the statistical modeling techniques required for cancer detection. Thus, taking these early-stage biomarkers into account can be predicted by the PRS model without directly measuring them, resulting in a more robust and accurate individual disease risk.
[0076] Additive Modeling: Figure 2E shows a more complex simulation. To model the complex interactions between biomarkers, PRS, and cancer risk, statistical modeling is used to model such interactions. In some embodiments, the statistical modeling uses a hidden Markov model (HMM). In some embodiments, the statistical modeling uses a machine learning model, such as a neural network. In Figure 2E, black arrows represent transitions between states and blue arrows represent observations. Biomarker data is observed over time, and the PRS contributes to the probability of state transitions.
[0077] A Viterbi (dynamic programming) algorithm is then used to calculate the most likely state sequence that produced the observation. generated quantities { array[T_unsup] int<lower=1, upper=K> y_star; real log_p_y_star; { array[T_unsup, K] int back_ptr; array[T_unsup, K] real best_logp; real best_total_logp; for (k in 1:K) { best_logp[1, k] = log(phi[k, u[1]]); } for (t in 2:T_unsup) { for (k in 1:K) { best_logp[t, k] = negative_infinity(); for (j in 1:K) { real logp; logp = best_logp[t - 1, j] + log(theta[j, k]) + log(phi[k, u[t]]); if (logp > best_logp[t, k]) { back_ptr[t, k] = j; best_logp[t, k] = logp; } } } } log_p_y_star = max(best_logp[T_unsup]); for (k in 1:K) { if (best_logp[T_unsup, k] == log_p_y_star) { y_star[T_unsup] = k; } } for (t in 1:(T_unsup - 1)) { y_star[T_unsup - t] = back_ptr[T_unsup - t + 1, y_star[T_unsup - t + 1]]; } } }
[0078] conclusion Many modifications and other embodiments of the inventions described herein will come to mind to one skilled in the art to which the inventions described herein pertain having the benefit of the teachings presented in the foregoing description and the associated drawings. It is, therefore, understood that the invention is not to be limited to the particular embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims. Furthermore, while the foregoing description and associated drawings describe exemplary embodiments in the context of certain exemplary combinations of elements and / or functions, it is understood that different combinations of elements and / or functions may be provided in alternative embodiments without departing from the scope of the appended claims. In this regard, combinations of elements and / or functions other than those expressly described above are contemplated, for example, as may be set forth in some of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A method for determining whether a subject is at increased risk of a disease, comprising: applying a polygenic risk model to the subject's genotype to generate a polygenic risk score (PRS) for the subject, the polygenic model being associated with a particular disease or group of diseases; Identifying one or more biomarker values for the subject, wherein one or more of the one or more biomarkers are associated with the specific disease or group of diseases; determining one or more recommended actions for the subject based at least in part on the PRS score and the one or more biomarker values; providing a notification regarding one or more of the one or more recommended actions for the subject, the one or more biomarker values of the subject, or the PRS of the subject; The method comprising:
2. assigning the subject to a risk category based on the PRS; In instances where the assigned risk category corresponds to a risk category associated with a high disease risk, determining a recommended biomarker action, the recommended biomarker action describing one or more recommended biomarker tests to be performed on the subject, the recommended biomarker action being included in the one or more recommended actions; The method of claim 1 further comprising:
3. The method of claim 2 , wherein the one or more biomarker values are identified based on results of the one or more recommended biomarker tests.
4. The method of claim 1 , wherein the disease states used in the polygenic risk model are determined using PRS log odds ratios.
5. 5. The method of claim 4, wherein the PRS log odds ratio value is based on an average population risk.
6. applying additional polygenic risk models to the subject genotype to generate one or more additional PRSs for the subject, wherein the subset of one or more biomarker values is associated with an additional specific disease or group of specific diseases associated with the additional polygenic risk model; The method of claim 1 further comprising:
7. 7. The method of claim 6, wherein one or more of the one or more biomarker values are associated with both the specific disease or specific disease group associated with the polygenic risk model and the additional specific disease or specific disease group associated with the additional polygenic risk model.
8. The method of claim 6 , wherein the additional specific disease or specific group of diseases is also associated with the polygenic risk model.
9. Applying the polygenic risk model to the subject genotypes further comprises: generating one or more early stage biomarker values for the specific disease or group of diseases associated with the polygenic risk model; determining the one or more recommended actions for the subject based at least in part on the PRS score, the one or more biomarker values, and the one or more early stage biomarker values; The method of claim 1 , comprising:
10. The method of claim 1 , further comprising estimating a joint probability distribution of the PRS and the one or more biomarker values.
11. The method of claim 1 , further comprising estimating a subset of the one or more biomarker values based on the PRS.
12. determining a probable disease status of the subject using statistical modeling, the probable disease status indicating whether the subject is presumed positive or presumed negative for the disease or diseases associated with the PRS; The method of claim 1 further comprising:
13. The method of claim 12 wherein the statistical modeling uses a Hidden Markov Model.
14. 14. An apparatus for determining whether a subject is at increased risk of disease, the apparatus comprising a processor and a memory having stored thereon software instructions that, when executed by the processor, cause the apparatus to perform the steps of any one of claims 1 to 13.
15. 14. A computer program product for determining whether a subject is at increased risk of a disease, the computer program product comprising at least one non-transitory computer-readable storage medium having software instructions stored thereon, the software instructions, when executed, causing an apparatus to perform the steps of any of claims 1 to 13.