Polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BOARD OF SUPERVISORS OF LOUISIANA STATE UNIV & AGRI & MECHANICAL COLLEGE
- Filing Date
- 2026-02-02
- Publication Date
- 2026-08-06
Smart Images

Figure US2026013520_06082026_PF_FP_ABST
Abstract
Description
[0001] Docket No. 2932719-000294-W01
[0002] Filed: February 2, 2026 INTERNATIONAL PATENT APPLICATION FOR:
[0003] POLYGENIC RISK SCORE WITH KLK3 SNP-SNP INTERACTION PAIRS FOR PREDICTING PROSTATE CANCER AGGRESSIVENESS
[0004] RELATED APPLICATIONS
[0005] This application claims priority to US Application No. 63 / 752,397, filed January 31, 2025.
[0006] TECHNICAL FIELD
[0007] The disclosure herein involves systems and methods for predicting prostate cancer aggressiveness.
[0008] BACKGROUND
[0009] Prostate cancer (PCa) is heterogeneous, making risk stratification essential for clinical care. Although polygenic risk scores (PRSs) with main effects of single-nucleotide polymorphisms (SNPs) can help identify individuals at high risk before biological and clinical onset, a PRS for predicting PCa aggressiveness remains underdeveloped.
[0010] INCORPORATION BY REFERENCE
[0011] Each patent, patent application, and / or publication mentioned in this specification is herein incorporated by reference in its entirety to the same extent as if each individual patent, patent application, and / or publication was specifically and individually indicated to be incorporated by reference.
[0012] SUMMARY OF THE INVENTION
[0013] A method is described herein comprising receiving a plurality of single nucleotide polymorphism pairs (SNP-SNP pairs) identified as significantly associated with prostate cancer, applying a SNP interaction identifier (SIPI) approach to analyze the SNP-SNP pairs using a discovery dataset to identify SNP-SNP pairs that interact to influence prostate cancer in subjects, wherein the analyzing identifies a first plurality of SNP-SNP pairs, where the first plurality ofDocket No. 2932719-000294-W01
[0014] Filed: February 2, 2026 SNP-SNP pairs comprises a subset of the SNP-SNP pairs, bootstrapping the first plurality of SNP-SNP pairs using the SIPI approach to identify a second plurality of SNP-SNP pairs, wherein the second plurality of SNP-SNP pairs comprises a subset of the first plurality of SNP-SNP pairs, removing SNP-SNP pairs from the second plurality of SNP-SNP pairs to create a third plurality of SNP-SNP pairs, wherein the removing comprises dropping a SNP-SNP pair when at least one of the SNPs comprise missing data, computing risk ratio scores for each of the third plurality of SNP-SNP pairs for each subject of the discovery dataset, wherein each SNP-SNP pair comprises pair genotypes, wherein a risk ratio score of a SNP-SNP pair for a subject comprises a sum of the subject specific pair genotype ratios across all pair genotypes present in the subject, wherein a pair genotype ratio comprises a numerator equal to prevalence of prostate cancer for the pair genotype in the discovery data set and the denominator equals the prevalence of prostate cancer for all pair genotypes in the discovery dataset, and using the risk ratio scores and stepwise logistic regression to reduce size of the third plurality of SNP-SNP pairs and identify a final set of SNP-SNP pairs, wherein the stepwise logistic regression is fit using the risk ratio scores to predict significance of the plurality of SNP-SNP pairs to predict prostate cancer.
[0015] In embodiments, the stepwise logistic regression comprises removing SNP-SNP pairs from the third plurality of SNP-SNP pairs based on a p-value greater than or equal to 0.05.
[0016] In embodiments, the method includes computing final risk ratio scores for each of the final set of SNP-SNP pairs for each subject in a validation dataset, wherein each SNP-SNP pair comprises pair genotypes, wherein a final risk ratio score of a SNP-SNP pair for a subject comprises a sum of the subject specific pair genotype ratios across all pair genotypes present in the subject, wherein a pair genotype ratio comprises a numerator equal to prevalence of prostate cancer for the pair genotype in the validation data set and the denominator equals the prevalence of prostate cancer for all pair genotypes in the validation dataset.
[0017] In embodiments, the method includes fitting a logistic regression model using a summation variable comprising information of the final set of SNP-SNP pairs and a polygenic risk score as predictors, wherein the logistic regression model comprises a predictive model for predicting a probability of a subject developing prostate cancer based on a corresponding summation variable value and polygenic risk score value.
[0018] In embodiments, the summation variable comprises a sum of final risk ratio scores across all SNP-SNP pairs in the final set of SNP-SNP pairs for each subject in the validation dataset.Docket No. 2932719-000294-W01
[0019] Filed: February 2, 2026
[0020] In embodiments, the SIPI approach implements the 3pRule.
[0021] In embodiments, the 3pRule comprises a p-pair significance of a SNP-SNP pair less than 1.6*10'6.
[0022] In embodiments, the 3pRule comprises a p-pair significance of the SNP-SNP pair less than the p-value significance of a first SNP of the SNP-SNP pair.
[0023] In embodiments, the 3pRule comprises a p-pair significance of the SNP-SNP pair less than the p-value significance a second SNP of the SNP-SNP pair.
[0024] In embodiments, the 3pRule comprises a p-pair significance the SNP-SNP pair is less than the p-value significance a first and second SNP of the SNP-SNP pair.
[0025] In embodiments, the bootstrapping comprising identifying SNP-SNP pairs that are identified as significant by the SIPI approach in 75 percent or more of five hundred bootstrap runs.
[0026] In embodiments, the polygenic risk score comprises a score computed using a subset of SNPs from known polygenic risk score PHS290.
[0027] BRIEF DESCRIPTION OF FIGURES
[0028] Figure 1 shows a process of selecting SNP-SNP interaction pairs based on the UK discovery dataset. SIPI: SNP (Single Nucleotide Polymorphism) Interaction Pattern Identifier approach; 3pRule: the 3 p-value rule; RR3x3: Risk Ratio score based on 3x3 genotype pair combinations; stepwise selection is based on logistic regression.
[0029] Figures 2A-2D illustrate distributions of four polygenic risk scores for predicting prostate cancer aggressiveness.
[0030] Figures 3A-3C illustrate prevalence of prostate cancer aggressiveness by 2 polygenic risk scores’ percentiles in the 3 study sets. UK: the United Kingdom set; UCA: the
[0031] US A / Canada / Australia set; and EU: the set of other European countries. The stars (*) are for p-value<0.05 compared with the middle 50% (25-74.9%) risk group based on the two-sided Wald chi-square tests in the logistic model adjusting for the 7 principal components of population stratification. The error bars represent 95% confidence intervals. The data can be found in Table 6.
[0032] Figure 4 illustrates a risk ratio table, under an embodiment.Docket No. 2932719-000294-W01
[0033] Filed: February 2, 2026
[0034] DETAILED DESCRIPTION
[0035] The KLK3, which encodes prostate-specific antigen (PSA), is linked to PCa aggressiveness. Recent findings on KLK3 SNP-SNP interactions show promise for predicting PCa aggressiveness. The objective of this study is to develop a PRS (PRS-KLK3int) by examining KLK3 SNP-SNP interaction pairs.
[0036] The PRS-KLK3int was developed based on a discovery set (10,836 PCa patients) and two validation sets with 14,348 and 16,584 patients of European ancestry. Atotal of 3,145 SNP pairs and two published PRSs were evaluated.
[0037] This study developed a PRS-KLK3int with 284 SNPs, combining an existing PRS with 270 SNPs and 12 SNP-SNP interaction pairs with 15 SNPs (one overlapped). All these 12 pairs were involved with at least one SNP from KLK3. The PRS-KLK3int outperformed two existing PRSs in predicting PCa aggressiveness (p-values: 3.5xl0'18, 9xl0’14, and 1.7xlO’20for the three sets). It effectively distinguished high-risk from low-risk groups across all datasets. The top 1% high-risk group had a higher prevalence of PCa aggressiveness than the middle 50% group (45.5% vs. 25.9%, OR=2.38, p=2.2*10'5) in the discovery set, and similar results were observed in validation sets (OR=2.56, p=4.3><10'6; OR=2.07, p=2.1 x 1 O'3). These findings support PRS-KLK3int as a valuable tool for PCa severity stratification, especially in identifying extremely high-risk PCa patients.
[0038] Prostate cancer (PCa) can range from mild to very aggressive, so predicting severity early is important. Genetic tools called polygenic risk scores (PRS) can estimate risk before biological signs appear, helping doctors act sooner. Current PRSs don’t predict aggressiveness well. This study developed a new score, PRS-KLK3int, using genetic interactions in the KLK3 gene, which is linked to prostate-specific antigen (PSA) and cancer severity. Researchers analyzed data from over 10,000 patients and confirmed results in two large groups. PRS-KLK3int, built from 284 genetic markers, outperformed older scores in identifying aggressive cancer. PCa patients in the top 1% risk group were about twice as likely to have aggressive cancer. This tool could guide early screening and prevention before PCa progression starts.
[0039] Precise risk classification is essential for guiding treatment or monitoring decisions for prostate cancer (PCa) patients. As a heterogeneous disease, PCa accounts for 11% of all cancer deaths, making it the second leading cause of death in American men in 20241. The risk of PCaDocket No. 2932719-000294-W01
[0040] Filed: February 2, 2026 and PCa-specific mortality rates are significantly different in different racial groups, likely due to a combination of environmental, biological, and genetic factors2. While the 5-year survival rate for localized and regional PCa is nearly 100%, that of PCa patients with distant metastasis is only 34%3. Due to the substantial clinical heterogeneity, distinguishing between indolent and aggressive PCa at diagnosis remains challenging4. Depending on the nature of selected biomarkers, risk stratification tools can be utilized at different stages of PCa disease development. Most current risk stratification tools target early detection of PCa aggressiveness in the biological or clinical onset stages. D'Amico Risk Classification and the Cancer of the Prostate Risk Assessment (CAPRA) Score are based on tumor characteristics (such as prostatespecific antigen (PSA), Gleason score, and clinical stage) for early detection in the clinical onset stage of PCa aggressiveness5. Gene expression signatures of tumors for PCa progression (such as Oncotype-DX6, 7, and Decipher7) are suitable for early detection in both biological and clinical onset stages. However, few reliable risk stratification tools were reported to predict PCa aggressiveness before its biological onset9.
[0041] Genetics plays an important role in PCa development. A large-scale twin study has shown that genetics contributes to 57% of PCa development10. A handful of inherited genetic variants or single nucleotide polymorphisms (SNPs) for predicting PCa aggressiveness have been identified in genome-wide association studies (GWAS)11,12Due to the limited predictive power of individual SNPs, many studies developed polygenic risk scores (PRSs), which are primarily calculated as a weighted sum of allele dosages from multiple GWAS-identified SN Ps13’14However, these conventional PRSs do not consider SNP-SNP interactions. Despite this limitation, PRSs are increasingly recognized as valuable tools in precision medicine for various phenotypes, including PCa.
[0042] While many PRSs have been developed to assess PCa risk, few specifically address PCa aggressiveness. One PRS with 290 SNPs predicts age at diagnosis of aggressive PCa, defined as with Gleason score >7, PSA> 10 ng / mL, T3-T4 stage, nodal metastases, or distant metastases, based on -72,000 men15. Another study developed a PRS of predicting age at PCa diagnosis with 166 SNPs using data from -76,000 men16. Because age at PCa diagnosis is positively associated with PCa-specific mortality17-18, 19, these two PRSs are rel ted to PCa aggressiveness. However, these scores’ utility in predicting PCa aggressiveness may be limited as they were designed with a focus on age at PCa diagnosis rather than disease severity itself, and included both non-PCaDocket No. 2932719-000294-W01
[0043] Filed: February 2, 2026 men and PC a patients in developing these scores Additionally, their focus on SNP main effects does not account for potential SNP-SNP interactions. This presents an opportunity for improvement.
[0044] SNP-SNP interaction is considered the key to improving prediction accuracy for PRSs. Our previous study identified 3,145 SNP-SNP interaction pairs for predicting PCa aggressiveness. The majority (98.6%) of them involved KLK3'’ . KLK3 encodes the PSA, which is used for PCa screening, diagnosis, and prognostic monitoring20. Increased levels of PSA in the bloodstream may suggest the existence of PCa or other prostate-related issues21. Previous studies showed that SNPs in KLK3 were significantly associated with PSA levels, PCa risk, and aggressiveness9-22, 23, 24‘25, 26. Given the substantial role KLK3 in PCa biology and the prevalence of SNP-SNP interaction pairs relating to KLK3, we aim to develop and validate a polygenic risk score (PRS-KLK3int) that integrates these interactions The objective of this study is to assess whether this PRS-KLK3int can improve the prediction of PCa aggressiveness compared to existing PRSs that are based exclusively on SNP main effects.
[0045] Methods
[0046] The main statistical methods utilized in this study include the SNP Interaction Pattern Identifier (SIPI) approach, the RR3x3 method (Risk Ratio scoring based on 3x3 genotype pair combinations), and logistic models. The details and usage of these methods are listed in the following sections. For the study population, this study included 41,768 PCa patients with European ancestry and valid PCa aggressiveness data at the OncoArray cohort from the PRACTICAL consortium9. For this study, the SNP data in the PCa OncoArray cohort were used. For independent datasets, we excluded the PCa patients in our previous study9, which was used to identify the 3,145 SNP interaction pairs. For enhancing the reproducibility of the proposed genetic risk scores, the participants are divided into 3 independent sets: the United Kingdom (UK), USA / Canada / Australia (UCA), and other European countries (EU) sets. We used the UK set as the discovery set (n=10,836), the UCA set (n=14,348), and the EU set (n=16,584) as the two validation sets. PCa aggressiveness is defined as a Gleason score> 8, PSA >100 ng / mL, disease-distant metastasis stage at diagnosis, or death from PCa, consistent with our previous study27.
[0047] SNP data and quality controlsDocket No. 2932719-000294-W01
[0048] Filed: February 2, 2026 This study used both genotyped and imputed SNP data. The genotyped SNP data were generated using the Infinium OncoArray (Illumina, San Diego, CA). The imputed SNP data were generated based on the genotyped data using three panels of the 1,000 Genomes Project as reference. Imputation was performed using SHAPEIT2 and IMPUTE2. One-step imputation has also been performed for regions surrounding known susceptibility loci. Strict quality control was applied for the OncoArray studies. SNPs were excluded if they had a call rate <95% and departed from Hardy -Weinberg equilibrium at p<l x 1012in PCa patients. Samples were excluded with a call rate of <95%. Ancestry analysis used a principal component analysis based on 2,318 ancestry informative markers. The details of population stratification and quality control criteria for the PRACTICAL Consortium cohort were reported previously27.
[0049] Selection of SNP-SNP interaction pairs
[0050] Our previous study identified 3,145 SNP-SNP pairs significantly associated with PCa aggressiveness9. (The 3,145 SNP-SNP pairs are identified in Lin HY, et al. KLK3 SNP-SNP interactions for prediction of prostate cancer aggressiveness. Sci Rep 11, 9264 (2021), which publication is incorporated herein in its entirety.) Among them, 3121 pairs were available in the PRACTICAL OncoArray SNP dataset. The SNP Interaction Pattern Identifier (SLPI) approach was utilized to test SNP-SNP interaction pairs in this study28. SIPI is a machine learning method to test 45 interaction models for each SNP pair by considering three key factors: model structure, inheritance mode, and coding direction. In order to increase accuracy for detecting SNP-SNP interactions, the 2-stage 3pRule+bootstrap approach is suggested29. Thus, we applied the 2-stage 3pRule+bootstrap approach for selecting SNP-SNP interaction pairs in the UK set, the discovery set. A subset of 3121 candidate pairs are used in the selection method described below. The selection process of Fig. 1 is applied to filter the SNP-SNP pairs using patient data from the UK discovery dataset. In the 1ststage, we applied the 3 p-value rule (3pRule), which is a modified rule to define significant SNP-SNP interaction pairs based on (1) p-pair of SNP1-SNP2 < p-pair-criterion (such as Bonferroni correction), (2) p-pair< p-main of SNP1, and (3) p-pair< p-main of SNP2. The p-pair is the p-value for a SNP-SNP interaction pair, and the p-main is the p-value of the 1stSNP or 2ndSNP in a pair. In this study, we applied the SIPI approach to get p-pairs. For testing SNP main effects, p-main values were obtained based on the minimum p-value of the 3 inheritance modes (dominant, recessive, and additive). In the 2ndstage, the bootstrap approach was applied to create datasets by resampling with replacement and with the same sample size asDocket No. 2932719-000294-W01
[0051] Filed: February 2, 2026 the original dataset. In the UK set, the p-pair-criterion was 1.6x1 O’5(=0.05 / 3121) for the 3pRule set and was IxlO’5for the 500 bootstrap runs. The approach lowers the p-pair criterion slightly for the 500 boot strap runs. And there were 99 pairs that were identified as “significant” in 75% (or more) of the 500 runs. The “SIPI,” “eval3pRule”, and “boot3p_SIPI” R functions in the SIPI package were applied.
[0052] RR3x3 scoring system for SNP-SNP interaction pairs
[0053] For scoring SNP-SNP interaction pairs, we developed RR3x3 (Risk Ratio scoring based on 3x3 genotype pair combinations). This RR3x3 scoring especially benefits in maintaining valuable signals for high-risk genotype groups, usually with a small sample size. Each SNP has 3 potential genotypes (such as GG, GA, and AA). With a pair of SNPs, a 3-by-3 (3x3) table can be constructed with 9 genotype combinations (called pair-genotypes). While SIPI is powerful in detecting SNP-SNP interactions, the SIPI-interaction patterns simplify the 9-genotype combinations for a SNP pair into a few groups (such as 2 groups of high vs. low-risk). This feature hinders the SIPI-interaction pattern’s prediction accuracy. In order to preserve a complete risk profile of a SNP pair, we used the RR3x3 scoring approach instead. In this RR3x3 scoring method, a risk ratio of a genotype combination in a SNP pair (called pair-genotype, such as AA-GG and AG-GG) is calculated using an outcome percentage in a specific pair-genotype divided by the overall outcome percentage in the study. For this study, the outcome percentage is the overall percentage of PCa aggressiveness (aggr%) in study participants, so the RR3x3 scores were calculated based on the aggr% for each pair-genotype divided by the overall aggr%. For the pairs selected based on the 3pRule+ Bootstrap, the corresponding RR3x3 scores were calculated accordingly. Using the stepwise selection (p<0.05) in the logistic model with the candidate pairs, the 12 pairs were selected. As shown in Equation 1, PRS-12pair was calculated based on the sum of the RR3x3 -scores of these selected 12 pairs.
[0054] PRS — 12pairi = Pair^ + Pairi2+ ••■ + Pairil2
[0055] where the pair score for each individual i and SNP pair J is:
[0056]
[0057] Docket No. 2932719-000294-W01
[0058] Filed: February 2, 2026
[0059]
[0060] and where for each individual i, SNP pair j, and genotype kG (1,2,...,9), the indicator function Ii]kequals 1 if subject i possesses genotype k within SNP pair j, and 0 otherwise
[0061] (Equation 1) Two PRSs with SNP main effects only
[0062] We are interested in comparing this PRS-12 pair with the 2 published PRSs, which only involved SNP main effects without considering SNP-SNP interactions. Two published PRSs with 166 SNPs16and 290 SNPs15were selected from the PGS catalog (PGS000742 and PGSOO3331)30. PGS000742 and PGS003331 were developed based on age at PCa diagnosis and age at aggressive PCa (Gleason score >7, PSA > 10 ng / mL, T3-T4 stage, nodal metastases, or distant metastases), respectively, using time-to-event analyses with both controls and PCa cases. In these 2 published scores, 155 and 270 biallelic SNPs (PRS-ml55 and PRS-m270) were tested in this study. Note that the 270 SNPs are a subset of the 290 SNPs by excluding SNPs with three or more nucleotides. This is done because corresponding SNP pairs of the PRS-12pair further described below are biallelic. The performance of the PRSs was evaluated based on the significance level of the selected PRS in logistic regression, adjusting for the first 7 principal components of population stratification as suggested by the PRACTICAL study31. For easy comparison, PRS-12pair, PRS-ml55, and PRS-m270 were standardized. To leverage the benefits of the two PRSs, we created a PRS-KLK3int with 284 SNPs based on the weighted sum of PRS-12pair and PRS-m270. (Or rather, the PRS 12pair and the PRS-m270 combined include 284 SNPs). The weights were generated based on logistic regression.
[0063] In addition to testing numeric PRSs, we also tested categorical PRSs with 7 groups. For clinical usage, it is crucial to identify distinct high- and low-risk groups based on PRSs. Thus, we classified PCa patients into 7 groups based on PRS’s distribution in each study set. These 7 groups are the bottom 1%, bottom 1-10%, bottom 11-25%, middle 50%, top 11-25%, top 1-10%, and top 1%., which were defined based on the PRS’s percentiles of [0, 1), [1, 10), [10, 25), [25, 75), [75, 90), [90, 99), and [99, 100], An open bracket “( )” indicates the interval does not include the number itself, and a closed bracket “[ ]” includes the number itself. The significance was defined based on a p-value <0.05. The Area Under the Receiver Operating CharacteristicDocket No. 2932719-000294-W01
[0064] Filed: February 2, 2026 Curve (AUC) values for the models were calculated. Furthermore, we were interested in evaluating PRS performance by adding age at diagnosis.
[0065] PRS-KLK3int Model
[0066] PRS-m270 is generated for a subject using the following equation:
[0067] PRS-m270= biSNPi+ b2SNP2+... + b27oSNP27O
[0068] (equation 2)
[0069] In this equation, there are 270 selected SNPs. SNPk is the kth SNP with the additive inheritance mode (coding as 0, 1, 2) as the count for the selected allele (often a minor allele), and bk is the weight for the kth SNP. A similar approach will be applied to calculating PRS-ml55. If needed, we will also compare other published PRSs for PCa progression / aggressiveness. PRS-KLK3int combines PRS-12pair and PRS-m270. The PRS-12pair is the sum of 12 SNP pairs, and each pair will be scored based on the novel R3x3 scoring method described herein.
[0070] RR3X3 scoring system for a SNP-SNP interaction pair: Each SNP has 3 potential genotypes (such as GG, GA, and AA). With a pair of SNPs, a 3-by-3 (3x3) table can be conducted with 9 genotype combinations (called pair genotypes). This RR3x3 scoring especially benefits in maintaining valuable signals for high-risk genotype groups, usually with a small sample size. In this RR3x3 scoring method, a risk ratio of a genotype combination in a SNP pair (called pair-genotype, such as AA-GG and AG-GG) was calculated using an outcome percentage in a specific pair-genotype divided by the overall outcome percentage in the study. An example of an interaction of SNP1 and SNP2 associated with an outcome is shown in Figure 4. For the CC-GG pair-genotype, the outcome prevalence is 0.24, and there are 2429 subjects with this CC-GG genotype. The AA-AA pair-genotype of SNP1-SNP2 is a high-risk group because the outcome prevalence is 0.38 based on 29 subjects with this pair-genotype. The RR3x3 scores for the subjects are listed below. The RR3x3 scores are 0.923 (=0.24 / 0.26) for subjects with the CC-GG pair-genotype and 1.462 (=0.38 / 0.26) for subjects with the AA-AA pair-genotype.
[0071] The PRS-KLK3int model is expressed as follows:
[0072] PRS-KLK3int= wti PRS-12pair + wt2PRS-m270Docket No. 2932719-000294-W01
[0073] Filed: February 2, 2026 where wti and wt2are the scoring weights for PRS-12pair and PRS-m270
[0074] (Equation 3) The PRS-KLK3int comprises 284 SNPs based on the weighted sum of PRS-12pair and PRS-m270. The weights were generated based on simple logistic regression. Ultimately the method described herein fits the above referenced model using simple logistic regression to produce scoring weights wti and wt2.
[0075] Results
[0076] Detection of SNP-SNP interaction pairs for PRS-12pair
[0077] The procedure of selecting SNP pairs based on the UK set for PRS is listed in Fig. 1. Among the 3,121 candidate SNP pairs, the 2-stage 3pRule+ bootstrap approach was applied to improve detection accuracy. By applying the 3pRule with p-pair< 1.6x1 O'5(=0.05 / 3121, Bonferroni correction) as the p-pair criterion, 613 pairs satisfied the 3pRule. In the 2ndstage, 99 pairs were satisfied with the 3pRule in >75% of 500 bootstrap datasets. These 99 pairs were grouped into the 6 clusters with 6 KLK3 SNPs as a hub. The cluster is defined as several SNP pairs sharing a common or hub SNP. Among these 99 pairs, 50, 37, 15, 6, 3, and 2 pairs had rsl7632542, rs2569735, rsl058205, rsl74776, rs2292185, and rs266876 as the hub, respectively. Some pairs are involved with 2 KLK3 SNPs. For example, there were 50 pairs involved with rsl7632542 (such as rsl7632542-SNPl, rsl7632542-SNP2, ..., rsl7632542-SNP50). In order to use more complete data for generating the scores, we excluded 16 pairs, which had missing values due to missing information for at least one SNP in the pair. For some individuals for a given SNP pair, the dataset may include data for SNP1 but does not include data for SNP2 so the RR3x3 score may not be computed for this pair. Then, we generated the RR3X3-scores for the remaining 83 pairs for further analyses. Using the RR3x3 score method, we got one score value for each pair. Thus, there were 83 scores for 83 pairs. After applying for stepwise selection in the logistic model with the outcome of prostate cancer aggressiveness in the UK dataset, 12 pairs with a p<0.05 were selected. The method computes the RR3x3 values of all 83 pairs for each patient in the entire UK or discovery dataset (i.e. in the 10,836 UK dataset) and regresses 83 pairs against presence of pCA in the discovery dataset. All SNP-SNP pairs that were greater than .05 in this single step are removed.Docket No. 2932719-000294-W01
[0078] Filed: February 2, 2026 Finally, PRS-12pair was calculated based on the sum of the RR3x3-scores of these 12
[0079] pairs. All 12 pairs were involved with KLK3 SNPs, and consist of only 15 SNPs. These 12 pairs and their gene information are listed in Table 1.
[0080] Table • 1 List of 12 SNP-SNP interaction pairs (withl5 SNPs) and their gene information in the
[0081] UK discovery.
[0082]
[0083] aMajor / minor allele; MAF; minor allele frequency
[0084] bChr: chromosome; position is based on the Genome Reference Consortium HumanBuild 37 (GRCh37
[0085] General results of 4 polygenic risk scores (PRSs) for PCa aggressiveness
[0086] The descriptive statistics of the 4 PRSs (PRS-12pair, PRS-ml55, PRS-m270, PRS- KLK3int) are listed in Table 2, and their distributions are shown in Supplementary Fig. 2.
[0087] Table -2 Descriptive statistics of the polygenic risk scores (PRS) with main effects only and their associations with prostate cancer aggressiveness by 3 study sets
[0088] "
[0089] "
[0090]
[0091] aThe United Kingdom (UK), USA / Canada / Australia (UCA), and other European countries (EU) sets
[0092] bStandard deviation; min: minimum; max: maximum
[0093] codds ratio and 95% confidence interval are per 1 standard deviation of the score. POvalues were based on the two-sided Wald chi-square tests in the logistic models.
[0094] dadjusted for the first 7 principal components of population stratificationDocket No. 2932719-000294-W01 Filed: February 2, 2026
[0095] The distributions of the 2 PRS-main scores were close to a normal distribution. However, the distribution of PRS-12pair was skewed to the right (positively skewed), indicating that some men had a considerably high score of PCa aggressiveness. This is a desired feature for PRS to identify extremely high-risk groups. However, a small variation was observed for the low-value part of PRS-12pair; therefore, PRS-12pair may not be an ideal tool for distinguishing low-risk aggressive groups. These results inspired us to create a combined score, PRS-KLK3int, by integrating PRS-12pair and PRS-m270 to take advantage of both scores.
[0096] Comparison of PRS-12pair with the 2 PRSs with SNP main effects only
[0097] For the 2 published PRSs, the associations between these 2 PRSs with SNP main effects only and PCa aggressiveness are listed in Table 3.
[0098] Table 3. Performance of polygenic risk scores (PRS) for predicting prostate cancer aggressiveness by study sites
[0099]
[0100] aThe United Kingdom (UK), USA / Canada / Australia (UCA), and other European countries (EU) sets
[0101] bAdjusted odds ratio (95% confidence interval) per 1 standard deviation generated based on a model with a PRS and the 7 principal components of population stratification. P-values were based on the two-sided Wald chi-square tests in the logistic models
[0102] PRS-ml55 and PRS-m270 can predict PCa aggressiveness for all 3 data sets (all p-values <0.05). Specifically, we observed that PRS-m270 performed better than PRS-ml55 for all 3 data sets, and these 2 PRSs performed better in the UK and EU sets than the UCA sets. Based on the adjusted results, the p-values for PRS-ml55 were 7.5xl0'4, 4.4xl0'4, and 4.8xl0'8for the UK, UCA, and EU sets, respectively, while the p-values for PRS-m270 were 8.2xl0-7, 3.3xl0-5, and 9.4x10’11, respectively. In addition, the performance of PRS-12pair was better in terms of significance level than PRS-ml55 and PRS-m270 in all 3 sets (Table 3). The p-values for PRS-12pair were 5.4xl0‘15, 4.0xl0‘8, and 1.8xl0’13for the UK, UCA, and EU sets, respectively.
[0103] We tested whether PRS-12pair can further improve the performance of the PRS-m270. In the model with PRS-12pair and PRS-m270 (Table 4), PRS-12pair performed better than PRS-m270 in all 3 sets. For the UK set, the p-values for PRS-m270 and PRS- 12pair were 1.4xl0‘4andDocket No. 2932719-000294-W01 Filed: February 2, 2026 2.9x1 O'12. Similar conditions were observed in the UCA and EU sets. After adjusting PRS-m270 and the 7 principal components, PRS-12pair can still better predict PCa aggressiveness than PRS-m270 (p=l.OxlO'10and 6.4x1 O'11for the UCA and EU sets, respectively). Further, PRS-m270was still significant (p=1.9xl0'2and 9.6x1 O'8for the UCA and EU sets, respectively).
[0104] Table 4, Performance of combined PRS for predicting prostate cancer aggressiveness
[0105]
[0106] aThe United Kingdom (UK), USA / Canada / Australia (UCA), and other European countries (EU) sets
[0107] bAdjusted odds ratio (95% confidence interval) per 1 standard deviation. All models were adjusted for the 7 principal components of population stratification. P-values were based on the two-sided Wald chi-square tests in the logistic models.cPRS-KLK3int is a score which combines PRS-12pair and PRS-m270
[0108] Performance evaluation of PRS-KLK3int (continuous score version)
[0109] The performance of PRS-KLK3int is shown in Table 4. As expected, the PRS-KLK3int performed better than PRS-12pair and PRS-m270 alone for predicting PCa aggressiveness. The p-values for PRS-KLK3int were 3.5xl0’18, 9xl0'14, and 1.7xlO'20for the UK, UCA, and EU sets, respectively. As shown in Table 5, the AUC values of PRS-KLK3int were 0.55, 0.58, and 0.59 after adjusting for the 7 principal components for the UK, UCA, and EU sets, respectively.
[0110] Significant improvements were made by adding the 12 pairs into the PRS-m270 in all 3 sets. The p-values of the AUC comparison between PRS-KLK3int and PRS-m270 were 7.3xl0’5, 0.004, and 0.051 for the UK, UCA, and EU sets, respectively.
[0111] Table 5. Area Under the Receiver Operating Characteristic Curve (AUC) values of the polygenic risk scores (PRSs)
[0112]
[0113] aThe United Kingdom (UK), USA / Canada / Australia (UCA), and other European countries (EU) sets.
[0114] bP-values were based on the two-sided DeEong testDocket No. 2932719-000294-W01
[0115] Filed: February 2, 2026
[0116] Performance evaluation of PRS-KLK3int (risk-group version)
[0117] We further identified distinct high- and low-risk groups based on PRS-KLK3int, PRS-m270, and PRS-ml55. For each study set, the PCa patients were classified into 7 groups based on PRS’s distribution. These 7 groups are bottom 1%, bottom 1-10%, bottom 11-25%, middle 50% (25-75%), top 11-25%, top 1-10%, and top 1%. PRS-KLK3int can distinguish high-risk PCa aggressiveness better than PRS-m270, and PRS-ml55. The performance of risk classification of PCa aggressiveness for the UK, US, and EU sets in terms of PRS percentiles is listed in Fig. 3 and Table 6.
[0118] Table 6. Prostate cancer aggressiveness proportions by the percentiles of PRS-KLK3int and PRS-m270.
[0119]
[0120] aThe United Kingdom (UK), USA / Canada / Australia (UCA), and other European countries (EU) sets
[0121] bN: sample size in the group, aggr%: percentage of prostate cancer aggressiveness
[0122] cBold for p-value<0.05 were based on the two-sided Wald chi-square tests in the logistic models. Adjusted odds ratio (95% confidence interval) with a reference group using the middle 50% PRS. All models had a PRS and the 7 principal components of population stratification.
[0123] In the UK set, PRS-KLK3int can significantly distinguish the top 1% and top 1-10% from the majority of the middle 50% group (see Table 7 as well as Table 6 and 8).
[0124] Table 7. Comparisons of genetic-guided risk groups based on the polygenic risk scores (PRSs) for the 3 study sets. (aThe United Kingdom (UK), USA / Canada / Australia (UCA), and other European countries (EU) sets.bAdjusted odds ratio (95% confidence interval) with aDocket No. 2932719-000294-W01
[0125] Filed: February 2, 2026 reference group using the middle 50% group. All models had a PRS and the 7 principal components of population stratification. Significant results with p-value<0.05 were bold).
[0126] Table 7. Comparisons of genetic-guided risk groups based on the polygenic risk scores (PRSs) for the 3 study sets. (aThe United Kingdom (UK), US A / Canada / Australia (UCA), and other European countries (EU) sets.bAdjusted odds ratio (95% confidence interval) with a reference group using the middle 50% group. All models had a PRS and the 7 principal components of population stratification. Significant results with p-value<0.05 were bold).
[0127]
[0128] Table 8. Prostate cancer aggressiveness proportions by the percentiles of PRS-ml55.
[0129]
[0130] Docket No. 2932719-000294-W01
[0131] Filed: February 2, 2026
[0132]
[0133] aThe United Kingdom (UK), USA / Canada / Australia (UCA), and other European countries (EU) sets
[0134] bN: sample size in the group, aggr%: percentage of prostate cancer aggressiveness
[0135] cBold for p-value<0.05 based on Hie two-sided Wald chi-square tests in the logistic models. Adjusted odds ratio (95% confidence interval) with a reference group using the middle 50% PRS. All models had a PRS and the 7 principal components of population stratification.
[0136] Based on PRS-KLK3int, the prevalence of PCa aggressiveness for the middle 50% group was 25.9%. PCa aggressiveness prevalence for the top 1% group (45.5%) was significantly higher than that in the middle 50% group (OR=2.38, p= 2.2xl0'5). This highlights that this PRS-KLk3int is able to categorize the top 1% of PCa patients at approximately 2-fold higher risk compared with the normal risk group. The top 1-10% group also had a significantly higher PCa aggressiveness prevalence than those in the middle 50% group (32.8% vs. 25.9%, OR=1.4, p= 2.2x1 O'3). For detecting the low-risk groups, PRS-KLK3int did not perform well in distinguishing the extremely low score group (bottom 1%), but it can significantly distinguish the bottom 10% (OR=0.81, p=0.010). For the UK set, PRS-ml55 and PRS-m270 cannot distinguish the top 1%, top 1-10%, and bottom 1% compared with the middle 50% group. However, PRS-ml55 can distinguish the bottom 1-10% vs the middle 50% (OR=0.81, p=0.012), but PRS-m270 cannot.
[0137] The performance of PRS-KLK3int in terms of PCa aggressiveness risk classification in the UCA and EU sets is similar to the UK set. For the UCA and EU sets, PRS-KLK3int can significantly distinguish the top 1% and top 1-10% from the majority of the middle 50% group. For PRS-KLK3int in the UCA set, PCa aggressiveness prevalence was 26.2%, 16.9%, and 12.3% for those in the top 1%, top 1-10% and middle 50% groups, respectively (Table 6). PCa aggressiveness prevalence for the top 1% and top 1-10% groups was significantly higher than that in the middle 50% group (OR=2.56, p= 4.3xl0'6for top 1%; OR.~1.45, p= 2.5xl0'5for top 1-10%). For PRS-KLK3int in the EU set, PCa aggressiveness prevalence was 37.8%, 29.3%, and 23.1% for those in the top 1%, top 1-10%, and middle 50% groups, respectively. PCa aggressiveness prevalence for the top 1% and top 1-10% groups was significantly higher than those in the middle 50% group (OR=2.07, p= 2. IxlO'5for top 1%; OR=1.38, p= 1. IxlO'6for top 1-10%). Similarly to the UK set, PRS-m270 and PRS-ml55 performance was inferior to PRS-KLK3int in the UCA and EU sets. For the UCA set, PRS-m270 could distinguish the top 1% andDocket No. 2932719-000294-W01
[0138] Filed: February 2, 2026 top 1-10% versus the middle 50% for PCa aggressiveness (OR=2.17, p=1.2xl0'4; OR=1.22, p=0.025, respectively). This similar trend of PRS-m270 can be observed in the EU set. PCa aggressiveness prevalence was significantly higher for the top 1% and top 1-10% compared with the middle 50% (OR=1.53, p=0.013; OR=1.23, p=0.001, respectively). ForPRS-ml55, PCa aggressiveness prevalence was significantly higher for the top 1% (OR=2.06, p=0.0004) in the UCA set and was significantly lower for the bottom 1% (OR=0.41, p=0.0003) in the EU set compared with the middle 50%. These results demonstrate that adding 12 pairs to PRS-m270 can significantly increase prediction accuracy for PCa aggressiveness, especially for the extremely high-risk group.
[0139] PRS-KLK3int and age at PCa diagnosis for predicting PCa aggressiveness
[0140] We evaluated PRS-KLK3int performance by adding age at PCa diagnosis, a known risk factor of PCa aggressiveness, in models. As shown in Table 9, age at diagnosis is positively associated with PCa aggressiveness for all 3 sets (UK: p=5.0xl0‘7; UCA: p=8.1xl0’46; EU: p=4.5xl0‘116). The age at PCa diagnosis was positively correlated with the risk of PCa aggressiveness. After adjusting age at diagnosis in the models, PRS-KLK3int remained highly significant (UK: p=8.0xl0'17; UCA: p=2.7xl0‘8; EU: p=5.8xl0'14). Adding age at PCa diagnosis to the models increased the AUC values significantly for the UCA and EU sets but not the UK set (Table 5). The AUC values were 0.56, 0.64, and 0.65 for the UK, UCA, and EU, respectively. Table 9. Model results of PRS-KLK3int by adjusting for age at prostate cancer diagnosis
[0141]
[0142] aThe United Kingdom (UK), USA / Canada / Australia (UCA), and other European countries (EU) sets
[0143] 11Adjusted odds ratio (95% confidence interval) per 1 standard deviation. Model adjusted for the first 7 principal components of population stratification. P-values were based on the two-sided Wald chi-square tests in the logistic models.
[0144] Discussion
[0145] In this study, we developed PRS-KLK3int, could significantly predict PCa aggressiveness susceptibility more accurately than the other 3 PRSs alone (PRS-12pair, PRS-m270, and PRS-ml55) when we evaluated the 3 large study sets (UK, UCA, and EU). To our knowledge, this is the first validated PRS involved SNP-SNP interactions for PCa aggressiveness. The major component of this score, PRS-12pair, contains 12 KLK3 interaction pairs consist ofDocket No. 2932719-000294-W01
[0146] Filed: February 2, 2026 only 15 SNPs. Notably, PRS-12pair outperformed the two other PRSs (PRS-ml55 andPRS-m270) for predicting PCa aggressiveness. The PRS-KLK3int can improve distinguishing PCa patients with an extremely high risk (the top 1% and top 1-10%) and bottom 10% compared with the middle-risk group of PCa aggressiveness. The p-values of PRS-KLK3int for comparing the top 1% with the middle-risk group are in a range of 4.3xl0'6to 2.2xl0'5for the 3 study sets. However, PRS-KLK3int did not perform well in differentiating the extremely low-risk group (bottom 1%) but can significantly predict the low-risk group in the bottom 10% based on its distribution. These findings demonstrated thatPRS-KLK3int combines the advantages of both scores of PRS-12pair and PRS-m270, allowing it to differentiate between high- and low-risk groups of PCa aggressiveness.
[0147] The PRS-12 pair contains 15 SNPs in 13 genes, including KLK3. Among them, there are 3 KLK3 SNPs (rsl7632542, rs2569735, and rsl74776). The genes linked to the 12 KLK3 SNP pairs exhibit biological functions associated with PCa aggressiveness. KLK3 is a gene that codes for PSA, which is a serine protease only secreted by the prostate gland32. In previous GWA studies, rsl 7632542 in KLK3 was associated with various PCa phenotypes, including PSA level, PCa risk, aggressiveness, and Gleason score22-33. In addition, SNP-SNP interactions involving several KLK3 SNPs (including rsl7632542, rs2569735, and rsl74776) can significantly predict PCa aggressiveness9. A meta-analysis study reported a significant association between rsl74776 and PCa risk34. Among other genes in PRS-12pair, most genes are involved in PCa and other cancers, such as SMAD3, CCND1, EPHB1, DEPTOR, JAZF1, PDCD6IP, and TPMELACTB. SMAD3 is overexpressed in advanced prostate tumors and regulates the expression of enzymes involved in the angiogenesis pathway33, 36. Moreover, rs35874463 in SMAD3 can predict height37, which is one of the known risk factors for PCa aggressiveness38. CCND1 and KLK3 are also included in the 8-gene expression signature to predict the survival of PCa patients. This gene signature is based on circulating tumor cells, and its performance is better than most clinical variables39. EPHB1 encodes a receptor tyrosine kinase involved in cell growth, differentiation, and apoptosis40, and its overexpression in prostate tumors suggests it may act as a potential oncogene41. In addition, expression of DEPTOR in prostate tumors was downregulated and correlated positively with PCa progression42. Low expression of DEPTOR promotes cell proliferation, survival, migration, and invasion, supporting its role as a potential tumor suppressor43, 44. Regarding JAZF1, it is associated with PCa cell growth and response toDocket No. 2932719-000294-W01
[0148] Filed: February 2, 2026 polyamine-targeted therapies45. Moreover, the JAZF1 rs849140 variant correlates with height46, a known risk factor for prostate cancer risk and aggressiveness. As for PDCD6IP, previous studies reported a positive correlation between the expression of PDCD6IP and Gleason Score48.
[0149] PDCD6IP expression was significantly different in prostate malignant tissues when compared to indolent and normal prostate tissues49. rs2729786 is located in the intergenic region between TPM1 and LACTB. TPM1 was identified as one of 27 genes for a PCa progression gene signature50. This evidence highlights the potential biological role of genes linked to the 12 KLK3 SNP pairs in PCa aggressiveness.
[0150] Although both PRS-m270 and PRS-ml55 are significantly associated with PCa aggressiveness, it is important to discuss why these two PRSs showed a protective effect (OR<1.0) on PCa aggressiveness. Several large-scale longitudinal studies have shown that men diagnosed with PCa at an older age have a higher risk of PCa-speciftc mortality, a terminal phenotype of PCa aggressiveness17, 18, 19. Our study results also indicated that age at PCa diagnosis is positively associated with PCa aggressiveness across all three sets. Consequently, a higher score of PRS-ml55 and PRS-m270 suggests a higher chance of a younger age at PCa diagnosis, which may correspond to a lower risk of PCa aggressiveness.
[0151] We developed a promising PRS-KLK3int by integrating 12 KLK3 SNP-SNP interaction pairs, which can be used in clinical settings as a valuable adjunct beyond existing tools primarily based on tumor gene expression (Decipher, Prolaris, and Oncotype-DX). This is because PRS-KLK3int can be used to classify high-risk PCa patients before biological onset, which will be beneficial for a large group of PCa patients under active surveillance. It has been shown that approximately 60% of low-risk PCa patients are under active surveillance53. Moreover, genetic scores can also play a crucial role in alleviating PCa patients' anxiety stemming from the uncertainty of long-term PCa outcomes. Studies have shown that PCa patients’ anxiety and depression are common and are significantly associated with poor prognosis (such as PC progression and survival), even after treatments based on longitudinal studies54, 53. The strength of this study lies in its rigorous study design and large discovery and validation datasets, which ensure the reliability and validity of the PRS findings. As shown in Fig. 1, the development of the PRS-12pair was through a solid process, including (1) the selection of candidate 3,145 SNP pairs based on our previous solid study9, (2) application of SIPI for testing SNP-SNP interactions28, (3) the 3pRule+bootstrap SNP pair selection process to reduce false positivity andDocket No. 2932719-000294-W01
[0152] Filed: February 2, 2026 (4) the RR3x3 scoring system based on the risk of 9 genotype combinations of a SNP pair. The processes 2-4 were developed by our team.
[0153] While these study findings are promising, we are aware that there are some limitations. First, PRS-KLK3int, developed in this study to predict PCa aggressiveness, was derived from PCa patients with European ancestry. Thus, its applicability to other racial and ethnic groups remains uncertain and warrants further investigation. Notably, men with African ancestry are known to have a higher risk of incidence, aggressiveness, and poor outcomes than those with European ancestry56, 57, underscoring the importance of validating and potentially adapting this genetic score in diverse populations. Second, the 12 SNP- SNP interaction pairs were selected based on SNP-SNP interactions using candidate genes in 4 major pathways9. A broader or genome-wide search level will be necessary to enhance the PRS of PCa aggressiveness. Third, while PRS-KLK3int shows promising results in terms of statistical significance, its AUC performance remains limited, indicating that it is not suitable for standalone use. However, it can be added to the existing tools to improve prediction accuracy. Finally, this study focused on genetic profiles without considering other factors, such as gene expression; therefore, its application in clinical settings needs further evaluation.
[0154] In conclusion, our findings highlight the PRS-KLK3int as a more effective tool for predicting PCa aggressiveness. This PRS supports risk stratification and personalized care to improve patient outcomes. In summary, PRS-KLK3 interaction effectively leverages interaction effects among biologically relevant SNPs to improve risk prediction for PCa aggressiveness Its unique ability to predict PCa aggressiveness or progression before biological onset offers a valuable complement to existing tools in clinical settings. By refining genetic risk stratification, this model has the potential to support early decision-making in clinical practice. Future studies should explore different models of genome-wide interactions and extend validation to various populations to enhance applicability.
[0155] Computer networks suitable for use with the embodiments described herein include local area networks (LAN), wide area networks (WAN), Internet, or other connection services and network variations such as the world wide web, the public internet, a private internet, a private computer network, a public network, a mobile network, a cellular network, a value-added network, and the like. Computing devices coupled or connected to the network may be any microprocessor controlled device that permits access to the network, including terminal devices, such as personalDocket No. 2932719-000294-W01
[0156] Filed: February 2, 2026 computers, workstations, servers, minicomputers, main-frame computers, laptop computers, mobile computers, palm top computers, hand held computers, mobile phones, TV set-top boxes, or combinations thereof. The computer network may include one of more LANs, WANs, Internets, and computers. The computers may serve as servers, clients, or a combination thereof.
[0157] The polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness can be a component of a single system, multiple systems, and / or geographically separate systems. The polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness can also be a subcomponent or subsystem of a single system, multiple systems, and / or geographically separate systems. The components of polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness can be coupled to one or more other components (not shown) of a host system or a system coupled to the host system.
[0158] One or more components of the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness and / or a corresponding interface, system or application to which the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness is coupled or connected includes and / or runs under and / or in association with a processing system. The processing system includes any collection of processorbased devices or computing devices operating together, or components of processing systems or devices, as is known in the art. For example, the processing system can include one or more of a portable computer, portable communication device operating in a communication network, and / or a network server. The portable computer can be any of a number and / or combination of devices selected from among personal computers, personal digital assistants, portable computing devices, and portable communication devices, but is not so limited. The processing system can include components within a larger computer system.
[0159] The processing system of an embodiment includes at least one processor and at least one memory device or subsystem. The processing system can also include or be coupled to at least one database. The term “processor” as generally used herein refers to any logic processing unit, such as one or more central processing units (CPUs), digital signal processors (DSPs), applicationspecific integrated circuits (ASIC), etc. The processor and memory can be monolithically integrated onto a single chip, distributed among a number of chips or components, and / or provided by some combination of algorithms. The methods described herein can be implemented in one orDocket No. 2932719-000294-W01
[0160] Filed: February 2, 2026 more of software algorithm(s), programs, firmware, hardware, components, circuitry, in any combination.
[0161] The components of any system that include the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness can be located together or in separate locations. Communication paths couple the components and include any medium for communicating or transferring files among the components. The communication paths include wireless connections, wired connections, and hybrid wireless / wired connections. The communication paths also include couplings or connections to networks including local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), proprietary networks, interoffice or backend networks, and the Internet. Furthermore, the communication paths include removable fixed mediums like floppy disks, hard disk drives, and CD-ROM disks, as well as flash RAM, Universal Serial Bus (USB) connections, RS-232 connections, telephone lines, buses, and electronic mail messages.
[0162] Aspects of the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness and corresponding systems and methods described herein may be implemented as functionality programmed into any of a variety of circuitry, including programmable logic devices (PLDs), such as field programmable gate arrays (FPGAs), programmable array logic (PAL) devices, electrically programmable logic and memory devices and standard cell-based devices, as well as application specific integrated circuits (ASICs). Some other possibilities for implementing aspects of the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness and corresponding systems and methods include: microcontrollers with memory (such as electronically erasable programmable read only memory (EEPROM)), embedded microprocessors, firmware, software, etc. Furthermore, aspects of the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness and corresponding systems and methods may be embodied in microprocessors having software-based circuit emulation, discrete logic (sequential and combinatorial), custom devices, fuzzy (neural) logic, quantum devices, and hybrids of any of the above device types. Of course the underlying device technologies may be provided in a variety of component types, e.g., metal-oxide semiconductor field-effect transistor (MOSFET) technologies like complementary metal-oxide semiconductor (CMOS), bipolar technologies likeDocket No. 2932719-000294-W01
[0163] Filed: February 2, 2026 emitter-coupled logic (ECL), polymer technologies (e.g., silicon-conjugated polymer and metal-conjugated polymer-metal structures), mixed analog and digital, etc.
[0164] It should be noted that any system, method, and / or other components disclosed herein may be described using computer aided design tools and expressed (or represented), as data and / or instructions embodied in various computer-readable media, in terms of their behavioral, register transfer, logic component, transistor, layout geometries, and / or other characteristics. Computer-readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, non-volatile storage media in various forms (e.g., optical, magnetic or semiconductor storage media) and carrier waves that may be used to transfer such formatted data and / or instructions through wireless, optical, or wired signaling media or any combination thereof. Examples of transfers of such formatted data and / or instructions by carrier waves include, but are not limited to, transfers (uploads, downloads, e-mail, etc.) over the Internet and / or other computer networks via one or more data transfer protocols (e.g., HTTP, FTP, SMTP, etc.). When received within a computer system via one or more computer-readable media, such data and / or instructionbased expressions of the above described components may be processed by a processing entity (e g., one or more processors) within the computer system in conjunction with execution of one or more other computer programs.
[0165] Unless the context clearly requires otherwise, throughout the description and the claims, the words “comprise,” “comprising,” and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in a sense of “including, but not limited to.” Words using the singular or plural number also include the plural or singular number respectively. Additionally, the words “herein,” “hereunder,” “above,” “below,” and words of similar import, when used in this application, refer to this application as a whole and not to any particular portions of this application. When the word “or” is used in reference to a list of two or more items, that word covers all of the following interpretations of the word: any of the items in the list, all of the items in the list and any combination of the items in the list.
[0166] The above description of embodiments of the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness is not intended to be exhaustive or to limit the systems and methods to the precise forms disclosed. While specific embodiments of, and examples for, the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness and corresponding systems and methods are described herein forDocket No. 2932719-000294-W01
[0167] Filed: February 2, 2026 illustrative purposes, various equivalent modifications are possible within the scope of the systems and methods, as those skilled in the relevant art will recognize. The teachings of the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness and corresponding systems and methods provided herein can be applied to other systems and methods, not only for the systems and methods described above.
[0168] The elements and acts of the various embodiments described above can be combined to provide further embodiments. These and other changes can be made to the polygenic risk score with KLK3 SNP-SNP interaction pairs for predicting prostate cancer aggressiveness and corresponding systems and methods in light of the above detailed description.
[0169] References
[0170] 1. Siegel RL, Giaquinto AN, Jemal A. Cancer statistics, 2024. CA Cancer J Clin 74, 12-49 (2024).
[0171] 2. Berglund A, et al. Epigenome-wide association study of prostate cancer in African American men identified differentially methylated genes. Cancer Med 13, e70044 (2024).
[0172] 3. American Cancer Society. Cancer Facts & Figures 2024 In: Survival Rates for Prostate Cancer). Atlanta, American Cancer Society (2024).
[0173] 4. Ye H, Sowalsky AG. Molecular correlates of intermediate- and high-risk localized prostate cancer. Urol Oncol 36, 368-374 (2018).
[0174] 5. May M, et al. Validity of the CAPRA score to predict biochemical recurrence-free survival after radical prostatectomy. Results from a european multi center survey of 1,296 patients. J Urol 178, 1957-1962; discussion 1962 (2007).
[0175] 6. Klein EA, et al. A 17-gene assay to predict prostate cancer aggressiveness in the context of Gleason grade heterogeneity, tumor multifocality, and biopsy undersampling. Eur Urol 66, 550-560 (2014).Docket No. 2932719-000294-W01
[0176] Filed: February 2, 2026 7. Cui F, Tang X, Man C, Fan Y. Prognostic value of 17-Gene genomic prostate score in patients with clinically localized prostate cancer: a meta-analysis. BMC Cancer 24, 628 (2024).
[0177] 8. Erho N, et al. Discovery and validation of a prostate cancer genomic classifier that predicts early metastasis following radical prostatectomy. PLoS One 8, e66855 (2013).
[0178] 9. Lin HY, et al. KLK3 SNP-SNP interactions for prediction of prostate cancer aggressiveness. Sci Rep 11, 9264 (2021).
[0179] 10. Mucci LA, et al. Familial Risk and Heritability of Cancer Among Twins in Nordic Countries. JAMA 315, 68-76 (2016).
[0180] 11. Berndt SI, et al. Two susceptibility loci identified for prostate cancer aggressiveness. Nat Commun 6, 6889 (2015).
[0181] 12. Amin Al Olama A, et al. A meta-analysis of genome-wide association studies to identify prostate cancer susceptibility loci associated with aggressive and non-aggressive disease. Hum Mol Genet 22, 408-415 (2013).
[0182] 13. Lewis CM, Vassos E. Polygenic risk scores: from research tools to clinical instruments. Genome Med 12, 44 (2020).
[0183] 14. Choi SW, Mak TS, O'Reilly PF. Tutorial: a guide to performing polygenic risk score analyses. NatProtoc 15, 2759-2772 (2020).
[0184] 15. Huynh-Le MP, et al. Prostate cancer risk stratification improvement across multiple ancestries with new polygenic hazard score. Prostate Cancer Prostatic Dis 25, 755-761 (2022).
[0185] 16. Karunamuni RA, et al. Additional SNPs improve risk stratification of a polygenic hazard score for prostate cancer. Prostate Cancer Prostatic Dis, (2021).Docket No. 2932719-000294-W01
[0186] Filed: February 2, 2026 17. Clark R, Vesprini D, Narod SA. The Effect of Age on Prostate Cancer Survival. Cancers (Basel) 14, (2022).
[0187] 18. Pettersson A, Robinson D, Garmo H, Holmberg L, Stattin P. Age at diagnosis and prostate cancer treatment and prognosis: a population-based cohort study. Ann Oncol 29, 377-385 (2018).
[0188] 19. Bernard B, Burnett C, Sweeney CJ, Rider JR, Sridhar SS. Impact of age at diagnosis of de novo metastatic prostate cancer on survival. Cancer 126, 986-993 (2020).
[0189] 20. Lawrence MG, Lai J, Clements JA. Kallikreins on steroids: structure, function, and hormonal regulation of prostate-specific antigen and the extended kallikrein locus. Endocr Rev 31, 407-446 (2010).
[0190] 21. Penney KL, et al. Association of KLK3 (PSA) genetic variants with prostate cancer risk and PSA levels. Carcinogenesis 32, 853-859 (2011).
[0191] 22. Sullivan J, et al. An analysis of the association between prostate cancer risk loci, PSA levels, disease aggressiveness and disease-specific mortality. Br J Cancer 113, 166-172 (2015).
[0192] 23. Kotarac N, Dobrijevic Z, Matijasevic S, Savic-Pavicevic D, Brajuskovic G. Association of KLK3, VAMP8 and MDM4 Genetic Variants within microRNA Binding Sites with Prostate Cancer: Evidence from Serbian Population. Pathol Oncol Res 26, 2409-2423 (2020).
[0193] 24. Chen C, Xin Z. Single-nucleotide polymorphism rs!058205 of KLK3 is associated with the risk of prostate cancer: A case-control study of Han Chinese men in Northeast China.
[0194] Medicine (Baltimore) 96, e6280 (2017).
[0195] 25. Helfand BT, et al. Associations of prostate cancer risk variants with disease aggressiveness: results of the NCI-SPORE Genetics Working Group analysis of 18,343 cases. Hum Genet 134, 439-450 (2015).Docket No. 2932719-000294-W01
[0196] Filed: February 2, 2026
[0197] 26. Batra J, O'Mara T, Patnala R, Lose F, Clements JA. Genetic polymorphisms in the human tissue kallikrein (KLK) locus and their implication in various malignant and non-malignant diseases. Biol Chem 393, 1365-1390 (2012).
[0198] 27. Amos CI, et al. The OncoArray Consortium: A Network for Understanding the Genetic Architecture of Common Cancers. Cancer Epidemiol Biomarkers Prev 26, 126-135 (2017).
[0199] 28. Lin HY, et al. SNP interaction pattern identifier (SIPI): an intensive search for SNP-SNP interaction patterns. Bioinformatics 33, 822-833 (2017).
[0200] 29. Lin HY, et al. Cluster effect for SNP-SNP interaction pairs for predicting complex traits. Sci Rep 14, 18677 (2024).
[0201] 30. Lambert SA, et al. The Polygenic Score Catalog as an open database for reproducibility and systematic evaluation. Nat Genet 53, 420-425 (2021).
[0202] 31. Schumacher FR, et al. Association analyses of more than 140,000 men identify 63 new prostate cancer susceptibility loci. Nat Genet 50, 928-936 (2018).
[0203] 32. Sutherland GR et al. Human prostate-specific antigen (APS) is a member of the glandular kallikrein gene family at 19q 13. Cytogenet Cell Genet 48, 205-207 (1988).
[0204] 33. Knipe DW, et al. Genetic variation in prostate-specific antigen-detected prostate cancer and the effect of control selection on genetic association studies. Cancer Epidemiol Biomarkers Prev 23, 1356-1365 (2014).
[0205] 34. Li H, Fei X, Shen Y, Wu Z. Association of gene polymorphisms of KLK3 and prostate cancer: A meta-analysis. Adv Clin Exp Med 29, 1001-1009 (2020).Docket No. 2932719-000294-W01
[0206] Filed: February 2, 2026 35. Lu S, Lee J, Revelo M, Wang X, Lu S, Dong Z. Smad3 is overexpressed in advanced human prostate cancer and necessary for progressive growth of prostate cancer cells in nude mice. Clin Cancer Res 13, 5692-5702 (2007).
[0207] 36. Jeon HY, et al. SMAD3 promotes expression and activity of the androgen receptor in prostate cancer. Nucleic Acids Res 51, 2655-2670 (2023).
[0208] 37. Lanktree MB, et al. Meta-analysis of Dense Genecentric Association Studies Reveals Common and Uncommon Variants Associated with Height. Am J Hum Genet 88, 6-18 (2011).
[0209] 38. Lophatananon A, et al. Height, selected genetic markers and prostate cancer risk: results from the PRACTICAL consortium. Br J Cancer 118, el6 (2018).
[0210] 39. Haas NB, et al. Blood-based gene expression signature associated with metastatic castrate-resistant prostate cancer patient response to abiraterone plus prednisone or enzalutamide. Prostate Cancer Prostatic Dis 24, 448-456 (2021).
[0211] 40. Pasquale EB. Eph receptors and ephrins in cancer: bidirectional signalling and beyond. Nat Rev Cancer 10, 165-180 (2010).
[0212] 41. Phan NN, et al. Overexpressed gene signature of EPH receptor A / B family in cancer patients-comprehensive analyses from the public high-throughput database. Int J Clin Exp Pathol 13, 1220-1242 (2020).
[0213] 42. Chen X, et al. DEPTOR is an in vivo tumor suppressor that inhibits prostate tumorigenesis via the inactivation of mTORCl / 2 signals. Oncogene 39, 1557-1571 (2020).
[0214] 43. Peterson TR, et al. DEPTOR is an mTOR inhibitor frequently overexpressed in multiple myeloma cells and required fortheir survival. Cell 137, 873-886 (2009).Docket No. 2932719-000294-W01
[0215] Filed: February 2, 2026 44. Cui D, et al. DEPTOR is a direct p53 target that suppresses cell growth and chemosensitivity. Cell Death Dis 11, 976 (2020).
[0216] 45. Rosario SR, Jacobi JJ, Long MD, Affronti HC, Rowsam AM, Smiraglia DJ. JAZF1 : A Metabolic Regulator of Sensitivity to a Polyamine-Targeted Therapy. Mol Cancer Res 21, 24-35 (2023).
[0217] 46. Johansson A, et al. Common variants in the JAZF1 gene associated with height identified by linkage and genome-wide association analysis. Hum Mol Genet 18, 373-380 (2009).
[0218] 47. Liao ZZ, Wang YD, Qi XY, Xiao XH. JAZF1, a relevant metabolic regulator in type 2 diabetes. Diabetes Metab Res Rev 35, e3148 (2019).
[0219] 48. Hamada S, et al. Elevated fatty acid synthase expression in prostate needle biopsy cores predicts upgraded Gleason score in radical prostatectomy specimens. Prostate 74, 90-96 (2014).
[0220] 49. Johnson IR, et al. Endosomal gene expression: a new indicator for prostate cancer patient prognosis? Oncotarget 6, 37919-37929 (2015).
[0221] 50. Wang R, Xiao Y, Pan M, Chen Z, Yang P. Integrative Analysis of Bulk RNA-Seq and Single-Cell RNA-Seq Unveils the Characteristics of the Immune Microenvironment and Prognosis Signature in Prostate Cancer. J Oncol 2022, 6768139 (2022).
[0222] 51. Zeng K, et al. LACTB suppresses liver cancer progression through regulation of ferroptosis. Redox Biol 75, 103270 (2024).
[0223] 52. Huang G, et al. Prognostic and predictive value of super-enhancer-derived signatures for survival and lung metastasis in osteosarcoma. J TranslMed 22, 88 (2024).Docket No. 2932719-000294-W01
[0224] Filed: February 2, 2026 53. Cooperberg M, Meeks W, Fang R, Gaylis F, Catalona W, Makarov D. Mp43-03 Active Surveillance for Low-Risk Prostate Cancer: Time Trends and Variation in the Aua Quality (Aqua) Registry. Journal of Urology 207, (2022).
[0225] 54. Li X, Zhang Y, Wang Y. A 5-year follow-up assessment of anxiety and depression in postoperative prostate cancer patients: longitudinal progression and prognostic value. Psychol Health Med 28, 529-539 (2023).
[0226] 55. Hu S, Li L, Wu X, Liu Z, Fu A. Post-surgery anxiety and depression in prostate cancer patients: prevalence, longitudinal progression, and their correlations with survival profiles during a 3-year follow-up. Ir J Med Set 190, 1363-1372 (2021).
[0227] 56. Powell IJ, Vigneau FD, Bock CH, Ruterbusch J, Heilbrun LK. Reducing prostate cancer racial disparity: evidence for aggressive early prostate cancer PSA testing of African American men. Cancer Epidemiol Biomarkers Prev 23, 1505-1511 (2014).
[0228] 57. Fuletra JG, Kamenko A, Ramsey F, Eun DD, Reese AC. African-American men with prostate cancer have larger tumor volume than Caucasian men despite no difference in serum prostate specific antigen. The Canadian journal of urology 25, 9193-9198 (2018).
Claims
Docket No. 2932719-000294-W01Filed: February 2, 2026 CLAIMS1. A method comprising,one or more applications running on at least one processor for providing,receiving a plurality of single nucleotide polymorphism pairs (SNP-SNP pairs) identified as significantly associated with prostate cancer;applying a SNP interaction identifier (SIPI) approach to analyze the SNP-SNP pairs using a discovery dataset to identify SNP-SNP pairs that interact to influence prostate cancer in subjects, wherein the analyzing identifies a first plurality of SNP-SNP pairs, where the first plurality of SNP-SNP pairs comprises a subset of the SNP-SNP pairs;bootstrapping the first plurality of SNP-SNP pairs using the SIPI approach to identify a second plurality of SNP-SNP pairs, wherein the second plurality of SNP-SNP pairs comprises a subset of the first plurality of SNP-SNP pairs;removing SNP-SNP pairs from the second plurality of SNP-SNP pairs to create a third plurality of SNP-SNP pairs, wherein the removing comprises dropping a SNP-SNP pair when at least one of the SNPs comprise missing data;computing risk ratio scores for each of the third plurality of SNP-SNP pairs for each subject of the discovery dataset, wherein each SNP-SNP pair comprises pair genotypes, wherein a risk ratio score of a SNP-SNP pair for a subject comprises a sum of the subject specific pair genotype ratios across all pair genotypes present in the subject, wherein a pair genotype ratio comprises a numerator equal to prevalence of prostate cancer for the pair genotype in the discovery data set and the denominator equals the prevalence of prostate cancer for all pair genotypes in the discovery dataset;using the risk ratio scores and stepwise logistic regression to reduce size of the third plurality of SNP-SNP pairs and identify a final set of SNP-SNP pairs, wherein the stepwise logistic regression is fit using the risk ratio scores to predict significance of the plurality of SNP-SNP pairs to predict prostate cancer.
2. The method of claim 1, wherein the stepwise logistic regression comprises removing SNP-SNP pairs from the third plurality of SNP-SNP pairs based on a p-value greater than or equal to 0.05.Docket No. 2932719-000294-W01Filed: February 2, 2026 3. The method of claim 2, computing final risk ratio scores for each of the final set of SNP-SNP pairs for each subject in a validation dataset, wherein each SNP-SNP pair comprises pair genotypes, wherein a final risk ratio score of a SNP-SNP pair for a subject comprises a sum of the subject specific pair genotype ratios across all pair genotypes present in the subject, wherein a pair genotype ratio comprises a numerator equal to prevalence of prostate cancer for the pair genotype in the validation data set and the denominator equals the prevalence of prostate cancer for all pair genotypes in the validation dataset.
4. The method of claim 2, fitting a logistic regression model using a summation variable comprising information of the final set of SNP-SNP pairs and a polygenic risk score as predictors, wherein the logistic regression model comprises a predictive model for predicting a probability of a subject developing prostate cancer based on a corresponding summation variable value and polygenic risk score value.
5. The method of claim 4, wherein the summation variable comprises a sum of final risk ratio scores across all SNP-SNP pairs in the final set of SNP-SNP pairs for each subject in the validation dataset.
6. The method of claim 1, wherein the SIPI approach implements the 3pRule.
7. The method of claim 6, wherein the 3pRule comprises a p-pair significance of a SNP- SNP pair less than 1.6* 10'6.
8. The method of claim 7, wherein the 3pRule comprises a p-pair significance of the SNP-SNP pair less than the p-value significance of a first SNP of the SNP-SNP pair.
9. The method of claim 8, wherein the 3pRule comprises a p-pair significance of the SNP-SNP pair less than the p-value significance a second SNP of the SNP-SNP pair.
10. The method of claim 9, wherein the 3pRule comprises a p-pair significance the SNP-SNP pair is less than the p-value significance a first and second SNP of the SNP-SNP pair.Docket No. 2932719-000294-W01Filed: February 2, 202611. The method of claim 1, wherein the bootstrapping comprising identifying SNP-SNP pairs that are identified as significant by the SIPI approach in 75 percent or more of five hundred bootstrap runs.
12. The method of claim 1, wherein the polygenic risk score comprises a score computed using a subset of SNPs from known polygenic risk score PHS290.