Multiple-ancestry polygenic risk assessment for breast cancer
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- MYRIAD GENETICS INC
- Filing Date
- 2024-07-05
- Publication Date
- 2026-05-13
AI Technical Summary
Conventional methods for assessing breast cancer risk using polygenic risk scores are limited by their inability to accurately predict risk across different heritage populations due to biases and inaccuracies, particularly when relying on genomic data from a single heritage group, leading to overestimation of risk and poor discrimination between low and high risk individuals.
A multiple-ancestry polygenic risk score system that utilizes a unique set of single nucleotide polymorphisms (SNPs) and genomic loci, selected through a synthetic stepwise regression methodology accounting for linkage disequilibrium, to provide accurate and calibrated risk assessments across various heritage groups, regardless of self-reported heritage information.
The system achieves superior prediction and calibration of breast cancer risks, effectively distinguishing between low and high risk individuals, improving patient outcomes and medical care by providing accurate risk assessments for all heritage populations.
Smart Images

Figure US2024036902_16012025_PF_FP_ABST
Abstract
Description
MULTIPLE-ANCESTRY POLYGENIC RISK ASSESSMENT FOR BREASTCANCERCROSS REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit under 35 U.S.C. § 119(e) of U.S. Provisional Application No. 63 / 525,534, filed July 7, 2023, the contents of which are incorporated herein by reference in their entirety.TECHNICAL FIELD
[0002] This disclosure relates to the fields of genetics and medicine. More particularly, this disclosure relates to methods for assessing and predicting polygenic traits and breast cancer risks for medical use, as well as treating breast cancer.BACKGROUND
[0003] It is desirable to use polygenic risk scores to assess the expectation of a clinical trait or condition in a subject such as the risk of a particular disease. Risk scores from genomic data depend on identifying polymorphic loci to be used.
[0004] Conventional methods for trait expectation, such as for breast cancer risk, have identified various breast cancer associated genes. However, breast cancer genetics is highly complex, and these conventional methods have limitations that cannot overcome or address these complexities, which reduces the accuracy of risk predictions. Also, conventional methods may rely on genomic data from a single heritage.
[0005] An important drawback in conventional methods for characterizing risk of a trait from genomic data is that baseline data for a trait in one particular population may not accurately predict the same trait in a different population of different heritage. Conventional methods using genomic data from a population drawn from one heritage can overestimate the risk of a particular trait in a different population. Overestimation of risk is a significant drawback, especially for disease traits.
[0006] Another drawback of conventional methods for determining traits such as cancer risk include the problem that calculations using genomic data often depend on self-reported heritage information. Errors in self-reported heritage information in genomic data can prevent appropriate determination of cancer risk for a global population.
[0007] A significant drawback of conventional methods for determining risk of a trait is a lack of discrimination between low risk and high risk of the trait for different populations. For example, conventional methods for breast cancer risk based on genomic data from one heritage may not be able to distinguish between low risk and high risk for a population of a different heritage. This drawback of conventional methods can confuse prevention and treatment strategies for a disease trait and jeopardize patient outcomes.
[0008] Conventional methods for polygenic risk scores may rely on SNPs discovered through genome-wide association studies (GWAS). However, such SNPs are usually not causal, but may be in linkage disequilibrium (LD) with causal variants. Historically, genomewide association studies have included predominantly European populations, resulting in miscalibrated and inaccurate PRS for non-Europeans. What is need is a set of SNPs for polygenic risk estimation that will discriminate risk for all heritage groups and populations. It is also desirable to develop a method of polygenic risk estimation that does not bias calculations by population, and provides accurate results for all heritage groups and populations.
[0009] What is needed is a highly calibrated and accurate method for determining polygenic risk scores for traits such as breast cancer risk to avoid overestimation. There is a need for such methods to be useful for all heritage populations, and regardless of selfreported patient data. An advantageous clinical risk algorithm can improve medical care and patient treatment.
[0010] There is an urgent need for methods to assess traits such as breast cancer risk with good discrimination of risk level for all populations regardless of heritage. There is a need for methods that can be efficiently brought to the point of medical care.BRIEF SUMMARY
[0011] This disclosure provides improved methods for determining polygenic traits, such as risks for breast cancer. The methods of this disclosure can be used in medicine, as well as for treating diseases for which risk is identified and / or assessed.
[0012] In some aspects, methods of this disclosure may provide superior prediction of clinical risk in breast cancer patients. The methods of this disclosure can provide polygenic risk prediction for breast cancer which can be applied globally to all patients of all heritage groups.
[0013] A multiple-ancestry polygenic risk score of this disclosure can be used to assess the expectation of a clinical trait or condition such as cancer.
[0014] Aspects of this disclosure can characterize an individual’s risk of a trait from genomic data obtained for the trait in one or more particular heritage groups or populations, where the individual may be of a different or the same heritage group or population. Embodiments of this disclosure can provide a multiple-ancestry polygenic risk score for a trait in an individual using genomic data from a population drawn from a different heritage than for the individual, without overestimating the risk of the trait in the individual.
[0015] In further aspects, this disclosure contemplates accurately determining a trait such as cancer risk using genomic data of individuals who self-report heritage information. A multiple-ancestry polygenic risk score of this disclosure can be used to accurately determine cancer risk for a global population, regardless of any errors in self-reported heritage information.
[0016] In additional aspects, this disclosure provides methods for determining risk of a trait with sufficient discrimination between low risk and high risk of the trait for different populations.
[0017] In some embodiments, this disclosure includes methods for breast cancer risk based on a multiple-ancestry polygenic risk score that can distinguish between low risk and high risk for a population or individual of any heritage. The methods of this disclosure can provide prevention and treatment strategies for a disease trait and improve patient outcomes.
[0018] In further embodiments, this disclosure provides highly calibrated and accurate methods for determining multiple-ancestry polygenic risk scores for traits such as breast cancer risk which avoid overestimation. The methods of this disclosure can be useful for all heritage populations, regardless of self-reported patient data, and can improve medical care and patient treatment.
[0019] In additional embodiments, this disclosure provides methods for assessing traits such as breast cancer risk with enhanced discrimination of risk level for all populations, regardless of heritage. Methods of this disclosure can be efficiently brought to the point of medical care.
[0020] Methods of this disclosure further contemplate using various trait risk markers, which may be single nucleotide polymorphisms (SNP) or a genomic loci. The SNPs of this disclosure may be associated with breast cancer risk in one or more different heritage groups. Combinations of SNPs can be used to provide a multiple-ancestry polygenic risk score (MA- PRS), which can stratify unaffected patients for breast cancer risk, irrespective of thepresence or absence of a family history of the disease. The genomic loci of this disclosure may be a specific portion of the genome that may include a single nucleotide polymorphism (SNP) associated with breast cancer risk in one or more different heritage groups and additional bases in high linkage disequilibrium with the breast cancer risk-associated SNP. These additional bases may be physically close to the SNP. These genomic loci may be within around 200 bases, around 100 bases, around 50 bases, around 40 bases, around 30 bases, around 20 bases, or around 10 bases of the breast cancer risk-associated SNP. Additionally or alternatively, these genomic loci may be within around 0.2 centimorgans (cM), around 0.1 cM, or around 0.05 cM of the breast cancer risk-associated SNP.
[0021] Aspects of this disclosure provide methods for polygenic risk scoring that rely on a unique set of SNPs or on a unique set of genomic loci discovered through designated criteria. In some embodiments, the unique set of SNP markers or unique set of genomic loci may be selected using a novel synthetic stepwise regression methodology that accounts for linkage disequilibrium. This unique set of SNPs or unique set of genomic loci for polygenic risk estimation can discriminate risk for all heritage groups and populations. The set of SNP markers for polygenic risk estimation disclosed herein provide accurate results for all heritage groups and populations, substantially without bias toward any population or heritage group.
[0022] Additional classes of markers or elements can include age, family history, breast density, and hormone exposure.
[0023] In certain aspects, the clinical utility of this disclosure may include superior prediction of clinical risk for breast cancer patients of all ancestries.
[0024] A multiple-ancestry polygenic score obtained by the methods of this disclosure can provide surprisingly increased accuracy in determining breast cancer risks.
[0025] Methods of this disclosure can provide surprisingly accurate determination of polygenic traits and risks by assessing and including contributions of a wide range of markers for different ancestries.
[0026] Embodiments of this disclosure contemplate determining the levels of polygenic traits and risks in the form of a score based on various genomic risk loci. The genomic risk loci can be discretely identified and defined, so that accurate determination can be done by genotyping subjects.
[0027] In certain aspects, the genomic risk loci can include genomic risk markers for breast cancer, which are combined with additional risk markers that can be specifically breast cancer-informative.
[0028] Embodiments of this disclosure include:
[0029] A method for assessing a risk of a trait in a subject, the method comprising: selecting a plurality trait-associated SNP markers and a plurality of ancestry- informative SNP markers; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject based on the trait-associated SNP markers and the plurality of ancestry-informative SNP markers.
[0030] The trait-associated SNPs can be selected using a synthetic stepwise regression methodology that accounts for linkage disequilibrium between variants. The trait-associated SNPs can comprise one or more of European breast cancer-associated SNPs, African breast cancer-associated SNPs, East Asian breast cancer-associated SNPs, and Amerindian breast cancer-associated SNPs.
[0031] The method above, wherein calculating the multiple-ancestry polygenic risk score for the risk of the trait in the subject with additional clinical variables of the subject. The additional clinical variables can be age, personal medical history, and family medical history of the subject.
[0032] The method above, wherein the trait is a risk of a disease in the subject. The disease may be cancer.
[0033] The plurality of ancestry-informative SNP markers can be from 10 to 50,000 SNP markers. The plurality of ancestry-informative SNP markers can be from 10 to 56 SNP markers.
[0034] The trait-associated SNP markers can be a plurality of cancer-associated SNP markers. The trait-associated SNP markers can be a plurality of from 10 to 50,000 breast cancer-associated SNP markers. The trait-associated SNP markers can be a plurality of from 10 to 329 breast cancer-associated SNP markers.
[0035] The method above, wherein the calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject may be done with training clinical data of a reference group. The calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject may be done with validating clinical data of a reference group.
[0036] The method above, wherein the genotype of the subject may be measured by NGS. The genotype of the subject may be determined with a sequencing chip.
[0037] The method above, wherein the plurality of ancestry-informative SNP markers may determine a fractional heritage in the genotype of the subject for each of four or more different heritage populations.
[0038] The method above, wherein the plurality of ancestry-informative SNP markers may determine a fractional heritage in the genotype of the subject for each of African, European, East Asian, and Amerindian heritage populations.
[0039] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can be accurate for subjects in three or more different heritage populations, even when the heritage populations are self-reported.
[0040] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can be accurate for subjects in African, European, East Asian, and Amerindian heritage populations, even when the heritage populations are self-reported.
[0041] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can be calibrated for subjects in four or more different heritage populations so that the risk of the trait is not overestimated in any heritage population.
[0042] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can be calibrated for subjects in African, European, East Asian, and Amerindian heritage populations so that the risk of the trait is not overestimated in any heritage population.
[0043] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can discriminate between low risk and high risk for subjects in four or more different heritage populations.
[0044] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can discriminate between low risk and high risk for subjects in African, European, East Asian, and Amerindian heritage populations.
[0045] The methods above, wherein the trait may be a risk of a disease in the subject, such as cancer.
[0046] The method above, wherein the calculating a multiple-ancestry polygenic risk score comprises using clinical cohorts of women of African self-reported ancestry, East Asian self-reported ancestry, and European self-reported ancestry.
[0047] The method above, wherein the calculating a multiple-ancestry polygenic risk score may comprise using the sum of ancestry specific polygenic risk scores weighted according to fractional ancestral composition.
[0048] The method above, wherein the multiple-ancestry polygenic risk score can be strongly associated with breast cancer in a reference cohort and in sub-cohorts defined by self-reported ancestry.
[0049] The method above, wherein the multiple-ancestry polygenic risk score can be combined with clinical and / or biological risk factors for accurate risk stratification for all women of all ancestries.
[0050] The method above, wherein the calculating a multiple-ancestry polygenic risk score may comprise calculating and combining: an African-specific PRS (PRSAT), East Asian-specific PRS (PRSEA), European- specific PRS (PRSEU), and Amerindian-specific PRS (PRSAI); an estimated weight for each ancestry including African (BAT), East Asian (BEA), European (BEU), and Amerindian (BAI); and, an Amerindian SNP genotype (xAm), according to the following equation: BATXPRSAT+ BEAXPRSEA + BEUXPRSEU + BA;XPRSA; + PAmxXAm.
[0051] A method for assessing a risk of a trait in a subject, the method comprising: selecting a plurality trait-associated genomic loci and a plurality of ancestry- informative genomic loci; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject based on the trait-associated genomic loci and the plurality of ancestry -informative genomic loci.
[0052] The trait-associated genomic loci can be selected using a synthetic stepwise regression methodology that accounts for linkage disequilibrium between variants. The trait- associated genomic loci can comprise one or more of European breast cancer-associated genomic loci, African breast cancer-associated genomic loci, East Asian breast cancer- associated genomic loci, and Amerindian breast cancer-associated genomic loci.
[0053] The method above, wherein the genomic loci are regions of the genome within around 200 bases, around 100 bases, around 50 bases, around 40 bases, around 30 bases, around 20 bases, or around 10 bases of the trait-associated SNPs. The method above, wherein the genomic loci are regions of the genome within around 0.2 centimorgans (cM), around 0.1 cM, or around 0.05 cM of the trait-associated SNPs. The method above, wherein the genomic loci are regions of the genome within around 200 bases, around 100 bases, around 50 bases, around 40 bases, around 30 bases, around 20 bases, or around 10 bases of the ancestry- informative genomic loci. The method above, wherein the genomic loci are regions of the genome within around 0.2 cM, around 0.1 cM, or around 0.05 cM of the ancestry-informative genomic loci.
[0054] The method above, wherein calculating the multiple-ancestry polygenic risk score for the risk of the trait in the subject with additional clinical variables of the subject. The additional clinical variables can be age, personal medical history, and family medical history of the subject.
[0055] The method above, wherein the trait is a risk of a disease in the subject. The disease may be cancer.
[0056] The plurality of ancestry-informative genomic loci can be from 10 to 50,000 genomic loci. The plurality of ancestry-informative genomic loci can be from 10 to 56 genomic loci.
[0057] The trait-associated genomic loci can be a plurality of cancer-associated genomic loci. The trait-associated genomic loci can be a plurality of from 10 to 50,000 breast cancer- associated genomic loci. The trait-associated genomic loci can be a plurality of from 10 to 329 breast cancer-associated genomic loci.
[0058] The method above, wherein the calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject may be done with training clinical data of a reference group. The calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject may be done with validating clinical data of a reference group.
[0059] The method above, wherein the genotype of the subject may be measured by NGS.
[0060] The method above, wherein the plurality of ancestry-informative genomic loci may determine a fractional heritage in the genotype of the subject for each of four or more different heritage populations.
[0061] The method above, wherein the plurality of ancestry-informative genomic loci may determine a fractional heritage in the genotype of the subject for each of African, European, East Asian, and Amerindian heritage populations.
[0062] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can be accurate for subjects in three or more different heritage populations, even when the heritage populations are self-reported.
[0063] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can be accurate for subjects in African, European, East Asian, and Amerindian heritage populations, even when the heritage populations are self-reported.
[0064] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can be calibrated for subjects in four or more different heritage populations so that the risk of the trait is not overestimated in any heritage population.
[0065] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can be calibrated for subjects in African, European, East Asian, and Amerindian heritage populations so that the risk of the trait is not overestimated in any heritage population.
[0066] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can discriminate between low risk and high risk for subjects in four or more different heritage populations.
[0067] The method above, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject can discriminate between low risk and high risk for subjects in African, European, East Asian, and Amerindian heritage populations.
[0068] The methods above, wherein the trait may be a risk of a disease in the subject, such as cancer.
[0069] The method above, wherein the calculating a multiple-ancestry polygenic risk score comprises using clinical cohorts of women of African self-reported ancestry, East Asian self-reported ancestry, and European self-reported ancestry.
[0070] The method above, wherein the calculating a multiple-ancestry polygenic risk score may comprise using the sum of ancestry specific polygenic risk scores weighted according to fractional ancestral composition.
[0071] The method above, wherein the multiple-ancestry polygenic risk score can be strongly associated with breast cancer in a reference cohort and in sub-cohorts defined by self-reported ancestry.
[0072] The method above, wherein the multiple-ancestry polygenic risk score can be combined with clinical and / or biological risk factors for accurate risk stratification for all women of all ancestries.
[0073] The method above, wherein the calculating a multiple-ancestry polygenic risk score may comprise calculating and combining: an African-specific PRS (PRSAT), East Asian-specific PRS (PRSEA), European- specific PRS (PRSEU), and Amerindian-specific PRS (PRSAI); an estimated weight for each ancestry including African (BAT), East Asian (BEA), European (BEU), and Amerindian (BAI); and, an Amerindian SNP genotype (xAm), according to the following equation:BATXPRSAT+ BEAXPRSEA + BEUXPRSEU + BA;XPRSA; + PAmxXAm.
[0074] Embodiments of this disclosure further contemplate methods for treating a disease in a subject in need thereof, the method comprising:selecting a plurality of disease-associated SNP markers and a plurality of ancestry-informative SNP markers; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the disease in the subject based on the plurality of the disease-associated SNP markers and ancestry- informative SNP markers and, wherein the score indicates a need for treating the subject; and administering to the subject a therapy for the disease.
[0075] The method above, may further comprise calculating the multiple-ancestry polygenic risk score with additional variables for age, personal medical history, and family medical history.
[0076] The method above, wherein the disease is cancer. In some embodiments, the therapy may be a cancer therapy selected from one or more of surgery, cryoablation, radiation therapy, bone marrow transplant, chemotherapy, immunotherapy, hormone therapy, stem cell therapy, drug therapy, biological therapy, and administration of a pharmaceutical, prophylactic or therapeutic compound. The disease may be breast cancer and the therapy may be a breast cancer therapy.
[0077] Embodiments of this disclosure further contemplate methods for treating a disease in a subject in need thereof, the method comprising: selecting a plurality of disease-associated genomic loci and a plurality of ancestry- informative genomic loci; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the disease in the subject based on the plurality of the disease-associated genomic loci and ancestry- informative genomic loci and, wherein the score indicates a need for treating the subject; and administering to the subject a therapy for the disease.
[0078] The method above, may further comprise calculating the multiple-ancestry polygenic risk score with additional variables for age, personal medical history, and family medical history.
[0079] The method above, wherein the disease is cancer. In some embodiments, the therapy may be a cancer therapy selected from one or more of surgery, cryoablation, radiation therapy, bone marrow transplant, chemotherapy, immunotherapy, hormone therapy, stem cell therapy, drug therapy, biological therapy, and administration of a pharmaceutical, prophylactic or therapeutic compound. The disease may be breast cancer and the therapy may be a breast cancer therapy.
[0080] This disclosure includes methods for diagnosing or prognosing a subject having a disease, the method comprising: selecting a plurality of disease-associated SNP markers and a plurality of ancestry-informative SNP markers; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the disease in the subject based on the plurality of the disease-associated SNP markers and ancestry- informative SNP markers, wherein the score indicates a diagnosis or prognosis for the subject. The disease may be cancer.
[0081] This disclosure includes methods for generating data for assessing a trait in a subject, the method comprising: selecting a plurality of disease-associated SNP markers and a plurality of ancestry-informative SNP markers; measuring a genotype of the subject; and measuring trait-associated SNP markers in the genotype of the subject.
[0082] The method above, may further comprise determining additional clinical variables of the subject such as age, personal medical history, and family medical history of the subject.
[0083] The method above, wherein the trait may be a risk of a disease in the subject, such as cancer.
[0084] The method above, wherein the plurality of ancestry-informative SNP markers are from 10 to 50,000 SNP markers or from 10 to 56 SNP markers.
[0085] The method above, wherein the trait-associated SNP markers are a plurality of cancer-associated SNP markers. The trait-associated SNP markers may be a plurality of from 10 to 50,000 breast cancer associated SNP markers or from 10 to 329 breast cancer associated SNP markers.
[0086] This disclosure includes methods for diagnosing or prognosing a subject having a disease, the method comprising: selecting a plurality of disease-associated genomic loci and a plurality of ancestry- informative genomic loci; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the disease in the subject based on the plurality of the disease-associated genomic loci and ancestry-informative genomic loci, wherein the score indicates a diagnosis or prognosis for the subject. The disease may be cancer.
[0087] This disclosure includes methods for generating data for assessing a trait in a subject, the method comprising: selecting a plurality of disease-associated genomic loci and a plurality of ancestry- informative genomic loci; measuring a genotype of the subject; and measuring trait-associated genomic loci in the genotype of the subject.
[0088] The method above, may further comprise determining additional clinical variables of the subject such as age, personal medical history, and family medical history of the subject.
[0089] The method above, wherein the trait may be a risk of a disease in the subject, such as cancer.
[0090] The method above, wherein the plurality of ancestry-informative genomic loci are from 10 to 50,000 genomic loci or from 10 to 56 genomic loci.
[0091] The method above, wherein the trait-associated genomic loci are a plurality of cancer-associated genomic loci. The trait-associated genomic loci may be a plurality of from 10 to 50,000 breast cancer genomic loci or from 10 to 329 breast cancer associated genomic loci.
[0092] The method above, wherein the genomic loci are regions of the genome within around 200 bases, around 100 bases, around 50 bases, around 40 bases, around 30 bases, around 20 bases, or around 10 bases of the trait-associated SNPs. The method above, wherein the genomic loci are regions of the genome within around 0.2 centimorgans (cM), around 0.1 cM, or around 0.05 cM of the trait-associated SNPs. The method above, wherein the genomic loci are regions of the genome within around 200 bases, around 100 bases, around 50 bases, around 40 bases, around 30 bases, around 20 bases, or around 10 bases of the ancestry- informative genomic loci. The method above, wherein the genomic loci are regions of the genome within around 0.2 cM, around 0.1 cM, or around 0.05 cM of the ancestry-informative genomic loci.
[0093] This disclosure further includes systems for assessing risk of a disease, such as cancer, in a subject, the system comprising: a processor for receiving a genotype of the subject; one or more processors for carrying out the steps: calculating a multiple-ancestry polygenic risk score for risk of the disease in thesubject based on a plurality of ancestry-informative SNP markers, a plurality of disease- associated SNP markers of the genotype, and additional variables for age, personal medical history, and family medical history; and a display for displaying and / or reporting the risk score.
[0094] This disclosure further includes systems for assessing risk of a disease, such as cancer, in a subject, the system comprising: a processor for receiving a genotype of the subject; one or more processors for carrying out the steps: calculating a multiple-ancestry polygenic risk score for risk of the disease in the subject based on a plurality of ancestry-informative genomic loci, a plurality of disease- associated genomic loci of the genotype, and additional variables for age, personal medical history, and family medical history; and a display for displaying and / or reporting the risk score.
[0095] Additional embodiments include non-transitory machine-readable storage mediums having stored therein instructions for execution by a processor which cause the processor to perform the steps of a method for assessing risk of a disease, such as cancer, in a subject, the method comprising: receiving a genotype of the subject; calculating a multiple-ancestry polygenic risk score for risk of the disease in the subject based on a plurality of ancestry-informative SNP markers, a plurality of disease- associated SNP markers of the genotype, and additional variables for age, personal medical history, and family medical history; and sending to a processor output for displaying and / or reporting the risk score.
[0096] Additional embodiments include non-transitory machine-readable storage mediums having stored therein instructions for execution by a processor which cause the processor to perform the steps of a method for assessing risk of a disease, such as cancer, in a subject, the method comprising: receiving a genotype of the subject; calculating a multiple-ancestry polygenic risk score for risk of the disease in the subject based on a plurality of ancestry-informative genomic loci, a plurality of disease- associated genomic loci of the genotype, and additional variables for age, personal medical history, and family medical history; and sending to a processor output for displaying and / or reporting the risk score.BRIEF DESCRIPTION OF THE DRAWINGS
[0097] FIG. 1 shows an illustration of ancestry in terms of contributions from different continents.
[0098] FIG. 2 shows an illustration of a distribution of genotypes based on ancestry for Hispanic, White / non-Hispanic, Black / African, and Asian genotypes.DETAILED DESCRIPTION OF THE DISCLOSURE
[0099] This disclosure includes methods for determining a multiple-ancestry polygenic risk score which can be predictive for a trait in a subject.
[0100] A multiple-ancestry polygenic risk score can be predictive for risk assessment for breast cancer. This multiple-ancestry polygenic risk score can utilize SNPs that are associated with disease presence and ancestry and / or genomic loci in linkage disequilibrium with these disease-associated or ancestry -informative SNPs to predict risk of disease, including of breast cancer.
[0101] In some aspects, this disclosure provides methods for multiple-ancestry polygenic risk prediction with surprisingly increased accuracy of risk assessment for a trait in a subject.
[0102] Embodiments of this disclosure further provide reliable breast cancer risk associations applicable to all populations of all ancestries.
[0103] This disclosure provides various methods for clinical risk management, risk magnitude assessment, as well as multiple-ancestry polygenic risk scores, and non-clinical trait prediction. Methods of this disclosure can provide predictive ability that is surprisingly accurate for populations of all ancestries.
[0104] Aspects of this disclosure include genotyping a subject using various markers associated with a disease and combining the genotypes in the form of a multiple-ancestry polygenic risk score to predict risk of a trait, such as a clinical condition or an extent of manifestation of a biological trait.
[0105] In further embodiments, a plurality of trait risk markers can be used to provide a multiple-ancestry polygenic risk prediction for the trait.
[0106] The plurality of trait risk markers may include various disease-associated gene markers.
[0107] In some embodiments, the plurality of trait risk markers may include from 1- 1,000,000 SNP markers.
[0108] In certain embodiments, the plurality of trait risk markers may include from 1- 10,000 SNP markers, or from 1-1000 SNP markers, or from 1-100 SNP markers. A plurality of trait risk markers may be from 1-1000 breast cancer SNP markers, or from 1-500 breast cancer SNP markers, or from 1-100 breast cancer SNP markers.
[0109] In certain embodiments, the plurality of trait risk markers may include 56 SNP markers to 385 SNP markers.
[0110] The method may also utilize detection of bases within genomic loci or genomic loci that are in linkage disequilibrium with the disease-associated SNPs in order to provide a multiple-ancestry polygenic risk prediction for the trait, as the bases in these genomic loci are highly associated with one another based on the reduced frequencies of crossing over between the bases in the genomic loci. In some embodiment, the plurality of trait risk markers may include genomic loci associated with from 1-1,000,000 SNP markers. In certain embodiments, the plurality of trait risk markers may include genomic loci associated with from 1-10,000 SNP markers, or genomic loci associated with from 1-1,000 SNP markers, or genomic loci associated with from 1-100 SNP markers. In some embodiments ,the plurality of trait risk markers may be genomic loci associated with 1-1,000 breast cancer SNP markers, or genomic loci associated with from 1-500 breast cancer SNP markers, or genomic loci associated with from 1-1000 breast cancer SNP markers. In certain embodiments, the plurality of trait risk markers may include genomic loci associated with 56 SNP markers to 385 SNP markers. In some embodiment, the plurality of trait risk markers may include bases within genomic loci associated with from 1-1,000,000 SNP markers. In certain embodiments, the plurality of trait risk markers may include bases within genomic loci associated with from 1-10,000 SNP markers, or bases within genomic loci associated with from 1-1,000 SNP markers, or bases within genomic loci associated with from 1-100 SNP markers. In some embodiments ,the plurality of trait risk markers may be bases within genomic loci associated with 1-1,000 breast cancer SNP markers, or bases within genomic loci associated with from 1-500 breast cancer SNP markers, or bases within genomic loci associated with from 1-1000 breast cancer SNP markers. In certain embodiments, the plurality of trait risk markers may include bases within genomic loci associated with 56 SNP markers to 385 SNP markers.
[0111] In some embodiments, the bases within the genomic loci or the genomic loci associated with the SNP markers may be within about 500 bases, within about 400 bases, within about 300 bases, within about 250 bases, within about 200 bases, within about 150 bases, within 100 bases, within about 90 bases, within about 80 bases, within about 70 bases, within about 60 bases, within about 50 bases, within about 40 bases, within about 30 bases,within about 20 bases, and within about 10 bases of the SNP marker. In some embodiments, the bases in the genomic loci or the genomic loci may be within 15 centimorgans (cM) or less, 10 cM or less, 9 cM or less, 8 cM or less, 7 cM or less, 6 cM or less, 5 cM or less, 4 cM or less, 3 cM or less, 2 cM or less, 1 cM or less, 0.75 cM or less, 0.5 cM or less, 0.25 cM or less, 0.2 cM or less, 0.1 cM or less, or 0.05 cM or less of the SNP marker. In some embodiments, the bases in the genomic loci or the genomic loci associated with the SNP markers may have a linkage disequilibrium value of above 0.1, above 0.2, above 0.5, above 0.6, above 0.7, above 0.8, above 0.9, or about 1.0 with the SNP marker. In some embodiments, the bases in the genomic loci or the genomic loci associated with the SNP markers may a logarithm of the odds score of at least above 2, at least above 3, at least above 4, at least above 5, at least above 6, at least above 7, at least above 8, at least above 9, at least above 10, at least above 20 at least above 30, at least above 40, at least above 50 with the SNP markers. In some embodiments, the bases in the genomic loci or the genomic loci associated with the SNP markers and the SNP markers may exhibit recombination during meiosis with each other at a frequency of less than or equal to about 20%, about 19%, about 18%, about 17%, about 16%, about 15%, about 14%, about 13%, about 12%, about 11%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.75%, about 0.5%, about 0.25%, or about 0.1% or less .
[0112] This disclosure provides methods for determining polygenic traits, such as risks for disease including breast cancer. The methods of this disclosure can be used for treating diseases for which risk is determined through polygenic scoring.
[0113] In some embodiments, methods of this disclosure may provide superior prediction of clinical risk in breast cancer patients. The methods of this disclosure can provide multipleancestry polygenic risk prediction for disease such as breast cancer which can be applied globally to all patients of all heritage groups.
[0114] A multiple-ancestry polygenic risk score of this disclosure can be used to assess the expectation of a clinical trait or condition such as cancer in a subject.
[0115] In certain embodiments, this disclosure can calculate an individual’s risk of a trait from genomic data obtained for the trait in one or more particular heritage groups or population, where the individual may be of the same or a different heritage group or population. Embodiments of this disclosure can therefore provide a multiple-ancestry polygenic risk score for a trait in an individual using genomic data from a population drawn from a plurality of different heritages, including heritages different than the heritage to whichthe individual belongs or has self-identified, without overestimating the risk of the trait in the individual.
[0116] In further embodiments, this disclosure contemplates accurately determining a trait such as cancer risk using genomic data of individuals who self-report heritage information. A multiple-ancestry polygenic risk score of this disclosure can be used to accurately determine cancer risk for a subject of any heritage, regardless of any errors in selfreported heritage information.
[0117] In additional embodiments, this disclosure provides methods for determining risk of a trait in an individual with sufficient discrimination between low risk and high risk for the trait, regardless of the heritage group or population to which the subject belongs or has selfidentified.
[0118] In some embodiments, this disclosure includes methods for breast cancer risk based on a multiple-ancestry polygenic risk score that can distinguish between low risk and high risk, surprisingly for an individual of any heritage.
[0119] The methods of this disclosure can provide prevention and treatment strategies for a disease trait to improve patient outcomes.
[0120] In further embodiments, this disclosure can provide multiple-ancestry polygenic risk scores that are highly calibrated and accurate. The multiple-ancestry polygenic risk scores can be used in methods for determining traits such as breast cancer risk in subjects which avoid overestimation. The methods of this disclosure can be useful for all heritage groups and / or populations, regardless of the use of self-reported patient data, and can improve medical care and patient treatment.
[0121] In additional embodiments, this disclosure provides methods for assessing traits such as breast cancer risk with enhanced discrimination of risk level for all populations, regardless of heritage. Methods of this disclosure can be efficiently brought to the point of medical care.
[0122] Methods of this disclosure further contemplate using various trait risk markers, which may be single nucleotide polymorphisms (SNP) or bases in genomic loci associated with the SNP or genomic loci associated with the SNP. The SNPs of this disclosure may be associated with breast cancer risk in one or more different heritage groups. Combinations of SNPs or bases within genomic loci associated with the SNP or genomic loci associated with the SNP can be used to provide a multiple-ancestry polygenic risk score (MA-PRS), which can stratify unaffected patients for breast cancer risk, irrespective of the presence or absence of a family history of the disease.
[0123] Additional classes of markers or elements can include age, family history, breast density, and hormone exposure.
[0124] In certain aspects, the clinical utility of this disclosure may include superior prediction of clinical risk for breast cancer patients of all ancestries.
[0125] A multiple-ancestry polygenic score obtained by the methods of this disclosure can provide surprisingly increased accuracy in determining breast cancer risks.
[0126] Methods of this disclosure can provide surprisingly accurate determination of polygenic traits and risks by assessing and including contributions of a wide range of markers for different ancestries.
[0127] Embodiments of this disclosure contemplate determining the levels of polygenic traits and risks in the form of a score based on various genomic risk loci. The genomic risk loci can be discretely identified and defined, so that accurate determination can be done by genotyping subjects.
[0128] In certain aspects, the genomic risk loci can include genomic risk markers for breast cancer, which are combined with additional risk markers that can be specifically breast cancer-informative.
[0129] In additional embodiments, the plurality of trait risk markers may include from 1- 100 family history elements, or from 1-20 family history elements, or from 1-10 family history elements.
[0130] Embodiments of this disclosure may include a plurality of trait risk markers such as from 1-100 clinical elements, or from 1-20 clinical elements, or from 1-10 clinical elements.
[0131] Embodiments herein can provide improved multiple-ancestry polygenic risk prediction for breast cancer.
[0132] Comprehensive risk assessment combining a polygenic SNP or SNP-associated genomic loci or bases within SNP-associated genomic loci scoring method with other risk factors and elements can improve the accuracy of risk estimates and facilitate decisionmaking for women with pathogenic variants in moderately penetrant genes.
[0133] In further aspects, a polygenic risk score of this disclosure may be surprisingly more accurate for breast cancer than using conventional methods.
[0134] In certain aspects, an association between the multiple-ancestry polygenic risk scores and breast cancer may be evaluated by fixed stratification methods. The fixed stratification may be adjusted for age and family history, among other variables and elements.
[0135] Embodiments of this disclosure can provide women an estimated lifetime risk for breast cancer with increased accuracy. Such risk estimation is useful to inform decisions based on a threshold for more aggressive screening, including consideration of breast magnetic resonance imaging (MRI).
[0136] In some aspects, disclosed herein are methods that can utilize breast cancer SNP markers or bases within genomic loci associated with breast cancer SNP markers or genomic loci associated with breast cancer SNP markers to provide a multiple-ancestry polygenic risk score for breast cancer.
[0137] Some examples of breast cancer risk markers are given in: Prediction of breast cancer risk based on profiling with common genetic variants, Mavaddat et al., J Natl Cancer Inst., 2015, April 8, Vol. 107(5), djv036.
[0138] Some examples of breast cancer risk markers are given in: Michailidou et al., Genome-wide association analysis of more than 120,000 individuals identifies 15 new susceptibility loci for breast cancer, Nat Genet., 2015, Vol. 47, pp. 373.
[0139] Some examples of breast cancer risk markers are given in Characterizing Genetic Susceptibility to Breast Cancer in Women of African Ancestry, Feng et al., Cancer Epidemiol Biomarkers Prev., 2017, July, Vol. 26(7), pp. 1016-1026.
[0140] Some examples of breast cancer risk markers are given in Rainville, I. et al., Breast Cancer Research and Treatment, 2020, Vol. 180, pp. 503-509.
[0141] Some examples of breast cancer risk markers are given in Early Diagnosis of Breast Cancer, Wang et al., Sensors (Basel), 2017, July, Vol. 17(7), p. 1572.
[0142] Some examples of genetic modifiers for breast cancer risk are given in Muranen TA, et al., Genetics in Medicine, 2017, Vol. 19(5), pp. 599-603.
[0143] Some examples of risk scores for breast cancer are given in Kuchenbaecker K, et al., J Natl Cancer Inst., 2017, Vol. 109(7), djw302.
[0144] Some examples for cancer risk are given in: Perencevich M, et al., Gastroenterology & Hepatology, 2011, Vol. 7(6), pp. 420-423.
[0145] Some examples for gene analysis are given in: Lek et al., Nature, 2016, Vol. 536.7616, pp. 285.Definitions
[0146] The following terms or definitions are provided solely to aid in the understanding of the disclosure. Additional definitions for other terms may be provided throughout this document. Further, terms given a general definition here in this section may be ascribed a more specific or different definition in another place of the disclosure that is applied to theindicated specific context. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the present disclosure. Practitioners are particularly directed to Sambrook el al. , Molecular Cloning: A Laboratory Manual, 2nded., Cold Spring Harbor Press, Plainsview, N.Y. (1989); and Ausubel etal., Current Protocols in Molecular Biology (Supplement 47), John Wiley & Sons, New York (1999), for definitions and terms of the art. Unless expressly defined otherwise herein, the terms used herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.
[0147] As used herein, unless stated to the contrary, “about” means + / - 10%, more preferably + / - 5%, more preferably + / - 1%, of the designated value.
[0148] As used herein, “algorithm” encompasses any formula, model, mathematical equation, algorithmic, analytical or programmed process, or statistical technique or classification analysis that takes one or more inputs or parameters, whether continuous or categorical, and calculates an output value, index, index value or score. Examples of algorithms include but are not limited to ratios, sums, regression operators such as exponents or coefficients, biomarker value transformations and normalizations (including, without limitation, normalization schemes that are based on clinical parameters such as age, gender, ethnicity, etc.), rules and guidelines, statistical classification models, and neural networks trained on populations. Also of use in the context of mutation load as described herein are linear and non-linear equations and statistical classification analyses to determine the relationship between (a) the number of mutations detected in a subject sample and (b) the level of the respective subject’s mutation load.
[0149] As used herein, “allele” means one of two or more different nucleotide sequences (DNA or RNA) that occur or are encoded at a specific locus, or two or more different polypeptide sequences encoded by such a locus. For example, a first allele can occur on one chromosome, while a second allele occurs on a second homologous chromosome, e.g., as occurs for different chromosomes of a heterozygous individual, or between different homozygous or heterozygous individuals in a population. In the context of the genotype at a particular locus (e.g., a SNP locus), an allele generally refers to the nucleotide base present on chromosome (out of the expected two) at that specific locus. For example, at one particular SNP locus a patient may have an adenine (A) one chromosome and a guanine (G) one the other, in which case it can be said that the patient has one A allele and one G allele.As used herein, “homozygous” means an individual or subject has only one type of allele at a given locus (e.g., a diploid individual has a copy of the same allele at a locus for each of two homologous chromosomes, such as A / A in the preceding example). An individual is “heterozygous” if more than one allele type is present at a given locus (e.g., a diploid individual with one copy each of two different alleles, such as A / G in the preceding example). The term “homogeneity” indicates the degree to which members of a group have the same genotype at one or more specific loci. In contrast, the term “heterogeneity” is used to indicate the degree to which individuals within the group differ in genotype at one or more specific loci (e.g., all homozygous, all the same type of heterozygosity, etc.). An allele “positively” correlates with a trait when it is linked to it and when presence of the allele is an indicator that the trait or trait form will occur in an individual comprising the allele. An allele “negatively” correlates with a trait when it is linked to it and when presence of the allele is an indicator that a trait or trait form will not occur in an individual comprising the allele.
[0150] “Allele frequency” refers to the frequency (e.g., proportion or percentage) at which an allele (e.g., adenine versus guanine in the example above) is present at a locus within an individual, within a line or within a population (or subpopulation). In the above example, for an allele “A”, diploid individuals of genotype “A / A”, “A / G,” or “G / G” have allele frequencies of 1.0, 0.5, or 0.0, respectively. One can estimate the allele frequency within a line or population (e.g., cases or controls) by averaging the allele frequencies of a sampling of individuals from that line or population. Similarly, one can calculate the allele frequency within a population of lines by averaging the allele frequencies of lines that make up the population. In some embodiments, the term “allele frequency” is used to define the minor allele frequency (MAF). MAF refers to the frequency at which the least common allele (where two alleles are observed) occurs in a given population, or the frequency at which the second most common allele (where more than two alleles are observed) occurs in a given population.
[0151] As used herein, “amplifying” in the context of nucleic acid amplification means any process or reaction whereby additional copies of a nucleic acid (or a transcribed form thereof) comprising a particular nucleotide sequence are produced. Amplification techniques include, but are not limited to, various polymerase based replication methods, including the polymerase chain reaction (PCR), ligase mediated methods such as the ligase chain reaction (LCR) and RNA polymerase based amplification (e.g., by transcription) methods. An“amplicon” is an amplified nucleic acid, e.g., a nucleic acid (or population of such nucleic acids, e.g., in solution) that is (or are) produced by amplifying a template nucleic acid by an amplification technique (e.g., PCR, LCR, transcription, or the like).
[0152] As used herein, the term “analyze” or “analyzing” generally includes “measure,” “measuring,” “detect,” “detecting,” “identify,” “identifying,” “assay,” “assaying,” “quantify,” or “quantifying,” and refers to the process of evaluating a biological sample (or a sample derived therefrom) for the presence, absence amount, level, or quality of some physical, chemical, or electromagnetic property(ies). This is often done by determining a value or set of values associated with such properties (e.g., number of sequencing reads in which a fluorescence signal indicating the presence of an adenine was observed at a particular position within the read corresponding to a particular position in a gene, chromosome or genome). Specific examples particularly relevant to the present disclosure include analyzing a sample to determining the sequence at one or more particular genomic loci in the sample, and may further comprise comparing test nucleotide sequence(s) detected in a patient’s sample against reference nucleotide sequence(s) and / or comparing the test number of any such test sequences to one or more reference numbers of such reference sequences.
[0153] As used herein, “breast cancer” encompasses any type of breast cancer that can develop in a subject. For example, the breast cancer may be characterized as Luminal A (ER+ and / or PR+, HER2-, low Ki67), Luminal B (ER+ and / or PR+, HER2+ (or HER2- with high Ki67), Triple negative / basal-like (ER-, PR-, HER2-) or HER2 type (ER-, PR-, HER2+). In another example, the breast cancer may be resistant to therapy or therapies such as alkylating agents, platinum agents, taxanes, vinca agents, anti-estrogen drugs, aromatase inhibitors, ovarian suppression agents, endocrine / hormonal agents, bisphophonate therapy agents or targeted biological therapy agents.
[0154] A locus (e.g., SNP or genomic locus) or allele is “correlated” or “associated” with a specified phenotype (e.g., increased risk of developing breast cancer) when it can be statistically linked (positively or negatively) to the phenotype. For example, a specified polymorphism may occur more commonly in a case population (e.g., breast cancer patients) than in a control population (e.g., individuals that do not have breast cancer). This correlation may suggest some natural or biological causal link (e.g., a natural law or phenomenon), but it typically does not prove or require such a link (i.e., the correlation is not such a law orphenomenon per se). As used herein, “correlation” refers instead to an artificial statistical linkage between a locus and a trait that underlies the phenotype.
[0155] A region or genomic locus may be “associated” with a SNP when it is statistically linked (positively or negatively) with the SNP. For example, a specified polymorphism in a region near the disease-associated SNP may occur more commonly in conjunction with the disease-specific SNP than other polymorphisms. This correlation may arise from linkage disequilibrium as the specified polymorphism and the disease-associate SNP as there is a low level of crossing over that occurs between the polymorphism and the disease associated SNP.
[0156] As used herein, the term “diagnosis” refers to methods by which a determination can be made as to whether an individual has or is likely to have a given clinical characteristic (e.g., risk of developing cancer). The skilled artisan often makes a diagnosis on the basis of one or more diagnostic indicators, e.g., a biomarker, the presence, absence, amount, or change in amount of which may indicate the presence, severity, or absence of the condition. Other diagnostic indicators can include patient history; physical symptoms, e.g., unexplained weight loss, fever, fatigue, pains, or skin anomalies; phenotype; genotype; or environmental or heredity factors. A skilled artisan will understand that the term “diagnosis” often refers to an increased probability or likelihood that given clinical characteristic is present or will occur; that is, that a clinical characteristic is more likely to be present or to occur in a patient exhibiting a given feature, e.g., the presence or level of a diagnostic indicator, when compared to individuals not exhibiting the feature. Diagnostic methods can be used independently, or in combination with other diagnosing methods known in the art to determine whether a clinical characteristic is present or is more likely to occur in a patient exhibiting a given feature.
[0157] As used herein, “disease” can encompass any disorder, condition, sickness, ailment, etc. that manifests in, e.g., a disordered or incorrectly functioning organ, part, structure, or system of the body, and results from, e.g., genetic or developmental errors, infection, poisons, nutritional deficiency or imbalance, toxicity, or unfavorable environmental factors.
[0158] As used herein, “genotype” means the genetic constitution of an individual (or group of individuals) at one or more genetic loci. Genotype is defined by the allele(s) of one or more known loci of the individual, typically, the compilation of alleles inherited from itsparents. In most aspects and embodiments of the present disclosure, the genotype will be the nucleotide (adenine, thymine (or uracil), cytosine, guanine) at a particular locus in either one or both (typically both) alleles of a subject’s genome or chromosomes. With respect to a particular nucleotide position or locus, the nucleotide(s) at that locus or equivalent thereof in one or both alleles form the genotype of the locus. A genotype can typically be homozygous (e.g., A / A) or heterozygous (e.g., A / B), though more complex genotypes are possible (e.g., AA / A, AA / B, etc.). Accordingly, “genotyping” or determining the genotype for a particular locus means determining the nucleotide(s) at a particular gene locus. One example of this is “detecting” the genotype at a locus, which means determining through a physical assay the physical presence (and optionally quantity) of the nucleotides at a given locus in a patient’s genome or chromosomes. Genotyping can also be done by determining the amino acid variant at a particular position of a protein which can be used to deduce the corresponding nucleotide variant(s).
[0159] As used herein, “haplotype” means the genotype of an individual at a plurality of genetic loci on a single DNA strand. Typically, the genetic loci described by a haplotype are physically and genetically linked, z.e., on the same chromosome strand.
[0160] As used herein, “high stringency hybridization conditions,” when used in connection with nucleic acid hybridization, means conditions capable of restricting hybridization between nucleic acid molecules in a reaction to only those molecules sufficiently homologous to hybridize under the following conditions: hybridization conducted overnight at 42°C in a solution containing 50% formamide, 5xSSC (750 mM NaCl, 75 mM sodium citrate), 50 mM sodium phosphate, pH 7.6, 5x Denhardt’s solution, 10% dextran sulfate, and 20 microgram / ml denatured and sheared salmon sperm DNA, with hybridization filters washed in O.lxSSC at about 65°C. In some embodiments “high stringency hybridization conditions” means the preceding hybridization conditions. The term “moderate stringent hybridization conditions,” when used in connection with nucleic acid hybridization, means conditions capable of restricting hybridization between nucleic acid molecules in a reaction to only those molecules sufficiently homologous to hybridize under the following conditions: hybridization conducted overnight at 37°C in a solution containing 50% formamide, 5xSSC (750 mM NaCl, 75 mM sodium citrate), 50 mM sodium phosphate, pH 7.6, 5x Denhardt’s solution, 10% dextran sulfate, and 20 microgram / ml denatured and sheared salmon sperm DNA, with hybridization filters washed in IxSSC at about 50°C. Insome embodiments “moderate stringency hybridization conditions” means the preceding hybridization conditions. It is noted that many other hybridization methods, solutions and temperatures can be used to achieve comparable stringent hybridization conditions as will be apparent to skilled artisans apprised of the present disclosure.
[0161] As used herein, a patient has an “increased risk” of a particular cancer if the probability of the patient developing that cancer (e.g., over the patient’s lifetime, over some defined period of time (e.g., within 10 years), etc.) exceeds some reference probability or value. The reference probability may be the probability (ie., prevalence) of the cancer across the general relevant patient population (e.g., all patients; all patients of a particular age, gender, ethnicity; patients having a particular cancer (and thus looking at the risk of a different cancer or an independent second primary of the same type as the first cancer); etc.). For example, if the lifetime probability of a particular cancer in the general population (or some specific subpopulation) is X% and a particular patient has been determined by the methods, systems or kits of the present disclosure to have a lifetime probability of that cancer of Y%, and if Y > X, then the patient has an “increased risk” of that cancer. Alternatively, the tested patient’s probability may only be considered “increased” when it exceeds the reference probability by some threshold amount (e.g, at least 0.5, 0.75, 0.85, 0.90, 0.95, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more fold or standard deviations greater than the reference probability; at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% greater than the reference probability).
[0162] The phrase “linkage disequilibrium” (LD) is used to describe the statistical correlation between two polymorphic genotypes (often neighboring). Typically, LD refers to the correlation between the alleles of a random gamete at the two loci, assuming Hardy- Weinberg equilibrium (statistical independence) between gametes. LD can be quantified with either Lewontin’s parameter of association (D’) or with Pearson correlation coefficient (r) (Devlin and Risch, 1995). Two loci with a LD value of 1 are generally said to be in complete LD. At the other extreme, two loci with a LD value of 0 are generally termed to be in linkage equilibrium. Linkage disequilibrium can be calculated following the application of the expectation maximization algorithm (EM) for the estimation of haplotype frequencies (Slatkin and Excoffier, 1996). LD values according to the present disclosure forgenotypes / loci are selected above 0.1, above 0.2, above 0.5, above 0.6, above 0.7, above 0.8, above 0.9, or about 1.0.
[0163] Another way one of skill in the art can identify SNPs in linkage disequilibrium with SNPs of the present disclosure is determining the LOD score for two loci. LOD stands for “logarithm of the odds”, a statistical estimate of whether two loci (e.g., or a locus and a disease locus) are likely to be located near each other on a chromosome and are therefore likely to be inherited together. A LOD score of between about 2-3 or higher is generally understood to suggest that two genes are located close to each other on the chromosome. In some embodiments, LOD values according to the present disclosure for genotypes / loci are selected at least above 2, at least above 3, at least above 4, at least above 5, at least above 6, at least above 7, at least above 8, at least above 9, at least above 10, at least above 20 at least above 30, at least above 40, at least above 50.
[0164] In some embodiments, SNPs in linkage disequilibrium with the SNPs of the present disclosure can have a specified genetic recombination distance of less than or equal to about 20 centimorgan (cM) or less. For example, 15 cM or less, 10 cM or less, 9 cM or less, 8 cM or less, 7 cM or less, 6 cM or less, 5 cM or less, 4 cM or less, 3 cM or less, 2 cM or less, 1 cM or less, 0.75 cM or less, 0.5 cM or less, 0.25 cM or less, or 0.1 cM or less. For example, two linked loci within a single chromosome segment can undergo recombination during meiosis with each other at a frequency of less than or equal to about 20%, about 19%, about 18%, about 17%, about 16%, about 15%, about 14%, about 13%, about 12%, about 11%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.75%, about 0.5%, about 0.25%, or about 0.1% or less.
[0165] In some embodiments, SNPs in linkage disequilibrium with the SNPs of the present disclosure are within at least 100 kb (which correlates in humans to about 0.1 cM, depending on local recombination rate), at least 50 kb, at least 20 kb or less of each other.
[0166] One example approach for the identification of surrogate markers for a particular SNP involves a strategy that presumes that SNPs surrounding the target SNP are in linkage disequilibrium and can therefore provide information about disease susceptibility. Thus, as described herein, surrogate markers can therefore be identified from publicly available databases, such as HAPMAP, by searching for SNPs fulfilling certain criteria whichhave been found in the scientific community to be suitable for the selection of surrogate marker candidates.
[0167] As used herein, “locus” or “genomic locus” or “genomic loci” means a specific position or site in a gene (or protein), chromosomal region, chromosome, or genome. As used herein, a “test locus” is a genomic locus (e.g., single nucleotide at a specified position within a chromosome) whose sequence or genotype is assessed according to the present disclosure. A test locus in the present disclosure is often, though not necessarily, a single nucleotide polymorphism. As used herein, “single nucleotide polymorphism” or “SNP” means a genetic variation between individuals; e.g., a single nitrogenous base position in the DNA of organisms that is polymorphic or variable. As used herein, “SNPs” is the plural of SNP. The identifier used herein for SNP loci (e.g., in Tables 1 and 2) is the “rs” identifier often used in the art. This identifier is used, e.g., in the dbSNP database available through the NCBI website and may be updated for changed for any given locus over time. Thus, any “rs” identifier used herein is expressly meant to include new or modified “rs” identifiers assigned to the same locus (i.e., the locus to which the “rs” identifier is assigned in Tables 1 or 2 and in dbSNP as of the date of the filing of this disclosure. References to DNA herein may include derivatives of any given source of DNA such as amplicons, RNA transcripts thereof, etc. As used herein, a “polymorphism” or “polymorphic” locus or position is a locus or position that is variable; that is, within a population, the nucleotide sequence at a polymorphism has more than one version or allele. One example of a polymorphism is a “single nucleotide polymorphism.”
[0168] As used herein, “marker,” “molecular marker” or “marker nucleic acid” means to a nucleotide sequence or encoded product thereof (e.g, a protein) used as a point of reference when identifying a locus or a linked locus. A marker can be derived from genomic nucleotide sequence or from expressed nucleotide sequences (e.g, from an RNA, nRNA, mRNA, a cDNA, etc.), or from an encoded polypeptide. The term also refers to nucleic acid sequences complementary to or flanking the marker sequences, such as nucleic acids used as probes or primer pairs capable of amplifying the marker sequence. A “marker probe” is a nucleic acid sequence or molecule that can be used to identify the presence of a marker locus, e.g., a nucleic acid probe that is complementary to a marker locus sequence or to sequences adjacent to or near such marker locus sequence. Nucleic acids are “complementary” when they specifically hybridize in solution, e.g., according to Watson-Crick base pairing rules and atcertain minimum hybridization conditions (e.g., medium stringency). A “marker locus” is a locus that can be used to track the presence of a second linked locus, e.g., a linked or correlated locus that encodes or contributes to the population variation of a phenotypic trait. For example, a marker locus can be used to monitor segregation of alleles at a locus that is genetically or physically linked to the marker locus. Thus, a “marker allele,” alternatively an “allele of a marker locus” is one of a plurality of polymorphic nucleotide sequences found at a marker locus in a population that is polymorphic for the marker locus. Each of the identified markers is expected to be in close physical and genetic proximity (resulting in physical and / or genetic linkage) to a genetic element that contributes to the relevant phenotype. Markers corresponding to genetic polymorphisms between members of a population can be analyzed (e.g., detected, measured, quantified) by several techniques. These include, e.g., PCR-based sequence specific amplification methods, detection of restriction fragment length polymorphisms (RFLP), detection of isozyme markers, detection of allele specific hybridization (ASH), detection of single nucleotide extension, detection of amplified variable sequences of the genome, detection of self-sustained sequence replication, detection of simple sequence repeats (SSRs), detection of single nucleotide polymorphisms (SNPs), or detection of amplified fragment length polymorphisms (AFLPs).
[0169] As used herein, “next generation sequencing” or “NGS” refers to a variety of high-throughput sequencing technologies that parallelize the sequencing process, producing thousands or millions of sequences at once. NGS is generally conducted with the following steps: First, DNA sequencing libraries are generated by clonal amplification by PCR in vitro,' second, the DNA is sequenced by synthesis, such that the DNA sequence is determined by the addition of nucleotides to the complementary strand rather through chain-termination chemistry typical of Sanger sequencing; third, the spatially segregated, amplified DNA templates are sequenced simultaneously in a massively parallel fashion, typically without the requirement for a physical separation step. NGS parallelization of sequencing reactions can generate hundreds of megabases to gigabases of nucleotide sequence reads in a single instrument run. Unlike conventional sequencing techniques, such as Sanger sequencing, which typically report the average genotype of an aggregate collection of molecules, NGS technologies typically digitally tabulate the sequence of numerous individual DNA fragments (sequence reads discussed in detail below), such that low frequency variants (e.g., variants present at less than about 10%, 5% or 1% frequency in a heterogeneous population of nucleic acid molecules) can be detected. The term “massively parallel” can also be used to refer tothe simultaneous generation of sequence information from many different template molecules by NGS.
[0170] NGS strategies can include several methodologies, including, but not limited to: (i) microelectrophoretic methods; (ii) sequencing by hybridization; (iii) real-time observation of single molecules, and (iv) cyclic-array sequencing. Cyclic-array sequencing refers to technologies in which a sequence of a dense array of DNA is obtained by iterative cycles of template extension and imaging-based data collection. Commercially available cyclic-array sequencing technologies include, but are not limited to, 454 sequencing, for example, used in 454 Genome Sequencers (Roche Applied Science; Basel), Solexa technology, for example, used in the Illumina Genome Analyzer, Illumina HiSeq, MiSeq, and NextSeq (San Diego, CA), the SOLiD platform (Applied Biosystems; Foster City, CA), the Polonator (Dover / Harvard) and HeliScope Single Molecule Sequencer technology (Helicos;Cambridge, MA). Other NGS methods include single molecule real time sequencing (e.g., Pacific Bio) and ion semiconductor sequencing (e.g., Ion Torrent sequencing). See, e.g., Shendure & Ji, Next Generation DNA Sequencing, NAT. BIOTECH. (2008) 26: 1135-1145 for a more detailed discussion of NGS sequencing technologies.
[0171] As used herein, “patient” or “individual” or “subject” refers to a human. A subject can be male or female.
[0172] As used herein, “sample” or “biological sample” refers to samples such as biopsy or tissue samples, frozen samples, blood and blood fractions or products (e.g., serum, platelets, red blood cells, and the like), tumor samples, sputum, bronchoalveolar lavage, cultured cells, e.g., primary cultures, explants, and transformed cells, stool, urine, etc. A “biopsy” refers to the process of removing a tissue sample for diagnostic or prognostic evaluation, and to the tissue specimen derived from such a process. Any suitable biopsy technique can be applied to the methods of the present disclosure. The biopsy technique applied will depend on the tissue type to be evaluated (e.g., lung etc.), the size and type of the tumor, among other factors. Representative biopsy techniques include, but are not limited to, excisional biopsy, incisional biopsy, needle biopsy, surgical biopsy, and bone marrow biopsy. An “excisional biopsy” refers to the removal of an entire tumor mass with a small margin of normal tissue surrounding it. An “incisional biopsy” refers to the removal of a wedge of tissue that includes a cross-sectional diameter of the tumor. A diagnosis made by endoscopy or fluoroscopy can require a “core-needle biopsy”, or a “fine-needle aspiration biopsy” whichgenerally obtains a suspension of cells from within a target tissue. A “bodily fluid” include all fluids obtained from a mammalian body, either processed (e.g., serum) or unprocessed, which can include, for example, blood, plasma, urine, lymph, gastric juices, bile, serum, saliva, sweat, and spinal and brain fluids. A biological sample is typically obtained from a subject. As used herein, “cancer cell samples” or “tumor sample” means a specimen comprising either at least one cancer cell or biomolecules derived therefrom, including without limitation, lung cancer (e.g, non-small cell lung cancer (NSCLC)), ovarian cancer, colorectal cancer, breast cancer, endometrial cancer, or prostate cancer. Non-limiting examples of such biomolecules include nucleic acids and proteins. Biomolecules “derived” from a cancer cell sample include molecules located within or extracted from the sample as well as artificially synthesized copies or versions of such biomolecules. One illustrative, non-limiting example of such artificially synthesized molecules includes PCR amplification products in which nucleic acids from the sample serve as PCR templates. “Nucleic acids of’ a cancer cell sample include nucleic acids located in a cancer cell or biomolecules derived from a cancer cell.
[0173] As used herein, “sequence read” means the sequence of an individual DNA molecule sequenced in a sequencing reaction. Especially in next-generation sequencing, individual DNA molecules used for sequencing can be relatively short (e.g., ranging from 50nt to l,000nt). These molecules are typically heavily overlapping in their sequences. Thus, any individual test locus is contained within numerous distinct DNA molecules in the sample. When each individual molecule is sequenced (often in parallel), the numerous resulting “sequence reads” can be aligned against each other and / or against a larger reference sequence (e.g., a reference human genome sequence such as the hgl9 version of the human genome assembly available at the University of California Santa Clara’s Genome Browser website). Generally speaking, a greater number of reliably sequenced (or “informative”) reads containing (or “covering”) any individual locus yields greater accuracy and confidence in the genotype / sequence at that locus. Thus, in some specific embodiments of each of the above aspects of the disclosure a test locus (or an allele at that locus) may be counted only if it is covered by at least some minimal number of sequence reads in the sequencing reaction(s).
[0174] As used herein, “score” means a value or set of values selected so as to provide a quantitative measure or assessment of a variable or characteristic of a subject or the subject’s condition or physiology. The value(s) comprising the score can be based on, derived from or incorporate, for example, quantitative data resulting in a measured amount of one or moresample constituents obtained from the subject. In certain embodiments the score can be derived from a single constituent, parameter or assessment, while in other embodiments the score is derived from multiple constituents, parameters and / or assessments. The score can be based upon or derived from an interpretation function; e.g., an interpretation function derived from a particular predictive model using any of various statistical algorithms. A “change in score” can refer to the absolute change in score, e.g. from one time point to the next, or the percent change in score, or the change in the score per unit time (i.e., the rate of score change).
[0175] As used herein, the term “treatment” or “therapy” or “therapeutic regimen” includes all clinical management of a subject and interventions, whether biological, chemical, physical, or a combination thereof, intended to sustain, ameliorate, improve, or otherwise alter the condition of a subject. These terms may be used synonymously herein. Treatments include but are not limited to administration of prophylactics or therapeutic compounds (including small molecule and biologic drugs), exercise regimens, physical therapy, dietary modification and / or supplementation, bariatric surgical intervention, administration of therapeutic compounds (prescription or over-the-counter), and any other treatments known in the art as efficacious in preventing, delaying the onset of, or ameliorating disease characterized by HML. A “response to treatment” includes a subject’s response to any of the above-described treatments, whether biological, chemical, physical, or a combination of the foregoing. A “treatment course” relates to the dosage, duration, extent, etc. of a particular treatment or therapeutic regimen. An initial therapeutic regimen as used herein is the first line of treatment.
[0176] As used herein, “variant allele ratio” means the proportion of informative sequence reads harboring a particular nucleotide at a specific locus as a proportion of the total sequence reads. For example, if a test locus is covered by 100 informative sequence reads in a particular sequencing reaction and 15 reads carry a particular nucleotide (e.g., a risk modifying allele), then the risk modifying allele ratio is 15%. In some contexts variant allele ratios that are too low or too high may indicate unreliability in an allele or genotype call (sometimes referred to herein as a call failure). For example, if the variant allele ratio is around 1%, this can in many cases be due to sequencing artifacts and noise (e.g., a small proportion of sequence reads simply contain sequencing errors). Thus, in some specific embodiments of each of the above aspects of the disclosure a test locus (or an allele orspecific nucleotide at that locus) may be counted only if the variant allele ratio is within a specific (e.g., pre-specified) range.Study Subjects
[0177] The study sets were derived from cohorts of women referred for hereditary cancer testing with a multi-gene panel. Patient data were eligible for inclusion if they were from women referred for hereditary cancer testing between ages 18 and 84 who tested negative for pathogenic variants in breast cancer-risk genes including one or more of BRCA1, BRCA2, TP53, PTEN, SIKH, CDH1, PAI. 2, CHEK2, ATM, NBN and BARD 1. Cases were defined as patients who had a personal history of invasive breast cancer (BC) and controls were defined as patients with no personal history of BC, ductal carcinoma in situ, lobular carcinoma in situ, atypical hyperplasia, or other breast disease at the time of consent. As part of the test request form, self-reported ancestry was collected according to the 1997 United States Office of Management and Budget standards on race and ethnicity.
[0178] Monogenic Pathogenic Mutation Screening and SNP Genotyping Methods
[0179] In certain aspects, study subjects were screened for the germline presence of pathogenic mutations in the following genes: APC, ATM, BARD1, BMPR1A, BRCA1, BRCA2, BRIP1, CDH1, CDK4, CDKN2A (pl4ARF, p 16), CHEK2, EPCAM, MLH1, MSH2, MSH6, MUTYH, NBN, PALB2, PMS2, PTEN, RAD51C, RAD51D, SMAD4, STK11, and TP53. Details of the screening methodology and statistical variant classification methods have been previously described. (Judkins et al., BMC Cancer 2015; 15:215; Pruss et al., Breast Cancer Res Treat 2014;147: 119-32.) In some embodiments, long-range and nested PCR were applied to segments of the CHEK2 gene to exclude pseudogene sequences.Sequencing may be performed using methods know in the art, for example using Illumina instruments (Illumina Inc., San Diego, CA) to identify both sequence variants and large rearrangements (including deletions and duplications).
[0180] SNP markers are genotyped using standard methods. For example, genotyping may be performed using hybrid selection of SNP targets for breast cancer risk and SNPs for genetic ancestry followed by NGS as described previously (Hughes et al., JCO Precision Oncology 2020:585-92). SNP markers associated with breast cancer risk are detailed in Table 2. SNPs used to determine genetic ancestry are listed in Table 1. For a subset of samples,ancestry variants were determined by targeted PCR with a custom rh-Amp genotyping pool (Integrated DNA Technologies, Coralville, Iowa) followed by NGS.Ancestry-informative SNPs
[0181] In general, a polygenic determination of a trait in a subject can be done with a set of polygenic SNP markers. In some embodiments, the trait can be ancestry.
[0182] Aspects of this disclosure provide advantages in characterizing the genotype of a subject according to ancestry.
[0183] In certain aspects, methods of this disclosure can use SNPs associated with one or more different heritage groups. In certain aspects, methods of this disclosure can use bases in genomic loci associated with SNPs associated with one or more different heritage groups. In certain aspects, methods of this disclosure can use genomic loci associated with SNPs associated with one or more different heritage groups. Combinations of SNPs or bases within genomic loci associated with these SNPs or genomic loci associated with these SNPs can be used for assessing the ancestry of a subject. A genotype of a subject can be determined based on fractional ancestry of one or more different heritage groups.
[0184] Embodiments of this disclosure provide methods for assessing ancestry of a subject by selecting a plurality of ancestry-informative SNP markers. Embodiments of this disclosure also provide methods for assessing ancestry of a subject by selecting a plurality of bases in genomic loci associated with ancestry-informative SNP markers. Embodiments of this disclosure also provide methods for assessing ancestry of a subject by selecting a plurality of genomic loci associated with ancestry-informative SNP markers. The ancestry- informative SNP markers can be based on one or more criteria such as the ability to substantially cover the entirety of the human genome, having at least 1% genomic frequency, and having different frequencies in different heritage populations. By obtaining the genotype of a subject, a fractional heritage in the genotype of the subject can be calculated for each of the different heritage populations based on the plurality of ancestry-informative SNP markers or bases in genomic loci associated with ancestry-informative SNP markers or genomic loci associated with ancestry -informative SNP markers.
[0185] In some embodiments, the ancestry-informative SNP markers can have different frequencies in four or more different heritage populations, such as in African, European, East Asian, and Amerindian heritage populations. A plurality of from 10 to 50,000 ancestryinformative SNP markers or 10 to 50,000 bases within genomic loci associated with ancestry- informative SNP markers or 10 to 50,000 genomic loci associated with ancestry -informative SNP markers can be used.
[0186] In some embodiments, the plurality of ancestry-informative SNP markers may include from 1-1,000,000 SNP markers.
[0187] In certain embodiments of this disclosure, a plurality of 10 to 56 ancestry- informative SNP markers can been used.
[0188] In some embodiments, the bases associated with ancestry-informative SNP markers may include 1-1,000,000 bases in genomic loci associated with SNP markers or 1- 1,000,000 genomic loci associated with the SNP markers. In certain embodiments, a plurality of 10 to 56 bases within genomic loci associated with ancestry -informative SNP markers can be used. In certain embodiments, a plurality of 10 to 56 genomic loci associated with ancestry -informative SNP markers can be used.
[0189] Methods of this disclosure can combine the use of the ancestry-informative SNP markers with additional SNP markers that may be associated with a biological trait.Additionally or alternatively, methods of this disclosure can combine the use of the bases in genomic loci associated with the ancestry -informative SNP or genomic loci associated with the ancestry-informative SNP markers with additional SNP markers that may be associated with a biological trait. Combinations of SNPs or bases within genomic loci associated with the informative SNPs can be used to provide a multiple-ancestry polygenic risk score (MA- PRS), which can stratify subjects for risk of the trait regardless of heritage. A multipleancestry polygenic risk score can inherently incorporate genomic information based on fractional ancestry. Additionally or alternatively, methods of this disclosure can combine the use of the ancestry-informative SNP markers with additional SNP markers that may be associated with a biological trait, bases in genomic loci associated with the additional SNP markers that may be associated with a biological trait, or genomic loci associated with the additional SNP markers that may be associated with a biological trait. Additionally or alternatively methods of this disclosure can combine the use of the bases in genomic loci associated with the ancestry-informative SNP or genomic loci associated with the ancestry- informative SNP markers with additional SNP markers that may be associated with a biological trait, bases in genomic loci associated with the additional SNP markers that may be associated with a biological trait, or genomic loci associated with the additional SNP markers that may be associated with a biological trait.
[0190] Aspects of this disclosure provide methods for multiple-ancestry polygenic risk scoring that rely on a unique set of SNPs or a unique set of bases in genomic loci associated with the informative SNPs or a unique set of genomic loci associated with the informative SNPs discovered through design criteria. This unique set of SNPs, bases in genomic loci associated with the informative SNPs, or genomic loci associated with the informative SNPs for multiple-ancestry polygenic risk estimation can discriminate risk for all heritage groups and populations. The unique set of SNP markers, bases in genomic loci associated with the informative SNPs, or genomic loci associated with the informative SNPs for polygenic risk estimation disclosed herein provide accurate results for all heritage groups and populations, substantially without bias toward any population or heritage group.
[0191] In certain aspects, a set of ancestry-informative SNP markers, bases in genomic loci associated with the informative SNPs, or genomic loci associated with the informative SNPs was discovered using estimates of individual SNP risk betas in different ancestry groups. In some embodiments, for each ancestry, individual SNP risk betas were determined from known values, from data obtained in myRisk patients, and through meta-analysis of combined data of the foregoing. Estimates of individual SNP risk betas in different ancestry groups can be used to determine a unique set of SNP markers, bases in genomic loci associated with the informative SNPs, or genomic loci associated with the informative SNPs that can provide a multiple-ancestry polygenic risk score for risk of a trait, such as cancer risk, and which can stratify unaffected patients for risk irrespective of ancestry.
[0192] For example, in some embodiments, African SNP risk betas can be determined from 1,000 or more, or from 5,000 or more, or from 10,000 or more myRisk measurements of patients of self-reported African ancestry. About seventy Asian SNP risk betas can be determined from Shu et al., Nat Commun., 2020, Vol. 11, pp. 1217-1226; Ho et al., Genetics in Medicine, 2022, pp. 586-600; and Ho et al., Nat. Commun., 2020, Vol. 11, pp. 3833. Hispanic SNP risk betas can be determined from 1,000 or more, or from 5,000 or more, or from 10,000 or more Hispanic myRisk measurements of patients of self-reported Hispanic ancestry. Amerindian SNP risk betas can be determined from myRisk measurements of patients of genetic Amerindian ancestry in Simmons et al., Prevention, Risk, Reduction, and Genetics, 2024, Vol. 42: 16, pp. 10533.
[0193] Embodiments of this disclosure can provide a multiple-ancestry polygenic risk score for a trait that can be clinically validated for all women of all heritage groups and populations.
[0194] A multiple-ancestry polygenic risk score of this disclosure can provide meaningful risk discrimination of a trait for all women of all heritage groups and populations.
[0195] A multiple-ancestry polygenic risk score of this disclosure can provide a statistical distribution of scores for a trait for a population, where the scores can be centered at zero with no bias for any ancestry-specific subpopulation.
[0196] FIG. 1 shows an illustration of ancestry in terms of contributions from different continents.
[0197] FIG. 2 shows an illustration of a distribution of genotypes based on ancestry for Hispanic, White / non-Hispanic, Bl ack / African, and Asian self-reported ancestry categories.
[0198] A unique set of ancestry -informative SNPs, bases in genomic loci associated with ancestry -informative SNPs, or genomic loci associated with ancestry-informative SNPs for multiple-ancestry polygenic risk estimation can be obtained by characterizing the ancestry of a subject in terms of contributions from different continents.
[0199] In further embodiments, a set of ancestry-informative SNP markers, bases in genomic loci associated with ancestry-informative SNPs, or genomic loci associated with ancestry-informative SNPs can be obtained by design criteria to distinguish between four continental ancestries: African, East Asian, European, and Amerindian. Examples of ancestry -informative SNPs may be found in Table 1.
[0200] Ancestry -informative SNPs of this disclosure include those in Table 1.Table 1 : Ancestry-informative SNPs
[0201] Bases in genomic loci associated with or genomic loci associated with these and other ancestry-informative SNPs can be determined by determining regions that are in linkage disequilibrium with the ancestry-informative SNPs. Determining linkage disequilibrium can be done by estimating the frequency of recombination during meiosis between the bases in the genomic loci associated with the SNP markers or the genomic loci associated with the SNP markers and the SNP markers. The frequency may be less than or equal to about 20%, about 19%, about 18%, about 17%, about 16%, about 15%, about 14%, about 13%, about 12%, about 11%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.75%, about 0.5%, about 0.25%, or about 0.1% or less. Additionally or alternatively, determining linkage disequilibrium can be done by identifying bases in genomic loci or genomic loci that are within 0.2 cM or less, 0.1 cM or less, or within 0.05 cM or less of the SNP marker. Additionally or alternatively, determining linkage disequilibrium can be done by identifying bases in genomic loci or genomic loci that have a linkage disequilibrium value of above 0.1, above 0.2, above 0.5, above 0.6, above 0.7, above 0.8, above 0.9, or about 1.0 with the SNP marker. Additionally or alternatively, determining linkage disequilibrium can be done by identifying bases in genomic loci or genomic loci that have a logarithm of odd score of at least above 2, at least above 3, at least above 4, at least above 5, at least above 6, at least above 7, at least above 8, at least above 9, at least above 10, at least above 20 at least above 30, at least above 40, at least above 50 with the SNP markers. Additionally or alternatively, determining linkage disequilibrium can be done by identifying bases in genomic loci or genomic loci that are within about 200 bases, within about 150 bases, within 100 bases, within about 90 bases, within about 80 bases, within about 70 bases, within about 60 bases, within about 50 bases, within about 40 bases, within about 30 bases, within about 20 bases, and within about 10 bases of the SNP marker.Multiple-ancestry polygenic risk estimation for cancer
[0202] Embodiments of this disclosure further contemplate combining a unique set of additional SNP markers, genomic loci associated with the additional SNP markers, or bases within genomic loci associated with the additional SNP markers associated with a trait such as risk of cancer with ancestry-informative SNP markers, genomic loci associated with ancestry -informative SNP markers, or bases within genomic loci associated with ancestry- informative SNP markers.
[0203] In some embodiments, methods of this disclosure can combine the use of ancestry -informative SNP markers, genomic loci associated with ancestry-informative SNP markers, or bases within genomic loci associated with ancestry -informative SNP markers with additional SNP markers, genomic loci associated with the additional SNP markers, or bases within genomic loci associated with the additional SNP markers that may be associated with cancer risk in one or more different heritage groups. Combinations of such SNPs, genomic loci associated with the ancestry -informative SNP markers, or bases within genomic loci associated with ancestry-informative SNP markers can be used to provide a multiple-ancestry polygenic risk score (MA-PRS) for cancer risk, which can stratify unaffected patients for cancer risk, irrespective of ancestry and the presence or absence of a family history of the disease.
[0204] In further embodiments, methods of this disclosure can combine the use of ancestry-informative SNP markers, genomic loci associated with the ancestry-informative SNP markers, or bases within genomic loci associated with ancestry-informative SNP markers with additional SNP markers, genomic loci associated with the additional SNP markers, or bases within genomic loci associated with the additional SNP markers that may be associated with breast cancer risk in women of one or more different heritage groups. Combinations of such SNPs, genomic loci associated with the ancestry -informative SNP markers, or bases within genomic loci associated with ancestry-informative SNP markers can be used to provide a multiple-ancestry polygenic risk score (MA-PRS) for breast cancer risk, which can stratify unaffected women for breast cancer risk, irrespective of ancestry and the presence or absence of a family history of the disease.
[0205] In some aspects, a multiple-ancestry polygenic risk estimation may characterize the risk of cancer in a subject regardless of the subject’s genetic ancestry by using a combination of ancestry -informative SNP markers, genomic loci associated with the ancestry -informative SNP markers, or bases within genomic loci associated with ancestry- informative SNP markers and cancer-associated SNPs, genomic loci associated with the cancer-informative SNP markers, or bases within genomic loci associated with cancer-informative SNP markers. The cancer-associated SNPs and the cancer-associated genomic loci may be derived from one or more different heritage groups or populations.
[0206] In some aspects, a multiple-ancestry polygenic risk estimation may characterize the risk of breast cancer in a woman regardless of genetic ancestry by using a combination of ancestry-informative SNP markers, genomic loci associated with the ancestry-informative SNP markers, or bases within genomic loci associated with ancestry-informative SNP markers and breast cancer-associated SNPs, genomic loci associated with the breast cancer- associated SNP markers, or bases within genomic loci associated with the breast cancer- associated SNP markers. The breast cancer-associated SNPs, genomic loci associated with the breast cancer-associated SNP markers, or bases within genomic loci associated with the breast cancer-associated SNP markers may be derived from one or more different heritage groups or populations.
[0207] In some embodiments, a multiple-ancestry polygenic risk score may characterize the risk of breast cancer in a woman regardless of genetic ancestry by using a combination of from 10 to 56 ancestry-informative SNP markers (Table 1), genomic loci associated with the ancestry -informative SNP markers, or bases within genomic loci associated with ancestry- informative SNP markers and from 10 to 329 breast cancer-associated SNPs (Table 2), genomic loci associated with the breast cancer-associated SNP markers, or bases within genomic loci associated with the breast cancer-associated SNP markers. The breast cancer- associated SNPs can include 93 previously published breast cancer-associated SNPs (Hughes et al., JCO Precis Oncol 6:e2200084), including up to 92 European breast cancer-associated SNPs and one Hispanic breast cancer SNP 6q25 (rsl40068132), as well as an additional 236 breast cancer-associated SNPs (Table 2), that were selected using a novel synthetic stepwise regression methodology that accounts for linkage disequilibrium.
[0208] In certain embodiments, a multiple-ancestry polygenic risk score may characterize the risk of breast cancer in a woman regardless of genetic ancestry by using a combination of 56 ancestry-informative SNP markers and 329 breast cancer-associated SNPs comprised of 94 European breast cancer-associated SNPs and 237 non-European breast cancer-associated SNPs including one Hispanic breast cancer SNP 6q25 (rsl40068132). In certain embodiments, a multiple-ancestry polygenic risk score may characterize the risk of breast cancer in a woman by measuring ancestry at each SNP or genomic loci and applying ancestry-specific risks and calibrate the overall risk score according to ancestry-specific frequencies.
[0209] A multiple-ancestry polygenic risk score of this disclosure can achieve a high level of accuracy for all women of all heritage groups and populations in terms of surprisingly high risk discrimination and superior accuracy of calibration.Table 2: Breast cancer-associated SNPsExemplary Statistical Methods
[0210] A description of statistical methods used for development and validation of MA- PRS is described below. Briefly, a small number of reference ancestries (African, East Asian, European, and Amerindian) were determined. A set of ancestral markers were selected to enable representation of genetic ancestry, for each individual, in terms of fractions attributable to each reference ancestry. For each reference ancestry, a PRS using ancestryspecific weights and frequencies of BC-associated SNPs, regions associated with the BC- associated SNPs, or bases within regions associated with the BC-associated SNPs was developed. MA-PRS was defined as a combination of ancestry-specific PRSs on the basis of genetic ancestral composition. Finally, the association of MA-PRS with BC risk was validated in a large, independent cohort of patients of diverse ancestry.
[0211] Ancestry-specific PRSs: SNP weights (or effect sizes) were estimated, as log odds ratios (ORs) for association with BC, for each reference ancestry on the basis of the largest accessible data set. African SNP weights were based on meta-analysis of 77,625 Black / African patients and literature (Du et al., JNCI, 2021, 1186-1176). East Asian weights were based on meta analysis of 150,201 East Asian patients and literature (Ho et al., Genetics in Medicine, 2022, 586-600; Ho et al., Nature Commun., 2020, 11 : 3833 ; and Shu et al., Nature Commun., 2020, 11 : 1217).
[0212] European SNP weights were determined from literature (Zhang et al., Nature Genetics, 2020, 52(6):572-581). For one Amerindian SNP (rsl40068132), which is prevalent only among Hispanics, a weight was estimated, denoted by pAm, on the basis of 6,718 selfreported Hispanic patients referred for hereditary cancer testing. At the individual SNP level, Amerindian SNP weights are set equal to the East Asian weights.
[0213] For each reference ancestry, an ancestry-specific PRS was defined as the sum of BC risk alleles, centered by the ancestry-specific allele frequency and weighted by the ancestry-specific effect size multiplied by the probability of the allele having been inherited from that ancestry.
[0214] Development of the MA-PRS'. The development set of 184,322 women was used to combine African-, East Asian-, European-, and Amerindian-specific PRSs (PRSAT, PRSEA, PRSEU, and PRSAI respectively) and the Amerindian SNP genotype (xAm = 0, 1, or 2) into a final MA-PRS. Weights for the African-, East Asian-, European-, and Amerindian-specific PRSs, denoted by BAT, BEA, BEU, and BAI, respectively, were estimated as log ORs from a multivariable logistic regression model with BC status as the dependent variable and age, personal and family cancer history, self-reported ancestry, genetic ancestry, and interaction between self-reported and genetic ancestry as independent variables.
[0215] The final MA-PRS was defined as:BATXPRSAT + BEAXPRSEA + BEUXPRSEU + BA;XPRS : + PAmxXAm.
[0216] Independent validation of the MA-PRS'. BC risk discrimination of the MA-PRS was evaluated in terms of ORs and P values from multivariable logistic regression models. ORs were normalized to the standard deviation of the MA-PRS distribution in women unaffected by BC and reported with 95% Wald Cis (Table 3).
[0217] The primary analysis evaluated the association of MA-PRS with BC after accounting for the same independent variables used in logistic regression models for MA- PRS development.
[0218] This study included two prespecified secondary analyses. First, the MA-PRS was evaluated for improved discrimination compared with a previously published MA-PRS-93 (Hughes et al., JCO Precis Oncol 6:e2200084), which overlaps with the 327 BC-associated SNPs used in the MA-PRS. Second, the primary and first secondary analyses were repeated within subcohorts defined by self-reported ancestry.
[0219] Goodness-of-fit of the relative risks associated with the MA-PRS was evaluated after accounting for clinical factors. The observed versus predicted effect of the MA-PRS on BC risk was assessed by comparing ORs from the continuous MA-PRS with those obtained from analysis of patients binned in categories according to MA-PRS percentiles.
[0220] All analyses were conducted using R version 3.5.3 or higher. P values were calculated from likelihood ratio chi-squared test statistics and are reported as two-sided. Cancer methods and treatment
[0221] Cancer therapy can include surgery, cryoablation, radiation therapy, bone marrow transplant, chemotherapy, immunotherapy, hormone therapy, stem cell therapy, drug therapy, biological therapy, and administration of a pharmaceutical, prophylactic or therapeutic compound including, for example, a biologic or exogenous active agent.
[0222] Examples of treatments include bariatric surgical intervention, physical therapy, diet, and diet supplementation.
[0223] Examples of a cancer biological therapy include adoptive cell transfer, angiogenesis inhibitors, bacillus Calmette-Guerin therapy, biochemotherapy, cancer vaccines, chimeric antigen receptor (CAR) T-cell therapy, cytokine therapy, gene therapy, immune checkpoint modulators, immunoconjugates, monoclonal antibodies, oncolytic virus therapy, and targeted drug therapy.
[0224] Examples of a cancer surgery include lumpectomy, partial mastectomy, total mastectomy, simple mastectomy, modified radical mastectomy, radical mastectomy, and Halsted radical mastectomy.
[0225] Examples of a cancer drug include drugs approved to prevent breast cancer including Evista (Raloxifene Hydrochloride), Raloxifene Hydrochloride, and Tamoxifen Citrate.
[0226] Examples of a cancer drug include drugs approved to treat breast cancer including, Abemaciclib, Abraxane (Paclitaxel Albumin-stabilized Nanoparticle Formulation), Ado-Trastuzumab Emtansine, Afinitor (Everolimus), Afinitor Disperz (Everolimus), Alpelisib, Anastrozole, Aredia (Pamidronate Disodium), Arimidex (Anastrozole), Aromasin (Exemestane), Atezolizumab, Capecitabine, Cyclophosphamide, Docetaxel, Doxorubicin Hydrochloride, Ellence (Epirubicin Hydrochloride), Enhertu (Fam-Trastuzumab Deruxtecan- nxki), Epirubicin Hydrochloride, Eribulin Mesylate, Everolimus, Exemestane, 5-FU (Fluorouracil Injection), Fam-Trastuzumab Deruxtecan-nxki, Fareston (Toremifene), Faslodex (Fulvestrant), Femara (Letrozole), Fluorouracil Injection, Fulvestrant, Gemcitabine Hydrochloride, Gemzar (Gemcitabine Hydrochloride), Goserelin Acetate, Halaven (Eribulin Mesylate), Herceptin Hylecta (Trastuzumab and Hyaluronidase-oysk), Herceptin (Trastuzumab), Ibrance (Palbociclib), Ixabepilone, Ixempra (Ixabepilone), Kadcyla (Ado- Trastuzumab Emtansine), Kisqali (Ribociclib), Lapatinib Ditosylate, Letrozole, Lynparza (Olaparib), Megestrol Acetate, Methotrexate, Neratinib Maleate, Nerlynx (Neratinib Maleate), Olaparib, Paclitaxel, Paclitaxel Albumin-stabilized Nanoparticle Formulation, Palbociclib, Pamidronate Disodium, Peijeta (Pertuzumab), Pertuzumab, Piqray (Alpelisib), Ribociclib, Talazoparib Tosylate, Talzenna (Talazoparib Tosylate), Tamoxifen Citrate, Taxotere (Docetaxel), Tecentriq (Atezolizumab), Thiotepa, Toremifene, Trastuzumab, Trastuzumab and Hyaluronidase-oysk, Trexall (Methotrexate), Tykerb (Lapatinib Ditosylate), Verzenio (Abemaciclib), Vinblastine Sulfate, Xeloda (Capecitabine), and Zoladex (Goserelin Acetate).
[0227] As used herein, the term “disease” includes any disorder, condition, sickness, ailment that manifests in, for example, a disordered or incorrectly functioning organ, part, structure, or system of the body.
[0228] As used herein, the term “sample” includes any biological sample that is isolated from a subject. A sample can include, without limitation, a single cell or multiple cells, fragments of cells, an aliquot of body fluid, whole blood, platelets, serum, plasma, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies, synovial fluid, lymphatic fluid, ascites fluid, and interstitial or extracellular fluid. The term “sample” also encompasses the fluid in spaces between cells, including synovial fluid, gingival crevicular fluid, bone marrow, cerebrospinal fluid (CSF), saliva, mucous, sputum, semen, sweat, urine, or any other bodily fluids. A blood sample can include whole blood or any fraction thereof, including blood cells, red blood cells, white blood cells or leucocytes, platelets, serum and plasma.
[0229] As used herein, the term “subject” includes humans. Humans generally include women and men and others such as non-binary.
[0230] In some embodiments, this disclosure can provide methods for recommending therapeutic regimens, including withdrawal from therapeutic regiments.
[0231] In further embodiments, an odds ratio can provide a clinician with a prognostic picture of a subject’s biological state. Such embodiments may provide subject-specific prognostic information, which can be informative for a therapy decision, and may also facilitate monitoring therapy response. Such embodiments may result in a surprisingly improved treatment, such as better control of a disease, or an increase in the proportion of subjects achieving amelioration of symptoms.
[0232] As used herein, the terms “biologic,” “biotherapy,” and / or “biopharmaceutical” can include pharmaceutical therapy products manufactured or extracted from a biological substance. A biologic can include vaccines, blood or blood components, allergenics, somatic cells, gene therapies, tissues, recombinant proteins, and living cells; and can be composed of sugars, proteins, nucleic acids, living cells or tissues, or combinations thereof.
[0233] As used herein, the terms “therapeutic regimen,” “therapy” and / or “treatment” can include any clinical management of a subject, as well as interventions, whether biological, chemical, physical, or a combination thereof, intended to sustain, ameliorate, improve, or otherwise alter the condition of a subject.
[0234] As used herein, the term “administering” can include the placement of a composition into a subject by a method or route that results in at least partial localization ofthe composition at a desired site such that a desired effect is produced. Routes of administration include both local and systemic administration. Generally, local administration results in more of the composition being delivered to a specific location as compared to the entire body of the subject, whereas systemic administration results in delivery to essentially the entire body of the subject. “Administering” also includes performing physical actions on a subject’s body, including physical therapy, as well as chiropractic care, massage and acupuncture.Devices and systems
[0235] As used herein, the term machine-readable storage medium can comprise, for example, a data storage material that is encoded with machine-readable data or data arrays. The data and machine-readable storage medium may be capable of being used for a variety of purposes, when using a machine programmed with instructions for using said data. Such purposes include storing, accessing and manipulating information relating to the risk of a subject or population over time, or risk in response to treatment, or for drug discovery for inflammatory disease. Data comprising genomic measurements can be implemented in computer programs that are executing on programmable computers, which may comprise a processor, a data storage system, one or more input devices, one or more output devices. Program code can be applied to the input data to perform the functions described herein, and to generate output information. Output information can then be applied to one or more output devices. A computer can be, for example, a personal computer, a microcomputer, or a workstation.
[0236] As used herein, the term computer program can be instruction code implemented in a high-level procedural or object-oriented programming language, to communicate with a computer system. The program may be implemented in machine or assembly language. The programming language can also be a compiled or interpreted language. Each computer program can be stored on storage media or a device such as ROM, or magnetic diskette, and can be readable by a programmable computer for configuring and operating the computer when the storage media or device is read by the computer to perform the described procedures. A health-related or genomic data management system can be considered to be implemented as a computer-readable storage medium, configured with a computer program, where the storage medium causes a computer to operate in a specific manner to perform various functions.Conclusion
[0237] All publications, patents and literature specifically mentioned herein are hereby incorporated by reference in their entirety for all purposes.
[0238] A reference SNP ID number, or “rs” ID, is an identification tag assigned by NCBI to a group or cluster of SNPs that map to an identical location. The rs ID number, or rs tag, is assigned after submission. A submitted SNP is evaluated to see if it maps to an identical location as previously submitted SNPs; if it does, then the submitted SNP is linked into the reference set of the existing reference SNP record. These SNP rs IDs are mapped to external resources or databases, including NCBI databases. The SNP rs ID number is noted on the records of these external resources and databases in order to point users back to the original dbSNP records. A reference SNP record has the format NCBI| rs<NCBI SNP ID>.
[0239] Words specifically defined herein have the meaning provided in the context of the present disclosure as a whole, and as are typically understood by those skilled in the art. As used herein, the singular forms “a,” “an,” and “the” include the plural.
[0240] While the present disclosure is described in conjunction with various embodiments, it is not intended that the present disclosure be limited to such embodiments. On the contrary, the present disclosure encompasses various alternatives, modifications, and equivalents, as will be appreciated by those of skill in the art.
[0241] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure, suitable methods and materials are described below. In addition, the materials, methods, and examples herein are illustrative only and not intended to be limiting.
[0242] Although the foregoing disclosure has been described in some detail by way of illustration and examples for purposes of clarity of understanding, it will be understood by persons of skill in the art that various changes and modifications may be practiced within the scope of the disclosure and the appended claims.EXAMPLES
[0243] Example 1 : Multiple-ancestry polygenic breast cancer risk assessment.
[0244] Accurate breast cancer risk assessment is essential to identify women for whom screening and preventative interventions may be lifesaving. Incorporation of polygenic risk scores (PRSs) into clinical models can improve risk prediction, but most PRS have shownsuboptimal performance among non-European subjects. Breast cancer risk assessment was determined by using single-nucleotide polymorphisms (SNPs) with small effects that were aggregated into a multiple-ancestry polygenic risk score (MA-PRS), multiple-ancestry developed and validated for populations of European, African, East Asian, and Amerindian descent. To make a MA-PRS meaningful for all women of all heritage groups and populations, a novel multiple-ancestry PRS (MA-PRS) was determined and validated that utilizes individual ancestral genetic composition.
[0245] A previous multiple MA-PRS (MA-149) based on 56 ancestry-informative and 93 breast cancer- (BC) associated single nucleotide polymorphisms (SNPs) was developed. MA- 149 achieved accuracy for diverse populations by characterizing genetic ancestry at each SNP in terms of fractions attributable to reference ancestries (African, East Asian, European) and applying ancestry-specific risks and frequencies.
[0246] MA-149 significantly outperformed the Tyrer-Cuzick (TC) model, and integration of MA-149 with TC improved predictive accuracy by roughly two-fold over TC alone.
[0247] In order to improve MA-149, the set of BC SNPs was expanded, and the ancestryspecific risks were refined.
[0248] To select an optimal set of BC-associated SNPs, a novel synthetic stepwise regression methodology was developed that accounts for linkage disequilibrium. Women referred for hereditary cancer testing and negative for pathogenic variants in BC genes were divided into consecutive cohorts to (1) refine ancestry-specific SNP risks (N = 58,191 Black / African; N = 27,160 East Asian), (2) develop the PRS (N = 184,322), and (3) conduct independent validation (N = 146,110),
[0249] Predictive accuracy and calibration of the new PRS were evaluated in the full cohort and subpopulations of different ancestry. A multivariable logistics regression adjusted for age, ancestry, and cancer history was used to test for improved BC risk prediction over clinical factors. Improvement over the previously reported MA-PRS-149, and a European PRS (Hughes et al., JCO Precis Oncol. 6:e2200084), was tested for by including additional PRS as covariates in multivariable models. Calibration was assessed through goodness-of-fit tests. Odds ratios (OR) from multivariable regression were reported per standard deviation.
[0250] A set of 385 SNPs (56 ancestry-informative and 329 BC-associated) were selected for the new PRS (MA-385) (Tables 1 and 2). Among non-Europeans, MA-385 was a better BC risk predictor (OR = 1.47, 95% CI: 1.42-1.52) than MA-149 (OR = 1.40, 95% CI: 1.35- 1.45). The strongest associations were observed in Ashkenazi Jewish and Hispanic women (Table 3). MA-385 identified more women at >2-fold increased risk than MA-149 (6.5% vs.2%). Goodness-of-fit tests showed that MA-385 was well-calibrated, while a 385-SNP PRS with European weights was miscalibrated for non-Europeans.Table 3: Association of MA-PRS-385 with BC risk after accounting for clinical factors
[0251] MA-385 was well-calibrated, improved upon clinical factors, and outperformed existing PRS in all tested ancestries. Incorporation of MA-385 into risk assessment could improve the early detection and prevention of BC.
Claims
WHAT IS CLAIMED IS:
1. A method for assessing a risk of a trait in a subject, the method comprising: selecting a plurality of trait-associated SNP markers and a plurality of ancestry- informative SNP markers; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject based on the trait-associated SNP markers and the plurality of ancestry-informative SNP markers.
2. The method of claim 1, wherein the trait-associated SNPs are selected using a synthetic stepwise regression methodology that accounts for linkage disequilibrium between variants.
3. The method of claim 1, wherein the trait-associated SNPs comprise one or more of European breast cancer-associated SNPs, African breast cancer-associated SNPs, East Asian breast cancer-associated SNPs, and Amerindian breast cancer-associated SNPs.
4. The method of claim 1, further comprising calculating the multiple-ancestry polygenic risk score for the risk of the trait in the subject with additional clinical variables of the subject.
5. The method of claim 4, wherein the additional clinical variables are age, personal medical history, and family medical history of the subject.
6. The method of claim 1, wherein the trait is a risk of a disease in the subject.
7. The method of claim 6, wherein the disease is cancer.
8. The method of claim 1, wherein the plurality of ancestry-informative SNP markers are from 10 to 50,000 SNP markers.
9. The method of claim 1, wherein the plurality of ancestry-informative SNP markers are from 10 to 56 SNP markers.
10. The method of claim 1, wherein the trait-associated SNP markers are a plurality of cancer-associated SNP markers.
11. The method of claim 1, wherein the trait-associated SNP markers are a plurality of from 10 to 50,000 breast cancer-associated SNP markers.
12. The method of claim 1, wherein the trait-associated SNP markers are a plurality of from 10 to 329 breast cancer-associated SNP markers.
13. The method of claim 1, wherein the calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject is done with training clinical data of a reference group.
14. The method of claim 1, wherein the calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject is done with validating clinical data of a reference group.
15. The method of claim 1, wherein the genotype of the subject is measured by NGS.
16. The method of claim 1, wherein the plurality of ancestry-informative SNP markers determine a fractional heritage in the genotype of the subject for each of four or more different heritage populations.
17. The method of claim 1, wherein the plurality of ancestry-informative SNP markers determine a fractional heritage in the genotype of the subject for each of African, European, East Asian, and Amerindian heritage populations.
18. The method of claim 1, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject is accurate for subjects in four or more different heritage populations, even when the heritage populations are self-reported.
19. The method of claim 1, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject is accurate for subjects in African, European, East Asian, and Amerindian heritage populations, even when the heritage populations are self-reported.
20. The method of claim 1, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject is calibrated for subjects in four or more different heritage populations so that the risk of the trait is not overestimated in any heritage population.
21. The method of claim 1, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject is calibrated for subjects in African, European, East Asian, andAmerindian heritage populations so that the risk of the trait is not overestimated in any heritage population.
22. The method of claim 1, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject discriminates between low risk and high risk for subjects in three or more different heritage populations.
23. The method of claim 1, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject discriminates between low risk and high risk for subjects in African, European, East Asian, and Amerindian heritage populations.
24. The method of claim 1, wherein the calculating a multiple-ancestry polygenic risk score comprises using clinical cohorts of women of African self-reported ancestry, East Asian self-reported ancestry, European self-reported ancestry, and Amerindian self-reported ancestry.
25. The method of claim 1, wherein the calculating a multiple-ancestry polygenic risk score comprises using the sum of ancestry specific polygenic risk scores weighted according to fractional ancestral composition.
26. The method of claim 1, wherein the multiple-ancestry polygenic risk score is strongly associated with breast cancer in a reference cohort and in sub-cohorts defined by selfreported ancestry.
27. The method of claim 1, wherein the multiple-ancestry polygenic risk score is combined with clinical and / or biological risk factors for accurate risk stratification for all women of all ancestries.
28. The method of claim 1, wherein the calculating a multiple-ancestry polygenic risk score comprises calculating and combining: an African-specific PRS (PRSAT), East Asian-specific PRS (PRSEA), European- specific PRS (PRSEU), and Amerindian-specific PRS (PRSAI); an estimated weight for each ancestry including African (BAT), East Asian (BEA), European (BEU), and Amerindian (BAI); and, an Amerindian SNP genotype (xAm), according to the following equation:BATXPRSAT + BEAXPRSEA + BEUXPRSEU + BA;XPRS I + PAmxXAm.
29. A method for assessing a risk of a trait in a subject, the method comprising: selecting a plurality of bases in genomic loci associated with trait-associated SNP markers and a plurality of bases in genomic loci associated with ancestry-informative SNP markers; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject based on the bases in genomic loci associated with trait-associated SNP markers and the plurality of bases in genomic loci associated with ancestry -informative SNP markers.
30. The method of claim 29, wherein the trait-associated SNPs are selected using a synthetic stepwise regression methodology that accounts for linkage disequilibrium between variants.
31. The method of claim 29, wherein the bases within genomic loci associated with trait- associated SNPs are selected using a synthetic stepwise regression methodology that accounts for linkage disequilibrium between variants.
32. The method of claim 29, wherein the bases within genomic loci associated with trait- associated SNPs are within about 200 bases, within about 150 bases, within 100 bases, within about 90 bases, within about 80 bases, within about 70 bases, within about 60 bases, within about 50 bases, within about 40 bases, within about 30 bases, within about 20 bases, and within about 10 bases of the trait-associated SNPs.
33. The method of claim 29, wherein the bases within genomic loci associated with trait- associated SNPs are within about 0.20 cM or less, 0.1 cM or less, or 0.05 cM or less of the trait-associated SNPs.
34. The method of claim 29, wherein the bases within genomic loci associated with trait- associated SNPs have a linkage disequilibrium value of above 0.1, above 0.2, above 0.5, above 0.6, above 0.7, above 0.8, above 0.9, or about 1.0 with the trait-associated SNPs.
35. The method of claim 29, wherein the bases within genomic loci associated with trait- associated SNPs have a logarithm of the odds score of at least above 2, at least above 3, at least above 4, at least above 5, at least above 6, at least above 7, at least above 8, at leastabove 9, at least above 10, at least above 20 at least above 30, at least above 40, at least above 50 with the trait-associated SNPs.
36. The method of claim 29, wherein the bases within genomic loci associated with trait- associated SNPs have a frequency of recombination during meiosis with the trait-associated SNPs of less than or equal to about 20%, about 19%, about 18%, about 17%, about 16%, about 15%, about 14%, about 13%, about 12%, about 11%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.75%, about 0.5%, about 0.25%, or about 0.1% or less.
37. The method of claim 29, wherein the bases within genomic loci associated with ancestry-informative SNPs are selected using a synthetic stepwise regression methodology that accounts for linkage disequilibrium between variants.
38. The method of claim 29, wherein the bases within genomic loci associated with ancestry -informative SNPs are within about 200 bases, within about 150 bases, within 100 bases, within about 90 bases, within about 80 bases, within about 70 bases, within about 60 bases, within about 50 bases, within about 40 bases, within about 30 bases, within about 20 bases, and within about 10 bases of the ancestry-informative SNPs.
39. The method of claim 29, wherein the bases within genomic loci associated with ancestry-informative SNPs are within about 0.2 cM or less, 0.1 cM or less, or 0.05 cM or less of the ancestry-informative SNPs.
40. The method of claim 29, wherein the bases within genomic loci associated with ancestry-informative SNPs have a linkage disequilibrium value of above 0.1, above 0.2, above 0.5, above 0.6, above 0.7, above 0.8, above 0.9, or about 1.0 with the ancestry- informative SNPs.
41. The method of claim 29, wherein the bases within genomic loci associated with ancestry-informative SNPs have a logarithm of the odds score of at least above 2, at least above 3, at least above 4, at least above 5, at least above 6, at least above 7, at least above 8, at least above 9, at least above 10, at least above 20 at least above 30, at least above 40, at least above 50 with the ancestry-informative SNPs.
42. The method of claim 29, wherein the bases within genomic loci associated with ancestry-informative SNPs have a frequency of recombination during meiosis with the ancestry -informative SNPs of less than or equal to about 20%, about 19%, about 18%, about 17%, about 16%, about 15%, about 14%, about 13%, about 12%, about 11%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, about 2%, about 1%, about 0.75%, about 0.5%, about 0.25%, or about 0.1% or less.
43. The method of claim 29, wherein the trait-associated SNPs comprise one or more of European breast cancer-associated SNPs, African breast cancer-associated SNPs, East Asian breast cancer-associated SNPs, and Amerindian breast cancer-associated SNPs.
44. The method of claim 29, further comprising calculating the multiple-ancestry polygenic risk score for the risk of the trait in the subject with additional clinical variables of the subject.
45. The method of claim 44, wherein the additional clinical variables are age, personal medical history, and family medical history of the subject.
46. The method of claim 29, wherein the trait is a risk of a disease in the subject.
47. The method of claim 46, wherein the disease is cancer.
48. The method of claim 29, wherein the plurality of bases in genomic loci associated with ancestry-informative SNP markers are from 10 to 50,000 bases in genomic loci associated with SNP markers.
49. The method of claim 29, wherein the plurality of bases in genomic loci associated with ancestry-informative SNP markers are from 10 to 56 bases in genomic loci associated with SNP markers.
50. The method of claim 29, wherein the bases in genomic loci associated with trait- associated SNP markers are a plurality of bases in genomic loci associated with cancer- associated SNP markers.
51. The method of claim 29, wherein the bases in genomic loci associated with trait- associated SNP markers are a plurality of from 10 to 50,000 bases in genomic loci associated with breast cancer-associated SNP markers.
52. The method of claim 29, wherein the bases in genomic loci associated with trait- associated SNP markers are a plurality of from 10 to 329 bases in genomic loci associated with breast cancer-associated SNP markers.
53. The method of claim 29, wherein the calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject is done with training clinical data of a reference group.
54. The method of claim 29, wherein the calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject is done with validating clinical data of a reference group.
55. The method of claim 29, wherein the genotype of the subject is measured by NGS.
56. The method of claim 29, wherein the plurality of bases in genomic loci associated with ancestry-informative SNP markers determine a fractional heritage in the genotype of the subject for each of four or more different heritage populations.
57. The method of claim 29, wherein the plurality of bases in genomic loci associated with ancestry-informative SNP markers determine a fractional heritage in the genotype of the subject for each of African, European, East Asian, and Amerindian heritage populations.
58. The method of claim 29, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject is accurate for subjects in four or more different heritage populations, even when the heritage populations are self-reported.
59. The method of claim 29, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject is accurate for subjects in African, European, East Asian, and Amerindian heritage populations, even when the heritage populations are self-reported.
60. The method of claim 29, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject is calibrated for subjects in four or more different heritage populations so that the risk of the trait is not overestimated in any heritage population.
61. The method of claim 29, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject is calibrated for subjects in African, European, East Asian, andAmerindian heritage populations so that the risk of the trait is not overestimated in any heritage population.
62. The method of claim 29, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject discriminates between low risk and high risk for subjects in three or more different heritage populations.
63. The method of claim 29, wherein the multiple-ancestry polygenic risk score for the risk of the trait in the subject discriminates between low risk and high risk for subjects in African, European, East Asian, and Amerindian heritage populations.
64. The method of any one of claims 1-63, wherein the trait is a risk of a disease in the subject.
65. The method of claim 64, wherein the disease is cancer.
66. The method of claim 29, wherein the calculating a multiple-ancestry polygenic risk score comprises using clinical cohorts of women of African self-reported ancestry, East Asian self-reported ancestry, European self-reported ancestry, and Amerindian self-reported ancestry.
67. The method of claim 29, wherein the calculating a multiple-ancestry polygenic risk score comprises using the sum of ancestry specific polygenic risk scores weighted according to fractional ancestral composition.
68. The method of claim 29, wherein the multiple-ancestry polygenic risk score is strongly associated with breast cancer in a reference cohort and in sub-cohorts defined by self-reported ancestry.
69. The method of claim 29, wherein the multiple-ancestry polygenic risk score is combined with clinical and / or biological risk factors for accurate risk stratification for all women of all ancestries.
70. The method of claim 29, wherein the calculating a multiple-ancestry polygenic risk score comprises calculating and combining: an African-specific PRS (PRSAT), East Asian-specific PRS (PRSEA), European- specific PRS (PRSEU), and Amerindian-specific PRS (PRSAI);an estimated weight for each ancestry including African (BAT), East Asian (BEA), European (BEU), and Amerindian (BAI); and, an Amerindian SNP genotype (xAm), according to the following equation: BATXPRSAT + BEAXPRSEA + BEUXPRSEU + BA;XPRS I + PAmxXAm.
71. A method for treating a disease in a subject in need thereof, the method comprising: selecting a plurality of disease-associated SNP markers and a plurality of ancestry- informative SNP markers; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the disease in the subject based on the plurality of the disease-associated SNP markers and ancestry- informative SNP markers and, wherein the score indicates a need for treating the subject; and administering to the subject a therapy for the disease.
72. The method of claim 71, further comprising calculating the multiple-ancestry polygenic risk score with additional variables for age, personal medical history, and family medical history.
73. The method of claim 71, wherein the disease is cancer.
74. The method of claim 73, wherein the therapy is a cancer therapy selected from one or more of surgery, cryoablation, radiation therapy, bone marrow transplant, chemotherapy, immunotherapy, hormone therapy, stem cell therapy, drug therapy, biological therapy, and administration of a pharmaceutical, prophylactic or therapeutic compound.
75. The method of claim 71, wherein the disease is breast cancer.
76. The method of claim 75, wherein the therapy is a breast cancer therapy.
77. A method for diagnosing or prognosing a subject having a disease, the method comprising: selecting a plurality of disease-associated SNP markers and a plurality of ancestry- informative SNP markers; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the disease in the subject based on the plurality of the disease-associated SNP markers and ancestry-informative SNP markers, wherein the score indicates a diagnosis or prognosis for the subject.
78. The method of claim 77, wherein the disease is cancer.
79. A method for generating data for assessing a trait in a subject, the method comprising: selecting a plurality of disease-associated SNP markers and a plurality of ancestry- informative SNP markers; measuring a genotype of the subject; and measuring trait-associated SNP markers in the genotype of the subject.
80. The method of claim 79, further comprising determining additional clinical variables of the subject.
81. The method of claim 80, wherein the additional clinical variables are age, personal medical history, and family medical history of the subject.
82. The method of claim 79, wherein the trait is a risk of a disease in the subject.
83. The method of claim 82, wherein the disease is cancer.
84. The method of claim 79, wherein the plurality of ancestry-informative SNP markers are from 10 to 50,000 SNP markers.
85. The method of claim 79, wherein the plurality of ancestry-informative SNP markers are from 10 to 56 SNP markers.
86. The method of claim 83, wherein the trait-associated SNP markers are a plurality of cancer-associated SNP markers.
87. The method of claim 83, wherein the trait-associated SNP markers are a plurality of from 10 to 50,000 breast cancer associated SNP markers.
88. The method of claim 83, wherein the trait-associated SNP markers are a plurality of from 10 to 329 breast cancer associated SNP markers.
89. A system for assessing risk of a disease in a subject, the system comprising: a processor for receiving a genotype of the subject; one or more processors for carrying out the steps:calculating a multiple-ancestry polygenic risk score for risk of the disease in the subject based on a plurality of ancestry-informative SNP markers, a plurality of disease- associated SNP markers of the genotype, and additional variables for age, personal medical history, and family medical history; and a display for displaying and / or reporting the risk score.
90. The system of claim 89, wherein the disease is cancer.
91. A non-transitory machine-readable storage medium having stored therein instructions for execution by a processor which cause the processor to perform the steps of a method for assessing risk of a disease in a subject, the method comprising: receiving a genotype of the subject; calculating a multiple-ancestry polygenic risk score for risk of the disease in the subject based on a plurality of ancestry-informative SNP markers, a plurality of disease- associated SNP markers of the genotype, and additional variables for age, personal medical history, and family medical history; and sending to a processor output for displaying and / or reporting the risk score.
92. The medium of claim 91, wherein the disease is cancer.
93. A method for treating a disease in a subject in need thereof, the method comprising: selecting a plurality of bases in genomic loci associated with trait-associated SNP markers and a plurality of bases in genomic loci associated with ancestry-informative SNP markers; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject based on the bases in genomic loci associated with trait-associated SNP markers and the plurality of bases in genomic loci associated with ancestry-informative SNP markers; and administering to the subject a therapy for the disease.
94. The method of claim 93, further comprising calculating the multiple-ancestry polygenic risk score with additional variables for age, personal medical history, and family medical history.
95. The method of claim 93, wherein the disease is cancer.
96. The method of claim 95, wherein the therapy is a cancer therapy selected from one or more of surgery, cryoablation, radiation therapy, bone marrow transplant, chemotherapy, immunotherapy, hormone therapy, stem cell therapy, drug therapy, biological therapy, and administration of a pharmaceutical, prophylactic or therapeutic compound.
97. The method of claim 95, wherein the disease is breast cancer.
98. The method of claim 97, wherein the therapy is a breast cancer therapy.
99. A method for diagnosing or prognosing a subject having a disease, the method comprising: selecting a plurality of bases in genomic loci associated with trait-associated SNP markers and a plurality of bases in genomic loci associated with ancestry-informative SNP markers; measuring a genotype of the subject; and calculating a multiple-ancestry polygenic risk score for the risk of the trait in the subject based on the bases in genomic loci associated with trait-associated SNP markers and the plurality of bases in genomic loci associated with ancestry -informative SNP markers, wherein the score indicates a diagnosis or prognosis for the subject.
100. The method of claim 99, wherein the disease is cancer.
101. A method for generating data for assessing a trait in a subject, the method comprising: selecting a plurality of bases in genomic loci associated with trait-associated SNP markers and a plurality of bases in genomic loci associated with ancestry-informative SNP markers; measuring a genotype of the subject; and measuring a plurality of bases in genomic loci associated with trait-associated SNP markers in the genotype of the subject.
102. The method of claim 101, further comprising determining additional clinical variables of the subject.
103. The method of claim 102, wherein the additional clinical variables are age, personal medical history, and family medical history of the subject.
104. The method of claim 101, wherein the trait is a risk of a disease in the subject.
105. The method of claim 104, wherein the disease is cancer.
106. The method of claim 101, wherein the plurality of bases within genomic loci associated with ancestry-informative SNP markers are from 10 to 50,000 bases within genomic loci associated SNP markers.
107. The method of claim 101, wherein the plurality of bases within genomic loci associated with ancestry -informative SNP markers are from 10 to 56 bases within genomic loci associated with SNP markers.
108. The method of claim 107, wherein the bases within genomic loci associated with trait- associated SNP markers are a plurality of bases within genomic loci associated with cancer- associated SNP markers.
109. The method of claim 107, wherein the bases within genomic loci associated with trait- associated SNP markers are a plurality of from 10 to 50,000 bases within genomic loci associated with breast cancer-associated SNP markers.
110. The method of claim 107, wherein the bases within genomic loci associated with trait- associated SNP markers are a plurality of from 10 to 329 bases within genomic loci associated with breast cancer-associated SNP markers.
111. A system for assessing risk of a disease in a subject, the system comprising: a processor for receiving a genotype of the subject; one or more processors for carrying out the steps: calculating a multiple-ancestry polygenic risk score for risk of the disease in the subject based on a plurality of bases in genomic loci associated with trait-associated SNP markers, a plurality of bases in genomic loci associated with ancestry-informative SNP markers, and additional variables for age, personal medical history, and family medical history; and a display for displaying and / or reporting the risk score.
112. The system of claim 111, wherein the disease is cancer.
113. A non-transitory machine-readable storage medium having stored therein instructions for execution by a processor which cause the processor to perform the steps of a method for assessing risk of a disease in a subject, the method comprising:receiving a genotype of the subject; calculating a multiple-ancestry polygenic risk score for risk of the disease in the subject based on a plurality of bases in genomic loci associated with trait-associated SNP markers, a plurality of bases in genomic loci associated with ancestry-informative SNP markers, and additional variables for age, personal medical history, and family medical history; and sending to a processor output for displaying and / or reporting the risk score.
114. The medium of claim 113, wherein the disease is cancer.