Prediction of polygenic traits using local ancestors
By calculating polygenic risk scores based on local ancestry, the method addresses the inaccuracy of conventional methods in diverse populations, enhancing the prediction and management of traits like cancer risk in mixed ancestry groups.
Patent Information
- Application Number
- JP2022564182
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-04-20
- Filing Date
- 2021-04-16
- Publication Date
- 2026-05-18
- Estimated Expiration
- 2041-04-16
AI Technical Summary
Conventional methods for determining polygenic risk scores are inaccurate in diverse populations due to reliance on polymorphic loci identified in specific populations, leading to reduced accuracy and predictive errors, particularly in mixed ancestry populations such as African Americans and US Latino Americans.
A method that calculates polygenic risk scores by considering local ancestry through the ancestral origin of alleles at risk loci, using genomic risk loci and additional ancestrally informative markers to adjust the scores based on the local ancestral origin, rather than relying on patient self-reporting or global ancestry composition.
This approach significantly improves the accuracy of polygenic risk score determination and prediction of traits like cancer risk in ancestrally mixed populations, providing reliable clinical risk management and treatment recommendations.
Smart Images

Figure 0007860901000012 
Figure 0007860901000013 
Figure 0007860901000014
Abstract
Description
[Technical Field]
[0001] This invention relates to the fields of genetics and medicine. More specifically, it relates to polygenic traits, methods for predicting risks for pharmaceutical use, and methods for treating diseases for which risks can be assessed. [Background technology]
[0002] It is desirable to use genomic measurements to determine the range or manifestation of various biological traits. Recently, genotyping of biological traits has been incorporated into predicting the risk of clinical conditions. Methods include genotyping polymorphic loci, determining polygenic risk scores, and characterizing predictions of clinical conditions.
[0003] Furthermore, it is desirable to use such polygenomic risk scores for predicting clinical status, regardless of ancestry. However, risk scores derived from genomic data depend on identifying the polymorphic loci to be used. Moreover, polygenic risk scores are specific to the particular population used for measurement and are therefore influenced by ancestry within that population.
[0004] A drawback of conventional methods for generating polygenic risk scores is that polymorphic loci identified in a particular population do not provide accurate polygenic risk scores in different populations.
[0005] For example, polymorphism loci identified in populations of European origin do not provide accurate polygenicity risk scores in populations under study and / or genetically diverse populations, including those with mixed ancestry over generations, such as those with both African and Eurasian ancestry. More specifically, polymorphism loci identified in populations of European origin do not provide accurate polygenicity risk scores for the US African American population and the US Latino American population.
[0006] One method of using mixed ancestral populations to identify polymorphic loci involved censoring certain mixed subjects from the defining population. However, this has the drawback of reducing the accuracy of the available information and scores.
[0007] Another approach is to adjust for the final score of the ancestral composition of the target population. Unfortunately, this has the drawback of relying on assumed genotype characteristics.
[0008] Generally, conventional methods only consider the entire ancestral composition of the subject of study and do not provide an accurate score.
[0009] What is needed is an effective and accurate method for determining polygenic risk scores that improve accuracy and reduce predictive error. A favorable clinical risk algorithm can improve medical care and patient management.
[0010] There is an urgent need for methods to assess clinical conditions, such as the risk of cancer. A method that can deliver these assessments efficiently, to the point where it could be considered a form of medical treatment, is required. [Overview of the Initiative] [Problems that the invention aims to solve]
[0011] This invention provides a method for determining the risks of polygenic traits and their pharmaceutical applications, and for treating diseases in which the risks are identified and / or assessed. [Means for solving the problem]
[0012] In some embodiments, the methods of the present invention can provide excellent prediction of clinical risk in patients with a wide variety of ancestry and / or mixed ancestry. The methods of the present invention can provide polygenic risk prediction that does not rely on patient self-reporting of ancestry. Furthermore, the methods of the present invention can provide polygenic risk prediction that does not rely on so-called "genetic" ancestry composition.
[0013] In certain embodiments, the clinical utility of the present invention includes excellent prediction of clinical risk in African American patients, patients of any mixed ancestry, and patients with non-European ancestry, including Latino and European American ancestry with partial African genetic roots, without relying on patient self-reporting of ancestry or being limited by so-called genetic ancestry composition.
[0014] In some embodiments, the method of the present invention can provide a polygenicity score that takes local ancestry into account. Local ancestry can be considered through the ancestral origin of alleles at risk loci.
[0015] The polygenicity score obtained by the method of the present invention can, surprisingly, lead to increased accuracy in determining the risk of polygenic traits and ancestrally mixed populations.
[0016] Examples of polygenic traits include the likelihood of cancer, such as breast cancer, and other diseases.
[0017] In a further embodiment, determining polygenic traits and risks may involve identifying and utilizing genomic risk loci. Genomic risk loci may be related to traits, if not indirectly related to genomic effects and traits.
[0018] In some embodiments of the present invention, the genomic risk loci that indirectly affect the trait may be associated with the trait only in a specific or local group of ancestors.
[0019] Surprisingly, the method of the present invention can lead to accurate determination of polygenic traits and risks by evaluating and incorporating the contribution of local ancestral groups.
[0020] Embodiments of the present invention contemplate determining the levels of polygenic traits and risks in the form of scores based on various genomic risk loci. Since the genomic risk loci can be separately identified and defined, an accurate determination can be made by genotyping the subject.
[0021] In certain embodiments, a genomic risk locus may include a genomic risk marker for a particular trait, which is combined with additional risk markers that may be ancestrally informative. The ancestrally informative marker may be a particular adjacent genomic risk marker and / or a flanking genomic risk marker. The ancestrally informative marker can provide information regarding the contribution of a group of local ancestries.
[0022] In certain embodiments, a score for a polygenic trait and / or risk may include determining the weights of additional ancestrally informative risk markers that will be combined with genomic risk markers.
[0023] Embodiments of the present invention include:
[0024] A method for evaluating a biological trait in a subject, comprising: Measuring the genotype in a sample derived from the subject, wherein the genotype has a window centered on a trait risk marker for the trait, and the window includes additional ancestrally informative markers flanking the risk marker; Phasing the genotype to determine the haplotypes within each window using a reference population having admixed ancestry; Calculating the odds of local ancestral origin for each window; and Calculating a polygenic risk score for the biological trait in the subject using the trait risk marker and the odds of local ancestral origin for each window, wherein the score is adjusted according to the local ancestral origin of the window. A method comprising.
[0025] The number of trait risk markers may range from 1 to 10,000. The process of calculating local ancestral odds may include dividing the stepped genotypes within each window into a sequence of non-overlapping tiles of up to approximately 300 additional ancestral knowledge markers, and calculating the ancestral odds for each haplotype in each tile using the empirical frequency of the haplotype in the reference population. Each tile may contain 1 to 100 additional ancestral knowledge markers. Each tile may contain 5 to 20 additional ancestral knowledge markers. The window may be approximately 1 MB wide. Genotypes can be determined by NGS. Genotypes can be determined by sequencing chips. Biological traits may be cancer likelihoods. Genomic risk markers may be cancer markers. Additional ancestral knowledge markers may be SNP markers or indel markers. Genomic risk markers may be breast cancer SNP markers. The process of calculating the polygenicity risk score may include calculating the incremental contribution of each allele to the polygenicity risk score as a local ancestor-specific risk effect beta obtained by multiplying several genotyped risk alleles, which is zero or one and less than the population-specific risk allele frequency.
[0026] The present invention further includes a method for recommending treatment for a subject having a disease, the method being: A step of measuring the genotype in a sample derived from a subject, wherein the genotype includes risk markers associated with the trait, and further includes additional ancestral knowledge markers that frank the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; The process of calculating the odds of local ancestral origin for each window; A step of calculating a polygeneity risk score for a biological trait in a subject using trait risk markers and the odds of local ancestral origin for each window, wherein the score is adjusted according to the local ancestral origin of the window; and A process of recommending treatment for a disease based on a risk score that exceeds a threshold level. This includes: The disease may be cancer or breast cancer. Treatment may be one of the following: treatment for the disease; treatment for the disease followed by a monitoring period; or tapering off of treatment for the disease. Treatment may be one or more of the following: surgery, cryoablation, radiotherapy, bone marrow transplantation, chemotherapy, immunotherapy, hormone therapy, stem cell therapy, drug therapy, biological therapy, and administration of pharmaceuticals, prophylactic or therapeutic compounds.
[0027] The present invention further includes a method for identifying subjects having a disease that would benefit from treatment, the method being: A step of measuring the genotype in a sample derived from a subject, wherein the genotype includes risk markers associated with the trait, and further includes additional ancestral knowledge markers that frank the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; The process of calculating the odds of local ancestral origin for each window; A step of calculating a polygeneity risk score for a biological trait in a subject using trait risk markers and the odds of local ancestral origin for each window, wherein the score is adjusted according to the local ancestral origin of the window; and A process for identifying individuals with a disease who would benefit from treatment, based on a risk score indicating the need for treatment or exceeding a threshold level. This includes the disease, which may be cancer or breast cancer.
[0028] Embodiments of the present invention further envision a method for treating a disease in a subject requiring treatment of the disease, the method being: A step of measuring the genotype in a sample derived from a subject, wherein the genotype includes risk markers associated with the trait, and further includes additional ancestral knowledge markers that frank the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; The process of calculating the odds of local ancestral origin for each window; A step of calculating a polygeneity risk score for a biological trait in a subject using trait risk markers and the odds of local ancestral origin for each window, wherein the score is adjusted according to the local ancestral origin of the window; and This includes: treatment for the disease; treatment for the disease following a monitoring period; or tapering off of treatment for the disease.
[0029] Additional embodiments include a method for monitoring the response of a subject having a disease, the method being: A step of measuring the genotype in a sample derived from a subject, wherein the genotype includes risk markers associated with the trait, and further includes additional ancestral knowledge markers that frank the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; The process of calculating the odds of local ancestral origin for each window; A process for calculating a polygeneity risk score for a biological trait in a subject using trait risk markers and the odds of local ancestral origin for each window, wherein the score is adjusted according to the local ancestral origin of the window. Includes.
[0030] The present invention includes a method for predicting the prognosis of a subject having a disease, the method being: A step of measuring the genotype in a sample derived from a subject, wherein the genotype includes risk markers associated with the trait, and further includes additional ancestral knowledge markers that frank the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; The process of calculating the odds of local ancestral origin for each window; A step of calculating a polygeneity risk score for a biological trait in a subject using trait risk markers and the odds of local ancestral origin for each window, wherein the score is adjusted according to the local ancestral origin of the window; and A process for predicting the prognosis of individuals with a poor prognosis for a disease, based on a risk score that indicates the need for treatment or exceeds a threshold level. Includes.
[0031] Further embodiments include a system for assessing the risk of disease in a subject, the system being: A processor for receiving genomic data from a target sample; A step of measuring the genotype in a sample derived from a subject, wherein the genotype has a window centered on trait risk markers for the trait, and the window includes additional ancestral knowledge markers that frank to the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; The process of calculating the odds of local ancestral origin for each window; and A process for calculating polygeneity risk scores for biological traits in a subject, using trait risk markers and local ancestral origin odds for each window. One or more processors for executing; and Display for presenting and / or reporting risk scores Includes.
[0032] Also included is a non-temporary machine-readable storage medium in which instructions for execution by the processor are stored, causing the processor to perform steps of a method for assessing the risk of disease in a subject, and the method is: The process of obtaining genomic data from a sample derived from the target; A step of measuring the genotype in a sample, wherein the genotype has a window centered on trait risk markers for a trait, and the window includes additional ancestral knowledge markers that frank to the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; The process of calculating the odds of local ancestral origin for each window; A process for calculating a polygenetic risk score for a biological trait in a subject using trait risk markers and the odds of local ancestral origin for each window; and The process of sending an output to the processor that presents and / or reports a risk score. Includes. [Brief explanation of the drawing]
[0033] [Figure 1] This figure shows the tiling of subcentimorgan regions into short local haplotypes. Haplotype frequencies can be used to accurately separate local African versus European ancestry from publicly available datasets. Nine windows for each of the 12 markers are shown. Likelihood of African local ancestry in African Americans (AfrAm). [Figure 2] This figure shows that some Europeans in the IBS (Iberian / Spanish) dataset (G1K) exhibit up to 2-3 African haplotypes per locus in local ancestry mapping. Some minor allele carriers in G1K appear to have a very African-like linkage disequilibrium (LD) pattern. Eur (dotted line), Afr (solid line). [Figure 3] This figure shows that the local / global ancestral distribution and AF are as expected. Predicted distribution of racial origin for rs132390. Empirical AfrAm (dotted line), G1K Afr (solid line), G1K Eur (dashed line). [Figure 4] This figure shows that the overall ancestral average for Europeans was 17%, which, as expected, was widely distributed. [Figure 5] This figure shows empirical MAF (y-axis) derived from local ancestral segments of Europeans versus G1K MAF values (x-axis). [Figure 6] This figure shows empirical MAF (y-axis) derived from local ancestral segments of Africans versus G1K MAF values (x-axis). [Figure 7] This figure shows that several markers may be required for local ancestor assignment of Eur / Afr. Surprisingly, a high accuracy of 90% to 95% can be achieved for local Afr / Eur ancestor identification using a relatively small number of franking markers (x axis). 95% accuracy (dashed line), 90% accuracy (solid line). [Figure 8] This figure shows that several markers may be required for local ancestry assignment of Eur / EAS. Surprisingly, a high accuracy of 90%–95% can be achieved in separating Europeans from East Asian ancestry using a relatively small number of franking markers (x-axis). 95% accuracy (dashed line), 90% accuracy (solid line). [Figure 9] This figure shows the error rates of local ancestral assignment based on unstacked genotypes. Figure 9 shows the results for three SNP loci, rs258809, rs10759243, and rs1550623, with the y-axis representing the error rate of local ancestral deconvolution as a function of the number of flanking loci. [Modes for carrying out the invention]
[0034] This invention includes a method for using local ancestry to predict polygenetic risk, thereby providing an accurate risk assessment for all subjects regardless of ancestry.
[0035] Embodiments of the present invention further provide reliable trait associations in ancestrally diverse populations and genetically mixed populations.
[0036] This disclosure provides various methods for related studies of multiple candidate gene loci that may be characterized by broader linkage disequilibrium (LD) patterns. These studies have little to no direct influence on traits.
[0037] Embodiments of the present invention can provide clinical risk management, risk magnitude assessment, and polygenic risk scoring and preclinical trait prediction. The methods of the present invention can provide accurate predictive capabilities for all subjects, and even for mixed genotypes.
[0038] Aspects of this disclosure include genotyping polymorphic loci and combining the genotypes in the form of polygenicity scores to predict the risk of a clinical condition or the range of signs of a biological trait.
[0039] In some embodiments, the prediction of polygeneity score may be specific to the population and ancestors in which it was found.
[0040] In a further embodiment, multiple trait risk markers can be used together with additional ancestral knowledge markers to provide predictions of polygenetic risk for a trait.
[0041] In a particular embodiment, the multiple trait risk markers may be 1 to 10,000 markers, 1 to 1,000 markers, or 1 to 100 markers.
[0042] Some risk markers may appear within a genotype window. This window can be approximately ±0.5 MB, which is a 1 MB window, or it may be ±1 MB or ±2 MB. In certain embodiments, the window can be any size useful for stratifying genotypes.
[0043] The present invention can provide risk prediction portability for research groups and genetically diverse groups. In certain embodiments, the present invention can provide risk prediction portability for ancestrally mixed populations, such as populations with both African and Eurasian ancestry, including African Americans and Latino Americans in the United States.
[0044] In a further embodiment, a portable method for predicting trait associations between branched populations is provided herein. Embodiments herein can result in improved prediction of polygeneity risk, even in the presence of the indirect nature of allelic influences on traits. In certain embodiments, various directly influencing loci may remain undiscovered but can be linked to loci discovered by linkage disequilibrium. Various directly influencing loci may be separated by a genetic distance of centimorgan units.
[0045] Aspects of the present invention may provide a unique method for identifying each allele of a genotype by its ancestry.
[0046] Further aspects of the present invention may provide a unique method for determining a polygenic risk score for predicting biological traits.
[0047] In some embodiments, larger-scale association studies can be used, which have a greater ability to decompose linkage disequilibrium (LD) patterns and concentrate on fewer loci. In some embodiments, the best single candidate SNP may not be a truly directly affecting allele.
[0048] In a further embodiment, LD patterns between candidate SNPs and the truly directly affecting alleles may be read within different ancestral populations.
[0049] In certain embodiments, one or more risk SNPs can be found within European populations. These risk SNPs may retain their significance within East Asian populations, and their odds ratios may be reduced. Some European SNPs can be used in African populations.
[0050] In some embodiments, the genetic distance resolution of a GWAS can be within the range of 0.1 cM, meaning that the LD pattern between candidate SNPs and true risk loci is expected to fade out over several thousand generations (20,000–60,000 years). For example, Eurasian populations may share a substantially common ancestor with each other, rather than with Africans, within this timeframe.
[0051] In further embodiments, both local and global genetic ancestors can be used, where the alleles are similar for different indirect effects. Alleles identified by GWAS studies are rarely unique to a single ancestral population. However, both odds ratios and mean score values can vary widely among ancestors. The gap between cultural or self-reported ancestors and genetic ancestors can be quite large. Genetic ancestors can accurately inform us of the magnitude of SNP trait effects.
[0052] In a particular embodiment, if the population is historically recently mixed with characteristic ancestral chromosome blocks >>0.1 cM, only the local ancestry of the DNA segment can inform the calculation of the effect.
[0053] Regarding the local ancestry of African Americans, the dominant source group is West Africans, accounting for over 80% on average, with wide variation, and Europeans may account for less than 20%. The average European segment size may be around 30 cM, and many segments may be shorter, possibly due to the inclusion of studies with lower marker density.
[0054] This invention aims to calculate local ancestors in a segmented genome.
[0055] In some embodiments, local haplotypes may be ancestral-knowledgeable.
[0056] In a further embodiment, the present invention results in tiling of subcentimorgan regions into short local haplotypes. In a particular embodiment, haplotype frequencies can be used to precisely separate local African ancestors from European ancestors from a reference dataset.
[0057] Embodiments of the present invention can reveal mixing that has occurred over thousands of years. Ancestral reference populations may themselves have experienced limited mixing (≧0.1 cM) within a historical timeframe. Low levels of African mixing may be observed within the Iberian Peninsula, and ancient DNA studies can date back to the Roman era, and in part to the Caliphate era. The IBS (Iberian / Spanish) reference dataset could be part of the genomes of 1000 Europeans and could provide up to 2-3 African haplotypes per locus in local ancestral mapping.
[0058] In some embodiments, a high-density gene array can be used. The array is 2 × 10⁻¹⁶. 5 ~5×10 6 It may contain individual SNPs.
[0059] In certain embodiments, a high-density AXIOM array can be used.
[0060] In a further embodiment, breast cancer risk markers can be used in conjunction with additional flanking SNPs within a window. The window may be approximately ±0.5 MB, which is a 1 MB window.
[0061] Some examples of breast cancer risk markers are given in "Prediction of breast cancer risk based on profiling with common genetic variants," Mavaddat et al., J Natl Cancer Inst., April 8, 2015, Vol. 107 (No. 5).
[0062] Some examples of breast cancer risk markers are given in "Characterizing Genetic Susceptibility to Breast Cancer in Women of African Ancestry," Feng et al., Cancer Epidemiol Biomarkers Prev., July 2017, Vol. 26 (No. 7), pp. 1016-1026.
[0063] Some examples of breast cancer risk markers are given in Early Diagnosis of Breast Cancer, Wang et al., Sensors (Basel), July 2017, Vol. 17 (No. 7), p. 1572.
[0064] In a specific manner, genotypes can be stratified using Beagle 5.1 with reference datasets from the African and European 1000 Genomes Project. Each stratified genotype can be divided into non-overlapping tiles of approximately 15 markers. The odds of African versus European origin can be calculated for each tile.
[0065] In a further embodiment, polygenicity risk scores adjusted for overall ancestry may, surprisingly, be more accurate than polygenicity risk scores adjusted for mean ancestry.
[0066] In some embodiments, polygenicity risk scores adjusted for local ancestry may, surprisingly, be more accurate than polygenicity risk scores adjusted for global ancestry.
[0067] In a further embodiment, the association between polygenic risk scores and breast cancer may be assessed by logistic regression. The logistic regression may be adjusted for age and family history among the variables. The logistic regression may include parameters or data for age and family history among the variables.
[0068] In an additional embodiment, flanking markers can be sequentially added from a stepwise set of 1000 genomic genotypes to maximize local ancestral segregation in each iteration.
[0069] In certain embodiments, it is advantageous to achieve a local African / European ancestry identification accuracy of 90–95% using a small number of flanking markers.
[0070] In a further embodiment, the challenging task of separating Europeans from East Asian ancestry may require fewer than 20 additional flanking SNP markers.
[0071] Polygenetic risk score One aspect of the present invention aims to provide a method for calculating a polygenicity score that takes into account the ancestral origin of the alleles of a risk locus, i.e., local ancestry.
[0072] The method of the present invention can use any number of trait risk markers. The method of determining a locally ancestral corrected polygenicity risk score of the present invention may use any three or more markers from the set of markers shown in Table 2, or any five or more markers, or any ten or more markers.
[0073] One aspect of the present invention aims to provide a method for calculating polygenicity scores by adjusting for the ancestral origin of the alleles of a risk locus, i.e., local ancestry.
[0074] In some embodiments, local ancestry adjustments can dramatically improve the portability of traits, such as the polygenic score predicting breast cancer, in ancestrally mixed populations.
[0075] In further embodiments, the indirectly influencing risk locus may be associated with a trait present in only one of the ancestral groups. In other embodiments, the risk allele may have a direct influence, with similar sizes across the entire spectrum of local ancestors.
[0076] Embodiments of the present invention provide a method for accurately measuring the contribution of alleles to local ancestral dependence on risk prediction. The method of the present invention can yield risk predictions with dramatically increased accuracy compared to conventional adjustments, based on the overall composition of the ancestor under study.
[0077] The methods disclosed herein may include the step of identifying the ancestral origin of all genotypes that are clinically relevant or related to a trait.
[0078] In a particular embodiment, the genotypes of an additional number of ancestral knowledgeable loci adjacent to the scoreable locus can be used.
[0079] In a further embodiment, specific weights can be applied to scoreable loci depending on ancestral origin, i.e., local ancestor.
[0080] The steps of the present invention may include genotyping of an unspecified African American subject. Genotyping can be performed by any method, including NDS, custom chips, or a combination of NGS and chips.
[0081] In some embodiments, a subset of known breast cancer risk markers can be used with additional markers within a window centered on each known risk marker. The subset may have higher allele frequencies in published African and European controls, and lower linkage disequilibrium among the additional markers.
[0082] Examples of additional markers include SNPs and indels.
[0083] In certain embodiments, the genotype may be a stratified haplotype. Genotype stratification can be performed by known methods.
[0084] Examples of methods for haplotype estimation include Hidden Markov models (HMMs), PHASE, Gibbs sampling, fastPHASE, BEAGLE, haplotype cluster modeling, IMPUTE2, MaCH, SHAPEIT1, HAPI-UR, and SHAPEIT2.
[0085] In additional embodiments, genotyping may be performed by mismatch based on similarity coefficients.
[0086] In a further step, the haplotype estimation method may use reference datasets for Africans and Europeans.
[0087] An example of a reference dataset is the 1000 Genomes Project. 5 Examples include publicly available and / or personal collections of more than one genome.
[0088] In an additional embodiment, the stepped genotype window may be divided into consecutive non-overlapping tiles. Each tile may contain multiple markers. In a particular embodiment, a tile may contain up to about 300 markers, or 1 to 100 markers, or 2 to 50 markers, or 5 to 40 markers, or 5 to 20 markers.
[0089] An additional step of the method of the present invention may include calculating the odds of African versus European origin for each haplotype in each tile. In a particular step, calculating the odds of African versus European origin for each haplotype may be done using the empirical frequency of the haplotype in a baseline set. For haplotypes absent in the baseline set, an observed frequency of 0.5 per set may be used. Haplotypes absent in all baseline sets can be equally considered likely to be either African or European.
[0090] In some embodiments, the odds for a locus that is African versus European in ancestral origin can be found according to formula I.
[0091]
number
[0092] In a further embodiment, the overall odds of African versus European ancestry for the entire segment containing each allele of a risk SNP can be calculated as the product of the odds of all tile haplotypes contained within the segment. The partial probability of each allele, representing a partial local ancestry of "ancestral European" versus "ancestral African," can be assigned according to Equation II.
[0093]
number
[0094] In an additional embodiment, the entire fraction of the European ancestry of each study subject, i.e., the overall ancestral GA, can be calculated as the average of the partial European local ancestry for all loci, according to Equation III.
[0095]
number
[0096] A further step of the method of the present invention may include calculating the incremental contribution of each allele to the polygenic risk score as the local ancestry-specific risk effect, beta, multiplied by the number of genotyped risk alleles, where GT SNP is zero or 1 and may be less than the population-specific risk allele frequency.
[0097] In some embodiments, the ancestry-specific beta of local ancestry for Europeans may be known. For risk markers that could not show significant and invariant effects in African studies, the African ancestry-specific beta may be assumed to be equal to zero. Also, the ancestry-specific beta of local ancestry for Africans may be known or may be assumed to be zero. In certain embodiments, for risk markers that have a somewhat significant effect in African population studies, the African ancestry-specific beta may be assumed to be the scaled-down beta from European studies.
[0098] In some aspects, the ancestral group β A,SNP and β E,SNP for each beta can be interpolated using equal weights for each overall ancestral mean, GAA for European ancestry, and (1 - GAA) for African ancestry, based on the assumption that a conventional comparative polygenic risk score (cPRS) can be calculated. The conventional score may be centered on the mixed population-specific minor allele frequency MAF AA,SNP .
[0099]
Number
[0100]
Number
[0101] In a particular embodiment, the comparison method or conventional method may include global ancestor adjustment, which may interpolate the ancestor beta using the actual global ancestor of each object according to formula VI.
[0102]
number
[0103] In a further embodiment, embodiments of the present invention can provide a method for adjusting local ancestry, which uses the partial likelihood LA of each specific allele of a European person according to formula VII, and the center of both components of the ancestral score, and the frequency MAF of each minor allele of the ancestral person. A,SNP and MAF E,SNP In this case, the ancestor beta can be interpolated.
[0104]
number
[0105] In a particular embodiment, the mean African American population-specific beta may be calculated by interpolation between the local ancestry-specific betas of Europeans and Africans, using the mean percentage of European versus African ancestry. The globally ancestral-adjusted beta can be calculated for each specific subject by interpolation between the local ancestry-specific betas of Europeans and Africans, using the actual percentage of European versus African ancestry.
[0106] The method of the present invention can use any number of trait risk markers. The method of determining a locally ancestral corrected polygenicity risk score of the present invention may use three or more markers from any set of markers shown in Table 2, or five or more markers from any set, or ten or more markers from any set.
[0107] Cancer score and treatment Cancer treatments may include surgery, cryoablation, radiotherapy, bone marrow transplantation, chemotherapy, immunotherapy, hormone therapy, stem cell therapy, drug therapy, biological therapy, and administration of pharmaceuticals, prophylactic or therapeutic compounds, such as biological or exogenous activators.
[0108] Examples of treatments include surgical intervention for obese individuals, physical therapy, diet, and dietary supplementation.
[0109] Examples of cancer biological therapies include adoptive cell transfer, angiogenesis inhibitors, Calmette-Guérin bacillus therapy, biochemotherapy, cancer vaccines, chimeric antigen receptor (CAR) T-cell therapy, cytokine therapy, gene therapy, immune checkpoint regulators, immune complexes, monoclonal antibodies, oncolytic virus therapy, and targeted drug therapy.
[0110] Examples of cancer surgery include breast tumor removal, partial mastectomy, total mastectomy, simple mastectomy, modified standard mastectomy, standard mastectomy, and Holstead standard mastectomy.
[0111] Examples of cancer drugs include those approved for the prevention of breast cancer, such as Evista (raloxifene hydrochloride), raloxifene hydrochloride, and tamoxifen citrate.
[0112] Examples of cancer drugs include those approved for treating breast cancer, such as abemaciclib, Abraxane (paclitaxel albumin-stabilized nanoparticle formulation), Ado-trastuzumab emtansine, Afinitor (everolimus), Afinitor Disperz (everolimus), alpelicib, anastrozole, Aredia (pamidronate disodium), Arimidex (anastrozole), Aromasin (exemestane), atezolizumab, capecitabine, cyclophosphamide, docetaxel, doxorubicin hydrochloride, Ellence (epirubicin hydrochloride), Enhertu (Fam-Trastuzumab Deruxtecan-nxki), epirubicin hydrochloride, eribulin mesylate, everolimus, exemestane, 5-FU (fluorouracil injection), and Fam-Trastuzumab Deruxtecan-nxki, Fareston (toremifene), Faslodex (fulvestrant), Femara (letrozole), fluorouracil injection, fulvestrant, gemcitabine hydrochloride, Gemzar (gemcitabine hydrochloride), goserelin acetate, Halaven (eribulin mesylate), HerceptinHylecta (trastuzumab and hyaluronidase-oysk), Herceptin (trastuzumab), Ibrance (palbociclib), Ixabepirone, Izempra (ixabepirone), Kadcyla (Ado-trastuzumab emtansine), Kisqali (ribociclib), lapatinib nitosylate, letrozole, Lynparza (olaparib), megestrol acetate, methotrexate, neratinib maleate, Nerlynx (neratinib maleate), olaparib, paclitaxel, paclitaxel-albumin-stabilized nanoparticle formulation, palbociclib, pamidronate nina Examples include Thorium, Perjeta (pertuzumab), pertuzumab, Piqray (alperisib), ribociclib, thalazoparib tosylate, Talzenna (thalazoparib tosylate), tamoxifen citrate, Taxotere (docetaxel), Tecentriq (atezolizumab), thiotepa, toremifene, trastuzumab, trastuzumab and hyaluronidase oysk, Trexall (methotrexate), Tykerb (lapatinib nitosylate), Verzenio (abemaciclib), vinblastine sulfate, Xeloda (capecitabine), and Zoladex (goserelin acetate).
[0113] As used herein, the term “disease” includes any disorder, condition, or illness, such as a chronic illness affecting an organ, part, structure, or system of the body that is not functioning properly or in a disordered manner.
[0114] As used herein, the term "sample" includes any biological sample isolated from a subject. Samples may include, but are not limited to, single or multiple cells, cell fragments, aliquots of body fluids, whole blood, platelets, serum, plasma, red blood cells, white blood cells (or leucocytes), endothelial cells, tissue biopsies, synovial fluid, lymph, ascites, and interstitial or extracellular fluid. The term "sample" also encompasses fluids within the intercellular spaces, including synovial fluid, gingival crevicular exudate, bone marrow, cerebrospinal fluid (CSF), saliva, mucus, sputum, semen, sweat, urine, or any other body fluid. Blood samples may include whole blood or any fraction thereof, and include blood cells, red blood cells, white blood cells (or leucocytes), platelets, serum, and plasma.
[0115] As used herein, the term "subject" includes humans and mammals.
[0116] In some embodiments, the present invention can provide a method for recommending a treatment regime, including withdrawal from the treatment regime.
[0117] In further embodiments, odds ratios can provide clinicians with a picture of the prognosis of a subject's biological condition. Such embodiments can provide subject-specific prognostic information that may be informative of treatment decisions and may also facilitate monitoring of treatment response. Surprisingly, such embodiments may result in an increased proportion of subjects achieving improved treatment, such as better management of the disease or remission of symptoms.
[0118] As used herein, the terms “biological,” “biotherapy,” and / or “biopharmaceutical” may include medicinal therapeutic products manufactured or extracted from biological substances. Biological products may include vaccines, blood or blood components, allergics, somatic cells, gene therapies, tissues, recombinant proteins, and living cells; they may consist of sugars, proteins, nucleic acids, living cells or tissues, or combinations thereof.
[0119] As used herein, the terms “treatment regimen,” “treatment,” and / or “procedure” may include any clinical management and intervention of a subject, whether biological, chemical, physical, or a combination thereof, that is intended to sustain, alleviate, improve, or otherwise alter the condition of that subject.
[0120] As used herein, the term “administer” may include the placement of a composition into a subject by a method or route that results in at least partial localization of the composition at a desired site, so as to produce the desired effect. Routes of administration include both topical and systemic administration. Typically, topical administration delivers more of the composition to a specific location compared to systemic administration, while systemic administration delivers it essentially to the entire body of the subject. “Administer” may also include performing physical actions on the body of the subject, such as physiotherapy, as well as chiropractic treatment, massage, and acupuncture.
[0121] Devices and Systems As used herein, the term "machine-readable storage medium" may include, for example, data storage material encoded by machine-readable data or data arrays. The data and machine-readable storage medium can be used for a variety of purposes when a machine programmed with instructions for using the data is used. Such purposes include storing, accessing, and manipulating information relating to the risks of a subject or population over time, or risks in response to treatment, or information for drug discovery of inflammatory diseases. The data, including genomic measurements, may be executed by a computer program running on a programmable computer that may include a processor, a data storage system, one or more input devices, and one or more output devices. Program code can be input to the input data to perform the functions described herein and generate output information. The output information can then be input to one or more output devices. The computer may be, for example, a personal computer, a microcomputer, or a workstation.
[0122] As used herein, a computer program can be a code of instructions executed in a high-level procedural language or an object-oriented programming language to communicate with a computer system. A program may be executed in a machine language or an assembly language. The programming language may be a compiled language or an interpreted language. Each computer program may be stored in a storage medium or device, such as a ROM or magnetic diskette, and may be readable by a programmable computer, in order to configure and operate a computer, such that the storage medium is read by the computer to perform the procedures described. A health-related or genome data management system can be thought of as running as a computer-readable storage medium composed of computer programs, where the storage medium causes the computer to operate in a particular manner, such as performing various functions.
[0123] conclusion All publications, patents, and documents specifically described herein are incorporated herein by reference in their entirety for all purposes.
[0124] The terms specifically defined herein are provided as a whole in the context of this disclosure and have meanings that will be typically understood by those skilled in the art. The singular forms “a,” “an,” and “the” as used herein include the plural forms.
[0125] While this disclosure is described in conjunction with various embodiments, it is not intended to be limited to such embodiments. Conversely, this disclosure includes various substitutes, modifications, and equivalents that will be recognized by those skilled in the art.
[0126] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art relating to this invention. Methods and materials similar to or equivalent to those described herein may be used in carrying out or testing the invention, but suitable methods and materials are listed below. In addition, the materials, methods, and examples herein are illustrative and not intended to be limiting.
[0127] While the foregoing disclosure has been described in some detail with explanations and examples to clarify understanding, it will be readily apparent to those skilled in the art that various changes and modifications may be made within the scope of the present invention and the appended claims.
[0128] [Examples] [Example 1] Genotype window for markers Polygenic risk scores were calculated by adjusting for the local ancestral origin of the alleles at the risk locus. To identify the ancestral origin of all clinically relevant or trait-related genotypes, genotypes from multiple additional ancestral-knowledgeable loci adjacent to the scoreable loci were used.
[0129] Genotypes were stratified using Beagle 5.1, with the 1000 Genomes Project's African and European reference datasets serving as the baseline. Each 1MB window of the stratified genotypes was divided into 15 consecutive, non-overlapping tiles of SNP markers.
[0130] Figure 1 shows the tiling of subcentimorgan regions into short local haplotypes. Local African versus European ancestors were accurately separated using haplotype frequencies from publicly available datasets. Figure 1 shows the use of 9 windows for 12 markers. Likelihood of African local ancestors in AfrAm.
[0131] [Example 2] Local Ancestry and Comparative Calculation Figure 2 shows that some Europeans in the IBS (Iberian / Spanish) dataset (G1K) exhibited up to 2-3 African haplotypes per locus in local ancestral mapping. Some minor allele carriers in G1K Europeans had a highly African-like LD pattern. Eur (dotted line), Afr (solid line).
[0132] Figure 3 shows a comparison of local and global ancestral distributions, where allele frequencies (AF) were as expected. Predicted distribution of racial origin for rs132390. Empirical AfrAm (dotted line), G1K Afr (solid line), G1K Eur (dashed line).
[0133] Figure 4 shows that the overall ancestral average for Europeans was 17%, which, as expected, was widely distributed.
[0134] Figure 5 shows empirical MAF (y-axis) versus G1K MAF values (x-axis) derived from local ancestral segments of Europeans. Figure 5 demonstrates that empirical ancestral allele frequencies (AF) derived from local ancestry estimate risk markers that closely match publicly available estimates of allele frequencies in Europeans and Africans. Thus, local ancestry calculation estimates the ancestral composition of chromosomal segments with remarkable accuracy and then assigns local ancestry to each risk allele.
[0135] Figure 6 shows empirical MAF (y-axis) versus G1K MAF values (x-axis) derived from local ancestral segments of Africans.
[0136] [Example 3] Calculation of local ancestor odds The odds of African versus European ancestral origin for each haplotype in each tile were calculated using Equation I, with the empirical frequency of the haplotype in the baseline set. Haplotypes absent in one of the baseline sets were assigned an observation frequency of 0.5 per set. Haplotypes absent in all of the baseline sets were equally assumed to be likely African versus European.
[0137] The overall odds of African versus European ancestry for the entire segment containing each allele of a risk SNP were calculated as the product of the odds of all tile haplotypes contained within the segment. The partial probability of each allele being "ancestrally European" versus "ancestrally African," or a partial local ancestry LA, was determined according to Equation II.
[0138] The entire ancestral fraction of each European subject, or the overall ancestral GA, was calculated as the average of the local ancestry of partial Europeans for all loci, according to Equation III.
[0139] The incremental contribution of each allele to the polygenic risk score was calculated as a local ancestor-specific risk effect (beta), obtained by multiplying it by the number of risk alleles determined by genotype (less than population-specific risk allele frequency, GT). SNP (= zero or 1).
[0140] Local ancestry-specific beta for the local ancestry of Europeans was obtained from Mavaddat (J Natl Cancer Inst., April 8, 2015, Vol. 107(No. 5)). Local ancestry-specific beta for the local ancestry of Africans was obtained either directly from Feng (Cancer Epidemiol Biomarkers Prev., July 2017, Vol. 26(No. 7), pp. 1016-1026) or estimated from Feng and Wang (Sensors (Basel), July 2017, Vol. 17(No. 7), p. 1572). Whenever the estimations from Feng and Wang for nominally significant breast cancer risk associations were consistent with Mavaddat's risk estimations, it was assumed that the latter were similarly appropriate for local and whole-ancestry risk markers in Africans. For Mavaddat risk markers that did not show significant or invariant effects in African studies, we assumed an African ancestry-specific beta of zero. For risk markers that sometimes show a slightly significant effect in African population studies, we assumed an African ancestry-specific beta of a scaled-down beta from European studies.
[0141] [Example 4] Polygenetic risk score Ancestral group β A,SNP and β E,SNP The conventional comparative polygenicity risk score (cPRS) was calculated based on the assumption that each beta is interpolated using weights equal to each overall ancestral mean, GAA for European ancestry, and (1-GAA) for African ancestry. The conventional score is calculated according to equations IV and V, using mixed population-specific minor allele frequencies (MAF). AA,SNP The center was there.
[0142] A conventional comparison method for global ancestor adjustment, which allows interpolation of the ancestor's beta using the actual global ancestors of each target, was calculated according to Equation VI.
[0143] Local ancestral adjustments were calculated, which can interpolate the ancestral beta using the partial likelihood LA of each specific allele that is European. Both components of the ancestral score are the frequency MAF of each minor ancestral allele according to Equation VII. A,SNP and MAF E,SNP The center was there.
[0144]
number
[0145] The mean African American population-specific beta was calculated by interpolation between the local ancestor-specific betas of Europeans and Africans, using the average percentage of European versus African ancestry. For each specific subject, the overall ancestry-adjusted beta was calculated by interpolation between the local ancestor-specific betas of Europeans and Africans, using the actual percentage of European versus African ancestry.
[0146] The method of the present invention can use any number of trait risk markers. The method of determining a locally ancestral corrected polygenicity risk score of the present invention may use three or more markers from any set of markers shown in Table 2, or five or more markers from any set, or ten or more markers from any set.
[0147] [Example 5] Accuracy of local ancestry-adjusted polygenic risk scores for breast cancer status Surprisingly, the local ancestral-adjusted polygenic risk score (PRS) calculated according to this invention was more accurate than the conventional PRS adjusted for "global" ancestry and the conventional average PRS adjusted for ancestry.
[0148] The association between polygenic risk scores and phenotypic breast cancer status was evaluated using age- and family history-adjusted logistic regression in a set of 4,615 unspecified African American patients.
[0149] The association between polygenic risk scores and phenotypic breast cancer status, calculated using the local ancestry-corrected method of the present invention, was surprisingly more accurate than that of the conventional "overall" ancestry-corrected methods and the conventional "mean" ancestry-corrected methods.
[0150] Table 1 shows the results obtained with a set of 64 risk markers, 19 of which were assumed to influence breast cancer risk in both European and African ancestral contexts. As shown in Table 1, the present invention's method for determining a polygenic risk score adjusted for local ancestry was surprisingly superior and more accurate than conventional "overall ancestry adjusted" or "mean ancestry adjusted" methods. The p-values in Table 1 show more than a threefold enhancement by the present invention's method for determining a polygenic risk score by adjustment or correction using local ancestry.
[0151] [Table 1]
[0152] Table 2 shows a set of 64 risk markers.
[0153] [Table 2-1] [Table 2-2]
[0154] [Example 6] Local ancestor markers Figure 7 shows that several markers may be required for local ancestor assignment of Eur / Afr. Surprisingly, a high accuracy of 90%–95% for local Afr / Eur ancestor identification can be achieved using a relatively small number of flanking markers (x-axis). 95% accuracy (dashed line), 90% accuracy (solid line).
[0155] Figure 8 shows that several markers may be required for local ancestry assignment of Eur / EAS. Surprisingly, a high accuracy of 90%–95% can be achieved in separating Europeans from East Asian ancestry using a relatively small number of flanking markers (x-axis). 95% accuracy (dashed line), 90% accuracy (solid line).
[0156] Figure 9 shows the error rates of local ancestral assignment by unstaggered genotypes. Figure 9 shows the results for three SNP loci, rs258809, rs10759243, and rs1550623, with the y-axis showing the error rate of local ancestral deconvolution as a function of the number of flanking loci.
Claims
1. A method for evaluating biological traits in a subject, A step of measuring the genotype in a sample derived from a target, wherein the genotype has a window centered on trait risk markers for the trait, and the window includes ancestral knowledge markers that frank to the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; A step of calculating the odds of local ancestral origin for each window, based at least on ancestral knowledge markers; and A process for calculating polygeneity risk scores for biological traits in a subject, using trait risk markers and local ancestral origin odds for each window. A polygenicity risk score, including values greater than the threshold level, is a method for characterizing the degree of a biological trait.
2. The method according to claim 1, wherein the number of trait risk markers is 1 to 10,000.
3. The method according to claim 1, wherein the step of calculating local ancestral odds comprises dividing the stratified genotypes within each window into a sequence of non-overlapping tiles of up to 300 ancestral knowledge markers, and calculating the ancestral odds for each haplotype in each tile using the empirical frequency of the haplotype in the reference population.
4. The method according to claim 3, wherein each tile includes 1 to 100 ancestor-knowledge markers or 5 to 20 ancestor-knowledge markers.
5. The method according to claim 1, wherein the window has a width of 1 MB.
6. The method according to claim 1, wherein the genotype is determined by NGS or by a sequencing chip.
7. The method according to claim 1, wherein the biological trait is the likelihood of cancer.
8. The method according to claim 1, wherein the genomic risk marker is a cancer marker.
9. The method according to claim 1, wherein the ancestral knowledge marker is an SNP marker or an indel marker.
10. The method according to claim 1, wherein the genomic risk marker is a breast cancer SNP marker.
11. The method according to claim 1, wherein the step of calculating a polygenicity risk score includes calculating the incremental contribution of each allele to the polygenicity risk score as a local ancestor-specific risk effect beta obtained by multiplying several genotyped risk alleles, wherein this is zero or 1 and less than the population-specific risk allele frequency.
12. A method for recommending treatment for individuals with a disease: A step of measuring the genotype in a sample derived from a subject, wherein the genotype has a window centered on trait risk markers for a biological trait, and the window includes ancestral knowledge markers that frank to the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; A step of calculating the odds of local ancestral origin for each window, based at least on ancestral knowledge markers; and A process for calculating polygeneity risk scores for biological traits in a subject, using trait risk markers and local ancestral origin odds for each window. A polygenic risk score that includes a threshold level greater than the threshold indicates the need for treatment, and the disease is cancer.
13. The method according to claim 12, wherein the disease is breast cancer.
14. Treatment is: Treatment for a disease, selected from one or more of the following: surgery, cryoablation, radiotherapy, bone marrow transplantation, chemotherapy, immunotherapy, hormone therapy, stem cell therapy, drug therapy, biological therapy, and administration of pharmaceuticals, prophylactic or therapeutic compounds; Treatment for the disease following the monitoring period; Gradual reduction of treatment for a disease The method according to claim 12, which is one of the methods.
15. A method for identifying subjects with diseases who would benefit from treatment: A step of measuring the genotype in a sample derived from a subject, wherein the genotype has a window centered on trait risk markers for a biological trait, and the window includes ancestral knowledge markers that frank to the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; A step of calculating the odds of local ancestral origin for each window, based at least on ancestral knowledge markers; and A process for calculating polygeneity risk scores for biological traits in a subject, using trait risk markers and local ancestral origin odds for each window. A polygenic risk score, including values greater than the threshold level, characterizes the degree of the biological trait as one that benefits from disease treatment, and the disease is cancer, by method.
16. The method according to claim 15, wherein the disease is breast cancer.
17. The procedure is: Treatment for a disease, selected from one or more of the following: surgery, cryoablation, radiotherapy, bone marrow transplantation, chemotherapy, immunotherapy, hormone therapy, stem cell therapy, drug therapy, biological therapy, and administration of pharmaceuticals, prophylactic or therapeutic compounds; Treatment for the disease following the monitoring period; Gradual reduction of treatment for a disease The method according to claim 15, which is one of the methods.
18. A method for monitoring the response of a person with a disease: A step of measuring the genotype in a sample derived from a subject, wherein the genotype has a window centered on trait risk markers for a biological trait, and the window includes ancestral knowledge markers that frank to the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; A step of calculating the odds of local ancestral origin for each window, based at least on ancestral knowledge markers; and A process for calculating polygeneity risk scores for biological traits in a subject, using trait risk markers and local ancestral origin odds for each window. A polygenicity risk score, including values greater than the threshold level, is a method for characterizing the degree of a biological trait.
19. A method for predicting the prognosis of a person with a disease: A step of measuring the genotype in a sample derived from a subject, wherein the genotype has a window centered on trait risk markers for a biological trait, and the window includes ancestral knowledge markers that frank to the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; A step of calculating the odds of local ancestral origin for each window, based at least on ancestral knowledge markers; and A process for calculating polygeneity risk scores for biological traits in a subject, using trait risk markers and local ancestral origin odds for each window. A method that includes a polygenic risk score, where a numerical value greater than a threshold level characterizes the biological trait as having a poor prognosis for the disease.
20. A system for evaluating the risk of disease in a subject: A processor for receiving genomic data from a target sample; A step of measuring the genotype in a sample derived from a target, wherein the genotype has a window centered on trait risk markers for the trait, and the window includes ancestral knowledge markers that frank to the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; A step of calculating the odds of local ancestral origin for each window, based at least on ancestral knowledge markers; and A process for calculating polygeneity risk scores for biological traits in a subject, using trait risk markers and local ancestral origin odds for each window. One or more processors for executing; and Display for presenting and / or reporting risk scores A system that includes this.
21. A non-temporary machine-readable storage medium storing instructions for execution by a processor causing the processor to perform steps of a method for evaluating the risk of disease in a subject, wherein the method is: The process of obtaining genomic data from a sample derived from the target; A step of measuring the genotype in a sample, wherein the genotype has a window centered on trait risk markers for a trait, and the window includes ancestral knowledge markers that frank to the risk markers; A process of stratifying genotypes and determining the haplotypes within each window using a reference population with mixed ancestry; A process of calculating the odds of local ancestral origin for each window, based at least on ancestral knowledge markers; A process for calculating a polygenetic risk score for a biological trait in a subject using trait risk markers and the odds of local ancestral origin for each window; and The process of sending an output to the processor that presents and / or reports a risk score. Non-temporary, machine-readable storage media, including [specific data / information].