Ancestry-adjusted polygenic risk score (PRS) models and model pipeline
Patent Information
- Application Number
- EP2024767648
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-09-01
- Filing Date
- 2024-03-01
- Publication Date
- 2026-01-14
AI Technical Summary
Current polygenic risk score (PRS) models primarily focus on single ancestry groups, particularly European populations, leading to biases and reduced effectiveness in predicting disease susceptibility across diverse ancestral groups, limiting their applicability in precision medicine and global health.
A computer-implemented method for generating ancestry-adjusted PRS models by processing genotype and phenotype data using genetic primary component analysis and genome-wide association studies, incorporating estimates of global ancestry to create models that can predict disease susceptibility in diverse populations, utilizing direct-to-consumer genetic data to account for individual ancestry.
The ancestry-adjusted PRS models provide improved predictive performance and accuracy across various ancestral groups, enabling early diagnosis and prevention of chronic diseases like Type 2 Diabetes and hypertension, and can be scaled for global precision health applications.
Smart Images

Figure US2024018190_12092024_PF_FP_ABST
Abstract
Description
ANCESTRY-ADJUSTED POLYGENIC RISK SCORE (PRS) MODELS AND MODELPIPELINECROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 449,703, filed March 3, 2023, and U.S. Provisional Patent Application No. 63 / 536,315, filed September 1, 2023, each of which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Field of the invention: the present disclosure provides systems, software, and methods for generating and evaluating polygenic risk score models. The systems, software, and methods of the present disclosure may have applications in the areas of polygenic risk score (PRS) models for predicting disease state, and other phenotypes, of individuals based on genetic data corresponding to the individuals. The systems, software, and methods of the present disclosure may have applications in the areas of genetics, precision medicine, and personalized medicine.
[0003] Early diagnosis and prevention of chronic modem diseases, including type 2 diabetes (T2D) and hypertension, may have the potential to make a significant impact in patient outcomes. Early diagnosis may allow for better allocation of intervention strategies that are effective at reducing the risk of disease progression. According to medical practitioners, challenges with insufficient screening may arise due to the fact that chronic diseases tend to progress slowly until they manifest clinically later in life.SUMMARY
[0004] Genetics of Type 2 Diabetes (T2D) may be extensively studied, with over 400 genetic variants found to be associated with the diseases. Polygenic risk scores (PRS) may be generated by operating a trained PRS model on genetic data corresponding to an individual. PRS represent a measure of the individual's overall genetic liability to a trait or disease. The rapid development of (PRS) stresses the importance of accurately assessing the ancestry makeup of participants in biomedical studies to avoid potential selection biases. However, current PRS models may mostly include single ancestry participants (typically European) which may not generalize across other ancestral groups.
[0005] In view of the above, there is a need for a PRS modelling system that generates predictive tools trained on heterogeneous datasets that are able to predict susceptibility using historical data available outside of clinical and research settings. Further, there is a need for a PRS modelling system that takes individual ancestry into account when generating PRS. These predictive tools may be effective for identifying individuals at risk of various undesirable outcomes, including risk of certain diseases. A major challenge to enabling precision health at aglobal scale is the bias between those who enroll in state sponsored genomic research and those suffering from chronic disease. The present disclosure provides a PRS programming pipeline that enables information contained in direct to consumer (DTC) databases of genetic and phenotype data to be used to generate PRS models. Enabling the use of DTC data is advantageous at least in that DTC databases include information corresponding to groups of individuals having diverse ancestries as well as individuals having unique, identified, specific ancestry. By using ancestry-diverse and ancestry-specific DTC data, the PRS programming pipeline is enabled to generate PRS models that provide improved ancestry-based performance of PRS models generated according to the technology disclosed herein.
[0006] In an aspect, the present disclosure provides a computer-implemented method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, comprising: (a) processing, using a genetic primary component analysis (PC A) model, genotype data and phenotype data corresponding to one or more populations of individuals to generate genetic primary components (PCs) corresponding to the genotype data and the phenotype data; (b) processing, using a genome-wide association study (GWAS) model, the genotype data and the phenotype data, thereby generating GWAS results, wherein the GWAS model comprises covariables comprising at least one of the genetic PCs; (c) obtaining estimates of global ancestry corresponding to individuals of the one or more populations of individuals; and (d) training a PRS model using at least the GWAS results and the estimates of global ancestry as covariates, thereby generating the trained ancestry-adjusted PRS model.
[0007] In some embodiments, (a) further comprises using reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the covariables comprise at least 10 of the genetic PCs.
[0008] In some embodiments, (c) further comprises receiving the estimates of global ancestry from a computer system.
[0009] In some embodiments, (c) further comprises processing the genotype data using a trained global ancestry model, thereby determining the estimates of global ancestry. In some embodiments, the trained global ancestry model is trained using reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the trained global ancestry model is trained using an instance of a Neural ADMIXTURE model. In some embodiments, the method further comprises training a global ancestry model using reference genomic data from a population with known ancestry or ancestry distribution, thereby generating the trained global ancestry model. In some embodiments, training the global ancestry model further comprises training an instance of a Neural ADMIXTURE model.
[0010] In some embodiments, the genotype data and the phenotype data comprise a merged array. In some embodiments, the method further comprises: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a direct-to-consumer (DTC) platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data, thereby harmonizing each of the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate the merged array. In some embodiments, performing the imputation on the genotype data further comprises using reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the method further comprises identifying individuals and variants for removal when genotype data and phenotype data corresponding to the individuals and variants fail to meet a QC threshold.
[0011] In some embodiments, the method further comprises storing the ancestry-adjusted PRS model in a repository.
[0012] In another aspect, the present disclosure provides a computer-implemented method for harmonizing data corresponding to one or more direct-to-consumer (DTC) platforms, comprising: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a DTC platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data using reference genomic data from a population with known ancestry or ancestry distribution, thereby harmonizing the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate a merged array comprising genotype and phenotype data from one or more populations of individuals.
[0013] In some embodiments, the method further comprises using the merged array to generate a PRS model.
[0014] In another aspect, the present disclosure provides a computer system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, upon execution by the one or more computer processors, implements a method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, the method comprising: (a) processing, using a genetic primary component analysis (PC A) model, genotype data and phenotype data corresponding to one or more populations of individuals to generate genetic primary components (PCs) corresponding to the genotype data and the phenotype data; (b) processing, using a genome-wide association study (GWAS) model, the genotype data and the phenotype data, thereby generating GWAS results, wherein the GWAS model comprises covariables comprising at least one of the genetic PCs; (c) obtaining estimates of global ancestry corresponding to individuals of the one or more populations of individuals; and (d) training a PRS model using at least the GWAS results and the estimates of global ancestry as covariates, thereby generating the trained ancestry-adjusted PRS model.
[0015] In some embodiments, (a) further comprises using reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the covariables comprise at least 10 of the genetic PCs.
[0016] In some embodiments, (c) further comprises receiving the estimates of global ancestry from a computer system.
[0017] In some embodiments, (c) further comprises processing the genotype data using a trained global ancestry model, thereby determining the estimates of global ancestry. In some embodiments, the trained global ancestry model is trained using reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the trained global ancestry model is trained using an instance of a Neural ADMIXTURE model. In some embodiments, the method further comprises training a global ancestry model using reference genomic data from a population with known ancestry or ancestry distribution, thereby generating the trained global ancestry model. In some embodiments, training the global ancestry model further comprises training an instance of a Neural ADMIXTURE model.
[0018] In some embodiments, the genotype data and the phenotype data comprise a merged array. In some embodiments, the method further comprises: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a direct-to-consumer (DTC) platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data andphenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data, thereby harmonizing each of the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate the merged array.
[0019] In some embodiments, performing the imputation on the genotype data further comprises using reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the method further comprises identifying individuals and variants for removal when genotype data and phenotype data corresponding to the individuals and variants fail to meet a QC threshold.
[0020] In some embodiments, the method further comprises storing the ancestry-adjusted PRS model in a repository.
[0021] In another aspect, the present disclosure provides a computer system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, upon execution by the one or more computer processors, implements a method for harmonizing data corresponding to one or more direct-to- consumer (DTC) platforms, the method comprising: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a DTC platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data using reference genomic data from a population with known ancestry or ancestry distribution, thereby harmonizing the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate a merged array comprising genotype and phenotype data from one or more populations of individuals.
[0022] In some embodiments, the method further comprises using the merged array to generate a PRS model.
[0023] In another aspect, the present disclosure provides a non-transitory computer readable medium comprising machine-executable code that, upon execution by one or more computer processors, implements a method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, the method comprising: (a) processing, using a genetic primary component analysis (PC A) model, genotype data and phenotype data corresponding to one or more populations of individuals to generate genetic primary components (PCs) corresponding to thegenotype data and the phenotype data; (b) processing, using a genome-wide association study (GWAS) model, the genotype data and the phenotype data, thereby generating GWAS results, wherein the GWAS model comprises covariables comprising at least one of the genetic PCs; (c) obtaining estimates of global ancestry corresponding to individuals of the one or more populations of individuals; and (d) training a PRS model using at least the GWAS results and the estimates of global ancestry as covariates, thereby generating the trained ancestry-adjusted PRS model.
[0024] In some embodiments, (a) further comprises using reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the covariables comprise at least 10 of the genetic PCs.
[0025] In some embodiments, (c) further comprises receiving the estimates of global ancestry from a computer system.
[0026] In some embodiments, (c) further comprises processing the genotype data using a trained global ancestry model, thereby determining the estimates of global ancestry. In some embodiments, the trained global ancestry model is trained using reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the trained global ancestry model is trained using an instance of a Neural ADMIXTURE model. In some embodiments, the method further comprises training a global ancestry model using reference genomic data from a population with known ancestry or ancestry distribution, thereby generating the trained global ancestry model. In some embodiments, training the global ancestry model further comprises training an instance of a Neural ADMIXTURE model.
[0027] In some embodiments, the genotype data and the phenotype data comprise a merged array. In some embodiments, the method further comprises: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a direct-to-consumer (DTC) platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data, thereby harmonizing each of the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate the merged array. In some embodiments, performing the imputation on the genotype data further comprises using reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the method further comprises identifying individuals and variants for removalwhen genotype data and phenotype data corresponding to the individuals and variants fail to meet a QC threshold.
[0028] In some embodiments, the method further comprises storing the ancestry-adjusted PRS model in a repository.
[0029] In another aspect, the present disclosure provides a non-transitory computer readable medium for harmonizing data corresponding to one or more direct-to-consumer (DTC) platforms, comprising, with one or more processors: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a DTC platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data using reference genomic data from a population with known ancestry or ancestry distribution, thereby harmonizing the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate a merged array comprising genotype and phenotype data from one or more populations of individuals.
[0030] In some embodiments, the method further comprises using the merged array to generate a PRS model.
[0031] In another aspect, the present disclosure provides a computer-implemented method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, comprising: (a) processing genotype data and phenotype data corresponding to populations of individuals to generate genetic primary components (PCs) and genome-wide association study (GWAS) results using covariables comprising at least one of the genetic PCs; and (b) training a PRS model using at least the GWAS results and estimates of global ancestry corresponding to individuals of the populations as covariates, thereby generating the trained ancestry-adjusted PRS model.
[0032] In another aspect, the present disclosure provides a computer-implemented method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, comprising generating genome-wide association study (GWAS) results from populations of individuals, and training a PRS model using at least the GWAS results and estimates of global ancestry corresponding to individuals of the populations.
[0033] In another aspect, the present disclosure provides a computer-implemented method comprising processing test genotype data of a test subject with an ancestry-adjusted polygenic risk score (PRS) model to determine an ancestry-adjusted polygenic risk score of the test subject.
[0034] In some embodiments, the ancestry-adjusted PRS model is obtained at least in part by: (a) processing, using a genetic primary component analysis (PCA) model, genotype data and phenotype data corresponding to one or more populations of individuals to generate genetic primary components (PCs) corresponding to the genotype data and the phenotype data; (b) processing, using a genome-wide association study (GWAS) model, the genotype data and the phenotype data, thereby generating GWAS results, wherein the GWAS model comprises covariables comprising at least one of the genetic PCs; (c) obtaining estimates of global ancestry corresponding to individuals of the one or more populations of individuals; and (d) training a PRS model using at least the GWAS results and the estimates of global ancestry as covariates, thereby generating the trained ancestry-adjusted PRS model.
[0035] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.
[0036] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
[0037] The above and other features of the invention including various novel details of construction and combinations of parts, and other advantages, will now be more particularly described with reference to the accompanying drawings and pointed out in the claims. It will be understood that the particular method and device embodying the invention are shown by way of illustration and not as a limitation of the invention. The principles and features of this invention may be employed in various and numerous embodiments without departing from the scope of the invention.
[0038] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0039] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. Tothe extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “figure” and “FIG.” herein), of which:
[0041] Figure 1 is an example schematic diagram of a PRS system.
[0042] Figure 2 illustrates an example of a process flow depicting and operating method.
[0043] Figure 3 A is an example schematic diagram illustrating a first embodiment of a PRS pipeline.
[0044] Figure 3B is an example schematic diagram illustrating a second embodiment of a PRS pipeline.
[0045] Figure 4 depicts data plots illustrating results of an example implementation of the process flow of Figure 2.
[0046] Figure 5A depicts tabulated data illustrating performance metrics corresponding to the example implementation of the process flow of Figure 2.
[0047] Figure 5B depicts data plots illustrating performance metrics corresponding to the example implementation of the process flow of Figure 2.
[0048] Figure 6 depicts data plots illustrating comparison of results of an example implementation of the process flow of Figure 2 with results of other PRS models.
[0049] Figure 7A depicts a correlation plot between two PGS models for gout.
[0050] Figure 7B depicts a QQ-plot between the two PGS models for gout of Figure 7A.
[0051] Figure 7C depicts a correlation plot between the two PGS models for gout of Figure7A with ancestry.
[0052] Figure 8 depicts correlation plots and agreement matrices between two pairs of PGS models for liver enzymes.
[0053] Figure 9 depicts a schematic of a workflow to evaluate existing European-focused PRS.
[0054] Figure 10 depicts an example of risk levels provided by PRS models.
[0055] Figures 11 A-l 1C depict proportion of individual per ancestry for a) UKBB, and b) Galatea Bio collection; c) Ancestry deconvolution for Galatea Bio collection generated using GB proprietary software.
[0056] Figure 12 shows a computer system 1201 that is programmed or otherwise configured to perform operations of the methods.DETAILED DESCRIPTION
[0057] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0058] The term “PLINK 2,” as used herein, generally refers to a program used to identify genetically related individuals (by applying the King algorithm), to perform primary component analysis (PC A) to identify global ancestry per individual, and to perform genome wide association studies (GWAS)
[0048] ,
[0059] The term “Beagle,” as used herein, generally refers to a program used to perform imputation on genotype data and arrays
[0049] ,
[0060] The term “Neural ADMIXTURE,” as used herein, generally refers to a trained neural network used to infer global ancestry per individual.
[0061] The term “1000 Genomes Project Consortium,” as used herein, generally refers to a provider of genotype data corresponding to a reference population having known ancestry used as a reference population for use in performing PCA and imputation and for training a Neural ADMIXTURE model
[0057] .
[0062] The term “BASIL algorithm,” as used herein, generally refers to a Batch Screening Iterative LASSO (BASIL) algorithm used to generate PRS models
[0050] ,
[0063] The term “pROC package,” as used herein, generally refers to a pROC package in R, which may be used for evaluating PRS models using the area under the curve (AUC) receiver operating characteristic (ROC) curves
[0058] ,
[0064] References:
[0065] 48. Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, Sham PC.PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet. 2007;81(3):559-75, which is incorporated by reference in its entirety.
[0066] 49. Ayres DL, Darling A, Zwickl DJ, Beerli P, Holder MT, Lewis PO, Suchard MA.BEAGLE: an application programming interface and high-performance computing library for statistical phylogenetics. Syst Biol. 2012;61 ( 1 ) : 170-3, which is incorporated by reference in its entirety.
[0067] 50. Qian J, Tanigawa Y, Du W, Aguirre M, Chang C, Tibshirani R, Hastie T. A fast and scalable framework for large-scale and ultrahigh-dimensional sparse regression withapplication to the UK Biobank. PLoS Genet. 2020; 16(10): el009141, which is incorporated by reference in its entirety.
[0068] 55. Mantes AD, Montserrat DM, Bustamante CD, Giro-i-Nieto X, loannidis AG.Neural ADMIXTURE: rapid population clustering with autoencoders; 2021. bioRxiv, which is incorporated by reference in its entirety.
[0069] 57. Genomes Project Consortium. A global reference for human genetic variation.Nature. 2015;526(7571):68, which is incorporated by reference in its entirety.
[0070] 58. Robin X, Turck N, Hainard A, Tiberti N, Lisacek F, Sanchez JC, Muller M. pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinform. 2011; 12(1): 1—8, which is incorporated by reference in its entirety.
[0071] More than 30 million people have been genotyped by direct-to-consumer (DTC) companies such as 23andMe, Ancestry DNA, and MyHeritage, providing a potential mechanism for democratizing access to medical interventions and thus catalyzing improvements in patient outcomes as the cost of data acquisition drops. However, much of these data are sequestered in the initial provider network, without the ability for the scientific community to either access or validate.
[0072] The present disclosure provides a novel genotype-phenotype platform that integrates heterogeneous data sources and applies learnings to common chronic disease conditions including, for example, Type 2 diabetes (T2D) and hypertension.
[0073] The present disclosure provides a PRS modelling system, or PRS system for short, that uses data from DTC platforms to invert research-based PRS models that are developed using genome sequencing and phenotypic data that is acquired as part of a designed research study. Quality control (QC) mechanisms of the present disclosure may enable traditional GWAS and PRS analyses based on DTC data. Inferred genetic ancestry information corresponding to individuals is used by the PRS system to adjust PRS models, thereby generating ancestry- adjusted PRS models. The genetics of ancestry-adjusted PRS models generated by the PRS system are supported by their ability to replicate known variants from publicly available independent GWAS studies.
[0074] The direct participation of individuals in providing data is used to leverage the potential to generate rich datasets enabling the creation of ancestry-adjusted PRS cardiometabolic and other disease state models. More importantly, federated learning of PRS from reuse of DTC data, as disclosed herein, provides a mechanism for scaling precision health care delivery beyond the small number of countries who can afford to finance these efforts directly.
[0075] The genetics of many disease states, for example T2D and hypertension, may be studied extensively in controlled datasets, and various polygenic risk scores (PRS) may bedeveloped. The present disclosure enables predictive tools for both phenotypes trained with heterogeneous genotypic and phenotypic data, e.g. DTC data, generated outside of the clinical environment.
[0076] Methods of the present disclosure recapitulate prior findings using various techniques with fidelity. The methods and systems of the present disclosure may be used to leverage DTC genetic repositories to identify individuals at risk of debilitating diseases based on their unique genetic landscape so that informed, timely clinical interventions can be incorporated. Ancestry- adjusted PRS models generated by the PRS system of the present disclosure can replicate previous findings and deliver enhanced discovery and single-variant resolution of causal T2D and hypertension risk and protection alleles. Additionally, the results provided by the present disclosure have confirmed the potential impact of DTC resources on mechanistic insights and clinical translation efforts.
[0077] The lack of ancestry diversity in current PRS models may result in systemic biases that threaten to widen existing health disparities among minority and majority populations in most developed countries. The overrepresentation of European individuals in genetic studies represents a major issue, hampering the translation of known PRS across populations.
[0078] Ancestry-adjusted PRS models of the present disclosure may provide improvements over current PRS models by defining and taking into account the genetic ancestry of participants. In some embodiments, genetic ancestry is embedded in as a covariate in the ancestry-adjusted PRS models. In other embodiments, genetic ancestry is embedded by scaling PRS scores as part of post-processing step. In this manner, the present disclosure provides more accurate models than traditional filtering to only European-based PRS models.
[0079] Furthermore, the present disclosure provides a validated DTC framework for validating and extending PRS models based on DTC data. This enables cost-effective approaches of enrolling understudied populations in complex disease genetics. An example embodiment of the DTC framework incorporates DTC data included in the Biobank Metaanalysis Network. The DTC framework provides a “straight to mobile instead of landlines” opportunity. This is advantageous for enabling low- and middle-income countries to harness advances in genomics for the study of their own populations where state-sponsored capacity for performing current methods is limited.
[0080] The PRS system is capable of generating ancestry-adjusted PRS models for one or more disease states or other phenotypes from heterogeneous datasets, for example from datasets housing a combination of genetic data and self-reported information from one or more DTC genetics companies. The ancestry-adjusted PRS models are able to identify subsets of users at substantially increased risk of presenting one or more disease states or other phenotypes ofinterest. This is particularly advantageous because it indicates that the ever-increasing availability of genetic data from DTC providers, most of it not annotated for traits of clinical relevance, can be leveraged by the present disclosure to generate predictive tools capable of improving diagnosis and prevention of diseases with genetic determinants.
[0081] DTC platforms can offer a wide range of information about personal wellness, ancestry, physical characteristics, and traits. Advances in genomic research may lead the DTC genomics industry to flourish and make accurate yet easy-to-interpret genomic results. Strict privacy policies of many companies may disallow them to share customers’ data without their consent. These platforms can serve as informative repositories giving actionable insights that aid traditional clinical approaches. The approach of subject recruitment for various complex phenotypes via online surveys may open up multiple avenues to complement conventional research and clinical strategies. DTC platforms also provide convenience along with a wider reach to recruit participants from various locations. They surpass barriers of single-point data collection centers to language restrictions thus allowing the aggregation of data from places with different ancestries and demographics. Democratizing the access to these genetic platforms and prediction tools may boost progress in precision medicine.
[0082] A PRS system according to the present disclosure may be able to replicate reported results with a very fast turnaround time. The participation of individual customers with the PRS system platform allows for the generation of a rich dataset that enables the creation of ancestry- adjusted PRS models according to the subject technology. The comparable predictive performance of the ancestry-adjusted PRS models enables the PRS system to quickly generate more ancestry-adjusted PRS models. The ancestry-adjusted DTC models replicate the findings seen in academic and government-funded biobanks. The clinical actionability of current PRS models has yet to be determined through pragmatic trials involving real-world data. Ancestry- adjusted PRS models according to the present disclosure provide a novel source of information that can shed light on this important issue, and to improve early diagnosis and prevention bringing precision medicine at scale for all.
[0083] Example PRS system
[0084] Figure 1 shows an example PRS system 100. The PRS system 100 is configured to generate PRS models, including PRS models that are scaled based on ancestry. The PRS system 100 is configured to operate one or more ancestry models to generate estimates of ancestry for individuals. The PRS system 100 is also configured to operate one or more PRS models to generate PRS results and to scale the PRS results based on the ancestry estimates.
[0085] The PRS system 100 is in communication with an external data source 200. The external data source 200 represents one or more external data sources, for example one or moreDTC platforms or repositories of data provided by the one or more DTC platforms, one or more sources of vcf files, one or more PRS model catalogs, and one or more other data stores, for example a data store comprising genotype data corresponding to one or more populations of known ancestry.
[0086] A data receiving interface 100 receives data from the data source 200, for example DTC data from one or more DTC platforms, and can communicate the data to a QC module 115 and to an internal data store 120. The QC module 115 is configured to perform one or more QC operations on the received data to generate QC’d data. For example, the QC module 115 can detect and eliminate outliers in DTC data and can screen the DTC data for replicated individuals, related individuals, and individuals for whom data necessary for further processing steps is missing. The QC module communicates QC’d data to and imputation module 125 and to an ancestry module 145. The QC module may also store QC’s data in the data store 120.
[0087] The imputation module 125 performs imputation on genotype data, for example on QC’d DTC data, to infer unobserved genotypes in the data. For example, if a genotyping array received from a first DTC platform includes one or more variants that are missing from a genotyping array received from a second DTC platform, the imputation module 125 may impute valued for the missing variants in the second DTC platform to harmonize the first and second genotyping arrays. The imputation module 125 may further merge imputed data sets, for example multiple imputed data sets based on DTC data received from multiple different DTC platforms, to generate merged genetic data, e.g. merged DTC data.
[0088] The ancestry module 145 operates one or more models or methods one received data, for example on DTC data, to generate ancestry-related data. For example, the ancestry module 145 can operate a primary components analysis (PCA) model to generate a set of primary components (PCs) corresponding to ancestry of a dataset (e.g., genetic PCs). In another example, the ancestry module 145 can operate a global ancestry model, for example a trained Neural ADMIXTURE model, on genetic data to generate individual ancestry estimates.
[0089] The PRS system 100 includes a GWAS module 130 that performs genome wide association studies (GWAS) on genetic data, for example on merged DTC data, to generate GWAS results. In some embodiments, the GWAS module 130 receives genetic PCs from the ancestry module 145 and uses them as covariables in a GWAS model. The PRS system 100 includes an output module 150, which can include, for example, one or more of a display device, a data communication interface, a data store, and a processor for further processing of data, for example to enable display on the display device. In some embodiments, the GWAS module 130 communicates GWAS results to the output module for display on a display device thereof.
[0090] A PRS generation module 135 receives GWAS results from the GWAS module 130 as well as ancestry information from the ancestry module 145, for example individual ancestry estimates. The PRS generation module creates one or more PRS models based on the GWAS results and the ancestry information. The PRS generation can communicate PRS models to the output module 150, for example for sharing with one or more external systems or model repositories.
[0091] The PRS system 100 includes a PRS execution module 140. The PRS module 140 is configured to receive or load a PRS model and operate the PRS model on individual genotype data to generate PRS. In a first example operation, the PRS module 140 receives and operates PRS models generated by the PRS system 100 from the PRS generation module 135. In a second example operation, the PRS module 140 loads and operates an external PRS model from the data store 120, or directly from the data receiving interface 110. The external PRS model includes, for example, a PRS model from a PRS model catalog. The PRS execution module 140 can also receive individual ancestry estimates from the ancestry module 145 and scale PRS generated by using a PRS model with the individual ancestry estimates.
[0092] The PRS system 100 include a user interface (UI) 160 to enable a user to enter information to configure one or more aspects of the PRS system. The PRS system 100 further includes at least one processor 170 and associated memory for storing and operating the various modules and data stores. It is noted that embodiments of the PRS system 100 can include multiple processors 170, one or more individual computing devices, multiple UIs 160, and one or more of each of the various modules and data stores.
[0093] Example method for generating ancestry-adjusted PRS models
[0094] Figure 2 shows an example method for generating PRS models based on DTC data. The method can be performed by a PRS system of the present disclosure.
[0095] At operation 1010, the PRS system receives, from a DTC source, DTC data. In an example embodiment, the DTC data include genetic data as well as self-reported answers to questions asked of participants corresponding to the genetic data, for example using a questionnaire or online survey. Non-limiting examples of questions include, for example, questions about general conditions like diabetes, blood pressure, lipid profile, and medication intake, and questions about with age, sex, weight, height.
[0096] At operation 1015, the PRS system performs on or more quality control (QC) operations on the DTC data to generate QC’d DTC data. The QC operations can include, for example, QC checks, outlier detections, and corresponding data correction and / or pruning of the received DTC data. Non-limiting examples of QC operations that may be performed by the PRS system include one or more of:
[0097] Excluding participants for whom genotype data is not available;
[0098] Excluding participants for whom answers for particular questions, for example age and sex, are not available;
[0099] Excluding variants with a call rate less than a threshold value, for example call rate of < 95%;
[0100] Excluding variants having palindromic markers (e.g. A / T, G / C, MAF > 0.4);
[0101] Excluding individuals having genotype call rates less than a threshold value, for example less than or equal to 97%,
[0102] Excluding individual that do not have matching between gender identification and chromosomal sex;
[0103] Excluding individuals having excess ancestry-adjusted heterozygosity; and
[0104] Excluding individuals genetically related to other individuals in the cohort and duplicates, for example by applying the King algorithm (-make-king, king estimate > 0.177; PLINK 2).
[0105] At operation 1020, the PRS system parses the QC’d DTC data to generate case and control groups for PRS model development. Case groups include individuals whose data indicate they may have a particular condition for which a particular PRS model is being generated. For example, when generating a hypertension PRS model, the system may include participants who report having high blood pressure, participants who are taking blood pressure medications, and / or participants reporting lab results that may be indicative of hypertension. Control groups include participants who did not report managing a health condition relevant to the particular PRS model, and those who did not report managing any health condition. Control groups can also include participants who report lab results or other information that is indicative of the participant not having the relevant condition. In some embodiments of the method 1000, the PRS generates case and control groups prior to performing some or all of the QC operations.
[0106] At operation 1030, the PRS system performs imputation to generate values for missing genomic information in the QC’d DTC datasets, thereby generating imputed DTC datasets, i.e. datasets that include the QC’s DTC data and imputed data. In some embodiments, the PRS system performs imputation using reference genomic data from a population with known ancestry or ancestry distribution and using an imputation method, for example the Beagle method. In practice, DTC genetic data may be generated using one or more of multiple array types and a database of DTC genetic data may include data generated by multiple array types, each of which may produce a dataset with different characteristics. In this case, the PRS system uses imputation across array types (up-imputing) to harmonize these diverse datasets.
[0107] At operation 1035, the PRS system merges the imputed DTC data to generate merged DTC data. In some embodiments, the PRS system merges multiple imputed DTC datasets, each of which may correspond to a different array type, to generate the merged DTC data. In some embodiments, the PRS system performs additional QC operations on the merged DTC data. In some embodiments, the DTC system only includes imputed makers having a call rate greater than a threshold value, for example greater than 0.95, in the merged DTC data.
[0108] At operation 1040, the PRS system identifies individual ancestry of participants included in the DTC dataset. In an example embodiment, the PRS system performs principal component analysis (PCA) to identify global ancestry per individual in the merged DTC data, using reference genomic data from a population with known ancestry or ancestry distribution. In an example embodiment, the PRS system performs PCA using a PCA method, for example PLINK 2, thereby generating a set of genetic principal components (PCs).
[0109] At operation 1045, the PRS system, or an operator of the system, selects variables to be included as covariables in a genome-wide association study (GWAS) model. In an example embodiment, the top ten genetic PCs, age, sex, and the merged DTC data (i.e. the merged genotyping arrays) are selected as covariables in the model.
[0110] At operation 1050, the PRS system operates the GWAS model to generate GWAS model output. GWAS model output may include genomic risk loci, e.g. genetic variants, for example SNPs, or blocks of correlated genetic variants that show a statistically significant association with a trait of interest. In an example embodiment, the GWAS is performed using an additive genetic model, for example using PLINK 2.
[0111] The PRS system may, optionally, at operation 1052 validate the GWAS results that are generated at operation 1050. For example, the PRS system may validate the GWAS results using one or more independent GWAS meta-analysis datasets, for example by comparing the p- values and the effect sizes for the variants assessed in the GWAS results that have identical chromosomal coordinates and alleles with the independent GWAS meta-analysis data sets.
[0112] At operation 1055, the GWAS results are displayed by the PRS system, for example using a GWAS results display program. In some embodiments, the results are communicated to a processor running the qqman package in R and the qqman is used to format the results for presentation on a display device, for example on a display screen of a computer system. The PRS system may also store the results in a local or external data store. The GWAS results are also communicated to a PRS model generation module or system for use in generating one or more PRS models.
[0113] At operation 1060, the PRS system determines estimates of global ancestry for use in adjusting PRS models for ancestry. In some embodiments, the PRS system utilizes the results ofglobal ancestry inference as a covariate in the training of the PRS models. In some embodiments, the global ancestry inference if generated, by the PRS system or by an external system, using an ancestry inference method, for example using Neural ADMIXTURE, with reference genomic data from a population with known ancestry or ancestry distribution. In some embodiments, the PRS system, or an external system, uses data from the 1000 Genomes Project Consortium for training a model in the supervised mode of Neural ADMIXTURE with default parameters.
[0114] At operation 1065, the PRS system generates one or more ancestry-adjusted PRS models based on the GWAS results generated at operation 1050 and, in some embodiments, using the results of global ancestry inference generated in operation 1065 as a covariate in the training of the PRS models. In some embodiments, the PRS system generates the one or more ancestry-adjusted PRS models using a method, for example using a meta-algorithm (algorithms that learn from the output of other algorithms), e.g. using the Batch Screening Iterative LASSO (BASIL) algorithm. In this manner, the present disclosure provides ancestry-adjusted PRS models that are more accurate than traditional filtering to only European-based PRS models.
[0115] In some embodiments, genetic ancestry is embedded by scaling PRS scores as part of post-processing operation when the PRS models are operated to generate PRS scores. The PRS scores that are scaled may be generated using various PRS models or using ancestry-adjusted PRS models.
[0116] At operation 1070, the PRS system optionally tests the performance of and validates the ancestry-adjusted PRS models using validation data previously partitioned from the case and control DTC data sets. Ancestry-adjusted PRS models that do not perform according to a validation standard may be rejected and / or retrained. In some embodiments, the PRS system evaluates the predictive ability of these PRS models using the area under the curve (AUC) receiver operating characteristic (ROC) curves using a model test method, for example using the pROC package in R.
[0117] At operation 1075, the PRS system provides the ancestry-adjusted PRS models to one or more output location which can include, for example, data stores (internal or external), PRS model repository, and a PRS execution engine, which operates a PRS model on genome data to generate PRS. In some embodiments, the PRS model is processed through a PRS pipeline of the present disclosure.
[0118] Example PRS processing pipelines for operating PRS models
[0119] Referring now to Figures 3 A and 3B, two example embodiments of PRS pipelines 2000 and 2010 for processing PGS models are shown. Both pipelines are configured to operate on inputs 2100, including a VCF file 2110 and a PGS model 2120 to generate PRS outputsincluding scores and bins 2510 and featured SNPs 2520. The VCF file 2110 typically includes genotype data, and more specifically variants such as SNPs, corresponding to a particular individual for whom a PRS is desired. A PGS model 2120 can include any model in PGS catalog format, for example a PGS model available from an external PGS catalog or an ancestry-adjusted PRS model generated according to the subject technology disclosure herein.
[0120] Referring once more to Figures 3 A and 3B, each PRS pipeline 2000 and 2110 includes a run models engine 2200 which operates a PGS model 2120 on a VCF file 2110 to generate PRS output. Each run models engine 2200 includes an impute missing variants engine 2210, a filter VCF variants engine 2220, and a calculate PRS engine 2230.
[0121] Each impute missing variants engine 2210 operates on genotype data, for example genotype data included in the VCF file 2110, to impute missing variants therein. Missing variants may include any variants that are included in the PGS model 2120 but not observed in the VCF file 2110.
[0122] Each filter VCF variants engine 2220 operates on genotype data to filter out variants that are not useful as inputs to the PGS model 2120. In some embodiments, the filter VCF variants engine 2220 performs filtering operations on imputed data generated by the impute missing VCF variants engine 2210. In other embodiments, the filter variants engine 2220 performs filtering operations prior to imputation by the impute missing variables engine 2210. In still further embodiments, the filter variants engine 2220 performs additional QC operations to remove outliers and other non-useful data. For example, and referring now to Figure 2, the filter variant engine 2220 may perform one or more of the QC operations described in relation to operation 1015 (Perform QC on received data) of method 1000. Referring once again to Figures 3 A and 3b, in some embodiments, the impute missing variants engine 2210 performs imputation to generate imputed values for variants that are filtered out by a QC operation performed by the filter VCF variants engine 2220 but that are included in the PGS model 2120.
[0123] Each compute PRS engine 2230 operates on imputed and filtered VCF data to generate PRS outputs, for example PRS outputs 2500.
[0124] Referring now to Figure 3B, PRS pipeline 2010 includes additional ancestry -related engines, including an ancestry prediction engine 2300 and an ancestry scaling engine 2400.
[0125] The ancestry prediction engine 2300 generates global ancestry predictions 2310. In some embodiments, the ancestry prediction engine 2300 operates a trained global ancestry prediction model on genotype data from a VCF file 2110 to generate the global ancestry predictions. In some embodiments, the ancestry prediction engine 2300 operates the trained global ancestry prediction model on filtered and / or imputed genotype data generated by the runmodels engine 2200. In some embodiments, the trained global ancestry prediction model includes an instance of a Neural ADMIXTURE model, trained as described elsewhere herein.
[0126] The ancestry scaling engine 2400 receives PRS outputs from the run models engine 2200 and global ancestry predictions 2310, e.g. PRS outputs 2500, from the ancestry prediction engine 2300. The ancestry scaling engine uses the global ancestry predictions 2310 to scale the PRS outputs, thereby generating ancestry-adjusted PRS outputs 2600 including scores and bins 2610 and featured SNPs 2620. In some embodiments, the ancestry scaling engine operates an ancestry scaling ML model trained on ancestry and PRS data to scale the PRS outputs using the global ancestry predictions. In some embodiments, a first embodiment of ancestry scaling ML model is trained using results generated by a particular PRS model 2120, global ancestry information, and risk scores associated with the global ancestry information. The trained first embodiment of the ancestry scaling ML model is then used to scale risk scores generated by the particular PRS models according to global ancestry predictions 2310 corresponding to a subject individual for whom VCF file 2110 is received by the PRS pipeline.
[0127] Referring now to Figures 1, 3 A, and 3B, in some embodiments, and referring now to Figure 1, the PRS pipelines include embodiments of PRS system 100. For example, the impute missing variant engine 2210 may include or correspond to an imputation module 125 of PRS system 100. In a similar manner: filter VCF variant engine 2220 may include or correspond to QC module 115; and calculate PRS engine 2230 may include or correspond to PRS execution module 140. Ancestry predictions engine 2310 and ancestry scaling module 2400 may each include or correspond to the same or separate instances of ancestry module 145.
[0128] Genotype data and / or phenotype data may be generated from a subject. The subject may be a suspected of a suffering from a disease or condition, such as Type 2 diabetes (T2D) and hypertension. In some cases, the subject may be asymptomatic for the disease or condition. For example, the disease or condition may not exhibit any symptoms and the subject may be unaware of the presence of disease or condition. The methods described herein may allow a disease or condition to be identified at an earlier stage than otherwise. The identification of the presence of the disease or condition at an earlier stage may allow a treatment option or recommendation to be determined at an earlier stage and may allow the subject to have an improved prognosis.
[0129] Biological samples may be obtained or derived from the subject. The biological sample may comprise nucleic acids. The biological sample be a cell-free deoxyribonucleic acid (cfDNA) sample or a cell-free ribonucleic acid (cfRNA) sample. The biological sample may comprise genomic DNA or germline DNA (gDNA). The nucleic acid may be a DNA (e.g. double-stranded DNA, single-stranded DNA, single-stranded DNA hairpins, cDNA, genomic DNA, germline DNA, circulating tumor DNA (ctDNA), cell-free DNA (cfDNA)), an RNA (e.g.cfRNA, mRNA, cRNA, miRNA, siRNA, miRNA, snoRNA, piRNA, tiRNA, snRNA), or a DNA / RNA hybrids. The biological sample may be a derived from or contain a biological fluid. For example, the biological sample may be a plasma sample, a serum sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a red blood cell sample, a urine sample, a saliva sample, or other body fluid sample. The biological sample may comprise or be a pleural fluid sample, peritoneal fluid sample, amniotic fluid sample, cerebrospinal fluid sample, lymphatic fluid sample, sweat sample, tear sample, semen sample, or any combination of biological fluid. In some case, the samples may comprise RNA and DNA. For example, a sample may comprise DNA or RNA, and may be analyzed by various methods (e.g., for determining genotype and / or phenotype).
[0130] The biological sample may be collected, obtained, or derived from the subject using a collection tube. The collection tube may be an ethylenediaminetetraacetic acid (EDTA) collection tube, a cell-free RNA collection tube, or a cell-free deoxyribonucleic acid (DNA) collection tube and CTC collection tubes, or other blood collection tube. The collection tube may comprise additional reagents for stabilizing the nucleic acid molecules or blood cells. The collection tube may allow the nucleic acid or blood cells to be stable such to minimize degradation of the biological sample prior to assaying. The additional reagents may comprise buffer salts or chelators.
[0131] The biological sample may be obtained or derived from a subject at a various times. The biological sample may be obtained or derived from a subject prior to the subject receiving a therapy for a disease or condition. The biological sample may be obtained or derived from a subject during receiving a therapy for the disease or condition. The biological sample may be obtained or derived from a subject after receiving a therapy for the disease or condition. The biological sample may be collected over 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 15, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or time points. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more hour period. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more day period. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more week period. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more month period. The time points may occur over a 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60 or more year period.
[0132] In various aspects as described herein, a clinical intervention or a therapy may be identified at least in part based on the identification of the presence, likelihood, or elevated risk of the disease or condition. The clinical intervention may be a plurality of clinical interventions. The clinical intervention may be selected from a plurality of clinical interventions. In some cases, the clinical interventions may be administered to the subject. After administration of the clinical intervention, a sample may be obtained or derived from the subject such to monitor the treatment. Additionally, by performing the methods or systems iteratively, therapies or clinical interventions may be updated based on the results of the methods. The monitoring of the treatment may include an assessment as well as a difference in assessment from a previously generated assessment . The difference in an assessment of the disease or condition in the subject among a plurality of time points (or samples) may be indicative of one or more clinical indications such as a diagnosis of the disease or condition, a prognosis of the disease or condition, or an efficacy or non-efficacy of a course of treatment for treating the disease or condition of said subject.
[0133] The biological samples may be subjected to additional reactions or conditions prior to assaying. For example, the biological sample may be subjected to conditions that are sufficient to isolate, enrich, or extract nucleic acids, such DNA molecules or RNA molecules.
[0134] The methods disclosed herein may comprise conducting one or more enrichment reactions on one or more nucleic acid molecules in a sample. The enrichment reactions may comprise contacting a sample with one or more beads or bead sets. The enrichment reactions may comprise one or more hybridization reactions. For example, the enrichment reactions may comprise contacting a sample with one or more capture probes or bait molecules that hybridize to a nucleic acid molecule of the biological sample. The enrichment reaction may comprise differential amplification of a set of nucleic acid molecules. The enrichment reaction may enrich for a plurality of genetic loci or sequences corresponding to genetic loci. For example, the enrichment reaction may enrich for sequences corresponding to certain genes. The enrichment reactions may comprise the use of primers or probes that may complementarity to sequences (or sequences upstream or downstream) of a sequence that is to be enriched. For example, a capture probe may comprise sequence complementarity to a set of genomic loci and allow the enrichment of the genomic loci. The enrichments reactions may comprise a plurality of probes or primers. A plurality of probes may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, or 180 different probes.
[0135] The methods disclosed herein may comprise conducting one or more isolation or purification reactions on one or more nucleic acid molecules in a sample. The isolation orpurification reactions may comprise contacting a sample with one or more beads or bead sets. The isolation or purification reaction may comprise one or more hybridization reactions, enrichment reactions, amplification reactions, sequencing reactions, or a combination thereof. The isolation or purification reaction may comprise the use of one or more separators. The one or more separators may comprise a magnetic separator. The isolation or purification reaction may comprise separating bead bound nucleic acid molecules from bead free nucleic acid molecules. The isolation or purification reaction may comprise separating capture probe hybridized nucleic acid molecules from capture probe free nucleic acid molecules. The isolation reactions may comprises removing or separating a group of nucleic acid molecules from another group of nucleic acids.
[0136] The methods disclosed herein may comprise conduction extraction reactions on one or more nucleic acids in a biological sample. The extraction reactions may lyse cells or disrupt nucleic acid interactions with the cell such that the nucleic acids may be isolated, purified, enriched or subjected to other reactions.
[0137] The methods disclosed herein may comprise amplification or extension reactions. The amplification reactions may comprise polymerase chain reaction. The amplification reaction may comprise PCR-based amplifications, non-PCR based amplifications, or a combination thereof. The one or more PCR-based amplifications may comprise PCR, qPCR, nested PCR, linear amplification, or a combination thereof. The one or more non-PCR based amplifications may comprise multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, circle-to-circle amplification or a combination thereof. The amplification reactions may comprise an isothermal amplification.
[0138] The method disclosed herein may comprise a barcoding reaction. A barcoding reaction may comprise the additional of a barcode or tag to the nucleic acid. The barcode may be a molecular barcode or a sample barcode . For example, a barcode nucleic acid may comprise a barcode sequence which may be a degenerate n-mer. The sequence may be randomly generated or generated such to synthesize a specific barcode sequence. The barcode nucleic acid may be added to a sample such to label the nucleic acid molecules in the sample. The barcodes may be specific to a sample. For example, a plurality of barcode nucleic acids may be added to a sample in which the barcode sequence is the same. Upon barcoding of the nucleic acids, those originating from a same sample may have a same barcode sequence, and may allow a nucleic acid to be identified as belonging to a particular or given sample. A molecular barcode may also be used such that each molecule (or a plurality of molecules) in a same volume have a different molecular barcode. This barcode may be subjected to amplification such that all ampliconsderived from a molecule have the same barcode. In this way, molecules originating from a same molecule may be identified. The sequences reads may be processed based on the barcode sequences. For example, the processing may reduce errors or allow a molecule to be tracked. Barcode sequences may be appended or otherwise added or incorporated into a sequence by various reactions, for example an amplification, extension, or ligation reaction, and may be performed enzymatically using a nucleic acid polymerase or ligase. The ligation may be an overhang or blunt end ligation and the barcodes may comprise complementarity to nucleic acids to be barcoded. This complementarity may be a sequence derived from the sample from the subject or may be constant sequence generated via a reaction performed on the nucleic acids in the sample.
[0139] In some cases, the biological sample may comprise multiple components. For example, the biological sample may be a whole blood sample. The biological sample may be subjected to reactions such to separate or fractionate a biological sample. For example, a whole blood sample may be a fractionated and cell free nucleic acids may be obtained. The whole blood sample may be fractionated using centrifugation such that blood cells may be separated from the plasma (which may contain cell free nucleic acid). A sample may be subjected to multiple rounds of separation or fractionation.
[0140] In various aspects described throughout the disclosure, the nucleic acids may be subjected to sequencing reactions. The sequencing the reactions may be used on DNA, RNA, or other nucleic acid molecules. Example of a sequencing reaction that may be used include capillary sequencing, next generation sequencing, Sanger sequencing, sequencing by synthesis, single molecule nanopore sequencing, sequencing by ligation, sequencing by hybridization, sequencing by nanopore current restriction, or a combination thereof. Sequencing by synthesis may comprise reversible terminator sequencing, processive single molecule sequencing, sequential nucleotide flow sequencing, or a combination thereof. Sequential nucleotide flow sequencing may comprise pyrosequencing, pH-mediated sequencing, semiconductor sequencing or a combination thereof. The sequencing reactions may comprise whole genome sequencing, whole exome sequencing, low-pass whole genome sequencing, targeted sequencing, methylation-aware sequencing, enzymatic methylation sequencing, bisulfite methylation sequencing. The sequencing reaction may be a transcriptome sequencing, mRNA-seq, totalRNA- seq, smallRNA-seq, exosome sequencing, or combinations thereof. Combinations of sequencing reactions may be used in the methods described elsewhere herein. For example, a sample may be subjected to whole genome sequencing and whole transcriptome sequencing. As the samples may comprise multiple types of nucleic acids (e.g. RNA and DNA), sequencing reactionsspecific to DNA or RNA may be used such to obtain sequence reads relating to the nucleic acid type.
[0141] The sequencing of nucleic acids may generate sequencing read data. The sequencing reads may be processed such to generate data of improved quality. The sequencing reads may be generated with a quality score. The quality score may indicate an accuracy of a sequence read or a level or signal above a nose threshold for a given base call. The quality scores may be used for filtering sequencing reads. For example, sequencing reads may be removed that do not meet a particular quality score threshold. The sequencing reads may be processed such to generate a consensus sequence or consensus base call. A given nucleic acid (or nucleic acid fragment) may be sequenced and errors in the sequence may be generated due to reactions prior or during sequencing. For example, amplification or PCR may generate error in amplicons such that the sequences are not identical to a parent sequence. Using sample barcodes or molecular barcodes, error correction may be performed. Error correction may include identifying sequence reads that do not corroborate with other sequences from a same sample or same original parent molecules. The use of barcodes may allow the identification or a same parent or sample. Additionally, the sequence reads may be processed by performing single strand consensus calling or double stranded consensus call, thereby reducing or suppressing error.
[0142] The methods and systems of the present disclosure may be implemented, at least in part, in hardware, software, firmware or any combination thereof. For example, various aspects of the described techniques may be implemented within one or more processors, including one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or any other equivalent integrated or discrete logic circuitry, as well as any combinations of such components. The term “processor” or “processing circuitry” may generally refer to any of the foregoing logic circuitry, alone or in combination with other logic circuitry, or any other equivalent circuitry. A control unit comprising hardware may also perform one or more of the techniques of this disclosure.
[0143] Such hardware, software, and firmware may be implemented within the same device or within separate devices to support the various operations and functions described in this disclosure. In addition, any of the described units, modules or components may be implemented together or separately as discrete but interoperable logic devices. Depiction of different features as modules or units is intended to highlight different functional aspects and does not necessarily imply that such modules or units must be realized by separate hardware or software components. Rather, functionality associated with one or more modules or units may be performed by separate hardware or software components or integrated within common or separate hardware or software components.
[0144] The methods and systems of the present disclosure may also be embodied or encoded in a computer-readable medium, such as a computer-readable storage medium, containing instructions. Instructions embedded or encoded in a computer-readable medium may cause a programmable processor, or other processor, to perform the method, e.g., when the instructions are executed. Computer-readable media may include non-transitory computer-readable storage media and transient communication media. Computer readable storage media, which is tangible and non-transitory, may include random access memory (RAM), read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electronically erasable programmable read only memory (EEPROM), flash memory, a hard disk, a CD-ROM, a floppy disk, a cassette, magnetic media, optical media, or other computer-readable storage media. It should be understood that the term “computer-readable storage media” refers to physical storage media, and not signals, carrier waves, or other transient media.
[0145] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 12 shows a computer system 1201 that is programmed or otherwise configured to perform operations of the methods, for example, generating trained ancestry-adjusted PRS models and harmonizing data. The computer system 1201 can regulate various aspects of methods and systems of the present disclosure, such as, for example, perform an algorithm, input training data, analyze data, store data in a repository, or output a result for the user. The computer system 1201 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.
[0146] The computer system 1201 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 1205, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 1201 also includes memory or memory location 1210 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 1215 (e.g., hard disk), communication interface 1220 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1225, such as cache, other memory, data storage and / or electronic display adapters. The memory 1210, storage unit 1215, interface 1220 and peripheral devices 1225 are in communication with the CPU 1205 through a communication bus (solid lines), such as a motherboard. The storage unit 1215 can be a data storage unit (or data repository) for storing data. The computer system 1201 can be operatively coupled to a computer network (“network”) 1230 with the aid of the communication interface 1220. The network 1230 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network1230 in some cases is a telecommunication and / or data network. The network 1230 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 1230, in some cases with the aid of the computer system 1201, can implement a peer-to-peer network, which may enable devices coupled to the computer system 1201 to behave as a client or a server.
[0147] The CPU 1205 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 1210. The instructions can be directed to the CPU 1205, which can subsequently program or otherwise configure the CPU 1205 to implement methods of the present disclosure. Examples of operations performed by the CPU 1205 can include fetch, decode, execute, and writeback.
[0148] The CPU 1205 can be part of a circuit, such as an integrated circuit. One or more other components of the system 1201 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0149] The storage unit 1215 can store files, such as drivers, libraries and saved programs. The storage unit 1215 can store user data, e.g., user preferences and user programs. The computer system 1201 in some cases can include one or more additional data storage units that are external to the computer system 1201, such as located on a remote server that is in communication with the computer system 1201 through an intranet or the Internet.
[0150] The computer system 1201 can communicate with one or more remote computer systems through the network 1230. For instance, the computer system 1201 can communicate with a remote computer system of a user (e.g., a medical professional or patient). Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 1201 via the network 1230.
[0151] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 1201, such as, for example, on the memory 1210 or electronic storage unit 1215. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 1205. In some cases, the code can be retrieved from the storage unit 1215 and stored on the memory 1210 for ready access by the processor 1205. In some situations, the electronic storage unit 1215 can be precluded, and machine-executable instructions are stored on memory 1210.
[0152] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code, or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a precompiled or as-compiled fashion.
[0153] Aspects of the systems and methods provided herein, such as the computer system 1201, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk.“Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0154] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, anyother magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0155] The computer system 1201 can include or be in communication with an electronic display 1235 that comprises a user interface (UI) 1240 for providing, for example, an input of data (e.g., genotype data and / or phenotype data), or an visual output. Examples of UI’s include, without limitation, a graphical user interface (GUI) and web-based user interface.
[0156] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 1205. The algorithm can, for example, generate trained ancestry-adjusted PRS models and harmonize data.EXAMPLES
[0157] The following Examples are provided to illustrate certain aspects of the present invention and to aid those of skill in the art in the art in practicing the invention. These Examples are in no way to be considered to limit the scope of the invention in any manner.
[0158] Example 1: Implementation of the PRS system for generation of Type 2 Diabetes (T2D) and Hypertension ancestry-adjusted PRS models
[0159] Overview
[0160] Using methods and systems of the present disclosure, genotyped data was collected from a novel DTC platform. Adult participants from an international genetic platform were invited to participate. Participants upload their genotype data files and were invited to self-report their health status and metabolic traits. In particular, they were invited to answer general health questionnaires regarding cardiometabolic traits over a period of 6 months.
[0161] The collected data included information corresponding to N= 4,550 (389 cases / 4,161 controls) who reported being affected or previously affected for T2D and N= 4,528 (1,027 cases / 3,501 controls) for hypertension. 164 out of 272 variants were identified showing identical effect direction to previously reported genome-significant findings in Europeans.
[0162] Quality control, imputation, and genome-wide association studies were performed on this dataset. The self-reported health status and metabolic traits information and uploaded genetic information were used in combination to explore their genetic susceptibility and to build polygenic risk scores (PRS) models regarding these traits using the BASIL algorithm.
[0163] In addition, ancestry estimations were generated using Neural ADMIXTURE for all individuals and ancestry-adjusted PRS models for type 2 diabetes and hypertension were generated and analyzed.
[0164] Performance metric of the ancestry-adjusted PRS models was AUC = 0.68, which is comparable to other PRS models obtained with larger datasets including clinical biomarkers.
[0165] Detailed description of T2D and hypertension ancestry-adjusted PRS model development and testing
[0166] Participants were drawn from a research database which offers a DTC genetic traits platform with hundreds of thousands of users globally. After uploading their genetic information, generated in other DTC platforms, users can be informed on their susceptibility to an extensive set of genetic traits. All participants created an account and agreed to a consent on the use of their data and legal agreement. Upon signing up, participants were invited to undertake a health online survey. Participants were redirected to the survey once they gave online consent to be a part of the research.
[0167] The online survey included questions about general conditions such as diabetes, blood pressure, lipid profile, and medication intake. It also included COVID-19, influenza and common cold-related questions along with age, sex, weight, height, and pandemic behavior. Data were collected over a period of six months, from May 01, 2021 to October 06, 2021.
[0168] Only the initial answers of each participant were included in the study, if genotype information was available, and if they answered the age and sex questions. Case-control groups were created following answers provided by participants. For both T2D and hypertension, participants were included as controls if they indicated that they had the condition or were taking medication to treat the condition. Additionally, for T2D, participants were included as cases if they reported having high levels of sugar in their blood work (above 126 ml / dL or 7mmol / L). Participants were defined as controls if they did not report managing health conditions listed in the questionnaire survey, or if they were not managing any health condition. Also, for the T2D cohort specifically, only participants who reported having normal sugar levels (100-125 mg / dL or 5.6 to 6.9 mml / L or Missing) were included as controls. Participants with missing values were excluded.
[0169] Genotype data: quality control, imputation, and GWAS
[0170] This study included seven independent genotyping arrays, comprising a total of 12,424 unrelated individuals.
[0171] Genotype-level data for each array were processed by applying identical quality control and imputation procedures. Briefly, variants with a call rate of < 95% and palindromicmarkers (A / T, G / C, MAF > 0.4) were excluded. An exact test was performed for Hardy- Weinberg equilibrium for individuals of the largest ancestral group (p < 1 x I O12, globally).
[0172] Individual quality control (QC) included genotype call rates > 97%, matching between gender identification and chromosomal sex, and no excess ancestry-adjusted heterozygosity. Samples genetically related to other individuals in the cohort and duplicates were detected and removed, by applying the King algorithm (-make-king, king estimate > 0.177; PLINK 2). Principal component analysis was performed to identify global ancestry per individual using 1000 genomes as reference population with PLINK 2.
[0173] Imputation was carried out with Beagle using data from the 1000 genomes project as a reference panel with Beagle. Next, a merged dataset was generated combining imputed genotypes (MAF > 0.01; imputation quality R2> 0.30) from available datasets. Imputed makers with call rate > 0.95 in the merged data were selected for downstream analysis.
[0174] The GWAS was performed for T2D and hypertension phenotypes (N= 4,550; N= 4,528, respectively) using an additive genetic model with PLINK 2 (-glm). the top ten principal components (PCs), age, sex, and the genotyping array were included as covariables in the model. The results were depicted using the qqman package in R.
[0175] PRS analysis
[0176] The Batch Screening Iterative LASSO (BASIL) algorithm is a meta-algorithm (algorithms that learn from the output of other algorithms), which employs a Lasso algorithm and enhances this output with another layer for faster variable selection in ultra-high-dimensional problems. Similar to the Lasso algorithm, BASIL may be used to find a parameter vector p whose components are the coefficients for the independent variable of the linear regression that approximates the solution of the problem.
[0177] BASIL solves the Lasso solution path in an iterative fashion, starting with a sequence of candidate parameters. From these candidate solutions, each iteration discards the ones that do not meet the requirements to be a suitable solution. The variables that are included in the final set for a viable solution are those that were also screened satisfying a desired threshold requirement, while the others are discarded (i.e., those solutions in which the coefficients in their positions inside the parameter are meant to be 0). This process is repeated until the optimum parameter 20 is found, which is the one that minimizes (20). The BASIL algorithm guarantees to find the exact solution and not only an approximation, via the Karush-Kuhn-Tucker condition (the first derivative necessary conditions for a solution to be optimal) which is verified along each iteration. This condition is necessary and sufficient to prove the exact solution.
[0178] Genetic ancestry
[0179] To address the confounding factor of population stratification in PRS estimations, various approaches may be taken. In some embodiments, a convention was followed of using the first 10 principal components of the PCA to the adjustment of the GWA study. For the correction of PRS models, estimates of global ancestry were used. For this purpose, Neural ADMIXTURE, a faster adaptation of the ADMIXTURE algorithm, was used with similar (or better) clustering results. Utilizing the Python implementation of Neural ADMIXTURE, data from the 1000 Genomes Project Consortium was used for training a model in the supervised mode of Neural ADMIXTURE with the default parameters, the results of global ancestry inference were utilized as a covariate in the training of our ancestry-adjusted PRS models.
[0180] Statistical analysis
[0181] First descriptive statistics were used to understand the characteristics of this cohort, including the geographic distribution of participants, age, sex, and comorbidities. This can become useful in future replication studies. Three ancestry-adjusted PRS models were built using BASIL, a lasso-based linear model: a) one model included genotype features alone; b) another model used only the covariates of age, sex, and the first ten genetic principal components (PCs); c) and a final (full) model used both genotypes and covariates. The inventors used the first 10 PCs to account for residual population micro-stratification as fixed effects. With the three models, their predictive performances were determined, as well as the performance of the full model against the covariate-only model (delta between models).
[0182] The predictive ability of these ancestry-adjusted PRS models was evaluated using the area under the curve (AUC) receiver operating characteristic (ROC) curves using the pROC package in R. To make comparisons between AUC curves from each model, a nonparametric method developed by DeLong et al. was used.
[0183] Genome-wide association results for diabetes type 2 and hypertension
[0184] Genome-wide association data was combined for: 1) 389 T2D cases and 4,161 controls (N= 4,550); and 2) 1,027 hypertension cases and 3,501 controls of European ancestry (N= 4,528). About 8 million variants were tested for T2D and hypertension passing quality control and imputation filters (MAF > 0.01, R2 > 0.3). Both results showed low inflation of test statistics (XGC = 1.01; XGC = 1.01).
[0185] Figure 4 shows genome-wide association results. A type 2 diabetes, and B hypertension. Top boxes show Q-Q plots, while bottom figures show Manhattan plots with two levels of significance ofp< 5 x 10'8(red line), and p < 1x10'6(blue line).
[0186] Nineteen T2D variants displayed significant evidence of replication (p < 0.05) in this dataset. Among them, variants were identified that are closely associated with genes linked totype 2 diabetes susceptibility (e.g., CDKAL1, KCNQ1 as well as variants in the FZO locus linked with both BMI and T2D. Overall, 164 out of 272 variants were identified as showing identical effect direction to genome-significant findings in Europeans. For the hypertension dataset, ten hypertension genetic markers were replicated and 230 out of 365 variants were identified as having identical effect direction.
[0187] The GWAS was validated using independent GWAS meta-analysis datasets from Mahajan et al. 2018 (74,124 T2D cases, 824,006 controls) and Evangelou et al. 2018 (757,201 individuals). The / ?-values and the effect sizes were compared for the variants assessed in both the studies that had identical chromosomal coordinates and alleles with the independent GWAS. The direction of the effect sizes (estimated as OR) was set to match the effect alleles in each study. It was observed that the effect sizes of the genome-wide significant variants in the independent GWAS were concordant in directionality in both our T2D and hypertension GWAS.
[0188] These observations highlight how carefully curated DTC repositories with ever increasing sample sizes and variant diversity can replicate previous findings and hold the potential of delivering enhanced discovery and single-variant resolution of causal T2D and hypertension risk and protection alleles. Additionally, these findings confirm the impact of DTC resources on mechanistic insights and clinical translation efforts.
[0189] Estimation of cardiometabolic PRS models using SNPnet
[0190] PRS models were built for each phenotype using the BASIL algorithm. The predictor variable was binary (presence or absence of diabetes or hypertension) as reported by participants. After tenfold cross-validation, the genotype-only models reported a predictive performance (AUC) of 0.56 for both diabetes and hypertension and increased to 0.68 for the full model (genotype and covariates together). Similarly, when filtering for participants of European ancestry, the genotype-only models reported a predictive performance of 0.57 and 0.53, increasing to 0.69 and 0.66 in the full model, respectively. Tabulated AUC results are shown in Figure 5A.
[0191] The performance of these models was compared using DeLong’s method
[0059] with no statistical significance. The imbalance ratio of both cohorts was an important factor that impacted the accuracy of these models. Additionally, the majority proportion of participants of mostly European ancestry also explains the small differences in performance between both types of models. Figure 5B shows the comparison between AUC curves in both phenotypes.
[0192] Figure 5B illustrates under the curve (AUC). Comparison of receiver operating characteristic (ROC) between two models Full model and European only model. Results for A) type 2 diabetes (T2D) were 0.68 and 0.69, respectively; while for B) hypertension were 0.68 and0.66, respectively. After applying the DeLong method of ROC comparison, the models were not significantly different.
[0193] The European Bioinformatics Institute (EBI) developed the PGS Catalog, which is an open resource of published polygenic scores (including variants, alleles, and weights). Those published PRS were investigated for T2D and hypertension. For those with reported AUC, the number of variants and the number of individuals whose data was used to train the model under various ancestry groups was obtained.
[0194] The ancestry-adjusted PRS models are comparable to other PRS models. Figure 6 shows the comparison in reported AUC for those models in the EBI PGS catalog including models generated by the inventors in accordance with the subject technology. The average AUC between these models was 0.70. The small number of variables used by the ancestry-adjusted PRS models (125 for T2D and 666 for hypertension) makes them comparable to described by Tanigawa et al.
[0043] using the BASIL algorithm. Likewise, the number of individuals whose data was used to train the models is modest in comparison with large academic and clinical databases. Nevertheless, the predictive performance does not seem to be overtly affected by the number of individuals in the study or the number of included variants highlighting that genetic array data from DTC repositories carry immense promise for the development of PRS tools aimed at improving early detection and prevention of T2D and hypertension.
[0195] Figure 6 illustrates comparison of PRS published in the EBI PGS Catalog for T2D and Hypertension. The color of the bubble represents the population ancestry that was included to build the PRS model. The size of the bubble represents the number of variables (variants) that ended up in the model after training. The x-axis shows the number of individuals used to train the model. The y-axis shows the AUC results as reported in the EBI PGS Catalog. The horizontal line shows the average AUC across all models.
[0196] In this study ancestry-adjusted PRS models were generated for T2D and hypertension from a heterogeneous dataset housing a combination of genetic data and self-reported information from a DTC genetics company. Despite a relatively modest predictive ability, these PRS models are able to identify subsets of users at substantially increased risk of presenting T2D or hypertension. This finding is remarkable because it demonstrates that the ever-increasing availability of genetic data from DTC providers, most of it not annotated for traits of clinical relevance, can be leveraged to generate predictive tools able to improve diagnosis and prevention of diseases with genetic determinants.
[0197] The study demonstrated the capability of inverting the model regarding genetic and phenotypic data acquisition in a research study. Individuals participating in the study shared their genetic array information from other DTC providers and were invited to take an online surveyregarding their general health condition. No difference was found in predictive performance between our trained models that included respondents from all inferred ancestries and those models with respondents from European heritage, due to the fact that 86% of our database was of European origin. The genetics of these PRS models for T2D and hypertension are supported by their ability to replicate known variants from publicly available independent GWAS studies.
[0198] Multiple array types were available in our database, and imputation across platforms (up-imputing) was necessary to harmonize these diverse datasets. The fact that individuals can self-report their genomic information may potentially corrupt the file being uploaded into our platform. However, applying appropriate quality control (QC) principles proved to successfully enable traditional GWAS and PRS analyses.
[0199] DTC platforms can offer a wide range of information about personal wellness, ancestry, physical characteristics, and traits. Advances in genomic research may lead the DTC genomics industry to flourish and make accurate yet easy-to-interpret genomic results. Strict privacy policies of many companies may disallow them to share customers’ data without their consent. These platforms can serve as informative repositories giving actionable insights that aid traditional clinical approaches. The approach of subject recruitment for various complex phenotypes via online surveys is opening up multiple avenues to complement conventional research and clinical strategies. DTC platforms also provide convenience along with a wider reach to recruit participants from various locations. They surpass barriers of single-point data collection centers to language restrictions thus allowing the aggregation of data from places with different ancestries and demographics. Democratizing the access to these genetic platforms and prediction tools may boost progress in precision medicine.
[0200] Federated learning approaches can further improve the possibility to increase the power of studies in DTC genomic analysis, and meta-analysis can be done in combination with academic and clinical datasets (including those from large consortiums).
[0201] As shown, the DTC platform and research strategy of the present disclosure are capable of replicating the reported results with a very fast turnaround time. The participation of individual customers in the platform allowed the generation of a rich dataset that enabled the creation of PRS cardiometabolic models. The comparable predictive performance of the ancestry-adjusted PRS models also is a great indication of how the present disclosure can be leveraged to quickly contribute more PRS models to the larger scientific community. Currently, publicly available PRS models for T2D and hypertension have an AUC of 0.7 on average as shown in Figure 6. This is still a low accuracy, and it is even lower when compared to the small difference between the full and genotype-only models. However, even the heterogeneous DTC platform was able to replicate the findings seen in academic and government-funded biobanks.
[0202] T2D and hypertension are multifactorial diseases that are impacted by genetic and environmental determinants, including lifestyle factors like nutrition and exercise habits. Nevertheless, PRS models may have limitations to provide accurate disease predictions, which compels the need to interpret these findings with caution, especially when they come from DTC genetic services. The clinical actionability of PRS models has yet to be determined through pragmatic trials involving real-world data. The present disclosure provides a novel source of information that can shed light on this important issue. Therefore, providing personalized information about T2D and hypertension predisposition is poised to improve early diagnosis and prevention bringing precision medicine at scale for all.
[0203] Example 2: Ancestry-adjusted PRS models
[0204] Figure 7A depicts a correlation plot between two PGS models for gout, PGS001248 and PGS002030, both of which can be found in the polygenic score (PGS) catalog pgscatalog.org. Figure 7B shows a Q-plot between the two models. Figure 7C shows a correlation plot between the two models with ancestry. Ancestry is indicated by shading of the points in the plot as well as by distribution curves per ancestry corresponding to each of the models.
[0205] Figure 8 depicts correlation plots between PGS models for liver enzymes, including correlation plots between PGS000670 and PGS002157 and between PGS000668 and PGS002158, all of which can be found in the PGS catalog. Figure 8 also depicts agreement matrices between the pairs of PGS models.
[0206] Example 3: Ancestry-adjusted PRS models in non-European populations
[0207] Polygenic risk scores (PRS) use genetic information from thousands (or millions) of markers to predict phenotypes such as complex disease risk. Genetic risk prediction can directly improve preventative health care strategies in various ways, such as increased screening for high- risk individuals (e.g., prostate cancer or breast cancer screening), or suggested lifestyle changes (e.g., for individuals at higher risk of developing chronic kidney disease or metabolic disease). While the effectiveness of PRS risk prediction varies by disease, presumably due to differences in the underlying genetic trait architectures, they can accurately predict type I diabetes (T1D), breast cancer (BC) and prostate cancer (PC) risk more accurately than current clinical models can. In addition, PRS outliers for coronary artery disease (CAD) may have a comparable or greater effect than monogenic mutations for familial hypercholesterolemia (FH).
[0208] Currently, the vast majority of PRS are developed using genetic and clinical data from individuals of European ancestry. However, these PRS models have substantially reduced predictive power when applied to individuals from other parts of the world. For example, there is an estimated average of a 2.5-fold lower accuracy in East Asian ancestry individuals and a 4.9-fold lower accuracy in African ancestry individuals, relative to European ancestry individuals. While there are still competing explanations for why PRS model accuracy is not portable across populations (e.g., different causal alleles, different effect sizes for causal alleles, local epistatic interactions, differential imputation accuracy), the problem remains that existing models are not clinically effective for the vast majority of the world’s people.
[0209] The lack of transferability of existing PRS models to non-European populations represents the single most important challenge in the PRS field. While part of this shortfall is due to the relative lack of large GWAS (and associated effect size estimates) in non-European populations, there is also a dearth of computational and statistical methods that have been developed specifically for individuals from diverse and / or admixed populations. In particular, the number of multiracial (i.e., likely admixed) individuals in the US increased from 9 million in 2010 to 33.8 million in 2020, and is likely to grow even more in the coming years, as are the numbers of US individuals with predominantly non-European ancestry. Further, individuals of non-European ancestry may have increased incidence rates of several common diseases, such as prostate cancer, asthma and cardiovascular disease. In this context, methods and systems of the present disclosure are applied toward quantifying and utilizing variation in genetic ancestry to better categorize the genetic contribution to disease risk.
[0210] Recognizing these challenges, the methods and systems of the present disclosure are used to develop and implement accurate PRS models in non-European populations. To do this, methods and systems of the present disclosure are used to dissect and “deconvolute” an individual’s genetic ancestry. Specifically, the methods and systems of the present disclosure use algorithms that can accurately and quickly estimate genetic ancestry across an individual’s chromosomes, and may be implement into both a standalone API and a DNAnexus App. In addition, a dashboard may be used to allow for the organization, manipulation and visualization of ancestry results. “Ancestry deconvolution” or local ancestry inference (LAI) techniques have become an increasingly important tool for better understanding the genetics of complex human diseases. For example, they have helped us understand the prevalence and pathogenicity of variants across population groups, helped with patient stratification for clinical trials, and suggested the importance of ancestry-specific criteria for evaluating renal function from laboratory measurements. Notably, self-reported ancestry is not always consistent with genetic ancestry, and thus is an imperfect proxy for assessing PRS predictive power. For example, in the UK Biobank, of the individuals who self identify as white British (field 21000 = 1001), 7.4% were flagged as genetic outliers indicative of high levels of non-European ancestry (field 22006 = “NA”).
[0211] Existing PRS models are systematically compared using the published effect sizes in the PGS catalog (pgscatalog.org) along with in-house trained models using three different methods to identify the best performing ones for European ancestry individuals. Then these are implemented as a companion to an ancestry API / App. Note that although European ancestry individuals may be a focus, the ancestry estimation component is still important because of the frequency of discordance between self-reported and genetic ancestry. Second, Machine Learning techniques are developed to estimate genetic risk that explicitly utilizes local ancestry estimates in its calculations; these are integrated directly with ancestry algorithms to produce a single, seamless product that simultaneously calculates genetic ancestry and polygenic risk scores for individuals of any background or ancestry. Our perspective here is that genetic ancestry is an important covariate regarding disease risk that must be accounted for. We do not make any assumptions about the reasons (i.e., causality) that it is important, but rather focus on models that are “ancestry-aware” in their genetic risk estimation.
[0212] The ability to incorporate LAI into PRS models depends in large part on the quality of the haplotype reference panel and reference genomes used for ancestry inference. Since existing reference panels suffer from a relative lack of indigenous American representation, individuals of primarily indigenous American ancestry are sequenced, and their genomes are incorporated into our ancestry reference panels. This sequencing utilizes separate funding streams and may utilize collaborations with hospitals and organizations throughout Latin America. They may help ensure that our genetic reference panel more fully represents the full scale of human genetic diversity.
[0213] Current PRS methods may not incorporate LAI, and thus are not suitable for admixed individuals. Our approach directly accounts for both diverse populations and admixture; it therefore solves the problem of lack of transferability of European-optimized PRS by developing the methodology that can accurately estimate genetic risk regardless of a person’s ancestry.
[0214] Our focus on ancestry-adjusted scores for a wide range of diseases is relevant to large potential markets. The medical testing market is more than $2 billion per year, and there are recent mandates for increasing diversity in biomedical studies. Our work to better predict disease risk for diverse populations may help both in identifying individuals who are a priority for preventative testing but also aid in patient stratification of non-European individuals for clinical trials.
[0215] Ancestry inference methods are developed using methods and systems of the present disclosure. Neural ADMIXTURE is a neural network autoencoder which can perform soft- clustering of genomic sequences while inferring its ancestry composition. This neural network adopts the theoretical framework of the widely used ADMIXTURE algorithm, but incorporates recent advances in deep learning to provide faster computational times, more accurate clustering,the capability to estimate the ancestry of new sequences after training with very high speed, and simultaneous prediction of clustering results using multiple cluster numbers. G-Nomix, our second algorithm, provides high-resolution ancestry predictions, where an ancestry label is predicted at each windowed region of the genetic sequence. The method makes use of multiple machine learning classifiers, including logistic regression, support vector machines with a novel string kernel, gradient boosting trees (e.g., XGBoost), conditional random fields, and convolutions, providing a two-stage process, with a multitude of classifiers providing an initial ancestry estimate within windowed regions of the chromosome, and a second classifier refining the initial predictions and correcting potential phasing errors. This approach provides improvements in both accuracy and speed compared with competing methods.
[0216] We have implemented versions of these programs that accept as input either genotype array data, low-pass whole genome sequence data or whole exome data, and with output consisting of global (Neural ADMIXTURE) or local (G-Nomix) ancestry estimates from either a 7 population model (of major continental or sub-continental ancestries) or a more detailed 22 population model with fine-scale regional ancestry results.
[0217] European PRS model validation: PRS models may be developed for European populations, but each new PRS is highly influenced by the characteristics of the discovery dataset. There is limited validation outside the training sets, and a lack of standardized approaches to compare and select the best available PRS. To address this problem, a framework is developed for harmonizing, testing, selecting and validating the best performing PRS for 13 phenotypes of medical interest, including cardiac conditions (atrial fibrillation [AF], CAD, hypertension), metabolic (T1D, T2D, hypothyroidism), gastric (celiac disease [CD], ulcerative colitis [UC], hemorrhoids), inflammatory (gout), and lab values (LDL, HDL and TG measures).
[0218] The following are evaluated: 1) available PRSs downloaded from the PGSCatalog, and 2) PRS trained in-house, using three available PRS software packages (Snpnet, LDpred2 and PRS-CSx) and a set of unrelated European individuals from UKBB (training set, N=285K). Then, to compare results from the different strategies in a harmonized context, a new, nonoverlapping set of European individuals from the UK biobank (testing set, N=95K) is used. Finally, the best performing PRS was validated in an independent cohort of Europeans (N=8,394). The following metrics are compared across models: 1) Area under the curve (AUC) or R2per PRS; 2) Odds ratio estimation stratified by PRS percentile; specifically, the enrichment of cases in the top 5% (or 10%) compared with cases from the middle of the PRS distribution (e.g., 40 - 60 %-tile) is quantified. This approach enables evaluation of all PRS models in a fair and consistent manner (Figure 9).
[0219] Figure 9 depicts a schematic describing an example workflow to evaluate existing European-focused PRS.
[0220] Integration and visualization: Results are graphically plotted using histograms and percentile plots corresponding to disease status, prevalence or odds ratio on the Y-axis, reflecting case enrichment at the upper end of the PRS distribution. Each individual sample is plotted relative to other available samples from the same ancestry; in the case of European-ancestry individuals, we are able to provide accurate results in a large data lake context. Additionally, each individual’s level of risk for a trait / phenotype is indicated on a scale with 5 - 10 categories (Figure 10).
[0221] Figure 10 depicts risk levels provided by PRS models.
[0222] The best-performing PRS model in European individuals are determined for a dozen of phenotypes of high medical interest by comparing predictability (AUC or R2) and enrichment of cases in the highest risk percentiles across different models and strategies in a harmonized context. Further, this enables the generation and integration of an interactive PRS dashboard.
[0223] Using methods and systems of the present disclosure, PRS models applicable to all individuals, regardless of ancestral background, are developed.
[0224] For example, new phenotypic and genomic cohorts for Latino populations are constructed, through a public-private supported research study, the Biobank of the Americas. Paths are established to recruit and clinically characterize individuals from Colombia, Paraguay, Mexico and different sites in the US (e.g. Stanford, Miami); the database contains linked genetic and clinical information from more than 20 thousand individuals with some non-European ancestry. Overall, the range of human diversity sampled is much broader than other studies such as the UK Biobank (Figures 11A-11C). Finally, DNA + EHR data from an additional 50 thousand individuals who self-identify as non-European are genetically characterized.
[0225] Figures 11A-11C depict proportion of individual per ancestry for a) UKBB, and b) Galatea Bio collection; c) Ancestry deconvolution for Galatea Bio collection generated using GB proprietary software.
[0226] Building a better ancestry reference panel: the studies focus on individuals from the Biobank of the Americas with >90% estimated indigenous American ancestry for further sequencing. High-coverage (~30X) whole genome sequence data is generated using suitable protocols in a CLIA-certified Center for Personalized Medicine (CPM).
[0227] After sequencing, suitable pipelines are used for read mapping, duplicate removal, recalibration and variant calling. The final variant calls are then combined with those from publicly available data sets as well as a collection of diverse, high-coverage genomes. Finally,population genetics tools such as Neural ADMIXTURE and UMAP are used to help define population-specific reference panels.
[0228] Genotype -phenotype data sets: Available data sets include:
[0229] UK Biobank (UKBB), a large-scale biomedical database and research resource, containing in-depth genetic and health information from half a million UK participants.
[0230] A DTC company and business partner with linked genotype - phenotype data from thousands of participants.
[0231] China Medical University Hospital (Taiwan), a collaborator with a database of linked genotype - EHR data from roughly 400,000 Taiwanese.
[0232] Rady, Los Angeles and Boston Children’s Hospitals have ancestrally diverse pediatric and family-based cohorts for a range of diseases.
[0233] Global Biobank Meta-Analysis Initiative (GBMI) Consortium, the largest international consortium studying human diseases. The Biobank of the Americas, as part of the GBMI consortium has access to GWAS summary statistics from the other consortium members.
[0234] A biobank is used with DNA samples from more than 500,000 unique individuals, including linked genotype - phenotype data from more than 50 thousand non-European individuals.
[0235] Ancestry inference: To characterize the genetic structure of the studied samples and datasets, both global and local ancestry are estimated using suitable methods. Ancestry clustering is performed using a supervised (k = 7) pre-trained model with Neural ADMIXTURE. LAI is performed using G-nomix software, using pre-phased data. These operations require a reference population panel representing the continental ancestries likely to appear in the test cohort. An extension collection of reference human genomes from around the world is compiled, including worldwide sampling from the 1000 Genomes Project and the Human Genome Diversity Project, more in-depth sampling from South, East and Southeast Asia from the GenomeAsia Project and other sources, and de novo generated reference genomes that more completely represent indigenous American diversity.
[0236] Ancestry-aware PRS methods: Method and systems are developed for accurate phenotype prediction by integrating PRS with individual ancestral background information. Multi -ancestry input data are split into two independent tranches, consisting of discovery (80%) and testing (20%) data sets. First, PRS models are trained for selected traits / diseases using a supervised learning paradigm, using genetic sequence (SNPs), LAI information for each SNP and phenotypic labels as inputs. Such models map input genetic sequences into predicted phenotypes. Multiple methods may be used, including linear techniques utilized by other PRS methods (e.g., Snpnet or PRS-CX), as well as machine learning-based non-linear models, including neuralnetworks or gradient boosting trees. Further, multiple training strategies are adopted, such as training models with samples coming from one unique population group (single-ancestry training), multiple population groups (multi-ancestry training), or from admixed individuals (admixed training), with population information detected with Neural ADMIXTURE and G- Nomix. Finally, trade-offs are characterized between data availability, ancestry transferability, predictive accuracy, and training / inference efficiency.
[0237] By training multiple predictive models, using different algorithms, training strategies and different phenotypes, a collection of PRS models is built, which is extended and compared to publicly available PRS models. In the testing phase, the performance of de novo generated models (i.e., with different algorithms and training strategies) is compared with existing ones that do not incorporate LAI. The following metrics are compared across models: 1) Area under the curve (AUC) or R2per PRS; 2) Odds ratio estimation stratified by PRS percentile - specifically, the enrichment of cases in the top 5% (or 10)% in the PRS distribution compared with individuals from the middle of the distribution (e.g., 40-th to 60-th %-tile).
[0238] A second stage is conducted using machine learning models, taking as input the predictions of the collection of PRS models of a sample for a given phenotype (singlephenotype), or for all available phenotypes (multi-phenotype), and all covariates (including principal components, high-resolution local ancestry inference predictions, and global ancestry predictions). The combined information of predicted PRS scores and predicted ancestry descriptors provides a low-dimensional genetic characterization of the phenotypic and ancestral information for each individual. This low-dimensional descriptor can be seen as a compression of a SNP sequence into a lower-dimensional phenotypic and ancestral-aware representation. Such machine learning “ensembling” models map the low-dimensional PRS+ancestry representations into predicted phenotypes. By incorporating this second “ensembling” stage, trained to predict one or multiple phenotypes, the machine learning model is able to provide additional robustness across ancestral populations, and boost the predictive performance of admixed individuals. Throughout the training of the models, performance is evaluated across multiple population groups, failure cases are analyzed, and explainable machine learning techniques are adopted to obtain better insights into what components are important in order to obtain accurate predictions.
[0239] Integration and visualization: A PRS dashboard is constructed for non-European individuals, applying the same methods previously described. Briefly, each sample is plotted together with samples from the same non-European population, or with similar ancestral proportions (for admixed individuals. Both internal GB samples and reference samples for comparison are used.
[0240] PRS models integrating LAI can increase the performance across 1) different ancestries, and 2) admixed individuals. For example, PRS models can be developed and trained to meet the following criteria: 1) increased AUC of LAI adjusted PRSs with respect to non-LAI adjusted PRS; and 2) increased percentage of cases identified in the tail of the PRS distribution.
[0241] Since the Biobank of the Americas contains samples from more than 500 thousand individuals, many of which have already been screened for genetic ancestry proportions, the sequencing and improved ancestry reference panel creation may be expected to proceed smoothly.
[0242] It will also be recognized by those skilled in the art that, while the invention has been described above in terms of preferred embodiments, it is not limited thereto. Various features and aspects of the above-described invention may be used individually or jointly. Further, although the invention has been described in the context of its implementation in a particular environment, and for particular applications (e.g. for generating and operating ancestry-adjusted PRS models), those skilled in the art will recognize that its usefulness is not limited thereto and that the present invention can be beneficially utilized in any number of environments and implementations where it is desirable to generate and operate PRS models that incorporate, in addition to or alternatively to ancestry, other potentially affecting variables, for example geographic location, lifestyle factors, and group identification of individuals comprising populations included as cohorts for developing PRS models. Accordingly, the claims set forth below should be construed in view of the full breadth and spirit of the invention as disclosed herein.
[0243] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A computer-implemented method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, comprising:(a) processing, using a genetic primary component analysis (PCA) model, genotype data and phenotype data corresponding to one or more populations of individuals to generate genetic primary components (PCs) corresponding to the genotype data and the phenotype data;(b) processing, using a genome-wide association study (GWAS) model, the genotype data and the phenotype data, thereby generating GWAS results, wherein the GWAS model comprises covariables comprising at least one of the genetic PCs;(c) obtaining estimates of global ancestry corresponding to individuals of the one or more populations of individuals; and(d) training a PRS model using at least the GWAS results and the estimates of global ancestry as covariates, thereby generating the trained ancestry-adjusted PRS model.
2. The computer-implemented method of claim 1, wherein (a) further comprises using reference genomic data from a population with known ancestry or ancestry distribution.
3. The computer-implemented method of claim 2, wherein the covariables comprise at least 10 of the genetic PCs.
4. The computer-implemented method of claim 1, wherein (c) further comprises receiving the estimates of global ancestry from a computer system.
5. The computer-implemented method of claim 1, wherein (c) further comprises processing the genotype data using a trained global ancestry model, thereby determining the estimates of global ancestry.
6. The computer-implemented method of claim 5, wherein the trained global ancestry model is trained using reference genomic data from a population with known ancestry or ancestry distribution.
7. The computer-implemented method of claim 6, wherein the trained global ancestry model is trained using an instance of a Neural ADMIXTURE model.
8. The computer-implemented method of claim 5, further comprising training a global ancestry model using reference genomic data from a population with known ancestry or ancestry distribution, thereby generating the trained global ancestry model.
9. The computer-implemented method of claim 8, wherein training the global ancestry model further comprises training an instance of a Neural ADMIXTURE model.
10. The computer-implemented method of claim 1, wherein the genotype data and the phenotype data comprise a merged array.
11. The computer-implemented method of claim 10, further comprising: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a direct-to-consumer (DTC) platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data, thereby harmonizing each of the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate the merged array.
12. The computer-implemented method of claim 11, wherein performing the imputation on the genotype data further comprises using reference genomic data from a population with known ancestry or ancestry distribution.
13. The computer-implemented method of claim 11, further comprising identifying individuals and variants for removal when genotype data and phenotype data corresponding to the individuals and variants fail to meet a QC threshold.
14. The computer-implemented method of claim 1, further comprising storing the ancestry- adjusted PRS model in a repository.
15. A computer-implemented method for harmonizing data corresponding to one or more direct-to-consumer (DTC) platforms, comprising: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a DTC platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data using reference genomic data from a population with known ancestry or ancestry distribution, thereby harmonizing the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate a merged array comprising genotype and phenotype data from one or more populations of individuals.
16. The computer-implemented method of claim 14, further comprising using the merged array to generate a PRS model.
17. A computer system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, upon execution by the one or more computer processors, implements a method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, the method comprising:(a) processing, using a genetic primary component analysis (PCA) model, genotype data and phenotype data corresponding to one or more populations of individuals to generate genetic primary components (PCs) corresponding to the genotype data and the phenotype data;(b) processing, using a genome-wide association study (GWAS) model, the genotype data and the phenotype data, thereby generating GWAS results, wherein the GWAS model comprises covariables comprising at least one of the genetic PCs;(c) obtaining estimates of global ancestry corresponding to individuals of the one or more populations of individuals; and(d) training a PRS model using at least the GWAS results and the estimates of global ancestry as covariates, thereby generating the trained ancestry-adjusted PRS model.
18. The computer system of claim 17, wherein (a) further comprises using reference genomic data from a population with known ancestry or ancestry distribution.
19. The computer system of claim 18, wherein the covariables comprise at least 10 of the genetic PCs.
20. The computer system of claim 17, wherein (c) further comprises receiving the estimates of global ancestry from a computer system.
21. The computer system of claim 17, wherein (c) further comprises processing the genotype data using a trained global ancestry model, thereby determining the estimates of global ancestry.
22. The computer system of claim 21, wherein the trained global ancestry model is trained using reference genomic data from a population with known ancestry or ancestry distribution.
23. The computer system of claim 22, wherein the trained global ancestry model is trained using an instance of a Neural ADMIXTURE model.
24. The computer system of claim 21, wherein the method further comprises training a global ancestry model using reference genomic data from a population with known ancestry or ancestry distribution, thereby generating the trained global ancestry model.
25. The computer system of claim 24, wherein training the global ancestry model further comprises training an instance of a Neural ADMIXTURE model.
26. The computer system of claim 17, wherein the genotype data and the phenotype data comprise a merged array.
27. The computer system of claim 26, wherein the method further comprises: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a direct-to-consumer (DTC) platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays;removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data, thereby harmonizing each of the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate the merged array.
28. The computer system of claim 27, wherein performing the imputation on the genotype data further comprises using reference genomic data from a population with known ancestry or ancestry distribution.
29. The computer system of claim 27, wherein the method further comprises identifying individuals and variants for removal when genotype data and phenotype data corresponding to the individuals and variants fail to meet a QC threshold.
30. The computer system of claim 17, wherein the method further comprises storing the ancestry-adjusted PRS model in a repository.
31. A computer system comprising one or more computer processors and computer memory coupled thereto, wherein the computer memory comprises machine-executable code that, upon execution by the one or more computer processors, implements a method for harmonizing data corresponding to one or more direct-to-consumer (DTC) platforms, the method comprising: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a DTC platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data using reference genomic data from a population with known ancestry or ancestry distribution, thereby harmonizing the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate a merged array comprising genotype and phenotype data from one or more populations of individuals.
32. The computer system of claim 31, wherein the method further comprises using the merged array to generate a PRS model.
33. A non-transitory computer readable medium comprising machine-executable code that, upon execution by one or more computer processors, implements a method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, the method comprising:(a) processing, using a genetic primary component analysis (PCA) model, genotype data and phenotype data corresponding to one or more populations of individuals to generate genetic primary components (PCs) corresponding to the genotype data and the phenotype data;(b) processing, using a genome-wide association study (GWAS) model, the genotype data and the phenotype data, thereby generating GWAS results, wherein the GWAS model comprises covariables comprising at least one of the genetic PCs;(c) obtaining estimates of global ancestry corresponding to individuals of the one or more populations of individuals; and(d) training a PRS model using at least the GWAS results and the estimates of global ancestry as covariates, thereby generating the trained ancestry-adjusted PRS model.
34. The non-transitory computer readable medium of claim 33, wherein (a) further comprises using reference genomic data from a population with known ancestry or ancestry distribution.
35. The non-transitory computer readable medium of claim 34, wherein the covariables comprise at least 10 of the genetic PCs.
36. The non-transitory computer readable medium of claim 33, wherein (c) further comprises receiving the estimates of global ancestry from a computer system.
37. The non-transitory computer readable medium of claim 33, wherein (c) further comprises processing the genotype data using a trained global ancestry model, thereby determining the estimates of global ancestry.
38. The non-transitory computer readable medium of claim 37, wherein the trained global ancestry model is trained using reference genomic data from a population with known ancestry or ancestry distribution.
39. The non-transitory computer readable medium of claim 38, wherein the trained global ancestry model is trained using an instance of a Neural ADMIXTURE model.
40. The non-transitory computer readable medium of claim 37, wherein the method further comprises training a global ancestry model using reference genomic data from a population with known ancestry or ancestry distribution, thereby generating the trained global ancestry model.
41. The non-transitory computer readable medium of claim 40, wherein training the global ancestry model further comprises training an instance of a Neural ADMIXTURE model.
42. The non-transitory computer readable medium of claim 33, wherein the genotype data and the phenotype data comprise a merged array.
43. The non-transitory computer readable medium of claim 42, wherein the method further comprises: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a direct-to-consumer (DTC) platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data, thereby harmonizing each of the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate the merged array.
44. The non-transitory computer readable medium of claim 43, wherein performing the imputation on the genotype data further comprises using reference genomic data from a population with known ancestry or ancestry distribution.
45. The non-transitory computer readable medium of claim 43, wherein the method further comprises identifying individuals and variants for removal when genotype data and phenotype data corresponding to the individuals and variants fail to meet a QC threshold.
46. The non-transitory computer readable medium of claim 33, wherein the method further comprises storing the ancestry-adjusted PRS model in a repository.
47. A non-transitory computer readable medium for harmonizing data corresponding to one or more direct-to-consumer (DTC) platforms, comprising, with one or more processors: receiving one or more arrays of genotype data and phenotype data, each of the one or more arrays corresponding to a population of individuals comprising a DTC platform cohort, each of the one or more arrays comprising a plurality of variants; performing one or more quality control (QC) operations on the genotype data and the phenotype data to identify individuals and variants for removal from the one or more arrays; removing, from the one or more arrays, genotype data and phenotype data corresponding to the identified individuals and variants; performing imputation on the genotype data using reference genomic data from a population with known ancestry or ancestry distribution, thereby harmonizing the one or more arrays, wherein each of the one or more harmonized arrays comprises the same plurality of variants as each other of the one or more arrays; and combining the one or more harmonized arrays to generate a merged array comprising genotype and phenotype data from one or more populations of individuals.
48. The non-transitory computer readable medium of claim 47, wherein the method further comprises using the merged array to generate a PRS model.
49. A computer-implemented method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, comprising:(a) processing genotype data and phenotype data corresponding to populations of individuals to generate genetic primary components (PCs) and genome-wide association study (GWAS) results using covariables comprising at least one of the genetic PCs; and(b) training a PRS model using at least the GWAS results and estimates of global ancestry corresponding to individuals of the populations as covariates, thereby generating the trained ancestry-adjusted PRS model.
50. A computer-implemented method for generating a trained ancestry-adjusted polygenic risk score (PRS) model, comprising generating genome-wide association study (GWAS) results from populations of individuals, and training a PRS model using at least the GWAS results and estimates of global ancestry corresponding to individuals of the populations.
51. A computer-implemented method comprising processing test genotype data of a test subject with an ancestry-adjusted polygenic risk score (PRS) model to determine an ancestry- adjusted polygenic risk score of the test subject.
52. The computer-implemented method of claim 51, wherein the ancestry-adjusted PRS model is obtained at least in part by:(a) processing, using a genetic primary component analysis (PCA) model, genotype data and phenotype data corresponding to one or more populations of individuals to generate genetic primary components (PCs) corresponding to the genotype data and the phenotype data;(b) processing, using a genome-wide association study (GWAS) model, the genotype data and the phenotype data, thereby generating GWAS results, wherein the GWAS model comprises covariables comprising at least one of the genetic PCs;(c) obtaining estimates of global ancestry corresponding to individuals of the one or more populations of individuals; and(d) training a PRS model using at least the GWAS results and the estimates of global ancestry as covariates, thereby generating the trained ancestry-adjusted PRS model.