System and method for incorporating quality control flags for genetic analysis
The hypometric genetics method transforms BLQ flags into binary traits for genetic analysis, addressing the underutilization of quality control flags to enhance genetic discovery and prediction, particularly for rare variants and extreme phenotypes.
Patent Information
- Application Number
- PCT/US2025/050503
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-11
- Filing Date
- 2025-10-10
- Publication Date
- 2026-04-16
AI Technical Summary
Traditional genetic analysis methods struggle to effectively utilize quality control flags, such as Below Limit of Quantification (BLQ) flags, leading to reduced statistical power in detecting associations with rare variants and overlooking important genetic information, particularly in large-scale genetic and metabolomic studies.
The hypometric genetics (hMG) approach transforms BLQ flags into binary traits for genetic analysis, incorporating them into joint multi-trait analysis with corresponding continuous traits to enhance genetic discovery and prediction.
This approach significantly improves the power of genetic discovery and prediction by leveraging previously underutilized data, particularly for rare variants and extreme phenotypes, enhancing the identification of genes and genetic variants associated with various traits and diseases.
Smart Images

Figure US2025050503_16042026_PF_FP_ABST
Abstract
Description
Our Ref. MIT.1010PCSystem and Method for Incorporating Quality Control Flags for Genetic AnalysisFEDERALLY SPONSORED RESEARCH AN D DEVELOPMENT
[0001] This invention was made with government support under AG067151, NS110453, AG074003, MH119509, NS129032, AG081017, MH109978, HG008155, AG062335, AG058002, AG062377, DA053631, NS115064, AG054012, and AG077227 awarded by the National Institutes of Health. The government has certain rights in the invention.BACKGROU N D
[0002] Genetic research has long focused on understanding the relationship between genetic variants and phenotypic traits. Traditional genome-wide association studies (GWAS) have been instrumental in identifying common genetic variants associated with various traits and diseases.However, these studies often face challenges in detecting associations with rare variants due to limited statistical power. Additionally, the presence of measurement errors and data points marked with quality control (Q.C) flags (e.g., below the limit of quantification, BLQ) can further complicate genetic analyses. These issues highlight the need for innovative approaches that can enhance the power of genetic discovery and provide more accurate predictions of phenotypic traits.
[0003] In recent years, the integration of omics data (e.g., metabolomics, proteomics, and transcriptomics data) with genetic studies has opened new avenues for understanding the genetic basis of complex traits. Metabolomics, the comprehensive study of metabolites in biological systems, provides a direct link between genetic variation and phenotypic expression. By analyzing metabolite levels, researchers can gain insights into the biochemical pathways influenced by genetic variants.However, the presence of data points marked with QC flags in metabolomics studies poses a significant challenge, as these data points can obscure accurate measurement values reflecting true underlying biology and reduce the overall statistical power of the genetic analysis.
[0004] Current approaches to genetic analysis typically focus on individuals sampled from populations , potentially overlooking important information at the extremes of trait distributions. While some methods have been developed to address this issue, such as extreme phenotype sampling, theyOur Ref. MIT.1010PC often require deliberate selection of samples with extreme values, which can be resource-intensive and may not fully capture the genetic basis of trait extremes.
[0005] Furthermore, existing methods for genetic analysis often struggle to effectively utilize quality control indicators and flags that are routinely generated during data collection processes. These indicators, such as those marking measurements below the limit of quantification, are frequently treated as nuisances to be removed or adjusted for, rather than potential sources of valuable genetic information.
[0006] In the field of metabolomics, particularly in large-scale studies using techniques like nuclear magnetic resonance (NMR) spectroscopy, a wealth of data is generated, including various quality control flags. However, current analytical approaches often fail to fully leverage this information, potentially missing important genetic associations and limiting the power of genetic discovery and prediction.
[0007] The limitations of existing methods become particularly apparent when dealing with rare variants, which are often of great interest in identifying causal genes and potential drug targets. Traditional approaches may lack the statistical power to detect associations with rare variants.
[0008] There remains an unmet need for innovative methods that can effectively incorporate all available information from large-scale genetic and metabolomic studies, including quality control indicators, to enhance genetic discovery and prediction. Such methods should be capable of improving statistical power for rare variant analysis, leveraging extreme phenotypic information, and providing more comprehensive insights into the genetic basis of complex traits and diseases.SUMMARY
[0009] An approach, referred to herein as the "hypometric genetics (hMG) approach," introduces a novel method for improving genetic discovery and prediction by leveraging previously underutilized data in large-scale genetic studies. The key innovation of hMG lies in its utilization of Below Limit of Quantification (BLQ) flags, and other quality control (QC) flags, as a valuable source of genetic information, transforming what was once considered noise into a powerful tool for genetic analysis. Although the term "hypometric," as used herein, refers to the use of genetic data that falls below standard measurement thresholds or limits of quantification, embodiments of the present invention are applicable more generally to data that is flagged with any quality control flag, whether or not that data falls below any particular measurement threshold or limit of quantification.
[0010] The hMG approach includes both a genetic discovery component and a genetic prediction model. The genetic discovery aspect focuses on identifying genes and genetic variants associated withOur Ref. MIT.1010PC various traits, including disease outcomes, phenotypes, biomarkers, and medically relevant measurements. By extracting relevant information from QC flags (e.g., converting BLQ fields into binary indicator variables) and performing genetic analysis (e.g., joint multi-trait analysis with corresponding continuous traits), hMG takes advantage of additional sources of genetic signals for significantly enhancing the power of genetic discovery.
[0011] The hMG approach also develops machine learning models for individual-level trait prediction. By generating polygenic risk scores using QC-derived traits, the method has shown the ability to predict original continuous traits with high accuracy. Moreover, polygenic risk scores trained on data points below BLQ showed improved predictive performance compared to models trained only on data points that are not marked with BLQ flags. In some cases, polygenic score models trained on BLQ traits alone may predict the corresponding quantitative traits with about 91% accuracy compared to models trained on the quantitative traits themselves. This innovative approach to genetic prediction leverages information from quality control flags that were previously overlooked, potentially leading to more comprehensive and accurate predictive models.
[0012] These two aspects of the invention are closely interrelated, with genetic discoveries informing the development of prediction models, and the performance of prediction models validating genetic basis and refining genetic discoveries. The hMG method's unique approach to incorporating QC data has shown particular promise in identifying metabolite-lowering associations for rare variants.
[0013] By addressing the limitations of current genetic analysis methods and harnessing the power of previously underutilized data, the hMG approach offers a significant advancement in the field of genetic research, with potential applications in drug target identification, pharmacokinetics studies, and personalized medicine.
[0014] Other features and advantages of various aspects and embodiments of the present invention will become apparent from the following description and from the claims.BRI EF DESCRI PTION OF TH E DRAWI NGS
[0015] FIG. 1 illustrate a system for identifying and prioritizing genes according to one embodiment of the present invention.
[0016] FIG. 2 illustrates a flowchart of a method that may be performed by embodiments of the present invention to generate / identify the QC-derived phenotypes shown in FIG. 1 based on the original phenotype shown in FIG. 1.Our Ref. MIT.1010PC
[0017] FIG. 3 illustrates a genetic analysis module, which may be used to implement any of the genetic analysis modules shown in FIG. 1.
[0018] FIG. 4 illustrates an embodiment of the common and rare genetic association module from FIG. 3.
[0019] FIG. 5 illustrates an embodiment of the polygenic score modeling module from FIG. 3.
[0020] FIG. 6 illustrates an embodiment of the gene prioritization module from FIG. 3.
[0021] FIG. 7A illustrates a comparison of genetic effect size estimates between a truncated quantitative trait and a corresponding binarized below-quantification-limit (BLQ) trait, demonstrating correlation of genetic associations according to one embodiment of the present invention.
[0022] FIG. 7B illustrates a comparison of statistical significance of genetic associations between a truncated quantitative trait and a corresponding binarized BLQ trait according to one embodiment of the present invention.
[0023] FIG. 7C illustrates predictive performance of polygenic score models trained on original traits (including BLQ measurements), truncated traits (excluding BLQ measurements), and binarized BLQ traits across multiple population groups according to one embodiment of the present invention.
[0024] FIG. 8A illustrates common variant associations (left panel) and rare variant associations (right panel) for a binarized below-quantification-limit (BLQ) trait according to one embodiment of the present invention.
[0025] FIG. 8B illustrates common variant associations (left panel) and rare variant associations (right panel) for a corresponding quantitative trait excluding BLQ measurements according to one embodiment of the present invention.
[0026] FIG. 8C illustrates polygenic score prediction performance for a binarized BLQ trait, showing odds ratios across polygenic score deciles according to one embodiment of the present invention.
[0027] FIG. 8D illustrates polygenic score prediction performance for a quantitative trait, showing covariate-adjusted phenotype values across polygenic score percentile bins according to one embodiment of the present invention.
[0028] FIG. 8E illustrates gene prioritization results from rare variant aggregation testing for a binarized BLQ trait, showing strength of statistical evidence for gene-phenotype associations according to one embodiment of the present invention.
[0029] FIG. 8F illustrates gene prioritization results from rare variant aggregation testing for a corresponding quantitative trait, showing strength of statistical evidence for gene-phenotype associations according to one embodiment of the present invention.
[0030] FIG. 9A illustrates a comparison of gene-phenotype association evidence (loglO Bayes Factor) between single-trait analysis and multi-trait analysis, demonstrating improved statistical powerOur Ref. MIT.1010PC through joint analysis of quality control flags and quantitative measurements according to one embodiment of the present invention.
[0031] FIG. 9B illustrates an expanded view of FIG. 9A, highlighting genes achieving prioritization only through multi-trait analysis according to one embodiment of the present invention.
[0032] FIG. 9C illustrates the number of prioritized genes across five metabolomics traits, comparing single-trait analysis versus multi-trait hypometric genetics analysis for non-synonymous and protein-truncating variants, demonstrating increased gene discovery according to one embodiment of the present invention.DETAI LED DESCRI PTION
[0033] FIG. 1 illustrates a system 100 for identifying and prioritizing genes according to one embodiment of the present invention. The system of FIG. 1 includes:• Original Phenotype 102: This represents the primary trait data collected in one or more omics studies, such as metabolomics, transcriptomics, proteomics, genomics, epigenomics, phenomics, lipidomics, glycomics, microbiomics, or exposomics data. The original phenotype data 102 may represent any one or more of these data types individually or in combination.• Phenotype(s) from Quality Control (QC) Flags 104: These phenotypes 104 are derived from one or more quality control (QC) indicators, including but not limited to BLQ flags, as described in more detail below in connection with FIG. 2 and elsewhere. The phenotypes derived from QC flags 104 enable the extraction of valuable genetic information from data points that are typically discarded or ignored in traditional analyses.• Genetic Analysis (FIG. 3): The genetic analysis component, shown and described in more detail below in connection with FIG. 3, may be applied to any one or more of the following, in any combination: a. The original phenotype data 102; b. The phenotypes derived from quality control flags 104; or c. A combination of both the original phenotype 102 and the QC flag-derived phenotypes 104.
[0034] The arrows in FIG. 1 indicate that the genetic analysis of FIG. 3 may be performed on each data type independently or in combination. This flexibility allows researchers to leverage quality control information in various ways, potentially uncovering genetic associations that may be missed by traditional methods focusing solely on original phenotype data.Our Ref. MIT.1010PC
[0035] By enabling the analysis of quality control flag-derived phenotypes 104 alongside or independent of original phenotype data 102, the system 100 of FIG. 1 facilitates a more comprehensive understanding of genetic influences on traits. This novel approach extracts valuable insights from previously underutilized data, potentially leading to enhanced genetic discovery and prediction capabilities.
[0036] Furthermore, as illustrated in FIG. 1, the genetic analysis component (FIG. 3) may not only analyze the original phenotype 102 and QC flag-derived phenotypes (104) independently, but may also utilize the outputs of these analyses as inputs for further genetic analysis. Specifically, the genetic analysis component (FIG. 3) may perform additional genetic analysis on the output of the genetic analysis performed on the original phenotype 102 and the output of the genetic analysis performed on the QC flag-derived phenotypes 104. This iterative approach allows for a more comprehensive examination of genetic associations by leveraging information from both the original phenotypes 102 and the quality control flag-derived phenotypes 104 in various combinations. This multi-layered analytical approach enhances the method's ability to identify rare variants, extreme phenotypes, and subtle genetic associations that might be missed by traditional genetic analysis methods focusing solely on original phenotype data.
[0037] FIG. 2 illustrates a flowchart of a method 200 that may be performed by embodiments of the present invention to generate / identify the QC-derived phenotypes 104 shown in FIG. 1 based on the original phenotype 102 shown in FIG. 1. This method 200 transforms quality control flags associated with the original phenotype data into derived phenotypes that can be used for enhanced genetic analysis.
[0038] The method 200 includes a quality control flag identification step 202, which receives original source data (such as the original phenotype 102 and optionally additional QC information) as input and generates one or more quality control (QC) flags 206 as output. This step is useful identifying measurements that may be unreliable or fall outside of standard quantification ranges. The QC flags 206 may be of any of the kinds disclosed herein, such as, but not limited to:• Below Limit of Quantification (BLQ) flags: These indicate that the concentration of a molecule of interest falls below the calibrated range of quantification for the measurement device.• Below Limit of Detection (BLD) flags: These indicate that the concentration of a molecule of interest falls below the detectable range of the measurement device, representing measurements where the analyte cannot be reliably distinguished from background noise or baseline signal.Our Ref. MIT.1010PC• High ethanol flags: These may indicate potential sample contamination or degradation.• Degraded sample flags: These suggest that the sample quality may have been compromised.• Polysaccharide flags: These may indicate the presence of interfering substances in the sample.• Unknown contamination flags: These suggest the presence of unidentified contaminants that may affect measurement accuracy.
[0039] The QC flag identification step 202 may identify the QC flags 206 using any of a variety of methods, such as any one or more of the following: comparing measured values to predefined thresholds for each type of measurement; analyzing signal-to-noise ratios in spectroscopic data to identify measurements below quantification limits; analyzing signal intensity relative to background noise levels to identify measurements below detection limits; applying automated quality control algorithms that assess multiple parameters of each measurement; or evaluating metadata associated with each sample, such as processing time or storage conditions, to identify potential quality issues. The QC flag identification step 202 may apply sample-level quality control criteria, such as removing samples with missingness greater than about 1%, excluding samples that fail Hardy-Weinberg equilibrium tests with p-values less than about 1.0x10 -7 for directly genotyped variants or less than about 1.0x10 -I for whole-exome sequencing variants, and filtering samples based on heterozygosity or missing rate outliers. The step 202 may also apply variant-level quality control procedures, including filtering variants with imputation quality scores (INFO scores) less than about 0.3, excluding variants with minor allele frequencies below about 0.01%, and removing variants that do not meet population-specific quality thresholds.
[0040] In addition to or instead of metabolomics applications, the hMG method may be applied to transcriptomics data involving levels of gene expression values. In transcriptomics studies, quality control flags may be generated to identify measurements that fall outside standard quantification ranges or exhibit other quality issues. For gene expression measurements, quality control flags may include, but are not limited to:• Below detection threshold flags: These indicate that the expression level of a gene falls below the detectable range of the measurement platform, such as RNA sequencing or microarray technologies.• Low read count flags: These may indicate genes with insufficient sequencing coverage for reliable quantification.• High variance flags: These suggest that gene expression measurements show unusually high variability across technical replicates.Our Ref. MIT.1010PC• Batch effect flags: These may indicate potential technical artifacts affecting gene expression measurements across different processing batches.• RNA degradation flags: These suggest that the RNA sample quality may have been compromised, affecting the accuracy of gene expression measurements.
[0041] The QC flag identification step 202 may identify quality control flags for gene expression data using methods such as: analyzing read count distributions to identify genes with expression levels below detection thresholds; evaluating technical replicate consistency to identify measurements with high variance; assessing RNA integrity scores and other sample quality metrics; or applying automated quality control pipelines that evaluate multiple parameters of gene expression measurements. By converting these transcriptomics-specific quality control flags into binary traits, the hMG method may extract valuable genetic information from gene expression data points that are typically excluded from traditional analyses.
[0042] Similarly, the hMG method may be applied to proteomics data involving levels of protein expression values. In proteomics studies, quality control flags may be generated to identify protein measurements that exhibit quality issues or fall outside standard quantification ranges. For protein expression measurements, quality control flags may include, but are not limited to:• Below limit of quantification (BLQ) flags: These indicate that the concentration of a protein falls below the calibrated range of quantification for the measurement platform, such as mass spectrometry or immunoassay-based methods.• Low signal-to-noise ratio flags: These may indicate protein measurements with insufficient signal strength for reliable quantification.• Missing peptide flags: These suggest that key peptides required for protein identification and quantification were not detected.• Protein degradation flags: These may indicate that protein samples show signs of degradation that could affect measurement accuracy.• Cross-reactivity flags: These suggest potential interference from other proteins or compounds that may affect the specificity of protein measurements.
[0043] The QC flag identification step 202 may identify quality control flags for protein expression data using methods such as: comparing measured protein concentrations to predefined detection limits for each measurement platform; analyzing peptide coverage and identification confidence scores; evaluating sample preparation quality metrics; or applying automated quality control algorithms that assess multiple parameters of protein measurements. By incorporating these proteomics-specific qualityOur Ref. MIT.1010PC control flags into the hMG analysis framework, the method may leverage previously underutilized information from protein expression studies to enhance genetic discovery and prediction capabilities.
[0044] For both transcriptomics and proteomics applications, the trait conversion step 208 may generate binary traits indicating the presence or absence of quality control flags, similar to the approach used for metabolomics data. These binary traits may capture information about individuals with extreme gene expression or protein expression phenotypes, potentially improving the statistical power for identifying genetic variants associated with expression quantitative trait loci (eQTLs) or protein quantitative trait loci (pQTLs). The joint multi-trait analysis capabilities of the hMG method may be particularly valuable for analyzing expression data, as it allows for the simultaneous consideration of both the continuous expression measurements and the binary traits derived from quality control flags.
[0045] By identifying these QC flags, step 202 provides crucial information for subsequent steps in the method 200, enabling the generation of derived phenotypes that incorporate valuable information from previously underutilized quality control indicators. In some cases, the QC flags may be associated with traits that exhibit significant heritability, such as observed-scale heritability values of greater than about 5%, including heritability values of about 5.3% to about 7.1%.
[0046] As used herein, the term "deriving" in the context of quality control flags may include extracting existing quality control information from data measurements. Other examples of "deriving" include generating, calculating, determining, or creating new quality control flags based on analysis of the data measurements. Deriving may involve identifying pre-existing quality control indicators within the data, computing quality control metrics from measurement parameters, applying algorithms to assess data quality, or any combination of these approaches. The deriving process may utilize raw measurement data, processed measurement data, metadata associated with measurements, or any other information available from the data source.
[0047] The method 200 also includes a trait conversion step 208, which receives the QC flags 206 as input and generates one or more traits 210 as output. Example types of traits include, but limited to: binary traits, continuous traits, censored variables, and other compound traits. This step is useful for transforming quality control indicators into a format suitable for genetic analysis. The traits 210 generated by step 208 may include, but are not limited to one or more:• BLQ indicators: A binary trait where 1 indicates the presence of a Below Limit of Quantification flag and 0 indicates its absence for a specific metabolite measurement.• Sample quality indicators: A binary trait where 1 indicates the presence of any quality control flag (e.g., high ethanol, degraded sample) and 0 indicates the absence of such flags.Our Ref. MIT.1010PC• Measurement reliability indicators: A binary trait where 1 indicates that all measurements for an individual across multiple time points have QC flags, and 0 indicates otherwise.• Measurement reliability score: A confidence score ranging from 0 to 1, indicating the strength of the belief of the measurement values, with 1.0 being the most confident, 0 being not confident, and the intermediate values reflecting intermediate strengths of beliefs.
[0048] As used herein, the term "auxiliary traits" refers to an umbrella term encompassing various types of traits derived from quality control information, including QC-derived traits, QC flag-derived traits, and binary traits. Auxiliary traits serve as an intermediate processing step between quality control flags and quality control-derived phenotypes, transforming raw quality control information into various trait formats suitable for genetic analysis. These auxiliary traits are supportive and complementary to the original phenotype data rather than being primary traits of interest, and may provide valuable genetic information that enhances the overall genetic analysis capabilities of the system.
[0049] The trait conversion step 208 may generate the traits in any of a variety of ways. For example, for each of the QC flags 206, the trait conversion step 208 may create a corresponding binary trait where 1 indicates the presence of the flag and 0 indicates its absence. As another example, for individuals with repeated measurements, the trait conversion step 208 may assign 1 to the binary trait if the QC flags 206 are present in all measurements, and 0 otherwise. As yet another example, for individuals with repeated measurements, the trait conversion step 208 may define traits based on whether an amount (e.g., a number or proportion) of QC flags 206 for a given measurement exceeds a predetermined threshold. As yet another example, for individuals with repeated measurements, the trait conversion step 208 may define traits based on whether an amount (e.g., a number or proportion) of QC flags 206 for a given measurement falls below a predetermined threshold. As yet another example, the trait conversion step 208 may define traits reflecting an amount (e.g., a number or proportion) of QC flags 206 for a given measurement. In some cases, the trait conversion step 208 may generate additional binarized traits at specific threshold values, such as at about 0.015 mmol / L, about 0.020 mmol / L, about 0.025 mmol / L, and about 0.030 mmol / L, which may correspond to different multiples of a limit of quantification threshold value.
[0050] The method 200 also includes a trait selection step 212, which receives the traits 210 as input, selects a subset of the traits 210, and outputs the selected traits as selected traits 214. This step is useful for focusing the analysis on the most informative traits derived from quality control flags.
[0051] The trait selection step 212 may select the selected traits 214 using any of a variety of techniques. For example, the trait selection step 212 may select traits that have at least a minimumOur Ref. MIT.1010PC percentage of measurements tagged with the corresponding quality control flag. In some cases, the minimum percentage may be about 1%, such that traits with at least 1% of measurements tagged with quality control flags are selected for further analysis.
[0052] As another example, the trait selection step 212 may select traits that have no more than a maximum percentage (e.g., 10%, 5%, or 2%) of measurements tagged with the corresponding quality control flag. This approach may focus the analysis on high-quality measurements with minimal quality control issues, may help avoid potential batch effects or systematic measurement problems that could confound genetic analysis, may identify traits suitable for use as baseline comparisons or control phenotypes in multi-trait analyses, may reduce noise in the genetic analysis by excluding traits with excessive quality control flags that might indicate widespread measurement unreliability, and may enable identification of traits with consistent measurement quality across different sample processing batches or time points.
[0053] As yet another example, the trait selection step 212 may apply LD score regression to estimate the heritability of each trait and select those with observed-scale heritability above a certain threshold (e.g., 5%). LD score regression (LDSC) may be used to estimate SNP-based heritability from GWAS summary statistics by leveraging the relationship between linkage disequilibrium and the variance explained by genetic variants. In some cases, traits may be selected that exhibit at least about 5% observed-scale heritability, such as traits showing heritability values of about 5.3%, about 6.0%, about 6.6%, about 7.0%, or about 7.1%. As yet another example, the trait selection step 212 may select traits that show lower levels of phenotypic correlation with other traits or with the original continuous traits. As another example, the trait selection step 212 may prioritize traits associated with metabolites or measurements known to be relevant to specific biological pathways or diseases of interest. Yet another example is that the trait selection step 212 may use a multi-trait analysis, which involves selecting sets of related traits that could benefit from joint multi-trait analysis, such as traits derived from the same metabolite measurements and their corresponding continuous traits.
[0054] The method 200 includes an optional covariate adjustment step 216, which receives the selected traits 214 as input and performs covariate adjustment to produce adjusted traits 218 as output. This step is useful for minimizing the influence of technical factors on the genetic analysis of the traits derived from quality control flags.
[0055] The covariate adjustment step 216 may perform covariate adjustment on the selected traits 214 to produce the adjusted traits 218 using any of a variety of techniques. For example, the covariate adjustment step 216 may adjust for technical factors that may influence the presence of quality controlOur Ref. MIT.1010PC flags, such as batch effects, sample processing time, or storage conditions. The covariate adjustment step 216 may account for demographic factors like age, sex, and ancestry that may affect the likelihood of observing certain quality control flags. The covariate adjustment step 216 may incorporate genotype principal components to account for population structure and potential confounding factors in the genetic analysis. The covariate adjustment step 216 may apply sample-level quality control criteria, including removing samples used to compute principal components, excluding samples with sex mismatches between genotype and phenotype data, filtering out samples reported as outliers for heterozygosity or missing rate, removing samples with sex chromosome aneuploidy, and excluding samples with ten or more third-degree relatives. The covariate adjustment step 216 may apply one or more regression models to remove the effects of covariates on the traits, similar to the adjustment performed on corresponding continuous traits. The covariate adjustment step 216 may use mixed- effects models to account for both fixed and random effects that may influence the traits derived from quality control flags.
[0056] The method 200 includes a phenotype generation step 220, which receives the adjusted traits 218 as input and generates the QC-derived phenotypes 104 as output. This step is useful for transforming the adjusted traits 218 into phenotypes 104 that can be used for enhanced genetic analysis according to embodiments of the present invention. The adjusted traits 218 may also be referred to as auxiliary traits, as they serve as intermediate processing elements that support the generation of the final QC-derived phenotypes 104.
[0057] The phenotype generation step 220 may generate the phenotypes 104 based on the adjusted traits 218 through various methods, such as any one or more of the following:• Direct conversion: Using the adjusted traits as new phenotypes, representing the presence or absence of specific quality control flags.• Aggregation: Combining multiple adjusted traits related to the same metabolite or measurement to create composite phenotypes.• Joint representation: Creating phenotypes that jointly represent both the trait and the corresponding continuous trait, allowing for multi-trait analysis.• Extreme phenotype identification: Generating phenotypes that specifically capture individuals with extremely low metabolite concentrations, as indicated by the BLQ flags.• Derived trait generation: Creating new quantitative traits based on the patterns of quality control flags across multiple measurements or metabolites.Our Ref. MIT.1010PC
[0058] Note that the phenotypes 104 generated in this step are the same phenotypes shown in the system 100 of FIG. 1 and are used by embodiments of the present invention in ways disclosed elsewhere herein. These QC-derived phenotypes 104 serve as valuable inputs for subsequent genetic analyses, including genome-wide association studies, polygenic risk score modeling, and rare variant aggregation tests. By incorporating information from quality control flags, these phenotypes enable the hypometric genetics (hMG) approach to leverage previously underutilized data, potentially improving the statistical power for discovering trait-associated genes and enhancing genetic prediction capabilities.
[0059] FIG. 3 illustrates a genetic analysis module 300, which may be used to implement any of the genetic analysis modules shown in FIG. 1. The genetic analysis module 300 may perform genetic analysis using any of a variety of techniques, such as any one or more of the following, each of which may be performed once or repeated any number of times:• Common and Rare Genetic Association 302: This component, detailed further in FIG. 4, focuses on identifying genetic associations across both common and rare variants.• Polygenic Score Modeling 304: Expanded upon in FIG. 5, this technique involves the development of machine learning models for individual-level trait prediction.• Gene Prioritization 306: Further elaborated in FIG. 6, this method aims to identify and prioritize genes associated with various traits.
[0060] FIG. 4 illustrates an embodiment of the common and rare genetic association module 302 from FIG. 3. This module is designed to identify common and rare genetic variants associated with the phenotype 408 by leveraging various types of input data. The common and rare genetic association module 302 may utilize any one or more of the following inputs to identify genetic variants:• Phenotypic Information 402: This includes the original continuous traits as well as the traits derived from quality control flags. The incorporation of binary traits from BLQ flags allows for the analysis of extreme phenotypes and potentially improves the power to detect rare variant associations.• Genetic Information 404: This encompasses both common and rare genetic variants, including directly genotyped and imputed variants, as well as whole-exome or whole-genome sequencing data. The module 302 may consider variants with different minor allele frequencies (MAF), focusing on both common (MAF > 1%) and rare (MAF < 1%) variants.• Other Information 406: This category includes covariates that may influence genetic associations, such as age, sex, ancestry, and technical factors like batch effects or sampleOur Ref. MIT.1010PC processing time. Adjusting for these covariates helps to isolate the genetic component of the traits and minimize confounding factors.
[0061] The common and rare genetic association module 302 may employ various techniques to identify genetic variants associated with the phenotype 408. For example, the common and rare genetic association module 302 may use genome-wide association studies (GWAS), which can be applied to both the original continuous traits and the traits derived from quality control flags. The module 302 may apply GWAS analysis using PLINK2 software. The module 302 may conduct GWAS analysis using individuals from multiple population groups, such as white British, non-British white, South Asian, and African populations, to enable population-specific analysis and subsequent meta-analysis across populations. The module 302 may apply variant quality control procedures, including filtering variants with missingness greater than about 1%, excluding variants with Hardy-Weinberg disequilibrium test p- values less than about l.Oxio-7for directly genotyped variants, about l.Oxio-4for imputed HLA allelotypes, or about 1.0x10 -15 for whole-exome sequencing variants. The module 302 may filter variants based on minor allele frequency thresholds of greater than about 0.01% for imputed variants, imputation quality scores (INFO scores) greater than about 0.3, and may prioritize variants present in reference datasets such as the HapMap Phase 3 dataset. As another example, the module 302 may implement methods like the Multiple Rare-variants and Phenotypes (MRP) approach, which allows for joint analysis of multiple phenotypes and cohorts. This Bayesian model comparison approach can aggregate genetic associations across genetic variants in the same gene and compare null and non-null alternative models. As yet another example, the module 302 may use a joint multi-trait analysis. By considering both the traits derived from quality control flags and the corresponding continuous traits, the module 302 may improve statistical power for discovering trait-associated genes, especially for rare variants.
[0062] FIG. 5 illustrates an embodiment of the polygenic score modeling module 304 from FIG. 3. This module 304 generates a polygenic score model 510 by integrating various inputs to enhance genetic prediction capabilities. The polygenic score modeling module 304 may, for example, apply batch screening iterative lasso (BASIL) implemented in the R snpnet package to perform variable selection and effect size estimation simultaneously. This approach finds the exact solution for LI- and L2-penalized multivariate regression (Elastic Net). The module 304 may assign different penalty factors to various types of genetic variants, prioritizing putative protein-truncating variants, pathogenic variants, and protein-altering variants. For example, the module 304 may assign a penalty factor of about 0.5 toOur Ref. MIT.1010PC putative protein-truncating variants and pathogenic variants, about 0.75 to putative protein-altering variants and likely-pathogenic variants, and about 1.2 for variants that are not present in the HapMap Phase 3 dataset. This allows for more targeted analysis of potentially functional genetic variations. The sparsity of the solution may be controlled by optimizing the tuning parameter A. based on the predictive performance on a validation set. The module 304 may generate polygenic risk scores using both the original continuous traits and the traits derived from quality control flags, enabling cross-trait prediction and validation. The module 304 may perform meta-analysis of GWAS results using inverse-variance weighted (IVW) meta-analysis implemented in METAL software to combine association summary statistics across multiple population groups. The module 304 may conduct population-specific analysis by performing separate GWAS within individual population groups and then combining the results through meta-analysis to improve statistical power and ensure generalizability across diverse ancestry groups. In another example, the polygenic score modeling module 304 may apply the clumping and thresholding method. In yet another example, the polygenic score modeling module 304 may apply a Bayesian polygenic score modeling method implemented, for example, in LDpred2, SBayesR, or PRS-CSx.
[0063] The polygenic score modeling module 304 may generate the polygenic score model 510 based on any one or more of the following: the phenotypic information 402, the genetic information 404, the genetic associations 408, reference data 502, and other information 510. Because the phenotypic information 402, the genetic information 404, and the genetic associations 408 were already described in connection with FIG. 4, they will not be described again in connection with FIG. 5.
[0064] The reference data in populations 502 may, for example, include population-specific data, such as allele frequency information or patterns of linkage-disequilibrium from specific population groups, such as white British, non-British white, South Asian, and African individuals in the UK Biobank. The reference data 502 may be used for model development and evaluation. The polygenic score modeling module 304 may derive principal component loadings of genetic variations from the reference data 502 to account for population structure in the modeling process. The module 304 may perform population-specific polygenic score modeling by training separate models for different population groups or by incorporating population-specific covariates and genetic principal components to account for ancestry and population structure differences.
[0065] The other information 504 may include biological knowledge that can enhance the polygenic score modeling, such as functional annotations, pathway information, and prior knowledge from previous studies.Our Ref. MIT.1010PC
[0066] FIG. 6 illustrates an embodiment of the gene prioritization module 306 from FIG. 3. The gene prioritization module 306 identifies prioritized genes 602 using any of a variety of inputs, such as the phenotypic information 402, genetic information 404, phenotype 408, reference data 502, and other information 504 described above.
[0067] The gene prioritization module 306 may employ a Bayesian model comparison approach called Multiple Rare-variants and Phenotypes (MRP) for rare variant aggregation tests. This method compares null and non-null alternative models using observed data, considering the ratio of marginal likelihoods (reported as Bayes Factor, BF) to indicate the strength of evidence supporting genephenotype relationships. The module 306 may perform joint analysis of traits derived from quality control flags (such as Below Limit of Quantification flags) and corresponding continuous traits. This multi-trait approach has been shown to improve statistical power in rare variant aggregation tests, resulting in a larger number of prioritized genes. The module 306 may apply LD score regression (LDSC) to estimate the SNP-based heritability of traits, which may be used to assess the genetic basis of quality control flag-derived phenotypes and guide trait selection for further analysis. In some cases, traits with observed-scale heritability above about 5% may be prioritized for analysis, including traits with heritability values ranging from about 5.3% to about 7.1%. In another example, the gene prioritization module 306 may employ the sequence kernel association test (SKAT) implemented in the SKAT R- package.
[0068] The module 306 may focus on rare (MAF < 1%) non-synonymous coding variants, allowing for targeted analysis of potentially functional genetic variations. The gene prioritization module 306 may aggregate genetic associations across the same gene while jointly analyzing multiple phenotypes and cohorts. The module 306 may conduct gene prioritization using data from specific population groups (e.g., white British individuals) or perform meta-analysis across multiple population groups in resources like the UK Biobank. The module 306 may apply population-specific analysis approaches by conducting separate analyses within individual population groups and then combining results through meta-analysis methods to improve statistical power and generalizability across diverse populations.
[0069] The module 306 may conduct gene prioritization using data from specific population groups (e.g., white British individuals) or perform meta-analysis across multiple population groups in resources like the UK Biobank. Genes may be prioritized based on a predefined threshold of statistical evidence, such as a Bayes Factor (BF) greater than 105. The module 306 may differentiate between different types of genetic variants, such as protein-truncating variants (PTVs) and protein-altering variants (PAVs), in the prioritization process. The module 306 may implement population-specific quality control proceduresOur Ref. MIT.1010PC and covariate adjustments, including the use of population-specific genotype principal components to account for population structure and potential confounding factors in the genetic analysis across different ancestry groups.
[0070] By integrating these methods, the gene prioritization module 306 can identify genes that are strongly associated with the analyzed traits, including those derived from quality control flags. This approach has demonstrated significant improvements in gene prioritization compared to traditional methods, particularly for rare variants and extreme phenotypes.
[0071] More generally, embodiments of the present invention provide a novel approach to genetic analysis, termed hypometric genetics (hMG), which significantly enhances both genetic discovery and prediction capabilities. By leveraging previously underutilized data in large-scale genetic studies, particularly quality control flags (such as Below Limit of Quantification (BLQ) flags), the hMG method offers a powerful new tool for researchers and clinicians in the field of genetics and personalized medicine. In this application, the term "quality control (QC) flag" refers to any indicator used to mark measurements that fall outside standard quantification ranges or meet other quality control criteria. Below Limit of Quantification (BLQ) flags are one specific example of quality control flags, and are used throughout this application as an illustrative example. However, the methods described herein are applicable to various types of quality control flags. As this implies, any reference herein to BLQ flags should be understood to be applicable more generally to any kind of quality control flag.
[0072] Examples of other quality control flags include:• High ethanol flags: These flags indicate samples with unusually high ethanol levels, which could be due to contamination or other factors affecting measurement quality.• Degraded sample flags: These flags mark samples that show signs of degradation, which could impact the accuracy of metabolite measurements.• Unidentified small molecule flags: These indicate the presence of unidentified compounds that could interfere with accurate quantification of known metabolites.• Citrate plasma flags: These flags mark samples where citrate plasma was used instead of EDTA plasma, which could affect metabolite measurements.
[0073] At its core, the hMG approach addresses the limitations of traditional genetic analysis methods by treating quality control flags, such as BLQ flags, as valuable sources of genetic information rather than as noise to be discarded. This paradigm shift allows for the extraction of meaningful genetic insights from data points that were previously overlooked or excluded from analysis.
[0074] The hMG method operates on two primary fronts:Our Ref. MIT.1010PC• Genetic Discovery: Embodiments of the invention enhance the identification of genes and genetic variants associated with a wide range of traits, including disease outcomes, phenotypes, biomarkers, and other medically relevant measurements. By converting quality control flags (e.g., BLQ flags) into traits and performing genetic analysis, the hMG approach can identify genetic basis, associated variants, and prioritized genes. Furthermore, through joint multi-trait gene analysis of quality control flags with corresponding continuous traits, the hMG approach significantly improves the power of genetic discovery, particularly for rare variants and extreme phenotypes.• Genetic Prediction: The hMG method develops machine learning models for individual-level trait prediction. By incorporating information from quality control flag-derived traits into polygenic risk scores, the approach demonstrates superior predictive capabilities compared to traditional methods.
[0075] Traits, within the context of embodiments of the present invention, may encompass any of a wide range of characteristics, such as disease outcomes, phenotypes, biomarkers, medically relevant measurements, and / or other characteristics, such as BLQ traits.
[0076] The complementarity of these two aspects of the invention creates a powerful synergy, where genetic discoveries inform and improve prediction models, while the performance of these models validates and refines genetic discoveries. This iterative process leads to continually improving results and insights.
[0077] The benefits of the hMG approach are substantial and wide-ranging. By increasing the number of prioritized genes compared to traditional methods, the hMG approach offers researchers a more comprehensive understanding of the genetic basis of complex traits and diseases. Furthermore, the hMG approach opens new avenues for drug target identification and personalized medicine approaches.
[0078] The hMG method's application to omics data (e.g., NMR-based metabolomics data) demonstrates its versatility and potential impact across various domains of genetic research. By extracting valuable information from quality control flags, embodiments of the present invention not only improve the utilization of existing data but also potentially reduces the need for additional costly and time-consuming data collection efforts.
[0079] In summary, the hMG approach represents a significant advancement in the field of genetic analysis, offering improved genetic discovery and prediction capabilities through innovative data utilization. Its potential applications span from basic research to clinical practice, promising to accelerateOur Ref. MIT.1010PC our understanding of complex genetic traits and pave the way for more targeted and effective therapeutic interventions.
[0080] Having described examples of the general purpose, operation, and benefits of embodiments of the present invention at a high level, certain embodiments of the present invention will now be described in more detail.
[0081] The hMG method of embodiments of the present invention converts quality control (e.g., BLQ) flags into binary traits for genetic analysis to leverage previously underutilized data in genetic studies. The method begins by defining binarized quality control traits for all NMR metabolite measurements. This involves creating a binary indicator (0 or 1) for each measurement, where 1 indicates the presence of a quality control (e.g., BLQ) flag and 0 indicates its absence.
[0082] After defining the binary traits, the method then identifies which NMR metabolite traits have at least some minimum amount of measurements tagged with quality control flags. In some cases, this minimum amount may be about 1% of measurements tagged with quality control flags. This step helps to focus the analysis on traits with a sufficient number of BLQ occurrences to be informative.
[0083] The selected traits with sufficient quality control occurrences are then used for subsequent genetic analysis steps, including joint multi-trait analysis and rare variant aggregation tests (described below).
[0084] In cases where individuals have repeated measurements (e.g., from different time points), the method assigns individuals to the "case" group (1) when the quality control flags are present in all measurements, and to the "control" group (0) otherwise. This approach ensures consistency in trait definition across multiple time points.
[0085] The binary quality control flag traits may be adjusted for the effects of technical covariates, similar to the adjustment performed on the corresponding continuous traits. This step helps to minimize the influence of technical factors on the genetic analysis.
[0086] The conversion of BLQ flags (and other quality control flags) into binary traits offers several advantages. For example, by treating BLQ flags as informative traits rather than noise, the hMG method extracts valuable genetic information from data points that are typically discarded or ignored in traditional analyses. Furthermore, the binary BLQ traits effectively capture information about individuals with extremely low metabolite concentrations, which may be particularly informative for identifying genetic variants with strong effects. The binarized BLQ traits also provide information that is complementary to the continuous metabolite measurements. This is evidenced by the lower levels of phenotypic correlation observed among BLQ traits compared to the original continuous traits, and byOur Ref. MIT.1010PC the high genetic correlation between BLQ and corresponding quantitative traits, which may exhibit Pearson's correlation coefficients of about 0.937 or higher for effect size estimates. The BLQ traits also show enrichment for metabolite-lowering associations, particularly for rare variants. This suggests that the binary traits are especially useful for discovering genetic variants associated with extremely low metabolite levels. Furthermore, by incorporating BLQ traits in joint multi-trait analyses, the hMG method improves the statistical power for discovering trait-associated genes, especially for rare variants.
[0087] The hMG method also performs joint multi-trait analysis of BLQ and corresponding continuous traits. This approach includes considering both the binarized BLQ traits and their corresponding continuous traits for analysis. This pairing allows for a more comprehensive examination of genetic associations. The hMG method may apply a Bayesian model comparison method called Multiple Rare-variants and Phenotypes (MRP) for rare variant aggregation tests. This method is capable of jointly analyzing multiple phenotypes and multiple cohorts.
[0088] The multi-trait MRP analysis takes into account the genetic correlation between traits when constructing the null model. This consideration assists in the joint analysis of correlated traits, such as BLQ and continuous traits derived from the same metabolite measurements.
[0089] The method compares the Bayes Factor (BF) of gene-phenotype associations from singletrait and multi-trait analyses. This comparison allows for the quantification of improvements in statistical power and prioritization of trait-associated genes.
[0090] The joint multi-trait analysis has demonstrated significant improvements in gene prioritization. For example, in the analysis of Free Cholesterol in Chylomicrons (CMs) and Extremely Large VLDL (XXL VLDL), the multi-trait MRP analysis of four related traits prioritized 21 genes, nearly doubling the number of prioritized genes compared to the single-trait analysis of the original trait alone. In some cases, the multi-trait analysis may result in an average of about 181% more prioritized genes compared to single-trait analysis approaches.
[0091] The joint analysis is particularly beneficial for rare variant analysis. It has shown improvements in gene prioritization across different population groups in the UK Biobank study, highlighting the benefits of hMG in boosting statistical power in rare variant aggregation tests. The hMG multi-trait analysis has shown consistent improvements in gene prioritization across various traits and different types of genetic variants (e.g., protein-altering variants and protein-truncating variants). The method may implement population-specific analysis by conducting separate rare variant aggregation tests within individual population groups and then performing meta-analysis across populations toOur Ref. MIT.1010PC combine evidence and improve overall statistical power while accounting for population-specific genetic architecture differences.
[0092] By combining information from both BLQ and continuous traits, the hMG method improves the ability to detect genetic associations, particularly for rare variants. Furthermore, the joint analysis captures a more complete picture of genetic influences on traits, including effects that may be more apparent in extreme phenotypes (as captured by BLQ traits). This approach also maximizes the use of available data by incorporating information from quality control flags that were previously underutilized. In summary, the joint multi-trait analysis of BLQ and corresponding continuous traits significantly enhances the ability of the hMG method to identify trait-associated genes and improve our understanding of the genetic basis of complex traits and diseases.
[0093] As mentioned above, The hMG method employs a Bayesian model comparison approach called Multiple Rare-variants and Phenotypes (MRP) for rare variant aggregation tests. This approach is flexible and can aggregate genetic associations across the same gene while jointly analyzing multiple phenotypes and cohorts. The MRP compares null and non-null alternative models using observed data, considering the ratio of marginal likelihoods (reported as Bayes Factor, BF) to indicate the strength of evidence supporting gene-phenotype relationships. A BF > 1 suggests stronger evidence for the non-null hypothesis. The MRP analysis may be conducted on summary statistics from unrelated individuals in specific population groups or across multiple groups, enabling meta-analysis of rare variant aggregation tests assuming similar effects across populations. The method may account for population structure by using population-specific genotype principal components and may perform separate analyses within individual population groups before combining results through meta-analysis approaches.
[0094] MRP analysis in hMG considers genetic correlations between traits when constructing the null model for joint analysis of correlated traits. It can be conducted on summary statistics from unrelated individuals in specific population groups or across multiple groups, enabling meta-analysis of rare variant aggregation tests assuming similar effects across populations. The approach focuses on rare (MAF < 1%) non-synonymous coding variants, allowing for targeted analysis of potentially functional genetic variations.
[0095] The multi-trait MRP analysis has demonstrated significant improvements in gene prioritization compared to single-trait analysis. For example, it nearly doubled the number of prioritized genes in the analysis of free cholesterol in CMs and XXL VLDL. The improvements in gene prioritization through hMG multi-trait analysis have been observed across various traits and different types of genetic variants (e.g., protein-altering variants and protein-truncating variants).Our Ref. MIT.1010PC
[0096] The hMG method may generate of polygenic risk scores (PGS) using BLQ-derived binary traits to enhance genetic prediction capabilities. In particular, the hMG method may apply inclusive polygenic score (iPGS) modeling using batch screening iterative lasso (BASIL) implemented in the R snpnet package. This approach performs variable selection and effect size estimation simultaneously by finding the exact solution for LI- and L2-penalized multivariate regression (Elastic Net). The method may incorporate meta-analysis of GWAS summary statistics using inverse-variance weighted (IVW) metaanalysis implemented in METAL software to combine results across multiple population groups before polygenic score modeling. In one embodiment, the method uses a dataset of 1,316,181 variants for 284,661 individuals in the UK Biobank, considering both common and rare genetic variants. The model may, for example, include age, sex, age2, age*sex, Townsend deprivation index, and the first 18 genotype principal components as unpenalized covariates to account for potential confounding factors.
[0097] The method may assign different penalty factors to various types of genetic variants, prioritizing putative protein-truncating variants, pathogenic variants, and protein-altering variants. For example, the method may assign a penalty factor of about 0.5 to putative protein-truncating variants and pathogenic variants, about 0.75 to putative protein-altering variants and likely-pathogenic variants, about 1.2 for variants that are not present in the HapMap Phase 3 dataset, and about 1.0 for other remaining variants. The method may apply quality control procedures for the dataset, including using variants that pass missingness thresholds of less than about 1%, Hardy-Weinberg equilibrium test p- values greater than about l.Oxio’7for directly genotyped variants, imputation quality scores (INFO scores) greater than about 0.3 for imputed variants, and minor allele frequency thresholds greater than about 0.01% for imputed variants. The tuning parameter , which controls the sparsity of the solution, may be optimized based on the predictive performance on a validation set. An Elastic Net parameter a of 0.99 may be used, as in some embodiments of the method.
[0098] The PGS models trained on BLQ traits may be used to predict both the original continuous traits and the binarized BLQ traits. This cross-trait prediction demonstrates the ability of BLQ-derived models to capture relevant genetic information. The predictive performance of the PGS models may be evaluated using metrics such as R2for continuous traits and area under the receiver operating characteristic curve (AUROC) for binary traits. These evaluations may be performed on held-out test sets. The method may implement population-specific modeling approaches by randomly splitting each population group into training, validation, and test sets, enabling separate model development and evaluation within each population group while maintaining consistent analytical approaches across different ancestry groups.Our Ref. MIT.1010PC
[0099] The generation of polygenic risk scores using BLQ-derived binary traits offers several advantages. For example, PGS trained on BLQ traits alone can predict the original continuous traits with high accuracy, demonstrating the value of incorporating BLQ information. In some cases, such cross-trait prediction may achieve about 91% of the accuracy of models trained directly on the quantitative traits. Furthermore, the ability to predict original traits using BLQ-derived models suggests that these binary traits capture complementary genetic information. In addition, the success in cross-trait prediction indicates the potential for using genetic information to improve the accuracy of imputing measurement values below the limit of quantification. In summary, the hMG method's approach to generating polygenic risk scores using BLQ-derived binary traits represents a novel way to leverage previously underutilized data for genetic prediction, potentially improving our ability to predict trait values and understand the genetic basis of complex traits.
[0100] One embodiment of the present invention identifies biologically relevant genetic associations by applying hypometric genetics (hMG) genetic analysis approaches to traits below the limit of quantification (BLQ) and corresponding original traits. This step leverages the hMG approach to uncover genetic associations that might be missed when only considering original traits. The hMG approach is effective in identifying additional signals for genetic discovery, especially from phenotypic extremes, which are traits that exhibit extreme values and are often below the limit of quantification.
[0101] The hMG genetic analysis may include at least one of polygenic score (PGS) modeling using inclusive PGS (iPGS) and gene-based rare variant aggregation tests. The inclusive PGS (iPGS) model may be trained on binarized BLQ traits and original traits, which helps in predicting the presence or absence of the BLQ QC flag in a held-out test set. This predictive performance may be assessed in individuals, demonstrating the model's accuracy in predicting original quantitative traits. In some cases, the genetic associations for BLQ traits and corresponding quantitative traits may show high correlation, such as Pearson's correlation coefficients of about 0.937 or higher, indicating substantial consistency between the genetic architectures of these trait types.
[0102] The hMG genetic analysis may be applied to a large cohort of individuals with metabolomics profiles. Specifically, the approach has been demonstrated on n=227,469 UK Biobank individuals, showcasing its applicability to large-scale datasets. This large cohort provides a robust dataset for validating the hMG approach and its effectiveness in genetic discovery.
[0103] The hMG genetic analysis approach may be trained on binarized BLQ traits and original traits. This training process may involve using both types of traits to enhance the model's ability toOur Ref. MIT.1010PC predict genetic associations accurately. The binarized BLQ traits may indicate the presence or absence of a below the limit of quantification flag for metabolomic traits, which is an aspect of the hMG approach.
[0104] The hMG approach demonstrates biologically relevant genetic associations by comparing the marginal likelihoods for the null and the non-null alternative models, indicating additional signals for genetic discovery. This approach leverages BLQ flags to identify genetic associations and improve the power of genetic discovery. The process involves assessing the predictive performance in individuals in a held-out test set, where the iPGS model trained on binarized BLQ traits predicts the presence or absence of the BLQ QC flag with high accuracy. The predictive performance is quantified using R2 and AUROC values, with the iPGS model showing an R2=0.1039 for the original traits.
[0105] In summary, embodiments of the present invention may be used to identify biologically relevant genetic associations. The approach may leverage large-scale datasets, binarized BLQ traits, and inclusive PGS modeling to enhance the power of genetic discovery, particularly for traits below the limit of quantification.
[0106] Embodiments of the present invention may be assessed for their predictive performance in a held-out test set. Such an assessment may involve evaluating how well the hMG approach can predict the presence or absence of below the limit of quantification (BLQ) flags and original quantitative traits. The hMG approach, which may include polygenic score (PGS) modeling using inclusive PGS (iPGS), may be applied to binarized BLQ traits and original traits. The predictive performance may be measured using metrics such as R2 and AUROC values. For instance, the iPGS model trained on binarized BLQ traits may predict the presence or absence of the BLQ QC flag in a held-out test set of individuals, showing a predictive performance of R2=0.1039 for the original traits. This indicates that the hMG approach can predict the original quantitative traits with high accuracy, achieving an 80% accuracy rate for common variants. In some cases, cross-trait prediction using BLQ-derived models may achieve about 91% of the accuracy compared to models trained directly on the quantitative traits. The bar charts (especially FIG. 3C) shown in Tanigawa et al., "Hypometric genetics: Improved power in genetic discovery by incorporating quality control flags," Am J Hum Genet, lll(ll):2478-2493 (2024), illustrate the R2 and AUROC values for each model, highlighting the effectiveness of the hMG approach in genetic discovery.
[0107] Embodiments of the present invention may apply gene-based rare variant aggregation tests that focus on non-synonymous rare variant associations. This step may be used for genetic discovery, particularly in addressing the challenge of limited statistical power when analyzing rare events. The gene-based rare variant aggregation tests may be designed to highlight the impacts of non-synonymousOur Ref. MIT.1010PC variants with large effects. These tests may be applied to identify metabolite-lowering associations among rare variants, thereby improving the power of rare variant aggregation tests.
[0108] Embodiments of the present invention may assess statistical evidence supporting genephenotype relationships. This may be achieved by comparing the marginal likelihoods for null and nonnull alternative models given the observed data, a process implemented using a Bayesian model comparison approach in MRP. The statistical evidence may be reported as Bayes Factor (BF), which quantifies the support for the gene-phenotype relationship.
[0109] Embodiments of the present invention may compare the marginal likelihoods for null and non-null alternative models, in order to determine the statistical significance of the observed genetic associations. This comparison helps in identifying additional signals for genetic discovery, particularly from phenotypic extremes.
[0110] Embodiments of the present invention may perform gene-based rare variant aggregation tests that address the challenge of limited statistical power in analyzing rare events. By focusing on non- synonymous rare variant associations, these tests enhance the ability to detect significant genetic associations that might otherwise be missed.
[0111] Embodiments of the present invention may jointly analyze traits below the limit of quantification ( BLQ) and original traits. The result of such analysis may be to prioritize more genes than approaches which analyze original traits alone.
[0112] For example, embodiments of the present invention may leverage phenotypic extremes to identify additional signals for genetic discovery. This embodiment highlights the utility of phenotypic extremes in enhancing the detection of genetic associations that might be missed when only considering average trait values. The hMG approach, by focusing on these extremes, can uncover rare but significant genetic variations that contribute to the observed traits. By jointly analyzing BLQ and original traits, leveraging phenotypic extremes, and improving the power of genetic discovery, the hMG approach offers an advancement in the field of genetics. In some cases, the multi-trait analysis approach may result in an average of about 181% more prioritized genes compared to single-trait analysis methods.
[0113] FIGS. 8A-8F illustrate comparative analyses demonstrating the effectiveness of the hypometric genetics approach in identifying genetic associations, genetic prediction of complex traits, and prioritizing therapeutic targets. These figures provide empirical evidence supporting the utility of incorporating quality control flags into genetic analysis.Our Ref. MIT.1010PC
[0114] FIGS. 8A and 8B demonstrate the genetic basis of both binarized BLQ traits and corresponding quantitative traits. FIG. 7A shows common variant associations (left panel) and rare variant associations (right panel) for a binarized BLQ trait, while FIG. 8B shows the corresponding associations for a quantitative trait excluding BLQ measurements. The comparison reveals that both trait types exhibit substantial genetic associations across common and rare variants, with the BLQ traits capturing complementary genetic information that may be particularly valuable for identifying rare variants with large effects.
[0115] FIGS. 8C and 8D illustrate the predictive performance of polygenic score models trained on different trait types. FIG. 8C shows polygenic score prediction performance for a binarized BLQ trait, displaying odds ratios across polygenic score deciles, demonstrating the ability of genetic variants to predict the presence or absence of BLQ flags. FIG. 8D shows polygenic score prediction performance for a quantitative trait, displaying covariate-adjusted phenotype values across polygenic score percentile bins. These figures validate that polygenic scores trained on BLQ traits can effectively predict both the binary BLQ status and the underlying quantitative phenotypes, achieving approximately 91% of the accuracy compared to models trained directly on quantitative traits.
[0116] FIGS. 8E and 8F demonstrate the gene prioritization capabilities of the hMG approach through rare variant aggregation testing. FIG. 8E shows gene prioritization results for a binarized BLQ trait, while FIG. 8F shows corresponding results for a quantitative trait. Both figures display the strength of statistical evidence for gene-phenotype associations, measured as Bayes Factors on a logarithmic scale. The comparison illustrates how genetic analysis on BLQ traits may prioritize genetic discovery for rare variants and extreme phenotypes, thereby facilitating therapeutic target discovery and validation based on human genetic evidence.
[0117] FIGS. 7A-7C provide additional validation of the comparative genetic discovery from BLQ traits and quantitative traits, further supporting the effectiveness of the hypometric genetics approach across different analytical perspectives and population groups. FIG. 7C illustrates how incorporating BLQ traits may significantly increase the predictive accuracy of genetic prediction models for complex traits, thereby demonstrating the advantage of incorporating quality control flag information in genetic analysis.
[0118] FIG. 7A demonstrates the strong correlation between genetic effect size estimates for truncated quantitative traits and corresponding binarized BLQ traits. The comparison shows that genetic variants exhibit consistent directional effects across both trait types, with effect size estimates showingOur Ref. MIT.1010PC high correlation. This correlation validates that BLQ traits capture meaningful genetic signals that are consistent with those observed in quantitative trait analysis, supporting the biological relevance of the genetic associations identified through the hMG approach.
[0119] FIG. 7B illustrates the relationship between statistical significance levels of genetic associations for truncated quantitative traits versus binarized BLQ traits. The comparison reveals patterns in the strength of statistical evidence across the two trait types, demonstrating how genetic variants may show different levels of statistical significance depending on the trait representation. This analysis may help identify variants that are particularly informative for extreme phenotypes captured by BLQ flags compared to those detected through quantitative trait analysis.
[0120] FIG. 7C presents a comprehensive comparison of predictive performance across different polygenic score modeling approaches and multiple population groups. The figure shows R2values for polygenic score models trained on original traits (including BLQ measurements), truncated traits (excluding BLQ measurements), and binarized BLQ traits across white British, non-British white, South Asian, and African population groups. This analysis illustrates how incorporating BLQ traits may significantly increase the predictive accuracy of genetic prediction models for complex traits. This analysis also demonstrates the cross-population applicability of the hMG approach and validates that models incorporating BLQ information maintain predictive performance across diverse ancestry groups. The results demonstrate the advantage of incorporating quality control flag information in genetic analysis and support the generalizability of the hypometric genetics method and its potential utility in multi-ancestry genetic studies.
[0121] Embodiments of the present invention may prioritize therapeutic targets based on the genetic analysis performed using the hMG approach. This prioritization may leverage the enhanced genetic discovery capabilities of the hMG method to identify genes and genetic variants that may serve as potential therapeutic targets for drug development or other medical interventions. The hMG approach's ability to identify metabolite-lowering associations for rare variants, particularly through the analysis of BLQ traits, provides valuable insights for therapeutic target identification and validation. For example, the method's demonstrated effectiveness in identifying genes such as APOC3, APOA5, and PDE3B, which have known roles in lipid metabolism and cardiovascular disease, illustrates the potential for discovering clinically relevant therapeutic targets.
[0122] The prioritization of therapeutic targets may comprise various activities related to therapeutic target development. For example, prioritizing therapeutic targets may include identifyingOur Ref. MIT.1010PC genes or gene products that show significant associations with traits of interest through the genetic analysis. The hMG method's joint multi-trait analysis approach, which has demonstrated an average of about 181% more prioritized genes compared to single-trait analysis methods , as illustrated in FIGS. 9A- 9C.
[0123] FIG. 9A shows a scatter plot comparing the strength of statistical evidence (loglO Bayes Factor) from single-trait analysis (x-axis) versus multi-trait analysis of four traits (y-axis), with genes labeled (APOC3, APOA5, LPL, etc.) and a diagonal line indicating y=x representing no improvement in hMG multi-trait analysis.
[0124] FIG. 9B shows an inset or zoomed view of the same comparison, highlighting genes like LPL, PLA2G12A, GCKR, ANGPTL3, SIK3, PDE3B, APOB, and APOA1 that achieve prioritization through multitrait analysis.
[0125] FIG. 9C shows bar charts displaying the number of prioritized genes (with Bayes Factor > 105) across five NMR metabolomics traits (CE in CMs and XXL VLDL, Free Choi, in CMs and XXL VLDL, Free Choi, in XL VLDL, PLs in CMs and XXL VLDL, and PLs in XL VLDL), comparing different trait combinations and shown separately for non-synonymous variants and protein-truncating variants, demonstrating that the joint analysis of BLQ and quantitative traits in multi-trait approaches significantly increases the number of prioritized genes compared to single-trait analyses, thereby enhancing the power of genetic discovery for rare variants and extreme phenotypes.
[0126] The hMG approach's ability of improving the statistical power, as illustrated in FIGS. 9A-9C. significantly expands the pool of potential therapeutic targets for consideration from the same amount of dataset. Alternatively or additionally, the improved statistical power with hMG methods may reduce the amount of dataset necessary to establish the comparable level of genetic evidence for therapeutic target discovery or validation. The prioritization may include validating these identified targets through additional analyses or experimental approaches to confirm their potential therapeutic relevance. This validation process may involve cross-referencing identified genes with known biological pathways, disease mechanisms, or existing therapeutic targets to assess their clinical potential. The prioritization may involve ranking genes or gene products based on factors such as the strength of genetic associations, as measured by Bayes Factors (BF) greater than 105, biological relevance to disease pathways, druggability assessments based on protein structure and function, safety profiles, and other criteria relevant to therapeutic target selection.Our Ref. MIT.1010PC
[0127] The hMG method's particular effectiveness in identifying rare variants with large effects, such as protein-truncating variants (PTVs) and protein-altering variants (PAVs), provides opportunities to prioritize targets that may have substantial therapeutic impact. The prioritization may also consider the enrichment of metabolite-lowering associations observed in BLQ traits, with about 97.3% of nominally significant associations showing metabolite-lowering effects compared to about 58.8% for corresponding continuous traits, which may be particularly valuable for identifying targets relevant to cardiometabolic disorders. The prioritization may include at least one of identifying, validating, or ranking the therapeutic targets (e.g., genes or gene products) based on the genetic analysis, with the ultimate goal of advancing the most promising candidates toward drug development and clinical applications.
[0128] Embodiments of the present invention may tailor clinical care for individuals based on the genetic analysis performed using the hMG approach. This tailoring may leverage the enhanced genetic prediction capabilities of the hMG method to personalize medical care and treatment decisions for individuals based on their genetic profiles and predicted genetic scores. The hMG approach's ability to generate polygenic risk scores with reliable accuracy, such as achieving about 91% of the predictive performance compared to models trained directly on truncated quantitative traits, supports the use of hMG approach in generating personalized predicted genetic scores. Moreover, the hMG approach's ability to incorporate BLQ information in polygenic risk scores with improved accuracy, such as achieving about 118% of the predictive performance compared to models training directly on truncated quantitative traits enables more precise risk stratification and personalized treatment recommendations. The method's particular effectiveness in identifying rare variants with large effects and metabolite-lowering associations provides valuable insights for tailoring interventions related to metabolic disorders and cardiovascular disease risk.
[0129] The tailoring of clinical care may comprise various activities related to personalized medicine and clinical decision-making. For example, tailoring clinical care may include enrolling individuals in clinical trials, treatment programs, or preventive care initiatives based on predicted genetic scores for the individuals. The hMG-derived polygenic scores may be particularly valuable for identifying individuals who are most likely to benefit from specific interventions, such as those with genetic predispositions to cardiometabolic disorders identified through the analysis of BLQ traits and corresponding metabolite measurements. The tailoring may include not enrolling individuals in certain clinical trials or treatment programs based on predicted genetic scores that indicate potential adverseOur Ref. MIT.1010PC outcomes or lack of therapeutic benefit. This selective enrollment approach may help optimize clinical trial efficiency and reduce exposure of individuals to potentially ineffective or harmful treatments. The tailoring may involve administering drugs to individuals based on predicted genetic scores that suggest favorable treatment responses or therapeutic efficacy. For instance, individuals with genetic variants identified through the hMG approach as associated with specific metabolic pathways may be prioritized for targeted therapies affecting those pathways. The tailoring may include not administering drugs to individuals based on predicted genetic scores that indicate potential adverse drug reactions or lack of therapeutic benefit, thereby reducing healthcare costs and minimizing unnecessary medication exposure.
[0130] The tailoring of clinical care may also involve adjusting the frequency of healthcare facility visits for individuals based on predicted genetic scores. This adjustment may include increasing the frequency of healthcare facility visits for individuals based on predicted genetic scores that indicate higher risk for disease development or progression, thereby enabling more intensive monitoring and early intervention. For example, individuals with high polygenic risk scores for cardiometabolic disorders, as determined through the hMG analysis of quality control flags and corresponding traits, may benefit from more frequent monitoring of metabolic biomarkers and cardiovascular risk factors. The adjustment may include decreasing the frequency of healthcare facility visits for individuals based on predicted genetic scores that indicate lower risk profiles, thereby optimizing healthcare resource allocation while maintaining appropriate care standards. This risk-stratified approach to healthcare delivery may be particularly valuable in managing chronic diseases where early intervention can significantly impact long-term outcomes.
[0131] The tailoring may include at least one of enrolling, not enrolling, administering drugs, not administering drugs, or adjusting the frequency of healthcare facility visits for the individuals based on the genetic analysis and predicted genetic scores derived from the hMG approach. The integration of hMG-derived insights into clinical decision-making represents a significant advancement in precision medicine, enabling healthcare providers to make more informed decisions about patient care based on comprehensive genetic risk profiles that incorporate previously underutilized quality control information. This approach may be particularly valuable for managing complex traits and diseases where traditional genetic analysis methods may have missed important associations, as demonstrated by the hMG method's ability to prioritize an average of about 181% more genes compared to single-trait analysis approaches.Our Ref. MIT.1010PC
[0132] It is to be understood that although the invention has been described above in terms of particular embodiments, the foregoing embodiments are provided as illustrative only, and do not limit or define the scope of the invention. Various other embodiments, including but not limited to the following, are also within the scope of the claims. For example, elements and components described herein may be further divided into additional components or joined together to form fewer components for performing the same functions.
[0133] Any of the functions disclosed herein may be implemented using means for performing those functions. Such means include, but are not limited to, any of the components disclosed herein, such as the computer-related components described below.
[0134] The techniques described above may be implemented, for example, in hardware, one or more computer programs tangibly stored on one or more computer-readable media, firmware, or any combination thereof. The techniques described above may be implemented in one or more computer programs executing on (or executable by) a programmable computer including any combination of any number of the following: a processor, a storage medium readable and / or writable by the processor (including, for example, volatile and non-volatile memory and / or storage elements), an input device, and an output device. Program code may be applied to input entered using the input device to perform the functions described and to generate output using the output device.
[0135] Embodiments of the present invention include features which are only possible and / or feasible to implement with the use of one or more computers, computer processors, and / or other elements of a computer system. Such features are either impossible or impractical to implement mentally and / or manually. For example, the hypometric genetics (hMG) method significantly improves the functionality of computer systems used for genetic analysis by enabling them to process and analyze previously underutilized data in the form of Below Limit of Quantification (BLQ) flags. This enhancement allows for more comprehensive and accurate genetic discoveries and predictions, representing a substantial improvement to the functioning of computer systems in the context of genetic analysis.
[0136] This improvement achieves this improvement using a variety of mechanisms. For example, the method converts BL flags into binary traits, effectively transforming previously discarded or ignored data points into valuable genetic information. This process allows computer systems to extract meaningful insights from a broader range of data, increasing the overall efficiency and effectiveness of genetic analysis. By incorporating BLQ-derived binary traits into genetic analyses, the hMG method significantly improves the statistical power for discovering trait-associated genes, particularly for rare variants and extreme phenotypes. This enhancement enables computer systems to identify geneticOur Ref. MIT.1010PC associations that may have been missed by traditional methods, thereby expanding the scope and accuracy of genetic discoveries. In addition, the hMG approach implements a joint multi-trait analysis of BLQ and corresponding continuous traits. This allows computer systems to capture a more complete picture of genetic influences on traits, including effects that may be more apparent in extreme phenotypes. The improved analytical capability results in a more comprehensive understanding of complex genetic relationships. The hMG method also trains and applies machine learning models for individual-level trait prediction by incorporating information from BLQ-derived binary traits into polygenic risk scores. This integration of previously underutilized data into predictive models enhances the accuracy and robustness of genetic predictions, representing a significant improvement in the predictive capabilities of computer systems used for genetic analysis.
[0137] By leveraging quality control flags, the hMG method maximizes the use of available data without requiring additional costly and time-consuming data collection efforts. This efficiency in data processing and analysis represents an improvement in the overall performance of computer systems in genetic studies. The hMG approach demonstrates particular effectiveness in identifying metabolitelowering associations for rare variants. This capability enhances the computer system's ability to analyze and interpret rare genetic variations, which are often challenging to study using traditional methods.
[0138] In summary, the hMG method's innovative approach to incorporating quality control flags into genetic analysis represents a significant improvement to the functionality of computer systems used in this field. By enabling these systems to process and analyze previously underutilized data, the hMG method enhances their ability to make comprehensive and accurate genetic discoveries and predictions, ultimately advancing our understanding of complex genetic traits and diseases.
[0139] Furthermore, the hMG method performs a transformation of data into a different state or thing, thereby further supporting its patent eligibility. For example, an embodiment of the hMG method transforms Below Limit of Quantification (BLQ) flags into binary traits, which are then utilized in novel ways for genetic analysis. This transformation process and its subsequent application in improving genetic discovery and prediction represent a significant data transformation that goes beyond mere abstract ideas.
[0140] The transformation process includes, for example, the following steps:• Binary Trait Definition: The hMG method first defines binarized BLQ traits for all NMR metabolite measurements. This involves creating a binary indicator (0 or 1) for each measurement, where 1 indicates the presence of a BLQ flag and 0 indicates its absence.• Data Selection: After defining the binary traits, the method identifies which NMR metaboliteOur Ref. MIT.1010PC traits have at least 1% of measurements tagged with BLQ QC flags. This step focuses the analysis on traits with a sufficient number of BLQ occurrences to be informative.• Covariate Adjustment: The binary BLQ traits are adjusted for the effects of technical covariates, similar to the adjustment performed on the corresponding continuous traits. This step helps to minimize the influence of technical factors on the genetic analysis.
[0141] This transformation and subsequent use of BLQ flags go beyond mere abstract ideas by:• Extracting valuable genetic information from data points that are typically discarded or ignored in traditional analyses.• Enabling the capture of information about individuals with extremely low metabolite concentrations, which may be particularly informative for identifying genetic variants with strong effects.• Providing complementary information to the continuous metabolite measurements, as evidenced by the lower levels of phenotypic correlation observed among BLQ traits compared to the original continuous traits.• Improving statistical power for discovering trait-associated genes, especially for rare variants and extreme phenotypes.
[0142] In summary, the hMG method's transformation of quality control flags into traits and their novel application in genetic analysis represents a significant data transformation that goes beyond abstract ideas. This approach enables the extraction of valuable genetic information from previously underutilized data, enhances genetic discovery, particularly for rare variants and extreme phenotypes, and contributes to more comprehensive genetic analyses.
[0143] Furthermore, the hypometric genetics (hMG) method integrates the use of quality control flags into a practical application that significantly improves genetic discovery and prediction. This integration goes beyond merely applying an abstract idea on a computer by providing tangible improvements in statistical power for identifying rare variants and extreme phenotypes.
[0144] The practical application of the hMG method includes, for example, the transformation of Below Limit of Quantification (BLQ) flags into informative binary traits, effectively capturing information about individuals with extremely low metabolite concentrations. This conversion allows for the extraction of valuable genetic information from data points that are typically discarded or ignored in traditional analyses. The joint multi-trait analysis implemented by the hMG method has demonstrated significant improvements in gene prioritization. For example, in the analysis of free cholesterol in Chylomicrons (CMs) and Extremely Large (XXL) VLDL, the multi-trait MRP analysis of four related traitsOur Ref. MIT.1010PC prioritized 11 genes, nearly quadrupling the number of prioritized genes from 3, compared to the single-trait analysis of the original trait alone.
[0145] Furthermore, the BLQ traits show enrichment for metabolite-lowering associations, particularly for rare variants. For protein-altering variant ( PAV) associations in prioritized genes, about 97.3% of nominally significant associations may have metabolite-lowering effects across BLQ traits, compared to about 58.8% for the corresponding continuous traits. This 65.4% enrichment highlights the advantage of hMG analysis in discovering rare variants with metabolite-lowering effects.
[0146] By integrating BLQ flags into a practical application that demonstrably improves genetic discovery and prediction, the hMG method goes beyond merely applying an abstract idea on a computer. It provides a novel, effective approach to leveraging previously underutilized data, resulting in tangible improvements in our ability to understand the genetic basis of complex traits and diseases.
[0147] In some aspects of the present invention, a method for genetic analysis is performed by at least one computer processor executing computer program instructions stored on at least one non- transitory computer-readable medium. The method comprises identifying an original phenotype from an original data source for phenotypic data, performing quality control flag identification on the original phenotypic data source to generate quality control flags, converting the identified quality control flags into auxiliary traits, generating quality control-derived phenotypes from the auxiliary traits, and performing genetic analysis on at least the quality control-derived phenotypes.
[0148] In some embodiments, generating the quality control flags comprises deriving the quality control flags from omics data measurements. The omics data measurements may comprise at least one of: nuclear magnetic resonance (NMR) spectroscopy measurements of metabolites, levels of gene expression values, or levels of protein expression values.
[0149] In some embodiments, converting the identified quality control flags into auxiliary traits comprises converting the quality control flags into binary traits indicating the presence or absence of each quality control flag. The binary traits may comprise Below Limit of Quantification (BLQ) indicators or Below Limit of Detection (BLD) indicators. In some embodiments, the binary traits comprise sample quality indicators. Generating the quality control-derived phenotypes from the converted traits may comprise creating phenotypes that capture information about individuals with extreme phenotypic values as indicated by the binary traits.
[0150] In some embodiments, for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether the quality control flags are present in all measurements, whether an amount of quality control flagsOur Ref. MIT.1010PC exceeds a predetermined threshold, or whether an amount of quality control flags falls below a predetermined threshold. Converting the quality control flags into binary traits may comprise selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags, or selecting traits that have no more than a percentage of measurements tagged with the quality control flags.
[0151] In some embodiments, performing genetic analysis comprises jointly analyzing both the original phenotype and the quality control-derived phenotypes in a multi-trait analysis. The multi-trait analysis may consider genetic correlations between traits when constructing a null model for joint analysis of correlated traits. Performing the genetic analysis by jointly analyzing both the original phenotype and the quality control-derived phenotypes in the multi-trait analysis may comprise performing a Bayesian model comparison approach that compares null and non-null alternative models using observed data. The Bayesian model comparison approach may comprise calculating a ratio of marginal likelihoods reported as a Bayes Factor to indicate strength of evidence supporting genephenotype relationships, or applying a Multiple Rare-variants and Phenotypes (MRP) approach for rare variant aggregation tests.
[0152] In some embodiments, performing genetic analysis comprises performing at least one of: rare variant aggregation tests, genome-wide association studies, or polygenic score modeling.
[0153] In some aspects of the present invention, a system for genetic analysis includes at least one non-transitory computer-readable medium having computer program instructions stored thereon, the computer program instructions being executable by at least one computer processor to perform a method. The method comprises identifying an original phenotype from an original data source for phenotypic data, performing quality control flag identification on the original phenotypic data source to generate quality control flags, converting the identified quality control flags into auxiliary traits, generating quality control-derived phenotypes from the auxiliary traits, and performing genetic analysis on at least the quality control-derived phenotypes.
[0154] In some aspects of the present invention, a method for nominating therapeutic targets using genetic analysis comprises identifying an original phenotype from an original data source for phenotypic data, performing quality control flag identification on the original phenotypic data source to generate quality control flags, converting the identified quality control flags into auxiliary traits, generating quality control-derived phenotypes from the auxiliary traits, performing genetic analysis on at least the quality control-derived phenotypes, and nominating therapeutic targets based on the genetic analysis.Our Ref. MIT.1010PC
[0155] In some embodiments, nominating the therapeutic targets comprises at least one of identifying, validating, or ranking the therapeutic targets based on the genetic analysis.
[0156] In some embodiments, generating the quality control flags comprises deriving the quality control flags from omics data measurements. The omics data measurements may comprise at least one of: nuclear magnetic resonance (NMR) spectroscopy measurements of metabolites, levels of gene expression values, or levels of protein expression values.
[0157] In some embodiments, converting the identified quality control flags into auxiliary traits comprises converting the quality control flags into binary traits indicating the presence or absence of each quality control flag. The binary traits may comprise Below Limit of Quantification (BLQ) indicators or Below Limit of Detection (BLD) indicators. In some embodiments, the binary traits comprise sample quality indicators. Generating the quality control-derived phenotypes from the converted traits may comprise creating phenotypes that capture information about individuals with extreme phenotypic values as indicated by the binary traits.
[0158] In some embodiments, for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether the quality control flags are present in all measurements, whether an amount of quality control flags exceeds a predetermined threshold, or whether an amount of quality control flags falls below a predetermined threshold. Converting the quality control flags into binary traits may comprise selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags, or selecting traits that have no more than a percentage of measurements tagged with the quality control flags.
[0159] In some embodiments, converting the quality control flags into binary traits comprises selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags.
[0160] In some embodiments, converting the quality control flags into binary traits comprises selecting traits that have no more than a percentage of measurements tagged with the quality control flags.
[0161] In some embodiments, performing genetic analysis comprises jointly analyzing both the original phenotype and the quality control-derived phenotypes in a multi-trait analysis, and nominating therapeutic targets is based on results of the multi-trait analysis.
[0162] In some embodiments, the multi-trait analysis considers genetic correlations between traits when constructing a null model for joint analysis of correlated traits.Our Ref. MIT.1010PC
[0163] In some embodiments, performing the genetic analysis by jointly analyzing both the original phenotype and the quality control-derived phenotypes in the multi-trait analysis comprises performing a Bayesian model comparison approach that compares null and non-null alternative models using observed data, and nominating therapeutic targets is based on results of the Bayesian model comparison approach.
[0164] In some embodiments, the Bayesian model comparison approach comprises calculating a ratio of marginal likelihoods reported as a Bayes Factor to indicate strength of evidence supporting gene-phenotype relationships.
[0165] In some embodiments, the Bayesian model comparison approach comprises applying a Multiple Rare-variants and Phenotypes (MRP) approach for rare variant aggregation tests.
[0166] In some embodiments, performing genetic analysis comprises performing at least one of: rare variant aggregation tests, genome-wide association studies, or polygenic score modeling, and nominating therapeutic targets is based on results of the genetic analysis.
[0167] In some aspects of the present invention, a system for nominating therapeutic targets using genetic analysis includes at least one non-transitory computer-readable medium having computer program instructions stored thereon, the computer program instructions being executable by at least one computer processor to perform a method. The method comprises identifying an original phenotype from an original data source for phenotypic data, performing quality control flag identification on the original phenotypic data source to generate quality control flags, converting the identified quality control flags into auxiliary traits, generating quality control-derived phenotypes from the auxiliary traits, performing genetic analysis on at least the quality control-derived phenotypes, and nominating therapeutic targets based on the genetic analysis.
[0168] In some aspects of the present invention, a method for tailoring clinical care for individuals based on genetic analysis comprises identifying an original phenotype from an original data source for phenotypic data, performing quality control flag identification on the original phenotypic data source to generate quality control flags, converting the identified quality control flags into auxiliary traits, generating quality control-derived phenotypes from the auxiliary traits, performing genetic analysis on at least the quality control-derived phenotypes, and tailoring clinical care for the individuals based on the genetic analysis.
[0169] In some embodiments, tailoring clinical care for the individuals comprises enrolling individuals based on predicted genetic scores for the individuals.
[0170] In some embodiments, tailoring clinical care for the individuals comprises not enrollingOur Ref. MIT.1010PC individuals based on predicted genetic scores for the individuals.
[0171] In some embodiments, tailoring clinical care for the individuals comprises administering drugs to the individuals based on predicted genetic scores for the individuals.
[0172] In some embodiments, tailoring clinical care for the individuals comprises not administering drugs to the individuals based on predicted genetic scores for the individuals.
[0173] In some embodiments, tailoring clinical care for the individuals comprises adjusting the frequency of healthcare facility visits for individuals based on predicted genetic scores for the individuals.
[0174] In some embodiments, adjusting the frequency of healthcare facility visits comprises increasing the frequency of healthcare facility visits for the individuals based on predicted genetic scores for the individuals.
[0175] In some embodiments, adjusting the frequency of healthcare facility visits comprises decreasing the frequency of healthcare facility visits for the individuals based on predicted genetic scores for the individuals.
[0176] In some embodiments, generating the quality control flags comprises deriving the quality control flags from omics data measurements.
[0177] In some embodiments, the omics data measurements comprise at least one of: nuclear magnetic resonance (NMR) spectroscopy measurements of metabolites, levels of gene expression values, or levels of protein expression values.
[0178] In some embodiments, converting the identified quality control flags into auxiliary traits comprises converting the quality control flags into binary traits indicating the presence or absence of each quality control flag.
[0179] In some embodiments, the binary traits comprise Below Limit of Quantification (BLQ) indicators.
[0180] In some embodiments, the binary traits comprise Below Limit of Detection (BLD) indicators.
[0181] In some embodiments, generating the quality control-derived phenotypes from the converted traits comprises creating phenotypes that capture information about individuals with extreme phenotypic values as indicated by the binary traits.
[0182] In some embodiments, the binary traits comprise sample quality indicators.
[0183] In some embodiments, for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether the quality control flags are present in all measurements.Our Ref. MIT.1010PC
[0184] In some embodiments, for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags exceeds a predetermined threshold.
[0185] In some embodiments, for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags falls below a predetermined threshold.
[0186] In some embodiments, converting the quality control flags into binary traits comprises selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags.
[0187] In some embodiments, converting the quality control flags into binary traits comprises selecting traits that have no more than a percentage of measurements tagged with the quality control flags.
[0188] In some embodiments, performing genetic analysis comprises jointly analyzing both the original phenotype and the quality control-derived phenotypes in a multi-trait analysis.
[0189] In some embodiments, the multi-trait analysis considers genetic correlations between traits when constructing a null model for joint analysis of correlated traits.
[0190] In some embodiments, performing the genetic analysis by jointly analyzing both the original phenotype and the quality control-derived phenotypes in the multi-trait analysis comprises performing a Bayesian model comparison approach that compares null and non-null alternative models using observed data.
[0191] In some embodiments, the Bayesian model comparison approach comprises calculating a ratio of marginal likelihoods reported as a Bayes Factor to indicate strength of evidence supporting gene-phenotype relationships.
[0192] In some embodiments, the Bayesian model comparison approach comprises applying a Multiple Rare-variants and Phenotypes (MRP) approach for rare variant aggregation tests.
[0193] In some embodiments, performing genetic analysis comprises performing at least one of: rare variant aggregation tests, genome-wide association studies, or polygenic score modeling.
[0194] In some aspects of the present invention, a system for tailoring clinical care for individuals based on genetic analysis includes at least one non-transitory computer-readable medium having computer program instructions stored thereon, the computer program instructions being executable by at least one computer processor to perform a method. The method comprises identifying an original phenotype from an original data source for phenotypic data, performing quality control flagOur Ref. MIT.1010PC identification on the original phenotypic data source to generate quality control flags, converting the identified quality control flags into auxiliary traits, generating quality control-derived phenotypes from the auxiliary traits, performing genetic analysis on at least the quality control-derived phenotypes, and tailoring clinical care for the individuals based on the genetic analysis.
[0195] Any claims herein which affirmatively require a computer, a processor, a memory, or similar computer-related elements, are intended to require such elements, and should not be interpreted as if such elements are not present in or required by such claims. Such claims are not intended, and should not be interpreted, to cover methods and / or systems which lack the recited computer-related elements. For example, any method claim herein which recites that the claimed method is performed by a computer, a processor, a memory, and / or similar computer-related element, is intended to, and should only be interpreted to, encompass methods which are performed by the recited computer- related element(s). Such a method claim should not be interpreted, for example, to encompass a method that is performed mentally or by hand (e.g., using pencil and paper). Similarly, any product claim herein which recites that the claimed product includes a computer, a processor, a memory, and / or similar computer-related element, is intended to, and should only be interpreted to, encompass products which include the recited computer-related element(s). Such a product claim should not be interpreted, for example, to encompass a product that does not include the recited computer-related element(s).
[0196] Each computer program within the scope of the claims below may be implemented in any programming language, such as assembly language, machine language, a high-level procedural programming language, or an object-oriented programming language. The programming language may, for example, be a compiled or interpreted programming language.
[0197] Each such computer program may be implemented in a computer program product tangibly embodied in a machine-readable storage device for execution by a computer processor. Method steps of the invention may be performed by one or more computer processors executing a program tangibly embodied on a computer-readable medium to perform functions of the invention by operating on input and generating output. Suitable processors include, by way of example, both general and special purpose microprocessors. Generally, the processor receives (reads) instructions and data from a memory (such as a read-only memory and / or a random access memory) and writes (stores) instructions and data to the memory. Storage devices suitable for tangibly embodying computer program instructions and data include, for example, all forms of non-volatile memory, such as semiconductor memory devices, including EPROM, EEPROM, and flash memory devices; magnetic disks such as internalOur Ref. MIT.1010PC hard disks and removable disks; magneto-optical disks; and CD-ROMs. Any of the foregoing may be supplemented by, or incorporated in, specially-designed ASICs (application-specific integrated circuits) or FPGAs (Field-Programmable Gate Arrays). A computer can generally also receive (read) programs and data from, and write (store) programs and data to, a non-transitory computer-readable storage medium such as an internal disk (not shown) or a removable disk. These elements will also be found in a conventional desktop or workstation computer as well as other computers suitable for executing computer programs implementing the methods described herein, which may be used in conjunction with any digital print engine or marking engine, display monitor, or other raster output device capable of producing color or gray scale pixels on paper, film, display screen, or other output medium.
[0198] Any data disclosed herein may be implemented, for example, in one or more data structures tangibly stored on a non-transitory computer-readable medium. Embodiments of the invention may store such data in such data structure(s) and read such data from such data structure(s).
[0199] Any step or act disclosed herein as being performed, or capable of being performed, by a computer or other machine, may be performed automatically by a computer or other machine, whether or not explicitly disclosed as such herein. A step or act that is performed automatically is performed solely by a computer or other machine, without human intervention. A step or act that is performed automatically may, for example, operate solely on inputs received from a computer or other machine, and not from a human. A step or act that is performed automatically may, for example, be initiated by a signal received from a computer or other machine, and not from a human. A step or act that is performed automatically may, for example, provide output to a computer or other machine, and not to a human.
[0200] The terms "A or B," "at least one of A or / and B," "at least one of A and B," "at least one of A or B," or "one or more of A or / and B" used in the various embodiments of the present disclosure include any and all combinations of words enumerated with it. For example, "A or B," "at least one of A and B" or "at least one of A or B" may mean: (1) including at least one A, (2) including at least one B, (3) including either A or B, or (4) including both at least one A and at least one B.
[0201] Although terms such as "optimize" and "optimal" are used herein, in practice, embodiments of the present invention may include methods which produce outputs that are not optimal, or which are not known to be optimal, but which nevertheless are useful. For example, embodiments of the present invention may produce an output which approximates an optimal solution, within some degree of error. As a result, terms herein such as "optimize" and "optimal" should beOur Ref. MIT.1010PC understood to refer not only to processes which produce optimal outputs, but also processes which produce outputs that approximate an optimal solution, within some degree of error.
Claims
Our Ref. MIT.1010PCClaimsWhat is claimed is:
1. A method for genetic analysis performed by at least one computer processor executing computer program instructions stored on at least one non-transitory computer-readable medium, the method comprising: identifying an original phenotype from an original data source for phenotypic data; performing quality control flag identification on the original phenotypic data source to generate quality control flags; converting the identified quality control flags into auxiliary traits; generating quality control-derived phenotypes from the auxiliary traits; and performing genetic analysis on at least the quality control-derived phenotypes.
2. The method of claim 1, wherein generating the quality control flags comprises deriving the quality control flags from omics data measurements.
3. The method of claim 2, wherein the omics data measurements comprise at least one of: nuclear magnetic resonance (NMR) spectroscopy measurements of metabolites, levels of gene expression values, or levels of protein expression values.
4. The method of claim 1, wherein converting the identified quality control flags into auxiliary traits comprises converting the quality control flags into binary traits indicating the presence or absence of each quality control flag.Our Ref. MIT.1010PC5. The method of claim 4, wherein the binary traits comprise Below Limit of Quantification (BLQ) indicators.
6. The method of claim 4, wherein the binary traits comprise Below Limit of Detection (BLD) indicators.
7. The method of claim 4, wherein generating the quality control-derived phenotypes from the converted traits comprises creating phenotypes that capture information about individuals with extreme phenotypic values as indicated by the binary traits.
8. The method of claim 4, wherein the binary traits comprise sample quality indicators.
9. The method of claim 4, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether the quality control flags are present in all measurements.
10. The method of claim 4, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags exceeds a predetermined threshold.
11. The method of claim 4, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags falls below a predetermined threshold.
12. The method of claim 4, wherein converting the quality control flags into binary traits comprises selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags.Our Ref. MIT.1010PC13. The method of claim 4, wherein converting the quality control flags into binary traits comprises selecting traits that have no more than a percentage of measurements tagged with the quality control flags.
14. The method of claim 1, wherein performing genetic analysis comprises jointly analyzing both the original phenotype and the quality control-derived phenotypes in a multi-trait analysis.
15. The method of claim 14, wherein the multi-trait analysis considers genetic correlations between traits when constructing a null model for joint analysis of correlated traits.
16. The method of claim 14, wherein performing the genetic analysis by jointly analyzing both the original phenotype and the quality control-derived phenotypes in the multi-trait analysis comprises performing a Bayesian model comparison approach that compares null and non-null alternative models using observed data.
17. The method of claim 16, wherein the Bayesian model comparison approach comprises calculating a ratio of marginal likelihoods reported as a Bayes Factor to indicate strength of evidence supporting gene-phenotype relationships.
18. The method of claim 16, wherein the Bayesian model comparison approach comprises applying a Multiple Rare-variants and Phenotypes (MRP) approach for rare variant aggregation tests.
19. The method of claim 1, wherein performing genetic analysis comprises performing at least one of: rare variant aggregation tests, genome-wide association studies, or polygenic score modeling.Our Ref. MIT.1010PC20. A system for genetic analysis, the system including at least one non-transitory computer- readable medium having computer program instructions stored thereon, the computer program instructions being executable by at least one computer processor to perform a method, the method comprising: identifying an original phenotype from an original data source for phenotypic data; performing quality control flag identification on the original phenotypic data source to generate quality control flags; converting the identified quality control flags into auxiliary traits; generating quality control-derived phenotypes from the auxiliary traits; and performing genetic analysis on at least the quality control-derived phenotypes.
21. The system of claim 20, wherein generating the quality control flags comprises deriving the quality control flags from omics data measurements.
22. The system of claim 21, wherein the omics data measurements comprise at least one of: nuclear magnetic resonance (NMR) spectroscopy measurements of metabolites, levels of gene expression values, or levels of protein expression values.
23. The system of claim 20, wherein converting the identified quality control flags into auxiliary traits comprises converting the quality control flags into binary traits indicating the presence or absence of each quality control flag.
24. The system of claim 23, wherein the binary traits comprise Below Limit of Quantification (BLQ) indicators.
25. The system of claim 23, wherein the binary traits comprise Below Limit of Detection (BLD) indicators.
26. The system of claim 23, wherein generating the quality control-derived phenotypes from the converted traits comprises creating phenotypes that capture information about individuals with extreme phenotypic values as indicated by the binary traits.
27. The system of claim 23, wherein the binary traits comprise sample quality indicators.Our Ref. MIT.1010PC28. The system of claim 23, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether the quality control flags are present in all measurements.
29. The system of claim 23, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags exceeds a predetermined threshold.
30. The system of claim 23, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags falls below a predetermined threshold.
31. The system of claim 23, wherein converting the quality control flags into binary traits comprises selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags.
32. The system of claim 23, wherein converting the quality control flags into binary traits comprises selecting traits that have no more than a percentage of measurements tagged with the quality control flags.
33. The system of claim 20, wherein performing genetic analysis comprises jointly analyzing both the original phenotype and the quality control-derived phenotypes in a multi-trait analysis.
34. The system of claim 33, wherein the multi-trait analysis considers genetic correlations between traits when constructing a null model for joint analysis of correlated traits.
35. The system of claim 33, wherein performing the genetic analysis by jointly analyzing both the original phenotype and the quality control-derived phenotypes in the multi-trait analysis comprises performing a Bayesian model comparison approach that compares null and non-null alternative models using observed data.Our Ref. MIT.1010PC36. The system of claim 35, wherein the Bayesian model comparison approach comprises calculating a ratio of marginal likelihoods reported as a Bayes Factor to indicate strength of evidence supporting gene-phenotype relationships.
37. The system of claim 35, wherein the Bayesian model comparison approach comprises applying a Multiple Rare-variants and Phenotypes (MRP) approach for rare variant aggregation tests.
38. The system of claim 20, wherein performing genetic analysis comprises performing at least one of: rare variant aggregation tests, genome-wide association studies, or polygenic score modeling.
39. A method for nominating therapeutic targets using genetic analysis, comprising: identifying an original phenotype from an original data source for phenotypic data; performing quality control flag identification on the original phenotypic data source to generate quality control flags; converting the identified quality control flags into auxiliary traits; generating quality control-derived phenotypes from the auxiliary traits; performing genetic analysis on at least the quality control-derived phenotypes; and nominating therapeutic targets based on the genetic analysis.
40. The method of claim 39, wherein nominating the therapeutic targets comprises at least one of identifying, validating, or ranking the therapeutic targets based on the genetic analysis.
41. The method of claim 39, wherein generating the quality control flags comprises deriving the quality control flags from omics data measurements.Our Ref. MIT.1010PC42. The method of claim 41, wherein the omics data measurements comprise at least one of: nuclear magnetic resonance (NMR) spectroscopy measurements of metabolites, levels of gene expression values, or levels of protein expression values.
43. The method of claim 39, wherein converting the identified quality control flags into auxiliary traits comprises converting the quality control flags into binary traits indicating the presence or absence of each quality control flag.
44. The method of claim 43, wherein the binary traits comprise Below Limit of Quantification (BLQ) indicators.
45. The method of claim 43, wherein the binary traits comprise Below Limit of Detection (BLD) indicators.
46. The method of claim 43, wherein generating the quality control-derived phenotypes from the converted traits comprises creating phenotypes that capture information about individuals with extreme phenotypic values as indicated by the binary traits.
47. The method of claim 43, wherein the binary traits comprise sample quality indicators.
48. The method of claim 43, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether the quality control flags are present in all measurements.Our Ref. MIT.1010PC49. The method of claim 43, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags exceeds a predetermined threshold.
50. The method of claim 43, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags falls below a predetermined threshold.
51. The method of claim 43, wherein converting the quality control flags into binary traits comprises selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags.
52. The method of claim 43, wherein converting the quality control flags into binary traits comprises selecting traits that have no more than a percentage of measurements tagged with the quality control flags.
53. The method of claim 39, wherein performing genetic analysis comprises jointly analyzing both the original phenotype and the quality control-derived phenotypes in a multi-trait analysis, and wherein nominating therapeutic targets is based on results of the multi-trait analysis.
54. The method of claim 53, wherein the multi-trait analysis considers genetic correlations between traits when constructing a null model for joint analysis of correlated traits.
55. The method of claim 53, wherein performing the genetic analysis by jointly analyzing both the original phenotype and the quality control-derived phenotypes in the multi-trait analysis comprises performing a Bayesian model comparison approach that compares null and non-null alternative modelsOur Ref. MIT.1010PC using observed data, and wherein nominating therapeutic targets is based on results of the Bayesian model comparison approach.
56. The method of claim 55, wherein the Bayesian model comparison approach comprises calculating a ratio of marginal likelihoods reported as a Bayes Factor to indicate strength of evidence supporting gene-phenotype relationships.
57. The method of claim 55, wherein the Bayesian model comparison approach comprises applying a Multiple Rare-variants and Phenotypes (MRP) approach for rare variant aggregation tests.
58. The method of claim 39, wherein performing genetic analysis comprises performing at least one of: rare variant aggregation tests, genome-wide association studies, or polygenic score modeling, and wherein nominating therapeutic targets is based on results of the genetic analysis.
59. A system for nominating therapeutic targets using genetic analysis, the system including at least one non-transitory computer-readable medium having computer program instructions stored thereon, the computer program instructions being executable by at least one computer processor to perform a method, the method comprising: identifying an original phenotype from an original data source for phenotypic data; performing quality control flag identification on the original phenotypic data source to generate quality control flags; converting the identified quality control flags into auxiliary traits; generating quality control-derived phenotypes from the auxiliary traits; performing genetic analysis on at least the quality control-derived phenotypes; and nominating therapeutic targets based on the genetic analysis.
60. The system of claim 59, wherein nominating the therapeutic targets comprises at least one of identifying, validating, or ranking the therapeutic targets based on the genetic analysis.Our Ref. MIT.1010PC61. The system of claim 59, wherein generating the quality control flags comprises deriving the quality control flags from omics data measurements.
62. The system of claim 61, wherein the omics data measurements comprise at least one of: nuclear magnetic resonance (NMR) spectroscopy measurements of metabolites, levels of gene expression values, or levels of protein expression values.
63. The system of claim 59, wherein converting the identified quality control flags into auxiliary traits comprises converting the quality control flags into binary traits indicating the presence or absence of each quality control flag.
64. The system of claim 63, wherein the binary traits comprise Below Limit of Quantification (BLQ) indicators.
65. The system of claim 63, wherein the binary traits comprise Below Limit of Detection (BLD) indicators.
66. The system of claim 63, wherein generating the quality control-derived phenotypes from the converted traits comprises creating phenotypes that capture information about individuals with extreme phenotypic values as indicated by the binary traits.
67. The system of claim 63, wherein the binary traits comprise sample quality indicators.
68. The system of claim 63, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether the quality control flags are present in all measurements.
69. The system of claim 63, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags exceeds a predetermined threshold.Our Ref. MIT.1010PC70. The system of claim 63, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags falls below a predetermined threshold.
71. The system of claim 63, wherein converting the quality control flags into binary traits comprises selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags.
72. The system of claim 63, wherein converting the quality control flags into binary traits comprises selecting traits that have no more than a percentage of measurements tagged with the quality control flags.
73. The system of claim 59, wherein performing genetic analysis comprises jointly analyzing both the original phenotype and the quality control-derived phenotypes in a multi-trait analysis, and wherein nominating therapeutic targets is based on results of the multi-trait analysis.
74. The system of claim 73, wherein the multi-trait analysis considers genetic correlations between traits when constructing a null model for joint analysis of correlated traits.
75. The system of claim 73, wherein performing the genetic analysis by jointly analyzing both the original phenotype and the quality control-derived phenotypes in the multi-trait analysis comprises performing a Bayesian model comparison approach that compares null and non-null alternative models using observed data, and wherein nominating therapeutic targets is based on results of the Bayesian model comparison approach.
76. The system of claim 75, wherein the Bayesian model comparison approach comprises calculating a ratio of marginal likelihoods reported as a Bayes Factor to indicate strength of evidence supporting gene-phenotype relationships.
77. The system of claim 75, wherein the Bayesian model comparison approach comprises applying a Multiple Rare-variants and Phenotypes (MRP) approach for rare variant aggregation tests.Our Ref. MIT.1010PC78. The system of claim 59, wherein performing genetic analysis comprises performing at least one of: rare variant aggregation tests, genome-wide association studies, or polygenic score modeling, and wherein nominating therapeutic targets is based on results of the genetic analysis.
79. A method for tailoring clinical care for individuals based on genetic analysis, comprising: identifying an original phenotype from an original data source for phenotypic data; performing quality control flag identification on the original phenotypic data source to generate quality control flags; converting the identified quality control flags into auxiliary traits; generating quality control-derived phenotypes from the auxiliary traits; performing genetic analysis on at least the quality control-derived phenotypes; and tailoring clinical care for the individuals based on the genetic analysis.
80. The method of claim 79, wherein tailoring clinical care for the individuals comprises enrolling individuals based on predicted genetic scores for the individuals.
81. The method of claim 79, wherein tailoring clinical care for the individuals comprises not enrolling individuals based on predicted genetic scores for the individuals.
82. The method of claim 79, wherein tailoring clinical care for the individuals comprises administering drugs to the individuals based on predicted genetic scores for the individuals.
83. The method of claim 79, wherein tailoring clinical care for the individuals comprises not administering drugs to the individuals based on predicted genetic scores for the individuals.Our Ref. MIT.1010PC84. The method of claim 79, wherein tailoring clinical care for the individuals comprises adjusting the frequency of healthcare facility visits for individuals based on predicted genetic scores for the individuals.
85. The method of claim 84, wherein adjusting the frequency of healthcare facility visits comprises increasing the frequency of healthcare facility visits for the individuals based on predicted genetic scores for the individuals.
86. The method of claim 84, wherein adjusting the frequency of healthcare facility visits comprises decreasing the frequency of healthcare facility visits for the individuals based on predicted genetic scores for the individuals.
87. The method of claim 79, wherein generating the quality control flags comprises deriving the quality control flags from omics data measurements.
88. The method of claim 87, wherein the omics data measurements comprise at least one of: nuclear magnetic resonance (NMR) spectroscopy measurements of metabolites, levels of gene expression values, or levels of protein expression values.
89. The method of claim 79, wherein converting the identified quality control flags into auxiliary traits comprises converting the quality control flags into binary traits indicating the presence or absence of each quality control flag.
90. The method of claim 89, wherein the binary traits comprise Below Limit of Quantification (BLQ) indicators.Our Ref. MIT.1010PC91. The method of claim 89, wherein the binary traits comprise Below Limit of Detection (BLD) indicators.
92. The method of claim 89, wherein generating the quality control-derived phenotypes from the converted traits comprises creating phenotypes that capture information about individuals with extreme phenotypic values as indicated by the binary traits.
93. The method of claim 89, wherein the binary traits comprise sample quality indicators.
94. The method of claim 89, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether the quality control flags are present in all measurements.
95. The method of claim 89, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags exceeds a predetermined threshold.
96. The method of claim 89, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags falls below a predetermined threshold.
97. The method of claim 89, wherein converting the quality control flags into binary traits comprises selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags.
98. The method of claim 89, wherein converting the quality control flags into binary traits comprises selecting traits that have no more than a percentage of measurements tagged with the quality control flags.Our Ref. MIT.1010PC99. The method of claim 79, wherein performing genetic analysis comprises jointly analyzing both the original phenotype and the quality control-derived phenotypes in a multi-trait analysis.
100. The method of claim 99, wherein the multi-trait analysis considers genetic correlations between traits when constructing a null model for joint analysis of correlated traits.
101. The method of claim 99, wherein performing the genetic analysis by jointly analyzing both the original phenotype and the quality control-derived phenotypes in the multi-trait analysis comprises performing a Bayesian model comparison approach that compares null and non-null alternative models using observed data.
102. The method of claim 101, wherein the Bayesian model comparison approach comprises calculating a ratio of marginal likelihoods reported as a Bayes Factor to indicate strength of evidence supporting gene-phenotype relationships.
103. The method of claim 101, wherein the Bayesian model comparison approach comprises applying a Multiple Rare-variants and Phenotypes (MRP) approach for rare variant aggregation tests.
104. The method of claim 79, wherein performing genetic analysis comprises performing at least one of: rare variant aggregation tests, genome-wide association studies, or polygenic score modeling.
105. A system for tailoring clinical care for individuals based on genetic analysis, the system including at least one non-transitory computer-readable medium having computer program instructions stored thereon, the computer program instructions being executable by at least one computer processor to perform a method, the method comprising: identifying an original phenotype from an original data source for phenotypic data;Our Ref. MIT.1010PC performing quality control flag identification on the original phenotypic data source to generate quality control flags; converting the identified quality control flags into auxiliary traits; generating quality control-derived phenotypes from the auxiliary traits; performing genetic analysis on at least the quality control-derived phenotypes; and tailoring clinical care for the individuals based on the genetic analysis.
106. The system of claim 105, wherein tailoring clinical care for the individuals comprises enrolling individuals based on predicted genetic scores for the individuals.
107. The system of claim 105, wherein tailoring clinical care for the individuals comprises not enrolling individuals based on predicted genetic scores for the individuals.
108. The system of claim 105, wherein tailoring clinical care for the individuals comprises administering drugs to the individuals based on predicted genetic scores for the individuals.
109. The system of claim 105, wherein tailoring clinical care for the individuals comprises not administering drugs to the individuals based on predicted genetic scores for the individuals.
110. The system of claim 105, wherein tailoring clinical care for the individuals comprises adjusting the frequency of healthcare facility visits for individuals based on predicted genetic scores for the individuals.
111. The system of claim 110, wherein adjusting the frequency of healthcare facility visits comprises increasing the frequency of healthcare facility visits for the individuals based on predicted genetic scores for the individuals.
112. The system of claim 110, wherein adjusting the frequency of healthcare facility visits comprises decreasing the frequency of healthcare facility visits for the individuals based on predicted genetic scores for the individuals.Our Ref. MIT.1010PC113. The system of claim 105, wherein generating the quality control flags comprises deriving the quality control flags from omics data measurements.
114. The system of claim 113, wherein the omics data measurements comprise at least one of: nuclear magnetic resonance (NMR) spectroscopy measurements of metabolites, levels of gene expression values, or levels of protein expression values.
115. The system of claim 105, wherein converting the identified quality control flags into auxiliary traits comprises converting the quality control flags into binary traits indicating the presence or absence of each quality control flag.
116. The system of claim 115, wherein the binary traits comprise Below Limit of Quantification (BLQ) indicators.
117. The system of claim 115, wherein the binary traits comprise Below Limit of Detection (BLD) indicators.
118. The system of claim 115, wherein generating the quality control-derived phenotypes from the converted traits comprises creating phenotypes that capture information about individuals with extreme phenotypic values as indicated by the binary traits.
119. The system of claim 115, wherein the binary traits comprise sample quality indicators.
120. The system of claim 115, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether the quality control flags are present in all measurements.
121. The system of claim 115, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags exceeds a predetermined threshold.Our Ref. MIT.1010PC122. The system of claim 115, wherein for individuals with repeated measurements, converting the quality control flags into binary traits comprises assigning values to the binary trait based on whether an amount of quality control flags falls below a predetermined threshold.
123. The system of claim 115, wherein converting the quality control flags into binary traits comprises selecting traits that have at least a minimum percentage of measurements tagged with the quality control flags.
124. The system of claim 115, wherein converting the quality control flags into binary traits comprises selecting traits that have no more than a percentage of measurements tagged with the quality control flags.
125. The system of claim 105, wherein performing genetic analysis comprises jointly analyzing both the original phenotype and the quality control-derived phenotypes in a multi-trait analysis.
126. The system of claim 125, wherein the multi-trait analysis considers genetic correlations between traits when constructing a null model for joint analysis of correlated traits.
127. The system of claim 125, wherein performing the genetic analysis by jointly analyzing both the original phenotype and the quality control-derived phenotypes in the multi-trait analysis comprises performing a Bayesian model comparison approach that compares null and non-null alternative models using observed data.
128. The system of claim 127, wherein the Bayesian model comparison approach comprises calculating a ratio of marginal likelihoods reported as a Bayes Factor to indicate strength of evidence supporting gene-phenotype relationships.
129. The system of claim 127, wherein the Bayesian model comparison approach comprises applying a Multiple Rare-variants and Phenotypes (MRP) approach for rare variant aggregation tests.
130. The system of claim 105, wherein performing genetic analysis comprises performing at least one of: rare variant aggregation tests, genome-wide association studies, or polygenic score modeling.