Causal inference method and device based on genetic variation, electronic equipment and medium
By directly acquiring whole-genome SNP data and proteome data from the detection platform for analysis, the problem of insufficient data specificity in Mendelian randomization technology is solved, achieving higher analytical accuracy and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-24
AI Technical Summary
Existing Mendelian randomization techniques rely on public databases, which cannot cover all specific populations and lack data specificity. Cross-platform data integration introduces batch effects, leading to reduced accuracy and reliability of analysis results.
By directly obtaining whole-genome SNP and proteomic data of the target population from the testing platform, genome-wide association analysis and quantitative trait locus analysis of proteins are performed. Combined with Mendelian randomization, the technical differences between platforms are eliminated, and the accuracy and reliability of the analysis results are improved.
By covering the entire chain of analysis from proteomics detection to Mendelian randomization, platform-to-platform differences are eliminated, improving the accuracy and reliability of Mendelian randomization analysis results.
Smart Images

Figure CN121922192A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of genome association analysis technology, and in particular to a method, apparatus, electronic device, and medium for causal inference based on genetic variation. Background Technology
[0002] Mendelian randomization (MR) is a causal inference method based on genetic variation (usually in the form of single nucleotide polymorphisms, SNPs). Its basic principle is to use the influence of randomly assigned genotypes on phenotypes in nature to infer the influence of biological factors on outcomes. By introducing an intermediate variable called an instrumental variable, it analyzes the causal relationship between exposure factors and outcomes, solving the problem that traditional experimental methods cannot effectively explain the causal relationship between exposure factors and outcome variables due to the presence of confounding factors. Its basic idea is to use genetic variations strongly correlated with exposure factors as instrumental variables to infer the causal relationship between exposure factors and outcome variables. Here, the term "exposure" is used to refer to the hypothetical causal risk factor, sometimes also called an intermediate phenotype. This can be a biomarker, physical measurement, or any other risk factor that may affect the outcome. Typically, the outcome is a disease, but it is not limited to diseases.
[0003] Current standard Mendelian randomization procedures primarily extract corresponding genetic variant SNP data from public databases and then perform Mendelian randomization analysis. However, existing Mendelian randomization techniques mainly rely on public databases (such as UK Biobank and Fenland), which cannot be replaced and do not cover all specific populations, resulting in insufficient data specificity. Furthermore, cross-platform data integration can introduce batch effects, thereby reducing the reliability of the results. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, apparatus, electronic device and medium for causal inference based on genetic variation, so as to improve the accuracy and reliability of Mendelian random analysis results.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a causal inference method based on genetic variation, comprising: acquiring whole-genome SNP data and proteomic data of test samples from a target population from a detection platform; performing genome-wide association analysis on the whole-genome SNP data to obtain outcome SNP data associated with the target phenotype; performing quantitative trait locus analysis of proteins based on the whole-genome SNP data and proteomic data to obtain exposure SNP data associated with the target exposure factor; performing data preprocessing on the outcome SNP data and exposure SNP data, and performing Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype.
[0006] Optionally, genome-wide association analysis (GWIA) can be performed on whole-genome SNP data to obtain outcome SNP data associated with the target phenotype. This includes: preprocessing the whole-genome SNP data to obtain preprocessed whole-genome SNP data; performing GWIA with phenotype data based on a general linear model (GLM) or a mixed linear model (MLM) to obtain the SNP effect value and significance index for each whole-genome SNP data; and determining the outcome SNP data associated with the target phenotype based on the SNP effect value and significance index.
[0007] Optionally, quantitative trait locus analysis of proteins based on genome-wide SNP data and proteomic data can be performed to obtain exposure SNP data associated with the target exposure factor. This includes: standardizing the proteomic data to obtain a standardized plasma or tissue protein expression matrix; filtering the genome-wide SNP data and normalizing the standardized plasma or tissue protein expression matrix; and performing quantitative trait locus analysis of proteins on the filtered genome-wide SNP data and the standardized plasma or tissue protein expression matrix to obtain exposure SNP data associated with the target exposure factor.
[0008] Optionally, data preprocessing is performed on the outcome SNP data and exposure SNP data, including: calculating statistical indicators for the outcome SNP data and exposure SNP data; wherein the statistical indicators include at least: p-value, F-test value, and R2 value; removing exposure SNP data with p-values greater than a first threshold and removing outcome SNP data with p-values less than the first threshold; removing outcome SNP data or exposure SNP data within a preset range whose R2 value with the target outcome SNP data or target exposure SNP data is greater than a second threshold; removing exposure SNP data with F-values greater than a third threshold; removing outcome SNP data whose allele frequencies are less than a fourth threshold; and integrating the format and structure of the outcome SNP data and exposure SNP data.
[0009] Optionally, Mendelian randomization analysis is performed on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype. This includes: performing Mendelian randomization analysis using a pre-defined regression method to obtain the estimated effect size, standard error, and p-value of the relationship between the exposure SNP data and the target exposure factor or target phenotype; and determining the causal relationship between the target exposure factor and the target phenotype based on the estimated effect size, standard error, confidence interval, and p-value.
[0010] Optionally, after performing Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype, the method further includes: performing heterogeneity tests and level pleiotropy tests on the exposure SNP data based on the estimated effect size and the p-value of the estimated effect size; and visualizing the exposure SNP data and the estimated effect size.
[0011] Secondly, the present invention provides a causal inference device based on genetic variation, comprising: a raw data acquisition module for acquiring whole-genome SNP data and proteomic data of test samples from a target population from a detection platform; an outcome SNP data acquisition module for performing genome-wide association analysis on the whole-genome SNP data to obtain outcome SNP data associated with the target phenotype; an exposure SNP data acquisition module for performing quantitative trait locus analysis of proteins based on the whole-genome SNP data and proteomic data to obtain exposure SNP data associated with the target exposure factor; and a Mendelian randomization analysis module for performing data preprocessing on the outcome SNP data and exposure SNP data, and performing Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype.
[0012] Optionally, the outcome SNP data acquisition module is specifically used for: preprocessing whole-genome SNP data to obtain preprocessed whole-genome SNP data; performing genome-wide association analysis with phenotypic data based on the general linear model GLM or the mixed linear model MLM to obtain the SNP effect value and significance index of each whole-genome SNP data, and determining the outcome SNP data associated with the target phenotype based on the SNP effect value and significance index.
[0013] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the steps of the method provided in any of the first aspects above.
[0014] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, performs the steps of the method provided in any of the first aspects above.
[0015] This invention brings the following beneficial effects: The causal inference method, apparatus, electronic device, and medium based on genetic variation provided by this invention first acquire whole-genome SNP data and proteomic data of test samples from the target population from a detection platform; then, perform genome-wide association analysis (GWIA) on the whole-genome SNP data to obtain outcome SNP data associated with the target phenotype; combined with protein quantitative trait locus analysis (MTTL) based on the whole-genome SNP data and proteomic data, obtain exposure SNP data associated with the target exposure factor; finally, perform data preprocessing on the outcome SNP data and exposure SNP data, and perform Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype. In this method, whole-genome SNP data and proteomic data of test samples from the target population obtained from the detection platform are directly used for GWIA and MTL analysis, covering the entire chain from proteomic detection to Mendelian randomization analysis. This eliminates reliance on public databases, thereby eliminating technical differences between platforms and improving the accuracy and reliability of Mendelian randomization analysis results.
[0016] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a causal inference method based on genetic variation provided in an embodiment of the present invention; Figure 2A schematic diagram of a causal inference device based on genetic variation provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Currently, existing Mendelian randomization techniques primarily rely on public databases. These databases are irreplaceable, and because they do not cover all specific populations, their data specificity is insufficient. It cannot be guaranteed that the population being analyzed shares similar characteristics or backgrounds with those in public databases, potentially leading to population stratification. Furthermore, the populations in public databases are mostly Western Europeans, whose genetic variations may differ significantly from East Asians due to dietary habits, climate, and other factors. This could result in different results within the same study, leading to false positives and unnecessary bias. In addition, public databases may use different detection technologies and platforms (e.g., proteins are analyzed using DIA, SomaScan, and Olink platforms). Cross-platform data integration introduces batch effects, reducing the reliability of the results.
[0022] Based on this, the present invention provides a method, apparatus, electronic device and medium for causal inference based on genetic variation, which can improve the accuracy and reliability of Mendelian random analysis results.
[0023] To facilitate understanding of this embodiment, a causal inference method based on genetic variation disclosed in this invention will first be described in detail. This method can be executed by electronic devices, such as smartphones, computers, and tablets. See also Figure 1 The flowchart shown illustrates a causal inference method based on genetic variation, indicating that the method mainly includes the following steps S101 to S104: Step S101: Obtain whole-genome SNP data and proteomic data of the target population from the detection platform.
[0024] In one implementation, SNP microarrays or high-throughput sequencing technologies can be used to obtain whole-genome SNP data and proteomic data of test samples from the target population. Single nucleotide polymorphisms (SNPs) refer to variations occurring at a single nucleotide position in the genome. These variations can be single base transitions, transversions, or insertions and deletions, such as a base pair mutating from AT to CG, causing a gene mutation. SNPs are one of the most common forms of genetic variation in the genome.
[0025] Step S102: Perform genome-wide association analysis on the whole-genome SNP data to obtain the outcome SNP data associated with the target phenotype.
[0026] In one implementation, genome-wide association study (GWAS) is a research method that detects the association between genetic variations (such as single nucleotide polymorphisms, SNPs) and phenotypes (such as diseases, physiological characteristics) across the entire genome. Its core idea is to identify genetic loci significantly associated with target traits by comparing genotypic differences among individuals and using statistical models.
[0027] In this embodiment of the invention, the whole genome SNP data can first be subjected to data quality control and filtering, then whole genome association analysis can be performed, and after multiple validation corrections, the outcome SNP data associated with the target phenotype can be obtained.
[0028] Step S103: Perform quantitative trait locus analysis of proteins based on whole-genome SNP data and proteome data to obtain exposure SNP data associated with the target exposure factor.
[0029] In one implementation, a protein quantitative trait locus (pQTL) is a genetic variation locus associated with protein expression levels. These genetic variations can affect processes such as transcription, translation, or degradation of proteins in specific genes, thereby influencing protein levels in the human body. pQTLs are a branch of Genome-Wide Association Studies (GWAS), primarily exploring the association between genotype and protein expression levels. Specifically, high-throughput sequencing technology is used to measure protein expression levels in a large number of individuals, and this data is compared with the genotype data of the corresponding individuals to identify genetic variations associated with specific protein levels. The discovery of pQTLs helps to understand the differences in protein levels among individuals and the relationship between these differences and biological processes such as health and disease. The discovery of these loci can also provide a foundation for studying the association between proteins and diseases, helping to reveal potential therapeutic targets and biological mechanisms.
[0030] In this embodiment of the invention, protein quantitative trait locus analysis can be performed based on the sample information of the test sample, whole genome SNP data and proteome data to obtain SNP data related to protein expression, that is, exposure SNP data associated with the target exposure factor.
[0031] Step S104: Perform data preprocessing on the outcome SNP data and exposure SNP data, and conduct Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype.
[0032] In one implementation, Mendelian randomization analysis is performed using regression methods to obtain the estimated effect size, standard error of the estimated effect size, and p-value of statistical significance of the relationship between the SNP and the exposure or outcome, thereby obtaining the causal relationship between the target exposure factor and the target phenotype.
[0033] The causal inference method based on genetic variation provided by this invention directly utilizes whole-genome SNP data and proteomic data of the target population samples obtained from the detection platform to perform genome-wide association analysis and protein quantitative trait locus analysis, covering the entire chain from proteomic detection to Mendelian randomization analysis. It no longer relies on public databases, thereby eliminating technical differences between platforms and improving the accuracy and reliability of Mendelian randomization analysis results.
[0034] In one implementation, for the aforementioned step S102, i.e., when performing genome-wide association analysis on whole-genome SNP data to obtain outcome SNP data associated with the target phenotype, the following methods may be used, including but not limited to: First, the whole-genome SNP data is preprocessed to obtain preprocessed whole-genome SNP data. In practice, data quality control and filtering are performed on the whole-genome SNP data, such as excluding SNPs and samples with high deletion rates, excluding SNPs with high minor allele frequencies, and filtering for multi-allele loci and sex chromosome SNPs, resulting in processed whole-genome SNP data. Specifically, the whole-genome SNP data can be filtered for an SNP deletion rate of 0.1%, a sample deletion rate of 0.05%, and a minor allele frequency of 0.01, and multi-allele loci and sex chromosomes can also be filtered.
[0035] Then, genome-wide association analysis was performed on the phenotypic data using the general linear model GLM or the mixed linear model MLM to obtain the SNP effect value and significance index of each genome-wide SNP data, and the outcome SNP data associated with the target phenotype were determined based on the SNP effect value and significance index.
[0036] In practice, tools such as PLINK and Gemma can be used to perform genome-wide association analysis based on general linear models (GLM) or mixed linear models (MLM) and phenotypic data (including categorical traits such as yield, quality, and disease status; and continuous traits such as height, which require multiple environments and repeated measurements to improve accuracy) to calculate SNP effect values (Beta / OR) and significance indicators (P-value).
[0037] Specifically, the Generalized Linear Model (GLM) is a flexible statistical model suitable for handling continuous or categorical phenotypic data. Its basic form is: Y = Xβ +
[0038] in, Y It is tabular data. X It is a covariate matrix (such as age, gender, principal genetic components, etc.). β It is the regression coefficient. This is the error term.
[0039] Mixture linear models (MLMs) introduce random effects into geometric linear models (GLMs) to handle complex data structures such as population stratification and correlation. Their form is: Y = Xβ+Zγ +
[0040] in, Z It is a random effects matrix (such as family structure or group stratification); γ It is the regression coefficient of random effects.
[0041] Specifically, the model can be selected based on the data. When the data structure is simple and the population stratification is not significant, GLM can be chosen as the first choice; when there is population stratification or individual correlation in the data, MLM should be chosen to reduce false positives. Specific steps include: (1) Perform quality control on whole-genome SNP data, standardize phenotypic data, and calculate principal components (PCA) to correct for population stratification. (2) If GLM is selected, construct a linear regression model with phenotypic data as the dependent variable and covariates (such as age, sex, and PCA principal components) as independent variables. If MLM is selected, introduce random effects (such as family structure or population stratification effects) on the basis of GLM and analyze using a mixed-effects model. (3) Perform GWAS analysis using statistical methods, calculate the SNP effect value and P value for each SNP, and screen for significantly associated loci (usually P < 5 × 10⁻⁶). 8 (This is the genome-wide significance threshold).
[0042] In one implementation, for the aforementioned step S103, i.e., when obtaining exposure SNP data associated with the target exposure factor through quantitative trait locus analysis of proteins based on whole-genome SNP data and proteomic data, the following methods may be used, including but not limited to: First, the proteomic data is standardized to obtain a standardized plasma or tissue protein expression matrix.
[0043] In practice, the proteomic data is first cleaned, including at least: removing missing values (removing proteins with expression levels below the detection limit or excessively high deletion rates); filtering low-variability proteins (retaining proteins with significant expression differences between different samples); and removing batch effects (correcting batch-to-batch systematic errors using methods such as Combat and Mean-based ratios). Then, data standardization is performed to eliminate dimensional differences and make the expression levels of different proteins comparable. Common methods include: centering (subtracting the mean from the expression level of each protein); scaling (dividing the expression level of each protein by its standard deviation (Z-score standardization); min-max standardization (scaling the expression level to the [0,1] interval); and robust standardization (scaling using the median and quartile ranges to reduce the impact of outliers). Finally, data transformation is performed, such as taking the logarithm of the expression levels (e.g., log2(x+1)), to obtain the standardized plasma or tissue protein expression matrix.
[0044] Then, the whole genome SNP data were filtered, and the standardized plasma or tissue protein expression matrix was normalized.
[0045] In practice, the data are first ensured to be on the same scale. Then, batch effects are estimated using the mean and variance between and within batches. Finally, the data are adjusted to remove batch effects. Specifically, whole-genome SNP data are filtered for SNP deletion rate of 0.1, sample deletion rate of 0.05, and minor allele frequency of 0.01. Multi-allele loci and sex chromosomes are also filtered out. Simultaneously, the standardized plasma or tissue protein expression matrix is normalized.
[0046] Finally, protein quantitative trait locus analysis was performed on the filtered whole-genome SNP data and the standardized plasma or tissue protein expression matrix to obtain exposure SNP data associated with the target exposure factor.
[0047] In practice, pQTL analysis is performed on the filtered data to obtain SNP data related to protein expression, that is, exposure SNP data associated with the target exposure factor.
[0048] In one implementation, for the aforementioned step S104, i.e., when performing data preprocessing on the outcome SNP data and the exposed SNP data, the following methods may be used, including but not limited to: (1) Calculate the statistical indicators for outcome SNP data and exposure SNP data; among which, the statistical indicators include at least: P value, F test value and R2 value.
[0049] (2) Remove exposure SNP data with P-values greater than the first threshold and outcome SNP data with P-values less than the first threshold. Specifically, perform correlation analysis on outcome SNP data and exposure SNP data, retaining only those with P-values ≤ 5 × 10⁻⁶. -8 Exposure SNP data and Pvalue ≥ 5 × 10 -8 The final SNP data, i.e., removing P-values greater than the first threshold (5 × 10⁻⁶). -8 Exposed SNP data and removal of P-values less than the first threshold (5 × 10⁻⁶) -8 The final SNP data.
[0050] (3) Remove outcome SNPs or exposure SNPs within a preset range whose R² value with the target outcome SNPs or target exposure SNPs is greater than the second threshold. Specifically, remove linkage imbalances by removing outcome SNPs or exposure SNPs within a 10000kb range (i.e., the preset range) whose R² value with the most significant SNP (i.e., the target outcome SNPs or target exposure SNPs) is greater than 0.001 (the second threshold).
[0051] (4) Remove exposure SNP data with F values greater than the third threshold. Specifically, remove weak instrumental variables, that is, remove exposure SNP data with F values > 10 (i.e., the third threshold) by calculating R2 values and F test values.
[0052] (5) Remove outcome SNP data with allele frequencies less than the fourth threshold. Specifically, filter outcome data and retain only outcome SNP data with allele frequencies greater than or equal to 0.01 (the fourth threshold), that is, remove outcome SNP data with allele frequencies less than the fourth threshold.
[0053] (6) Integrate the format and structure of outcome SNP data and exposed SNP data. Specifically, coordinate and integrate exposed SNP data and outcome SNP data to ensure that they have the same format and structure, such as: coordinating SNP directions, removing palindromic SNPs whose direction cannot be determined, and incompatible SNPs.
[0054] In one implementation, for the aforementioned step S104, i.e., when performing Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype, the following methods may be used, including but not limited to: first, performing Mendelian randomization analysis using a preset regression method to obtain the estimated effect value, standard error of the estimated effect value, and p-value of the relationship between the exposure SNP data and the target exposure factor or target phenotype; then, determining the causal relationship between the target exposure factor and the target phenotype based on the estimated effect value, standard error of the estimated effect value, confidence interval, and p-value.
[0055] In practice, Mendelian randomization analysis is performed using five regression methods (MR Egger, Weighted median, Inverse variance weighted, Simple mode, and Weighted mode; the most important factor is the p-value of the Inverse variance weighted method). This yields the estimated effect size (representing the strength and direction of the association between the SNP data and the exposure (target exposure factor) or outcome (target phenotype); the standard error of the estimated effect size (standard error is a key variable in calculating confidence intervals and p-values); and the p-value for statistical significance (the p-value is used to assess the statistical significance of the relationship between the SNP data and the exposure or outcome. The smaller the p-value, the more likely the relationship between the exposure or outcome is to be real).
[0056] In one implementation, after performing Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype, the method further includes: performing heterogeneity tests and level pleiotropy tests on the exposure SNP data based on the estimated effect value and the p-value of the estimated effect value; and visualizing the exposure SNP data and the estimated effect value.
[0057] In practice, to ensure the accuracy of Mendelian randomization results, sensitivity analysis and visualization analysis can be performed on the exposed SNP data (i.e., instrumental variables) used. Specifically, this includes: (1) Heterogeneity test: This test is mainly used to assess whether there is heterogeneity in the exposure SNP data used. The p-value is significantly smaller than the commonly used significance level (0.05). Heterogeneity refers to the fact that different SNPs may have different relationships between exposure and outcome, which may affect the reliability and interpretation of MR analysis.
[0058] Although some heterogeneity may be observed in the analysis, the presence of heterogeneity does not have the same serious impact as when using a fixed-effects model, because the random-effects model (inverse variance weighted, IVW) is mainly used. The random-effects model can take into account the heterogeneity between studies and assign a specific weight to each study, thereby reducing the impact of heterogeneity on the analysis results.
[0059] (2) Horizontal pleiotropy test: This method is mainly used to detect whether the exposed SNP data has an effect on multiple outcomes. Pleiotropy refers to the phenomenon that a SNP has an effect on multiple phenotypes or traits. The horizontal pleiotropy test is mainly used to detect and evaluate this potential pleiotropy. If the variable tool does not affect the outcome through exposure, this violates Mendel's hypothesis, indicating that horizontal pleiotropy exists.
[0060] The intercept or bias term in the MR Egger method is used to detect level pleiotropy. Generally, if this value is significantly non-zero, it may indicate the presence of level pleiotropy. Simply put, when the intercept is non-zero, meaning the exposure has an effect on the outcome (i.e., the effect size is zero), this effect may originate from other confounding factors, thus indicating the presence of level pleiotropy.
[0061] (3) Visual analysis Scatter plot: Used to illustrate the relationship between genetic variation and exposure factors. Each point represents a Special National Product (SNP), with the horizontal axis representing the effect size of the SNP and the vertical axis representing the effect size of the exposure factor. If a causal relationship exists between the exposure factor and the outcome variable, there is a linear relationship between the effect sizes of the SNP and the exposure factor, and the slope represents the estimated causal effect.
[0062] Forest plot: Used to display the causal effect estimate and confidence interval for each SNP, mainly focusing on the overall effect. Each entry represents a SNP, the x-axis value of which represents the effect size estimate for that SNP, and the line segment represents the confidence interval. If the line segment crosses 0, it indicates that the effect size is not significant.
[0063] Leave-one-out plot: Each time, a SNP is removed from the dataset, and then a subset of the remaining SNPs is used for MR analysis to obtain an effect estimate without that SNP. This process is performed once for each SNP in the dataset. Finally, the effect estimates after the leave-one-out method are compared with the effect estimates of the complete dataset to assess the contribution and stability of each SNP to the overall MR effect.
[0064] Each entry represents the MR analysis result of the subset of SNPs after removing that SNP, with the corresponding x-axis representing its effect size estimate (i.e., the effect size estimate after leaving one out). The impact of each SNP on the MR analysis results is assessed; if outliers are found, they need to be removed and the analysis repeated.
[0065] Leave-one-out cross-validation provides a method to assess the stability and contribution of each SNP to the overall MR effect. The leave-one-out plot visually illustrates the impact of each SNP on the overall effect, identifies potentially "sensitive" SNPs, and evaluates the stability of the overall MR effect.
[0066] Funnel plot: Used to examine publication bias or other biases in research results; essentially, it's a visualization of heterogeneity. If bias exists, meaning there is heterogeneity, a funnel plot can show the relationship between effect size and its precision (usually expressed as standard error or confidence interval width).
[0067] In a funnel plot, each point represents the effect size of a single SNP. The horizontal axis typically represents the effect size, and the vertical axis represents precision (e.g., standard error or confidence interval width). Ideally, the funnel plot should be symmetrical, meaning that the precision of the effect size (e.g., standard error or confidence interval width) is independent of the effect size. If the funnel plot exhibits a degree of asymmetry, it may indicate heterogeneity or other biases, requiring further consideration of the existence and possible causes of heterogeneity.
[0068] In addition, pay attention to outliers that deviate significantly from other points; these points may require special attention. If all points are at the bottom of the funnel but not at the top or middle, this may indicate publishing bias or other systematic errors.
[0069] The method provided in this invention constructs pQTL analysis based on different detection platforms, eliminates technical differences between platforms, covers the entire chain from proteomics detection to Mendelian randomization analysis, no longer depends on whether there is corresponding data in public databases, and directly performs analysis based on sample data, thereby improving the accuracy and reliability of Mendelian randomization analysis results.
[0070] In addition to the causal inference method based on genetic variation provided in the foregoing embodiments, this invention also provides a causal inference device based on genetic variation, see [link to previous embodiment]. Figure 2 The schematic diagram of a causal inference device based on genetic variation shown includes: The raw data acquisition module 201 is used to acquire whole-genome SNP data and proteomic data of test samples from the target population from the detection platform; The outcome SNP data acquisition module 202 is used to perform genome-wide association analysis on whole-genome SNP data to obtain outcome SNP data associated with the target phenotype; The exposure SNP data acquisition module 203 is used to perform quantitative trait locus analysis of proteins based on whole-genome SNP data and proteome data to obtain exposure SNP data associated with the target exposure factor. The Mendelian randomization analysis module 204 is used to preprocess the outcome SNP data and exposure SNP data, and to perform Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype.
[0071] The causal inference device based on genetic variation provided by the present invention directly utilizes the whole genome SNP data and proteome data of the target population obtained from the detection platform to perform genome-wide association analysis and protein quantitative trait locus analysis, covering the entire chain from proteome detection to Mendelian randomization analysis. It no longer relies on public databases, thereby eliminating technical differences between platforms and improving the accuracy and reliability of Mendelian randomization analysis results.
[0072] In one embodiment, the aforementioned outcome SNP data acquisition module 202 is specifically used for: preprocessing whole-genome SNP data to obtain preprocessed whole-genome SNP data; performing genome-wide association analysis with phenotypic data based on a general linear model (GLM) or a mixed linear model (MLM) to obtain the SNP effect value and significance index of each whole-genome SNP data, and determining the outcome SNP data associated with the target phenotype based on the SNP effect value and significance index.
[0073] In one embodiment, the above-mentioned exposure SNP data acquisition module 203 is specifically used to: standardize proteomic data to obtain a standardized plasma or tissue protein expression matrix; filter whole-genome SNP data and normalize the standardized plasma or tissue protein expression matrix; and perform protein quantitative trait locus analysis on the filtered whole-genome SNP data and the standardized plasma or tissue protein expression matrix to obtain exposure SNP data associated with the target exposure factor.
[0074] In one embodiment, the Mendelian randomization analysis module 204 is specifically used for: calculating statistical indicators of outcome SNP data and exposure SNP data; wherein the statistical indicators include at least: p-value, F-test value, and R² value; removing exposure SNP data with p-values greater than a first threshold and removing outcome SNP data with p-values less than the first threshold; removing outcome SNP data or exposure SNP data within a preset range whose R² value with respect to target outcome SNP data or target exposure SNP data is greater than a second threshold; removing outcome SNP data and exposure SNP data with F-values greater than a third threshold; removing outcome SNP data with allele frequencies less than a fourth threshold; and integrating the format and structure of outcome SNP data and exposure SNP data.
[0075] In one implementation, the Mendelian randomization analysis module 204 is specifically used to: perform Mendelian randomization analysis using a preset regression method to obtain the estimated effect value, standard error of the estimated effect value, and p-value of the relationship between the exposed SNP data and the target exposure factor or target phenotype; and determine the causal relationship between the target exposure factor and the target phenotype based on the estimated effect value, standard error of the estimated effect value, confidence interval, and p-value.
[0076] In one embodiment, the above-mentioned apparatus further includes a testing module for: performing heterogeneity testing and level pleiotropy testing on the exposure SNP data based on the estimated effect value and the P-value of the estimated effect value; and visualizing the exposure SNP data and the estimated effect value.
[0077] It should be noted that the device provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment. The specific numerical values provided in this embodiment are merely exemplary and are not intended to limit the scope of the invention.
[0078] This invention also provides an electronic device, specifically, the electronic device includes a processor and a storage device; the storage device stores a computer program, and the computer program, when run by the processor, executes the method described in any of the above embodiments.
[0079] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 30, a memory 31, a bus 32 and a communication interface 33. The processor 30, the communication interface 33 and the memory 31 are connected through the bus 32. The processor 30 is used to execute executable modules, such as computer programs, stored in the memory 31.
[0080] The memory 31 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 33 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0081] Bus 32 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0082] The memory 31 is used to store programs. After receiving an execution instruction, the processor 30 executes the program. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 30 or implemented by the processor 30.
[0083] Processor 30 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 30 or by instructions in software form. Processor 30 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 31. The processor 30 reads the information in memory 31 and, in conjunction with its hardware, completes the steps of the above method.
[0084] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.
[0085] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0086] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A causal inference method based on genetic variation, characterized in that, include: Obtain whole-genome SNP data and proteomic data of test samples from the target population from the testing platform; Genome-wide association analysis was performed on the whole-genome SNP data to obtain outcome SNP data associated with the target phenotype; Based on the whole-genome SNP data and the proteome data, protein quantitative trait locus analysis was performed to obtain exposure SNP data associated with the target exposure factor; The outcome SNP data and the exposure SNP data were preprocessed, and Mendelian randomization analysis was performed on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype.
2. The method according to claim 1, characterized in that, Genome-wide association analysis was performed on the genome-wide SNP data to obtain outcome SNP data associated with the target phenotype, including: The whole-genome SNP data were preprocessed to obtain preprocessed whole-genome SNP data; Genome-wide association analysis was performed on phenotypic data using a general linear model (GLM) or a mixed linear model (MLM) to obtain the SNP effect value and significance index for each genome-wide SNP. Based on the SNP effect value and significance index, the outcome SNP data associated with the target phenotype were determined.
3. The method according to claim 1, characterized in that, Based on the whole-genome SNP data and the proteomic data, quantitative trait locus analysis of proteins was performed to obtain exposure SNP data associated with the target exposure factor, including: The proteomic data were standardized to obtain a standardized plasma or tissue protein expression matrix; The whole-genome SNP data were filtered, and the standardized plasma or tissue protein expression matrix was normalized. Protein quantitative trait locus analysis was performed on the filtered whole-genome SNP data and the standardized plasma or tissue protein expression matrix to obtain exposure SNP data associated with the target exposure factor.
4. The method according to claim 1, characterized in that, Data preprocessing is performed on the outcome SNP data and the exposure SNP data, including: Calculate statistical indicators for the outcome SNP data and the exposure SNP data; wherein the statistical indicators include at least: p-value, F-test value, and R2 value; Remove exposed SNP data with P-values greater than the first threshold and remove outcome SNP data with P-values less than the first threshold; Remove SNP data or exposure SNP data within a preset range whose R2 value is greater than the second threshold compared to the target outcome SNP data or target exposure SNP data; Remove exposed SNP data with an F-value greater than the third threshold; Remove outcome SNP data with allele frequencies below the fourth threshold; The format and structure of the integrated outcome SNP data and the exposed SNP data are determined.
5. The method according to claim 1, characterized in that, Mendelian randomization analysis was performed on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype, including: Mendelian randomization analysis was performed using a pre-defined regression method to obtain the estimated effect size, standard error of the estimated effect size, and p-value of the relationship between the exposure SNP data and the target exposure factor or the target phenotype. The causal relationship between the target exposure factor and the target phenotype is determined based on the estimated effect size, the standard error of the estimated effect size, the confidence interval, and the p-value.
6. The method according to claim 5, characterized in that, After performing Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype, the analysis also includes: Based on the estimated effect size and the p-value of the estimated effect size, heterogeneity test and level pleiotropy test are performed on the exposure SNP data; The exposure SNP data and the estimated effect value are visualized.
7. A causal inference device based on genetic variation, characterized in that, include: The raw data acquisition module is used to acquire whole-genome SNP data and proteomic data of test samples from the target population from the testing platform; The outcome SNP data acquisition module is used to perform genome-wide association analysis on the whole genome SNP data to obtain outcome SNP data associated with the target phenotype; The exposure SNP data acquisition module is used to perform protein quantitative trait locus analysis based on the whole genome SNP data and the proteome data to obtain exposure SNP data associated with the target exposure factor; The Mendelian randomization analysis module is used to preprocess the outcome SNP data and the exposure SNP data, and to perform Mendelian randomization analysis on the preprocessed outcome SNP data and exposure SNP data to obtain the causal relationship between the target exposure factor and the target phenotype.
8. The apparatus according to claim 7, characterized in that, The SNP data acquisition module is specifically used for: The whole-genome SNP data were preprocessed to obtain preprocessed whole-genome SNP data; Genome-wide association analysis was performed on phenotypic data using a general linear model (GLM) or a mixed linear model (MLM) to obtain the SNP effect value and significance index for each genome-wide SNP. Based on the SNP effect value and significance index, the outcome SNP data associated with the target phenotype were determined.
9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of the method described in any one of claims 1 to 6.