Method and system for automated inference of cross-ethnic lifestyle habit and disease causation

CN116312806BActive Publication Date: 2026-09-25SHENZHEN ARTS CHANGHUA INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310289368.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-16
Publication Date
2026-09-25
Estimated Expiration
2043-03-16

AI Technical Summary

Technical Problem

在现实情况下,存在小样本的关联关系统计量,直接使用上述提及的模型会导致结果的准确度不可信,可解释性不足

Benefits of technology

[0017]与现有技术相比,本发明上述技术方案,基于孟德尔随机化模型对获得的统计数据进行分析,基于大样本人种的大量数据获得的荟萃分析结果与小样本人种的孟德尔随机化分析结果融合,进而对融合后的数据进行荟萃分析,并基于预设条件判断当前分析的生活习惯与指定疾病是否有因果关系,由此可知,通过上述方案,由于大样本人种数据的参与,可增加小样本人种因果分析的准确度和可解释性,另外,还根据单核苷酸多态性的显著性对获得的小样本人种的生活习惯相关的关联关系统计量进行清洗,以筛除掉具有多效性的遗传变异,只使用有效的遗传变异进行后续的孟德尔分析,从而进一步提升分析结果的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312806B_ABST
    Figure CN116312806B_ABST
Patent Text Reader

Abstract

The application discloses a kind of cross ethnic life habit and disease causal relationship automated inference method and system, wherein the method comprises: obtaining the corresponding associated statistical measurement of large sample ethnic group and small sample ethnic group, and the statistical data obtained is cleaned;The associated statistical measurement of large sample ethnic group is matched as a whole, to obtain multiple pairs of target data pairs;Multiple pairs of target data pairs are subjected to Mendelian randomization analysis to obtain a first randomization result;The first randomization result is subjected to meta-analysis to obtain a first analysis result;The associated statistical measurement of small sample ethnic group is subjected to Mendelian randomization analysis to obtain a second randomization result;When each of the above processing results meets the preset condition, the current analysis of small sample ethnic group is output. The life habit and the specified disease have causal relationship;The above method can effectively increase the accuracy and explainability of the causal analysis between the life system of small sample ethnic group and disease.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of causal inference technology in bioinformatics and genetics, and in particular to an automated inference method and system for cross-ethnic causal relationships of lifestyle habits and diseases based on Mendelian randomization. Background Technology

[0002] Causal inference between variables is a common problem in biology, especially the causal relationship between lifestyle habits and complex diseases. If the causal link between the two can be obtained, people will be able to intervene in certain unhealthy lifestyle habits in advance, thereby reducing the risk of disease.

[0003] In this regard, simple regression analysis can only obtain the correlation between variables but cannot determine the causal relationship because: ① the correlation coefficient cannot determine whether the causal direction is A->B or B->A; ② there may be an unobserved (confounding) factor C that is related to both variables, leading to A->C->B.

[0004] Randomized controlled trials (RCTs) are considered the gold standard for evaluating causal effects. Their basic method involves randomly assigning participants to different groups and administering different interventions to each group, measuring the effectiveness of these interventions under these stringent conditions. When a sufficient number of participants are available, this method can offset the influence of known and unknown confounding factors on each group. However, due to ethical constraints, participant compliance, and study duration, RCTs are difficult to implement. Furthermore, limitations in inclusion and exclusion criteria in RCTs can lead to heterogeneity between the research sample and the real-world population, thus raising questions about the validity of the conclusions drawn.

[0005] In contrast, observational studies and non-randomized controlled trials have more readily available data and their sample selection is closer to real-world conditions. However, both observational and non-randomized controlled trials require appropriate computational models to infer the causal relationship between exposure and outcome variables. For this, the Mendelian randomization model is widely used in the field. Mendelian randomization is a statistical model that uses genetic variation as an instrumental variable (commonly single nucleotide polymorphisms). Its basic results are shown in the attached figure. Figure 3 As shown, Z: instrumental variable; U: confounding variable; X: exposure variable; Y: outcome variable. The Mendelian randomization model works by stating that if the instrumental variable is correlated with the exposure variable but not with any confounding variable correlated with the outcome variable, and the instrumental variable has no other way of influencing the outcome variable besides through the exposure variable, then the causal effect of the exposure variable on the outcome variable can be estimated.

[0006] Currently existing Mendelian analysis models include single-sample and two-sample models. For single-sample models, a single genetic variation is used as an instrumental variable to estimate the causal effect of the exposure variable on the outcome variable. For two-sample models, association statistics between genetic variation and the exposure variable, as well as association statistics between genetic variation and the outcome variable, are derived from two independent, non-overlapping samples. In reality, association statistics exist in small samples, and directly using the aforementioned models can lead to unreliable accuracy and insufficient interpretability of the results. Summary of the Invention

[0007] The purpose of this invention is to provide an automated inference method for cross-ethnic causal relationships between lifestyle habits and diseases by using Mendelian randomization to analyze small sample populations, thereby improving the accuracy and interpretability of the causal relationships found between lifestyle habits and diseases.

[0008] To achieve the above objectives, this invention discloses an automated method for inferring the causal relationship between cross-ethnic lifestyle habits and diseases, comprising: Obtain association statistics, wherein the association data includes a first association statistic of the association between single nucleotide polymorphisms (SNPs) of one or more large-sample ethnic groups and a specified lifestyle, a second association statistic of the association between single nucleotide polymorphisms of two or more large-sample ethnic groups and a specified disease, a third association statistic of the association between single nucleotide polymorphisms of one small-sample ethnic group and a specified lifestyle, and a fourth association statistic of the association between single nucleotide polymorphisms of one small-sample ethnic group and a specified disease; Determine whether the significance of each single nucleotide polymorphism in the first association statistic and the third association statistic exceeds a preset threshold. If yes, the single nucleotide polymorphism is screened out; otherwise, it is retained. Pair the primary and secondary association statistics belonging to a large sample of races into one-to-many or many-to-many pairs to obtain multiple target data pairs. Based on the Mendelian randomization model, Mendelian randomization analysis was performed on multiple pairs of the target data to obtain multiple sets of first randomization results; Meta-analysis was performed on multiple sets of the first randomization results to obtain the first analysis results; Based on the Mendelian randomization model, Mendelian randomization analysis was performed on the third association statistic and the fourth association statistic to obtain the second randomization result; The second randomization result and the first analysis result are treated as a whole and subjected to another meta-analysis to obtain the second analysis result; When multiple sets of the first randomization results, the first analysis results, the second randomization results, and the second analysis results all meet the preset conditions, the output shows that the current lifestyle habits of a small sample population are causally related to the specified disease.

[0009] Preferably, the preset conditions include: For multiple sets of the first randomization results, at least one set has a significance level of less than 0.05; For the first analysis results, the fixed effect significance is less than 0.05, the square of I is less than 0.05, and the equivalent significance is greater than 0.05; For the second randomization result, the significance is less than 0.05; For the second analysis results, the significance of the fixed effect is less than 0.05.

[0010] Preferably, the method for obtaining the third association statistic is as follows: To obtain raw lifestyle and genotypic data of a small sample of ethnic groups; Remove occasional lifestyle-related data items from the original lifestyle data; Based on the total number of samples from a small sample group in the original lifestyle data, lifestyle-related data items with a missing number greater than or equal to a preset value are filtered out. The missing number is the number of people in the total number of samples that are not currently covered by the lifestyle data. For lifestyle habits with fewer missing values ​​than the preset value, if the lifestyle habit is numerical data, the median method is used to fill in the relevant data items; if the lifestyle habit is categorical data, the highest frequency method is used to fill in the relevant data items to obtain the target lifestyle habit data. Genome-wide association analysis tools were used to calculate and analyze genotype data of a small sample of ethnic groups and the target lifestyle data to obtain the third association statistic.

[0011] Preferably, the intersection of the lifestyle habits in the first association statistics and the third association statistics is also calculated to filter out data items related to non-public lifestyle habits in the first association statistics and the third association statistics.

[0012] This invention also discloses an automated inference system for cross-ethnic lifestyle habits and disease causal relationships, comprising: The basic data acquisition module is used to acquire correlation statistics, which include the first correlation statistics of the association between single nucleotide polymorphisms of one or more large sample ethnic groups and specified lifestyle habits, the second correlation statistics of the association between single nucleotide polymorphisms of two or more large sample ethnic groups and specified diseases, the third correlation statistics of the association between single nucleotide polymorphisms of one small sample ethnic group and specified lifestyle habits, and the fourth correlation statistics of the association between single nucleotide polymorphisms of one small sample ethnic group and specified diseases. The data cleaning module is used to screen out single nucleotide polymorphisms whose significance exceeds a preset threshold in the first association statistics and the third association statistics. The pairing module is used to perform one-to-many or many-to-many pairing of the first association statistics and the second association statistics belonging to a large sample of races, in order to obtain multiple pairs of target data. The first data processing module is used to perform Mendelian randomization analysis on multiple pairs of target data based on the Mendelian randomization model to obtain multiple sets of first randomization results; The second data processing module is used to perform meta-analysis on multiple sets of the first randomization results to obtain the first analysis result. The third data processing module is used to perform Mendelian randomization analysis on the third association statistic and the fourth association statistic based on the Mendelian randomization model to obtain the second randomization result; The fourth data processing module is used to perform meta-analysis again on the second randomization result and the first analysis result as a whole to obtain the second analysis result; The confirmation module is used to compare multiple sets of the first randomization results, the first analysis results, the second randomization results, and the second analysis results with preset conditions to confirm whether there is a causal relationship between the current lifestyle habits of a small sample population and the specified disease.

[0013] Preferably, the preset conditions include: For multiple sets of the first randomization results, at least one set has a significance level of less than 0.05; For the first analysis results, the fixed effect significance is less than 0.05, the square of I is less than 0.05, and the equivalent significance is greater than 0.05; For the second randomization result, the significance is less than 0.05; For the second analysis results, the significance of the fixed effect is less than 0.05.

[0014] Preferably, it also includes an optimization module, which is used to find the intersection of the lifestyle habits in the first association statistics and the third association statistics, so as to filter out non-public lifestyle-related data items in the first association statistics and the third association statistics.

[0015] This invention also discloses an automated inference system for cross-ethnic lifestyle habits and disease causal relationships, comprising: One or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including instructions for performing the automated inference method for cross-ethnic lifestyle and disease causality as described above.

[0016] The present invention also discloses a computer-readable storage medium comprising a computer program that can be executed by a processor to perform the method described above for automating the inference of cross-ethnic lifestyle and disease causal relationships.

[0017] Compared with existing technologies, the above-mentioned technical solution of the present invention analyzes the obtained statistical data based on the Mendelian randomization model, integrates the meta-analysis results obtained from a large sample of ethnic groups with the Mendelian randomization analysis results from a small sample of ethnic groups, and then performs meta-analysis on the integrated data. Based on preset conditions, it determines whether there is a causal relationship between the current lifestyle habits and the specified disease. Thus, it can be seen that, through the above solution, the accuracy and interpretability of causal analysis of small sample ethnic groups can be increased due to the participation of large sample ethnic group data. In addition, the statistical quantities of associations related to lifestyle habits of small sample ethnic groups are cleaned according to the significance of single nucleotide polymorphisms to screen out pleiotropic genetic variations, and only valid genetic variations are used for subsequent Mendelian analysis, thereby further improving the accuracy of the analysis results. Attached Figure Description

[0018] Figure 1 This is a structural diagram illustrating the principle of automated causal relationship inference in an embodiment of the present invention.

[0019] Figure 2 This is a flowchart of the automated causal relationship inference method in an embodiment of the present invention.

[0020] Figure 3 This is a basic structural diagram of the Mendelian randomized model. Detailed Implementation

[0021] To illustrate the technical content, structural features, objectives, and effects of the present invention in detail, the following description is provided in conjunction with the embodiments and accompanying drawings.

[0022] This embodiment discloses an automated inference method for cross-ethnic causal relationships between lifestyle habits and diseases based on Mendelian randomization. Specifically, it uses a Mendelian randomization model to process statistically obtained sample data to identify lifestyle systems causally related to complex diseases. This embodiment utilizes statistical measures of the association between genetic variation and lifestyle habits, and between genetic variation and disease, from a large sample population to assist a small sample population, improving the accuracy and interpretability of the results. In this embodiment, it is assumed that the causal relationship between lifestyle habits and complex diseases does not vary significantly across different ethnic groups.

[0023] like Figure 1 and Figure 2 The inference method in this embodiment includes the following steps: S1: Obtain association statistics. The association data includes the first association statistics of single nucleotide polymorphisms (SNPs) in one or more large-sample populations and their association with specified lifestyle habits; the second association statistics of single nucleotide polymorphisms in two or more large-sample populations and their association with specified diseases; the third association statistics of single nucleotide polymorphisms in one small-sample population and their association with specified lifestyle habits; and the fourth association statistics of single nucleotide polymorphisms in one small-sample population and their association with specified diseases.

[0024] S2: Clean the obtained statistical data, that is, determine whether the significance of each single nucleotide polymorphism in the first association statistic and the third association statistic exceeds the preset threshold. If it does, the single nucleotide polymorphism is screened out; otherwise, it is retained.

[0025] S3: Perform one-to-many or many-to-many pairings on the first and second association statistics belonging to a large sample of ethnicities to obtain multiple target data pairs. For example, if one first association statistic is obtained, forming file A, and two second association statistics are obtained, forming files B and C, then there are two target data pairs formed by pairing: A&B and A&C.

[0026] S4: Based on the Mendelian randomization model, perform Mendelian randomization analysis on multiple pairs of target data to obtain multiple sets (two sets in this embodiment) of first randomization results for multiple large-sample ethnic groups.

[0027] S5: Perform a meta-analysis on the above two groups of first randomization results to obtain the first analysis results.

[0028] S6: Based on the Mendelian randomization model, Mendelian randomization analysis was performed on the third and fourth association statistics to obtain a second randomization result for a small sample of ethnic groups.

[0029] S7: Perform a meta-analysis again on the second randomization results and the first analysis results as a whole to obtain the second analysis results.

[0030] S8: Determine whether the two sets of first randomization results, first analysis results, second randomization results, and second analysis results simultaneously meet the preset conditions. If so, output that the current lifestyle habits of the small sample population are causally related to the specified disease. Otherwise, if any of the first randomization results, first analysis results, second randomization results, and second analysis results does not meet the preset conditions, it indicates that the causal relationship between the current lifestyle habits of the small sample population and the specified disease is highly questionable.

[0031] Furthermore, the preset conditions include the following: For multiple sets of first randomization results, at least one set has a significance level of less than 0.05; For the results of the first analysis, the fixed effect significance is less than 0.05, the square of I is less than 0.05, and the equivalent significance is greater than 0.05; For the second randomization result, the significance is less than 0.05; For the second analysis results, the significance of the fixed effect is less than 0.05.

[0032] Furthermore, the method for obtaining the third correlation statistic is as follows: (1) Obtain raw lifestyle and genotype data of a small sample of ethnic groups; (2) Remove occasional lifestyle-related data items from the original lifestyle data; (3) Based on the total number of samples of small sample ethnic groups in the original lifestyle data, filter out lifestyle-related data items with missing numbers greater than or equal to the preset value. The missing number is the number of people in the total number of samples not covered by the current lifestyle. For example, if the total number of samples of small sample ethnic groups is 3099, and only 309 or fewer people have the current lifestyle, then the lifestyle will be filtered out. (4) For lifestyle-related data items with a missing number less than the preset value, if the lifestyle is numerical data, the median method is used to fill the data; if the lifestyle is categorical data, the most frequent method is used to fill the data, so as to obtain the target lifestyle data.

[0033] (5) Use genome-wide association analysis tools to calculate and analyze the genotype data and target lifestyle data of a small sample of ethnic groups to obtain the third association statistics.

[0034] Furthermore, to improve the correlation between large and small sample data, the intersection of lifestyle habits in the first and third association statistics is calculated to filter out data items related to non-public lifestyle habits in the first and third association statistics.

[0035] Specifically, the following example illustrates the use of the above inference method, including the data preparation stage and the causal inference stage.

[0036] The data preparation phase includes the following steps: 1.1) Download two large-sample statistical analyses of the association between single nucleotide polymorphisms (SNPs) and coronary heart disease, and two large-sample statistical analyses of the association between SNPs and type 2 diabetes in Europeans (in this example, there are two designated diseases, namely coronary heart disease and type 2 diabetes), from "https: / / www.ebi.ac.uk / gwas / home" (GWASCatalog database). Download one large-sample statistical analysis of the association between SNPs and lifestyle habits in Europeans from "https: / / gwas.mrcieu.ac.uk / " (openGWAS database) (a total of 268 different lifestyle data were downloaded).

[0037] 1.2) Lifestyle and genotypic data of East Asians were extracted from the UK Biobank, and the extracted lifestyle data underwent data preprocessing. First, occasional lifestyle-related data items were removed, such as whether alcohol was consumed or other unusual foods were eaten yesterday. Then, since the extracted East Asian sample size was 3099, lifestyle-related data items with more than 309 missing values ​​were removed. Next, missing values ​​were processed: for numerical data, median values ​​were used for imputation; for categorical data, the highest frequency values ​​were used. Ultimately, 120 lifestyle habits were retained.

[0038] In addition, in this embodiment, the number of East Asians with coronary heart disease and the number of East Asians with type 2 diabetes were extracted from the UK Biobank were 78.

[0039] 1.3) Download a genome-wide association analysis tool (such as PLINK software) from "https: / / www.cog-genomics.org / plink / ", and use the genotype data of East Asians in the UK Biobank and the data extracted in step 2) above to calculate the statistical relationships between single nucleotide polymorphisms (SNPs) and lifestyle habits, as well as the statistical relationships between SNPs and coronary heart disease and type 2 diabetes.

[0040] 1.4) The intersection of East Asian lifestyle habits extracted from the UK Biobank and European lifestyle habits downloaded (76 in total) was analyzed, and subsequent causal inferences were made on the overlapping lifestyle habits.

[0041] In the causal inference stage, the first step is to infer the causal relationship between cheese consumption frequency and coronary heart disease, including the following steps: 2.1) The association statistics between single nucleotide polymorphisms (SNPs) and cheese consumption frequency in large-sample Europeans and small-sample East Asians, and the association statistics between SNPs and coronary heart disease, were processed into data packages; 2.2) Data cleaning was performed on the statistical correlation between single nucleotide polymorphisms and cheese consumption frequency in the data package formed in step 2.1). The cleaning method is as follows: Single nucleotide polymorphisms that are strongly correlated with the exposure variable are extracted, and a threshold for the significance P-value is set. Single nucleotide polymorphisms that are higher than the threshold will be screened out. The "clump_data" method in the "TwoSampleMR" package of the R language is used to ensure that the remaining single nucleotide polymorphisms are independent of each other and to screen out single nucleotide polymorphisms with a minor allele frequency greater than 0.01. Using the R language package "PhenoScanner", we can detect whether a single nucleotide polymorphism is strongly correlated with multiple exposure variables by setting a threshold for the significance P-value, and then screen it out if it is.

[0042] 2.3) First, the association statistics between a single nucleotide polymorphism (SNP) and cheese intake in Europeans were paired with the association statistics between two SNPs and coronary heart disease to obtain two target data pairs. Then, Mendelian randomization analysis was performed on the two target data pairs using a Mendelian randomization model to obtain two sets of first randomization results. Then, a meta-analysis was performed on the two groups of first randomization results to obtain the first analysis results; Next, Mendelian randomization analysis was performed on the association statistics between single nucleotide polymorphisms and lifestyle habits and the association statistics between single nucleotide polymorphisms and coronary heart disease in East Asians to obtain the second randomization results. Then, a meta-analysis was performed using the meta-analysis results of Europeans (first analysis results) and the Mendelian randomization results of East Asians (second randomization results) to obtain the second analysis results.

[0043] 2.4) Result Judgment. Determine whether the two sets of first randomization results, first analysis results, second randomization results, and second analysis results in step 2.3) above meet the above preset conditions. If so, output the confirmation result, that is, for East Asians, the frequency of cheese intake is causally related to coronary heart disease.

[0044] Then, inferring the causal relationship between poultry consumption frequency and type 2 diabetes involves the following steps: 3.1) The association statistics between single nucleotide polymorphisms (SNPs) and poultry intake frequency in large-sample Europeans and small-sample East Asians, and the association statistics between SNPs and type 2 diabetes, were processed into data packages; 3.2) Perform data cleaning on the data packets formed in step 2.1) above. The cleaning method is as described above and will not be repeated here.

[0045] 3.3) First, the association statistics of one single nucleotide polymorphism (SNP) with poultry intake in Europeans were paired with the association statistics of two SNPs with type 2 diabetes to obtain two target data pairs. Then, Mendelian randomization analysis was performed on the two target data pairs using a Mendelian randomization model to obtain two sets of first randomization results. Then, a meta-analysis was performed on the two first randomization results to obtain the first analysis results; Next, Mendelian randomization analysis was performed on the association statistics between single nucleotide polymorphisms and lifestyle habits and the association statistics between single nucleotide polymorphisms and type 2 diabetes in East Asians to obtain the second randomization results. Then, a meta-analysis was performed using the meta-analysis results of Europeans (first analysis results) and the Mendelian randomization results of East Asians (second randomization results) to obtain the second analysis results.

[0046] 3.4) Result Judgment. Determine whether the two sets of first randomization results, first analysis results, second randomization results, and second analysis results in step 3.3) above meet the above preset conditions. If so, output the confirmation result, that is, for East Asians, there is a causal relationship between poultry consumption frequency and type 2 diabetes.

[0047] Following this logic, the analysis ultimately revealed that for East Asians, 46 lifestyle habits are causally related to coronary heart disease, and 33 lifestyle habits are causally related to type 2 diabetes.

[0048] The automated inference method for cross-ethnic causal relationships between lifestyle habits and diseases disclosed in the above embodiments of the present invention analyzes the obtained statistical data based on a Mendelian randomization model, and fuses the meta-analysis results obtained from a large sample of ethnic groups with the Mendelian randomization analysis results from a small sample of ethnic groups. Then, a meta-analysis is performed on the fused data, and a causal relationship between the currently analyzed lifestyle habits and a specified disease is determined based on preset conditions. The participation of large sample ethnic data increases the accuracy and interpretability of causal analysis for small sample ethnic groups. Furthermore, the statistical data related to lifestyle habits obtained from small sample ethnic groups are cleaned based on the significance of single nucleotide polymorphisms to remove pleiotropic genetic variations, using only valid genetic variations for subsequent Mendelian analysis, thereby further improving the accuracy of the analysis results.

[0049] The present invention also discloses an automated inference system for cross-ethnic lifestyle habits and disease causal relationships, which includes a basic data acquisition module, a data cleaning module, a matching module, a first data processing module, a second data processing module, a third data processing module, a fourth data processing module, and a confirmation module.

[0050] The basic data acquisition module is used to obtain association statistics. The association data includes the first association statistics of the association between single nucleotide polymorphisms of one or more large-sample ethnic groups and specified lifestyle habits, the second association statistics of the association between single nucleotide polymorphisms of two or more large-sample ethnic groups and specified diseases, the third association statistics of the association between single nucleotide polymorphisms of one small-sample ethnic group and specified lifestyle habits, and the fourth association statistics of the association between single nucleotide polymorphisms of one small-sample ethnic group and specified diseases.

[0051] The data cleaning module is used to screen out single nucleotide polymorphisms whose significance exceeds a preset threshold in the first association statistics and the third association statistics.

[0052] The pairing module is used to perform one-to-many or many-to-many pairings on the first and second association statistics belonging to a large sample of ethnic groups to obtain multiple pairs of target data.

[0053] The first data processing module is used to perform Mendelian randomization analysis on multiple pairs of target data based on the Mendelian randomization model to obtain multiple sets of first randomization results.

[0054] The second data processing module is used to perform meta-analysis on multiple sets of first randomization results to obtain the first analysis results.

[0055] The third data processing module is used to perform Mendelian randomization analysis on the third and fourth association statistics based on the Mendelian randomization model to obtain the second randomization result.

[0056] The fourth data processing module is used to perform a meta-analysis on the second randomization result and the first analysis result as a whole to obtain the second analysis result.

[0057] The confirmation module is used to compare multiple sets of first randomization results, first analysis results, second randomization results, and second analysis results with preset conditions to confirm whether there is a causal relationship between the current lifestyle habits of a small sample population and the specified disease.

[0058] Specifically, the preset conditions include: For multiple sets of first randomization results, at least one set has a significance level of less than 0.05; For the results of the first analysis, the fixed effect significance is less than 0.05, the square of I is less than 0.05, and the equivalent significance is greater than 0.05; For the second randomization result, the significance is less than 0.05; For the second analysis results, the significance of the fixed effect is less than 0.05.

[0059] Furthermore, the system also includes an optimization module, which is used to find the intersection of lifestyle habits in the first association statistics and the third association statistics in order to filter out data items related to non-public lifestyle habits in the first association statistics and the third association statistics.

[0060] This invention also discloses another automated causal relationship inference system, comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including instructions for performing the automated causal relationship inference method as described above. The processor may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, used to execute the relevant programs to implement the functions required by the modules in the automated causal relationship inference system of this application embodiment, or to execute the automated causal relationship inference method of this application embodiment.

[0061] This invention also discloses a computer-readable storage medium comprising a computer program executable by a processor to perform the automated causal reasoning method described above. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be read-only memory (ROM), random access memory (RAM), or magnetic media, such as floppy disks, hard disks, magnetic tapes, magnetic disks, or optical media, such as digital versatile discs (DVDs), or semiconductor media, such as solid-state disks (SSDs).

[0062] This application also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the aforementioned automated causal relationship inference method.

[0063] The above-disclosed embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of the present invention. Therefore, any equivalent variations made in accordance with the claims of the present invention are still within the scope of the present invention.

Claims

1. An automated method for inferring the causal relationship between cross-ethnic lifestyle habits and diseases, characterized in that, include: Obtain association statistics, wherein the association data includes a first association statistic of the association between single nucleotide polymorphisms (SNPs) of one or more large-sample ethnic groups and a specified lifestyle, a second association statistic of the association between single nucleotide polymorphisms of two or more large-sample ethnic groups and a specified disease, a third association statistic of the association between single nucleotide polymorphisms of one small-sample ethnic group and a specified lifestyle, and a fourth association statistic of the association between single nucleotide polymorphisms of one small-sample ethnic group and a specified disease; Determine whether the significance of each single nucleotide polymorphism in the first association statistic and the third association statistic exceeds a preset threshold. If yes, the single nucleotide polymorphism is screened out; otherwise, it is retained. Pair the primary and secondary association statistics belonging to a large sample of races into one-to-many or many-to-many pairs to obtain multiple target data pairs. Based on the Mendelian randomization model, Mendelian randomization analysis was performed on multiple pairs of the target data to obtain multiple sets of first randomization results; Meta-analysis was performed on multiple sets of the first randomization results to obtain the first analysis results; Based on the Mendelian randomization model, Mendelian randomization analysis was performed on the third association statistic and the fourth association statistic to obtain the second randomization result; The second randomization result and the first analysis result are treated as a whole and subjected to another meta-analysis to obtain the second analysis result; When multiple sets of the first randomization results, the first analysis results, the second randomization results, and the second analysis results all meet the preset conditions, the output shows that the current lifestyle habits of a small sample population are causally related to the specified disease.

2. The automated inference method for cross-ethnic lifestyle habits and disease causal relationships according to claim 1, characterized in that, The method for obtaining the third correlation statistic: To obtain raw lifestyle and genotypic data of a small sample of ethnic groups; Remove occasional lifestyle-related data items from the original lifestyle data; Based on the total number of samples from a small sample group in the original lifestyle data, lifestyle-related data items with a missing number greater than or equal to a preset value are filtered out. The missing number is the number of people in the total number of samples that are not currently covered by the lifestyle data. For lifestyle habits with fewer missing values ​​than the preset value, if the lifestyle habit is numerical data, the median method is used to fill in the relevant data items; if the lifestyle habit is categorical data, the highest frequency method is used to fill in the relevant data items to obtain the target lifestyle habit data. Genome-wide association analysis tools were used to calculate and analyze genotype data of a small sample of ethnic groups and the target lifestyle data to obtain the third association statistic.

3. The automated inference method for cross-ethnic lifestyle habits and disease causal relationships according to claim 1, characterized in that, Furthermore, the intersection of the lifestyle habits in the first association statistics and the third association statistics is calculated to filter out data items related to non-public lifestyle habits in the first association statistics and the third association statistics.

4. An automated inference system for cross-ethnic lifestyle habits and disease causal relationships, characterized in that, include: The basic data acquisition module is used to acquire correlation statistics, which include the first correlation statistics of the association between single nucleotide polymorphisms of one or more large sample ethnic groups and specified lifestyle habits, the second correlation statistics of the association between single nucleotide polymorphisms of two or more large sample ethnic groups and specified diseases, the third correlation statistics of the association between single nucleotide polymorphisms of one small sample ethnic group and specified lifestyle habits, and the fourth correlation statistics of the association between single nucleotide polymorphisms of one small sample ethnic group and specified diseases. The data cleaning module is used to screen out single nucleotide polymorphisms whose significance exceeds a preset threshold in the first association statistics and the third association statistics. The pairing module is used to perform one-to-many or many-to-many pairing of the first association statistics and the second association statistics belonging to a large sample of races, in order to obtain multiple pairs of target data. The first data processing module is used to perform Mendelian randomization analysis on multiple pairs of target data based on the Mendelian randomization model to obtain multiple sets of first randomization results; The second data processing module is used to perform meta-analysis on multiple sets of the first randomization results to obtain the first analysis result. The third data processing module is used to perform Mendelian randomization analysis on the third association statistic and the fourth association statistic based on the Mendelian randomization model to obtain the second randomization result; The fourth data processing module is used to perform meta-analysis again on the second randomization result and the first analysis result as a whole to obtain the second analysis result; The confirmation module is used to compare multiple sets of the first randomization results, the first analysis results, the second randomization results, and the second analysis results with preset conditions to confirm whether there is a causal relationship between the current lifestyle habits of a small sample population and the specified disease.

5. The automated inference system for cross-ethnic lifestyle habits and disease causal relationships according to claim 4, characterized in that, It also includes an optimization module, which is used to find the intersection of the lifestyle habits in the first association statistics and the third association statistics, so as to filter out the non-public lifestyle related data items in the first association statistics and the third association statistics.

6. An automated inference system for cross-ethnic lifestyle habits and disease causal relationships, characterized in that, include: One or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including instructions for performing the automated inference method for cross-ethnic lifestyle and disease causality as described in any one of claims 1 to 3.

7. A computer-readable storage medium, characterized in that, Includes a computer program that can be executed by a processor to perform the automated inference method for cross-ethnic lifestyle and disease causality as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Construction method of pulmonary thromboembolism risk prediction model based on single nucleotide polymorphism, SNP locus combination and application

    CN112553327A

  • Method for evaluating causal relationship between micronutrients and mental diseases based on Mendel randomization

    CN113113141A

  • Machine learning platform for generating risk models

    US20220044761A1