Construction method of complex disease genetic risk assessment model, model and application thereof
By combining whole-genome sequencing and statistical analysis with ensemble learning algorithms, a genetic risk assessment model for complex diseases is constructed, which solves the problems of data scarcity and insufficient generalization ability in existing technologies, and achieves high-precision risk assessment and prediction for complex diseases.
Patent Information
- Application Number
- CN202210861556.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-20
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2042-07-20
AI Technical Summary
Existing technologies lack universally applicable strategies for assessing the risk of complex diseases, making accurate assessment particularly difficult when data is scarce, and existing models have weak generalization capabilities.
By using whole-genome sequencing and statistical analysis, risk genes with significant mutational differences are identified. An integrated learning algorithm is then used to construct a genetic risk assessment model for complex diseases. Considering the genetic background of families and sporadic populations, a multi-index standard is used to screen characteristic genes and construct a nonlinear risk assessment model.
It achieves high-precision risk assessment of complex diseases with small sample sizes, has good generalization ability and predictive effect, and is applicable to the clinical diagnosis and risk prediction of a variety of complex diseases.
Smart Images

Figure CN115295075B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of gene detection, specifically involving the construction method, model, and application of genetic risk assessment models for complex diseases. Background Technology
[0002] Complex diseases are those caused by multiple genes and environmental factors, with a high incidence rate in the population and exhibiting genetic heterogeneity and phenotypic complexity. Clinical studies have shown that genetic factors play a crucial role in the development of complex diseases; therefore, elucidating the genotype-phenotype relationship of complex diseases is helpful in studying their pathogenesis. Due to the complexity of multi-gene interactions, these diseases often lack a clear inheritance pattern, making it difficult to determine genetic characteristics and risk assessment models for clinical diagnosis and treatment.
[0003] For example, venous thromboembolism (VTE) has been shown to be primarily driven by genetic risk and exhibit ethnic specificity. Early risk assessment is a crucial step in VTE prevention. Currently, the CHEST guidelines and the Caprini risk assessment scale are widely used in clinical practice and have been validated to significantly reduce the incidence of deep vein thrombosis and pulmonary embolism in hospitalized patients. However, these risk assessment scales only detect two specific genetic risk genes (FV Leiden mutation and prothrombin G20210A mutation), which is inconsistent with the complex genetic mechanisms of VTE. This results in poor accuracy in clinical application of these scales and makes them unsuitable for capturing the multimodal characteristics of VTE. A deeper understanding of the pathogenesis and etiology of complex diseases requires the construction of multi-gene statistical models for prevention and treatment.
[0004] Currently, in the study of the genetic characteristics of complex diseases, with the development of high-throughput sequencing technology, researchers have proposed using genome-wide association studies (GWAS) to discover disease-associated genes. This combination of epidemiological study design and molecular genetic analysis techniques has made significant contributions to the genetic assessment of complex diseases. However, GWAS analysis based on mixed linear models still struggles to uncover key rare mutations that contribute to the etiology of complex diseases. In risk assessment studies of complex diseases, researchers construct weighted linear risk assessment models by designing test statistics such as OR values. This method is simple to apply, but its assessment accuracy is low. To improve the accuracy of model predictions, with the development of artificial intelligence, researchers have applied various machine learning algorithms, such as neural networks, to simulate learning and improve model assessment capabilities. However, because the accuracy of complex network algorithms depends on sample size, current deep learning algorithm development is limited to common complex diseases such as cancer, for which large-scale datasets are readily available, and cannot achieve risk assessment for complex diseases with scarce samples. Besides using complex algorithms to optimize models, researchers have attempted to stratify populations based on disease characteristics or add environmental factors for correction. While this has improved the accuracy of risk assessment to some extent, it has significantly reduced the model's generalization ability and foresight.
[0005] Therefore, current research indicates that using statistical algorithms based on genetic background to assess and predict the risk of complex diseases is feasible and has promising applications. However, it requires a large sample size and has weak generalization ability, lacking a universally applicable risk assessment strategy for complex diseases. Summary of the Invention
[0006] Therefore, the technical problem to be solved by the present invention is to provide a method for constructing a genetic risk assessment model for complex diseases, the model itself, and its application, thereby solving the technical problem of the lack of universally applicable risk assessment strategies for complex diseases in the prior art.
[0007] One technical solution provided by this invention is a method for constructing a genetic risk assessment model for complex diseases, comprising the following steps:
[0008] S1. Collect research samples, including samples from patients and samples from healthy individuals;
[0009] S2, genome sequencing and data processing, including whole genome sequencing and mutation site analysis, annotation of mutation sites, and identification of rare genetic variations that have a destructive impact on protein function;
[0010] S3. Statistical analysis to obtain characteristic genes that are significantly associated with whether or not a complex disease is present;
[0011] S4. Construct genetic risk assessment models for complex diseases;
[0012] The statistical analysis includes:
[0013] (1) Based on whole-genome sequencing data, Fisher's exact one-sided test algorithm was used to identify genes with high mutation frequency in diseased members at 95% confidence level;
[0014] (2) Based on whole-genome sequencing data, the Logistic test algorithm was used to identify risk genes that showed significant mutational differences between diseased members and healthy controls at a 95% confidence level;
[0015] (3) Based on the risk genes obtained in steps (1) and (2), an integrated analysis method is used to mine risk genes that are significantly associated with complex diseases.
[0016] Preferably, the statistical analysis also includes:
[0017] (4) The AUC value of true positives is greater than that of false positives in single gene discrimination, which is greater than 0.5;
[0018] (5) The relative risk of disease for a single gene is greater than 1 (OR).
[0019] Preferably, in step S1, the research samples include a family study cohort and a population validation cohort; the family study cohort is selected from family samples with a family genetic history that conforms to Mendelian inheritance laws, including diseased members and healthy members; the population validation cohort is selected from independent samples without genetic background, including diseased members without any cause and healthy control groups.
[0020] Preferably, the identification of rare genetic variations that have a destructive effect on protein function includes:
[0021] (1) In all the samples studied, the frequency of minor alleles was less than 5%;
[0022] (2) Disrupt the variants of the protein-coding sequence, namely stop gain, initiation loss, frameshift, or canonical splicing site alteration;
[0023] (3) Destructive missense variants are predicted by a polymorphic phenotype computer prediction algorithm.
[0024] Another technical solution provided by the present invention is a genetic risk assessment model for complex diseases. The model is based on all characteristic genes that are significantly related to whether or not a complex disease is contracted, obtained by any of the above methods. The characteristic genes are subjected to feature screening and dimensionality reduction, and one gene or a combination of genes is selected according to the characteristic gene sorting.
[0025] Preferably, a principal component regression algorithm is used to build a comprehensive evaluation index to achieve dimensionality reduction in the screening of feature genes.
[0026] Preferably, the LASSO regression algorithm is used to construct a nonlinear risk assessment model.
[0027] Preferably, the random forest algorithm is used to construct the integrated risk assessment model.
[0028] Preferably, environmental influencing factors are superimposed in the model.
[0029] The present invention also provides a technical solution for the application of a complex disease genetic risk assessment model in complex disease genetic risk assessment products or in pathogenesis research.
[0030] Beneficial effects:
[0031] This invention provides a method for constructing a genetic risk assessment model for complex diseases. Addressing the clinical diagnostic challenges of complex diseases, this method proposes a risk assessment strategy that combines clinical experience and statistical inference for complex diseases with scarce data. This invention utilizes completely independent sporadic population data to test the constructed optimal risk assessment model, verifying the generalizability and application prospects of the complex disease risk assessment strategy.
[0032] This invention presents a novel risk assessment system for constructing a genetic risk assessment model for complex diseases. Due to the high heterogeneity of risk loci among different patients in rare and complex diseases caused by genetic factors, existing technologies using SNPs as the genetic risk assessment unit struggle to uncover statistically significant general patterns. Therefore, this invention accumulates rare mutation sites at the gene level to identify risk factors with significant mutational differences between affected individuals and healthy controls, constructing a universally applicable risk assessment model and demonstrating good predictive accuracy in independent samples.
[0033] Since the OR value only represents the absolute risk of a single genetic factor for a disease, using the OR value to rank genetic factors for iterative modeling contradicts the complex molecular mechanisms of multi-gene interactions in diseases. Therefore, this invention uses phenotypic labeled family samples and employs multi-gene factor analysis, regression analysis, and ensemble learning algorithms for statistical modeling to identify characteristic genes significantly associated with the disease. The modeling is then iteratively performed after ranking the genes by their relative risk levels using corresponding statistical indicators.
[0034] Because the assessment of a patient's family history of disease in clinical diagnosis is highly subjective, independent risk assessment of patients with a family history of disease and sporadic patients without such a history is prone to systematic errors in clinical application and is not conducive to standardizing clinical diagnostic criteria. Therefore, this invention trains a model using families with a family history of disease and tests it in an independent sporadic population. Considering that a family history of disease is an important genetic factor in complex diseases, the algorithm that can fit the genetic family sample well and has a certain predictive ability in the independent sporadic sample is finally selected to construct a universally applicable risk assessment model, thereby enhancing the application value of the risk assessment model in clinical diagnosis.
[0035] This invention presents a genetic risk assessment model for complex diseases, addressing the technical challenge of accurately assessing risks with small sample sizes in existing technologies. Current risk assessment models for complex polygenic diseases primarily focus on cancer, employing machine learning to construct complex interaction networks. These models require large-scale datasets for training, resulting in poor training accuracy for risk assessment models of complex diseases with scarce data. Therefore, this invention fully utilizes data resources, employing regression analysis and ensemble learning algorithms suitable for small sample learning to construct a risk assessment model, demonstrating its good stability and generalization ability.
[0036] This invention constructs a knowledge-driven statistical model. Two feature selection methods are employed: directly selecting experimentally validated risk loci and using association analysis to uncover risk loci significantly associated with phenotypes. While directly selecting experimentally validated risk loci ensures the reliability of the risk model, existing techniques suffer from poor fitting results due to experimental constraints and are not conducive to exploring new and complex disease risk mechanisms. Furthermore, existing techniques for uncovering disease-related risk genes sometimes result in known associated genes being replaced by other statistically significant genes, leading to poor model interpretability. Therefore, this invention constructs a novel knowledge-driven statistical model that, based on known genes with genetic risk, gradually incorporates unverified, significantly differentially expressed genes. This ensures model interpretability while adding highly reliable potential risk genes, thereby improving the model's fitting and predictive performance. Attached Figure Description
[0037] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0038] Figure 1 This is a schematic diagram of the risk assessment model construction method in Embodiment 2 of the present invention;
[0039] Figure 2 This is a graph showing the test results of the risk assessment model in Embodiment 3 of the present invention;
[0040] Figure 3 This is a graph showing the test results of the risk assessment model in Embodiment 4 of the present invention;
[0041] Figure 4 This is a graph showing the test results of the risk assessment model in Embodiment 5 of the present invention;
[0042] Figure 5 This is a graph showing the test results of the risk assessment model in Embodiment 6 of the present invention;
[0043] Figure 6 This is a graph showing the verification results of the proportional risk assessment model of this invention. Detailed Implementation
[0044] To explain in detail the technical content, objectives, and effects of the present invention, the following description is provided in conjunction with the embodiments.
[0045] The meanings of English abbreviations and explanations of English terms in this article:
[0046] OR: odds ratio. Risk level value;
[0047] AUC: area under curve.
[0048] p-value: A parameter used to determine the result of hypothesis testing. In this invention, Fisher's test is used to calculate the p-value of genes with a high mutation frequency in diseased members.
[0049] Example 1
[0050] This embodiment describes a method for constructing a genetic risk assessment model for complex diseases, including the following steps:
[0051] S1. Collect research samples, including samples from patients and samples from healthy individuals;
[0052] The study samples include a family study cohort and a population validation cohort; the family study cohort is selected from family samples with a family history of inheritance in accordance with Mendelian laws of inheritance, including diseased members and healthy members; the population validation cohort is selected from independent samples without genetic background, including diseased members without any cause and healthy controls.
[0053] S2, genome sequencing and data processing, including whole genome sequencing and mutation site analysis, annotation of mutation sites, and identification of rare genetic variations that have a destructive impact on protein function;
[0054] Bioinformatics procedures were applied uniformly across all samples, with FastQC (v0.11.4) (BaseSpaceLabs, Illumina, San Diego, CA, USA) and Trimmomatic (v0.35) used for quality control. DNA sequences were aligned to the UCSC human reference genome (hg19, GRCh37) using Burrows-Wheeler alignment software (v0.7.17); different calls were made using the same bioinformatics pipeline according to GATK best practices (v3.5); and variant annotation was performed using ANNOVAR (v2018Apr16).
[0055] The following criteria were used to identify rare genetic variants predicted to have a disruptive effect on protein function:
[0056] (1) In all the populations studied, the frequency of minor alleles was less than 5%;
[0057] (2) Variations of protein-coding sequences that disrupt the sequence (i.e., stop gain, initiation loss, frameshift, or changes in the canonical splicing site);
[0058] (3) Destructive missense variants are predicted by the polymorphic phenotype (polyphen2) computer prediction algorithm.
[0059] The filter for small allele frequencies less than 5% was applied because genetic variants with small allele frequencies greater than 5% were well represented in the previously used GWAS genotyping arrays. All bioinformatics filters were pre-specified and uniformly applied in the discovery and validation cohorts using PLINK (version 1.9).
[0060] S3. Statistical analysis to obtain characteristic genes that are significantly associated with whether or not a complex disease is present;
[0061] An ensemble model was constructed using Fisher's exact one-sided test and logistic regression test to identify characteristic genes significantly associated with the occurrence of complex diseases. These include:
[0062] (1) Based on whole-genome sequencing data, Fisher's exact one-sided test algorithm was used to identify genes with high mutation frequency in diseased members at 95% confidence level;
[0063] (2) Based on whole-genome sequencing data, the Logistic test algorithm was used to identify risk genes that showed significant mutational differences between diseased members and healthy controls at a 95% confidence level;
[0064] (3) Based on the risk genes obtained in steps (1) and (2), an integrated analysis method is used to mine risk genes that are significantly associated with complex diseases;
[0065] In the statistical analysis process, a preferred technical solution may also include:
[0066] (4) The AUC value of true positives is greater than that of false positives in single gene discrimination, which is greater than 0.5;
[0067] (5) The relative risk of disease for a single gene is greater than 1 (OR).
[0068] Using the screened genes related to complex diseases and the collected samples as units, rare mutation sites in the exon regions and alternative splicing regions of the samples were statistically analyzed according to mutation type. The number of mutation sites of each gene in each sample for each mutation type was obtained. The whole gene was sorted according to OR value, AUC, and p-value in that order. Then, it was filtered according to the criteria of OR>1, p-value<0.05, and AUC>0.5 to obtain characteristic genes that are significantly associated with whether or not a complex disease is present.
[0069] The method of this invention differs from the traditional GWAS analysis process. The multi-algorithm evaluation model based on hypothesis testing avoids random errors caused by algorithmic bias to a certain extent and reduces the impact of sample size on model accuracy. The multi-index standard evaluation can obtain genes with large differences and strictness.
[0070] S4. Construct genetic risk assessment models for complex diseases;
[0071] Based on all the characteristic genes that are significantly associated with the occurrence of complex diseases, a genetic risk assessment model for complex diseases is constructed. The characteristic genes are then subjected to feature screening and dimensionality reduction, and one or more gene combinations are selected by sorting the characteristic genes.
[0072] Unlike traditional GWAS analysis processes, multi-algorithm evaluation models based on hypothesis testing avoid random errors caused by algorithmic bias to a certain extent and reduce the impact of sample size on model accuracy. Multi-index standard evaluation can obtain genes with large differences and strictness.
[0073] Example 2
[0074] This embodiment uses venous thromboembolism (VTE) as an example. See [link to example]. Figure 1 This paper demonstrates a sequencing and analysis method for identifying genes associated with complex diseases by utilizing significant differences in gene mutations between diseased and healthy samples. The method for constructing a VTE genetic risk assessment model includes the following steps:
[0075] S1. Collect research samples, including samples from patients and samples from healthy individuals;
[0076] The study sample included a family cohort and a population validation cohort. The family cohort collected 105 samples from 35 families with Mendelian inheritance histories, of whom 76 were affected and 29 were asymptomatic. The population validation cohort consisted of independent samples without a genetic background, including 99 patients and 90 asymptomatic individuals. All participants were from the Chinese Pulmonary Thromboembolism Registry (CURES), encompassing patients diagnosed with acute pulmonary embolism at over 100 hospitals in China since 2011. Specimen collection was approved with research consent and by the Institutional Accreditation Committee of the China-Japan Friendship Hospital (No. 2016-SSW-7). All information obtained was protected and de-identified.
[0077] S2, genome sequencing and data processing;
[0078] This embodiment primarily employs the BWA+GATK workflow for whole-genome sequencing and preliminary mutation analysis. First, the raw FASTQ data files are preprocessed and aligned to the human reference genome GRCh37 / hg19. Mutation detection is performed on individual samples to generate gVCF files, followed by joint mutation detection of 105 samples to generate the original VCF files. After VQSR correction, a reliable VCF result file is generated. Based on the ANNOVAR annotation tool, mutation sites are annotated, including gene-based annotation, region-based annotation, and population frequency-based annotation (including 1000 Genomes, ExAC, ESP6500, CG46, and gnomADgenome).
[0079] The following criteria were used to identify rare genetic variants predicted to have a disruptive effect on protein function:
[0080] (1) In all the populations studied, the frequency of minor alleles was less than 5%;
[0081] (2) Variations of protein-coding sequences that disrupt the sequence (i.e., stop gain, initiation loss, frameshift, or changes in the canonical splicing site);
[0082] (3) Destructive missense variants are predicted by the polymorphic phenotype (polyphen2) computer prediction algorithm.
[0083] The filter for small allele frequencies less than 5% was applied because genetic variants with small allele frequencies greater than 5% were well represented in the previously used GWAS genotyping arrays. All bioinformatics filters were pre-specified and uniformly applied in the discovery and validation cohorts using PLINK (version 1.9).
[0084] S3. Statistical analysis to obtain characteristic genes that are significantly associated with whether or not a complex disease is present;
[0085] An ensemble model was constructed using Fisher's exact one-sided test and logistic regression test to identify characteristic genes significantly associated with the occurrence of complex diseases. These include:
[0086] (1) Based on whole-genome sequencing data, Fisher's exact one-sided test algorithm was used to identify genes with high mutation frequency in diseased members at 95% confidence level;
[0087] (2) Based on whole-genome sequencing data, the Logistic test algorithm was used to identify risk genes that showed significant mutational differences between diseased members and healthy controls at a 95% confidence level;
[0088] (3) Based on the risk genes obtained in steps (1) and (2), an integrated analysis method is used to mine risk genes that are significantly associated with complex diseases;
[0089] (4) The AUC value of true positives is greater than that of false positives in single gene discrimination, which is greater than 0.5;
[0090] (5) The relative risk of disease for a single gene is greater than 1 (OR).
[0091] Using data from a Han Chinese population with a family history of VTE, a complex disease risk gene mining model was employed. From 14,829 mutated genes obtained through WGS sequencing, Fisher's exact one-sided test algorithm was used to identify 67 genes with high mutation frequencies in VTE patients. The Logistic regression algorithm was used to identify 38 genes with significant mutation differences between affected members and the control group. The potential risk genes obtained from different algorithms were integrated to identify 35 statistically significant VTE risk genes: SERPINC1, F5, AQP9, and L... RRC6, KIF6, DHX34, REXO1, TPBGL, MFAP1, RTKN2, ASPN, MKI67, KIAA1522, FGFR4, PGLYRP4, CCBL2, ATP5S, ENDOG, DMBT1, DBR1, PNPLA7, TBC1D24, CSF1R, GSTA5, ERICH6, C7orf25, GALK1, TMEM143, OXR1, SETD5, AQP7, ERMAP, ZNF831, ENOSF1, GPR142. SERPINC1 and F5 have been experimentally validated. Furthermore, of the 39 experimentally validated VTE pathogenic genes, only 17 are known to pose a pathogenic risk (OR) in Chinese families. knwon >1).
[0092] Example 3
[0093] This embodiment is a VTE genetic risk assessment model, based on the VTE genetic risk assessment model of 35 characteristic genes obtained in Example 2.
[0094] A total of 51,827 rare (population frequency less than 0.05) mutation sites were found in the whole exon region plus alternative splicing region in 105 samples from 35 families, involving 14,829 genes. The number of mutation sites for each gene in each mutation type was counted in each sample, generating a mutation count matrix with gene samples as rows and mutation types as columns. The mutation count matrix has 1,557,045 rows (14,829 genes multiplied by 105 samples) and nine columns, including gene name, sample number, and the number of alternative splicing mutations, frameshift insertion / deletion mutations, nonsense mutations, missense mutations, non-frameshift insertion / deletion mutations, synonymous mutations, and unknown mutations. Based on the known gene mutation screening results and gene load test results, synonymous mutations and unknown mutations were disregarded, and only five mutation types—alternative splicing, frameshift insertion / deletion, nonsense mutations, missense mutations, and non-frameshift insertion / deletion—were retained.
[0095] Gene features were ranked according to OR, p-value, and AUC. SERPINC1, F5, AQP9, LRRC6, and KIF6 were used. Logistic regression modeling was performed on the top 5 feature genes, with an AUC of 0.872, precision of 1, recall of 0.579, and an overall F1 score of 0.733.
[0096] As shown in Table 1, 5-fold cross-validation was performed on the family samples. Both the training and validation sets were sampled from families, and the model precision, recall, and F1 score were calculated for each fold.
[0097] Table 1. Results of 5-fold cross-validation of the characteristic gene model.
[0098]
[0099] Example 4
[0100] This embodiment is a VTE genetic risk assessment model, based on the VTE genetic risk assessment model of 35 characteristic genes obtained in Example 2.
[0101] Risk assessment was performed stepwise on the 35 characteristic genes in descending order of OR value, and the AUC value was calculated to test the assessment effect. When the AUC value no longer increased after the seventh step of adding ASPN, the inclusion of new characteristic genes was stopped. Finally, SERPINC1, LRRC6, KIF6, TPBGL, MFAP1, RTKN2, and ASPN were selected as characteristic risk genes for risk assessment (AUC = 0.83). See [link to relevant documentation]. Figure 2 As shown in Table 2, the VTE genetic risk assessment model in this embodiment outperforms the risk assessment of known genes using the same algorithm. This demonstrates that the risk features mined using statistical algorithms in this invention have stronger risk perception capabilities.
[0102] Example 5
[0103] This embodiment is a VTE genetic risk assessment model, based on the VTE genetic risk assessment model of 35 characteristic genes obtained in Example 2.
[0104] This embodiment uses principal component analysis (PCA) to reduce the dimensionality of 35 feature genes. A scree plot reveals that 12 principal components can effectively explain the original data. Risk assessment is performed stepwise according to the principal components, and AUC values are calculated to verify the assessment effectiveness. After PC1, the AUC value no longer increases, and new components are no longer included. Analysis of the PC1 components uses seven orthogonally related genes—SERPINC1, CCBL2, TMEM143, GALK1, FGFR4, C7orf25, and ERMAP—that play a major assessment role as feature risk models. See [link to relevant documentation]. Figure 3 As shown in Table 2, the detection AUC in sporadic populations can reach 0.89, which indicates that the model in this embodiment achieves good predictive performance in sporadic populations.
[0105] Example 6
[0106] This embodiment is a VTE genetic risk assessment model, based on the VTE genetic risk assessment model of 35 characteristic genes obtained in Example 2.
[0107] This embodiment employs a non-linear penalty term-addition algorithm to screen 35 feature genes. Feature ranking is used to progressively assess risk, and the AUC value is calculated to verify the assessment effect. When the AUC value stops increasing after step 7 (up to MFAP1), new feature genes are no longer included. Finally, AQP9, SERPINC1, ASPN, LRRC6, DHX34, RTKN2, and MFAP1 are selected as the feature risk genes for risk assessment, with an AUC of 0.88. (See...) Figure 4 As shown in Table 2, the evaluation model in this embodiment achieves dimensionality reduction while outperforming the evaluation using OR values and eigenfactors.
[0108] Example 7
[0109] This embodiment is a VTE genetic risk assessment model, based on the VTE genetic risk assessment model of 35 characteristic genes obtained in Example 2.
[0110] This embodiment employs a random forest algorithm based on ensemble learning, ranking gene risk levels based on support. After iteration 7 to REXO1, the AUC value stops increasing, and new feature genes are no longer included. Ultimately, AQP9, SERPINC1, ERICH6, ENDOG, ASPN, TPBGL, and REXO1 are selected as feature risk genes for risk assessment. In the fitting effect on pedigree training data, the AUC is 0.83, weaker than the penalized term algorithm, but its prediction effect on independent sporadic populations is significantly better than the penalized term algorithm, greatly improving the robustness of the risk prediction model. See [link to relevant documentation]. Figure 5 And Table 2.
[0111] Comparative Example
[0112] This comparative study constructed a VTE risk assessment model based on known VTE genes with potential risk. Known VTE genes with potential risk were sorted in descending order of OR value. The number of mutations was used as a risk index, and the results were cumulatively summed for 35 families. The AUC value was calculated to test the assessment effectiveness. When the AUC value no longer increased after the addition of SERPINC1 in step seven, the inclusion of new known genes was stopped. Finally, F5, SCARA5, TSPAN15, MPHOSPH9, F2, GRK5, and SERPINC1 were selected as characteristic risk genes for risk assessment. The AUC was 0.75 in families and 0.56 in sporadic populations. (See [link to relevant documentation]). Figure 6 And Table 2.
[0113] Table 2. Validation results of the VTE risk assessment models constructed in Examples 4-7 and comparative examples of the present invention.
[0114]
[0115] Example 8
[0116] This embodiment presents a VTE genetic risk assessment model based on the 35 characteristic genes obtained in Example 2. Furthermore, considering that the combined genetic and environmental factors may increase the risk of VTE, a hybrid VTE risk assessment model based on both characteristic genes and environmental factors was constructed.
[0117] This embodiment, in addition to genetic factors, incorporates the influence of sex and age on VTE incidence. When sex is added, the characteristic gene model AUC increases from 0.872 to 0.887, but the improvement is not significant. When age is added, the characteristic gene model AUC increases from 0.872 to 0.952, showing a significant improvement. If both sex and age are considered, the characteristic gene model AUC increases from 0.872 to 0.956, a slight improvement compared to the model with only age added.
[0118] In summary, the genetic risk assessment model for complex diseases developed using the method of this invention, based on genes as units, is applicable to a variety of complex diseases, such as cardiovascular diseases including heart failure, hypertension, and cardiomyopathy, and metabolic diseases including type 2 diabetes, hypercholesterolemia, and thyroid dysfunction. It is applicable to a wider range of populations in different regions around the world and can more robustly and accurately predict the genetic risk of diseases. It has significant value for the prediction of risk and the study of disease mechanisms in complex diseases.
[0119] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for constructing a genetic risk assessment model for complex diseases, characterized in that, Includes the following steps: S1. Collect research samples, which include a family study cohort and a population validation cohort; the family study cohort is selected from family samples with a family history that conforms to Mendelian inheritance, including diseased members and healthy members; the population validation cohort is selected from independent samples without a genetic background, including diseased members without any cause and healthy control groups. S2, genome sequencing and data processing, including whole genome sequencing and mutation site analysis, annotation of mutation sites, and identification of rare genetic variations that have a destructive impact on protein function; S3. Statistical analysis to obtain characteristic genes that are significantly associated with whether or not a complex disease is present; The statistical analysis includes: (1) Based on whole genome sequencing data, Fisher's exact one-sided test algorithm was used to mine genes with high mutation frequency in diseased members at 95% confidence level; (2) Based on whole-genome sequencing data, the Logistic test algorithm was used to identify risk genes that showed significant mutation differences between diseased members and healthy controls at a 95% confidence level; (3) Based on the risk genes obtained in steps (1) and (2), ensemble analysis methods are used to mine risk genes that are significantly associated with complex diseases; S4. Construct genetic risk assessment models for complex diseases; The feature genes are subjected to feature screening and dimensionality reduction, and one gene or a combination of multiple genes is selected by sorting the feature genes. Alternatively, a comprehensive evaluation index can be constructed using principal component regression algorithm to achieve dimensionality reduction in the screening of feature genes; or, an integrated risk assessment model can be constructed using random forest algorithm.
2. The method for constructing a genetic risk assessment model for complex diseases according to claim 1, characterized in that, The statistical analysis also includes: (4) The AUC value of true positives is greater than that of false positives in single gene discrimination, which is greater than 0.5; (5) The relative risk of disease for a single gene is greater than 1 (OR).
3. The method for constructing a genetic risk assessment model for complex diseases according to claim 2, characterized in that, The identification of rare genetic variations that have a disruptive effect on protein function includes: (1) In all the samples studied, the frequency of minor alleles was less than 5%; (2) Disrupting variants of the protein-coding sequence, i.e., stop gain, initiation loss, frameshift, or canonical splicing site alteration; (3) Destructive missense variants are predicted by a polymorphic phenotype computer prediction algorithm.
4. A genetic risk assessment model for complex diseases, characterized in that, Obtained by the method according to any one of claims 1 to 3.
5. The genetic risk assessment model for complex diseases according to claim 4, characterized in that, The model incorporates environmental influencing factors.
6. The genetic risk assessment model for complex diseases according to claim 4 or 5 in the context of complex diseases Applications in disease genetic risk assessment products or in pathogenesis research.
Citation Information
Patent Citations
Next generation sequencing based coronary heart disease genetic risk evaluation method
CN105279369A
Analyzing method for inheritance mutation mode of trio family on basis of second-generation high throughput sequencing
CN109545281A