Screening methods and systems for risk genes and risk mutations in polygenic inherited diseases
By designing three screening strategies and analyzing significant mutation sites, the problem of incorrect or missed screening in the screening of polygenic genetic diseases was solved, and the screening of risk genes and mutations with higher accuracy was achieved. It is applicable to the screening of polygenic genetic diseases such as cardiovascular diseases and metabolic diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-20
- Publication Date
- 2026-04-03
AI Technical Summary
Current technologies for screening risk genes and risk mutations for polygenic genetic diseases suffer from misscreening and omissions, and their accuracy needs to be further improved. In particular, it is difficult to detect rare mutations and comprehensively analyze risk mutations from multiple families in complex polygenic genetic diseases.
Three different screening strategies were employed: Filtering strategy 1 targeted genes associated with confirmed polygenic inherited diseases, Filtering strategy 2 targeted candidate genes that may be related, and Filtering strategy 3 targeted other genes. Cosegregating variants were screened by combining GWAS studies and functional experiments. Sanger validation was used to analyze significant mutation sites and test gene load. Data from multiple families were used for screening.
It improves the accuracy of screening risk genes and risk mutations for polygenic inherited diseases, avoids missing rare mutations, can analyze non-core families, discover potential risk genes, and provides a more comprehensive basis for genetic risk assessment.
Smart Images

Figure CN115171784B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of gene detection, specifically relating to a method and system for screening risk genes and risk mutations of polygenic inherited diseases. Background Technology
[0002] Polygenic diseases are hereditary diseases caused by the cumulative effect of genetic information through two or more pairs of disease-causing genes, and their genetic effects are largely influenced by environmental factors. Common polygenic diseases include congenital heart disease, childhood schizophrenia, familial intellectual disability, spina bifida, anencephaly, juvenile diabetes mellitus, congenital hypertrophic pyloric stenosis, severe myasthenia gravis, congenital megacolon, tracheoesophageal fistula, congenital cleft palate, congenital hip dislocation, congenital esophageal atresia, clubfoot, primary epilepsy, bipolar disorder, hypospadias, congenital asthma, incomplete testicular descent, and hydrocephalus.
[0003] Unlike single-gene inherited diseases, polygenic inherited diseases are not determined solely by genetic factors, but rather by the combined effects of genetic and environmental factors. The degree to which genetic factors play a role compared to environmental factors is called heritability, expressed as a percentage. For example, schizophrenia, the most common and most detrimental mental illness, is a polygenic inherited disease with a heritability of 80%. This means that genetic factors play a significant role in the development of schizophrenia, while environmental factors have a relatively smaller role. Another example is venous thromboembolism (VTE), which has been proven to be a polygenic inherited disease primarily driven by genetic risk and exhibiting racial specificity. Polygenic inherited diseases generally have a familial tendency; for instance, the incidence rate among close relatives of schizophrenia patients is several times higher than in the general population, and the closer the blood relationship to the patient, the higher the incidence rate.
[0004] Identifying and confirming the pathogenic genes and mutation sites that cause disease is crucial for the prevention and improvement of polygenic inherited diseases. Existing technologies include many gene mutation pathogenicity prediction software programs for preliminary assessment of gene mutation pathogenicity, such as the early SIFT, PolyPhen, and MutationTaster, and more recently developed programs like LRT, FATHMM, PROVEAN, VEST3, FATHMM-MKL, MetaSVM, and MetaLR. However, these programs may contain mutations that do not conform to familial inheritance patterns, and some rare pathogenic mutations may remain undetected.
[0005] A Chinese patent application with publication number CN11091867A discloses a "Method and System for Screening Gene Variant Sites." This method includes: obtaining a first dataset containing gene variant sites from a specified population; clustering the gene variant sites in the first dataset to obtain multiple clusters of gene variant sites; scoring the gene variant sites in each cluster; and screening out gene variant sites with scores greater than a preset threshold. This method uses whole-genome sequencing data at 30x sequencing depth from a certain number of Chinese individuals as the basic dataset. It uses the GATK tool to extract gene variant sites from the basic dataset to obtain the original dataset, and then uses Affymetrix software to screen out the first gene variant sites from the original dataset, followed by clustering and subsequent scoring. However, this method lacks targeted samples and still misses rare mutations in complex polygenic genetic diseases.
[0006] Another Chinese patent application, CN112375815A, discloses a "method for screening pathogenic mutations in genetic diseases using high-throughput sequencing based on core pedigrees." This method includes: obtaining a genome sample to be tested, the genome sample originating from offspring and their parents; determining the DNA sequence of the genome sample using high-throughput methods; comparing and annotating the sequencing results with a human reference genome to form an annotation file; analyzing and screening the annotation file using a binary classification method to remove high-frequency mutations and classify the detected gene mutations; and removing non-pathogenic mutations based on the classification and annotation of the detected gene mutations. This method has several drawbacks: it only supports core pedigrees and cannot calculate mutations in non-core pedigrees; it only calculates mutations for a single pedigree and cannot analyze multiple pedigrees; and it only screens for pathogenic mutation sites, not risk genes. This deficiency exists because some rare pathogenic mutations are often difficult to replicate in other pedigrees due to their low population frequency, but risk mutation genes are consistent across different pedigrees. Summary of the Invention
[0007] Therefore, the technical problem to be solved by this invention is to provide a method and system for screening risk genes and risk mutations for polygenic inherited diseases. This addresses the technical problems in existing technologies where the screening of risk genes and risk mutations for polygenic inherited diseases suffers from incorrect or missed screening, and the accuracy needs further improvement.
[0008] The present invention provides a technical solution for screening risk genes and risk mutations of polygenic inherited diseases, including screening known genetic risk factors and screening risk variants in research samples; the risk variants are formulated with three different filtering strategies based on different gene sets, namely filtering strategy 1, filtering strategy 2 and filtering strategy 3; filtering strategy 1 is a mutation screening filtering strategy for genes confirmed to be related to the polygenic inherited disease to be tested; filtering strategy 2 is a mutation screening filtering strategy for genes that may be related to the polygenic inherited disease to be tested; filtering strategy 3 is a mutation screening filtering strategy for other genes.
[0009] Preferably, the research sample includes a family study cohort and a population validation cohort; the family study cohort is selected from families of a specific region and ethnicity, including diseased members and non-diseased members; the population validation cohort is selected from the general population of the same region and ethnicity as the family study cohort, including sporadic cases without cause and healthy controls.
[0010] Preferably, the risk variants are all co-segregating variants in the screening family study cohort.
[0011] Preferably, the filtering strategy 1 obtains all identified genetic risk factors (i.e., confirmatory genes) for the polygenic hereditary diseases to be tested through GWAS studies or functional experiments, and the filtering rules are as follows:
[0012] a) Screen for variations in all identified gene regions.
[0013] b) Filter out non-bicelestem mutations or synonymous mutation sites.
[0014] c) The mutation is segregated in at least one family lineage, and there is no more than one mutation carrier among healthy members in other families;
[0015] This yields a list of candidate sites for filtering strategy 1, which is then used to screen for loss-of-function mutations or rare non-synonymous mutations for sequencing verification.
[0016] Preferably, the filtering strategy 2 collects other known genes or genes that may be related to the polygenic hereditary disease to be tested, i.e., candidate genes, based on existing literature, and the filtering rules are as follows:
[0017] a) Variations in the candidate genes obtained through screening.
[0018] b) Filter out non-bicelestem mutations or synonymous mutation sites.
[0019] c) The mutation is segregated in at least one family lineage, and there is no more than one mutation carrier in healthy members of other families;
[0020] This yields a candidate site list for filtering strategy 2, which is then used to screen for loss-of-function mutations, rare pathogenic nonsynonymous mutations shared by at least two families, or rare pathogenic gene nonsynonymous mutations shared by at least three families, and then validated by sequencing.
[0021] Preferably, the filtering rules of filtering strategy 3 are as follows:
[0022] a) Screen for rare variants located in exon regions and alternative splicing regions.
[0023] b) Filter out non-bicelestem mutations or synonymous mutation sites.
[0024] c) The mutation is segregated in at least one family lineage, and there is no more than one mutation carrier in healthy members of other families;
[0025] This yields a candidate site list for filtering strategy 3. Based on the pathogenicity prediction of the site and the results of screening in the population cohort, significant non-homozygous sites in the population are selected to obtain new gene sites that may be related to the polygenic genetic diseases to be tested.
[0026] Preferably, after obtaining new gene loci that are potentially associated with the polygenic hereditary disease to be tested, significant mutation site analysis and gene load testing are performed, including association analysis between the variant and the polygenic hereditary disease to be tested and load analysis of the association between the gene and the polygenic hereditary disease to be tested.
[0027] Preferably, the association analysis between the variant and the polygenic genetic disease to be tested is as follows: for each variant obtained by screening with three different filtering strategies, the allele frequencies of cases and control groups in the family study mutation dataset and the population validation mutation dataset are analyzed, and the association analysis between each variant and whether or not the disease is present is performed based on logistic regression model and Fisher's exact test, respectively. An association with p<0.05 is defined as a significant mutation.
[0028] The association load analysis between the genes and the polygenic genetic diseases to be tested was conducted as follows: mutations were accumulated on a gene-by-gene basis and statistically analyzed to determine the difference in the number of gene mutations between patients and healthy controls; mutation types included alternative splicing, nonsense mutations, frameshift insertions / deletions, missense mutations, non-frameshift insertions / deletions, synonymous mutations, and unknown mutations; gene load analysis was performed on each combination of mutation types based on logistic regression models and Fisher's exact test to obtain the risk gene ranking results under different combinations.
[0029] Preferably, the study also includes whole-genome sequencing and whole-exome sequencing of the research samples, followed by mutation site analysis, annotation and filtering to select mutation sites in exon regions and alternative splicing regions, as well as rare mutation sites with a population frequency of less than 0.05.
[0030] The present invention also provides a technical solution for a screening system for risk genes and risk mutations of polygenic inherited diseases, including a computer system, which is programmed to perform the steps of the above-described screening method for risk genes and risk mutations of polygenic inherited diseases.
[0031] Beneficial effects:
[0032] The present invention provides a method for screening risk genes and risk mutations for polygenic inherited diseases. By designing and applying three different screening strategies, all cosegregating variants are screened from confirmatory genes, candidate genes and other genes. Then, the mutation sites are verified by Sanger, thereby obtaining the risk genes and risk mutations of specific polygenic inherited diseases, laying the foundation for subsequent assessment of the genetic risk of the disease and development of related products.
[0033] This invention proposes a screening strategy for risk mutations and risk genes in polygenic inherited diseases. This strategy can analyze not only core families but also non-core families. It employs different strategies at different levels (confirming genes, candidate genes, and other genes) to effectively screen for risk mutations. Furthermore, based on the final risk mutation site, potential risk genes are analyzed. Multiple families can be analyzed collectively; for example, if two or more families coexist the same risk mutation, or if a rare mutation is not found to be repeated in two or more families but the mutated gene is consistent across different families, potential risk genes can be identified, preventing missed detections. Attached Figure Description
[0034] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0035] Figure 1 This is a schematic diagram of the queue of research samples in an embodiment of the present invention.
[0036] Figure 2 This invention presents three different gene mutation filtering strategies for screening risk variants of venous thromboembolism in 35 families with VTE disease. Detailed Implementation
[0037] To explain in detail the technical content, objectives, and effects of the present invention, the following description is provided in conjunction with the embodiments.
[0038] The meanings of English abbreviations and explanations of English terms in this article:
[0039] MAF: Minor allele frequency in public database.
[0040] SIFT: Sorting Intolerant From Tolerant. SIFT is a software for predicting the harmfulness of nonsynonymous mutations; it's a prediction algorithm based on the conservatism of base substitution probabilities.
[0041] Polyphen: Polymorphism Phenotyping. A software for predicting the harmfulness of nonsynonymous mutations. Polyphen: Nucleotide polymorphism and phenotype.
[0042] ACMG: American College of Medical Genetics and Genomics.
[0043] ACMG (InterVar_automated): InterVar is a bioinformatics software tool that provides clinical interpretation of genetic variations according to ACMG guidelines.
[0044] Missense: a misspelling or misinterpretation;
[0045] Splicing: Variable shearing;
[0046] Nonframeshift insertion;
[0047] Nonframeshift deletion: Deletion without frameshifting;
[0048] Frameshift deletion: Frameshift deletion;
[0049] Frameshift insertion;
[0050] Stop gain: terminates mutation;
[0051] Uncertain significance: pathogenicity is uncertain;
[0052] Likely pathogenic: potentially pathogenic;
[0053] Pathogenic: causative.
[0054] The present invention provides a method for screening risk genes and risk mutations for polygenic inherited diseases, comprising the following steps:
[0055] S1. Collection of research samples and gene testing experiments
[0056] The study sample includes a family study cohort and a population validation cohort; the family study cohort is selected from families of a specific region and ethnicity, including both diseased and unaffected members; the population validation cohort is selected from the general population of the same region and ethnicity as the family study cohort, including sporadic cases without apparent cause and healthy controls.
[0057] Genetic testing experiments include whole-genome sequencing and whole-exome sequencing.
[0058] S2. Perform mutation site analysis, annotation, and filtering to select mutation sites in exon regions and alternative splicing regions, as well as rare mutation sites with a population frequency of less than 0.05.
[0059] S3. Screening known genetic risk factors and identifying risk variants in research samples.
[0060] Among them, screening for known genetic risk factors involves screening for genes related to known polygenic inherited diseases, including confirmed genes and candidate genes;
[0061] Risk variants are all cosegregating variants in the screening family study cohort. Three different filtering strategies are developed based on different gene sets: filtering strategy 1, filtering strategy 2, and filtering strategy 3. Filtering strategy 1 is a mutation screening filtering strategy for genes confirmed to be associated with the polygenic genetic disease to be tested; filtering strategy 2 is a mutation screening filtering strategy for genes that may be associated with the polygenic genetic disease to be tested; and filtering strategy 3 is a mutation screening filtering strategy for other genes.
[0062] Specifically, filtering strategy 1 involves obtaining all identified genetic risk factors (i.e., confirmatory genes) for the polygenic hereditary diseases to be tested through GWAS studies or functional experiments. The filtering rules are as follows:
[0063] a) Screen for variations in all identified gene regions.
[0064] b) Filter out non-bicelestem mutations or synonymous mutation sites.
[0065] c) The mutation is segregated in at least one family lineage, and there is no more than one mutation carrier in healthy members of other families;
[0066] This yields a list of candidate sites for filtering strategy 1, which is then used to screen for loss-of-function mutations or rare non-synonymous mutations for sequencing verification.
[0067] Filtering strategy 2 involves collecting other known genes or genes that may be associated with the polygenic hereditary disease being investigated, i.e., candidate genes, based on existing literature. The filtering rules are as follows:
[0068] a) Variations in the candidate genes obtained through screening.
[0069] b) Filter out non-bicelestem mutations or synonymous mutation sites.
[0070] c) The mutation is segregated in at least one family lineage, and there is no more than one mutation carrier among healthy members in other families;
[0071] This yields a candidate site list for filtering strategy 2, which is then used to screen for loss-of-function mutations, rare pathogenic nonsynonymous mutations shared by at least two families, or rare pathogenic gene nonsynonymous mutations shared by at least three families, and then validated by sequencing.
[0072] The filtering rules for filtering strategy 3 are as follows:
[0073] a) Screen for rare variants located in exon regions and alternative splicing regions.
[0074] b) Filter out non-bicelestem mutations or synonymous mutation sites.
[0075] c) The mutation is segregated in at least one family lineage, and there is no more than one mutation carrier among healthy members in other families;
[0076] This yields a candidate site list for filtering strategy 3. Based on the pathogenicity prediction of the site and the results of screening in the population cohort, significant non-homozygous sites in the population are selected to obtain new gene sites that may be related to the polygenic genetic diseases to be tested.
[0077] S4. Perform significant mutation site analysis and gene load testing.
[0078] This includes association analysis between variants and the polygenic genetic disease to be tested, and load analysis of the association between genes and the polygenic genetic disease to be tested.
[0079] The association analysis between variants and the polygenic genetic diseases to be tested was as follows: For each variant obtained by screening with three different filtering strategies, the allele frequencies of cases and control groups in the family study mutation dataset and the population validation mutation dataset were analyzed. The association between each variant and the disease status was analyzed based on logistic regression model and Fisher's exact test, respectively. An association with p < 0.05 was defined as a significant mutation.
[0080] The association load analysis between genes and the polygenic hereditary diseases under investigation was conducted as follows: mutations were accumulated on a gene-by-gene basis and statistically analyzed to determine the difference in the number of gene mutations between patients and healthy controls. Mutation types included alternative splicing, nonsense mutations, frameshift insertions / deletions, missense mutations, non-frameshift insertions / deletions, synonymous mutations, and unknown mutations. Gene load analysis was performed on each combination of mutation types based on logistic regression models and Fisher's exact test to obtain the risk gene ranking results under different combinations.
[0081] Polygenic inherited diseases are complex conditions caused by multiple genes and environmental factors. They have a high incidence rate in the population and exhibit genetic heterogeneity and phenotypic complexity. Clinical studies have shown that genetic factors play a crucial role in the development of complex diseases; therefore, elucidating the genotype-phenotype relationship of complex diseases is helpful in studying their pathogenesis. Due to the complexity of polygenic interactions, these diseases often lack a clear inheritance pattern, making it difficult to determine the genetic characteristics used for clinical diagnosis and treatment. Furthermore, the pathogenic SNPs in these diseases are mostly rare mutation sites, and even SNPs discovered experimentally cannot be reproduced.
[0082] This invention establishes targeted research samples for specific diseases, designs three different filtering strategies to screen for disease-related confirmatory genes, candidate genes, and other genes, and then uses significant mutation site analysis to finally identify gene loci significantly associated with the detected polygenic hereditary diseases. This method achieves higher accuracy in screening for gene mutations significantly associated with polygenic hereditary diseases, avoiding the missed detection of rare mutations.
[0083] The method of this invention is applicable to the screening of risk genes and risk mutations for polygenic inherited diseases, such as cardiovascular diseases, including heart failure, hypertension, cardiomyopathy, etc., and metabolic diseases, including type 2 diabetes, hypercholesterolemia, thyroid dysfunction, etc.
[0084] Example 1
[0085] This embodiment takes venous thromboembolism as an example to screen for its risk genes and risk mutations, including the following steps:
[0086] S1. Collection of research samples and gene testing experiments
[0087] The study sample included samples from VTE patients and healthy individuals. Both VTE patients and healthy individuals included members of a family lineage and members of other populations outside that family lineage.
[0088] like Figure 1 As shown, the research samples in this embodiment were selected from a family study cohort of 216 samples (including 87 patients with venous thromboembolism and 129 unaffected members) from 35 Han Chinese families with venous thromboembolism and a population validation cohort of 1,598 samples (including 99 sporadic cases of venous thromboembolism without cause and 1,499 healthy controls) from Han Chinese.
[0089] The inclusion criteria for the family study cohort were: confirmed venous thromboembolism according to clinical guidelines, with two or more individuals in a family lineage having the disease. This clinical study was approved by the Institutional Review Committee of China-Japan Friendship Hospital (No. 2016-SSW-7). Clinical information and peripheral blood samples (confirmed VTE cases and other healthy members) of all enrolled VTE family members were collected, and DNA was extracted and cryopreserved. All information obtained was protected for privacy.
[0090] Samples were selected from the pedigree study cohort for whole-genome sequencing (WGS, >40x). The selection criteria were: a) core family members (at least one child and one parent diagnosed with venous thromboembolism); or b) non-core family members (the proband and their siblings). After WGS data analysis and variant screening, all available samples from 216 individuals across 35 families were validated for variants using Sanger sequencing.
[0091] The population validation cohort was derived from other Chinese genomic research datasets, including 99 cases of non-occurring incidental venous thromboembolism from the Chinese Pulmonary Thromboembolism Registry and 1499 healthy controls from the Chinese Pulmonary Hypertension Study Cohort (CPH). Whole exome sequencing (WES, >100x) data from the 99 patients and WGS data (>10x) data from the 1499 controls were obtained.
[0092] S2, Mutation Site Analysis, Annotation, and Filtering
[0093] This embodiment primarily employs the BWA+GATK workflow for whole-genome sequencing and preliminary mutation analysis. First, the raw FASTQ data is preprocessed and aligned to the human reference genome GRCh37 / hg19. Mutation detection is performed on individual samples to generate gVCF files, followed by joint mutation detection of 105 samples to generate the original VCF files. After VQSR correction, a reliable VCF result file is generated. Annotations are performed on mutation sites using the ANNOVAR annotation tool, including gene-based annotation, region-based annotation, and population frequency-based annotation (including 1000 Genomes, ExAC, ESP6500, CG46, and gnomADgenome). The harmfulness of non-synonymous mutations is assessed using algorithms such as SIFT (sorts intolerant from tolerant) and Polyphen-2 (Polymorphism Phenotyping v2). Finally, the mutation sites were filtered. First, they were filtered by region, that is, mutation sites in exon regions and alternative splicing regions were selected for study. Then, they were filtered by population frequency, that is, mutation sites whose frequencies in the population database were no more than 0.05 were selected.
[0094] S3, VTE genetic risk genes and risk mutation screening
[0095] S31. Screen for known genetic risk factors;
[0096] We collected 210 known VTE-related genes, of which 55 were confirmed VTE genetic risk factors validated by GWAS studies or functional experiments. These genes were referred to as confirmed genes, and the six associated biological pathways included: coagulation, anticoagulation, platelets, inflammation, erythrocytes, and unknown functional pathways. The remaining 156 known genes that may be associated with VTE were referred to as candidate genes. By searching OMIM, HPO, and existing literature, we identified only genes that matched the search terms but had no reported direct association with VTE. Matching keywords included: "venous thromboembolism," "venous thrombosis," "thromboembolism," and "pulmonary embolism."
[0097] S32. Screening for risk variations of venous thromboembolism in 35 families;
[0098] like Figure 2 The specific filtering strategy illustrated employs three different filtering strategies based on different gene sets: Filtering Strategy 1 is a mutation screening strategy for genes confirmed to be related to VTE; Filtering Strategy 2 is a mutation screening strategy for genes possibly related to VTE; and Filtering Strategy 3 is a mutation screening strategy for other genes. Mutation screening is then performed for both autosomal dominant and X-linked inheritance patterns.
[0099] Specifically, for filtering strategy 1, a list of 55 identified genetic risk factors for VTE, i.e., a confirmatory gene list, was established through GWAS studies or functional experiments. The filtering rules are as follows:
[0100] a) Screen for variants in 55 identified gene regions;
[0101] b) Filter out non-biselenate or synonymous mutation sites (non-alternative splicing regions);
[0102] c) The mutation is segregated in at least one VTE family and in other families there is no more than one mutation carrier among healthy members older than 45 years.
[0103] Based on the above three conditions, a list of candidate sites for filtering strategy 1 is obtained. Finally, loss-of-function mutations or rare non-synonymous mutations are screened and submitted for Sanger sequencing verification.
[0104] For filtering strategy 2, 156 other known genes or genes that may be associated with VTE (i.e., candidate genes) were collected based on existing literature. The filtering rules are as follows:
[0105] a) Screen for variants in 156 candidate genes;
[0106] b) Filter out non-biselenate or synonymous mutation sites (non-alternative splicing regions);
[0107] c) The mutation is segregated in at least one VTE family and in other families there is no more than one mutation carrier among healthy members older than 45 years.
[0108] Based on the above three conditions, a candidate site list for filtering strategy 2 is obtained. Finally, loss-of-function mutations or rare pathogenic nonsynonymous mutations shared by at least two families or rare pathogenic gene nonsynonymous mutations shared by at least three families are screened and submitted for Sanger sequencing verification.
[0109] For screening strategy 3, the filtering rules for other genes are as follows:
[0110] a) Screening for rare variants located in exon regions and alternative splicing regions;
[0111] b) Filter out non-biselenate or synonymous mutation sites (non-alternative splicing regions);
[0112] c) The mutation is segregated in at least one VTE family and in other families there is no more than one mutation carrier among healthy members older than 45 years.
[0113] Based on the above three conditions, a candidate site list for filtering strategy 3 is obtained. Then, based on the pathogenicity prediction of the site and the results of screening in the population cohort, significant non-homozygous sites in the population are selected to obtain potential new gene sites related to VTE.
[0114] After applying the three screening strategies described above, all cosegregating variants screened from confirmatory genes, candidate genes, and other genes were summarized, especially the mutation sites validated by Sanger, to explain the genetic risk of VTE in each family.
[0115] S4. Analysis of significant mutation sites and gene load testing
[0116] S41. Perform an association analysis between the variation and venous thromboembolism;
[0117] For each variant selected using three different filtering strategies, the allele frequencies of cases and controls in the pedigree mutation dataset and the population validation mutation dataset were analyzed. Association analyses between each variant and disease status were performed using logistic regression models and Fisher's exact test, with an association p < 0.05 defined as a significant mutation.
[0118] S42. Perform a load analysis of the association between genes and venous thromboembolism;
[0119] Mutations were accumulated at the gene level for statistical analysis, examining the differences in mutation counts between VTE patients and healthy controls. For differential analysis, only rare mutations (population frequency less than 0.05) in exons and alternative splicing regions were considered. Mutation types were categorized into seven types: alternative splicing, nonsense mutations, frameshift insertion / deletion, missense mutations, non-frameshift insertion / deletion, synonymous mutations, and unknown mutations. Mutation types with loss of function were selected, with non-frameshift insertion / deletion, synonymous mutations, and unknown mutations considered acceptable. This study considered gene load testing under 11 combinations of mutation types: 7 independent mutations, loss-of-function mutations (alternative splicing, nonsense, frameshift insertion / deletion), loss-of-function mutations plus missense mutations, loss-of-function mutations plus missense mutations and non-frameshift insertion / deletion mutations, and single nucleotide mutations (missense mutations, synonymous mutations). Gene load analysis was performed on each mutation type combination using logistic regression and Fisher's exact test to obtain risk gene ranking results for different combinations. Finally, based on prior knowledge (the known gene sequence associated with VTE), a more reasonable result was selected.
[0120] In this embodiment, the pedigree study cohort collected 216 samples from 35 families, including 22 core families, i.e., at least one child and one parent were diagnosed with VTE, and 13 non-core families. In each core family, three family members who met the criteria were called core family members, i.e., one parent was diagnosed with the disease and the child / child was diagnosed with the disease.
[0121] In this embodiment, 105 samples from VTE families (including 76 confirmed VTE cases and 29 healthy controls) were selected and sent for whole-genome sequencing. A total of 68 variants were validated using the Sanger assay on all available samples from each family. Of these, 41 variants were located on 16 confirmatory genes, and 27 variants were located on 14 candidate genes. In this embodiment, analysis of VTE-related genetic risk genes revealed 68 mutation sites in 30 genes that conformed to a partial family co-segregation pattern. These genes included A4GALT, ACE, CASR, CDH23, CFTR, COL6A2, F2, F5, F8, FGA, GP6, GRK5, JAK2, KNG1, LYST, MYRF, NLRP2, NRAP, PIEZO1, PLACG2, PROC, PROS1, SCARA5, SERPINC1, SH2B3, SPATC1L, SPINK1, TET2, TSPAN15, and VWF.
[0122] Among them, GP6c.G1094A:p.R365H, JAK2c.G380A:p.G127D, and TET2c.G3451T:p.E1151X were co-founded in two or more families. The nonsynonymous mutation GP6c.G1094A:p.R365H was found in families F03 and F32, and this mutation site was also associated with VTE in the population validation cohort (p-value = 0.02). JAK2c.G380A:p.G127D was found in families F32 and F33. The nonsense mutation TET2c.G3451T:p.E1151X was found in families F11 and F34. However, there was no significant difference in the mutation sites of JAK2 and TET2 in the population cohort, as shown in Tables 1 and 2.
[0123] Table 1. Commonly mutated gene loci in two or more families and their significance in the population validation cohort.
[0124]
[0125]
[0126] *Significant (P<0.05).
[0127] Table 3 shows the mutation status of gene mutation sites validated by Sanger in various family samples. A total of 68 mutation sites from 30 genes were found to conform to a partial family co-segregation pattern. Specifically, GP6 c.G1094A:p.R365H, JAK2 c.G380A:p.G127D, and TET2 c.G3451T:p.E1151X were co-founded in two or more families. The nonsynonymous mutation GP6 c.G1094A:p.R365H was found in families F03 and F32, and this mutation site was also associated with VTE in the population validation cohort (p-value = 0.02). JAK2 c.G380A:p.G127D was found in families F32 and F33. The nonsense mutation TET2 c.G3451T:p.E1151X was found in families F11 and F34. However, there was no significant difference in the mutation sites of JAK2 and TET2 in the population cohort (see Tables 2 and 3).
[0128] Table 3. Analysis of genetic risk genes and risk mutations in 35 VTE families and population validation cohorts.
[0129]
[0130]
[0131]
[0132] In this embodiment, see [reference] Figure 2 After WGS data analysis and variant screening, all available samples from 216 individuals across 35 families were validated using Sanger sequencing, yielding 13,954,729 mutation sites. Three screening strategies were applied: 49 genes with 14,794 mutation sites were obtained from 55 confirmed genes; 144 genes with 52,845 mutation sites were obtained from 156 candidate genes; and 12,238 genes with 31,769 mutation sites were obtained from other genes. All co-segregating variants screened from confirmed genes, candidate genes, and other genes were summarized. Specifically, 49 genes with 2,426 mutation sites were screened from confirmed genes; 142 genes with 8,334 mutation sites were screened from candidate genes; and other genes... A total of 5,455 genes and 7,845 mutation sites were screened. In particular, mutation sites were screened using the Sanger assay, resulting in 18 genes and 46 mutation sites in the confirmed genes, 19 genes and 37 mutation sites in the candidate genes, and 60 genes and 66 mutation sites in other genes suspected of being related to VTE. Finally, 68 variants were verified by Sanger sequencing, of which 41 variants were located in 16 confirmed genes and 27 variants were located in 14 candidate genes, involving 32 families.
[0133] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for screening risk genes and risk mutations for polygenic inherited diseases, characterized in that, This includes screening known genetic risk factors and screening risk variants in research samples. Three different filtering strategies are developed based on different gene sets for these risk variants: Filtering Strategy 1, Filtering Strategy 2, and Filtering Strategy 3. Filtering Strategy 1 is a mutation screening strategy for genes confirmed to be associated with the polygenic genetic disease under investigation; Filtering Strategy 2 is a mutation screening strategy for genes possibly associated with the polygenic genetic disease under investigation; and Filtering Strategy 3 is a mutation screening strategy for other genes. The research sample includes a family study cohort and a population validation cohort; the family study cohort is selected from families of a specific region and ethnicity, including diseased members and non-diseased members; the population validation cohort is selected from the general population of the same region and ethnicity as the family study cohort, including sporadic diseased individuals without cause and healthy controls. The risk variants are all cosegregating variants in the screening family study cohort; The filtering strategy 1 involves obtaining all identified genetic risk factors (i.e., confirmatory genes) for the polygenic genetic diseases to be tested through GWAS studies or functional experiments. The filtering rules are as follows: a) Screen for variations in all identified gene regions. b) Filter out non-bicelestem mutations or synonymous mutation sites. c) The mutation is segregated in at least one family lineage, and there is no more than one mutation carrier among healthy members in other families; This yields a list of candidate sites for filtering strategy 1, which is then used to screen for loss-of-function mutations or rare non-synonymous mutations for sequencing verification.
2. The method for screening risk genes and risk mutations for polygenic inherited diseases according to claim 1, characterized in that, The filtering strategy 2 involves collecting other known genes or genes that may be related to the polygenic genetic disease to be tested, i.e., candidate genes, based on existing literature. The filtering rules are as follows: a) Variations in the candidate genes obtained through screening. b) Filter out non-bicelestem mutations or synonymous mutation sites. c) The mutation is segregated in at least one family lineage, and there is no more than one mutation carrier among healthy members in other families; This yields a candidate site list for filtering strategy 2, which is then used to screen for loss-of-function mutations, rare pathogenic nonsynonymous mutations shared by at least two families, or rare pathogenic gene nonsynonymous mutations shared by at least three families, and then validated by sequencing.
3. The method for screening risk genes and risk mutations for polygenic inherited diseases according to claim 1, characterized in that, The filtering rules for filtering strategy 3 are as follows: a) Screen for rare variants located in exon regions and alternative splicing regions. b) Filter out non-bicelestem mutations or synonymous mutation sites. c) The mutation is segregated in at least one family lineage, and there is no more than one mutation carrier among healthy members in other families; This yields a candidate site list for filtering strategy 3. Based on the pathogenicity prediction of the site and the results of screening in the population cohort, significant non-homozygous sites in the population are selected to obtain new gene sites that may be related to the polygenic genetic diseases to be tested.
4. The method for screening risk genes and risk mutations for polygenic inherited diseases according to any one of claims 1 to 3, characterized in that, After obtaining new gene loci that are potentially associated with the polygenic genetic disease to be tested, significant mutation site analysis and gene load testing are performed, including association analysis between variants and the polygenic genetic disease to be tested and load analysis of the association between genes and the polygenic genetic disease to be tested.
5. The method for screening risk genes and risk mutations for polygenic inherited diseases according to claim 4, characterized in that, The association analysis between the variants and the polygenic genetic diseases to be tested was as follows: for each variant obtained by screening with three different filtering strategies, the allele frequencies of cases and control groups in the family study mutation dataset and the population validation mutation dataset were analyzed. The association between each variant and whether or not the disease is present was analyzed based on logistic regression model and Fisher's exact test, respectively. An association with p < 0.05 was defined as a significant mutation. The association load analysis between the genes and the polygenic genetic diseases to be tested was conducted as follows: mutations were accumulated on a gene-by-gene basis and statistically analyzed to determine the difference in the number of gene mutations between patients and healthy controls; mutation types included alternative splicing, nonsense mutations, frameshift insertions / deletions, missense mutations, non-frameshift insertions / deletions, synonymous mutations, and unknown mutations; gene load analysis was performed on each combination of mutation types based on logistic regression models and Fisher's exact test to obtain the risk gene ranking results under different combinations.
6. The method for screening risk genes and risk mutations for polygenic inherited diseases according to claim 1, characterized in that, It also includes whole-genome sequencing and whole-exome sequencing of the research samples, followed by mutation site analysis, annotation and filtering, selecting mutation sites in exon regions and alternative splicing regions and rare mutation sites with a population frequency of less than 0.
05.
7. A screening system for risk genes and risk mutations of polygenic inherited diseases, characterized in that, The invention includes a computer system programmed to perform the steps of the screening method for risk genes and risk mutations of polygenic inherited diseases as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Genetic disease high-throughput sequencing pathogenic mutation screening method based on core family
CN112375815A
High-throughput sequencing variation risk grouping screening method and system
CN113793642A