Genomics-based irritable bowel syndrome risk marker, application and early screening kit
By conducting genome-wide association studies and meta-analyses in multi-ethnic cohorts, novel genetic susceptibility loci and risk variants associated with irritable bowel syndrome were identified, addressing the issue of incomplete genetic structure in different populations and providing a method for early identification and treatment targets.
Patent Information
- Application Number
- CN202511258061.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2026-01-16
AI Technical Summary
The current understanding of the genetic structure of irritable bowel syndrome (IBS) in different populations is incomplete, and the insufficient sample size of the European population has hindered the discovery of the subtype-specific genetic basis of IBS.
By conducting genome-wide association studies (GWAS) and meta-analyses in multiple large-scale, ethnic cohorts, we identified novel genetic susceptibility loci and risk variants associated with irritable bowel syndrome and its subtypes. We used functional annotation, variant level and gene level analysis to screen for potential pathogenic genes and developed an early screening kit for detecting pathogenic genes.
Five new genetic risk variants and potential pathogenic genes were identified, providing the possibility for early identification of irritable bowel syndrome (IBS), demonstrating the role of large-scale GWAS in elucidating the genetic etiology of IBS, and discovering new therapeutic targets.
Smart Images

Figure CN121344178A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a genomics-based risk biomarker for irritable bowel syndrome, its application, and a corresponding early screening kit. Background Technology
[0002] Irritable bowel syndrome (IBS) is a classic gut-brain interaction disorder affecting approximately 10% of the global population, with significant differences in reported prevalence across different regions. This global disparity highlights its widespread impact, although data remains scarce in some areas.
[0003] Clinically, irritable bowel syndrome (IBS) is characterized by persistent or recurrent abdominal pain, bloating, mucus in stool, and changes in bowel habits, often interfering with daily functioning. This condition imposes a significant burden on patients' quality of life, comparable to chronic diseases such as diabetes and hepatitis, and its psychological and social consequences further exacerbate this impact. These data highlight IBS as a major global public health challenge requiring inclusive, multi-ethnic research.
[0004] Most small-scale studies on irritable bowel syndrome (IBS) subtypes have been limited to European populations. Furthermore, the last large-scale genome-wide association study (GWAS) of IBS was conducted a considerable period of time ago. These factors have resulted in an incomplete understanding of the genetic structure of IBS in different populations. Previous GWAS studies on IBS subtypes in European populations, due to insufficient sample size, have hindered the discovery of the specific genetic basis of different IBS subtypes.
[0005] This invention aims to elucidate the genetic etiology of irritable bowel syndrome (IBS). Through comprehensive GWAS and meta-analyses conducted in multiple large-scale, multi-ethnic cohorts (including the All of Us research program, the UK Biobank, etc.), this invention aims to identify novel genetic susceptibility loci and risk variants associated with IBS and its subtypes. A multi-layered approach, including functional annotation, variant-level, and gene-level analyses, is employed to prioritize the screening of potential pathogenic variants and genes for IBS. Summary of the Invention
[0006] To address the shortcomings of existing technologies, the main objective of this invention is to provide a genomics-based risk biomarker for irritable bowel syndrome and its application.
[0007] To achieve the above-mentioned main objectives, on the one hand, the present invention provides genomics-based risk biomarkers for irritable bowel syndrome, which include the following six pathogenic genes: CADM2, PHF2, PCLO, SHISA6, LRP1B, and TANK.
[0008] According to another specific embodiment of the present invention, the method for selecting risk markers includes the following steps:
[0009] A. Meta-analysis of genome-wide association studies (GWAS) of irritable bowel syndrome (IBS) and its subtypes from independent cohorts (e.g., UK Biobanks);
[0010] B. Functionally annotate risk sites for irritable bowel syndrome to identify candidate genes, and prioritize screening for potential pathogenic variants through systematic variation level analysis.
[0011] C. Screen for potential pathogenic genes through gene-level analysis.
[0012] According to another specific embodiment of the present invention, the variation level analysis includes variation annotation, functional genomics, and expression quantitative trait locus analysis.
[0013] According to another specific embodiment of the present invention, gene-level analysis includes transcriptome-wide association studies, co-localization analysis, and Mendelian randomization.
[0014] On the other hand, the present invention provides an application of the above-mentioned genomics-based irritable bowel syndrome risk biomarkers, which are used in the preparation of an early screening kit (prediction kit) for irritable bowel syndrome.
[0015] In another aspect, the present invention provides an early screening kit for predicting irritable bowel syndrome, the kit comprising detection reagents for detecting the following six pathogenic genes: CADM2, PHF2, PCLO, SHISA6, LRP1B and TANK.
[0016] Irritable bowel syndrome (IBS) is the most common gastrointestinal disorder worldwide, significantly impacting the quality of life of 10%-15% of the population. Although some related factors have been identified, its pathophysiological mechanisms remain unclear.
[0017] This invention utilizes genome-wide association studies (GWAS) to screen for the most likely pathogenic genes of irritable bowel syndrome (IBS), enabling early identification of IBS. This invention offers the following advantages:
[0018] This invention not only demonstrates the role of the latest large-scale GWAS in elucidating the genetic etiology of irritable bowel syndrome (IBS), but also identifies five new genetic risk variants, potential unreported pathogenic genes, and therapeutic targets for IBS.
[0019] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description
[0020] Figure 1This is a summary of the cohorts used, the genome-wide association studies (GWAS) conducted, and the loci detected. The analyzed population datasets comprised 1,095,668 individuals from four sources: All of Us (10,685 patients, 234,703 controls), FinnGen (9,323 patients, 301,931 controls), GERA (3,117 patients, 53,520 controls), and the UK Biobank (UKB) and Belgian cohorts (53,400 patients, 433,201 controls). Genome-wide association studies (Manhattan plot) and meta-analyses identified significant loci (genomic regions associated with irritable bowel syndrome), with novel loci representing previously unreported associations. Functional / causal analyses (transcriptome-wide association studies, colocalization analysis, multigene priority scoring, Mendelian randomization) explored biological mechanisms, and pharmacodynamic tools (DRUGBAI, Open Targets, DGIdb) assessed the therapeutic potential of related genes.
[0021] Figure 2 is a schematic diagram of a genome-wide association study, in which:
[0022] Figure 2A This is a Manhattan plot of genome-wide association studies. Genome-wide association results from a meta-analysis of irritable bowel syndrome (IBS) (76,525 patients and 1,023,355 controls). Twelve identified risk loci are marked with rsID.
[0023] Figure 2B This study reveals gene expression patterns associated with the etiology of irritable bowel syndrome (IBS) by examining the causal effects of relevant variants on gene expression in various tissues. Dark blue indicates higher gene expression levels in that tissue.
[0024] Figure 3 shows the results of transcriptome-wide association and proteome-wide association, displaying only the transcriptome-wide significant genes after Bonferroni correction for P-value thresholds. LINC02803 showed the most significant association in the transcriptome-wide association study, among which:
[0025] Figure 3A The text represents genes that are fully significant in the transcriptome in the Irritable Bowel Syndrome Meta-Dataset Transcriptome-Wide Association Study. The red dashed line represents the proteome significance threshold (z-score <-4.64 or >4.64).
[0026] Figure 3B This represents all significant genes in the transcriptome of a transcriptome-wide association study of a mixed irritable bowel syndrome dataset.
[0027] Figure 3C This represents the fully significant transcriptome genes in the transcriptome-wide association study of the constipation-predominant irritable bowel syndrome dataset.
[0028] Figure 4 The gene prioritization results are presented as candidate genes supported by at least four pieces of evidence, along with the evidence provided for each analysis. A total of 10 pieces of evidence are used for gene prioritization, with each column representing a type of supporting evidence. Mapped_gene: Genes closest to the dominant SNP; eGene: Genes whose expression is associated with the dominant SNP in the quantitative trait locus dataset; Coloc: Colocalization evidence (colocalization PP4 > 0.8); Magma: Genes enriched in specific tissues and cell types; PoPS: Genes in the top 2% of multigene prioritization scores; SMR: Genes with Mendelian randomization p-values below the Bonferroni-corrected p-threshold based on Summary data; TWAS: Genes with p-values below the Bonferroni-corrected p-threshold from transcriptome-wide association studies; Functional_genomics: Genes near functional SNPs identified by functional genomics; Annov: The closest gene annotated to risk SNPs and functional variants; Drug_target: Genes with documented clinical drug targets.
[0029] Figure 5 This indicates the genetic and phenotypic associations of irritable bowel syndrome (IBS) with other traits; phenotypic (causal association, left) and genetic risk associations (shared genetic structure, right) between IBS and other traits. Detailed Implementation
[0030] Example 1
[0031] Irritable bowel syndrome (IBS) is the most common gastrointestinal disorder worldwide, significantly impacting the quality of life of 10%-15% of the population. Although some related factors have been identified, its pathophysiological mechanisms remain unclear.
[0032] This invention utilizes data from the All of Us study to conduct genome-wide association studies (GWAS) on the overall irritable bowel syndrome (IBS) cohort and its subtypes. Statistical adjustments were made for sex, age, and the top 10 principal components to control for potential confounding factors. A GWAS meta-analysis of IBS and its subtypes was performed on 76,525 IBS patients and 1,019,143 controls from five independent cohorts (All of Us study, UK Biobank, Bellygenes, the Genetic Epidemiology of Aging Study, and FinnGen). Functional annotation of IBS risk sites was performed to identify candidate genes. Potential pathogenic variants were prioritized through systematic level-of-variance analysis (including variant annotation, functional genomics, and quantitative trait locus analysis). Potential pathogenic genes were screened through gene-level analysis (including transcriptome-wide association studies, co-localization analysis, and Mendelian randomization).
[0033] In the "All of Us Study" cohort of 10,685 patients and 234,703 controls, this invention identified two genome-wide significant differences (P < 5 × 10⁻⁶). -8 Independent irritable bowel syndrome (IBS) susceptibility loci were identified. A GWAS meta-analysis revealed 12 new independent signals (P < 5 × 10⁻⁶). -8 ) and 5 unreported loci (rs143348218, rs3748618, rs6432674, rs9536395, rs976714). Subtype specificity based on subtype-specific GWAS meta-dataset (P<5×10) -8 The study results included 3 loci (3 new independent loci) of constipation-predominant irritable bowel syndrome (IBS-C) and 44 loci (2 new independent loci) of mixed irritable bowel syndrome (IBS-M).
[0034] By integrating multi-level analysis evidence, this invention screened out the most likely pathogenic genes for irritable bowel syndrome (IBS), including CADM2, PHF2, PCLO, SHISA6, LRP1B, and TANK. Notably, the TANK gene has not been previously reported to be associated with IBS.
[0035] I. Methods
[0036] 1. "All of Us Research Program" dataset
[0037] The All of Us Study's Genome Center generates genomic data from biological samples provided by participants. The irritable bowel syndrome (IBS) study population includes individuals from multiple ethnic groups (10,685 patients and 234,703 controls).
[0038] 2. Public GWAS aggregated statistics
[0039] The publicly available GWAS summary statistics used in the meta-analysis of this study are as follows: Figure 1 As shown.
[0040] 3. FinnGen project dataset
[0041] FinnGen is a biobank launched in Finland. https: / / www.finngen.fi / This invention downloaded the irritable bowel syndrome dataset from FinnGen (R9), which includes 9,323 patients and 301,931 controls (gs: / / finngen-public-data-r9 / summary_stats / ). This data is publicly available and requires no additional approval.
[0042] 4. UK Biobank (UKB) and Bellygenes datasets
[0043] This cohort combined patients with irritable bowel syndrome (IBS) identified in the UK Biobank and the Bellygenes program. IBS GWAS data from the UK Biobank (40,548 patients and 290,220 controls) and the Bellygenes program (12,852 patients and 139,981 controls) were used for the meta-analysis. The IBS dataset was downloaded from the European Bioinformatics Institute GWAS catalog (https: / / www.ebi.ac.uk / gwas / ), accession number GCST90016564. For the subtype IBS GWAS meta-analysis, this invention used subtype IBS datasets from two independent European cohorts (UK Biobank and Lifelines), accession numbers GCST90243959, GCST90243960, and GCST90243961.
[0044] 5. Data sets for genetic epidemiological studies of the elderly
[0045] The Genetic Epidemiology of Adult Health and Aging (GERA) study is an important resource for genetic research on aging. The irritable bowel syndrome (IBS) study population included 3,117 patients of European descent and 53,520 controls of European descent.
[0046] 6. Quality control procedures and GWAS for genetic data from the "All of Us" cohort.
[0047] To ensure data quality and result accuracy, a series of quality control procedures were performed on genetic variations using PLINK v2.0
[18] . For association analysis, only SNPs with a detection rate > 0.98, an imputation quality score INFO ≥ 0.98, a minor allele frequency ≥ 0.005, no insertion / deletion, no multiple alleles, and located on an autosome were used. At the same time, SNPs that deviated from Hardy-Weinberg equilibrium (P < 1 × 10⁻⁶) were removed. -6 The variance of ) was then assessed. Linkage disequilibrium (LD) pruning was performed using a sliding window of 50 SNPs with a step size of 5 SNPs to remove paired variances with r² > 0.2 within the window. Heterozygosity was calculated, and individuals whose values deviated from the mean by more than three standard deviations were excluded. Patient and control definitions were determined based on patient IDs identified via ICD-10 codes. Subtypes were defined as constipation-predominant irritable bowel syndrome (IBS-C), diarrhea-predominant irritable bowel syndrome (IBS-D), and mixed irritable bowel syndrome (IBS-M). A GWAS was then performed using PLINK v2.0 on a logistic regression model, adjusting for sex, age, and the top 10 principal components as covariates.
[0048] 7. Statistical Analysis
[0049] Meta-analysis
[0050] For the meta-analysis of irritable bowel syndrome (IBS) cohorts, this invention used METAL software to integrate five datasets: Allof Us, FinnGen, GERA, the UK Biobank, and Bellygenes. The meta-analysis employed a sample size-weighted Z-score to address case-control imbalance. For IBS subtype groups, this invention integrated subtype data from Allof Us, the UK Biobank, and Lifelines databases.
[0051] Identification of genome-wide significant signals and loci: A locus is defined as a genomic segment with ≥1 dominant variant, which defines a genetically independent signal. Any segment containing only ≤1 dominant variant achieves genome-wide significance (P<5×10⁻⁶). -8 Regions of SNPs in the region were ignored in further site definition or analysis because single variant associations may be unreliable. LD-based clustering was performed using PLINK, utilizing an ancestry-matched 1000 Genomes Project reference panel and r 2 A threshold of 0.01 was used to identify dominant SNPs within each of the above salient loci. Salient loci were named according to their chromosomal and the start and end positions of their salient regions (in Mb) (e.g., chr5:129.52Mb–131.88Mb). A region association map (LocusZoom) for each locus was created using the R script “LocusZoom-like plots”. Gene start and end positions and gene names (HGNC gene symbols) in the gene trajectory information were defined using a GENCODE Human Genome Construction 37 (version 38) gene annotation file. The complete gene annotation file was reduced to unique gene names. Gene trajectory information was colorized according to the p-value of MAGMA-based gene association analysis. After creating these region association maps, each locus with multiple dominant SNPs was visually examined to determine whether the labeled SNPs represented a single association signal or multiple independent spatially separated association signals.
[0052] Genetic correlation analysis
[0053] Genetic correlation analysis was performed using LDSC[20,23]. LDSC was performed according to the guidelines in the official documentation and using the default parameters.
[0054] Previously unreported genome-wide significant sites and signals
[0055] To identify newly detected sites and signals, this invention retrieved all previously reported literature on SNPs associated with irritable bowel syndrome (IBS). On May 22, 2025, a GWAS catalog was searched using the keywords “irritable bowel syndrome” and “IBS symptom measurement.” The results revealed 60 studies published between May 2014 and November 2024. All significant SNPs identified in these 60 studies were downloaded. Site boundaries were converted to build 38 genomic locations using Liftover. These build 38 genomic locations were then used to identify whether any previous GWAS signals fell within the aforementioned significant site boundaries. The LD between the dominant SNP and all previously identified SNPs within that site was calculated using a 1000 Genomes reference panel in LDlinkR matched with the dominant SNP. LD values were categorized into high (r² ≥ 0.8), medium (0.8 > r² ≥ 0.5), low (0.5 > r² ≥ 0.1), or weak LD (r² < 0.1) groups to assess the similarity between previously reported associated signals and the dominant SNP. A locus was considered previously unreported only if none of the 60 irritable bowel syndrome-related GWAS studies reported a significant SNP within the locus boundaries, or if only a weak LD existed between the dominant SNP and any previously reported SNP.
[0056] Functional and genetic annotation of GWAS loci
[0057] Functional and genetic annotations of all genome-wide significant SNPs and their marker SNPs in the irritable bowel syndrome meta-analysis (GWAS) were performed using ANNOVAR
[26] integrated in FUMA. All annotations were based on the hg19 (GRCh37) reference genome, with systematic characterization of genomic functional regions and variant types. Specifically, the hg19 refGene gene-based annotation database included in ANNOVAR was used (raw data from RefSeq Gene, last updated on August 22, 2020, in the UCSC Genome Browser). This invention uses annotate_variation.pl, with default parameters, to annotate the most recent gene that dominates SNPs and functional variants.
[0058] Transcriptome-wide association study
[0059] This invention utilizes a transcriptome-wide association study (TWAS) by integrating meta-analysis GWAS data for irritable bowel syndrome with pre-calculated GTEx v8 tissue-specific weights (gastrointestinal, mental health, and whole blood weight files). The official LD reference panel is used to interpret linkage disequilibrium. FDR correction is applied to the results for each chromosome separately.
[0060] Colocalization and study of genetically associated signals in irritable bowel syndrome
[0061] The dominant SNPs of independent signals in irritable bowel syndrome (IBS) GWAS were queried from GTEx data (gastrointestinal tissues, mental tissues, and whole blood weighted files) using meta-analysis to identify cis-expressed quantitative trait loci (cis-eQTLs). For all SNPs with significant eQTLs, colocalization analysis was performed using the “coloc” R package, using only SNPs within a 1 Mb window centered on the dominant variant region
[29] . A posterior probability (PP.H4) ≥ 0.8 was considered strong evidence of colocalization. This hypothesis (H4) indicated that the association between IBS and gene expression was due to the same functional variant. Genes with strong evidence of colocalization in all included tissues were summarized to identify colocalized genes associated with IBS.
[0062] Mendelian randomization based on summary data
[0063] SMR analysis was used to integrate BrainMeta V2 human brain eQTL data, CAGE eQTL pooled data, and GWAS meta-analysis data for irritable bowel syndrome. The SMR binary (BESD) eQTL dataset was downloaded from the SMR website and included 2,865 human cortical samples (https: / / yanglab.westlake.edu.cn / software / smr / ). Genes with PHEIDI > 0.05 and FDR < 0.05 were considered candidate pathogenic genes.
[0064] eQTL annotations
[0065] This invention performs eQTL annotation to identify potential target genes regulated by dominant or functional SNPs. eQTL annotation analysis was performed using the FUMA platform with default parameters to identify genes associated with irritable bowel syndrome (IBS) expression. Six IBS-related datasets were systematically compiled, including brain tissue, digestive system tissue, and whole blood samples from GTEx v8, as well as datasets from the DICE, PsychENCODE, eQTLGen, CMC, and BIOSQTL databases.
[0066] Gene priority
[0067] To identify high-confidence candidate genes for irritable bowel syndrome (IBS), a two-step gene prioritization strategy integrating ten complementary methods was employed. First, candidate genes were prioritized using the Multigene Priority Score (PoPS), with the top 2% of genes in terms of PoPS score considered as potential risk genes. Second, evidence from other analyses, including genomic location, eQTL, TWAS, co-localization analysis, SMR, MAGMA, and PoPS, was also used for prioritization. Each piece of evidence was assigned a score, and genes supported by at least four pieces of evidence were considered high-confidence risk genes.
[0068] Causal associations with other phenotypes and metabolites
[0069] This invention investigated the potential causal associations between irritable bowel syndrome (IBS) and all systemic phenotypes of the disease. IBS datasets from a meta-analysis were used as exposure factors, and causal inference was performed using the TwoSampleMR package. Inverse variance weighting (IVW) was employed as the primary analytical method. Associations with a p-value less than 0.05 were considered significant causal relationships. Sensitivity analyses were performed to assess the robustness of the results. For significant associations that failed the sensitivity analysis, Steiger filtering was used to test for directionality, and the MR-PRESSO method was used to handle level pleiotropy to exclude outliers (SNPs). Associations with heterogeneity p-values ≥ 0.05 were considered reliable and retained as positive results. Subsequently, LDSC was performed on these positive phenotypes identified by MR to assess the genetic associations between IBS and each phenotype. Genetic associations with p-values < 0.05 were considered statistically significant, indicating the existence of shared genetic structures and further supporting the potential biological link between IBS and the corresponding phenotypes.
[0070] Heritability estimation of irritable bowel syndrome and genetic correlation between irritable bowel syndrome subtypes
[0071] SNP-based heritability (h²SNP) and genetic correlation (rg) were estimated using LDSC. Furthermore, SNP-based heritability was estimated using LDAKv5.2 under the BLD-LDAK model. Pairwise genetic correlation was calculated using LDSC, and rg between each subtype pair was estimated by regressing the product of the SNP-level Z-scores onto their corresponding LD scores. Significant rg values were interpreted as evidence of shared genetic structure among irritable bowel syndrome subtypes.
[0072] Drug availability of candidate pathogenic genes
[0073] To assess the druggability of candidate genes for irritable bowel syndrome (IBS), genes with high levels of evidence were first screened from gene prioritization analysis. These genes were then cross-referenced with three major drug databases: DGIdb (https: / / www.dgidb.org / ), DrugBank (https: / / www.drugbank.ca / ), and Open Targets (https: / / www.opentargets.org / ). Genes with genetic associations to approved or investigational therapies in these databases were considered potentially druggable for IBS, and this evidence was incorporated into the gene prioritization framework as additional supporting criteria.
[0074] II. Results
[0075] 1. Genome-wide association studies identified two new risk sites in the "All of Us" cohort.
[0076] This study first conducted a genome-wide association study (GWAS) on the irritable bowel syndrome (IBS) cohort within the "All of Us Study Program" cohort, which included 245,388 participants. After rigorous quality control, the genotypes of 10,685 patients and 234,703 controls were retained for the GWAS. Two genome-wide significant loci were identified, one located on chromosome 5 (near the MCC gene, dominated by single nucleotide polymorphism rs1459986827, P = 1.677 × 10⁻⁶). -8 The odds ratio was 1.57. Another one was located on chromosome 22 (no previously reported associated gene, dominant single nucleotide polymorphism rs1345599196, P = 5.832 × 10⁻⁶). -9 The ratio is 1.14. Figure 1 ).
[0077] In constipation-predominant irritable bowel syndrome, three novel genome-wide significant loci were detected, one of which is located on chromosome 10 (no previously reported associated gene, dominant single nucleotide polymorphism rs558309812, P = 9.184 × 10⁻⁶). -9 The odds ratio was 3.15. Another one was located on chromosome 7 (near the TMEM196 gene, dominated by single nucleotide polymorphism rs117869341, P = 2.568 × 10⁻⁶). -9 The odds ratio was 2.34. There was also one located on chromosome 8 (near the TRAPPC9 gene), with the dominant single nucleotide polymorphism rs145822365, P = 1.260 × 10⁻⁶. -8 The ratio is 2.14. Figure 1In mixed irritable bowel syndrome, two genome-wide significant loci were identified, one located on chromosome 14 (no previously reported associated gene, dominant single nucleotide polymorphism rs142881434, P = 3.305 × 10⁻⁶). -8 The odds ratio was 2.34. Another one was located on chromosome 2 (no previously reported associated gene, dominant single nucleotide polymorphism rs148231700, P = 3.592 × 10⁻⁶). -8 The ratio is 2.27. Figure 1 ).
[0078] No genome-wide significant loci were identified in diarrhea-predominant irritable bowel syndrome. All of the above loci remained independent and significant after linkage disequilibrium pruning.
[0079] 2. Meta-analysis of multi-cohort genome-wide association studies and annotation of genomic risk sites
[0080] To gain a deeper understanding of the genetic structure of irritable bowel syndrome (IBS) in all populations, this invention further conducted the largest meta-analysis of genome-wide association studies (GWAS) for IBS, integrating data from five cohorts, including (1) the "Allof Us" cohort, (2) the Elderly Genetic Epidemiology Study cohort, (3) the UK Biobank cohort, (4) the Bellygenes cohort, and (5) the FinnGen cohort, encompassing a total of 76,525 patients and 1,023,355 controls. Figure 1 A meta-analysis across ancestry identified 152 risk loci, of which 119 were previously unreported as associated with irritable bowel syndrome (IBS). Genome-wide significant single nucleotide polymorphism (SNP) associations were detected in a meta-analysis of genome-wide association studies of IBS across all populations (P < 5 × 10⁻⁶). -8 ), corresponding to 12 independent loci, tagged by rs1248825 (CADM2), rs10156602 (PHF2), rs3748618 (RFWD2), rs976714 (PCLO), rs67427799 (LRP1B), rs6432674 (TANK), rs1546559 (SHISA6), rs143348218 (CACNA2D3), rs34209273, rs34365748, rs451637 (C4B), and rs9536395. Figure 2A(See Table 1). Among them, five loci (rs143348218, rs3748618, rs6432674, rs9536395, and rs976714) were identified as novel loci that had not been previously reported to be associated with irritable bowel syndrome. After excluding pseudogenes, three new risk genes were identified: CACNA2D3, C4B, and TANK.
[0081] Meta-analyses of genome-wide association studies (GWAS) of various subtypes of irritable bowel syndrome (IBS) were conducted on (1) the “All of Us” cohort, (2) the UK Biobank cohort, and (3) the Lifelines cohort. These included 14,552 patients with mixed IBS, 9,901 with diarrhea-predominant IBS, 5,930 with constipation-predominant IBS, and 321,117 controls with mixed IBS, 319,117 controls with diarrhea-predominant IBS, and 320,937 controls with constipation-predominant IBS. Figure 1 Meta-analysis identified 44 risk loci, and ultimately identified 4 independent risk loci with genome-wide significance in mixed irritable bowel syndrome (P<5×10⁻⁶). -8 The markers were rs142881434, rs148231700, rs2048419 (MFHAS1), and rs7718889. Simultaneously, three new independent risk sites were identified in constipation-predominant irritable bowel syndrome, marked by rs117869341 (TMEM196), rs145822365 (TRAPPC9), and rs558309812.
[0082] 3. Specific genetic relationships among irritable bowel syndrome subtypes
[0083] To investigate the genetic associations among different subtypes of irritable bowel syndrome (IBS), this invention used LD score regression to assess the genetic relationships among diarrhea-predominant IBS, mixed IBS, and constipation-predominant IBS. LD score regression showed a positive correlation between constipation-predominant IBS and diarrhea-predominant IBS (rg = 0.514, P = 0.003), while a strong correlation existed between constipation-predominant IBS and mixed IBS (rg = 0.80, P = 9.74 × 10⁻⁶). -9 Notably, the highest genetic association was observed between mixed irritable bowel syndrome (IBS) and diarrhea-predominant IBS (rg = 0.86, P = 2.87 × 10⁻⁶). -17 ).
[0084] 4. Integrated analysis and annotation of irritable bowel syndrome risk genes
[0085] Given that risk variants identified by genome-wide association studies (GWAS) are primarily located in non-coding regions, most of these variants likely influence irritable bowel syndrome (IBS) by regulating gene or protein expression levels. Therefore, transcriptome-wide association studies were used to identify risk genes. In transcriptome-wide association studies, 231 risk genes were identified through Bonferroni correction. Figure 3A LINC02803 was the highest-ranking risk gene (P = 1.74 × 10⁻⁶). -9 To investigate whether shared pathogenic variants drive quantitative trait loci and genome-wide association study (GWAS) signals, colocalization analysis was performed using the COLOC R package [37,38]. Colocalization analysis identified 40 significant colocalization signals when the posterior probability of colocalization hypothesis 4 (i.e., PP4, which indicates that colocalized quantitative trait loci are associated with GWAS) was >0.8.
[0086] Finally, Mendelian randomization based on Summary data was used to infer potential pathogenic genes. Instrumental variable-dependent heterogeneity was used to filter heterogeneity caused by linkage disequilibrium. Mendelian randomization analysis based on Summary data identified 25 significant candidate genes at 29 irritable bowel syndrome risk sites, all of which passed the heterogeneity test (PHEIDI>0.05).
[0087] 5. Prioritize gene sequencing for annotation of key risk genes
[0088] Through a series of integrated analyses, including transcriptome-wide association studies, colocalization analyses, and Mendelian randomization based on summary data, multiple risk genes / proteins for irritable bowel syndrome (IBS) were identified. Using functional genomics and fine mapping methods, combined with the results of variant and gene-level analyses, this invention nominates potential pathogenic variants for IBS.
[0089] Specifically, for dominant single nucleotide polymorphisms (SNPs) and potentially pathogenic variants, candidate target genes were assigned to these variants based on the most recent gene and quantitative trait locus annotations from ANNOVAR. In the meta-analysis, genes or proteins reaching significance thresholds in transcriptome-wide association studies, co-localization analyses, Mendelian randomization based on summary data, and co-localization were considered candidate pathogenic genes. Furthermore, a multi-gene priority score was used to nominate potential risk genes for genome-wide association study signals (the top 2% of genes in the multi-gene priority score). Each piece of evidence was assigned 1 point, and the gene with the highest score was considered more likely to be a potential pathogenic gene for irritable bowel syndrome (IBS).
[0090] Based on all the evidence, this invention preferentially screened 23 high-confidence risk genes for irritable bowel syndrome, such as CADM2, PHF2, PCLO, SHISA6, LRP1B, and TANK. Figure 4 Among these high-confidence irritable bowel syndrome (IBS) risk genes, 11 genes, such as CADM2, PHF2, PCLO, SHISA6, FAM120A, and LRP1B, were also preferentially screened in previous studies, further supporting these candidate genes as potential pathogenic genes for IBS. Notably, the remaining 14 genes are new findings in this study. Particularly noteworthy is the TANK gene, which has not been previously reported to be associated with IBS.
[0091] 6. Causal and genetic links in irritable bowel syndrome
[0092] In Mendelian randomization analysis, using irritable bowel syndrome (IBS) as an exposure factor, causal relationships between IBS and 27 diseases or metabolites were identified. Figure 5 Of particular note is gastroesophageal reflux disease (P = 1.25 × 10⁻⁶). -10 Asthenia (P = 2.08 × 10⁻⁶) -10 ) and major depressive disorder (P = 3.94 × 10) -8 Subsequently, based on the positive results of the Mendelian randomization analysis, a regression analysis of the LD score was performed to explore genetic associations. This analysis revealed genetic associations between 11 diseases or metabolites and irritable bowel syndrome (IBS). These diseases or metabolites cover multiple categories, including mental illnesses (such as attention deficit hyperactivity disorder and schizophrenia), digestive system diseases (such as gastroesophageal reflux disease and acute pancreatitis), metabolic / cardiovascular related indicators (such as low-density lipoprotein and C-reactive protein), and hypothyroidism. Specifically, in the LD score regression analysis, the p-value for gastroesophageal reflux disease was 1.25 × 10⁻⁶. - 10. The genetic correlation was 62.82%, highlighting a close genetic link between it and irritable bowel syndrome. Figure 5 ).
[0093] 7. Analysis of drug target database for candidate genes of irritable bowel syndrome
[0094] To assess the druggability of candidate genes for irritable bowel syndrome (IBS), three major drug databases—DGIdb, DrugBank, and Open Targets—were queried to identify genes genetically associated with approved or investigational therapies. Eight candidate genes (CACNA2D3, P2RY12, RTN4, LRP1B, SLC5A6, TANK, CADM2, and COP1) were identified as genetically associated with compounds with established or potential therapeutic applications.
[0095] III. Discussion
[0096] This study pioneered the first genome-wide association study (GWAS) for irritable bowel syndrome (IBS) using the newly released All of Us research program database. This database contains multi-ethnic population data from 10,685 patients and 234,703 controls, identifying two novel risk loci and five subtype-specific risk loci for mixed IBS-M and constipation-predominant IBS-C. To further deepen our understanding, this study conducted the largest IBS GWAS meta-analysis to date, integrating data from five cohorts to form a massive sample of 76,525 patients and 1,023,355 controls (total sample size N = 1,095,668), identifying 119 previously unreported risk loci. Notably, this study also validated six known risk loci (CADM2, PHF2, PCLO, SHISA6, FAM120A, and LRP1B) reported in previous studies, confirming their importance. Overall, these findings identified five new genetic risk loci (rs143348218, rs3748618, rs6432674, rs9536395, rs976714), six validated pathogenic genes (CADM2, PHF2, PCLO, SHISA6, LRP1B, and TANK), and 14 new potential pathogenic genes, greatly advancing our understanding of the genetic mechanisms of irritable bowel syndrome with an unprecedented sample size and statistical power.
[0097] Building upon large-scale GWAS meta-analysis, this invention also conducted a series of in-depth post-GWAS analyses to identify potential causal risk variants and genes for irritable bowel syndrome (IBS). Precisely locating pathogenic variants from genetic risk loci remains a major challenge in IBS research. These analyses identified potential pathogenic variants from reported loci, providing an important starting point for further mechanistic studies.
[0098] In addition to the analysis at the variation level, this invention also conducted a series of gene-level analyses, including transcriptome-wide association studies (TWAS), Mendelian randomization (SMR) based on summary data, colocalization analysis, and polygenic priority scoring (PoPS), to identify potential pathogenic genes. Several genes previously reported in irritable bowel syndrome (IBS)-related GWAS studies play multifaceted roles in the disease. CADM2 is a synaptic cell adhesion molecule involved in the formation of neural circuits and associated with substance dependence disorders, which link the gut-brain axis. PHF2 is known to be associated with neuroticism, depression, and autism; it also contributes to brain development and is expressed in enteric nerve fibers, suggesting a possible role in bidirectional communication between the gut and brain. PCLO is essential for presynaptic cytoskeleton matrix and synaptic vesicle transport and is associated with bipolar disorder and major depressive disorder, further suggesting its potential involvement in the pathogenesis of IBS. SHISA6 regulates synaptic transmission and neuronal development and may affect IBS by modulating glutamate receptor activity and signaling pathways. These genes highlight the complex genetic overlap between neurological disorders and irritable bowel syndrome, warranting further investigation into their causal role in the disease.
[0099] Genes involved in newly identified loci associated with irritable bowel syndrome (IBS) either influence mood or anxiety disorders, are expressed in the nervous system, or both. This study identified the TANK gene as a risk locus for IBS, suggesting its potential involvement in immune signaling and inflammation mechanisms. As a regulator of the TRAF protein, TANK inhibits NF-κB activation by isolating TRAF2, which may disrupt mucosal immunity and contribute to typical IBS features such as visceral hypersensitivity. TANK's role in antiviral immunity and DNA damage responses further suggests its potential involvement in intestinal epithelial repair and neuroimmune communication. NSUN2, essential for tRNA methylation, may affect the gut-brain axis; PDIA4, involved in endoplasmic reticulum protein folding, may contribute to IBS through gastrointestinal inflammatory pathways. ZC3H7B may regulate IBS-related immune responses. These findings highlight novel genetic targets for IBS research, warranting further investigation to clarify their roles in disease pathogenesis.
[0100] This study has several limitations. The insufficient sample size from non-European populations may have hindered the identification of population-specific genetic associations and limited the generalizability of the findings. Furthermore, limitations of current databases prevented protein-level analyses, which could have provided crucial insights into the functional significance of identified genetic variations. Future research should prioritize establishing larger, more diverse cohorts, particularly targeting underrepresented ethnic groups, and incorporating proteomics data through self-built cohorts to validate genetic findings and elucidate the pathophysiological mechanisms of irritable bowel syndrome (IBS).
[0101] This study not only demonstrates the role of the latest large-scale genome-wide association studies (GWAS) in elucidating the genetic etiology of irritable bowel syndrome (IBS), but also identifies five novel genetic risk variants, potential unreported pathogenic genes, and therapeutic targets for IBS. These findings provide new insights into the etiology of IBS and highlight potential therapeutic intervention targets.
[0102] While the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the scope of the invention. Any person skilled in the art can make modifications without departing from the scope of the invention; all equivalent modifications made in accordance with the invention should be covered by the scope of the invention.
Claims
1. A genomic-based irritable bowel syndrome risk marker characterized in that, The risk markers include the following six pathogenic genes: CADM2, PHF2, PCLO, SHISA6, LRP1B and TANK.
2. The irritable bowel syndrome risk marker of claim 1, wherein, The selection method of the risk markers includes the following steps: A, performing a meta-analysis of whole genome association studies of irritable bowel syndrome and its subtypes on irritable bowel syndrome patients and control persons from independent cohorts; B, performing functional annotation on irritable bowel syndrome risk loci to identify candidate genes, and preferentially screening potential pathogenic variations through systematic variation level analysis; C, screening potential pathogenic genes through gene level analysis.
3. The irritable bowel syndrome risk marker of claim 2, wherein, The variation level analysis includes variation annotation, functional genomics and expression quantitative trait locus analysis.
4. The irritable bowel syndrome risk marker of claim 3, wherein, The gene level analysis includes transcriptome-wide association study, colocalization analysis and Mendelian randomization.
5. Use of a genomic-based irritable bowel syndrome risk marker as claimed in claim 1, wherein, The risk markers are applied to the preparation of an irritable bowel syndrome early screening kit.
6. An early screening kit for irritable bowel syndrome prediction, characterized by, The kit includes detection reagents for detecting the following six pathogenic genes: CADM2, PHF2, PCLO, SHISA6, LRP1B and TANK.