Method for analyzing and identifying oral cancer biomarker related to inflammatory bowel disease

By analyzing the causal relationship between IBD and oral cancer, combined with single-cell RNA sequencing and functional enrichment analysis, the NFKBIA gene was identified as a potential therapeutic target for IBD-related oral cancer, solving the problem of lack of effective biomarkers in the prior art, and achieving a deeper understanding of IBD-related oral cancer and possible treatment directions.

CN120138155APending Publication Date: 2025-06-13GUANGDONG GENERAL HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510416314.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, the causal relationship between inflammatory bowel disease (IBD) and oral cancer is not clear, and effective diagnostic and predictive biomarkers are lacking.

Method used

By obtaining the summary data of the entire genome association study of IBD, the Mendel randomization method was used to analyze the causal relationship between IBD and oral cancer, and IBD-related genes were determined. Then, oral cancer tissue samples were collected, single-cell RNA sequencing was performed, cluster analysis was performed using dimensionality reduction technology, different cell types were identified, and gene expression differences between different cell types were compared. The biological pathways and functions involved in the NFKBIA gene were determined through functional enrichment analysis, and combined with Mendel randomization analysis and single-cell RNA sequencing results, the potential IBD-related oral cancer treatment targets of the NFKBIA gene were identified.

Benefits of technology

By single-cell RNA sequencing of oral cancer specimens, delineating cell clusters and gene expression patterns, revealing cellular heterogeneity and gene expression dynamics during oral cancer disease progression, and identifying possible therapeutic targets related to IBD.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120138155A_ABST
    Figure CN120138155A_ABST
Patent Text Reader

Abstract

The invention discloses a method for analyzing and identifying an inflammatory bowel disease related oral cancer biomarker, and belongs to the technical field of bioengineering.The method comprises the following steps that 1, IBD whole genome correlation research summary data is obtained; 2, analyzing a causal relationship between IBD and the oral cancer by using a Mendel randomization method, and determining IBD related genes; 3, collecting an oral cancer tissue sample, and carrying out single-cell RNA sequencing; according to the invention, single-cell RNA sequencing is carried out on an oral cancer specimen, and a cell cluster and a gene expression mode are described. Cell clusters and types are displayed by using t-SNE and UMAP, key modes of gene expression are defined by using a lattice diagram, and pathways related to NFKBIA expression are analyzed through GSEA. The experimental result explains the cell heterogeneity and gene expression kinetics in the oral cancer disease progression process, and reveals possible therapeutic targets related to IBD.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioengineering technology, and specifically to a method for analyzing and identifying biomarkers related to oral cancer in inflammatory bowel disease. Background Art

[0002] Inflammatory bowel disease (IBD) is a chronic inflammatory state of the gastrointestinal tract, characterized by two different diseases, Crohn's disease and ulcerative colitis. It leads to a high risk of different types of cancers, including oral cancer. Nowadays, the intersection of inflammatory bowel disease (IBD) and cancer research has attracted great attention, especially through Mendelian randomization (MR) studies. MR provides a powerful framework for exploring causal relationships by using genetic variants as instrumental variables.

[0003] However, in the prior art, the causal relationship between IBD and oral cancer is not clear, and there is a lack of effective diagnostic and predictive biomarkers. Therefore, a method for analyzing and identifying biomarkers related to oral cancer in inflammatory bowel disease is proposed.

[0004] The above information disclosed in this background art is only used to increase the understanding of the background art of the present invention. Therefore, it may include prior art that is not known to those of ordinary skill in the art. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. For this reason, an object of the present invention is to propose a method for analyzing and identifying biomarkers related to oral cancer in inflammatory bowel disease.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A method for analyzing and identifying biomarkers related to oral cancer in inflammatory bowel disease, comprising the following steps:

[0008] Step 1: Obtain the summary data of the genome-wide association study of IBD;

[0009] Step 2: Use the Mendelian randomization method to analyze the causal relationship between IBD and oral cancer, and determine the IBD-related genes;

[0010] Step 3: Collect oral cancer tissue samples and perform single-cell RNA sequencing;

[0011] Step 4: Use dimensionality reduction technology to perform clustering analysis on the single-cell data and identify different cell types;

[0012] Step 5: Compare the gene expression differences between different cell types and identify the expression pattern of the NFKBIA gene in oral cancer;

[0013] Step 6: Use functional enrichment analysis to determine the biological pathways and functions involved in the NFKBIA gene;

[0014] Step 7: Combine the results of Mendelian randomization analysis and single-cell RNA sequencing to identify potential treatment targets for IBD-related oral cancer in the NFKBIA gene.

[0015] As a further optimization of the present invention, in Step 2, the Mendelian randomization method includes: selecting genetic variations related to IBD as instrumental variables, and using the instrumental variables to estimate the causal effect between IBD and oral cancer.

[0016] As a further optimization of the present invention, in Step 3, the specific steps of single-cell RNA sequencing are as follows:

[0017] Use single-cell RNA sequencing technology to sequence oral cancer tissue samples, preprocess and quality control the sequencing data, and use dimensionality reduction technology to perform dimensionality reduction processing on the single-cell data.

[0018] As a further optimization of the present invention, in Step 6, the specific steps of functional enrichment analysis are as follows:

[0019] Use gene ontology classification and KEGG pathway enrichment analysis to determine the biological pathways and functions involved in the NFKBIA gene.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0021] In the present invention, by performing single-cell RNA sequencing on oral cancer specimens, cell clusters and gene expression patterns are depicted. t-SNE and UMAP are used to display cell clusters and types, dot plots are used to define key patterns of gene expression, and pathways related to NFKBIA expression are analyzed by GSEA. The experimental results clarify the cell heterogeneity and gene expression dynamics during the progression of oral cancer disease, and reveal possible treatment targets related to IBD.

[0022] The above summary is only for the purpose of the specification and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the present invention will be readily apparent by reference to the drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a result diagram of the Mendelian randomization analysis of the present invention;

[0024] Figure 2 It is a result diagram of the leave-one-out sensitivity analysis of the Mendelian randomization of the present invention;

[0025] Figure 3This is the potential causal relationship diagram between IBD and increased risk of oral tumors in the present invention;

[0026] Figure 4 This is the single-cell data analysis result diagram of the present invention;

[0027] Figure 5 This is the expression pattern diagram of specific genes among cells in the present invention;

[0028] Figure 6 This is the NFKBIA expression and related gene set enrichment analysis result diagram of the present invention;

[0029] Figure 7 This is the weighted gene co-expression network analysis result diagram of oral cancer in the present invention;

[0030] Figure 8 This is the research background diagram under different conditions of the present invention;

[0031] Figure 9 This is the diagram of immune cell composition, model performance, and interactions between immune cells in the present invention;

[0032] Figure 10 This is the correlation analysis diagram between NRG2 expression and immune cell infiltration in the present invention. Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0034] A method for identifying IBD-related oral cancer biomarkers based on single-cell RNA sequencing and Mendelian randomization analysis, including:

[0035] Data source: RNA expression profiles and clinical information were downloaded from TCGA and GEO for a large amount of RNA-seq data and related clinicopathological features of lung adenocarcinoma patients. And the GWAS data was retrieved from the recently released comprehensive metabolomics dataset of UK Biobank.

[0036] GWAS data source for IBD: Summary data from a genome-wide association study (GWAS) of inflammatory bowel disease (IBD) analyzed 475,638 individuals of European descent. This study is part of the Integrative Epidemiology Unit (IEU) project, which can be accessed online through IDIEU-a-31 and finngen_R11_C3_ORALCAVITY_EXALLC.

[0037] Ethical statement: Informed consent for the use of GWAS summary statistics was obtained according to protocols approved by the respective institutional review boards.

[0038] Selection of instrumental variables: The significance threshold for instrumental variables associated with each immune trait was set at 1×10 -5 established. The clumping method in PLINK (version 1.90) was used. Single nucleotide polymorphisms (SNPs) with an r2 linkage disequilibrium threshold below 0.1 within a 500-kilobase window were removed, and the 1000 Genomes Project was used as the reference dataset for these calculations.

[0039] For the analysis of HCC, the significance level was adjusted to a more stringent threshold of 5×10 -8 , and the proportion of phenotypic variance explained and the F statistic were calculated for each instrumental variable to evaluate the strength of the instrumental variable and mitigate the impact of weak instrument bias. After excluding instrumental variables with an F value below 10, the remaining instrumental variables associated with HCC were used for subsequent analysis.

[0040] Single-cell clustering and annotation: The single-cell data of GSM were filtered, quality-controlled, normalized, dimension-reduced, and clustered using the Seurat package (version 4.4.0).

[0041] First, the single-cell sequencing data were screened using the CreateSeuratObject function, retaining genes with expression data in at least 3 cells and cells with more than 350 detected genes (min.cell = 3, min.features = 350). The mitochondrial gene score and ribosomal gene score were calculated based on mitochondrial genes and ribosomal genes, respectively, using the PercentageFeatureSet function. Cells with mitochondrial gene scores and ribosomal gene scores below 20% were retained. The SCTransform function was used to search for 3000 highly variable genes for normalization (variable.features.n = 3000), followed by principal component analysis and PCA dimension reduction. An elbow plot was drawn to identify the available dimensions of the data, and the principal components before the elbow were selected for analysis. The first 45 PCs were selected for subsequent analysis, and the batch effect was removed using the CCA method in Seurat.

[0042] Then, the FindNeighbors and FindClusters functions of the Seurat package were used to perform unsupervised clustering analysis on the cell data after batch removal. The optimal resolution was determined by the clustering tree function, with the optimal resolution parameter being 0.7, and a total of 26 clusters were obtained. UMAP was used for clustering visualization. The cells were manually annotated with reference to the marker genes provided by cellmarker. The proportion of the annotated cells was counted, and a bar chart was plotted using ggplot2 for visualization.

[0043] Cell communication analysis: The CellChat package software was used to predict the differences in cell communication between different groups with high and low Miller scores, with the threshold P-value less than 0.05. After obtaining different ligand-receptor pairs, the ligand-receptor pairs belonging to membrane proteins were selected for analysis to observe the expression of ligand-receptor pairs in different cells.

[0044] Establishment and validation of machine learning models:

[0045] 1. Dataset splitting: The oral cancer dataset was divided into two independent cohorts (TCGA and GEO) to evaluate the generalization ability of the model on different data sources. Ten algorithms and 101 algorithm combinations were used to train the model, including Random Survival Forest (RSF), Elastic Net (Enet), Lasso, Ridge regression, Stepwise Cox regression, Cox boost, Partial Least Squares Regression (plsRcox), Supervised Principal Components (SuperPC), Gradient Boosting Machine (GBM), and Survival Support Vector Machine (Survival-svm).

[0046] 2. Model selection: Based on the maximum average Harrell’s concordance index (C-index), ranging from [0,1], the closer the value is to 1, the better the model performance. The risk score was calculated individually for each sample using the formula "risk score = Σ(coefficient × expression value)".

[0047] 3. Visualization: The R software package "ggplot2" was used for data visualization, including a Sankey diagram to visualize the connection between the two risk clusters. The dataset was divided into overall, tcga-specific, and geo-specific subsets for Kaplan-Meier survival analysis. This included 10-fold cross-validation to plot the ROC curve and Decision Curve Analysis (DCA) to verify the robustness of the model.

[0048] 4. Model development: k-fold cross-validation was adopted to test the model accuracy, where k is the number of data splits.

[0049] Immune infiltration analysis: The roles of differentially expressed genes (DEGs) in the ARDS and non-ARDS groups were explored through functional enrichment and immune infiltration analysis. Gene Ontology (GO) classification and KEGG pathway enrichment analysis were used to determine the functions of these genes in biological processes and pathways, and an FDR value less than 0.05 was set as the statistical significance threshold.

[0050] The CIBERSORT and ESTIMATE algorithms were applied to identify differences between different cohorts. The CIBERSORT algorithm, based on linear support vector regression, estimates the composition and abundance of immune cells from gene expression data. The ESTIMATE algorithm uses single-sample gene set enrichment analysis (ssGSEA) to calculate stromal and immune scores and predict the levels of infiltrating stromal and immune cells in tumor tissues.

[0051] The R software package ggplot2 was used for data visualization, including Sankey diagrams, to visually display the distribution of immune cell subtypes in oral cancer samples.

[0052] Statistical analysis: The Student's t-test or Wilcoxon rank-sum test was used to evaluate continuous data, and the choice was based on the distribution characteristics of the data. Spearman's rank correlation analysis was used to evaluate the relationship between variables. Statistical significance was determined for results with p-values below the 0.05 threshold.

[0053] Experimental results:

[0054] For the results of Mendelian randomization (MR) analysis of the relationship between inflammatory bowel disease (IBD) and various types of tumors, as shown in Figure 1 Figures 1A - 1D, each figure represents a different type of tumor (such as uterine body, head and neck, esophagus, oral cavity). Listed along the y-axis are the genetic variants used as instrumental variables in the MR analysis. The MR effect sizes are shown on the x-axis, representing the estimated effect of each SNP on the risk of IBD for tumors. The red diamonds represent the overall MR effect estimated using the inverse-variance weighted and MR Egger methods. The position relative to the vertical line (no effect) indicates the direction and significance of the association. Whether there is a potential causal relationship between IBD and the risk of developing each specific type of tumor is shown in the figure, and the position and confidence interval of the red diamond provide a summary of the overall effect.

[0055] In 1E-1H, each represents a different tumor type. The x-axis represents the estimated causal effect of IBD on the tumor risk of each genetic variant. The y-axis represents the estimated precision (the reciprocal of the standard error). Each point corresponds to a single nucleotide polymorphism (SNP) used in the analysis. The vertical lines represent the overall effect estimates using two methods: inverse variance weighting (IVW) and MREgger. The symmetry of the points around the line indicates the presence or absence of bias. Asymmetry may indicate potential bias or pleiotropy affecting the results.

[0056] For the leave-one-out sensitivity analysis of Mendelian randomization, as shown in Figure 2 Figures 2A - 2D, each figure corresponds to a specific tumor type (e.g., uterine corpus, head and neck, esophagus, oral cavity). The SNPs listed on the y-axis are excluded one at a time to assess their impact on the overall effect estimate. The x-axis represents the estimated causal effect of IBD on tumor risk. The black dots and black lines represent the effect estimates and confidence intervals after excluding each SNP. The red diamond represents the overall effect estimate including all SNPs.

[0057] In 2E - EH, the consistency of the effect estimates when excluding a single SNP indicates robustness. When certain SNPs are removed, significant changes may indicate their impact on the results. The slope of each line represents the estimated causal effect. The consistency between methods indicates reliable causal inference. Differences may indicate potential bias or pleiotropy. This helps to evaluate the reliability and consistency of the causal relationship between IBD and tumor risk.

[0058] For the potential causal relationship between IBD and an increased risk of oral tumors, as shown in Figure 3 Figures: Uterine corpus: Most methods show a reduced risk (OR < 1). Head and neck: Increased risk (OR > 1), with significant results for all methods except MR Egger. Esophagus: Increased risk (OR > 1), with significant results for all methods except MR Egger. Oral cavity: Increased risk (OR > 1), with significant results for all methods except MR Egger.

[0059] In summary, the experimental results indicate a potential causal relationship between IBD and an increased risk of head and neck, esophagus, and oral tumors.

[0060] For single-cell data analysis, highlighting cell clusters and gene expression patterns, as shown in Figure 4 Figure 4A, the t-SNE plot (clusters): shows the cell clusters identified by Seurat, with different colors representing different clusters. Figure 4BZHONG, t-SNE plot (cell type): shows the distribution of cell types, such as fibroblasts, T cells, endothelial cells, myeloid cells, mast cells, and malignant cells. Figure 4 In C, UMAP plot (cluster): similar to panel A, but uses UMAP dimensionality reduction to show cell clusters. UMAP plot (cell type): similar to 4B, shows the distribution of cell types in the UMAP space. Dot Plot: illustrates the gene expression of different cell types. Genes such as IL7R, LYZ, and ACTB are shown in the figure. The size of the dot represents the percentage of cells expressing the gene, and the color represents the average expression level.

[0061] For the expression of a specific gene among cells, as attached Figure 5 shown, in ATG5, the region showing ATG5 expression is shown, with higher expression in some clusters. In Figure 5B, RB1CC1: highlights the expression of RB1CC1, with different levels in different cell groups. In Figure 5C, IL24: shows IL24 expression, with some clusters showing higher levels. Figure 5 In dTMEM59, it represents the expression location of TMEM59, with significant expression in specific regions. In 5E ATG12, shows the expression pattern of ATG12, with higher expression levels in some clusters. In 5F NFKBIA, shows the expression of NFKBIA, which is obvious in one cluster. The color gradient from gray to red represents increasing expression levels, with red indicating higher expression.

[0062] For NFKBIA expression and related gene set enrichment analysis, as attached Figure 6 shown, in 6A, t-SNE plot: shows cells based on NFKBIA expression, with red indicating high expression and blue indicating low expression, highlighting the expression distribution of NFKBIA in the cell population. In 6B, ridge plot: shows the enriched pathways related to NFKBIA expression, such as chemokine-mediated signaling and cytokine activity. The figure shows the significant pathways after p-value adjustment. In 6C, GSEA result dot plot: shows the results of gene set enrichment analysis (GSEA). Each dot represents a gene set, the size represents the number of genes, and the color represents the significance level (p.adjust). In 6D, enrichment plot: shows the running enrichment score of specific gene sets, and the figure shows the distribution of these gene sets in the ranked dataset.

[0063] For oral cancer weighted gene co-expression network analysis (WGCNA), as attached Figure 7As shown, in 7A, the scale-free topology fitting index is evaluated according to the soft thresholding power, which is generally higher than 0.8. In 7B, the average connectivity is shown as a function of the soft thresholding power, which helps to select the optimal power that balances network connectivity and scale independence. In 7C, the dendrogram represents the hierarchical clustering of genes, and the branches represent modules. Each color corresponds to a different module identified in the data. In 7D, the module eigengenes are aggregated to summarize the expression profiles of the modules. Modules with similar patterns are grouped together. In 7E, the relationship between module membership (the degree of fit between genes and modules) and gene significance (the correlation with oral cancer characteristics) is illustrated. A strong correlation indicates the relevance of the module to the trait. In 7F, the correlation between module eigengenes and traits is shown. Each cell represents the correlation coefficient and p-value, indicating the strength of the association between the module and the trait. In 7G, the bar chart shows the gene significance between different modules. The higher the bar, the stronger the correlation with the oral cancer trait. In 7H, the relationship between module membership and gene significance of a specific module is depicted. A high correlation indicates an important module related to oral cancer.

[0064] As shown in the appendix Figure 8 In 8A, this table ranks various models according to the area under the curve (AUC) value. A higher AUC value indicates that the model performs better in distinguishing between the control group and the treatment group. In 8B, this figure shows the differentially expressed genes. The x-axis shows the log fold change, and the y-axis shows the statistical significance (-log10 p-value). The genes of interest are highlighted, indicating significant changes in expression. In 8C, (GSE31056) shows the performance of the model on the test dataset. The matrix represents true positives, true negatives, false positives, and false negatives to evaluate the prediction accuracy. In 8D, (Train): shows the performance of the model on the training dataset, showing similar metrics to evaluate how well the model learns from the data. In 8E, the expression levels of selected genes in different samples are visualized. The genes and samples are clustered to show the expression patterns, and the colors represent different expression levels. In 8F, the receiver operating characteristic (ROC) curve of the GSE31056 dataset shows the trade-off between sensitivity and specificity. The AUC value is 0.784, indicating a good discrimination level and providing a 95% confidence interval for reliability.

[0065] For immune cell composition, model performance, and interactions between immune cells, as shown in the appendix Figure 9 In 9A, the relative abundances of various immune cell types in the control and treated samples are shown. Each color represents a different cell type, highlighting the differences in immune cell distribution between groups. Figure 9In panel B, the ROC curves of different models or datasets show their performance in distinguishing between the control group and the treatment group. The higher the curve, the better the sensitivity and specificity, and the AUC values are shown for comparison. In panel 9C, the fractions of different immune cell types between the control group and the treatment group are compared. Significant differences are marked with asterisks, indicating which cell types are significantly different between the groups. In panel 9D, the correlations between different immune cell types are shown. Positive correlations are shown in red and negative correlations in blue, which helps to identify the relationships and interactions between cell types. In panel 9E, scatter plots of pairwise comparisons between immune cell types are shown, and the distributions and trend lines help to identify patterns and correlations.

[0066] For the correlation analysis of NRG2 expression and immune cell infiltration, as shown Figure 10 in Figure 10A, the correlations between NRG2 expression and various immune cell types are shown. The color coding indicates the significance of the correlations: green indicates a positive correlation and red indicates a negative correlation. In Figure 10B, NRG2 expression is positively correlated with memory B cells (R = 0.26, p < 2e-16). In Figure 10C, it is negatively correlated with activated dendritic cells (R = -0.17, p = 0.0067). In Figure 10D, it is negatively correlated with neutrophils (R = -0.15, p = 0.00067). In Figure 10E, it is negatively correlated with resting NK cells (R = -0.12, p = 0.04). In Figure 10F, it is negatively correlated with CD4 memory T cells (R = -0.13, p = 0.027). In Figure 10G, it is positively correlated with regulatory T cells (Tregs) (R = 0.13, p = 0.028). In Figure 10H, the correlation matrix shows the interactions between immune cell types, and the colors and lines indicate the strength and direction of the correlations, which helps to identify the interactions between different immune cells.

[0067] Parts not involved in the present invention are the same as or can be implemented using the prior art. Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for analyzing and identifying biomarkers of oral cancer associated with inflammatory bowel disease, characterized in that: The following steps are involved: Step 1: Obtain summary data of IBD genome-wide association studies; Step 2: Use Mendelian randomization to analyze the causal relationship between IBD and oral cancer and identify IBD-related genes; Step 3: Collect oral cancer tissue samples and perform single-cell RNA sequencing; Step 4: Use dimensionality reduction technology to cluster the single-cell data and identify different cell types; Step 5: Compare the gene expression differences between different cell types and identify the expression pattern of NFKBIA gene in oral cancer; Step 6: Use functional enrichment analysis to determine the biological pathways and functions involved in the NFKBIA gene; Step 7: Combine Mendelian randomization analysis and single-cell RNA sequencing results to identify the NFKBIA gene as a potential therapeutic target for IBD-related oral cancer.

2. The method for analyzing and identifying biomarkers of oral cancer associated with inflammatory bowel disease according to claim 1, characterized in that: In step 2, the Mendelian randomization method includes: selecting genetic variants associated with IBD as instrumental variables and using instrumental variables to estimate the causal effect between IBD and oral cancer.

3. The method for analyzing and identifying inflammatory bowel disease-related oral cancer biomarkers according to claim 1, characterized in that: In step 3, the specific steps of single-cell RNA sequencing are: Oral cancer tissue samples were sequenced using single-cell RNA sequencing technology, the sequencing data were preprocessed and quality controlled, and the single-cell data were reduced in dimensionality using dimensionality reduction technology.

4. The method for analyzing and identifying biomarkers of oral cancer associated with inflammatory bowel disease according to claim 1, characterized in that: In step six, the specific steps of functional enrichment analysis are: Gene ontology classification and KEGG pathway enrichment analysis were used to determine the biological pathways and functions involved in the NFKBIA gene.