Ulcerative colitis prediction model, construction method and prediction kit
By screening nine acetylation-related core genes in UC patients, a predictive model for ulcerative colitis was constructed and a predictive kit was developed, which solved the problem of insufficient sensitivity and specificity of UC diagnostic methods and achieved accurate diagnosis of UC.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WEIFANG MEDICAL UNIV
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-02
AI Technical Summary
Existing diagnostic methods for UC lack sensitivity and specificity, failing to effectively screen out acetylation-related core genes with clinical application value. This results in traditional diagnosis relying on symptom observation and endoscopic examination, lacking precise diagnostic tools.
Nine acetylation-related core genes were screened from the gene expression data of UC patients using machine learning algorithms. A predictive model for ulcerative colitis was constructed, and a risk scoring model was established by combining logistic regression analysis. A predictive kit was also developed for diagnosis.
This method provides an improved diagnostic accuracy for UC by using objective and quantitative criteria to address the shortcomings of traditional diagnostic methods. It is suitable for rapid diagnosis of clinical samples and improves the sensitivity and specificity of diagnosis.
Smart Images

Figure CN122135993A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ulcerative colitis prediction technology, specifically to an ulcerative colitis prediction model, its construction method, and a prediction kit. Background Technology
[0002] Ulcerative colitis (UC) is a chronic inflammatory bowel disease primarily affecting the sigmoid colon and rectum. Its global incidence and prevalence are steadily increasing, particularly among young people in China, where the incidence rate is approximately 11.6 cases per 100,000 people, with the highest proportion in the 20-40 age group. The disease has a long course and is difficult to cure. Long-term chronic inflammation not only leads to permanent intestinal damage but also significantly increases the risk of hospitalization, surgery, and colorectal cancer, severely reducing patients' quality of life. It has become a pressing public health issue, necessitating the development of precise diagnostic tools.
[0003] The pathogenic factors of ulcerative colitis (UC) have been confirmed to be related to genetic susceptibility and environmental factors, but the exact pathogenesis has not been fully elucidated. In recent years, the role of epigenetic regulation in disease development has gradually attracted attention. Among them, histone modifications (including acetylation, methylation, ubiquitination, etc.) play a central role in immune response regulation and the pathogenesis of inflammatory diseases as a key mechanism of gene transcriptional regulation. Numerous studies have shown that abnormal acetylation modifications are closely related to the pathological process of UC: for example, excessive activation of the long non-coding RNA DANCR by H3K27 acetylation has an oncogenic effect, activation of histone deacetylase 1 / 2 (HDAC1 / 2) can improve intestinal mucosal integrity, and engineered probiotics can also reduce intestinal inflammation by promoting histone acetylation.
[0004] However, current research on acetylation related to UC still has significant limitations: First, the global pattern of acetylation modification in UC patients and the specific mechanisms of action of related genes have not been clarified, and the impact of different acetylation modifications on disease phenotypes lacks systematic analysis; Second, existing studies have failed to effectively screen out core acetylation-related genes with clinical application value as biomarkers, resulting in clinical diagnosis still relying on traditional symptom observation and endoscopic examination, with limited sensitivity and specificity.
[0005] With the rapid development of artificial intelligence technology, machine learning has become an important tool for analyzing large-scale omics data and identifying disease biomarkers. Currently, algorithms such as random forests, elastic networks, and support vector machines have been used in UC-related research for scenarios such as microbiome data analysis and early diagnostic biomarker screening. However, single algorithms are prone to selection bias and insufficient model generalization ability. Therefore, integrating multiple machine learning techniques to build consensus models and combining multi-dimensional verification to improve the reliability of results has become an effective way to discover acetylation-related core biomarkers and build accurate diagnostic models for UC, providing technical support for solving the current accuracy problem in UC diagnosis. Summary of the Invention
[0006] This invention overcomes the shortcomings of existing technologies by providing a predictive model for ulcerative colitis, a method for constructing such a model, and a predictive kit to address the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a predictive model for ulcerative colitis, wherein the model is: risk score for each sample = 68.798 – 11.031 * HMGCS2 relative expression level – 46.586 * GCNT2 relative expression level + 20.391 * B4GALNT2 relative expression level + 5.719 * NAT8B relative expression level + 26.562 * LPCAT1 relative expression level + 17.042 * PPARGC1A relative expression level – 10.500 * ACSS2 relative expression level + 3.110 * PDK2 relative expression level – 77.421 * HADHA relative expression level.
[0008] Furthermore, a risk score less than zero is defined as non-ulcerative colitis, while a risk score greater than zero is defined as ulcerative colitis.
[0009] A method for constructing a predictive model for ulcerative colitis. (1) Gene expression data of ulcerative colitis patients and healthy controls were obtained from the database. Machine learning algorithms combined with LOOCV (leave-one-out cross-validation) were used to screen out acetylation-related core genes and perform data preprocessing. (2) Analyze the core genes related to acetylation that were screened; (3) A risk scoring model was constructed for acetylation-related core genes, and its ability to distinguish and diagnose ulcerative colitis was validated internally using the dataset; (4) Use clinical samples to externally validate the risk scoring model’s ability to identify and diagnose ulcerative colitis.
[0010] Furthermore, the core acetylation-related genes identified were: HMGCS2, GCNT2, B4GALNT2, NAT8B, LPCAT1, PPARGC1A, ACSS2, PDK2, and HADHA.
[0011] Further, the specific steps of (1) are as follows: obtain the dataset from the GEO gene expression comprehensive database, obtain the gene expression data of colonic mucosal tissue of patients with ulcerative colitis and healthy controls, use the R language limma package to screen differentially expressed genes in the data; screen acetylation-related genes from MSigDB and GeneCards databases, obtain candidate genes by taking the intersection with differentially expressed genes, integrate machine learning algorithms, construct and optimize the prediction model under the leave-one-out cross-validation framework, and finally determine the core acetylation-related genes after feature screening.
[0012] Further, the specific steps of (2) are as follows: analyze the inter-gene correlation and chromosomal localization of core acetylation genes; construct a miRNA-transcription factor-core ARGs regulatory network using the NetworkAnalyst platform, combined with miRTarBase v9.0 and the TRRUST database, and screen potential drug targets; draw calibration curves and DCA using the R language rmda package, calculate the area under the ROC curve using the pROC package, and evaluate the diagnostic ability of 9 acetylation-related core genes for ulcerative colitis; perform GSEA gene set enrichment analysis based on acetylation-related core genes to explore associated signaling pathways, use LM22 feature analysis and cibersort algorithm to detect immune cell abundance, and analyze the difference in immune cell infiltration between high and low expression groups of core genes; use the R language ConsensusClusterPlus unsupervised clustering package to divide ulcerative colitis sample subgroups based on the expression profile of acetylation-related core genes, verify subgroup differences through PCA, and perform cross-reference analysis to verify subtype stability; use Mendelian randomization method, with SNPs as instrumental variables, and perform IVW Using the MR-Egger method, we explored the causal relationship between acetylation-related core genes and susceptibility to ulcerative colitis.
[0013] Further, the specific steps of (3) are as follows: based on logistic regression analysis of the relative expression levels of genes HMGCS2, GCNT2, B4GALNT2, NAT8B, LPCAT1, PPARGC1A, ACSS2, PDK2 and HADHA, a risk scoring model is established, samples are selected from the dataset, and internal validation is performed by combining logistic regression with the relative expression levels of core ARGs. The diagnostic efficacy of the model is evaluated by confusion matrix, sensitivity, specificity, positive predictive value, negative predictive value and ROC curve.
[0014] Furthermore, the specific steps of (4) are as follows: take clinical colonic mucosal tissue samples from normal people and patients with ulcerative colitis, detect the relative expression level of genes using qRT-PCR, and clinically validate the model by combining logistic regression with the relative expression level of core genes. The validation indicators include confusion matrix, sensitivity, specificity, positive predictive value, negative predictive value, AUC value and Youden index, to further confirm the clinical applicability and diagnostic efficacy of the model.
[0015] A kit for predicting ulcerative colitis, comprising: The reagents for detecting the relative expression levels of the acetylation-related core genes, including HMGCS2, GCNT2, B4GALNT2, NAT8B, LPCAT1, PPARGC1A, ACSS2, PDK2, and HADHA, as well as the aforementioned ulcerative colitis prediction model.
[0016] Furthermore, The reagents for detecting the relative expression levels of the acetylation-related core genes include primers, internal controls, and qRT-PCR detection reagents.
[0017] Compared with the prior art, the beneficial effects of the present invention are: This application provides a novel method for diagnosing ulcerative colitis (UC). Through machine learning, nine acetylation-related core genes are screened from the genes of UC patients, and a predictive model is constructed to improve the predictability of UC diagnosis in clinical settings. This model is applicable to the rapid diagnosis of clinical samples, providing objective and quantitative criteria for UC diagnosis, solving the problems of insufficient sensitivity and specificity of traditional diagnostic methods, and providing precise reference for clinical diagnosis and treatment. Attached Figure Description
[0018] Figure 1 This is the result of differential gene identification in this invention; Figure 2 This is the result of the functional enrichment analysis of the present invention; Figure 3 These are the screening results for acetylation-related genes in this invention; Figure 4 The results of the intergene correlation and chromosomal localization analysis of the core acetylation gene of this invention are shown. Figure 5 The results of the assessment of the diagnostic capability of the core acetylation gene of this invention for ulcerative colitis; Figure 6 This is the result of verification of the functional pathway and immune association of the core acetylation gene in this invention; Figure 7The identification results of nine acetylation-related subgroups in the UC sample of this invention; Figure 8 The results of cluster analysis of UC samples of 22 acetylation-related genes screened from UC samples in this invention; Figure 9 This is the result of the difference in immune cell abundance between the hyperacetylation and hypoacetylation gene expression groups of the core acetylation gene of this invention; Figure 10 This invention presents the association analysis results between small molecule drugs based on standardized connectivity scores and nine acetylation-related gene targets. Figure 11 This is the result of the Mendelian randomization study in this invention; Figure 12 This is a Venn diagram showing the correspondence between differentially expressed genes and 364 acetylated genes in subgroups A and B of this invention. Figure 13 The results are for the validation of the regulatory network and target of the core acetylation gene in this invention. Figure 14 The confusion matrix for ulcerative colitis (UC) diagnosis is used for internal model validation of the GSE107499 and GSE87466 datasets in this invention. Figure 15 ROC curves for ulcerative colitis (UC) diagnosis used for internal model validation of the GSE107499 and GSE87466 datasets in this invention; Figure 16 Scatter plot of risk score distribution for UC patients and healthy controls for internal model validation of the GSE107499 and GSE87466 datasets of this invention; Figure 17 A visual comparison chart of risk scores between UC patients and healthy controls for internal model validation using the GSE107499 and GSE87466 datasets of this invention; Figure 18 This is the diagnostic confusion matrix for ulcerative colitis (UC) in the internal model validation of the GSE92415 dataset of this invention. Figure 19 ROC curves for ulcerative colitis (UC) diagnosis used for internal model validation on the GSE92415 dataset of this invention; Figure 20 Scatter plot of risk score distribution for UC patients and healthy controls for internal model validation of the GSE92415 dataset of this invention; Figure 21 A visual comparison chart of risk scores between UC patients and healthy controls for internal model validation using the GSE92415 dataset of this invention; Figure 22This is a diagnostic confusion matrix for ulcerative colitis (UC) that is externally validated for this invention. Figure 23 ROC curve for external validation of the ulcerative colitis (UC) diagnosis of this invention; Figure 24 Scatter plot of risk score distribution in UC patients and healthy controls for external validation of this invention; Figure 25 This is a visual comparison chart of risk scores between UC patients and healthy controls used for external validation of this invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1
[0020] A predictive model for ulcerative colitis, where the risk score for each sample is calculated as follows: 68.798 – 11.031 * HMGCS2 relative expression level – 46.586 * GCNT2 relative expression level + 20.391 * B4GALNT2 relative expression level + 5.719 * NAT8B relative expression level + 26.562 * LPCAT1 relative expression level + 17.042 * PPARGC1A relative expression level – 10.500 * ACSS2 relative expression level + 3.110 * PDK2 relative expression level – 77.421 * HADHA relative expression level.
[0021] The relative expression levels were standardized and normalized, and GAPDH was used as an internal reference for the determination of relative expression levels.
[0022] A risk score less than zero is defined as non-ulcerative colitis, while a risk score greater than zero is defined as ulcerative colitis.
[0023] A method for constructing a predictive model for ulcerative colitis: 1. Gene expression data from ulcerative colitis patients and healthy controls were obtained from the database. Machine learning algorithms combined with LOOCV were used to screen for acetylation-related core genes. (1) Data source: Two public datasets, GSE107499 and GSE92415, were obtained from the GEO Gene Expression Comprehensive Database. The samples were colonic mucosal tissues from patients with ulcerative colitis (UC) and healthy controls.
[0024] (2) Data preprocessing: The limma package in R language was used to perform DEGs screening on normal and UC samples in GSE107499 and GSE92415, and the threshold was set to |logFC|>0.585 and p value<0.05.
[0025] Differential gene identification results: such as Figure 1 As shown, A represents the principal component analysis results of the expression profiles of healthy controls (HC) and ulcerative colitis (UC) samples in the GSE92415 dataset; B represents the principal component analysis results of the expression profiles of healthy controls (HC) and ulcerative colitis (UC) samples in the GSE107499 dataset; C represents the comparative analysis results of the volcano plot of differentially expressed genes between HC and UC patients in the GSE92415 dataset; D represents the comparative analysis results of the volcano plot of differentially expressed genes between HC and UC patients in the GSE107499 dataset; and E represents the Venn diagram of common differentially expressed genes in the GSE107499 and GSE92415 datasets.
[0026] GSE92415 and GSE107499 identified 2452 and 2165 DEGs, respectively, with a total of 1446 DEGs in the two datasets. PCA showed significant differences in clustering patterns between UC patients and healthy controls, effectively distinguishing the two groups (Figures 1A and 1B). Volcano plots showed the distribution of DEGs in the two datasets (Figures 1C and 1D), and Venn diagrams presented the total number of DEGs. Figure 1 E).
[0027] (3) Preliminary analysis of differential gene function: GO and KEGG enrichment analysis was performed using the R language ClusterProfiler package / David online platform, and a DEGs co-expression network was constructed using GeneMANIA.
[0028] Functional enrichment analysis results: like Figure 2 As shown, differentially expressed gene function enrichment analysis is performed on gene ontology biological processes (A), gene ontology cellular components (B), gene ontology molecular functions (C), and the KEGG pathway (D).
[0029] Gene-ontology biological processes are enriched in inflammatory responses, immune responses, neutrophil chemotaxis, lipopolysaccharide-induced cellular responses, chemotaxis, signal transduction, antimicrobial peptide-mediated humoral immune responses, chemokine signaling pathways, killing of other organisms' cells, and innate immune responses (Figure 2A). Gene-ontology cellular components involve the extracellular space, plasma membrane, extracellular exosomes, outer plasma membrane, plasma membrane integrated components, membrane integrated components, extracellular matrix, cell surface, and apical membrane (Figure 2B). Gene-ontology molecular functions are mainly enriched in chemokine activity, calcium ion binding, transmembrane signal receptor activity, chemokine receptor (CXCR) binding, extracellular matrix structural components, and receptor binding (Figure 2C). The KEGG pathway is enriched in rheumatoid arthritis, cytokine-cytokine receptor interactions, Staphylococcus aureus infection, peroxisome proliferator-activated receptor (PPAR) signaling, chemokine signaling, interleukin (IL)-17 signaling, tumor necrosis factor (TNF) signaling, and the cell adhesion molecule KEGG pathway (Figure 2D).
[0030] (4) Initial screening of acetylation-related genes: Using “acetylation” as the search term, we screened acetylation-related genes from the MSigDB molecular feature database and the GeneCards human gene database, set a minimum relevance score of 5.0, and took the intersection with the common DEGs.
[0031] Initial screening results of acetylation-related genes: 364 unique acetylation-related genes were obtained from the MSigDB and GeneCards databases. After intersecting with 1446 common DEGs, 29 candidate acetylation-related genes were obtained. Figure 3 A).
[0032] (5) Final identification of acetylation-related core genes: A machine learning-based ensemble approach was used to analyze 29 genes to construct consensus features related to acetylation. Leave-one-out cross-validation (LOOCV) framework was used with 10 machine learning algorithms, including Lasso, Ridge, stepwise Cox, CoxBoost, Random Survival Forest (RSF), Elastic Net (Enet), Cox Partial Least Squares Regression (plsRcox), Supervised Principal Component Analysis (SuperPC), Generalized Boosting Regression Model (GBM), and Survival Support Vector Machine (survival-SVM). Ten-fold cross-validation was performed on each algorithm, and the model was selected based on the highest receiver operating characteristic (AUC) obtained from the two datasets. 101 prediction models were constructed, and the consistency index (C-index) of each model was calculated on all datasets (GSE107499, GSE92415, and GSE87466 public datasets obtained from the GEO Gene Expression Comprehensive Database; learning was performed on GSE107499 and GSE87466 public datasets, and validation was performed on GSE92415).
[0033] Final gene screening results: Random Forest (RSF) with Enet [α=0.3], RSF with Enet [α=0.9], RSF with Enet [α=0.1], RSF with Enet [α=0.2], RSF with Enet [α=0.5], RSF with Enet [α=0.6], RSF with Enet [α=0.7], RSF with Enet [α=0.8], and Random Forest (RSF) with Lasso regression model performed best on each dataset, all achieving a C-index of 0.99 (Figure 3B). Nine common acetylation-related genes appeared as hub genes in all models: HMGCS2, GCNT2, B4GALNT2, NAT8B, PPARGC1A, ACSS2, PDK2, LPCAT1, and HADHA. These nine differentially expressed acetylation-related core genes are the finally identified acetylation-related core genes.
[0034] 2. Analyze and validate the screened acetylation-related core genes. (1) The correlation between genes and chromosomal location of nine acetylation-related core genes were analyzed.
[0035] The results are as follows Figure 4 As shown, A represents the chromosomal localization of nine differentially expressed acetylation-related genes; box plots BD show the expression of these nine acetylation-related genes between the UC group and the control group in the GSE107499, GSE92415, and GSE87466 datasets, respectively. E represents the associations among the nine acetylation-related genes.
[0036] The chromosomal locations of nine acetylation-related core genes were identified (Figure 4A). Compared with the healthy control group, the expression of eight genes, including HMGCS2 and GCNT2, was significantly downregulated in UC patients, while only LPCAT1 was significantly upregulated (Figures 4B-D). Correlation analysis showed that the nine acetylation-related core genes were closely associated, and only LPCAT1 was positively correlated with other genes (Figure 4E).
[0037] (2) Using the NetworkAnalyst platform, combined with miRTarBase v9.0 and the TRRUST database, a miRNA-transcription factor-core ARGs regulatory network was constructed to screen potential drug targets.
[0038] Figure 13 A shows the regulatory network and target validation results of the core acetylation gene; B shows the gene-transcription factor network construction diagram; C shows the gene-microRNA network construction diagram.
[0039] High-abundance transcription factors associated with HADHA, PDK2, and ACSS2 were identified (Figure 13A), as well as high-abundance miRNAs associated with HADHA, GCNT2, and LPCAT1. Figure 13 B); Identify ACSS2 and HMGCS2 as potential targets for small molecule drugs ( Figure 13 C).
[0040] (3) The calibration curve and DCA were plotted using the rmda package in R language, and the area under the ROC curve (AUC) was calculated using the pROC package to evaluate the diagnostic ability of the nine acetylation-related core genes for ulcerative colitis.
[0041] The results are as follows Figure 5 As shown, A represents the construction of a nomogram model based on nine acetylation-related genes. B shows the calibration plot evaluating the robustness of the nomogram predictions. C represents the decision curve analysis of the established nomogram. DF shows the receiver operating characteristic (ROC) curves for using acetylation-related features in the diagnosis of ulcerative colitis (UC).
[0042] The nomogram model shows the distribution of gene point values (Figure 5A); the calibration curve shows that the predicted results have no significant deviation from the actual observed values, with a mean absolute error of 0.012 (Figure 5B); the DCA curve shows that the model can provide significant net benefit over a wide threshold probability range of 0.12–1.00 (Figure 5C); ROC curve analysis shows that the AUC of the multi-marker combination of the nine acetylation-related core genes reaches 0.960, and the AUC of each individual acetylation-related core gene is at a high level (Figures 5D-F).
[0043] (4) Validation of associated signaling pathways: Based on the core ARGs, GSEA gene set enrichment analysis was performed to explore associated signaling pathways. The abundance of 22 immune cells was detected by LM22 feature analysis and cibersort algorithm, and the differences in immune cell infiltration between high and low expression groups of acetylation-related core genes were analyzed.
[0044] Figure 6 A represents the GSEA samples with high and low expression of acetylation-related core genes, B represents the proportion of 22 immune cell types, and C represents the association between 9 core ARGs and 22 immune cell types.
[0045] Figure 9 This represents the difference in immune cell abundance between the high-acetylation and low-acetylation gene expression groups.
[0046] GSEA analysis showed that core acetylation-related genes were significantly associated with UC-related immune pathways such as cytokine-cytokine receptor interactions and the TNF signaling pathway (Figure 6A).
[0047] In UC samples, the infiltration rates of activated CD4+ T cells and γδ T cells were significantly increased, while the infiltration rates of CD8+ T cells and regulatory T cells were significantly decreased (Figure 6B). Within UC patient samples, after dividing the samples into high- and low-expression subgroups based on the expression levels of acetylation-related core genes, except for LPCAT1, the high-expression groups of the other eight acetylation-related core genes showed more significant infiltration of CD8+ T cells, regulatory T cells, and M2 macrophages (Figure 9); acetylation-related core genes showed significant correlations with 22 types of immune cells (Figure 6C).
[0048] (5) Different acetylation modification patterns identified by differentially expressed genes related to acetylation: Using the ConsensusClusterPlus unsupervised clustering package in R language, UC sample subgroups were divided based on the expression profiles of core genes related to acetylation. The subgroup differences were verified by PCA, and cross-reference analysis was performed to verify the stability of the subtypes.
[0049] Figure 7 To identify nine acetylation-related subsets in UC samples, A shows subset analysis based on the nine acetylation-related genes. B shows principal component analysis. C shows box plots illustrating the expression levels of the nine acetylation-related genes in subsets A and B. D shows gene ontology analysis of the nine differentially expressed acetylation-related genes in the UC subtype. E shows KEGG-based pathway enrichment analysis. F shows box plots illustrating the differences in immune infiltration between the two subsets.
[0050] Initial clustering divided the UC samples into two subgroups (Group A, n=75; Group B, n=87). PCA confirmed significant differences in gene expression between the two groups (Figure 7B). Nine acetylation-related core genes showed different expression characteristics between the subgroups (Figures 7A and 7C). Box plots showed that LPCAT1 was highly expressed in Group A, while HMGCS2, GCNT2, B4GALNT2, NAT8B, PPARGC1A, ACSS2, PDK2, and HADHA were enhanced in Group B. Subgroup GO / KEGG pathway analysis (Figures 7D and 7E) and differences in immune cell infiltration (Figure 7F) both showed unique characteristics.
[0051] Figure 12 A Venn diagram showing the correspondence between differentially expressed genes and 364 acetylated genes in subgroups A and B; Figure 8 for Figure 12 Cluster analysis of UC samples, focusing on 22 acetylation-related genes screened from the UC samples, is presented. A represents subgroup analysis based on these 22 acetylation-related genes. B represents principal component analysis. C is a box plot showing the expression levels of nine core acetylation-related genes in subgroups A and B. D represents gene ontology analysis of nine differentially expressed acetylation-related genes in UC subtypes. E represents KEGG-based pathway enrichment analysis. F is a box plot showing the differences in immune infiltration between the two subgroups.
[0052] Cross-referenced analysis identified 22 ARGs (Figure 12). Further clustering divided UC patients into two subgroups (Group A n=97 and Group B n=65). PCA verified the differences in gene expression (Figure 8B). The expression of 9 core genes differed significantly between the subgroups (Figures 8A and 8C). Furthermore, the GO / KEGG pathway analysis (Figures 8D and 8E) and immune cell infiltration profile analysis (Figure 8F) were highly consistent with the initial acetylation pattern.
[0053] (6) Causal relationship verification: Using Mendelian randomization method and SNPs as instrumental variables, the causal relationship between core ARGs and UC susceptibility was explored by combining IVW with MR-Egger method.
[0054] Causal relationship verification results: like Figure 11The results of the Mendelian randomization study are shown below. A is a scatter plot showing the causal effect of NAT8B on UC risk. B is a scatter plot showing the causal effect of PPARGC1A on UC risk. C is a Forest plot showing the causal effect of each SNP in NAT8B on UC risk. D is a Forest plot showing the causal effect of each SNP in PPARGC1A on UC risk. E is a leave-one-out plot with one SNP retained, used to visualize the causal effect of NAT8B on UC risk. F is a leave-one-out plot with one SNP retained, used to visualize the causal effect of PPARGC1A on UC risk.
[0055] Scatter plots (Figures 11A and 11B), Forest plots (Figures 11C and 11D), and leave-one-out verification plots (Figures 11A and 11B) Figure 11 Both E and 11F support the robustness of the results.
[0056] (7) Validation of potential therapeutic drugs: Screening for potential therapeutic small molecule compounds that target acetylation-related core genes through CMap connectivity mapping analysis.
[0057] Figure 10 This is the result of an association analysis between small molecule drugs and nine acetylation-related core gene targets based on standardized connectivity scores.
[0058] Five small molecule compounds with potential therapeutic potential for UC were screened using CMap analysis. Reciquimod and chloromethathiazole were significantly associated with seven acetylation-related core genes (Figure 10).
[0059] 3. Construct a risk scoring model for the core acetylation gene and internally validate its ability to discriminate and diagnose ulcerative colitis using a dataset: (1) Risk scoring model construction: Based on logistic regression analysis of the relative expression levels of genes HMGCS2, GCNT2, B4GALNT2, NAT8B, LPCAT1, PPARGC1A, ACSS2, PDK2 and HADHA, a risk scoring model for identifying patients with ulcerative colitis was established.
[0060] The specific process of logistic regression analysis is as follows: First, the binary sample group (Control / UC) is defined as the dependent variable, and the relative expression levels of nine acetylation-related core genes are defined as the independent variables. The binomial distribution is selected using the `glm` function, and the `logit` link function is specified to fit the binary logistic regression model. The intercept term and the regression coefficients corresponding to each gene are solved using the maximum likelihood estimation method. Then, the statistical test results of the coefficients (such as Z-values and P-values) are output using the `summary` function to determine the significance and direction of each gene's influence on UC risk (positive coefficients represent risk genes, and negative coefficients represent protective genes). Finally, based on the linear formula "intercept term + relative expression level of each gene × corresponding coefficient", the parameters of the regression model are transformed into quantifiable sample risk scores. A risk score greater than zero indicates ulcerative colitis, and a risk score less than zero indicates non-ulcerative colitis. The final risk score for each sample was calculated as follows: 68.798 – 11.031 * HMGCS2 relative expression level – 46.586 * GCNT2 relative expression level + 20.391 * B4GALNT2 relative expression level + 5.719 * NAT8B relative expression level + 26.562 * LPCAT1 relative expression level + 17.042 * PPARGC1A relative expression level – 10.500 * ACSS2 relative expression level + 3.110 * PDK2 relative expression level – 77.421 * HADHA relative expression level. A risk score less than zero was defined as non-ulcerative colitis, and a risk score greater than zero was defined as ulcerative colitis.
[0061] (2) Use the dataset to perform internal validation of the model. Validation method: Samples were selected from three datasets, and logistic regression was used in combination with the relative expression levels of core ARGs for two internal validations. The diagnostic efficacy of the model was evaluated by confusion matrix, sensitivity, specificity, positive predictive value, negative predictive value and ROC curve.
[0062] a. Internal validation samples were selected from datasets GSE107499 and GSE87466. 119 samples were randomly selected from the two datasets, including 45 healthy controls and 74 UC samples.
[0063] The results are as follows Figure 14-17 As shown, Figure 14 Confusion matrix for ulcerative colitis (UC) diagnosis during model internal validation for the GSE107499 and GSE87466 datasets; Figure 15 ROC curves for ulcerative colitis (UC) diagnosis, internally validated for the GSE107499 and GSE87466 datasets; Figure 16 Scatter plot of risk score distribution for UC patients and healthy controls for internal model validation on the GSE107499 and GSE87466 datasets; Figure 17 A visual comparison of risk scores between UC patients and healthy controls for internal model validation on the GSE107499 and GSE87466 datasets; The confusion matrix showed that 66 cases of UC were correctly predicted as UC and 8 cases were incorrectly predicted as HC in the UC sample; 44 cases of HC were correctly predicted as HC and 1 case was incorrectly predicted as UC in the HC sample; the model's sensitivity was 0.985, specificity was 0.846, positive predictive value was 0.891, negative predictive value was 0.978, and the ROC curve AUC reached 0.981; the risk score distribution plot clearly distinguished UC patients from healthy controls, further confirming the model's high diagnostic accuracy in the internal dataset.
[0064] b. Internal validation samples were selected from dataset GSE92415. 183 samples were randomly selected from the dataset, including 21 healthy controls and 162 UC samples.
[0065] The results are as follows Figure 18-21 As shown, Figure 18 The confusion matrix for ulcerative colitis (UC) diagnosis during model internal validation on the GSE92415 dataset; Figure 19 ROC curves for ulcerative colitis (UC) diagnosis used for internal model validation on the GSE92415 dataset; Figure 20 Scatter plot of risk score distribution for UC patients and healthy controls for internal model validation on the GSE92415 dataset; Figure 21 A visual comparison of risk scores between UC patients and healthy controls for internal model validation on the GSE92415 dataset; The confusion matrix showed that all 162 UC samples and 21 healthy controls were correctly identified, with no misdiagnosis. The model's sensitivity, specificity, positive predictive value, and negative predictive value were all 1.0, and the ROC curve AUC reached 1.0. The risk scores of the samples showed significant separation by group, and the boundaries between the two groups were clear after sorting by risk scores, further confirming the high diagnostic accuracy of the model in the internal dataset.
[0066] 4. External validation of the risk scoring model's ability to distinguish and diagnose ulcerative colitis using clinical samples: Clinical colonic mucosal tissue samples were used as external validation samples, and relative gene expression levels were detected by qRT-PCR. Seventeen colonic mucosal tissue samples were collected from healthy individuals and patients with ulcerative colitis (5 healthy controls (HC) and 12 patients with ulcerative colitis (UC)). The tissues were thoroughly ground using a micro-volume human tissue universal total RNA extraction kit to obtain total RNA. Genomic DNA was removed from the total RNA using the HiScript II Q RT SuperMix for qRT-PCR (+gDNA wiper) kit, followed by reverse transcription to convert the total RNA into cDNA. Using the cDNA as a template, qRT-PCR amplification was performed using the SYBR Green method. The reaction system contained SYBR Mix-U, forward and reverse primers, and the DNA template. GAPDH was used as an internal reference gene. A two-step amplification procedure was employed, melting curves were plotted, and the relative expression level of the target gene was calculated using the 2^(-ΔΔCt) method.
[0067] The primer sequences used in the above process are: GAPDH (Internal Reference): Upstream primer (SEQ NO:1): ctgggctacactgagcacc Downstream primer (SEQ NO:2): aagtggtcgttgagggcaatg HMGCS2: Upstream primer (SEQ NO:3): ttgccctggaggtctattttcc Downstream primer (SEQ NO:4): gaagcccatacgggtctgg GCNT2: Upstream primer (SEQ NO:5): gtgaaactgcgatacaaccc Downstream primer (SEQ NO:6): cctgaagctataaacgtggtc B4GALNT2: Upstream primer (SEQ NO:7): acctgttccagaatgccagg Downstream primer (SEQ NO:8): ggtccctggagtttagtggc NAT8B: Upstream primer (SEQ NO:9): caccttccggcgattactgaa Downstream primer (SEQ NO:10): gagggccagaatccaggag LPCAT1: Upstream primer (SEQ NO:11): cgcctcactcgtcctacttc Downstream primer (SEQ NO:12): ttccccagatcgggatgtctc PPARGC1A: Upstream primer (SEQ NO:13): gctttctgggtggactcaagt Downstream primer (SEQ NO:14): gagggcaatccgtcttcatcc ACSS2: Upstream primer (SEQ NO:15): gccagtactgtctgcccatt Downstream primer (SEQ NO:16): ggtgggaacagaggacaagg PDK2: Upstream primer (SEQ NO:17): aggggcacccaagtacatc Downstream primer (SEQ NO:18): tgccggaggaaagtgaatgac HADHA: Upstream primer (SEQ NO:19): tgacccgaagaagctgaatt Downstream primer (SEQ NO:20): tagctacatccacaccaact The risk score for each sample was calculated using the model. The relative gene expression levels and risk scores for each sample are shown in Table 1. Table 1. Relative gene expression levels and risk scores in clinical samples.
[0068] The model was clinically validated by combining logistic regression with the relative expression levels of core genes. Validation metrics included confusion matrix, sensitivity, specificity, positive predictive value, negative predictive value, AUC value, and Youden index, further confirming the model's clinical applicability and diagnostic efficacy.
[0069] The results are as follows Figure 22-25 As shown, Figure 22 Confusion matrix for externally validated diagnosis of ulcerative colitis (UC); Figure 23 ROC curve for externally validated diagnosis of ulcerative colitis (UC); Figure 24 Scatter plot of risk score distribution for UC patients and healthy controls for external validation; Figure 25 A visual comparison chart of risk scores between UC patients and healthy controls for external validation; The confusion matrix showed 5 true negatives (correctly identified HC), 0 false positives (HC misdiagnosed as UC), 1 false negative (UC misdiagnosed as HC), and 11 true positives (correctly identified UC). The core diagnostic indicators were: sensitivity 0.92, specificity 1.00, positive predictive value 1.00, negative predictive value 0.83, accuracy 0.94, Youden index 0.92, and AUC value 0.98. This model is applicable to colonic mucosal tissue sample testing, especially for rapid diagnosis of clinical samples. It provides objective and quantitative criteria for UC diagnosis, solving the problems of insufficient sensitivity and specificity in traditional diagnostic methods, and providing accurate reference for clinical diagnosis and treatment.
[0070] Example 2: A kit for predicting ulcerative colitis, comprising: The reagents for detecting the relative expression levels of the acetylation-related core genes include primers, internal controls, and qRT-PCR detection reagents. The ulcerative colitis prediction model from Example 1 is also included.
[0071] The core genes involved in acetylation include: HMGCS2, GCNT2, B4GALNT2, NAT8B, LPCAT1, PPARGC1A, ACSS2, PDK2, and HADHA.
[0072] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A predictive model for ulcerative colitis, characterized in that, The model is as follows: Risk score for each sample = 68.798 – 11.031 * HMGCS2 relative expression level – 46.586 * GCNT2 relative expression level + 20.391 * B4GALNT2 relative expression level + 5.719 * NAT8B relative expression level + 26.562 * LPCAT1 relative expression level + 17.042 * PPARGC1A relative expression level – 10.500 * ACSS2 relative expression level + 3.110 * PDK2 relative expression level – 77.421 * HADHA relative expression level.
2. The ulcerative colitis prediction model according to claim 1, characterized in that: A risk score less than zero is defined as non-ulcerative colitis, while a risk score greater than zero is defined as ulcerative colitis.
3. A method for constructing a predictive model for ulcerative colitis according to any one of claims 1-2, characterized in that: (1) Gene expression data of ulcerative colitis patients and healthy controls were obtained from the database. The acetylation-related core genes were screened out using machine learning algorithms combined with LOOCV and the data were preprocessed. (2) Analyze the core genes related to acetylation that were screened; (3) A risk scoring model was constructed for acetylation-related core genes, and its ability to distinguish and diagnose ulcerative colitis was validated internally using the dataset; (4) Use clinical samples to externally validate the risk scoring model’s ability to identify and diagnose ulcerative colitis.
4. The method for constructing a predictive model for ulcerative colitis according to claim 3, characterized in that: The core acetylation-related genes identified were: HMGCS2, GCNT2, B4GALNT2, NAT8B, LPCAT1, PPARGC1A, ACSS2, PDK2, and HADHA.
5. The method for constructing a predictive model for ulcerative colitis according to claim 3, characterized in that, The specific steps of (1) are as follows: obtain the dataset from the GEO gene expression comprehensive database, obtain the gene expression data of colonic mucosal tissue of patients with ulcerative colitis and healthy controls, and use the R language limma package to screen for differentially expressed genes in the data; Acetylation-related genes were screened from the MSigDB and GeneCards databases, and candidate genes were obtained by taking the intersection with differentially expressed genes. Machine learning algorithms were integrated, and a prediction model was constructed and optimized under the leave-one-out cross-validation framework. Finally, the core acetylation-related genes were determined through feature screening.
6. The method for constructing a predictive model for ulcerative colitis according to claim 3, characterized in that, The specific steps of (2) are as follows: Analyze the inter-gene correlation and chromosomal localization of acetylation-related core genes; construct a regulatory network of miRNAs, transcription factors, and core acetylation-related genes using the NetworkAnalyst platform, combined with miRTarBase v9.0 and the TRRUST database, and screen potential drug targets; draw calibration curves and clinical decision curves using the R language rmda package, calculate the area under the ROC curve using the pROC package, and evaluate the diagnostic ability of 9 acetylation-related core genes for ulcerative colitis; perform GSEA gene set enrichment analysis based on core ARGs to explore associated signaling pathways, use LM22 feature analysis and cibersort algorithm to detect immune cell abundance, and analyze the difference in immune cell infiltration between high and low expression groups of acetylation-related core genes; use the R language ConsensusClusterPlus unsupervised clustering package to divide ulcerative colitis sample subgroups based on the expression profile of core ARGs, verify subgroup differences through principal component analysis, and perform cross-reference analysis to verify subtype stability; use Mendelian randomization method, based on SNPs Using the instrumental variable, we explored the causal relationship between acetylation-related core genes and susceptibility to ulcerative colitis through the inverse variance weighting method combined with the MR-Egger method.
7. The method for constructing a predictive model for ulcerative colitis according to claim 3, characterized in that, The specific steps of (3) are as follows: Based on logistic regression analysis of the relative expression levels of genes HMGCS2, GCNT2, B4GALNT2, NAT8B, LPCAT1, PPARGC1A, ACSS2, PDK2 and HADHA, a risk scoring model is established. Samples are selected from the dataset, and internal validation is performed by combining logistic regression with the relative expression levels of core ARGs. The diagnostic efficacy of the model is evaluated by using the confusion matrix, sensitivity, specificity, positive predictive value, negative predictive value and ROC curve.
8. The method for constructing a predictive model for ulcerative colitis according to claim 3, characterized in that, The specific steps of (4) are as follows: take clinical colonic mucosal tissue samples from normal people and patients with ulcerative colitis, use qRT-PCR to detect the relative expression level of acetylation-related core genes, and use logistic regression combined with the relative expression level of core genes to clinically validate the model. The validation indicators include confusion matrix, sensitivity, specificity, positive predictive value, negative predictive value, area under the curve and Youden index, to further confirm the clinical applicability and diagnostic efficacy of the model.
9. A kit for predicting ulcerative colitis, characterized in that, The kit includes: The reagent for detecting the relative expression levels of the acetylation-related core genes, including HMGCS2, GCNT2, B4GALNT2, NAT8B, LPCAT1, PPARGC1A, ACSS2, PDK2, and HADHA, and the ulcerative colitis prediction model according to any one of claims 1-2.
10. The ulcerative colitis prediction kit according to claim 9, characterized in that: The reagents for detecting the relative expression levels of the acetylation-related core genes include primers, internal controls, and qRT-PCR detection reagents.