Ulcerative colitis diagnosis marker based on expression quantity of basp1 and s100a11 genes and application thereof
By screening the BASP1 and S100A11 gene combinations using Lasso regression and multiple machine learning algorithms, a non-invasive diagnostic model for ulcerative colitis was constructed. This model addresses the issues of high invasiveness and insufficient efficacy in existing diagnostic techniques, achieving efficient and accurate diagnosis of ulcerative colitis and providing support for biological mechanisms.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU PROVINCE HOSPITAL (THE FIRST AFFILIATED HOSPITAL OF NANJING MEDICAL UNIVERSITY)
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-21
AI Technical Summary
Existing technologies for the diagnosis of ulcerative colitis are highly invasive, complex to operate, costly, and have poor patient compliance. Furthermore, existing biomarker screening strategies are limited, have high false positive rates, and fail to delve into the biological mechanisms involved, resulting in insufficient diagnostic efficacy.
We used a Lasso regression model combined with multiple machine learning algorithms to screen for BASP1 and S100A11 gene combinations. We conducted differential analysis using peripheral blood transcriptome data, constructed a binary logistic regression model, developed a non-invasive diagnostic kit, and provided biological mechanism support by combining immune cell infiltration analysis and GSEA enrichment analysis.
It improves the robustness and accuracy of biomarker screening, achieves efficient non-invasive diagnosis, with an AUC value as high as 0.861, sensitivity of 79.45%, and specificity of 78.72%, and provides biological mechanism support, forming a complete technology chain.
Smart Images

Figure CN122428031A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a diagnostic biomarker, specifically a biomarker based on... BASP1 and S100A11 Diagnostic biomarkers for ulcerative colitis based on gene expression levels and their applications. Background Technology
[0002] Ulcerative colitis (UC) is a chronic, relapsing, nonspecific inflammatory disease that primarily affects the colorectal mucosa and is one of the main types of inflammatory bowel disease (IBD). Clinical manifestations of UC include diarrhea, bloody and mucous stools, abdominal pain, and tenesmus. The disease is protracted and difficult to cure, severely impacting patients' quality of life, and long-term UC is associated with an increased risk of colorectal cancer.
[0003] Currently, the clinical diagnosis of ulcerative colitis (UC) mainly relies on a comprehensive assessment of clinical manifestations, endoscopic examinations, imaging studies, and histopathological biopsy results. Colonoscopy combined with mucosal biopsy is considered the gold standard for UC diagnosis. However, this method has drawbacks such as high invasiveness, complex procedures, high costs, and poor patient compliance, making it particularly unsuitable for early screening, dynamic monitoring of the disease, and evaluation of treatment efficacy. Therefore, the development of molecular diagnostic biomarkers based on non-invasive or minimally invasive samples such as peripheral blood and stool has become a hot topic and urgent need in UC diagnostic technology research and development.
[0004] In recent years, with the development of high-throughput sequencing and bioinformatics technologies, researchers have screened multiple differentially expressed genes or proteins related to ulcerative colitis (UC) using techniques such as transcriptomics and proteomics. However, existing research still has significant shortcomings in biomarker screening strategies and clinical translation: First, screening methods are relatively simple, mostly using traditional differential expression analysis, resulting in a large number of screened genes, a high false positive rate, and a lack of systematic evaluation of the diagnostic efficacy of gene combinations; Second, robust validation and ensemble screening strategies are lacking, with many studies using only single statistical methods or simple models, failing to comprehensively utilize the advantages of multiple machine learning algorithms for cross-validation and feature ensemble, leading to insufficient generalization ability of the screened biomarkers on independent datasets; Third, there is a disconnect between technology and mechanism, with existing technologies typically only reaching the stage of biomarker discovery and preliminary validation, failing to conduct in-depth correlation analysis between key genes and changes in the immune microenvironment and potential molecular pathways in UC, thus weakening the biological significance and clinical application persuasiveness of the biomarkers.
[0005] Although existing technologies have investigated the roles of S100 protein family members (such as S100A8 / A9, S100A12) or other individual genes in inflammatory bowel disease, there is a lack of research that systematically screens and cross-validates gene combinations with high diagnostic performance from peripheral blood transcriptome data by integrating Lasso regression with multiple machine learning algorithms. BASP1 and S100A11 Reports on this topic have not yielded a complete protocol for using this combination to construct a non-invasive diagnostic model for UC and to validate its clinical efficacy. Therefore, developing a robust, efficient, and biologically supported combination of non-invasive diagnostic biomarkers for UC has significant clinical and market value. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method based on BASP1 and S100A11 Diagnostic biomarkers for ulcerative colitis based on gene expression levels and their applications.
[0007] The technical solution of this invention to solve the above technical problems: A combination of biomarkers for the diagnosis of ulcerative colitis, comprising the biomarker combination detected by... BASP1 Genes and S100A11 The reagent composition for gene expression levels.
[0008] More preferably, the reagent includes specific recognition BASP1 Genes and S100A11 Primers, probes, or antibodies for genes.
[0009] A kit for diagnosing ulcerative colitis, comprising detection... BASP1 Genes and S100A11 Primer pairs for gene expression levels.
[0010] Further preferred primer pairs include the following sequences: BASP1 -F: SEQIDNO:1, BASP1 -R: SEQIDNO:2; S100A11 -F: SEQIDNO:3, S100A11 -R:SEQIDNO:4.
[0011] A method for screening diagnostic markers for ulcerative colitis includes the following steps: (1) Obtain peripheral blood transcriptome data from patients with ulcerative colitis and healthy controls, and perform standardization processing; (2) Perform differential expression analysis on the standardized data to obtain the differentially expressed gene set; (3) The Lasso regression model was used to reduce the dimensionality and perform initial screening of the differentially expressed gene set to obtain the first candidate gene set; (4) Apply at least two machine learning algorithms to evaluate the feature importance of the first candidate gene set and screen out the genes that rank highly in all algorithms; (5) Take the intersection of the screening results of different algorithms in step (4) to determine the core diagnostic markers.
[0012] More preferably, the machine learning algorithm described in step (4) is selected from at least two of the following: generalized linear model, random forest, support vector machine and extreme gradient boosting.
[0013] Further preferred, the core diagnostic biomarker is BASP1 Genes and S100A11 Gene.
[0014] A method for constructing a diagnostic model for ulcerative colitis includes the following steps: (1) In the test sample BASP1 Genes and S100A11 Gene expression levels; (2) Using the expression level as an input variable, construct a binary logistic regression model and output the calculated value.
[0015] Further preferred samples are peripheral blood samples.
[0016] A sort of BASP1 Genes and S100A11 The application of genes as biomarkers in the preparation of kits for diagnosing ulcerative colitis.
[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. Improved robustness and accuracy of biomarker selection: This invention abandons the single-model selection strategy and adopts an integrated strategy of "Lasso regression denoising + consensus selection of multiple machine learning algorithms," effectively reducing the risk of bias and overfitting from a single algorithm. The resulting biomarkers... BASP1 and S100A11 The gene combination demonstrated stable diagnostic efficacy (with an AUC value as high as 0.861) across multiple independent external validation datasets, proving the reliability of the screening method.
[0018] 2. Achieving a complete closed loop for biomarker R&D: This invention not only provides methods for screening and validating biomarkers, but also transforms them into specific qPCR detection kits, and validates their effectiveness through clinical samples (AUC 0.846, sensitivity 79.45%, specificity 78.72%), forming a complete technology chain from "bioinformatics mining" to "experimental validation" and then to "product prototype".
[0019] 3. Providing a non-invasive diagnostic solution supported by biological mechanisms: This invention reveals, through immune cell infiltration analysis and GSEA enrichment analysis, that... BASP1 and S100A11 The association with the infiltration levels of macrophages and T cells and specific inflammatory pathways in the intestinal inflammatory microenvironment of UC provides a solid biological basis for the clinical application of biomarkers and enhances the persuasiveness of the technical solution. Attached Figure Description
[0020] Figures 1-2 Volcano plot of differential gene expression analysis in peripheral blood samples from UC patients and healthy controls.
[0021] Figures 3-4 A diagram illustrating the feature selection process for 10-fold cross-validation of differentially expressed genes based on a Lasso regression model.
[0022] Figures 5-7 Heatmap / bar chart of feature importance scores for candidate genes using four machine learning algorithms (GLM, RF, SVM, XGB).
[0023] Figures 8-12 Analysis of key genes: core diagnostic biomarkers BASP1 and S100A11 Expression levels in UC patients and healthy controls.
[0024] Figures 13-21 External validation of key genes: BASP1 and S100A11 Diagnostic efficacy of gene combinations in three independent external datasets.
[0025] Figures 22-23 . BASP1 and S100A11 A heatmap showing the correlation between expression levels and the proportion of major immune cells infiltrating the peripheral blood of UC patients.
[0026] Figures 24-25 .based on BASP1 and S100A11 Figure 1. Gene set enrichment analysis (GSEA) results of UC patient samples grouped by expression level.
[0027] Figures 26-30 .joint BASP1 and S100A11 ROC curve of the diagnostic model in clinical validation samples. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with the accompanying drawings and embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0029] Example 1: Diagnostic markers for ulcerative colitis BASP1 and S100A11 Bioinformatics screening and validation 1. Data Sources and Preprocessing Download the following four independent datasets from the Gene Expression Omnibus (GEO) database: GSE3365, GSE94648, GSE126124, and GSE119600. GSE3365 was used as the training set for biomarker screening; GSE94648, GSE126124, and GSE119600 served as external validation sets for independent validation of biomarker diagnostic efficacy. All datasets included peripheral blood transcriptome data from UC patients and healthy controls.
[0030] The downloaded raw data is standardized, and the ComBat algorithm is used to remove batch effects to ensure comparability between data.
[0031] 2. Screening for differentially expressed genes The limma package in R (version 4.4.2) was used to compare the UC group and the healthy control group in the training set. The selection criteria were set as |log2FC|>1 and corrected P-value (adj.Pvalue)<0.05. A total of 54 differentially expressed genes were obtained, including 28 upregulated genes and 26 downregulated genes. A volcano plot was generated, and the results are shown below. Figure 1 As shown.
[0032] 3. Initial screening using Lasso regression features To reduce noise and overfitting risks, Lasso regression analysis was performed on the 54 differentially expressed genes. Using the glmnet package in R, 10-fold cross-validation was employed to determine the optimal penalty parameter λ (lambda.min), and genes with non-zero regression coefficients were selected. The results showed that 10 genes remained in the Lasso regression model, serving as the first candidate gene set. The feature selection process for Lasso regression is as follows: Figure 2 As shown.
[0033] 4. Consensus screening of multiple machine learning algorithms Using the expression levels of 10 candidate genes as input features, four machine learning classification models based on different principles were constructed: Generalized Linear Model (GLM), Random Forest (RF), Support Vector Machine (SVM), and eXtreme Gradient Boosting (XGBoost). The models were constructed using the DALEX package and related algorithm packages in the R language.
[0034] Ten candidate genes were ranked by importance using the feature importance evaluation functions built into each model. The selection criteria were set as the top 5 genes in importance across at least three models. The intersection of genes selected from different models was then used to find the gene that was uniquely shared. BASP1 and S100A11 Four algorithms score the feature importance of candidate genes as follows: Figure 3 As shown.
[0035] 5. Validation of diagnostic performance on independent datasets In the GSE3365 training set, respectively, verification was performed. BASP1 and S100A11 Differences in expression were observed. Results showed that, compared to the healthy control group, UC patients had significantly higher levels of [specific expression level] in their peripheral blood. BASP1 and S100A11 The expression levels of [specific substances] were significantly increased (P<0.05), as shown in the results. Figure 4 As shown.
[0036] To further evaluate the diagnostic efficacy of the combined two-gene approach, a combined diagnostic model was constructed by fitting the expression levels of the two genes using binary logistic regression. Receiver operating characteristic (ROC) curves were plotted, and the area under the curve (AUC) was calculated. The results showed that the AUC of this combined diagnostic model on the three validation datasets were 0.844 (GSE94648), 0.861 (GSE126124), and 0.782 (GSE119600), respectively, significantly higher than the AUC values of either single gene, indicating that... BASP1 and S100A11 Combined use results in superior diagnostic performance, as shown in the following figures. Figure 5 As shown.
[0037] 6. Immune cell infiltration analysis and pathway enrichment analysis To explore BASP1 and S100A11 To investigate the relationship between immune microenvironment and UC, the infiltration proportions of 22 immune cell types in peripheral blood samples from UC patients were analyzed using the CIBERSORT algorithm via deconvolution. Correlation analysis showed that... BASP1In UC, there was a significant positive correlation between infiltration of neutrophils and M0 macrophages (P<0.05). S100A11 Positive correlation with neutrophils and resting mast cells, and negative correlation with resting CD4+ memory T cells and M1 macrophages, as shown in the results. Figure 6 As shown.
[0038] Furthermore, gene set enrichment analysis (GSEA) was used to explore related signaling pathways to gain a deeper understanding. BASP1 and S100A11 Molecular role in UC. Results showed that high expression... BASP1 It participates in pathways such as B cell receptor signaling pathway, chemokine signaling pathway, and cytokine-receptor interaction; S100A11 It participates in pathways such as chemokine signaling, cytokine receptor interaction, and the cytotoxic effects of natural killer cells, with results as follows: Figure 7 As shown. These results provide, from a mechanistic perspective, BASP1 and S100A11 It provides biological support as a diagnostic biomarker for UC.
[0039] Example 2: Targeting BASP1 and S100A11 Development and clinical performance evaluation of qPCR detection kits 1. Primer design and synthesis Targeting people BASP1 Gene (GeneID:10409) and S100A11 Specific primers for the conserved sequence region of the gene (GeneID: 6282) were designed using Primer-BLAST software, and BLAST alignment confirmed the absence of nonspecific amplification. The primers were synthesized by Sangon Biotech (Shanghai) Co., Ltd., and their sequences are as follows: BASP1 -F: 5'-AGGGGAACCCAAAAAGACTGA-3' (SEQ ID NO: 1); BASP1 -R: 5'-GGTGTGGAACTAGGCGCTTC-3' (SEQ ID NO: 2); S100A11 -F: 5'-CTGAGCGGTGCATCGAGTC-3' (SEQ ID NO: 3); S100A11 -R: 5'-TGTGAAGGCAGCTGTTCTGTA-3' (SEQ ID NO: 4); Meanwhile, primers for the internal reference gene GAPDH were designed as a control.
[0040] 2. Clinical Sample Collection and RNA Extraction Peripheral blood samples were collected from 120 patients with active ulcerative colitis (UC) diagnosed by colonoscopy and pathology and hospitalized in the Department of Gastroenterology, First Affiliated Hospital of Nanjing Medical University between June and December 2025, and peripheral blood samples were collected from 120 healthy volunteers matched for age and sex as controls. This study was approved by the Ethics Committee of the First Affiliated Hospital of Nanjing Medical University (2025-SR-1243), and all participants signed informed consent forms.
[0041] Total RNA was extracted from peripheral blood using the Trizol method. RNA concentration and purity (A260 / A280 ratio between 1.8 and 2.0) were determined using Nanodrop 2000, and RNA integrity was detected by 1.5% agarose gel electrophoresis.
[0042] 3. Reverse transcription and real-time quantitative PCR detection 1 μg of total RNA was used to synthesize cDNA via reverse transcription using the PrimeScript™ RT reagent Kit with gDNA Eraser (TaKaRa). qPCR reactions were performed using a TB Green® Premix Ex Taq™ II (TaKaRa) on an ABI 7500 real-time quantitative PCR instrument (ABI, USA). The reaction volume was 20 μL, and the reaction conditions were: 95℃ pre-denaturation for 30 seconds; 95℃ denaturation for 5 seconds; 60℃ annealing / extension for 30 seconds, for a total of 40 cycles. Each sample was tested in triplicate. GAPDH was used as an internal reference gene, and the relative expression level of the target gene was calculated using the 2^−ΔΔCt method.
[0043] 4. Clinical validation analysis of the diagnostic model Peripheral blood samples from 120 UC patients and 120 healthy controls were tested using qPCR. BASP1 and S100A11 The expression levels of [the substance] were analyzed, and the results showed that compared with the healthy control group, [the expression levels of the substance were significantly lower]. BASP1 and S100A11 The expression was significantly elevated in the peripheral blood of UC patients (P<0.001). Furthermore, based on the modified Mayo score, 120 patients were divided into 47 in the inactive phase and 73 in the active phase. qPCR results showed that the expression of expression was significantly elevated in the peripheral blood of UC patients in the active phase. BASP1 and S100A11 The expression level was significantly elevated (P<0.001), suggesting... BASP1 and S100A11 It has greater diagnostic value for active UC.
[0044] by BASP1 and S100A11The ΔCt value was used as a covariate, and a binary logistic regression model was constructed: logit(P) = β0 + β1 * ΔCt_BASP1 + β2 * ΔCt_S100A11. The optimal cutoff value was determined through ROC curve analysis. The results showed that the AUC value of this combined diagnostic model for UC was 0.846 (95% confidence interval: 0.769–0.905), with a sensitivity of 79.45% and a specificity of 78.72%. Figure 8 As shown.
[0045] 5. Kit Components Based on the above results, the present invention provides a real-time quantitative PCR kit for detecting ulcerative colitis, comprising the following components: (1) Specific primer mixture: containing primer pairs shown in SEQ ID NO:1-4, at a concentration of 10 μM; (2) Internal reference primers: GAPDH upstream and downstream primers; (3) SYBR Green qPCR premix; (4) Water without RNase; (5) Positive control: plasmids containing BASP1 and S100A11 gene fragments; (6) Negative control: water without RNase.
Claims
1. A combination of biomarkers for diagnosing ulcerative colitis, characterized in that, The combination of markers is detected BASP1 Genes and S100A11 The reagent composition for gene expression levels.
2. The marker combination according to claim 1, characterized in that, The reagent includes specific recognition. BASP1 Genes and S100A11 Primers, probes, or antibodies for genes.
3. A kit for diagnosing ulcerative colitis, characterized in that, Includes detection BASP1 Genes and S100A11 Primer pairs for gene expression levels.
4. The reagent kit according to claim 3, characterized in that, The primer pair includes the following sequences: BASP1 -F: SEQ ID NO:1, BASP1 -R: SEQ ID NO:2; S100A11 -F: SEQ ID NO:3, S100A11 -R: SEQ ID NO:
4.
5. A method for screening diagnostic markers for ulcerative colitis, characterized in that, Includes the following steps: (1) Obtain peripheral blood transcriptome data from patients with ulcerative colitis and healthy controls, and perform standardization processing; (2) Perform differential expression analysis on the standardized data to obtain the differentially expressed gene set; (3) The Lasso regression model was used to reduce the dimensionality and perform initial screening of the differentially expressed gene set to obtain the first candidate gene set; (4) Apply at least two machine learning algorithms to evaluate the feature importance of the first candidate gene set and screen out the genes that rank highly in all algorithms; (5) Take the intersection of the screening results of different algorithms in step (4) to determine the core diagnostic markers.
6. The method according to claim 5, characterized in that, The machine learning algorithm described in step (4) is selected from at least two of the following: generalized linear model, random forest, support vector machine and extreme gradient boosting.
7. The method according to claim 5, characterized in that, The core diagnostic biomarker is BASP1 Genes and S100A11 Gene.
8. A method for constructing a diagnostic model for ulcerative colitis, characterized in that, Includes the following steps: (1) In the test sample BASP1 Genes and S100A11 Gene expression levels; (2) Using the expression level as an input variable, construct a binary logistic regression model and output the calculated value.
9. The method according to claim 8, characterized in that, The sample was a peripheral blood sample.
10. A kind BASP1 Genes and S100A11 The application of genes as biomarkers in the preparation of kits for diagnosing ulcerative colitis.