Application of biomarker in preparation of inflammatory bowel disease screening or subtype detection product
By using biomarkers such as GALNT14 and MICAL2, combined with Point-biserial correlation coefficient test and logistic regression model, the problems of low accuracy and insufficient specificity of IBD diagnosis in the prior art were solved, and high accuracy and specificity of inflammatory bowel disease screening and subtype detection were achieved.
Patent Information
- Application Number
- CN202510061556.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
AI Technical Summary
The lack of specificity of existing IBD peripheral blood biomarkers for the diagnosis of ulcerative colitis and Crohn's enteritis, resulting in a low diagnostic accuracy and the inability to distinguish between the two in early conditions.
Biomarkers such as GALNT14 and MICAL2 were analyzed by Point-biserial correlation coefficient test and machine learning model (logistic regression) to prepare inflammatory bowel disease screening and subtype detection products.
It improves the accuracy of screening and subtype detection of inflammatory bowel disease, has good specificity, can effectively distinguish ulcerative colitis from Crohn's enteritis, and reduces invasive examinations and psychological burden on patients.
Smart Images

Figure CN119955922A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biomarkers, and in particular to the application of biomarkers in the preparation of inflammatory bowel disease screening or subtype detection products. Background Art
[0002] Inflammatory bowel disease (IBD) is a long-term inflammatory disease that affects the rectum and colon. It is clinically divided into ulcerative colitis (UC) and Crohn's colitis (CD). IBD is difficult to cure completely and often accompanies patients throughout their lives. Patients often experience alternating improvements and deteriorations in their condition. This repeated course of disease not only seriously reduces the patient's quality of life, but also further increases the patient's physical and psychological burden. At the same time, the treatment and course management of IBD requires long-term and frequent medical interventions, which not only imposes an economic burden on patients, but also leads to a huge investment in medical resources.
[0003] However, there is a lack of a "gold standard" for the diagnosis of IBD. Since there may be a mismatch between the patient's subjective symptoms and the actual situation of intestinal inflammation, current diagnosis and treatment monitoring and follow-up mainly rely on digestive endoscopy. In order to accurately determine the severity of the disease and adjust medication accordingly, patients often need to undergo digestive endoscopy at least once a year or more. However, this invasive examination method has the risk of intestinal perforation and is costly. When the cause of the patient's disease is difficult to identify, colonoscopic biopsy often brings a serious psychological burden to the patient. In addition, ulcerative colitis and Crohn's enteritis are similar in the early stages of clinical symptoms and are easily missed or misdiagnosed, and up to 15% of IBD patients are defined as "unexplained colitis" in the early stages of the disease. This hinders IBD patients from formulating treatment plans and intervention strategies in a timely manner, resulting in prolonged and aggravated disease. Therefore, there is an urgent need to find sensitive and reliable IBD diagnostic methods to win precious time for patient treatment.
[0004] The use of peripheral blood biomarkers to diagnose diseases has the advantages of being convenient, non-invasive and light in burden. However, existing IBD peripheral blood biomarkers, including C-reactive protein (CRP) and p-antineutrophil cytoplasmic antibody (p-ANCA), lack specificity for the diagnosis of UC and CD, resulting in low diagnostic accuracy. Among them, CRP is an acute phase protein that can increase rapidly in the acute phase of systemic inflammatory response. When an inflammatory response occurs in the human body, interleukin-6 released by cells can promote the synthesis of CRP in the liver. References 1~3 (Turner, D., et al. C-Reactive Protein (CRP), Erythrocyte Sedimentation Rate (ESR) or Both? A Systematic Evaluation in Pediatric Ulcerative Colitis. Journal of Crohn's&Colitis,2011,5,423-429,doi:10.1016 / j.crohns.2011.05.003;Khanna,R.,et al.Endoscopic Scoring Indices forEvaluation of Disease Activity in Crohn's Disease.The Cochrane Database ofSystematic Reviews,2016,8,CD010642,doi:10.1002 / 14651858.CD010642.;Yoon,JY,et al.Correlations of C-Reactive Protein Levels and Erythrocyte SedimentationRates with Endoscopic Activity Indices in Patients with Ulcerative Colitis. Digestive Diseases and Sciences, 2014, 59, 829-837. doi: 10.1007 / s10620-013-2907-3.) showed that in both adult and pediatric cohorts, CRP levels were found to be correlated with clinical disease activity and endoscopic activity. However, the above markers lack sensitivity and specificity, and are often affected by extra-gastrointestinal inflammation, resulting in low diagnostic efficiency and inability to distinguish CD from UC. Therefore, it is urgent to find new sensitive and reliable peripheral blood biomarkers for the early diagnosis of IBD and the distinction between UC and CD, so as to assist in the clinical diagnosis and treatment of IBD. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a biomarker, which can be applied to clinical screening or subtype detection of inflammatory bowel disease with high accuracy.
[0006] The application of biomarkers in the preparation of inflammatory bowel disease screening or subtype detection products, wherein the biomarker used to prepare inflammatory bowel disease screening products is GALNT14, and the biomarker used for inflammatory bowel disease subtype detection products is MICAL2.
[0007] In the present invention, the Point-biserial correlation coefficient was used for testing, and the test results are shown in Table 1. The results indicate that the biomarkers (GALNT14, MICAL2) are closely related to the clinical data of IBD patients. The expression level of GLANT14 in IBD samples showed significant differences compared with normal samples, and the expression level of MICAL2 in UC and CD samples showed significant differences ( Figure 1 ).
[0008] Preferably, the biomarkers for preparing inflammatory bowel disease screening products also include one or more of NUSAP1, SIK1, and INPP5B.
[0009] In the present invention, the expression levels of NUSAP1 and SIK1 in IBD samples showed significant differences compared with those in normal samples, while the expression levels of INPP5B in IBD samples and normal samples showed no significant difference ( Figure 1 ), but the prediction accuracy can still be improved when INPP5B is combined with other differentiated biomarkers.
[0010] Preferably, the biomarkers for inflammatory bowel disease subtype detection products also include OSM or CHST7.
[0011] In the present invention, the Point-biserial correlation coefficient was used for testing, and the test results are shown in Table 1. The results suggest that OSM and CHST7 are closely related to the clinical data of IBD patients. The expression levels of OSM and CHST7 in UC and CD samples showed significant differences ( Figure 1 ).
[0012] Preferably, the inflammatory bowel disease screening product is a product used to distinguish between inflammatory bowel disease patients and healthy people.
[0013] Preferably, the inflammatory bowel disease subtype detection product is a detection product for distinguishing ulcerative colitis from Crohn's enteritis.
[0014] Preferably, the inflammatory bowel disease screening product comprises specific primers for amplifying biomarkers for preparing the inflammatory bowel disease screening product.
[0015] In the present invention, the forward and reverse primer sequences of GALNT14, NUSAP1, SIK1 and INPP5B are shown as SEQ ID NOs. 1 to 8, respectively.
[0016] Preferably, the inflammatory bowel disease subtype detection product comprises specific primers for amplifying biomarkers of the inflammatory bowel disease subtype detection product.
[0017] In the present invention, the forward and reverse primer sequences of MICAL2, OSM and CHST7 are shown in SEQ ID NOs. 11 to 16, respectively.
[0018] Preferably, when used for inflammatory bowel disease screening, the following steps are included:
[0019] Step S11, collecting a blood sample from a subject, and detecting the gene expression level of a biomarker in the blood sample for preparing an inflammatory bowel disease screening product;
[0020] Step S12, based on the logistic regression machine learning model, determines whether the subject suffers from inflammatory bowel disease through the gene expression level of the biomarker obtained in step S11.
[0021] Further preferably, the calculation formula for inflammatory bowel disease screening is:
[0022]
[0023] in, is the predicted probability of inflammatory bowel disease, that is, Under the condition of y G1 =1,
[0024] e is the natural base, is the gene expression level of biomarker x1, is the weight of biomarker x1, and b is the bias term;
[0025] If the predicted probability is less than 0.5, the subject is judged to be a healthy subject, and if the predicted probability is greater than or equal to 0.5, the subject is judged to be a patient with inflammatory bowel disease.
[0026] Preferably, when used for detecting inflammatory bowel disease subtypes, the method comprises the following steps:
[0027] Step S21, collecting a blood sample from a subject, and detecting the gene expression level and / or protein expression level of a biomarker for an inflammatory bowel disease subtype detection product in the blood sample;
[0028] Step S22, based on the logistic regression machine learning model, distinguishing the ulcerative colitis patients and Crohn's enteritis patients in the subjects through the gene expression level and / or protein expression level of the biomarker obtained in step S21.
[0029] Further preferably, when only the gene expression level of the biomarker obtained in step S21 is used, the calculation formula for detecting inflammatory bowel disease subtypes is:
[0030]
[0031] in, is the predicted probability of inflammatory bowel disease subtype, that is, Under the condition of y G2 =1,
[0032] e is the natural base, is the gene expression level of biomarker x2, is the weight of biomarker x2, and b is the bias term;
[0033] If the predicted probability is less than 0.5, the patient is judged to be a patient with ulcerative colitis, and if the predicted probability is greater than or equal to 0.5, the patient is judged to be a patient with Crohn's enteritis.
[0034] More preferably, when the gene expression level and protein expression level of the biomarker obtained in step S21 are used for judgment at the same time, the calculation formula for detecting inflammatory bowel disease subtypes is:
[0035]
[0036] Among them, p G3 is the predicted probability of inflammatory bowel disease subtype,
[0037] e is the natural base, is the gene expression level of biomarker x2, is the weight of biomarker x2, b is the bias term, is the weight of biomarker x3, is the protein expression level of biomarker x3;
[0038] If the predicted probability is less than 0.5, the patient is judged to be a patient with ulcerative colitis, and if the predicted probability is greater than or equal to 0.5, the patient is judged to be a patient with Crohn's enteritis.
[0039] In the present invention, after adjusting the calculation formula using the protein expression level, it has better diagnostic performance.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] The biomarkers of the present invention can be used in the preparation of inflammatory bowel disease screening or subtype detection products. The biomarkers can be selected in different combinations, and the detection has high accuracy, good specificity, and little damage to the subjects, and is easy to accept. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 Are the expression levels of biomarkers in IBD patients or patients with different IBD subtypes, where A is the difference in expression levels of GALNT14, NUSAP1, SIK1 and INPP5B between healthy subjects and IBD patients; B is the difference in expression levels of MICAL2, OSM and CHST7 between UC patients and CD patients.
[0043] Figure 2 Figure 2 is the principal component analysis, ROC curve analysis and confusion matrix of biomarkers in IBD screening of subjects, where A and B are the ROC curve analysis and confusion matrix of GALNT14 in IBD screening of subjects, respectively; C to E are the principal component analysis, ROC curve analysis and confusion matrix of the combined use of GALNT14 and NUSAP1 in IBD screening of subjects, respectively; F to H are the principal component analysis, ROC curve analysis and confusion matrix of the combined use of GALNT14, NUSAP1 and SIK1 in IBD screening of subjects, respectively; I to K are the principal component analysis, ROC curve analysis and confusion matrix of the combined use of GALNT14, NUSAP1, SIK1 and INPP5B in IBD screening of subjects, respectively.
[0044] Figure 3 These are the ELISA analysis results of biomarkers detected in IBD subtypes of subjects, where A and B are the differences in protein expression levels of OSM and CHST7 in healthy subjects and IBD patients, respectively; C and D are the differences in protein expression levels of OSM and CHST7 in UC patients and CD patients, respectively.
[0045] Figure 4 Figure 2 is the principal component analysis, ROC curve analysis and confusion matrix of biomarkers in detecting IBD subtypes in subjects, where A and B are the ROC curve analysis and confusion matrix of MICAL2 in detecting IBD subtypes in subjects, respectively; C to E are the principal component analysis, ROC curve analysis and confusion matrix of the combined use of MICAL2 and CHST7 in detecting IBD subtypes in subjects, respectively; F to H are the principal component analysis, ROC curve analysis and confusion matrix of the combined use of MICAL2 and OSM in detecting IBD subtypes in subjects, respectively.
[0046] Figure 5Figure 3 is the principal component analysis, ROC curve analysis and confusion matrix of biomarkers in detecting IBD subtypes in subjects, wherein A to C are the principal component analysis, ROC curve analysis and confusion matrix of the combined use of MICAL2 gene expression level and OSM protein expression level in detecting IBD subtypes in subjects; D to F are the principal component analysis, ROC curve analysis and confusion matrix of the combined use of MICAL2 gene expression level and CHST7 protein expression level in detecting IBD subtypes in subjects. DETAILED DESCRIPTION
[0047] The present invention will be further described in detail below in conjunction with the examples, but the embodiments of the present invention are not limited to the following examples.
[0048] The raw materials used in the present invention are all commercially available.
[0049] Point-biserial correlation coefficient of clinical sample validation data
[0050] Point-Biserial Correlation is a statistical method specifically used to measure the linear correlation between a binary variable and a continuous variable. Its calculation formula is:
[0051]
[0052] Where r is the point-by-point correlation coefficient, is the mean of the continuous variable in category 1 (such as “IBD patients”), is the mean of the continuous variable in category 0 (such as "normal subjects"), s n is the population standard deviation of all continuous variables, n1 is the number of samples in category 1, n0 is the number of samples in category 0, and n is the total number of samples. The value of r ranges between -1 and 1. A positive correlation (r>0) means that category 1 corresponds to a larger continuous variable value and category 0 corresponds to a smaller continuous variable value. A negative correlation (r<0) means that category 1 corresponds to a smaller continuous variable value and category 0 corresponds to a larger continuous variable value. The closer the absolute value is to 1, the stronger the correlation is; when the absolute value is close to 0, the correlation is weaker.
[0053] Table 1: Point-biserial correlation coefficients of biomarkers
[0054] Biomarkers Correlation P AUC GALNT14 0.4700 0.0000 0.794 NUSAP1 0.2790 0.0270 0.650 SIK1 -0.2610 0.0390 0.644 INPP5B -0.0060 0.9610 0.441 MICAL2 0.670 0.0000 0.883 CHST7 -0.533 0.0000 0.812 OSM 0.320 0.0370 0.312
[0055] Example 1: Screening for inflammatory bowel disease
[0056] (1) Sample collection
[0057] The subject information is as follows:
[0058] Table 2: Subject information for inflammatory bowel disease screening
[0059] Subject information IBD patients (43 cases) Normal subjects (20 cases) Gender (male / female) 31 / 12 10 / 10 Age (years) 38.7±12.9 33.3±12.0 Number of patients diagnosed for the first time (%) 7(16.2%) /
[0060] The venous blood of each subject was collected into a vacuum blood collection tube containing an anticoagulant, which was gently inverted several times to mix thoroughly, and the blood collection tube was placed in an ice water bath for temporary storage.
[0061] (2) mRNA extraction
[0062] Red blood cell lysis solution was added to the peripheral blood samples and then incubated at room temperature for 30 minutes. 1×PBS buffer solution was added to mix evenly. After centrifugation at 500g at room temperature, red blood cell lysis solution was added again. The above steps were repeated until the peripheral blood sample solution was transparent or slightly red. After centrifugation at 500g, the supernatant was discarded, Trizol reagent (Takara, 109) was added to the precipitate, and total RNA of the peripheral blood samples was extracted using Trizol reagent.
[0063] (3) PCR amplification
[0064] The extracted mRNA samples were reverse transcribed using the TransScript kit (TransGen Biotech, AT311-03). qRT-PCR quantitative detection was performed using Taq Pro Universal SYBR qPCR Master Mix (Vazyme, Q712-02) and the program was run on the QuantStudio 6Flex Real-time Fluorescence Quantitative PCR System (Applied Biosystems, Carlsbad, California, USA). The gene expression levels of biomarkers were calculated using Calculate the gene expression level of each biomarker Subtract the expression level of internal reference β-actin (CT β-actin ),Right now The specific PCR primers were synthesized by Shangya Biotechnology, and the sequences are shown in the table below:
[0065] Table 3: Biomarker primer sequences for inflammatory bowel disease screening
[0066] Primer name Primer sequence (5' to 3') GALNT14-F CCTGTCAGTCATCACCTTGTT(SEQ ID NO.1) GALNT14-R GGATGCTATGTGCTCGATGT(SEQ ID NO.2) NUSAP1-F GCAGTCTTCTGCTAGCCAAT(SEQ ID NO.3) NUSAP1-R GGCCTTTCTATCCCAGCTTAC(SEQ ID NO.4) SIK1-F CTTCAGCTACTTCGGCTTCTT(SEQ ID NO.5) SIK1-R CCCTCAAGTAACTGGGTCAATC(SEQ ID NO.6) INPP5B-F GAGAGGAGGAACCAGGACTATAA(SEQ ID NO.7) INPP5B-R CCAGCCACAAGATCACATCA(SEQ ID NO.8) β-ACTIN-F CACCATTGGCAATGAGCGGTTC(SEQ ID NO.9) β-ACTIN-R AGGTCTTTGCGGATGTCCACGT(SEQ ID NO.10)
[0067] The PCR system was as follows: the total system volume was 20 μL, specifically, cDNA was 2 μL, F-primer was 0.4 μL, R-primer was 0.4 μL, DEPC water was 7.2 μL, and SYBR qPCR Master Mix was 10 μL.
[0068] The PCR program is shown in the following table:
[0069] Table 4: PCR program parameters
[0070]
[0071] The level of Figure 1 As shown in Figure A, the gene expression levels of GALNT14, NUSAP1 and SIK1 in IBD patients were significantly different from those in normal subjects.
[0072] (4) A machine learning model based on inflammatory bowel disease screening determines whether the subject has inflammatory bowel disease
[0073] Based on the above four biomarkers, the inflammatory bowel disease screening model using the logistic regression algorithm was used for screening, where the input features were the gene expression levels of the four biomarkers. The input label is whether the sample is a patient with inflammatory bowel disease. The model calculates the formula Calculated for the subject sample, where is the predicted probability of inflammatory bowel disease, that is, Under the condition of y G1 =1, e is the natural base, is the gene expression level of biomarker x1, is the weight of biomarker x1, and b is the bias term; if the predicted probability is less than 0.5, it is judged as a healthy subject, and if the predicted probability is greater than or equal to 0.5, it is judged as a patient with inflammatory bowel disease.
[0074] The predicted probability output was used to determine whether the patient was an inflammatory bowel disease patient, and the diagnostic performance of the above biomarkers was analyzed using the receiver operating characteristic (ROC) curve and confusion matrix.
[0075] Analyze the results
[0076] (1) GALNT14 was selected as a biomarker for IBD diagnosis, and the model calculation formula (in W GALNT14 ×ΔCT GALNT14 , w GALNT14 is 1.275) analysis and ROC analysis (ROC curve is as Figure 2 A), the accuracy of the diagnostic model is 0.778, the AUC value is 0.794, and the classification results of the confusion matrix are as follows Figure 2 As shown in B, using only GALNT14 can accurately predict all IBD samples in the test sample (100%), and only 30% of the healthy control samples are correctly predicted.
[0077] (2) GALNT14 and NUSAP1 were selected as the biomarker combination for IBD diagnosis. PCA results showed that this biomarker combination had good classification performance ( Figure 2 C), calculated by the model formula (in w GALNT14 ×ΔCT GALNT14 +W NUSAP1 ×ΔCT NUSAP1 , w under this biomarker combination GALNT14 is 2.301, w NUSAP1 1.306) analysis and ROC analysis, the accuracy of the diagnostic model was 0.857, and the AUC value was 0.809 ( Figure 2 D), the classification results of the confusion matrix are as follows Figure 2 As shown in E, using GALNT14 and NUSAP1 for prediction, the model's predictive ability for healthy control samples was significantly improved; 93% of IBD samples and 70% of control samples were correctly predicted.
[0078] (3) GALNT14, NUSAP1 and SIK1 were selected as the biomarker combination for IBD diagnosis. The PCA results showed that this biomarker combination had good classification performance ( Figure 2 F), the calculation formula of the economic model ( is the sum of the product of the weights of the biomarkers GALNT14, NUSAP1, and SIK1 and their gene expression levels, where w GALNT14 is 2.269, W NUSAP1 is 1.283, w SIK1 -0.830) analysis and ROC analysis, the accuracy of the diagnostic model was 0.873, and the AUC value was 0.836 ( Figure 2 G), the classification results of the confusion matrix are as follows Figure 2 As shown in H, using GALNT14, NUSAP1 and SIK1 for prediction, the model's predictive ability for healthy control samples was further improved; 93% of IBD samples and 75% of control samples were correctly predicted.
[0079] (4) GALNT14, NUSAP1, SIK1 and INPP5B were selected as the biomarker combination for IBD diagnosis. The PCA results showed that this biomarker combination had good classification performance ( Figure 2 I), calculated by the model formula ( is the sum of the product of the weights of the biomarkers GALNT14, NUSAP1, SIK1 and INPP5B and their gene expression levels, where w GALNT14is 2.435, W NUSAP1 is 0.998, w SIK1 is -0.632, W INPP5B -0.551) analysis and ROC analysis, the accuracy of the diagnostic model was 0.889, and the AUC value was 0.849 ( Figure 2 J), the classification results of the confusion matrix are as follows Figure 2 As shown in Figure K, using GALNT14, NUSAP1, SIK1 and INPP5B for prediction, the model's predictive ability for IBD patient samples was slightly improved; 98% of IBD samples and 70% of control samples were correctly predicted.
[0080] Example 2: Detection of inflammatory bowel disease subtypes
[0081] (1) Sample collection
[0082] The subject information is as follows:
[0083] Table 5: Subject information for inflammatory bowel disease subtype testing
[0084] Subject information CD patients (21 cases) UC patients (22 cases) Gender (male / female) 18 / 3 13 / 9 Age (years) 36.4±12.4 41.0±13.2 Number of patients diagnosed for the first time (%) 0(0%) 7(31.8%)
[0085] The venous blood of each subject was collected into a vacuum blood collection tube containing an anticoagulant, which was gently inverted several times to mix thoroughly, and the blood collection tube was placed in an ice water bath for temporary storage.
[0086] (2) mRNA extraction
[0087] Extraction was performed with reference to "Step (2) mRNA extraction" in Example 1.
[0088] (3) PCR amplification
[0089] Refer to the step "Step (3) PCR amplification" in Example 1.
[0090] Gene expression levels of biomarkers used Calculate the gene expression level of each biomarker Subtract the expression level of internal reference Gapdh (CT Gapdh ),Right now The specific PCR primers were synthesized by Shangya Biotechnology, and the sequences are shown in the table below:
[0091] Table 6: Biomarker primer sequences for inflammatory bowel disease subtype detection
[0092] Primer name Primer sequence (5' to 3') MICAL2-F GCCAACTACAGCTCATCCTATT(SEQ ID NO.11) MICAL2-R GCTCTAGAACCTTCACGAACTC(SEQ ID NO.12) OSM-F CTCTTTGTGAAGCTAGGGAGTT(SEQ ID NO.13) OSM-R GCACCACCTGTCCTGATTTA(SEQ ID NO.14) CHST7-F CATTTCAACAAGGCATCCTCAC(SEQ ID NO.15) CHST7-R ACTGCACTTGGCCCTTATT(SEQ ID NO.16) Gapdh-F GGTGTGAACCATGAGAAGTATGA(SEQ ID NO.17) Gapdh-R GAGTCCTTCCACGATACCAAAG(SEQ ID NO.18)
[0093] The PCR system and PCR procedure were the same as those in Example 1.
[0094] The level of Figure 1 As shown in Figure 2B, the gene expression levels of MICAL2, OSM, and CHST7 in CD patients were significantly different from those in UC patients.
[0095] (4) Use the product number E-EL-H2247 and E13998h The protein levels of biomarkers OSM and CHST7 were detected by the kits according to the manufacturer's instructions. The results are as follows Figure 3 shown.
[0096] Protein expression level analysis: Figure 3 As shown, the protein expression levels of CHST7 in CD patients were significantly different from those in UC patients, while there was no significant difference in the protein expression level of OSM.
[0097] (5) A machine learning model based on inflammatory bowel disease screening to determine whether the subject has inflammatory bowel disease
[0098] Based on the above four biomarkers, the inflammatory bowel disease screening model with logistic regression algorithm was used for screening.
[0099] i) If only the gene expression levels of biomarkers are used for prediction, the input features are the gene expression levels of the four biomarkers The input label is whether the sample is a patient with inflammatory bowel disease. The model calculates the formula Calculated for the subject sample, where is the predicted probability of inflammatory bowel disease subtype, that is, Under the condition of y G2 =1, e is the natural base, is the gene expression level of biomarker x2, is the weight of the biomarker x2, and b is the deviation term; if the predicted probability is less than 0.5, the patient is judged to be a patient with ulcerative colitis, and if the predicted probability is greater than or equal to 0.5, the patient is judged to be a patient with Crohn's enteritis.
[0100] ii) If both the gene expression level and protein expression level of the biomarker are used, the protein expression level needs to be embedded in the LR model in the form of weights to adjust the calculation formula. The corrected calculation formula is: p G3 is the predicted probability of inflammatory bowel disease subtype, e is the natural base, is the gene expression level of biomarker x2, is the weight of biomarker x2, b is the bias term, is the weight of biomarker x3, is the protein expression level of biomarker x3; if the predicted probability is less than 0.5, the patient is judged to be a patient with ulcerative colitis; if the predicted probability is greater than or equal to 0.5, the patient is judged to be a patient with Crohn's enteritis.
[0101] The predicted probability output was used to determine whether the patient was an inflammatory bowel disease patient, and the diagnostic performance of the above biomarkers was analyzed using the receiver operating characteristic (ROC) curve and confusion matrix.
[0102] Analyze the results
[0103] (1) MICAL2 was selected as a biomarker for IBD subtype detection, and the model calculation formula (in W MICAL2 ×ΔCT MICAL2 , w MICAL2 is 0.327) analysis and ROC analysis (ROC curve is as Figure 4 A), the accuracy of the diagnostic model is 0.791, the AUC value is 0.833, and the classification results of the confusion matrix are as follows Figure 4 As shown in B, using only MICAL2 for prediction, the model has a strong ability to predict UC patients among IBD patients; 91% of UC samples and 67% of CD samples were correctly predicted.
[0104] (2) MICAL2 and CHST7 were selected as the biomarker combination for IBD subtype detection. PCA results showed that this biomarker combination had good classification performance ( Figure 4 C), calculated by the model formula (in w MICAL2 ×ΔCT MICAL2 +w CHST7 ×ΔCT CHST7 , w under this biomarker combination MICAL2 is 0.461, w CHST7 The accuracy of the diagnostic model was 0.605, and the AUC value was 0.729 ( Figure 4 D), the classification results of the confusion matrix are as follows Figure 4 As shown in E, using MICAL2 and CHST7 for prediction, the model's ability to predict CD patients in IBD patients was significantly improved; 82% of UC samples and 81% of CD samples were correctly predicted.
[0105] (3) MICAL2 and OSM were selected as the biomarker combination for IBD subtype detection. PCA results showed that this biomarker combination had good classification performance ( Figure 4 F), calculated by the model formula (in wMICAL2 ×ΔCT MICAL2 +w OSM ×ΔCT OSM , w under this biomarker combination MICAL2 is 0.296, w OSM -0.189) analysis and ROC analysis, the accuracy of the diagnostic model was 0.814, and the AUC value was 0.844 ( Figure 4 G), the classification results of the confusion matrix are as follows Figure 4 As shown in Figure H, using MICAL2 and OSM for prediction, the model's predictive ability for UC patient samples in IBD patients was slightly improved (86%), and only 33% of CD samples were correctly predicted.
[0106] (4) MICAL2 and OSM were selected as the biomarker combination for IBD subtype detection. Among them, OSM used protein expression levels. The PCA results showed that this biomarker combination had good classification performance ( Figure 5 A), calculated by the model formula (in w MICAL2 ×ΔCT MICAL2 , w under this biomarker combination MICAL2 is 0.327, w OSM 0.0001) analysis and ROC analysis, the accuracy of the diagnostic model was 0.814, and the AUC value was 0.883 ( Figure 5 B), the classification results of the confusion matrix are as follows Figure 5 As shown in C, using the gene expression level of MICAL2 and the protein expression level of OSM for prediction, the model has a strong predictive ability for UC patient samples in IBD patients; 91% of UC samples and 71% of CD samples were correctly predicted.
[0107] (5) MICAL2 and CHST7 were selected as the biomarker combination for IBD subtype detection, among which CHST7 used protein expression level. The PCA results showed that this biomarker combination had good classification performance ( Figure 5 D), calculated by the model formula (in w MICAL2 ×ΔCT MICAL2 , w under this biomarker combination MICAL2 is 0.327, w CHST7 0.0001) analysis and ROC analysis, the accuracy of the diagnostic model was 0.837, and the AUC value was 0.911 ( Figure 5 E), the classification results of the confusion matrix are as follows Figure 5As shown in F, using the gene expression level of MICAL2 and the protein expression level of CHST7 for prediction, the model's predictive ability for CD patient samples in IBD patients was slightly improved; 91% of UC samples and 76% of CD samples were correctly predicted.
[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention is described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. The use of biomarkers in the preparation of inflammatory bowel disease screening and / or subtype detection products, characterized in that: The biomarker used to prepare the inflammatory bowel disease screening product is GALNT14, and the biomarker used for the inflammatory bowel disease subtype detection product is MICAL2.
2. The use according to claim 1, characterized in that: The biomarkers used to prepare inflammatory bowel disease screening products also include one or more of NUSAP1, SIK1, and INPP5B.
3. The use according to claim 1, characterized in that: The biomarkers for inflammatory bowel disease subtype detection products also include OSM or CHST7.
4. The use according to claim 1 or 2, characterized in that: The inflammatory bowel disease screening product comprises specific primers for amplifying biomarkers for preparing the inflammatory bowel disease screening product.
5. The use according to claim 1 or 3, characterized in that: The inflammatory bowel disease subtype detection product comprises specific primers for amplifying biomarkers of the inflammatory bowel disease subtype detection product.
6. The use according to claim 1 or 2, characterized in that: When applying inflammatory bowel disease screening, the following steps are involved: Step S11, collecting a blood sample from a subject, and detecting the gene expression level of a biomarker in the blood sample for preparing an inflammatory bowel disease screening product; Step S12, based on the logistic regression machine learning model, determines whether the subject suffers from inflammatory bowel disease through the gene expression level of the biomarker obtained in step S11.
7. The use according to claim 6, characterized in that: The calculation formula for inflammatory bowel disease screening is: in, is the predicted probability of inflammatory bowel disease, that is, Under the condition of y G1 =1, e is the natural base, is the gene expression level of biomarker x1, is the weight of biomarker x1, and b is the bias term; If the predicted probability is less than 0.5, the subject is judged to be a healthy subject, and if the predicted probability is greater than or equal to 0.5, the subject is judged to be a patient with inflammatory bowel disease.
8. The use according to claim 1 or 3, characterized in that: When applying the inflammatory bowel disease subtype test, the following steps are included: Step S21, collecting a blood sample from a subject, and detecting the gene expression level and / or protein expression level of a biomarker for an inflammatory bowel disease subtype detection product in the blood sample; Step S22, based on the logistic regression machine learning model, distinguishing the ulcerative colitis patients and Crohn's enteritis patients in the subjects through the gene expression level and / or protein expression level of the biomarker obtained in step S21.
9. The use according to claim 8, characterized in that: When only the gene expression level of the biomarker obtained in step S21 is used, the calculation formula for detecting inflammatory bowel disease subtypes is: in, is the predicted probability of inflammatory bowel disease subtype, that is, Under the condition of y G2 =1, e is the natural base, is the gene expression level of biomarker x2, is the weight of biomarker x2, and b is the bias term; If the predicted probability is less than 0.5, the patient is judged to be a patient with ulcerative colitis, and if the predicted probability is greater than or equal to 0.5, the patient is judged to be a patient with Crohn's enteritis.
10. The use according to claim 9, characterized in that: When the gene expression level and protein expression level of the biomarker obtained in step S21 are used simultaneously, the calculation formula for detecting inflammatory bowel disease subtypes is: Among them, p G3 is the predicted probability of inflammatory bowel disease subtype, e is the natural base, is the gene expression level of biomarker x2, is the weight of biomarker x2, b is the bias term, is the weight of biomarker x3, is the protein expression level of biomarker x3; If the predicted probability is less than 0.5, the patient is judged to be a patient with ulcerative colitis, and if the predicted probability is greater than or equal to 0.5, the patient is judged to be a patient with Crohn's enteritis.
Citation Information
Patent Citations
biomarker
CN107530431A
Construction method of depression risk prediction model
CN118430820A
Uses of salt-inducible kinase (SIK) inhibitors
US20170224700A1
Treatment and detection methods for inflammatory bowel disease
US20240103008A1