Diagnosis means for detecting active inflammatory bowel disease

A non-invasive method using protein expression analysis in stool samples via SWATH-MS and machine learning addresses the limitations of current IBD diagnosis and monitoring, achieving high accuracy and reducing the need for invasive procedures.

WO2025107068A1PCT designated stage expired Publication Date: 2025-05-30SOCIETE DE COMMERCIALISATION DES PRODUITS DE LA RECHERCHE APPLIQUEE SOCPRA ET HUMAINES S E C +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2024/051530
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-21
Filing Date
2024-11-20
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Current methods for diagnosing inflammatory bowel disease (IBD) are invasive and not always accurate, particularly in distinguishing between Crohn's disease and ulcerative colitis, and in monitoring treatment efficacy.

Method used

A non-invasive method involving the measurement of protein expression levels in stool samples using SWATH-MS and machine learning to develop a predictive model for diagnosing IBD, distinguishing between Crohn's disease and ulcerative colitis, and monitoring treatment efficacy.

Benefits of technology

The method achieves high sensitivity and specificity in diagnosing active IBD and distinguishing between Crohn's disease and ulcerative colitis, and effectively monitors treatment efficacy, reducing the need for invasive procedures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2024051530_30052025_PF_FP_ABST
    Figure CA2024051530_30052025_PF_FP_ABST
Patent Text Reader

Abstract

It is provided a method of detecting inflammatory bowel disease (IBD) in a patient comprising the step of measuring in a sample of said patient protein expression from the sample, and determining from the measured expression the presence or absence in the patient of inflammatory bowel disease. The method comprises measuring the protein expression level measured of S100-A9, neutral ceramidase, serum albumin, chymotrypsin-C, protein S100-A4, alpha-1-acid glycoprotein 1, neprilysin, lactotransferrin, immunoglobulin lambda-like polypeptide 5, immunoglobulin heavy variable 4-28, protein S100-A8, chymotrypsin-like elastase family member 3A, IgGFc-binding protein, mucin-2, antithrombin-l 11, myeloblastin, zymogen granule membrane protein 16, annexin A2, glyceraldehyde-3-phosphate dehydrogenase, chloride anion exchanger, and / or a combination thereof.
Need to check novelty before this filing date? Find Prior Art

Description

DIAGNOSIS MEANS FOR DETECTING ACTIVE INFLAMMATORYBOWEL DISEASECROSS-REFERENCE TO RELATED APPLICATION

[0001] The present application is claiming priority from U.S. Provisional Application No. 63 / 601 ,410 filed November 21 , 2023, the content of which is hereby incorporated by reference in their entirety.TECHNICAL FIELD

[0002] It is provided a method of identifying patients who have inflammatory bowel disease by measuring proteins expression levels.BACKGROUND

[0003] Inflammatory bowel disease (IBD) is a chronic disorder of the gastrointestinal tract which affects millions of people worldwide. It is characterized by inflammation of the intestinal mucosa, leading to symptoms such as abdominal pain, diarrhea, rectal bleeding, and weight loss. During flare-ups, patients require drug treatment, such as steroids, immunosuppressants, and biological therapies, to reduce inflammation and promote healing. Several other diseases and conditions can present symptoms similar to those of IBD, including celiac disease, irritable bowel syndrome (IBS), and infectious colitis. However, each of these diseases requires different treatments. Consequently, rapid and accurate diagnosis of IBD flare-ups is essential to ensure appropriate treatment and management of this condition. This is especially true as IBD is associated with both unique and severe complications, sometimes requiring hospitalization and intestinal resection.

[0004] Currently, the gold standard for diagnosing and monitoring IBD is colonoscopy and biopsy, invasive procedures that can be uncomfortable and present risks of complications. Moreover, IBD is a lifelong disease, and repeated colonoscopies are necessary for disease follow-up, representing a significant burden for patients. Stool biomarkers have emerged as a promising non-invasive approach for IBD diagnosis and monitoring because they are in direct contact with the affected area of inflammation and pathology in IBD and can be utilized repeatedly as required. Among stool biomarkers, protein biomarkers have several advantages over other molecules since they are more stable in stool samples and can provide information on the activity and severity of the disease. Calprotectin is a calcium-binding protein that is released byinflammatory cells and is highly elevated in the feces of patients with IBD. Calprotectin is a common clinically used fecal biomarker to monitor disease activity and response to treatment and to distinguish between IBD and other gastrointestinal conditions that may have similar symptoms. However, it is not always accurate, and false-positive or falsenegative results can occur. Especially when the calprotectin value falls within the range of 100 to 300 g / g, it can be challenging to predict the transition from the remission phase to the flare-up phase of IBD.

[0005] It is thus highly desired to be provided with new non-invasive means for IBD diagnosis.SUMMARY

[0006] It is provided a method of detecting inflammatory bowel disease (IBD) in a patient comprising the step of measuring in a sample of said patient protein expression from said sample, wherein the protein expression level measured is of actin alpha skeletal muscle, actin cytoplasmic 1 , adenosine deaminase, alkaline phosphatase placental type, alpha-1-acid glycoprotein 1, alpha-1-antichymotrypsin, alpha-1- antitrypsin, alpha-amylase 1, alpha-amylase 2B, angiotensin-converting enzyme, annexin A11 , annexin A2, antithrombin-l 11, AP-1 complex subunit beta-1 , azurocidin, bis(5'-adenosyl)-triphosphatase ENPP4, cadherin-related family member 2, calcium- activated chloride channel regulator 1 , calcium-activated chloride channel regulator 4, carboxypeptidase A1 , carboxypeptidase A2, carboxypeptidase B, carcinoembryonic antigen-related cell adhesion molecule 7, chloride anion exchanger, chymotrypsin-C, chymotrypsin-like elastase family member 2A, chymotrypsin-like elastase family member 3A, chymotrypsin-like elastase family member 3B, chymotrypsinogen B, chymotrypsinogen B2, complement C3, complement C4-A, Copine-3, creatine kinase U-type mitochondrial, CUB and zona pellucida-like domain-containing protein 1 , CUB and zona pellucida-like domain-containing protein 1(27), cystatin-A, Defensin alpha 5, deleted in malignant brain tumors 1 protein, dermcidin, dipeptidase 1 , dipeptidyl peptidase 4, ectonucleotide pyrophosphatase / phosphodiesterase family member 7, elongation factor Tu mitochondrial, F-box only protein 3, ferritin light chain, galectin-3- binding protein, galectin-8, glyceraldehyde-3-phosphate dehydrogenase, hemoglobin subunit beta, hemoglobin subunit delta, histone H3.1 , histone H4, IgGFc-binding protein, immunoglobulin heavy constant alpha 1 , immunoglobulin heavy constant alpha 2, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 1(6), immunoglobulin heavy constant gamma 2, immunoglobulin heavy constantgamma 2(7), immunoglobulin heavy constant mu, immunoglobulin heavy variable 1-69, immunoglobulin heavy variable 4-28, immunoglobulin heavy variable 4-34, immunoglobulin heavy variable 4-4, immunoglobulin heavy variable 6-1 , immunoglobulin J chain, immunoglobulin kappa constant, immunoglobulin kappa variable 1-5, immunoglobulin kappa variable 1 D-17, immunoglobulin kappa variable 2D-29, immunoglobulin kappa variable 3-20, immunoglobulin kappa variable 3D-20, immunoglobulin lambda constant 3, immunoglobulin lambda constant 7, immunoglobulin lambda-like polypeptide 5, intelectin-2, intestinal-type alkaline phosphatase, IQ domain-containing protein K, isoform 12 of Titin, isoform 3 of E3 ubiquitin-protein ligase RNF170, isoform 3 of Krueppel-like factor 7, isoform alpha of pancreatic secretory granule membrane major, glycoprotein GP2, kallikrein-1 , kinesin- like protein KIF20B, lactotransferrin, low-density lipoprotein receptor-related protein 1 B, lysozyme C, maltase-glucoamylase, meprin A subunit alpha, meprin A subunit beta, mesothelin-like protein, mucin-2, myeloblastin, myeloperoxidase, myosin-1 , myosin-8, neprilysin, neutral ceramidase, neutrophil defensin 1 , neutrophil gelatinase-associated lipocalin, olfactomedin-4, pancreatic alpha-amylase, pancreatic triacylglycerol lipase, phospholipase A2, phospholipase B-like 1 , plasminogen, polymeric immunoglobulin receptor, polyubiquitin-B, probable non-functional immunoglobulin kappa variable 2D- 24, prosaposin, protein S100-A4, protein S100-A8, protein S100-A9, pyrin and HIN domain-containing protein 1 , serine protease 1 , serum albumin, serum amyloid P- component, SRC kinase signaling inhibitor 1 , superoxide dismutase [Cu-Zn], superoxide dismutase [M n] mitochondrial, synaptotagmin-like protein 4, trefoil factor 3, Xaa-Pro aminopeptidase 2, zinc finger protein 790, zymogen granule membrane protein 16; or a combination thereof.

[0007] It is also provided a method of detecting inflammatory bowel disease (IBD) in a patient comprising the step of measuring in a sample of said patient protein expression from said sample, wherein the protein expression level measured is of at least two proteins selected from the group consisting of alpha-1 -acid glycoprotein 1 , azurocidin, ferritin light chain, hemoglobin subunit delta, immunoglobulin lambda constant 3, myeloblastin, neutral ceramidase, phospholipase B-like 1, protein S100-A8, protein S100-A9, synaptotagmin-like protein 4, complement C3, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 2, kinesin-like protein KIF20B, lactotransferrin, myeloperoxidase, neprilysin. antithrombin-l 11, serum amyloid P-component, angiotensin-converting enzyme, annexin A2, carboxypeptidase A1 , carboxypeptidase A2, carcinoembryonic antigen-related cell adhesion molecule,chloride anion exchanger, chymotrypsin-like elastase family member 3B, elongation factor Tu mitochondrial, hemoglobin subunit beta, histone H4.16, neutrophil defensin 1 , olfactomedin-4, pancreatic alpha-amylase, phospholipase A2, zinc finger protein 790, adenosine deaminase, annexin A11 , bis(5'-adenosyl)-triphosphatase ENPP4, Copine- 3, dipeptidyl peptidase 4, F-box only protein 3, galectin-8, immunoglobulin lambda-like polypeptide 5, isoform 3 of Krueppel-like factor 7, mesothelin-like protein, myosin-8, pyrin and HIN domain-containing protein 1.

[0008] In an embodiment, the method provided herein further comprises the step of comparing the measured protein expression level to a detection cut-off value.

[0009] In an embodiment, the method provided herein further comprises the step of comparing the measured protein expression level with a protein expression level from a reference sample.

[0010] In another embodiment, the reference sample is from a healthy and / or unaffected patient.

[0011] In an embodiment, the sample is obtained by a non-surgically invasive procedure.

[0012] In a further embodiment, the sample is a stool sample.

[0013] In another embodiment, the sample is blood, serum, plasma, fecal, or a urine sample.

[0014] In an embodiment, the protein expression level measured is of S100A9, azurocidin (AZU1), immunoglobulin lambda constant 3, hemoglobin subunit delta, phospholipase B-like 1 (PLBD1 ), alpha-1-acid glycoprotein 1 (alpha 1-AGP), neutral ceramidase (ASAH2), or a combination thereof.

[0015] In another embodiment, the protein expression level measured is of alpha- 1-acid glycoprotein 1 , azurocidin, ferritin light chain, hemoglobin subunit delta, immunoglobulin lambda constant 3, myeloblastin, neutral ceramidase, phospholipase B-like 1 , protein S100-A8, protein S100-A9, synaptotagmin-like protein 4, complement C3, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 2, kinesin-like protein KIF20B, lactotransferrin, myeloperoxidase, neprilysin, antithrombin-l 11, serum amyloid P-component, or a combination thereof.

[0016] In an embodiment, the method provided herewith further comprises distinguishing between a patient with Crohn's disease and the patient with ulcerative colitis by measuring the protein expression level of at least two proteins selected from the group consisting of angiotensin-converting enzyme, annexin A2, carboxypeptidase A1 , carboxypeptidase A2, carcinoembryonic antigen-related cell adhesion molecule, chloride anion exchanger, chymotrypsin-like elastase family member 3B, elongation factor Tu mitochondrial, hemoglobin subunit beta, histone H4.16, neutrophil defensin 1 , olfactomedin-4, pancreatic alpha-amylase, phospholipase A2, serum amyloid P- component, and zinc finger protein 790.

[0017] In another embodiment, the method provided herewith further comprises monitoring treatment efficiency or active remission in the patient diagnosed with inflammatory bowel disease by measuring the protein expression level of at least two proteins selected from the group consisting of adenosine deaminase, alpha-1-acid glycoprotein 1 , annexin A11 , antithrombin-l 11 , azurocidin, Bis(5'-adenosyl)- triphosphatase ENPP4, chloride anion exchanger, Copine-3, dipeptidyl peptidase 4, F- box only protein 3, galectin-8, hemoglobin subunit beta, hemoglobin subunit delta, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 2, immunoglobulin lambda-like polypeptide 5, isoform 3 of Krueppel-like factor 7, lactotransferrin, mesothelin-like protein, myeloblastin, myeloperoxidase, myosin-8, neutral ceramidase, neutrophil defensin 1 , phospholipase B-like 1 , protein S100-A8, protein S100-A9, pyrin and HIN domain-containing protein 1 , serum amyloid P- component, and synaptotagmin-like protein 4.

[0018] In another embodiment, the protein expression level is measured by contacting at least two ligands to the sample.

[0019] In a further embodiment, the ligands are antibodies or bindings fragments thereof.

[0020] In an embodiment, the ligands are immobilized upon a substrate.

[0021] It is also provided a kit for detecting inflammatory bowel disease (IBD) in a patient following a method as defined herein, wherein said kit comprises means for measuring the protein expression level of at least two proteins.

[0022] In an embodiment, the means for measuring the level of the at least two proteins is an assay cartridge.

[0023] In another embodiment, wherein the level of the at least two proteins is measured by mass spectrometry.BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Reference will now be made to the accompanying drawings.

[0025] Fig. 1 illustrates a schematic representation outlines of the experimental procedure as described herein, which consists of three main steps: 1 ) sample processing and SWATH-MS analysis, involving obtaining proteome data from stool samples. 2) Data Processing, Training, and Optimizing Machine Learning Model: in this phase, a machine learning model is trained and optimized using 78 training samples. 3) Validating Model Performance: The final step involves testing the model's performance on 45 prospective samples.

[0026] Fig. 2 illustrates identification and correction of batch effect. In Fig. 2A, box plot illustrating protein distribution in unprocessed log-transformed data format across three sample batches. Fig. 2B, box plot displaying the quantile normalized data. Fig. 2C, PCA analysis of the normalized data, revealing clear clustering due to batch effects, and the arrows show the replicated samples in different batches. Fig. 2D box plot representing the influence of ComBat batch effect correction on the initial data. Fig. 2E box plot of data after batch correction on normalized data. Fig. 2F PCA analysis indicating the successful elimination of batch effects by the close representation of the replicated samples in different batches.

[0027] Fig. 3 illustrates a volcano plot summarizing differentially expressed proteins (DEPs) detected in samples obtained from active IBD patients vs Ctrl. Among the approximately 300 proteins detected as being consistently present in samples obtained from IBD patients, 48 were initially identified either as reduced or increased. Group comparisons between alBD and Ctrl were calculated by Welch's t-test, Log2(FC)| > 0.70, p-value > 1.3 in ProStar.

[0028] Fig. 4 illustrates functional analysis of differentially expressed proteins in two groups. Fig. 4A shows the significant gen ontology analysis of DEPs proteins (p- value <0.05), whereas Fig. 4B KEGG pathway enrichment analysis for DEPs proteins in active IBD and controls.

[0029] Fig. 5 illustrates a Protein Correlation Heatmap. This heatmap reveals two prominent clusters: one comprising upregulated proteins and the other containing downregulated proteins. It visually represents correlation strength, with stronger correlations and weaker ones depicted. Notably, the seven selected proteins for the model are highlighted in the map, each mainly associated with distinct clusters exhibiting low intercorrelations.

[0030] Fig. 6 illustrates unsupervised group classification via Principal component analysis (PCA) across three datasets, wherein Fig. 6A showing the original dataset comprising 250 proteins; in Fig. 6B the dataset following DEPs analysis with 48 proteins; and in Fig. 6C the dataset featuring the 7 proteins selected through the feature selection process.

[0031] Fig. 7 illustrates in Fig. 7A ROC curve analysis involved applying an optimized SVM model to the training dataset to determine the threshold, resulting in 93% sensitivity and 97% specificity. In Fig. 7B, in the ROC curve analysis of the model applied to the testing dataset, an AUG of 0.96 was obtained. Using a previously selected threshold, the model achieved 96% sensitivity and 76% specificity. In Fig. 7C, the confusion matrix displays the counts of true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN). The results highlight the high accuracy and exceptional performance of the generated model in accurately classifying the unseen data.

[0032] Fig. 8 illustrates a plot showing the expression levels of 16 selected proteins in active Crohn's disease (aCD) and active ulcerative colitis (aUC) patients. Statistical significance between the two groups was assessed using the Mann-Whitney test. Significant differences are marked by asterisks (*** q-value < 0.001 , ** q-value < 0.01 , * q-value <0.05) and "ns" denotes non-significant differences. Horizontal lines represent the median with standard deviation (SD) protein expression levels for each group. Despite these significant differences, there is no single threshold that can reliably distinguish between the groups, indicating partial overlap. This underscores the necessity of using machine learning for effective classification.DETAILED DESCRIPTION

[0033] In accordance with the present description, it is provided a method of identifying patients who have inflammatory bowel disease by measuring proteins expression levels.

[0034] Inflammatory bowel disease (IBD) flare-ups exhibit symptoms that are similar to other diseases and conditions, making diagnosis and treatment complicated. As encompassed herein, inflammatory bowel disease encompasses Crohn's disease and ulcerative colitis for example. Currently, the gold standard for diagnosing and monitoring IBD is colonoscopy and biopsy, invasive and uncomfortable procedures, and the fecal calprotectin test, which is not sufficiently accurate. Therefore, it is necessary to develop an alternative method. Enclosed is a proof of concept for the application of Sequential Window Acquisition of All Theoretical Mass Spectra (SWATH) mass spectrometry (MS) and machine learning to develop a non-invasive and accurate predictive model using the stool proteome to distinguish between active IBD and non- IBD controls. Proteome profiles of 120 samples were obtained and data processing procedures were optimized to select an appropriate pipeline. The differentially abundant analysis identified 48 proteins. Utilizing Correlation-based Feature Selection (CFS) 7 proteins were selected for proceeding steps. To identify the most appropriate predictive machine learning model, five of the most popular methods including Support Vector Machines (SVMs), Random Forests, Logistic Regression, Naive Bayes, and k- nearest neighbors (KNN) were assessed. The generated model was validated by implementing the algorithm on 42 prospective unseen datasets, the results showed a sensitivity of 96% and a specificity of 76%, indicating its effectiveness. The effectiveness of utilizing stool proteome obtained through SWATH-MS in accurately diagnosing active IBD via a machine learning mode is described herein.

[0035] To this aim, 120 samples were collected with less restrictive SOP requirements for clinical stool sampling (e.g. 4°C for up to 48h of storage before laboratory testing), then used high-throughput SWATH mass spectrometry for proteome profiling. Differentially expressed biomarkers were discovered, and then the best set was selected based on comparing different machine learning and feature selection methods. The final model was validated by the test cohort with 96% sensitivity and 76% specificity, confirming the selected markers' efficacy.

[0036] From the 120 stool samples collected, it included 68 active IBD and 52 gastrointestinal symptomatic non-IBD controls. All IBD patients had a confirmed diagnosis based on imaging, endoscopy and histological data. The age distribution of the samples ranged from 18 to 90 years and sex distributions in each group lay approximately in the equal range with no statistical difference.

[0037] The samples were analyzed using SWATH-MS in four distinct batches with 3 samples in all batches for batch effect diagnosis. Initially, batches 1-3 including 78 samples were used for retrospective analysis and model training while keeping aside batch-4 with 42 samples as a prospective validation group. For accurate peptide identification, the combined library (DDA and DIA) was utilized in conjunction with MBR (Match Between Runs) within the DIA-NN software. The DIA-NN software employs collections of deep neural networks to enhance the ability to match DIA fragmentation patterns with spectral libraries, thereby improving sensitivity. Enabling the match between runs (MBR) parameter lead to an increase in the average number of identified entities and significantly improved data completeness by reducing the occurrence of missing values (https: / / github.com / vdemichev / DiaNN). An estimated 1250 proteins and 9000 peptides were identified and quantified.

[0038] To obtain a precise differential expression protein (DEP), it is necessary to conduct accurate data analysis of quantitative proteomic studies. This involves various key steps in data processing, including normalization, batch effect correction, imputation of missing values, and appropriate statistical analysis. Since there is currently no established standard procedure for data processing in quantitative proteomics, to ensure accurate biomarker analysis, each analytical step was optimized and an appropriate pipeline was identified, as summarized in Fig. 1. To begin analysis, the data was first purified by eliminating any potential contaminants and proteins that had less than 70% valid values in each batch. After completing this step, a total of 250 proteins was left for further analysis. Afterward, logarithm transformation was applied to the intensity values as common practice for normalizing skewed data and approximating a normal distribution. To evaluate the data structure initially, a box plot was used to observe differences in variance and means (see Fig. 2A). From Fig. 2A, it is clear that there is significant variation among the samples and also batches which indicates the batch effect. In order to eliminate the unwanted non-biological variability caused by differences in procedures of sample collection, storage, preparation, and spectral data acquisition, a normalization process was necessary.

[0039] Fig. 2B shows the intensity distributions after quantile normalization displaying high similarity, which is desirable in experiments in which most features are expected to remain constant. Although normalization improves comparability among samples, it primarily focuses on aligning their overall patterns. Consequently, even after normalization, batch effects that specifically impact particular proteins or protein groups can remain a significant source of variance. To explore if data was affected by batches,a Principal component analysis (PCA) was performed using batch labels. The results depicted in Fig. 2C highlight the considerable influence of batch on the samples distribution and clustering of samples. Moreover, the replicated samples were generally not closely grouped, except for replicate D, which could randomly position. This clustering can be caused by slight differences in mass spectrometer parameters for batch 1 and also slight differences in conditions for each batch. In addition, we attempted to apply median-centering normalization as an alternative to quantile normalization to assess its impact on the batch effect. However, the PCA results did not demonstrate any noticeable improvement with this approach. To remove the batch effect, the ComBat method was used which is a popular and widely used method for gene expression data but also applicable to proteomics data (Johnson et al., 2007, Biostatistics, 8: 118-127). ComBat offers an enhanced variant of the mean shift that makes use of a Bayesian framework. The application of the ComBat algorithm on normalized data yielded a substantial improvement in correcting batch effects as seen in Figs. 2D and E. This improvement was evident in the closest representation of the replicated samples of each batch, as observed in the PCA analysis shown in Fig. 2F.

[0040] Missing values (MV) are commonly encountered in quantitative proteomics datasets, primarily due to the limitations of protein detection and random fluctuations that occur during the process of data acquisition. The presence of MV necessitates the consideration of their removal or imputation. To determine the most appropriate approach for handling missing values, it is crucial to identify the origin and type of these missing values. In general, MVs can be categorized into three types: missing values not at random (MNAR), missing values at random (MAR), and missing values completely at random (MCAR). The analysis of the data in each batch and condition, categorized as control or active IBD, revealed the absence of intentionally missing values. In other words, no proteins were exclusively present under just one condition. Additionally, comparing the replicated samples confirmed the random nature of the MVs. Random forest (RF) (Stekhoven and Buhlmann, 2012, Bioinformatics, 28: 112- 118) and K-nearest neighbors (KNN) (Troyanskaya et al., 2001 , Bioinformatics, 17: 520-525) are commonly recommended for addressing random missingness (Spratt and Ju, 2016, Modern Proteomics-Sample Preparation, Analysis and Practical Applications, 463-492).

[0041] Notably, Wang et al. have introduced the NAguideR toolkit (Wang et al., 2020, Nucleic acids research, 48: p. e-83), which incorporates 23 commonly used imputation methods and provides evaluation criteria to assist researchers in selectingthe most suitable method for their dataset. When the toolkit was applied to the dataset, RF and KNN ranked first and second as the most appropriate imputation methods. After evaluating both methods, neither of them were found significantly superior to the other. However, it was decided to proceed with the kNN method, considering N=5. The kNN method utilizes a machine learning algorithm to estimate missing values based on the values of their five closest neighbors in the feature space.

[0042] Following data cleaning and preprocessing, protein abundance data were prepared for further statistical analysis and downstream investigation. The goal was to identify the subset of proteins that demonstrated significant changes between the two conditions among the pool of 250 proteins. The t-test and Limma (Ritchie et al., 2015, Nucleic acids research, 43: p.e47) are two widely used hypothesis testing methods. In the analysis, the ProStar software was used (Wiecznorak et al., Bioinformatics, 2017, 33: 135-136), which incorporates both of these methods and provides options for both the t-tests (Student and Welch) and Limma. Considering the assumption of varying data variation and different sample size in two study groups, Welch's t-test was employed (West, 2021, Annals of Clinical Biochemistry, 58: 267-269). To identify differentially expressed proteins, two criteria were applied to identify differentially expressed proteins: a fold change (FC) ratio of at least 1.6 (i.e., |Log2(FC)| > 0.70) and a p-value less than 0.05 i.e., LoglO(p-value) > 1.3). These criteria allowed to achieve a false discovery rate (FDR) below 1% by selecting the best estimator with a piO value of less than 0.05. Using these criteria for the training group, 48 DEPs were identified as shown in the volcano plot (Fig. 3). Compared to Controls, there are 32 proteins presented as upregulated and 16 proteins shown as down regulated.

[0043] KEGG and gene ontology enrichment analysis were performed in order to gain insight into the functional role of dysregulated proteins. Gene ontology analysis revealed that these 48 proteins are mainly involved in the biological processes related to inflammatory and immune response, particularly neutrophil and myeloid cell-related processes as shown in Fig. 4A, which confirms the upregulation of inflammatory genes in patients with active IBD compared to control. Moreover, KEGG pathway enrichment analysis results revealed 7 pathways which are significantly affected by DEPS, including the Renin-angiotensin system, protein digestion and absorption, mineral absorption, the complement and coagulation cascade, the IL-17 signaling pathway, pancreatic secretion and Chagas disease. The "Renin-angiotensin system (RAS)" pathway is highly enriched and likely plays a significant role in IBD as illustrated in Fig. 4B. Some studies have reported altered levels and activities of RAS components in theinflamed mucosa. These studies suggest that RAS inhibition can have antiinflammatory effects in IBD. That is why pharmacologically inhibiting the classic RAS pathway using ACE inhibitors and angiotensin II receptor blockers (ARBs) has been a well-established strategy to treat hypertension. The next high fold enriched pathway includes protein digestion and absorption, reflect alterations in digestive functions and nutrient absorption in IBD patients. Moreover, the role of the IL-17 pathway in the pathogenesis of IBD and its involvement in inflammatory cytokine production, neutrophil recruitment and tissue remodeling has been demonstrated.

[0044] To construct a predictive model, it is essential to select relevant features while removing redundant and irrelevant ones through feature selection. This process reduces data dimensionality, improves model performance and reduces overfitting. As provided herewith, 'features' refer to the ‘proteins’. Two common classic feature selection models are the filter and wrapper methods. The main difference between them is that a filter model selects features based on intrinsic data properties, while a wrapper model involves a learning algorithm in determining feature quality. To identify the most relevant features among the 48 DEPs, five well-known feature selection methods were assessed, including correlation-based feature selection (Cfs), Boruta, information gain, gain ratio and the wrapper method.

[0045] The 7-feature subset obtained by the Cfs method showed better prediction performance than the others. Cfs is a filter-based feature selection method that chooses features based on their maximum correlation with the class variable and minimum intercorrelation. Fig. 5 illustrates the correlation heatmap among proteins, with the selected ones highlighted. This visualization demonstrates that the selected proteins are primarily chosen from distinct clusters, confirming their low intercorrelation. These 7 proteins including 6 upregulated and 1 downregulated protein as display in Table 1. Fig. 6 illustrates the enhancement in unsupervised group classification via PCA across three datasets: the original dataset comprising 250 proteins, the dataset following DEPs analysis with 48 proteins and the dataset featuring the 7 proteins selected through the feature selection process. Based on these plots, it is visually evident how effectively the two groups separate as we reduce the number of proteins and additionally, the cumulative proportion of variance explained by the first two principal components significantly increases from 22% to 61%.Table 1 : Characteristics of selected proteins:Protein name G _ene . Fold .. P.- Attribute change Value we .ig.ht.sProtein S100-A9 S100A9 6.9 0.0000 -3.9813Azurocidin AZU1 4.5 0.0000 -2.7925Immunoglobulin lambda|GLC3constant 3Hemoglobin subunit delta HBB 5.4 0.0000 -2.2529Phospholipase B-like 1 PLBD1 1.7 0.0000 -1.6708Alpha-1-acid glycoprotein 1 ORM1 2.8 0.0000 -0.9056Neutral ceramidase ASAH2 -2.6 0.0000 1.1675

[0046] The list of the final 7 selected proteins for a prediction model including their fold change and P-value across two groups is listed in Table 1. The attribute weights show different levels of importance to different features (proteins) during the model's training and classification process.

[0047] Machine learning (ML) is a powerful tool in bioinformatic analysis. Supervised machine learning refers to using quantitative proteome data with known clinical conditions to train a model for prediction of prospective samples. To identify the most appropriate predictive classifier according to the data nature, the five most popular machine learning methods including support vector machines (SVM), random forests (RF), logistic regression (LR), k-nearest neighbors (kNN) and naive Bayes (NB) were evaluated. Table 2 displays the performance metrics of five classifiers for predicting alBD from non-IBD controls in terms of accuracy, precision, recall, F-score, area under ROC curve (AU-ROC), area under precision and recall curve (AU-PRC). Detailed information on each of these parameters can be found in Table 3. These metrics were obtained using WEKA. The results indicate that the SVM classifier outperforms the others in all criteria.Table 2: Performance Metrics of ClassifiersClassifiers AccuracyJPrecision Recall scFore AU-ROC AU-PRCSVM 95% 0.97 0.93 0.96 0.95 0.96NB 90% 0.94 0.90 0.92 0.93 0.94LR 88% 0.89 0.91 0.90 0.92 0.92KNN 88% 0.91 0.89 0.90 0.90 0.89RF 87% 0.89 0.89 0.89 0.94 0.93

[0048] Table 2 compares the performance metrics of each classifier based on the prediction results obtained from the training and validation of 78 samples using 10-fold cross-validation. The SVM model outperforms the other classifiers based on the first four metrics. However, when considering the area under the curve (AUG), the RF classifier outperforms the others. Further analysis confirms that the SVM model works better for this data and provides more accurate predictions for unseen data.Table 3: General definition of different performance metrics for model selection.

[0049] In the context of machine learning, there is a risk of a model becoming overly proficient at learning from the training data, a phenomenon known as overfitting. This entails not only capturing the inherent patterns but also incorporating noise or random variations present in the data. Consequently, an overfitted model performs exceptionally well on the training data but faces challenges when attempting to apply its knowledge to new and unfamiliar data. Two strategies that help to avoid overfitting are cross validation and hyperparameter tuning. In this study 10-fold cross-validation was used for the training data where the data divided into 10 subsets, with 9 parts of the data used for training and 1 part for validation in each fold. The experiment was then repeated 10 times, with each of these subsets serving as the validation group. The final result indicated the average across the 10 folds, which provides a more realistic assessment of the model's performance. Moreover, most machine learning algorithms have parameters that can be adjusted, referred to as hyperparameters. These hyperparameters are critical for building robust and accurate models, as they help find the balance between bias and variance, thereby preventing the model from overfitting. Two common effective techniques for hyperparameter tuning are Grid search and Random search (Bischl et al., 2023, 12: p. e1484). In this analysis “tuneLength = 10” was employed as a function for each classifier as a grid search method. This means the system will perform hyperparameter tuning by randomly selecting 10 different combinations of hyperparameters for each classifier and evaluating their performanceusing cross-validation to find the best hyperparameter settings. The analysis indicates that, for an SVM classifier with a polynomial kernel of degree 1 , a scale parameter of 0.001 , and a cost parameter of 8 (C=8), it outperforms other configurations and achieves an accuracy of 0.95, a sensitivity of 0.93, and a specificity of 0.97 in classifying the training dataset.

[0050] To validate the optimized model, it was applied to 40 blind samples from batch 4. The prediction results indicated 96% sensitivity and 76% specificity, as shown in the confusion matrix in Fig. 7. Fig. 7 also illustrates that the area under the ROC curve was equal to 0.96. The results highlight the high accuracy and performance of the generated model in accurately classifying the unseen data.

[0051] It is thus demonstrated the potential use of SWATH-DIA proteomic profiling of stool samples as a tool for diagnosing active-IBD from non-IBD controls. This was achieved by employing machine learning techniques to develop a robust predictive model. To accomplish this, an experiment was designed with three main steps of 1 ) data acquisition and processing 2) training and optimizing a machine learning model based on 78 retrospective samples and 3) validating the model’s performance on 45 prospective samples. Achieving 96% sensitivity with a 0.96 AUG on the unseen dataset confirmed the model's robustness, also indicating our ability to successfully and effectively process the data obtained from four separate batches with different collection times. The processing steps included the successful removal of batch effects and employed effective methods for normalization and missing value imputation.

[0052] The batch effect was corrected by using the ComBat method. ComBat starts by adjusting each batch of data separately to have similar means and variances, then calculates the differences between the batches and uses this information to "harmonize" the data. ComBat adjusts the data for each sample in a way that minimizes the batch-related differences while preserving the true biological differences.

[0053] To impute missing values, it is crucial to understand the nature of the data and determine the reasons for their absence, which will guide the selection of an appropriate imputation method. Upon comparing the replicated samples, it was observed that the missing values were missing at random. “Zero”, “Mean” and “minimum value” are the straightforward imputation methods commonly used, they may not always be suitable, especially when the missing values occur randomly and not due to limits of detection or actual missing data. In such cases, imputing them with thesemethods could introduce bias into the analysis. Therefore, the K-nearest neighbors (KNN) imputation method was used with a setting of 5 neighbors. This implies that it leverages information from the 5 most similar samples in the dataset to estimate the missing values.

[0054] Among the differentially expressed proteins (DEPs), the highest-scoring proteins in the volcano plot were S100A8 and S100A9, which are well-known neutrophil-derived proteins, predominantly found as the S100A8 / S100A9 complex, also known as calprotectin. This finding further confirms the correctness of the analysis pathway.

[0055] Utilizing all 48 differentially expressed proteins as biomarker signatures for classification may not be practical. Therefore, the number of biomarkers needed to be reduce without compromising prediction accuracy. However, selecting only the best proteins and combining them based on previous studies does not guarantee an improvement in overall classification performance. Furthermore, in machine learning, a specific coefficient is assigned to each biomarker, known as weight, based on their importance and effect on classification in order to achieve an optimal result. For instance, Mooiweer et al. (2014, Inflammatory bowel diseases, 20: 307-314) found that the combination of fecal hemoglobin and calprotectin did not enhance their predictive accuracy compared to using fecal Hb and FC individually . Similarly, Schroder et al. (2007, Alimentary pharmacology & therapeutics, 26(7): 1035-1042) found that the combination of calprotectin, lactoferrin and neutrophile elastase did not increase predictive accuracy when compared with calprotectin alone. In this regard, using correlation-based feature selection helped to only keep the 7 most relevant proteins with maximum correlation with the class variable and minimum intercorrelation. For instance, the retaining of both S100A9 and S100A8 proteins does not provide significant additional informative value because both of them are subunits of calprotectin and exhibit a high correlation with each other. Moreover, the correlation heatmap in Fig. 5 indicates that S100A9 and S100A8 also share a high correlation with lactoferrin, and also there's a noticeable correlation between azurocidin, myeloblastin and myeloperoxidase. Although all of them were identified previously as potential IBD markers, keeping one of them would give similar results.

[0056] The seven selected proteins include up-regulated S100A9, azurocidin (AZU1 ), immunoglobulin lambda constant 3, hemoglobin subunit delta, phospholipase B-like 1 (PLBD1 ), alpha-1-acid glycoprotein 1 (alpha 1-AGP), and downregulatedneutral ceramidase (ASAH2). Two of these proteins, S100A9 and AZU1 are associated with neutrophils and play a key role in the host's defense against bacterial infections. S100A9 is, in fact, a subunit of calprotectin, accounting for approximately 60% of the total soluble proteins in the cytosol fraction of neutrophils, while AZU1 is found in the azurophilic granules of neutrophils, alongside other proteins. Hemoglobin delta is linked to occult intestinal bleeding in IBD patients and previous research has highlighted a correlation between fecal hemoglobin and calprotectin. Immunoglobulin lambda is a light chain of hemoglobin and can be indicative of an active immune system in IBD patients. The increase of free light chains (FLCs), including kappa and lambda immunoglobulins in plasma, has previously been shown in diabetes and immune system abnormalities, as well as autoimmune-based inflammatory diseases. However, the dysregulation of lambda light chains in stool and their relevance to IBD have not been studied in detail. PLBD1 is a phospholipase that can generate lipid mediators of inflammation and was first identified in neutrophils. However, to the best of our knowledge, its relationship with IBD has not been specifically investigated. Alpha 1- AGP is one of the major acute phase proteins in humans, and its serum concentration increases in response to systemic tissue injury, inflammation or infection. It is known that there is a significant increase in fecal alpha 1-AGP in active IBD patients compared to non-active patients, suggesting alpha 1-AGP as a potential biomarker for evaluating IBD activity. ASAH2 is involved in breaking down ceramides to sphingosines. Its downregulation in IBD causes ceramide accumulation in microdomains of cholesterol and sphingolipid-enriched membranes, resulting in impairment of the barrier function of the gut. The loss of ASAH2 causes elevated levels of sphingosine-1 -phosphate and systemic inflammation in ASAH2 knockout mice. These proteins collectively offer insights into the complex molecular mechanisms and potential biomarkers associated with IBD.

[0057] The superiority of SVM over other models can be attributed to various factors, including the characteristics of the data, the nature of the classes, the distribution of the features, and the inherent strengths and weaknesses of each algorithm. Some advantages of SVM over other classifiers include being less prone to overfitting due to its optimization process and regularization (controlled by the parameter C and gamma) and greater robustness to outliers and noisy data.

[0058] SVM serves as a robust technique for constructing a classifier. Its primary objective is to establish a decision boundary between two classes, facilitating the classification of data points based on their features. This decision boundary, referred toas a hyperplane, is positioned in a manner that maximizes its distance from the nearest data points of each class, which are known as support vectors. As proved herewith, the selection of the optimal kernel function is part of the hyperparameter tuning process. Depending on the nature of the data, one kernel (with a degree of 1 ) outperforms the others. This configuration is commonly referred to as "Linear SVM" or "SVM with a Linear Kernel". This setup assumes that the data is linearly separable which could be considered an advantage in simplifying the model complexity.

[0059] Accordingly, it is provided the proof of-concept that SWATH-based MS analysis can be advantageously used as an additional tool for assisting the gastroenterologist. This, in turn, can significantly enhance the effectiveness of IBD therapy and overall disease management. Moreover, this approach offers substantial advantages in terms of expediting and improving the precision of IBD diagnoses, thereby preventing the deterioration of the patient's condition due to delayed colonoscopy or inaccurate diagnosis. It also ensures the optimal prescription of drugs from the outset, maximizing treatment efficacy. Additionally, by reducing the necessity for unnecessary colonoscopies, it not only carries financial benefits but also minimizes patient discomfort and anxiety, saves time, enhances convenience, and streamlines the diagnosis and monitoring process.

[0060] It is provided a means to identify patients who have inflammatory bowel disease (including Crohn's disease and ulcerative colitis).

[0061] Although Crohn's disease and ulcerative colitis symptoms are similar, their pathological features and clinical treatments differ. Currently, distinguishing between these diseases involves invasive procedures such as colonoscopy and histopathology. The use of fecal proteins as non-invasive biomarkers offers a promising alternative due to their stability and proximity to inflamed tissues. It is provided herein a demonstration of the effectiveness of using the stool proteome for accurately distinguishing between active IBD patients and symptomatic non-IBD patients. Using SWATH / DIA proteomic profiling of stool samples from active IBD patients, combined with machine learning algorithms, a set of sixteen specific proteins was identified that effectively discriminate between CD and UC. This integration of proteomics with machine learning significantly enhances diagnostic accuracy. The model’s practicality was confirmed through validation with an independent set of samples, and its robustness was demonstrated by its ability to process data from multiple batches collected over time. This underscores the model’s real-world applicability. Importantly, all stool samples were collected underclinically compatible standard operating procedures (SOPs), ensuring relevance and predictability to clinical practice. An estimated 1 ,250 human proteins and 9,000 peptides were identified and quantified. After removing contaminants and those with less than 70% valid values among all samples, 250 proteins remained. Following the pre-processing data treating steps including normalization, batch effect correction and missing value imputation, significantly different proteins were identified between the two groups based on p-value and fold change. A total of 51 proteins were identified as differentially expressed (DEPs) between aCD and aUC and are listed below. Specifically, 32 proteins were higher in aUC compared to aCD, while 19 proteins were higher in aCD compared to aUC.

[0062] The expression levels of these 16 proteins in aCD and aUC samples are depicted in Fig. 8. Significant differences are observed. ANXA2 showed the most significant differences and Zinc Finger Protein 790 (ZNF790) exhibited no significant difference based on the Mann-Whitney U test. However, no single protein's expression level consistently distinguished the two groups, as partial overlap between aCD and aUC samples is evident for each protein. This necessitates the use of machine learning for effective classification.

[0063] To assess model performance, the average F1 scores and balanced accuracy were compared, alongside their variance across the folds, with the corresponding best-tuned hyperparameters. This approach was necessary because cross-validation randomly divides the dataset into separate training and validation sets for each fold, leading to slight differences in model performance due to variations in the training data within each split. While the XGBoost and Random Forest (RF) models achieved 100 percent accuracy, their lower mean of balanced accuracy and F1 scores point to overfitting. This indicates that although these models performed perfectly on the training data, they struggled to generalize to unseen data. In contrast, Naive Bayes (NB) and KNN exhibited high performance with less variance across folds compared to other classifiers. As a result, these two models were selected as a potential model for further evaluation on unseen data to assess their generalization ability.

[0064] The performance of the NB and KNN models on the unseen test dataset, consisting of 16 samples. The KNN model achieved a Training AUC of 0.95 and a Test AUC of 0.90, indicating some evidence of overfitting, as reflected by the drop in the F1 score from 0.93 (training) to 0.84 (test). On the other hand, the NB model shows more consistent performance, with a Training AUC of 0.96 and a Test AUC of 0.96,suggesting no signs of overfitting. Additionally, the sensitivity and specificity of the NB model on the training dataset were 0.86 and 0.89, respectively, with an overall accuracy of 0.87. On the test dataset, the sensitivity was 0.82, specificity was 0.80, and accuracy was 0.81. Based on this performance, the NB model was selected as the final classifier for distinguishing between aCD and aUC patients.

[0065] To understand the contribution of each of the 16 features to the predictions made by the Naive Bayes model, SHAP (SHapley Additive exPlanations) was used. The results highlight that ANXA2 and PLA2G1 B are the most influential features, with low values strongly associated with predicting Crohn’s disease (aCD). Similarly, for the third most important feature, APCS, lower values have a strong association with predicting ulcerative colitis (aUC). The remaining features show moderate impacts, suggesting that the model captures correlations between these features.

[0066] To determine whether the selected proteins are part of interconnected networks, a network analysis was performed on the 51 differentially expressed proteins (DEPs). The analysis revealed that 15 of the selected proteins are indeed interconnected. Interestingly, many of these selected proteins are located on the periphery of the network rather than at its core. This peripheral positioning suggests that while these proteins are part of the broader biological network, they might not serve a central hubs but still play important roles in specific processes or pathways. As shown in Table 4, proteins like ACE, AMY2A, CPA1 and APCS exhibit relatively high centrality and betweenness, highlighting their importance within the network, while others with lower centrality values may contribute to more specialized functions. The selected biomarker proteins, such as PLA2G1 B, ACE, CPA1-2 and CELA3B are associated with various biological processes, including antibacterial humoral response and proteolysis, suggesting their potential involvement in immune response and metabolic functions.Table 4: Network centrality and betweenness values for DEPs in the protein interaction network.Proteins Centrality BetweenessLYZ 0.3529 0.3076GAPDH 0.3235 0.2789LCN2 0.2353 0.1433ACE 0.2353 0.1075PLG 0.2059 0.0402AMY2B 0.1765 0.1231AMY2A 0.1765 0.1231DPP4 0.1765 0.1076S100A9 0.1765 0.0492CPA1 0.1471 0.1658H3-3B 0.1471 0.0556APCS 0.1471 0.0517PRTN3 0.1471 0.0163PIGR 0.1471 0.0137SERPINA3 0.1471 0.0137AZU1 0.1471 0.0122SERPINC1 0.1471 0.0105OLFM4 0.1176 0.2139RPS27A 0.1176 0.0174DEFA1 B 0.1176 0.0052CLCA1 0.0882 0.1658SLC26A3 0.0882 0.0285CLCA4 0.0882 0.0285HBB 0.0882 0.0049CELA3B 0.0882 0CPA2 0.0882 0PLA2G1B 0.0882 0JCHAIN 0.0588 0TUFM 0.0588 0ANXA2 0.0588 0H4-16 0.0588 0CEACAM7 0.0588 0DMBT1 0.0588 0SOD2 0.0588 0MEP1A 0.0294 0

[0067] To determine whether some gene ontology terms are enriched within that Network, a gene ontology terms enrichment analysis was ran on the 51 differentially expressed proteins. Key biological processes associated with the proteins that aredifferentially expressed between aCD and aUC and the 16 selected proteins were highlighted among the contributing proteins of each term. Among these processes, the most notable are antibacterial immune response and proteolysis, involving five of the identified proteins. The prominence of the antibacterial immune response suggests that the immune system's defense against bacteria differs between CD and UC, likely due to how each disease interacts with the gut microbiota. On the other hand, proteolysis, a process critical for immune regulation and controlling inflammation, appears to be more active in CD, where tissue damage tends to be deeper, compared to the more surfacelevel inflammation in UC. These differences help explain the distinct ways the two diseases progress and affect the body.

[0068] Accordingly, it is provided that ANXA2, CEACAM7 and PLA2G1 B consistently ranked among the top features across most selection methods. Additionally, the Mann-Whitney U test conducted on 16 selected proteins confirmed that ANXA2, CEACAM7 and PLA2G1 B exhibited the most significant difference. Furthermore, protein APCS and TUFM also emerged as the next top contributors in this analysis. SHAP value analysis reinforced this finding showing that these proteins were among the most impactful features influencing our final model’s decision-making.

[0069] ANXA2 (Annexin A) is an important member of the annexin family with a role in cancer progression and inflammation. PLA2G1 B (Phospholipase A2 Group 1 B) which showed higher expression in UC in comparison to CD in this study, is a key pancreatic enzyme in the digestive tract, crucial for fat metabolism, immune functions, and linked to gastrointestinal inflammation and tumors. While PLA2 group 2 enzymes are known to be synthesized at sites of active inflammation in Crohn's disease, there is limited research specifically on PLA2 group 1 in differentiating between UC and CD. CEACAM7 (Carcinoembryonic Antigen-Related Cell Adhesion Molecule 7), a member of the CEACAM family, is distinct in being specifically expressed in pancreatic and colorectal epithelial cells, particularly on the apical surface of differentiated intestinal epithelial cells (lECs). It plays an essential role in normal cellular differentiation and proliferation. APCS (Amyloid P component, serum, also known as SAP), is a key biomarker that consistently ranked highest in multiple selection methods, and it has higher expression in CD compared to UC. APCS helps regulate inflammation by interacting with the complement system to mediate immune cell activation, conditions (42). ANXA2, CEACAM7, PLA2G1 B, APCS, and TUFM consistently ranked as top features, supported by SHAP value, correlation, and network analyses.

[0070] It is thus encompassed a means to distinguish between patients with Crohn's disease and those with ulcerative colitis.

[0071] It is further encompassed a means to monitor treatment efficiency in patients already diagnosed with inflammatory bowel disease, and thus distinguishing patients who show signs of activity (poor response to medications) from those who are in remission (good response). This includes the group of patients who have an ambiguous test result for calprotectin, the current non-invasive standard.

[0072] Accordingly, it is provided a method of detecting IBD in a patient comprising the steps of obtaining a sample from the patient; measuring protein expression from said sample, and determining from the measured expression the presence or absence in the patient of inflammatory bowel disease. Determination may comprise comparing this measured level to a detection cut-off value. In an embodiment, the determination is made comparing the measured expression level with those of a reference sample.

[0073] As used herein, "reference sample" means any sample, standard, or level used for comparison purposes. In one embodiment, the reference sample is obtained from a healthy and / or unaffected patient, or part of the body of the same subject or patient. In another embodiment, the reference sample is obtained from untreated tissue and / or cells of the same subject or patient body.

[0074] The terms "level of expression" or "expression level" are generally used interchangeably and generally refer to the amount of polynucleotide or amino acid product or protein in a biological sample. "Expression" generally refers to the process by which genetic coding information is converted into structures that are present and operative in a cell. Thus, as used herein, "expression" of a gene can mean transcription into a polynucleotide, translation into a protein, or even post-translational modification of a protein. A transcribed polynucleotide, a translated protein, or a post-translationally modified protein fragment may also be derived from a transcript produced by alternative splicing, or a degraded transcript, or a protein by proteolysis, for example.

[0075] The expression level can be an increased or a decreased of expression.

[0076] A variety of samples may be analyzed. In certain embodiments, the samples may be obtained by a non-surgically invasive procedure from a human patient and may, for example, include blood, serum, plasma, fecal, or urine samples.

[0077] In an embodiment, the protein expression level measured is of actin alpha skeletal muscle, actin cytoplasmic 1 , adenosine deaminase, alkaline phosphatase placental type, alpha-1-acid glycoprotein 1 , alpha-1-antichymotrypsin, alpha-1- antitrypsin, alpha-amylase 1, alpha-amylase 2B, angiotensin-converting enzyme, annexin A11 , annexin A2, antithrombin-ill, AP-1 complex subunit beta-1, azurocidin, bis(5'-adenosyl)-triphosphatase ENPP4, cadherin-related family member 2, calcium- activated chloride channel regulator 1 , calcium-activated chloride channel regulator 4, carboxypeptidase A1 , carboxypeptidase A2, carboxypeptidase B, carcinoembryonic antigen-related cell adhesion molecule 7, chloride anion exchanger, chymotrypsin-C, chymotrypsin-like elastase family member 2A, chymotrypsin-like elastase family member 3A, chymotrypsin-like elastase family member 3B, chymotrypsinogen B, chymotrypsinogen B2, complement C3, complement C4-A, Copine-3, creatine kinase U-type mitochondrial, CUB and zona pellucida-like domain-containing protein 1 , CUB and zona pellucida-like domain-containing protein 1(27), cystatin-A, Defensin alpha 5, deleted in malignant brain tumors 1 protein, dermcidin, dipeptidase 1 , dipeptidyl peptidase 4, ectonucleotide pyrophosphatase / phosphodiesterase family member 7, elongation factor Tu mitochondrial, F-box only protein 3, ferritin light chain, galectin-3- binding protein, galectin-8, glyceraldehyde-3-phosphate dehydrogenase, hemoglobin subunit beta, hemoglobin subunit delta, histone H3.1 , histone H4, IgGFc-binding protein, immunoglobulin heavy constant alpha 1 , immunoglobulin heavy constant alpha 2, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 1(6), immunoglobulin heavy constant gamma 2, immunoglobulin heavy constant gamma 2(7), immunoglobulin heavy constant mu, immunoglobulin heavy variable 1-69, immunoglobulin heavy variable 4-28, immunoglobulin heavy variable 4-34, immunoglobulin heavy variable 4-4, immunoglobulin heavy variable 6-1 , immunoglobulin J chain, immunoglobulin kappa constant, immunoglobulin kappa variable 1-5, immunoglobulin kappa variable 1 D-17, immunoglobulin kappa variable 2D-29, immunoglobulin kappa variable 3-20, immunoglobulin kappa variable 3D-20, immunoglobulin lambda constant 3, immunoglobulin lambda constant 7, immunoglobulin lambda-like polypeptide 5, intelectin-2, intestinal-type alkaline phosphatase, IQ domain-containing protein K, isoform 12 of Titin, isoform 3 of E3 ubiquitin-protein ligase RNF170, isoform 3 of Krueppel-like factor 7, isoform alpha of pancreatic secretory granule membrane major, glycoprotein GP2, kallikrein-1 , kinesin- like protein KIF20B, lactotransferrin, low-density lipoprotein receptor-related protein 1 B, lysozyme C, maltase-glucoamylase, meprin A subunit alpha, meprin A subunit beta, mesothelin-like protein, mucin-2, myeloblastin, myeloperoxidase, myosin-1 , myosin-8,neprilysin, neutral ceramidase, neutrophil defensin 1 , neutrophil gelatinase-associated lipocalin, olfactomedin-4, pancreatic alpha-amylase, pancreatic triacylglycerol lipase, phospholipase A2, phospholipase B-like 1 , plasminogen, polymeric immunoglobulin receptor, polyubiquitin-B, probable non-functional immunoglobulin kappa variable 2D- 24, prosaposin, protein S100-A4, protein S100-A8, protein S100-A9, pyrin and HIN domain-containing protein 1 , serine protease 1 , serum albumin, serum amyloid P- component, SRC kinase signaling inhibitor 1 , superoxide dismutase [Cu-Zn], superoxide dismutase [Mn] mitochondrial, synaptotagmin-like protein 4, trefoil factor 3, Xaa-Pro aminopeptidase 2, ubiquitin-ribosomal protein es31 fusion protein, zinc finger protein 790, zymogen granule membrane protein 16; or, a combination thereof.

[0078] Particularly, the protein expression level measured is alpha-1 -acid glycoprotein 1, azurocidin, ferritin light chain, hemoglobin subunit delta, immunoglobulin lambda constant 3, myeloblastin, neutral ceramidase, phospholipase B-like 1, protein S100-A8, protein S100-A9, synaptotagmin-like protein 4, complement C3, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 2, kinesin-like protein KIF20B, lactotransferrin, myeloperoxidase, neprilysin. antithrombin- ill, serum amyloid P-component, angiotensin-converting enzyme, annexin A2, carboxypeptidase A1 , carboxypeptidase A2, carcinoembryonic antigen-related cell adhesion molecule, chloride anion exchanger, chymotrypsin-like elastase family member 3B, elongation factor Tu mitochondrial, hemoglobin subunit beta, histone H4.16, neutrophil defensin 1 , olfactomedin-4, pancreatic alpha-amylase, phospholipase A2, zinc finger protein 790, adenosine deaminase, annexin A11 , bis(5'-adenosyl)- triphosphatase ENPP4, Copine-3, dipeptidyl peptidase 4, F-box only protein 3, galectin- 8, immunoglobulin lambda-like polypeptide 5, isoform 3 of Krueppel-like factor 7, mesothelin-like protein, myosin-8, pyrin and HIN domain-containing protein 1 , or a combination thereof.

[0079] In an embodiment, the protein expression level measured for diagnosing inflammatory bowel disease is of alpha-1 -acid glycoprotein 1 , azurocidin, ferritin light chain, hemoglobin subunit delta, immunoglobulin lambda constant 3, myeloblastin, neutral ceramidase, phospholipase B-like 1 , protein S100-A8, protein S100-A9, synaptotagmin-like protein 4, complement C3, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 2, kinesin-like protein KIF20B, lactotransferrin, myeloperoxidase, neprilysin, antithrombin-l 11, serum amyloid P- component, or a combination thereof.

[0080] In another embodiment, the protein expression level measured is of S100A9, azurocidin (AZU1 ), immunoglobulin lambda constant 3, hemoglobin subunit delta, phospholipase B-like 1 (PLBD1), alpha-1-acid glycoprotein 1 (alpha 1-AGP), neutral ceramidase (ASAH2), or a combination thereof.

[0081] In a further embodiment, the protein expression level measured is of protein S100-A9, neutral ceramidase, serum albumin, chymotrypsin-C, protein S100-A4, alpha- 1-acid glycoprotein 1 , neprilysin, lactotransferrin, immunoglobulin lambda-like polypeptide 5, immunoglobulin heavy variable 4-28, protein S100-A8, chymotrypsin-like elastase family member 3A, IgGFc-binding protein, mucin-2, antithrombin-l 11, myeloblastin, zymogen granule membrane protein 16, annexin A2, glyceraldehyde-3- phosphate dehydrogenase, chloride anion exchanger, or a combination thereof.

[0082] In another embodiment, the protein expression level measured is of protein S100-A9, neutral ceramidase, serum albumin, chymotrypsin-C, protein S100-A4, alpha- 1-acid glycoprotein 1 , neprilysin, lactotransferrin, immunoglobulin lambda-like polypeptide 5, immunoglobulin heavy variable 4-28, or a combination thereof.

[0083] As encompassed herein, it is provided a means to distinguish between patients with Crohn's disease and those with ulcerative colitis by measuring protein expression level of angiotensin-converting enzyme, annexin A2, carboxypeptidase A1 , carboxypeptidase A2, carcinoembryonic antigen-related cell adhesion molecule, chloride anion exchanger, chymotrypsin-like elastase family member 3B, elongation factor Tu mitochondrial, hemoglobin subunit beta, histone H4.16, neutrophil defensin 1 , olfactomedin-4, pancreatic alpha-amylase, phospholipase A2, serum amyloid P- component, zinc finger protein 790, or a combination thereof.

[0084] It is further encompassed a means to monitor treatment efficiency in patients (active remission) already diagnosed with inflammatory bowel disease by measuring protein expression level of adenosine deaminase, alpha-1-acid glycoprotein 1 , annexin A11, antithrombin-l 11, azurocidin, Bis(5'-adenosyl)-triphosphatase ENPP4, chloride anion exchanger, Copine-3, dipeptidyl peptidase 4, F-box only protein 3, galectin-8, hemoglobin subunit beta, hemoglobin subunit delta, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 2, immunoglobulin lambdalike polypeptide 5, isoform 3 of Krueppel-like factor 7, lactotransferrin, mesothelin-like protein, myeloblastin, myeloperoxidase, myosin-8, neutral ceramidase, neutrophil defensin 1, phospholipase B-like 1 , protein S100-A8, protein S100-A9, pyrin and HINdomain-containing protein 1 , serum amyloid P-component, synaptotagmin-like protein 4, or a combination thereof.

[0085] In an embodiment, a combination of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, or 47 protein expression levels is measured.

[0086] In another embodiment, a combination of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, or 31 protein expression levels is measured.

[0087] In a further embodiment, a combination of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 protein expression levels is measured.

[0088] a combination of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, or 17 protein expression levels is measured.

[0089] As provided herewith a kit for detecting inflammatory bowel disease (IBD) in a patient following a method as described herein, the kit comprising means for measuring the protein expression level of the at least two proteins, and optionally instructions for use. The means can be for e.g. at least two ligands such as antibodies or bindings fragments thereof.EXAMPLE I Sample collection and analysisSample collection and Research Ethics

[0090] A total of 120 samples were obtained from the Clinical Biochemistry Lab of the CIUSSS de I'Estrie-CHUS in the context of the f-cal testing program. The research protocol for accessing stool samples from patients that have been tested for f-cal includes a reverse consent procedure for using residual stool samples and accessing the related clinical data on the Ariane network for diagnosis. This protocol has been approved by the Research Ethics Committee of the CIUSSS de I'Estrie-CHUS. Patients under 18 years were excluded from the study. When prescribed an f-cal test by their doctor, patients are instructed to collect a stool sample at home and bring it to the hospital within 24h. In the lab, a special device is used to collect a fixed amount of stool (~ 50 mg) and perform the extraction to be tested for calprotectin by ELISA. Theremaining stool samples were stored frozen at -80°C and waited for confirmation of the patient’s lack of objection from the Archive Division before being stored in the lab.Sample Preparation

[0091] Sample preparation was implemented as previously described (Gagne et al., 2022, International Journal of Molecular Sciences, 23(19): 11601 ). Briefly, 100 mg of frozen stool specimens were solubilized in 1 ml of Tris buffer (25mM Tris, pH 7.5) and centrifuged. Then the aqueous phase between the pellet and the floating residuals was recovered and stored at -80 °C until preparation for LC-MS / MS analysis. The concentration of solubilized proteins in the individual samples was measured using a BCA test. For reduction, the samples were treated with 10 mM of dithiothreitol (DTT) and for alkylation, the samples were exposed to 15 mM iodoacetamide. Subsequently, the quenching step was implemented using 10 mM DTT. The proteins were precipitated with cold acetone and methanol and digested by Trypsin / Lys-C. The cleaning and recovery of the peptides were done with a reverse-phase Strata-X polymeric SPE sorbent column (Phenomenex) according to the manufacturer’s instructions. The recovered peptides were dried under nitrogen flow at 37 °C for 45 min and stored at 4 °C until being resuspended in 20 pL of mobile phase solvent A (0.2% v / v formic acid and 3% DMSO v / v in water) before LC-MS / MS analysis.SWATH-MS Data Acquisition

[0092] The acquisition of LC-MS / MS data was conducted at the proteomics facilities located at Allumiqs in Sherbrooke, Quebec, Canada. Samples were analyzed by Eksigent pUHPLC (Eksigent, Redwood City, CA, USA) coupled to an ABSciex TripleTOF 6600 mass spectrometer equipped with an electrospray interface with a 25 pm iD capillary. Data-independent Acquisition (DIA) Sequential Window Acquisition of All Theoretical Mass Spectra (SWATH) acquisition mode was used to acquire raw data on the individual samples. Source voltage was set to 5.5 kV and maintained at 325 °C, the curtain gas was set at 35 psi, gas one was set at 27 psi and gas two was set at 10 psi. Separation was performed on a reverse-phase Kinetex XB column 0.3 mm i d., 2.6 pm particles, 150 mm (Phenomenex), which was maintained at 60 °C. Samples were injected by loop overfilling into a 5 pL loop. For the 60 min LC gradient, the mobile phase consisted of the following: solvent A (0.2% v / v formic acid and 3% DMSO v / v in water) and solvent B (0.2% v / v formic acid and 3% DMSO in EtOH) at a flow rate of 3 pL / min.Spectral Library Generation

[0093] To generate an ion library, extracted proteins from a representative pool of samples (3 IBD and 3 non-IBD controls) were separated on a 4-20% polyacrylamide gel and then reduced, alkylated, and digested in the gel. Peptides were extracted from the gel by successive rounds of dehydration and sonication and purified using reversephase SPE. Data-Dependent Acquisition (DDA) mode was used to acquire raw data on 12 gel fractions of a pooled sample. The spectral library was created following the procedure outlined in a previous study (Gagne et al., 2022, International Journal of Molecular Sciences, 23(19): p. 11601 ). Briefly, the obtained raw data (.wiff) files of DDA and DIA mode were converted into mzML format with MSConvert (GUI) from ProteoWizard (v3.0.22074). Subsequently, FragPipe software was utilized (https: / / fragpipe.nesvilab.org / ) to search the MS / MS spectra against the human proteome reviewed database (UP000005640; including isoforms and contaminants; accessible at www.uniprot.org as of March 15, 2022, containing 20,411 reviewed proteins) via MSFragger search engine. This search was conducted with default open search parameters, which included specifying a peptide length between 6 and 42, using strict trypsin as the enzyme with a maximum of 1 missed cleavage allowed, setting the maximum fragment charge to 4, designating methionine oxidation as a variable modification, and specifying carbamidomethylating as a fixed modification. The mass tolerance for precursor ions was set on ±20 ppm, and for fragment ions at 20 ppm. The false discovery rate (FDR) for both peptide and protein identifications was set at 5%. The DDA and DIA-based libraries were merged and carefully filtered to remove duplicated precursors and counted a total of 2000 proteins. This integration increased the proteome coverage of the library.Label Free Quantification Analysis

[0094] All DIA-converted data in mzML format were processed using DIA-NN software (version 1.8.1 ) with the following parameters: a fragment ion m / z range of 200 to 1800, a precursor m / z range of 300 to 1800, a precursor false-discovery rate (FDR) threshold of 1%, automatic settings for mass accuracy at both the MS2 and MS1 levels, and the scan window. Protein inference was set to 'Genes,' and the quantification strategy was 'robust LG (high accuracy).' Cross-run normalization was disabled, while match-between-runs (MBR) was enabled.Statistical and Modeling Analysis

[0095] The statistical analysis was conducted with R software (version 4.2.2) and the basement of RStudio included packages: ggplot2 for visualization, limma for normalization, sva for batch effect correction and impute for imputation. Differentially expressed proteins were identified by ProStar software (version 1.30.5). Machine learning and feature selection analysis was mainly performed using freely available WEKA software (https: / / www.cs.waikato.ac.nz / ml / weka / , version 3.8.6) and using R packages Caret (Classification And Regression Training), caretEnsemble and Boruta.EXAMPLE II Biomarkers for the Classification of Crohn's Disease and Ulcerative ColitisStool sample collection

[0096] A total of 69 stool samples were collected, comprising 46 from patients with active-CD (aCD) and 23 from active-UC (aUC) patients. Samples were obtained from the Clinical Hematology Lab of the Centre integre universitaire de sante et de services sociaux (CIUSSS) de I'Estrie - Centre hospitalier universitaire de Sherbrooke (CHUS) through the fecal calprotectin (f-cal) testing program. Patients collected stool samples at home and brought them to the hospital within 24 hours (2 hours at room temperature, up to 24 hours refrigerated at 4°C). In the lab, ~50 mg of stool was collected for calprotectin testing by ELISA. Remaining samples were stored at -80°C and included in the study after confirming no patient objection. Criteria for inclusion were that patients were adults with active disease, defined by the presence of symptoms and endoscopy activity by the clinician. Histologic activity was considered when available although histologic activity without clinical or endoscopic activity was not considered as aCD or aUC. Samples with ambiguous diagnoses were thus excluded. The age range of participants was from 19 to 83 years, with a mean age of 49.5 years. The sex distribution was nearly equal across both groups, with 53% female and 47% male, showing no significant statistical difference.Sample Preparation for Mass Spectrometry Analysis

[0097] 100 mg of frozen stool samples were dissolved in 1 ml of lysis buffer (25 mM Tris, 1% SDS, pH 7.5) and then centrifuged. The protein concentration in the aqueous phase was determined using a BCA assay. The samples underwent reduction with 10 mM DTT, alkylation with 15 mM iodoacetamide, and quenching with 10 mM DTT. Proteins were precipitated using cold acetone and methanol, then digested with Trypsin / Lys-C enzyme. Peptides were cleaned and recovered with a reverse-phaseStrata-X SPE column, dried under nitrogen at 37°C for 45 minutes, and stored at - 80°C. Prior to LC-MS / MS analysis, peptides were resuspended in 20 L of a mobile phase solvent of 0.2% formic acid and 3% DMSO in water.Identifying and Quantifying Proteins and Peptide in Stool Samples

[0098] LC-MS / MS data acquisition was performed at the Allumiqs Solutions proteomics facility in Sherbrooke, Quebec, Canada. An Eksigent UHPLC system coupled with an ABSciex TripleTOF 6600 mass spectrometer in Data-independent Acquisition (DIA) mode was used. Chromatographic separation was achieved on a reverse-phase Kinetex XB (0.3 mm i.d., 2.6 pm particles, 150 mm) at 60°C with a 30- minute LC gradient. The mobile phase included solvent A (0.2% formic acid and 3% DMSO in water) and solvent B (0.2% formic acid and 3% DMSO in ethanol) at a flow rate of 3 pL / min. Various parameter combinations for SWATH were optimized using the SWATH Variable Window Calculator (Sciex). Spectral library generation involved processing proteins from pooled samples on a polyacrylamide gel, followed by peptide extraction. Data were acquired in DDA mode and searched against the human proteome database (UP000005640 including isoforms; Uniprot.org) using MSFragger from the FragPipecomputational platform (version 19.1 ), resulting in an initial library of 2000 proteins. Label-free quantification on the SWATH MS data was performed with DIA-NN software (version 1.8.1 : ref), using the Match Between Run (MBR) mode to reannotate the library with gene-based protein inference at a 1% FDR threshold, and with a robust LC quantification strategy, while cross-run normalization was disabled.Statistical Analysis for Discovering Potential Biomarkers

[0099] A total of 69 samples were analyzed in four separate batches by mass spectrometry. 80 percent of the samples, including 35 active Crohn's disease (aCD) and 18 active ulcerative colitis (aUC) samples, were used as a training dataset to develop a predictive model. The remaining 23 samples were set aside as a blind testing group to validate the finalized model. Data were pre-processed to enhance the detection of differentially expressed proteins and improve the performance of the developed predictive machine learning model. The data was cleaned to remove proteins and peptides with valid values in less than 70% of all samples, Iog2 transformed, quantile normalized using the preprocessCore package in R, and batch effects were detected and corrected using the combat function from the sva package. Missing values were imputed using the KNN method from the impute package.

[0100] Differentially expressed proteins (DEPs) and peptides were identified using ProStar software (version 1.30.5) through a Welch’s t-test, with a fold change (FC) threshold of at least 1.5 and a p-value less than 0.05. The p-value was adjusted to control the false discovery rate based on the method of st. boot. Four distinct feature selection algorithms were utilized in R including: (I) Recursive Feature Elimination (RFE) and (II) Boruta from Boruta package, (III) Random Forest (RF) and (IV) Regularized Random Forest (RRF) from RRF package. These algorithms were used to compute and compare the statistical significance of potential biomarkers in our dataset.Implementation of Machine Learning Algorithms

[0101] Six machine learning algorithms were trained for binary classification: k- Nearest Neighbors (K NN), Naive Bayes (NB), extreme Gradient Boosting (XGBoost), Random Forest (RF), Support Vector Machine (SVM), and Lasso and Regularized Logistic Regression (glmnet) from caret package in R. Training these mentioned algorithms based on training dataset, allowed us to compare the models based on their performance metrics including sensitivity, specificity, precision, recall and F1-score (shows the harmonic mean of precision and recall) and balanced accuracy to distinguishing aCD from aUC. The analysis was conducted using 10-fold cross- validation with three repeats to minimize the risk of overfitting. Additionally, a tuneGrid was set up with different hyperparameters for each model. This allowed testing various combinations of these parameters to find the best-performing setup for each model. SMOTE (Synthetic Minority Over-sampling Technique) was applied to address class imbalance in datasets. After comparing their performance metrics, the best-performing model were selected and applied it to a blind testing dataset to evaluate its performance.Network and gene ontology analysis of differentiating proteins

[0102] An extensive network analysis was performed of the 51 differentiating proteins and highlighted the key metrics such as centrality and betweenness for selected proteins, which provides insights into their importance within the network. The network was built using the STRING database (version 12.0). Additionally, biological processes and molecular functions associated with these proteins were investigated, utilizing Gene Ontology (GO) terms to categorize their roles in cellular activities. This integrated approach helps in understanding how these selected proteins interact within biological networks and highlighted the potential implications of their functionalsignificance. A GO terms enrichment analysis was also performed. To do so, the python package GOATOOLS (version 1.2.3) and the 2024-08-01 GO annotations release and the 2024-06-17 GO definitions release were used. As background set, the whole human proteome from UniProtKB (UP000005640), and, as query set, all the 51 differentially expressed proteins previously identified were used.

[0103] While the present disclosure has been described in connection with specific embodiments thereof, it will be understood that it is capable of further modifications and this application is intended to cover any variations, uses, or adaptations and including such departures from the present disclosure as come within known or customary practice within the art and as may be applied to the essential features hereinbefore set forth, and as follows in the scope of the appended claims.

Claims

WHAT IS CLAIMED IS:

1. A method of detecting inflammatory bowel disease (IBD) in a patient comprising the step of measuring in a sample of said patient protein expression from said sample, wherein the protein expression level measured is of at least two proteins selected from the group consisting of alpha-1-acid glycoprotein 1, azurocidin, ferritin light chain, hemoglobin subunit delta, immunoglobulin lambda constant 3, myeloblastin, neutral ceramidase, phospholipase B-like 1 , protein S100-A8, protein S100-A9, synaptotagmin-like protein 4, complement C3, immunoglobulin heavy constant gamma1. immunoglobulin heavy constant gamma 2, kinesin-like protein KIF20B, lactotransferrin, myeloperoxidase, neprilysin. antithrombin-l 11 , serum amyloid P- component, angiotensin-converting enzyme, annexin A2, carboxypeptidase A1 , carboxypeptidase A2, carcinoembryonic antigen-related cell adhesion molecule, chloride anion exchanger, chymotrypsin-like elastase family member 3B, elongation factor Tu mitochondrial, hemoglobin subunit beta, histone H4.16, neutrophil defensin 1 , olfactomedin-4, pancreatic alpha-amylase, phospholipase A2, zinc finger protein 790, adenosine deaminase, annexin A11 , bis(5'-adenosyl)-triphosphatase ENPP4, Copine- 3, dipeptidyl peptidase 4, F-box only protein 3, galectin-8, immunoglobulin lambda-like polypeptide 5, isoform 3 of Krueppel-like factor 7, mesothelin-like protein, myosin-8, pyrin and HIN domain-containing protein 1.

2. The method of claim 1 , further comprising the step of comparing the measured protein expression level to a detection cut-off value.

3. The method of claim 1 , further comprising the step of comparing the measured protein expression level with a protein expression level from a reference sample.

4. The method of claim 3, wherein the reference sample is from a healthy and / or unaffected patient.

5. The method of any one of claims 1-4, wherein the sample is obtained by a non- surgically invasive procedure.

6. The method of any one of claims 1-5, wherein the sample is a stool sample.

7. The method of any one of claims 1-5, wherein the sample is blood, serum, plasma, fecal, or a urine sample.

8. The method of any one of claims 1-7, wherein the protein expression level measured is of S100A9, azurocidin (AZU1 ), immunoglobulin lambda constant 3, hemoglobin subunit delta, phospholipase EB-like 1 (PLBD1), alpha-1-acid glycoprotein 1 (alpha 1- AGP), neutral ceramidase (ASAH2), or a combination thereof.

9. The method of any one of claims 1-7, wherein the protein expression level measured is of alpha-1-acid glycoprotein 1 , azurocidin, ferritin light chain, hemoglobin subunit delta, immunoglobulin lambda constant 3, myeloblastin, neutral ceramidase, phospholipase EB-like 1 , protein S100-A8, protein S100-A9, synaptotagmin-like protein 4, complement 03, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 2, kinesin-like protein KIF20B, lactotransferrin, myeloperoxidase, neprilysin, antithrombin-l 11, serum amyloid P-component, or a combination thereof.

10. The method of any one of claims 1-7, further comprising distinguishing between a patient with Crohn's disease and the patient with ulcerative colitis by measuring the protein expression level of at least two proteins selected from the group consisting of angiotensin-converting enzyme, annexin A2, carboxypeptidase A1 , carboxypeptidase A2, carcinoembryonic antigen-related cell adhesion molecule, chloride anion exchanger, chymotrypsin-like elastase family member 3B, elongation factor Tu mitochondrial, hemoglobin subunit beta, histone H4.16, neutrophil defensin 1 , olfactomedin-4, pancreatic alpha-amylase, phospholipase A2, serum amyloid P- component, and zinc finger protein 790.

11. The method of any one of claims 1-7, further comprising monitoring treatment efficiency or active remission in the patient diagnosed with inflammatory bowel disease by measuring the protein expression level of at least two proteins selected from the group consisting of adenosine deaminase, alpha-1-acid glycoprotein 1, annexin A11 , antithrombin-l 11, azurocidin, Bis(5'-adenosyl)-triphosphatase ENPP4, chloride anion exchanger, Copine-3, dipeptidyl peptidase 4, F-box only protein 3, galectin-8, hemoglobin subunit beta, hemoglobin subunit delta, immunoglobulin heavy constant gamma 1 , immunoglobulin heavy constant gamma 2, immunoglobulin lambda-like polypeptide 5, isoform 3 of Krueppel-like factor 7, lactotransferrin, mesothelin-like protein, myeloblastin, myeloperoxidase, myosin-8, neutral ceramidase, neutrophildefensin 1 , phospholipase B-like 1 , protein S100-A8, protein S100-A9, pyrin and HIN domain-containing protein 1 , serum amyloid P-component, and synaptotagmin-like protein 4.

12. The method of any one of claims 1-11, wherein the protein expression level is measured by contacting at least two ligands to the sample.

13. The method of claim 12, wherein the ligands are antibodies or bindings fragments thereof.

14. The method of claim 12 or 13, wherein said ligands are immobilized upon a substrate.

15. A kit for detecting inflammatory bowel disease (IBD) in a patient following a method as defined in any one of claims 1-11 , wherein said kit comprises means for measuring the protein expression level of said at least two proteins.

16. The kit of claim 15, wherein the protein expression level is measured by contacting at least two ligands to the sample.

17. The kit of 16, wherein the ligands are antibodies or bindings fragments thereof.

18. The kit of claim 16 or 17, wherein said ligands are immobilized upon a substrate.

19. The kit of any one of claims 15-18, wherein the means for measuring the level of the at least two proteins is an assay cartridge.

20. The kit of any one of claims 15-18, wherein the level of the at least two proteins is measured by mass spectrometry.

Citation Information

Patent Citations

  • Diagnosis of crohn´s disease and ulcerative colitis

    WO2019074432A1