System and method for predicting a response to nutritional therapy in newly diagnosed children with crohn's disease
A machine learning module using multiomics data predicts EEN response in children with Crohn's disease, addressing the lack of accurate prediction tools by achieving high accuracy in identifying responders and non-responders, thus optimizing treatment adherence.
Patent Information
- Application Number
- PCT/IL2025/050184
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-22
- Filing Date
- 2025-02-23
- Publication Date
- 2025-08-28
AI Technical Summary
Current methods lack an accurate prediction tool to determine a child's response to exclusive enteral nutrition (EEN) for treating Crohn's disease, leading to unnecessary inconvenience due to the lengthy and challenging nature of the treatment.
A machine learning module utilizing multiomics data, including serum and fecal metabolomics, microbiome composition, demographics, and clinical parameters, is trained to predict a patient's response to EEN, employing a random forest algorithm to identify key metabolites and microbial signatures.
The system achieves high accuracy in predicting EEN response, with an AUROC of 0.98, sensitivity of 0.95, and specificity of 0.92, effectively identifying responders and non-responders, thereby optimizing treatment adherence and reducing unnecessary treatment duration.
Smart Images

Figure IL2025050184_28082025_PF_FP_ABST
Abstract
Description
[0001] SYSTEM AND METHOD FOR PREDICTING A RESPONSE TO NUTRITIONAL THERAPY IN NEWLY DIAGNOSED CHILDREN WITH CROHN’S DISEASE
[0002] RELATED APPLICATIONS
[0003] This application claims priority under 35 U.S.C. 119(e) from U.S. provisional application No. 63 / 556,410 filed on February 22, 2024, the disclosure of which is incorporated herein by reference.
[0004] FIELD OF THE DISCLOSURE
[0005] The present disclosure relates to a system and method of predicting a response to nutritional therapy for treating Crohn’s Disease (CD), and more specifically predicting based on multiomics data with a machine learning module.
[0006] BACKGROUND
[0007] The pathogenesis of Crohn's disease (CD) is thought to result from an interplay of several factors, most notably genetics, microbiome, mucosal immunology and environment factors. Of the latter, nutritional factors have been repeatedly shown to be associated with the development and progression of CD. Dietary intervention affects the immune system directly and indirectly by manipulating the microbiome, as well as the gut barrier and mucos layer. Exclusive enteral nutrition (EEN), is considered the first line treatment to induce remission in pediatric CD, achieving clinical remission in most treated of patients (around 70%). EEN is preferred over corticosteroids to promote mucosal healing, restore bone mineral density, and improve growth.
[0008] Alongside its clear advantages, EEN poses significant challenges to the patient as it requires a firm commitment of 6-8 weeks of formula only. Adherence to the treatment is difficult and includes barriers as cost, bad taste of the formula, social isolation, monotony, and excessive duration of treatment. Therefore, there is a need to identify ahead of the treatment those most likely to respond, so to avoid unnecessary inconvenience to the child as it is known that one third of children with CD will fail treatment with exclusive enteral nutrition (EEN). However, despite several efforts currently there is no prediction tool available to accurately predict response to EEN prior to treatment.
[0009] Different hypotheses assume that CD is mainly affected by different factors, such as genetics, environment, host signals or microbial. Accordingly, researchers analyze different avenues based on their hypothesis related to the main cause.
[0010] The following disclosure aims to describe use of a machine learning module based on multiomics data to predict an EEN response in treatment of naive children with CD. The disclosure is based on the hypothesis that a comprehensive biological approach using multiomics data would be more successful in predicting response to nutritional therapy, given that EEN affect various biological and microbial pathways. Metabolomics offers a holistic approach in the determination of metabolites in the biological system, and these may be closely associated with the microbiome signature.
[0011] SUMMARY
[0012] An embodiment of the disclosure relates to a system and method for predicting a response to a nutritional therapy in newly diagnosed people with Crohn’s disease. The system includes a machine learning program that is executed by a processor on a computer. The machine learning program is trained by providing it with multiomics data of patients that were treated by an exclusive enteral nutrition (EEN) program. Once trained the machine learning program accepts multiomics data of a patient and predicts if the patient will respond to EEN treatment or not.
[0013] There is thus provided according to an embodiment of the disclosure, a system for predicting a response to nutritional therapy for treating Crohn’s Disease (CD), comprising: a computer with a processor and memory; a machine learning module; wherein the machine learning module is programed to train a model by accepting a training data set comprising multiomics data of subjects with Crohn’s disease that were treated by exclusive enteral nutrition (EEN) and a determination if each subject was a responder that successfully responded to the treatment or was a non-responder that unsuccessfully responded to the treatment; and wherein the machine learning module uses the trained model to accept multiomics data of a patient and predict if the patient is a responder or a non-responder; wherein the multiomics data includes identified metabolites from serum metabolomics.
[0014] In an embodiment of the disclosure, the multiomics data additionally includes identified metabolites from fecal metabolomics. Optionally, the fecal metabolomics were analyzed by an LC-MS / MS analysis to identify metabolites. In an embodiment of the disclosure, the multomics data additionally includes identified metabolites from one or more of the following: fecal metabolomics, microbiome composition, demographics, a food frequency questionnaire, and clinical parameters.
[0015] In an embodiment of the disclosure, some of the identified metabolites indicate that the patient is a responder. Optionally, some of the identified metabolites indicate that the patient is a non-responder. Alternatively, or Additionally, some of the identified metabolites indicate that the patient is a responder and some of the identified metabolites indicate that the patient is a non-responder.
[0016] In an embodiment of the disclosure, the multiomics data includes identifying microbes. Optionally, the machine learning module uses a random forest algorithm to provide the prediction. In an embodiment of the disclosure, ranking of the identified metabolites are adjusted based on the food frequency questionnaire, demographics and / or clinical parameters.
[0017] There is further provided according to an embodiment of the disclosure, a method of predicting a response to nutritional therapy for treating Crohn’s Disease (CD), comprising: installing a machine learning module on a computer having a processor and memory; training a model with the machine learning module, by accepting a training data set comprising multiomics data of subjects with Crohn’s disease that were treated by exclusive enteral nutrition (EEN) and a determination if each subject was a responder that successfully responded to the treatment or was a non-responder that unsuccessfully responded to the treatment; and accepting multiomics data of a patient and predicting if the patient is a responder or a non-responder; wherein the multiomics data includes identified metabolites from serum metabolomics.
[0018] In an embodiment of the disclosure, a non-transitory computer readable medium storing a machine learning module program including instruction which, when executed by a processor, predict if a patient will be a responder or a non-responder according to the above method. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Some embodiments of the invention are herein described, by way of example only, with reference to the accompanying drawings. With specific reference now to the drawings in detail, it is stressed that the particulars shown are by way of example and for purposes of illustrative discussion of embodiments of the invention. In this regard, the description taken with the drawings makes apparent to those skilled in the art how embodiments of the invention may be practiced.
[0020] In the drawings:
[0021] Fig. 1 is a schematic illustration of a system for predicting a response to nutritional therapy for treating Crohn’s Disease (CD) in newly diagnosed children, according to an embodiment of the disclosure;
[0022] Figures 2A-2F are graphs of serum and fecal metabolomics data for predicting a response to exclusive enteral nutrition (EEN), according to an embodiment of the disclosure;
[0023] Figures 3A-3F are graphs of microbiome diversity and random forest (RF) analysis to predict response to exclusive enteral nutrition (EEN), according to an embodiment of the disclosure;
[0024] Figures 4A-4B are graphs of multiomics random forest (RF) analysis producing a highly accurate metabolomics-based model to predict response to exclusive enteral nutrition (EEN), according to an embodiment of the disclosure;
[0025] Figures 5A-5B are graphs of serum and fecal metabolomics data adjusted to covariates, according to an embodiment of the disclosure; and
[0026] Figures 6A-6B are graphs of Food-Frequency Questionnaire (FFQ) random forest (RF) analysis to predict response to exclusive enteral nutrition (EEN), according to an embodiment of the disclosure. DETAILED DESCRIPTION
[0027] Fig. 1 is a schematic illustration of a system 100 for predicting a response to nutritional therapy for treating Crohn’s Disease (CD) in newly diagnosed children, according to an embodiment of the disclosure. System 100 includes a computer 110 with a processor 114 and memory 116 for executing a machine learning model 120. Optionally, the computer 110 includes I / O components such as a display 112, a keyboard 118 and a pointing device 119 (e.g., a mouse).
[0028] In an embodiment of the disclosure, the machine learning module 120 is programmed to accept a training data set 105 comprising lists of patients and respective multiomics data and a determination if the patient was responsive to an Exclusive Enteral Nutrition (EEN) treatment program. The training data set 105 is used to train a model 122 for use by the machine learning module 120. Once trained the machine learning module 120 is capable of receiving query data 115 comprising multiomics data and to provide a response 125, which includes a determination if the patient is expected to successfully respond to EEN treatment (a responder) or unsuccessfully respond to EEN treatment (a non-responder).
[0029] In an embodiment of the disclosure, the model is based on identifying the most predictive metabolites (“top ranked metabolites”) from fecal metabolomics, serum metabolomics, microbiome composition, demographics, a Food-Frequency Questionnaire (FFQ)) and clinical parameters (e.g., including details of disease activity) all serving as multiomics data. Optionally, also microbes may be identified from the multiomics data. The top ranked metabolites may indicate that the patient is a responder or non-responder responsive to the patient having or lacking specific metabolites of a significant amount relative to other people. Optionally a machine learning random forest model is used to determine if the patient is to be considered a responder or a non-responder.
[0030] In some embodiments of the disclosure, only serum metabolomics or only fecal metabolomics or only microbiome are used to generate model 122. Alternatively, a combination of two of the three are used to generate model 122. Optionally, a combination including at least serum metabolomics is used to achieve the most accurate prediction. In some embodiments of the disclosure, model 122 may be instructed to use at least a selected set of metabolites for serum metabolomics or for fecal metabolomics or for microbiome. Optionally, the model 122 may use the selected set of metabolites for the respective metabolomics (or microbiome) and only identify metabolites for the others. For example, if provided that valeric acid, alpha aminobutyric acid and 2-hydroxyglutraic acid are predictive metabolites for serum metabolomics, these metabolites may be used in addition to identifying others from serum metabolomics or instead of identifying others from serum metabolomics, while generating other metabolites from fecal metabolomics or microbiome. Optionally, the selected set of metabolites are given a higher weight when identifying additional metabolites.
[0031] In an embodiment of the disclosure, patients commenced on EEN at CD onset were followed through 8 weeks. Stool was collected for microbiome and metabolomics, and serum for metabolomics. Disease indices and clinical data were recorded including a Food- Frequency Questionnaire (FFQ). A targeted quantitative metabolomics approach was applied to analyze fecal and serum samples using a combination of direct injection mass spectrometry with a reverse-phase LC-MS / MS custom assay. DNA extracted from stool was sequenced and 16S rRNA gene was amplified. Optionally, feature selection is performed using Recursive Feature Elimination with Cross-Validation (RFECV). For each omic analysis and the multiomics integration, selected components are fitted into a machine learning random forest (RF) model. Cross validation may include using stratified 3-fold or leave one out. Optionally, metabolite ranking may be adjusted responsive to the food frequency questionnaire (FFQ), demographics and / or clinical parameters.
[0032] In an exemplary embodiment of the disclosure, a study included 51 children (aged 14.8+2.7 yrs) treated with EEN, of whom 34 (66%) were responders; 35 (69%) had serum metabolomics, 19 (37%) fecal metabolomics and 31 (61%) microbiome data while 17 (33%) had all components. Clinical data and standard labs did not predict a response. The individual omics models ranked several key metabolites and microbes and were able to predict EEN response with an accuracy ranging at 0.806-0.850. Serum 3-carboxy-4- methyl-5-propyl-2-furanpropionic acid (CMPF), fecal phenylethylamine and fecal alphaaminobutyric acid were higher in non-responders while serum malic acid and fecal 3- methyladipic acid were higher in responders. Moreover, levels of Veillonela and Bifidobacterium species alongside Blautia genus were higher in responders while Lachnospiraceae and Fusicatenibacter genus were higher in non-responders. The final multiomics model which included also clinical data, retained components from serum and fecal metabolomics but none from the microbiome; it predicted EEN response with high accuracy of 0.882. Higher level of serum 2-hydroxyglutraic acid, 2-hydroxy-2- methylbutyric acid and fecal 3 -methyladipic acid were detected in responders. In non- responders there were higher levels of serum alpha- aminobutyric acid, valeric acid, isovaleric acid, alpha- aminobutyric acid and kynrenine and alongside with increased fecal kynrenine / tryptophan ratio, both part of the kynrenine tryptophan pathway, a metabolomics-based multiomics model 122 was generated for machine learning module 120 to predict a response to EEN.
[0033] METHODS
[0034] According to an embodiment of the disclosure, children (0-18 years) diagnosed with CD were prospectively enrolled since 2014 into a prospective cohort study which its overall goal was to identify predictors of response to treatments and disease course. At inclusion, patients underwent colonoscopy and small-bowel imaging (magnetic resonance enterography [MRE], ultrasonography, or capsule endoscopy). Bio- samples, including serum and stool, as well as explicit clinical, laboratory and demographic data, and medications were prospectively recorded on a REDCap electronic platform. Disease activity is quantified by standard indices of colonoscopic degree of inflammation (Simple Endoscopic Score in CD [SES-CD]), disease classifications (IBD-classes criteria and Paris classification), disease activity scores (weight pediatric CD activity index (wPCDAI) and the MINI score), as well as patient-reported measures including Food- Frequency Questionnaire (FFQ), and IMPACT-III quality of life scale. Children were prospectively followed and treatment success was assessed between 6-8 after initiating EEN intervention, while collecting extensive clinical, outcome and therapeutic data.
[0035] Sample handling Stool samples without additives were aliquoted, labeled and stored at -80°C for metabolomics analysis. A separated stool sample was collected in tubes containing 95% ethanol, centrifuged to pellet stool, then sub-aliquoted into 2-ml cryovials in ~100-200mg aliquots and frozen for future microbiome studies. For serum collection, blood was drawn into two 6-ml SST tubes and centrifuged at 1800Xg. Serum was then aliquoted and stored at -80°C until future analyses of metabolomics.
[0036] Data generation
[0037] Serum and fecal metabolomics were generated in an appropriate laboratory. Details of the analyses can be found in the Supplementary Methods below. Briefly, for metabolomics, a targeted quantitative approach was used to analyze the samples using a combination of direct injection mass spectrometry with a reverse-phase LC-MS / MS custom assay. This custom assay, in combination with an ABSciex 5500 QTrap mass spectrometer, was used for the targeted identification and quantification of 637 different endogenous metabolites. Similarly, bile acid quantification analysis was applied for stool samples, using UHPLC system for LC-MS / MS analysis with an AB Sciex QTrap 5500 mass spectrometer. For the microbiome, DNA was extracted from stool for libraries preparation, targeting the V4 region of 16S rRNA gene for bacterial workflow. Libraries were cleaned, quantified and equimolar amounts were pooled and sequenced using Illumina Miseq v2 Kit to generate paired-end reads. Finally, FFQs filled by the subjects were sent to pre-analysis at The S. Daniel Abraham International Center for Health and Nutrition, Ben-Gurion University in the Negev, Israel.
[0038] Data processing
[0039] Metabolomics: Metabolites which were under limit of detection (LOD) values in at least 75% of samples were excluded. The remaining metabolites with LOD values were imputed as half of LOD value of each metabolite. To make the data easier to work with and not lose important information, we conducted fatty acids grouping into saturated and unsaturated chains in respect to their length and type (see supplementary materials, Supp Table 1 below). Finally, to equalize signal intensities of metabolites concentrations, data was normalized to subjects’ median and scaled using auto scaling (z- score normalization). Microbiome: A custom R script was used to identify and remove sequencing primers. DADA2 R package was then employed to process raw sequences to normalized taxonomy tables, using a standard workflow as demonstrated in http s : / / benj j neb . ithub . io / dada2. DADA2 processing commences with quality filtration, setting for this project maxEE (maximal expected errors allowed) to 5, and truncating sequence length at 210. DADA2 default setting were applied for the following step: error learning (an estimation of error rate per sequencing run), dereplication, inference of unique Accurate Sequence Variants (AS Vs), paired-end merging and chimera removal. Taxonomic assignment was performed against Silva database v 138. To avoid bias related to sequence depth, data were rarified to an equal depth of 10000 seq / sample. Taxonomy features that accounted for more than 0.01% of all reads were included. Downstream analysis was conducted using R version 4.3.2, specifically phyloseq and vegan packages. Microbiome counts were normalized using total sum scaling and transformed using center log ratio.
[0040] Food-Frequency Questionnaire (FFQ): The analyzed data retrieved contained the frequency of consumption per day, alongside a consumption portion weight, of 110 food items, 131 estimated nutrient concentrations and total calories intake in each subject per day according to the FFQ report. These food items were grouped under 9 food categories (according to dietetics food groups): starch (sourced in bread / grains), starch (sourced in legumes), meat, dairy, low fat dairy, vegetables, fruits, fats, sugar / sweets and others. Food intakes were normalized by the formula:
[0041] Statistical Analysis
[0042] Metabolomics: Univariant analysis was applied for each metabolite, for normally distributed metabolites student T-test was calculated and for non-normal distribution mann- whitney U test was calculated (implemented in scipy python package). Next, analysis of covariance (ANCOVA, implemented in pingouin package in python) was applied with clinical parameters at baseline as co variates. Co variates included disease io activity markers (i.e. C-reactive protein-CRP, wPCDAI and SES-CD), disease location and behavior according to Paris classification (i.e. location: LI or L3 - no L2 subject were available, L4 or non-L4 - L4a and L4b were lumped due to small number of only L4b subjects, and behavior: Bl or B2 - only 1 subject with B2B3 phenotype), and additional demographic characteristics (age, BMI -Body mass index and sex). For all statistical steps multiple test corrections were applied using Benjamini Hochberg false discovery rate (FDR, implemented in statsmodels package in python).
[0043] Microbiome: Bacterial a-diversity, which quantifies the intra-sample diversity, i.e., the distribution of species abundances in a given sample, was estimated using chaol richness and Shannon diversity index, statistical tests used to compare between groups was Wilcoxon test (R phyloseq package). 0-diversity, which measures dis- similarities between samples, was calculated using the Bray-Curtis dissimilarity index and statistical test that was used included Kruskal Wallis test and PERMANOVA (Permutational Multivariate Analysis of Variance) test implemented in R vegan package.
[0044] FFQ: Univariant analysis was applied for each food category and for each nutrient, for normally distributed metabolites student T-test was calculated and for non-normal distribution Mann- Whitney U test was calculated (implemented in scipy python package). For all statistical steps multiple test correction was applied using Benjamini Hochberg false discovery rate (FDR, implemented in statsmodels package in python).
[0045] Machine Eearning analysis: Feature selection for each ‘omic was applied using RFECV function implemented in scikit-learn package in Python. The algorithm combines the concepts of recursive feature elimination (RFE) and cross-validation (CV) to find the optimal set of features for a machine learning classifier / model. RFECV combines RFE with CV using a classifier, chosen here to be random forest (RF), as previously described by Chen et al (PMID: 34683469) and Mossotto et al (PMID: 35982195). It performs RFE while considering the performance of the random forest model through cross-validation at each step. The algorithm starts with all features, fits the RF model, and evaluates its performance using CV and selected evaluation metric. The least important feature is then removed, and the process is repeated until the optimal evaluation metric over CV iterations is reached, yielding a set of selected features. To address any imbalance in our cohort, weighted-Fl evaluation metric was selected; the Fl -score is a composite metric that harmonizes precision (the accuracy of positive predictions) and recall (the ability to capture all relevant positive instances). For serum metabolomics and microbiome we used stratified 3-fold CV and Leave One Out CV (LOOCV) for fecal metabolomics and the integrated multiomics analysis to account for the small sample size. Finally, a random forest classifier (250 decision trees, other parameter were set to default) was fitted to serum metabolomics, fecal metabolomics, microbiome, FFQ and the integrated multiomics analysis utilizing selected features generated by RFECV. Additional metrics were also calculated, including area under the ROC curve (AUROC), accuracy, sensitivity, specificity, as well as positive and negative predictive values (PPV and NPV). 95% Confidence intervals were calculated using 10-repeated cross validation.
[0046] RESULTS
[0047] A total of 51 children met the eligibility criteria, of whom 34 (66%) were responders (Table 1). Serum metabolomics were available for 35 (69%) children (containing 178 metabolites after data processing, as described above), fecal metabolomics for 19 (37%) children (containing 213 metabolites after data processing) and fecal microbiome for 31 (61%) children (containing 330 microbial taxonomies); 17 (33%) children had all three components.
[0048] None of the clinical and demographic variables recorded at baseline predicted response to nutrition treatment, including age, body mass index (BMI), sex, C-reactive protein (CRP), wPCDAI, SES-CD, disease location and disease behavior (Table 1).
[0049] Serum and fecal Metabolomics
[0050] Several serum metabolites were differentially abundant between responders and nonresponders (Figure 2A), four were found to be more abundant in non-responders including 3-carboxy-4-methyl-5-propyl-2-furanpropionic acid (CMPF), 3 -methyladipic acid, unsaturated very long chain fatty acids cholesterol esters and picolinic acid (See Supplementary Table 1). Twelve others were less abundant in non-responders including malic acid, 5-hydroxy lysine, putrescine, fumaric acid, lactic acid, phenylalanine, 2- hydroxy glutaric acid, pyruvic acid, 3 -methyladipic acid, arginine, malonic acid, 4- hydroxyphenylpyruvic acid and caproic acid (Supplementary Table 2). When using ANCOVA adjusted for disease activity using covariates wPCDAI, SES-CD, CRP, for disease location and behavior using covariates L1 / L3, non-L4 / L4, B1 / B2 and for other demographic characteristics, age, sex and BMI, there were slight changes in the highlighted serum metabolites (Supplementary Table 1). One additional metabolite was found to be more abundant in non-responders and seven others were found to be less abundant, three metabolites from this group lost their significance due to this adjustment (figure 5A). None of the serum metabolites reached statistical significance after FDR correction for multiple comparisons, whether adjusted to clinical and demographic variables or not (Supplementary Table 1 and 2).
[0051] Several fecal metabolites were differentially abundant between responders and non- responders (Figure 2B), two were found to be more abundant in non-responders including alpha-aminobutyric acid and phenylethylamine. Four others were less abundant in non-responders including 3-methylipidic acid, benzoic acid, gamma-aminobutyric acid and orotic acid (Supplementary Table 3). When using ANCOVA adjusted for disease activity, for disease location and behavior for other demographic characteristics there were major changes in the highlighted fecal metabolites (Supplementary Table 4). All fecal metabolites prior to the adjustment lost their significance and only one metabolite were found to be more abundant in non-responders while two metabolites were found to be less abundant in non-responders (Figure 2B). None of the fecal metabolites reached statistical significance after FDR correction for multiple comparisons, whether adjusted to clinical and demographic variables or not (Supplementary Table 3 and 4).
[0052] The traditional statistics applied above relies on fixed rules, while machine learning analysis adapts, learns and able to recognize patterns directly from data, offering a dynamic approach to understand and predict outcomes in diverse datasets. Machine learning random forest model was generated utilizing selected features by recursive feature elimination and cross-validation (RFECV) pipeline (see methods), predicted EEN response from serum / fecal metabolomics yielding AUROC 0.87 / 0.82 (Figure 2C-2D, 95% CI [0.86-0.87] / [0.80-0.82]), accuracy 0.84 / 0.4 (95% CI [0.82-0.85] / [0.82-0.86]), specificity 0.81 / 0.58 (95% CI [0.78-0.84] / [0.53-0.62]), sensitivity 0.85 / 0.94 (95% CI [0.84-0.87] / [0.92-0.95]), NPV 0.77 / 0.77 (95% CI [0.75-0.79] / [0.70-0.83]), and PPV 0.88 / 0.86 (95% CI [0.87-0.90] / [0.85-0.87]), respectively (Table 2). Top influencing features on the model’s accuracy in the serum were malic acid, CMPF and putrescine and, in the stool, 3 -methyladipic acid, alpha- aminobutyric acid, and Taurodeoxycholic acid (TDCA) (Figure 2E-2F). These discoveries match with the serum metabolites we identified in the previous traditional analyses. The machine learning analysis random forest results do not just include statistical significant metabolites but also included some metabolites that didn't show strong statistical differences. The stool metabolites chosen for the model were linked to important metabolites from both the disease activity, location-adjusted simpler analysis and the original analysis without adjustments.
[0053] Microbiome
[0054] Higher Shannon index diversity (a diversity) was observed in non-responders (2.5 [IQR 2.3-2.7]) compared with responders (2.3 [IQR 1.9-2.4], p=0.02). Chaol richness index was similar in trend between responders and non-responders (32 [IQR 30-36]) than responders (30 [IQR 26-35], p=0.25) (Figure 3A-3B) as well as community composition (P diversity) (Figure 3C, PERMANOVA test, Fi, 29=1.1, p=0.32). Machine learning random forest analysis produced a model to differentiate responders and non-responders with AUROC 0.92 (95% CI [0.91-0.92], Figure 3D), accuracy 0.85 (95% CI [0.82- 0.87]), sensitivity 0.85 (95% CI [0.83-0.87]), specificity 0.84 (95% CI [0.79-0.88]), PPV 0.94 (95% CI [0.92-0.95]), and NPV 0.66 (95% CI [0.62-0.70]) (Table 2). The bacterial taxonomies that contributed most for the random forest model’s accuracy, were assigned to genera Vellionella and Bifidobacterium with higher abundance in responders. Higher abundance in non-responders were assigned to order Bacteroidales (which include the genera Prevotella and Bacteroides) and genus Lachnospiraceae ND3007 group (Figure 3E).
[0055] FFQ Only a trend was observed in meat intake (responders: 28.7 [IQR: 13.3-37.3], nonresponders: 63.3 [IQR: 27.5-76.5], p=0.07), dairy intake (responders: 115.3 [IQR:67.9- 357.6], non-responders: 35.7[IQR: 16.4-73.4] , p=0.09) and sugar / sweets intake (responders: 23.5 [IQR: 15.3-58.4], non-responders: 62.7 [IQR: 42.8-99.2], p=0.14). FFQ analysis also measured estimated 131 nutrient concentrations according to the reported consumed food per day. Prior FDR correction several nutrient concentrations were different between responders and non-responders including docosahexanoic acid (responders: 0.06 [IQR: 0.03-0.08], non-responders: 0.1 [IQR: 0.08-0.15], p=0.02), docosapentaenoic acid (responders: 0.007 [IQR: 0.005-0.01], non-responders: 0.014 [IQR: 0.10-0.20], p=0.04), erucic acid (responders: 0.02 [IQR: 0.01-0.02], non- responders: 0.03 [IQR: 0.02-0.04], p=0.04) and eicosapentaenoic acid (responders: 0.02 [IQR: 0.004-0.03], non-responders: 0.03 [IQR: 0.03-0.05], p=0.04) (Supplementary Table 5 and 6). A machine learning random forest analysis, with feature selected from both food group intake and nutrients (Figure 6A), was applied, yielding a predictive model yet with very low specificity to non-responders (negative) cases. The features included were docosahexanoic acid, erucic acid and meat intake which were higher in non-responders and dairy intake, fruits intake and tocopherol delta nutrient, higher in responders. Evaluation metrics: AUROC 0.79 (95% CI [0.0.77-0.79], Figure 6B), accuracy 0.79 (95% CI [0.77-0.81]), sensitivity 0.96 (95% CI [0.94-0.99]), specificity 0.29 (95% CI [0.29-0.29]), PPV 0.80 (95% CI [0.80-0.80]), and NPV 0.77 (95% CI [0.62-0.92]) (Table 2).
[0056] Multiomics
[0057] Finally, multicomponent model was tested, by using early integrating approach [PMID: 34285775], thus a simple concatenation of features (i.e. metabolites, microbes and etc.) across all previously described omics. We included only the 17 patients who had all ‘omics data available. The early integrated multiomics data contained 599 features, including 178 serum metabolites, 213 stool metabolites, 201 microbial taxonomies and 8 clinical and demographic characteristics as described above. These features were introduced to the RFECV pipeline which retained 6 features,(l%) of the features as predictors of EEN response. The final machine learning random forest model included one fecal metabolome and five serum metabolomes, while, in agreement with the univariate analysis, none of the microbiome components nor clinical records contributed to the model. The model was highly accurate to predict response to nutritional treatment with a AUROC 0.98 (95% CI [0.97-0.99]), accuracy 0.95 (95% CI [0.92-0.98]), sensitivity 0.95 (95% CI [0.3-0.98]), specificity 0.92 (95% CI [0.84-1]), PPV 0.98 (95%
[0058] CI [0.95-1]), and NPV 0.87 (95% CI [0.78-0.95]) (Table 2). The top contributing features to the model were serum 2-hydroxyglutaric acid, malic acid, alpha-ketoglutaric acid, and 2-hydroxy-2-methylbutyric acid as well as fecal 3-methyladipic acid, all higher in responders and serum 3 -Methyladipic acid which was higher in non-responders (Figure 4A). Reassuringly, most of the features in the multiomics model were also included in the individual omics random forest models (i.e. serum malic acid, serum 2-hydroxyglutaric acid, serum and fecal 3-methyladipic acid).
[0059] Details relating to the figures
[0060] Figure 2A-2F. Serum and fecal metabolomics data to predict EEN response.
[0061] Footnote: (2A-2B) Volcano plots of (2A) serum and (2B) fecal metabolites, top right corner are metabolites that are more abndant in non-responders / less abundant in responders to EEN treatment. (2C-2D) Area under the ROC curves indicates machine learning random forest model performance for (2C) serum metabolomics and (2D) fecal metabolomics. (2E-2F) Features were selected previously by the RFECV algorithm, features importance, which indicates how each feature influence on the decrease in mean accuracy is shown for each model (2E) serum or (2F) fecal metabolomics. Light green - higher median values in responders, light red- higher median values in non-responders. RFECV- Recursive Feature Elimination with Cross-Validation, EEN - exclusive enteral nutrition, RF - random forest, AUC - area under the receiver operating characteristic (ROC) curve.
[0062] Figure 3A-3E. Microbiome diversity and random forest (RF) analysis to predict response to exclusive enteral nutrition (EEN).
[0063] Footnote: (3A-3B) Microbiome alpha-diversity, (3A) Shannon diversity index and (3B) chaol richness index shows bacterial taxonomical genus diversity and richness for EEN responders and non-responders, (3C) Microbiome beta-diversity calculated using Bray- Curtis distance, (3C) Principal Coordinate Analysis (PCoA) plot of EEN responders and non-responders microbiome community at baseline. (3D) Features were selected previously by the RFECV algorithm, features importance, which indicates how each feature influence on the decrease in mean accuracy is shown for the microbiome-based machine learning random forest model, (3E) Area under the ROC curve (AUROC) indicates machine learning random forest model performance using microbiome taxonomy (Phylum-Genus) for EEN response prediction. Light bars indicate -higher median values in responders, dark bars indicate - higher median values in non- responders. RFECV- Recursive Feature Elimination with Cross-Validation
[0064] Figure 4A-4B. Multiomics random forest (RF) analysis produce highly accurate metabolomics-based model to predict response to exclusive enteral nutrition (EEN). Footnote: Multicomponent RF model was based on demographics, disease activity data (i.e. CRP, wPCDAI, SESCD), serum metabolomics, fecal metabolomics and microbiome composition. (4A) Features were selected previously by the RFECV algorithm, features importance, which indicates how each feature influences the decrease in mean accuracy is shown for the microbiome-based machine learning random forest model (4B) AUC- ROC curve indicates model performance for the resulting metabolomics-based model. Light bars indicate -higher median values in responders, dark bars indicate- higher median values in non-responders. CRP - C-reactive protein, wPCDAI- weighted pediatric Crohn’s disease activity index, SESCD - Simple endoscopic score in Crohn’s disease, RFECV- Recursive Feature Elimination with Cross-Validation, AUC - area under the curve, ROC curve - receiver operating characteristic curve.
[0065] Figure 5A-5B. Serum and fecal metabolomics data adjusted to covariates.
[0066] Footnote: (5A-5B) Volcano plots of (5A) serum and (5B) fecal metabolites, top right corner are metabolites that are more abundant in non-responders / less abundant in responders to EEN treatment. Adjusted for covariates including CRP, wPCDAI, SESCD, age, sex, BMI, disease location (L1 / L3 and L4 / Non-L4) and disease behavior (B1 / B2). CRP - C-reactive protein, wPCDAI- weighted pediatric Crohn’s disease activity index, SESCD - Simple endoscopic score in Crohn’s disease, BMI -Body mass index. Disease behavior and location according Paris classification: Bl - Non- structuring Nonpenetrating, B2- Structuring, LI - Distal 1 / 3 ileal ± limited cecal disease, L3 - Ileocolonic, L4 - Upper disease proximal / distal to ligament of Treitz.
[0067] Figure 6-6B. Food-Frequency Questionnaire (FFQ) random forest (RF) analysis to predict response to exclusive enteral nutrition (EEN).
[0068] Footnote: FFQ model was based on reported food consumption and estimated nutrient abundance. (6A) Features were selected previously by the RFECV algorithm, features importance, which indicates how each feature influence on the decrease in mean accuracy is shown for the FFQ machine learning random forest model (6B) AUC-ROC curve indicates model performance. Light bars indicate -higher median values in responders, dark bars indicate - higher median values in non-responders. RFECV- Recursive Feature Elimination with Cross-Validation, AUC - area under the curve, ROC curve - receiver operating characteristic curve.
[0069] TABLES
[0070] Table 1. Baseline characteristics of the included cohort (count (%), mean± SD, or medians (IQR) are indicated as appropriate)
[0071] Footnote: Columns represent baseline clinical characteristics for Non-Responders and Responders to EEN diet therapy. P-values were calculated using one-way t-test for normal distributed characteristics and Mann- Whitney u-test for non-normal distributed. EEN- exclusive enteral nutrition, wPCDAI- weight pediatric Crohn’s disease activity index, CRP- C-reactive protein, SESCD- Simple Endoscopic Score in CD, BMI- Body mass index. Disease behavior and location according Paris classification: Bl - Non- stricturing Non-penetrating, B2- Stricturing, B2B3 - Both penetrating and stricturing, LI
[0072] - Distal 1 / 3 ileal ± limited cecal disease, L3 - Ileocolonic, L4 - Upper disease proximal / distal to legament of Treitz.AIncluded in clinical data that was integrated with all other omics for random forest classification model. Table 2. Evaluation metrics for random forest models of the individual omics and the multiomic
[0073] Footnote: Each column represents evaluation metrics for a machine learning random forest model created using selected components by Recursive Feature Elimination with Cross-Validation algorithm. ROC- Receiver operating characteristic. Clinical data included in the multiomics model - Age, Sex, BMI, wPCDAI, SESCD, Disease location (LI or L3, L4 or non-L4), Disease behavior (Bl or B2) and CRP. Serum metabolomics include - Malic acid, CMPF (3-carboxy-4-methyl-5-propyl-2-furanpropionic acid), Putrescine, 3 -Methyladipic acid, Malonic acid, Fumaric acid and 2-Hydroxyglutaric acid. Fecal metabolites include -3 -Methyladipic acid, Alpha-aminobutyric acid, TDCA (Taurodeoxycholic acid), 4-Hydroxyphenylacetic acid, Gamma-aminobutyric acid, Benzoic acid, Betaine and Uric acid. Microbiome components included - Family Bifidobacteriaceae, Genus Veillonella, Genus Bifidobacterium, Order Bacteroidales,
[0074] Genus Lachnospiraceae ND3007 group, , Genus Lachnospiraceae UCG-010, Genus Erysipelotrichaceae UCG-003, Family Fusobacteriaceae, Order Fusobacteriales. FFQ components included - docosahexanoic acid, Meat intake, Erucic acid, Dairy intake, Tocopherol delta and Fruit intake. Multiomics early integrated model included - Serum 2-Hydroxyglutaric acid, Serum Malic acid, Serum Alpha-aminobutyric acid, Serum 2- Hydroxy 2-methylbutyric acid, Fecal 3 -Methyladipic acid, Serum 3 -Methyladipic acid and Serum Fumaric acid. All features are sorted from the most contributing (highest impact on mean decrease in accuracy) to the model to the least.
[0075] SUPPLEMENTARY METHODS
[0076] Metabolomics
[0077] TMIC MEGA Assay DI / LC-MS / MS Method
[0078] A targeted quantitative metabolomics approach was applied to analyze the samples using a combination of direct injection mass spectrometry with a reverse-phase LC-MS / MS custom assay. This custom assay, in combination with an ABSciex 5500 QTrap (Applied Biosystems / MDS Sciex) mass spectrometer, was used for the targeted identification and quantification of upto 600 different endogenous metabolites including amino acids or amino acid related, biogenic amines, ceramides, cholesterol esters, diacylglycerols, acylcarnitines, glycerophospholipids, sphingomyelins, triacylglycerols, organic acids and nucleotide / nucleosides. The method combines the derivatization and extraction of analytes, and the selective mass-spectrometric detection using multiple reaction monitoring (MRM) pairs. Isotope-labeled internal standards and other internal standards are used for metabolite quantification. The custom assay contains a 96 deep-well plate with a filter plate attached with sealing tape, and reagents and solvents used to prepare the plate assay. First 14 wells are used for one blank, three zero samples, seven standards and three quality control samples. For all metabolites except organic acid, samples were thawed on ice, vortexed and centrifuged at 13,000x g. Sample were loaded onto the center of the filter on the upper 96-well plate and dried in a stream of nitrogen. Subsequently, phenyl-isothiocyanate was added for derivatization. After incubation, the filter spots will be dried again using an evaporator. Extraction of the metabolites was achieved by adding 300 pL of extraction solvent. The extracts were obtained by centrifugation into the lower 96-deep well plate, followed by a dilution step with MS running solvent.
[0079] For organic acid analysis, 150 pL of ice-cold methanol and 10 pL of isotope-labeled internal standard mixture will be added to the samples for overnight protein precipitation. Then it was centrifuged at 13000x g for 20 min. 50 pL of supernatant was loaded into the center of wells of a 96-deep well plate, followed by the addition of 3- nitrophenylhydrazine (NPH) reagent. After incubation for 2h, BHT stabilizer and water were added before LC-MS injection.
[0080] Mass spectrometric analysis was performed on an ABSciex 5500 Qtrap® tandem mass spectrometry instrument (Applied Biosystems / MDS Analytical Technologies, Foster City, CA) equipped with an Agilent 1290 series UHPLC system (Agilent Technologies, Palo Alto, CA). The samples were delivered to the mass spectrometer by a LC method followed by a direct injection (DI) method. Data analysis was done using Analyst 1.6.2.
[0081] Sample Preparation for Bile Acid Analysis
[0082] Samples were thawed on ice, in dark, before use. For bile acid analysis, to a 96-well filter plate, 10 pL of the internal standard (I STD) mixture solution was added to each well except Well Al which was used as double blank. The whole plate was then dried under nitrogen for 5 min. After that, 10 pL of the samples (PBS as blank sample, calibration standards, QC standards and samples) were pipetted directly onto the center of each corresponding spot in the 96-well plate and again dried under nitrogen for 20 min, followed by the addition of 100 pL of methanol to each well. The whole plate was then covered and shaken at 600 rpm for 20 min, and spun in a centrifuge for 2 min at 500xg. To each well of the bottom collection plate, 60 pL of water was added and mixed thoroughly, and then 5 pL was injected into an UHPLC-equipped QTrap 5500 mass spectrometer for LC-MS / MS analysis.
[0083] LC-MS / MS Method for Bile Acid Analysis
[0084] An Agilent 1290 series UHPLC system (Agilent, Palo Alto, CA) was used for LC- MS / MS analysis with an AB Sciex QTrap 5500 mass spectrometer (Sciex Canada, Concord, ON). The controlling software for the LC-MS system was Analyst 1.6.3. For the UHPLC work, solvent A was prepared by mixing 25 mL of 200 mM ammonium acetate aqueous solution, 475 mL of water and 75 pL of formic acid; and solvent B was prepared by mixing 25 mL of 200 mM ammonium acetate aqueous solution, 150 mL of methanol, 325 mL of acetonitrile and 75 mL of formic acid. The gradient profile for the UHPLC solvent run was set as follows: t = 0 min, 35% B, flow rate = 0.5 mL / min; t = 0.25 min, 35% B, flow rate = 0.5 mL / min; t = 0.35 min, 40% B, flow rate = 0.5 mL / min; t = 1.90 min, 45% B, flow rate = 0.8 mL / min; t = 2.10 min, 55% B, flow rate = 0.8 mL / min; t = 3.30 min, 65% B, flow rate = 1.0 mL / min; t = 3.50 min, 100% B, flow rate = 1.0 mL / min; t = 4.00 min, 100% B, flow rate = 1.0 mL / min; t = 4.10 min, 35% B, flow rate = 0.9 mL / min; and t = 5.0 min, 35% B, flow rate = 0.5 mL / min. The sample injection volume was 5 pL. The mass spectrometer was set to a negative electrospray ionization mode with multiple reaction monitoring (MRM). The lonSpray voltage was set at -4500 V and the temperature at 600°C. The curtain gas (CUR), ion source gas 1 (GAS1), ion source gas 2 (GAS2) and collision gas (CAD) were set at 20, 40, 50 and medium, respectively. The entrance potential (EP) was set at -10 V. Likewise, the declustering potential (DP), collision energy (CE), collision cell exit potential (CXP), MRM QI and Q3 were set individually for each analyte and ISTD.
[0085] Microbiome
[0086] DNA extraction and sequencing and processing of raw data. DNA extraction and deep amplicon sequencing of the bacterial 16S rRNA gene (V4 region) were performed at HyLabs (Rehovot, Israel), Next- Generation Sequencing (NGS) unit. Briefly, DNA was extracted using Magcore Nucleic Acid Extractor (RBC©) and the MagCore DNA Tissue kit according to the manufacturer’s instructions, with the following adaptations: 10-50 mg of sample were transferred to a tube containing glass lysis beads (Bead Beating Tube Type C [Soil]; Geneaid); the exact amount taken was recorded to allow absolute DNA quantification. Buffer GT from the Magcore DNA Tissue kit was added to the tubes, which were then placed in a Bead Beater (BioSpec, USA) for two minutes. After removal of lysis glass beads and cell debris by centrifugation, the lysate was incubated with Proteinase K, and further processed according to manufacturer’s instructions. Libraries were prepared using a two-step PCR protocol, with the first step utilizing primers tailed with CS1 / CS2 common sequences and targeting either the V4 region of 16S rRNA gene (515F: 5’-GTGCCAGCMGCCGCGGTAA , 806R: 5'-GGACTACHVGGGTWTCTAAT, 25 cycles) for bacterial workflow. Access Array primers for Illumina (Fluidigm) were used in a second, 10 cycle PCR to add barcode, adaptor, and index sequences to each sample. Libraries were cleaned using Kapa Pure beads (Kapa), quantified by Qubit, and equimolar amounts were pooled and sequenced using Illumina Miseq v2 Kit (500 cycles) to generate 2x250 paired-end reads.
[0087] Supplementary Table 1. Analysis of variance (ANOVA) for serum metabolites (mean± SD, or medians (IQR) are indicated as appropriate)
[0088] Median / Mean IQR / STD
[0089] Median / Mean IQR / STD
[0090] ANOVA p-val FDR NonNon- Responders Responders responders responders
[0091] Malic acid 0.002 0.22 2.10 (1.8-2.8) 1.70 (1.6-1.9) 5-Hydroxylysine 0.002 0.22 0.64 ±0.15 0.55 ±0.11 Putrescine 0.004 0.22 0.15 ±0.06 0.11 ±0.046
[0092] CMPF 0.005 0.22 0.28 (0.01-0.98) 1.46 (0.91-1.71)
[0093] Fumaric acid 0.01 0.30 0.48 (0.39-0.62) 0.40 (0.37-0.46)
[0094] Lactic acid 0.01 0.30 1622.00 (1450-2106) 1309.00 (1224-1462)
[0095] Phenylalanine 0.01 0.30 51.01 ±9.6 44.84 ±6.6
[0096] 2-hydroxyglutaric acid 0.01 0.30 0.35 (0.32-0.44) 0.27 (0.25-0.36)
[0097] Pyruvic acid 0.02 0.30 48.29 (38-64) 39.00 (34-43)
[0098] 3-Methyladipic acid 0.02 0.30 0.01 (0.009-0.02) 0.03 (0.021-0.035)
[0099] VLCFA unsaturated CE 0.02 0.30 21.39 (14.8-45.3) 49.23 (29-65) Arginine 0.02 0.30 101 ±19 92 ±14 Malonic acid* 0.04 0.49 0.05 (0.036-0.060) 0.08 (0.06-0.09) 4-Hydroxyphenylpyruvic
[0100] 0.05 0.54 2.89 ±0.88 2.46 ±0.75 acid* Caproic acid* 0.05 0.54 0.78 (0.65-1.01) 1.00 (0.93-1.34) Picolinic acid 0.05 0.55 0.04 (0.036-0.13) 0.16 (0.092-0.32)
[0101] Footnote ANOVA p-values and FDR (False discovery rate) correction between responders and non-responders to EEN (exclusive enteral nutrition) therapy. IQR -
[0102] Interquartile range, STD- Standard deviation. Astrick(*) represents metabolite that did not maintained significance after adjusting for covariates
[0103] Supplementary Table 2. Analysis of covariance (ANCOVA) for serum metabolites
[0104] (mean± SD, or medians (IQR) are indicated as appropriate)
[0105] Median Median
[0106] IQR / STD IQR / STD Non
[0107] ANCOVA p-val FDR / Mean / Mean NonResponders responders Responders responders
[0108] Malic acid 0.0008 0.14 2.10 (1.8-2.8) 1.70 (1.6-1.9) Fumaric acid 0.002 0.19 0.48 (0.39-0.62) 0.40 (0.37-0.46) 5-Hydroxylysine 0.004 0.19 0.64 ±0.15 0.55 ±0.11 Putrescine 0.004 0.19 0.15 ±0.06 0.11 ±0.046 2-hydroxyglutaric acid 0.006 0.22 0.35 (0.32-0.44) 0.27 (0.25-0.36) Hippuric acid* 0.01 0.26 0.016 (0.016-0.016) 0.40 (0.016-1.34) VLCFA unsaturated CE 0.01 0.26 21.39 (14.8-45.3) 49.23 (29-65) 3-Methyladipic acid 0.01 0.26 0.01 (0.009-0.02) 0.027 (0.021-0.035) Lactic acid 0.02 0.29 1622.00 (1450-2106) 1309.00 (1224-1462) Pyruvic acid 0.02 0.29 48.29 (38-64) 39.00 (34-43) Phenylalanine 0.02 0.31 51.01 ±9.6 44.84 ±6.6 alpha-Ketoglutaric acid 0.02 0.32 6.17 ±1.91 5.62 ±1.82 N-Acetyl-Alanine* 0.02 0.32 1.20 ±0.2 1.15 ±0.26 CMPF 0.03 0.34 0.28 (0.01-0.98) 1.46 (0.91-1.71) Arginine 0.03 0.34 101 19 92 14 Picolinic acid 0.03 0.34 0.04 (0.036-0.13) 0.16 (0.092-0.32) Alanine* 0.04 0.37 249 (207-351) 238 (191-262) N-Acetyl-Serine * 0.04 0.37 1.08 (0.98-1.21) 1.09 (0.90-1.16) Ethanolamine* 0.04 0.39 7.42 (6.66-9.05) 6.64 (5.81-8.00) 2-Hydroxy-2-
[0109] 0.05 0.39 0.27 ±0.051 0.26 ±0.035 methylbutyric acid* Beta- Alanine* 0.05 0.39 0.47
[0110] Footnote ANCOVA p-values and FDR (False discovery rate) correction between responders and non-responders to EEN (exclusive enteral nutrition) therapy. Covariants included here are: Age, Sex, BMI - Body mass index, wPCDAI- weight pediatric Crohn’s disease activity index, CRP- C-reactive protein, SESCD- Simple Endoscopic Score in CD, Disease location (according to Paris classification, L 1 / L3 and L4 / non-L4) and Disease behavior (according to Paris classification, B1 / B2). LI - Distal 1 / 3 ileal ± limited cecal disease, L3 - Ileocolonic, L4 - Upper disease proximal / distal to legament of Treitz, Bl - Non-stricturing Non-penetrating, B2- Stricturing. IQR - Interquartile range, STD- Standard deviation. Astrick(*) represents metabolite that maintained significance only after adjusting for covariates.
[0111] Supplementary Table 3. Analysis of variance (ANOVA) for fecal metabolites (mean±
[0112] SD, or medians (IQR) are indicated as appropriate)
[0113] Median / Mean IQR / STD
[0114] ANOVA P- FDR ^^c(^an^^canIQR / STD Non- Nonval Responders Responders responders responders
[0115] 3-Methyladipic 0.25 (0.22-0.30) 0.008 acid*
[0116] Benzoic acid* 0.02 1 0.43 (0.37-1.4) 0.02 (0.022-0.4
[0117] Alpha aminobutyric 0.03 1 152.00 (70-296) 411.00 acid*
[0118] Gamma- aminobutyric 0.03 1 19.4 (14.1-155.5) 6.85 (6.2-7.3) acid*
[0119] Orotic acid* 0.03 1 2.27 (1.06-10.47) 0.09 (0.09-0.56)
[0120] Phenylethylamine* 0.04 1 1.45 (0.17-5.05) 7.57 (7.41-9.26)
[0121] Footnote. ANOVA p-values and FDR (False discovery rate) correction between responders and non-responders to EEN (exclusive enteral nutrition) therapy. IQR - Interquartile range, STD- Standard deviation. Astrick(*) represents metabolite that did not maintained significance after adjusting for covariates
[0122] Supplementary Table 4. Analysis of covariance (ANCOVA) for fecal metabolites (mean±
[0123] SD, or medians (IQR) are indicated as appropriate)
[0124] Median / Mean IQR / STD Median / Mean IQR / STD Non-
[0125] ANCOVA p-val FDR Responders Responders Non-responders responders
[0126] N6-Acetyl-
[0127] Lysine* 0.007 0.91 41 (0.08-61.3) 0.08 (0.08-33)
[0128] N-Acetyl-
[0129] Leucine* 0.01 0.91 73 (26-104) 121.2 (56-150)
[0130] N-Acetyl-
[0131] Alanine* 0.05 0.91 96 (77-124) 202 (96-211)
[0132] Footnote-. ANCOVA p-values and FDR (False discovery rate) correction between responders and non-responders to EEN (exclusive enteral nutrition) therapy. Covariants included here are: Age, Sex, BMI - Body mass index, wPCDAI- weight pediatric Crohn’s disease activity index, CRP- C-reactive protein, SESCD- Simple Endoscopic Score in CD, Disease location (according paris classifiaction, L1 / L3 and L4 / non-L4) and Disease behavior (according to Paris classification, B 1 / B2). LI - Distal 1 / 3 ileal ± limited cecal disease, L3 - Ileocolonic, L4 - Upper disease proximal / distal to legament of Treitz, B l - Non-stricturing Non-penetrating, B2- Stricturing. IQR - Interquartile range, STD- Standard deviation. Astrick(*) represents metabolite that maintained significance only after adjusting for covariates.
[0133] Supplementary Table 5. Analysis of variance (ANOVA) for food frequency questionnaire
[0134] (FFQ) food categories, normalized to energy intake (mean± SD, or medians (IQR) are indicated as appropriate)
[0135] Median / Mean IQR / STD
[0136] Median / Mean IQR / STD
[0137] ANOVA p-val FDR Non- Non- Responders Responders responders responders
[0138] Meat 0.07 0.4 28.7 (13.3. 37.3) 63.3 (27.5, 76.5)
[0139] Dairy 0.09 0.4 115.3 (67.9, 357.6) 35.7 (16.4, 73.4)
[0140] Sugar / sweets 0.14 0.42 23.5 (15.3. 58.4) 62.7 (42.8, 99.2)
[0141] Supplementary Table 6. Analysis of variance (ANOVA) for food frequency questionnaire
[0142] (FFQ) analyized nutreints, normalized to energy intake (mean± SD, or medians (IQR) are indicated as appropriate)
[0143] Median / Mean IQR / STD
[0144] Median / Mean IQR / STD
[0145] ANOVA p-val FDR Non- Non- Responders Responders responders responders docosahexanoic 0.02 0.95 0.06 (0.03, 0.08) 0.10 (0.08, 0.15) docosapentaenoic 0.04 0.95 0.007 (0.005, 0.01) 0.014 (0.010, 0.02) erucic 0.04 0.95 0.02 (0.01, 0.02) 0.03 (0.02, 0.04) eicosapentaenoic 0.04 0.95 0.02 (0.004, 0.03) 0.03 (0.03, 0.05) The terms “comprises”, “comprising”, “includes”, “including”, “has”, “having” and their conjugates mean “including but not limited to”.
[0146] The term “consisting of’ means “including and limited to”.
[0147] The term “consisting essentially of’ means that the composition, method or structure may include additional ingredients, steps and / or parts, but only if the additional ingredients, steps and / or parts do not materially alter the basic and novel characteristics of the claimed composition, method or structure.
[0148] As used herein, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a compound” or “at least one compound” may include a plurality of compounds, including mixtures thereof.
[0149] Throughout this application, embodiments of this disclosure may be presented with reference to a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the disclosure. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as “from 1 to 6” should be considered to have specifically disclosed subranges such as “from 1 to 3”, “from 1 to 4”, “from 1 to 5”, “from 2 to 4”, “from 2 to 6”, “from 3 to 6”, etc.; as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0150] Whenever a numerical range is indicated herein (for example “10-15”, “10 to 15”, or any pair of numbers linked by these another such range indication), it is meant to include any number (fractional or integral) within the indicated range limits, including the range limits, unless the context clearly dictates otherwise. The phrases “range / ranging / ranges between” a first indicate number and a second indicate number and “range / ranging / ranges from” a first indicate number “to”, “up to”, “until” or “through” (or another such range-indicating term) a second indicate number are used herein interchangeably and are meant to include the first and second indicated numbers and all the fractional and integral numbers therebetween. Unless otherwise indicated, numbers used herein and any number ranges based thereon are approximations within the accuracy of reasonable measurement and rounding errors as understood by persons skilled in the art.
[0151] Programs described in the disclosure can be stored on a non-transitory computer readable medium such as a diskonkey, microSD, DVD or other types of storage elements. The programs may be transferred with the computer readable medium to be executed on a general-purpose computer to implement the system as described in the disclosure.
[0152] As used herein the term “method” refers to manners, means, techniques and procedures for accomplishing a given task including, but not limited to, those manners, means, techniques and procedures either known to, or readily developed from known manners, means, techniques and procedures by practitioners of the chemical, pharmacological, biological, biochemical and medical arts.
[0153] It is appreciated that certain features of the invention, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination or as suitable in any other described embodiment of the invention. Certain features described in the context of various embodiments are not to be considered essential features of those embodiments, unless the embodiment is inoperative without those elements.
[0154] Although the invention has been described in conjunction with specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art. Accordingly, it is intended to embrace all such alternatives, modifications and variations that fall within the spirit and broad scope of the appended claims.
[0155] All publications, patents and patent applications mentioned in this specification are herein incorporated in their entirety into the specification, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated herein by reference. In addition, citation or identification of any reference in this application shall not be construed as an admission that such reference is available as prior art to the present invention. To the extent that section headings are used, they should not be construed as necessarily limiting.
[0156] In addition, any priority document(s) of this application is / are hereby incorporated herein by reference in its / their entirety.
Claims
CLAIMS1. A system for predicting a response to nutritional therapy for treating Crohn’s Disease (CD), comprising: a computer with a processor and memory; a machine learning module; wherein the machine learning module is programed to train a model by accepting a training data set comprising multiomics data of subjects with Crohn’s disease that were treated by exclusive enteral nutrition (EEN) and a determination if each subject was a responder that successfully responded to the treatment or was a non-responder that unsuccessfully responded to the treatment; and wherein the machine learning module uses the trained model to accept multiomics data of a patient and predict if the patient is a responder or a non-responder; wherein the multiomics data includes identified metabolites from serum metabolomics.
2. The system of claim 1, wherein the multiomics data additionally includes identified metabolites from fecal metabolomics.
3. The system of claim 2, wherein the fecal metabolomics were analyzed by an LC- MS / MS analysis to identify metabolites.
4. The system of claim 1, wherein the multomics data additionally includes identified metabolites from one or more of the following: fecal metabolomics, microbiome composition, demographics, a food frequency questionnaire, and clinical parameters.
5. The system of claim 1, wherein some of the identified metabolites indicate that the patient is a responder.
6. The system of claim 1, wherein some of the identified metabolites indicate that the patient is a non-responder.
7. The system of claim 1, wherein some of the identified metabolites indicate that the patient is a responder and some of the identified metabolites indicate that the patient is a non-responder.
8. The system of claim 1, wherein the multiomics data includes identifying microbes.
9. The system of claim 1, wherein the machine learning module uses a random forest algorithm to provide the prediction.
10. The system of claim 4, wherein ranking of the identified metabolites are adjusted based on the food frequency questionnaire, demographics and / or clinical parameters.
11. A method of predicting a response to nutritional therapy for treating Crohn’s Disease (CD), comprising: installing a machine learning module on a computer having a processor and memory; training a model with the machine learning module, by accepting a training data set comprising multiomics data of subjects with Crohn’s disease that were treated by exclusive enteral nutrition (EEN) and a determination if each subject was a responder that successfully responded to the treatment or was a non-responder that unsuccessfully responded to the treatment; and accepting multiomics data of a patient and predicting if the patient is a responder or a non-responder; wherein the multiomics data includes identified metabolites from serum metabolomics.
12. The method of claim 11, wherein the multiomics data additionally includes identified metabolites from fecal metabolomics.
13. The method of claim 12, wherein the fecal metabolomics were analyzed by an LC- MS / MS analysis to identify metabolites.
14. The method of claim 11, wherein the multomics data additionally includes identified metabolites from one or more of the following: fecal metabolomics, microbiome composition, demographics, a food frequency questionnaire, and clinical parameters.
15. The method of claim 11, wherein some of the identified metabolites indicate that the patient is a responder.
16. The method of claim 11, wherein some of the identified metabolites indicate that the patient is a non-responder.
17. The method of claim 11, wherein some of the identified metabolites indicate that the patient is a responder and some of the identified metabolites indicate that the patient is a non-responder.
18. The method of claim 11, wherein the multiomics data includes identifying microbes.
19. The method of claim 11, wherein the machine learning module uses a random forest algorithm to provide the prediction.
20. A non-transitory computer readable medium storing a machine learning module program including instruction which, when executed by a processor, predict if a patient will be a responder or a non-responder according to the method of claim 11.
Citation Information
Patent Citations
Crohn disease patient diet management system and method
CN118969197A
Method and system for microbiome-derived diagnostics and therapeutics
US20170286619A1
Generating predicted values of biomarkers for scoring food
US20190251861A1