Multi-OMIC patient stratification in inflammatory bowel disease treatment
By integrating multi-omic data and using machine learning for unsupervised clustering, the method addresses the lack of precision in IBD treatment, enabling personalized therapy predictions and improved patient outcomes.
Patent Information
- Application Number
- PCT/US2025/020865
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-23
- Filing Date
- 2025-03-21
- Publication Date
- 2025-09-25
AI Technical Summary
Current methods for treating Inflammatory Bowel Disease (IBD) lack precision and fail to account for individual variability, leading to inconsistent treatment responses among patients.
A method utilizing multi-omic data integration and machine learning algorithms to identify biomarkers associated with treatment responses, employing unsupervised clustering to stratify patient populations into phenotypic groups, and defining likely responders based on these biomarkers.
Enhances personalized treatment strategies by accurately predicting patient responses to therapies, improving treatment efficacy and outcomes for IBD patients.
Smart Images

Figure US2025020865_25092025_PF_FP_ABST
Abstract
Description
MULTI-OMIC PATIENT STRATIFICATION IN INFLAMMATORY BOWEL DISEASE TREATMENTCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority and benefit of U.S. Provisional Patent Application No. 63 / 568,422 filed on March 21, 2024 and U.S. Provisional Patent Application No. 63 / 674,797 filed on July 23, 2024. The contents of this application are incorporated herein by reference in their entirety for all purposes.TECHINCAL FIELD
[0002] The present disclosure generally relates to the field of precision medicine, and more specifically, to methods and systems for stratifying patient populations in Inflammatory Bowel Disease (IBD) treatment based on multi-omic data analysis.BACKGROUND
[0003] Inflammatory Bowel Disease (IBD) is a term that encompasses a group of disorders that cause inflammation in the digestive tract. The two main types of IBD are Crohn's disease and ulcerative colitis. These conditions are characterized by chronic inflammation of the gastrointestinal tract, leading to symptoms such as abdominal pain, diarrhea, rectal bleeding, and weight loss. The exact cause of IBD is unknown, but it is thought to result from a combination of genetic and environmental factors, as well as an abnormal immune response.
[0004] Precision medicine is an emerging approach for disease treatment and prevention that takes into account individual variability in genes, environment, and lifestyle for each person. This approach allows doctors and researchers to predict more accurately which treatment and prevention strategies for a particular disease will work in which groups of people. It is in contrast to a one-size-fits-all approach, in which disease treatment and prevention strategies are developed for the average person, with less consideration for the differences between individuals.
[0005] Multi-omic data refers to the integration of multiple types of biological data, including genomics, tran scrip tomics, and proteomics. Genomics is the study of the complete set of genes within an organism (the genome), and how these genes interact with each other and the environment. Transcriptomics is the study of the complete set of RNA transcriptsproduced by the genome under specific circumstances or in a specific cell. Proteomics is the large-scale study of proteins, particularly their structures and functions.
[0006] Machine learning is a type of artificial intelligence that enables computers to learn from and make decisions or predictions based on data. Machine learning algorithms can analyze large amounts of data and identify patterns or trends that may not be immediately apparent to human analysts. These algorithms can be used to analyze multi-omic data and identify biomarkers, which are measurable substances in an organism whose presence is indicative of some phenomenon such as disease, infection, or environmental exposure.
[0007] Unsupervised clustering is a type of machine learning algorithm used to group data points into clusters based on their similarity. This technique can be used to stratify patient populations into phenotypic groups based on identified biomarkers.
[0008] In the context of IBD treatment, these techniques can potentially be used to predict patient responses to specific therapies, enabling more personalized and effective treatment strategies. However, the development and validation of such methods present considerable challenges and are the subject of ongoing research.BRIEF SUMMARY
[0009] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.
[0010] According to an aspect of the present disclosure, a method for stratifying a patient population in Inflammatory Bowel Disease (IBD) treatment includes accessing a multi-omic dataset comprising at least two of a genomic profile of patient data, a transcriptomic profile of patient data, or a proteomic profile of patient data. The method further includes employing a machine learning algorithm to analyze the dataset and identify biomarkers associated with a response to a therapy for treating IBD. The patient population is stratified into phenotypic groups based on the identified biomarkers using unsupervised clustering. A patient population that is likely to respond to this therapy is then defined based on the stratification and further analysis of patient metadata.
[0011] Methods according to any embodiment of any aspect of the disclosure may be computer-implemented unless context indicates otherwise, such as e.g. where sample handling steps are involved. Treatment may include administering a therapeutic agent, as describedherein. As the skilled person understands, the methods according to the present aspect can also be used to determine whether a patient with IBD is likely to respond to treatment. Further, the identification of biomarkers associated with a response to a therapy for treating IBD may be implicit in the training of the machine learning model. In other words, the configuration (e.g. predictive features and / or parameters) of the trained machine learning model may implicitly or explicitly identify biomarkers from the multi-omic dataset that are associated with response to therapy that the machine learning algorithm has been trained to predict. Therefore, the step of accessing a multi-omic dataset comprising at least two of a genomic profile of patient data, a transcriptomic profile of patient data, or a proteomic profile of patient data can comprise accessing a multi-omic dataset comprising at least two of a genomic profile for the patient, a transcriptomic profile for the patient, and a proteomic profile for the patient. Further, the step of employing a machine learning algorithm to analyze the dataset and identify biomarkers associated with a response to a therapy for treating IBD may comprise one or both of: (i) training a machine learning model to predict a response to a therapy for treating IBD, and (ii) providing the multi-omic data set for a patient as input to a machine learning model that has been trained to predict a response to a therapy for treating IBD, thereby obtaining a prediction indicative of whether the patient is likely to respond to the therapy. The machine learning model may have been trained to predict a response to a therapy for treating IBD (thereby implicitly identifying biomarkers associated with this response). Therefore, the machine learning model may have been trained using data comprising, for each of a plurality of patients with IBD in a training patient population: (i) a multi-omic dataset comprising at least two of a genomic profile for the patient, a transcriptomic profile for the patient, and a proteomic profile for the patient, and (ii) a ground truth predicted response to therapy. In embodiments comprising training of the machine learning model, the step of accessing a multi-omic dataset comprising at least two of a genomic profile of patient data, a transcriptomic profile of patient data, or a proteomic profile of patient data may comprise accessing said profiles for each of the plurality of patients in the patient population. As the skilled person understands, as the patient population (also referred to herein as “training patient population”) may be stratified into phenotypic groups using unsupervised clustering where a patient population that is likely to respond to this therapy can be defined based on the stratification and further analysis of patient metadata, the stratification also implicitly identified the biomarkers that differentiate between the patient population that is likely to respond to the therapy and patients that are unlikely to respond to therapy. Thus, the method may comprise stratifying the training patient population into phenotypic groups based on the multi-omic datasets using unsupervised clustering (andimplicitly, the identified biomarkers); an defining a patient population predicted to respond to a therapy based on the stratification. This can in turn be used as ground truth labels for the training of the machine learning model.
[0012] In embodiments, the stratifying the patient population (e.g. training patient population) into phenotypic groups using unsupervised clustering may comprise using a plurality of biomarkers that have been identified using a multi-omic analysis method that jointly models at least two of genomic profiles, transcriptomic profiles, and proteomic profiles for a patient population and identifies latent factors associated with sources of variability in the datasets. The multi-omic analysis method may be a dimensionality reduction method, such as t-SNE, UMPA, PCA, or a Multi-Omics Factor Analysis (MOFA). Thus, the biomarkers used for unsupervised clustering may be summarized biomarkers that combine information across a plurality of features of the genomic profiles, transcriptomic profiles, and / or proteomic profiles. Thus, also described according to other aspects of the disclosure is a method for stratifying a patient population in Inflammatory Bowel Disease (IBD) treatment includes accessing a multi- omic dataset comprising at least two of a genomic profile of patient data, a transcriptomic profile of patient data, or a proteomic profile of patient data. The method further includes employing a multi-omic analysis method (i.e. a method that jointly models at least two of genomic profiles, transcriptomic profiles, and proteomic profiles for a patient population and identifies latent factors associated with sources of variability in the datasets) to analyze the dataset and identify biomarkers (e.g. “summary biomarkers”). These biomarkers may be associated with a response to a therapy for treating IBD. The patient population is stratified into phenotypic groups based on the identified biomarkers using unsupervised clustering. A patient population that is likely to respond to this therapy is then defined based on the stratification and further analysis of patient metadata (thereby confirming that the biomarkers based on which the unsupervised clustering was done are associated with a response to a therapy for treating IBD).
[0013] According to other aspects of the present disclosure, the method may include validating the identified biomarkers using an independent patient dataset. The method may also include analyzing tissue types separately or using appropriate statistical methods like fixed / random effects models to account for repeated sampling. The method may further include checking if any results are confounded by medication usage.
[0014] According to another aspect of the present disclosure, a system for precision medicine in Inflammatory Bowel Disease (IBD) treatment includes a data storage unitconfigured to store a multi-omic dataset comprising genomic, transcriptomic, and proteomic profiles of patient data. The system also includes a processor configured to employ machine learning algorithms to analyze the dataset to identify biomarkers associated with a response to a drug for treating IBD, stratify the patient population into phenotypic groups based on the identified biomarkers using unsupervised clustering, and define a patient population predicted to respond to the drug based on the stratification and further analysis of patient metadata.
[0015] According to other aspects of the present disclosure, the system may include a user interface for inputting additional patient metadata. The system may also be integrated with electronic health records to automatically access and update patient metadata. The system may further be configured to provide alerts when new biomarkers are identified that may affect the stratification.
[0016] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform a method for stratifying a patient population for precision medicine in Inflammatory Bowel Disease (IBD) treatment. The method includes accessing a multi-omic dataset comprising genomic, transcriptomic, and proteomic profiles of patient data, employing machine learning algorithms to analyze the dataset and identify biomarkers associated with a response to a drug for treating IBD, stratifying the patient population into phenotypic groups based on the identified biomarkers using unsupervised clustering, and defining a patient population predicted to respond to the drug based on the stratification and further analysis of patient metadata.
[0017] According to other aspects of the present disclosure, the instructions may further cause the processor to validate the identified biomarkers using a cross-validation technique. The instructions may also cause the processor to adjust the stratification based on feedback from clinical outcomes. The instructions may further cause the processor to perform feature selection to identify the biomarkers with the greatest predictive power. The instructions may also cause the processor to integrate additional omic data types, such as metabolomic or lipidomic data.
[0018] The foregoing general description of the illustrative embodiments and the following detailed description thereof are merely exemplary aspects of the teachings of this disclosure and are not restrictive.BRIEF DESCRIPTION OF THE DRAWINGS
[0019] FIG. 1 illustrates a workflow for multi-omics data integration and analysis in inflammatory bowel disease research, according to aspects of the present disclosure.
[0020] FIG. 2 shows scatter plots representing principal component analysis of transcriptomics data, in accordance with example embodiments.
[0021] FIG. 3 depicts performance metrics and important features of a multi-omics classification model for inflammatory bowel disease subtypes, according to an embodiment.
[0022] FIG. 4 presents visualizations of disease severity and biological pathways in inflammatory bowel disease, according to aspects of the present disclosure.
[0023] FIG. 5 shows data visualizations related to multi-omics analysis in inflammatory bowel disease, in accordance with example embodiments.
[0024] FIG. 6 shows distributions of relevant metadata information in the SPARC dataset. Specifically, a subset of samples for which the patient has been diagnosed with UC or CD.
[0025] FIGs. 7A-7D show examples of genes identified significantly down- / up-regulated in inflamed samples. CXCL9 (proteomics) (7 A), C2orf88 (transcriptomics) (7B), SLC35A3 (transcriptomics) (7C) and CDH17 (transcriptomics) (7D). Blue letters indicate proteomics while green letters indicate transcriptomics panels.
[0026] FIG. 8 shows examples of genes identified significantly up-regulated in inflamed samples in Cluster B.
[0027] FIG. 9 depicts an illustrative implementation of a computer system 900 that may be used in connection with some embodiments of the technology described herein.DETAILED DESCRIPTION
[0028] The following description sets forth exemplary aspects of the present disclosure. It should be recognized, however, that such description is not intended as a limitation on the scope of the present disclosure. Rather, the description also encompasses combinations and modifications to those exemplary aspects described herein.
[0029] The present disclosure elaborates on a method for stratifying patient populations for precision medicine in the treatment of Inflammatory Bowel Disease (IBD), a condition affecting millions and characterized by a high degree of heterogeneity in patient response to treatments. The disclosed method leverages a comprehensive multi-omic dataset, integrating genomic, transcriptomic, and proteomic data, to identify subpopulations of patients who arelikely to benefit from a therapeutic agent, such as a compound that functions as an inflammasome pathway inhibitor with multi-cytokine effects.
[0030] The method leverages comprehensive data sets to enhance its analytical capabilities and precision in stratifying patient populations for IBD treatment. These data sets may encompass a wide range of biological and clinical information, including but not limited to genetic profiles, gene expression data, protein abundance measurements, and patient metadata. The types of data considered can span from quantitative measurements (e.g., biomarker levels, disease activity scores) to qualitative assessments (e.g., endoscopic findings, patient-reported outcomes) and longitudinal data (e.g., treatment response over time, disease progression patterns). By integrating these diverse data types, the method aims to capture the multifaceted nature of IBD and identify meaningful patterns that can inform personalized treatment strategies.
[0031] The method may utilize data sets from various sources, including but not limited to: electronic health records (EHRs) from hospitals and clinics; patient-reported outcome measures (PROMs) collected through surveys or mobile applications; genomic sequencing data from blood or tissue samples; transcriptomic profiles obtained from intestinal biopsies; proteomic data derived from blood, stool, or urine samples; metabolomic profiles from blood or urine samples; microbiome sequencing data from stool samples; imaging data such as endoscopy, MRI, or CT scans; dietary and lifestyle information collected through questionnaires or wearable devices; environmental exposure data from geographical information systems or personal monitoring devices; pharmacogenomic data related to drug metabolism and response; immunological profiles including cytokine levels and immune cell populations; epigenetic data such as DNA methylation patterns; longitudinal clinical data on disease progression and treatment outcomes; biobank samples linked to clinical information; data from clinical trials and observational studies; public databases such as the Gene Expression Omnibus (GEO) or The Cancer Genome Atlas (TCGA); multi-center collaborative research initiatives focused on IBD; and population-based cohort studies with long-term follow-up data.
[0032] In some embodiments, the dataset utilized in this method is sourced from the IBD Plexus platform, specifically the Study of a Prospective Adult Research Cohort with IBD (SPARC IBD), encompassing data from 3,133 individuals diagnosed with Crohn's disease (CD) or ulcerative colitis (UC). The genomic data may include, but is not limited to, whole exome sequencing and / or genome-wide association studies, while the transcriptomic data isderived from tissue biopsies (e.g. intestine biopsies) and the proteomic data from blood plasma samples. Alternative sources of data, such as microbiomic profiles from stool samples, may also be incorporated to enhance the multi-omic dataset.
[0033] Normalization of transcriptomic data can be achieved through a multi-step process involving within-batch and / or across-batch normalization techniques. For within-batch normalization, DESeq2 may be employed. DESeq2 is a method for differential expression analysis that uses negative binomial generalized linear models. It generally functions by (1) estimating size factors to account for differences in sequencing depth between samples, (2) estimating gene- wise dispersion parameters, (3) fitting a negative binomial model for each gene, and (4) performing statistical tests for differential expression. DESeq2 takes raw count data as input and outputs normalized count data and differential expression results. Various implementations may use different dispersion estimation methods or statistical tests. For across-batch normalization, ComBat-seq may be utilized. ComBat-seq is an empirical Bayes method for removing batch effects in sequencing count data. It generally operates by (1) standardizing the data across genes, (2) estimating batch and condition effects using a generalized linear model, and (3) adjusting the data to remove batch effects while preserving biological variation. ComBat-seq takes normalized count data and batch information as input, outputting batch-corrected count data. Different implementations may use various prior distributions or parameter estimation techniques. Alternative methods such as edgeR or limma can be employed for similar purposes. edgeR uses negative binomial models and empirical Bayes methods, while limma employs linear models and empirical Bayes methods for differential expression analysis. Multi-omics factor analysis (MOFA; Argelaguet et al. Molecular Systems Biology, Vol. 14, No. 6, June 2018, 14:e8124) may be employed to integrate the omics datasets, providing a holistic view of the molecular underpinnings of IBD. MOFA is an unsupervised method for integrating multi-omics data. It generally functions by (1) modeling each omics dataset as a matrix factorization (specifically, the product of a matrix of factors for each sample, with dimensions number of factors * numbers of samples, that is common across modalities, and a matrix of weights, with dimensions numbers of features in a specific modality * number of samples, that is specific to each modality), (2) sharing factors across datasets to capture common patterns, (3) using sparsity constraints to encourage interpretable factors, and (4) performing variational inference to estimate model parameters. The factors identified may also be referred to as “latent factors”. MOFA takes multiple omics datasets as input and outputs latent factors that explain sources of variation across datasets.Various implementations may use different optimization algorithms or factor selection methods. In embodiments using the MOFA model, the MFOA model is trained on metadata categories that may include, but are not limited to, disease state, treatment response, patient demographics, and environmental exposures. These metadata categories serve as covariates in the model, allowing for the identification of factors associated with specific clinical or environmental variables. In embodiments using the MOFA model, one or more factors identified by the MOFA model may be identified as being associated with specific clinical or environmental variables (together referred to as “metadata categories”), using e.g. a regression model (e.g. a Cox regression model). For example, one or more factors identified by the MOFA model can be used as predictors in models that predict metadata categories associated with clinical outcome, such as disease state and treatment response. Metadata categories such as patient demographics and environmental exposures may be used as covariates in such a model.
[0034] Subsequent to the integration, patient subgroups are identified through unsupervised clustering techniques such as hierarchical clustering, with alternatives including k-means clustering or Gaussian mixture models. In embodiments, the unsupervised clustering is applied to integrated multi-omics data, such as e.g. a matrix of latent factors for each sample identified using MOFA. In embodiments, one or more molecular features characteristic of each cluster may then be identified based on one or more of the original genomic, transcriptomic and proteomic data.
[0035] In some embodiments, the unsupervised clustering may be performed using at least one of hierarchical clustering, k-means clustering, or Gaussian mixture models. Hierarchical clustering builds a hierarchy of clusters, either in a top-down (divisive) or bottom-up (agglomerative) approach. It can create a tree-like structure of nested clusters. Agglomerative clustering starts with individual data points and merges them, while divisive clustering starts with all data in one cluster and recursively divides it. This approach can provide a clear visualization of the cluster hierarchy, allowing for intuitive interpretation of patient subgroups at different levels of granularity. K-means clustering is an iterative algorithm that partitions data into K pre-defined clusters by minimizing the within-cluster sum of squares. Implementations may use different initialization methods (e.g., random, k-means++), or incorporate fuzzy logic for soft clustering. This method is computationally efficient for large datasets, making it suitable for analyzing multi-omic data from numerous patients. Gaussian mixture models is a probabilistic model that assumes data points are generated from a mixture of a finite number of Gaussian distributions with unknown parameters, and infers theparameters of these distributions. This approach can be implemented with different covariance structures (e.g., spherical, diagonal, full) or using variational inference for faster computation. It provides a soft clustering approach, allowing for uncertainty in subgroup assignments, which can be beneficial when patient characteristics overlap between groups. In exemplary embodiments, the molecular features characteristic of each cluster may be determined using mutual information and partial least squares discriminant analysis, and gene sets for each cluster may be aggregated to gene ontology terms using GSEApy's EnrichR module or other pathway analysis tools such as DAVID or Reactome. Mutual information is a measure of the mutual dependence between two variables, quantifying the amount of information obtained about one variable by observing the other. It can be implemented using different estimation techniques, such as k-nearest neighbors or kernel density estimation, and may capture nonlinear relationships between features and cluster assignments, potentially revealing complex biological interactions.
[0036] Partial Least Squares Discriminant Analysis (PLS-DA) is a supervised method that combines dimensionality reduction with classification, projecting predictors and response variables into a new space to maximize covariance. Extensions may include orthogonal PLS (OPLS) or multi-block PLS for handling multiple data types simultaneously. This method can be effective for high-dimensional data, allowing for the identification of key molecular features that distinguish between patient subgroups. GSEApy's EnrichR module is a Python implementation of the Enrichr tool, which performs gene set enrichment analysis to identify biological pathways or functions overrepresented in a set of genes. It can use different gene set databases (e.g., GO, KEGG, Reactome) or custom gene sets specific to IBD research, and can provide a comprehensive analysis of the biological significance of identified patient subgroups, linking molecular features to known pathways and functions. DAVID (Database for Annotation, Visualization and Integrated Discovery) is a web-based tool that integrates functional genomic annotations with intuitive graphical summaries. It can offer multiple analysis modules, including functional annotation clustering and gene functional classification, and may provide a user-friendly interface for researchers to explore the biological context of identified patient subgroups, facilitating hypothesis generation. Reactome is a free, open- source, curated and peer-reviewed pathway database. It can be used through web interface, R package, or API for programmatic access, and may offer detailed molecular-level pathway information, allowing for in-depth exploration of the biological processes underlying patient subgroups in IBD.
[0037] Prophylactic evaluation in the context of inflammatory bowel disease (IBD) research involves assessing potential treatments before the onset of symptoms or disease progression. This approach is crucial for understanding how compounds may prevent or mitigate IBD, potentially leading to more effective early intervention strategies. By evaluating compounds prophylactically, researchers can gain insights into their mechanisms of action and identify patient populations that might benefit most from preventive therapies, ultimately improving long-term outcomes for IBD patients.
[0038] In some embodiments, the therapeutic efficacy of a compound or therapeutic treatment is evaluated prophylactically in a dextran sodium sulfate (DSS)-induced colitis model, with colon tissue subjected to RNA sequencing to ascertain gene expression changes. In embodiments, the RNA sequencing data is subject to unsupervised clustering, and a cluster of samples characterized by upregulation of omics features related to inflammasome activation and cytokine signaling is identified, suggesting a population (e.g., with similar features as a DSS-induced colitis model) that could uniquely benefit from a compound or therapeutic treatment. This cluster also exhibits elevated disease severity scores, indicating a potential for substantial therapeutic impact.
[0039] In other embodiments, the therapeutic efficacy could be evaluated in alternative animal models of colitis, such as the TNBS (2,4,6-trinitrobenzenesulfonic acid)-induced colitis model or the IL- 10 knockout mouse model. These models may provide complementary insights into different aspects of IBD pathogenesis. In addition to RNA sequencing, other omics approaches can be employed to assess changes in the colon tissue, including proteomics analysis to directly measure protein levels and post-translational modifications, metabolomics profiling to identify changes in metabolite concentrations, and epigenomics analysis, such as ChlP-seq or ATAC-seq, to examine changes in chromatin structure and gene regulation. The cluster analysis can be expanded to include multi-omics data integration, combining transcriptomics with proteomics and / or metabolomics data to provide a more comprehensive molecular profile of the responsive population. For example, multi-omics data may be analysed as described herein and multi-omics factors that are associated with a responsive population may be identified.
[0040] In certain embodiments, machine learning algorithms, such as random forests or support vector machines, are applied to the multi-omics data to improve the identification and characterization of responsive patient clusters. In embodiments, a machine learning model may be trained to take as input one or more of genomic data, proteomic data and transcriptomic datafor a subject, and classify the subject between at least a first class that is likely to respond to a treatment and a second class that is less likely to respond to the treatment. The machine learning model may use training data comprising, for each of a plurality of subjects in a training cohort, one or more of genomic data, proteomic data and transcriptomic data, and a known class label derived from unsupervised clustering of the integrated multi-omic data for the training cohort. For example, the unsupervised clustering may identify a first set of one or more clusters that comprise subjects that are likely to respond to the treatment, and a second set of one or more clusters that comprise subjects that are less likely to respond to the treatment. Such identification may be based on e.g. associations between multi-omic latent features and clinical outcome metrics, and / or molecular characteristics of the subjects in the clusters. Subjects in the first set of one or more clusters may be associated with a known class label corresponding to the first class. Subjects in the second set of one or more clusters may be associated with a known class label corresponding to the second class. The compound or therapeutic treatment can be evaluated in combination with standard-of-care treatments for IBD, such as corticosteroids or immunomodulators, to assess potential synergistic effects. Time-course experiments can be conducted to track changes in gene expression and other molecular markers over the course of treatment, providing insights into the kinetics of the therapeutic response. Ex vivo studies using patient-derived organoids or explant cultures could be performed to validate the findings from animal models in a more human-relevant context.
[0041] In some embodiments, the identified cluster of responsive patients is further characterized using single-cell RNA sequencing to reveal cell type-specific responses to the treatment. Biomarkers associated with the responsive cluster can be developed into a diagnostic assay to predict treatment efficacy in clinical settings, potentially using less invasive sampling methods such as blood or stool analysis.
[0042] The method further includes the step of validating the identified patient clusters and biomarkers using independent patient datasets, which may be sourced from different cohorts or geographical regions. This validation ensures the robustness and generalizability of the stratification model. Additionally, the method contemplates the monitoring of patient populations for changes in biomarkers over time, which may necessitate the adjustment of therapeutic regimens.
[0043] In some embodiments, the method includes validating the identified patient clusters and biomarkers using independent patient datasets. For example, the validation may involve applying the stratification model developed using the SPARC IBD cohort to data from otherlarge-scale IBD studies, such as the 1000IBD project or the International Inflammatory Bowel Disease Genetics Consortium dataset. This cross-cohort validation may help assess the robustness of the identified biomarkers and patient subgroups across different patient populations. The validation process may include comparing the performance metrics of the machine learning classifier, such as accuracy, precision, recall, and AUC-ROC, between the original SPARC IBD dataset and the independent validation datasets. Additionally, the consistency of the identified biomarkers and their associations with clinical outcomes may be evaluated across these different cohorts to ensure the generalizability of the findings.
[0044] In some aspects, alternative validation approaches may be employed to further strengthen the robustness of the stratification model. These may include techniques such as bootstrapping or jackknife resampling to generate multiple subsets of the original data for internal validation. Another approach may involve temporal validation, where the model is trained on data from earlier time points and validated on more recent data from the same cohort, assessing its predictive power over time. In some cases, external validation may be performed using data from clinical trials or real-world evidence databases, providing insights into the model's performance in different clinical settings. Furthermore, the validation process may be extended to include functional validation of key biomarkers through in vitro or in vivo experiments, corroborating their biological relevance in IBD pathogenesis or treatment response.
[0045] Commercially relevant alternatives to the disclosed method may include the use of different machine learning algorithms for data analysis, such as support vector machines, neural networks, or ensemble methods. The choice of algorithm may be influenced by factors such as dataset size, feature dimensionality, and computational resources. Furthermore, the method may be adapted to include additional omic data types, such as metabolomic or lipidomic data, to further refine patient stratification.
[0046] The disclosed method and system provide an approach to precision medicine in IBD treatment, offering the potential to improve patient outcomes through the targeted delivery of a compound to those subpopulations predicted to respond favorably. This approach exemplifies the integration of multi-omic data analysis with advanced machine learning techniques to address the complexities of IBD and enhance the efficacy of therapeutic interventions.
[0047] The present disclosure relates to methods, systems, and computer-readable media for stratifying a patient population in the context of Inflammatory Bowel Disease (IBD) treatment. More specifically, the disclosure pertains to the integration and analysis of multi- omic data, including genomic, transcriptomic, and proteomic profiles, to identify biomarkers associated with a response to a specific IBD therapy. The identified biomarkers may then be used to stratify the patient population into phenotypic groups, thereby enabling the definition of a patient population that is likely to respond to the therapy. This approach leverages machine learning algorithms and unsupervised clustering techniques to analyze and categorize the patient data, providing a more personalized and potentially effective treatment strategy for IBD.
[0048] In some aspects, the disclosure also encompasses a system for precision medicine in IBD treatment. This system may include a data storage unit configured to store the multi- omic dataset and a processor configured to employ machine learning algorithms to analyze the dataset, identify biomarkers, stratify the patient population, and define a patient population predicted to respond to the therapy. The system may further include a user interface for inputting additional patient metadata and may be integrated with electronic health records to automatically access and update patient metadata. Also disclosed herein is a system comprising one or more processors and one or more non-transitory computer readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to implement a method of the disclosure.
[0049] In other aspects, the disclosure relates to a non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method of the disclosure, such as a method for stratifying a patient population for precision medicine in IBD treatment. This method may involve accessing a multi-omic dataset, employing machine learning algorithms to analyze the dataset and identify biomarkers, stratifying the patient population into phenotypic groups based on the identified biomarkers, and defining a patient population predicted to respond to the therapy.
[0050] Furthermore, the disclosure may also encompass a method for predicting a response to a treatment regimen for IBD in a subject. This method may involve obtaining a biological sample from the subject, contacting the biological sample with a set of probes capable of detecting a panel of biomarkers (e.g., selected from biomarkers described herein), detecting the presence or absence of the biomarkers in the sample, analyzing the pattern ofbiomarkers detected to determine an enrichment score for the sample, and predicting the subject's response to the treatment regimen based on the enrichment score.
[0051] In yet another aspect, the disclosure relates to a method for categorizing patients based on multi-omic biological information. This method may involve preprocessing metadata associated with biological samples (e.g., obtaining pre-processed data from an online database (e.g., SPARC IBD cohort)), assigning biological information to the preprocessed metadata, annotating the biological information, correcting batch effects in the biological information, applying Multi-Omics Factor Analysis (MOFA) to integrate the biological information, training a MOFA model, analyzing the integrated biological information using the trained MOFA model, identifying patient clusters, and validating the identified patient clusters through pathway analysis.
[0052] These and other aspects of the present disclosure provide a comprehensive and personalized approach to IBD treatment, potentially improving patient outcomes and advancing the field of precision medicine.
[0053] In some aspects, the method for stratifying a patient population in Inflammatory Bowel Disease (IBD) treatment involves accessing a multi-omic dataset. This dataset may comprise at least two of a genomic profile of patient data, a transcriptomic profile of patient data, or a proteomic profile of patient data. In some cases, the multi-omic dataset may further comprise microbiomic profiles (e.g., a description of the type and / or quantity of microbes in the patient’s gut) of the patient data. The genomic, transcriptomic, proteomic, and microbiomic profiles may be derived from biological samples collected from the patients, such as blood, stool, tissue biopsy, or saliva samples.
[0054] In some embodiments, machine learning algorithms are employed to analyze the multi-omic dataset and identify biomarkers associated with a response to a therapy for treating IBD. The machine learning algorithms used to analyze the dataset and identify biomarkers can include at least one of a support vector machine, a neural network, a decision tree, a random forest, or a gradient boosting machine. These algorithms may be selected based on their ability to handle high-dimensional data, their predictive accuracy, or other relevant factors.
[0055] In some cases, the identified biomarkers are validated using an independent patient dataset. This validation step may involve comparing the performance of the machine learning algorithms on the original dataset with their performance on the independent dataset, assessingthe consistency of the identified biomarkers across different datasets, or other suitable validation techniques.
[0056] In some embodiments, the patient population is stratified into phenotypic groups based on the identified biomarkers using unsupervised clustering. The unsupervised clustering may be performed using at least one of k-means clustering, hierarchical clustering, or Gaussian mixture models. These clustering techniques may be selected based on their ability to identify distinct groups within the patient population, their computational efficiency, or other relevant factors.
[0057] In some cases, a patient population is defined that is predicted to respond to the therapy based on the stratification and further analysis of patient metadata (e.g., stratification into cluster A or cluster B). The patient metadata may include at least one of age, gender, disease duration, prior medication history, or lifestyle factors. This metadata may be used to refine the stratification, identify potential confounding factors, or provide additional context for interpreting the identified biomarkers.
[0058] In some embodiments, the therapy for treating IBD may involve the administration of a therapy, such as a small molecule compound, optionally in combination with at least one other IBD treatment. Non-limiting examples of suitable therapies include a therapeutic agent useful in the treatment of inflammasome-related diseases / disorders, immune diseases, inflammatory diseases, auto-immune diseases, auto-inflammatory diseases, or cancer, as disclosed herein.
[0059] In various embodiments, the therapeutic agent is selected from famesoid X receptor(FXR) agonists; anti-steatotics; anti-fibrotic s; JAK inhibitors; checkpoint inhibitors; chemotherapy, radiation therapy and surgical procedures; urate-lowering therapies; anabolics and cartilage regenerative therapy; blockade of IL- 17; complement inhibitors; Bruton's tyrosine Kinase inhibitors (BTK inhibitors); Toll Like receptor inhibitors (TLR7 / 8 inhibitors); CAR-T therapy; anti-hypertensive agents; cholesterol lowering agents; leukotriene A4 hydrolase (LTAH4) inhibitors; SGLT2 inhibitors; P 2- agonists; anti-inflammatory agents; nonsteroidal anti-inflammatory drugs (“NSAIDs”); acetylsalicylic acid drugs (ASA) including aspirin; paracetamol; regenerative therapy treatments; cystic fibrosis treatments; and atherosclerotic treatment.
[0060] In some embodiments, the therapy is one or more of the compounds described in paragraphs
[0047] to
[0083] in International Application No PCT / US2024 / 030111 filed on October 2, 2024, which is incorporated herein by reference in its entirety.
[0061] In some cases, the method further comprises the step of monitoring the patient population for changes in the biomarkers over time. This monitoring step may involve repeated collection and analysis of biological samples from the patients, tracking changes in the identified biomarkers, or other suitable monitoring techniques. The results of this monitoring step may be used to adjust the therapy, refine the stratification, or provide additional insights into the progression of IBD.
[0062] In some embodiments, the identified biomarkers are further analyzed to determine their genetic pathways and interactions. This analysis may involve mapping the biomarkers onto known genetic pathways, identifying potential interactions between the biomarkers, or other suitable analysis techniques. The results of this analysis may provide additional insights into the mechanisms of action of the therapy, the pathophysiology of IBD, or other relevant aspects.
[0063] In some cases, the method further comprises the step of providing personalized treatment recommendations for each stratified phenotypic group. These recommendations may be based on the predicted response of the phenotypic group to the therapy, the specific characteristics of the phenotypic group, or other relevant factors.
[0064] In some embodiments, the stratification is used to exclude patients from the patient population who are predicted not to respond to the drug. This exclusion step may involve identifying patients who fall into phenotypic groups with a low predicted response to the therapy, removing these patients from the patient population, or other suitable exclusion techniques. The results of this exclusion step may be used to focus the therapy on patients who are likely to benefit, potentially improving the overall effectiveness of the therapy and reducing unnecessary treatment.
[0065] In some aspects, the method involves accessing a multi-omic dataset that includes genomic, transcriptomic, and proteomic profiles of patient data. The genomic profile may include data derived from whole exome sequencing or global screening array (e.g. genome wide SNP arrays), and may be processed to remove single nucleotide polymorphisms (SNPs) where there is less than 1% variation. The transcriptomic profile may include data derived from RNA sequencing of colon tissue, and may be processed using normalization techniques suchas DESeq2 normalization for within-batch normalization and ComBat-seq for across batch normalization. The proteomic profile may include data derived from blood samples, and may be processed using normalization techniques appropriate for proteomic data.
[0066] In some cases, the multi-omic dataset may be stored in a data storage unit, such as a cloud-based storage system or a local server. The data storage unit may be configured to receive and store updates to the multi-omic dataset, allowing for continuous updating and refinement of the dataset as new patient data is collected. The data storage unit may also be integrated with electronic health records to automatically access and update patient metadata, such as age, gender, disease duration, prior medication history, or lifestyle factors.
[0067] In some embodiments, a processor may be configured to employ machine learning algorithms to analyze the multi-omic dataset and identify biomarkers associated with a response to a therapy for treating IBD. The machine learning algorithms may include at least one of a support vector machine, a neural network, a decision tree, a random forest, or a gradient boosting machine. The processor may be further configured to perform cross- validation of the machine learning algorithms using a separate validation dataset, to assess the robustness and generalizability of the identified biomarkers. In embodiments, a machine learning model may be trained to take as input genomic data (e.g. a genomic profile), proteomic data (e.g. a proteomic profile) and transcriptomic data (e.g. a transcriptomic profile) for a subject, and classify the subject between at least a first class that is likely to respond to a treatment and a second class that is less likely to respond to the treatment. The machine learning model may use training data comprising, for each of a plurality of subjects in a training cohort, genomic data, proteomic data and transcriptomic data, and a known class label indicating whether the subject is considered likely to respond to the treatment. The known class label may be identified based on one or more clinical outcome metrics included in metadata associated with subjects in the training cohort.
[0068] In some cases, the processor may be further configured to generate reports summarizing the stratification and predicted responses of the patient population. These reports may include visualizations of the stratification, lists of the identified biomarkers, descriptions of the associated genetic pathways and interactions, and other relevant information. The reports may be generated in a format suitable for clinical interpretation, such as tables, graphs, or heat maps.
[0069] In some embodiments, the system may further comprise a user interface for inputting additional patient metadata. The user interface may be designed to be user-friendly and intuitive, allowing clinicians or researchers to easily input and update patient metadata. The user interface may also provide access to the reports generated by the processor, allowing users to view and interpret the stratification and predicted responses of the patient population.
[0070] In some cases, the system may be configured to provide alerts when new biomarkers are identified that may affect the stratification. These alerts may be delivered via the user interface, email, text message, or other suitable notification methods. The alerts may include information about the new biomarkers, their potential impact on the stratification, and suggested actions for the user to take.
[0071] In some embodiments, the system may include a module for simulating patient responses to the treatment (e.g. administering a drug) based on the stratification. This module may use mathematical models, machine learning algorithms, or other suitable techniques to predict how individual patients or phenotypic groups are likely to respond to the drug. The results of these simulations may be used to refine the stratification, adjust the therapy, or provide additional insights into the mechanisms of action of the drug.
[0072] In some cases, the system may be configured to operate in a cloud computing environment. This configuration may allow for scalable storage and processing capacity, remote access to the system, and other benefits associated with cloud computing. The cloud computing environment may be secured using encryption, access controls, or other suitable security measures to protect the confidentiality and integrity of the patient data.
[0073] In some embodiments, the system may be further configured to provide decision support for selecting combination therapies based on the stratification. This decision support may involve suggesting additional drugs or therapies that may enhance the effectiveness of the primary therapy, based on the identified biomarkers and their associated genetic pathways and interactions. The decision support may also involve suggesting adjustments to the dosage or administration schedule of the primary therapy, based on the predicted responses of the patient population.
[0074] In some cases, the system may be further configured to track the outcomes of patients and refine the stratification model based on real- world data. This tracking may involve collecting follow-up data on the patients' responses to the therapy, their disease progression, their quality of life, or other relevant outcomes. The follow-up data may be used to update themulti-omic dataset, retrain the machine learning algorithms, adjust the stratification, or perform other suitable updates to the system.
[0075] In some embodiments, the method involves employing machine learning algorithms to analyze the multi-omic dataset and identify biomarkers associated with a response to a therapy for treating IBD. The machine learning algorithms may include, but are not limited to, a support vector machine, a neural network, a decision tree, a random forest, or a gradient boosting machine. These algorithms may be selected based on their ability to handle high-dimensional data, their predictive accuracy, or other relevant factors. The machine learning algorithms may be implemented on a computing device, such as a server, a personal computer, a laptop, a tablet, or a smartphone, and may be executed using a suitable programming language, such as Python, R, Java, or C++.
[0076] In some cases, the identified biomarkers are validated using an independent patient dataset. This validation step may involve comparing the performance of the machine learning algorithms on the original dataset with their performance on the independent dataset, assessing the consistency of the identified biomarkers across different datasets, or other suitable validation techniques. The independent patient dataset may be obtained from a different cohort of IBD patients, a different geographical region, a different time period, or other suitable sources.
[0077] In some embodiments, the patient population is stratified into phenotypic groups based on the identified biomarkers using unsupervised clustering. For example, the identified biomarkers may be inputted into a trained machine learning model (e.g., a classifier) that classifies a patient into a phenotypic group based on the biomarkers of the patient. The unsupervised clustering may be performed using at least one of k-means clustering, hierarchical clustering, or Gaussian mixture models. These clustering techniques may be selected based on their ability to identify distinct groups within the patient population, their computational efficiency, or other relevant factors. The unsupervised clustering may be implemented using a suitable software package, such as the scikit-leam library in Python, the hclust function in R, or other suitable tools.
[0078] In some cases, a patient population is defined that is predicted to respond to the therapy based on the stratification and further analysis of patient metadata. The patient metadata may include at least one of age, gender, disease duration, prior medication history, orlifestyle factors. This metadata may be used to refine the stratification, identify potential confounding factors, or provide additional context for interpreting the identified biomarkers.
[0079] In some cases, the method further comprises the step of monitoring the patient population for changes in the biomarkers over time. This monitoring step may involve repeated collection and analysis of biological samples from the patients, tracking changes in the identified biomarkers, or other suitable monitoring techniques. The results of this monitoring step may be used to adjust the therapy, refine the stratification, or provide additional insights into the progression of IBD.
[0080] In some embodiments, the identified biomarkers (e.g., biomarkers associated with a cluster described herein) are further analyzed to determine their genetic pathways and interactions. This analysis may involve mapping the biomarkers onto known genetic pathways, identifying potential interactions between the biomarkers, or other suitable analysis techniques. The results of this analysis may provide additional insights into the mechanisms of action of the therapy, the pathophysiology of IBD, or other relevant aspects.
[0081] In some cases, the method further comprises the step of providing personalized treatment recommendations for each stratified phenotypic group. These recommendations may be based on the predicted response of the phenotypic group to the therapy, the specific characteristics of the phenotypic group, or other relevant factors.
[0082] In some embodiments, the stratification is used to exclude patients from the patient population who are predicted not to respond to the drug. This exclusion step may involve identifying patients who fall into phenotypic groups with a low predicted response to the therapy, removing these patients from the patient population, or other suitable exclusion techniques. The results of this exclusion step may be used to focus the therapy on patients who are likely to benefit, potentially improving the overall effectiveness of the therapy and reducing unnecessary treatment.
[0083] In some embodiments, the method involves stratifying the patient population into phenotypic groups based on the identified biomarkers using unsupervised clustering. Unsupervised clustering is a machine learning technique that groups data points based on their similarity without any prior knowledge of the data labels. This technique can be particularly useful in the context of precision medicine, as it allows for the identification of distinct patient subgroups that may exhibit different responses to a given therapy.
[0084] In some cases, the unsupervised clustering may be performed using at least one of k-means clustering, hierarchical clustering, or Gaussian mixture models. K-means clustering is a method that partitions the data into k distinct clusters based on the distance between data points and the centroid of the cluster. Hierarchical clustering, on the other hand, creates a treelike model of the data where each data point is linked to its nearest neighbors. Gaussian mixture models are a probabilistic model that assumes all the data points are generated from a mixture of a finite number of Gaussian distributions with unknown parameters.
[0085] In certain embodiments, the method involves accessing a multi-omic dataset comprising at least two of a genomic profile of patient data, a transcriptomic profile of patient data, or a proteomic profile of patient data. The multi-omic dataset may be obtained from biological samples collected from subjects. These biological samples may be selected from the group of blood, stool, tissue biopsy (e.g. intestine biopsy), saliva, or a combination thereof.
[0086] The genomic profile may include data derived from whole exome sequencing, whole genome sequencing, or global screening array (e.g. whole genome SNP arrays). In some cases, VCFtools may be used to process the genomics data. VCFtools may reduce the raw data to biallelic sites with a minor allele frequency of at least 0.05. VCFtools is a software package designed for working with Variant Call Format (VCF) files, which are commonly used to store genetic variation data. This versatile toolkit provides a wide range of functions for manipulating, filtering, and analyzing VCF files. VCFtools can perform tasks such as merging or comparing multiple VCF files, extracting specific regions or variants of interest, calculating various population genetic statistics, and converting VCF data to other file formats. In genomics research, VCFtools may be used to process raw sequencing data, reducing it to biallelic sites with a specified minor allele frequency threshold. This preprocessing step helps to focus analyses on more informative genetic variants and can improve the efficiency of downstream computational tasks. VCFtools is particularly valuable in large-scale genomics studies, where it can handle the processing and analysis of data from thousands of individuals efficiently. The term “genomic profile” may refer to information identifying the presence / absence or variant allele fraction of one or more genetic variants in a sample (or a subject from which the sample has been previously obtained). The sample may be a blood sample, a plasma sample, a tissue biopsy, a saliva sample, a cell swab sample, etc. Generally, any sample from which germline genetic variation present in the subject can be identified may be used.
[0087] PyEnsembl may be employed for annotating the genomics data, using the Ensembl release of the human genome. PyEnsembl is a Python library that provides a convenient interface for accessing and querying genomic annotations from the Ensembl database. It allows researchers and bioinformaticians to programmatically retrieve information about genes, transcripts, exons, and other genomic features. In the context of the invention, PyEnsembl may be employed for annotating the genomics data, using a specific release of the Ensembl human genome database. This annotation process is crucial for interpreting the genetic variants identified in the multi-omic dataset and understanding their potential functional implications in inflammatory bowel disease (IBD).
[0088] In some embodiments, PyEnsembl may be used after the initial processing of genomic data with VCFtools. Once the genetic variants have been filtered and processed, PyEnsembl can be utilized to map these variants to specific genes or genomic regions, providing additional context for their potential biological significance. This annotation step enhances the interpretability of the genomic data, allowing researchers to connect genetic variations with known genes, regulatory elements, or other functional genomic features that may be relevant to IBD pathogenesis or treatment response. By integrating PyEnsembl into the analysis pipeline, the invention leverages up-to-date genomic annotations to improve the accuracy and biological relevance of the multi-omic data integration and subsequent patient stratification efforts.
[0089] The transcriptomic profile may include data derived from RNA sequencing of tissue samples. The tissue sample may be an intestine biopsy. In some cases, the transcriptomic data may be processed using normalization techniques such as DESeq2 normalization for within-batch normalization and ComBat-seq for across batch normalization. The transcriptomic profile may include data derived from transcriptomic data. Transcriptomic data may include RNA sequencing data, array based gene expression data, qRT-PCR data, or data from any other technology that can quantify the level of expression of a plurality of genes or transcripts. In embodiments, the transcriptomics data includes RNA sequencing data (RNA- seq). The term “transcriptomic profile” may refer to information that quantifies the level of expression of one or more genes or transcripts in a sample.
[0090] The proteomic profile may include data derived from blood (e.g. plasma) samples. In some cases, the proteomic data may be processed using normalization techniques appropriate for proteomic data. Proteomic data refers to data indicative of the level of one or more signature proteins in a sample. The level of a protein in a sample may be measured usingan affinity based assay, such as a multiplex affinity based assay, or any assay using antigen binding reagents (e.g. antibodies or fragments thereof) to specifically detect and / or quantify the presence of one or more proteins. The level of a protein in a sample may be measured using an ELISA assay, Luminex assay or an Olink Proximity Extension Assay (PEA). The term “proteomic profile” may refer to information that quantifies the level of one or more proteins in a sample.
[0091] In some embodiments, the multi-omic dataset further comprises microbiomic profiles of the patient data. These microbiomic profiles may be derived from stool samples and may provide information about the gut microbiome composition and function.
[0092] In some cases, the biological sample may be contacted with a set of probes capable of detecting a panel of biomarkers. The presence or absence of these biomarkers in the sample may then be detected. This process may involve various laboratory techniques such as polymerase chain reaction (PCR), enzyme-linked immunosorbent assay (ELISA), or mass spectrometry.
[0093] The term “omics feature” refers to individual data points of a multi-omics data set (e.g. individual data points of a genomics profile, a transcriptomic profile or a proteomic profile). These can be used as input to a machine learning model, and / or as input to a Multi- Omics Factor Analysis. Multi-Omics Factor Analysis (MOFA) may be used for integrating multiple omics data sources. MOFA may be an unsupervised method that models each omics dataset as a matrix factorization, sharing factors across datasets to capture common patterns. This integration may provide a holistic view of the molecular underpinnings of IBD. MOFA identifies latent factors that are weighted combinations of the omics features in the original omics data.
[0094] In some embodiments, metabolomics may be used as an additional omics data type for patient stratification. Metabolomics may provide information about small molecule metabolites in biological samples, offering insights into metabolic processes that may be altered in IBD.
[0095] The multi-omic dataset may be stored in a data storage unit, such as a cloud-based storage system or a local server. The data storage unit may be configured to receive and store updates to the multi-omic dataset, allowing for continuous updating and refinement of the dataset as new patient data is collected.
[0096] In some cases, the method for stratifying a patient population in Inflammatory Bowel Disease (IBD) treatment may employ a machine learning algorithm to analyze the multi- omic dataset and identify biomarkers associated with a response to a therapy for treating IBD. The machine learning algorithm may include at least one of a support vector machine, a neural network, a decision tree, a random forest, or a gradient boosting machine.
[0097] The machine learning algorithm may be a support vector machine (SVM). An SVM constructs a hyperplane or set of hyperplanes in a high-dimensional space, which may be used for classification, regression, or other tasks. The SVM may aim to find the hyperplane that has the largest distance to the nearest training data points of any class, as this may generally lead to better generalization of the classifier to unseen data.
[0098] In certain embodiments, SVM may be employed as one of the machine learning techniques to analyze the multi-omic dataset and identify biomarkers associated with response to IBD therapy. In this case, it may be used to distinguish between responders and nonresponders to specific IBD treatments based on their multi-omic profiles. SVM is particularly effective when dealing with high-dimensional data, making it well-suited for analyzing complex genomic, transcriptomic, and proteomic features. The algorithm can handle non-linear relationships by using kernel functions to map the data into higher-dimensional spaces where linear separation becomes possible. In the invention, SVM may be used alongside other machine learning algorithms to build robust predictive models for patient stratification, potentially improving the accuracy of identifying patient subgroups that are likely to respond to particular therapies.
[0099] In other embodiments, the machine learning algorithm may be a neural network. A neural network is composed of layers of interconnected nodes, each of which may perform a simple computation. The network may learn complex patterns in the data by adjusting the strengths of the connections between nodes based on the error of the network's predictions.
[0100] In some cases, the machine learning algorithm may be a decision tree. A decision tree algorithm breaks down a dataset into smaller subsets while simultaneously developing an associated decision tree. The final result is a tree with decision nodes and leaf nodes. Decision trees may be easy to interpret and may handle both numerical and categorical data.
[0101] In other cases, the machine learning algorithm may be a random forest. A random forest is an ensemble learning method that may construct multiple decision trees during training and may output the class that is the mode of the classes (for classification) or mean prediction(for regression) of the individual trees. Random forests may correct for decision trees' habit of overfitting to their training set.
[0102] In some embodiments, the machine learning algorithm may be a gradient boosting machine. Gradient boosting is a technique where new models may be added to correct the errors made by existing models. Models may be added sequentially until no further improvements can be made. Gradient boosting machines may combine weak learners into a single strong learner in an iterative fashion.
[0103] In some cases, the machine learning algorithm may be a gradient boosted decision tree model, such as XGBoost (extreme Gradient Boosting), CatBoost or AdaBoost. XGBoost is an implementation of gradient boosted decision trees designed for speed and performance. XGBoost may use more regularized model formalization to control over-fitting, which may give it better performance. XGBoost may be employed as a key machine learning algorithm for analyzing the multi-omic dataset and identifying biomarkers associated with IBD therapy response. XGBoost works by building an ensemble of weak prediction models, typically decision trees, in a sequential manner. Each new model aims to correct the errors made by the previous models, gradually improving the overall predictive performance.
[0104] In various embodiments, XGBoost may be used to classify patients into different response categories or to predict treatment outcomes based on their multi-omic profiles. The algorithm's ability to handle high-dimensional data, capture complex non-linear relationships, and provide feature importance rankings makes it particularly suitable for integrating diverse omics data types. In an exemplary embodiment, XGBoost is used to train a classifier that distinguishes between UC and CD patients, achieving high accuracy and providing insights into the most important features (biomarkers) for this classification task. This approach allows for the identification of key molecular signatures that may be predictive of disease subtype and potentially treatment response in IBD patients.
[0105] The machine learning algorithm may be employed to analyze the multi-omic dataset, which may include genomic, transcriptomic, and proteomic data. The algorithm may be trained on this dataset to identify patterns and relationships between the multi-omic features and the response to IBD therapy. For example, XGBoost (or another machine learning algorithm) may be used to train a classifier that distinguishes between patients that have different clinical outcome metrics (e.g. responsive or non-responsive to a therapy), achievinghigh accuracy and providing insights into the most important features (biomarkers) for this classification task.
[0106] In some embodiments, the machine learning algorithm may be trained using a supervised learning approach. In this approach, the algorithm is provided with a labeled dataset, where each data point may be associated with a known outcome (e.g., response or non-response to therapy). The algorithm may learn to predict the outcome based on the input features.
[0107] In other cases, the machine learning algorithm may be trained using an unsupervised learning approach. In this approach, the algorithm is provided with an unlabeled dataset and attempts to find inherent patterns or structures in the data without predefined outcomes.
[0108] The machine learning algorithm may be evaluated using techniques such as cross- validation, where the dataset may be divided into training and testing sets. The algorithm may be trained on the training set and its performance may be evaluated on the testing set. This process may be repeated multiple times with different partitions of the data to ensure robust performance.
[0109] In some embodiments, feature selection techniques may be applied before or during the machine learning process to identify the most relevant biomarkers. These techniques may include methods such as recursive feature elimination, Lasso regularization, or principal component analysis.
[0110] The output of the machine learning algorithm may be a set of biomarkers that are most strongly associated with the response to IBD therapy. Alternatively, the structure of the trained machine learning model may be analysed to identify said biomarkers. For example, methods such as feature importance analysis (e.g. Shapley values analysis) may be used to identify biomarkers (e.g. omics or multi-omic features) that are more significantly associated with the prediction made by the machine learning model, such as e.g. response to IBD. As another example, feature selection methods including recursive feature elimination, and regularization (e.g. Lasso regularization) can be used to train the machine learning model, thereby identifying a reduced set of biomarkers (e.g. omics or multi-omic features) that are more significantly associated with the prediction made by the machine learning model, such as e.g. response to IBD. These biomarkers may then be used to stratify the patient population into phenotypic groups that may be more or less likely to respond to the therapy.
[0111] In some cases, the method for stratifying a patient population in Inflammatory Bowel Disease (IBD) treatment may involve identifying biomarkers associated with a response to IBD therapy. The identification of these biomarkers may be achieved through the analysis of a multi-omic dataset using machine learning algorithms.
[0112] The machine learning algorithm may analyze the multi-omic dataset, which may include genomic, transcriptomic, and proteomic data, to identify patterns and relationships between the multi-omic features and the response to IBD therapy. In some cases, the algorithm may employ feature selection techniques to identify the most relevant biomarkers. These techniques may include methods such as recursive feature elimination, Lasso regularization, or principal component analysis.
[0113] The panel of biomarkers identified by the machine learning algorithm may comprise at least two biomarkers from a specific group. This group may include, but may not be limited to, IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, ABHD5, CIDEC, RPS26, TMEM25, LY96, IL12B, BCL2, ANGPTL3, EPPK1, FGF19, INSL5, FKBP7, PILRA, CCL20, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P-STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L-DT, and NRG3.
[0114] In some cases, the panel of biomarkers may comprise at least one biomarker from a specific subgroup. This subgroup may include, but may not be limited to, RPS26, TMEM25, ANGPTL3, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P-STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L-DT, and NRG3.
[0115] The identified biomarkers may be validated using an independent patient dataset. This validation process may involve applying the machine learning algorithm to a separate dataset and comparing the results to those obtained from the original dataset. The validation may help ensure the robustness and generalizability of the identified biomarkers.
[0116] In some cases, ANOVA and chi-square tests may be applied to identify characteristic features of patient subpopulations. ANOVA may be used for continuous variables, such as gene expression levels or protein concentrations, while chi-square tests may be used for categorical variables, such as the presence or absence of specific genetic variants.
[0117] The identified biomarkers may comprise at least two of the biomarkers selected from a specific group. This group may include, but may not be limited to, IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, ABHD5, CIDEC, RPS26, TMEM25, LY96, IL12B, BCL2, ANGPTL3, EPPK1, FGF19, INSL5, FKBP7, PILRA, CCL20, IFT43-TTLL5,FAM9B, SLC24A3, CASTOR3P-STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L- DT, and NRG3.
[0118] In some cases, GSEApy may be used for pathway enrichment analysis. GSEApy is a Python implementation of Gene Set Enrichment Analysis (GSEA). This analysis may help identify biological pathways that may be overrepresented in the set of identified biomarkers, providing insights into the underlying biological mechanisms of IBD and treatment response. For example, GSEA may be used to identify one or more biological pathways that are overrepresented in omics features (e.g. genes and / or proteins) that have high weights in one or more multi-omics latent factors of interest. For example, omics features that have highest absolute weights in a factor of a MOFA analysis may be extracted and GSEA may be applied to this subset to identify pathways that are enriched in this subset relative to the complete set of omics features in the multi-omics data.
[0119] The enrichment score for a sample may be determined using a scoring algorithm that weights biomarkers. This algorithm may assign different weights to each biomarker based on its predictive value or biological significance. The weighted scores may then be combined to produce an overall enrichment score for the sample. In some embodiments, determining an enrichment score comprises determining a Gene Set Enrichment Analysis (GSEA) score of a sample compared to a control sample (e.g., from a healthy subject, from a subject that responded to pharmaceutical treatment of an inflammatory bowel disorder, or from a subject that did not respond to pharmaceutical treatment of an inflammatory bowel disorder) based on a set of biomarkers (e.g., a set of biomarkers described herein). In some cases, the biomarker identification process may involve multiple iterations and refinements. The machine learning algorithm may be retrained with new data or adjusted parameters to improve its performance in identifying relevant biomarkers. The identified biomarkers may be further validated through experimental studies or clinical trials to confirm their association with IBD treatment response.
[0120] In some cases, the method for stratifying a patient population in Inflammatory Bowel Disease (IBD) treatment may involve unsupervised clustering techniques to group patients based on identified biomarkers. Unsupervised clustering may be performed using at least one of k-means clustering, hierarchical clustering, or Gaussian mixture models.
[0121] K-means clustering partitions the patient population into a predetermined number of clusters based on the similarity of their biomarker profiles. The algorithm may iteratively assign patients to clusters and update cluster centroids until convergence. In some cases, theoptimal number of clusters may be determined using techniques such as the elbow method or silhouette analysis.
[0122] Hierarchical clustering creates a tree-like structure of nested clusters, either in a top-down (divisive) or bottom-up (agglomerative) approach. This method may not require a pre- specified number of clusters and may provide a dendrogram visualization of the patient groupings. In some cases, the final number of clusters may be determined by cutting the dendrogram at a specific level.
[0123] Gaussian mixture models assume that the patient population consists of a mixture of Gaussian distributions. This probabilistic approach may allow for soft clustering, where patients may have a probability of belonging to multiple clusters. In some cases, the optimal number of components (clusters) may be determined using criteria such as the Bayesian Information Criterion (BIC) or Akaike Information Criterion (AIC).
[0124] The unsupervised clustering may be applied to the biomarkers identified through the analysis of the multi-omic dataset. These biomarkers may include genomic, transcriptomic, and proteomic features associated with IBD and treatment response. In some cases, dimensionality reduction techniques such as Principal Component Analysis (PCA) or t- Distributed Stochastic Neighbor Embedding (t-SNE) may be applied prior to clustering to improve computational efficiency and visualization.
[0125] The clustering analysis may reveal distinct phenotypic groups within the patient population. In some cases, two distinct inflammation phenotypes may be identified in Crohn's disease patients based on factor analysis. These phenotypes may be characterized by different patterns of biomarker expression and may potentially respond differently to treatment.
[0126] In some cases, the distribution of Human Leukocyte Antigen (HLA) genes may be analyzed across the identified patient subgroups. For example, HLA gene expression may be analyzed in the transcriptomic and / or proteomic data. HLA genes play a crucial role in immune function and may be associated with IBD susceptibility and treatment response. The analysis of HLA gene distribution (e.g. differences in gene and / or protein expression for HLA genes between patient subgroups) may provide insights into the immunological characteristics of different patient subgroups.
[0127] The unsupervised clustering approach can identify of patient subgroups without relying on predefined clinical categories. This data-driven stratification can reveal novel patient subgroups that may not be apparent through traditional clinical classification methods. In somecases, the identified subgroups may be further characterized by analyzing their clinical characteristics, treatment outcomes, and other relevant metadata.
[0128] The results of the unsupervised clustering can be validated using various techniques. In some cases, internal validation measures such as the Calinski-Harabasz index or the Davies-Bouldin index may be used to assess the quality of the clustering. External validation may involve comparing the identified subgroups to known clinical subtypes or treatment response patterns.
[0129] The stratification of the patient population into phenotypic groups based on unsupervised clustering of biomarker data can provide a foundation for personalized treatment approaches in IBD. In some cases, the identified subgroups can be used to guide treatment decisions, predict treatment response, or design targeted clinical trials.
[0130] In some cases, the patient population is stratified into phenotypic groups based on the identified biomarkers using unsupervised clustering techniques. The unsupervised clustering may be performed using at least one of k-means clustering, hierarchical clustering, or Gaussian mixture models. These clustering techniques may group patients with similar biomarker profiles together, potentially revealing distinct subpopulations within the broader IBD patient population.
[0131] The pattern of biomarkers detected in a patient's sample may be analyzed to determine an enrichment score. This enrichment score may provide a quantitative measure of the patient's biomarker profile and may be used to predict the patient's response to a treatment regimen. In some cases, the enrichment score may be calculated using a weighted algorithm that takes into account the relative importance of each biomarker in predicting treatment response.
[0132] A patient population predicted to respond to a therapy may be defined based on the stratification results. This may involve identifying clusters or subgroups of patients with biomarker profiles associated with a positive response to the therapy. In some cases, the enrichment score may be compared to a threshold value to categorize a subject's response as either likely to respond or unlikely to respond to the therapy.
[0133] The stratification results may be used to provide personalized treatment recommendations for each phenotypic group. For example, patients in a subgroup characterized by high expression of one or more inflammatory markers may be recommendedfor more aggressive anti-inflammatory therapies, while patients in a subgroup with lower inflammatory marker expression may be recommended for milder treatments.
[0134] In some cases, the Simple Endoscopic Score for Crohn's Disease (SES-CD) is used to assess disease severity in conjunction with the biomarker-based stratification. The SES-CD may provide an additional clinical measure to complement the molecular data, potentially enhancing the accuracy of patient stratification and treatment recommendations.
[0135] The stratification may also be used to exclude subjects from the patient population who are predicted not to respond to a therapy. This approach can help focus treatment on patients who are more likely to benefit, potentially improving overall treatment efficacy and reducing unnecessary exposure to ineffective therapies.
[0136] Patient metadata, including at least one of age, gender, disease duration, prior medication history, or lifestyle factors, may be incorporated into the stratification process (e.g., see FIG. 6). These factors may provide additional context for interpreting the biomarker data and may help refine the phenotypic groupings.
[0137] In some cases, the enrichment score may be compared to a threshold value to categorize the subject's response. For example, an enrichment score below a certain threshold may indicate that the subject is more likely to respond to the therapy than a subject with an enrichment score above the threshold. This categorization may be used to guide treatment decisions in clinical practice.
[0138] The stratification approach may be applied in various clinical scenarios. For instance, in a newly diagnosed IBD patient, the biomarker profile may be used to predict the likely disease course and guide initial treatment selection. In a patient with established IBD who has failed multiple therapies, the stratification may help identify alternative treatment options that may be more effective based on the patient's molecular profile.
[0139] In some cases, the stratification results may be used to design more targeted clinical trials for new IBD therapies. By selecting patients with specific biomarker profiles, these trials may be able to demonstrate efficacy more efficiently and potentially lead to the development of more personalized treatment approaches.
[0140] The patient stratification approach may be continuously refined as new data becomes available. This may involve periodically re-analyzing the biomarker data, updating the clustering algorithms, and adjusting the enrichment score calculations to improve the accuracy of treatment response predictions.
[0141] In some cases, the method for stratifying a patient population in Inflammatory Bowel Disease (IBD) treatment may involve the integration and interaction of multiple elements. The workflow may begin with the input of a multi-omic dataset, which may comprise genomic, transcriptomic, and proteomic data from patient samples. This dataset may be preprocessed, e.g. using batch effect correction, normalization, etc., to ensure data quality and consistency.
[0142] In some cases, the transcriptomic data may be subjected to batch effect correction using pyComBat. PyComBat may be an empirical Bayes method for removing batch effects in sequencing count data. The method may standardize the data across genes, estimate batch and condition effects using a generalized linear model, and adjust the data to remove batch effects while preserving biological variation.
[0143] A machine learning algorithm may then be employed to analyze the (optionally preprocessed) multi-omic dataset. This algorithm may identify biomarkers associated with a response to a therapy for treating IBD. The biomarker identification process may involve feature selection techniques and may result in a panel of biomarkers that are predictive of treatment response.
[0144] In some cases, unsupervised clustering techniques are applied to the identified biomarkers to stratify the patient population into phenotypic groups. These techniques may include k-means clustering, hierarchical clustering, or Gaussian mixture models. The resulting stratification may reveal distinct subgroups of patients with similar molecular profiles.
[0145] Based on the stratification results, a patient population predicted to respond to a therapy may be defined. This may involve identifying clusters or subgroups of patients with biomarker profiles associated with a positive response to the therapy. In some cases, the therapy may be administered to a subject or recommended for administration to a subject of the patient population based on this stratification. The therapy may comprise a compound, a therapeutic treatment, or a combination thereof.
[0146] The method may also include monitoring the patient population for changes in the biomarkers over time. This monitoring process may involve collecting and analyzing biological samples from patients at regular intervals. The results of this monitoring may be used to track disease progression, assess treatment efficacy, and inform decisions about treatment adjustments.
[0147] In some cases, the subject's response to the therapy may be re-evaluated by repeating the steps of the method at subsequent time points. This re-evaluation process may involve collecting new biological samples or receiving multi-omic data derived from biological samples previously obtained from the subject, analyzing the multi-omic data, identifying biomarkers, and performing unsupervised clustering or prediction with a machine learning model. The results of this re-evaluation may provide updated information about the patient's response to treatment and may guide further therapeutic decisions.
[0148] The treatment regimen may be adjusted based on the predicted response and the results of ongoing monitoring and re-evaluation. This adjustment may involve changes in dosage, frequency of administration, or even a switch to an alternative therapy if the current treatment is not producing the desired response.
[0149] In some cases, the integration and interaction of these elements may create a dynamic, iterative process for personalized IBD treatment. The continuous input of new data and re-evaluation of patient responses may allow for ongoing refinement of the stratification model and improvement in treatment outcomes. This approach may enable a more precise and adaptive method for managing IBD, potentially leading to improved patient outcomes and more efficient use of therapeutic resources.
[0150] With reference to FIG. 1, in certain embodiments, a comprehensive workflow for integrating and analyzing multi-omics data from IBD patients was described. The workflow as illustrated leverages data from the Study of a Prospective Adult Research Cohort with IBD (SPARC IBD), a component of the Crohn's & Colitis Foundation's IBD Plexus research platform (although other sources and cohorts may be used). More detail on this workflow can be found in Preto et al. (Multi-omics data integration identifies novel biomarkers and patient subgroups in inflammatory bowel disease, Journal of Crohn's and Colitis, Volume 19, Issue 1, January 2025), which is incorporated herein by reference in its entirety.
[0151] The workflow as illustrated begins with the collection of biological samples from IBD patients, including plasma samples for genomics and proteomics analysis, and intestinal biopsy samples for transcriptomics analysis. These samples are processed using appropriate experimental techniques to generate raw data for each omics modality. For the purpose of this work, the full cohort are subset to samples of patients characterized with three omics modalities: genomics (specifically, Illumina Infinium Global Screening Array and wholeexome sequencing), transcriptomics (specifically, RNA sequencing), and proteomics (specifically, Olink® Proximity Extension Assay (PEA)).
[0152] Table 1 summarizes the different batches across these three omics modalities for patients with all three omics modalities available.
[0153] In some aspects, the genomics data may be derived from whole exome sequencing or genome-wide association studies. The transcriptomics data may be obtained from RNA sequencing of intestinal biopsy samples. The proteomics data may be generated from blood plasma samples using mass spectrometry-based techniques, or affinity based assays (e.g. Olink proteomics assays). In embodiments, the genomics data may comprise information identifying the presence of one or more SNPs identified in one or more genome-wide association studies. For example, the one or more genome-wide association studies (GWAS) may be GWAS identifying SNPs that are significantly associated with one or more of: a risk of developing IBD, a risk of developing an inflammatory disease, a risk of developing ulcerative colitis, a risk of developing Crohn’s disease, etc.
[0154] In some embodiments, the raw data from each omics modality then underwent preprocessing and quality control steps. For genomics data, this may involve variant calling and filtering. Transcriptomics data may be normalized and batch-corrected. Proteomics data may be processed to remove low-quality features and normalize intensities.
[0155] In some embodiments, the preprocessed data from each omics modality may be integrated into a multi-omics dataset. This integration step may involve matching samplesacross modalities based on patient identifiers and collection dates. The integrated dataset may then be used for downstream analyses.
[0156] The workflow as illustrated in FIG. 1 includes two main analysis branches: patient stratification and machine learning classification. For patient stratification, the multi-omic data was integrated using MOFA and the integrated multi-omics data was analyzed separately for Crohn's disease (CD) and ulcerative colitis (UC) patients. This stratification involved unsupervised clustering techniques to identify distinct patient subgroups based on their molecular profiles.
[0157] In parallel, a machine learning classifier was trained on the integrated multi-omics data to distinguish between CD and UC samples. This classifier may utilize algorithms such as random forests, support vector machines, or gradient boosting machines. In the illustrated embodiment, the classifier uses a model based on an ensemble of decision trees, such as specifically a gradient boosted decision tree model (e.g. XGBoost). Any other classifier known in the art may be used.
[0158] In some embodiments, the results from both the patient stratification and machine learning classification analyses can be used to gain insights into IBD patient subgroups, identify potential biomarkers, and elucidate disease mechanisms. These insights can inform the development of personalized treatment strategies and improve patient outcomes.
[0159] In some cases, the workflow may be iterative, with new data being incorporated as it becomes available, allowing for continuous refinement of the patient stratification and classification models. The workflow may also be adaptable to include additional omics modalities or clinical data as they become available.
[0160] This multi-omics data integration and analysis workflow provides a comprehensive approach to understanding the molecular basis of IBD and may contribute to the development of precision medicine strategies for IBD treatment.
[0161] In some embodiments, the method for stratifying a patient population in Inflammatory Bowel Disease (IBD) treatment includes correcting batch effects in multi-omics data analysis. The transcriptomics data from the SPARC IBD Cohort initially exhibited a strong batch effect that hinders the identification of biological patterns. This batch effect can be visualized using principal component analysis (PCA), as shown in Figures 2A-2C. In particular, PCA of the first 2 components of the transcriptomics samples is provided, colored by batch (FIG. 2A), tissue (FIG. 2B). and diagnosis (FIG. 2C).
[0162] To address this issue, a batch correction method was applied using ComBat, an empirical Bayes-based algorithm. Other batch correction methods may be used. This correction can be visualized using principal component analysis (PCA), as shown in Figures 2D-2F. PCA of the first 2 components of the transcriptomics samples after correcting for batch effect using batch variance and pyComBat is provided, colored by batch (FIG. 2D), tissue (FIG. 2E), and diagnosis (FIG. 2F). The samples may cluster primarily based on their batch rather than biological factors, with a high silhouette score on the PCA plot. After applying ComBat, the batch effect was drastically reduced, as evidenced by a decrease in the silhouette score (FIG. 2D). The effectiveness of the batch correction can be further corroborated by calculating the Mutual Information (MI) scores between the features and the batch. Prior to batch correction, a high percentage of the transcriptomics features may have MI scores above a certain threshold. After correction, this percentage may drop significantly, indicating a substantial reduction in batch-associated variation.
[0163] Following the batch effect correction, PCA revealed a slight association between the samples and tissue type (FIG. 2E). This association may be expected, given that transcription patterns are known to be tissue- specific. The separation between colon and small intestine samples became more apparent after batch correction, suggesting that the biological signal related to tissue type is preserved and enhanced.
[0164] In some aspects, the method includes using multi-omics signatures to differentiate between Crohn's Disease (CD) and Ulcerative Colitis (UC) patients. Specifically, a machine learning classifier, which was an XGBoost classifier in this specific example, was trained on the multi-omics dataset to predict whether a sample originated from a UC or CD patient. The classifier's performance was evaluated using a train-test split and cross-validation. The classifier achieved high accuracy in distinguishing between UC and CD samples, with metrics such as accuracy, precision, recall, and AUC-ROC being reported (FIG. 3A and 3B). Other machine learning models known in the art may be used.
[0165] Analysis of the most predictive features (e.g. using feature importance analysis) revealed several known and potentially novel biomarkers for differentiating between UC and CD (FIG. 3C). In the transcriptomics data, genes such as IGKV6D-21, GALNT6, TBX3, NRG4, and USP38 were found to be among the top predictors. In the proteomics data, several inflammation-related proteins were identified as important predictors, including INSL5, EGFR, IL12B, FGF19, LY96, CCL20, and BCL2. Any feature selection and / or feature importance analysis known in the art may be used.
[0166] In some embodiments, the method includes identifying molecular profiles correlated with disease severity in Ulcerative Colitis. Specifically, multi-omics samples from UC patients were analyzed using Multi-Omics Factor Analysis (MOFA). The factors from MOFA were then clustered using unsupervised clustering to identify characteristic molecular profiles correlated with clinical phenotypes such as disease severity and macroscopic appearance. A correlation was observed between a MOFA factor and disease severity based on endoscopy (FIG. 4A). This correlation suggests that the features associated with this factor can potentially be used to distinguish between patients with severe or moderate disease and patients with mild disease or in remission. Any other multi-omic feature integration and unsupervised clustering method known in the art may be used, such as e.g. a dimensionality reduction method such as PCA, t-SNE, or UMAP followed by unsupervised clustering.
[0167] Several potential biomarkers for disease severity were identified across the three omics types, using ANOVA for proteomic and transcriptomic features, and chi-square tests for genomic features, in each case testing for top 150 features (although any number of top features may be used) of the above MOFA factor significantly associated with severe disease. In genomics, certain alleles of genes including TAF6L, ZNF268, and CEP164 were found to be associated with disease severity (FIG. 4B). In proteomics, disease severity was found to correlate with levels of proteins including IL17A, EPO, and REGIB (Fig. 4D). Transcriptomics analysis reveal genes including EMILIN2 and DLD, which exhibit positive and inverse correlations with disease severity, respectively (FIG. 4F).
[0168] Pathway enrichment analysis was performed on the significant features for each omics modality independently. For proteomics, interleukin and cytokine signaling pathways were found to be among the enriched pathways (FIG. 4C). Transcriptomics analysis revealed enriched pathways related to metabolic processes, hemostasis, and response to pathogens (FIG. 4E). Any enrichment analysis method known in the art may be used, such as e.g. any gene set / pathway enrichment analysis method, including but not limited to GSEA and derivatives thereof.
[0169] In some aspects, the method includes investigating molecular phenotypes of inflamed samples in Crohn's Disease. Multi-omics samples from CD patients were analyzed using MOFA to identify molecular profiles correlated with clinical phenotypes. The analysis was focused on colon transcriptomics samples due to the strong tissue effect observed.
[0170] Hierarchical clustering was performed across all samples (all though any number of samples may be used) based on all the factors yielded by MOFA to assess whether any of the factors correlated with variables of clinical interest (FIG. 5A). Any other unsupervised clustering method may be used. A clear correlation was observed between a MOFA factor and macroscopic appearance (FIG. 5B). Non-inflamed samples were found to have values for this factor close to zero, while inflamed samples showed either highly positive or highly negative factor values, suggesting two distinct inflammation phenotypes.
[0171] Analysis of the top 150 features (although any number of top features may be used) in each omics modality for the above factor identified by MOFA (i.e. features with top 150 MOFA weights for this factor) revealed several markers differentiating between inflamed and normal samples, such as CXCL9, HLA-DRA, INHBB, SERPINA3, and NOS2 (FIG. 5D). These features show potential as biomarkers for intestinal inflammation in CD.
[0172] Pathway analysis of the characteristic features of the two inflamed sample subpopulations identified by clustering on MOFA factors revealed enriched pathways related to the innate immune system for one cluster of inflamed samples (e.g., natural killer cell proliferation and activation) (FIG. 51). The other cluster showed enriched pathways related to antigen processing and presentation, immunoglobulin production, and MHC class.
[0173] These exemplary embodiments demonstrate the potential of multi-omics integration to identify distinct molecular phenotypes of inflammation in IBD, which may provide insights into disease mechanisms and could potentially be used for patient stratification and personalized treatment approaches. Specifically, the approaches above were used to identify subpopulations of patients in integrated multi-omics data, classify CD and UC patients, and identify biomarkers that are associated with the subpopulations, biomarkers that are associated with the classification of CD vs UC, and biomarkers that are associated with disease severity. The data above and further described in the examples below therefore demonstrates that the approaches proposed can be applied to the development of stratification biomarkers for response to a treatment (since the data contains information indicative of disease severity and disease phenotype subgroups).
[0174] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the disclosure. Accordingly, other implementations are within the scope of the following claims.
[0175] Reference to “about” a value or parameter herein includes (and describes) embodiments that are directed to that value or parameter per se.For example, description referring to “about X” includes description of “X”. In some embodiments, the term “about” when used in association with a measurement, or used to modify a value, a unit, a constant, or a range of values, refers to variations of + / - 10%, 5%, 2%, or 1%.
[0176] Reference to “between” two values or parameters herein includes (and describes) embodiments that include those two values or parameters per se. For example, description referring to “between x and y” includes description of “x” and “y” per se.EXAMPLES
[0177] The presently disclosed subject matter will be better understood by reference to the following examples, which are provided as exemplary of the invention, and not by way of limitation.Example 1Batch Effect Correction in Multi-Omics Data Analysis for Inflammatory Bowel Disease
[0178] In this example, a method for correcting batch effects in multi-omics data analysis for inflammatory bowel disease (IBD) was demonstrated. The study utilized transcriptomics data from the SPARC IBD cohort, which included samples from patients with Crohn's disease (CD) and ulcerative colitis (UC).
[0179] Initially, the transcriptomics data exhibited a strong batch effect that was the predominant source of variation, hindering the identification of biological patterns. This batch effect was visualized using principal component analysis (PCA), as shown in Figure 2A. The samples clustered primarily based on their batch rather than biological factors, with a silhouette score of 0.28 on the PCA plot.
[0180] Figure 17B and 17C show the same PCA plot as Figure 2A, but with the data points colored by tissue type and diagnosis, respectively. In Figure 2B, the clustering by batch is still apparent, but some separation between colon and small intestine samples can be observed. Figure 2C demonstrates that the batch effect obscures any potential clustering by diagnosis (CD vs. UC).
[0181] To address this issue, a batch correction method was applied using ComBat, an empirical Bayes-based algorithm. After applying ComBat, the batch effect was drastically reduced, as evidenced by the decrease in the silhouette score from 0.28 to -0.05 (Figure 2D).
[0182] The effectiveness of the batch correction was further corroborated by calculating the Mutual Information (MI) scores between the features and the batch. Prior to batch correction, 22.28% of the transcriptomics features had MI scores above 0.2. After correction, this percentage dropped to 3.55%, indicating a significant reduction in batch-associated variation.
[0183] Following the batch effect correction, the PCA revealed a slight association between the samples and tissue type (Figure 2E). This association was expected, given thattranscription patterns are known to be tissue-specific, and most of the samples in the dataset were derived from colon and small intestine tissues. The separation between colon and small intestine samples became more apparent after batch correction, suggesting that the biological signal related to tissue type was preserved and enhanced.
[0184] Importantly, after batch correction, other variables such as diagnosis (CD vs. UC) did not show an apparent association in the PCA (Figure 2F). This suggested that the batch correction process did not introduce artificial differences between the disease subtypes, while still allowing for the detection of biologically relevant signals.
[0185] This example demonstrated the importance of addressing batch effects in multi- omics data analysis for IBD. By effectively correcting for these technical variations, the method enabled more reliable downstream analyses, including the identification of disease- associated biomarkers and patient stratification. The approach described here may be applicable to other multi-omics studies in IBD and other complex diseases, where batch effects can confound biological signal detection.Example 2Multi-omics Classification of Ulcerative Colitis and Crohn's Disease
[0186] In this example, a machine learning approach was used to differentiate between Ulcerative Colitis (UC) and Crohn's Disease (CD) using multi-omics data. The study utilized genomic, transcriptomic, and proteomic data from 1,023 patients (703 CD and 320 UC) in the SPARC IBD cohort.
[0187] An XGBoost classifier was trained on the multi-omics dataset to predict whether a sample originated from a UC or CD patient. The model's performance was evaluated using a 80 / 20 train-test split and 5-fold cross-validation. The classifier achieved high accuracy in distinguishing between UC and CD samples, with an overall accuracy of 0.77 ± 0.02 and an AUC-ROC of 0.75 ± 0.02 (Figure 3 A and 3B).
[0188] The model demonstrated better performance in identifying CD samples compared to UC samples. For CD, the precision was 0.77 ± 0.02 and recall was 0.87 ± 0.02. For UC, the precision was 0.75 ± 0.02 and recall was 0.57 ± 0.02 (Figure 3B).
[0189] Analysis of the most predictive features revealed several known and potentially novel biomarkers for differentiating between UC and CD (Figure 18C). In the transcriptomicsdata, genes such as IGKV6D-21, GALNT6, TBX3, NRG4, and USP38 were among the top predictors. These genes have been previously associated with inflammation or IBD. Interestingly, some genes like RPS26 and TMEM25 had not been clearly linked to IBD before, suggesting potential new avenues for investigation.
[0190] In the proteomics data, several inflammation-related proteins were identified as important predictors, including INSL5, EGFR, IL12B, FGF19, LY96, CCL20, and BCL2. These proteins have been previously associated with IBD or disease activity. The protein ANGPTL3 was also identified as a top predictor but had not been previously linked to IBD, warranting further investigation.
[0191] Notably, genomic features did not appear among the top predictors for distinguishing between UC and CD in this model.
[0192] This example demonstrated the potential of multi-omics data integration and machine learning approaches for differentiating between UC and CD. The identified biomarkers, both known and novel, may provide insights into the molecular differences between these two forms of IBD and could potentially be used to aid in diagnosis, particularly for patients with indeterminate colitis. Furthermore, this approach highlighted the value of combining different types of molecular data for improving disease classification and understanding.Example 3Identifying Molecular Profiles Correlated with Disease Severity in Ulcerative Colitis
[0193] In this example, multi-omics samples from ulcerative colitis (UC) patients were analyzed using Multi-Omics Factor Analysis (MOFA) to identify characteristic molecular profiles correlated with clinical phenotypes such as disease severity and macroscopic appearance. The goal was to identify biomarkers that could be used to define patient subpopulations.
[0194] The analysis revealed a correlation (RA2 = 0.40) between factor 1 from MOFA and disease severity based on endoscopy (Figure 4A). This correlation was further corroborated by a Kendall's Tau test, which showed a moderate yet significant correlation of 0.30 (p-value = 8xlOA-4). This suggested that the features associated with factor 1 could potentially distinguish between patients with severe or moderate disease from those with mild disease or in remission.
[0195] To investigate this further, the top 150 features across the three omics modalities were selected based on their weights for factor 1. ANOVA was applied to proteomics and transcriptomics features, while chi-square analysis was used for genomics features (p-value < 0.01), comparing severe cases against all others.
[0196] Several potential biomarkers for disease severity were identified across the three omics types. In genomics, all severe patients were found to have the reference allele (0) of TAF6L (chrl l_bp62771388) (Figure 4B), while severe patients typically had the first alternate allele (1) for ZNF268 (chrl2_bpl33202004) and CEP164 (chrl l_bp 117403235).
[0197] In proteomics, disease severity was found to correlate with levels of IL17A (Figure 4D), EPO, and REGIB. Transcriptomics analysis revealed several genes, including EMILIN2 and DLD, which exhibited positive and inverse correlations with disease severity, respectively (Figure 4F).
[0198] Pathway enrichment analysis was performed on the significant features for each omics modality independently. For proteomics, interleukin and cytokine signaling pathways were among the enriched pathways (Figure 4C). Transcriptomics analysis revealed enriched pathways related to metabolic processes, hemostasis, and response to pathogens (Figure 4E).
[0199] This example demonstrated the potential of multi-omics integration to identify molecular signatures associated with disease severity in UC. The identified biomarkers and pathways may provide insights into disease mechanisms and could potentially be used for patient stratification and personalized treatment approaches in UC.Example 4Investigating Molecular Phenotypes of Inflamed Samples in Crohn's Disease
[0200] In this example, multi-omics samples from Crohn's disease (CD) patients were analyzed using Multi-Omics Factor Analysis (MOFA) to identify molecular profiles correlated with clinical phenotypes. The study focused on colon transcriptomics samples due to the strong tissue effect observed.
[0201] Hierarchical clustering was performed across all samples based on the factors yielded by MOFA to assess correlations with variables of clinical interest (Figure 5A). Factor 3, which showed relevant contributions from multiple omics modalities, exhibited a clear correlation with macroscopic appearance (Figure 5B). Non-inflamed samples had factor 3values close to zero, while inflamed samples showed either highly positive or highly negative factor 3 values, suggesting two distinct inflammation phenotypes.
[0202] The distribution of factor 3 values across the three main identified clusters revealed two subpopulations of inflamed samples (Figure 5C). Clusters 1 and 3 were combined as cluster A, while cluster 2 was designated as cluster B, based on their similar inflammation patterns.
[0203] Analysis of the top 150 features in each omics modality based on the absolute weights of factor 3 revealed several markers differentiating between inflamed and normal samples, such as CXCL9, HLA-DRA, INHBB, SERPINA3, and NOS2 (Figure 5D) (see also FIGs 7-8). These features showed potential as biomarkers for intestinal inflammation in CD.
[0204] ANOVA was applied to the top proteomics and transcriptomics features, and chi- square analysis was used for genomics features (p-value < 0.01) to identify characteristics of the two inflamed subpopulations. Cluster B was found to have a large majority of Human Leukocyte Antigen (HLA) genes among its characteristic transcripts (Figure 5G). Additionally, the TSBP1-AS1 transcript, characteristic of cluster B, corresponded to the most occurring SNPs in the genomics data (Figure 5H).
[0205] Pathway analysis of the characteristic features revealed enriched pathways related to the innate immune system for cluster A (e.g., natural killer cell proliferation and activation) (Figure 51). Cluster B showed enriched pathways related to antigen processing and presentation, immunoglobulin production, and MHC class.
[0206] This example demonstrated the potential of multi-omics integration to identify distinct molecular phenotypes of inflammation in CD. The results suggested that the differentiation between the two subpopulations of inflamed samples was due to cluster B developing a more robust adaptive immune response characterized by high expression of HLA genes and other inflammation-related pathways. These findings may provide insights into disease mechanisms and could potentially be used for patient stratification and personalized treatment approaches in CD. Exemplary biomarkers upregulated in Cluster B are shown in FIG. 8. Cluster B also includes the following features ANG, CD79B, CD80, CLEC4G, CNDP1, CNTN3, FAM20A, GZMB, HLA-E, HTR1A, INHBB, MANSC1, MZB1, NOS2, PRG2, PRG3, PTGDS, RGMA, SERPINA3, TNFRSF8, TTR, A1BG-AS1, ADAMTS10, ADARB1, ADCYAP1R1, ADGRA2, AMOTL1, AP5M1, ARHGAP44-AS1, ARHGAP44, ARL10, BHLHE22, BOC, CAND2, CCDC80, CDH17, CEROX1, CH25H, CKMT1A,CLSTN3, C0L14A1, CSRP2, CX3CL1, CXADR, DAZL, DCLK1, DLG4, DTX3, DZIP1, EBF1, EBB, EFS, ELN, EPB41L4B, EPS8, EVC2, FBLN1, FBXL7, FGFR1, FLU, FLRT2, FZD10-AS1, GEM, GLI2, GLB, GLIS2, GPD1L, GRID1, GRIK5, HHLA2, HHLA3, HIGD1A, HLA-A, HLA-B, HLA-C, HLA-DMB, HLA-DPA1, HLA-DPB1, HLA-DPB2, HLA-DQA1, HLA-DQA2, HLA-DQB1-AS1, HLA-DQB1, HLA-DQB2, HLA-DRA, HLA- DRB1, HLA-DRB5, HLA-DRB6, HLA-DRB9, HLA-E, HLA-F, HLA-H, HLA-L, HLA-U, HOOK1, INPP4A, IRAG1, KCND1, KIAA1755, KRT20, LCNL1, LETR1, LGR4, LIMA1, LIMS2, LIX1L, LRRC19, LRRC4B, LSAMP, LTBP3, MAGI2-AS3, MAP3K12, MAP3K3, MAPK10, MFAP4, MIR100HG, MN1, MOXD1, MPC2, MPDZ, MRPL35, MYO16, MYO5A, NCKAP5L, NR2F2-AS1, OCLN, PACS1, PALD1, PHC1, PKIB, PTGDS, PTPRS, RAB3IL1, RBPMS, RECK, RSPO1, RSPO3, RUNX1T1, SDC3, SH3BGRL2, SLC24A3, SLC35D1, SOBP, ST8SIA1, STARD9, SYNGAP1, TCF7L1, TGFB1I1, TMCC2, TMEM200B, TMEM87B, TNFSF12, TNS2, TRERF1, TSBP1-AS1, TSPAN3, TTYH2, UGP2, WASL-DT, WASL, XK, ZCCHC24, ZEB2, ZNF154, ZNF532, which are described in Preto, Antonio Jose, et al. "Multi-omics data integration identifies novel biomarkers and patient subgroups in inflammatory bowel disease." Journal of Crohn's and Colitis (2025): jjael97, which is incorporated by reference in its entirety (definitions in this disclosure control over definitions in references incorporated by reference).
[0207] Additional Embodiments,1. A method for stratifying a patient population in Inflammatory Bowel Disease (IBD) treatment, the method comprising: accessing a multi-omic dataset comprising at least two of a genomic profile of patient data, a transcriptomic profile of patient data, or a proteomic profile of patient data; employing a machine learning algorithm to analyze the dataset and identify biomarkers associated with a response to a therapy for treating IBD; stratifying the patient population into phenotypic groups based on the identified biomarkers using unsupervised clustering; and defining a patient population predicted to respond to a therapy based on the stratification.2. The method of embodiment 1, wherein the machine learning algorithm includes at least one of a support vector machine, a neural network, a decision tree, a random forest, or a gradient boosting machine.3. The method of embodiment 1, further comprising validating the identified biomarkers using an independent patient dataset.4. The method of embodiment 1, wherein the unsupervised clustering is performed using at least one of k-means clustering, hierarchical clustering, or Gaussian mixture models.5. The method of embodiment 1, wherein the patient metadata includes at least one of age, gender, disease duration, prior medication history, or lifestyle factors.6. The method of embodiment 1, further comprising administering the therapy to a subject of the patient population based on the stratification, wherein the therapy comprises a compound, a therapeutic treatment, or a combination thereof.7. The method of embodiment 1, further comprising the step of monitoring the patient population for changes in the biomarkers over time.8. The method of embodiment 1, wherein the multi-omic dataset further comprises microbiomic profiles of the patient data.9. The method of embodiment 1, wherein the identified biomarkers comprise at least two of the biomarkers selected from the group of IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, ABHD5, CIDEC, RPS26, TMEM25, LY96, IL12B, BCL2, ANGPTL3, EPPK1, FGF19, INSL5, FKBP7, PILRA, CCL20, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P- STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L-DT, and NRG3.10. The method of embodiment 1, further comprising the step of providing personalized treatment recommendations for each stratified phenotypic group.11. The method of embodiment 1, wherein the stratification is used to exclude subjects from the patient population who are predicted not to respond to a therapy.12. A method of predicting a response to a therapy for an inflammatory bowel disease (IBD) in a subject in need thereof, the method comprising: a. obtaining a biological sample from the subject; b. contacting the biological sample with a set of probes capable of detecting a panel of biomarkers, wherein the panel of biomarkers comprises at least five biomarkers selected from the group of regenerating family member 1 beta (REGIB), fatty acid binding protein 6 (FABP6), regenerating family member 1 alpha (REGIA), major histocompatibility complex, class II, DQ beta 1 (HLA-DQB1), major histocompatibility complex, class II, DQ alpha 1 (HLA-DQA1), complement factor I (CFI), serpin family A member 1 (SERPINA1), indoleamine 2,3-dioxygenase 1 (IDO1), sodium channel epithelial 1 beta subunit (SCNN1B), deleted in malignant brain tumors 1 (DMBT1), suppressor of cytokine signaling 3 (SOCS3),guanylate binding protein 4 (GBP4), C-X-C motif chemokine ligand 1 (CXCL1), CXCL5, CXCL9, CXCL10, CXCL11, dual oxidase 2 (DU0X2), apolipoprotein Al (APOA1), carbonic anhydrase 4 (CA4), ubiquitin D (UBD), guanylate binding protein 1 (GBP1), interferon induced protein with tetratricopeptide repeats 3 (IFIT3), t-box 3 (TBX3), transglutaminase 2 (TGM2), vanin 1 (VNN1), peptidase inhibitor 3 (PI3), complement C3 (C3), cytochrome P450 family 2 subfamily B member 7, pseudogene (CYP2B7P), C-C motif chemokine ligand 20 (CCL20), lipocalin 2 (LCN2), major histocompatibility complex, class II, DP alpha 1 (HLA-DPA1), interleukin 13 receptor subunit alpha 2 (IL13RA2), Immunoglobulin Kappa Variable 6D-21 (IGKV6D-21), IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, ABHD5, CIDEC, RPS26, TMEM25, LY96, IL12B, BCL2, ANGPTL3, EPPK1, FGF19, INSL5, FKBP7, PILRA, CCL20, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P-STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L-DT, and NRG3; c. detecting the presence or absence of the biomarkers in the sample; d. analyzing the pattern of biomarkers detected to determine an enrichment score for the sample; and e. predicting the subject's response to the treatment regimen based on the enrichment score, wherein an enrichment score less than zero indicates that the subject is more likely to respond to the therapy than a subject with an enrichment score greater than zero.13. The method of embodiment 12, wherein the panel of biomarkers comprises at least two of the biomarkers selected from the group of IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, ABHD5, CIDEC, RPS26, TMEM25, LY96, IL12B, BCL2, ANGPTL3, EPPK1, FGF19, INSL5, FKBP7, PILRA, CCL20, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P- STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L-DT, and NRG3.14. The method of embodiment 13, wherein the panel of biomarkers comprises at least one of the biomarkers selected from the group of RPS26, TMEM25, ANGPTL3, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P-STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L- DT, and NRG3.15. The method of embodiment 12, further comprising comparing the enrichment score to a threshold value to categorize the subject's response as responsive or non-responsive to the therapy.16. The method of embodiment 12, wherein the biological sample is selected from the group of blood, stool, tissue biopsy, saliva, or a combination thereof.17. The method of embodiment 12, wherein the therapy includes administration of a compound, a therapeutic treatment, or a combination thereof.18. The method of embodiment 12, further comprising the step of adjusting the treatment regimen based on the predicted response.19. The method of embodiment 12, wherein the enrichment score is determined using a scoring algorithm that weights the biomarkers based on their predictive value.20. The method of embodiment 12, further comprising the step of re-evaluating the subject's response to the therapy by repeating steps (a) through (e) at subsequent time points.
[0208] An illustrative implementation of a computer system 900 that may be used in connection with any of the embodiments of the technology described herein (e.g., such as the process of FIG. 1) is shown in FIG. 9. The computer system 900 includes one or more processors 904 and one or more articles of manufacture that comprise non-transitory computer- readable storage media (e.g., memory 910 and one or more non-volatile storage media 906). The processor 904 may control writing data to and reading data from the memory 910 and the non-volatile storage device 906 in any suitable manner, as the aspects of the technology described herein are not limited to any particular techniques for writing or reading data. To perform any of the functionality described herein, the processor 904 may execute one or more processor-executable instructions stored in one or more non-transitory computer-readable storage media (e.g., the memory 910), which may serve as non-transitory computer-readable storage media storing processor-executable instructions for execution by the processor 904.
[0209] Computer system device 900 may also include a network input / output (I / O) interface 902 via which the computer system may communicate with other computing devices (e.g., over a network), and may also include one or more user VO interfaces 908, via which the computer system may provide output to and receive input from a user. The user VO interfaces 908 may include devices such as a keyboard, a mouse, a microphone, a display device (e.g., a monitor or touch screen), speakers, a camera, and / or various other types of VO devices.
[0210] The above-described embodiments can be implemented in any of numerous ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor (e.g., a microprocessor) or collection of processors, whether provided in a single computing device or distributed among multiple computing devices. Further, it should be appreciated that a computer may be embodied in any of a number of forms, such asa rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer, as nonlimiting examples. Additionally, a computer may be embedded in a device not generally regarded as a computer but with suitable processing capabilities, including a Personal Digital Assistant (PDA), a smartphone, a tablet, or any other suitable portable or fixed electronic device.
[0211] In this respect, it should be appreciated that one implementation of the embodiments described herein comprises at least one non-transitory computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or other tangible, non-transitory computer-readable storage medium) encoded with a computer program (i.e., a plurality of executable instructions) that, when executed on one or more processors, performs the abovedescribed functions of one or more embodiments (e.g., part of or all of the processes described above with reference to FIG. 1). The computer-readable medium may be transportable such that the program stored thereon can be loaded onto any computing device to implement aspects of the techniques described herein. In addition, it should be appreciated that the reference to a computer program which, when executed, performs any of the above-described functions, is not limited to an application program running on a host computer. Rather, the terms computer program and software are used herein in a generic sense to reference any type of computer code (e.g., application software, firmware, microcode, or any other form of computer instruction) that can be employed to program one or more processors to implement aspects of the techniques described herein. Additionally, it should be appreciated that according to one aspect, one or more computer programs that when executed perform methods of the present disclosure need not reside on a single computer or processor, but may be distributed in a modular fashion among a number of different computers or processors to implement various aspects of the present disclosure.
[0212] Having thus described several aspects and embodiments of the technology set forth in the disclosure, it is to be appreciated that various alterations, modifications, and improvements will readily occur to those skilled in the art. Such alterations, modifications, and improvements are intended to be within the spirit and scope of the technology described herein. For example, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemedto be within the scope of the embodiments described herein. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described. In addition, any combination of two or more features, systems, articles, materials, kits, and / or methods described herein, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.
[0213] The above-described embodiments can be implemented in any of numerous ways. One or more aspects and embodiments of the present disclosure involving the performance of processes or methods may utilize program instructions executable by a device (e.g., a computer, a processor, or other device) to perform, or control performance of, the processes or methods. In this respect, various inventive concepts may be embodied as a computer readable storage medium (or multiple computer readable storage media) (e.g., a computer memory, one or more floppy discs, compact discs, optical discs, magnetic tapes, flash memories, circuit configurations in Field Programmable Gate Arrays or other semiconductor devices, or other tangible computer storage medium) encoded with one or more programs that, when executed on one or more computers or other processors, perform methods that implement one or more of the various embodiments described above. In some embodiments, computer readable media may be non-transitory media.
[0214] Computer-executable instructions may be in many forms, such as program modules, executed by one or more computers or other devices. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Typically, the functionality of the program modules may be combined or distributed as desired in various embodiments.
[0215] Also, data structures may be stored in computer-readable media in any suitable form. For simplicity of illustration, data structures may be shown to have fields that are related through location in the data structure. Such relationships may likewise be achieved by assigning storage for the fields with locations in a computer-readable medium that convey a relationship between the fields. However, any suitable mechanism may be used to establish a relationship between information in fields of a data structure, including through the use of pointers, tags or other mechanisms that establish a relationship between data elements.
[0216] Also, a computer may have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that can be used for a user interface include keyboards, and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer may receive input information through speech recognition or in other audible formats.
[0217] Such computers may be interconnected by one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network, and intelligent network (IN) or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks or fiber optic networks.
[0218] Also, as described, some aspects may be embodied as one or more methods. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.
[0219] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.
[0220] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”
[0221] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only(optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.
[0222] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.
[0223] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively.
Claims
CLAIMSWhat is claimed is:
1. A computer- implemented method for determining whether a patient with Inflammatory Bowel Disease (IBD) is likely to respond to treatment with a therapeutic agent, the method comprising: accessing a multi-omic dataset comprising at least two of a genomic profile for the patient, a transcriptomic profile for the patient, and a proteomic profile for the patient; providing the multi-omic data set as input to a machine learning model that has been trained to predict a response to a therapy for treating IBD, thereby obtaining a prediction indicative of whether the patient is likely to respond to the therapy; wherein the machine learning model has been trained using data comprising, for each of a plurality of patients with IBD in a training patient population: (i) a multi-omic dataset comprising at least two of a genomic profile for the patient, a transcriptomic profile for the patient, and a proteomic profile for the patient, and (ii) a ground truth predicted response to therapy, wherein the ground truth predicted response to therapy is associated with a cluster label that has been obtained by: stratifying the training patient population into phenotypic groups based on the multi- omic datasets using unsupervised clustering; and defining a patient population predicted to respond to a therapy based on the stratification.
2. The method of claim 1, wherein the phenotypic groups comprise a cluster of samples characterized by upregulation of transcriptomic and / or proteomic features related to inflammasome activation and cytokine signaling and / or a cluster of sample associated with significantly higher disease severity scores than other clusters, wherein the samples in said cluster are predicted to respond to the therapy.
3. The method of claim 1 or claim 2, wherein the machine learning model includes at least one of a support vector machine, a neural network, a decision tree, a random forest, or a gradient boosting machine.
4. The method of any preceding claim, further comprising training the machine learning model, identifying one or more biomarkers associated with the prediction made by the machine learning model, obtaining the ground truth predicted response to therapy is associated with a cluster label, validating the machine learning model using an independent patient dataset and / or validating the identified biomarkers using an independent patient dataset.
5. The method of any preceding claim, wherein the unsupervised clustering is performed using at least one of k-means clustering, hierarchical clustering, or Gaussian mixture models.
6. The method of any preceding claim, wherein the patient metadata includes at least one of: age, gender, disease duration, prior medication history, and lifestyle factors; optionally wherein the machine learning model further takes as input patient metadata.
7. The method of any preceding claim, further comprising administering the therapy to the patient or recommending the therapy for administration to the patient when the patient is predicted to be likely to respond to the therapy, wherein the therapy comprises a compound, a therapeutic treatment, or a combination thereof.
8. The method of any preceding claim, further comprising the step of monitoring the patient for changes in the prediction over time, and / or monitoring the training patient population for changes in one or more biomarkers associated with the stratification of patients over time.
9. The method of any preceding claim, wherein the multi-omic dataset further comprises microbiomic profiles of the patient data.
10. The method of any preceding claim, wherein the machine learning model takes as input data for at least two of the biomarkers selected from the group of IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, ABHD5, CIDEC, RPS26, TMEM25, LY96, IL12B, BCL2, ANGPTL3, EPPK1, FGF19, INSL5, FKBP7, PILRA, CCL20, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P-STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L- DT, and NRG3.
11. The method of claim 10, wherein the biomarkers selected from IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, RPS26, TMEM25 are transcriptomics biomarkers, and / or wherein the biomarkers selected from LY96, IL12B, BCL2, FGF19, ANGPTL3, INSL5, CCL20 are proteomics biomarkers.
12. The method of claim 10, wherein the identified biomarkers comprise at least two of the biomarkers selected from the group of IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, ABHD5, CIDEC, RPS26, TMEM25, LY96, IL12B, BCL2, ANGPTL3, EPPK1, FGF19, INSL5, FKBP7, PILRA, CCL20, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P- STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L-DT, and NRG3.
13. The method of any preceding claim, further comprising the step of providing personalized treatment recommendations based on the prediction from the machine learning model and / or for each stratified phenotypic group.
14. The method of any preceding claim, wherein the stratification is used to exclude subjects from the patient population who are predicted not to respond to a therapy; and / or wherein the method comprises excluding the patient from treatment with the therapy when the patient is predicted to be unlikely to respond to the therapy.
15. A method of predicting a response to a therapy for an inflammatory bowel disease (IBD) in a subject in need thereof, the method comprising: a. obtaining a biological sample from the subject; b. contacting the biological sample with a set of probes capable of detecting a panel of biomarkers, wherein the panel of biomarkers comprises at least five biomarkers selected from the group of regenerating family member 1 beta (REGIB), fatty acid binding protein 6 (FABP6), regenerating family member 1 alpha (REGIA), major histocompatibility complex, class II, DQ beta 1 (HLA-DQB1), major histocompatibility complex, class II, DQ alpha 1 (HLA-DQA1), complement factor I (CFI), serpin family A member 1 (SERPINA1), indoleamine 2,3-dioxygenase 1 (IDO1), sodium channel epithelial 1 beta subunit (SCNN1B), deleted in malignant brain tumors 1 (DMBT1), suppressor of cytokine signaling 3 (SOCS3), guanylate binding protein 4 (GBP4), C-X-C motif chemokine ligand 1 (CXCL1), CXCL5, CXCL9, CXCL10, CXCL11, dual oxidase 2 (DUOX2), apolipoprotein Al (APOA1), carbonic anhydrase 4 (CA4), ubiquitin D (UBD), guanylate binding protein 1 (GBP1), interferon induced protein with tetratricopeptide repeats 3 (IFIT3), t-box 3 (TBX3), transglutaminase 2 (TGM2), vanin 1 (VNN1), peptidase inhibitor 3 (PI3), complement C3 (C3), cytochrome P450 family 2 subfamily B member 7, pseudogene (CYP2B7P), C-C motif chemokine ligand 20 (CCL20), lipocalin 2 (LCN2), major histocompatibility complex, class II, DP alpha 1 (HLA- DPA1), interleukin 13 receptor subunit alpha 2 (IL13RA2), Immunoglobulin Kappa Variable6D-21 (IGKV6D-21), IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, ABHD5, CIDEC, RPS26, TMEM25, LY96, IL12B, BCL2, ANGPTL3, EPPK1, FGF19, INSL5, FKBP7, PILRA, CCL20, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P-STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L-DT, and NRG3; c. detecting the presence or absence or level of expression of the biomarkers in the sample; and d. predicting the subject's response to the therapy based on the results of the detecting.
16. The method of claim 15, wherein predicting the subject's response to the therapy based on the results of the detecting comprises determining an enrichment score for the sample based on the results of the detecting, wherein an enrichment score less than zero indicates that the subject is more likely to respond to the therapy than a subject with an enrichment score greater than zero.
17. The method of 15 or claim 16, wherein the panel of biomarkers comprises at least two of the biomarkers selected from the group of IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, ABHD5, CIDEC, RPS26, TMEM25, LY96, IL12B, BCL2, ANGPTL3, EPPK1, FGF19, INSL5, FKBP7, PILRA, CCL20, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P- STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L-DT, and NRG3.
18. The method of claim 17, wherein the panel of biomarkers comprises at least one of the biomarkers selected from the group of RPS26, TMEM25, ANGPTL3, IFT43-TTLL5, FAM9B, SLC24A3, CASTOR3P-STAG3, MAP2, KCNQ5, MREG, ARID4A-TOMM20L-DT, and NRG3.
19. The method of any of claims 15-18, wherein the biological sample is selected from the group of blood, stool, tissue biopsy, saliva, or a combination thereof.
20. The method of any of claims 15-19, wherein the therapy includes administration of a compound, a therapeutic treatment, or a combination thereof.
21. The method of any of claims 15-20, further comprising the step of adjusting the treatment regimen based on the predicted response.
22. The method of claim 16, wherein the enrichment score is determined using a scoring algorithm that weights the biomarkers based on their predictive value.
23. The method of any of claims 15 to 22, further comprising the step of re-evaluating the subject's response to the therapy by repeating steps (a) through (d) at subsequent time points.
24. The method of any of claims 15 to 23, wherein the biomarkers selected from IGKV6D-21, GALNT6, TBX3, NRG4, USP38, FNDC1, RPS26, TMEM25 are transcriptomics biomarkers, and / or wherein the biomarkers selected from LY96, IL12B, BCL2, FGF19, ANGPTL3, INSL5, CCL20 are proteomics biomarkers.
25. The method of claim 24, wherein transcriptomic and / or genomic biomarkers are detected by contacting the biological sample with a set of nucleic acid probes, and proteomic biomarkers are detected by contacting the biological sample with a set of antigen binding reagents.
Citation Information
Patent Citations
Acute ischemic stroke clinical phenotype construction method, and key biomarker screening method and application
CN113851216A
Multi-omics data hierarchical classification structure learning system based on machine learning
CN117556334A
Multi-omic search engine for integrative analysis of cancer genomic and clinical data
US20210319907A1
System for predicting therapy resistance and its molecular mechanisms in rectal cancer before treatment
US20240062915A1
Ex VIVO gastrointestinal biopsy platform to evaluate multi-OMIC signatures for screening candidate therapeutics
WO2022031954A1
Cited By
Multi-element machine learning model-based cross-species lung disease feature gene screening method and system, electronic system and storage device
CN121072811A