System and method for predicting at least one outcome in a subject suffering from an inflammatory bowel disease
The system uses a machine learning model trained on capsule endoscopy frames to predict Crohn's disease outcomes, addressing the lack of predictive markers by providing accurate and actionable insights for treatment planning.
Patent Information
- Application Number
- PCT/IL2025/050617
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-18
- Filing Date
- 2025-07-17
- Publication Date
- 2026-01-22
AI Technical Summary
Current systems lack robust and clinically actionable predictive markers for disease progression and outcomes in inflammatory bowel disease, particularly Crohn's disease, which has a highly variable clinical course, limiting effective treatment strategies.
A system and method utilizing a machine learning model trained on capsule endoscopy frames, employing self-supervised pre-training and pretext tasks to mimic expert diagnostic patterns, for predicting disease relapse, need for biological therapy, surgery, or Lewis score, leveraging advanced Al algorithms for comprehensive analysis of CE video data.
Enables accurate prediction of disease outcomes and tailored treatment strategies by extracting nuanced visual data from CE recordings, enhancing clinical decision-making and improving patient management.
Smart Images

Figure IL2025050617_22012026_PF_FP_ABST
Abstract
Description
[0001] SYSTEM AND METHOD FOR PREDICTING AT LEAST ONE OUTCOME IN A SUBJECT SUFFERING FROM AN INFLAMMATORY BOWEL DISEASE
[0002] TECHNICAL FIELD
[0003] The present invention relates to the field of medical diagnostics, and more particularly to systems and methods for predicting one or more clinical outcomes in subjects suffering from inflammatory bowel disease (IBD).
[0004] BACKGROUND
[0005] Inflammatory Bowel Disease (IBD) is a chronic and relapsing inflammatory condition of the gastrointestinal (GI) tract that primarily includes two major clinical entities: Crohn’s Disease (CD) and Ulcerative Colitis (UC). These conditions are characterized by persistent inflammation that leads to progressive damage of the GI tract, often resulting in debilitating symptoms, impaired quality of life, and significant healthcare burdens.
[0006] The global market for IBD treatment was valued at approximately USD 19.2 billion in 2020, and it is projected to grow at a compound annual growth rate (CAGR) of 4.8% from 2021 through 2028. This growth is largely driven by the increasing prevalence of IBD worldwide, particularly in industrialized nations, and by the growing demand for advanced treatment modalities. The therapeutic landscape is further evolving with the adoption of biologic therapies, as well as the development of promising pipeline drugs such as Upadacitinib, Risankizumab, Tofacitinib, Ustekinumab, among others, which aim to modulate immune responses more precisely and effectively.
[0007] In parallel, artificial intelligence (Al) has emerged as a transformative tool in the field of gastroenterology. Recent advancements have demonstrated the ability of Al algorithms to accurately detect and classify abnormalities in still images, particularly in capsule endoscopy (CE). Furthermore, Al-enabled real-time polyp detection during colonoscopy has been shown to significantly improve diagnostic performance, with FDA- approved tools such as GI Genius™ (Medtronic, Dublin, Ireland) now being used in clinical practice. Despite these advancements, the integration of Al in the context of inflammatory bowel disease remains in its infancy. A major unmet clinical need lies in the ability to predict disease progression and outcomes, particularly in Crohn’s Disease, where the clinical course can be highly variable and difficult to manage. Although substantial research has been dedicated to personalized treatment approaches, there remains a critical gap in translating these findings into actionable, real-world clinical decision-making tools.
[0008] Given these challenges and opportunities, there exists a compelling need for improved systems and methods that leverage modern computational techniques, particularly Al, to predict one or more outcomes in individuals diagnosed with IBD.
[0009] GENERAL DESCRIPTION
[0010] In accordance with a first aspect of the presently disclosed subject matter, there is provided a system for predicting at least one outcome in a subject suffering from an inflammatory bowel disease, said system comprising a processing circuitry configured to: obtain (i) a series of frames of said subject's gastrointestinal tract, and (ii) a machine learning model capable of receiving a series of frames of the gastrointestinal tract of a given subject suffering from inflammatory bowel disease and predicting one or more outcomes associated with said disease course in said given subject; and, predict, utilizing said obtained series of frames of said subject's gastrointestinal tract and said machine learning model, at least one outcome in said subject suffering from said inflammatory bowel disease, wherein said at least one outcome includes at least one of: (a) relapse of said disease, (b) need for biological therapy, (c) need for surgery, or (d) a Lewis score reflecting a level of inflammatory activity in said subject.
[0011] In some cases, the machine learning model is trained using a self-supervised pretrained model.
[0012] In some cases, the self-supervised pre-trained model is a teacher-student model.
[0013] In some cases, the teacher-student model is trained using pretext tasks designed to mimic the diagnostic patterns of expert CE readers.
[0014] In some cases, the pretext tasks include at least one of: (a) identifying and / or prioritizing anatomical distribution of inflammation, (b) identifying morphology, (c) identifying extent of ulcerations and strictures, or (d) predicting a next frame in a sequence and identifying a correct temporal order of frames to ensure GI anatomical understanding.
[0015] In some cases, the inflammatory bowel disease is one of: Crohn's disease or ulcerative colitis.
[0016] In some cases, the frames are generated by a capsule endoscopy camera traveling through the subject's gastrointestinal tract.
[0017] In some cases, the machine learning model is trained based on labeled training data containing a plurality of series of frames of the gastrointestinal tract of a plurality of subjects suffering from said inflammatory bowel disease, wherein each series of frames is associated with at least one outcome of said (a) to (d).
[0018] In accordance with a second aspect of the presently disclosed subject matter, there is provided a method for predicting at least one outcome in a subject suffering from an inflammatory bowel disease comprising: obtaining (i) a series of frames of said subject's gastrointestinal tract, and (ii) a machine learning model capable of receiving a series of frames of the gastrointestinal tract of a given subject suffering from inflammatory bowel disease and predicting one or more outcomes associated with said disease course in said given subject; and, predicting, utilizing said obtained series of frames of said subject's gastrointestinal tract and said machine learning model, at least one outcome in said subject suffering from said inflammatory bowel disease, wherein said at least one outcome includes at least one of: (a) relapse of said disease, (b) need for biological therapy, (c) need for surgery, or (d) a Lewis score reflecting a level of inflammatory activity in said subject.
[0019] In some cases, the machine learning model is trained using a self-supervised pretrained model.
[0020] In some cases, the self-supervised pre-trained model is a teacher-student model.
[0021] In some cases, the teacher-student model is trained using pretext tasks designed to mimic the diagnostic patterns of expert CE readers.
[0022] In some cases, the pretext tasks include at least one of: (a) identifying and / or prioritizing anatomical distribution of inflammation, (b) identifying morphology, (c) identifying extent of ulcerations and strictures, or (d) predicting a next frame in a sequence and identifying a correct temporal order of frames to ensure GI anatomical understanding. In some cases, the inflammatory bowel disease is one of: Crohn's disease or ulcerative colitis.
[0023] In some cases, the frames are generated by a capsule endoscopy camera traveling through the subject's gastrointestinal tract.
[0024] In some cases, the machine learning model is trained based on labeled training data containing a plurality of series of frames of the gastrointestinal tract of a plurality of subjects suffering from said inflammatory bowel disease, wherein each series of frames is associated with at least one outcome of said (a) to (d).
[0025] In accordance with a third aspect of the presently disclosed subject matter, there is provided a non-transitory computer readable storage medium having computer readable program code embodied therewith, the computer readable program code, executable by at least one processor to perform a method for predicting at least one outcome in a subject suffering from an inflammatory bowel disease, the method for predicting at least one outcome in a subject suffering from an inflammatory bowel disease comprising: obtaining (i) a series of frames of said subject's gastrointestinal tract, and (ii) a machine learning model capable of receiving a series of frames of the gastrointestinal tract of a given subject suffering from inflammatory bowel disease and predicting one or more outcomes associated with said disease course in said given subject; and, predicting, utilizing said obtained series of frames of said subject's gastrointestinal tract and said machine learning model, at least one outcome in said subject suffering from said inflammatory bowel disease, wherein said at least one outcome includes at least one of: (a) relapse of said disease, (b) need for biological therapy, (c) need for surgery, or (d) a Lewis score reflecting a level of inflammatory activity in said subject.
[0026] BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to understand the presently disclosed subject matter and to see how it may be carried out in practice, the subj ect matter will now be described, by way of non-limiting examples only, with reference to the accompanying drawings, in which:
[0028] Fig. 1 is an exemplary three-stage process performed by the system and method of the presently disclosed subject matter;
[0029] Fig- 2 is an exemplary illustration of a Smart Sampling ViT Video Classification, in accordance with the presently disclosed subject matter; Fig- 3 is an exemplary illustration of a ViT Interval Sampling and Majority Voting, in accordance with the presently disclosed subject matter;
[0030] Fig- 4 is an exemplary illustration of a Transform er-CNN Ensemble Feature Fusion, in accordance with the presently disclosed subject matter;
[0031] Fig- 5 is an exemplary illustrations of images generated using MixUp and CutMix data augmentation techniques, which are employed to enhance the performance and robustness of machine learning models, in accordance with the presently disclosed subject matter;
[0032] Fig. 6 is a block diagram schematically illustrating one example of a system for predicting at least one outcome in a subject suffering from an inflammatory bowel disease, in accordance with the presently disclosed subject matter; and,
[0033] Fig. 7 is an exemplary flowchart illustrating an example of a sequence of operations carried out by a system for predicting at least one outcome in a subject suffering from an inflammatory bowel disease, in accordance with the presently disclosed subject matter.
[0034] DETAILED DESCRIPTION
[0035] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the presently disclosed subject matter. However, it will be understood by those skilled in the art that the presently disclosed subject matter may be practiced without these specific details. In other instances, well- known methods, procedures, and components have not been described in detail so as not to obscure the presently disclosed subject matter.
[0036] In the drawings and descriptions set forth, identical reference numerals indicate those components that are common to different embodiments or configurations.
[0037] Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the specification discussions utilizing terms such as “obtaining^ “predicting”, “utilizing“, “generating”, “training”, or the like, include action and / or processes of a computer that manipulate and / or transform data into other data, said data represented as physical quantities, e.g., such as electronic quantities, and / or said data representing the physical objects. The terms “computer”, “processor”, “processing resource”, “processing circuitry”, and “controller” should be expansively construed to cover any kind of electronic device with data processing capabilities, including, by way of non-limiting example, a personal desktop / laptop computer, a server, a computing system, a communication device, a smartphone, a tablet computer, a smart television, a processor (e.g. digital signal processor (DSP), a microcontroller, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), a group of multiple physical machines sharing performance of various tasks, virtual servers co- residing on a single physical machine, any other electronic computing device, and / or any combination thereof.
[0038] The operations in accordance with the teachings herein may be performed by a computer specially constructed for the desired purposes or by a general-purpose computer specially configured for the desired purpose by a computer program stored in a non- transitory computer readable storage medium. The term "non-transitory" is used herein to exclude transitory, propagating signals, but to otherwise include any volatile or nonvolatile computer memory technology suitable to the application.
[0039] As used herein, the phrase "for example," "such as", "for instance" and variants thereof describe non-limiting embodiments of the presently disclosed subject matter. Reference in the specification to "one case", "some cases", "other cases" or variants thereof means that a particular feature, structure or characteristic described in connection with the embodiment s) is included in at least one embodiment of the presently disclosed subject matter. Thus, the appearance of the phrase "one case", "some cases", "other cases" or variants thereof does not necessarily refer to the same embodiment s).
[0040] It is appreciated that, unless specifically stated otherwise, certain features of the presently disclosed subject matter, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the presently disclosed subject matter, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination.
[0041] In embodiments of the presently disclosed subject matter, fewer, more and / or different stages than those shown in Fig. 7 may be executed. In embodiments of the presently disclosed subject matter one or more stages illustrated in Fig. 7 may be executed in a different order and / or one or more groups of stages may be executed simultaneously. Each module in Fig. 6 can be made up of any combination of software, hardware and / or firmware that performs the functions as defined and explained herein. The modules in Fig. 6 may be centralized in one location or dispersed over more than one location. In other embodiments of the presently disclosed subject matter, the system may comprise fewer, more, and / or different modules than those shown in Fig. 6.
[0042] Any reference in the specification to a method should be applied mutatis mutandis to a system capable of executing the method and should be applied mutatis mutandis to a non-transitory computer readable medium that stores instructions that once executed by a computer result in the execution of the method.
[0043] Any reference in the specification to a system should be applied mutatis mutandis to a method that may be executed by the system and should be applied mutatis mutandis to a non-transitory computer readable medium that stores instructions that may be executed by the system.
[0044] Any reference in the specification to a non-transitory computer readable medium should be applied mutatis mutandis to a system capable of executing the instructions stored in the non-transitory computer readable medium and should be applied mutatis mutandis to method that may be executed by a computer that reads the instructions stored in the non-transitory computer readable medium.
[0045] By way of introduction, CD is a chronic, relapsing, and life-long inflammatory disorder of the gastrointestinal tract, characterized by a highly heterogeneous clinical course. Disease severity may range from mild or asymptomatic cases to severe, progressive forms requiring multiple hospitalizations, surgeries, and intensive management. One of the key challenges in clinical practice is the early identification of patients at high risk for adverse outcomes. Early detection of such individuals would enable timely initiation of aggressive, targeted therapies, whereas those with a more favorable prognosis might be safely managed with conservative measures or close observation. Unfortunately, robust and clinically actionable predictive markers for disease course in CD remain largely lacking.
[0046] A central pillar in the modern therapeutic strategy for IBD is the achievement of mucosal healing. Numerous studies have demonstrated that patients who achieve mucosal healing consistently exhibit superior short and long term outcomes, including reduced relapse rates, fewer complications, and decreased need for surgery or corticosteroid use. Consequently, mucosal healing has become a widely accepted treatment target in the management of IBD.
[0047] CE represents a highly valuable tool in both the diagnosis and monitoring of
[0048] Crohn's disease. By enabling non-invasive, panenteric visualization of the small and large bowel, CE offers unique advantages over traditional imaging modalities. Notably, CE is capable of detecting mild-to-moderate mucosal inflammation in over 80% of CD patients in clinical remission, including those also meeting biomarker remission criteria, underscoring its utility in subclinical disease assessment.
[0049] A typical CE recording comprises approximately 10,000 images, yet current scoring indices for assessing inflammatory severity rely on a limited set of semi- quantitative variables, such as:
[0050] • The number of ulcers (e.g., none / few / multiple),
[0051] • Subjective estimates of ulcerated area (“eyeballing”),
[0052] • Presence or absence of strictures.
[0053] These simplified scoring systems inherently process only a fraction of the rich visual data captured during the CE examination, thereby overlooking nuanced features that could be of prognostic or diagnostic significance.
[0054] In this context, Al offers a compelling opportunity to extract and analyze the full spectrum of optical information present in CE recordings. Al-driven analysis may provide a comprehensive and objective assessment of the inflammatory burden by evaluating not only ulceration extent and morphology, but also the spatial distribution of lesions throughout the gastrointestinal tract. This distribution, inferred from the temporal position of images within the CE sequence (i.e., location between the first and last frames), may have important implications for disease classification and prognosis.
[0055] Taken together, full-length CE videos, when systematically analyzed using advanced Al algorithms, may offer significant clinical value in predicting disease outcomes and tailoring treatment strategies in Crohn’s disease.
[0056] With the above in mind, the presently disclosed subject matter employs a multimodel ensemble approach with models each pre-trained on different clinically informed pretext tasks. These models are then integrated into a unified system and method to provide comprehensive insights into the CE video data.
[0057] Attention is drawn to Fig. 1 illustrating an exemplary three-stage process performed by the presently disclosed subject matter.
[0058] It is to noted that the following steps serve as a mere example, not intended in any way to limit the scope of the presently disclosed subject matter, and that other steps, alternatively or additionally to the following, or a different number of steps may also be applicable. As illustrated in Fig.l, the exemplary three-step process includes the following stages:
[0059] 1. Stage I - Automatic analysis of single CE images for normal vs. pathological findings.
[0060] 2. Stage II - Automatic analysis of complete CE videos of CD patients in order to train a medically-informed self-supervised model.
[0061] 3. Stage III - Automatic analysis of complete CE videos of CD patients for automatic calculation of four exemplary clinical outcomes.
[0062] Stage I
[0063] By way of introduction, this stage is directed to the automated analysis of individual capsule endoscopy (CE) images, with the objective of distinguishing between normal and pathological findings. This stage serves as a foundational step, facilitating more advanced downstream processing, particularly smart sampling for the subsequent Vision Transformer (ViT)-based video classification in Stage II.
[0064] To perform image classification, a Convolutional Neural Network (CNN)-based model is employed. It should be noted that the use of CNNs and the described approach are provided merely by way of example, and are not intended to limit the scope of the presently disclosed subject matter. Alternative techniques and models known in the art may also be applicable.
[0065] CNNs represent a class of deep learning architectures specifically designed for structured, grid-like data such as images. These models have become a cornerstone of modern computer vision due to their capacity to automatically learn hierarchical features, thereby enabling robust performance in a wide range of tasks, including image recognition, fine-grained classification, facial verification, and even recommendation systems.
[0066] In the context of the present disclosure, the system and method may utilize a CNNbased neural network, such as the EfficientNet architecture (though other architectures may be employed), selected for its optimal trade-off between computational efficiency and predictive accuracy. This balance is particularly critical in CE applications, given the substantial number of images (often exceeding 10,000) generated during a single procedure. The CNN model may be trained on a proprietary dataset comprising over 30,000 manually labeled CE images, annotated by expert readers. This training allows the model to accurately learn discriminative features that differentiate between normal and pathological findings.
[0067] Once trained, the CNN can rapidly process new CE images, classifying them based on the patterns it has learned. This automated analysis offers substantial advantages, including improved time efficiency and reduced human error during the initial stages of image interpretation.
[0068] Importantly, the purpose of this classification is not limited to standalone analysis. It plays a pivotal role in facilitating smart sampling, i.e., the strategic selection of the most informative and diagnostically relevant frames, which are then provided as input to the ViT-based video analysis in Stage II. This process ensures that the transformer models focus on the most meaningful content, thereby enhancing the efficiency, relevance, and accuracy of the subsequent full-video analysis for Crohn’s disease diagnosis and treatment planning.
[0069] Stage
[0070] By way of introduction, this stage focuses on a dual-pronged approach that integrates self-supervised pre-training and transfer learning using Vision Transformer (ViT) architectures. Originally developed for natural language processing (NLP), transformer models have been successfully adapted for video analysis by treating sequences of video frames analogously to words in a sentence, thereby enabling the learning of complex spatiotemporal relationships.
[0071] It should be noted that the approach described herein is provided solely by way of example and is not intended to limit the scope of the presently disclosed subject matter. Other techniques and models known in the art may also be applicable.
[0072] In the experiments conducted, the transformer model employed for video classification was the TimeSformer architecture, developed by Facebook in 2021. While additional transformer variants were considered, TimeSformer was selected for its unique dual-attention mechanism that concurrently attends to both spatial (intra-frame) and temporal (inter-frame) dimensions. This capability renders TimeSformer particularly well-suited for analyzing capsule endoscopy (CE) videos, which contain subtle but diagnostically critical visual transitions across thousands of sequential frames. The process initially involved transfer learning, where a TimeSformer model pretrained on human action recognition was adapted for CE video analysis. This was achieved by replacing the model’s original classification head with a task-specific head tailored for binary classification of patients as either in Crohn's disease remission or not in remission. Feeding the model with labeled CE videos demonstrated a successful domain adaptation, effectively repurposing the original model for a novel clinical application. This proof-of-concept (POC) confirmed the feasibility of leveraging preexisting transformer models for endoscopic data interpretation.
[0073] Encouraged by this initial success, we pursued a more advanced strategy: selfsupervised pre-training. Unlike transfer learning, which depends on labeled datasets, selfsupervised learning involves training on large volumes of unlabeled CE videos through pretext tasks generated by the model itself. In the context of the current system, these are referred to as "medically informed tasks", crafted to emulate the nuanced visual assessments typically performed by human CE readers.
[0074] Examples of such tasks may include: (a) predicting the next frame in a temporal sequence; (b) reconstructing a scrambled video segment, (c) detecting anatomical landmarks, (d) identifying transitions between different segments of the gastrointestinal tract.
[0075] These tasks compel the model to develop an intrinsic understanding of the spatiotemporal structure and visual dynamics of CE videos, without reliance on manual labeling. This capability is particularly valuable in the IBD domain, where labeled datasets are limited due to the expertise and time required for annotation.
[0076] To further enhance the self-supervised learning process, a student-teacher architecture may be employed. In this paradigm, a teacher model generates tasks or representations that act as learning targets, while a student model is trained to replicate or surpass these targets. This structure guides the student model more effectively, accelerating learning and improving its ability to extract clinically meaningful patterns from the data in an unsupervised setting.
[0077] The resulting models trained via self-supervised learning tend to exhibit greater generalizability, making them well-suited for real-world CE datasets that include a broad spectrum of pathological presentations. Moreover, such models serve as a robust foundation for fine-tuning on specific downstream tasks, such as predicting remission status or detecting pathological regions, with only limited amounts of labeled data required. The representations learned during pre-training encapsulate a deep, contextual understanding of endoscopic content, enhancing diagnostic precision and clinical relevance.
[0078] In the presently disclosed system and method, the integration of Stage I and Stage II culminates in a critical operational feature: smart sampling. Specifically, the CNN from Stage I may be used to automatically select 64 diagnostically significant keyframes from each CE video. These frames are chosen based on the CNN’s confidence in identifying regions indicative of disease presence and severity. This targeted sampling enables the transformer model to focus its analysis on the most informative frames, thereby maximizing computational efficiency while preserving diagnostic integrity.
[0079] Stage III
[0080] By way of introduction, this stage is directed to the deployment of various adaptable and tunable implementations, each configured to perform distinct downstream tasks utilizing the self-supervised pre-trained models described in Stage II. The general framework may comprise three key components: smart sampling, feature extraction, and a flexible architecture that serves as a foundation for fine-tuning across multiple clinical applications.
[0081] Illustrated below are four representative operational modes, each corresponding to a specific downstream task. These examples are provided to demonstrate the versatility and modularity of the system and method of the presently disclosed subject matter. In each instance, a distinct architecture may be fine-tuned for a targeted clinical outcome using the shared pre-trained model as a backbone.
[0082] It should be noted that the operational modes presented herein are provided merely by way of example, and are not intended to limit the scope of the presently disclosed subject matter in any way. Alternative or additional modes of operation, whether directed to different architectures, tasks, or configurations, may also fall within the scope of the invention.
[0083] 1. Prediction of Relapse in Patients with Quiescent Crohn’s Disease Architecture Overview: • Smart Sampling: A trained CNN model is employed to select 64 key frames from each CE video. These frames are considered the most indicative of subtle mucosal changes or early inflammatory signs that could predict relapse.
[0084] • Feature Extraction with Self-supervised Model: The selected frames are processed through a pre-trained self-supervised model to extract deep spatial- temporal features.
[0085] • Classification Head: A classification layer is integrated atop the selfsupervised model to classify patients into risk categories for relapse based on extracted features.
[0086] 2. Prediction of the Need for Biological Therapy
[0087] Architecture Overview:
[0088] • Smart Sampling: A smart sampling is applied to extract frames showcasing areas with severe inflammation, ulcers, or other markers indicating a potentially aggressive disease course requiring biological therapy.
[0089] • SAM Integration: A Segment Anything Model (SAM) is incorporated for an enhanced understanding of segmentation-aware features within the selected frames, focusing on the severity and extent of disease involvement.
[0090] • Predictive Modeling: The features extracted by the self-supervised and SAM models are fed into a predictive model that assesses the need for biological therapy.
[0091] 3. Prediction of the Need for Surgery
[0092] Architecture Overview:
[0093] • Smart Sampling for Surgery Indicators: Focus smart sampling on frames showing strictures, deep ulcers, or fistulas that are traditional indicators necessitating surgical intervention.
[0094] • Feature Extraction: Use of the self-supervised model for feature extraction, emphasizing structural abnormalities indicative of advanced disease.
[0095] • YOLO for Localizing Critical Findings: Employ YOLO (You Only Look Once) or a similar object detection model to precisely localize and assess the extent of critical findings within the selected frames. • Surgical Need Assessment: Aggregate the extracted features and localization data to predict the immediate or future need for surgical intervention, possibly through a decision-making neural network layer that considers both the severity and the spread of the identified abnormalities.
[0096] 4. Automated Scoring (Lewis Score) of the Full CE Videos
[0097] Architecture Overview:
[0098] • Video Analysis: Instead of smart sampling, sequential clips are processed through the self-supervised model to identify and classify pathological findings.
[0099] • SAM for Detailed Segmentation: To refine the understanding of characteristics of pathological finding, aiding in the precise quantification required for scoring.
[0100] • Lewis Score Calculation: Maps the results to a Lewis Score, aggregating each marker with its scoring criteria.
[0101] To further enhance the automated scoring of Lewis Score from full-length CE videos as described in Operational Mode 4 above, the presently disclosed system may additionally employ an exemplary hierarchical Al pipeline composed of both frame-level and video-level modules. This exemplary hierarchical Al pipeline raises accuracy, sharply cuts review time, and surfaces interpretable metrics.
[0102] 1. Frame analysis: A ConvNeXt network, trained on 18K labeled frames with MixUp + CutMix (p = 0.5 each) augmentations and focal loss, achieves 0.87 accuracy / 0.85 precision / 0.85 recall. Any frame whose highest class-probability falls below 0.40, as well as all unanalyzable frames, is discarded immediately.
[0103] 2. Temporal aggregation: Remaining frames are grouped into overlapping 100-frame windows. For each window, 50 engineered statistics (moments, percentiles, entropy, temporal dynamics, cross-class correlations, continuity cues) are computed and averaged per video. 3. Video-level prediction: XGBoost and a bidirectional LSTM-attention head operate on the aggregated features. At the binary threshold classification tasks, all models score a high AUC. The LSTM and XGboost model offer comparable results in the ternary classification.
[0104] Effective management of CD depends on precise evaluation of mucosal inflammation. The Lewis Score (LS) converts CE findings into a standardized numerical scale that informs therapeutic decisions. However, generating the Lewis Score requires expert review of every video frame, and inter-reader agreement remains only moderate. While previous machine learning approaches have primarily targeted lesion detection, few have extended their utility to reliably predicting LS severity categories. The above hierarchical Al pipeline introduces a two-stage system: (1) frame-level ConvNeXt with balanced training via MixUp, CutMix, and focal loss; (2-3) statistical video-level learners that map temporal patterns to LS categories.
[0105] Materials and Methods
[0106] Data Acquisition and Splitting
[0107] • Frame dataset: 18 k still images manually labeled as normal mucosa, ulcer, stricture, or unanalyzable, sourced from CE studies.
[0108] • Video corpus: >1000 full-length CE videos (-10000 frames each) with gastroenterologist-assigned LS values.
[0109] • Splits: 65 % train, 15 % validation, 20 % test; stratified by LS category without patient overlap.
[0110] Frame-Level Classification, Data Augmentation, and Focal Loss
[0111] The frame classifier adopts the ConvNeXt architecture pretrained on ImageNet. Each mini -batch undergoes MixUp and CutMix with probabilities 0.5 each: • MixUp blends pixel intensities and label vectors of two images via linear interpolation, encouraging the network to learn smoother decision boundaries and reducing over-confidence.
[0112] • CutMix pastes a random rectangular patch from one image into another and combines the labels in proportion to the patch area, forcing the model to associate localized visual cues with mixed labels (see, Fig. 5, left section).
[0113] After augmentation, focal loss automatically down-weights easy normal frames and concentrates gradient updates on hard ulcer / stricture examples, improving learning for under-represented classes without explicit resampling. Frames with maximum class probability below 0.40 or predicted as unanalyzable are excluded from downstream processing.
[0114] Temporal Windowing
[0115] Each video is segmented into 100-frame windows with 50% overlap. Only complete windows are kept.
[0116] Feature Engineering
[0117] A 50-dimensional vector is extracted per window:
[0118] Window vectors are averaged per video and standardized using training-set statistics.
[0119] Video-Level Learners
[0120] • Logistic Regression (Log Reg) is a linear classifier that models the log-odds of each LS class as a weighted sum of features; its simplicity yields fast training and transparent coefficients.
[0121] • Extreme Gradient Boosting (XGBoost) builds an ensemble of decision trees sequentially, where each new tree corrects errors made by the previous ensemble; built-in regularization controls overfitting and offers strong non-linear modeling power.
[0122] • Bidirectional LSTM with Attention (Bi-LSTM) processes the sequence of window features in both forward and backward temporal directions, then applies an attention layer to focus on the most informative windows; this architecture captures long-range dependencies intrinsic to CE videos but requires more compute.
[0123] All models were hyper-parameter optimized via five-fold validation, and class weights were set inversely proportional to class frequency.
[0124] Evaluation Metrics
[0125] • Binary tasks: AUROC, sensitivity, specificity, and Fl score.
[0126] • Ternary tasks: Macro-Fl and balanced accuracy.
[0127] Results
[0128] Binary Classification Across Thresholds
[0129] Table 1 - AUROC on the test set (best learner per threshold)
[0130] Threshold 900 balances sensitivity and specificity for distinguishing moderate-to-severe inflammation from mild disease.
[0131] Ternary Classification
[0132] From the foregoing, it is apparent that MixUp and CutMix, combined with focal loss, expose the ConvNeXt filters to diverse lesion-background juxtapositions and steer gradients toward under-represented ulcer and stricture classes, improving validation accuracy. The XGBoost learner provides a good overall classifier for both binary and ternary tasks.
[0133] While the Bi-LSTM that excels when temporal order is critical - lags behind XGBoost for three-way severity stratification. Limitations include single-center data and moderate sensitivity caused by the relative scarcity of severe disease cases.
[0134] Web-Based Frame Labeling Interface
[0135] Early iterations of relied on a stand-alone Python GUI that had to be wrapped into an executable for distribution. The application required the full image dataset to be copied locally, offered only single-label annotation, and demanded manual source-code edits whenever new lesion classes were added, version control and centralized metrics were absent.
[0136] To overcome these bottlenecks, a secure, invitation-only web platform is implemented, which exposes the labeling workflow: • Browser interface: clinicians label frames directly in any modern browser; no local setup is needed.
[0137] • Multi-label support: checkboxes enable simultaneous assignment of ulcer, stricture and quality flags, accelerating edge-case curation.
[0138] • Labeling dashboard: pie charts and tables summarize class distribution, remaining unlabeled frames, gives the user ability to filter and search for specific frames in videos.
[0139] • Collaboration & versioning: versioning is updated on every change, making it easy to revert to an older dataset in case it is required.
[0140] With respect to the operational modes described above, it is noteworthy that each mode has been specifically designed to effectively address its corresponding clinical task, with an emphasis on both clinical relevance and model explainability. Looking forward, there exists considerable potential to further enhance the interpretability of model outputs through the integration of a Large Language Model (LLM). Such an integration could serve to generate natural language explanations that articulate the rationale behind the model’s predictions, thereby providing clinicians with transparent, context-aware insights into the underlying decision-making process.
[0141] Attention is now drawn to Fig. 6, depicting a block diagram schematically illustrating one example of the system for predicting at least one outcome in a subject suffering from an inflammatory bowel disease 500, in accordance with the presently disclosed subject matter.
[0142] In accordance with the presently disclosed subject matter, the system for predicting at least one outcome in a subject suffering from an inflammatory bowel disease 500 (also interchangeably referred to herein as “system 500”) may comprise a network interface 506. The network interface 506 (e.g., a network card, a Wi-Fi client, 3G / 4G client, or any other component), enables system 500 to communicate over a network with external systems and handles inbound and outbound communications from such systems. For example, system 500 may receive, through network interface 506, one or more series of frames of one or more subjects' gastrointestinal tract.
[0143] System 500 may further comprise or be otherwise associated with a data repository 504 (e.g., a database, a storage system, a memory including Read Only Memory - ROM, Random Access Memory - RAM, or any other type of memory, etc.) configured to store data. Some examples of data that may be stored in the data repository 504 include:
[0144] • One or series of frames of one or more subjects' gastrointestinal tract;
[0145] • One or more machine learning models capable of receiving a series of frames of the gastrointestinal tract of one or more subjects suffering from inflammatory bowel disease and predicting one or more outcomes associated with the disease course in said one or more subjects;
[0146] • One or more outcomes associated with the disease course in one or more subjects;
[0147] • One or more Lewis scores reflecting a level of inflammatory activity in one or more subjects;
[0148] • One or more labeled training data containing a plurality of series of frames of the gastrointestinal tract of a plurality of subjects suffering from an inflammatory bowel disease; etc.
[0149] Data repository 504 may be further configured to enable retrieval and / or update and / or deletion of the stored data. It is to be noted that in some cases, data repository 504 may be distributed, while the system 500 has access to the information stored thereon, e.g., via a wired or wireless network to which system 500 is able to connect (utilizing its network interface 506).
[0150] System 500 further comprises processing circuitry 502. Processing circuitry 502 may be one or more processing units (e.g., central processing units), microprocessors, microcontrollers (e.g., microcontroller units (MCUs)) or any other computing devices or modules, including multiple and / or parallel and / or distributed processing units, which are adapted to independently or cooperatively process data for controlling relevant system 500 resources and for enabling operations related to system’s 500 resources.
[0151] The processing circuitry 502 comprises a prediction module 508, configured to perform a prediction process, as further detailed herein, inter alia with reference to Fig. 7.
[0152] Turning to Fig. 7 there is shown a flowchart illustrating one example of operations carried out by the system for predicting at least one outcome in a subject suffering from an inflammatory bowel disease 500, in accordance with the presently disclosed subject matter.
[0153] Accordingly, the system for predicting at least one outcome in a subject suffering from an inflammatory bowel disease 500 (also interchangeably referred to hereafter as “system 500”) may be configured to perform a predication process 600, e.g., using predication module 508.
[0154] For this purpose, system 500 obtains (i) a series of frames of said subject's gastrointestinal tract, and (ii) a machine learning model capable of receiving a series of frames of the gastrointestinal tract of a given subject suffering from inflammatory bowel disease and predicting one or more outcomes associated with said disease course in said given subject (block 602).
[0155] Once obtained, system 500 predicts, utilizing said obtained series of frames of said subject's gastrointestinal tract and said machine learning model, at least one outcome in said subject suffering from said inflammatory bowel disease (block 604).
[0156] In some cases, the at least one outcome may include, for example, at least one of: (a) relapse of said disease, (b) need for biological therapy, (c) need for surgery, or (d) a Lewis score reflecting a level of inflammatory activity in said subject.
[0157] It is to be of note that the above outcomes serve as mere examples, not intended in any way to limit the scope of the presently disclosed subject matter, and that other outcomes, alternatively or additionally to the above, may also be applicable.
[0158] In some cases, the obtained series of frames may be generated, for example, by a capsule endoscopy camera traveling through the subject's gastrointestinal tract.
[0159] It is to be noted that the above serves as a mere example and that other methods or techniques to obtained series of frames may also be applicable, mutatis mutandis.
[0160] In some cases, the machine learning model may be trained, for example, based on labeled training data containing a plurality of series of frames of the gastrointestinal tract of a plurality of subjects suffering from said inflammatory bowel disease, where each series of frames is associated with at least one outcome of said (a) to (d).
[0161] In some cases, as indicated above, the machine learning model may be trained, for example, using a self-supervised pre-trained model. In such cases, the self-supervised pre-trained model may be, for example, a teacher- student model.
[0162] In some cases, the teacher-student model may be trained, for example, using pretext tasks designed to mimic the diagnostic patterns of expert CE readers. In such cases, the pretext tasks may include, for example, at least one of: (a) identifying and / or prioritizing anatomical distribution of inflammation, (b) identifying morphology, (c) identifying extent of ulcerations and strictures, or (d) predicting a next frame in a sequence and identifying a correct temporal order of frames to ensure GI anatomical understanding.
[0163] It is to be of note that the above pretext tasks serve as mere examples, not intended in any way to limit the scope of the presently disclosed subject matter, and that other pretext tasks, alternatively or additionally to the above, may also be applicable.
[0164] It is to be noted, with reference to Fig. 7, that some of the blocks can be integrated into a consolidated block or can be broken down to a few blocks and / or other blocks may be added. It is to be further noted that some of the blocks are optional. It should be also noted that whilst the flow diagram is described also with reference to the system elements that realizes them, this is by no means binding, and the blocks can be performed by elements other than those described herein.
[0165] It is to be understood that the presently disclosed subject matter is not limited in its application to the details set forth in the description contained herein or illustrated in the drawings. The presently disclosed subject matter is capable of other embodiments and of being practiced and carried out in various ways. Hence, it is to be understood that the phraseology and terminology employed herein are for the purpose of description and should not be regarded as limiting. As such, those skilled in the art will appreciate that the conception upon which this disclosure is based may readily be utilized as a basis for designing other structures, methods, and systems for carrying out the several purposes of the present presently disclosed subject matter.
[0166] It will also be understood that the system according to the presently disclosed subject matter can be implemented, at least partly, as a suitably programmed computer. Likewise, the presently disclosed subject matter contemplates a computer program being readable by a computer for executing the disclosed method. The presently disclosed subject matter further contemplates a machine-readable memory tangibly embodying a program of instructions executable by the machine for executing the disclosed method.
Claims
CLAIMS:
1. A system for predicting at least one outcome in a subject suffering from an inflammatory bowel disease, said system comprising a processing circuitry configured to: obtain (i) a series of frames of said subject's gastrointestinal tract, and (ii) a machine learning model capable of receiving a series of frames of the gastrointestinal tract of a given subject suffering from inflammatory bowel disease and predicting one or more outcomes associated with said disease course in said given subject; and, predict, utilizing said obtained series of frames of said subject's gastrointestinal tract and said machine learning model, at least one outcome in said subject suffering from said inflammatory bowel disease, wherein said at least one outcome includes at least one of: (a) relapse of said disease, (b) need for biological therapy, (c) need for surgery, or (d) a Lewis score reflecting a level of inflammatory activity in said subject.
2. The system of claim 1, wherein said machine learning model is trained using a selfsupervised pre-trained model.
3. The system of claim 2, wherein said self-supervised pre-trained model is a teacherstudent model.
4. The system of claim 3, wherein said teacher-student model is trained using pretext tasks designed to mimic the diagnostic patterns of expert CE readers.
5. The system of claim 4, wherein said pretext tasks include at least one of: (a) identifying and / or prioritizing anatomical distribution of inflammation, (b) identifying morphology, (c) identifying extent of ulcerations and strictures, or (d) predicting a next frame in a sequence and identifying a correct temporal order of frames to ensure GI anatomical understanding.
6. The system of claim 1, wherein said inflammatory bowel disease is one of: Crohn's disease or ulcerative colitis.
7. The system of claim 1, wherein said frames are generated by a capsule endoscopy camera traveling through the subject's gastrointestinal tract.
8. The system of claim 1, wherein said machine learning model is trained based on labeled training data containing a plurality of series of frames of the gastrointestinal tract of a plurality of subjects suffering from said inflammatory bowel disease, wherein each series of frames is associated with at least one outcome of said (a) to (d).
9. A method for predicting at least one outcome in a subject suffering from an inflammatory bowel disease comprising: obtaining (i) a series of frames of said subject's gastrointestinal tract, and (ii) a machine learning model capable of receiving a series of frames of the gastrointestinal tract of a given subject suffering from inflammatory bowel disease and predicting one or more outcomes associated with said disease course in said given subject; and, predicting, utilizing said obtained series of frames of said subject's gastrointestinal tract and said machine learning model, at least one outcome in said subject suffering from said inflammatory bowel disease, wherein said at least one outcome includes at least one of: (a) relapse of said disease, (b) need for biological therapy, (c) need for surgery, or (d) a Lewis score reflecting a level of inflammatory activity in said subject.
10. The method of claim 9, wherein said machine learning model is trained using a selfsupervised pre-trained model.
11. The method of claim 10, wherein said self-supervised pre-trained model is a teacherstudent model.
12. The method of claim 11, wherein said teacher-student model is trained using pretext tasks designed to mimic the diagnostic patterns of expert CE readers.
13. The method of claim 12, wherein said pretext tasks include at least one of: (a) identifying and / or prioritizing anatomical distribution of inflammation, (b) identifying morphology, (c) identifying extent of ulcerations and strictures, or (d) predicting a next frame in a sequence and identifying a correct temporal order of frames to ensure GI anatomical understanding14. The method of claim 9, wherein said inflammatory bowel disease is one of: Crohn's disease or ulcerative colitis.
15. The method of claim 9, wherein said frames are generated by a capsule endoscopy camera traveling through the subject's gastrointestinal tract.
16. The method of claim 9, wherein said machine learning model is trained based on labeled training data containing a plurality of series of frames of the gastrointestinal tract of a plurality of subjects suffering from said inflammatory bowel disease, wherein each series of frames is associated with at least one outcome of said (a) to (d).
17. A non-transitory computer readable storage medium having computer readable program code embodied therewith, the computer readable program code, executable by at least one processor to perform a method for predicting at least one outcome in a subject suffering from an inflammatory bowel disease, the method for predicting at least one outcome in a subject suffering from an inflammatory bowel disease comprising: obtaining (i) a series of frames of said subject's gastrointestinal tract, and (ii) a machine learning model capable of receiving a series of frames of the gastrointestinal tract of a given subject suffering from inflammatory bowel disease and predicting one or more outcomes associated with said disease course in said given subject; and, predicting, utilizing said obtained series of frames of said subject's gastrointestinal tract and said machine learning model, at least one outcome in said subject suffering from said inflammatory bowel disease, wherein said at least one outcome includes at least one of: (a) relapse of said disease, (b) need for biologicaltherapy, (c) need for surgery, or (d) a Lewis score reflecting a level of inflammatory activity in said subject.
Citation Information
Patent Citations
Methods for treatment of inflammatory bowel disease
US20220028550A1
Self-supervised representation learning paradigm for medical images
US20230111306A1
Ai platform for processing speech and video information collected during a medical procedure
US20230298589A1