Automated disease severity assessment based on analysis of medical videos
A computerized deep-learning system processes medical videos to enhance disease severity assessment in IBD, addressing the limitations of human-read scoring systems by providing granular, automated measurements.
Patent Information
- Application Number
- PCT/IB2025/051837
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-15
- Filing Date
- 2025-02-20
- Publication Date
- 2025-08-28
AI Technical Summary
Current human-read scoring systems for Inflammatory Bowel Disease (IBD) clinical trials, such as the Mayo Endoscopic Score (MES) and Ulcerative Colitis Endoscopic Index of Severity (UCEIS), lack sufficient granularity to accurately measure disease severity differences before and after treatment, limiting their discriminatory power in assessing treatment efficacy.
A computerized deep-learning system using self-supervised learning and attention-based deep learning networks processes medical videos to generate frame-level and video-level inferences for disease severity assessment, incorporating pre-processing, frame embeddings, and video-level analysis to provide continuous scores or segment-level assessments.
Enhances the ability to measure treatment efficacy with improved granularity and discriminatory power by automating disease severity assessment in IBD, providing continuous scores and detailed frame-level and video-level insights.
Smart Images

Figure IB2025051837_28082025_PF_FP_ABST
Abstract
Description
Automated Disease Severity Assessment Based on Analysis of Medical VideosCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application 63 / 555,883, filed on February 20, 2024. This application also shares some subject matter with International Application PCT / IB2024 / 057930, filed on August 16, 2024, and PCT / IB2024 / 061440, filed on November 15, 2024. The contents of these applications are incorporated by reference herein.BACKGROUND
[0002] This disclosure relates generally to computerized technology for assessing disease severity of a patient based on medical videos of an area of interest of the patient.SUMMARY
[0003] Endoscopy-based disease severity assessment in Inflammatory Bowel Disease (IBD) clinical trials is typically done using human-read scoring systems such as the Mayo Endoscopic Score (MES) or Ulcerative Colitis Endoscopic Index of Severity (UCEIS). Computer vision and artificial intelligence (Al) have the potential to automate and improve upon these measurements. Moreover, in some cases, typical scoring methods do not reflect differences between pre-treatment and post-treatment disease with sufficient granularity to measure treatment efficacy with desirable levels of discriminatory power.
[0004] In some implementations of the present disclosure, computerized deep-learning systems and methods are disclosed for processing medical videos to assess disease severity of a patient. One or more implementations comprise a pre-processor configured to pre- process video data to obtain a plurality of video frames corresponding to one or more medical videos corresponding to the patient. An encoder is configured for encoding a plurality of video frames of the medical videos to obtain respective frame embeddingscorresponding to respective frames of the plurality of video frames, wherein the encoder has been pre-trained using self-supervised learning. One or more video frame classifiers, each comprising an attention-based deep learning network, processes the respective frame embeddings and computes on a frame-by-frame basis, respective frame-level inferences corresponding to the respective frames. A video analyzer comprises an attention-based deep learning network and uses the respective frame-level inferences and the respective frame embeddings to calculate at least one disease severity assessment of the patient.
[0005] In some implementations, a continuous score may be computed. In some implementations, this continuous score may be within a score range of standard scale such as the Mayo Endoscopic subscore (MES). In other embodiments, it may be within a nonstandard range. The score may be computed by processing all frames of a video (or all frames corresponding to a withdrawal path) together. Additionally, or alternately, frames corresponding to one of a plurality of segments are processed together to obtain segment level scores which can be averaged to obtain a video-level patient score. These and other variations consistent with the present disclosure are more fully disclosed below.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] FIG. 1 illustrates a medical video analysis system computing frame-level and video-level inferences for assessing disease severity of a patient in accordance with one or more embodiments of the present disclosure.
[0007] FIG. 2 illustrates a medical video analysis system computing frame-level and video-level inferences for assessing disease severity of a patient in accordance with one or more embodiments of the present disclosure.
[0008] FIG. 3 illustrates a medical video analysis system computing frame-level and video-level inferences for assessing disease severity of a patient in accordance with one or more embodiments of the present disclosure.
[0009] FIG. 4 illustrates a medical video processing system configured to train an attention-based deep learning network to automatically compute inferences for frames of a medical video in accordance with one or more embodiments of the present disclosure.
[0010] FIG. 5 illustrates a method for processing medical videos to train an attentionbased deep learning network to automatically compute inferences for frames of a medical video in accordance with one or more embodiments of the present disclosure.
[0011] FIG. 6 illustrates a video-level augmentation process to enhance training in accordance with one or more embodiments of the present disclosure.
[0012] FIG. 7 illustrates a method for processing medical videos to compute inferences for frames of a medical video using an attention-based deep learning network trained in accordance with one or more embodiments of the present disclosure.
[0013] FIG. 8 illustrates a medical video processing system configured to train an attention-based deep learning network to automatically analyze and compute one or more inferences for a medical video.
[0014] FIG. 9 illustrates a medical video processing system for computing one or more inferences for a medical video using a trained attention-based network in accordance with one or more embodiments of the present disclosure.
[0015] FIG. 10 illustrates a medical video processing system for computing one or more inferences for a medical video using a trained attention-based network in accordance with one or more embodiments of the present disclosure.
[0016] FIG. 11 illustrates a medical video processing system for computing one or more inferences for a medical video using a trained attention-based network in accordance with one or more embodiments of the present disclosure.
[0017] FIG. 12 illustrates a method for processing medical videos to compute inferencesfor a medical video using an atention-based deep learning network trained in accordance with one or more embodiments of the present disclosure.
[0018] FIG. 13 shows an example of a video review interface in accordance with one or more embodiments of the present disclosure.
[0019] FIG. 14 shows an example of a computer system, one or more of which may be used to implement one or more of the apparatuses, systems, and methods illustrated herein.
[0020] While embodiments of the present disclosure are described with reference to the above drawings, the drawings are intended to be illustrative. Other embodiments are consistent with the spirit and scope of the disclosure.DETAILED DESCRIPTION
[0021] The various embodiments now will be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific examples of practicing the embodiments. This specification may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this specification will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. Among other things, this specification may be embodied as methods or devices. Accordingly, any of the various embodiments herein may take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. The following specification is, therefore, not to be taken in a limiting sense.
[0022] FIG. 1 illustrates a medical video analysis system 1000 in accordance with an embodiment of the present disclosure. This and other embodiments will be described with reference to endoscopy videos. But the underlying principles of the disclosure are applicable to other medical image / video analysis applications.
[0023] This example shows system 1000 for analyzing a medical video 101, such as endoscopy video. System 1000 processes a plurality of frames 12, which, taken together, comprise medical video 101. Pre-processor 110 pre-processes the video data corresponding to medical video 101 by resizing frames and / or masking any annotations to generate resized and masked frames 103. Block 110 processes the video data such that each frame has uniform dimensions D 1. In this example, if there are a total of m 1 frames in the video, then the size of the resized and masked frames data 103 is DI x ml after pre-processing. In one implementation of this example, after pre-processing, each frame 12 in a sequence corresponding to medical video 101 has dimensions DI of 224 x 224 x 3 (if using RGB color classes). Other dimensions can be used in different implementations. Pre-processing block 110 prepares a sequence of frames 103, the sequence having a total of ml frames, corresponding to medical video 101. Thus, the full sequence of frames corresponding to medical video 101 has dimension DI x ml as discussed above, where ml is total number of frames obtained from the video for processing.
[0024] Self-Supervised Learning (SSL) pre-trained encoder 120 processes frames 103 to obtain a sequence of frame embeddings 104. There is one frame embedding 104 for each frame 103 processed by SSL pre-trained encoder 120, and the sequential order of frame embeddings 104 is the same as the sequential order of frames 12 in medical video 101 and in the sequence of resized and masked frames 103. Thus, if the size of each frame embedding is of dimension D2, then the full sequence of frame embeddings has dimensions D2 x ml. The value of D2 depends on the underlying SSL model used for SSL pre-training of encoder 120, as discussed below with respect to FIG. 4.
[0025] Next, one or more frame-by-frame classifiers 130 are applied. Each frame-by- frame classifier 130 infers a classification for each of one or more frame embeddings 104 to generate one or more subsets 105 of frame embeddings having a particular classification.Therefore, in one example, subset 105 has dimensions D2 x m3, where m3 is a total number of frame embeddings selected from the full set of ml frames. Thus, in this example, m3 <= ml.
[0026] Although particular types of classification inferences are described below with respect to FIGs. 2-4, a computed “inference” can be any one or more of a prediction, estimate, score, suggestion, categorization, assessment, calculation, or other inference. In some embodiments, rather than using a softmax layer to provide class probabilities, one or more outputs of a multi-layer perceptron (MLP) can be used to provide regression output in the form of, for example, a predicted score within a range of values.
[0027] In some examples, one or more frame-by-frame classifiers 130 may be configured for computing classification inferences for a medical video 101 using a trained attentionbased network comprising attention-based encoder and inference network as discussed below with respect to FIG. 4. In the illustrated example, the frame-by-frame classifier 130 comprises a trained attention-based network including a trained version of the attentionbased encoder 440 of system 4000 illustrated in FIG. 4 and a trained version of the inference network 450 of system 4000 illustrated in FIG. 4.
[0028] As described below in the context of FIG. 4, the number of inferences output for each frame will depend on the number of classes relevant to a particular application.
[0029] Video analyzer 140 analyzes frame embeddings 105 to generate one or more video-level scores or class inferences. In addition, frame-level attention and model information can also be generated, such as attention vectors and / or values representing attention weighting. In some examples, all of such information can be analyzed and presented to a user (e.g., a physician, health care professional, and / or patient) in video review interface 150 for further analysis. For example, for an endoscopic medical video 101 of a patient’s colon, the video review interface could display frames 12 showing an image offrame 12, along with relevant information such as segment classification, subsegment location, scope location, scope direction, a biopsy location (if any), and relevant disease severity scoring information (e.g., using any number of scoring systems known in the art or otherwise, such as Mayo Endoscopic Subscore (MES) or Ulcerative Colitis Endoscopic Index of Severity (UCEIS)). Continuous scores, having non-integer values, can also be displayed on video review interface 150, along with the ability to view attention maps and high attention regions, and the ability to scroll directly to frames showing features of interest for disease severity assessment, including MES data, bleeding data, erosion data, and vascular pattern data. For Crohn’s disease patients, further data on size, surface, and / or narrowing of ulcers, and the percentage of area of an ulcer could also be displayed.
[0030] FIG. 2 illustrates a medical video processing system 2000 for computing classification inferences for frames of a medical video using trained attention-based networks to identify frames on, e.g., the withdrawal path (sometimes referred to herein as the “backward path”) of an endoscopy video and in the left colon area for further analysis.
[0031] This example, like FIG. 1, shows system 2000 for analyzing a medical video 101, such as endoscopy video. System 2000 processes a plurality of frames 12, which, taken together, comprise medical video 101. Pre-processor 110 pre-processes the video data corresponding to medical video 101 by generating and resizing frames, and masking the frames to remove annotations to generate resized and masked frames 103. Block 110 processes the video data such that each frame has uniform dimensions D 1.
[0032] As discussed with respect to FIG. 1, Self-Supervised Learning (SSL) pre-trained encoder 120 processes frames 103 to obtain a sequence of frame embeddings 104. There is one frame embedding 104 for each frame 103 processed by SSL pre-trained encoder 120, and the sequential order of frame embeddings 104 is the same as the sequential order of frames 12 in medical video 101 and in the sequence of resized and masked frames 103.Thus, if the size of each frame embedding is of dimension D2, then the full sequence of frame embeddings has dimensions D2 x ml. The value of D2 depends on the underlying SSL model used for SSL pre-training of encoder 120, as discussed below with respect to FIG. 4.
[0033] Next, three frame-by-frame classifiers 230A, 230B, and 230C are applied. Like frame-by-frame classifier 130 in FIG. 1, each frame-by-frame classifier 230A-230C infers a classification for each of one or more frame embeddings 104 to generate one or more subsets 104A and 104B of frame embeddings having a particular classification. Therefore, after frame embeddings 104 are processed through FWD-BWD classifier / filter 230A, a subset 104A of frame embeddings representing the frames that were taken during the endoscope’s backward path through a patient’s colon are retained. Subset 104A has dimensions D2 x m2, where m2 is a total number of frame embeddings selected from the full set of ml frames. Thus, in this example, m2 <= m 1.
[0034] Next, frame embeddings subset 104A are processed through LEFT-RIGHT classifier / filter 230B, a subset 104B of frame embeddings representing the frames that were taken during the endoscope’s backward path through a left side of the patient’s colon are generated. Subset 104B has dimensions D2 x m3, where m3 is a total number of frame embeddings classified as left colon, selected from the subset of m2 backward path frames. Thus, in this example, m3<= m2 <= ml .
[0035] Next, frame embeddings subset 104B are processed through segment classifier 230C, so that the subset 104B of frame embeddings representing the frames that were taken during the endoscope’s backward path through a left side of the patient’s colon are classified as belonging to one of the following three left colon segments: rectum, sigmoid colon and descending colon. Subset 104B, as mentioned above, has dimensions D2 x m3.
[0036] Video analyzer 140 analyzes frame embeddings 104B and frame class inferences105 for the frames represented by frame embedding subset 104B to generate one or more video-level scores or class inferences for the left colon segments. In addition, frame-level attention and model information can also be generated, such as attention vectors and / or representations of attention weighting. All of such information can be analyzed and presented to a user (e.g., a physician, health care professional, and / or patient) in video review interface 150 for further analysis. For example, for an endoscopic medical video 101 of a patient’s colon, the video review interface could display frames 12 showing an image of frame 12, along with relevant information such as segment classification, subsegment location, scope location, scope direction, a biopsy location (if any), and relevant disease severity scoring information (e.g., using any number of scoring systems known in the art or otherwise, such as Mayo Endoscopic Subscore (MES) or Ulcerative Colitis Endoscopic Index of Severity (UCEIS)). Continuous scores, having non-integer values, can also be displayed on video review interface 150, along with the ability to view attention maps and high attention regions, and the ability to scroll directly to frames showing features of interest for disease severity assessment, including MES data, bleeding data, erosion data, and vascular pattern data. For Crohn’s disease patients, further data on size, surface, and / or narrowing of ulcers, and the percentage of area of an ulcer could also be displayed.
[0037] FIG. 3 illustrates a medical video processing system 3000 for computing classification inferences for frames of a medical video using one or more trained attentionbased networks to identify frames on the withdrawal path of e.g., an endoscopy video of a patient’s colon for further analysis.
[0038] This example, like FIGs. 1 and 2, shows a system 3000 for analyzing a medical video 101, such as endoscopy video. System 3000 processes a plurality of frames 12, which, taken together, comprise medical video 101. Pre-processor 110 pre-processes the video data corresponding to medical video 101 by generating and resizing frames and masking theframes to remove annotations to generate resized and masked frames 103. Block 110 processes the video data such that each frame has uniform dimensions D 1. In this example, if there are a total of m 1 frames in the video, then the size of the resized and masked frames 103 is DI x ml after pre-processing. In one implementation of this example, after preprocessing, each frame 12 in a sequence corresponding to medical video 101 has dimensions DI.
[0039] As discussed with respect to FIGs. 1 and 2, Self-Supervised Learning (SSL) pretrained encoder 120 processes frames 103 to obtain a sequence of frame embeddings 104. There is one frame embedding 104 of dimension D2 for each frame 103 processed by SSL pre-trained encoder 120, and the sequential order of frame embeddings 104 is the same as the sequential order of frames 12 in medical video 101 and in the sequence of resized and masked frames 103. Again, the value of D2 depends on the underlying SSL model used for SSL pre-training of encoder 120, as discussed below with respect to FIG. 4.
[0040] Next, two frame-by-frame classifiers 230A and 230C are applied. Like frame-by- frame classifier 130 in FIG. 1, each frame-by-frame classifier 230A and 230C infers a classification for each of one or more frame embeddings 104 to generate subsets 104A and 104B of frame embeddings having particular classifications. Therefore, after frame embeddings 104 are processed through FWD-BWD classifier / filter 230A, a subset 104A of frame embeddings representing the frames that were taken during the endoscope’s backward path through a patient’s colon are generated. Subset 104A has dimensions D2 x m2, where m2 is a total number of frame embeddings selected from the full set having a total of ml frames. Thus, in this example, m2 <= ml .
[0041] Next, in one example, frame embeddings subset 104A are processed through segment classifier 230C, so that the subset 104A of frame embeddings representing the frames that were taken during the endoscope’s backward path as belonging to the one of thefollowing representative colon segments: ileum, ascending colon, transverse colon, descending colon, sigmoid colon and rectum, although more or fewer colon segments and / or subsegments could be represented depending on which colon segments were recorded on video during the endoscope’s path through the patient’s colon. Subset 104A, as mentioned above, has dimensions D2 x m2, where m2 is a total number of frame embeddings selected from the backward path frames having a total of m2 frames. Thus, in this example, m2 <= ml.
[0042] Video analyzer 140 analyzes frame embeddings 104B and frame class inferences 105 for the frames represented by frame embedding subset 104B to generate one or more video-level scores or class inferences for the left colon segments. In addition, frame-level attention and model information can also be generated, such as attention vectors and feature vectors. All of such information can be analyzed and presented to a user (e.g., a physician, health care professional, and / or patient) in video review interface 150 for further analysis. An example video review interface 1300 is capable of displaying some or all of the information discussed above and / or other relevant information as discussed below with respect to FIG. 13.
[0043] FIG. 4 illustrates a medical video processing system including an attention-based deep learning network (e.g., frame classifier 230A, 230B , and / or 230C), which is trained to automatically compute inferences for frames of a medical video. In some examples, a medical video 401 may be one of many training videos in a training set to train downstream attention-based deep learning network comprising attention-based encoder 440 and inference network 450 comprising MLP 450-1 on an inference task (Softmax layer 450-2 is part of the network but is not trained). In one example, a training set includes over 21 million labeled frames from 1,335 endoscopy videos corresponding to 753 patients fortraining blocks 440 and 450, and 4160 endoscopy videos from four clinical trials and over 61 million frames fortraining block 120.
[0044] Similar to FIGs. 1-3, pre-processing block 110 pre-processes the video data corresponding to a video 401 by generating and resizing frames and removing annotations. Block 110 processes the video data such that each frame has uniform dimensions, DI. In this example, after pre-processing, each frame 43 in a sequence 403 corresponding to video 401 has dimensions DI of 224 x 224 x 3 (if using RGB color channels). Other dimensions can be used in different implementations. Pre-processing block 110 prepares a sequence 403 of frames 43 corresponding to a medical video 401. Thus, the full sequence of frames 43 corresponding to medical video 401 has dimensions of N X DI where N is the total number of frames obtained from the video for processing and DI is the size of each frame.
[0045] Self-Supervised Learning (SSL) pre-trained encoder 120 processes frames 43 to obtain frame embeddings 404. There is one frame embedding 44 for each frame 43 processed by SSL pre-trained encoder 120, and the sequential order of frame embeddings 44 is the same as the sequential order of frames 43 in frame sequence 403. Each frame embedding 44 has a dimension D2.
[0046] The value of D2 depends on the underlying SSL model used for SSL pre-training of encoder 120. In one example, a vision transformer encoder (ViT) is pre-trained using a DIN0v2 objective (see “DIN0v2” approach described by Oquab et al. in “DIN0v2: Learning Robust Visual Features without Supervision” (2023)). However, other SSL pretrained encoders and pre-training techniques can be used without necessarily departing from the spirit and scope of the present disclosure. In another example, an SSL pre-trained encoder, such as encoder 120, can be pre-trained using a DINOvl objective (see “DINOvl” approach described in Caron et al. in “Emerging Properties in Self-Supervised Vision Transformers” (2021)). In yet another example, a convolutional neural network (CNN) encoder such as a ResNet (CNN with residual connections) is used with the “SimCLR”approach (see Chen et al. in “A Simple Framework for Contrastive Learning of Visual Representations” (2020)). All three papers referenced in this paragraph are incorporated by reference herein.
[0047] In one example, SSL pre-trained encoder 120 has been pre-trained using publicly available clinical trial endoscopy videos. In one example, over 61 million frames for over 4,000 endoscopy videos are used for SSL pre-training. In one example, the SSL pre-trained encoder is a ViT-B / 16 encoder pre-trained using the DINOv2 approach. In one example, a batch size of 256 is used for pre-training with a cosine decayed learning rate of 8 x 10-4with an Adam optimizer. In one example, pretraining is done for 15 epochs on four NVIDIA A10G GPUs. Also see, Mobadersany, Pooya et al., "Harnessing Temporal Information for Precise Frame-Level Predictions in Endoscopy Videos,” cited and incorporated by reference above. Also see, International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer Nature Switzerland, 2024, for additional details. This paper is hereby incorporated by reference.
[0048] In one such DINOv2 example, the value of D2 is 768 (one-dimensional embedding). Thus, the full sequence 404 of frame embeddings 44 corresponding to a medical video 401 has dimensions of N X D2 where N is the total number of frames obtained from the video for processing.
[0049] Video-level temporal augmenter (VLA) 425 processes frame embeddings 44 of frame embedding sequence 404. VLA 425 may randomly apply one or more video-level augmentations to a sequence 404 of frame embeddings 44. Additionally, or alternately, VLA 425 may randomly determine whether to apply a first temporal modification operation to the sequence 404 of embeddings 44 and then randomly determines whether to apply a second temporal modification operation to the set of embeddings. One example of such an operation is splitting, i.e., cropping the sequence, in which a subset of the frame embeddingsin the sequence are kept and the rest are discarded. Another such example is reversing the sequence of frame embeddings. In one example, the first temporal modification is splitting and the second temporal modification is reversing, as described further below in the context of FIG. 6.
[0050] Continuing with the description of FIG. 4, VLA 425 outputs augmented sequence 405 of frame embeddings 45. Sequence 405 has dimensions NA X D2 where NA is the number of frame embeddings after augmentation operations. In this example, if VLA 130 did not apply a splitting operation to sequence 12, then NA = N. Otherwise, NA is less than N. The frame-embeddings are processed by an attention-based deep learning network which, in this example, comprises attention-based encoder 440 and inference network 450. Attention-based encoder 440 processes sequence 405 of frame embeddings 45 and outputs sequence 406 of attention vectors 46 having dimensions NA X D2. Each attention vector 46 corresponds to a frame embedding 45 that has been processed by attention-based encoder 440.
[0051] In one example, attention-based encoder 440 is a ViT encoder based on the model described in Dosovitskly et al. “An Image is Worth 16X16 Words”, 2021, hereby incorporated by reference herein. In Vaswani, an extra classification token is added and is used to provide a summarized output token to be used for classifying an image corresponding to a collection of input tokens (which, in Vaswani, each correspond to patches of the image to be analyzed). However, frame-by-frame classifications may be desired, and, in that case, the extra classification token used in Vaswani may not be needed and the output tokens for each input frame embedding are used directly by the inference network. One example of encoder 440 uses a ViT encoder with four layers and eight self-attention heads in each layer.
[0052] Each attention vector 46 of sequence 406 is processed by inference network 450including multi-layer perceptron (MLP) 450-1 and softmax layer 450-2 when doing classification. When doing regression, the Softmax layer 450-2 is not needed. In one example, the attention encoder has a dropout of 0.25 and the classification network has a dropout of 0.5, a batch size of 1 is used, and an Adam optimizer is used with a learning rate of 10-5and a weight decay of 10-6. Network 450 outputs a set of computed inferences 47. In classification problems, inferences 47 may be in the form of probabilities for each class to which a frame may be inferred to belong, and in regression problems, they can be the regression values corresponding to each frame. The dimensions of the set of inferences 47 are NA X C where C is the number of classes (or regression outputs in the case of regression inferences), which depends on the application.
[0053] In one example as shown in FIG. 3, frame classifier 230C is trained to infer colon segment class probabilities corresponding to the five bowel segments: rectum (RM), left colon / sigmoid (LC) (in which descending colon and sigmoid colon are bundled together as one segment), transverse colon (TC), right colon (RC), and ileum (IL). In such an application, “C” is equal to 5, meaning that, for each frame, classification network 450 outputs five class probabilities. In another example shown in FIG. 3, frame classifier 230A is trained to infer whether the frame is taken from a forward path or a withdrawal path of the procedure. In such an application, “C” is equal to 2, meaning that, for each frame, inference network 450 outputs two class probabilities. In another example as shown in FIG. 2, frame classifier 230B is trained to infer whether the frame is taken from the left side or right side of the colon. In such an application, C is also equal to 2. In another example shown in FIG. 2, frame classifier 230C is trained to infer, for each frame from the left side of the colon: rectum (RM), sigmoid colon (SC), and descending colon (DC), which of the three colon segments (RM, SC, or DC) the frame is taken from. In such a 3-class example, C is 3.
[0054] In another example, multiple models are combined together in an end-to-endfashion to operate as follows. First, frames corresponding to a withdrawal path are selected by a first trained model. Then, of those, frames corresponding to the left side of the colon are selected by another model. Then, the selected frames (which correspond to a withdrawal path and the left side of the colon) are analyzed by one or more inference models to perform one or more inference tasks such as frame-by-frame segment classification and / or regression to, for example, output a disease severity score or other possible analysis results relevant to studying the colon.
[0055] Learning module 460 implements typical learning processing by comparing class inferences 47 to training data labels to compute a loss (error) according to a selected loss function and then back propagating that loss to adjust learnable parameters in MLP 450-1 and in attention-based network 440.
[0056] Although particular types of classification inferences are described above, a computed “inference”can be any one or more of a prediction, estimate, score, suggestion, categorization, assessment, calculation, or other inference. In some embodiments, rather than using a softmax layer to provide class probabilities, one or more outputs of an MLP can be used to provide regression output in the form of, for example, a predicted score within a range of values.
[0057] VLA 425 is used during training. However, when the trained network is used for post-training inference tasks, VLA 425 is not used (and, if present, may be bypassed).
[0058] FIG. 5 illustrates a method 5000 for processing medical videos to train an attention-based deep learning network, such as one comprising attention-based encoder 440 and inference network 450 shown in FIG. 4, to automatically compute classification inferences for frames of a medical video in accordance with an embodiment of the present disclosure.
[0059] Step 501 pre-processes medical videos in a training set to generate frames, re-sizethem to a uniform size, and mask any annotations. Step 502 encodes each frame of a medical video to obtain a sequence of frame embeddings corresponding to the medical video, using an SSL pre-trained encoder. Step 503 selectively applies one or more temporal augmentations to the sequence of frame embeddings. As illustrated, the one or more temporal augmentations include splitting (i.e., cropping the video to a selected sub-sequence of the sequence of frame embeddings) and reversing. Step 503 outputs an augmented sequence of frame embeddings (which, for some sequences, might be the same as the sequence prior to application of step 503 in the case that it is randomly determined whether one or more operations are applied or not applied).
[0060] Step 504 processes the augmented sequence of frame embeddings through an attention-based network to compute an attention-based vector for each frame embedding of the augmented sequence. Step 505 processes the attention-based vectors to compute classification or regression inferences on a frame-by-frame basis. The number of outputs depends on the application, as previously discussed.
[0061] Step 506 uses frame labels and the computed inferences (e.g., class probabilities) to compute a loss value that quantifies the error in the inferences using a selected loss function. In one example, a cross-entropy loss function is used for a classification application.However, other loss functions may be used, depending on the application. Step 507 determines whether error is now minimized. If yes, then training method 5000 ends. If no, then step 508 adjusts learnable parameters of the classification and attention-based deep learning network to further reduce error and the method returns to step 502.
[0062] FIG. 6 illustrates a method 6000 that may be carried out by VLA 425 (FIG. 4) to selectively perform one or more temporal-augmentation operations. As illustrated, step 601 determines, via random selection, whether to select a sequence Fi of frame embeddings for a splitting (time-cropping, or time trimming) operation. Step 602 determines whethersequence Fi has been selected. If the result of step 602 is no, then step 603 sets the augmented sequence, Faug, equal to the initial sequence Fi. If the result of step 602 is yes, then step 604 initiates the splitting operation by, for example, randomly selecting a starting frame number Rs (an integer) between 0 and N / 2 - L / 2 where N is the number of frames in the sequence Fi and L is an integer greater than 0 and less than or equal to Ni and randomly selecting an ending frame number between N / 2 + L / 2 and N. Step 605 then sets Faug equal to the frame sequence from frame rsto frame re. Thus, after the splitting operation, Faug is a subset of Fi (which might be Fi, or a smaller subset of Fi) In one example, L is set to be 2, but other numbers can also be used.
[0063] Step 606 determines, randomly, whether to select Faug for sequence reversal. Step607 determined whether step 606 has selected Faug for reversal. If the result of step 607 is no, then the frame sequence Faug is not reversed. If the result of step 607 is yes, then step608 reverses the sequence of Faug. Step 609 then outputs Faug. Those skilled in the art will appreciate that, when the sequence is reversed, label data may also need to be adjusted to facilitate matching labels to downstream inferences for purposes of supervised (or weakly supervised) learning. Such adjustments to label data may also be necessary in view of other temporal augmentations such as time-cropping (splitting).
[0064] Method 6000 of FIG. 6 can, in one example, be expressed more formally as the below algorithm.Algorithm 1INPUT: Matrix of embeddings in ith video during training (Fi)OUTPUT: Augmented subset of Fi (F“U9)rbetween 0 and 1; if rispllt> 0.5 then start Random integer between 0 and Ni / 2 - L / 2 where L = {1 e Z 1 1 > 0 and 1 < Ni} riSend Random integer between 0 and 12 Ni - L / 2 where L = {1 e Z 1 1 > 0 and 1 < Ni} j-.augFnsplit , splitri Rows ri:ptartto rndfrom Fielse|FaugF. endReverse Ranc|omf]oat number between 0 and 1 ; embeddings (rows) in FU9end return F9
[0065] Those skilled in the art will appreciate that the example of FIG. 6 illustrates one example of introducing temporal augmentations to training video sequences. In the above example, a given video sequence processed by method 6000 may be split (cropped) only, reversed only, split and reversed, or left unchanged based on randomizing operations in the method. Alternatively, these and / or other operations may be performed based on random or pre-defined selection techniques. In one general example for time-cropping operations, a function “min” is defined in a way that there are at least min(A, N) frames available in the video after time cropping, where A is defined based on the severity of the desired augmentation and N is the number of frames available in the video. In one example, the lower A is, the more severe the augmentation, i.e., the more aggressive the time cropping.
[0066] FIG. 7 illustrates a method 7000 for processing medical videos to train an attention-based deep learning network, such as one comprising attention-based encoder 440 and inference network 450 shown in FIG. 4, to automatically compute classification inferences for frames of a medical video in accordance with an embodiment of the present disclosure.
[0067] Step 701 pre-processes a medical video to generate frames and re-size them to a uniform size. Step 702 encodes each frame of a medical video to obtain a sequence of frame embeddings corresponding to the medical video, using an SSL pre-trained encoder.
[0068] Step 703 processes the sequence frame embeddings corresponding to the medical video through an attention-based network to compute an attention-based vector for eachframe embedding of the sequence. Step 704 processes the attention-based vectors to compute classification inferences on a frame-by-frame basis. The number of classes depends on the application, as previously discussed.
[0069] FIG. 8 illustrates a medical video processing system configured to train an attention-based deep learning network (video analyzer 140) to automatically analyze and compute one or more inferences for a medical video. Frame-embeddings 804 each having dimension D2 x n from a training set are processed by an attention-based deep learning network which, in this example, comprises attention-based encoder 810 and attention weighting network 820. Attention-based encoder 810 processes sequence 804 of frame embeddings and outputs sequence 805 of attention vectors having dimensions D2 x n. Each attention vector corresponds to a frame embedding that has been processed by attentionbased encoder 810. An attention score for each frame is output at 806 and input into aggregator 830 along with the attention vectors for each frame 805, to generate an aggregated vector 807 for the patient video having dimension D2 x 1.
[0070] In one example, attention-based encoder 810 is a ViT encoder based on the model described in Dosovitskly et al. “An Image is Worth 16X16 Words”, 2021, hereby incorporated by reference herein. In Vaswani, an extra classification token is added and is used to provide a summarized output token to be used for classifying an image corresponding to a collection of input tokens (which, in Vaswani, each correspond to patches of the image to be analyzed). However, in the case that frame-by-frame classifications are desired, the extra classification token used in Vaswani may not be needed and the output tokens for each input frame embedding are used directly by the inference network. One example of encoder 440 uses a ViT encoder with four layers and eight self-attention heads in each layer.
[0071] The aggregated vector 807 for the patient video is processed by multi-layerperceptron (MLP) 840 to generate a video level score at 808. When doing regression, a Softmax layer is not needed. In one example, the attention encoder has a dropout of 0.25 and the classification network has a dropout of 0.5, a batch size of 1 is used, and an Adam optimizer is used with a learning rate of 10-5and a weight decay of 10-6.
[0072] Learning module 460 implements typical learning processing by comparing class inferences 47 to training data labels to compute a loss (error) according to a selected loss function and then back propagating that loss to adjust learnable parameters in MLP 840 and in attention weighting 820 and attention encoder 810.
[0073] FIG. 9 illustrates a medical video processing system (e.g., video analyzer 140) for computing classification inferences for frames of a medical video using a trained attentionbased network in accordance with an embodiment of the present disclosure. Frameembeddings 104A that comprise video frames from the medical video showing the endoscopic video in a backward path through the colon are processed through trained attention encoder 810. Each frame embedding 104A has dimension D2 x n. Attention-based encoder 810 processes sequence 104A of frame embeddings and outputs sequence 905 of attention vectors having dimensions D2 x n. Note that, in one example, this assumes that attention encoder 810 has the same output size as does SSL pre-trained encoder 120 referenced in earlier figures. In one example, that applies. However, in alternative examples, depending on the encoders used at each stage of processing, the output dimensions of a downstream attention encoder such as encoder 810 may be different than those of the earlier SSL pre-trained encoder such as encoder 120.
[0074] Each attention vector corresponds to a frame embedding that has been processed by attention-based encoder 810. Attention weighting block 820 assigns an attention score 806 for each attention vector 905. Scores 906 and attention vectors 905 are processed by aggregator 830 to generate an aggregated vector 907 for the patient video having dimensionD2 x 1. The aggregated vector 907 for the patient video is processed by multi-layer perceptron (MLP) 840 to generate a video-level score at 908.
[0075] FIG. 10 illustrates a medical video processing system (e.g., video analyzer 140) for computing classification inferences for frames of a medical video using a trained attentionbased network in accordance with an embodiment of the present disclosure, e.g., as shown in FIG. 2. Frame-embeddings 104B that comprise video frames from the medical video showing the endoscopic video in a backward path through the colon are first sorted into groups for different segments of the colon based on the segment classifier results (e.g., from frame-by-frame classifier 130). In one example illustrated in FIG. 10, only frames classified as from the backward path and of segments from the left side of the colon are processed through trained attention encoders 810, attention weighting networks 820, aggregators 830, and MLPs 840 in respective groups 104B-1, 104B-2, and 104B-3 corresponding to each of the (e.g., three) colon segments comprising the left side of the colon. Each group 104B-1, 104B-2, and 104B-3 is processed separately through an attention encoder 810, attention weighting network 820, aggregator 830, and MLP 840. Each frame embedding 104B-1, 104B-2, and 104B-3 has dimension D2 x n. Attention-based encoder 810 processes the sequences of frame embeddings and outputs attention vectors 905 having dimensions D2 x n. Each attention vector corresponds to a frame embedding that has been processed by attention-based encoder 810. Attention weighting block 820 assigns an attention score 906 for each attention vector 905. Aggregator 830 processes attention vectors 905 along with attention scores 906 for each frame 905, to generate an aggregated vector 907 for the segment having dimension D2 x 1. The aggregated vector 907 for each segment is processed by an MLP 840 to generate a segment 1 score, a segment 2 score, and a segment 3 score, which can be averaged (or weighted averaged) to generate a patient level score. In this example, attention scores per frame 906 and the segment-level scores, in addition to thepatient-level score are input into video review interface 150 for display as discussed further below.
[0076] FIG. 11 illustrates a medical video processing system (e.g., video analyzer 140) for computing classification inferences for frames of a medical video using a trained attentionbased network in accordance with an embodiment of the present disclosure as shown in FIG. 3. Frame-embeddings 104A that comprise video frames from the medical video showing the endoscopic video in a backward path through the colon are first sorted into groups for different segments of the colon based on the segment classifier results (e.g., from frame-by- frame classifiers 130). In one example illustrated in FIG. 11, frames classified as from the backward path and of segments from both the left and right sides of the colon are processed through trained attention encoders 810, attention weighting networks 820, aggregators 830, and MLPs 840 in respective groups 104A-1, 104A-2, 104A-3, 104A-4 and 104A-5 corresponding to each of the (e.g., five or six) colon segments comprising the full colon. A colon is generally understood to have six anatomical segments: rectum, sigmoid colon, descending colon, transverse colon, ascending colon, and ileum. FIG. 5 shows five groups of frame embeddings (with corresponding processing paths) rather than six for ease of illustration and because, in some implementations, certain segments can be bundled together for classification purposes. The illustrated example assumes the sigmoid and descending segments are bundled together for classification purposes. Each group 104A-1, 104A-2, 104A-3, 104A-4, and 104A-5 is processed separately through an attention encoder 810, attention weighting network 820, aggregator 830, and MLP 840. Each frame embedding 104A-1, 104A-2, 104A-3, 104A-4 and 104A-5 has dimension D2 x n. Attention-based encoder 810 processes the sequences of frame embeddings and outputs sequence 905 of attention vectors having dimensions D2 x n. Each attention vector corresponds to a frame embedding that has been processed by attention-based encoder 810. Attention weightingblock 820 assigns an atention score 906 for each atention vector 905. Aggregator 830 processes atention vectors 905 along with atention scores 906 for each frame 905, to generate an aggregated vector 907 for the segment having dimension D2 x 1. The aggregated vector 907 for each segment is processed by an MLP 840 to generate a segment 1 score, a segment 2 score, a segment 3 score, a segment 4 score, and a segment 5 score which can be averaged to generate a patient level score. In this example, atention scores per frame 906 and the segment-level scores, in addition to the patient-level score are input into video review interface 150 for display as discussed further below.
[0077] Those skilled in the art will appreciate that, in a regression inference application, the output of an MLP such as MLP 840 can be from a single artificial neuron to provide a single value, which can be over a range of continuous values using a desired number of significant digits. Moreover, those skilled in the art will further appreciate that a simple scaling operation can convert that number to an appropriate value in a desired range, such as from 0 to 3, as in the case of an MES score, or other ranges depending on the application. Therefore, such scaling operations are not further illustrated or detailed herein. Various ranges can be used for endoscopic applications or other disease-assessment applications.
[0078] FIG. 12 illustrates a method 1200 for processing medical videos to compute inferences for frames of a medical video using an atention-based deep learning network trained in accordance with an embodiment of the present disclosure. At step 1201, each frame of the video is encoded using an SSL pre-trained encoder to obtain a sequence of frame embeddings corresponding to the medical video. At step 1202, a frame-by-frame classifier is used to determine which frame embeddings are from the endoscope withdrawal path. Next, at 1202 it is determined whether the whole colon is being examined and scored in the medical video. If the whole colon is being examined, then all of the frame embeddings are retained for further analysis, and processing proceeds directly to determining whether toassign a score for each segment at 1205.
[0079] If not, then the frame embeddings are submitted at 1203 to a frame-level left / right colon classification network, and frames of interest are retained at step 1204 before proceeding to determining whether to assign a score for each segment at step 1205. If the result of step 1205 is no, then the method proceed to 1209 and the attention-based deep learning network is used to process all backward path (withdrawal path) frame embeddings together to obtain a video-level score.
[0080] If segment-based scoring is desired (result of step 1205 is yes), then step 1206 submits the frame embeddings to a segment classifier. Based on the segment classifications, at step 1207 an attention-based deep learning network is used to process each segment’s backward path frame embeddings to obtain scores for each segment. Then a video-level score is obtained by averaging (or weighted averaging) the scores for each segment at step 1208.
[0081] FIG. 13 shows an example of a video review interface 1300 for display on one or more devices, in accordance with an embodiment of the present disclosure. For example, for an endoscopic medical video 101 of a patient’s colon, the video review interface could display video frame 12 at 1301, along with video timestamp data 1303 and frame data 1304, along with relevant information 1302. In the example shown, information 1302 includes an automated (Al) MES score, a standard MES score, a segment classification, and bleeding, erosion, and vascular scores. In alternatives, various information can be provided including subsegment location, scope location, scope direction, a biopsy location (if any), and relevant disease severity scoring information (e.g., using any number of scoring systems known in the art or otherwise, such as Mayo Endoscopic Subscore (MES) or Ulcerative Colitis Endoscopic Index of Severity (UCEIS)). Continuous scores, having non-integer values having one or more significant digits beyond the decimal point, can also be displayed onvideo review interface 150, and may be displayed either as numbers 1302 via heat maps 1305 and 1306, along with the ability to view attention maps and high attention regions at 1307, which provides the user with the ability to scroll directly to frames having the highest impact on the AI-MES score (e.g., the MES score calculated using embodiments described in the present disclosure). The video review interface 1300 can also show other features of interest for disease severity assessment shown in the displayed frame, including MES data, bleeding data, erosion data, and vascular pattern data, as shown in exemplary information bar 1302. Other relevant data may also be included. For example, for Crohn’s disease patients, further data on size, surface, and / or narrowing of ulcers, and the percentage of area of an ulcer could also be displayed.TRAINING DATA AND SELECTED RESULT EXAMPLES
[0082] Models consistent with embodiments presented herein were trained and validated using a large collection of clinical trials including UNIFI: NCT02407236, JAK- UC: NCT01959282, SEA VUE: NCT03464136, and TRIDENT: NCT02877134. These trials account for over one thousand hours of endoscopy video (sixty million frames). Models consistent with embodiments presented herein show good performance. For example, in some demonstrations, AUC = 0.8 for MES, AUC=0.79 for UCEIS, and AUC=0.93 for segment-wise endoscope location. Also, in a study in which embodiments consistent with the present disclosure were trained on ulcerative colitis clinical trial data (UNIFI: NCT02407236, Phase 3, 965 subjects, 3,128 videos) and a continuous MES score was auto-generated by those embodiments, discriminatory power greater than human MES scoring was demonstrated for detecting differences between a placebo and treatment with 200mg of guselkumab. Further details of relevant results are disclosed in U.S. Provisional Application No. 63 / 555,883, fded on February 20, 2024, incorporated herein by reference.
[0083] FIG. 14 shows an example of a computer system 1400, one or more of which maybe used to implement one or more of the apparatuses, systems, and methods illustrated herein. Computer system 1400 executes instruction code contained in a computer program product 1460. Computer program product 1460 comprises executable code in an electronically readable medium that may instruct one or more computers such as computer system 1400 to perform processing that accomplishes the exemplary method steps performed.
[0084] The electronically readable medium may be any transitory or non-transitory medium that stores information electronically and may be accessed locally or remotely, for example via a network connection. The medium may include a plurality of geographically dispersed media each configured to store different parts of the executable code at different locations and / or at different times. The executable instruction code in an electronically readable medium directs the illustrated computer system 1400 to carry out various exemplary tasks described herein. The executable code for directing the carrying out of tasks described herein would be typically realized in software. However, it will be appreciated by those skilled in the art, that computers or other electronic devices might utilize code realized in hardware to perform many or all the identified tasks. Those skilled in the art will understand that many variations on executable code may be found that implement exemplary methods within the spirit and the scope of the disclosure.
[0085] The code or a copy of the code contained in computer program product 1460 may reside in one or more storage persistent media (not separately shown) communicatively coupled to system 1400 for loading and storage in persistent storage device 1470 and / or memory 1410 for execution by processor 1420. Computer system 1400 also includes I / O subsystem 1430 and peripheral devices 1440. I / O subsystem 1430, peripheral devices 1440, processor 1420, memory 1410, and persistent storage device 1470 are coupled via bus 1450. Like persistent storage device 1470 and any other persistent storage that might containcomputer program product 1460, memory 1410 is a non-transitory media (even if implemented as a typical volatile computer memory device). Moreover, those skilled in the art will appreciate that in addition to storing computer program product 1460 for carrying out processing described herein, memory 1410 and / or persistent storage device 1470 may be configured to store the various data elements referenced and illustrated herein.
[0086] Those skilled in the art will appreciate computer system 1400 illustrates just one example of a system in which a computer program product in accordance with the disclosure may be implemented. To cite but one example, execution of instructions contained in a computer program product may be distributed over multiple computers, such as, for example, over the computers of a distributed computing network.
[0087] Instructions for implementing an artificial neural network or other deep learning network may reside in computer program product 1460. When processor 1420 is executing the instructions of computer program product 1460, the instructions, or a portion thereof, are typically loaded into working memory 1410 from which the instructions are readily accessed by processor 1420.
[0088] Processor 1420 may comprise multiple processors which may comprise respective additional working memories (additional processors and memories not individually illustrated) including one or more graphics processing units (GPUs) comprising at least thousands of arithmetic logic units supporting parallel computations on a large scale. GPUs are often utilized in deep learning applications because they can perform the relevant processing tasks more efficiently than typical general-purpose processors (CPUs). Processor 1420 may additionally or alternatively comprise one or more specialized processing units comprising systolic arrays and / or other hardware arrangements that support efficient parallel processing. Such specialized hardware may work in conjunction with a CPU and / or GPU to carry out the various processing described herein. Such specialized hardware may compriseapplication specific integrated circuits and the like (which may refer to a portion of an integrated circuit that is application-specific), field programmable gate arrays and the like, or combinations thereof. However, a processor such as processor 1420 may be implemented as one or more general purpose processors (preferably having multiple cores) without necessarily departing from the spirit and scope of the present disclosure.
[0089] While the word inference or infer may be variously used herein, one having skill in the art will understand that the systems and methods described herein are not so limited and that the term inference herein may indicate the performance of any of a variety of calculations and / or generation of a variety of outputs which may include, without limitation, any one or more of inferences, scores, estimates, predictions, projections, suggestions, recommendations, classifications, categorizations, annotations, conclusions, or the like or any combination of the foregoing.ADDITIONAL EMBODIMENTS
[0090] The present disclosure further contemplates one or more additional embodiments, some of which are set forth as example embodiments below.
[0091] Embodiment 1. A computer system comprising: one or more processors; a storage medium storing (a) instructions and (b) a dataset representing video data, wherein the video data comprises endoscopic video data; an encoding module that uses the one or more processors to encode a set of frames from the dataset as highly-dimensional vector data, wherein (a) the set of frames represents data comprising a plurality of views of an area of interest, and (b) the highly-dimensional vector data contains information representing each of a plurality of views of the area of interest; one or more attention modules, comprising at least one of: i. a first attention module that uses the one or more processors to generate, from the highly-dimensional vector data, a video-levelatention map. ii. a second atention module that uses the one or more processors to generate, from the highly-dimensional vector data, a frame-level atention map; iii. a third atention module that uses the one or more processors to generate, from the highly-dimensional vector data, a pixel-level atention map, and iv. a fourth atention module that uses the one or more processors to generate, from the highly dimensional vector data, a temporal- sequence-level atention map; and a prediction module that uses the one or more processors to output an assessment based on one or more of: the video-level atention map, the frame-level atention map, the pixellevel atention map, and the temporal-sequence-level atention map.
[0092] Embodiment 2. The computer system of embodiment 1, wherein the assessment comprises at least one of: a disease severity estimation, an assessment of patient health, a measurement of disease progression, and a medical diagnostic determination.
[0093] Embodiment 3. The computer system of embodiment 2, wherein the assessment relates to one or more of: inflammatory bowel disease, Crohn’s disease, and ulcerative colitis.
[0094] Embodiment 4. The computer system of embodiment 2, wherein the assessment comprises a continuous disease severity score.
[0095] Embodiment 5. A method of using one or more computers to analyze one or more medical videos corresponding to a patient exam, the method comprising: a. pre-processing video data to obtain a plurality of video frames correspondingto the one or more medical videos corresponding to the patient exam; b. encoding the plurality of video frames using a pre-trained encoder to obtain a plurality of frame embeddings corresponding to respective frames of the plurality of video frames, wherein the pre-trained video encoder has been pretrained using self-supervised learning; c. concatenating the plurality of frame embeddings; d. processing the plurality of frame embeddings in an attention-based deep learning network to obtain a summary vector resulting from attention processing that attends to each of the plurality of frame embeddings relative to other of the plurality of frame embeddings; and e. submitting the summary vector to a classification network to obtain one or more outputs regarding a patient corresponding to the medical videos from the patient exam.
[0096] Embodiment 6. The method of embodiment 5, wherein the one or more medical videos comprise a plurality of videos such that the plurality of frame embeddings that are concatenated for further processing include frame embeddings from each of the plurality of videos.
[0097] Embodiment 7. The method of any of embodiments 5-6 wherein the patient exam is a transthoracic echocardiogram.
[0098] Embodiment 8. The method of any of embodiments 5-7 wherein the pretrained encoder comprises a vision transformer (ViT).
[0099] Embodiment 9. The method of any of embodiments 5-7 wherein the pretrained encoder comprises a convolutional neural network.
[0100] Embodiment 10. The method of any of embodiments 5-9 wherein the pre-trained encoder is trained using SimCLR.
[0101] Embodiment 11. The method of any of embodiments 5-9 wherein the pretrained encoder is trained using DINO.
[0102] Embodiment 12. The metho of any of embodiments 5-9 wherein the pretrained encoder is trained using DIN0v2.
[0103] Embodiment 13. The method of any of embodiments 5-12 wherein the attention-based deep learning network is a transformer multi -head attention network.
[0104] Embodiment 14. The method of embodiment 13 wherein the transformer multihead attention network is a Set Transformer.
[0105] Embodiment 15. The method of any of embodiments 5-12 wherein the attention-based deep learning network is a multi-instance attention network.
[0106] Embodiment 16. The method of any of embodiment 5-12 wherein the attentionbased deep learning network comprises a convolutional neural network, an attention block, and an aggregator.
[0107] Embodiment 17. The method of embodiment 15 wherein the attention block associates respective learnable parameters with respective frame embedding outputs of the convolutional neural network.
[0108] The method of embodiment 17 wherein the aggregator applies the respective learnable parameters to the respective frame embedding outputs and sums the resulting attention-scored frame embedding output to obtain the summary vector.
[0109] The method of any of embodiments 5-6 and 8-18 wherein the patient exam is an endoscopy.
[0110] While the present disclosure has been particularly described with respect to the illustrated embodiments, it will be appreciated that various alterations, modifications, andadaptations may be made based on the disclosure and are intended to be within the scope of the disclosure. While the disclosure has been described in connection with what are presently considered to be the most practical and preferred embodiments, it is to be understood that the present disclosure is not limited to the disclosed embodiments but, on the contrary, is intended to cover various modifications and equivalent arrangements included within the scope of the underlying principles of the invention as described by the various embodiments referenced above and below.
Claims
CLAIMSWhat is claimed is:
1. A computerized deep-learning system for processing medical videos to assess disease severity of a patient, the computerized deep-learning system comprising: a pre-processor configured to pre-process video data to obtain a plurality of video frames corresponding to one or more medical videos corresponding to the patient; an encoder configured for encoding a plurality of video frames to obtain respective frame embeddings corresponding to respective frames of the plurality of video frames, wherein the encoder has been pre-trained using self-supervised learning; one or more video frame classifiers, each video frame classifier comprising an attention-based deep learning network configured to compute respective attention vectors by processing the respective frame embeddings and to use the respective attention vectors to compute, on a frame-by-frame basis, respective frame-level inferences corresponding to the respective frames; and a video analyzer comprising an attention-based deep learning network configured to use at least some of the respective frame-level inferences and the respective frame embeddings to calculate at least one disease severity assessment of the patient.
2. The computerized deep-learning system of claim 1 wherein the one or more video frame classifiers comprises a first frame classifier configured to compute first respective frame-level inferences by processing the respective frame embeddings and using the first respective frame-level inferences to identify a first set of frames from the medical videos for further analysis.
3. The computerized deep learning system of claim 2 wherein the one or more video frame classifiers further comprises:a second frame classifier configured to compute second respective frame-level inferences by processing the respective frame embeddings corresponding to the first set of frames and using the second respective frame-level inferences to identify a second set of frames from the medical videos for further analysis.
4. The computerized deep-learning system of any of claims 2-3 wherein: the one or more medical videos comprise an endoscopy video; and the first respective frame-level inferences are regarding whether the frames are from a forward path or a withdrawal path of the endoscopy video and the first set of frames are inferred to be from the withdrawal path of the endoscopy video.
5. The computerized deep-learning system of any of claims 3-4 wherein: the one or more medical videos comprise an endoscopy video; and the second respective frame-level inferences are regarding whether the frames are from a left colon area or a right colon area and the second set of frames are inferred to be from a left colon area.
6. The computerized deep-learning system of any of claims 3-5 wherein the one or more video frame classifiers further comprises a third frame classifier configured to compute third respective frame-level inferences by processing the respective frame embeddings corresponding to the second set of frames and using the second respective frame-level inferences to identify third respective framelevel inferences comprising segment classifications corresponding to left colon segments.
7. The computerized deep-learning system of claim 6, wherein the left colon segments comprise one or more of: a descending colon, a sigmoid colon, and a rectum.
8. The computerized deep-learning system of claim 4 wherein: the one or more medical videos comprise an endoscopy video; andthe second computed inferences comprise segment classifications corresponding to colon segments.
9. The computerized deep-learning system of claim 8 wherein the colon segments comprise one or more of: an ileum, an ascending colon, a transverse colon, a descending colon, a sigmoid colon, and a rectum.
10. The computerized deep-learning system of any of claims 6-7, wherein the at least one disease severity assessment comprises an average of disease severity scores for each of the left colon segments.
11. The computerized deep-learning system of any of claims 8-9, wherein the at least one disease severity assessment comprises an average of disease severity scores for each of the colon segments.
12. The computerized deep-learning system of claim 11, wherein the average of disease severity scores for each of the left colon segments comprises a weighted average.
13. The computerized deep-learning system of claim 12, wherein the average of disease severity scores for each of the colon segments comprises a weighted average.
14. The computerized deep-learning system of any of claims 1-13, further comprising a video review interface displayed on a user device, wherein the video review interface is configured to interactively display the one or more video frames and the one or more disease severity assessments.
15. The computerized deep-learning system of any of claims 1-14, wherein the at least one disease severity assessment comprises a continuous disease severity score.
16. A method of using one or more computers to assess disease severity of a patient, the method comprising: pre-processing video data to obtain a plurality of video frames corresponding to toone or more medical videos of the patient; encoding the plurality of video frames using a pre-trained encoder to obtain respective frame embeddings corresponding to respective frames of the plurality of video frames, wherein the encoder has been pre-trained using self-supervised learning; computing one or more respective frame-level inferences corresponding to the respective frames of the plurality of video frames from a plurality of respective attention vectors, wherein the respective frame embeddings are processed using a frame-level attention-based deep learning network to generate the plurality of respective attention vectors and compute on a frame-by-frame basis, respective frame-level inferences corresponding to the respective frames; and calculating at least one disease severity assessment of the patient using a video-level attention-based deep learning network based on at least some of the respective frame-level inferences and the respective frame embeddings.
17. The method of claim 16, wherein the one or more respective frame-level inferences comprises first respective frame-level inferences computed by processing the respective frame embeddings to identify a first set of frames from the medical video for further analysis.
18. The method of claim 17, wherein the one or more respective frame-level inferences comprises second respective frame-level inferences by processing the respective frame embeddings corresponding to the first set of frames to identify one or more second sets of frames from the medical video for further analysis.
19. The method of claim 18, wherein the one or more respective frame-level inferences comprises third respective frame-level inferences by processing the respective frame embeddings corresponding to the second set of frames to identify one or more third sets of frames.
20. The method of any of claims 17-19 wherein: the one or more medical videos comprise endoscopy videos; and the first respective frame-level inferences are regarding whether the frames are from a forward path or a backward path of the endoscopy video and the first set of frames are inferred to be from the backward path of the endoscopy video.
21. The method of any of claims 18-20 wherein: the one or more medical videos comprise endoscopy videos; and the second respective frame-level inferences are regarding whether the frames are from a left side of the colon or a right side of the colon and the second set of frames are inferred to be from the left side of the colon.
22. The method of any of claims 19-21 wherein: the one or more medical videos comprise endoscopy videos; and the third respective frame-level inferences are regarding whether the frames are from one or more of: a descending colon, a sigmoid colon, and a rectum.
23. The method of claim 20 wherein the second respective frame-level inferences are regarding whether the frames are inferred to be from one or more of: a rectum, a sigmoid colon, a descending colon, a transverse colon, an ascending colon, or an ileum.
24. The method of any of claims 16-23 wherein one or more temporal augmentation operations are selectively applied to sequences of frame embeddings corresponding to training videos during training of the attention-based deep learning encoder.
25. The method of claim 24 wherein the one or more temporal augmentations comprise time-cropping and reversing.
26. The method of any of claims 24-25 wherein the one or more temporal augmentations are selectively applied based on a selection function such that, for each sequence of aplurality of sequences of frame embeddings, none, one, or more than one of the temporal augmentations is applied.
27. The method of any of claims 24-26 wherein the selection function randomly selects whether to apply each of the one or more temporal augmentations to a sequence of frame embeddings that is currently being processed.
28. The method of any of claims 16-27, wherein the at least one disease severity assessment comprises a continuous disease severity score.
29. The method of any of claims 16-28, further comprising displaying the one or more video frames and the one or more disease severity assessments on a user interface of a user device.
30. A non-transitory computer readable medium comprising a plurality of computer readable instructions, which upon execution by at least one processor, performs the following operations to assess disease severity of a patient: pre-processing video data to obtain a plurality of video frames corresponding to one or more medical videos of the patient; encoding the plurality of video frames using a pre-trained encoder to obtain respective frame embeddings corresponding to respective frames of the plurality of video frames, wherein the encoder has been pre-trained using self-supervised learning; computing one or more respective frame-level inferences corresponding to the respective frames of the plurality of video frames from a plurality of respective attention vectors, wherein the respective frame embeddings are processed using a frame-level attention-based deep learning network to generate the plurality of respective attention vectors and compute on a frame-by-frame basis, respective frame-level inferences corresponding to the respective frames; andcalculating at least one disease severity assessment of the patient using a video-level attention-based deep learning network based on at least some of the respective frame-level inferences and the respective frame embeddings.
Citation Information
Patent Citations
US202463555883P