Frame-level automated analysis for medical videos

WO2025104705A3PCT designated stage expired Publication Date: 2025-07-17JANSSEN RESEARCH & DEVELOPMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2024/061440
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-20
Filing Date
2024-11-15
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing medical video analysis methods struggle to accurately identify anatomical regions in endoscopy videos due to inter-patient variations in segment length and bowel elasticity, and they have limitations in capturing long-range temporal dependencies.

Method used

The proposed solution involves a frame-level encoder pre-trained using self-supervised learning techniques to obtain frame-level embeddings, which are then processed by an attention-based network comprising a vision transformer encoder and a multi-layer perceptron for inference tasks such as classification and segmentation. Additionally, a video-level augmenter is used to apply temporal augmentations during training.

Benefits of technology

This approach effectively addresses the challenges of inter-patient variations and long-range temporal dependencies, achieving improved accuracy in anatomical region identification and classification in medical videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024061440_17072025_PF_FP_ABST
    Figure IB2024061440_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of computerized analysis systems and methods of medical videos are disclosed. In some embodiments, a downstream attention-based network comprises an attention-based encoder and an inference network. In some embodiments, the attention-based encoder comprises a vision transformer (ViT) encoder that processes those frame embeddings to obtain attention-based vector representations, one corresponding to each frame, which can then be processed through the MLP or other neural network for inference tasks such as classification, regression, and / or segmentation. In some embodiments, during training of the downstream attention-based network, a video-level augmenter is used to selectively apply temporal augmentations to the sequence of frame embeddings before they are processed by the downstream attention-based network. Some embodiments are trained and adapted specifically for analysis of endoscopy videos. These and other aspects of the disclosure are further detailed herein.
Need to check novelty before this filing date? Find Prior Art

Description

Frame-Level Automated Analysis for Medical VideosCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application 63 / 599,994, filed on November 16, 2023 and U.S. Provisional Application 63 / 555,883, filed on February 20, 2024. This application also shares some subject matter with International Application PCT / IB2024 / 057930, filed on August 16, 2024. The contents of these applications are incorporated by reference herein.BACKGROUND

[0002] This disclosure relates generally to computerized technology for processing videos of an area of interest of a patient.SUMMARY

[0003] In some medical video analysis applications, it is helpful to automatically identify the anatomical region in video frames as a camera or other sensor traverses one or more areas of interest of a patient. Endoscopy analysis for evaluating inflammatory bowel disease (IBD) is one such application. For example, in typical Crohn’s Disease assessment, severity is evaluated using the Simple Endoscopic Score for Crohn’s Disease (SES-CD), determined across five anatomic segments: rectum (RM), left colon / sigmoid (LC), transverse colon (TC), right colon (RC), and ileum (IL). Automated endoscopy analysis relies in part on automated classification of video frames across these five anatomic segments.

[0004] Several past methods have used fixed points and distance traveled to map current cameral location to bowel segments based on templates for colon segmentation. However, these approaches cannot easily account for inter-patient variations in segment length and bowel elasticity.

[0005] Some methods have avoided these issues by trying to directly classify framesegments based on the image information in the frame using convolutional neural networks (CNNs) or Long Short-Term Memory (LSTM) networks. However, such methods have limitations in their ability to effectively capture long-range temporal dependencies.

[0006] Transformers have shown promise in other contexts for capturing long-range temporal dynamics. However, the memory cost associated with transformers makes their use in the context of long medical videos challenging. For example, an endoscopy video can exceed 60,000 frames.

[0007] Embodiments of the present disclosure address these challenges through systems and methods that leverage a frame-level encoder that is pre-trained using self-supervised learning (SSL) techniques to obtain frame-level embeddings corresponding to a sequence of frames in a medical video. In some embodiments, a downstream attention-based network comprises an attention-based encoder and an inference network, the inference network may comprise a multi-layer perceptron (MLP) or other neural network. In some embodiments, the attention-based encoder comprises a vision transformer (ViT) encoder that processes those frame embeddings to obtain attention-based vector representations, one corresponding to each frame, which can then be processed through the MLP or other neural network for inference tasks such as classification, regression, and / or segmentation. In some embodiments, during training of the downstream attention-based network, a video-level augmenter is used to selectively apply temporal augmentations to the sequence of frame embeddings before they are processed by the downstream attention-based network. These and other variations on embodiments consistent with the present disclosure are more folly disclosed below.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] FIG. 1 illustrates a medical video processing system for training an attention-basednetwork in accordance with an embodiment of the present disclosure.

[0009] FIG. 2 illustrates a medical video processing system for computing classification inferences for frames of a medical video using a trained attention-based network in accordance with an embodiment of the present disclosure.

[0010] FIG. 3 illustrates a method for processing medical videos to train an attentionbased deep learning network to automatically compute inferences for frames of a medical video in accordance with an embodiment of the present disclosure.

[0011] FIG. 4 illustrates a video-level augmentation process to enhance training in accordance with an embodiment of the present disclosure.

[0012] FIG. 5 illustrates a method for processing medical videos to compute inferences for frames of a medical video using an attention-based deep learning network trained in accordance with an embodiment of the present disclosure.

[0013] FIG. 6 shows an example of a computer system, one or more of which may be used to implement one or more of the apparatuses, systems, and methods illustrated herein.

[0014] While embodiments of the present disclosure are described with reference to the above drawings, the drawings are intended to be illustrative. Other embodiments are consistent with the spirit and scope of the disclosure.DETAILED DESCRIPTION

[0015] The various embodiments now will be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, specific examples of practicing the embodiments. This specification may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this specification will be thorough and complete, and will fully convey the scope of thedisclosure to those skilled in the art. Among other things, this specification may be embodied as methods or devices. Accordingly, any of the various embodiments herein may take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. The following specification is, therefore, not to be taken in a limiting sense.

[0016] FIG. 1 illustrates a medical video processing system 1000 in accordance with an embodiment of the present disclosure. This and other embodiments will be described with reference to endoscopy videos. But the underlying principles of the disclosure are applicable to other medical image / video analysis applications.

[0017] This example shows system 1000 for processing a medical video 10, such as endoscopy video, from a medical video training set 100. System 1000 processes training set 100 comprising medical videos 10 to train downstream attention-based deep learning network comprising attention-based encoder 140 and inference network 150 on an inference task. In one example, training set 100 includes over 21 million labeled frames from 1,335 endoscopy videos corresponding to 753 patients fortraining blocks 140 and 150, and 4160 endoscopy videos from four clinical trials and over 61 million frames fortraining block 120.

[0018] Pre-processing block 110 pre-processes the video data corresponding to a video 10 by generating and resizing frames and removing annotations. Block 110 processes the video data such that each frame has uniform dimensions, DI. In this example, after preprocessing, each frame 11 in a sequence 101 corresponding to video 10 has dimensions DI of 224 x 224 x 3 (if using RGB color channels). Other dimensions can be used in different implementations. Pre-processing block 110 prepares a sequence 101 of frames 11 corresponding to a medical video 10. Thus, the full sequence 101 of frames 11 corresponding to medical video 10 has dimensions of N X DI where N is the total number of frames obtained from the video for processing.

[0019] Self-Supervised Learning (SSL) pre-trained encoder 120 processes frames 11 to obtain sequence 102 of frame embeddings 12. There is one frame embedding 12 for each frame 11 processed by SSL pre-trained encoder 120, and the sequential order of frame embeddings 12 is the same as the sequential order of frames 11 in frame sequence 101. Each frame embedding 12 has a dimension D2.

[0020] The value of D2 depends on the underlying SSL model used for SSL pre-training of encoder 120. In one example, a vision transformer encoder (ViT) is pre-trained using a DIN0v2 objective (see “DIN0v2” approach described by Oquab et al. in “DIN0v2: Learning Robust Visual Features without Supervision” (2023)). However, other SSL pretrained encoders and pre-training techniques can be used without necessarily departing from the spirit and scope of the present disclosure. In another example, an SSL pre-trained encoder, such as encoder 120, can be pre-trained using a DINOvl objective (see “DINOvl” approach described in Caron et al. in “Emerging Properties in Self-Supervised Vision Transformers” (2021)). In yet another example, a convolutional neural network (CNN) encoder such as a ResNet (CNN with residual connections) is used with the “SimCLR” approach (see Chen et al. in “A Simple Framework for Contrastive Learning of Visual Representations” (2020)). All three papers referenced in this paragraph are incorporated by reference herein.

[0021] In one example, SSL pre-trained encoder 120 has been pre-trained using publicly available clinical trial endoscopy videos. In one example, over 61 million frames for over 4,000 endoscopy videos are used for SSL pre-training. In one example, the SSL pre-trained encoder is a ViT-B / 16 encoder pre-trained using the DINOv2 approach. In one example, a batch size of 256 is used for pre-training with a cosine decayed learning rate of 8 x 10-4with an Adam optimizer. In one example, pretraining is done for 15 epochs on four NVIDA A10G GPUs. Also see, Mobadersany, Pooya et al., "Harnessing Temporal Information forPrecise Frame-Level Predictions in Endoscopy Videos,” cited and incorporated by reference above. International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer Nature Switzerland, 2024, for additional details. This paper is hereby incorporated by reference.

[0022] In one such DIN0v2 example, the value of D2 is 768 (one-dimensional embedding). Thus, the full sequence 102 of frame embeddings 12 corresponding to a medical video 10 has dimensions of N X D2 where N is the total number of frames obtained from the video for processing.

[0023] Video-level augmenter (VLA) 130 processes frame embeddings 12 of frame embedding sequence 102. In one embodiment, VLA 130 randomly applies one or more video-level augmentations to a sequence 102 of frame embeddings 12. In one embodiment, VLA 130 randomly determines whether to apply a first temporal modification operation to the sequence 102 of embeddings 12 and then randomly determines whether to apply a second temporal modification operation to the set of embeddings. One example of such an operation is splitting, i.e., cropping the sequence, in which a subset of the frame embeddings in the sequence are kept and the rest are discarded. Another such example is reversing the sequence of frame embeddings. In one example, the first temporal modification is splitting and the second temporal modification is reversing, as described further below in the context of FIG. 4.

[0024] Continuing with the description of FIG. 1, VLA 130 outputs augmented sequence 103 of frame embeddings 13. Sequence 103 has dimensions NA X D2 where NA is the number of frame embeddings after augmentation operations. In this example, if VLA 130 did not apply a splitting operation to sequence 12, then NA = N. Otherwise, NA is less than N. The frame -embeddings are processed by an attention-based deep learning network which, in this example, comprises attention-based encoder 140 and inference network 150.Atention-based encoder 140 processes sequence 103 of frame embeddings 13 and outputs sequence 104 of atention vectors 14 having dimensions NA X D2. Each atention vector 14 corresponds to a frame embedding 12 that has been processed by atention-based encoder 140.

[0025] In one example, atention-based encoder 140 is a ViT encoder based on the model described in Dosovitskly et al. “An Image is Worth 16X16 Words”, 2021, hereby incorporated by reference herein. In Vaswani, an extra classification token is added and is used to provide a summarized output token to be used for classifying an image corresponding to a collection of input tokens (which, in Vaswani, each correspond to patches of the image to be analyzed). However, in the illustrated embodiment, frame-by-frame classifications are desired, and, therefore, the extra classification token used in Vaswani is not needed and the output tokens for each input frame embedding are used directly by the inference network. One example of encoder 140 uses a ViT encoder with four layers and eight self-atention heads in each layer.

[0026] Each atention vector 14 of sequence 104 is processed by inference network 150 including multi-layer perceptron (MLP) 150-1 and softmax layer 150-2 when doing classification. When doing regression, the Softmax layer 150-2 is not needed. In one example, the atention encoder has a dropout of 0.25 and the classification network has a dropout of 0.5, a batch size of 1 is used, and an Adam optimizer is used with a learning rate of 10-5and a weight decay of 10-6. Network 150 outputs a set of computed inferences 15. In classification problems, inferences 15 may be in the form of probabilities for each class to which a frame may be inferred to belong, and in regression problems, they can be the regression values corresponding to each frame. The dimensions of the set of inferences 15 are NA X C where C is the number of classes (or regression outputs in the case of regression inferences), which depends on the application.

[0027] In one example, system 1000 is trained to infer colon segment class probabilities corresponding to the five bowel segments: rectum (RM), left colon / sigmoid (LC), transverse colon (TC), right colon (RC), and ileum (IL). In such an application, “C” is equal to 5, meaning that, for each frame, classification network 150 outputs five class probabilities. In another example, system 1000 is trained to infer whether the frame is taken from a forward path or a withdrawal path of the procedure. In such an application, “C” is equal to 2, meaning that, for each frame, inference network 150 outputs two class probabilities. In another example, system 1000 is trained to infer whether the frame is taken from the left side or right side of the colon. In such an application, C is also equal to 2. In another example, system 1000 is trained to infer, for each frame from the left side of the colon: rectum (RM), sigmoid colon (SC), and descending colon (DC), which of the three colon segments (RM, SC, or DC) the frame is taken from. In such a 3-class example, C is 3.

[0028] In another example, multiple models are combined together in an end-to-end fashion to operate as follows. First, frames corresponding to a withdrawal path are selected by a first trained model. Then, of those, frames corresponding to the left side of the colon are selected by another model. Then, the selected frames (which correspond to a withdrawal path and the left side of the colon) are analyzed by one or more inference models to perform one or more inference tasks such as frame-by-frame segment classification and / or regression to, for example, output a disease severity score or other possible analysis results relevant to studying the colon.

[0029] Learning module 160 implements typical learning processing by comparing class inferences 15 to training data labels to compute a loss (error) according to a selected loss function and then back propagating that loss to adjust learnable parameters in MLP 150-1 and in attention-based network 140.

[0030] Although particular types of classification inferences are described above, acomputed “inference”, in some embodiments, can be a prediction, estimate, score, suggestion, categorization, or other inference. In some embodiments, rather than using a softmax layer to provide class probabilities, one or more outputs of an MLP can be used to provide regression output in the form of, for example, a predicted score within a range of values.

[0031] FIG. 2 illustrates a medical video processing system 2000 for computing classification inferences for a medical video 20 using a trained attention-based network comprising attention-based encoder 140 and inference network 150 in accordance with an embodiment of the present disclosure. In the illustrated example, the trained attention-based encoder 140 is a trained version of the attention-based encoder 140 of system 1000 illustrated in FIG. 1 and the trained inference network 150 is a trained version of the inference network 150 of system 1000 illustrated in FIG. 1.

[0032] Inference system 2000 is similar to system 1000 except without the VLA block 130 and learning block 160 illustrated in FIG. 1. In the example illustrated in FIG. 2, a medical video 20 (such as an endoscopy video) is processed by pre-processing block 210 to generate frames and to re-size the frames to a common size having dimension DI, e.g., 224 x 224 x 3. Pre-processing block 210 outputs a frame sequence 201 of frames 21 corresponding to medical video 20. The same SSL pre-trained encoder 120 used by system 1000 during training (illustrated in FIG. 1) is used during inference by system 2000 illustrated in FIG. 2. In system 2000, SSL pre-trained encoder outputs a sequence 202 of frame embeddings 22. Each frame embedding has a dimension D2 (e.g., 768 for a DIN0v2 SSL pre-trained encoder). Trained attention-based encoder 140 generates a sequence 204 of attention vectors24 which are then processed by inference network 150 to compute classification inferences25 for each frame. As described above in the context of FIG. 1, the number of inferences 25 output for each frame will depend on the number of classes relevant to a particularapplication or the number of outputs relevant to a particular regression applications such as disease severity scoring for each frame.

[0033] FIG. 3 illustrates a method 3000 for processing medical videos to train an attention-based deep learning network, such as one comprising attention-based encoder 140 and inference network 150 shown in FIG. 1, to automatically compute classification inferences for frames of a medical video in accordance with an embodiment of the present disclosure.

[0034] Step 301 pre-processes medical videos in a training set to generate frames, re-size them to a uniform size, and mask any annotations. Step 302 encodes each frame of a medical video to obtain a sequence of frame embeddings corresponding to the medical video, using an SSL pre-trained encoder. Step 303 selectively applies one or more temporal augmentations to the sequence of frame embeddings. In the illustrated embodiment, the one or more temporal augmentations include splitting (i.e., cropping the video to a selected subsequence of the sequence of frame embeddings) and reversing. Step 303 outputs an augmented sequence of frame embeddings (which, for some sequences, might be the same as the sequence prior to application of step 303 in an embodiment in which it is randomly determined whether one or more operations are applied or not applied).

[0035] Step 304 processes the augmented sequence of frame embeddings through an attention-based network to compute an attention-based vector for each frame embedding of the augmented sequence. Step 305 processes the attention-based vectors to compute classification or regression inferences on a frame-by-frame basis. The number of outputs depends on the application, as previously discussed.

[0036] Step 306 uses frame labels and the computed inferences (e.g., class probabilities) to compute a loss value that quantifies the error in the inferences using a selected loss function. In one example, a cross-entropy loss function is used. However, other loss functions may beused. Step 307 determines whether error is now minimized. If yes, then training method 3000 ends. If no, then step 308 adjusts learnable parameters of the classification and attention-based deep learning network to further reduce error and the method returns to step 302.

[0037] FIG. 4 illustrates a method 4000 that may, in one embodiment, be carried out by VLA 130 of the embodiment of FIG. 1 to selectively perform one or more temporalaugmentation operations. As illustrated, step 401 determines, via random selection, whether to select a sequence Fi of frame embeddings for a splitting (time-cropping, or time trimming) operation. Step 402 determines whether sequence Fi has been selected. If the result of step 402 is no, then step 403 sets the augmented sequence, Faug, equal to the initial sequence Fi. If the result of step 402 is yes, then step 404 initiates the splitting operation by, for example, randomly selecting a starting frame number Rs (an integer) between 0 and N / 2 - L / 2 where N is the number of frames in the sequence Fi and L is an integer greater than 0 and less than or equal to Ni and randomly selecting an ending frame number re between N / 2 + L / 2 and N. Step 405 then sets Faug equal to the frame sequence from frame rsto frame re. Thus, after the splitting operation, Faug is a subset of Fi (which might be Fi, or a smaller subset of Fi) In one example, L is set to be 2, but other numbers can also be used.

[0038] Step 406 determines, randomly, whether to select Faug for sequence reversal. Step407 determined whether step 406 has selected Faug for reversal. If the result of step 407 is no, then the frame sequence Faug is not reversed. If the result of step 407 is yes, then step408 reverses the sequence of Faug. Step 409 then outputs Faug. Those skilled in the art will appreciate that, when the sequence is reversed, label data may also need to be adjusted to facilitate matching labels to downstream inferences for purposes of supervised (or weakly supervised) learning. Such adjustments to label data may also be necessary in view of other temporal augmentations such as time-cropping (splitting).

[0039] Method 4000 of FIG. 4 can, in one example, be expressed more formally as the below algorithm.Algorithm 1INPUT: Matrix of embeddings in ith video during training (Fi)OUTPUT: Augmented subset of Fi (F^119)r between 0 and 1; ifr°plit> 0.5 then rt Random integer between 0 and Ni / 2 - L / 2 where L = {1 e Z 1 1 > 0 and 1 < Ni} Random integer between 0 and 12 Ni - L / 2 where L = {1 e Z 1 1 > 0 and 1 < Ni} Rows riiStartto ri endfrom Fie. seend ween 0 and 1 ; embeddings (rows) in FU9end return paug

[0040] Those skilled in the art will appreciate that the example of FIG. 4 illustrates one example of introducing temporal augmentations to training video sequences. In the above example, a given video sequence processed by method 4000 may be split (cropped) only, reversed only, split and reversed, or left unchanged based on randomizing operations in the method. In alternative embodiments, these and / or other operations may be performed based on random or pre-defined selection techniques. In one general example for time-cropping operations, a function “min” is defined in a way that there are at least min(A, N) frames available in the video after time cropping, where A is defined based on the severity of the desired augmentation and N is the number of frames available in the video. In one example, the lower A is, the more severe the augmentation, i.e., the more aggressive the time cropping.

[0041] FIG. 5 illustrates a method 5000 for processing medical videos to train anatention-based deep learning network, such as one comprising atention-based encoder 140 and inference network 150 shown in FIG. 2, to automatically compute classification inferences for frames of a medical video in accordance with an embodiment of the present disclosure.

[0042] Step 501 pre-processes a medical video to generate frames and re-size them to a uniform size. Step 502 encodes each frame of a medical video to obtain a sequence of frame embeddings corresponding to the medical video, using an SSL pre-trained encoder.

[0043] Step 503 processes the sequence frame embeddings corresponding to the medical video through an atention-based network to compute an atention-based vector for each frame embedding of the sequence. Step 504 processes the atention-based vectors to compute classification inferences on a frame-by-frame basis. The number of classes depends on the application, as previously discussed.TRAINING DATA AND SELECTED RESULT EXAMPLES

[0044] In some embodiments, data from four different clinical trials involving 1847 patients and 4160 videos was used for training, validation, and testing. In a specific example, two datasets comprised patients with Crohn’s disease (CD) (ClinicalTrials.gov IDs: NCT03464136, NCT02877134), while the other two datasets included patients with ulcerative colitis (UC) (ClinicalTrials.gov IDs: NCT02407236, NCT01959282). The UC datasets lack anatomic segment labels due to the global disease severity scoring used for UC, which does not include the Simple Endoscopic Score for Crohn’s Disease (SES-CD) criteria. Therefore, UC datasets were only used for SSL pre-training of the initial encoder and not for training the downstream atention-based network. Automatic text extraction using an OCR-based algorithm was used followed by manual review and refinement. This process was applied toall frames from the CD clinical trial endoscopy videos. Approximately 50% of these frames did not have textual annotations for anatomic segments, mainly due to their placement in the forward path. These frames were categorized as “Unknown” and excluded from the downstream supervised learning task. The remaining frames were mapped to standardized anatomic segment labels: IL, RC, TC, LC, and RM. In one example, over 21 million labeled frames from 1335 videos of 753 patients were used fortraining the downstream attentionbased network and for validation and testing, using a 70: 10:20% split ratio, respectively. The 20% of CD data allocated for testing was excluded from use for pre-training the initial encoder.

[0045] Performance of a model consistent with embodiments of the present disclosure was tested against pre-existing as well as other models. Results are shown below in Table 1. Tests measured five different models: Template-based, CNN-based, LTSM-based, foundation-based, and, finally, a model consistent with embodiments of the present disclosure (referred to using the name “EndoFormer” for convenience). All models were optimized on the same training set used for training a downstream attention-based network consistent with the present disclosure.

[0046] The template-based model relied on Yao, H., et al., “Motion-based camera localization system in colonoscopy videos” in Medical Image Analysis 73, 102180 (2021) and used Gunnar Fameback’s dense OF method (see: Fameback, G. Two-frame motion estimation based on polynomial expansion,” in Image Analysis: 13th Scandinavian Conference, SCIA 2003 Halmstad, Sweden, June 29-July 2, 2003 Proceedings 13. pp. 363- 370. Springer (2003)). The CNN-based model relied on Azagra et al., “Endomapper dataset of complete calibrated endoscopy procedures,” Scientific Data 10(1), 671 (2023). The foundation-based model used an SSL pre-trained ViT encoder as an initial encoder, consistent with EndoFormer, but replaced the downstream attention-based deep learningnetwork with linear layers. The LSTM-based model was based on Jin, Y., et al., “Temporal memory relation network for workflow recognition from surgical video. IEEE Transactions on Medical Imaging 40(7), 1911-1923 (2021). It also used an SSL pre-trained ViT encoder as an initial encoder, consistent with EndoFormer, but replaced the downstream attentionbased network with an LSTM network. Except for EndoFormer, none of the models were trained using video-level augmentations such as, for example, those illustrated in FIG. 4. The models were evaluated using area under the curve (AUC), Fl score, accuracy, and adjacent accuracy.TABLE 1

[0047] Variations on embodiments of the present disclosure were tested. In one variation, an SSL pre-trained encoder such as encoder 120 in FIGs. 1 and 2 was implemented using a ViT encoder with a DINOv2 training objective. In another variation, an SSL pre-trained encoder such as encoder 120 in FIGs. 1 and 2 was implemented using a ResNet encoder and pre-trained using a SimCLR objective. Tests showed that both variations performed similarly in terms of accuracy (the latter performing with only slightly less accuracy).However, the ResNet / SimCLR variation took longer to train. Also, variations were tested with and without use of a video-level augmenter such as VLA 130 of FIG. 1. For videos without perturbation during inference (i.e., no random splitting or random reversals), both variations performed similarly. For videos with only random splitting variations during inference, the variation using a VLA during training performed somewhat better. For videos with random reversals during inference, the variation using a VLA during trainingperformed significantly better. For additional details, please see Mobadersany et al., "Harnessing Temporal Information for Precise Frame-Level Predictions in Endoscopy Videos,” cited and incorporated by reference above. International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer Nature Switzerland, 2024, for additional details. This paper is hereby incorporated by reference.

[0048] FIG. 6 shows an example of a computer system 6000, one or more of which may be used to implement one or more of the apparatuses, systems, and methods illustrated herein. Computer system 6000 executes instruction code contained in a computer program product 660. Computer program product 660 comprises executable code in an electronically readable medium that may instruct one or more computers such as computer system 6000 to perform processing that accomplishes the exemplary method steps performed.

[0049] The electronically readable medium may be any transitory or non-transitory medium that stores information electronically and may be accessed locally or remotely, for example via a network connection. The medium may include a plurality of geographically dispersed media each configured to store different parts of the executable code at different locations and / or at different times. The executable instruction code in an electronically readable medium directs the illustrated computer system 6000 to carry out various exemplary tasks described herein. The executable code for directing the carrying out of tasks described herein would be typically realized in software. However, it will be appreciated by those skilled in the art, that computers or other electronic devices might utilize code realized in hardware to perform many or all the identified tasks. Those skilled in the art will understand that many variations on executable code may be found that implement exemplary methods within the spirit and the scope of the disclosure.

[0050] The code or a copy of the code contained in computer program product 660 may reside in one or more storage persistent media (not separately shown) communicativelycoupled to system 6000 for loading and storage in persistent storage device 670 and / or memory 610 for execution by processor 620. Computer system 600 also includes I / O subsystem 630 and peripheral devices 640. I / O subsystem 630, peripheral devices 640, processor 620, memory 610, and persistent storage device 670 are coupled via bus 650. Like persistent storage device 670 and any other persistent storage that might contain computer program product 660, memory 610 is a non-transitory media (even if implemented as a typical volatile computer memory device). Moreover, those skilled in the art will appreciate that in addition to storing computer program product 660 for carrying out processing described herein, memory 610 and / or persistent storage device 670 may be configured to store the various data elements referenced and illustrated herein.

[0051] Those skilled in the art will appreciate computer system 6000 illustrates just one example of a system in which a computer program product in accordance with the disclosure may be implemented. To cite but one example, execution of instructions contained in a computer program product may be distributed over multiple computers, such as, for example, over the computers of a distributed computing network.

[0052] Instructions for implementing an artificial neural network or other deep learning network may reside in computer program product 660. When processor 620 is executing the instructions of computer program product 660, the instructions, or a portion thereof, are typically loaded into working memory 610 from which the instructions are readily accessed by processor 620.

[0053] Processor 620 may comprise multiple processors which may comprise respective additional working memories (additional processors and memories not individually illustrated) including one or more graphics processing units (GPUs) comprising at least thousands of arithmetic logic units supporting parallel computations on a large scale. GPUs are often utilized in deep learning applications because they can perform the relevantprocessing tasks more efficiently than typical general-purpose processors (CPUs). Processor 620 may additionally or alternatively comprise one or more specialized processing units comprising systolic arrays and / or other hardware arrangements that support efficient parallel processing. Such specialized hardware may work in conjunction with a CPU and / or GPU to carry out the various processing described herein. Such specialized hardware may comprise application specific integrated circuits and the like (which may refer to a portion of an integrated circuit that is application-specific), field programmable gate arrays and the like, or combinations thereof. However, a processor such as processor 620 may be implemented as one or more general purpose processors (preferably having multiple cores) without necessarily departing from the spirit and scope of the present disclosure.

[0054] While the word inference or infer may be variously used herein, one having skill in the art will understand that the systems and methods described herein are not so limited and that the term inference herein may indicate the performance of any of a variety of calculations and / or generation of a variety of outputs which may include, without limitation, any one or more of inferences, scores, estimates, predictions, projections, suggestions, recommendations, classifications, categorizations, annotations, conclusions, or the like or any combination of the foregoing.While the present disclosure has been particularly described with respect to the illustrated embodiments, it will be appreciated that various alterations, modifications, and adaptations may be made based on the disclosure and are intended to be within the scope of the disclosure. While the disclosure has been described in connection with what are presently considered to be the most practical and preferred embodiments, it is to be understood that the present disclosure is not limited to the disclosed embodiments but, on the contrary, is intended to cover various modifications and equivalent arrangements included within the scope of the underlying principles of the invention as described by the variousembodiments referenced above and below.

Claims

CLAIMSWhat is claimed is:

1. A method of using one or more computers to analyze one or more medical videos corresponding to one or more patient exams, the method comprising: pre-processing video data to obtain a plurality of video frames corresponding to a medical video of the one or more medical videos, the medical videos corresponding to a patient exam of the one or more patient exams; encoding the plurality of video frames using a pretrained encoder to obtain respective frame embeddings corresponding to respective frames of the plurality of video frames, wherein the pre-trained video encoder has been pre-trained using self-supervised learning; processing the respective frame embeddings using an attention-based deep learning encoder to obtain respective attention vectors resulting from attention processing, the respective attention vectors corresponding to the respective frames of the plurality of video frames; and submitting the respective attention vectors to an inference network to obtain one or more computed inferences corresponding to the respective frames of the of the plurality of video frames.

2. The method of claim 1 wherein the one or more computed inferences comprise at least one computed inference corresponding to each respective frame.

3. The method of any of claims 1-2 wherein the one or more computed inferences comprise classification inferences.

4. The method of any of claims 1-3 wherein the one or more computed inferences comprise regression inferences.

5. The method of any of claims 1-4 wherein the plurality of medical videos compriseendoscopy videos.

6. The method of claim 2 wherein the plurality of medical videos comprise endoscopy videos and the at least one computed inference for each frame comprises at least one classification inference.

7. The method of claim 6 wherein the at least one classification inference comprises a classification with respect to classes comprising rectum, left colon / sigmoid, transverse colon, right colon, and ileum.

8. The method of claim 6 wherein the at least one classification inference comprises an inference with respect to classes comprising a left side of a colon and a right side of a colon.

9. The method of any of claims 1-8 wherein the attention-based deep learning encoder and the inference network is trained using supervised learning.

10. The method of any of claims 1-8 wherein the attention-based deep learning encoder and the inference network is trained using weakly supervised learning.

11. The method of any of claims 1-10 wherein the pretrained encoder comprises a vision transformer (ViT) encoder.

12. The method of any of claims 1-10 wherein the pretrained encoder comprises a convolutional neural network.

13. The method of claim 11 wherein the pretrained encoder is trained using DINOv2.

14. The method of claim 12 wherein the pretrained encoder is trained using SimCLR.

15. The method of any of claims 1-14 wherein one or more temporal augmentation operations are selectively applied to sequences of frame embedding corresponding to training videos during training of the attention-based deep learning encoder and the inference network.

16. The method of claim 15 wherein the one or more temporal augmentations comprisetime-cropping and reversing.

17. The method of any of claims 15-16 wherein the one or more temporal augmentations are selectively applied based on a selection function such that, for each sequence of a plurality sequences of frame embeddings, none, one, or more than one of the temporal augmentations is applied.

18. The method of claim 17 wherein the selection function randomly selects whether to apply each of the one or more temporal augmentations to a sequence of frame embeddings that is currently being processed.

19. A computerized deep-learning system for processing medical videos comprising: a pre-trained encoder configured for encoding a plurality of video frames corresponding to a medical video to obtain respective frame embeddings corresponding to respective frames of the plurality of video frames, wherein the pre-trained video encoder has been pre-trained using self-supervised learning; an attention-based deep learning encoder configured to process the respective frame embeddings obtain respective attention vectors resulting from attention processing, the respective attention vectors corresponding to the respective frames of the plurality of video frames; and an inference network configured to process the respective attention vectors and generate one or more computed inferences corresponding to the respective frames of the of the plurality of video frames.

20. A computer program product stored in a non-transitory tangible medium and comprising instructions executable on one or more processors of one or more computers to implement processing to analyze one or more medical videos, the processing comprising executing the method of any of claims 1-18.

21. A computerized deep-learning system for processing medical videos comprising a plurality of computerized deep-learning systems according to claim 19 wherein: a first computerized deep-learning system of the plurality of computerized deeplearning systems is configured to generate first computed inferences on a frame-by-frame basis to identify a first set of frames from a medical video for further analysis by a second computerized deep-learning system of the plurality of computerized deep learning systems; and the second computerized deep-learning system is configured to generate one or more second computed inferences regarding the first set of frames.

22. The computerized system according to claim 21, further wherein: the one or more second computed inferences regarding the first set of frames are inferences computed on a frame-by-frame basis to identify a second set of frames that is a subset of the first set of frames, the second set of frames being identified for further analysis by a third computerized deep-learning system of the plurality of deep-learning systems; and the third computerized deep-learning system of the plurality of deep-learning system is configured to generate one or more third computed inferences regarding the second set of frames.

23. The computerized system of claim 22 wherein: the medical video is an endoscopy video; the computed inferences are regarding whether frames are from a forward path or a withdrawal path of the endoscopy video and the first set of frames are inferred to be from a backward path of the endoscopy video; the second computed inferences are regarding whether frames are from a left colon area or a right colon area and the second set of frames are inferred to be from a left colonarea; and the one or more third computed inferences comprise a disease severity inference based on the second set of frames.