Automated disease severity assessment based on analysis of medical videos

CN122804257APending Publication Date: 2026-09-22JANSSEN RES & DEV LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580015844.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-15
Filing Date
2025-02-20
Publication Date
2026-09-22

Smart Images

  • Figure CN122804257A_ABST
    Figure CN122804257A_ABST
Patent Text Reader

Abstract

Embodiments of computerized deep learning systems and methods for processing medical videos to assess disease severity of a patient are disclosed. In one or more embodiments, an encoder is configured to encode a plurality of video frames corresponding to a medical video to obtain respective frame embeddings corresponding to respective frames of the plurality of video frames, wherein the encoder has been pre-trained using self-supervised learning. One or more video frame classifiers, each comprising an attention-based deep learning network, process the respective frame embeddings and compute respective frame-level inferences corresponding to the respective frames on a frame-by-frame basis. A video analyzer comprises an attention-based deep learning network and uses the respective frame-level inferences and the respective frame embeddings to compute at least one disease severity assessment of the patient.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications This application claims the benefit of U.S. Provisional Application 63 / 555,883, filed February 20, 2024. This application also shares some subject matter with International Applications PCT / IB2024 / 057930, filed August 16, 2024, and PCT / IB2024 / 061440, filed November 15, 2024. The contents of these applications are incorporated herein by reference. Background Technology

[0002] This disclosure relates in its entirety to computerized techniques for assessing the severity of a patient's disease based on medical videos of the patient's region of interest. Summary of the Invention

[0003] Endoscopic-based assessments of disease severity in inflammatory bowel disease (IBD) clinical trials are typically performed using human reading scoring systems such as the Mayo Endoscopic Severity Index (MES) or the Ulcerative Colitis Endoscopic Severity Index (UCEIS). Computer vision and artificial intelligence (AI) have the potential to automate and improve these measurements. Furthermore, in some cases, typical scoring methods fail to reflect the differences in disease severity before and after treatment with sufficient granularity to measure treatment efficacy with the required level of discriminative power.

[0004] In some specific embodiments of this disclosure, a computerized deep learning system and method for processing medical videos to assess the severity of a patient's disease are disclosed. One or more embodiments include a preprocessor configured to preprocess video data to obtain a plurality of video frames corresponding to one or more medical videos, each corresponding to a patient. An encoder is configured to encode the plurality of video frames of the medical videos to obtain corresponding frame embeddings corresponding to corresponding frames among the plurality of video frames, wherein the encoder has been pre-trained using self-supervised learning. One or more video frame classifiers (each video frame classifier including an attention-based deep learning network) process the corresponding frame embeddings and compute corresponding frame-level inferences on a frame-by-frame basis. A video analyzer includes an attention-based deep learning network and uses the corresponding frame-level inferences and corresponding frame embeddings to compute at least one assessment of the severity of the patient's disease.

[0005] In some implementations, a continuous score can be calculated. In some implementations, this continuous score can be within the range of a standard scale (such as the Mayo Endoscopic Score (MES)). In other implementations, it can be within a non-standard range. The score can be calculated by processing all frames of the video together (or all frames corresponding to the withdrawal path). Alternatively or additionally, frames corresponding to one of a plurality of segments can be processed together to obtain a segment-level score, and the average of the segment-level scores can be taken to obtain a video-level patient score. These and other variations consistent with this disclosure are disclosed more fully below. Attached Figure Description

[0006] Figure 1 An example is illustrated of a medical video analytics system that calculates frame-level and video-level inferences for assessing the severity of a patient's disease according to one or more embodiments of this disclosure.

[0007] Figure 2 An example is illustrated of a medical video analytics system that calculates frame-level and video-level inferences for assessing the severity of a patient's disease according to one or more embodiments of this disclosure.

[0008] Figure 3 An example is illustrated of a medical video analytics system that calculates frame-level and video-level inferences for assessing the severity of a patient's disease according to one or more embodiments of this disclosure.

[0009] Figure 4 An example is illustrated of a medical video processing system configured to train an attention-based deep learning network to automatically compute inferences about frames of a medical video, according to one or more embodiments of the present disclosure.

[0010] Figure 5 An example is illustrated of a method for processing medical videos to train an attention-based deep learning network to automatically compute inferences about frames of the medical videos, according to one or more embodiments of the present disclosure.

[0011] Figure 6 A video-level augmentation process for augmentation training according to one or more embodiments of this disclosure is illustrated.

[0012] Figure 7 Methods for processing medical videos to compute inferences about frames of the medical video using a trained attention-based deep learning network, according to one or more embodiments of the present disclosure, are illustrated.

[0013] Figure 8 An example of a medical video processing system is illustrated, which is configured to train an attention-based deep learning network to automatically analyze and compute one or more inferences about medical videos.

[0014] Figure 9 A medical video processing system according to one or more embodiments of the present disclosure is illustrated for computing one or more inferences about a medical video using a trained attention-based network.

[0015] Figure 10 A medical video processing system according to one or more embodiments of the present disclosure is illustrated for computing one or more inferences about a medical video using a trained attention-based network.

[0016] Figure 11 A medical video processing system according to one or more embodiments of the present disclosure is illustrated for computing one or more inferences about a medical video using a trained attention-based network.

[0017] Figure 12 Methods for processing medical videos to compute inferences about the medical videos using a trained attention-based deep learning network, according to one or more embodiments of the present disclosure, are illustrated.

[0018] Figure 13 An example of a video review interface according to one or more embodiments of this disclosure is shown.

[0019] Figure 14 Examples of computer systems are shown, one or more of which can be used to implement one or more of the devices, systems and methods illustrated herein.

[0020] While embodiments of the present disclosure have been described with reference to the accompanying drawings, the drawings are intended to be illustrative. Other embodiments are consistent with the spirit and scope of this disclosure. Detailed Implementation

[0021] Various embodiments will now be described more fully below with reference to the accompanying drawings, which form part of this document and illustrate specific examples of practical embodiments by way of illustration. However, this specification may be embodied in many different forms and should not be construed as limited to the embodiments listed herein; rather, these embodiments are provided so that this specification will be comprehensive and complete and will fully convey the scope of this disclosure to those skilled in the art. Among other things, this specification may also be embodied as a method or apparatus. Therefore, any of the various embodiments described herein may take the form of a completely hardware implementation, a completely software implementation, or an implementation combining software and hardware aspects. Therefore, the following description should not be construed as limiting.

[0022] Figure 1A medical video analysis system 1000 according to an embodiment of this disclosure is illustrated. This embodiment and other embodiments will be described with reference to endoscopic video. However, the basic principles of this disclosure are applicable to other medical image / video analysis applications.

[0023] This example illustrates a system 1000 for analyzing medical video 101 (such as endoscopic video). System 1000 processes multiple frames 12, which together constitute medical video 101. A preprocessor 110 preprocesses the video data corresponding to medical video 101 by resizing the frames and / or masking any annotations to generate resized and masked frames 103. Block 110 processes the video data such that each frame has a uniform dimension D1. In this example, if there are a total of m1 frames in the video, the size of the resized and masked frame data 103 after preprocessing is D1 × m1. In one specific implementation of this example, after preprocessing, each frame 12 in the sequence corresponding to medical video 101 has a dimension D1 of 224 × 224 × 3 (if RGB color class is used). Other dimensions may be used in different specific implementations. Preprocessing block 110 prepares frame sequence 103, which has a total of m1 frames corresponding to medical video 101. Therefore, the complete sequence of frames corresponding to medical video 101 has the dimension D1×m1 as described above, where m1 is the total number of frames obtained from the video for processing.

[0024] A self-supervised learning (SSL) pre-trained encoder 120 processes frame 103 to obtain a sequence of frame embeddings 104. For each frame 103 processed by the SSL pre-trained encoder 120, there exists a frame embedding 104, and the order of the frame embeddings 104 is the same as the order of frames 12 in the sequence of medical video 101 and in the resized and masked sequence of frames 103. Therefore, if the size of each frame embedding is dimension D2, the complete sequence of frame embeddings has a dimension D2×m1. The value of D2 depends on the underlying SSL model used for SSL pre-training of encoder 120, as follows: Figure 4 The subject of discussion.

[0025] Next, one or more frame-by-frame classifiers 130 are applied. Each frame-by-frame classifier 130 infers a classification for each frame embedding in one or more frame embeddings 104 to generate one or more subsets 105 of frame embeddings with a specific classification. Thus, in one example, subset 105 has a dimension D2×m3, where m3 is the total number of frame embeddings selected from the complete set of m1 frames. Therefore, in this example, m3<=m1.

[0026] Despite the following about Figures 2 to 4This describes a specific type of classification inference, but the computed "inference" can be any one or more of a prediction, estimate, rating, suggestion, classification, evaluation, calculation, or other inference. In some implementations, instead of using a softmax layer to provide class probabilities, one or more outputs of a multilayer perceptron (MLP) can be used to provide a regression output, for example, a predicted rating within a range of values.

[0027] In some examples, one or more frame-by-frame classifiers 130 can be configured to compute classification inferences for medical video 101 using a trained attention-based network, which includes an attention-based encoder and an inference network, as described below. Figure 4 The discussion focuses on this. In the illustrated example, the frame-by-frame classifier 130 includes a trained attention-based network, which includes... Figure 4 The illustrated system 4000's attention-based encoder 440 is a trained version and Figure 4 The illustrated system 4000 is a trained version of the inference network 450.

[0028] As shown below Figure 4 As described in the context, the number of inferences for each frame output will depend on the number of categories associated with a particular application.

[0029] Video analyzer 140 analyzes frame embeddings 105 to generate one or more video-level scores or category inferences. Additionally, frame-level attention and model information, such as attention vectors and / or values ​​representing attention weighting, can be generated. In some examples, all such information can be analyzed and presented to a user (e.g., a physician, healthcare professional, and / or patient) in a video review interface 150 for further analysis. For example, for an endoscopic medical video 101 of a patient's colon, the video review interface could display frame 12, showing an image of frame 12 along with relevant information such as segment classification, subsegment location, endoscope location, endoscope orientation, biopsy location (if any), and relevant disease severity scoring information (e.g., using any number of scoring systems or other methods known in the art, such as the Mayo Endoscopic Severity Index (MES) or the Ulcerative Colitis Endoscopic Severity Index (UCEIS)). Continuous scores with non-integer values ​​can also be displayed on the video review interface 150, along with the ability to view attention maps and high-attention areas, and the ability to scroll directly to frames displaying features of interest (including MES data, bleeding data, erosion data, and vascular pattern data) assessing disease severity. For Crohn's disease patients, additional data on ulcer size, surface and / or narrowing, and percentage of ulcer area can also be displayed.

[0030] Figure 2An example of a medical video processing system 2000 is illustrated, which uses a trained attention-based network to compute classification inferences for frames in medical videos to identify, for example, frames in the withdrawal path (sometimes referred to herein as the “backward path”) and the left colon region of an endoscopic video for further analysis.

[0031] and Figure 1 Similarly, this example illustrates a system 2000 for analyzing medical video 101 (such as endoscopic video). System 2000 processes multiple frames 12, which together constitute medical video 101. Preprocessor 110 preprocesses the video data corresponding to medical video 101 by generating frames, resizing frames, and masking frames to remove annotations, to generate resized and masked frames 103. Block 110 processes the video data such that each frame has a uniform dimension D1.

[0032] Such as about Figure 1 The self-supervised learning (SSL) pre-trained encoder 120, as discussed, processes frame 103 to obtain a sequence of frame embeddings 104. For each frame 103 processed by the SSL pre-trained encoder 120, there exists a frame embedding 104, and the order of the frame embeddings 104 is the same as the order of frames 12 in the sequence of medical video 101 and in the resized and masked sequence of frames 103. Therefore, if the size of each frame embedding is dimension D2, the complete sequence of frame embeddings has a dimension D2×m1. The value of D2 depends on the underlying SSL model used for SSL pre-training of the encoder 120, as follows: Figure 4 The subject of discussion.

[0033] Next, three frame-by-frame classifiers, 230A, 230B, and 230C, are applied. Figure 1 Similar to the frame-by-frame classifier 130, each frame-by-frame classifier 230A-230C infers a classification for each frame embedding in one or more frame embeddings 104 to generate one or more subsets 104A and 104B of frame embeddings with a specific classification. Therefore, after processing the frame embeddings 104 by the FWD-BWD classifier / filter 230A, a subset 104A of frame embeddings representing frames acquired during the posterior path of the endoscope through the patient's colon is retained. Subset 104A has a dimension D2×m2, where m2 is the total number of frame embeddings selected from the complete set of m1 frames. Therefore, in this example, m2<=m1.

[0034] Next, the frame embedding subset 104A is processed by a left-right classifier / filter 230B to generate a frame embedding subset 104B representing the frames acquired during the posterior path of the endoscope through the left side of the patient's colon. Subset 104B has a dimension D2×m3, where m3 is the total number of frame embeddings classified as belonging to the left colon, selected from a subset of m2 posterior path frames. Therefore, in this example, m3 <= m2 <= m1.

[0035] Next, the frame embedding subset 104B is processed by segment classifier 230C such that the frame embedding subset 104B, representing frames acquired during the posterior path of the endoscope through the left side of the patient's colon, is classified as belonging to one of the following three left colonic segments: rectum, sigmoid colon, and descending colon. As described above, subset 104B has a dimension of D²×m³.

[0036] Video analyzer 140 analyzes the frame embeddings 104B and frame category inferences 105 of frames represented by a subset of frame embeddings 104B to generate one or more video-level scores or category inferences for the left colon segment. Additionally, frame-level attention and model information, such as attention vectors and / or attention-weighted representations, can be generated. All such information can be analyzed and presented to users (e.g., physicians, healthcare professionals, and / or patients) in video review interface 150 for further analysis. For example, for endoscopic medical video 101 of a patient's colon, the video review interface can display frame 12, showing an image of frame 12 along with relevant information such as segment classification, subsegment location, endoscope location, endoscope orientation, biopsy location (if any), and relevant disease severity scoring information (e.g., using any number of scoring systems or other methods known in the art, such as the Mayo Endoscopic Severity Index (MES) or the Ulcerative Colitis Endoscopic Severity Index (UCEIS)). Continuous scores with non-integer values ​​can also be displayed on the video review interface 150, along with the ability to view attention maps and high-attention areas, and the ability to scroll directly to frames displaying features of interest (including MES data, bleeding data, erosion data, and vascular pattern data) assessing disease severity. For Crohn's disease patients, additional data on ulcer size, surface and / or narrowing, and percentage of ulcer area can also be displayed.

[0037] Figure 3 An example of a medical video processing system 3000 is illustrated, which uses one or more trained attention-based networks to compute classification inferences for frames of medical videos to identify, for example, frames on the withdrawal path of an endoscopic video of a patient's colon for further analysis.

[0038] and Figure 1 and Figure 2Similarly, this example illustrates a system 3000 for analyzing medical video 101 (such as endoscopic video). System 3000 processes multiple frames 12 that together constitute medical video 101. Preprocessor 110 preprocesses the video data corresponding to medical video 101 by generating frames, resizing frames, and masking frames to remove annotations, to generate resized and masked frames 103. Block 110 processes the video data such that each frame has a uniform dimension D1. In this example, if there are a total of m1 frames in the video, the size of the resized and masked frame 103 after preprocessing is D1 × m1. In a specific implementation of this example, after preprocessing, each frame 12 in the sequence corresponding to medical video 101 has a dimension D1.

[0039] Such as about Figure 1 and Figure 2 The self-supervised learning (SSL) pre-trained encoder 120, as discussed, processes frame 103 to obtain a sequence of frame embeddings 104. For each frame 103 processed by the SSL pre-trained encoder 120, there exists a frame embedding 104 of dimension D2, and the order of the frame embeddings 104 is the same as the order of frames 12 in the sequence of medical video 101 and in the resized and masked frames 103. Again, the value of D2 depends on the underlying SSL model used for SSL pre-training of the encoder 120, as described below regarding... Figure 4 The subject of discussion.

[0040] Next, two frame-by-frame classifiers, 230A and 230C, are applied. Figure 1 Similar to the frame-by-frame classifier 130, each frame-by-frame classifier 230A and 230C infers a classification for each frame embedding in one or more frame embeddings 104 to generate subsets 104A and 104B of frame embeddings with specific classifications. Therefore, after processing the frame embeddings 104 by the FWD-BWD classifier / filter 230A, a subset 104A of frame embeddings representing frames acquired during the posterior path of the endoscope through the patient's colon is generated. Subset 104A has a dimension D2×m2, where m2 is the total number of frame embeddings selected from the complete set with a total of m1 frames. Therefore, in this example, m2<=m1.

[0041] Next, in one example, the frame embedding subset 104A is processed by segment classifier 230C such that the frame embedding subset 104A representing frames acquired during the backward path of the endoscope belongs to one of the following representative colonic segments: ileum, ascending colon, transverse colon, descending colon, sigmoid colon, and rectum. However, more or fewer colonic segments and / or subsegments can be represented based on which colonic segments were recorded on the video during the endoscope's path through the patient's colon. As mentioned above, subset 104A has a dimension D2×m2, where m2 is the total number of frame embeddings selected from the backward path frames having a total of m2 frames. Therefore, in this example, m2<=m1.

[0042] Video analyzer 140 analyzes the frame embeddings 104B and frame category inferences 105 of frames represented by a subset of frame embeddings 104B to generate one or more video-level scores or category inferences for the left colon segment. Additionally, frame-level attention and model information, such as attention vectors and feature vectors, can be generated. All such information can be analyzed and presented to users (e.g., physicians, healthcare professionals, and / or patients) in video review interface 150 for further analysis. Example video review interface 1300 is capable of displaying some or all of the information discussed above and / or as follows regarding… Figure 13 Other relevant information discussed.

[0043] Figure 4 An example of a medical video processing system is illustrated, comprising an attention-based deep learning network (e.g., frame classifiers 230A, 230B, and / or 230C) trained to automatically compute inferences for frames in a medical video. In some examples, medical video 401 may be one of many training videos in a training set to train a downstream attention-based deep learning network comprising an attention-based encoder 440 and an inference network 450 comprising MLP 450-1 (Softmax layer 450-2 is part of the network but not trained) on an inference task. In one example, the training set comprises over 21 million labeled frames from 1,335 endoscopic videos corresponding to 753 patients for training blocks 440 and 450, and over 61 million frames from 4,160 endoscopic videos from four clinical trials and for training block 120.

[0044] Similar to Figures 1 to 3Preprocessing block 110 preprocesses the video data corresponding to video 401 by generating frames, resizing the frames, and removing annotations. Block 110 processes the video data so that each frame has a uniform dimension D1. In this example, after preprocessing, each frame 43 in the sequence 403 corresponding to video 401 has a dimension D1 of 224×224×3 (if RGB color channels are used). Other dimensions may be used in different implementations. Preprocessing block 110 prepares the sequence 403 of frames 43 corresponding to medical video 401. Therefore, the complete sequence of frames 43 corresponding to medical video 401 has a dimension of N×D1, where N is the total number of frames obtained from the video for processing, and D1 is the size of each frame.

[0045] A self-supervised learning (SSL) pre-trained encoder 120 processes frame 43 to obtain frame embedding 404. For each frame 43 processed by the SSL pre-trained encoder 120, there exists a frame embedding 44, and the order of the frame embeddings 44 is the same as the order of the frames 43 in the frame sequence 403. Each frame embedding 44 has a dimension D2.

[0046] The value of D2 depends on the underlying SSL model used for SSL pre-training of encoder 120. In one example, the DINOv2 objective is used to pre-train the visual transformer encoder (ViT) (see the “DINOv2” method described by Oquab et al. in “DINOv2: Learning Robust Visual Features without Supervision” (2023). However, other SSL pre-trained encoders and pre-training techniques may be used without departing from the spirit and scope of this disclosure. In another example, the DINOv1 objective can be used to pre-train an SSL pre-trained encoder, such as encoder 120 (see the “DINOv1” method described by Caron et al. in “Emerging Properties in Self-Supervised VisionTransformers” (2021). In yet another example, a convolutional neural network (CNN) encoder, such as ResNet (a CNN with residual connections), is used with the “SimCLR” method (see Chen et al. in “ASimple Framework for Contrastive Learning of Visual Representations” (2020)). All three papers cited in this paragraph are incorporated into this paper by way of citation.

[0047] In one example, the SSL pre-trained encoder 120 has been pre-trained using publicly available clinical trial endoscopy videos. In one example, over 61 million frames from more than 4,000 endoscopy videos were used for SSL pre-training. In one example, the SSL pre-trained encoder is a ViT-B / 16 encoder pre-trained using the DINOv2 method. In one example, a batch size of 256 was used for pre-training, with a cosine decay learning rate of [missing information] in the case of the Adam optimizer. In one example, 15 pre-training iterations were performed on four NVIDIA A10G GPUs. See also Mobadersany, Pooya et al., “Harnessing Temporal Information for Precise Frame-Level Predictions in Endoscopy Videos,” cited above and incorporated herein by reference. For further details, see the International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer Nature Switzerland, 2024. This paper is incorporated herein by reference.

[0048] In one such DINOv2 example, the value of D2 is 768 (one-dimensional embedding). Therefore, the complete sequence 404 of the frame embedding 44 corresponding to the medical video 401 has a dimension of N×D2, where N is the total number of frames obtained from the video for processing.

[0049] A video-level temporal enhancer (VLA) 425 processes frame embeddings 44 in a frame embedding sequence 404. The VLA 425 may randomly apply one or more video-level enhancements to the sequence 404 of frame embeddings 44. Alternatively, the VLA 425 may randomly determine whether to apply a first temporal modification operation to the sequence 404 of embeddings 44, and then randomly determine whether to apply a second temporal modification operation to the set of embeddings. An example of such an operation is splitting, i.e., cropping the sequence, where a subset of the frame embeddings in the sequence is retained and the remainder is discarded. Another example of such an operation is reversing the frame embedding sequence. In one example, the first temporal modification is splitting, and the second temporal modification is reversing, as follows: Figure 6 Further details are provided in the background.

[0050] continue Figure 4The description states that the VLA 425 output frame embeds an enhanced sequence 405 of 45. Sequence 405 has a dimension N. A ×D2, where N A This is the number of frame embeddings after the enhancement operation. In this example, if VLA 130 did not apply the splitting operation to sequence 12, then N A =N. Otherwise, N A The dimension is less than N. The frame embeddings are processed by an attention-based deep learning network, which in this example includes an attention-based encoder 440 and an inference network 450. The attention-based encoder 440 processes the sequence 405 of the frame embeddings 45 and outputs a sequence of dimensions N. A A sequence 406 of ×D2 attention vectors 46. Each attention vector 46 corresponds to a frame embedding 45 that has been processed by the attention-based encoder 440.

[0051] In one example, the attention-based encoder 440 is a ViT encoder based on the model described by Dosovitskly et al. in “An Image is Worth 16x16 Words”, 2021, which is incorporated herein by reference. In Vaswani's paper, additional classification tokens are added, and these additional tokens are used to provide a summary output token for classifying the image corresponding to the set of input tokens (in Vaswani's paper, each input token corresponds to a patch of the image to be analyzed). However, frame-by-frame classification may be desired, and in this case, the additional classification tokens used in Vaswani's paper may not be necessary, and the output tokens embedded in each input frame can be used directly by the inference network. An example of encoder 440 uses a ViT encoder with four layers and eight self-attention heads in each layer.

[0052] When performing classification, each attention vector 46 of sequence 406 is processed by an inference network 450 comprising a multilayer perceptron (MLP) 450-1 and a softmax layer 450-2. When performing regression, the softmax layer 450-2 is not required. In one example, the dropout rate of the attention encoder is 0.25, the dropout rate of the classification network is 0.5, the batch size used is 1, and the learning rate is... And the weight decays to The Adam optimizer. Network 450 outputs a computed inference set 47. In classification problems, inferences 47 can be in the form of the probability that a frame can be inferred to belong to each class, and in regression problems, they can be the regression values ​​corresponding to each frame. The dimension of inference set 47 is N. A×C, where C is the number of categories (or, in the case of regression inference, the number of regression outputs), depending on the application.

[0053] In such Figure 3 In one example shown, the frame classifier 230C is trained to infer colon segment class probabilities corresponding to five intestinal segments: rectum (RM), left colon / sigmoid colon (LC) (where the descending colon and sigmoid colon are bundled together as one segment), transverse colon (TC), right colon (RC), and ileum (IL). In this type of application, "C" equals 5, meaning that for each frame, the classification network 450 outputs five class probabilities. Figure 3 In another example shown, frame classifier 230A is trained to infer whether a frame was obtained from the forward or retreat path of the procedure. In this type of application, "C" equals 2, meaning that for each frame, the inference network 450 outputs two class probabilities. Figure 2 In another example shown, frame classifier 230B is trained to infer whether a frame was acquired from the left or right side of the colon. In this type of application, C also equals 2. Figure 2 In another example shown, frame classifier 230C is trained to infer which of the three colonic segments (RM, SC, or DC) the frame comes from for each frame from the left side of the colon: rectum (RM), sigmoid colon (SC), and descending colon (DC). In this type of 3-class example, C is 3.

[0054] In another example, multiple models are combined in an end-to-end manner as follows: First, a first trained model selects frames corresponding to the withdrawal path. Then, another model selects frames that correspond to the left side of the colon. Then, one or more inference models analyze the selected frames (which correspond to the withdrawal path and the left side of the colon) to perform one or more inference tasks, such as frame-by-frame classification and / or regression, thereby outputting, for example, a disease severity score or other possible analytical results relevant to the colon being studied.

[0055] The learning module 460 performs typical learning processing by comparing the category inference 47 with the training data labels to calculate the loss (error) according to the selected loss function and then backpropagating the loss to tune the learnable parameters in the MLP 450-1 and the attention-based network 440.

[0056] Although a specific type of classification inference has been described above, the calculated "inference" can be any one or more of a prediction, estimate, rating, suggestion, classification, evaluation, calculation, or other inference. In some implementations, instead of using a softmax layer to provide class probabilities, one or more outputs of an MLP can be used to provide a regression output in the form of a predicted rating within a range of values.

[0057] VLA 425 is used during training. However, VLA 425 is not used when the trained network is used for post-training inference tasks (and if present, it can be bypassed).

[0058] Figure 5 An embodiment of the present disclosure is illustrated for processing medical videos to train an attention-based deep learning network (such as including...) Figure 4 The method 5000 for automatically calculating the classification inference of frames in medical videos (shown as an attention-based deep learning network with an attention-based encoder 440 and an inference network 450) is as follows.

[0059] Step 501 preprocesses the medical videos in the training set to generate frames, resizes them to a uniform size, and masks any annotations. Step 502 encodes each frame of the medical videos using an SSL pre-trained encoder to obtain a frame embedding sequence corresponding to the medical video. Step 503 selectively applies one or more temporal enhancements to the frame embedding sequence. As shown, one or more temporal enhancements include splitting (i.e., cropping the video into selected subsequences of the frame embedding sequence) and inversion. Step 503 outputs the enhanced sequence of frame embeddings (for some sequences, this may be the same as the sequence before the application of one or more operations, depending on whether one or more operations are applied randomly).

[0060] Step 504 processes the augmented sequence of frame embeddings using an attention-based network to compute an attention-based vector for each frame embedding of the augmented sequence. Step 505 processes the attention-based vectors to compute classification or regression inference on a frame-by-frame basis. As previously discussed, the number of outputs depends on the application.

[0061] Step 506 uses the frame labels and the computed inferences (e.g., class probabilities) to calculate a loss value, which quantifies the error in the inference using a chosen loss function. In one example, for a classification application, the cross-entropy loss function is used. However, depending on the application, other loss functions may be used. Step 507 determines whether the error is now minimized. If yes, training method 5000 ends. If not, step 508 adjusts the learnable parameters of the classification and attention-based deep learning networks to further reduce the error, and the method returns to step 502.

[0062] Figure 6 Method 6000 is illustrated, which can be derived from VLA 425 ( Figure 4 The process involves selectively performing one or more time-enhancing operations. As shown in the figure, step 601 determines whether to select the frame embedding sequence F via random selection. iPerform a splitting (time trimming or time pruning) operation. Step 602 determines whether sequence F has been selected. i If the result of step 602 is negative, then step 603 will enhance sequence F. aug Set to equal the initial sequence F i If the result of step 602 is yes, then step 604 initiates the splitting operation by: for example, randomly selecting a starting frame number R between 0 and N / 2-L / 2. s (integers), where N is the sequence F i The number of frames in the dataset, where L is greater than 0 and less than or equal to N. i An integer, and randomly selects an end frame number between N / 2 + L / 2 and N. Then step 605 will F aug Set to equal from frame r s to frame r e The frame sequence. Therefore, after the splitting operation, F aug It is F i A subset of (which may be F) i , or F i (A smaller subset of L). In one example, L is set to 2, but other numbers can also be used.

[0063] Step 606: Randomly determine whether to select F. aug Perform sequence reversal. Step 607 determines whether F has been selected in step 606. aug Perform the reversal. If the result of step 607 is negative, then frame sequence F... aug No reversal. If the result of step 607 is yes, then step 608 will F aug The sequence is reversed. Then, step 609 outputs F. aug Those skilled in the art will understand that when the sequence is reversed, the label data may also need to be modulated to facilitate matching the labels with downstream inferences for supervised (or weakly supervised) learning purposes. Such modulation of the label data may also be necessary given other temporal augmentations (such as temporal pruning (splitting)).

[0064] In one example Figure 6 Method 6000 can be more formally represented as the following algorithm.

[0065] Algorithm 1 Those skilled in the art will understand that Figure 6The example illustrates an instance of introducing temporal augmentation into a training video sequence. In the example above, a given video sequence processed by method 6000 can be split (cropped), reversed, split and reversed, or left unchanged based on randomization operations within the method. Alternatively, these and / or other operations can be performed based on random or predefined selection techniques. In a general example of the temporal cropping operation, the function "min" is defined such that at least min(A, N) frames are available in the video after temporal cropping, where A is defined based on the desired severity of augmentation, and N is the number of available frames in the video. In one example, a lower value of A indicates more severe augmentation, i.e., a more aggressive temporal cropping.

[0066] Figure 7 An embodiment of the present disclosure is illustrated for processing medical videos to train an attention-based deep learning network (such as including...) Figure 4 The method 7000 for automatically calculating the classification inference of frames in medical videos (shown as an attention-based deep learning network with an attention-based encoder 440 and an inference network 450) is as follows.

[0067] Step 701 preprocesses the medical video to generate frames and resizes them to a uniform size. Step 702 encodes each frame of the medical video using an SSL pre-trained encoder to obtain a frame embedding sequence corresponding to the medical video.

[0068] Step 703 processes the frame embedding sequence corresponding to the medical video using an attention-based network to compute an attention-based vector for each frame embedding in the sequence. Step 704 processes the attention-based vectors to compute classification inference on a frame-by-frame basis. As previously discussed, the number of categories depends on the application.

[0069] Figure 8 An example of a medical video processing system is illustrated, configured to train an attention-based deep learning network (video analyzer 140) to automatically analyze and compute one or more inferences from medical videos. Frame embeddings 804, each with dimension D2×n, from the training set are processed by the attention-based deep learning network, which in this example includes an attention-based encoder 810 and an attention-weighted network 820. The attention-based encoder 810 processes the sequence 804 of frame embeddings and outputs a sequence 805 of attention vectors with dimension D2×n. Each attention vector corresponds to a frame embedding that has been processed by the attention-based encoder 810. At 806, an attention score for each frame is output and, together with the attention vector for each frame 805, is input into an aggregator 830 to generate an aggregated vector 807 of the patient video with dimension D2×1.

[0070] In one example, the attention-based encoder 810 is a ViT encoder based on the model described by Dosovitskly et al. in “An Image is Worth 16x16 Words”, 2021, which is incorporated herein by reference. In Vaswani's paper, additional classification tokens are added, and these additional tokens are used to provide a summary output token for classifying the image corresponding to the set of input tokens (in Vaswani's paper, each input token corresponds to a patch of the image to be analyzed). However, in cases where frame-by-frame classification is desired, the additional classification tokens used in Vaswani's paper may not be necessary, and the output tokens embedded in each input frame can be used directly by the inference network. An example of encoder 440 uses a ViT encoder with four layers and eight self-attention heads in each layer.

[0071] At position 808, the aggregated vector 807 of the patient video is processed by a multilayer perceptron (MLP) 840 to generate a video-level score. A softmax layer is not required when performing regression. In one example, the dropout rate of the attention encoder is 0.25, the dropout rate of the classification network is 0.5, the batch size used is 1, and the learning rate is... And the weight decays to Adam optimizer.

[0072] The learning module 460 performs typical learning processing by comparing the category inference 47 with the training data labels to calculate the loss (error) according to the selected loss function and then backpropagating the loss to adjust the learnable parameters in the MLP 840 and in the attention weighting 820 and attention encoder 810.

[0073] Figure 9A medical video processing system (e.g., video analyzer 140) according to an embodiment of the present disclosure is illustrated for computing classification inferences of frames of a medical video using a trained attention-based network. Frame embeddings 104A comprising video frames from a medical video showing an endoscopic video in a backward path through the colon are processed by a trained attention encoder 810. Each frame embedding 104A has a dimension D2×n. The attention-based encoder 810 processes the sequence 104A of frame embeddings and outputs a sequence 905 of attention vectors having a dimension D2×n. Note that in one example, this assumes that the attention encoder 810 has the same output size as the SSL pre-trained encoder 120 referenced in the earlier figures. In one example, this applies. However, in alternative examples, depending on the encoder used at each processing stage, the output dimension of the downstream attention encoder (such as encoder 810) may differ from the output dimension of the earlier SSL pre-trained encoder (such as encoder 120).

[0074] Each attention vector corresponds to a frame embedding that has already been processed by the attention-based encoder 810. An attention weighting block 820 assigns an attention score 806 to each attention vector 905. The score 906 and attention vector 905 are processed by an aggregator 830 to generate an aggregated vector 907 of the patient video with dimension D2×1. At 908, the aggregated vector 907 of the patient video is processed by a multilayer perceptron (MLP) 840 to generate a video-level score.

[0075] Figure 10 An example is illustrated of a medical video processing system (e.g., video analyzer 140) according to an embodiment of the present disclosure for computing classification inferences of frames of a medical video using a trained attention-based network, such as... Figure 2 As shown. First, based on the results of a segment classifier (e.g., from a frame-by-frame classifier 130), frame embeddings 104B, including video frames from a medical video showing an endoscopic video along the posterior path through the colon, are grouped for different segments of the colon. Figure 10In one illustrated example, within corresponding groups 104B-1, 104B-2, and 104B-3, which correspond to each colonic segment including (e.g., three) segments on the left side of the colon, frames classified only as originating from the backward path and segments from the left side of the colon are processed by a trained attention encoder 810, attention-weighted network 820, aggregator 830, and MLP 840. Each group 104B-1, 104B-2, and 104B-3 is processed separately by the attention encoder 810, attention-weighted network 820, aggregator 830, and MLP 840. Each frame embedding 104B-1, 104B-2, and 104B-3 has a dimension of D2×n. The attention-based encoder 810 processes the sequence of frame embeddings and outputs an attention vector 905 with a dimension of D2×n. Each attention vector corresponds to a frame embedding that has already been processed by the attention-based encoder 810. Attention weighting block 820 assigns an attention score 906 to each attention vector 905. Aggregator 830 processes the attention vector 905 along with the attention scores 906 for each frame 905 to generate an aggregated vector 907 of segments with dimension D2×1. The aggregated vector 907 of each segment is processed by MLP 840 to generate segment 1 scores, segment 2 scores, and segment 3 scores, which can be averaged (or weighted averaged) to generate a patient-level score. In this example, in addition to the patient-level score, the attention score 906 for each frame and the segment-level score are also input into the video review interface 150 for display, as discussed further below.

[0076] Figure 11 An example of a medical video processing system (e.g., video analyzer 140) according to an embodiment of the present disclosure is illustrated, which uses a trained attention-based network to compute classification inferences for frames of a medical video. Figure 3 As shown. First, based on the segment classifier results (e.g., from frame-by-frame classifier 130), frame embeddings 104A, including video frames from a medical video showing an endoscopic video along the posterior path through the colon, are grouped for different segments of the colon. Figure 11 In one illustrated example, in corresponding groups 104A-1, 104A-2, 104A-3, 104A-4, and 104A-5, which correspond to each colonic segment in a group comprising (e.g., five or six) colonic segments of a complete colon, frames classified as either from the backward path or from the left or right side of the colon are processed by a trained attention encoder 810, attention-weighted network 820, aggregator 830, and MLP 840. The colon is generally understood to have six anatomical segments: rectum, sigmoid colon, descending colon, transverse colon, ascending colon, and ileum. For ease of illustration, and because in some specific implementations certain segments may be bundled together for classification purposes, Figure 5 Five sets of frame embeddings (with corresponding processing paths) are shown instead of six. The illustrated example assumes that the sigmoid colon and descending colon segments are bundled together for classification purposes. Each set 104A-1, 104A-2, 104A-3, 104A-4, and 104A-5 are processed separately by an attention encoder 810, an attention weighting network 820, an aggregator 830, and an MLP 840. Each frame embedding 104A-1, 104A-2, 104A-3, 104A-4, and 104A-5 has a dimension D2×n. The attention-based encoder 810 processes the sequence of frame embeddings and outputs a sequence 905 of attention vectors with a dimension D2×n. Each attention vector corresponds to a frame embedding that has been processed by the attention-based encoder 810. The attention weighting block 820 assigns an attention score 906 to each attention vector 905. Aggregator 830 processes attention vector 905 along with attention scores 906 for each frame 905 to generate aggregated vectors 907 for segments with dimension D2×1. The aggregated vector 907 for each segment is processed by MLP 840 to generate segment 1, segment 2, segment 3, segment 4, and segment 5 scores, which can be averaged to generate a patient-level score. In this example, in addition to the patient-level score, the attention score 906 for each frame and the segment-level score are also input into the video review interface 150 for display, as discussed further below.

[0077] Those skilled in the art will understand that in regression inference applications, the output of an MLP (such as the MLP 840) can be derived from a single artificial neuron to provide a single value that can be within a range of consecutive values ​​using a desired number of significant digits. Furthermore, those skilled in the art will further understand that simple scaling operations can convert this number to a desired range (such as 0 to 3, as in the case of MES scoring) or an appropriate value within another range depending on the application. Therefore, such scaling operations will not be further described or detailed herein. Various ranges can be used for endoscopic applications or other disease assessment applications.

[0078] Figure 12 A method 1200 for processing medical video to compute inferences about frames of the medical video using a trained attention-based deep learning network, according to an embodiment of the present disclosure, is illustrated. At step 1201, each frame of the video is encoded using an SSL pre-trained encoder to obtain a sequence of frame embeddings corresponding to the medical video. At step 1202, a frame-by-frame classifier is used to determine which frame embeddings originate from the endoscope withdrawal path. Next, at 1202, it is determined whether the entire colon is being examined in the medical video and scored accordingly. If the entire colon is being examined, all frame embeddings are preserved for further analysis, and processing proceeds directly to determining at 1205 whether a score is assigned to each segment.

[0079] If not, the frame embeddings are submitted to the frame-level left / right colon classification network at step 1203, and the frames of interest are retained at step 1204. Then, at step 1205, it is determined whether to assign a score to each segment. If the result of step 1205 is no, the method proceeds to step 1209, and an attention-based deep learning network is used to process all backward path (retreat path) frame embeddings together to obtain a video-level score.

[0080] If a segment-based rating is desired (if the result of step 1205 is yes), then step 1206 submits the frame embeddings to the segment classifier. Based on the segment classification, at step 1207, an attention-based deep learning network is used to process the backward path frame embeddings for each segment to obtain a rating for each segment. Then, at step 1208, a video-level rating is obtained by averaging (or weighted averaging) the ratings for each segment.

[0081] Figure 13 An example of a video review interface 1300 for display on one or more devices according to an embodiment of the present disclosure is shown. For example, for an endoscopic medical video 101 of a patient's colon, the video review interface may display video frame 12 at 1301, along with video timestamp data 1303 and frame data 1304, along with related information 1302. In the example shown, information 1302 includes automated (AI) MES scoring, standard MES scoring, segment classification, and bleeding, erosion, and vascularity scores. In alternative embodiments, various information may be provided, including segment location, endoscope location, endoscope orientation, biopsy location (if any), and relevant disease severity scoring information (e.g., using any number of scoring systems or other methods known in the art, such as the Mayo Endoscopic Severity Index (MES) or the Ulcerative Colitis Endoscopic Severity Index (UCEIS)). Continuous scores with non-integer values ​​(one or more significant digits after the decimal point) can also be displayed on the video review interface 150 and can be shown as number 1302 via heatmaps 1305 and 1306, along with the ability to view attention maps and high-attention areas at 1307. This provides the user with the ability to directly scroll to the frame that has the highest impact on the AI-MES score (e.g., the MES score calculated using the implementation described in this disclosure). The video review interface 1300 can also display other features of interest used for assessing the severity of the disease shown in the displayed frames, including MES data, bleeding data, erosion data, and vascular pattern data, as shown in the exemplary information bar 1302. Other relevant data may also be included. For example, for patients with Crohn's disease, additional data regarding ulcer size, surface and / or narrowing, and percentage of ulcer area may also be displayed.

[0082] Examples of training data and selected results Extensive clinical trials (including UNIFI: NCT02407236, JAK-UC: NCT01959282, SEAVUE: NCT03464136, and TRIDENT: NCT02877134) were used to train and validate models consistent with the implementation presented herein. These trials considered over one thousand hours of endoscopic video (sixty million frames). Models consistent with the implementation presented herein demonstrated good performance. For example, in some demonstrations, AUC=0.8 for MES, AUC=0.79 for UCEIS, and AUC=0.93 for segmented endoscopic positions. Furthermore, in one study, embodiments consistent with this disclosure were trained with clinical trial data on ulcerative colitis (UNIFI: NCT02407236, Phase 3, 965 participants, 3,128 videos), and continuous MES scores were automatically generated by these embodiments, demonstrating greater discriminative power than human MES scores in detecting differences between placebo and treatment with 200 mg guslkumab. Further details of the relevant results are disclosed in U.S. Provisional Application No. 63 / 555,883, filed February 20, 2024, which is incorporated herein by reference.

[0083] Figure 14 An example of computer system 1400 is shown, and one or more computer systems in this system may be used to implement one or more of the devices, systems, and methods illustrated herein. Computer system 1400 executes instruction code contained in computer program product 1460. Computer program product 1460 includes executable code in an electronically readable medium that may instruct one or more computers (such as computer system 1400) to perform processing to implement the exemplary method steps performed.

[0084] An electronically readable medium can be any temporary or non-temporary medium that stores information electronically and can be accessed locally or remotely, for example, via a network connection. The medium may include multiple geographically dispersed media, each configured to store different portions of executable code at different locations and / or at different times. The executable instruction code in the electronically readable medium instructs the illustrated computer system 1400 to perform the various exemplary tasks described herein. The executable code used to instruct the performance of the tasks described herein will typically be implemented in software. However, those skilled in the art will understand that a computer or other electronic device may utilize code implemented in hardware to perform many or all of the identified tasks. Those skilled in the art will understand that many variations of the executable code implementing the exemplary methods can be found within the spirit and scope of this disclosure.

[0085] Code or copies of code contained in computer program product 1460 may reside in one or more persistent storage media (not shown separately) communicatively linked to system 1400 for loading and storage into persistent storage device 1470 and / or memory 1410 for execution by processor 1420. Computer system 1400 also includes I / O subsystem 1430 and peripheral devices 1440. I / O subsystem 1430, peripheral devices 1440, processor 1420, memory 1410, and persistent storage device 1470 are linked via bus 1450. Similar to persistent storage device 1470 and any other persistent storage device that may contain computer program product 1460, memory 1410 is a non-transitory medium (even if implemented as a typical volatile computer memory device). Furthermore, those skilled in the art will understand that, in addition to storing computer program product 1460 for performing the processes described herein, memory 1410 and / or persistent storage device 1470 may also be configured to store various data elements referenced and illustrated herein.

[0086] Those skilled in the art will understand that computer system 1400 is merely one example of a system that can implement a computer program product according to this disclosure. By way of example only, execution of instructions contained in the computer program product can be distributed across multiple computers, such as, for example, across computers in a distributed computing network.

[0087] Instructions for implementing artificial neural networks or other deep learning networks may reside in computer program product 1460. When processor 1420 is executing the instructions of computer program product 1460, the instructions or portions thereof are typically loaded into working memory 1410 from which processor 1420 can easily access the instructions.

[0088] Processor 1420 may include multiple processors, which may include corresponding additional working memory (the additional processors and memory are not separately illustrated), the additional working memory including one or more graphics processing units (GPUs), the one or more GPUs including at least several thousand arithmetic logic units supporting massively parallel computing. GPUs are frequently used in deep learning applications because they can perform related processing tasks more efficiently than typical general-purpose processors (CPUs). Processor 1420 may additionally or alternatively include one or more dedicated processing units, the one or more dedicated processing units including systolic arrays and / or other hardware arrangements supporting efficient parallel processing. Such dedicated hardware may work in conjunction with CPUs and / or GPUs to perform the various processing described herein. Such dedicated hardware may include application-specific integrated circuits (ASICs) and the like (which may refer to a portion of an ASIC), field-programmable gate arrays and the like, or combinations thereof. However, a processor (such as processor 1420) may be implemented as one or more general-purpose processors (preferably having multiple cores) without necessarily departing from the spirit and scope of this disclosure.

[0089] While the term “inference” may be used in various ways herein, those skilled in the art will understand that the systems and methods described herein are not limited thereto, and that the term “inference” herein may refer to the execution of any computation in various calculations and / or the generation of various outputs, which may include, but are not limited to, any one or more of inference, rating, estimation, prediction, estimation, suggestion, recommendation, classification, categorization, annotation, conclusion, etc., or any combination of the foregoing.

[0090] Additional Implementation Plan This disclosure further contemplates one or more additional embodiments, some of which are set forth below as example embodiments.

[0091] Implementation Scheme 1. A computer system comprising: one or more processors; a storage medium storing (a) instructions and (b) a dataset representing video data, wherein the video data includes endoscopic video data; an encoding module using the one or more processors to encode a set of frames from the dataset into high-dimensional vector data, wherein (a) the set of frames represents data including a region of interest (ROI), and (b) the high-dimensional vector data contains information representing each of the ROI's multiple views; and one or more attention modules comprising at least one of the following: i. A first attention module, wherein the first attention module uses the one or more processors to generate a video-level attention map from the high-dimensional vector data; ii. A second attention module, which uses the one or more processors to generate a frame-level attention map from the high-dimensional vector data; iii. A third attention module, which uses the one or more processors to generate a pixel-level attention map from the high-dimensional vector data; and iv. A fourth attention module, which uses the one or more processors to generate a time-series level attention map from the high-dimensional vector data; The prediction module uses one or more processors to output an evaluation based on one or more of the following: the video-level attention map, the frame-level attention map, the pixel-level attention map, and the time-series-level attention map.

[0092] Implementation Scheme 2. The computer system according to Implementation Scheme 1, wherein the assessment includes at least one of the following: disease severity estimation, patient health assessment, disease progression measurement, and medical diagnosis determination.

[0093] Implementation Scheme 3. The computer system according to Implementation Scheme 2, wherein the assessment relates to one or more of the following: inflammatory bowel disease, Crohn's disease, and ulcerative colitis.

[0094] Implementation Scheme 4. The computer system according to Implementation Scheme 2, wherein the assessment includes a continuous disease severity score.

[0095] Implementation Scheme 5. A method for analyzing one or more medical videos corresponding to a patient examination using one or more computers, the method comprising: a. Preprocess the video data to obtain multiple video frames corresponding to the one or more medical videos, which correspond to the patient examination; b. Encode the plurality of video frames using a pre-trained encoder to obtain a plurality of frame embeddings corresponding to the corresponding frames in the plurality of video frames, wherein the pre-trained video encoder has been pre-trained using self-supervised learning; c. Concatenate the multiple frames; d. Process the plurality of frame embeddings in an attention-based deep learning network to obtain a summary vector generated by the attention processing, the attention processing focusing on each of the plurality of frame embeddings relative to the other frame embeddings in the plurality of frame embeddings; and e. Submit the aggregated vector to a classification network to obtain one or more outputs about the patient corresponding to the medical video from the patient's examination.

[0096] Implementation Scheme 6. The method according to Implementation Scheme 5, wherein the one or more medical videos comprise a plurality of videos, such that the plurality of frame embeddings cascaded for further processing comprise frame embeddings from each of the plurality of videos.

[0097] Implementation Scheme 7. The method according to any one of Implementation Schemes 5 to 6, wherein the patient examination is transthoracic echocardiography.

[0098] Implementation Scheme 8. The method according to any one of Implementation Schemes 5 to 7, wherein the pre-trained encoder includes a visual transformer (ViT).

[0099] Implementation Scheme 9. The method according to any one of Implementation Schemes 5 to 7, wherein the pre-trained encoder comprises a convolutional neural network.

[0100] Implementation Scheme 10. The method according to any one of Implementation Schemes 5 to 9, wherein the pre-trained encoder is trained using SimCLR.

[0101] Implementation Scheme 11. The method according to any one of Implementation Schemes 5 to 9, wherein the pre-trained encoder is trained using DINO.

[0102] Implementation Scheme 12. The method according to any one of Implementation Schemes 5 to 9, wherein the pre-trained encoder is trained using DINOv2.

[0103] Implementation Scheme 13. The method according to any one of Implementation Schemes 5 to 12, wherein the attention-based deep learning network is a transformer multi-head attention network.

[0104] Implementation Scheme 14. The method according to Implementation Scheme 13, wherein the multi-head attention network of the transformer is a set transformer.

[0105] Implementation Scheme 15. The method according to any one of Implementation Schemes 5 to 12, wherein the attention-based deep learning network is a multi-instance attention network.

[0106] Implementation Scheme 16. The method according to any one of Implementation Schemes 5 to 12, wherein the attention-based deep learning network comprises a convolutional neural network, an attention block, and an aggregator.

[0107] Implementation Scheme 17. The method according to Implementation Scheme 15, wherein the attention block associates a corresponding learnable parameter with the corresponding frame embedding output of the convolutional neural network.

[0108] According to the method of implementation scheme 17, the aggregator applies the corresponding learnable parameters to the corresponding frame embedding output and sums the resulting attention score frame embedding outputs to obtain the summary vector.

[0109] The method according to any one of embodiments 5 to 6 and 8 to 18, wherein the patient examination is an endoscopic examination.

[0110] While this disclosure has been described in detail with respect to the illustrated embodiments, it should be understood that various changes, modifications, and adaptations may be made based on this disclosure, and such changes, modifications, and adaptations are intended to be within the scope of this disclosure. Although this disclosure has been described in conjunction with embodiments currently considered to be most practical and preferred, it should be understood that this disclosure is not limited to the disclosed embodiments, but rather, is intended to cover various modifications and equivalent arrangements that fall within the scope of the basic principles of the invention as described in the various embodiments referenced above and below.

Claims

1. A computerized deep learning system for processing medical videos to assess the severity of a patient's disease, the computerized deep learning system comprising: A preprocessor configured to preprocess video data to obtain a plurality of video frames corresponding to one or more medical videos corresponding to the patient; An encoder configured to encode a plurality of video frames to obtain corresponding frame embeddings corresponding to corresponding frames among the plurality of video frames, wherein the encoder has been pre-trained using self-supervised learning; One or more video frame classifiers, each video frame classifier including an attention-based deep learning network configured to compute a corresponding attention vector by processing the corresponding frame embedding, and to compute a corresponding frame-level inference corresponding to the corresponding frame on a frame-by-frame basis using the corresponding attention vector; and A video analyzer comprising an attention-based deep learning network configured to compute at least one disease severity assessment of the patient using at least some of the corresponding frame-level inferences and the corresponding frame embeddings.

2. The computerized deep learning system according to claim 1, wherein, The one or more video frame classifiers include a first frame classifier configured to compute a first corresponding frame-level inference by processing the corresponding frame embedding, and to use the first corresponding frame-level inference to identify a first set of frames from the medical video for further analysis.

3. The computerized deep learning system according to claim 2, wherein, The one or more video frame classifiers further include: A second frame classifier is configured to compute a second corresponding frame-level inference by processing the corresponding frame embeddings corresponding to the first set of frames, and to use the second corresponding frame-level inference to identify a second set of frames from the medical video for further analysis.

4. The computerized deep learning system according to any one of claims 2 to 3, wherein: The one or more medical videos include endoscopic videos; and The first corresponding frame-level inference is about whether the frame comes from the forward path or the retraction path of the endoscopic video, and the first set of frames is inferred to come from the retraction path of the endoscopic video.

5. The computerized deep learning system according to any one of claims 3 to 4, wherein: The one or more medical videos include endoscopic videos; and The second corresponding frame-level inference is about whether the frame comes from the left colon region or the right colon region, and the second set of frames is inferred to come from the left colon region.

6. The computerized deep learning system according to any one of claims 3 to 5, wherein, The one or more video frame classifiers further include: A third frame classifier is configured to compute a third corresponding frame-level inference by processing the corresponding frame embeddings corresponding to the second set of frames, and to use the second corresponding frame-level inference to identify a third corresponding frame-level inference that includes a segment classification corresponding to the left colon segment.

7. The computerized deep learning system according to claim 6, wherein, The left colonic segment includes one or more of the following: descending colon, sigmoid colon, and rectum.

8. The computerized deep learning system according to claim 4, wherein: The one or more medical videos include endoscopic videos; and The calculated second inference includes segment classifications corresponding to colonic segments.

9. The computerized deep learning system according to claim 8, wherein, The colonic segment includes one or more of the following: ileum, ascending colon, transverse colon, descending colon, sigmoid colon, and rectum.

10. The computerized deep learning system according to any one of claims 6 to 7, wherein, The at least one disease severity assessment includes the average of the disease severity scores for each of the left colon segments.

11. The computerized deep learning system according to any one of claims 8 to 9, wherein, The at least one disease severity assessment includes the average of the disease severity scores for each of the colonic segments.

12. The computerized deep learning system according to claim 11, wherein, The average of the disease severity scores for each of the left colon segments includes a weighted average.

13. The computerized deep learning system according to claim 12, wherein, The average of the disease severity scores for each of the colonic segments includes a weighted average.

14. The computerized deep learning system according to any one of claims 1 to 13, further comprising a video review interface displayed on a user device, wherein, The video review interface is configured to interactively display one or more video frames and one or more disease severity assessments.

15. The computerized deep learning system according to any one of claims 1 to 14, wherein, The at least one disease severity assessment includes a continuous disease severity score.

16. A method for assessing the severity of a patient's disease using one or more computers, the method comprising: The video data is preprocessed to obtain multiple video frames corresponding to one or more medical videos of the patient; The plurality of video frames are encoded using a pre-trained encoder to obtain corresponding frame embeddings corresponding to corresponding frames in the plurality of video frames, wherein the encoder has been pre-trained using self-supervised learning. Based on multiple corresponding attention vectors, one or more corresponding frame-level inferences corresponding to corresponding frames in the plurality of video frames are computed, wherein the corresponding frame embeddings are processed using a frame-level attention-based deep learning network to generate the multiple corresponding attention vectors and compute corresponding frame-level inferences corresponding to the corresponding frames on a frame-by-frame basis; and Based on at least some of the corresponding frame-level inferences and corresponding frame embeddings, a video-level attention-based deep learning network is used to compute at least one disease severity assessment for the patient.

17. The method according to claim 16, wherein, The one or more corresponding frame-level inferences include a first corresponding frame-level inference, which is calculated by processing the corresponding frame embedding to identify a first set of frames from the medical video for further analysis.

18. The method according to claim 17, wherein, The one or more corresponding frame-level inferences include a second corresponding frame-level inference, which is performed by processing the corresponding frame embeddings corresponding to the first set of frames to identify one or more second sets of frames from the medical video for further analysis.

19. The method according to claim 18, wherein, The one or more corresponding frame-level inferences include a third corresponding frame-level inference, which is performed by processing the corresponding frame embeddings corresponding to the second set of frames to identify one or more third sets of frames.

20. The method according to any one of claims 17 to 19, wherein: The one or more medical videos include endoscopic videos; and The first corresponding frame-level inference is about whether the frame comes from the forward path or the backward path of the endoscopic video, and the first set of frames is inferred to come from the backward path of the endoscopic video.

21. The method according to any one of claims 18 to 20, wherein: The one or more medical videos include endoscopic videos; and The second corresponding frame-level inference is about whether the frame comes from the left or right side of the colon, and the second set of frames is inferred to come from the left side of the colon.

22. The method according to any one of claims 19 to 21, wherein: The one or more medical videos include endoscopic videos; and The third corresponding frame-level inference is about whether the frame comes from one or more of the following: descending colon, sigmoid colon, and rectum.

23. The method of claim 20, wherein, The second corresponding frame-level inference is about whether the frame is inferred to originate from one or more of the following: rectum, sigmoid colon, descending colon, transverse colon, ascending colon, or ileum.

24. The method according to any one of claims 16 to 23, wherein, One or more temporal augmentation operations are selectively applied to the frame embedding sequence corresponding to the training video during the training of the attention-based deep learning encoder.

25. The method according to claim 24, wherein, The one or more time enhancements include time clipping and inversion.

26. The method according to any one of claims 24 to 25, wherein, The one or more temporal enhancements are applied selectively based on a selection function, such that for each sequence in the plurality of frame embedding sequences, one or more of the temporal enhancements are not applied, or are applied.

27. The method according to any one of claims 24 to 26, wherein, The selection function randomly selects whether to apply each of the one or more temporal enhancements to the currently being processed frame embedding sequence.

28. The method according to any one of claims 16 to 27, wherein, The at least one disease severity assessment includes a continuous disease severity score.

29. The method according to any one of claims 16 to 28, further comprising: The one or more video frames and the one or more disease severity assessments are displayed on the user interface of the user device.

30. A non-transitory computer-readable medium comprising a plurality of computer-readable instructions, which, when executed by at least one processor, perform the following operations to assess the severity of a patient's disease: The video data is preprocessed to obtain multiple video frames corresponding to one or more medical videos of the patient; The plurality of video frames are encoded using a pre-trained encoder to obtain corresponding frame embeddings corresponding to corresponding frames among the plurality of video frames, wherein, The encoder has been pre-trained using self-supervised learning; Based on multiple corresponding attention vectors, one or more corresponding frame-level inferences corresponding to corresponding frames in the plurality of video frames are computed, wherein the corresponding frame embeddings are processed using a frame-level attention-based deep learning network to generate the multiple corresponding attention vectors and compute corresponding frame-level inferences corresponding to the corresponding frames on a frame-by-frame basis; and Based on at least some of the corresponding frame-level inferences and corresponding frame embeddings, a video-level attention-based deep learning network is used to compute at least one disease severity assessment for the patient.