Endoscopy video timeline interest level prediction
The system enhances endoscopic video analysis by generating intelligent interest prediction indicators and selecting keyframes, enabling efficient navigation and compression, thus improving the review process for physicians.
Patent Information
- Application Number
- JP2025013020
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2025-01-29
- Publication Date
- 2025-08-12
AI Technical Summary
Endoscopic procedures generate large volumes of video data, making it difficult for physicians to efficiently review and analyze important sections, with existing AI and machine learning techniques lacking in assisting physicians in quickly navigating recorded endoscopic videos.
A system that generates intelligent interest prediction indicators and selects keyframes as thumbnails based on video analysis, allowing physicians to efficiently navigate and identify salient aspects of endoscopic procedures using a user interface with predictive indicators and dynamic compression.
Improves the efficiency, accuracy, and speed at which physicians can review and identify important segments in endoscopic procedures by providing intelligent navigation and compression of video recordings.
Smart Images

Figure 2025117565000001_ABST
Abstract
Description
[Technical Field]
[0001] Priority claim This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 626,778, filed January 30, 2025, the contents of which are incorporated herein by reference.
[0002] This document relates generally, but not exclusively, to medical imaging systems, and more particularly to endoscopy systems. [Background technology]
[0003] Endoscopy is a medical procedure that allows physicians to view the inside of the body without making large incisions. The procedure involves the use of an endoscope, a flexible tube equipped with a light and a camera. The endoscope can be inserted through a natural opening in the body, such as the mouth, or through a small incision. Endoscopy is used for a variety of purposes, including diagnosing and treating conditions within the digestive tract, respiratory system, or other organs. It is a minimally invasive method that provides a detailed view of internal organs, allowing for more accurate diagnosis and targeted treatment. Common types of endoscopic procedures include gastroscopy to examine the upper digestive tract, colonoscopy for the lower intestine, and bronchoscopy for the lungs.
[0004] Endoscopic procedures are highly beneficial in modern medicine due to their minimally invasive nature, which leads to shortened recovery times and a lower risk of complications compared to traditional surgery. These procedures are generally safe and are performed under local or general anesthesia to ensure patient comfort. Physicians not only diagnose conditions via endoscopes, but may also perform various treatments, such as biopsies, polyp removal, and certain surgical procedures, directly through the endoscope. This versatility has made endoscopy an essential tool in many medical fields, including gastroenterology, pulmonology, and oncology. Advances in endoscopic technology continue to increase its effectiveness, providing high-resolution images and new techniques for treatment and diagnosis. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] WO 2011 / 140118A1 Summary of the Invention [Means for solving the problem]
[0006] This disclosure describes various techniques for analyzing video recordings for endoscopic and other types of medical imaging. In some examples, the system generates intelligent interest prediction indicators that are displayed in association with a timeline search bar of the video recording. The intelligent interest prediction indicators are quantified and scored on the timeline based on one or more parameters. In other examples, the system intelligently selects keyframes from different video segments to display as thumbnails associated with those segments. These techniques may improve the efficiency, accuracy, and speed with which physicians review and identify salient aspects of recorded endoscopic procedures.
[0007] In some aspects, the present disclosure is directed to a system for navigating frames of a segment of a video recording of a medical procedure, the system comprising: a user interface including a display; and a processing unit configured to select a thumbnail image for the segment based on an assessment of potential interest, display the selected thumbnail image on the user interface, and receive input from a user selecting the displayed thumbnail image on the user interface.
[0008] In some aspects, the present disclosure is directed to a system for navigating frames of a segment of a video recording of a medical procedure, the system comprising: a user interface including a display; and a processing unit configured to determine a predictive indicator for the segment of the video recording based on an assessment of potential interest and a current selection of one or more user-selectable parameters, display the predictive indicator on the user interface, and align the predictive indicator to a timeline of the segment.
[0009] In some aspects, the present disclosure is directed to a system for navigating frames of a segment of a video recording of a medical procedure, the system comprising a processing unit configured to determine a predictive indicator for the segment of the video recording based on an assessment of potential interest, dynamically adjust a compression rate of the segment based on the predictive indicator, and store the video recording at the adjusted compression rate.
[0010] In the drawings, which are not necessarily drawn to scale, like numerals may describe like components in different views. Like numerals with different letter prefixes may represent different instances of like components. The drawings illustrate generally by way of example, but not by way of limitation, various embodiments discussed in this document. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a schematic diagram of an example of an endoscopic system including an imaging and control system and an endoscope. [Figure 2] 2 is a schematic diagram of the endoscopic system of FIG. 1 including an endoscope connected to a control unit of the imaging and control system. [Figure 3] 1A-1C illustrate examples of frames from a segment of a video recording of a medical procedure and corresponding thumbnail images within the segment from the video recording. [Figure 4A]FIG. 10 illustrates an example graphical display of a predictive indicator displayed with various user-selectable parameters. [Figure 4B] FIG. 10 illustrates another example of a graphical display of a predictive indicator displayed with various user-selectable parameters. [Figure 5] 1 is a schematic diagram of an example of a computer-based tissue sample analyzer. [Figure 6] FIG. 1 is a schematic diagram of an example of a trained machine learning model. [Figure 7] 1 is a flow diagram of an example method for navigating frames of a segment of a video recording of a medical procedure. [Figure 8] 1 is a flow diagram of an example method for navigating frames of a segment of a video recording of a medical procedure. [Figure 9] 1 is a flow diagram of an example method for navigating frames of a segment of a video recording of a medical procedure. [Figure 10] FIG. 1 is a block diagram illustrating an example machine on which one or more examples may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0012] Endoscopic procedures are important diagnostic and therapeutic tools in the medical field. Advances in endoscopic devices and imaging technology have enabled high-quality video capture during endoscopic examinations. However, increasing video volumes make it difficult for physicians to efficiently review and analyze the recorded footage.
[0013] Manually scrolling through these large video files to identify important sections can be time-consuming and inefficient. Artificial intelligence and machine learning algorithms have shown promise in the automated analysis of medical images and videos. However, techniques for integrating these algorithms to assist physicians in quickly navigating recorded endoscopic videos are lacking.
[0014] Therefore, the present inventors have recognized a need for an improved graphical user interface system and method that leverages analysis of endoscopic video content to assist physicians in quickly locating important video segments and key images. Intelligent analysis combined with a dynamic interface may assist physician workflow and improve patient care.
[0015] This disclosure describes various techniques for analyzing video recordings for endoscopic and other types of medical imaging. In some examples, the system generates intelligent interest prediction indicators that are displayed in association with a timeline search bar of the video recording. The intelligent interest prediction indicators are quantified and scored on the timeline based on one or more parameters. In other examples, the system intelligently selects keyframes from different video segments to display as thumbnails associated with those segments. These techniques may improve the efficiency, accuracy, and speed with which physicians review and identify salient aspects of recorded endoscopic procedures.
[0016] FIG. 1 is a schematic diagram of an example endoscopic system 10 including an imaging and control system 12 and an endoscope 14. The endoscopic system 10 is suitable for use with the systems, devices, and methods described below, such as modular endoscopic systems, modular endoscopes, and methods for designing, constructing, and disassembling endoscopes. According to some examples, the endoscope 14 may be insertable into an anatomical region for imaging and / or to provide for the passage of one or more sampling devices for biopsies or one or more therapeutic devices for treatment of a disease condition associated with the anatomical region. The endoscope 14 may advantageously interface and connect with the imaging and control system 12. In the illustrated example, the endoscope 14 includes a colonoscope, although other types of endoscopes may be used with the features and teachings of the present disclosure.
[0017] The imaging and control system 12 may include a controller 16, a user interface including an output unit 18 (e.g., a display) and an input unit 20, a light source 22, a fluid source 24, and a suction pump 26. The imaging and control system 12 may include various ports for coupling with the endoscope system 10. For example, the controller 16 may include a data input / output port for receiving data from and transmitting data to the endoscope 14.
[0018] The light source 22 may include an output port for transmitting light to the endoscope 14, such as via a fiber optic link. The fluid source 24 may include a port for delivering fluid to the endoscope 14. The fluid source 24 may comprise a pump and a tank of fluid, or may be connected to an external tank, container, or storage unit. The suction pump 26 may include a port used to draw a vacuum from the endoscope 14 to generate suction, such as to draw fluid from the anatomical region into which the endoscope 14 is inserted. The output unit 180 and the input unit 20 may be used by the operator of the endoscopic system 10 to control functions of the endoscopic system 10 and to view the output of the endoscope 14.
[0019] The controller 16 may additionally be used to generate signals or other outputs from treating the anatomical region into which the endoscope 14 is inserted. In some examples, the controller 16 may generate electrical outputs, acoustic outputs, fluid outputs, etc. to treat the anatomical region using, for example, cauterization, cutting, freezing, etc.
[0020] The endoscope 14 may include an insertion portion 28, a functional portion 30, and a handle portion 32, which may be coupled to a cable portion 34 and a coupler portion 36. The insertion portion 28 may extend distally from the handle portion 32, and the cable portion 34 may extend proximally from the handle portion 32. The insertion portion 28 may be elongated and include a bent portion and a distal end to which the functional portion 30 may be attached. The bent portion may be controllable (e.g., by a control knob 38 on the handle portion 32) for maneuvering the distal end through tortuous anatomical passageways (e.g., the stomach, duodenum, kidney, urinary tract, etc.). The insertion portion 28 may also be elongated and include one or more working channels (e.g., internal lumens) that may support insertion of one or more therapeutic tools of the functional portion 30. The working channels may extend between the handle portion 32 and the functional portion 30. Additional functions such as fluid passageways, guidewires, and pullwires may also be provided by insert 28 (eg, via aspiration or irrigation passageways, etc.).
[0021] The handle module 32 may include a knob 38 as well as a port 40. The knob 38 may be coupled to a pull wire extending through the insert 28. The port 40 may be configured to couple various electrical cables, fluid tubing, etc. to the handle module 32 for coupling with the insert 28.
[0022] The imaging and control system 12, according to some examples, may be mounted on a moving platform (e.g., cart 41) having shelves for housing the light source 22, suction pump 26, image processing unit 42, etc. Alternatively, some components of the imaging and control system 12 shown in Figures 1 and 2 may be mounted directly on the endoscope 14 to make the endoscope "self-contained."
[0023] Figure 2 is a schematic diagram of the endoscopic system 10 of Figure 1, including an imaging and control system 12 and an endoscope 14. Figure 2 shows schematically the components of the imaging and control system 12 coupled to the endoscope 14, which in the illustrated example includes a colonoscope. The imaging and control system 12 may include or be coupled to an image processing unit 42, a treatment generator 44, and a drive unit 46, as well as a light source 22, an input unit 20, and an output unit 18. The image processing unit 42 includes one or more processors that may be distributed, such as locally and remotely, or located in one location.
[0024] The image processing unit 42 and the light source 22 may each interface with the endoscope 14 by a wired or wireless electrical connection. Thus, the imaging and control system 12 may illuminate an anatomical region, collect signals representative of the anatomical region, process the signals representative of the anatomical region, and display an image representative of the anatomical region on the display unit 18. The imaging and control system 12 may include a light source 22 that illuminates the anatomical region using a desired spectrum of light (e.g., broadband white light, narrowband imaging using a preferred electromagnetic wavelength, etc.). The imaging and control system 12 may connect to the endoscope 14 (e.g., via an endoscope connector) for signal transmission (e.g., light output from the light source, video signals from the imaging system at the distal tip, etc.).
[0025] Fluid source 24 may include a source of air, saline, or other fluid, as well as associated fluid passageways (e.g., air channels, irrigation channels, aspiration channels) and connectors (barbed fittings, fluid seals, valves, etc.). Imaging and control system 12 may also include a drive unit 46, which may be an optional component. Drive unit 46 may include a motorized drive for advancing the distal portion of endoscope 14, as described in PCT Publication No. WO 2011 / 140118 A1 to Frassica et al., entitled "Rotate-to-Advance Catheterization System," which is incorporated herein by reference in its entirety.
[0026] 3 shows an example image 300 of a segment of a video recording of a medical procedure. As previously mentioned, this disclosure describes a system that intelligently selects keyframes from different video segments to display as thumbnails associated with the segments. These techniques may improve the efficiency, accuracy, and speed with which a physician can review and identify salient aspects of a recorded endoscopic procedure.
[0027] A processing unit of the system, such as image processing unit 42 of endoscopic system 10 of FIG. 2 or another processing unit not associated with endoscopic system 10, may divide the video recording of the medical procedure into multiple segments, such as, in a non-limiting example, eight segments. Then, based on an assessment of potential interest, the processing unit may select a thumbnail image for each segment.
[0028] In some examples, the assessment of potential interest is determined based on an analysis of the video recording. For example, the analysis of the video recording may include detection scores for various parameters. These parameters may include disease detection (e.g., Computer-Aided Detection (CADe) and Computer-Aided Diagnosis (CADx) output) and potential associated confidence levels, activation of optical imaging modes (e.g., white light, narrow band imaging (NBI)) at various stages, endoscope withdrawal speed, bowel cleanliness score (which may be determined by mucosal detection rate), tool detection in the image, blood detection in the image, and foreign body detection in the image. In this way, the processing unit may intelligently select key frames from different video segments of the video recording of the medical procedure.
[0029] CADe includes algorithms or systems designed to assist radiologists or other medical professionals by highlighting areas of interest within medical images. These areas may indicate the presence of abnormalities such as tumors, fractures, or other pathological changes. CADe systems do not make diagnoses; they simply draw attention to areas that may require further analysis. CADx systems go a step further by not only detecting abnormalities but also providing an interpretation or possible diagnosis. These systems analyze medical images and offer diagnostic suggestions based on visual patterns they detect. CADx systems can be used to aid the decision-making process by providing additional information or a second opinion to radiologists or physicians.
[0030] Based on the evaluation of potential interest for each corresponding segment, the processing unit selects a corresponding thumbnail image for each of the segments, which is shown in thumbnail ribbon 302 as thumbnail image 304a-304h. The system displays the selected thumbnail image 304a-304h on a user interface, such as on display 18 of FIG. 1 or on another display unrelated to endoscopy system 10, such as a personal computing device or tablet computing device. In some examples, the selected thumbnail image is displayed along timeline 306 of the video recording. In this way, the clinician knows where the selected thumbnail image occurred within the video recording of the medical procedure.
[0031] Next, via a user interface, such as input unit 20 of Figure 1, the system receives input from the user selecting a displayed thumbnail image. For example, in Figure 3, the user selects displayed thumbnail image 304e, which the processing unit identifies as a keyframe and displays above as larger image 300. If the user selects another keyframe that the processing unit selected from a segment, such as thumbnail image 304b, thumbnail image 304b is displayed above as larger image 300, thereby allowing the user to navigate between segments by selecting displayed thumbnail images.
[0032] As mentioned above, in some examples, the system generates intelligent predictive interest indicators 308. A processing unit of the system, such as image processing unit 42 of endoscopic system 10 of FIG. 2, may determine predictive indicators 308 for segments of a video recording based on an assessment of potential interest.
[0033] In some examples, the assessment of potential interest is determined based on an analysis of the video recording. The intelligent interest prediction indicators are quantified and scored on a timeline based on one or more parameters. For example, the analysis of the video recording includes a detection score for various parameters. As above, in some examples, the analysis of the video recording includes a detection score for one or more of the following parameters: disease detection, activation of optical imaging mode, scope speed, bowel cleanliness, presence of tools, presence of foreign bodies, and presence of blood. In some examples, as shown in detail in FIG. 4, the parameters are user-selectable.
[0034] The processing unit may then display a predictive indicator 308 on a user interface, such as display 18 of Figure 1. The predictive indicator 308 includes a line 310 having peaks and valleys, where the peaks represent higher levels of assessed potential interest and the valleys represent lower levels of assessed potential interest. The processing unit may align the predictive indicator 308 with the timeline 306 of the video recording.
[0035] Additionally, the thumbnail image may be aligned with the predictive indicator 308. For example, thumbnail image 304e is shown aligned with the highest peak of the predictive indicator 308. After viewing the predictive indicator 308, a user, e.g., a clinician, may select thumbnail image 304e to further review that segment of the video recording.
[0036] In some examples, the processing unit is configured to dynamically adjust the compression rate of the segments based on the predictive indicators and store the video recording at the adjusted compression rate. For example, it may be desirable to increase the compression rate of segments having predictive indicators below a threshold, which may allow the system to store more data if assessed to be of low interest. In some examples, the system may increase the compression rate of segments having predictive indicators below a threshold after a configurable period of time has elapsed. Similarly, it may be desirable to decrease the compression rate of segments having predictive indicators above a threshold, which may allow the system to improve the quality of the stored video recording by saving more data.
[0037] In some examples, the intelligent keyframe selection and / or assessment of potential interest for the intelligent interest prediction indicator is determined based on a trained machine learning model, such as described below with respect to Figures 5 and 6. In some such examples, the trained machine learning model is trained using past behavior of a clinician.
[0038] It should be noted that in other examples, the analysis of the video recordings need not be performed by the endoscopy system 10. Instead, the video recordings may be stored on a mass storage device external to the endoscopy system 10, such as a central server within a hospital or a remotely located computing device, e.g., a cloud-based computing device. The techniques of the present disclosure may then be used, and the results displayed on a user interface not associated with the endoscopy system, such as a tablet, laptop computer, desktop computer, or some other device having a display.
[0039] 4A shows an example of a graphical display 400 of a predictive indicator 308 displayed with various user-selectable parameters 402. A processing unit may generate the graphical display 400 on a user interface, such as the display 18 of FIG.
[0040] The parameters used by the processing unit to determine the potential interest rating may be user-selectable to dynamically update the intelligent interest prediction indicator score based only on a currently selected subset of available parameters. In this way, if the user remembers detecting blood and wishes to review the associated video, the user can deselect all parameters except for detecting blood. This may immediately draw the user's attention to the particular portion of the recording that shows blood.
[0041] 4A, user-selectable parameters 402 are displayed and data representing the user-selectable parameters is aligned to the segment's timeline, e.g., activated lighting mode 406 data 404 is aligned to the video timeline 408.
[0042]
[0023] Figure 4B shows another example of a graphical display of a predictive indicator displayed with various user-selectable parameters. Figure 4B includes similar features to those shown and described above with respect to Figure 4A, and like reference numerals are used for such features. For the sake of brevity, those features will not be described again in detail.
[0043] In the example shown in FIG. 4A, all of the user-selectable parameters 402 have been selected. In contrast, the graphical display 410 shown in FIG. 4B has only a subset of the user-selectable parameters 402 selected, namely, disease detection parameters and velocity parameters. In other examples, additional or alternative parameters may be selected. Selecting or deselecting one or more of the user-selectable parameters 402 may dynamically adjust the predictive indicators 308 to indicate where a processing unit, e.g., the image processing unit 42 of FIG. 2, identified anomalies, so that a physician reviewing the video may easily find these instances. For example, the dynamic adjustment of the predictive indicators 308 may shift the peaks and valleys of the line 310, indicating changes in the assessed level of potential interest associated with the multiple thumbnail images in the thumbnail ribbon 302. In this manner, the processing unit determines the predictive indicators for a segment of the video recording based on the assessment of potential interest and the current selection of one or more user-selectable parameters. Upon receiving user input that changes the current selection to an updated selection, the processing unit dynamically adjusts the predictive indicators based on the updated selection.
[0044] 5 shows a schematic diagram of an example computer-based assessment of potential interest analyzer 500. The computer-based assessment of potential interest analyzer 500 is configured to, among other things, select corresponding thumbnail images based on an assessment of potential interest, such as based on an analysis of video footage, for one or more segments.
[0045] In some examples, the computer-based assessment analyzer of potential interest 500 analyzes the video recording and generates detection scores for various parameters. In some examples, the computer-based assessment analyzer of potential interest 500 may include an input interface 502 through which the various parameters are provided as input features to a trained machine learning (ML) or AI model, such as a trained artificial intelligence (AI) model 504. One or more relevant input parameters 510, which may be extracted from various component outputs 512, are applied to the AI model to generate predicted outputs from the AI model inference 506. The relevant input parameters 510 may include, but are not limited to, one or more of the following parameters: presence of disease, activation of an optical imaging mode, scope speed, bowel cleanliness, presence of a foreign body, presence of a tool, and presence of blood.
[0046] For example, one or more relevant input parameters 510, which may be extracted from sensor data generated by component outputs 512 of various components of the endoscopic system 10 of FIG. 1, may be applied to the trained AI model 504, which may generate confidence scores 514 for the various parameters. Then, based on an assessment of potential interest, the AI model 504 selects a thumbnail image for display on a user interface for the segment.
[0047] In another example, the computer-based assessment of potential interest analyzer 500 determines a predictive indicator for a segment of a video recording based on the assessment of potential interest, such as predictive indicator 308 of Figure 3. A processing unit, such as image processing unit 42 of Figure 2, displays the predictive indicator on a user interface and aligns the predictive indicator with the timeline of the segment, as shown in Figure 4A.
[0048] In some embodiments, the input interface 502 may be a direct data link between the computer-based assessment analyzer of potential interest 500 and one or more medical devices (e.g., the endoscopy system 10 of FIG. 1 ) that generate at least some of the input parameters. Additionally or alternatively, the input interface 502 may be a conventional user interface that facilitates interaction between a user and the computer-based assessment analyzer of potential interest 500. For example, the input interface 502 may facilitate a user interface through which a user may manually input information.
[0049] Based on one or more of the input parameters, the predicted output from the AI model inference 506 generates a rating of potential interest and performs an inference operation using the AI model 504 to select a corresponding thumbnail image for a segment of the video recording based on the rating of potential interest. For example, the input interface 502 may deliver input parameters to the input layer of the AI model 504, which propagates these input parameters through the AI model 504 to the output layer. The AI model 504 may provide a computer system with the ability to perform tasks without being explicitly programmed by making inferences based on patterns found in analyzing data. The AI model 504 explores the research and construction of algorithms (e.g., machine learning algorithms) that can learn from existing data and make predictions on new data. Such algorithms operate by building an AI model from example training data to make data-driven predictions or decisions, which are expressed as outputs or ratings.
[0050] There are two general modes of machine learning (ML): supervised ML and unsupervised ML. Supervised ML uses prior knowledge (e.g., examples that associate inputs with outputs or outcomes) to learn the relationship between inputs and outputs. The goal of supervised ML is to learn a function that best approximates the relationship between training inputs and outputs, given some training data, so that an ML model can implement the same relationship given the input to generate the corresponding output. Unsupervised ML is the training of ML algorithms using uncategorized or unlabeled information, allowing the algorithm to act on that information without guidance. Supervised ML is useful in exploratory analysis because it can automatically identify structure in data.
[0051] Common tasks in supervised ML are classification and regression problems. Classification problems, also called categorization problems, aim to classify items into one of several categorical values (e.g., is this object an apple or an orange?). Regression algorithms aim to quantify some items (e.g., by providing a score for the value of some input). Some examples of commonly used supervised ML algorithms are logistic regression (LR), naive Bayes, random forest (RF), neural network (NN), deep neural network (DNN), matrix decomposition, and support vector machine (SVM).
[0052] Some common tasks in unsupervised ML include clustering, representation learning, and density estimation. Some examples of commonly used unsupervised ML algorithms are K-means, principal component analysis, and autoencoders.
[0053] Another type of ML is federated learning (also known as collaborative learning), in which an algorithm is trained across multiple distributed devices that hold local data without exchanging data. This approach contrasts with traditional centralized machine learning techniques, in which all local datasets are uploaded to a single server, as well as more classical distributed approaches that often assume that local data samples are identically distributed. Federated learning allows multiple actors to build a common, robust machine learning model without sharing data, thus addressing important issues such as data privacy, data security, data access rights, and access to heterogeneous data.
[0054] In some examples, the AI model may be continuously or periodically trained prior to execution of the inference operation with predicted outputs from the AI model inference 506. Then, during the inference operation, the patient-specific input features provided to the AI model may be propagated from the input layer, through one or more hidden layers, and finally to the output layer.
[0055] Using these techniques, a processing unit such as image processing unit 42 of FIG. 2 may select thumbnail images for a segment based on an assessment of potential interest, display the selected thumbnail images on a user interface, and receive input on the user interface from a user selecting the displayed thumbnail images.
[0056] In other examples, the processing unit may determine a predictive indicator for a segment of the video recording based on the assessment of potential interest, display the predictive indicator on a user interface, and align the predictive indicator with a timeline of the segment.
[0057] FIG. 6 shows a schematic diagram of an example trained machine learning model 600. One approach to training data for developing the trained machine learning model 600 is to utilize annotated endoscopic video footage. This training data includes sample videos that have been reviewed by clinical experts and annotated to indicate timestamps at which various parameters of interest occurred. For example, training videos are annotated to tag segments indicating the presence of specific disease states, activation of specific imaging modes, changes in scope speed, bowel cleanliness scores, detection of foreign bodies, presence of tools, and presence of blood. The annotations indicate the start time and duration of each of these events of interest within the video.
[0058] By training a machine learning model on a dataset containing these expert-annotated videos, the model learns to predict the likelihood of these various events occurring in new, unlabeled endoscopic footage. The trained model outputs a detection score for each parameter over time based on the patterns learned from the annotated training data.
[0059] The imaging processing unit 42 of Figure 2 may be used to generate data. For example, the imaging processing unit 42 generates data regarding one or more of the following parameters: presence of disease, activation of optical imaging mode, scope speed, bowel cleanliness, presence of foreign body, presence of tool, and presence of blood. For example, an imaging device 602, such as the imaging and control system 12 of Figure 1, generates data such as blood presence data 604, disease data 604, and speed data 608. The data is used to generate N sets of video training data 610, such as one or more of disease data N612, scope speed data N614, and blood presence data N616.
[0060] Before the video training data 610 is ready to be used as training data 618, one or more signal processing steps, such as sampling, feature extraction, filtering, etc., may be performed on the video training data 610. The training data 618 may include N sets of training data based on the video training data 610. In addition, the training data 618 may include annotation training data 620. For example, the annotation training data 620 may include a set of labeling data N 622 and timestamp data N 624 generated by a medical professional 626. The neural network structure 628 may include labels associated with one or more parameters. The timestamp data N 624 includes timestamps at which various parameters of interest occur.
[0061] The training data 618 is used to train a trained machine learning model 600, e.g., an AI or machine learning model such as the AI model 504 of FIG. 5. The training data 618 may be applied to a neural network structure 628, such as a DNN, that includes an input layer, one or more hidden layers, and an output layer. The training data 618 and annotated training data 620 may be fed to the input layer of the neural network structure 628, which propagates the input data or data features through one or more hidden layers to the output layer, which outputs weights and biases to form the trained machine learning model 600. The trained machine learning model 600 can perform tasks without being explicitly programmed by making inferences based on patterns found in an analysis of data.
[0062] 7 shows a flow diagram of an example method 700 for navigating frames of a segment of a video recording of a medical procedure. At block 702, the method 700 includes selecting a thumbnail image for the segment based on an assessment of potential interest. At block 704, the method 700 includes displaying the selected thumbnail image on a user interface. At block 706, the method 700 includes receiving input from a user on the user interface to select the displayed thumbnail image.
[0063] 8 shows a flow diagram of an example method 800 for navigating frames of a segment of a video recording of a medical procedure. At block 802, the method 800 includes determining a predictive indicator for the segment of the video recording based on an assessment of potential interest and a current selection of one or more user-selectable parameters. At block 804, the method 800 includes displaying the predictive indicator on a user interface. At block 806, the method 800 includes aligning the predictive indicator to a timeline of the segment.
[0064] 9 shows a flow diagram of an example method 900 for navigating frames of a segment of a video recording of a medical procedure. At block 902, the method 900 includes determining a predictive indicator for the segment of the video storage based on an assessment of potential interest. At block 904, the method 900 includes dynamically adjusting a compression rate for the segment based on the predictive indicator. At block 906, the method 900 includes storing the video recording at the adjusted compression rate.
[0065] FIG. 10 shows a block diagram of an example machine 1000 on which any one or more of the techniques (e.g., methodologies) discussed herein may be implemented. Examples described herein may include or be operated by logic or certain components or mechanisms within machine 1000. Circuitry (e.g., processing circuitry) is a collection of circuitry implemented in a tangible entity of machine 1000, including hardware (e.g., simple circuits, gates, logic, etc.). The membership of circuitry may be flexible over time. Circuitry includes members that, when operating, may perform specific operations, either alone or in combination. In one example, circuitry hardware may be invariably designed (e.g., hard-wired) to perform specific operations. In one example, circuitry hardware may include variably connected physical components (e.g., execution units, transistors, simple circuits, etc.) that include machine-readable media that have been physically modified (e.g., magnetically, electrically, with a variable arrangement of invariant massed particles, etc.) to encode instructions for specific operations. When connecting physical components, the underlying electrical properties of the hardware components are changed, for example, from insulator to conductor, or vice versa. The instructions enable the embedded hardware (e.g., an execution unit or loading mechanism) to create circuitry members within the hardware through variable connections to perform certain portions of an operation during operation. Thus, in one example, a machine-readable medium element is part of the circuitry or is communicatively coupled to other components of the circuitry when the device is operating. In one example, any of the physical components may be used in two or more members of two or more circuitries. For example, during operation, an execution unit may be used in a first circuit of a first circuitry at one time and reused by a second circuit within the first circuitry or a third circuit within the second circuitry at a different time. Additional examples of these components with respect to machine 1000 follow.
[0066] In alternative examples, machine 1000 may operate as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, machine 1000 may operate as a server machine, a client machine, or both in a server-client network environment. In one example, machine 1000 may function as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. Machine 1000 may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a network router, switch, or bridge, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken by the machine. Additionally, although only a single machine is shown, the term "machine" shall be taken to include any collection of machines that individually or collectively execute a set (or sets) of instructions to perform any one or more of the methodologies discussed herein, such as cloud computing, software as a service (SaaS), other computer cluster configurations, etc.
[0067] The machine 1000 may include a hardware processor 1002 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 1004, static memory (e.g., memory or storage for firmware, microcode, basic-input-output (BIOS)), and mass storage 1008 (e.g., a hard drive, tape drive, flash storage, or other block device), some or all of which may communicate with each other via an interlink 1030 (e.g., a bus). The machine 1000 may further include a display unit 1010, an alphanumeric input device 1012 (e.g., a keyboard), and a user interface (UI) navigation device 1014 (e.g., a mouse). In one example, the display unit 1010, the input device 1012, and the UI navigation device 1014 may be touchscreen displays. The machine 1000 may additionally include a signal generating device 1018 (e.g., a speaker), a network interface device 1020, and one or more sensors 1016, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensor. The machine 1000 may include an output controller 1028, such as a serial (e.g., universal serial bus (USB)), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection, for communicating with or controlling one or more peripheral devices (e.g., printers, card readers, etc.).
[0068] The processor 1002, the main memory 1004, the static memory 1006, or the registers of the mass storage 1008 may be or include a machine-readable medium 1022 on which is stored one or more sets of data structures or instructions 1024 (e.g., software) that embody or are utilized by any one or more of the techniques or functions described herein. The instructions 1024 may also reside, completely or at least partially, within any of the processor 1002, the main memory 1004, the static memory 1006, or the registers of the mass storage 1008 during their execution by the machine 1000. In one example, one or any combination of the hardware processor 1002, the main memory 1004, the static memory 1006, or the mass storage 1008 may constitute the machine-readable medium 1022. Although the machine-readable medium 1022 is shown as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) configured to store one or more instructions 1024.
[0069] The term "machine-readable medium" may include any medium capable of storing, encoding, or carrying instructions for execution by machine 1000, causing machine 1000 to perform any one or more of the techniques of this disclosure, or storing, encoding, or carrying data structures used by or related to such instructions. Non-limiting examples of machine-readable media may include solid-state memory, optical media, magnetic media, and signals (e.g., radio frequency signals, other photon-based signals, acoustic signals, etc.). In one example, a non-transitory machine-readable medium comprises a machine-readable medium having a plurality of particles with an unchanging (e.g., stationary) mass and is thus a composition of matter. Thus, a non-transitory machine-readable medium is a machine-readable medium that does not include a transitory, propagating signal. Specific examples of non-transitory machine-readable media may include non-volatile memory such as semiconductor memory devices (e.g., Electrically Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM)) and flash memory devices, magnetic disks such as internal hard disks and removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0070] The information stored or otherwise provided on the machine-readable medium 1022 may represent the instructions 1024, such as the instructions 1024 themselves or a format from which the instructions 1024 may be derived. The format from which the instructions 1024 may be derived may include source code, encoded instructions (e.g., in compressed or encrypted form), packaged instructions (e.g., divided into multiple packages), etc. The information representing the instructions 1024 in the machine-readable medium 1022 may be processed by a processing circuit into instructions for implementing any of the operations discussed herein. For example, deriving the instructions 1024 from information (e.g., processing by a processing circuit) may include compiling (e.g., from source code, object code, etc.), interpreting, loading, organizing (e.g., dynamically or statically linking), encoding, decoding, encrypting, decrypting, packaging, unpackaging, or otherwise processing the information into the instructions 1024.
[0071] In one example, deriving the instructions 1024 may include assembling, compiling, or interpreting information (e.g., by a processing circuit) to create the instructions 1024 from some intermediate or preprocessed format provided by the machine-readable medium 1022. If the information is provided in multiple parts, it may be combined, unpacked, and modified to create the instructions 1024. For example, the information may reside in multiple compressed source code packages (or object code, or binary executable code, etc.) on one or more remote servers. The source code packages may be encrypted when transferred over a network, decrypted, decompressed, assembled (e.g., linked) as needed, compiled or interpreted on the local machine (e.g., into a library, a standalone executable, etc.), and executed by the local machine.
[0072] The instructions 1024 may further be transmitted or received over a communications network 1026 using a transmission medium via the network interface device 1020 utilizing any one of several transport protocols (e.g., frame relay, internet protocol (IP), transmission control protocol (TCP), user datagram protocol (UDP), hypertext transfer protocol (HTTP), etc.). Exemplary communication networks may include, among others, a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), a LoRa / LoRaWAN, or satellite communication network, a cellular network (e.g., a cellular network such as one conforming to the 3G, 4G LTE / LTE-A, or 5G standards), a Plain Old Telephone (POTS) network, and a wireless data network (e.g., the Institute of Electrical and Electronics Engineers (IEEE) 702.11 family of standards known as Wi-Fi®, the IEEE 1002.15.4 family of standards, a peer-to-peer (P2P) network). In one example, the network interface device 1020 may include one or more physical jacks (e.g., an Ethernet jack, a coaxial jack, or a telephone jack) or one or more antennas for connecting to the communication network 1026.In one example, network interface device 1020 may include multiple antennas for communicating wirelessly using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) techniques. The term "transmission medium" is intended to include any intangible medium capable of storing, encoding, or carrying instructions for execution by machine 1000, including digital or analog communication signals or other intangible media that facilitate the communication of such software. Transmission media are machine-readable media.
[0073] Various notes Each of the non-limiting claims and examples set forth herein may stand on its own or may be combined with one or more of the other examples in various permutations or combinations.
[0074] The above detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings show, by way of illustration, specific embodiments in which the invention may be practiced. These embodiments are also referred to herein as "examples." Such examples may include elements in addition to those shown or described. However, the inventors also contemplate examples in which only those elements shown or described are provided. Furthermore, the inventors also contemplate examples using any combination or permutation of those elements shown or described, either with respect to the particular example (or one or more claims thereof) or any other example (or one or more claims thereof) shown or described herein.
[0075] In the event of a conflict of usage between this document and any document incorporated by reference, the usage in this document takes precedence.
[0076] As used herein, the terms "a" or "an" are used, as is common in patent documents, to include one or more, regardless of any other instance or usage of "at least one" or "one or more." As used herein, the term "or" refers to a non-exclusive or, unless otherwise indicated, such that "A or B" includes "A but not B," "B but not A," and "A and B." As used herein, the terms "including" and "in which" are used as the plain English equivalents of the terms "comprising" and "wherein," respectively. Also, in the following claims, the terms "including" and "comprising" are open-ended, i.e., systems, devices, articles, compositions, formulations, or processes that include elements in addition to those listed after such terms in a claim are still deemed to be within the scope of that claim. Moreover, in the following claims, terms such as "first," "second," and "third" are used merely as labels and are not intended to impose numerical requirements on their objects.
[0077] Examples of the methods described herein may be at least partially machine- or computer-implemented. Some examples may include computer-readable or machine-readable media encoded with instructions operable to configure an electronic device to perform the methods described in the above examples. Implementations of such methods may include code, such as microcode, assembly language code, high-level language code, etc. Such code may include computer-readable instructions for performing various methods. The code may form part of a computer program product. Further, in one example, the code may be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media, such as during execution or at other times. Examples of these tangible computer-readable media may include, but are not limited to, hard disks, removable magnetic disks, removable optical disks (e.g., compact disks and digital video disks), magnetic cassettes, memory cards or sticks, random access memories (RAM), read-only memories (ROM), etc.
[0078] The above description is intended to be illustrative, not limiting. For example, the above-described examples (or one or more claims thereof) may be used in combination with each other. Other embodiments may be utilized by those of ordinary skill in the art upon reviewing the above description. The Abstract is provided to comply with 37 CFR §1.72(b) to allow the reader to quickly ascertain the nature of the technical disclosure. This Abstract is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Also, in the above Detailed Description, various features may be grouped together to streamline the disclosure. This should not be construed as intending that an unclaimed disclosed feature is essential to any claim. Rather, inventive subject matter may lie in fewer than all features of a particular disclosed embodiment. Accordingly, the following claims are hereby contemplated, with each claim standing on its own as a separate embodiment, incorporated into the Detailed Description as an example or embodiment, and such embodiments may be combined with each other in various combinations or permutations. The scope of the invention should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. [Explanation of symbols]
[0079] 10 Endoscopy System 12 Imaging and Control System 14 Endoscopy 16 Controllers 18 Output unit, display unit, display 20 input units 22 Light source 24 Fluid source 26 Suction pump 28 Insertion section 30 Functional Section 32 Handle section, handle module 34 Cable section 36 Coupler Section 38 Control knob, knob 40 ports 41 Cart 42 Image Processing Unit 44 Treatment Generator 46 Drive Unit 300 images 302 Thumbnail Ribbon 304a~304h thumbnail images 306 Timeline 308 Intelligent Interest Prediction Indicator, Prediction Indicator 310 line 400 Graphical Display 402 User Selectable Parameters 404 Data 406 Lighting Modes 408 Video Timeline 500 Computer-Based Assessment Analyzer of Potential Interest 502 input interface 504 trained AI models, AI models 506 AI Model Inference 510 Related Input Parameters 512 Component Output 514 Confidence Score 600 trained machine learning models 602 Imaging Device 604 Blood Presence Data 606 Disease Data 608 Speed Data 610 Video Training Data 612 Disease Data N 614 Scope Speed Data N 616 Blood Presence Data N 618 training data 620 annotated training data 622 Labeling Data N 624 Timestamp Data N 626 medical professionals 628 Neural Network Structure 1000 machines 1002 Hardware Processor 1004 main memory 1006 Static Memory 1008 Mass Storage 1010 Display Unit 1012 Alphanumeric Input Device 1014 User interface (UI) navigation device, UI navigation device 1016 Sensor 1018 Signal Generating Device 1020 Network Interface Device 1022 Machine-readable medium 1024 Data Structures or Instructions, Instructions 1026 Communication Network 1028 Output Controller 1030 mutual links
Claims
1. 1. A system for navigating frames of a segment of a video recording of a medical procedure, the system comprising: a user interface including a display; For a segment, selecting a thumbnail image based on an assessment of potential interest; displaying the selected thumbnail image on the user interface; receiving input from a user on the user interface to select the displayed thumbnail image; A processing unit configured as follows: A system for navigating a frame, comprising:
2. The system for navigating a frame of claim 1 , wherein the assessment of potential interest is determined based on an analysis of the video recording.
3. The analysis of the video recordings is based on the following parameters: the presence of disease, Activation of optical imaging mode, Scope speed, intestinal cleanliness, The presence of foreign bodies, The presence of tools, and Presence of blood a detection score for one or more of 3. A system for navigating a frame according to claim 2.
4. the video recording includes a plurality of segments, and the processing unit: for each of the plurality of segments, selecting a corresponding thumbnail image based on an assessment of potential interest in the corresponding segment; further configured to display, on the user interface, the selected thumbnail images aligned with their corresponding segments; The user interface: Allowing the user to navigate between segments by selecting the displayed thumbnail images It was configured as follows:
10. A system for navigating a frame according to claim 1.
5. The system for navigating a frame of claim 1 , wherein the assessment of potential interest is determined based on a trained machine learning model.
6. 6. The system for navigating a frame of claim 5, wherein the trained machine learning model is trained using past behavior of a clinician.
7. 1. A system for navigating frames of a segment of a video recording of a medical procedure, the system comprising: a user interface including a display; determining a predictive indicator for the segment of the video recording based on an assessment of potential interest and a current selection of one or more user-selectable parameters; displaying the predictive indicator on the user interface; Aligning the predictive indicator to the timeline of the segment A processing unit configured as follows: A system for navigating a frame, comprising:
8. The system for navigating a frame of claim 7 , wherein the assessment of potential interest is determined based on an analysis of the video footage.
9. The analysis of the video recording is based on the following user selectable parameters: the presence of disease, Activation of optical imaging mode, Scope speed, intestinal cleanliness, The presence of tools the presence of foreign bodies, and Presence of blood a detection score for one or more of 9. A system for navigating a frame according to claim 8.
10. The processing unit displaying data on the user interface representative of the user-selectable parameters; Aligning the data representing the user-selectable parameters to the timeline of the segment. It was configured as follows:
10. A system for navigating a frame according to claim 9.
11. 8. The system for navigating a frame of claim 7, wherein the predictive indicator comprises a line having peaks and valleys, the peaks representing higher levels of assessed potential interest and the valleys representing lower levels of assessed potential interest.
12. The processing unit dynamically adjusting a compression ratio for the segment based on the predictive indicator; Storing the video recording at the adjusted compression ratio. It was configured as follows:
8. A system for navigating a frame according to claim 7.
13. The processing unit configured to increase the compression rate of the segments having a predictive indicator below a threshold.
13. A system for navigating a frame according to claim 12.
14. The system for navigating a frame of claim 7 , wherein the assessment of potential interest is determined based on a trained machine learning model.
15. 15. The system for navigating a frame of claim 14, wherein the trained machine learning model is trained using past behavior of a clinician.
16. The processing unit receiving a user input that changes the current selection to an updated selection; Dynamically adjusting the predictive indicator based on the updated selection. It was configured as follows:
8. A system for navigating a frame according to claim 7.
17. 1. A system for navigating frames of a segment of a video recording of a medical procedure, the system comprising: determining a predictive indicator for the segment of the video footage based on an assessment of potential interest; dynamically adjusting a compression ratio for the segment based on the predictive indicator; Storing the video recording at the adjusted compression ratio. A processing unit configured as follows: A system for navigating a frame, comprising:
18. The processing unit Increasing the compression rate of said segments having predictive indicators below a threshold. It was configured as follows:
18. A system for navigating a frame according to claim 17.
19. Increasing the compression rate of said segments having predictive indicators below a threshold. The processing unit configured as described above configured to increase the compression rate of the segments having the predictive indicator below the threshold after a configurable period of time has elapsed.
20. A system for navigating a frame according to claim 18.
20. The system for navigating a frame of claim 17 , wherein the assessment of potential interest is determined based on an analysis of the video footage.
21. The analysis of the video recordings is based on the following parameters: the presence of disease, Activation of optical imaging mode, Scope speed, intestinal cleanliness, The presence of foreign bodies, The presence of tools, and Presence of blood a detection score for one or more of 21. A system for navigating a frame according to claim 20.
Citation Information
Patent Citations
Medical device for examining the neck
JP2014508021A
System and method for generating and displaying information for review of an in vivo image stream
JP2022505154A
Medical support operation method, device and computer program product
JP2023509075A
Learning device, learning method, image processing apparatus, endocope system, and program
WO2022049901A1
Program, information processing method, and information processing device
WO2022185651A1