PREDICTING INTEREST IN ENDOSCOPY VIDEOS IN TIMELINE
The system enhances endoscopic video review by using an intelligent interest prediction indicator and key image thumbnails, addressing inefficiencies in manual navigation and improving physician workflow.
Patent Information
- Application Number
- DE102025102490
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2025-01-23
- Publication Date
- 2025-07-31
AI Technical Summary
Endoscopic video recordings are voluminous, making manual review for salient parts time-consuming and inefficient, and existing AI techniques lack integration to assist physicians in rapid navigation.
A system generates an intelligent interest prediction indicator and selects key images as thumbnails based on potential interest scores, using machine learning to enhance navigation efficiency and accuracy.
The system significantly improves the efficiency, accuracy, and speed at which physicians can review and identify salient aspects of endoscopic procedures by intelligently navigating through video recordings.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
PRIORITY CLAIM
[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 626,778, filed January 30, 2025, the contents of which are incorporated herein by reference. AREA OF REVELATION
[0002] This document refers generally, but not restrictively, to medical imaging systems and in particular to endoscopy systems. BACKGROUND
[0003] Endoscopy is a medical procedure that allows doctors to view the inside of the body without making large incisions. This procedure uses an endoscope, a flexible tube with a light and camera attached. The endoscope can be inserted through natural body openings, such as the mouth, or through small incisions. Endoscopy is used for a variety of purposes, including the diagnosis and treatment of diseases of the gastrointestinal tract, respiratory tract, and other organs. It is a minimally invasive method that provides a detailed view of the internal organs, allowing for more accurate diagnosis and targeted treatments. Common endoscopic procedures include gastroscopy for examining the upper digestive tract, colonoscopy for the lower intestine, and bronchoscopy for the lungs.
[0004] Endoscopic procedures are of great benefit in modern medicine because they are minimally invasive, resulting in shorter recovery times and a lower risk of complications compared to traditional surgeries. These procedures are generally safe and are performed under local or general anesthesia for patient comfort. Doctors can not only diagnose through endoscopy but also perform various treatments such as biopsies, polyp removal, and even some surgical procedures directly through the endoscope. This versatility makes endoscopy an indispensable tool in many fields of medicine, including gastroenterology, pulmonology, and oncology. Advances in endoscopic technology are constantly improving its effectiveness, providing high-resolution images and new techniques for treatment and diagnosis. SUMMARY OF REVELATION
[0005] This disclosure describes various techniques for analyzing video recordings for endoscopic and other types of medical imaging. In some examples, a system generates an intelligent interest prediction indicator displayed in conjunction with a seek bar on the timeline of the video recording. The intelligent interest prediction indicator is quantified and scored across the timeline based on one or more parameters. In other examples, a system intelligently selects key frames from various video segments to display as thumbnails in conjunction with those segments. These techniques can increase the efficiency, accuracy, and speed with which a physician can review and identify salient aspects of a recorded endoscopic procedure.
[0006] In some aspects, this disclosure is directed to a system for navigating frame by frame of a segment of a video recording of a medical procedure, the system comprising: a user interface including a display; and a processing unit configured to: select a thumbnail for the segment based on a rating of potential interest; display the selected thumbnail on the user interface; and receive, on the user interface, input from a user selecting the displayed thumbnail.
[0007] In some aspects, this disclosure is directed to a system for navigating frame rates of a segment of a video recording of a medical procedure, the system comprising: a user interface including a display; and a processing unit configured to: determine a predictive indicator for the segment of the video recording based on a rating of potential interest and a current selection of one or more user-selectable parameters; display the predictive indicator on the user interface; and align the predictive indicator with a timeline of the segment.
[0008] In some aspects, this disclosure is directed to a system for navigating individual frames of a segment of a video recording of a medical procedure, the system comprising: a processing unit configured to: Determining a predictive indicator for the segment of the video recording based on a rating of potential interest; dynamically setting a compression rate of the segment based on the predictive indicator; and saving the video recording with the set compression rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In the drawings, which are not necessarily to scale, like numbers may describe similar components in different views. Like numbers with different letter suffixes may represent different variations of similar components. The drawings generally illustrate, by way of example but not limitation, various embodiments covered in this document. Fig. Figure 1 is a schematic diagram of an endoscopy system including an imaging and control system and an endoscope. Fig. 2 is a schematic diagram of the endoscopy system of Fig. 1, which comprises the endoscope connected to a control unit of the imaging and control system. Fig. 3 shows an example of an image of a segment of a video recording of a medical procedure and a corresponding thumbnail image within the segment of the video recording. Fig. Figure 4A shows an example graphical representation of a predictive indicator displayed along with various user-selectable parameters. Fig. Figure 4B shows another example of a graphical representation of a predictive indicator displayed along with various user-selectable parameters. Fig. Figure 5 is a schematic diagram of an example of a computer-assisted tissue sample analyzer. Fig. Figure 6 shows a schematic diagram of an example of a trained machine learning model. Fig. 7 shows a flowchart of an example method for navigating images of a segment of a video recording of a medical procedure. Fig. 8 shows a flowchart of an example method for navigating images of a segment of a video recording of a medical procedure. Fig. 9 shows a flowchart of an example method for navigating images of a segment of a video recording of a medical procedure. Fig. 10 is a block diagram showing an example machine on which one or more examples may be implemented. DETAILED DESCRIPTION
[0010] Endoscopy procedures are important diagnostic and therapeutic tools in the medical field. Advances in endoscopes and imaging technology have enabled high-quality video recording during endoscopic examinations. However, the increasing volume of video presents physicians with the challenge of efficiently reviewing and analyzing the recorded material.
[0011] Manually scrolling through these large video files to identify suspicious portions can be time-consuming and inefficient. Artificial intelligence and machine learning algorithms have shown promise in automatically analyzing medical images and videos. However, there is a lack of techniques for integrating these algorithms to assist physicians in quickly navigating recorded endoscopy videos.
[0012] Therefore, the inventors recognized a need for improved graphical user interface systems and methods that leverage endoscopic video content analysis to assist physicians in locating important video segments and key images. Intelligent analytics combined with dynamic interfaces can support physician workflows and improve patient care.
[0013] This disclosure describes various techniques for analyzing video recordings for endoscopic and other types of medical imaging. In some examples, a system generates an intelligent interest prediction indicator displayed in conjunction with a seek bar on the timeline of the video recording. The intelligent interest prediction indicator is quantified and scored across the timeline based on one or more parameters. In other examples, a system intelligently selects key frames from various video segments to display as thumbnails in conjunction with those segments. These techniques can increase the efficiency, accuracy, and speed with which a physician can review and identify salient aspects of a recorded endoscopic procedure.
[0014] Fig. 1 is a schematic representation of an example of an endoscopy system 10 including an imaging and control system 12 and an endoscope 14. The endoscopy system 10 is suitable for use with the systems, devices, and methods described below, such as modular endoscopy systems, modular endoscopes, and methods for designing, assembling, and disassembling endoscopes. According to some examples, the endoscope 14 may be inserted into an anatomical area for imaging and / or for delivering one or more sampling devices for biopsies or one or more therapeutic devices for treating a disease state associated with the anatomical area. The endoscope 14 may advantageously be connected and interfaced with the imaging and control system 12.In the example shown, the endoscope 14 is a colonoscope, although other types of endoscopes may be used with the features and teachings of the present disclosure.
[0015] The imaging and control system 12 may include a control unit 16, a user interface with an output unit 18 (e.g., a display) and an input unit 20, a light source 22, a fluid source 24, and a suction pump 26. The imaging and control system 12 may include various ports for connecting to the endoscopy system 10. For example, the control unit 16 may include a data input / output port for receiving data from and transmitting data to the endoscope 14.
[0016] The light source 22 may include an output port for transmitting light to the endoscope 14, for example, via a fiber optic connection. The fluid source 24 may include a port for transmitting fluid to the endoscope 14. The fluid source 24 may consist of a pump and a fluid reservoir or be connected to an external tank, container, or storage unit. The suction pump 26 may have a port used to create a vacuum from the endoscope 14, for example, to suction fluid from the anatomical region into which the endoscope 14 is inserted. The output unit 18 and the input unit 20 may be used by an operator of the endoscopy system 10 to control functions of the endoscopy system 10 and to display the output of the endoscope 14.
[0017] The controller 16 may additionally be used to generate signals or other outputs for treating the anatomical region into which the endoscope 14 has been inserted. In some examples, the controller 16 may generate an electrical output, an acoustic output, a fluid output, and the like for treating the anatomical region, e.g., by cauterizing, cutting, freezing, and the like.
[0018] The endoscope 14 may include an insertion portion 28, a functional portion 30, and a handle portion 32, which may be connected to a cable portion 34 and a coupling portion 36. The insertion portion 28 may extend distally from the handle portion 32, and the cable portion 34 may extend proximally from the handle portion 32. The insertion portion 28 may be elongated and include a bending portion and a distal end to which a functional portion 30 may be attached. The bending portion may be controllable (e.g., by a control knob 38 on a handle portion 32) to maneuver the distal end through tortuous anatomical passages (e.g., stomach, duodenum, kidney, ureter, etc.). The insertion section 28 may also include one or more working channels (e.g., an inner lumen) that may be elongated and assist in the insertion of one or more therapeutic tools of the functional section 30.The working channel may extend between the handle portion 32 and the functional portion 30. Additional functions, such as fluid passages, guide wires, and pull wires, may also be provided through the insertion portion 28 (e.g., via suction or irrigation channels, and the like).
[0019] The handle module 32 may include a knob 38 and connectors 40. The knob 38 may be connected to a pull wire extending through the insertion section 28. The connectors 40 may be configured to connect various electrical cables, fluid hoses, and the like to the handle module 32 for coupling to the insertion section 28.
[0020] The imaging and control system 12 may, according to some examples, be provided on a mobile platform (e.g., a cart 41) with compartments for accommodating the light source 22, the suction pump 26, the image processing unit 42, etc. Alternatively, several components of the imaging and control system 12 that are included in the Fig. shown, are attached directly to the endoscope 14 so that the endoscope is “self-contained”.
[0021] Fig. 2 is a schematic representation of the endoscopy system 10 of Fig. 1, including the imaging and control system 12 and the endoscope 14. Fig. Figure 2 schematically shows the components of the imaging and control system 12 coupled to the endoscope 14, which in the illustrated example comprises a colonoscope. The imaging and control system 12 may include the controller 16, which may contain or be connected to the image processing unit 42, the treatment generator 44, and the drive unit 46, as well as the light source 22, the input unit 20, and the output unit 18. The image processing unit 42 includes one or more processors, which may be distributed, e.g., locally and remotely controlled, or co-located.
[0022] The image processing unit 42 and the light source 22 may each be connected to the endoscope 14 through wired or wireless electrical connections. The imaging and control system 12 may accordingly illuminate an anatomical region, collect signals representative of the anatomical region, process signals representative of the anatomical region, and display images representative of the anatomical region on the display unit 18. The imaging and control system 12 may include the light source 22 to illuminate the anatomical region with light of the desired spectrum (e.g., broadband white light, narrowband imaging with preferred electromagnetic wavelengths, and the like). The imaging and control system 12 may be connected (e.g., via an endoscope port) to the endoscope 14 to transmit signals (e.g., light from the light source, video signals from the imaging system at the distal end, and the like).
[0023] The fluid source 24 may include one or more sources of air, saline, or other fluids, as well as associated fluid pathways (e.g., air channels, irrigation channels, suction channels) and ports (fittings, fluid seals, valves, etc.). The imaging and control system 12 may also include the drive unit 46, which may be an optional component. The drive unit 46 may include a motorized drive for advancing a distal portion of the endoscope 14, as described in PCT Pub. No. WO 2011 / 140118 A1 by Frassica et al., entitled "Rotate-to-Advance Catheterization System," which is hereby incorporated by reference in its entirety.
[0024] Fig. Figure 3 shows an example image 300 of a segment of a video recording of a medical procedure. As previously mentioned, this disclosure describes a system that intelligently selects key frames from various video segments, which are displayed as thumbnails of those segments. These techniques can increase the efficiency, accuracy, and speed with which a physician can review and identify salient aspects of a recorded endoscopic procedure.
[0025] A processing unit of a system, such as the image processing unit 42 of the endoscopy system 10 of Fig. 2 or another processing unit not connected to the endoscopy system 10 may divide a video recording of a medical procedure into a plurality of segments, such as eight segments in one non-limiting example. The processing unit may then select a thumbnail image for each segment based on an assessment of potential interest.
[0026] In some examples, the assessment of potential interest is performed based on an analysis of the video recording. The analysis of the video recording includes, for example, a detection score for various parameters. These parameters may include: disease detection (e.g., CADe (computer-aided detection) and CADx (computer-aided diagnosis) output and a possible associated confidence level); activation of light imaging modes at various stages (e.g., white light, narrowband imaging (NBI)); endoscope withdrawal speed; bowel cleanliness score (which may be determined by the mucosal detection rate); detection of a tool in the image; detection of blood in the image; and detection of a foreign body in the image. In this way, the processing unit can intelligently select key frames from different video segments of a video recording of a medical procedure.
[0027] CADe refers to algorithms or systems designed to assist radiologists or other medical professionals by highlighting areas of interest in medical images. These areas may indicate the presence of abnormalities such as tumors, fractures, or other pathological changes. CADe systems do not provide a diagnosis but merely point out areas that may require further analysis. CADx systems go a step further by not only detecting abnormalities but also providing an interpretation or a possible diagnosis. These systems analyze medical images and offer diagnostic suggestions based on the visual patterns detected. CADx systems can be used as decision support by providing the radiologist or physician with additional information or a second opinion.
[0028] Based on the assessment of potential interest for each corresponding segment, the processing unit selects a corresponding thumbnail for each of the segments, which is represented in the thumbnail band 302 as thumbnail 304a through thumbnail 304h. The system displays the selected thumbnail 304a through thumbnail 304h, e.g., on the user interface, e.g., on the display 18 of Fig. 1, or on another display not connected to the endoscopy system 10, such as a personal computing device or a tablet computing device. In some examples, the selected thumbnail images are displayed along a timeline 306 of the video recording. This allows the physician to know at what point in the video recording of the medical procedure the selected thumbnail image occurred.
[0029] Next, the system receives via the user interface, such as the input unit 20 in Fig. 1, Input from a user who selects the displayed thumbnail. In Fig. 3, a user has selected the displayed thumbnail 304e, which the processing unit has identified as a keyframe, and which is displayed above the larger image 300. If the user selects a different keyframe that the processing unit has selected from a segment, e.g., thumbnail 304b, then the thumbnail 304b is displayed above the larger image 300, allowing the user to navigate between the segments by selecting the displayed thumbnails.
[0030] As mentioned above, in some examples, the system generates an intelligent interest prediction indicator 308. A processing unit of a system, such as the image processing unit 42 of the endoscopy system 10 of Fig. 2, may determine a predictive indicator 308 for a segment of the video recording based on an assessment of potential interest.
[0031] In some examples, the assessment of potential interest is performed based on an analysis of the video recording. The intelligent indicator for interest prediction is quantified and evaluated based on one or more parameters over the time axis. For example, the analysis of the video recording includes a detection assessment for various parameters. As described above, in some examples, the analysis of the video recording includes a detection assessment for one or more of the following parameters: presence of disease; activation of light imaging modes; speed of the endoscope; cleanliness of the bowel; presence of tools; presence of foreign objects; and presence of blood. In some examples, such as those detailed in Fig. As shown in Figure 4A, the parameters are user selectable.
[0032] The processing unit may then display the prediction indicator 308 on the user interface, e.g., on the display 18 of Fig. 1. The predictive indicator 308 includes a line 310 with peaks and valleys, where the peaks represent higher values of the assessed potential interest and the valleys represent lower values of the assessed potential interest. The processing unit can compare the predictive indicator 308 with a timeline 306 of the video recording.
[0033] Furthermore, the thumbnail images can be aligned to the predictive indicator 308. For example, the thumbnail 304e is aligned to the highest point of the predictive indicator 308. A user, such as a physician, can select the thumbnail 304e after viewing the predictive indicator 308 to further review that segment of the video recording.
[0034] In some examples, the processing unit is configured to dynamically adjust the compression rate of the segment based on the prediction indicator and save the video recording at the adjusted compression rate. For example, it may be desirable to increase the compression rate for segments whose prediction indicator is below a threshold, which may allow the system to store more data if they are determined to be less interesting. In some examples, the system may increase the compression rate for segments whose prediction indicator is below the threshold after a configurable period of time. Likewise, it may be desirable to decrease the compression rate for segments with a prediction indicator greater than a threshold, which may allow the system to improve the quality of the saved video recording by storing more data.
[0035] In some examples, the potential interest score for the intelligent keyframe selection and / or the intelligent interest prediction indicator is determined based on a trained machine learning model, as described below with respect to Fig. 5 and Fig. 6. In some such examples, the trained machine learning model is trained using a clinician's past behavior.
[0036] It should be noted that in other examples, the analysis of the video recording need not be performed by the endoscopy system 10. Instead, the video recordings may be stored on a mass storage device external to the endoscopy system 10, such as a central server at the hospital or on a remote computing device, e.g., a cloud-based computing device. The techniques of this disclosure may then be used, and the results may be displayed on a user interface that is not connected to the endoscopy system, e.g., on a tablet, a laptop, a desktop computer, or other device with a display.
[0037] Fig. 4A shows an example of a graphical display 400 of a prediction indicator 308 displayed along with various user-selectable parameters 402. The processing unit may display the graphical display 400 on a user interface, such as the display 18 in Fig. 1, generate.
[0038] The parameters used by the processing unit to determine the potential interest rating can be selected by the user to dynamically update the intelligent interest prediction indicator's rating based only on a currently selected subset of the available parameters. If a user recalls a blood detection and wants to view the associated video, they can deselect all parameters except blood detection. This can immediately draw the user's attention to the specific part of the recording showing the blood.
[0039] In the Fig. In the example shown in Figure 4A, the user-selectable parameters 402 are displayed, and the data representing the user-selectable parameters is aligned to the segment's timeline. For example, the data 404 of the activated lighting mode 406 is aligned to the video timeline 408.
[0040] Fig. Figure 4B shows another example of a graphical display of a prediction indicator displayed along with various user-selectable parameters. Fig. 4B contains similar functions to those in Fig. 4A, and similar reference numbers are used for these functions. For brevity, these features will not be described again in detail.
[0041] In the Fig. In the example shown in Figure 4A, all user-selectable parameters 402 were selected. In contrast, in the example shown in Fig. 4B, only a subset of the user-selectable parameters 402 is selected, namely the disease detection parameter and the speed parameter. In other examples, additional or alternative parameters may be selected.
[0042] The selection or deselection of one or more of the user-selectable parameters 402 can dynamically adjust the prediction indicator 308 to indicate where the processing unit, e.g., the image processing unit 42 of Fig. 2, has identified anomalies so that a physician reviewing the video can easily find those times. For example, dynamic adjustment of the predictive indicator 308 may cause the peaks and troughs of line 310 to shift, indicating changes in the levels of assessed potential interest associated with particular thumbnail images of the thumbnail tape 302. In this way, the processing unit determines a predictive indicator for the segment of the video recording based on an assessment of potential interest and a current selection of one or more user-selectable parameters. Upon receiving user input that changes the current selection to an updated selection, the processing unit dynamically adjusts the predictive indicator based on the updated selection.
[0043] Fig. Figure 5 shows a schematic diagram of an example of a computerized potential interest rating analyzer. The computerized potential interest rating analyzer (500) is configured, among other things, to select a corresponding thumbnail image for one or more segments based on a potential interest rating, for example, based on an analysis of the video recording.
[0044] In some examples, the computer-based analyzer 500 analyzes the video recording and generates a recognition score for various parameters. In some examples, the computer-based analyzer 500 for assessing potential interest may include an input interface 502 through which various parameters are provided as input features to a trained machine learning (ML) or artificial intelligence (AI) model, e.g., a trained AI model 504. One or more relevant input parameters 510, which may be extracted from various component outputs 512, are applied to the AI model to generate an output predicted by the AI model inference 506.The relevant input parameters 510 may include, but are not limited to, one or more of the following parameters: presence of disease, activation of light imaging modes, speed of the endoscope, cleanliness of the bowel, presence of foreign bodies, presence of tools, and presence of blood.
[0045] For example, one or more relevant input parameters 510 that can be extracted from sensor data received from component outputs 512 of various components of the endoscopy system 10 of Fig. 1 are applied to the trained AI model 504, and the AI model 504 can generate confidence values 514 for the various parameters. Based on an assessment of potential interest, the AI model 504 then selects a thumbnail image for a segment to display on a user interface.
[0046] In other examples, the computerized potential interest rating analyzer 500 determines a predictive indicator for the segment of the video recording based on the potential interest rating, such as the predictive indicator 308 in Fig. 3. The processing unit, e.g., the image processing unit 42 of Fig. 2, displays the prediction indicator on the user interface and aligns the prediction indicator to a timeline of the segment, as shown in Fig. 4A shown.
[0047] In some embodiments, the input interface 502 may provide a direct data connection between the computerized potential interest assessment analyzer 500 and one or more medical devices (e.g., the endoscopy system 10 of Fig. 1) that generates at least some of the input parameters. Additionally or alternatively, the input interface 502 may be a conventional user interface that facilitates interaction between a user and the computer-based analyzer 500 for evaluating potential interests. For example, the input interface 502 may be a user interface through which the user can manually enter information.
[0048] Based on one or more input parameters, the output predicted by AI model inference 506 performs an inference operation using AI model 504 to create a potential interest score and, based on the potential interest score, select a corresponding thumbnail for the segment of the video recording. For example, input interface 502 may provide the input parameters to an input layer of AI model 504, which passes these input parameters through AI model 504 to an output layer. AI model 504 may enable a computer system to perform tasks without being explicitly programmed by drawing conclusions based on patterns found when analyzing data. AI model 504 is engaged in the study and construction of algorithms (e.g.,Machine learning algorithms (Machine learning algorithms) that can learn from existing data and make predictions about new data. Such algorithms build an AI model from training data to make data-driven predictions or decisions, expressed as outcomes or scores.
[0049] There are two common machine learning (ML) techniques: supervised ML and unsupervised ML. Supervised ML uses prior knowledge (e.g., examples that correlate inputs with outputs or outcomes) to learn the relationships between the inputs and outputs. The goal of supervised ML ( ) is to learn a function that, given training data, best approximates the relationship between the training inputs and outputs so that the ML model can implement the same relationships when given inputs to produce the corresponding outputs. Unsupervised ML is the training of an ML algorithm using information that is neither classified nor labeled, allowing the algorithm to act on that information without guidance. Unsupervised ML is useful in exploratory analysis because of its ability to automatically detect structure in data.
[0050] Common tasks for supervised ML are classification problems and regression problems. Classification problems, also called categorization problems, aim to classify objects into one of several categories (e.g., Is this object an apple or an orange?). Regression algorithms aim to quantify some element (e.g., by assigning a score to the value of an input). Some examples of commonly used supervised ML algorithms are logistic regression (LR), Naive Bayes, random forest (RF), neural networks (NN), deep neural networks (DNN), matrix factorization, and support vector machines (SVM).
[0051] Common unsupervised ML tasks include clustering, representation learning, and density estimation. Some examples of commonly used unsupervised ML algorithms are K-means clustering, principal component analysis, and autoencoders.
[0052] Another type of machine learning is federated learning (also known as collaborative learning), in which an algorithm is trained on multiple decentralized devices using local data without sharing the data. This approach contrasts with traditional centralized machine learning techniques, which upload all local datasets to a server, as well as with more classic decentralized approaches, which often assume that local data samples are identically distributed. Federated learning allows multiple actors to build a common, robust machine learning model without sharing data, thus addressing critical issues such as privacy, data security, data access rights, and access to heterogeneous data.
[0053] In some examples, the AI model may be trained continuously or periodically before performing the inference operation using the output predicted by the AI model inference 506. Then, during the inference operation, the patient-specific input features provided to the AI model may be propagated from an input layer through one or more hidden layers and finally to an output layer.
[0054] Using these techniques, a processing unit, such as the image processing unit 42 of Fig. 2, select a thumbnail for a segment based on an assessment of potential interest, display the selected thumbnail on the user interface, and receive input on the user interface from a user selecting the displayed thumbnail.
[0055] In other examples, the processing unit may determine a predictive indicator for the segment of the video recording based on a potential interest rating, display the predictive indicator on the user interface, and align the predictive indicator with a timeline of the segment.
[0056] Fig. Figure 6 shows a schematic diagram of an example of a trained machine learning model 600. One approach to training data for developing the trained machine learning model 600 is to leverage annotated endoscopy video footage. This training data includes example videos that have been reviewed and annotated by clinical experts to indicate times at which various parameters of interest occur. For example, the training videos are annotated to mark segments that demonstrate the presence of certain disease states, the activation of certain imaging modes, changes in endoscope speed, bowel cleanliness values, foreign body detection, the presence of tools, and the presence of blood. The annotations indicate the start time and duration of each of these events of interest in the videos.
[0057] By training a machine learning model on a dataset containing these expert-annotated videos, the model learns to predict the probability of these various events occurring in new, unannotated endoscopy images. Over time, the trained model outputs a detection score for each parameter based on patterns learned from the annotated training data.
[0058] The image processing unit 42 of Fig. 2 can be used to generate data. For example, the image processing unit 42 generates data for one or more of the following parameters: presence of disease, activation of the light imaging modes, speed of the endoscope, cleanliness of the intestine, presence of a foreign body, presence of tools, and presence of blood. For example, an imaging device 602, such as the imaging and control system 12 of Fig. 1, data such as blood presence data 604, disease data 606, and speed data 608. The data is used to generate N sets of video training data 610, e.g., one or more disease data N 612, speed data N 614, and blood presence data N 616.
[0059] One or more signal processing steps may be performed on the video training data 610, such as sampling, feature extraction, filtering, and the like, before the video training data 610 can be used as training data 618. The training data 618 may include N sets of training data based on the video training data 610. Furthermore, the training data 618 may include annotation training data 620. For example, the annotation training data 620 may include sets of label data N 622 and timestamp data N 624 generated by a medical practitioner 626. The structure of the neural network 628 may include labels associated with one or more parameters. The timestamp data N 624 contains timestamps at which various parameters of interest occur.
[0060] The training data 618 is used to train an AI or machine learning model, such as the trained machine learning model 600, e.g., the AI model 504 of Fig. 5. The training data 618 can be applied to a neural network structure 628, such as a DNN, which includes an input layer, one or more hidden layers, and an output layer. The training data 618 and the annotated training data 620 can be fed into the input layer of the neural network structure 628, which passes the input data or data features through one or more hidden layers to the output layer, which outputs weights and offsets to form the trained machine learning model 600. The trained machine learning model 600 is capable of performing tasks without being explicitly programmed by drawing conclusions based on patterns found during data analysis.
[0061] Fig. 7 shows a flowchart of an example method 700 for navigating images of a segment of a video recording of a medical procedure. At block 702, the method 700 includes selecting a thumbnail for a segment based on a rating of potential interest. At block 704, the method 700 includes displaying the selected thumbnail on the user interface. At block 706, the method 700 includes receiving input at the user interface from a user selecting the displayed thumbnail.
[0062] Fig. 8 shows a flowchart of an example method 800 for navigating images of a segment of a video recording of a medical procedure. At block 802, the method 800 includes determining a predictive indicator for the segment of the video recording based on an assessment of potential interest and a current selection of one or more user-selectable parameters. At block 804, the method 800 includes displaying the predictive indicator on the user interface. At block 806, the method 800 includes aligning the predictive indicator with a timeline of the segment.
[0063] Fig. 9 shows a flowchart of an example method 900 for navigating images of a segment of a video recording of a medical procedure. At block 902, the method 900 includes determining a predictive indicator for the segment of the video recording based on an assessment of potential interest. At block 904, the method 900 includes dynamically adjusting a compression rate of the segment based on the predictive indicator. At block 906, the method 900 includes saving the video recording with the adjusted compression rate.
[0064] Fig.10 shows a block diagram of an example machine 1000 on which one or more of the techniques (e.g., methods) described herein may be performed. The examples described herein may include logic or a series of components or mechanisms in or operated by the machine 1000. Circuits (e.g., processing circuits) are a collection of circuits implemented in tangible units of the machine 1000 and include hardware (e.g., simple circuits, gates, logic, etc.). Circuit membership may be flexible over time. Circuits include elements that, individually or in combination, can perform specific operations when in operation. In one example, the hardware of the circuit may be immutably designed (e.g., hard-wired) to perform a specific operation.In one example, the circuit hardware may comprise variably connected physical components (e.g., execution units, transistors, simple circuits, etc.), including a machine-readable medium that is physically modified (e.g., magnetically, electrically, movable placement of particles with fixed mass, etc.) to encode instructions for the specific operation. Connecting the physical components changes the underlying electrical properties of a hardware component, e.g., from an insulator to a conductor or vice versa. The instructions allow embedded hardware (e.g., the execution units or a loading mechanism) to create elements of the circuit in hardware via the variable connections to perform portions of the specific operation during operation.Accordingly, in one example, the elements of the machine-readable medium are part of the circuit or communicatively coupled to the other components of the circuit when the device is in operation. In one example, each of the physical components may be used in more than one element of more than one circuit. For example, during operation, execution units may be used at one time in a first circuit of a first circuit and reused at another time by a second circuit of the first circuit or by a third circuit of a second circuit. Further examples of these components with respect to machine 1000 follow.
[0065] In alternative examples, machine 1000 may operate as a standalone device or be connected (e.g., networked) to other machines. In a networked deployment, machine 1000 may operate as a server, as a client, or both in a server-client network environment. In one example, machine 1000 may act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. Device 1000 may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a web appliance, a network router, switch, or bridge, or any device capable of executing commands (sequential or otherwise) that specify actions to be performed by that device.Although only a single machine is depicted, the term "machine" also includes any collection of machines that individually or collectively execute a set (or sets) of instructions to perform one or more of the methods discussed herein, e.g., cloud computing, software as a service (SaaS), or other computer cluster configurations.
[0066] The machine 1000 may include a hardware processor 1002 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), main memory 1004, static memory (e.g., memory or storage for firmware, microcode, a BIOS (basic input / output)), and mass storage 1008 (e.g., hard drives, tape drives, flash memory, or other block devices), some or all of which may communicate with each other via an interconnect 530 (e.g., bus). The machine 1000 may further include a display unit 1010, an alphanumeric input device 1012 (e.g., a keyboard), and a user interface (UI) navigation device 1014 (e.g., a mouse). In one example, the display unit 1010, the input device 1012, and the UI navigation device 1014 may be a touchscreen display. The machine 1000 may additionally include a signal generating device 1018 (e.g.,a speaker), a network interface device 1020, and one or more sensors 1016, such as a Global Positioning System (GPS) sensor, a compass, an accelerometer, or other sensor. The device 1000 may include an output controller 1028, such as a serial (e.g., Universal Serial Bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near-field communication (NFC), etc.) connection for communicating with or controlling one or more peripheral devices (e.g., a printer, card reader, etc.).
[0067] The registers of processor 1002, main memory 1004, static memory 1006, or mass storage 1008 may be or include a machine-readable medium 1022 on which one or more sets of data structures or instructions 1024 (e.g., software) are stored that embody or are used by one or more of the techniques or functions described herein. Instructions 1024 may also reside, in whole or in part, in one of the registers of processor 1002, main memory 1004, static memory 1006, or mass storage 1008 while being executed by machine 1000. In one example, one or any combination of hardware processor 1002, main memory 1004, static memory 1006, or mass storage 1008 may constitute machine-readable medium 1022.While the machine-readable medium 1022 is illustrated as a single medium, the term “machine-readable medium” may include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) configured to store the one or more instructions 1024.
[0068] The term "machine-readable medium" may include any medium capable of storing, encoding, or carrying instructions for execution by machine 1000 that cause machine 1000 to perform one or more of the techniques of the present disclosure, or capable of storing, encoding, or carrying data structures used by or associated with such instructions. Non-limiting examples of machine-readable media may include solid-state storage, optical media, magnetic media, and signals (e.g., radio frequency signals, other photon-based signals, audio signals, etc.). In one example, a non-transitory machine-readable medium includes a machine-readable medium having a plurality of particles that have an unchanging (e.g., rest) mass and are thus composed of matter.Accordingly, non-transitory machine-readable media are machine-readable media that do not contain transitory propagation signals. Specific examples of non-transitory machine-readable media include: non-volatile memories such as semiconductor memories (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)) and flash memory; magnetic disks such as internal hard disks and removable disks; magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0069] In one example, the information stored or otherwise provided on the machine-readable medium 1022 may be representative of the instructions 1024, such as the instructions 1024 themselves or a format from which the instructions 1024 may be derived. This format from which the instructions 1024 may be derived may include source code, encrypted instructions (e.g., in compressed or encrypted form), packaged instructions (e.g., divided into multiple packages), or the like. The information representative of the instructions 1024 in the machine-readable medium 1022 may be processed into the instructions by processing circuitry to implement any of the operations discussed herein. Deriving the instructions 1024 from the information (e.g., processing by the processing circuitry) may include, for example, compiling (e.g., from source code, object code, etc.), interpreting, loading, organizing (e.g.,dynamic or static linking), encoding, decoding, encrypting, decrypting, packing, unpacking, or otherwise manipulating the information on the 1024 instructions.
[0070] In one example, deriving the instructions 1024 may include assembling, compiling, or interpreting the information (e.g., by the processing circuitry) to create the instructions 1024 from an intermediate or preprocessing format provided by the machine-readable medium 1022. The information, if present in multiple pieces, may be combined, unpacked, and modified to create the instructions 1024. For example, the information may be present in multiple compressed source code packages (or object code, or binary executable code, etc.) on one or more remote servers. The source code packages may be encrypted when transmitted over a network and decrypted, decompressed, assembled (e.g., linked) if necessary, and compiled or interpreted (e.g., into a library, a standalone executable, etc.) on a local machine.) and run from the local machine.
[0071] The instructions 1024 may further be transmitted or received over a communication network 1026 using a transmission medium via the network interface device 1020 using any transmission protocol (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.). Examples of communication networks can include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), LoRa / LoRaWAN or satellite communication networks, mobile phone networks (e.g., cellular networks (e.g., those conforming to 3G, 4G LTE / LTE-A, or 5G standards), Plain Old Telephone (POTS) networks, and wireless data networks (e.g., Institute of Electrical and Electronics Engineers (IEEE) 1002.11 family of standards known as Wi-Fi®, IEEE 1002.15.4 family of standards, peer-to-peer (P2P) networks, and others.In one example, network interface device 1020 may include one or more physical jacks (e.g., Ethernet, coaxial, or telephone jacks) or one or more antennas for connecting to communications network 1026. In one example, network interface device 1020 may include a plurality of antennas to communicate wirelessly using at least one of SIMO (Single-Input Multiple-Output), MIMO (Multiple-Input Multiple-Output), or MISO (Multiple-Input Single-Output) techniques. The term "transmission medium" encompasses any intangible medium capable of storing, encoding, or transmitting instructions for execution by machine 1000, and includes digital or analog communication signals or other intangible media for facilitating communication of such software. A transmission medium is a machine-readable medium. Various notes
[0072] Each of the non-limiting claims or examples described herein may stand alone or may be combined in various permutations or combinations with one or more of the other examples.
[0073] The above detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings show, by way of illustration, certain embodiments in which the invention may be practiced. These embodiments are also referred to herein as "examples." Such examples may include elements in addition to those shown or described. However, the inventors of the present invention also contemplate examples in which only the elements shown or described are provided. Furthermore, the inventors also contemplate examples in which any combination or permutation of the elements shown or described (or one or more claims thereof) is used, either with respect to a particular example (or one or more claims thereof) or with respect to other examples shown or described (or one or more claims thereof) herein.
[0074] In case of conflicting usages between this document and the documents incorporated by reference, the usage in this document shall prevail.
[0075] In this document, the terms "a" or "an" are used, as is customary in patent documents, to include one or more than one, regardless of other instances or uses of "at least one" or "one or more." In this document, the term "or" is used to refer to a non-exclusive "or," such that "A or B" includes "A but not B," "B but not A," and "A and B," unless otherwise specified. In this document, the terms "including" and "in which" are used as plain English equivalents of the respective terms "comprising" and "wherein." Also in the following claims, the terms "including" and "comprising" are open-ended, i.e.A system, device, article, composition, formulation, or method that includes elements in addition to those listed after such a term in a claim is still within the scope of the claim. Furthermore, in the following claims, the terms "first," "second," "third," etc., are used merely as identifiers and are not intended to impose numerical requirements on their subject matter.
[0076] The method examples described herein may be implemented, at least in part, by machine or computer. Some examples may include a computer-readable medium or a machine-readable medium encoded with instructions that can configure an electronic device to perform the methods described in the above examples. An implementation of such methods may include code, such as microcode, assembly language code, high-level language code, or the like. Such code may include computer-readable instructions for performing various methods. The code may form portions of computer program products. Furthermore, in one example, the code may be stored on one or more transient, non-transitory, or non-transitory tangible computer-readable media, for example, during execution or at other times.Examples of these tangible computer-readable media include, but are not limited to, hard disks, removable magnetic disks, removable optical disks (e.g., compact discs and digital video disks), magnetic cassettes, memory cards or sticks, RAMs (Random Access Memories), ROMs (Read-Only Memories), and the like.
[0077] The above description is illustrative and not restrictive. For example, the examples described above (or one or more claims thereof) may be used in combination with one another. Other embodiments may be used, as will be determined by one of ordinary skill in the art after reviewing the above description. The Abstract is provided in accordance with 37 CFR §1.72(b) to enable the reader to quickly appreciate the nature of the technical disclosure. It is presented with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. The above Detailed Description may also group together various features to simplify the disclosure. This should not be construed to mean that any unclaimed disclosed feature is essential to any claim.Rather, the subject matter of the invention may reside in fewer than all features of a particular disclosed embodiment. Therefore, the following claims are hereby incorporated into the detailed description as examples or embodiments, each claim on its own representing a separate embodiment, and it is intended that these embodiments may be combined with one another in various combinations or permutations. The scope of the invention should be determined by reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] US 63 / 626,778
[0001] WO 2011 / 140118 A1
[0023]
Claims
[1] A system for navigating individual frames of a segment of a video recording of a medical procedure, the system comprising: a user interface with a display; and a processing unit configured to: Selecting a thumbnail for the segment based on an assessment of potential interest; Displaying the selected thumbnail on the user interface; and Receiving input from a user on the user interface that selects the displayed thumbnail. [2] The frame navigation system of claim 1, wherein the potential interest rating is determined based on an analysis of the video recording. [3] A system for navigating in individual frames according to claim 2, wherein the analysis of the video recording comprises a recognition value for one or more of the following parameters: presence of a disease; Activation of light imaging modes; Endoscope speed; intestinal cleanliness; Presence of a foreign body; presence of tools; and Presence of blood. [4] The frame navigation system of claim 1, wherein the video recording comprises a plurality of segments, the processing unit further configured to: Selecting a corresponding thumbnail image for each of the plurality of segments based on an assessment of potential interest for the corresponding segment; display the selected thumbnails on the user interface aligned with their corresponding segment, and where the user interface is set up to: to allow the user to navigate between segments by selecting the displayed thumbnails. [5] The system for navigating image frames according to claim 1, wherein the potential interest score is determined based on a trained machine learning model. [6] The system for navigating frame by frame according to claim 5, wherein the trained machine learning model is trained using past behavior of a clinician. [7] A system for navigating individual frames of a segment of a video recording of a medical procedure, the system comprising: a user interface comprising a display; and a processing unit configured to to determine a predictive indicator for the segment of the video recording based on an assessment of potential interest and a current selection of one or more user-selectable parameters; display the prediction indicator on the user interface; and align the prediction indicator with a timeline of the segment. [8] The frame navigation system of claim 7, wherein the potential interest rating is determined based on an analysis of the video recording. [9] A system for navigating in individual frames according to claim 8, wherein the analysis of the video recording comprises a recognition value for one or more of the following user-selectable parameters: Presence of an illness; Activation of light imaging modes; Speed of the endoscope; Intestinal cleanliness; Presence of tools; Presence of a foreign body; and Presence of blood. [10] A system for navigating in individual images according to claim 9, wherein the processing unit is arranged to: Displaying data representing the user-selectable parameters on the user interface; and Align the data representing the user-selectable parameters with the segment timeline. [11] The frame navigation system of claim 7, wherein the prediction indicator comprises a line with peaks and passes, wherein peaks represent higher levels of assessed potential interest and passes represent lower levels of assessed potential interest. [12] A system for navigating in individual images according to claim 7, wherein the processing unit is arranged to: dynamically adjust a compression rate of the segment based on the prediction indicator; and to save the video recording with the adjusted compression rate. [13] A system for navigating in individual images according to claim 12, wherein the processing unit is arranged to: increase the compression rate for segments that have a prediction indicator smaller than a threshold. [14] The system for navigating image frames according to claim 7, wherein the potential interest score is determined based on a trained machine learning model. [15] The system for navigating frame by frame according to claim 14, wherein the trained machine learning model is trained using past behavior of a clinician. [16] A system for navigating in individual images according to claim 7, wherein the processing unit is arranged to: receive user input that changes the current selection to an updated selection; and dynamically adjust the forecast indicator based on the updated selection. [17] A system for navigating individual frames of a segment of a video recording of a medical procedure, the system comprising: a processing unit configured to: Determining a predictive indicator for the segment of the video recording based on an assessment of potential interest; dynamically adjusting a compression rate of the segment based on the prediction indicator; and Save the video recording with the adjusted compression rate. [18] A system for navigating in individual images according to claim 17, wherein the processing unit is arranged to: increase the compression rate for segments that have a prediction indicator smaller than a threshold. [19] The system for navigating in individual images according to claim 18, wherein the processing unit is configured to increase the compression rate for segments having a prediction indicator that is less than a threshold value, and is configured to increase the compression rate for segments having a prediction indicator that is less than the threshold value after a configurable period of time has elapsed. [20] The frame navigation system of claim 17, wherein the potential interest rating is determined based on an analysis of the video recording. [21] A system for navigating in individual frames according to claim 20, wherein the analysis of the video recording comprises a recognition value for one or more of the following parameters: Presence of an illness; Activation of light imaging modes; Speed of the endoscope; Intestinal cleanliness; Presence of a foreign body; presence of tools; and Presence of blood.
Citation Information
Patent Citations
US-PATENTANMELDUNGSERIALNO.63/626,778
Rotate-to-advance catheterization system
WO2011140118A1