Endoscopic video timeline interest level prediction
By generating intelligent interest prediction indicators and selecting keyframes, the problem of low efficiency in endoscopic video recording processing is solved, and intelligent navigation of rapid positioning of important video segments and key images is achieved, and inspection efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510126581.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2025-01-27
- Publication Date
- 2025-08-01
AI Technical Summary
The processing efficiency of existing endoscopic video recordings is low, making it difficult for doctors to quickly locate important video segments and key images, and lacks intelligent navigation technology.
By generating intelligent interest prediction indicators and selecting keyframes associated with the video recording timeline, analyzing video content using machine learning models, dynamically adjusting compression rates and displaying thumbnails, providing an intelligent navigation system.
Improve the efficiency, accuracy and speed of endoscopic video examination, helping doctors quickly locate important video segments and key images and optimize workflow.
Smart Images

Figure CN120391960A_ABST
Abstract
Description
Technical Field
[0001] This document generally relates to, but is not limited to, medical imaging systems, and more particularly to endoscopic systems. Background Art
[0002] Endoscopy is a medical procedure that allows a doctor to view the inside of the body without making large incisions. The procedure involves the use of an endoscope, a flexible tube with a light and a camera device attached to it. The endoscope can be inserted through a natural opening in the body, such as the mouth, or through a small incision. Endoscopy is used for various purposes, including diagnosing and treating diseases in the gastrointestinal, respiratory, and other organs. It is a minimally invasive method that provides a detailed view of internal organs, enabling more accurate diagnosis and more targeted treatment. Common types of endoscopic procedures include: gastroscopy for examining the upper digestive tract, colonoscopy for the lower intestine, and bronchoscopy for the lungs.
[0003] Endoscopic procedures are very beneficial in modern medicine due to their minimally invasive nature. Compared to traditional surgical procedures, endoscopic procedures result in shorter recovery times and a lower risk of complications. These procedures are generally safe and are performed under local or general anesthesia to ensure patient comfort. Doctors can not only diagnose conditions through the endoscope but also directly perform various treatments through the endoscope, such as biopsies, polyp removals, and even some forms of surgery. Such versatility makes the endoscope an important tool in many medical fields, including gastroenterology, pulmonology, and oncology. Advancements in endoscopic technology are continuously improving its effectiveness, providing high-resolution images and new techniques for treatment and diagnosis. Summary of the Invention
[0004] This disclosure describes various techniques for analyzing video recordings for endoscopic medical imaging and other types of medical imaging. In some examples, the system generates intelligent interest prediction metrics that are displayed in association with a timeline search bar of the video recording. The intelligent interest prediction metrics are quantified and scored along the timeline based on one or more parameters. In other examples, the system intelligently selects key frames from different video segments to be displayed as thumbnails associated with those segments. These techniques can improve efficiency, accuracy, and speed, and using these techniques, doctors can examine and identify important aspects of the recorded endoscopic procedures.
[0005] In some aspects, this disclosure is directed to a system for navigating frames in segments of video recordings of medical procedures, the system including: a user interface that includes a display; and a processing unit configured to: for a segment, select a thumbnail image based on an assessment of potential interest; display the selected thumbnail image on the user interface; and receive an input from the user selecting the displayed thumbnail image on the user interface.
[0006] In some aspects, the present disclosure is directed to a system for navigating frames within segments of a video recording of a medical procedure, the system comprising: a user interface including a display; and a processing unit configured to: determine a prediction metric for a segment of the video recording based on an assessment of potential interest and a current selection of one or more user-selectable parameters; display the prediction metric on the user interface; and align the prediction metric with the timeline of the segment.
[0007] In some aspects, the present disclosure is directed to a system for navigating frames within segments of a video recording of a medical procedure, the system comprising: a user interface including a display; and a processing unit configured to: determine a prediction metric for a segment of the video recording based on an assessment of potential interest; dynamically adjust the compression rate of the segment based on the prediction metric; and store the video recording at the adjusted compression rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In the drawings (which are not necessarily to scale), like reference numerals may describe similar components in different views. Like reference numerals with different alphabetical suffixes may represent different instances of similar components. The drawings generally illustrate, by way of example and not limitation, the various embodiments discussed in this document.
[0009] Figure 1 is a schematic illustration of an example of an endoscopic system that includes an imaging and control system and an endoscope.
[0010] Figure 2 is Figure 1 a schematic illustration of an endoscopic system that includes an endoscope connected to a control unit of an imaging and control system.
[0011] Figure 3 depicts an example of frames of a segment of a video recording of a medical procedure and corresponding thumbnail images within the segment of the video recording.
[0012] Figure 4A shows an example of a graphical display of a prediction metric displayed with various user-selectable parameters.
[0013] Figure 4B shows another example of a graphical display of a prediction metric displayed with various user-selectable parameters.
[0014] Figure 5 is a schematic illustration of an example of a computer-based tissue sample analyzer.
[0015] Figure 6 shows a schematic illustration of an example of a trained machine learning model.
[0016] Figure 7 A flowchart depicting an example of a method for navigating frames in a segment of a video recording of a medical procedure.
[0017] Figure 8 A flowchart depicting an example of a method for navigating frames in a segment of a video recording of a medical procedure.
[0018] Figure 9 A flowchart depicting an example of a method for navigating frames in a segment of a video recording of a medical procedure.
[0019] Figure 10 A block diagram showing an example of a machine on which one or more examples may be implemented. Detailed Description
[0020] Endoscopic procedures are important diagnostic and therapeutic tools in the medical field. Advancements in endoscopic devices and imaging technologies have enabled high-quality video capture during endoscopy. However, the increased volume of video poses challenges for doctors to effectively examine and analyze the recorded segments.
[0021] Manually scrolling through these large video files to identify important parts can be time-consuming and inefficient. Artificial intelligence and machine learning algorithms have shown potential in the automatic analysis of medical images and videos. However, there is a lack of techniques for integrating these algorithms to assist doctors in quickly navigating the recorded endoscopic videos.
[0022] Accordingly, the present inventors have recognized a need for improved graphical user interface systems and methods that utilize the analysis of endoscopic video content to assist doctors in easily locating important video segments and key images. Intelligent analysis combined with a dynamic interface can help doctors' workflows and improve patient care.
[0023] The present disclosure describes various techniques for analyzing video recordings for endoscopic medical imaging and other types of medical imaging. In some examples, the system generates intelligent interest prediction metrics that are displayed in association with a timeline search bar of the video recording. The intelligent interest prediction metrics are quantified and scored along the timeline based on one or more parameters. In other examples, the system intelligently selects key frames from different video segments to be displayed as thumbnails associated with those segments. These techniques can improve efficiency, accuracy, and speed, and using these techniques, doctors can examine and identify important aspects of the recorded endoscopic procedures.
[0024] Figure 1Schematic diagram of an example of an endoscopic system 10, the system including an imaging and control system 12 and an endoscope 14. The endoscopic system 10 is suitable for use with the systems, devices, and methods described below, such as modular endoscopic systems, modular endoscopes, and methods for designing, constructing, and deconstructing endoscopes. According to some examples, the endoscope 14 can be inserted into an anatomical region for imaging and / or providing access for one or more sampling devices for biopsy, or providing access for one or more treatment devices for treating a disease state associated with the anatomical region. In an advantageous aspect, the endoscope 14 can interact with and be connected to the imaging and control system 12. In the example shown, the endoscope 14 includes a colonoscope, however, other types of endoscopes can be used with the features and teachings of the present disclosure.
[0025] The imaging and control system 12 can include: a controller 16, a user interface including an output unit 18 (such as a display) and an input unit 20, a light source 22, a fluid source 24, and a suction pump 26. The imaging and control system 12 can include various ports for coupling with the endoscopic system 10. For example, the controller 16 can include a data input port for receiving data from the endoscope 14 and a data output port for transmitting data to the endoscope 14.
[0026] The light source 22 can include an output port for transmitting light to the endoscope 14, for example, via an optical fiber link. The fluid source 24 can include a port for transmitting fluid to the endoscope 14. The fluid source 24 can include a pump and a fluid tank, or can be connected to an external tank, container, or storage unit. The suction pump 26 can include a port for creating a suction relative to the endoscope 14 to generate suction, for example, for withdrawing fluid from the anatomical region into which the endoscope 14 is inserted. The output unit 18 and the input unit 20 can be used by an operator of the endoscopic system 10 to control the functions of the endoscopic system 10 and view the output of the endoscope 14.
[0027] The controller 16 can additionally be used to generate a signal or other output from processing the anatomical region into which the endoscope 14 is inserted. In some examples, the controller 16 can generate an electrical output, an acoustic output, a fluid output, etc., for processing the anatomical region using, for example, cauterization, cutting, freezing, etc.
[0028] The endoscope 14 can include an insertion section 28, a functional section 30, and a handle section 32, which can be coupled to a cable section 34 and a coupler section 36. The insertion section 28 can extend distally from the handle section 32, and the cable section 34 can extend proximally from the handle section 32. The insertion section 28 can be elongate and include a bending section and a distal end, to which the functional section 30 can be attached. The bending section can be controllable (e.g., via a control knob 38 on the handle section 32) to maneuver the distal end through tortuous anatomical passageways (e.g., the stomach, duodenum, kidney, ureter, etc.). The insertion section 28 can also include one or more working channels (e.g., lumens), which can be elongate and support the insertion of one or more treatment tools of the functional section 30. The working channels can extend between the handle section 32 and the functional section 30. The insertion section 28 can also provide additional functions such as fluid passageways, guide wires, and pull wires (e.g., providing the above-attached functions via a suction or irrigation passageway, etc.).
[0029] The handle module 32 can include a knob 38 and a port 40. The knob 38 can be coupled to a pull wire that extends through the insertion portion 28. The port 40 can be configured to couple various cables, fluid tubes, etc. to the handle module 32 for coupling to the insertion section 28.
[0030] According to some examples, the imaging and control system 12 can be disposed on a mobile platform (e.g., a cart 41) having a shelf for housing a light source 22, a suction pump 26, an image processing unit 42, etc. Alternatively, Figure 1 and Figure 2 certain components of the imaging and control system 12 shown in can be disposed directly on the endoscope 14 to make the endoscope "self - contained".
[0031] Figure 2 is Figure 1 a schematic diagram of an endoscope system 10 that includes an imaging and control system 12 and an endoscope 14. Figure 2 Schematically shown are the components of an imaging and control system 12 coupled to an endoscope 14, which in the illustrated example includes a colonoscope. The imaging and control system 12 can include a controller 16, which can include or be coupled to an image processing unit 42, a treatment generator 44, and a drive unit 46, as well as a light source 22, an input unit 20, and an output unit 18. The image processing unit 42 includes one or more processors that can be distributed, for example, locally and remotely or co - located in one location.
[0032] The image processing unit 42 and the light source 22 can each interact with the endoscope 14 via a wired or wireless connection. The imaging and control system 12 can thus illuminate the anatomical region, collect signals representative of the anatomical region, process the signals representative of the anatomical region, and display an image representative of the anatomical region on the display unit 18. The imaging and control system 12 can include a light source 22 to illuminate the anatomical region with light of a desired spectrum (e.g., broadband white light, narrow band imaging using preferred electromagnetic wavelengths, etc.). The imaging and control system 12 can be connected (e.g., via an endoscope connector) to the endoscope 14 for signal transmission (e.g., light output from the light source, video signal from the imaging system at the distal end, etc.).
[0033] The fluid source 24 can include one or more sources of air, saline, or other fluids, as well as associated fluid paths (e.g., air channels, irrigation channels, suction channels) and connectors (barb fittings, fluid seals, valves, etc.). The imaging and control system 12 can also include a drive unit 46, which can be an optional component. The drive unit 46 can include a motorized driver for advancing the distal segment of the endoscope 14, as described in Frassica et al. in PCT Publication No. WO 2011 / 140118 A1, titled "Rotate-to-Advance Catheterization System," which is incorporated herein by reference in its entirety.
[0034] Figure 3 An example of an image 300 of a segment of a video recording of a medical procedure is depicted. As described above, the present disclosure describes a system that intelligently selects key frames from different video segments for display as thumbnails associated with those segments. These techniques can improve efficiency, accuracy, and speed, and using these techniques, a doctor can examine and identify important aspects of a recorded endoscopy procedure.
[0035] The processing unit of the system, for example Figure 2 the image processing unit 42 of the endoscope system 10 or another processing unit not associated with the endoscope system 10, can divide the video recording of the medical procedure into multiple segments, for example, into eight segments in a non-limiting example. Then, the processing unit can select a thumbnail image for each segment based on an assessment of potential interest.
[0036] In some examples, an assessment of potential interest is determined based on an analysis of a video record. As an example, the analysis of the video record includes detection scores for various parameters. These parameters can include: disease detection (e.g., CADe (Computer-Aided Detection) and CADx (Computer-Aided Diagnosis) outputs) and potential associated confidence levels; activation of light imaging modes at different stages (e.g., white light, narrow band imaging (NBI)); endoscope withdrawal speed; bowel cleanliness score (which may be determined by mucosal detection rate); detection of tools in the image; detection of blood in the image; and detection of foreign objects in the image. In this way, the processing unit can intelligently select key frames from different video segments of the video record of the medical procedure.
[0037] CADe involves algorithms or systems that assist radiologists or other medical professionals by highlighting regions of interest in medical images. These regions can indicate the presence of anomalies such as tumors, fractures, or other pathological changes. CADe systems do not make a diagnosis; they simply draw attention to areas that may require further analysis. CADx systems go a step further and can not only detect anomalies but also provide an interpretation or a reasonable diagnosis. These systems analyze medical images based on the visual patterns they detect and provide diagnostic suggestions. CADx systems can be used to assist in the decision-making process by providing additional information or a second opinion to a radiologist or a doctor.
[0038] For each segment, the processing unit selects a corresponding thumbnail image based on an assessment of the potential interest of each respective segment, which is shown in the thumbnail strip 302 such as thumbnails 304a to 304h. The system displays the selected thumbnails 304a to 304h, for example, on a user interface (e.g., Figure 1 the display 18) or on another display not associated with the endoscope system 10 (e.g., a personal computing device or a tablet computing device). In some examples, the selected thumbnail images are displayed along the timeline 306 of the video record. In this way, the clinician knows the position in the video record of the medical procedure where the selected thumbnail image appears.
[0039] Next, the system receives, via a user interface such as Figure 1 the input unit 20, an input from the user selecting one of the displayed thumbnail images. For example, in Figure 3 , the user has selected the displayed thumbnail image 304e, and the thumbnail image 304e is recognized by the processing unit as a key frame and is displayed above as a larger image 300. If the user selects another key frame selected by the processing unit from the segment, such as the thumbnail image 304b, then the thumbnail image 304b will be displayed as the large image 300 above, enabling the user to navigate between segments by selecting the displayed thumbnail images.
[0040] As described above, in some examples, the system generates an intelligent interest prediction metric 308. The processing unit of the system, such as Figure 2 the image processing unit 42 of the endoscope system 10, can determine the prediction metric 308 for a segment of the video recording based on an assessment of the latent interest.
[0041] In some examples, the assessment of the latent interest is determined based on an analysis of the video recording. The intelligent interest prediction metric is quantified and scored over time based on one or more parameters. As an example, the analysis of the video recording includes detection scores for various parameters. As described above, in some examples, the analysis of the video recording includes detection scores for one or more of the following parameters: presence of a disease; activation of an optical imaging mode; probe speed; bowel cleanliness; presence of a tool; presence of a foreign object; and presence of blood. In some examples, as Figure 4A shown in detail, the parameters are selectable by the user.
[0042] Then, the processing unit can display the prediction metric 308 on a user interface such as Figure 1 the display 18. The prediction metric 308 includes a line 310 with peaks and valleys, where the peaks represent higher levels of the assessed latent interest and the valleys represent lower levels of the assessed latent interest. The processing unit can align the prediction metric 308 with the timeline 306 of the video recording.
[0043] In addition, a thumbnail image can be aligned with the prediction metric 308. For example, the thumbnail 304e is shown aligned with the highest peak of the prediction metric 308. A user, such as a clinician, after seeing the prediction metric 308, can select the thumbnail image 304e to further view the segment of the video recording.
[0044] In some examples, the processing unit is configured to: dynamically adjust the compression rate of a segment based on the prediction metric and store the video recording at the adjusted compression rate. For example, it may be desirable to increase the compression rate for segments with a prediction metric below a threshold, which can allow the system to store more data when the segment is evaluated as less interesting. In some examples, after a configurable time period has passed, the system can increase the compression rate for segments with a prediction metric below a threshold. Similarly, it may be desirable to decrease the compression rate for segments with a prediction metric above a threshold, which can allow the system to improve the quality of the stored video recording by storing more data.
[0045] In some examples, the assessment of the latent interest for intelligent keyframe selection and / or the intelligent interest prediction metric is determined based on a trained machine learning model, such as described below with respect to Figure 5 and Figure 6Description. In some such examples, the past behavior of a clinician is used to train a trained machine learning model.
[0046] It should be noted that in other examples, the analysis of the video recording does not need to be performed by the endoscopic system 10. Instead, the video recording can be stored on a mass storage device external to the endoscopic system 10, such as stored on a hospital's central server or on a remotely located computing device (e.g., a cloud-based computing device). Then, the techniques of the present disclosure can be used, and the results can be displayed on a user interface not associated with the endoscopic system, such as a tablet computer, a laptop computer, a desktop computer, or some other device with a display.
[0047] Figure 4A An example of a graphical display 400 of a prediction metric 308 displayed together with various parameters 402 that a user can select is shown. The processing unit can generate the graphical display 400 on a display 18 of a user interface, such as Figure 1 the display 18.
[0048] The parameters used by the processing unit to determine the evaluation of potential interest can be selected by the user to dynamically update the intelligent interest prediction metric score based on only a subset of the currently selected available parameters. In this way, if a user invokes a blood test and expects to view the associated video, the user can deselect all parameters other than the blood test. This can immediately draw the user's attention to a specific part of the recording showing the blood.
[0049] In Figure 4A the example shown, the parameters 402 that a user can select are shown, and the data representing the parameters that a user can select is aligned with the timeline of the segment. For example, the data 404 of the activated light mode 406 is aligned with the video timeline 408.
[0050] Figure 4B Another example of a graphical display of a prediction metric displayed together with various parameters that a user can select is shown. Figure 4B Includes features similar to those shown and described above with respect to Figure 4A and uses similar reference numerals for such features. For the sake of brevity, these features are not described in detail again.
[0051] In Figure 4A the example shown, all parameters 402 that a user can select are selected. In contrast, Figure 4BThe graphical display 410 shown in [figure] only selects a subset of the parameters 402 that the user can select, namely the disease detection parameter and the speed parameter. Additional or alternative parameters can be selected in other examples. Selecting or deselecting one or more of the parameters 402 that the user can select can dynamically adjust the prediction metric 308 to indicate the location where the processing unit (e.g., Figure 2 the image processing unit 42) of [device] identifies anomalies, so that the doctor viewing the video can easily find these times. For example, the dynamic adjustment of the prediction metric 308 can cause the peaks and valleys of the line 310 to move, indicating a change in the evaluated potential interest level associated with a particular thumbnail image of the thumbnail strip 302. In this way, the processing unit determines the prediction metric for a segment of the video recording based on the evaluation of potential interest and the current selection of one or more parameters that the user can select. When receiving user input that changes the current selection to an updated selection; the processing unit dynamically adjusts the prediction metric based on the updated selection.
[0052] Figure 5 A schematic diagram showing an example of a computer-based potential interest assessment analyzer is shown. In addition, the computer-based potential interest assessment analyzer 500 is configured to: for one or more segments, based on the evaluation of potential interest, such as based on the analysis of the video recording, select the corresponding thumbnail images.
[0053] In some examples, the computer-based potential interest assessment analyzer 500 analyzes the video recording and generates detection scores for various parameters. In some examples, the computer-based potential interest assessment analyzer 500 may include an input interface 502, through which various parameters are provided as input features to a trained machine learning (ML) model or an artificial intelligence (AI) model, such as the trained AI model 504. One or more relevant input parameters 510 that can be extracted from various component outputs 512 are applied to the AI model to generate an output 506 for inferring predictions from the AI model. The relevant input parameters 510 may include, but are not limited to, one or more of the following parameters: the presence of a disease; the activation of an optical imaging mode; the probe speed; the bowel cleanliness; the presence of foreign objects; the presence of tools; and the presence of blood.
[0054] For example, one or more relevant input parameters 510 that can be extracted from the sensor data generated by the Figure 1 component outputs 512 of the various components of the endoscopic system 10 can be applied to the trained AI model 504, and the AI model 504 can generate confidence scores 514 for various parameters. Then, the artificial intelligence model 504 selects thumbnail images for display on the user interface based on the evaluation of potential interest for the segment.
[0055] In other examples, the computer-based potential interest assessment analyzer 500 determines a prediction metric for a segment of a video recording based on an assessment of potential interest, such as Figure 3 the prediction metric 308. The processing unit, such as Figure 2 the image processing unit 42, displays the prediction metric on a user interface and aligns the prediction indicator with the timeline of the segment, as shown, for example, in Figure 4A what is shown.
[0056] In some embodiments, the input interface 502 can be a direct data link between the computer-based potential interest assessment analyzer 500 and one or more medical devices (e.g., [[ID=4 the endoscope system 10), which generates at least some of the input parameters. Additionally or alternatively, the input interface 502 can be a classical user interface that facilitates interaction between the user and the computer-based potential interest assessment analyzer 500. For example, the input interface 502 can enhance the user interface through which the user can manually input information.
[0057] Based on one or more input parameters, the AI model 504 performs an inference operation using the AI model 504 to generate an assessment of potential interest and selects a corresponding thumbnail image for the segment of the video recording based on the assessment of potential interest. For example, the input interface 502 can pass the input parameters to the input layer of the AI model 504, and the input layer of the AI model 504 propagates these input parameters through the AI model 504 to the output layer. The AI model 504 can provide the computer system with the ability to perform tasks by making inferences based on patterns found in data analysis without being explicitly programmed. The AI model 504 explores the research and construction of algorithms (e.g., machine learning algorithms) that can learn from existing data and make predictions about new data. Such algorithms operate by building an AI model from example training data to make data-driven predictions or decisions expressed as outputs or assessments.
[0058] There are two common paradigms for machine learning (ML): supervised ML and unsupervised ML. Supervised ML uses prior knowledge (e.g., examples that relate inputs to outputs or results) to learn the relationship between inputs and outputs. The goal of supervised ML is to learn a function that, given some training data, best approximates the relationship between the training inputs and outputs so that the ML model can achieve the same relationship given an input to generate a corresponding output. Unsupervised ML is training an ML algorithm using information that is neither classified nor labeled and allowing the algorithm to operate on that information without guidance. Unsupervised ML is useful in exploratory analysis because it can automatically identify structures in the data.
[0059] Common tasks for supervised ML are classification problems and regression problems. Classification problems - also known as categorization problems - aim to classify items into one of several categorical values (e.g., is the object an apple or an orange?). Regression algorithms aim to quantify some items (e.g., by providing a score for some input values). Some examples of commonly used supervised ML algorithms are logistic regression (LR), Naive Bayes, random forest (RF), neural network (NN), deep neural network (DNN), matrix factorization, and support vector machine (SVM).
[0060] Common tasks for unsupervised ML include clustering, representation learning, and density estimation. Some examples of commonly used unsupervised ML algorithms are K-means clustering, principal component analysis, and autoencoders.
[0061] Another type of ML is federated learning (also known as collaborative learning), which trains algorithms on multiple decentralized devices that hold local data without exchanging data. This approach contrasts with traditional centralized machine learning techniques where all local datasets are uploaded to a server, and also with more classical decentralized methods that typically assume local data samples are identically distributed. Federated learning enables multiple participants to build a common, robust machine learning model without sharing data, thus allowing it to address key issues such as data privacy, data security, data access rights, and access to heterogeneous data.
[0062] In some examples, an AI model can be continuously or periodically trained by inferring a predicted output 506 from the AI model before performing an inference operation. Then, during the inference operation, patient-specific input features provided to the AI model can propagate from the input layer, through one or more hidden layers, and ultimately to the output layer.
[0063] By using these techniques, a processing unit such as the image processing unit 42, can select a thumbnail image for a segment based on an assessment of potential interest, display the selected thumbnail image on a user interface, and receive an input from the user selecting the displayed thumbnail image on the user interface.
[0064] In other examples, the processing unit can determine a prediction metric for a segment of a video recording based on an assessment of potential interest, display the prediction metric on a user interface, and align the prediction metric with the timeline of the segment.
[0065] A schematic diagram showing an example of a trained machine learning model 600 is presented. One method for training data to develop the trained machine learning model 600 is to utilize annotated endoscopic video segments. The training data contains sample videos that are examined and annotated by clinical experts to indicate the timestamps at which various parameters of interest occur. For example, the training videos are annotated to mark segments that show the presence of a specific disease state, the activation of certain imaging modalities, changes in probe speed, bowel cleanliness scores, detection of foreign objects, the presence of tools, and the presence of blood. The annotations indicate the start time and duration of each of these events of interest within the video.
[0066] By training a machine learning model on a dataset of videos containing these expert annotations, the model learns to predict the likelihood of these different events occurring within new, unlabeled endoscopic segments. The trained model outputs a detection score that varies over time for each parameter, based on the patterns learned from the annotated training data.
[0067] The image processing unit 42 can be used to generate data. For example, the image processing unit 42 generates data for one or more of the following parameters: the presence of a disease; the activation of a light imaging modality; probe speed; bowel cleanliness; the presence of foreign objects; the presence of tools; and the presence of blood. For example, the imaging and control system 12 of the imaging device 602 generates data, such as blood presence data 604, disease data 606, and speed data 608. This data is used to generate N sets of video training data 610, such as one or more of disease data N 612, probe speed data N 614, and blood presence data N 616. Before the video training data 610 is ready to be used as training data 618, one or more signal processing steps, such as sampling, feature extraction, filtering, etc., can be performed on the video training data 610. The training data 618 can include N sets of training data based on the video training data 610. Additionally, the training data 618 can include annotated training data 620. For example, the annotated training data 620 can include multiple sets of labeled data N 622 and multiple sets of timestamp data N 624 generated by medical practitioners 626. The neural network structure 628 can include labels associated with one or more parameters. The timestamp data N 624 includes the timestamps at which various parameters of interest occur.
[0068] The training data 618 is used to train an AI model or a machine learning model, such as the trained machine learning model 600, for example
[0069] The training data 618 is used to train an AI model or a machine learning model, such as the trained machine learning model 600, for example The AI model 504. The training data 618 can be applied to a neural network structure 628 including an input layer, one or more hidden layers, and an output layer, such as a DNN. The training data 618 and the annotated training data 620 can be fed into the input layer of the neural network structure 628, and the input layer propagates the input data or data features through one or more hidden layers to the output layer of output weights and biases to form the trained machine learning model 600. The trained machine learning model 600 can perform tasks by making inferences based on patterns found in data analysis without being explicitly programmed.
[0070] A flowchart depicting an example of a method 700 for navigating frames in a segment of a video recording of a medical procedure. At block 702, the method 700 includes selecting a thumbnail image for the segment based on an assessment of potential interest. At block 704, the method 700 includes displaying the selected thumbnail image on a user interface. At block 706, the method 700 includes receiving, on the user interface, an input from the user selecting the displayed thumbnail image.
[0071] A flowchart depicting an example of a method 800 for navigating frames in a segment of a video recording of a medical procedure. At block 802, the method 800 includes determining a prediction metric for the segment of the video recording based on an assessment of potential interest and a current selection of one or more parameters that a user can select. At block 804, the method 800 includes displaying the prediction metric on a user interface. At block 806, the method 800 includes aligning the prediction metric with the timeline of the segment.
[0072] A flowchart depicting an example of a method 900 for navigating frames in a segment of a video recording of a medical procedure. At block 902, the method 900 includes determining a prediction metric for the segment of the video recording based on an assessment of potential interest. At block 904, the method 900 includes dynamically adjusting the compression rate of the segment based on the prediction metric. At block 906, the method 900 includes storing the video recording at the adjusted compression rate.
[0073] A block diagram of an exemplary machine 1000 is shown, on which any one or more of the techniques (e.g., methods) discussed herein may be performed. As described herein, an example may include logic or multiple components or mechanisms in the machine 1000, or may be operated by logic or multiple components or mechanisms in the machine 2100. A circuit system (e.g., a processing circuit system) is a collection of circuits implemented in a tangible entity of the machine 1000 that includes hardware (e.g., simple circuits, gates, logic, etc.). The circuit system component relationships may be flexible over time. A circuit system includes components that can perform specified operations individually or in combination when operating. In an example, the hardware of the circuit system may be invariantly designed to perform a specific operation (e.g., hardwired). In an example, the hardware of the circuit system may include physically variable components (e.g., execution units, transistors, simple circuits, etc.) to encode instructions for a specific operation, and the physically variable components include machine-readable media that are physically modified (e.g., magnetically, electrically, movable placement of invariantly aggregated particles, etc.). When connecting physical components, the potential electrical characteristics of the hardware components are changed, e.g., from an insulator to a conductor or from a conductor to an insulator. The instructions enable the embedded hardware (e.g., an execution unit or a loading mechanism) to create components of the circuit system in the hardware via variable connections to perform parts of a specific operation when operating. Thus, in an example, the machine-readable media element is part of the circuit system or communicatively coupled to other components of the circuit system when the device is operating. In an example, any physical component may be used in more than one component of more than one circuit system. For example, in operation, an execution unit may be used in a first circuit of a first circuit system at one point in time and reused by a second circuit in the first circuit system, or reused by a third circuit in a second circuit system at a different time. Additional examples of these components of the machine 1000 are provided below.
[0074] In an alternative example, machine 1000 can operate as a stand-alone device or can be connected (e.g., networked) to other machines. In a networked deployment, machine 1000 can operate in a server-client network environment as a server machine, a client machine, or both. In an example, machine 1000 can act as a peer machine in a peer-to-peer (P2P) (or other distributed) network environment. Machine 1000 can be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, web device, network router, switch or bridge, or any machine capable of executing instructions (sequentially or otherwise) that specify actions to be taken by that machine. Further, although only a single machine is shown, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein, such as cloud computing, software as a service (SaaS), other computer cluster configurations.
[0075] Machine 1000 can include a hardware processor 1002 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 1004, a static memory (e.g., for firmware, microcode, basic input / output (BIOS)), and a mass storage device 1008 (e.g., a hard disk drive, a tape drive, a flash memory, or other block device), some or all of which can communicate with each other via an interconnection (e.g., a bus) 530. Machine 1000 can also include a display unit 1010, an alphanumeric input device 1012 (e.g., a keyboard), and a user interface (UI) navigation device 1014 (e.g., a mouse). In an example, the display unit 1010, the input device 1012, and the UI navigation device 1014 can be a touch screen display. Machine 1000 can additionally include a signal generation device 1018 (e.g., a speaker), a network interface device 1020, and one or more sensors 1016, such as a global positioning system (GPS) sensor, a compass, an accelerometer, or other sensors. Machine 1000 can include an output controller 1028, such as a serial (e.g., universal serial bus (USB)), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection to communicate with or control one or more peripheral devices (e.g., a printer, a card reader, etc.).
[0076] Registers of the processor 1002, main memory 1004, static memory 1006, or mass storage device 1008 may be, or may include, a machine-readable medium 1022 having stored thereon a set or more sets of data structures or instructions 1024 (e.g., software) that embody any one or more of the techniques or functions described herein or are utilized by any one or more of the techniques or functions described herein. During execution of instructions 1024 by the machine 1000, the instructions 1024 may also reside, completely or at least partially, in any register of the processor 1002, main memory 1004, static memory 1006, or mass storage device 1008. In an example, one or any combination of the hardware processor 1002, main memory 1004, static memory 1006, or mass storage device 1008 may constitute a machine-readable medium 1022. Although the machine-readable medium 1022 is shown as a single medium, the term "machine-readable medium" may include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) configured to store one or more instructions 1024.
[0077] The term "machine-readable medium" may include any medium that can store, encode, or carry instructions for execution by the machine 1000 and that cause the machine 1000 to perform any one or more of the techniques of the present disclosure or that can store, encode, or carry data structures used by or associated with such instructions. Non-limiting examples of machine-readable media may include solid-state memory, optical media, magnetic media, and signals (e.g., radio frequency signals, other photon-based signals, sound signals, etc.). In an example, a non-transitory machine-readable medium includes a machine-readable medium having a plurality of particles with invariant (e.g., stationary) mass and is thus a composition of matter. Thus, a non-transitory machine-readable medium is a machine-readable medium that does not include transitory propagated signals. Specific examples of non-transitory machine-readable media may include: non-volatile memory such as semiconductor storage devices (e.g., electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM)), and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[0078] In an example, the information stored or otherwise provided on the machine-readable medium 1022 can represent instructions 1024, such as the instructions 1024 themselves or a format from which the instructions 1024 can be derived. Such a format from which the instructions 1024 can be derived can include source code, encoded instructions (e.g., in a compressed or encrypted form), packaged instructions (e.g., split into multiple packages), and the like. The information representing the instructions 1024 in the machine-readable medium 1022 can be processed by processing circuitry into instructions to implement any of the operations discussed herein. For example, deriving the instructions 1024 from the information (e.g., by processing circuitry) can include compiling, interpreting, loading, organizing (e.g., dynamically or statically linking), encoding, decoding, encrypting, decrypting, packaging, unpackaging, or otherwise manipulating the information into the instructions 1024.
[0079] In an example, the derivation of the instructions 1024 can include the assembly, compilation, or interpretation of the information (e.g., by processing circuitry) to create the instructions 1024 according to some intermediate or preprocessed format provided by the machine-readable medium 1022. The information can be combined, unpacked, and modified when provided in multiple parts to create the instructions 1024. For example, the information can be in multiple compressed source code packages (or object code, or binary executable code, etc.) on one or several remote servers. The source code packages can be encrypted when transmitted over a network and, if necessary, decrypted, decompressed, assembled (e.g., linked), and compiled or interpreted (e.g., compiled or interpreted into libraries, stand-alone executables, etc.) at the local machine and executed by the local machine.
[0080] Any of a plurality of transport protocols (e.g., Frame Relay, Internet Protocol (IP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), etc.) can also be used via the network interface device 1020 to transmit or receive the instructions 1024 over a communication network 1026 using a transmission medium. Example communication networks can include local area networks (LANs), wide area networks (WANs), packet data networks (e.g., the Internet), LoRa / LoRaWAN or satellite communication networks, mobile telephone networks (e.g., cellular networks, e.g., cellular networks compliant with 3G, 4G LTE / LTE-A, or 5G standards), plain old telephone (POTS) networks, and wireless data networks (e.g., referred to as IEEE 702.11 standard family, IEEE 702.15.4 standard family, peer-to-peer (P2P) networks, etc. The network interface device 1020 may include one or more physical jacks (e.g., Ethernet, coaxial, or telephone jacks) or one or more antennas to connect to the communication network 1026. The network interface device 1020 may include multiple antennas to perform wireless communication using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multiple-input single-output (MISO) techniques. The term "transmission medium" should be considered to include any medium that can store, encode, or carry instructions for execution by the machine 1000, and includes digital or analog communication signals or other media to facilitate the communication of such software. The transmission medium is a machine-readable medium.
[0081] Various notes
[0082] Each non-limiting claim or each example described herein may exist independently, or may be combined with one or more other examples in various permutations or combinations.
[0083] The detailed description above includes references to the drawings, which form a part of the detailed description. The drawings illustrate, by way of example, specific embodiments in which the invention may be practiced. Such embodiments are also referred to herein as "examples". Such examples may include elements in addition to those shown or described. However, the inventors also contemplate examples in which only those elements shown or described are provided. In addition, the inventors also contemplate examples (or one or more claims thereof) using any combination or arrangement of those elements shown or described with respect to a particular example (or one or more claims thereof) herein or with respect to other examples (or one or more claims thereof).
[0084] In the event of any inconsistency in usage between this document and any document incorporated by reference, the usage in this document shall prevail.
[0085] In this document, as is common in patent documents, the term "a" or "an" is used to include one or more than one, independent of any other instances or uses of "at least one" or "one or more". In this document, unless otherwise specified, the term "or" is used to refer to a non-exclusive or, such that "A or B" includes "A but not B", "B but not A", and "A and B". In this document, the terms "including" and "in which" are used as the plain English equivalents of the respective terms "comprising" and "wherein". Further, in the appended claims, the terms "including" and "comprising" are open-ended, i.e., a system, apparatus, article, composition, formulation, or process that includes elements other than those listed after such a term in a claim is still considered to fall within the scope of that claim. Additionally, in the appended claims, the terms "first", "second", "third", etc. are used only as labels and are not intended to impose numerical requirements on their objects.
[0086] The method examples described herein can be implemented, at least in part, by a machine or a computer. Some examples can include a computer-readable medium or a machine-readable medium encoded with instructions that can operate to configure an electronic device to perform the methods described in the above examples. Implementations of such methods can include code, such as microcode, assembly language code, high-level language code, etc. Such code can include computer-readable instructions for performing various methods. The code can form part of a computer program product. Further, in an example, the code can be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media, during execution or at other times. Examples of such tangible computer-readable media can include, but are not limited to, hard disks, removable disks, removable optical disks (e.g., compact disks and digital video disks), magnetic tape cartridges, memory cards or memory sticks, random access memory (RAM), read only memory (ROM), etc.
[0087] The foregoing description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of their claims) may be used in combination with each other. For instance, other embodiments may be utilized by those of ordinary skill in the art after reading the above description. The abstract is provided to comply with 37 C.F.R. § 1.72(b) to allow the reader to quickly ascertain the nature of the technical disclosure. The abstract is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the above detailed description, various features may be grouped together to streamline the disclosure. This should not be construed as intending that the disclosed features not claimed are essential to any claim. Rather, the inventive subject matter may lie in less than all of the features of a particular disclosed embodiment. Thus, the appended claims are hereby incorporated as examples or embodiments into the detailed description, where each claim stands on its own as a separate embodiment, and it is contemplated that such embodiments may be combined with each other in various combinations or permutations. The scope of the invention should be determined with reference to the appended claims along with the full scope of equivalents to which such claims are entitled.
Claims
1. A system for navigating frames in a segment of a video recording of a medical procedure, the system comprising: A user interface including a display; and A processing unit configured to: Select thumbnail images for the segment based on an assessment of potential interest; Display the selected thumbnail images on the user interface; and Receive, on the user interface, an input from a user selecting one of the displayed thumbnail images.
2. The system for navigating a frame according to claim 1, wherein, Determine the assessment of potential interest based on an analysis of the video recording.
3. The system for navigating frames according to claim 2, wherein, The analysis of the video recording includes detection scores for one or more of the following parameters: Presence of a disease; Activation of a light imaging mode; Probe speed; Bowel cleanliness; Presence of foreign objects; Presence of tools; and Presence of blood.
4. The system for navigating frames according to claim 1, wherein, The video recording includes a plurality of segments, wherein the processing unit is further configured to: For each of the plurality of segments, select a corresponding thumbnail image based on an assessment of potential interest for the corresponding segment; Display the selected thumbnail images aligned with their corresponding segments on the user interface, and Wherein the user interface is configured to: Enable the user to navigate between segments by selecting one of the displayed thumbnail images.
5. The system for navigating a frame according to claim 1, wherein, Determine the assessment of potential interest based on a trained machine learning model.
6. The system for navigating a frame according to claim 5, wherein, Train the trained machine learning model using past behavior of clinicians.
7. A system for navigating frames in a segment of a video recording of a medical procedure, the system comprising: A user interface including a display; and A processing unit configured to: Determine a prediction metric for the segment of the video recording based on an assessment of potential interest and a current selection of one or more parameters that a user can select; Display the prediction metric on the user interface; and Align the prediction metric with the timeline of the segment.
8. The system for navigating a frame according to claim 7, wherein, Determine the assessment of potential interest based on an analysis of the video recording.
9. The system for navigating frames according to claim 8, wherein, The analysis of the video recording includes detection scores for one or more of the following parameters that a user can select: Presence of a disease; Activation of a light imaging mode; Probe speed; Bowel cleanliness; Presence of tools; Presence of foreign objects; And Presence of blood.
10. The system for navigating frames according to claim 9, wherein, The processing unit is configured to: Display data representing the parameters that a user can select on the user interface; and Align the data representing the parameters that a user can select with the timeline of the segment.
11. The system for navigating a frame according to claim 7, wherein, The prediction metric includes a line having peaks and valleys, where a peak represents a higher level of assessed potential interest and a valley represents a lower level of assessed potential interest.
12. The system for navigating a frame according to claim 7, wherein, The processing unit is configured to: Dynamically adjust the compression rate of the segment based on the prediction metric; and Store the video recording at the adjusted compression rate.
13. The system for navigating a frame according to claim 12, wherein, The processing unit is configured to: Increase the compression rate for segments having a prediction metric less than a threshold.
14. The system for navigating frames according to claim 7, wherein, Determine the assessment of potential interest based on a trained machine learning model.
15. The system for navigating a frame according to claim 14, wherein, Train the trained machine learning model using past behavior of clinicians.
16. The system for navigating frames according to claim 7, wherein, The processing unit is configured to: Receive user input to change the current selection to an updated selection; and Dynamically adjust the prediction metric based on the updated selection.
17. A system for navigating frames in a segment of a video recording of a medical procedure, the system comprising: A processing unit configured to: Determine a prediction metric for the segment of the video recording based on an assessment of potential interest; Dynamically adjust the compression rate of the segment based on the prediction metric; and Store the video recording at the adjusted compression rate.
18. The system for navigating frames according to claim 17, wherein, The processing unit is configured to: Increase the compression rate for segments having a prediction metric less than a threshold.
19. The system for navigating frames according to claim 18, wherein, The processing unit configured to increase the compression rate for segments having a prediction metric less than a threshold is configured to: Increase the compression rate for segments having a prediction metric less than the threshold after a configurable period of time has elapsed.
20. The system for navigating frames according to claim 17, wherein, Determine the assessment of potential interest based on an analysis of the video recording.
21. The system for navigating a frame according to claim 20, wherein, The analysis of the video recording includes detection scores for one or more of the following parameters: Presence of a disease; Activation of a light imaging mode; Probe speed; Bowel cleanliness; Presence of a foreign object; Presence of a tool; and Presence of blood.
Citation Information
Patent Citations
Rotate-to-advance catheterization system
WO2011140118A1
Cited By
Medical video content semantic segmentation labeling method fused with knowledge graph
CN122416455A
A medical video content semantic segmentation labeling method fusing a knowledge graph
CN122416455B