Hierarchical segmentation of surgical scenarios
Through the hierarchical segmentation method and machine learning model, the fuzzy segmentation problem of anatomical structures in surgical video frames is solved, the segmentation accuracy of key structures such as gallbladder artery and gallbladder duct is improved, and the safety of surgery is enhanced.
Patent Information
- Application Number
- CN202380078695.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-13
- Filing Date
- 2023-11-14
- Publication Date
- 2025-07-08
AI Technical Summary
The segmentation of surgical video frames is blurred due to factors such as similar anatomical structures, blood, visceral fat and smoke, which makes it difficult to predict and segment the anatomical category.
Using a hierarchical segmentation method, by generating multi-label probability maps and hierarchies, image pixels are processed to generate leaf-level segmentation maps, and class labels are updated to higher parent classes until a predetermined confidence threshold is reached, and machine learning models are trained and segmented.
Improve the accuracy of segmentation and detection of anatomical structures, reduce inter-class confusion, and improve the segmentation effect of surgical scenarios, especially the identification of gallbladder artery and gallbladder duct.
Smart Images

Figure CN120283270A_ABST
Abstract
Description
Background Art
[0001] The present disclosure generally relates to computing technologies, and more particularly to computing technologies for hierarchical segmentation of video frames in surgical videos.
[0002] Computer-aided systems, particularly computer-aided surgical systems (CAS), rely on video data digitally captured during surgery. Such video data can be stored and / or streamed. In some cases, the video data can be used to enhance a person's physical senses, perception, and reaction capabilities. For example, such systems can effectively provide information corresponding to a field of view that is extended in time and space, which enables a person to adjust current and future actions based on portions of the environment that are not included in his or her physical field of view. Alternatively or additionally, the video data can be stored and / or transmitted for several purposes, such as archiving, training, postoperative analysis, and / or patient consultation.
[0003] Segmentation of surgical scenes can provide valuable information for real-time guidance and postoperative analysis of robot-assisted laparoscopy. However, unfortunately, due to ambiguities caused by similar appearances of anatomical structures, occlusions caused by blood, visceral fat, and / or smoke, and weakened anatomical references due to camera pose, segmentation of surgical video frames is challenging. This can lead to missed detections or incorrect predictions of anatomical categories. Summary of the Invention
[0004] According to one aspect, a computer-implemented method for hierarchical segmentation of video frames in a surgical video is provided. The method includes: obtaining an image of an anatomical structure, wherein the image includes a plurality of image pixels; generating a multi-label probability map for each node in a predefined segmentation class hierarchy; processing the plurality of image pixels to generate a leaf-level segmentation map; and processing each leaf-level segmentation to update the class label of each leaf-level segmentation to a higher parent class until a predetermined prediction confidence threshold is reached.
[0005] According to another aspect, a system includes: a data memory that includes video data associated with a surgical procedure; and a machine learning training system for training a hierarchical model for hierarchical segmentation of video frames in a surgical video. The system is configured to: obtain an image of an anatomical structure from the video data, wherein the image includes a plurality of image pixels; generate a multi-label probability map for each node in a predefined segmentation class hierarchy; process the plurality of image pixels to generate a leaf-level segmentation map; and process each leaf-level segmentation to update the class label of each leaf-level segmentation to a higher parent class until a predetermined prediction confidence threshold is reached.
[0006] According to one aspect, a computer program product is provided that includes a memory device storing computer-executable instructions that, when executed by one or more processors, cause the one or more processors to perform a plurality of operations for hierarchically segmenting video frames in a surgical video. The plurality of operations include the following: obtaining an image of an anatomical structure, where the image includes a plurality of image pixels; generating multi-label probability maps for two or more nodes in a predefined segmentation class hierarchy; processing the plurality of image pixels to generate a leaf-level segmentation map; and updating the class label of at least one leaf-level segmentation to a higher parent class until a predetermined prediction confidence threshold is reached.
[0007] When taken in conjunction with the accompanying drawings, the above and other features and advantages of the present disclosure will be apparent from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The details of the proprietary rights described herein are particularly pointed out and distinctly claimed in the claims at the end of the specification. The foregoing and other features and advantages of aspects of the present disclosure will be apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:
[0009] Figure 1 A computer-assisted surgery (CAS) system in accordance with one or more aspects is depicted;
[0010] Figure 2 A surgical system in accordance with one or more aspects is depicted;
[0011] Figure 3 A system for analyzing video and data in accordance with one or more aspects is depicted;
[0012] Figure 4A A visual flowchart showing hierarchical model inference for laparoscopic frames in accordance with one or more aspects is depicted;
[0013] Figure 4B A visual flowchart showing the mixing of cystic arteries using a trained hierarchical model in accordance with one or more aspects is depicted;
[0014] Figure 5 A hierarchical structure diagram that can be used for hierarchical training and inference in accordance with one or more aspects is depicted;
[0015] Figure 6 A visual comparison of a classification cross-entropy (CCE) baseline with hierarchical training and inference in accordance with one or more aspects is depicted;
[0016] Figure 7Depicts a visual comparison of CCE baselines with hierarchical training and inference over a time series of frames according to one or more aspects;
[0017] Figure 8 Depicts a flowchart of a method for hierarchical segmentation of video frames in a surgical video; and
[0018] Figure 9 Depicts a block diagram of a computer system according to one or more aspects.
[0019] The figures depicted herein are illustrative. Without departing from the spirit of the various aspects, many variations of these figures and / or operations described herein are possible. For example, actions may be performed in a different order, or actions may be added, deleted, or modified. Additionally, the term "coupled" and its variations describe having a communication path between two elements and do not imply a direct connection between the elements without intermediate elements / connections. All such variations are considered to be part of this specification. Detailed Description
[0020] Exemplary aspects of the technical solutions described herein include systems and methods for hierarchical segmentation of video frames in a surgical video.
[0021] To improve segmentation in surgical video analysis, aspects of the technical solutions are described herein, and these technical solutions can use a segmentation hierarchy and an associated hierarchical inference scheme to allow prediction of grouped anatomical structures when fine-grained categories cannot be reliably distinguished. This disclosure provides unique and novel technical solutions rooted in computational techniques and provides an improvement over current segmentation capabilities to achieve better results than current segmentation techniques. In this disclosure, a multi-label segmentation loss proposed through a hierarchy of anatomical categories is formulated and used to train a network. Subsequently, a leaf-to-root inference scheme ("Hiera-Mix") can be used to determine the trade-off between label confidence and granularity in a given scene. This method can be applied to any segmentation model and can be evaluated using a large dataset (such as, by way of example, a laparoscopic cholecystectomy dataset with 65,000 labeled frames).
[0022] This document describes technical solutions for addressing such technical challenges. In particular, the technical solutions herein can facilitate improving the segmentation and detection accuracy of "critical structures" (e.g., the cystic artery and cystic duct) when evaluating along the entire hierarchical path. This can correspond to a significant improvement in the segmentation output while reducing inter-class confusion. For other anatomical categories, which benefit less from the hierarchy, segmentation and detection are not affected. Additionally, the embodiments described herein provide a hierarchical method for improving the segmentation of surgical scenes in frames with ambiguous anatomical structures. This can be achieved by more appropriately reflecting the model's parsing of the scene and may be beneficial for applications of surgical scene segmentation, including advancements in computer-assisted intraoperative guidance.
[0023] Laparoscopic cholecystectomy is a minimally invasive surgical procedure that can be used to remove the gallbladder. The procedure involves dissecting the "critical area" to expose the "critical structures" (e.g., the cystic artery and cystic duct) that keep the gallbladder attached to the body, and clamping and separating these critical structures after they are exposed. However, adverse consequences can occur during the procedure, including death. In a small number of cases, major bile duct injuries can occur due to accidental division of the common bile duct and cystic duct. Therefore, official guidelines encourage surgeons to establish a "critical safety view" (CVS) before clamping and separating the critical structures. After achieving the CVS, both the cystic artery and cystic duct are separated and clearly distinguishable, such that it can be seen that the cystic artery and cystic duct are the only two structures entering the gallbladder, and thus they can be easily traced as they enter the gallbladder.
[0024] Robot-assisted surgery (RAS) (such as robotic laparoscopy) enables surgeons to perform minimally invasive surgeries with a higher precision rate, thereby reducing pain, scarring, blood loss, and accelerating recovery. A key component of RAS that allows for increased surgical precision is the provision of visual feedback via integrated imaging and display technologies. Such technologies generally enable the surgeon to obtain high-resolution magnified views of the internal anatomy of interest and the surgical tools being used. Additionally, protocols have been developed that allow for postoperative analysis of the recorded surgical videos. However, the interpretation of surgical videos can be challenging for various reasons, including but not limited to occlusion caused by blood, visceral fat, inflammation, smoke (e.g., smoke caused by electrocautery), and other anatomical structures that are not the target structures of interest, anatomical reference attenuation due to magnification / camera angle, and specular reflections. Deep learning-based semantic segmentation of anatomical structures in surgical video frames has the potential to enhance surgical safety and workflow.
[0025] According to the embodiments disclosed herein, hierarchical semantic segmentation (HSS) for segmenting anatomical structures in surgical video frames can be used for any segmentation problem where the classes can be arranged in a hierarchy. HSS can improve segmentation performance in computer vision datasets such as street scene parsing and human parsing. Performance improvement can also be achieved by imposing a hierarchy on the anatomical classes. A hierarchical inference method ("Hiera-Mix") can predict grouped confusable structures until they can be confidently distinguished from each other. The embodiments disclosed herein facilitate the adoption of a hierarchical approach to improve cross-sectional anatomical segmentation of key structures such as the cystic artery and cystic duct, and the unanatomized fat covering the cystic artery and cystic duct, and the common bile duct below the cystic artery and cystic duct.
[0026] Turning now to Figure 1 , an example computer-aided system (CAS) system 100 is generally shown in accordance with one or more aspects. The CAS system 100 includes at least a computing system 102, a video recording system 104, and a surgical instrument system 106. As Figure 1 shown, the actor 112 can be a medical staff member performing a surgical procedure on a patient 110 using the CAS system 100. The actor 112 can be any medical staff member such as a surgeon, an assistant, a nurse, an administrator, or any other actor interacting with the CAS system 100 in a surgical environment. The surgical procedure can be any type of surgery such as, but not limited to, cataract surgery, laparoscopic cholecystectomy, endoscopic transsphenoidal approach (eTSA) pituitary adenoma resection, or any other surgical procedure. In other examples, the actor 112 can be a technician, an administrator, an engineer, or any other such person interacting with the CAS system 100. For example, the actor 112 can record data from the CAS system 100, configure / update one or more attributes of the CAS system 100, check the past performance of the CAS system 100, repair the CAS system 100, etc., including combinations and / or pluralities thereof.
[0027] The surgical procedure can include multiple phases, and each phase can include one or more surgical actions. A "surgical action" can include cutting, compressing, stapling, clamping, suturing, cauterizing, sealing, or any other such action performed to complete a phase of a surgical procedure. A "phase" represents a surgical event consisting of a series of steps (e.g., closure). A "step" refers to accomplishing a specified surgical goal (e.g., hemostasis). During each step, certain surgical instruments 108 (e.g., forceps) are used to achieve a specific goal by performing one or more surgical actions. Additionally, a specific anatomical structure of the patient can be the target of the (multiple) surgical actions.
[0028] The video recording system 104 includes one or more cameras 105, such as operating room cameras, endoscopic cameras, laparoscopic cameras, etc., including combinations and / or multiple ones thereof. The camera 105 captures video data of a surgical procedure being performed. The video recording system 104 includes one or more video capture devices, and the one or more video capture devices may include cameras 105 placed in the operating room to capture events around (i.e., outside the body) a patient undergoing surgery. The video recording system 104 further includes cameras 105 (e.g., endoscopic cameras) that pass through the body of the patient 110 to capture endoscopic data. The endoscopic data provides videos and images of the surgical procedure.
[0029] The computing system 102 includes one or more memory devices, one or more processors, user interface devices, and other components. Figure 1 All or a portion of the computing system 102 shown may be implemented, for example, by Figure 9 all or a portion of the computer system 800. The computing system 102 may execute one or more computer-executable instructions. Execution of the instructions facilitates the computing system 102 to perform one or more methods, including the methods described herein. The computing system 102 may communicate with other computing systems via wired and / or wireless networks. In one or more examples, the computing system 102 includes one or more trained machine learning models, and the one or more trained machine learning models may detect and / or predict features of a surgical procedure being performed or that has been performed previously. Features may include structures (such as anatomical structures) in the captured surgical video, surgical instruments 108. Features may further include events such as stages and / or actions in the surgical procedure. Detected features may further include actors 112 and / or patients 110. In one or more examples, based on the detection, the computing system 102 may provide suggestions for subsequent actions for the actor 112 to take. Alternatively or additionally, the computing system 102 may provide one or more reports based on the detection. The detection performed by the machine learning model may be executed in an autonomous or semi-autonomous manner.
[0030] The machine learning model may include an artificial neural network (such as a deep neural network, a convolutional neural network, a recurrent neural network), a vision transformer, an encoder, a decoder, or any other type of machine learning model. The machine learning model may be trained in a supervised, unsupervised, or hybrid manner. The machine learning model may be trained to perform detection and / or prediction using one or more types of data acquired by the CAS system 100. For example, the machine learning model may use video data captured via the video recording system 104. Alternatively or additionally, the machine learning model uses surgical instrument data from the surgical instrument system 106. In still other examples, the machine learning model uses a combination of video data and surgical instrument data.
[0031] Additionally, in some examples, the machine learning model may also use audio data captured during a surgical procedure. The audio data may include sounds emitted by the surgical instrument system 106 when activating one or more surgical instruments 108. Alternatively or additionally, the audio data may include voice commands, snippets, or conversations from one or more actors 112. The audio data may further include sounds made by the surgical instrument 108 during its use.
[0032] In one or more examples, the machine learning model may detect surgical actions, surgical phases, anatomical structures, surgical instruments, and various other features from data associated with a surgical procedure. In some examples, the detection may be performed in real time. Alternatively or additionally, the computing system 102 analyzes the surgical data (i.e., various types of data captured during a surgical procedure) in an offline manner (e.g., after the surgery). In one or more examples, the machine learning model detects the surgical phase based on detecting some of these features (such as anatomical structures, surgical instruments, etc., including their combinations and / or multiple of them).
[0033] The data collection system 150 may be used to store surgical data, including the (multiple) videos captured during a surgical procedure. The data collection system 150 includes one or more storage devices 152. The data collection system 150 may be a local storage system, a cloud-based storage system, or a combination thereof. Further, the data collection system 150 may use any type of cloud-based storage architecture, such as a public cloud, a private cloud, a hybrid cloud, etc., including their combinations and / or multiple of them. In some examples, the data collection system may use distributed storage, i.e., the storage devices 152 are located at different geographical locations. The storage device 152 may include any type of electronic data storage medium for recording machine-readable data, such as semiconductor-based storage media, magnetic-based storage media, optical-based storage media, etc., including their combinations and / or multiple of them. For example, the data storage medium may include a flash-based solid-state drive (SSD), a magnetic-based hard disk drive, magnetic tape, an optical disc, etc., including their combinations and / or multiple of them.
[0034] In one or more examples, the data collection system 150 can be part of the video recording system 104, and vice versa. In some examples, the data collection system 150, the video recording system 104, and the computing system 102 can communicate with each other via a communication network, which can be wired, wireless, or a combination thereof. Communication between the systems can include data (e.g., video data, instrument data, etc., including combinations and / or multiple items thereof) transmission, data manipulation commands (e.g., browse, copy, paste, move, delete, create, compress, etc., including combinations and / or multiple items thereof), data manipulation results, etc., including combinations and / or multiple items thereof. In one or more examples, the computing system 102 can manipulate data that has been stored / is being stored in the data collection system 150 based on the output from one or more machine learning models (e.g., stage detection, anatomical structure detection, surgical tool detection, etc., including combinations and / or multiple items thereof). Alternatively or additionally, the computing system 102 can manipulate data that has been stored / is being stored in the data collection system 150 based on information from the surgical instrument system 106.
[0035] In one or more examples, the video captured by the video recording system 104 is stored on the data collection system 150. In some examples, the computing system 102 collates portions of the video data stored on the data collection system 150. In some examples, the computing system 102 filters the video before it is stored on the data collection system 150 by the video recording system 104. Alternatively or additionally, the computing system 102 filters the video after it is stored on the data collection system 150 by the video recording system 104.
[0036] Now turning to Figure 2 , a surgical system 200 is generally shown in accordance with one or more aspects. Figure 2 Examples of Figure 1 depict a surgical support system 202 that can include or can be coupled to Figure 1When performing a surgical operation on patient 110 in [system name], the surgical support system 202 can store, access, and / or update surgical data 214 associated with a training dataset and / or real-time data. The surgical support system 202 can store, access, and / or update surgical objectives 216 to assist in training and guiding one or more surgical operations. The user configuration 218 can track and store user preferences.
[0037] Now turning to Figure 3 , a system 300 for analyzing video and data according to one or more aspects is generally shown. According to various aspects, video and data are captured from Figure 1 the video recording system 104. The analysis can enable the use of machine learning to predict features in the video data including surgical stages and structures (e.g., instruments, anatomical structures, etc., including combinations and / or multiple items thereof). In one or more examples, the system 300 can be Figure 1 the computing system 102 or a part thereof. According to some aspects, the system 300 uses a data stream in the surgical data to identify the surgical state.
[0038] The system 300 includes a data receiving system 305 that collects surgical data, including video data and surgical instrument data. The data receiving system 305 can include one or more devices (e.g., one or more user devices and / or servers) located within and / or associated with the surgical operating room and / or control center. The data receiving system 305 can receive surgical data in real time, i.e., while the surgical operation is being performed. Alternatively or additionally, the data receiving system 305 can receive or access surgical data in an offline manner, e.g., by accessing data stored in Figure 1 the data collection system 150 in [system name].
[0039] System 300 further includes a machine learning processing system 310 that uses one or more machine learning models to process surgical data to identify one or more features in the surgical data, such as surgical stages, instruments, anatomical structures, etc., including combinations thereof and / or multiple items therein. It will be understood that the machine learning processing system 310 may include one or more devices (e.g., one or more servers), and each device may be configured to include a part or all of one or more of the depicted components of the machine learning processing system 310. In some instances, part or all of the machine learning processing system 310 is located in the cloud and / or away from the operating room and / or physical location corresponding to part or all of the data receiving system 305. It should be understood that several components of the machine learning processing system 310 are depicted and described herein. However, these components are only an example structure of the machine learning processing system 310, and in other examples, different combinations of these components may be used to construct the machine learning processing system 310. Such variations in the combination of components are included in the technical solutions described herein.
[0040] The machine learning processing system 310 includes a machine learning training system 325, which may be a separate device (e.g., a server) that stores its output as one or more trained machine learning models 330. The machine learning models 330 can be accessed by the machine learning execution system 340. In some examples, the machine learning execution system 340 may be separate from the machine learning training system 325. In other words, in some aspects, the device that "trains" the model is separate from the device that "infers" (i.e., uses the trained machine learning model 330 to perform real-time processing of surgical data).
[0041] In some examples, the machine learning processing system 310 further includes a data generator 315 that combines real images and video data from the video recording system 104 to generate simulated surgical data (such as a set of synthetic images and / or synthetic videos) to generate the trained machine learning model 330. The data generator 315 can access (read / write) the data memory 320 to record data, including multiple images and / or multiple videos. The images and / or videos may include images and / or videos collected during one or more surgeries (e.g., one or more surgical procedures). For example, the images and / or videos may have been Figure 1 collected by the actors 112 (e.g., surgeons, surgical nurses, anesthesiologists, etc., including combinations thereof and / or multiple items therein) wearing user devices during the surgery, non-wearable imaging devices located in the operating room, endoscope cameras inserted Figure 1 into the patient 110 within the body, etc. (including combinations thereof and / or multiple items therein). In some examples, the data memory 320 is associated with Figure 1separated from the data collection system 150. In other examples, the data memory 320 is part of the data collection system 150.
[0042] Each image and / or video recorded in the data memory 320 for performing training (e.g., generating the machine learning model 330) can be defined as a base image and can be associated with other data characterizing the associated surgical and / or rendering specifications. For example, the other data can identify the type of surgery, the location of the surgery, one or more individuals involved in performing the surgery, the surgical objective, and / or the surgical outcome. Alternatively or additionally, the other data can indicate the surgical stage corresponding to the image or video, the rendering specifications corresponding to the image or video, and / or the type of imaging device that captured the image or video (e.g., if the device is a wearable device, and / or the role of the specific person wearing the device, etc., including combinations and / or pluralities thereof). Further, the other data can include image segmentation data identifying and / or characterizing one or more objects depicted in the image or video (e.g., tools, anatomical objects, etc., including combinations and / or pluralities thereof). The characterization can indicate the location, orientation, or pose of the object in the image. For example, the characterization can indicate a set of pixels corresponding to the object and / or the state of the object resulting from past or current user operations. Various techniques for identifying an object in one or more coordinate systems can be used to perform the positioning.
[0043] The machine learning training system 325 uses the recorded data in the data memory 320 to generate a trained machine learning model 330, and the recorded data can include simulated surgical data (e.g., a set of synthetic images and / or synthetic videos) and / or actual surgical data. The trained machine learning model 330 can be defined based on a model type and a set of hyperparameters (e.g., defined based on an input from a client device). The trained machine learning model 330 can be configured based on a set of parameters, and the set of parameters can be dynamically defined based on (e.g., continuous or repeated) training (i.e., learning, parameter tuning). The machine learning training system 325 can use one or more optimization algorithms to define the set of parameters to minimize or maximize one or more loss functions. The (learned) set of parameters can be stored using a specific data structure of a certain trained machine learning model in the trained machine learning model 330 as part of the trained machine learning model 330. The data structure can also include one or more non-learnable variables (e.g., hyperparameters and / or model definitions).
[0044] The machine learning execution system 340 can access the (multiple) data structures of the trained machine learning model 330 and configure the trained machine learning model 330 accordingly for inference (e.g., prediction, classification, etc., including combinations and / or multiple items thereof). The trained machine learning model 330 can include, for example, fully convolutional network adaptation, adversarial network models, encoders, decoders, or other types of machine learning models. The type of the trained machine learning model 330 can be indicated in the corresponding data structure. The trained machine learning model 330 can be configured according to one or more hyperparameters and a set of learned parameters.
[0045] The trained machine learning model 330 receives surgical data to be processed as input during execution and then generates one or more inferences according to the training. For example, the video data captured by the video recording system 104 in Figure 1 can include a data stream (e.g., an array of intensity, depth, and / or RGB values) for each of a single image or a set of frames (e.g., including multiple images or images with sequence data), and the set of frames represents a fixed or variable length time window in the video. The video data captured by the video recording system 104 can be received by the data receiving system 305, which can include one or more devices located in the operating room where the surgical operation is being performed. Alternatively, the data receiving system 305 can include a device located remotely, and during the execution of the surgical operation, the captured video data is streamed to the device in real time. Alternatively or additionally, the data receiving system 305 accesses data from the data collection system 150 or any other data source (e.g., local or remote storage device) in an offline manner.
[0046] The data receiving system 305 can process the received video and / or data. When a video stream in an encoded format is received, the processing can include decoding so that the data of the image sequence can be extracted and processed. The data receiving system 305 can also process other types of data included in the input surgical data. For example, the surgical data can include additional data streams (such as audio data, RFID data, text data, measurements from one or more surgical instruments / sensors, etc., including combinations and / or multiple items thereof), which can represent the stimulation / surgical state in the operating room. The data receiving system 305 synchronizes different inputs from different devices / sensors before inputting them into the machine learning processing system 310.
[0047] Once trained, the trained machine learning model 330 can analyze the input surgical data and, in one or more respects, predict and / or characterize features (e.g., structures) included in video data including the surgical data. The video data can include sequential images and / or encoded video data (e.g., encoded using a digital video file / stream format and / or codec such as MP4, MOV, AVI, WEBM, AVCHD, OGG, etc., including combinations and / or pluralities thereof). The prediction and / or characterization of the features can include segmenting the video data or predicting the localization of structures with a probability heatmap. In some instances, one or more preprocessing or enhancements (e.g., intensity normalization, resizing, cropping, etc., including combinations and / or pluralities thereof) are performed before or associated with segmenting the video data by the one or more trained machine learning models 330. The output of the one or more trained machine learning models 330 can include image segmentation or probability heatmap data indicating which, if any, of a defined set of structures are predicted within the video data, the location and / or localization and / or pose of the (multiple) structures within the video data, and / or the state of the (multiple) structures. The location can be a set of coordinates in an image / frame in the video data. For example, the coordinates can provide a bounding box. The coordinates can provide the boundaries around the predicted (multiple) structures. In one or more examples, the trained machine learning model 330 is trained to perform higher-level prediction and tracking, such as predicting the stage of a surgical procedure and tracking one or more surgical instruments used in the surgical procedure.
[0048] Although some techniques for predicting the surgical stage (“stage”) in a surgical procedure are described herein, it should be understood that any other stage prediction techniques that do not affect aspects of the technical solutions described herein can be used. In some examples, the machine learning processing system 310 includes a detector 350 that uses the trained machine learning model 330 to identify various items or states within a surgical procedure (“surgery”). The detector 350 can use a specific surgical tracking data structure 355 from a list of surgical tracking data structures. The detector 350 can select the surgical tracking data structure 355 based on the type of surgical procedure being performed. In one or more examples, the type of surgical procedure can be predetermined or input by the actor 112. For example, the surgical tracking data structure 355 can identify a set of potential stages that may correspond to a portion of a specific type of surgery as “stage predictions,” where the detector 350 is a stage detector.
[0049] In some examples, the surgical tracking data structure 355 can be a graph that includes a set of nodes and a set of edges, where each node corresponds to a potential stage. These edges can provide a directed connection between the nodes, and this directed connection (via the direction) indicates the expected order in which the stages will be encountered throughout the program iteration. The surgical tracking data structure 355 can include one or more branch nodes that feed into multiple next nodes and / or can include one or more divergence points and / or convergence points between the nodes. In some instances, a stage indicates a surgical action that is being performed or has been performed (e.g., a surgical maneuver) and / or a combination of actions that have been performed. In some instances, a stage is related to the biological state of a patient undergoing a surgical procedure. For example, the biological state can indicate a complication (e.g., a blood clot, an arterial / venous blockage, etc., including combinations and / or multiple items thereof), a precondition (e.g., a lesion, a polyp, etc., including combinations and / or multiple items thereof). In some examples, the trained machine learning model 330 is trained to detect "abnormal conditions" such as bleeding, arrhythmia, vascular anomalies, etc., including combinations and / or multiple items thereof.
[0050] Each node within the surgical tracking data structure 355 can identify one or more characteristics of the stage corresponding to that node. These characteristics can include visual characteristics. In some instances, a node identifies one or more tools that are typically being used or are available (e.g., on a tool tray) during that stage. The node also identifies one or more roles of the person who is typically performing the surgical task, typical types of movements (e.g., hand or tool movements), etc., including combinations and / or multiple items thereof. Thus, the detector 350 can use the segmentation data generated by the machine learning execution system 340 that indicates the presence and / or characteristics of specific objects within the field of view to identify the estimated node corresponding to the real image data. The identification of the node (i.e., the stage) can be further based on the previously detected stages of a given program iteration and / or other detected inputs (e.g., including verbal audio data that includes a person-to-person request or comment, an explicit identification of the current or past stage, a request for information, etc., including combinations and / or multiple items thereof).
[0051] The detector 350 can output predictions, such as phase predictions associated with a portion of the video data analyzed by the machine learning processing system 310. The phase prediction is associated with the portion of the video data by identifying the start time and end time of the video portion analyzed by the machine learning execution system 340. The phase prediction output can include segmentation of the video, where each segment corresponds to and includes an identification of a surgical phase detected by the detector 350 based on the output of the machine learning execution system 340. Further, in one or more examples, the phase prediction can include additional data dimensions, such as, but not limited to, an identification of structures (e.g., instruments, anatomical structures, etc., including combinations and / or pluralities thereof) identified by the machine learning execution system 340 in the analyzed video portion. The phase prediction can also include a confidence score for the prediction. Other examples can include various other types of information in the output phase prediction. Further, other types of outputs of the detector 350 can include status information or other information for generating an audio output, a visual output, and / or a command. For example, the output can trigger an alert, enhance visualization, identify a current condition of the prediction, identify a future condition of the prediction, command control of a device, and / or cause other such data / commands to be transmitted to support system components, such as, for example, via Figure 2 the surgical support system 202.
[0052] It should be noted that although some of the figures depict endoscopic video being analyzed, the technical solutions described herein can also be applicable to analyzing video and image data captured by a camera that is not an endoscope (i.e., a camera outside the patient) during the performance of an open surgery (i.e., a non-laparoscopic surgery). For example, the video and image data can be captured by a camera mounted on one or more individuals (e.g., a surgeon) in the operating room. Alternatively or additionally, the camera can be mounted on a surgical instrument, a wall, or other locations in the operating room. Alternatively or additionally, the video can be images captured by other imaging modalities, such as ultrasound.
[0053] Now turning to Figure 4A and Figure 4B , a method for hierarchical segmentation of video frames in a surgical video is depicted according to one or more aspects. As shown, Figure 4A hierarchical model inference for a laparoscopic frame 400 is depicted, and Figure 4BDepicts a hybrid 450 shown only for the cystic artery, where box 452 corresponds to a trained hierarchical model that processes an image to give a multi-label probability map for each node in a predefined hierarchical structure of segmentation classes. Box 454 corresponds to a higher-level key structure class that is used to indicate the uncertainty between the cystic artery and the cystic duct, where root-to-leaf and inference are performed for each pixel to give a leaf-level segmentation map. Box 456 corresponds to a root-level key region class that groups the key structures and the unanatomized regions below them, where a post-processing step is performed for each leaf-level anatomical segmentation, and thereby, the associated class labels are updated to successively higher parent class labels in the hierarchy until sufficient prediction confidence is obtained. Box 458 corresponds to an "unknown" class that is used to indicate the uncertainty at the root level of the hierarchy.
[0054] It should be understood that the HSS can include a hierarchy that arranges the segmentation classes in a tree-like structure for the purpose of enhancing learning by exploiting hierarchical relationships. The hierarchy T can be composed of nodes and edges (V, Ε). Each node v ∈ V represents a class, and each edge (u, v) ∈ Ε represents a hierarchical relationship between two classes u, v ∈ V, where v is the parent node of the child node u. In T, it can be assumed that each class node is both its own parent node and its own child node, i.e., (v, v) ∈ E. Additionally, in T, the root node V R represents the most general class, while the leaf node V L represents the most fine-grained class. It should also be understood that a typical hierarchy-independent segmentation model maps the graph to a dense feature tensor Subsequently, the Softmax operator is used to map F to a dense probability tensor Y ∈ [0, 1] H×W×|V L | . The hierarchy-independent segmentation model is typically optimized using the categorical cross-entropy (CCE) loss:
[0055]
[0056] where, T ∈ {0, 1} H×W×|V L | represents the ground truth tensor. During inference, the argmax operation can be used to obtain the prediction map P ∈ V L H×W .
[0057] Note that hierarchical semantic segmentation may require a change from the multi-class classification formulation described above for the hierarchy-agnostic model to a multi-label classification formulation, i.e., now each pixel is mapped to a class for each level of the hierarchy rather than mapping each pixel to a single class in the set of leaf nodes. If it is assumed that T has N levels and each level is "complete", i.e., contains classes that account for all possible objects in the image, then the probability tensor output can be defined by the hierarchical model as:
[0058] S ∈ [0,1]H×W×|V|,
[0059] where S is the union of the probability tensors Yi for each level of the hierarchy (i.e., S = Y1 ∪ Y2 ∪ …… ∪ YN). Thus, the hierarchical model is trained using the sum of the CCE losses, which can be referred to as "Hiera CCE" and can be described by:
[0060] L HieraCCE = Σ n∈N L CCE (Y n ,T n ), (2)
[0061] During inference, the leaf node classes can be used to obtain a granular prediction, but considering the highest scoring root-to-leaf path for each pixel i in the hierarchy, as described by:
[0062]
[0063] where P is the set of root-to-leaf paths in the hierarchy and is the highest scoring root-to-leaf path, where the leaf node class can be assigned to each pixel to give the leaf-level prediction P L .
[0064] It should be understood that while equation (3) ensures that pixel predictions consider the hierarchy during the inference phase, the Hiera CCE loss described by equation (2) does not enforce the hierarchical relationship during the training phase. One way to address this issue involves applying the "tree-min" loss ("Hiera TM") method, where Hiera TM enforces the following two properties:
[0065] a. Positive T property: For each pixel, if a class is labeled positive, then all of its parent nodes in T should be labeled positive.
[0066] b. Negative T property: For each pixel, if a class is labeled negative, then all of its child nodes in T should be labeled negative.
[0067] Based on these two characteristics, for the pixel prediction vector s = [s v v∈V ∈ [0,1] |V| the following two constraints are as follows:
[0068] c. Positive T constraint: For each pixel, if v is labeled positive and u is the parent node of v, then s v ≤ s u should hold.
[0069] d. Negative T constraint: For each pixel, if v is labeled negative and u is the child node of v, then 1 - s v ≤ 1 - s u should hold.
[0070] These two hierarchical constraints can be incorporated into the Hiera TM loss. Thus, given the score vector s = {s v} v∈V ∈ [0,1] |V| and the associated ground truth binary label vector t = {t v} v∈V ∈ {0,1} |V| :
[0071]
[0072] where A v and C v respectively represent the subset of the parent nodes and the subset of the child nodes of v in T.
[0073] It should be understood that in order to better utilize the hierarchy imposed on the class labels, a post - processing inference method (i.e., Hiera - Mix) can be performed to update the class labels in the fine - grained leaf - level prediction map to more general class labels in the hierarchy's parent based on a prediction confidence threshold. For example, referring to Figure 4B , the gallbladder artery segmentation can be updated to "critical structure" (box 454), "critical region" (box 456), or "unknown" (box 458), and the update stops only when the prediction confidence threshold is met or exceeded. More formally, for an N - level hierarchy T, the leaf - level prediction map P L can be obtained by using the highest - scoring root - to - leaf node inference scheme, and each class in P L is iterated. For the class v L in P N , the binary mask can be defined as B N , and the score maps of each class in the root - to - leaf path of the class v N can be defined as S1,..., S N and the associated classes v1,..., vN Thus, for each v i the class confidence m can be calculated using the masked mean given by i :
[0074]
[0075] where H and W are the dimensions of P L Using a predefined confidence threshold T, the class label v N is reassigned to v i* where the index i can be determined as follows:
[0076] i * = max{i | m i ≥ T, i ∈ 1, …, N}. (6)
[0077] It should be understood that if the threshold condition is not met, i.e., there is not enough confidence at the root level, the class label is reassigned to "unknown".
[0078] To validate this hierarchical approach, according to one aspect, an experimental setup was created and evaluated using 65,000 labeled frames from 1107 independent internally collected laparoscopic cholecystectomy videos. Frames were sampled from the videos in a 30 - second window at a rate of one frame per second (fps), and the window was selected from the entire anatomical region containing the key structures (hereinafter referred to as the "key region"). The following structures were labeled: the key region, the cystic artery, the cystic duct, the liver, the Rouviere sulcus, the gallbladder, and the intestinal structures. Additionally, the labeled frames were divided into training frames, validation frames, and test frames (80% / 10% / 10%) according to the videos.
[0079] According to one aspect, a segmentation network with a Swin Base (Swin-B) transformer backbone (Swin Seg) was evaluated. Additionally, HRNet was also used for comparison in the ablation experiments. It should be understood that both networks provide a common baseline for segmentation and both networks can be implemented using PyTorch 1.12. The model was optimized using the AdamW optimizer, a learning rate of 0.0001, and a "1Cycle" scheduler (such as "OneCycleLR" in PyTorch). The model was trained for 40 epochs, where the batch size was 8, and the model that converged at epoch 40 was used for evaluation. A "balanced" sampler was used to select the training examples in each epoch, where each epoch included 2500 samples for each class label. In the example, it took approximately 24 hours to train each model on a 48G NVIDIA Graphics Processing Unit (GPU). During training, the model used random image augmentations (e.g., padding, cropping, flipping, blurring, rotation, and noise). This is merely an example, and many variations can be implemented in accordance with aspects of the present disclosure.
[0080] It should be understood that, in one aspect, for baseline evaluation, the classification cross-entropy (CCE) loss can be used to train the model. Referring to Figure 5 , a hierarchy 500 that can be used for hierarchical training and inference is shown. In the example, three hierarchical loss variants were considered during the ablation experiments: Hiera CCE, Hiera TM, and a mixture of these losses (Hiera TM + CCE), which is given by:
[0081] L HieraTM+CCE = λ1L HieraTM + λ2L HieraCCE , (7)
[0082] where the values of λ1 = 5 and λ2 = 1 were used in the experiment.
[0083] In the example and referring to Figure 5 , a hierarchy 500 that can be used for hierarchical model training and inference is shown. This hierarchy groups the cystic artery and cystic duct together under a key structure, while the key region corresponds to the union of the unanatomized peritoneal covering area containing the key structure before exposure and the key structure and the unanatomized area below it after exposure. Through this construction, a hierarchical path is created from the cystic artery and cystic duct to the key structure and then to the key region, thus allowing Hiera-Mix to trade off between more easily segmentable but coarser labels (key region and key structure) and more difficult to segment but fine-grained labels (cystic artery and cystic duct).
[0084] The segmentation performance can be evaluated using the Dice score, precision, and recall per pixel. In the examples, the F1 score, precision, and recall per structure are used to evaluate the frame-level presence detection. In this case, for an anatomical structure to be detected as a true positive in a frame, the Dice score against the ground truth annotation needs to be 0.5. To evaluate the proposed Hiera-Mix method, hierarchical segmentation and detection metrics are designed, where higher-level classes in the hierarchical path are allowed to be counted as true positives. For example, when calculating the metrics for the cystic artery class, its parent classes (key structures and key regions) are counted as true positives. An "H" is added to denote the hierarchical metric, e.g., "Dice-H".
[0085] Referring to Table 1 below, the impact of Hiera-Mix on the cystic artery and cystic duct using the hierarchical segmentation and detection metrics disclosed herein is shown. With Hiera-Mix, it is observed that both the Dice-H per pixel and the detection F1-H for both the cystic artery and cystic duct are improved in both the validation set and the test set. This is mainly attributed to a significant increase in precision-H and a small increase in recall-H compared to the CCE baseline (where key regions are counted as valid true positives to allow for a fair comparison). Thus, Table 1 shows that Hiera-Mix improves the segmentation (top) and detection (bottom) of the cystic artery and cystic duct. In this regard, the metrics assume that the key structure class and the key region class are valid predictions for the cystic artery and cystic duct, where the improvements are shown in green.
[0086]
[0087] Table 1
[0088] Figure 6 A visual comparison 600 of the CCE baseline and Hiera-Mix is shown in which the top row 602 shows a frame in which the CCE model incorrectly classifies the cystic artery as the cystic duct, while Hiera-Mix more correctly identifies it as a key structure of the uncertain class. This is a difficult example because the artery is located to the left of the duct in the frame, which is atypical. The middle row 604 shows a frame in which the model trained with CCE misses the cystic artery, while the hierarchical model with Hiera-Mix has detected the cystic artery as a key structure. The third row 606 shows that the model trained with CCE detects the cystic duct, while the hierarchical model labels it as a key region, as in the GT. The bottom row 608 shows a frame in which the model trained with CCE has segmented the cystic duct before sufficient dissection, but the hierarchical model has labeled it as the gallbladder, as in the ground truth. It should be understood that in Figure 6In it, the cystic artery can be shown in light green, the cystic duct can be shown in off-white, the key structure can be shown in blue, the key area can be shown in dark purple, the gallbladder can be shown in dark green, the liver can be shown in light brown, and the Rouviere sulcus can be shown in light purple.
[0089] Figure 7 Another aspect of Hiera-Mix is shown in which the CCE model 650 misses the cystic artery or under-segments it in the first four frames. In contrast, Hiera-Mix uses key structure labels to more accurately capture the extent of the cystic artery over the entire sequence in the box. As can be seen, it is observed that the CCE model has omissions and under-segmentations of the cystic artery, while Hiera-Mix uses cystic artery and key structure labels to better capture the extent of the cystic artery over the entire sequence. It should be understood that in Figure 7 , the cystic artery can be shown in light green, the cystic duct can be shown in off-white, the key structure can be shown in blue, the key area can be shown in dark purple, the gallbladder can be shown in dark green, and the liver can be shown in light brown.
[0090] In addition, Table 2 below shows the performance measured using non-hierarchical segmentation and detection metrics (predictions of only leaf labels constitute true positives). In this case, it is observed that when examining the single-class segmentation overlap of the cystic artery and cystic duct, the mixture improves the precision of both segmentation (top) and detection (bottom) and reduces the recall, which is due to applying a confidence threshold that promotes more conservative behavior compared to the CCE baseline. Also, while the segmentation Dice is slightly reduced for Hiera-Mix, the detection F1 is improved. Positive differences can be shown in green, while negative differences can be shown in red.
[0091]
[0092]
[0093] Table 2
[0094] Referring to Table 3 below, the segmentation and detection performance of all classes of Swin Seg trained with CCE loss and Hiera TM+CCE loss are shown. Importantly, for classes without hierarchical relationships, it is observed that the performance of these two losses is roughly similar on both the validation set and the test set. Ablation experiments are first run to determine the model and hierarchical loss to be used in further experiments. Table 4 below shows the average Dice scores of all classes. It is observed that the best configuration is Swin Seg trained with the hybrid hierarchical loss Hiera TM+CCE. Therefore, the results shown in Table 1, Table 2, and Table 3 and Figures 4 and Figure 5 used this method.
[0095]
[0096] Table 3
[0097]
[0098] Table 4
[0099] Having a Hybrid Hierarchical Split (Hiera-Mix) allows the segmentation model to reflect class label uncertainty in its segmentation output, e.g., labeling a structure as "critical structure" when it is unclear whether the anatomical structure is the cystic artery or the cystic duct. When evaluated on a sub-hierarchy for each structure, this method can improve the segmentation and detection accuracy of the cystic artery and the cystic duct. Compared with a model trained using the standard categorical cross-entropy (CCE) loss, the increased precision achieved using Hiera-Mix means a reduction in false positive predictions, while the increased recall indicates that Hiera-Mix more often classifies the cystic artery and the cystic duct as at least "critical regions". The latter advantage may be due to the strengthening of the attribution of the cystic artery and the cystic duct to critical regions during hierarchical model training.
[0100] The hierarchy of laparoscopic cholecystectomy creates a path from critical regions to critical structures and then to distinguishable cystic artery and cystic duct. Hiera-Mix applied to laparoscopic cholecystectomy aims to enforce the attribution of critical structures to critical regions, reduce the premature detection of critical structures, and reduce the misidentification of the cystic artery and the cystic duct. With Hiera-Mix, it can be observed that there is an improvement in Dice-H per pixel and Detection F1-H, which is attributed to a significant increase in precision and a small increase in recall compared to the CCE baseline, where critical regions are also counted as valid true positives to allow for a fair comparison.
[0101] It should be understood that Hiera-Mix can allow the segmentation model to handle class label uncertainty in its segmentation, e.g., labeling an anatomical structure as a critical structure when it is unclear whether the anatomical structure is the cystic artery or the cystic duct. By evaluating on a sub-hierarchy for each structure, it is observed that the hierarchical method improves the segmentation and detection of the cystic artery and the cystic duct. In this regard, the increased precision-H of the hierarchical method may indicate a reduction in false positives, reflecting a more conservative behavior compared to the CCE-trained model baseline, while the increased recall-H may indicate that Hiera-Mix more often assigns valid labels compared to the CCE-trained model. The latter advantage may be due to the hierarchical method allowing the cystic artery and the cystic duct to assume classes at multiple levels of the hierarchy during training.
[0102] Now turning to Figure 8, generally shows a flowchart of a method 700 for segmenting anatomical structures in surgical video frames using hierarchical semantic segmentation (HSS). All or part of method 700 can be implemented, for example, by all or part of the CAS system 100 of Figure 1 and / or Figure 9 the computer system 800 of
[0103] Referring to Figure 8 , according to some aspects, a method 700 for segmenting anatomical structures in surgical video frames using hierarchical semantic segmentation (HSS) is shown, and the method includes obtaining an image of an anatomical structure and / or region of interest, as shown by operation box 702. Referring again to Figure 4A and Figure 4B , hierarchical model inference for laparoscopic frame 400 is shown, where only the endoscopic image of the cystic artery is obtained. The image can be processed to generate a multi-label probability map for each node in a predefined hierarchical structure of segmentation classes, as shown by operation box 704. In this case, a trained hierarchical model can be used to process the image to generate a multi-label probability map for each node in a predefined hierarchical structure of segmentation classes. A leaf-level segmentation map can be generated by performing root-to-leaf and inference on each of the image pixels, as shown by operation box 706. Each leaf-level anatomical segmentation can be processed, and each class label can be updated to successively higher parent classes until sufficient prediction confidence is achieved, as shown by operation box 708.
[0104] Figure 8 The processing shown in Figure 8 is not intended to indicate that the operations are to be performed in any particular order, nor that all the operations shown in Figure 8 are to be included in every case. Additionally,
[0105] Now turning to Figure 9, in accordance with one aspect, computer system 800 is generally shown. As described herein, computer system 800 can be an electronic computer framework that includes and / or employs any number of computing devices and networks utilizing various communication technologies and combinations thereof. Computer system 800 can be easily expanded, extended, and modularized, and has the ability to be changed to different services or reconfigure some features independently of other features. Computer system 800 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smartphone. In some examples, computer system 800 can be a cloud computing node. Computer system 800 can be described in the general context of computer-executable instructions, such as program modules, executed by the computer system. Generally, program modules can include routines, programs, objects, components, logics, data structures, etc. that perform specific tasks or implement specific abstract data types. Computer system 800 can be practiced in a distributed cloud computing environment where tasks are executed by remote processing devices linked by a communication network. In a distributed cloud computing environment, program modules can be located in local and remote computer system storage media including memory storage devices.
[0106] As Figure 9 shown, computer system 800 has one or more central processing units (CPUs) 801a, 801b, 801c, etc. (collectively or generically referred to as (multiple) processors 801). Processor 801 can be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. Processor 801 can be any type of circuit capable of executing instructions. Processor 801 is also referred to as a processing circuit, which is coupled to system memory 803 and various other components via system bus 802. System memory 803 can include one or more memory devices, such as read-only memory (ROM) 804 and random access memory (RAM) 805. ROM 804 is coupled to system bus 802 and can include a basic input / output system (BIOS) that controls certain basic functions of computer system 800. RAM is a read-write memory coupled to system bus 802 for use by processor 801. System memory 803 provides temporary storage space for the operation of the instructions during operation. System memory 803 can include random access memory (RAM), read-only memory, flash memory, or any other suitable memory system.
[0107] Computer system 800 includes an input / output (I / O) adapter 806 and a communication adapter 807 coupled to system bus 802. I / O adapter 806 can be a small computer system interface (SCSI) adapter that communicates with hard disk 808 and / or any other similar components. I / O adapter 806 and hard disk 808 are collectively referred to as mass storage device 810 herein.
[0108] Software 811 for execution on a computer system 800 may be stored in a mass storage device 810. The mass storage device 810 is an example of a tangible storage medium readable by a processor 801, where the software 811 is stored as instructions for execution by the processor 801 to operate the computer system 800, as described below with reference to the various figures. Examples of computer program products and the execution of such instructions are discussed in more detail herein. A communication adapter 807 interconnects the system bus 802 with a network 812, which may be an external network that enables the computer system 800 to communicate with other such systems. In one aspect, a portion of the system memory 803 and the mass storage device 810 together store an operating system, which may be any suitable operating system that coordinates Figure 9 the functions of the various components shown.
[0109] Additional input / output devices are shown connected to the system bus 802 via a display adapter 815 and an interface adapter 816. In one aspect, adapters 806, 807, 815, and 816 may be connected to one or more I / O buses that are connected to the system bus 802 via an intermediate bus bridge (not shown). A display 819 (e.g., a screen or display monitor) is connected to the system bus 802 by the display adapter 815, which may include a graphics controller and a video controller for enhancing the performance of graphics-intensive applications. A keyboard, mouse, touch screen, one or more buttons, speakers, etc. may be interconnected with the system bus 802 via the interface adapter 816, which may include, for example, a super I / O chip that integrates multiple device adapters into a single integrated circuit. Suitable I / O buses for connecting peripheral devices such as hard disk controllers, network adapters, and graphics adapters typically include common protocols such as Peripheral Component Interconnect (PCI). Thus, as Figure 9 configured, the computer system 800 includes processing capabilities in the form of a processor 801, storage capabilities including the system memory 803 and the mass storage device 810, input devices such as buttons, touch screens, etc., and output capabilities including speakers 823 and a display 819.
[0110] In some aspects, the communication adapter 807 may transmit data using any suitable interface or protocol (such as Internet Small Computer System Interface, etc.). The network 812 may be a cellular network, a radio network, a Wide Area Network (WAN), a Local Area Network (LAN), or the Internet, etc. External computing devices may be connected to the computer system 800 via the network 812. In some examples, the external computing device may be an external network server or a cloud computing node.
[0111] It should be understood that Figure 9 the block diagram is not intended to indicate that the computer system 800 will includeFigure 9 all of the components shown. Instead, computer system 800 can include Figure 9 any suitable fewer or additional components not shown therein (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). Further, aspects described herein with respect to computer system 800 can be implemented with any suitable logic, where, in various aspects, the logic as referred to herein can include any suitable hardware (e.g., processors, embedded controllers, or application specific integrated circuits, etc.), software (e.g., application programs, etc.), firmware, or any suitable combination of hardware, software, and firmware. Aspects can be combined to include two or more of the aspects described herein.
[0112] Aspects disclosed herein can be systems, methods, and / or computer program products at any possible level of integration of technical details. A computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to execute aspects.
[0113] A computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read only memory (ROM), an erasable programmable read only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium as used herein should not be construed as being a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0114] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network), or to an external computer or external storage device. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.
[0115] The computer-readable program instructions for performing the operations of this disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and high-level languages such as Python, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer, and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet, using an Internet service provider). In some aspects, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by using the state information of the computer-readable program instructions to personalize the electronic circuit in order to perform aspects of this disclosure.
[0116] Aspects are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to aspects of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0117] These computer-readable program instructions can be provided to a processor of a computer system, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions which implement aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0118] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0119] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various aspects. In this regard, each box in the flowchart or block diagram may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified (multiple) logical function. In some alternative implementations, the functions noted in the box may not occur in the order noted in the figures. For example, two boxes shown in succession may, in fact, be executed substantially concurrently, or the boxes may sometimes be executed in the reverse order, depending on the functionality involved. It will also be noted that each box of the block diagrams and / or flowchart illustration, and combinations of boxes in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.
[0120] The description of the various aspects is presented for purposes of illustration, but is not intended to be exhaustive or limited to the disclosed aspects. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described aspects. The terms used herein were chosen to best explain the principles of the aspects, the practical application, or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the aspects described herein.
[0121] Aspects are described herein with reference to the accompanying drawings. Alternative aspects may be devised without departing from the scope of this disclosure. Various connections and positional relationships between elements (e.g., above, below, adjacent, etc.) are set forth in the following description and the drawings. Unless otherwise stated, these connections and / or positional relationships may be direct or indirect, and this disclosure is not intended to be limiting in this regard. Accordingly, the coupling of entities may refer to direct or indirect coupling, and the positional relationship between entities may be a direct or indirect positional relationship. In addition, the various tasks and process steps described herein may be incorporated into a more comprehensive program or process having additional steps or functions not detailed herein.
[0122] The following definitions and abbreviations are used to interpret the claims and the specification. As used herein, the terms "comprises / comprising / includes / including", "has / having", "contains / containing", or any other variation thereof are intended to cover non-exclusive inclusion. For example, a composition, mixture, process, method, article, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such composition, mixture, process, method, article, or apparatus.
[0123] In addition, the term "exemplary" is used herein to mean "serving as an example, instance, or illustration". Any aspect or design described herein as "exemplary" is not necessarily to be construed as more preferred or advantageous than other aspects or designs. The terms "at least one" and "one or more" can be understood to include any integer greater than or equal to one, i.e., one, two, three, four, etc. The term "plurality" can be understood to include any integer greater than or equal to two, i.e., two, three, four, five, etc. The term "connected" may include indirect "connection" and direct "connection".
[0124] The terms "about", "substantially", "approximately", and variations thereof are intended to include the degree of error associated with measuring a particular quantity based on the equipment available at the time of filing the application. For example, "about" may include a range of ±8% or 5% or 2% of a given value.
[0125] For the sake of brevity, conventional techniques related to the manufacture and use of aspects may or may not be described in detail herein. In particular, aspects of the computing systems and specific computer programs for implementing the various technical features described herein are well known. Accordingly, for the sake of brevity, many conventional implementation details are only briefly mentioned or completely omitted herein without providing well-known system and / or process details.
[0126] It should be understood that the various aspects and / or portions of the various aspects disclosed herein can be combined in combinations different from those specifically presented in the specification and the drawings. It should also be understood that, according to examples, certain acts or events of any of the processes or methods described herein can be performed in a different order, can be added, combined, or entirely omitted (e.g., all of the described acts or events may not be necessary to implement the techniques). Additionally, although for clarity some aspects of the present disclosure are described as being performed by a single module or unit, it should be understood that the techniques of the present disclosure can be performed by a combination of units or modules associated with, for example, a medical device.
[0127] In one or more examples, the described techniques can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium can include a non-transitory computer-readable medium corresponding to a tangible medium such as a data storage medium (e.g., RAM, ROM, EEPROM, flash memory, or any other medium that can be used to store the required program code in the form of instructions or data structures and that can be accessed by a computer).
[0128] The instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), graphics processing units (GPUs), microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, as used herein, the term "processor" can refer to any of the foregoing structures or any other physical structure suitable for implementing the described techniques. Additionally, the techniques can be implemented entirely in one or more circuits or logic elements.
[0129] Although the invention has been described with reference to aspects, those of ordinary skill in the art will understand that various changes can be made and equivalents can be substituted for its elements without departing from the scope of the invention. Additionally, aspects or portions of aspects can be combined, in whole or in part, without departing from the scope of the invention. Further, many modifications can be made to adapt a particular situation or material to the teachings of the invention without departing from the scope of the invention. Accordingly, the invention is not intended to be limited to the specific aspects disclosed as contemplated for implementing the invention, but the invention will include all aspects falling within the scope of the appended claims. Additionally, unless otherwise specified, any use of the terms first, second, etc. does not denote any order or importance, but the terms first, second, etc. are used to distinguish one element from another.
Claims
1. A computer-implemented method for hierarchical segmentation of video frames in a surgical video, the method comprising: Obtaining an image of an anatomical structure, wherein the image includes a plurality of image pixels; Generating a multi-label probability map for each node in a predefined segmentation class hierarchy; Processing the plurality of image pixels to generate a leaf-level segmentation map; and Processing each leaf-level segmentation and updating the class label of each leaf-level segmentation to a higher parent class until a prediction confidence threshold is reached.
2. The computer-implemented method according to claim 1, wherein, Obtaining the image includes obtaining a video stream from one or more of a camera external to the patient's body, an endoscopic camera, and a laparoscopic camera.
3. The computer-implemented method according to any one of claims 1 or 2, wherein, Generating the multi-label probability map includes mapping each of the plurality of image pixels to a class at each hierarchy level.
4. The computer-implemented method according to any one of claims 1, 2 or 3, wherein, Processing the plurality of image pixels to generate the leaf-level segmentation map includes performing root-to-leaf and inference on each of the plurality of image pixels.
5. The computer-implemented method according to any one of claims 1, 2, 3 or 4, wherein, Processing each leaf-level segmentation includes determining a prediction confidence threshold level using the following formula: where H and W are the dimensions of the prediction map P L and iI11-dimensional meters {im≥T, i∈1,, N}.
6. The computer-implemented method according to any one of claims 1 to 5, wherein, Processing each leaf-level segmentation includes performing a processing inference method to update the class label in a fine-grained leaf-level prediction map to a more general class label based on the prediction confidence threshold.
7. The computer-implemented method according to any one of the preceding claims, wherein, Processing each leaf-level segmentation includes repeatedly updating each leaf-level segmentation until the prediction confidence threshold is reached.
8. A system, comprising: A data memory that includes video data associated with a surgical procedure; And A machine learning training system for training a hierarchical model to perform hierarchical segmentation of video frames in a surgical video, the system being configured to: Obtain an image of an anatomical structure from the video data, wherein the image includes a plurality of image pixels; Generate a multi-label probability map for each node in a predefined segmentation class hierarchy; Process the plurality of image pixels to generate a leaf-level segmentation map; And Process each leaf-level segmentation and update the class label of each leaf-level segmentation to a higher parent class until a predetermined prediction confidence threshold is reached.
9. The system according to claim 8, wherein, The system is configured to obtain an image using at least one of a camera external to the patient's body, an endoscopic camera, and a laparoscopic camera.
10. The system according to claim 8 or 9, wherein, The system is configured to generate the multi-label probability map by mapping each of the plurality of image pixels to a class at each hierarchy level.
11. The system according to any one of claims 8, 9 or 10, wherein, The system is configured to process the plurality of image pixels by performing root-to-leaf and inference on each of the plurality of image pixels.
12. The system according to any one of claims 8 to 11, wherein, The system is configured to process each leaf-level segmentation to determine a prediction confidence threshold level based on the following formula: where H and W are the dimensions of the prediction map P L and 1 = maX{im; T, i ∈ 1,, N}.
13. A computer program product, comprising a memory device having computer-executable instructions stored thereon that, when executed by one or more processors, cause the one or more processors to perform a plurality of operations for hierarchical segmentation of video frames in a surgical video, the plurality of operations including the following: Obtaining an image of an anatomical structure, wherein The image includes a plurality of image pixels; Generating a multi-label probability map for two or more nodes in a predefined segmentation class hierarchy; Processing the plurality of image pixels to generate a leaf-level segmentation map; And Updating the class label of at least one leaf-level segmentation to a higher parent class until a predetermined prediction confidence threshold is reached.
14. The computer program product according to claim 13, wherein The one or more processors are configured to obtain an image using at least one of a camera, an endoscopic camera, and a laparoscopic camera disposed outside of a patient's body.
15. The computer program product according to claim 13 or 14, wherein The one or more processors are configured to perform at least one of the following: generate a multi-label probability map by mapping one or more of the plurality of image pixels to classes at each hierarchical level, process the plurality of image pixels by performing root-to-leaf and inference on one or more of the plurality of image pixels, and process each leaf-level segmentation to determine a prediction confidence threshold level using dimensions of the prediction map.