Surgical robotic system for unobstructive display of intraoperative ultrasound images as an overlay
The surgical robotic system uses a machine learning algorithm to dynamically position ultrasound images as an overlay on the laparoscopic camera feed, addressing workflow disruptions and tissue obstruction issues by enhancing visibility during surgeries.
Patent Information
- Application Number
- PCT/IB2025/051908
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2025-02-21
- Publication Date
- 2025-09-04
AI Technical Summary
In surgical robotic systems, the display of intraoperative ultrasound images as a secondary feed causes inconvenience and disruption to the surgical workflow due to the need for the surgeon to turn away from the primary display, and static overlay placement can lead to potential tissue damage from visual obstruction.
A surgical robotic system with a machine learning algorithm that segments surgical instruments from tissue, dynamically updating the position of the ultrasound image feed as an overlay on the laparoscopic camera image to minimize tissue obstruction, using a convolutional neural network for image segmentation and tracking.
The system ensures continuous visibility of the surgical site by minimizing visual obstruction, allowing surgeons to maintain focus on the primary display without accidentally damaging healthy tissue.
Smart Images

Figure IB2025051908_04092025_PF_FP_ABST
Abstract
Description
SURGICAL ROBOTIC SYSTEM FOR UNOBSTRUCTIVE DISPLAY OF INTRAOPERATIVE ULTRASOUND IMAGES AS AN OVERLAYCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 559,428, filed February 29, 2024, the entire content of which is incorporated herein by reference.BACKGROUND
[0002] Surgical robotic systems are currently being used in a variety of medical procedures, including minimally invasive surgical procedures. Some surgical robotic systems include a surgeon console controlling a surgical robotic arm and a surgical instrument having an end effector (e.g., forceps or grasping instrument) coupled to and actuated by the robotic arm. In operation, the robotic arm is moved to a position over a patient and then guides the surgical instrument into a small incision via a surgical port or a natural orifice of a patient to position the end effector at a work site within the patient’s body. A variety of instrument types are used with surgical robotic systems that are designed to perform specific functions, such as ultrasound probes, which may be used to generate intraoperative ultrasound images for use by the surgeon.
[0003] Intraoperative ultrasound (U / S) is commonly used in robotic -assisted partial nephrectomy. For example, a surgeon may use a first display including a 3D image of the surgical site provided by a laparoscopic, steresocopic camera, and a secondary display including a ultrasound image. However, because the ultrasound image is generally displayed on the secondary display, the surgeon has to turn away from the main display to view the ultrasound image, causing inconvenience and disruption to the workflow during a surgical procedure.SUMMARY
[0004] The present disclosure provides for a surgical robotic system having one or more robotic arms each having an instrument, a laparoscopic camera, and a laparoscopic ultrasound probe for providing intraoperative images. The intraoperative images are processed by an image processing device to generate an image feed, which is displayed on a display of a surgeon console . The surgeon console further includes hand controllers for manipulating the instrument, the laparoscopic camera, and the laparoscopic ultrasound probe.
[0005] A “picture-in-picture” mode has been prototyped to integrate the image feed with the primary live laparoscopic view. However, the image feed is statically located on a portionof the display (e.g., a left bottom comer) throughout the procedure based on a user’s setting, which may cause the portion of the display (e.g., display of a surgical region) to be partially unobservable, resulting a potential risk of accidentally damaging healthy tissue.
[0006] A method is provided herein for display of an image feed as an overlay on a laparaoscopic camera image on a surgical display, which minimizes visual obstruction of tissue on a surgical display by one or more surgical instruments. First, a machine learning algorithm is trained to segment the one or more surgical instruments from the tissue on the surgical display (e.g., highlight in blue). Second, the displayed image feed is attached to the shaft of one of the surgical instruments. As this surgical instrument moves, a location of the image feed dynamically changes to maximize overlap (e.g., occupation of pixels) with the instrument shaft while minimizing overlap with the tissue.
[0007] According to one aspect of the present disclosure, a surgical system is disclosed. The surgical system includes a laparoscopic camera, an image processing device, a surgeon console, a processor, and a memory. The laparoscopic camera is configured to capture a video feed of tissue and a surgical instrument. The image processing device is configured to receive the video feed and ultrasound data of the tissue and generate an image feed. The surgeon console includes a display for displaying a graphical user interface for outputting the image feed and the video feed. The memory includes instructions stored thereon, which when executed by the processor cause the system to: receive the image feed and the video feed; determine, using a machine learning model, a position of the surgical instrument in the video feed based on image segmentation; track the position of the surgical instrument in the video feed; and display the video feed and the image feed on the graphical user interface. The image feed is displayed as an overlay on a portion of the video feed.
[0008] In another aspect of the present disclosure, the image segmentation may create at least one segmentation mask.
[0009] In yet another aspect of the present disclosure, the at least one segmentation mask may represent the surgical instrument or the tissue.
[0010] In a further aspect of the present disclosure, the machine learning model may include a convolutional neural network.
[0011] In yet a further aspect of the present disclosure, the instructions, when executed by the processor, may further cause the system to dynamically update a position of the image feed based on the tracked position of the surgical instrument. The image feed may replace a portion of the surgical instrument on the video feed.
[0012] In another aspect of the present disclosure, the image feed replaces a shaft of the surgical instrument on the video feed.
[0013] In yet another aspect of the present disclosure, the instructions, when executed by the processor, may further cause the system to: determine, using the machine learning model, a position of a tissue segment corresponding to a least observable segment of the tissue; and track the position of the tissue segment.
[0014] In a further aspect of the present disclosure, the instructions, when executed by the processor, may further cause the system to dynamically update a position of the image feed based on the tracked position of the tissue segment. The image feed replaces a portion of the tissue segment on the video feed.
[0015] In yet a further aspect of the present disclosure, the graphical user interface may include a user interface element for adjusting a view of the image feed.
[0016] In another aspect of the present disclosure, the user interface elements may be configured to adjust a tracking mode, a transparency, and / or a size of the ultrasound image.
[0017] In yet another aspect of the present disclosure, the surgical system may further include an ultrasound probe configured to obtain the ultrasound data of the tissue.
[0018] In a further aspect of the present disclosure, the image feed may include an ultrasound image feed generated by the image processing device based on the received ultrasound data of the tissue.
[0019] In yet a further aspect the present disclosure, the image feed may include an ultrasound image, a pre-operative magnetic resonance image (MRI), a computerized tomography (CT) scan, a three-dimensional reconstructed organ model image, and / or an optical coherence tomography (OCT) image.BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Various embodiments of the present disclosure are described herein with reference to the drawings, wherein:
[0021] FIG. 1 is a perspective view of a surgical robotic system including a control tower, a console, and one or more surgical robotic arms, each disposed on a movable cart according to an aspect of the present disclosure;
[0022] FIG. 2 is a perspective view of a surgical robotic arm of the surgical robotic system of FIG. 1 according to an aspect of the present disclosure;
[0023] FIG. 3 is a perspective view of a movable cart having a setup arm with the surgical robotic arm of the surgical robotic system of FIG. 1 according to an aspect of the present disclosure;
[0024] FIG. 4 is a schematic diagram of a computer architecture of the surgical robotic system of FIG. 1 according to an aspect of the present disclosure;
[0025] FIG. 5 is a plan schematic view of movable carts of FIG. 1 positioned about a surgical table according to an aspect of the present disclosure;
[0026] FIG. 6 is a side view of an ultrasound probe for use with the surgical robotic system according to an aspect of the present disclosure;
[0027] FIG. 7 is a perspective view of the ultrasound probe being gripped by a gripper surgical instrument according to an aspect of the present disclosure;
[0028] FIG. 8 shows a schematic diagram of the ultrasound probe and a laparoscopic camera according to an aspect of the present disclosure;
[0029] FIG. 9 is a flow chart of a method for display of an ultrasound image as overlay on a laparoscopic camera image on a surgical display according to an embodiment of the disclosure;
[0030] FIG. 10 shows a laparoscopic camera image with a statically placed image feed to according to an aspect of the present disclosure;
[0031] FIG. 11 shows a laparoscopic camera image with a dynamically placed image feed according to an aspect of the present disclosure;
[0032] FIG. 12 is a schematic diagram of a machine learning processing system according to an embodiment of the present disclosure.DETAILED DESCRIPTION
[0033] Embodiments of the presently disclosed surgical robotic system are described in detail with reference to the drawings, in which like reference numerals designate identical or corresponding elements in each of the several views.
[0034] With reference to FIG. 1, a surgical robotic system 10 includes a control tower 20, which is communicatively coupled to all of the components of the surgical robotic system 10 including a surgeon console 30 and one or more movable carts 60. Each of the movable carts 60 includes a robotic arm 40 having a surgical instrument 50 coupled thereto. The robotic arms 40 also couple to the movable carts 60. The robotic system 10 may include any number of movable carts 60 and / or robotic arms 40.
[0035] The surgical instrument 50 is configured for use during minimally invasive surgical procedures. In embodiments, the surgical instrument 50 may be configured for open surgical procedures. One of the robotic arms 40 may include a laparoscopic camera 51 configured to capture video of the surgical site. The laparoscopic camera 51 may be a stereoscopic endoscope configured to capture two side-by-side (i.e., left, and right) images of the surgical site to produce a video stream of the surgical scene. The laparoscopic camera 51 is coupled to an image processing device 56, which may be disposed within the control tower 20. The image processing device 56 may be any computing device as described below configured to receive the video feed from the laparoscopic camera 51 and output the processed video stream.
[0036] The surgeon console 30 includes a first display 32, which displays a video feed of the surgical site provided by a camera 51 disposed on the robotic arm 40, and a second display 34, which displays a user interface for controlling the surgical robotic system 10. The first display 32 and the second display 34 may be touchscreens allowing for displaying various graphical user inputs selectable or movable by the user.
[0037] The surgeon console 30 also includes a plurality of user interface devices, such as foot pedals 36 and a pair of hand controllers 38a and 38b, which are used by a user to remotely control the robotic arms 40. The surgeon console further includes an armrest 33 used to support clinician’s arms while the clinician is operating the hand controllers 38a and 38b.
[0038] The control tower 20 can also include a display 23, which may be a touchscreen, and outputs on the graphical user interfaces (GUIs). The control tower 20 also acts as an interface between the surgeon console 30 and one or more of the robotic arms 40. In particular, the control tower 20 is configured to control the robotic arms 40, such as to move the robotic arms 40 and the corresponding surgical instrument 50, based on a set of programmable instructions and / or input commands from the surgeon console 30. In response to the instructions and / or input, the robotic arms 40 and the surgical instrument 50 execute a desired movement sequence in response to input from the foot pedals 36 and the hand controllers 38a and 38b. The system 10 can be configured so that the foot pedals 36 may be used to affect one or more of a wide variety of system functions, such as to enable and lock the hand controllers 38a and 38b, reposition camera movement, and activate / deactivate an electrosurgical instrument. In particular, the foot pedals 36 may be used to perform a clutching action on the hand controllers 38a and 38b. Clutching is initiated by pressing one of the foot pedals 36, which disconnects (i.e., prevents movement inputs from) the hand controllers 38a and / or 38b such that the robotic arm 40 and corresponding instrument 50 or camera 51 are not actuated. This allows the user to reposition the hand controllers 38a and 38b without moving the roboticarm(s) 40 and the instrument 50 and / or camera 51. This is useful when reaching control boundaries of the surgical space, for instance.
[0039] Each of the control tower 20, the surgeon console 30, and the robotic arm 40 includes a respective computer 21, 31, 41. The computers 21 , 31 , 41 are interconnected to each other using any suitable communication network based on wired or wireless communication protocols. The term “network,” whether plural or singular, as used herein, denotes a data network, including, but not limited to, the Internet, Intranet, a wide area network, or a local area network. Suitable protocols include, but are not limited to, transmission control protocol / intemet protocol (TCP / IP), datagram protocol / intemet protocol (UDP / IP), and / or datagram congestion control protocol (DCCP). Wireless communication may be achieved via one or more wireless configurations, e.g., radio frequency (RF), optical, Wi-Fi, Bluetooth (an open wireless protocol for exchanging data over short distances, using short-length radio waves, from fixed and mobile devices, creating personal area networks (PANs), ZigBee® (a specification for a suite of high-level communication protocols using small, low-power digital radios based on the IEEE 122.15.4-1203 standard for wireless personal area networks (WPANs)).
[0040] The computers 21, 31, 41 may include any suitable processor (not shown) connected operably to a memory (not shown), which may include one or more of volatile, nonvolatile, magnetic, optical, or electrical media, such as read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM (EEPROM), non-volatile RAM (NVRAM), or flash memory. The processor may be any suitable processor (e.g., control circuit) adapted to perform the operations, calculations, and / or set of instructions described in the present disclosure, such as a hardware processor, a field programmable gate array (FPGA), a digital signal processor (DSP), a central processing unit (CPU), a microprocessor, and combinations thereof. Those skilled in the art will appreciate that the processor may be substituted for by using any logic processor (e.g., control circuit) adapted to execute algorithms, calculations, and / or set of instructions described herein.
[0041] With reference to FIG. 2, each of the robotic arms 40 may include a plurality of links 42a, 42b, 42c, which are interconnected at joints 44b and 44c, respectively. Other configurations of links and joints may be utilized as known by those skilled in the art. The joint 44a is configured to secure the robotic arm 40 to the movable cart 60 and defines a first longitudinal axis. With reference to FIG. 3, the movable cart 60 includes a lift 67 and a setup arm 61, which provides a base for mounting of the robotic arm 40. The lift 67 allows for vertical movement of the setup arm 61 and, thereby, of the robotic arms 40 mounted on the setup arm61. The movable cart 60 also includes a display 69 for displaying information pertaining to the robotic arm 40. In embodiments, the robotic arms 40 may include any type and / or number of joints.
[0042] With further reference to FIG. 3, the setup arm 61 includes a first link 62a, a second link 62b, and a third link 62c, which provide for lateral maneuverability of the robotic arms 40. The links 62a, 62b, 62c are interconnected at joints 63a and 63b, each of which may include an actuator (not shown) for rotating the links 62b and 62b relative to each other and the link 62c. In particular, the links 62a, 62b, 62c are movable in corresponding lateral planes, which are parallel to each other, thereby allowing for extension of the robotic arm 40 relative to the patient (e.g., surgical table). In embodiments, the robotic arm 40 may be coupled to the surgical table (not shown). The setup arm 61 includes controls 65 for adjusting movement of the links 62a, 62b, 62c as well as the lift 67. In embodiments, the setup arm 61 may include any type and / or number of joints.
[0043] The third link 62c may include a rotatable base 64 having two degrees of freedom. In particular, the rotatable base 64 includes a first actuator 64a and a second actuator 64b. The first actuator 64a is rotatable about a first stationary arm axis, which is perpendicular to a plane defined by the third link 62c. And the second actuator 64b is rotatable about a second stationary arm axis which is transverse to the first stationary arm axis. The first and second actuators 64a and 64b along with the lift 67 allow for full three-dimensional orientation of the robotic arm 40.
[0044] Returning to FIG. 2, the actuator 48b of the joint 44b is coupled to the joint 44c via the belt 45a, and the joint 44c is in turn coupled to the joint 46b via the belt 45b. Joint 44c may include a transfer case coupling the belts 45a and 45b, such that the actuator 48b is configured to rotate each of the links 42b, 42c and a holder 46 relative to each other. More specifically, links 42b, 42c, and the holder 46 are passively coupled to the actuator 48b which enforces rotation about a pivot point “P” which lies at an intersection of the first axis defined by the link 42a and the second axis defined by the holder 46. In other words, the pivot point “P” is a remote center of motion (RCM) for the robotic arm 40. Thus, the actuator 48b controls the angle 0 between the first and second axes allowing for orientation of the surgical instrument 50. Due to the interlinking of the links 42a, 42b, 42c, and the holder 46 via the belts 45a and 45b, the angles between the links 42a, 42b, 42c, and the holder 46 are also adjusted in order to achieve the desired angle 0. In embodiments, some or all of the joints 44a, 44b, 44c may include an actuator to obviate the need for mechanical linkages.
[0045] The joints 44a and 44b include respective actuators 48a and 48b configured to drive the joints 44a, 44b, 44c relative to each other through a series of belts 45a and 45b or other mechanical linkages such as a drive rod, a cable, or a lever and the like. In particular, the actuator 48a is configured to rotate the robotic arm 40 about a longitudinal axis defined by the link 42a.
[0046] With reference to FIG. 2, the holder 46 defines a second longitudinal axis and configured to receive an instrument drive unit (IDU) 52 (FIG. 1). The IDU 52 is configured to couple to an actuation mechanism of the surgical instrument 50 and the camera 51 and is configured to move (e.g., rotate) and actuate the instrument 50 and / or the camera 51. IDU 52 transfers actuation forces from its actuators to the surgical instrument 50 to actuate components of an end effector 49 of the surgical instrument 50. The holder 46 includes a sliding mechanism 46a, which is configured to move the IDU 52 along the second longitudinal axis defined by the holder 46. The holder 46 also includes a joint 46b, which rotates the holder 46 relative to the link 42c. During endoscopic procedures, the instrument 50 may be inserted through an endoscopic access port 55 (FIG. 3) held by the holder 46. The holder 46 also includes a port latch 46c for securing the access port 55 to the holder 46 (FIG. 2).
[0047] The IDU 52 is attached to the holder 46, followed by a sterile interface module (SIM) 43 being attached to a distal portion of the IDU 52. The SIM 43 is configured to secure a sterile drape (not shown) to the IDU 52. The instrument 50 is then attached to the SIM 43. The instrument 50 is then inserted through the access port 55 by moving the IDU 52 along the holder 46. The SIM 43 includes a plurality of drive shafts configured to transmit rotation of individual motors of the IDU 52 to the instrument 50 thereby actuating the instrument 50. In addition, the SIM 43 provides a sterile barrier between the instrument 50 and the other components of the robotic arm 40, including the IDU 52.
[0048] The robotic arm 40 also includes a plurality of manual override buttons 53 (FIG. 1) disposed on the IDU 52 and the setup arm 61, which may be used in a manual mode. The user may press one or more of the buttons 53 to move the component associated with the button 53.
[0049] With reference to FIG. 4, each of the computers 21, 31, 41 of the surgical robotic system 10 may include a plurality of controllers, which may be embodied in hardware and / or software. The computer 21 of the control tower 20 includes a controller 21a and safety observer 21b. The controller 21a receives data from the computer 31 of the surgeon console 30 about the current position and / or orientation of the hand controllers 38a and 38b and the state of the foot pedals 36 and other buttons. The controller 21a processes these input positions to determine desired drive commands for each joint of the robotic arm 40 and / or the IDU 52 andcommunicates these to the computer 41 of the robotic arm 40. The controller 21a also receives the actual joint angles measured by encoders of the actuators 48a and 48b and uses this information to determine force feedback commands that are transmitted back to the computer 31 of the surgeon console 30 to provide haptic feedback through the hand controllers 38a and 38b. The safety observer 21b performs validity checks on the data going into and out of the controller 21a and notifies a system fault handler if errors in the data transmission are detected to place the computer 21 and / or the surgical robotic system 10 into a safe state.
[0050] The computer 41 includes a plurality of controllers, namely, a main cart controller 41a, a setup arm controller 41b, a robotic arm controller 41c, and an instrument drive unit (IDU) controller 4 Id. The main cart controller 4 la receives and processes joint commands from the controller 21a of the computer 21 and communicates them to the setup arm controller 41b, the robotic arm controller 41c, and the IDU controller 41d. The main cart controller 41a also manages instrument exchanges and the overall state of the movable cart 60, the robotic arm 40, and the IDU 52. The main cart controller 41a also communicates actual joint angles back to the controller 21a.
[0051] Each of joints 63a and 63b and the rotatable base 64 of the setup arm 61 are passive joints (i.e., no actuators are present therein) allowing for manual adjustment thereof by a user. The joints 63a and 63b and the rotatable base 64 include brakes that are disengaged by the user to configure the setup arm 61. The setup arm controller 41b monitors slippage of each of joints 63a and 63b and the rotatable base 64 of the setup arm 61, when brakes are engaged or can be freely moved by the operator when brakes are disengaged, but do not impact controls of other joints. The robotic arm controller 41c controls each joint 44a and 44b of the robotic arm 40 and calculates desired motor torques required for gravity compensation, friction compensation, and closed loop position control of the robotic arm 40. The robotic arm controller 41c calculates a movement command based on the calculated torque. The calculated motor commands are then communicated to one or more of the actuators 48a and 48b in the robotic arm 40. The actual joint positions are then transmitted by the actuators 48a and 48b back to the robotic arm controller 41c.
[0052] The IDU controller 4 Id receives desired joint angles for the surgical instrument 50, such as wrist and jaw angles, and computes desired currents for the motors in the IDU 52. The IDU controller 4 Id calculates actual angles based on the motor positions and transmits the actual angles back to the main cart controller 41a.
[0053] The robotic arm 40 is controlled in response to a pose of the hand controller controlling the robotic arm 40, e.g., the hand controller 38a, which is transformed into a desiredpose of the robotic arm 40 through a hand eye transform function executed by the controller 21a. The hand eye function, as well as other functions described herein, is / are embodied in software executable by the controller 21a or any other suitable controller described herein. The pose of one of the hand controllers 38a may be embodied as a coordinate position and rollpitch-yaw (RPY) orientation relative to a coordinate reference frame, which is fixed to the surgeon console 30. The desired pose of the instrument 50 is relative to a fixed frame on the robotic arm 40. The pose of the hand controller 38a is then scaled by a scaling function executed by the controller 21a. In embodiments, the coordinate position may be scaled down and the orientation may be scaled up by the scaling function. In addition, the controller 21a may also execute a clutching function, which disengages the hand controller 38a from the robotic arm 40. In particular, the controller 21a stops transmitting movement commands from the hand controller 38a to the robotic arm 40 if certain movement limits or other thresholds are exceeded and in essence acts like a virtual clutch mechanism, e.g., limits mechanical input from effecting mechanical output.
[0054] The desired pose of the robotic arm 40 is based on the pose of the hand controller 38a and is then passed by an inverse kinematics function executed by the controller 21a. The inverse kinematics function calculates angles for the joints 44a, 44b, 44c of the robotic arm 40 that achieve the scaled and adjusted pose input by the hand controller 38a. The calculated angles are then passed to the robotic arm controller 41c, which includes a joint axis controller having a proportional-derivative (PD) controller, the friction estimator module, the gravity compensator module, and a two-sided saturation block, which is configured to limit the commanded torque of the motors of the joints 44a, 44b, 44c.
[0055] With reference to FIG. 5, the surgical robotic system 10 is set up around a surgical table 90. The system 10 includes movable carts 60a-d, which may be numbered “1” through “4.” During setup, each of the carts 60a-d are positioned around the surgical table 90. Position and orientation of the carts 60a-d depends on a plurality of factors, such as placement of a plurality of access ports 55a-d, which in turn, depends on the surgery being performed. Once the port placements are determined, the access ports 55a-d are inserted into the patient, and carts 60a-d are positioned to insert instruments 50 and the laparoscopic camera 51 into corresponding ports 55a-d.
[0056] During use, each of the robotic arms 40a-d is attached to one of the access ports 55a-d that is inserted into the patient by attaching the latch 46c (FIG. 2) to the access port 55 (FIG. 3). The IDU 52 is attached to the holder 46, followed by the SIM 43 being attached to a distal portion of the IDU 52. Thereafter, the instrument 50 is attached to the SIM 43. Theinstrument 50 is then inserted through the access port 55 by moving the IDU 52 along the holder 46.
[0057] FIG. 6 shows an ultrasound probe 170, which may be a drop-in, laparoscopic ultrasound probe. The ultrasound probe 170 is configured to generate ultrasound data by emitting acoustic waves into the tissue and to receive the reflected acoustic waves, i.e., ultrasound data. The received reflected acoustic waves are then processed by the image processing device 56 to identify various properties of the tissues through which the acoustic wave traveled, such as a density of the tissue. The laparoscopic camera 51 is coupled to the image processing device 56, which may include any computing device configured to receive the ultrasound data from the ultrasound probe 170 and to generate ultrasound images.
[0058] The ultrasound probe 170 includes a shaft 172 and an end effector 174 including an ultrasound transducer 176 disposed at a distal end portion of the shaft 172. The end effector 174 may be articulated relative to the shaft 172 (e.g., about pivot pin). Any suitable articulating mechanism used in surgical instruments may be used. The ultrasound probe 170 may be a drop- in ultrasound probe that is configured to be grasped at a protrusion 178 by the instrument 50 as shown in FIGS. 6 and 7. In embodiments, the ultrasound probe 170 may be configured as an instrument attachable to the IDU 52 and controllable directly through the surgeon console 30 in the same manner as any of the instruments 50 and / or camera 51. FIG. 8 shows such an embodiment of the ultrasound probe 170, which is inserted through the first access port 55a. The camera 51 is inserted through the second access port 55b such that the ultrasound transducer 176 is within the field of view of the camera 51 and is visible on the display screen 32. In embodiments where the ultrasound probe 170 is a drop-in probe, the instrument 50 is inserted through the third access port 55c and is used to maneuver the ultrasound probe 170 in contact and over tissue to obtain ultrasound images. The robotic control of the ultrasound probe 170 described below via the system 10 applies to both embodiments, where the ultrasound probe 170 is directly controllable by the system 10 (i.e., via the IDU 52) or grasped by the instrument 50.
[0059] The system 10 is initially configured for ultrasound operation by inserting the ultrasound probe 170 into the body cavity and navigating the ultrasound probe 170 to the tissue that is to be scanned by ultrasound. Navigation may be done by grasping the ultrasound sound probe 170 using the instrument 50 and / or controlling the ultrasound probe 170 directly, depending on the type of the ultrasound probe 170. While movement of the ultrasound probe 170 is described below using only the instrument 50, the same control commands may be provided to the IDU 52 and the robotic arm 40 controlling the instrument 50 to control themovement of ultrasound probe 170 directly. In addition, the camera 51 is also positioned within the body cavity and pointed towards the surgical site such that the ultrasound probe 170 and instrument 50 and their respective end effectors are within the field of view of the camera 51.
[0060] FIG. 9 shows a method 400 for display of an ultrasound image (e.g., ultrasound feed “UF”) as an overlay on a laparaoscopic camera image (e.g., a video feed “VF”) on a surgical display, such as the first display screen 32 of the surgeon console 30. The method may be implemented as software instructions (e.g., obstruction avoidance software) stored in amemory and executable by a processor (e.g., controller 21a, image processing device 56, etc.). The obstruction avoidance software may be initiated either manually or automatically. For example, a GUI may be displayed on one of the display screens 32 and / or 34, which may be used to manually start the obstruction avoidance software. The obstruction avoidance software may be started automatically in response to occurrence of a triggering event, e.g., insertion of the ultrasound probe 170 through the access port 55a, or a phase, which may be detected using a machine learning (ML) processing system 310 of FIG. 12.
[0061] At step 402, the ultrasound data and the video feed “VF” are received by the image processing device 56. While the ultrasound probe 170 is moved into contact with tissue, the image processing device 56 receives and processes the ultrasound data from the ultrasound probe 170 to generate an ultrasound image feed “UF.” The image feed “IF” may be dynamically updated as the ultrasound probe 170 captures data throughout the surgical procedure. In aspects, various parameters, e.g., a playback speed, ultrasound depth, etc. of the image feed “IF” may be adjusted by the clinician. At the same time, the image processing device 56 receives and processes the video feed “VF” from the laparoscopic camera 51 of the ultrasound probe 170 contacting the tissue. Generally, the video feed “VF” includes a view of the surgical site, including images of the tissue and the surgical instrument(s) 50 being used, as shown in FIG. 10.
[0062] At step 404, the video feed “VF” is analyzed to determine a position of one or more surgical instruments 50 in the video feed “VF”. A machine learning model may be used to perform image segmentation on the video feed “VF.” In aspects, the machine learning model may be a convolutional neural network (CNN) employing a classification system to classify various portions of the video feed. The image segmentation process may create one or more segmentation masks representative of different objects within the video feed. For example, a first segmentation mask “SMI” may be representative of the instrument(s) 50 within the video feed “VF”, and a second segmentation mask SM2 may be representative of tissue within the video feed “VF”, as shown in FIG. 11. To create the segmentation masks, the machine learningmodel may identify a first set of pixels in the video feed “VF” that correspond to the instrument(s) 50, and a second set of pixels in the video feed “VF” that correspond to the tissue. The first set of pixels and the second set of pixels may then be used by the image segmentation process to create the first segmentation mask “SMI” and the second segmentation mask “SM2,” respectively. The image processing device 56 may then associate the pixels of the first segmentation mask “SMI” with the instrument 50, and the pixels of the second segmentation mask “SM2” with the tissue, to identify the instrument 50 and the tissue within the video feed at any given time during the surgical procedure.
[0063] At step 406, the processor causes the image processing device 56 to employ image tracking, which tracks a position of an object in the video feed “VF”. The image processing device 56 may track a position(s) of the instrument(s) 50. The image tracking process monitors the movement of pixels representative of the first, (i.e., the instrument) segmentation mask “SMI,” and updates a position(s) of the instrument(s) 50 in memory. In aspects, the image processing device 56 may be configured to track a central point of the instrument(s) 50. The central point of the instrument(s) 50 may be calculated using a cluster of coordinates. The image processing device 56 may average the coordinates to determine the central point. In another example, the image processing device 56 may track a position of the tissue, such as a least observable segment of tissue within the surgical field. This approach may be useful when the instruments 50 move too rapidly, or where the clinician desires a more fixed point without obstructing their view of the site. It will be understood that step 406 may be repeated throughout the process as a position of the tracked object updates.
[0064] At step 408, the video feed “VF’ and an image feed “IF” are displayed on a GUI on one of the display screens 32 and / or 34. The image feed “IF” may include an ultrasound image, a pre-operative magnetic resonance image (MRI), a computerized tomography (CT) scan, a three-dimensional reconstructed organ model image, and / or an optical coherence tomography (OCT) image. In doing so, the image feed “IF” is displayed as an overlay on an optimal portion of the video feed in a manner intended to minimize visual obstruction of the surgical site, e.g., the tissue of the patient. As used herein, an optimal portion of the video feed includes a portion that is least critical and / or least observable at a particular stage of surgery. A least critical portion of the video feed may include a portion of the video feed which does not occlude tissue or anatomy (e.g., tissue) critical to the surgical procedure. A least observable portion of the video feed may include a portion of the video feed that contains a visual obstruction. A visual obstruction includes a visual blockage (e.g., clarity, lighting, size, etc.) and / or object (e.g., asurgical instrument, surgeon hand, etc.), which blocks visibility of a portion of the tissue within a surgical site. Examples of both approaches are described further below.
[0065] For example, to overlay a least critical portion of the video feed, the image feed “IF” may be overlaid on a portion of the video feed closest to the surgical site (e.g., near a tumor). This approach may ease the task of mental registration and / or the ability of the surgeon to relate an ultrasound image coordination system to the target anatomy of the surgical site (e.g., target anatomy as viewed from the endoscope). In addition, the surgeon may be able to adjust attributes of the image feed “IF” such as size and / or transform the image feed “IF” (e.g., rotation, mirroring, shearing) to enable ease of interpretation. Such adjustments may be performed using user interface elements on a GUI, as discussed further below.
[0066] In another example, to overlay the least observable portion of the video feed, the image feed “IF” may be overlaid on a non-working portion of the instrument 50 (FIG. 11), such as a shaft of the instrument 50 distal of the end effector of the instrument, which forms a working portion of the instrument 50. As the instrument 50 is already blocking visualization of a portion of the tissue, overlaying the image feed “IF” over the instrument will effectively minimize obstruction of visible tissue. The image processing device 56 may replace a portion of the pixels in the first segmentation mask “SMI” with the image feed “IF”, maximizing overlap of the image feed “IF” with the instrument 50. As the tracked position(s) of the instrument(s) 50 updates, the position of the image feed “IF” will synchronously update, thereby continuing to overlap with the instrument 50. The image feed “IF” and / or the video feed “VF” may display details related to the surgical procedure, such a depth of instrument 50, a surgical path, keep-out zones, etc.
[0067] In another example, to overlay the least observable portion of the video feed, the image processing device 56 may use the machine learning model to dynamically predict which portion of the tissue is least observable at any given time. Factors determining whether tissue is observable may include clarity, brightness, sharpness, secondary obstructions (e.g., clinician’s hand), etc. The image processing device 56 may replace a portion of the pixels in the second segmentation mask “SM2” with the image feed “IF”, maximizing overlap of the image feed “IF”. As the tracked position of the tissue updates, the position of the image feed “IF” will synchronously update, thereby continuing to overlap with the least observable portion of the tissue.
[0068] The GUI may include various user interface elements for adjusting display settings (not shown). For example, an ultrasound image feed element may enable a clinician to adjust a view of the image feed “IF”, including transparency, brightness, size, etc. The ultrasoundimage feed element may include an interactive icon for setting a “tracking mode,” which will allow a user to set the tracked portion of the screen which the image feed “IF” is overlaid on. A clinician may toggle between tracking the instrument 50 and a portion of the tissue, and / or may set which portion of each is used as a reference point for tracking (e.g., a central point).
[0069] With reference to FIG. 12, the surgical robotic system 10 may include a machine learning (ML) processing system 310 that processes the surgical data using one or more ML models to identify one or more features, such as surgical phase, instrument, anatomical structure, etc., in the surgical data. The ML processing system 310 includes a ML training system 325, which may be a separate device (e.g., server) that stores its output as one or more trained ML models 330. The ML models 330 are accessible by a ML execution system 340. The ML execution system 340 may be separate from the ML training system 325, namely, devices that “train” the models are separate from devices that “infer,” i.e., perform real-time processing of surgical data using the trained ML models 330.
[0070] System 10 includes a data reception system 305 that collects surgical data, including the video data and surgical instrumentation data. The data reception system 305 can include one or more devices (e.g., one or more user devices and / or servers) located within and / or associated with a surgical operating room and / or control center. The data reception system 305 can receive surgical data in real-time, i.e., as the surgical procedure is being performed.
[0071] The ML processing system 310, in some examples, may further include a data generator 315 to generate simulated surgical data, such as a set of virtual or masked images, or record the video data from the image processing device 56, to train the ML models 330 as well as other sources of data, e.g., user input, arm movement, etc. Data generator 315 can access (read / write) a data store 320 to record data, including multiple images and / or multiple videos.
[0072] The ML processing system 310 also includes a phase detector 350 that uses the ML models to identify a phase within the surgical procedure. Phase detector 350 uses a particular procedural tracking data structure 355 from a list of procedural tracking data structures. Phase detector 350 selects the procedural tracking data structure 355 based on the type of surgical procedure that is being performed. In one or more examples, the type of surgical procedure is predetermined or input by user. The procedural tracking data structure 355 identifies a set of potential phases that may correspond to a part of the specific type of surgical procedure.
[0073] In some examples, the procedural tracking data structure 355 may be a graph that includes a set of nodes and a set of edges, with each node corresponding to a potential phase. The edges may provide directional connections between nodes that indicate (via the direction) an expected order during which the phases will be encountered throughout an iteration of thesurgical procedure. The procedural tracking data structure 355 may include one or more branching nodes that feed to multiple next nodes and / or may include one or more points of divergence and / or convergence between the nodes. In some instances, a phase indicates a procedural action (e.g., surgical action) that is being performed or has been performed and / or indicates a combination of actions that have been performed. In some instances, a phase relates to a biological state of a patient undergoing a surgical procedure. For example, the biological state may indicate a complication (e.g., blood clots, clogged arteries / veins, etc.), pre-condition (e.g., lesions, polyps, etc.). In some examples, the ML models 330 are trained to detect an “abnormal condition,” such as hemorrhaging, arrhythmias, blood vessel abnormality, etc.
[0074] The phase detector 350 outputs the phase prediction associated with a portion of the video data that is analyzed by the ML processing system 310. The phase prediction is associated with the portion of the video data by identifying a start time and an end time of the portion of the video that is analyzed by the ML execution system 340. The phase prediction that is output may include an identity of a surgical phase as detected by the phase detector 350 based on the output of the ML execution system 340. Further, the phase prediction, in one or more examples, may include identities of the structures (e.g., instrument, anatomy, etc.) that are identified by the ML execution system 340 in the portion of the video that is analyzed. The phase prediction may also include a confidence score of the prediction. Other examples may include various other types of information in the phase prediction that is output.
[0075] It will be understood that various modifications may be made to the embodiments disclosed herein. Therefore, the above description should not be construed as limiting, but merely as exemplifications of various embodiments. Those skilled in the art will envision other modifications within the scope and spirit of the claims appended thereto.
Claims
WHAT IS CLAIMED IS:
1. A surgical system comprising: a laparoscopic camera configured to capture a video feed of tissue and a surgical instrument; an image processing device for receiving the video feed and ultrasound data of the tissue, the image processing device configured to generate an image feed; a surgeon console including a display for displaying a graphical user interface for outputting the image feed and the video feed; a processor; and a memory including instructions stored thereon, which when executed by the processor cause the system to: receive the image feed and the video feed; determine, using a machine learning model, a position of the surgical instrument in the video feed based on image segmentation; track the position of the surgical instrument in the video feed; and display the video feed and the image feed on the graphical user interface, wherein the image feed is displayed as an overlay on a portion of the video feed.
2. The surgical system of claim 1, wherein the image segmentation creates at least one segmentation mask.
3. The surgical system of claim 2, wherein the at least one segmentation mask represents the surgical instrument or the tissue.
4. The surgical system of claim 1, wherein the machine learning model includes a convolutional neural network.
5. The surgical system of claim 1, wherein the instructions, when executed by the processor, further cause the surgical system to: dynamically update a position of the image feed based on the tracked position of the surgical instrument, wherein the image feed replaces a portion of the surgical instrument on the video feed.
6. The surgical system of claim 5, wherein the portion of the surgical instrument on the video feed is a shaft.
7. The surgical system of claim 1, wherein the instructions, when executed by the processor, further cause the system to: determine, using the machine learning model, a position of a tissue segment corresponding to a least observable segment of the tissue; and track the position of the tissue segment.
8. The surgical system of claim 7, wherein the instructions, when executed by the processor, further cause the system to: dynamically update a position of the image feed based on the tracked position of the tissue segment, wherein the image feed replaces a portion of the tissue segment on the video feed.
9. The surgical system of claim 1, wherein the graphical user interface includes a user interface element for adjusting a view of the image feed.
10. The surgical system of claim 9, wherein the user interface element is configured to adjust at least one of a tracking mode, a transparency, or a size of the image feed.
11. The surgical system of claim 1, further comprising: an ultrasound probe configured to obtain the ultrasound data of the tissue.
12. The surgical system of claim 11, wherein the image feed includes an ultrasound image feed generated by the image processing device based on the received ultrasound data of the tissue.
13. The surgical system of claim 1, wherein the image feed is selected from the group consisting of an ultrasound image, a pre-operative magnetic resonance image (MRI), a computerized tomography (CT) scan, a three-dimensional reconstructed organ model image, and / or an optical coherence tomography (OCT) image
Citation Information
Patent Citations
Interactive user interfaces for robotic minimally invasive surgical systems
US20130245375A1
Hybrid hardware and computer vision-based tracking system and method
US20200197102A1
Composite medical imaging systems and methods
WO2020243425A1
Generating augmented visualizations of surgical sites using semantic surgical representations
WO2022195304A1