Digitalization of the operating room
Through the integration of depth cameras with surgical robot systems, 3D point cloud data is generated and analyzed, minimally invasive surgical data storage and analysis problems are solved, digital reconstruction and real-time guidance of the surgery are realized, and surgical quality and safety are improved.
Patent Information
- Application Number
- CN202080101730.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-05
- Filing Date
- 2020-06-16
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-06-16
AI Technical Summary
The prior art is difficult to effectively store and analyze large amounts of data in minimally invasive surgical procedures, making it difficult to achieve digital playback, offline analysis and intraoperative guidance of the surgery.
Through the integration of depth cameras with surgical robot systems, 3D point cloud data is generated and fused with robot system data, key information is stored using semantic segmentation technology, object identification and analysis is combined with machine learning algorithms, and digital assets are generated to support surgical replay and intraoperative warnings.
Digital reconstruction of surgical procedures has been achieved, supporting preoperative planning, postoperative analysis and intraoperative guidance, and improving the quality and safety of the surgery.
Smart Images

Figure CN115699198B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to surgical robotic systems and, more particularly, to digitizing surgical actions performed using surgical robotic systems. Background Art
[0002] Minimally invasive surgery (MIS) such as laparoscopic surgery involves techniques designed to reduce tissue damage during a surgical procedure. For example, a laparoscopic procedure typically involves creating a plurality of small incisions in a patient's body (e.g., in the abdomen) and introducing one or more tools and at least one camera into the patient's body through the incisions. A surgical procedure can then be performed by using the introduced surgical tools, where visual assistance is provided by the camera.
[0003] Generally speaking, MIS provides multiple benefits such as reduced patient scarring, less patient pain, shorter patient recovery periods, and lower medical costs associated with patient recovery. MIS can be performed using a surgical robotic system that includes one or more robotic arms configured to manipulate surgical tools based on commands from a remote operator. The surgical robotic arms can, for example, support various devices at their distal ends, such as surgical end effectors, imaging devices, cannulas for providing access to a patient's body cavity and organs, and the like. Control of such a robotic system may require control inputs from a user (e.g., a surgeon or other operator) via one or more user interface devices that translate manipulations or commands from the user into control of the robotic system. For example, when a surgical tool is positioned at a surgical site of a patient, a tool driver having one or more motors can actuate one or more degrees of freedom of the surgical tool in response to a user command. Thus, the surgical robotic arms can assist in performing a surgical procedure.
[0004] Performing a surgical procedure using a surgical robotic system can benefit from preoperative planning, postoperative analysis, and intraoperative guidance. A system that captures image data can allow a user to view and interact with a recorded rendering of a surgical action. The captured data can be used for preoperative planning, postoperative analysis, and intraoperative guidance. Summary of the Invention
[0005] To establish consistency within the surgical field and develop objective standards for surgical procedures for novice surgeons to follow, extensive data collection can be performed to find commonalities among surgeons who have been trained at different universities, residencies, hospitals, and countries. Regardless of whether these commonalities ultimately stem from machine intelligence or individual detection and human intelligence, both approaches have a common problem: the need to digitize surgical procedures so that they can be replayed and quantitatively analyzed using offline analysis.
[0006] In addition, the digitization of the operating room should ideally eliminate the storage of raw sensor data, as each surgical procedure will generate terabytes of information that are difficult for humans and computers to store, replay, and view. This disclosure describes the digitization of the OR and surgical actions, which supports surgical procedure replay, offline analysis, and intraoperative guidance in a manner that enables large-scale deployment of surgical robotic systems to hospitals operating in different countries and environments.
[0007] This disclosure describes methods for reconstructing the key temporal and spatial aspects of a surgical action such that the surgical action can be digitally reconstructed for preoperative replay and planning, postoperative analysis, and intraoperative intervention. Depth cameras (e.g., RGBD sensors) can be integrated with a surgical robotic system (e.g., with a control tower) and integrated onto an unconstrained portable electronic device (e.g., a handheld computer tablet) to create a fused point cloud data stream. For example, one or more processors fuse the continuous RGBD point cloud stream with the recorded robotic data from the surgical robotic system to create a common coordinate system shared between an augmented reality tablet, a control tower, and an operating table.
[0008] The point cloud data for a multi-hour surgical procedure may be too large to be stored in the cloud for each surgical procedure, so semantic segmentation can be used to store only the portions of the point cloud data that are not associated with predefined or classified objects identified in the scene. This set of classified, identified objects and the unclassified point cloud data are stored in the cloud for offline analysis.
[0009] Additionally, online analysis of the identifiable objects in the scene identified by a trained semantic segmentation machine learning model can be used to alert operating room personnel of important events during the procedure. In summary, the OR digitization methods described herein provide operating room personnel with the ability to view replays of past surgical procedures, receive postoperative analysis of surgical procedures that have been analyzed by artificial or human intelligence, and receive intraoperative alerts regarding critical surgical events.
[0010] In some embodiments, a method for generating a digital asset that includes a record of a surgical action is described, the digital asset being fused from 3D point cloud data and recorded robotic system data. A surgical action is performed using a surgical robotic system. The performance of the surgical action is sensed by one or more depth cameras (also referred to as 3D scanners), which may be a real surgical procedure or a demonstration. At the same time, robotic system data associated with the surgical robotic system is recorded. The robotic system data can include the position, orientation, and movement of the robotic arm, as well as other notable indicia associated with the system or procedure.
[0011] Using computer vision, objects can be identified in the image data generated by a depth camera. However, some objects may not be identifiable. The identified objects can be memorized by storing their positions and orientations as sensed in their environment. Thus, their memory footprint can be reduced. For unidentifiable objects, image data such as 3D point cloud data or mesh data can be stored. Robot system data is also stored. The stored data can be saved as digital assets and synchronized over time (e.g., with timestamps) and / or managed in one or more sequences of frames. The stored data can be used to reconstruct surgical actions for playback, provide real-time alerts during surgery, and for offline analysis.
[0012] A playback device can reconstruct a surgical action by rendering the identified objects at the positions and orientations specified in the digital assets. The stored image data (e.g., 3D point cloud or mesh) can be rendered together with the identified objects, sharing a common 3D coordinate system, and synchronized over time. The stored robot system data can be used to reconstruct the precise positions, orientations, and movements of the surgical robots (e.g., robotic arms and surgical tools) rendered for playback or for real-time analysis.
[0013] Events serving as "jump points" for notable surgical events can be generated based on human activities and / or robot system data. For example, human activities and / or robot system data can indicate in the data when a UID is used for a surgical procedure or when a surgeon or instrument is changed, when a trocar is inserted, etc. A user can select an event during playback to jump to the relevant part of the surgical action. Additional aspects and features are described in this disclosure. Events can be generated by automatic analysis of one or more data streams in the data flow and / or by analysis performed by a person. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A surgical robot system according to one embodiment is shown.
[0015] Figure 2 and Figure 3 A system for capturing a surgical action according to some embodiments is shown.
[0016] Figure 4 An example of a digitized surgical action according to some embodiments is shown.
[0017] Figure 5 Accessible events of a digitized surgical action according to some embodiments are shown.
[0018] Figure 6 Digitized 3D data of a surgical action according to some embodiments is shown.
[0019] Figure 7 A method for digitizing surgical actions according to some embodiments is shown. DETAILED DESCRIPTION
[0020] Examples of various aspects and variations of the present invention are described herein and shown in the accompanying drawings. The following description is not intended to limit the present invention to these embodiments, but rather to enable those skilled in the art to make and use the present invention.
[0021] The following specification and drawings are illustrative of the present disclosure and should not be construed as limiting the present disclosure. Many specific details are described to provide a thorough understanding of the various embodiments of the present disclosure. However, in some instances, well-known or conventional details are not described so as to simplify the discussion of the embodiments of the present disclosure.
[0022] The phrase "an embodiment" or "embodiments" referred to in the specification means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present disclosure. The phrase "in an embodiment" appearing throughout the specification does not necessarily all refer to the same embodiment.
[0023] Reference Figure 1 , which is a pictorial view of an exemplary surgical robot system 1 in an operating site. The robot system 1 includes a user console 2, a control tower 3, and one or more surgical arms 4 at a surgical robot platform 5 (such as a table, a bed, etc.). The system 1 can be combined with any number of devices, tools, or accessories for performing a surgical operation on a patient 6. For example, the system 1 can include one or more surgical tools 7 for performing a surgical operation. The surgical tool 7 can be an end effector attached to the distal end of the surgical arm 4 for performing a surgical procedure.
[0024] Each surgical tool 7 can be manually manipulated, robotically manipulated, or both during a surgical operation. For example, the surgical tool 7 can be a tool for accessing, viewing, or manipulating the internal anatomy of the patient 6. In one embodiment, the surgical tool 7 is a gripper that can grasp the tissue of the patient. The surgical tool 7 can be manually controlled by a bedside operator 8; or it can be robotically controlled by actuation of the surgical arm 4 to which it is attached. The arm 4 is shown as a table-mounted system, but in other configurations, the arm 4 can be mounted on a cart, on the ceiling or sidewall, or in another suitable structural support.
[0025] Generally, a remote operator 9 (such as a surgeon or other operator) can use the user console 2 to remotely manipulate the arm 4 and / or the attached surgical tool 7, such as remote operation. The user console 2 can be located in the same operating room as the rest of the system 1, as Figure 1As shown. However, in other environments, the user console 2 may be located in an adjacent or nearby room, or it may be located at a remote location, e.g., in a different building, city, or country. The user console 2 may include a seat 10, foot controls 13, one or more handheld user input devices (UIDs) 14, and at least one user display 15 configured to display, e.g., a view of the surgical site within the patient 6. In an exemplary user console 2, the remote operator 9 sits in the seat 10 and views the user display 15 while manipulating the foot controls 13 and the handheld UID 14 to remotely control the arm 4 and the surgical tool 7 (mounted on the distal end of the arm 4).
[0026] In some variations, the bedside operator 8 may also operate the system 1 in a "bedside" mode, where the bedside operator 8 (the user) is now located on one side of the patient 6 and simultaneously manipulates the robot-driven tool (the end effector attached to the arm 4), e.g., holding the handheld UID 14 and a manual laparoscopic tool with one hand. For example, the left hand of the bedside operator may manipulate the handheld UID to control the robotic components, while the right hand of the bedside operator may manipulate the manual laparoscopic tool. Thus, in these variations, the bedside operator 8 may perform both robot-assisted minimally invasive surgery and manual laparoscopic surgery on the patient 6.
[0027] During an exemplary procedure (surgery), the patient 6 is prepared for surgery and draped in a sterile manner to achieve anesthesia. The initial access to the surgical site (to facilitate access to the surgical site) may be performed manually while the arms of the robotic system 1 are in a stowed configuration or a retracted configuration. Once the access is completed, the initial positioning or setup of the robotic system 1 including its arm 4 may be performed. The surgery then continues, where the remote operator 9 at the user console 2 uses the foot controls 13 and the UID 14 to manipulate the various end effectors and possibly an imaging system to perform the surgery. Manual assistance may also be provided by a bedside person (e.g., the bedside operator 8) wearing a sterile gown at the operating table or surgical table, who may perform tasks on one or more of the robotic arms 4, such as retracting tissue, performing manual repositioning, and tool changes. There may also be a non-sterile person to assist the remote operator 9 at the user console 2. When the procedure or surgery is completed, the system 1 and the user console 2 may be configured or set to a state to facilitate the completion of the post-operative procedure, such as cleaning or disinfection and entering or printing a health record via the user console 2.
[0028] In one embodiment, the remote operator 9 holds and moves the UID 14 to provide an input command to move the robotic arm actuator 17 in the robotic system 1. The UID 14 may be communicatively coupled to the remainder of the robotic system 1 via, for example, a console computer system 16. The UID 14 may generate a spatial state signal corresponding to the movement of the UID 14, such as the position and orientation of the handheld housing of the UID, and the spatial state signal may be an input signal for controlling the movement of the robotic arm actuator 17. The robotic system 1 may use a control signal derived from the spatial state signal to control the proportional movement of the actuator 17. In one embodiment, a console processor of the console computer system 16 receives the spatial state signal and generates a corresponding control signal. Based on these control signals for how the actuator 17 is powered to move a section or link of the arm 4, the movement of the corresponding surgical tool attached to the arm may mimic the movement of the UID 14. Similarly, the interaction between the remote operator 9 and the UID 14 may generate, for example, a gripping control signal that closes the jaws of the gripper of the surgical tool 7 and grips the tissue of the patient 6.
[0029] The surgical robotic system 1 may include a number of UIDs 14, and a corresponding control signal is generated for each UID to control the actuators and surgical tools (end effectors) of the respective arms 4. For example, the remote operator 9 may move a first UID 14 to control the movement of the actuator 17 located in the left robotic arm, where the actuator responds by moving linkages, gears, etc. in the arm 4. Similarly, the movement of a second UID 14 by the remote operator 9 controls the movement of another actuator 17, which in turn moves other linkages, gears, etc. of the robotic system 1. The robotic system 1 may include a right arm 4 fixed to a bed or table on the right side of the patient, and a left arm 4 located on the left side of the patient. The actuator 17 may include one or more motors that are controlled such that they drive the joints of the arm 4 to rotate to, for example, change the orientation of an endoscope or gripper of the surgical tool 7 attached to the arm relative to the patient. The movement of several actuators 17 in the same arm 4 may be controlled by a spatial state signal generated from a particular UID 14. The UID 14 may also control the movement of the corresponding surgical tool gripper. For example, each UID 14 may generate a corresponding gripping signal to control the movement of an actuator (e.g., a linear actuator) that opens or closes the jaws of the gripper at the distal end of the surgical tool 7 to grip tissue within the patient 6.
[0030] In some embodiments, communication between the platform 5 and the user console 2 may be through the control tower 3, which may convert user commands received from the user console 2 (and more specifically from the console computer system 16) into robotic control commands transmitted to the arm 4 on the robotic platform 5. The control tower 3 may also transmit status and feedback from the platform 5 back to the user console 2. The communication connections between the robotic platform 5, the user console 2, and the control tower 3 may be via wired and / or wireless links, using any suitable data communication protocol among various data communication protocols. Any wired connection may optionally be built into the floor and / or walls or ceiling of the operating room. The robotic system 1 may provide video output to one or more displays, including displays within the operating room and remote displays accessible via the Internet or other networks. The video output or feed may also be encrypted to ensure privacy, and all or part of the video output may be saved to a server or an electronic health record system.
[0031] The surgical robotic arm may have movable, articulating, and / or motorized members having multiple degrees of freedom that may hold various tools or attachments at the distal end. Example systems include the da Vinci(r) surgical system, which may be used for minimally invasive surgery (e.g., urological procedures, general laparoscopic surgical procedures, gynecological laparoscopic surgical procedures, general non-cardiovascular thoracoscopic surgical procedures, and thoracoscopic-assisted cardiothoracotomy procedures). A "virtual surgical robotic arm" may be a computer-generated model of a robotic arm rendered on captured video set by the user. The virtual surgical robotic arm may be a complex 3D model of a real robotic arm. Alternatively or additionally, the virtual surgical robotic arm may include visual aids such as arrows, tool tips, or other representations involving providing pose information about the robotic arm, such as a geometrically simplified version of the real robotic arm.
[0032] Figure 2 A system for digitizing surgical actions is shown, which includes one or more depth cameras (e.g., 3D scanning sensors 21, 22, and 23). These depth cameras may be arranged on Figure 1 any part of the surgical robotic system described (e.g., the control tower 3, the operating table 5, and / or the user console 2) or integrated with any such part. Additionally, the depth cameras may be integrated with a portable electronic device 20 and / or mounted on a bracket, wall, or ceiling of the operating room. The depth cameras sense surgical actions performed using the surgical robotic system. The surgical actions may be real procedures performed on a real patient using the surgical robotic system as shown, demonstrations on a real patient or a mannequin, or other training actions. Figure 1 shown.
[0033] In some embodiments, the system includes three to five depth cameras. The depth cameras can generate synchronized RGBD (color and depth) frames. These RGBD frames can be fused to produce a singular sequence of wide-angle RGBD frames for OR. The frame rate of the sensors can be dynamically changed by the system, depending on which camera is considered most important for retrieving data at any given point in time. The camera considered most important can have an increased frame rate.
[0034] The portable electronic device 20 can be a computer tablet operated by a surgical personnel during the procedure, and the portable electronic device houses a 3D scanning sensor that retrieves RGBD (color + depth) frames. These RGBD frames captured by the portable electronic device can be localized in 3D space by a continuously running SLAM (Simultaneous Localization and Mapping) algorithm that specifies the position of the tablet relative to the surgical robotic system in 3D space. The tablet sends this image data (RGBD frames) to the processor 24, which can be done at various transfer rates, depending on the speed of the Wi-Fi connection or other communication protocols in the OR.
[0035] During the procedure, robotic system data associated with the surgical robotic system is recorded. The recording can be concurrent and synchronized with the image data captured by the depth cameras. The robotic system data stream can include user controls from the user console, such as the input status, position, or rotation of a user interface device (such as the UID 14 in Figure 1 ), or the status of a footswitch (such as footswitch 13), which are used to operate the surgical robotic system. The robotic system data can also include real-time information describing the status, position, or rotation of a surgical instrument attached to the robotic arm (such as whether a gripper or cutter is open or closed or how far it extends into the surgical workspace) or the force or energy applied using the surgical instrument. Additionally, the movement, position, and orientation of the robotic arm or the surgical table can be recorded, such as the joint values (describing joint positions) or modes (draped, not draped) of each robotic arm in the robotic arm. Other information can also be included, such as the network connectivity of the surgical robotic system, user login data, graphical user interface status (e.g., what screen is active on the user display 15 shown in Figure 1 ), and the status of touch points on the surgical robotic system. The robotic system data can be continuously recorded, timestamped, and stored long-term in synchronization with the timestamped point cloud data.
[0036] The processor 24 can coordinate the robotic system data with the image data 40 based on the alignment of timestamps. The image data can be a collection of RGBD frames such that the combined data (digital asset 30) can be reconstructed for the surgical robotic system (including Figure 13D meshes of the surgical robot, control tower, user console), and any pre-classified objects in the operating room, such as Mayo stands, walls, floors, stools, shelves, lighting booms, and screens / monitors.
[0037] In some embodiments, the endoscope 31 may generate an endoscope feed that is also received by the processor and fused with the digitized surgical actions. The endoscope feed may also be associated with and synchronized with the robotic system data and the image data for playback. The processor may compress the image data, the robotic system data, and optionally the endoscope feed into digital assets 30 that can be used for playback. The digital assets may be stored in an electronic device (such as the network-connected server 26). The playback device 28 may retrieve the digital assets for playback. As described in other sections, the processor may format the digital assets such that the positions and orientations of the identified objects are stored instead of the original image data or point cloud data.
[0038] For the identified objects, a common coordinate system may be used, the same as for the unidentified objects, walls, ceilings, and floors of the surgical robotic system, to determine the position and orientation. The playback device may access the models 32 of the identified objects to render them into the positions and orientations corresponding to the unidentified objects to render the OR and the surgical actions (e.g., to the display) during playback.
[0039] In addition to synchronizing and compressing the data of the surgical actions for playback, the system may also generate user feedback during the surgical actions based on the robotic system data or the positions and orientations of the identified objects. The user feedback may indicate at least one of the following: the user interface device for the surgical procedure, the replacement of the surgeon, the replacement of the instrument, the docking of the trocar, the insertion of the trocar, the phase change, the slow execution of the surgical procedure, the improper position of the trocar, the improper position of the patient, the arm collision, and the improper position of the OR personnel.
[0040] This feedback may be automatically generated using machine learning algorithms and / or computer-implemented logic based on data analysis. For example, based on the positions and movements of the surgical instruments and robotic arms, the endoscope feed, or other data, the processor may determine that a trocar is being inserted, the insertion angle, and whether the insertion is appropriate. In another example, based on computer vision, the processor may determine whether the patient is improperly positioned on the operating table.
[0041] In another example, based on an evaluation of telemetry of robotic system data and comparison with known benchmarks or previous surgical actions, the processor can determine that a surgical action is being performed too slowly or too quickly. If it is determined that a part of the surgical procedure involving cutting or removal of an organ is being performed faster or slower than an average or predetermined time by more than a threshold amount of time, the system can provide feedback to the surgeon and personnel so that the surgical action can be adjusted if appropriate.
[0042] In yet another example, the processor can analyze the position of surgical personnel in the OR and determine if they are properly positioned. For example, a person may be standing at the feet of the patient. The processor is able to identify the person and her position in the OR relative to the operating table and, based on reference to known benchmarks or predetermined parameters, determine that the person is not standing in the appropriate position. Feedback can be provided in real time to alert the person to move along the operating table towards the middle part of the patient. All feedback can be given in real time, for example, with a minimum delay (including inevitable delays such as those caused by data processing, communication delays, and buffering), to provide real-time feedback during the surgical procedure.
[0043] Figure 3 Processing of surgical actions for forming a digital asset 30 according to some embodiments is shown. Raw image data, which can be 3D point cloud data, is received for processing. Point cloud data is a set of data points in space generated by a 3D image capture device. The point cloud data represents measurements of many points on the outer surface of an object sensed by the 3D image capture device.
[0044] When collecting the raw image data 40 (e.g., from Figure 2 cameras 21 to 23), the data can be processed at a pre-processor block 42 before analysis. Here, the point cloud data across frames can be denoised, downsampled, and merged so that the density of the 3D point cloud is consistent in 3D space. A probability-based processing filter can be applied to determine which points should be maintained over time and which should be discarded since objects in the scene are not guaranteed to be static.
[0045] As the common 3D point cloud is continuously updated, another process (which may be slower and operate in parallel with the pre-processor) is performed at block 44, block 48, and block 50 to analyze the updated points. The stable portion of the points in the point cloud can be separated from the unstable portion to accelerate point cloud analysis. At the mesh converter block 48 (optional), a background process can generate a 3D mesh of the point cloud, which has a reduced set of points representing the vertices of 3D polygon shapes. In regions (e.g., pixels) of the point cloud that appear to include known objects such as an operating table and robotic arm, user console, control tower, or other known and classified objects, the simplified 3D mesh of the point cloud is not computed. Instead, a component transformation that describes the position and orientation of the identified object in space replaces the identified object. The position and orientation are aligned with the original point cloud. It should be understood that the orientation describes the direction or rotation of the object in the common 3D coordinate system.
[0046] In some embodiments, portions of the point cloud with significant updates will be processed to update the 3D meshes of those portions, while those portions with no updates or below a threshold change amount will not be updated. The 3D meshes of portions of the point cloud with unidentified objects can be maintained (updated at a sufficient rate) for consumer-facing features, such as, for example, in the case where a robotic arm travels around a patient's body and in the case of determining trocar placement on a patient taking into account the patient's size and procedure.
[0047] To identify objects, at block 44, a computer vision analyzer 44 can detect and identify objects in the original image data, such as but not limited to Figure 1 the shown control tower, operating table, robotic arm, and user console. The computer vision analyzer can apply one or more machine learning algorithms (e.g., neural networks), which are trained to detect and identify (with a certain probability) objects belonging to one or more classifications. The machine learning algorithms can perform classification (e.g., identifying whether an object exists in the image), object identification (e.g., identifying an object in a specific region), and / or semantic segmentation (e.g., identifying objects at the pixel level). The computer vision algorithms can include, for example, trained deep neural networks, trained convolutional neural networks, edge detection, corner detection, feature detection, patch detection, or other equivalent techniques that can be deployed to detect and identify objects. An API, such as, for example, Tensorflow, ImageAI, or OpenCV, can be utilized.
[0048] Using similar image processing and computer vision algorithms, the processor can determine the position (e.g., position in a common 3D coordinate system) and orientation of the identified object at block 50. Robot system data can be used to determine the detailed positioning of the robotic arm, surgical instrument, and surgical table. For example, in addition to other data, the robot system data can include the angle and height of the surgical table and the joint angles of each robotic arm. An object identifier (such as a name or unique identifier) can be associated with each of the identified objects for reference. For example, the playback device can render a computer model of the surgical table based on the identifier type and render the model in the common coordinate system based on the stored position and orientation associated with the object identifier.
[0049] In some embodiments, humans such as surgical assistants and other personnel are identified and converted to a more compressed data representation, such as a spline curve representation. This can further reduce the footprint of the digital assets. It should be understood that digital assets can be digital files or groups of files that are stored long-term for playback, rather than being stored short-term or temporarily for processing purposes.
[0050] To reduce the data storage size and improve subsequent retrievability and readability, the raw image data (e.g., 3D point cloud data) for objects that are not identified or classified in the surgical environment is stored as raw point cloud data, but this is not the case for the identified objects as discussed in other sections (which can be stored as position and orientation or as spline data). In some cases, the 3D point cloud data for the unidentified objects can be converted to mesh data (e.g., polygon mesh, triangle mesh) or spline data at block 48. The processor can execute one or more conversion algorithms (e.g., Delaunay triangulation, alpha shape, rolling ball method, and other equivalent techniques) to convert the 3D point cloud data to mesh data.
[0051] In some embodiments, unidentified objects can be stored and further analyzed offline at block 46. In this context, offline means that the process of providing real-time feedback and generating digital assets for playback does not wait for or depend on this offline analysis. At this block, continuous offline analysis of the unidentified objects can be performed during subsequent surgical procedures until the unidentified objects are classified. The point cloud data including the unidentified objects can be stored in a computer-readable memory (e.g., in a networked server) and analyzed by an automated computing platform (using computer vision algorithms) and / or by scientists and others to identify and classify the unidentified objects.
[0052] The point cloud data can include timestamp data that allows the data to be temporally fused with other sensor data retrieved from the surgical robotic system, such as microphones or cameras on the surgeon's bridge and the robot logs. The fused data can further assist in classifying and identifying objects. The new classifications can be used to train machine learning algorithms for computer vision, which can then be used by the processor to "identify" these objects in the digitization of future surgical procedures.
[0053] For example, referring to Figure 1 , some Mayo stands (e.g., stand 19), first assistants 8, chairs or shelf units 18, or lighting booms may not have been identified in one or more earlier procedures. These objects can be classified offline and then used to identify these or similar objects in later procedures, thus improving online analysis and reducing the footprint of digital assets with repetition. As described in other sections, these objects can be represented by their position and orientation in the OR for storage purposes and replaced with computer-generated models or representations during playback, just like other identified objects. Computer-based object identification for unclassified objects can occur in real-time in the OR or offline in the cloud using offline analysis after the data has been uploaded to the server.
[0054] Additionally or alternatively, the critical surgical events described in other sections can also be identified offline by an automated computing platform and / or by scientists or other personnel. The computing platform can train machine learning algorithms to identify further surgical events useful for postoperative analysis based on the position, orientation, and movement of surgical personnel, robotic arms, UIDs, and surgical instruments / tools.
[0055] Referring back to Figure 3 , the combiner 52 can combine various data into one or more digital assets for future playback. In some embodiments, the point cloud data of unidentifiable objects, the position and orientation associated with one or more identified objects, and the robotic system data are stored as digital assets. In some embodiments, the digital assets also include endoscopic images captured during the surgical procedure. The data is synchronized temporally (e.g., with timestamps) such that the data resembles the original surgical procedure when played back.
[0056] In some embodiments, high-resolution data (spatial resolution and / or temporal resolution), such as raw 3D point cloud data or high-resolution mesh data generated from 3D point cloud data, will be stored in the cloud for service teams, debugging, and data analysis. This high-resolution digital asset can be analyzed offline. However, to replay this data on consumer-facing devices, the combiner can downsample the robotic system data. Additionally, the point cloud data or mesh data can be streamed "on demand" to reduce the computational complexity for low-power consumer devices and reduce the network traffic required to view different parts of the surgery. Examples of low-power consumer devices include laptops, smartphones running mobile apps, or Wi-Fi-connected mobile, unconstrained virtual reality (VR) devices.
[0057] A buffer or staging server 26 can host the digital asset in a consumer-facing application, which can be provided as a surgical data stream, and the consumer-facing application allows the consumer to replay a surgery that has had its data fused by timestamp. The digital asset can be streamed to the playback device, for example, via a communication network (e.g., TCP / IP, Wi-Fi, etc.). Examples of digital assets are shown in Figure 4 below.
[0058] Figure 4 FIG. shows a digital asset 30 according to some embodiments. The digital asset can include image data (e.g., 3D point cloud or mesh data) containing unrecognized objects, compressed data associated with the recognized objects, recorded robotic surgical data, and an endoscopic feed. The data can be fused by timestamp in a single data set. The endoscopic feed can be a stereoscopic video feed captured by an endoscope used during a surgical procedure. The endoscopic feed can have image data of the interior of the patient.
[0059] The digital asset can include robotic system data recorded during a surgical procedure. The recorded robotic system data can include at least one of the following: UID status, including position, rotation, and raw electromagnetic (EM) sensor data; footswitch status; instrument status, including position rotation and applied energy / force; robotic arm status, including joint values and mode; operating table status, including position and rotation; system errors or information, e.g., Wi-Fi connection, which surgeon logged in, GUI status / information, and which touch points on the robot were pressed / pushed.
[0060] The digital asset includes recognized objects extracted from 3D point cloud data. The recognized objects can include Figure 1Components of the surgical robot system, such as, for example, the user console, the operating table and the robotic arm, the control tower, and other parts, position and orient the mobile electronic device 20 described above. Additionally, with the advancement of the system's intelligent platform (e.g., through offline analysis), additional objects can be identified and stored in the digital asset. For example, instead of Figure 1 the Mayo stand, the display, and other devices that are part of the surgical robot system shown above can increasingly be automatically identified by the system and included as the identified objects. The identified objects can include humans stored as spline curve data and / or for the positions and orientations and joint values of leg joints, shoulders, necks, arm joints, and other human joints.
[0061] In addition, the digital asset can also include 3D point cloud data or mesh data for unidentified and unknown objects, such as, for example, used bandages on the floor, unknown surgical instruments and devices, etc. Thus, the identified objects can be represented in a compressed form, while the unidentified objects are represented with compressed data.
[0062] The playback device can render a model corresponding to the identified object based on the position and orientation of the identified object. For example, the model of the operating table and the robotic arm can be rendered on the point cloud or mesh data that defines the walls, floor, and ceiling of the OR. The model is rendered using the position and orientation specified by the digital asset in a common coordinate system. The detailed positions of the operating table and the robotic arm can be memorized in the recorded robotic system data and rendered based on that data.
[0063] As described in other parts, the digital asset can include events 70 associated with the timeline or frames of the digital asset. For example, as Figure 5 shown, the events can provide playback access points that the user can jump to during playback. Noteworthy events can be automatically determined based on the analysis of the digitized surgical actions. For example, Figure 2 or Figure 3 a processor or an offline computing platform can analyze the surgical robot system data (such as the joint values of the robotic arm or the movement and status of the surgical instrument, the position and / or movement of the assistant, UID input, endoscopic feed images, etc.) to determine critical events. Machine learning, decision trees, programming logic, or any combination thereof can be used to perform the analysis to determine critical events.
[0064] For example, robotic system data (e.g., telemetry) can be analyzed to directly indicate events such as when a UID is used for a surgery, when a surgeon is changed, or when an instrument is changed. An offline platform can apply one or more machine learning algorithms (e.g., offline) to endoscopic feeds, robotic system data, and image data to generate events indicating at least one of the following: when a trocar is inserted, phase changes, abnormally slow portions of a surgical procedure, adverse patient positions, robotic arm collisions, incorrect OR personnel positions, improper trocar positions, or insertion port positions. Each event can be associated with a time point or frame of a digital asset. Thus, a user can directly jump to an event of interest or search a set of actions to extract actions that show a particular event.
[0065] In addition to jump points, digital assets can be accelerated or decelerated. For example, a user can select to play an entire surgical procedure at a rate that exceeds the real-time playback speed of the surgical procedure (e.g., 1.5X, 2.0X, or 3.0X) by requesting incremental data about the surgical procedure from a grading server.
[0066] Referring Figure 6 , digital asset 30 can include a sequence of frames 72 generated based on 3D image data captured by a 3D scanner. Frame 74 can include point cloud data or mesh data of an unrecognized object and / or room geometry such as a floor, walls, and ceiling. Key frames 76 can be temporally scattered in the sequence of frames 72 that include the position and orientation of one or more recognized objects. Additionally, it is not necessary for each key frame to contain every recognized object. For example, a key frame can include only those recognized objects that have values (e.g., position and orientation or spline data) that have changed since the previous key frame. This can further reduce the digital asset footprint and improve computational efficiency.
[0067] Referring Figure 7 , a method or process 80 for digitizing surgical actions according to some embodiments is shown. For example, as Figure 2 and Figure 3 shown, the method can be executed by one or more processors.
[0068] At operation 81, the process includes sensing surgical actions performed using a surgical robotic system via one or more depth cameras. As described, one or more depth cameras (also referred to as 3D scanners) can be integrated with components of the surgical robotic system, mounted on a wall or a bracket in the OR, and / or integrated with a portable electronic device (e.g., a tablet).
[0069] At operation 82, the process includes recording robotic system data associated with a surgical robot of a surgical robotic system. The robotic system data can define the position of the robotic arm (e.g., joint values), the height and angle of the surgical table, and various other telemetry data of the surgical robotic system, as described in other sections.
[0070] At operation 83, the process includes identifying one or more identified objects in image data generated by one or more depth cameras. This can be performed using computer vision algorithms. As discussed in other sections, the identified objects can include components of the surgical robotic system as well as other general components. The library of identifiable objects can grow with repetition as these classifications are extended to cover new objects.
[0071] At operation 64, the process includes storing in an electronic memory a) the image data of the un-identified objects, b) the positions and orientations associated with one or more identified objects, and c) the robotic system data. The image data can be stored as 3D point cloud data, 3D mesh data, or other equivalent representations.
[0072] In some embodiments, operation 64 alternatively includes storing in an electronic memory a) the spline curves, mesh representations, or the positions, orientations, and joint values of one or more humans identified in the image data, b) the 3D point cloud data (or mesh data) of the un-identified objects, and c) the robotic system data, as digital assets for playback of surgical actions. One or more events can be determined based on the position or movement of one or more humans or the robotic system data or the robotic system data, and the one or more events are used as jump points directly accessible during playback of the digital assets, as described in other sections.
[0073] In summary, the embodiments of the present disclosure allow for digital reconstruction of the operating room, which supports preoperative planning and replay of surgical procedures. It improves the awareness of the position of physical objects in the scene and what a person is doing at any given point in time. It also supports postoperative analysis. It enhances the offline analysis of critical events during surgery. It can provide explanations for the causes of different problems during surgery. It generates data such that operating room personnel can identify problems and their causes. It also addresses intraoperative guidance. The embodiments can capture data outside of the endoscope and robotic logs. Endoscopic analysis and robotic logs can be integrated with operating room analysis. It can alert OR personnel of critical events (e.g., in real time) to avoid surgical errors, increase surgical speed, improve patient outcomes, and improve the ergonomics / human factors of the surgical robotic system for its users.
[0074] In some embodiments, the various blocks and operations may be described by one or more processors executing instructions stored on a computer-readable medium. Each processor may include a single processor or multiple processors, each of which includes a single processor core or multiple processor cores. Each processor may represent one or more general-purpose processors, such as a microprocessor, a central processing unit (CPU), etc. More specifically, each processor may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. Each processor may also be one or more special-purpose processors, such as an application-specific integrated circuit (ASIC), a cellular or baseband processor, a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, a graphics processor, a communication processor, a cryptographic processor, a coprocessor, an embedded processor, or any other type of logic capable of processing instructions.
[0075] Modules, components, and other features such as the algorithms or method steps described herein may be implemented by a microprocessor, discrete hardware components, or integrated into the functionality of hardware components such as an ASIC, FPGA, DSP, or similar device. Additionally, such features and components may be implemented as firmware or functional circuitry within a hardware device; however, such details are not closely related to the embodiments of the present disclosure. It should also be understood that network computers, handheld computers, mobile computing devices, servers, and / or other data processing systems with fewer or potentially more components may also be used in conjunction with the embodiments of the present disclosure.
[0076] Some portions of the detailed description above have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. Those skilled in the data processing arts use these algorithmic descriptions and representations to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, considered to be a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulation of physical quantities.
[0077] However, it should be borne in mind that these and similar terms are to be associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless explicitly stated otherwise from the foregoing discussion, it is to be understood that throughout the specification, discussions using terms such as those set forth in the following claims refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities within the registers and memories of the computer system and transforms it into other data similarly represented as physical quantities within the computer system memory or registers or other such information storage, transmission, or display devices.
[0078] Embodiments of the present disclosure also relate to apparatuses for performing the operations herein. Such computer programs are stored in a non-transitory computer-readable medium. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium (e.g., read-only memory (“ROM”), random access memory (“RAM”), magnetic disk storage media, optical storage media, flash memory devices).
[0079] The processes or methods depicted in the foregoing figures may be executed by processing logic that includes hardware (e.g., circuitry, dedicated logic, etc.), software (e.g., embodied on a non-transitory computer-readable medium), or a combination of both. Although the processes or methods are described above in accordance with some sequential operations, it should be understood that some of the operations described may be performed in a different order. In addition, some operations may be performed in parallel rather than sequentially.
[0080] Embodiments of the present disclosure have been described without reference to any particular programming language. It should be understood that various programming languages may be used to implement the teachings of the embodiments of the present disclosure described herein.
[0081] In the foregoing specification, embodiments of the present disclosure have been described with reference to specific exemplary embodiments of the present disclosure. It will be apparent that various modifications may be made to the present disclosure without departing from the broader spirit and scope as set forth in the following claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. For purposes of explanation, the foregoing description uses specific nomenclature to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that practicing the invention does not require specific details. For purposes of illustration and description, the foregoing description of specific embodiments of the invention has been provided. They are not intended to be exhaustive or to limit the invention to the specific forms disclosed; various modifications and changes are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, to thereby enable others skilled in the art to best utilize the invention and various embodiments with various modifications as are suited to the particular use contemplated. The following claims and their equivalents are intended to define the scope of the invention.
Claims
1. A method executed by one or more processors, comprising: sensing a surgical action performed using a surgical robotic system via one or more depth cameras; recording robotic system data associated with the surgical robotic system; identifying one or more identified objects in point cloud data generated by the one or more depth cameras; and storing in an electronic memory: a) point cloud data of unidentified objects, b) positions and orientations associated with the one or more identified objects, and c) the robotic system data, the robotic system data being synchronized in time and stored as a digital asset for playback.
2. The method according to claim 1, further comprising extracting the point cloud data of the unidentified objects from the point cloud data generated by the one or more depth cameras for storage in the digital asset.
3. The method according to claim 1, wherein The digital asset further includes endoscopic images captured during the surgical action.
4. The method according to claim 1, wherein The digital asset includes: a) consecutive frames, each consecutive frame including the point cloud data of the unidentified objects; and b) key frames, the key frames being temporally dispersed among the consecutive frames, the key frames including the positions and orientations of the one or more identified objects.
5. The method according to claim 1, further comprising associating events with the digital asset, the events indicating at least one of the following: insertion of a trocar, phase change, slow execution of a surgical procedure, improper trocar position, improper patient position, arm collision, improper OR personnel position, wherein the events are directly accessible during playback of the digital asset.
6. The method according to claim 1, further comprising associating events with one or more frames of the digital asset, the events indicating at least one of the following: user interface device for a surgical procedure, change of surgeon, change of instrument, docking of a trocar, wherein the events are directly accessible during playback.
7. The method according to claim 1, wherein The point cloud data of the unidentified objects is stored as mesh data.
8. The method according to claim 1, wherein The robotic system data includes at least one of the following: position or rotation of a user interface device controlling the operation of the surgical robotic system, electromagnetic sensor data, status of a foot switch controlling the operation of the surgical robotic system, position or rotation of a surgical instrument attached to a robotic arm, force or energy applied using the surgical instrument, joint values or modes of the robotic arm, network connectivity of the surgical robotic system, user login data, graphical user interface status, and status of touch points on the surgical robotic system.
9. The method according to claim 1, wherein The one or more identified objects include at least one of the following: a robotic arm, one or more persons, a user console, a control station, a surgical table.
10. The method according to claim 9, wherein The one or more persons are represented by joint values or spline curves.
11. The method according to claim 1, wherein The one or more depth cameras are RGB depth cameras.
12. The method according to claim 11, wherein The one or more depth cameras include a depth camera disposed on a control tower of the surgical robot system and a second depth camera disposed on a portable electronic device, and the one or more depth cameras are used in an operating room for the surgical action.
13. The method according to claim 1, wherein, The point cloud data of the unrecognized object is stored in an electronic memory and used to train a machine learning algorithm for future recognition of the unrecognized object.
14. The method according to claim 1, further comprising generating user feedback during the surgical action based on the robot system data or the position and orientation of the recognized object, the user feedback including at least one of the following: a user interface device for a surgical operation, replacement of a surgeon, replacement of an instrument, docking of a trocar, insertion of a trocar, phase change, slow execution of a surgical operation, improper trocar position, improper patient position, arm collision, and improper OR personnel position.
15. The method according to claim 1, wherein The playback of the digital asset includes: rendering a model corresponding to the recognized object based on the position and orientation of the recognized object.
16. A surgical robot system, comprising: a plurality of depth cameras disposed on a control tower and a mobile electronic device to capture a surgical action performed by a surgical robot, and one or more processors configured to perform the following operations: recording robot system data associated with the surgical robot; identifying one or more recognized objects in point cloud data generated by the one or more depth cameras; and storing in an electronic memory: a) point cloud data of an unrecognized object, b) the position and orientation associated with the one or more recognized objects, and c) the robot system data as a digital asset for playback.
17. The surgical robot system according to claim 16, wherein, The point cloud data of the unrecognized object, the position and orientation associated with the one or more recognized objects, and the robot system data are synchronized in time.
18. The surgical robot system according to claim 16, wherein, The point cloud data of the unrecognized object is stored as mesh data.
19. A surgical robot system, comprising: a plurality of depth cameras disposed in an operating room to capture a surgical action performed by a surgical robot, and one or more processors configured to perform the following operations: recording robot system data associated with the surgical robot; identifying one or more humans based on point cloud data generated by the one or more depth cameras; and storing in an electronic memory: a) a spline curve representation of the one or more humans, b) the point cloud data of the unrecognized object, and c) the robot system data as a digital asset for playback of the surgical action.
20. The surgical robot system according to claim 19, wherein, The digital asset includes one or more events that can be directly accessed during the playback of the digital asset, the one or more events being determined based on the position or movement of the one or more humans or the robotic system data, the one or more events indicating at least one of the following: insertion of a trocar, change of phase, slow execution of a surgical procedure, improper trocar position, improper patient position, robotic arm collision, improper position of surgical personnel.
Citation Information
Patent Citations
Robot localization in a workspace via detection of a datum
CN111149067A
Surgical field camera system
US20190038362A1
Automatic endoscope video augmentation
US20200110936A1