Systems and methods for determining in vivo depth perception in surgical robotic systems

The surgical robotic system generates and combines depth maps with confidence values to enhance precision and safety in surgical robotics by providing reliable depth perception, addressing the lack of sufficient depth information in conventional systems.

JP7794450B2Active Publication Date: 2026-01-06VICARIOUS SURGICAL INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022547862
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-06
Filing Date
2021-02-08
Publication Date
2026-01-06
Estimated Expiration
2041-02-08

AI Technical Summary

Technical Problem

Conventional augmented and virtual reality systems in surgical robotics lack sufficient depth information, hindering accurate visual representation and feedback, which is crucial for safe and efficient robotic surgery.

Method used

A surgical robotic system that generates multiple depth maps using cameras with autofocus and disparity data, combines these maps into a single composite depth map, and assigns confidence values to distance data, enabling precise control of robotic components.

Benefits of technology

Enhances surgical precision by providing reliable depth perception, allowing safe and efficient robotic surgery by accurately controlling robotic movements and reducing collisions outside the operator's field of view, thereby improving surgical outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794450000001
    Figure 0007794450000001
  • Figure 0007794450000002
    Figure 0007794450000002
  • Figure 0007794450000003
    Figure 0007794450000003
Patent Text Reader

Abstract

A system and method for generating a depth map from image data in a surgical robotic system uses a robotic subsystem having a camera assembly with first and second cameras for generating the image data. The system and method generates multiple depth maps based on the image data and then converts the multiple depth maps into a single composite depth map having associated distance data. The system and method can then control the camera assembly based on the distance data in the single composite depth map.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] <Related Applications> This application claims priority to U.S. Provisional Patent Application No. 62 / 971,097, filed February 6, 2020, entitled "DEPTH PERCEPTION IN VIVO," the contents of which are incorporated herein by reference.

[0002] The present invention relates to surgical robotic systems, and more particularly to surgical robotic systems that use a camera assembly. [Background technology]

[0003] Minimally invasive surgery (MIS) has proven beneficial to patient outcomes when compared to open surgery, where surgeons operate by hand through large incisions, by significantly reducing patient recovery time, risk of infection, and future hernia incidence, and allowing more outpatient procedures to be performed.

[0004] Despite the advent of manual and robotic MIS systems, open surgery remains the standard procedure for many indications due to the complexity of the procedure and general limitations of current MIS solutions, including the amount of training and practice required to become proficient in MIS and the limited abdominal access from a single entry point.

[0005] U.S. Patent No. 10,285,765, entitled "Virtual Reality Surgical Device," U.S. Patent Publication No. 2019 / 0142531, entitled "Virtual Reality Wrist Assembly," and U.S. Patent Publication No. 2019 / 0076199, entitled "Virtual Reality Surgical Camera System," disclose advanced surgical systems that allow access to any location in the abdomen with a single MIS incision. Natural interfaces and poseable viewpoints create surgical systems that require minimal instrument-specific training. Surgeons interact with the robot as if it were their own hands and eyes, which, combined with high-quality manufacturing and a lower cost per system, allows surgeons to focus on providing high-quality care to their patients.

[0006] One of the biggest problems with traditional MIS systems is the possibility of injuries occurring outside the operator's field of view. Generating a detailed three-dimensional (3D) map of the surgical stage allows the system to proactively prevent collisions and help surgeons plan optimal surgical paths for robotic components, such as robotic arms.

[0007] Furthermore, depth data will enable the use of autonomous surgical robots that can interact closely with human tissue while navigating and performing surgical procedures within a patient in a safe and efficient manner. Intelligent, autonomous surgical systems capable of performing basic surgical procedures require detailed three-dimensional knowledge of the patient's interior. Therefore, depth knowledge is fundamental to this goal.

[0008] The robotic surgery system disclosed in the aforementioned publication blends miniaturized robotics with augmented reality. In some embodiments, the system includes two eight-degree-of-freedom (DOF) robotic arms, plus a stereoscopic camera assembly or head, which is inserted through a single incision and deployed within the patient's abdominal cavity. The surgeon controls the robotic arms using two six-axis handheld controllers while visualizing the complete surgical situation via the robotic head and a virtual reality headset.

[0009] The system's unique architecture offers opportunities and capabilities not realized by any other surgical method or system, yet allows the surgeon to rely solely on the patient's internal visual feedback. The brain naturally uses a variety of low- and high-level sources to obtain reliable and robust distance estimates. This is because humans rely heavily on depth cues to interact with their environment. Making these cues accessible through the robotic head enhances the surgeon's capabilities, creating a much richer and more effective experience.

[0010] However, available conventional and commercially available augmented reality and virtual reality systems are unable to provide sufficient depth information to the surgeon to ensure accurate visual representation and feedback. Summary of the Invention

[0011] The present invention is directed to a surgical robotic system that uses a depth perception subsystem to generate multiple depth maps and then combine or merge the depth maps into a single depth map. The depth perception subsystem also generates a set of confidence values ​​associated with distance data in the single composite depth map. The confidence values ​​are indicative of the confidence or likelihood that distance data associated with selected points or portions of the depth maps is correct or accurate. The depth perception subsystem generates depth maps associated with cameras of a camera assembly. Specifically, the depth perception subsystem generates depth maps associated with the camera's autofocus mechanism, disparity data associated with each camera, and(parallax) data, and between the image data from each camera (disparity) The depth perception subsystem then processes all of the depth maps to generate a single combined depth map. The depth map data and confidence values ​​can be used by the system to move one or more components of the robotic subsystem.

[0012] The present invention is directed to a surgical robotic system comprising: a robot subsystem having a camera assembly with first and second cameras for generating image data; a computing unit comprising: a processor for processing the image data; a control unit for controlling the robot subsystem; and a depth perception subsystem that receives the image data generated by the first and second cameras, generates multiple depth maps based on the image data, and then converts the multiple depth maps into a single composite depth map with associated distance data. The robot subsystem further comprises multiple robot arms and a motor unit for controlling movement of the multiple robot arms and the camera assembly. The control unit uses the distance data associated with the single combined depth map to control one of the camera assembly and the robot arm. The depth perception subsystem further comprises a depth map conversion unit that receives the multiple depth maps and then converts the depth maps into a single composite depth map. The depth map conversion unit generates the single composite depth map using a region convolutional neural network (R-CNN) technique.

[0013] Additionally, each of the first and second cameras includes an image sensor for receiving optical data and generating image data in response thereto, a lens and optical system having one or more lens elements optically coupled with the image sensor for focusing the optical data onto the image sensor, and an autofocus mechanism associated with the lens and optical system for automatically adjusting the one or more lens elements and generating autofocus data.

[0014] The depth perception subsystem of the present invention includes any combination of the following: a first autofocus conversion unit for receiving autofocus data from a first camera and converting the autofocus data into a first autofocus depth map; a second autofocus conversion unit for receiving autofocus data from a second camera and converting the autofocus data into a second autofocus depth map; a first parallax conversion unit for receiving image data from the first camera and converting the image data into a first parallax depth map; a second parallax conversion unit for receiving image data from the second camera and converting the image data into a second parallax depth map; and a second parallax conversion unit for receiving image data from the first camera and the second camera and converting the image data into a second parallax depth map accordingly. (disparity) Generate a depth map, (disparity) Conversion unit.

[0015] The first and second disparity units may be configured to acquire first and second consecutive images in the image data and then measure the amount by which each portion of the first image moves relative to the second image. Furthermore, each of the first and second disparity conversion units may include: a separation unit that receives image data from a respective camera, divides the image data into a plurality of segments, and then generates shifted image data in response to the plurality of segments; a movement determination unit that receives position data from the respective camera, and then generates camera movement data indicative of the camera's position in response thereto; and a distance conversion unit that receives the image data and the camera movement data and then converts the image data and the camera movement data into a respective disparity depth map. The distance conversion unit generates the respective disparity depth map using a region convolutional neural network (R-CNN) technique. Also, (disparity) The conversion unit converts between an image of the image data received from the first camera and an image of the image data received from the second camera. (disparity) Analyze the following.

[0016] The depth perception subsystem includes a first autofocus depth map, a second autofocus depth map, a first parallax depth map, and a second parallax depth map.(disparity) and a depth map conversion unit for receiving the depth maps, forming a received depth map, and then converting the received depth maps into a single composite depth map. The depth map conversion unit also includes a depth map generation unit for, after receiving the received depth maps, converting the received depth maps into the single composite depth map, and a confidence value generation unit for generating from the received depth maps a confidence value associated with each of the distance values ​​associated with each point in the single composite depth map. The confidence value indicates a confidence level of the distance values ​​associated with the single combined depth map.

[0017] The present invention also relates to a method for generating depth maps from image data in a surgical robotic system, the method comprising: providing a robotic subsystem having a camera assembly having first and second cameras for generating image data; generating multiple depth maps based on the image data from the first and second cameras; converting the multiple depth maps into a single composite depth map having associated distance data; and controlling the camera assembly based on the distance data in the single composite depth map.

[0018] The method also includes one or more, or any combination thereof, of the following: converting autofocus data from the first camera into a first autofocus depth map; converting autofocus data from the second camera into a second autofocus depth map; converting image data from the first camera into a first disparity depth map; converting image data from the second camera into a second disparity depth map; and converting the image data from the first camera and the image data from the second camera into a second disparity depth map. (disparity) Generating a depth map.

[0019] The method also includes generating a first autofocus depth map, a second autofocus depth map, a first disparity depth map, a second disparity depth map, and (disparity)The method includes receiving a depth map, forming a received depth map, and then converting the received depth map into a single composite depth map. The method further includes generating, from the received depth map, a confidence value associated with each of the distance values ​​associated with each point in the single composite depth map, the confidence value indicating a confidence level of the distance values ​​associated with the single combined depth map. [Brief explanation of the drawings]

[0020] These and other features and advantages of the present invention will be more fully understood by reference to the following detailed description taken in conjunction with the accompanying drawings, in which like reference numerals refer to like elements throughout the different drawings, illustrating the principles of the invention and showing relative dimensions, although not to scale. (Figure 1) FIG. 1 is a schematic block diagram of a surgical robotic system suitable for use with the present invention. (Figure 2) FIG. 2 is a schematic diagram of a depth perception subsystem in accordance with the teachings of the present invention. (Figure 3) FIG. 2 is a schematic block diagram of a disparity conversion unit of the depth perception subsystem of the present invention; (Figure 4) FIG. 2 is a schematic block diagram of a depth map conversion unit of a depth perception subsystem in accordance with the teachings of the present invention. (Figure 5) 3 is an exemplary schematic diagram of a processing technique employed by a depth map conversion unit of a depth perception subsystem in accordance with the teachings of the present invention. (Figure 6) 1 is an illustration of an exemplary depth map in accordance with the teachings of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0021] In the following description, numerous specific details are set forth with respect to the systems and methods of the present invention, as well as the environments in which the systems and methods may operate, to provide a thorough understanding of the disclosed subject matter. However, it will be apparent to those skilled in the art that the disclosed subject matter may be practiced without such specific details, and that certain features well known in the art have not been described in detail to avoid complicating the disclosed subject matter. In addition, it is understood that any examples provided below are merely illustrative and should not be construed as limiting, and that it is contemplated by the inventors that other systems, devices, and / or methods can be used to implement the teachings of the present invention and are considered to be within the scope of the present invention.

[0022] Although the systems and methods of the present invention can be designed for use with one or more surgical robotic systems used as part of a virtual reality surgery, the systems of the present invention can be used with any type of surgical system, including, for example, robotic surgical systems, straight stick surgical systems, and laparoscopic systems. In addition, the systems of the present invention can be used in other non-surgical systems where a user needs access to a large amount of information while controlling a device or apparatus.

[0023] The systems and methods disclosed herein can be incorporated into and utilized, for example, with the robotic surgical devices and associated systems disclosed in U.S. Patent No. 10,285,765 and PCT Patent Application No. PCT / US20 / 39203, and / or the camera system disclosed in U.S. Patent Application Publication No. 2019 / 0076199. The entire contents and teachings of the aforementioned applications and publications are incorporated herein by reference. A surgical robotic system forming part of the present invention can comprise a surgical system including a user workstation, a robotic support system (RSS), a motor unit, and an embedded surgical robot including one or more robotic arms and one or more camera assemblies. The embedded robotic arms and camera assemblies can form part of a single support axis robotic system or can form part of a split-arm architecture robotic system.

[0024] The robotic arm can have movement-related portions or regions that correspond to the user's shoulder, elbow, wrist, and fingers. For example, a robotic elbow can follow the position and orientation of a human elbow, and a robotic wrist can follow the position and orientation of a human wrist. The robotic arm can also have an associated end region, terminating in an end effector that follows the movement of one or more of the user's fingers, such as the index finger when the user pinches the index finger and thumb. The robotic shoulder is fixed in place while the robotic arm follows the movement of the user's arm. In one embodiment, the position and orientation of the user's torso are subtracted from the position and orientation of the user's arm. This subtraction allows the user to move their torso without moving the robotic arm. The robotic arm can be configured to reach all areas of the surgical site and can operate in different configurations and tight spaces.

[0025] The present invention is directed to a surgical robotic system employing a depth perception subsystem that generates multiple depth maps and then combines or merges the depth maps into a single composite depth map. The depth perception subsystem also generates a set of confidence values ​​associated with distance data in the single composite depth map. The confidence values ​​are indicative of the confidence or likelihood that distance data associated with selected points or portions of the depth map is correct or accurate. The depth perception subsystem generates depth maps associated with cameras of a camera assembly. Specifically, the depth perception subsystem generates depth maps associated with the camera's autofocus mechanism, disparity data associated with each camera, and (parallax data) , and the relationship between the image data from each camera (disparity) The depth perception subsystem then processes all of the depth maps to generate a single combined depth map. The depth map data and confidence values ​​are used by the system to move one or more components of the robotic subsystem.

[0026] FIG. 1 is a schematic block diagram illustration of a surgical robotic system 10 in accordance with the teachings of the present invention. System 10 includes a display device or unit 12, a virtual reality (VR) computing unit 14, a sensing and tracking unit 16, a computing unit 18, and a robotic subsystem 20. Display unit 12 can be any selected type of display for displaying information, images, or videos generated by VR computing unit 14, computing unit 18, and / or robotic subsystem 20. Display unit 12 can include, for example, a head-mounted display (HMD), a screen or display, a three-dimensional (3D) screen, etc. The display unit can also include an optional sensor and tracking unit 16A, such as those found in commercially available head-mounted displays. Sensing and tracking units 16 and 16A can include one or more sensors or detectors coupled to a user of the system, such as a nurse or surgeon. Sensors can be coupled to the user's arms, and if a head-mounted display is not used, additional sensors can also be coupled to the user's head and / or neck region. The sensors in this configuration are represented by sensing and tracking unit 16. If the user uses a head-mounted display, eye, head, and / or neck sensors and associated tracking technology can be incorporated into or used within the device and thus form part of the optional sensing and tracking unit 16A. The sensors of the sensing and tracking unit 16 coupled to the surgeon's arm can be coupled to selected regions of the arm, such as the shoulder region, elbow region, wrist or hand region, and, if desired, the fingers. The sensors generate position data indicative of the position of selected parts of the user. The sensing and tracking unit 16 and / or 16A can be utilized to control the camera assembly 44 and robotic arm 42 of the robotic subsystem 20. The position data 34 generated by the sensors of the sensing and tracking unit 16 is conveyed to the computing unit 18 for processing by the processor 22.From the position data, the computing unit 18 can determine or calculate the position and / or orientation of each part of the surgeon's arm and communicate this data to the robotic subsystem 20. According to an alternative embodiment, the sensing and tracking unit 16 can use sensors coupled to the surgeon's torso or any other body part. Furthermore, the sensing and tracking unit 16 can use, in addition to sensors, an inertial momentum unit (IMU) with, for example, an accelerometer, a gyroscope, a magnetometer, and a motion processor. The addition of a magnetometer is standard practice in the field, and magnetic orientation allows for the reduction of sensor drift around a vertical axis. Alternative embodiments also include sensors placed within surgical materials, such as gloves, surgical scrubs, or a surgical gown. The sensors may be reusable or disposable. Additionally, sensors can be located external to the user, such as in a fixed location in a room, such as an operating room. The external sensors can generate external data 36 that is processed by the computing unit and thus used by the system 10. According to another embodiment, when display unit 12 is a head-mounted device using an associated sensing and tracking unit 16A, the device generates tracking and position data 34A that is received and processed by VR computation unit 14. Additionally, sensing and tracking unit 16 may include hand controllers, if desired.

[0027] In embodiments in which the display is an HMD, the display unit 12 may be, for example, a virtual reality head-mounted display such as the Oculus Rift, Varjo VR-1, or HTC Vive Pro Eye. The HMD may provide the user with a display coupled to or worn on the user's head, lenses that enable a focused view of the display, and a sensor and / or tracking system 16A that provides tracking of the position and orientation of the display. The position and orientation sensor system may include, for example, accelerometers, gyroscopes, magnetometers, motion processors, infrared tracking, eye tracking, computer vision, emitting and sensing alternating magnetic fields, and any other method or combination of tracking at least one of position and orientation. As is known, the HMD may provide image data for the surgeon's right and left eyes from a camera assembly 44. To maintain a virtual reality experience for the surgeon, the sensor system may track the position and orientation of the surgeon's head and then relay the data to the VR computing unit 14 and, if necessary, to the computing unit 18. The computing unit 18 can further adjust the pan and tilt of the camera assembly 44 of the robotic subsystem 20 via the motor unit 40 to follow the movements of the user's head.

[0028] When associated with the display unit 12, sensor or position data generated by the sensors can be communicated to the computing unit 18, either directly or via the VR computing unit 14. Similarly, tracking and position data 34 generated by other sensors in the system, such as from the sensing and tracking unit 16, which may be associated with the user's arms and hands, can be communicated to the computing unit 18. The tracking and position data 34, 34A can be processed by the processor 22 and stored, for example, in the storage unit 24. The tracking and position data 34, 34A can also be used by the control unit 26 to responsively generate control signals for controlling one or more portions of the robotic subsystem 20. The robotic subsystem 20 can include a user workstation, a robotic support system (RSS), a motor unit 40, and an implantable surgical robot including one or more robotic arms 42 and one or more camera assemblies 44. The embedded robotic arm and camera assembly can form part of a single support axis robotic system such as that shown and described in U.S. Pat. No. 10,285,765, or can form part of a split arm architecture robotic system such as that shown and disclosed in PCT Patent Application No. PCT / US20 / 39203, the contents of which are incorporated herein by reference.

[0029] The control signals generated by the control unit 26 can be received by the motor unit 40 of the robot subsystem 20. The motor unit 40 can include a series of servo motors configured to separately drive the robot arm 42 and the camera assembly 44. The robot arm 42 can be controlled to follow the scaled-down movements or motions of the surgeon's arm as sensed by associated sensors. The robot arm 42 can have portions or regions associated with the movements associated with the user's shoulder, elbow, wrist, and fingers. For example, a robotic elbow can follow the position and orientation of a human elbow, and a robotic wrist can follow the position and orientation of a human wrist. The robotic arm 42 can also have associated end regions, such as terminating in an end effector that follows the movement of one or more of the user's fingers, such as the index finger when the user pinches the index finger and thumb together. The robot's shoulder is fixed in place while the robot's arm follows the movement of the user's arm. In one embodiment, the position and orientation of the user's torso is subtracted from the position and orientation of the user's arm. This subtraction allows the user to move their torso without moving the robotic arm.

[0030] The robotic camera assembly 44 is configured to provide image data 48, such as a live video feed of the procedure or surgical site, to the surgeon and to allow the surgeon to operate and control the cameras that make up the camera assembly 44. The camera assembly 44 preferably includes a pair of cameras whose optical axes are axially separated by a selected distance (known as the inter-camera distance) to provide a stereoscopic view of the surgical site. The surgeon can control the movement of the cameras through the movement of a head-mounted display, via sensors coupled to the surgeon's head, or by using a hand controller or sensors that track the user's head or arm movements, thereby enabling the surgeon to obtain a desired view of the surgical site in an intuitive and natural manner. The cameras are movable in multiple directions, including, for example, yaw, pitch, and roll, as is known. The components of the stereoscopic camera can be configured to provide a user experience that feels natural and comfortable. In some embodiments, the inter-axis distance between the cameras can be modified to adjust the depth of the surgical site perceived by the user.

[0031] The camera assembly 44 is actuated by the surgeon's head movement. For example, during surgery, if the surgeon wants to view an object located above the current field of view, the surgeon looks upward, resulting in the stereoscopic camera being rotated upward about the pitch axis from the user's perspective. Image or video data 48 generated by the camera assembly 44 can be displayed on the display unit 12. If the display unit 12 is a head-mounted display, the display may include a built-in tracking and sensor system that captures raw orientation data for the HMD's yaw, pitch, and roll directions, as well as the HMD's position data in Cartesian space (x, y, z). However, alternative tracking systems may be used to provide supplemental position and orientation tracking data for the display instead of or in addition to the HMD's built-in tracking system. Examples of camera assemblies suitable for use in the present invention include the camera assemblies disclosed in U.S. Pat. No. 10,285,765 and U.S. Patent Application Publication No. 2019 / 0076199, the contents of which are incorporated herein by reference.

[0032] Image data 48 generated by the camera assembly 44 is communicated to the virtual reality (VR) computation unit 14 and processed by the VR or image rendering unit 30. The image data 48 may include still photographs or image data as well as video data. The VR rendering unit 30 may include appropriate hardware and software for processing the image data and then rendering the image data for display by the display unit 12, as known in the art. Additionally, the VR rendering unit 30 may combine the image data received from the camera assembly 44 with information related to the position and orientation of the cameras in the camera assembly and information related to the position and orientation of the surgeon's head. This information enables the VR rendering unit 30 to generate and transmit an output video or image rendering signal to the display unit 12. That is, the VR rendering unit 30 renders readings of the position and orientation of the hand controllers and the surgeon's head position for display on the display unit, such as in an HMD worn by the surgeon.

[0033] The VR computation unit 14 may also include a virtual reality (VR) camera unit 38 for generating one or more virtual reality (VR) cameras for use or placement in the VR world displayed within the display unit 12. The VR camera unit may generate one or more virtual cameras in the virtual world, which may be used by the system 10 to render images for the head-mounted display. This ensures that the VR camera always renders the same view as a user wearing the head-mounted display sees in the cubemap. In one embodiment, a single VR camera may be used, while in another embodiment, separate left-eye and right-eye VR cameras may be used to render onto separate left-eye and right-eye cubemaps in the display, providing a stereo view. The VR camera's FOV setting may be self-configuring relative to the FOV exposed by the camera assembly 44. In addition to providing a contextual background for the live camera view or image data, the cubemap may be used to generate dynamic reflections on virtual objects. This effect allows reflective surfaces on virtual objects to pick up reflections from the cubemap, making these objects appear to the user as if they actually reflect their real-world environment.

[0034] The robotic arm 42 may be comprised of multiple mechanically linked actuating sections or portions that can be constructed and combined for rotational and / or hinged motion to emulate different portions of a human arm, such as, for example, the shoulder region, elbow region, and wrist region of the arm. The actuator sections of the robotic arm may be configured, for example, to provide cable-driven rotational motion, but within reasonable rotational limits. The actuator sections may be configured to provide maximum torque and maximum speed with minimal size.

[0035] The present invention relates to generating and providing depth perception-related data (e.g., distance data and / or depth map data) to a robotic subsystem so that the data can be used to assist a surgeon in controlling the movement of one or more components, such as a robotic arm or camera, of the subsystem. The depth perception-related data is important because it allows a surgeon to determine the amount of movement the robot can safely perform at the surgical site before and during a surgical procedure. Additionally, the data can be used for automated operations without a surgeon. The present invention also relates to the generation and provision of depth perception-related data (e.g., lens focus data, image data, image (disparity) Software and hardware (e.g., processors, memory, storage devices, etc.) may be used to calculate or determine a three-dimensional (3D) distance map, or depth map, from a number of different data and image sources, including image data and image disparity-related data. According to other embodiments, other types of depth cues may be used as input to depth perception subsystem 50 of the present invention.

[0036] The present invention can use a variety of computing elements and sensors to determine or extract depth and related distance information. There are various hardware and software that can be used with the system of the present invention to extract depth perception information or distance data. For example, hardware sensors, such as structured light or time-of-flight sensors, measure changes in physical parameters to estimate distance. Software sensors are used to infer distance by analyzing specific features within one or more images in time and space. The system can use a variety of computing elements and sensors to generate or convert input images or other types of data into depth-related data. (disparity), epipolar geometry, Structure from Motion (SfM), and other techniques can be used. While our system can extract depth-related information from a single cue or source, our system can also consider additional inputs or data sources when constructing a final combined three-dimensional (3D) depth map. Our final combined depth map essentially combines multiple lower-quality depth maps into a single final combined depth map of any selected scene that is more robust to noise, occlusion, and ambiguity, such as the interior of a human body.

[0037] The computing unit 18 of the surgical robotic system 10 may include a depth perception subsystem 50, as shown in FIG. 2 , for example. The depth perception subsystem 50 is configured to interact with one or more components of the robotic subsystem 20, such as the camera assembly 44 and the robot arm 42. The camera assembly 44 may include, for example, a pair of stereoscopic cameras, including a left camera 44A and a right camera 44B. The left camera 44A may include, for example, a lens and optics system 54A including one or more lenses and associated optical elements for receiving optical or image information, among other components. The camera 44A may also include an image sensor 58A for capturing optical or image data and an autofocus mechanism 62A for providing the camera with autofocus functionality. The autofocus mechanism 62A interacts with and automatically changes or adjusts the optics, such as the lenses in the lens and optics system 54A, to focus an image on the image sensor 58A. The image sensor surface may typically correspond to the focal plane. Similarly, the right camera 44B may include a lens and optics system 54B including one or more lenses and associated optical elements. The camera 44B may also include an image sensor 58B for capturing optical or image data and an autofocus mechanism 62B for providing autofocus functionality for the camera.

[0038] Cameras 44A, 44B with autofocus capabilities can provide information that can be converted into an initial, rough depth map of the area directly monitored by the camera. Thus, each camera 44A, 44B constantly monitors the input stream of image data from its image sensor 58A, 58B to maintain the observed environment or object in focus. For each image portion in the image data stream, a subset of the image's pixels can be maintained in focus by the corresponding image sensor. As is known, autofocus mechanisms can generate effort signals, which are adjustments required by the autofocus hardware and software. The effort signals can be converted into control signals that can be used to mechanically or electrically alter the geometry of the optical system. Furthermore, for a given image, any subset of in-focus pixels can be associated with a control signal. The control signals can be converted into approximate depths by autofocus conversion units 70A, 70B, thereby generating a depth map of in-focus pixels.

[0039] The illustrated depth perception subsystem 50 may further include an autofocus conversion unit 70A for receiving autofocus data 64A generated by the autofocus mechanism 62A. The autofocus conversion unit 70A is responsible for converting the autofocus data into distance data, which is displayed as or forms part of a depth map 72A. As used herein, the term “depth map” or “distance map” is intended to include an image, image channel, or map containing information regarding or about the distance between one or more objects or image surfaces from a selected viewpoint or viewpoint in the overall scene. The depth map may be created from source image or image data and may be presented in any selected color, such as grayscale, and may include one or more color variations or hues, each variation or hue corresponding to various or different distances of the image or object from the viewpoint in the overall scene. Similarly, the autofocus mechanism 62B generates autofocus data received by an autofocus conversion unit 70B. In response, the autofocus conversion unit 70B converts the autofocus data into distance data, which is displayed as or forms part of a separate depth map 72B. In some embodiments, the depth of focus can be intentionally varied over time to generate a more detailed depth map of different pixels.

[0040] The depth perception subsystem 50 may further include parallax conversion units 80A and 80B for converting image data into distance data. Specifically, the left and right cameras 44A and 44B generate camera data 74A and 74B, respectively, which may include, for example, image data and camera position data, and are transmitted to and received by the parallax conversion units 80A and 80B, respectively, which then convert the data into distance data that may form part of the separate depth maps 76A and 76B. As is known, the parallax effect typically exists in both natural and artificial optical systems and can cause objects farther from the image sensor to appear to move more slowly than objects closer to the image sensor when the image sensor is moved. In some embodiments, measuring the parallax effect is achieved by measuring, for two images taken at successive intervals, how much each part of the image has moved relative to its corresponding part in the previous interval. The more a part of the image moves between intervals, the closer it is to the camera.

[0041] Details of the disparity conversion units 80A and 80B are shown, for example, in FIG. 3 . Because the disparity conversion units 80A and 80B are identical, only the disparity conversion unit 80A will be described below for simplicity and clarity. The camera data 74A generated by the camera 44A may include image data 74C and camera position data 74D. The camera position data 74D corresponds to the vertical and horizontal position of the camera based on the position measured or commanded by on-board sensors and electronics. The image data 74C is introduced and received by the separation unit 130. The separation unit 130 divides the image data 74C into multiple patches or segments and then generates shifted image data 132 in response thereto by comparing sets of typically consecutive images within the image data. The camera position data 74D generated by the camera 44A is received by a movement determination unit 134, which determines the position of the camera and then generates camera movement data 136. The camera movement data 136 relates to the amount of movement or rotation of the camera, measured by on-board sensors, estimated based on kinematics, or simply estimated based on commands. The shifted image data 132 and the movement data 136 are then introduced to a distance conversion unit that converts the two types of input data 132, 136 into distance data, which may be represented in the form of a separate depth map 76A. An example of how to determine distance data from image data is described in: Active estimation of distance in a robotic system that replicates human eye movement, Santini et al., Robotics and Autonomous Systems, August 2006, the contents of which are incorporated herein by reference.

[0042] The distance transformation unit 140 can use known processing techniques, such as, for example, a multi-layer domain convolutional neural network (R-CNN) technique. For example, according to one embodiment, training data for the network is generated using image pairs selected from the image data. Furthermore, the separation unit 130 can segment an image (e.g., an image captured at time t) into smaller image segments. For each image segment, a position, which may be on an image from the same image sensor but at a different time (e.g., time t+1), is calculated using a normalized cross-correlation technique. Because the depth perception subsystem 50 can easily determine the motion actuated during the preceding time interval and the position difference of the image segment before and after the motion, the distance of 3D points included in the image segment can be calculated or determined via known optical considerations and known analytical formulations and techniques.

[0043] Referring again to Figure 2, depth perception subsystem 50 converts image data 78A received from camera 44A and image data 78B received from camera 44B into distance data. (disparity) It further comprises a transformation unit 90 , and the distance data may form part of a depth map 92 . (disparity) The conversion unit 90 converts the image data received from the cameras 44A and 44B into a difference between the images. (disparity) Specifically, (disparity) The transformation unit 90 analyzes the same image section of each input image and determines the differences between them. The differences between the images recorded by the optical system and image sensors of the cameras 44A, 44B by observing the scene from different viewpoints are (disparity) This can be used in conjunction with the known geometry and placement of the camera's optical system to convert the information into distance information. According to one embodiment, the distance between the images from the left and right cameras 44A, 44B can be calculated. (disparity) is computed using a well-layered regional convolutional neural network (R-CNN) that simultaneously considers and processes all pixels from the image. (disparity)The transformation unit 90 may be trained, for example, using images from each camera 44A, 44B selected from a real-time image feed. (disparity) can be calculated using the formula: (disparity) The value (d) can be converted to a depth value (Z) by the following formula: Z= R * f / d where f is the focal length of the cameras and T is the baseline distance between the cameras.

[0044] (disparity) The depth map or distance data 92 generated by the transformation unit 90 and corresponding to the input image data under consideration can be generated and refined by using epipolar geometry. For example, the likely location of point A on the left image received by the left camera 44A can be estimated on the right image received from the right camera 44B. For example, a normalized cross-correlation between the left and right portions of the image around a selected point, such as point A, is performed to obtain a more accurate location estimate. Then, (disparity) and depth information is derived using well-known analytical formulas. (disparity) The depth map produced by transformation unit 90 can be further improved by a manual refinement process that removes artifacts and outliers that are not easily detected by the automated functions of the depth perception subsystem.

[0045] The inventors have used autofocus conversion units 70A and 70B, parallax conversion units 80A and 80B, and (disparity)It has been recognized that the depth maps 72A, 72B, 76A, 76B, and 92 generated by the conversion unit 90 may be inherently unreliable when used separately to determine distance. Individual depth maps may not contain all of the image data and associated position data necessary to properly and appropriately control the robotic subsystem 20. To address this unreliability, the depth perception subsystem 50 may employ a depth map conversion unit 100. The depth map conversion unit 100 is configured to receive all of the depth maps and associated distance data generated by the depth perception subsystem 50 and combine or merge the depth maps into a single composite depth map 122. The depth map conversion unit 100 may employ one or more different types of processing techniques, including, for example, a regional convolutional neural network (R-CNN)-based encoder-decoder architecture.

[0046] The depth map conversion unit 100 is shown in more detail in FIGS. 3 and 4. As shown in FIG. 3, the depth map conversion unit 100 may include a depth map generation unit 120 that combines input depth maps and generates a single combined output depth map 122 from the depth maps. The depth map conversion unit 100 also includes a confidence value generation unit 110 for generating one or more confidence values ​​112 associated with each distance or point on the depth map or associated with a portion or segment of the depth map. As used herein, the term "confidence value" or "likelihood value" is intended to include any value that quantifies the accuracy or veracity of a given parameter and provides a way to communicate its certainty or reliability. In this embodiment, the value is associated with the confidence of a distance measurement or value, such as a distance value, associated with a depth map. The value may be expressed in any selected range, preferably between 0 and 1, with zero representing the lowest or minimum confidence level or value and one representing the highest or maximum confidence level or value. Furthermore, the confidence value may be expressed as a distance or distance range for a given depth. The confidence interval can be determined by statistical analysis of the data, which can take into account the spread of depth values ​​from depth maps from various depth cues within a given region of the composite depth map, or their variation over time.

[0047] See FIG. 5. The depth map conversion unit 100 can import multiple input data streams 116, such as depth maps, and then use a regional convolutional neural network (CNN)-based encoder-decoder architecture 114 that processes the depth map data using a series of CNN filters or stages. The CNN filters can be arranged at the input as an encoder stage or a series of CNN filters 118A. In this case, the data on the depth map is downsampled to reduce the distance and image data from the best or highest quality pixel or image segment. The data is then upsampled in a decoder stage of a CNN filter 118B that uses a series of aligned CNN filters, and the data is combined with other data from the input side to form or create a combined image, such as a single combined depth map 122. The encoder-decoder CNN architecture 114 removes noise from the input data, thereby helping to generate more accurate output data, such as a single composite depth map 122. The encoder-decoder CNN architecture 114 can also have a parallel upsampling or decoding stage of a CNN filter 118C, which also upsamples the input data and removes any accompanying noise. The data can then be processed through a Softmax function 124 to generate or create a confidence value 112. As is known, the Softmax function is a multidimensional generalization of the logistic function and can be used in multinomial logistic regression as the final activation function of a neural network to normalize the network's output into a probability distribution over predicted output classes. In this embodiment, the output of the Softmax function can be used to calculate the probability that the depth value for a particular point in the depth map is accurate. The encoder-decoder CNN architecture 114 merges the input depth maps with a probabilistic approach that minimizes or reduces noise in the single combined depth map 122 and estimates the associated true distance value better than the data contained in each individual input depth map.In some other embodiments, the estimates of 112 and 118A may also take into account any confidence data generated by 72A, 72B, 76A, 76B, and 92 during processing. The encoder-decoder CNN architecture 114 may be trained using the expected outputs of samples in one or more training sets of input depth or distance cues or maps, and the resulting depth and likelihood maps may be analytically calculated from the input depth maps. Other methods known in the art for combining depth maps, such as Kalman filters, particle filters, etc., may also be utilized.

[0048] As mentioned above, noise in the input depth map can arise in a variety of ways. For example, with regard to autofocus mechanisms, cameras tend to follow laws of physics and optics, such as focal length and lens mechanisms, to inadvertently amplify measurement errors and thus invalidate estimates. The systems and methods of the present invention use depth maps 72A, 72B in conjunction with other depth maps to remove spurious or outlier readings. Furthermore, (disparity) Noise in the computation of is generally related to the diversity of the observed environment. A large amount of unique features in an image strongly reduces the probability of ambiguity (e.g., this type of cue or noise source in a depth map). When considered independently from other input data sources, (disparity) Depth sources using cues or depth maps do not have a means for resolving ambiguity. The system of the present invention can resolve this ambiguity by considering which other sources estimate the same region of the image and discarding erroneous or unlikely possibilities. Disparity cues or depth maps are computationally complex because they depend on several noisy parameters, such as accurate measurements of actuated camera motion and knowledge of the geometric relationship between the left and right imaging sensors 58A, 58B. The fusion approach of the present invention allows the system 10 to reduce the effects of noise in these cues or depth maps and produce a single, combined depth map that is less noisy than if the input sources or depth maps were considered individually.

[0049] FIG. 6 is an example of a composite depth map 122 generated by the depth map conversion unit 100 in accordance with the teachings of the present invention. The illustrated depth map 122 is formed by combining all of the input depth maps. The depth map 122 includes a scene 144, which includes a series of pixels forming an image within the scene. The image may have different hues indicating different distances or depths from the viewpoint. The current scene 144 is represented in grayscale, although other colors may be used. Lighter hues 146 may represent pixels or segments of the image within the scene 144 that are closer to the viewpoint, while darker hues 148 may represent pixels or segments of the image that are farther from the viewpoint. Thus, pixels in the depth map are associated with distance values, and the depth map generator may also generate confidence values ​​112 associated with each distance value. Thus, each point or pixel in the depth map may have a depth value and an associated confidence value. The confidence values ​​may be stored separately within the system.

[0050] Referring again to FIG. 2 , the depth map 122 and confidence value 112 generated by the depth map conversion unit 100 are introduced to the control unit 26. The distance and confidence values ​​are used by the control unit 26 to control the movement of the cameras 44 a, 44 b and / or the robot arm 42 of the robotic subsystem 20. The confidence values ​​provide a reasonable degree of confidence for the distances shown in the depth map 122 that, for example, when a surgeon moves the robotic arm, the arm will not touch a surface before or after the distance measurement. For delicate surgical procedures, having confidence in the distance values ​​in the depth map is important because the surgeon needs to know whether the instructions sent to the robotic arm and camera are accurate. The depth map is important for automatically alerting or preventing the surgeon from accidentally touching the surgical environment or anatomical structures, thus enabling the system to automatically traverse the surgical environment and interact with anatomical components without intervention by the surgeon. Furthermore, the depth perception subsystem can enable “guide rails” to be virtually placed within the surgical environment to aid in the control and / or direct movement of the robotic subsystem. Similarly, depth maps allow surgical environments to be augmented or virtual objects to be more accurately placed within the environment (e.g., by overlaying a preoperative scan of a patient's vitals on the surgical site), as disclosed or illustrated by way of example in International Patent Application No. PCT / US2020 / 059137, the contents of which are incorporated herein by reference. Depth maps can also be used with computer vision and artificial intelligence to help identify anatomical and abnormal structures. Furthermore, depth maps can be used in combination with advanced sensory information (e.g., multi-wavelength images to detail the vasculature) or patient images (e.g., MRIs, CAT scans, etc.) to create rich three-dimensional maps that surgeons can use to plan future procedures.

[0051] It will thus be seen that the present invention efficiently attains the objects set forth above, among those that will become apparent from the foregoing description. Because certain changes can be made in the above-described construction without departing from the scope of the invention, it is intended that all matter contained in the above description or shown in the accompanying drawings be interpreted as illustrative and not restrictive.

[0052] It will be understood that the claims are intended to cover all of the generic and specific features of the invention that have been individually described, as well as any statements of the scope of the invention that may be omitted as a matter of language.

[0053] Having described the invention, what is claimed as new and desired to be protected by Letters Patent is set forth in the appended claims.

Claims

1. a robotic subsystem having a camera assembly having first and second cameras for generating image data; A computing unit; Equipped with The computing unit a processor for processing the image data; a control unit for controlling the robotic subsystem; a depth perception subsystem for receiving the image data generated by the first and second cameras, generating at least two different types of depth maps based on the image data, selected from an autofocus depth map, a parallax depth map, or a disparity depth map, and converting the at least two different types of depth maps into a single composite depth map having associated distance data; having Surgical robotic system.

2. the robotic subsystem further comprises: a plurality of robot arms; and a motor unit for controlling movement of the plurality of robot arms and the camera assembly; The surgical robotic system of claim 1 .

3. The surgical robotic system of claim 2 , wherein the control unit uses the distance data associated with the single composite depth map to control one of the camera assembly or the robotic arm.

4. The depth perception subsystem further comprises: a depth map conversion unit for receiving the two different types of depth maps and converting the depth maps into the single composite depth map. The surgical robotic system of claim 1 .

5. The surgical robot system of claim 4 , wherein the depth map conversion unit uses a regional convolutional neural network (R-CNN) technique to generate the single composite depth map.

6. The first and second cameras each include: an image sensor that receives optical data and responsively generates said image data; a lens and optical system having one or more lens elements optically coupled with the image sensor for focusing the optical data onto the image sensor; an autofocus mechanism associated with the lens and optical system that automatically adjusts the one or more lens elements and generates autofocus data; Equipped with The surgical robotic system of claim 1 .

7. the depth perception subsystem: a first autofocus transformation unit for receiving the autofocus data from the first camera and transforming the autofocus data into a first camera autofocus depth map; a second autofocus transformation unit for receiving the autofocus data from the second camera and transforming the autofocus data into a second camera autofocus depth map; Equipped with The surgical robot system of claim 6.

8. The depth perception subsystem further comprises: a first disparity transformation unit for receiving image data from the first camera and transforming the image data into a first camera disparity depth map; a second disparity transformation unit for receiving image data from the second camera and transforming the image data into a second camera disparity depth map; Equipped with The surgical robotic system of claim 7.

9. The depth perception subsystem further comprises: a disparity transformation unit for receiving image data from the first camera and image data from the second camera and for generating the disparity depth map in response thereto; The surgical robotic system of claim 8.

10. the depth perception subsystem: a first autofocus transformation unit for receiving the autofocus data from the first camera and transforming the autofocus data into a first camera autofocus depth map; a second autofocus transformation unit for receiving the autofocus data from the second camera and transforming the autofocus data into a second camera autofocus depth map; a first disparity transformation unit for receiving image data from the first camera and transforming the image data into a first camera disparity depth map; a second disparity transformation unit for receiving image data from the second camera and transforming the image data into a second camera disparity depth map; a disparity transformation unit for receiving image data from the first camera and for receiving image data from the second camera and for generating a disparity depth map in response thereto; comprising one or more of: The surgical robot system of claim 6.

11. 11. The surgical robot system of claim 10, wherein each of the first and second parallax units is configured to acquire first and second successive images in the image data and measure the amount that each portion of the first image moves relative to the second image.

12. Each of the first and second cameras generates position data, and each of the first and second parallax conversion units: a separation unit for receiving the image data from each of the cameras, dividing the image data into a plurality of segments, and generating shifted image data in response to the plurality of segments; a movement determination unit for receiving the position data from each of the cameras and responsively generating camera movement data indicative of the positions of the cameras; a distance transformation unit that receives the image data and the camera movement data and transforms the image data and the camera movement data into each of the disparity depth maps; Equipped with The surgical robotic system of claim 11.

13. The surgical robotic system of claim 12 , wherein the distance transformation unit uses a region-specific convolutional neural network (R-CNN) technique to generate each of the disparity depth maps.

14. The surgical robot system of claim 10 , wherein the image disparity conversion unit analyzes image disparity between an image in the image data received from the first camera and an image in the image data received from the second camera.

15. 15. The surgical robotic system of claim 14, wherein the image disparity between the images from the first and second cameras is determined using a Layered Local Convolutional Neural Network (R-CNN) technique.

16. The depth perception subsystem further comprises: a depth map conversion unit configured to receive the first camera autofocus depth map, the second camera autofocus depth map, the first camera disparity depth map, the second camera disparity depth map, and the disparity depth map, form a received depth map, and convert the received depth map into the single composite depth map; The surgical robotic system of claim 10.

17. 17. The surgical robotic system of claim 16, wherein the depth map conversion unit uses a regional convolutional neural network (R-CNN) based encoder-decoder architecture to generate the single composite depth map.

18. Each point in each of the received depth maps has a distance value associated with it, and the depth map conversion unit: a depth map generation unit for receiving the received depth maps and converting the received depth maps into the single composite depth map; a confidence value generation unit for generating from the received depth map a confidence value associated with each of the distance values ​​associated with each point of the single composite depth map; Equipped with the confidence value indicates a degree of confidence in the distance value associated with the single composite depth map.

18. The surgical robotic system of claim 17.

19. 1. A method of generating a depth map from image data in a surgical robotic system, comprising: providing a robotic subsystem having a camera assembly having first and second cameras for generating image data; generating at least two different types of depth maps based on image data from the first and second cameras, the depth maps being selected from an autofocus depth map, a parallax depth map, or a disparity depth map; converting the at least two different types of depth maps into a single composite depth map having distance data associated therewith; controlling the camera assembly based on the distance data in the single composite depth map; A method having the following.

20. the robot subsystem further comprises a plurality of robotic arms and a motor unit for controlling movement of the plurality of robotic arms and the camera assembly; The method further comprises controlling the robot arm based on the distance data in the single composite depth map.

20. The method of claim 19.

21. Each of the first and second cameras an image sensor for receiving optical data and responsively generating said image data; a lens and optical system having one or more lens elements optically coupled with the image sensor for focusing the optical data onto the image sensor; an autofocus mechanism associated with the lens and optical system for automatically adjusting the one or more lens elements and generating autofocus data; Equipped with 20. The method of claim 19.

22. The method further comprises: converting the autofocus data from the first camera into a first camera autofocus depth map; The autofocus data from the second camera is converted into a second camera autofocus depth map. Step 2:

22. The method of claim 21, comprising:

23. The method further comprises: converting the image data from the first camera into a first camera disparity depth map; converting the image data from the second camera into a second camera disparity depth map; 23. The method of claim 22, comprising:

24. 24. The method of claim 23, further comprising generating a disparity depth map from the image data from the first camera and the image data from the second camera.

25. The method further comprises: converting the autofocus data from the first camera into a first camera autofocus depth map; converting the autofocus data from the second camera into a second camera autofocus depth map; converting the image data from the first camera into a first camera disparity depth map; converting the image data from the second camera into a second camera disparity depth map; generating a disparity depth map from the image data from the first camera and the image data from the second camera; 22. The method of claim 21, comprising one or more of:

26. Transforming the image data from the first camera into a first camera disparity depth map includes: acquiring first and second successive images within the image data; measuring the amount that each portion of the first image moves relative to the second image; 26. The method of claim 25, comprising:

27. Transforming the image data from the second camera into a second camera disparity depth map includes: acquiring first and second successive images within the image data; measuring the amount that each portion of the first image moves relative to the second image; 27. The method of claim 26, comprising:

28. The first camera generates position data, and the method further comprises: dividing the image data from the first camera into a plurality of segments and generating shifted image data responsive to the plurality of segments; generating from the first camera movement data indicative of the position of the camera in response to the position data; converting the image data and the camera movement data into the first disparity depth map; 26. The method of claim 25, comprising:

29. The second camera generates position data, and the method further comprises: dividing the image data from the second camera into a plurality of segments and generating shifted image data responsive to the plurality of segments; generating from the second camera movement data indicative of the position of the camera in response to the position data; converting the image data and the camera movement data into the second disparity depth map; 29. The method of claim 28, comprising:

30. 26. The method of claim 25, wherein generating the disparity depth map from the image data from the first camera and the image data from the second camera further comprises analyzing the disparity between images in the image data received from the first camera and images in the image data received from the second camera.

31. The method further comprises: receiving the first camera autofocus depth map, the second camera autofocus depth map, the first camera disparity depth map, the second camera disparity depth map, and the disparity depth map to form a received depth map; converting the received depth maps into the single composite depth map; 26. The method of claim 25, comprising:

32. Each point in each of the receive depth maps has a distance value associated therewith, the method further comprising: generating from the received depth maps a confidence value associated with each of the distance values ​​associated with each point of the single composite depth map; the confidence value indicates a degree of confidence in the distance value associated with the single composite depth map.

32. The method of claim 31.

33. 11. The surgical robot system of claim 10, wherein the two different types of depth maps include at least two of the first camera autofocus depth map, the second camera autofocus depth map, the first camera parallax depth map, the second camera parallax depth map, and a disparity depth map.

34. 26. The method of claim 25, wherein the two different types of depth maps include at least two of the first camera autofocus depth map, the second camera autofocus depth map, the first camera disparity depth map, the second camera disparity depth map, and a disparity depth map.

Citation Information

Patent Citations

  • Data processor, imaging device, and data processing method

    JP2016085637A

  • Depth prediction from image data using statistical models

    JP2019526878A

  • Control device and medical image pickup system

    WO2016185952A1

  • Medical image processing device, system, method, and program

    WO2017138209A1

  • Medical arm system, control device, and control method

    WO2018159328A1